eb051ea69636a5696523d47d09fe5360bc462762
188 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
eb051ea696 |
feat(#1123,#1124): enforce duplicate-producer invariant + fail-loud loadCentralConfigKeys in gen-capability-registry (#1131)
Closes #1123 Closes #1124 Refs #857 |
||
|
|
7f1d49935c |
ci(#1104): keep next package.json in sync with the last published release (#1109)
* ci(#1104): sync next package.json version to the last published release next rested on a -dev stream per ADR-660 (1.3.1-dev.0) — a never-published placeholder that leaked to source/dev installs. Make every release type write its exact published version back to next: - finalize/hotfix (push main): auto-backmerge sets next's version to main's released version, folded into the existing back-merge PR (+ pinned setup-node). - rc (no main push): the rc job opens + admin-merges a sync PR after publish. Shared, fail-closed scripts/sync-next-version.cjs stamps package.json + the runtime manifests via the npm version hook and refuses any non-release version. Amends ADR-660 (supersedes the -dev stream decision). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * ci(#1104): harden next-version sync against post-publish failure modes Review hardening (Codex + code-review gates) on the #1104 sync helper and its workflow callers: - release.yml rc Sync step: continue-on-error so a post-publish sync hiccup cannot fail an already-published release (npm immutability would block re-run). - auto-backmerge.yml inline sync: set -euo pipefail + validate VERSION before any shell use (closes a ${VERSION}-in-commit-message injection vector); git add -u instead of -A. - sync-next-version.cjs: reuse an existing open PR instead of failing gh pr create on rc re-runs; regex-parse the PR number and fail loud; discriminate the git diff --cached --quiet exit code (only status 1 == has-diff, else rethrow); git add -u to avoid sweeping runner artifacts into next; tolerate already-merged on admin merge. - tests: +2 (existing-PR reuse, non-diff rethrow); 14/14 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3e836fef0d |
feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (
|
||
|
|
e4f0910d62 |
test(#1074): agent-size-budget per-file baseline + line→byte rebase (PR 3/3) (#1097)
Extends the #1074 scheme to tests/agent-size-budget.test.cjs, which used the same assertTightCeiling tier ratchet but was still line-based (never rebased in #717). Completes the migration — the last part of the #1074 epic. - Rebase agent sizing from lines to LF-normalized bytes (#717/#683). - Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests); add a per-agent baseline (tests/agent-size-baseline.json) as the primary anti-creep, and byte hard caps (XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB), each above its tier high-water with real headroom. No separate new-file cap: a net-new agent is DEFAULT-tier, already bounded by the DEFAULT cap. - Keep the agent-classification tests verbatim. - scripts/workflow-size.cjs: add generic measureMdFiles(dir, predicate) (workflows + agents share one byte-measurement path); measureWorkflows now delegates to it. - scripts/update-size-baseline.cjs: one 'npm run size:baseline' now regenerates BOTH the workflow and agent baselines (gsd-* filter for agents). Rebased onto next after PR 2/3 (#1096) merged: replicate the scripts/lib/workflow-size.cjs -> scripts/workflow-size.cjs move (PR 1/3) across the generator and the agent test's require; regenerate the agent baseline against current agents (a uniform +170 B preamble drift on all 33 since authoring). Addresses the #1097 review (trek-e): - BLOCKER (acceptance criterion 5): document the agent contract in CONTEXT.md. Adds RULESET.AGENT_SIZE_BUDGET (caps 57344/49152/24576, per-agent baseline, dual size:baseline, shared measureMdFiles seam) and disambiguates it from the separate DEFECT.AGENT-FILE-SIZE-CAP-BREACH 45K-CHAR guard (two units, two purposes). - Docs: now that #1096's docs/TESTING-SUITES.md "Workflow size budget" section is in next, fold in the agent coverage here (renamed to "Workflow & agent size budget"): agent caps + per-agent baseline + the how-to + reference rows, and the disambiguation from the 45K-char guard. - Minor (negative proof): add a boundary-fixture test exercising the hard-cap comparison at cap-1/cap/cap+1 through the real lfByteCount path, so a future threshold/operator edit can't silently neuter a cap. - Nit: align the tier test name wording ("stays within") with the <= operator. Negative proof on a real tracked agent (gsd-planner): baseline catches +10 B; XL hard cap catches 57,516 > 57,344 with the baseline current. Closes #1095 (PR 3/3 child); landing this completes the #1074 epic. |
||
|
|
74d7bc8239 |
test(#1074): add additive per-file workflow size baseline guard (PR 1/3) (#1089)
* test(#1074): add additive per-file workflow size baseline guard (PR 1/3) Introduces a committed per-file size baseline scheme alongside (not replacing) the existing tier anti-creep tests. Green by construction — the baseline records current sizes, so both schemes pass side by side during migration. - scripts/lib/allowlist-ratchet.cjs: add assertFileBaseline (third pure helper, same injected-fail style) — per-file growth/shrink/add/remove diff vs baseline. - scripts/workflow-size.cjs: single source of truth for LF-normalized byte counting (#683) + workflow enumeration, shared by the guard and the generator so they can never measure differently. Lives in scripts/ root (NOT scripts/lib/) because it is dev/CI-only tooling — scripts/lib/ is bundled into the installed runtime, scripts/ root is not, so this keeps it out of the shipped payload. - scripts/update-size-baseline.cjs + npm run size:baseline: regenerate the snapshot (sorted keys, trailing newline, idempotent). - tests/workflow-size-baseline.json: generated snapshot (88 workflows). - tests/workflow-size-budget.test.cjs: import the shared counter (drops the duplicated local byteCount) and add the per-file baseline describe block. - Tests for the helper, the shared module, and the generator (incl. round-trip and fault-injection cases). Refs #1074. Part 1 of 3; PR 2 swaps enforcement, PR 3 covers the agent test. * test(#1074): regenerate workflow baseline after Update-branch merge with next The 'Update branch' merge (652a916b) pulled in next's update.md change (#1090) without regenerating the snapshot, leaving the per-file baseline stale by one file. Re-ran `npm run size:baseline` so the committed baseline matches the merged workflow files. Refs #1074. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
94872662e9 |
feat(#1082): complete phase 5 — descriptor-drive all install surfaces + materialize the InstallPlan — ADR-857/1016/58 (#1080)
* feat(#1077): phase 5f-2 — drive the hookEvents dialect (PostToolUse/AfterTool) from the descriptor postToolEvent (bin/install.js) and preToolEvent (applySettingsJsonHooks in runtime-hooks-surface.cts) now select the event-name dialect from registry.runtimes[id].runtime.hookEvents instead of the hardcoded (runtime === 'gemini' || runtime === 'antigravity') check: hookEvents === 'gemini' → AfterTool/BeforeTool; else → PostToolUse/PreToolUse. hookEvents threaded into the applySettingsJsonHooks opts bag. Equivalence-preserving (Codex-verified): hookEvents 'gemini' is exactly {gemini, antigravity}, 'claude' the rest; undefined → claude dialect (matches the old else). The per-event SET guards (isQwen||claude → SubagentStop/Stop/PreCompact; runtime==='claude' → FileChanged; isGemini → Gemini agent-events) stay HARDCODED — hookEvents (2-value) is too coarse to drive them (the event set differs within hookEvents='claude'); per-event-set drive tracked in #1076. Registry-parity test (enh-1077): asserts BOTH post-tool (AfterTool/PostToolUse) AND pre-tool (BeforeTool/PreToolUse) dialects are a pure function of hookEvents, for gemini/antigravity/claude/augment — non-vacuous (catches a broken hookEvents thread). Closes #1077 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1077): build hooks/dist in before() so dialect-drive test passes in scoped CI hooks/dist is gitignored and absent in scoped/windows CI jobs that do not pre-run build:hooks. Without it, install() finds no hook files and all AfterTool/BeforeTool/PostToolUse/PreToolUse event arrays come back empty, failing every hook-presence assertion. Added an idempotent ensureHooksDist() called in a top-level before() — mirrors the pattern from bug-376-claude-js-hook-gsd-rewriter.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1055): add installSurface/writesSharedSettings/permissionWriter/extendedHookEvents to runtime descriptors Purely additive: four new fields on all 16 runtime capability.json descriptors, validator extended with three new closed-vocab sets, registry regenerated. Test fixtures (VALID_RUNTIME_CAP and makeRuntimeCap) updated to include the new required fields so all 255 capability-registry tests continue to pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): drive per-event hook guards from extendedHookEvents descriptor Replace hardcoded runtime-name checks (isQwen||runtime==='claude', runtime==='claude', isGemini) in applySettingsJsonHooks with a single descriptor-driven extendedEvents array derived from the new opts field. Remove isQwen and isGemini derivations (no remaining uses after the three guard blocks are migrated). Wire extendedHookEvents from the capability registry in bin/install.js call site. Add behavioral regression test (enh-1076-extended-hook-events-drive.test.cjs) confirming the drive is purely descriptor-based and runtime-name-agnostic. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1055): drive resolveRuntimeConfigIntent from the runtime descriptor; retire hand-kept REGISTRY - Rewrites src/runtime-config-adapter-registry.cts to require capability-registry.cjs and read installSurface / writesSharedSettings / permissionWriter from runtimes[id].runtime; deletes the hand-kept REGISTRY const (ADR-857 phase 5g drive 2). - ALLOWED_CONFIG_RUNTIMES is now derived from descriptor entries that have installSurface. - Fixes the configFormat parity gate in scripts/gen-capability-registry.cjs to read installSurface directly from capMap descriptor bodies, breaking the require cycle (adapter now requires the generated registry; gen-script must not require the adapter). - Adds golden-master test tests/enh-1055-config-intent-descriptor-drive.test.cjs (41 tests) pinning all 16 runtimes' return shapes and the TypeError-on-unknown contract. - Updates scripts/lint-test-file-count.allowlist.json (config module, +1 file). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): make hooksSurface descriptor load-bearing for the settings-json hook-skip - Adds hooksSurface?: string to ApplySettingsJsonHooksOpts and destructuring in applySettingsJsonHooks (src/runtime-hooks-surface.cts). - Replaces the hardcoded !isOpencode && !isKilo hook-skip guard with hooksSurface !== 'none'; removes the now-unused isOpencode/isKilo derivations (ADR-857 phase 5g drive 3). - Passes hooksSurface from the runtime descriptor at the applySettingsJsonHooks call site in bin/install.js using the established _capabilityRegistry?.runtimes?.[runtime]?.runtime?.hooksSurface idiom. - Extends tests/enh-1076-extended-hook-events-drive.test.cjs with two new suites proving: (a) hooksSurface:'none' writes no hooks regardless of runtime name; (b) hooksSurface:'settings-json' writes hooks even for 'opencode' (previously hardcoded to skip). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: record installSurface/writesSharedSettings/permissionWriter/extendedHookEvents descriptor axes in ADR-1016 Add Decision 7a documenting the four axes added in the 5f-completion pass, update axis counts from "six" to "twelve", note 5f-completion drives as done in Decision 8's ladder, update Out of scope to reflect #1055/#1076 are done and 5g (InstallPlan capstone) remains the only open phase. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1055): parity gate must fire on configFormat↔installSurface mismatch (read installSurface at the descriptor level) The test fixture makeRuntimeCapMap did not include installSurface in the runtime object, so the gate's typeof r.installSurface !== 'string' guard always skipped the entry and never threw. Added installSurface as an optional third parameter to makeRuntimeCapMap and passed the correct installSurface values ('settings-json' for claude, 'codex-toml' for codex) to the two THROWS tests. The gate implementation already reads r.installSurface correctly from the descriptor level. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): add installSurface↔hooksSurface + extendedHookEvents↔hookEvents consistency gates with rejection tests GATE A: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES map in validateRuntimeBody enforces that a runtime's hooksSurface is valid for its installSurface (e.g. profile-marker-only only allows none, codex-toml only allows codex-hooks-json). Derived from the 16 real runtime descriptors. GATE B: validateRuntimeBody checks that if extendedHookEvents contains Gemini agent-events (BeforeAgent/AfterAgent/BeforeModel), hookEvents must be 'gemini'; if it contains Claude-family events (SubagentStop/Stop/PreCompact/FileChanged), hookEvents must be 'claude'. Added 10 rejection tests in suite 27 covering each gate + each new field validator. All 16 real runtimes satisfy both gates (verified before coding). Exports: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES, VALID_INSTALL_SURFACES, VALID_EXTENDED_HOOK_EVENTS, VALID_PERMISSION_WRITERS, validateRuntimeBody. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1076): strengthen hooksSurface-drive assertions; defensive hooksSurface fallback; drop vacuous dup 1. bin/install.js: add explicit literal fallback for hooksSurface when the committed capability registry fails to load (opencode/kilo → 'none', all others → 'settings-json'). The descriptor is always the source of truth in normal operation. 2. enh-1076 Suite 7: change SessionStart assertion from key-presence (hasOwnProperty) to at least-one-command (hasHooksFor), so the test fails if hooks are initialized-but-empty. ensureHooksDist() in before() guarantees hook files exist. 3. enh-1055 Test 2: remove vacuous duplicate suite that re-asserted intent.runtime === row.runtime already fully covered by Test 1's deepStrictEqual over all four fields. 4. capability-registry.test.cjs: fix stale comments in the grok-skip test that claimed the parity gate uses the adapter registry; gate reads purely from the descriptor (installSurface absent → typeof r.installSurface !== 'string' → soft-skip). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1082): materialize the InstallPlan — collect install-level descriptor axes into resolveInstallPlan; route install()/finishInstall() through it (ADR-58/5g) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: record 5g InstallPlan materialization (ADR-58 Accepted, ADR-1016 phase-5 complete) ADR-1016 Decision 8 step 7 updated to DONE: resolveInstallPlan(runtime) in runtime-config-adapter-registry collects install-level descriptor axes into the typed InstallPlan consumed by install()/finishInstall(). Out-of-scope section updated: 5g capstone is complete, phase 5 fully materialized. ADR-1016 line ~20 updated: InstallPlan IS now materialized (both halves). ADR-58 Implementation note added (2026-06-11): realized in runtime-config-adapter-registry (co-located with adapter-selection). CONTEXT.md Runtime Config Adapter Registry entry extended to document resolveInstallPlan and both-halves realization. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1082): update install drift guard to the resolveInstallPlan seam (5g) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
837991d5a2 |
chore(#1087): Windows test-portability lint + n/no-path-concat + LF normalization (#1088)
* chore(#1087): add Windows test-portability lint + DEFECT.WINDOWS-TEST-PORTABILITY Local gsd-test runs Mac+Linux only, so Windows-only test failures (Git Bash msys2 not honoring Node's chmod exec bit for PATH-executing extension-less scripts; `/` vs `\` path assertions) surface for the first time in CI's windows lanes — repeatedly (most recently PR #1084's #381 fix). Add scripts/lint-windows-test-portability.cjs: a high-signal, low-false- positive tripwire that flags any tests/**/*.test.cjs combining a chmod exec bit with a `sh -c`/`bash -c` invocation and no process.platform guard, unless annotated `// windows-portability-ok: <reason>`. Wired into lint:ci and runnable as `npm run lint:windows-test-portability`. Clean against all 721 current test files (zero pre-existing violations). Document the broader anti-pattern as DEFECT.WINDOWS-TEST-PORTABILITY in CONTEXT.md (.symptom/.examples/.detect/.fix-forward/.prevention): local gsd-test cannot substitute for the CI windows lane — watch it green before declaring a PR done. tests/lint-windows-test-portability.test.cjs covers the scanContent matrix. Closes #1087 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1087): enable n/no-path-concat + add .gitattributes/.editorconfig (LF) Two cross-platform "free wins" complementing the Windows test-portability lint: - Enable eslint-plugin-n's n/no-path-concat (already-installed plugin) as 'error' — flags string path concatenation (the / vs \ separator class). Zero existing violations, so it's a clean ratchet, not a refactor. - Add .gitattributes (* text=auto eol=lf + binary exemptions) and .editorconfig (LF, UTF-8, final newline, trim whitespace) to normalize line endings and kill the CRLF-only-fails-on-Windows class at the source. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0cc37a94c2 |
feat(#1056): phase 5e — close ConverterName enum + configFormat↔installSurface parity guard (#1057)
Two gen-time validation tightenings (validation-only; serialized registry content
unchanged; bin/install.js + adapter + descriptors untouched):
Part B: validateArtifactKindEntry now requires artifactLayout[].converter ∈
VALID_CONVERTER_NAMES (15 names, all exported by install.js) ∪ {null} — a typo'd
converter fails at gen time instead of silently → installExports[name]===undefined
at install time.
Part A: a HARD buildRegistry parity gate asserts each runtime descriptor's
configFormat agrees with the adapter registry's installSurface via a fixed mapping
(cursor-hooks-json/profile-marker-only→none, codex-toml→toml, copilot-instructions→
markdown, cline-rules→markdown-dir, settings-json→settings-json) — keeps configFormat
from drifting; prerequisite-validation for the deferred full drive (#1055).
The full config-writing drive (retire resolveRuntimeConfigIntent) is deferred to
#1055: configFormat is lossy vs installSurface (cursor vs profile-marker both → none;
opencode/kilo permissionWriter has no descriptor field) → needs schema extension.
Closes #1056
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
4698b3e349 |
fix(#1051): force-exit + per-chunk timeout for the windows full-test lane; close leaked test handles (#1054)
The `full test (windows-latest, 22)` job intermittently got CANCELLED at its 20m wall-clock cap with no failed test step — a false-negative gate (recurrence of #869). Root cause: a unit test leaves an open event-loop handle, so the chunk's `node --test` child hangs ~150s on Windows after its last test prints; two such stalls push the already-~13m job past 20m. Fix (defense in depth): - run-tests.cjs: pass --test-force-exit (Node >=22; engines requires >=22.0.0) so the runner exits once all tests finish regardless of lingering handles — the durable backstop. Account for the flag in the argv-length ceiling. - run-tests.cjs: add a per-chunk execFileSync timeout (default 600000ms, env RUN_TESTS_CHUNK_TIMEOUT_MS) that fails loudly with a diagnostic naming the chunk's files, so a hung chunk can never silently eat the job budget. - perf-316 test: terminate both Worker threads on all paths (afterEach + finally) so they cannot outlive the test. - locking-bugs test: kill spawned children in a finally that wraps the whole spawn -> waitFor -> barrier-release -> Promise.all sequence, so a barrier timeout no longer leaks live child processes. - Refresh the stale synckit comment (synckit/SDK bridge was removed). Regression tests in run-tests-harness: a hung chunk hits the per-chunk timeout and fails with a clear message; force-exit lets a chunk with a leaked handle exit cleanly. Closes #1051 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
58ed55683e |
feat(#1049): phase 5d — drive artifactLayout from the runtime descriptor (retire the 128-LOC switch) (#1053)
resolveRuntimeArtifactLayout now builds Layout from
registry.runtimes[id].runtime.artifactLayout[scope] — a loop dispatching each
ArtifactKind through the SAME 5 builders (commandsKind/agentsKind/skillsKind/
convertedCommandsKind/kimiAgentsKind, unchanged) by (kind, converter, nesting) —
replacing the hardcoded switch(runtime). Equivalence-preserving for all 16 runtimes
× {global, local} (Codex-verified, no divergence). -43 LOC; bin/install.js + the
converters + the install loop untouched. getInstallExports()[converterName]
resolution, configDir threading, scope default, unknown-runtime guard all preserved.
Driving the local scope surfaced a 5a gap: the old switch had no scope branch for 13
runtimes (cursor/gemini/codex/copilot/antigravity/windsurf/augment/trae/qwen/hermes/
codebuddy/opencode/kilo) → local == global for them, but 5a authored local:[].
Backfilled local=global for those 13 (descriptor-faithful; a fall-through shim would
wrongly give cline/kimi local=global). claude/cline/kimi scope-gating untouched.
validateArtifactKindEntry tightened: destSubpath/prefix/nesting/converter required
(ConverterName enum still open — 5e). New 39-case deep-equal golden equivalence test.
Closes #1049
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
cec7e704d6 |
feat(#435): expand workflow-policy linter to full Cartesian matrix cross-product (#1050)
`expandRunsOn` enumerated a multi-axis `strategy.matrix` one key at a time,
producing partial realization contexts. A true `os × shell` matrix therefore
left `${{ matrix.shell }}` unresolvable against any `{ os: ... }`-only context,
firing spurious `UNRESOLVABLE_MATRIX` violations and leaving shell-pinning
coverage incomplete on Cartesian jobs.
Enumerate the full GitHub Actions cross-product of all base-list matrix keys
(every `matrix.<k>` array, excluding the `include`/`exclude` control keys) via
a named `cartesianProduct` helper. Each realization's context now carries a
value for every matrix key, so `${{ matrix.<key> }}` resolves per realization.
Single-axis matrices keep byte-for-byte identical output; only multi-axis
matrices change shape. The `include` and `exclude` blocks are unchanged
(full tuple-aware exclude is a documented out-of-scope follow-up).
Tests: updated the `os × shell` test to assert post-fix behavior (4 step
realizations, 2 WRONG_SHELL_FOR_OS, 0 UNRESOLVABLE_MATRIX); added a compliant
`os × node-version` cross-product test; added a fast-check property test that
the realization count equals the product of axis lengths and that no axis key
is dropped from any context.
Closes #435
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
1fab2e10ba |
fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver (#1045)
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker, gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks. On a shim-only install — where gsd-tools.cjs exists under the runtime home but gsd-tools is NOT on PATH — those calls fail with "command not found" and the agent silently skips init/state/validate/commit ceremony, deferring to the orchestrator or bypassing GSD bookkeeping entirely. were never migrated, so it persisted on Claude Code and every other runtime that consumes the source agents directly. Only gsd-phase-researcher.md carried a resolver — and a stale, claude-only truncated one. Fix (all runtimes): - Inject the canonical multi-runtime gsd_run preamble (byte-equal to _runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/ augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every command-position bare gsd-tools to gsd_run. - Upgrade gsd-phase-researcher.md's stale resolver to the canonical one. - Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the sync caught and corrected a mis-placed preamble during development). - Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards to agents/ so no runtime can silently regress. Closes #1041 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1041): backfill changeset PR number to 1045 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fbd62cd84f |
feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).
Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).
Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).
Closes #1035
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
2ac6592096 |
feat(#1023): first phase-6 cutover — ui-review (verify:post) inline → loop.render-hooks dispatch (#1024)
Replace the inlined ui-review invocation in autonomous.md §3d.5 with a loop.render-hooks verify:post dispatch — the first workflow to consume render-hooks and fire a skill from it (closes the #1018 live-execution residual as real wiring). Capability-driven, equivalence-preserving for the current registry (only ui-review at verify:post, default on): fires gsd-ui-review under the same precondition (UI-SPEC exists via consumes-gate + workflow.ui_review). Gate findings (real pattern issues, fixed so every future cutover inherits them): - bug-2643 static "Skill() references a real skill" check vs templated Skill(skill="gsd-${ref.skill}") dispatch → skip ${...}-templated names. - Coverage moved, not lost: gen-capability-registry now validates steps[].ref.skill in skills + ref.agent in agents + rejects gsd- double-prefix. - Tightened §3d.5 tests; markdown clarity (consumes rule, LLM-native JSON read, UI-REVIEW.md score hint). gsd-ui-review skill + §3a.5/ui-phase untouched. §5.6/ui-phase cutover deferred (#1022 step-can-halt-vs-gate model question). Closes #1023 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
19edab21da |
fix(#1006): rc CHANGELOG preview crash on malformed changeset fragment + validate fragment content at the gate (#1007)
* fix(#1006): harden render --preview against fragment parse failures `render --preview` wrote `report.preview` unconditionally. When a `.changeset` fragment fails to parse, `cmdRender` early-returns with `{exitCode:1, report: {failures}}` and NO `preview` key, so `process.stdout.write(undefined)` threw ERR_INVALID_ARG_TYPE and the rc release job's "Preview CHANGELOG" step died with a cryptic TypeError that masked the real cause. Guard the preview write on `typeof report.preview === 'string'` (ADR-227: shape, not just type); when absent, fall through to the existing failure reporter that names the offending fragment and exits non-zero — identical to a non-preview render. Also backfills the stray placeholder `pr: 0` -> `pr: 939` in .changeset/936-convergence-inline-plan-phase.md that triggered the live failure. Regression test (red-then-green verified) added at the render --preview seam. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1006): validate changeset fragment content at the Changeset Required gate The `Changeset Required` gate (scripts/changeset/lint.cjs) only checked that a `.changeset/*.md` fragment EXISTS in the PR diff; it never validated the fragment's contents. So a malformed fragment (e.g. an un-backfilled `pr: 0` placeholder) silently merged to `next` and only detonated later in the rc release job. This is the upstream prevention for #1006 — the crash hardening turns the failure into a clear message, this stops the bad fragment ever reaching the release path. evaluateLint now accepts `fragmentFailures` and fails with the typed reason `fail_invalid_fragment` (naming each offending file) before the existence/ opt-out checks — a malformed fragment beats `no-changelog`, since it will break the render regardless. main() reads + parseFragment()s every changed fragment: a deleted fragment (not on disk) is skipped, a present-but-unreadable one fails closed. Tests assert on the typed LINT_REASON enum (no raw-text matching), a precedence case over the opt-out label, and an end-to-end suite that drives the real main() against a temp git repo (malformed -> fail, valid -> pass, deleted -> skipped) so the wiring is regression-proof. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1006): assert the typed --json report in the preview regression test Code review flagged the preview parse-failure regression test for positive raw-text matching on CLI output (`combined.includes('bad-fragment.md')` / `'invalid_pr'`), which this repo's testing standards forbid. Keep the non-json `runRenderRaw` call for the negative crash proof (the ERR_INVALID_ARG_TYPE crash lives only on the non-json stdout.write path), and add a `--json` invocation that asserts the offending fragment + typed `invalid_pr` reason via the structured `report.failures[]` surface instead of rendered prose. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
adaf3e17d8 |
fix(#1001): make bug-969 hardening tests hermetic + move build tsbuildinfo out of shipped tree (regression from #996) (#1002)
* fix(#969): make bug-969 hardening tests hermetic and move build tsbuildinfo out of shipped tree (regression from #996) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#969): self-heal legacy bin-local tsbuildinfo and make sentinel test hermetic (adversarial-review follow-ups) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(changeset): set pr number to 1002 * docs(#1001): record DEFECT.SHARED-ARTIFACT-MUTATION-IN-CONCURRENT-TEST anti-pattern in CONTEXT.md --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
88e30d5342 |
test(#969): fix stale-build flake (incremental + re-emit-on-missing) and make runGsdTools retry-once before surfacing subprocess kills (#996)
Closes #969 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
1fd5c86a1e |
fix(#983): rewrite bare .claude paths in Trae/Windsurf converters (Codex/Cline parity) (#995)
* fix(#983): rewrite bare .claude paths in Trae/Windsurf converters (Codex/Cline parity) Both convertClaudeToWindsurfMarkdown and convertClaudeToTraeMarkdown only handled trailing-slash .claude/ forms; bare ~/.claude and $HOME/.claude references (e.g. configDir = ~/.claude, RUNTIME_CONFIG_DIR=".../$HOME/.claude") survived conversion and pointed users at the wrong config dir. Fix: add bare-form replacements using negative lookahead (?![\w-]) to protect .claude-plugin and .claudeignore, mirroring Cline (#782) and Codex (#570) precedent. Also adds CLAUDE_CONFIG_DIR -> WINDSURF_CONFIG_DIR / TRAE_CONFIG_DIR rewrite. _applyRuntimeRewrites windsurf case gets matching \b-anchored bare-form lines, mirroring the existing trae case. Closes #983 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(#983): backfill changeset pr number (995) --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
36b68ac81d |
fix(#977): map ephemeral fnm multishell execPath to a stable fnm alias in normalizeNodePath (#992)
* fix(#977): map ephemeral fnm multishell execPath to a stable fnm alias in normalizeNodePath Closes #977 * chore(#977): backfill changeset pr number (992) --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
972a41a528 |
fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) (#990)
* fix(#967): make verify key-links docs author-strict (from:/to: are file paths; symbols go in via:) Closes #967 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(#967): backfill changeset pr number (990) --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
921a7cd618 |
fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op (#986)
* fix(#974): error on graphify --budget with missing/non-numeric value instead of silent NaN no-op When `--budget` was the last arg or followed by a non-numeric token, parseInt(undefined/NaN-string, 10) produced NaN. NaN is falsy so both the router check and applyBudget gate silently skipped budget trimming. The query ran unbounded with no warning. Fix: guard in graphify-command-router.cts — if args[budgetIdx+1] is absent or parses to NaN, emit ERROR_REASON.USAGE and return early. Defensive fix in graphify.cts: tighten `if (!budgetTokens)` → `if (budgetTokens == null)` and `if (options.budget)` → `if (options.budget != null)` so a real 0/NaN caller is handled predictably by both independent guards. Regression tests: 16 cases (unit/mock, subprocess, property-based) covering boundary inputs: missing value, non-numeric, valid integers, and fast-check properties over the budget parse contract. Closes #974 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(#974): backfill changeset pr number (986) * fix(#974): convert static property (b) to real fc.assert property with constrained generator + deterministic seed The test named "property: --budget as last arg always produces usage error" was a static test with no fc.assert — it only checked a single hardcoded term ("someterm") and could never flake or produce a fast-check path. This is a generator/property bug (case b): the test was mislabeled as a property test but lacked parameterization. Root-cause of CI path "46:0:0" / seed 42 failure: a naive parameterized version without the !startsWith('--') filter could feed term='--budget', causing args.indexOf('--budget') to hit index 2 (the term slot) rather than index 3 (the flag slot), placing the router in a different code path. The property still holds — NaN detection fires on rawBudget='--budget' — but the assertion text referenced the wrong invariant, making the failure appear spurious. Fix: constrain the generator to non-flag terms (filter out strings starting with '--') and pass explicit { seed: 42, numRuns: 200 } to fc.assert so the test is fully deterministic in CI regardless of GSD_FC_SEED. No change to src/graphify-command-router.cts (router is correct). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#974): constrain non-numeric budget property generator to genuinely-NaN values and pin fast-check seeds (CI determinism) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#974): use a strict valid-term generator and pin seeds so budget property tests are deterministic Replace fc.string({minLength:1}).filter(!startsWith('--')) term generators in properties (b) and (d) with a shared validTerm = fc.stringMatching(/^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/) that is alphanumeric-leading and contains no whitespace, flags, or sign-numerics. This eliminates the class of CI failures where the old generator produced out-of- contract inputs (" ", "+5", empty) that the router legitimately rejects for reasons outside the budget-parse contract under test. Pin distinct seeds (1001–1004) on every fc.assert for CI determinism. Stress-tested at numRuns=100000 per property and across 10 seeds (1,2,7,13,42,43,44,99,12345,999999) at numRuns=2000 — all pass. No src change: gsd-core/bin/lib/graphify-command-router.cjs is correct. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#974): replace flaky term-fuzzing properties with deterministic examples; keep value-fuzzing properties The validTerm regex /^[A-Za-z0-9][A-Za-z0-9_.-]{0,29}$/ admitted single-digit strings like "0" which are falsy; the router's `if (!term)` guard fires before the budget-missing-value path, producing a spurious errFn call. Properties (b) and (d), which test the BUDGET contract (not term handling), are replaced with deterministic example loops over fixed valid terms. Properties (a) and (c), which fuzz the BUDGET VALUE with a fixed term, are kept unchanged (seed+numRuns pinned). The validTerm generator is fully removed. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
626575cbc5 |
fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works (#982)
* fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works The dispatcher built `{ name, archivePhases }` but never parsed `--force`, so `options.force` was always `undefined` and the guard inside `cmdMilestoneComplete` (which tells users to "Re-run with --force to override") could never be bypassed. Add `const force = args.includes('--force')` and pass it into the options object. The guard already honors `options.force` — no changes to milestone.cts needed. Closes #978 * chore(#978): backfill changeset pr number (982) --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
caca4d255c |
feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4) (#988)
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4) Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the route fn. capabilities/intel/capability.json declares the intel command family; commands-only (skills:[]), declares the existing intel.enabled gate (default false — behavior unchanged). Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical; existing intel.test.cjs passes unchanged. Completes the first-party command- family cutover sequence (graphify/audit/intel). Closes #985 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI) The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning` expectation while the router builds it with path.join(cwd, '.planning') → backslashes on Windows, so the planningDir-arg assertions (query/status/diff/ snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute the expectation with path.join (cross-platform); production router unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9a03539c2d |
feat(#981): audit-uat + audit-open command cutover — commands-only capability (ADR-857 phase 4d-impl-3) (#984)
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs case arms to a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring the graphify cutover (#972). New src/audit-command-router.cts exports routeAuditUat/routeAuditOpen, each lazily requiring only its backing module (uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads. capabilities/audit/capability.json declares the two command families; commands-only (skills:[]), no config gate, audit_review cluster untouched. Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical; existing audit regression tests pass unchanged. Closes #981 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
77bfd943dd |
feat(#972): graphify command cutover — first capability owning a command family (ADR-857 phase 4d-impl-2) (#975)
Migrate graphify into a Capability that owns its `graphify` command family, dispatched via the registry (#961 mechanism) instead of a hardcoded case. graphify is now an enable/disable plug-in. - src/graphify-command-router.cts: routeGraphifyCommand (standard route*Command), reproduces the removed case EXACTLY (query +--budget, status, diff, build, hidden build snapshot, usage/unknown errors); injectable _graphify test seam. - capabilities/graphify/capability.json: role feature, tier:full, skills:[graphify], config:{graphify.enabled default false}, commands:[{family:graphify, module, router:routeGraphifyCommand}]. - removed case 'graphify' from gsd-tools.cjs; graphify now flows default -> dispatchCapabilityCommand -> commandFamilies.graphify -> router. - regenerated registry (commandFamilies/bySkill/configSchema/profileMembership/ capabilityClusters for graphify); tier:full keeps 4c install/surface a no-op. Equivalence-proven: 8 recording-mock unit tests assert the exact fn+args per subcommand (budget, snapshot-vs-build); 9 subprocess tests assert distinguishing output shapes; existing graphify tests pass unchanged through the new path. Surfaced (not silently accepted): the pre-existing --budget-no-value NaN no-op quirk, preserved for equivalence, filed separately. Closes #972 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
354e0e1b94 | fix(ratchet): add --update drift repair + inherited-drift guidance to regression-name lint (#971) | ||
|
|
0c567c17e4 |
feat(#961): capability command mechanism (commandFamilies index + default-case dispatch) — ADR-857 phase 4d-impl-1 (#964)
* feat(#961): capability command mechanism — commandFamilies index + default-case dispatch (ADR-857 phase 4d-impl-1) Build the capability command mechanism per ADR-959: the `commands` declaration field on the feature role, the registry commandFamilies index, and a real dispatchCapabilityCommand consulted in runCommand's default case (replacing the dead _dispatchNonFamily shim's role). The registry DISCOVERS a standard route*Command (no rebuilt handler table). On a default-case command, dispatch enforces a bare-.cjs-basename, resolves the module under gsd-core/bin/lib/, asserts confinement, requires the resolved path, and own-property-guards the router export before calling it. A require.main===module guard makes gsd-tools.cjs importable for tests; the CLI path is unchanged. Additive: commandFamilies is empty today, so the default case is behavior- preserving for every command; the 10 dead _dispatchNonFamily sites are untouched (future migration markers). The graphify cutover is the separate next step. Closes #961 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#961): surface capability router failures structurally + enforce sync contract Review follow-up: dispatchCapabilityCommand now wraps the router invocation so an unexpected (non-ExitError) throw is converted to a structured, attributed error(msg, SDK_FAIL_FAST) — honoring --json-errors — instead of escaping as a raw stack trace; an intentional ExitError propagates unchanged. An async router (returns a thenable) is rejected loudly with a structured error (the contract is synchronous, like the 12 host routers). Not shipped as "consistent with existing behavior": the host's pre-existing version of this gap is filed as #965. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e8cfb560b2 |
fix(#851): correct Codex quick adapter for generic multi_agent_v1 schema (#958)
* fix(#851): correct Codex adapter for generic multi_agent_v1 schema
The Codex skill adapter header in getCodexSkillAdapterHeader() documented
typed spawn_agent(agent_type=...) as a direct, unconditional mapping for all
Task()/Agent() calls. In sessions exposing only the generic multi_agent_v1
schema (message/items/fork_context — no agent_type field), this mapping is
silently invalid: the orchestrator cannot natively dispatch typed gsd-planner/
gsd-executor agents and may fall back to inline execution or produce errors.
Fix: Section C now requires schema detection before spawning. It documents the
typed mapping as conditional on the agent_type-capable schema (e.g. multi_agent_v2)
and introduces an explicitly-labeled generic-agent workaround for multi_agent_v1
sessions — read the agent TOML, inject its instructions as a role-preamble, and
call spawn_agent(message=...) — clearly marking the result as NOT equivalent to
typed gsd-planner/gsd-executor execution.
Regression test: tests/bug-851-codex-quick-adapter-agent-type-fallback.test.cjs
asserts schema-awareness language, the multi_agent_v1 fallback, the workaround
label, and backward compat with the existing bug-279 typed-spawn contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add changeset for PR #958
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#851): keep Codex adapter block consistent with materialized skill surface
The `~/.codex/agents/<agent-name>.toml` literal introduced in the #851 prose
was being rewritten to the real install path by `_applyRuntimeRewrites` (the
`~/.codex/` → pathPrefix substitution) before the SKILL.md was written to
disk. `getCodexSkillAdapterHeader()` still returned `~/.codex/agents/...` so
the test assertion (exact match between builder output and materialized file)
always failed.
Fix: replace the `~/.codex/agents/` literal with the runtime-neutral form
`agents/<agent-name>.toml` plus a parenthetical naming `$CODEX_HOME/` — which
is not matched by any rewrite pattern and survives the path-substitution step
unchanged. The #851 schema-detection + generic-subagent-fallback intent is
fully preserved.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#851): resolve active Codex config root in fallback; strengthen tests (adversarial review)
- Rewrites the generic-agent workaround step 1 to explicitly describe
active config root resolution (priority: $CODEX_HOME → --config-dir →
--local .codex → default global dir) without the literal ~/.codex/
substring that _applyRuntimeRewrites replaces, preventing bug-3582
divergence.
- Replaces OR/loose-includes test assertions in bug-851 with AND-logic
checks covering all four required elements: (a) schema-detection step,
(b) active-config-root resolution for the TOML path including all three
override mechanisms, (c) NOT-equivalent-to-typed-gsd-planner/gsd-executor
label, and (d) fail-closed rule when typed dispatch is mandatory.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#851): register bug-851/947/948/950 in lint-regression-test-names allowlist
The ratchet (
|
||
|
|
46967baae8 |
fix(#948): guard STATE.md no-op writes; preserve milestone_name/stopped_at (closes #944) (#952)
* test(#948): add regression tests for no-op write guard and record-session auto-create (#944) Red before fix: 11/15 tests fail. Green after: 15/15. Covers zero-match patch byte-identity, milestone_name preservation, stopped_at frontmatter-wins, record-session auto-create fallback, and adversarial fixtures (CRLF, empty body, non-canonical labels). Also registers bug-948-state-noop-write-guard.test.cjs in the state bucket of lint-test-file-count.allowlist.json. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#948): guard STATE.md no-op writes; preserve milestone_name/stopped_at (closes #944) Shared root cause: `readModifyWriteStateMd` wrote STATE.md unconditionally even when the transform produced no change, and `syncStateFrontmatter` re-derived frontmatter from the possibly-stale body on every write. Three coordinated fixes in src/state.cts: 1. readModifyWriteStateMd: add no-op guard — when transform result === input content, skip the write entirely (no platformWriteSync, no last_updated bump, no frontmatter re-derive). Fixes #948 zero-match phantom write and the #944 phantom last_updated bump. 2. syncStateFrontmatter: extend existing-frontmatter preserve logic — fall back to existingFm['milestone_name'] / existingFm['milestone'] when the derived value is the template placeholder 'milestone' (getMilestoneInfo returns this literal when it cannot match the version in ROADMAP.md); prefer existingFm['stopped_at'] / existingFm['paused_at'] over a body-derived value (the frontmatter value, written by the canonical record-session path, wins over stale historical body lines). Mirrors the fallback already in cmdStateJson. 3. cmdStateRecordSession: when --stopped-at / --resume-file are supplied but body labels are absent, DWIM auto-create a canonical ## Session section (mirroring how add-decision / add-blocker / record-metric auto-create their sections). Never return a silent recorded:false when the caller supplied values. SDK check: no sdk/src/state.ts exists in this repo (the comment in cmdStateSnapshot references a sibling concern in the TypeScript SDK codebase, which is a separate repo not present here). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: add changeset for PR #952 (fix #948/#944) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#948): correct stopped_at preserve rule; adjust test for sync behaviour The "always prefer frontmatter stopped_at" rule in syncStateFrontmatter was too aggressive — it broke phase.complete which intentionally updates stopped_at in the body and expects syncStateFrontmatter to pick it up. The primary fix (no-op guard in readModifyWriteStateMd) already prevents the stale-body-overwrites-frontmatter scenario from #948: the file is not written when the transform produces no change, so syncStateFrontmatter never runs on a zero-match patch. The body-derived value can only win when an actual write occurs, which means the body was legitimately updated. Reverted to the original #905 rule for stopped_at/paused_at: fall back to existing frontmatter only when the derived value is absent (empty/null). Also adjusted the sync-suite test to assert what state sync actually does (milestone_name preservation) rather than a stopped_at-wins property that state sync does not have by design. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#944): update existing session block in place (adversarial review) HIGH finding: the DWIM auto-create in cmdStateRecordSession was appending a second ## Session block unconditionally, even when one already existed with non-canonical content (e.g. a markdown table). Both buildStateFrontmatter and cmdStateSnapshot read only the FIRST ## Session block via regex, so the newly-written Stopped at / Resume file values landed in the second, invisible block — frontmatter stopped_at stayed stale and state-snapshot returned nulls. Fix: check for an existing ## Session heading. When one is present, normalize that section in place by replacing its body with canonical **Last session:** / **Stopped at:** / **Resume file:** bold-label lines. Only append a brand-new section when NO ## Session heading exists. LOW finding: the auto-create scaffold emits **Last session:** but cmdStateSnapshot only matched **Last Date:**, so session.last_date was null after auto-create despite a valid timestamp being written. Fix: extend the lastDateMatch regex in cmdStateSnapshot to also accept **Last session:** / Last session: (the form the scaffold writes). Tests: 3 new tests added to bug-948-state-noop-write-guard.test.cjs that confirmed failure against the previous HEAD and pass after this fix: - exactly one ## Session block after record-session with non-canonical existing block - state-snapshot sees correct stopped_at via first Session block (not a duplicate) - state-snapshot session.last_date is non-null after auto-create on body-less file Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#944): improve in-place section replace to cleanly remove old body content The previous regex `/(^## Session[ \t]*$)([\s\S]*?)(?=\n^## |\n*$)/im` with a lazy match consumed nothing after the heading, so old non-canonical body content (e.g. table rows) remained after the new canonical lines. While functionally correct (parsers found the canonical lines first in the FIRST ## Session block), it left stale content in the section. Replace with a negative-lookahead per-line pattern that consumes all content from the heading up to (but not including) the next ## heading, producing a clean section with only the canonical bold-label lines. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
533b518553 |
fix(security-scan): update scanner self-exemption allowlists for renamed suite files
The three shell scanners exempt their own adversarial test fixtures by exact filename; the *.security.test.cjs renames broke those entries, so the PR diff scan flagged the scanners' own test payloads. Verified locally with all three scanners in --diff origin/next mode (0 findings) and the security suite (207/207). The .sh files were missed in the original reference sweep because the rename grep filtered to .cjs/.yml/.json/.md extensions. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8bb6784009 |
chore(ratchet): regenerate bug-* allowlist after rebase onto next (244 -> 257)
13 bug-* files landed upstream between the audit baseline and this branch's rebase; they predate the ratchet policy, so they are grandfathered. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
622e4be6d8 |
test(ratchet): ban new top-level bug-NNNN test files via identity allowlist
244 one-off bug-* files (~38% of the suite) are grandfathered in lint-regression-test-names.allowlist.json; new ones fail lint with fold-into-module guidance, and deletions force allowlist pruning so the baseline only shrinks. Wired into npm run lint:ci (new single entry point for every CI lint). Policy documented in docs/TESTING-SUITES.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
cd5db1f8db |
test(suites): seed security/slow/integration suites via measured retags
Renames (git mv) with all references updated (ci-test-scope RULES, windows-parity allowlist, test-file-count allowlist, docs in 6 locales): - 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step ran zero files since the suite taxonomy landed; it is now honest. - graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite; e2e gsd-tools spawns) — runs on full-matrix lanes and push to next. - installer-migration-install-integration -> *.integration.test.cjs (13s; an integration test by its own name). Coverage gate measured after retags: 88.55% lines (gate 70%). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a647053dcf |
ci(scope): narrow #494 invariant — changed tests join windows lane, not full matrix
full_matrix fired on 15/15 sampled PRs because any tests/** change forced it, costing ~25 runner-minutes each. A changed test file now always joins the scoped windows lane (covering the #482 OS-specific failure class per-file) and still runs on ubuntu 22/24 via targeted_tests; the residual macOS / windows-node-22 cross-product is covered on every push to next. Also narrows WINDOWS_HINTS from 6 substrings (102/633 files, a ~10-minute scoped lane) to windows/win32/shell/path — the dropped hints (workflow, install, hook) are either platform-independent lint tests or already covered by fullMatrix rules. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ed467cd8f2 |
feat(#942): tier→profile/cluster derivation + consistency gate (ADR-857 phase 4a) (#943)
Establish `tier` as the source of install-profile + cluster membership
(ADR-894 §4). The registry now derives two views: capabilityClusters
(capId → its skills) and profileMembership (capId → {tier, profiles}, where
profiles is the PROFILE_RANK suffix from the capability's tier). A consistency
gate cross-checks them against the hand-authored PROFILES/CLUSTERS: HARD (throws)
on a capId-matching-a-cluster-name with a different skill set; SOFT
(pending-reconciliation stderr warning, not serialized) when a capability skill
isn't yet in the closure-resolved hand-authored profile.
The SOFT gate loads the real skills manifest and resolves each profile's closure
once so transitively-included skills don't false-warn; warnings are de-duped to
one per (capability, skill); both derived views are scoped to skill-owning
capabilities; serialized with a global capId sort for determinism; reserved-name
guards at every write site; lazy requires of the built constants.
Behavior-preserving: install/surface untouched; derived views consumed by
nothing. Full generation rides along with the phase-6 migration.
Closes #942
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
dcb0d8a28d |
fix(#935): install changeset CLI so /gsd-update changelog preview works (#938)
- bin/install.js now copies scripts/changeset/ and scripts/lib/ into <configDir>/scripts/ so $GSD_DIR/scripts/changeset/cli.cjs resolves at runtime; aborts install with an explicit failure if the source directory is missing from the package. - gsd-core/workflows/update.md: corrected path from gsd-core/scripts/changeset/cli.cjs to scripts/changeset/cli.cjs; added an explicit [ ! -f ] guard so a missing CLI surfaces a clear message rather than silently swallowing the error; stderr captured via 2>&1 sentinel so node errors are visible in the preview output. - release.yml's changeset-CLI invocations (node scripts/changeset/cli.cjs) remain at the repo-root path and are unaffected by this change. Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6dbd895028 |
feat(#910): federated config merge in config-loader (ADR-857 phase 3b) (#914)
Build the federated config merge: each capability owns its config-key slice
(ADR-857 decision 3 / ADR-894), and loadConfig merges them defensively. The
registry now emits a full configSchema index ({key:{owner,type,default,
description}}, generator-validated); a new src/federated-config.cts resolves
federated keys defensively (skip central keys -> pending-migration warning,
skip malformed slices -> warning never throw, else type-checked user override
?? default, with nested dotted-path lookup and enum validation); and loadConfig
applies the overlay on every return path.
Wired as a provably-empty no-op channel: every UI-pilot key is still central,
so validKeys is empty and loadConfig returns byte-identical output on all paths
(identity return when the overlay is empty; shared CONFIG_DEFAULTS never
mutated). Registry-only; no key is cut over; nothing in the live loop changes.
Closes #910
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
d12809985e |
fix(#905): preserve STATE.md frontmatter scalars in syncStateFrontmatter (#907)
syncStateFrontmatter was silently dropping current_phase, current_phase_name, current_plan, and progress when body annotations were absent (e.g. after an agent or tool rewrote the body). These scalars can only be derived from body annotations — when absent, buildStateFrontmatter returns nothing for those keys. Added existingFm fallbacks mirroring the same pattern already applied in cmdStateJson, so every writeStateMd call preserves the existing values instead of stripping them. Also extended cmdStateJson with the same fallbacks for the three non-progress scalars. Adds regression test (7 cases) + lint-test-file-count allowlist entry. Closes #905 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
808df9110c |
fix(#892): parse checklist-style roadmap phases in validate/verify (#908)
buildRoadmapPhaseVariants() only matched heading-style phases (## Phase N:), silently skipping the supported checklist format (- [x] **Phase N: name**). This caused W007 false-positives for every on-disk phase dir when the project uses a checklist ROADMAP. Fix adds a second regex pass (mirroring the existing buildNotStartedPhaseVariants() approach). Also refactors the duplicate inline heading-only regex in cmdValidateConsistency() to delegate to buildRoadmapPhaseVariants() (DRY). Regression test in tests/bug-892-validate-checklist-roadmap-phases.test.cjs covers both paths. Closes #892 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
48cc27bd84 |
feat(#903): generate Loop Host Contract from workflow markers (ADR-857 phase 3a-impl-2) (#906)
Replace the inline LOOP_HOST_CONTRACT constant in the Capability Registry generator with a generated-from-workflows contract (ADR-894 §3). The contract is now derived from inert `<!-- gsd:loop-host ... -->` marker blocks in the five step workflows, emitted as the committed gsd-core/bin/lib/loop-host-contract.cjs, and required by gen-capability-registry.cjs — one source of truth, no drift. Drift guards in gen-loop-host-contract.cjs: per-step point ownership (each step must declare exactly its canonical loop points), multiple-block + duplicate-key hard errors, and a word-boundary agent-role cross-check. Contract content is byte-identical to the former constant; registry-only, nothing wired into the live loop. Closes #903 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ad754ca6cd |
feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl) (#902)
* feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl) First phase-3 code: the Capability Registry generation pipeline, built against the ADR-894 contract and NOT wired into the live loop (registry-only, per the staged-cutover design). - capabilities/ui/capability.json — the UI pilot (ADR-894 worked example): 2 skills, 2 agents, 3 config keys, 2 steps + 1 gate, with `when` activation. - scripts/gen-capability-registry.cjs — --write/--check generator. Hand-rolled schema validation (envelope + role-typed feature/runtime bodies + typed steps/contributions/gates + when + gate-check variants); cross-capability invariants (single ownership; requires exist+acyclic+tier-monotone; config-key ownership exclusive, collision-vs-central as a pending-migration warning); hooks validated against an inline LOOP_HOST_CONTRACT (3a-impl-2 swaps its source to the generated-from-workflows contract); GLOBAL point-ordered consumes-satisfiability; materialized byLoopPoint ordering (produces/consumes topo-sort); emits gsd-core/bin/lib/capability-registry.cjs (role-partitioned indexes + requiresClosure). Prototype-pollution guards (Object.create(null) + inline literal key checks) + fragment.path traversal guard. - gsd-core/bin/lib/capability-registry.cjs — committed generated artifact (mirrors package-identity.cjs: script-generated, tracked, linted, regenerated on build, drift-tested), wired via the new `gen:capability-registry` build step. - tests/capability-registry.test.cjs — 72 tests: schema + invariant + hook + ordering + adversarial (path-traversal, proto-pollution, runtime body, self-consume, cycles, collisions) + committed-file staleness guard. New-CLI-module checklist (INVENTORY 97->98, MANIFEST, ARCHITECTURE), CONTEXT.md "Capability Registry" un-[Planned]'d. Nothing wired into install/surface/loop. Gates: lint, code-review (4 bugs fixed), security-review (path-traversal + prototype-pollution fixed), codex adversarial-review ×3 (8+ findings fixed, confirmed sound), clean-build docker 13190 pass / 0 fail. Closes #896 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#896): CRLF-agnostic --check for capability-registry staleness (Windows) The committed capability-registry.cjs staleness guard failed on Windows CI only: git checks out the committed .cjs as CRLF (autocrlf, no .gitattributes) while the generator emits LF, so the byte-for-byte --check comparison mismatched. Normalize line endings on both sides of the --check comparison (no .gitattributes change, no change to the LF the generator writes). Adds a regression test simulating the Windows CRLF checkout. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0a11d361ca |
feat(#69): nest concrete skills under namespace routers at install (#883)
Emit the 6 gsd-ns-* routers as the only top-level skill bundles and nest the ~61 concrete skills under <router>/skills/<name>/SKILL.md on runtimes with confirmed non-recursive skill loaders (claude global, cline, qwen, hermes, augment, trae, antigravity). Router bodies rewrite their routing tables from Skill-tool dispatch to a Read skills/<name>/SKILL.md pattern. Recursive/unconfirmed loaders (cursor, codex, copilot, windsurf, codebuddy, opencode, kilo) keep the flat layout. Completes the v1.40 namespace architecture (#2792) so the eager skill listing drops to ~6 entries. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a480510f54 |
fix(#872): make roadmap-phase-fallback tests hermetic against ambient GSD env (#873)
extractCurrentMilestone reads STATE.md via planningDir(cwd), which is workstream-aware (honours GSD_PROJECT/GSD_WORKSTREAM). The fixtures write STATE.md to the plain <tmp>/.planning/STATE.md, so a developer shell inside a GSD workstream (GSD_WORKSTREAM exported) redirected the read to a non-existent workstream subdir -> version=null -> closed milestone sections leaked into the slice and assertions failed. Clean CI/Docker env never hit it. Not a Node-26 regex bug; reproduces identically on any Node with GSD_WORKSTREAM set. - scripts/run-tests.cjs: strip GSD_PROJECT/GSD_WORKSTREAM before spawning test children so the local runner env matches clean CI/Docker. - tests/roadmap-phase-fallback.test.cjs: file-level beforeEach/afterEach save/delete/restore of both vars; new regression test pinning workstream-aware STATE.md resolution. - tests/run-tests-harness.test.cjs: guard asserting the runner strips both vars (so removing the deletion fails clean CI). Closes #872 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
606363c416 |
chore(#846): remove unused PR-size labeler (size/S–XL) workflow (#848)
The PR Gate workflow's only job, size-check, labeled every PR with size/S–size/XL based on lines changed. Those labels aren't used in any review, triage, or automation flow, so the workflow was pure noise. - Delete .github/workflows/pr-gate.yml - Drop size-check from required status checks in both rulesets so PRs don't block forever on a check that never reports - Remove pr-gate.yml from INERT_WORKFLOWS (ci-test-scope.cjs) and the knownInert list (ci-test-scope.test.cjs) - Remove "PR Gate / size-check" from setup-branch-protection.sh Closes #846 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
543e51e71f |
fix(#844): sync runtime manifest versions on npm version bump (#845)
* fix(#844): sync runtime manifest versions on npm version bump The release workflow bumps package.json via `npm version` but never stamped the runtime-integration manifests that must track it (.claude-plugin/plugin.json #766, gemini-extension.json #775), so the first RC/finalize whose version diverged from the -dev stream failed the test suite before tagging/publishing. Add scripts/sync-manifest-versions.cjs (single VERSIONED_MANIFESTS registry) wired to a `version` npm lifecycle hook that stamps + stages the manifests on every `npm version` — covering all four release bump sites and local bumps with no workflow edits. A regression guard test fails if any repo JSON whose version matches package.json is not registered, forcing future version-bearing manifests into the sync. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#844): add changeset for manifest version sync fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1b6bd66f2c |
feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged) (#821)
* feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged) Wire three new context-tracking events (SubagentStop, Stop, PreCompact) to gsd-context-monitor so context-headroom warnings surface at model-stop and subagent-finalisation moments — not just on PostToolUse. Add a new FileChanged hook (gsd-config-reload.js) that hot-reloads .planning/config.json context mid-session when the user edits it, injecting a config summary as hookSpecificOutput.additionalContext. Updates plugin manifest hooks.json, managed-hooks-registry, installer-migration-report allowlist, and shell-command-projection cleanup tables. Tests: 21 new assertions in enh-770-claude-hook-events.test.cjs; enh-788 and issue-766 test suites updated. Closes #770 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#770): document newly-registered Claude Code lifecycle hooks Add a Hook coverage table to the Claude Code npm installer section of docs/how-to/install-on-your-runtime.md describing SubagentStop, Stop, PreCompact, and the new FileChanged (gsd-config-reload.js) hook that hot-reloads .planning/config.json mid-session. Also fixes the changeset frontmatter (adds type: Added + pr: 821) so docs-lint can consume the fragment. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): add gsd-config-reload.js to INVENTORY.md and regenerate manifest The feat commit added hooks/gsd-config-reload.js but did not bump the Hooks count in docs/INVENTORY.md (14→15) or add the new row, and did not regenerate docs/INVENTORY-MANIFEST.json. Both inventory-counts and inventory-manifest-sync tests failed across the full CI matrix. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): make lifecycle-hook tests deterministic on scoped runner Replace the shared hooks/dist/ ensemble setup (ensureHooksDist / teardownHooksDist) in the Claude hook tests with per-test isolation: pre-populate each test's own tmpDir/.claude/hooks/ with stub files and pass installerMigrations:[] to install() so the first-time-baseline migration does not remove the stubs before the copy step can run. Root cause: hooks/dist/ is gitignored and absent on a fresh npm ci. ensureHooksDist() created it and teardownHooksDist() deleted it, but with --test-concurrency=4 both test files ran concurrently as separate Node.js worker processes sharing the same filesystem. One file's afterEach teardown deleted hooks/dist/ while the other file's install() was copying from it, producing an ENOENT (reproduced 2/10 runs locally). The additional issue: even with pre-placed stubs surviving the copy race, the 000-first-time-baseline migration classified hooks/gsd-*.js as bundled-gsd-hook artifacts, auto-removed them, and the copy step never re-ran (hooks/dist/ absent) — leaving contextMonitorFile missing and all hook registrations silently skipped (the 'got: []' symptom). Fix: pre-populate targetDir/hooks/ per-test (isolated temp dir) AND pass installerMigrations:[] so the baseline scan is skipped. The Qwen suites already used this pattern correctly; the Claude suites are aligned to it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): ship gsd-config-reload.js by adding it to build-hooks HOOKS_TO_COPY The #770 feature added hooks/gsd-config-reload.js and registered it in MANAGED_HOOKS, the installer, INVENTORY, and the test EXPECTED_ALL_HOOKS list — but never added it to scripts/build-hooks.js HOOKS_TO_COPY. As a result the hook was never copied into hooks/dist/ during the build, so: - the hook would never ship to users (real production bug — the FileChanged config-reload feature was dead-on-arrival), and - install-minimal-hooks.test.cjs #1755 ("all expected hooks are copied from hooks/dist/ to target", ".js hooks are executable after copy", "manifest contains .js hook entries") failed on any environment with a clean checkout (no pre-existing hooks/dist/): coverage, full test macos-22/macos-24, test ubuntu-24. The failures were masked locally only by a stale hooks/dist/ left from a prior build (build-hooks copies into dist without clearing it). On CI's fresh `npm ci` there is no dist, so the omission surfaced. Fix: add 'gsd-config-reload.js' to HOOKS_TO_COPY so build-hooks stages it into hooks/dist/ alongside the other JS hooks. Verified by removing hooks/dist/ and rerunning the full suite green (0 fail). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): make config prototype-pollution beforeEach deterministic on scoped runner Root cause: the #663 and alert-#26 prototype-pollution describe blocks seeded .planning/config.json in beforeEach via a bare runGsdTools('config-ensure-section') whose result was discarded. That command runs in a spawned gsd-tools child; on the scoped CI lane (--test-concurrency=4, config.test.cjs scheduled alongside the heavy install/tarball suites that #770 pulled into the targeted set) the child can be transiently killed under resource pressure (non-zero exit, empty stderr — an OS-level kill, not an app error). The swallowed failure left config.json absent, so the first subtest's readConfig() threw ENOENT opening <tmp>/.planning/config.json. Only 1 of 4 subtests failed, confirming a per-invocation transient, not a deterministic miss; the full suite schedules files differently so config.test.cjs did not collide with those heavy neighbors → passed there. Fix: add ensureConfigReady(tmpDir) which retries config-ensure-section on ANY failure or missing file and throws a clear diagnostic if it still cannot create config.json, then use it in both prototype-pollution beforeEach blocks. Setup is now deterministic under load; the #663/alert-#26 security assertions are unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d32b8db635 |
feat(#836): no-LLM duplicate-issue detection + challenge + 1-day auto-close (#843)
* feat(#836): no-LLM duplicate-issue detection + challenge + 1-day auto-close Adds a deterministic (no-LLM) duplicate-issue governance lifecycle: - scripts/issue-dedupe.cjs: pure, unit-tested module (tokenize, Sørensen–Dice title similarity, scoreCandidates, renderChallengeComment, shouldClose) with fail-safe destructive-action guards. - duplicate-check.yml (issues:opened): scores new-issue title against open issues, posts a challenge comment + applies the pending `possible-duplicate` label on a clear match. - duplicate-sweep.yml (daily cron): closes possible-duplicate issues whose challenge comment is >24h old with no human reply and no 👎 veto; honors exempt labels; re-checks the label immediately before close (TOCTOU guard); strips the label on close to avoid reopen loops. - remove-duplicate-label.yml (issue_comment:created): clears the label and applies needs-maintainer-review when any human responds. - bug_report.yml / docs_issue.yml: add the required "I searched existing issues" preflight checkbox so all five forms force a pre-search attestation. - docs/agents/triage-labels.md: document the label + lifecycle. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#836): add changeset fragment for duplicate-issue detection Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
988024c1a3 |
fix(#837): three-dot diff in ci-test-scope so docs-only PRs skip the heavy matrix (#841)
CI test-scope detection diffed changed files with a two-dot `git diff --name-only base head`, where base is the moving tip of `next`. A PR branch cut from a slightly older `next` surfaced every product file `next` had gained since the merge-base, flipping product_changed/full_matrix and running the full Windows/macOS matrix + coverage on docs-only PRs. Switch to a three-dot `git diff --name-only base...head` (vs the merge-base), matching GitHub's PR "Files changed" semantics. Add a regression test that builds a stale-base topology, plus a guard test pinning `fetch-depth: 0` on the `changes` job (required for the merge-base to be locally available). Closes #837 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1040fb792e |
feat(#777): register Cursor-native hooks (.cursor/hooks.json) for session-start/post-tool parity (#831)
* feat(#777): register Cursor-native hooks (.cursor/hooks.json) for session-start/post-tool parity - Add gsd-cursor-session-start.js: injects STATE.md presence reminder (or new-project nudge) into Cursor sessions via the sessionStart hook event - Add gsd-cursor-post-tool.js: emits an additional_context nudge when write-class tool calls touch .planning/ files (postToolUse hook event) - Add 'cursor-hooks-json' installSurface to runtime-config-adapter-registry; writeCursorHooksJson/reconcileCursorHooksJson write the canonical { version: 1, hooks: { sessionStart, postToolUse } } JSON shape with idempotent reconciliation that preserves user-owned hook entries - Hook scripts are copied with /gsd:→gsd- rewrite so installed files contain no colon-form slash-command refs (bug-376 invariant) - 20 new tests in tests/cursor-hooks.test.cjs cover all reconciler paths, entry helpers, removal, runtime adapter surface, and hook script behavior - Update CONTEXT.md, ARCHITECTURE.md, installer-migrations.md, and 000-first-time-baseline.cts to include Cursor hooks.json surface Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#777): build hooks/dist on demand in bug-376 test for scoped/windows CI hooks/dist is gitignored and only produced by `npm run build:hooks`. The CI scoped (ubuntu-latest/node-22) and windows (windows-latest/node-24) test jobs do NOT run build:hooks before executing tests, so bug-376's prerequisite suite was failing with "hooks/dist not found" on both legs. Add ensureHooksDist() helper (mirrors bug-3357 pattern) that builds hooks/dist on demand in the before() hooks of prerequisite and Suite 3. Also add ensureHooksDist() call to Suite 3's before() so the snapshot step is also hermetic. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e04e757672 |
feat(#776): Gemini hook events (BeforeAgent/AfterAgent/BeforeModel) + hooksConfig.enabled check (#829)
* feat(#776): Gemini hook events (BeforeAgent/AfterAgent/BeforeModel) + hooksConfig.enabled check Register three new Gemini-CLI hook events on install: - BeforeAgent: fires before agent planning; wired to gsd-context-monitor - AfterAgent: fires after final response generation; wired to gsd-context-monitor - BeforeModel: fires before each LLM call (per-turn); wired to gsd-context-monitor All three reuse gsd-context-monitor.js (no new hook files). Uninstall cleanup loop extended to remove the new events. Non-array guard added for robustness against malformed settings. Also detect hooksConfig.enabled:false in Gemini settings and emit a clear warning — without this check, all registered hooks silently do nothing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: update changeset pr: 829 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#776): document Gemini hook events Add hook coverage table to the Gemini CLI section of install-on-your-runtime.md, covering the three new events (BeforeAgent/AfterAgent/BeforeModel wired to gsd-context-monitor) plus a callout for the hooksConfig.enabled:false silent failure mode detected by the installer. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |