Classify edge-probe/prohibition-probe predicate-generation as core
verification substrate (not an off-by-default Feature Capability),
settled before ADR-857 phase 6 (Migrate) freezes the core/plug-in line.
Prompted by @davesienkowski's boundary analysis on #857.
- ADR-857: amendment note + top-level decision carve-out + Loop Extension
Points exemption + new "Verification substrate vs. plug-in tier"
subsection + 2 Alternatives rows + phase-6 rollout exception + Open
Questions resolution.
- ADR-550: cross-reference pinning the core-substrate classification;
ties "exogenous grading" to Decision 4 (judgment-tier) and Decision 5
(test the contract, not the classifier).
- CONTEXT.md: glossary — Probe Core / Edge Probe reclassified (probe-core
on the contract side, adapters as the generator) + new "Verification
substrate (predicate boundary)" term.
Decomposition: the verifier<->predicate contract is core/non-toggleable;
the generator (probe adapters) is core-default but independently
versionable. Docs-only; no user-facing change.
Closes#1120
Refs #857
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* ci(#1104): sync next package.json version to the last published release
next rested on a -dev stream per ADR-660 (1.3.1-dev.0) — a never-published
placeholder that leaked to source/dev installs. Make every release type write
its exact published version back to next:
- finalize/hotfix (push main): auto-backmerge sets next's version to main's
released version, folded into the existing back-merge PR (+ pinned setup-node).
- rc (no main push): the rc job opens + admin-merges a sync PR after publish.
Shared, fail-closed scripts/sync-next-version.cjs stamps package.json + the
runtime manifests via the npm version hook and refuses any non-release version.
Amends ADR-660 (supersedes the -dev stream decision).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* ci(#1104): harden next-version sync against post-publish failure modes
Review hardening (Codex + code-review gates) on the #1104 sync helper and
its workflow callers:
- release.yml rc Sync step: continue-on-error so a post-publish sync hiccup
cannot fail an already-published release (npm immutability would block re-run).
- auto-backmerge.yml inline sync: set -euo pipefail + validate VERSION before
any shell use (closes a ${VERSION}-in-commit-message injection vector); git
add -u instead of -A.
- sync-next-version.cjs: reuse an existing open PR instead of failing gh pr
create on rc re-runs; regex-parse the PR number and fail loud; discriminate
the git diff --cached --quiet exit code (only status 1 == has-diff, else
rethrow); git add -u to avoid sweeping runner artifacts into next; tolerate
already-merged on admin merge.
- tests: +2 (existing-PR reuse, non-diff rethrow); 14/14 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(spec-phase): name tie-breaking/rounding mode in the edge-probe precision probe (#1102)
The precision edge probe fired on numeric-range requirements but its text never
named the most common rounding failure mode — tie-breaking / rounding mode
(half-up vs half-to-even, ceil/floor/truncate). In an A/B experiment the weak
tier (haiku) false-passed 50% of tie-class defects against the surfaced-
unresolved spec; naming the rule in the probe text drove honest abstention
83%→100% (false-pass 17%→0%, haiku 50%→0%) with no category regression
(precision still fires only on numeric-range).
Sharpen the single TAXONOMY probe string and update every rendering together so
the doc↔fixture↔machine contract (edge-probe-docs-fixtures.test.cjs) stays green:
src/edge-probe.cts, the edge-probe.md taxonomy table + 2 worked examples, the
01-round-half-even and 04-money-rounding fixtures, and the resolve-edge-coverage
how-to. Prose-only — no consumer keys on the probe text; SHAPE_CUES firing
unchanged; fully backward compatible.
* chore(changeset): Changed fragment for edge-probe precision probe text (#1108)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (e4f0910d), which replaced the line-based tier-max
size ratchet (#597) with the byte-based per-file baseline guard (#1074).
The edge-probe feature legitimately grows two workflows:
- spec-phase.md 15131 -> 23094 (+7963): Step 5.5 Edge-Completeness Probe
- plan-phase.md 93135 -> 94253 (+1118): covered/backstop edge lift into
must_haves.truths (the live <downstream_consumer> block)
Both remain under their tier hard caps (plan-phase XL 98304, ~4KB
headroom; spec-phase DEFAULT 40960). Growth is real inline workflow
content the feature requires at that step — not eager @-import proxy
gaming. Drops the obsolete line-based XL_BUDGET 93000->94000 bump
(superseded by the byte baseline) via rebase.
* chore(#550): reconcile INVENTORY headline counts after rebase onto next
Rebase onto current next dropped the prior reconcile commit (stale counts
refs 68 / CLI 107). Current next + the edge-probe additions yield:
- References (68 -> 69 shipped): + gsd-core/references/edge-probe.md
- CLI Modules (107 -> 109 shipped): + edge-probe.cjs + probe-core.cjs
Rows for all three already present; only the headline counts were stale.
Caught by tests/inventory-counts.test.cjs (CI ubuntu-24 leg).
Extends the #1074 scheme to tests/agent-size-budget.test.cjs, which used the
same assertTightCeiling tier ratchet but was still line-based (never rebased in
#717). Completes the migration — the last part of the #1074 epic.
- Rebase agent sizing from lines to LF-normalized bytes (#717/#683).
- Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests);
add a per-agent baseline (tests/agent-size-baseline.json) as the primary
anti-creep, and byte hard caps (XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB),
each above its tier high-water with real headroom. No separate new-file cap:
a net-new agent is DEFAULT-tier, already bounded by the DEFAULT cap.
- Keep the agent-classification tests verbatim.
- scripts/workflow-size.cjs: add generic measureMdFiles(dir, predicate)
(workflows + agents share one byte-measurement path); measureWorkflows now
delegates to it.
- scripts/update-size-baseline.cjs: one 'npm run size:baseline' now regenerates
BOTH the workflow and agent baselines (gsd-* filter for agents).
Rebased onto next after PR 2/3 (#1096) merged: replicate the
scripts/lib/workflow-size.cjs -> scripts/workflow-size.cjs move (PR 1/3) across
the generator and the agent test's require; regenerate the agent baseline
against current agents (a uniform +170 B preamble drift on all 33 since
authoring).
Addresses the #1097 review (trek-e):
- BLOCKER (acceptance criterion 5): document the agent contract in CONTEXT.md.
Adds RULESET.AGENT_SIZE_BUDGET (caps 57344/49152/24576, per-agent baseline,
dual size:baseline, shared measureMdFiles seam) and disambiguates it from the
separate DEFECT.AGENT-FILE-SIZE-CAP-BREACH 45K-CHAR guard (two units, two
purposes).
- Docs: now that #1096's docs/TESTING-SUITES.md "Workflow size budget" section
is in next, fold in the agent coverage here (renamed to "Workflow & agent
size budget"): agent caps + per-agent baseline + the how-to + reference rows,
and the disambiguation from the 45K-char guard.
- Minor (negative proof): add a boundary-fixture test exercising the hard-cap
comparison at cap-1/cap/cap+1 through the real lfByteCount path, so a future
threshold/operator edit can't silently neuter a cap.
- Nit: align the tier test name wording ("stays within") with the <= operator.
Negative proof on a real tracked agent (gsd-planner): baseline catches +10 B;
XL hard cap catches 57,516 > 57,344 with the baseline current.
Closes#1095 (PR 3/3 child); landing this completes the #1074 epic.
Completes the #1074 migration for workflows. The per-file baseline (PR 1) is
now the primary anti-creep guard, so the tier-max tighten-only ceilings are
retired here.
- Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests) and
the GRACE constant — per-file baseline already guards every file by name,
strictly stronger than max(tier).
- Convert the per-file tier test into 'SIZE: workflow tier hard caps': absolute
red lines (XL 96 KiB / LARGE 60 KiB / DEFAULT 40 KiB) that mean 'extract, do
not raise', each with real headroom above its high-water file.
- Add a 32 KiB (Codex project_doc_max_bytes) cap for net-new workflow files not
yet in the baseline and not explicitly tiered.
- Rewrite the header doc comment for the new two-guard model; drop the
assertTightCeiling import (now unused here; still exported + used by the agent
test until PR 3).
Addresses the three PR-2 items from the #1089 review:
- Finish the enumeration consolidation (Minor #1): the tier hard-cap and
new-file guards now read both their file list and byte sizes from
measureWorkflows()/listWorkflowStems(), removing the inline readdirSync +
per-file byteCount split-brain. Enumeration and measurement share one source.
- Rewrite CONTEXT.md RULESET.WORKFLOW_SIZE_BUDGET (Minor #2): stale caps
(XL<=90000/LARGE<=54000/DEFAULT<=38000) replaced with the baseline-first
model, current hard caps, the new-file anchor, and the size:baseline
remediation.
- Ship docs (contract-change requirement): a Diataxis how-to + reference for
the size guard in docs/TESTING-SUITES.md, beside the sibling regression-name
ratchet (the 'file grew, CI red -> npm run size:baseline, commit the one-line
diff, justify or extract lazily' workflow). CONTRIBUTING.md's workflow-tree
note is updated to the baseline model and points at the new guide.
Negative proof: each guard fails independently when violated (new-file cap at
33 KB; hard cap with baseline current; baseline on any per-file growth).
Refs #1074. Part 2 of 3.
Fresh windsurf/devin-desktop workspace installs write skills under .devin/ (legacy .windsurf/ recognized); global ~/.codeium/windsurf/ unchanged. Also threads real isGlobal through _applyRuntimeRewrites so global skill content references the codeium path. Closes#1085.
* feat(#1077): phase 5f-2 — drive the hookEvents dialect (PostToolUse/AfterTool) from the descriptor
postToolEvent (bin/install.js) and preToolEvent (applySettingsJsonHooks in
runtime-hooks-surface.cts) now select the event-name dialect from
registry.runtimes[id].runtime.hookEvents instead of the hardcoded
(runtime === 'gemini' || runtime === 'antigravity') check: hookEvents === 'gemini'
→ AfterTool/BeforeTool; else → PostToolUse/PreToolUse. hookEvents threaded into the
applySettingsJsonHooks opts bag. Equivalence-preserving (Codex-verified): hookEvents
'gemini' is exactly {gemini, antigravity}, 'claude' the rest; undefined → claude
dialect (matches the old else).
The per-event SET guards (isQwen||claude → SubagentStop/Stop/PreCompact;
runtime==='claude' → FileChanged; isGemini → Gemini agent-events) stay HARDCODED —
hookEvents (2-value) is too coarse to drive them (the event set differs within
hookEvents='claude'); per-event-set drive tracked in #1076.
Registry-parity test (enh-1077): asserts BOTH post-tool (AfterTool/PostToolUse) AND
pre-tool (BeforeTool/PreToolUse) dialects are a pure function of hookEvents, for
gemini/antigravity/claude/augment — non-vacuous (catches a broken hookEvents thread).
Closes#1077
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1077): build hooks/dist in before() so dialect-drive test passes in scoped CI
hooks/dist is gitignored and absent in scoped/windows CI jobs that do not
pre-run build:hooks. Without it, install() finds no hook files and all
AfterTool/BeforeTool/PostToolUse/PreToolUse event arrays come back empty,
failing every hook-presence assertion. Added an idempotent ensureHooksDist()
called in a top-level before() — mirrors the pattern from
bug-376-claude-js-hook-gsd-rewriter.test.cjs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1055): add installSurface/writesSharedSettings/permissionWriter/extendedHookEvents to runtime descriptors
Purely additive: four new fields on all 16 runtime capability.json descriptors,
validator extended with three new closed-vocab sets, registry regenerated.
Test fixtures (VALID_RUNTIME_CAP and makeRuntimeCap) updated to include the new
required fields so all 255 capability-registry tests continue to pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1076): drive per-event hook guards from extendedHookEvents descriptor
Replace hardcoded runtime-name checks (isQwen||runtime==='claude',
runtime==='claude', isGemini) in applySettingsJsonHooks with a single
descriptor-driven extendedEvents array derived from the new opts field.
Remove isQwen and isGemini derivations (no remaining uses after the three
guard blocks are migrated). Wire extendedHookEvents from the capability
registry in bin/install.js call site. Add behavioral regression test
(enh-1076-extended-hook-events-drive.test.cjs) confirming the drive is
purely descriptor-based and runtime-name-agnostic.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1055): drive resolveRuntimeConfigIntent from the runtime descriptor; retire hand-kept REGISTRY
- Rewrites src/runtime-config-adapter-registry.cts to require capability-registry.cjs
and read installSurface / writesSharedSettings / permissionWriter from
runtimes[id].runtime; deletes the hand-kept REGISTRY const (ADR-857 phase 5g drive 2).
- ALLOWED_CONFIG_RUNTIMES is now derived from descriptor entries that have installSurface.
- Fixes the configFormat parity gate in scripts/gen-capability-registry.cjs to read
installSurface directly from capMap descriptor bodies, breaking the require cycle
(adapter now requires the generated registry; gen-script must not require the adapter).
- Adds golden-master test tests/enh-1055-config-intent-descriptor-drive.test.cjs (41 tests)
pinning all 16 runtimes' return shapes and the TypeError-on-unknown contract.
- Updates scripts/lint-test-file-count.allowlist.json (config module, +1 file).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1076): make hooksSurface descriptor load-bearing for the settings-json hook-skip
- Adds hooksSurface?: string to ApplySettingsJsonHooksOpts and destructuring in
applySettingsJsonHooks (src/runtime-hooks-surface.cts).
- Replaces the hardcoded !isOpencode && !isKilo hook-skip guard with
hooksSurface !== 'none'; removes the now-unused isOpencode/isKilo derivations
(ADR-857 phase 5g drive 3).
- Passes hooksSurface from the runtime descriptor at the applySettingsJsonHooks
call site in bin/install.js using the established
_capabilityRegistry?.runtimes?.[runtime]?.runtime?.hooksSurface idiom.
- Extends tests/enh-1076-extended-hook-events-drive.test.cjs with two new suites
proving: (a) hooksSurface:'none' writes no hooks regardless of runtime name;
(b) hooksSurface:'settings-json' writes hooks even for 'opencode' (previously
hardcoded to skip).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: record installSurface/writesSharedSettings/permissionWriter/extendedHookEvents descriptor axes in ADR-1016
Add Decision 7a documenting the four axes added in the 5f-completion pass,
update axis counts from "six" to "twelve", note 5f-completion drives as done
in Decision 8's ladder, update Out of scope to reflect #1055/#1076 are done
and 5g (InstallPlan capstone) remains the only open phase.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1055): parity gate must fire on configFormat↔installSurface mismatch (read installSurface at the descriptor level)
The test fixture makeRuntimeCapMap did not include installSurface in the runtime object,
so the gate's typeof r.installSurface !== 'string' guard always skipped the entry and never threw.
Added installSurface as an optional third parameter to makeRuntimeCapMap and passed the correct
installSurface values ('settings-json' for claude, 'codex-toml' for codex) to the two THROWS tests.
The gate implementation already reads r.installSurface correctly from the descriptor level.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1076): add installSurface↔hooksSurface + extendedHookEvents↔hookEvents consistency gates with rejection tests
GATE A: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES map in validateRuntimeBody enforces that a
runtime's hooksSurface is valid for its installSurface (e.g. profile-marker-only only allows none,
codex-toml only allows codex-hooks-json). Derived from the 16 real runtime descriptors.
GATE B: validateRuntimeBody checks that if extendedHookEvents contains Gemini agent-events
(BeforeAgent/AfterAgent/BeforeModel), hookEvents must be 'gemini'; if it contains Claude-family
events (SubagentStop/Stop/PreCompact/FileChanged), hookEvents must be 'claude'.
Added 10 rejection tests in suite 27 covering each gate + each new field validator.
All 16 real runtimes satisfy both gates (verified before coding).
Exports: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES, VALID_INSTALL_SURFACES,
VALID_EXTENDED_HOOK_EVENTS, VALID_PERMISSION_WRITERS, validateRuntimeBody.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1076): strengthen hooksSurface-drive assertions; defensive hooksSurface fallback; drop vacuous dup
1. bin/install.js: add explicit literal fallback for hooksSurface when the committed
capability registry fails to load (opencode/kilo → 'none', all others → 'settings-json').
The descriptor is always the source of truth in normal operation.
2. enh-1076 Suite 7: change SessionStart assertion from key-presence (hasOwnProperty)
to at least-one-command (hasHooksFor), so the test fails if hooks are initialized-but-empty.
ensureHooksDist() in before() guarantees hook files exist.
3. enh-1055 Test 2: remove vacuous duplicate suite that re-asserted intent.runtime === row.runtime
already fully covered by Test 1's deepStrictEqual over all four fields.
4. capability-registry.test.cjs: fix stale comments in the grok-skip test that claimed the
parity gate uses the adapter registry; gate reads purely from the descriptor (installSurface
absent → typeof r.installSurface !== 'string' → soft-skip).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#1082): materialize the InstallPlan — collect install-level descriptor axes into resolveInstallPlan; route install()/finishInstall() through it (ADR-58/5g)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: record 5g InstallPlan materialization (ADR-58 Accepted, ADR-1016 phase-5 complete)
ADR-1016 Decision 8 step 7 updated to DONE: resolveInstallPlan(runtime) in
runtime-config-adapter-registry collects install-level descriptor axes into
the typed InstallPlan consumed by install()/finishInstall(). Out-of-scope
section updated: 5g capstone is complete, phase 5 fully materialized.
ADR-1016 line ~20 updated: InstallPlan IS now materialized (both halves).
ADR-58 Implementation note added (2026-06-11): realized in
runtime-config-adapter-registry (co-located with adapter-selection).
CONTEXT.md Runtime Config Adapter Registry entry extended to document
resolveInstallPlan and both-halves realization.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1082): update install drift guard to the resolveInstallPlan seam (5g)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
windsurf now also answers to devin-desktop (CLI --devin-desktop) for the Windsurf→Devin Desktop rebrand; all paths unchanged. The .devin/skills/ workspace migration is split to #1085. Closes#792.
Install-time injection of a disallowedTools deny-list into Claude copies of read-only verifier/auditor agents (mirrors the #443 effort injection); source agents stay runtime-neutral so Gemini/Qwen/Hermes are unaffected. Closes#767.
* refactor(#1059): phase 5f-1 — extract standalone hook-surface writers into a module
Extract the structurally-isolated hook-surface writer functions (cline/cursor/
copilot/codex-hooks-json + buildHookCommand + atomicWriteFileSync + node/bash runner
resolvers) out of bin/install.js into a new src/runtime-hooks-surface.cts module
(-693 LOC from install.js). Behavior-preserving: install.js requires + re-exports
the moved functions (module.exports surface preserved); no descriptor reads, no
behavior change. Prerequisite for the descriptor-drive (5f-2), mirroring ADR-3660's
artifactLayout extract→drive split.
Review caught + fixed 3 coupling issues: (HIGH) the module's atomicWriteFileSync
dropped the shared __atomicWrittenTmps temp-tracking → now ONE shared set (module
owns it, install.js aliases it, both cleanups read it); (drift) buildHookCommand
called resolveNodeRunner(opts) vs the original resolveNodeRunner() → reverted; two
source-grep tests (workflow-guard, sh-hook-paths) that scanned install.js for the
moved functions → made behavioral/non-vacuous; duplicate runner resolvers consolidated.
Settings-json hook block (~648 LOC) deferred to 5f-1b; descriptor-drive to 5f-2.
New-module checklist done. ~62 hook test files green; gsd-test 17592/0.
Closes#1059
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1059): reconcile CLI Modules count after merging next (uat-predicate)
Merging current next (which added uat-predicate.cjs via #247) alongside this
branch's runtime-hooks-surface.cjs put the filesystem at 107 bin/lib modules, but
both sides had independently bumped the INVENTORY headline 105→106 so the merge
under-counted. Set "CLI Modules (107 shipped)" + regenerate INVENTORY-MANIFEST.json.
Both module rows already present. Fixes inventory-counts.test.cjs (the only CI red).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results
Wire the already-reserved `phase.uat-passed` alias (subcommand `uat-passed`,
mutation:false) into the phase command router with a new markdown-aware
predicate that evaluates HUMAN-UAT results and reports pass only when every
required check passes. Post-SDK-retirement (ADR-0174/#174) successor to the
SDK-framed #70, with no SDK-specific API surface.
New pure module src/uat-predicate.cts:
- stripFalsePositiveContexts: frontmatter -> HTML-comment -> CommonMark-style
fenced-block state machine (tracks delimiter char+length) -> blockquote,
each a small composable step, so a `result: passed` inside frontmatter, a
fenced/~~~ block (incl. ~~~ nested in a ``` fence), a comment, or a
blockquote is never counted.
- parseUatResultItems: heading-block parser, column-0-anchored same-line
result; a heading with no result -> `missing` (fail-closed).
- analyzeMarkdown: unterminated fence/comment detection (malformed -> blocker).
- evaluateUatPassed: allowlist pass/verification semantics; passed = no
blockers && >=1 check && all passing; no_uat_artifacts discriminator (no
vacuous pass); optional requireVerification policy hook.
Thin cmdPhaseUatPassed handler in phase.cts; router closure rejects unknown
flags via makeInvalidArgs. Hardened across two Codex adversarial passes
(vacuous pass, dropped failing tests, permissive verification status,
nested-fence escape, cross-line result value, masked unterminated comment) —
all fixed fail-closed. New unit + CLI-integration suites incl. a fast-check
property test; docs, CONTEXT glossary, inventory, and changeset updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#247): backfill changeset PR number (#1063)
* fix(#247): indexOf paired-scan for unterminated-comment detection
CodeQL js/incomplete-multi-character-sanitization (high) flagged the
`raw.replace(/<!--[\s\S]*?-->/g,'')`-then-`.includes('<!--')` detection in
analyzeMarkdown as incomplete sanitization (a single regex pass can leave a
residual `<!--`). Replace it with a paired left-to-right indexOf scan that
contains no `.replace()` of the comment token — CodeQL-clean and strictly
more correct (a closed earlier comment can never mask a later unterminated
one). Behaviour unchanged; 98 predicate tests + scoped docker run green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Convert the planner's soft comment-text guideline into a plan-write-time
HARD GATE. When an acceptance criterion negative-greps for a literal
(`grep -c 'LIT' file == 0`) and that same literal appears verbatim in an
`<action>` body (JSDoc samples, head-comment references, "what NOT to do"
snippets), the executor's commit-time verify gate later fails on the
comment echo rather than a real regression — wasting cycles and training
the executor to distrust the gate.
`verify.plan-structure` (the `validate_plan` step) now scans for this:
- confidently-extracted (quoted) negative-grep literal echoed in an
<action> → error (valid:false), failing plan creation
- unquoted/ambiguous grep target → warning (fallback policy)
- `<!-- planner-discipline-allow: LIT -->` escape hatch skips a literal
- positive-count gates (`== N`) and `!= 0`/`>= 0` are out of scope
Adds the `<comment_text_discipline>` block to gsd-planner.md, the full
rules + allowlist example to planner-antipatterns.md, and regression
fixtures for downstream incidents 12-04, 11-04, 12-02 (plus a boundary
case proving positive-count gate 11-02 is not flagged).
Closes#429
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver
Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker,
gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks.
On a shim-only install — where gsd-tools.cjs exists under the runtime home but
gsd-tools is NOT on PATH — those calls fail with "command not found" and the
agent silently skips init/state/validate/commit ceremony, deferring to the
orchestrator or bypassing GSD bookkeeping entirely.
were never migrated, so it persisted on Claude Code and every other runtime that
consumes the source agents directly. Only gsd-phase-researcher.md carried a
resolver — and a stale, claude-only truncated one.
Fix (all runtimes):
- Inject the canonical multi-runtime gsd_run preamble (byte-equal to
_runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/
augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of
the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every
command-position bare gsd-tools to gsd_run.
- Upgrade gsd-phase-researcher.md's stale resolver to the canonical one.
- Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the
sync caught and corrected a mis-placed preamble during development).
- Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards
to agents/ so no runtime can silently regress.
Closes#1041
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1041): backfill changeset PR number to 1045
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer
The /gsd-code-review structural pre-pass invoked fallow with flags no published
fallow version accepts (--json, --profile, --stdin-files), so it failed on every
run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any
fallow version. Three compounding defects:
1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet,
--changed-since/--base for changed-files scoping (no file-list input), and
--max-crap for thresholds. There is no --profile or --stdin-files.
2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail),
0 when clean. The pre-pass treated any non-zero exit as a crash and discarded
the output — i.e. it threw away exactly the findings it exists to surface.
Success is now decided by whether a valid fallow JSON report was produced,
not by the exit code.
3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema
(unusedExports/duplicates/circularDependencies) fallow never shipped, and was
dead code (the workflow embedded raw JSON; its tests asserted the fictional
schema, one even calling a non-existent runFallowAudit and passing vacuously).
Fixes: align the invocation to fallow's documented agent-facing pattern; map the
profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase
runs via --changed-since with a repo-scope fallback; rewrite the normalizer to
fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies
+ duplication.clone_groups) and wire it into the workflow so the reviewer
receives normalized findings; replace the fictional-schema fixtures and tests
with real-schema ones and delete the vacuous runFallowAudit test.
Closes#1012
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1012): backfill changeset PR number to 1044
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1013): resolve worktree.baseRef from user/global settings cascade
cmdWorktreeBaseCheck resolved worktree.baseRef from the project checkout's
.claude/ only (settings.local.json then settings.json). A user/global
worktree.baseRef:"head" — the layer /config writes and the harness honors,
and the only sensible place for a machine-wide preference — was invisible. On a
phase lane (HEAD ahead of origin/HEAD, or no origin/HEAD symref) base-check
returned shouldDegrade:true and execute-phase forced sequential execution,
silently losing the parallel worktree execution the user configured.
CLAUDE_CONFIG_DIR (relocated user config dir) was also ignored.
resolveEffectiveBaseRef now accepts an optional user/global config dir and reads
its settings.json as a third, lowest-precedence layer (project local > project
shared > user/global). cmdWorktreeBaseCheck resolves it via
getGlobalConfigDir('claude'), which honors CLAUDE_CONFIG_DIR. The existing
project-level reads stay as higher-precedence overrides and the injectable
readFile seam is preserved, keeping the unit tests hermetic.
Closes#1013
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#1013): backfill changeset PR number to 1038
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).
Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).
Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).
Closes#1035
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Maintainer governance decision (waiving a standalone ADR for Qoder, PR #1021):
registering a runtime that reuses the existing profile-marker-only install
surface + a single skills kind is an enhancement governed by ADR-3660 via
addendum. Records the addendum-vs-new-ADR qualifying criteria, the normative
agent-frontmatter contract (name+description only — the sibling converter
shape), and Qoder as the first runtime logged under this path.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Record the resolution of #1022 (surfaced scoping the §5.6 ui-phase cutover):
a step is purely additive and never halts the host; host-blocking preconditions
are gates (blocking/onError:halt) — no hook-model change. Runtime/mode context
(auto/chain vs manual) self-gates in the skill, not via when (config-only).
§5.6 decomposes into the existing plan:pre step (ui-phase, self-gates on
frontend + pipeline) + a new plan:pre gate (frontend-and-no-UI-SPEC → halt,
when: workflow.ui_safety_gate) that blocks planning in manual mode — preserving
the "run /gsd:ui-phase first" UX (maintainer call: pipelines-only auto-fire).
The render-hooks dispatch template grows to handle gates, not just steps.
Recorded in ADR-894 (clarification) + CONTEXT.md
(RULESET.CAPABILITY.step-additive-gate-blocks). Unblocks the §5.6 cutover.
Closes#1022
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Realize ADR-857 Branch 8 (host-CLI support as role:runtime Capabilities) +
materialize ADR-58's InstallPlan. Closes the 6-axis descriptor vocabulary,
absorbs the hard-case runtimes as data, and stages the install migration as a
5a→5g ladder (4/6 axes already modularized; InstallPlan is the 5g capstone,
reachable by collection, not a big-bang rewrite).
Amended after a plugin-side design grill (rubber-duck + grill-with-docs vs
#956 MemPalace / #999 Impeccable):
- New term Connected Capability (CONTEXT.md) — a Capability whose integration
shape brings its own external process/service/state; orthogonal to authorship.
Named, tracked gap (vehicle #956); current schema does not express it.
- Narrowed the dogfood claim: the descriptor dogfoods the runtime interface
only, not the feature-plugin/Connected path.
- Structural "off means off" rule (CONTEXT.md RULESET.CAPABILITY.off-means-off):
the host derives shared outputs from active hooks; a hook adds/is-counted,
never mutates host source. Ratify in ADR-894; proven by spike #1018.
- Hand-waves resolved by code: sandboxTier real-but-thin; model-catalog
orthogonal (not a 7th axis); converters closed into a ConverterName enum +
added the kimi-agents artifact kind; configHome is pure read-only.
- Hook-firing path is unproven (render-hooks built, never consumed) → spike
#1018 must prove render-hooks→live-workflow execution before phase-5 build.
Design-only; no code. Status: Proposed.
Closes#1016
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(templates): add optional Business Context section to PROJECT.md template
Adds an optional `## Business Context` section (Customer, Revenue model,
Success metric, Strategy notes) between Core Value and Requirements, for
monetized or customer-facing projects. Optional by default — an HTML comment
tells non-business projects to delete it; capped at four one-line fields to
stay a constraint reference, not a business plan. The milestone evolution
review in complete-milestone.md checks it only when the section is present.
Refs #72
* chore(changeset): set pr number for #72 fragment
* test(#72): add source-text-is-the-product exemption marker
Addresses review Minor #1 on PR #756. The contract test reads the
PROJECT.md template and complete-milestone workflow .md files and
asserts on their content (the local/no-source-grep pattern). Those
.md files ARE the product surface, so this is a valid
source-text-is-the-product case. Add the explicit // allow-test-rule
marker per RULESET.TESTS.no-source-grep.exemption so intent is
audit-traceable before the rule promotes to error (#453).
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#973): add Edit to gsd-planner tools and forbid whole-file Write of ROADMAP.md
gsd-planner shipped Write but not Edit — the same writer-agent gap fixed for six
agents in #571/#581. Without Edit, an in-place ROADMAP update fell back to a
whole-file Write that truncated committed milestone history (292→16 lines in a
real incident).
Changes:
- agents/gsd-planner.md: add Edit to tools: frontmatter (adjacent to Write)
- agents/gsd-planner.md: update_roadmap step now directs Edit (scoped), with an
explicit blocking prohibition on whole-file Write of ROADMAP.md or any existing
curated .planning/ file
- agents/gsd-planner.md: Write contract section clarifies Write is authorized only
for net-new PLAN.md creation; existing files must use Edit
- tests/agent-frontmatter.test.cjs: extend SECTION_WRITER_AGENTS list (#581 test)
to cover gsd-planner — fails before fix, passes after
- .changeset/973-gsd-planner-edit-tool.md: Fixed changeset, pr:0
Closes#973
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore(#973): backfill changeset pr number (989)
* fix(#973): trim gsd-planner.md prose under agent size cap (keep Edit + scoped-Edit-for-ROADMAP rule)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4)
Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to
a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring
the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts
exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the
status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the
route fn. capabilities/intel/capability.json declares the intel command family;
commands-only (skills:[]), declares the existing intel.enabled gate (default
false — behavior unchanged).
Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing intel.test.cjs passes unchanged. Completes the first-party command-
family cutover sequence (graphify/audit/intel).
Closes#985
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI)
The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning`
expectation while the router builds it with path.join(cwd, '.planning') →
backslashes on Windows, so the planningDir-arg assertions (query/status/diff/
snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute
the expectation with path.join (cross-platform); production router unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs
case arms to a registry-dispatched Capability (commandFamilies mechanism, #961),
mirroring the graphify cutover (#972). New src/audit-command-router.cts exports
routeAuditUat/routeAuditOpen, each lazily requiring only its backing module
(uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads.
capabilities/audit/capability.json declares the two command families;
commands-only (skills:[]), no config gate, audit_review cluster untouched.
Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical;
existing audit regression tests pass unchanged.
Closes#981
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Migrate graphify into a Capability that owns its `graphify` command family,
dispatched via the registry (#961 mechanism) instead of a hardcoded case.
graphify is now an enable/disable plug-in.
- src/graphify-command-router.cts: routeGraphifyCommand (standard route*Command),
reproduces the removed case EXACTLY (query +--budget, status, diff, build,
hidden build snapshot, usage/unknown errors); injectable _graphify test seam.
- capabilities/graphify/capability.json: role feature, tier:full, skills:[graphify],
config:{graphify.enabled default false}, commands:[{family:graphify, module,
router:routeGraphifyCommand}].
- removed case 'graphify' from gsd-tools.cjs; graphify now flows
default -> dispatchCapabilityCommand -> commandFamilies.graphify -> router.
- regenerated registry (commandFamilies/bySkill/configSchema/profileMembership/
capabilityClusters for graphify); tier:full keeps 4c install/surface a no-op.
Equivalence-proven: 8 recording-mock unit tests assert the exact fn+args per
subcommand (budget, snapshot-vs-build); 9 subprocess tests assert distinguishing
output shapes; existing graphify tests pass unchanged through the new path.
Surfaced (not silently accepted): the pre-existing --budget-no-value NaN no-op
quirk, preserved for equivalence, filed separately.
Closes#972
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* test(#947): add regression tests and update stale Hermes assertions
- Add bug-947-hermes-gsd-prefix.test.cjs: 12 TDD tests covering fresh
install canonical layout, bare-stem migration, manifest key format,
and non-Hermes runtime isolation
- Update hermes-skills-migration.test.cjs: bare-stem → gsd-prefixed
path and name assertions (#947 canonical layout)
- Update install-nested-layout.test.cjs: Hermes NEST matrix prefix ''
→ 'gsd-'
- Update install-regressions.test.cjs: Defect #1 now seeds bare-stem
dirs (help/, quick/) and asserts gsd-help/ canonical output; use
real GSD stems so readGsdCommandNames() migration finds them
- Update install-runtime-artifacts.test.cjs: Hermes nested layout and
legacy migration assertions align with gsd- prefix
- Update install.test.cjs: Hermes install test uses gsd- prefixed paths
- Update runtime-artifact-layout.test.cjs: prefix '' → 'gsd-'
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#947): restore gsd- prefix on Hermes skills for canonical dispatch
Hermes skills were installing under bare-stem paths
(skills/gsd/<stem>/SKILL.md, name: <stem>) due to prefix: '' set in
ADR-3660 / #3664. This broke /gsd-<stem> dispatch and forced users to
invoke skills without the gsd- namespace prefix.
- src/runtime-artifact-layout.cts: change Hermes skillsKind prefix
from '' to 'gsd-'; skills now land at skills/gsd/gsd-<stem>/SKILL.md
with name: gsd-<stem>
- bin/install.js _runLegacyInstallMigrations: invert the #3664
migration — remove stale bare-stem dirs (using readGsdCommandNames()
to distinguish GSD-owned stems from user content), keep gsd-* dirs
which are now canonical
- bin/install.js _runLegacyUninstallCleanup: also remove bare-stem
dirs on uninstall for clean teardown
- bin/install.js uninstallRuntimeArtifacts: post-cleanup removes
DESCRIPTION.md and empty skills/gsd/ category dir on Hermes
- bin/install.js: remove skillListPrefix Hermes exception (now uses
shared 'gsd-' path)
- docs/adr/3660-runtime-artifact-layout-module.md: document #947
reversal of the bare-stem sub-decision
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* chore: add changeset for #947 fix (#955)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(#947): remove ALL pre-migration bare-stem Hermes skills on reinstall (adversarial review)
Replace readGsdCommandNames()-based bare-stem cleanup (which missed skills
not in the commands source tree, e.g. dev-preferences) with
_removeHermesBareStemDirs(), called AFTER the install loop when the exact
set of installed gsd-<stem>/ dirs is authoritative. For every gsd-<stem>/
written this run, the corresponding bare skills/gsd/<stem>/ is removed.
User-owned bare dirs with no gsd-<stem> counterpart are preserved.
Add two adversarial-review regression tests that FAIL on old code:
- bare skills/gsd/dev-preferences/ removed when gsd-dev-preferences/ installed
- user-owned bare dir with no gsd-<stem> counterpart is preserved (no over-deletion)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Review-pass fixes: lint:ci composes npm run lint (one eslint invocation
home); the scripts/ coverage floor moves to package.json
(test:coverage:scripts-floor) so both thresholds live together; the ratchet
test uses helpers.createTempDir; TESTING-SUITES.md clarifies what the
Windows scoped lane runs and why feat-*/enh-* files are exempt from the
bug-* ratchet.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Renames (git mv) with all references updated (ci-test-scope RULES,
windows-parity allowlist, test-file-count allowlist, docs in 6 locales):
- 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step
ran zero files since the suite taxonomy landed; it is now honest.
- graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite;
e2e gsd-tools spawns) — runs on full-matrix lanes and push to next.
- installer-migration-install-integration -> *.integration.test.cjs
(13s; an integration test by its own name).
Coverage gate measured after retags: 88.55% lines (gate 70%).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Design how a Capability contributes a gsd-tools CLI command family and how the
hardcoded 73-case runCommand switch opens to registry-driven dispatch. Realizes
ADR-857 decision 7's reserved commands/module field (deferred by ADR-894).
Grilled to its leanest form: the registry DISCOVERS a standard route*Command
(no rebuilt handler table, no new arg convention); dispatch sits in the default
case (collision structurally impossible, no shadowing gate needed); graphify is
the first real cutover (lowest blast radius, has skill+cluster+config gate,
full-only so 4c stays no-op), proven equivalent and serving as the phase-6
template.
Design-only; CONTEXT.md gains a "Capability Command Family [Planned]" entry.
Closes#959
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Combined PRD + ADR for wiring MemPalace (local-first AI memory) into the
GSD loop as an ADR-857 feature capability. Bidirectional sync, three
selectable memory-relationship modes (augment/kg_backend/replace),
loop-point recall+capture map, opt-in tier:full, MCP-primary/CLI-fallback.
Marked Pre-Proposal: the first-party-plugin proposal standard is not yet
established and ADR-857 phase-6 loop wiring is pending. First of a planned
series; PRD/ADR format is provisional pending PM-method evaluation.
Refs #956
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Add a read-side query composing the three toggle systems into one
per-capability view. resolveCapabilityState({registry, installedSkills,
surfacedSkills, config, cwd}) reports installed (skills ⊆ resolved install
profile), surfaced (skills ⊆ resolved surface), and per-hook active (no when →
active; non-empty-string when → resolved via _resolveActivationValue; empty/
non-string → inactive), with no forced composite verdict. cmdCapabilityState
does the I/O (resolveProfile + resolveSurface + loadConfig), resolves the
runtime config dir via the canonical getGlobalConfigDir (--config-dir override),
and surfaces resolution failures as warnings rather than a false installed='*'.
Routed as `gsd-tools capability state`.
Additive: install/surface/workflows untouched; consumed by nothing.
Closes#945
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#930): remove self-masking next dist-tag repoint from release finalize
The "Clean up next dist-tag" step silently failed under OIDC trusted
publishing (which can't write dist-tags) while unconditionally reporting
success via || true + an echo. It also violated the release model by
trying to repoint @next→stable; @next is managed exclusively by the rc
job's --tag next publish.
Closes#930
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: update ADR-660 to reflect removal of next dist-tag repoint
The finalize job no longer runs `npm dist-tag add … next`; update the
ADR-660 description of step 4 to match the new behavior — @next is
managed exclusively by the rc job.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
`context: fork` strips the `Agent` tool from a subagent's environment.
Spawning orchestrators (`/gsd-autonomous`, `/gsd-execute-phase`,
`/gsd-plan-phase`) depend on `Agent` to dispatch sub-agents; running
them forked silently disables the core capability they exist to provide
(#921). Remove `context: fork` from all three command frontmatter files.
`effort: xhigh` (introduced by #769) is preserved.
The `<runtime_compatibility>` Agent-availability guard added by #913 was
checking whether `Agent` was present *before* attempting the call. On
runtimes where the tool list is dynamically resolved this produced
false-negative aborts in sessions that have the tool (#922). Replace
the introspection-based pattern with an attempt-based gate: always
attempt the `Agent()` call; stop only if a real tool-unavailable error
is returned. This preserves #853's backgrounded-session close-off and
#913's intent of preventing inline role-collapse, while eliminating
false negatives.
Tests updated: enh-769-context-fork-effort.install.test.cjs asserts the
three orchestrators lack `context: fork` and that the converter still
passes the field through for non-orchestrator commands; plan-phase-drift-
guard.test.cjs adds four assertions for the attempt-based gate language;
workflow-size-budget unchanged (budgets not exceeded).
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Add the loop.render-hooks resolver: the first registry-consuming query.
`gsd-tools loop render-hooks <point>` validates the point against the
authoritative canonical 12, reads the registry's materialized byLoopPoint
hooks, filters them by activation, and emits a JSON envelope {point,
activeHooks, rendered} with ordered markdown.
Activation resolves each hook's `when` key by precedence: loadConfig value
(post-cutover federated) -> raw config.json workstream/root single-key lookup
(pre-cutover central override) -> registry configSchema default (so a
default:true capability hook is active out-of-the-box) -> inactive. Guarded
single-value reads only (no merged object built from untrusted keys).
Registry-only: no workflow calls the resolver yet (wiring is the phase-6
cutover). Completes the phase-3 trio.
Closes#918
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Build the federated config merge: each capability owns its config-key slice
(ADR-857 decision 3 / ADR-894), and loadConfig merges them defensively. The
registry now emits a full configSchema index ({key:{owner,type,default,
description}}, generator-validated); a new src/federated-config.cts resolves
federated keys defensively (skip central keys -> pending-migration warning,
skip malformed slices -> warning never throw, else type-checked user override
?? default, with nested dotted-path lookup and enum validation); and loadConfig
applies the overlay on every return path.
Wired as a provably-empty no-op channel: every UI-pilot key is still central,
so validKeys is empty and loadConfig returns byte-identical output on all paths
(identity return when the overlay is empty; shared CONFIG_DEFAULTS never
mutated). Registry-only; no key is cut over; nothing in the live loop changes.
Closes#910
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
`loadConfig` in configuration.cts was superseded by config-loader.cts
(ADR-857 phase 2e, #885). Exhaustive grep confirms no caller imports
loadConfig from configuration.cjs — all live callers use config-loader.cjs
or the core.cjs back-compat re-export. configuration.cts now provides only
the pure normalization and defaults primitives (normalizeLegacyKeys,
mergeDefaults, migrateOnDisk, CONFIG_DEFAULTS) that config-loader.cts
depends on. Updated CONTEXT.md and docs/INVENTORY.md to reflect the
narrowed module surface.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Replace the inline LOOP_HOST_CONTRACT constant in the Capability Registry
generator with a generated-from-workflows contract (ADR-894 §3). The contract
is now derived from inert `<!-- gsd:loop-host ... -->` marker blocks in the five
step workflows, emitted as the committed gsd-core/bin/lib/loop-host-contract.cjs,
and required by gen-capability-registry.cjs — one source of truth, no drift.
Drift guards in gen-loop-host-contract.cjs: per-step point ownership (each step
must declare exactly its canonical loop points), multiple-block + duplicate-key
hard errors, and a word-boundary agent-role cross-check. Contract content is
byte-identical to the former constant; registry-only, nothing wired into the
live loop.
Closes#903
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl)
First phase-3 code: the Capability Registry generation pipeline, built against
the ADR-894 contract and NOT wired into the live loop (registry-only, per the
staged-cutover design).
- capabilities/ui/capability.json — the UI pilot (ADR-894 worked example): 2
skills, 2 agents, 3 config keys, 2 steps + 1 gate, with `when` activation.
- scripts/gen-capability-registry.cjs — --write/--check generator. Hand-rolled
schema validation (envelope + role-typed feature/runtime bodies + typed
steps/contributions/gates + when + gate-check variants); cross-capability
invariants (single ownership; requires exist+acyclic+tier-monotone; config-key
ownership exclusive, collision-vs-central as a pending-migration warning);
hooks validated against an inline LOOP_HOST_CONTRACT (3a-impl-2 swaps its
source to the generated-from-workflows contract); GLOBAL point-ordered
consumes-satisfiability; materialized byLoopPoint ordering (produces/consumes
topo-sort); emits gsd-core/bin/lib/capability-registry.cjs (role-partitioned
indexes + requiresClosure). Prototype-pollution guards (Object.create(null) +
inline literal key checks) + fragment.path traversal guard.
- gsd-core/bin/lib/capability-registry.cjs — committed generated artifact
(mirrors package-identity.cjs: script-generated, tracked, linted, regenerated
on build, drift-tested), wired via the new `gen:capability-registry` build step.
- tests/capability-registry.test.cjs — 72 tests: schema + invariant + hook +
ordering + adversarial (path-traversal, proto-pollution, runtime body,
self-consume, cycles, collisions) + committed-file staleness guard.
New-CLI-module checklist (INVENTORY 97->98, MANIFEST, ARCHITECTURE), CONTEXT.md
"Capability Registry" un-[Planned]'d. Nothing wired into install/surface/loop.
Gates: lint, code-review (4 bugs fixed), security-review (path-traversal +
prototype-pollution fixed), codex adversarial-review ×3 (8+ findings fixed,
confirmed sound), clean-build docker 13190 pass / 0 fail.
Closes#896
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#896): CRLF-agnostic --check for capability-registry staleness (Windows)
The committed capability-registry.cjs staleness guard failed on Windows CI only:
git checks out the committed .cjs as CRLF (autocrlf, no .gitattributes) while the
generator emits LF, so the byte-for-byte --check comparison mismatched. Normalize
line endings on both sides of the --check comparison (no .gitattributes change,
no change to the LF the generator writes). Adds a regression test simulating the
Windows CRLF checkout.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#894): ADR-894 Capability declaration format + registry generation
ADR-857 rollout phase 3a (design-only). Resolve ADR-857's deferred open
question — the on-disk Capability declaration format — as a reviewable design
ADR before any generator code.
Specifies: the capabilities/<id>/capability.json folder layout (migration-staged
ownership — declarations reference existing stems until the phase-6 move); the
capability.json schema for role:feature (skills/agents/hooks/federated config/
loopHooks) and role:runtime (the six closed projection-primitive axes); the 12
named Loop Extension Points; the gen-capability-registry.cjs generator design
(validation + cross-capability invariants + --write/--check drift gate, mirroring
gen-inventory-manifest); the generated capability-registry.cjs shape (by-id /
by-skill / by-loop-point indexes + requires-closure); and a full worked example
(the UI capability: ui-phase + ui-review + agents + config + two loop hooks).
No code — design artifact only; the generator build, federated config loader
(3b), and loop seam (3c) implement against this contract.
Closes#894
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#894): amend ADR-894 with grilled capability declaration format
Stress-tested the declaration format before merge; the format changed
materially. Amendments:
- loopHooks[] -> three typed arrays (steps/contributions/gates), each with its
own shape (step: ref+produces/consumes; contribution: fragment+into agent-role;
gate: check+blocking).
- Add the Loop Host Contract (§3): each step publishes its points, agent roles,
and core artifacts so the generator validates hooks against reality, not
trusted strings.
- requires = capability ids only (host implicit); add tier-monotone invariant;
drop the requires:["plan"] error from the example.
- Config federation = atomic move: a migrated key leaves the central schema in
the same PR; presence in both is a collision (invariant stays).
- One registry, role-partitioned indexes (feature indexes vs runtimes index).
- Rework the UI worked example to the split-array shape (2 steps + 1 gate) +
a contribution illustration.
Adds a "Grilling amendments" section recording the six changes.
* docs(#894): amend ADR-894 with round-2 grilling (operational reality)
Second design-grill round, folded in before merge:
- Loop Host Contract is GENERATED from structured workflow markers
(<loop-point>/<agent-role>/<loop-artifact>) via gen-loop-host-contract.cjs —
it can't drift from the real workflows.
- Hook activation `when`: cheap deterministic config-level gating evaluated by
loop.render-hooks; deeper phase-context applicability self-gates inside the
dispatched skill (no phase-context vocabulary to keep honest).
- `tier` is the source of install-profile + cluster membership; profiles and
clusters are generated from tier + requires-closure (/gsd:surface operates on
capabilities) — collapses ADR-857's multiple toggle systems.
- Gate `check` = query | declarative-predicate | agentVerdict; agentVerdict is
forced advisory; only deterministic checks may block.
- byLoopPoint ordering is materialized in the registry; render-hooks filters the
active set + renders. Same-capability hooks degrade gracefully when an entry
step self-gates.
- Rollout: registry-only until atomic per-feature cutover (no double-execution
with still-inlined workflow features).
Updates the Grilling amendments / Consequences / Alternatives / Open questions
sections; reworks the UI example with `when`.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>