bd349aa88ea83e9408d7f05cfb2cd48449af2d47
1554 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0c0a8966ba |
fix(#1151): drive codex sandbox_mode emission from runtime descriptor sandboxTier axis (#1152)
* fix(#1151): drive codex sandbox_mode emission from runtime descriptor sandboxTier axis The sandboxTier runtime-capability axis was cosmetic: declared and validated on all 16 descriptors but read by nothing. The codex per-agent sandbox_mode line was emitted unconditionally from the hardcoded CODEX_AGENT_SANDBOX map, so the descriptor field drove no behaviour (ADR-857 audit finding F10; the "rides along in 5e/5g" promise in ADR-1016 §8 never landed). Make the axis load-bearing: - resolveInstallPlan projects sandboxTier as a 7th InstallPlan axis and fails loud (throws) on a missing/invalid value rather than coercing to 'none'. - installCodexConfig / generateCodexAgentToml gate sandbox_mode emission on sandboxTier !== 'none'. - The per-agent CODEX_AGENT_SANDBOX map is kept: it is GSD agent policy, not a runtime-descriptor property (different layer). Full removal of that map is tracked under #1138 (phase-6 descriptor-residue removal). For codex (sandboxTier === 'codex-agent-sandbox') the emitted TOML is byte-identical to before; for 'none' runtimes sandbox_mode is omitted. Adds leaf, projection, and installCodexConfig threading-seam regression tests; updates the enh-1082 InstallPlan golden master with sandboxTier for all 16 runtimes. Confirmed hypothesis: schema-first vocabulary closure outran consumer wiring, with no conformance gate to catch the orphaned axis. Closes #1151 Refs #857, #1138 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1151): stamp changeset with PR number 1152 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
607813f5d0 |
feat(#1136): consume resolved capability state (#1153)
* feat(#1136): consume resolved capability state * chore(#1136): add capability state changeset |
||
|
|
b431b1fab4 |
fix(#1140): implement state add-roadmap-evolution CJS handler (#1148)
* fix(#1140): implement state add-roadmap-evolution CJS handler `query state.add-roadmap-evolution` was unreachable: the CJS state router listed it in the `unsupported` map with a circular message ("...is SDK-only. Use: gsd-tools query state.add-roadmap-evolution ...") and no CJS handler existed after the SDK retirement (ADR-0174). Every `/gsd:phase insert` and `/gsd:phase --edit` run hit a dead end recording Roadmap Evolution. Re-implement `cmdStateAddRoadmapEvolution` in CJS (src/state.cts) and wire it into the state router; remove the now-stale `unsupported` entry. The handler appends a single-line bullet under `## Accumulated Context` → `### Roadmap Evolution` (creating the subsection/section if missing, deduping identical entries), scoping every lookup to the Accumulated Context body so a decoy heading in an unrelated section is never targeted, and flattening multiline notes to a single bullet. Section-boundary regexes mirror the sibling add-decision/add-blocker handlers and preserve following sections on CRLF input. Regression cases live in tests/state.test.cjs (per the no-new-bug-NNNN-files policy) and cover the literal issue repro plus the CLI/parser QA matrix (missing/empty/whitespace note, flag-shaped value, duplicate flags, hostile shell metacharacters, Unicode, decoy section, CRLF, missing STATE.md). Closes #1140 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1140): backfill changeset PR number (1148) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1b880fefd7 |
feat(#1137): migrate review verification hooks to capabilities (#1147)
* feat(#1137): migrate review verification hooks to capabilities * chore(#1147): add changeset |
||
|
|
44024aa535 |
fix(#1133): honor model_policy on the claude runtime (forward-port to next) (#1144)
Forward-port of the #1133 hotfix (commit 22f237a5, 1.4.5 hotfix line) onto next. The hotfix was authored against src/core.cts (v1.4.4); on next the resolver logic lives in src/model-resolver.cts (ADR-457 extraction, #888), so the patch is re-applied there rather than cherry-picked. resolveModelInternal step 2.5 now honors model_policy on the claude runtime: the policy-resolved full model ID is mapped back to a Claude Code agent alias via CLAUDE_POLICY_ID_TO_ALIAS (reverse of MODEL_ALIAS_MAP + claude-fable-5 -> fable). Bare aliases (opus/sonnet/haiku/fable) pass through; an ID with no Claude alias warns once to stderr (deduped by agentType::policyModel::tier) and falls back to the configured tier alias. Non-claude runtimes return full IDs verbatim (unchanged). resolveModelForTier is intentionally unchanged. The warn-dedupe cache lives in model-resolver.cts; core.cts composes the exported _resetRuntimeWarningCacheForTests to clear both that cache and the config-loader warning cache (config-loader cannot import model-resolver -- circular dependency). Ports the 6 #1133 tests (rewriting the old claude-no-op test that asserted the bug) plus one added test covering the MODEL_ALIAS_MAP reverse-map path. Forward-port of #1133 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4ab5c7b3f2 |
feat(#1135): migrate planning hooks to capabilities (#1141)
* feat(#1135): migrate planning hooks to capabilities * chore(#1135): add phase 6 planning capabilities changeset * fix(#1135): satisfy lint for agent hook rendering |
||
|
|
fd01e7a12e |
feat(#1132): complete contribution hook prerequisite
Closes #1132 |
||
|
|
1d90ad3c30 |
fix(commands): route nested sub_repos by longest prefix, not array order (#1130)
groupFilesBySubrepo selected the first sub_repos entry in array order
whose prefix matched a file, so a file under a more-specific nested
sub-repo (e.g. packages/core/widget.js with sub_repos
["packages", "packages/core"]) was mis-routed to the less-specific
parent ("packages").
Select the longest (most-specific) matching prefix within each
first-segment bucket instead, making routing independent of sub_repos
array order. String()-guard the length comparison so non-string entries
still never throw (preserves the #311 tolerance). Update the stale doc
comment that claimed first-match semantics.
Closes #391.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
eb051ea696 |
feat(#1123,#1124): enforce duplicate-producer invariant + fail-loud loadCentralConfigKeys in gen-capability-registry (#1131)
Closes #1123 Closes #1124 Refs #857 |
||
|
|
827011b865 |
fix(#1098): guard generate-claude-md against clobbering hand-crafted files; redirect default to .claude/CLAUDE.md (#1118)
/gsd-new-project wrote a repo-root CLAUDE.md full of broad project docs, overwriting/diluting a hand-crafted instruction file. --force was parsed but silently dropped, and nothing guarded an existing non-GSD file. - Guard: an existing instruction file with no `<!-- GSD:<section>-start` markers (hand-crafted) is left untouched; report action:"skipped". --force (now wired through CmdGenerateClaudeMdOptions) overwrites intentionally. The marker check uses /<!-- GSD:[a-z]+-start/ so a file merely documenting GSD syntax is safe. - Redirect: the Claude-family default output is now ./.claude/CLAUDE.md (a valid auto-loaded project-memory location) instead of repo-root ./CLAUDE.md. Aligned across the handler default, config-defaults.manifest.json, buildNewProjectConfig, the config template, new-project.md, and cmdGenerateClaudeProfile; advisory read-CLAUDE.md hints in plan-phase/quick/profile-user updated. Codex still writes AGENTS.md. Closes #1098 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fc37ae0c4f |
fix(#1115): capability-probe codex hook-trust bypass flag in /gsd:review; fail loud on empty output (#1122)
On codex-cli < 0.137.0 the review.md `codex exec` invocation passed --dangerously-bypass-hook-trust (added in 0.137) unconditionally and discarded stderr, so codex exited "unexpected argument" before reading the prompt and the empty output file was treated as a completed review — a silent degraded review. - Capability-probe the flag (`codex exec --help | grep`) and apply it via $CODEX_BYPASS_FLAG only when supported (works fine without it on older CLIs). - Capture codex stderr to a .err file instead of /dev/null, and replace an empty output with a diagnostic so a broken reviewer is surfaced (mirrors Cursor/OpenCode). - Update enh-773 enforcement test to require the capability gate + fail-loud guard instead of the unconditional flag. Closes #1115 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1c073c1c81 |
fix(#1107): progress consults verification.status before reporting a phase complete (#1116)
/gsd-progress derived phase completeness from plan/summary counts only and never consulted the verification.status query (the #651 seam), so a phase whose VERIFICATION.md ended human_needed or gaps_found was reported complete and routing skipped to the next phase. Add Step 1.7 (consult verification.status for the current phase) and routing rows that send gaps_found to plan-phase --gaps (Route V.gaps) and human_needed to verify-work (Route V.human) before the generic complete row. passed/missing/unknown still route as complete so unverified phases are not falsely blocked. Closes #1107 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fb8e7a3a65 |
fix(#1114): write-profile resolves the active runtime config home for USER-PROFILE.md (#1119)
Under Codex, `write-profile` wrote ~/.claude/gsd-core/USER-PROFILE.md while Codex discuss-phase advisor-mode (installed under ~/.codex) checked the Codex home and never found the profile, so advisor-mode silently stayed disabled. Resolve the default output via the runtime-aware getGlobalConfigDir (GSD_RUNTIME / config.runtime), mirroring cmdGenerateDevPreferences. Claude unchanged; --output still wins. Closes #1114 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e837bfc1de |
fix(#1101): record-session updates ## Session Continuity in place, no duplicate block (#1113)
The reported symptom (recorded:false yet STATE.md frontmatter still mutated) was already resolved on next by #944/#948. This fixes the residual: the DWIM auto-create recognised only the canonical `## Session` heading, so a bootstrap `## Session Continuity` section fell through to the append branch and produced a second `## Session` block. Insert only the missing canonical fields after the `## Session Continuity` heading — preserving the heading and any prose (e.g. "Next recommended action") — and teach the snapshot / frontmatter readers to recognise that heading (the optional ` Continuity` group still excludes `## Session Continuity Archive`, preserving #2444 scoping). Closes #1101 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4beceb02fe |
fix(#1103): preserve newline before Plans: header in roadmap.annotate-dependencies (#1111)
The match regex's `(?:^|\n)` anchor consumes a leading newline on mid-string matches (the common case where a `**Plans:** N plans` summary line precedes a bare `Plans:` block). The replacement dropped that newline, fusing the summary line onto the header — e.g. `**Plans:** 3 plansPlans:`. Re-emit the consumed newline so markdown line boundaries are preserved. Closes #1103 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7f1d49935c |
ci(#1104): keep next package.json in sync with the last published release (#1109)
* ci(#1104): sync next package.json version to the last published release next rested on a -dev stream per ADR-660 (1.3.1-dev.0) — a never-published placeholder that leaked to source/dev installs. Make every release type write its exact published version back to next: - finalize/hotfix (push main): auto-backmerge sets next's version to main's released version, folded into the existing back-merge PR (+ pinned setup-node). - rc (no main push): the rc job opens + admin-merges a sync PR after publish. Shared, fail-closed scripts/sync-next-version.cjs stamps package.json + the runtime manifests via the npm version hook and refuses any non-release version. Amends ADR-660 (supersedes the -dev stream decision). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * ci(#1104): harden next-version sync against post-publish failure modes Review hardening (Codex + code-review gates) on the #1104 sync helper and its workflow callers: - release.yml rc Sync step: continue-on-error so a post-publish sync hiccup cannot fail an already-published release (npm immutability would block re-run). - auto-backmerge.yml inline sync: set -euo pipefail + validate VERSION before any shell use (closes a ${VERSION}-in-commit-message injection vector); git add -u instead of -A. - sync-next-version.cjs: reuse an existing open PR instead of failing gh pr create on rc re-runs; regex-parse the PR number and fail loud; discriminate the git diff --cached --quiet exit code (only status 1 == has-diff, else rethrow); git add -u to avoid sweeping runner artifacts into next; tolerate already-merged on admin merge. - tests: +2 (existing-PR reuse, non-diff rethrow); 14/14 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3e836fef0d |
feat(spec-phase): spec-completeness edge-probe (#550) (#584)
* feat(spec-phase): spec-completeness edge-probe (#550) — relocated to gsd-core/
Rebased onto current next and relocated the whole feature from get-shit-done/ to
gsd-core/ per #615 (trek-e re-review #4, option 1). The artifact now builds to
gsd-core/bin/lib/edge-probe.cjs; all hard-coded path strings (tests, workflow
@-refs, run-tests.cjs sentinel, eslint ADR-457 ignore, .gitignore) updated.
Content conflicts in .gitignore / eslint.config.mjs / run-tests.cjs resolved.
Feature: Step 5.5 edge-completeness probe walks each SPEC requirement against a
closed 8-category edge taxonomy, proposes applicable candidate edges, and resolves
each (covered/dismissed/backstop/unresolved). covered/backstop criteria are lifted
by plan-phase into must_haves.truths, extending the goal-backward verifier's reach
to boundary edges no requirement was written for. Engine authored as strict TS
(src/edge-probe.cts, ADR-457), compiled to the gitignored gsd-core/bin/lib/edge-probe.cjs.
Folds in every prior review round on PR #584:
- RR-01..03: plan-phase resolves the phase *-SPEC.md and injects {SPEC_PATH} into the
planner; must_haves<->Edge-Coverage quality_gate; held-out planner-contract test.
- RR-04/11: Step 5.5 invokes the compiled engine at runtime (npm --prefix-pinned,
source-checkout-gated build fallback) instead of LLM re-derivation; the engine
capture is exit-checked and the report JSON-validated before use (fail closed).
- RR-05..10: all six fixtures embedded + count-equality; backstop/covered require a
resolution; Array.isArray(shapes); duplicate-resolution rejection; CLI JSON exit(2);
per-artifact build sentinel.
- Authored-shape validation: invalid (non-empty) shapes fail closed (VALID_SHAPES).
- Adversarial-review hardening: orphan/typo resolution rejection, requirement input
validation (id/text/shapes, duplicate id, non-array), zero-applicable guard.
Full suite 0 failures; npm run lint 0 errors; edge-probe suite 72/72.
* test(#550): RED — status×verification re-cut + probe-core engine specs
Re-cut the edge-probe resolution model onto two orthogonal axes per
ADR-550 Decision 7 (trek-e #644 comments 2026-06-03 14:36 + 14:44):
status: resolved | dismissed | unresolved (lifecycle, shared)
verification: explicit | backstop | null (only when resolved)
- tests/probe-core.test.cjs (new): behavioral specs for the generic engine
to be extracted — validateResolution(r, validators), validateRequirement,
analyzeCoverage(items, resolutions?, validators), byVerification rollup,
runProbeCli I/O scaffold (injected-io unit tests).
- tests/edge-probe.test.cjs: covered→{resolved,explicit}, backstop→
{resolved,backstop}; coverage gains byVerification.{explicit,backstop};
proposeEdges items gain verification:null.
- 6 fixtures re-genned + re-embedded in edge-probe.md; coverage.resolved
COUNT preserved on every fixture (closed set = resolved+dismissed; doc
line: 'adjacency=covered + ordering=dismissed' -> 2). edge-probe.md prose
rewritten to the two-axis model.
Fails as expected: probe-core.cjs has no source yet; edge-probe still
emits the old covered/backstop enum (27/61 edge specs red).
* feat(#550): extract probe-core seam + refactor edge-probe onto it (ADR-550 D7)
Extract the generic resolution model into src/probe-core.cts (the shared
seam the prohibition probe #644 is born on) and refactor edge-probe.cts
into its first adapter.
probe-core owns (probe-agnostic):
- the status×verification re-cut: status: resolved|dismissed|unresolved ×
verification: <probe-defined>|null
- validateResolution(r, validators) / validateRequirement (generic id+text)
- analyzeCoverage(items, resolutions?, validators) over ALREADY-PROPOSED
items[] (core never assumes propose is deterministic — edge resolves via
LLM, #644 proposes via LLM), with merge / dup-reject / orphan-reject
- byVerification rollup; coverage.resolved = closed set (resolved+dismissed),
count-preserved from the pre-re-cut engine
- runProbeCli I/O scaffold (injected io; one bin per probe)
- hybrid typing: generic params + injected runtime validators
{categories, verification, requiredFieldsByVerification} (ADR-550 #5)
edge-probe keeps ONLY the edge cluster: Shape/SHAPE_CUES/VALID_SHAPES/
classifyShape/TAXONOMY/applicableCategories/proposeEdges + EDGE_VALIDATORS
{explicit,backstop}; delegates merge/rollup/CLI to probe-core. Every shipped
#584 guarantee preserved (fail-closed shapes, orphan/dup rejection, input
validation, CLI exit 2). 104/104 edge+probe-core+docs+contract specs green.
* chore(#550): register probe-core.cjs artifact in ledgers + inventory
New gitignored build artifact gsd-core/bin/lib/probe-core.cjs (compiled
from src/probe-core.cts) needs registering in every artifact ledger:
- .gitignore + eslint.config.mjs ADR-457 ignore: lint the .cts source,
never the emitted .cjs.
- scripts/run-tests.cjs per-artifact build sentinel: build if probe-core.cjs
is missing on a clean checkout.
- docs/INVENTORY.md: CLI Modules 83 -> 84, new probe-core.cjs row, and the
edge-probe.cjs row updated to reflect it is now the first probe-core adapter.
- docs/INVENTORY-MANIFEST.json: regenerated (gen-inventory-manifest.cjs --write).
probe-core.test.cjs is a single test file (under the 2-file cap), so no
lint-test-file-count allowlist entry is needed.
* docs(adr-550): spec-phase probe pattern + prohibition contract [Accepted]
trek-e's final ADR-550 body, verbatim (open-gsd/gsd-core#644 comment
2026-06-03T15:23Z), Accepted by both maintainer and #550 author. Lands on
PR #584 alongside the probe-core extraction (Decision 7) it governs, so the
contract and its first implementation arrive together.
Decisions: probe packaging (3 layers); recall->precision protocol;
prohibition home = SPEC acceptance criteria + optional must_haves.prohibitions:
(truths untouched, no polarity); tiered verification test|judgment
(judgment = mode-dependent soft-gate-with-flags, never silent pass / never
hard-halt); CI tests the contract not the classifier; secure-phase ownership
seam; and Decision 7 — probe-core seam + status×verification re-cut (7a-7e),
which this PR implements.
* feat(#550): fail-closed probe-core across full status×verification + runProbeCli structural guard
Re-review #5 (trek-e) seam-hardening on the generic probe-core contract #644 inherits:
- validateResolution now enforces the 'verification is null unless resolved'
invariant for EVERY status (not just resolved): a dismissed/unresolved
resolution carrying a verification tier is rejected instead of merging verbatim.
- An unresolved resolution carrying a resolution/reason payload is rejected
(was silently dropped into the unresolved count).
- runProbeCli structurally validates the report an adapter returns before writing
it (was: any malformed object stringified as green output) — fails closed → exit 2.
- coverage.resolved kept count-preserved (closed set) per the blessed migration
contract; a new test locks that an all-dismissed run is NOT affirmatively covered
(byVerification is the honest gate).
Tests: probe-core 37/37; full edge-probe suite 113/113; full suite 1816/1816; lint 0.
* docs(adr-550): annotate Decision 5 #584/#644 scope + correct 7a coverage.resolved semantics
Re-review #5 (trek-e) clarity edits:
- Decision 5: annotate that only contract item (a) ships on #584 (the edge
adapter's parse+validate test); (b)–(d) are #644 scope, matching Consequences.
- Decision 7a: correct the 'coverage.resolved is preserved (status === resolved)'
parenthetical — the blessed/implemented semantics are count-preserved = the
CLOSED set (resolved + dismissed = applicable − unresolved), with byVerification
carrying the per-tier resolved-status breakdown. The old parenthetical
contradicted the shipped count.
* test(#550): cover runProbeCli structural-guard numeric-count branch
Second-pass coverage audit found the 'coverage object present but counts
non-numeric' branch of isValidReport (built probe-core.cjs:60-61) unexercised —
the {nope:true} malformed test fails earlier at the items[] check. Add a report
with well-formed items[] + a coverage object carrying non-numeric counts so the
numeric branch is hit. No source change; closes the line gap.
* fix(#550): reject edge requirement with missing/empty text when no shapes override (M2)
The edge adapter's `text` is the classification signal and a required field, but
core `validateRequirement` left it optional, so a `{ id }` requirement classified to
zero shapes -> zero edges -> was silently DROPPED from coverage with no signal -- the
exact fail-open this feature exists to eliminate. Reject missing/empty text unless an
authored `shapes` override (incl. `[]`) opts out of prose classification.
* fix(#550): validate verbatim items in analyzeCoverage shared seam (m1)
A proposed item with no matching author resolution is rolled up VERBATIM, but its own
status/fields were never validated -- an item carrying an out-of-enum status (the dropped
"covered") or `dismissed` without a reason would be counted closed. The edge adapter only
proposes `unresolved` items, but the prohibition adapter (#644) proposes LLM-generated
items that arrive populated. An Item is structurally a superset of a Resolution, so reuse
validateResolution to fail closed. ADR-550 Decision 5 hardens this shared seam.
* fix(#550): move edge-coverage lift instruction to runtime planner surface (M1)
templates/planner-subagent-prompt.md is loaded by nothing at runtime (no @-import in
agents/gsd-planner.md; plan-phase.md spawns the planner from its own inline
<planning_context>), so the precise covered/backstop -> must_haves.truths lift instruction
this PR added there never reached the planner -- and the RR-02 contract test asserted it in
that dead file, giving false green. Move the instruction into plan-phase.md's runtime
<downstream_consumer> block (where the rest of the wire already lives), revert the dead-template
edit, and retarget RR-02 to the loaded surface with a guard against re-orphaning.
* test(#550): lock machine<->SPEC vocabulary mapping against drift (m2)
The machine contract uses orthogonal status x verification; the SPEC table renders a flat
covered/dismissed/backstop/unresolved. The migration map (ADR-550 Decision 7a) was prose-only
with no test, so the layers could silently drift. The SPEC table is LLM-rendered (no JS
renderer to round-trip), so pin the canonical bijection as code AND ground it in every doc
surface that renders the vocabulary (ADR migration clause, spec.md legend, reference mapping
table) -- a rename or remap now fails the suite.
* docs(#550): add how-to for resolving edge-coverage findings (B1)
Feature shipped reference coverage (FEATURES.md, COMMANDS.md, references/edge-probe.md) but
no how-to -- reference-only does not satisfy the Diataxis docs standard for a user-facing
capability. Add a single-mode how-to (imperative, goal-directed) walking each resolution
state (specify/dismiss/backstop/defer), the soft gate, and --auto, with taxonomy/concepts
linked out to the reference. Register it in the docs/how-to index.
* docs(#550): add Probe Core + Edge Probe glossary entries to CONTEXT.md (N1)
trek-e re-review #7 N1 (Major): adding probe-core/edge-probe as src/*.cts-derived
seam modules (ADR-550 Decision 7) requires CONTEXT.md Domain-terms glossary entries
per the maintainer-enforced new-seam gate. Adds '### Probe Core Module' and
'### Edge Probe Module' with exports, generated source paths, and the ADR-550 seam
contract, placed beside the Research Module feature-seam entries.
* test(#550): add fast-check property suite for probe-core analyzeCoverage (N2)
trek-e re-review #7 N2 (RULESET.TESTS.property-based-testing): analyzeCoverage is a
transformation/rollup module, the class the property-testing predicate covers, and
fast-check is already a dependency with an established *.property.test.cjs pattern.
Adds 5 properties over the algebraic invariants: closed-set identity
(applicable === resolved + unresolved), byVerification sums ≤ resolved, per-tier
recount + resolved-status-only counting, rollup determinism, and stable orphan
rejection. 200 runs/seed 42 via helpers/fast-check-setup.cjs.
* test(#550): align allow-test-rule tokens to canonical runtime-contract-is-the-product (N3)
trek-e re-review #7 N3 (Nit): the // allow-test-rule: tokens (source-text-is-the-product,
docs-parity) differed from CONTEXT.md's canonical exemption category
'runtime-contract-is-the-product' (RULESET.TESTS.no-source-grep.exemption, CONTEXT.md:240).
All three tests assert deployed runtime-contract surfaces (spec-phase.md Step 5.5, the
plan-phase.md planner prompt, the rendered reference/SPEC/ADR vocabulary), so the canonical
category fits; each now carries a one-line justification per the ruleset format. Free-text
reason — lint behavior unchanged.
* test(#550): re-baseline plan-phase + spec-phase byte sizes for edge-probe
Rebased onto next (
|
||
|
|
e4f0910d62 |
test(#1074): agent-size-budget per-file baseline + line→byte rebase (PR 3/3) (#1097)
Extends the #1074 scheme to tests/agent-size-budget.test.cjs, which used the same assertTightCeiling tier ratchet but was still line-based (never rebased in #717). Completes the migration — the last part of the #1074 epic. - Rebase agent sizing from lines to LF-normalized bytes (#717/#683). - Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests); add a per-agent baseline (tests/agent-size-baseline.json) as the primary anti-creep, and byte hard caps (XL 56 KiB / LARGE 48 KiB / DEFAULT 24 KiB), each above its tier high-water with real headroom. No separate new-file cap: a net-new agent is DEFAULT-tier, already bounded by the DEFAULT cap. - Keep the agent-classification tests verbatim. - scripts/workflow-size.cjs: add generic measureMdFiles(dir, predicate) (workflows + agents share one byte-measurement path); measureWorkflows now delegates to it. - scripts/update-size-baseline.cjs: one 'npm run size:baseline' now regenerates BOTH the workflow and agent baselines (gsd-* filter for agents). Rebased onto next after PR 2/3 (#1096) merged: replicate the scripts/lib/workflow-size.cjs -> scripts/workflow-size.cjs move (PR 1/3) across the generator and the agent test's require; regenerate the agent baseline against current agents (a uniform +170 B preamble drift on all 33 since authoring). Addresses the #1097 review (trek-e): - BLOCKER (acceptance criterion 5): document the agent contract in CONTEXT.md. Adds RULESET.AGENT_SIZE_BUDGET (caps 57344/49152/24576, per-agent baseline, dual size:baseline, shared measureMdFiles seam) and disambiguates it from the separate DEFECT.AGENT-FILE-SIZE-CAP-BREACH 45K-CHAR guard (two units, two purposes). - Docs: now that #1096's docs/TESTING-SUITES.md "Workflow size budget" section is in next, fold in the agent coverage here (renamed to "Workflow & agent size budget"): agent caps + per-agent baseline + the how-to + reference rows, and the disambiguation from the 45K-char guard. - Minor (negative proof): add a boundary-fixture test exercising the hard-cap comparison at cap-1/cap/cap+1 through the real lfByteCount path, so a future threshold/operator edit can't silently neuter a cap. - Nit: align the tier test name wording ("stays within") with the <= operator. Negative proof on a real tracked agent (gsd-planner): baseline catches +10 B; XL hard cap catches 57,516 > 57,344 with the baseline current. Closes #1095 (PR 3/3 child); landing this completes the #1074 epic. |
||
|
|
b055ca14e2 |
test(#1074): swap workflow size enforcement to baseline + loose hard caps (PR 2/3) (#1096)
Completes the #1074 migration for workflows. The per-file baseline (PR 1) is now the primary anti-creep guard, so the tier-max tighten-only ceilings are retired here. - Delete the 'SIZE: tier anti-creep' describe (3 assertTightCeiling tests) and the GRACE constant — per-file baseline already guards every file by name, strictly stronger than max(tier). - Convert the per-file tier test into 'SIZE: workflow tier hard caps': absolute red lines (XL 96 KiB / LARGE 60 KiB / DEFAULT 40 KiB) that mean 'extract, do not raise', each with real headroom above its high-water file. - Add a 32 KiB (Codex project_doc_max_bytes) cap for net-new workflow files not yet in the baseline and not explicitly tiered. - Rewrite the header doc comment for the new two-guard model; drop the assertTightCeiling import (now unused here; still exported + used by the agent test until PR 3). Addresses the three PR-2 items from the #1089 review: - Finish the enumeration consolidation (Minor #1): the tier hard-cap and new-file guards now read both their file list and byte sizes from measureWorkflows()/listWorkflowStems(), removing the inline readdirSync + per-file byteCount split-brain. Enumeration and measurement share one source. - Rewrite CONTEXT.md RULESET.WORKFLOW_SIZE_BUDGET (Minor #2): stale caps (XL<=90000/LARGE<=54000/DEFAULT<=38000) replaced with the baseline-first model, current hard caps, the new-file anchor, and the size:baseline remediation. - Ship docs (contract-change requirement): a Diataxis how-to + reference for the size guard in docs/TESTING-SUITES.md, beside the sibling regression-name ratchet (the 'file grew, CI red -> npm run size:baseline, commit the one-line diff, justify or extract lazily' workflow). CONTRIBUTING.md's workflow-tree note is updated to the baseline model and points at the new guide. Negative proof: each guard fails independently when violated (new-file cap at 33 KB; hard cap with baseline current; baseline on any per-file growth). Refs #1074. Part 2 of 3. |
||
|
|
9e2ef2c94d |
fix(#1091): thread install scope into skill converters so local Antigravity/Copilot installs use workspace paths (#1092)
The skills layout wrapper (skillsKind) invoked every per-runtime skill converter as realConverter(content, skillName, runtime, cmdNames). The 3rd positional arg is overloaded: claude/kimi/cline converters read `runtime` there, but the copilot/antigravity converters read `isGlobal` there — so they received the truthy runtime string and always took the global path branch, leaking ~/.gemini/antigravity/ and ~/.copilot/ into local/workspace installs instead of .agent/ and .github/. Thread `scope` from resolveRuntimeArtifactLayout -> dispatchKindEntry -> skillsKind, derive isGlobal = scope === 'global', and pass it as a non-colliding 5th positional arg. Move isGlobal out of the colliding 3rd slot in the two converter signatures (3rd/4th become ignored _runtime/_cmdNames, matching the kimi convention). The fix flows through the shared ArtifactKind.stage closure, so applySurface re-apply inherits it via the same seam. Regression test exercises the wrapper seam (installRuntimeArtifacts at local scope) for both runtimes and asserts workspace paths, not global. Closes #1091 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
74d7bc8239 |
test(#1074): add additive per-file workflow size baseline guard (PR 1/3) (#1089)
* test(#1074): add additive per-file workflow size baseline guard (PR 1/3) Introduces a committed per-file size baseline scheme alongside (not replacing) the existing tier anti-creep tests. Green by construction — the baseline records current sizes, so both schemes pass side by side during migration. - scripts/lib/allowlist-ratchet.cjs: add assertFileBaseline (third pure helper, same injected-fail style) — per-file growth/shrink/add/remove diff vs baseline. - scripts/workflow-size.cjs: single source of truth for LF-normalized byte counting (#683) + workflow enumeration, shared by the guard and the generator so they can never measure differently. Lives in scripts/ root (NOT scripts/lib/) because it is dev/CI-only tooling — scripts/lib/ is bundled into the installed runtime, scripts/ root is not, so this keeps it out of the shipped payload. - scripts/update-size-baseline.cjs + npm run size:baseline: regenerate the snapshot (sorted keys, trailing newline, idempotent). - tests/workflow-size-baseline.json: generated snapshot (88 workflows). - tests/workflow-size-budget.test.cjs: import the shared counter (drops the duplicated local byteCount) and add the per-file baseline describe block. - Tests for the helper, the shared module, and the generator (incl. round-trip and fault-injection cases). Refs #1074. Part 1 of 3; PR 2 swaps enforcement, PR 3 covers the agent test. * test(#1074): regenerate workflow baseline after Update-branch merge with next The 'Update branch' merge (652a916b) pulled in next's update.md change (#1090) without regenerating the snapshot, leaving the per-file baseline stale by one file. Re-ran `npm run size:baseline` so the committed baseline matches the merged workflow files. Refs #1074. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
b4a7eabaae |
feat(#1085): migrate windsurf workspace skills to .devin/ + fix global content refs (#1093)
Fresh windsurf/devin-desktop workspace installs write skills under .devin/ (legacy .windsurf/ recognized); global ~/.codeium/windsurf/ unchanged. Also threads real isGlobal through _applyRuntimeRewrites so global skill content references the codeium path. Closes #1085. |
||
|
|
77a671ec53 |
feat(#791): migrate antigravity workspace base dir .agent → .agents (#1090)
Fresh antigravity workspace installs write under the canonical .agents/ (plural) base; legacy .agent/ stays recognized (dual-read). Global ~/.gemini/antigravity/ path unchanged. Closes #791. |
||
|
|
94872662e9 |
feat(#1082): complete phase 5 — descriptor-drive all install surfaces + materialize the InstallPlan — ADR-857/1016/58 (#1080)
* feat(#1077): phase 5f-2 — drive the hookEvents dialect (PostToolUse/AfterTool) from the descriptor postToolEvent (bin/install.js) and preToolEvent (applySettingsJsonHooks in runtime-hooks-surface.cts) now select the event-name dialect from registry.runtimes[id].runtime.hookEvents instead of the hardcoded (runtime === 'gemini' || runtime === 'antigravity') check: hookEvents === 'gemini' → AfterTool/BeforeTool; else → PostToolUse/PreToolUse. hookEvents threaded into the applySettingsJsonHooks opts bag. Equivalence-preserving (Codex-verified): hookEvents 'gemini' is exactly {gemini, antigravity}, 'claude' the rest; undefined → claude dialect (matches the old else). The per-event SET guards (isQwen||claude → SubagentStop/Stop/PreCompact; runtime==='claude' → FileChanged; isGemini → Gemini agent-events) stay HARDCODED — hookEvents (2-value) is too coarse to drive them (the event set differs within hookEvents='claude'); per-event-set drive tracked in #1076. Registry-parity test (enh-1077): asserts BOTH post-tool (AfterTool/PostToolUse) AND pre-tool (BeforeTool/PreToolUse) dialects are a pure function of hookEvents, for gemini/antigravity/claude/augment — non-vacuous (catches a broken hookEvents thread). Closes #1077 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1077): build hooks/dist in before() so dialect-drive test passes in scoped CI hooks/dist is gitignored and absent in scoped/windows CI jobs that do not pre-run build:hooks. Without it, install() finds no hook files and all AfterTool/BeforeTool/PostToolUse/PreToolUse event arrays come back empty, failing every hook-presence assertion. Added an idempotent ensureHooksDist() called in a top-level before() — mirrors the pattern from bug-376-claude-js-hook-gsd-rewriter.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1055): add installSurface/writesSharedSettings/permissionWriter/extendedHookEvents to runtime descriptors Purely additive: four new fields on all 16 runtime capability.json descriptors, validator extended with three new closed-vocab sets, registry regenerated. Test fixtures (VALID_RUNTIME_CAP and makeRuntimeCap) updated to include the new required fields so all 255 capability-registry tests continue to pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): drive per-event hook guards from extendedHookEvents descriptor Replace hardcoded runtime-name checks (isQwen||runtime==='claude', runtime==='claude', isGemini) in applySettingsJsonHooks with a single descriptor-driven extendedEvents array derived from the new opts field. Remove isQwen and isGemini derivations (no remaining uses after the three guard blocks are migrated). Wire extendedHookEvents from the capability registry in bin/install.js call site. Add behavioral regression test (enh-1076-extended-hook-events-drive.test.cjs) confirming the drive is purely descriptor-based and runtime-name-agnostic. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1055): drive resolveRuntimeConfigIntent from the runtime descriptor; retire hand-kept REGISTRY - Rewrites src/runtime-config-adapter-registry.cts to require capability-registry.cjs and read installSurface / writesSharedSettings / permissionWriter from runtimes[id].runtime; deletes the hand-kept REGISTRY const (ADR-857 phase 5g drive 2). - ALLOWED_CONFIG_RUNTIMES is now derived from descriptor entries that have installSurface. - Fixes the configFormat parity gate in scripts/gen-capability-registry.cjs to read installSurface directly from capMap descriptor bodies, breaking the require cycle (adapter now requires the generated registry; gen-script must not require the adapter). - Adds golden-master test tests/enh-1055-config-intent-descriptor-drive.test.cjs (41 tests) pinning all 16 runtimes' return shapes and the TypeError-on-unknown contract. - Updates scripts/lint-test-file-count.allowlist.json (config module, +1 file). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): make hooksSurface descriptor load-bearing for the settings-json hook-skip - Adds hooksSurface?: string to ApplySettingsJsonHooksOpts and destructuring in applySettingsJsonHooks (src/runtime-hooks-surface.cts). - Replaces the hardcoded !isOpencode && !isKilo hook-skip guard with hooksSurface !== 'none'; removes the now-unused isOpencode/isKilo derivations (ADR-857 phase 5g drive 3). - Passes hooksSurface from the runtime descriptor at the applySettingsJsonHooks call site in bin/install.js using the established _capabilityRegistry?.runtimes?.[runtime]?.runtime?.hooksSurface idiom. - Extends tests/enh-1076-extended-hook-events-drive.test.cjs with two new suites proving: (a) hooksSurface:'none' writes no hooks regardless of runtime name; (b) hooksSurface:'settings-json' writes hooks even for 'opencode' (previously hardcoded to skip). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: record installSurface/writesSharedSettings/permissionWriter/extendedHookEvents descriptor axes in ADR-1016 Add Decision 7a documenting the four axes added in the 5f-completion pass, update axis counts from "six" to "twelve", note 5f-completion drives as done in Decision 8's ladder, update Out of scope to reflect #1055/#1076 are done and 5g (InstallPlan capstone) remains the only open phase. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1055): parity gate must fire on configFormat↔installSurface mismatch (read installSurface at the descriptor level) The test fixture makeRuntimeCapMap did not include installSurface in the runtime object, so the gate's typeof r.installSurface !== 'string' guard always skipped the entry and never threw. Added installSurface as an optional third parameter to makeRuntimeCapMap and passed the correct installSurface values ('settings-json' for claude, 'codex-toml' for codex) to the two THROWS tests. The gate implementation already reads r.installSurface correctly from the descriptor level. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): add installSurface↔hooksSurface + extendedHookEvents↔hookEvents consistency gates with rejection tests GATE A: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES map in validateRuntimeBody enforces that a runtime's hooksSurface is valid for its installSurface (e.g. profile-marker-only only allows none, codex-toml only allows codex-hooks-json). Derived from the 16 real runtime descriptors. GATE B: validateRuntimeBody checks that if extendedHookEvents contains Gemini agent-events (BeforeAgent/AfterAgent/BeforeModel), hookEvents must be 'gemini'; if it contains Claude-family events (SubagentStop/Stop/PreCompact/FileChanged), hookEvents must be 'claude'. Added 10 rejection tests in suite 27 covering each gate + each new field validator. All 16 real runtimes satisfy both gates (verified before coding). Exports: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES, VALID_INSTALL_SURFACES, VALID_EXTENDED_HOOK_EVENTS, VALID_PERMISSION_WRITERS, validateRuntimeBody. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1076): strengthen hooksSurface-drive assertions; defensive hooksSurface fallback; drop vacuous dup 1. bin/install.js: add explicit literal fallback for hooksSurface when the committed capability registry fails to load (opencode/kilo → 'none', all others → 'settings-json'). The descriptor is always the source of truth in normal operation. 2. enh-1076 Suite 7: change SessionStart assertion from key-presence (hasOwnProperty) to at least-one-command (hasHooksFor), so the test fails if hooks are initialized-but-empty. ensureHooksDist() in before() guarantees hook files exist. 3. enh-1055 Test 2: remove vacuous duplicate suite that re-asserted intent.runtime === row.runtime already fully covered by Test 1's deepStrictEqual over all four fields. 4. capability-registry.test.cjs: fix stale comments in the grok-skip test that claimed the parity gate uses the adapter registry; gate reads purely from the descriptor (installSurface absent → typeof r.installSurface !== 'string' → soft-skip). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1082): materialize the InstallPlan — collect install-level descriptor axes into resolveInstallPlan; route install()/finishInstall() through it (ADR-58/5g) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: record 5g InstallPlan materialization (ADR-58 Accepted, ADR-1016 phase-5 complete) ADR-1016 Decision 8 step 7 updated to DONE: resolveInstallPlan(runtime) in runtime-config-adapter-registry collects install-level descriptor axes into the typed InstallPlan consumed by install()/finishInstall(). Out-of-scope section updated: 5g capstone is complete, phase 5 fully materialized. ADR-1016 line ~20 updated: InstallPlan IS now materialized (both halves). ADR-58 Implementation note added (2026-06-11): realized in runtime-config-adapter-registry (co-located with adapter-selection). CONTEXT.md Runtime Config Adapter Registry entry extended to document resolveInstallPlan and both-halves realization. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1082): update install drift guard to the resolveInstallPlan seam (5g) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
93c5ecd645 |
feat(#792): add devin-desktop runtime alias for windsurf (#1086)
windsurf now also answers to devin-desktop (CLI --devin-desktop) for the Windsurf→Devin Desktop rebrand; all paths unchanged. The .devin/skills/ workspace migration is split to #1085. Closes #792. |
||
|
|
e644819d59 |
feat(#767): inject disallowedTools deny-list into Claude read-only agents (#1081)
Install-time injection of a disallowedTools deny-list into Claude copies of read-only verifier/auditor agents (mirrors the #443 effort injection); source agents stay runtime-neutral so Gemini/Qwen/Hermes are unaffected. Closes #767. |
||
|
|
837991d5a2 |
chore(#1087): Windows test-portability lint + n/no-path-concat + LF normalization (#1088)
* chore(#1087): add Windows test-portability lint + DEFECT.WINDOWS-TEST-PORTABILITY Local gsd-test runs Mac+Linux only, so Windows-only test failures (Git Bash msys2 not honoring Node's chmod exec bit for PATH-executing extension-less scripts; `/` vs `\` path assertions) surface for the first time in CI's windows lanes — repeatedly (most recently PR #1084's #381 fix). Add scripts/lint-windows-test-portability.cjs: a high-signal, low-false- positive tripwire that flags any tests/**/*.test.cjs combining a chmod exec bit with a `sh -c`/`bash -c` invocation and no process.platform guard, unless annotated `// windows-portability-ok: <reason>`. Wired into lint:ci and runnable as `npm run lint:windows-test-portability`. Clean against all 721 current test files (zero pre-existing violations). Document the broader anti-pattern as DEFECT.WINDOWS-TEST-PORTABILITY in CONTEXT.md (.symptom/.examples/.detect/.fix-forward/.prevention): local gsd-test cannot substitute for the CI windows lane — watch it green before declaring a PR done. tests/lint-windows-test-portability.test.cjs covers the scanContent matrix. Closes #1087 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1087): enable n/no-path-concat + add .gitattributes/.editorconfig (LF) Two cross-platform "free wins" complementing the Windows test-portability lint: - Enable eslint-plugin-n's n/no-path-concat (already-installed plugin) as 'error' — flags string path concatenation (the / vs \ separator class). Zero existing violations, so it's a clean ratchet, not a refactor. - Add .gitattributes (* text=auto eol=lf + binary exemptions) and .editorconfig (LF, UTF-8, final newline, trim whitespace) to normalize line endings and kill the CRLF-only-fails-on-Windows class at the source. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7c07fce70f |
fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes (#1084)
* fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes On runtimes that execute each fenced bash block in a separate shell process (e.g. Claude Code — documented behavior: each Bash command is a separate process; inline shell functions and exported vars do not persist between calls), the once-per-file gsd_run() function was undefined in every block after the preamble block, and the call was swallowed by `2>/dev/null || echo "{}"` into silent empty state. Fix (budget-neutral session-level resolution): - Ship gsd-core/bin/gsd_run, a POSIX sh wrapper that symlink-resolves its own location and execs the co-located gsd-tools.cjs. Exposed on PATH via the npm `bin` field (global installs) and shipped to local installs via the recursive gsd-core/ copy. - The per-file launcher preamble now appends `export PATH='<bindir>':"$PATH"` to the file named by $CLAUDE_ENV_FILE (Claude Code's documented env-persistence mechanism) so later fresh-shell blocks resolve gsd_run from PATH. Guarded as a strict no-op when CLAUDE_ENV_FILE is unset; the inline gsd_run() definition remains the fallback for all other runtimes. The single-quoted dir neutralizes shell metacharacters at source time. - Propagated via scripts/sync-runtime-launcher.cjs to all launcher-using files. - XL workflow byte budget 93000 -> 93200 (the ~130B clause pushes plan-phase.md to 93135; legitimate content growth, ratchet-up per #717). Regression tests (I)/(J) in runtime-launcher-parity.test.cjs cover wrapper delegation and end-to-end PATH persistence (sourcing the env file with a space-bearing install path). Known limitation: an install path containing a literal single-quote yields a malformed env-file line and falls back to the status quo (no regression); rare on sanitized home directories. Closes #381 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#381): add changeset for gsd_run fresh-shell reachability fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#381): scope test (J) bare-PATH execution to POSIX (Windows Git Bash exec bit) Windows Git Bash (msys2) does not honor Node's chmod exec bit for PATH-executing extension-less scripts, so the bare `gsd_run` command lookup failed there even though the env-file PATH persistence was correct. The env-file content assertions (the fix's actual cross-platform logic) still run on every platform; only the final source-and-execute sub-step is gated to non-win32. Global installs on Windows are covered by npm's generated bin shim. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2023f47c64 |
fix(#1058): cross-reference install manifest in validate agents to catch .md/.toml pair drift (#1079)
* fix(#1058): cross-reference install manifest in validate agents to catch pair drift `validate agents` considered an agent installed if ANY supported file format was present on disk. The Codex installer generates a per-agent PAIR (agents/gsd-*.md AND agents/gsd-*.toml) and records both in gsd-file-manifest.json, so a partial generated install — one side of the pair missing — was reported as healthy (agents_found: true, missing: []), masking an incomplete Codex agent install. checkAgentsInstalled now cross-references the install manifest beside the agents dir (path.dirname(agentsDir)/gsd-file-manifest.json): for each expected agent, if the manifest tracks files for it and any tracked file is absent on disk, the agent is reported in a new `incomplete` list and agents_found becomes false. The check no-ops when no manifest is present (preserves bundled/claude behavior) and is scoped to expected agents so retired/stale manifest entries cannot false-flag. Regression cases added to tests/agent-install-validation.test.cjs cover the drift case, the complete-pair (no false positive), and the no-manifest no-op. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1058): add changeset for validate-agents manifest pair-drift fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9855ea3f39 |
fix(#1070): recognize "Complete ✓" terminal status in planned-phase transition (#1078)
* fix(#1070): recognize "Complete ✓" terminal status in planned-phase transition LLM phase executors (e.g. OpenCode) may write `Status: Complete ✓` into STATE.md when finishing a phase. `state planned-phase` then failed to advance the Status field on both the frontmatter `**Status:**` line and the Current Position `Status:` line, because `Complete ✓` matched neither KNOWN_TEMPLATE_DEFAULTS['Status'] nor any KNOWN_STATUS_PATTERNS entry — so it was preserved as an executor-authored value and the state machine stayed stuck on the prior phase. Add a narrow, fully-anchored pattern `/^Complete\s*[✓✔✅☑]?\s*$/i` to KNOWN_STATUS_PATTERNS so a bare `Complete` / `Complete ✓` terminal marker yields to the next phase's `Ready to execute`. Both Status writers consult this array, so the single addition fixes both paths. Caveat-bearing statuses like `Complete but needs manual QA` are not matched and remain preserved. Regression cases added to tests/state.test.cjs (planned-phase block) exercising both code paths plus the preservation guarantee. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1070): add changeset for planned-phase Complete-status fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f11462e58e |
refactor(#1067): phase 5f-1b — extract the settings-json hook block (applySettingsJsonHooks) — ADR-857/1016 (#1075)
* refactor(#1067): promote referencesHook to runtime-hooks-surface module scope referencesHook was declared as a local function inside install() but also called in finishInstall() (module scope), meaning JS hoisting was the only thing making it work from finishInstall. Move it to src/runtime-hooks-surface.cts, export it, and have both call sites in install.js use the module's copy. This is the prerequisite for COMMIT 2 (applySettingsJsonHooks extraction) per ADR-857 phase 5f-1b. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(#1067): extract applySettingsJsonHooks to runtime-hooks-surface module ADR-857 phase 5f-1b: move the ~457-line settings.json hook-registration block from install() into applySettingsJsonHooks(settings, opts) in src/runtime-hooks-surface.cts. install() replaces the block with a single call. Behavior-preserving: all runtime=== guards, isGemini/isQwen/isOpencode/isKilo derivations, postToolEvent/preToolEvent dialect branches, idempotency checks, fs.existsSync guards, and console.log/warn messages are verbatim. Opts bag: 13 fields — runtime, isGlobal, targetDir, postToolEvent, updateCheckCommand, contextMonitorCommand, promptGuardCommand, readGuardCommand, readInjectionScannerCommand, configReloadCommand, hookOpts, localCmd, localShellCmd. preToolEvent computed inside (from runtime). workflowGuardCommand / worktreePathGuardCommand / validateCommitCommand / graphifyUpdateCommand / sessionStateCommand / phaseBoundaryCommand / contextMonitorFile also computed inside. settings.hooks-only mutations confirmed. 5 source-scan tests updated to read runtime-hooks-surface.cts alongside install.js (concatenated), so structural regression guards remain valid at their new canonical location. install.js: 12700 → 12254 lines (−446). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
58bfae9d6a |
refactor(#1059): phase 5f-1 — extract standalone hook-surface writers into a module — ADR-857/1016 (#1064)
* refactor(#1059): phase 5f-1 — extract standalone hook-surface writers into a module Extract the structurally-isolated hook-surface writer functions (cline/cursor/ copilot/codex-hooks-json + buildHookCommand + atomicWriteFileSync + node/bash runner resolvers) out of bin/install.js into a new src/runtime-hooks-surface.cts module (-693 LOC from install.js). Behavior-preserving: install.js requires + re-exports the moved functions (module.exports surface preserved); no descriptor reads, no behavior change. Prerequisite for the descriptor-drive (5f-2), mirroring ADR-3660's artifactLayout extract→drive split. Review caught + fixed 3 coupling issues: (HIGH) the module's atomicWriteFileSync dropped the shared __atomicWrittenTmps temp-tracking → now ONE shared set (module owns it, install.js aliases it, both cleanups read it); (drift) buildHookCommand called resolveNodeRunner(opts) vs the original resolveNodeRunner() → reverted; two source-grep tests (workflow-guard, sh-hook-paths) that scanned install.js for the moved functions → made behavioral/non-vacuous; duplicate runner resolvers consolidated. Settings-json hook block (~648 LOC) deferred to 5f-1b; descriptor-drive to 5f-2. New-module checklist done. ~62 hook test files green; gsd-test 17592/0. Closes #1059 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1059): reconcile CLI Modules count after merging next (uat-predicate) Merging current next (which added uat-predicate.cjs via #247) alongside this branch's runtime-hooks-surface.cjs put the filesystem at 107 bin/lib modules, but both sides had independently bumped the INVENTORY headline 105→106 so the merge under-counted. Set "CLI Modules (107 shipped)" + regenerate INVENTORY-MANIFEST.json. Both module rows already present. Fixes inventory-counts.test.cjs (the only CI red). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9223f2f4c8 |
feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results (#1063)
* feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results Wire the already-reserved `phase.uat-passed` alias (subcommand `uat-passed`, mutation:false) into the phase command router with a new markdown-aware predicate that evaluates HUMAN-UAT results and reports pass only when every required check passes. Post-SDK-retirement (ADR-0174/#174) successor to the SDK-framed #70, with no SDK-specific API surface. New pure module src/uat-predicate.cts: - stripFalsePositiveContexts: frontmatter -> HTML-comment -> CommonMark-style fenced-block state machine (tracks delimiter char+length) -> blockquote, each a small composable step, so a `result: passed` inside frontmatter, a fenced/~~~ block (incl. ~~~ nested in a ``` fence), a comment, or a blockquote is never counted. - parseUatResultItems: heading-block parser, column-0-anchored same-line result; a heading with no result -> `missing` (fail-closed). - analyzeMarkdown: unterminated fence/comment detection (malformed -> blocker). - evaluateUatPassed: allowlist pass/verification semantics; passed = no blockers && >=1 check && all passing; no_uat_artifacts discriminator (no vacuous pass); optional requireVerification policy hook. Thin cmdPhaseUatPassed handler in phase.cts; router closure rejects unknown flags via makeInvalidArgs. Hardened across two Codex adversarial passes (vacuous pass, dropped failing tests, permissive verification status, nested-fence escape, cross-line result value, masked unterminated comment) — all fixed fail-closed. New unit + CLI-integration suites incl. a fast-check property test; docs, CONTEXT glossary, inventory, and changeset updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#247): backfill changeset PR number (#1063) * fix(#247): indexOf paired-scan for unterminated-comment detection CodeQL js/incomplete-multi-character-sanitization (high) flagged the `raw.replace(/<!--[\s\S]*?-->/g,'')`-then-`.includes('<!--')` detection in analyzeMarkdown as incomplete sanitization (a single regex pass can leave a residual `<!--`). Replace it with a paired left-to-right indexOf scan that contains no `.replace()` of the comment token — CodeQL-clean and strictly more correct (a closed earlier comment can never mask a later unterminated one). Behaviour unchanged; 98 predicate tests + scoped docker run green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8813ee5f95 |
feat(#429): HARD GATE on negative-grep literals echoed in plan <action> bodies (#1062)
Convert the planner's soft comment-text guideline into a plan-write-time HARD GATE. When an acceptance criterion negative-greps for a literal (`grep -c 'LIT' file == 0`) and that same literal appears verbatim in an `<action>` body (JSDoc samples, head-comment references, "what NOT to do" snippets), the executor's commit-time verify gate later fails on the comment echo rather than a real regression — wasting cycles and training the executor to distrust the gate. `verify.plan-structure` (the `validate_plan` step) now scans for this: - confidently-extracted (quoted) negative-grep literal echoed in an <action> → error (valid:false), failing plan creation - unquoted/ambiguous grep target → warning (fallback policy) - `<!-- planner-discipline-allow: LIT -->` escape hatch skips a literal - positive-count gates (`== N`) and `!= 0`/`>= 0` are out of scope Adds the `<comment_text_discipline>` block to gsd-planner.md, the full rules + allowlist example to planner-antipatterns.md, and regression fixtures for downstream incidents 12-04, 11-04, 12-02 (plus a boundary case proving positive-count gate 11-02 is not flagged). Closes #429 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0cc37a94c2 |
feat(#1056): phase 5e — close ConverterName enum + configFormat↔installSurface parity guard (#1057)
Two gen-time validation tightenings (validation-only; serialized registry content
unchanged; bin/install.js + adapter + descriptors untouched):
Part B: validateArtifactKindEntry now requires artifactLayout[].converter ∈
VALID_CONVERTER_NAMES (15 names, all exported by install.js) ∪ {null} — a typo'd
converter fails at gen time instead of silently → installExports[name]===undefined
at install time.
Part A: a HARD buildRegistry parity gate asserts each runtime descriptor's
configFormat agrees with the adapter registry's installSurface via a fixed mapping
(cursor-hooks-json/profile-marker-only→none, codex-toml→toml, copilot-instructions→
markdown, cline-rules→markdown-dir, settings-json→settings-json) — keeps configFormat
from drifting; prerequisite-validation for the deferred full drive (#1055).
The full config-writing drive (retire resolveRuntimeConfigIntent) is deferred to
#1055: configFormat is lossy vs installSurface (cursor vs profile-marker both → none;
opencode/kilo permissionWriter has no descriptor field) → needs schema extension.
Closes #1056
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
4698b3e349 |
fix(#1051): force-exit + per-chunk timeout for the windows full-test lane; close leaked test handles (#1054)
The `full test (windows-latest, 22)` job intermittently got CANCELLED at its 20m wall-clock cap with no failed test step — a false-negative gate (recurrence of #869). Root cause: a unit test leaves an open event-loop handle, so the chunk's `node --test` child hangs ~150s on Windows after its last test prints; two such stalls push the already-~13m job past 20m. Fix (defense in depth): - run-tests.cjs: pass --test-force-exit (Node >=22; engines requires >=22.0.0) so the runner exits once all tests finish regardless of lingering handles — the durable backstop. Account for the flag in the argv-length ceiling. - run-tests.cjs: add a per-chunk execFileSync timeout (default 600000ms, env RUN_TESTS_CHUNK_TIMEOUT_MS) that fails loudly with a diagnostic naming the chunk's files, so a hung chunk can never silently eat the job budget. - perf-316 test: terminate both Worker threads on all paths (afterEach + finally) so they cannot outlive the test. - locking-bugs test: kill spawned children in a finally that wraps the whole spawn -> waitFor -> barrier-release -> Promise.all sequence, so a barrier timeout no longer leaks live child processes. - Refresh the stale synckit comment (synckit/SDK bridge was removed). Regression tests in run-tests-harness: a hung chunk hits the per-chunk timeout and fails with a clear message; force-exit lets a chunk with a leaked handle exit cleanly. Closes #1051 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
58ed55683e |
feat(#1049): phase 5d — drive artifactLayout from the runtime descriptor (retire the 128-LOC switch) (#1053)
resolveRuntimeArtifactLayout now builds Layout from
registry.runtimes[id].runtime.artifactLayout[scope] — a loop dispatching each
ArtifactKind through the SAME 5 builders (commandsKind/agentsKind/skillsKind/
convertedCommandsKind/kimiAgentsKind, unchanged) by (kind, converter, nesting) —
replacing the hardcoded switch(runtime). Equivalence-preserving for all 16 runtimes
× {global, local} (Codex-verified, no divergence). -43 LOC; bin/install.js + the
converters + the install loop untouched. getInstallExports()[converterName]
resolution, configDir threading, scope default, unknown-runtime guard all preserved.
Driving the local scope surfaced a 5a gap: the old switch had no scope branch for 13
runtimes (cursor/gemini/codex/copilot/antigravity/windsurf/augment/trae/qwen/hermes/
codebuddy/opencode/kilo) → local == global for them, but 5a authored local:[].
Backfilled local=global for those 13 (descriptor-faithful; a fall-through shim would
wrongly give cline/kimi local=global). claude/cline/kimi scope-gating untouched.
validateArtifactKindEntry tightened: destSubpath/prefix/nesting/converter required
(ConverterName enum still open — 5e). New 39-case deep-equal golden equivalence test.
Closes #1049
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
cec7e704d6 |
feat(#435): expand workflow-policy linter to full Cartesian matrix cross-product (#1050)
`expandRunsOn` enumerated a multi-axis `strategy.matrix` one key at a time,
producing partial realization contexts. A true `os × shell` matrix therefore
left `${{ matrix.shell }}` unresolvable against any `{ os: ... }`-only context,
firing spurious `UNRESOLVABLE_MATRIX` violations and leaving shell-pinning
coverage incomplete on Cartesian jobs.
Enumerate the full GitHub Actions cross-product of all base-list matrix keys
(every `matrix.<k>` array, excluding the `include`/`exclude` control keys) via
a named `cartesianProduct` helper. Each realization's context now carries a
value for every matrix key, so `${{ matrix.<key> }}` resolves per realization.
Single-axis matrices keep byte-for-byte identical output; only multi-axis
matrices change shape. The `include` and `exclude` blocks are unchanged
(full tuple-aware exclude is a documented out-of-scope follow-up).
Tests: updated the `os × shell` test to assert post-fix behavior (4 step
realizations, 2 WRONG_SHELL_FOR_OS, 0 UNRESOLVABLE_MATRIX); added a compliant
`os × node-version` cross-product test; added a fast-check property test that
the realization count equals the product of axis lengths and that no axis key
is dropped from any context.
Closes #435
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
9e3b056b15 |
fix(#779): correct stale model-catalog model IDs verified against live providers (#1047)
Verify-first audit: gemini opus gemini-3-pro→gemini-3.1-pro-preview (undefined in gemini-cli source), codex sonnet gpt-5.3-codex→gpt-5.4 (deprecated per OpenAI); qwen3-coder-next verified valid, unchanged. Adds a regression guard + sourcing note. Closes #779. |
||
|
|
734c56ccfe |
feat(#1046): phase 5c — drive commandStyle from the runtime descriptor (runtime-slash codex-check → lookup) (#1048)
formatGsdSlash now reads commandStyle from registry.runtimes[id].runtime.commandStyle (lazy require of the committed capability-registry.cjs) instead of the hardcoded if (rt === 'codex'). Equivalence-preserving (Codex-verified): codex (shell-var) → $gsd- + lowercased token; all 15 others (slash-hyphen) + unknown → /gsd- + case-preserved token. canonicalizeRuntimeName + input normalization + claude default preserved; no circular load (mirrors 5b runtime-homes pattern). Added a registry-parity test: 16 parametrized sub-tests derive the expected prefix + lowercasing from each runtime's commandStyle, proving the prefix is a pure function of the registry (catches future hardcode-vs-registry divergence). Closes #1046 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6968e04d8a |
feat(#1040): phase 5b — drive configHome from the runtime descriptor — ADR-857/1016 (#1043)
* feat(#1040): phase 5b — drive configHome from the runtime descriptor (runtime-homes switch → lookup) getGlobalConfigDir now resolves configHome from registry.runtimes[id].runtime.configHome via a single resolveConfigHomeFromDescriptor(configHome, {env, home, existsSync}) (dot-home / dot-home-nested / xdg / generic-agents-root), replacing the hardcoded 16-runtime switch. Equivalence-preserving: byte-identical config dirs for all 16 runtimes + grok + default + explicitDir (Codex-verified, no divergence). Nuances preserved: xdg env[1] is a FILE path → path.dirname; existsSync injection seam keeps antigravity/kimi probe tests hermetic; grok stays hardcoded (GROK_AGENTS_HOME → ~/.agents, not in the 16); copilot two-env fallback; explicitDir short-circuit. getGlobalSkillsBase unchanged (out of scope). Lazy require of the committed capability-registry.cjs (no circular load). Test env-clearing lists (install.test ENV_KEYS, bug-3126 envKeys) now derived from the registry runtime configHome.env arrays — auto-correct, closes the missing KIMI_CONFIG_DIR gap. New 81-case golden equivalence test. Closes #1040 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1040): make windsurf golden path Windows-portable (path.join, not POSIX literal) The dot-home-nested windsurf equivalence case hardcoded '/home/u/.codeium/windsurf' but the resolver builds it via path.join(home,parent,name) → backslashes on Windows. Use path.join for the expected value. Test-only; production resolver unchanged. Defensive scan confirmed it was the only path.join-derived hardcoded literal. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1fab2e10ba |
fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver (#1045)
* fix(#1041): route all source agents through the canonical multi-runtime gsd-tools resolver Source agents/*.md (gsd-planner, gsd-executor, gsd-verifier, gsd-plan-checker, gsd-intel-updater, gsd-debugger, …) called bare "gsd-tools …" in shell blocks. On a shim-only install — where gsd-tools.cjs exists under the runtime home but gsd-tools is NOT on PATH — those calls fail with "command not found" and the agent silently skips init/state/validate/commit ceremony, deferring to the orchestrator or bypassing GSD bookkeeping entirely. were never migrated, so it persisted on Claude Code and every other runtime that consumes the source agents directly. Only gsd-phase-researcher.md carried a resolver — and a stale, claude-only truncated one. Fix (all runtimes): - Inject the canonical multi-runtime gsd_run preamble (byte-equal to _runtime-launcher.snippet.sh — claude/codex/cursor/gemini/copilot/windsurf/ augment/trae/qwen/cline/opencode/kilo/hermes/antigravity homes) at the top of the first gsd_run block of all 12 gsd-tools-calling agents, and rewrite every command-position bare gsd-tools to gsd_run. - Upgrade gsd-phase-researcher.md's stale resolver to the canonical one. - Extend scripts/sync-runtime-launcher.cjs to maintain agents/ in parity (the sync caught and corrected a mis-placed preamble during development). - Extend the bare-gsd-tools (#2851) and launcher-parity (#373) regression guards to agents/ so no runtime can silently regress. Closes #1041 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1041): backfill changeset PR number to 1045 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e4dfa6b9ea |
fix(#1012): invoke fallow with its real CLI and wire the report normalizer (#1044)
* fix(#1012): invoke fallow with its real CLI and wire the report normalizer The /gsd-code-review structural pre-pass invoked fallow with flags no published fallow version accepts (--json, --profile, --stdin-files), so it failed on every run and degraded silently per REQ-FALLOW-02 — the feature never delivered on any fallow version. Three compounding defects: 1. Invalid flags. Real fallow audit uses --format json (not --json), -q/--quiet, --changed-since/--base for changed-files scoping (no file-list input), and --max-crap for thresholds. There is no --profile or --stdin-files. 2. Exit-code handling. fallow audit exits 1 when it FINDS issues (verdict=fail), 0 when clean. The pre-pass treated any non-zero exit as a crash and discarded the output — i.e. it threw away exactly the findings it exists to surface. Success is now decided by whether a valid fallow JSON report was produced, not by the exit code. 3. Schema mismatch. normalizeFallowReport parsed a fictional top-level schema (unusedExports/duplicates/circularDependencies) fallow never shipped, and was dead code (the workflow embedded raw JSON; its tests asserted the fictional schema, one even calling a non-existent runFallowAudit and passing vacuously). Fixes: align the invocation to fallow's documented agent-facing pattern; map the profile preset (minimal/standard/strict) to --max-crap (50/30/15); scope phase runs via --changed-since with a repo-scope fallback; rewrite the normalizer to fallow's real schema (dead_code.unused_exports/unused_files/circular_dependencies + duplication.clone_groups) and wire it into the workflow so the reviewer receives normalized findings; replace the fictional-schema fixtures and tests with real-schema ones and delete the vacuous runFallowAudit test. Closes #1012 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1012): backfill changeset PR number to 1044 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1e3ce6df05 |
fix(#222): orchestrator self-heal when synthesizer returns SUMMARY.md inline (#1042)
* fix(#222): orchestrator self-heal when synthesizer returns SUMMARY.md inline The gsd-research-synthesizer agent intermittently hits an LLM false-refusal: instead of writing .planning/research/SUMMARY.md with the Write tool, it returns the SUMMARY.md content inline and fabricates a non-existent write restriction (e.g. "the runtime is blocking file writes"). The shipped prompt hardening (#240) is necessary but insufficient — the false-refusal recurs under some context loads, and a drifting subagent then leaves gsd-roadmapper to fail with "SUMMARY.md not found". Adds an orchestrator-level self-heal to new-project.md and new-milestone.md: after the synthesizer returns, verify .planning/research/SUMMARY.md exists; if it is missing but the agent returned content inline, the orchestrator persists that content with the Write tool (logging a warning) before spawning gsd-roadmapper; if missing with no content, surface the error and stop rather than proceed against a missing SUMMARY.md. This absorbs the failure mode deterministically instead of depending on the subagent never drifting. Closes #222 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#222): backfill changeset PR number to 1042 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ab8b84286e |
fix(#1013): resolve worktree.baseRef from user/global settings cascade (#1038)
* fix(#1013): resolve worktree.baseRef from user/global settings cascade cmdWorktreeBaseCheck resolved worktree.baseRef from the project checkout's .claude/ only (settings.local.json then settings.json). A user/global worktree.baseRef:"head" — the layer /config writes and the harness honors, and the only sensible place for a machine-wide preference — was invisible. On a phase lane (HEAD ahead of origin/HEAD, or no origin/HEAD symref) base-check returned shouldDegrade:true and execute-phase forced sequential execution, silently losing the parallel worktree execution the user configured. CLAUDE_CONFIG_DIR (relocated user config dir) was also ignored. resolveEffectiveBaseRef now accepts an optional user/global config dir and reads its settings.json as a third, lowest-precedence layer (project local > project shared > user/global). cmdWorktreeBaseCheck resolves it via getGlobalConfigDir('claude'), which honors CLAUDE_CONFIG_DIR. The existing project-level reads stay as higher-precedence overrides and the injectable readFile seam is preserved, keeping the unit tests hermetic. Closes #1013 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1013): backfill changeset PR number to 1038 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
12d6285b4f |
fix(#1000): align gsd-intel-updater output to canonical intel filenames (#1037)
* fix(#1000): align gsd-intel-updater output to canonical intel filenames The intel-updater agent was instructed to write short names (files.json, apis.json, deps.json) and a markdown arch.md, but the intel library + gsd-tools intel CLI read only the canonical long names from INTEL_FILES (file-roles.json, api-map.json, dependency-graph.json, arch-decisions.json as JSON). After /gsd:map-codebase --query refresh the agent output was orphaned — intel status and validate reported the canonical files missing and intel query returned nothing. Renames every short reference to its INTEL_FILES canonical name and converts the arch output from markdown to queryable arch-decisions.json. Adds a drift-proof regression test (derived from the exported INTEL_FILES map) in the owning module's test file tests/intel.test.cjs, reviving the maintainer-approved approach from closed PR #608. Closes #1000 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1000): backfill changeset PR number to 1037 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5599e5a9cf |
fix(#1004): detect http-route hook registrations in installer presence check (#1032)
* fix(#1004): detect http-route hook registrations in installer presence check referencesHook only inspected h.command and h.args, so a managed hook re-registered as a type:"http" entry (local hook-server routing) — whose identity lives only in h.url — was invisible. The installer then appended a stock command duplicate on every install/update, running the hook twice per event. Adds the h.url arm, mirroring the #976 args-form fix. Closes #1004 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1004): backfill changeset PR number to 1032 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fbd62cd84f |
feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).
Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).
Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).
Closes #1035
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
8bd5e07c58 |
feat(#1031): autonomous §3a.5 plan:pre ui-phase cutover — last inlined ui-phase site (step-only) (#1033)
Cut over autonomous.md §3a.5 (autonomous plan:pre ui-phase step) to the loop.render-hooks plan:pre dispatch — completing the ui-phase migration begun in #1026 (plan-phase.md §5.6). Step-only, non-blocking: autonomous is always pipeline, so it fires active kind==step hooks and never runs the manual-only plan:pre blocking gate. Skip condition keys on "no active step hooks" (not empty activeHooks), so the gate-only {ui_phase:false, ui_safety_gate:true} case skips silently with no spurious warning — matching OLD §3a.5. Fires gsd-ui-phase under the identical precondition (frontend + no UI-SPEC + workflow.ui_phase active), bare ${PHASE_NUM} args. Replaces the inline ui-safety-gate.cjs probe + config-get with render-hooks + the ui.plan-gate check verb. Codex caught the gate-only spurious-warning divergence on the first pass; fixed + re-confirmed equivalence-preserving. gsd-ui-phase skill, §5.6, §3d.5 untouched. Closes #1031 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9ed8c7d574 |
feat(#1026): §5.6/ui-phase cutover — first gate dispatch (plan:pre step + blocking gate) (#1028)
Replace plan-phase.md §5.6 (a 6-branch inline UI gate) with a capability-driven
loop.render-hooks plan:pre dispatch — the FIRST gate dispatch in any workflow.
A step (ui-phase, when:workflow.ui_phase) + a new blocking gate
(when:workflow.ui_safety_gate). New ui.plan-gate check verb returns
{frontend, hasUiSpec, block}; the dispatch runs it unconditionally then fires
the active step (pipeline) or halts on the active blocking gate (manual). The
gate-handling (run check.query; halt if blocking+block) is the reusable
phase-6 template for blocking-gate cutovers.
Config semantics fixed per #1022 + maintainer call: ui_phase gates plan-time
UI-SPEC generation, ui_safety_gate gates the planning block. Common case + all
ui_phase=false cases are equivalence-preserving; the one intended change is
{ui_phase:true, ui_safety_gate:false} now auto-generating in pipelines.
Review found it broken twice (non-generic dispatch, phase-lookup divergence,
then the step-only check nested in a gate loop) — fixed; final Codex pass
verified all 8 (ui_phase,ui_safety_gate)x{pipeline,manual} cases correct.
gsd-ui-phase skill + autonomous §3a.5 untouched (§3a.5 deferred).
getRoadmapPhaseWithFallback mirrors cmdRoadmapGetPhase for lookup parity.
Closes #1026
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|