9e2ef2c94dea3f1267deba570c9a423fb38c2f85
37 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
94872662e9 |
feat(#1082): complete phase 5 — descriptor-drive all install surfaces + materialize the InstallPlan — ADR-857/1016/58 (#1080)
* feat(#1077): phase 5f-2 — drive the hookEvents dialect (PostToolUse/AfterTool) from the descriptor postToolEvent (bin/install.js) and preToolEvent (applySettingsJsonHooks in runtime-hooks-surface.cts) now select the event-name dialect from registry.runtimes[id].runtime.hookEvents instead of the hardcoded (runtime === 'gemini' || runtime === 'antigravity') check: hookEvents === 'gemini' → AfterTool/BeforeTool; else → PostToolUse/PreToolUse. hookEvents threaded into the applySettingsJsonHooks opts bag. Equivalence-preserving (Codex-verified): hookEvents 'gemini' is exactly {gemini, antigravity}, 'claude' the rest; undefined → claude dialect (matches the old else). The per-event SET guards (isQwen||claude → SubagentStop/Stop/PreCompact; runtime==='claude' → FileChanged; isGemini → Gemini agent-events) stay HARDCODED — hookEvents (2-value) is too coarse to drive them (the event set differs within hookEvents='claude'); per-event-set drive tracked in #1076. Registry-parity test (enh-1077): asserts BOTH post-tool (AfterTool/PostToolUse) AND pre-tool (BeforeTool/PreToolUse) dialects are a pure function of hookEvents, for gemini/antigravity/claude/augment — non-vacuous (catches a broken hookEvents thread). Closes #1077 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1077): build hooks/dist in before() so dialect-drive test passes in scoped CI hooks/dist is gitignored and absent in scoped/windows CI jobs that do not pre-run build:hooks. Without it, install() finds no hook files and all AfterTool/BeforeTool/PostToolUse/PreToolUse event arrays come back empty, failing every hook-presence assertion. Added an idempotent ensureHooksDist() called in a top-level before() — mirrors the pattern from bug-376-claude-js-hook-gsd-rewriter.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1055): add installSurface/writesSharedSettings/permissionWriter/extendedHookEvents to runtime descriptors Purely additive: four new fields on all 16 runtime capability.json descriptors, validator extended with three new closed-vocab sets, registry regenerated. Test fixtures (VALID_RUNTIME_CAP and makeRuntimeCap) updated to include the new required fields so all 255 capability-registry tests continue to pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): drive per-event hook guards from extendedHookEvents descriptor Replace hardcoded runtime-name checks (isQwen||runtime==='claude', runtime==='claude', isGemini) in applySettingsJsonHooks with a single descriptor-driven extendedEvents array derived from the new opts field. Remove isQwen and isGemini derivations (no remaining uses after the three guard blocks are migrated). Wire extendedHookEvents from the capability registry in bin/install.js call site. Add behavioral regression test (enh-1076-extended-hook-events-drive.test.cjs) confirming the drive is purely descriptor-based and runtime-name-agnostic. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1055): drive resolveRuntimeConfigIntent from the runtime descriptor; retire hand-kept REGISTRY - Rewrites src/runtime-config-adapter-registry.cts to require capability-registry.cjs and read installSurface / writesSharedSettings / permissionWriter from runtimes[id].runtime; deletes the hand-kept REGISTRY const (ADR-857 phase 5g drive 2). - ALLOWED_CONFIG_RUNTIMES is now derived from descriptor entries that have installSurface. - Fixes the configFormat parity gate in scripts/gen-capability-registry.cjs to read installSurface directly from capMap descriptor bodies, breaking the require cycle (adapter now requires the generated registry; gen-script must not require the adapter). - Adds golden-master test tests/enh-1055-config-intent-descriptor-drive.test.cjs (41 tests) pinning all 16 runtimes' return shapes and the TypeError-on-unknown contract. - Updates scripts/lint-test-file-count.allowlist.json (config module, +1 file). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): make hooksSurface descriptor load-bearing for the settings-json hook-skip - Adds hooksSurface?: string to ApplySettingsJsonHooksOpts and destructuring in applySettingsJsonHooks (src/runtime-hooks-surface.cts). - Replaces the hardcoded !isOpencode && !isKilo hook-skip guard with hooksSurface !== 'none'; removes the now-unused isOpencode/isKilo derivations (ADR-857 phase 5g drive 3). - Passes hooksSurface from the runtime descriptor at the applySettingsJsonHooks call site in bin/install.js using the established _capabilityRegistry?.runtimes?.[runtime]?.runtime?.hooksSurface idiom. - Extends tests/enh-1076-extended-hook-events-drive.test.cjs with two new suites proving: (a) hooksSurface:'none' writes no hooks regardless of runtime name; (b) hooksSurface:'settings-json' writes hooks even for 'opencode' (previously hardcoded to skip). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: record installSurface/writesSharedSettings/permissionWriter/extendedHookEvents descriptor axes in ADR-1016 Add Decision 7a documenting the four axes added in the 5f-completion pass, update axis counts from "six" to "twelve", note 5f-completion drives as done in Decision 8's ladder, update Out of scope to reflect #1055/#1076 are done and 5g (InstallPlan capstone) remains the only open phase. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1055): parity gate must fire on configFormat↔installSurface mismatch (read installSurface at the descriptor level) The test fixture makeRuntimeCapMap did not include installSurface in the runtime object, so the gate's typeof r.installSurface !== 'string' guard always skipped the entry and never threw. Added installSurface as an optional third parameter to makeRuntimeCapMap and passed the correct installSurface values ('settings-json' for claude, 'codex-toml' for codex) to the two THROWS tests. The gate implementation already reads r.installSurface correctly from the descriptor level. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1076): add installSurface↔hooksSurface + extendedHookEvents↔hookEvents consistency gates with rejection tests GATE A: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES map in validateRuntimeBody enforces that a runtime's hooksSurface is valid for its installSurface (e.g. profile-marker-only only allows none, codex-toml only allows codex-hooks-json). Derived from the 16 real runtime descriptors. GATE B: validateRuntimeBody checks that if extendedHookEvents contains Gemini agent-events (BeforeAgent/AfterAgent/BeforeModel), hookEvents must be 'gemini'; if it contains Claude-family events (SubagentStop/Stop/PreCompact/FileChanged), hookEvents must be 'claude'. Added 10 rejection tests in suite 27 covering each gate + each new field validator. All 16 real runtimes satisfy both gates (verified before coding). Exports: INSTALL_SURFACE_TO_ALLOWED_HOOKS_SURFACES, VALID_INSTALL_SURFACES, VALID_EXTENDED_HOOK_EVENTS, VALID_PERMISSION_WRITERS, validateRuntimeBody. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1076): strengthen hooksSurface-drive assertions; defensive hooksSurface fallback; drop vacuous dup 1. bin/install.js: add explicit literal fallback for hooksSurface when the committed capability registry fails to load (opencode/kilo → 'none', all others → 'settings-json'). The descriptor is always the source of truth in normal operation. 2. enh-1076 Suite 7: change SessionStart assertion from key-presence (hasOwnProperty) to at least-one-command (hasHooksFor), so the test fails if hooks are initialized-but-empty. ensureHooksDist() in before() guarantees hook files exist. 3. enh-1055 Test 2: remove vacuous duplicate suite that re-asserted intent.runtime === row.runtime already fully covered by Test 1's deepStrictEqual over all four fields. 4. capability-registry.test.cjs: fix stale comments in the grok-skip test that claimed the parity gate uses the adapter registry; gate reads purely from the descriptor (installSurface absent → typeof r.installSurface !== 'string' → soft-skip). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1082): materialize the InstallPlan — collect install-level descriptor axes into resolveInstallPlan; route install()/finishInstall() through it (ADR-58/5g) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: record 5g InstallPlan materialization (ADR-58 Accepted, ADR-1016 phase-5 complete) ADR-1016 Decision 8 step 7 updated to DONE: resolveInstallPlan(runtime) in runtime-config-adapter-registry collects install-level descriptor axes into the typed InstallPlan consumed by install()/finishInstall(). Out-of-scope section updated: 5g capstone is complete, phase 5 fully materialized. ADR-1016 line ~20 updated: InstallPlan IS now materialized (both halves). ADR-58 Implementation note added (2026-06-11): realized in runtime-config-adapter-registry (co-located with adapter-selection). CONTEXT.md Runtime Config Adapter Registry entry extended to document resolveInstallPlan and both-halves realization. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1082): update install drift guard to the resolveInstallPlan seam (5g) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
93c5ecd645 |
feat(#792): add devin-desktop runtime alias for windsurf (#1086)
windsurf now also answers to devin-desktop (CLI --devin-desktop) for the Windsurf→Devin Desktop rebrand; all paths unchanged. The .devin/skills/ workspace migration is split to #1085. Closes #792. |
||
|
|
7c07fce70f |
fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes (#1084)
* fix(#381): make gsd_run launcher reachable in fresh-shell-per-block runtimes On runtimes that execute each fenced bash block in a separate shell process (e.g. Claude Code — documented behavior: each Bash command is a separate process; inline shell functions and exported vars do not persist between calls), the once-per-file gsd_run() function was undefined in every block after the preamble block, and the call was swallowed by `2>/dev/null || echo "{}"` into silent empty state. Fix (budget-neutral session-level resolution): - Ship gsd-core/bin/gsd_run, a POSIX sh wrapper that symlink-resolves its own location and execs the co-located gsd-tools.cjs. Exposed on PATH via the npm `bin` field (global installs) and shipped to local installs via the recursive gsd-core/ copy. - The per-file launcher preamble now appends `export PATH='<bindir>':"$PATH"` to the file named by $CLAUDE_ENV_FILE (Claude Code's documented env-persistence mechanism) so later fresh-shell blocks resolve gsd_run from PATH. Guarded as a strict no-op when CLAUDE_ENV_FILE is unset; the inline gsd_run() definition remains the fallback for all other runtimes. The single-quoted dir neutralizes shell metacharacters at source time. - Propagated via scripts/sync-runtime-launcher.cjs to all launcher-using files. - XL workflow byte budget 93000 -> 93200 (the ~130B clause pushes plan-phase.md to 93135; legitimate content growth, ratchet-up per #717). Regression tests (I)/(J) in runtime-launcher-parity.test.cjs cover wrapper delegation and end-to-end PATH persistence (sourcing the env file with a space-bearing install path). Known limitation: an install path containing a literal single-quote yields a malformed env-file line and falls back to the status quo (no regression); rare on sanitized home directories. Closes #381 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#381): add changeset for gsd_run fresh-shell reachability fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#381): scope test (J) bare-PATH execution to POSIX (Windows Git Bash exec bit) Windows Git Bash (msys2) does not honor Node's chmod exec bit for PATH-executing extension-less scripts, so the bare `gsd_run` command lookup failed there even though the env-file PATH persistence was correct. The env-file content assertions (the fix's actual cross-platform logic) still run on every platform; only the final source-and-execute sub-step is gated to non-win32. Global installs on Windows are covered by npm's generated bin shim. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
58ed55683e |
feat(#1049): phase 5d — drive artifactLayout from the runtime descriptor (retire the 128-LOC switch) (#1053)
resolveRuntimeArtifactLayout now builds Layout from
registry.runtimes[id].runtime.artifactLayout[scope] — a loop dispatching each
ArtifactKind through the SAME 5 builders (commandsKind/agentsKind/skillsKind/
convertedCommandsKind/kimiAgentsKind, unchanged) by (kind, converter, nesting) —
replacing the hardcoded switch(runtime). Equivalence-preserving for all 16 runtimes
× {global, local} (Codex-verified, no divergence). -43 LOC; bin/install.js + the
converters + the install loop untouched. getInstallExports()[converterName]
resolution, configDir threading, scope default, unknown-runtime guard all preserved.
Driving the local scope surfaced a 5a gap: the old switch had no scope branch for 13
runtimes (cursor/gemini/codex/copilot/antigravity/windsurf/augment/trae/qwen/hermes/
codebuddy/opencode/kilo) → local == global for them, but 5a authored local:[].
Backfilled local=global for those 13 (descriptor-faithful; a fall-through shim would
wrongly give cline/kimi local=global). claude/cline/kimi scope-gating untouched.
validateArtifactKindEntry tightened: destSubpath/prefix/nesting/converter required
(ConverterName enum still open — 5e). New 39-case deep-equal golden equivalence test.
Closes #1049
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
9e3b056b15 |
fix(#779): correct stale model-catalog model IDs verified against live providers (#1047)
Verify-first audit: gemini opus gemini-3-pro→gemini-3.1-pro-preview (undefined in gemini-cli source), codex sonnet gpt-5.3-codex→gpt-5.4 (deprecated per OpenAI); qwen3-coder-next verified valid, unchanged. Adds a regression guard + sourcing note. Closes #779. |
||
|
|
fbd62cd84f |
feat(#1035): phase 5a — author 16 role:runtime capability descriptors (registry-only) (#1039)
Author capabilities/<runtime>/capability.json for all 16 runtimes (role:runtime),
populating the registry's runtimes index ({} -> 16). Each carries the 6 ADR-1016
axes (configHome structured, configFormat, artifactLayout structured, commandStyle,
hooksSurface + hookEvents, sandboxTier, supportTier), extracted from the live
modules (runtime-homes/runtime-slash/runtime-artifact-layout/runtime-config-adapter
+ CODEX_AGENT_SANDBOX). validateRuntimeBody tightened to enforce the closed
vocabularies + structured configHome/artifactLayout (rejects old string configHome;
env required; skillsHome recursively validated).
Registry-only: bin/install.js + src/*.cts untouched, zero behavior change. The 5b-5f
drive steps consume these one axis at a time (ADR-1016 §8).
Codex accuracy review caught + fixed: opencode/kilo register ZERO lifecycle hooks
(install.js skips the whole block) -> hooksSurface:none (configFormat stays
settings-json); kimi probe selects on <root>/skills existence -> probeExists:"skills".
ArtifactKind required-field strictness + ConverterName closing deferred to 5d/5e.
ADR-1016 (Proposed) amended to match (Decisions 1 + 5).
Closes #1035
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
9ed8c7d574 |
feat(#1026): §5.6/ui-phase cutover — first gate dispatch (plan:pre step + blocking gate) (#1028)
Replace plan-phase.md §5.6 (a 6-branch inline UI gate) with a capability-driven
loop.render-hooks plan:pre dispatch — the FIRST gate dispatch in any workflow.
A step (ui-phase, when:workflow.ui_phase) + a new blocking gate
(when:workflow.ui_safety_gate). New ui.plan-gate check verb returns
{frontend, hasUiSpec, block}; the dispatch runs it unconditionally then fires
the active step (pipeline) or halts on the active blocking gate (manual). The
gate-handling (run check.query; halt if blocking+block) is the reusable
phase-6 template for blocking-gate cutovers.
Config semantics fixed per #1022 + maintainer call: ui_phase gates plan-time
UI-SPEC generation, ui_safety_gate gates the planning block. Common case + all
ui_phase=false cases are equivalence-preserving; the one intended change is
{ui_phase:true, ui_safety_gate:false} now auto-generating in pipelines.
Review found it broken twice (non-generic dispatch, phase-lookup divergence,
then the step-only check nested in a gate loop) — fixed; final Codex pass
verified all 8 (ui_phase,ui_safety_gate)x{pipeline,manual} cases correct.
gsd-ui-phase skill + autonomous §3a.5 untouched (§3a.5 deferred).
getRoadmapPhaseWithFallback mirrors cmdRoadmapGetPhase for lookup parity.
Closes #1026
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
76f42ddb4b | feat(#1014): add Claude Fable 5 model config (#1015) | ||
|
|
626575cbc5 |
fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works (#982)
* fix(#978): parse --force in milestone complete dispatcher so the guard's documented override works The dispatcher built `{ name, archivePhases }` but never parsed `--force`, so `options.force` was always `undefined` and the guard inside `cmdMilestoneComplete` (which tells users to "Re-run with --force to override") could never be bypassed. Add `const force = args.includes('--force')` and pass it into the options object. The guard already honors `options.force` — no changes to milestone.cts needed. Closes #978 * chore(#978): backfill changeset pr number (982) --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
caca4d255c |
feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4) (#988)
* feat(#985): intel command cutover — commands-only capability, last first-party family (ADR-857 phase 4d-impl-4) Migrate the intel CLI command family from a hardcoded gsd-tools.cjs case arm to a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring the graphify (#972) and audit (#984) cutovers. New src/intel-command-router.cts exports routeIntelCommand reproducing all 9 subcommands verbatim (incl. the status timeAgo non-raw post-processing), lazily requiring intel.cjs inside the route fn. capabilities/intel/capability.json declares the intel command family; commands-only (skills:[]), declares the existing intel.enabled gate (default false — behavior unchanged). Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical; existing intel.test.cjs passes unchanged. Completes the first-party command- family cutover sequence (graphify/audit/intel). Closes #985 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#985): compute expected planningDir via path.join in intel cutover unit tests (Windows CI) The intel-command-cutover unit-mock assertions hardcoded a POSIX `/.planning` expectation while the router builds it with path.join(cwd, '.planning') → backslashes on Windows, so the planningDir-arg assertions (query/status/diff/ snapshot/validate/update/api-surface) failed only on windows-latest CI. Compute the expectation with path.join (cross-platform); production router unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9a03539c2d |
feat(#981): audit-uat + audit-open command cutover — commands-only capability (ADR-857 phase 4d-impl-3) (#984)
Migrate the audit-uat + audit-open CLI commands from hardcoded gsd-tools.cjs case arms to a registry-dispatched Capability (commandFamilies mechanism, #961), mirroring the graphify cutover (#972). New src/audit-command-router.cts exports routeAuditUat/routeAuditOpen, each lazily requiring only its backing module (uat.cjs/audit.cjs) inside the route fn — matching the old per-case lazy loads. capabilities/audit/capability.json declares the two command families; commands-only (skills:[]), no config gate, audit_review cluster untouched. Behavior-CHANGING (dispatch path) but equivalence-proven: CLI output identical; existing audit regression tests pass unchanged. Closes #981 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
77bfd943dd |
feat(#972): graphify command cutover — first capability owning a command family (ADR-857 phase 4d-impl-2) (#975)
Migrate graphify into a Capability that owns its `graphify` command family, dispatched via the registry (#961 mechanism) instead of a hardcoded case. graphify is now an enable/disable plug-in. - src/graphify-command-router.cts: routeGraphifyCommand (standard route*Command), reproduces the removed case EXACTLY (query +--budget, status, diff, build, hidden build snapshot, usage/unknown errors); injectable _graphify test seam. - capabilities/graphify/capability.json: role feature, tier:full, skills:[graphify], config:{graphify.enabled default false}, commands:[{family:graphify, module, router:routeGraphifyCommand}]. - removed case 'graphify' from gsd-tools.cjs; graphify now flows default -> dispatchCapabilityCommand -> commandFamilies.graphify -> router. - regenerated registry (commandFamilies/bySkill/configSchema/profileMembership/ capabilityClusters for graphify); tier:full keeps 4c install/surface a no-op. Equivalence-proven: 8 recording-mock unit tests assert the exact fn+args per subcommand (budget, snapshot-vs-build); 9 subprocess tests assert distinguishing output shapes; existing graphify tests pass unchanged through the new path. Surfaced (not silently accepted): the pre-existing --budget-no-value NaN no-op quirk, preserved for equivalence, filed separately. Closes #972 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0c567c17e4 |
feat(#961): capability command mechanism (commandFamilies index + default-case dispatch) — ADR-857 phase 4d-impl-1 (#964)
* feat(#961): capability command mechanism — commandFamilies index + default-case dispatch (ADR-857 phase 4d-impl-1) Build the capability command mechanism per ADR-959: the `commands` declaration field on the feature role, the registry commandFamilies index, and a real dispatchCapabilityCommand consulted in runCommand's default case (replacing the dead _dispatchNonFamily shim's role). The registry DISCOVERS a standard route*Command (no rebuilt handler table). On a default-case command, dispatch enforces a bare-.cjs-basename, resolves the module under gsd-core/bin/lib/, asserts confinement, requires the resolved path, and own-property-guards the router export before calling it. A require.main===module guard makes gsd-tools.cjs importable for tests; the CLI path is unchanged. Additive: commandFamilies is empty today, so the default case is behavior- preserving for every command; the 10 dead _dispatchNonFamily sites are untouched (future migration markers). The graphify cutover is the separate next step. Closes #961 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#961): surface capability router failures structurally + enforce sync contract Review follow-up: dispatchCapabilityCommand now wraps the router invocation so an unexpected (non-ExitError) throw is converted to a structured, attributed error(msg, SDK_FAIL_FAST) — honoring --json-errors — instead of escaping as a raw stack trace; an intentional ExitError propagates unchanged. An async router (returns a thenable) is rejected loudly with a structured error (the contract is synchronous, like the 12 host routers). Not shipped as "consistent with existing behavior": the host's pre-existing version of this gap is filed as #965. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dc139b38e0 |
feat(#949): install/surface consume derived profiles/clusters (ADR-857 phase 4c) (#954)
Make install + surface read the registry's derived profileMembership/ capabilityClusters so a capability's tier drives what installs + surfaces. resolveProfile (when given the registry) unions capability skills for the profiles its tier implies before the requires: closure; resolveSurface merges capabilityClusters into the cluster map. bin/install.js, /gsd:surface, and the capability-state resolver all thread the registry. Shipped as a proven no-op: the UI capability is reconciled to tier:full (its skills were full-only in the hand-authored profiles), so it contributes only to the full profile (already the '*' sentinel) and core/standard are unchanged. Equivalence tests prove resolveProfile/resolveSurface/listSurface/staging/ capability-state are identical with vs without the registry; the core-alias staging path is verified equivalent (empty manifest → raw PROFILES.core). Closes #949 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
85cfa5dc13 |
feat(#945): unified capability-state resolver (ADR-857 phase 4b) (#946)
Add a read-side query composing the three toggle systems into one
per-capability view. resolveCapabilityState({registry, installedSkills,
surfacedSkills, config, cwd}) reports installed (skills ⊆ resolved install
profile), surfaced (skills ⊆ resolved surface), and per-hook active (no when →
active; non-empty-string when → resolved via _resolveActivationValue; empty/
non-string → inactive), with no forced composite verdict. cmdCapabilityState
does the I/O (resolveProfile + resolveSurface + loadConfig), resolves the
runtime config dir via the canonical getGlobalConfigDir (--config-dir override),
and surfaces resolution failures as warnings rather than a false installed='*'.
Routed as `gsd-tools capability state`.
Additive: install/surface/workflows untouched; consumed by nothing.
Closes #945
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
ed467cd8f2 |
feat(#942): tier→profile/cluster derivation + consistency gate (ADR-857 phase 4a) (#943)
Establish `tier` as the source of install-profile + cluster membership
(ADR-894 §4). The registry now derives two views: capabilityClusters
(capId → its skills) and profileMembership (capId → {tier, profiles}, where
profiles is the PROFILE_RANK suffix from the capability's tier). A consistency
gate cross-checks them against the hand-authored PROFILES/CLUSTERS: HARD (throws)
on a capId-matching-a-cluster-name with a different skill set; SOFT
(pending-reconciliation stderr warning, not serialized) when a capability skill
isn't yet in the closure-resolved hand-authored profile.
The SOFT gate loads the real skills manifest and resolves each profile's closure
once so transitively-included skills don't false-warn; warnings are de-duped to
one per (capability, skill); both derived views are scoped to skill-owning
capabilities; serialized with a global capId sort for determinism; reserved-name
guards at every write site; lazy requires of the built constants.
Behavior-preserving: install/surface untouched; derived views consumed by
nothing. Full generation rides along with the phase-6 migration.
Closes #942
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
74a121bb4f |
fix(#934): reapply verifier handles missing pristine baseline post-rename (#937)
* fix(#934): reapply verifier handles missing pristine baseline post-rename Gap 1 (verify-reapply-patches.cjs): when backup-meta.json records a pristine_hash for a file but gsd-pristine/ has no corresponding snapshot on disk, the verifier fell to over-broad mode and produced false FAIL_USER_LINES_MISSING. Fix: return advisory OK_NO_BASELINE (non-blocking, exit 0) so the verifier does not block on files it cannot reason about. Gap 2 (new migration 004): migration 003 removed legacy get-shit-done/ runtime files but left gsd-pristine/get-shit-done/ orphan snapshots in place. Those stale snapshots referenced get-shit-done/... key paths that no longer match the active gsd-core/... layout. Fix: add migration 004-prune-stale-pristine-get-shit-done (NOT editing 003, preserving its checksum — ref #670 guard) to remove all files under gsd-pristine/get-shit-done/ as GSD-managed pristine snapshots. Includes tests: bug-934 OK_NO_BASELINE assertions in the verifier test, new installer-migration-prune-stale-pristine.test.cjs, updated installer-migrations baseline-lock checksum for 004. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#934): rename migration to satisfy legacy-name guard + mark intentional path refs Rename src/installer-migrations/004-prune-stale-pristine-get-shit-done.cts → 004-prune-stale-pristine-snapshots.cts so the filename no longer contains the forbidden token. Update .gitignore and eslint.config.mjs to track the new built path. Add gsd-allow-legacy-name markers to the remaining intentional uses of the legacy path string in the migration body (lines 3 and 100) and in tests (installer-migration-prune-stale-pristine.test.cjs lines 202 and 226; and the baseline-lock key in installer-migrations.test.cjs:1469). Update the baseline checksum for migration 2026-06-09-prune-stale-pristine-get-shit-done to reflect the two new marker comments added to its body. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c6c3a51277 | Merge branch 'next' into kimi-runtime-support | ||
|
|
5670feaef5 |
feat(#918): loop.render-hooks resolver — consume the Capability Registry (ADR-857 phase 3c) (#920)
Add the loop.render-hooks resolver: the first registry-consuming query.
`gsd-tools loop render-hooks <point>` validates the point against the
authoritative canonical 12, reads the registry's materialized byLoopPoint
hooks, filters them by activation, and emits a JSON envelope {point,
activeHooks, rendered} with ordered markdown.
Activation resolves each hook's `when` key by precedence: loadConfig value
(post-cutover federated) -> raw config.json workstream/root single-key lookup
(pre-cutover central override) -> registry configSchema default (so a
default:true capability hook is active out-of-the-box) -> inactive. Guarded
single-value reads only (no merged object built from untrusted keys).
Registry-only: no workflow calls the resolver yet (wiring is the phase-6
cutover). Completes the phase-3 trio.
Closes #918
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
6dbd895028 |
feat(#910): federated config merge in config-loader (ADR-857 phase 3b) (#914)
Build the federated config merge: each capability owns its config-key slice
(ADR-857 decision 3 / ADR-894), and loadConfig merges them defensively. The
registry now emits a full configSchema index ({key:{owner,type,default,
description}}, generator-validated); a new src/federated-config.cts resolves
federated keys defensively (skip central keys -> pending-migration warning,
skip malformed slices -> warning never throw, else type-checked user override
?? default, with nested dotted-path lookup and enum validation); and loadConfig
applies the overlay on every return path.
Wired as a provably-empty no-op channel: every UI-pilot key is still central,
so validKeys is empty and loadConfig returns byte-identical output on all paths
(identity return when the overlay is empty; shared CONFIG_DEFAULTS never
mutated). Registry-only; no key is cut over; nothing in the live loop changes.
Closes #910
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
bf8813b7a0 | Merge branch 'next' into kimi-runtime-support | ||
|
|
48cc27bd84 |
feat(#903): generate Loop Host Contract from workflow markers (ADR-857 phase 3a-impl-2) (#906)
Replace the inline LOOP_HOST_CONTRACT constant in the Capability Registry generator with a generated-from-workflows contract (ADR-894 §3). The contract is now derived from inert `<!-- gsd:loop-host ... -->` marker blocks in the five step workflows, emitted as the committed gsd-core/bin/lib/loop-host-contract.cjs, and required by gen-capability-registry.cjs — one source of truth, no drift. Drift guards in gen-loop-host-contract.cjs: per-step point ownership (each step must declare exactly its canonical loop points), multiple-block + duplicate-key hard errors, and a word-boundary agent-role cross-check. Contract content is byte-identical to the former constant; registry-only, nothing wired into the live loop. Closes #903 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ad754ca6cd |
feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl) (#902)
* feat(#896): Capability Registry generator + UI pilot (ADR-857 phase 3a-impl) First phase-3 code: the Capability Registry generation pipeline, built against the ADR-894 contract and NOT wired into the live loop (registry-only, per the staged-cutover design). - capabilities/ui/capability.json — the UI pilot (ADR-894 worked example): 2 skills, 2 agents, 3 config keys, 2 steps + 1 gate, with `when` activation. - scripts/gen-capability-registry.cjs — --write/--check generator. Hand-rolled schema validation (envelope + role-typed feature/runtime bodies + typed steps/contributions/gates + when + gate-check variants); cross-capability invariants (single ownership; requires exist+acyclic+tier-monotone; config-key ownership exclusive, collision-vs-central as a pending-migration warning); hooks validated against an inline LOOP_HOST_CONTRACT (3a-impl-2 swaps its source to the generated-from-workflows contract); GLOBAL point-ordered consumes-satisfiability; materialized byLoopPoint ordering (produces/consumes topo-sort); emits gsd-core/bin/lib/capability-registry.cjs (role-partitioned indexes + requiresClosure). Prototype-pollution guards (Object.create(null) + inline literal key checks) + fragment.path traversal guard. - gsd-core/bin/lib/capability-registry.cjs — committed generated artifact (mirrors package-identity.cjs: script-generated, tracked, linted, regenerated on build, drift-tested), wired via the new `gen:capability-registry` build step. - tests/capability-registry.test.cjs — 72 tests: schema + invariant + hook + ordering + adversarial (path-traversal, proto-pollution, runtime body, self-consume, cycles, collisions) + committed-file staleness guard. New-CLI-module checklist (INVENTORY 97->98, MANIFEST, ARCHITECTURE), CONTEXT.md "Capability Registry" un-[Planned]'d. Nothing wired into install/surface/loop. Gates: lint, code-review (4 bugs fixed), security-review (path-traversal + prototype-pollution fixed), codex adversarial-review ×3 (8+ findings fixed, confirmed sound), clean-build docker 13190 pass / 0 fail. Closes #896 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#896): CRLF-agnostic --check for capability-registry staleness (Windows) The committed capability-registry.cjs staleness guard failed on Windows CI only: git checks out the committed .cjs as CRLF (autocrlf, no .gitattributes) while the generator emits LF, so the byte-for-byte --check comparison mismatched. Normalize line endings on both sides of the --check comparison (no .gitattributes change, no change to the LF the generator writes). Adds a regression test simulating the Windows CRLF checkout. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
54191f6c6c | Merge branch 'next' into kimi-runtime-support | ||
|
|
40d48c0508 |
feat(#815): add /gsd-update --next to install the @next RC channel (#839)
Adds an opt-in --next (alias --rc) flag to /gsd-update targeting the @next RC dist-tag (ADR #660), with a {latest,next} allowlist enforced at three layers, channel-aware version check + banner, and byte-for-byte unchanged default @latest behavior. Closes #815 |
||
|
|
7f868dcc6b | fix: close Kimi runtime review gaps | ||
|
|
4e12967683 | Merge remote-tracking branch 'upstream/next' into kimi-runtime-support | ||
|
|
4056d830bc |
refactor(#651): consolidate verification-status routing into one queryable seam (#755)
* refactor(#651): consolidate verification-status routing into one queryable seam The passed/gaps_found/human_needed verification status was re-encoded as bare strings across three prose surfaces (gsd-verifier emits, execute-phase routes, ship gates), each independently deciding the per-status next action with no parity coupling — the DEFECT.GENERATIVE-FIX class. Give the enum one home: src/verification.cts (-> bin/lib/verification.cjs) exposing `gsd_run query verification.status <phaseDir>` returning a typed {status, next_action, next_command}. ship.md and execute-phase.md now consume the query instead of re-deriving the routing in prose; gsd-verifier.md points at the shared vocabulary as the single emitter (values unchanged). Also fixes the latent broad-grep status misread (DEFECT.FRONTMATTER-SCALAR- BROAD-GREP): execute-phase.md read `grep "^status:"` over the whole report, so a body `status:` line could misroute a valid phase. Extraction is now frontmatter-scoped in one place. A parity test fails if a verifier status gains no route. Lands the two CONTEXT.md DEFECT entries captured on the issue. Closes #651 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#651): set changeset pr to 755 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f7e902f1cf |
feat(#52): add agent_skills_security.trusted_global_roots allowlist for global skills (#754)
* feat(#52): add agent_skills_security.trusted_global_roots allowlist Opt-in allowlist so a global: agent skill whose SKILL.md realpath resolves outside the default global skills base (e.g. ~/.claude/skills) is accepted when its real target lies under a user-declared trusted root. Default [] is byte-identical to prior behavior; the symlink-escape guard is preserved and simply re-applied against each declared root. - src/security.cts: loadTrustedGlobalRoots — tilde-expand (~ and ~/), reject project-relative and dangerously broad roots (filesystem/UNC root, homedir), realpath-canonicalize each root every run and drop non-existent ones. - src/init.cts: on base-check failure the guard consults the trusted roots (hoisted out of the loop); emits a stderr NOTE when a skill is accepted via a trusted root so the widened boundary is visible. - src/core.cts: thread agent_skills_security through loadConfig. - config-schema.manifest.json: allow the new key path. - docs/CONFIGURATION.md: document the option and its security model. - tests/agent-skills.test.cjs: unit + end-to-end CLI coverage (regression, feature, negative, broad-root hardening, stderr NOTE). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#52): add changeset fragment for trusted_global_roots (#754) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
7e76f1a736 |
feat(#703): add --granularity override flag to /gsd:plan-phase (#750)
* feat(#703): add --granularity override flag to /gsd:plan-phase Add a `--granularity <coarse|standard|fine>` flag to /gsd:plan-phase that overrides the configured planning granularity for a single invocation. The override is a new highest-priority tier above the existing precedence chain (granularities[phaseType] -> granularity -> planning.granularity -> 'standard') in resolveGranularityInternal; when the flag is absent, resolution is byte-for-byte unchanged. cmdInitPlanPhase now resolves with phaseType 'planning' so granularities.planning participates, and emits the resolved value in the init JSON, which the plan-phase workflow forwards to the planner prompt. Invalid values are rejected at the CLI boundary via a shared assertValidGranularityOverride helper. Closes #703 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#703): set changeset pr to 750 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cf8bd3cd5e |
fix(#683): auto-degrade phase execution to sequential on worktree base mismatch (#749)
* fix(#683): auto-degrade phase execution to sequential on worktree base mismatch Claude Code forks worktree-isolated executors off the repository default branch (origin/HEAD), not the orchestrator's HEAD. Running /gsd-execute-phase on a branch diverged from the default (unmerged milestone/feature branch) left every executor without the phase's plan files and tripped the worktree-branch-check guard with `exit 42` — 100% reproducible, all OSes. - New module src/worktree-base-ref.cts: HEAD-vs-fork-base drift detection (origin/HEAD with symbolic-ref fallback) and no-clobber worktree.baseRef management, exposed as `worktree base-check` / `worktree set-baseref`. - execute-phase.md: pre-dispatch, for Claude Code with worktrees enabled, auto-degrades the run to sequential on the main tree when a base mismatch is detected, recommending worktree.baseRef:"head". The exit-42 guard stays as a backstop. - Installer: fresh local Claude installs set worktree.baseRef:"head" in .claude/settings.local.json (no-clobber, respecting an explicit shared settings.json value); upgrades print an opt-in notice pointing at `gsd-tools worktree set-baseref`. - Docs: how-to guide, CLI/config reference, planning-config cross-ref. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#683): auto-apply worktree.baseRef on upgrade; gate fresh+upgrade on use_worktrees Per maintainer direction: on a local Claude Code UPGRADE, set worktree.baseRef:"head" automatically (no opt-in notice) when the project's workflow.use_worktrees is enabled, instead of merely printing a remediation notice. For consistency the FRESH path is now gated the same way: both paths compute worktrees-enabled once (bounded walk-up read of .planning/config.json, default enabled unless workflow.use_worktrees === false) and apply the no-clobber baseRef only when enabled — never overwriting an explicit value in settings.local.json or a shared settings.json. gsd-tools worktree set-baseref remains for manual use. Docs + changeset updated; tests hardened (file-exists assertions, fresh+disabled case, upgrade idempotency). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#683): measure workflow byte-budget on LF, fixing Windows-only CI failure The workflow-size-budget test failed only on Windows: git checks out the .md files as CRLF (no eol=lf in .gitattributes) and byteCount used fs.statSync().size (raw on-disk bytes), counting an extra \r per line. That inflated execute-phase.md — the XL high-water-mark file pinned near its ceiling by the tighten-only ratchet — from 88492 LF bytes to ~90245 on Windows, over the 90000 XL ceiling, while passing on the LF-checkout Mac/Linux runners. The ceilings are explicitly "calibrated against raw `wc -c`" on an LF checkout, so the measurement should be LF-based on every platform. byteCount now reads the file and counts Buffer.byteLength after stripping CR, making the budget platform-independent (a no-op on LF checkouts; verified statSync === normalized for all 88 workflow files). No ceilings changed. Added a regression test asserting CRLF and LF content of the same file count identically. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#683): make worktree-base-ref test path mocks Windows-safe (path.join) tests/worktree-base-ref.test.cjs keyed its injected readFile/writeFile mocks (and a few expected `file` values) with forward-slash template literals like `${claudeDir}/settings.local.json`. The module composes those paths with path.join(), which emits backslashes on Windows, so the mock keys never matched the module's lookup → readFile returned null → resolveEffectiveBaseRef / cmdWorktreeBaseCheck / cmdWorktreeSetBaseRef (and the JSONC variants) failed on the Windows full-test runner only (they passed on Mac/Linux, and the install tests passed because they use the real filesystem). The module is correct; only the test fixtures hardcoded '/'. All mock keys and path assertions now use path.join(base, ...) mirroring the module, so they match on every platform (no-op on POSIX). 19 path references across 16 lines. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
86836ebd2a | Merge branch 'open-gsd:next' into kimi-runtime-support | ||
|
|
b79b767629 |
refactor(bin): replace process.exit() in CLI entrypoints + ratchet rule to error (#738) (#741)
Part 2 of 2 of the n/no-process-exit cleanup (completes umbrella #738; part 1 was #739/scripts). Converts the 20 flagged process.exit() calls in the three hand-written gsd-core/bin CLI entrypoints and flips n/no-process-exit to error. - New src/cli-exit.cts -> gsd-core/bin/lib/cli-exit.cjs (ExitError + runMain), the gsd-core-side equivalent of scripts/lib/cli-exit.cjs; registered in .gitignore, eslint ignores, and the inventory manifest like its siblings. - gsd-tools.cjs: 13 apply-prompt-budget exits -> throw ExitError; main()->runMain. - verify-reapply-patches.cjs: 6 exits -> throw ExitError / return verdict; runMain. - check-latest-version.cjs: 1 exit -> return verdict; runMain. - eslint.config.mjs: n/no-process-exit warn -> error. Scope note: the gsd-core/bin/lib/*.cjs modules (core, state, profile-pipeline, roadmap-command-router, adr-parser, ui-safety-gate) are tsc-generated and eslint-ignored (ADR-457), so their process.exit calls were never flagged and are intentionally left untouched. Only the linted hand-written entrypoints are in scope. Exit codes verified unchanged for all three entrypoints. Closes #738 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d008c83ec9 |
feat(01-02): add Kimi runtime name policy
- Register canonical kimi in runtime alias manifest and fallback policy - Add focused canonicalization coverage without extra aliases |
||
|
|
ba231ecbfc |
chore: clean up clear-cut ESLint warnings (#732) (#734)
Pay down pre-existing error→warn lint debt. Removes dead imports/vars, unused functions, redundant regex/string escapes, and stale eslint-disable directives; converts unused `catch (_e)` to optional catch binding (src/*.cts). No behavior change. Lint 345→125 warnings (0 errors); deferred categories (n/no-process-exit, test-sleeps, control-regex) tracked in #732 for follow-up. Full test suite green (0 failures); code-review verified all removals unused and all escape fixes semantics-preserving. Closes #732 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
11afca2968 |
feat(#656): Research module — content-addressed cache + provider seam + registry-API legitimacy (#664)
* feat(#656): add Research Store module (content-addressed cache, TTL staleness) Content-addressed research cache behind a clock seam: researchKey (sha256, deterministic), putResearch/getResearch ({hit,stale}, never throws), ttlForSource (curated HIGH 30d / MED 7d / web LOW 1d), two-tier resolveStorePath (curated -> ~/.gsd/research-cache, web/synthesis -> project .planning/research/.cache). 28 behavioral + property tests; boundary coverage at ttl-1/ttl/ttl+1. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): add Research Provider module (waterfall + confidence + plan) Single source of truth for the Balanced provider waterfall (docs Context7->Ref->Jina, web Exa+Tavily, fallback Perplexity/Brave, Firecrawl scrape-only). classifyConfidence stamps HIGH|MEDIUM|LOW by provider (never throws). providerAvailability maps config flags to usable providers. planResearch checks the Research Store (injected seam) and returns cache-hits + a per-question fetch plan, falling through the waterfall to the always-available websearch terminal. 22 behavioral + property tests. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): add Package Legitimacy module (registry-API verdicts, slopcheck optional) Replaces the pip-install-or-degrade slopcheck prose gate with code: classifyPackage (pure, never throws) computes OK|SUS|SLOP from tunable thresholds (minAgeDays 30, minWeeklyDownloads 1000, requireRepo). checkPackages queries injectable npm/PyPI/crates registry adapters (real https with 5s timeout, degraded-not-thrown on failure); slopcheck is one optional adapter that can only escalate severity, never degrade to [ASSUMED]. 34 behavioral + property tests; boundary coverage on age and downloads (limit-1/limit/limit+1). Known follow-up: real npm adapter must add api.npmjs.org last-week downloads fetch (currently null -> unknown-downloads). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): detect Tavily/Ref/Perplexity/Jina provider keys; complete npm downloads adapter config: add tavily_search/ref_search/perplexity/jina availability flags (env var or ~/.gsd/<x>_api_key), mirroring brave_search/exa_search/firecrawl, so the Research Provider waterfall can gate them. package-legitimacy: real npm adapter now fetches api.npmjs.org last-week downloads (bounded, degraded-not-thrown) so weeklyDownloads is populated. +12 config tests; 34 legitimacy tests unchanged. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): expose Research seam via gsd-tools query (research-plan, research-store, package-legitimacy) Routes the L2-hybrid surface so agents reach it as CLI: 'query research-store get/put' (cache, HOME-sandboxable), 'query research-plan --input' (cache-hits + fetch plan from planResearch), 'query package-legitimacy check --ecosystem' (async registry verdicts). Commands skip .planning root resolution and appear in top-level usage. 5 behavioral runGsdTools tests; command-contract unchanged (335). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#656): document Research module (CONTEXT predicates, ADR-0656, architecture, changeset) Adds GSD-RESEARCH.* + DEFECT.RESEARCH-PROVIDER-PROSE-DRIFT predicates to CONTEXT.md, ADR-0656 recording the L2-hybrid seam decision, a docs/ARCHITECTURE.md Research Module subsection, and an Added changeset fragment (pr:0, backfill on PR). Notes the #657 deferrals (agent collapse + install.js MCP mapping). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): sync inventory for research modules Regenerate INVENTORY-MANIFEST.json and bump docs/INVENTORY.md CLI Modules count 82->85 with rows for research-store/research-provider/package-legitimacy (DEFECT.INVENTORY-DRIFT). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): eslint-ignore generated research .cjs artifacts (ADR-457) research-store/research-provider/package-legitimacy .cjs are tsc-generated from src/*.cts, so they belong in the ESLint ignore block (lint the .cts source, not the emitted .cjs). Fixes tests/551-eslint-bin-lib-coverage. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): backfill changeset pr number to #664 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): satisfy eslint lint-tests gate Fix 20 eslint errors in the new research files: use helpers.cleanup() instead of raw fs.rmSync() in tests (local/no-raw-rmsync-in-tests, Windows-EBUSY retry budget); drop redundant '| string' union members and unnecessary type assertions; deterministic object normalization in researchKey (no-base-to-string). Logic unchanged; 6180 tests still green. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): harden package legitimacy per review (W1/W2/I3/I4) W1: httpsGet now reads statusCode; npm/PyPI/crates map 404 -> exists:false -> SLOP (registry-existence is the #1 slopsquatting defense; previously only npm caught it). Transport made injectable (_setHttpGet) for hermetic 404 tests. W2: suspicious-postinstall is now terminal SLOP independent of the optional slopcheck adapter, and the regex drops the bare https?:// arm (over-fired on esbuild/sharp/node-gyp) for shell-exec/download-exec signatures only. I3: checkPackages now threads version to registry.lookup and adapters verify that specific version exists. I4: moreServerVerdict -> moreSevereVerdict. +11 regression tests (all RED-first); 45 total green. Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): research-store tier coherence + freshness + version TTL (W4/I1/I2/I4) I1: tier now derives from source (curated -> user ~/.gsd, else -> project .planning), not kind, so put-tier and get-tier can't diverge; kind is a key component only. W4: getResearch searches both tiers and returns the freshest (non-stale preferred), never letting a stale curated entry shadow a fresh web one; blank version caps TTL at 1 day (no 30d on version-blind keys). I2: atomic platformWriteSync instead of raw fs.writeFileSync on the shared global path. I4: dropped the dead ttlForSource arm. CLI get now searches both tiers. +5 RED-first regression tests; 38 green. Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): expose classifyConfidence as a CLI route, killing dead code (W3) Adds 'gsd-tools query classify-confidence --provider X [--verified]' so research agents get the confidence tier FROM CODE (provider waterfall + verification lever) instead of asserting it in prose. classifyConfidence previously had no runtime caller. HIGH means 'trusted provider'; --verified raises web results to MEDIUM (verification semantics documented in ADR-0656). +4 behavioral tests. Addresses review by @davesienkowski on #664 (W3). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): close Codex adversarial-review findings (path-traversal, version-age, malformed-cache) HIGH: research key must be 64-hex sha256 (isValidResearchKey) + resolved-path containment check in put/get + CLI validation -> blocks '../../x' arbitrary-file-write. HIGH: package legitimacy now derives publishedAt from the REQUESTED version (npm time[version], PyPI releases[version] upload_time, crates versions[].created_at) so a new malicious version of an old package can't inherit old age and evade 'too-new'. MEDIUM: getResearch validates entry shape (finite fetched_at + positive ttl + required fields) -> malformed cache entry is a miss, not fresh-forever. +regression tests (RED-first); 111 green. Codex adversarial review (required pre-PR gate). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): close code-review correctness findings (1) package-legitimacy CLI now rejects unknown --flags instead of silently consuming the following package as a flag value; only --ecosystem takes a value. (2) crates recent_downloads (90-day) normalized to a weekly figure before the minWeeklyDownloads threshold (was ~13x too lenient). (3) research-plan --input validates parsed JSON is an object with an Array questions before destructuring -> clean usage error instead of an uncaught TypeError on null/bad input. (4) research-store put rejects a flag value that is itself a --flag (no more storing '--source' as content). (5) planResearch skips questions whose text is not a non-empty string instead of emitting question:undefined. +13 RED-first regression tests; 143 green. Code-review gate. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(#657): extract researcher documentation_lookup to shared @-reference 6 researcher agents carried a near-duplicate <documentation_lookup> block; consolidate into gsd-core/references/research-documentation-lookup.md (@-included). Unifies the ctx7 CLI fallback to the safer 'command -v ctx7' guard (drops silent 'npx --yes ctx7@latest' execution in 5 agents). Behavior-preserving dedup; inventory 63->64 references. Phase A of the agent collapse. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(#657): extract researcher philosophy + verification-protocol to shared @-references philosophy and the pitfalls+pre-submission-checklist common-core were near-duplicated in project/phase researchers; consolidate into gsd-core/references/research-{philosophy,verification-protocol}.md (@-included). phase-researcher keeps its 3 extra checklist items inline. Pre-submission domains checklist made agent-agnostic so project-researcher doesn't lose features/architecture coverage. Write-contract intentionally left inline (bug-214 tests assert it verbatim). Inventory 64->66 refs. Behavior-preserving. Phase A. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#657): wire gsd-phase-researcher to the Research seam (Phase B / S1) The phase researcher now CALLS the code seam instead of carrying inline mechanics: provider waterfall -> 'gsd-tools query research-plan' (+ research-store put to cache digests); confidence-tier prose -> 'gsd-tools query classify-confidence'; slopcheck pip-install protocol -> 'gsd-tools query package-legitimacy check'. This makes the Research module a real runtime consumer (validates the seam end-to-end, addresses reviewer S1) and removes the duplicated waterfall/confidence/slopcheck prose. RESEARCH.md output contract, commit step, structured returns, and Phase-A @-includes unchanged. package-legitimacy-gate.test.cjs rewritten prose-grep -> behavioral (asserts the seam invocation). Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#657): wire gsd-project-researcher to the seam + add tavily/ref/jina MCP tools (Phase C.1) project-researcher now calls gsd-tools query research-plan / classify-confidence (+ research-store put) instead of the inline provider waterfall + confidence-tier prose (mirrors the phase-researcher rewire; no package-legitimacy — phase-only). Output contract (STACK/FEATURES/ARCHITECTURE/PITFALLS/SUMMARY.md + sections, no-commit, structured returns, Phase-A @-includes) unchanged. Adds mcp__tavily/ref/jina__* to the project/phase/ui researcher tools frontmatter (Balanced provider set) so install.js MCP mapping (C.2) has a consumer. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#657): cover tavily/ref/jina MCP install handling + frontmatter parity guard (Phase C.2) Investigation: exa/firecrawl have no explicit per-runtime tool-mapping — every mcp__<server>__* except context7 rides the generic passthrough (Copilot lowercases; OpenCode/Cursor/Windsurf/Augment keep as-is; Gemini auto-discovers). tavily/ref/jina are handled identically, no install path broken. Added 12 copilot-install passthrough tests + a mcp-tool-inheritance parity guard (tavily co-declared with exa, jina with firecrawl, ref present across the 3 web researchers) so the MCP set can't drift. No io.github registry ids invented (none sourceable in-repo); documented as a follow-up. 488 tests green. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#657): profiles as source of truth for researcher agents + drift-guard (Phase C.3) scripts/research-profiles.cjs declares each of the 7 researcher agents' identity + contract (name, description, color, tools, required @-includes, required gsd-tools seam calls, output-contract markers). scripts/gen-research-agents.cjs --check validates every committed agent against its profile; --write regenerates ONLY the frontmatter from profiles (body untouched) and is a verified no-op against the current agents (zero diff = fidelity). tests/research-agent-profiles.test.cjs is the DEFECT.GENERATIVE-FIX drift guard. Design note: profiles govern the generatable/contract surface rather than destructively regenerating the disparate operational prose bodies (those were deduped via @-includes in Phase A). scripts/ is not inventoried (no inventory change). Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#657): complete agent provider-dispatch + parity guard; align legitimacy field; validate profiles Adversarial-review findings: (HIGH) the seam-wired agents' Step-C dispatch only mapped 6 providers, so a planResearch result of jina/ref/perplexity/brave (reachable via the waterfall fallbacks) had no handling -> agent stall; completed both agents' dispatch to all 9 PROVIDER_WATERFALL ids + a catch-all, and added a parity test asserting agent dispatch stays in sync with research-provider PROVIDER_WATERFALL (DEFECT.GENERATIVE-FIX). (MEDIUM) phase-researcher package-legitimacy JSON example used 'package' but the module returns 'name' -> aligned. (LOW) gen-research-agents checkAgent now returns a clear failure for a malformed profile instead of throwing. +parity/validation tests (RED-first). Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): make classifyConfidence verification-evidence-driven (W3) Confidence conflated provider authority with claim verification — context7/ref stamped HIGH purely by provider identity, and the only verification lever was a self-set --verified flag. Split into two axes: provider authority (static) + verification evidence (code-computed). HIGH now requires ground-truth corroboration (legitimacyVerdict OK), independent of provider; authority alone caps at MEDIUM; SLOP caps at LOW; the self-reported --verified is demoted to a MEDIUM-only web lever. HIGH = corroborated-against-authoritative-source, not a correctness guarantee. Adds --legitimacy-verdict to the classify-confidence CLI; updates CONTEXT.md predicate + ADR-0656 (tier set unchanged, ADR-consistent). Addresses davesienkowski's W3 review on #664. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#656): bind classify-confidence verdict to code, closing CLI self-grading Adversarial review found the new --legitimacy-verdict flag was caller-supplied, so an agent could self-assert OK->HIGH without any real legitimacy check — reintroducing the exact self-grading hole W3 closes. Remove the free flag; the CLI now computes the verdict via checkPackages only when --package/--ecosystem is given (code-computed, not agent-asserted). Update the stale CLI test (context7 alone -> MEDIUM) and extend the property test to vary legitimacyVerdict. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
463cffd894 |
chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/ Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary (`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers are unaffected. Mechanical (bulk, ~90% of the diff): - `git mv get-shit-done gsd-core` - Swept path/identifier references across the repo via `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead preserves the five legitimate slug variants that are NOT the directory: get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names). - Build/manifest wiring: package.json (bin, files, coverage globs), tsconfig.build.json (outDir), ~86 .gitignore build-output entries, stryker.config.mjs, scan-ignore files, install.js path strings. - Frozen (not rewritten): CHANGELOG.md history; translated docs (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/). New logic (review here): - src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper ADR-0008 installer migration. On upgrade it walks the legacy `~/.claude/get-shit-done/` tree, classifies each file via the prior install manifest, and emits remove-managed / backup-and-remove for managed files while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked root and symlinked entries; bounds-checks every path under configDir). The framework rolls back on install failure. Emptied dirs may remain (framework has no recursive dir-removal primitive) — documented. - scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare `get-shit-done` directory token (split token to avoid self-match; case- insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines). Wired into the lint-tests CI job. - Restored scripts/lint-package-identity-drift.cjs detection regexes (the mechanical sweep had wrongly rewritten the old-name patterns it exists to detect) and marked them as intentional legacy references. - TDD tests for the migration and the guard; do.md slash-command guard regex tightened so a `/gsd-core/bin` path segment is not mistaken for a command; changeset + docs/installer-migrations.md row added. Breaking: the installed runtime path moves `~/.claude/get-shit-done/` -> `~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed files (preserving user files) on upgrade. Users with custom hooks/configs hardcoding the old path must update them. Closes #604 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unsweep pending changesets + allowlist injection-example docs CI fixes for the rename PR: - Do not sweep pending .changeset/*.md (ephemeral release-note fragments, like CHANGELOG); reverted those body edits so 5 pre-existing malformed fragments (missing type/pr) no longer enter the PR diff and trip docs-lint. Allowlisted .changeset/ in the legacy-name guard accordingly. - Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in prompt-injection-scan.sh: they contain intentional injection examples / security-model prose; the path-reference rewrites are kept. CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR; none in the new migration/guard) and are out of scope for the rename. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): resolve CodeQL alerts surfaced on this PR The rename diff touched files carrying pre-existing CodeQL findings; per the no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving them off. All behavior-preserving: - scripts/ci-test-scope.cjs: build the config-path match from string .includes() instead of a RegExp over an arg-derived value (js/regex-injection). - src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName so the table-cell escape is complete (js/incomplete-sanitization). - tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization). - tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace, keep the meaningful POSIX-class conversion (js/identity-replacement). Verified: build:lib green; the touched test files + ci-test-scope + profile-output suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization) The prior commit's fixes for two alerts were ineffective: - ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file` reaching static regex `.test(file)` calls (not the config rule). Removed ALL regex over file/t — startsWith/includes/=== string checks + an isWindowsHint helper — so there is no regex sink for the tainted value. - js/incomplete-multi-character-sanitization (3 test files): a single `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint loop (replace until stable) plus a final bare-opener strip. Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL CodeQL flags the regex PATTERNS syntactically (regex-injection on the --files arg split; incomplete-multi-character-sanitization on the <!--...--> replace), so loop fixes do not satisfy it. Made these paths regex-free: - ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/). - 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)). Behavior preserved; ci-test-scope + the 3 suites pass; guard clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unblock security base64 scan on the large rename diff The security job hit its 10m timeout: base64-scan.sh choked on the binary test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/ non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings), and the ~800-file rename diff is slow to scan regardless. - scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they can't carry base64-obfuscated *text* and feeding NUL bytes through the per-line scanner is pathologically slow. collect_files already filtered binary *extensions*; this catches binary *content* in text extensions. - .github/workflows/security-scan.yml: raise the security job timeout 10m->30m to accommodate very large diffs (the scan itself is unchanged). Verified locally: scan skips the fixture, 0 "ignored null byte" warnings, 0 findings, exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): sweep get-shit-done refs introduced by merging next The branch was updated with next (#614/#384/#618 etc.), which reference the get-shit-done/ dir (still named that on next). Swept the stale references in the merged files to gsd-core so the rename stays consistent and lint:legacy-name passes: - commands/gsd/discuss-phase.md (runtime-launcher shim paths) - src/core.cts (getAgentsDir layout comments) - tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib) Verified: guard 0 violations; build green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant The #614 runtime-launcher shim added to discuss-phase.md references `${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it mis-read the directory path as a dangling `/gsd-core` command ref (same class as the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path segments are not treated as slash-command references. Verified locally on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22 image) full suite: 0 failures - bug-3683 + bug-2954 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI) CI intermittently failed state.test's gsd-tools subprocess with "findProjectRoot is not a function" (flip-flopping across legs; not reproducible on mac full suite, gsd-test linux full suite, test:unit, or state.test x8). findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs); binding it via destructure at module-load can be undefined under a load-ordering edge. Resolve it lazily at call time via a small wrapper so the lookup happens after core.cjs is fully initialized. Verified green on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22) full suite: 0 failures - state.test.cjs: 106/106; gsd-tools loads cleanly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): allowlist verification-patterns.md placeholder examples in secret scan The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var examples (illustrative Stripe test-key / database-URL / API-key placeholders) — not real credentials. Added it to .secretscanignore with the strict annotation, mirroring the existing gsd-core/workflows/plan-phase.md exception. Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next exits 0 with 0 findings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |