bd613566cbba1e6808329c97c08946cac371c7b9
276 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bd613566cb |
feat(#2100): drive Windsurf through the EoS descriptor + wire Cascade's blocking hook bus (ADR-1239)
Fold all 10 residual isWindsurf branches in bin/install.js onto descriptor-driven hostBehaviors (byte-parity — no fold changes any install output): - 2 dead destructures dropped (uninstall, finishInstall); the dead `else if (isWindsurf)` legacy agent-loop arm removed (windsurf ∈ _DESCRIPTOR_AGENTS_RUNTIMES → unreachable). - skipSharedHooksInstall:true folds the two `!isWindsurf` shared-hooks exclusions. - legacyDevinSkillsCleanup:true folds the `.devin`→`.windsurf` one-time cleanup gate. - installsCommandBodiesForWorkflowDelegation:true folds the #1629 command-body copy (workflow-delegation target — load-bearing; local-install verified intact). - verificationStyle:"windsurf-workflows" folds the workflow-count report. - Corrected stale _LEGACY_SCAN_SUBDIR_NAMES + hooks-json manifest comments (cursor + windsurf). Zero live runtime==='windsurf'/isWindsurf branches remain across bin/install.js, install-engine.cts, surface.cts, runtime-artifact-conversion.cts (AC2 guard scans all four). UPGRADE (Cascade hook bus): wire GSD's write/command safety guards into Windsurf's native hook bus. New hooksSurface 'windsurf-hooks-json' (VALID_HOOKS_SURFACES 7→8, GATE A profile-marker-only allowlist, the HooksSurface union) + writeWindsurfHooksJson (Cursor-templated, Cascade's flat {hooks:{<event>:[{command}]}} shape) writing .windsurf/hooks.json with two BLOCKING pre-hooks: - pre_write_code → gsd-windsurf-pre-write.js: blocks writes to a file outside the active git worktree / into .git internals. - pre_run_command → gsd-windsurf-pre-command.js: conservative destructive-command deny-list (rm -rf of root/home incl. sudo/env/path-prefixed forms; fork bombs; force-push refspec forms — HEAD:main, +main, --force/-f — to main/master/next). Both use Cascade's protocol (stdin JSON, exit 2 + stderr to block, exit 0 to allow, fail-open on error/timeout). Tokenize-based classifier (no catastrophic-backtracking regex; 4096-char cap) with the fail-closed false-positives fixed post-review. The 4 advisory GSD guards + pre_mcp_tool_use + 5 post_* logging events are deliberately NOT wired: Cascade has no context-injection channel for advisory hooks and GSD has no MCP guard — porting them would be non-functional padding (documented; codebuddy #2098 / copilot #2099 faithful-subset precedent). extendedHookEvents stays []. Golden: the 2 guard scripts ship in the shared hook bundle (HOOKS_TO_COPY + the shared managed-hooks-registry), exactly like cursor's 6 gsd-cursor-*.js scripts — so the 8 shared-bundle runtimes' fixtures gain the 2 inert windsurf scripts + the registry hash (functionally inert for non-windsurf; the established cursor pattern). No install-output change beyond that (the folds are byte-parity; skip-bundle runtimes untouched). New scripts registered in managed-hooks-registry + build-hooks + INVENTORY. Tests: declarative-reference- windsurf (adapter/axes/fail-closed + AC2 guard) + windsurf-hooks-bridge (live exit-2 blocking + allow/fail-open + ReDoS-bound + writer/reconcile/remove idempotency); VALID_HOOKS_SURFACES pin updated to 8. Matrix hookBus delta + changeset (Changed). capability-registry regenerated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ab04916682 |
feat(#2095): migrate Kimi CLI onto EoS imperative adapter + native hook-bus + background dispatch (ADR-1239)
Fold all runtime==='kimi'/isKimi logic branches into descriptor-driven hostBehaviors (localInstallDeferred, verificationStyle, agentManifestStyle, reapplyCommand, doneBannerStyle) + add 'kimi' to _DESCRIPTOR_AGENTS_RUNTIMES. Kimi's skills/kimi-agents dispatch was already descriptor-driven (converter-by- name + kimi-agents kind). Zero isKimi/runtime==='kimi' branches remain. UPGRADE 1 (native hook bus): new hooksSurface 'kimi-hooks-toml' + a marker- delimited config.toml [[hooks]] emitter (buildKimiHooksTomlBlock/writeKimiHooksToml in runtime-hooks-surface.cts; resolveKimiHooksTomlDir in runtime-homes.cts). GSD's lifecycle hooks now wire into Kimi's native ~/.kimi/config.toml (Context7- confirmed path) at SessionStart/PreToolUse/Stop/PreCompact/SubagentStart/ SubagentStop — kimi becomes a hooks/ consumer (the 3 && !isKimi exclusion guards removed). config.toml holds absolute install paths so it's golden-excluded via an exact relative-path (.kimi/config.toml), not a basename (which would blind Codex's config.toml). New hooksSurface value added to the closed enum in capability-validator + runtime-config-adapter-registry. UPGRADE 2 (background dispatch): flip dispatch.backgroundDispatch true (Kimi's Agent tool takes run_in_background; root agent already gets the Agent tool), so negotiation no longer flattens dispatch. subagentToolkit stays 'undocumented' per AC (coder/explore/plan have distinct tool policies). MCP transport explicitly deferred (no installer-driven MCP for any runtime). Golden: only kimi.json changes (hooks/ scripts now installed); all 15 others + claude-local byte-identical (kilo/zcode keep their own exclusions). Tests: kimi-imperative-reference (adapter/axes/fail-closed/hostBehaviors + source-grep guard) + kimi-upgrades (config.toml [[hooks]] SessionStart + marker idempotency + backgroundDispatch negotiation). CONTEXT.md glossary + matrix + how-to updated; changeset (Added). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a1de52d71b |
fix(#2128): sanctions must be a dedicated // comment line (decoy-proof)
Re-review found the `// phase-id-owner:` suppression treated a `//` embedded in a string literal as a comment — help/doc text quoting the sanction syntax (the exact string the scanner's own main() prints) would silently suppress a real re-derivation. Require the marker to LEAD its own comment line (`^\s*//…`), so a `//` inside a string or trailing a code line never counts. All 5 real sanctions are already dedicated lines (scanRepo stays green); trailing same-line sanctions are no longer honored — put the comment on the line directly above. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e2eaa5b046 |
fix(#2128): address review — migrate 9 mis-allowlisted sites, harden scanner + guards
Correctness review of the Phase 4 guard found the allowlist over-broad and the scanner/guards evadable. Fixed all findings: - Migrate 9 sites that were wrongly sanctioned: their regex is the PURE canonical token (`\d+[A-Z]?(?:\.\d+)*`, no variant), byte-identical to already-migrated siblings. The old justification argued against swapping to the extractPhaseToken() FUNCTION (behavior-risky) — but the guard only wants the same regex built from the SOURCE string (byte-equal, zero risk). Coverage is now 32 migrated / 5 sanctioned, not the overstated 23 / 14 (audit.cts x3, uat.cts, init.cts x4, roadmap-upgrade.cts). Each conversion proven byte-equal (.source + .flags). - Harden the drift detector: also catch the `[0-9]`-in-place-of-`\d` variant; document the accepted limits (cross-line split, semantic restructuring — covered by the identity guard + review, not a text scan). - Sanction robustness: a `phase-id-owner:` marker now counts only inside a `//` comment (a bare substring in a string no longer suppresses a real flag), and the preceding-line window skips blank lines (an auto-formatter's blank line no longer reactivates the flag). - roadmap-parser.cts:462 comment: corrected — that regex carries no /i flag, so its [A-Za-z] class does real case work (matches state.cts:1409's rationale). - Identity guard: surface require failures instead of silently skipping, and floor coverage at >75% of consumer modules (inspects 156/157). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
09be501eb7 |
feat(#2128): phase-id anti-divergence guard — canonical token source + drift scanner + guards
Phase 4 of epic #2121 (ADR-2121 Decision 7), closing the recurrence loop that produced #2111 / #2114 / #2104: no module outside src/phase-id.cts may re-implement phase-ID parsing without failing CI. - phase-id.cts: add PHASE_NUMBER_TOKEN_SOURCE — the canonical phase-number-token grammar (\d+[A-Z]?(?:\.\d+)*) for enumeration/scan call sites, the ANY-phase counterpart to phaseMarkdownRegexSource(n)'s known-number lookup. Extend-only (never touches normalizePhaseName; blast radius 79 fns / CRITICAL). - scripts/lint-phase-id-drift.cjs: pure findPhaseIdRegexDrift(text) + scanRepo(root), wired to `npm run check:phase-id-drift`. Flags a literal re-derivation of the canonical token (both /\d/ and new-RegExp `\\d` escaping, plus the [A-Za-z] and [.-] near-variants) anywhere in src/** outside phase-id.cts, unless sanctioned with `// phase-id-owner: <reason>`. Narrow by design: bare \d+, digits-only captures, \w ids, status-message text and pipe-tables are not flagged. - tests/phase-id-drift-guard.test.cjs: fail-first drift cases (AC1) + live scanRepo(ROOT) zero-drift (AC3) + identity guard — phase-id.cjs exports the complete locked surface and no consumer re-exports a divergent copy (AC2). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
24896ddac7 | fix(#2089): register 4 new cursor hook scripts in build + managed-hooks whitelists | ||
|
|
8b99f4f3c3 |
fix(#2088): weight install-heavy test files so they spread across chunks
The targeted CI lane runs changed files UNSHARDED; #2088 touched 13 install-heavy test files that all landed in one chunk, blowing the 600s per-chunk backstop on the slow Windows runner (pure slowness, not a leak — per run-tests.cjs's own comment). Weight install*/codex-* files (~10x a unit file) toward the per-chunk budget so they spread across chunks instead of clustering; light-file chunking is unchanged (weight 1). Adds harness regression tests (heavy split vs light control). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
080bacdb4b | Merge branch 'next' into codex/gsd-onboard | ||
|
|
97972ca7fd |
fix(#1575): lower MAX_FILES_PER_CHUNK from 90 to 60 to fix macOS Node 22 timeout
Shard 2/3 chunk 2 (~80 files including state.test.cjs, perf-*, worktree-cleanup) exceeded the 600s per-chunk timeout on macOS Node 22. Reducing the cap from 90 to 60 splits this into two ~40-file chunks, each well within the 600s budget. Three chunks at ~5 min each = ~15 min, safely under the 20m job cap. |
||
|
|
4ceedafe9f |
chore(#1925): generalize golden-install-parity generator to all runtimes + args
Support regenerating a single runtime (node scripts/gen-golden-install-parity-zcode.cjs zcode) or all runtimes (no args). Needed when a shared payload file changes and every fixture must be recaptured. |
||
|
|
7acbdc71b8 |
test(#1925): register zcode across installer surfaces + fluidify count pins
Resolve the zcode test cascade exposed by gsd-test: - model-catalog.json: add zcode runtimeTierDefaults (null tiers, matching the other profile-marker-only runtimes) so KNOWN_RUNTIMES stays parity with allRuntimes. - runtime-name-policy: add zcode to GLOBAL_CONFIG_HOME_FRAGMENTS (~/.zcode) so getGlobalConfigHomeFragment stops falling through to the .claude default. - installer-migration-install: add the zcode fresh-install contract (flat-skills, no settings, no package.json — same shape as trae). - golden-install-parity: capture the zcode fixture via a standalone generator script (not node --test — the gate stays gsd-test). - capability-matrix.md: regenerate so zcode appears as a first-party row. Fluidify the remaining count-pinned guards so adding a runtime no longer trips a hand-pinned snapshot: gemini-runtime-removed (flag count derives from the registry), non-claude-runtimes-registry-derivation (golden list derived from the registry). |
||
|
|
68a5258d45 |
fix(#2017): grant mcp__plugin_context7_context7__* for plugin-marketplace context7 (8 agents) (#2029)
* fix(#2017: grant mcp__plugin_context7_context7__* for plugin-marketplace context7 The 8 context7-using agents granted only mcp__context7__* (standalone server form). Claude Code's plugin-marketplace context7 install names tools mcp__plugin_context7_context7__*, so the grant never matched and every researcher/planner/executor silently lost doc lookup (fell back to WebSearch). - 8 agents: add mcp__plugin_context7_context7__* alongside mcp__context7__*. - scripts/research-profiles.cjs: update the researcher profile tools to match. - tests/context7-plugin-grant-parity.test.cjs: regression guard — no agent grants the standalone form without the plugin form. Closes #2017 * docs(#2017): backfill changeset pr 2029 |
||
|
|
425d2f2c8b |
fix(#1990): keep onboard resolver delegation in sync with canonical launcher
onboard.md delegates the gsd_run preamble to
gsd-core/references/gsd-run-resolver.md via @-include, but two guards
regressed once the canonical launcher snippet advanced to
${CLAUDE_CONFIG_DIR:-$HOME/.claude} (#2024):
- runtime-launcher-parity (B2): the resolver reference still shipped the
old $HOME/.claude arm. references/ is not covered by
sync-runtime-launcher.cjs, so refresh the resolver bash block to be
byte-equal to _runtime-launcher.snippet.sh.
- /gsd:onboard command contract: sync-runtime-launcher.cjs had re-inlined
the preamble into onboard.md (a delegating file). Teach the sync
transform to strip-but-never-inline files that @-include the resolver,
mirroring the exemption already in the parity test (B/B2).
Regenerate golden install fixtures and the workflow size baseline for the
smaller onboard.md.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
|
||
|
|
bd77b40107 |
feat(#1825): configurable graphify graph location (graphify.graph_path) (#2013)
* feat(#1825): configurable graphify graph location (graphify.graph_path) Add a graphify.graph_path config key (.planning/config.json) that overrides where /gsd-graphify query|status|diff read the knowledge graph, so one curated umbrella-level cross-repo graph can serve multiple sibling projects without N drifting ~5 MB mirror copies. Previously the graph location was hardcoded to <cwd>/.planning/graphs/. - src/graphify.cts: resolveGraphLocation(cwd, planningDir) honors the key (resolved relative to project root; absolute paths honored via path.resolve); falls back to the historical .planning/graphs/graph.json when unset/blank/ non-string (byte-identical). Wired into graphifyQuery, graphifyStatus, graphifyDiff (snapshot travels with the configured graph via dirname), and writeSnapshot. Configured-but-missing -> actionable error naming the path. Build stays project-scoped (skill hardcodes the cp dest); umbrella graph is built in the umbrella project, sub-projects only READ it. - config-schema.manifest.json: register graphify.graph_path in validKeys. - tests/graphify-graph-path.test.cjs: boundary matrix (unset byte-identical, set+present reads configured graph not default, set+missing actionable error, relative resolved vs project root, blank treated as unset, snapshot alongside configured graph, diff from configured dir, build project-scoped) + VALID_CONFIG_KEYS registration. - docs: CONFIGURATION.md row, FEATURES.md REQ-GRAPH-06, CONTEXT.md module note, .changeset (Added). Closes #1825 * docs(#1825): backfill changeset pr number 2013 |
||
|
|
97730e59a1 |
fix(#1164): wire external-job config keys + document wave:post choice (#2006)
Refinements A-D to the external-job capability (PR #1998 follow-up): A. Document why the contribution registers at execute:wave:post: #1164 asks for wave:pre, but execute-phase.md only dispatches wave:post today (wave:pre is declared in the loop host contract but not rendered). Wiring wave:pre is a core-loop change #1164 puts out of scope; the executor honors the runtime_budget classification guidance before running any tagged task. B. external_job.artifact_dir is now consumed (was declared but unused): the adapter resolves it via the canonical capability-config seam and surfaces the resolved root in submit output. C. external_job.submit_timeout_ms / poll_timeout_ms are now read from config (were shadowed by env-only reads). Precedence: env > config > registry default; non-numeric config values fall back (no guessing, no NaN). D. CLI surface gains unit coverage: parseFlags, findPlanningDir, resolveExternalJobSettings, formatShowReport. Regenerates capability-registry.cjs from the updated capability.json. |
||
|
|
8f2ebbe9bf |
feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap). --gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): backfill changeset PR number (#1996) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): drop Gemini CLI from issue templates (review nit) Removes the sunset Gemini CLI runtime from the two GitHub issue-template runtime lists that the removal PR missed, per @davesienkowski's review nit: - feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise request a feature for a runtime GSD no longer supports) - bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json retrieval-help line Leaves the post-removal templates fully consistent with the Antigravity redirect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5ae4ea4c84 |
feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) (#1998)
* feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) The async external-job consumer half (#1165) shipped long ago: the core loop reads .planning/async-jobs/<job>.json manifests and treats a non-terminal one as the legal external_job_waiting half-state. The PRODUCER half (#1164) was the remaining unimplemented piece of #1105. This adds the producer as a default-off capability: - capabilities/external-job/ — capability.json (execute:wave:post -> executor, plan:post -> planner contributions, external_job.* config keys, default-off) + fragments teaching runtime-budget classification and externalization. - src/external-job.cts -> gsd-core/bin/lib/external-job.cjs — pure producer module: SLURM state -> manifest-status map (no guessing), manifest build/validate (versioned stability contract), sbatch/squeue/sacct parsers, and a fail-closed manifest writer (refuses a second non-terminal job for a plan_id already in flight; refuses to clobber a malformed manifest). fs/clock seams for deterministic tests. - scripts/slurm-adapter.cjs — operator CLI (submit/poll/show) wrapping bounded sbatch/squeue/sacct subprocesses; surfaces manifest commands for confirmation and never auto-runs them (trust boundary). - tests/external-job.test.cjs — 23 behavioral + fast-check property tests. - docs/reference/long-running-operations.md + docs/how-to/async-external-jobs.md. - CONTEXT.md glossary entry for the External-job Capability. - Regenerated capability-registry.cjs; pruned the now-stale test-file-count allowlist entry (external-job is at the 2-file cap). * chore(#1105): backfill PR number in changeset * fix(#1105): sync capability artifacts + update registry shape-pin tests gsd-test caught that adding the external-job capability requires its dependent artifacts regenerated and its registry-shape drift absorbed: - sync-manifest-versions: stamp 1.7.0-rc.2 into capability.json (was 1.0.0). - gen-capability-matrix --write: regenerate docs/reference/capability-matrix.md. - gen-inventory-manifest --write: regenerate docs/INVENTORY-MANIFEST.json. - check-gap-analysis-plan-post-e2e: plan:post now has 1 contribution (external-job planner fragment) instead of 0. - execute-wave-post-gate-pipeline-e2e: execute:wave:post now has 2 contributions (mempalace + external-job) instead of 1. * fix(#1105): regenerate capability-registry after version stamp sync-manifest-versions re-stamped external-job/capability.json from 1.0.0 to 1.7.0-rc.2 after the last registry regeneration, leaving the committed capability-registry.cjs stale (CI gen-capability-registry --check failed). gsd-test masked this because its setup runs the full 'npm run build' (which regenerates the registry); CI's 'npm test' pretest only runs build:lib. |
||
|
|
b0bd2f7a48 |
chore: move committed-generated-artifact freshness checks to lint:ci (#2000)
gsd-test's build leg runs the full 'npm run build' (which regenerates capability-registry.cjs, loop-host-contract.cjs, package-identity.cjs, etc.), so committed-freshness guards that lived in the unit suite were masked there: gsd-test passed a stale-commit that CI's shard-1/3 test then red-flagged (caught live on PR #1998). The mandated pre-push gate was green on a commit CI correctly flagged. Move the committed-state --check guards into a new 'lint:generated-sync' script wired into lint:ci (the single orchestrated entry point the lint-tests CI job already runs on a build:lib-only tree, so the committed artifacts are checked without regeneration). gsd-test no longer contains these guards, so it can no longer mask them. - package.json: add lint:generated-sync (7 generators --check); wire into lint:ci. - generate-package-identity.cjs: add --check mode (was the only generator without it); no-arg behaviour unchanged (still writes, as build expects). - Remove the committed-freshness guards from the unit suite, keeping all behavioral/structural tests: - capability-registry.test.cjs: drop the --check describe. - loop-host-contract.test.cjs: drop the committed-file staleness test (keep the normalizeLineEndings unit test). - capability-matrix-sync.test.cjs: drop --check + byte-for-byte (keep the architectural content invariants: every cap appears, security ship:pre). - issue-844-manifest-version-sync.test.cjs: drop describe D (--check). - issue-498-package-identity.test.cjs: drop the drift-check test (keep behavioral module-export tests); drop the now-unused render import and its allow-test-rule exemption (allowlist ratcheted 175 -> 174). |
||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
241b08fa18 |
ci(#1975): exclude release-tarball-smoke from scoped lane (fixes Windows chunk timeout)
This consolidation PR's breadth (28 changed test files) exposed a scoped-test-lane capacity limit: ci-test-scope pulls the 3–6 min release-tarball-smoke.install.test.cjs (npm pack + npm install -g, 10MB/1499 files) into the targeted+windows lane whenever install files change AND when it is itself a changed file — bundling it with the other 27 files overran the 600s per-chunk timeout on windows-latest-24 (deterministic). release-tarball-smoke has its OWN dedicated workflow (.github/workflows/install-smoke.yml, triggered on the production install paths), so its scoped-lane run is redundant. Add a SCOPED_LANE_EXCLUDE guard that drops it from both targeted_tests and windows_tests however it entered (matched rule OR changed-file), and remove it from the install rule's tests list. Update the ci-test-scope.test.cjs assertion accordingly. No coverage lost. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6d072435d0 |
test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks, read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim block-scoped describe wrappers; 427 subtests conserved 1:1. Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec, so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard. Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6, docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md (EN + ja/ko/pt/zh) and ADR-0002. lint:ci green. Part of epic #1969. Closes #1975. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
697cbb1f05 |
test(#1977): consolidate 22 misc + repo-invariant regression tests
Final epic-#1969 batch. Fold 22 issue-named files: the 4 genuine repo-wide invariant scans (551-eslint-bin-lib-coverage, bug-3054 stale /gsd-next, bug-3810 no-gsd-sdk-runtime-refs, feat-3593 cli-negative-universal) into a NEW shared repo-invariants.test.cjs; the other 18 as singletons into their nearest module suite (model-resolver, codex-config, runtime-converters, security, state-transition, worktree-safety, roadmap-parser, etc.). Verbatim block-scoped describe wrappers; 334 subtests conserved 1:1. Host-env pre-check (B2+B6): the 6 CLI folds into GSD_TEST_MODE-setting hosts (model-resolver/ codex-config/runtime-converters) are benign — each origin independently sets GSD_TEST_MODE=1 itself (idempotent), unlike the B6 real-install case. Regenerates regression-name allowlist (222->213), ratchets file-count allowlist (state 17->16), makes 7 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids). Repoints 2 tests/ refs in docs/TESTING-SUITES.md. lint:ci green. Part of epic #1969. Closes #1977. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cb7ca4f3dc |
test(#1976): consolidate 17 agent regression tests into agent-contract suites
Fold 17 issue-named agent-markdown regression files into their canonical agent suites: agent-frontmatter (shared agent-body-contract catch-all: read-loop guards, retired-slash-ref bans, sycophancy hardening, doc-writer/code-fixer/review-fix/mapper contracts, write-truncation), planner-language-regression (planner directive/grep/phase contract), executor-mvp-tdd-section, verifier-behavior-unverified, intel, and research-agent-profiles. Verbatim block-scoped describe wrappers; 98 subtests conserved 1:1. All unit (markdown-content) tests — no CLI/env surface. No new files. Regenerates regression-name allowlist (222->207), ratchets file-count allowlist (intel entry removed, phase 4->3), makes 13 relocated allow-test-rule exemptions issue-ref- compliant (ADR-456; prunes stale ids). No CONTEXT.md/ADR refs named the moved files. lint:ci green. Part of epic #1969. Closes #1976. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0cc7a1a426 |
test(#1974): consolidate 27 installer/hooks remainder tests into module suites
Fold 27 issue-named installer/hooks/statusline/migration/reapply regression files into their canonical module suites (installer-migrations, installer-migration-report, gsd-statusline, reapply-verify-hunks, install-*, gsd-check-update-worker-platform-gate, etc.). Verbatim block-scoped describe wrappers; 276 subtests conserved 1:1. No new files. The one subdir origin (tests/installer-migrations/001-legacy-orphan-files) moved up one level into installer-migrations.test.cjs; its single ../../ module require corrected to ../ so it resolves from tests/ root (verified). Host-env pre-check: no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. Regenerates regression-name allowlist (222->205), ratchets file-count allowlist (verify 11->8, validate entry removed), makes 16 relocated allow-test-rule exemptions issue-ref- compliant (ADR-456; prunes stale ids). Repoints 15 tests/ references across state-md.md (EN + ja/ko/pt/zh). lint:ci green. Part of epic #1969. Closes #1974. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2c32e8d890 |
test(#1973): consolidate 48 workflow regression tests into workflow-aspect suites
Fold 48 issue-named workflow-markdown regression files into the canonical test that owns each workflow aspect (execute-phase-*, plan-phase-*, quick-*, discuss-*, worktree-cleanup, secure-phase, verify, update, settings, etc.), across 32 existing suites. Verbatim block-scoped describe wrappers; 281 subtests conserved 1:1. No new test files. Host-env pre-check: only gsd-settings-advanced spawns CLI and it sets no GSD_WORKSTREAM/GSD_PROJECT value — no leak risk. Regenerates regression-name allowlist (222->181), ratchets file-count allowlist (verify 11->10), makes 30 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 30 stale ids). lint:ci green. Part of epic #1969. Closes #1973. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
85ed50cc4f |
test(#1972): consolidate 94 command/module regression tests into subject suites
Fold 94 issue-named command/module regression files into the canonical test file that owns each subject-under-test, across 52 existing suites (state, config, frontmatter, roadmap-parser, capability-registry, shell-command-projection-dispatch, plan-phase-drift-guard, health-validation, runtime-converters, commands, etc.). Verbatim block-scoped describe wrappers; 881 subtests conserved 1:1. No new test files. Host-env pre-check (per B2): the only GSD_WORKSTREAM/GSD_PROJECT-touching destinations (intel, planning-workspace) clear those vars hermetically, so folded CLI tests are safe. Regenerates regression-name allowlist (222->162), ratchets file-count allowlist across 8 buckets (validate entry removed after dropping <=2), makes 34 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 34 stale ids). Repoints CONTEXT.md + ADR-0002/443/1235/3524 test-file references. lint:ci green. Part of epic #1969. Closes #1972. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
de3ba45d00 |
test(#1971): consolidate 48 gsd-tools CLI regression tests into subcommand suites
Fold 48 issue-named gsd-tools CLI regression files into the canonical test file that owns each subcommand subject (state, roadmap, phase, milestone, audit, config, router/dispatch, stats, verify, health, etc.), preserving every assertion and its origin issue number as provenance (block-scoped describe wrappers, 299 subtests conserved 1:1). No monolithic gsd-tools.test.cjs created — routes into 18 existing per-subject suites. Removes 48 tests/ files. Regenerates regression-name allowlist (271->231), ratchets the file-count allowlist across 6 buckets (audit/milestone/phase/roadmap/state/verify), and makes 10 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 10 stale ids). Repoints one CONTEXT.md symptom ref and ADR-3524's parity-test ref. lint:ci green. Part of epic #1969. Closes #1971. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4f779eda43 |
test(#1970): consolidate 61 install-suite regression tests into function suites
Fold 61 issue-named install/codex/runtime regression files into the canonical test file that owns each subject-under-test, preserving every assertion and its origin issue number as provenance (block-scoped describe wrappers, zero assertion loss — 652 subtests conserved 1:1). Routes: - codex-config.test.cjs +19 (codex config/toml/hooks/adapter/skill surface) - install.test.cjs +18 (node-runner norm, manifest, arg parse, finishInstall) - install-runtime-artifacts +13 (per-runtime conversion + emission) - install-minimal-hooks +6 (hook-event dialects + guards) - path-replacement +2 (opencode absolute pathPrefix) - install-write-confinement +2 (pristine dir writes) - install-regressions +1 (user-artifact preservation) Removes 61 tests/ files → 61 fewer CI processes. Regenerates the regression-name allowlist (271→222), ratchets the file-count allowlist (config 10→9, install 12→9), and makes 13 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 13 stale allowlist ids). Repoints ADR-0009's moved-test list. lint:ci green. Part of epic #1969. Closes #1970. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5b5dbe059b |
feat(#1942): VS Code extension — repo-local IDE host with reachability proof (#1966)
* feat(#1942): VS Code extension — repo-local IDE host with reachability proof The VS Code extension (vscode/extension.js + vscode/package.json) is a repo-local, buildable extension (not Marketplace-published). The gsd.invoke handler dispatches in-process through the GSD command-routing hub. The handler is exported separately from activate() so it is testable without a VS Code host. tests/vscode-extension-reachability.test.cjs: proves the handler dispatches through the hub + returns a result (keystone wired); the manifest declares the command + engine; resolveEngineRoot finds the gsd-core/ dir. * docs(changeset): VS Code extension (#1966) * fix(#1942): register vscode/package.json in VERSIONED_MANIFESTS (#844) * fix(#1942): align vscode/package.json version with package.json (1.7.0-rc.1) |
||
|
|
b188c0d085 |
fix(#1967): build hooks/dist once upfront in run-tests to close scoped-CI empty-dir race (#1968)
* fix(#1967): build hooks/dist once upfront in run-tests to close scoped-CI empty-dir race hooks/dist/ is gitignored and not built by prepare (build:lib only), so the scoped CI lane starts with it absent. The first install test's before() hook triggers build-hooks.js, which creates DIST_DIR empty then fills it file-by- file; a concurrently-spawned install.js reader can observe the empty window and fail with 'Failed to install hooks: directory is empty' (intermittently failing e.g. bug-3683-workflow-colon-namespace-leak on scoped legs). Add ensureBuiltHooks() to scripts/run-tests.cjs — the same upfront chokepoint as ensureBuiltArtifacts — to build hooks/dist once, single-process, before any concurrent test spawns install.js. Completeness is checked against build-hooks.js HOOKS_TO_COPY (absent/empty/partial/zero-byte -> rebuild; complete -> no-op). Folds regression coverage into bug-969-test-infra-flake-hardening (Part C), proven fail-first (ensureBuiltHooks undefined on next). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1967): set changeset pr to 1968 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2391632973 |
feat(#1855): add Claude plugin marketplace manifest
Add .claude-plugin/marketplace.json so Claude-plugin-compatible runtimes
(ZCODE et al.) discover gsd-core from a custom marketplace source. The
canonical version lives at plugins[0].version and tracks package.json via
the release version-sync.
Refactor scripts/sync-manifest-versions.cjs so VERSIONED_MANIFESTS entries
are {path, versionKey} dot-path descriptors (default 'version'); register
marketplace.json with versionKey 'plugins.0.version'. getByPath/setByPath
reject __proto__/constructor/prototype (prototype-pollution guard).
plugin.json / gemini-extension.json behavior is unchanged.
- tests/issue-1855-marketplace-manifest.test.cjs: schema + version-sync guard
- tests/issue-844-manifest-version-sync.test.cjs: updated for descriptor shape
- VERSIONING.md + auto-backmerge VERSION_STAMP_MANIFESTS: include marketplace.json
- docs/how-to/install-on-your-runtime.md: marketplace discovery how-to
|
||
|
|
d8bc1bc636 |
fix(#1831): add state-rebuild test files to lint-test-file-count allowlist
PRs #1829 and #1830 added tests/state-rebuild.test.cjs and tests/state-rebuild-cli.test.cjs but did not add them to the lint-test-file-count allowlist for the 'state' module. The lint-test-file-count.test.cjs harness reported FAIL_NOVEL_FILES with the two files listed as novel. Verified via gsd-test (--base origin/main --head origin/next): without this fix, 2/22097 tests fail (both lint-test-file-count allowlist cases). The state module has 17 test files; allowlist entry now lists all 17 alphabetically. Process note: the original PRs should have updated this allowlist in the same commit that added the test files. Recapture of the golden-install-parity fixtures is NOT required — those fixtures on next are correct (verified: no diff from UPDATE_GOLDEN=1 locally). |
||
|
|
79c69d443b |
refactor(#1793): ADR-1769 Phase 7 — sync, prune, update migrations (#1760, #1761) (#1794)
Migrate cmdStateSync, cmdStatePrune, cmdStateUpdate (state.cts) onto the STATE.md Transition Module substrate and close the maintenance bug pair #1760/#1761 (ADR-1769, epic #1769). Completes the substrate: all 10 lifecycle/maintenance transitions now route through transitionCore. - Add {kind: 'sync'|'prune'|'update'} to StateTransitionIntent with syncCore, pruneCore, updateCore in src/state-transition.cts. - updateCore: body-only single-field update (strip/reassemble), mirroring the pre-migration cmdStateUpdate contract. - pruneCore: pure content→content section pruning (Decisions / Recently Completed / resolved Blockers / Performance Metrics rows at or below cutoff), byte-identical tokenizeHeadings splicing. Adapter owns currentPhase, dry-run, and STATE-ARCHIVE.md. - syncCore: body writes (Total Plans in Phase, Progress bar, Last Activity) given disk-derived numbers. Adapter owns the disk scan + roadmap scope. - #1760: cmdStatePrune now derives currentPhase from 'Current Phase' OR 'Phase' (the canonical template emits 'Phase: X of Y'), so prune engages on template-conformant STATE.md instead of bailing 'Only 0 phases'. - #1761: cmdStateSync skips the Progress write (percent=null) when a milestone version is set in frontmatter but the ROADMAP has no versioned heading for it (milestone cannot be bounded). Projects without a milestone version are unaffected. Leaves progress untouched rather than silently writing fallback- derived wrong values. - Regression: bug-1760 (prune engages on template field) + bug-1761 (sync leaves progress untouched when unbounded). ADR-1769 marked Accepted. All 82 transition + 515 regression tests pass. Closes #1793 |
||
|
|
064f63b299 |
refactor(#1791): ADR-1769 Phase 6 — patch migration + curated current_phase_name preserve (#1695) (#1792)
Migrate cmdStatePatch (state.cts) onto the STATE.md Transition Module substrate and close the curated-field clobber bug class #1743/#1695 (ADR-1769, epic #1769). - Add {kind: 'patch'} to StateTransitionIntent, with patchCore in src/state-transition.cts. Applies each caller-supplied {field:value} pair via stateReplaceField over the full content (body + frontmatter), tracking updated vs. failed. data.updated/data.failed mirror the CLI output shape. - Collapse cmdStatePatch to a transitionCore dispatch. Field-name validation (security) and the resync-progress decision stay in the adapter. - #1695/#1743 fix: extend the #1230 delta heuristic in readModifyWriteStateMd to the curated current_phase_name, table-driven via getFieldClassification('current_phase_name').preservation === 'preserve-always'. When a write does NOT change the body Phase: source line, the curated frontmatter value wins over syncStateFrontmatter's body re-derivation (which harvests a wrong parenthetical aside — #1695). begin/planned/complete-phase rewrite their body Phase line, so the delta does not fire for them and current_phase_name still advances. - Regression: bug-1695-state-patch-clobbers-phase-name.test.cjs — unrelated patch preserves curated current_phase_name; patching the Phase source advances. All 70 transition + 493 regression tests pass (state/phase/milestone + #905/#397/#3242 lineages). Closes #1791 |
||
|
|
4525bc6f4d |
fix(#1749): close epic #1702 audit gaps — drift-guard bin/install.js, ci-test-scope wiring, ADR divergence (#1751)
Post-merge coverage-audit follow-ups to epic #1702 (found by an independent gpt-5.5/high audit + adr-phase-coverage cross-reference). None are CRITICAL — the 9-rule enforcement shipped and works; these close completeness/integrity gaps between ADR-1703's promises and the as-built reality. 1. drift-guard bin/install.js scope (ADR-1703 L114-119): the drift guard covered src/runtime-homes.cts only; the ADR named bin/install.js too. Phase 6's glob expansion made bin/install.js a covered surface. Extended tests/portability-vocab-drift.test.cjs with TWO sound checks: (a) any bin/install.js top-level function that directly returns path.*() must be in PATH_RETURNING_FNS (tight, 0 FP — the body-contains heuristic is unsound here, ~33 FPs); (b) a curated two-way existence lock on the installer path helpers (catches a rename making a vocab entry stale; keeps the curation in sync with PATH_RETURNING_FNS). The residual new-resolver boundary (temp-var shape) is documented. 2. ci-test-scope wiring (the Phase 6 portability selection rule was ineffective): eslint-rules/ was not in the product-code prefix list, so an eslint-rules-only change set code_changed=false and CLEARED the matched tests (reproduced: targeted_tests=[]). Added eslint-rules/ to the prefix list and the P1-P4 RuleTester suites to the selection rule (it previously listed only P5/P6). Verified: code_changed=true, 11 tests selected. 3. ADR-1703 acceptance note amended to record the two further as-built divergences: the disable-ban shipped as an out-of-band test (not the specified local/no-portability-disable meta-rule — the test runs outside ESLint so it cannot be self-disabled, at least as strong); and the drift-guard bin/install.js scope resolution above. Epic #1702 all eight phase boxes now checked. No runtime change; no-changelog (contributor tooling + docs). Closes #1749 Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
871621c3c8 |
feat(#1740): require-fs-op-fallback production AST rule + Windows transient-lock retry (Phase 6) (#1742)
* feat(#1740): require-fs-op-fallback production AST rule + Windows transient-lock retry (Phase 6) ADR-1703 Phase 6 of the cross-platform portability epic (#1702). Adds the second production-code portability AST rule + the ADR-mandated glob expansion to bin/install.js and scripts/build-hooks.js. - eslint-rules/require-fs-op-fallback.cjs: flags an unguarded fs.rename / fs.renameSync (the atomic-publish primitive named first in DEFECT.WINDOWS-FS-OPS.symptom) that is NOT inside a try/catch whose handler references a transient errno ('EPERM'/'EBUSY'/'EACCES' or a *RETRY_ERRNOS set) AND NOT behind a Windows platform guard. A catch that silently swallows or cleans-up-and-rethrows without an errno check does NOT satisfy the defect's 'never silently swallow' clause. copyFile/unlink are deliberately not flagged (they are the fallback primitives per the defect's own fix-forward). Scope narrowed to rename per Phase 5's precision discipline; documented on #1740. - src/shell-command-projection.cts: export retryRenameSync(from, to) — the drop-in bounded-retry helper over the existing atomicRenameWithRetry. - 27 bare fs.renameSync sites across 11 modules routed through retryRenameSync (capability-lifecycle/lock/source, installer-migrations, milestone, phase, planning-workspace, roadmap-upgrade, runtime-hooks-surface, state, workstream). Idempotent on POSIX; resilient to AV/indexer transient locks on Windows. - eslint.config.mjs: register rule at error on src/**/*.cts; new focused portability-rules block covering bin/install.js + scripts/build-hooks.js (ADR-1703 L124-126 glob expansion — both files are compliant: zero rename violations). - tests: 15-case RuleTester suite; portability-rule-disable-ban extended (PROTECTED_RULES + scans bin/install.js/build-hooks.js with shebang handling); ci-test-scope portability-lint selection rule. - CONTEXT.md DEFECT.WINDOWS-FS-OPS predicate rewritten to point at the rule; docs/contributing/cross-platform-portability-rules.md reference + how-to. Closes #1740 * chore(#1740): backfill changeset pr:1742 * fix(#1740): tighten require-fs-op-fallback precision (codex review HIGH-1/HIGH-2) Addresses two false-negative findings from the codex (gpt-5.5/high) adversarial review of PR #1742: HIGH-1 — a catch that REFERENCES a transient errno but only rethrows (no retry/fallback) was marked compliant. The DEFECT.WINDOWS-FS-OPS fix-forward requires retry, not just recognition. Fix: catchHandlerHasRetrySignal now requires a loop `continue` backedge OR a `return <call>` delegation; a bare rethrow is flagged. The misleading `/* retry logic */` valid test is replaced with a real retry loop, and the rethrow-only shape is added as invalid. HIGH-2 — the nested-try ancestor walk treated an OUTER errno-catch as protecting the rename even when an INNER catch intercepted/swallowed the error (the outer catch is unreachable). Fix: isInsideTransientErrnoTryCatch now stops at the NEAREST enclosing TryStatement WITH A CATCH HANDLER whose block contains the rename (try-finally is skipped — it doesn't catch); outer catches are no longer consulted. The unsound nested-try valid test is converted to invalid, and a try-finally-skipped valid case is added. Verified: 17 RuleTester cases pass; zero new production violations (the 27 fixed sites use retryRenameSync; the real retry loops — atomicRenameWithRetry, capability-ledger/consent, build-hooks — remain compliant via continue/errno); lint:ci green; disable-ban + vocab-drift green. --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
9d52043f50 |
feat(#1733): normalize-path-in-content production AST rule + fix Windows agent-skills content leak (Phase 5) (#1736)
* feat(#1733): normalize-path-in-content production AST rule (Phase 5) ADR-1703 Phase 5 — the first production-code rule. local/normalize-path-in-content (src/**/*.cts, @typescript-eslint/parser): flags a path-returning fn result (path.basename excluded — returns a separator-less filename) interpolated into an @-reference / config-dir markdown body without .replace(/\\/g,'/') normalization, per RULESET.CONTENT-PATH-NORMALIZATION / DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT. Build-and-assess found the canonical defect site (computePathPrefix) already compliant and only 1 src/ hit — a false positive (path.basename in a status message) — eliminated by narrowing (exclude basename; require a real @-ref/ config-dir marker, not bare .md). 0 src/ violations: clean forward-prevention. The out-of-band disable-ban now scans src/**/*.cts too (typescript-estree) so the production rule also cannot be eslint-disabled. Registered (error) + PROTECTED_RULES; CONTEXT.md predicates + how-to doc updated. - RuleTester suite (26 cases) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1733): add changeset for Windows agent-skills path-leak fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: harden mutation-matrix.cjs stdin read against EAGAIN on non-blocking pipe scripts/mutation-matrix.cjs read piped stdin via readFileSync(process.stdin.fd). On macOS libuv marks the stdin pipe fd non-blocking, so a synchronous read can throw EAGAIN before the writer fills the pipe — intermittently, under heavy CI shard load — aborting the script (status 2) and flaking mutation-matrix-ratchet. Replace with readStdinSync(): an fs.readSync loop that retries on EAGAIN (1ms synchronous Atomics.wait yield), stops on 0-byte/EOF, and rethrows other errors. Deterministic regression test injects EAGAIN via an fs.readSync monkeypatch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * ci: re-run golden-install-parity on src/lib + installer changes (close drift guard) golden-install-parity hashes every installed bin/lib/*.cjs per runtime, so it must re-run whenever the built lib could change. ci-test-scope selected it for neither src/** nor installer changes, so a source-only edit (e.g. #1691's milestone.cts/roadmap.cts) recompiled bin/lib and silently drifted the golden fixtures past the scoped lane. Add golden-install-parity.test.cjs to both the 'TS runtime sources' and 'installer and package layout' selection rules, with behavioral regression tests for each. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
f08b177215 |
feat(#1726): G1-G6 portability AST rules; fix all offenders; delete the ratchet (Phase 4) (#1731)
Phase 4 of epic #1702. Closes #1726. |
||
|
|
41dfeed45a |
feat(#1724): complete install write-confinement (copyWithPathReplacement, installCodexConfig) (ADR-1239 Phase B) (#1725)
* feat(#1724): complete install write-confinement (copyWithPathReplacement, installCodexConfig) ADR-1239 Phase B (parent #1679). PR #1706 (2a) confined the layout-driven plan path and _copyStaged's inline guard; this completes the destSubpath write-confinement acceptance criterion for the two remaining write sites and canonicalizes _copyStaged. - copyWithPathReplacement: new required confinementRoot param; a fail-closed gate (assertDestWithinConfigHome + hasExistingSymlinkBetween) runs BEFORE the rmSync/mkdirSync; root threaded through recursion + all 4 call sites (stageRoot for pristine staging, targetDir for the 3 install sites); writes go through the validated absolute path. Exported for behavioral testing. - installCodexConfig: confines config.toml, agents/, and per-agent agents/<name>.toml (name from agent frontmatter) via the canonical gate + symlink-escape guard (parity with the other two functions). - _copyStaged: fail-closed when configDir omitted (all callers pass it); delegates strict-subpath to the canonical gate, keeps its symlink guard, writes through the validated absolute path. Reuses the existing assertDestWithinConfigHome (handles absolute dests via path.resolve) and hasExistingSymlinkBetween — no new module. Behavioral regression tests (escape/dest==root/fail-closed/symlink/name-injection), red-first proven; cross-platform symlink tests use t.skip not bare return. Closes #1724 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1724): backfill changeset PR number (#1725) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c57e0d56c2 |
feat(#1720): no-unguarded-nonportable-exec AST rule; retire the regex script (Phase 3) (#1723)
Phase 3 of epic #1702. Closes #1720. |
||
|
|
30d4b85de5 |
feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1684): validate and document host-integration axes (16 runtimes) Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add host-integration capability matrix and adr amendment New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1684): harden dispatch negotiation edge cases Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1684): register host-integration.cjs in lint-ignore and inventory New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add changeset fragment for host-integration interface Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add how-to for sourcing a host's integration axes Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1a46109b97 |
enhance(#1579): deterministic gsd-tools query eval.score verb (#1583)
* feat(#1579): deterministic gsd-tools query eval.score verb Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb (coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain; gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt. Non-breaking — additive only. arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR). * fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline - C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives) - C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws - C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface) - C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access) - inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows) - size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step) - eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage) * fix(#1579): register eval family in alias-drift gates Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs families and to familyArrayKeys in the manifest-coverage test, so the eval family lands under the same drift guard as every sibling family (state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review. check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10. * docs(#1579): use half-open verdict band ranges in CLI-TOOLS overall_score is fractional and thresholds are >=80/>=60/>=40, so a score in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit. * fix(#1579): validate eval.score CLI inputs Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior. |
||
|
|
a63684c222 |
enhance(#1577): WebFetch/WebSearch injection isolation + opt-in blocking (#1585)
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the largest untrusted channel) in gsd-read-injection-scanner; shared untrusted-input-boundary reference @-included by the 8 ingest agents (randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring); opt-in security.injection_blocking (default advisory — non-breaking). arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth). * fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized - A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already in the transcript. The prompt-level data/instruction boundary is the primary control. - A2: registered security.injection_blocking in the config schema + defaults manifests (default false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads. - A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention). - A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale). - A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input. - Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) + drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as the noted pre-existing follow-up. * fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate The new reference quotes injection phrases ('ignore previous instructions', 'you are now…') as examples agents must NOT comply with, tripping the repo's own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red on HEAD). Allowlist it alongside the other security docs (security-model.md, TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS scanner test doesn't scan references/, so only the shell gate needed it. Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15. * fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so no named web-ingress agent is uncovered, keeping the two justified additions (gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10. - gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset. - gsd-assumptions-analyzer reads 5-15 codebase source files (external/source- document ingress per the boundary), though it has no web tools. INGEST_AGENTS in the isolation test now asserts all 10; size baselines regenerated (+60 bytes each, both well under the DEFAULT cap); changeset reworded 8 -> 10. Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39. * docs(#1577): document security.injection_blocking + boundary seam trek-e Major 2 + Minor: - docs/CONFIGURATION.md: add the top-level security.injection_blocking key to the Full Schema and a Security Settings subsection, distinguishing it from the workflow.security_* namespace; honest circuit-breaker-not-redactor framing matching ADR-1577 / security-model. - CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry. Verified: lint:docs ok; config-field-docs + contributor-standards green. * test(#1577): make read-injection property test git-text, not binary trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as degenerate-edge inputs. The NUL is what actually made git classify it binary (git binary = NUL in first 8K). Replace both with text-safe escapes that keep the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File now diffs/blames line-by-line. Verified: property test 2/2; no NUL/raw-noncharacter bytes remain. * docs(#1577): align untrusted boundary docs Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner. * docs(#1577): align ADR ingest agent count Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
0c4d570541 |
fix(#1628): type-safe config-set validation — close JSON-coercion enum bypass + enforce capability schema
Three related defects in cmdConfigSet, all 'config-set stores invalid values silently': 1. Missing guards: workflow.security_block_on (enum) and workflow.security_asvs_level (integer 1-3) had no store-time validation. 2. Systemic JSON-coercion bypass: every string-enum guard used VALID_X.includes(String(parsedValue)). Because the value is JSON-parsed before validation, String(["member"]) === "member" let a JSON array slip through and an array was stored in a scalar key. Reproduced on human_verify_mode, statusline.context_position, context_guard_mode, fallow.scope/profile, source_grounding_authority, drift_action, context. 3. Unvalidated capability keys: 32 capability-registry-owned keys (4 enum, 25 boolean, 2 number, 1 string) had no hardcoded guard, so any value — including coerced arrays/objects and out-of-enum strings like code_review_depth=garbage — was stored silently. Fix: a type-safe assertEnumValue() helper (requires typeof === 'string' before membership), routed through all nine central string-enum guards (messages preserved byte-for-byte); plus a generic capability-registry validation block that validates every capability key against its declared type/values (enum via the registry's values — single source of truth — boolean, number, string). Behavioral regression tests cover every central enum key and representative capability keys (array + object coercion rejected, out-of-enum rejected, valid accepted) with boundary coverage for the security keys. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c63fa35b0a |
fix(#1615): allowlist windsurf-conversion.test.cjs in prompt-injection-scan
The commandName validation tests legitimately contain real injection payloads (newline + system-role override phrases, fake [SYSTEM] tags, jailbreak strings) to prove the validator rejects them. The scanner cannot distinguish a test fixture asserting rejection from an actual injection attempt, so CI failed on the test that adds the security control.
Added tests/windsurf-conversion.test.cjs to scripts/prompt-injection-scan.sh ALLOWLIST with a comment citing the defect class.
Also added DEFECT.PROMPT-INJECTION-SCAN-COLLISION-WITH-TESTS to CONTEXT.md so the pattern is documented. Initial draft of that predicate ITSELF triggered the scanner (it quoted the literal injection phrase as an example) — reworded to use descriptive references ('scanner-matching payload', 'instruction-override phrase') since the scanner scans CONTEXT.md too. That meta-collision is now called out in the fix-forward and prevention subkeys.
|
||
|
|
652142521b |
enhance(#1549): validate PR-title issue-ref convention at open time (#1576)
* enhance(#1549): validate PR-title issue-ref convention at open time The release changelog is title-driven: release.yml generates "What's Changed" from PR titles, then format-github-release-notes.cjs buckets each line by its conventional-commit prefix and relies on a `(#<issue>)` in the title to render the issue link. Both rules were enforced only socially, so titles like `fix(core): ...` (no issue link) and `[security] fix(...): ...` (leading tag defeats the `^fix` bucket anchor -> mis-filed under Enhancement) silently broke the changelog, landing on the maintainer as release-time cleanup. Extract the title matcher into one shared module consumed by BOTH the changelog classifier and a new PR-title CI gate, so a title that passes the gate cannot mis-bucket in the changelog (single source of truth). - scripts/lib/conventional-title.cjs (new): classifyBucket + evaluatePrTitle + the anchored regexes. One matcher, two consumers. - scripts/release-notes/format-github-release-notes.cjs: classifyTitle now delegates to classifyBucket (behavior preserved; existing tests green). - .github/workflows/pr-title-validator.yml (new): runs evaluatePrTitle on pull_request opened/edited/reopened/synchronize, for ALL authors (the drift came from member PRs). Trusted base-ref checkout; WARN_ONLY knob for rollout. - tests/conventional-title.test.cjs (new): bucket + gate cases incl. the leading-tag mis-bucket (backfills the untested classifyTitle case) and a cross-check that the classifier delegates to the shared matcher. - CONTRIBUTING.md: document the `type(#<issue>):` rule and no-leading-tag. Claude-Session: https://claude.ai/code/session_01UMV5Qr3H4oFikbuiEauGQk * fix(#1549): check out the PR in pr-title-validator so the new matcher resolves The workflow checked out the base branch (next) as a trusted policy source, but the shared matcher (scripts/lib/conventional-title.cjs) is introduced by this PR and does not exist on next yet — so require() failed and validate-title errored on its own introducing PR. Check out the PR's merge ref instead: the matcher under review is present, the check is self-consistent, and a fork pull_request runs read-only with no secrets, so running the PR's own pure-string regex is safe. * fix(#1549): move conventional-title.cjs out of installed scripts/lib/ bin/install.js bundles every file under scripts/lib/ into the user-installed payload (the changeset CLI's dependencies), and install.test.cjs (#935) asserts that exact set. The new matcher is release/CI tooling that must NOT ship to users, so placing it in scripts/lib/ both broke the install manifest test and would have shipped dead code. Relocate it next to its consumer in scripts/release-notes/ (which the installer does not copy) and update the three require paths (classifier, workflow, test) + the CONTRIBUTING reference. install.test.cjs now 125/125; conventional-title + release-notes suites green; lint:ci clean. * fix(#1549): load title matcher from trusted base ref, not PR code Addresses review (Solvely-Colin + trek-e): the gate checked out the PR merge ref and require()'d evaluatePrTitle from PR-controlled code, so any future PR could edit conventional-title.cjs to return { valid: true } and wave its own malformed title through — a self-bypassable required check. Load the matcher from a base-branch checkout instead (ref: github.event.pull_request.base.ref), the same trusted-policy-source pattern pr-target-validator.yml already uses. The PR can change its title but not the ruler that measures it. An existsSync bootstrap guard skips the check when the matcher isn't on the base branch yet (the introducing PR); every PR after merge is fully gated. This keeps the single shared matcher (#1549's whole point) rather than forking the regex into the workflow. Also per review: - add tests/conventional-title.property.test.cjs (fast-check): any `type(#n): summary` round-trips to valid; evaluatePrTitle/classifyBucket are total functions (never throw). - pin the `fix(#):` zero-digit boundary as missing-issue-ref. Claude-Session: https://claude.ai/code/session_01VqUHNQCh71pEqjo96zkgQL --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
da4a86d8c1 |
feat(#1596): ship GSD skills via .claude-plugin/plugin.json
Phase B-provide of epic #1258. Adds a build-generated skills/ dir + a skills manifest field so plugin-installed GSD exposes gsd-core:<skill> the native Claude Code way. Closes the gap where plugin-only installs lacked the skill surface because bin/install.js never ran. - scripts/gen-plugin-skills.cjs: build step converting commands/gsd/*.md to skills/gsd-<stem>/SKILL.md via convertClaudeCommandToClaudeSkill - .claude-plugin/plugin.json: add "skills": "./skills/" - package.json: add skills to files, gen:plugin-skills to build chain - tests/issue-766-plugin-manifest.test.cjs: Section H conformance (manifest field + dir + frontmatter + count parity) + C2 skills symlink - docs/adr/766-*.md: dated amendment adding skills surface row - .changeset/rapid-bears-hum.md: type Added - skills/: 69 generated gsd-<stem>/SKILL.md files (build-committed) Closes #1596 |
||
|
|
e12a2abfd8 |
feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed), enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way to browse or audit parked seeds on demand. This adds a read-only listing, following the established --list → workflow pattern (per the approved scope on - gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the seeds dir, returns { count, seeds[], summary } JSON with each seed's id, slug, status, scope, trigger_when, planted, title. Optional case-insensitive status filter. User-controlled content is sanitized (sanitizeForDisplay) and every path validated (requireSafePath); read-only. Independent of audit.scanSeeds, which only returns unimplemented seeds for the milestone surface. - /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that renders the seed table. Closes #441 * chore(#441): point changeset fragment at PR #722 * test(#441): allowlist list-seeds test in prompt-injection scan The test asserts that list-seeds neutralizes injection payloads (<system>, [INST]) embedded in seed content, so the fixtures legitimately contain those patterns — same as the sibling security tests already on the allowlist. * fix(#441): use canonical /gsd:capture colon form in list-seeds workflow Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new list-seeds workflow used the hyphen form. * docs(#441): sync help full.md + INVENTORY for --list-seeds Adds the --list-seeds entry to the help reference (help/modes/full.md, per bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json. * docs(#441): add --list-seeds how-to + drop phantom statuses Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers): - USER-GUIDE.md Seeds section (how-to): extend the task to cover auditing parked seeds on demand via --list-seeds, including the status filter — kept task-oriented per Diataxis how-to mode. - CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected from the list-seeds filter vocabulary; the system only produces dormant|active|triggered (src/audit.cts scanSeeds). Reference must be factually accurate and complete. * fix(#441): guard non-scalar status frontmatter in cmdListSeeds A seed with a bare `status:` line (extractFrontmatter yields {}) or a `status: [a, b]` value (yields an array) crashed the whole audit list: `(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string. Coerce every frontmatter read through a `fmStr` helper (mirrors the existing `typeof fm.id === 'string'` guard), so a non-scalar status falls back to dormant and non-scalar scope/trigger_when/title can no longer leak a raw array/object into the JSON contract. Title is now capped symmetrically. Adds regression coverage for empty and array `status:` and non-scalar fields. Refs #441 * docs(#441): align list-seeds workflow status vocabulary The load_seeds step listed `implemented` as an example status filter, but the real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds); `implemented` has no producer. Matches the earlier CLI-TOOLS.md correction. Refs #441 * refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested in-process (review minor #1). No behavior change. Filter comparison now matches the raw lowercased status (both sides already normalized) instead of sanitizeForDisplay(status); sanitization is for output, not matching (review nit #3). * test(#441): add fast-check property coverage and count=1 boundary for list-seeds Adds tests/list-seeds.property.test.cjs with four fast-check properties over deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing (review minor #1). Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2). * chore(#441): sync runtime launcher snippet into list-seeds workflow Propagate the current _runtime-launcher.snippet.sh (with non-Claude runtime home probes) into the new list-seeds.md workflow via scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation. * test(#441): record list-seeds.md in workflow size baseline (#1074) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
fb24d1c7ea |
test(#1496): add behavioral validateCapability check for capability tutorial manifests (#1497)
* test(#1464): add behavioral manifest-validation test for capability tutorial docs Extracts JSON capability manifests from tutorial/reference docs and validates them through the real validateCapability — closing the test gap that let issue #1464's broken tutorial manifest (step missing ref) pass undetected. Adds fix-1464-docs-manifest-validation.test.cjs with: - Suite 1: build/install/reference docs' complete manifests pass validateCapability - Suite 2: adversarial fixtures prove the original #1464 bug shapes are caught (step without ref → "steps[0].ref must be an object…"; id/folderId mismatch) - Suite 3: extractManifests helper unit tests (complete vs. partial block filtering) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(#1496): allowlist docs module for 3-file test cluster docs-parity-live-registry, docs-update, and the new fix-1464-docs-manifest-validation sit on the same docs production module; add the docs entry to lint-test-file-count.allowlist.json so the novel-offender CI gate passes. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
33ccf5f89d |
fix(#1367): project-local install uses flat gsd-<cmd>.md layout (fixes /gsd: colon namespace) (#1489)
* fix(#1367): project-local install uses flat gsd-<cmd>.md layout Claude Code project-local installs now write command files as flat gsd-<cmd>.md at .claude/commands/ level instead of commands/gsd/<cmd>.md (subdirectory), so Claude Code registers /gsd-<cmd> (hyphen form) matching hooks, statusline, and all cross-command references. - capabilities/claude/capability.json: local destSubpath commands/gsd → commands - bin/install.js else branch: flat gsd-<stem>.md loop with runtime rewrites - bin/install.js uninstall (1c): remove flat files + legacy subdir cleanup - bin/install.js writeManifest: record flat commands/gsd-<cmd>.md keys - legacy migration: preserves dev-preferences.md across reinstall and uninstall - gsd-core/bin/lib/capability-registry.cjs: regenerated - 6 new regression tests (L0–L5) in bug-1367-*.test.cjs - Updated E suite in bug-3683 + bug-1736, layout + surface + descriptor tests - scripts/lint-regression-test-names.allowlist.json: grandfathered bug-1367 test Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#1367): add issue reference to allow-test-rule comment Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |