b2f4aa94356dc4859946f508965a2fa8e40bc93b
4646 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b2f4aa9435 |
docs(#2420): clean stale get-shit-done/ path refs in translated docs (#2421)
After the package/repo rename in #604, the English docs were updated to use gsd-core/... paths, but the four translated doc trees (ja-JP, zh-CN, ko-KR, pt-BR) and .changeset/README.md were never updated and still referenced the pre-rename get-shit-done/ runtime directory, which no longer exists. This commit brings the translations in line with the English docs: - docs/{ja-JP,zh-CN,ko-KR,pt-BR}/**/*.md (57 files): get-shit-done/ -> gsd-core/ (path references) #references-get-shit-donereferencesmd -> #references-gsd-corereferencesmd (anchor in INVENTORY -> ARCHITECTURE links) - .changeset/README.md:9 issue URL: open-gsd/get-shit-done-redux -> open-gsd/gsd-core Legacy references intentionally preserved (historical record): - CHANGELOG.md, .changeset/archived/*, docs/RELEASE-NOTES-LEGACY.md - docs/cleanup-get-shit-done-cc.md, docs/adr/*, docs/research/* - docs/{ja-JP,ko-KR}/superpowers/plans/2026-03-18-* (developer's local paths) - docs/{INVENTORY,README,FEATURES,installer-migrations}.md (rename-history descriptions, some tagged <!-- gsd-allow-legacy-name -->) - Code/tests implementing or testing legacy-cleanup logic (bin/install.js, gsd-core/bin/lib/legacy-cleanup.cjs, scripts/lint-legacy-dir-name.cjs, migration sources/tests) No source code changes — documentation only. Fixes #2420 |
||
|
|
873bdf51e5 |
fix(#2352): expand tilde paths in review scope before the deleted-file filter (#2419)
* fix(#2352): tilde-expand SUMMARY.md key-files paths before deleted-file filter compute_file_scope's "Filter deleted files" step tested the literal `~/...` value from SUMMARY.md key-files entries with `[ -f "$file" ]`, which bash never tilde-expands (only a literal `~` in source text expands, not one arriving as an already-expanded variable value). Real files recorded with a `~/...` path were silently misclassified as deleted and dropped from REVIEW_FILES, and a phase whose every recorded file used a tilde path hit the empty-scope skip as a false negative. Adds a tilde-normalization loop as step 1 of post-processing (all tiers), before the deleted-file filter, rewriting a leading `~/` to `${HOME}/...` so downstream existence checks, the empty-scope short-circuit, and the FILES_TO_READ/CONFIG_FILES construction all see a real, openable path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2352): regenerate fixtures + lint gate-prep * chore(#2352): add Fixed changeset fragment (pr 2419) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dd5a2211c9 |
enhance(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) (#2416)
* test(#1964): add failing-first semantic-recall contract tests Epic #1957 Phase 3C (final). Source-text-is-the-product contract tests: semantic recall via MemPalace (top-k meaning-similar prior resolutions, catches same-root-cause/different-wording cases), indexing resolved sessions at archive, graceful degradation to keyword matching when MemPalace is absent, knowledge-base.md stays the durable plain-text source of truth, agent Phase 0 / Matching Logic is semantic-first (the stale 'keyword overlap, not semantic similarity' claim must go), and no new embedding/vector infra (reuse MemPalace). Failing-first: reference, the Matching Logic reframe, the Phase 0 consolidation, and the archive indexing step do not yet exist. * feat(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) Epic #1957 Phase 3C (FINAL). Replaces keyword-overlap matching with semantic recall: at Phase 0 the debugger queries MemPalace with the current symptoms and surfaces the top-k meaning-similar prior resolutions, catching the same-root-cause/different-wording cases keyword overlap missed (the self-noted 'keyword overlap, not semantic similarity' limitation). Resolved sessions are indexed into MemPalace at archive (symptoms + root_cause(s) + fix + recurrence guard). knowledge-base.md remains the durable plain-text source of truth; when MemPalace is absent the debugger falls back to keyword-overlap matching (logged, never a silent skip). No new embedding/vector infrastructure — MemPalace is reused. Size-neutral agent edits: the Matching Logic section reframed (keyword-only -> semantic-first + keyword-fallback + @-include); Phase 0's three keyword bullets consolidated into one semantic-first bullet; one MemPalace-indexing step added at archive. Agent at 57222 B (122 B headroom — final phase). Full rules in gsd-core/references/debugger-semantic-recall.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1964): address orthogonal review (invocation mechanism, index Resolution-not-symptoms + redaction, fallback detail) - HIGH: the 'query MemPalace' instruction was WHAT-level only; the agent has no MCP tools. Added an Invocation section naming the Bash CLI (mempalace search --wing <wing>) + MCP-when-registered + wing resolution (config.mempalace.wing -> project_code -> project dir), matching every other MemPalace integration. Without this the feature silently degraded to keyword matching even when MemPalace was present. - MEDIUM (security x2): index the agent-authored Resolution summary (root_cause + fix + recurrence_guard), NOT raw user-supplied Symptoms — excludes attacker-controlled prose from the cross-session index AND reduces secret/PII leakage. Redact secret-shaped values before indexing. Stated the write order (KB append + commit MUST succeed before indexing). - LOW: restored 'identifiers' + 'case-insensitive' to the keyword fallback; added a test asserting the fallback mechanics survived the Phase 0 consolidation (Error patterns field, 2+ token overlap, identifiers, case-insensitive). * chore(#1964): ratchet agent-size baseline downward (leaner archive bullet shrank gsd-debugger.md 57222->57197) * chore(#1964): backfill changeset pr number (PR #2416) |
||
|
|
c67f301867 |
feat(#1963): emit blameless-postmortem Prevention block at resolution (#2410)
* test(#1963): add failing-first prevention/postmortem contract tests Epic #1957 Phase 3B. Source-text-is-the-product contract tests: blameless 5-Whys that BRANCHES per Phase 2A RCA (not a single-cause chain; treats agent error as 'why was that possible?'), the 'why wasn't this caught?' question, the recurrence-guard taxonomy (regression test / assertion / lint rule / KB pattern), the KB-entry why_not_caught + recurrence_guard fields with backward compat, the session-manager prevention summary line, and the Zawinski scope-boundary (a block, not a subsystem). Failing-first: reference, archive_session edit, KB schema extension, and session-manager summary do not yet exist. * feat(#1963): emit blameless-postmortem Prevention block at resolution Epic #1957 Phase 3B. At archive_session the debugger now produces a Prevention block with three blame-free components: a branching 5-Whys causal chain (branches per Phase 2A RCA, not a single chain; 'agent error' prompts 'why was that possible?', never blame), a 'why wasn't this caught?' answer naming the missed gate (test/typecheck/lint/review/verify), and a concrete recurrence guard (regression test / assertion / lint rule / KB pattern). The knowledge-base entry gains two structured fields (why_not_caught + recurrence_guard) so future Phase-0 recall surfaces the prior prevention, not just the prior fix. Additive: old entries without the fields still load. The session-manager compact summary surfaces a one-line prevention summary. Full rules extracted to gsd-core/references/debugger-prevention.md (slim archive_session step + 2 KB fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1963): address orthogonal review (CRITICAL append-template drift + Phase-0 consumption + parity test) - CRITICAL: the archive_session KB append template omitted Why not caught + Recurrence guard (only the Entry Format had them) — the feature's core deliverable silently did not happen. Added both fields to the append template the agent actually follows (nearest-instruction wins). - HIGH: Phase 0 (KB read) only surfaced root_cause + fix; the new fields were dead data. Extended the Phase 0 Evidence line to consume why_not_caught + recurrence_guard when present (absent on old entries — backward compat holds). - MEDIUM: added a cross-section parity test (every Entry-Format field must also appear in the append template — the guard that would have caught the Critical) + a Phase-0-consumption assertion. - MEDIUM: the 'branches per Phase 2A' claim is now wired — reuses reasoning_checkpoint.candidate_causes across the four categories. - MEDIUM: recurrence-guard taxonomy gains type refinement + config-default change; LOW: added 'build' gate to both surfaces for parity. - NIT: compact-summary fallback shape ('no gate existed'); verify the guard artifact exists before recording it. * test(#1963): anchor Phase-0 consumption test on the specific heading The regex /Phase 0[\s\S]{0,1200}/ matched the first 'Phase 0' in the file (in knowledge_base_protocol prose), not the Phase 0 block in investigation_loop. Anchor on '**Phase 0: Check knowledge base**' and widen to 1500 chars. * chore(#1963): backfill changeset pr number (PR #2410) |
||
|
|
36a311c5bb |
enhance(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) (#2409)
* test(#1962): add failing-first repro-hardening contract tests Epic #1957 Phase 3A. Source-text-is-the-product contract tests: PBT shrinking (fast-check/Hypothesis, minimized seed, manual-minimization degradation), the four oracle types (specified/derived/metamorphic/implicit with implicit flagged weakest), boundary neighbors (off-by-one/min-max/empty-singleton tied to the equivalence class), oracle_type in DEBUG Resolution, and the Phase 1A tie-in (minimized seed + real oracle => the mutation guardrail bites). Failing-first: reference, agent cross-refs, and template field do not yet exist. * feat(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) Epic #1957 Phase 3A. Extends Minimal Reproduction (shrinking) and Test-First Debugging (oracle classification + boundary neighbors): - Shrinking: wrap an input-space failing input in a property (fast-check JS/TS, Hypothesis Python) and store the MINIMIZED counterexample as the regression seed; degrade to manual minimization when no PBT framework is present. - Oracle classification: state specified / derived (contract/model) / metamorphic / implicit (crash, weakest) before writing the assertion; record under Resolution.oracle_type; never default to implicit silently. - Boundary neighbors: off-by-one, min/max, empty/singleton around the fixed defect's equivalence class. Together they turn the regression test into a root-cause check — what the Phase 1A mutation guardrail needs to bite. Full rules extracted to gsd-core/references/ debugger-repro-hardening.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1962): address orthogonal review (bounding, provenance, oracle scope, sufficient-triple) - HIGH: added a 'Bound the property/shrink run' section (60s timeout, degrade- to-manual on timeout, do-not-raise-default-run-limits, argv-not-shell) — the gauntlet violation the sibling references already honored. - Medium: test-provenance caveat (the failing input often comes from the bug report — author the generator from a sanitized description, cross-ref debugger-fix-acceptance.md). - Medium: oracle scope note — the 4 types cover deterministic bugs; non- deterministic failures re-route to stability-stress per bug-taxonomy. - Medium: Phase 1A tie-in corrected — seed+oracle is necessary not sufficient; boundary neighbors close the adjacent-input escape; the sufficient triple is seed+oracle+neighbors. - Low: preserve the original noisy repro as a secondary reference; operationalize 'equivalence class' (the predicate the fix draws). Nit: degradation reworded. * chore(#1962): backfill changeset pr number (PR #2409) --------- Co-authored-by: sim <sim@local> |
||
|
|
6baa2a8182 |
feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger (#2407)
* test(#1961): add failing-first bug-taxonomy routing contract tests Epic #1957 Phase 2B. Source-text-is-the-product contract tests (3 taxonomy classes, explicit class->technique routing table, Bohrbug->repro+SBFL+bisect, Heisenbug->record-replay/stability+SKIP-SBFL, Concurrency->atomicity/order/ deadlock checklist, bug_class in DEBUG Current Focus, supersede-not-append) plus a routing-table specification object pinning the documented decisions (SBFL forbidden on Heisenbug is the load-bearing 1B/2B seam). Failing-first: reference, Phase 1.75, and routing-table reframe do not yet exist. * feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger Epic #1957 Phase 2B (reliability-critical). Adds Phase 1.75: classify the failure as Bohrbug / Heisenbug-Mandelbug / Concurrency, then route the investigation technique via an explicit class->technique table (Kernighan: no opaque heuristic). Bohrbug -> reproduction + SBFL (Phase 1.25) + git bisect; Heisenbug/Mandelbug -> record-replay (rr) + stability-stress + statistical sampling, with SBFL explicitly SKIPPED (a flaky spectrum poisons the Ochiai ranking — the load-bearing 1B/2B seam); Concurrency -> the atomicity/order/deadlock checklist first. Reframes (supersedes, not appends — Zawinski) the flat 'Technique Selection by situation' table into a class-routed table; the 11 techniques remain as routed targets. bug_class recorded in Current Focus (DEBUG template); common-bug- patterns catalog cross-referenced to the taxonomy. Full rules extracted to gsd-core/references/debugger-bug-taxonomy.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1961): address orthogonal review (phase-name drift, General lane, revoke framing, row-scoped tests, bounding) - HIGH: reference said 'Phase 1B' (epic shorthand); corrected to the deployed 'Phase 1.25' (matches the agent + SBFL reference). - HIGH: 6 of 11 techniques (Rubber duck, Delta, Working backwards, Differential, Comment-out, Follow-the-indirection) were orphaned by the situation-table reframe. Added a 'General (any class, situation-cued)' lane to BOTH the reference routing table and the agent's Technique Selection table that re-homes them — supersede-not-append now holds. - MEDIUM: the SBFL-skip is structurally retroactive (Phase 1.25 runs before Phase 1.75 classification), so reframed the table column from 'Do NOT use' to 'Revoke if already run' + an explicit 'retroactive revocation, not proactive skip' note stating the ordering honestly. - MEDIUM: contract tests are now row-scoped (parse the table by class, assert per-row) instead of presence-only; added a guard that the previously- orphaned techniques now have a General-lane route. - LOW: pinned the canonical bug_class value form (lowercase-kebab: bohrbug|heisenbug-mandelbug|concurrency; prose may use title-case). - NIT: added a 'Bound the Heisenbug-chase runs' note (rr/stability/sampling timeouts) per the unbounded-subprocess gauntlet. * chore(#1961): backfill changeset pr number (PR #2407) |
||
|
|
f8b16d1874 |
enhance(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger (#2405)
* test(#1960): add failing-first RCA-branching contract + schema-invariant tests Epic #1957 Phase 2A. Source-text-is-the-product contract tests (fishbone >=2 categories, AND-gate, multi-cause root_cause, backward compat, reasoning checkpoint candidate_causes+and_gate fields, debugger-philosophy single-cause note, DEBUG template) plus behavioral schema-invariant checks on two fixtures: two contributing causes (AND-gate yes) -> both recorded; single-cause (AND-gate no) -> one root_cause, identical to today. Failing-first: reference, agent edits, and template note do not yet exist. * feat(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger Epic #1957 Phase 2A. Guards against 5-Whys single-cause bias: before committing root_cause, the debugger enumerates candidate causes across >=2 Ishikawa categories (code/config/environment/data) and explicitly answers an AND-gate question. When the AND-gate fires, every contributing cause is recorded, so a multi-cause fix no longer recurs via the unaddressed second cause. Resolution.root_cause may hold one OR a small set (additive; single-cause sessions are byte-identical to today). The Structured Reasoning Checkpoint gains candidate_causes + and_gate fields; debugger-philosophy.md adds the single-cause-bias trap. Full rules extracted to gsd-core/references/debugger-rca-branching.md (slim Phase 2 routing + 2 checkpoint fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1960): address orthogonal review (AND-gate self-consistency, parity guard, narrowed claim, ripples) - Reference: the collapse rule now enforces AND-gate self-consistency — and_gate=yes with a single confirmed cause is flagged as incomplete (return to Phase 3); a race/timing note clarifies such bugs bridge categories; the 'byte-identical' backward-compat claim narrowed to 'root_cause shape unchanged; reasoning_checkpoint gains 2 fields in every session'. - DEBUG.md: stale 'five-field' mirror prose -> seven-field (parallel-surface drift the reviewer flagged); new debug-session-management parity test pins the field-count claim to the gsd-debugger.md YAML keys (CRLF-safe). - Scalar-assuming consumers of set-valued root_cause updated: session-manager compact summaries (319/332), diagnose-only return (1062), archive entry (1216), ROOT CAUSE FOUND return (1322). - Test: added the AND-gate-yes/single-cause invariant + fixture; rephrased the fixture describe block honestly as a schema-invariant specification. - Phase 2 bullet phrasing clarified ('at hypothesis formation, before the Phase 4 commit'). * test(#1960): parity regex accepts word-form count ('seven-field' or '7-field') * test(#1960): parity regex counts array-valued YAML keys (no inline value) * chore(#1960): backfill changeset pr number (PR #2405) |
||
|
|
50efae13ce |
fix(#2305): stage the shared guard hooks Kilo's native plugin spawns (#2327)
* fix(#2305): stage shared guard hooks for Kilo — drop skipSharedHooksInstall Kilo's capability descriptor declared BOTH hostBehaviors.nativePlugin (a plugin that spawns the shared PreToolUse guard scripts as subprocesses) AND hostBehaviors.skipSharedHooksInstall:true, which suppresses staging of hooks/*.js into the Kilo config dir. The plugin's runHook treats an absent hook script as a silent allow, so every guard it spawned (gsd-prompt-guard, gsd-read-guard, gsd-worktree-path-guard) no-opped on every Kilo install. OpenCode uses the byte-identical plugin with hook staging on and is unaffected — it is the reference shape. The skip flag predates Kilo's plugin surface: it dates to #1821 (hooks were dead weight for a runtime with no hook consumer), and #2093 added the hooks-dependent nativePlugin without revisiting it. - capabilities/kilo/capability.json: remove skipSharedHooksInstall (regenerated gsd-core/bin/lib/capability-registry.cjs accordingly) - bin/install.js: correct the stale #1821 comments claiming Kilo has no plugin surface - tests/kilo-upgrades.test.cjs: install-fixture tests (global + local) asserting the guard scripts land where the plugin's walk-up resolves them; an end-to-end test driving a disallowed out-of-worktree write through the REAL installed Kilo tree and asserting the guard rejects it; a cross-runtime descriptor invariant (nativePlugin and skipSharedHooksInstall:true must never coexist) - tests/kilo-imperative-reference.test.cjs: flip the pinned assertion - golden fixtures regenerated (kilo now stages the 24 hook files, same set as OpenCode) Fixes #2305 * fix(#2305): warn loudly when a guard hook script is missing (runHook) runHook's absent-file branch returned a silent exit-0 allow — the mechanism that let #2305 ship undetected: with the hooks bundle never staged on Kilo, every PreToolUse guard the plugin spawned resolved to "file not found → allow" with zero signal anywhere. Keep the adapter's design contract (a missing hook must never break the tool call — pinned by the existing adapter test) but make the absence loud: console.error once per hook file, naming the unresolved path and the remediation. Applied identically to .kilo/ and .opencode/ plugin copies (byte-parity guard). Golden parity fixtures regenerated (the installed plugin file's hash changed). Fixes #2305 * chore(#2305): add changeset fragment * test(#2305): include gsd-workflow-guard.js in the staged-guards regression list The native plugin spawns four guards on write-like tool calls — the regression test's PLUGIN_GUARD_HOOKS list covered three. Staging itself was already asserted via the golden fixtures (the full bundle), but the named per-guard assertion should cover every guard the plugin actually dispatches. Surfaced by cross-AI review of PR #2327. * test(#2305): update the #1821 tests that encoded Kilo's false no-plugin premise The #1821 hook-copy test asserted Kilo must receive no staged hooks — the exact behavior this PR reverses (and the cause of all 8 CI failures). Kilo moves from the ZCode "no dead hooks" loop to the OpenCode group, with positive assertions on the new contract: the three guard hooks the plugin spawns, hooks/lib/git-cmd.js, and plugins/gsd-core.js all staged. The integration runtime contract flips kilo packageJson to true (the CommonJS marker ships with the bundle), and the pi contract comment no longer cites Kilo as a no-plugin runtime. * chore(#2305): scope the queued #1821 changeset fragment to ZCode only The fragment still claimed the installer skips hooks for Kilo — rendering both it and this PR's fragment into the same release would ship two contradictory statements about Kilo's install behavior. It now claims ZCode only and notes that #2327 reverses the Kilo half. * chore(#2305): rename changeset fragment to the generator naming convention 2305-kilo-stage-guard-hooks.md -> loud-guard-hooks.md, matching the <adjective>-<noun>-<noun> shape npm run changeset generates (review nit). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
13d181aedf |
fix(#2349): exclude status: superseded plans from phase completion counts (#2404)
Adds a status: superseded plan-frontmatter marker that scanPhasePlans excludes from both plan and summary counts, so a phase with a deliberately-unexecuted plan no longer reads incomplete forever (the plan-level analogue of #1514). Includes all-superseded completion handling and a bounded, symlink-safe frontmatter read. Fixes #2349. |
||
|
|
56a5c6404c |
feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger (#2403)
* test(#1959): add failing-first SBFL contract + Ochiai correctness tests Epic #1957 Phase 1B. Source-text-is-the-product contract tests (Ochiai formula documented, Tarantula fallback, top-N seeding, no-coverage skip logged, ranking->Evidence, Bohrbug gating) plus a behavioral Ochiai formula-correctness section: bound [0,1], max-score invariant, a known-fault fixture proving the fault ranks #1 (criterion 2), clean degradation on zero failing tests, and two fast-check properties. Failing-first: reference file and agent routing do not yet exist. * feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger Epic #1957 Phase 1B. When a runnable test suite with per-test coverage exists (>=1 failing AND >=1 passing test), the debugger computes an Ochiai suspiciousness ranking over the coverage spectrum and seeds the top-N suspicious locations into Evidence as first-class hypothesis candidates, narrowing the search space deterministically before LLM reasoning. Tarantula documented as fallback. Degrades cleanly (logged, never silent) when there is no test suite, no failing tests, or no per-test coverage, and is explicitly not trusted on flaky/Heisenbug spectra (pairs with Phase 2B bug-taxonomy). Full rules extracted to gsd-core/references/debugger-sbfl.md (slim Phase 1.25 routing kept in the agent to respect the size cap). No new coverage framework — reuses the project's existing test/coverage runner. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * test(#1959): bound property generators to valid coverage counts The [0,1] property generated failedExec independently of totalFailed, but Ochiai's score is only bounded by 1 under the coverage invariant failedExec <= totalFailed (a failing test that executed s is one of the totalFailed failing tests). Out-of-domain inputs (failedExec=100, totalFailed=5) make the formula correctly return >1. Bound failedExec by totalFailed via fc.chain so the property tests the real domain. Also cleaned up the ranking property (removed dead code). * fix(#1959): address orthogonal review (monotonicity property, degradation row, coverage bounding) - Replace vacuous ranking property (true-by-sort-construction) with a non-trivial monotonicity property: holding totalFailed + passedExec fixed, ochiai is non-decreasing in failedExec. An inverted formula would fail it. - Add the missing 'no passing tests' degradation row (preconditions require >=1 passing test; Tarantula would divide by totalPassed=0). - Bound the coverage subprocess (CLAUDE.md gauntlet): cap the coverage run, degrade-to-skip on timeout, never hang the debug session. - Reword 'discard the ranking' -> 'mark the Evidence entry as revoked (do not delete)' per Kernighan auditability. * test(#1959): bound monotonicity-property generator to valid coverage (failedExecA <= totalFailed) * chore(#1959): backfill changeset pr number (PR #2403) |
||
|
|
863a54ec82 |
fix(#2350): pass --raw to config-get in every build/test gate (#2399)
Adds --raw to config-get workflow.build_command|test_command reads in the post-merge, regression, verify-phase, and audit-fix gates so an unset key is a genuinely empty string, not the literal "" — restoring the auto-detect cascade and graceful skip instead of a false exit-127 failure. Regression guard sweeps all four gate files. Fixes #2350. |
||
|
|
5e52350736 |
feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger (#2396)
* test(#1958): add failing-first guardrail contract tests Epic #1957 Phase 1A. Adds source-text-is-the-product tests asserting the 5-signal fix-acceptance guardrail contract (target test, mutation check, no-op/deletion detector, adjacent tests, revert-and-reconfirm), graceful degradation, FIX REJECTED BY GUARDRAIL return path, per-signal debug-file recording, and subprocess bounding. Failing-first: reference file and agent sections do not yet exist. * feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger Epic #1957 Phase 1A. Prevents accepting a fix that merely greens the test (Goodhart defense / APR overfitting). Adds a 5-signal gate run before fix acceptance: target test, mutation check (Stryker), no-op/behavior-deleting detector, adjacent/held-out tests, revert-and-reconfirm. Degrades gracefully when Stryker or a test suite is absent (each skip logged, never a silent pass), records per-signal results under Resolution.verification, and returns a FIX REJECTED BY GUARDRAIL outcome the session-manager surfaces for revise / accept-as-debt / abandon. Full rules extracted to gsd-core/references/debugger-fix-acceptance.md (slim routing kept in the agent to respect the agent-size cap). Debug template + INVENTORY + manifest + agent-size baseline + AGENTS.md updated. * test(#1958): correct newline-tolerant assertion + regen install-parity goldens The revert-and-reconfirm assertion collapsed whitespace before matching so markdown line-wrapping does not break it. Regenerated the golden-install-parity and install-tree fixtures (npm run gen:golden) to absorb the intentional gsd-debugger.md / gsd-debug-session-manager.md / DEBUG.md / new reference-file changes to the installed artifact tree. * fix(#1958): tighten guardrail per orthogonal review Addresses the isolated reviewer's findings: - signal 5 now states its recorded-repro dependency and routes the no-repro case to the degradation row; revert mechanism specified (git stash / git revert -n); minimality flag tied to diff structure, not revert-ability. - bounded-subprocesses section now bounds the git subprocess (5-30s) too, requires argv-array argument passing, and scopes Stryker to the driving regression test (a mutant killed only by a non-driving test is a finding). - new test-provenance (security) clause: the driving test must be agent-authored; bug-report repro scripts are DATA, never executed verbatim. - tightened 3 contract assertions to bind to specific clauses (guardrail_verdict field, deletion-reject-unless-RCA, 60s+git bounding). - Goodhart framing softened to 'partially-independent'; DEBUG.md template verification field notes the nested map shape. * chore(#1958): backfill changeset pr number (PR #2396) * fix(#1958): add issue ref to allow-test-rule annotation (ADR-456) CI lint-allow-test-rule-refs requires every allow-test-rule exemption to carry a 'see #NNN' issue ref per ADR-456. The new test file's annotation lacked it; this adds (see #1958). |
||
|
|
2c54f219c9 |
fix(#2348): derive verification staleness from git commit time, not mtime (#2394)
readVerificationStatus() decided a phase's verification was `stale` (a *-SUMMARY.md newer than the *-VERIFICATION.md) by comparing filesystem mtimes. mtimes are assigned at checkout time and are not preserved by `git clone` / `cp -R`, and any unrelated `touch` / reformat / editor-save re-stales a valid report — so a committed phase declaring `status: passed` could silently read `stale` on a fresh clone purely from checkout order, falsely rewriting a ROADMAP row and blocking milestone close (#2022 gate). Each file's effective "last changed" time is now its git commit time when the file is committed AND clean, and its mtime otherwise (uncommitted or working-tree-dirty). Both are real wall-clock change times, so a summary committed after — or edited after — the verification reads stale, while a clean fresh clone stays passed. Git commit time is content-tied and clone- stable; mtime is retained only where it is the true last-changed signal. Implementation: - Two bounded git calls per phase (never one-per-file): `git log --first-parent --format=%ct --name-only` for commit times, and `git diff --name-only HEAD` to drop dirty files. readVerificationStatus runs per-phase in the init/roadmap listing loops, so per-file spawning would fan out to P×(S+1) git processes ("Unbounded Subprocesses"). - `--first-parent` so merge commits report their file lists (plain `--name-only` omits merge diffs and would under-date merge-landed content). - The dirty-check fails SAFE: if `git diff` is inconclusive (errors / exits non-zero) the commit times are discarded so every file falls back to mtime, never trusting a possibly-stale commit time (no false "not stale"). - Paths matched back by `/`-bounded suffix (root vs nested `plans/` can't collide) and passed after `--` (dash-named files can't be read as flags). - A phase with no summaries skips git entirely; the scan short-circuits on the first stale summary. A `phaseCleanCommitTimesMs` seam keeps the unit tests hermetic (no git spawn); the resolver's two-call error handling is unit-tested via an injected execGit; two real-git integration tests lock the end-to-end path, the committed-then-edited (dirty) regression, and the `--` argv guard. |
||
|
|
f2c077df38 |
chore(#2387): refactor CONTEXT.md legacy content + add glossary drift gate (#2391)
* chore(#2387): refactor CONTEXT.md legacy content + add glossary drift gate Apply the audit-and-enforce concept from the ADR index (#2356) to CONTEXT.md: correct stale facts, and add a CI gate so the machine-verifiable claims can't silently re-rot. CONTEXT.md was entirely hand-maintained with nothing checking its claims against the shipped tree, so it had rotted. An audit against live code (Memtrace + filesystem + gh), each finding adversarially re-verified, drove 38 factual corrections + 1 surfaced by the new gate: - Dead references: Package Identity named @opengsd/get-shit-done-redux (package is @opengsd/gsd-core); Shell Command Projection named run-git/run-npm/run-tool (real exports execGit/execNpm/execTool); a partial docs/adr/1606 ref; retired sdk/ framing. - Superseded facts: allRuntimes 15 -> 17 (pi #2102, zcode); "seven nested-loader runtimes" -> five (claude reverted flat #924, antigravity flat); stacked-PR examples rebasing onto main -> next; QUOTA_SENTINELS precedence corrected to match src/agent-command-router.cts. - Drifted CONTRIBUTING.md line citations refreshed. Per CONTRIBUTING.md:179, only stale FACTS were corrected -- no maintainer intent, lesson, or opinion was rewritten, and the append-only session log is untouched except one dated in-place superseding note. The three tests that assert on CONTEXT.md content (phase6-capstone-conformance, tracer-bullet, external-job-waiting) keep all their anchors. New scripts/check-glossary-refs.cjs (--check, wired into lint:generated-sync): - Check A: every backticked file reference under a TRACKED_PREFIXES allowlist resolves on disk. Generated gsd-core/bin/lib/*.cjs (77 refs, gitignored), ~/-paths, .planning/, and bare filenames are deliberately skipped so a clean CI checkout never false-fails. - Check B: the allRuntimes count + member set in the glossary prose match bin/install.js's allRuntimes literal (drifts on every runtime addition). tests/check-glossary-refs.test.cjs covers both, including the false-positive guard that a missing bin/lib/*.cjs ref does NOT trip the gate. Closes #2387 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2387): confine glossary-gate file refs to ROOT (no `..` traversal) Pre-PR security review finding (low): extractTrackedRefs fed tokens straight to fs.existsSync(path.join(ROOT, token)), and PATH_TOKEN_RE admits `.` in a segment, so a CONTEXT.md token like `src/../../../etc/passwd` passed the `src/` prefix check and normalized to an out-of-tree absolute path — turning the doc lint into a filesystem-existence oracle on the CI host (existsSync only; CONTEXT.md is a trusted committed file, hence low severity, but a defense-in-depth gap). Add isWithinRoot() confinement in extractTrackedRefs: a token is dropped unless path.resolve(ROOT, token) stays within ROOT. A CONTEXT.md reference is always a plain in-repo path, so a `..` escape is never legitimate. Regression test asserts a `..`-bearing token is skipped and never named in output. Refs #2387 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2387): drop legacy `get-shit-done` name from a CONTEXT.md defect entry CI lint-legacy-dir-name failed: the line-928 upstream-issue re-point I applied wrote the historical provenance as "gsd-build/get-shit-done#3545", and scripts/lint-legacy-dir-name.cjs forbids the legacy `get-shit-done` name. Reword to "moved from #3545 in the predecessor repo" — same provenance, no legacy name. Caught by `npm run lint:ci` (the CI lint chain), which I had not run locally — lint:generated-sync + eslint do not include lint-legacy-dir-name. Refs #2387 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
81f7ab4df1 |
refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) (#2392)
* refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) Cutover all 55 remaining case arms from runCommand's switch to HOST_COMMAND_ROUTERS. runCommand now contains only its default case (~40 lines): the three-layer dispatch (capability → overlay → host table) plus the unknown-command diagnostic. The 73-case switch is dissolved. Each case body was relocated verbatim to a module-scope route*Command function via a brace-matching extractor; inner break; statements (from _dispatchNonFamily early-exit patterns) were converted to return; (5 arms affected); loop break; statements preserved. Closes #2384 * chore: retrigger CI |
||
|
|
d91e32b3ce |
fix(#2383): untrack node_modules — accidentally committed as a hardcoded absolute-path symlink (#2385)
* fix(#2383): untrack node_modules — accidentally committed as a hardcoded absolute-path symlink |
||
|
|
f15eb5f5c9 |
fix(#2347): make the decision-shape evidence test format-agnostic (#2389)
#1365's fail-loud guard reused the parser's own D- grammar as its evidence test, so a populated <decisions> block using any other ID prefix (e.g. D5-01) was invisible to both parser and guard, collapsing could-not-parse into a clean none-present pass. Add an ID-shaped bold-lead-in probe as format-agnostic evidence on both parse paths; empty/prose scaffolds stay none-present. Graduates the #2371 d5-prefix representative fixture to its expected* assertion. Closes #2347. Admin-merged (self-review bypass) with full green CI. |
||
|
|
062f3fda90 |
chore(#2371): representative gate-fixture corpus + document-shaped property test (#2380)
* chore(#2371): representative gate-fixture corpus + document-shaped property test Adds tests/fixtures/representative/ — a permanent corpus of verbatim, incident-sourced fixtures (never author-invented) from #2286, #2347, #2365, #2366, each labeled with its expected gate verdict in a MANIFEST.json and driven through the real CLI gate entrypoint via tests/representative-corpus.test.cjs. Adds a document-shaped fast-check property test alongside the existing writer-seeded bijection test in tests/api-coverage.test.cjs: the existing generator produces rows and renders them through the writer, so the document shape is a constant and it cannot fail against a decoy table; the new one generates the document space instead. Two gates (#2365, #2347) are still open, so their corpus/property assertions are marked with node:test's official `todo` option — the test executes and reports its failure without affecting the process exit code (https://nodejs.org/api/test.html#test-options). The audit-uat corpus (#2286, fixed by #2317) is a normal passing assertion, proving the methodology works end to end and not just cataloguing gaps. Records the fixture-provenance rule in CONTRIBUTING.md: a gate's fixtures may not be derived from the gate's own writer, grammar, or docstring examples; a negative fixture must come from a source that doesn't know the gate exists. No production src/*.cts changes — validation only, per #2371's scope. Closes #2371 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): address orthogonal-review findings — dedup generators, fix field naming, wire dead fields Standards-axis review findings, all fixed: - Deduplicated the row-shape generators (capabilityGen/rowGen/validRowGen) that were copy-pasted between the parse/render bijection test and the new document-shaped property test in tests/api-coverage.test.cjs — a future edit to one could have silently desynced the two properties. Hoisted to a single module-scope declaration both tests reference. - Renamed decision-coverage-guard/MANIFEST.json's expectedOutcome -> expectedReason. It asserted against the gate's `reason` field, but this codebase already has a real, different `outcome` field at parser altitude (extractDecisions' DecisionOutcome) — naming the manifest field after the wrong altitude's term was exactly the ambiguity the "Fixture provenance" rule this PR adds exists to eliminate. - Removed the unused `role` field from three MANIFEST.json files (never read by any test) and wired the previously-dead per-fixture `expectedMinItems` in audit-uat/MANIFEST.json into a real per-file assertion in tests/representative-corpus.test.cjs, using cmdAuditUat's `results` array — catches a regression that moves items between the two fixture files while preserving the aggregate total, which the existing total_items check alone would miss. No changes to test intent or coverage — same assertions, correctly named and fully wired. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: remove dead capabilityState/capabilityWriter requires from gsd-tools.cjs Surfaced by the mandatory pre-PR lint gate (no-unused-vars) while preparing this PR — unrelated to #2371's own changes, but a defect found while working is fixed in place rather than deferred. Leftover from #2368/#2370 (merged just before this branch rebased onto it): the case 'capability' arm that needed these two requires was relocated to bin/lib/capability-command-router.cjs, which already requires both modules directly (lines 24-25) and is their only real consumer (cmdCapabilityState, resolveCapabilityRuntimeState, cmdCapabilitySet). The two requires left behind in gsd-tools.cjs had zero other references in the file and were never re-exported — confirmed via grep across the file and its module.exports. Behavior-preserving: Node's require cache means the underlying modules still load exactly once via capability-command-router.cjs's own requires; gsd-tools.cjs never used its now-removed local bindings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): replace todo-marked assertions with characterization tests gsd-test's own JSONL result parser (gsd-test-runner's internal/pipeline/parse.go, verified directly against that repo's source) has no concept of node:test's `todo` option — it only recognizes kind:"pass"|"fail" and hard-errors on anything else. A { todo: true } test whose body throws is counted as a real failure in gsd-test's own verdict, exactly as if it weren't marked todo — proven by an actual gsd-test run against this branch, which reported outcome:"failed" with all six todo-marked assertions (the property test plus five representative-corpus fixtures) in the failure list, each carrying the correct raw node:test `todo` field the tool's parser simply doesn't read. Replaces todo with characterization: MANIFEST.json now carries both the correct target verdict (expected*) and the exact current observed verdict (currentBuggyOutput, directly verified against live CLI output for all five fixtures). Tests assert currentBuggyOutput — an honest, non-vacuous pin of today's known-broken reality that passes today and will fail loudly the moment the referenced fix changes the observed output, at which point the assertion should be flipped to expected* and currentBuggyOutput deleted. The document-shaped property test switches from throwing fc.assert to non-throwing fc.check (returns RunDetails per fast-check's own docs) and asserts report.failed === true directly, for the same reason. Updates all prose (CONTRIBUTING.md, the fixture READMEs) that previously claimed todo would be respected — that claim was factually wrong for this repo's actual tooling and must not ship. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: regenerate golden-install-parity fixtures after rebase onto next Rebasing onto the current next (which now includes #2381's todo-severity changes to gsd-core/bin/gsd-tools.cjs) produced real conflicts in all 18 golden-install-parity fixtures — expected, since both branches changed the same gsd-tools.cjs hash entry. Resolved by taking one side to unblock the rebase, then regenerating fresh from source via npm run gen:golden and verifying the result; every file's diff is exactly the one hash line for gsd-core/bin/gsd-tools.cjs, correcting a stale intermediate hash from the arbitrary conflict pick. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
58028eaf56 |
fix(#2341): de-dup Cursor / menu by marking skills user-invocable:false (#2386)
Cursor installs both a skills and a commands surface and shows both in '/', duplicating every /gsd-*. Extend the #789 CodeBuddy de-dup to Cursor: convertClaudeCommandToCursorSkill (in both src and the live bin/install.js) now emits user-invocable:false, so the skill stays model-invocable while the commands surface is the single '/' entry point. Closes #2341. Admin-merged (self-review bypass) with full green CI. |
||
|
|
e6c16efa6d |
refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3) (#2382)
* refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3) Relocate 13 case arms from runCommand's switch to HOST_COMMAND_ROUTERS: - resolve: resolve-model, resolve-granularity, resolve-execution (3) - git: git (1) - config: config-ensure-section, config-set, config-set-model-profile, config-get, config-new-project, config-path, migrate-config (7) - research: research-store, research-plan (2) Each case body moved verbatim to a module-scope route*Command function (closures over config/commands/output/_dispatchNonFamily preserved). dispatchHostCommand extended to pass defaultValue + workstreamContext (needed by config-get and config-path). No new files, no logic change. Closes #2373 * chore: retrigger CI |
||
|
|
b0f672f88c |
fix(#2337): capture and surface todo severity (#2381)
add-todo.md gains a confirm-based infer_severity step (infer from the blocker/major/minor/cosmetic taxonomy, confirm via AskUserQuestion with TEXT_MODE fallback, before writing) and a severity frontmatter field. cmdListTodos and cmdInitTodos now surface severity, backward-compatible (key omitted when absent), in parity. Closes #2337. Admin-merged (self-review bypass) with full green CI. |
||
|
|
dcb4954131 |
fix(#2335): normalize volta node image paths to the stable shim (#2375)
Adds a volta branch to normalizeNodePath() that rewrites the version-pinned node image path to volta's stable shim, so managed hooks survive a volta node prune (the fnm/Homebrew/mise class, now covered for volta). Also unifies the release-smoke install timeout into one shared 600s constant across before() and runSmoke() so slow benches no longer spuriously time out. Closes #2335. Admin-merged (self-review bypass) with full green CI. |
||
|
|
b302f53ee6 |
refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) (#2370)
* refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) Behavior-preserving relocation of the 706-line case 'capability': arm from gsd-tools.cjs into a new hand-authored bin/lib/capability-command-router.cjs (sibling of ensure-runtime-build.cjs). The 15 bin/-relative require paths are rewritten to sibling-relative (correct for bin/lib/). dispatchHostCommand is now async (capability's install/upgrade/consent ops await the lifecycle); sync routers (state/phase/…) pass through await unchanged. case 'capability': removed; capability dispatches via HOST_COMMAND_ROUTERS. Validated by the existing capability-lifecycle / -consent / -trust / -loader test suites (no logic changed). Golden install-parity fixtures regenerated. Closes #2368 (Slice 1 — relocation). Probe consolidation (capHostVersion→ readHostVersion, capReadStrict dedup) deferred to a follow-up slice. * fix(#2368): add capabilityState/capabilityWriter requires + INVENTORY row The relocated capability arm references capabilityState (cmdCapabilityState, resolveCapabilityRuntimeState) and capabilityWriter (cmdCapabilitySet) — both module-scope requires in gsd-tools.cjs (L288/289) that the initial closure-dep scan missed. Added as sibling requires to capability-command-router.cjs. Also adds the new cli module to docs/INVENTORY.md + regenerates the manifest. * fix(#2368): correct capHostVersion __dirname depth for bin/lib/ relocation capHostVersion's VERSION/package.json paths were bin/-relative ('..' and '..','..'); on relocation to bin/lib/ they resolved one level too deep, so capHostVersion returned 0.0.0 and capability install failed the engines.gsd gate (#1920). Added one more '..' to each (now resolves gsd-core/VERSION and repo-root package.json correctly). * test(#2368): drop capability from the invocation loop (async/FS vs /fake/cwd) capability is async and does FS/config reads, so invoking it against the unit test's /fake/cwd is fragile. The 6 sync Tier-1 routers stay in the invocation loop; capability is covered by the non-invoking registry- ownership assertion + the dedicated capability-* test suites. * chore: retrigger CI (no-changelog label now present) |
||
|
|
67a9243cf1 |
chore(#2356): make the ADR index a generated artifact and enforce ADR lifecycle invariants (#2367)
* chore: rebuild ADR index as a generated artifact and enforce lifecycle invariants
The ADR index in docs/adr/README.md was hand-maintained with nothing checking
it, and had drifted to 40 of 65 ADRs. The absent rows included the entire
capability family (857/894/959/1016/1143/1213/1244) and ADR-1239 (EoS) itself,
so the decisions a reader most needed were the ones they could not find.
Make the index a derived artifact, matching the repo's existing generated-file
idiom (lint:generated-sync), and enforce the corpus' lifecycle invariants:
- scripts/gen-adr-index.cjs generates the index between markers and validates
the status vocabulary (Accepted/Proposed/Superseded/Legacy/Retired),
successor links, id/filename agreement, and supersession symmetry.
- Wire --check into lint:generated-sync so drift fails CI.
Correct the lifecycle metadata the gate surfaced, without flipping any status:
- ADR-1239 (EoS) declared it subsumed ADR-1016/58/3660/894; none recorded it.
Add reciprocal "Subsumed by" pointers + dated amendments. Subsumption keeps
the target Accepted -- these are live adapters, not dead decisions.
- ADR-857/894 carry dated status caveats: they read Proposed while the
capability system shipped and epic #857 is closed. Ratification is a
maintainer act and is deliberately left open.
- Link ADR-0005/0007/0012/3524 -> ADR-0174 and ADR-0010 -> ADR-0009; record
the reciprocal Supersedes on ADR-0009.
- ADR-218 declared itself "ADR-0175" -- an unfinished rename.
- The 0011 PRD moves from the non-canonical "Draft" to "Legacy".
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: capture stderr via spawnSync; record ADR-0010 draft supersession
Two fixes surfaced by the first gsd-test run and by regenerating the index:
- tests/adr-index-gate.test.cjs used execFileSync, which only surfaces stderr
through the thrown error on non-zero exit. The `--write` path exits 0 while
reporting outstanding violations on stderr, so the helper always saw ''.
spawnSync captures both streams on both outcomes.
- The hand-maintained index recorded 0010-skill-surface-budget-module.md as
"earlier draft superseded by ADR-0011" while the file itself still said
Proposed. Deriving the index from the files would have dropped that
assertion and resurrected a superseded draft as a live decision, so it is
recorded at its source, with the reciprocal Supersedes on ADR-0011.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: drop the dead sdk/ model-catalog candidate retired by ADR-0174
src/model-catalog.cts resolved model-catalog.json through three candidates, the
second being sdk/shared/model-catalog.json three levels up. That was the legacy
source-repo fallback kept by the #3288 fix ("check the co-located path FIRST,
before the legacy source-repo path").
ADR-0174 then retired the @opengsd/gsd-sdk package boundary and deleted the sdk/
tree (
|
||
|
|
cf004df678 |
refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) (#2364)
* refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) Pilot cutover for ADR-2346 Phase 1 (epic #2345). Introduces the Layer-2 host dispatch table — dispatchHostCommand + HOST_COMMAND_ROUTERS, consulted in runCommand's default case after capability/overlay dispatch, before the unknown-command error. Migrates 'state' as the pilot: removes the hardcoded case 'state': arm; state now dispatches default -> dispatchHostCommand -> routeStateCommand, byte-identical to the old path (proven by the new state-command-cutover equivalence test, 5-category template). Host commands are NOT capabilities (core, non-toggleable, no tier/activationKey) — the capability registry stays reserved for toggleable feature bundles per ADR-959. This is the host-vs-capability distinction the merged ADR-2346 lacked; the ADR is corrected here alongside the code that realizes it. - gsd-core/bin/gsd-tools.cjs: HOST_COMMAND_ROUTERS + dispatchHostCommand (prototype-pollution-safe); wired into default case; case 'state': removed; dispatchHostCommand + HOST_COMMAND_ROUTERS exported for tests. - tests/state-command-cutover.test.cjs: UNIT/DISPATCH/BEHAVIOR/REGISTRY equivalence (recording-mock + runGsdTools end-to-end + pollution guard). - docs/adr/2346-*.md: refine Decision 1/2 to the host-table vs capability- registry model (correction that did not land in the merged #2355). Behavior-preserving. Subsequent P1b/c PRs migrate phase/init/roadmap/validate/ verify using this proven template. Closes #2360. * test(#2360): regenerate golden fixtures + allowlist for state cutover Bookkeeping for the gsd-tools.cjs change: npm run gen:golden regenerates the install-parity fixtures (gsd-tools.cjs content hash changed), and the new tests/state-command-cutover.test.cjs is added to the lint-test-file-count allowlist under the 'state' prefix. * refactor(#2360): migrate remaining Tier-1 routers (phase/init/roadmap/validate/verify) Completes P1: all 6 Tier-1 host routers now dispatch via HOST_COMMAND_ROUTERS (state landed in the pilot commit). init preserves its #1688 warnIfStaleBake pre-hook; validate binds the output emitter. Cutover test extended to assert all 6 are consumed + owned. Golden install-parity fixtures regenerated. |
||
|
|
f1a91072b6 |
docs(#2357): fix Registry Discussions category name and state its format (#2361)
The submission process told an admin to create a `Registry` Discussions category and never stated its format. Both were wrong in a way a correct reading of the docs could not catch. Name: the category is `EoS Registry`, and it carries threads for both registries — `discussion` is a required field on Capability entries as well as EoS entries. A reader following the old text would create a second, duplicate category. Format: `discussion` being required means the thread must exist before the entry's PR, opened by the entry's author — an outside contributor holding neither `maintain` nor `admin`. GitHub's Announcement format restricts starting discussions to those two levels, so an admin could pick it, block every external submission at the first step, and see nothing wrong. Records open-ended as the required format, and why Announcement and Question/Answer do not work. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ed06b6a4b9 |
fix(#2329): write opencode slash commands to commands/ (plural), migrate legacy command/ (#2354)
* test(#2329): fail-first tests for opencode commands/ (plural) command dir Red phase, empirically probed: global/local install lands in command/ (singular) with 71 gsd-*.md files and no commands/; the manifest records 71 keys under command/ and zero under commands/; all four declaring sites report 'command'. Migration coverage is black-box (two sequential install runs against one configDir) so it holds regardless of how the fix implements cleanup. The Kilo guard passes today by design — a forward-looking no-collateral check. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): write opencode commands to commands/ (plural), migrate legacy command/ OpenCode discovers slash commands from commands/ (plural); the installer wrote them to command/ (singular), so none of the ~71 /gsd-* commands appeared in the TUI. Five sites declared the directory and all had to agree: - capabilities/opencode/capability.json: both artifactLayout destSubpath entries (global + local) and hostBehaviors.flatCommandDir - bin/install.js: the manifest prefix was a SEPARATE hardcoded 'command/' literal, so the manifest would have diverged from the descriptor even after a rename. It now derives from _hostBehaviors(runtime).flatCommandDir. - src/install-engine.cts installOpencodeFamilyArtifacts: the actual write target, which bypasses resolveRuntimeArtifactLayout via combinedFamilyInstall. This was a fifth site the issue did not list — without it the descriptor change alone would not have moved a single file. Migration: an upgrade over a pre-fix install removes only manifest-proven GSD-managed files from the legacy command/ dir and rmdirs it once empty. Unmanifested user files are preserved, never deleted. Kilo shares the opencode family install path and is explicitly unaffected — pinned by a no-collateral test. Note on the tests: the migration cases originally built their legacy fixture by running the installer and relying on it to produce command/ — i.e. they depended on the bug to set up the fixture, and became unsatisfiable the moment it was fixed (block 1 requires command/ to be absent after a fresh install). They now fabricate the legacy layout explicitly, including rewriting the manifest keys to the command/ prefix — which is load-bearing, since the migration only removes manifest-proven files and an unrewritten fixture would silently no-op and pass even against a broken migration. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): regenerate opencode install golden after rebase onto next The golden conflicted on rebase because #2322 also regenerated it. Resolved by regenerating from the merged source rather than hand-merging a generated file; the only delta is the 71 command/gsd-*.md -> commands/gsd-*.md key renames. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): update stale tests that pinned opencode's singular command/ dir Seven tests encoded the old contract (opencode: command/gsd-help.md exists, the descriptor's flatCommandDir, the install-integration contract, and the resolveRuntimeArtifactLayout golden). They passed in the red phase precisely because they pinned the buggy singular dir; the fix intentionally changes that contract, so these are stale-test corrections, not regressions. Kilo shares the opencode family install path and is deliberately NOT changing — it stays on command/ (singular). The shared opencode/kilo test is now split via an explicit per-runtime dir map so the two cannot be conflated, and Kilo's own layout test is untouched. tests/opencode-command-dir-plural.test.cjs independently pins Kilo unchanged end-to-end. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): changeset for opencode commands/ dir fix Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): backfill PR number 2354 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): correct the changeset — do not assert opencode ignores command/ The changeset repeated the issue's stated mechanism ("OpenCode discovers them from commands/ ... a clean install produced no usable commands in the TUI at all"). OpenCode's source contradicts that: packages/core/src/v1/config/command.ts globs {command,commands}/**/*.md, so BOTH names resolve, and its own skill doc still calls .opencode/command/ typical. Shipping that claim as a release note would document a mechanism that does not exist. The change is still right, for the stronger reason: OpenCode's config docs list plural as the convention and singular as backwards compatibility, so GSD was shipping on the alias the vendor may withdraw. Reworded to describe it as the alignment it is, decided on OpenCode's source and docs rather than on bug reports in either repo. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): baseline opencode's commands/ surface — closes a data-loss path this PR opened Not a bookkeeping gap. Moving opencode's command dir to commands/ moved the install destination to a surface the first-time baseline scan does not cover: 000-first-time-baseline's RUNTIME_SURFACES.opencode lists ['gsd-core','command', 'skills','agents'] — no 'commands'. installOpencodeFamilyCommands unconditionally unlinks every gsd-*.md under its destination before writing the fresh set (install-engine.cts:870-873), with zero manifest or migration involvement. The only thing that protects a pre-existing file is assertInstallerMigrationsUnblocked, which runs before materialization and halts when the baseline scan flags an unknown file at a KNOWN surface. Probed: a pre-existing commands/gsd-plan.md is silently destroyed (install exits 0). The identical file under the legacy, already-baselined command/ surface correctly halts the install with "installer migration blocked pending user choice". So this PR would have traded a protected surface for an unprotected one. Fixed with a NEW fix-forward migration rather than editing 000, per docs/installer-migrations.md:131-134 — an applied migration never re-runs, so editing 000 would only protect fresh installs and leave every existing machine exposed. A new id runs for both populations and drifts no shipped checksum; adding its entry to EXPECTED_CHECKSUMS is the case that test explicitly sanctions. All five pre-existing shipped checksums verified byte-identical. Kilo is excluded by the migration's runtimes filter and keeps command/. This was previously deferred as a PR-body note claiming "low impact — nothing else acts on baseline-scan misses". That claim was never probed and was wrong. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): drop the parenthetical product description from the changeset The product-name purity guard (#1777) rejects "Kilo (which still uses command/)" — fragment prose renders verbatim into CHANGELOG.md, so a product name must not carry a parenthetical. Reworded to a plain sentence; the meaning is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
15b3cc8690 |
docs(#2346): Command Dispatch Completion ADR + graduate ADR-959 to Accepted (#2355)
Records the decision (ADR-2346) to dissolve runCommand's 73-case switch into a two-layer dispatch (registry families + leaf-verb table filling the prepared _dispatchNonFamily seam), collapsing it to ~15 lines. Covers the four decisions ADR-959 leaves open: full dissolution, family/leaf classification rule, shared parseFamilyArgs, and the capability-arm extraction shape. Phased under epic #2345 (P1-P4). Behavior-preserving; each cutover proven by the audit-command-cutover equivalence template. - docs/adr/2346-command-dispatch-completion.md (new) - docs/adr/959-*.md: Status Proposed -> Accepted + amendment section - docs/adr/README.md: index rows for 959 + 2346 - docs/ARCHITECTURE.md: forward-reference note under Command Routing Hub - CONTEXT.md: seed glossary entry Closes #2346 (docs-only; no production code). |
||
|
|
23a65c4a3d |
fix(#2322): materialize installed third-party capability skills (#2340)
* test(#2322): fail-first tests for third-party capability skill materialization Red phase: tests (1) and (6) fail — resolveSurface reports the third-party stem surfaced (#2045) but no SKILL.md is ever written to disk. The other four are controls that must keep holding: first-party-wins collision, profile-tier filter, nested-router layout unperturbed, and absent/malformed capability must not throw. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): materialize installed third-party capability skills A capability could report installed:true, surfaced:true, active:true and still never exist as an invocable command. #2045 fixed the registry layer — resolveSurface unions registry.capabilityClusters into the resolved skill set — but the materialization layer never got the matching fix. stageSkillsForRuntimeAsSkills only ever read gsd-core's own bundled commands/gsd/*.md and silently skipped any stem it couldn't find there, so a third-party skill living at <GSD_HOME>/.gsd/capabilities/<id>/skills/<stem>/ was never copied. Registry said surfaced; disk had nothing. Installed capability skills are now staged alongside the first-party ones, copied verbatim (they are authored complete for their target runtime and need no converter). First-party stems always win a collision, the profile filter still applies, and an absent or malformed capability degrades rather than throwing. Security: capability.json's skills[] entries are validated only as non-empty non-reserved strings (capability-validator.cjs:503-514) — no path shape is enforced upstream — so stems are sanitized (rejecting separators, '..', absolute paths, NUL) with an independent isPathConfined check on both the read and write paths. A '../../evil' stem writes nothing outside the capability's own dir. Also fixes a defect this surfaced in pruneSkillDirs: a materialized capability skill dir has no first-party manifest entry, so every apply logged "preserving (user-owned or unknown)" for a live GSD-managed dir. The retained check now precedes the manifest gate; no deletion outcome changes, and genuinely unknown gsd-* dirs still warn and are preserved. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): address security review — bind skills to declaring capability, fix full profile An independent security review BLOCKED the first pass. Both blockers were mine. BLOCKER 1 (security): readInstalledCapabilitySkill scanned every capability dir and returned the first sorted match, never checking that a capability DECLARES the stem — ownership was inferred from attacker-controlled filesystem layout. Since install copies the whole bundle and the validator only checks DECLARED entries, a capability declaring `skills: []` could ship an undeclared skills/deploy/SKILL.md and win the `deploy` stem on sort order, supplying the agent-invocable instructions the user believed came from the registered capability. Stems are now bound to their owning capId via registry.capabilityClusters, and only that capability's dir is read. BLOCKER 2: the fill-in pass was gated `skills !== '*'` on the premise that applySurface materializes `full` into a concrete Set. True for applySurface — false for the installer, which is the default path: resolveProfile returns the '*' sentinel and bin/install.js passes it straight to staging. So #2322 survived on the default `full` profile, i.e. the fix didn't fix the reported bug. The registry is now plumbed to staging, and '*' stages all capability-cluster stems. Wiring this surfaced a second gap: the ADR-1239 imperative adapter (the primary install path) never threaded its registry either, which would have silently defeated the fix on the real default install. HIGH: staged capability skills were never prunable — pruneSkillDirs gates on the first-party manifest, so uninstalling a capability left its instructions live in the agent's context forever. Staged skills now carry a marker making them GSD-owned and prunable; genuinely unknown gsd-* dirs still warn and are preserved. MEDIUM: the "staged verbatim" claim was false — applySurface rewrites bodies over the whole stage dir. The tests asserted byte-equality and passed only because their fixtures contained no rewrite triggers. Claim dropped; tests now assert the rewrite against triggering content. LOW: isPathConfined is lexical, not realpath (symlink-defeatable, currently unreachable because install rejects symlinks) — comment corrected. The validator does not enforce non-empty, so isSafeCapabilitySkillStem is the sole defense, not a second layer — comment corrected and it now has traversal/NUL/absolute/empty test coverage (previously mutating it to `return true` left every test green). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2322): pin that the imperative adapter forwards a capability registry The delegation-args test deep-equalled the exact argv to installRuntimeArtifacts, so threading the composed capability registry through the ADR-1239 imperative adapter (required for #2322 — without it the default `full` install path never materializes third-party capability skills) failed it. The contract legitimately gained a parameter, so this is a stale-test correction, not a regression. Rather than deep-equalling the whole composed registry (brittle — it embeds the full agent/profile map), the test pins the leading args exactly and asserts only that a registry-shaped value is forwarded. That still fails if the adapter stops threading it, which is the regression the test exists to catch. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2322): backfill PR number 2340 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
19c1f54a2a |
fix(#2316): stop phase complete silently dropping ghost requirement IDs (#2339)
* test(#2316): fail-first tests for ghost REQ-IDs, v-heading over-match, all-orphan gap check
Red phase: 4 of 10 fail against current source (#2316-1 ghost-ID warning,
-3 requirements_updated honesty, -4a v1-heading suppression, -6b all-orphan gap
rows). The other 6 are controls/boundaries that must keep passing — including the
#1159 deferred-heading guard and the literal "TBD" placeholder boundary.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA
* fix(#2316): stop phase complete silently dropping ghost requirement IDs
phase complete parses a phase's `**Requirements**:` line from ROADMAP.md and
reconciles it into REQUIREMENTS.md. When a cited ID was registered nowhere, every
branch degraded to a no-op and the report was indistinguishable from a run that
applied every update: `requirements_updated: true, warnings: [], has_warnings:
false`, file byte-for-byte unchanged.
Four defects on that path, all long-standing (traced to
|
||
|
|
ada79bee97 |
fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md (#2338)
* fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md Step 4 rewrote the `## Current Milestone` heading in the shared root PROJECT.md unconditionally. references/workstream-flag.md marks PROJECT.md `# Shared`, and per-workstream milestone state already lives in the workstream's own STATE.md / ROADMAP.md / REQUIREMENTS.md. With parallel milestones — the sanctioned design — whichever workstream ran new-milestone last silently won the shared heading. Step 4 is now skipped when a workstream is active; step 6 no longer stages PROJECT.md in that mode (cmdCommit returns nothing_to_commit rather than failing when a staged path is unchanged). Also fixes a second defect found while diagnosing this, same root cause (the workflow was workstream-unaware): step 1 parsed only --reset-phase-numbers and the milestone name, so GSD_WS was never set — yet ${GSD_WS} was interpolated at the routing lines. It always expanded to empty, so `/gsd:new-milestone --ws x` suggested `/gsd:discuss-phase [N]` with the workstream scope silently dropped, violating the routing-propagation contract. Step 1 now parses --ws using the established idiom from verify-work.md. Guard is keyed on GSD_WS, not $GSD_WORKSTREAM: the runtime launcher does not export the latter and it is only priority 2 of 5 in resolution, so it would miss the --ws flag case that is the actual repro. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2308): regenerate install goldens for the new-milestone workflow change gsd-core/workflows/ ships as an installed artifact, so new-milestone.md's content hash is pinned in all 18 runtime golden fixtures. Only that hash changed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2308): address review — inert step-6 guard, dropped Evolution repair, tautological tests Independent review found the first pass was partly cosmetic: 1. The step-6 `if [ -n "$GSD_WS" ]` branch was INERT. GSD_WS is assigned in step 1's shell and each step's bash block runs in its own shell — this file already proves it, since step 5 round-trips OUTGOING_MILESTONE through a file for exactly that reason (#2288). The guard read an unset variable, always took the flat branch, and staged PROJECT.md anyway. Rather than re-deriving GSD_WS in step 6, the branch is removed entirely: step 4 Part A's guard is what protects the shared heading, so post-guard the only content PROJECT.md can carry is Part B's idempotent Evolution backfill — which must be staged, not stranded. A regression test now asserts no cross-step GSD_WS branch returns. 2. Skipping ALL of step 4 also dropped the `## Evolution` structural repair — a shared, idempotent backfill that is not workstream state. A pre-Evolution project running only `--ws` would never get the section that transition and complete-milestone expect. Step 4 is now split: Part A (milestone-state write) is workstream-guarded; Part B (Evolution) always runs. 3. The tests were tautological prose-pinning — including one asserting a comment mentions "#2308". The step-6 test asserted the guard's TEXT was present, so it passed on the inert guard it existed to catch. Replaced with executable tests that extract the step-1 and step-6 fences and run them under bash with stubbed gsd_run, asserting real parse and --files behavior. 4. --ws is now stripped from the milestone name (step 1 previously left "--ws search" in the remaining text), and documented in argument-hint, help/modes/full.md, and docs/COMMANDS.md. 5. Changeset no longer overstates: --ws reaches the prose guard and routing hints only, not the SDK calls (state.milestone-switch/phases.clear/init.new-milestone still take no ${GSD_WS} — out of scope here). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * chore(#2308): regenerate SKILL.md, goldens, and size baseline for the argument-hint change skills/gsd-new-milestone/SKILL.md is generated from commands/gsd/new-milestone.md, so documenting --ws in the argument-hint made it stale (caught by lint:ci's gen-plugin-skills --check). Regenerated it plus the install goldens and workflow size baseline, since commands/, skills/, and gsd-core/workflows/ all ship. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2308): backfill PR number 2338 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b2961c3f69 |
fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) (#2336)
* test(#2070): fail-first tests for adaptive model_profile and models tier validation Encodes the three acceptance criteria from #2070 plus the boundary cases the resolver silently ignores today (non-string values, empty string, mistyped phase-type key), and pins VALID_TIERS to a catalog-derived set. Red phase: these fail against current src/ by design. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) W004 sourced its profile list from a hand-maintained literal that predated the adaptive profile, so `"model_profile": "adaptive"` was false-flagged. It now reads VALID_PROFILES, which model-catalog.cts derives from model-catalog.json. models.<phase_type> was validated nowhere: the resolver's tier gate silently drops unknown values, so a typo like `"planning": "opuss"` was an undiagnosable no-op. A new W022 flags unknown phase-type keys and invalid tier values (including non-string values, which the same gate also drops). VALID_TIERS moves from a function-local literal in model-resolver.cts to a catalog-derived export, so health and the resolver cannot disagree by construction rather than by parity test. Object.values(adaptiveTierMap) is ['opus','sonnet','haiku'] plus 'inherit' — identical to the previous literal, so resolution behavior is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2070): changeset for validate health adaptive profile + W022 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2070): close review findings — malformed models, tier-list duplication, changeset gate Review of the initial fix surfaced three real defects, folded in per the no-defer rule: 1. verify.cts: the W022 guard skipped a top-level `models` that is present but not a plain object (`[]`, `"opus"`, `5`, `true`). The resolver ignores those identically, so they were the same undiagnosable no-op #2070 targets — just one level up. They now warn; absent/null/{} stay silent. 2. config-loader.cts: RUNTIME_OVERRIDE_TIERS was a second hardcoded copy of the tier vocabulary this change had just de-hardcoded elsewhere. It now derives from the catalog via ADAPTIVE_TIER_VALUES (no 'inherit' — runtime overrides resolve to a concrete tier). Byte-equivalent to the old literal. 3. scripts/changeset/lint.cjs: USER_FACING_PREFIXES omitted `src/`. Post-ADR-457 the product source is src/*.cts compiled to a gitignored gsd-core/bin/lib, so the `gsd-core/` prefix is dead coverage for library code and a src/-only PR could merge with no release note — including this one. Adding `src/` closes the gate; tests/ stays non-user-facing. Also corrects a false docstring in the VALID_TIERS test: value-equality cannot detect a re-hardcoded literal, so the test no longer claims it does. Two review findings were rejected with evidence rather than actioned: - W021 double-allocation is governed by ADR-612 ("W021 renumber -> void ... kept, message-disambiguated"), not a defect. - Global-defaults validation would be a false-positive generator: config-loader reads ~/.gsd/defaults.json only on the "no .planning/" branch, and health early-returns E001 without .planning/, so those values provably never affect resolution in any context health can run. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2070): regenerate install goldens for the changeset-lint change scripts/ ships as an installed artifact, so scripts/changeset/lint.cjs's content hash is pinned in all 18 runtime golden fixtures. Adding 'src/' to USER_FACING_PREFIXES changed that hash and tripped every golden parity check. Regenerated via `npm run gen:golden`; the only delta is the lint.cjs hash. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2070): backfill PR number 2336 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a30fb75b51 |
fix(#2068): dynamic routing escalates the model per --attempt (#2334)
* fix(#2068): dynamic routing escalates the model per --attempt, not just effort cmdResolveExecution resolved the model via resolveModelInternal (which ignores dynamic_routing), so retries escalated effort but the model stayed pinned to the default tier. Resolve the model via resolveModelForTier when --attempt is given, gated identically to the effort resolution so model and effort stay symmetric — an omitted --attempt keeps the classic profile model (unchanged for everyone, incl. dynamic_routing users who don't pass --attempt). Escalation is capped at max_escalations. Falls back to resolveModelInternal when dynamic_routing is off. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2068): backfill PR number 2334 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1794acb255 |
chore(#2331): trigger PR-policy workflows on pull_request_target so fork PRs get the verdict (#2333)
* chore(#2331): trigger PR-policy workflows on pull_request_target Three PR-policy workflows (pr-title-validator, pr-target-validator, require-issue-link) triggered on plain `pull_request`, so a fork PR's GITHUB_TOKEN was downgraded to read-only regardless of the declared `permissions:`. Each one comments on the PR and THEN emits its verdict, so the createComment 403 killed the github-script step before core.setFailed ran: the contributor saw an API stack trace instead of the instructions the comment exists to deliver. Confirmed on PR #2084 (job 86573823878), whose title has been non-compliant since 2026-07-08 while the explanatory comment 403'd on every run. Switches all three to pull_request_target (base-repo context, write-capable token), matching the three siblings that already do this correctly (pr-template-format, close-draft-prs, auto-close-unsolicited-prs). Safe: the only checkouts are BASE-branch with persist-credentials: false, and every PR-controlled input is read as data — no head code executes. Also wraps each comment in try/catch so a comment failure can never again suppress the verdict. Extends tests/workflow-maintainer-skip.test.cjs with the trigger lock already applied to close-draft-prs.yml (:32-42) for this same defect class. Closes #2331 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2331): strip backticks before echoing untrusted text into bot comments Found by the orthogonal security review of this change. pr-title-validator and pr-target-validator echo attacker-controlled text (the PR title; the fork's branch name) into an inline-code span in a comment posted by github-actions[bot]. A single backtick closes the span early and the remainder renders as live Markdown — GFM autolinks a bare URL — so a fork author could make our own bot post an arbitrary clickable link into a PR thread, borrowing the bot's credibility for phishing. This interpolation is unchanged from next, but it was NOT previously reachable from forks: the createComment call 403'd and the comment was never posted. The trigger switch in the parent commit is what makes it reachable by untrusted authors for the first time, using the write token it grants — so it is in scope here and fixed here rather than deferred. A PR title has no charset restriction, so that vector is fully exploitable. The branch-name vector is weaker (check-ref-format forbids space, ':', '[' and '*', so no bare URL, link or emphasis is expressible) but is the same class and is stripped identically rather than left to the charset to police. Stripping the backtick is complete: it is the only character that can break out of an inline-code span. Only the rendered body needs this — core.warning/setFailed go to the job log, where @actions/core already escapes workflow commands. Also strengthens the try/catch test to assert core.setFailed sits AFTER the catch block rather than merely existing, so moving the verdict inside the try (the exact inversion #2331 fixes) fails the test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2331): assert verdict ordering on code, not on comment prose The first cut of the verdict-ordering guard failed against correct code. It used indexOf('core.setFailed') on raw source, and these workflows name core.setFailed in their own comments while explaining the bug — at lines 29/118/147, 16 and 6, all BEFORE the catch block. So the assertion compared a comment to the call and reported the inversion it was written to catch. gsd-test caught it: 4 unique failures across linux-node22/24. The code was right; the test was measuring the wrong text. Fixes: - readWorkflowCode() strips whole-line YAML/JS comments so positional assertions see only executable text. - The ordering check is extracted to verdictSurvivesCommentFailure() and exercised against BOTH a good and an inverted sample, so the guard is proven non-vacuous rather than merely passing. - A test pins the trap itself: raw source really does mention core.setFailed before the catch, while the stripped view puts the real call after it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2331): drop the unnecessary permission widening; the trigger was the whole bug Both orthogonal review passes flagged the permissions block, from opposite directions — one said issues:write was dead surface on require-issue-link, the other said it was the load-bearing scope the two validators lacked. Neither is right, and the repo's own history settles it: - pr-title-validator declares pull-requests:write ONLY, and its sticky comment has posted 26 times. - pr-target-validator declares pull-requests:write ONLY — posted 8 times. - require-issue-link declares issues:write ONLY — posted on same-repo PRs #106, #164, #232, #259. So GitHub accepts EITHER scope for issues.createComment when the target is a PR, and all three files already declared a sufficient one. The 403 was purely the fork token downgrade. My added scopes fixed nothing and widened privilege on precisely the workflows now running as pull_request_target — the context where surplus scope matters most. Reverted: permissions are byte-identical to next, and the diff is now trigger + try/catch + sanitizer only. The permission test previously used an (issues|pull-requests) alternation, so it passed on the pre-fix tree and would not have caught removal of the scope that matters. It now asserts each file's SPECIFIC scope and, more usefully, asserts the absence of the other — locking the least-privilege property against a future 'add it to be safe' regression. It is a forward lock, not a #2331 fails-first test; the trigger assertion is the fails-first one. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9ad2bab4be |
fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime (#2332)
* fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime The installer writes resolve_model_ids:"omit" for non-alias runtimes into the machine-wide ~/.gsd/defaults.json (#1156); any runtime read it back, so install order silently flipped Claude's adaptive tier aliases (executor->sonnet, planner->opus) to '' in no-project sessions. Resolution is now scoped to the runtime actually resolving, identified by a new per-install <install>/gsd-core/.gsd-runtime marker (installer writes it beside VERSION). The "omit" branch returns '' only when the PROJECT explicitly set omit (honored for all runtimes, #2517 finding #4) OR the active runtime lacks native aliases. Claude ignores a global-defaults-only omit and keeps its aliases; the active runtime is canonicalized (GSD_RUNTIME -> config.runtime -> marker -> claude) so alias/case spellings can't defeat the check; explicit project omit is workstream/project-scope aware; explicit true still materializes IDs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2297): backfill PR number 2332 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1bb724048a |
fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist (#2325)
* fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist The convergence reviewer-flag whitelist predated the 1.7.0 Antigravity CLI adapter and silently dropped --agy/--antigravity, so convergence fell back to --codex only and the working adapter was unreachable (worse after Gemini CLI's upstream shutdown). Add both flags to the workflow grep whitelist, the command argument-hint + flag docs, and the regenerated SKILL.md; they pass through to /gsd-review unchanged. --gemini behavior is untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2293): backfill PR number 2325 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5ea3401d4f |
fix(#2289): context-monitor emits injection envelope only for supported events (#2324)
* fix(#2289): context-monitor emits injection envelope only for supported events gsd-context-monitor is wired to Codex Stop/SubagentStart/SubagentStop/PreCompact (#772), but it emitted a hookSpecificOutput.additionalContext envelope for every event. Codex's Stop schema rejects that shape ("hook returned invalid stop hook JSON output") exactly when context is low. Use a positive allowlist: emit only for context-injection events (PostToolUse, AfterTool, and the pre-existing Gemini missing-name fallback); exit 0 silently for Stop and every other event. Debounce and critical-session bookkeeping still run on silenced events. Behavioral regression tests drive the real hook. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2289): backfill PR number 2324 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2285): de-flake per-plan use_worktree property test on duplicate briefs The fast-check property in claude-orchestration.test.cjs located each plan's agent() call via indexOf on the brief, but only plan IDs were unique — briefs could collide (fast-check shrinks toward short strings). On a colliding seed the lookup found the first duplicate's line and misattributed its isolation, failing "use_worktree:true must carry isolation" intermittently (surfaced on macOS CI). emitWorkflowScript is correct for duplicate briefs (verified); suffix the unique id onto each agent() label so the test probe is unambiguous. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
52fab7d9d7 |
fix(#2288): archive phase history under the outgoing milestone version (#2323)
* fix(#2288): archive phase history under the outgoing milestone version phases.clear derived its archive directory from a live getMilestoneInfo() read, but new-milestone.md switches the milestone BEFORE phases.clear runs, so phase history was filed under the NEW milestone's <version>-phases/ dir. Add a --archive-version override (threaded from new-milestone.md, captured before the switch) with precedence override -> live read -> dated label. Harden the version label against path traversal on both phases.clear and the sibling milestone-complete sink (the label is a moved directory name), and persist the outgoing version via a file + quoted shell expansion so untrusted STATE.md content is never re-parsed by the shell. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2288): backfill PR number 2323 into changesets Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b041f101fb |
fix(#2287): surface unresolved deferred-items.md entries in progress + audit-uat (#2318)
The executor SCOPE BOUNDARY convention (agents/gsd-executor.md) logs out-of-scope discoveries to a phase directory's deferred-items.md, but no reader ever consumed it — forensic_audit, cmdAuditUat, and capture --list all skipped it — so deferred items were permanently invisible. cmdAuditUat (src/uat.cts) now scans each phase dir's deferred-items.md via a new parseDeferredItems (reusing the collectSection/splitGapsEntries/ extractGapEntryFields seams) and surfaces entries whose status != resolved (fail-safe: a missing/garbled status surfaces rather than hides, matching the false-negative-averse posture of #2286). forensic_audit (gsd-core/workflows/progress.md) gains Check 7 that globs .planning/phases/*/deferred-items.md and reports unresolved entries. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a22333034b |
fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml (#2312)
* fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml generateCodexAgentToml embedded a per-agent `model_overrides` value verbatim as the Codex `.toml` `model`, leaking GSD/Claude tier aliases (opus/sonnet/haiku/fable) and `claude-*` ids. Codex/ChatGPT rejects those (400 "The 'sonnet' model is not supported when using Codex with a ChatGPT account"), and since spawn_agent has no inline model param, the model is baked into the .toml at install time — so the orchestrator could not recover and fell back to the non-equivalent generic-agent workaround. Translate a GSD tier alias through the Codex tier map (sonnet -> gpt-5.6-terra); drop with a deduped warning any Anthropic-flavored value with no Codex mapping (fable) or a `claude-*` id, so emission falls through to the runtime-aware resolver or Codex's default. A final safety gate blocks an Anthropic-flavored model from the runtime- resolver path too (runtime/target mismatch). Mirrors the Claude-side override guard (#2041). Real Codex/OpenAI model ids in model_overrides still pass through verbatim (#2256 preserved); runtime:"codex" tier resolution unchanged (#2517). Adds regression + fast-check property tests in tests/codex-config.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2310): backfill changeset PR number to #2312 * fix(#2310): Codex passive-model posture — omit Anthropic-flavored model (all namespacings) Adopt ADR-1239's passive/session-only posture for Codex model handling: a Codex agent .toml `model` is embedded ONLY for an explicit real-Codex model_overrides pin; any Anthropic-flavored value is omitted so the agent inherits the always- available session model (never a 400). - model_overrides tier alias (opus/sonnet/haiku/fable) or a Claude model id → omit (was: translate to gpt-*); an explicit real-Codex model id → embed verbatim (#2256 preserved). - Detect ALL Anthropic namespacings, not just `claude-*`: single-source the canonical CLAUDE_AGENT_ALIASES from model-resolver.cts and treat any id whose value contains "claude" (case-insensitive) as Anthropic-flavored — catching `anthropic/claude-*` and `us.anthropic.claude-*` (the forms the catalog assigns to opencode/hermes/kilo), which reach a Codex .toml via the runtime-resolver path on a mixed-runtime + Codex install. - The final safety gate applies to the runtime-resolver path too. The full passive posture (removing #2517's runtime-resolver per-tier embedding + a correctness health-check + a Codex TOML sync path) is tracked as the ADR-2310 epic #2313. Regression + fast-check property tests in tests/codex-config.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
636316f720 |
fix(#2286): audit-uat surfaces Gaps section + frontmatter/heading verification items (#2317)
parseUatItems only scanned '### N.' expected/result blocks and parseVerificationItems only recognized table/bullet/numbered shapes, so audit-uat returned a false-clean total_items:0 when a file recorded open findings in a '## Gaps' section, declared items in a frontmatter human_verification: array, or used the '### N. <label>'+bold-paragraph verification shape. parseUatItems now also scans '## Gaps' (via collectSection + iterateBullets) and surfaces any entry whose status != resolved. parseVerificationItems now treats the frontmatter human_verification: array (via extractFrontmatter) as the primary source when present, and adds a tokenizeHeadings fallback for the '### N.'+bold-paragraph shape, preserving the existing table/bullet/numbered recognition (no double-count, no regression). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ff9cb6069f |
fix(#2285): wire claude-orchestration Workflow backend into execute-phase (#2314)
The claude-orchestration capability (#1143) shipped registered 'active' but fully inert: detectWorkflowBackend/emitWorkflowScript had no caller outside their own CLI router, and execute-phase.md declared an execute:wave:pre hook point that the workflow body never rendered — so claude_orchestration.enabled:true had zero effect on real runs. Approach B (maintainer-chosen): - execute-phase.md now renders the execute:wave:pre hook (gsd_run loop render-hooks execute:wave:pre) at a new step 2.75, immediately before each wave's Agent() dispatch — fixing the latent dead-hook gap for any pre-wave capability. - Move the claude-orchestration contribution execute:wave:post -> execute:wave:pre (a pre-wave backend selector belongs before dispatch, not after); rename fragments/execute-wave-post.md -> execute-wave-pre.md with prose instructing the orchestrator to call resolve-wave-dispatch before step 3. Unrelated wave:post contributions (ui.safety-gate, drift, external-job, mempalace) untouched. - New .cts seam resolveWaveDispatch(input) composes detectWorkflowBackend + emitWorkflowScript into one {backend:'inline'|'workflow', ...} result; exposed as gsd-tools claude-orchestration resolve-wave-dispatch. This is a real non-CLI-router, non-test caller of both functions. Fail-closed: any gate miss (disabled, non-Claude runtime, Workflow tool absent, SDK below floor, execution_backend:inline, malformed input) or an emit failure resolves to inline with a byte-identical result shape — no regression to the default-off execute-phase path. Regression tests (tests/fix-2285-*) cover happy-path activation + SDK-floor BVA, the fail-closed gate-miss table with detectWorkflowBackend parity, a fast-check composition property, capability.json contribution assertions, and a source-contract guard that execute:wave:pre is now actually rendered. Dependent registry-shape assertions updated in-scope. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
75bedf16fd |
fix(#2284): project Hermes named dispatch onto delegate_task; protect comparison tables (#2309)
Hermes installs brand-swapped "Claude Code" -> "Hermes Agent" in shipped workflows/*.md but never projected the Agent(...) dispatch calls onto Hermes's delegate_task contract, so installed workflows kept literal Agent(...) syntax and falsely asserted "The Agent tool IS available" (Hermes exposes delegate_task, not Agent). Dispatch projection: a generic named-dispatch engine (projectNamedDispatchToStructuralDelegate) wired into the per-runtime RUNTIME_CONTENT_DISPATCH.hermes.md converter, branching entirely on the documentation-sourced hostIntegration.dispatch facts read via _hostIntegrationDispatch (capability.json unchanged): namedDispatch:false -> resolve the gsd-* role and embed a load-its-prompt instruction in the payload; background:true -> map onto delegate_task background; read-only / maxDepth:1 -> no nested delegation to leaf roles; per-call model dropped. Span detection uses literal Agent( scanning + local balanced paren/quote matching (immune to upstream document quote imbalance) and handles all three corpus call forms (multi-line, object-literal, single-line compact). An independent, mask-free post-projection guard fails the install loud on any residual Agent(/subagent_type/leaked model. Fail-closed: install throws if a literal gsd-* role reference cannot be resolved. commands -> skill path untouched. No literal Agent( survives in installed Hermes workflows. Folded in (maintainer-directed) a pre-existing cross-cutting branding defect: the "Claude Code" -> brand swap corrupted <runtime_compatibility> comparison tables (where "Claude Code" is a compared-runtime label) for every branding runtime. New shared applyClaudeCodeBrandSwap helper protects <runtime_compatibility> regions via split-and-rejoin (no sentinel token) while still rebranding genuine self-references; adopted by all six branding .md converters. Also a surgical prose-consistency fix so plan-review-convergence.md's dispatch-adjacent terminology is coherent post-projection (no broad bare-word rename). Golden install-parity regenerated for the six branding runtimes (dispatch/branding scope only); other runtimes unchanged. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
612fcb00f7 |
fix(#2232): cap phase-token continuation segments at exactly 2 digits (all sites) (#2254)
* fix(#2232): cap phase-token continuation segments at exactly 2 digits (all sites) A phase whose slug's first word is a ≥2-digit number (dir 14-2026-photos-performance, roadmap phase "2026 Photos & Performance" → slug 2026-photos-…) had its phase token over-collected as "14-2026" instead of "14", so every phase-locating verb (init.plan-phase, init.execute-phase, phase-plan-index, state.planned-phase, roadmap.annotate-dependencies) resolved phase_dir=null / plan_count=0 while the directory existed. This is the residual case #2043 explicitly scoped out: its ≥2-digit continuation gate (\d{2,}) distinguishes single-digit slug words but not multi-digit ones (years, counts). The structural distinguisher: getPhaseDirFromPhaseId writes sub-phase and plan continuation segments zero-padded to EXACTLY 2 digits, so a genuine continuation's digit run is exactly 2 — \d{2}(?!\d). The (?!\d) guard caps the run without anchoring what follows, so each call site keeps its own trailing grammar (letter suffixes, dotted sub-phases, boundaries). Shared-source, not hand-synced: the grammar lives once in phase-id.cts as PHASE_CONTINUATION_SEGMENT_SOURCE / isPhaseContinuationSegment (the #2121 single-owner seam), consumed by all five #2043 sites: - phase-id.cts extractPhaseToken (the reported repro) - validate.cts PHASE_TOKEN_FROM_DIR_RE + canonicalPlanStem - roadmap-parser.cts isDirInMilestone numericRe (hyphenated mode) - core-utils.cts + phase.cts extractCanonicalPlanId (paired plan component only — the LEADING phase component keeps unbounded \d{2,}; phase numbers ≥100 are legitimate) Digit-width policy, resolved per triage and locked by boundary tests at 1/2/3/4-digit continuation widths across all sites: sub-phase/plan numbers ≥100 are out of the dir-token grammar. validate.cts phaseDirNameRe's leading \d{2,} is intentionally untouched — it encodes the write-side padding of the leading dir number, not the continuation heuristic, and has no year collision. Fixes #2232 Claude-Session: https://claude.ai/code/session_017KaYUJnfzV3JVVuQnhkcjg * chore(#2232): add changeset for PR #2254 Claude-Session: https://claude.ai/code/session_017KaYUJnfzV3JVVuQnhkcjg * test(#2232): parity gate + fast-check properties for the continuation cap Addresses trek-e's review on PR #2254 (M1, M2, B1). Test-only — the fix itself was verified as a true root-cause fix, so no source changes. M1 — drift/parity enforcement for the new shared constant. scripts/lint-phase-id-drift.cjs guards PHASE_NUMBER_TOKEN_SOURCE only; its TOKEN_DRIFT_RE cannot match a bare \d{2,} re-derivation, so a future edit reintroducing a raw digit-cap at a consuming site would pass lint + CI silently. Extending the lint was rejected: \d{2,} legitimately appears at the intentionally-unbounded LEADING-token sites (validate phaseDirNameRe, core-utils/phase tokenRe), so a textual guard would need sanctions on correct code and would flag by spelling rather than by behaviour. Instead, per the repo's *-parity.test.cjs precedent, added tests/phase-continuation-parity.test.cjs: a shared digit-width corpus (1/2/3/4/5) asserting every consuming surface's notion of "is this segment absorbed" equals isPhaseContinuationSegment(). Covers all five #2043 sites: extractPhaseToken, PHASE_TOKEN_FROM_DIR_RE, canonicalPlanStem, extractCanonicalPlanId (paired component), and roadmap isDirInMilestone (hyphenated mode, on a real ROADMAP fixture). The corpus states the policy independently of the regex, so it fails on divergence rather than mirroring whatever the code does. Failing-first verified: reverting PHASE_TOKEN_FROM_DIR_RE to \d{2,} fails 3 parity tests; reverting the owner constant itself fails 11 across parity + properties + examples. M2 — fast-check properties for the changed parser (4 added to phase-id.test.cjs, following its existing inline fc precedent): - biconditional: a segment is absorbed IFF its digit run is exactly 2 - the owner agrees with observable extraction for every digit run - metamorphic: a write-side getPhaseDirFromPhaseId dir round-trips to its own normalizePhaseName id — ties the cap to the zero-padding convention it mirrors, so a change to the write-side width fails loudly - metamorphic: the round-trip holds when the phase name leads with a year (the #2232 bug itself, generatively) Digit runs are generated as digit strings (not String(int)) so leading-zero forms like "02" — the whole point of the rule — are actually exercised. B1 — GitGuardian red. The session-trailer hypothesis is disproven: the same Claude-Session trailer rides 3 commits now merged to next via #2173, whose GitGuardian check PASSED. GitGuardian's own comment names tests/phase-id.test.cjs:260 — the synthetic dir literal 'M1-14-2026-photos' tripping the generic high-entropy detector. Composed it from parts; the assertion is unchanged, only the source spelling. Refs #2232 Claude-Session: https://claude.ai/code/session_019SkiJk38YWAbmxHrGxEmuU * test(#2232): name the parity gate after the invariant, not the phase module CI caught two failures from the new parity test, both one root cause: lint-test-file-count caps each production module at 2 test files (primary + one integration, per the #3740 consolidation). The file was named phase-continuation-parity.test.cjs, and the linter clusters a test to a production module by name prefix — "phase-*" bound it to src/phase.cts, whose cluster (phase.test.cjs + phase-dependency-levels.test.cjs) was already at the cap, making 3. That tripped the lint-tests job AND the ubuntu-24 unit lane, where tests/lint-test-file-count.test.cjs is a meta-test asserting the linter exits 0 against the real repo. Renamed to continuation-grammar-parity.test.cjs, matching the convention the repo's other cross-cutting parity gates already follow: they are named after the INVARIANT, not a module — capability-precedence-parity, agent-classification-parity, and runtime-launcher-parity all have no corresponding src/*.cts, so they cluster to nothing. The gate tests a grammar shared ACROSS phase-id/validate/core-utils/roadmap-parser rather than the phase module specifically, so the invariant-name is also the semantically correct home. Not allowlisted: a novel offender belongs under the cap, not ratcheted into the exemption list. Content unchanged — same 12 assertions across the same 5 surfaces. Refs #2232 Claude-Session: https://claude.ai/code/session_019SkiJk38YWAbmxHrGxEmuU --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
4a9833d3e3 |
fix(#2278): use Edit() not Write() for Claude allow-permissions + migrate legacy (#2302)
GSD_CLAUDE_ALLOW_PERMISSIONS pre-populated Claude Code settings.json with Write(.planning/*) and Write(STATE.md). Claude Code has no standalone Write permission gate — file-editing tools are gated collectively via Edit(pattern) — so those rules never matched, fresh installs still hit first-run approval prompts for .planning/* and STATE.md, and Claude Code emitted a session-start warning about the unmatched rules. Swap the two entries to Edit(.planning/*) / Edit(STATE.md). Add a GSD_CLAUDE_LEGACY_ALLOW_PERMISSIONS list of the retired Write(...) forms, consulted by mergeClaudePermissions (actively remove stale entries when adding current ones, idempotent, user entries preserved) and by the uninstall cleanup filter (still removes the legacy form). Sample settings.json in docs/USER-GUIDE.md corrected to match. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f74442310d |
fix(#2257): auto-resume debug on non-terminal session-manager return (#2300)
The /gsd-debug orchestrator handled the gsd-debug-session-manager return with only two literal-string checks (DEBUG SESSION COMPLETE, ABANDONED) and no else branch, so a usable-but-non-terminal progress summary (the manager's own turn/context budget exhausted mid-loop, with a valid on-disk checkpoint) fell through to the user as if the debug were complete. Same gap at the continue subcommand. Callee side (agents/gsd-debug-session-manager.md): add an explicit non-terminal CONTINUE_REQUIRED return marker, distinct from the two terminal shapes and from a genuine user-input checkpoint. Orchestrator (gsd-core/workflows/debug.md Sections 4 and 1c): classify returns exhaustively — recognized terminal markers behave as before, anything else is non-terminal and auto-resumes by re-spawning the session manager from the same slug/checkpoint. Anti-loop guard: after two consecutive no-progress resumes (unchanged next_action/updated), emit a blocker report instead of looping. Regression test (source-text contract guard, fix-2196 idiom) asserts both sections' non-terminal/auto-resume branch, the CONTINUE_REQUIRED marker, and the anti-loop bound. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9c65a2ea02 |
fix(#2256): resolve capability-registry configSchema defaults in config-get (#2299)
cmdConfigGet resolved absent keys through only the 4-key SCHEMA_DEFAULTS map, so the ~42 registry-declared configSchema defaults (including the workflow.security_enforcement security gate, default true) returned 'Key not found' (rc=1) — diverging from the runtime's own resolveConfigKey Level-4 resolver and letting '... || echo false' guards silently read the gate as disabled. Add a resolveSchemaDefault helper that layers SCHEMA_DEFAULTS over the already-imported getCapabilityConfigSchema(cwd) accessor, wired into all three absent-key branches. --default flag precedence, the legacy 4 keys, and 'Key not found' for genuinely unknown keys are preserved. Two pre-existing, security-relevant defects in the same surface, found while writing the regression tests, are fixed inline (no-defer policy): - The --default fallback path never masked secret-named keys, printing e.g. 'config-get brave_search --default <secret>' in plaintext. All six default-emission sites now route through emitResolvedDefault, which applies the same isSecretKey/maskSecret masking the found-key path uses. - Dotted-key traversal used raw bracket access, so 'config-get __proto__' / 'constructor' walked the JS prototype chain and returned internals at rc=0 instead of erroring. Each segment is now own-property-gated. Regression tests folded into tests/config-get-default.test.cjs cover registry defaults (boolean/enum/number, read live from the registry), the no-file/mid-traversal/final-undefined branches, --default and legacy precedence, prototype-pollution keys, secret masking, and the Key-not-found vs No-config-file negative cases. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
315d94f6d4 |
feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode. - gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top. - gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer. - --no-tracer flag wired through plan-phase workflow/command/help/skill. - CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled. - tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1945): backfill changeset PR number to 2294 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a68f1be10e |
ci(#2280): fix release-pipeline workflow defects (finalize timeout + auto-backmerge build:lib) (#2281)
Closes #2280 - release.yml: finalize timeout 10 -> 30 (match rc) - auto-backmerge.yml: npm ci + build:lib before version-sync so the version hook can require the gitignored capability-ledger.cjs |
||
|
|
c4237df8e6 |
docs(#2276): 1.7.0 release documentation — what's-new, EoS explanation, feature index (#2282)
Add a curated 1.7.0 release-highlights page (docs/whats-new-1.7.0.md) and a conceptual Embeddable Orchestration System (EoS) explanation (docs/explanation/embeddable-orchestration-system.md), extend docs/FEATURES.md with a v1.7.0 feature section, and wire both new docs into the docs index (docs/README.md) and the root README. Covers the release's marquee changes: the ADR-1239 Host-Integration Interface / EoS (Embeddable Orchestration System) runtime expansion, the Capability + EoS discoverability registries, the gsd-mcp-server companion, model-catalog advances (GPT-5.6, (1M) badge), statusline enhancements, the compact GSD-state format, plus a themed summary of the 100 fixes and 4 security hardenings. Also corrects a stale CONTEXT.md glossary entry: the Capability Registry Overlay now documents the #2009 fail-open behavior for a load-failed gate-declaring capability (previously described as fail-closed). American house style; no parity-gated reference docs hand-edited. Refs #2276, #1678 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |