81f7ab4df129cddcaacba765aaec4b1566b437da
4632 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
81f7ab4df1 |
refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) (#2392)
* refactor(#2384): leaf dispatch table + runCommand collapse (ADR-2346 P4) Cutover all 55 remaining case arms from runCommand's switch to HOST_COMMAND_ROUTERS. runCommand now contains only its default case (~40 lines): the three-layer dispatch (capability → overlay → host table) plus the unknown-command diagnostic. The 73-case switch is dissolved. Each case body was relocated verbatim to a module-scope route*Command function via a brace-matching extractor; inner break; statements (from _dispatchNonFamily early-exit patterns) were converted to return; (5 arms affected); loop break; statements preserved. Closes #2384 * chore: retrigger CI |
||
|
|
d91e32b3ce |
fix(#2383): untrack node_modules — accidentally committed as a hardcoded absolute-path symlink (#2385)
* fix(#2383): untrack node_modules — accidentally committed as a hardcoded absolute-path symlink |
||
|
|
f15eb5f5c9 |
fix(#2347): make the decision-shape evidence test format-agnostic (#2389)
#1365's fail-loud guard reused the parser's own D- grammar as its evidence test, so a populated <decisions> block using any other ID prefix (e.g. D5-01) was invisible to both parser and guard, collapsing could-not-parse into a clean none-present pass. Add an ID-shaped bold-lead-in probe as format-agnostic evidence on both parse paths; empty/prose scaffolds stay none-present. Graduates the #2371 d5-prefix representative fixture to its expected* assertion. Closes #2347. Admin-merged (self-review bypass) with full green CI. |
||
|
|
062f3fda90 |
chore(#2371): representative gate-fixture corpus + document-shaped property test (#2380)
* chore(#2371): representative gate-fixture corpus + document-shaped property test Adds tests/fixtures/representative/ — a permanent corpus of verbatim, incident-sourced fixtures (never author-invented) from #2286, #2347, #2365, #2366, each labeled with its expected gate verdict in a MANIFEST.json and driven through the real CLI gate entrypoint via tests/representative-corpus.test.cjs. Adds a document-shaped fast-check property test alongside the existing writer-seeded bijection test in tests/api-coverage.test.cjs: the existing generator produces rows and renders them through the writer, so the document shape is a constant and it cannot fail against a decoy table; the new one generates the document space instead. Two gates (#2365, #2347) are still open, so their corpus/property assertions are marked with node:test's official `todo` option — the test executes and reports its failure without affecting the process exit code (https://nodejs.org/api/test.html#test-options). The audit-uat corpus (#2286, fixed by #2317) is a normal passing assertion, proving the methodology works end to end and not just cataloguing gaps. Records the fixture-provenance rule in CONTRIBUTING.md: a gate's fixtures may not be derived from the gate's own writer, grammar, or docstring examples; a negative fixture must come from a source that doesn't know the gate exists. No production src/*.cts changes — validation only, per #2371's scope. Closes #2371 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): address orthogonal-review findings — dedup generators, fix field naming, wire dead fields Standards-axis review findings, all fixed: - Deduplicated the row-shape generators (capabilityGen/rowGen/validRowGen) that were copy-pasted between the parse/render bijection test and the new document-shaped property test in tests/api-coverage.test.cjs — a future edit to one could have silently desynced the two properties. Hoisted to a single module-scope declaration both tests reference. - Renamed decision-coverage-guard/MANIFEST.json's expectedOutcome -> expectedReason. It asserted against the gate's `reason` field, but this codebase already has a real, different `outcome` field at parser altitude (extractDecisions' DecisionOutcome) — naming the manifest field after the wrong altitude's term was exactly the ambiguity the "Fixture provenance" rule this PR adds exists to eliminate. - Removed the unused `role` field from three MANIFEST.json files (never read by any test) and wired the previously-dead per-fixture `expectedMinItems` in audit-uat/MANIFEST.json into a real per-file assertion in tests/representative-corpus.test.cjs, using cmdAuditUat's `results` array — catches a regression that moves items between the two fixture files while preserving the aggregate total, which the existing total_items check alone would miss. No changes to test intent or coverage — same assertions, correctly named and fully wired. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: remove dead capabilityState/capabilityWriter requires from gsd-tools.cjs Surfaced by the mandatory pre-PR lint gate (no-unused-vars) while preparing this PR — unrelated to #2371's own changes, but a defect found while working is fixed in place rather than deferred. Leftover from #2368/#2370 (merged just before this branch rebased onto it): the case 'capability' arm that needed these two requires was relocated to bin/lib/capability-command-router.cjs, which already requires both modules directly (lines 24-25) and is their only real consumer (cmdCapabilityState, resolveCapabilityRuntimeState, cmdCapabilitySet). The two requires left behind in gsd-tools.cjs had zero other references in the file and were never re-exported — confirmed via grep across the file and its module.exports. Behavior-preserving: Node's require cache means the underlying modules still load exactly once via capability-command-router.cjs's own requires; gsd-tools.cjs never used its now-removed local bindings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2371): replace todo-marked assertions with characterization tests gsd-test's own JSONL result parser (gsd-test-runner's internal/pipeline/parse.go, verified directly against that repo's source) has no concept of node:test's `todo` option — it only recognizes kind:"pass"|"fail" and hard-errors on anything else. A { todo: true } test whose body throws is counted as a real failure in gsd-test's own verdict, exactly as if it weren't marked todo — proven by an actual gsd-test run against this branch, which reported outcome:"failed" with all six todo-marked assertions (the property test plus five representative-corpus fixtures) in the failure list, each carrying the correct raw node:test `todo` field the tool's parser simply doesn't read. Replaces todo with characterization: MANIFEST.json now carries both the correct target verdict (expected*) and the exact current observed verdict (currentBuggyOutput, directly verified against live CLI output for all five fixtures). Tests assert currentBuggyOutput — an honest, non-vacuous pin of today's known-broken reality that passes today and will fail loudly the moment the referenced fix changes the observed output, at which point the assertion should be flipped to expected* and currentBuggyOutput deleted. The document-shaped property test switches from throwing fc.assert to non-throwing fc.check (returns RunDetails per fast-check's own docs) and asserts report.failed === true directly, for the same reason. Updates all prose (CONTRIBUTING.md, the fixture READMEs) that previously claimed todo would be respected — that claim was factually wrong for this repo's actual tooling and must not ship. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: regenerate golden-install-parity fixtures after rebase onto next Rebasing onto the current next (which now includes #2381's todo-severity changes to gsd-core/bin/gsd-tools.cjs) produced real conflicts in all 18 golden-install-parity fixtures — expected, since both branches changed the same gsd-tools.cjs hash entry. Resolved by taking one side to unblock the rebase, then regenerating fresh from source via npm run gen:golden and verifying the result; every file's diff is exactly the one hash line for gsd-core/bin/gsd-tools.cjs, correcting a stale intermediate hash from the arbitrary conflict pick. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
58028eaf56 |
fix(#2341): de-dup Cursor / menu by marking skills user-invocable:false (#2386)
Cursor installs both a skills and a commands surface and shows both in '/', duplicating every /gsd-*. Extend the #789 CodeBuddy de-dup to Cursor: convertClaudeCommandToCursorSkill (in both src and the live bin/install.js) now emits user-invocable:false, so the skill stays model-invocable while the commands surface is the single '/' entry point. Closes #2341. Admin-merged (self-review bypass) with full green CI. |
||
|
|
e6c16efa6d |
refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3) (#2382)
* refactor(#2373): cutover resolve/git/config/research host routers (ADR-2346 P3) Relocate 13 case arms from runCommand's switch to HOST_COMMAND_ROUTERS: - resolve: resolve-model, resolve-granularity, resolve-execution (3) - git: git (1) - config: config-ensure-section, config-set, config-set-model-profile, config-get, config-new-project, config-path, migrate-config (7) - research: research-store, research-plan (2) Each case body moved verbatim to a module-scope route*Command function (closures over config/commands/output/_dispatchNonFamily preserved). dispatchHostCommand extended to pass defaultValue + workstreamContext (needed by config-get and config-path). No new files, no logic change. Closes #2373 * chore: retrigger CI |
||
|
|
b0f672f88c |
fix(#2337): capture and surface todo severity (#2381)
add-todo.md gains a confirm-based infer_severity step (infer from the blocker/major/minor/cosmetic taxonomy, confirm via AskUserQuestion with TEXT_MODE fallback, before writing) and a severity frontmatter field. cmdListTodos and cmdInitTodos now surface severity, backward-compatible (key omitted when absent), in parity. Closes #2337. Admin-merged (self-review bypass) with full green CI. |
||
|
|
dcb4954131 |
fix(#2335): normalize volta node image paths to the stable shim (#2375)
Adds a volta branch to normalizeNodePath() that rewrites the version-pinned node image path to volta's stable shim, so managed hooks survive a volta node prune (the fnm/Homebrew/mise class, now covered for volta). Also unifies the release-smoke install timeout into one shared 600s constant across before() and runSmoke() so slow benches no longer spuriously time out. Closes #2335. Admin-merged (self-review bypass) with full green CI. |
||
|
|
b302f53ee6 |
refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) (#2370)
* refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) Behavior-preserving relocation of the 706-line case 'capability': arm from gsd-tools.cjs into a new hand-authored bin/lib/capability-command-router.cjs (sibling of ensure-runtime-build.cjs). The 15 bin/-relative require paths are rewritten to sibling-relative (correct for bin/lib/). dispatchHostCommand is now async (capability's install/upgrade/consent ops await the lifecycle); sync routers (state/phase/…) pass through await unchanged. case 'capability': removed; capability dispatches via HOST_COMMAND_ROUTERS. Validated by the existing capability-lifecycle / -consent / -trust / -loader test suites (no logic changed). Golden install-parity fixtures regenerated. Closes #2368 (Slice 1 — relocation). Probe consolidation (capHostVersion→ readHostVersion, capReadStrict dedup) deferred to a follow-up slice. * fix(#2368): add capabilityState/capabilityWriter requires + INVENTORY row The relocated capability arm references capabilityState (cmdCapabilityState, resolveCapabilityRuntimeState) and capabilityWriter (cmdCapabilitySet) — both module-scope requires in gsd-tools.cjs (L288/289) that the initial closure-dep scan missed. Added as sibling requires to capability-command-router.cjs. Also adds the new cli module to docs/INVENTORY.md + regenerates the manifest. * fix(#2368): correct capHostVersion __dirname depth for bin/lib/ relocation capHostVersion's VERSION/package.json paths were bin/-relative ('..' and '..','..'); on relocation to bin/lib/ they resolved one level too deep, so capHostVersion returned 0.0.0 and capability install failed the engines.gsd gate (#1920). Added one more '..' to each (now resolves gsd-core/VERSION and repo-root package.json correctly). * test(#2368): drop capability from the invocation loop (async/FS vs /fake/cwd) capability is async and does FS/config reads, so invoking it against the unit test's /fake/cwd is fragile. The 6 sync Tier-1 routers stay in the invocation loop; capability is covered by the non-invoking registry- ownership assertion + the dedicated capability-* test suites. * chore: retrigger CI (no-changelog label now present) |
||
|
|
67a9243cf1 |
chore(#2356): make the ADR index a generated artifact and enforce ADR lifecycle invariants (#2367)
* chore: rebuild ADR index as a generated artifact and enforce lifecycle invariants
The ADR index in docs/adr/README.md was hand-maintained with nothing checking
it, and had drifted to 40 of 65 ADRs. The absent rows included the entire
capability family (857/894/959/1016/1143/1213/1244) and ADR-1239 (EoS) itself,
so the decisions a reader most needed were the ones they could not find.
Make the index a derived artifact, matching the repo's existing generated-file
idiom (lint:generated-sync), and enforce the corpus' lifecycle invariants:
- scripts/gen-adr-index.cjs generates the index between markers and validates
the status vocabulary (Accepted/Proposed/Superseded/Legacy/Retired),
successor links, id/filename agreement, and supersession symmetry.
- Wire --check into lint:generated-sync so drift fails CI.
Correct the lifecycle metadata the gate surfaced, without flipping any status:
- ADR-1239 (EoS) declared it subsumed ADR-1016/58/3660/894; none recorded it.
Add reciprocal "Subsumed by" pointers + dated amendments. Subsumption keeps
the target Accepted -- these are live adapters, not dead decisions.
- ADR-857/894 carry dated status caveats: they read Proposed while the
capability system shipped and epic #857 is closed. Ratification is a
maintainer act and is deliberately left open.
- Link ADR-0005/0007/0012/3524 -> ADR-0174 and ADR-0010 -> ADR-0009; record
the reciprocal Supersedes on ADR-0009.
- ADR-218 declared itself "ADR-0175" -- an unfinished rename.
- The 0011 PRD moves from the non-canonical "Draft" to "Legacy".
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: capture stderr via spawnSync; record ADR-0010 draft supersession
Two fixes surfaced by the first gsd-test run and by regenerating the index:
- tests/adr-index-gate.test.cjs used execFileSync, which only surfaces stderr
through the thrown error on non-zero exit. The `--write` path exits 0 while
reporting outstanding violations on stderr, so the helper always saw ''.
spawnSync captures both streams on both outcomes.
- The hand-maintained index recorded 0010-skill-surface-budget-module.md as
"earlier draft superseded by ADR-0011" while the file itself still said
Proposed. Deriving the index from the files would have dropped that
assertion and resurrected a superseded draft as a live decision, so it is
recorded at its source, with the reciprocal Supersedes on ADR-0011.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: drop the dead sdk/ model-catalog candidate retired by ADR-0174
src/model-catalog.cts resolved model-catalog.json through three candidates, the
second being sdk/shared/model-catalog.json three levels up. That was the legacy
source-repo fallback kept by the #3288 fix ("check the co-located path FIRST,
before the legacy source-repo path").
ADR-0174 then retired the @opengsd/gsd-sdk package boundary and deleted the sdk/
tree (
|
||
|
|
cf004df678 |
refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) (#2364)
* refactor(#2360): host dispatch table + state cutover pilot (ADR-2346 P1) Pilot cutover for ADR-2346 Phase 1 (epic #2345). Introduces the Layer-2 host dispatch table — dispatchHostCommand + HOST_COMMAND_ROUTERS, consulted in runCommand's default case after capability/overlay dispatch, before the unknown-command error. Migrates 'state' as the pilot: removes the hardcoded case 'state': arm; state now dispatches default -> dispatchHostCommand -> routeStateCommand, byte-identical to the old path (proven by the new state-command-cutover equivalence test, 5-category template). Host commands are NOT capabilities (core, non-toggleable, no tier/activationKey) — the capability registry stays reserved for toggleable feature bundles per ADR-959. This is the host-vs-capability distinction the merged ADR-2346 lacked; the ADR is corrected here alongside the code that realizes it. - gsd-core/bin/gsd-tools.cjs: HOST_COMMAND_ROUTERS + dispatchHostCommand (prototype-pollution-safe); wired into default case; case 'state': removed; dispatchHostCommand + HOST_COMMAND_ROUTERS exported for tests. - tests/state-command-cutover.test.cjs: UNIT/DISPATCH/BEHAVIOR/REGISTRY equivalence (recording-mock + runGsdTools end-to-end + pollution guard). - docs/adr/2346-*.md: refine Decision 1/2 to the host-table vs capability- registry model (correction that did not land in the merged #2355). Behavior-preserving. Subsequent P1b/c PRs migrate phase/init/roadmap/validate/ verify using this proven template. Closes #2360. * test(#2360): regenerate golden fixtures + allowlist for state cutover Bookkeeping for the gsd-tools.cjs change: npm run gen:golden regenerates the install-parity fixtures (gsd-tools.cjs content hash changed), and the new tests/state-command-cutover.test.cjs is added to the lint-test-file-count allowlist under the 'state' prefix. * refactor(#2360): migrate remaining Tier-1 routers (phase/init/roadmap/validate/verify) Completes P1: all 6 Tier-1 host routers now dispatch via HOST_COMMAND_ROUTERS (state landed in the pilot commit). init preserves its #1688 warnIfStaleBake pre-hook; validate binds the output emitter. Cutover test extended to assert all 6 are consumed + owned. Golden install-parity fixtures regenerated. |
||
|
|
f1a91072b6 |
docs(#2357): fix Registry Discussions category name and state its format (#2361)
The submission process told an admin to create a `Registry` Discussions category and never stated its format. Both were wrong in a way a correct reading of the docs could not catch. Name: the category is `EoS Registry`, and it carries threads for both registries — `discussion` is a required field on Capability entries as well as EoS entries. A reader following the old text would create a second, duplicate category. Format: `discussion` being required means the thread must exist before the entry's PR, opened by the entry's author — an outside contributor holding neither `maintain` nor `admin`. GitHub's Announcement format restricts starting discussions to those two levels, so an admin could pick it, block every external submission at the first step, and see nothing wrong. Records open-ended as the required format, and why Announcement and Question/Answer do not work. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ed06b6a4b9 |
fix(#2329): write opencode slash commands to commands/ (plural), migrate legacy command/ (#2354)
* test(#2329): fail-first tests for opencode commands/ (plural) command dir Red phase, empirically probed: global/local install lands in command/ (singular) with 71 gsd-*.md files and no commands/; the manifest records 71 keys under command/ and zero under commands/; all four declaring sites report 'command'. Migration coverage is black-box (two sequential install runs against one configDir) so it holds regardless of how the fix implements cleanup. The Kilo guard passes today by design — a forward-looking no-collateral check. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): write opencode commands to commands/ (plural), migrate legacy command/ OpenCode discovers slash commands from commands/ (plural); the installer wrote them to command/ (singular), so none of the ~71 /gsd-* commands appeared in the TUI. Five sites declared the directory and all had to agree: - capabilities/opencode/capability.json: both artifactLayout destSubpath entries (global + local) and hostBehaviors.flatCommandDir - bin/install.js: the manifest prefix was a SEPARATE hardcoded 'command/' literal, so the manifest would have diverged from the descriptor even after a rename. It now derives from _hostBehaviors(runtime).flatCommandDir. - src/install-engine.cts installOpencodeFamilyArtifacts: the actual write target, which bypasses resolveRuntimeArtifactLayout via combinedFamilyInstall. This was a fifth site the issue did not list — without it the descriptor change alone would not have moved a single file. Migration: an upgrade over a pre-fix install removes only manifest-proven GSD-managed files from the legacy command/ dir and rmdirs it once empty. Unmanifested user files are preserved, never deleted. Kilo shares the opencode family install path and is explicitly unaffected — pinned by a no-collateral test. Note on the tests: the migration cases originally built their legacy fixture by running the installer and relying on it to produce command/ — i.e. they depended on the bug to set up the fixture, and became unsatisfiable the moment it was fixed (block 1 requires command/ to be absent after a fresh install). They now fabricate the legacy layout explicitly, including rewriting the manifest keys to the command/ prefix — which is load-bearing, since the migration only removes manifest-proven files and an unrewritten fixture would silently no-op and pass even against a broken migration. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): regenerate opencode install golden after rebase onto next The golden conflicted on rebase because #2322 also regenerated it. Resolved by regenerating from the merged source rather than hand-merging a generated file; the only delta is the 71 command/gsd-*.md -> commands/gsd-*.md key renames. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): update stale tests that pinned opencode's singular command/ dir Seven tests encoded the old contract (opencode: command/gsd-help.md exists, the descriptor's flatCommandDir, the install-integration contract, and the resolveRuntimeArtifactLayout golden). They passed in the red phase precisely because they pinned the buggy singular dir; the fix intentionally changes that contract, so these are stale-test corrections, not regressions. Kilo shares the opencode family install path and is deliberately NOT changing — it stays on command/ (singular). The shared opencode/kilo test is now split via an explicit per-runtime dir map so the two cannot be conflated, and Kilo's own layout test is untouched. tests/opencode-command-dir-plural.test.cjs independently pins Kilo unchanged end-to-end. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): changeset for opencode commands/ dir fix Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): backfill PR number 2354 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): correct the changeset — do not assert opencode ignores command/ The changeset repeated the issue's stated mechanism ("OpenCode discovers them from commands/ ... a clean install produced no usable commands in the TUI at all"). OpenCode's source contradicts that: packages/core/src/v1/config/command.ts globs {command,commands}/**/*.md, so BOTH names resolve, and its own skill doc still calls .opencode/command/ typical. Shipping that claim as a release note would document a mechanism that does not exist. The change is still right, for the stronger reason: OpenCode's config docs list plural as the convention and singular as backwards compatibility, so GSD was shipping on the alias the vendor may withdraw. Reworded to describe it as the alignment it is, decided on OpenCode's source and docs rather than on bug reports in either repo. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): baseline opencode's commands/ surface — closes a data-loss path this PR opened Not a bookkeeping gap. Moving opencode's command dir to commands/ moved the install destination to a surface the first-time baseline scan does not cover: 000-first-time-baseline's RUNTIME_SURFACES.opencode lists ['gsd-core','command', 'skills','agents'] — no 'commands'. installOpencodeFamilyCommands unconditionally unlinks every gsd-*.md under its destination before writing the fresh set (install-engine.cts:870-873), with zero manifest or migration involvement. The only thing that protects a pre-existing file is assertInstallerMigrationsUnblocked, which runs before materialization and halts when the baseline scan flags an unknown file at a KNOWN surface. Probed: a pre-existing commands/gsd-plan.md is silently destroyed (install exits 0). The identical file under the legacy, already-baselined command/ surface correctly halts the install with "installer migration blocked pending user choice". So this PR would have traded a protected surface for an unprotected one. Fixed with a NEW fix-forward migration rather than editing 000, per docs/installer-migrations.md:131-134 — an applied migration never re-runs, so editing 000 would only protect fresh installs and leave every existing machine exposed. A new id runs for both populations and drifts no shipped checksum; adding its entry to EXPECTED_CHECKSUMS is the case that test explicitly sanctions. All five pre-existing shipped checksums verified byte-identical. Kilo is excluded by the migration's runtimes filter and keeps command/. This was previously deferred as a PR-body note claiming "low impact — nothing else acts on baseline-scan misses". That claim was never probed and was wrong. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): drop the parenthetical product description from the changeset The product-name purity guard (#1777) rejects "Kilo (which still uses command/)" — fragment prose renders verbatim into CHANGELOG.md, so a product name must not carry a parenthetical. Reworded to a plain sentence; the meaning is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
15b3cc8690 |
docs(#2346): Command Dispatch Completion ADR + graduate ADR-959 to Accepted (#2355)
Records the decision (ADR-2346) to dissolve runCommand's 73-case switch into a two-layer dispatch (registry families + leaf-verb table filling the prepared _dispatchNonFamily seam), collapsing it to ~15 lines. Covers the four decisions ADR-959 leaves open: full dissolution, family/leaf classification rule, shared parseFamilyArgs, and the capability-arm extraction shape. Phased under epic #2345 (P1-P4). Behavior-preserving; each cutover proven by the audit-command-cutover equivalence template. - docs/adr/2346-command-dispatch-completion.md (new) - docs/adr/959-*.md: Status Proposed -> Accepted + amendment section - docs/adr/README.md: index rows for 959 + 2346 - docs/ARCHITECTURE.md: forward-reference note under Command Routing Hub - CONTEXT.md: seed glossary entry Closes #2346 (docs-only; no production code). |
||
|
|
23a65c4a3d |
fix(#2322): materialize installed third-party capability skills (#2340)
* test(#2322): fail-first tests for third-party capability skill materialization Red phase: tests (1) and (6) fail — resolveSurface reports the third-party stem surfaced (#2045) but no SKILL.md is ever written to disk. The other four are controls that must keep holding: first-party-wins collision, profile-tier filter, nested-router layout unperturbed, and absent/malformed capability must not throw. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): materialize installed third-party capability skills A capability could report installed:true, surfaced:true, active:true and still never exist as an invocable command. #2045 fixed the registry layer — resolveSurface unions registry.capabilityClusters into the resolved skill set — but the materialization layer never got the matching fix. stageSkillsForRuntimeAsSkills only ever read gsd-core's own bundled commands/gsd/*.md and silently skipped any stem it couldn't find there, so a third-party skill living at <GSD_HOME>/.gsd/capabilities/<id>/skills/<stem>/ was never copied. Registry said surfaced; disk had nothing. Installed capability skills are now staged alongside the first-party ones, copied verbatim (they are authored complete for their target runtime and need no converter). First-party stems always win a collision, the profile filter still applies, and an absent or malformed capability degrades rather than throwing. Security: capability.json's skills[] entries are validated only as non-empty non-reserved strings (capability-validator.cjs:503-514) — no path shape is enforced upstream — so stems are sanitized (rejecting separators, '..', absolute paths, NUL) with an independent isPathConfined check on both the read and write paths. A '../../evil' stem writes nothing outside the capability's own dir. Also fixes a defect this surfaced in pruneSkillDirs: a materialized capability skill dir has no first-party manifest entry, so every apply logged "preserving (user-owned or unknown)" for a live GSD-managed dir. The retained check now precedes the manifest gate; no deletion outcome changes, and genuinely unknown gsd-* dirs still warn and are preserved. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): address security review — bind skills to declaring capability, fix full profile An independent security review BLOCKED the first pass. Both blockers were mine. BLOCKER 1 (security): readInstalledCapabilitySkill scanned every capability dir and returned the first sorted match, never checking that a capability DECLARES the stem — ownership was inferred from attacker-controlled filesystem layout. Since install copies the whole bundle and the validator only checks DECLARED entries, a capability declaring `skills: []` could ship an undeclared skills/deploy/SKILL.md and win the `deploy` stem on sort order, supplying the agent-invocable instructions the user believed came from the registered capability. Stems are now bound to their owning capId via registry.capabilityClusters, and only that capability's dir is read. BLOCKER 2: the fill-in pass was gated `skills !== '*'` on the premise that applySurface materializes `full` into a concrete Set. True for applySurface — false for the installer, which is the default path: resolveProfile returns the '*' sentinel and bin/install.js passes it straight to staging. So #2322 survived on the default `full` profile, i.e. the fix didn't fix the reported bug. The registry is now plumbed to staging, and '*' stages all capability-cluster stems. Wiring this surfaced a second gap: the ADR-1239 imperative adapter (the primary install path) never threaded its registry either, which would have silently defeated the fix on the real default install. HIGH: staged capability skills were never prunable — pruneSkillDirs gates on the first-party manifest, so uninstalling a capability left its instructions live in the agent's context forever. Staged skills now carry a marker making them GSD-owned and prunable; genuinely unknown gsd-* dirs still warn and are preserved. MEDIUM: the "staged verbatim" claim was false — applySurface rewrites bodies over the whole stage dir. The tests asserted byte-equality and passed only because their fixtures contained no rewrite triggers. Claim dropped; tests now assert the rewrite against triggering content. LOW: isPathConfined is lexical, not realpath (symlink-defeatable, currently unreachable because install rejects symlinks) — comment corrected. The validator does not enforce non-empty, so isSafeCapabilitySkillStem is the sole defense, not a second layer — comment corrected and it now has traversal/NUL/absolute/empty test coverage (previously mutating it to `return true` left every test green). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2322): pin that the imperative adapter forwards a capability registry The delegation-args test deep-equalled the exact argv to installRuntimeArtifacts, so threading the composed capability registry through the ADR-1239 imperative adapter (required for #2322 — without it the default `full` install path never materializes third-party capability skills) failed it. The contract legitimately gained a parameter, so this is a stale-test correction, not a regression. Rather than deep-equalling the whole composed registry (brittle — it embeds the full agent/profile map), the test pins the leading args exactly and asserts only that a registry-shaped value is forwarded. That still fails if the adapter stops threading it, which is the regression the test exists to catch. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2322): backfill PR number 2340 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
19c1f54a2a |
fix(#2316): stop phase complete silently dropping ghost requirement IDs (#2339)
* test(#2316): fail-first tests for ghost REQ-IDs, v-heading over-match, all-orphan gap check
Red phase: 4 of 10 fail against current source (#2316-1 ghost-ID warning,
-3 requirements_updated honesty, -4a v1-heading suppression, -6b all-orphan gap
rows). The other 6 are controls/boundaries that must keep passing — including the
#1159 deferred-heading guard and the literal "TBD" placeholder boundary.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA
* fix(#2316): stop phase complete silently dropping ghost requirement IDs
phase complete parses a phase's `**Requirements**:` line from ROADMAP.md and
reconciles it into REQUIREMENTS.md. When a cited ID was registered nowhere, every
branch degraded to a no-op and the report was indistinguishable from a run that
applied every update: `requirements_updated: true, warnings: [], has_warnings:
false`, file byte-for-byte unchanged.
Four defects on that path, all long-standing (traced to
|
||
|
|
ada79bee97 |
fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md (#2338)
* fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md Step 4 rewrote the `## Current Milestone` heading in the shared root PROJECT.md unconditionally. references/workstream-flag.md marks PROJECT.md `# Shared`, and per-workstream milestone state already lives in the workstream's own STATE.md / ROADMAP.md / REQUIREMENTS.md. With parallel milestones — the sanctioned design — whichever workstream ran new-milestone last silently won the shared heading. Step 4 is now skipped when a workstream is active; step 6 no longer stages PROJECT.md in that mode (cmdCommit returns nothing_to_commit rather than failing when a staged path is unchanged). Also fixes a second defect found while diagnosing this, same root cause (the workflow was workstream-unaware): step 1 parsed only --reset-phase-numbers and the milestone name, so GSD_WS was never set — yet ${GSD_WS} was interpolated at the routing lines. It always expanded to empty, so `/gsd:new-milestone --ws x` suggested `/gsd:discuss-phase [N]` with the workstream scope silently dropped, violating the routing-propagation contract. Step 1 now parses --ws using the established idiom from verify-work.md. Guard is keyed on GSD_WS, not $GSD_WORKSTREAM: the runtime launcher does not export the latter and it is only priority 2 of 5 in resolution, so it would miss the --ws flag case that is the actual repro. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2308): regenerate install goldens for the new-milestone workflow change gsd-core/workflows/ ships as an installed artifact, so new-milestone.md's content hash is pinned in all 18 runtime golden fixtures. Only that hash changed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2308): address review — inert step-6 guard, dropped Evolution repair, tautological tests Independent review found the first pass was partly cosmetic: 1. The step-6 `if [ -n "$GSD_WS" ]` branch was INERT. GSD_WS is assigned in step 1's shell and each step's bash block runs in its own shell — this file already proves it, since step 5 round-trips OUTGOING_MILESTONE through a file for exactly that reason (#2288). The guard read an unset variable, always took the flat branch, and staged PROJECT.md anyway. Rather than re-deriving GSD_WS in step 6, the branch is removed entirely: step 4 Part A's guard is what protects the shared heading, so post-guard the only content PROJECT.md can carry is Part B's idempotent Evolution backfill — which must be staged, not stranded. A regression test now asserts no cross-step GSD_WS branch returns. 2. Skipping ALL of step 4 also dropped the `## Evolution` structural repair — a shared, idempotent backfill that is not workstream state. A pre-Evolution project running only `--ws` would never get the section that transition and complete-milestone expect. Step 4 is now split: Part A (milestone-state write) is workstream-guarded; Part B (Evolution) always runs. 3. The tests were tautological prose-pinning — including one asserting a comment mentions "#2308". The step-6 test asserted the guard's TEXT was present, so it passed on the inert guard it existed to catch. Replaced with executable tests that extract the step-1 and step-6 fences and run them under bash with stubbed gsd_run, asserting real parse and --files behavior. 4. --ws is now stripped from the milestone name (step 1 previously left "--ws search" in the remaining text), and documented in argument-hint, help/modes/full.md, and docs/COMMANDS.md. 5. Changeset no longer overstates: --ws reaches the prose guard and routing hints only, not the SDK calls (state.milestone-switch/phases.clear/init.new-milestone still take no ${GSD_WS} — out of scope here). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * chore(#2308): regenerate SKILL.md, goldens, and size baseline for the argument-hint change skills/gsd-new-milestone/SKILL.md is generated from commands/gsd/new-milestone.md, so documenting --ws in the argument-hint made it stale (caught by lint:ci's gen-plugin-skills --check). Regenerated it plus the install goldens and workflow size baseline, since commands/, skills/, and gsd-core/workflows/ all ship. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2308): backfill PR number 2338 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b2961c3f69 |
fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) (#2336)
* test(#2070): fail-first tests for adaptive model_profile and models tier validation Encodes the three acceptance criteria from #2070 plus the boundary cases the resolver silently ignores today (non-string values, empty string, mistyped phase-type key), and pins VALID_TIERS to a catalog-derived set. Red phase: these fail against current src/ by design. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2070): accept adaptive model_profile in validate health; warn on invalid models tiers (W022) W004 sourced its profile list from a hand-maintained literal that predated the adaptive profile, so `"model_profile": "adaptive"` was false-flagged. It now reads VALID_PROFILES, which model-catalog.cts derives from model-catalog.json. models.<phase_type> was validated nowhere: the resolver's tier gate silently drops unknown values, so a typo like `"planning": "opuss"` was an undiagnosable no-op. A new W022 flags unknown phase-type keys and invalid tier values (including non-string values, which the same gate also drops). VALID_TIERS moves from a function-local literal in model-resolver.cts to a catalog-derived export, so health and the resolver cannot disagree by construction rather than by parity test. Object.values(adaptiveTierMap) is ['opus','sonnet','haiku'] plus 'inherit' — identical to the previous literal, so resolution behavior is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2070): changeset for validate health adaptive profile + W022 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2070): close review findings — malformed models, tier-list duplication, changeset gate Review of the initial fix surfaced three real defects, folded in per the no-defer rule: 1. verify.cts: the W022 guard skipped a top-level `models` that is present but not a plain object (`[]`, `"opus"`, `5`, `true`). The resolver ignores those identically, so they were the same undiagnosable no-op #2070 targets — just one level up. They now warn; absent/null/{} stay silent. 2. config-loader.cts: RUNTIME_OVERRIDE_TIERS was a second hardcoded copy of the tier vocabulary this change had just de-hardcoded elsewhere. It now derives from the catalog via ADAPTIVE_TIER_VALUES (no 'inherit' — runtime overrides resolve to a concrete tier). Byte-equivalent to the old literal. 3. scripts/changeset/lint.cjs: USER_FACING_PREFIXES omitted `src/`. Post-ADR-457 the product source is src/*.cts compiled to a gitignored gsd-core/bin/lib, so the `gsd-core/` prefix is dead coverage for library code and a src/-only PR could merge with no release note — including this one. Adding `src/` closes the gate; tests/ stays non-user-facing. Also corrects a false docstring in the VALID_TIERS test: value-equality cannot detect a re-hardcoded literal, so the test no longer claims it does. Two review findings were rejected with evidence rather than actioned: - W021 double-allocation is governed by ADR-612 ("W021 renumber -> void ... kept, message-disambiguated"), not a defect. - Global-defaults validation would be a false-positive generator: config-loader reads ~/.gsd/defaults.json only on the "no .planning/" branch, and health early-returns E001 without .planning/, so those values provably never affect resolution in any context health can run. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2070): regenerate install goldens for the changeset-lint change scripts/ ships as an installed artifact, so scripts/changeset/lint.cjs's content hash is pinned in all 18 runtime golden fixtures. Adding 'src/' to USER_FACING_PREFIXES changed that hash and tripped every golden parity check. Regenerated via `npm run gen:golden`; the only delta is the lint.cjs hash. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2070): backfill PR number 2336 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a30fb75b51 |
fix(#2068): dynamic routing escalates the model per --attempt (#2334)
* fix(#2068): dynamic routing escalates the model per --attempt, not just effort cmdResolveExecution resolved the model via resolveModelInternal (which ignores dynamic_routing), so retries escalated effort but the model stayed pinned to the default tier. Resolve the model via resolveModelForTier when --attempt is given, gated identically to the effort resolution so model and effort stay symmetric — an omitted --attempt keeps the classic profile model (unchanged for everyone, incl. dynamic_routing users who don't pass --attempt). Escalation is capped at max_escalations. Falls back to resolveModelInternal when dynamic_routing is off. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2068): backfill PR number 2334 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1794acb255 |
chore(#2331): trigger PR-policy workflows on pull_request_target so fork PRs get the verdict (#2333)
* chore(#2331): trigger PR-policy workflows on pull_request_target Three PR-policy workflows (pr-title-validator, pr-target-validator, require-issue-link) triggered on plain `pull_request`, so a fork PR's GITHUB_TOKEN was downgraded to read-only regardless of the declared `permissions:`. Each one comments on the PR and THEN emits its verdict, so the createComment 403 killed the github-script step before core.setFailed ran: the contributor saw an API stack trace instead of the instructions the comment exists to deliver. Confirmed on PR #2084 (job 86573823878), whose title has been non-compliant since 2026-07-08 while the explanatory comment 403'd on every run. Switches all three to pull_request_target (base-repo context, write-capable token), matching the three siblings that already do this correctly (pr-template-format, close-draft-prs, auto-close-unsolicited-prs). Safe: the only checkouts are BASE-branch with persist-credentials: false, and every PR-controlled input is read as data — no head code executes. Also wraps each comment in try/catch so a comment failure can never again suppress the verdict. Extends tests/workflow-maintainer-skip.test.cjs with the trigger lock already applied to close-draft-prs.yml (:32-42) for this same defect class. Closes #2331 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2331): strip backticks before echoing untrusted text into bot comments Found by the orthogonal security review of this change. pr-title-validator and pr-target-validator echo attacker-controlled text (the PR title; the fork's branch name) into an inline-code span in a comment posted by github-actions[bot]. A single backtick closes the span early and the remainder renders as live Markdown — GFM autolinks a bare URL — so a fork author could make our own bot post an arbitrary clickable link into a PR thread, borrowing the bot's credibility for phishing. This interpolation is unchanged from next, but it was NOT previously reachable from forks: the createComment call 403'd and the comment was never posted. The trigger switch in the parent commit is what makes it reachable by untrusted authors for the first time, using the write token it grants — so it is in scope here and fixed here rather than deferred. A PR title has no charset restriction, so that vector is fully exploitable. The branch-name vector is weaker (check-ref-format forbids space, ':', '[' and '*', so no bare URL, link or emphasis is expressible) but is the same class and is stripped identically rather than left to the charset to police. Stripping the backtick is complete: it is the only character that can break out of an inline-code span. Only the rendered body needs this — core.warning/setFailed go to the job log, where @actions/core already escapes workflow commands. Also strengthens the try/catch test to assert core.setFailed sits AFTER the catch block rather than merely existing, so moving the verdict inside the try (the exact inversion #2331 fixes) fails the test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2331): assert verdict ordering on code, not on comment prose The first cut of the verdict-ordering guard failed against correct code. It used indexOf('core.setFailed') on raw source, and these workflows name core.setFailed in their own comments while explaining the bug — at lines 29/118/147, 16 and 6, all BEFORE the catch block. So the assertion compared a comment to the call and reported the inversion it was written to catch. gsd-test caught it: 4 unique failures across linux-node22/24. The code was right; the test was measuring the wrong text. Fixes: - readWorkflowCode() strips whole-line YAML/JS comments so positional assertions see only executable text. - The ordering check is extracted to verdictSurvivesCommentFailure() and exercised against BOTH a good and an inverted sample, so the guard is proven non-vacuous rather than merely passing. - A test pins the trap itself: raw source really does mention core.setFailed before the catch, while the stripped view puts the real call after it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2331): drop the unnecessary permission widening; the trigger was the whole bug Both orthogonal review passes flagged the permissions block, from opposite directions — one said issues:write was dead surface on require-issue-link, the other said it was the load-bearing scope the two validators lacked. Neither is right, and the repo's own history settles it: - pr-title-validator declares pull-requests:write ONLY, and its sticky comment has posted 26 times. - pr-target-validator declares pull-requests:write ONLY — posted 8 times. - require-issue-link declares issues:write ONLY — posted on same-repo PRs #106, #164, #232, #259. So GitHub accepts EITHER scope for issues.createComment when the target is a PR, and all three files already declared a sufficient one. The 403 was purely the fork token downgrade. My added scopes fixed nothing and widened privilege on precisely the workflows now running as pull_request_target — the context where surplus scope matters most. Reverted: permissions are byte-identical to next, and the diff is now trigger + try/catch + sanitizer only. The permission test previously used an (issues|pull-requests) alternation, so it passed on the pre-fix tree and would not have caught removal of the scope that matters. It now asserts each file's SPECIFIC scope and, more usefully, asserts the absence of the other — locking the least-privilege property against a future 'add it to be safe' regression. It is a forward lock, not a #2331 fails-first test; the trigger assertion is the fails-first one. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9ad2bab4be |
fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime (#2332)
* fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime The installer writes resolve_model_ids:"omit" for non-alias runtimes into the machine-wide ~/.gsd/defaults.json (#1156); any runtime read it back, so install order silently flipped Claude's adaptive tier aliases (executor->sonnet, planner->opus) to '' in no-project sessions. Resolution is now scoped to the runtime actually resolving, identified by a new per-install <install>/gsd-core/.gsd-runtime marker (installer writes it beside VERSION). The "omit" branch returns '' only when the PROJECT explicitly set omit (honored for all runtimes, #2517 finding #4) OR the active runtime lacks native aliases. Claude ignores a global-defaults-only omit and keeps its aliases; the active runtime is canonicalized (GSD_RUNTIME -> config.runtime -> marker -> claude) so alias/case spellings can't defeat the check; explicit project omit is workstream/project-scope aware; explicit true still materializes IDs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2297): backfill PR number 2332 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1bb724048a |
fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist (#2325)
* fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist The convergence reviewer-flag whitelist predated the 1.7.0 Antigravity CLI adapter and silently dropped --agy/--antigravity, so convergence fell back to --codex only and the working adapter was unreachable (worse after Gemini CLI's upstream shutdown). Add both flags to the workflow grep whitelist, the command argument-hint + flag docs, and the regenerated SKILL.md; they pass through to /gsd-review unchanged. --gemini behavior is untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2293): backfill PR number 2325 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5ea3401d4f |
fix(#2289): context-monitor emits injection envelope only for supported events (#2324)
* fix(#2289): context-monitor emits injection envelope only for supported events gsd-context-monitor is wired to Codex Stop/SubagentStart/SubagentStop/PreCompact (#772), but it emitted a hookSpecificOutput.additionalContext envelope for every event. Codex's Stop schema rejects that shape ("hook returned invalid stop hook JSON output") exactly when context is low. Use a positive allowlist: emit only for context-injection events (PostToolUse, AfterTool, and the pre-existing Gemini missing-name fallback); exit 0 silently for Stop and every other event. Debounce and critical-session bookkeeping still run on silenced events. Behavioral regression tests drive the real hook. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2289): backfill PR number 2324 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#2285): de-flake per-plan use_worktree property test on duplicate briefs The fast-check property in claude-orchestration.test.cjs located each plan's agent() call via indexOf on the brief, but only plan IDs were unique — briefs could collide (fast-check shrinks toward short strings). On a colliding seed the lookup found the first duplicate's line and misattributed its isolation, failing "use_worktree:true must carry isolation" intermittently (surfaced on macOS CI). emitWorkflowScript is correct for duplicate briefs (verified); suffix the unique id onto each agent() label so the test probe is unambiguous. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
52fab7d9d7 |
fix(#2288): archive phase history under the outgoing milestone version (#2323)
* fix(#2288): archive phase history under the outgoing milestone version phases.clear derived its archive directory from a live getMilestoneInfo() read, but new-milestone.md switches the milestone BEFORE phases.clear runs, so phase history was filed under the NEW milestone's <version>-phases/ dir. Add a --archive-version override (threaded from new-milestone.md, captured before the switch) with precedence override -> live read -> dated label. Harden the version label against path traversal on both phases.clear and the sibling milestone-complete sink (the label is a moved directory name), and persist the outgoing version via a file + quoted shell expansion so untrusted STATE.md content is never re-parsed by the shell. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2288): backfill PR number 2323 into changesets Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b041f101fb |
fix(#2287): surface unresolved deferred-items.md entries in progress + audit-uat (#2318)
The executor SCOPE BOUNDARY convention (agents/gsd-executor.md) logs out-of-scope discoveries to a phase directory's deferred-items.md, but no reader ever consumed it — forensic_audit, cmdAuditUat, and capture --list all skipped it — so deferred items were permanently invisible. cmdAuditUat (src/uat.cts) now scans each phase dir's deferred-items.md via a new parseDeferredItems (reusing the collectSection/splitGapsEntries/ extractGapEntryFields seams) and surfaces entries whose status != resolved (fail-safe: a missing/garbled status surfaces rather than hides, matching the false-negative-averse posture of #2286). forensic_audit (gsd-core/workflows/progress.md) gains Check 7 that globs .planning/phases/*/deferred-items.md and reports unresolved entries. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a22333034b |
fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml (#2312)
* fix(#2310): guard Codex agent model_overrides so Anthropic aliases never leak into .toml generateCodexAgentToml embedded a per-agent `model_overrides` value verbatim as the Codex `.toml` `model`, leaking GSD/Claude tier aliases (opus/sonnet/haiku/fable) and `claude-*` ids. Codex/ChatGPT rejects those (400 "The 'sonnet' model is not supported when using Codex with a ChatGPT account"), and since spawn_agent has no inline model param, the model is baked into the .toml at install time — so the orchestrator could not recover and fell back to the non-equivalent generic-agent workaround. Translate a GSD tier alias through the Codex tier map (sonnet -> gpt-5.6-terra); drop with a deduped warning any Anthropic-flavored value with no Codex mapping (fable) or a `claude-*` id, so emission falls through to the runtime-aware resolver or Codex's default. A final safety gate blocks an Anthropic-flavored model from the runtime- resolver path too (runtime/target mismatch). Mirrors the Claude-side override guard (#2041). Real Codex/OpenAI model ids in model_overrides still pass through verbatim (#2256 preserved); runtime:"codex" tier resolution unchanged (#2517). Adds regression + fast-check property tests in tests/codex-config.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2310): backfill changeset PR number to #2312 * fix(#2310): Codex passive-model posture — omit Anthropic-flavored model (all namespacings) Adopt ADR-1239's passive/session-only posture for Codex model handling: a Codex agent .toml `model` is embedded ONLY for an explicit real-Codex model_overrides pin; any Anthropic-flavored value is omitted so the agent inherits the always- available session model (never a 400). - model_overrides tier alias (opus/sonnet/haiku/fable) or a Claude model id → omit (was: translate to gpt-*); an explicit real-Codex model id → embed verbatim (#2256 preserved). - Detect ALL Anthropic namespacings, not just `claude-*`: single-source the canonical CLAUDE_AGENT_ALIASES from model-resolver.cts and treat any id whose value contains "claude" (case-insensitive) as Anthropic-flavored — catching `anthropic/claude-*` and `us.anthropic.claude-*` (the forms the catalog assigns to opencode/hermes/kilo), which reach a Codex .toml via the runtime-resolver path on a mixed-runtime + Codex install. - The final safety gate applies to the runtime-resolver path too. The full passive posture (removing #2517's runtime-resolver per-tier embedding + a correctness health-check + a Codex TOML sync path) is tracked as the ADR-2310 epic #2313. Regression + fast-check property tests in tests/codex-config.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
636316f720 |
fix(#2286): audit-uat surfaces Gaps section + frontmatter/heading verification items (#2317)
parseUatItems only scanned '### N.' expected/result blocks and parseVerificationItems only recognized table/bullet/numbered shapes, so audit-uat returned a false-clean total_items:0 when a file recorded open findings in a '## Gaps' section, declared items in a frontmatter human_verification: array, or used the '### N. <label>'+bold-paragraph verification shape. parseUatItems now also scans '## Gaps' (via collectSection + iterateBullets) and surfaces any entry whose status != resolved. parseVerificationItems now treats the frontmatter human_verification: array (via extractFrontmatter) as the primary source when present, and adds a tokenizeHeadings fallback for the '### N.'+bold-paragraph shape, preserving the existing table/bullet/numbered recognition (no double-count, no regression). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ff9cb6069f |
fix(#2285): wire claude-orchestration Workflow backend into execute-phase (#2314)
The claude-orchestration capability (#1143) shipped registered 'active' but fully inert: detectWorkflowBackend/emitWorkflowScript had no caller outside their own CLI router, and execute-phase.md declared an execute:wave:pre hook point that the workflow body never rendered — so claude_orchestration.enabled:true had zero effect on real runs. Approach B (maintainer-chosen): - execute-phase.md now renders the execute:wave:pre hook (gsd_run loop render-hooks execute:wave:pre) at a new step 2.75, immediately before each wave's Agent() dispatch — fixing the latent dead-hook gap for any pre-wave capability. - Move the claude-orchestration contribution execute:wave:post -> execute:wave:pre (a pre-wave backend selector belongs before dispatch, not after); rename fragments/execute-wave-post.md -> execute-wave-pre.md with prose instructing the orchestrator to call resolve-wave-dispatch before step 3. Unrelated wave:post contributions (ui.safety-gate, drift, external-job, mempalace) untouched. - New .cts seam resolveWaveDispatch(input) composes detectWorkflowBackend + emitWorkflowScript into one {backend:'inline'|'workflow', ...} result; exposed as gsd-tools claude-orchestration resolve-wave-dispatch. This is a real non-CLI-router, non-test caller of both functions. Fail-closed: any gate miss (disabled, non-Claude runtime, Workflow tool absent, SDK below floor, execution_backend:inline, malformed input) or an emit failure resolves to inline with a byte-identical result shape — no regression to the default-off execute-phase path. Regression tests (tests/fix-2285-*) cover happy-path activation + SDK-floor BVA, the fail-closed gate-miss table with detectWorkflowBackend parity, a fast-check composition property, capability.json contribution assertions, and a source-contract guard that execute:wave:pre is now actually rendered. Dependent registry-shape assertions updated in-scope. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
75bedf16fd |
fix(#2284): project Hermes named dispatch onto delegate_task; protect comparison tables (#2309)
Hermes installs brand-swapped "Claude Code" -> "Hermes Agent" in shipped workflows/*.md but never projected the Agent(...) dispatch calls onto Hermes's delegate_task contract, so installed workflows kept literal Agent(...) syntax and falsely asserted "The Agent tool IS available" (Hermes exposes delegate_task, not Agent). Dispatch projection: a generic named-dispatch engine (projectNamedDispatchToStructuralDelegate) wired into the per-runtime RUNTIME_CONTENT_DISPATCH.hermes.md converter, branching entirely on the documentation-sourced hostIntegration.dispatch facts read via _hostIntegrationDispatch (capability.json unchanged): namedDispatch:false -> resolve the gsd-* role and embed a load-its-prompt instruction in the payload; background:true -> map onto delegate_task background; read-only / maxDepth:1 -> no nested delegation to leaf roles; per-call model dropped. Span detection uses literal Agent( scanning + local balanced paren/quote matching (immune to upstream document quote imbalance) and handles all three corpus call forms (multi-line, object-literal, single-line compact). An independent, mask-free post-projection guard fails the install loud on any residual Agent(/subagent_type/leaked model. Fail-closed: install throws if a literal gsd-* role reference cannot be resolved. commands -> skill path untouched. No literal Agent( survives in installed Hermes workflows. Folded in (maintainer-directed) a pre-existing cross-cutting branding defect: the "Claude Code" -> brand swap corrupted <runtime_compatibility> comparison tables (where "Claude Code" is a compared-runtime label) for every branding runtime. New shared applyClaudeCodeBrandSwap helper protects <runtime_compatibility> regions via split-and-rejoin (no sentinel token) while still rebranding genuine self-references; adopted by all six branding .md converters. Also a surgical prose-consistency fix so plan-review-convergence.md's dispatch-adjacent terminology is coherent post-projection (no broad bare-word rename). Golden install-parity regenerated for the six branding runtimes (dispatch/branding scope only); other runtimes unchanged. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
612fcb00f7 |
fix(#2232): cap phase-token continuation segments at exactly 2 digits (all sites) (#2254)
* fix(#2232): cap phase-token continuation segments at exactly 2 digits (all sites) A phase whose slug's first word is a ≥2-digit number (dir 14-2026-photos-performance, roadmap phase "2026 Photos & Performance" → slug 2026-photos-…) had its phase token over-collected as "14-2026" instead of "14", so every phase-locating verb (init.plan-phase, init.execute-phase, phase-plan-index, state.planned-phase, roadmap.annotate-dependencies) resolved phase_dir=null / plan_count=0 while the directory existed. This is the residual case #2043 explicitly scoped out: its ≥2-digit continuation gate (\d{2,}) distinguishes single-digit slug words but not multi-digit ones (years, counts). The structural distinguisher: getPhaseDirFromPhaseId writes sub-phase and plan continuation segments zero-padded to EXACTLY 2 digits, so a genuine continuation's digit run is exactly 2 — \d{2}(?!\d). The (?!\d) guard caps the run without anchoring what follows, so each call site keeps its own trailing grammar (letter suffixes, dotted sub-phases, boundaries). Shared-source, not hand-synced: the grammar lives once in phase-id.cts as PHASE_CONTINUATION_SEGMENT_SOURCE / isPhaseContinuationSegment (the #2121 single-owner seam), consumed by all five #2043 sites: - phase-id.cts extractPhaseToken (the reported repro) - validate.cts PHASE_TOKEN_FROM_DIR_RE + canonicalPlanStem - roadmap-parser.cts isDirInMilestone numericRe (hyphenated mode) - core-utils.cts + phase.cts extractCanonicalPlanId (paired plan component only — the LEADING phase component keeps unbounded \d{2,}; phase numbers ≥100 are legitimate) Digit-width policy, resolved per triage and locked by boundary tests at 1/2/3/4-digit continuation widths across all sites: sub-phase/plan numbers ≥100 are out of the dir-token grammar. validate.cts phaseDirNameRe's leading \d{2,} is intentionally untouched — it encodes the write-side padding of the leading dir number, not the continuation heuristic, and has no year collision. Fixes #2232 Claude-Session: https://claude.ai/code/session_017KaYUJnfzV3JVVuQnhkcjg * chore(#2232): add changeset for PR #2254 Claude-Session: https://claude.ai/code/session_017KaYUJnfzV3JVVuQnhkcjg * test(#2232): parity gate + fast-check properties for the continuation cap Addresses trek-e's review on PR #2254 (M1, M2, B1). Test-only — the fix itself was verified as a true root-cause fix, so no source changes. M1 — drift/parity enforcement for the new shared constant. scripts/lint-phase-id-drift.cjs guards PHASE_NUMBER_TOKEN_SOURCE only; its TOKEN_DRIFT_RE cannot match a bare \d{2,} re-derivation, so a future edit reintroducing a raw digit-cap at a consuming site would pass lint + CI silently. Extending the lint was rejected: \d{2,} legitimately appears at the intentionally-unbounded LEADING-token sites (validate phaseDirNameRe, core-utils/phase tokenRe), so a textual guard would need sanctions on correct code and would flag by spelling rather than by behaviour. Instead, per the repo's *-parity.test.cjs precedent, added tests/phase-continuation-parity.test.cjs: a shared digit-width corpus (1/2/3/4/5) asserting every consuming surface's notion of "is this segment absorbed" equals isPhaseContinuationSegment(). Covers all five #2043 sites: extractPhaseToken, PHASE_TOKEN_FROM_DIR_RE, canonicalPlanStem, extractCanonicalPlanId (paired component), and roadmap isDirInMilestone (hyphenated mode, on a real ROADMAP fixture). The corpus states the policy independently of the regex, so it fails on divergence rather than mirroring whatever the code does. Failing-first verified: reverting PHASE_TOKEN_FROM_DIR_RE to \d{2,} fails 3 parity tests; reverting the owner constant itself fails 11 across parity + properties + examples. M2 — fast-check properties for the changed parser (4 added to phase-id.test.cjs, following its existing inline fc precedent): - biconditional: a segment is absorbed IFF its digit run is exactly 2 - the owner agrees with observable extraction for every digit run - metamorphic: a write-side getPhaseDirFromPhaseId dir round-trips to its own normalizePhaseName id — ties the cap to the zero-padding convention it mirrors, so a change to the write-side width fails loudly - metamorphic: the round-trip holds when the phase name leads with a year (the #2232 bug itself, generatively) Digit runs are generated as digit strings (not String(int)) so leading-zero forms like "02" — the whole point of the rule — are actually exercised. B1 — GitGuardian red. The session-trailer hypothesis is disproven: the same Claude-Session trailer rides 3 commits now merged to next via #2173, whose GitGuardian check PASSED. GitGuardian's own comment names tests/phase-id.test.cjs:260 — the synthetic dir literal 'M1-14-2026-photos' tripping the generic high-entropy detector. Composed it from parts; the assertion is unchanged, only the source spelling. Refs #2232 Claude-Session: https://claude.ai/code/session_019SkiJk38YWAbmxHrGxEmuU * test(#2232): name the parity gate after the invariant, not the phase module CI caught two failures from the new parity test, both one root cause: lint-test-file-count caps each production module at 2 test files (primary + one integration, per the #3740 consolidation). The file was named phase-continuation-parity.test.cjs, and the linter clusters a test to a production module by name prefix — "phase-*" bound it to src/phase.cts, whose cluster (phase.test.cjs + phase-dependency-levels.test.cjs) was already at the cap, making 3. That tripped the lint-tests job AND the ubuntu-24 unit lane, where tests/lint-test-file-count.test.cjs is a meta-test asserting the linter exits 0 against the real repo. Renamed to continuation-grammar-parity.test.cjs, matching the convention the repo's other cross-cutting parity gates already follow: they are named after the INVARIANT, not a module — capability-precedence-parity, agent-classification-parity, and runtime-launcher-parity all have no corresponding src/*.cts, so they cluster to nothing. The gate tests a grammar shared ACROSS phase-id/validate/core-utils/roadmap-parser rather than the phase module specifically, so the invariant-name is also the semantically correct home. Not allowlisted: a novel offender belongs under the cap, not ratcheted into the exemption list. Content unchanged — same 12 assertions across the same 5 surfaces. Refs #2232 Claude-Session: https://claude.ai/code/session_019SkiJk38YWAbmxHrGxEmuU --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
4a9833d3e3 |
fix(#2278): use Edit() not Write() for Claude allow-permissions + migrate legacy (#2302)
GSD_CLAUDE_ALLOW_PERMISSIONS pre-populated Claude Code settings.json with Write(.planning/*) and Write(STATE.md). Claude Code has no standalone Write permission gate — file-editing tools are gated collectively via Edit(pattern) — so those rules never matched, fresh installs still hit first-run approval prompts for .planning/* and STATE.md, and Claude Code emitted a session-start warning about the unmatched rules. Swap the two entries to Edit(.planning/*) / Edit(STATE.md). Add a GSD_CLAUDE_LEGACY_ALLOW_PERMISSIONS list of the retired Write(...) forms, consulted by mergeClaudePermissions (actively remove stale entries when adding current ones, idempotent, user entries preserved) and by the uninstall cleanup filter (still removes the legacy form). Sample settings.json in docs/USER-GUIDE.md corrected to match. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f74442310d |
fix(#2257): auto-resume debug on non-terminal session-manager return (#2300)
The /gsd-debug orchestrator handled the gsd-debug-session-manager return with only two literal-string checks (DEBUG SESSION COMPLETE, ABANDONED) and no else branch, so a usable-but-non-terminal progress summary (the manager's own turn/context budget exhausted mid-loop, with a valid on-disk checkpoint) fell through to the user as if the debug were complete. Same gap at the continue subcommand. Callee side (agents/gsd-debug-session-manager.md): add an explicit non-terminal CONTINUE_REQUIRED return marker, distinct from the two terminal shapes and from a genuine user-input checkpoint. Orchestrator (gsd-core/workflows/debug.md Sections 4 and 1c): classify returns exhaustively — recognized terminal markers behave as before, anything else is non-terminal and auto-resumes by re-spawning the session manager from the same slug/checkpoint. Anti-loop guard: after two consecutive no-progress resumes (unchanged next_action/updated), emit a blocker report instead of looping. Regression test (source-text contract guard, fix-2196 idiom) asserts both sections' non-terminal/auto-resume branch, the CONTINUE_REQUIRED marker, and the anti-loop bound. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9c65a2ea02 |
fix(#2256): resolve capability-registry configSchema defaults in config-get (#2299)
cmdConfigGet resolved absent keys through only the 4-key SCHEMA_DEFAULTS map, so the ~42 registry-declared configSchema defaults (including the workflow.security_enforcement security gate, default true) returned 'Key not found' (rc=1) — diverging from the runtime's own resolveConfigKey Level-4 resolver and letting '... || echo false' guards silently read the gate as disabled. Add a resolveSchemaDefault helper that layers SCHEMA_DEFAULTS over the already-imported getCapabilityConfigSchema(cwd) accessor, wired into all three absent-key branches. --default flag precedence, the legacy 4 keys, and 'Key not found' for genuinely unknown keys are preserved. Two pre-existing, security-relevant defects in the same surface, found while writing the regression tests, are fixed inline (no-defer policy): - The --default fallback path never masked secret-named keys, printing e.g. 'config-get brave_search --default <secret>' in plaintext. All six default-emission sites now route through emitResolvedDefault, which applies the same isSecretKey/maskSecret masking the found-key path uses. - Dotted-key traversal used raw bracket access, so 'config-get __proto__' / 'constructor' walked the JS prototype chain and returned internals at rc=0 instead of erroring. Each segment is now own-property-gated. Regression tests folded into tests/config-get-default.test.cjs cover registry defaults (boolean/enum/number, read live from the registry), the no-file/mid-traversal/final-undefined branches, --default and legacy precedence, prototype-pollution keys, secret masking, and the Key-not-found vs No-config-file negative cases. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
315d94f6d4 |
feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode. - gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top. - gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer. - --no-tracer flag wired through plan-phase workflow/command/help/skill. - CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled. - tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1945): backfill changeset PR number to 2294 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a68f1be10e |
ci(#2280): fix release-pipeline workflow defects (finalize timeout + auto-backmerge build:lib) (#2281)
Closes #2280 - release.yml: finalize timeout 10 -> 30 (match rc) - auto-backmerge.yml: npm ci + build:lib before version-sync so the version hook can require the gitignored capability-ledger.cjs |
||
|
|
c4237df8e6 |
docs(#2276): 1.7.0 release documentation — what's-new, EoS explanation, feature index (#2282)
Add a curated 1.7.0 release-highlights page (docs/whats-new-1.7.0.md) and a conceptual Embeddable Orchestration System (EoS) explanation (docs/explanation/embeddable-orchestration-system.md), extend docs/FEATURES.md with a v1.7.0 feature section, and wire both new docs into the docs index (docs/README.md) and the root README. Covers the release's marquee changes: the ADR-1239 Host-Integration Interface / EoS (Embeddable Orchestration System) runtime expansion, the Capability + EoS discoverability registries, the gsd-mcp-server companion, model-catalog advances (GPT-5.6, (1M) badge), statusline enhancements, the compact GSD-state format, plus a themed summary of the 100 fixes and 4 security hardenings. Also corrects a stale CONTEXT.md glossary entry: the Capability Registry Overlay now documents the #2009 fail-open behavior for a load-failed gate-declaring capability (previously described as fail-closed). American house style; no parity-gated reference docs hand-edited. Refs #2276, #1678 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
20ff405cb3 |
feat(#2162): opt-in compact GSD-state format for the statusline (#2175)
* feat(#2162): opt-in compact GSD-state format for the statusline New statusline.state_format config, enum full|compact (default full — existing rendering untouched). "compact" renders the state segment as "<version> · P<phase>/<total> · <status>", e.g. "v1.12 · P7/12 · executing" — dropping the milestone name and progress bar (the two biggest width costs) and collapsing narrative statuses to a single keyword. Per the #2162 approval conditions, the keyword set is the canonical vocabulary from normalizeStateStatus() in state-document.cjs (discussing/planning/executing/verifying/completed/paused) — no parallel hand-rolled list, so the vocabularies can't drift — and the canonical stuck state "paused" renders uppercase as PAUSED (no new "blocked" lifecycle state). Statuses the normalizer passes through unrecognized fall back to their first word capped at 16 chars. Lifecycle scenes preserved: active_phase wins over the body phase number, milestone completion renders "complete", idle-with-next-action renders "next <action> <phases>". Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2162): changeset fragment for PR #2175 * fix(#2162): review fixes — ENUM_KEYS coverage, cap boundary tests, changeset format - register statusline.state_format in the fix-1628 coercion-bypass matrix - 15/16/17-char boundary tests for the shortGsdStatus fallback cap - changeset body ends with the (#2162) citation per house convention Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): round-2 review fixes — scene exclusivity, direct config-set coverage - compact renderer gates the milestone-complete scene behind the absence of an in-flight phase id, mirroring formatGsdState's if/else precedence (Scene 1 beats Scene 3); regression test covers the non-atomic active_phase + percent=100 STATE.md shape - direct config-set accept/reject test for statusline.state_format plain strings (ENUM_KEYS matrix covers only the JSON coercion shapes) Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the statusline hook change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): complete-scene gate matches formatGsdState exactly (+property tests) Re-review Major: gating done on !phaseId held completion back for the legacy phaseNum shape — formatGsdState reaches Scene 3 on percent=100 regardless of phaseNum, so compact must too. Gate is now !s.activePhase. The phaseNum-only test now expects 'complete' and cross-checks the full renderer; a parity test feeds identical inputs to both renderers. Re-review Minor: shortGsdStatus gets fast-check property coverage (totality, canonical fixed points, separator safety, fallback shape). Golden fixtures regenerated for the hook byte change. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg |
||
|
|
eec9efc351 |
feat(#2160): collapse verbose '(1M context)' model suffix to compact (1M) badge (#2173)
* feat(#2160): collapse verbose '(1M context)' model suffix to compact (1M) badge Claude Code appends " (1M context)" to the model display name in long-context sessions, eating 12 characters of statusline width. Collapse it to " (1M)" — the signal stays, the width doesn't. Any other display name passes through unchanged. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2160): changeset fragment for PR #2173 * fix(#2160): review fixes — ctx variant, boundary test, changeset format - broaden the suffix match with a context|ctx alternation (approval-condition variant the regex missed) - pin non-context parentheticals ((beta), (deprecated)) as untouched - changeset body ends with the (#2160) citation per house convention Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the statusline hook change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
fc913b37a5 |
refactor(#2268): gen:golden one-command fixture regenerator (#2275)
Phase 3 (convenience form) of golden-parity redesign (epic #2264). Adds npm run gen:golden (regenerates both fixture sets) and points the golden-parity/tree failure messages at it. Full CI-auto-comment deferred (documented in ADR-2264). Closes #2268. |
||
|
|
89b1bef881 |
refactor(#2267): golden-parity file-set snapshot + anti-staleness CI selection (#2274)
Phase 2 of golden-parity redesign (epic #2264). Adds an install file-set snapshot (golden-install-tree) and a ci-test-scope rule selecting golden-parity whenever any installed-source path changes, closing the silent-staleness hole behind the #2266 red. ADR-2264 amended (the copy/transform split premise was unsound). Closes #2267. |
||
|
|
6a474db3aa |
refactor(#2266): single-source golden-parity manifest builder + fixture correction (#2273)
Phase 1 of golden-install-parity redesign (epic #2264). Consolidates buildParityManifest + exclusion constants into tests/helpers/install-shared.cjs (fixes realRoot divergence), adds anti-divergence guard, corrects 12 stale golden fixtures to portable values. Closes #2266. |
||
|
|
ef5a5bc15d |
docs(#2265): ADR-2264 golden-install-parity redesign (#2270)
Phase 0 of golden-install-parity redesign epic #2264. Adds docs/adr/2264-golden-parity-redesign.md + index entry. Closes #2265. |
||
|
|
5f609762ed |
fix(#2237): fail loud on ambiguous bare-number phase directory collision (#2262)
* fix(#2237): fail loud on ambiguous bare-number phase directory collision When two unrelated projects share a .planning/phases/ tree, a bare phase number silently resolved to the first 0N-* directory found — risking cross-project file writes. The fix detects multiple matches for the same phase number and surfaces an ambiguous_matches result instead of silently taking the first. Changes: - src/phase-locator.cts: searchPhaseInDir uses filter() + ambiguity check - src/phase.cts: cmdFindPhase same pattern - src/init.cts: cmdInitPhaseOp surfaces ambiguous_matches in the result - tests/phase-locator.test.cjs: 3 regression tests * docs: backfill changeset PR number (#2262) * merge: keep up to date with next * fix: regenerate stale capability-registry after next merge |
||
|
|
e35534c3ec |
fix(#2236): add hookShell parameter for PowerShell call operator (#2261)
* fix(#2236): add hookShell parameter for PowerShell call operator Windows Claude Code with a PowerShell hook runner failed on every hook with 'Unexpected token' — the installer emitted bare quoted paths that PowerShell parses as string literals, not commands. The root cause was that hookCommandNeedsPowerShellCallOperator returned false unconditionally — the &/no-& decision was keyed on (platform, runtime) only, with no signal for the effective hook-execution shell. One runtime (Claude Code) hosts either Git Bash or PowerShell on Windows. Fix: thread a new hookShell parameter through the projection chain. When hookShell='powershell', the & call operator is prepended. Default (Git Bash, no prefix) is unchanged and regression-locked. * docs: backfill changeset PR number (#2261) * merge: keep up to date with next * fix: regenerate stale capability-registry after next merge |
||
|
|
27415c780e |
fix(#2252): exclude PLAN-REVIEW artifacts from plan count (#2263)
* fix(#2252): exclude PLAN-REVIEW artifacts from plan count The loose /PLAN/i fallback in isRootPlanFile matched *-PLAN-REVIEW.md, inflating plan counts. Added PLAN_REVIEW_RE exclusion before the fallback. * docs: backfill changeset PR number (#2263) * fix: regenerate stale capability-registry after next merge |
||
|
|
d8af61be44 |
fix(#2220): replace invalid mempalace mine --room with detect_room() staging (#2260)
* fix(#2220): replace invalid 'mine --room' with detect_room() staging approach mempalace mine has no --room flag (only search does) — verified against MemPalace 3.5.0 official docs (mempalaceofficial.com/reference/cli.html). The headless capture path used --room, causing every headless/no-MCP run to fail with 'unrecognized arguments: --room' and silently skip capture. Fix: replace the flag with a staging-based approach that uses detect_room()'s documented folder-path match — stage the artifact under a room-named subfolder with a mempalace.yaml room taxonomy, then run 'mempalace mine <stage> --wing'. Docs sources cited in-file: - CLI reference: https://mempalaceofficial.com/reference/cli.html - Mining guide: https://mempalaceofficial.com/guide/mining.html - Config guide: https://mempalaceofficial.com/guide/configuration.html Changes: - skills/gsd-mempalace-capture/SKILL.md: headless staging instructions - commands/gsd/mempalace-capture.md: same - capabilities/mempalace/fragments/capture-problems.md: reference staging - .gitignore: exclude .planning/.mempalace-stage/ - tests/mempalace-capture-headless-invocation.test.cjs: regression test - Golden install parity fixtures + workflow-size baseline regenerated * docs: backfill changeset PR number (#2260) * fix: regenerate golden fixtures after next merge |
||
|
|
8b70db343b |
fix(#2204): phase-completion writes 'All phases complete' per ADR-2207 (#2259)
* fix(#2204): phase-completion writes 'All phases complete' per ADR-2207 completePhaseCore was writing the overloaded bare 'Milestone complete' on the last phase — the same string space the milestone-close verb owns for terminal state. Per ADR-2207, phase-completion now writes the existing intermediate value 'All phases complete' (already used in gsd2-import.cts). Milestone termination ('<version> milestone complete' / 'Awaiting next milestone') remains solely with milestoneCompleteCore. Status lifecycle: Ready to plan → All phases complete → <version> milestone complete → Awaiting next milestone. Changes: - src/state-transition.cts: completePhaseCore status value - src/phase.cts: #2028 guard comment - tests/state-transition.test.cjs: assertion + test name - tests/phase.test.cjs: 8 assertion updates (positive + negative) - tests/state.test.cjs: normalizeStateStatus test case + reset regex - tests/workstream.test.cjs: fixture status to terminal value - gsd-core/workflows/progress.md: Route D label - gsd-core/workflows/transition.md: Route B label - CONTEXT.md: Status lifecycle glossary entry (ADR-2207) - .changeset/brave-geese-jump.md * test(#2204): regenerate golden-install-parity fixtures + workflow-size baseline Workflow file edits (progress.md, transition.md) changed install payload hashes and pushed past the committed workflow-size baseline. Regenerated all 17 golden-install-parity fixtures + claude-local via the standalone gen script (which now also covers the local-scope claude layout). Updated workflow-size-baseline.json and agent-size-baseline.json via size:baseline. * fix(#2204): correct claude-local golden hashes + document gen-script limitation The gen-script's claude-local generation produces macOS-specific hashes incompatible with Linux CI (local-scope install embeds platform-varying node-runner paths). Reverted to manual update using Linux FAILURES.md +actual hashes for the 2 changed workflow files. Added explanatory comment in the gen script. * test(#2204): add isCompletedInventory coverage + clarify CONTEXT.md glossary Addresses orthogonal code-review findings (Medium #1 + #2): - Add isCompletedInventory test cases for ADR-2207 status lifecycle (terminal 'milestone complete' → true; intermediate 'All phases complete' → false; archived → true; active statuses → false) - Clarify CONTEXT.md glossary: note that isCompletedInventory intentionally excludes the intermediate value * docs: backfill changeset PR number (#2259) * docs(#2204): add Status lifecycle table to state-md reference (ADR-2207) |
||
|
|
c1885df9e5 |
chore(#2143): prohibition-with-teeth + migrate remaining ad-hoc table sites — Phase 4 (final) (#2253)
* chore(#2143): prohibition-with-teeth + migrate remaining table sites — Phase 4 Phase 4 of epic #2143 (ADR-2143 §7). Completes the markdown table/mutation consolidation by (a) giving the ad-hoc-parsing prohibition teeth and (b) migrating the last ad-hoc table sites onto the shared seam. - src/markdown-table.cts: new formatting-preserving `updateTableCell` primitive (self-contained, ragged-row-tolerant header/delimiter/cell-range scan; splices only the target cell's raw span, preserving all other bytes incl. padding/CRLF; no-op-preserves-padding when a transformer returns the current value). Exports splitTableRow/isDelimiterRow/findTableStartOffset for tolerant reuse. - eslint-rules/no-adhoc-markdown-parsing.cjs: TABLE-REGEX detector extended to `new RegExp(<literal|static-template>)`; new `.replace()`-mutation detector for roadmap/state/content receivers with a table/section-shaped pattern. - scripts/lint-table-schema-drift.cjs (wired into lint:ci): fails if a TABLE_SCHEMA header drifts from its authored table; tests import its logic (single source). - Migrated onto the seam (behaviour-preserving vs pre-Phase-4 HEAD, verified byte-diff old-vs-new): roadmap.cts cmdRoadmapUpdatePlanProgress, phase.cts cmdPhaseComplete + traceability, milestone.cts cmdRequirementsMarkComplete, uat.cts read path, state.cts metrics/decisions/By-Phase. - Incidental correctness gains from the migration: a decoy table can no longer swallow a phase-progress update (## Progress scoping); a ragged neighbouring row no longer silently aborts an edit; completing integer phase N no longer touches a decimal sub-phase N.x row; record-metric no longer drops trailing section content or duplicates the ## Performance Metrics section. - Kept justified allow-adhoc-markdown markers only where genuinely not a table (security.cts <|role|> token) or a loose non-GFM section (uat human-verify). Two orthogonal isolated reviews (correctness/adversarial + security) passed; correctness found 4 behaviour regressions in the first migration pass, all fixed and re-verified byte-identical-or-better vs OLD. Surfaced for maintainer (pre-existing, ambiguous domain logic, NOT changed here): templates/state.md places a By-Phase table under ## Performance Metrics while cmdStateRecordMetric assumes a Plan|Duration|Tasks|Files table. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): match traceability row by first-cell value, not Requirement header Phase 4's migration matched the REQUIREMENTS.md traceability row by a column literally named `Requirement` (`row['Requirement']`), but real tables head that column `REQ-ID`. The by-name lookup found nothing, so `phase complete` and `requirements mark-complete` left the Status cell `Pending` (regressed #2769 / #2203, caught by gsd-test — 8 failures, both node 22/24). - src/phase.cts, src/milestone.cts: match the row by its FIRST cell's value (the requirement-ID column) regardless of that column's HEADER name, via `Object.values(row)[0]` (updateTableCell builds the record in header order). This mirrors OLD's first-cell `\|\s*<id>\s*\|` anchor, restoring header-name independence while keeping the seam. - src/milestone.cts hasTable: broadened from `Requirement`-only to also recognize `Requirement ID` / `REQ-ID` / `REQ ID` headers, kept in sync with the now-positional rowMatch/hasRow so a REQ-ID-headed table participates in the ADR-2143 §6 write-set and the #2140 table_unmatched drift check (it was silently omitted before — a checkbox-only partial reconcile against a REQ-ID table could report as fully reconciled). The `Requirement`-headed path is byte-identical to OLD. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2143): replace stale structural milestone guards with behavioural suite The `milestone.cjs regex global state fix` block was a source-structure guard (allow-test-rule: structural-regression-guard) — it readFileSync'd the compiled milestone.cjs and asserted removed regex idioms (`tablePattern.test`, `afterTable !== reqContent`, `doneTable = new RegExp(...)`). Phase 4's migration deleted those regexes (table update is now updateTableCell), making the assertions obsolete. Per the Test Cleanup rule, replace them in-PR with a behavioural suite driving the compiled CLI: - multi-ID mark-complete flips all IDs (guards the lastIndex/global-state class), - Pending->Complete flip under both `REQ-ID` and `Requirement` headers (#2769), - idempotent already_complete detection with no corruption, - REQ-ID-headed table participates in write_set (traceability entry, applied), - REQ-ID-headed table trips #2140 table_unmatched drift on a missing row. Pruned the now-nonexistent structural-regression-guard entry from the lint-allow-test-rule-refs allowlist (the source-text-is-the-product entry for the same file remains valid). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): backfill PR number 2253 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): record-metric targets its own metrics table, not By-Phase velocity `state record-metric` appended its per-plan row (`| Phase 1 P1 | 5min | 3 tasks | 4 files |`) into the FIRST table under `## Performance Metrics` — which on a real template-derived STATE.md is the By-Phase velocity table `| Phase | Plans | Total | Avg/Plan |`, polluting it on EVERY plan completion (execute-plan.md:414 is a per-plan call). The command's own metrics table is `| Plan | Duration | Tasks | Files |`, which the template does not ship, so the row never reached it; the scaffold branch also emitted a wrong `| Phase | Plan | Duration | Notes |` header matching neither the row nor the canonical table. Pre-existing (predates Phase 4); surfaced while migrating this site and fixed here per no-defer, on the user's explicit go-ahead. - src/state.cts cmdStateRecordMetric: locate the metrics table by its own header shape (`Plan|Duration|Tasks|Files`, via splitTableRow/isDelimiterRow) rather than "first table in the section". When the section exists but has no metrics table (only the By-Phase table), self-heal by appending a fresh **Per-Plan Metrics:** table to the END of the section body — By-Phase table, Recent Trend and footer preserved verbatim, no duplicate `## Performance Metrics` heading, created stays false. Absent-section scaffold header corrected to the canonical `| Plan | Duration | Tasks | Files |`. Ragged-tolerance + None-yet preserved. - Not touching templates/state.md (golden-install-parity hashed) — record-metric self-creates the table on first use instead. Failing-first regression test (tests/state.test.cjs) demonstrates the By-Phase pollution on the pre-fix build, then green after. Verified: no pollution, self- heal idempotency, both-tables isolation, content/heading preservation, flags, None-yet, corrected scaffold header (23-check adversarial harness + all existing record-metric scenarios). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#2143): deleteSection seam primitive (level-bounded whole-section removal) ADR-2143 §4 shipped withSection/collectSection (replace a section BODY) but no way to DELETE a section (heading + body). Phase 4 suppressed the phase-remove section delete instead of building it. deleteSection(content, predicate, opts) locates the section via the collectSection machinery and splices out from the heading's start offset to the next same-or-higher-level heading — so a level-3 `### Phase N` delete stops at a following level-2 `## Progress`, never past it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): phase remove no longer deletes ## Progress on last-phase removal updateRoadmapAfterPhaseRemoval deleted a `### Phase N` detail section with a greedy raw regex whose lazy scan, on the LAST phase, ran to EOF and destroyed the following `## Progress` heading and its entire tracking table — silent data loss, uncovered by tests (removal tests only exercised a middle phase). Migrated onto the new deleteSection seam (level-bounded, stops at `## Progress`); dropped the allow-adhoc-markdown SECTION-DELETION suppression. Failing-first regression (tests/phase.test.cjs) removes the LAST phase and asserts the ## Progress heading + table survive; middle-phase removal is byte-identical. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#2143): deleteTableRow seam primitive (row removal, ragged-tolerant) Sibling of updateTableCell: locates the first GFM table, matches a DATA row by predicate (ragged-tolerant record build, header order), and splices out that row's whole line preserving every other byte. Returns {ok:false,reason} on no table / no match. Enables migrating the phase-remove Progress-table row delete off its ad-hoc regex (ADR-2143 §7 — the "future row-delete seam" Phase 4 punted). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): phase remove deletes the Progress row via deleteTableRow The Progress-table row delete used a whole-document regex with two defects: (a) `\.?\s` required whitespace after the phase number, so a COMPACT row `|2|Beta|` was never deleted (stale row left behind); (b) unscoped — it could strike a row in a different table (e.g. an earlier `| Phase | Requirements |` table). Migrated onto deleteTableRow, scoped to the `## Progress` section (mirrors deriveProgressFromRoadmap), matching the row by first-cell phase number (integer zero-pad-insensitive; decimal exact; removing `2` never touches `2.5`). Both allow-adhoc-markdown suppressions removed. New behavioural tests: compact unpadded row deleted; padded byte-parity on the surviving rows (their ordinal correctly renumbers via the pre-existing renumber block). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): deleteTableRow leaves no dangling newline on last EOL-less row Deleting the final row of a table with no trailing EOL sliced from the row's start to end-of-string, stranding the newline that terminated the previous line. Back rowStart over the preceding \r?\n in that branch so the table ends cleanly. (Caught by the primitive's own unit test on gsd-test; local scenario checks missed the no-trailing-EOL edge.) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): migrate read-only section-collects onto collectSection Six hand-rolled `## Section` read-extract regexes replaced by the collectSection seam (behaviour-preserving; extracted bodies feed the same downstream parsers): state.cts matchSessionSection (## Session / ## Session Continuity) + ## Blockers, smart-entry.cts ## Blockers, audit.cts ## Current Focus + ## Open Questions. Removes 6 allow-adhoc-markdown "pending #1372" suppressions. Incidental fix: the old Session regex `## Session[ \t]*\n` silently failed on a CRLF `## Session\r\n` heading (Windows STATE.md), nulling all session fields; collectSection is CRLF-safe, so session state now resolves on Windows. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): fence-safe state-transition section writes + dedup stripFrontmatter - milestoneCompleteCore's `## Current Position` and `## Operator Next Steps` section resets used fence-blind raw regexes that a fenced `##` inside the body could truncate/mis-target (#2130/#2067/#2080 class). Migrated onto a fence-aware tokenizeHeadings-based helper (resetSectionVerbatim) that is byte-identical to the old output on the canonical path (9/9 fixtures) and correctly ignores a fenced fake heading (proven robustness gain). - mutateCurrentPositionFirstTime: hand-rolled locate+splice → collectSection + replaceSection (byte-parity). - stripFrontmatter was inlined byte-identically in state.cts AND state-transition.cts; hoisted the single canonical copy into frontmatter.cts (both call sites now import it) + unit tests — eliminates the divergence risk per CLAUDE.md "Generative Fix Divergence". Removes 3 allow-adhoc-markdown / #1372 markers. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): name-address By-Phase sum + uat parse, eslint recall hole, catches - state.cts By-Phase "Total plans completed" sum: positional 2nd-cell regex → name-addressed splitTableRow read (correct on a reordered header, where the old code silently summed the wrong column). Marker removed. - uat.cts parseVerificationItems: loose pipe regex → splitTableRow within the existing table/numbered/bullet union scan (item list byte-identical; does NOT reintroduce the reverted strict-parseMarkdownTable item-drop). Marker removed. - eslint no-adhoc-markdown-parsing: close the `new RegExp(identifier)` recall hole — resolve a const-declared table-shaped regex identifier (mirrors the .replace() detector) + RuleTester cases; param/call args stay out (boundary). - commands.cts: delete a lying comment that claimed the scaffold date "stays on raw UTC / deferred" — #2136 already moved it to realClock.localToday(). - Empty catches (classified, not blind-swept): removed 4 dead try/catch; fixed 3 error-hiding (phase-insert decimal-dir I/O collision now fails loud; phase-remove rename partial-failure surfaced; milestone-archive true count via finally); left best-effort swallows with justification comments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): extractFencedBlock seam + migrate api-coverage named fence parseCoverageMatrix extracted its ```coverage fenced block with an ad-hoc regex (the last real allow-adhoc-markdown suppression). Added extractFencedBlock to the markdown-sectionizer seam (reuses stripFencedCode's CommonMark fence engine — info-string match, ~~~/backtick, nesting, indent) and migrated onto it; byte- parity on the parsed CoverageMatrix across 8 fixtures. Only security.cts:367 (a genuine `<|role|>` protocol-token false-positive, not a GFM table) remains marked in src/ — the "prohibition with teeth" goal (nothing grandfathered but a true FP) is met. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): By-Phase row insert is name-addressed (insertTableRow seam) updatePerformanceMetricsSection's INSERT-new-row branch located the By-Phase table with a canonical-column-order-only regex + a hardcoded positional row literal, so on a reordered header it silently inserted nothing — inconsistent with the now name-addressed UPDATE and SUM halves of the same function. Added insertTableRow (markdown-table seam sibling of updateTableCell/deleteTableRow: name-addressed, header-order-agnostic, EOL-preserving) and migrated the branch onto it, mapping By-Phase values by column NAME. Canonical-order output is byte-identical; a reordered header now inserts a correctly-mapped row; a pre-existing CRLF mixed-EOL splice glitch is incidentally fixed. Retired the now-dead byPhaseTablePattern const. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): phase-list checkbox flip via updateBullet seam Added updateBullet (markdown-sectionizer): a fence-aware, offset-tracked single-bullet write primitive (GFM 1–4-space marker tolerance) — the write counterpart to read-only iterateBullets. Migrated mutateMilestonePhase's phase-list checkbox flip (`- [ ] Phase N …` → `- [x] … (completed <date>)`) off its whole-slice regex onto it, same milestone-slice scope + clock seam. Byte-identical across simple / idempotent / metachar-title / double-space / CRLF scenarios. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): scope the Progress-ordinal renumber to ## Progress via seam phase remove's integer-renumber decremented Progress-table phase ordinals with a whole-document `content.replace(/(\|\s*)(\d+)(\.\s)/g, …)` — unscoped, so it also rewrote any `| N. …` cell in an unrelated/decoy table (same class as the batch-2 row-delete scoping bug). Migrated onto updateTableCell, scoped to the ## Progress section, decrementing each affected row's leading phase ordinal by column name. Byte-identical on canonical Progress tables + multi-row + decimal-sibling cases; a decoy `| 3. … |` row before ## Progress is now correctly left untouched. The sibling heading / checkbox-bullet / PLAN.md-filename / Depends-on-prose renumbers are not GFM-table mutations (outside ADR-2143's table/section mandate) — left as-is. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2143): review fixes — scope traceability write, restore Current Position H3-stop Adversarial review of the remediation (BLOCK verdict) — all 9 findings fixed: - F1 (BLOCKER): requirements mark-complete / phase complete flipped the checkbox but NOT the traceability row on the shipped template, because updateTableCell bound to the FIRST table (## Out of Scope, no Status column) instead of the ## Traceability table — the #2140 silent-divergence class, re-introduced by the seam migration and missed by tests (fixtures had Traceability first). Scoped the write + hasRow probe to the ## Traceability section slice (updateTraceability Cell helper) in milestone.cts + phase.cts. Failing-first tests on the Out-of-Scope-before-Traceability layout; the #2769 first-cell match preserved. - F2 (MAJOR): mutateCurrentPositionFirstTime restored to locateCurrentPosition (STOP_H2_PLUS) — collectSection's default H2-stop swallowed a level-3 subsection and the field regexes clobbered it (#2130 class). - F3/F8: Progress-ordinal renumber re-escapes via escapeCell + keys padding recovery by row index (was de-escaping `\|` and losing padding on dup values). - F4: insertTableRow escapes cell values internally. - F5: updateBullet accepts a tab after the marker (`[ \t]{1,4}`). - F7: resetSectionVerbatim consumes CRLF blank lines (byte-parity on CRLF). - F6/F9: corrected two misleading comments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(changeset): data-loss + CRLF-session user-facing fixes (#2253) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2143): de-flake the G10 windsurf ReDoS-guard wall-clock assertion The G10 test asserted `elapsedMs < 1000` for a 200k-char payload — a wall-clock assertion (CLAUDE.md: never assert on wall-clock time) that flaked on a loaded node24 bench at ~1.1s. It was redundant: runHook's spawnSync `timeout: 10000` already SIGKILLs a catastrophic-backtracking hook, so the exit-0 assertion is the real ReDoS guard. Removed the timing assertion; kept exit-0 + documented the subprocess-timeout mechanism. Surfaced (not caused) by this branch's gsd-test runs loading the bench; unrelated to the markdown-parsing changes but fixed in place per the no-flaky-tests rule. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2cbf186420 |
chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 (#2251)
* chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 Phase 3 of epic #2143 (ADR-2143 §5/§6). The three target bugs (#2140, #2112, #2118) were already fixed tactically on next; this introduces the reusable structural contracts and rewires the primary #2140 site onto them. - src/write-set.cts (new): the parse `Result<T> = {ok,value|reason}` (§5) and the per-surface write-set (`WriteOutcome {surface, applied, requirement?}`, `WriteSet`, `writeSetComplete`) (§6). markdown-table.cts now imports + re-exports `Result` from here (single source; distinct from command-routing-hub's Result). - requirements mark-complete (src/milestone.cts): returns a PER-REQUIREMENT, per-surface write-set; `write_set_complete` is true only if every surface of every requirement applied — structurally forbidding the #2140 OR-into-one-flag masking, including across a multi-ID batch (adversarial-review regression). Pre-existing output fields unchanged (behaviour-preserving; #2140 already fixed). - deriveProgressFromRoadmap (src/phase-lifecycle.cts): removed the vestigial null-swallowing try/catch (findTableWithColumns never throws) — ADR §5 no-swallow; RoadmapProgress return contract unchanged. - commit --files (#2112) and milestone complete --dry-run (#2118) left as-is (single-surface commit / pre-mutation preview — not genuine multi-surface writes). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2244): backfill changeset PR number (#2251) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
efd04716da |
chore(#2143): withSection bounded-mutation seam + phase.cts migration — Phase 2 (#2250)
* chore(#2143): bounded-mutation seam (withSection/withPhaseSection) + phase.cts migration — Phase 2 Phase 2 of epic #2143 (ADR-2143 §4): add a bounded-mutation primitive so a per-phase ROADMAP edit is structurally confined to that phase's own section, and migrate the phase-scoped mutation sites in `phase.cts` onto it. - `src/markdown-sectionizer.cts`: `withSection(content, target, edit, opts?)` — resolves a section via `collectSection` and applies `edit` to ONLY that section's body, re-serialising via `replaceSection`. The edit callback sees only the section body, so any regex it runs is physically confined. - `src/roadmap-parser.cts`: `withPhaseSection(content, phaseId, edit)` — resolves a phase's `### Phase N` detail-section heading via the #2121 phase-id source and delegates to `withSection`. Heading match is anchored to the heading start (a sibling phase whose title mentions the number is not hijacked) and bounds at the next ATX heading of any level (`levelBounded:false`). - `src/phase.cts`: `mutateMilestonePhase`'s plan-count and per-plan-checkbox writes now route through `withPhaseSection` — structurally retiring the #2130 / #2067 / #2080 boundary-crossing class for these sites. The phase-LIST checkbox is intentionally left milestone-slice-scoped (it lives outside any `### Phase N` detail section). Cross-phase renumbering is untouched. - Property test (fast-check): editing phase k leaves every sibling section byte-identical; regression tests for title-collision + mixed heading depth. Behaviour-preserving (verified by old-vs-new differential runs on real fixtures). Extend-never-mutate (ADR-2143 §2). Registration: CONTEXT.md + docs/INVENTORY.md export lists. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2243): backfill changeset PR number (#2250) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |