720301140022dfd1e323d29e2982109b90d12154
192 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7203011400 |
feat(#3072): ship the deferred MCP served catalog (resources + prompts) (#3083)
* test(#3072): add failing-first coverage for the mcp served catalog 55 input-class rows from the phase test matrix, across four suites: the catalog module over injected readFile/readDir seams, the protocol surface through handleMessage, the install-vs-catalog parity gate, and fast-check properties for uri round-trip, traversal refusal, and pagination partition. src/mcp-catalog.cts lands as a skeleton whose functions throw, so the suites fail on BEHAVIOR rather than on a missing module. The REASON enum is real so tests assert typed codes instead of message prose. Hostile coverage for the one client-controlled path surface (resources/read): dot-dot and backslash traversal, percent- and double-encoded traversal, absolute posix and windows paths, file:// scheme, null byte, symlink escape, unindexed sibling, non-string and empty uri, wrong root segment. IO faults are injected by monkeypatching the seam, never chmod 0o000 - root bypasses mode bits, so a permission-based test silently passes with zero coverage in root CI. Refs #3072 * feat(#3072): serve the mcp catalog as resources and prompts gsd-mcp-server now serves GSD's own content alongside its three tools: the workflow, reference and command tree as MCP resources (resources/list, cursor paginated, and resources/read over gsd://<segment>/<relpath> uris) and the 71 commands/gsd/*.md as MCP prompts keyed by bare command name. initialize advertises resources and prompts, and deliberately does not advertise subscribe or listChanged - the catalog is fixed for a server process lifetime, so declaring a notification we never send would be a lie a host acts on. Composition scope is SHARED, not re-declared. shouldCompose lives in src/mcp-catalog.cts and bin/install.js now imports it instead of carrying its own regex, so the served catalog and the installed file floor cannot drift on what gets composed. Proven behavior-preserving across all 2871 tracked paths plus windows-backslash, absolute and near-miss-prefix cases: zero mismatches. tests/mcp-catalog-parity.test.cjs asserts served text equals the installer composition-stage text over the real tree, with anti-vacuity guards requiring both a marker-bearing workflow and a non-composed file in the comparison set. Two measurements corrected the literal issue text. Composition is scoped to gsd-core/workflows/ only, because a reference or command that documents marker syntax with an unfenced example would otherwise be parsed as carrying a real marker and have that line lossily dropped. And parity is asserted at the composition stage rather than against an emitted runtime tree, since install applies per-runtime path rewrites afterwards and the catalog is host-agnostic, so byte equality with any one runtime would be false by construction. resources/read is the one client-controlled path surface and is guarded in two independent layers: the uri must be an exact key in the prebuilt index, which defeats every traversal string by construction, and the mapped path is then re-checked with validatePath so a symlink planted inside a root after indexing is still refused. Also fixes a real drift defect found while here: SERVER_VERSION was hardcoded 1.7.0 while the package is at 1.9.1. It now resolves lazily from VERSION or package.json, reusing the precedent in runtime-artifact-conversion. Closes #3072 * test(#3072): make the catalog parity gate drive the real installer Review found the parity gate vacuous: it never imported or spawned bin/install.js, and recomputed the installer side with the SAME shouldCompose and composeWorkflow the catalog calls internally. It therefore proved only that src/mcp-catalog.cts is self-consistent. The old row 52 compared shouldCompose against a regex literal frozen in the test file rather than against the installer at all. An inline divergent regex re-added to bin/install.js - the exact regression ADR-1671 asks this gate to catch - would have left the suite green. The gate now spawns a real bin/install.js and compares the composition DECISION, observed as gsd:section marker survival, against what the catalog serves for the same files. Marker presence is the right observable because the installer applies per-runtime path rewrites after composing while the catalog applies none, so raw byte equality between the two surfaces is false by construction and must not be asserted. Sensitivity was proven, not assumed: overlaying the shouldCompose export that bin/install.js imports so it always returns false makes a real spawned install leave autonomous.md's markers in place while the catalog still strips them, and the row 48 assertion diverges. Anti-vacuity guards are kept and extended - the comparison set must be non-empty, must contain a workflow that actually carries markers, must contain a file the predicate declines to compose, and the install must have emitted a non-zero file count. The marker-documenting reference case has no instance in the real tree, so it uses an overlay fixture built with the same technique workflow-fragments-emission.install.test.cjs already uses. Renamed to .install.test.cjs so it lands in the install suite it now belongs to. Refs #3072 * test(#3072): retarget the unknown-method assertion off a now-implemented method tests/gsd-mcp-server.test.cjs used 'resources/read' as its example of an UNKNOWN JSON-RPC method. The served catalog implements that method, so it now returns -32602 (no uri supplied) rather than -32601. The remote runner caught it deterministically on both linux lanes: -32602 !== -32601. The test's intent is still correct and worth keeping, so it is corrected rather than deleted or weakened. It now uses 'resources/subscribe', which the server deliberately does not implement and deliberately does not advertise in initialize's capabilities, because it never sends the corresponding notification. That turns the assertion into a real contract - the advertised capability surface and the implemented method surface agree - instead of an arbitrary method name a future feature could invalidate the same way. Swept the rest of the suite for other assertions pinning the newly implemented methods; this was the only one. Refs #3072 * chore(#3072): backfill changeset PR number 3083 * test(#3072): make the catalog fake fs separator-agnostic for windows CI caught this on windows-latest (22 and 24): every catalog fixture indexed ZERO entries, surfaced by the anti-vacuity guards as 'fixture catalog must actually index resources for this property to mean anything'. Mechanism: makeFakeFs keyed its dirMap/fileMap on POSIX-joined paths (${root}/${rel}), while production buildCatalog looks paths up with path.join, which is backslash-separated on Windows. Every lookup missed, tryReadDir returned null, and the catalog came back empty. Production is NOT at fault and is unchanged. The same CI run proves it: on windows-latest the real-filesystem tests all passed, including 'installer composition decision matches the served catalog for every file in the real installed tree' and the row-51 non-vacuity proof against a real spawned installer. A real Windows fs accepts both separators; the FAKE did not, so the fake was the unfaithful one and is what changed. Lookup keys are now normalized unconditionally with .replace(/\\/g,'/') in readDir and readFile - never path.sep-conditional, never platform-gated. The row-42/43 injected-fault wrappers got the same treatment, since they compared raw production paths against POSIX-literal fixtures. No assertion was weakened, and the anti-vacuity guards that caught this are untouched - they are the reason this surfaced as a loud failure instead of a suite that silently asserted nothing on Windows. Refs #3072 --------- Co-authored-by: sim <sim@local> |
||
|
|
8f75e27554 |
fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag Every isolation gate already resolved correctly. The resolved value then reached the executor through a prose instruction telling the model to substitute it into a call the model composes itself, and nothing verified the substitution. When it was dropped, the executor edited and committed in the user's primary checkout with no consent and no warning. A prose backstop would be the same class of artifact as the defect, so this is a shipped PreToolUse hook on the Agent tool. It fires at the instant of the call rather than being read once at the top of a workflow, which is the only placement the model cannot skip. The guard is inert unless it can positively establish that this is a GSD project, that the project resolves to harness isolation, and that the dispatch targets an executor. A non-GSD repo has no invariant to enforce. Where it cannot read the configuration at all, it denies rather than assuming, with its own reason -- a guard that cannot verify must not answer safe. A malformed payload allows rather than throwing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3045): extend the isolation guard to Cursor Cursor is the second of only two runtimes that resolve harness isolation, so shipping the guard for Claude alone left half the exposed surface unguarded while the changeset implied it was covered. The two runtimes fail differently. On Claude the harness flag is a per-dispatch kwarg the model must copy into a call it composes, and the defect is that it can be dropped. On Cursor the flag is --worktree, which applies to the whole session, and the subagent-start payload carries no isolation field at all. There is no flag to check, so the guard verifies the effective state instead: whether the workspace is genuinely running outside the user's primary checkout. That is a stronger check than the Claude one because it tests reality rather than intent, and it is commented so nobody later rewrites it into a flag check. Isolation is established two ways, either sufficient: the workspace resolves to a linked git worktree, or it sits under the worktree root Cursor manages. The second matters because a directory Cursor placed there is a legitimate isolated session even before it becomes a distinct git worktree, where linkage alone would report no repository. Detecting linkage required a new primitive rather than the existing context resolver. That resolver short-circuits on finding a local .planning directory before it ever compares the git directory to the common one -- and an isolation worktree normally has its own checked-out .planning. Reusing it would have read a correctly isolated session as unisolated and denied it, which is the failure direction that gets a guard switched off. The comparison is now its own shortcut-free function that the resolver delegates to after its own shortcut, so existing behavior is unchanged, and the case that would have broken is pinned. The subagent type is checked before any configuration is read, so an unreadable config cannot deny a dispatch this guard would never have enforced against. The input-schema comment on the Cursor hook documented only the fields common to every event and omitted the ones specific to this one. That omission cost a halt during this work; it now documents both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): enforce the resolved dispatch decision, not the host capability The guard keyed on the registry's dispatch.isolation, which says only that a runtime is CAPABLE of harness worktrees. The decision that actually governs a dispatch is the one the workflow resolves after gating, and that legitimately comes out as sequential in three documented cases: a project setting use_worktrees false, a per-plan submodule intersection, and the base-check auto-degrade. The workflow tells the model to omit the flag in exactly those cases, and the guard was denying every one of them. The third case matters most. The preceding fix made the base-check degrade on git timeouts and a missing git binary, where it had previously answered "safe". That correction is right, and it means a transient hang now degrades to sequential far more often than before -- so the two changes composed into a trap where the workflow behaved exactly as designed and the guard blocked it. The workflow already resolves isolation in shell, deterministically, which is what makes it a trustworthy source in a way the model-authored call is not. It now records that resolved value through a dedicated verb, and both guards read it first. A fresh record is authoritative, so sequential dispatches pass untouched. Absent or stale, the guards fall back to the capability check combined with the project's use_worktrees setting, which still covers the case that never reaches the workflow. Also widened the matcher to accept Task alongside Agent, since a host that names the tool Task would otherwise leave the guard silently inert while implying coverage; stopped assuming Claude when no runtime is declared, which is the shipped default and would have demanded a Claude-only argument elsewhere; and made a non-git project inert rather than denied, since advising a worktree session is not actionable without a repository. The original diagnosis never modeled sequential mode as legitimate. That omission is what let this through, and it is now recorded there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): record at resolution and bind the record to its dispatch Two independent reviews converged on the same failure: the guard was fail-open in a default install, so it did not catch the defect it exists to catch. A shipped project carries no runtime key, which made "runtime not confidently known" the common case rather than a corner one. A record asserting that isolation was required but carrying no flag then fell through to a capability lookup that answered "none", and the dispatch was allowed. The flag itself only arrived from a second shell block -- the same block a model dropping the argument would also skip. A test had pinned that behavior as intended. The record is now written by the resolver, as an unavoidable consequence of asking for the value, rather than by a step the model is told in prose to go and run. A guard against a prose-carried value cannot itself depend on prose. Mode, flag and identifiers are written together and atomically, so the flagless window is gone, and a record asserting isolation with no resolvable flag now denies instead of degrading. Runtime is also resolved from the installer's own recorded default, which makes confident resolution the normal case. The per-plan submodule gate degrades after the phase-level decision and never re-recorded, so a plan that legitimately ran sequentially was denied against a still-fresh phase record. It now records its own, scoped to the plan. A record also authorized any dispatch for four hours. One phase degrading to sequential could silently license an unisolated dispatch in the next. Records now carry phase and plan, the guards require them to match, and the window is minutes rather than hours -- the resolver rewrites it before every dispatch, so a long window bought nothing and only widened the hole. The flag validator rejected any value beginning with two dashes, which is exactly the form Cursor and Windsurf declare, so their real value could never have been stored. Writer and reader also derived the record path differently and diverged inside a linked worktree without local planning state. The predictable path remains a way to silence the control without leaving a trace in the diff. It grants no access an agent with shell does not already have, so it is documented as accepted rather than redesigned around. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): correct the staleness boundary and unmask a vacuous parity test The remote runner returned twenty failures. One was a real production defect the boundary case existed to catch: a record whose age exactly equalled the staleness window was treated as fresh, so it stayed authoritative for one tick past its own expiry. Freshness is now strictly inside the window. The parity test meant to stop the two guards' executor lists from drifting could never have failed. Its project fixture was a bare directory rather than a repository, so the non-git inert branch answered before the executor list was ever consulted. It asserted agreement it never actually measured. The fixture is now a real repository, like every sibling in the file. A test also asserted that Windsurf declares the worktree flag. It does not -- Windsurf resolves to no isolation by design, having no named concurrent dispatch to isolate. The test claimed a registry fact that was never true, and a comment in the resolver repeated it. Both corrected, and the test now proves what it should have all along: that the parser accepts any bare flag value, rather than one runtime's supposed value. The new guard was missing from the bundled-hook whitelist, which is the surface that decides what actually ships, and the per-plan gate had gained calls to the launcher without the preamble those calls require. The changeset carried parenthetical product descriptions the purity rule forbids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3045): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3045): make the guard tests hold on Windows Two tests redirect HOME to control where the installer-persisted runtime default is read from. Node resolves the home directory from USERPROFILE on Windows and never consults HOME, so both silently read the real runner profile, found no recorded runtime, and asserted against a project the hook had not recognised. The production code was already correct in asking the platform rather than the variable; only the tests were wrong to assume one variable answers everywhere. The helpers now mirror the override onto both. The symlink spoofing test also created a directory symlink unconditionally, which needs elevated privileges on Windows. It survived on this runner, but it would fail on any host without them, so the creation is now attempted and the test skips explicitly when it cannot be done -- a bare return would have counted as a pass and hidden the gap. Skipping alone would have left the platform uncovered, so the behaviour it proves is now also driven in-process through an injected realpath, following the seam already used for the clock. That case no longer depends on privileges at all, and the end-to-end test keeps its original assertions wherever symlinks work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
da062c0e0d |
chore(#2996): inventory the workflow fragment tree as its own manifest families (#3061)
* feat(#2996): inventory the workflow fragment tree as its own families Epic #1671 Phase 6.5, the epic's last deliverable. 47 step files across 15 workflows and 13 mode files were invisible to docs/INVENTORY-MANIFEST.json. Not through a missed row — through construction: buildManifest walks each family with a flat readdirSync + isFile() and never recurses, so nothing under gsd-core/workflows/<wf>/ could ever appear. modes/ has been invisible that way since #717 without any gate firing, which is the evidence that this is a generator gap rather than someone forgetting a row. Two new families, workflow_steps and workflow_modes, keyed by <workflow>/<subdir>/<file> rather than a bare basename. That is deliberate: two workflows may each own a regression-gate.md, and a step file may share a name with a top-level workflow. The manifest is compared by JSON equality, so a basename collision would silently drop an entry and read as "up to date". Recursion is bounded at exactly one named subdirectory, and a limit+1 test pins that bound so it cannot quietly become a general walk. tests/inventory-manifest-sync.test.cjs carried its OWN duplicate copy of the FAMILIES table — the DEFECT.GENERATIVE-FIX divergence class. Adding a family to the generator alone would have left that test verifying six of eight families while still reporting green. The table now lives once in the generator and is imported, so the two surfaces cannot drift; runMain is guarded behind require.main so importing does not execute the CLI. The per-file roster stays in the generated manifest rather than being copied into INVENTORY.md: 60 hand-maintained rows in lockstep with a generated artifact is precisely the drift this file exists to catch. CONTEXT.md's RULESET.MANIFEST-CANONICAL-KEY and DEFECT.INVENTORY-DRIFT both said "six families" and now say eight, with the two key shapes and the import rule recorded. The non-shipping example index was regenerated for the same edits. Note on scope: this issue also asked for a one-fragment-edit proof. That landed independently as PR #3046 and is not rebuilt here. Refs #2996 * fix(#2996): correct a fabricated roster and an inert coverage pragma Isolated review returned one blocker and three lesser findings. All four were real; all four are fixed. BLOCKER — docs/INVENTORY.md claimed the workflow_modes roster was "discuss-phase, sketch". There is no gsd-core/workflows/sketch/ and never has been; the second member is `help` (4 mode files), exactly as the manifest generated by this same diff already listed. A doc contradicting the manifest it describes, in the PR whose whole purpose is closing doc/reality drift. The adjacent hand-maintained "15 workflows" count is also removed: an unenforced number in a table cell is the same staleness class this file exists to catch, and no test guards table-cell counts. MAJOR — the CLI entry guard carried `/* istanbul ignore next */`, which excludes nothing here. This repo measures coverage with c8 (test:coverage:scripts-floor, 55% floor over scripts/**/*.cjs), and c8/v8-to-istanbul honors only `/* c8 ignore next */`. The pragma looked like it was doing something and was not — the same failure shape as a marker that looks like working gating. MINOR — collectNested called statSync/readdirSync unguarded, so a dangling symlink or an EACCES directory under any workflow's steps/ would throw uncaught and red the manifest gate for the entire repo. An entry that cannot be statted is, for inventory purposes, not a countable file — the same disposition as "not a directory". Row 13c pins the behavior with a real dangling symlink. Refs #2996 * chore(#2996): backfill changeset pr number to 3061 * test(#2996): guard the dangling-symlink row on Windows fs.symlinkSync throws EPERM on Windows without elevation or Developer Mode, so row 13c would red the Windows lane. Guarded with the repo's idiom — a process.platform check plus a genuine t.skip() carrying its reason, never a bare return, which node:test counts as a PASS and would hide the gap. Worth recording why this was not caught here: CI classified this PR's diff as inert (no bin/, gsd-core/, or src/ changes), so the full test matrix was SKIPPED entirely — the 'full test (${{ matrix.os }}, ...)' job shows as skipping with its matrix expression unexpanded. The Windows lane never ran. It would have fired on the next PR that does touch core code, in someone else's change. --------- Co-authored-by: sim <sim@local> |
||
|
|
ed360cd99f |
chore(#2995): extend fragment emission to agents/ and reclaim size-cap headroom (#3058)
* feat(#2995): extend fragment emission to agents/ across every read point Epic #1671 Phase 6.4. `composeWorkflow` stripped `<!-- gsd:section -->` markers only for `gsd-core/workflows/`, so a marked agent shipped its markers verbatim into every runtime — and agent text is loaded into a subagent's context on every dispatch. The issue proposed widening the `copyWithPathReplacement` guard. That is a no-op for agents: agents never traverse that function. Agent content is read for emission at five independent points, and the obvious chokepoint `stageAgentsForProfile` short-circuits on the DEFAULT `full` profile (`skills === '*'` returns the real unstaged directory), so a hook placed there is dead code on most installs. Composition now happens at two call sites instead of five parallel surfaces: `stageAgentsForRuntimeWithConverter` (with `agentsKind` and `kimiAgentsKind` routed through it via an identity converter) and the inline agent loop in bin/install.js. Both compose BEFORE any path rewrite, so a `.claude/` -> `.windsurf/` regex can never reach inside a marker attribute — the ordering #2930 established for workflows. `installCodexConfig` was the fifth read point: Codex embeds each agent's prompt into a per-agent `.toml` via its own readFileSync. Call-graph analysis missed it; the exhaustive per-runtime emission sweep found it. That is why the new guard is behavioral rather than structural — a sixth read point fails the sweep without anyone remembering to extend a list. tests/agent-fragments-emission.install.test.cjs spawns a real installer for every runtime at every agent-bearing scope, derived from RUNTIME_META and the capability registry at run time so a new runtime cannot be silently under-covered. It asserts markers are absent AND the `when="always"` body is retained, so marker-absence cannot be satisfied by dropping content. An identity-composer negative control proves the assertion can fail. Verified: 0 install failures, 0 marker leaks, body retained on 27 runtime/scope paths; red before the wiring on claude(global+local), zcode(global+local), kimi, codex and opencode. Refs #2995 * chore(#2995): give the tightest agents headroom and correct the design lock Epic #1671 Phase 6.4, second half. `agents/gsd-verifier.md` had 12 bytes of headroom under its 49,152-byte LARGE cap and `agents/gsd-debugger.md` had 147 under its 57,344-byte XL cap. Both now extract reference material to `gsd-core/references/` behind an @-reference — the documented DEFECT.AGENT-FILE-SIZE-CAP-BREACH remedy: gsd-verifier 49,140 -> 46,371 B headroom 12 -> 2,781 gsd-debugger 57,197 -> 48,851 B headroom 147 -> 8,493 Byte accounting proves no content was lost: the combined agent+reference delta is exactly the new files' headers plus the agents' slim replacement blocks. Each agent keeps its routing table and a one-line summary per entry, so it degrades gracefully on a runtime that does not inline @-references. `agents/gsd-planner.md` is untouched and still passes both char guards (49,130 < 49,152); it needed no change, so it took none. The other nine LARGE/XL agents carry NO gsd:section markers, and that is deliberate, not deferred. `when=` selection is read from gsd-core/workflows/section-manifest.json, which gen-section-manifest.cjs derives from gsd-core/workflows/*.md only — shape `{workflows: ...}`, no per-agent key, no per-agent init entry point. An agent atom therefore fails admission gate (2) ("a fact the init seam demonstrably computes at a real entry point") and would evaluate false forever while looking like working gating. Marking agents would manufacture exactly the silent-inertness rot the frozen vocabulary exists to prevent. ADR-1671 gains three amendments, two of which close gaps /adr-phase-coverage found against what actually merged: - The 19 -> 29 vocabulary widening shipped in #2994 with no coordinated ADR amendment, which that bullet's own rule forbids. Recorded now. - `flag:--verify-only` was one of six atoms #2992 withheld and deferred to "the LARGE/XL rollout phase". Five shipped; this one is permanently rejected, and that disposition lived only in a merged PR body. - Phase 6.4's own finding: emission extends to agents/, gating does not. CONTEXT.md's glossary was stale on both seams — Workflow Fragments Module still listed the original 4-atom vocabulary and described when= as "not yet acted on", and Section Manifest Module still described InvocationFacts as {waveFlag, phaseNumber, hasPriorPhases}. Both now match the shipped contract. Inventory manifest regenerated AFTER build:lib per the documented ordering landmine; 19 install-tree fixtures pick up the two new references. Refs #2995 * chore(#2995): correct the compose-site count and mark the raw stager Self-review found two comment defects in the prior commit. The agentsKind comment claimed composition lands at TWO call sites; it is three, since installCodexConfig's per-agent .toml writer was added after that comment was written. And stageAgentsForProfile is now production-dead — both callers route through the composing stager — while staying exported and unit-tested, which makes it a trap: it does a raw copyFileSync and short-circuits to the unstaged source directory under the default profile, so a future caller would silently reintroduce the marker-shipping path. Its JSDoc now says so. * test(#2995): guard the marker-documenting-doc class for agents Widening the composer's scope to agents/ makes reachable the exact class #2930 narrowed scope to avoid: a file that DOCUMENTS the marker syntax with an unfenced example is indistinguishable from a real marker, so the composer drops that line from the emitted artifact. Three rows. A fenced example must compose byte-identically. No shipped agent may carry a marker outside a fence — asserted by parsing every real agent and requiring zero explicit sections, which is what makes the fence protection load-bearing rather than decorative. And a non-vacuity row asserts an UNFENCED marker IS parsed as a real marker, so if that ever stops being true the second row is guarding nothing. Also applies two review findings: stageAgentsForProfile's new JSDoc claimed it had no production caller, which is false — bin/install.js's _stageAgents still calls it, and its consumers compose before writing. Corrected to state the invariant instead. And a let/const nit in the emission sweep. * fix(#2995): keep verifier status vocabulary in the agent, fix a wrong fixture The first remote run came back red with three failures. Both root causes were mine. 1. tests/agent-frontmatter.test.cjs requires agents/gsd-verifier.md to literally contain HOLLOW and DISCONNECTED. The Step 4b extraction moved that status vocabulary into gsd-core/references/verifier-wiring-patterns.md, so the agent no longer had it. Byte accounting said no content was lost, and byte-wise that was true — but a contract required those tokens to live IN THE AGENT. That is ADR-1671:66's flexReserve floor stated concretely: a load-bearing fragment must not be trimmed out of its host, and "the bytes still exist somewhere" is not the test. The two status tables are restored to the agent and deliberately mirrored in the reference with a note saying so, so the procedure there still reads standalone. gsd-verifier lands at 47,069 B — headroom 12 -> 2,083, rather than the 2,781 the first attempt claimed. 2. Row 12b of the new marker-documentation guard asserted that an unfenced marker example parses as a real marker, and threw instead: "unmatched /gsd:section close marker". The grammar is WHOLE-LINE only. The fixture had put the OPEN marker inline mid-sentence, so it was correctly not recognised as an open while the close, on its own line, was. That is a real refinement of the hazard this guard exists for: only a marker on its OWN line is mis-parsed — which is exactly how a documentation example is normally written. Row 12b now uses a whole-line marker, and a new row 12c pins the inline case as explicitly NOT a marker. No test was weakened to accommodate the change; the change was corrected to satisfy the tests. Refs #2995 * chore(#2995): backfill changeset pr number to 3058 --------- Co-authored-by: sim <sim@local> |
||
|
|
ef823ca9d9 |
fix(#2830): propagate a halted plan to its transitive dependents (#3038)
* test(#2830): add failing regression tests for halted-plan dependent blocking Add tests/fix-2830-halted-plan-dependents.test.cjs covering direct, transitive (2 and 3 hop), and diamond dependents of a halted plan across both independent "which plans are incomplete" readers (phase-plan-index's cmdPhasePlanIndex and findPhaseInternal/searchPhaseInDir), the negative case (an unrelated decoupled plan stays runnable), and a parity check that the two readers agree. Uses only modules that already exist at this commit (gsd-tools.cjs via subprocess, the pre-existing phase-locator.cjs) so the test file loads and runs cleanly on a fresh clone of this exact commit. These fail against current behavior: neither reader has any concept of a halted plan or a blocked_by/runnable view yet. * fix(#2830): a halted plan no longer leaves its dependents on the runnable work list A plan that reaches a designed stop still writes a SUMMARY, so both "which plans are incomplete" readers saw it as an ordinary completion and reported its dependents as ordinary runnable work — never checking whether an upstream plan had halted rather than finished. - New `status: halted` frontmatter value, documented in all four SUMMARY templates alongside the existing `status: complete`. - New shared src/plan-dependency-graph.cts: a single computeHaltPropagation pass that both phase.cts's cmdPhasePlanIndex (wave-grouping) and phase-locator.cts's searchPhaseInDir (the phase-location primitive, ~50 dependent symbols across 5 command routers) now call, so the two-implementation divergence that caused this bug cannot recur. It accepts an optional precomputedOrder so cmdPhasePlanIndex — which already runs Kahn's algorithm in computeDependencyLevels for wave assignment — passes that order straight through instead of a second traversal; searchPhaseInDir (no prior traversal) lets the module derive its own. The two small duplicated predicates each reader would otherwise carry (is this status "halted"?, which summary file matches which plan id?) are centralized in the same module as isHaltedStatus/buildSummaryFileIndex. - Additive fields only: `halted`/`blocked_by`/`runnable` on cmdPhasePlanIndex's plans[] and top level, `halted_plans`/`blocked_by`/ `runnable_plans` on searchPhaseInDir's result. The pre-existing `incomplete`/`incomplete_plans` fields are unchanged in meaning and membership. - execute-phase.md's discover_and_group_plans step now also skips any plan whose `blocked_by` is non-empty, reporting it by name with its blocking chain, in addition to (not instead of) the existing has_summary skip rule. Extends tests/fix-2830-halted-plan-dependents.test.cjs (introduced in the prior commit) with direct unit coverage of computeHaltPropagation (including the precomputedOrder call shape) and a fast-check property test — both only possible once this commit's new module exists. Closes #2830 * fix(#2830): surface the halt-aware view from init execute-phase The adopted work made phase-locator compute halted_plans / blocked_by / runnable_plans, but cmdInitExecutePhase builds its output by explicitly enumerating fields, so all three were computed and then silently dropped at the exact consumer the issue names as regressed. Forwards them additively -- incomplete_plans and incomplete_count keep their name, type and semantics byte-for-byte -- and adds the same three empty defaults to the roadmap-only fallback so the shape is consistent in both branches. Covered by a new test that drives the real CLI end to end rather than the locator function, since the locator already worked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2830): fail closed on dependency cycles and stop the templates inviting the defect Three review findings, all fixed: - BLOCKER (isolated adversarial). Cycle participants never reach indegree 0 in the Kahn pass, so they were excluded from the topological order, never visited by the forward pass, and vanished from blocked_by entirely -- i.e. reported as runnable. The wave-grouping reader hard-fails on a cycle so it never hit this, but the phase-location reader does not, so init execute-phase offered a plan depending directly on a halted plan. Reproduced, then fixed in the shared engine so every consumer is safe regardless of pre-checks: a node absent from the order is now blocked with a deterministic, non-empty named cause. A plan silently missing from both blocked_by and runnable is the exact disappearance this issue exists to prevent. - MAJOR (isolated adversarial). All four summary templates showed the field as an inline comment on the value line. Frontmatter parsing does not strip trailing comments, so an executor copying the templates' own presentation wrote a halt that parsed as a non-halted string, silently reproducing the original bug. Guidance moved off the value line, and the halt predicate now tolerates an unquoted trailing comment. - HARD standards violation. A test regex-matched child-process stderr prose for /cycle/i, which CONTRIBUTING bans. Replaced with the structured failure signal plus a differential assertion (same fixture without the cycle edge must succeed), so it stays cycle-specific without matching prose. Also folds the duplicated read-summary-and-check-halted wrapper out of both readers into the shared module -- centralizing only the predicate left the exact two-copies-that-drift pattern the module exists to prevent -- and commits the artifact-types documentation for the new status value. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2830): stop the property generator hanging the whole suite The remote runner did not fail -- it hung. Two containers sat in this file for 31+ minutes, and an earlier attempt ran 9 hours before I killed it. The runner passes --test-timeout=0, so nothing ever reaps it: this would have hung CI indefinitely, not reported a failure. Root cause: the DAG generator built edges by rejection -- from: fc.integer({ min: 0, max: n - 1 }) to: fc.integer({ min: 0, max: n - 1 }) .filter(({ from, to }) => from < to) With n === 1 both integers are forced to 0, so the predicate is unsatisfiable and fast-check retries value generation forever. n is drawn from 1..12 and fast-check biases toward boundary values, so n === 1 is reached almost at once. This also explains why the failing-first run completed normally while the fixed run hung: before the fix the graph module did not exist, so the property test threw on import and never reached generation. It only starts hanging once the code under test works. Generates the DAG by construction instead -- `to` is drawn strictly above `from`, with the degenerate single-node case short-circuited to an empty edge list -- so no rejection sampling is involved. Switches the import to the shared fast-check setup so the seed and run count are pinned per CONTRIBUTING, and adds a bounded regression guard that samples the arbitrary directly, so a future reintroduction fails loudly instead of hanging. Verified: the file now completes in 2 seconds, 29 tests started and 29 finished, zero failures, against an indefinite hang before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2830): restore the depends_on display contract and acknowledge the workflow growth Full-suite run surfaced two things the focused harnesses could not. 1. Regression of a pinned pre-existing contract (#3785). A refactor routed the EMITTED depends_on field through the new dependency resolver, which also consults the canonical-prefix map. The original consulted the plan map only, so a short canonical prefix passed through verbatim -- '24-01' stayed '24-01' rather than becoming '24-01-auth-hardening'. The emitted field is a DISPLAY mapping, not the DAG resolution, and #3785 pins that. Reverted with a comment recording why it must not use the resolver; full resolution is still used for the wave DAG and halt propagation, which is what needs it. 2. The workflow file grew 518 bytes without an acknowledgment, from the halt-aware skip rule and the widened parse contract. Acknowledged. Note on where the acknowledgment landed: the guidance is to add a NEW fragment, but execute-phase.md is already named by an existing fragment and the linter hard-fails when two ack sources name the same path. Appending to the owning fragment, following its own established multi-PR pattern, was the only lint-clean option. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2830): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ff4a57b78c |
chore(#1671): migrate the remaining 13 LARGE/XL workflows to the fragment model — Phase 6.3 (#3030)
* chore(#2994): fragmentize progress.md forensic audit onto the fragment model Extract the --forensic-gated forensic_audit step to workflows/progress/steps/forensic-audit.md behind a section marker, and repair progress.md's init line to forward --forensic so the atom is actually true in production rather than only under direct CLI tests. progress.md shrinks 32630 -> 27207 bytes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize the four manifest-wired workflows new-project, quick, new-milestone and progress each already had a dedicated cmdInit* entry point but zero marked sections. Extract nine gated bodies to workflows/<wf>/steps/ behind section markers and repair each init line to forward its flags. Fold --full into the discuss/research/validate facts inside cmdInitQuick so the when= grammar never sees an OR, per the chunked-mode precedent. Fixes found while working, per the no-defer rule: - cmdInitProgress passed no phase info to buildSectionManifestField, so state:phase-mvp-mode was permanently false — an atom in the vocabulary whose fact could never be computed. - the quick init router folded flag tokens into the free-text description, which the new forwarding would have corrupted. - a #2508 dispatch note was nested inside quick.md's Agent(prompt=) fence, leaking orchestrator guidance into the subagent prompt. - progress.md had a 3-vs-4 backtick outer-fence imbalance. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize verify-work.md and admit state:ui-phase-active Wire cmdInitVerifyWork to buildSectionManifestField — it was a dedicated entry point that never emitted a manifest — and mark two sections. state:ui-phase-active folds (plan:pre hooks include an active ui step) OR (the phase dir holds a *-UI-SPEC.md) into one boolean in init.cts, so the grammar still sees a single operator-free atom. The inner Playwright-MCP check stays as prose inside the fragment: it is live session state and no init seam can precompute it. The MVP false-branch note is a real fallback, not redundant prose, so it sits outside the marker — gating it away would delete the text needed precisely when MVP mode is off. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): follow moved workflow content in drift guards Retarget every guard that asserted on content this branch moved into workflows/<wf>/steps/, mirroring 815b3d897. Each retargeted assertion was verified to still fail when its step file is blanked, so none was weakened into vacuity. Three assertions in verify-mvp-uat were genuinely red. Three more were worse than red — passing for the wrong reason: - quick-commit-boundary and worktree-cleanup anchored on indexOf('Step 5.6'), which matched a later cross-reference and sliced 16069 chars that coincidentally held the asserted substrings. Replaced with an expandWorkflowSections helper that splices step content back in place. - phase6-review-capabilities lost its end boundary and widened to EOF. - playwright-ui-verify matched 'UI' in an unrelated bullet and 'fall back' in a subagent-dispatch line after the real content moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize code-review and complete-milestone, admit three atoms Add dedicated cmdInitCodeReview and cmdInitCompleteMilestone entry points alongside the shared generic ones rather than modifying them — init.phase-op and init.manager carry a CRITICAL blast radius (179 dependents, 24 processes) and stay byte-identical for their other callers. Admit flag:--fix, state:fallow-enabled and state:git-create-tag, each with a consuming section and a fact its own entry point computes. Both sections had the resolver-in-body hazard: the fallow config-gate and the git.create_tag check each sat inside the very block being gated, so gating would have disabled the resolver that decides the gate. Both are hoisted into init and the bodies now consume the resolved fact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): retarget code-review and milestone drift guards, fix two red tests Retarget guards that asserted on content moved into steps/, proving non-vacuity by blanking each step file and confirming failure. Also fixes two genuinely red tests found while working, per the no-defer rule: - workflow-fragments' frozen-vocabulary lock was missing state:ui-phase-active, so commit 7ef7f8336 shipped red. Lint and build both passed over it, which is why neither is sufficient verification. - code-review's quick.md capability-hook assertion carried a stale delimiter after the 18ff35d20 extraction. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize autonomous.md and admit state:plan-strategy-converge Five sections share one atom, the pattern plan-phase already uses for flag:--research-phase. The atom folds --converge OR --cross-ai into a single boolean in cmdInitAutonomous so the grammar stays operator-free. cmdInitAutonomous is additive; init.milestone-op, init.manager and init.phase-op are untouched and still consumed. The $PLAN_STRATEGY bash resolver is deliberately retained — ungated local-planning bullets still read it, so the init-side fact supplements it rather than replacing it. converge-fail-fast required splitting one bash fence so the always-run CONVERGENCE_ARGS construction stays outside the marker. All three flag-absent fallbacks were left outside their markers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize review and discuss-phase-assumptions Admit state:reviewer-instances-configured (two peripheral notes share it; the core reviewer-lane dispatch stays unmarked — it is the workflow's primary always-evaluated logic, not an optional branch) and state:auto-advance-active, which folds --auto OR two config keys into one boolean so the grammar stays operator-free. discuss-phase-assumptions was the highest-risk edit in this PR. Its auto_advance step is a full if/elif/else; gating it whole would have deleted the flag-absent fallback needed exactly when --auto is off. Split verified exact: resolvers 636-651 and the 'End here' fallback 668-669 both stay outside the marker; only 653-667 is gated. Adds emitted-drift acks for the two files that grew — review.md (+55 B) and autonomous.md (+737 B from 80799211c, which had none and would have red-gated the push. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize docs-update, update, transition and new-milestone Part A Completes the 13-workflow rollout. Three of these had no init call at all and gained a dedicated entry point plus their first gsd_run query line. Admits state:is-monorepo and adds state:next-channel, state:workstream-active and state:flat-mode. Vocabulary 26 -> 30 atoms. Part A of new-milestone applies when NO workstream is active — the negation of state:workstream-active. Rather than teach the grammar negation, which is the Greenspun drift the frozen list exists to prevent, it gets a separate positively-phrased atom whose fact is the inverse. Part B, which always runs, stays outside the marker. flag:--verify-only is deliberately NOT admitted: docs-update has no contiguous purely-additive region for it, and an atom without a consuming section is dead vocabulary. Evidence recorded in the slice report. update.md reuses its existing resolved $GSD_TOOLS rather than prepending the canonical preamble, which would have clobbered it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): stop automated-ui-verification re-resolving its own gate, retire dead vocabulary Two defects the new tests caught. The automated-ui-verification step re-ran gsd_run loop render-hooks and recomputed UI_PHASE_ACTIVE inside a body that is only read when that fact is already true — the circular self-disabling pattern this design forbids, introduced by 3c654b168. cmdInitVerifyWork now exposes ui_phase_active and the step consumes it. Its launcher preamble goes too: no gsd_run remains. The Playwright-MCP check stays as prose — that is live session state. Dead vocabulary predating this PR: flag:--full and state:needs-codebase-map were admitted with a gate-1 claim that never materialized. flag:--full is removed, redundant once quick folds it into discuss/research/validate. state:needs-codebase-map gets the real consumer it always lacked, gating new-project's codebase-map offer. Vocabulary 30 -> 29, and no atom is now without a consuming section. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): add the atom-admission, inversion and resolver-hoist gates The two existing parity guards prove vocabulary/predicate symmetry but never that a fact is computed — an atom no cmdInit* assembles evaluates false forever. These close that hole: - per-atom satisfiability for all 29 atoms, plus an anti-vacuity assertion so the loop cannot silently cover zero atoms - dead-vocabulary check against the shipped manifest - inversion guard: the flag-absent fallbacks in discuss-phase-assumptions and verify-work must stay outside their markers - data-driven resolver-hoist guard over the shipped manifest, so a future extraction cannot reintroduce the circular class - compound-fold coverage (--full, --cross-ai, --rc, config-only --auto) - null-vs-[] degraded/computed distinction, and flag value shapes Also repairs the frozen-vocabulary lock, which was stale and red for the seven atoms earlier commits on this branch shipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2994): add changeset for the fragment-model rollout Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): cite the issue on the two new allow-test-rule exemptions ADR-456 requires an issue ref on the same line as the annotation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2994): correct the atom-count claims after retiring flag:--full The vocabulary doc comments still said 30 entries; it is 29 since flag:--full was removed as dead vocabulary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): dedupe the phase-fallback block and harden --ws parsing Review findings. MAJOR: the three new init entry points each pasted a verbatim copy of the guardedFindPhase/guardedGetRoadmapPhase fallback, taking the repo from four copies to seven — DEFECT.GENERATIVE-FIX. Extracted applyRoadmapFallback and folded six of the seven; each call site keeps its own field-set via a closure. Duplication removed rather than papered over with a parity test. cmdInitPhaseOp stays out: its fallback omits has_reviews, so it is not a byte-identical copy, and it is CRITICAL-radius. LOW, pre-existing: GSD_WS captured [^[:space:]]+ and expands unquoted, so a workstream name holding glob metacharacters would expand against the filesystem. Narrowed to [A-Za-z0-9._-]+. The unquoted expansion is kept — it must word-split into two args and vanish when empty. Also restores the vocabulary ordering convention, and fixes a masked test bug the mandated run surfaced: the flag-forwarding guard checked only the first init line per workflow, but new-milestone has two, so a real failure was reporting exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): drop the stale new-milestone emitted-drift ack new-milestone.md was acked for a +406 B growth measured against an intermediate commit. Net against origin/next it SHRANK by 8 bytes, so nothing needed the ack and it explained nothing — which the differential attribution check reports as a stale acknowledgment, not a pass. update.md's entry stays: it genuinely grew +703 B. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): resolve the 15 failures from the full matrix run All 15 were real and identical on both lanes. REAL REGRESSION: autonomous.md hit 41479 chars against the #2196 guard's 40960 cap — a CHARS cap distinct from the LARGE tier byte cap, which the five section stubs pushed it over. Extracted the 3a.5 UI Design Contract body to references/; now 39968 chars, and the file nets -795 B vs base, so its growth ack is deleted rather than left stale. REAL DEFECT: docs referenced /gsd-transition, which is not a live registered command. Reworded. STALE FIXTURE: the emission byte-identity test hardcoded two marked workflows; this branch legitimately marks fifteen. Fixture corrected — the source was right. The rest were drift guards over the eight workflows the earlier sweep did not cover, retargeted at where the content now lives with non-vacuity proven by blanking each step file and confirming failure. The GSD_WS forwarding guard was checked as a possible real break and is not one: the charclass narrowing is intact and forwarding works end to end. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): drop the ack for a newly-added reference file A new file's emitted ripple is attributable to the diff that adds it, so the acknowledgment explained nothing and the differential check reports it as stale. Removing the last entry removes the fragment — an empty one signals nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): retarget the UI-contract guards and clear two transitive advisories The §3a.5 extraction that brought autonomous.md under the #2196 char cap moved its body to references/autonomous-ui-design-contract.md, so ten guards in autonomous-ui-steps and check-ui-safety-gate were asserting it against the host. Retargeted via a combined read, each proven non-vacuous by blanking the reference file and confirming failure. This class had already bitten twice on this branch because each sweep was scoped to the workflows touched at that moment, so this one was exhaustive: ~70 test files across all 13 workflows, zero further broken or vacuous assertions found. Also clears two high transitive advisories the matrix flagged on one lane — fast-uri GHSA-7p8r-x3mc-p8w7 and three ip-address SSRF/trust-boundary issues. Both pre-date this branch: package-lock.json was untouched until now, so the production tree was byte-identical to the base. Lockfile-only, package.json unchanged, verified against a real npm ci install. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): backfill changeset pr number to 3030 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
88f6d9bd1b |
fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu * fix: preserve installer executable mode * chore: add changeset for PR #2812 * test(#2644): acknowledge Cursor emission changes * test(#2644): drop spent emitted drift acknowledgments * fix(#2644): remove retired Cursor command converter --------- Co-authored-by: clezcoding <clezcoding@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a987cf2731 |
chore(#2932): emit a per-invocation section manifest from the init bundle (#2987)
* chore(#2932): emit a per-invocation section manifest from init Extends the init bundle with a typed per-invocation section manifest so an invocation loads only the branch guidance it will actually take. The three flag/state-gated branches in execute-phase.md move into their own step files; the parent keeps its gsd:section markers wrapping a one-line on-demand reference, so each section's prose lives in exactly one file and the parent shrinks 93369 -> 89507 bytes. A new drift-guarded generator derives the shipped section manifest from those markers, and a new pure evaluator maps invocation facts to applicable section ids. The evaluator is a lookup over the frozen WHEN_VOCABULARY, never a parser (Greenspun's Tenth Rule, ADR-1671:69); a parity test asserts the vocabulary and the predicate map stay exhaustively in sync. Closes #2932 * fix(#2932): fail closed on prototype-chain when values An isolated adversarial review found WHEN_PREDICATES[section.when] was a bracket lookup on a plain-prototype object, so inherited Object.prototype members resolved as predicates: "constructor"/"toString"/"valueOf"/ "hasOwnProperty" returned truthy and SILENTLY INCLUDED the section, and "__proto__" threw an untyped TypeError carrying no .reason. Both violate the module's documented fail-closed contract, and the manifest is read from disk at run time so it cannot be assumed trustworthy. Builds the predicate map on a null prototype and guards the lookup with an explicit Object.hasOwn check. Adds table-driven coverage for nine Object.prototype-shaped keys asserting the TYPED reason (asserting only that it throws would still pass while broken) plus a fast-check property injecting a hostile value at an arbitrary document position. * test(#2932): retarget execute-phase step assertions at extracted step files * fix(#2932): emit typed reasons for generator lib-load and write failures * fix(#2932): restore launcher preamble in extracted steps and refresh derived fixtures * chore(#2932): backfill changeset pr number to 2987 --------- Co-authored-by: sim <sim@local> |
||
|
|
c61dd49d95 |
enhance(#2255): blocking catastrophic-shrink guard for curated .planning/ writes (#2301)
* feat(#2255): blocking catastrophic-shrink guard for .planning writes Adds hooks/gsd-write-guard.js, a PreToolUse hook that hard-blocks (decision: 'block', exit 2) a whole-file Write collapsing a curated .planning/ artifact (ROADMAP.md, .planning/milestones/*-ROADMAP.md, STATE.md) below 40% of its on-disk line count. Files under 40 lines are exempt; GSD_ALLOW_PLANNING_SHRINK=1 (named in the block message) bypasses for legitimate milestone resets. Fix 3 of #973 — the only defense independent of per-agent tool config. Registered on the Claude plugin surface (hooks.json), settings-json runtimes (runtime-hooks-surface.cts, self-contained pattern), Kimi spec, and the OpenCode/Kilo plugin buses. Golden install fixtures and INVENTORY regenerated; regression tests negative-controlled (16/16 RED with the hook absent, 16/16 GREEN with it present). * chore(#2255): backfill changeset pr number to 2301 * enhance(#2255): address review — fail-closed reads, typed block output, registration, property test Review fixes for trek-e's CHANGES_REQUESTED on PR #2301: - Blocker 2: register gsd-write-guard.js in BUNDLED_GSD_HOOK_FILES (no-shipping-drift test). - Blocker 3: update the always-on hook enumerations in ADR-766 and CONTEXT.md from six to seven. - Major 4: fail CLOSED on non-ENOENT read errors — only a missing file (new-file Write) passes; EACCES/EISDIR/ELOOP/etc now block, with a typed readError field and the override still honored. Tested, with a negative control against the pre-fix hook. - Major 5: fast-check property test for the SHRINK_RATIO/FLOOR_LINES budget contract (blocked ⟺ newLines < oldLines*SHRINK_RATIO above the floor; sub-floor always exempt), boundary examples pinned. - Major 6: block output now carries typed oldLines/newLines/ overrideEnvVar fields; tests assert on those instead of regexing the free-form reason string. - Minor: CURATED_PATTERNS are case-insensitive (case-insensitive-FS bypass on macOS/Windows); limit+1 boundary tests added for both the floor and the ratio. * enhance(#2255): engage the write guard on Kimi's native payload shape The guard shipped with Claude-vocabulary checks (tool_name 'Write', tool_input.file_path), which #2304 showed leaves a guard dormant on Kimi: the [[hooks]] matcher is registered pre-translated but kimi-cli forwards its native payload verbatim — tool_name 'WriteFile' (bare or module-qualified) and tool_input.path per its tool schemas (src/kimi_cli/tools/file/write.py). The guard matched, saw an unknown name, and exited 0. Apply the same per-guard normalization PR #2326 gives the three sibling guards (name + field mapping, inlined — hook scripts stage as standalone files), and write the block reason to stderr as well as stdout JSON: Kimi feeds stderr, not stdout, back to the model on exit 2, so a stdout-only reason blocks without telling the model why or naming the documented override. Regression tests pipe Kimi-shaped payloads (engage, qualified-name, stderr-reason) plus exemption pins (StrReplaceFile stays out of scope by design; non-curated paths pass) — verified red against the pre-fix guard, green after. * enhance(#2255): rebase onto next; regenerate golden-parity fixtures * enhance(#2255): wire the escape hatch into complete-milestone's reorganize step Review Blocker 1: the guard hard-blocked /gsd:complete-milestone's ROADMAP reorganize — the tree's only legitimate milestone reset and the exact caller GSD_ALLOW_PLANNING_SHRINK was built for. The reorganize step now performs the rewrite through a shell write with the hatch set on the command (a hook inherits the runtime env, so a bare Write cannot carry a per-step override), and a binding test derives the env var name from the guard's typed output and asserts (a) the workflow step sets it and (b) the guard passes the identical catastrophic payload under it — so the next complete-milestone.md edit cannot silently re-break the wiring. * enhance(#2255): drop dead Edit-class mapping from normalizeKimiPayload Review Major 1: StrReplaceFile -> 'Edit' and the old_string/new_string reconstruction were unreachable-by-effect — the guard exits 0 for any tool_name !== 'Write', so nothing ever read the fields they set, leaving guaranteed-surviving mutants against the Stryker bar. The map now carries only WriteFile -> 'Write'; the StrReplaceFile exemption test message states the fall-through it actually exercises. * enhance(#2255): review minors — American spellings; writeSync before exit(2) Minor 1: normalised/normalise -> American house style. Minor 2: the two block paths wrote stdout+stderr via async pipe writes then exit(2) — async-on-Windows, unflushed at exit; fs.writeSync(1/2, ...) makes the block payload durable. * enhance(#2255): assert stderr equals the typed reason, not raw prose Minor 3: the last raw-text match in the suite pinned override-name prose on stderr. The contract is "stderr carries the reason Kimi feeds back" — now asserted as stderr non-empty and byte-equal to the parsed stdout.reason. * enhance(#2255): bind the write-guard's Kimi normalization into the parity test Review Major 2: the guard's normalizeKimiPayload is a 4th inlined copy with nothing binding it. This extends PR #2326's kimi-guard-normalization-parity test (same path and helpers, authored as a superset so either merge order resolves cleanly): sibling byte-parity is existence-gated zero-or-all — trivially green until #2326 lands, full-strength after — and the write-guard copy is bound semantically (map is the value-inverse of convertKimiToolName; the Kimi name for Write must map, or the guard is dormant on Kimi; the path -> file_path half must be present). Byte-parity is deliberately not asserted for this copy: it legitimately omits the Edit-class mapping (Major 1 — dead code in a Write-only guard). * enhance(#2255): refresh golden-parity fixtures for revised guard + workflow * chore(#2255): regenerate golden fixtures after rebase onto next The committed fixture hashes were generated against a tree predating next's latest 11 commits, which independently modified the same install-parity surface. Rebased onto next and regenerated with `npm run gen:golden`. Verified: against upstream/next the regenerated fixtures differ by exactly this PR's own entries -- hooks/gsd-write-guard.js (new), hooks/managed-hooks-registry.cjs, plugins/gsd-core.js, and gsd-core/workflows/complete-milestone.md. No unrelated drift. * fix(#2255): regenerate workflow size baseline for complete-milestone `complete-milestone.md` grew 31071 -> 32061 (+990) when the round-2 review fix bound GSD_ALLOW_PLANNING_SHRINK=1 into the reorganize step, but tests/workflow-size-baseline.json was never regenerated. The per-file workflow baseline test (issue #1074) failed on ubuntu-latest/22 and both macOS shard 1/3 jobs. The growth is justified: it is the escape-hatch binding requested in review round 2 (the guard must not hard-block the tree's only legitimate milestone reset), not incidental bloat. Regenerated via `npm run size:baseline`; the diff is exactly the one entry. * chore(#2255): regenerate golden fixtures and size baseline after rebase onto next * enhance(#2255): bind the shrink escape hatch mechanically — single-use sentinel the guard consumes Round-5 M1: the per-step `GSD_ALLOW_PLANNING_SHRINK=1 tee` prefix was inert (no PreToolUse hook exists on Bash in this family; the write succeeded by dodging the guard, not by the override firing) and the protection was prose. The hatch is now a transport code consults: complete-milestone's reorganize step arms `.planning/.gsd-allow-shrink` with the target's path, keeps the Write tool as the sanctioned path, and the guard — at the block point only — verifies the sentinel is fresh (15 min) and names the pending target, then CONSUMES it and allows that one write. Path-bound + single-use + freshness keep it from becoming a standing unlock. The env var remains as the interactive transport, where it can actually reach the hook. Regression tests written first (negative control: 3 failed pre-fix): the armed-sentinel Write passes and consumes; stale does not exempt; a token for a different file neither exempts nor is consumed; the binding test now takes the sentinel name from the guard's typed output (overrideSentinel), asserts the step arms it, and asserts the step no longer routes the rewrite around Write via a shell pipe. Also in this commit, same file: - m2: block emission is exception-safe — emitBlock() wraps both writeSync sites in their own try/catch that still exits 2, so an EPIPE can no longer convert fail-closed into the outer catch's fail-open. - Header discloses the two reviewed design limits (cumulative sequential shrink; lexical match vs symlinked paths) per round-5 scoping. * docs(#2255): document the sentinel transport across guard surfaces; changeset ends with the (#2255) parenthetical (m4) USER-GUIDE bullet, INVENTORY row (en + ja/ko/pt/zh), the runtime-hooks-surface registration comment, and the changeset now describe both hatches — the single-use sentinel for workflow steps and the env var for interactive use — instead of implying a per-step env can reach a hook. The changeset's trailing `Resolves #2255.` prose becomes the `(#2255)` parenthetical the repo's fragments use (round-5 m4). * chore(#2255): regenerate derived families on the rebased tree (full sweep) Full generator sweep after rebasing onto next @ the body-parser-patched lockfile: build, gen-inventory-manifest, gen:golden, size:baseline. Every regen delta verified to be either a PR-owned entry (gsd-write-guard.js, complete-milestone.md, INVENTORY/USER-GUIDE) or exact convergence to next's committed value for entries our arbitrary-side conflict resolution had left stale (all 18 runtime fixtures checked mechanically). * test(#2255): use helpers.cleanup for sentinel teardown, not raw fs.rmSync The repo's local/no-raw-rmsync-in-tests rule exists for the Windows-EBUSY retry budget; the sentinel disarm now rides it like every other teardown. * chore(#2255): regenerate derived families after rebase onto next Full sweep on the rebased tree (build -> gen-inventory-manifest -> gen:golden -> size:baseline). Every delta is either a PR-owned entry (hooks/gsd-write-guard.js, its registration surfaces hooks/managed-hooks-registry.cjs and the two plugin buses, gsd-core/workflows/complete-milestone.md) or exact convergence to next's committed value across all 18 runtime fixtures. * chore(#2255): regenerate derived families after rebase onto next @ |
||
|
|
cc3ee301a7 |
fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root (#2593)
* fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root installSharedHooksBundle wrote `{"type":"commonjs"}` over <configRoot>/package.json unconditionally — no existence check, no merge, no backup — on every install and every /gsd-update re-install. On the 11 affected runtimes that file is often user-owned; on OpenCode and Kilo it is the documented place to declare local-plugin npm dependencies, so a user's name/type/dependencies/scripts were destroyed on each run. The uninstall path already read the file and unlinked it only on an exact content match. That asymmetry was the defect: the discipline existed in the codebase, it just was not applied on the write side. Move the marker into the directories GSD creates and fills with its own .js files — hooks/ (all shared-hooks runtimes, incl. Kimi's own root) and the nativePlugin dir (plugins/ for OpenCode+Kilo, extensions/ for pi) — and stop writing the config root entirely. New src/commonjs-marker.cts owns the marker string plus one ownership predicate (absent / gsd-owned / foreign, fail-closed on an unreadable file) shared by ensureCommonJsMarker and removeCommonJsMarker, so install and uninstall cannot drift apart again. Nothing else depended on the config-root marker: package identity is baked at build time (#378/#498) and version resolution prefers gsd-core/VERSION and already tolerates a missing root package.json (#1383) — Codex has installed without one all along. A package.json in plugins/ or extensions/ is inert to plugin discovery, which globs *.{ts,js} only (see installer-migration 006). Uninstall retires the pre-fix config-root marker, so upgrading users are cleaned up on removal, and still never touches a file it did not write. * fix(#2544): point the changeset fragment at the filed PR The fragment's `pr:` field is only knowable after `gh pr create` returns. * fix(#2544): register commonjs-marker.cjs in the tsc-generated ESLint ignore set bin/lib/commonjs-marker.cjs is tsc output (src/commonjs-marker.cts is the linted source), so it belongs in the ADR-457 ignore list like its siblings. Clears the lint-tests no-var failure and the repo-invariants "linted xor ignored" migration-state test. * fix(#2544): pin the kimi CommonJS marker to hooks/, not the ~/.kimi root The UPGRADE 1 test still asserted the pre-#2544 marker location (~/.kimi/package.json). The marker now lives inside ~/.kimi/hooks — the directory GSD itself creates — matching the updated golden-install-parity and install-tree fixtures. Also asserts the root marker is NOT written. * fix(#2544): make the CommonJS marker write path non-fatal Review round 2, Major 3 + Minor 1 + the stagedHooks nit. ensureCommonJsMarker rethrew any non-EEXIST write error and neither call site caught it, so EACCES on a read-only hooks/, EROFS, or ENOSPC aborted the whole install with a raw stack trace. Every other marker interaction in the module is best-effort — removeCommonJsMarker swallows unlink failures, classifyMarker swallows read failures — and this was the write path, i.e. the one most likely to fail on a locked-down config dir. It now returns a new 'failed' outcome and both call sites warn and continue. Sibling found while sweeping for the same defect class: fs.mkdirSync sat OUTSIDE the try block, so an unwritable parent threw past the guard entirely. Creating the directory is the same environmental hazard as writing into it, so it moved inside. Also in this file: - The hooks marker is now gated on `stagedHooks && hooksOk`, not stagedHooks alone. stagedHooks is computed from the SOURCE listing before the copy loop, so it stays true when the copies land but verifyInstalled() then fails — marking a hooks/ GSD did not successfully populate claims an ownership the install did not earn. - The uninstall rmdir of the native plugin dir is gated on GSD having actually removed something from it. Hoisting it out of the adapter-exists guard (so the marker-only case could prune) had silently widened it into deleting a user-created but empty plugins/ or extensions/ dir — the same "don't touch territory GSD didn't fill" principle this issue is about, inverted. - Kimi's pre-#2544 marker at its native hook root (~/.kimi) is retired at the same call site that writes its replacement. That path is outside kimi's configDir, so installer-migration 007 structurally cannot reach it. * fix(#2544): retire the stale config-root marker via installer-migration 007 Review round 2, Major 1 — the PR's headline claim was false for existing installs. Upgraders kept BOTH markers: the new one under hooks/ and the stale {"type":"commonjs"} at the config root, so their config root stayed pinned to CommonJS and their dependency manifest stayed gone until they uninstalled. The migration is unusual in one way, and it is the part worth reviewing: the config-root marker was never recorded in gsd-file-manifest.json (writeManifest records hooks/, agents/, commands/, scripts/ and the native plugin, never a root package.json), so classifyArtifact answers 'unknown' for it and the planner's own guard downgrades a remove-managed on an 'unknown' classification to preserve-user. 007 therefore supplies the "purpose-built detector for an old GSD-owned shape" that docs/installer-migrations.md#remove-managed sanctions — exact content match, the same predicate removeCommonJsMarker has always used — and declares the resulting classification on the action. A package.json with any other content is left untouched, and there is deliberately no backup-and-remove branch: a non-matching file here is not a patched GSD artifact, it is somebody else's file. Scope is all runtimes. The `runtimes` field is OMITTED rather than `[]`: validateStringArray requires the field to be non-empty WHEN PRESENT, while the runtime filter treats an empty array as "all" — so `runtimes: []` throws at plan time and the migration never runs. The metadata test pins this. Kimi is a deliberate carve-out, named in the migration's own header: its marker lived at ~/.kimi, outside kimi's configDir, and migration relPaths are structurally confined to configDir. It is retired by the installer instead. Registration: shipped-migrations table, .gitignore for the emitted .cjs, the EXPECTED_CHECKSUMS baseline, and the ESLint ignore set. That last one is not copied from migration 006 by rote — 006 needs no entry because it imports nothing, while 007 imports node builtins, so tsc emits its __importDefault helper and the `var` in it trips no-var. This is the same lint gate that made round 1 red. * test(#2544): fault-injection and multi-runtime marker coverage Review round 2, Major 2 + Minors 4 and 5. Major 2 — CONTRIBUTING.md:514-531 is mandatory for install/uninstall flows and the suite had no fs monkeypatching at all. Every branch now covered is one whose doc comment claims it as the module's safety posture: - classifyMarker non-ENOENT lstat error -> 'foreign' (the fail-closed rule), with an ENOENT control alongside it so the test discriminates rather than just asserting one side - classifyMarker readFileSync throw -> 'foreign' (present-but-unreadable never downgrades to the permissive answer) — the fixture's bytes are exactly GSD's marker, so the test fails if the code ever answers on content it could not read - a DIRECTORY at the marker path (CONTRIBUTING:521; the symlink case was already covered with a real symlink, the directory case needs no injection at all) - the ensureCommonJsMarker TOCTOU EEXIST branch — the entire reason for flag:'wx' - the new 'failed' outcome, for both writeFileSync (EACCES/EROFS/ENOSPC) and the mkdirSync that used to sit outside the guard - removeCommonJsMarker unlink throw -> false These save and restore fs methods in `finally` rather than using chmod 0o000, which does not fault under root and would pass vacuously in root Docker and CI. Minor 4 — uninstall was driven for opencode only. pi's extensions/ and both kimi locations now have behavioral coverage, install and uninstall, each paired with a user-authored-file case proving GSD leaves it alone. Minor 5 — the stagedHooks gate had no assertion behind its stated reason. A pre-existing, GSD-untouched hooks/ directory is now driven through a runtime that declares skipSharedHooksInstall and asserted to stay marker-free, with its user content intact. Also regression-tests the uninstall rmdir gate from the previous commit: an empty plugin dir GSD removed nothing from must survive. * docs(#2544): correct stale marker prose, register the module, document the trade-off Review round 2, Minors 2, 3 and 6. Minor 2 — six files asserted the installed ROOT ships the synthetic marker. None was load-bearing (all three walk-up consumers are VERSION-first with try/catch and the marker never carried a `version`), but ADR-457:52 is the rationale for keeping a generated module, so a future reader would mis-derive the constraint from it. Each site is corrected to what is now true: the installed tree carries no package.json with a .name at all, because the only ones GSD stages are {"type":"commonjs"} markers and they now live in GSD's own directories. Two of the six needed more than a location swap. hooks/gsd-check-update-worker.js and the platform-gate test both described `require('../package.json').name` resolving to undefined; post-#2544 that require does not resolve at all, so the history is kept accurate and the present-tense claim corrected rather than just moved. And src/runtime-artifact-conversion.cts described the no-root-package.json case as Codex-only — it is now every runtime, which strengthens that comment's own argument for lazy resolution. The generated .cjs sibling needs no edit: it is gitignored build output, not a tracked file. Minor 3 — src/commonjs-marker.cts had no CONTEXT.md entry, unlike every peer module, and CONTEXT.md is the #2 co-change partner of bin/install.js. Added, including the fail-closed posture and the never-throws contract. Minor 6 — the plugins//extensions/ marker shadows the config root for all .js siblings, so an OpenCode/Kilo user's ESM plugin/*.js stays broken. That is exactly what #2544's Fix section prescribed and it is disclosed in the PR body, but the PR body is not documentation. It now lives in the OpenCode section of docs/how-to/install-on-your-runtime.md, stated as a real constraint rather than a pure improvement, with the .ts mitigation and a fallback for ESM plugins. * test(#2544): attribute the CommonJS marker in the emitted-provenance rules The differential emitted-attribution gate (#2723, landed on `next` after this branch was cut) went red on the macOS shards once this PR rebased onto it. Two distinct causes, both real gaps rather than noise: 1. `plugins/package.json` and `extensions/package.json` matched NO rule — the `native-plugin` rule covers `*.{js,cjs,mjs}` only, so the marker read as an unattributed emitted family. 2. `hooks/package.json` fell through to `hooks-built`, which attributes an emitted `hooks/<X>` to a repo source `hooks/<X>`. There is no `hooks/package.json` in the repo, so it resolved to a nonexistent path. Cause 2 is exactly the failure already documented three lines above it for Copilot's `gsd-session.json` — "a code literal, not a built script" — so the fix follows that precedent rather than inventing one: `package.json` is excluded from `hooks-built` the same way, and a dedicated `commonjs-marker` rule attributes the family across all four roots it can appear in (both hooks roots plus `plugins`/`extensions`) to the sources that actually emit it. Deliberately a RULE, not an entry in tests/emitted-drift-ack.json. An ack is for a one-off ripple and goes stale by design — the gate fails a stale ack precisely so it cannot pre-clear the next change on that path. These markers are a permanent part of the emitted tree from #2544 onward, so they need standing attribution. Verified by reproducing the CI failure locally with GSD_EMITTED_BASE: 3 provenance errors + 12 unattributed paths before, 35/35 green after. * fix(#2544): route the #2717 hooks-surface marker helpers through commonjs-marker #2717 landed a second copy of ensureCommonJsMarker/removeCommonJsMarkerIfGsdOwned in src/runtime-hooks-surface.cts for the runtimes that stage .js hooks via dedicated paths (cursor/windsurf/codex). That copy had drifted from this PR's module on the two properties that matter: - ownership probe: `fs.existsSync` FOLLOWS symlinks and reports false for a DANGLING one, so a dangling package.json symlink classified as absent and the write went straight through it. Demonstrated: against the pre-fix copy, ensureCommonJsMarker() on a hooks/ dir holding a dangling package.json symlink returns true and creates {"type":"commonjs"} OUTSIDE that directory. - create: a plain writeFileSync leaves the classify->write window open, where commonjs-marker creates with flag:'wx' (O_EXCL). Both helpers now delegate to src/commonjs-marker.cts, which is what this PR's own docstring already claimed was the single place these rules are enforced. Exported signatures are unchanged (still boolean), so bin/install.js and the #2717 tests are unaffected. The new subtest is the only coverage that fails if the duplicate is ever reintroduced — the two implementations agree on every non-adversarial input, so the existing suites pass against both. * test(#2544): pin the stagedHooks gate on zcode, not windsurf The Minor-5 coverage picked windsurf because hostBehaviors.skipSharedHooksInstall kept it out of the shared hooks bundle, so GSD staged nothing into hooks/ and the marker was correctly absent. #2717 changed that premise: cursor/windsurf/codex now stage their .js hooks via dedicated paths and get the marker beside those scripts. Measured on this tree, windsurf stages 2 .js hooks and receives a marker — so the assertion was pinning behaviour that is now wrong, not the gate it was written for. ZCode is the durable choice: per #1821 it has hooksSurface:'none' AND no plugin surface to spawn hooks, so GSD stages no .js there by either route (measured: 0 staged, no marker). The property under test is unchanged — a user-created hooks/ directory GSD never fills stays marker-free. * test(#2544): use the shared cleanup helper in the migration test Addresses the review's Major 1. The suppression's stated reason — "no helpers import available" — was not correct: tests/helpers.cjs exports cleanup, and the other test file added in this same PR imports it (tests/commonjs-marker.test.cjs). The local reimplementation dropped two protections that are live on this repo's windows-latest lane: the CWD guard (Windows cannot remove a directory that is the current working directory) and the 20 x 250ms retry budget that absorbs the deferred-scan handle Windows Defender holds on newly-written files. Local function and suppression both removed; local/no-raw-rmsync-in-tests now passes without one. * test(#2544): expect hooks/package.json for the #2717 runtimes The fresh-install contract table predates #2717, which stages cursor/windsurf/ codex .js hooks via dedicated paths and writes the CommonJS marker beside them. All three therefore now receive hooks/package.json legitimately. Measured on this tree: codex stages 3 .js hooks, cursor 6, windsurf 2 — each with the marker; cline/copilot/trae/zcode stage none and get none, so their contracts are unchanged. * fix(#2544): gate the #2717 marker writes on having staged something The three dedicated marker writers #2717 added ran unconditionally. Each one mkdirs hooks/ up front and stages its scripts conditionally on the source existing, so with an absent or empty hook source they created a directory, filled it with nothing, and marked it as GSD's anyway. That is the same write-into-someone-else's-territory this issue is about, and installSharedHooksBundle already guards the identical case with `stagedHooks`. The dedicated paths now carry the matching gate: - cursor / windsurf: `installedScripts.size > 0` - codex: a new `codexStagedHooks` flag. The enclosing guard only proves that hooks/dist EXISTS; it says nothing about whether any CODEX_HOOKS_TO_COPY entry landed. Covered for cursor and windsurf by driving each writer against a src tree whose hooks/ dir is empty. The codex leg is defensive and deliberately uncovered: its trigger state needs a package tree where hooks/dist exists but holds none of the allowlist, which is not constructible from a real checkout. * test(#2544): scope the commonjs-marker sources per root The rule declared one flat source list for every marker root, so `extensions/package.json` was attributed to runtime-hooks-surface.cts (which never writes there) and `.kimi/hooks/package.json` to install-engine.cts. That is not merely untidy. emitted-diff.cjs accepts the FIRST satisfied source, so a flat list containing bin/install.js let any change anywhere in that 13k-line file authorise marker drift for every root — the blanket escape hatch this file's own agents-verbatim comment refuses for exactly the same reason. Sources are now derived per root from ctx.rel. Note the rule ctx is `{ rel, runtime }` and carries no `root`, so keying on ctx.root would have sent every path down one branch silently. * test(#2544): state precisely what the zcode assertion pins The comment claimed the test pinned installSharedHooksBundle's `stagedHooks` gate. It does not, and neither did the windsurf version it replaced: zcode declares skipSharedHooksInstall, so the outer guard skips that helper entirely and the gate is never evaluated. The test passes on the runtime exclusion. What it does pin — the outcome a pre-existing, GSD-untouched hooks/ stays marker-free — is still worth having, and is what the review asked for. The two `staging zero hook scripts` tests are the ones that pin a real staged-nothing gate. Comment corrected rather than left implying coverage that is not there. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
640eaee16e |
chore(#2930): fragmentize execute-phase.md and prove per-runtime composed emission (#2972)
* feat(#2930): fragmentize plan-phase.md workflow into per-runtime-composed sections Adds src/workflow-fragments.cts (in-file <!-- gsd:section --> marker parser/composer, ADR-1671 epic #1671 Phase 3), wires it into bin/install.js's copyWithPathReplacement emission path, and pilots the marker grammar on gsd-core/workflows/plan-phase.md. Bookkeeping ripple for the new src/*.cts module: .gitignore, eslint.config.mjs, docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json, and a CONTEXT.md glossary entry. Amends ADR-1671 with open questions 1 and 2 resolutions and records the closed when= applicability grammar. Adds docs/reference/workflow-fragments.md and an ARCHITECTURE.md section documenting the marker authoring model. * fix(#2930): put allow-test-rule issue ref on the same line as the marker lint-allow-test-rule-refs.cjs requires the #NNN issue reference on the same source line as `allow-test-rule:`; it was one line below and read as an unreferenced novel exemption. * docs(#2930): link the orphaned gate-predicates reference from the docs index Found while adding the workflow-fragments reference doc: docs/reference/gate-predicates.md shipped without an entry in docs/README.md, so it was unreachable from the docs index. Fixed inline rather than deferred. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2930): scope composition to workflows, add typed failure reasons Review findings from two orthogonal passes: - Scope composeWorkflow to gsd-core/workflows/ only. It previously ran on every .md the installer copied, so a future agent/command/reference doc documenting the marker syntax with an unfenced example would have been mis-parsed and silently stripped — a lossy drop the phase forbids. - Add a frozen REASON enum; failures attach a typed .reason and tests assert on it instead of matching free-form message text (CONTRIBUTING.md:635-694). - Derive the property generator's when= values from WHEN_VOCABULARY instead of duplicating them (DEFECT.GENERATIVE-FIX). - Add adversarial parser fixtures: Unicode headings, NUL, U+FFFD, BOM, fence-within-fence, tilde and indented fences, lone-CR marker line. - Document why --mvp is structurally unmarkable: its content is interleaved, not sectioned, so the whole-line grammar cannot reach it. Also fixes two stale tests on this branch, each reproduced on the unmodified tree before correction. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2930): retarget the pilot from plan-phase to execute-phase The full remote matrix went red on both Linux lanes. Root cause was ours: tests/phase6-capstone-conformance.test.cjs holds a PRE_PHASE6 ceiling of 94519 bytes for plan-phase.md, asserting an ADR-857 Phase-6 completion property. That is a third size gate beyond the tier caps and the differential ratchet, and it left plan-phase.md just 36 bytes of headroom rather than the 3821 computed from the XL cap. The 330 marker bytes overran it by 294. Raising the ceiling is not an option: it is a red line certifying another ADR's completion. plan-phase.md is reverted to byte-identical origin/next and the pilot moves to execute-phase.md, which has 728 bytes of headroom under its own ceiling and lands at 93147 with 3 marker pairs. The vocabulary narrows to the atoms actually used: always, flag:--wave, state:gap-closure-phase, state:has-prior-phases. Recorded in the ADR: every branch the epic names lives in plan-phase.md, which cannot be fragmentized until caps move from source to emitted bytes. That is direct evidence for the epic's premise and may reorder phases 3-4. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2930): backfill changeset PR number (#2972) * fix(#2930): make the emission install tests portable on Windows The windows-latest lane went red on three tests in the new install suite; Linux was green. Both causes were in the test harness, not the module. Root normalization: the opencode converter always embeds the install root forward-slashed, but the tests stripped it with the native-separator string from mkdtemp. On Windows that never matched, so the root leaked through unstripped — and because the real and stub install roots have different prefix lengths, that length difference landed directly in the byte-delta assertion (344 observed vs 275 expected). Normalize both text and root to one separator form before stripping. @-ref resolution: the helper stripped only the @~/ and @$HOME/ forms, so a Windows absolute ref (@C:/Users/...) fell through and was joined onto the root, producing ...\@C:\Users\... Strip the @ first, then detect absoluteness from the token's own shape (POSIX, drive-letter, or UNC) with no platform branching, so every OS takes the same path. Neither assertion was weakened; the exact-equality byte check is the point of the test and still holds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2930): document every REASON member and guard the doc/enum parity Code review found the reference doc's 'Fails closed' list covering 10 of the 11 frozen REASON members — MALFORMED_ATTRIBUTES (parseAttrs rejects malformed key="value" syntax) had no bullet, and it is distinct from UNRECOGNIZED_ATTRIBUTE, which is valid syntax with an unknown key. Two parallel surfaces sharing one constant with nothing asserting they agree is the DEFECT.GENERATIVE-FIX class, so the same commit adds the parity assertion: the test derives the enum side from the built module and the doc side by parsing the reference page, keyed on the reason IDENTIFIER rather than prose so a reworded bullet does not break it, and reports set differences in both directions by name. Proven non-vacuous: removing the MALFORMED_ATTRIBUTES bullet turns the suite red naming that exact member; restoring it returns 44/44. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9ac0dfad58 |
chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared context-composer seam. Its success condition is that review-prompt output does not change, and the only authority on "did not change" is the behavior that shipped before the refactor. Capture that behavior now, while it is still the live implementation. 47 characterization cases, every `expected` value computed by executing the current implementation rather than hand-authored — the independence CONTRIBUTING.md "Fixture provenance (#2371)" asks for. A corpus is only worth what it can detect, so this one was validated by mutation rather than assumed. Five deliberate defects were injected and each must be caught by at least one case: - the note reserve deducted unconditionally instead of only under pressure - the pressure test relaxed from `>` to `>=` - a no-op head-shrink still setting the shrunk flag - the per-plan floor dropped from the proportional share - drop order reversed Two of those exposed real holes in the first cut of this corpus, and the cases that close them exist because of it: - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a floored plan group, and the 1024-char floor absorbed the entire trim, so the mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap, which makes the strict inequality observable as context kept vs omitted. - No case reached proportional-truncate at all — B6 and B7 both hard-failed the min-set pre-check first, leaving planTruncationPct at 0 across every case and the floor semantics entirely unexercised. Rebudgeted to 700 and 1100 so the min-set fits and the truncate step is actually reached; they now record 40.20% and 48.80%. The A4/A10 families sweep the pressure boundary from both sides, which is where this function has regressed before: CONTEXT.md's LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band, because the suite paired a trivially-fitting budget with a trivially-overflowing one and never sampled between them. A4 pins that nothing is trimmed from the cap down to 81 tokens under it; A10 pins that pressure fires at +1. Together with A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures. Two facts the corpus establishes that the design notes had wrong: - "" and null sections are NOT distinguished. applyBudget uses truthy checks throughout, so an empty-string section is treated as absent: not rendered, not dropped, never recorded in `omitted`. B13b pins this while the ladder is actively trimming, where only the non-empty `research` is dropped. - Sizing matters. B12/B13 were first written at a budget where both hard-failed the min-set check and returned "", so comparing them compared two empty strings and proved nothing. Committed as its own commit, ahead of the refactor, and regenerated against the pre-refactor implementation, so the oracle is demonstrably independent of the change it will adjudicate. Refs #2929 * refactor(#2929): extract the context-composer seam from prompt-budget Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer — per-runtime artifact emission — but it is walled inside the cross-AI review pipeline. Lift it into a shared seam so later phases can call it, without changing what the review pipeline emits. ADR-1671 specifies the composer as "priority + binary-search cutoff to a per-runtime budget". Read against the code it generalizes, that contract cannot express the thing being generalized. applyBudget is not a cutoff: it is a fixed five-step ladder in which each section carries its own shrink strategy, and only three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor; instructions and roadmap are never touched at all. A cutoff composer sorts by priority and discards the tail — it has no way to say "shrink this one", "truncate that one but never below 1 KB each", or "these three are the only droppables, in this order". Building to the literal contract and routing prompt-budget through it would have silently changed review-prompt output, which is the one outcome this phase forbids. So shrink strategies are the core abstraction here, and cutoff becomes one strategy among them — the right one for per-runtime emission in Phases 3-4, not for this ladder. That is an elaboration of the ADR's intent, not a departure from it, and ADR-1671 is updated to say so. Three decisions worth stating: - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan of surviving fragments and never a string. assemblePrompt's rendering is prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and owning it in the composer would force emission to adopt prompt-shaped rendering. The split is what lets one seam serve both consumers. - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires for emission caps. The existing code converts a token budget to a character budget with a hardcoded `* 4`; that assumption is now an explicit `charsPerUnit` inverse, which is precisely what a byte unit needs in order to reuse this. - The entry point is `composeWithinBudget`, not `applyBudget`. That name already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter being an unrelated graph-edge budget. A third would make every symbol search in this repo ambiguous, and it already misresolves: preflight and impact queries for "applyBudget" return graphify's. Behavior is unchanged and proven so: all 47 characterization cases reproduce byte-identically, and the corpus is mutation-validated rather than merely green (see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and from eighteen mutable accumulators to two, both inside a helper copied verbatim. estimateTokens deliberately stays in prompt-budget and keeps its exact math: src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins plan estimates and recorded actuals to that same scale, so moving or changing it would silently break the calibration loop. Refs #2929 * docs(#2929): document the context-composer seam and amend ADR-1671 Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new domain modules), and a mutation-matrix entry for the new module. The ADR amendment is the substantive part. ADR-1671 specified the composer as "priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2 established that a cutoff alone cannot express the function the platform generalizes, so the ADR now records shrink strategies as the core abstraction with cutoff as one strategy among them, reserved for per-runtime emission in Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned against that contract and would otherwise be planned against a mechanism that does not work. The mutation-matrix entry is not bookkeeping. Stryker scores per module against a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the extracted code unmeasured while prompt-budget's own score floated free of the logic it used to cover. context-composer gets its own entry at the same floor. Refs #2929 * test(#2929): pin the effectiveBudget rounding mode in the parity corpus An isolated correctness review found a real blind spot: mutating `Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator happened to produce a whole number, so floor, round and ceil all agreed and the rounding mode was entirely unpinned by a corpus whose whole job is to pin observable behavior. Three cases fix that by straddling the .5 boundary: A11 95 * 0.90 = 85.5 floor 85, round 86 -> the two disagree A12 97 * 0.90 = 87.3 floor and round agree; ceil (88) does not A13 93 * 0.85 = 79.05 same guard at a non-multiple-of-10 margin, so the margin arithmetic is exercised and not just the budget A11 alone catches the round mutation; all three catch ceil. Regenerated against the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so the expanded corpus keeps the independence property the original capture had. The corpus is now mutation-validated against seven injected defects, every one caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink setting its flag, the truncate floor ignored, drop order reversed, and both rounding-mode changes. Refs #2929 * feat(#2929): flexReserve floors and the byte-stable isolate prefix Two of issue #2929's "Done when" items were unimplemented rather than deferred, and an isolated review flagged them alongside my own audit. Both are part of ADR-1671's composer contract, so shipping the seam without them would have left Phases 3-4 building against a contract that does not exist yet. flexReserve is a per-fragment floor in measure units that every strategy must respect, which is what makes it different from the pre-existing floorChars: that one is a chars-denominated detail of proportional-truncate alone and is retained unchanged. A floored fragment is never dropped, is never head-shrunk below its floor, and raises its own proportional cap. A fragment already smaller than its floor is untouchable outright. Metadata gains `floored`, listing the ids whose floor actually prevented a trim — a guarantee no caller can observe is a guarantee no test can hold you to. isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed, never dropped, but still counted, because a prefix excluded from accounting would silently under-count real context. Metadata gains `isolatePrefix` so a caller can hash or assert on the exact bytes. Declaring an isolate fragment after a non-isolate one throws: a prefix that is not at the front is not a prefix, and accepting it would make the cross-runtime stability claim meaningless. Adds tests/context-composer.test.cjs for the exact new semantics and tests/context-composer.property.test.cjs for the five invariants, including the budget-monotonicity property the issue names explicitly. Both are registered in the mutation matrix, since coverage does not migrate with relocated code. prompt-budget uses neither feature, and its output is unchanged: all 50 corpus cases still reproduce byte-identically. Refs #2929 * chore(#2929): allowlist the prompt-budget parity suite The parity corpus needs its own test file and that makes prompt-budget a three-file module against a limit of two. The lint offers consolidation or an allowlist entry with justification; the entry is the right call here. Consolidation would mean folding the characterization suite into prompt-budget.test.cjs, which is the one thing that should not happen to it. The parity suite is a distinct concern with a distinct lifecycle: it is generated rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own scoring target, and its failure means something categorically different from a unit-test failure — not "this behavior is wrong" but "observable output moved". Burying it inside a general unit file would obscure exactly that signal. The allowlist is an identity ratchet, so this entry pins today's three exact filenames: adding a fourth still fails, and dropping back to two requires removing the entry. Refs #2929 * fix(#2929): register the new module with two gates it was missing The remote matrix caught three defects that no local check could, because the local runner is blocked in this repo and these suites had therefore never executed. Eight failures, identical on node22 and node24, so nothing environment-shaped. Two are the new-module ripple. A net-new src/*.cts lands in six places and this change had reached four of them — .gitignore, INVENTORY, the manifest, and the CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet baseline (a deliberate review-visible mirror of the matrix floors, which every COVERED module must carry). Both are now registered, the ratchet at the same floor of 66 the matrix declares. The third was a test asserting an outcome it had made impossible. It set budget:1 alongside a 400-char required fragment, so the group budget came out at -99 and the proportional-truncate step was skipped entirely — the deliberate "non-positive group budget is skipped, never clamped" rule inherited from the original ladder. Nothing was trimmed, and the test then asserted a truncation. Rebudgeted so the step actually runs, with the arithmetic written out in a comment so the next reader does not have to re-derive why 120 rather than 80. Fixing that surfaced a genuine bug in the composer. `floored` is documented as recording fragments whose flexReserve prevented a trim that would otherwise have happened, but the push sat in the else-branch of "content did not change", so it only fired when nothing was trimmed at all. A fragment truncated to a reserve-raised cap has also had a trim prevented — 40 characters' worth in the test above — and was silently absent from the field that exists to make the guarantee observable. The condition was already right; it was in the wrong branch. Now recorded on both paths: a drop prevented outright, and a truncation capped higher than the share alone would have allowed. Parity is unaffected — prompt-budget never sets flexReserve, so the branch is unreachable from every corpus path, and all 50 cases still match. Refs #2929 * chore(#2929): backfill changeset PR number (#2958) * chore(#2929): correct the corpus case count in the changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
05b170e448 |
chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam Productionizes the ADR-1671 Option-E reference example as a real module: src/context-predicates.cts (parser + selector + index builder) compiled to gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's --check/--write drift-guard idiom and wired into lint:generated-sync. Parser behavior is deliberately prototype-equivalent in this commit so the next commit's regression matrix binds to the real defects rather than to a missing module. Two locked design deviations from the prototype: - duplicates carry a count, not line numbers - the committed index carries no line field at all, resolving ADR-1671 open question 4: an artifact without line numbers cannot drift on a line shift, so promoting --check to a CI gate does not make it routinely red Also reconciles the one remaining duplicate predicate ID (RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording is removed) so the gate can land fail-closed on duplicates. Refs #1671 * test(#2928): failing-first matrix for the predicate fact-store Adds the regression matrix from the phase test plan: parser declaration forms, fence and comment regions, ID/value grammar boundaries at limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard CLI, the selector query surface, and four document-shaped fast-check properties. Seven rows are RED for behavioral reasons against the ported parser: indented-bare, star-list, plus-list and numbered-list declaration forms are dropped; a tilde fence and a four-backtick fence containing a shorter fence are not skipped; and a multi-line HTML comment is parsed as live. Eleven selector rows are RED because the query surface is not wired yet. Negative fixtures come from real repo documents that predate the grammar (CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the fixture-provenance rule, and the property generators are document-shaped rather than seeded from our own serializer. Refs #1671 * fix(#2928): consume the shared fence scanner, relocate the index, wire the selector Drives the failing-first matrix green. Parser: replaces the ported naive triple-backtick toggle with the shared markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord gain an export keyword — the only change to that module, which has 71 upstream dependents — because it already returns line-indexed spans, which is exactly what a line-reporting parser needs. It also already documents itself as the second copy of the fence state machine pending consolidation; adding a third copy here would have been the generative-fix divergence this repo warns about. A parity suite now pins predicate fence-skipping against that scanner across eight fence shapes. HTML-comment skipping stays local because the sectionizer has no comment scanner. Declaration forms widen to indented-bare, star, plus and numbered list items. Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The remote matrix run caught the original choice — a committed .cjs there ships ~120KB of CONTEXT.md prose into a runtime module, and two content guards fired truthfully on it (a leaked .claude install path, and four hardcoded package-name literals). Neither guard was allowlisted; the artifact moved instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to require it — it is a drift-detection artifact, so the selector parses CONTEXT.md live and is always current. Generator: adds a frozen REASON enum and --check --json so the gate's outcome is asserted structurally instead of by matching prose, and --context-path/--index-path so tests drive the real CLI against a temp tree with no filesystem monkeypatching. Selector: gsd_run query context-predicates with --class/--prefix/--contains, structured output carrying a matched count, own-property guards, and no project-root resolution. Registering it exposed that the query dispatch table and the usage string had drifted: a new parity test found 20 routed commands missing from the usage list, all added here rather than deferred. Refs #1671 * test(#2928): lock the newly-public scanFencedBlocks contract Exporting scanFencedBlocks made it public API for the first time, so it needs its own contract test independent of the consumer that motivated the export. Memtrace's co-change analysis flagged the gap: this suite changes together with markdown-sectionizer.cts 8 times in 90 days and was absent from the diff. Covers the documented rules: 0-based indices, -1 for an unterminated fence, the same-char/>=length/no-trailing-text closer rule, a shorter fence inside a longer one staying content, CommonMark 4.5 backtick-in-info-string, and <=3-space indent tolerance. Refs #1671 * fix(#2928): address both isolated review passes Two independent reviewers (correctness axis and security axis, neither the author) found seven findings. All are fixed here with regression tests; none deferred. BLOCKER — comment-blind fence scanning caused silent, permanent predicate loss. The HTML-comment scan and the fence scan ran as two independent passes, and the fence scanner is comment-blind, so a fence delimiter inside an HTML comment with no later close read as an unterminated fence and skipped every remaining line to EOF. Worse, the drift-guard could not catch it: it diffs against a baseline produced by the same corrupted parse. The two constructs now interleave in a single pass so each suppresses the other's boundary detection while active, covered in both directions. The parity suite still binds this scanner to markdown-sectionizer's for comment-free documents, so the two cannot diverge unnoticed. BLOCKER — the selector was not consumed anywhere, leaving the phase's acceptance criterion unmet. Now wired into the pre-work predicate-citation step in contributor-standards, which is the repo's actual brief-assembly path; no code-level brief assembler exists to wire into. MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id regex nested a dot-containing character class inside a dot-prefixed repeat, so N consecutive dots had exponentially many partitions: 40 dots took 565ms and growth was exponential. CI runs this parser over a pull request's own CONTEXT.md, so any contributor could have hung a shared runner with one line. Replaced with linear per-segment validation. Doubled-dot ids are now rejected; the real document contains none. MAJOR — the duplicate-id gate had only ever been proven on synthetic fixtures. A test now re-inserts the exact line this branch removed and asserts the real generator names it. MAJOR — --check together with --write silently let write win, turning the gate into a writer; a missing path value resolved to the cwd and leaked an EISDIR stack trace. Both are now clean usage errors. MINOR — the hoisted skip-list was exported as a live mutable Set; replaced with a read-only predicate. MINOR — flag-shaped selector values were unmatchable; the inline --flag=value form now provides the escape hatch. Refs #1671 * chore(#2928): backfill changeset PR number 2938 --------- Co-authored-by: sim <sim@local> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9a76ca6783 |
fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter
extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.
Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.
The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.
Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (
|
||
|
|
46ba02acde |
feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs * fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions * fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests * chore(#2630): backfill changeset pr to 2661 |
||
|
|
6ee4349272 |
fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) (#2642)
* fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) * chore(#2537): backfill changeset pr to 2642 |
||
|
|
f654c24a3e |
feat(#2505): Phase 4 — runtime-aware subagent dispatch (Option A; resolve-dispatch-type query) (#2525)
* feat(#2508): Phase 4 Option A — runtime-aware subagent dispatch via resolve-dispatch-type query (#2505) * fix(#2508): prose-variant preamble (avoid scanner-tripping literals) + namedDispatch===false-only mapping * fix(#2508): remove leftover old-preamble lines (keep prose variant only) * fix #2508: prose-only reference file * test #2508: regen golden install parity after workflow preamble additions * fix #2508: remove preamble from plan-phase.md (Phase 6 capstone ceiling); regen size+golden baselines * docs(changeset): backfill PR #2525 for Phase 4 (#2508) |
||
|
|
c5e0371775 |
feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging Red phase for issue #1951 (reversibility tagging: classify decisions by undo cost, gate one-way doors behind a checkpoint:decision). Tests assert, per the issue's acceptance criteria: - discuss-phase CONTEXT.md template records a **Reversibility:** field with a rationale on captured decisions, and states it is optional - gsd-planner @-references planner-reversibility.md and stays under the 49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs) - a one-way rating inserts a checkpoint:decision before the dependent task; reversible inserts none; costly is flagged but never blocks - the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard) and inserting a checkpoint implies autonomous: false - docs/reference/plan-md.md documents <reversibility> as optional with all three ratings - --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected into the planner prompt, and is advertised in the command argument-hint and help full mode (argument-hint parity) - the override suppresses the gate but still persists the rating - cmdVerifyPlanStructure accepts every rating and the absent case (additive-validator guarantee, behavioral via runGsdTools) - parity: thinking-models-planning.md #4 adopts the canonical three-level taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone - no content loss from the planner extraction made to fit under the cap Prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately — regression guards proving the validator already accepts unknown optional tags. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1951): reversibility tagging — gate one-way-door decisions Classify planning decisions by what undoing them would cost, and give a one-way door a human beat before the agent walks through it (issue #1951, The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way door framing). Acceptance criteria met: - discuss-phase records an optional reversibility rating with a rationale on <decisions> entries in the phase CONTEXT.md template. Unrated decisions are treated as reversible, so existing phases are unaffected. - a one-way rating makes gsd-planner insert a checkpoint:decision before the task that implements the decision, reusing the existing checkpoint mechanism -- no new checkpoint machinery. - reversible ratings trigger no checkpoint; costly ratings are flagged in the plan but never block. - the rating persists on the task as the optional <reversibility rating=> element. cmdVerifyPlanStructure accepts every rating and the absent case; the structural validator does not reject unknown optional tags. - --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings -- the override changes what stops the run, not what the plan remembers. Single taxonomy, not two: references/thinking-models-planning.md #4 already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the canonical three-level vocabulary and now points at planner-reversibility.md as the taxonomy owner, with a parity test that fails if the surfaces diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE). agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the checkpoint DO/DON'T guidance was relocated verbatim into planner-antipatterns.md -- already @-referenced from the same section for the same topic, so the planner still loads it and nothing was dropped. A test guards the relocation against content loss. Files: gsd-core/references/planner-reversibility.md (NEW, canonical taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase workflow/command/help (flag wiring + parity), plan-md.md schema, discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest, size baselines, install goldens, plugin skills regen, changeset. Closes #1951 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1951): address orthogonal review findings Two isolated reviewers (correctness + security), neither of which authored the change. Every finding fixed: Security — the rationale is untrusted input (ADR-1577). It originates in conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each hop an LLM reading the previous hop's output, with no validation on the path. planner-reversibility.md and the discuss-phase template now state it is data and never instructions, and name the </reversibility> early-termination hazard explicitly -- a rationale that closes its own element injects sibling structure the executor reads as real tasks. Four tests guard it. Correctness 1 — nothing machine-enforced the feature's own promise: a task rated one-way with no preceding checkpoint:decision validated as fully clean, so a planner error silently reopened the gap this feature exists to close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A warning, not an error: <reversibility> stays additive and the plan stays valid. Four tests cover ungated (warns), gated (silent), still-valid, and reversible/costly never flagged. Correctness 2 — pass-always test. The --no-reversibility-gates parse test substring-matched the whole workflow file, and plan-phase.md prose mentions both tokens in one sentence, so it passed with the bash conditional deleted: it was testing the documentation, not the parser. Now scoped to the fenced bash blocks and matched as one physical line, with a negative control confirming prose alone cannot satisfy it. Correctness 3 — costly had no itemized emission rule, only one-way did, so two agents could diverge on whether to tag costly at all. Correctness 4 — template convention break: the example ratings were bare while every sibling field uses [...] to signal substitution, inviting an LLM to copy one-way/costly forward as boilerplate. Now bracketed. Correctness 5 — latent false-green: .includes('reversible') also matches inside irreversible/irreversibility, which appear in anti-pattern prose, so a surface that dropped the real taxonomy entry would still pass. Now word-boundary matched. ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md 1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on next). The ceiling may only rise for privileged host machinery, and reversibility gating is optional-feature logic, so the wiring was slimmed to its minimum and the explanatory prose moved to the reference files the planner already loads. plan-phase.md is now 94400 bytes -- 119 under the ceiling and 70 bytes SMALLER than on next, so the host loop shrank while gaining the feature, which is what phase 6 ratchets toward. The tracer contract (tests/tracer-bullet.test.cjs) is unchanged. Lint — fixed an unnecessary non-null assertion in verify.cts and a CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST- PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture must carry the common task elements The gated-one-way fixture built a checkpoint:decision task from the abbreviated skeleton in gsd-planner.md, which shows only the checkpoint-specific elements (<decision>/<context>/<resume-signal>). cmdVerifyPlanStructure requires <name> and <action> on EVERY task regardless of type, so the fixture failed validation for reasons that had nothing to do with reversibility: errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"] Caught by gsd-test on 14d14a39 (2 failures, both this fixture). The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task carries <name>/<files>/<action>/<verify> like any other. Fixture corrected to match. Verified behaviorally against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Not a product defect: the validator's every-task contract is intentional and pre-existing, and docs/reference/plan-md.md scopes its required-element list to type=auto/tracer only because those are the elements a planner must author, not because checkpoints are exempt from <name>. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1951): backfill changeset pr number to 2471 * fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision Both CI failures were real defects in code this PR added, not false positives. CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 — the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`, which escapes the hyphen but not backslash, so the escape was incomplete. It was also unnecessary: `-` carries no special meaning outside a character class. Replaced with a complete metacharacter escape (backslash included). Word-boundary behavior verified unchanged across all three ratings — notably that "irreversible" prose still does not satisfy a "reversible" match, which is the false-green this helper exists to prevent. Prompt injection scan — the checkpoint fixture used the human-verification child element inside <verify>. That tag name is a fake-instruction-boundary pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed files, so copying the shape from tests/verify.test.cjs (unflagged only because it is not in this diff) tripped the gate. Switched to the documented plain-prose <verify> form. The first attempt at that fix failed the same gate a second time: the comment explaining the collision quoted the offending tag literally. The comment now names it in prose instead — the scanner does not care whether a match is code or commentary, which is the whole point of the DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md. Verified locally before push: scan reports 0 findings across 57 changed files, eslint clean, and both fixtures still validate as designed (gated one-way silent, ungated one-way warns, neither errors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): record measured cost and halve gsd-tools spawns The Windows shard 1/3 job timeout was traced to the sharding layer, not to this PR's assertions — see #2472. Two contributing factors were this file's own, and are fixed here. 1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs, so scripts/run-tests.cjs weighted it at the table's median fallback (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x under-weight. Recorded the measured value from the green gsd-test run (max across the node22/node24 lanes, per gen-test-timings.cjs's convention). Only this one entry: a full regen churns 634 entries of run-to-run drift, and the table is explicitly advisory and un-gated, so a 637-line diff does not belong in a feature PR. 2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost. Spawns cut from 9 to 6 with no coverage lost: - the ungated-one-way warning and its stays-valid assertion now share one plan instead of building the same plan twice; - the reversible/costly never-flagged-as-ungated test was strictly subsumed by the additive suite, which already runs those two ratings ungated and asserts no /reversibilit/ warning at all — and the gate warning's text contains both "reversibility" and "one-way", so the broader assertion catches it. It only re-spawned gsd-tools twice to prove the same thing. Both are symptom fixes. The shard imbalance itself (19/11/10 minutes against a 20-minute cap, from a cost-blind round-robin partition that also reshuffles downstream files whenever one is inserted) is tracked in #2472. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture adopts the #2444 type-branched contract Surfaced by rebasing onto next, which gained #2444 (branch plan-structure validation on task type=checkpoint:*) while this PR was in review. cmdVerifyPlanStructure no longer applies one required-element set to every task. A checkpoint:decision now requires <name> + <resume-signal> + <decision> + <options>, and is exempt from the <action>/<verify>/<done>/ <files> set that auto and tracer tasks carry. The gated-one-way fixture predated that split and failed on the new requirement: errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"] Fixture rewritten to mirror the checkpoint:decision contract exactly — real <options> with two <option> children — rather than padding it with fields checkpoints no longer need. That also drops the plain-prose <verify> the earlier revision carried purely to dodge the prompt-injection scan; a checkpoint task has no <verify> requirement at all, so the workaround is moot. Verified against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|
||
|
|
b6e6a22fce |
fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage (seven commits, never pushed) onto current origin/next as a single squashed commit. The original work was substantial and correct; this commit preserves its full scope, trimmed where rebase conflicts + workflow size budgets required it. Three independent layers where response_language was being dropped are closed: Layer 1 — orchestrator-facing directives across workflows. Adds the strong "All user-facing output in this workflow MUST be presented in {response_language}; technical terms, code, paths, and subagent prompts stay in English" directive to ~40 workflows that previously either lacked it entirely (verify-work, new-project, new-milestone, quick, manager, and ~35 more) or carried only the weak subagent-prompt-only form (plan-phase, execute-phase). The directive covers narration between tool calls and banner output, not just the AskUserQuestion prompts. Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts an optional responseLanguage parameter and renders the frame strings ("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.") in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/ Chinese/Korean/Italian) with an alias table covering ~30 input variants (en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads config.response_language via loadConfig(cwd) and passes it through, so the byte-for-byte block verify-work.md reprints verbatim is already localized when written — preserving the anti-injection hygiene rule at verify-work.md (the model is forbidden to translate after the fact). CJK display width is computed by East Asian Width property ranges (W/F) so the right ║ border of the banner stays aligned for full-width characters. English fallback is byte-identical to the pre-fix behavior when response_language is unset or unrecognized. Layer 3 — literal English report templates in execute-phase. The top-of- workflow directive covers all template sites (templates are a structural source, not literal output). Inline render-language notes that previously sat at each template site were removed during the squash because they pushed execute-phase.md over its frozen pre-phase-6 byte ceiling (93600 — ADR-857 Phase 6 capstone). The single top directive covers the same surface with fewer bytes. Also extends src/docs.cts and src/init.cts to propagate response_language into the init JSON bundle of the additional workflows so the directive can read it. Tests added: - tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls back to English default; recognized language swaps only the two frame strings while structural lines stay untouched; CJK display-width regression (independent recomputation of East Asian Width W/F ranges). - tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language wiring through docs.cts/init.cts. References: #2402; reporter's three-layer triage + Layer-4 follow-up; the byte-for-byte anti-injection hygiene rule at verify-work.md (the reason Layer 2 must be renderer-side, not model-translated). This is a squash of the in-flight bot branch — seven commits representing the original implementation plus its subsequent fix/CJK-padding/test/ changeset/regen cycles, none of which were ever pushed or PR'd. The squash captures the final coherent state. * chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md * chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451) Rebase conflicts were entirely in generated artifacts (golden-install-parity fixtures + workflow-size-baseline.json). After taking theirs during rebase, regenerated cleanly against the merged source tree. |
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
1720aacf0c |
feat(#1949): <precondition> task element — Design by Contract (#2422)
* test(#1949): add failing-first tests for <precondition> element Red phase for issue #1949 (Design by Contract: <precondition> element asserted before task execution). Tests assert: - docs/reference/plan-md.md documents the new <precondition> element - agents/gsd-planner.md @-references planner-preconditions.md and stays under the 49152-char cap (progressive-disclosure requirement) - gsd-core/references/planner-preconditions.md exists and documents the three emission cases mandated by the issue (user_setup / prior-phase artifact / env-var) and the contract triad mapping - agents/gsd-executor.md asserts <precondition> before task execution and routes unmet preconditions through existing checkpoint machinery - cmdVerifyPlanStructure (behavioral via runGsdTools) accepts plans both with and without <precondition> — the additive-validation guarantee - Parity assertion: plan-md.md and planner-preconditions.md agree on the canonical tag spelling (DEFECT.GENERATIVE-FIX-DIVERGENCE guard) Most prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately (regression guards proving the validator already accepts unknown optional tags). * feat(#1949): <precondition> task element — Design by Contract Add an optional <precondition> element to <task> in PLAN.md (issue #1949, The Pragmatic Programmer Topic 23). The front-of-task side of the plan contract — preconditions (before) ↔ postconditions (<verify>/<done>/ <acceptance_criteria>, after) ↔ invariants (must_haves.truths, across the whole plan). Together with the tracer-bullet proposal (#1945), this closes both ends of the 'outrunning your headlights' failure mode for an autonomous AI executor. Acceptance criteria met: - <precondition> is an optional element on <task>; plans that omit it validate unchanged (cmdVerifyPlanStructure checks for presence of required tags, does not reject unknown optional tags). - gsd-executor evaluates the precondition before any other task work. Unmet halts execution with a checkpoint:human-verify and no partial commit; met or absent produces no visible change to execution flow. Unmet is never auto-approved under AUTO_CFG=true — a missing prerequisite is a fact the executor cannot establish on its own. - gsd-planner emits <precondition> in exactly the three cases the issue mandates: user_setup consumption, prior-phase artifact dependency, and env-var/runtime-config dependency. - Tests cover met, unmet, and absent preconditions plus the additive- validator guarantee. Files: - gsd-core/references/planner-preconditions.md (NEW): full emission rules, the three cases with worked examples, format guidance, anti-patterns, the contract triad mapping, and the executor assertion contract. Progressive disclosure. - agents/gsd-planner.md: slim <precondition> note in Task Anatomy with @-reference to the new file. To stay under the 49152-char agent-file cap (27-char headroom before this change), the inline <comment_text_discipline> and <region_scoped_negative_gate> summaries are compressed to one-line pointers — their full rules already live in planner-antipatterns.md, so no content is lost. - agents/gsd-executor.md: new step 0 'Precondition check' in the execute_tasks loop, before the type dispatch, routing unmet through checkpoint_return_format. - docs/reference/plan-md.md: new Preconditions section in the schema reference, with the canonical example and the three emission cases. - CONTEXT.md: Precondition glossary entry as a sibling of Tracer Bullet. - docs/INVENTORY.md + INVENTORY-MANIFEST.json: row for the new references/planner-preconditions.md (regen via gen-inventory-manifest). - tests/precondition-element.test.cjs: failing-first tests covering schema docs, planner emission contract, executor assertion contract, reference-file presence + the three cases, behavioral additive- validator guarantee, and a parity assertion (DEFECT.GENERATIVE-FIX- DIVERGENCE guard). - .changeset/quick-hawks-bark.md: Added fragment. Companion to #1945 (tracer bullets). * chore(#1949): regen agent-size baseline + install-tree goldens Documented baseline regenerations required by the feat(#1949) prose changes (RULESET.AGENT_SIZE_BUDGET + golden-install-parity): - npm run size:baseline — locks in the new gsd-executor.md size (+1050 bytes: the precondition-check step 0 block). gsd-planner.md is net smaller (-142 bytes: compressed two inline summary blocks whose full rules already lived in planner-antipatterns.md to make room for the slim <precondition> pointer). No hard-cap breach. - npm run gen:golden — pick up the new references/planner-preconditions.md + the two changed agent files across all 18 runtime install trees. Both regens are CI-mandated after intentional agent/reference changes; see CLAUDE.md 'RULESET.AGENT_SIZE_BUDGET' and the comments in tests/golden-install-parity.test.cjs. * fix(#1949): bound <precondition> checks to read-only (security review) Apply the security-review finding (LOW, isolated /security-review subagent): the executor's 'run the cheapest check' phrasing for a plan-author-controlled prose line was broader than ideal — a hostile plan author could craft a <precondition> whose 'cheapest check' is side-effecting (curl to an attacker host under the guise of verification, rm -rf before checking, secret emission). The risk is inherited from GSD's existing plan-trust model (<verify>, <action>, <done> already direct the executor to run arbitrary shell), so <precondition> does not materially expand it. But the new prose actively directs execution ('run the check') rather than passively consuming the element, so the bound is worth making explicit. Tightened across all four surfaces that describe the check shape: - agents/gsd-executor.md step 0: 'Verify with read-only checks only — file existence, env var presence (no value output), idempotent GET /health-style pings. Do NOT run commands with side effects (writes, network POSTs, secret emission) as the check; if a side-effecting check seems required, halt and surface via checkpoint instead.' - gsd-core/references/planner-preconditions.md Format section: same bound, plus the halt-and-surface escape hatch. - docs/reference/plan-md.md Preconditions section: mirrored. - CONTEXT.md Precondition glossary entry: mirrored. Regenerated agent-size baseline (executor grew 46186 -> 46440; still under the 49152 cap) and install-tree goldens. * chore(#1949): backfill changeset pr number 2422 Per CONTRIBUTING.md changeset workflow + feature-builder directive Step 8.7: backfill the placeholder pr:0 with the real PR number immediately after gh pr create returns. Avoids the fail_invalid_fragment gate. * fix(#1949): cite [#1949] on allow-test-rule exemption (ADR-456) CI's lint:ci runs lint-allow-test-rule-refs which per ADR-456 requires every // allow-test-rule: exemption on a NEW test file to carry an issue reference (#NNN or URL). My earlier push omitted it. Local 'npm run lint' (eslint) does NOT run this check — only 'npm run lint:ci' does. CLAUDE.md explicitly warns: 'lint:ci ≠ lint — CI runs lint:ci; a local pass is not the gate.' I should have run lint:ci before pushing; correcting now. Pattern matches the companion feature's test file: tests/tracer-bullet.test.cjs:1 // allow-test-rule: source-text-is-the-product [#1945] |
||
|
|
8d2f8bcb23 |
fix(#2388): gate shared requirement completion on sibling plans, revert on gaps (#2424)
* fix(#2388): gate shared-ID requirement marking and revert on gaps_found Adds requirements.ready-ids (execute-plan.md's update_requirements step) so a requirement ID declared by multiple plans in a phase only marks Complete once every declaring plan has produced a SUMMARY.md, and requirements.revert-phase (execute-phase.md's gaps_found branch) so a gaps_found verdict reverts the phase's own prematurely-Complete IDs before the gap report renders. Single-plan IDs still mark immediately. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2388): regenerate fixtures + lint gate-prep * fix(#2388): repair failing tests after gate verification * chore(#2388): add changeset (#2424) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dd5a2211c9 |
enhance(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) (#2416)
* test(#1964): add failing-first semantic-recall contract tests Epic #1957 Phase 3C (final). Source-text-is-the-product contract tests: semantic recall via MemPalace (top-k meaning-similar prior resolutions, catches same-root-cause/different-wording cases), indexing resolved sessions at archive, graceful degradation to keyword matching when MemPalace is absent, knowledge-base.md stays the durable plain-text source of truth, agent Phase 0 / Matching Logic is semantic-first (the stale 'keyword overlap, not semantic similarity' claim must go), and no new embedding/vector infra (reuse MemPalace). Failing-first: reference, the Matching Logic reframe, the Phase 0 consolidation, and the archive indexing step do not yet exist. * feat(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) Epic #1957 Phase 3C (FINAL). Replaces keyword-overlap matching with semantic recall: at Phase 0 the debugger queries MemPalace with the current symptoms and surfaces the top-k meaning-similar prior resolutions, catching the same-root-cause/different-wording cases keyword overlap missed (the self-noted 'keyword overlap, not semantic similarity' limitation). Resolved sessions are indexed into MemPalace at archive (symptoms + root_cause(s) + fix + recurrence guard). knowledge-base.md remains the durable plain-text source of truth; when MemPalace is absent the debugger falls back to keyword-overlap matching (logged, never a silent skip). No new embedding/vector infrastructure — MemPalace is reused. Size-neutral agent edits: the Matching Logic section reframed (keyword-only -> semantic-first + keyword-fallback + @-include); Phase 0's three keyword bullets consolidated into one semantic-first bullet; one MemPalace-indexing step added at archive. Agent at 57222 B (122 B headroom — final phase). Full rules in gsd-core/references/debugger-semantic-recall.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1964): address orthogonal review (invocation mechanism, index Resolution-not-symptoms + redaction, fallback detail) - HIGH: the 'query MemPalace' instruction was WHAT-level only; the agent has no MCP tools. Added an Invocation section naming the Bash CLI (mempalace search --wing <wing>) + MCP-when-registered + wing resolution (config.mempalace.wing -> project_code -> project dir), matching every other MemPalace integration. Without this the feature silently degraded to keyword matching even when MemPalace was present. - MEDIUM (security x2): index the agent-authored Resolution summary (root_cause + fix + recurrence_guard), NOT raw user-supplied Symptoms — excludes attacker-controlled prose from the cross-session index AND reduces secret/PII leakage. Redact secret-shaped values before indexing. Stated the write order (KB append + commit MUST succeed before indexing). - LOW: restored 'identifiers' + 'case-insensitive' to the keyword fallback; added a test asserting the fallback mechanics survived the Phase 0 consolidation (Error patterns field, 2+ token overlap, identifiers, case-insensitive). * chore(#1964): ratchet agent-size baseline downward (leaner archive bullet shrank gsd-debugger.md 57222->57197) * chore(#1964): backfill changeset pr number (PR #2416) |
||
|
|
c67f301867 |
feat(#1963): emit blameless-postmortem Prevention block at resolution (#2410)
* test(#1963): add failing-first prevention/postmortem contract tests Epic #1957 Phase 3B. Source-text-is-the-product contract tests: blameless 5-Whys that BRANCHES per Phase 2A RCA (not a single-cause chain; treats agent error as 'why was that possible?'), the 'why wasn't this caught?' question, the recurrence-guard taxonomy (regression test / assertion / lint rule / KB pattern), the KB-entry why_not_caught + recurrence_guard fields with backward compat, the session-manager prevention summary line, and the Zawinski scope-boundary (a block, not a subsystem). Failing-first: reference, archive_session edit, KB schema extension, and session-manager summary do not yet exist. * feat(#1963): emit blameless-postmortem Prevention block at resolution Epic #1957 Phase 3B. At archive_session the debugger now produces a Prevention block with three blame-free components: a branching 5-Whys causal chain (branches per Phase 2A RCA, not a single chain; 'agent error' prompts 'why was that possible?', never blame), a 'why wasn't this caught?' answer naming the missed gate (test/typecheck/lint/review/verify), and a concrete recurrence guard (regression test / assertion / lint rule / KB pattern). The knowledge-base entry gains two structured fields (why_not_caught + recurrence_guard) so future Phase-0 recall surfaces the prior prevention, not just the prior fix. Additive: old entries without the fields still load. The session-manager compact summary surfaces a one-line prevention summary. Full rules extracted to gsd-core/references/debugger-prevention.md (slim archive_session step + 2 KB fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1963): address orthogonal review (CRITICAL append-template drift + Phase-0 consumption + parity test) - CRITICAL: the archive_session KB append template omitted Why not caught + Recurrence guard (only the Entry Format had them) — the feature's core deliverable silently did not happen. Added both fields to the append template the agent actually follows (nearest-instruction wins). - HIGH: Phase 0 (KB read) only surfaced root_cause + fix; the new fields were dead data. Extended the Phase 0 Evidence line to consume why_not_caught + recurrence_guard when present (absent on old entries — backward compat holds). - MEDIUM: added a cross-section parity test (every Entry-Format field must also appear in the append template — the guard that would have caught the Critical) + a Phase-0-consumption assertion. - MEDIUM: the 'branches per Phase 2A' claim is now wired — reuses reasoning_checkpoint.candidate_causes across the four categories. - MEDIUM: recurrence-guard taxonomy gains type refinement + config-default change; LOW: added 'build' gate to both surfaces for parity. - NIT: compact-summary fallback shape ('no gate existed'); verify the guard artifact exists before recording it. * test(#1963): anchor Phase-0 consumption test on the specific heading The regex /Phase 0[\s\S]{0,1200}/ matched the first 'Phase 0' in the file (in knowledge_base_protocol prose), not the Phase 0 block in investigation_loop. Anchor on '**Phase 0: Check knowledge base**' and widen to 1500 chars. * chore(#1963): backfill changeset pr number (PR #2410) |
||
|
|
36a311c5bb |
enhance(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) (#2409)
* test(#1962): add failing-first repro-hardening contract tests Epic #1957 Phase 3A. Source-text-is-the-product contract tests: PBT shrinking (fast-check/Hypothesis, minimized seed, manual-minimization degradation), the four oracle types (specified/derived/metamorphic/implicit with implicit flagged weakest), boundary neighbors (off-by-one/min-max/empty-singleton tied to the equivalence class), oracle_type in DEBUG Resolution, and the Phase 1A tie-in (minimized seed + real oracle => the mutation guardrail bites). Failing-first: reference, agent cross-refs, and template field do not yet exist. * feat(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) Epic #1957 Phase 3A. Extends Minimal Reproduction (shrinking) and Test-First Debugging (oracle classification + boundary neighbors): - Shrinking: wrap an input-space failing input in a property (fast-check JS/TS, Hypothesis Python) and store the MINIMIZED counterexample as the regression seed; degrade to manual minimization when no PBT framework is present. - Oracle classification: state specified / derived (contract/model) / metamorphic / implicit (crash, weakest) before writing the assertion; record under Resolution.oracle_type; never default to implicit silently. - Boundary neighbors: off-by-one, min/max, empty/singleton around the fixed defect's equivalence class. Together they turn the regression test into a root-cause check — what the Phase 1A mutation guardrail needs to bite. Full rules extracted to gsd-core/references/ debugger-repro-hardening.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1962): address orthogonal review (bounding, provenance, oracle scope, sufficient-triple) - HIGH: added a 'Bound the property/shrink run' section (60s timeout, degrade- to-manual on timeout, do-not-raise-default-run-limits, argv-not-shell) — the gauntlet violation the sibling references already honored. - Medium: test-provenance caveat (the failing input often comes from the bug report — author the generator from a sanitized description, cross-ref debugger-fix-acceptance.md). - Medium: oracle scope note — the 4 types cover deterministic bugs; non- deterministic failures re-route to stability-stress per bug-taxonomy. - Medium: Phase 1A tie-in corrected — seed+oracle is necessary not sufficient; boundary neighbors close the adjacent-input escape; the sufficient triple is seed+oracle+neighbors. - Low: preserve the original noisy repro as a secondary reference; operationalize 'equivalence class' (the predicate the fix draws). Nit: degradation reworded. * chore(#1962): backfill changeset pr number (PR #2409) --------- Co-authored-by: sim <sim@local> |
||
|
|
6baa2a8182 |
feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger (#2407)
* test(#1961): add failing-first bug-taxonomy routing contract tests Epic #1957 Phase 2B. Source-text-is-the-product contract tests (3 taxonomy classes, explicit class->technique routing table, Bohrbug->repro+SBFL+bisect, Heisenbug->record-replay/stability+SKIP-SBFL, Concurrency->atomicity/order/ deadlock checklist, bug_class in DEBUG Current Focus, supersede-not-append) plus a routing-table specification object pinning the documented decisions (SBFL forbidden on Heisenbug is the load-bearing 1B/2B seam). Failing-first: reference, Phase 1.75, and routing-table reframe do not yet exist. * feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger Epic #1957 Phase 2B (reliability-critical). Adds Phase 1.75: classify the failure as Bohrbug / Heisenbug-Mandelbug / Concurrency, then route the investigation technique via an explicit class->technique table (Kernighan: no opaque heuristic). Bohrbug -> reproduction + SBFL (Phase 1.25) + git bisect; Heisenbug/Mandelbug -> record-replay (rr) + stability-stress + statistical sampling, with SBFL explicitly SKIPPED (a flaky spectrum poisons the Ochiai ranking — the load-bearing 1B/2B seam); Concurrency -> the atomicity/order/deadlock checklist first. Reframes (supersedes, not appends — Zawinski) the flat 'Technique Selection by situation' table into a class-routed table; the 11 techniques remain as routed targets. bug_class recorded in Current Focus (DEBUG template); common-bug- patterns catalog cross-referenced to the taxonomy. Full rules extracted to gsd-core/references/debugger-bug-taxonomy.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1961): address orthogonal review (phase-name drift, General lane, revoke framing, row-scoped tests, bounding) - HIGH: reference said 'Phase 1B' (epic shorthand); corrected to the deployed 'Phase 1.25' (matches the agent + SBFL reference). - HIGH: 6 of 11 techniques (Rubber duck, Delta, Working backwards, Differential, Comment-out, Follow-the-indirection) were orphaned by the situation-table reframe. Added a 'General (any class, situation-cued)' lane to BOTH the reference routing table and the agent's Technique Selection table that re-homes them — supersede-not-append now holds. - MEDIUM: the SBFL-skip is structurally retroactive (Phase 1.25 runs before Phase 1.75 classification), so reframed the table column from 'Do NOT use' to 'Revoke if already run' + an explicit 'retroactive revocation, not proactive skip' note stating the ordering honestly. - MEDIUM: contract tests are now row-scoped (parse the table by class, assert per-row) instead of presence-only; added a guard that the previously- orphaned techniques now have a General-lane route. - LOW: pinned the canonical bug_class value form (lowercase-kebab: bohrbug|heisenbug-mandelbug|concurrency; prose may use title-case). - NIT: added a 'Bound the Heisenbug-chase runs' note (rr/stability/sampling timeouts) per the unbounded-subprocess gauntlet. * chore(#1961): backfill changeset pr number (PR #2407) |
||
|
|
f8b16d1874 |
enhance(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger (#2405)
* test(#1960): add failing-first RCA-branching contract + schema-invariant tests Epic #1957 Phase 2A. Source-text-is-the-product contract tests (fishbone >=2 categories, AND-gate, multi-cause root_cause, backward compat, reasoning checkpoint candidate_causes+and_gate fields, debugger-philosophy single-cause note, DEBUG template) plus behavioral schema-invariant checks on two fixtures: two contributing causes (AND-gate yes) -> both recorded; single-cause (AND-gate no) -> one root_cause, identical to today. Failing-first: reference, agent edits, and template note do not yet exist. * feat(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger Epic #1957 Phase 2A. Guards against 5-Whys single-cause bias: before committing root_cause, the debugger enumerates candidate causes across >=2 Ishikawa categories (code/config/environment/data) and explicitly answers an AND-gate question. When the AND-gate fires, every contributing cause is recorded, so a multi-cause fix no longer recurs via the unaddressed second cause. Resolution.root_cause may hold one OR a small set (additive; single-cause sessions are byte-identical to today). The Structured Reasoning Checkpoint gains candidate_causes + and_gate fields; debugger-philosophy.md adds the single-cause-bias trap. Full rules extracted to gsd-core/references/debugger-rca-branching.md (slim Phase 2 routing + 2 checkpoint fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1960): address orthogonal review (AND-gate self-consistency, parity guard, narrowed claim, ripples) - Reference: the collapse rule now enforces AND-gate self-consistency — and_gate=yes with a single confirmed cause is flagged as incomplete (return to Phase 3); a race/timing note clarifies such bugs bridge categories; the 'byte-identical' backward-compat claim narrowed to 'root_cause shape unchanged; reasoning_checkpoint gains 2 fields in every session'. - DEBUG.md: stale 'five-field' mirror prose -> seven-field (parallel-surface drift the reviewer flagged); new debug-session-management parity test pins the field-count claim to the gsd-debugger.md YAML keys (CRLF-safe). - Scalar-assuming consumers of set-valued root_cause updated: session-manager compact summaries (319/332), diagnose-only return (1062), archive entry (1216), ROOT CAUSE FOUND return (1322). - Test: added the AND-gate-yes/single-cause invariant + fixture; rephrased the fixture describe block honestly as a schema-invariant specification. - Phase 2 bullet phrasing clarified ('at hypothesis formation, before the Phase 4 commit'). * test(#1960): parity regex accepts word-form count ('seven-field' or '7-field') * test(#1960): parity regex counts array-valued YAML keys (no inline value) * chore(#1960): backfill changeset pr number (PR #2405) |
||
|
|
56a5c6404c |
feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger (#2403)
* test(#1959): add failing-first SBFL contract + Ochiai correctness tests Epic #1957 Phase 1B. Source-text-is-the-product contract tests (Ochiai formula documented, Tarantula fallback, top-N seeding, no-coverage skip logged, ranking->Evidence, Bohrbug gating) plus a behavioral Ochiai formula-correctness section: bound [0,1], max-score invariant, a known-fault fixture proving the fault ranks #1 (criterion 2), clean degradation on zero failing tests, and two fast-check properties. Failing-first: reference file and agent routing do not yet exist. * feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger Epic #1957 Phase 1B. When a runnable test suite with per-test coverage exists (>=1 failing AND >=1 passing test), the debugger computes an Ochiai suspiciousness ranking over the coverage spectrum and seeds the top-N suspicious locations into Evidence as first-class hypothesis candidates, narrowing the search space deterministically before LLM reasoning. Tarantula documented as fallback. Degrades cleanly (logged, never silent) when there is no test suite, no failing tests, or no per-test coverage, and is explicitly not trusted on flaky/Heisenbug spectra (pairs with Phase 2B bug-taxonomy). Full rules extracted to gsd-core/references/debugger-sbfl.md (slim Phase 1.25 routing kept in the agent to respect the size cap). No new coverage framework — reuses the project's existing test/coverage runner. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * test(#1959): bound property generators to valid coverage counts The [0,1] property generated failedExec independently of totalFailed, but Ochiai's score is only bounded by 1 under the coverage invariant failedExec <= totalFailed (a failing test that executed s is one of the totalFailed failing tests). Out-of-domain inputs (failedExec=100, totalFailed=5) make the formula correctly return >1. Bound failedExec by totalFailed via fc.chain so the property tests the real domain. Also cleaned up the ranking property (removed dead code). * fix(#1959): address orthogonal review (monotonicity property, degradation row, coverage bounding) - Replace vacuous ranking property (true-by-sort-construction) with a non-trivial monotonicity property: holding totalFailed + passedExec fixed, ochiai is non-decreasing in failedExec. An inverted formula would fail it. - Add the missing 'no passing tests' degradation row (preconditions require >=1 passing test; Tarantula would divide by totalPassed=0). - Bound the coverage subprocess (CLAUDE.md gauntlet): cap the coverage run, degrade-to-skip on timeout, never hang the debug session. - Reword 'discard the ranking' -> 'mark the Evidence entry as revoked (do not delete)' per Kernighan auditability. * test(#1959): bound monotonicity-property generator to valid coverage (failedExecA <= totalFailed) * chore(#1959): backfill changeset pr number (PR #2403) |
||
|
|
5e52350736 |
feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger (#2396)
* test(#1958): add failing-first guardrail contract tests Epic #1957 Phase 1A. Adds source-text-is-the-product tests asserting the 5-signal fix-acceptance guardrail contract (target test, mutation check, no-op/deletion detector, adjacent tests, revert-and-reconfirm), graceful degradation, FIX REJECTED BY GUARDRAIL return path, per-signal debug-file recording, and subprocess bounding. Failing-first: reference file and agent sections do not yet exist. * feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger Epic #1957 Phase 1A. Prevents accepting a fix that merely greens the test (Goodhart defense / APR overfitting). Adds a 5-signal gate run before fix acceptance: target test, mutation check (Stryker), no-op/behavior-deleting detector, adjacent/held-out tests, revert-and-reconfirm. Degrades gracefully when Stryker or a test suite is absent (each skip logged, never a silent pass), records per-signal results under Resolution.verification, and returns a FIX REJECTED BY GUARDRAIL outcome the session-manager surfaces for revise / accept-as-debt / abandon. Full rules extracted to gsd-core/references/debugger-fix-acceptance.md (slim routing kept in the agent to respect the agent-size cap). Debug template + INVENTORY + manifest + agent-size baseline + AGENTS.md updated. * test(#1958): correct newline-tolerant assertion + regen install-parity goldens The revert-and-reconfirm assertion collapsed whitespace before matching so markdown line-wrapping does not break it. Regenerated the golden-install-parity and install-tree fixtures (npm run gen:golden) to absorb the intentional gsd-debugger.md / gsd-debug-session-manager.md / DEBUG.md / new reference-file changes to the installed artifact tree. * fix(#1958): tighten guardrail per orthogonal review Addresses the isolated reviewer's findings: - signal 5 now states its recorded-repro dependency and routes the no-repro case to the degradation row; revert mechanism specified (git stash / git revert -n); minimality flag tied to diff structure, not revert-ability. - bounded-subprocesses section now bounds the git subprocess (5-30s) too, requires argv-array argument passing, and scopes Stryker to the driving regression test (a mutant killed only by a non-driving test is a finding). - new test-provenance (security) clause: the driving test must be agent-authored; bug-report repro scripts are DATA, never executed verbatim. - tightened 3 contract assertions to bind to specific clauses (guardrail_verdict field, deletion-reject-unless-RCA, 60s+git bounding). - Goodhart framing softened to 'partially-independent'; DEBUG.md template verification field notes the nested map shape. * chore(#1958): backfill changeset pr number (PR #2396) * fix(#1958): add issue ref to allow-test-rule annotation (ADR-456) CI lint-allow-test-rule-refs requires every allow-test-rule exemption to carry a 'see #NNN' issue ref per ADR-456. The new test file's annotation lacked it; this adds (see #1958). |
||
|
|
b302f53ee6 |
refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) (#2370)
* refactor(#2368): extract capability arm to capability-command-router (ADR-2346 P2) Behavior-preserving relocation of the 706-line case 'capability': arm from gsd-tools.cjs into a new hand-authored bin/lib/capability-command-router.cjs (sibling of ensure-runtime-build.cjs). The 15 bin/-relative require paths are rewritten to sibling-relative (correct for bin/lib/). dispatchHostCommand is now async (capability's install/upgrade/consent ops await the lifecycle); sync routers (state/phase/…) pass through await unchanged. case 'capability': removed; capability dispatches via HOST_COMMAND_ROUTERS. Validated by the existing capability-lifecycle / -consent / -trust / -loader test suites (no logic changed). Golden install-parity fixtures regenerated. Closes #2368 (Slice 1 — relocation). Probe consolidation (capHostVersion→ readHostVersion, capReadStrict dedup) deferred to a follow-up slice. * fix(#2368): add capabilityState/capabilityWriter requires + INVENTORY row The relocated capability arm references capabilityState (cmdCapabilityState, resolveCapabilityRuntimeState) and capabilityWriter (cmdCapabilitySet) — both module-scope requires in gsd-tools.cjs (L288/289) that the initial closure-dep scan missed. Added as sibling requires to capability-command-router.cjs. Also adds the new cli module to docs/INVENTORY.md + regenerates the manifest. * fix(#2368): correct capHostVersion __dirname depth for bin/lib/ relocation capHostVersion's VERSION/package.json paths were bin/-relative ('..' and '..','..'); on relocation to bin/lib/ they resolved one level too deep, so capHostVersion returned 0.0.0 and capability install failed the engines.gsd gate (#1920). Added one more '..' to each (now resolves gsd-core/VERSION and repo-root package.json correctly). * test(#2368): drop capability from the invocation loop (async/FS vs /fake/cwd) capability is async and does FS/config reads, so invoking it against the unit test's /fake/cwd is fragile. The 6 sync Tier-1 routers stay in the invocation loop; capability is covered by the non-invoking registry- ownership assertion + the dedicated capability-* test suites. * chore: retrigger CI (no-changelog label now present) |
||
|
|
2cbf186420 |
chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 (#2251)
* chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 Phase 3 of epic #2143 (ADR-2143 §5/§6). The three target bugs (#2140, #2112, #2118) were already fixed tactically on next; this introduces the reusable structural contracts and rewires the primary #2140 site onto them. - src/write-set.cts (new): the parse `Result<T> = {ok,value|reason}` (§5) and the per-surface write-set (`WriteOutcome {surface, applied, requirement?}`, `WriteSet`, `writeSetComplete`) (§6). markdown-table.cts now imports + re-exports `Result` from here (single source; distinct from command-routing-hub's Result). - requirements mark-complete (src/milestone.cts): returns a PER-REQUIREMENT, per-surface write-set; `write_set_complete` is true only if every surface of every requirement applied — structurally forbidding the #2140 OR-into-one-flag masking, including across a multi-ID batch (adversarial-review regression). Pre-existing output fields unchanged (behaviour-preserving; #2140 already fixed). - deriveProgressFromRoadmap (src/phase-lifecycle.cts): removed the vestigial null-swallowing try/catch (findTableWithColumns never throws) — ADR §5 no-swallow; RoadmapProgress return contract unchanged. - commit --files (#2112) and milestone complete --dry-run (#2118) left as-is (single-surface commit / pre-mutation preview — not genuine multi-surface writes). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2244): backfill changeset PR number (#2251) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d49ac81306 |
chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 (#2248)
* chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 Phase 1 of epic #2143 (ADR-2143): consolidate markdown table parsing onto a canonical seam and migrate the pilot reader. - Add src/markdown-table.cts: parseMarkdownTable (GFM tables -> typed {columns, rows} addressed by column NAME; ragged rows are typed parse errors, not silent), a single-source TABLE_SCHEMAS registry (RoadmapProgress / RequirementsTraceability / QuickTasks / Security, with variants under one id), matchTableSchema, and findTableBySchema. Result<T> is scoped to this seam (distinct from the dispatch Result). - Migrate deriveProgressFromRoadmap (src/phase-lifecycle.cts) off the position-anchored regex to name-based resolution via the seam — fixes #2137 (the 5-column milestone-grouped Progress table previously returned all-null). - Add a schema-backed `gsd-tools quick-tasks-append` subcommand and route fast.md's log_to_state through it, retiring the inline `awk NF-2` column arithmetic — fixes #2133 (addresses #2012, #2119). Cell values are escaped (| and newlines) and the STATE.md read-modify-write is atomic under readModifyWriteStateMd (lost-update race, cf. #500/#905/#1230). - Writer/reader/template parity test guards TABLE_SCHEMAS against drift (ADR-2143 §3 Generative-Fix-Divergence). Registration: .gitignore, eslint.config.mjs, docs/INVENTORY.md + INVENTORY-MANIFEST.json, CONTEXT.md glossary, docs/CLI-TOOLS.md. Behaviour-preserving for the canonical 4-column Progress table; the named bugs are driven fail-first. Extend-never-mutate (ADR-2143 §2). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2242): backfill changeset PR number (#2248) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): escape backslash before pipe in markdown-table cell escaping CodeQL js/incomplete-sanitization (high): escapeCell escaped | -> \| but not the backslash itself. Now escapes \ -> \\ before | -> \|, and splitTableRow unescapes both \\ -> \ and \| -> | symmetrically so cell values (incl. literal backslashes) round-trip exactly. Added backslash round-trip tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): read ROADMAP Progress table by column name — supersede #2168 ad-hoc scan Rebase reconciliation with #2168 (the tactical #2137 fix that marked itself "pending #2143"). deriveProgressFromRoadmap now resolves the Progress table via a new seam helper findTableWithColumns (first table whose header is a superset of Phase/Plans Complete/Status/Completed, any order, extra columns ignored) and reads cells by NAME — order/injection-invariant per ADR-2143 §3 — instead of the exact TABLE_SCHEMAS match. This satisfies #2168's column-invariance property test while staying seam-based and preserving its `## Progress` scoping (#2012/#1445). Ragged Progress tables now resolve to null (ADR-2143 fail-loud); updated the stale state.test.cjs assertion that predated the Phase-1 migration. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bd613566cb |
feat(#2100): drive Windsurf through the EoS descriptor + wire Cascade's blocking hook bus (ADR-1239)
Fold all 10 residual isWindsurf branches in bin/install.js onto descriptor-driven hostBehaviors (byte-parity — no fold changes any install output): - 2 dead destructures dropped (uninstall, finishInstall); the dead `else if (isWindsurf)` legacy agent-loop arm removed (windsurf ∈ _DESCRIPTOR_AGENTS_RUNTIMES → unreachable). - skipSharedHooksInstall:true folds the two `!isWindsurf` shared-hooks exclusions. - legacyDevinSkillsCleanup:true folds the `.devin`→`.windsurf` one-time cleanup gate. - installsCommandBodiesForWorkflowDelegation:true folds the #1629 command-body copy (workflow-delegation target — load-bearing; local-install verified intact). - verificationStyle:"windsurf-workflows" folds the workflow-count report. - Corrected stale _LEGACY_SCAN_SUBDIR_NAMES + hooks-json manifest comments (cursor + windsurf). Zero live runtime==='windsurf'/isWindsurf branches remain across bin/install.js, install-engine.cts, surface.cts, runtime-artifact-conversion.cts (AC2 guard scans all four). UPGRADE (Cascade hook bus): wire GSD's write/command safety guards into Windsurf's native hook bus. New hooksSurface 'windsurf-hooks-json' (VALID_HOOKS_SURFACES 7→8, GATE A profile-marker-only allowlist, the HooksSurface union) + writeWindsurfHooksJson (Cursor-templated, Cascade's flat {hooks:{<event>:[{command}]}} shape) writing .windsurf/hooks.json with two BLOCKING pre-hooks: - pre_write_code → gsd-windsurf-pre-write.js: blocks writes to a file outside the active git worktree / into .git internals. - pre_run_command → gsd-windsurf-pre-command.js: conservative destructive-command deny-list (rm -rf of root/home incl. sudo/env/path-prefixed forms; fork bombs; force-push refspec forms — HEAD:main, +main, --force/-f — to main/master/next). Both use Cascade's protocol (stdin JSON, exit 2 + stderr to block, exit 0 to allow, fail-open on error/timeout). Tokenize-based classifier (no catastrophic-backtracking regex; 4096-char cap) with the fail-closed false-positives fixed post-review. The 4 advisory GSD guards + pre_mcp_tool_use + 5 post_* logging events are deliberately NOT wired: Cascade has no context-injection channel for advisory hooks and GSD has no MCP guard — porting them would be non-functional padding (documented; codebuddy #2098 / copilot #2099 faithful-subset precedent). extendedHookEvents stays []. Golden: the 2 guard scripts ship in the shared hook bundle (HOOKS_TO_COPY + the shared managed-hooks-registry), exactly like cursor's 6 gsd-cursor-*.js scripts — so the 8 shared-bundle runtimes' fixtures gain the 2 inert windsurf scripts + the registry hash (functionally inert for non-windsurf; the established cursor pattern). No install-output change beyond that (the folds are byte-parity; skip-bundle runtimes untouched). New scripts registered in managed-hooks-registry + build-hooks + INVENTORY. Tests: declarative-reference- windsurf (adapter/axes/fail-closed + AC2 guard) + windsurf-hooks-bridge (live exit-2 blocking + allow/fail-open + ReDoS-bound + writer/reconcile/remove idempotency); VALID_HOOKS_SURFACES pin updated to 8. Matrix hookBus delta + changeset (Changed). capability-registry regenerated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c1756d0cd5 |
chore(#1867): register ui-consideration probe + regen install cascade (SHIP-01)
Ship-safe registration + regenerated snapshots for the #1867 UI-consideration probe (Phase 3, SHIP-01): - CONTEXT.md: PROBE.ui.{verification,axis,seam} predicates + ui-consideration -probe added to PROBE.family (machine-canon for the 3rd adapter, MIXED axis). - agents/gsd-ui-{researcher,checker}.md: one @-include of references/ui-consideration-probe.md each (both under the 24576 agent cap). - docs/INVENTORY.md + INVENTORY-MANIFEST.json: register the reference doc and the compiled ui-consideration-probe.cjs (inventory-manifest-sync green). - tests/fixtures/golden-install-parity/*.json (16 runtimes): recaptured against a clean full build — folds in the deferred Phase-1 (ref doc, plan-phase lift) and Phase-2 (ui-phase step, UI-SPEC section) install-surface changes. - tests/agent-size-baseline.json: ratcheted the two grown UI agents. - .changeset/vivid-orcas-chatter.md: type Added (pr updated at PR-open). Inventory/golden/size gates green; lint:ci + lint:docs + lint:changeset green. The plan-phase.md PRE_PHASE6 ceiling stays RED pending #1852 (unchanged). Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS |
||
|
|
303a796579 | docs(changeset): #2089 cursor host-integration migration + golden fixture | ||
|
|
015c3a7fda |
fix(#2071): extract install-time effort resolvers so effort sync stops requiring the un-shipped bin/install.js
`gsd-tools effort sync` crashed in every installed runtime (e.g. ~/.claude/gsd-core/) with `Cannot find module '../../../bin/install.js'`: cmdEffortSync (src/commands.cts) required the package-root bin/install.js for its install-time effort resolvers, but the installer only copies the gsd-core/ subtree into a runtime home — bin/install.js is never present there. So `effort` config changes silently never reached installed agents without a full reinstall (exactly the gap #488 was meant to close). 4th instance of the recurring "runtime code under gsd-core/ requires a file outside the shipped subtree via ../../../" anti-pattern (#1223/#1920/#1383 were the prior three, all already mitigated). Fix (ADR-457 direction — extract, single source): move readGsdEffectiveEffortConfig + resolveInstallTimeEffort (with their _getGsdEffortCatalog + _readGsdConfigFile helpers) out of the hand-authored bin/install.js into a new src/install-effort-resolver.cts that compiles into the shipped gsd-core/bin/lib/install-effort-resolver.cjs. commands.cts now requires it as a sibling (`./install-effort-resolver.cjs`) — always present in the installed tree — instead of `../../../bin/install.js`. bin/install.js imports the same four symbols back from the new module (it still calls them + re-exports them), so there is one source of truth and no duplication/drift. The lazy manifest read is repointed from the package-root layout (`.., gsd-core, bin, shared`) to the bin/lib layout (`.., shared`). Scope note: this is one of four instances of the anti-pattern; the other three are already shipped/guarded. A build-time guard rejecting new cross-boundary requires whose target isn't in the installer copy manifest (to prevent instance #5) is recommended on the issue but kept out of this fix. Tests: tests/effort-sync-installed-runtime.test.cjs does a real minimal install into a temp home (the golden-parity helper) and runs the issue's exact repro (`gsd-tools effort sync --config-dir <temp>`), asserting no MODULE_NOT_FOUND for bin/install.js. Fail-first verified: against pristine next the same test throws `Cannot find module '../../../bin/install.js'` at cmdEffortSync; post-fix it syncs cleanly. New module registered in .gitignore (ADR-457), eslint ignores, docs/INVENTORY.md + INVENTORY-MANIFEST.json. bin/install.js is not shipped and the new module is under bin/lib (excluded from golden parity), so no golden fixtures change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6addeccd19 |
feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates an external API/SDK/service can no longer seal without a decided coverage matrix. - src/api-coverage.cts: deterministic detector (compound verb+noun signal + <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix parse/validate/render with field-length caps. - check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as a token under .planning/phases/ only (traversal-neutralized); validates COVERAGE.md or blocks iff a strong integration signal is detected and no matrix exists; fail-closed when phases tree exists but phase unresolvable. - capabilities/ai-integration: workflow.api_coverage_gate config key (default true), plan:pre contribution, blocking verify:pre gate. Data-driven. - gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch. - Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e. Code+security review findings fixed (stopword FP, scope containment, pipe/cap rejection, prompt-injection message hygiene). - Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs. Closes #1562 |
||
|
|
603593d41d |
fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest defaults to WATCH mode in an interactive TTY — exactly where a user runs `gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never exited and the orchestrator waited indefinitely. Recovery needed the user to manually prompt "something blocking?". Fix — one shared helper + a bounded, surfacing timeout on the test-command gates: - New pure module src/normalize-test-command.cts + `gsd-tools query normalize-test-command` verb: rewrites a resolved command to a best-effort one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`; a package-manager `test` script whose package.json runner is watch-vitest → `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged — never double-flagged). Named `normalize-test-command` (not `test-*`) so the file does not match node --test's default `test-*` discovery glob. - The three gates that HUNG or silently-continued route through that ONE helper and bound execution with `timeout $(config-get workflow.test_gate_timeout)` (new config key, default 600s): the regression gate (extracted to execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen — it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124. - verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint on 124, staying under its frozen 40960-byte tier cap. Security hardening (review): the normalizer only rewrites a runner named as a standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never mangled), is length-capped and uses only linear-time split-based scanning (no super-linear backtracking on an adversarial `workflow.test_command`), and reads package.json only when it is a regular file (never blocks on a FIFO via `--dir`). Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores, inventory manifest/index. All 16 golden-install-parity fixtures + workflow size baseline regenerated for the changed shipped files; bin/lib is excluded from the parity manifest. Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route through the shared helper + configured timeout + exit-124 watch-mode hint; verify-phase asserted as normalize-only/already-bounded). tests/execute-phase-active-flags.test.cjs repointed at the extracted step; tests/planner-language-regression.test.cjs allowlist comment updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ea4378063f | Merge branch 'next' into feat/1820-specless-predicate-rail | ||
|
|
080bacdb4b | Merge branch 'next' into codex/gsd-onboard | ||
|
|
e3262d94d3 |
feat(capabilities): add claude-orchestration capability (Workflow backend) (#1143)
Default-off, BETA, claude-only capability adopting Claude Code's Workflow tool (/effort ultracode, Agent SDK >= v0.3.149) as an optional parallel-execution backend for the GSD loop. Restores the wave parallelism + plan-checker + verifier that #853 forces inline on Claude Code, and folds gsd-ultraplan-phase under one runtime gate. - Pure fail-closed core (src/claude-orchestration.cts): detectWorkflowBackend (gate ladder: enabled -> Claude -> backend != inline -> nested+background host -> valid Agent SDK -> SDK >= floor; every miss degrades to inline) and emitWorkflowScript (waves -> parallel() barriers, plans -> gsd-executor + worktree, files_modified overlap -> separate stages, resumeFromRunId, budget). All interpolated identifiers validated script-safe; briefs JSON-quoted. - claude-orchestration command family (gsd-tools claude-orchestration detect-backend|emit-workflow) for orchestrator invocation. - Two gated loop contributions at wired points (execute:wave:post, plan:post); federated config keys (enabled/execution_backend/min_agent_sdk_version). - ADR-1143 implementation amendment; CONTEXT.md glossary entry; explanation doc. On any runtime lacking the Workflow tool, behaviour is byte-identical to today. closes #1143 |
||
|
|
7ef834cabc | feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them | ||
|
|
4c673e51f3 |
chore(#1990): resync runtime launcher and regenerate artifacts after rebase onto next
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> |
||
|
|
8de2ff9121 |
feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator Third-party capability gates declared via check.predicate were rendered for display but never evaluated (only built-in check.query gates fired; the security capability's gate worked solely via a hard-coded ship.md branch). Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts) that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a bounded sh -c command at the project root (via shell-command-projection.execTool), inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed. Wire a 'check predicate' subcommand into check-command-router.cts and extend the three generic workflow gate-dispatch sites (execute:wave:post, execute:post, plan:post) to route check.predicate gates to the new evaluator. The two-step gate contract (command-failure => onError; block => halt) is unchanged. - src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible - src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags - docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md - tests: 38 unit + integration tests (exit mapping, timeout, interpolation, property-based bijection, malformed-predicate fail-closed, real subprocess e2e) Closes #2008 * docs(#2008): backfill changeset pr number 2011 |
||
|
|
5ae4ea4c84 |
feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) (#1998)
* feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) The async external-job consumer half (#1165) shipped long ago: the core loop reads .planning/async-jobs/<job>.json manifests and treats a non-terminal one as the legal external_job_waiting half-state. The PRODUCER half (#1164) was the remaining unimplemented piece of #1105. This adds the producer as a default-off capability: - capabilities/external-job/ — capability.json (execute:wave:post -> executor, plan:post -> planner contributions, external_job.* config keys, default-off) + fragments teaching runtime-budget classification and externalization. - src/external-job.cts -> gsd-core/bin/lib/external-job.cjs — pure producer module: SLURM state -> manifest-status map (no guessing), manifest build/validate (versioned stability contract), sbatch/squeue/sacct parsers, and a fail-closed manifest writer (refuses a second non-terminal job for a plan_id already in flight; refuses to clobber a malformed manifest). fs/clock seams for deterministic tests. - scripts/slurm-adapter.cjs — operator CLI (submit/poll/show) wrapping bounded sbatch/squeue/sacct subprocesses; surfaces manifest commands for confirmation and never auto-runs them (trust boundary). - tests/external-job.test.cjs — 23 behavioral + fast-check property tests. - docs/reference/long-running-operations.md + docs/how-to/async-external-jobs.md. - CONTEXT.md glossary entry for the External-job Capability. - Regenerated capability-registry.cjs; pruned the now-stale test-file-count allowlist entry (external-job is at the 2-file cap). * chore(#1105): backfill PR number in changeset * fix(#1105): sync capability artifacts + update registry shape-pin tests gsd-test caught that adding the external-job capability requires its dependent artifacts regenerated and its registry-shape drift absorbed: - sync-manifest-versions: stamp 1.7.0-rc.2 into capability.json (was 1.0.0). - gen-capability-matrix --write: regenerate docs/reference/capability-matrix.md. - gen-inventory-manifest --write: regenerate docs/INVENTORY-MANIFEST.json. - check-gap-analysis-plan-post-e2e: plan:post now has 1 contribution (external-job planner fragment) instead of 0. - execute-wave-post-gate-pipeline-e2e: execute:wave:post now has 2 contributions (mempalace + external-job) instead of 1. * fix(#1105): regenerate capability-registry after version stamp sync-manifest-versions re-stamped external-job/capability.json from 1.0.0 to 1.7.0-rc.2 after the last registry regeneration, leaving the committed capability-registry.cjs stale (CI gen-capability-registry --check failed). gsd-test masked this because its setup runs the full 'npm run build' (which regenerates the registry); CI's 'npm test' pretest only runs build:lib. |
||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
1388e7a362 |
feat(#1683): published Host-Integration SDK surface + smoke test — Slice 2 (#1939)
The SDK entry (src/host-integration-sdk.cts) is the single PUBLIC surface a host-plugin author imports: the negotiated schema + classification, the five adapters (declarative/imperative/model/hook/state), and the serialized handshake. Frozen so the public shape cannot be mutated. Everything else in gsd-core stays internal. tests/sdk-smoke.test.cjs imports ONLY from the SDK entry and builds a third-party host-plugin end-to-end (compose adapters + handshake + classify) — proving an external author can wire a host without gsd-core internals (#1683 AC). ESLint-ignore + inventory manifest kept in sync for the new tsc-emitted module. |