c0b2a05d2f310adc0a1f35fd71fbc9f28f4e4977
64 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c0b2a05d2f |
fix(#4594): one canonical dispatch-identity owner — the emitted format and the parser that reads it back (#4693)
* fix(#4594): give dispatch identity one owner for the emitted format and its parser The isolation guards decided whether a run-scoped sentinel applied to a dispatch by regex-scraping model-authored prose. The scrape returned values in a different namespace from the ones the sentinel records, so the comparison could never succeed: sentinel { phase: "03", plan: "03-02-hardening" } <- $PHASE_NUMBER / $plan_id prose "Execute plan 02 of phase 03-auth." scraped { phase: "03-auth.", plan: "02" } <- greedy (\S+), both wrong #4594 reports only the phase half. Measured against a real phase-plan-index run, plans[].id is phase-prefixed, plan-numbered AND slugged, while the prose carries a bare in-phase plan number — so the plan field mismatches too, and the Claude path is dead rather than latent. A fresh sentinel was therefore discarded on every executor dispatch and every legitimate ISOLATION=none degrade was denied, leaving the work unrun. hooks/lib/dispatch-identity.js is now the single owner of both halves. The two prompt-body producers emit a canonical marker carrying the same shell values the sentinel records, so producer and consumer agree by construction. The prose frame stays as a fallback, bounded by the phase-token grammar ADR-2121 owns and deliberately reporting no plan — an absent identifier means "cannot compare" and is safe; a wrong one is a false mismatch and is not. The prose sentence itself is byte-identical: the executor agent reads it too, so the marker is purely additive (Hyrum's Law). An inapplicable sentinel is now named in the guards' deny reason instead of being dropped silently — the silence is why this survived three producers and two consumers unnoticed. Interpolated values come from a sentinel file and from prompt text, so both are length-bounded and stripped of control characters. ADR-4630 locks the seam and maps the epic's three phases. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4594): resolve eight review findings across the dispatch-identity seam Three orthogonal review engines ran on 43418af144 — the code-review skill's Standards and Spec axes, and an isolated adversarial security pass — plus a self-review of the committed diff. Every finding is fixed here; none deferred. F1 (major, reproduced). A keyless or unknown-key-only marker — the literal "[gsd:dispatch]" or "[gsd:dispatch run=..]" — matched the marker grammar and returned source:'marker' with both fields null, suppressing the prose fallback entirely. Any prompt text containing that literal silently disabled identity narrowing, so a fresh sentinel applied to a dispatch it was never scoped to, defeating #3045 SECURITY F2. Prompt text is attacker-influenceable. A marker that yields neither recognized key is no longer a marker: the scan continues to later markers, then later texts, then prose. Forward-compatible tolerance of unknown keys is unchanged. F2/F3 (major). The first cut duplicated sanitizeForReason, describeSentinelDiscard and REASON_INTERPOLATION_MAX_LEN byte-for-byte across both guards — the exact defect class this epic exists to delete, and with no cold-load justification, since both hooks already require hooks/lib/. They now live in hooks/lib/isolation-deny-reason.js, and buildSentinelDiscard lives in isolation-sentinel.js beside the comparison it mirrors, returning the nested {sentinel:{phase,plan}, dispatch:{phase,plan}} shape instead of a bespoke four-field bag that renamed the pairs already flowing through the seam. F4 (hard violation). The visibility test asserted on the deny reason's prose. CONTRIBUTING.md prohibits raw text matching on hook output, which is why every deny carries a stable reason_code. The discard is now a structured sentinel_discarded field on each hook's stdout JSON, and the test asserts that; the sentence stays for the operator but is no longer the contract. F5 (hard violation). The 64-character truncation limit had no boundary coverage. 63/64/65 are now exercised against the single consolidated helper. F6 (minor). sanitizeForReason stripped C0/C1 controls but not U+2028/U+2029 or the bidi overrides, so a crafted value could still reflow or reverse the message. Both classes are stripped, with a test each. F7 (major). The producer/template parity test was vacuous — it rendered a marker and re-parsed its own output, and would have passed with both templates deleted. It now reads the two workflow templates, extracts each marker line, substitutes the measured values and asserts the owner's parser returns them. Proven red by deleting one template's marker line before being proven green. F8 (doc). ADR-4630 and the design notes claimed the marker is guaranteed on the orchestrator-worktree path because that prompt is built in shell. It is not: executor-isolation-dispatch.md:131 says plainly that those are template placeholders, not shell variables, so {plan_id} is model-substituted there too. A false guarantee in a design lock is worse than a stated limit. Both documents now say the marker is model-substituted on both paths and that the prose fallback is the real floor everywhere. The "3 workflow templates" count was also wrong — 3 prose sites across 2 files, 2 of which carry the marker. Refs #4630 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): refresh the compact-content baseline and acknowledge execute-phase.md growth Refs #4630. The dispatch-identity marker and its substitution note grew gsd-core/workflows/execute-phase.md by 525 bytes (91846 -> 92371), which drifts two real-tree guards that lint:ci does not run: - tests/benchmark-compact-content.test.cjs asserts the committed baseline is "up to date"; the split for execute-phase.md moved off 25827 -> 25952 and on 23576 -> 23701, taking its compaction reduction 8.72% -> 8.67%. Baseline regenerated with scripts/benchmark-compact-content.cjs --write. - tests/emitted-attribution.test.cjs requires a growth acknowledgment trailer for any emitted file that grows, keyed on the bare filename. Added below. The growth is two additions and no rewrites: the [gsd:dispatch ...] marker line inside the Agent() prompt's <objective>, and the note telling the orchestrator to substitute {plan_id} with the plan's id verbatim. Both are load-bearing -- the marker is what lets a guard hook match a dispatch to the sentinel the per-plan gate wrote, and without the note the orchestrator has no instruction telling it the value must not be paraphrased. Emitted-Drift-Ack-Growth: execute-phase.md — adds the canonical [gsd:dispatch] identity marker and its {plan_id} substitution note, which the isolation guards compare verbatim against the run-scoped sentinel (#4594) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4594): set changeset fragment pr to 4693 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1e47560e34 |
feat(#4593): add a macOS-specific conformance tier, final phase of epic #4589 (#4607)
test-conformance's macos-latest leg (Phase 2, #4591) has been running the same 546-file, Windows-oriented conformance-tier list as windows-latest -- built from signals like windows-shell-token/windows-env-var that have nothing to do with macOS. Issue #4593 asked for macOS coverage sized to its own evidence-backed surface (zsh dispatch, case-sensitivity, darwin- specific behavior) instead. Issue #4593 was filed before Phase 5 (#4603) existed and referenced updating test-full's macOS legs -- that job is gone. Corrected the issue's body before any code was touched: the "shrink from full replay" half of the original ask was already done by Phase 5; what remained was narrowing the still-Windows-oriented tier macOS was inheriting. Two design assumptions were measured and rejected before accepting a design (documented in docs/adr/4593-macos-conformance-tier-architecture.md): - Reusing the general tier's signals minus its 3 Windows-specific categories barely narrows anything (546 -> 424, 78% retained) -- most files match multiple signals and only need one to survive exclusion. - A standalone CRLF/autocrlf signal, despite the issue naming "CRLF-checkout behavior": even narrowed to /\bCRLF\b|autocrlf/i it hit 143/930 files. Root cause: CRLF is primarily a Windows checkout concern in this codebase (ADR-1703 files it under DEFECT.WINDOWS-TEST- PORTABILITY), so the signal was really re-selecting Windows-relevant files already covered by the general tier, not narrowing macOS specifically. Built 5 new, genuinely macOS-specific signals instead: darwin-literal (darwin alone, not the general tier's win32-OR-darwin), zsh-dispatch, case-sensitivity, plus chmod-mode-bit and symlink-keyword reused verbatim from the general tier (genuinely Unix-relevant, not Windows-motivated). Measured against the real tree: 196 of 930 eligible unit-suite files (21%), versus the general tier's 546 (59%) -- a real, evidence-backed narrowing. scripts/gen-platform-conformance-tier.cjs gains classifyMacosContent/ classifyMacosTree/renderMacosGeneratedFile and a --target windows (default, unchanged)/--target macos CLI flag, so the same generator produces two independent, gated outputs rather than needing a second script. New committed output: scripts/lib/macos-conformance-tier. generated.cjs. .github/workflows/test.yml's test-conformance job: only the macos-latest leg's file-list source changes; windows-latest is byte-for-byte untouched. New shipped-file ripples handled proactively (19 install-tree fixtures regenerated, bin/install.js registered). An isolated code-review pass found one real defect: the ADR's per- category count table had drifted by 1 (zsh-dispatch, case-sensitivity) because the new test file's own fixture strings joined the tree it classifies after the table was authored -- fixed, with the union total (196, what CI actually gates on) confirmed unaffected. An isolated security-review pass found no qualifying findings. The ADR also records an explicit requirement for any future widening proposal: check whether the motivating regression is already covered by Phase 1's no-rendered-text-length-assert lint rule (#4590) before re-proposing full macOS/Linux parity, since that is exactly what #4421's root cause was (a rendered-text-length assertion, not a real behavioral divergence). Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bcd99696d3 | chore(#4591): add platform-conformance-tier classifier + gate CI on it (#4598) | ||
|
|
37b965c0d1 |
enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code (#4553)
* enhance(#4139): Phase 7 — the agent-skill seam picks the payload in code ADR-4139 stream 2. The non-Claude `#2454` persona fallback in cmdAgentSkills (src/init.cts) now selects between a canonical agents/<name>.md and a token-minimized agents/<name>.compact.md sibling based on workflow.compact_content, resolved in code (a real function call with a real exit code) rather than a prose config-get gate — the same precedent stream 1's spine/detail split established for a load-bearing seam, applied here because this seam already runs through TypeScript instead of an eager @-include. A missing compact sibling falls back to the canonical persona and discloses the fallback in the served payload itself (a leading HTML-comment provenance line), so the Done-when contract — compact when on, canonical when off, never silent or empty — holds even for an agent nobody has compacted yet. Authored a .compact.md sibling for all 35 shipped agents (agents/gsd-*.md), each an independent, complete rewrite (not an extraction — nothing is "moved" the way spine/detail moves text) that preserves frontmatter, every @-include, every output-format contract, and every guardrail verbatim while cutting restatement and verbose framing. Verified mechanically: every pair registers (a canonical sibling exists), every compact file is strictly smaller, and the full @-include set matches canonical's — including which references are standalone eager-load lines versus inline prose mentions, since demoting one to inline changes what the host actually substitutes. Traced the install path before writing any code (.gsd/phase/.../40-design.md): stageAgentsForRuntimeWithConverter glob-copies every agents/*.md file with no stem filtering under the default full profile, so the new .compact.md files install for free with zero installer changes — matching issue #4407's stated scope. A tiered agent profile that doesn't stage a compact sibling degrades through the same fallback-with-provenance path already required for an unauthored one, so no installer change is needed there either. Extends tests/helpers/compact-content-variant.cjs with an AGENTS_ROOT export (deliberately not folded into DEFAULT_VARIANT_ROOTS, since agent variants are reached by a generic code construction rather than a literal path in prose, and checkReachability's markdown-search shape has nothing to find there). Reachability is instead proven behaviorally: tests/agent-skills.test.cjs's new "#4407 compact payload selection" describe block spawns gsd_run agent-skills against real compact/canonical fixture pairs and asserts on the served payload, which can only pass if the seam genuinely wires through. Fixed a pre-existing test whose agents/*.md glob incidentally matched the new .compact.md siblings (tests/agent-skills.test.cjs's Skill-frontmatter drift guard) and added the 35 new agents/*.compact.md entries to docs/INVENTORY.md's roster, both real, unrelated-to-content defects the new files' mere existence surfaced. Regenerated: install-tree fixtures (19 runtimes now ship 35 more agent files under the full profile), INVENTORY-MANIFEST.json, and the variant-swap token benchmark baseline (npm run benchmark:compact-content-variants --write). Closes #4407. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): apply orthogonal review findings from the compact-payload seam Standards axis of /code-review: extracted readNonEmptyFileOrNull(filePath) to collapse the duplicated read-and-empty-check shape between the compact and canonical branches in cmdAgentSkills, and updated the adjacent comment enumerating flat JSON extras to name agent_payload_variant alongside source/degraded (added by the prior commit, comment left stale). Security review and the Spec axis found no defects requiring a code change; their non-blocking observations (a pre-existing, unmodified path-construction pattern; the reasoned, documented substitution of a behavioral test for the literal reachability check) are recorded in .gsd/phase/enhance-4407-agent-skill-seam/60-review.json. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): repo-wide roster/cap fixes surfaced by shipping .compact.md agents Root-caused via a real gsd-test run (93 failures) rather than guessing which tests glob agents/ naively. Two classes of defect, both genuine: 1. Identity-roster confusion (11 files/areas): many tests and one production script derive "the set of GSD agents" from `readdirSync(agentsDir).filter(f => f.endsWith('.md'))`, which incidentally matched the new .compact.md variant siblings too — a compact file is a rendering of an EXISTING agent identity, not a new one. Fixed at the shared root (tests/helpers/agent-roster.cjs's listAgentFiles, which several tests already consolidated on) and at each independent glob that didn't use it: agent-size-budget.test.cjs (tier-cap lookup now strips the .compact suffix before checking XL/LARGE membership, so a compact file inherits its canonical sibling's tier instead of silently falling through to DEFAULT), agent-skills-bootstrap.test.cjs, check-contract-drift.test.cjs (the actual script, not just its test), codex-config.test.cjs (confirmed directly against generateCodexAgentToml that a compact role's derived sandbox_mode is byte-identical to its canonical sibling's before excluding it — not assumed), and copilot-install.test.cjs (two counts that legitimately DO need both files — an installed-file count and a full-conversion smoke test — fixed to expect 70, not stay pinned to 35). no-bare-gsd-tools-command-position.test.cjs needed the opposite kind of fix: two compact files reproduce descriptive prose already allowlisted at their canonical file's line number; added matching entries at the compact files' own line numbers rather than excluding them from the scan (a genuine bare gsd-tools command-position bug in a compact file would be as real a defect as in canonical). 2. A hard, non-ackable cap (found via emitted-attribution.test.cjs's real-tree run): six agents' compact renditions (gsd-debugger, gsd-executor, gsd-phase-researcher, gsd-plan-checker, gsd-planner, gsd-verifier) exceed the 32,768-byte NEW_FILE_CAP (ADR-1610) even after aggressive compaction — confirmed structural, not a compaction-quality gap: each is dominated by content this phase's own rules require verbatim (the ~2.6 KB gsd_run bootstrap preamble runtime-launcher-parity.test.cjs requires inlined in every agent that calls gsd_run, output-format contracts, guardrails). ADR-4139's prescribed remedy (spine + lazily-read parts) has no landing spot in cmdAgentSkills's single-file synchronous read. Removed these 6 compact files rather than ship an over-cap file or invent a multi-part read mechanism out of scope for this phase; recorded by name with the reason in .gsd/phase/enhance-4407-agent-skill-seam/40-design.md and 50-test-matrix.md, per #4407's own "or explicitly recorded as not worth covering" allowance. Their canonical personas are served correctly today via the fallback-with-disclosed-provenance path this phase's own Done-when #2 already requires — 29 of 35 agents now have a compact variant. Also fixes an unrelated, genuinely pre-existing defect this gsd-test run surfaced: gsd-core/workflows/execute-plan.md sat 21 bytes over its own DEFAULT-tier hard cap (40,960 bytes) at the branch point, before any change in this PR touched it — confirmed via `git show <merge-base>:...execute-plan.md | wc -c`. Per CLAUDE.md's no-deferral rule, fixed inline rather than filed: two meaning-preserving trims in the <success_criteria> block (a repeated parenthetical replaced with a same-exception reference; one redundant qualifier dropped) bring it to 40,940 bytes. Regenerated install-tree fixtures, INVENTORY-MANIFEST.json, and the variant benchmark baseline to reflect the 6 removed files. Docs/INVENTORY.md's 6 now-orphaned roster rows removed alongside them. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4407): make .compact.md-aware roster checks resilient to partial coverage Round 2 of the gsd-test-driven roster fixes: two checks assumed every agent has a compact sibling (true for 29 of 35 after the NEW_FILE_CAP exception), breaking once 6 stems legitimately have none. - tests/agent-classification-parity.test.cjs: the INVENTORY.md parser was picking up the "### Compact Payload Variants" subsection's rows as phantom/uncounted entries in the primary/advanced/inventory-only classification this test validates — a compact row documents an existing agent's alternate rendition and never gets its own AGENTS.md heading, so it was never meant to participate in that classification. Excluded at the parser, not per-assertion. - tests/copilot-install.test.cjs: the derived expected-file-list generator assumed every listAgentFiles() stem has a .compact.md source sibling; checks disk per stem now instead. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4407): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8b5e347377 |
chore: regenerate golden install-tree fixtures for the new lib file
scripts/lib/npm-version-check-diagnosis.cjs (added in the previous commit) is a new shipped file under the "scripts" files-glob, so it needs to appear in every runtime's golden install-tree fixture. Confirmed by gsd-test: 23 tests/golden-install-tree.test.cjs failures, one per runtime, each showing the same single added path. Ran npm run gen:install-tree; diff is exactly one line per fixture file, matching the new file. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a27cb6b2fa |
enhance(#4139): Phase 6 — the lazily-read remainder and the artifact templates (#4540)
* enhance(#4406): the lazily-read remainder and the artifact templates ADR-4139 Decision 3, Phase 6 of the #4139 Compact Content epic. Covers stream 1b (gsd-core/workflows/<name>/{modes,steps,templates}/*.md) and stream 4 (gsd-core/templates/**) with a variant-swap mechanism, confirmed with the user: two independent, complete files per covered path (canonical + .compact.md sibling), with the gate picking which one gets Read at the call site. This is a different shape from Phase 5's spine+detail partition, and is safe here specifically because these files are already reached only by a runtime Read — a missed Read already means zero overlay content today, with or without workflow.compact_content, so selecting between two independently-complete files introduces no new failure mode (documented in gsd-core/references/compact-content-gate.md's new "Streams 1b and 4" section). Disposition, after inspecting every candidate rather than trusting a byte-size threshold (same rigor Phase 5 applied to review.md): - Stream 1b: 1 of 78 files compacted (help/modes/full.md, a user-facing reference doc emitted verbatim, not orchestrator instruction). The other 9 size-threshold candidates are dominated by fail-closed guards, exact CLI invocations, or output-format contracts (AskUserQuestion blocks) — recorded not-worth-compacting, same reasoning as Phase 5's review.md. - Stream 4: a ground-truth reachability audit replaced the initial size-only candidate list. Two files (summary.md, user-setup.md) got compact variants; a third (spec.md) was drafted, then dropped after discovering its only two call sites are eager @-includes, not a runtime Read — stream-1 material hiding under gsd-core/templates/, not stream-4's actual mechanism. summary.md itself has 3 eager call sites and only 1 genuine runtime-Read call site (execute-plan.md); only that one was wired, so the compact variant's savings apply to the sequential single-plan execution path only. - Discovered while auditing reachability: 12 gsd-core/templates/** files with zero references anywhere in workflow/agent/command prose, compiled source, or tests — dead scaffolding predating this phase. Deleted in this same PR per this repo's no-defer policy, after re-verifying against a computed path.join(...) pattern (not just a plain-string search) that nearly caused two genuinely load-bearing templates (user-profile.md, dev-preferences.md) to be misclassified as dead. New checker (tests/helpers/compact-content-variant.cjs): registration, reachability, protected-content-preserved, size-smaller — replacing Phase 3/5's disjointness/completeness checks, which assume a partition rather than two deliberately-overlapping documents. The reachability check's own "unprefixed match" guard had a real bug (rejected the repo's own `~/.claude/gsd-core/...` convention), caught by running it against the already-wired help/modes/full.compact.md pair rather than only synthetic fixtures — fixed to anchor on the nearest `gsd-core` path segment instead. Template consumer parity (tests/compact-content-template-variant-parity.test.cjs): proves each compact variant's `## File Template` fenced block — the actual output-format contract a generated SUMMARY.md/USER-SETUP.md is parsed against — is byte-identical to the canonical file, then runs the one real deterministic consumer (gsd-core/bin/lib/coverage.cjs's classifyContent, backing `gsd-tools uat classify-coverage`) against content built from that shared contract. Added a sibling benchmark script (scripts/benchmark-compact-content-variants.cjs) rather than extending the existing spine/detail one — different data shape, and the existing script's own contract deliberately isolates it from a test-only helper's shape changing. Emitted-drift acknowledgement: not needed. Every changed/added path in this diff is hand-authored and present in the diff itself, so diffEmitted's attribution loop resolves `via` to the path's own source before reaching the ack-lookup branch (same reasoning Phase 5 verified for its own diff). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * enhance(#4406): address code-review findings on the variant-swap gate - docs/CONFIGURATION.md and gsd-core/references/planning-config.md's workflow.compact_content entries described only the spine+detail mechanism (Phase 5) and were missing this phase's variant-swap mechanism and its benchmark:compact-content-variants script entirely — required since this PR's changeset is type Added (CLAUDE.md's "Missing Docs for Changesets" rule). Both now describe both mechanisms and which call sites are wired. - Added the missing RED^-1/no-op fixture for checkProtectedContentPreserved: a canonical file with zero <!-- gsd:protected --> blocks must be a no-op, not a violation — the only branch of that function the existing fixtures didn't exercise. - Collapsed findCompactFiles/findMarkdownFiles in tests/helpers/compact-content-variant.cjs into one findFilesWithSuffix helper — the two were identical recursive walks differing only in the extension predicate (minor Duplicated-Code finding). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4406): restore copilot-instructions.md, a false-positive dead-template classification gsd-test caught this, not static analysis: 10 real failures in tests/copilot-install.test.cjs, tests/installer-migration-install.integration.test.cjs, and tests/repo-layout.test.cjs — all downstream of bin/install.js's Copilot install path, which does fs.readFileSync(path.join(targetDir, 'gsd-core', 'templates', 'copilot-instructions.md')) after copying gsd-core/templates/** into the target project, then merges it into both .github/copilot-instructions.md and (local installs) AGENTS.md. The reachability audit that flagged this file as dead checked src/*.cts and gsd-core/bin/*.cjs but never the repo-root bin/install.js — a separately maintained installer bundle outside the src/-to-gsd-core/bin/lib/ compiled-output convention. The fs.existsSync guard around that read degrades to a silent skip rather than a crash when the template is missing, which is why this surfaced only once the real E2E install test ran, not from any static check. Re-verified the remaining 11 deleted filenames against bin/install.js specifically (plain substring and quoted-filename search) before trusting that list — all 11 have zero hits there, confirmed dead by the same standard this one file failed. Regenerated the installer emitted-tree goldens (tests/fixtures/install-tree/*.json) to reflect the restored file, and corrected the "Removed" changeset (jolly-lynx-sprint.md) and the phase design doc from 12 to 11 deleted files. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Emitted-Drift-Ack-Growth: execute-plan.md — call-site wiring for the summary.md and user-setup.md .compact.md variants Emitted-Drift-Ack-Growth: help.md — call-site wiring for full.compact.md, same variant-resolution rule * docs(#4406): backfill changeset PR numbers Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4406): resolve removed-but-needed lint findings on the dead-template deletion CI's own full-test matrix (not gsd-test's matrix, which does not run this check) caught 4 more false-positive dead-template classifications via tests/removed-but-needed-lint.test.cjs / scripts/lint-removed-but-needed.cjs — a literal, word-boundary basename check across .github/workflows/, gsd-core/, and docs/ (excluding docs/adr/** and docs/research/**) for every file a PR deletes. It has no semantic awareness, so a deleted template's basename colliding with something else entirely still fires: - claude-md.md: gsd-core/templates/README.md had a stale table row claiming /gsd-profile reads this template to generate CLAUDE.md. Verified false (no code reads it anywhere, same search that already covered bin/install.js) — fixed the row to *(inline)*, matching every other command-generated artifact in that table. File stays deleted. - codebase/testing.md: collided with docs/guides/testing.md, an illustrative example row in docs-update.md's sample output table (an unrelated real generated-docs path). Swapped the example topic to "contributing" — the row is illustrative, any topic works. File stays deleted. - codebase/architecture.md, codebase/stack.md: collided with docs/reference/ planning-artifacts.md's directory listing of a user's own generated .planning/codebase/architecture.md and stack.md output — the same semantic mismatch already investigated and dismissed as unrelated earlier in this phase's audit, now caught by a gate instead of judgment. That listing repeats across 5 locale copies of the doc. - continue-here.md: collided with the real .continue-here.md pause-work artifact, referenced across 15+ locale and workflow files. For the last two, the lint's own error message offers "restore the file or update every consumer in the same commit." Rewording 15+ files across languages I cannot verify translation quality for, to shave 2 already-tiny templates that were merely presumed dead, is disproportionate to this PR's actual scope — restored codebase/architecture.md, codebase/stack.md, and continue-here.md instead, and corrected docs/ARCHITECTURE.md's Templates section accordingly. Final confirmed-dead set: claude-md.md, codebase/concerns.md, codebase/conventions.md, codebase/integrations.md, codebase/structure.md, codebase/testing.md, debug-subagent-prompt.md, discovery.md — 8 files, down from the original 12. Verified locally: GSD_REMOVED_BUT_NEEDED_BASE=next node scripts/lint-removed-but-needed.cjs now passes clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Emitted-Drift-Ack-Growth: docs-update.md — swapped an illustrative example-table topic (testing -> contributing) to avoid a removed-but-needed basename collision with the deleted codebase/testing.md template; net +10 bytes * fix(#4406): split codex-config.test.cjs to fix a genuine Windows CI timeout Root cause of the `full test (windows-latest, 24, shard 2/3)` failure the user asked to be actually fixed, not just re-run past: PR #4497 (landed 2026-09-07, one day before this PR's CI run) isolated tests/codex-config.test.cjs into its own dedicated chunk because its measured weight (17.87, ~45% of the post-cut Windows budget) made it unsafe to share a chunk with any other file. That isolation was necessary but not sufficient — even alone, with zero companion-file contention, the file's real Windows execution time sits right at the 600s per-chunk ceiling. Two independent CI runs on two unrelated PRs (this one and #4154) were both killed within ~1.4s of the identical 600000ms mark — not random contention, a deterministic near-miss the isolation fix couldn't address because it never reduced the file's own cost, only removed the risk of a companion file's cost stacking on top of it (which the PR #4497 comment explicitly anticipated: "if a future profiling pass genuinely speeds up codex-config.test.cjs itself, this isolation can be revisited"). The file itself explains why it's this heavy: 11,262 lines / 433 tests / 79 describe blocks, accumulated over dozens of bug-fix PRs (#2695, #2760, #3245, #3285, #3346, #3426, #3427, #3562, #3566, #3582, #3808, and more), several of which are explicitly documented as "folded" in from separate files that were never actually split back out ("Verified non-duplicate against both the pre-existing target and the other three folded sources"). Split into 4 files by top-level AST statement boundaries (never a naive column-0 regex — an early attempt at that overcounted 79 apparent "describe(" matches when only 21 are genuinely top-level; the rest are nested inside a handful of large folded-in blocks, which a regex can't tell apart from real top-level statements). Verified lossless twice: the split script asserts byte-for-byte reconstruction of every source character, and independently, total test()/describe() call counts match exactly between the original file and the sum across all 4 new files (433/79 both sides). Each new file carries the complete original shared header (imports/helpers) for safety; per-file unused-import warnings from that duplication are resolved via ESLint-precise alias renames (`{ foo: _foo }`, the standard form for an intentionally-unused destructured binding — never a bare `{ _foo }`, which would destructure a different, nonexistent property). No change needed to scripts/run-tests.cjs's ISOLATED_HEAVY_FILES or its pinned test in tests/run-tests-harness.test.cjs: the file that keeps the original name (tests/codex-config.test.cjs) is now only ~28% of the original's size and safely isolated in its own chunk as before; the other three new files re-enter normal weight-balanced packing, none individually close to disproportionate. Confirmed no other file hardcodes the hardcoded filename anywhere that would silently stop these tests from running (the CI test-selection scripts determine scope algorithmically, not by literal filename). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e03921c7d8 | enhance(#4405): split the rest of the eager-window workflows worth splitting (#4536) | ||
|
|
a0f8f956c4 |
enhance(#4139): Phase 3 — partition rules + the five checks (#4497)
* enhance(#4139): Phase 3 — partition rules + the five checks ADR-4139 Decision 5, epic #4139 Phase 3. Issue #4403's own "Proposed behavior" section lists four checks; the ADR's Decision 5 and its own phase table ("partition rules + the five checks") list five — the same four plus "boundary moves are declared, ongoing". Same issue-vs-ADR drift Phase 2 hit on the detail.md vs detail/*.md layout: the ADR is the locked, reviewed document, so it wins. This PR implements all five. docs/PARTITION-RULES.md (new) is the partition-rules document: the partition rule itself, the protected-content list and <!-- gsd:protected --> sentinel syntax (relocated unchanged from gsd-core/references/compact-content-protected-content.md, now deleted — it was never referenced by any runtime workflow Read, only by the predecessor test as documentation, so nothing at runtime regresses, and removing it from gsd-core/references/ also drops it from all 19 installed-project shipped-content trees for a file nothing ever read), and the five checks explained for a human reader. Referenced from a new CONTRIBUTING.md subsection under "Editing shipped content". tests/helpers/compact-content-split.cjs (new) is the shared mechanics: split discovery (any gsd-core/workflows/<name>/detail/*.md paired with <name>.md — no registry file, a pair is registered by existing on disk), line normalization (carries forward Phase 2's bare-label-line isTrivial fix and the canonical gsd_run-launcher-preamble exclusion), sentinel extraction, and a Boundary-Move-Declared commit-trailer reader that is a direct structural port of tests/helpers/emitted-runtime.cjs's Emitted-Drift-Ack-Hash/-Growth trailer reader (ADR-3942) — same merge-base range, same fail-closed throw on an uncomputable range, same dedupe/conflict rules. tests/compact-content-partition-guard.test.cjs (new) is the actual guard, superseding tests/plan-phase-compact-split.test.cjs (deleted — its per-pair checks are now the general guard's job for plan-phase specifically). Checks 2 (disjointness) and 3 (registration + size cap) run unconditionally against every registered split. Checks 1 (completeness, fires once per split on the PR that introduces a new detail/ path), 4 (protected content — no trailer can ever excuse this one, unlike check 5) and 5 (boundary moves declared) are PR-diff-scoped against the resolved base ref and skip cleanly when there's nothing to compare (a fresh clone, no PR in flight) — a deliberate asymmetry from check 5's trailer reader, which must throw rather than silently pass when ITS range is uncomputable, since that function is answering "did this PR declare its moves" rather than "is there even a diff to look at". Each of the five checks carries a RED (deliberately broken fixture) / GREEN (fixed) test pair, built against synthetic temp files or real throwaway git repos, per this repo's rule that a guard nobody has seen go red is not yet a guard. Building the real fixtures caught and fixed one real bug before it shipped: check 4's line-presence test was using the trivial-line-filtered normalizer, so a byte-identical spine falsely reported its own protected code-fence line as "deleted" — fixed with a non-filtering membership check. Extending docs/INVENTORY.md's "Workflow Sub-Files" table for `detail` surfaced a pre-existing, unrelated gap in the SAME area: gsd-core/workflows/<name>/templates/*.md is a fourth workflow sub-file kind that already existed on disk and was already known to lint-response-language-coverage.cjs's FRAGMENT_DIRS, but was invisible to gen-inventory-manifest.cjs and undocumented in that table. Fixed alongside it, same pattern, same PR, rather than deferred. Also, mechanically required by the new fourth sub-file kind: - scripts/lint-response-language-coverage.cjs: `detail` added to FRAGMENT_DIRS alongside modes/steps/templates — a detail/<part>.md inherits its parent's response_language coverage through the same per-file proof, not a parallel one. - tests/workflow-size-budget.test.cjs: explicit regression test locking that detail/ files are governed solely by the hard, non-waivable NEW_FILE_CAP (tests/helpers/emitted-diff.cjs) and never by the XL/LARGE/DEFAULT spine tiers — true by construction (measureWorkflows/listWorkflowStems don't recurse), made explicit per the issue's own Done-when item rather than left true-by-omission. - scripts/gen-inventory-manifest.cjs: `workflow_detail` and `workflow_templates` NESTED_FAMILIES entries; docs/INVENTORY-MANIFEST.json regenerated (plan-phase/detail/elaboration.md, discuss-phase/templates/*.md now tracked); docs/INVENTORY.md's table updated to four kinds. Verified: `npm run lint:ci` clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4403): review findings + a real gsd-test failure in the new guard Two orthogonal review passes (Standards + Spec, isolated sub-agents) plus a separate security review ran against the prior commit. Fixed everything each surfaced: - Security (Low, path-traversal existence oracle): checkRegistration's dangling-reference check extracted detail-path-shaped substrings from spine PROSE via a regex that permits `.`/`/` freely, then joined them onto repoRoot and probed fs.existsSync with no containment check — a spine file containing `../../../etc/detail/passwd.md`-shaped text could make the guard test file existence outside the repo. Added a path.relative-based containment check before the fs.existsSync call; anything that resolves outside repoRoot is now reported as a dangling reference directly, never probed on disk. - Standards (Boundary Coverage): the size-cap fixtures covered NEW_FILE_CAP and NEW_FILE_CAP-1 but not NEW_FILE_CAP+1 — added the third boundary-point case CLAUDE.md's TEST RULES require (limit-1/limit/limit+1). - Standards (Property-Based Testing): extractProtectedBlocks (a sentinel parser) and the new parseBoundaryMoveTrailerValues (a declare/dedupe/conflict parser, bijective-shaped) had no fast-check property test. Added three: a render/parse bijectivity property for the trailer parser (mirroring the exact ADR-3942 sibling test's alphabet/idiom), a dedupe-is-idempotent property for the same parser, and a well-formed-sentinel-round-trips property for extractProtectedBlocks. Then dispatched gsd-test on the resulting commit. It found a real bug the reviews couldn't have caught (none of them can run inside gsd-test's sandbox): checks 4/5's real-repo assertion failed against plan-phase's own split, reporting DISK_PLANS/#3218-comment lines as "undeclared boundary moves" — content Phase 2 (#4402) legitimately moved into detail/elaboration.md months before this PR's Boundary-Move-Declared mechanism existed to require a trailer for it. Root cause: `resolveBase()`'s own doc comment already documents that no `origin/*` remote-tracking ref exists inside the gsd-test sandbox container, and its fallback candidate (a bare `next` branch) can resolve to a point in history that predates an already-merged, already-reviewed split — making that split look "newly introduced" from the sandbox's vantage point. Check 1 (completeness) already scopes itself correctly to only genuinely-new detail paths (git diff status 'A'); checks 4 and 5 did not share that scoping, so a stale base made them re-litigate a settled split retroactively. Fixed by having checks 4/5 skip any split name check 1 already counted as newly-split — their own premise ("did an EXISTING split shed/undeclare something") does not apply to a split that is, from the resolved base's vantage point, brand new; that is check 1's domain alone. Verified locally (25/25 tests pass via a direct `node -e` require, since `node --test` is blocked in this repo) and via re-reasoning through the exact real-repo scenario the gsd-test failure showed. Also regenerated all 19 tests/fixtures/install-tree/*.json goldens — the prior commit's deletion of gsd-core/references/compact-content-protected-content.md was never reflected there, which is what golden-install-tree.test.cjs's other 19 failures in the same gsd-test run were. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4403): backfill changeset pr number to 4497 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4403): isolate codex-config.test.cjs into its own chunk, root-causing the Windows CI failure PR #4497's "full test (windows-latest, 24, shard 2/3)" job failed: run-tests killed chunk 3/8 at the 600s per-chunk backstop, with codex-config.test.cjs (weight 17.87, by far the chunk's dominant cost) packed alongside 39 other files. Traced, not assumed: - scripts/run-tests.cjs's own timeout-headroom comment for the OUTER per-shard timeout documents that "adding one test file reshuffled 115 of 268 unit files between shards" — shard/chunk composition is architecturally known to be unstable to single-file additions, which is exactly what this PR's own new tests/compact-content-partition-guard.test.cjs is. - A second comment, dated 2026-09-06 (one day before this PR, PR #4428's own CI), already documents the SAME chunk hitting the SAME 600s backstop with the SAME file (codex-config.test.cjs, "a genuinely MEASURED weight of 17.87 — not a stale-table miss") dominating it — the fix then was cutting the Windows per-chunk budget from 60 to 40. That cut clearly was not enough: two documented incidents in two days, at two different budget settings, both centered on one file that alone consumes ~45% of even the reduced Windows budget. - tests/test-timings.json's own header confirms its source data (test-events-linux-node22/24.jsonl) is Linux-only, and run-tests.cjs's own chunk-timeout diagnostic already prints "real Windows cost runs ~2.2x the recorded figure" — the packer's weight-balancing is working off data that is both stale (table last regenerated 2026-08-07) and known to underestimate the platform where the failure occurs. Given codex-config.test.cjs is disproportionately heavy AND every companion sharing its chunk is decided by a packing algorithm already documented as reshuffling unpredictably on any new file, tuning the shared budget a third time only moves the marginal line to wherever the next new file happens to land — it does not remove the gamble. Isolating codex-config.test.cjs into its own dedicated single-file chunk, unconditionally and on every platform, removes it at the source: the file never enters the pool packChunks balances, so no other file's packing changes, and no future single-file addition (mine or anyone else's) can silently reintroduce this exact failure by landing in its chunk. Extracted as a small pure function, partitionIsolatedFiles (mirroring this file's existing pattern of pulling packing/analysis logic out of main() for in-process unit coverage — see computeSweepProtectSet, analyzeChunkEvents), with 6 new tests in tests/run-tests-harness.test.cjs covering basename matching across path separators, near-miss non-matches, the empty-list case, and the isolated-set contents. Root cause is now closed rather than papered over with a retry: this failure is a property of one specific heavy file's chunk placement, not something that recurs randomly. If codex-config.test.cjs itself is ever genuinely sped up, this isolation can be revisited — this is a packing-side mitigation for a known file's cost, not a claim the cost is irreducible. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8b7a0b696b |
enhance(#4139): Phase 2 — one shared gate, one pilot split, one accuracy spot-check (#4471)
* enhance(#4402): split plan-phase into a spine + detail, add the shared compact-content gate ADR-4139 Decisions 3-5, Phase 2 of the #4139 Compact Content epic. Pilot split for plan-phase.md, the largest of the 58 eagerly-@-included workflow files (98,290 bytes): the spine keeps every happy-path step, every protected-content block (planner/checker prompt templates, quality gates, the failing-direction few-shot example, the two ScheduleWakeup guardrail paragraphs — each marked with a <!-- gsd:protected --> sentinel), and condensed one-paragraph summaries of five rare/opt-in fallback paths (planner and checker filesystem-hang recovery, phase-split recommendation, source-audit gaps, the thinking-partner conditional, and plan bounce). The full text of those five moves verbatim to gsd-core/workflows/plan-phase/detail.md (9.9KB, well under the 32,768-byte NEW_FILE_CAP), read by the spine only when workflow.compact_content is false (the default) — the exact same resolution rule now stated once in the new shared gsd-core/references/compact-content-gate.md, which every future split references instead of restating. Verified mechanically (tests/plan-phase-compact-split.test.cjs, scoped to this one split — Phase 3/#4403 owns the generalized guard): the union of spine + detail contains every non-trivial line the parent commit carried (0 missing), no non-trivial line is duplicated between them (0 duplicated), and every declared protected block is well-formed and non-empty. The spine shrinks from 98,290 to 93,206 bytes (-5.2% of the eager-window cost this epic exists to reduce); detail.md's 9,853 bytes are only ever paid by a project that has NOT opted in. Verified live, end to end, twice, against this actual repo (not a synthetic fixture) — real gsd-planner and gsd-plan-checker subagent spawns, real PLAN.md output: - workflow.compact_content=false: planned a real disposable phase (a docs/how-to page for enabling the key itself); planner returned PLANNING COMPLETE, checker returned VERIFICATION PASSED, all fact-checks against real repo state confirmed. - workflow.compact_content=true (detail.md never read): planned a second real disposable phase; planner returned PLANNING COMPLETE with frontmatter.validate and verify.plan-structure both clean, again fully grounded against real repo state. The five condensed fallback sections were independently re-read spine-only and confirmed sufficient to act on correctly without detail.md's elaboration. Also drafts gsd-core/references/compact-content-protected-content.md — the protected-content category list and <!-- gsd:protected --> sentinel syntax ADR-4139 Decision 5 calls for, written to move to Phase 3 (#4403) unchanged once it lands there. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4402): move detail.md into the ADR-4139-mandated detail/ subdirectory Two independent review sub-agents (Standards and Spec axes of /code-review) caught the same structural defect: ADR-4139 Decision 6 mandates gsd-core/workflows/<name>/detail/*.md ("one or more parts... individually skippable"), and this PR had shipped a flat plan-phase/detail.md instead, copying issue #4402's own (inconsistent) restatement rather than the locked ADR text. Fixed by git-mv to plan-phase/detail/elaboration.md and updating every cross-reference (the spine's step 0.5 gate pointer, the shared compact-content-gate.md's own resolution-rule wording, and the completeness test's path constants). Also, from the same review pass: - docs/CONFIGURATION.md and gsd-core/references/planning-config.md's workflow.compact_content rows said "nothing branches on it yet" — no longer true now that plan-phase.md's spine does. Updated both to name plan-phase as the pilot and note the rest of the corpus is still pending. - Regenerated all 19 tests/fixtures/install-tree/*.json golden fixtures (npm run gen:install-tree) — the three new shipped files were missing from the installer emitted-tree goldens. - Found via a cache-busted `eslint . --max-warnings 0` (this repo's eslint --cache has produced false-greens before): the split test's `git show` call had a bare `timeout: 10000` literal, tripping local/no-adhoc-timeout-literal. Extracted to the existing GIT_TIMEOUT_MS constant from tests/helpers/timeouts.cjs instead of a second guessed copy of the same class of timeout. Verified NOT needed, by tracing the actual mechanism rather than asserting (tests/helpers/emitted-provenance.cjs's gsd-core-verbatim rule attributes every gsd-core/{workflows,references}/** path to itself as an identity source): an Emitted-Drift-Ack-Hash/-Growth trailer. Every changed/added path in this diff is hand-authored and present in the diff itself, so diffEmitted's attribution loop resolves `via` to the path's own source before ever reaching the ack-lookup branch — there is no unattributed delta to acknowledge. The spine also shrank (98,290 to 93,206 bytes), so the growth ratchet has nothing to ack either. Re-verified after these changes: the completeness/disjointness self-check (0 missing, 0 duplicated) still holds against the relocated detail file, and a full `npm run lint:ci` passes clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4402): restore literal content the pre-existing drift guards pin on The first gsd-test run against this split (19 failures) surfaced real regressions: several pre-existing structural guards pin the EXACT text of the sections this split condensed, and paraphrasing broke them. - tests/plan-phase-drift-guard.test.cjs expects the literal `DISK_PLANS=$(gsd_run query find-phase ...)` bash assignment inside plan-phase.md itself, not a prose description of the same check. Restored the exact line into both §9a and §11a's spine summaries. - tests/thinking-partner.test.cjs expects plan-phase.md to literally offer "No, I'll decide" as the skip option. Restored that exact phrase into the condensed thinking-partner paragraph. - Both restores would have duplicated the same text into plan-phase/detail/elaboration.md (which still carries the full elaboration). Removed the now-redundant restatements from the detail file instead of leaving them duplicated — the spine already computes DISK_PLANS before the detail elaboration is ever read, so the detail file references it rather than recomputing it. - Re-running scripts/sync-runtime-launcher.cjs after that edit found the canonical gsd_run preamble had also become an unintentional spine/detail duplicate (both files call gsd_run and each is required, by runtime-launcher-parity's own contract, to carry its own copy). That's sanctioned duplication under a DIFFERENT contract, not lost/copy-pasted content, so tests/plan-phase-compact-split.test.cjs now excludes it from the disjointness check the same way it already excludes trivial fences/headings. - Applied the adversarial-review finding on tests/plan-phase-compact-split.test.cjs's own isTrivial(): a blanket `line.length <= 15` cutoff silently swallowed real content (e.g. the 14-char `<quality_gate>` sentinel). Replaced it with a specific bare-label-line pattern (`Options:`, `Display banner:` etc.) — verified 0 missing / 0 duplicated against the actual split, an improvement over both the original cutoff and a naive full removal (which produces false-positive "duplicates" on generic recurring labels). - gsd-core/references/planning-config.md's own workflow.compact_content row used `/gsd-plan-phase` (hyphen). That file is Claude-facing source text (gsd-core/references/), which tests/slash-command-namespace.test.cjs requires in colon form; docs/CONFIGURATION.md's use of the hyphen form is correct as-is since docs/ is human-facing and outside that test's scanned directories. Fixed to `/gsd:plan-phase`. - tests/plan-phase-compact-split.test.cjs's own `git show` of the parent commit failed inside the gsd-test sandbox ("detected dubious ownership") because the checkout is mounted under a UID the invoking user doesn't own. Scoped `-c safe.directory=<repo-root>` to that one git invocation rather than touching global git config. - docs/INVENTORY.md still had one outstanding "detail.md part" wording fix from the earlier adversarial-review pass, staged now. Re-verified locally against the exact assertions in all four affected test files (all pass) before dispatching a fresh gsd-test run — no change here should have broken any of the other 18 gates; `npm run lint` is clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4402): restore the full marker enumeration to §9a's spine trigger line The isolated Spec-axis review flagged that §9a's "Triggered when" line was condensed to "Agent() returns but the return contains no recognized marker" — dropping the literal `## PLANNING COMPLETE` / `## PHASE SPLIT RECOMMENDED` / `## ⚠ Source Audit` / `## CHECKPOINT REACHED` / `## PLANNING INCONCLUSIVE` enumeration, which is exactly the "machine- parsed structural headings" category compact-content-protected-content.md lists as protected. The load-bearing use of that same list (the gsd_stall_watch call and the Handle Planner Return bullets a few lines above) was never touched — only this one descriptive restatement was genericized — but leaving any instance of a protected category unsentineled is the silent erosion ADR-4139 Decision 4(c) warns sufficiency isn't machine-checkable enough to catch on its own. Restored the full enumeration into the spine. That reintroduced an exact duplicate into plan-phase/detail/elaboration.md, which still stated the same trigger sentence verbatim. Reworded the detail file's version to reference the spine's trigger condition instead of restating it, since the spine is now the single place that sentence lives in full — mirroring the DISK_PLANS/"already computed above" pattern from the previous commit. Re-verified locally: completeness/disjointness (0 missing, 0 duplicated) and all previously-fixed literal-content assertions still hold. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4402): backfill changeset pr number to 4471 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
476394689a |
fix(#4254): pin sequential executor to the orchestrator's validated root (#4476)
* test(#4254): sequential executor root pin — failing-first regression + matrix The new suite executes the shipped supplied-root-pin guard against real git fixtures (drifted primary-checkout cwd halts before the write and the FATAL names both roots; matching cwd permits it; unexpanded/empty pins halt; normalization forms; submodule and sibling boundaries; metacharacter quoting; drive-letter form gate) and locks the dispatch contract across execute-phase.md, its sequential-root-pin step fragment, and worktree-path-safety.md. The #2772 per-plan serialization assertion retargets to the fragment that now carries those rules (ADR-857 Phase 6 ceiling), plus the host-step wiring. * fix(#4254): pin sequential executor to the orchestrator's validated root Sequential-mode dispatch told the executor to self-derive PROJECT_ROOT from its own cwd; every existing guard is worktree-mode-only or self-referential, so an executor spawned with a drifted cwd committed onto the wrong checkout silently. - worktree-path-safety.md step 0p: mode-agnostic supplied-root pin guard, composed by the orchestrator at build time with the literal $ORCHESTRATOR_WT (git-vs-git comparison on both sides — representation-safe on Windows, the #4296 lesson), fail-closed on empty/unexpanded pins, registered-submodule allowance, warn-and-proceed only when the dispatch carries no pin block. - execute-phase.md sequential branch: build-time embed of the bound <project_root_pin> via the new execute-phase/steps/sequential-root-pin.md fragment (ADR-857 Phase 6 frozen ceiling — the host step cannot grow; the wave serialization rules move with the fragment, verbatim in substance) plus the per-write/commit pin instruction in <sequential_execution>. Worktree-mode dispatch untouched (its self-derived toplevel IS correct there). - INVENTORY rows (5 locales) + INVENTORY-MANIFEST + install-tree goldens regenerated for the new fragment; changeset added. * chore(#4254): backfill changeset PR number * fix(#4254): accept backslash-separated Windows drive pins CI on windows-latest showed every permit-path test failing with "Actual root: <none>": pins composed from Node's path.join arrive in the backslash drive form (C:\Users\RUNNER~1\...), which the guard's absolute-form gate rejected before the cwd-side root was ever computed — a legitimate matching pin could never pass. The gate now accepts either separator ([A-Za-z]:[\\/]); git -C resolves both forms (and 8.3 short names) to the same canonical toplevel, so the git-vs-git comparison is unaffected. Form-gate tests cover the emitted (C:/…) and produced (C:\…) spellings plus short names. * fix(#4254): portable drive-form gate for MSYS bash The bracket class [\\/] that accepted backslash drive pins parses inconsistently on MSYS bash (the Windows CI leg still rejected C:\ pins — every permit-path test red with "Actual root: <none>"). Replace it with standard pattern escaping outside brackets: [A-Za-z]:/*|[A-Za-z]:\\* — the escape form is version- and build-portable. Verified across all forms: both drive spellings accepted; bare "C:", relative, empty, and unexpanded rejected. * fix(#4254): runtime-generated backslash comparator + self-describing FATAL The Windows CI legs failed every #4254 permit-path row with 'Actual root: <none>' across two prior pattern spellings ([\\/] and \\*). Stage misattribution: <none> appears whenever the FATAL fires BEFORE the cwd-side capture assigns ACTUAL_ROOT — the absolute-form gate was what fired. Mechanism: the test harness spawns bash -c <script> through the Windows command-line boundary; that round-trip applies one extra shell-quoting pass with double-quote semantics — a backslash written twice in the script text arrives halved, while a lone backslash survives (the pin displays intact; row 9's pure-bash gate independently showed the halved pattern rejecting C:\ while C:/ still passed its surviving arm). On windows-latest every pin carries backslashes (os.tmpdir() is the 8.3 short form C:\Users\RUNNER~1\...), so the gate ate every pin before the actual root was ever computed. Fix, robust by construction: - the drive-form gate generates its backslash comparator at RUNTIME (BS=$(printf '\134'); match [A-Za-z]:"$BS"*) — the shipped guard now contains no doubled backslash anywhere, enforced by a regression assertion on the extracted guard text; - the FATAL self-describes: Guard stage (pin-unbound / form-gate / actual-capture / pinned-capture / root-mismatch) plus a Diagnostic line carrying git's own stderr for capture failures and both compared values for mismatches — future platform failures name their stage in the log; - row 9's hand-rolled duplicate case gate (transit-fragile copy, #4296 Minor 1 duplication smell) is replaced by driving the SHIPPED guard and asserting the stage; rows 2/4 pin the new stage machinery. Validated on darwin across drift/match/relative/unbound/empty/bare-drive/ forward-and-backslash drive forms, each also re-run under a simulated Windows transit (every doubled backslash halved) with identical outcomes. * fix(#4254): close the empty-comparator fail-open seam in the drive-form gate Self-review of the runtime-generated backslash comparator: if printf's octal escape ever returned empty, the drive arm [A-Za-z]:"$BS"* would widen to drive-RELATIVE pins (C:foo) — the construction's one theoretical fail-open path. Fail closed with a self-describing diagnostic instead of trusting the shell's printf. --------- Co-authored-by: sim <sim@local> |
||
|
|
2cf119f57e |
fix(#4217): reconcile artifacts before classifying an abnormally-ended executor (#4442)
* fix(#4217): reconcile artifacts before classifying abnormal ends * test(#4217): pin the completion-reconciliation contract * chore(#4217): regen derived inventory and install-tree fixtures * test(#4217): follow the #4003 anchoring pins into the reconciliation fragment Emitted-Drift-Ack-Growth: execute-phase.md — the runtime-neutral completion-reconciliation pointer, the two Codex wait-rule bindings, and the step-7 reconcile-first gate net +33 bytes over the extracted fallback block (#4217) * chore(#4217): add changeset fragment * chore(#4217): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
19b66c3ec8 |
fix(#4218): stop the orchestrator steering an executor that is still working (#4391)
* fix(#4218): stop the orchestrator steering an executor that is still working An executor with recent RED/GREEN/REFACTOR commits and passing verification had not yet written its SUMMARY because it was finishing closeout. The parent saw no local OS test/build process, inferred an "idle tail", and sent "Finalize immediately" into a working child; in CLI runs the same inference interrupted an executor before GREEN, leaving a RED commit and an uncommitted edit. The stall block said only "if no completion signal, no SUMMARY.md, and no expected-branch commits appear for N minutes" — it never said what to do when commits DO exist and only the SUMMARY is outstanding, never defined the threshold as a period without progress rather than a total runtime, and never ruled out a process listing as an idleness signal. Four rules close that: - the threshold measures a period WITHOUT MEANINGFUL PROGRESS, from the last sign of progress, not from dispatch — a long verification tail is not a stall; - commits + missing SUMMARY + recent activity resolves to KEEP WAITING, with steering, interrupting and re-dispatching each named and forbidden; - urgency/finalization messages ("Finalize immediately" and family) are forbidden outright — they arrive mid-verification and truncate a correct run. The existing user-facing pause is the only sanctioned stop, and `kill and retry` is a clean restart, not a nudge; - the absence of a local OS test/build process is NOT idleness: a native subagent runs in the runtime's own session, and an executor between two tool calls shows no process at all. Progress is judged only by the signals this workflow names. Five prose-contract assertions in tests/execute-phase-wave.test.cjs, all red on next. * fix(#4218): extract the progress policy to a step fragment CI's #1168 gate caught it: execute-phase.md sits 77 bytes under a frozen 93600 ceiling and the four rules added ~2.3 KB. "Extract, not bump" is the repo's stated remedy, and this workflow already carries policy detail that way. execute-phase/steps/executor-progress-policy.md owns the policy. The worktree-recovery arm moved with it — `kill and switch to inline execution` qualifies the stop this policy governs, so it belongs beside the rule about when stopping is sanctioned at all, not stranded in the host. The #3212 recovery OPTIONS stay in the host, where tests/config.test.cjs pins them. The host keeps what must be read before the orchestrator acts: the verdict, the threshold definition, and a pointer that fires before any message is sent to the child. execute-phase.md is now 93475 bytes — 48 SMALLER than next. * chore: add changeset for #4218 * chore(#4218): regenerate the inventory manifest for the new step fragment docs/INVENTORY-MANIFEST.json is the authoritative per-file list behind INVENTORY.md's `<workflow>/steps/*.md` row, so a new fragment has to appear there or gen-inventory-manifest --check reds the lint-tests lane. * chore(#4218): restore the issue ref on the allow-test-rule marker ADR-456 requires a #NNN on a new exemption; the block rewrite that moved the policy into the fragment dropped it. * chore(#4218): regenerate the install-tree fixtures for the new step fragment The fragment ships with the workflow, so every runtime's golden install tree gains one path — gen:install-tree is the generator that owns those fixtures. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
4c60879b5d |
fix(#4132): verify durable runtime surface sources (#4182)
* fix(#4132): verify durable runtime surface sources * chore(#4132): record PR number in changeset * test(#4132): cover rejected commands source alias * fix(#4132): reject aliased package fallback * test(#4132): cover rejected agents source alias * test(#4132): cover partially aliased marker provider * fix(#4132): reject partially aliased source providers * test(#4132): cover routed source identity probes * fix(#4132): route installed source identity probes * refactor(#4132): tighten installer source metadata * test(#4132): cover corpus trust boundary attacks * fix(#4132): close installed corpus trust gaps * refactor(#4132): keep installer authority private * fix(#4132): preserve private installer fallback * test(#4132): preserve fixture source authority * fix(#4132): reject overlapping source fallback * fix(#4132): avoid redundant installed corpus reads * refactor(#4132): simplify provider resolution * test(#4132): sync install tree fixtures after rebase --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
0fca71eaae |
enhance(#2529): cover every workflow with response-language directives + CI lint (#2558)
* enhance(#2529): cover every workflow with response-language directives + CI lint Every workflow now carries response-language coverage in one of three forms, and a CI lint keeps it that way. - 43 workflows load the new shared reference, `gsd-core/references/response-language-directive.md`, by eager `@`-import. - Lazy-loaded modes/steps/templates, which cannot rely on an eager import, carry an exact inline directive; 35 such paths are pinned by exact path. - Fragments dispatched by a covered parent inherit coverage, proven per file rather than granted per directory. The 45 workflows whose directive covered only "questions, prompts, and explanations" now name inter-tool narration, which is the defect #2529 reports: the running commentary between tool calls stayed English while the answers around it were translated. `scripts/lint-response-language-coverage.cjs` enforces it and fails closed on three independent discovery failures (unreadable catalog, empty catalog, unfollowed symlink). It resolves which reference a workflow imports and applies the same four-predicate test to that file, so a weakened shared reference uncovers its importers instead of passing silently, reported once as a systemic failure rather than 43 times. The walk follows symlinked subtrees with a realpath cycle bound. `lint:ci` invokes it by name. REQ-LANG-03 and REQ-LANG-04 state the contract in docs/FEATURES.md; REQ-LANG-04 names the two forms that satisfy it ("narration", "between tool calls") rather than enumerating class members an author cannot use verbatim, and a test pins that text to what the matcher accepts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): register the coverage test in the docs-guard lane `107eb8c1` (#3787) landed the docs-guard lane on `next` while this PR was open: a test that reads a `docs/` path must be named in `scripts/docs-guard-registry.cjs` or carry a `docs-guard-exempt` marker, so the guards that read a doc run on the PR that changes it. `tests/response-language-coverage.test.cjs` reads `docs/FEATURES.md` -- it extracts every form REQ-LANG-04 offers an author and runs each through the matcher that enforces it. Registration, not exemption, is the correct side of that gate: a reword of the requirement with no code change is precisely the diff this test exists to catch, and it is the diff the lane would otherwise skip. Registered narrowly (`['docs/FEATURES.md']`) rather than with the `'*'` sentinel, so an unrelated docs change does not pull this test into the lane. Verified: lint-docs-guard-registration 0 violations, tests/ci-docs-guard-registry.test.cjs 51/51, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): consolidate this PR's emitted-growth acks into its own fragment This PR ripples emitted bytes across 85 workflow paths. Until now each ripple was acknowledged by appending to whichever live fragment owned that path, because two ack sources may never name the same path. `a84f7563` (#3078) swept all 45 fully-spent fragments off `next`. Forty-two of the paths this PR grows were owned by swept fragments, so those keys are now unowned and this PR's own fragment declares them directly -- one path, one source, and no dependence on a fragment that no longer exists. Each adopted entry keeps its measurement and records where it came from. Two paths are handled differently, because the sweep did not free them: - `review.md` is now owned by `3034-parallel-reviewer-lanes.json`, which landed on `next` after the sweep. Its entry is live, so the old route still applies: this PR's note is appended to that entry rather than declared a second time. - `plan-review-convergence.md` keeps the arrangement made in round 24. Result: 3 fragments in the directory, 85 keys in this PR's own, 0 cross-source duplicates. `lint-emitted-drift-ack` exit 0, `tests/emitted-attribution.test.cjs` green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): move REQ-LANG-03/04 into the feature fragment that now generates them `36375513` (#3845) made docs/FEATURES.md a generated projection of docs/features/*.md, marked "do not edit by hand". This PR wrote REQ-LANG-03 and REQ-LANG-04 straight into the generated file, so the rebase left the requirement present in the projection and absent from its source -- the next regeneration would have deleted both, and `tests/features-index-gate.test.cjs` was already red on the mismatch. Both requirements now live in docs/features/response-language-config.md alongside REQ-LANG-01 and -02. Regenerating produces a docs/FEATURES.md that is byte-identical to the committed one, so the text this PR shipped is unchanged -- only its source of truth moved to where #3840 put it. The docs-guard registration is widened to name the fragment as well as the projection. The requirement's source is the fragment now, and an edit there that skips regeneration would otherwise reach this guard through neither path. Verified: features-index-gate 68/68, lint-docs-guard-registration 0 violations, ci-docs-guard-registry + response-language-coverage 142/142. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): hand the plan-phase ack back to its new live owner `c933184b` (#3825) landed `3172-stated-failing-direction.json` on `next` after fragment had adopted that path when the sweep left it unowned, so the merged tree named it from two sources -- a hard failure in `scripts/lint-emitted-drift-ack.cjs`. The path has a live owner again, so the append route applies: this PR's note joins that entry, carrying its own measurement, and the key is dropped from this PR's fragment (84 keys left, the others untouched). The provenance sentence written for the swept-fragment case is removed rather than reused -- this path was never orphaned, so that account of it would be false. Same shape as `review.md` and `plan-review-convergence.md`: ownership is a property of the merged tree, and a fragment landing upstream after a push can reclaim a key no local check would have flagged. Verified: lint-emitted-drift-ack exit 0, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): state byte figures that are true against the tree The reference claimed `execute-phase.md` has "2 bytes of headroom under the ceiling named below". That was true when the sentence was written -- the file sat at 93398 against the 93400 comfort assert -- and upstream has since shrunk it to 91493 against a 93600 hard ceiling, so the figure now understates the headroom by three orders of magnitude. The rationale the sentence supports does not depend on the number, so the number is gone rather than refreshed: a restated figure would go stale again on the next upstream edit, and nothing parses it. Audited every other numeric claim this PR ships the same way, mechanically against the merge base: all 82 FILE-delta claims in the ack fragment match the real per-file delta exactly, and the 1,629-byte reference and 63-byte import line check out. One class was imprecise: the 41 notes for workflows whose inline directive was rewritten in place quoted the conversion counterfactual as "+1,692 bytes more loaded context", which is the reference form's whole weight, not the increase over the inline directive those files already carry. Each now names both quantities and the net (+1,605 / +1,609 / +1,584). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): one rule for pinned vs inherited coverage, and the docs to pick it Review measured that 14 of the 35 pinned fragments would pass by inheritance anyway, and that the PR asserted both readings at once: inheritance is real coverage (so those 14 pins are noise) or it is not (so 30 inheriting fragments are green-but-uncovered). Only one can be true. Inheritance is real: the predicate proves it per file -- the parent must dispatch this exact path from a read/execute context AND be covered itself -- so the parent's directive is in the loaded context by the time the fragment is read. The 14 pins are therefore removed along with the directive lines they pinned, and those files inherit like the 30 structurally identical ones. The rule is now stated where the set is declared, and enforced from the other side by a test: no member of the pinned set may be one that would have inherited. That is what decides the form for the next fragment. - pinned set 35 -> 21; 14 workflow files revert to their base content - `findViolations` no longer returns early on a pinned path: a file that becomes eagerly loaded and takes the shared reference is strictly better off, and the gate must not red that. The reference form is admitted because its own wording is validated in turn; an arbitrary reworded inline line still fails. - the reference-directive cache is keyed by size and mtime, not by path alone, so a rewritten reference re-asked in one process no longer returns the stale verdict - `carriesInlineDirective` names its negation blindness: four independent hits read vocabulary, not polarity - the real-tree scan asserts each source produced files instead of `> 152`, a constant that read as the workflow count and would have passed a scan that lost one of its two directories - the pinned-set size assertion goes the same way: the size follows from the rule, so the rule is what the suite asserts Docs, for the gate that now governs every future workflow: - `docs/contributing/response-language-coverage.md` -- why the narration class is the discriminator, the four coverage forms, the decision order that picks one, the pinned line, and what each failure message means - a row in CONTRIBUTING.md's CI checks table, matching the docs-guard row - `docs/CONFIGURATION.md` points at it from the `response_language` entry Also: the changeset said 45 reworded workflows; it is 44 (42 @-reference + 21 pinned + 44 rewritten = 107 touched). That text ships to CHANGELOG.md. `3707-parse-gap-reporting.json` landed on `next` reclaiming `audit-uat.md` and `progress.md`; both handed back by the append route, leaving 82 keys here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): correct the reference-taker count, 43 -> 42 The ack notes said the import line is byte-identical "in each of the 43 workflows that take the reference" and that the alternative would be "43 inline copies". The shared reference has 42 importers; the 43rd file in review's table is `execute-phase.md`, which imports the OTHER reference. Corrected in all 41 notes that carry the sentence, across this PR's fragment and the two it appends to. Found by re-running the numeric audit from the previous round after the rebase, which also re-verified all 84 FILE-delta claims against the new base -- all exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): migrate the emitted-drift ack from a fragment to commit trailers ADR-3942 (#3954) landed while this PR was open: the acknowledgment is now a commit trailer and tests/emitted-drift-acks/ no longer exists. The fragment is deleted and each key it declared becomes one trailer, reasons unchanged. The four keys this PR had handed to 3034-*, 3172-* and 3707-* under the one-source rule come home here. That rule was the whole reason for the hand-backs, and the trailer model has no shared namespace to collide in -- five of this PR's rounds were spent on exactly those collisions. Emitted-Drift-Ack-Growth: add-backlog.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-tests.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: add-todo.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: ai-integration-phase.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by |
||
|
|
2f4f7538e9 | fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) | ||
|
|
2f64e6230a |
feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local> |
||
|
|
cad70f4f3e |
fix(#4120): replace shellcheck npm dep with dependency-free downloader (#4121)
* fix(#4120): replace shellcheck npm dep with dependency-free downloader The `shellcheck` devDependency (added in #4109) pulled in decompress@4.2.1 for archive extraction, which carries an unpatched CRITICAL zip-slip vulnerability (GHSA-mp2f-45pm-3cg9, CVSS 9.1) plus two moderate findings. decompress's latest published version IS the vulnerable one -- no patched release exists upstream, so npm audit fix cannot resolve this by upgrading. Removes the shellcheck package entirely and replaces its role with scripts/lib/shellcheck-fetch.cjs: a small downloader using only Node's built-in https/zlib plus a hand-written tar-entry reader, fetching a pinned koalaman/shellcheck release directly from GitHub releases. The reader never uses an archive-supplied name as a filesystem path (the exact defect class decompress had) -- it only returns the matched entry's bytes; the caller writes those bytes to a path it constructs itself. Bounds the download with a 30s-per-hop timeout, consistent with the ShellCheck subprocess's own timeout. Covers linux/darwin on x86_64/aarch64, matching this repo's actual CI (lint-tests runs only on ubuntu-latest) and local dev needs; Windows fails with a clear, honest error rather than silently misbehaving. Adds tests/lint-workflow-shellcheck-fetch.test.cjs covering the tar-parser (unit cases plus a fast-check property test per CLAUDE.md's parser-testing requirement), a security behavioral pin confirming traversal-style entry names are treated as opaque strings never filesystem paths, and boundary coverage for the redirect-following logic's MAX_REDIRECTS limit (limit-1/limit/limit+1, via an injectable transport, no real network I/O). npm audit: 0 vulnerabilities (was 1 critical + 5 moderate). The lint script reproduces the identical result against the current tree: "212 pre-existing finding(s) from baseline, 0 new" -- no behavior regression, no baseline changes needed. * fix(#4120): register shellcheck-fetch.cjs with the installer scripts/lib/shellcheck-fetch.cjs shipped without being added to GSD_SCRIPTS_LIB_FILES in bin/install.js, which would have left it orphaned on uninstall and broken the golden install-tree fixtures for every runtime. Adds the entry and regenerates the 19 affected fixtures via npm run gen:install-tree. * docs(#4120): add changeset for the decompress CVE fix --------- Co-authored-by: sim <sim@local> |
||
|
|
8c9265d4e5 |
fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a blocker" but tagged severity: warning — the tier plan-phase's revision loop counts as must-fix — and the planner is never taught the rule, so every multi-wave phase touching shared mutable state replans at least once, and intentionally coupled plans re-flag identically every iteration to the stall prompt. Three coordinated changes: - gsd-plan-checker: retag 3b to severity: info, the tier references/revision-loop.md already exempts by design; recognize a coupling_justified frontmatter declaration in the Do-NOT-flag list so deliberate pairs converge. Additions are offset by trimming 3b motivation prose — the checker sits 45 bytes under its LARGE hard cap. - plan-phase step 12: INFO-only accept — an issues block with zero BLOCKER/WARNING entries accepts the plan and surfaces the advisories instead of re-entering the revision loop. Real blockers and warnings still gate unconditionally. - gsd-planner: slim pointer in assign_waves to the new progressive-disclosure reference gsd-core/references/planner-coupling.md (the planner sits 19 chars under its own cap), which carries the shared-mutable-state rule and the coupling_justified escape hatch so first-pass plans avoid the finding when the coupling is unintentional. Documented the coupling_justified field in docs/reference/plan-md.md. Growth acks per #2914; inventory manifest and install-tree fixtures regenerated for the new reference file. Closes #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): pin Dimension 3b at severity: info The severity retag makes the old assertion (severity: warning) stale; lock the advisory tier from both directions — info must be present, warning must not — so a future edit cannot silently re-arm the revision-loop trigger. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * chore(#3724): changeset fragment for PR #3758 Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * docs(#3724): roster planner-coupling.md in docs/INVENTORY.md The new reference was enumerated in the manifest and all 19 install-tree fixtures but missing its row in the Modular Planner Decomposition table — the roster half the manifest-sync test cannot check. (Review Blocker.) Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): cover all four acceptance criteria (review round 1) - plan-checker-coupling: the 3b severity assertion is now a PARITY check deriving the exempt tier from revision-loop.md's flow instead of hardcoding info — editing either side alone reds the suite. New describe pins the other three criteria: plan-phase's INFO-only accept clause (proven failing-first), the BLOCKER + WARNING count staying intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and the planner pointer + planner-coupling.md content. - ack fragment: $comment's plan-phase figure corrected to +79B; the 2775 pin note carried forward into the gsd-planner.md entry, updated for upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves verbatim). The parallel-dependent-plans re-anchor this commit originally carried was superseded by upstream #3764 during review; this branch no longer touches that file. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 2 — align the stance enumeration, complete the template contract MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it. Funded by extracting the inline <examples> block to the new progressive- disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined from the same spot; #1949 precedent), which also restores the 3b motivation clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base (49107 -> 48486) — the extraction the byte pressure was owed. MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and the field's shape becomes one 'plan-id: reason' string per coupled peer so a plan justified against two peers can express it; docs/reference/plan-md.md's Type column names the shape. NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow ratchet an 'XL tier'. Acks and derived artifacts updated accordingly (checker entry removed — a shrink needs no ack; INVENTORY roster row + regen:derived for the new file). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): derive the 3b negative severity assertion (review round 2) Every severity token in the 3b span must BE the tier revision-loop.md exempts, replacing the hardcoded severity:warning negative — if the loop's exemption ever moves, the failure names the real conflict instead of blaming the agent file with a mutually-unsatisfiable pair. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): refit the planner coupling pointer under the char cap Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the base, leaving 5 chars of headroom where the +16-char pointer was measured against 13 more. The pointer prose shortens to 'Non-file coupling:' — 49150 chars, back under the strict 49152-char cap — and the ack figures follow. The @-path the tests pin is unchanged. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep Upstream #3078/#3823 deleted all fully-spent ack fragments, including 3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md +79B append. Per the collision remedy that sweep added: take the deletion and home the still-live entry in this PR's own fragment. Figures re-measured at this merge base (90871 -> 90950 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack Upstream #3825 shipped 3172-stated-failing-direction.json naming only plan-phase.md, now spent at the base — colliding with this PR's live plan-phase entry. Per the #3003 pattern the fully-spent single-path fragment is deleted and this fragment stays the path's one source; figures re-measured at this base (93073 -> 93152 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 3 — true up the ack figures, restore the wave comment The fragment's absolute sizes are re-measured and anchored to base |
||
|
|
6beaa66b25 |
enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate Content-assertion suite for the Step 7 re-verification evidence gate (agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md). Committed before the implementation to prove RED via gsd-test. * enhance(#3304): gate re-verification blockers on deterministic evidence Step 7's anti-pattern scan re-runs at full, unbounded scope on every re-verification pass, independent of the must-haves established in Step 2. A blocker it finds — other than the self-evidencing debt-marker check — previously reverted a completed gap-closure round and started another --gaps cycle on nothing more than the verifier's own new judgment call, with no bound on how many times that could repeat. A Step 7 blocker now blocks unconditionally in re-verification mode only if it is a carried-forward gap (present in the prior VERIFICATION.md's gaps: list) or the flagged file was git-modified since the prior pass (a regression; fails closed toward blocking when history is unresolvable). Otherwise it predates the gap-closure round unflagged and needs deterministic evidence — a named test run red, or another concrete reproducible artifact — to stay blocking. Unevidenced, it downgrades to a new advisory: frontmatter list and report section instead of setting status: gaps_found, and never reverts a completed must-have. Maintainer approval was narrowed to this evidence condition only, explicitly rejecting the broader "advisory whenever untraceable to a requirement/decision/prior-gap" proposal — implemented and pinned by tests/verifier-evidence-gate.test.cjs and documented as rejected in gsd-core/references/verifier-evidence-gate.md so it can't silently re-expand. Closes #3304 * fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself (not the production prose): a {0,600} match window was shorter than the 724-char paragraph it was scanning (the "exclude from Step 9 Rule 1" phrase starts at offset 662), and two regexes assumed no indentation after a markdown list-continuation line break. All three phrases are confirmed unique across agents/gsd-verifier.md, so the windowed submatches are replaced with direct whole-string assertions instead of just widening the window. Also acknowledges the deliberate byte growth in agents/gsd-verifier.md that the differential-attribution check (ADR-2719) correctly flagged. Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152). * docs(#3304): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
8487f0ed42 |
enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage - pin configured, absent, and malformed branch-list behavior - require opposite CLI and execute warning outcomes * feat(01-01): warn on configured protected branches - resolve the base branch union configured protected branch names - expose exact boolean CLI comparison output for workflow callers - keep execute-phase warning advisory and within its byte budget * test(01-01): add failing protected branch config coverage - cover valid list persistence and null unset - reject hostile shapes while preserving the prior value * feat(01-01): validate protected branch configuration - register git.protected_branches as a canonical config key - require a non-empty array of non-blank branch names * test(01-02): add failing ship protected-branch controls - Execute both workflow warning blocks with exact predicate arguments - Require true and false results to produce opposite warning outcomes - Preserve the none-strategy feature-branch offer contract * feat(01-02): warn at ship on protected branches - Reuse the typed protected-branch predicate in ship preflight - Keep raw base resolution for PR targeting and advisory branch creation - Prove execute and ship warning blocks with opposite-result controls * test(01-02): add failing protected-branch docs parity - Require the canonical schema key in both English config references - Pin the non-empty string-array type and absent default - Require synchronized multi-branch examples and advisory semantics * feat(01-02): publish protected branch configuration contract - Document the optional non-empty string-array field in both references - Explain resolved-base union and absent-field compatibility - Keep execute and ship warnings advisory under branching_strategy none * fix(01): CR-01 honor active workstream branch policy * fix(01): WR-01 assert protected config path selection * docs: add changeset fragment for #3648 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx * fix(#3648): resolve base_branch precedence inversion and round-1 findings Blocker 1/2: production config resolution was flat-first, so a project that migrated to git.base_branch but still carried a stale flat base_branch got the old value back. Add base_branch to normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos pattern: canonical nested wins) and route readEffectiveGitConfig's test seam through the same normalization so it can't silently diverge from production again. Adds a regression test with both keys set that fails without the fix. Blocker 3/4/5: restore the handle_branching case-selector prose and "none" contract sentence that #3389's tests anchor on, and revert the unrelated prose/comment compaction in the same step — both were drive-by edits outside #3552's scope. Also addresses review majors/minors: delete readConfigBaseBranch and readConfigProtectedBranches (dead in production, only self-tested); --is-protected now fails closed (reports protected) instead of silently answering false when the base branch can't be verified; trim configured protected-branch names; fix HOME-without-USERPROFILE vacuous isolation on Windows; correct the drift-ack's byte accounting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N * test(#3648): add failing legacy-key hoist safety coverage Round-2 review found normalizeLegacyKeys block 5 records a normalization carrying the DISCARDED flat value on the canonical-wins branch. Probing that turned up a second, unreported defect in the same helper shape: blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no object guard, so a config whose section key holds a string is spread into index keys — {"git":"main","base_branch":"release"} -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}} The resolved value is accidentally still correct, so nothing fails and no diagnostic fires. But normalizations.length > 0 sets configDirty, and config-loader then serializes that shape back into the user's config.json — a read that silently corrupts config. The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"} explicitly; this is the input it would have caught. Covers both defects across blocks 1 and 5, with object/array/null negative controls that must stay green in both phases, and a fast-check property over arbitrary `git` values. * test(#3648): pin fail-closed handling of malformed protected_branches Replaces the test that pinned the fail-OPEN behaviour. The old assertion — ['develop', 42] yields isProtected === false for 'develop' — locked in the exact failure #3552 exists to close: config-set validation is bypassable by a direct edit of .planning/config.json, so a user who believes 'develop' is protected got a silent false and no warning. It was also inconsistent with the fail-CLOSED direction twelve lines away, where an unverified base reports protected and writes a diagnostic. A protection predicate must not have two opposite failure directions depending on which input is bad (#3648 review Blocker 3). New coverage: a bad element drops only itself, a non-array contributes no names, an empty list is well-formed rather than malformed, and --is-protected surfaces the rejection. Both negative controls — a clean list reports nothing rejected and writes no diagnostic — must stay green in either phase, so the reject channel cannot fire unconditionally. * fix(#3648): drop only invalid protected_branches and report them Partition git.protected_branches instead of discarding the whole list on one bad element, and carry the rejections out through ProtectedBranchStatus so --is-protected can name them on stderr. Valid names keep protecting; the user finds out the rest were ignored. A non-array value still contributes no names — a bare string is not a list of branch names — but is now reported rather than swallowed. An empty array stays silent: declaring no extra protected branches is a valid choice, not a misconfiguration. writeDiagnostic is hoisted out of the unverified-base branch since both arms now use it. * test(#3648): prove the predicate diagnostic survives both call sites The workflow bash stub now emits a stderr diagnostic the way the real command does, which is what makes a swallowed `2>/dev/null` visible to a test — previously the stub was silent on stderr, so discarding it changed no observable behaviour and the call sites could drop the explanation undetected. Adds the Minor 2 binding check as well: ship must expose the predicate result as IS_PROTECTED rather than only echoing a warning, asserted by running the extracted bash and reading the bound value, not by grepping the workflow source. Both tests carry opposite-outcome controls — an empty diagnostic must leave the text absent, and a false predicate must bind false. * fix(#3648): surface the predicate diagnostic and bind ship's result Drop `2>/dev/null` from the --is-protected call at both call sites. The fail-closed explanation and the new rejected-entry warning both go to stderr, so discarding it left the user with a bare "protected branch" warning on a branch that is not protected and no way to tell a real match from a degraded-git guess. `git branch --show-current` keeps its own redirect — that one is genuine noise. ship.md binds IS_PROTECTED and its prose now branches on the variable, so the following steps have evaluable state instead of having to infer it from warning text in tool output. execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth 319 bytes (was 331 before the redirect came out). Baseline re-verified against the current rebase base by blob id; the ceiling check passes with 755 bytes of margin. * test(#3648): restore negative space for the readFile config seam The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm it pinned survives verbatim in readEffectiveGitConfig's readFile branch — the JSON.parse catch, the non-object guard, the git-section object guard, .trim() and blank-string rejection — and the four surviving readFile injections were positive-path only. protected_branches was never driven through this seam at all. Restores nine cases against the seam, including protected_branches partitioning, plus a control proving loadConfig still wins when both seams are supplied. Records honestly what the suite pins. Mutating the built lib shows .trim() is KILLED, while the non-object guard and the blank-string rejection SURVIVE — both are unreachable through this entry point for the same reasons the deleted suite documented against its own equivalents: a JSON-parsed non-object carries no relevant own-property either way, and a blank value is rejected a second time downstream by the resolver's truthiness check. They stay as defence-in-depth and are labelled known-unkillable rather than left looking like coverage this suite does not provide. * test(#3648): distinguish detached HEAD from a missing branch argument `args[1] ?? ''` collapsed two different situations into one: a detached HEAD, where `git branch --show-current` legitimately prints nothing, and the flag being called with no argument at all. Both answered false, so the right outcome arrived by an unintentional path and a caller bug was indistinguishable from normal operation. Asserts the detached case stays silent and the missing-argument case reports, with a control that the two diagnostics differ. * fix(#3648): report a missing --is-protected branch argument Answer false either way, but say so when the flag arrives with no argument. A detached HEAD passes an explicit empty string and stays silent, since that is a normal state rather than a misconfiguration. * docs(#3648): state exact-name matching and per-entry rejection isProtected is exact string equality, so a git-flow project must enumerate every release/* and hotfix/* by name. #3552 only asked for an integration-branch field, so the implementation satisfies the letter of the issue while leaving its git-flow motivation partly unserved — say so where users will meet it rather than leaving them to discover it. Also documents the Blocker 3 behaviour change: an invalid entry is ignored with a warning naming it and the remaining names still apply. Both statements land in docs/CONFIGURATION.md and gsd-core/references/planning-config.md, and the config-field-docs parity test asserts each in both so the two cannot drift. * refactor(#3648): extract isValidProtectedBranches for cross-surface pinning The `git.protected_branches` check inside `cmdConfigSet` and the resolver's per-entry filter in `git-base-branch.cts` are deliberately different shapes — all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot fail the guard open. Nothing structural keeps their two definitions of "usable branch name" in step. Lifting the write-side check into a named, exported predicate lets a property test ask both surfaces about the same value and assert they agree, which is the fast-check gap the round-2 review flagged. No behaviour change: the predicate is the same expression, called from the same place. * fix(#3648): stop --is-protected rewriting the config it is asking about `gsd_run query git.base-branch --is-protected` runs on every execute-phase and every ship. It resolved config through `loadConfig`, whose normalize-then-write path rewrites `.planning/config.json` whenever any legacy key normalizes — so a boolean question was silently editing the user's checked-in config. This PR had widened the trigger by adding a fifth normalization block (top-level `base_branch` -> `git.base_branch`), making it fire for exactly the projects the feature targets. `loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution is unchanged, only the two write-back side effects are suppressed. The predicate passes `persist: false`; the ~30 other callers are untouched, so a legacy config is still migrated by ordinary use. Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and reflows whitespace even when the values are equivalent. Three tests, each with its own control: the end-to-end CLI leaves the file byte-identical while still answering `true` from the legacy key (proving the config WAS read); an ordinary persisting load of the same fixture DOES change the bytes (proving the fixture is live rather than inert); and `persist:false` vs default over one directory returns deep-equal config while differing on the write. Reverting the one-line `persist: false` fails the first of those and only that one. Also from the review: - `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through the same precedence authority production uses". It does not, and cannot — it reproduces two of production's steps over a single file. The comment now names what the seam covers and what it does NOT (root/workstream deep merge, builtin and global defaults, federated merge), and the seam now applies production's flat-then-nested lookup so it stops disagreeing about a surviving flat key. - The missing-argument diagnostic promised "answering false", which the fail-closed guard on the same call can contradict by printing `true`. It now states what it did with the argument and leaves the answer to stdout. * test(#3648): re-pin block 5 on #3760's refusal contract #3767 landed on next while this PR was in review and fixed the non-object config-section defect properly: a present-but-non-object section now BLOCKS its own migration — value preserved, no Normalization pushed, refusal reported via `skipped[]` — rather than being rebuilt from a plain-object view. That supersedes this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread but still dropped the section value silently, and which the round-3 review correctly called out as destruction in place of corruption. The rebase drops that commit and routes block 5 through the upstream helper. This file's tests asserted the superseded design, so they are rewritten to pin block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's suite was written — against the contract that now governs it: ordinary hoist into an absent/null/object section, canonical-nested-wins, and refusal for each of string/number/boolean/array sections with the exact `skipped` entry. Two controls keep it from passing vacuously: the refusal must be scoped to block 5 (an unrelated block still normalizes in the same call), and a property over arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually exclusive per key, that a refusal leaves both the section and the legacy key untouched, and that a hoist manufactures no index key the input did not carry. * docs(#3648): correct the Git Query and Config Loader module contracts CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct `.planning/config.json` read. Since this PR it is the EFFECTIVE configuration resolved by the Config Loader — a materially different authority, carrying the root/workstream deep merge, flat-then-nested lookup and builtin/federated defaults. The `--is-protected` predicate, `git.protected_branches`, and the two invariants that distinguish the predicate from the plain query (fails closed on an unverified base; must not write) were undocumented entirely. The Config Loader entry now states that loading is not side-effect-free by default and documents `options.persist`. docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write` was run and produced no diff: the manifest indexes roster NAMES, not row prose, so a description edit cannot move it. Also closes the global-defaults minor: `git.protected_branches` is inert in `~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is section-wide and predates this PR, so the fix is to state the scope where users meet it rather than to quietly extend the resolution set for two new keys. * fix(#3648): close four defects found by the round-4 external review Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially against this branch. Four findings reproduced against source; each is fixed with a failing-first test and a control, and each fix was verified by reverting it and watching exactly the intended test fail. 1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was only half closed. `loadConfigResolved` re-enters itself with a bare `{ workstream: null }` when a workstream has no config.json of its own, and that literal discarded every other option — so the recursive pass ran at the DEFAULT persistence and rewrote the ROOT config. Reproduced: with GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected` rewrote `.planning/config.json` despite `persist:false`. Both recursions now forward `options` and override only `workstream`; the explicit override still wins the hasOwnProperty check, so spreading cannot let `workstreamContext` reintroduce a workstream. 2. Both workflow call sites failed OPEN, and aborted under `set -e` (both reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty string when the query fails, so `[ "$X" = true ]` was simply false: no warning, no trace — a silent hole in the guard whose only job is to warn. The bare assignment also aborted the step under `set -e`. Both sites now degrade VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the check did not run. Deliberately not fail-closed — claiming "protected" on no evidence would warn on every branch whenever gsd-tools is unavailable. 3. `isValidProtectedBranches` and the resolver disagreed on a sparse array (antigravity). `.every()` skips holes; the resolver's `for...of` yields `undefined` for them, so `["main", , "develop"]` was accepted by config-set and rejected by the resolver. The cross-surface property passed only because `fc.array` cannot generate a hole. The predicate now indexes, and the generator punches holes so that axis is actually falsifiable. JSON cannot express a hole, so this is unreachable in production — but two definitions of one predicate must not contradict each other. 4. A top-level `protected_branches` silently outranked `git.protected_branches` (antigravity). Routing the key through `get(key, {section, field})` gave it flat-then-nested precedence, which is back-compat for keys `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has no legacy form, so that invented an undocumented alias. It now resolves nested-only through a new `getNested`, in production and in the test seam. `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's refusal path can leave behind — and a control pins that distinction. Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`); a git command that runs and exits non-zero counts as a clean negative, so a cwd that is not a repository answers `false`, not `true`. Verified pre-existing on next @ |
||
|
|
370cfc6680 |
enhance(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% (#4043)
* feat(#4036): persist CI shard/job timeout-vs-cap trending, warn at 90% Adds two new mechanisms plus an audit-coverage extension: - scripts/lib/ci-job-timing.cjs: shared elapsed-vs-cap arithmetic - scripts/ci-check-job-near-cap.cjs: in-job advisory near-cap check, wired into test/test-full/mutate/smoke as each job's last step - scripts/ci-timeout-report.cjs + .github/workflows/ci-timeout-report.yml: scheduled REST-API poll that appends new records to tests/ci-timeout-budget-history.jsonl and opens a small data-only PR - tests/ci-test-job-timeout-budget.test.cjs: extended to cover mutate (mutation.yml) and smoke (install-smoke.yml), which previously had no headroom-factor gate coverage at all Does not change any timeout-minutes value, shard composition, or shard-1 contents — those stay maintainer policy calls per the issue's own scope. * fix(#4036): address two-orthogonal-review findings - Parity tests guarding the two hand-duplicated literals this design cannot single-source through GH Actions YAML: CI_JOB_TIMEOUT_MINUTES vs each job's own timeout-minutes, and ci-timeout-report.cjs's JOB_RULES name-prefixes vs each job's actual name: template. - Thread run.event through as runEvent on every persisted record, so PR-context and push-context install-smoke timings (genuinely different matrix shape) are distinguishable in the history rather than silently conflated under one job name. - Replace the Windows near-cap start-time step's ambiguous PowerShell +/>> precedence with GitHub's documented string-interpolation form. - Move github.run_id out of direct ${{ }} shell interpolation into an env: var in the new scheduled workflow, per this repo's own expression-injection-safe convention. * test(#4036): regenerate golden install-tree fixtures for scripts/lib/ci-job-timing.cjs npm run gen:install-tree — scripts/ ships wholesale into the installed package (per ADR/known-defect precedent from #4012's own PR history: a new scripts/lib/*.cjs file needs its golden entry regenerated or every runtime's install-tree test fails). Confirmed via gsd-test: this was the sole cause of the first real verification run's 25 failures (all in tests/golden-install-tree.test.cjs, one per runtime). Top-level scripts/*.cjs files (ci-check-job-near-cap.cjs, ci-timeout-report.cjs) are not individually tracked in these fixtures — consistent with every other existing top-level scripts/*.cjs file, so no entry was expected or added for those two. * fix(#4036): register new lib file with installer, fix H1 shell policy - bin/install.js: add ci-job-timing.cjs to GSD_SCRIPTS_LIB_FILES (a hand-maintained registry, not generated — tests/install.test.cjs asserts every scripts/lib/ file is enumerated here) - test.yml: replace the two OS-conditional "Record job start time" step pairs (test + test-full jobs) with a single unconditional `node -e` step. The prior pair's Windows variant declared an explicit shell: pwsh, which scripts/workflow-policy.cjs's H1 checker statically flags against every OS a job's matrix can realize, independent of the step's own if: gate. A single Node one-liner needs no shell override at all — it's syntactically valid and behaves identically under bash, zsh, and pwsh — which is both H1 compliant and removes the last OS-specific shell syntax from this change entirely. Both defects were found by a real gsd-test run, not local gates — lint:ci and build:lib were clean throughout because neither the scripts/lib/ install-manifest parity check nor the H1 shell-policy baseline runs as part of lint:ci; both are gsd-test-only suites. * docs(#4036): how-to for reading CI timeout budget signals The phase-gate docs check correctly flagged the enablement sequence as 3 real steps (read the near-cap warning, find the accumulated trend file, pick the right maintainer lever) — a reference table can't carry a sequence. Adds docs/how-to/read-ci-timeout-signals.md, indexed from docs/README.md. * chore(#4036): backfill changeset PR number (4043) --------- Co-authored-by: sim <sim@local> |
||
|
|
11c24ce973 |
chore(#4012): regenerate the 19 install-tree goldens
scripts/ ships, so a new file under scripts/lib/ moves every install-tree golden. The remote runner caught this as 26 failures across all 19 runtimes. My fault in the dispatch: I limited the implementation's verification to build:lib and lint:ci and left regen:derived out. Both of those passed, which is precisely why a green local gate is not a substitute for the runner — the goldens are only exercised there. One line per golden: scripts/lib/ndjson-reporter.cjs joining the shipped tree. Refs #4012 |
||
|
|
2ea5efc151 |
enhance(#3911): hooks declare their crash policy (#3960)
* enhance(#3911): give hooks an exit seam that needs no build ADR-3889 Phase 7 foundation. The 19 shipped enforcement hooks hold 91 of the epic's 128 terminators and cannot reach `terminateNow` today. The obvious route — requiring `gsd-core/bin/lib/cli-exit.cjs`, as gsd-agent-isolation-guard.js already does for two other modules — is rejected. That precedent carries its own warning (#3582): those files are tsc output, gitignored and absent on a raw plugin-marketplace or git-clone install, so the hook must first call ensureRuntimeBuild() to self-heal. Making the module a hook needs IN ORDER TO TERMINATE depend on a build inverts the dependency, and its failure mode is precisely the fail-open this phase exists to remove: a guard that cannot terminate cannot deny. `lint-hooks-runtime-build-seam` already encodes that concern, and Design B would have had to add an ensureRuntimeBuild() call to all 19 hooks to satisfy it. So `hooks/lib/` becomes a third emit location for cli-exit and a fifth for the registry, preserving the invariant `src/cli-exit.cts`'s own header states: it imports nothing but node:fs and its sibling registry, and the generator dual-emits that sibling alongside each copy so a relative require resolves next to whichever copy loaded it. Shipping needed no change — build-hooks.js already declares HOOKS_SUBDIRS_TO_COPY = ['lib']. Proven, not asserted: the two files are copied into an otherwise-empty tmpdir and a child process requires them and terminates — PASS exits 0, HOOK_DENY exits 2 with the payload on both stdout and stderr. That test fails the moment the hooks copy gains a require reaching outside hooks/lib/. Also fixed inline: the registry's fifth target let any `--write` test overwrite the real committed hooks/lib/exit-code-registry.js, because the test helper derived only three of the other output paths. It now redirects all five, and a regression test asserts every committed artifact is byte-identical after a redirected write. Install-tree goldens pick up the two new shipped paths across 11 runtimes — insertions only, no removals. lint:ci was green while they were stale, so this was found by regenerating rather than by a gate. Verification runs on the remote runner. Refs #3911 * enhance(#3911): declare a crash policy, and migrate the write guard Adds `hooks/lib/hook-exit.js` — the hook-facing vocabulary over `terminateNow`, hand-written because the cli-exit copy beside it is generated: allow(payload) exit 0 deny(payload, stderr?) exit 2 crash(onCrash, payload) whichever the hook DECLARED `crash()` takes the policy as a required argument with no default, which is the whole mechanism: fail-open by accident stops being expressible. A hook must name ALLOW or DENY at the call site, and an unrecognized value terminates INTERNAL rather than guessing. Fail-open stays legal; fail-open by omission does not. `gsd-write-guard.js` is the first hook migrated, all 12 sites, and it exposed a gap in the seam. `terminateNow`'s doc comment justified its fd-2 write by citing this hook's `emitBlock` — but modeled it as sending the same bytes to both streams, when `emitBlock` actually sends full JSON to stdout and only the bare `reason` string to stderr, because Kimi's hook bus feeds stderr verbatim back to the model. Migrating as written would have turned a readable sentence into a JSON blob for Kimi-backed agents. #3911 requires both "all 19 hooks terminate through terminateNow" and "no hook's effective default changes". Those are jointly satisfiable only by teaching the seam to carry a distinct stderr payload, so `terminateNow` gains an optional third argument: omitted, behavior is byte-for-byte what it was; a string is written raw, which is exactly the Kimi case. The doc comment's inaccurate claim about emitBlock is corrected in place. Proven rather than asserted: the pre-migration file is reconstructed from HEAD and driven with the same catastrophic-shrink payload as the migrated one — exit code, stdout and stderr all byte-identical. Verification runs on the remote runner. Refs #3911 * enhance(#3911): all 19 hooks terminate through the seam Migrates the remaining 18 enforcement hooks onto allow/deny/crash. An AST walk now reports zero `process.exit(` call sites across every `hooks/*.js` — down from the 91 the census measured. Each hook with an outer catch declares its policy once, at module top, with the reason that policy is right for that specific guard: a read guard that cannot scan must not block the read; a statusline that renders every prompt must degrade rather than crash; an injection scanner must not retroactively block a result already returned. Those sentences are the deliverable — they are what turns fail-open-by-accident into fail-open-on-purpose. No hook's effective default changed. Wiring exposed two defects, both fixed here rather than noted. A SECOND stdout/stderr-splitting site turned up in `gsd-workflow-guard.js`'s `emitForceAddBlock`, matching the pattern already known from the write guard — full JSON to stdout, bare reason to stderr for the Kimi bus. It uses the `stderrPayload` argument added in the previous commit, which is now carrying its second real caller rather than one special case. More seriously, `terminateNow` emitted both streams inside ONE try, so a payload that failed to serialize aborted before the stderr write ever ran. The two windsurf guards write nothing to stdout on a block and only a reason string to stderr, so `deny(undefined, reason)` exited 2 with EMPTY stderr — a deny that silently loses its reason, which is the exact "fails with success" class this epic exists to close. The streams are now emitted independently, each with its own guard, and `undefined` means "nothing to write for this stream" rather than an error. Regression tests inject a throwing write on one fd and assert the other still receives its payload; they fail against the single-try version. Byte-identity was proven per hook, not assumed: each pre-change file is reconstructed from HEAD and driven side by side with the migrated one across its normal path, its deny path, malformed stdin and empty stdin — exit code, stdout and stderr compared. Verification runs on the remote runner. Refs #3911 * enhance(#3911): harden the three shell hooks, and pin every hook's policy `gsd-phase-boundary.sh`, `gsd-session-state.sh` and `gsd-validate-commit.sh` gain `set -euo pipefail`. The expected hazard did not materialize, and that is worth recording: every intentionally-non-zero command in all three is already the condition of an `if`/`elif`, which `set -e` never fires on, and none of them reads a possibly-unset variable or pipes through a grep that may legitimately match nothing. No `|| true` guards were needed. Each hook was still checked command-by-command before the flags went in rather than after. Twenty-one before/after cases across the three hooks — disabled and enabled, planning and non-planning, missing STATE.md, malformed JSON, the Kimi payload shape, quoted and unquoted `-m`, valid and over-long Conventional Commits — all match on exit code, stdout and stderr. The hardening is shown to actually fire, not merely added: with a stubbed `node` that fails at the JSON-emit step, phase-boundary and session-state go from silently exiting 0 with empty stdout to failing visibly with the error surfaced. No such case could be constructed for `gsd-validate-commit.sh`, whose every statement already sits inside an if-condition — recorded as unproven rather than claimed. `tests/hooks-crash-policy.test.cjs` adds the per-hook coverage the issue asks for, table-driven over all 19 hooks rather than 76 hand-written cases: normal allow, deny where a deny path exists, crash-honors-the-declared-policy, and an unclosed-stdin case — the one `process.exitCode` structurally cannot serve. The deny assertions encode each hook's ACTUAL stream split rather than a uniform shape, since four of the six deliberately differ. A drift guard enumerates `hooks/*.js` and fails if a terminating hook is ever added without a row. Writing those tests surfaced two hooks that emit a block decision in their JSON body and exit 0. Both were checked rather than assumed, and neither is a fails-with-success: `gsd-read-injection-scanner.js` is PostToolUse, where the tool has already run and exit 2 has no meaning, and `gsd-cursor-subagent-start.js` follows Cursor's JSON-body protocol. They are deliberately left alone — a mechanical sweep to `deny()` would have broken exactly these two. Verification runs on the remote runner. Refs #3911 * fix(#3838): the commit validator says when it could not validate #3911 claims to subsume #3838. Measurement said otherwise, so this closes it for real rather than by assertion. `set -euo pipefail`, added earlier on this branch, does NOT fix #3838: bash exempts a command used as an `if` condition from `set -e`, and all three of the hook's swallow-and-pass sites are exactly that shape. Verified against the hardened hook with a node shim that fails only the classifier call — a non-conforming commit still exited 0 with empty stdout AND empty stderr, indistinguishable from "your commit conforms". That is the defect verbatim. All three sites named in #3838 now capture the real exit status instead of consuming it as a condition, and each distinguishes its genuine negative from "could not run": - the classifier: 0 = is a git commit, 1 = genuinely not one, anything else = could not classify. Its `node -e` now wraps the require and the call in try/catch and exits 3 on a throw, so a broken require chain can never be mistaken for `isGitSubcommand` legitimately returning false — which is the arm that matters, since `token-scanner.cjs` is a gitignored build artifact and a fresh checkout lands there. - the opt-in config read and the JSON command extraction get the same treatment. On "could not run" the hook emits a diagnostic to stderr naming which check failed and why, then exits 0. The issue confirms this is safe — it is a PreToolUse hook, so stderr does not disturb the JSON protocol — and ranks it the smallest sufficient fix. The gate still fails open, but it can no longer do so silently, which is the whole complaint: a validator that disables itself quietly costs more than one that is absent, because it is trusted. Both controls are unchanged and pinned by tests: a conforming commit still passes silently, a non-conforming one still exits 2 with its existing block payload. The defect test asserts stderr is non-empty and names the failure; it fails against the pre-fix hook. Verification runs on the remote runner. Refs #3911, #3838 * docs(#3911): document the hook crash-policy contract Reference and Explanation via a new docs/features fragment (FEATURES.md is generated from it), INVENTORY rows for the three new hooks/lib files, and an ARCHITECTURE note on the hooks section. How-To: docs/how-to/declare-a-hook-crash-policy.md, indexed from docs/README.md — a hook author now has to choose and declare a crash policy, which is more than one step and crosses into which harness protocol their hook speaks. It covers allow/deny/crash, writing an ON_CRASH reason that is actually useful, when a deny needs a distinct stderr payload, the two hooks whose harness reads a JSON-body decision and must NOT use deny(), and what to do when a check cannot run at all — with #3838 as the worked example. Refs #3911 * test(#3911): prove the seam actually ships, and stop hand-rolling temp cleanup Two review findings. The acceptance criterion 'hooks/dist/** stays in parity via the build seam (lint:hooks-runtime-build-seam)' was misstated and unmet: that lint checks something else — that a hook requiring a compiled gsd-core/bin/lib module also calls ensureRuntimeBuild(). Nothing exercised that the three new hooks/lib files reach hooks/dist/lib at all. That gap is not theoretical: #770 is a recorded ship-blocking bug where a new hook never shipped because a copy list missed it. The suite now builds dist through the repo's own ensureBuiltHooks(), byte-compares each shipped copy against its source, and spawns a child that requires the SHIPPED dist copy and denies — which is what catches a copy that exists but cannot resolve its sibling registry. gsd-validate-commit.sh hand-duplicated mktemp/run/rm three times; one idempotent trap on EXIT replaces them, guarded so cleanup cannot alter the exit status. Behavior-neutral across five cases, with temp-file counts taken before and after each run. Refs #3911 * fix(#3911): stage transitive hook lib requires, not just one level The remote run returned 7 failures across 3 real causes. The important one is a PRODUCTION bug this phase exposed rather than caused. `writeCursorHooksJson` scanned each hook script for `./lib/X` requires exactly one level deep and never re-scanned the lib files it staged for their own sibling requires. Nothing had a transitive lib dependency before, so the gap was invisible. Adding hook-exit.js -> cli-exit.js -> exit-code-registry.js made real Cursor installs ship a bundle that dies at require time with MODULE_NOT_FOUND. It now walks to a fixed point, and a real installed Cursor hook runs to completion. The staging harness in shared-hooks-dir-resolution hand-copied its fixture, so the injection scanner crashed at require time and its exit-1 was being read as a policy decision. Migrated to copyScriptWithDeps, which walks the require graph — the repo's recorded rule for this class, since adding another copyFileSync keeps it alive for the next person. The missing-lib-source test in cursor-hook-workspace-roots hardcoded which lib file it expected to be named in the abort message; the same throw now fires for a different file first. Its assertion is unchanged in substance — staging still must abort rather than ship a broken hook — only the name is no longer pinned. The last one was my own test asserting an uppercase reason code. Measured against origin/next: the pre-change hook emits the same lowercase 'config_unreadable', so the test was wrong, not the migration. Corrected to the real value rather than making the code match the test. Verification runs on the remote runner. Refs #3911 * chore(#3911): regenerate the cursor install-tree golden The staging fix means a Cursor install now correctly carries the two transitive lib files it was silently missing. Additive only — no path was removed. The golden diff is the evidence the packaging defect was real. Refs #3911 * chore(#3911): backfill the changeset PR number Refs #3911 * fix(#3911): a git probe that timed out is not a negative A macOS CI lane failed three deny cases at 2084ms, 2112ms and 2177ms — just past the 2000ms budget these hooks give their git probes. The three that passed took 72ms, 595ms and 651ms. Under shard contention `git rev-parse` overruns, the hook reads the non-zero result as "not a git repo", and allows with exit 0 and empty stdout AND empty stderr. Under load, the guards silently stop guarding. That is ADR-3889's thesis exactly, sitting inside the security hooks this phase is about. The repo had already recognized the class in one place — gsd-cursor-subagent-start.js fail-closed-denies on `git_timed_out` (#3045) — but nowhere else. `hooks/lib/git-probe.js` classifies a probe's outcome, distinguishing a real non-zero exit from ETIMEDOUT, a signal kill, and a spawn failure, rather than folding all four into `status !== 0`. Three guards route their eight git probes through it. The resolution is the same shape #3838 took, and the same one that issue endorsed as smallest-sufficient: fail open, but loudly. **No exit code changes on any path** — a developer on a loaded machine is still not blocked, which keeps #3911's declaration-pass contract intact for exit codes. What changes is that the hook now says on stderr which probe could not answer, instead of presenting silence as a clean verdict. Scope was checked across every hooks/*.js, not just the three that failed: gsd-agent-isolation-guard spawns no git; gsd-statusline's two probes gate only a cosmetic display segment, not an allow/deny decision, and are left alone. The C2 deny assertion was a real-race test — it demanded exit 2 while a slow git legitimately yields 0. It now requires the hook to either deny, or allow with a diagnostic naming the probe that could not run; a silent allow still fails, so the assertion is not vacuous. A deterministic regression stubs git on PATH to sleep past the budget rather than waiting for load to reproduce it. Verification runs on the remote runner. Refs #3911 * test(#3911): a PATH shim cannot intercept the hooks' git spawn on Windows The deterministic timeout regression stubbed git on PATH and asserted the guard reports rather than silently allows. It passes on Linux and macOS and failed on Windows in 83ms and 176ms — the stub was never invoked at all. Mechanism: the hooks call spawnSync('git', args) with no shell:true, so on Windows CreateProcess resolves git.exe only and never a PATH .cmd shim. The git.cmd branch could not have worked and is removed rather than left implying a Windows path that does. Adding shell:true to the hooks to serve a test would change product behavior and widen an injection surface, so the case is skipped on win32 only, with the mechanism written into the skip reason so a future reader does not 'fix' it that way. Linux and macOS keep the coverage, and macOS is where the underlying fail-open was actually caught. Refs #3911 --------- Co-authored-by: sim <sim@local> |
||
|
|
1e67ec9737 |
enhance(#3908): the scanners distinguish an empty diff from one they could not compute (#3937)
* feat(#3908): the scanners distinguish an empty diff from one they could not compute collect_files ended 2>/dev/null || true, which destroyed the evidence three ways: the redirect discarded git's diagnostic, the pipe replaced git's status with grep's, and || true forced success regardless. Four distinct conditions - an established-empty diff, a bad ref, no repository, and a repository with no commits - all reported clean, and a secret scanner reporting clean because git failed is indistinguishable from an all-clear to any gate consuming it. git now runs separately from the filter so its status and diagnostic both survive. An established-empty diff exits NO_INPUT; a scope that could not be established exits UNAVAILABLE; the usage sites move off 2 to USAGE. || true is retained on the filter alone, where it is correct: a diff of only images is empty, not failed. Codes are sourced from a generated shell fragment rather than written into three scripts, so a re-allocation cannot desync them, and a missing fragment fails loudly instead of falling back to literals. The security workflow is updated in the same change: without it, a docs-only PR would newly fail the job. * fix(#3908): keep scanner stderr out of the file list, and drop try/finally from test bodies Capturing git and find output with 2>&1 was right for the failure path but wrong for the success path: a warning emitted alongside a successful diff flowed into the file list and was treated as a filename. stderr is now captured separately, forwarded as a warning on success and as the diagnostic on failure, and never folded into the list. Also converts the control tests' try/finally blocks to t.after(), which CONTRIBUTING bans inside a test body because it masks failures. * chore(#3908): backfill changeset pr number * docs(#3908): record the scanners' four-outcome exit contract SECURITY.md is root-level, so the docs gate correctly held: a Changed fragment owes a file under docs/. The contract also belongs where the feature is described, as REQ-SCAN-INJ-05. docs/FEATURES.md is GENERATED from per-feature fragments (#3840) - the first edit went into the generated file and gen-features --check caught it, which is the same edit-the-output drift this epic exists to close. The fragment is the source; FEATURES.md is regenerated. --------- Co-authored-by: sim <sim@local> |
||
|
|
941b62249e |
enhance(#3906): two terminators over one registry, with a versioned exit projection (#3924)
* feat(#3906): two terminators over one registry, with a versioned projection Adds terminateNow (write-then-terminate, for callers that cannot wait for the event loop) beside runMain (drain-then-exit), both projecting through one shared function so they cannot disagree - the parity the ADR makes mandatory. A failed write does not change the exit code: letting it propagate would fail a hook open, which is what the fail-closed branches exist to prevent. The projection is versioned. v1 reproduces today's integers, including keeping a payload-carried degraded result at exit 0 - ADR-2980 ratified that across 60 sites and declined normalizing it on measured blast radius. v2 applies the registry. --exit-contract=v2 or GSD_EXIT_CONTRACT=v2 selects it; an unrecognized version throws rather than silently defaulting. The registry is now emitted beside both copies of the exit module, so it resolves as a sibling in the built tree and in the committed scripts/ copy that must load on an unbuilt clone. * fix(#3906): actually restrict code 2 to terminateNow, and generate the registry's type The claim that terminateNow is the only place 2 can be produced was false: runMain's outcome arm applied no guard, so runMain(()=>'HOOK_DENY') set exitCode 2 through the drain path - and the parity matrix demonstrated it while calling it parity. runMain now refuses any outcome projecting to the hook-protocol code, gated on the code rather than the name so an alias cannot slip past, and the matrix asserts the restriction instead of contradicting it. The ambient type for the generated registry was hand-written with no gate against the generator's actual output - the declared-surface-diverges-from-runtime defect class this epic exists to close, reintroduced inside it. It is now a third generated artifact covered by the same --check. Also converts every test-body try/finally to t.after(). * test(#3906): derive the glossary fixture's dependencies instead of hand-listing them Adding a require to scripts/lib/cli-exit.cjs broke 31 tests in one suite that built its fixture from a hand-written dependency list, so the new sibling was absent and the copied script could not load. copyScriptWithDeps walks the require graph and exists for exactly this class - #3412 paid the same bill when one new require broke 82 tests across two suites. Migrating rather than adding another copyFileSync line keeps the class closed. The other nine suites referencing that path were triaged; none copies-and-spawns, so none needed migrating. * fix(#3906): enumerate the new shipped file, drop a vendor name from shipped data, and fix three test defects install: scripts/lib/exit-code-registry.cjs was missing from GSD_SCRIPTS_LIB_FILES, so it shipped to every install and orphaned on uninstall. The registry gave HOOK_DENY a meaning naming one harness, and that string ships into every runtime's tree - a guard correctly caught it leaking into the hermes and qwen installs. The registry is runtime-neutral infrastructure; the vendor name belongs in the ADR, not in shipped data. Two more fixture harnesses built their trees from hand-listed dependencies and broke on the new require; both migrated to the derived helper, and all 23 copy-and-spawn candidates were enumerated so the class is closed rather than patched. One generator test used a fixture code that collided with a real allocation, so the generator correctly reported a duplicate where the test expected drift. The large-payload test embedded a 256KB literal in the child's argv, exceeding Linux's 128KiB MAX_ARG_STRLEN so the child never started - it now builds the payload inside the child. * chore(#3906): backfill changeset pr number * docs(#3906): document the exit-code contract selector P2 is the first phase of this epic with a user-invocable surface, so the flag and env var owe a reference entry. Records what actually differs between v1 and v2 today (one outcome), that an unrecognized value is rejected rather than silently defaulted, and the fail-safe property that makes switching safe. --------- Co-authored-by: sim <sim@local> |
||
|
|
878f25025c |
enhance(#3905): the exit-code registry — one number, one meaning, enforced at build (#3920)
* feat(#3905): the exit-code registry — one number, one meaning, enforced at build A generated registry replaces locally-invented exit codes. Every entry records code, name, meaning, owning module and the decision that authorized it. The generator refuses to build a table where two entries claim one code, two claim one name, a code falls in a range Node or the shell reserves, 2 is claimed by anything but the hook adapter, or an allocation carries no justification. exitCodeFor is pure and total: it throws rather than returning undefined, including for prototype-chain names. Inert by design — nothing emits a registered code until #3906. Every registered code is non-zero, asserted over the whole table, so a caller testing for failure behaves identically for pass and trips for everything else. * feat(#3905): make the registry generator's failures machine-readable Adds a --json mode carrying {ok, reason, context, detail}, where context is a typed payload naming the specifics the prose embedded - which code collided and under which names, which band rejected a code, which field was missing. The tests now assert on that structure instead of regex-matching the generator's stderr, which CONTRIBUTING prohibits, and the CONTEXT.md glossary gains the entry the issue's scope requires. * test(#3905): refresh the install-tree fixtures for the new declaration The registry declaration ships in the install tree, so all 19 golden fixtures needed regenerating. Caught by the remote matrix, not by lint:ci - the install-tree goldens are verified by a test rather than a lint, so a newly shipped file clears every local gate and fails only under the suite. * chore(#3905): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
c933184b97 |
enhance(#3172): require a stated failing direction for every automated acceptance command (#3825)
* test(#3172): failing-first suite for the stated failing-direction probe Pins the <fails_when> pairing walk, placeholder denylist, MISSING sentinel exemption, degraded-read contract, CLI arm and the plan-authoring contract text. RED by construction: the module exports it requires do not exist yet. Executed on the remote runner. * feat(#3172): require a stated failing direction for every automated acceptance command Every runnable <automated> command now carries a <fails_when> sibling naming what output constitutes failure. A command with no expressible failure mode is not an acceptance test: it reads as rigour and is not falsifiable. - verify-command-grounding gains a failing-direction probe sharing the existing <automated> grammar, MISSING sentinel and walk guard rather than copying them - gsd-tools check verify-failure-directions <N> backs it; plan-phase dispatches it and hands the JSON to gsd-plan-checker check 8f - Dimension 8 detail extracted to references to stay under the agent size cap Verified on the remote runner. * fix(#3172): close four review findings in the failing-direction probe - MISSING_SENTINEL_RE matched an env-var assignment prefix (MISSING=1 cmd), so a real command was exempted from the new blocking gate. Tightened the SHARED constant rather than adding a second copy. - Both token regexes scanned to EOF on unclosed openers (O(n^2), 1562ms at 40k). Bodies are now non-crossing; 1ms, byte-identical on well-formed input. The pre-existing AUTOMATED_BLOCK_RE carried the same defect and is fixed here too. - probePhaseFailingDirections reported status 'ok' when one plan was unreadable, conflating 'could not look' with 'nothing to report'. - Extracted the phase-resolution block both check arms had copied verbatim. Also corrects a docs/AGENTS.md dimension list stale since #2401. Verified on the remote runner. * fix(#3172): project the planner rule onto the spawn contract, settle emitted bookkeeping The remote runner refuted the planner-side edit. agents/gsd-planner.md is frozen under a 49152-LF-char cap asserted by four suites and sat at 49,146 — six chars of headroom — so the +537 of authoring rule blew it. #3297/#3645 already settled where such a rule goes: the planner spawn contract in plan-phase.md, beside <tracked_source_paths>. The agent file is reverted to origin/next verbatim. - plan-phase.md gains <failing_direction_contract>; tests row 30 now asserts the contract there and row 30b guards the freeze in both directions - plan-phase.md growth acknowledged by APPENDING to the 3409 fragment, per the precedent that two ack sources may never name the same path - install-tree fixtures regenerated for the three new reference files Verified on the remote runner. * chore(#3172): backfill PR number into the changeset fragment pr:0 -> pr:3825 now that the PR exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
14679b866b |
enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability Binds the approved triage shape before any of it exists: - containment — the execute:wave:post hook must not render unless workflow.live_dom_uat is true AND the capability resolves active (fail-closed on a missing state entry, and on a non-boolean value) - criterion 4 — agents/gsd-executor.md carries no browser MCP family; asserted as an absence, which is the only way it is observable - Hyrum guard — the pre-existing mcp__playwright__* branch must stay outside the key-gated block, or upgrading silently removes working automated UI verification for every current Playwright-MCP user - parity — the browser glob list now lives in two surfaces (agent frontmatter + workflow detection block); the assertion fails if either gains or loses a family without the other Red by construction: the capability, agent and workflow block do not exist yet. Verified on the remote runner. Refs #2856 * enhance(#2856): add default-off live-DOM UAT capability A phase whose acceptance criteria needed a live DOM could not be finished by the agent that executed it: gsd-executor carries no browser tools, so it correctly returned checkpoint:human-action even though the work was not human-only, just tool-less. Every such phase degraded to "executed, then finished by hand in the orchestrator", and autonomous: false could not distinguish "a human must judge this" from "the executor lacks the tool". Implements the shape approved at triage, not the one reported. The executor's tools: line is NOT widened, in any configuration: for a first-party agent the static list is the only control that exists (ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one default-off capability owns the key, the agent, and the step: - capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat (boolean, default false), one additive step at execute:wave:post (onError: skip, gates: []), so it can never halt a wave - agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP globs, in its own tools: line, with no Bash - verify-work automated_ui_verification — a gsd:live-dom-families block naming both new families AND the key; presence alone never activates Two independent fail-closed gates: isCapabilityActive renders a hook only on state.active === true, plus the step's own `when`. The pre-existing mcp__playwright__* branch keeps the gating it already had and stays outside the new block. Pulling it behind a default-off key would have silently removed working automated UI verification from every current Playwright-MCP user on upgrade. Also closes a host gap this surfaced: execute:wave:post dispatched only contribution + gate, so ANY registered step was declared and silently never run — exactly the single-kind hand-roll loop-hook-dispatch.md names. Step 5.75 now dispatches every kind == "step". The browser-profile lock is tolerated, not coordinated: --isolated is a flag on the operator's own MCP-server registration that GSD neither launches nor parameterizes, so the verifier reports could_not_look / profile_locked, names the flag, and stops. DOM-VERIFY.md keeps could_not_look and nothing_to_report distinct behind a closed reason enum — collapsing them is the ambiguous-run-notes defect reported. Verified on the remote runner. Closes #2856 * fix(#2856): apply review findings from the orthogonal passes Correctness pass (blocker): - delete detectionBlockIsCrlfSafe. It was pass-always: it read the file, replaced LF with CRLF, then indexOf'd marker strings that contain no newline, so the replacement could not change the result and the assertion could never fail for the reason it stated. There is no real CRLF risk on this surface either — the gsd:live-dom-families block has no parser, only human and agent readers. Deleted rather than replaced, per the repo's pass-always-test rule. Isolated security pass (two minors, both real): - execute-phase.md step 5.75: this change is what first activates kind == "step" dispatch at execute:wave:post, which newly opens the ref.command shell path at that loop point. Our own step uses ref.agent and never touches it, but the door is now open, so the step-dispatch line carries the same in-context validate-before-shell warning the sibling gate-dispatch line directly below it already carries. - gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker influenced. Require it wrapped in inline code or a fence, kept short, and never left reading as a directive to the next reader. Verified on the remote runner. Refs #2856 * fix(#2856): settle the new-agent roster ripple Checkpoint 2 returned 28 failures, none in the new suite — all of them the guards that exist to make adding an agent a deliberate act. Each is a real boundary that had to move: - docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526), so the browser globs lose their backticks; primary-agent counts 21->22, roster 33/34->34/35, Verifiers category 1->2 - docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md to be classified exactly once - gsd-dom-verifier: add the anti-heredoc instruction and the commented hooks: frontmatter pattern both agent gates require - gsd-core/bin/shared/model-catalog.json: every shipped agent needs a profile entry (#3229) - copilot-install / kilo-upgrades / qwen-upgrades: expected agent list and the 34->35 roster boundary - execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately carries one step now. Asserted as an exact shape — one step, capId live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a real guard against accidental change rather than being relaxed Two findings worth naming: mcp-tool-inheritance (#2526) rejected the agent for documenting mcp__playwright__* while its tools: line withholds it — a dead instruction that invites the agent to claim a path it cannot take. The prose now names the Playwright MCP family without the dispatchable token, in both the agent and the capability fragment. runtime-launcher-parity rejected the new gsd_run call: each fenced block is its own shell, so a workflow step file invoking gsd_run needs its own canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs. That script also normalizes explore.md, which is unrelated pre-existing drift the parity check tolerates, so it is reverted to keep this diff scoped. The emitted-drift ack supersedes the spent #3370 entry for execute-phase.md — it is merged into next, so its ripple is absorbed at the base and it can no longer clear anything. That is the same supersede the #3370 entry itself performed on the spent #3324 fragment. Its unrelated execute-plan.md entry is untouched. Verified on the remote runner. Refs #2856 * fix(#2856): drop the stale emitted-drift ack entry The automated-ui-verification.md entry was written speculatively rather than from a reported growth, and the check names that precisely: an ack "written or reworded in THIS diff, but nothing here needed it, so it explains nothing". The growth tier keys on the bare filename as it appears under gsd-core/workflows/ or agents/. automated-ui-verification.md is nested under verify-work/steps/, so it was never in the tracked set — only execute-phase.md was ever reported, both before and after the launcher preamble landed. Only ack what the check actually reports. Verified on the remote runner. Refs #2856 * chore(#2856): backfill changeset pr number pr:0 -> 3716. The placeholder fails both changeset-lint (fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by design and can only be resolved once the PR number exists. Both now report ok against GITHUB_BASE_REF=next. Refs #2856 --------- Co-authored-by: sim <sim@local> |
||
|
|
ea594300d9 |
fix(#3606): validate hook-kind coverage at call sites and dispatch generically (#3687)
* test(#3606): pin hook-kind coverage in the wired guard * fix(#3606): validate hook-kind coverage at call sites and dispatch generically * fix(#3606): address review - segment-granular narrowing, zero-coverage diagnosis, quick.md, fragment extraction * fix(#3606): drop stale shrink-ack, export HOOK_GROUP_KINDS, dedupe scanner regex * chore(#3606): regenerate install-tree fixtures for new wave-post fragment * chore(#3606): sync canonical launcher preamble into new fragment * fix(#3606): keep fragment preamble ahead of first gsd_run mention * fix(#3606): revert sync script's preamble move in explore.md * chore(#3606): regenerate derived manifests post-rebase * chore(#3606): allowlist peer test files - base was red on the count lane * chore(#3606): regenerate inventory for peer's verify-command-grounding doc * chore(#3606): grounding test maps to its own module by longest prefix * chore(#3606): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
79781e68eb |
enhance(#2401): ground verify-command paths and inherit prior-phase commands (#3678)
* feat(#2401): ground <automated> verify-command paths and inherit prior-phase commands Adds a deterministic resolvability probe over each PLAN.md <automated> verify command and surfaces the nearest prior phase's proven commands to the planner at every context window. - src/verify-command-grounding.cts: recognizer (not a shell interpreter) that grounds a leading cd <literal> chain and npm --prefix <literal>, and reports unresolvable rather than guessing. Never executes command text. - gsd-tools check verify-command-paths <N>: per-phase probe, wired into plan-phase.md before the plan-check pass. - init.plan-phase gains prior_verify_commands, ungated by context_window. - gsd-plan-checker: new Verify Command Path Resolvability dimension that reports the failing target and never prescribes a replacement. Also fixes first-match-wins prefix bucketing in scripts/lint-test-file-count.cjs (readdir order is not stable across platforms, so a module whose name extends another's with a hyphen bucketed differently on Linux than on macOS). Closes #2401 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2401): ground the canonical --prefix form, quoted paths, and absolute cd resets Independent review found three defects in the recognizer: - npm --prefix DIR run SCRIPT never reached the script-existence check, because the pattern required npm and run to be adjacent. That is the form the docs tell planners to prefer, so script_missing never fired for it. The prefix flag and its value are now stripped before matching. - --prefix captured with \S+, so a quoted path containing a space was truncated to a stray opening quote and reported as a missing directory - a false blocker, worse than the bug this feature fixes. The capture is now quote-aware. - A chained cd whose later segment was absolute concatenated instead of resetting, producing a nonsense path and another false blocker. The fold now resets on an absolute segment. Also replaces the bespoke phase-directory regex with the canonical phase-id helpers. Real phase directories are NN-slug, not phase-N-slug, so the prior-command harvest matched nothing outside its own fixtures and the planner-inheritance half of this feature was dead code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(#2401): source task blocks from the canonical sectionizer The module carried its own copy of the <task>-block grammar - a fourth hand-rolled mirror of the one markdown-sectionizer owns. verify.cts keeps its copy only because it needs the type= attribute the canonical helper discards; this module never reads that attribute, so it can share the owner outright instead of adding a test around a copy. extractAutomatedCommands now takes task bodies from extractTaggedBlocks and the out-of-task remainder from stripTaggedBlocks. A task-grammar parity test pins the attributed task-name set against the canonical helper across six awkward task shapes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2401): extract agent-file overflow to references and repair the property arbitrary The remote matrix run came back red with 19 failures, four root causes: - agents/gsd-plan-checker.md and agents/gsd-planner.md both blew the 49152 agent cap. Their bodies move to gsd-core/references/, leaving @-reference stubs, per the documented overflow pattern. - The new checker dimension invoked gsd_run before the canonical preamble that defines it. The call is deleted outright: plan-phase.md already runs the probe and hands the result in as {VERIFY_PATHS}, so the dimension consumes that rather than re-running anything. - fc.fullUnicodeString does not exist in fast-check 4.8.0. Replaced with fc.string({ unit: 'binary' }), which covers the same 0000-10FFFF range. - Three runtime-loaded files grew; acknowledged in the existing ack fragments that already own those bare filenames, since two ack sources may never name the same path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2401): regenerate golden install-tree fixtures for the new references Adding two files under gsd-core/references/ changes what the installer emits into every runtime's tree, so all 19 golden install-parity fixtures went stale. Regenerated with npm run gen:install-tree; the delta is exactly the two new reference paths per runtime, no removals. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2401): backfill changeset pr number to 3678 * fix(#2401): treat ~ as a home expansion only at the start of a path Windows CI caught this on both shards; the Linux-only remote matrix cannot see it. The dynamic-path refusal rejected ~ anywhere, and a GitHub Windows runner's tmpdir is an 8.3 short name - C:\Users\RUNNER~1\AppData\Local\Temp - so a valid absolute Windows path came back unresolvable/dynamic_path. This was a production bug, not a test artifact: any Windows user whose project path carries an 8.3 short name, or any literal ~, silently lost the probe entirely - every command degrading to unresolvable with no explanation. ~ is a home expansion only at the start of a path; elsewhere it is an ordinary literal. The check is now split: $, backtick, *, ? and newline stay refused anywhere (substitution and globs, and the glob characters are illegal in Windows path components regardless), while ~ is refused only leading, tolerating one leading quote since the check runs before quote stripping. The prior tests only caught this on Windows because only Windows puts a ~ in tmpdir. Four new tests pin it on every platform via a fixture directory literally named RUNNER~1. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bf2332e67c |
fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in its fail-closed catch and is misreported as an unreadable dispatch-isolation configuration, blocking every executor dispatch. These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor guard and update worker, plus the seam's actionable build error surfacing instead of the generic misreport. Also adds the drift lint the acceptance criteria require, with a fixture proving it CAN fail — a guard never shown to fail is worthless. It is red here by design: it flags today's unfixed hooks, which is exactly the defect. * fix(3582): route every hook's compiled-module require through the self-heal seam RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard, cursor guard and update worker, the fail-closed-with-actionable-message assertion, and the lint's own real-tree check. The compiled runtime library is produced by build:lib and gitignored (ADR-457, build-at-publish). The npm package builds before publishing; a plugin-marketplace or git-clone install materializes the raw tree and never does, so on that channel those modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly this, and the CLI entrypoint already calls it — no hook did. The isolation guard's Cannot-find-module therefore landed in its fail-closed catch and was reported as 'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was misdiagnosed as an unreadable project config and every executor dispatch was blocked. All SEVEN affected files now call the seam before their first compiled require. The issue named four; a scan found six; implementing it surfaced a seventh — the shared isolation sentinel helper, used by BOTH guards, which requires two compiled modules itself and would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so fixed here rather than left as a known-broken remainder. Failure posture is deliberately split by hook kind: - Gates (agent isolation guard, cursor subagent start) surface the seam's actionable build error distinctly instead of swallowing it into the generic text, and stay fail-closed — a genuinely unreadable project config still DENIES exactly as before. - Cosmetic and detached hooks (statusline, update worker, update check, update banner) DEGRADE rather than crash: the statusline draws on every render and the worker is a detached process, so a build failure there must not take down the prompt. The npm path is untouched: the seam's already-built fast path returns immediately, so prebuilt installs pay nothing and behave bit-for-bit as before. Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than remembered — without it the next hook to add a compiled require reintroduces the class silently. It is proven able to fail: a fixture hook requiring a compiled module without the seam is flagged, and one that uses the seam is not. Verified directly — on the unfixed tree it named all seven offenders; with the fix it passes. While writing the lint's comment stripper, a naive whole-text block-comment regex ate its own fixture, because this repo's comments legitimately spell the compiled-lib glob whose star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression test pinning that case. * fix(3582): test the three untested seam call sites and assert typed reason codes Two independent reviews converged on the same major gap: the fix wired the seam into seven files but only four had cold-tree tests. The adversarial pass put it plainly — deleting the shared isolation-sentinel helper's seam call would not have failed any test in the diff. That file was my own addition beyond the issue's four, so it shipped untested; that is now closed. - Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT directly under cwd, and every existing cold-tree fixture puts it there, so the early return always fired first. Now covered, and proven load-bearing by mutation: with the call removed the spy records zero seam invocations and the test fails. - update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED VERDICT — the fallback cache filename, and silent suppression when the package name degrades to null — rather than merely 'did not throw'. The banner hook previously had no test file at all. Standards violation fixed: two tests asserted on free-form prose via assert.match against a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch it. Both isolation guards now emit a machine-readable reason_code from a frozen enum, following the repo's existing REASON convention, and the tests assert that instead. The human-readable message is unchanged for operators; only the assertion target moved. The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT extracted, and the reason is recorded at each site: both viable shapes — a path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file literal co-occurrence check, so extracting would require the lint to special-case its own helper. Triplication is the lesser evil while the lint stays a co-occurrence scan. The lint's header now states what it does and does not catch (literal quoted requires only; hooks/ scan root), so a future reader does not over-trust a guard that a concatenated path or a require inside a non-hooks helper would evade. * chore(3582): regenerate the committed install-tree fixtures Adding a new shipped hook helper changed the install tree, and those fixtures are committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>' tests failed on 541a1913. Regenerated rather than hand-edited. The delta across all 15 runtime fixtures is exactly two lines — the new helper under both its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in no unrelated drift. This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from lint:ci, which passed both before and after. * chore(3582): backfill changeset PR number (#3629) --------- Co-authored-by: sim <sim@local> |
||
|
|
9448736872 |
fix(#3547): exercise the real global config-home shape in the install harness (#3567)
* test(#3547): failing-first regression for collapsed global install shape * fix(#3547): exercise the real global config-home shape in the install harness * test(#3547): align ripple suites with the real global install shape * test(#3547): update stale collapsed-shape pins in provenance and migration suites * fix(#3547): bump emitted-baseline schema version for the real install shape --------- Co-authored-by: sim <sim@local> |
||
|
|
c5b83cb050 |
chore(#3560): delete two unreachable workflows, gate workflow reachability in lint (#3564)
* chore(#3560): delete two unreachable workflows, gate reachability in lint discovery-phase.md and plan-milestone-gaps.md shipped to all 19 runtime install trees with no command, agent, or skill referencing them. plan-milestone-gaps' command was deleted by #2790 and the workflow was left behind; discovery-phase's own header claimed a caller in plan-phase.md's mandatory_discovery step, and that step does not exist — plan-phase.md contains zero occurrences of "discovery". docs/INVENTORY.md asserted discovery-phase.md was an alternate entry for /gsd-new-project. new-project.md never referenced it. The row and the matching note sentence are removed across all five locales rather than corrected. Adds rule 6 to lint-command-contract: every shipped workflow must be reachable from a loader, walking the transitive closure over the three reference shapes this repo uses. The closure seeds ONLY from commands/agents/skills, so a workflow that references only itself and a pair that reference only each other are both correctly reported rather than satisfying themselves; a visited set makes reference cycles terminate. The measure is a mention in a LOADER — docs/ and install-tree fixtures deliberately do not count, because scan.md proved a file can be documented and shipped while entirely unreached. Ships blocking, not report-only: #3561 is in this branch's base, so the tree reports 0 unreachable from the start. Closes #3560 * test(#3560): drive rule 6 end-to-end, sweep a stale allowlist, update ADR-0002 Review findings. Rule 6 had no end-to-end coverage: the tests exercised the pure closure with in-memory data, so the wiring — file collection, exit code, diagnostic — was unproven, and #3560's acceptance list explicitly wants a fixture showing the rule FAILS on a planted orphan. Adds an optional --root to lint-command-contract (default behavior unchanged) and four tests driving the real CLI through the process seam against a temp fixture: clean=0, planted orphan=1, orphan referenced only from docs/=1, orphan reachable transitively=0. The docs/ case is what pins the Goodhart defense — a mention outside a loader must not confer reachability. Deletes two tests that were byte-identical to a third and could not assert anything loader-specific, since the closure is source-agnostic by design; that distinction lives in the lint script's file collection and is now covered above. Removes a stale ALLOWLIST entry for discovery-phase.md in planner-language-regression — the exact sweep-miss class rule 6 exists to catch, found in the PR that adds the rule. ADR-0002 described five per-file frontmatter checks; rule 6 is a repo-level reachability graph, so the Decision section now says so. Refs #3560 * test(#3560): cut the bug-3298 test pin on the deleted plan-milestone-gaps workflow The remote runner went red with four failures: tests/phase.test.cjs asserted the plan-milestone-gaps workflow exists and checked its mkdir patterns, so deleting the file broke the test that pinned it. This is the fence the epic describes — the content-sync test IS what keeps an unreachable file alive — and cutting the coupling is what makes the deletion safe. Removes only that arm. The bug-3298 block guards three workflows against phase-dir prefix drift; the import and add-backlog arms and both shared mkdir-pattern helpers are untouched. Worth recording where the sweep failed: my reachability walk covered commands, agents, skills, gsd-core and docs, and lint-removed-but-needed covers .github/workflows, gsd-core, docs and package.json. Neither looks at tests/, so a test-pinned deletion is invisible to both and surfaces only on the remote runner. The how-to added by this PR names that gap explicitly so the next deletion searches tests/ by hand. Refs #3560 * docs(#3560): add a how-to for resolving unreachable-workflow findings * chore(#3560): backfill changeset pr number to 3564 --------- Co-authored-by: sim <sim@local> |
||
|
|
7976b1ca0d |
feat(#1689): per-plan agent_hint executor routing (#3417)
* feat(#1689): per-plan agent_hint executor routing Option A per-plan specialist routing: a plan with an `agent_hint:` frontmatter field is dispatched to that subagent instead of gsd-executor when it resolves on the active runtime; absent/unresolved/disabled falls back to gsd-executor (byte-identical). Default-on via workflow.agent_hint_routing. - src/phase.cts: parse agent_hint into the plan-index JSON (plan_json.agent_hint) - agent-install-check.cts: resolveAgentHint() reuses getAgentsDir + runtime filename variants; probes project + global agent dirs; fails closed; rejects path-traversing names - gsd-tools.cjs: 'resolve-agent' query route (fail-closed to gsd-executor; --raw/--json) - execute-phase.md: lean per-plan reference + {EXECUTOR_TYPE} placeholder (host stays under the ADR-857 Phase 6 byte ceiling) - execute-phase/steps/per-plan-executor-routing.md: resolution logic (Agent()-based dispatch; advisory on orchestrator-worktree) - config: workflow.agent_hint_routing (validKey, default-on via SCHEMA_DEFAULTS, boolean validator) - docs (CONFIGURATION.md, plan-md.md), changeset, tests/agent-hint-routing-1689.test.cjs (17 tests) * chore(#1689): backfill changeset PR number (#3417) * chore(#1689): regenerate install-tree fixtures for new workflow fragment * chore(#1689): ack deliberate execute-phase.md growth (agent_hint routing) * test(#1689): SPAWN contract allows parameterized subagent_type placeholder agent-frontmatter's spawn-type checks scanned subagent_type="..." as a concrete agent name. execute-phase now uses subagent_type="{EXECUTOR_TYPE}" (a runtime placeholder resolved via resolve-agent, default gsd-executor). Skip {TOKEN} placeholders in both the known-type and <available_agent_types> checks; execute-phase still lists the built-in roster incl. gsd-executor. * fix(#1689): CI conformance for the routing fragment - per-plan-executor-routing.md: add the canonical runtime-launcher preamble to its gsd_run block (runtime-launcher-parity #373), matching sibling step fragments. - agent-install-check.cts: drop a literal ~/.claude/agents path from the resolveAgentHint JSDoc so it does not leak into the compiled engine .cjs (cline install leak guard). --------- Co-authored-by: sim <sim@local> |
||
|
|
d30c99bc92 |
chore(#3421): delete orphan verify-phase workflow, migrate live gates to verifier (#3422)
* chore(#1892): delete orphan verify-phase workflow, migrate live gates to verifier reference * test(#1892): retarget structural suites from verify-phase.md to verifier-phase-gates.md * chore(#1892): reword retired-workflow mentions for removed-but-needed lint * test(#1892): correct stale surface labels in retargeted suites * docs(#1892): add verifier-phase-gates row to locale inventories * chore(#3421): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
2076d450d7 |
fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not the runtime name (#2728)
* fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not runtime name quick.md and diagnose-issues.md kept the pre-#2584 `RUNTIME != "claude"` worktree gate, so every non-Claude runtime failed closed regardless of the capability it negotiated — including Codex, which declares orchestrator-worktree. Route both through the negotiated dispatch.isolation seam via a new shared reference, and migrate the two execute-phase reference fragments that carried the same runtime-name gate. - new gsd-core/references/dispatch-isolation-gate.md: canonical ISOLATION resolution, harness-flag resolution, single-agent degrade rule - quick.md / diagnose-issues.md read the gate; dispatch uses the {harnessFlag} placeholder rather than a hardcoded isolation="worktree" - execute-phase-wave-guard.md / execute-phase-between-wave-reset.md: migrate [ "$RUNTIME" = "claude" ] -> [ "$ISOLATION" = "harness-worktree" ] - every degrade site now clears BOTH USE_WORKTREES and ISOLATION; clearing one dispatched an isolated agent with no base guard and no manifest - parity guard in host-integration.test.cjs scans workflows AND references and matches six reintroduction shapes - migrate four tests that pinned the pre-#2584 runtime-name contract Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): use the /gsd:<cmd> namespace in the isolation degrade messages The degrade warnings cited /gsd-execute-phase, the retired hyphen form that slash-command-namespace.test.cjs rejects in Claude-facing source. Same length, so the quick.md size budget is unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2652): add changeset for PR #2728 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): normalize dispatch-site paths to forward slashes for Windows path.relative() returns backslash-separated paths on Windows, so the #2652 dispatch-site parity test compared "gsd-core\workflows\quick.md" against the hardcoded forward-slash literal "gsd-core/workflows/quick.md" and failed on every windows-latest CI lane. Normalize with .replace(/\\/g, '/'), matching the existing convention used elsewhere in this suite (e.g. tests/branch-no-track-guard.test.cjs:37). * test(#2652): restore the size-growth acknowledgment The rebase dropped tests/emitted-drift-ack.json. #2757/#2758 fixed the ATTRIBUTION axis, but the SIZE-GROWTH axis is independent: diagnose-issues.md (+2086) and quick.md (+230) still need an ack naming them and saying why. Verified: 65/66 without it (both files named), 66/66 with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): convert execute-plan.md Pattern A onto the dispatch-isolation gate Pattern A hardcoded `isolation="worktree"` — Claude Code's own literal — gated only on `workflow.use_worktrees`, with no capability negotiation at all. It is the same defect #2652 fixes at the other four sites, just a different shape: the file contains no RUNTIME variable, so the new detector correctly does not flag it. Concrete break: a Codex user who follows this PR's own newly-documented pattern and sets `workflow.use_worktrees: true` to get isolated dispatch via /gsd:quick then runs a plan through /gsd-execute-plan Pattern A, and hits an unconverted path — either an Agent() call erroring on an unrecognized parameter or silent unisolated execution, depending on host tolerance. Pattern A is a single-agent dispatch site through the host's own subagent tool, so it takes the same treatment as quick.md and diagnose-issues.md: resolve ISOLATION/HARNESS_FLAG through the canonical reference, degrade to sequential on orchestrator-worktree hosts, and substitute the host's declared {harnessFlag} instead of Claude Code's literal. while the area was open. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2652): add the INVENTORY row for dispatch-isolation-gate.md, refresh CONTEXT Two bookkeeping gaps flagged in review: INVENTORY.md had no row for the new gsd-core/references/dispatch-isolation-gate.md. INVENTORY-MANIFEST.json was regenerated correctly and its --check only diffs a live directory scan against the committed manifest, so CI passed regardless — but gen-inventory-manifest.cjs's own stderr guidance says to add the matching INVENTORY.md row. This is the repo's named "Inventory Drift" pattern. Placed with the dispatch/isolation cluster (worktree-branch-check, runtime-aware-dispatch) rather than alphabetically, matching how that table is grouped. CONTEXT.md's Host-Integration Interface entry still described dispatch.isolation as "declared and negotiated but not yet consumed by any scheduler — Phase 1 of #2584". That was already stale before this PR (execute-phase graduated in Phase 3) and more so now with three single-agent dispatch sites consuming it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): detect reversed-operand runtime gates; add a permutation property All five reintroduction regexes assumed $RUNTIME on the LEFT of the comparison, so `[ "claude" != "$RUNTIME" ]` — the same gate written backwards — evaded every one of them. Verified against the old patterns before fixing: all four reversed shapes (single bracket, double bracket, test builtin, JS template) scored EVADED. Each comparison shape is now generated in both operand orders from a single template, so a shape cannot be added in one order and forgotten in the other. The mutation table gains the four reversed cases. Also adds the fast-check property review suggested in place of the hand-rolled cases: it generates the cross product of the axes an author actually varies — bracket form, operator, operand order, quoting, spacing, runtime id — so a permutation the hand-written patterns miss surfaces here rather than in production. The 11 explicit cases stay as named regression anchors. execute-plan.md joins the scan's required-identities list now that it is a converted dispatch site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): acknowledge the execute-plan.md size growth The Pattern A conversion adds 811 bytes to an emitted workflow. Per #2719 the size axis needs its own acknowledgment, independent of attribution. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): repin the execute-plan.md PROSE_ALLOWLIST line after the rebase The #2751 command-position gate pins its prose exemptions by line number. This branch inserts the dispatch-isolation resolution above the `validated downstream by gsd-tools uat classify-coverage` sentence, moving it from execute-plan.md:387 to :397 — which fired the gate twice for one displacement (an un-allowlisted mention at 397, a stale entry at 387). The prose itself is unchanged from next; only the pin moves. Fixes #2652 * fix(#2652): gate the #2649 base-check on ISOLATION in diagnose-issues.md The rebase onto next merged #2649's pre-dispatch base-check textually, but its degrade flipped USE_WORKTREES after ISOLATION was already resolved, so the degrade never reached the dispatch decision. Gate the block on ISOLATION = "harness-worktree" and degrade ISOLATION itself, the same pairing quick.md already uses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): key quick.md post-dispatch bookkeeping on ISOLATION, not the Claude literal Review Blocker: the manifest append (l.822), worktree merge-back (l.825), and its skip clause (l.839) all conditioned on the literal isolation="worktree" — Claude Code's own rendering of {harnessFlag}. Cursor renders --worktree, so a newly-unblocked isolated Cursor run created a worktree whose committed work was never merged back and never cleaned up, silently. All three now key on ISOLATION = "harness-worktree" at dispatch. The existing parity detector cannot catch this class (its ISOLATION_TOKEN treats the literal as a legitimate marker), so this adds a dedicated literal-condition detector with a discrimination proof against both pre-fix sentences, a benign-mention control, and a positive pin on all three re-keyed conditions. Verified fail-first against the pre-fix quick.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): scope the use_worktrees=false install stamp to isolation=none runtimes `_stampNonClaudeRuntimeDefaults` rewrote every non-Claude runtime's `workflow.use_worktrees` read to `--default false`. That default resolved before `gsd_run query dispatch-isolation` was ever consulted, so the five runtimes that declare worktree support — cursor (harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree) — got ISOLATION=none regardless of what they negotiated. The gate this PR migrates dispatch onto was therefore still deciding isolation by runtime name, one layer down. The stamp's #1521 premise was that worktree isolation *was* Claude Code's isolation="worktree" spawn parameter, which no other host honored. #2584 replaced that premise with the negotiated capability. The stamp is now scoped to runtimes whose negotiated isolation really is `none`, where the default it writes is the outcome the resolver reaches anyway. `_negotiatedDispatchIsolation` mirrors routeDispatchIsolation's resolution against the same registry — closed vocabulary, a harness-worktree host must declare its flag, an orchestrator-worktree host must carry a descriptor that resolves — and fails closed to `none` on anything else, so an undeclared or unknown runtime keeps today's behavior. Two #1515 tests pinned the superseded premise for codex and are re-pointed at the new contract rather than deleted: the safety property they protect is now held by the isolation gate's fail-closed resolution, not by a name-scoped install-time default. Verified fail-first — all five assertions red against the pre-fix source, green after. * test(#2652): acknowledge the emitted ripple and re-point the end-to-end stamp proof Scoping the use_worktrees stamp changes emitted output, and two gates caught it. `gsd-core/workflows/execute-phase.md` now differs at emit time for the five hosts that declare worktree support (cursor harness-worktree; codex, opencode, kimi, kimi-code orchestrator-worktree) — the source file is byte-identical, only the stamp is gone. Acknowledged in this PR's fragment. `tests/install.test.cjs`'s real-install assertion pinned the superseded premise end-to-end, asserting codex receives `--default false`. Re-pointed rather than deleted, matching the two unit tests: it now proves codex keeps the unstamped `true` read. A second arm installs windsurf — which declares isolation `none` — and asserts the false stamp is still applied there, so the change cannot silently degrade into "never stamp" without a test noticing. The ack entry collides with `2658-trae-instruction-file-path.json`, which is fully spent (merged via #2925, so all 25 of its entries are present at base and gate nothing) and is pruned for the same reason and by the same rule as the spent `2649-*` fragment this PR already removed. #2566 prunes the same file for the same collision on `new-project.md`; a delete/delete merges cleanly either way, and the base-side cleanup would make both unnecessary. * fix(#2652): re-record the sentinel when a dispatch site degrades isolation Review Blocker B1/B2/B3. Every isolation degrade in a dispatch site is decided in shell, where routeDispatchIsolation cannot see it. That resolver persists whatever it resolved to the run-scoped sentinel as an unconditional side effect (#3045), so a degrade that only reassigns $ISOLATION leaves the sentinel asserting harness-worktree while the dispatch correctly omits the harness flag. The shipped PreToolUse guard reads the sentinel at the instant of the Agent() call and denies that mismatch with exit 2 — the work does not run unisolated, it does not run at all. Latent on this branch and lands on rebase, since |
||
|
|
4b69dc346b |
fix(#2725): repoint the pre-commit alias guard at sources git can actually stage, drop nine dead ones (#3273)
* test(#2725): failing-first coverage for the inert pre-commit alias guard Replaces two stale tests that asserted on `sdk/src/query/command-manifest.phase.ts` — a path retired with the SDK boundary (ADR-0174), so both passed forever while guarding nothing. The new matrix drives .githooks/pre-commit through its GIT_OVERRIDE/NPM_OVERRIDE seams and asserts on the real tracked sources the drift checker reads. Red until the guard is repointed off the gitignored build outputs it currently watches. * fix(#2725): repoint the pre-commit alias guard at sources git can actually stage `.githooks/pre-commit` carried ten staged-path guards and not one of them could do its job. Nine anchored on `^sdk/…`, a tree retired by ADR-0174, and invoked npm scripts that no longer exist (`check:state-document-fresh` and eight siblings). Their only reachable behavior was to abort the commit with `Missing script` — which required matching a path that cannot exist, so they were dead twice over. The tenth was the real defect. Its npm target does exist, but it matched `^gsd-core/bin/lib/command-aliases\.cjs$` — a gitignored build output (.gitignore:172). An ignored path never appears in `git diff --cached --name-only`, so the guard was not stale, it was structurally unmatchable: it watched the derived layer instead of the source layer, from the day it was written. Repointed at the nine tracked `src/*.cts` sources `scripts/check-alias-drift.cjs` actually reads. The family table moves to `scripts/lib/alias-drift-families.cjs` so the checker and the hook derive their surface from one list, and a new parity test fails if a family is added to the checker without the hook learning to watch its source — the rot mechanism, not just this instance of it. Matching is now `grep -Fxqf` (fixed strings, whole line): exact path equality, no regex anchors to get wrong as the list grows. Staged paths are collected into a variable before matching, because `grep -q` exits on first match and would SIGPIPE its upstream, which under `set -o pipefail` turns a successful match into a non-zero pipeline status. The two CONTRIBUTING.md recipes pasted copies of the hook bodies inline — a third parallel surface, and one that had already drifted: the pre-commit copy carried the same dead `sdk/` patterns, and the pre-push copy would have overwritten the committed hook with a paraphrase that drops the GIT_OVERRIDE seam its test drives. Both now just point git at the committed files. The pre-push recipe's `'@example-corp\\.com$'` also never matched anything — inside single quotes bash keeps both backslashes. `.githooks/pre-push` was audited for the same rot and has none: it keys on no paths, no-ops unless GSD_BLOCKED_AUTHOR_REGEX is set, and is covered. Unchanged. Hooks stay opt-in. Nothing registers `core.hooksPath` for you, per CONTRIBUTING.md's documented one-time setup. * test(#2725): make the hook/checker parity assertion bidirectional Review finding from the isolated adversarial pass: the parity row only caught the hook UNDER-watching relative to scripts/lib/alias-drift-families.cjs. Drop a family from the module and the hook would keep watching its source with nothing to notice — the same divergence class this change exists to close, just pointed the other way. The new row takes every `src/*-command-router.cts` on disk as the universe and asserts the hook stays silent for the eight routers the drift check does not read. Both directions are now covered by running the real hook, not by comparing two lists in the test. Also corrects a CONTRIBUTING.md overclaim caught by the standards pass: 9 of the 11 watched paths derive from the module, not all 11 — bash cannot require a CJS module, so the hook carries literals and the test is what binds them. * fix(#2725): ship the shared family table and fix the mock that hid its own rows Three defects the remote runner caught that local probing did not. The mock `git` in the regression test emitted its staged-path payload as `printf '%s' "src/command-aliases.cts\n"`. Bash does not expand `\n` inside a double-quoted string and printf does not expand escapes in a `%s` argument, so the mock produced one unterminated line containing a literal backslash-n. No whole-line match could ever succeed, and every row that expects the hook to FIRE failed while every row that expects silence passed — which is exactly the signature the run reported: 9 failures, all of them fire-expecting rows. The payload now goes through a file the mock `cat`s, which is byte-exact and is what makes the CR-terminated and empty-staged-list rows mean what they say. `scripts/lib/alias-drift-families.cjs` was not enumerated in `GSD_SCRIPTS_LIB_FILES` (bin/install.js:377), so the installer never copied it. That is not cosmetic: `scripts/check-alias-drift.cjs` ships, and it now requires this module — an installed tree would have failed with MODULE_NOT_FOUND the first time a consumer ran `check:alias-drift`. Added to the manifest, which is what the #3184 install/uninstall parity tests assert against `readdirSync`. Regenerated the 19 committed install-tree fixtures via `npm run gen:install-tree` to record the new emitted path. The diff is +1 line per fixture and nothing else. `npm run lint:ci` exits 0. The earlier claim that `scripts/` is outside the emitted surface was wrong: `scripts/lib/` is copied into every runtime's install tree, which is why 19 golden-install-tree cases moved. * chore(#2725): backfill changeset pr: 3273 --------- Co-authored-by: sim <sim@local> |
||
|
|
342590c70e |
refactor(#3184): milestone windowing has one owner and a decidable failure signal (#3209)
* test(#3184): failing-first milestone-window single-owner suite Covers the 50 input classes in the phase test matrix: scope classification (genuinely-empty vs truncated vs unscoped vs unreadable), the section-end owner's level boundaries, consumer-output identity per ADR-3180 Decision 4(c), the milestone.complete refusal with negative proof that no directory moved, the version-token boundary defect, drift-guard behavior, and three fast-check properties over document-shaped generators. Committed alone so the remote runner records the failure before the fix lands. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * refactor(#3184): milestone windowing routes through one owner Three copies of the milestone section-end walk lived in roadmap-parser.cts — two distinct computeSectionEnd function nodes plus an inline third in getMilestonePhaseFilter's versionOverride branch. computeMilestoneSectionEnd is now the sole owner and the other two are deleted, not kept in sync by comment. The whole-repo drift guard found what the epic did not: state.cts held three more re-derivations of the same vocabulary — two byte-identical milestone bounding checks carrying a defect neither reported copy has (no boundary after the version token, so v2.0 matched inside v2.0.1), and a milestone-sectioning predicate. All three route through the owner now. A composition-level duplicate appeared inside this change's own first pass: getMilestonePhaseFilter and cmdMilestoneComplete each re-assembled a window out of the owner's primitives, and had already diverged on whether to skip a closed milestone heading. sliceMilestoneWindow is the one composition. Windows now carry the ADR-3180 SCOPE discriminator, so a truncated window is distinguishable from a genuinely empty milestone — those were output-identical, which is the whole failure class. roadmap analyze emits it (#3165), and milestone complete refuses to archive on anything but COMPLETE rather than pass-all moving every phase directory on disk (#3166). The pass-all degrade is preserved where its premise holds: making the filter deny-all would trade a silent over-inclusive answer for a silent under-inclusive one on the read paths that count with it. extractCurrentMilestone keeps its signature — 200+ affected symbols across 41 files and 25 process flows — and is a one-line wrapper over the scoped owner. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): fence-aware phase detection and one heading-selection owner Review fixes from the two orthogonal passes. The blocker: hasPhaseEntries matched ATX phase headings fence-aware via tokenizeHeadings but tested the #2199 bullet form against un-stripped markdown, so a fenced EXAMPLE of the bullet syntax counted as a real phase. A genuinely empty milestone then classified TRUNCATED and milestone complete refused a legitimate archive — a false positive in the destructive direction, worse than the defect this phase set out to fix. Both that path and getMilestonePhaseFilter own pre-existing bullet scan now run on stripFencedCode, since leaving one meant the owner file gave two different answers to the same question. The selection rule — locate, prefer the non-closed heading, else the first — had been written three more times inside the file whose thesis is single ownership. selectMilestoneHeading owns it; all three sites route through it. The copies were behaviorally identical, so this is de-duplication with no observable change, verified by probing that all three paths select the same heading. roadmap analyze emitting a scope no consumer read left #3165's actual symptom alive, so Route 0 in next.md now treats a non-complete scope as scan-failed rather than as a clean empty scan, and the ADR amendment no longer overstates what shipped. Also: the scope refusal moved above the archive-directory create, so a refusal leaves nothing on disk; the versionOverride comment names all four consumers; COMMANDS.md documents the new guard beside its sibling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#2658): exclude the changelog from the malformed-path scan The gate walks every emitted .md/.js/.cjs file in an installed tree and asserts none contains `.claude/.trae/rules` or `.trae/.trae/rules`. CHANGELOG.md ships into that tree, and its #2658 entry quotes both malformed paths while describing the fix that removed them — so the release note documenting the fix trips the fix's own regression test. Red on next before this branch. The installer is correct: a probe over a real --trae --local install found 621 emitted files, exactly one hit, and it was gsd-core/CHANGELOG.md. The scan scope was the defect, not the product. Excluded by exact relative path rather than by loosening the patterns or skipping all markdown — the emitted agent and command markdown is precisely what #2658 was about, so the gate stays strong everywhere it matters. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): regenerate install-tree fixtures for the shared drift scanner scripts/lib/ ships in the npm package and installer, so extracting the shared tree-walk into scripts/lib/drift-scan.cjs adds one path to every runtime's install tree. Regenerated via npm run gen:install-tree; the delta is exactly that one path per fixture. The two drift guards themselves do not ship (scripts/lint-*.cjs is excluded), so only the extracted library moves. This matches the existing scripts/lib/allowlist-ratchet.cjs precedent, which is likewise a lint-only helper carried in the shipped tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): restore the #730 sub-milestone boundary and narrow the refusal The remote runner caught two regressions this branch introduced. Both were mine, and neither review pass found them — only running the existing suite did. The version-token boundary. I replaced locateMilestoneHeadings' \b with (?![\w.-]), reasoning that v2.0 matching inside v2.0.1 was the same defect #2562 fixed in isMilestoneShippedInRoadmap. It is not the same question. A milestone state of v8.0 legitimately selects the '## v8.0-B' sub-milestone section over a closed v8.0-A sibling (#730), and \b is what allows it while the stricter boundary forbids it — nine tests in roadmap-phase-fallback said so. Reverted to \b; the state.cts consolidation is now a straight merge with no behavior change, and the v2.0/v2.0.1 ambiguity is left exactly as it was. The ADR amendment and the design doc no longer claim otherwise. The refusal scope. I refused whenever the window was not COMPLETE, but #3166 is about the TRUNCATED window specifically — the heading is found and the section closes before the phase region, so pass-all archives everything. UNREADABLE and UNSCOPED are pre-existing, legitimately handled states, and refusing on them broke 'handles missing ROADMAP.md gracefully' and three archive tests. Narrowed to TRUNCATED; docs corrected to match. One of the new tests was also wrong: its fixture gave the shipped and current milestones' phases the same numeric id, and the filter matches on that id, so it could not have distinguished the two windows. Fixture corrected to exercise what it claims to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): enumerate drift-scan.cjs for uninstall The installer copies scripts/lib/ wholesale, but uninstall removes an explicit set — deliberately, so a user's own helpers in that directory survive. The extracted drift-scan.cjs was copied in and never enumerated, so it outlived uninstall, left the directory non-empty, and the rmdir that follows failed. Added to GSD_SCRIPTS_LIB_FILES, following allowlist-ratchet.cjs, which is likewise a lint-only helper that ships there and is enumerated. Verified with a real install-then-uninstall into a temp target: scripts/lib/ held exactly the three GSD files and was gone afterwards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): assert install and uninstall agree on scripts/lib and scripts/changeset Found while shipping this phase, and fixed here rather than noted. install() copies scripts/lib/ and scripts/changeset/ into the target WHOLESALE — the comment at the copy site literally says "and any future lib helpers". uninstall() removes them by hardcoded enumeration, deliberately, so a user's own helpers in those directories survive. A wholesale writer paired with an enumerated remover cannot stay in sync by construction: any file added to either directory ships to every user and is then orphaned in their repo forever, since it survives uninstall, leaves the directory non-empty, and the rmdir that follows fails. Nothing reported this. 31,225 tests were green over it. That is the same divergence class this epic exists to delete, sitting in the installer, so it gets the same remedy CLAUDE.md prescribes for it: a parity assertion that fails the moment the two surfaces disagree. The test compares each directory's real contents against its enumeration and names the offending file plus the constant to add it to. Both enumerations are hoisted to module scope and exported, so the test asserts on the actual arrays rather than pattern-matching the installer's source — no allow-test-rule annotation needed. Proven non-vacuous both ways: empty diff on the current tree, correct report when an unenumerated file is injected. scripts/changeset/ turned out to carry the identical defect and is covered too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * chore(#3184): backfill changeset PR number Also narrows the wording to match the shipped behavior: the refusal fires on a truncated window specifically, not on any non-complete scope. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8f75e27554 |
fix(#3045): fail closed when an executor dispatch drops its resolved isolation (#3069)
* feat(#3045): deny an executor dispatch that drops its isolation flag Every isolation gate already resolved correctly. The resolved value then reached the executor through a prose instruction telling the model to substitute it into a call the model composes itself, and nothing verified the substitution. When it was dropped, the executor edited and committed in the user's primary checkout with no consent and no warning. A prose backstop would be the same class of artifact as the defect, so this is a shipped PreToolUse hook on the Agent tool. It fires at the instant of the call rather than being read once at the top of a workflow, which is the only placement the model cannot skip. The guard is inert unless it can positively establish that this is a GSD project, that the project resolves to harness isolation, and that the dispatch targets an executor. A non-GSD repo has no invariant to enforce. Where it cannot read the configuration at all, it denies rather than assuming, with its own reason -- a guard that cannot verify must not answer safe. A malformed payload allows rather than throwing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3045): extend the isolation guard to Cursor Cursor is the second of only two runtimes that resolve harness isolation, so shipping the guard for Claude alone left half the exposed surface unguarded while the changeset implied it was covered. The two runtimes fail differently. On Claude the harness flag is a per-dispatch kwarg the model must copy into a call it composes, and the defect is that it can be dropped. On Cursor the flag is --worktree, which applies to the whole session, and the subagent-start payload carries no isolation field at all. There is no flag to check, so the guard verifies the effective state instead: whether the workspace is genuinely running outside the user's primary checkout. That is a stronger check than the Claude one because it tests reality rather than intent, and it is commented so nobody later rewrites it into a flag check. Isolation is established two ways, either sufficient: the workspace resolves to a linked git worktree, or it sits under the worktree root Cursor manages. The second matters because a directory Cursor placed there is a legitimate isolated session even before it becomes a distinct git worktree, where linkage alone would report no repository. Detecting linkage required a new primitive rather than the existing context resolver. That resolver short-circuits on finding a local .planning directory before it ever compares the git directory to the common one -- and an isolation worktree normally has its own checked-out .planning. Reusing it would have read a correctly isolated session as unisolated and denied it, which is the failure direction that gets a guard switched off. The comparison is now its own shortcut-free function that the resolver delegates to after its own shortcut, so existing behavior is unchanged, and the case that would have broken is pinned. The subagent type is checked before any configuration is read, so an unreadable config cannot deny a dispatch this guard would never have enforced against. The input-schema comment on the Cursor hook documented only the fields common to every event and omitted the ones specific to this one. That omission cost a halt during this work; it now documents both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): enforce the resolved dispatch decision, not the host capability The guard keyed on the registry's dispatch.isolation, which says only that a runtime is CAPABLE of harness worktrees. The decision that actually governs a dispatch is the one the workflow resolves after gating, and that legitimately comes out as sequential in three documented cases: a project setting use_worktrees false, a per-plan submodule intersection, and the base-check auto-degrade. The workflow tells the model to omit the flag in exactly those cases, and the guard was denying every one of them. The third case matters most. The preceding fix made the base-check degrade on git timeouts and a missing git binary, where it had previously answered "safe". That correction is right, and it means a transient hang now degrades to sequential far more often than before -- so the two changes composed into a trap where the workflow behaved exactly as designed and the guard blocked it. The workflow already resolves isolation in shell, deterministically, which is what makes it a trustworthy source in a way the model-authored call is not. It now records that resolved value through a dedicated verb, and both guards read it first. A fresh record is authoritative, so sequential dispatches pass untouched. Absent or stale, the guards fall back to the capability check combined with the project's use_worktrees setting, which still covers the case that never reaches the workflow. Also widened the matcher to accept Task alongside Agent, since a host that names the tool Task would otherwise leave the guard silently inert while implying coverage; stopped assuming Claude when no runtime is declared, which is the shipped default and would have demanded a Claude-only argument elsewhere; and made a non-git project inert rather than denied, since advising a worktree session is not actionable without a repository. The original diagnosis never modeled sequential mode as legitimate. That omission is what let this through, and it is now recorded there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): record at resolution and bind the record to its dispatch Two independent reviews converged on the same failure: the guard was fail-open in a default install, so it did not catch the defect it exists to catch. A shipped project carries no runtime key, which made "runtime not confidently known" the common case rather than a corner one. A record asserting that isolation was required but carrying no flag then fell through to a capability lookup that answered "none", and the dispatch was allowed. The flag itself only arrived from a second shell block -- the same block a model dropping the argument would also skip. A test had pinned that behavior as intended. The record is now written by the resolver, as an unavoidable consequence of asking for the value, rather than by a step the model is told in prose to go and run. A guard against a prose-carried value cannot itself depend on prose. Mode, flag and identifiers are written together and atomically, so the flagless window is gone, and a record asserting isolation with no resolvable flag now denies instead of degrading. Runtime is also resolved from the installer's own recorded default, which makes confident resolution the normal case. The per-plan submodule gate degrades after the phase-level decision and never re-recorded, so a plan that legitimately ran sequentially was denied against a still-fresh phase record. It now records its own, scoped to the plan. A record also authorized any dispatch for four hours. One phase degrading to sequential could silently license an unisolated dispatch in the next. Records now carry phase and plan, the guards require them to match, and the window is minutes rather than hours -- the resolver rewrites it before every dispatch, so a long window bought nothing and only widened the hole. The flag validator rejected any value beginning with two dashes, which is exactly the form Cursor and Windsurf declare, so their real value could never have been stored. Writer and reader also derived the record path differently and diverged inside a linked worktree without local planning state. The predictable path remains a way to silence the control without leaving a trace in the diff. It grants no access an agent with shell does not already have, so it is documented as accepted rather than redesigned around. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3045): correct the staleness boundary and unmask a vacuous parity test The remote runner returned twenty failures. One was a real production defect the boundary case existed to catch: a record whose age exactly equalled the staleness window was treated as fresh, so it stayed authoritative for one tick past its own expiry. Freshness is now strictly inside the window. The parity test meant to stop the two guards' executor lists from drifting could never have failed. Its project fixture was a bare directory rather than a repository, so the non-git inert branch answered before the executor list was ever consulted. It asserted agreement it never actually measured. The fixture is now a real repository, like every sibling in the file. A test also asserted that Windsurf declares the worktree flag. It does not -- Windsurf resolves to no isolation by design, having no named concurrent dispatch to isolate. The test claimed a registry fact that was never true, and a comment in the resolver repeated it. Both corrected, and the test now proves what it should have all along: that the parser accepts any bare flag value, rather than one runtime's supposed value. The new guard was missing from the bundled-hook whitelist, which is the surface that decides what actually ships, and the per-plan gate had gained calls to the launcher without the preamble those calls require. The changeset carried parenthetical product descriptions the purity rule forbids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3045): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3045): make the guard tests hold on Windows Two tests redirect HOME to control where the installer-persisted runtime default is read from. Node resolves the home directory from USERPROFILE on Windows and never consults HOME, so both silently read the real runner profile, found no recorded runtime, and asserted against a project the hook had not recognised. The production code was already correct in asking the platform rather than the variable; only the tests were wrong to assume one variable answers everywhere. The helpers now mirror the override onto both. The symlink spoofing test also created a directory symlink unconditionally, which needs elevated privileges on Windows. It survived on this runner, but it would fail on any host without them, so the creation is now attempted and the test skips explicitly when it cannot be done -- a bare return would have counted as a pass and hidden the gap. Skipping alone would have left the platform uncovered, so the behaviour it proves is now also driven in-process through an injected realpath, following the seam already used for the clock. That case no longer depends on privileges at all, and the end-to-end test keeps its original assertions wherever symlinks work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ed360cd99f |
chore(#2995): extend fragment emission to agents/ and reclaim size-cap headroom (#3058)
* feat(#2995): extend fragment emission to agents/ across every read point Epic #1671 Phase 6.4. `composeWorkflow` stripped `<!-- gsd:section -->` markers only for `gsd-core/workflows/`, so a marked agent shipped its markers verbatim into every runtime — and agent text is loaded into a subagent's context on every dispatch. The issue proposed widening the `copyWithPathReplacement` guard. That is a no-op for agents: agents never traverse that function. Agent content is read for emission at five independent points, and the obvious chokepoint `stageAgentsForProfile` short-circuits on the DEFAULT `full` profile (`skills === '*'` returns the real unstaged directory), so a hook placed there is dead code on most installs. Composition now happens at two call sites instead of five parallel surfaces: `stageAgentsForRuntimeWithConverter` (with `agentsKind` and `kimiAgentsKind` routed through it via an identity converter) and the inline agent loop in bin/install.js. Both compose BEFORE any path rewrite, so a `.claude/` -> `.windsurf/` regex can never reach inside a marker attribute — the ordering #2930 established for workflows. `installCodexConfig` was the fifth read point: Codex embeds each agent's prompt into a per-agent `.toml` via its own readFileSync. Call-graph analysis missed it; the exhaustive per-runtime emission sweep found it. That is why the new guard is behavioral rather than structural — a sixth read point fails the sweep without anyone remembering to extend a list. tests/agent-fragments-emission.install.test.cjs spawns a real installer for every runtime at every agent-bearing scope, derived from RUNTIME_META and the capability registry at run time so a new runtime cannot be silently under-covered. It asserts markers are absent AND the `when="always"` body is retained, so marker-absence cannot be satisfied by dropping content. An identity-composer negative control proves the assertion can fail. Verified: 0 install failures, 0 marker leaks, body retained on 27 runtime/scope paths; red before the wiring on claude(global+local), zcode(global+local), kimi, codex and opencode. Refs #2995 * chore(#2995): give the tightest agents headroom and correct the design lock Epic #1671 Phase 6.4, second half. `agents/gsd-verifier.md` had 12 bytes of headroom under its 49,152-byte LARGE cap and `agents/gsd-debugger.md` had 147 under its 57,344-byte XL cap. Both now extract reference material to `gsd-core/references/` behind an @-reference — the documented DEFECT.AGENT-FILE-SIZE-CAP-BREACH remedy: gsd-verifier 49,140 -> 46,371 B headroom 12 -> 2,781 gsd-debugger 57,197 -> 48,851 B headroom 147 -> 8,493 Byte accounting proves no content was lost: the combined agent+reference delta is exactly the new files' headers plus the agents' slim replacement blocks. Each agent keeps its routing table and a one-line summary per entry, so it degrades gracefully on a runtime that does not inline @-references. `agents/gsd-planner.md` is untouched and still passes both char guards (49,130 < 49,152); it needed no change, so it took none. The other nine LARGE/XL agents carry NO gsd:section markers, and that is deliberate, not deferred. `when=` selection is read from gsd-core/workflows/section-manifest.json, which gen-section-manifest.cjs derives from gsd-core/workflows/*.md only — shape `{workflows: ...}`, no per-agent key, no per-agent init entry point. An agent atom therefore fails admission gate (2) ("a fact the init seam demonstrably computes at a real entry point") and would evaluate false forever while looking like working gating. Marking agents would manufacture exactly the silent-inertness rot the frozen vocabulary exists to prevent. ADR-1671 gains three amendments, two of which close gaps /adr-phase-coverage found against what actually merged: - The 19 -> 29 vocabulary widening shipped in #2994 with no coordinated ADR amendment, which that bullet's own rule forbids. Recorded now. - `flag:--verify-only` was one of six atoms #2992 withheld and deferred to "the LARGE/XL rollout phase". Five shipped; this one is permanently rejected, and that disposition lived only in a merged PR body. - Phase 6.4's own finding: emission extends to agents/, gating does not. CONTEXT.md's glossary was stale on both seams — Workflow Fragments Module still listed the original 4-atom vocabulary and described when= as "not yet acted on", and Section Manifest Module still described InvocationFacts as {waveFlag, phaseNumber, hasPriorPhases}. Both now match the shipped contract. Inventory manifest regenerated AFTER build:lib per the documented ordering landmine; 19 install-tree fixtures pick up the two new references. Refs #2995 * chore(#2995): correct the compose-site count and mark the raw stager Self-review found two comment defects in the prior commit. The agentsKind comment claimed composition lands at TWO call sites; it is three, since installCodexConfig's per-agent .toml writer was added after that comment was written. And stageAgentsForProfile is now production-dead — both callers route through the composing stager — while staying exported and unit-tested, which makes it a trap: it does a raw copyFileSync and short-circuits to the unstaged source directory under the default profile, so a future caller would silently reintroduce the marker-shipping path. Its JSDoc now says so. * test(#2995): guard the marker-documenting-doc class for agents Widening the composer's scope to agents/ makes reachable the exact class #2930 narrowed scope to avoid: a file that DOCUMENTS the marker syntax with an unfenced example is indistinguishable from a real marker, so the composer drops that line from the emitted artifact. Three rows. A fenced example must compose byte-identically. No shipped agent may carry a marker outside a fence — asserted by parsing every real agent and requiring zero explicit sections, which is what makes the fence protection load-bearing rather than decorative. And a non-vacuity row asserts an UNFENCED marker IS parsed as a real marker, so if that ever stops being true the second row is guarding nothing. Also applies two review findings: stageAgentsForProfile's new JSDoc claimed it had no production caller, which is false — bin/install.js's _stageAgents still calls it, and its consumers compose before writing. Corrected to state the invariant instead. And a let/const nit in the emission sweep. * fix(#2995): keep verifier status vocabulary in the agent, fix a wrong fixture The first remote run came back red with three failures. Both root causes were mine. 1. tests/agent-frontmatter.test.cjs requires agents/gsd-verifier.md to literally contain HOLLOW and DISCONNECTED. The Step 4b extraction moved that status vocabulary into gsd-core/references/verifier-wiring-patterns.md, so the agent no longer had it. Byte accounting said no content was lost, and byte-wise that was true — but a contract required those tokens to live IN THE AGENT. That is ADR-1671:66's flexReserve floor stated concretely: a load-bearing fragment must not be trimmed out of its host, and "the bytes still exist somewhere" is not the test. The two status tables are restored to the agent and deliberately mirrored in the reference with a note saying so, so the procedure there still reads standalone. gsd-verifier lands at 47,069 B — headroom 12 -> 2,083, rather than the 2,781 the first attempt claimed. 2. Row 12b of the new marker-documentation guard asserted that an unfenced marker example parses as a real marker, and threw instead: "unmatched /gsd:section close marker". The grammar is WHOLE-LINE only. The fixture had put the OPEN marker inline mid-sentence, so it was correctly not recognised as an open while the close, on its own line, was. That is a real refinement of the hazard this guard exists for: only a marker on its OWN line is mis-parsed — which is exactly how a documentation example is normally written. Row 12b now uses a whole-line marker, and a new row 12c pins the inline case as explicitly NOT a marker. No test was weakened to accommodate the change; the change was corrected to satisfy the tests. Refs #2995 * chore(#2995): backfill changeset pr number to 3058 --------- Co-authored-by: sim <sim@local> |
||
|
|
ff4a57b78c |
chore(#1671): migrate the remaining 13 LARGE/XL workflows to the fragment model — Phase 6.3 (#3030)
* chore(#2994): fragmentize progress.md forensic audit onto the fragment model Extract the --forensic-gated forensic_audit step to workflows/progress/steps/forensic-audit.md behind a section marker, and repair progress.md's init line to forward --forensic so the atom is actually true in production rather than only under direct CLI tests. progress.md shrinks 32630 -> 27207 bytes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize the four manifest-wired workflows new-project, quick, new-milestone and progress each already had a dedicated cmdInit* entry point but zero marked sections. Extract nine gated bodies to workflows/<wf>/steps/ behind section markers and repair each init line to forward its flags. Fold --full into the discuss/research/validate facts inside cmdInitQuick so the when= grammar never sees an OR, per the chunked-mode precedent. Fixes found while working, per the no-defer rule: - cmdInitProgress passed no phase info to buildSectionManifestField, so state:phase-mvp-mode was permanently false — an atom in the vocabulary whose fact could never be computed. - the quick init router folded flag tokens into the free-text description, which the new forwarding would have corrupted. - a #2508 dispatch note was nested inside quick.md's Agent(prompt=) fence, leaking orchestrator guidance into the subagent prompt. - progress.md had a 3-vs-4 backtick outer-fence imbalance. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize verify-work.md and admit state:ui-phase-active Wire cmdInitVerifyWork to buildSectionManifestField — it was a dedicated entry point that never emitted a manifest — and mark two sections. state:ui-phase-active folds (plan:pre hooks include an active ui step) OR (the phase dir holds a *-UI-SPEC.md) into one boolean in init.cts, so the grammar still sees a single operator-free atom. The inner Playwright-MCP check stays as prose inside the fragment: it is live session state and no init seam can precompute it. The MVP false-branch note is a real fallback, not redundant prose, so it sits outside the marker — gating it away would delete the text needed precisely when MVP mode is off. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): follow moved workflow content in drift guards Retarget every guard that asserted on content this branch moved into workflows/<wf>/steps/, mirroring 815b3d897. Each retargeted assertion was verified to still fail when its step file is blanked, so none was weakened into vacuity. Three assertions in verify-mvp-uat were genuinely red. Three more were worse than red — passing for the wrong reason: - quick-commit-boundary and worktree-cleanup anchored on indexOf('Step 5.6'), which matched a later cross-reference and sliced 16069 chars that coincidentally held the asserted substrings. Replaced with an expandWorkflowSections helper that splices step content back in place. - phase6-review-capabilities lost its end boundary and widened to EOF. - playwright-ui-verify matched 'UI' in an unrelated bullet and 'fall back' in a subagent-dispatch line after the real content moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize code-review and complete-milestone, admit three atoms Add dedicated cmdInitCodeReview and cmdInitCompleteMilestone entry points alongside the shared generic ones rather than modifying them — init.phase-op and init.manager carry a CRITICAL blast radius (179 dependents, 24 processes) and stay byte-identical for their other callers. Admit flag:--fix, state:fallow-enabled and state:git-create-tag, each with a consuming section and a fact its own entry point computes. Both sections had the resolver-in-body hazard: the fallow config-gate and the git.create_tag check each sat inside the very block being gated, so gating would have disabled the resolver that decides the gate. Both are hoisted into init and the bodies now consume the resolved fact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): retarget code-review and milestone drift guards, fix two red tests Retarget guards that asserted on content moved into steps/, proving non-vacuity by blanking each step file and confirming failure. Also fixes two genuinely red tests found while working, per the no-defer rule: - workflow-fragments' frozen-vocabulary lock was missing state:ui-phase-active, so commit 7ef7f8336 shipped red. Lint and build both passed over it, which is why neither is sufficient verification. - code-review's quick.md capability-hook assertion carried a stale delimiter after the 18ff35d20 extraction. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize autonomous.md and admit state:plan-strategy-converge Five sections share one atom, the pattern plan-phase already uses for flag:--research-phase. The atom folds --converge OR --cross-ai into a single boolean in cmdInitAutonomous so the grammar stays operator-free. cmdInitAutonomous is additive; init.milestone-op, init.manager and init.phase-op are untouched and still consumed. The $PLAN_STRATEGY bash resolver is deliberately retained — ungated local-planning bullets still read it, so the init-side fact supplements it rather than replacing it. converge-fail-fast required splitting one bash fence so the always-run CONVERGENCE_ARGS construction stays outside the marker. All three flag-absent fallbacks were left outside their markers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize review and discuss-phase-assumptions Admit state:reviewer-instances-configured (two peripheral notes share it; the core reviewer-lane dispatch stays unmarked — it is the workflow's primary always-evaluated logic, not an optional branch) and state:auto-advance-active, which folds --auto OR two config keys into one boolean so the grammar stays operator-free. discuss-phase-assumptions was the highest-risk edit in this PR. Its auto_advance step is a full if/elif/else; gating it whole would have deleted the flag-absent fallback needed exactly when --auto is off. Split verified exact: resolvers 636-651 and the 'End here' fallback 668-669 both stay outside the marker; only 653-667 is gated. Adds emitted-drift acks for the two files that grew — review.md (+55 B) and autonomous.md (+737 B from 80799211c, which had none and would have red-gated the push. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize docs-update, update, transition and new-milestone Part A Completes the 13-workflow rollout. Three of these had no init call at all and gained a dedicated entry point plus their first gsd_run query line. Admits state:is-monorepo and adds state:next-channel, state:workstream-active and state:flat-mode. Vocabulary 26 -> 30 atoms. Part A of new-milestone applies when NO workstream is active — the negation of state:workstream-active. Rather than teach the grammar negation, which is the Greenspun drift the frozen list exists to prevent, it gets a separate positively-phrased atom whose fact is the inverse. Part B, which always runs, stays outside the marker. flag:--verify-only is deliberately NOT admitted: docs-update has no contiguous purely-additive region for it, and an atom without a consuming section is dead vocabulary. Evidence recorded in the slice report. update.md reuses its existing resolved $GSD_TOOLS rather than prepending the canonical preamble, which would have clobbered it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): stop automated-ui-verification re-resolving its own gate, retire dead vocabulary Two defects the new tests caught. The automated-ui-verification step re-ran gsd_run loop render-hooks and recomputed UI_PHASE_ACTIVE inside a body that is only read when that fact is already true — the circular self-disabling pattern this design forbids, introduced by 3c654b168. cmdInitVerifyWork now exposes ui_phase_active and the step consumes it. Its launcher preamble goes too: no gsd_run remains. The Playwright-MCP check stays as prose — that is live session state. Dead vocabulary predating this PR: flag:--full and state:needs-codebase-map were admitted with a gate-1 claim that never materialized. flag:--full is removed, redundant once quick folds it into discuss/research/validate. state:needs-codebase-map gets the real consumer it always lacked, gating new-project's codebase-map offer. Vocabulary 30 -> 29, and no atom is now without a consuming section. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): add the atom-admission, inversion and resolver-hoist gates The two existing parity guards prove vocabulary/predicate symmetry but never that a fact is computed — an atom no cmdInit* assembles evaluates false forever. These close that hole: - per-atom satisfiability for all 29 atoms, plus an anti-vacuity assertion so the loop cannot silently cover zero atoms - dead-vocabulary check against the shipped manifest - inversion guard: the flag-absent fallbacks in discuss-phase-assumptions and verify-work must stay outside their markers - data-driven resolver-hoist guard over the shipped manifest, so a future extraction cannot reintroduce the circular class - compound-fold coverage (--full, --cross-ai, --rc, config-only --auto) - null-vs-[] degraded/computed distinction, and flag value shapes Also repairs the frozen-vocabulary lock, which was stale and red for the seven atoms earlier commits on this branch shipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2994): add changeset for the fragment-model rollout Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): cite the issue on the two new allow-test-rule exemptions ADR-456 requires an issue ref on the same line as the annotation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2994): correct the atom-count claims after retiring flag:--full The vocabulary doc comments still said 30 entries; it is 29 since flag:--full was removed as dead vocabulary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): dedupe the phase-fallback block and harden --ws parsing Review findings. MAJOR: the three new init entry points each pasted a verbatim copy of the guardedFindPhase/guardedGetRoadmapPhase fallback, taking the repo from four copies to seven — DEFECT.GENERATIVE-FIX. Extracted applyRoadmapFallback and folded six of the seven; each call site keeps its own field-set via a closure. Duplication removed rather than papered over with a parity test. cmdInitPhaseOp stays out: its fallback omits has_reviews, so it is not a byte-identical copy, and it is CRITICAL-radius. LOW, pre-existing: GSD_WS captured [^[:space:]]+ and expands unquoted, so a workstream name holding glob metacharacters would expand against the filesystem. Narrowed to [A-Za-z0-9._-]+. The unquoted expansion is kept — it must word-split into two args and vanish when empty. Also restores the vocabulary ordering convention, and fixes a masked test bug the mandated run surfaced: the flag-forwarding guard checked only the first init line per workflow, but new-milestone has two, so a real failure was reporting exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): drop the stale new-milestone emitted-drift ack new-milestone.md was acked for a +406 B growth measured against an intermediate commit. Net against origin/next it SHRANK by 8 bytes, so nothing needed the ack and it explained nothing — which the differential attribution check reports as a stale acknowledgment, not a pass. update.md's entry stays: it genuinely grew +703 B. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): resolve the 15 failures from the full matrix run All 15 were real and identical on both lanes. REAL REGRESSION: autonomous.md hit 41479 chars against the #2196 guard's 40960 cap — a CHARS cap distinct from the LARGE tier byte cap, which the five section stubs pushed it over. Extracted the 3a.5 UI Design Contract body to references/; now 39968 chars, and the file nets -795 B vs base, so its growth ack is deleted rather than left stale. REAL DEFECT: docs referenced /gsd-transition, which is not a live registered command. Reworded. STALE FIXTURE: the emission byte-identity test hardcoded two marked workflows; this branch legitimately marks fifteen. Fixture corrected — the source was right. The rest were drift guards over the eight workflows the earlier sweep did not cover, retargeted at where the content now lives with non-vacuity proven by blanking each step file and confirming failure. The GSD_WS forwarding guard was checked as a possible real break and is not one: the charclass narrowing is intact and forwarding works end to end. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): drop the ack for a newly-added reference file A new file's emitted ripple is attributable to the diff that adds it, so the acknowledgment explained nothing and the differential check reports it as stale. Removing the last entry removes the fragment — an empty one signals nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): retarget the UI-contract guards and clear two transitive advisories The §3a.5 extraction that brought autonomous.md under the #2196 char cap moved its body to references/autonomous-ui-design-contract.md, so ten guards in autonomous-ui-steps and check-ui-safety-gate were asserting it against the host. Retargeted via a combined read, each proven non-vacuous by blanking the reference file and confirming failure. This class had already bitten twice on this branch because each sweep was scoped to the workflows touched at that moment, so this one was exhaustive: ~70 test files across all 13 workflows, zero further broken or vacuous assertions found. Also clears two high transitive advisories the matrix flagged on one lane — fast-uri GHSA-7p8r-x3mc-p8w7 and three ip-address SSRF/trust-boundary issues. Both pre-date this branch: package-lock.json was untouched until now, so the production tree was byte-identical to the base. Lockfile-only, package.json unchanged, verified against a real npm ci install. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): backfill changeset pr number to 3030 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
88f6d9bd1b |
fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu * fix: preserve installer executable mode * chore: add changeset for PR #2812 * test(#2644): acknowledge Cursor emission changes * test(#2644): drop spent emitted drift acknowledgments * fix(#2644): remove retired Cursor command converter --------- Co-authored-by: clezcoding <clezcoding@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
067a4d1c6c |
fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls (#3015)
* test(#2650): add failing-first regression for plan-phase stall detection Regression test for gsd_stall_should_recover / gsd_stall_watch and the planner.stall_* config keys, none of which exist yet — proves RED before the fix lands in the next commit. * fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls Mirrors the already-shipped executor.stall_* pattern (execute-phase.md, bug #3212) but with a dispatch change the executor's prose-only surveillance lacks: the standard planner spawn, chunked-outline planner spawn, chunked-per-plan planner spawn, plan-checker spawn, and revision-loop planner respawn now dispatch with run_in_background=true and are followed by a real, bounded bash poll (gsd_stall_watch) that returns control to the orchestrator on its own schedule instead of waiting indefinitely on a subagent that may never return. On stall, the existing accept-plans/retry/ stop recovery menu (9a/11a) is auto-surfaced instead of requiring a manual interrupt. New config keys planner.stall_detect_interval_minutes (default 5) / planner.stall_threshold_minutes (default 10) mirror executor.stall_*. The helper functions (gsd_stall_should_recover, gsd_stall_watch) live in a new lazily-loaded gsd-core/workflows/plan-phase/steps/stall-detection- helpers.md rather than inline, and per-site prose is kept minimal, because plan-phase.md is frozen under the ADR-857 Phase 6 PRE_PHASE6 gate (tests/phase6-capstone-conformance.test.cjs) with ~36 bytes of headroom at baseline; the net effect is plan-phase.md.md ships slightly SMALLER than before (the old unconditional-wait ORCHESTRATOR RULE sentences are gone at the five touched sites, superseded by the bounded watcher). Also fixes a stale doc comment in tests/workflow-size-budget.test.cjs that still described the per-file workflow-size-baseline.json guard removed by #2724 (ADR-2719 Phase 4) as if it were still the enforcement mechanism — discovered while verifying this fix's own byte budget. Researcher and pattern-mapper spawns are untouched (out of scope per the issue's Agent Brief). * fix(#2650): make gsd_stall_watch single-cycle; harden numeric config inputs Two review findings addressed on top of the prior commit: 1. gsd_stall_watch previously looped internally for the full threshold+interval duration inside ONE Bash tool call (up to 15 min at defaults) — a single call blocking that long risks the host tool's own timeout killing it before it ever prints a result, silently defeating the fix. Redesigned to a single sleep-and-check cycle per call, taking an explicit dispatch_ts so the orchestrator prose can repeat the (short, default 5 min) call until it resolves; the outer threshold is now enforced by dispatch_ts accumulating across calls, not by one call's duration. Documented the resulting trade-off (up to one interval of added latency on the success path) in the changeset and reference doc. 2. PLANNER_STALL_INTERVAL_MINUTES/THRESHOLD_MINUTES are config-controlled values that flow into bash arithmetic ($(( ))). A review flagged this as command injection; empirically verified against both macOS bash 3.2.57 and Docker bash:5 that this is NOT actually exploitable (bash hard-errors on a `$(cmd)`-shaped arithmetic operand rather than invoking it) — but an unvalidated malformed value WOULD abort the stall-watcher itself with that bash error, silently defeating the exact hang-recovery this issue ships. Added integer validation with safe-default fallback, both at the config-resolution point and defensively inside gsd_stall_should_recover. Also adds the previously-missing integration coverage for gsd_stall_watch's real execution (grep/find/date plumbing), not just the pure classifier. * fix(#2650): correct AC2 self-test — helpers doc may name teams-status in prose The AC2 regression test asserted the stall-detection-helpers.md step file never contains the substring "teams-status" at all, but the file's own prose explicitly documents its independence from that guard (containing the word by design). Narrowed the assertion to what actually matters: no second `query teams-status` call site and no gating on it, not a blanket absence of the word. * test(#2650): regenerate golden install-tree fixtures for the new step file gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md is an emitted file (installed for every runtime), so adding it changes the install tree even though it is invisible to docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (both explicitly scope to non-recursive gsd-core/workflows/*.md — verified against the execute-phase #2930 and pre-existing plan-phase step-file precedent, which are equally absent from both inventory artifacts). The golden install tree snapshots the sorted list of emitted relative paths per runtime, so a file invisible to the inventory is still visible here. Regenerated via `npm run gen:install-tree` — one line added per runtime fixture (19 files), no other drift. * fix(#2650): restore 7 ORCHESTRATOR RULE labels; sync runtime-launcher preamble Two more consequences of extracting helper bodies out of plan-phase.md, both caught by verification (0017e1a78, 9 unique failures): 1. tests/plan-phase-drift-guard.test.cjs (#913) requires at least 7 "ORCHESTRATOR RULE — ALL RUNTIMES" labels in plan-phase.md itself, one per agent spawn site. Moving the full explanatory blocks to plan-phase/steps/stall-detection-helpers.md carried 5 of the 7 labels out with them (only the untouched researcher/pattern-mapper sites kept theirs). Restored a short label at each of the 5 stall-watch sites, trimmed a few more redundant words ("Per 7.99, " — already established by the adjacent step-7.99 pointer) to stay under the frozen PRE_PHASE6 cap (94497 bytes, 21 bytes headroom). 2. tests/runtime-launcher-parity.test.cjs (#373) requires exactly one canonical gsd_run preamble, byte-equal to gsd-core/workflows/_runtime-launcher.snippet.sh, before the first gsd_run call in any workflow .md that calls it (recursive scan under gsd-core/workflows/, unlike the non-recursive inventory/step-tag-balance checks). The new step file's config-get calls use gsd_run without one. Fixed via `node scripts/sync-runtime-launcher.cjs`, verified: exactly 1 preamble occurrence, before the first call, including the .claude/ and .codex/ home fallback arms. Also verified (no fix needed, evidence recorded): the generic `gsd-core-verbatim` identity rule in tests/helpers/emitted-provenance.cjs (roots: ['gsd-core'], pattern matching workflows/.+) self-attributes any new gsd-core/workflows/** path to itself, so the new step file needs no drift-ack entry — consistent with plan-phase.md's own net shrinkage requiring none either. * test(#2650): acknowledge plan-phase.md's +14 byte drift Restoring the 5 ORCHESTRATOR RULE — ALL RUNTIMES labels (#913) flipped plan-phase.md from -142 bytes (post-extraction) to +14 bytes net growth against baseline (94483 -> 94497), which the differential attribution size ratchet (tests/emitted-attribution.test.cjs) correctly flags as unacknowledged growth. Added tests/emitted-drift-acks/2650-plan-phase- stall-detection.json, keyed on the bare filename plan-phase.md per the existing fragment schema (see tests/emitted-drift-acks/2649-diagnose- execute-plan-base-check.json), explaining the growth as exactly the 5 restored labels — still verified under the PRE_PHASE6 cap (94497 < 94519) and satisfying #913's 7-label requirement. * fix(#2650): bind {outputFile} from the real Agent() return — was dead code Independent review blocker: PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE were read by every gsd_stall_watch call but never assigned anywhere in the diff. With the variable permanently empty, `[ -f "$output_file" ]` was always false, marker_found could never become true, and marker_received was unreachable — the marker-based detection path was permanently dead. Worse for the plan-checker spawn specifically: a checker that PASSES touches no *-PLAN.md files, so it had no working completion signal at all without the marker path. A healthy plan-checker finishing cleanly in two minutes would be declared stalled once planner.stall_threshold_minutes elapsed and the recovery menu would fire on an already-succeeded agent — worse than the original unbounded hang. Fixed by replacing the dead bash variable with the `{outputFile}` orchestrator-substitution token, the same convention docs-update.md:471 already uses for a real run_in_background=true Agent() return ("Read tool: file_path: `{outputFile from README agent result}`"). This is a net BYTE SAVING at each site (`"{outputFile}"` is shorter than `"$PLANNER_OUTPUT_FILE"`), which funded moving the full binding explanation — including why plan-checker's *-PLAN.md glob alone is not a working completion signal — into the lazily-loaded reference file to stay under the frozen PRE_PHASE6 cap (94496 bytes, 22 headroom; net +13 over baseline, acknowledged in tests/emitted-drift-acks/2650-plan-phase- stall-detection.json). Added a regression test asserting plan-phase.md itself binds {outputFile} at all 5 spawn sites and contains no dangling $PLANNER_OUTPUT_FILE / $CHECKER_OUTPUT_FILE reference — the previous test suite only exercised gsd_stall_watch's behavior when handed a valid argument, which is why the dead production wiring survived two rounds of review. Also fixed tests/fix-2650-plan-phase-stall-detection.test.cjs:170-195's raw try/finally to use t.after(), per CONTRIBUTING's test-cleanup convention. * chore(#2650): backfill changeset PR number to 3015 * fix: normalize CRLF at the read boundary in all .md-bash-extraction tests Maintainer-authorized scope expansion, folded into this PR rather than deferred: the Windows CI lane on this PR's own tests/fix-2650-plan-phase- stall-detection.test.cjs exposed DEFECT.TEST-SHELL-PIPELINE-NONPORTABLE (CONTEXT.md; recurring since #1700) as a repo-wide latent class, not a one-off. Ten test files parse a fenced ```bash block out of a workflow .md file and execute it via spawnSync/execFileSync; a Windows checkout can yield CRLF line endings despite .gitattributes eol=lf, and bash then treats the trailing \r on every extracted line as part of the token — "unexpected EOF while looking for matching `"'" or a bare syntax error, partway through the script. Added tests/helpers.cjs:readFileNormalized() — strips \r\n -> \n at the read boundary, before any fence-slicing or regex runs, so every downstream operation is correct by construction. Migrated all ten call sites to it: Previously broken (fs.readFileSync with no normalization anywhere between read and spawn): - tests/worktree-cleanup.test.cjs (extractCwdGuardBash) — also fixes a misleading comment claiming the fence regex alone was "CRLF-safe"; it protected only the fence delimiters, never the captured body. - tests/new-milestone-clear-phases.test.cjs (extractFenceBetween, extractFenceContaining) - tests/code-review-pipeline-regression.test.cjs (extractPostProcessingScript) - tests/drift-detection.test.cjs (readGate/bashBlock, plus the snippet-file comparison read in the same test) - tests/graphify-visualization.test.cjs (extractStep3Block) - tests/pause-work-improvements.test.cjs (extractCheckBlock) - tests/plan-review-convergence.test.cjs (extractReviewerFlagsParseBlock and the inline post-config-gate resolution-block slices) Already correct (split(/\r?\n/) then join('\n')), migrated to the shared helper for consistency rather than a fourth/fifth/sixth copy of the same fix: - tests/git-base-branch.test.cjs (extractHandleBranchingBash) - tests/quick-branching.test.cjs (extractStep25Bash) - tests/runtime-launcher-parity.test.cjs (extractResolverSnippet) Verified against a simulated Windows CRLF checkout (not assumed): for both the worktree-cleanup.test.cjs and new-milestone-clear-phases.test.cjs extraction shapes, confirmed the pre-fix code produces a real bash syntax error on CRLF input and the post-fix code does not. One eslint follow-up: local/no-crlf-fragile-split statically flags any bare `\n` inside a markdown-fence-shaped regex, regardless of whether the receiver was already normalized — it cannot see the readFileNormalized() data-flow. Kept `\r?\n` in extractCwdGuardBash's fence regex (redundant but harmless on pre-normalized input) rather than fight the rule. Scope note: this diff is broader than issue #2650's own change (plan- phase.md stall detection) because the Windows lane surfaced a genuine repo-wide defect class while verifying that fix, and the maintainer authorized fixing it here rather than filing it separately and shipping a known-broken pattern. Runtime impact: none — this is a test-harness-only defect. The live orchestrator (Claude Code or another runtime) does not do a byte-exact extract-and-pipe of .md content into a shell the way these tests do; it reads the instructions and generates its own bash invocation text, which does not reproduce a raw CRLF pass-through the same way. Not touched: tests/plan-review-convergence.test.cjs's separate, tracked spawnSync ETIMEDOUT flake under bench load (#3005, reproduced on unmodified next) — unrelated load-sensitivity, not a CRLF symptom. * fix(#2650): remove stale drift-ack fragment — plan-phase.md is self-explaining tests/emitted-drift-acks/2650-plan-phase-stall-detection.json acknowledged plan-phase.md's own emitted-path hash move, but plan-phase.md is directly edited in this diff. Per the emitted-attribution law (ADR-2719, tests/emitted-attribution.test.cjs), a workflow's emitted key equals its own source path (gsd-core-verbatim identity rule), so a direct edit to the source is self-explaining and auto-attributed — no ack was ever needed. Verified via the pre-merge lint (scripts/lint-emitted-drift-ack.cjs, run through npm run lint:ci with a fully cleared eslint cache): it passes clean with the fragment removed, confirming no contradiction between the lint and the runtime attribution gate — this was simply an unnecessary fragment. * fix(#2650): restore plan-phase.md drift-ack — size ratchet demands it against next tests/emitted-drift-acks/2650-plan-phase-stall-detection.json was deleted in the previous commit because, against an earlier verification base, it was inert: it explained a moved emitted hash that a direct edit to plan-phase.md already self-attributes. Against origin/next@f1af47766a the demand is different: plan-phase.md is 13 bytes larger than the base copy, which trips the emitted-attribution size ratchet — a job this same ack also performs. Recreated in the documented shape, keyed on the bare filename plan-phase.md (not the full path, and not restating the byte delta per review guidance), describing the actual change: the {outputFile} binding fix for the dead PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE variables and the 5 restored ORCHESTRATOR RULE labels required by #913, both at the stall-watch spawn sites, with explanatory bodies living in the lazily-loaded gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md reference. Confirmed no other fragment (on this branch or on next) claims the bare key "plan-phase.md" before recreating — scripts/lint-emitted-drift-ack.cjs's duplicate check is an exact string match, and the only other mention of plan-phase.md in tests/emitted-drift-acks/ (2658-trae-instruction-file-path.json) uses the full path as its key, so there is no collision. * fix(#2650): real cause of Windows CI failure — bash -c argv-transport, not CRLF The CRLF diagnosis for PR #3015's Windows failure was wrong. Proven wrong, not assumed: .gitattributes' blanket `* text=auto eol=lf` means a Windows checkout never receives CRLF for stall-detection-helpers.md, and the extracted fence's line 64 is byte-identical and correctly balanced on every platform. The real cause: runShouldRecover() passed a 70+ line, quote-dense script as ONE argv element to `spawnSync('bash', ['-c', script, arg0, ...])` PLUS four more positional args. Windows has no execve — Node serializes that whole argv into a single CreateProcess command-line string, and Git Bash's MSYS layer re-splits and unescapes it with its own rules. The boundary between the script and the trailing args was not stable across that round trip (live evidence: one failure's stderr was prefixed `gsd_stall_should_recover_test:` — arg0 arrived — another `/usr/bin/bash:` — arg0 did not). Fixed by writing the script to a temp file and running `bash <file> <args>` instead — the four values are now normal, quote-free positional args, and the script itself never enters argv transport at all. Mirrors tests/quick-branching.test.cjs's extractStep25Bash/runStep, which already uses this exact shape and is green on Windows on `next`. tests/worktree-cleanup.test.cjs's extractCwdGuardBash/runGuard stays on `bash -c` but never appends extra positional args beyond the script itself, so it never hits the same boundary — checked both siblings per review, not assumed. Corrected the now-actively-misleading CRLF comment in extractStallHelpersBash(), and corrected the changeset's claim that the repo-wide CRLF-normalization fix (folded into this branch, maintainer- authorized) explains this PR's own Windows failure — it doesn't, though it remains defensible on its own merits as general test-portability hardening. Separately, while auditing the shipped (non-test) gsd_stall_watch for Windows portability per review request, found and fixed a second, real user-facing defect: the artifact-freshness check used GNU find's `-newermt "@<epoch>"` shorthand, which the BSD find(1) actually shipped on macOS does NOT understand ("Can't parse date/time: @<epoch>", verified live against /usr/bin/find on both a stale and a genuinely fresh file). With the adjacent `2>/dev/null`, that failed silently and permanently degraded artifact_fresh to false on every macOS run — a plan-checker or planner actively writing plan files could still be reported "stalled." Replaced with `find $glob -mmin -N` ("modified less than N minutes ago"), which needs no date-string parsing and is supported identically by GNU find and BSD find; verified live that the old shape fails and the new shape passes against the same real fresh file. Added a real-execution regression test (gsd_stall_watch with `sleep` stubbed to a no-op so the test doesn't actually wait, but the real `find ... -mmin` line still runs) proving the fix, replacing the prior "not integration-tested" note for that path. Note: the remote gsd-test runner is Linux-only, so it cannot itself confirm the Windows fix — only the actual windows-latest CI lane can. * fix(#2650): route the third bash -c call site through the same temp-file seam runWatch() and a `-mmin` regression test still passed their script via `bash -c <script>` after the previous commit only converted runShouldRecover() — live Windows CI on 4b86cc57f confirmed the mechanism: failures went 11 -> 4, and `full test (windows-latest, 22, shard 1/3)` and `shard 2/3` flipped from fail to pass, but the remaining 4 failures (all in this file, all still `bash: -c:`) were exactly the gsd_stall_watch describe block, which runWatch() serves. runWatch() passes NO extra positional args at all, so this also rules out the trailing-args theory from the prior commit: the ~73-line, quote-dense script itself is what does not survive Windows argv serialization when passed as a single `-c` element, regardless of how many (if any) further argv elements follow it. Extracted one shared runBashScript(script, args, opts) helper — write to a fs.mkdtempSync'd file, run `bash <file> [args...]`, clean up in `finally` — and routed all three bash-invoking call sites in this file through it (runShouldRecover, runWatch, and the -mmin freshness test that builds its own script inline for the `sleep` stub). One transport seam means a fourth call site in this file cannot silently reintroduce the bug in isolation, which is exactly what happened here with a second call site. Corrected extractStallHelpersBash()'s doc comment a second time to state the mechanism precisely (script content, not argv-element count) and cite the live evidence (11->4 failures, shards 1 and 2 flipping green) so the next reader does not have to rediscover it. Audited every other bash-invoking call site in files this branch touches, per review request: - tests/code-review-pipeline-regression.test.cjs (runPostProcessing), tests/graphify-visualization.test.cjs (runBlock), and tests/drift-detection.test.cjs (two execFileSync('bash', ['-c', ...]) sites, one of them carrying the same giant runtime-launcher preamble text) — all pre-existing, UNCHANGED by this branch (only touched for the readFileNormalized() CRLF swap), and already exercised on `next`'s last six Windows CI runs per the reviewer's own citation. Left as-is: no evidence of failure, and converting untested pre-existing code outside #2650's scope on an unverifiable guess would be its own risk. - tests/git-base-branch.test.cjs (runHandleBranchingStep) and tests/quick-branching.test.cjs (runStep) already use the same temp-file pattern. No action needed. - tests/runtime-launcher-parity.test.cjs (runResolver) uses `bash -c` but is explicitly `if (process.platform === 'win32') return '';` guarded off on Windows entirely, for an unrelated extension-less-PATH-stub reason — never reaches Windows argv transport at all. No action needed. - tests/worktree-cleanup.test.cjs (runGuard) confirmed by the reviewer as correct and verified; not touched, per instruction. Do not touch: the -mmin fix, the drift-ack fragment, the changeset — all three confirmed correct in prior rounds and left untouched here. Note: the remote gsd-test runner is Linux-only and cannot confirm this; only the windows-latest lanes on #3015 can. * fix(#2650): give runBashScript a default timeout runShouldRecover() was the only one of the three call sites through runBashScript() with no timeout — runWatch() and the -mmin test both pass timeout: 10000 explicitly. Not a regression (this path never had a bound before), but CONTEXT.md's unbounded-subprocess guidance applies directly, and runShouldRecover() is driven repeatedly by a fast-check property test: one pathological input that fails to terminate would hang CI indefinitely instead of failing. timeout: 10000 is now the helper's own default, with ...opts spread after it so the two existing explicit timeout: 10000 call sites are unchanged and any future caller inherits a bound automatically. * fix(#2650): build the -mmin freshness test's glob with forward slashes Windows CI on d6ddda6ea reported the last failure: the -mmin regression test expected 'active' but got 'waiting' — find matched nothing, the same silent-degradation shape as the macOS -newermt defect, but this time in the test's own fixture rather than the shipped bash. Traced what production actually passes: every gsd_stall_watch call site in plan-phase.md builds artifact_glob as `"${PHASE_DIR}"'/*-PLAN.md'` — PHASE_DIR is a POSIX-style .planning/phases/NN-slug value, and the whole thing runs under Git Bash regardless of host OS, so production's glob is always forward-slash. The test instead built it with `path.join(tmp, '*-PLAN.md')`, which on Windows yields a backslash path (C:\Users\RUNNER~1\...\*-PLAN.md). In bash pathname expansion a backslash escapes the next character, so that pattern can never match a real path — find silently returns empty under the existing 2>/dev/null, same shape as the macOS bug. Confirmed as a test artifact, not a production defect: production never constructs the glob this way, so no Windows user is affected. Fixed by forward-slashing the tmp dir before appending the glob suffix, matching production's own convention, with a comment recording why (so a future "simplify this back to path.join" edit doesn't silently reintroduce the failure). The shipped bash's unquoted $artifact_glob is untouched — quoting it would break the multi-file glob expansion it exists for. Note: the remote runner is Linux-only and already passed clean at d6ddda6ea (0/29,603, both node lanes); only the windows-latest lanes on #3015 can confirm this fix. * fix(#2650): forward-slash the three remaining runWatch globs (vacuous-pass CR) The :353 fix (833c11da9) only converted the -mmin freshness test's glob. Three sibling tests in the same describe block still built theirs with path.join(tmp, '*-PLAN.md'), which yields a backslash path on Windows. Two of those three were silently passing for the wrong reason: the '-> stalled' and '-> waiting' tests both expect the glob to match nothing, and on Windows a backslash path matches nothing regardless of whether the directory is actually empty (bash eats each backslash as an escape before the pattern is even evaluated). They would have passed identically with glob expansion completely broken, which is a vacuous pass — not exercising what they claim to. The third ('-> marker_received') is outcome-independent of the glob, so it was merely inconsistent rather than wrong. Converted all three to the same `${tmp.replace(/\\/g, '/')}/*-PLAN.md` construction already used at the -mmin test, so every glob in the file now matches production's own forward-slash `"${PHASE_DIR}"'/*-PLAN.md'` shape, and the two negative tests are meaningful on Windows instead of accidentally correct. Reworded the trailing comment on the 'stalled' test's glob line: it now describes the fixture (the tmp dir contains no *-PLAN.md files) rather than the pattern, since "matches nothing" read as a property of the glob syntax when it's a property of what's on disk. No assertion, the sleep stub, runBashScript, or the shipped bash changed. Smoke-tested all three updated tests manually before committing (not via node --test): marker_received / stalled / waiting, all correct. * fix(#2650): fix own regression tests for #2993's plan-phase.md relocation 531101843's merge with origin/next brought in #2993 (unrelated, epic #1671 Phase 6.2), which extracted plan-phase.md's whole "Chunked Planning Mode" section into gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md, leaving a <!-- gsd:section --> pointer behind. tests/plan-phase-drift-guard. test.cjs (#913) was already updated to read the combined surface (host file + every steps/*.md) so its label count didn't go blind — my own #2650 regression tests were not, and searched plan-phase.md alone for the two chunked spawn sites' headings, which no longer exist there. Two tests failed outright (indexOf returning -1); a third ("standard planner spawn") was silently weakened to an unbounded slice-to-EOF by the same relocation, since its own end-boundary heading also moved — passing by accident rather than by testing what it claimed. Promoted the drift guard's local readPlanPhaseCombined() to a shared, exported tests/helpers.cjs readWorkflowCombined(workflowPath) (host file + sorted steps/*.md, CRLF-normalized at the read boundary) so a second, divergent implementation is never written — the drift guard now delegates to it via a same-named local wrapper, unchanged at every existing call site. Fixed the three affected tests in tests/fix-2650-plan-phase-stall-detection. test.cjs: - "standard planner spawn (step 8)": end boundary changed from the now-gone "## 8.5. Chunked Planning Mode" heading to "## 9. Handle Planner Return", which still exists in plan-phase.md. - "chunked outline spawn (8.5.1)" / "chunked per-plan spawn (8.5.2)": now read gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md directly (not the generic multi-file combined blob, whose file-sort ordering would put unrelated step files between 8.5.2's slice and any downstream anchor) — the same heading-to-heading slicing as before still works because the file is small and self-contained. - Extended the "no unbound $PLANNER_OUTPUT_FILE/$CHECKER_OUTPUT_FILE" check to also scan chunked-planning-mode.md, since two of the five spawn sites now live there. - Added a new count-based test asserting exactly 5 (not "at least one") `gsd_stall_watch "$TS" "{outputFile}"` invocations across the combined surface, mirroring #913's own label-count guard, so every one of the five spawns stays provably bounded and a future relocation can't silently drop one without a test noticing. Also added a small positive test that plan-phase.md's <!-- gsd:section --> pointer to chunked-planning-mode.md exists (#2993 is unrelated to #2650 but its presence is now load-bearing for where 2 of the 5 spawn sites live). Audited every other test file in the repo for a stale reference to content #2993 relocated (searched for the moved headings/prose and for "chunked-planning-mode"/"CHUNKED_MODE" across all *.test.cjs): only this file and the drift guard needed changes. tests/issue-2762-plan-reviews-chunked.test.cjs already reads chunked-planning-mode.md directly (brought in correct by the same merge). gen-section-manifest.test.cjs, init.test.cjs, and workflow-fragments.test.cjs reference "chunked-planning-mode" only as a manifest/section-id fixture value for #2993 itself, not as a stale pointer to relocated content. Did not touch: the ported ORCHESTRATOR RULE lines, run_in_background=true, the glob constructions, runBashScript, the -mmin change, the timeout default, or the drift-ack fragment (confirmed correct against the stale local `next` ref two rounds ago and left alone). --------- Co-authored-by: sim <sim@local> |
||
|
|
ad3b9ec486 |
chore(#1671): fragmentize plan-phase.md and repair flag forwarding to the init bundle — Phase 6.2 (#3019)
* chore(#2993): fragmentize plan-phase.md onto the fragment model Epic #1671 Phase 6.2. plan-phase.md is the largest workflow in the repo and carried zero markers; it was deferred out of the Phase 3 pilot for two reasons, both now dead. The 36-byte PRE_PHASE6 headroom was never the blocker it looked like — fragmentizing is net-negative on host source, so the trim is what creates the room. The --mvp interleaving was resolved by measurement in #2992 and no sub-line mechanism is built. - widen WHEN_VOCABULARY 14 -> 19 via a second coordinated ADR-1671 amendment: flag:--ingest, flag:--prd, flag:--research-phase, flag:--reviews, state:chunked-mode - state:chunked-mode is `--chunked` OR config workflow.plan_chunked, and that disjunction is resolved in the FACT, never in the grammar, so a compound condition never becomes an operator - parse the new flags on the plan-phase route; extract six gated bodies to gsd-core/workflows/plan-phase/steps/ behind manifest-gated stubs - prd-express-path.md was already extracted but read unconditionally; its wrapper is now gated, so the existing extraction finally pays off plan-phase.md 94,483 -> 87,575 bytes (cap 94,519): headroom goes from 36 bytes to 6,944. Also closes a surfaced docs gap: five real plan-phase flags (--chunked, --skip-ui, --bounce, --skip-bounce, --granularity) were documented in neither the argument-hint nor help. Making --chunked load-bearing without fixing its siblings would leave the defect class half-open. Refs #2993 * fix(#2993): forward flags to the init bundle so section gating actually fires Blocker found by the correctness review, confirmed directly, and missed by both the isolated reviewer and every test in this branch. Neither workflow forwarded its flags to the init CLI: plan-phase.md:71 INIT=$(gsd_run query init.plan-phase "$PHASE" $GRAN_PARAM) execute-phase.md:84 INIT=$(gsd_run query init.execute-phase "${PHASE_ARG}") So every flag: atom was permanently false in production and its section permanently excluded. For plan-phase that made the PRD express path UNREACHABLE — a regression, since it was an unconditional read before. For execute-phase this is PRE-EXISTING: #2932 shipped `flag:--wave` gating that has never once been true, so `--wave` silently dropped its own wave-filtering guidance. Fixed here under the no-defer rule. Why every test missed it: they drive the init CLI directly with flags, which works. Production goes through the workflow's bash line, which did not pass them — the exact "assert against the shape production uses" trap this branch's own test matrix warns about. - parse and forward --prd/--ingest/--research-phase/--reviews/--chunked (plan-phase) and --wave (execute-phase), using the anchored regex idiom the neighbouring GRAN_PARAM line already uses - add a regression guard DERIVED FROM THE MANIFEST: for every flag:--X section, the owning workflow's init line must forward --X. It fails against the pre-fix files and covers any future atom, rather than spot-checking today's six. Verified through the workflow shape, not the CLI shape: `3 --prd spec.md` now yields ["prd-express-gate"] (was []), `2 --wave 2` yields ["partial-wave"] (was []). Refs #2993 * test(#2993): acknowledge the execute-phase ripple and regenerate install-tree fixtures Remote matrix was red with 46 unique failures, identical on both lanes. Both causes are mechanical consequences of changing shipped workflow content, and neither is visible to any local gate. - emitted-attribution: execute-phase.md grew 163 bytes from the WAVE_PARAM forwarding fix and was unacknowledged, while the ack fragment named plan-phase.md, which SHRANK and therefore needed no ack at all — a stale entry is itself a failure. The reason now names the real ripple. The entry had to merge into the existing 2930 fragment: the ack linter does unconditional cross-fragment duplicate-key detection with no spent/live exception, so a second fragment declaring execute-phase.md collides even when the first is already merged and inert. Resolved per the linter's own guidance and that file's precedent of appending successive ripple reasons to one entry. - golden-install-tree: tests/fixtures/install-tree/*.json are committed and deliberately excluded from the ADR-2719 attribution cutover, so they must be regenerated when shipped tree content changes. Regenerated after build:lib per the ordering landmine. 19 runtimes each gained exactly the six new plan-phase step files; zero paths removed, which is the absolute failure shape those fixtures exist to catch. Refs #2993 * fix(#2993): restore the launcher preamble in an extracted step and follow moved content in its drift guards Second red run: 26 unique failures, identical on both lanes, in two classes. RUNTIME BUG (runtime-launcher-parity, 7 failures) — chunked-planning-mode.md calls gsd_run but carried no canonical launcher preamble, which is what DEFINES gsd_run(). On any non-Claude runtime that step would fail outright. The preamble is now copied verbatim from the canonical source of truth, gsd-core/workflows/_runtime-launcher.snippet.sh, and the fence dedented to column 0 to match the prd-express-path.md sibling (a list-continuation indent breaks the byte-equal preamble match). prd-express-path.md already had a correct one. This is the same defect #2932 hit when it extracted steps; the parity test caught a real bug, not a stale assertion. DRIFT GUARDS (plan-phase-drift-guard, issue-2762-plan-reviews-chunked, skill-frontmatter-contract) — these assert plan-phase.md contains content this branch moved into step files. Retargeted at where the content now lives, with the asserted property unchanged; the ALL-RUNTIMES label COUNT test now reads host + every step file so the count is preserved across the split rather than reduced. Each retargeted guard was verified to still fail when its step file is stripped, so none was weakened into vacuity. No emitted-drift ack was needed: currentSizes() enumerates gsd-core/workflows/*.md non-recursively, so files under plan-phase/steps/ are never in the size ratchet's scope. Refs #2993 * chore(#2993): backfill changeset pr number to 3019 --------- Co-authored-by: sim <sim@local> |
||
|
|
a987cf2731 |
chore(#2932): emit a per-invocation section manifest from the init bundle (#2987)
* chore(#2932): emit a per-invocation section manifest from init Extends the init bundle with a typed per-invocation section manifest so an invocation loads only the branch guidance it will actually take. The three flag/state-gated branches in execute-phase.md move into their own step files; the parent keeps its gsd:section markers wrapping a one-line on-demand reference, so each section's prose lives in exactly one file and the parent shrinks 93369 -> 89507 bytes. A new drift-guarded generator derives the shipped section manifest from those markers, and a new pure evaluator maps invocation facts to applicable section ids. The evaluator is a lookup over the frozen WHEN_VOCABULARY, never a parser (Greenspun's Tenth Rule, ADR-1671:69); a parity test asserts the vocabulary and the predicate map stay exhaustively in sync. Closes #2932 * fix(#2932): fail closed on prototype-chain when values An isolated adversarial review found WHEN_PREDICATES[section.when] was a bracket lookup on a plain-prototype object, so inherited Object.prototype members resolved as predicates: "constructor"/"toString"/"valueOf"/ "hasOwnProperty" returned truthy and SILENTLY INCLUDED the section, and "__proto__" threw an untyped TypeError carrying no .reason. Both violate the module's documented fail-closed contract, and the manifest is read from disk at run time so it cannot be assumed trustworthy. Builds the predicate map on a null prototype and guards the lookup with an explicit Object.hasOwn check. Adds table-driven coverage for nine Object.prototype-shaped keys asserting the TYPED reason (asserting only that it throws would still pass while broken) plus a fast-check property injecting a hostile value at an arbitrary document position. * test(#2932): retarget execute-phase step assertions at extracted step files * fix(#2932): emit typed reasons for generator lib-load and write failures * fix(#2932): restore launcher preamble in extracted steps and refresh derived fixtures * chore(#2932): backfill changeset pr number to 2987 --------- Co-authored-by: sim <sim@local> |
||
|
|
4f6935e29b |
fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks (#2846)
* test(#2717): CommonJS marker for cursor/windsurf/codex staged .js hooks Cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) stage .js hook scripts via dedicated paths that bypass installSharedHooksBundle — the only writer of the {"type":"commonjs"} marker. Under a config root declaring {"type":"module"}, Node loaded those scripts as ESM and every require() failed with 'require is not defined', silently disabling the runtime's hooks. Adds regression tests (RED first, fix lands next commit): - parametrized cursor/windsurf/codex install asserts hooks/package.json exists with exactly GSD's marker content; - end-to-end: a cursor require()-using hook loads under a planted ESM-typed config root without the require-is-not-defined error; - the ensureCommonJsMarker / removeCommonJsMarkerIfGsdOwned contract: GSD markers are removed on uninstall, user-authored package.json is never touched. * fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks The {"type":"commonjs"} marker lived only inside installSharedHooksBundle, which cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) never reach. Their .js hooks are staged by dedicated paths, so under a config root declaring {"type":"module"} Node loaded them as ESM and every require() failed with 'require is not defined', silently disabling those runtimes' hooks. Decouple the marker write into a shared helper so any code path that stages .js hooks can ensure it lands in the SAME directory as the scripts: - src/runtime-hooks-surface.cts: add ensureCommonJsMarker(dir) + removeCommonJsMarkerIfGsdOwned(dir) (byte-identical content to installSharedHooksBundle's marker; preserves a user-authored package.json on both write and uninstall). Call ensureCommonJsMarker(hooksDir) from writeCursorHooksJson + writeWindsurfHooksJson; call removeCommonJsMarkerIfGsdOwned on their matching remove paths. Export both. - bin/install.js: call hooksSurface.ensureCommonJsMarker after the codex hook copy; call hooksSurface.removeCommonJsMarkerIfGsdOwned in the generic hooks-removal loop (safe no-op where no marker exists). No change to which runtimes receive the shared bundle, the !isCodex gate, skipSharedHooksInstall, or kimi/kimi-code/cline/copilot/trae/zcode (all unchanged — audit in the diagnosis). RED @ dbb7d2bb (6 failures: 3 missing markers + the ESM require error + missing helpers); GREEN pending. * docs(#2717): changeset fragment (pr:0, backfilled post-PR) * chore(#2717): regen codex/cursor/windsurf install-tree fixtures + attribution ack The fix adds hooks/package.json to those three runtimes' install trees (the new CommonJS marker), so the golden install-tree fixtures gain one path each (regenerated via npm run gen:install-tree). emitted-attribution (ADR-2719) flags the 3 emitted hooks/package.json paths under the hooks-built rule; acknowledge them. Also drops 5 spent ack entries left by now-merged PRs (#2694 code-review.md/code-review-fix.md, #2695 worker/registry, #2794 review.md) — they are stale on this branch (base already carries them). * fix(#2717): codex ESM-root behavioral test + hooks-built provenance for package.json Two review-driven follow-ups on the #2717 fix: - Adversarial review noted the ESM-root behavioral test covered only cursor; refactor it into a helper and add a codex case (the !isCodex-gated path most likely to regress, whose marker write lives in bin/install.js). gsd-check-update.js require()s at module load, so it surfaces the ESM failure immediately. - emitted-provenance flagged hooks/package.json as 'attributed source does not exist' — the marker is code-derived (a fixed literal emitted by ensureCommonJsMarker at install time), not built from a tracked source. Route the hooks-built rule's sources/transforms for package.json to the surface source file, mirroring the existing .cmd-shim sub-family. * chore(#2717): drop now-redundant hooks/package.json attribution ack The hooks-built provenance routing (prior commit) now self-attributes the emitted hooks/package.json to src/runtime-hooks-surface.cts, which IS in this diff — so the attribution is self-explaining and the emitted-drift-ack entry became stale. Delete the (now-empty) ack file per ADR-2719's empty-file rule. * docs(changeset): backfill #2717 PR number to 2846 |
||
|
|
c7c2fe3c2b |
fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd (#2680)
* fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd gsd-cursor-session-start.js and gsd-cursor-stop.js both resolved the project as path.join(process.cwd(), '.planning', 'STATE.md'). Under the cursor-agent CLI, hooks are invoked with cwd set to the Cursor config dir (~/.cursor), not the workspace — so the lookup always missed. sessionStart could only ever emit the "no .planning/ workflow found" nudge and stop's verify-work reminder could never fire, even with .planning/STATE.md sitting in the workspace. Slash commands were unaffected, which is why only the hook layer looked blind. Both hooks already buffered stdin into `raw` and never parsed it; the payload's workspace_roots carries the real path. Multi-root was left open in the report ("first root vs any root"). Resolved forward: prefer the first root that actually carries .planning/STATE.md, so a workspace whose GSD project is not the first root still resolves — strictly better than first-root-only and identical to it in the single-root CLI case. Falls back to roots[0], then to cwd, keeping IDE behavior unchanged if the IDE ever invokes hooks from the workspace. The resolver is duplicated verbatim across the two scripts rather than shared via hooks/lib/: these hooks ship standalone, and a new hooks/lib/ file must be registered in the GENERATED installer's GSD_HOOK_LIB_FILES allowlist — the installer-omits-shipped-file class that yields MODULE_NOT_FOUND at runtime. Per CLAUDE.md "Generative Fix Divergence", the duplication carries a parity assertion so the copies cannot drift. Failing-first, demonstrated by direct invocation with cwd != workspace: pre-fix sessionStart -> "no .planning/ workflow found" stop -> {} post-fix sessionStart -> ".planning/STATE.md is present" stop -> reminder tests/fix-2587-cursor-hook-workspace-roots.test.cjs spawns the real scripts as child processes with a cwd lacking .planning/ and workspace_roots pointing at it. Boundary coverage on the roots array (0 / 1 / 2 entries), plus malformed-JSON fail-open, junk-entry filtering, the parity assertion, and a guard that neither script resolves .planning from cwd again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): extend workspace_roots fix to subagentStart; keep cwd a candidate Three findings from the isolated review, all fixed. 1. MISSED SITE (high). gsd-cursor-subagent-start.js carried the identical defect at line 43 — its own header documents workspace_roots in the input schema, but it resolved .planning/ from process.cwd() anyway. Under the cursor-agent CLI that meant every Cursor subagent (planner, executor, verifier) started with "no .planning/ workflow found" and no phase context. The report named only sessionStart and stop; the defect class was wider. Verified pre-fix vs post-fix by direct invocation with cwd != workspace. 2. SEMANTIC NARROWING (medium). The first cut searched only workspace_roots and fell back to cwd solely when the array was EMPTY. So when roots were supplied but none carried .planning/ while cwd did, the hook reported absent — where the pre-fix code, which always used cwd, reported present. That contradicted the fallback's own stated intent of preserving IDE behavior. cwd is now a CANDIDATE in the search (`[...roots, process.cwd()]`), so the fix is a strict superset of both the old behavior and the CLI fix, never a narrowing. 3. STALE GOLDEN FIXTURES (high, would have failed CI). The golden-install-parity fixtures store a content hash per installed file; these three hooks appear in 13 of the 19 runtime fixtures. Regenerated via `npm run gen:golden` — the diff is exactly the three hook hashes in exactly those 13 runtimes. Tests extended: subagentStart resolution via workspace_roots; the stop hook's absent branch (previously only session-start's was covered); an explicit regression guard that a project at cwd is still found when roots miss; parity now asserts all THREE copies byte-identical; and the cwd guard sweeps the whole RESOLVING_HOOKS list so a future hook in this family cannot be left on cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * refactor(#2587): extract cursor workspace resolution to a shared hooks/lib module The duplicate-plus-parity-test approach was the wrong call. The reported issue named two hooks; a third (subagentStart) had the identical defect. That is the signature of a systemic problem, and three copies of a resolver guarded by a parity assertion is a divergence risk maintained by hand rather than a fix. hooks/lib/cursor-workspace.js is now the single implementation. All three Cursor hooks require it; none defines a local copy. Divergence is prevented structurally instead of by asserting three copies stay byte-identical. The reason duplication looked necessary was real, and is fixed properly here rather than worked around: Cursor sets hostBehaviors.skipSharedHooksInstall (#2089), so it never reaches the installer's bulk hooks/lib copy — it was the ONE runtime shipping these hooks WITHOUT hooks/lib (verified against all 19 golden fixtures: cursor had the hook scripts, no lib). A naive require would have thrown MODULE_NOT_FOUND at load, BEFORE each hook's own try/catch, wedging every session on precisely the runtime this bug is about. writeCursorHooksJson (src/runtime-hooks-surface.cts) now stages the hooks/lib helpers the staged scripts actually require, discovered by scanning their require('./lib/…') calls rather than a hardcoded name — so a future helper cannot be silently omitted. This is narrower than flipping skipSharedHooksInstall, which would wrongly pull in every shared hook. cursor-workspace.js is also added to GSD_HOOK_LIB_FILES so uninstall and the manifest manage it for the runtimes that do receive hooks/lib. Verified against a REAL install (runMinimalInstall, cursor/global): the helper is staged, and all three INSTALLED hooks resolve the workspace end-to-end from a cwd that is not the project. Also closes the review gap that the stop hook was excluded from the cwd-candidate regression loop — it now sweeps RESOLVING_HOOKS. The byte-parity test is replaced by a structural guard (every hook requires the shared module, none redefines it) plus a new install test asserting the helper is staged and the installed hook actually loads against it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): fail loud on a missing hook lib source; drop unsubstituted version marker Two findings from the installer-focused review. H1 — the staging step's `if (!fs.existsSync(libSrc)) continue;` silently defeated the very guarantee it was added for. Reproduced: delete hooks/lib/cursor-workspace.js from source, run the cursor install — it exits 0, prints "Done!", and ships the three hook scripts with an EMPTY hooks/lib/. The installed hook then throws `Cannot find module './lib/cursor-workspace.js'` at load, before its own try/catch, wedging every session — and nothing surfaces until a user hits it. The scan protected against a required-but-UNLISTED helper while leaving required-but-MISSING wide open (typo, bad rebase, an accidental delete). It now throws: a missing helper source is a packaging bug and aborts the install. M1 — hooks/lib/cursor-workspace.js carried a `gsd-hook-version: <placeholder>` marker that NOTHING substitutes: copyLibDir stamps .sh files only, and writeCursorHooksJson's staging applies just the colon-to-dash rewrite. Verified the literal was reaching disk on both the bulk (--claude) and Cursor (--cursor) paths. hooks/lib/git-cmd.js — the only pre-existing hooks/lib/*.js — carries no such marker, so this was newly introduced, not inherited. Marker removed, matching that precedent, with a note on why. (The explanatory comment deliberately does not spell the token out, or it would reintroduce the literal.) M2 — the require-scan regex demanded the exact compact form, so `require( "./lib/x.js" )` would silently fail to stage its helper and compound H1. Now tolerant of interior whitespace and either quote style. Regression test added for H1 — the reviewer confirmed the invariant had zero coverage repo-wide: a source tree carrying the hooks but no hooks/lib/ must make writeCursorHooksJson throw rather than produce a broken install. Re-verified end to end: the missing-source case throws, no unsubstituted literal ships, and the installed hook still resolves the workspace from a foreign cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2587): backfill changeset pr number (#2680) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6ee4349272 |
fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) (#2642)
* fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) * chore(#2537): backfill changeset pr to 2642 |
||
|
|
6ad30f74b6 |
feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler. harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run. Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard. Closes #2627 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f654c24a3e |
feat(#2505): Phase 4 — runtime-aware subagent dispatch (Option A; resolve-dispatch-type query) (#2525)
* feat(#2508): Phase 4 Option A — runtime-aware subagent dispatch via resolve-dispatch-type query (#2505) * fix(#2508): prose-variant preamble (avoid scanner-tripping literals) + namedDispatch===false-only mapping * fix(#2508): remove leftover old-preamble lines (keep prose variant only) * fix #2508: prose-only reference file * test #2508: regen golden install parity after workflow preamble additions * fix #2508: remove preamble from plan-phase.md (Phase 6 capstone ceiling); regen size+golden baselines * docs(changeset): backfill PR #2525 for Phase 4 (#2508) |