c19d3d7bdae02aaf836c79f27982d8be0451f507
202 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c5e0371775 |
feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging Red phase for issue #1951 (reversibility tagging: classify decisions by undo cost, gate one-way doors behind a checkpoint:decision). Tests assert, per the issue's acceptance criteria: - discuss-phase CONTEXT.md template records a **Reversibility:** field with a rationale on captured decisions, and states it is optional - gsd-planner @-references planner-reversibility.md and stays under the 49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs) - a one-way rating inserts a checkpoint:decision before the dependent task; reversible inserts none; costly is flagged but never blocks - the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard) and inserting a checkpoint implies autonomous: false - docs/reference/plan-md.md documents <reversibility> as optional with all three ratings - --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected into the planner prompt, and is advertised in the command argument-hint and help full mode (argument-hint parity) - the override suppresses the gate but still persists the rating - cmdVerifyPlanStructure accepts every rating and the absent case (additive-validator guarantee, behavioral via runGsdTools) - parity: thinking-models-planning.md #4 adopts the canonical three-level taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone - no content loss from the planner extraction made to fit under the cap Prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately — regression guards proving the validator already accepts unknown optional tags. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1951): reversibility tagging — gate one-way-door decisions Classify planning decisions by what undoing them would cost, and give a one-way door a human beat before the agent walks through it (issue #1951, The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way door framing). Acceptance criteria met: - discuss-phase records an optional reversibility rating with a rationale on <decisions> entries in the phase CONTEXT.md template. Unrated decisions are treated as reversible, so existing phases are unaffected. - a one-way rating makes gsd-planner insert a checkpoint:decision before the task that implements the decision, reusing the existing checkpoint mechanism -- no new checkpoint machinery. - reversible ratings trigger no checkpoint; costly ratings are flagged in the plan but never block. - the rating persists on the task as the optional <reversibility rating=> element. cmdVerifyPlanStructure accepts every rating and the absent case; the structural validator does not reject unknown optional tags. - --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings -- the override changes what stops the run, not what the plan remembers. Single taxonomy, not two: references/thinking-models-planning.md #4 already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the canonical three-level vocabulary and now points at planner-reversibility.md as the taxonomy owner, with a parity test that fails if the surfaces diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE). agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the checkpoint DO/DON'T guidance was relocated verbatim into planner-antipatterns.md -- already @-referenced from the same section for the same topic, so the planner still loads it and nothing was dropped. A test guards the relocation against content loss. Files: gsd-core/references/planner-reversibility.md (NEW, canonical taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase workflow/command/help (flag wiring + parity), plan-md.md schema, discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest, size baselines, install goldens, plugin skills regen, changeset. Closes #1951 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1951): address orthogonal review findings Two isolated reviewers (correctness + security), neither of which authored the change. Every finding fixed: Security — the rationale is untrusted input (ADR-1577). It originates in conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each hop an LLM reading the previous hop's output, with no validation on the path. planner-reversibility.md and the discuss-phase template now state it is data and never instructions, and name the </reversibility> early-termination hazard explicitly -- a rationale that closes its own element injects sibling structure the executor reads as real tasks. Four tests guard it. Correctness 1 — nothing machine-enforced the feature's own promise: a task rated one-way with no preceding checkpoint:decision validated as fully clean, so a planner error silently reopened the gap this feature exists to close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A warning, not an error: <reversibility> stays additive and the plan stays valid. Four tests cover ungated (warns), gated (silent), still-valid, and reversible/costly never flagged. Correctness 2 — pass-always test. The --no-reversibility-gates parse test substring-matched the whole workflow file, and plan-phase.md prose mentions both tokens in one sentence, so it passed with the bash conditional deleted: it was testing the documentation, not the parser. Now scoped to the fenced bash blocks and matched as one physical line, with a negative control confirming prose alone cannot satisfy it. Correctness 3 — costly had no itemized emission rule, only one-way did, so two agents could diverge on whether to tag costly at all. Correctness 4 — template convention break: the example ratings were bare while every sibling field uses [...] to signal substitution, inviting an LLM to copy one-way/costly forward as boilerplate. Now bracketed. Correctness 5 — latent false-green: .includes('reversible') also matches inside irreversible/irreversibility, which appear in anti-pattern prose, so a surface that dropped the real taxonomy entry would still pass. Now word-boundary matched. ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md 1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on next). The ceiling may only rise for privileged host machinery, and reversibility gating is optional-feature logic, so the wiring was slimmed to its minimum and the explanatory prose moved to the reference files the planner already loads. plan-phase.md is now 94400 bytes -- 119 under the ceiling and 70 bytes SMALLER than on next, so the host loop shrank while gaining the feature, which is what phase 6 ratchets toward. The tracer contract (tests/tracer-bullet.test.cjs) is unchanged. Lint — fixed an unnecessary non-null assertion in verify.cts and a CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST- PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture must carry the common task elements The gated-one-way fixture built a checkpoint:decision task from the abbreviated skeleton in gsd-planner.md, which shows only the checkpoint-specific elements (<decision>/<context>/<resume-signal>). cmdVerifyPlanStructure requires <name> and <action> on EVERY task regardless of type, so the fixture failed validation for reasons that had nothing to do with reversibility: errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"] Caught by gsd-test on 14d14a39 (2 failures, both this fixture). The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task carries <name>/<files>/<action>/<verify> like any other. Fixture corrected to match. Verified behaviorally against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Not a product defect: the validator's every-task contract is intentional and pre-existing, and docs/reference/plan-md.md scopes its required-element list to type=auto/tracer only because those are the elements a planner must author, not because checkpoints are exempt from <name>. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1951): backfill changeset pr number to 2471 * fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision Both CI failures were real defects in code this PR added, not false positives. CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 — the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`, which escapes the hyphen but not backslash, so the escape was incomplete. It was also unnecessary: `-` carries no special meaning outside a character class. Replaced with a complete metacharacter escape (backslash included). Word-boundary behavior verified unchanged across all three ratings — notably that "irreversible" prose still does not satisfy a "reversible" match, which is the false-green this helper exists to prevent. Prompt injection scan — the checkpoint fixture used the human-verification child element inside <verify>. That tag name is a fake-instruction-boundary pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed files, so copying the shape from tests/verify.test.cjs (unflagged only because it is not in this diff) tripped the gate. Switched to the documented plain-prose <verify> form. The first attempt at that fix failed the same gate a second time: the comment explaining the collision quoted the offending tag literally. The comment now names it in prose instead — the scanner does not care whether a match is code or commentary, which is the whole point of the DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md. Verified locally before push: scan reports 0 findings across 57 changed files, eslint clean, and both fixtures still validate as designed (gated one-way silent, ungated one-way warns, neither errors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): record measured cost and halve gsd-tools spawns The Windows shard 1/3 job timeout was traced to the sharding layer, not to this PR's assertions — see #2472. Two contributing factors were this file's own, and are fixed here. 1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs, so scripts/run-tests.cjs weighted it at the table's median fallback (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x under-weight. Recorded the measured value from the green gsd-test run (max across the node22/node24 lanes, per gen-test-timings.cjs's convention). Only this one entry: a full regen churns 634 entries of run-to-run drift, and the table is explicitly advisory and un-gated, so a 637-line diff does not belong in a feature PR. 2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost. Spawns cut from 9 to 6 with no coverage lost: - the ungated-one-way warning and its stays-valid assertion now share one plan instead of building the same plan twice; - the reversible/costly never-flagged-as-ungated test was strictly subsumed by the additive suite, which already runs those two ratings ungated and asserts no /reversibilit/ warning at all — and the gate warning's text contains both "reversibility" and "one-way", so the broader assertion catches it. It only re-spawned gsd-tools twice to prove the same thing. Both are symptom fixes. The shard imbalance itself (19/11/10 minutes against a 20-minute cap, from a cost-blind round-robin partition that also reshuffles downstream files whenever one is inserted) is tracked in #2472. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture adopts the #2444 type-branched contract Surfaced by rebasing onto next, which gained #2444 (branch plan-structure validation on task type=checkpoint:*) while this PR was in review. cmdVerifyPlanStructure no longer applies one required-element set to every task. A checkpoint:decision now requires <name> + <resume-signal> + <decision> + <options>, and is exempt from the <action>/<verify>/<done>/ <files> set that auto and tracer tasks carry. The gated-one-way fixture predated that split and failed on the new requirement: errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"] Fixture rewritten to mirror the checkpoint:decision contract exactly — real <options> with two <option> children — rather than padding it with fields checkpoints no longer need. That also drops the plain-prose <verify> the earlier revision carried purely to dodge the prompt-injection scan; a checkpoint task has no <verify> requirement at all, so the workaround is moot. Verified against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bb97ffb5aa |
fix(#2431): self-suppress TDD Audit section when all commits are missing (#2467)
* fix(#2431): self-suppress TDD Audit section when all commits are missing The TDD Audit section in ship.md step 8 was always emitted — but the execute pipeline only writes gate_status: git trailers when TDD mode is active. Without TDD mode (the default), every commit's trailer is absent and the section normalizes to 100% missing, producing a noise table with no way to disable it. Fix: add a self-suppress instruction at the point where gate_status values are normalized. When every commit in the scan normalizes to 'missing', skip both step 8 (TDD Audit section) and step 9 (aggregate gate_status trailer) entirely. Only emit when at least one commit carries a real value (skill, fallback, or exempt). This is data-driven, NOT config-gated. An earlier iteration used inline 'gsd_run query config-get workflow.tdd_mode' — but workflow.tdd_mode is owned by the tdd capability, and ADR-857 Phase 6 forbids host loop workflows from reading capability-owned keys via inline config-get. The self-suppress approach avoids any config-get entirely; it checks the actual trailer data and skips when there's nothing real to report. Matches the triage's suggested approach: 'have the audit gracefully degrade (skip the section)' when there is no real signal. Tests: tests/workflow-compat.test.cjs gains 3 #2431 assertions: - documents self-suppress when every commit is missing - step 9 (aggregate trailer) is also gated on real values existing - does NOT read workflow.tdd_mode inline (ADR-857 Phase 6 compliant) The existing feat-41 assertions still pass — the section content is preserved, only gated by the self-suppress instruction. References: #2431; PR #585 (consumer shipped, producer never wired); ADR-857 Phase 6 (capability-owned config keys must not be read inline by host loop workflows). * chore(#2431): backfill pr:2467 in .changeset/clever-moles-frolic.md |
||
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|
||
|
|
b6e6a22fce |
fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage (seven commits, never pushed) onto current origin/next as a single squashed commit. The original work was substantial and correct; this commit preserves its full scope, trimmed where rebase conflicts + workflow size budgets required it. Three independent layers where response_language was being dropped are closed: Layer 1 — orchestrator-facing directives across workflows. Adds the strong "All user-facing output in this workflow MUST be presented in {response_language}; technical terms, code, paths, and subagent prompts stay in English" directive to ~40 workflows that previously either lacked it entirely (verify-work, new-project, new-milestone, quick, manager, and ~35 more) or carried only the weak subagent-prompt-only form (plan-phase, execute-phase). The directive covers narration between tool calls and banner output, not just the AskUserQuestion prompts. Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts an optional responseLanguage parameter and renders the frame strings ("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.") in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/ Chinese/Korean/Italian) with an alias table covering ~30 input variants (en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads config.response_language via loadConfig(cwd) and passes it through, so the byte-for-byte block verify-work.md reprints verbatim is already localized when written — preserving the anti-injection hygiene rule at verify-work.md (the model is forbidden to translate after the fact). CJK display width is computed by East Asian Width property ranges (W/F) so the right ║ border of the banner stays aligned for full-width characters. English fallback is byte-identical to the pre-fix behavior when response_language is unset or unrecognized. Layer 3 — literal English report templates in execute-phase. The top-of- workflow directive covers all template sites (templates are a structural source, not literal output). Inline render-language notes that previously sat at each template site were removed during the squash because they pushed execute-phase.md over its frozen pre-phase-6 byte ceiling (93600 — ADR-857 Phase 6 capstone). The single top directive covers the same surface with fewer bytes. Also extends src/docs.cts and src/init.cts to propagate response_language into the init JSON bundle of the additional workflows so the directive can read it. Tests added: - tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls back to English default; recognized language swaps only the two frame strings while structural lines stay untouched; CJK display-width regression (independent recomputation of East Asian Width W/F ranges). - tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language wiring through docs.cts/init.cts. References: #2402; reporter's three-layer triage + Layer-4 follow-up; the byte-for-byte anti-injection hygiene rule at verify-work.md (the reason Layer 2 must be renderer-side, not model-translated). This is a squash of the in-flight bot branch — seven commits representing the original implementation plus its subsequent fix/CJK-padding/test/ changeset/regen cycles, none of which were ever pushed or PR'd. The squash captures the final coherent state. * chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md * chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451) Rebase conflicts were entirely in generated artifacts (golden-install-parity fixtures + workflow-size-baseline.json). After taking theirs during rebase, regenerated cleanly against the merged source tree. |
||
|
|
352876ff0c |
fix(#2315): respect review.default_reviewers in bare convergence invocation (#2451)
* test(#2315): regression test for review.default_reviewers precedence A bare /gsd-plan-review-convergence invocation (no reviewer flags) is supposed to let users configure a persistent reviewer lineup via review.default_reviewers and just run the loop. Instead, the orchestrator's argument parser silently discards that configuration and forces --codex on every no-flag invocation — with no warning that the configured reviewers were ignored. This commit adds a regression test that fails against the pre-fix workflow (the buggy unconditional --codex fallback is still present at this commit) and passes after the fix lands: - Structural: the workflow must NOT contain an unconditional 'if [ -z "$REVIEWER_FLAGS" ]; then REVIEWER_FLAGS="--codex"; fi' line before the workflow.plan_review_convergence config gate. - Structural: the workflow must query review.default_reviewers AFTER the config gate and document that empty REVIEWER_FLAGS lets gsd-review apply the default. - Behavioral: matrix across {configured, unset, empty-array, explicit-flag} invoking the actual deployed parse + resolution blocks with a stubbed gsd_run. Also updates two existing tests whose assertions the fix makes stale: - #2293 behavioral: endMarker was the buggy unconditional fallback line; the bare invocation assertion was 'run("5") === "--codex"'. Both flip post-fix (endMarker is now the last --all grep line; bare invocation returns empty from the parse block, default applied later in step 1.5). - command-default-claim: the pre-fix command documented '--codex (default if no reviewer specified)' which was the user-facing mirror of the bug. The assertion now requires the command to document the review.default_reviewers precedence. * fix(#2315): respect review.default_reviewers in bare convergence invocation Root cause: plan-review-convergence.md step 1 (Parse and Normalize Arguments) contained an unconditional fallback that set REVIEWER_FLAGS=\"--codex\" whenever no explicit reviewer flag was supplied. This value was then interpolated verbatim into the gsd-review args, so gsd-review saw --codex as an explicit flag (precedence rule 1) and never reached rule 3 (review.default_reviewers). The same path silently dropped any configured review.reviewer_instances (instances participate ONLY via review.default_reviewers per ADR-1517). Fix: - Remove the unconditional --codex fallback from step 1. - Add a config-gated resolution in step 1.5 (after CONVERGENCE_ENABLED check) that queries review.default_reviewers and either leaves REVIEWER_FLAGS empty (letting gsd-review apply its own rule-3 default) or falls back to --codex when no default is configured — preserving the pre-fix default for unconfigured users (AC3). - Replace the banner {REVIEWER_FLAGS} token with {REVIEWER_DISPLAY} so the startup banner reflects what will actually run (AC4), not a hardcoded value. The fix upholds the documented precedence contract (ADR-0011, ADR-0015) that the bug was actively violating. Explicit-flag invocations (--gemini, --all, etc.) are unaffected (AC5). References: #2315; ADR-0011 (review.default_reviewers precedence); ADR-0015 (autonomous cross-AI convergence); ADR-1517 (reviewer instances). * chore(#2315): bump plan-review-convergence.md baseline + changeset - Bump plan-review-convergence.md size baseline 23713 → 25536 (the new step-1.5 default-resolution block). - Add .changeset/plucky-yaks-roar.md documenting the user-visible change. * fix(#2315): restore /gsd: colon syntax + skip behavioral test when jq missing Two follow-ups to the #2315 fix discovered by gsd-test: 1. The fix commit accidentally regressed the slash-command namespace in the disabled-feature exit message: /gsd:plan-review-convergence (correct, from PR #3452) became /gsd-plan-review-convergence (retired dash syntax). The slash-command-namespace invariant test caught this. Restored the colon form. 2. The behavioral test exercises the deployed reviewer-resolution block, which pipes through jq. jq is a documented production dependency (review.md:244 "install jq if missing") and is present in every production deployment, but is NOT on PATH in the gsd-test linux-node{22,24} containers (same constraint as tests/opencode-review-reconstruction. property.test.cjs). Without jq, the printf|jq pipeline fails silently, the ||echo 0 fallback yields DEFAULT_REVIEWERS_COUNT=0, and the resolution falls through to the --codex branch — producing a false negative. Added a jqAvailable guard at module load (matching the existing pattern) and skip the behavioral test when jq is absent. The structural tests (no bash execution) still run and validate the fix. * chore(#2315): regenerate golden-install-parity fixtures Source changes to plan-review-convergence.md (workflow + skill mirror + command doc) changed the install-tree hashes. Regenerated via 'npm run gen:golden' after rebuilding gsd-core/bin/lib/install-engine.cjs ('npm run build:lib') — the on-disk lib was stale relative to src/install-engine.cts (isSymlinkedDestOptIn) and blocked fixture gen. * test(#2315): address review findings — strengthen structural tests + property test Code-review + security-review (isolated subagent passes) surfaced Low/Nit findings; this commit addresses the test-side findings: - Structural test 1 ("unconditional --codex one-liner") now asserts the buggy line is absent EVERYWHERE, not just before the config gate. The earlier assertion allowed a maintainer to re-add the line after the gate (passing the structural test) while the bug would still bite at runtime before the gate runs. - Structural test 4 ("banner uses REVIEWER_DISPLAY") now asserts the LITERAL banner placeholder "Reviewers: {REVIEWER_DISPLAY}" and forbids "Reviewers: {REVIEWER_FLAGS}". The earlier workflow.includes( "REVIEWER_DISPLAY") was satisfied by a comment mention. - New structural test for command/skill content parity: both files must document the review.default_reviewers precedence on the --codex flag (catches a manual edit to one that the gen:plugin-skills mirror misses). - Behavioral test stub now passes default_reviewers via env var ($GSD_TEST_DEFAULT_REVIEWERS) instead of inline-interpolating into a bash single-quoted string. Removes the (currently-safe) fragility where a future test input containing a single quote would close the bash quote and execute as bash under execFileSync. - New property test (fast-check, numRuns=25) for the JSON-classification contract: non-empty arrays of slugs -> empty REVIEWER_FLAGS; empty array / scalar JSON / malformed JSON -> --codex fallback. Locks the parser contract per CLAUDE.md mandate. * fix(#2315): address review findings — defensive jq-missing warning + banner cleanup Code-review surfaced two Low-severity workflow-side findings: - Defensive jq-missing warning: if jq is not on PATH in production (it is a documented dependency per review.md:244, but the dependency can be absent in degraded environments), the printf|jq pipeline fails silently to "0" and a user with review.default_reviewers configured gets --codex with no indication their configured default was unreadable. Added a command -v jq guard at the top of the resolution that falls back to --codex AND emits a stderr warning explaining the reason. This makes the failure diagnosable instead of silently reproducing the #2315 override. - Banner leading-space cleanup: REVIEWER_FLAGS accumulates with a leading space ("$REVIEWER_FLAGS --gemini" from ""), so the explicit-flag branch of REVIEWER_DISPLAY="$REVIEWER_FLAGS" rendered "Reviewers: --gemini" (double space). Pre-existing but worth fixing alongside the AC4 banner work. Strip one leading space with the ${VAR# } parameter expansion in the explicit-flag branch only (the configured-default and --codex branches already produce clean strings). * chore(#2315): bump plan-review-convergence.md baseline + regen golden fixtures The defensive jq-missing warning grew plan-review-convergence.md (25536 -> 26285 bytes). Bumps the per-file baseline snapshot and regenerates the golden-install-parity fixtures for the resulting install-tree hash changes. * chore(#2315): backfill pr:2451 in .changeset/plucky-yaks-roar.md |
||
|
|
eb45fc0e8c |
fix(#2415): close_phase_todos stages the pending/ deletion alongside completed/ (#2447)
* fix(#2415): close_phase_todos stages the pending/ deletion alongside completed/ Bug: the workflow step moved resolved todos from .planning/todos/pending/ to .planning/todos/completed/ with a plain 'mv', then committed listing ONLY the destination directory in --files: mv "$TODO_FILE" "$COMPLETED_DIR/" gsd_run query commit '...' --files .planning/todos/completed/ .planning/STATE.md Git's index still tracked the moved file at its old pending/<name>.md path. The commit therefore only staged the new completed/<name>.md copy — the deletion at pending/ was never staged, never committed, and lingered as an unstaged deletion in git status indefinitely until some later broad 'git add -A' caught it. The phase genuinely closed the todo, but the working tree was never clean. Fix: add .planning/todos/pending/ to the --files list. 'git add' of that directory (since git 2.0) stages deletions of tracked files in the pathspec, so the moved-away file is staged as a deletion atomically with the new completed/ copy in the same commit. Chose plain mv + two-dir --files over 'git mv' because git mv FAILS on: - untracked todos (new todo file not yet committed) - non-git .planning dirs (worktree safety / pre-init projects) Plain mv has neither failure mode. Regression tests in tests/close-phase-todos-stage-deletion.test.cjs (source-text-is-the-product: workflow .md text IS what the runtime loads) cover: - the commit --files list includes BOTH completed/ AND pending/ - the move uses plain 'mv' (not 'git mv') so untracked + non-git cases work * chore(#2415): trim commit subject to keep execute-phase.md under byte ceiling The fix added '.planning/todos/pending/' (~26 bytes) to the commit --files list. To stay under the ADR-857 Phase 6 pre-phase-6 byte ceiling margin (93400 bytes, hard ceiling 93600), shortened the commit subject from 'auto-close N todo(s) resolved by this phase' to 'close N resolved todo(s)'. Net change vs origin/next: +6 bytes (93384 → 93390), well under the margin. Also regenerates the golden-install-parity fixtures (execute-phase.md content-hash update across all runtimes). * chore(#2415): bump execute-phase.md workflow-size baseline (93384 → 93390) The +pending/ fix added 6 net bytes (93384 → 93390), still well under the ADR-857 Phase 6 pre-phase-6 byte ceiling margin (93400). * chore(changeset): backfill pr:2447 in .changeset/sturdy-wasps-run.md * fix(#2415): add issue ref to allow-test-rule annotation (ADR-456) CI lint-allow-test-rule-refs failed on the prior commit — ADR-456 requires '// allow-test-rule: <category> see #NNNN' so every exemption is traceable to an issue. Added 'see #2415' to the source-text-is-the-product annotation. |
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
1a46bc068a |
fix(#2376): emit absolute subagent-facing paths from init/state, convert workflow literals (#2428)
* fix(#2376): emit absolute subagent-facing init/state paths Make init.* and state.* path fields absolute rather than cwd-relative so subagent prompts resolve correctly regardless of working directory. Adds intel_dir/conflicts_path/requirements_path/roadmap_path/state_path to cmdInitIngestDocs, an absolute debug_dir to cmdStateLoad, and replaces bare .planning/... literals in 12 workflow Agent() prompt blocks with the absolute init-JSON path fields. Includes decoy-cwd regression tests and realpath'd tmpdir fixtures for macOS. Squashed rebase of the #2376 commit series onto a fresh origin/next (previous merge ee25543a1 was against a now-stale next). * chore(#2376): add changeset * chore(#2376): regenerate golden fixtures + workflow size baseline Regenerated after rebasing the absolute-path fix onto current next (picks up #2351's run-with-timeout content in execute-phase.md too). * fix(#2376): trim execute-phase.md redundancy to stay under the size margin * chore(#2376): regenerate golden/size baseline after rebase onto next |
||
|
|
d0bacc2517 |
fix(#2351): replace hardcoded timeout with portable run-with-timeout (#2426)
* fix(#2351): replace hardcoded gnu timeout with portable run-with-timeout Stock macOS ships neither `timeout` nor `gtimeout` (GNU coreutils). The 10 hardcoded `timeout <n> <cmd>` calls across the workflow/agent/reference gates exited 127 ("command not found") on such hosts, and the gates — which only distinguish 0/124/other — misreported a passing build or test as a FAILURE. Fix: a single Node-based `gsd_run run-with-timeout <secs> [--] <cmd…>` verb in gsd-tools.cjs. Coreutils-independent (stock macOS AND Windows), keeps GNU `timeout`'s exit-code contract (124 timeout, passthrough, 127/126 ENOENT/EACCES, 128+signum on signal), inherits stdio so pipes/redirects work, and reaps the whole process group so a watch-mode runner cannot outlive its budget. Runs before gsd-tools' flag parsing so the wrapped argv stays opaque. Hardened per adversarial review: - On timeout, SIGKILL the group SYNCHRONOUSLY before resolving — a descendant that traps SIGTERM was otherwise orphaned holding stdout, hanging captured gates (the exact watch-mode hang the feature prevents). - Forward SIGINT/SIGTERM to the child tree instead of dying and orphaning it. - Reject blank/whitespace <seconds> (was a silent unbounded run); clamp the timer to the 32-bit setTimeout ceiling (was a spurious immediate timeout). - Lint detector: catch GNU long options / `-k5` / `$((...))`; anchor to command position so prose "timeout 30 seconds" no longer false-positives. Resolution lives once in the CLI; all 10 sites call the shared verb. A parity guard (scripts/lint-portable-timeout.cjs, wired into lint:ci) fails the build if a bare `timeout`/`gtimeout` execution reappears (the portable `command -v timeout` probe form is intentionally allowed). Also fixes the identical bug in the zh-CN checkpoints translation, updates the tests that asserted the old strings, trims a redundant phrase in gsd-verifier.md to keep it under its size hard cap, and refreshes the size baselines + golden install-parity fixtures. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2351): add changeset (#2426) * chore: regenerate golden/size baseline after rebase onto next --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
40ce95f882 |
fix(#2358): scope review.md and ship.md temp files to a per-run mktemp directory (#2433)
* fix(#2358): scope review workflow temp files to a per-run mktemp dir /gsd-review wrote every prompt/section/output temp file to a hardcoded /tmp path keyed only on the phase number, so two GSD projects sharing a small phase number collide on the exact same path and a crashed run's leftover file becomes bait a later, unrelated run can silently read. ship.md's external peer-review stderr capture was strictly worse — one shared, unqualified path across every project/phase/run. Thread a single mktemp -d "${TMPDIR:-/tmp}/gsd-review.XXXXXX" run directory through every review.md temp path (67 sites) via a new {run_dir}/$RUN_DIR placeholder, mirroring the existing {phase} substitution mechanism, and clean it up at the end of the run. Route ship.md's stderr capture through a per-run mktemp file the same way. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2358): regenerate fixtures + lint gate-prep * fix(#2358): repair failing tests after gate verification * fix(#2358): thread RUN_DIR scoping into reviewer-instances.md (#1517) review.md's own invoke_reviewers step lazily loads gsd-core/references/reviewer-instances.md for the review.reviewer_instances codepath, but that doc was missed when review.md and ship.md were moved to the run-scoped {run_dir} temp directory. It still read the combined prompt from the old /tmp/gsd-review-prompt-{phase}.md (which build_prompt no longer writes, breaking reviewer-instances functionality outright) and wrote each instance's output to the old unscoped /tmp/gsd-review-${INSTANCE_NAME}-{phase}.md, leaving the exact cross-project temp-file collision bug open for that code path. Both paths now thread through {run_dir}, matching every other reviewer block in review.md. Extends the existing #2358 regression test with assertions pinning reviewer-instances.md's prompt read and output write to {run_dir}, and adds the Fixed changeset fragment. Regenerated the golden-install-parity content hashes for reviewer-instances.md via `npm run gen:golden` (paths unchanged; only the modified file's hash moved). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2358): backfill changeset pr (#2433) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a7d83dc234 |
fix(#2390): warn on goal-shaped phase.add titles, correct auto-detect docs (#2425)
* fix(#2390): phase.add title warning + auto-detect doc fix phase.add now returns a `warning` field when a description reads as goal-shaped (>80 chars and/or multi-sentence) rather than title-shaped, instead of silently writing the whole paragraph verbatim as the `### Phase N:` header. The CLI still creates the phase as-is (the strict two-layer slash-vs-CLI interface is unchanged); the warning just surfaces the gap. Also clarifies six doc sites (command argument hints, workflow detection steps, and how-to/reference docs) that described the phase-number argument as "auto-detecting" the next unplanned phase -- that detection is an orchestrating-workflow/LLM step reading ROADMAP.md (concretely: `query roadmap.analyze`'s `next_phase` field), not a `gsd-tools.cjs` CLI feature. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2390): regenerate fixtures + lint gate-prep * fix(#2390): repair failing tests after gate verification * chore(#2390): add changeset (#2425) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8d2f8bcb23 |
fix(#2388): gate shared requirement completion on sibling plans, revert on gaps (#2424)
* fix(#2388): gate shared-ID requirement marking and revert on gaps_found Adds requirements.ready-ids (execute-plan.md's update_requirements step) so a requirement ID declared by multiple plans in a phase only marks Complete once every declaring plan has produced a SUMMARY.md, and requirements.revert-phase (execute-phase.md's gaps_found branch) so a gaps_found verdict reverts the phase's own prematurely-Complete IDs before the gap report renders. Single-plan IDs still mark immediately. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2388): regenerate fixtures + lint gate-prep * fix(#2388): repair failing tests after gate verification * chore(#2388): add changeset (#2424) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
873bdf51e5 |
fix(#2352): expand tilde paths in review scope before the deleted-file filter (#2419)
* fix(#2352): tilde-expand SUMMARY.md key-files paths before deleted-file filter compute_file_scope's "Filter deleted files" step tested the literal `~/...` value from SUMMARY.md key-files entries with `[ -f "$file" ]`, which bash never tilde-expands (only a literal `~` in source text expands, not one arriving as an already-expanded variable value). Real files recorded with a `~/...` path were silently misclassified as deleted and dropped from REVIEW_FILES, and a phase whose every recorded file used a tilde path hit the empty-scope skip as a false negative. Adds a tilde-normalization loop as step 1 of post-processing (all tiers), before the deleted-file filter, rewriting a leading `~/` to `${HOME}/...` so downstream existence checks, the empty-scope short-circuit, and the FILES_TO_READ/CONFIG_FILES construction all see a real, openable path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2352): regenerate fixtures + lint gate-prep * chore(#2352): add Fixed changeset fragment (pr 2419) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
863a54ec82 |
fix(#2350): pass --raw to config-get in every build/test gate (#2399)
Adds --raw to config-get workflow.build_command|test_command reads in the post-merge, regression, verify-phase, and audit-fix gates so an unset key is a genuinely empty string, not the literal "" — restoring the auto-detect cascade and graceful skip instead of a false exit-127 failure. Regression guard sweeps all four gate files. Fixes #2350. |
||
|
|
b0f672f88c |
fix(#2337): capture and surface todo severity (#2381)
add-todo.md gains a confirm-based infer_severity step (infer from the blocker/major/minor/cosmetic taxonomy, confirm via AskUserQuestion with TEXT_MODE fallback, before writing) and a severity frontmatter field. cmdListTodos and cmdInitTodos now surface severity, backward-compatible (key omitted when absent), in parity. Closes #2337. Admin-merged (self-review bypass) with full green CI. |
||
|
|
ada79bee97 |
fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md (#2338)
* fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md Step 4 rewrote the `## Current Milestone` heading in the shared root PROJECT.md unconditionally. references/workstream-flag.md marks PROJECT.md `# Shared`, and per-workstream milestone state already lives in the workstream's own STATE.md / ROADMAP.md / REQUIREMENTS.md. With parallel milestones — the sanctioned design — whichever workstream ran new-milestone last silently won the shared heading. Step 4 is now skipped when a workstream is active; step 6 no longer stages PROJECT.md in that mode (cmdCommit returns nothing_to_commit rather than failing when a staged path is unchanged). Also fixes a second defect found while diagnosing this, same root cause (the workflow was workstream-unaware): step 1 parsed only --reset-phase-numbers and the milestone name, so GSD_WS was never set — yet ${GSD_WS} was interpolated at the routing lines. It always expanded to empty, so `/gsd:new-milestone --ws x` suggested `/gsd:discuss-phase [N]` with the workstream scope silently dropped, violating the routing-propagation contract. Step 1 now parses --ws using the established idiom from verify-work.md. Guard is keyed on GSD_WS, not $GSD_WORKSTREAM: the runtime launcher does not export the latter and it is only priority 2 of 5 in resolution, so it would miss the --ws flag case that is the actual repro. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2308): regenerate install goldens for the new-milestone workflow change gsd-core/workflows/ ships as an installed artifact, so new-milestone.md's content hash is pinned in all 18 runtime golden fixtures. Only that hash changed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2308): address review — inert step-6 guard, dropped Evolution repair, tautological tests Independent review found the first pass was partly cosmetic: 1. The step-6 `if [ -n "$GSD_WS" ]` branch was INERT. GSD_WS is assigned in step 1's shell and each step's bash block runs in its own shell — this file already proves it, since step 5 round-trips OUTGOING_MILESTONE through a file for exactly that reason (#2288). The guard read an unset variable, always took the flat branch, and staged PROJECT.md anyway. Rather than re-deriving GSD_WS in step 6, the branch is removed entirely: step 4 Part A's guard is what protects the shared heading, so post-guard the only content PROJECT.md can carry is Part B's idempotent Evolution backfill — which must be staged, not stranded. A regression test now asserts no cross-step GSD_WS branch returns. 2. Skipping ALL of step 4 also dropped the `## Evolution` structural repair — a shared, idempotent backfill that is not workstream state. A pre-Evolution project running only `--ws` would never get the section that transition and complete-milestone expect. Step 4 is now split: Part A (milestone-state write) is workstream-guarded; Part B (Evolution) always runs. 3. The tests were tautological prose-pinning — including one asserting a comment mentions "#2308". The step-6 test asserted the guard's TEXT was present, so it passed on the inert guard it existed to catch. Replaced with executable tests that extract the step-1 and step-6 fences and run them under bash with stubbed gsd_run, asserting real parse and --files behavior. 4. --ws is now stripped from the milestone name (step 1 previously left "--ws search" in the remaining text), and documented in argument-hint, help/modes/full.md, and docs/COMMANDS.md. 5. Changeset no longer overstates: --ws reaches the prose guard and routing hints only, not the SDK calls (state.milestone-switch/phases.clear/init.new-milestone still take no ${GSD_WS} — out of scope here). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * chore(#2308): regenerate SKILL.md, goldens, and size baseline for the argument-hint change skills/gsd-new-milestone/SKILL.md is generated from commands/gsd/new-milestone.md, so documenting --ws in the argument-hint made it stale (caught by lint:ci's gen-plugin-skills --check). Regenerated it plus the install goldens and workflow size baseline, since commands/, skills/, and gsd-core/workflows/ all ship. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2308): backfill PR number 2338 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1bb724048a |
fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist (#2325)
* fix(#2293): recognize --agy/--antigravity in plan-review-convergence whitelist The convergence reviewer-flag whitelist predated the 1.7.0 Antigravity CLI adapter and silently dropped --agy/--antigravity, so convergence fell back to --codex only and the working adapter was unreachable (worse after Gemini CLI's upstream shutdown). Add both flags to the workflow grep whitelist, the command argument-hint + flag docs, and the regenerated SKILL.md; they pass through to /gsd-review unchanged. --gemini behavior is untouched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2293): backfill PR number 2325 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
52fab7d9d7 |
fix(#2288): archive phase history under the outgoing milestone version (#2323)
* fix(#2288): archive phase history under the outgoing milestone version phases.clear derived its archive directory from a live getMilestoneInfo() read, but new-milestone.md switches the milestone BEFORE phases.clear runs, so phase history was filed under the NEW milestone's <version>-phases/ dir. Add a --archive-version override (threaded from new-milestone.md, captured before the switch) with precedence override -> live read -> dated label. Harden the version label against path traversal on both phases.clear and the sibling milestone-complete sink (the label is a moved directory name), and persist the outgoing version via a file + quoted shell expansion so untrusted STATE.md content is never re-parsed by the shell. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2288): backfill PR number 2323 into changesets Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
b041f101fb |
fix(#2287): surface unresolved deferred-items.md entries in progress + audit-uat (#2318)
The executor SCOPE BOUNDARY convention (agents/gsd-executor.md) logs out-of-scope discoveries to a phase directory's deferred-items.md, but no reader ever consumed it — forensic_audit, cmdAuditUat, and capture --list all skipped it — so deferred items were permanently invisible. cmdAuditUat (src/uat.cts) now scans each phase dir's deferred-items.md via a new parseDeferredItems (reusing the collectSection/splitGapsEntries/ extractGapEntryFields seams) and surfaces entries whose status != resolved (fail-safe: a missing/garbled status surfaces rather than hides, matching the false-negative-averse posture of #2286). forensic_audit (gsd-core/workflows/progress.md) gains Check 7 that globs .planning/phases/*/deferred-items.md and reports unresolved entries. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ff9cb6069f |
fix(#2285): wire claude-orchestration Workflow backend into execute-phase (#2314)
The claude-orchestration capability (#1143) shipped registered 'active' but fully inert: detectWorkflowBackend/emitWorkflowScript had no caller outside their own CLI router, and execute-phase.md declared an execute:wave:pre hook point that the workflow body never rendered — so claude_orchestration.enabled:true had zero effect on real runs. Approach B (maintainer-chosen): - execute-phase.md now renders the execute:wave:pre hook (gsd_run loop render-hooks execute:wave:pre) at a new step 2.75, immediately before each wave's Agent() dispatch — fixing the latent dead-hook gap for any pre-wave capability. - Move the claude-orchestration contribution execute:wave:post -> execute:wave:pre (a pre-wave backend selector belongs before dispatch, not after); rename fragments/execute-wave-post.md -> execute-wave-pre.md with prose instructing the orchestrator to call resolve-wave-dispatch before step 3. Unrelated wave:post contributions (ui.safety-gate, drift, external-job, mempalace) untouched. - New .cts seam resolveWaveDispatch(input) composes detectWorkflowBackend + emitWorkflowScript into one {backend:'inline'|'workflow', ...} result; exposed as gsd-tools claude-orchestration resolve-wave-dispatch. This is a real non-CLI-router, non-test caller of both functions. Fail-closed: any gate miss (disabled, non-Claude runtime, Workflow tool absent, SDK below floor, execution_backend:inline, malformed input) or an emit failure resolves to inline with a byte-identical result shape — no regression to the default-off execute-phase path. Regression tests (tests/fix-2285-*) cover happy-path activation + SDK-floor BVA, the fail-closed gate-miss table with detectWorkflowBackend parity, a fast-check composition property, capability.json contribution assertions, and a source-contract guard that execute:wave:pre is now actually rendered. Dependent registry-shape assertions updated in-scope. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f74442310d |
fix(#2257): auto-resume debug on non-terminal session-manager return (#2300)
The /gsd-debug orchestrator handled the gsd-debug-session-manager return with only two literal-string checks (DEBUG SESSION COMPLETE, ABANDONED) and no else branch, so a usable-but-non-terminal progress summary (the manager's own turn/context budget exhausted mid-loop, with a valid on-disk checkpoint) fell through to the user as if the debug were complete. Same gap at the continue subcommand. Callee side (agents/gsd-debug-session-manager.md): add an explicit non-terminal CONTINUE_REQUIRED return marker, distinct from the two terminal shapes and from a genuine user-input checkpoint. Orchestrator (gsd-core/workflows/debug.md Sections 4 and 1c): classify returns exhaustively — recognized terminal markers behave as before, anything else is non-terminal and auto-resumes by re-spawning the session manager from the same slug/checkpoint. Anti-loop guard: after two consecutive no-progress resumes (unchanged next_action/updated), emit a blocker report instead of looping. Regression test (source-text contract guard, fix-2196 idiom) asserts both sections' non-terminal/auto-resume branch, the CONTINUE_REQUIRED marker, and the anti-loop bound. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
315d94f6d4 |
feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode. - gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top. - gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer. - --no-tracer flag wired through plan-phase workflow/command/help/skill. - CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled. - tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1945): backfill changeset PR number to 2294 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8b70db343b |
fix(#2204): phase-completion writes 'All phases complete' per ADR-2207 (#2259)
* fix(#2204): phase-completion writes 'All phases complete' per ADR-2207 completePhaseCore was writing the overloaded bare 'Milestone complete' on the last phase — the same string space the milestone-close verb owns for terminal state. Per ADR-2207, phase-completion now writes the existing intermediate value 'All phases complete' (already used in gsd2-import.cts). Milestone termination ('<version> milestone complete' / 'Awaiting next milestone') remains solely with milestoneCompleteCore. Status lifecycle: Ready to plan → All phases complete → <version> milestone complete → Awaiting next milestone. Changes: - src/state-transition.cts: completePhaseCore status value - src/phase.cts: #2028 guard comment - tests/state-transition.test.cjs: assertion + test name - tests/phase.test.cjs: 8 assertion updates (positive + negative) - tests/state.test.cjs: normalizeStateStatus test case + reset regex - tests/workstream.test.cjs: fixture status to terminal value - gsd-core/workflows/progress.md: Route D label - gsd-core/workflows/transition.md: Route B label - CONTEXT.md: Status lifecycle glossary entry (ADR-2207) - .changeset/brave-geese-jump.md * test(#2204): regenerate golden-install-parity fixtures + workflow-size baseline Workflow file edits (progress.md, transition.md) changed install payload hashes and pushed past the committed workflow-size baseline. Regenerated all 17 golden-install-parity fixtures + claude-local via the standalone gen script (which now also covers the local-scope claude layout). Updated workflow-size-baseline.json and agent-size-baseline.json via size:baseline. * fix(#2204): correct claude-local golden hashes + document gen-script limitation The gen-script's claude-local generation produces macOS-specific hashes incompatible with Linux CI (local-scope install embeds platform-varying node-runner paths). Reverted to manual update using Linux FAILURES.md +actual hashes for the 2 changed workflow files. Added explanatory comment in the gen script. * test(#2204): add isCompletedInventory coverage + clarify CONTEXT.md glossary Addresses orthogonal code-review findings (Medium #1 + #2): - Add isCompletedInventory test cases for ADR-2207 status lifecycle (terminal 'milestone complete' → true; intermediate 'All phases complete' → false; archived → true; active statuses → false) - Clarify CONTEXT.md glossary: note that isCompletedInventory intentionally excludes the intermediate value * docs: backfill changeset PR number (#2259) * docs(#2204): add Status lifecycle table to state-md reference (ADR-2207) |
||
|
|
d49ac81306 |
chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 (#2248)
* chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 Phase 1 of epic #2143 (ADR-2143): consolidate markdown table parsing onto a canonical seam and migrate the pilot reader. - Add src/markdown-table.cts: parseMarkdownTable (GFM tables -> typed {columns, rows} addressed by column NAME; ragged rows are typed parse errors, not silent), a single-source TABLE_SCHEMAS registry (RoadmapProgress / RequirementsTraceability / QuickTasks / Security, with variants under one id), matchTableSchema, and findTableBySchema. Result<T> is scoped to this seam (distinct from the dispatch Result). - Migrate deriveProgressFromRoadmap (src/phase-lifecycle.cts) off the position-anchored regex to name-based resolution via the seam — fixes #2137 (the 5-column milestone-grouped Progress table previously returned all-null). - Add a schema-backed `gsd-tools quick-tasks-append` subcommand and route fast.md's log_to_state through it, retiring the inline `awk NF-2` column arithmetic — fixes #2133 (addresses #2012, #2119). Cell values are escaped (| and newlines) and the STATE.md read-modify-write is atomic under readModifyWriteStateMd (lost-update race, cf. #500/#905/#1230). - Writer/reader/template parity test guards TABLE_SCHEMAS against drift (ADR-2143 §3 Generative-Fix-Divergence). Registration: .gitignore, eslint.config.mjs, docs/INVENTORY.md + INVENTORY-MANIFEST.json, CONTEXT.md glossary, docs/CLI-TOOLS.md. Behaviour-preserving for the canonical 4-column Progress table; the named bugs are driven fail-first. Extend-never-mutate (ADR-2143 §2). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2242): backfill changeset PR number (#2248) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): escape backslash before pipe in markdown-table cell escaping CodeQL js/incomplete-sanitization (high): escapeCell escaped | -> \| but not the backslash itself. Now escapes \ -> \\ before | -> \|, and splitTableRow unescapes both \\ -> \ and \| -> | symmetrically so cell values (incl. literal backslashes) round-trip exactly. Added backslash round-trip tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): read ROADMAP Progress table by column name — supersede #2168 ad-hoc scan Rebase reconciliation with #2168 (the tactical #2137 fix that marked itself "pending #2143"). deriveProgressFromRoadmap now resolves the Progress table via a new seam helper findTableWithColumns (first table whose header is a superset of Phase/Plans Complete/Status/Completed, any order, extra columns ignored) and reads cells by NAME — order/injection-invariant per ADR-2143 §3 — instead of the exact TABLE_SCHEMAS match. This satisfies #2168's column-invariance property test while staying seam-based and preserving its `## Progress` scoping (#2012/#1445). Ragged Progress tables now resolve to null (ADR-2143 fail-loud); updated the stale state.test.cjs assertion that predated the Phase-1 migration. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
98e4233ce9 |
fix(#2176): ground the Antigravity reviewer in the repo under review (#2184)
* fix(#2176): ground the Antigravity reviewer in the repo under review - capability-probe --add-dir (mirrors the Codex bypass-flag probe) and pass the repo root on both invocation arms - anchor _AGY_PROMPT to the absolute repo root; mandate a REVIEWED-WITHOUT-REPO-ACCESS self-report when the repo is unreadable - stamp a [reviewed-without-repo-access] marker on self-reported or scratch-anchored output; Consensus Summary down-weights marked reviews - apply the same absolute-root anchor to the cursor-agent prompt (AC5) Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2176): changeset fragment for PR #2184 Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2176): review fixes — size baseline, cursor root anchor, anchored blind tells - regenerate tests/workflow-size-baseline.json for review.md's growth - cursor anchor uses git rev-parse --show-toplevel (bare pwd resolved the wrong root from a repo subdirectory) - blind-review tells anchored: self-report to the first lines of output, scratch tell to a workspace-declaration phrasing — a grounded review quoting either string is no longer mis-stamped Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2176): round-2 review fixes — scratch-tell bridge, behavioral test, changeset Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the review.md change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test(#2176): pass the transcript path to bash with forward slashes The behavioral detection test substitutes a mkdtemp path into the bash compound; on Windows runners that path contains backslashes, which bash strips, so the transcript is never found and the first assertion fails (windows-latest/24 lane). Git Bash accepts D:/-style paths. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2176): use /gsd:review namespace syntax in workflow comment The slash-command namespace invariant (#3443) bans retired /gsd-<cmd> references in Claude-facing sources; a cursor-anchor comment used /gsd-review. Size baseline + golden fixtures regenerated for the byte change. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test(#2176): derive the POSIX path via path.sep, not a hardcoded separator Review finding: out.replaceAll('\\', '/') hardcodes both separators; use the separator-safe out.split(path.sep).join(path.posix.sep) idiom. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test(#2176): use the merged toPosixPath seam for the bash path Per maintainer note: #2247's shell-command-projection now centralizes running-OS → POSIX path conversion; import it instead of the inline split/join idiom. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg |
||
|
|
3592697bed |
fix(#2107): orchestrator honors gate="blocking-human" checkpoints in auto-mode (#2113)
* fix(execute-phase): honor gate="blocking-human" in auto-mode checkpoint handling The package-legitimacy gate (#2827) spans two layers. gsd-executor refuses to auto-approve a gate="blocking-human" checkpoint and escalates it so a human can vet the package. execute-phase's checkpoint_handling step then dispatched purely on checkpoint *type* and never read gate -- so under --auto/--chain it auto-approved the checkpoint the executor had just refused to auto-approve. Net effect: the slopsquatting defence was inert in exactly the unattended mode where it matters. An [ASSUMED]/[SUS] package reached install with no human ever seeing the prompt. - gsd-core/workflows/execute-phase.md: carve out gate="blocking-human" (and the package-legitimacy what-built markers) ahead of every auto-mode branch. - gsd-core/references/checkpoints.md: document the gate attribute and its two values. blocking-human previously appeared nowhere outside gsd-executor.md, so no planner had a documented way to author a non-auto-approvable checkpoint. - tests/package-legitimacy-gate.test.cjs: the existing regression test asserted the executor half only, which is why it stayed green while the gate was open. Now asserts the orchestrator half too. * chore(changeset): link to issue #2107 * chore(changeset): backfill PR number 2113 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa * test(#2107): refresh golden-install-parity hashes for edited gsd-core files The golden fixtures pin content hashes for gsd-core/references/checkpoints.md and gsd-core/workflows/execute-phase.md, both edited by this fix. Regenerated via UPDATE_GOLDEN=1; only those two keys change across all 17 runtime fixtures. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa * fix(#2107): keep the carve-out inside the ADR-857 host-loop budget The ADR-857 phase-6 ratchet pins execute-phase.md below 93600 LF bytes so optional-feature logic keeps migrating out of the host loop. The carve-out first landed 623 bytes over that ceiling. Move the two-layer rationale (why gsd-executor escalates these checkpoints) into references/checkpoints.md, where the gate is now documented, and reduce the workflow to the operative rule. execute-phase.md is 93589 bytes, under the ceiling; the gate token and both <what-built> marker strings are kept because the orchestrator matches on them. Refresh the two baselines the edit invalidates: golden-install-parity fixtures (only the checkpoints.md and execute-phase.md hashes move) and workflow-size-baseline.json (one line). The ADR-857 ceiling itself is untouched. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JNR8m2pv5U7ubn4iiXVrMa * fix(#2107): executor honors blocking-human on the decision branch + gate transport Review found the fix incomplete one layer down. Two executor-layer gaps: 1. Blocker — agents/gsd-executor.md auto-mode dispatch gated checkpoint:human-verify on gate="blocking-human" but the checkpoint:decision branch below auto-selected the first option with no gate check. The executor resolves a decision itself (auto-selects and continues) without returning it, so the orchestrator carve-out never runs for it. A planner following the new checkpoints.md rule 6 ("gate a decision whose default would be wrong to assume") would have it silently auto-selected under --auto/--chain — the exact #2107 harm, one checkpoint type over. The decision branch now STOPs and returns for an explicit human decision when gate="blocking-human". 2. Major (transport) — checkpoint_return_format carried no field conveying the gate to the freshly-spawned orchestrator, so recognition of the proactive pre-install checkpoint rested on freeform prose. Added a **Gate:** field to the return format and re-pointed the execute-phase carve-out at it ("If the returned Gate: is blocking-human"). Net byte-negative: execute-phase.md drops 93589 -> 93583, widening ADR-857 headroom from 11 to 17 bytes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2107): cover decision carve-out + gate transport, de-vacuum conditional tests - New: 'auto mode does not auto-select a blocking-human decision checkpoint' asserts the executor decision branch STOPs on blocking-human. Verified red on the pre-fix executor (2 fail), green with the fix (27 pass). - New: 'checkpoint_return_format transports the gate ...' asserts the **Gate:** field carries blocking-human across the executor->orchestrator boundary. - New: 'auto-select rule for decision is conditional' — orchestrator-side mirror of the human-verify conditional test, for the execute-phase decision branch. - Fix vacuous test: both conditional tests now assert the anchor matched (length > 0) before iterating, so anchor drift can no longer pass with zero assertions. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2107): refresh golden + size baselines for executor + execute-phase edits Regenerated via UPDATE_GOLDEN=1 and update-size-baseline.cjs. Only the gsd-executor.md and gsd-core/workflows/execute-phase.md hashes move across the runtime fixtures (35 ins / 35 del, no keys added or removed); checkpoints.md is unchanged this round. Size baselines: gsd-executor.md 43607 -> 43973, execute-phase.md 93589 -> 93583 (still under the ADR-857 ceiling). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
b5ce72f729 |
fix(#2119): single SECURITY.md writer — auditor is return-only (#2154)
* fix(2119): single SECURITY.md writer — auditor is return-only The gsd-security-auditor held Write/Edit and was instructed to write SECURITY.md (no <N>- prefix, no template frontmatter), while the orchestrator's Step 6 also wrote the correct padded <N>-SECURITY.md from templates/SECURITY.md. Two writers, two naming conventions, two shapes — the auditor's unprefixed file was invisible to the workflow's *-SECURITY.md glob detector and unparseable for the threats_open gate. Fix (option 1 from the issue): make the auditor return-only. - Remove Write/Edit from auditor's tools - Rewrite all 'Write SECURITY.md' instructions to 'Return structured verdict' with threats_open count - Add explicit constraint in workflow Step 5 spawn prompt - Update existing test (was asserting Write in tools — now asserts absence) - Add new regression test for single-writer contract - Update docs/AGENTS.md stale Tools/Produces rows - Regenerate golden fixtures + agent size baseline * docs(changeset): backfill PR number (#2154) * chore(#2119): regenerate pi/qwen golden fixtures after next merge The single-writer change edits gsd-core/workflows/secure-phase.md and agents/gsd-security-auditor.md; pi.json (added on next) and qwen.json (merge straggler) were the only runtime fixtures still holding pre-change hashes for those files. All other runtimes already reflect the change. Regenerated via the sanctioned gen-golden-install-parity script. * merge origin/next — regenerate goldens + baseline for merged state * fix slash-command syntax: /gsd-secure-phase → /gsd:secure-phase (#2154 CI fix) |
||
|
|
4bb846b67a |
fix(#2112): scope commit to --files pathspec, not entire index (#2148)
* fix(2112): scope commit to --files pathspec, not entire index cmdCommit/cmdCommitToSubrepo/cmdPrSubrepo staged exactly the files named in --files but then ran a bare 'git commit' with no pathspec, absorbing anything else in the index into a commit whose message described only the named files (#2112). Fix: append '-- ...stagedPaths' to the commit args when the caller declared a scope. Three guards are load-bearing: - stagedPaths (not filesToStage) excludes skipped missing files (#2014) - explicitFiles gate keeps the default .planning/ path byte-identical - MERGE_HEAD check via 'git rev-parse' falls back to bare commit during merge - --amend is left without pathspec (different operation) cmdPrSubrepo pathspec uses changedFiles (old+new for renames) so the full rename is captured atomically. Also fixes workflow markdown in spec-phase.md and add-tests.md. All-files-missing now short-circuits to nothing_to_commit instead of absorbing the entire index under a message describing files that were not committed. * docs(changeset): backfill PR number (#2148) * test: update golden-install-parity fixtures for workflow markdown changes (#2112) * test: update golden fixtures + workflow baselines for #2112 changes - claude-local.json golden fixture (now generated via gen script) - workflow-size-baseline.json (add-tests.md +16, spec-phase.md +42 bytes) - Extended gen-golden-install-parity-zcode.cjs to also regenerate the claude local-layout fixture |
||
|
|
f8c5c1590f | fix(#2196): declare the debug session-manager spawn foreground + no-TaskOutput + recovery (#2227) | ||
|
|
24c324cbfa |
fix(#2194): add Bash timeout guidance for prompt-fed reviewers in review.md
The Gemini, Claude, and Codex reviewer blocks invoked the CLIs with no explicit timeout, so each inherited the host default (~2 min on Claude Code). A source- grounded review of a large plan set takes ~570s (Codex xhigh) / ~525s (headless Claude) — both exceed that window, so the lane is killed mid-review, its output is empty, and the cross-AI review silently proceeds with fewer lanes. CodeRabbit and OpenCode already documented a timeout; the four main lanes did not. Add a shared timeout-guidance note directing a high Bash timeout (>= 900000; 1200000 for Codex xhigh / headless Claude), referencing BASH_MAX_TIMEOUT_MS for the Claude Code host cap, and framing a slow-lane empty output as a timeout kill (not the 0xc0000142 crash it gets misdiagnosed as) so operators re-run with more time instead of diagnosing a CLI failure. Closes #2194 Recaptures the 18 golden-install-parity fixtures + workflow-size baseline (only the review.md entry changed in each; review is LARGE-tier, 47168 < 61440). |
||
|
|
924ff6822e |
fix(#2138): surface track_shipping push failures (review)
Review (MEDIUM): `git push ... 2>&1` did not check exit code, so a silent push failure (auth-token expiry, network blip, non-fast-forward) would proceed to the report step and declare success — silently reproducing the exact #2138 defect. Add a fallback warning naming the rerun command so a failed push is visible (best-effort: the PR already exists, so we still report it rather than abort). Recaptures goldens + size baseline. |
||
|
|
aaf74878f7 |
fix(#2138): push the track_shipping ship-note onto the PR branch [ci skip]
track_shipping committed the STATE ship-note ('Phase N shipped — PR #N') AFTER
create_pr and never pushed it, so the commit stayed local-only. When the GitHub
PR merged (especially fast/auto-merge) the ship-note was not in the source branch
and never reached the default branch — STATE's ship-status was silently lost,
recoverable only by STATE self-heal on the next /gsd-start.
Push the ship-note commit onto the PR branch with a [ci skip] trailer. GitHub
honors [ci skip]/[skip ci], so this lands the note on merge without triggering a
redundant pipeline, and preserves the PR number in STATE.
Recaptures the 18 golden-install-parity fixtures + the workflow-size baseline
(only the ship.md entry changed in each).
Closes #2138
|
||
|
|
78b9b04f1d |
fix(#2133): correct fast.md log_to_state column-count gate (NF-2)
The guard used `awk -F'|' '{print NF-1}'` but a markdown header has a leading
and trailing pipe, so NF counts (real columns + 2); NF-1 is always one too
high. The `-eq 5` test was therefore unsatisfiable for the very 5-column
header quick.md writes, so /gsd-fast has never appended a Quick Task row since
PR #85 (regression closing #27).
- Count real columns with NF-2 (5 for the 5-col header, 6 for the 6-col).
- Accept 5 OR 6 columns (quick.md Step 7b writes both shapes).
- Select the appended row template by the detected count so its cell count
always matches the header — keeps #27 fixed for the validate-mode table.
Closes #2133
|
||
|
|
b55a2e7655 |
fix(#2117): distinguish not-yet-validated phase from validated failure in audit-milestone
audit-milestone's Nyquist scan classified a phase from `nyquist_compliant` alone, so a phase seeded by plan-phase but never run through validate-phase read PARTIAL — identical to a phase that validated and genuinely failed. The template's `status` field could discriminate the two, but no workflow ever promoted it off `draft`, so it was dead. Make `status` live and read it: - validate-phase.md §6: set `status: validated` in both the create (State B) and update (State A) VALIDATION.md paths. - audit-milestone.md §5.5: parse `status`; add a distinct NOT-VALIDATED bucket keyed on `status: draft`, gate COMPLIANT/PARTIAL on `status: validated`, and report `not_validated_phases` in the audit YAML. - VALIDATION.md template: document the draft → validated lifecycle. Tests & generated artifacts: - Regression test folded into policy-138 (owning workflow-contract file); fail-first verified vs origin/next (0 matches pre-fix). - Regenerate golden-install-parity fixtures cleanly: adds the previously-missed qwen.json and removes a contaminated `settings.local.json` entry that had leaked into claude-local.json (the harness excludes hook-config files). - Correct a stale validate-phase.md workflow-size-baseline entry. Closes #2117 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
79d7657eff |
feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.
Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
slash-programmatic / active-model / native-extension / bun) + hostBehaviors
{nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
(global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
markdown); the 16 other fixtures + claude-local change only by the shared
model-catalog hash line.
Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
(every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
consumes params; getArgumentCompletions; before_provider_request active-model
steering (fail-open on null resolution); functional session_start /
before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
vocabulary.
Docs (host-integration matrix + how-to) + changeset (Added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
f014ec83bd |
feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads: finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks copy), and a skills converter-name registry (the artifactLayout.converter field is now load-bearing, not decorative). frontmatterDialect stays the documented dispatch key for frontmatter (no descriptor field for it). Dead isKilo destructure bindings removed. Byte-identical golden parity for all 16 runtimes (opencode, which shares kilo's combined-family path, verified clean). UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin + extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus). UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented' per AC so dispatch degrades to 'degraded' by design. Model-catalog single-source edit ripples the shared model-catalog.json hash into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale- bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/ codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error. Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/ hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades (plugin parity+load, model-override converter, agents dispatch surface, MCP doc). Matrix + how-to + config docs updated; changeset added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a8ff8fbb30 | docs(#1867): specify propose-then-confirm --auto behavior in ui-phase Step 9.5 (review Minor) | ||
|
|
31a500b970 | docs(#1867): replace stale plan-phase.md:921 line-pointer with section-name reference (review #6) | ||
|
|
830670f288 |
docs(#1867): clarify Step 9.5 manual element-extraction is by-design (review Minor 1)
trek-e's re-review asked to either mechanize the Step 9.5 element extraction or add an inline note justifying why it stays manual, so a future maintainer doesn't read it as an oversight. Verified the premise: the requirement-side edge-probe path (spec-phase Step 5.5) it mirrors is ALSO a hand-populated heredoc + fail-loud <replace:> placeholder guard — not a mechanical parse. The UI probe mirrors that idiom verbatim. Mechanizing would be worse: a UI-SPEC has no single machine-parseable "elements" column (surfaces are spread across the design-token tables, Copywriting, and researcher-named prose), so a regex/table parse would fail-OPEN (miss a prose-named surface, or feed a design-token row as a bogus element). Expanded the inline comment to state the parity + the fail-open rationale explicitly. No logic change; regenerated the workflow size baseline and the 16 golden-install-parity fixtures for the +847 B comment. Refs #1867 Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ |
||
|
|
bea7196c4c |
feat(#1867): wire ui-consideration probe into ui-phase (WIRE-01)
Add the live ui-phase producer path for the UI-consideration probe (Phase 2, WIRE-01). Two small exports on the Phase-1 adapter — proposeElements (the propose-then-confirm view of detected kinds + applicable categories) and autoResolve (the deterministic --auto floor that never dismisses and never auto-backstops an unclassified item, #1110) — plus a post-verification '## 9.5 UI-Consideration Probe' step in ui-phase.md mirroring spec-phase 5.5's RUNTIME_DIR shim + fatal-invoke/malformed-report/zero-applicable fail-closed guards, propose-then-confirm (the partial-cue recall mitigation), and the '## UI Considerations' write-back in the shipped plan-phase.md:921 lift format. autoResolve is the CODE floor; the covered-upgrade stays workflow prose (the two-layer --auto). Un-upgraded backstops route to insufficient_spec -> human_needed at verify, never a silent pass (#1154). Tests: +8 typed (proposeElements shape/determinism, autoResolve never-dismiss, partial-cue strict-subset) — structured-value only. ui-phase.md size baseline ratcheted 15477->24447 (under DEFAULT cap). The plan-phase.md PRE_PHASE6 ceiling stays RED pending #1852 (unchanged from Phase 1). Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS |
||
|
|
6d7450c465 |
feat(#1867): lift UI Considerations into plan-phase must_haves (LIFT-01)
Separate *-UI-SPEC.md glob (edge-coverage glob exclusion at :746 left intact
- Hyrum/D-08); a terse lift bullet reusing the ## Edge Coverage rule verbatim
(covered -> truths string, backstop -> flat scalar {statement, verification:
backstop}, unresolved -> assumption; no new verb - ADR-550 #1278/#1154); and a
no-silent-drop checklist line. Lift logic lives ONLY in the plan-phase
workflow, never gsd-planner.md (D-10, agent-size cap).
NOTE: this grows plan-phase.md +862B, over the #1168 PRE_PHASE6 ceiling (94519)
by ~802B, so the workflow-size gates are RED until #1852's lazy-split lands the
-21KB headroom. Deliberately does NOT raise the ceiling constant (would collide
with #1820/#1835's in-flight raise). #1867 sequences after #1852; rebase +
regenerate the size baseline then.
|
||
|
|
185abe2d66 |
feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna), advancing from the superseded GPT-5.4/5.5 generation. Model IDs verified against OpenAI developer API docs: - gpt-5.6-sol: flagship, /, reasoning xhigh - gpt-5.6-terra: balanced, .50/, reasoning medium - gpt-5.6-luna: fast/cheap, /, reasoning medium Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast), so profile semantics are unchanged — only the underlying IDs advance. Updates: catalog JSON, test assertions (catalog defaults), docs (CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings), and changeset. Closes #2122 |
||
|
|
dbc730d8de |
fix(#2073): capability-probe external killer (timeout/gtimeout) for macOS
Code review (HIGH): a hardcoded 'timeout 600 agy' fails with rc 127 on stock macOS (no GNU timeout/gtimeout), silently losing the agy reviewer. Probe for 'timeout'/'gtimeout' via command -v and fall back to agy's native --print-timeout alone when neither exists (mirrors scripts/base64-scan.sh). External cap (600s) stays >= --print-timeout (540s) so it only backstops a pre-session stall. Factor the prompt into _AGY_PROMPT to avoid duplicating the long -p string across both branches. Update the agy + #687 tests to assert the probe + bound + fallback, regen the 17 goldens + size baseline, refresh the maintainer-note version stamp to 1.0.16. |
||
|
|
af9069f865 |
fix(#2073): harden agy reviewer block (arg overflow, 404, pre-session stall)
Three failure modes on agy 1.0.16, all fixed by mirroring the Cursor block's
invocation discipline:
* file-reference prompt instead of inline "$(cat)" — a large review prompt
overflowed the exec arg list (rc 126).
* external 'timeout 600' wrapper — --print-timeout cannot fire before agy
creates a session, so a pre-session stall hung unbounded.
* --model from review.models.agy when set — escape hatch for a pinned model
that 404s (exit 0, empty stdout + transcript).
* stdin </dev/null so agy never blocks on a tty.
Also enrich the Step 3 empty-output stub to grep agy cli.log for a
model-availability diagnostic, and correct the stale 'no --model flag' note
plus the 'review.models.agy reserved for future' comment (the config key was
already read but never passed through).
|
||
|
|
1ee00320c7 | Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns | ||
|
|
4483300253 |
fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.
Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
→ ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
→ REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
→ FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
reviewer's own override was ignored); init.quick now resolves `reviewer_model`
(gsd-code-reviewer) and the spawn threads it.
resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.
Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.
Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.
Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
6addeccd19 |
feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates an external API/SDK/service can no longer seal without a decided coverage matrix. - src/api-coverage.cts: deterministic detector (compound verb+noun signal + <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix parse/validate/render with field-length caps. - check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as a token under .planning/phases/ only (traversal-neutralized); validates COVERAGE.md or blocks iff a strong integration signal is detected and no matrix exists; fail-closed when phases tree exists but phase unresolvable. - capabilities/ai-integration: workflow.api_coverage_gate config key (default true), plan:pre contribution, blocking verify:pre gate. Data-driven. - gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch. - Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e. Code+security review findings fixed (stopword FP, scope containment, pipe/cap rejection, prompt-injection message hygiene). - Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs. Closes #1562 |
||
|
|
603593d41d |
fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest defaults to WATCH mode in an interactive TTY — exactly where a user runs `gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never exited and the orchestrator waited indefinitely. Recovery needed the user to manually prompt "something blocking?". Fix — one shared helper + a bounded, surfacing timeout on the test-command gates: - New pure module src/normalize-test-command.cts + `gsd-tools query normalize-test-command` verb: rewrites a resolved command to a best-effort one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`; a package-manager `test` script whose package.json runner is watch-vitest → `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged — never double-flagged). Named `normalize-test-command` (not `test-*`) so the file does not match node --test's default `test-*` discovery glob. - The three gates that HUNG or silently-continued route through that ONE helper and bound execution with `timeout $(config-get workflow.test_gate_timeout)` (new config key, default 600s): the regression gate (extracted to execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen — it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124. - verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint on 124, staying under its frozen 40960-byte tier cap. Security hardening (review): the normalizer only rewrites a runner named as a standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never mangled), is length-capped and uses only linear-time split-based scanning (no super-linear backtracking on an adversarial `workflow.test_command`), and reads package.json only when it is a regular file (never blocks on a FIFO via `--dir`). Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores, inventory manifest/index. All 16 golden-install-parity fixtures + workflow size baseline regenerated for the changed shipped files; bin/lib is excluded from the parity manifest. Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route through the shared helper + configured timeout + exit-124 watch-mode hint; verify-phase asserted as normalize-only/already-bounded). tests/execute-phase-active-flags.test.cjs repointed at the extracted step; tests/planner-language-regression.test.cjs allowlist comment updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ea4378063f | Merge branch 'next' into feat/1820-specless-predicate-rail | ||
|
|
d7129222c0 | Merge branch 'next' into codex/gsd-onboard |