diff --git a/.changeset/tidy-birds-march.md b/.changeset/tidy-birds-march.md new file mode 100644 index 000000000..9c809103e --- /dev/null +++ b/.changeset/tidy-birds-march.md @@ -0,0 +1,5 @@ +--- +type: Fixed +pr: 3390 +--- +**Interactive runs no longer stop for a checkpoint after every tracer task** — under the `end-of-phase` default a tracer whose `` is automated-only is re-run and expansion continues with no `checkpoint:human-verify`; `mid-flight`, tracers carrying ``, and any tracer carrying `gate="blocking-human"` still stop for a human, and a failing tracer still halts. (#3299) diff --git a/CONTEXT.md b/CONTEXT.md index 4e805c794..96442be8d 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -483,7 +483,7 @@ Phase 1 deliverable under `--mvp` on a new project — the Phase-1 whole-applica Single-feature task that moves one user capability from open-to-close (happy path) end-to-end. Contrast with the horizontal layer (all models, then all APIs, then all UI). The default planning unit under tracer-first decomposition (the leading task is a Tracer Bullet); SPIDR Splitting axes (Spike, Paths, Interfaces, Data, Rules) are the canonical decomposition tools when a slice is too large for one phase. ### Tracer Bullet -The default GSD decomposition lead: a **permanent, production-quality, minimal end-to-end slice** that wires one path through every layer a phase touches and becomes part of the skeleton of the final system — written for keeps, not thrown away. Contrast with a **prototype** (throwaway reconnaissance code, deleted once its lesson is learned): a tracer's *functionality* gaps are acceptable but its *architectural* gaps are not; stubs are allowed only where they can later be filled without an architectural change. GSD ships tracers, never prototypes — which is why `gsd-planner` LEADS every plan with a `type="tracer"` task (default; `--no-tracer` / `TRACER_MODE=false` opts back into horizontal layers) and `gsd-executor` runs an early integration feedback gate on the tracer's `` before expansion tasks (autonomous: halt-on-fail; interactive: `checkpoint:human-verify`). Origin: *The Pragmatic Programmer* "Tracer Bullets" (#1945); the Walking Skeleton is the Phase-1 whole-application special case. See Vertical Slice, MVP Mode, Walking Skeleton, Precondition. +The default GSD decomposition lead: a **permanent, production-quality, minimal end-to-end slice** that wires one path through every layer a phase touches and becomes part of the skeleton of the final system — written for keeps, not thrown away. Contrast with a **prototype** (throwaway reconnaissance code, deleted once its lesson is learned): a tracer's *functionality* gaps are acceptable but its *architectural* gaps are not; stubs are allowed only where they can later be filled without an architectural change. GSD ships tracers, never prototypes — which is why `gsd-planner` LEADS every plan with a `type="tracer"` task (default; `--no-tracer` / `TRACER_MODE=false` opts back into horizontal layers) and `gsd-executor` runs an early integration feedback gate on the tracer's `` before expansion tasks (autonomous: halt-on-fail; interactive: honors `workflow.human_verify_mode` — under `end-of-phase` an automated-only `` continues with no checkpoint, otherwise `checkpoint:human-verify`, #3299). Origin: *The Pragmatic Programmer* "Tracer Bullets" (#1945); the Walking Skeleton is the Phase-1 whole-application special case. See Vertical Slice, MVP Mode, Walking Skeleton, Precondition. ### Precondition The front-of-task side of the GSD plan contract (issue #1949, *The Pragmatic Programmer* Topic 23 — Design by Contract). An optional `` element on `` stating, in a single line of runnable/checkable prose, what must already be true for the task to begin safely — e.g. "`OPENAI_API_KEY` is set", "`dist/schema.json` from Phase 02 exists", "server responds to GET /health". `gsd-executor` evaluates it before any other task work, using **read-only checks only** (file existence, env var presence with no value output, idempotent `GET /health`-style pings — no writes, no network POSTs, no secret emission; if a side-effecting check seems required, the executor halts and surfaces a checkpoint rather than running it): met-or-absent is a no-op (back-compat for every existing plan); unmet returns a `checkpoint:human-verify` with no partial commit, and is NEVER auto-approved under `AUTO_CFG=true` (a missing prerequisite is a fact the executor cannot establish on its own, not a verification step). `gsd-planner` emits `` in exactly three cases — `user_setup` consumption, prior-phase artifact dependency, or env-var/runtime-config dependency — when the assumption is not already guaranteed by `depends_on` ordering. Closes the contract triad whose other two sides are postconditions (``/``/``) and invariants (`must_haves.truths`). The structural validator (`cmdVerifyPlanStructure`) does not reject unknown optional tags, so adding `` passes plan-structure validation unchanged. Canonical schema reference: `docs/reference/plan-md.md` → Preconditions; emission rules + anti-patterns: `gsd-core/references/planner-preconditions.md`. The architectural-end companion is the Tracer Bullet (#1945); together they close both ends of the "outrunning your headlights" failure mode. See Tracer Bullet. diff --git a/agents/gsd-executor.md b/agents/gsd-executor.md index b8c047cad..8cc5870d4 100644 --- a/agents/gsd-executor.md +++ b/agents/gsd-executor.md @@ -158,11 +158,12 @@ For each task: - Commit (see task_commit_protocol) - Track completion + commit hash for Summary -2. **If `type="tracer"`:** (the leading thin end-to-end slice — production-quality, never a throwaway) - - Execute and commit exactly like `type="auto"` (real implementation, real ``, atomic commit). - - **Then run the tracer feedback gate BEFORE any expansion task** — an early integration checkpoint on the proven slice: - - **Autonomous run (auto mode active — `AUTO_CHAIN` or `AUTO_CFG` is `"true"`, per ``):** re-run the tracer's `` end-to-end. If it **fails**, HALT and surface it (deviation Rule 1) — do NOT proceed to expansion tasks. Pouring more layers onto a broken foundation is exactly the failure this gate prevents. If it passes, log `⚡ Tracer verified end-to-end — expanding` and continue. - - **Interactive run (auto mode not active):** immediately after committing the tracer, STOP and return a `checkpoint:human-verify` for the tracer's `` (the working slice) using checkpoint_return_format, before any expansion task. +2. **If `type="tracer"`:** (production-quality, never a throwaway) + - Execute and commit exactly like `type="auto"`. + - **Then run the tracer feedback gate BEFORE any expansion task** — an early integration checkpoint on the proven slice. In order (full chain: "Tracer feedback gate", checkpoints.md): + - **`gate="blocking-human"` → STOP**, return a `checkpoint:human-verify`. Every mode, auto included (golden rule 6). + - **Auto mode active** (`AUTO_CHAIN`/`AUTO_CFG` is `"true"`, per ``): re-run `` end-to-end. Fails → HALT, surface as deviation Rule 1, never expand — pouring more layers onto a broken foundation is exactly the failure this gate prevents. Passes → log `⚡ Tracer verified end-to-end — expanding`, continue. + - **Interactive:** per `HUMAN_VERIFY_MODE` — `end-of-phase` (default) + automated-only `` → re-run; fails → HALT as above, passes → continue, no checkpoint; else STOP → `checkpoint:human-verify` (#3299). 3. **If `type="checkpoint:*"`:** - STOP immediately — return structured checkpoint message @@ -306,6 +307,7 @@ Check if auto mode is active at executor start (chain flag or user preference): ```bash AUTO_CHAIN=$(gsd_run query config-get workflow._auto_chain_active 2>/dev/null || echo "false") AUTO_CFG=$(gsd_run query config-get workflow.auto_advance 2>/dev/null || echo "false") +HUMAN_VERIFY_MODE=$(gsd_run query config-get workflow.human_verify_mode --default end-of-phase --raw 2>/dev/null || echo "end-of-phase") ``` Auto mode is active if either `AUTO_CHAIN` or `AUTO_CFG` is `"true"`. Store the result for checkpoint handling below. @@ -322,7 +324,7 @@ For full automation-first patterns, server lifecycle, CLI handling: **Quick reference:** Users NEVER run CLI commands. Users ONLY visit URLs, click UI, evaluate visuals, provide secrets. Claude does all automation. -**Tracer feedback gate:** a `type="tracer"` task is followed by an early integration checkpoint on the proven slice (see `` → `execute_tasks`) — in autonomous runs a failing tracer `` HALTS before any expansion task; in interactive runs the executor emits a `checkpoint:human-verify` for the tracer immediately after committing it. +**Tracer feedback gate:** synthesized after a `type="tracer"` task; `gate="blocking-human"` STOPs in every mode. Branch in `` → `execute_tasks`; full chain in checkpoints.md (#3299). --- diff --git a/agents/gsd-planner.md b/agents/gsd-planner.md index 538f05366..599493fbb 100644 --- a/agents/gsd-planner.md +++ b/agents/gsd-planner.md @@ -265,7 +265,7 @@ Exceptions where `tdd="true"` is not needed: `type="checkpoint:*"` tasks, config End-to-end "[capability]" — one path only [one file per layer the phase touches] Wire ONE entry point through every layer to the far end of the stack. No other call sites, no batching. Real error handling on the single path. - [a real, runnable END-TO-END check of the one path — not a per-layer unit test] + [a real END-TO-END check of the one path — not a per-layer unit test] The single happy path works end-to-end and is committed. ``` diff --git a/docs/AGENTS.md b/docs/AGENTS.md index fb4eaa4c7..baea495e3 100644 --- a/docs/AGENTS.md +++ b/docs/AGENTS.md @@ -221,7 +221,7 @@ GSD uses a multi-agent architecture where thin orchestrators (workflow files) sp - Follows XML task instructions precisely - Atomic git commit per completed task - Handles task types: auto, tracer, checkpoint (human-verify, decision, human-action) -- Tracer feedback gate: after a `tracer` slice, verifies it end-to-end before expansion tasks — autonomous runs halt on failure; interactive runs emit a human-verify checkpoint +- Tracer feedback gate: after a `tracer` slice, verifies it end-to-end before expansion tasks — autonomous runs halt on failure; interactive runs honor `workflow.human_verify_mode` (under the `end-of-phase` default an automated-only `` continues with no checkpoint; otherwise a human-verify checkpoint is emitted, #3299) - Reports deviations from plan in SUMMARY.md - Invokes node repair on verification failure diff --git a/docs/reference/plan-md.md b/docs/reference/plan-md.md index 862be6c75..63c21b45b 100644 --- a/docs/reference/plan-md.md +++ b/docs/reference/plan-md.md @@ -248,7 +248,7 @@ Full taxonomy, emission rules, and anti-patterns (chiefly: rating everything `on | Type | Use | Autonomy | |---|---|---| | `auto` | Everything the executor can do independently. | Fully autonomous. | -| `tracer` | The leading thin end-to-end slice a plan starts with by default (tracer-first) — production-quality, wired through every layer, with a real end-to-end ``. | Fully autonomous; after committing, the executor runs the tracer's `` as an early integration gate — autonomous runs halt on failure before expansion, interactive runs present a `checkpoint:human-verify`. | +| `tracer` | The leading thin end-to-end slice a plan starts with by default (tracer-first) — production-quality, wired through every layer, with a real end-to-end ``. | Fully autonomous; after committing, the executor runs the tracer's `` as an early integration gate. A tracer carrying `gate="blocking-human"` STOPs for a human in every mode, auto included. Otherwise autonomous runs halt on failure before expansion, and interactive runs honor `workflow.human_verify_mode` (#3299): under the `end-of-phase` default a `` carrying only `` is re-run and, on success, expansion continues with **no** checkpoint (failure still halts); under `mid-flight`, or when the tracer carries ``, a `checkpoint:human-verify` is presented. Full precedence chain: `gsd-core/references/checkpoints.md` → "Tracer feedback gate". | | `checkpoint:human-verify` | Visual or functional verification that requires a human to look at a running UI or service. | Pauses execution; presents to the developer; resumes on approval. | | `checkpoint:decision` | Implementation choices that arose during execution and require the developer's input. | Pauses execution; presents options; resumes on selection. | | `checkpoint:human-action` | Truly unavoidable manual steps (account creation, hardware interaction). Used sparingly. | Pauses execution; resumes on confirmation. | @@ -284,7 +284,7 @@ Plans that contain any checkpoint task must set `autonomous: false` in frontmatt | `` | Every file the task creates or modifies. The executor writes only these files. | | `` | Files the executor must read before touching anything — the file being modified, any source-of-truth pattern file, any file whose types or conventions must be replicated. | | `` | Concrete instructions with exact identifiers, file paths, function signatures, and expected values. Never says "align X with Y" without specifying the target state. Never contains fenced code blocks or full implementations. | -| `` | A runnable command or check that proves the task succeeded. Must distinguish pass from fail — `echo "done"` is not valid. | +| `` | A runnable command or check that proves the task succeeded. Must distinguish pass from fail — `echo "done"` is not valid. Accepts either the wrapped form (`cmd`) or the legacy bare-text form (`cmd`); both are valid. **On a `type="tracer"` task, prefer the wrapped form:** the tracer feedback gate's auto-continue (#3299) requires a `` carrying only ``, so a bare-text tracer verify falls through to the STOP fallback and still presents a `checkpoint:human-verify` in interactive runs even under `end-of-phase`. | | `` | Verifiable conditions: grep-verifiable strings, command exit codes, observable behaviours. No subjective language ("looks correct", "properly configured"). Negative greps (`! grep -Eq 'PAT' file`) are file-scoped — region-scope them (`sed -n`/`awk` range, then grep) when a sibling task needs the construct elsewhere in the same file (#968). | | `` | A short measurable statement of the completed outcome. | diff --git a/gsd-core/references/checkpoints.md b/gsd-core/references/checkpoints.md index 925c79bc7..25d616ec8 100644 --- a/gsd-core/references/checkpoints.md +++ b/gsd-core/references/checkpoints.md @@ -30,7 +30,7 @@ The gate spans two layers, and both must honor it. `gsd-executor` refuses to aut **When:** Claude completed automated work, human confirms it works correctly. -> **Default mode (#3309): `workflow.human_verify_mode = end-of-phase`.** New projects do NOT halt mid-flight at `checkpoint:human-verify`. The planner suppresses those task emissions and embeds the verification details into the relevant `auto` task's `` block; the verifier harvests every `` at end-of-phase (Step 8) and consolidates them into the existing `human_needed` → `{phase_num}-UAT.md` flow in `workflows/execute-phase.md`. The user reviews everything in one batch. +> **Default mode (#3309): `workflow.human_verify_mode = end-of-phase`.** New projects do NOT halt mid-flight at *planner-emitted* `checkpoint:human-verify` tasks. (The executor-synthesized **tracer feedback gate** is the one runtime checkpoint this mode also governs, with its own precedence chain — see "Tracer feedback gate (#3299)" below.) The planner suppresses those task emissions and embeds the verification details into the relevant `auto` task's `` block; the verifier harvests every `` at end-of-phase (Step 8) and consolidates them into the existing `human_needed` → `{phase_num}-UAT.md` flow in `workflows/execute-phase.md`. The user reviews everything in one batch. > > **Why this is the default:** every mid-flight halt costs a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) because subagent context is discarded across the pause. A plan with N human-verify checkpoints pays the cold-start cost N+1 times — measured at "tens of thousands of tokens" per round-trip on real projects. > @@ -107,6 +107,30 @@ The gate spans two layers, and both must honor it. `gsd-executor` refuses to aut Type "approved" or describe issues ``` + +### Tracer feedback gate (#3299) + +A `type="tracer"` task is followed by an early integration checkpoint on the proven slice, run BEFORE any expansion task. This checkpoint is **synthesized by the executor at runtime** — no planner emits it — so planner-side `human_verify_mode` suppression cannot reach it. It must therefore consult the mode itself. + +Evaluate the rows **in order** and take the first that matches — they are a precedence chain, not independent conditions: + +| # | Run | Tracer `` | Behavior | +|---|---|---|---| +| 1 | **Any run, any mode** (incl. auto) | task carries `gate="blocking-human"` | **STOP → `checkpoint:human-verify`.** Never auto-continued. | +| 2 | Auto mode active (`AUTO_CHAIN`/`AUTO_CFG`) | any (row 1 already took `blocking-human`) | Re-run verify; HALT on failure, continue on success. **Pre-existing behavior — unchanged by #3299.** | +| 3 | Interactive, `end-of-phase` (default) | only `` | Re-run verify; HALT on failure, continue to expansion on success — **no checkpoint** | +| 4 | Interactive, `end-of-phase` | carries `` | STOP → `checkpoint:human-verify` | +| 5 | Interactive, `mid-flight` | any | STOP → `checkpoint:human-verify` | + +**Carve-outs — the #3299 auto-continue (row 3) applies ONLY when all three hold:** the run is interactive, the mode is `end-of-phase`, and the tracer's `` contains only ``. Anything else STOPs or falls to the pre-existing auto-mode branch. HALT-on-failure is unconditional in rows 2 and 3 alike: a failing tracer never becomes an approvable checkpoint and never proceeds to expansion, because layering expansion onto a broken slice is the failure this gate exists to prevent. + +Row 1 is deliberately **not** scoped to interactive runs. Golden rule 6 above states that `gate="blocking-human"` stops for a human in *every* mode including auto-mode, and a precedence chain that let an autonomous run continue past it would make this file assert two incompatible rules about the same gate. No planner emits `gate` on a `type="tracer"` task today, but `src/verify.cts` parses only `type` and does not consult `gate` on non-checkpoint tasks, so a hand-authored, imported, or externally-generated `PLAN.md` can carry it and validate — unreachable by our planner is not unreachable. + +Read `HUMAN_VERIFY_MODE` with an explicit default — `workflow.human_verify_mode` is absent from `SCHEMA_DEFAULTS`, so a bare `config-get` exits non-zero with `Key not found` on any project whose `config.json` predates #3309: + +```bash +HUMAN_VERIFY_MODE=$(gsd_run query config-get workflow.human_verify_mode --default end-of-phase --raw 2>/dev/null || echo "end-of-phase") +``` diff --git a/gsd-core/references/planner-human-verify-mode.md b/gsd-core/references/planner-human-verify-mode.md index c8ffb923f..b9a56b17c 100644 --- a/gsd-core/references/planner-human-verify-mode.md +++ b/gsd-core/references/planner-human-verify-mode.md @@ -50,8 +50,22 @@ Choose `mid-flight` when you genuinely need the work to stop before any subseque `checkpoint:decision` and `checkpoint:human-action` tasks are still emitted in `end-of-phase` mode. Those gate the work itself (a choice the executor needs from the user, or an auth step only the user can perform), not post-hoc verification of completed work. Only `checkpoint:human-verify` is suppressed. +## The tracer feedback gate (executor-side, #3299) + +This mode is not purely a planner concern. The **tracer feedback gate** — the executor's early integration checkpoint after a `type="tracer"` task, in `workflows/execute-plan.md` and `agents/gsd-executor.md` — synthesizes a `checkpoint:human-verify` at runtime that no planner ever emitted, so planner-side suppression cannot reach it. + +That gate predates this mode (added by #2294; `end-of-phase` became the default in #3309, whose scope was the planner and verifier only), and until #3299 it branched on auto-mode alone. The result was that under the documented default, an interactive run halted after **every** tracer whose evidence was purely a test verdict, asking the user to retype a result the executor had just computed. + +The gate now honors `human_verify_mode`. + +The full precedence chain lives in `gsd-core/references/checkpoints.md` → "Tracer feedback gate (#3299)"; it is evaluated in order, and `gate="blocking-human"` outranks everything. Summary: an interactive `end-of-phase` run with an automated-only `` re-runs it and continues with no checkpoint (HALT on failure, unconditionally); `mid-flight`, ``, and `blocking-human` all still STOP; the auto-mode branch is unchanged. + +**Why a tracer carrying `` still halts rather than deferring to the end-of-phase UAT batch.** Deferring would be the more uniform reading of this mode — `` on an `auto` task defers, so arguably it should defer on a tracer too. It deliberately does not, for three reasons. First, and decisively: **the end-of-phase harvest does not cover tracers.** `agents/gsd-verifier.md` collects `` blocks from `auto` tasks; deferring a tracer's human evidence without first widening that seam would drop the evidence on the floor entirely — strictly worse than halting. Second, the tracer gate exists to stop expansion being layered onto an unproven slice; deferring its human evidence would let every expansion task build on a slice no human has confirmed, the exact failure the gate was introduced to prevent. Third, the reported defect is scoped to tracers with *no* human-observable evidence, and fail-closed is the safe direction outside that scope. If uniformity is later preferred, the harvest must be widened to tracers in the same change — record that decision here rather than re-deriving it. + +`workflow.human_verify_mode` is **absent from `SCHEMA_DEFAULTS`** in `src/config.cts`, so `query config-get workflow.human_verify_mode` exits non-zero with `Key not found` on any project whose `config.json` predates #3309 — it does not resolve the documented `end-of-phase` default. Every consumer must therefore pass `--default end-of-phase` explicitly. + ## Compatibility with other modes - **`workflow.tdd_mode`**: orthogonal. TDD tasks still emit `tdd="true"` and ``; the `` block carries the human-check sub-element when `human_verify_mode = end-of-phase`. - **`MVP_MODE`**: orthogonal. Vertical-slice ordering is unchanged. The first task remains a failing end-to-end test; later auto tasks may carry `` instead of standalone checkpoint tasks. -- **`workflow.auto_advance` / `_auto_chain_active`**: in mid-flight mode these auto-approve checkpoint:human-verify halts. In end-of-phase mode there are no halts to auto-approve, so the flags have no effect on this code path. +- **`workflow.auto_advance` / `_auto_chain_active`**: in mid-flight mode these auto-approve checkpoint:human-verify halts. In end-of-phase mode there are no *planner-emitted* halts to auto-approve, so the flags have no effect on the planner's output. They are not inert at execution time, though: the executor-side tracer feedback gate above synthesizes its own checkpoint, and the auto-mode branch takes precedence over `human_verify_mode` there — except for `gate="blocking-human"`, which is evaluated first and STOPs in every mode (#3299). diff --git a/gsd-core/workflows/execute-plan.md b/gsd-core/workflows/execute-plan.md index 3281aeaf7..62eb3789b 100644 --- a/gsd-core/workflows/execute-plan.md +++ b/gsd-core/workflows/execute-plan.md @@ -94,6 +94,7 @@ TASK_COUNT=$(grep -cE '^\s*]' .planning/phases/XX-name/{phase}-{ INLINE_THRESHOLD=$(gsd_run query config-get workflow.inline_plan_threshold 2>/dev/null || echo "2") USE_WORKTREES=$(gsd_run query config-get workflow.use_worktrees --raw 2>/dev/null || echo "true") RUNTIME=$(gsd_run query config-get runtime --default claude --raw 2>/dev/null || echo "claude") +HUMAN_VERIFY_MODE=$(gsd_run query config-get workflow.human_verify_mode --default end-of-phase --raw 2>/dev/null || echo "end-of-phase") grep -n "type=\"checkpoint" .planning/phases/XX-name/{phase}-{plan}-PLAN.md ``` @@ -219,7 +220,7 @@ Deviations are normal — handle via rules below. 3. Per task: - **MANDATORY read_first gate:** If the task has a `` field, you MUST read every listed file BEFORE making any edits. This is not optional. Do not skip files because you "already know" what's in them — read them. The read_first files establish ground truth for the task. - `type="auto"`: if `tdd="true"` → TDD execution. Implement with deviation rules + auth gates. Verify done criteria. Commit (see task_commit). Track hash for Summary. - - `type="tracer"`: execute like `type="auto"` (production-quality, real ``, commit), then run the tracer feedback gate BEFORE any expansion task — an early integration checkpoint. Auto mode active (`AUTO_CHAIN` or `AUTO_CFG`): re-run the tracer ``; on failure HALT and surface (deviation) — do NOT start expansion tasks. Interactive: STOP → return a `checkpoint:human-verify` for the tracer via checkpoint_protocol before expansion. + - `type="tracer"`: execute like `type="auto"` (production-quality, real ``, commit), then run the tracer feedback gate BEFORE any expansion task — an early integration checkpoint. Evaluate in order (#3299). First, `gate="blocking-human"` → STOP → return a `checkpoint:human-verify` via checkpoint_protocol — every mode, auto included (golden rule 6, checkpoints.md). Next, Auto mode active (`AUTO_CHAIN` or `AUTO_CFG`): re-run the tracer ``; on failure HALT and surface (deviation) — do NOT start expansion tasks. Next, `HUMAN_VERIFY_MODE` is `end-of-phase` (default) AND the tracer's `` carries only `` (no ``) → re-run the tracer ``; on failure HALT and surface as a deviation exactly as in the auto-mode branch — never a checkpoint; on success log `⚡ Tracer verified end-to-end — expanding` and continue to expansion, do NOT synthesize a checkpoint. Otherwise (`mid-flight`, or the tracer carries genuine human-observable evidence) → STOP → return a `checkpoint:human-verify` for the tracer via checkpoint_protocol before expansion. - `type="checkpoint:*"`: STOP → checkpoint_protocol → wait for user → continue only after confirmation. - **HARD GATE — acceptance_criteria verification:** After completing each task, if it has ``, you MUST run a verification loop before proceeding: 1. For each criterion: execute the grep, file check, or CLI command that proves it passes diff --git a/tests/emitted-drift-acks/2775-planner-package-legitimacy-gate.json b/tests/emitted-drift-acks/2775-planner-package-legitimacy-gate.json index cfdd7d7a7..e530aa1ab 100644 --- a/tests/emitted-drift-acks/2775-planner-package-legitimacy-gate.json +++ b/tests/emitted-drift-acks/2775-planner-package-legitimacy-gate.json @@ -2,7 +2,7 @@ "version": 1, "paths": { "gsd-planner.md": { - "reason": "#2775: the STRIDE supply-chain row for npm/pip/cargo installs was rewritten from 'slopcheck + blocking human checkpoint for [ASSUMED]/[SUS]' to 'package-legitimacy gate + blocking human checkpoint for [ASSUMED]/[SUS]', matching ADR-0656 (registry-API verdicts are the gate; slopcheck is an optional escalate-only adapter no shipped configuration wires). The +14 bytes is the corrected mitigation description agents read at plan time, not incidental prose growth. #3565 appends the fenced `## Return Markers` section (~1.2 KB) enumerating the six exact stall-watch dispatch markers plan-phase.md matches on — the new registry lint requires every declared marker to be emitted in-fence by its producer, and the planner previously documented none of them anywhere. #2401 append (merged into this fragment because two ack sources may never name the same path): the 'Inherit the command that already worked' paragraph is now a short pointer to the extracted gsd-core/references/planner-verify-command-grounding.md (prior_verify_commands inheritance, npm --prefix grounding rules, and the guessing prohibition all moved to the reference). The file is now 49274 bytes, well under the XL tier's 57344-byte cap, and 49077 chars (LF-normalized) under the separate 49152-char cap asserted by tests/planner-decomposition.test.cjs, tests/precondition-element.test.cjs, tests/reversibility-tagging.test.cjs, and tests/security.test.cjs. #3003 append (merged into this fragment because two ack sources may never name the same path): the implicit-dependency wave rule now computes same-wave overlap over files_modified PLUS files_deleted, so a plan deleting a file another same-wave plan edits is pushed to a later wave instead of racing it -- one branch removing what the other is writing is the sharpest conflict there is, and reading files_modified alone scored that pair conflict-free. Deliberately terse (+56 LF chars, landing at 49133 against the 49152-char cap named above, 19 chars of headroom). The `# Implicit dependency: files_modified overlap forces a later wave.` comment is left VERBATIM: tests/parallel-dependent-plans.test.cjs matches that exact unbackticked substring, so rewording it to name the new channel reds the suite -- the files_deleted addition rides in the pseudocode and the Rule sentence instead: the rationale is carried in docs/reference/plan-md.md, which has no cap. The next contributor who needs room in this agent must extract, not add prose." + "reason": "#2775: the STRIDE supply-chain row for npm/pip/cargo installs was rewritten from 'slopcheck + blocking human checkpoint for [ASSUMED]/[SUS]' to 'package-legitimacy gate + blocking human checkpoint for [ASSUMED]/[SUS]', matching ADR-0656 (registry-API verdicts are the gate; slopcheck is an optional escalate-only adapter no shipped configuration wires). The +14 bytes is the corrected mitigation description agents read at plan time, not incidental prose growth. #3565 appends the fenced `## Return Markers` section (~1.2 KB) enumerating the six exact stall-watch dispatch markers plan-phase.md matches on — the new registry lint requires every declared marker to be emitted in-fence by its producer, and the planner previously documented none of them anywhere. #2401 append (merged into this fragment because two ack sources may never name the same path): the 'Inherit the command that already worked' paragraph is now a short pointer to the extracted gsd-core/references/planner-verify-command-grounding.md (prior_verify_commands inheritance, npm --prefix grounding rules, and the guessing prohibition all moved to the reference). The file was 49274 bytes (49077 chars) immediately before the #3003 and #3299 arms below, both of which supersede that figure, well under the XL tier's 57344-byte cap, and 49077 chars (LF-normalized) under the separate 49152-char cap asserted by tests/planner-decomposition.test.cjs, tests/precondition-element.test.cjs, tests/reversibility-tagging.test.cjs, and tests/security.test.cjs. #3003 append (merged into this fragment because two ack sources may never name the same path): the implicit-dependency wave rule now computes same-wave overlap over files_modified PLUS files_deleted, so a plan deleting a file another same-wave plan edits is pushed to a later wave instead of racing it -- one branch removing what the other is writing is the sharpest conflict there is, and reading files_modified alone scored that pair conflict-free. Deliberately terse (+56 LF chars, landing at 49133 against the 49152-char cap named above, 19 chars of headroom). The `# Implicit dependency: files_modified overlap forces a later wave.` comment is left VERBATIM: tests/parallel-dependent-plans.test.cjs matches that exact unbackticked substring, so rewording it to name the new channel reds the suite -- the files_deleted addition rides in the pseudocode and the Rule sentence instead: the rationale is carried in docs/reference/plan-md.md, which has no cap. The next contributor who needs room in this agent must extract, not add prose. — #3299 append: the tracer task-shape template's now wraps its command in . The Nyquist Rule earlier in the file already required every to carry , but this tracer-specific template still showed the legacy bare-text form, so a planner following its own most specific template emitted tracers the #3299 feedback gate can never auto-continue on (the gate requires a carrying only ) — making the fix unreachable on the default tracer-first path. Caught in peer review. Written on ONE line, matching the one-line example under docs/reference/plan-md.md's \"## Reversibility\" heading (cited by section, not line: an earlier \":207\" citation was accurate when written and drifted with a later merge of next). Sizes re-measured against this merge of next, which landed the #3003 arm above: the file is 49146 chars after this change (49343 bytes), +13 chars on next's 49133, leaving 6 chars under the 49152-CHAR cap asserted independently by tests/planner-decomposition.test.cjs, tests/precondition-element.test.cjs, tests/reversibility-tagging.test.cjs and tests/security.test.cjs (all four assert `< 49152`, so 49146 passes). The 49090 / 62-chars-of-headroom figures this arm carried in earlier rounds were measured before #3003 and are superseded. Note the cap counts CHARACTERS, not bytes, and `wc -m` silently reports bytes under LC_ALL=C, so measure in a UTF-8 locale. The three-line form crossed the cap; the placeholder lost the word 'runnable' to buy that room — 'a real END-TO-END check' still carries the operative requirement, and the neighbouring prose already says the tracer carries the same and validation as any auto task. With 6 chars left this file is effectively full: the next contributor needing room here must extract to a reference file, not bump the cap." } } } diff --git a/tests/emitted-drift-acks/3370-execute-phase-gate-conflation.json b/tests/emitted-drift-acks/3370-execute-phase-gate-conflation.json index 5cc939f98..7d404c976 100644 --- a/tests/emitted-drift-acks/3370-execute-phase-gate-conflation.json +++ b/tests/emitted-drift-acks/3370-execute-phase-gate-conflation.json @@ -1,6 +1,6 @@ { "version": 1, "paths": { - "execute-plan.md": "#3370: the Pattern A dispatch prompt spec gained the gate-semantics clause (gate=\"blocking\" (the default) is auto-approvable in auto-mode per the executor's own checkpoint protocol, gate=\"blocking-human\" always surfaces to a human; add no instruction overriding that protocol), closing the identically-shaped dispatch-time gap on the single-plan path named in the issue. Growth ~248 bytes (38913 -> 39161, still under the DEFAULT 40 KiB ceiling). Supersedes the spent #2652 fragment (merged into next), which also named execute-plan.md and would otherwise double-ack the same path. #3659 appends --mode \"$ISOLATION\" to the #2649 pre-dispatch base-check and corrects the baseRef restore-advice to the orchestrator/harness split (+154B; the harness does not read baseRef, #48). Deliberate growth." + "execute-plan.md": "#3370: the Pattern A dispatch prompt spec gained the gate-semantics clause (gate=\"blocking\" (the default) is auto-approvable in auto-mode per the executor's own checkpoint protocol, gate=\"blocking-human\" always surfaces to a human; add no instruction overriding that protocol), closing the identically-shaped dispatch-time gap on the single-plan path named in the issue. Growth ~248 bytes (38913 -> 39161, still under the DEFAULT 40 KiB ceiling). Supersedes the spent #2652 fragment (merged into next), which also named execute-plan.md and would otherwise double-ack the same path. #3659 appends --mode \"$ISOLATION\" to the #2649 pre-dispatch base-check and corrects the baseRef restore-advice to the orchestrator/harness split (+154B; the harness does not read baseRef, #48). Deliberate growth. — #3299 append: the tracer feedback gate's interactive branch, which read `Interactive: STOP -> return a checkpoint:human-verify` and keyed on auto-mode alone, now branches on HUMAN_VERIFY_MODE — under the documented `end-of-phase` default an automated-only tracer `` is re-run and continues to expansion with no checkpoint. Growth is 796 bytes (39315 -> 40111, still under the DEFAULT_CAP of 40960 LF bytes, 849 bytes of headroom; the earlier 39161 -> 39957 / 1003 figures were measured before the 2026-08-22 merge of next and are superseded): that branch plus a `HUMAN_VERIFY_MODE=` read alongside the existing RUNTIME/USE_WORKTREES reads in parse_segments. The read carries an explicit `--default end-of-phase` because `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS (src/config.cts): a bare config-get exits non-zero with `Key not found` on any project whose config.json predates #3309, which is every pre-existing project and the reporter's exact config. The three carve-outs (``, `gate=\"blocking-human\"`, `mid-flight`) are stated inline here rather than delegated to gsd-core/references/dispatch-isolation-gate.md, because this file is the inline dispatch path used for step-by-step / non-Claude-Code execution, where agents/gsd-executor.md is never loaded. This entry carries #3370's own reason forward verbatim above rather than replacing it — that growth is in the base and still needs its account. The blocking-human STOP is evaluated BEFORE the auto-mode branch here, matching golden rule 6 in checkpoints.md — the earlier ordering let an autonomous run continue past a blocking-human tracer." } } diff --git a/tests/no-bare-gsd-tools-command-position.test.cjs b/tests/no-bare-gsd-tools-command-position.test.cjs index ee60a43b1..a46930819 100644 --- a/tests/no-bare-gsd-tools-command-position.test.cjs +++ b/tests/no-bare-gsd-tools-command-position.test.cjs @@ -82,11 +82,11 @@ const BARE_COMMAND_RE = new RegExp( // Each entry MUST carry a one-line reason; the test prints the allowlist on // failure so a reviewer can see exactly what is sanctioned. const PROSE_ALLOWLIST = [ - { file: 'agents/gsd-executor.md', line: 793, reason: 'describes the SDK return envelope of `gsd-tools query commit`; not an instruction to run the bare word' }, + { file: 'agents/gsd-executor.md', line: 795, reason: 'describes the SDK return envelope of `gsd-tools query commit`; not an instruction to run the bare word' }, { file: 'agents/gsd-phase-researcher.md', line: 33, reason: 'package-legitimacy provenance rule names the command as the source of an OK verdict; descriptive' }, { file: 'agents/gsd-roadmapper.md', line: 642, reason: 'parenthetical "e.g." naming SDK queries a user *could* run; not an agent instruction' }, { file: 'agents/gsd-intel-updater.md', line: 40, reason: 'cross-platform note names the `gsd-tools intel ` CLI surface descriptively ("CLI invocations go through..."); not an agent instruction' }, - { file: 'gsd-core/workflows/execute-plan.md', line: 414, reason: 'describes the downstream SDK validation step (`validated downstream by ...`); names the mechanism, does not instruct the agent to type it' }, + { file: 'gsd-core/workflows/execute-plan.md', line: 415, reason: 'describes the downstream SDK validation step (`validated downstream by ...`); names the mechanism, does not instruct the agent to type it' }, { file: 'agents/gsd-research-synthesizer.md', line: 65, reason: 'a code comment inside a fenced block explaining what the commit step loads (`# Planning config loaded via gsd-tools query ...`); descriptive, not an invocation — and explicitly names gsd-tools.cjs as the alternative' }, ]; diff --git a/tests/tracer-bullet.test.cjs b/tests/tracer-bullet.test.cjs index 30d974e04..fa5bbf2b0 100644 --- a/tests/tracer-bullet.test.cjs +++ b/tests/tracer-bullet.test.cjs @@ -91,15 +91,47 @@ function parseExecutorContract(md) { // Autonomous: halt-on-fail before any expansion task. // Keyed on the file's own auto-mode definition (AUTO_CHAIN or AUTO_CFG), // not AUTO_CFG alone — see . + // #3299 B3: `gate="blocking-human"` is evaluated BEFORE the auto-mode + // branch and binds in every mode (golden rule 6). Ordering is the whole + // point — a chain that reached the auto branch first would auto-continue + // past the repo's strongest gate in exactly the unattended mode where it + // matters most, which is the defect this ordering fixes. Asserted by + // POSITION, not presence: both clauses existing in the wrong order passes + // any presence check and still ships the bypass. + blockingHumanPrecedesAutoMode: (() => { + const iGate = md.indexOf('`gate="blocking-human"` \u2192 STOP'); + const iAuto = md.indexOf('**Auto mode active**'); + return iGate > -1 && iAuto > iGate && /Every mode, auto included/i.test(md); + })(), autoHaltsOnFailure: - /Autonomous run \(auto mode active/i.test(md) && - /`AUTO_CHAIN` or `AUTO_CFG`/.test(md) && - /HALT and surface it/i.test(md) && - /do NOT proceed to expansion tasks/i.test(md), - // Interactive: emit checkpoint:human-verify immediately after the tracer. + /\*\*Auto mode active\*\* \(`AUTO_CHAIN`\/`AUTO_CFG`/.test(md) && + /HALT/.test(md) && + /never expand/i.test(md), + // Interactive: the branch exists and still names checkpoint:human-verify. interactiveHumanVerify: - /Interactive run \(auto mode not active\)/i.test(md) && + /\*\*Interactive:\*\*/.test(md) && /checkpoint:human-verify/.test(md), + // #3299: that checkpoint is now the FALLBACK, not the unconditional result. + // Merely finding HUMAN_VERIFY_MODE on the line proves nothing — peer review + // showed a branch can name the variable and still checkpoint unconditionally. + // Require the ordered clause markers AND that the auto-continue clause is + // free of any STOP outcome, which is what "conditional" actually means here. + interactiveIsConditional: (() => { + const m = md.match(/\*\*Interactive:\*\*([^\n]*)/); + if (!m) return false; + const body = m[1]; + // Both outcomes must be present and separated, and the auto-continue + // half must contain no STOP — naming HUMAN_VERIFY_MODE proves nothing on + // its own, as a branch can cite the variable and still checkpoint + // unconditionally (that exact shape passed an earlier revision). + const iElse = body.indexOf('else'); + if (iElse < 0) return false; + const autoContinue = body.slice(0, iElse); + return /HUMAN_VERIFY_MODE/.test(body) + && /no checkpoint/i.test(autoContinue) + && !/\bSTOP\b/.test(autoContinue) + && /STOP \u2192 `checkpoint:human-verify`/.test(body.slice(iElse)); + })(), // Cross-referenced in the checkpoint protocol section too. documentedInCheckpointProtocol: /\*\*Tracer feedback gate:\*\*/.test(md), }; @@ -194,9 +226,36 @@ describe('#1945 executor: post-tracer feedback gate', () => { assert.ok(c.autoHaltsOnFailure, 'autonomous run must halt (surfaced) before expansion when the tracer verify fails'); }); - // Acceptance: interactive run presents a human-verify checkpoint after the tracer. - test('interactive run emits checkpoint:human-verify after the tracer (acceptance #4)', () => { - assert.ok(c.interactiveHumanVerify, 'interactive run must emit checkpoint:human-verify immediately after the tracer'); + // Golden rule 6 (checkpoints.md): `gate="blocking-human"` stops for a human in + // EVERY mode, auto included. The precedence chain is first-match, so this is an + // ordering property, not a presence one — an auto-mode branch evaluated first + // silently swallows a blocking-human tracer in exactly the unattended run where + // the gate matters most, with both clauses still present in the file. + test('gate="blocking-human" is evaluated BEFORE the auto-mode branch (golden rule 6)', () => { + assert.ok( + c.blockingHumanPrecedesAutoMode, + 'the blocking-human STOP must appear before the auto-mode branch and state that it binds in every mode', + ); + }); + + // Acceptance #4, as narrowed by #3299. Originally "an interactive run ALWAYS + // emits checkpoint:human-verify after the tracer". That is no longer the + // contract: under human_verify_mode=end-of-phase an automated-only tracer + // verify auto-continues with no checkpoint. What survives of #1945's + // acceptance is that the interactive branch still HAS a checkpoint outcome — + // it is now the fallback rather than the unconditional result. + // + // Left as a bare `interactiveHumanVerify` substring check this test kept + // passing after #3299 purely because the strings still appear in the fallback + // clause, while its NAME asserted the opposite of shipped behavior — the same + // one-copy-stale drift #3299 itself is about. Suite 6 owns the conditional + // contract; this one is scoped to what #1945 still guarantees. + test('interactive run retains a checkpoint:human-verify outcome (acceptance #4, narrowed by #3299)', () => { + assert.ok(c.interactiveHumanVerify, 'interactive branch must still exist and still name checkpoint:human-verify'); + assert.ok( + c.interactiveIsConditional, + 'post-#3299 the interactive checkpoint is CONDITIONAL — the branch must consult HUMAN_VERIFY_MODE, not emit unconditionally', + ); }); test('gate is cross-referenced in the checkpoint protocol', () => { @@ -263,10 +322,6 @@ describe('#1945 glossary + docs', () => { assert.match(COMMANDS_DOC, /\| `--no-tracer` \|/, 'COMMANDS.md flag table must include --no-tracer'); }); - test('docs/reference/plan-md.md task-types table includes tracer', () => { - assert.match(PLAN_MD_REF, /\| `tracer` \|/, 'plan-md.md Task types table must include a tracer row'); - }); - test('docs/how-to and docs/AGENTS reflect tracer-first + the executor gate', () => { assert.match(HOWTO, /tracer/i, 'how-to must mention tracer-first'); assert.match(HOWTO, /--no-tracer/, 'how-to must mention the --no-tracer opt-out'); @@ -368,3 +423,468 @@ describe('#1945 behavioral: verify plan-structure accepts type="tracer" (accepta ); }); }); + +// ─── Suite 6: #3299 — the tracer gate must honor workflow.human_verify_mode ─── + +// The tracer feedback gate (#2294) predates human_verify_mode (#3309), whose scope +// was the planner + verifier only. Until #3299 the gate branched on auto-mode ALONE, +// so under the documented `end-of-phase` default an interactive run halted after +// EVERY tracer — synthesizing a checkpoint:human-verify no planner ever emitted and +// asking the user to retype a verdict the executor had just computed. +// +// These assertions are prose-shaped because the gate itself is prose: it is executed +// by an agent reading agents/gsd-executor.md and gsd-core/workflows/execute-plan.md. +// The two files duplicate the rule and MUST stay in sync — a fix landing in only one +// leaves the defect live on whichever dispatch path reads the other. +describe('#3299 regression: tracer feedback gate honors workflow.human_verify_mode', () => { + const HV_REF = read('gsd-core/references/planner-human-verify-mode.md'); + const CHECKPOINTS = read('gsd-core/references/checkpoints.md'); + const { parsePredicates } = require('../gsd-core/bin/lib/context-predicates.cjs'); + + // ── Why the selection layer looks like this ──────────────────────────────── + // + // These files are prose an agent executes, so a regression test must prove the + // OPERATIVE text is right — not that correct-looking text exists somewhere in + // the file. Six rounds of adversarial review defeated weaker shapes, each + // leaving the suite green while the reported bug shipped: + // + // 1. keyword presence -> reverting the rule entirely passed + // 2. presence + names -> `blocking-human AND mid-flight` passed + // 3. blacklisting `STOP` -> "pause and invoke checkpoint_protocol" passed + // 4. exact-pin one clause -> an override sentence ABOVE the clause passed + // 5. + hand-rolled `` (which comments the real rule through EOF), + // and a normal fenced documentation example of the rule combined with a + // whitespace-only reformat of the live list item. + // + // Rather than hand-roll a third scanner, defer to the repo's own interleaved + // fence/comment scanner via the PUBLIC `parsePredicates` export: instrument + // candidate lines as throwaway predicate declarations and let it tell us which + // ones are operative. Verified: fenced, balanced-commented, and + // after-unclosed-comment candidates are all correctly excluded. + function operativeLineIndexes(md, candidateRe) { + // Guard the CLASS, not just the one caller that got it wrong (review round + // 9). This helper REPLACES the candidate line. Deleting an ordinary content + // line is harmless, but deleting a fence DELIMITER leaves its partner behind + // to become an opener, inverting fence parity for the entire remainder of the + // document — `computeSkippedLineFlags` is a strict FORWARD state machine, so + // every marker after the deletion then lands alternately inside and outside a + // phantom fence. Ask about a fence delimiter's position with + // `isOperativePosition` instead, which INSERTS and therefore perturbs nothing. + // + // Checked against the lines this regex ACTUALLY matches in THIS document, not + // against a sample of delimiter spellings. A fixed probe list was the first + // attempt and it is not the class: `~~~xml`, ```` ```json ````, longer tilde + // runs and info strings all walk straight past any list short enough to + // write down. Matching on the real data cannot go stale, and cannot pass a + // delimiter it has not thought of. (Codex review, round 9.) + const lines = md.split(/\r?\n/); + for (const [i, line] of lines.entries()) { + if (candidateRe.test(line) && /^\s*(?:`{3,}|~{3,})/.test(line)) { + throw new Error( + `operativeLineIndexes: the candidate regex ${candidateRe} matches the fence delimiter ` + + `${JSON.stringify(line)} at 0-based line ${i}. Replacing a delimiter inverts fence parity ` + + 'for the rest of the document and misclassifies LIVE fences downstream. ' + + 'Use isOperativePosition(md, idx).'); + } + } + const injected = new Set(); + const instrumented = lines + .map((line, i) => { + if (!candidateRe.test(line)) return line; + // A ONE-LINE `` carries both delimiters, so the scanner's + // multi-line comment tracking never opens for it — and replacing the + // line with a predicate marker STRIPS the delimiters, promoting the + // commented text to operative. Re-test with complete same-line spans + // removed: if the candidate only matched inside one, it is not live. + if (!candidateRe.test(line.replace(//g, ''))) return line; + // Same trap, unclosed form: `/g, '').replace(/` + // wrapper comments the candidate's CONTENT while leaving the position outside + // every span, so the probe alone would answer "live" for ``. + // Reject that here rather than relying on each caller's own shape test to + // happen to exclude it. + // + // The rule is the SCANNER's, not a tighter one of our own: it skips an entire + // line whose TRIMMED text starts with ` real content`, which the + // scanner skips outright. Agreeing with it beats out-reasoning it. + // (Codex review, round 9.) + if (lines[idx].trimStart().startsWith('/g, '').replace(/', + '', + '', // 10 + '', // 11 same-line-commented opener + ].join('\n'); + const FENCE = /^\s*```xml(?:\s.*)?$/; + + assert.throws(() => operativeLineIndexes(md, FENCE), /matches the fence delimiter/, + 'a candidate regex matching a fence delimiter must be refused outright, not answered wrongly — ' + + 'this helper REPLACES the candidate, so removing one delimiter inverts parity downstream'); + + // The spellings that defeated the first attempt at this guard, which probed a + // fixed list of delimiter strings. Each is a real fence opener and none of + // them appears in any list short enough to write down — which is why the + // guard now matches the document's own lines instead. (Codex review, round 9.) + for (const [label, doc, re] of [ + ['~~~xml', 'p\n~~~xml\n\n~~~\n', /^\s*~~~xml$/], + ['```json', 'p\n```json\n{}\n```\n', /^\s*```json$/], + ['~~~~ (4 tildes)', 'p\n~~~~\nx\n~~~~\n', /^\s*~~~~$/], + ['``` with info string', 'p\n```xml title=a\n\n```\n', /^\s*```xml\s.*$/], + ]) { + assert.throws(() => operativeLineIndexes(doc, re), /matches the fence delimiter/, + `${label}: the guard must cover the delimiter CLASS, not a sample of its spellings — a regex ` + + 'it waves through reintroduces the exact parity inversion this row exists for'); + } + + // Codex review, round 9. `isOperativePosition` must agree with the scanner, + // which skips an ENTIRE line whose trimmed text starts with ` - `X.B=2`\n', 1), false, + 'a line that STARTS with a comment opener is skipped wholesale by the scanner, balanced or not — ' + + 'this probe must not claim it is live'); + + assert.deepStrictEqual( + [1, 5, 9, 11].map((i) => isOperativePosition(md, i)), + [true, true, false, false], + 'both live openers must read operative (the second is the one the deletion route lost), and ' + + 'neither the block-commented nor the same-line-commented opener may'); + + // Non-vacuity: the probe must not simply answer "true" for every position it + // is handed, and must stay correct as content shifts above it. + assert.strictEqual(isOperativePosition(md, 2), false, 'a line INSIDE a fence is not operative'); + assert.strictEqual(isOperativePosition(md, 0), true, 'plain prose is operative'); + const shifted = ['extra', '', ...md.split('\n')].join('\n'); + assert.deepStrictEqual([3, 7].map((i) => isOperativePosition(shifted, i)), [true, true], + 'both openers must still read operative after unrelated lines are inserted above them — the ' + + 'defect this replaces was a parity coincidence that a shift like this flipped'); + }); + + const OPERATIVE = [ + ['agents/gsd-executor.md', () => regionFrom(EXECUTOR, + soleOperativeIndex('agents/gsd-executor.md', EXECUTOR, EXEC_ANCHOR, 'tracer task branch'), + /^\s*3\.\s+\*\*If\s/), "2. **If `type=\"tracer\"`:** (production-quality, never a throwaway) - Execute and commit exactly like `type=\"auto\"`. - **Then run the tracer feedback gate BEFORE any expansion task** \u2014 an early integration checkpoint on the proven slice. In order (full chain: \"Tracer feedback gate\", checkpoints.md): - **`gate=\"blocking-human\"` \u2192 STOP**, return a `checkpoint:human-verify`. Every mode, auto included (golden rule 6). - **Auto mode active** (`AUTO_CHAIN`/`AUTO_CFG` is `\"true\"`, per ``): re-run `` end-to-end. Fails \u2192 HALT, surface as deviation Rule 1, never expand \u2014 pouring more layers onto a broken foundation is exactly the failure this gate prevents. Passes \u2192 log `\u26a1 Tracer verified end-to-end \u2014 expanding`, continue. - **Interactive:** per `HUMAN_VERIFY_MODE` \u2014 `end-of-phase` (default) + automated-only `` \u2192 re-run; fails \u2192 HALT as above, passes \u2192 continue, no checkpoint; else STOP \u2192 `checkpoint:human-verify` (#3299)."], + ['gsd-core/workflows/execute-plan.md', () => { + const i = soleOperativeIndex('gsd-core/workflows/execute-plan.md', EXECUTE_PLAN, EP_ANCHOR, 'tracer dispatch line'); + return EXECUTE_PLAN.split(/\r?\n/)[i].replace(/\s+/g, ' ').trim(); + }, "- `type=\"tracer\"`: execute like `type=\"auto\"` (production-quality, real ``, commit), then run the tracer feedback gate BEFORE any expansion task \u2014 an early integration checkpoint. Evaluate in order (#3299). First, `gate=\"blocking-human\"` \u2192 STOP \u2192 return a `checkpoint:human-verify` via checkpoint_protocol \u2014 every mode, auto included (golden rule 6, checkpoints.md). Next, Auto mode active (`AUTO_CHAIN` or `AUTO_CFG`): re-run the tracer ``; on failure HALT and surface (deviation) \u2014 do NOT start expansion tasks. Next, `HUMAN_VERIFY_MODE` is `end-of-phase` (default) AND the tracer's `` carries only `` (no ``) \u2192 re-run the tracer ``; on failure HALT and surface as a deviation exactly as in the auto-mode branch \u2014 never a checkpoint; on success log `\u26a1 Tracer verified end-to-end \u2014 expanding` and continue to expansion, do NOT synthesize a checkpoint. Otherwise (`mid-flight`, or the tracer carries genuine human-observable evidence) \u2192 STOP \u2192 return a `checkpoint:human-verify` for the tracer via checkpoint_protocol before expansion."], + ]; + + test('the complete operative gate region is pinned in both copies', () => { + for (const [name, extract, expected] of OPERATIVE) { + assert.strictEqual(extract(), expected, + `${name}: the tracer gate's decision region drifted from its pinned contract.\n\n` + + `Pinned WHOLE on purpose: pinning only the auto-continue clause let an unconditional override ` + + `sentence be added beside it with every assertion still passing. If the behavior genuinely ` + + `changed, update the prose AND this expected string together; do not narrow the assertion.`); + } + }); + + test('the canonical checkpoints.md tracer section is pinned whole', () => { + const i = soleOperativeIndex('checkpoints.md', CHECKPOINTS, CK_ANCHOR, 'tracer-gate heading'); + assert.strictEqual(regionFrom(CHECKPOINTS, i, /^\s*###\s|^\s*` | Behavior | |---|---|---|---| | 1 | **Any run, any mode** (incl. auto) | task carries `gate=\"blocking-human\"` | **STOP \u2192 `checkpoint:human-verify`.** Never auto-continued. | | 2 | Auto mode active (`AUTO_CHAIN`/`AUTO_CFG`) | any (row 1 already took `blocking-human`) | Re-run verify; HALT on failure, continue on success. **Pre-existing behavior \u2014 unchanged by #3299.** | | 3 | Interactive, `end-of-phase` (default) | only `` | Re-run verify; HALT on failure, continue to expansion on success \u2014 **no checkpoint** | | 4 | Interactive, `end-of-phase` | carries `` | STOP \u2192 `checkpoint:human-verify` | | 5 | Interactive, `mid-flight` | any | STOP \u2192 `checkpoint:human-verify` | **Carve-outs \u2014 the #3299 auto-continue (row 3) applies ONLY when all three hold:** the run is interactive, the mode is `end-of-phase`, and the tracer's `` contains only ``. Anything else STOPs or falls to the pre-existing auto-mode branch. HALT-on-failure is unconditional in rows 2 and 3 alike: a failing tracer never becomes an approvable checkpoint and never proceeds to expansion, because layering expansion onto a broken slice is the failure this gate exists to prevent. Row 1 is deliberately **not** scoped to interactive runs. Golden rule 6 above states that `gate=\"blocking-human\"` stops for a human in *every* mode including auto-mode, and a precedence chain that let an autonomous run continue past it would make this file assert two incompatible rules about the same gate. No planner emits `gate` on a `type=\"tracer\"` task today, but `src/verify.cts` parses only `type` and does not consult `gate` on non-checkpoint tasks, so a hand-authored, imported, or externally-generated `PLAN.md` can carry it and validate \u2014 unreachable by our planner is not unreachable. Read `HUMAN_VERIFY_MODE` with an explicit default \u2014 `workflow.human_verify_mode` is absent from `SCHEMA_DEFAULTS`, so a bare `config-get` exits non-zero with `Key not found` on any project whose `config.json` predates #3309: ```bash HUMAN_VERIFY_MODE=$(gsd_run query config-get workflow.human_verify_mode --default end-of-phase --raw 2>/dev/null || echo \"end-of-phase\") ``` ", + 'checkpoints.md tracer-gate section drifted. Pinned whole so behavior-bearing prose cannot be ' + + 'added around the table (an "ignore row 3, always wait" line below it previously passed).'); + }); + + // Asserting only that the row EXISTS is what let the canonical schema table + // drift out of sync with shipped behavior after #3299 without CI noticing — + // CONTEXT.md names this file the canonical reference for the task-type + // contract, so a wrong row here is the authoritative wrong answer. Keyword + // presence is not enough either: peer review defeated an earlier revision by + // APPENDING "Nevertheless, interactive runs always present a + // checkpoint:human-verify." — every required keyword still matched, so the + // reference could contradict itself with CI green. Hence the EXACT pin, which + // is deliberately brittle: a wording change must be a conscious edit in both + // places. Selection routes through soleOperativeIndex rather than a raw + // startsWith find — an earlier duplicate of this test used the raw form, which + // this suite records at :477 as a defeated round-1 shape, and it was removed in + // review of #3390 rather than left as a second hand-maintained copy of the + // same canonical string. + test('docs/reference/plan-md.md tracer row Autonomy cell matches shipped behavior exactly', () => { + const i = soleOperativeIndex('plan-md.md', PLAN_MD_REF, ROW_ANCHOR, 'tracer table row'); + const cells = PLAN_MD_REF.split(/\r?\n/)[i].trim().replace(/^\|/, '').replace(/\|$/, '').split('|').map((c) => c.trim()); + assert.strictEqual(cells.length, 3, `tracer row must have 3 cells, got ${cells.length}`); + assert.strictEqual(cells[2].replace(/\s+/g, ' ').trim(), "Fully autonomous; after committing, the executor runs the tracer's `` as an early integration gate. A tracer carrying `gate=\"blocking-human\"` STOPs for a human in every mode, auto included. Otherwise autonomous runs halt on failure before expansion, and interactive runs honor `workflow.human_verify_mode` (#3299): under the `end-of-phase` default a `` carrying only `` is re-run and, on success, expansion continues with **no** checkpoint (failure still halts); under `mid-flight`, or when the tracer carries ``, a `checkpoint:human-verify` is presented. Full precedence chain: `gsd-core/references/checkpoints.md` \u2192 \"Tracer feedback gate\".", + 'plan-md.md tracer Autonomy cell drifted. CONTEXT.md names this table the canonical schema ' + + 'reference — update the cell AND this expected string together.'); + }); + + // Structural, not copy-pinned: the contract is the SHAPE of the verify, so a + // wording improvement to the placeholder must not false-fail (round-5 Minor). + test('planner tracer template emits exactly one -wrapped verify', () => { + const i = soleOperativeIndex('agents/gsd-planner.md', PLANNER, PLANNER_ANCHOR, 'tracer task shape marker'); + const lines = PLANNER.split(/\r?\n/); + // The fence OPENER must be operative AND the first non-blank line after the + // marker. Matching the first raw ```xml in the remainder let a commented-out + // decoy template be selected while the live one regressed (review round 8) — + // this was the one selection in the suite that was not fence/comment aware. + let openIdx = -1; + for (let j = i + 1; j < lines.length; j++) { + if (lines[j].trim() === '') continue; + openIdx = j; + break; + } + assert.notStrictEqual(openIdx, -1, 'the tracer task shape marker must be followed by content'); + // Shape off the RAW line, liveness off the position. The round-8 fix asked + // `operativeLineSet` — which deletes the opener it is asking about, inverting + // fence parity downstream: on the head planner it reported the live "Task-level + // TDD" fence at 0-based 233 as NON-operative, and the assertion only passed + // because 0-based 262 happened to land in a surviving parity slot. One extra + // live ```xml example anywhere earlier in the file flipped it to a false + // FAILURE blaming a decoy that does not exist (review round 9). + assert.ok( + /^\s*```xml(?:\s.*)?$/.test(lines[openIdx]) && isOperativePosition(PLANNER, openIdx), + 'the first non-blank line after the tracer task shape marker must be a LIVE ```xml fence opener — ' + + 'a commented-out or non-adjacent decoy template must not be selectable', + ); + const after = lines.slice(openIdx).join('\n'); + const fence = after.match(/```xml[^\r\n]*\r?\n([\s\S]*?)```/); + assert.ok(fence, 'the tracer task shape must be followed by a fenced xml block'); + const verifies = fence[1].match(/[\s\S]*?<\/verify>/g) || []; + assert.strictEqual(verifies.length, 1, `the tracer template must contain exactly ONE , found ${verifies.length}`); + const inner = verifies[0].replace(/^/, '').replace(/<\/verify>$/, '').replace(/\s+/g, ' ').trim(); + assert.match(inner, /^[^<>]+<\/automated>$/, + "the tracer template's body must be exactly one non-empty child — the #3299 " + + 'gate auto-continues only on an automated-only verify, so a bare-text template makes the fix ' + + 'unreachable for every tracer the planner generates'); + }); + + test('every site reading the mode passes an explicit --default end-of-phase', () => { + // Requires the ASSIGNMENT, not merely the command. Matching the config-get + // substring alone let `IGNORED_MODE=$(gsd_run query config-get ...)` keep this + // row green while nothing defines HUMAN_VERIFY_MODE — the gate then falls + // through to STOP and #3299 is back with the regression suite still passing. + // A test that survives the regression it exists to catch is not a test. + // The lookahead after `end-of-phase` closes the other half: the bare prefix + // also accepted `--default end-of-phase-wrong`. (Codex review, round 9.) + const READ = /^\s*HUMAN_VERIFY_MODE=\$\(gsd_run query config-get workflow\.human_verify_mode --default end-of-phase(?=\s|$)/; + // NOT operativeLineIndexes here, deliberately. All three reads live inside a + // ```bash fence, which is their correct executable form in these files, and + // that selector excludes fenced lines by design — using it would assert the + // opposite of the shipped shape. What the original bare whole-file + // assert.match genuinely could not catch is a read present ONLY inside an + // HTML comment, or a second drifted copy alongside the live one. Pin both: + // exactly one occurrence, inside a live fence, outside any comment. + const BASH_FENCES = new Set(['bash', 'sh', 'shell', 'zsh']); + const liveFencedReads = (md) => { + let fenceLang = null, inComment = false, hits = 0; + for (const raw of md.split(/\r?\n/)) { + let line = raw; + if (inComment) { + const end = line.indexOf('-->'); + if (end === -1) continue; + line = line.slice(end + 3); + inComment = false; + } + // Strip COMPLETE spans first: a one-line comment carries both + // delimiters, so an open/close test that only looks for an unpaired `/g, ''); + const open = line.indexOf('