From 622f43353c07d4a0e45191418e750934a1d5b7b6 Mon Sep 17 00:00:00 2001 From: Behruz Nassre Esfahani <20915308+behruznassre@users.noreply.github.com> Date: Sun, 23 Aug 2026 15:43:53 -0700 Subject: [PATCH] fix(#3299): tracer feedback gate honors workflow.human_verify_mode (#3390) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * fix(#3299): tracer feedback gate honors workflow.human_verify_mode The tracer feedback gate (#2294) predates `workflow.human_verify_mode` (#3309, whose scope was the planner and verifier only), and branched on auto-mode alone. Under the documented `end-of-phase` default an interactive run therefore halted after EVERY `type="tracer"` task, synthesizing a `checkpoint:human-verify` no planner ever emitted and asking the user to retype a verdict the executor had just computed — at the cost of a full executor cold-start each time. Planner-side suppression cannot reach this halt because the executor synthesizes it at runtime, which is why #3309 did not close it. The gate now branches on HUMAN_VERIFY_MODE in the interactive path: under `end-of-phase` an automated-only tracer `` is re-run and, on success, expansion continues with no checkpoint. HALT-on-failure is unchanged. `mid-flight`, `gate="blocking-human"`, and tracers carrying genuine `` evidence all still stop; the autonomous branch is untouched. `--default end-of-phase` on the config read is load-bearing, not decorative: `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS, so a bare `config-get` exits non-zero with `Key not found` on any project whose config.json predates #3309 — which is the reporter's exact config and every pre-existing project. Both copies of the rule (workflows/execute-plan.md and agents/gsd-executor.md) are updated together; the reference doc records the seam and the human-check-still-halts rationale so it cannot recur. Fixes #3299 Co-Authored-By: Claude Opus 5 (1M context) * chore(#3299): add changeset Co-Authored-By: Claude Opus 5 (1M context) * fix(#3299): reconcile the canonical schema table and the stale acceptance test Review round 1 (trek-e) — three items, all in the drift class this PR is about, two of them landed inside this PR's own diff. 1. docs/reference/plan-md.md:233 — CONTEXT.md names this file the canonical schema reference for the tracer task-type contract, and its Task-types row still claimed interactive runs unconditionally present a checkpoint:human-verify. CONTEXT.md and docs/AGENTS.md were updated in the first round; this one was missed, so the authoritative reference was the wrong answer. The row now carries the human_verify_mode-conditional behavior and points at the canonical precedence chain. 2. tests/tracer-bullet.test.cjs — the docs assertion only checked that a tracer ROW EXISTS, never its content, which is why CI could not see the drift. It now asserts the row's actual claims and rejects the pre-#3299 wording. Separately, the #1945 acceptance test named 'interactive run emits checkpoint:human-verify after the tracer' kept passing only because its substrings still occur in the fallback clause, while its name asserted the opposite of shipped behavior. Renamed and narrowed to what #1945 still guarantees, plus a new interactiveIsConditional pin so the unconditional prose cannot be restored under a passing substring check. 3. plan-md.md's row now documents that the legacy bare-text form (valid, and still shown at :179) does not reach the #3299 auto-continue — only a carrying does — so the benefit is silently unreachable for tracers using that format. Mutation-verified: reverting the plan-md row fails 1 test; reverting the executor's interactive branch fails 4. Co-Authored-By: Claude Opus 5 (1M context) * fix(#3299): make the tracer gate reachable from the planner template, and bind the assertions Peer review round 3 found two Majors, both verified by reproducing the mutation before fixing. MAJOR 1 — the fix was largely inert on its own default path. agents/gsd-planner.md's Nyquist Rule (:191) says every includes , but the tracer-specific template twelve lines later emitted the legacy bare-text form. The gate auto-continues only on a carrying only , so every tracer produced from the canonical template fell to the STOP fallback and #3299's benefit was unreachable for exactly the task type it targets. Template now wraps in ; a contract assertion pins it so the two cannot drift apart again. MAJOR 2 — the new assertions did not bind condition to action. Appending 'Nevertheless, interactive runs always present a checkpoint:human-verify' to the canonical row, and 'then immediately STOP and return a checkpoint:human-verify' to the auto-continue clause in BOTH operative copies, restored unconditional interactive checkpointing and left the suite 35/35 green. Every required keyword still matched. Fixed by: - clause 2 must now contain no STOP outcome and emit no checkpoint at all — 'never a checkpoint' has to be true OF the clause, not merely stated in it; - interactiveIsConditional replaced with the ordered-clause parse plus the same no-STOP property, instead of proving only that HUMAN_VERIFY_MODE appears somewhere on the line; - the plan-md.md Autonomy cell is now pinned EXACTLY rather than by keyword presence. Deliberately brittle: CONTEXT.md names that table the canonical schema reference, so a wording change must be a conscious edit in both places. Mutation-verified after the fix: the combined semantic regression now fails 3 tests; reverting the planner template fails 1. Co-Authored-By: Claude Opus 5 (1M context) * test(#3299): exact-pin the safety clauses instead of blacklisting outcome verbs Peer review round 4. Blacklisting did not hold, twice over: - Round 3 banned literal STOP and the 'return a'/'present a' checkpoint forms in the auto-continue clause. Round 4 defeated that by appending 'then pause and invoke checkpoint_protocol with a checkpoint:human-verify before expansion' — none of the banned tokens, same restored interruption after every successful tracer. 36/36 passed. - The planner guard looked for anywhere inside , so '[...]' satisfied it while leaving the legacy bare form operative. 107/107 passed across tracer, planner and the three size-cap suites. Synonyms are unbounded; the clauses are not. Both are now pinned exactly on normalized whitespace, the same approach already proven on the plan-md.md Autonomy cell, with defence-in-depth checks behind them: no checkpoint-emitting or blocking outcome in any wording inside clause 2, and the planner's body must be exactly one non-empty child with no commented markup. These pins are deliberately brittle. Each is a safety contract, so changing the behavior must be a conscious edit in both the prose and the expectation. Mutation-verified: the synonym-checkpoint mutation fails 1; the commented-out wrapper fails 1; the round-3 literal-STOP + contradictory-doc-row regression fails 3. Co-Authored-By: Claude Opus 5 (1M context) * test(#3299): strip comments, require uniqueness, pin whole regions Peer review round 5. Exact-pinning one clause was still bypassable two ways, both reproduced before fixing (each left the suite fully green): - COMMENTED DECOYS. Put the correct text in an HTML comment followed by a live wrong copy: every extractor selected the commented decoy. Worked against the planner template, the canonical plan-md.md row, and both executor branches. - SURROUNDING OVERRIDE. Insert 'after every tracer, pause and invoke checkpoint_protocol before expansion, regardless of the mode-specific rules below' immediately ABOVE the pinned clause, or 'ignore row 3; always wait for approval' below the canonical table. The pinned text was untouched, so equality held while the shipped meaning inverted. The shape that holds, applied to every operative surface: 1. strip HTML comments BEFORE selecting, so a decoy cannot be chosen; 2. require the structural anchor to occur EXACTLY ONCE, so a live second copy cannot hide behind a correct first one; 3. pin the ENTIRE decision region, not one clause, so no unparsed prefix or suffix can override what the pin proves. Applied to: the executor's whole tracer branch, execute-plan.md's whole dispatch line, checkpoints.md's whole precedence section, and plan-md.md's Autonomy cell. Also addresses the round-5 Minor: the planner template is now asserted STRUCTURALLY (exactly one in the fenced block, body exactly one non-empty child) rather than pinning the descriptive placeholder verbatim, so behavior-preserving wording changes no longer false-fail. The clause and section pins keep their exact form — those have a safety rationale the placeholder copy does not. Mutation-verified, all six rounds: override-above-clause 1; commented decoy row 1; commented decoy branch 1; ignore-row-3 override 1; synonym checkpoint 1; commented-out wrapper 2. Co-Authored-By: Claude Opus 5 (1M context) * test(#3299): drop the superseded exact-placeholder planner assertion Peer review round 6, Minor. The round-5 brittleness fix ADDED a structural planner assertion but left the old exact-placeholder one in place, so the over-brittleness it was meant to remove was still live: rewording the descriptive placeholder while preserving exactly one non-empty direct child failed the old test and passed the new one. Removed the old test. The structural assertion is the real contract — the gate auto-continues on the SHAPE of the verify, not on the wording of a placeholder. Verified both directions: a behavior-preserving reword now passes; reverting the template to bare still fails. Co-Authored-By: Claude Opus 5 (1M context) * test(#3299): select operative prose via parsePredicates, not a hand-rolled scanner Peer review round 7. I had judged the round-6 selector bypass adversarial-only and out of scope, intending to disclose it. Both premises were wrong, and the review said so: - 'Needs new src API' — false. parsePredicates is ALREADY a public export and internally uses the repo's interleaved fence/comment scanner. Instrumenting candidate lines as throwaway predicate declarations borrows that scanner with no src change at all. - 'Adversarial-only' — false, and this is the part that mattered. Two ORDINARY edits silently turned the guards into decoy checks: * a forgotten '-->' comments the live rule through to EOF, and the balanced-only stripper still saw and accepted the commented rule; * a normal fenced documentation example of the rule, plus a whitespace-only reformat of the live list item, made the selector choose the example. Neither needs intent. A dangling comment is a typo; a fenced example is good documentation. Together they reproduce exactly the accidental drift #3299 came from — with CI green. The selection layer now defers to parsePredicates for operativeness, uses whitespace-tolerant anchors so a reformat cannot decouple the live line from its pin, extracts regions by operative line index rather than string search, and carries a self-guard test proving fenced / balanced-commented / after-unclosed-comment copies are all excluded. The helper also ignores indexes it did not inject, so a pre-existing GSDTEST.CANDIDATE line cannot pollute it. Verified both ordinary-edit scenarios now fail the suite (each was green before). Co-Authored-By: Claude Opus 5 (1M context) * test(#3299): close the operative-selection gaps the maintainer blocked on trek-e's Blocker: the operative-line selection layer had three gaps, all reachable by ordinary future doc edits rather than sabotage. He independently found a fourth I had not disclosed. All are fixed. 1. INDENTATION PROMOTION (his find, not in my disclosure). The instrumentation replaced a matched candidate with an UNINDENTED marker regardless of the original line's indentation. A 4-space-indented CommonMark code block is not skipped by parsePredicates (it accepts indented declarations by design), so stripping the indent PROMOTED an indented decoy to operative — the exact inversion of the guard's purpose. The marker now preserves the original indent, and a candidate that is itself indented 4+ spaces is never injected. 2. NO SET MEMBERSHIP. The filter accepted any in-range integer, so a pre-existing literal GSDTEST.CANDIDATE= in source text could pollute the count. Now filters on a Set of the indexes actually injected on this call. 3. RAW FENCE SELECTION (planner). The template test matched the first raw ```xml fence after the marker with no fence/comment awareness — the one selection in the suite that was not operative-aware — so a commented-out decoy template between the marker and the real one would be selected while the live template regressed. The opener must now be operative AND the first non-blank line after the marker. 4. RAW END ANCHOR (regionFrom). The end anchor was tested against raw lines, so a fenced example containing a ### / * test(#3299): allow-list operative indentation; pin marker provenance Review round 9. BLOCKER — the round-8 indentation guard was written as a DENY-list, /^(?: {4,}|\t)/, and CommonMark has more indented-code forms than that enumerates: " \t", " \t" and " \t" all open an indented code block and all slipped through, so an indented decoy was still promoted to operative while the live rule regressed (34/34 green). Inverted to an allow-list — only 0-3 literal spaces is ordinary block indentation; anything else is code. Enumerating the bad shapes was the error, not the specific regex. MINOR — the injected-index Set validated the marker's VALUE but not its SOURCE. A pre-existing literal `GSDTEST.CANDIDATE=` could name an index that some other (skipped) candidate had contributed to the set, and be accepted. Now also requires p.line - 1 === Number(p.value): the predicate must have been parsed from the line it names. MINOR (false negative) — ```xml title=x is a valid CommonMark info string, and requiring exactly ```xml failed the suite (33/34) on a behavior-preserving edit. Both the opener assertion and the extraction now accept an info string. Mutation-verified: the mixed " \t" decoy and the forged-provenance marker each now fail; the info-string fence no longer false-fails. KNOWN LIMITATION, disclosed on the PR rather than papered over: parsePredicates is a predicate parser, not a general CommonMark operativeness oracle. Two standards-valid constructs still read as operative — a lazy blockquote continuation line (state opens only on a line that literally starts with ">"), and a comment opened mid-line ("prose `), rather than leaving that to each caller's own shape test to happen to exclude. `operativeLineIndexes` now THROWS when its candidate regex matches a fence delimiter, so the unsound route cannot be reached again by a future caller rather than only being fixed at the one site that got it wrong. Verified with the same extra-example scenario above: with the fix, all 35 rows stay green. Teeth: reverting the call site to `operativeLineSet` turns the tracer-template row red on the new guard. The regression row pins both live openers (the second is the one the deletion route lost), the block-commented and same-line-commented openers, a line inside a fence, and re-checks both openers after unrelated lines shift above them. Only tests/tracer-bullet.test.cjs changes — no agent file is touched, so the 5-char gsd-planner.md and 19-byte gsd-executor.md headroom are unaffected. Verified: `npm run lint:ci` exit 0; full `npm test` 31307 tests / 31292 pass / 0 fail / 14 skipped, TMPDIR unset, against a freshly synced origin/next. Co-Authored-By: Claude Opus 5 (1M context) * fix(#3299): guard the delimiter class, match the scanner, pin the assignment Codex full-PR review of #3390, run against the round-9 head. Three defects, two of them in the code that round added. 1. The mode-read pin survived the regression it exists to catch. `READ` matched the config-get substring only, so rewriting the shipped line as `IGNORED_MODE=$(gsd_run query config-get ...)` kept the row green while nothing defined HUMAN_VERIFY_MODE — the gate falls through to STOP and #3299 is back with the suite passing. The regex now requires the assignment. A lookahead after `end-of-phase` closes the other half: the bare prefix also accepted `--default end-of-phase-wrong`. Proven by mutation: renaming the variable in agents/gsd-executor.md now turns that row red, and did not before. 2. The round-9 fence-delimiter guard was a SAMPLE of the class, not the class. It probed a fixed list of five delimiter strings. `~~~xml`, ```json, `~~~~` and arbitrary info strings all walk past any list short enough to write down — the guard was added precisely because one such regex had already slipped through. Now matched against the lines the regex actually selects in the document, which cannot go stale and cannot miss a spelling nobody thought of. Four such spellings pinned as rows. 3. `isOperativePosition` disagreed with the scanner it delegates to. For ` real content` it stripped the span, found surviving content, and answered "live". `computeSkippedLineFlags` skips an ENTIRE line whose trimmed text starts with `` (which comments the real rule through EOF), + // and a normal fenced documentation example of the rule combined with a + // whitespace-only reformat of the live list item. + // + // Rather than hand-roll a third scanner, defer to the repo's own interleaved + // fence/comment scanner via the PUBLIC `parsePredicates` export: instrument + // candidate lines as throwaway predicate declarations and let it tell us which + // ones are operative. Verified: fenced, balanced-commented, and + // after-unclosed-comment candidates are all correctly excluded. + function operativeLineIndexes(md, candidateRe) { + // Guard the CLASS, not just the one caller that got it wrong (review round + // 9). This helper REPLACES the candidate line. Deleting an ordinary content + // line is harmless, but deleting a fence DELIMITER leaves its partner behind + // to become an opener, inverting fence parity for the entire remainder of the + // document — `computeSkippedLineFlags` is a strict FORWARD state machine, so + // every marker after the deletion then lands alternately inside and outside a + // phantom fence. Ask about a fence delimiter's position with + // `isOperativePosition` instead, which INSERTS and therefore perturbs nothing. + // + // Checked against the lines this regex ACTUALLY matches in THIS document, not + // against a sample of delimiter spellings. A fixed probe list was the first + // attempt and it is not the class: `~~~xml`, ```` ```json ````, longer tilde + // runs and info strings all walk straight past any list short enough to + // write down. Matching on the real data cannot go stale, and cannot pass a + // delimiter it has not thought of. (Codex review, round 9.) + const lines = md.split(/\r?\n/); + for (const [i, line] of lines.entries()) { + if (candidateRe.test(line) && /^\s*(?:`{3,}|~{3,})/.test(line)) { + throw new Error( + `operativeLineIndexes: the candidate regex ${candidateRe} matches the fence delimiter ` + + `${JSON.stringify(line)} at 0-based line ${i}. Replacing a delimiter inverts fence parity ` + + 'for the rest of the document and misclassifies LIVE fences downstream. ' + + 'Use isOperativePosition(md, idx).'); + } + } + const injected = new Set(); + const instrumented = lines + .map((line, i) => { + if (!candidateRe.test(line)) return line; + // A ONE-LINE `` carries both delimiters, so the scanner's + // multi-line comment tracking never opens for it — and replacing the + // line with a predicate marker STRIPS the delimiters, promoting the + // commented text to operative. Re-test with complete same-line spans + // removed: if the candidate only matched inside one, it is not live. + if (!candidateRe.test(line.replace(//g, ''))) return line; + // Same trap, unclosed form: `/g, '').replace(/` + // wrapper comments the candidate's CONTENT while leaving the position outside + // every span, so the probe alone would answer "live" for ``. + // Reject that here rather than relying on each caller's own shape test to + // happen to exclude it. + // + // The rule is the SCANNER's, not a tighter one of our own: it skips an entire + // line whose TRIMMED text starts with ` real content`, which the + // scanner skips outright. Agreeing with it beats out-reasoning it. + // (Codex review, round 9.) + if (lines[idx].trimStart().startsWith('/g, '').replace(/', + '', + '', // 10 + '', // 11 same-line-commented opener + ].join('\n'); + const FENCE = /^\s*```xml(?:\s.*)?$/; + + assert.throws(() => operativeLineIndexes(md, FENCE), /matches the fence delimiter/, + 'a candidate regex matching a fence delimiter must be refused outright, not answered wrongly — ' + + 'this helper REPLACES the candidate, so removing one delimiter inverts parity downstream'); + + // The spellings that defeated the first attempt at this guard, which probed a + // fixed list of delimiter strings. Each is a real fence opener and none of + // them appears in any list short enough to write down — which is why the + // guard now matches the document's own lines instead. (Codex review, round 9.) + for (const [label, doc, re] of [ + ['~~~xml', 'p\n~~~xml\n\n~~~\n', /^\s*~~~xml$/], + ['```json', 'p\n```json\n{}\n```\n', /^\s*```json$/], + ['~~~~ (4 tildes)', 'p\n~~~~\nx\n~~~~\n', /^\s*~~~~$/], + ['``` with info string', 'p\n```xml title=a\n\n```\n', /^\s*```xml\s.*$/], + ]) { + assert.throws(() => operativeLineIndexes(doc, re), /matches the fence delimiter/, + `${label}: the guard must cover the delimiter CLASS, not a sample of its spellings — a regex ` + + 'it waves through reintroduces the exact parity inversion this row exists for'); + } + + // Codex review, round 9. `isOperativePosition` must agree with the scanner, + // which skips an ENTIRE line whose trimmed text starts with ` - `X.B=2`\n', 1), false, + 'a line that STARTS with a comment opener is skipped wholesale by the scanner, balanced or not — ' + + 'this probe must not claim it is live'); + + assert.deepStrictEqual( + [1, 5, 9, 11].map((i) => isOperativePosition(md, i)), + [true, true, false, false], + 'both live openers must read operative (the second is the one the deletion route lost), and ' + + 'neither the block-commented nor the same-line-commented opener may'); + + // Non-vacuity: the probe must not simply answer "true" for every position it + // is handed, and must stay correct as content shifts above it. + assert.strictEqual(isOperativePosition(md, 2), false, 'a line INSIDE a fence is not operative'); + assert.strictEqual(isOperativePosition(md, 0), true, 'plain prose is operative'); + const shifted = ['extra', '', ...md.split('\n')].join('\n'); + assert.deepStrictEqual([3, 7].map((i) => isOperativePosition(shifted, i)), [true, true], + 'both openers must still read operative after unrelated lines are inserted above them — the ' + + 'defect this replaces was a parity coincidence that a shift like this flipped'); + }); + + const OPERATIVE = [ + ['agents/gsd-executor.md', () => regionFrom(EXECUTOR, + soleOperativeIndex('agents/gsd-executor.md', EXECUTOR, EXEC_ANCHOR, 'tracer task branch'), + /^\s*3\.\s+\*\*If\s/), "2. **If `type=\"tracer\"`:** (production-quality, never a throwaway) - Execute and commit exactly like `type=\"auto\"`. - **Then run the tracer feedback gate BEFORE any expansion task** \u2014 an early integration checkpoint on the proven slice. In order (full chain: \"Tracer feedback gate\", checkpoints.md): - **`gate=\"blocking-human\"` \u2192 STOP**, return a `checkpoint:human-verify`. Every mode, auto included (golden rule 6). - **Auto mode active** (`AUTO_CHAIN`/`AUTO_CFG` is `\"true\"`, per ``): re-run `` end-to-end. Fails \u2192 HALT, surface as deviation Rule 1, never expand \u2014 pouring more layers onto a broken foundation is exactly the failure this gate prevents. Passes \u2192 log `\u26a1 Tracer verified end-to-end \u2014 expanding`, continue. - **Interactive:** per `HUMAN_VERIFY_MODE` \u2014 `end-of-phase` (default) + automated-only `` \u2192 re-run; fails \u2192 HALT as above, passes \u2192 continue, no checkpoint; else STOP \u2192 `checkpoint:human-verify` (#3299)."], + ['gsd-core/workflows/execute-plan.md', () => { + const i = soleOperativeIndex('gsd-core/workflows/execute-plan.md', EXECUTE_PLAN, EP_ANCHOR, 'tracer dispatch line'); + return EXECUTE_PLAN.split(/\r?\n/)[i].replace(/\s+/g, ' ').trim(); + }, "- `type=\"tracer\"`: execute like `type=\"auto\"` (production-quality, real ``, commit), then run the tracer feedback gate BEFORE any expansion task \u2014 an early integration checkpoint. Evaluate in order (#3299). First, `gate=\"blocking-human\"` \u2192 STOP \u2192 return a `checkpoint:human-verify` via checkpoint_protocol \u2014 every mode, auto included (golden rule 6, checkpoints.md). Next, Auto mode active (`AUTO_CHAIN` or `AUTO_CFG`): re-run the tracer ``; on failure HALT and surface (deviation) \u2014 do NOT start expansion tasks. Next, `HUMAN_VERIFY_MODE` is `end-of-phase` (default) AND the tracer's `` carries only `` (no ``) \u2192 re-run the tracer ``; on failure HALT and surface as a deviation exactly as in the auto-mode branch \u2014 never a checkpoint; on success log `\u26a1 Tracer verified end-to-end \u2014 expanding` and continue to expansion, do NOT synthesize a checkpoint. Otherwise (`mid-flight`, or the tracer carries genuine human-observable evidence) \u2192 STOP \u2192 return a `checkpoint:human-verify` for the tracer via checkpoint_protocol before expansion."], + ]; + + test('the complete operative gate region is pinned in both copies', () => { + for (const [name, extract, expected] of OPERATIVE) { + assert.strictEqual(extract(), expected, + `${name}: the tracer gate's decision region drifted from its pinned contract.\n\n` + + `Pinned WHOLE on purpose: pinning only the auto-continue clause let an unconditional override ` + + `sentence be added beside it with every assertion still passing. If the behavior genuinely ` + + `changed, update the prose AND this expected string together; do not narrow the assertion.`); + } + }); + + test('the canonical checkpoints.md tracer section is pinned whole', () => { + const i = soleOperativeIndex('checkpoints.md', CHECKPOINTS, CK_ANCHOR, 'tracer-gate heading'); + assert.strictEqual(regionFrom(CHECKPOINTS, i, /^\s*###\s|^\s*` | Behavior | |---|---|---|---| | 1 | **Any run, any mode** (incl. auto) | task carries `gate=\"blocking-human\"` | **STOP \u2192 `checkpoint:human-verify`.** Never auto-continued. | | 2 | Auto mode active (`AUTO_CHAIN`/`AUTO_CFG`) | any (row 1 already took `blocking-human`) | Re-run verify; HALT on failure, continue on success. **Pre-existing behavior \u2014 unchanged by #3299.** | | 3 | Interactive, `end-of-phase` (default) | only `` | Re-run verify; HALT on failure, continue to expansion on success \u2014 **no checkpoint** | | 4 | Interactive, `end-of-phase` | carries `` | STOP \u2192 `checkpoint:human-verify` | | 5 | Interactive, `mid-flight` | any | STOP \u2192 `checkpoint:human-verify` | **Carve-outs \u2014 the #3299 auto-continue (row 3) applies ONLY when all three hold:** the run is interactive, the mode is `end-of-phase`, and the tracer's `` contains only ``. Anything else STOPs or falls to the pre-existing auto-mode branch. HALT-on-failure is unconditional in rows 2 and 3 alike: a failing tracer never becomes an approvable checkpoint and never proceeds to expansion, because layering expansion onto a broken slice is the failure this gate exists to prevent. Row 1 is deliberately **not** scoped to interactive runs. Golden rule 6 above states that `gate=\"blocking-human\"` stops for a human in *every* mode including auto-mode, and a precedence chain that let an autonomous run continue past it would make this file assert two incompatible rules about the same gate. No planner emits `gate` on a `type=\"tracer\"` task today, but `src/verify.cts` parses only `type` and does not consult `gate` on non-checkpoint tasks, so a hand-authored, imported, or externally-generated `PLAN.md` can carry it and validate \u2014 unreachable by our planner is not unreachable. Read `HUMAN_VERIFY_MODE` with an explicit default \u2014 `workflow.human_verify_mode` is absent from `SCHEMA_DEFAULTS`, so a bare `config-get` exits non-zero with `Key not found` on any project whose `config.json` predates #3309: ```bash HUMAN_VERIFY_MODE=$(gsd_run query config-get workflow.human_verify_mode --default end-of-phase --raw 2>/dev/null || echo \"end-of-phase\") ``` ", + 'checkpoints.md tracer-gate section drifted. Pinned whole so behavior-bearing prose cannot be ' + + 'added around the table (an "ignore row 3, always wait" line below it previously passed).'); + }); + + // Asserting only that the row EXISTS is what let the canonical schema table + // drift out of sync with shipped behavior after #3299 without CI noticing — + // CONTEXT.md names this file the canonical reference for the task-type + // contract, so a wrong row here is the authoritative wrong answer. Keyword + // presence is not enough either: peer review defeated an earlier revision by + // APPENDING "Nevertheless, interactive runs always present a + // checkpoint:human-verify." — every required keyword still matched, so the + // reference could contradict itself with CI green. Hence the EXACT pin, which + // is deliberately brittle: a wording change must be a conscious edit in both + // places. Selection routes through soleOperativeIndex rather than a raw + // startsWith find — an earlier duplicate of this test used the raw form, which + // this suite records at :477 as a defeated round-1 shape, and it was removed in + // review of #3390 rather than left as a second hand-maintained copy of the + // same canonical string. + test('docs/reference/plan-md.md tracer row Autonomy cell matches shipped behavior exactly', () => { + const i = soleOperativeIndex('plan-md.md', PLAN_MD_REF, ROW_ANCHOR, 'tracer table row'); + const cells = PLAN_MD_REF.split(/\r?\n/)[i].trim().replace(/^\|/, '').replace(/\|$/, '').split('|').map((c) => c.trim()); + assert.strictEqual(cells.length, 3, `tracer row must have 3 cells, got ${cells.length}`); + assert.strictEqual(cells[2].replace(/\s+/g, ' ').trim(), "Fully autonomous; after committing, the executor runs the tracer's `` as an early integration gate. A tracer carrying `gate=\"blocking-human\"` STOPs for a human in every mode, auto included. Otherwise autonomous runs halt on failure before expansion, and interactive runs honor `workflow.human_verify_mode` (#3299): under the `end-of-phase` default a `` carrying only `` is re-run and, on success, expansion continues with **no** checkpoint (failure still halts); under `mid-flight`, or when the tracer carries ``, a `checkpoint:human-verify` is presented. Full precedence chain: `gsd-core/references/checkpoints.md` \u2192 \"Tracer feedback gate\".", + 'plan-md.md tracer Autonomy cell drifted. CONTEXT.md names this table the canonical schema ' + + 'reference — update the cell AND this expected string together.'); + }); + + // Structural, not copy-pinned: the contract is the SHAPE of the verify, so a + // wording improvement to the placeholder must not false-fail (round-5 Minor). + test('planner tracer template emits exactly one -wrapped verify', () => { + const i = soleOperativeIndex('agents/gsd-planner.md', PLANNER, PLANNER_ANCHOR, 'tracer task shape marker'); + const lines = PLANNER.split(/\r?\n/); + // The fence OPENER must be operative AND the first non-blank line after the + // marker. Matching the first raw ```xml in the remainder let a commented-out + // decoy template be selected while the live one regressed (review round 8) — + // this was the one selection in the suite that was not fence/comment aware. + let openIdx = -1; + for (let j = i + 1; j < lines.length; j++) { + if (lines[j].trim() === '') continue; + openIdx = j; + break; + } + assert.notStrictEqual(openIdx, -1, 'the tracer task shape marker must be followed by content'); + // Shape off the RAW line, liveness off the position. The round-8 fix asked + // `operativeLineSet` — which deletes the opener it is asking about, inverting + // fence parity downstream: on the head planner it reported the live "Task-level + // TDD" fence at 0-based 233 as NON-operative, and the assertion only passed + // because 0-based 262 happened to land in a surviving parity slot. One extra + // live ```xml example anywhere earlier in the file flipped it to a false + // FAILURE blaming a decoy that does not exist (review round 9). + assert.ok( + /^\s*```xml(?:\s.*)?$/.test(lines[openIdx]) && isOperativePosition(PLANNER, openIdx), + 'the first non-blank line after the tracer task shape marker must be a LIVE ```xml fence opener — ' + + 'a commented-out or non-adjacent decoy template must not be selectable', + ); + const after = lines.slice(openIdx).join('\n'); + const fence = after.match(/```xml[^\r\n]*\r?\n([\s\S]*?)```/); + assert.ok(fence, 'the tracer task shape must be followed by a fenced xml block'); + const verifies = fence[1].match(/[\s\S]*?<\/verify>/g) || []; + assert.strictEqual(verifies.length, 1, `the tracer template must contain exactly ONE , found ${verifies.length}`); + const inner = verifies[0].replace(/^/, '').replace(/<\/verify>$/, '').replace(/\s+/g, ' ').trim(); + assert.match(inner, /^[^<>]+<\/automated>$/, + "the tracer template's body must be exactly one non-empty child — the #3299 " + + 'gate auto-continues only on an automated-only verify, so a bare-text template makes the fix ' + + 'unreachable for every tracer the planner generates'); + }); + + test('every site reading the mode passes an explicit --default end-of-phase', () => { + // Requires the ASSIGNMENT, not merely the command. Matching the config-get + // substring alone let `IGNORED_MODE=$(gsd_run query config-get ...)` keep this + // row green while nothing defines HUMAN_VERIFY_MODE — the gate then falls + // through to STOP and #3299 is back with the regression suite still passing. + // A test that survives the regression it exists to catch is not a test. + // The lookahead after `end-of-phase` closes the other half: the bare prefix + // also accepted `--default end-of-phase-wrong`. (Codex review, round 9.) + const READ = /^\s*HUMAN_VERIFY_MODE=\$\(gsd_run query config-get workflow\.human_verify_mode --default end-of-phase(?=\s|$)/; + // NOT operativeLineIndexes here, deliberately. All three reads live inside a + // ```bash fence, which is their correct executable form in these files, and + // that selector excludes fenced lines by design — using it would assert the + // opposite of the shipped shape. What the original bare whole-file + // assert.match genuinely could not catch is a read present ONLY inside an + // HTML comment, or a second drifted copy alongside the live one. Pin both: + // exactly one occurrence, inside a live fence, outside any comment. + const BASH_FENCES = new Set(['bash', 'sh', 'shell', 'zsh']); + const liveFencedReads = (md) => { + let fenceLang = null, inComment = false, hits = 0; + for (const raw of md.split(/\r?\n/)) { + let line = raw; + if (inComment) { + const end = line.indexOf('-->'); + if (end === -1) continue; + line = line.slice(end + 3); + inComment = false; + } + // Strip COMPLETE spans first: a one-line comment carries both + // delimiters, so an open/close test that only looks for an unpaired `/g, ''); + const open = line.indexOf('