* fix(#3299): tracer feedback gate honors workflow.human_verify_mode
The tracer feedback gate (#2294) predates `workflow.human_verify_mode`
(#3309, whose scope was the planner and verifier only), and branched on
auto-mode alone. Under the documented `end-of-phase` default an
interactive run therefore halted after EVERY `type="tracer"` task,
synthesizing a `checkpoint:human-verify` no planner ever emitted and
asking the user to retype a verdict the executor had just computed —
at the cost of a full executor cold-start each time.
Planner-side suppression cannot reach this halt because the executor
synthesizes it at runtime, which is why #3309 did not close it.
The gate now branches on HUMAN_VERIFY_MODE in the interactive path:
under `end-of-phase` an automated-only tracer `<verify>` is re-run and,
on success, expansion continues with no checkpoint. HALT-on-failure is
unchanged. `mid-flight`, `gate="blocking-human"`, and tracers carrying
genuine `<human-check>` evidence all still stop; the autonomous branch
is untouched.
`--default end-of-phase` on the config read is load-bearing, not
decorative: `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS,
so a bare `config-get` exits non-zero with `Key not found` on any
project whose config.json predates #3309 — which is the reporter's
exact config and every pre-existing project.
Both copies of the rule (workflows/execute-plan.md and
agents/gsd-executor.md) are updated together; the reference doc records
the seam and the human-check-still-halts rationale so it cannot recur.
Fixes #3299
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(#3299): add changeset
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): reconcile the canonical schema table and the stale acceptance test
Review round 1 (trek-e) — three items, all in the drift class this PR is
about, two of them landed inside this PR's own diff.
1. docs/reference/plan-md.md:233 — CONTEXT.md names this file the canonical
schema reference for the tracer task-type contract, and its Task-types row
still claimed interactive runs unconditionally present a
checkpoint:human-verify. CONTEXT.md and docs/AGENTS.md were updated in the
first round; this one was missed, so the authoritative reference was the
wrong answer. The row now carries the human_verify_mode-conditional
behavior and points at the canonical precedence chain.
2. tests/tracer-bullet.test.cjs — the docs assertion only checked that a
tracer ROW EXISTS, never its content, which is why CI could not see the
drift. It now asserts the row's actual claims and rejects the pre-#3299
wording. Separately, the #1945 acceptance test named 'interactive run emits
checkpoint:human-verify after the tracer' kept passing only because its
substrings still occur in the fallback clause, while its name asserted the
opposite of shipped behavior. Renamed and narrowed to what #1945 still
guarantees, plus a new interactiveIsConditional pin so the unconditional
prose cannot be restored under a passing substring check.
3. plan-md.md's <verify> row now documents that the legacy bare-text form
(valid, and still shown at :179) does not reach the #3299 auto-continue —
only a <verify> carrying <automated> does — so the benefit is silently
unreachable for tracers using that format.
Mutation-verified: reverting the plan-md row fails 1 test; reverting the
executor's interactive branch fails 4.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): make the tracer gate reachable from the planner template, and bind the assertions
Peer review round 3 found two Majors, both verified by reproducing the
mutation before fixing.
MAJOR 1 — the fix was largely inert on its own default path.
agents/gsd-planner.md's Nyquist Rule (:191) says every <verify> includes
<automated>, but the tracer-specific template twelve lines later emitted the
legacy bare-text form. The gate auto-continues only on a <verify> carrying
only <automated>, so every tracer produced from the canonical template fell
to the STOP fallback and #3299's benefit was unreachable for exactly the task
type it targets. Template now wraps in <automated>; a contract assertion pins
it so the two cannot drift apart again.
MAJOR 2 — the new assertions did not bind condition to action.
Appending 'Nevertheless, interactive runs always present a
checkpoint:human-verify' to the canonical row, and 'then immediately STOP and
return a checkpoint:human-verify' to the auto-continue clause in BOTH
operative copies, restored unconditional interactive checkpointing and left
the suite 35/35 green. Every required keyword still matched. Fixed by:
- clause 2 must now contain no STOP outcome and emit no checkpoint at all —
'never a checkpoint' has to be true OF the clause, not merely stated in it;
- interactiveIsConditional replaced with the ordered-clause parse plus the
same no-STOP property, instead of proving only that HUMAN_VERIFY_MODE
appears somewhere on the line;
- the plan-md.md Autonomy cell is now pinned EXACTLY rather than by keyword
presence. Deliberately brittle: CONTEXT.md names that table the canonical
schema reference, so a wording change must be a conscious edit in both
places.
Mutation-verified after the fix: the combined semantic regression now fails 3
tests; reverting the planner template fails 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): exact-pin the safety clauses instead of blacklisting outcome verbs
Peer review round 4. Blacklisting did not hold, twice over:
- Round 3 banned literal STOP and the 'return a'/'present a' checkpoint
forms in the auto-continue clause. Round 4 defeated that by appending
'then pause and invoke checkpoint_protocol with a checkpoint:human-verify
before expansion' — none of the banned tokens, same restored interruption
after every successful tracer. 36/36 passed.
- The planner guard looked for <automated> anywhere inside <verify>, so
'<verify>[...]<!--<automated>--></verify>' satisfied it while leaving the
legacy bare form operative. 107/107 passed across tracer, planner and the
three size-cap suites.
Synonyms are unbounded; the clauses are not. Both are now pinned exactly on
normalized whitespace, the same approach already proven on the plan-md.md
Autonomy cell, with defence-in-depth checks behind them: no checkpoint-emitting
or blocking outcome in any wording inside clause 2, and the planner's <verify>
body must be exactly one non-empty <automated> child with no commented markup.
These pins are deliberately brittle. Each is a safety contract, so changing the
behavior must be a conscious edit in both the prose and the expectation.
Mutation-verified: the synonym-checkpoint mutation fails 1; the commented-out
wrapper fails 1; the round-3 literal-STOP + contradictory-doc-row regression
fails 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): strip comments, require uniqueness, pin whole regions
Peer review round 5. Exact-pinning one clause was still bypassable two ways,
both reproduced before fixing (each left the suite fully green):
- COMMENTED DECOYS. Put the correct text in an HTML comment followed by a live
wrong copy: every extractor selected the commented decoy. Worked against the
planner template, the canonical plan-md.md row, and both executor branches.
- SURROUNDING OVERRIDE. Insert 'after every tracer, pause and invoke
checkpoint_protocol before expansion, regardless of the mode-specific rules
below' immediately ABOVE the pinned clause, or 'ignore row 3; always wait for
approval' below the canonical table. The pinned text was untouched, so
equality held while the shipped meaning inverted.
The shape that holds, applied to every operative surface:
1. strip HTML comments BEFORE selecting, so a decoy cannot be chosen;
2. require the structural anchor to occur EXACTLY ONCE, so a live second copy
cannot hide behind a correct first one;
3. pin the ENTIRE decision region, not one clause, so no unparsed prefix or
suffix can override what the pin proves.
Applied to: the executor's whole tracer branch, execute-plan.md's whole
dispatch line, checkpoints.md's whole precedence section, and plan-md.md's
Autonomy cell.
Also addresses the round-5 Minor: the planner template is now asserted
STRUCTURALLY (exactly one <verify> in the fenced block, body exactly one
non-empty <automated> child) rather than pinning the descriptive placeholder
verbatim, so behavior-preserving wording changes no longer false-fail. The
clause and section pins keep their exact form — those have a safety rationale
the placeholder copy does not.
Mutation-verified, all six rounds: override-above-clause 1; commented decoy row
1; commented decoy branch 1; ignore-row-3 override 1; synonym checkpoint 1;
commented-out wrapper 2.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): drop the superseded exact-placeholder planner assertion
Peer review round 6, Minor. The round-5 brittleness fix ADDED a structural
planner assertion but left the old exact-placeholder one in place, so the
over-brittleness it was meant to remove was still live: rewording the
descriptive placeholder while preserving exactly one non-empty direct
<automated> child failed the old test and passed the new one.
Removed the old test. The structural assertion is the real contract — the gate
auto-continues on the SHAPE of the verify, not on the wording of a placeholder.
Verified both directions: a behavior-preserving reword now passes; reverting the
template to bare <verify> still fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): select operative prose via parsePredicates, not a hand-rolled scanner
Peer review round 7. I had judged the round-6 selector bypass adversarial-only
and out of scope, intending to disclose it. Both premises were wrong, and the
review said so:
- 'Needs new src API' — false. parsePredicates is ALREADY a public export and
internally uses the repo's interleaved fence/comment scanner. Instrumenting
candidate lines as throwaway predicate declarations borrows that scanner with
no src change at all.
- 'Adversarial-only' — false, and this is the part that mattered. Two ORDINARY
edits silently turned the guards into decoy checks:
* a forgotten '-->' comments the live rule through to EOF, and the
balanced-only stripper still saw and accepted the commented rule;
* a normal fenced documentation example of the rule, plus a whitespace-only
reformat of the live list item, made the selector choose the example.
Neither needs intent. A dangling comment is a typo; a fenced example is good
documentation. Together they reproduce exactly the accidental drift #3299 came
from — with CI green.
The selection layer now defers to parsePredicates for operativeness, uses
whitespace-tolerant anchors so a reformat cannot decouple the live line from its
pin, extracts regions by operative line index rather than string search, and
carries a self-guard test proving fenced / balanced-commented /
after-unclosed-comment copies are all excluded. The helper also ignores indexes
it did not inject, so a pre-existing GSDTEST.CANDIDATE line cannot pollute it.
Verified both ordinary-edit scenarios now fail the suite (each was green before).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): close the operative-selection gaps the maintainer blocked on
trek-e's Blocker: the operative-line selection layer had three gaps, all
reachable by ordinary future doc edits rather than sabotage. He independently
found a fourth I had not disclosed. All are fixed.
1. INDENTATION PROMOTION (his find, not in my disclosure). The instrumentation
replaced a matched candidate with an UNINDENTED marker regardless of the
original line's indentation. A 4-space-indented CommonMark code block is not
skipped by parsePredicates (it accepts indented declarations by design), so
stripping the indent PROMOTED an indented decoy to operative — the exact
inversion of the guard's purpose. The marker now preserves the original
indent, and a candidate that is itself indented 4+ spaces is never injected.
2. NO SET MEMBERSHIP. The filter accepted any in-range integer, so a
pre-existing literal GSDTEST.CANDIDATE=<valid index> in source text could
pollute the count. Now filters on a Set of the indexes actually injected on
this call.
3. RAW FENCE SELECTION (planner). The template test matched the first raw
```xml fence after the marker with no fence/comment awareness — the one
selection in the suite that was not operative-aware — so a commented-out
decoy template between the marker and the real one would be selected while
the live template regressed. The opener must now be operative AND the first
non-blank line after the marker.
4. RAW END ANCHOR (regionFrom). The end anchor was tested against raw lines, so
a fenced example containing a ### / <type line truncated the pinned region
early — a false FAILURE on a legitimate doc edit. End anchors now go through
the same operative filter as start anchors.
Mutation-verified: the indented-decoy + whitespace-varied-anchor combination
and the commented-out fence decoy each now fail the suite (both passed clean
before). Truncation is confirmed fixed by extraction — the region spans the
full section and retains the content following a fenced example, where it
previously stopped at it.
Note on the remaining brittleness: adding a fenced example INSIDE a pinned
region still fails the whole-region exact pin. That is the intended tradeoff
for a safety contract, not the truncation defect, and is called out as such.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): allow-list operative indentation; pin marker provenance
Review round 9.
BLOCKER — the round-8 indentation guard was written as a DENY-list,
/^(?: {4,}|\t)/, and CommonMark has more indented-code forms than that
enumerates: " \t", " \t" and " \t" all open an indented code block and all
slipped through, so an indented decoy was still promoted to operative while the
live rule regressed (34/34 green). Inverted to an allow-list — only 0-3 literal
spaces is ordinary block indentation; anything else is code. Enumerating the
bad shapes was the error, not the specific regex.
MINOR — the injected-index Set validated the marker's VALUE but not its SOURCE.
A pre-existing literal `GSDTEST.CANDIDATE=<n>` could name an index that some
other (skipped) candidate had contributed to the set, and be accepted. Now also
requires p.line - 1 === Number(p.value): the predicate must have been parsed
from the line it names.
MINOR (false negative) — ```xml title=x is a valid CommonMark info string, and
requiring exactly ```xml failed the suite (33/34) on a behavior-preserving edit.
Both the opener assertion and the extraction now accept an info string.
Mutation-verified: the mixed " \t" decoy and the forged-provenance marker each
now fail; the info-string fence no longer false-fails.
KNOWN LIMITATION, disclosed on the PR rather than papered over: parsePredicates
is a predicate parser, not a general CommonMark operativeness oracle. Two
standards-valid constructs still read as operative — a lazy blockquote
continuation line (state opens only on a line that literally starts with ">"),
and a comment opened mid-line ("prose <!--", where state opens only when the
trimmed line STARTS with "<!--"). Closing those means either teaching the shared
src/context-predicates.cts about container/lazy-continuation state — a change to
a module every health rule consumes, well outside a tracer-gate fix — or
hand-rolling a CommonMark parser inside a test, which is how this suite got into
trouble in the first place. Left for the maintainer to scope.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(#3299): re-arm the execute-plan.md emitted-drift ack after the base merge
The #3299 ack rode on tests/emitted-drift-acks/2652-quick-diagnose-dispatch-isolation.json,
which upstream retired in 362d0434b (#3370) once #2728's entries were spent.
#3370's own fragment now owns execute-plan.md at the base, so a new
3299-*.json naming that path would collide — mergeAckSources rejects a
duplicate key across fragments rather than silently last-winning.
Re-arms #3370's entry instead, the mechanism the gate is built for (a spent
ack whose reason changes in the diff is live again), carrying #3370's own
reason forward verbatim so the base growth keeps its account.
Verified: emitted-attribution 175/175 against origin/next@be9329b10.
* fix(#3299): honor golden rule 6 in the tracer gate, extract the chain
Addresses the review on #3390 (B1-B3, M1-M4, minors).
B3 — checkpoints.md asserted two incompatible rules about the same gate.
Golden rule 6 says gate="blocking-human" stops for a human in every mode;
the precedence table scoped row 1 to interactive runs, so a first-match
chain let an auto-mode tracer carrying that gate fall to row 2 and
auto-continue. Rule 6 wins: row 1 is now "Any run, any mode", the
justification sentence it falsified is gone, and the STOP is evaluated
before the auto-mode branch at all three dispatch sites — gsd-executor.md,
execute-plan.md and the plan-md.md schema row. Unreachable by our planner
is not unreachable: src/verify.cts parses only `type` and never consults
`gate` on non-checkpoint tasks, so an imported PLAN.md can carry it.
B1 — the LARGE-tier cap. gsd-executor.md is 49150 on next against a 49152
cap, so this PR could not add a byte. Extracted rather than trimmed: the
precedence chain now lives only in checkpoints.md (already @-imported by
<checkpoint_protocol>, so no new load), and the duplicate summary inside
that protocol section is a pointer. The rationale the earlier trim
deleted is restored — "production-quality, never a throwaway" and
"Pouring more layers onto a broken foundation...". Result 49097: 55 bytes
under the cap and a net 53-byte REDUCTION against next, so the PR returns
headroom instead of consuming it.
B2 — merged upstream/next and resolved all three drift-ack conflicts.
2775 changed shape upstream (string -> {reason}); adopted the new form.
M1 — the 2775 ack claimed the Nyquist Rule sat "twelve lines earlier"; it
is ~75 lines. Corrected to "earlier in the file".
M2 — ack arithmetic restated from measurement, not from a stale base. The
2943 #3299 append is DELETED: with gsd-executor.md now shrinking there is
no ripple to acknowledge, and emitted-attribution correctly flagged the
entry as stale.
M3 — changeset rewritten to the documented bold-lead + em-dash one-liner.
M4 — the two self-defeated shapes are gone. The planner-human-verify-mode
presence checks now go through operativeLineIndexes. The config-get check
does NOT: all three reads live inside ```bash fences, which is their
correct executable form, and that selector excludes fenced lines by
design. It instead pins exactly one live, uncommented, fenced read per
file — mutation-tested against both a commented-out read and a duplicate.
Minors — dangling colon lead-in dropped, a "below" pointer that pointed
above corrected, and the `(default)` asymmetry between the two dispatch
copies aligned.
Two defects the merge surfaced, both caught only by the full suite:
the new #3576 gate rejected this PR's own bare `references/checkpoints.md`
cite in planner-human-verify-mode.md (rewritten to the canonical
gsd-core/ form), and the line-keyed PROSE_ALLOWLIST entry for
gsd-executor.md needed 794 -> 795 after this change shifted the line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): correct the size record the 08-22 merge falsified
Review round: one Major, four Minors.
Major — the #3299 arm's arithmetic was measured before the merge and is
now wrong in a document whose whole purpose is to be an accurate size
record. Re-measured at head: execute-plan.md is 39315 B on next and
40111 B here, so the 796-byte delta was right but the endpoints and the
headroom were not (849 bytes against DEFAULT_CAP 40960, not 1003). The
superseded figures are named rather than silently replaced. Confirmed
the workflow cap counts LF BYTES while the agent cap counts CHARACTERS —
two caps in two units, one per file.
Minor 1 — 2943-context7-tool-name.json reverted to next. JSON.parse of
both sides was already identical; the diff was an em-dash/times-sign
re-serialization left over from adding and then removing the #3299 arm.
No business in this PR.
Minor 2 — the duplicated `tracer row Autonomy cell` test is gone. Both
copies were new here and carried the same ~8-line canonical string; the
one removed selected its row with a raw startsWith find, the shape this
suite records at :477 as defeated in round 1. Its rationale — why the
cell is pinned EXACTLY, and the append-a-contradiction attack that
defeated keyword matching — is carried onto the surviving fence-aware
copy rather than deleted with it.
Minor 3 — the executor's condensed interactive clause said only "re-run,
continue", which does not distinguish pass from fail; read in isolation
it invites expansion onto a broken slice, the outcome the gate exists to
prevent. Now "re-run; fails → HALT as above, passes → continue, no
checkpoint". The pinned expected string moved with it. Executor at
48,905 chars, 247 under the cap.
Minor 4 — 2775 asserted two different current sizes for gsd-planner.md.
The stale half is next's own text taken wholesale, so the contradiction
was inherited; it now reads as a before-figure rather than a current one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): cite the plan-md example by section, not by a drifting line
Review round 7, Nit N-1. The 2775 ack fragment justified its one-line
formatting with "matching docs/reference/plan-md.md:207's own example
style". At head, :207 is prose; the one-line <verify><automated>
example it means is at :222. The citation was accurate when written
(77c2fda, f23205c) and drifted with a later merge of next.
Re-pointed by section rather than by line — it has already drifted
once, and the fragment's whole purpose is to be an accurate record —
and the drift itself is recorded inline so the correction does not
quietly overwrite what the earlier number said.
Also narrows the changeset's "any task with gate=blocking-human" to
"any tracer carrying gate=blocking-human" (found by Codex in the
whole-PR pass). Golden rule 6 and the #3299 decision table both scope
that gate to checkpoints and to the tracer feedback gate; the normal
type="auto" branch never inspects `gate`, so the wider claim promised
behavior the implementation does not have.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): answer fence-delimiter liveness by insertion, not replacement
Review round 9. The round-8 fence-awareness fix was itself unsound, in the same
class it was added to close.
`operativeLineIndexes` detects operative lines by REPLACING each candidate with
a throwaway predicate declaration and asking `parsePredicates` which survived.
Sound for ordinary content lines. Not sound for a fence DELIMITER, which is
exactly what the tracer-template selection passed it: deleting every ```xml
OPENER leaves each matching closer to become an opener, and since
`computeSkippedLineFlags` is a strict FORWARD state machine, fence parity
inverts for the whole remainder of the document.
Measured against the real file rather than argued:
agents/gsd-planner.md has 3 live top-level ```xml openers — 0-based 180, 232,
262. operativeLineIndexes reported 180 and 262. Line 232, the "Task-level TDD"
example, read NON-OPERATIVE — a wrong answer from a helper whose only job is
that question.
It passed only by parity coincidence, and one extra live example anywhere
earlier flipped it to a false FAILURE blaming a decoy that does not exist:
HEAD as-is | anchor 260 | openIdx 262 | ASSERTION PASSES
+1 unrelated ```xml example | anchor 265 | openIdx 267 | ASSERTION *** FAILS ***
Fixed by asking the question a way that perturbs nothing. `isOperativePosition`
INSERTS a marker on its own line immediately before the candidate instead of
replacing it. Insertion preserves every delimiter, and because the skip-state
machine runs strictly forward, a line inserted at `idx` observes exactly the
fence/comment state the candidate observes, with nothing but the marker between
them — so marker-operative IS the candidate's position-liveness.
The review's suggested direction (substitute a same-shaped opener that still
opens a fence) cannot work here: the marker would then be inside the fence and
would never parse as a predicate at all.
Position-liveness is not content-liveness, so the helper also rejects a line
that is entirely comment (`<!-- ```xml -->`), rather than leaving that to each
caller's own shape test to happen to exclude.
`operativeLineIndexes` now THROWS when its candidate regex matches a fence
delimiter, so the unsound route cannot be reached again by a future caller
rather than only being fixed at the one site that got it wrong.
Verified with the same extra-example scenario above: with the fix, all 35 rows
stay green. Teeth: reverting the call site to `operativeLineSet` turns the
tracer-template row red on the new guard. The regression row pins both live
openers (the second is the one the deletion route lost), the block-commented
and same-line-commented openers, a line inside a fence, and re-checks both
openers after unrelated lines shift above them.
Only tests/tracer-bullet.test.cjs changes — no agent file is touched, so the
5-char gsd-planner.md and 19-byte gsd-executor.md headroom are unaffected.
Verified: `npm run lint:ci` exit 0; full `npm test` 31307 tests / 31292 pass /
0 fail / 14 skipped, TMPDIR unset, against a freshly synced origin/next.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): guard the delimiter class, match the scanner, pin the assignment
Codex full-PR review of #3390, run against the round-9 head. Three defects,
two of them in the code that round added.
1. The mode-read pin survived the regression it exists to catch.
`READ` matched the config-get substring only, so rewriting the shipped line
as `IGNORED_MODE=$(gsd_run query config-get ...)` kept the row green while
nothing defined HUMAN_VERIFY_MODE — the gate falls through to STOP and #3299
is back with the suite passing. The regex now requires the assignment. A
lookahead after `end-of-phase` closes the other half: the bare prefix also
accepted `--default end-of-phase-wrong`. Proven by mutation: renaming the
variable in agents/gsd-executor.md now turns that row red, and did not before.
2. The round-9 fence-delimiter guard was a SAMPLE of the class, not the class.
It probed a fixed list of five delimiter strings. `~~~xml`, ```json, `~~~~`
and arbitrary info strings all walk past any list short enough to write down
— the guard was added precisely because one such regex had already slipped
through. Now matched against the lines the regex actually selects in the
document, which cannot go stale and cannot miss a spelling nobody thought of.
Four such spellings pinned as rows.
3. `isOperativePosition` disagreed with the scanner it delegates to.
For `<!-- closed --> real content` it stripped the span, found surviving
content, and answered "live". `computeSkippedLineFlags` skips an ENTIRE line
whose trimmed text starts with `<!--`, balanced or not, before it considers
fences at all. Verified directly against parsePredicates. It now applies the
scanner's own rule instead of out-reasoning it. Latent for the present caller
(its anchored ```xml shape cannot match a comment-prefixed line), real in
general.
Disclosed rather than fixed, and raised with the maintainer: the exact executor
region pin ends before the second operative tracer-gate paragraph at
agents/gsd-executor.md:327, which is only heading-checked — so contradictory
later instructions could ship. How much of that file to pin is a call for its
owner.
Independently probed isOperativePosition across 19 edge cases before the review
(line 0, CRLF, tab / 4-space / mixed " \t" indentation, 0-3 space fences, nested
fences, ~~~ fences, info strings, bounds); all correct. That probe is what
surfaced finding 2, which the review then confirmed from the other direction.
Verified: `npm run lint:ci` exit 0; full `npm test` 31296 tests / 31281 pass /
0 fail / 14 skipped, TMPDIR unset. One caveat stated rather than smoothed over:
in that run tests/planning-snapshot.test.cjs was truncated by concurrency after
row A5 — 11 tests did not execute, which a 0-fail aggregate cannot show. Re-run
in isolation it is 87 tests / 87 pass / 0 fail, and it is untouched by this
change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
36 KiB
Core principle: Claude automates everything with CLI/API. Checkpoints are for verification and decisions, not manual work.
Golden rules:
- If Claude can run it, Claude runs it - Never ask user to execute CLI commands, start servers, or run builds
- Claude sets up the verification environment - Start dev servers, seed databases, configure env vars
- User only does what requires human judgment - Visual checks, UX evaluation, "does this feel right?"
- Secrets come from user, automation comes from Claude - Ask for API keys, then Claude uses them via CLI
- Auto-mode bypasses verification/decision checkpoints — When
workflow._auto_chain_activeorworkflow.auto_advanceis true in config: human-verify auto-approves, decision auto-selects first option, human-action still stops (auth gates cannot be automated) gate="blocking-human"is never auto-approved — a checkpoint carrying this gate stops for a human in every mode, including auto-mode, regardless of its type. Rule 5 does not apply to it. The executor's precondition-unmet checkpoint (a task's<precondition>evaluated false — unmetuser_setupstep, missing env var, absent prior-phase artifact) reports this gate (#3210).
The gate attribute:
| Value | Auto-mode behavior | Use for |
|---|---|---|
gate="blocking" |
Bypassed per rule 5 (human-verify auto-approves, decision auto-selects) | The default. Post-hoc verification and implementation choices that are safe to take the recommended path on when unattended. |
gate="blocking-human" |
Never bypassed. Stops for a human in auto-mode too. | Irreversible or trust-establishing steps a human must actually see: package-legitimacy verification before install, any decision whose default answer would be wrong to assume, and unmet <precondition> facts the executor cannot establish on its own (#3210). |
Reach for gate="blocking-human" whenever auto-approving the checkpoint would defeat its purpose. If the checkpoint exists because a human must decide something, blocking is the wrong gate — auto-mode will decide it for them.
The gate spans two layers, and both must honor it. gsd-executor refuses to auto-approve a gate="blocking-human" checkpoint and escalates it via checkpoint_return_format precisely so a human sees it; execute-phase's checkpoint_handling step then decides what the user is actually shown. An orchestrator that dispatches on checkpoint type alone would auto-approve the very checkpoint the executor just refused to auto-approve, nullifying that refusal one layer up and letting an unattended --auto / --chain run install a package no human ever vetted.
<checkpoint_types>
## checkpoint:human-verify (Most Common - 90%)When: Claude completed automated work, human confirms it works correctly.
Default mode (#3309):
workflow.human_verify_mode = end-of-phase. New projects do NOT halt mid-flight at planner-emittedcheckpoint:human-verifytasks. (The executor-synthesized tracer feedback gate is the one runtime checkpoint this mode also governs, with its own precedence chain — see "Tracer feedback gate (#3299)" below.) The planner suppresses those task emissions and embeds the verification details into the relevantautotask's<verify><human-check>block; the verifier harvests every<verify><human-check>at end-of-phase (Step 8) and consolidates them into the existinghuman_needed→{phase_num}-UAT.mdflow inworkflows/execute-phase.md. The user reviews everything in one batch.Why this is the default: every mid-flight halt costs a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) because subagent context is discarded across the pause. A plan with N human-verify checkpoints pays the cold-start cost N+1 times — measured at "tens of thousands of tokens" per round-trip on real projects.
Set
workflow.human_verify_mode = mid-flightin.planning/config.jsonto opt back into the pre-#3309 behavior of halting at every checkpoint.checkpoint:decisionandcheckpoint:human-actionare unaffected by either value — those gate the work itself, not post-hoc verification.
Use for:
- Visual UI checks (layout, styling, responsiveness)
- Interactive flows (click through wizard, test user flows)
- Functional verification (feature works as expected)
- Audio/video playback quality
- Animation smoothness
- Accessibility testing
Structure:
<task type="checkpoint:human-verify" gate="blocking">
<what-built>[What Claude automated and deployed/built]</what-built>
<how-to-verify>
[Exact steps to test - URLs, commands, expected behavior]
</how-to-verify>
<resume-signal>[How to continue - "approved", "yes", or describe issues]</resume-signal>
</task>
Example: UI Component (shows key pattern: Claude starts server BEFORE checkpoint)
<task type="auto">
<name>Build responsive dashboard layout</name>
<files>src/components/Dashboard.tsx, src/app/dashboard/page.tsx</files>
<action>Create dashboard with sidebar, header, and content area. Use Tailwind responsive classes for mobile.</action>
<verify>npm run build succeeds, no TypeScript errors</verify>
<done>Dashboard component builds without errors</done>
</task>
<task type="auto">
<name>Start dev server for verification</name>
<action>Run `npm run dev` in background, wait for "ready" message, capture port</action>
<verify>fetch http://localhost:3000 returns 200</verify>
<done>Dev server running at http://localhost:3000</done>
</task>
<task type="checkpoint:human-verify" gate="blocking">
<what-built>Responsive dashboard layout - dev server running at http://localhost:3000</what-built>
<how-to-verify>
Visit http://localhost:3000/dashboard and verify:
1. Desktop (>1024px): Sidebar left, content right, header top
2. Tablet (768px): Sidebar collapses to hamburger menu
3. Mobile (375px): Single column layout, bottom nav appears
4. No layout shift or horizontal scroll at any size
</how-to-verify>
<resume-signal>Type "approved" or describe layout issues</resume-signal>
</task>
Example: Xcode Build
<task type="auto">
<name>Build macOS app with Xcode</name>
<files>App.xcodeproj, Sources/</files>
<action>Run `xcodebuild -project App.xcodeproj -scheme App build`. Check for compilation errors in output.</action>
<verify>Build output contains "BUILD SUCCEEDED", no errors</verify>
<done>App builds successfully</done>
</task>
<task type="checkpoint:human-verify" gate="blocking">
<what-built>Built macOS app at DerivedData/Build/Products/Debug/App.app</what-built>
<how-to-verify>
Open App.app and test:
- App launches without crashes
- Menu bar icon appears
- Preferences window opens correctly
- No visual glitches or layout issues
</how-to-verify>
<resume-signal>Type "approved" or describe issues</resume-signal>
</task>
Tracer feedback gate (#3299)
A type="tracer" task is followed by an early integration checkpoint on the proven slice, run BEFORE any expansion task. This checkpoint is synthesized by the executor at runtime — no planner emits it — so planner-side human_verify_mode suppression cannot reach it. It must therefore consult the mode itself.
Evaluate the rows in order and take the first that matches — they are a precedence chain, not independent conditions:
| # | Run | Tracer <verify> |
Behavior |
|---|---|---|---|
| 1 | Any run, any mode (incl. auto) | task carries gate="blocking-human" |
STOP → checkpoint:human-verify. Never auto-continued. |
| 2 | Auto mode active (AUTO_CHAIN/AUTO_CFG) |
any (row 1 already took blocking-human) |
Re-run verify; HALT on failure, continue on success. Pre-existing behavior — unchanged by #3299. |
| 3 | Interactive, end-of-phase (default) |
only <automated> |
Re-run verify; HALT on failure, continue to expansion on success — no checkpoint |
| 4 | Interactive, end-of-phase |
carries <human-check> |
STOP → checkpoint:human-verify |
| 5 | Interactive, mid-flight |
any | STOP → checkpoint:human-verify |
Carve-outs — the #3299 auto-continue (row 3) applies ONLY when all three hold: the run is interactive, the mode is end-of-phase, and the tracer's <verify> contains only <automated>. Anything else STOPs or falls to the pre-existing auto-mode branch. HALT-on-failure is unconditional in rows 2 and 3 alike: a failing tracer never becomes an approvable checkpoint and never proceeds to expansion, because layering expansion onto a broken slice is the failure this gate exists to prevent.
Row 1 is deliberately not scoped to interactive runs. Golden rule 6 above states that gate="blocking-human" stops for a human in every mode including auto-mode, and a precedence chain that let an autonomous run continue past it would make this file assert two incompatible rules about the same gate. No planner emits gate on a type="tracer" task today, but src/verify.cts parses only type and does not consult gate on non-checkpoint tasks, so a hand-authored, imported, or externally-generated PLAN.md can carry it and validate — unreachable by our planner is not unreachable.
Read HUMAN_VERIFY_MODE with an explicit default — workflow.human_verify_mode is absent from SCHEMA_DEFAULTS, so a bare config-get exits non-zero with Key not found on any project whose config.json predates #3309:
HUMAN_VERIFY_MODE=$(gsd_run query config-get workflow.human_verify_mode --default end-of-phase --raw 2>/dev/null || echo "end-of-phase")
When: Human must make choice that affects implementation direction.
Use for:
- Technology selection (which auth provider, which database)
- Architecture decisions (monorepo vs separate repos)
- Design choices (color scheme, layout approach)
- Feature prioritization (which variant to build)
- Data model decisions (schema structure)
Structure:
<task type="checkpoint:decision" gate="blocking">
<decision>[What's being decided]</decision>
<context>[Why this decision matters]</context>
<options>
<option id="option-a">
<name>[Option name]</name>
<pros>[Benefits]</pros>
<cons>[Tradeoffs]</cons>
</option>
<option id="option-b">
<name>[Option name]</name>
<pros>[Benefits]</pros>
<cons>[Tradeoffs]</cons>
</option>
</options>
<resume-signal>[How to indicate choice]</resume-signal>
</task>
Example: Auth Provider Selection
<task type="checkpoint:decision" gate="blocking">
<decision>Select authentication provider</decision>
<context>
Need user authentication for the app. Three solid options with different tradeoffs.
</context>
<options>
<option id="supabase">
<name>Supabase Auth</name>
<pros>Built-in with Supabase DB we're using, generous free tier, row-level security integration</pros>
<cons>Less customizable UI, tied to Supabase ecosystem</cons>
</option>
<option id="clerk">
<name>Clerk</name>
<pros>Beautiful pre-built UI, best developer experience, excellent docs</pros>
<cons>Paid after 10k MAU, vendor lock-in</cons>
</option>
<option id="nextauth">
<name>NextAuth.js</name>
<pros>Free, self-hosted, maximum control, widely adopted</pros>
<cons>More setup work, you manage security updates, UI is DIY</cons>
</option>
</options>
<resume-signal>Select: supabase, clerk, or nextauth</resume-signal>
</task>
Example: Database Selection
<task type="checkpoint:decision" gate="blocking">
<decision>Select database for user data</decision>
<context>
App needs persistent storage for users, sessions, and user-generated content.
Expected scale: 10k users, 1M records first year.
</context>
<options>
<option id="supabase">
<name>Supabase (Postgres)</name>
<pros>Full SQL, generous free tier, built-in auth, real-time subscriptions</pros>
<cons>Vendor lock-in for real-time features, less flexible than raw Postgres</cons>
</option>
<option id="planetscale">
<name>PlanetScale (MySQL)</name>
<pros>Serverless scaling, branching workflow, excellent DX</pros>
<cons>MySQL not Postgres, no foreign keys in free tier</cons>
</option>
<option id="convex">
<name>Convex</name>
<pros>Real-time by default, TypeScript-native, automatic caching</pros>
<cons>Newer platform, different mental model, less SQL flexibility</cons>
</option>
</options>
<resume-signal>Select: supabase, planetscale, or convex</resume-signal>
</task>
When: Action has NO CLI/API and requires human-only interaction, OR Claude hit an authentication gate during automation.
Use ONLY for:
- Authentication gates - Claude tried CLI/API but needs credentials (this is NOT a failure)
- Email verification links (clicking email)
- SMS 2FA codes (phone verification)
- Manual account approvals (platform requires human review)
- Credit card 3D Secure flows (web-based payment authorization)
- OAuth app approvals (web-based approval)
Do NOT use for pre-planned manual work:
- Deploying (use CLI - auth gate if needed)
- Creating webhooks/databases (use API/CLI - auth gate if needed)
- Running builds/tests (use Bash tool)
- Creating files (use Write tool)
Structure:
<task type="checkpoint:human-action" gate="blocking">
<action>[What human must do - Claude already did everything automatable]</action>
<instructions>
[What Claude already automated]
[The ONE thing requiring human action]
</instructions>
<verification>[What Claude can check afterward]</verification>
<resume-signal>[How to continue]</resume-signal>
</task>
Example: Email Verification
<task type="auto">
<name>Create SendGrid account via API</name>
<action>Use SendGrid API to create subuser account with provided email. Request verification email.</action>
<verify>API returns 201, account created</verify>
<done>Account created, verification email sent</done>
</task>
<task type="checkpoint:human-action" gate="blocking">
<action>Complete email verification for SendGrid account</action>
<instructions>
I created the account and requested verification email.
Check your inbox for SendGrid verification link and click it.
</instructions>
<verification>SendGrid API key works: curl test succeeds</verification>
<resume-signal>Type "done" when email verified</resume-signal>
</task>
Example: Authentication Gate (Dynamic Checkpoint)
<task type="auto">
<name>Deploy to Vercel</name>
<files>.vercel/, vercel.json</files>
<action>Run `vercel --yes` to deploy</action>
<verify>vercel ls shows deployment, fetch returns 200</verify>
</task>
<!-- If vercel returns "Error: Not authenticated", Claude creates checkpoint on the fly -->
<task type="checkpoint:human-action" gate="blocking">
<action>Authenticate Vercel CLI so I can continue deployment</action>
<instructions>
I tried to deploy but got authentication error.
Run: vercel login
This will open your browser - complete the authentication flow.
</instructions>
<verification>vercel whoami returns your account email</verification>
<resume-signal>Type "done" when authenticated</resume-signal>
</task>
<!-- After authentication, Claude retries the deployment -->
<task type="auto">
<name>Retry Vercel deployment</name>
<action>Run `vercel --yes` (now authenticated)</action>
<verify>vercel ls shows deployment, fetch returns 200</verify>
</task>
Key distinction: Auth gates are created dynamically when Claude encounters auth errors. NOT pre-planned — Claude automates first, asks for credentials only when blocked. </checkpoint_types>
<execution_protocol>
When Claude encounters type="checkpoint:*":
- Stop immediately - do not proceed to next task
- Display checkpoint clearly using the format below
- Wait for user response - do not hallucinate completion
- Verify if possible - check files, run tests, whatever is specified
- Resume execution - continue to next task only after confirmation
For checkpoint:human-verify:
╔═══════════════════════════════════════════════════════╗
║ CHECKPOINT: Verification Required ║
╚═══════════════════════════════════════════════════════╝
Progress: 5/8 tasks complete
Task: Responsive dashboard layout
Built: Responsive dashboard at /dashboard
How to verify:
1. Visit: http://localhost:3000/dashboard
2. Desktop (>1024px): Sidebar visible, content fills remaining space
3. Tablet (768px): Sidebar collapses to icons
4. Mobile (375px): Sidebar hidden, hamburger menu appears
────────────────────────────────────────────────────────
→ YOUR ACTION: Type "approved" or describe issues
────────────────────────────────────────────────────────
For checkpoint:decision:
╔═══════════════════════════════════════════════════════╗
║ CHECKPOINT: Decision Required ║
╚═══════════════════════════════════════════════════════╝
Progress: 2/6 tasks complete
Task: Select authentication provider
Decision: Which auth provider should we use?
Context: Need user authentication. Three options with different tradeoffs.
Options:
1. supabase - Built-in with our DB, free tier
Pros: Row-level security integration, generous free tier
Cons: Less customizable UI, ecosystem lock-in
2. clerk - Best DX, paid after 10k users
Pros: Beautiful pre-built UI, excellent documentation
Cons: Vendor lock-in, pricing at scale
3. nextauth - Self-hosted, maximum control
Pros: Free, no vendor lock-in, widely adopted
Cons: More setup work, DIY security updates
────────────────────────────────────────────────────────
→ YOUR ACTION: Select supabase, clerk, or nextauth
────────────────────────────────────────────────────────
For checkpoint:human-action:
╔═══════════════════════════════════════════════════════╗
║ CHECKPOINT: Action Required ║
╚═══════════════════════════════════════════════════════╝
Progress: 3/8 tasks complete
Task: Deploy to Vercel
Attempted: vercel --yes
Error: Not authenticated. Please run 'vercel login'
What you need to do:
1. Run: vercel login
2. Complete browser authentication when it opens
3. Return here when done
I'll verify: vercel whoami returns your account
────────────────────────────────────────────────────────
→ YOUR ACTION: Type "done" when authenticated
────────────────────────────────────────────────────────
</execution_protocol>
<authentication_gates>
Auth gate = Claude tried CLI/API, got auth error. Not a failure — a gate requiring human input to unblock.
Pattern: Claude tries automation → auth error → creates checkpoint:human-action → user authenticates → Claude retries → continues
Gate protocol:
- Recognize it's not a failure - missing auth is expected
- Stop current task - don't retry repeatedly
- Create checkpoint:human-action dynamically
- Provide exact authentication steps
- Verify authentication works
- Retry the original task
- Continue normally
Key distinction:
- Pre-planned checkpoint: "I need you to do X" (wrong - Claude should automate)
- Auth gate: "I tried to automate X but need credentials" (correct - unblocks automation)
</authentication_gates>
<automation_reference>
The rule: If it has CLI/API, Claude does it. Never ask human to perform automatable work.
Service CLI Reference
| Service | CLI/API | Key Commands | Auth Gate |
|---|---|---|---|
| Vercel | vercel |
--yes, env add, --prod, ls |
vercel login |
| Railway | railway |
init, up, variables set |
railway login |
| Fly | fly |
launch, deploy, secrets set |
fly auth login |
| Stripe | stripe + API |
listen, trigger, API calls |
API key in .env |
| Supabase | supabase |
init, link, db push, gen types |
supabase login |
| Upstash | upstash |
redis create, redis get |
upstash auth login |
| PlanetScale | pscale |
database create, branch create |
pscale auth login |
| GitHub | gh |
repo create, pr create, secret set |
gh auth login |
| Node | npm/pnpm |
install, run build, test, run dev |
N/A |
| Xcode | xcodebuild |
-project, -scheme, build, test |
N/A |
| Convex | npx convex |
dev, deploy, env set, env get |
npx convex login |
Environment Variable Automation
Env files: Use Write/Edit tools. Never ask human to create .env manually.
Dashboard env vars via CLI:
| Platform | CLI Command | Example |
|---|---|---|
| Convex | npx convex env set |
npx convex env set OPENAI_API_KEY sk-... |
| Vercel | vercel env add |
vercel env add STRIPE_KEY production |
| Railway | railway variables set |
railway variables set API_KEY=value |
| Fly | fly secrets set |
fly secrets set DATABASE_URL=... |
| Supabase | supabase secrets set |
supabase secrets set MY_SECRET=value |
Secret collection pattern:
<!-- WRONG: Asking user to add env vars in dashboard -->
<task type="checkpoint:human-action">
<action>Add OPENAI_API_KEY to Convex dashboard</action>
<instructions>Go to dashboard.convex.dev → Settings → Environment Variables → Add</instructions>
</task>
<!-- RIGHT: Claude asks for value, then adds via CLI -->
<task type="checkpoint:human-action">
<action>Provide your OpenAI API key</action>
<instructions>
I need your OpenAI API key for Convex backend.
Get it from: https://platform.openai.com/api-keys
Paste the key (starts with sk-)
</instructions>
<verification>I'll add it via `npx convex env set` and verify</verification>
<resume-signal>Paste your API key</resume-signal>
</task>
<task type="auto">
<name>Configure OpenAI key in Convex</name>
<action>Run `npx convex env set OPENAI_API_KEY {user-provided-key}`</action>
<verify>`npx convex env get OPENAI_API_KEY` returns the key (masked)</verify>
</task>
Dev Server Automation
| Framework | Start Command | Ready Signal | Default URL |
|---|---|---|---|
| Next.js | npm run dev |
"Ready in" or "started server" | http://localhost:3000 |
| Vite | npm run dev |
"ready in" | http://localhost:5173 |
| Convex | npx convex dev |
"Convex functions ready" | N/A (backend only) |
| Express | npm start |
"listening on port" | http://localhost:3000 |
| Django | python manage.py runserver |
"Starting development server" | http://localhost:8000 |
Server lifecycle:
# Run in background, capture PID
npm run dev &
DEV_SERVER_PID=$!
# Wait for ready (max 30s) — uses fetch() for cross-platform compatibility
gsd_run run-with-timeout 30 -- bash -c 'until node -e "fetch(\"http://localhost:3000\").then(r=>{process.exit(r.ok?0:1)}).catch(()=>process.exit(1))" 2>/dev/null; do sleep 1; done'
Port conflicts: Kill stale process (lsof -ti:3000 | xargs kill) or use alternate port (--port 3001).
Server stays running through checkpoints. Only kill when plan complete, switching to production, or port needed for different service.
CLI Installation Handling
| CLI | Auto-install? | Command |
|---|---|---|
| npm/pnpm/yarn | No - ask user | User chooses package manager |
| vercel | Yes | npm i -g vercel |
| gh (GitHub) | Yes | brew install gh (macOS) or apt install gh (Linux) |
| stripe | Yes | npm i -g stripe |
| supabase | Yes | npm i -g supabase |
| convex | No - use npx | npx convex (no install needed) |
| fly | Yes | brew install flyctl or curl installer |
| railway | Yes | npm i -g @railway/cli |
Protocol: Try command → "command not found" → auto-installable? → yes: install silently, retry → no: checkpoint asking user to install.
Pre-Checkpoint Automation Failures
| Failure | Response |
|---|---|
| Server won't start | Check error, fix issue, retry (don't proceed to checkpoint) |
| Port in use | Kill stale process or use alternate port |
| Missing dependency | Run npm install, retry |
| Build error | Fix the error first (bug, not checkpoint issue) |
| Auth error | Create auth gate checkpoint |
| Network timeout | Retry with backoff, then checkpoint if persistent |
Never present a checkpoint with broken verification environment. If the local server isn't responding, don't ask user to "visit localhost:3000".
Cross-platform note: Use
node -e "fetch('http://localhost:3000').then(r=>console.log(r.status))"instead ofcurlfor health checks.curlis broken on Windows MSYS/Git Bash due to SSL/path mangling issues.
<!-- WRONG: Checkpoint with broken environment -->
<task type="checkpoint:human-verify">
<what-built>Dashboard (server failed to start)</what-built>
<how-to-verify>Visit http://localhost:3000...</how-to-verify>
</task>
<!-- RIGHT: Fix first, then checkpoint -->
<task type="auto">
<name>Fix server startup issue</name>
<action>Investigate error, fix root cause, restart server</action>
<verify>fetch http://localhost:3000 returns 200</verify>
</task>
<task type="checkpoint:human-verify">
<what-built>Dashboard - server running at http://localhost:3000</what-built>
<how-to-verify>Visit http://localhost:3000/dashboard...</how-to-verify>
</task>
Automatable Quick Reference
| Action | Automatable? | Claude does it? |
|---|---|---|
| Deploy to Vercel | Yes (vercel) |
YES |
| Create Stripe webhook | Yes (API) | YES |
| Write .env file | Yes (Write tool) | YES |
| Create Upstash DB | Yes (upstash) |
YES |
| Run tests | Yes (npm test) |
YES |
| Start dev server | Yes (npm run dev) |
YES |
| Add env vars to Convex | Yes (npx convex env set) |
YES |
| Add env vars to Vercel | Yes (vercel env add) |
YES |
| Seed database | Yes (CLI/API) | YES |
| Click email verification link | No | NO |
| Enter credit card with 3DS | No | NO |
| Complete OAuth in browser | No | NO |
| Visually verify UI looks correct | No | NO |
| Test interactive user flows | No | NO |
</automation_reference>
<writing_guidelines>
DO:
- Automate everything with CLI/API before checkpoint
- Be specific: "Visit https://myapp.vercel.app" not "check deployment"
- Number verification steps
- State expected outcomes: "You should see X"
- Provide context: why this checkpoint exists
DON'T:
- Ask human to do work Claude can automate ❌
- Assume knowledge: "Configure the usual settings" ❌
- Skip steps: "Set up database" (too vague) ❌
- Mix multiple verifications in one checkpoint ❌
Placement:
- After automation completes - not before Claude does the work
- After UI buildout - before declaring phase complete
- Before dependent work - decisions before implementation
- At integration points - after configuring external services
Bad placement: Before automation ❌ | Too frequent ❌ | Too late (dependent tasks already needed the result) ❌ </writing_guidelines>
Example 1: Database Setup (No Checkpoint Needed)
<task type="auto">
<name>Create Upstash Redis database</name>
<files>.env</files>
<action>
1. Run `upstash redis create myapp-cache --region us-east-1`
2. Capture connection URL from output
3. Write to .env: UPSTASH_REDIS_URL={url}
4. Verify connection with test command
</action>
<verify>
- upstash redis list shows database
- .env contains UPSTASH_REDIS_URL
- Test connection succeeds
</verify>
<done>Redis database created and configured</done>
</task>
<!-- NO CHECKPOINT NEEDED - Claude automated everything and verified programmatically -->
Example 2: Full Auth Flow (Single checkpoint at end)
<task type="auto">
<name>Create user schema</name>
<files>src/db/schema.ts</files>
<action>Define User, Session, Account tables with Drizzle ORM</action>
<verify>npm run db:generate succeeds</verify>
</task>
<task type="auto">
<name>Create auth API routes</name>
<files>src/app/api/auth/[...nextauth]/route.ts</files>
<action>Set up NextAuth with GitHub provider, JWT strategy</action>
<verify>TypeScript compiles, no errors</verify>
</task>
<task type="auto">
<name>Create login UI</name>
<files>src/app/login/page.tsx, src/components/LoginButton.tsx</files>
<action>Create login page with GitHub OAuth button</action>
<verify>npm run build succeeds</verify>
</task>
<task type="auto">
<name>Start dev server for auth testing</name>
<action>Run `npm run dev` in background, wait for ready signal</action>
<verify>fetch http://localhost:3000 returns 200</verify>
<done>Dev server running at http://localhost:3000</done>
</task>
<!-- ONE checkpoint at end verifies the complete flow -->
<task type="checkpoint:human-verify" gate="blocking">
<what-built>Complete authentication flow - dev server running at http://localhost:3000</what-built>
<how-to-verify>
1. Visit: http://localhost:3000/login
2. Click "Sign in with GitHub"
3. Complete GitHub OAuth flow
4. Verify: Redirected to /dashboard, user name displayed
5. Refresh page: Session persists
6. Click logout: Session cleared
</how-to-verify>
<resume-signal>Type "approved" or describe issues</resume-signal>
</task>
<anti_patterns>
❌ BAD: Asking user to start dev server
<task type="checkpoint:human-verify" gate="blocking">
<what-built>Dashboard component</what-built>
<how-to-verify>
1. Run: npm run dev
2. Visit: http://localhost:3000/dashboard
3. Check layout is correct
</how-to-verify>
</task>
Why bad: Claude can run npm run dev. User should only visit URLs, not execute commands.
✅ GOOD: Claude starts server, user visits
<task type="auto">
<name>Start dev server</name>
<action>Run `npm run dev` in background</action>
<verify>fetch http://localhost:3000 returns 200</verify>
</task>
<task type="checkpoint:human-verify" gate="blocking">
<what-built>Dashboard at http://localhost:3000/dashboard (server running)</what-built>
<how-to-verify>
Visit http://localhost:3000/dashboard and verify:
1. Layout matches design
2. No console errors
</how-to-verify>
</task>
❌ BAD: Asking human to deploy / ✅ GOOD: Claude automates
<!-- BAD: Asking user to deploy via dashboard -->
<task type="checkpoint:human-action" gate="blocking">
<action>Deploy to Vercel</action>
<instructions>Visit vercel.com/new → Import repo → Click Deploy → Copy URL</instructions>
</task>
<!-- GOOD: Claude deploys, user verifies -->
<task type="auto">
<name>Deploy to Vercel</name>
<action>Run `vercel --yes`. Capture URL.</action>
<verify>vercel ls shows deployment, fetch returns 200</verify>
</task>
<task type="checkpoint:human-verify">
<what-built>Deployed to {url}</what-built>
<how-to-verify>Visit {url}, check homepage loads</how-to-verify>
<resume-signal>Type "approved"</resume-signal>
</task>
❌ BAD: Too many checkpoints / ✅ GOOD: Single checkpoint
<!-- BAD: Checkpoint after every task -->
<task type="auto">Create schema</task>
<task type="checkpoint:human-verify">Check schema</task>
<task type="auto">Create API route</task>
<task type="checkpoint:human-verify">Check API</task>
<task type="auto">Create UI form</task>
<task type="checkpoint:human-verify">Check form</task>
<!-- GOOD: One checkpoint at end -->
<task type="auto">Create schema</task>
<task type="auto">Create API route</task>
<task type="auto">Create UI form</task>
<task type="checkpoint:human-verify">
<what-built>Complete auth flow (schema + API + UI)</what-built>
<how-to-verify>Test full flow: register, login, access protected page</how-to-verify>
<resume-signal>Type "approved"</resume-signal>
</task>
❌ BAD: Vague verification / ✅ GOOD: Specific steps
<!-- BAD -->
<task type="checkpoint:human-verify">
<what-built>Dashboard</what-built>
<how-to-verify>Check it works</how-to-verify>
</task>
<!-- GOOD -->
<task type="checkpoint:human-verify">
<what-built>Responsive dashboard - server running at http://localhost:3000</what-built>
<how-to-verify>
Visit http://localhost:3000/dashboard and verify:
1. Desktop (>1024px): Sidebar visible, content area fills remaining space
2. Tablet (768px): Sidebar collapses to icons
3. Mobile (375px): Sidebar hidden, hamburger menu in header
4. No horizontal scroll at any size
</how-to-verify>
<resume-signal>Type "approved" or describe layout issues</resume-signal>
</task>
❌ BAD: Asking user to run CLI commands
<task type="checkpoint:human-action">
<action>Run database migrations</action>
<instructions>Run: npx prisma migrate deploy && npx prisma db seed</instructions>
</task>
Why bad: Claude can run these commands. User should never execute CLI commands.
❌ BAD: Asking user to copy values between services
<task type="checkpoint:human-action">
<action>Configure webhook URL in Stripe</action>
<instructions>Copy deployment URL → Stripe Dashboard → Webhooks → Add endpoint → Copy secret → Add to .env</instructions>
</task>
Why bad: Stripe has an API. Claude should create the webhook via API and write to .env directly.
</anti_patterns>
## checkpoint:tdd-review (TDD Mode Only)When: All waves in a phase complete and workflow.tdd_mode is enabled. Inserted by the execute-phase orchestrator after aggregate_results.
Purpose: Collaborative review of TDD gate compliance across all type: tdd plans in the phase. Advisory — does not block execution.
Use for:
- Verifying RED/GREEN/REFACTOR commit sequence for each TDD plan
- Surfacing gate violations (missing RED or GREEN commits)
- Reviewing test quality (tests fail for the right reason)
- Confirming minimal GREEN implementations
Structure:
<task type="checkpoint:tdd-review" gate="advisory">
<what-checked>TDD gate compliance for {count} plans in Phase {X}</what-checked>
<gate-results>
| Plan | RED | GREEN | REFACTOR | Status |
|------|-----|-------|----------|--------|
| {id} | ✓ | ✓ | ✓ | Pass |
</gate-results>
<violations>[List of gate violations, or "None"]</violations>
<resume-signal>Review complete — proceed to phase verification</resume-signal>
</task>
Auto-mode behavior: When workflow._auto_chain_active or workflow.auto_advance is true, the TDD review checkpoint auto-approves (advisory gate — never blocks).
Checkpoints formalize human-in-the-loop points for verification and decisions, not manual work.
The golden rule: If Claude CAN automate it, Claude MUST automate it.
Checkpoint priority:
- checkpoint:human-verify (90%) - Claude automated everything, human confirms visual/functional correctness
- checkpoint:decision (9%) - Human makes architectural/technology choices
- checkpoint:human-action (1%) - Truly unavoidable manual steps with no API/CLI
When NOT to use checkpoints:
- Things Claude can verify programmatically (tests, builds)
- File operations (Claude can read files)
- Code correctness (tests and static analysis)
- Anything automatable via CLI/API