* fix(#3299): tracer feedback gate honors workflow.human_verify_mode
The tracer feedback gate (#2294) predates `workflow.human_verify_mode`
(#3309, whose scope was the planner and verifier only), and branched on
auto-mode alone. Under the documented `end-of-phase` default an
interactive run therefore halted after EVERY `type="tracer"` task,
synthesizing a `checkpoint:human-verify` no planner ever emitted and
asking the user to retype a verdict the executor had just computed —
at the cost of a full executor cold-start each time.
Planner-side suppression cannot reach this halt because the executor
synthesizes it at runtime, which is why #3309 did not close it.
The gate now branches on HUMAN_VERIFY_MODE in the interactive path:
under `end-of-phase` an automated-only tracer `<verify>` is re-run and,
on success, expansion continues with no checkpoint. HALT-on-failure is
unchanged. `mid-flight`, `gate="blocking-human"`, and tracers carrying
genuine `<human-check>` evidence all still stop; the autonomous branch
is untouched.
`--default end-of-phase` on the config read is load-bearing, not
decorative: `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS,
so a bare `config-get` exits non-zero with `Key not found` on any
project whose config.json predates #3309 — which is the reporter's
exact config and every pre-existing project.
Both copies of the rule (workflows/execute-plan.md and
agents/gsd-executor.md) are updated together; the reference doc records
the seam and the human-check-still-halts rationale so it cannot recur.
Fixes #3299
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(#3299): add changeset
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): reconcile the canonical schema table and the stale acceptance test
Review round 1 (trek-e) — three items, all in the drift class this PR is
about, two of them landed inside this PR's own diff.
1. docs/reference/plan-md.md:233 — CONTEXT.md names this file the canonical
schema reference for the tracer task-type contract, and its Task-types row
still claimed interactive runs unconditionally present a
checkpoint:human-verify. CONTEXT.md and docs/AGENTS.md were updated in the
first round; this one was missed, so the authoritative reference was the
wrong answer. The row now carries the human_verify_mode-conditional
behavior and points at the canonical precedence chain.
2. tests/tracer-bullet.test.cjs — the docs assertion only checked that a
tracer ROW EXISTS, never its content, which is why CI could not see the
drift. It now asserts the row's actual claims and rejects the pre-#3299
wording. Separately, the #1945 acceptance test named 'interactive run emits
checkpoint:human-verify after the tracer' kept passing only because its
substrings still occur in the fallback clause, while its name asserted the
opposite of shipped behavior. Renamed and narrowed to what #1945 still
guarantees, plus a new interactiveIsConditional pin so the unconditional
prose cannot be restored under a passing substring check.
3. plan-md.md's <verify> row now documents that the legacy bare-text form
(valid, and still shown at :179) does not reach the #3299 auto-continue —
only a <verify> carrying <automated> does — so the benefit is silently
unreachable for tracers using that format.
Mutation-verified: reverting the plan-md row fails 1 test; reverting the
executor's interactive branch fails 4.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): make the tracer gate reachable from the planner template, and bind the assertions
Peer review round 3 found two Majors, both verified by reproducing the
mutation before fixing.
MAJOR 1 — the fix was largely inert on its own default path.
agents/gsd-planner.md's Nyquist Rule (:191) says every <verify> includes
<automated>, but the tracer-specific template twelve lines later emitted the
legacy bare-text form. The gate auto-continues only on a <verify> carrying
only <automated>, so every tracer produced from the canonical template fell
to the STOP fallback and #3299's benefit was unreachable for exactly the task
type it targets. Template now wraps in <automated>; a contract assertion pins
it so the two cannot drift apart again.
MAJOR 2 — the new assertions did not bind condition to action.
Appending 'Nevertheless, interactive runs always present a
checkpoint:human-verify' to the canonical row, and 'then immediately STOP and
return a checkpoint:human-verify' to the auto-continue clause in BOTH
operative copies, restored unconditional interactive checkpointing and left
the suite 35/35 green. Every required keyword still matched. Fixed by:
- clause 2 must now contain no STOP outcome and emit no checkpoint at all —
'never a checkpoint' has to be true OF the clause, not merely stated in it;
- interactiveIsConditional replaced with the ordered-clause parse plus the
same no-STOP property, instead of proving only that HUMAN_VERIFY_MODE
appears somewhere on the line;
- the plan-md.md Autonomy cell is now pinned EXACTLY rather than by keyword
presence. Deliberately brittle: CONTEXT.md names that table the canonical
schema reference, so a wording change must be a conscious edit in both
places.
Mutation-verified after the fix: the combined semantic regression now fails 3
tests; reverting the planner template fails 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): exact-pin the safety clauses instead of blacklisting outcome verbs
Peer review round 4. Blacklisting did not hold, twice over:
- Round 3 banned literal STOP and the 'return a'/'present a' checkpoint
forms in the auto-continue clause. Round 4 defeated that by appending
'then pause and invoke checkpoint_protocol with a checkpoint:human-verify
before expansion' — none of the banned tokens, same restored interruption
after every successful tracer. 36/36 passed.
- The planner guard looked for <automated> anywhere inside <verify>, so
'<verify>[...]<!--<automated>--></verify>' satisfied it while leaving the
legacy bare form operative. 107/107 passed across tracer, planner and the
three size-cap suites.
Synonyms are unbounded; the clauses are not. Both are now pinned exactly on
normalized whitespace, the same approach already proven on the plan-md.md
Autonomy cell, with defence-in-depth checks behind them: no checkpoint-emitting
or blocking outcome in any wording inside clause 2, and the planner's <verify>
body must be exactly one non-empty <automated> child with no commented markup.
These pins are deliberately brittle. Each is a safety contract, so changing the
behavior must be a conscious edit in both the prose and the expectation.
Mutation-verified: the synonym-checkpoint mutation fails 1; the commented-out
wrapper fails 1; the round-3 literal-STOP + contradictory-doc-row regression
fails 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): strip comments, require uniqueness, pin whole regions
Peer review round 5. Exact-pinning one clause was still bypassable two ways,
both reproduced before fixing (each left the suite fully green):
- COMMENTED DECOYS. Put the correct text in an HTML comment followed by a live
wrong copy: every extractor selected the commented decoy. Worked against the
planner template, the canonical plan-md.md row, and both executor branches.
- SURROUNDING OVERRIDE. Insert 'after every tracer, pause and invoke
checkpoint_protocol before expansion, regardless of the mode-specific rules
below' immediately ABOVE the pinned clause, or 'ignore row 3; always wait for
approval' below the canonical table. The pinned text was untouched, so
equality held while the shipped meaning inverted.
The shape that holds, applied to every operative surface:
1. strip HTML comments BEFORE selecting, so a decoy cannot be chosen;
2. require the structural anchor to occur EXACTLY ONCE, so a live second copy
cannot hide behind a correct first one;
3. pin the ENTIRE decision region, not one clause, so no unparsed prefix or
suffix can override what the pin proves.
Applied to: the executor's whole tracer branch, execute-plan.md's whole
dispatch line, checkpoints.md's whole precedence section, and plan-md.md's
Autonomy cell.
Also addresses the round-5 Minor: the planner template is now asserted
STRUCTURALLY (exactly one <verify> in the fenced block, body exactly one
non-empty <automated> child) rather than pinning the descriptive placeholder
verbatim, so behavior-preserving wording changes no longer false-fail. The
clause and section pins keep their exact form — those have a safety rationale
the placeholder copy does not.
Mutation-verified, all six rounds: override-above-clause 1; commented decoy row
1; commented decoy branch 1; ignore-row-3 override 1; synonym checkpoint 1;
commented-out wrapper 2.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): drop the superseded exact-placeholder planner assertion
Peer review round 6, Minor. The round-5 brittleness fix ADDED a structural
planner assertion but left the old exact-placeholder one in place, so the
over-brittleness it was meant to remove was still live: rewording the
descriptive placeholder while preserving exactly one non-empty direct
<automated> child failed the old test and passed the new one.
Removed the old test. The structural assertion is the real contract — the gate
auto-continues on the SHAPE of the verify, not on the wording of a placeholder.
Verified both directions: a behavior-preserving reword now passes; reverting the
template to bare <verify> still fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): select operative prose via parsePredicates, not a hand-rolled scanner
Peer review round 7. I had judged the round-6 selector bypass adversarial-only
and out of scope, intending to disclose it. Both premises were wrong, and the
review said so:
- 'Needs new src API' — false. parsePredicates is ALREADY a public export and
internally uses the repo's interleaved fence/comment scanner. Instrumenting
candidate lines as throwaway predicate declarations borrows that scanner with
no src change at all.
- 'Adversarial-only' — false, and this is the part that mattered. Two ORDINARY
edits silently turned the guards into decoy checks:
* a forgotten '-->' comments the live rule through to EOF, and the
balanced-only stripper still saw and accepted the commented rule;
* a normal fenced documentation example of the rule, plus a whitespace-only
reformat of the live list item, made the selector choose the example.
Neither needs intent. A dangling comment is a typo; a fenced example is good
documentation. Together they reproduce exactly the accidental drift #3299 came
from — with CI green.
The selection layer now defers to parsePredicates for operativeness, uses
whitespace-tolerant anchors so a reformat cannot decouple the live line from its
pin, extracts regions by operative line index rather than string search, and
carries a self-guard test proving fenced / balanced-commented /
after-unclosed-comment copies are all excluded. The helper also ignores indexes
it did not inject, so a pre-existing GSDTEST.CANDIDATE line cannot pollute it.
Verified both ordinary-edit scenarios now fail the suite (each was green before).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): close the operative-selection gaps the maintainer blocked on
trek-e's Blocker: the operative-line selection layer had three gaps, all
reachable by ordinary future doc edits rather than sabotage. He independently
found a fourth I had not disclosed. All are fixed.
1. INDENTATION PROMOTION (his find, not in my disclosure). The instrumentation
replaced a matched candidate with an UNINDENTED marker regardless of the
original line's indentation. A 4-space-indented CommonMark code block is not
skipped by parsePredicates (it accepts indented declarations by design), so
stripping the indent PROMOTED an indented decoy to operative — the exact
inversion of the guard's purpose. The marker now preserves the original
indent, and a candidate that is itself indented 4+ spaces is never injected.
2. NO SET MEMBERSHIP. The filter accepted any in-range integer, so a
pre-existing literal GSDTEST.CANDIDATE=<valid index> in source text could
pollute the count. Now filters on a Set of the indexes actually injected on
this call.
3. RAW FENCE SELECTION (planner). The template test matched the first raw
```xml fence after the marker with no fence/comment awareness — the one
selection in the suite that was not operative-aware — so a commented-out
decoy template between the marker and the real one would be selected while
the live template regressed. The opener must now be operative AND the first
non-blank line after the marker.
4. RAW END ANCHOR (regionFrom). The end anchor was tested against raw lines, so
a fenced example containing a ### / <type line truncated the pinned region
early — a false FAILURE on a legitimate doc edit. End anchors now go through
the same operative filter as start anchors.
Mutation-verified: the indented-decoy + whitespace-varied-anchor combination
and the commented-out fence decoy each now fail the suite (both passed clean
before). Truncation is confirmed fixed by extraction — the region spans the
full section and retains the content following a fenced example, where it
previously stopped at it.
Note on the remaining brittleness: adding a fenced example INSIDE a pinned
region still fails the whole-region exact pin. That is the intended tradeoff
for a safety contract, not the truncation defect, and is called out as such.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): allow-list operative indentation; pin marker provenance
Review round 9.
BLOCKER — the round-8 indentation guard was written as a DENY-list,
/^(?: {4,}|\t)/, and CommonMark has more indented-code forms than that
enumerates: " \t", " \t" and " \t" all open an indented code block and all
slipped through, so an indented decoy was still promoted to operative while the
live rule regressed (34/34 green). Inverted to an allow-list — only 0-3 literal
spaces is ordinary block indentation; anything else is code. Enumerating the
bad shapes was the error, not the specific regex.
MINOR — the injected-index Set validated the marker's VALUE but not its SOURCE.
A pre-existing literal `GSDTEST.CANDIDATE=<n>` could name an index that some
other (skipped) candidate had contributed to the set, and be accepted. Now also
requires p.line - 1 === Number(p.value): the predicate must have been parsed
from the line it names.
MINOR (false negative) — ```xml title=x is a valid CommonMark info string, and
requiring exactly ```xml failed the suite (33/34) on a behavior-preserving edit.
Both the opener assertion and the extraction now accept an info string.
Mutation-verified: the mixed " \t" decoy and the forged-provenance marker each
now fail; the info-string fence no longer false-fails.
KNOWN LIMITATION, disclosed on the PR rather than papered over: parsePredicates
is a predicate parser, not a general CommonMark operativeness oracle. Two
standards-valid constructs still read as operative — a lazy blockquote
continuation line (state opens only on a line that literally starts with ">"),
and a comment opened mid-line ("prose <!--", where state opens only when the
trimmed line STARTS with "<!--"). Closing those means either teaching the shared
src/context-predicates.cts about container/lazy-continuation state — a change to
a module every health rule consumes, well outside a tracer-gate fix — or
hand-rolling a CommonMark parser inside a test, which is how this suite got into
trouble in the first place. Left for the maintainer to scope.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(#3299): re-arm the execute-plan.md emitted-drift ack after the base merge
The #3299 ack rode on tests/emitted-drift-acks/2652-quick-diagnose-dispatch-isolation.json,
which upstream retired in 362d0434b (#3370) once #2728's entries were spent.
#3370's own fragment now owns execute-plan.md at the base, so a new
3299-*.json naming that path would collide — mergeAckSources rejects a
duplicate key across fragments rather than silently last-winning.
Re-arms #3370's entry instead, the mechanism the gate is built for (a spent
ack whose reason changes in the diff is live again), carrying #3370's own
reason forward verbatim so the base growth keeps its account.
Verified: emitted-attribution 175/175 against origin/next@be9329b10.
* fix(#3299): honor golden rule 6 in the tracer gate, extract the chain
Addresses the review on #3390 (B1-B3, M1-M4, minors).
B3 — checkpoints.md asserted two incompatible rules about the same gate.
Golden rule 6 says gate="blocking-human" stops for a human in every mode;
the precedence table scoped row 1 to interactive runs, so a first-match
chain let an auto-mode tracer carrying that gate fall to row 2 and
auto-continue. Rule 6 wins: row 1 is now "Any run, any mode", the
justification sentence it falsified is gone, and the STOP is evaluated
before the auto-mode branch at all three dispatch sites — gsd-executor.md,
execute-plan.md and the plan-md.md schema row. Unreachable by our planner
is not unreachable: src/verify.cts parses only `type` and never consults
`gate` on non-checkpoint tasks, so an imported PLAN.md can carry it.
B1 — the LARGE-tier cap. gsd-executor.md is 49150 on next against a 49152
cap, so this PR could not add a byte. Extracted rather than trimmed: the
precedence chain now lives only in checkpoints.md (already @-imported by
<checkpoint_protocol>, so no new load), and the duplicate summary inside
that protocol section is a pointer. The rationale the earlier trim
deleted is restored — "production-quality, never a throwaway" and
"Pouring more layers onto a broken foundation...". Result 49097: 55 bytes
under the cap and a net 53-byte REDUCTION against next, so the PR returns
headroom instead of consuming it.
B2 — merged upstream/next and resolved all three drift-ack conflicts.
2775 changed shape upstream (string -> {reason}); adopted the new form.
M1 — the 2775 ack claimed the Nyquist Rule sat "twelve lines earlier"; it
is ~75 lines. Corrected to "earlier in the file".
M2 — ack arithmetic restated from measurement, not from a stale base. The
2943 #3299 append is DELETED: with gsd-executor.md now shrinking there is
no ripple to acknowledge, and emitted-attribution correctly flagged the
entry as stale.
M3 — changeset rewritten to the documented bold-lead + em-dash one-liner.
M4 — the two self-defeated shapes are gone. The planner-human-verify-mode
presence checks now go through operativeLineIndexes. The config-get check
does NOT: all three reads live inside ```bash fences, which is their
correct executable form, and that selector excludes fenced lines by
design. It instead pins exactly one live, uncommented, fenced read per
file — mutation-tested against both a commented-out read and a duplicate.
Minors — dangling colon lead-in dropped, a "below" pointer that pointed
above corrected, and the `(default)` asymmetry between the two dispatch
copies aligned.
Two defects the merge surfaced, both caught only by the full suite:
the new #3576 gate rejected this PR's own bare `references/checkpoints.md`
cite in planner-human-verify-mode.md (rewritten to the canonical
gsd-core/ form), and the line-keyed PROSE_ALLOWLIST entry for
gsd-executor.md needed 794 -> 795 after this change shifted the line.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): correct the size record the 08-22 merge falsified
Review round: one Major, four Minors.
Major — the #3299 arm's arithmetic was measured before the merge and is
now wrong in a document whose whole purpose is to be an accurate size
record. Re-measured at head: execute-plan.md is 39315 B on next and
40111 B here, so the 796-byte delta was right but the endpoints and the
headroom were not (849 bytes against DEFAULT_CAP 40960, not 1003). The
superseded figures are named rather than silently replaced. Confirmed
the workflow cap counts LF BYTES while the agent cap counts CHARACTERS —
two caps in two units, one per file.
Minor 1 — 2943-context7-tool-name.json reverted to next. JSON.parse of
both sides was already identical; the diff was an em-dash/times-sign
re-serialization left over from adding and then removing the #3299 arm.
No business in this PR.
Minor 2 — the duplicated `tracer row Autonomy cell` test is gone. Both
copies were new here and carried the same ~8-line canonical string; the
one removed selected its row with a raw startsWith find, the shape this
suite records at :477 as defeated in round 1. Its rationale — why the
cell is pinned EXACTLY, and the append-a-contradiction attack that
defeated keyword matching — is carried onto the surviving fence-aware
copy rather than deleted with it.
Minor 3 — the executor's condensed interactive clause said only "re-run,
continue", which does not distinguish pass from fail; read in isolation
it invites expansion onto a broken slice, the outcome the gate exists to
prevent. Now "re-run; fails → HALT as above, passes → continue, no
checkpoint". The pinned expected string moved with it. Executor at
48,905 chars, 247 under the cap.
Minor 4 — 2775 asserted two different current sizes for gsd-planner.md.
The stale half is next's own text taken wholesale, so the contradiction
was inherited; it now reads as a before-figure rather than a current one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): cite the plan-md example by section, not by a drifting line
Review round 7, Nit N-1. The 2775 ack fragment justified its one-line
formatting with "matching docs/reference/plan-md.md:207's own example
style". At head, :207 is prose; the one-line <verify><automated>
example it means is at :222. The citation was accurate when written
(77c2fda, f23205c) and drifted with a later merge of next.
Re-pointed by section rather than by line — it has already drifted
once, and the fragment's whole purpose is to be an accurate record —
and the drift itself is recorded inline so the correction does not
quietly overwrite what the earlier number said.
Also narrows the changeset's "any task with gate=blocking-human" to
"any tracer carrying gate=blocking-human" (found by Codex in the
whole-PR pass). Golden rule 6 and the #3299 decision table both scope
that gate to checkpoints and to the tracer feedback gate; the normal
type="auto" branch never inspects `gate`, so the wider claim promised
behavior the implementation does not have.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): answer fence-delimiter liveness by insertion, not replacement
Review round 9. The round-8 fence-awareness fix was itself unsound, in the same
class it was added to close.
`operativeLineIndexes` detects operative lines by REPLACING each candidate with
a throwaway predicate declaration and asking `parsePredicates` which survived.
Sound for ordinary content lines. Not sound for a fence DELIMITER, which is
exactly what the tracer-template selection passed it: deleting every ```xml
OPENER leaves each matching closer to become an opener, and since
`computeSkippedLineFlags` is a strict FORWARD state machine, fence parity
inverts for the whole remainder of the document.
Measured against the real file rather than argued:
agents/gsd-planner.md has 3 live top-level ```xml openers — 0-based 180, 232,
262. operativeLineIndexes reported 180 and 262. Line 232, the "Task-level TDD"
example, read NON-OPERATIVE — a wrong answer from a helper whose only job is
that question.
It passed only by parity coincidence, and one extra live example anywhere
earlier flipped it to a false FAILURE blaming a decoy that does not exist:
HEAD as-is | anchor 260 | openIdx 262 | ASSERTION PASSES
+1 unrelated ```xml example | anchor 265 | openIdx 267 | ASSERTION *** FAILS ***
Fixed by asking the question a way that perturbs nothing. `isOperativePosition`
INSERTS a marker on its own line immediately before the candidate instead of
replacing it. Insertion preserves every delimiter, and because the skip-state
machine runs strictly forward, a line inserted at `idx` observes exactly the
fence/comment state the candidate observes, with nothing but the marker between
them — so marker-operative IS the candidate's position-liveness.
The review's suggested direction (substitute a same-shaped opener that still
opens a fence) cannot work here: the marker would then be inside the fence and
would never parse as a predicate at all.
Position-liveness is not content-liveness, so the helper also rejects a line
that is entirely comment (`<!-- ```xml -->`), rather than leaving that to each
caller's own shape test to happen to exclude.
`operativeLineIndexes` now THROWS when its candidate regex matches a fence
delimiter, so the unsound route cannot be reached again by a future caller
rather than only being fixed at the one site that got it wrong.
Verified with the same extra-example scenario above: with the fix, all 35 rows
stay green. Teeth: reverting the call site to `operativeLineSet` turns the
tracer-template row red on the new guard. The regression row pins both live
openers (the second is the one the deletion route lost), the block-commented
and same-line-commented openers, a line inside a fence, and re-checks both
openers after unrelated lines shift above them.
Only tests/tracer-bullet.test.cjs changes — no agent file is touched, so the
5-char gsd-planner.md and 19-byte gsd-executor.md headroom are unaffected.
Verified: `npm run lint:ci` exit 0; full `npm test` 31307 tests / 31292 pass /
0 fail / 14 skipped, TMPDIR unset, against a freshly synced origin/next.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): guard the delimiter class, match the scanner, pin the assignment
Codex full-PR review of #3390, run against the round-9 head. Three defects,
two of them in the code that round added.
1. The mode-read pin survived the regression it exists to catch.
`READ` matched the config-get substring only, so rewriting the shipped line
as `IGNORED_MODE=$(gsd_run query config-get ...)` kept the row green while
nothing defined HUMAN_VERIFY_MODE — the gate falls through to STOP and #3299
is back with the suite passing. The regex now requires the assignment. A
lookahead after `end-of-phase` closes the other half: the bare prefix also
accepted `--default end-of-phase-wrong`. Proven by mutation: renaming the
variable in agents/gsd-executor.md now turns that row red, and did not before.
2. The round-9 fence-delimiter guard was a SAMPLE of the class, not the class.
It probed a fixed list of five delimiter strings. `~~~xml`, ```json, `~~~~`
and arbitrary info strings all walk past any list short enough to write down
— the guard was added precisely because one such regex had already slipped
through. Now matched against the lines the regex actually selects in the
document, which cannot go stale and cannot miss a spelling nobody thought of.
Four such spellings pinned as rows.
3. `isOperativePosition` disagreed with the scanner it delegates to.
For `<!-- closed --> real content` it stripped the span, found surviving
content, and answered "live". `computeSkippedLineFlags` skips an ENTIRE line
whose trimmed text starts with `<!--`, balanced or not, before it considers
fences at all. Verified directly against parsePredicates. It now applies the
scanner's own rule instead of out-reasoning it. Latent for the present caller
(its anchored ```xml shape cannot match a comment-prefixed line), real in
general.
Disclosed rather than fixed, and raised with the maintainer: the exact executor
region pin ends before the second operative tracer-gate paragraph at
agents/gsd-executor.md:327, which is only heading-checked — so contradictory
later instructions could ship. How much of that file to pin is a call for its
owner.
Independently probed isOperativePosition across 19 edge cases before the review
(line 0, CRLF, tab / 4-space / mixed " \t" indentation, 0-3 space fences, nested
fences, ~~~ fences, info strings, bounds); all correct. That probe is what
surfaced finding 2, which the review then confirmed from the other direction.
Verified: `npm run lint:ci` exit 0; full `npm test` 31296 tests / 31281 pass /
0 fail / 14 skipped, TMPDIR unset. One caveat stated rather than smoothed over:
in that run tests/planning-snapshot.test.cjs was truncated by concurrency after
row A5 — 11 tests did not execute, which a 0-fail aggregate cannot show. Re-run
in isolation it is 87 tests / 87 pass / 0 fail, and it is untouched by this
change.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
25 KiB
PLAN.md schema reference
A per-plan PLAN.md is GSD Core's executable unit of work — a structured document that tells an executor agent exactly what to build and how to verify it was built correctly. This page documents its structure. See docs index.
Overview
Plans live inside phase directories at:
.planning/phases/<NN>-<slug>/<NN>-<PP>-PLAN.md
For example: .planning/phases/03-post-feed/03-02-PLAN.md (Phase 3, Plan 2).
Plans are produced by the gsd-planner agent (spawned by /gsd-plan-phase) and consumed by execute-phase. A phase typically contains between one and four plans; plans within a phase are assigned to execution waves so that independent work runs in parallel.
YAML frontmatter
Every PLAN.md opens with a YAML frontmatter block between --- delimiters.
Annotated example
---
phase: 03-post-feed
plan: 02
type: execute
wave: 2
depends_on: ["03-01"]
files_modified:
- src/components/PostFeed.tsx
- src/components/PostCard.tsx
- src/app/feed/page.tsx
files_deleted:
- src/components/LegacyFeed.tsx
autonomous: true
requirements: ["FEED-01", "FEED-03"]
user_setup: []
must_haves:
truths:
- "User can scroll through posts from followed accounts"
- "Each post shows author avatar, name, timestamp, and content"
- "Empty state appears when no posts exist"
artifacts:
- path: "src/components/PostFeed.tsx"
provides: "Scrollable post list"
min_lines: 40
- path: "src/components/PostCard.tsx"
provides: "Individual post card"
exports: ["PostCard"]
key_links:
- from: "src/components/PostFeed.tsx"
to: "src/app/api/feed/route.ts"
via: "fetch in useEffect — calls /api/feed endpoint"
pattern: "fetch.*api/feed"
---
Frontmatter field reference
| Field | Required | Type | Purpose |
|---|---|---|---|
phase |
Yes | string | Phase identifier, e.g. 03-post-feed. |
plan |
Yes | string | Plan number within the phase, e.g. 02. |
type |
Yes | execute or tdd |
execute for standard plans; tdd for test-driven plans where tests are written before implementation. |
wave |
Yes | integer | Execution wave. Plans in wave 1 run in parallel (no dependencies). Plans in wave 2+ wait for all plans in the previous wave to complete. Pre-computed at plan time by gsd-planner. |
depends_on |
Yes | array of plan IDs | Plans this plan must wait for. Empty array = wave 1. Example: ["03-01"] means this plan runs after Plan 01 in Phase 3. |
files_modified |
Yes | array of paths | Every file this plan creates or modifies. Used by the plan-checker to detect same-wave file conflicts and by execute-phase for merge tracking. |
files_deleted |
No | array of paths | Every file this plan deliberately removes. The post-wave cleanup gauntlet blocks the merge of any executor branch whose diff deletes a file — a net against a mass-deletion accident — and this field is the opt-in that names the exceptions. Matching is exact per path after separator normalization: a declared path merges, an undeclared one still blocks that plan's entry (and only that entry). There are no globs and no directory prefixes, so a declaration can never authorize more than it literally lists. Omit the field and the guard's original unconditional block stays in force, which is why absence is always the safe default (#3003). Counts toward same-wave conflict detection alongside files_modified: a plan deleting a file another plan in the same wave is editing is the sharpest conflict there is — one branch removes what the other is writing — so the two plans are pushed into different waves regardless of which side holds the deletion. |
autonomous |
Yes | boolean | true when all tasks are type auto. false when the plan contains any checkpoint:* task that requires human interaction. |
requirements |
Yes | array of IDs | Requirement IDs from ROADMAP.md that this plan addresses. Every phase requirement ID must appear in at least one plan's requirements field. Empty arrays are a BLOCKER. |
user_setup |
No | array of objects | External-service setup steps that Claude cannot automate (account creation, secret retrieval, dashboard configuration). When present, execute-phase generates a USER-SETUP.md checklist for the developer. |
status |
No | superseded |
Marks a plan that was deliberately reassigned or abandoned mid-phase and will never be executed. A status: superseded plan is excluded from the phase's plan and summary counts, so it never holds the phase below 100%. See Superseded plans. Any other value (or the field's absence) has no effect on counting. |
estimate |
No | object | Projected execution cost: {tokens, raw_tokens, tasks, confidence} (#2631, ADR-2629). tokens is an estimateTokens-scale projection with the project's calibration factor already applied (which is why the plan-checker passes --calibrated to estimate-check — re-applying it would square the correction); confidence (low/med/high) is derived from the calibration sample count, never self-rated. Additive and optional — a plan without it behaves exactly as before. A plan estimated above workflow.smart_zone_tokens is flagged with a split recommendation at plan time; the flag is advisory and never blocks. |
must_haves |
Yes | object | Goal-backward verification criteria. See below. |
agent_hint |
No | string | Per-plan specialist executor routing (#1689). Name of a subagent that shares the gsd-executor execution contract (reads execute-plan.md, atomic-commit protocol). When the named agent resolves on the active runtime (an agent file exists in the runtime's agent dir), execute-phase dispatches it instead of gsd-executor. Unset/unresolved → gsd-executor, byte-identical to today. Default-on via workflow.agent_hint_routing; set false to disable. See Per-plan executor routing. |
gap_closure |
Only in gap-closure mode | string, exact match | Must be exactly the literal lowercase true — validated as a string comparison, not a YAML boolean, so True, TRUE, yes, and 1 are all rejected. Required on every plan generated by /gsd-plan-phase --gaps, checked by the plan-gap-closure schema (src/frontmatter.cts) rather than plan. /gsd-execute-phase --gaps-only filters strictly on this field, so an omitted or wrong-valued gap_closure on a gap-closure plan means it is silently skipped — zero executors spawned, no error (#2847). Standard and reviews-mode plans validate against the unmodified plan schema, which neither requires nor checks this field (nothing rejects it as an extra field either, if present). |
Per-plan executor routing
A plan can opt into a specialist executor by setting agent_hint: to the name of a subagent that shares the gsd-executor execution contract — it reads execute-plan.md, follows the atomic-commit protocol, and carries Read/Edit/Write/Bash. A Flutter specialist, for example:
---
agent_hint: well-me-flutter-engineer
---
At dispatch, execute-phase resolves the hint against the active runtime's agent directory (both project-local and user-global, across the runtime's filename variants — .md, .agent.md, .toml, …) and dispatches the named subagent via subagent_type. If the field is absent, blank, or the named agent does not resolve, the plan dispatches to gsd-executor — byte-identical to behavior without the field. Routing is gated by workflow.agent_hint_routing (default-on; see CONFIGURATION).
The specialist agent is an ordinary agent file (e.g. agents/well-me-flutter-engineer.md on Claude Code); there is no separate registration manifest.
Superseded plans
A phase reads complete when every *-PLAN.md has a matching *-SUMMARY.md. When a plan is reassigned or dropped mid-phase — its work folded into a later plan — it will never gain a summary, and without a marker it would pin the phase below 100% forever (the plan-level analogue of a retired phase). Add status: superseded to that plan's frontmatter to exclude it from both the plan count (denominator) and the summary count (numerator):
---
phase: 05-api
plan: "12"
type: execute
status: superseded
---
A phase with 13 plans, two of them superseded, then reads 11/11 → complete — no fabricated summary required. The match is case-insensitive. Plans without the marker are counted exactly as before.
must_haves field
must_haves captures what must be observably true for the phase goal to be achieved. It is derived during planning and verified after execution by the gsd-verifier agent.
Sub-fields
| Sub-field | Type | Purpose |
|---|---|---|
truths |
array of strings | Observable behaviours from the user's perspective. Each must be verifiable. Example: "User can send a message", not "WebSocket library installed". |
artifacts |
array of objects | Files that must exist with substantive implementation (not stubs). |
artifacts[].path |
string | File path relative to project root. |
artifacts[].provides |
string | What capability this file delivers. |
artifacts[].min_lines |
integer (optional) | Minimum line count to be considered non-stub. |
artifacts[].exports |
array of strings (optional) | Expected named exports to verify. |
artifacts[].contains |
string (optional) | Regex or literal pattern that must appear in the file. |
key_links |
array of objects | Critical connections between artifacts — the wiring that makes the system work end-to-end. |
key_links[].from |
string | Source file (relative path from project root). Must be a literal file path — describe components or symbols in via:. |
key_links[].to |
string | Target file (relative path from project root). Must be a literal file path — describe endpoints, modules, or APIs in via:. |
key_links[].via |
string | Description of how they connect, including any endpoint, component, or symbol name (e.g. fetch in useEffect — calls /api/feed, Prisma query via prisma.message, import). |
key_links[].pattern |
string (optional) | Regex to verify the connection exists in source. |
Body structure
After frontmatter, the plan body uses named XML-style blocks read by the executor agent.
<objective>
States what the plan delivers and why it matters for the project:
<objective>
Implement the post feed as a scrollable card list.
Purpose: Core display feature for the social feed phase.
Output: PostFeed and PostCard components wired to /api/feed.
</objective>
<execution_context>
Lists the workflow files associated with executing the plan. Always includes the execute-plan workflow; adds the checkpoints reference when the plan contains checkpoint tasks:
<execution_context>
@~/.claude/gsd-core/workflows/execute-plan.md
@~/.claude/gsd-core/templates/summary.md
</execution_context>
These @ paths point at the local GSD install, not at repository files. The prefix shown here (~/.claude/gsd-core/…) is the Claude global-install location; other runtimes and local installs resolve to their own install directory — for example .cursor/gsd-core/…, or an absolute project path for a --local install. Because the prefix is install-relative, this block is not clone-portable: a committed plan carries whichever prefix the authoring install had. Execution does not depend on it — /gsd-execute-phase loads the execute-plan workflow from its own installed copy — so the block records the execution context rather than resolvable repository references. Contrast <context> (below), whose repository-relative @ paths resolve after a git clone.
<context>
References source files the executor needs to read. Includes project-level planning docs and any source files whose patterns or types the plan must replicate. Prior plan SUMMARY.md files are included only when there is a genuine dependency (imported types, shared decision) — not reflexively:
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@.planning/STATE.md
@src/components/UserCard.tsx
</context>
<tasks>
Contains one or more <task> elements. Every task element must carry <name>, <files>, <read_first>, <action>, <verify>, <acceptance_criteria>, and <done> for type="auto" and type="tracer" tasks. Optional <precondition> (see Preconditions) and <reversibility> (see Reversibility) elements may sit between <name> and <files>.
Preconditions
<precondition> is an optional element on <task> (issue #1949, The Pragmatic Programmer Topic 23 — Design by Contract). It states, in a single line of runnable/checkable prose, what must already be true for the task to begin safely. It closes the front-of-task side of the contract triad — preconditions (before) ↔ postconditions (<verify>/<done>/<acceptance_criteria>, after) ↔ invariants (must_haves.truths, across the whole plan).
<task type="auto">
<name>Add /reveal endpoint handler</name>
<precondition>server bootstraps and responds to GET /health (from the tracer slice)</precondition>
<files>server/reveal.ts</files>
<action>…</action>
<verify>curl /reveal?path=… opens the OS file manager</verify>
<done>Endpoint committed and manually verified</done>
</task>
Optional and back-compat: a plan that omits <precondition> on every task behaves exactly as today — the executor skips the check with no visible change. Adding <precondition> to a task tells the executor to assert it before any other task work (read-only checks only: file existence, env var presence, idempotent health pings; no side-effecting checks — halt and surface a checkpoint if one seems required) and halt (returning a checkpoint:human-verify, no partial commit) on an unmet precondition. Plans that include <precondition> pass verify plan-structure unchanged — the structural validator checks for the presence of required tags and does not reject unknown optional tags.
Emission cases (planner-side): emit <precondition> only when a task relies on state the plan's own depends_on ordering does not already guarantee. Three cases cover every legitimate use:
- External service setup (
user_setupfrontmatter) — the consuming task ties a specific setup step to itself so the executor halts if the setup was skipped. - Prior-phase artifact dependency — a generated schema, a migration's dist output, a contract file from an earlier phase. Cross-phase
depends_ondoes not cross phase boundaries, so<precondition>is the explicit pointer. - Environment variable / runtime configuration — a tool, API, or script the task invokes requires an env var or runtime config that exists now, not at plan time.
Full emission rules, anti-patterns ("the system is ready" is not checkable; do not use <precondition> for intra-plan sequencing — that is what depends_on is for), and the contract triad mapping: see gsd-core/references/planner-preconditions.md.
Reversibility
<reversibility> is an optional element on <task> (issue #1951, The Pragmatic Programmer Topic 15 — "Reversibility"). It records how costly the decision the task implements would be to undo, so a one-way-door choice gets a human beat before the agent walks through it. The rating attribute carries the classification; the body carries a one-line rationale.
<task type="auto">
<name>Define the on-disk event log format</name>
<reversibility rating="one-way">Phases 4-6 read this file; changing the
format after they land requires a migration for every existing project.</reversibility>
<files>src/event-log.cts</files>
<action>…</action>
<verify><automated>npm run test:unit -- event-log</automated></verify>
<done>Format documented and written by the writer under test</done>
</task>
| Rating | Meaning | Effect on the plan |
|---|---|---|
reversible |
Undo is local and cheap. | None. This is the default when no rating is given. |
costly |
Undo touches many call sites or needs a coordinated change. | Flagged in the plan so the reader sees the weight. Never blocks. |
one-way |
Undo requires a migration, breaks a published contract, or is impossible. | The planner inserts a checkpoint:decision immediately before the dependent task. |
Optional and back-compat: a plan that omits <reversibility> on every task behaves exactly as today — no flag, no checkpoint. Plans that include it pass verify plan-structure unchanged; the structural validator checks for the presence of required tags and does not reject unknown optional tags.
Autonomy: inserting a checkpoint:decision means the plan contains a checkpoint, so its frontmatter must set autonomous: false.
Override: /gsd-plan-phase --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses checkpoint insertion for intentionally-unattended runs. Ratings are still recorded and costly items are still flagged — the override changes what stops the run, not what the plan remembers.
Full taxonomy, emission rules, and anti-patterns (chiefly: rating everything one-way produces checkpoint fatigue; prefer removing irreversibility over gating it): see gsd-core/references/planner-reversibility.md.
Task types
| Type | Use | Autonomy |
|---|---|---|
auto |
Everything the executor can do independently. | Fully autonomous. |
tracer |
The leading thin end-to-end slice a plan starts with by default (tracer-first) — production-quality, wired through every layer, with a real end-to-end <verify>. |
Fully autonomous; after committing, the executor runs the tracer's <verify> as an early integration gate. A tracer carrying gate="blocking-human" STOPs for a human in every mode, auto included. Otherwise autonomous runs halt on failure before expansion, and interactive runs honor workflow.human_verify_mode (#3299): under the end-of-phase default a <verify> carrying only <automated> is re-run and, on success, expansion continues with no checkpoint (failure still halts); under mid-flight, or when the tracer carries <human-check>, a checkpoint:human-verify is presented. Full precedence chain: gsd-core/references/checkpoints.md → "Tracer feedback gate". |
checkpoint:human-verify |
Visual or functional verification that requires a human to look at a running UI or service. | Pauses execution; presents to the developer; resumes on approval. |
checkpoint:decision |
Implementation choices that arose during execution and require the developer's input. | Pauses execution; presents options; resumes on selection. |
checkpoint:human-action |
Truly unavoidable manual steps (account creation, hardware interaction). Used sparingly. | Pauses execution; resumes on confirmation. |
Plans that contain any checkpoint task must set autonomous: false in frontmatter.
auto task structure
<task type="auto">
<name>Task 1: Create PostCard component</name>
<files>src/components/PostCard.tsx</files>
<read_first>src/components/UserCard.tsx, src/types/post.ts</read_first>
<action>Create PostCard component accepting a Post prop (id, authorId, content, createdAt,
reactionCount). Render author avatar using UserAvatar from UserCard pattern. Show timestamp
using date-fns formatDistanceToNow. Export as named export PostCard.</action>
<verify>npx tsc --noEmit</verify>
<acceptance_criteria>
- src/components/PostCard.tsx exports named export PostCard
- PostCard.tsx contains "reactionCount" prop usage
- npx tsc --noEmit exits 0
</acceptance_criteria>
<done>PostCard renders post content with author and timestamp</done>
</task>
Required fields for auto tasks
| Field | Rule |
|---|---|
<files> |
Every file the task creates or modifies. The executor writes only these files. |
<read_first> |
Files the executor must read before touching anything — the file being modified, any source-of-truth pattern file, any file whose types or conventions must be replicated. |
<action> |
Concrete instructions with exact identifiers, file paths, function signatures, and expected values. Never says "align X with Y" without specifying the target state. Never contains fenced code blocks or full implementations. |
<verify> |
A runnable command or check that proves the task succeeded. Must distinguish pass from fail — echo "done" is not valid. Accepts either the wrapped form (<verify><automated>cmd</automated></verify>) or the legacy bare-text form (<verify>cmd</verify>); both are valid. On a type="tracer" task, prefer the wrapped form: the tracer feedback gate's auto-continue (#3299) requires a <verify> carrying only <automated>, so a bare-text tracer verify falls through to the STOP fallback and still presents a checkpoint:human-verify in interactive runs even under end-of-phase. |
<acceptance_criteria> |
Verifiable conditions: grep-verifiable strings, command exit codes, observable behaviours. No subjective language ("looks correct", "properly configured"). Negative greps (! grep -Eq 'PAT' file) are file-scoped — region-scope them (sed -n/awk range, then grep) when a sibling task needs the construct elsewhere in the same file (#968). |
<done> |
A short measurable statement of the completed outcome. |
Plan quality dimensions
The gsd-plan-checker agent reviews every PLAN.md across 12 dimensions before execution begins. A plan that fails any BLOCKER-severity check is returned to gsd-planner for revision (up to 3 iterations):
| Dimension | What it checks |
|---|---|
| 1 — Requirement Coverage | Every phase requirement ID from ROADMAP.md appears in at least one plan's requirements frontmatter field and has covering task(s). |
| 2 — Task Completeness | Every auto task carries all required fields (<files>, <action>, <verify>, <acceptance_criteria>, <done>). No vague or empty fields. |
| 3 — Dependency Correctness | depends_on references are valid, acyclic, and consistent with wave numbers. Wave N plan depends only on plans in waves < N. |
| 4 — Key Links Planned | Artifacts in must_haves.key_links have corresponding tasks that implement the wiring — not just the artifact creation. |
| 5 — Scope Sanity | Plans stay within context budget: 2–3 tasks per plan (4 = warning, 5+ = BLOCKER), ≤ 8–10 files per plan (15+ = BLOCKER). |
| 6 — Verification Derivation | must_haves.truths are user-observable behaviours, not implementation details. Artifacts map to truths. Key links cover critical wiring. |
| 7 — Context Compliance | Every D-NN decision from CONTEXT.md is addressed by at least one task. No task implements anything from <deferred>. |
| 7b — Scope Reduction Detection | Task actions do not silently reduce a locked decision to a "v1", "stub", or "future enhancement" without delivering the full decision scope. Always a BLOCKER when found. |
| 7c — Architectural Tier Compliance | Tasks assign capabilities to the correct tier per the RESEARCH.md Architectural Responsibility Map (when present). Security-sensitive capabilities in the wrong tier are BLOCKERs. |
| 8 — Nyquist Compliance | When workflow.nyquist_validation is enabled and RESEARCH.md exists, every task has an <automated> verify command, no consecutive window of 3 tasks lacks coverage, and VALIDATION.md is present. |
| 9 — Cross-Plan Data Contracts | When plans share data pipelines, their transformations are compatible — no plan strips data that another plan needs in original form. |
| 10 — CLAUDE.md Compliance | Plans respect project-specific conventions, forbidden patterns, required tools, and security requirements from ./CLAUDE.md. |
| 11 — Research Resolution | When RESEARCH.md exists, its ## Open Questions section is marked (RESOLVED) before planning proceeds. |
| 12 — Pattern Compliance | When PATTERNS.md exists, tasks reference the correct analog patterns for each new or modified file. |
Wave execution model
Wave numbers are pre-computed during planning. Execute-phase groups plans by wave number and runs each wave's plans in parallel:
Wave 1: Plan 01, Plan 02, Plan 03 (all run simultaneously — no dependencies)
Wave 2: Plan 04 (waits for Wave 1 to complete)
Wave 3: Plan 05 (waits for Wave 2 to complete)
Plans within a wave that modify overlapping files must not be in the same wave — the plan-checker's Dimension 3 flags this as a BLOCKER.
Plan output
After a plan executes successfully, the executor writes a SUMMARY.md at:
.planning/phases/<NN>-<slug>/<NN>-<PP>-SUMMARY.md
The SUMMARY.md is the canonical record of what was built. Subsequent plans in the same phase may reference it when they have a genuine dependency on its types or decisions.