Files
msd-core/get-shit-done/references/planner-human-verify-mode.md
Tom Boucher 25fb81d01e feat(3309): workflow.human_verify_mode = end-of-phase (new default; mid-flight opt-back-in) (#3325)
* test(3309): red — workflow.human_verify_mode contract

New behavioral test file covers:
- workflow.human_verify_mode is a recognized config key (VALID_CONFIG_KEYS)
- defaults to 'mid-flight' (preserves current behavior)
- config-set / config-get round-trips for both values
- persists in config.json as string
- planner agent file references the flag with canonical wording, couples
  end-of-phase mode with the rule that checkpoint:human-verify is not
  emitted, and documents the <verify><human-check> deferred-item shape
- verifier agent file references harvesting <verify><human-check> blocks
- references/checkpoints.md documents the cost-control alternative

Source-text assertions on agent .md files are exempted via
allow-test-rule: source-text-is-the-product — those files ARE the
runtime contract loaded by AI runtimes, so asserting their wording is
the only way to verify the agents will respect the flag.

Fails 10/11 against current source. Will pass after the fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(3309): add workflow.human_verify_mode = end-of-phase opt-out

Each mid-flight checkpoint:human-verify halt costs a full executor
cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on every
respawn) because subagent context is discarded across the pause. A plan
with N human-verify checkpoints pays the cold-start cost N+1 times. The
reporter (rentanything-nb) measured this at "tens of thousands of tokens"
per round-trip and "hundreds of thousands per week."

This adds workflow.human_verify_mode (default 'mid-flight') with an
'end-of-phase' value that:
- instructs gsd-planner to NOT emit <task type="checkpoint:human-verify">
  tasks; verification details go into a <verify><human-check> sub-block
  on the relevant auto task instead
- instructs gsd-verifier (Step 8) to harvest those <verify><human-check>
  blocks at end-of-phase and merge them into its own human-verification
  list
- the existing human_needed → HUMAN-UAT.md flow in execute-phase.md is
  the single sink — no new file/writer is created

checkpoint:decision and checkpoint:human-action are unaffected — those
gate the work itself, not post-hoc verification.

Surfaces touched:
- bin/lib/config-schema.cjs, bin/lib/config.cjs — register key + default
- sdk/src/config.ts, sdk/src/query/config-schema.ts — SDK parity
- agents/gsd-planner.md — slim Detection section + reference link
- agents/gsd-verifier.md — Step 8 harvest instruction
- get-shit-done/references/planner-human-verify-mode.md — full rules,
  loaded conditionally to keep planner.md under its size budget
- get-shit-done/references/checkpoints.md — surface the alternative
- docs/CONFIGURATION.md — config table row
- docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json — track new reference

Tag name <human-check> chosen instead of <human> to avoid the
prompt-injection scan pattern that flags <system|assistant|human> tags.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(3309): align changeset pr: to actual PR number

The pr: field was authored as 3319 (a guess at the next number) before
the PR was opened. Actual PR is #3325.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(3309): flip workflow.human_verify_mode default to end-of-phase

Per maintainer direction on PR #3325, end-of-phase is the new project
default. Mid-flight checkpoint:human-verify halts cost a full executor
cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per
round-trip — reported at "tens of thousands of tokens" per round-trip,
"hundreds of thousands per week" on real projects. The cost-control
mode is what new projects should get out of the box.

mid-flight remains a one-line opt-back-in via:

    gsd config-set workflow.human_verify_mode mid-flight

Behavior change for existing projects: the new default takes effect
when .planning/config.json is rewritten (config-set, fresh project).
Existing in-flight PLAN.md files with checkpoint:human-verify tasks
continue to work in either mode — the flag only changes what the
planner emits next time it runs.

Surfaces updated:
- bin/lib/config.cjs, sdk/src/config.ts — default flipped
- sdk/src/config.ts docstring — describes new default + opt-back-in
- agents/gsd-planner.md — Detection section explains new default
- references/planner-human-verify-mode.md — reordered modes; added
  guidance on when to opt back into mid-flight
- references/checkpoints.md — surface the default flip and the why
- docs/CONFIGURATION.md — table row reflects new default + reason
- tests/feat-3309-human-verify-mode.test.cjs — default test asserts
  end-of-phase
- .changeset/fierce-geese-march.md — describes the default flip and
  the migration semantics

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: address human verify mode review

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 23:29:11 -04:00

3.5 KiB

Planner — Human Verification Mode

Loaded by gsd-planner when deciding whether to emit <task type="checkpoint:human-verify"> tasks. Read workflow.human_verify_mode from .planning/config.json (default end-of-phase since #3309).

The two modes

end-of-phase (default — issue #3309)

Do not emit any <task type="checkpoint:human-verify"> tasks. Every mid-flight halt costs a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) because subagent context is discarded across the pause; a plan with N human-verify checkpoints pays the cold-start cost N+1 times — measured at "tens of thousands of tokens" per round-trip on real projects. This is the default for that reason.

Instead, fold each would-be verification step into the relevant auto task using a <verify><human-check> sub-block:

<task type="auto">
  <name>Wire dashboard route</name>
  <files>app/dashboard/page.tsx, app/api/dashboard/route.ts</files>
  <action>...</action>
  <verify>
    <automated>npm test -- --filter=dashboard</automated>
    <human-check>
      <test>Visit http://localhost:3000/dashboard</test>
      <expected>Sidebar left, content right on desktop &gt;1024px; collapses to hamburger at 768px</expected>
      <why_human>Visual layout — grep cannot verify breakpoint behavior</why_human>
    </human-check>
  </verify>
  <done>Layout renders correctly across breakpoints</done>
</task>

The verifier (Step 8) harvests every <verify><human-check> block at end-of-phase and consolidates them into the existing human_needed → HUMAN-UAT.md path in workflows/execute-phase.md. The user reviews everything in one batch instead of paying a cold-start cost per item.

mid-flight (opt-back-in — pre-#3309 behavior)

Set gsd config-set workflow.human_verify_mode mid-flight to restore the canonical mid-flight pattern: emit <task type="checkpoint:human-verify"> tasks at the points where human confirmation is required, and the executor halts at each one to ask the user.

<task type="checkpoint:human-verify" gate="blocking">
  <what-built>Dev server running at http://localhost:3000</what-built>
  <how-to-verify>
    1. Visit /dashboard
    2. Sidebar collapses at 768px
  </how-to-verify>
  <resume-signal>"approved" or describe issues</resume-signal>
</task>

Choose mid-flight when you genuinely need the work to stop before any subsequent task runs (e.g., the next task depends on visual confirmation of the previous one), and you accept the cold-start cost as the price of that hard barrier.

What is not affected

checkpoint:decision and checkpoint:human-action tasks are still emitted in end-of-phase mode. Those gate the work itself (a choice the executor needs from the user, or an auth step only the user can perform), not post-hoc verification of completed work. Only checkpoint:human-verify is suppressed.

Compatibility with other modes

  • workflow.tdd_mode: orthogonal. TDD tasks still emit tdd="true" and <behavior>; the <verify> block carries the human-check sub-element when human_verify_mode = end-of-phase.
  • MVP_MODE: orthogonal. Vertical-slice ordering is unchanged. The first task remains a failing end-to-end test; later auto tasks may carry <verify><human-check> instead of standalone checkpoint tasks.
  • workflow.auto_advance / _auto_chain_active: in mid-flight mode these auto-approve checkpoint:human-verify halts. In end-of-phase mode there are no halts to auto-approve, so the flags have no effect on this code path.