feat(3309): workflow.human_verify_mode = end-of-phase (new default; mid-flight opt-back-in) (#3325)

* test(3309): red — workflow.human_verify_mode contract

New behavioral test file covers:
- workflow.human_verify_mode is a recognized config key (VALID_CONFIG_KEYS)
- defaults to 'mid-flight' (preserves current behavior)
- config-set / config-get round-trips for both values
- persists in config.json as string
- planner agent file references the flag with canonical wording, couples
  end-of-phase mode with the rule that checkpoint:human-verify is not
  emitted, and documents the <verify><human-check> deferred-item shape
- verifier agent file references harvesting <verify><human-check> blocks
- references/checkpoints.md documents the cost-control alternative

Source-text assertions on agent .md files are exempted via
allow-test-rule: source-text-is-the-product — those files ARE the
runtime contract loaded by AI runtimes, so asserting their wording is
the only way to verify the agents will respect the flag.

Fails 10/11 against current source. Will pass after the fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(3309): add workflow.human_verify_mode = end-of-phase opt-out

Each mid-flight checkpoint:human-verify halt costs a full executor
cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on every
respawn) because subagent context is discarded across the pause. A plan
with N human-verify checkpoints pays the cold-start cost N+1 times. The
reporter (rentanything-nb) measured this at "tens of thousands of tokens"
per round-trip and "hundreds of thousands per week."

This adds workflow.human_verify_mode (default 'mid-flight') with an
'end-of-phase' value that:
- instructs gsd-planner to NOT emit <task type="checkpoint:human-verify">
  tasks; verification details go into a <verify><human-check> sub-block
  on the relevant auto task instead
- instructs gsd-verifier (Step 8) to harvest those <verify><human-check>
  blocks at end-of-phase and merge them into its own human-verification
  list
- the existing human_needed → HUMAN-UAT.md flow in execute-phase.md is
  the single sink — no new file/writer is created

checkpoint:decision and checkpoint:human-action are unaffected — those
gate the work itself, not post-hoc verification.

Surfaces touched:
- bin/lib/config-schema.cjs, bin/lib/config.cjs — register key + default
- sdk/src/config.ts, sdk/src/query/config-schema.ts — SDK parity
- agents/gsd-planner.md — slim Detection section + reference link
- agents/gsd-verifier.md — Step 8 harvest instruction
- get-shit-done/references/planner-human-verify-mode.md — full rules,
  loaded conditionally to keep planner.md under its size budget
- get-shit-done/references/checkpoints.md — surface the alternative
- docs/CONFIGURATION.md — config table row
- docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json — track new reference

Tag name <human-check> chosen instead of <human> to avoid the
prompt-injection scan pattern that flags <system|assistant|human> tags.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(3309): align changeset pr: to actual PR number

The pr: field was authored as 3319 (a guess at the next number) before
the PR was opened. Actual PR is #3325.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(3309): flip workflow.human_verify_mode default to end-of-phase

Per maintainer direction on PR #3325, end-of-phase is the new project
default. Mid-flight checkpoint:human-verify halts cost a full executor
cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per
round-trip — reported at "tens of thousands of tokens" per round-trip,
"hundreds of thousands per week" on real projects. The cost-control
mode is what new projects should get out of the box.

mid-flight remains a one-line opt-back-in via:

    gsd config-set workflow.human_verify_mode mid-flight

Behavior change for existing projects: the new default takes effect
when .planning/config.json is rewritten (config-set, fresh project).
Existing in-flight PLAN.md files with checkpoint:human-verify tasks
continue to work in either mode — the flag only changes what the
planner emits next time it runs.

Surfaces updated:
- bin/lib/config.cjs, sdk/src/config.ts — default flipped
- sdk/src/config.ts docstring — describes new default + opt-back-in
- agents/gsd-planner.md — Detection section explains new default
- references/planner-human-verify-mode.md — reordered modes; added
  guidance on when to opt back into mid-flight
- references/checkpoints.md — surface the default flip and the why
- docs/CONFIGURATION.md — table row reflects new default + reason
- tests/feat-3309-human-verify-mode.test.cjs — default test asserts
  end-of-phase
- .changeset/fierce-geese-march.md — describes the default flip and
  the migration semantics

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: address human verify mode review

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-05-09 23:29:11 -04:00
committed by GitHub
parent e7942c21b3
commit 25fb81d01e
13 changed files with 294 additions and 2 deletions

View File

@@ -0,0 +1,5 @@
---
type: Added
pr: 3325
---
**`workflow.human_verify_mode = end-of-phase` is now the default** — the planner no longer emits `<task type="checkpoint:human-verify">` tasks for new projects; verification details are embedded into `<verify><human-check>` blocks on `auto` tasks and the verifier consolidates them at end-of-phase into the existing HUMAN-UAT.md flow. The previous mid-flight behavior cost a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per `checkpoint:human-verify` round-trip — measured at "tens of thousands of tokens" per round-trip on real projects. Set `workflow.human_verify_mode = mid-flight` in `.planning/config.json` to restore the pre-#3309 behavior. `checkpoint:decision` and `checkpoint:human-action` are unaffected by either value. **Behavior change for existing projects:** the new default takes effect when `.planning/config.json` is rewritten (e.g. via `gsd config-set` or first run on a new GSD version). Existing in-flight PLAN.md files with `checkpoint:human-verify` tasks continue to work in either mode — the flag only changes what the planner emits next time it runs. (#3309)

View File

@@ -304,6 +304,8 @@ This prevents the "scavenger hunt" anti-pattern where executors explore the code
Exceptions where `tdd="true"` is not needed: `type="checkpoint:*"` tasks, configuration-only files, documentation, migration scripts, glue code wiring existing tested components, styling-only changes.
`workflow.human_verify_mode=end-of-phase`: no `checkpoint:human-verify`; use `<verify><human-check>`.
## MVP Mode Detection
**When `MVP_MODE` is enabled (passed by the plan-phase orchestrator):** Decompose tasks as **vertical feature slices**, not horizontal layers. Required reading: `@~/.claude/get-shit-done/references/planner-mvp-mode.md` (loaded conditionally by the orchestrator).

View File

@@ -491,6 +491,20 @@ npm test -- --grep "$PHASE_TEST_PATTERN" 2>&1 | grep -q "passing"
**Needs human if uncertain:** Complex wiring grep can't trace, dynamic state behavior, edge cases.
**Harvest deferred items from PLAN.md (#3309 / `workflow.human_verify_mode = end-of-phase`):** Scan every PLAN file in the phase for `<verify><human-check>` blocks on `auto` tasks. These are verification items the planner deliberately deferred from `checkpoint:human-verify` to end-of-phase to avoid the executor cold-start cost. Each block has the same shape used by the planner:
```xml
<verify>
<human-check>
<test>What to do</test>
<expected>What should happen</expected>
<why_human>Why grep can't verify</why_human>
</human-check>
</verify>
```
Merge those harvested items into the same human verification list as your own analysis. Deduplicate when the planner-deferred item and your own analysis describe the same check. The downstream `human_needed` → HUMAN-UAT.md path in `workflows/execute-phase.md` is the single sink — no separate file is created.
**Format:**
```markdown

View File

@@ -209,6 +209,7 @@ All workflow toggles follow the **absent = enabled** pattern. If a key is missin
| `workflow.plan_chunked` | boolean | `false` | Enable chunked planning mode. When `true` (or when `--chunked` flag is passed to `/gsd-plan-phase`), the orchestrator splits the single long-lived planner Task into a short outline Task followed by N short per-plan Tasks (~3-5 min each). Each plan is committed individually for crash resilience. If a Task hangs and the terminal is force-killed, rerunning with `--chunked` resumes from the last completed plan. Particularly useful on Windows where long-lived Tasks may hang on stdio. Added in v1.38 |
| `workflow.code_review_command` | string | (none) | Shell command for external code review integration in `/gsd-ship`. Receives changed file paths via stdin. Non-zero exit blocks the ship workflow. Added in v1.36 |
| `workflow.tdd_mode` | boolean | `false` | Enable TDD pipeline as a first-class execution mode. When `true`, the planner aggressively applies `type: tdd` to eligible tasks (business logic, APIs, validations, algorithms) and the executor enforces RED/GREEN/REFACTOR gate sequence. An end-of-phase collaborative review checkpoint verifies gate compliance. Added in v1.36 |
| `workflow.human_verify_mode` | string | `'end-of-phase'` | Controls human verification checkpoints. `'end-of-phase'` (default since #3309) suppresses `checkpoint:human-verify` tasks and embeds checks into `<verify><human-check>` blocks for end-of-phase review. `'mid-flight'` restores blocking checkpoint tasks. `checkpoint:decision` and `checkpoint:human-action` are unaffected. See [Checkpoints Reference](../get-shit-done/references/checkpoints.md#checkpoint_types). |
| `workflow.cross_ai_execution` | boolean | `false` | Delegate phase execution to an external AI CLI instead of spawning local executor agents. Useful for leveraging a different model's strengths for specific phases. Added in v1.36 |
| `workflow.cross_ai_command` | string | (none) | Shell command template for cross-AI execution. Receives the phase prompt via stdin. Must produce SUMMARY.md-compatible output. Required when `cross_ai_execution` is `true`. Added in v1.36 |
| `workflow.cross_ai_timeout` | number | `300` | Timeout in seconds for cross-AI execution commands. Prevents runaway external processes. Added in v1.36 |

View File

@@ -223,6 +223,7 @@
"planner-antipatterns.md",
"planner-chunked.md",
"planner-gap-closure.md",
"planner-human-verify-mode.md",
"planner-mvp-mode.md",
"planner-reviews.md",
"planner-revision.md",

View File

@@ -261,7 +261,7 @@ Full roster at `get-shit-done/workflows/*.md`. Workflows are thin orchestrators
---
## References (59 shipped)
## References (60 shipped)
Full roster at `get-shit-done/references/*.md`. References are shared knowledge documents that workflows and agents `@-reference`. The groupings below match [`docs/ARCHITECTURE.md`](ARCHITECTURE.md#references-get-shit-donereferencesmd) — core, workflow, thinking-model clusters, and the modular planner decomposition.
@@ -350,11 +350,12 @@ The `gsd-planner` agent is decomposed into a core agent plus reference modules t
| `planner-revision.md` | Plan revision patterns for iterative refinement. |
| `planner-source-audit.md` | Planner source-audit and authority-limit rules. |
| `planner-mvp-mode.md` | Vertical-slice planning rules for MVP mode. |
| `planner-human-verify-mode.md` | Rules for `workflow.human_verify_mode = end-of-phase`: suppress `checkpoint:human-verify` task emission and route deferred items via `<verify><human-check>`. |
| `skeleton-template.md` | SKELETON.md template emitted for new-project Walking Skeleton (Phase 1 + `--mvp`). |
| `user-story-template.md` | User story format for MVP planning — "As a / I want to / So that" structured fields. |
| `spidr-splitting.md` | SPIDR splitting decomposition rules for handling large user stories in MVP mode. |
> **Subdirectory:** `get-shit-done/references/few-shot-examples/` contains additional few-shot examples (`plan-checker.md`, `verifier.md`) that are referenced from specific agents. These are not counted in the 59 top-level references.
> **Subdirectory:** `get-shit-done/references/few-shot-examples/` contains additional few-shot examples (`plan-checker.md`, `verifier.md`) that are referenced from specific agents. These are not counted in the 60 top-level references.
---

View File

@@ -20,6 +20,7 @@ const VALID_CONFIG_KEYS = new Set([
'workflow.nyquist_validation', 'workflow.ai_integration_phase', 'workflow.ui_phase', 'workflow.ui_safety_gate',
'workflow.auto_advance', 'workflow.node_repair', 'workflow.node_repair_budget',
'workflow.tdd_mode',
'workflow.human_verify_mode',
'workflow.text_mode',
'workflow.research_before_questions',
'workflow.discuss_mode',

View File

@@ -109,6 +109,7 @@ function buildNewProjectConfig(userChoices) {
ui_safety_gate: true,
ai_integration_phase: true,
tdd_mode: false,
human_verify_mode: 'end-of-phase',
text_mode: false,
research_before_questions: false,
discuss_mode: 'discuss',
@@ -354,6 +355,12 @@ function cmdConfigSet(cwd, keyPath, value, raw) {
}
}
// Human verification checkpoint mode (#3309)
const VALID_HUMAN_VERIFY_MODES = ['mid-flight', 'end-of-phase'];
if (keyPath === 'workflow.human_verify_mode' && !VALID_HUMAN_VERIFY_MODES.includes(String(parsedValue))) {
error(`Invalid workflow.human_verify_mode '${value}'. Valid values: ${VALID_HUMAN_VERIFY_MODES.join(', ')}`);
}
const setConfigValueResult = setConfigValue(cwd, keyPath, parsedValue);
// Mask secrets in both JSON and text output. The plaintext is written

View File

@@ -18,6 +18,12 @@ Plans execute autonomously. Checkpoints formalize interaction points where human
**When:** Claude completed automated work, human confirms it works correctly.
> **Default mode (#3309): `workflow.human_verify_mode = end-of-phase`.** New projects do NOT halt mid-flight at `checkpoint:human-verify`. The planner suppresses those task emissions and embeds the verification details into the relevant `auto` task's `<verify><human-check>` block; the verifier harvests every `<verify><human-check>` at end-of-phase (Step 8) and consolidates them into the existing `human_needed` → HUMAN-UAT.md flow in `workflows/execute-phase.md`. The user reviews everything in one batch.
>
> **Why this is the default:** every mid-flight halt costs a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) because subagent context is discarded across the pause. A plan with N human-verify checkpoints pays the cold-start cost N+1 times — measured at "tens of thousands of tokens" per round-trip on real projects.
>
> Set `workflow.human_verify_mode = mid-flight` in `.planning/config.json` to opt back into the pre-#3309 behavior of halting at every checkpoint. `checkpoint:decision` and `checkpoint:human-action` are unaffected by either value — those gate the work itself, not post-hoc verification.
**Use for:**
- Visual UI checks (layout, styling, responsiveness)
- Interactive flows (click through wizard, test user flows)

View File

@@ -0,0 +1,57 @@
# Planner — Human Verification Mode
> Loaded by `gsd-planner` when deciding whether to emit `<task type="checkpoint:human-verify">` tasks. Read `workflow.human_verify_mode` from `.planning/config.json` (default `end-of-phase` since #3309).
## The two modes
### `end-of-phase` (default — issue #3309)
Do **not** emit any `<task type="checkpoint:human-verify">` tasks. Every mid-flight halt costs a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) because subagent context is discarded across the pause; a plan with N human-verify checkpoints pays the cold-start cost N+1 times — measured at "tens of thousands of tokens" per round-trip on real projects. This is the default for that reason.
Instead, fold each would-be verification step into the relevant `auto` task using a `<verify><human-check>` sub-block:
```xml
<task type="auto">
<name>Wire dashboard route</name>
<files>app/dashboard/page.tsx, app/api/dashboard/route.ts</files>
<action>...</action>
<verify>
<automated>npm test -- --filter=dashboard</automated>
<human-check>
<test>Visit http://localhost:3000/dashboard</test>
<expected>Sidebar left, content right on desktop &gt;1024px; collapses to hamburger at 768px</expected>
<why_human>Visual layout — grep cannot verify breakpoint behavior</why_human>
</human-check>
</verify>
<done>Layout renders correctly across breakpoints</done>
</task>
```
The verifier (Step 8) harvests every `<verify><human-check>` block at end-of-phase and consolidates them into the existing `human_needed` → HUMAN-UAT.md path in `workflows/execute-phase.md`. The user reviews everything in one batch instead of paying a cold-start cost per item.
### `mid-flight` (opt-back-in — pre-#3309 behavior)
Set `gsd config-set workflow.human_verify_mode mid-flight` to restore the canonical mid-flight pattern: emit `<task type="checkpoint:human-verify">` tasks at the points where human confirmation is required, and the executor halts at each one to ask the user.
```xml
<task type="checkpoint:human-verify" gate="blocking">
<what-built>Dev server running at http://localhost:3000</what-built>
<how-to-verify>
1. Visit /dashboard
2. Sidebar collapses at 768px
</how-to-verify>
<resume-signal>"approved" or describe issues</resume-signal>
</task>
```
Choose `mid-flight` when you genuinely need the work to stop before any subsequent task runs (e.g., the next task depends on visual confirmation of the previous one), and you accept the cold-start cost as the price of that hard barrier.
## What is *not* affected
`checkpoint:decision` and `checkpoint:human-action` tasks are still emitted in `end-of-phase` mode. Those gate the work itself (a choice the executor needs from the user, or an auth step only the user can perform), not post-hoc verification of completed work. Only `checkpoint:human-verify` is suppressed.
## Compatibility with other modes
- **`workflow.tdd_mode`**: orthogonal. TDD tasks still emit `tdd="true"` and `<behavior>`; the `<verify>` block carries the human-check sub-element when `human_verify_mode = end-of-phase`.
- **`MVP_MODE`**: orthogonal. Vertical-slice ordering is unchanged. The first task remains a failing end-to-end test; later auto tasks may carry `<verify><human-check>` instead of standalone checkpoint tasks.
- **`workflow.auto_advance` / `_auto_chain_active`**: in mid-flight mode these auto-approve checkpoint:human-verify halts. In end-of-phase mode there are no halts to auto-approve, so the flags have no effect on this code path.

View File

@@ -25,6 +25,18 @@ export interface WorkflowConfig {
nyquist_validation: boolean;
/** Mirrors gsd-tools flat `config.tdd_mode` (from `workflow.tdd_mode`). */
tdd_mode: boolean;
/**
* Issue #3309. `end-of-phase` (default) suppresses mid-flight
* `<task type="checkpoint:human-verify">` task emission; the planner
* embeds verification details into the relevant `auto` task's
* `<verify><human-check>` block and the verifier harvests them at
* end-of-phase into the existing HUMAN-UAT.md path. `mid-flight`
* restores the pre-#3309 behavior where the executor halts at each
* `checkpoint:human-verify` task and pays a full executor cold-start
* cost (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per
* round-trip.
*/
human_verify_mode: 'mid-flight' | 'end-of-phase';
auto_advance: boolean;
/** Internal auto-chain flag used by workflow routing. */
_auto_chain_active?: boolean;
@@ -94,6 +106,7 @@ export const CONFIG_DEFAULTS: GSDConfig = {
verifier: true,
nyquist_validation: true,
tdd_mode: false,
human_verify_mode: 'end-of-phase',
auto_advance: false,
node_repair: true,
node_repair_budget: 2,

View File

@@ -22,6 +22,7 @@ export const VALID_CONFIG_KEYS: ReadonlySet<string> = new Set([
'workflow.nyquist_validation', 'workflow.ai_integration_phase', 'workflow.ui_phase', 'workflow.ui_safety_gate',
'workflow.auto_advance', 'workflow.node_repair', 'workflow.node_repair_budget',
'workflow.tdd_mode',
'workflow.human_verify_mode',
'workflow.text_mode',
'workflow.research_before_questions',
'workflow.discuss_mode',

View File

@@ -0,0 +1,183 @@
// allow-test-rule: source-text-is-the-product
// Planner and verifier agent .md files ARE the runtime contract loaded by
// the AI runtimes. Asserting that the canonical wording for the new
// `workflow.human_verify_mode` flag is present in those files is the only
// way to verify the agents will respect the flag at runtime.
/**
* Enhancement #3309: workflow.human_verify_mode = end-of-phase
*
* "mid-flight" preserves the pre-#3309 behavior — the planner emits
* `<task type="checkpoint:human-verify">` tasks, and the executor halts at
* each one. Each halt costs a full executor cold-start (CLAUDE.md, MEMORY.md,
* STATE.md, plan re-read) because subagent context is discarded across the
* pause.
*
* "end-of-phase" (the new default) instructs the planner NOT to emit
* `checkpoint:human-verify` tasks and instead embed the verification details
* into the relevant `auto` task's `<verify><human-check>` block. The verifier
* (Step 8) harvests these blocks at end-of-phase and consolidates them into the existing
* `human_needed` → HUMAN-UAT.md path, restoring the v1.35-shaped behavior
* the reporter wanted without resurrecting the v1.35 writer.
*/
'use strict';
const { test, describe, beforeEach, afterEach } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('fs');
const path = require('path');
const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs');
function readConfig(tmpDir) {
const configPath = path.join(tmpDir, '.planning', 'config.json');
return JSON.parse(fs.readFileSync(configPath, 'utf-8'));
}
const REPO_ROOT = path.join(__dirname, '..');
// ─── Schema registration ──────────────────────────────────────────────────────
describe('workflow.human_verify_mode in VALID_CONFIG_KEYS', () => {
test('is a recognized config key', () => {
const { VALID_CONFIG_KEYS } = require('../get-shit-done/bin/lib/config.cjs');
assert.ok(
VALID_CONFIG_KEYS.has('workflow.human_verify_mode'),
'workflow.human_verify_mode should be in VALID_CONFIG_KEYS',
);
});
});
// ─── Default value (CJS) ──────────────────────────────────────────────────────
describe('workflow.human_verify_mode default value', () => {
let tmpDir;
beforeEach(() => { tmpDir = createTempProject(); });
afterEach(() => { cleanup(tmpDir); });
test('defaults to end-of-phase in new project config', () => {
const result = runGsdTools('config-ensure-section', tmpDir, { HOME: tmpDir });
assert.ok(result.success, `config-ensure-section failed: ${result.error}`);
const config = readConfig(tmpDir);
assert.strictEqual(
config.workflow.human_verify_mode,
'end-of-phase',
'workflow.human_verify_mode should default to "end-of-phase" — the cost-control mode is the project default; opt back into the pre-#3309 mid-flight behavior with config-set',
);
});
});
// ─── Round-trip ──────────────────────────────────────────────────────────────
describe('workflow.human_verify_mode config round-trip', () => {
let tmpDir;
beforeEach(() => {
tmpDir = createTempProject();
runGsdTools('config-ensure-section', tmpDir, { HOME: tmpDir });
});
afterEach(() => { cleanup(tmpDir); });
test('config-set end-of-phase persists to config.json', () => {
const setResult = runGsdTools('config-set workflow.human_verify_mode end-of-phase', tmpDir);
assert.ok(setResult.success, `config-set failed: ${setResult.error}`);
const config = readConfig(tmpDir);
assert.strictEqual(config.workflow.human_verify_mode, 'end-of-phase');
});
test('config-set mid-flight overwrites end-of-phase in config.json', () => {
runGsdTools('config-set workflow.human_verify_mode end-of-phase', tmpDir);
const setResult = runGsdTools('config-set workflow.human_verify_mode mid-flight', tmpDir);
assert.ok(setResult.success, `config-set failed: ${setResult.error}`);
const config = readConfig(tmpDir);
assert.strictEqual(config.workflow.human_verify_mode, 'mid-flight');
});
test('persists in config.json as string', () => {
runGsdTools('config-set workflow.human_verify_mode end-of-phase', tmpDir);
const config = readConfig(tmpDir);
assert.strictEqual(config.workflow.human_verify_mode, 'end-of-phase');
assert.strictEqual(typeof config.workflow.human_verify_mode, 'string');
});
test('rejects invalid mode values', () => {
const result = runGsdTools('config-set workflow.human_verify_mode midflight', tmpDir);
assert.strictEqual(result.success, false);
assert.match(result.error, /Invalid workflow\.human_verify_mode 'midflight'/);
assert.match(result.error, /mid-flight, end-of-phase/);
});
});
// ─── Planner agent contract ──────────────────────────────────────────────────
describe('agents/gsd-planner.md acknowledges workflow.human_verify_mode', () => {
let plannerSrc;
test('loads', () => {
plannerSrc = fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-planner.md'), 'utf-8');
assert.ok(plannerSrc.length > 0);
});
test('mentions workflow.human_verify_mode by canonical name', () => {
plannerSrc = plannerSrc || fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-planner.md'), 'utf-8');
assert.ok(
plannerSrc.includes('workflow.human_verify_mode'),
'planner must reference the flag by canonical key so the runtime can resolve config-driven behavior',
);
});
test('explains the end-of-phase behavior (do NOT emit checkpoint:human-verify)', () => {
plannerSrc = plannerSrc || fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-planner.md'), 'utf-8');
// The planner must instruct: when end-of-phase, do NOT emit checkpoint:human-verify
assert.ok(
/end-of-phase[\s\S]{0,400}checkpoint:human-verify/i.test(plannerSrc) ||
/checkpoint:human-verify[\s\S]{0,400}end-of-phase/i.test(plannerSrc),
'planner must couple "end-of-phase" mode with the rule that checkpoint:human-verify tasks are not emitted',
);
});
test('routes deferred verification through the <verify><human-check> block on auto tasks', () => {
plannerSrc = plannerSrc || fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-planner.md'), 'utf-8');
assert.ok(
/`?<verify>`?\s*[\s\S]{0,200}`?<human-check>`?/i.test(plannerSrc) ||
plannerSrc.includes('<verify><human-check>') ||
plannerSrc.includes('`<verify><human-check>`'),
'planner must document the <verify><human-check>...</human-check></verify> shape so the verifier can harvest deferred items',
);
});
});
// ─── Verifier agent contract ─────────────────────────────────────────────────
describe('agents/gsd-verifier.md harvests deferred human verification items', () => {
test('Step 8 mentions harvesting <verify><human-check> blocks from PLAN.md', () => {
const verifierSrc = fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-verifier.md'), 'utf-8');
assert.ok(
verifierSrc.includes('<verify><human-check>') || /<verify>[\s\S]{0,200}<human-check>/i.test(verifierSrc),
'verifier must instruct itself to harvest <verify><human-check> blocks from PLAN.md when human_verify_mode = end-of-phase',
);
assert.ok(
verifierSrc.includes('human_verify_mode'),
'verifier must reference the flag by canonical key',
);
});
});
// ─── References doc parity ───────────────────────────────────────────────────
describe('references/checkpoints.md documents the flag', () => {
test('mentions workflow.human_verify_mode in the human-verify section', () => {
const refSrc = fs.readFileSync(
path.join(REPO_ROOT, 'get-shit-done', 'references', 'checkpoints.md'),
'utf-8',
);
assert.ok(
refSrc.includes('workflow.human_verify_mode'),
'checkpoints reference must document the new flag so users know the cost-control alternative exists',
);
});
});