feat(3309): workflow.human_verify_mode = end-of-phase (new default; mid-flight opt-back-in) (#3325)
* test(3309): red — workflow.human_verify_mode contract New behavioral test file covers: - workflow.human_verify_mode is a recognized config key (VALID_CONFIG_KEYS) - defaults to 'mid-flight' (preserves current behavior) - config-set / config-get round-trips for both values - persists in config.json as string - planner agent file references the flag with canonical wording, couples end-of-phase mode with the rule that checkpoint:human-verify is not emitted, and documents the <verify><human-check> deferred-item shape - verifier agent file references harvesting <verify><human-check> blocks - references/checkpoints.md documents the cost-control alternative Source-text assertions on agent .md files are exempted via allow-test-rule: source-text-is-the-product — those files ARE the runtime contract loaded by AI runtimes, so asserting their wording is the only way to verify the agents will respect the flag. Fails 10/11 against current source. Will pass after the fix. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(3309): add workflow.human_verify_mode = end-of-phase opt-out Each mid-flight checkpoint:human-verify halt costs a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on every respawn) because subagent context is discarded across the pause. A plan with N human-verify checkpoints pays the cold-start cost N+1 times. The reporter (rentanything-nb) measured this at "tens of thousands of tokens" per round-trip and "hundreds of thousands per week." This adds workflow.human_verify_mode (default 'mid-flight') with an 'end-of-phase' value that: - instructs gsd-planner to NOT emit <task type="checkpoint:human-verify"> tasks; verification details go into a <verify><human-check> sub-block on the relevant auto task instead - instructs gsd-verifier (Step 8) to harvest those <verify><human-check> blocks at end-of-phase and merge them into its own human-verification list - the existing human_needed → HUMAN-UAT.md flow in execute-phase.md is the single sink — no new file/writer is created checkpoint:decision and checkpoint:human-action are unaffected — those gate the work itself, not post-hoc verification. Surfaces touched: - bin/lib/config-schema.cjs, bin/lib/config.cjs — register key + default - sdk/src/config.ts, sdk/src/query/config-schema.ts — SDK parity - agents/gsd-planner.md — slim Detection section + reference link - agents/gsd-verifier.md — Step 8 harvest instruction - get-shit-done/references/planner-human-verify-mode.md — full rules, loaded conditionally to keep planner.md under its size budget - get-shit-done/references/checkpoints.md — surface the alternative - docs/CONFIGURATION.md — config table row - docs/INVENTORY.md, docs/INVENTORY-MANIFEST.json — track new reference Tag name <human-check> chosen instead of <human> to avoid the prompt-injection scan pattern that flags <system|assistant|human> tags. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(3309): align changeset pr: to actual PR number The pr: field was authored as 3319 (a guess at the next number) before the PR was opened. Actual PR is #3325. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(3309): flip workflow.human_verify_mode default to end-of-phase Per maintainer direction on PR #3325, end-of-phase is the new project default. Mid-flight checkpoint:human-verify halts cost a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per round-trip — reported at "tens of thousands of tokens" per round-trip, "hundreds of thousands per week" on real projects. The cost-control mode is what new projects should get out of the box. mid-flight remains a one-line opt-back-in via: gsd config-set workflow.human_verify_mode mid-flight Behavior change for existing projects: the new default takes effect when .planning/config.json is rewritten (config-set, fresh project). Existing in-flight PLAN.md files with checkpoint:human-verify tasks continue to work in either mode — the flag only changes what the planner emits next time it runs. Surfaces updated: - bin/lib/config.cjs, sdk/src/config.ts — default flipped - sdk/src/config.ts docstring — describes new default + opt-back-in - agents/gsd-planner.md — Detection section explains new default - references/planner-human-verify-mode.md — reordered modes; added guidance on when to opt back into mid-flight - references/checkpoints.md — surface the default flip and the why - docs/CONFIGURATION.md — table row reflects new default + reason - tests/feat-3309-human-verify-mode.test.cjs — default test asserts end-of-phase - .changeset/fierce-geese-march.md — describes the default flip and the migration semantics Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: address human verify mode review --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
5
.changeset/fierce-geese-march.md
Normal file
5
.changeset/fierce-geese-march.md
Normal file
@@ -0,0 +1,5 @@
|
||||
---
|
||||
type: Added
|
||||
pr: 3325
|
||||
---
|
||||
**`workflow.human_verify_mode = end-of-phase` is now the default** — the planner no longer emits `<task type="checkpoint:human-verify">` tasks for new projects; verification details are embedded into `<verify><human-check>` blocks on `auto` tasks and the verifier consolidates them at end-of-phase into the existing HUMAN-UAT.md flow. The previous mid-flight behavior cost a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per `checkpoint:human-verify` round-trip — measured at "tens of thousands of tokens" per round-trip on real projects. Set `workflow.human_verify_mode = mid-flight` in `.planning/config.json` to restore the pre-#3309 behavior. `checkpoint:decision` and `checkpoint:human-action` are unaffected by either value. **Behavior change for existing projects:** the new default takes effect when `.planning/config.json` is rewritten (e.g. via `gsd config-set` or first run on a new GSD version). Existing in-flight PLAN.md files with `checkpoint:human-verify` tasks continue to work in either mode — the flag only changes what the planner emits next time it runs. (#3309)
|
||||
@@ -304,6 +304,8 @@ This prevents the "scavenger hunt" anti-pattern where executors explore the code
|
||||
|
||||
Exceptions where `tdd="true"` is not needed: `type="checkpoint:*"` tasks, configuration-only files, documentation, migration scripts, glue code wiring existing tested components, styling-only changes.
|
||||
|
||||
`workflow.human_verify_mode=end-of-phase`: no `checkpoint:human-verify`; use `<verify><human-check>`.
|
||||
|
||||
## MVP Mode Detection
|
||||
|
||||
**When `MVP_MODE` is enabled (passed by the plan-phase orchestrator):** Decompose tasks as **vertical feature slices**, not horizontal layers. Required reading: `@~/.claude/get-shit-done/references/planner-mvp-mode.md` (loaded conditionally by the orchestrator).
|
||||
|
||||
@@ -491,6 +491,20 @@ npm test -- --grep "$PHASE_TEST_PATTERN" 2>&1 | grep -q "passing"
|
||||
|
||||
**Needs human if uncertain:** Complex wiring grep can't trace, dynamic state behavior, edge cases.
|
||||
|
||||
**Harvest deferred items from PLAN.md (#3309 / `workflow.human_verify_mode = end-of-phase`):** Scan every PLAN file in the phase for `<verify><human-check>` blocks on `auto` tasks. These are verification items the planner deliberately deferred from `checkpoint:human-verify` to end-of-phase to avoid the executor cold-start cost. Each block has the same shape used by the planner:
|
||||
|
||||
```xml
|
||||
<verify>
|
||||
<human-check>
|
||||
<test>What to do</test>
|
||||
<expected>What should happen</expected>
|
||||
<why_human>Why grep can't verify</why_human>
|
||||
</human-check>
|
||||
</verify>
|
||||
```
|
||||
|
||||
Merge those harvested items into the same human verification list as your own analysis. Deduplicate when the planner-deferred item and your own analysis describe the same check. The downstream `human_needed` → HUMAN-UAT.md path in `workflows/execute-phase.md` is the single sink — no separate file is created.
|
||||
|
||||
**Format:**
|
||||
|
||||
```markdown
|
||||
|
||||
@@ -209,6 +209,7 @@ All workflow toggles follow the **absent = enabled** pattern. If a key is missin
|
||||
| `workflow.plan_chunked` | boolean | `false` | Enable chunked planning mode. When `true` (or when `--chunked` flag is passed to `/gsd-plan-phase`), the orchestrator splits the single long-lived planner Task into a short outline Task followed by N short per-plan Tasks (~3-5 min each). Each plan is committed individually for crash resilience. If a Task hangs and the terminal is force-killed, rerunning with `--chunked` resumes from the last completed plan. Particularly useful on Windows where long-lived Tasks may hang on stdio. Added in v1.38 |
|
||||
| `workflow.code_review_command` | string | (none) | Shell command for external code review integration in `/gsd-ship`. Receives changed file paths via stdin. Non-zero exit blocks the ship workflow. Added in v1.36 |
|
||||
| `workflow.tdd_mode` | boolean | `false` | Enable TDD pipeline as a first-class execution mode. When `true`, the planner aggressively applies `type: tdd` to eligible tasks (business logic, APIs, validations, algorithms) and the executor enforces RED/GREEN/REFACTOR gate sequence. An end-of-phase collaborative review checkpoint verifies gate compliance. Added in v1.36 |
|
||||
| `workflow.human_verify_mode` | string | `'end-of-phase'` | Controls human verification checkpoints. `'end-of-phase'` (default since #3309) suppresses `checkpoint:human-verify` tasks and embeds checks into `<verify><human-check>` blocks for end-of-phase review. `'mid-flight'` restores blocking checkpoint tasks. `checkpoint:decision` and `checkpoint:human-action` are unaffected. See [Checkpoints Reference](../get-shit-done/references/checkpoints.md#checkpoint_types). |
|
||||
| `workflow.cross_ai_execution` | boolean | `false` | Delegate phase execution to an external AI CLI instead of spawning local executor agents. Useful for leveraging a different model's strengths for specific phases. Added in v1.36 |
|
||||
| `workflow.cross_ai_command` | string | (none) | Shell command template for cross-AI execution. Receives the phase prompt via stdin. Must produce SUMMARY.md-compatible output. Required when `cross_ai_execution` is `true`. Added in v1.36 |
|
||||
| `workflow.cross_ai_timeout` | number | `300` | Timeout in seconds for cross-AI execution commands. Prevents runaway external processes. Added in v1.36 |
|
||||
|
||||
@@ -223,6 +223,7 @@
|
||||
"planner-antipatterns.md",
|
||||
"planner-chunked.md",
|
||||
"planner-gap-closure.md",
|
||||
"planner-human-verify-mode.md",
|
||||
"planner-mvp-mode.md",
|
||||
"planner-reviews.md",
|
||||
"planner-revision.md",
|
||||
|
||||
@@ -261,7 +261,7 @@ Full roster at `get-shit-done/workflows/*.md`. Workflows are thin orchestrators
|
||||
|
||||
---
|
||||
|
||||
## References (59 shipped)
|
||||
## References (60 shipped)
|
||||
|
||||
Full roster at `get-shit-done/references/*.md`. References are shared knowledge documents that workflows and agents `@-reference`. The groupings below match [`docs/ARCHITECTURE.md`](ARCHITECTURE.md#references-get-shit-donereferencesmd) — core, workflow, thinking-model clusters, and the modular planner decomposition.
|
||||
|
||||
@@ -350,11 +350,12 @@ The `gsd-planner` agent is decomposed into a core agent plus reference modules t
|
||||
| `planner-revision.md` | Plan revision patterns for iterative refinement. |
|
||||
| `planner-source-audit.md` | Planner source-audit and authority-limit rules. |
|
||||
| `planner-mvp-mode.md` | Vertical-slice planning rules for MVP mode. |
|
||||
| `planner-human-verify-mode.md` | Rules for `workflow.human_verify_mode = end-of-phase`: suppress `checkpoint:human-verify` task emission and route deferred items via `<verify><human-check>`. |
|
||||
| `skeleton-template.md` | SKELETON.md template emitted for new-project Walking Skeleton (Phase 1 + `--mvp`). |
|
||||
| `user-story-template.md` | User story format for MVP planning — "As a / I want to / So that" structured fields. |
|
||||
| `spidr-splitting.md` | SPIDR splitting decomposition rules for handling large user stories in MVP mode. |
|
||||
|
||||
> **Subdirectory:** `get-shit-done/references/few-shot-examples/` contains additional few-shot examples (`plan-checker.md`, `verifier.md`) that are referenced from specific agents. These are not counted in the 59 top-level references.
|
||||
> **Subdirectory:** `get-shit-done/references/few-shot-examples/` contains additional few-shot examples (`plan-checker.md`, `verifier.md`) that are referenced from specific agents. These are not counted in the 60 top-level references.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -20,6 +20,7 @@ const VALID_CONFIG_KEYS = new Set([
|
||||
'workflow.nyquist_validation', 'workflow.ai_integration_phase', 'workflow.ui_phase', 'workflow.ui_safety_gate',
|
||||
'workflow.auto_advance', 'workflow.node_repair', 'workflow.node_repair_budget',
|
||||
'workflow.tdd_mode',
|
||||
'workflow.human_verify_mode',
|
||||
'workflow.text_mode',
|
||||
'workflow.research_before_questions',
|
||||
'workflow.discuss_mode',
|
||||
|
||||
@@ -109,6 +109,7 @@ function buildNewProjectConfig(userChoices) {
|
||||
ui_safety_gate: true,
|
||||
ai_integration_phase: true,
|
||||
tdd_mode: false,
|
||||
human_verify_mode: 'end-of-phase',
|
||||
text_mode: false,
|
||||
research_before_questions: false,
|
||||
discuss_mode: 'discuss',
|
||||
@@ -354,6 +355,12 @@ function cmdConfigSet(cwd, keyPath, value, raw) {
|
||||
}
|
||||
}
|
||||
|
||||
// Human verification checkpoint mode (#3309)
|
||||
const VALID_HUMAN_VERIFY_MODES = ['mid-flight', 'end-of-phase'];
|
||||
if (keyPath === 'workflow.human_verify_mode' && !VALID_HUMAN_VERIFY_MODES.includes(String(parsedValue))) {
|
||||
error(`Invalid workflow.human_verify_mode '${value}'. Valid values: ${VALID_HUMAN_VERIFY_MODES.join(', ')}`);
|
||||
}
|
||||
|
||||
const setConfigValueResult = setConfigValue(cwd, keyPath, parsedValue);
|
||||
|
||||
// Mask secrets in both JSON and text output. The plaintext is written
|
||||
|
||||
@@ -18,6 +18,12 @@ Plans execute autonomously. Checkpoints formalize interaction points where human
|
||||
|
||||
**When:** Claude completed automated work, human confirms it works correctly.
|
||||
|
||||
> **Default mode (#3309): `workflow.human_verify_mode = end-of-phase`.** New projects do NOT halt mid-flight at `checkpoint:human-verify`. The planner suppresses those task emissions and embeds the verification details into the relevant `auto` task's `<verify><human-check>` block; the verifier harvests every `<verify><human-check>` at end-of-phase (Step 8) and consolidates them into the existing `human_needed` → HUMAN-UAT.md flow in `workflows/execute-phase.md`. The user reviews everything in one batch.
|
||||
>
|
||||
> **Why this is the default:** every mid-flight halt costs a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) because subagent context is discarded across the pause. A plan with N human-verify checkpoints pays the cold-start cost N+1 times — measured at "tens of thousands of tokens" per round-trip on real projects.
|
||||
>
|
||||
> Set `workflow.human_verify_mode = mid-flight` in `.planning/config.json` to opt back into the pre-#3309 behavior of halting at every checkpoint. `checkpoint:decision` and `checkpoint:human-action` are unaffected by either value — those gate the work itself, not post-hoc verification.
|
||||
|
||||
**Use for:**
|
||||
- Visual UI checks (layout, styling, responsiveness)
|
||||
- Interactive flows (click through wizard, test user flows)
|
||||
|
||||
57
get-shit-done/references/planner-human-verify-mode.md
Normal file
57
get-shit-done/references/planner-human-verify-mode.md
Normal file
@@ -0,0 +1,57 @@
|
||||
# Planner — Human Verification Mode
|
||||
|
||||
> Loaded by `gsd-planner` when deciding whether to emit `<task type="checkpoint:human-verify">` tasks. Read `workflow.human_verify_mode` from `.planning/config.json` (default `end-of-phase` since #3309).
|
||||
|
||||
## The two modes
|
||||
|
||||
### `end-of-phase` (default — issue #3309)
|
||||
|
||||
Do **not** emit any `<task type="checkpoint:human-verify">` tasks. Every mid-flight halt costs a full executor cold-start (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) because subagent context is discarded across the pause; a plan with N human-verify checkpoints pays the cold-start cost N+1 times — measured at "tens of thousands of tokens" per round-trip on real projects. This is the default for that reason.
|
||||
|
||||
Instead, fold each would-be verification step into the relevant `auto` task using a `<verify><human-check>` sub-block:
|
||||
|
||||
```xml
|
||||
<task type="auto">
|
||||
<name>Wire dashboard route</name>
|
||||
<files>app/dashboard/page.tsx, app/api/dashboard/route.ts</files>
|
||||
<action>...</action>
|
||||
<verify>
|
||||
<automated>npm test -- --filter=dashboard</automated>
|
||||
<human-check>
|
||||
<test>Visit http://localhost:3000/dashboard</test>
|
||||
<expected>Sidebar left, content right on desktop >1024px; collapses to hamburger at 768px</expected>
|
||||
<why_human>Visual layout — grep cannot verify breakpoint behavior</why_human>
|
||||
</human-check>
|
||||
</verify>
|
||||
<done>Layout renders correctly across breakpoints</done>
|
||||
</task>
|
||||
```
|
||||
|
||||
The verifier (Step 8) harvests every `<verify><human-check>` block at end-of-phase and consolidates them into the existing `human_needed` → HUMAN-UAT.md path in `workflows/execute-phase.md`. The user reviews everything in one batch instead of paying a cold-start cost per item.
|
||||
|
||||
### `mid-flight` (opt-back-in — pre-#3309 behavior)
|
||||
|
||||
Set `gsd config-set workflow.human_verify_mode mid-flight` to restore the canonical mid-flight pattern: emit `<task type="checkpoint:human-verify">` tasks at the points where human confirmation is required, and the executor halts at each one to ask the user.
|
||||
|
||||
```xml
|
||||
<task type="checkpoint:human-verify" gate="blocking">
|
||||
<what-built>Dev server running at http://localhost:3000</what-built>
|
||||
<how-to-verify>
|
||||
1. Visit /dashboard
|
||||
2. Sidebar collapses at 768px
|
||||
</how-to-verify>
|
||||
<resume-signal>"approved" or describe issues</resume-signal>
|
||||
</task>
|
||||
```
|
||||
|
||||
Choose `mid-flight` when you genuinely need the work to stop before any subsequent task runs (e.g., the next task depends on visual confirmation of the previous one), and you accept the cold-start cost as the price of that hard barrier.
|
||||
|
||||
## What is *not* affected
|
||||
|
||||
`checkpoint:decision` and `checkpoint:human-action` tasks are still emitted in `end-of-phase` mode. Those gate the work itself (a choice the executor needs from the user, or an auth step only the user can perform), not post-hoc verification of completed work. Only `checkpoint:human-verify` is suppressed.
|
||||
|
||||
## Compatibility with other modes
|
||||
|
||||
- **`workflow.tdd_mode`**: orthogonal. TDD tasks still emit `tdd="true"` and `<behavior>`; the `<verify>` block carries the human-check sub-element when `human_verify_mode = end-of-phase`.
|
||||
- **`MVP_MODE`**: orthogonal. Vertical-slice ordering is unchanged. The first task remains a failing end-to-end test; later auto tasks may carry `<verify><human-check>` instead of standalone checkpoint tasks.
|
||||
- **`workflow.auto_advance` / `_auto_chain_active`**: in mid-flight mode these auto-approve checkpoint:human-verify halts. In end-of-phase mode there are no halts to auto-approve, so the flags have no effect on this code path.
|
||||
@@ -25,6 +25,18 @@ export interface WorkflowConfig {
|
||||
nyquist_validation: boolean;
|
||||
/** Mirrors gsd-tools flat `config.tdd_mode` (from `workflow.tdd_mode`). */
|
||||
tdd_mode: boolean;
|
||||
/**
|
||||
* Issue #3309. `end-of-phase` (default) suppresses mid-flight
|
||||
* `<task type="checkpoint:human-verify">` task emission; the planner
|
||||
* embeds verification details into the relevant `auto` task's
|
||||
* `<verify><human-check>` block and the verifier harvests them at
|
||||
* end-of-phase into the existing HUMAN-UAT.md path. `mid-flight`
|
||||
* restores the pre-#3309 behavior where the executor halts at each
|
||||
* `checkpoint:human-verify` task and pays a full executor cold-start
|
||||
* cost (CLAUDE.md, MEMORY.md, STATE.md, plan re-read on respawn) per
|
||||
* round-trip.
|
||||
*/
|
||||
human_verify_mode: 'mid-flight' | 'end-of-phase';
|
||||
auto_advance: boolean;
|
||||
/** Internal auto-chain flag used by workflow routing. */
|
||||
_auto_chain_active?: boolean;
|
||||
@@ -94,6 +106,7 @@ export const CONFIG_DEFAULTS: GSDConfig = {
|
||||
verifier: true,
|
||||
nyquist_validation: true,
|
||||
tdd_mode: false,
|
||||
human_verify_mode: 'end-of-phase',
|
||||
auto_advance: false,
|
||||
node_repair: true,
|
||||
node_repair_budget: 2,
|
||||
|
||||
@@ -22,6 +22,7 @@ export const VALID_CONFIG_KEYS: ReadonlySet<string> = new Set([
|
||||
'workflow.nyquist_validation', 'workflow.ai_integration_phase', 'workflow.ui_phase', 'workflow.ui_safety_gate',
|
||||
'workflow.auto_advance', 'workflow.node_repair', 'workflow.node_repair_budget',
|
||||
'workflow.tdd_mode',
|
||||
'workflow.human_verify_mode',
|
||||
'workflow.text_mode',
|
||||
'workflow.research_before_questions',
|
||||
'workflow.discuss_mode',
|
||||
|
||||
183
tests/feat-3309-human-verify-mode.test.cjs
Normal file
183
tests/feat-3309-human-verify-mode.test.cjs
Normal file
@@ -0,0 +1,183 @@
|
||||
// allow-test-rule: source-text-is-the-product
|
||||
// Planner and verifier agent .md files ARE the runtime contract loaded by
|
||||
// the AI runtimes. Asserting that the canonical wording for the new
|
||||
// `workflow.human_verify_mode` flag is present in those files is the only
|
||||
// way to verify the agents will respect the flag at runtime.
|
||||
|
||||
/**
|
||||
* Enhancement #3309: workflow.human_verify_mode = end-of-phase
|
||||
*
|
||||
* "mid-flight" preserves the pre-#3309 behavior — the planner emits
|
||||
* `<task type="checkpoint:human-verify">` tasks, and the executor halts at
|
||||
* each one. Each halt costs a full executor cold-start (CLAUDE.md, MEMORY.md,
|
||||
* STATE.md, plan re-read) because subagent context is discarded across the
|
||||
* pause.
|
||||
*
|
||||
* "end-of-phase" (the new default) instructs the planner NOT to emit
|
||||
* `checkpoint:human-verify` tasks and instead embed the verification details
|
||||
* into the relevant `auto` task's `<verify><human-check>` block. The verifier
|
||||
* (Step 8) harvests these blocks at end-of-phase and consolidates them into the existing
|
||||
* `human_needed` → HUMAN-UAT.md path, restoring the v1.35-shaped behavior
|
||||
* the reporter wanted without resurrecting the v1.35 writer.
|
||||
*/
|
||||
|
||||
'use strict';
|
||||
|
||||
const { test, describe, beforeEach, afterEach } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs');
|
||||
|
||||
function readConfig(tmpDir) {
|
||||
const configPath = path.join(tmpDir, '.planning', 'config.json');
|
||||
return JSON.parse(fs.readFileSync(configPath, 'utf-8'));
|
||||
}
|
||||
|
||||
const REPO_ROOT = path.join(__dirname, '..');
|
||||
|
||||
// ─── Schema registration ──────────────────────────────────────────────────────
|
||||
|
||||
describe('workflow.human_verify_mode in VALID_CONFIG_KEYS', () => {
|
||||
test('is a recognized config key', () => {
|
||||
const { VALID_CONFIG_KEYS } = require('../get-shit-done/bin/lib/config.cjs');
|
||||
assert.ok(
|
||||
VALID_CONFIG_KEYS.has('workflow.human_verify_mode'),
|
||||
'workflow.human_verify_mode should be in VALID_CONFIG_KEYS',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Default value (CJS) ──────────────────────────────────────────────────────
|
||||
|
||||
describe('workflow.human_verify_mode default value', () => {
|
||||
let tmpDir;
|
||||
beforeEach(() => { tmpDir = createTempProject(); });
|
||||
afterEach(() => { cleanup(tmpDir); });
|
||||
|
||||
test('defaults to end-of-phase in new project config', () => {
|
||||
const result = runGsdTools('config-ensure-section', tmpDir, { HOME: tmpDir });
|
||||
assert.ok(result.success, `config-ensure-section failed: ${result.error}`);
|
||||
|
||||
const config = readConfig(tmpDir);
|
||||
assert.strictEqual(
|
||||
config.workflow.human_verify_mode,
|
||||
'end-of-phase',
|
||||
'workflow.human_verify_mode should default to "end-of-phase" — the cost-control mode is the project default; opt back into the pre-#3309 mid-flight behavior with config-set',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Round-trip ──────────────────────────────────────────────────────────────
|
||||
|
||||
describe('workflow.human_verify_mode config round-trip', () => {
|
||||
let tmpDir;
|
||||
beforeEach(() => {
|
||||
tmpDir = createTempProject();
|
||||
runGsdTools('config-ensure-section', tmpDir, { HOME: tmpDir });
|
||||
});
|
||||
afterEach(() => { cleanup(tmpDir); });
|
||||
|
||||
test('config-set end-of-phase persists to config.json', () => {
|
||||
const setResult = runGsdTools('config-set workflow.human_verify_mode end-of-phase', tmpDir);
|
||||
assert.ok(setResult.success, `config-set failed: ${setResult.error}`);
|
||||
|
||||
const config = readConfig(tmpDir);
|
||||
assert.strictEqual(config.workflow.human_verify_mode, 'end-of-phase');
|
||||
});
|
||||
|
||||
test('config-set mid-flight overwrites end-of-phase in config.json', () => {
|
||||
runGsdTools('config-set workflow.human_verify_mode end-of-phase', tmpDir);
|
||||
|
||||
const setResult = runGsdTools('config-set workflow.human_verify_mode mid-flight', tmpDir);
|
||||
assert.ok(setResult.success, `config-set failed: ${setResult.error}`);
|
||||
|
||||
const config = readConfig(tmpDir);
|
||||
assert.strictEqual(config.workflow.human_verify_mode, 'mid-flight');
|
||||
});
|
||||
|
||||
test('persists in config.json as string', () => {
|
||||
runGsdTools('config-set workflow.human_verify_mode end-of-phase', tmpDir);
|
||||
|
||||
const config = readConfig(tmpDir);
|
||||
assert.strictEqual(config.workflow.human_verify_mode, 'end-of-phase');
|
||||
assert.strictEqual(typeof config.workflow.human_verify_mode, 'string');
|
||||
});
|
||||
|
||||
test('rejects invalid mode values', () => {
|
||||
const result = runGsdTools('config-set workflow.human_verify_mode midflight', tmpDir);
|
||||
assert.strictEqual(result.success, false);
|
||||
assert.match(result.error, /Invalid workflow\.human_verify_mode 'midflight'/);
|
||||
assert.match(result.error, /mid-flight, end-of-phase/);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Planner agent contract ──────────────────────────────────────────────────
|
||||
|
||||
describe('agents/gsd-planner.md acknowledges workflow.human_verify_mode', () => {
|
||||
let plannerSrc;
|
||||
|
||||
test('loads', () => {
|
||||
plannerSrc = fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-planner.md'), 'utf-8');
|
||||
assert.ok(plannerSrc.length > 0);
|
||||
});
|
||||
|
||||
test('mentions workflow.human_verify_mode by canonical name', () => {
|
||||
plannerSrc = plannerSrc || fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-planner.md'), 'utf-8');
|
||||
assert.ok(
|
||||
plannerSrc.includes('workflow.human_verify_mode'),
|
||||
'planner must reference the flag by canonical key so the runtime can resolve config-driven behavior',
|
||||
);
|
||||
});
|
||||
|
||||
test('explains the end-of-phase behavior (do NOT emit checkpoint:human-verify)', () => {
|
||||
plannerSrc = plannerSrc || fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-planner.md'), 'utf-8');
|
||||
// The planner must instruct: when end-of-phase, do NOT emit checkpoint:human-verify
|
||||
assert.ok(
|
||||
/end-of-phase[\s\S]{0,400}checkpoint:human-verify/i.test(plannerSrc) ||
|
||||
/checkpoint:human-verify[\s\S]{0,400}end-of-phase/i.test(plannerSrc),
|
||||
'planner must couple "end-of-phase" mode with the rule that checkpoint:human-verify tasks are not emitted',
|
||||
);
|
||||
});
|
||||
|
||||
test('routes deferred verification through the <verify><human-check> block on auto tasks', () => {
|
||||
plannerSrc = plannerSrc || fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-planner.md'), 'utf-8');
|
||||
assert.ok(
|
||||
/`?<verify>`?\s*[\s\S]{0,200}`?<human-check>`?/i.test(plannerSrc) ||
|
||||
plannerSrc.includes('<verify><human-check>') ||
|
||||
plannerSrc.includes('`<verify><human-check>`'),
|
||||
'planner must document the <verify><human-check>...</human-check></verify> shape so the verifier can harvest deferred items',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Verifier agent contract ─────────────────────────────────────────────────
|
||||
|
||||
describe('agents/gsd-verifier.md harvests deferred human verification items', () => {
|
||||
test('Step 8 mentions harvesting <verify><human-check> blocks from PLAN.md', () => {
|
||||
const verifierSrc = fs.readFileSync(path.join(REPO_ROOT, 'agents', 'gsd-verifier.md'), 'utf-8');
|
||||
assert.ok(
|
||||
verifierSrc.includes('<verify><human-check>') || /<verify>[\s\S]{0,200}<human-check>/i.test(verifierSrc),
|
||||
'verifier must instruct itself to harvest <verify><human-check> blocks from PLAN.md when human_verify_mode = end-of-phase',
|
||||
);
|
||||
assert.ok(
|
||||
verifierSrc.includes('human_verify_mode'),
|
||||
'verifier must reference the flag by canonical key',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── References doc parity ───────────────────────────────────────────────────
|
||||
|
||||
describe('references/checkpoints.md documents the flag', () => {
|
||||
test('mentions workflow.human_verify_mode in the human-verify section', () => {
|
||||
const refSrc = fs.readFileSync(
|
||||
path.join(REPO_ROOT, 'get-shit-done', 'references', 'checkpoints.md'),
|
||||
'utf-8',
|
||||
);
|
||||
assert.ok(
|
||||
refSrc.includes('workflow.human_verify_mode'),
|
||||
'checkpoints reference must document the new flag so users know the cost-control alternative exists',
|
||||
);
|
||||
});
|
||||
});
|
||||
Reference in New Issue
Block a user