* enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select Auto-mode used to auto-select a checkpoint:decision's first <option> unconditionally, making a decision checkpoint's safety depend on option presentation order rather than an authored choice. Add an optional auto_select="<option-id>" attribute on the <task> tag: absent, auto-mode now escalates to a human exactly like gate="blocking-human" does; present, it names the option auto-mode selects; naming an id with no matching <option id> is a hard structural-validation error at plan-parse time rather than a silent fallback to the first option. gate="blocking-human" continues to win over everything, unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): anchor auto_select/id attribute regexes past hyphenated decoys An isolated adversarial review of the auto_select work found that both new attribute regexes used \b as their left anchor, which is a word boundary, not a "start of attribute name" boundary. A decoy attribute ending in the same word (e.g. data-id="...") sitting before the real id="..." on the same <option> tag matched first, silently corrupting the extracted option id. Anchor on (?:^|\s) instead so only the real attribute name can match. Adds a regression test reproducing the exact decoy-attribute shape, plus a Unicode option-id test from the same review pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4095): register auto-select-attribute.test.cjs in the docs-guard lane lint-docs-guard-registration failed: the new test reads docs/reference/ plan-md.md but was not registered, so a future edit to that doc could silently desync from the test without the guard catching it on the PR that changed the doc. Registered alongside its direct precedents (precondition-element.test.cjs, reversibility-tagging.test.cjs), which read the same file for the same reason. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): fit the decision bullet under execute-phase.md's frozen byte ceiling The remote gsd-test run caught what local checks missed: execute-phase.md carries a frozen ADR-857 Phase-6 byte ceiling (93600) with only 36 bytes of headroom before this change, and the original checkpoint:decision wording pushed it to 93772 (over the ceiling). Cascaded into failures in phase6-capstone-conformance, execute-phase-completion-reconciliation, claude-orchestration, and the compact-content drift-report test. Also caught: tests/package-legitimacy-gate.test.cjs anchors a "decision is conditional, not unconditional" safety check on the literal phrase "first option" in the decision bullet — which #4095 deliberately removes, since there is no more unconditional first-option pick. The test was asserting an assumption this change intentionally makes obsolete; re-anchored on tokens that still identify the bullet ('decision', 'auto-spawn') without weakening what the test actually verifies (the bullet must still carry a blocking-human carve-out). Also fixed a word-order mismatch between my own new test's regex and the actual doc text it was asserting against (tests/auto-select-attribute.test.cjs). Regenerated the compact-content benchmark baseline (tests/fixtures/compact-content-benchmark-baseline.json) to match the new byte counts. Emitted-Drift-Ack-Growth: gsd-executor.md — +7 bytes (49138 -> 49145), from the auto_select carve-out added to the checkpoint:decision auto-mode bullet; already trimmed once to fit the 49152 hard cap. Emitted-Drift-Ack-Growth: execute-phase.md — +12 bytes (93564 -> 93576), from the same carve-out in the orchestrator's decision bullet; kept 24 bytes under the frozen 93600 ADR-857 ceiling after two rounds of trimming for clarity vs. margin. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4095): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
5
.changeset/bold-bears-zip.md
Normal file
5
.changeset/bold-bears-zip.md
Normal file
@@ -0,0 +1,5 @@
|
||||
---
|
||||
type: Changed
|
||||
pr: 4912
|
||||
---
|
||||
**`checkpoint:decision` auto-selection is now opt-in via `auto_select`** — in auto-mode, a decision checkpoint with no `auto_select="<option-id>"` attribute now escalates to a human instead of silently picking the first `<option>`. Add `auto_select` naming the intended option's `id` to keep a plan fully unattended; an `auto_select` that names a non-existent option id now fails `verify plan-structure` at plan-parse time. (#4095)
|
||||
@@ -545,6 +545,9 @@ The front-of-task side of the GSD plan contract (issue #1949, *The Pragmatic Pro
|
||||
### Reversibility Rating
|
||||
Classification of a planning decision by what undoing it would cost later (issue #1951, *The Pragmatic Programmer* Topic 15 — "Reversibility"; Bezos's one-way/two-way door framing). Three closed values: `reversible` (undo is local and cheap — one file, one function, an implementation swapped behind a stable interface), `costly` (undo touches many call sites or needs a coordinated change — shared interface shape, cross-module contract, dependency major bump), `one-way` (undo requires a data migration, breaks a published contract, or is impossible — on-disk/wire format, public API shape, external-service lock-in). Surfaced two places: `discuss-phase` records it inline on `<decisions>` entries in phase CONTEXT.md as `— **Reversibility:** <rating> — <rationale>` (optional; an unrated decision is treated as `reversible`), and `gsd-planner` carries it onto the implementing task as the optional `<reversibility rating="…">` element. Planner behavior is rating-dependent: `one-way` inserts a `checkpoint:decision` BEFORE the dependent task (and forces `autonomous: false`), `costly` is flagged in the plan but never blocks, `reversible` flows normally. Reuses the existing `checkpoint:decision` mechanism — no new checkpoint machinery. Override `--no-reversibility-gates` (`REVERSIBILITY_GATES=false`, `/gsd:plan-phase`) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings, so the signal survives a run that chose not to stop for it. The structural validator (`cmdVerifyPlanStructure`) does not reject unknown optional tags, so `<reversibility>` passes plan-structure validation unchanged. Canonical taxonomy owner: `gsd-core/references/planner-reversibility.md`; schema reference: `docs/reference/plan-md.md` → Reversibility; the Reversibility Test thinking model (`references/thinking-models-planning.md` #4) is the reasoning step that produces the rating and consumes this taxonomy rather than defining a second one. Primary anti-pattern: rating everything `one-way` (checkpoint fatigue) — default to `reversible` when unsure, and prefer *removing* irreversibility (writer seam, versioned contract, vendor adapter) over gating it. The decision-risk companion to the Precondition (#1949), which guards implementation assumptions. See Precondition, Tracer Bullet.
|
||||
|
||||
### Decision Auto-select
|
||||
Opt-in mechanism for `checkpoint:decision` auto-mode selection (issue #4095). `auto_select="<option-id>"` is an optional attribute on `<task type="checkpoint:decision">` naming the `<option id="…">` that unattended (`workflow._auto_chain_active` / `workflow.auto_advance`) execution should pick when the checkpoint is reached. Before this attribute existed, auto-mode always selected the FIRST `<option>` — combined with the planner convention of front-loading the recommended choice, this made a decision checkpoint's safety depend on option presentation order, a property no plan author was told was load-bearing. Semantics: `auto_select` absent → auto-mode escalates to a human, the same treatment `gate="blocking-human"` already gets; `auto_select` naming a real option id → that option is selected and logged; `auto_select` naming an id with no matching `<option id="…">` → `verify plan-structure` fails at plan-parse time, never a silent fallback to the first option. `gate="blocking-human"` continues to win over everything, unchanged. The structural validator (`cmdVerifyPlanStructure`) does not reject a plan that omits `auto_select` — only a *declared* `auto_select` with no matching option id is an error, so every pre-#4095 plan remains structurally valid; what changes is auto-mode's runtime behavior on such a plan (escalate instead of guessing), not its plan-structure validity. Canonical schema reference: `docs/reference/plan-md.md` → Auto-select; behavioral reference: `gsd-core/references/checkpoints.md` → `checkpoint:decision`. See Reversibility Rating.
|
||||
|
||||
### Behavior-Adding Task
|
||||
Predicate over a PLAN.md task: `tdd="true"` frontmatter AND `<behavior>` block names a user-visible outcome AND `<files>` includes at least one non-`*.md` / non-`*.json` / non-`*.test.*` source file. Pure doc/config/test-only tasks are exempt. The MVP+TDD Gate (in `references/execute-mvp-tdd.md`) only halts execution on this predicate; the gsd-executor agent applies all three checks at runtime. Currently a prose-only specification — no shared utility.
|
||||
|
||||
|
||||
@@ -331,7 +331,7 @@ For full automation-first patterns, server lifecycle, CLI handling:
|
||||
**Auto-mode checkpoint behavior** (when `AUTO_CFG` is `"true"`):
|
||||
|
||||
- **checkpoint:human-verify** → Auto-approve **except package-legitimacy checkpoints**. If checkpoint has `gate="blocking-human"` OR its purpose indicates package legitimacy verification (`what-built` mentions `Package verification required before install` or `Package install failed — human verification required`), do **not** auto-approve. STOP and return checkpoint_return_format for explicit human confirmation. Precondition-unmet checkpoints report `blocking-human` — never auto-approved.
|
||||
- **checkpoint:decision** → If checkpoint has `gate="blocking-human"`, do **not** auto-select — STOP and return checkpoint_return_format for an explicit human decision (a `blocking-human` decision exists because its default answer would be wrong to assume). Otherwise auto-select first option (planners front-load the recommended choice), log `⚡ Auto-selected: [option name]`, continue to next task.
|
||||
- **checkpoint:decision** → If checkpoint has `gate="blocking-human"`, do **not** auto-select — STOP and return checkpoint_return_format for an explicit human decision (a `blocking-human` decision exists because its default answer would be wrong to assume). Otherwise auto-select `auto_select`'s option (absent → STOP like `blocking-human`, #4095), log `⚡ Auto-selected: [option]`, continue to next task.
|
||||
- **checkpoint:human-action** → STOP normally. Auth gates cannot be automated — return structured checkpoint message using checkpoint_return_format.
|
||||
|
||||
**Standard checkpoint behavior** (when `AUTO_CFG` is not `"true"`):
|
||||
|
||||
@@ -244,6 +244,40 @@ Full taxonomy, emission rules, and anti-patterns (chiefly: rating everything `on
|
||||
|
||||
---
|
||||
|
||||
## Auto-select
|
||||
|
||||
`auto_select` is an **optional** attribute on a `<task type="checkpoint:decision">` element (issue #4095). It names the `id` of the `<option>` that auto-mode (`workflow._auto_chain_active` / `workflow.auto_advance`) should select when the checkpoint is reached unattended.
|
||||
|
||||
```xml
|
||||
<task type="checkpoint:decision" gate="blocking" auto_select="nextauth">
|
||||
<decision>Select authentication provider</decision>
|
||||
<options>
|
||||
<option id="supabase"><name>Supabase Auth</name></option>
|
||||
<option id="clerk"><name>Clerk</name></option>
|
||||
<option id="nextauth"><name>NextAuth.js</name></option>
|
||||
</options>
|
||||
<resume-signal>Select: supabase, clerk, or nextauth</resume-signal>
|
||||
</task>
|
||||
```
|
||||
|
||||
**Semantics:**
|
||||
|
||||
| `auto_select` | Auto-mode behavior |
|
||||
|---|---|
|
||||
| Absent | Escalates to a human — same treatment as `gate="blocking-human"`. Auto-mode does not guess an answer from option order. |
|
||||
| Names a real `<option id="…">` | Selects that option and logs `⚡ Auto-selected: [option id]`, then continues. |
|
||||
| Names an id that does not match any `<option id="…">` | `verify plan-structure` fails at plan-parse time. Never a silent fallback to the first option. |
|
||||
|
||||
`gate="blocking-human"` continues to win over everything, unchanged — it stops for a human in every mode regardless of `auto_select`.
|
||||
|
||||
**Why:** before this attribute existed, auto-mode always picked the first `<option>`, and the planner convention of front-loading the recommended choice meant the safety of every decision checkpoint depended on presentation order — a detail no plan author was told was load-bearing. `auto_select` makes the unattended answer an authored decision instead of a byproduct of layout.
|
||||
|
||||
**Optional and back-compat for the structural validator:** a plan that omits `auto_select` on every `checkpoint:decision` task still passes `verify plan-structure` — the validator only rejects a *declared* `auto_select` that doesn't match any option id. What changes is auto-mode's *runtime* behavior (escalate instead of guessing), not plan-structure validity.
|
||||
|
||||
Full behavioral reference: `gsd-core/references/checkpoints.md` → `checkpoint:decision`.
|
||||
|
||||
---
|
||||
|
||||
## Task types
|
||||
|
||||
| Type | Use | Autonomy |
|
||||
@@ -251,7 +285,7 @@ Full taxonomy, emission rules, and anti-patterns (chiefly: rating everything `on
|
||||
| `auto` | Everything the executor can do independently. | Fully autonomous. |
|
||||
| `tracer` | The leading thin end-to-end slice a plan starts with by default (tracer-first) — production-quality, wired through every layer, with a real end-to-end `<verify>`. | Fully autonomous; after committing, the executor runs the tracer's `<verify>` as an early integration gate. A tracer carrying `gate="blocking-human"` STOPs for a human in every mode, auto included. Otherwise autonomous runs halt on failure before expansion, and interactive runs honor `workflow.human_verify_mode` (#3299): under the `end-of-phase` default a `<verify>` carrying only `<automated>` is re-run and, on success, expansion continues with **no** checkpoint (failure still halts); under `mid-flight`, or when the tracer carries `<human-check>`, a `checkpoint:human-verify` is presented. Full precedence chain: `gsd-core/references/checkpoints.md` → "Tracer feedback gate". |
|
||||
| `checkpoint:human-verify` | Visual or functional verification that requires a human to look at a running UI or service. | Pauses execution; presents to the developer; resumes on approval. |
|
||||
| `checkpoint:decision` | Implementation choices that arose during execution and require the developer's input. | Pauses execution; presents options; resumes on selection. |
|
||||
| `checkpoint:decision` | Implementation choices that arose during execution and require the developer's input. | Pauses execution; presents options; resumes on selection. In auto-mode, an `auto_select="<option-id>"` attribute lets the plan name the unattended answer — see [Auto-select](#auto-select). |
|
||||
| `checkpoint:human-action` | Truly unavoidable manual steps (account creation, hardware interaction). Used sparingly. | Pauses execution; resumes on confirmation. |
|
||||
|
||||
Plans that contain any checkpoint task must set `autonomous: false` in frontmatter.
|
||||
|
||||
@@ -8,14 +8,14 @@ Plans execute autonomously. Checkpoints formalize interaction points where human
|
||||
2. **Claude sets up the verification environment** - Start dev servers, seed databases, configure env vars
|
||||
3. **User only does what requires human judgment** - Visual checks, UX evaluation, "does this feel right?"
|
||||
4. **Secrets come from user, automation comes from Claude** - Ask for API keys, then Claude uses them via CLI
|
||||
5. **Auto-mode bypasses verification/decision checkpoints** — When `workflow._auto_chain_active` or `workflow.auto_advance` is true in config: human-verify auto-approves, decision auto-selects first option, human-action still stops (auth gates cannot be automated)
|
||||
5. **Auto-mode bypasses verification checkpoints, and decision checkpoints with a declared `auto_select`** — When `workflow._auto_chain_active` or `workflow.auto_advance` is true in config: human-verify auto-approves; a `checkpoint:decision` carrying `auto_select="<option-id>"` auto-selects that named option; a `checkpoint:decision` with NO `auto_select` escalates to a human instead of guessing by position (#4095); human-action still stops (auth gates cannot be automated)
|
||||
6. **`gate="blocking-human"` is never auto-approved** — a checkpoint carrying this gate stops for a human in *every* mode, including auto-mode, regardless of its type. Rule 5 does not apply to it. The executor's precondition-unmet checkpoint (a task's `<precondition>` evaluated false — unmet `user_setup` step, missing env var, absent prior-phase artifact) reports this gate (#3210).
|
||||
|
||||
**The `gate` attribute:**
|
||||
|
||||
| Value | Auto-mode behavior | Use for |
|
||||
|-------|--------------------|---------|
|
||||
| `gate="blocking"` | Bypassed per rule 5 (human-verify auto-approves, decision auto-selects) | The default. Post-hoc verification and implementation choices that are safe to take the recommended path on when unattended. |
|
||||
| `gate="blocking"` | Bypassed per rule 5 (human-verify auto-approves; decision auto-selects only when `auto_select` names an option, else escalates) | The default. Post-hoc verification and implementation choices that are safe to take the recommended path on when unattended. |
|
||||
| `gate="blocking-human"` | **Never bypassed.** Stops for a human in auto-mode too. | Irreversible or trust-establishing steps a human must actually see: package-legitimacy verification before install, any decision whose default answer would be wrong to assume, and unmet `<precondition>` facts the executor cannot establish on its own (#3210). |
|
||||
|
||||
Reach for `gate="blocking-human"` whenever auto-approving the checkpoint would defeat its purpose. If the checkpoint exists because a human must *decide* something, `blocking` is the wrong gate — auto-mode will decide it for them.
|
||||
@@ -147,7 +147,7 @@ HUMAN_VERIFY_MODE=$(gsd_run query config-get workflow.human_verify_mode --defaul
|
||||
|
||||
**Structure:**
|
||||
```xml
|
||||
<task type="checkpoint:decision" gate="blocking">
|
||||
<task type="checkpoint:decision" gate="blocking" auto_select="option-a">
|
||||
<decision>[What's being decided]</decision>
|
||||
<context>[Why this decision matters]</context>
|
||||
<options>
|
||||
@@ -166,6 +166,8 @@ HUMAN_VERIFY_MODE=$(gsd_run query config-get workflow.human_verify_mode --defaul
|
||||
</task>
|
||||
```
|
||||
|
||||
`auto_select` is optional and names the `id` of the `<option>` auto-mode should pick when this checkpoint is reached unattended. Omit it and auto-mode escalates to a human instead of guessing from option order — the same treatment `gate="blocking-human"` already gets. An `auto_select` value that doesn't match any `<option id="…">` fails `verify plan-structure` at plan-parse time rather than silently falling back to the first option.
|
||||
|
||||
**Example: Auth Provider Selection**
|
||||
```xml
|
||||
<task type="checkpoint:decision" gate="blocking">
|
||||
|
||||
@@ -1059,7 +1059,7 @@ AUTO_MODE=$(gsd_run query check auto-mode --pick active 2>/dev/null)
|
||||
|
||||
When executor returns a checkpoint AND `AUTO_MODE` is `true`:
|
||||
- **human-verify** → Auto-spawn continuation agent with `{user_response}` = `"approved"`. Log `⚡ Auto-approved checkpoint`. **Except `blocking-human`.**
|
||||
- **decision** → Auto-spawn continuation agent with `{user_response}` = first option from checkpoint details. Log `⚡ Auto-selected: [option]`. **Except `blocking-human`.**
|
||||
- **decision** → `auto_select` present: auto-spawn with `{user_response}` = that option, log `⚡ Auto-selected: [option]`. Absent: present to user (#4095). **Except `blocking-human`.**
|
||||
- **human-action** → Present to user (existing behavior below). Auth gates cannot be automated.
|
||||
|
||||
<!-- gsd:protected -->
|
||||
|
||||
@@ -174,6 +174,7 @@ const DOCS_GUARD_TESTS = {
|
||||
// entry).
|
||||
'tests/learnings.test.cjs': ['docs/FEATURES.md'],
|
||||
'tests/analyze-dependencies.test.cjs': ['docs/COMMANDS.md'],
|
||||
'tests/auto-select-attribute.test.cjs': ['docs/reference/plan-md.md'],
|
||||
'tests/autonomous-converge.test.cjs': [
|
||||
'docs/COMMANDS.md',
|
||||
'docs/how-to/run-phases-autonomously.md',
|
||||
|
||||
@@ -1016,6 +1016,10 @@ interface PlanTaskInfo {
|
||||
// checkpoint:decision fields
|
||||
hasDecision: boolean;
|
||||
hasOptions: boolean;
|
||||
/** `auto_select` attribute value from the opening `<task>` tag, or null if absent (#4095). */
|
||||
autoSelect: string | null;
|
||||
/** `id` attribute of every `<option id="…">` found inside `<options>…</options>`, in document order (#4095). */
|
||||
optionIds: string[];
|
||||
// checkpoint:human-action fields
|
||||
hasInstructions: boolean;
|
||||
hasVerification: boolean;
|
||||
@@ -1060,6 +1064,15 @@ function extractPlanTaskInfos(content: string): PlanTaskInfo[] {
|
||||
const hasName = nameArr.length > 0;
|
||||
const name = hasName ? nameArr[0].trim() : '';
|
||||
|
||||
// `(?:^|\s)` (not `\b`) so a hyphenated attribute ending in `auto_select`
|
||||
// can never be mistaken for the real attribute — the same defensive
|
||||
// anchor as extractOptionIds' `id` match below (#4095).
|
||||
const autoSelectMatch = attrs.match(/(?:^|\s)auto_select\s*=\s*["']([^"']*)["']/);
|
||||
const autoSelect = autoSelectMatch ? autoSelectMatch[1] : null;
|
||||
|
||||
const optionsArr = extractTaggedBlocks(body, 'options');
|
||||
const optionIds = optionsArr.length > 0 ? extractOptionIds(optionsArr[0]) : [];
|
||||
|
||||
infos.push({
|
||||
name,
|
||||
type,
|
||||
@@ -1077,6 +1090,8 @@ function extractPlanTaskInfos(content: string): PlanTaskInfo[] {
|
||||
hasHowToVerify: /<how-to-verify[\s>]/.test(body),
|
||||
hasDecision: /<decision[\s>]/.test(body),
|
||||
hasOptions: /<options[\s>]/.test(body),
|
||||
autoSelect,
|
||||
optionIds,
|
||||
hasInstructions: /<instructions[\s>]/.test(body),
|
||||
hasVerification: /<verification[\s>]/.test(body),
|
||||
hasResumeSignal: /<resume-signal[\s>]/.test(body),
|
||||
@@ -1090,6 +1105,36 @@ function extractPlanTaskInfos(content: string): PlanTaskInfo[] {
|
||||
return infos;
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract the `id` attribute of every `<option id="…">` opening tag found in
|
||||
* `optionsBody` (the inner text of one `<options>…</options>` block), in
|
||||
* document order. Bounded attribute scan (`[^>]{0,500}`), mirroring the same
|
||||
* ReDoS-safe idiom `extractPlanTaskInfos` uses for the `<task type="…">`
|
||||
* attribute string — this file's established pattern for reading an
|
||||
* attribute value without a general XML parser (#4095).
|
||||
*/
|
||||
function extractOptionIds(optionsBody: string): string[] {
|
||||
const ids: string[] = [];
|
||||
if (typeof optionsBody !== 'string' || optionsBody.length === 0) return ids;
|
||||
|
||||
const OPTION_OPEN_RE = /<option(\s[^>]{0,500})?>/g;
|
||||
let match: RegExpExecArray | null;
|
||||
while ((match = OPTION_OPEN_RE.exec(optionsBody)) !== null) {
|
||||
const attrs = match[1] ?? '';
|
||||
// `(?:^|\s)` (not `\b`) so a decoy attribute like `data-id="…"` inside
|
||||
// the same opening tag cannot be mistaken for the real `id` — `\b`
|
||||
// matches at the `-`→`i` boundary too, which `.match()`'s
|
||||
// first-hit-wins semantics would then silently prefer (#4095).
|
||||
const idMatch = attrs.match(/(?:^|\s)id\s*=\s*["']([^"']{1,200})["']/);
|
||||
if (idMatch) ids.push(idMatch[1]);
|
||||
|
||||
if (match.index === OPTION_OPEN_RE.lastIndex) {
|
||||
OPTION_OPEN_RE.lastIndex++;
|
||||
}
|
||||
}
|
||||
return ids;
|
||||
}
|
||||
|
||||
function isCheckpointType(type: string): boolean {
|
||||
return type.startsWith('checkpoint:');
|
||||
}
|
||||
@@ -1131,6 +1176,16 @@ function validatePlanTaskStructure(task: PlanTaskInfo): { errors: string[]; warn
|
||||
case 'checkpoint:decision':
|
||||
if (!task.hasDecision) errors.push(`Task '${taskName}' missing <decision>`);
|
||||
if (!task.hasOptions) errors.push(`Task '${taskName}' missing <options>`);
|
||||
if (task.autoSelect !== null) {
|
||||
if (task.autoSelect.length === 0) {
|
||||
errors.push(`Task '${taskName}' auto_select is empty — name an <option id="…">`);
|
||||
} else if (task.hasOptions && !task.optionIds.includes(task.autoSelect)) {
|
||||
errors.push(
|
||||
`Task '${taskName}' auto_select="${task.autoSelect}" does not match any `
|
||||
+ `<option id="…"> (available: ${task.optionIds.join(', ') || 'none'})`,
|
||||
);
|
||||
}
|
||||
}
|
||||
break;
|
||||
case 'checkpoint:human-action':
|
||||
if (!task.hasAction) errors.push(`Task '${taskName}' missing <action>`);
|
||||
|
||||
366
tests/auto-select-attribute.test.cjs
Normal file
366
tests/auto-select-attribute.test.cjs
Normal file
@@ -0,0 +1,366 @@
|
||||
// allow-test-rule: source-text-is-the-product [#4095]
|
||||
// Agent .md, workflow .md, reference .md, and docs/reference/*.md files — their text IS what
|
||||
// the runtime loads. Per CONTRIBUTING.md exception matrix, asserting these files document the
|
||||
// auto_select contract tests the deployed surface, not derived behavior. The behavioral test
|
||||
// (cmdVerifyPlanStructure) asserts the validator's actual parse-time logic. Issue #4095.
|
||||
|
||||
'use strict';
|
||||
|
||||
const { describe, test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs');
|
||||
const { lfByteCount } = require('../scripts/workflow-size.cjs');
|
||||
|
||||
const ROOT = path.resolve(__dirname, '..');
|
||||
const PLAN_MD_DOC = path.join(ROOT, 'docs', 'reference', 'plan-md.md');
|
||||
const EXECUTOR = path.join(ROOT, 'agents', 'gsd-executor.md');
|
||||
const EXECUTE_PHASE_WORKFLOW = path.join(ROOT, 'gsd-core', 'workflows', 'execute-phase.md');
|
||||
const CHECKPOINTS_REF = path.join(ROOT, 'gsd-core', 'references', 'checkpoints.md');
|
||||
|
||||
/** Agent-file hard red line (tests/agent-size-budget.test.cjs LARGE_CAP). */
|
||||
const LARGE_CAP = 49152;
|
||||
|
||||
function read(file) {
|
||||
return fs.readFileSync(file, 'utf-8').replace(/\r\n/g, '\n').replace(/\r/g, '\n');
|
||||
}
|
||||
|
||||
// ─── Schema documentation (docs/reference/plan-md.md) ────────────────────────
|
||||
|
||||
describe('issue #4095: plan-md.md documents auto_select', () => {
|
||||
test('plan-md.md has an Auto-select section', () => {
|
||||
const doc = read(PLAN_MD_DOC);
|
||||
assert.match(
|
||||
doc,
|
||||
/^## Auto-select$/m,
|
||||
'docs/reference/plan-md.md must have a "## Auto-select" section',
|
||||
);
|
||||
});
|
||||
|
||||
test('plan-md.md states auto_select is optional', () => {
|
||||
const doc = read(PLAN_MD_DOC);
|
||||
assert.match(
|
||||
doc,
|
||||
/`auto_select`[^\n]*\*\*optional\*\*|\*\*optional\*\*[^\n]*`auto_select`/,
|
||||
'plan-md.md must describe auto_select as optional',
|
||||
);
|
||||
});
|
||||
|
||||
test('plan-md.md documents that an unmatched auto_select fails at plan-parse time (not a silent fallback)', () => {
|
||||
const doc = read(PLAN_MD_DOC);
|
||||
assert.match(
|
||||
doc,
|
||||
/`verify plan-structure` fails at plan-parse time/,
|
||||
'plan-md.md must state an unmatched auto_select fails verify plan-structure at plan-parse time',
|
||||
);
|
||||
assert.match(
|
||||
doc,
|
||||
/[Nn]ever a silent fallback to the first option/,
|
||||
'plan-md.md must explicitly rule out silently falling back to the first option',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Executor bypass contract (agents/gsd-executor.md) ───────────────────────
|
||||
|
||||
describe('issue #4095: gsd-executor.md routes absent auto_select through checkpoint_return_format', () => {
|
||||
test('executor mentions auto_select', () => {
|
||||
const exec = read(EXECUTOR);
|
||||
assert.match(exec, /auto_select/, 'gsd-executor.md must reference auto_select');
|
||||
});
|
||||
|
||||
test('checkpoint:decision auto-mode bullet routes an absent auto_select like blocking-human', () => {
|
||||
const exec = read(EXECUTOR);
|
||||
const bullet = exec.split('\n').find(
|
||||
(l) => l.includes('**checkpoint:decision**') && /[Aa]uto-select/.test(l),
|
||||
);
|
||||
assert.ok(bullet, 'gsd-executor.md must keep the checkpoint:decision auto-mode bullet');
|
||||
assert.match(
|
||||
bullet,
|
||||
/auto_select/,
|
||||
'the checkpoint:decision auto-mode bullet must mention auto_select',
|
||||
);
|
||||
assert.match(
|
||||
bullet,
|
||||
/blocking-human/,
|
||||
'the checkpoint:decision auto-mode bullet must route an absent auto_select the same way as blocking-human',
|
||||
);
|
||||
});
|
||||
|
||||
test('executor is under the 49152-byte cap after adding auto_select content', () => {
|
||||
const bytes = lfByteCount(EXECUTOR);
|
||||
assert.ok(
|
||||
bytes < LARGE_CAP,
|
||||
`gsd-executor.md is ${bytes} bytes, must be < ${LARGE_CAP} (LF-normalized, measured the same way tests/agent-size-budget.test.cjs does)`,
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Orchestrator contract (gsd-core/workflows/execute-phase.md) ─────────────
|
||||
|
||||
describe('issue #4095: execute-phase.md decision bullet and carve-out', () => {
|
||||
test('the decision bullet mentions auto_select', () => {
|
||||
const wf = read(EXECUTE_PHASE_WORKFLOW);
|
||||
const bullet = wf.split('\n').find((l) => l.trim().startsWith('- **decision** →'));
|
||||
assert.ok(bullet, 'execute-phase.md must keep the "- **decision** →" bullet');
|
||||
assert.match(bullet, /auto_select/, 'the decision bullet must mention auto_select');
|
||||
});
|
||||
|
||||
test('the protected carve-out paragraph is unchanged', () => {
|
||||
const wf = read(EXECUTE_PHASE_WORKFLOW);
|
||||
assert.ok(
|
||||
wf.includes(
|
||||
'**Carve-out — overrides all branches above.** If the returned `Gate:` is `blocking-human`',
|
||||
),
|
||||
'execute-phase.md must keep the carve-out paragraph verbatim — it is a <!-- gsd:protected --> '
|
||||
+ 'section and must not be touched by the auto_select change',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── checkpoints.md contract ──────────────────────────────────────────────────
|
||||
|
||||
describe('issue #4095: checkpoints.md golden rule 5 and checkpoint:decision example', () => {
|
||||
test('golden rule 5 no longer makes a bare unconditional "decision auto-selects first option" claim', () => {
|
||||
const ref = read(CHECKPOINTS_REF);
|
||||
assert.doesNotMatch(
|
||||
ref,
|
||||
/decision auto-selects first option/,
|
||||
'checkpoints.md must not claim decision checkpoints auto-select the first option unconditionally',
|
||||
);
|
||||
});
|
||||
|
||||
test('golden rule 5 escalates to a human when auto_select is absent', () => {
|
||||
const ref = read(CHECKPOINTS_REF);
|
||||
const rule5 = ref.split('\n').find((l) => /^5\. \*\*Auto-mode bypasses/.test(l));
|
||||
assert.ok(rule5, 'checkpoints.md must keep golden rule 5');
|
||||
assert.match(rule5, /escalates to a human/, 'golden rule 5 must state that an absent auto_select escalates to a human');
|
||||
assert.match(rule5, /auto_select/, 'golden rule 5 must mention auto_select');
|
||||
});
|
||||
|
||||
test('the checkpoint:decision example shows auto_select= on the opening <task> tag', () => {
|
||||
const ref = read(CHECKPOINTS_REF);
|
||||
assert.match(
|
||||
ref,
|
||||
/<task type="checkpoint:decision"[^\n>]*auto_select="[^"]+"/,
|
||||
'checkpoints.md must show an auto_select="…" attribute on a checkpoint:decision <task> tag',
|
||||
);
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Behavioral test: cmdVerifyPlanStructure additive + validating ───────────
|
||||
//
|
||||
// checkpoint:decision requires <decision>, <options>, <resume-signal> per the
|
||||
// existing validator, and the plan frontmatter must set autonomous: false
|
||||
// because the plan contains a checkpoint (src/verify.cts's
|
||||
// "Has checkpoint tasks but autonomous is not false" rule).
|
||||
|
||||
function planWith({
|
||||
autoSelect = undefined,
|
||||
optionIds = ['a', 'b', 'c'],
|
||||
includeOptions = true,
|
||||
optionAttrsById = {},
|
||||
} = {}) {
|
||||
const attrs = ['type="checkpoint:decision"', 'gate="blocking"'];
|
||||
if (autoSelect !== undefined) {
|
||||
attrs.push(`auto_select="${autoSelect}"`);
|
||||
}
|
||||
const lines = [
|
||||
`<task ${attrs.join(' ')}>`,
|
||||
' <name>Task 1: Pick the thing</name>',
|
||||
' <decision>Pick the thing</decision>',
|
||||
' <context>Later phases depend on this.</context>',
|
||||
];
|
||||
if (includeOptions) {
|
||||
lines.push(' <options>');
|
||||
for (const id of optionIds) {
|
||||
const extraAttrs = optionAttrsById[id] || '';
|
||||
lines.push(
|
||||
` <option ${extraAttrs}id="${id}">`,
|
||||
` <name>Option ${id}</name>`,
|
||||
' <pros>Pro</pros>',
|
||||
' <cons>Con</cons>',
|
||||
' </option>',
|
||||
);
|
||||
}
|
||||
lines.push(' </options>');
|
||||
}
|
||||
lines.push(
|
||||
' <resume-signal>Select: ' + optionIds.join(', ') + '</resume-signal>',
|
||||
'</task>',
|
||||
'',
|
||||
);
|
||||
return [
|
||||
'---',
|
||||
'phase: 01-test',
|
||||
'plan: 01',
|
||||
'type: execute',
|
||||
'wave: 1',
|
||||
'depends_on: []',
|
||||
'files_modified: [src/x.ts]',
|
||||
'autonomous: false',
|
||||
'must_haves:',
|
||||
' truths:',
|
||||
' - "something is true"',
|
||||
'---',
|
||||
'',
|
||||
'<tasks>',
|
||||
'',
|
||||
...lines,
|
||||
'</tasks>',
|
||||
].join('\n');
|
||||
}
|
||||
|
||||
function planWithAutoTask({ autoSelect } = {}) {
|
||||
const attrs = ['type="auto"'];
|
||||
if (autoSelect !== undefined) {
|
||||
attrs.push(`auto_select="${autoSelect}"`);
|
||||
}
|
||||
const lines = [
|
||||
`<task ${attrs.join(' ')}>`,
|
||||
' <name>Task 1: Test</name>',
|
||||
' <files>src/x.ts</files>',
|
||||
' <action>Do the thing.</action>',
|
||||
' <verify><automated>echo ok</automated></verify>',
|
||||
' <done>Done</done>',
|
||||
'</task>',
|
||||
'',
|
||||
];
|
||||
return [
|
||||
'---',
|
||||
'phase: 01-test',
|
||||
'plan: 01',
|
||||
'type: execute',
|
||||
'wave: 1',
|
||||
'depends_on: []',
|
||||
'files_modified: [src/x.ts]',
|
||||
'autonomous: true',
|
||||
'must_haves:',
|
||||
' truths:',
|
||||
' - "something is true"',
|
||||
'---',
|
||||
'',
|
||||
'<tasks>',
|
||||
'',
|
||||
...lines,
|
||||
'</tasks>',
|
||||
].join('\n');
|
||||
}
|
||||
|
||||
function verifyPlan(tmpDir, content) {
|
||||
const rel = path.join('.planning', 'phases', '01-test', '01-01-PLAN.md');
|
||||
fs.mkdirSync(path.join(tmpDir, '.planning', 'phases', '01-test'), { recursive: true });
|
||||
fs.writeFileSync(path.join(tmpDir, rel), content);
|
||||
const result = runGsdTools(`verify plan-structure ${rel}`, tmpDir);
|
||||
assert.ok(result.success, `verify plan-structure failed to run: ${result.error}`);
|
||||
return JSON.parse(result.output);
|
||||
}
|
||||
|
||||
describe('issue #4095: cmdVerifyPlanStructure validates auto_select', () => {
|
||||
test('auto_select absent, options present → valid, no errors (back-compat)', (t) => {
|
||||
const tmp = createTempProject();
|
||||
t.after(() => cleanup(tmp));
|
||||
|
||||
const out = verifyPlan(tmp, planWith({ autoSelect: undefined }));
|
||||
assert.strictEqual(out.valid, true, `errors: ${JSON.stringify(out.errors)}`);
|
||||
assert.deepStrictEqual(out.errors, [], 'an absent auto_select must not be flagged (back-compat)');
|
||||
});
|
||||
|
||||
test('auto_select="b" matches an existing <option id="b"> → valid, no auto_select errors', (t) => {
|
||||
const tmp = createTempProject();
|
||||
t.after(() => cleanup(tmp));
|
||||
|
||||
const out = verifyPlan(tmp, planWith({ autoSelect: 'b', optionIds: ['a', 'b', 'c'] }));
|
||||
assert.strictEqual(out.valid, true, `errors: ${JSON.stringify(out.errors)}`);
|
||||
assert.ok(
|
||||
!out.errors.some((e) => /auto_select/i.test(e)),
|
||||
`a matching auto_select must not error; got: ${JSON.stringify(out.errors)}`,
|
||||
);
|
||||
});
|
||||
|
||||
test('auto_select with a Unicode option id matches correctly → valid, no auto_select errors', (t) => {
|
||||
const tmp = createTempProject();
|
||||
t.after(() => cleanup(tmp));
|
||||
|
||||
const out = verifyPlan(tmp, planWith({ autoSelect: '日本語', optionIds: ['a', '日本語', 'c'] }));
|
||||
assert.strictEqual(out.valid, true, `errors: ${JSON.stringify(out.errors)}`);
|
||||
assert.ok(
|
||||
!out.errors.some((e) => /auto_select/i.test(e)),
|
||||
`a Unicode auto_select matching a Unicode option id must not error; got: ${JSON.stringify(out.errors)}`,
|
||||
);
|
||||
});
|
||||
|
||||
test('a decoy attribute ending in "id" (e.g. data-id) on another option does not shadow the real id', (t) => {
|
||||
const tmp = createTempProject();
|
||||
t.after(() => cleanup(tmp));
|
||||
|
||||
const out = verifyPlan(tmp, planWith({
|
||||
autoSelect: 'decoy-holder',
|
||||
optionIds: ['decoy-holder', 'real'],
|
||||
optionAttrsById: { 'decoy-holder': 'data-id="not-a-real-option" ' },
|
||||
}));
|
||||
assert.strictEqual(out.valid, true, `errors: ${JSON.stringify(out.errors)}`);
|
||||
assert.ok(
|
||||
!out.errors.some((e) => /auto_select/i.test(e)),
|
||||
`auto_select="decoy-holder" must match <option id="decoy-holder"> even though that ` +
|
||||
`same option also carries a data-id attribute; got: ${JSON.stringify(out.errors)}`,
|
||||
);
|
||||
});
|
||||
|
||||
test('auto_select="nope" matches no option → invalid, error names auto_select and nope', (t) => {
|
||||
const tmp = createTempProject();
|
||||
t.after(() => cleanup(tmp));
|
||||
|
||||
const out = verifyPlan(tmp, planWith({ autoSelect: 'nope', optionIds: ['a', 'b', 'c'] }));
|
||||
assert.strictEqual(out.valid, false, `errors: ${JSON.stringify(out.errors)}`);
|
||||
assert.ok(
|
||||
out.errors.some((e) => /auto_select/.test(e) && /nope/.test(e)),
|
||||
`an unmatched auto_select must produce an error naming auto_select and nope; got: ${JSON.stringify(out.errors)}`,
|
||||
);
|
||||
});
|
||||
|
||||
test('auto_select="" (empty) → invalid, error mentions auto_select', (t) => {
|
||||
const tmp = createTempProject();
|
||||
t.after(() => cleanup(tmp));
|
||||
|
||||
const out = verifyPlan(tmp, planWith({ autoSelect: '', optionIds: ['a', 'b', 'c'] }));
|
||||
assert.strictEqual(out.valid, false, `errors: ${JSON.stringify(out.errors)}`);
|
||||
assert.ok(
|
||||
out.errors.some((e) => /auto_select/i.test(e)),
|
||||
`an empty auto_select must produce an error mentioning auto_select; got: ${JSON.stringify(out.errors)}`,
|
||||
);
|
||||
});
|
||||
|
||||
test('auto_select="a" with <options> entirely omitted → the existing missing <options> error still fires', (t) => {
|
||||
const tmp = createTempProject();
|
||||
t.after(() => cleanup(tmp));
|
||||
|
||||
const out = verifyPlan(tmp, planWith({ autoSelect: 'a', includeOptions: false, optionIds: ['a'] }));
|
||||
assert.strictEqual(out.valid, false, `errors: ${JSON.stringify(out.errors)}`);
|
||||
assert.ok(
|
||||
out.errors.some((e) => /missing <options>/.test(e)),
|
||||
`omitting <options> must still produce the missing <options> error; got: ${JSON.stringify(out.errors)}`,
|
||||
);
|
||||
});
|
||||
|
||||
test('a non-checkpoint:decision task carrying auto_select is ignored → valid, no error', (t) => {
|
||||
const tmp = createTempProject();
|
||||
t.after(() => cleanup(tmp));
|
||||
|
||||
const out = verifyPlan(tmp, planWithAutoTask({ autoSelect: 'nope' }));
|
||||
assert.strictEqual(out.valid, true, `errors: ${JSON.stringify(out.errors)}`);
|
||||
assert.deepStrictEqual(out.errors, [], 'auto_select on a non-checkpoint:decision task must be inert');
|
||||
});
|
||||
});
|
||||
|
||||
// ─── Parity: plan-md.md and checkpoints.md spell the attribute identically ───
|
||||
|
||||
describe('issue #4095: parity between plan-md.md and checkpoints.md', () => {
|
||||
test('both surfaces spell the attribute auto_select=', () => {
|
||||
const doc = read(PLAN_MD_DOC);
|
||||
const ref = read(CHECKPOINTS_REF);
|
||||
assert.ok(doc.includes('auto_select='), 'plan-md.md must spell auto_select=');
|
||||
assert.ok(ref.includes('auto_select='), 'checkpoints.md must spell auto_select=');
|
||||
});
|
||||
});
|
||||
@@ -18,8 +18,8 @@
|
||||
"reductionPct": 16.51
|
||||
},
|
||||
"execute-phase": {
|
||||
"offTokens": 26344,
|
||||
"onTokens": 24054,
|
||||
"offTokens": 26355,
|
||||
"onTokens": 24065,
|
||||
"reductionPct": 8.69
|
||||
},
|
||||
"new-project": {
|
||||
@@ -39,8 +39,8 @@
|
||||
}
|
||||
},
|
||||
"aggregate": {
|
||||
"offTokens": 109344,
|
||||
"onTokens": 92657,
|
||||
"offTokens": 109355,
|
||||
"onTokens": 92668,
|
||||
"reductionPct": 15.26
|
||||
}
|
||||
}
|
||||
|
||||
@@ -718,7 +718,7 @@ describe('execute-phase.md — orchestrator honors the blocking-human gate', ()
|
||||
|
||||
test('auto-select rule for decision is conditional, not unconditional', () => {
|
||||
const autoSelectLines = lineIndexes(model.lines, (line) =>
|
||||
hasAllTokens(line, ['decision', 'auto-spawn', 'first', 'option'])
|
||||
hasAllTokens(line, ['decision', 'auto-spawn'])
|
||||
);
|
||||
|
||||
assert.ok(
|
||||
|
||||
Reference in New Issue
Block a user