* test(#3172): failing-first suite for the stated failing-direction probe Pins the <fails_when> pairing walk, placeholder denylist, MISSING sentinel exemption, degraded-read contract, CLI arm and the plan-authoring contract text. RED by construction: the module exports it requires do not exist yet. Executed on the remote runner. * feat(#3172): require a stated failing direction for every automated acceptance command Every runnable <automated> command now carries a <fails_when> sibling naming what output constitutes failure. A command with no expressible failure mode is not an acceptance test: it reads as rigour and is not falsifiable. - verify-command-grounding gains a failing-direction probe sharing the existing <automated> grammar, MISSING sentinel and walk guard rather than copying them - gsd-tools check verify-failure-directions <N> backs it; plan-phase dispatches it and hands the JSON to gsd-plan-checker check 8f - Dimension 8 detail extracted to references to stay under the agent size cap Verified on the remote runner. * fix(#3172): close four review findings in the failing-direction probe - MISSING_SENTINEL_RE matched an env-var assignment prefix (MISSING=1 cmd), so a real command was exempted from the new blocking gate. Tightened the SHARED constant rather than adding a second copy. - Both token regexes scanned to EOF on unclosed openers (O(n^2), 1562ms at 40k). Bodies are now non-crossing; 1ms, byte-identical on well-formed input. The pre-existing AUTOMATED_BLOCK_RE carried the same defect and is fixed here too. - probePhaseFailingDirections reported status 'ok' when one plan was unreadable, conflating 'could not look' with 'nothing to report'. - Extracted the phase-resolution block both check arms had copied verbatim. Also corrects a docs/AGENTS.md dimension list stale since #2401. Verified on the remote runner. * fix(#3172): project the planner rule onto the spawn contract, settle emitted bookkeeping The remote runner refuted the planner-side edit. agents/gsd-planner.md is frozen under a 49152-LF-char cap asserted by four suites and sat at 49,146 — six chars of headroom — so the +537 of authoring rule blew it. #3297/#3645 already settled where such a rule goes: the planner spawn contract in plan-phase.md, beside <tracked_source_paths>. The agent file is reverted to origin/next verbatim. - plan-phase.md gains <failing_direction_contract>; tests row 30 now asserts the contract there and row 30b guards the freeze in both directions - plan-phase.md growth acknowledged by APPENDING to the 3409 fragment, per the precedent that two ack sources may never name the same path - install-tree fixtures regenerated for the three new reference files Verified on the remote runner. * chore(#3172): backfill PR number into the changeset fragment pr:0 -> pr:3825 now that the PR exists. --------- Co-authored-by: sim <sim@local>
3.1 KiB
Stated Failing Direction (#3172)
Reference file for the gsd-planner agent. Loaded on-demand via
@reference from the<failing_direction_contract>block of the planner spawn prompt ingsd-core/workflows/plan-phase.md— NOT fromagents/gsd-planner.md, which is frozen under a 49152-LF-char cap, so planner-side rules are projected onto its spawn contract (the #3297 / #3645 precedent).
Every runnable <automated> command needs a <fails_when> sibling naming what output
constitutes failure. A command with no expressible failure mode is not an acceptance test.
<verify>
<automated>npm --prefix apps/api test -- auth.spec.ts</automated>
<fails_when>non-zero exit, or "0 passed" in the summary line</fails_when>
</verify>
Why. #3172: six plans shipped 21 <automated> commands that could not run at all — a
--lib target against a binary-only package. They sat inside the very blocks that decide whether
work is done, so the acceptance criteria for those plans were improvised at execution time by
three separate executors instead of reviewed at planning time. Cargo happened to exit non-zero,
so it failed loudly. The identical mistake with a command that exits 0 on a no-op — a test-name
filter matching nothing — passes green and silently. Naming the failure signal is what makes the
difference visible while you are still authoring the plan.
The authoring test, applied to yourself: if this command were silently doing nothing, what in its output would tell me? If you cannot answer, you do not yet have an acceptance command — you have a command. Fix the command, do not invent a statement for it.
Rules
- One statement per runnable command, placed immediately after it. Within a task, each
<fails_when>binds to the nearest preceding<automated>, and the first statement after a command is the binding one. Two commands need two statements. - Name an observable signal, not the word "failure".
non-zero exit,"0 passed" in the summary,the coverage line is absent,stderr contains "ECONNREFUSED"are signals. "the command fails", "it doesn't work", "an error occurs" are restatements and will be flagged. - Short is fine.
non-zero exitis complete. There is no minimum length and no required keyword. TBD,TODO,N/A,none,unknown,?,-are rejected outright as whole values. A statement you cannot write is a command you should not ship.- Any characters are safe.
exit code > 0,stderr contains "FAIL" && exit != 0are ordinary prose here — that is exactly why this is an element and not an attribute. - The
MISSING — Wave 0 must create …sentinel is exempt. It is not a runnable command, so it has no failure mode to state. Do not attach a<fails_when>to one.
Where the failing direction comes from
Prefer the signal the tool actually emits over one you imagine. When
prior_verify_commands supplies a command a prior phase already proved, the failure signal that
command produces is the one to state — you have seen its output. When you author a new command,
name the signal from the tool's documented output shape, not from a guess about it.