Files
msd-core/gsd-core/references/tdd.md
0xdhx 092d9256b8 fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it (#4768)
* test(#4748): pin the letter-axis defect at the seven shell sites outside #4660's six

Extends tests/nsegment-phase-grammar.test.cjs one class over: for each of the
seven sites the live shell lines are read off disk by anchor and executed in
bash against a letter-suffixed fixture. The four `$((10#$PHASE_INT))` split
sites must yield PHASE_N without a shell error for `03A` / `12A` / `3A` /
`03A.1.2` and the commit-scope ERE they build must match both `feat(3A-01):`
and `feat(03A-1):`; the review-file lookup must bind init's `padded_phase`
rather than re-pad in shell; the `--from`/`--to`/`--only` and
plan-review-convergence extractions must return `12A` / `23A.1.2` (and
`23.1.2`) whole; the legacy normalizer must pad `3A` to `03A` and must not
mangle an already-padded `08`. Every pre-existing shape (`06`, `08.5`,
`23.1.2`, `36.14`) is a regression control.

tests/init.test.cjs asserts `init execute-phase` emits `padded_phase` for a
directory-backed `03A`, a ROADMAP-only `4B` (→ `04B`), the existing ROADMAP
fallback `1` (→ `01`), and `null` when the phase is not found.

Negative control against the unfixed tree: 41 failures in the grammar file,
exactly the "(fails before the fix)" cases and the three derived from them
(scope ERE, three-flag extraction, the `08` octal trap); 2 in init.test.cjs,
both the new assertions. Every regression control already green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): carry a letter-suffixed phase id through the seven shell sites that aborted or truncated it

The canonical phase-number grammar (src/phase-id.cts) is digits, an optional
uppercase letter, then dotted segments — `12A`, `3A`, `23A.1.2` are documented
shapes that `init`, `phase-id.cts` and `phase remove` renumbering already
round-trip. Seven shell sites in shipped workflows and references still
assumed digits-and-dots. Four classes, one fix each:

Class 1 — `PHASE_INT=${PHASE_NUMBER%%.*}; $((10#$PHASE_INT))` (execute-phase.md
×2, completion-reconciliation.md, tdd.md). The post-#4619 split stops at the
first DOT, so on `03A` the "integer" is `03A` and bash aborts with `value too
great for base`. Split at the first NON-DIGIT instead (`%%[!0-9]*`): the
integer half is a pure digit run, and the letter rides along in the rest the
way the dotted fraction already did — `03A.1.2` → PHASE_N `3A\.1\.2`, so the
#4003 zero-pad-tolerant scope ERE matches both `feat(3A-01):` and
`feat(03A-1):`. Byte-identical output for every id that worked before.

Class 2 — `PADDED=$(printf "%02d" "${PHASE_NUMBER}")` before the REVIEW.md
lookup (execute-phase.md). `printf` cannot pad a letter id (prints `03`,
exits 1) — and cannot even re-pad an already-padded `08`, which bash reads as
an invalid octal and prints as `00`, so the lookup resolved phases 08 and 09
to `00-REVIEW.md` today. The disk path hands the workflow the directory's
padded number but the ROADMAP fallback hands it the heading's bare one, which
is why the re-pad existed. `cmdInitExecutePhase` now emits `padded_phase`
through `normalizePhaseName`, exactly as the plan-phase and code-review inits
do, and the workflow binds `{padded_phase}` instead of re-deriving.

Class 3 — `grep -oE '[0-9]+\.?[0-9]*'` (autonomous.md `--from`/`--to`/`--only`,
plan-review-convergence.md). Stops at the letter, so `--from 12A` ran from
phase 12 with no error. Now the canonical ERE `[0-9]+[A-Z]?(\.[0-9]+)*`, which
also closes the single-segment dot-axis gap the same shape carried (`23.1.2`
→ `23.1`, #4568's class in a spelling neither lint saw).

Class 4 — the legacy manual normalizer (phase-argument-parsing.md, reached
from mvp-phase.md). Its two branches (`^[0-9]+$`, `^[0-9]+\.[0-9]+$`) left
`12A` unpadded and never padded `3A` to the `03A` a directory carries; its
integer branch also hit the same `printf` octal trap on `08`. One branch for
the whole canonical token now, padding the digit run via `$((10#…))`.
Whether this legacy surface should instead be retired in favour of `init`'s
normalization is the maintainer call the issue names; extending it keeps the
documented contract true either way.

Driven end to end: `init execute-phase 3A` on a fixture with a
`03A-letter-variant/` directory emits `phase_number: "03A"` and now
`padded_phase: "03A"`; on a ROADMAP-only `### Phase 4B:` it emits `"4B"` /
`"04B"`. The issue's own evidence line claimed `padded_phase` was already in
the execute-phase init output — it was not; that key is emitted by the
code-review / plan-phase inits, which is where the claim was read from.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): extend lint-phase-id-drift with three ratchets for letter-hostile phase-id consumers

The rules that landed with #4619, #4568 and #4660 police grammar MIRRORS —
regexes that describe a phase id. The #4748 sites are CONSUMERS of one, and
every existing rule reported clean on them: the shell-arithmetic rule's
`_INT` escape trusts a NAME the dot-only split did not earn on `03A`; the
`[0-9]+\.?[0-9]*` shape is neither the bounded form the single-segment rule
bans nor the unbounded form the letterless rule inspects; and nothing looked
at `printf "%02d"` at all. Three narrow additions, one per shape:

- findDotOnlyIntegerSplitDrift — `X_INT=${<phase-var>%%.*}`; the safe split
  is `%%[!0-9]*`. Keys on the SOURCE variable being phase-carrying.
- findLooseDottedPhaseRegexDrift — `[0-9]+\.?[0-9]*` / `\d+\.?\d*` on a
  phase-carrying line; the canonical form is `[0-9]+[A-Z]?(\.[0-9]+)*`.
  Disjoint from the two sibling regex rules by construction.
- findShellPhasePrintfPadDrift — `printf "%0Nd" …` whose arguments name a
  phase-carrying, non-`_INT` variable; a pad of an `_INT` via `$((10#…))`
  and a `{padded_phase}` binding are the sanctioned shapes.

Same `<!-- phase-id-owner: … -->` sanction, same scan roots as their nearest
sibling (shell idioms over workflows + references, the regex shape over
workflows + references + agents), same documented limit of a per-line
textual scan. The post-#4619 comment that described the `_INT` convention
as proven by `%%.*` is corrected to name the digit-run split. Confirmed
against the base commit: each rule fires on exactly its own unfixed sites
(2+1+1, 3+1, 1+1) and zero violations remain on the fixed tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* docs(#4748): add Fixed changeset

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline and acknowledge emitted growth

The three top-level workflow files below grew by the letter-aware split, the
canonical extraction ERE, the `{padded_phase}` binding, and the comment lines
that name the grammar each site now honours. The committed compact-content
benchmark moved with them; refreshed with `benchmark-compact-content.cjs
--write` (aggregate reduction 15.47% -> 15.45%).

Emitted-Drift-Ack-Growth: execute-phase.md — #4748: first-non-digit PHASE_INT split at the plan-selection and TDD-gate sites, `{padded_phase}` binding at the REVIEW.md lookup, and the comments naming why (482 bytes)
Emitted-Drift-Ack-Growth: autonomous.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the --from/--to/--only extractions plus one comment naming the grammar (249 bytes)
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #4748: canonical `[0-9]+[A-Z]?(\.[0-9]+)*` at the phase extraction plus one comment naming the grammar (160 bytes)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* fix(#4748): name padded_phase in execute-phase.md's init parse list

A `{field}` token inside a workflow bash block is substituted from the init
JSON only for fields the workflow tells the model to parse. `phase_number`
is on that list; `padded_phase` was not, so the `PADDED="{padded_phase}"`
binding at the review lookup would have been a literal — for every phase,
not only letter ones. Found by the pre-file adversarial review (claim 2, the
author's own named suspicion); the test now asserts the parse list carries
the field beside `phase_number`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on its source and widen the printf rule to any %d form

Two false negatives from the pre-file adversarial review of the three #4748
ratchets: `PHASE_PREFIX=${PHASE_NUMBER%%.*}` escaped the split rule because
the destination did not end in `_INT` (the defect is the split, not the
name it lands in), and `printf '%02d'` / `printf "%2d"` escaped the printf
rule because it required double quotes and the zero flag (`%d` cannot parse
a letter id under any width). Both rules now key on the phase-carrying
SOURCE alone; base-site firing counts are unchanged (2+1+1, 1+1) and the
fixed tree stays at zero. The `[[:digit:]]` spelling and the `/phase/i`
heuristic remain the sibling rules' documented limits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): refresh the compact-content benchmark baseline after the parse-list edit

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): compose init's emitted padded_phase through the live REVIEW.md lookup

The Class 2 site is a `{padded_phase}` template token, which no test can
execute as written. This substitutes the value init emits
(`normalizePhaseName`) into the three live lookup lines and runs them
against a fixture, so the emitted value, the binding, the path construction
and the status extraction are exercised together — `03A-REVIEW.md` and
`08-REVIEW.md` each resolve to their own status. Suggested by the resumed
adversarial review pass (claim C).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): move the #4619 and #4003 source-parity pins to the letter-safe split

tests/execute-phase-decimal-arithmetic.test.cjs and
tests/safe-resume-gate-anchoring.test.cjs pin the four Class 1 sites'
snippet byte-for-byte, so the first-non-digit split reddened both in the
whole-suite run (scripts/ci-test-scope.cjs does not select either file for
a workflow edit — the scoped run was green). The pinned snippet is now the
shipped one, and the behavioural half of the #4619 file gains the letter
case (`03A` → `3A`, `23A.1.2` → `23A\.1\.2`) beside its decimal cases.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4634): key the dot-only split rule on the _INT destination again, tolerating the quoted spelling

Keying on the source alone (the previous commit's widening, from a review
probe) flags `PARENT_PHASE="${PHASE_NUMBER%%.*}"` in
gap-closure-artifacts.md — a correct derivation that wants everything
before the first dot, letter included. The defect this rule polices is a
dot split INTO the name the shell-arithmetic rule trusts as an integer, so
`_INT` is the discriminator on purpose; the quoted spelling that site uses
is now tolerated so the same shape into an `_INT` cannot hide behind it.
Base-site firing unchanged (2+1+1), zero on the fixed tree, and the
parent-phase line is pinned as a silent case.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* test(#4748): use t.after() for the composition test's fixture cleanup

CONTRIBUTING forbids try/finally inside a test body; the per-test cleanup
form is `t.after(() => cleanup(dir))`. Flagged by the filing driver's
test-ruleset gate before the PR was created.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

* chore(#4748): set changeset fragment pr to 4768

* chore(#4748): refresh the compact-content benchmark baseline after rebasing onto next

Regenerated with `node scripts/benchmark-compact-content.cjs --write` on the
rebased tree (base 0d6bf19bf); `--check` confirms it matches the live recompute.
Only the execute-phase split and the aggregate totals differ from next's copy.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cjzdZtYjcBAa3Lqh2VrLK

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-15 23:31:36 -04:00

13 KiB

TDD is about design quality, not coverage metrics. The red-green-refactor cycle forces you to think about behavior before implementation, producing cleaner interfaces and more testable code.

Principle: If you can describe the behavior as expect(fn(input)).toBe(output) before writing fn, TDD improves the result.

Key insight: TDD work is fundamentally heavier than standard tasks—it requires 2-3 execution cycles (RED → GREEN → REFACTOR), each with file reads, test runs, and potential debugging. TDD features get dedicated plans to ensure full context is available throughout the cycle.

<when_to_use_tdd>

When TDD Improves Quality

TDD candidates (create a TDD plan):

  • Business logic with defined inputs/outputs
  • API endpoints with request/response contracts
  • Data transformations, parsing, formatting
  • Validation rules and constraints
  • Algorithms with testable behavior
  • State machines and workflows
  • Utility functions with clear specifications

Skip TDD (use standard plan with type="auto" tasks):

  • UI layout, styling, visual components
  • Configuration changes
  • Glue code connecting existing components
  • One-off scripts and migrations
  • Simple CRUD with no business logic
  • Exploratory prototyping

Heuristic: Can you write expect(fn(input)).toBe(output) before writing fn? → Yes: Create a TDD plan → No: Use standard plan, add tests after if needed </when_to_use_tdd>

<tdd_plan_structure>

TDD Plan Structure

Each TDD plan implements one feature through the full RED-GREEN-REFACTOR cycle.

---
phase: XX-name
plan: NN
type: tdd
---

<objective>
[What feature and why]
Purpose: [Design benefit of TDD for this feature]
Output: [Working, tested feature]
</objective>

<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@relevant/source/files.ts
</context>

<feature>
  <name>[Feature name]</name>
  <files>[source file, test file]</files>
  <behavior>
    [Expected behavior in testable terms]
    Cases: input → expected output
  </behavior>
  <implementation>[How to implement once tests pass]</implementation>
</feature>

<verification>
[Test command that proves feature works]
</verification>

<success_criteria>
- Failing test written and committed
- Implementation passes test
- Refactor complete (if needed)
- All 2-3 commits present
</success_criteria>

<output>
After completion, create SUMMARY.md with:
- RED: What test was written, why it failed
- GREEN: What implementation made it pass
- REFACTOR: What cleanup was done (if any)
- Commits: List of commits produced
</output>

One feature per TDD plan. If features are trivial enough to batch, they're trivial enough to skip TDD—use a standard plan and add tests after. </tdd_plan_structure>

<execution_flow>

Red-Green-Refactor Cycle

RED - Write failing test:

  1. Create test file following project conventions
  2. Write test describing expected behavior (from <behavior> element)
  3. Run test - it MUST fail intentionally (#3770): the TARGET test you named must be the test that fails, on an assertion for the planned behavior. A nonzero exit alone is NOT RED — syntax errors, zero-test discovery, fixture crashes, parser errors, and unrelated assertions are INVALID_RED and must not authorize GREEN.
  4. Persist the RED evidence record (command, exit code, failing test, expected result, actual result) and verify it: gsd_run check tdd-red-evidence <record.json>. Only verdict RED_EVIDENCE_OK satisfies the RED gate; INVALID_RED blocks GREEN until the RED phase is fixed.
  5. If test passes: feature exists or test is wrong. Investigate.
  6. Commit: test({phase}-{plan}): add failing test for [feature]

GREEN - Implement to pass:

  1. Write minimal code to make test pass
  2. No cleverness, no optimization - just make it work
  3. Run test - it MUST pass
  4. Commit: feat({phase}-{plan}): implement [feature]

REFACTOR (if needed):

  1. Clean up implementation if obvious improvements exist
  2. Run tests - MUST still pass
  3. Only commit if changes made: refactor({phase}-{plan}): clean up [feature]

Result: Each TDD plan produces 2-3 atomic commits. </execution_flow>

<test_quality>

Good Tests vs Bad Tests

Test behavior, not implementation:

  • Good: "returns formatted date string"
  • Bad: "calls formatDate helper with correct params"
  • Tests should survive refactors

One concept per test:

  • Good: Separate tests for valid input, empty input, malformed input
  • Bad: Single test checking all edge cases with multiple assertions

Descriptive names:

  • Good: "should reject empty email", "returns null for invalid ID"
  • Bad: "test1", "handles error", "works correctly"

No implementation details:

  • Good: Test public API, observable behavior
  • Bad: Mock internals, test private methods, assert on internal state </test_quality>

<framework_setup>

Test Framework Setup (If None Exists)

When executing a TDD plan but no test framework is configured, set it up as part of the RED phase:

1. Detect project type:

# JavaScript/TypeScript
if [ -f package.json ]; then echo "node"; fi

# Python
if [ -f requirements.txt ] || [ -f pyproject.toml ]; then echo "python"; fi

# Go
if [ -f go.mod ]; then echo "go"; fi

# Rust
if [ -f Cargo.toml ]; then echo "rust"; fi

2. Install minimal framework:

Project Framework Install
Node.js Jest npm install -D jest @types/jest ts-jest
Node.js (Vite) Vitest npm install -D vitest
Python pytest pip install pytest
Go testing Built-in
Rust cargo test Built-in

3. Create config if needed:

  • Jest: jest.config.js with ts-jest preset
  • Vitest: vitest.config.ts with test globals
  • pytest: pytest.ini or pyproject.toml section

4. Verify setup:

# Run empty test suite - should pass with 0 tests
npm test  # Node
pytest    # Python
go test ./...  # Go
cargo test    # Rust

5. Create first test file: Follow project conventions for test location. The RED-commit gate (workflows/execute-phase.md) recognises the patterns below at any depth, repo root included. Note the gate's pathspec is deliberately wider than the project types the detection step above enumerates — it costs nothing to recognise a convention the onboarding flow does not yet auto-detect, and a project using one should not have its RED commits go unseen:

  • *.test.* / *.spec.* next to source — JS/TS and anything sharing the convention
  • __tests__/ directory
  • tests/ directory at root
  • *_test.go — Go
  • test_*.py / *_test.py — Python (in addition to tests/)
  • *_test.exs — Elixir
  • *_spec.rb / *_test.rb — Ruby

Cost of the broad pathspec (#4379). *.spec.* can match a non-test file that happens to carry the word — api.spec.json, openapi.spec.yaml — which lets the RED gate pass on a commit touching only that. This is not new: the previous **/*.spec.* already matched those at any nested path, so dropping the **/ prefix only extends the same false-positive class to the repo root. It is accepted rather than narrowed, because narrowing it is a behaviour change to the currently-supported case and not part of making other languages visible.

Known gap — Rust (#4379). #[test] conventionally lives inside the implementation file, so a Rust RED commit touches src/*.rs and no path-based gate can distinguish it from ordinary source. Widening the pathspec to cover it would match all source and make the gate meaningless. cargo test works; the RED-commit gate cannot see it, so a Rust project using workflow.tdd_mode should expect the gate to trip.

Framework setup is a one-time cost included in the first TDD plan's RED phase. </framework_setup>

<error_handling>

Error Handling

Test doesn't fail in RED phase:

  • Feature may already exist - investigate
  • Test may be wrong (not testing what you think)
  • Fix before proceeding

Test doesn't pass in GREEN phase:

  • Debug implementation
  • Don't skip to refactor
  • Keep iterating until green

Tests fail in REFACTOR phase:

  • Undo refactor
  • Commit was premature
  • Refactor in smaller steps

Unrelated tests break:

  • Stop and investigate
  • May indicate coupling issue
  • Fix before proceeding </error_handling>

<commit_pattern>

Commit Pattern for TDD Plans

TDD plans produce 2-3 atomic commits (one per phase):

test(08-02): add failing test for email validation

- Tests valid email formats accepted
- Tests invalid formats rejected
- Tests empty input handling

feat(08-02): implement email validation

- Regex pattern matches RFC 5322
- Returns boolean for validity
- Handles edge cases (empty, null)

refactor(08-02): extract regex to constant (optional)

- Moved pattern to EMAIL_REGEX constant
- No behavior changes
- Tests still pass

Comparison with standard plans:

  • Standard plans: 1 commit per task, 2-4 commits per plan
  • TDD plans: 2-3 commits for single feature

Both follow same format: {type}({phase}-{plan}): {description}

Benefits:

  • Each commit independently revertable
  • Git bisect works at commit level
  • Clear history showing TDD discipline
  • Consistent with overall commit strategy </commit_pattern>

<gate_enforcement>

Gate Enforcement Rules

When workflow.tdd_mode is enabled in config, the RED/GREEN/REFACTOR gate sequence is enforced for all type: tdd plans.

Gate Definitions

Gate Required Commit Pattern Validation
RED Yes test({phase}-{plan}): ... Test exists AND fails before implementation — intentionally: check tdd-red-evidence returns RED_EVIDENCE_OK (target test failed on an assertion for the behavior; anything else is INVALID_RED)
GREEN Yes feat({phase}-{plan}): ... Test passes after implementation
REFACTOR No refactor({phase}-{plan}): ... Tests still pass after cleanup

Fail-Fast Rules

  1. Unexpected GREEN in RED phase: If the test passes before any implementation code is written, STOP. The feature may already exist or the test is wrong. Investigate before proceeding.
  2. INVALID_RED in RED phase (#3770): A nonzero exit is not RED by itself. Zero-test discovery, fixture/load crashes, nonzero exits with no failing test, unrelated failing tests, and unexpected greens all classify as INVALID_RED (gsd_run check tdd-red-evidence). STOP and fix the RED phase — do NOT proceed to GREEN.
  3. Missing RED commit: If no test(...) commit precedes the feat(...) commit, the TDD discipline was violated. Flag in SUMMARY.md.
  4. REFACTOR breaks tests: Undo the refactor immediately. Commit was premature — refactor in smaller steps.

Executor Gate Validation

After completing a type: tdd plan, the executor validates the git log:

# The commit protocol promises no zero-padding for ${PHASE}/${PLAN} — strip both and
# match the commit-scope position anchored (#4003). #4619: PHASE may be decimal/
# N-segment; zero-strip only the leading integer segment, escape the rest.
# #4748: it may also carry a letter suffix (03A), so split at the first non-digit.
PHASE_INT=${PHASE%%[!0-9]*}; PHASE_REST=${PHASE#"$PHASE_INT"}
PHASE_N="$((10#$PHASE_INT))${PHASE_REST//./\\.}"
PLAN_N=$((10#${PLAN}))
# Check for RED gate commit
git log --oneline -E --grep="^test\((0*${PHASE_N})-(0*${PLAN_N})\):" | head -1
# Check for GREEN gate commit  
git log --oneline -E --grep="^feat\((0*${PHASE_N})-(0*${PLAN_N})\):" | head -1
# Check for optional REFACTOR gate commit
git log --oneline -E --grep="^refactor\((0*${PHASE_N})-(0*${PLAN_N})\):" | head -1

If RED or GREEN gate commits are missing, add a ## TDD Gate Compliance section to SUMMARY.md with the violation details. </gate_enforcement>

<end_of_phase_review>

End-of-Phase TDD Review Checkpoint

When workflow.tdd_mode is enabled, the execute-phase orchestrator inserts a collaborative review checkpoint after all waves complete but before phase verification.

Review Checkpoint Format

### TDD REVIEW — Phase {X}

TDD Plans: {count} | Gate violations: {count}

| Plan | RED | GREEN | REFACTOR | Status |
|------|-----|-------|----------|--------|
| {id} |  ✓  |   ✓   |    ✓     | Pass   |
| {id} |  ✓  |   ✗   |    —     | FAIL   |

{If violations exist:}
⚠ Gate violations are advisory — review before advancing.

What the Review Checks

  1. Gate sequence: Each TDD plan has RED → GREEN commits in order
  2. Test quality: RED phase tests fail for the right reason (not import errors or syntax)
  3. Minimal GREEN: Implementation is minimal — no premature optimization in GREEN phase
  4. Refactor discipline: If REFACTOR commit exists, tests still pass

This checkpoint is advisory — it does not block phase completion but surfaces TDD discipline issues for human review. </end_of_phase_review>

<context_budget>

Context Budget

TDD plans target ~40% context usage (lower than standard plans' ~50%).

Why lower:

  • RED phase: write test, run test, potentially debug why it didn't fail
  • GREEN phase: implement, run test, potentially iterate on failures
  • REFACTOR phase: modify code, run tests, verify no regressions

Each phase involves reading files, running commands, analyzing output. The back-and-forth is inherently heavier than linear task execution.

Single feature focus ensures full quality throughout the cycle. </context_budget>