* fix(#4619): execute-phase computes decimal/N-segment phase numbers without breaking shell arithmetic $((10#${PHASE_NUMBER})) is a hard bash/zsh syntax error when PHASE_NUMBER is decimal (01.1, from an inserted phase) or N-segment (23.1.2) — neither is valid shell-arithmetic syntax at all, and the failed expansion aborts the rest of the snippet in a non-interactive shell. safe_resume_gate runs unconditionally before trusting STATE.md or dispatching any executor, so execute-phase failed at its own gate before the first executor on any decimal phase, regardless of workflow.tdd_mode. Regression from #4194. Fixes all 4 sites: safe_resume_gate and the TDD gate in workflows/execute-phase.md, the completion-signal spot-check fallback in workflows/execute-phase/steps/completion-reconciliation.md, and the executor gate validation example in references/tdd.md. Each now zero-strips only the leading integer segment into a *_INT variable (via %%.* / # parameter expansion — always valid shell syntax regardless of what follows) and keeps the remainder as an escaped-dot string for the anchored commit- scope regex, exactly as issue #4619 verified in both bash and zsh. A plain integer phase (12, 01) computes byte-identically to before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4619): pin the decimal/N-segment fix and characterize the pre-fix bug Behavioral coverage via real bash execution: the old $((10#01.1)) form throws (characterizes the bug, matching the issue's own reproduction); the new form resolves 01.1 -> 1\.1 and 23.1.2 -> 23\.1\.2, unchanged for plain integers (12 -> 12, 01 -> 1); the resulting anchored ERE matches feat(01.1-03):/test(1.1-3): and correctly rejects feat(01-03):, feat(01.2-03):, feat(011-03):, feat(12-03): for a decimal phase — mirroring issue #4619's own verified table exactly. Updates safe-resume-gate-anchoring.test.cjs's 4 existing source-text assertions (one per site) to the new fixed text. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4634): refine the shell-arith drift detector to distinguish safe from unsafe arithmetic With #4619's fix in place, the guard's original "ban $((10#... outright, match any occurrence" was too blunt: it flagged a comment merely mentioning the pattern in prose, the now-safe $((10#$PHASE_INT)) arithmetic on an already-%%.*-stripped integer, and the always-safe plan-id arithmetic (plan ids are plain integers, never decimal). Refines the detector to skip full-line comments and to only flag a captured variable/placeholder name that contains "phase" and does NOT end in _INT/_int — the naming convention the #4619 fix establishes at all four sites for "already reduced to a safe integer." A plan-id variable was never phase-number arithmetic in the first place and is excluded on the same basis. This closes epic #4634's D6 ("lint-phase-id-drift... passes with no new exemptions") and D7 ("a decimal and N-segment phase id survive an end-to-end execute-phase selection without error") for real — the guard now reports zero violations across all five .cts/.md rules. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore: regenerate conformance-tier manifests for the new test file Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4619): cover the plain-padded-integer near-miss matrix too Review found the anchored-ERE near-miss coverage only exercised the decimal case (PHASE_NUMBER=01.1); issue #4619's own worked table also verifies the plain padded-integer case (01 -> PHASE_N=1) against its own near-miss set (matches 01-03, rejects 01.1-03/011-03/12-03). Adds the missing assertion. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4619): add Fixed changeset Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4619): correct JS backslash-escaping in safe-resume-gate anchoring test The test's string-literal assertions for the PHASE_FRAC//./\\.} pattern wrote only 2 backslash characters in JS source, which single-quoted-string parsing collapses to 1 real backslash at runtime -- but the workflow/reference files actually contain 2 raw backslash bytes at that position (needed so bash's ${var//pattern/replacement} produces the correct single-backslash output). Write 4 backslash characters in the JS source at all 4 occurrences so the runtime string matches the files' real bytes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4619): refresh the committed compact-content benchmark baseline The new PHASE_INT/PHASE_FRAC arithmetic lines added to gsd-core/workflows/execute-phase.md shifted its committed compaction-ratio baseline. Regenerate via `node scripts/benchmark-compact-content.cjs --write`. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4619): note the safe_resume_gate arithmetic growth in the test header The emitted-attribution gate flags execute-phase.md growing 91253 -> 91846 bytes (593 bytes). The growth is the fix: the safe_resume_gate and TDD RED block now derive PHASE_INT/PHASE_FRAC before computing PHASE_N, so a decimal/N-segment phase number (e.g. 01.1, 2.3.1) zero-strips its leading integer segment via base-10 arithmetic instead of forcing the whole value through $((10#...)) and hitting a hard shell syntax error on the first dot. A blank line previously separated the Emitted-Drift-Ack-Growth trailer from the Co-Authored-By trailer below it, which splits git's trailer-block detection: only the last contiguous non-blank run of Key: Value lines at the end of a commit message is recognized as trailers, so the growth ack was silently read as ordinary body text and the differential-attribution gate failed with the growth unacknowledged. Joining the two trailers into one contiguous block fixes it. Emitted-Drift-Ack-Growth: execute-phase.md — adds PHASE_INT/PHASE_FRAC derivation to the safe_resume_gate and TDD RED commit-scope grep so a decimal/N-segment phase number zero-strips its leading integer segment via base-10 arithmetic instead of failing on a non-numeric value (#4619) Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4208): replace chmod-based restore-failure injection with a root-proof git shim `tests/commit-files-deletion.test.cjs`'s two restore-failure tests simulated an unwritable index via a `post-index-change` hook running `chmod a-w` on the git dir. That relies on the OS enforcing the *owner's own* permission bits against itself, which uid 0 (a routine identity inside this repo's Docker-based gsd-test benches) does not: every DAC check short-circuits true for root, so the write the chmod meant to block silently succeeds, the restore comes back clean, and the disclosure/rollback behavior under test never actually gets exercised. This is CLAUDE.md's own named anti-pattern for I/O-failure injection ("Cross-platform test IO-failure injection" — chmod tricks fail under root Docker/CI). It is confirmed as the actual root cause here, not a production defect: `src/commands.cts`'s `restoreRemovedEntries`/rollback-disclosure logic (added by #4253, merged just before this run) was hand-traced and manually reproduced end to end on an unprivileged workstation against a freshly built `gsd-core/bin/lib/commands.cjs`, and it already produces exactly the `staging_failed` + "could not be restored" / "could NOT be restored during rollback" results both tests assert. The other `post-index-change`-based tests in this file (a `sleep` to force a timeout; a real `update-index` to flip a restored entry's mode) are unaffected because neither depends on a permission check — consistent with only the two chmod-based tests failing on the real remote run. Replaces the chmod fixture with a fake `git` placed ahead of the real one on PATH that fails only `update-index --add --cacheinfo` — the one call the restore makes — unconditionally, regardless of privilege level. Every other git invocation execs straight through to the real binary, so the rest of each scenario (`rm --cached`, the restore's own `ls-files` verification, etc.) is exercised exactly as before. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4619): backfill changeset pr number to 4644 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4619): feed the bash fixture script via stdin, not argv, to fix Windows CI Passing the script as a `-c "<script>"` argv element made it subject to Windows' CreateProcess command-line argument encoding, which silently dropped the escaped-dot backslashes before bash ever saw them (observed on PR #4644's windows-latest CI shard: `1\.1` came back as `1.1`). Feeding the same script via stdin instead removes argv entirely from the transport, so there is nothing for Windows to re-encode. POSIX behavior is unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
12 KiB
Principle: If you can describe the behavior as expect(fn(input)).toBe(output) before writing fn, TDD improves the result.
Key insight: TDD work is fundamentally heavier than standard tasks—it requires 2-3 execution cycles (RED → GREEN → REFACTOR), each with file reads, test runs, and potential debugging. TDD features get dedicated plans to ensure full context is available throughout the cycle.
<when_to_use_tdd>
When TDD Improves Quality
TDD candidates (create a TDD plan):
- Business logic with defined inputs/outputs
- API endpoints with request/response contracts
- Data transformations, parsing, formatting
- Validation rules and constraints
- Algorithms with testable behavior
- State machines and workflows
- Utility functions with clear specifications
Skip TDD (use standard plan with type="auto" tasks):
- UI layout, styling, visual components
- Configuration changes
- Glue code connecting existing components
- One-off scripts and migrations
- Simple CRUD with no business logic
- Exploratory prototyping
Heuristic: Can you write expect(fn(input)).toBe(output) before writing fn?
→ Yes: Create a TDD plan
→ No: Use standard plan, add tests after if needed
</when_to_use_tdd>
<tdd_plan_structure>
TDD Plan Structure
Each TDD plan implements one feature through the full RED-GREEN-REFACTOR cycle.
---
phase: XX-name
plan: NN
type: tdd
---
<objective>
[What feature and why]
Purpose: [Design benefit of TDD for this feature]
Output: [Working, tested feature]
</objective>
<context>
@.planning/PROJECT.md
@.planning/ROADMAP.md
@relevant/source/files.ts
</context>
<feature>
<name>[Feature name]</name>
<files>[source file, test file]</files>
<behavior>
[Expected behavior in testable terms]
Cases: input → expected output
</behavior>
<implementation>[How to implement once tests pass]</implementation>
</feature>
<verification>
[Test command that proves feature works]
</verification>
<success_criteria>
- Failing test written and committed
- Implementation passes test
- Refactor complete (if needed)
- All 2-3 commits present
</success_criteria>
<output>
After completion, create SUMMARY.md with:
- RED: What test was written, why it failed
- GREEN: What implementation made it pass
- REFACTOR: What cleanup was done (if any)
- Commits: List of commits produced
</output>
One feature per TDD plan. If features are trivial enough to batch, they're trivial enough to skip TDD—use a standard plan and add tests after. </tdd_plan_structure>
<execution_flow>
Red-Green-Refactor Cycle
RED - Write failing test:
- Create test file following project conventions
- Write test describing expected behavior (from
<behavior>element) - Run test - it MUST fail intentionally (#3770): the TARGET test you named must be the test that fails, on an assertion for the planned behavior. A nonzero exit alone is NOT RED — syntax errors, zero-test discovery, fixture crashes, parser errors, and unrelated assertions are INVALID_RED and must not authorize GREEN.
- Persist the RED evidence record (command, exit code, failing test, expected result, actual result) and verify it:
gsd_run check tdd-red-evidence <record.json>. Only verdictRED_EVIDENCE_OKsatisfies the RED gate;INVALID_REDblocks GREEN until the RED phase is fixed. - If test passes: feature exists or test is wrong. Investigate.
- Commit:
test({phase}-{plan}): add failing test for [feature]
GREEN - Implement to pass:
- Write minimal code to make test pass
- No cleverness, no optimization - just make it work
- Run test - it MUST pass
- Commit:
feat({phase}-{plan}): implement [feature]
REFACTOR (if needed):
- Clean up implementation if obvious improvements exist
- Run tests - MUST still pass
- Only commit if changes made:
refactor({phase}-{plan}): clean up [feature]
Result: Each TDD plan produces 2-3 atomic commits. </execution_flow>
<test_quality>
Good Tests vs Bad Tests
Test behavior, not implementation:
- Good: "returns formatted date string"
- Bad: "calls formatDate helper with correct params"
- Tests should survive refactors
One concept per test:
- Good: Separate tests for valid input, empty input, malformed input
- Bad: Single test checking all edge cases with multiple assertions
Descriptive names:
- Good: "should reject empty email", "returns null for invalid ID"
- Bad: "test1", "handles error", "works correctly"
No implementation details:
- Good: Test public API, observable behavior
- Bad: Mock internals, test private methods, assert on internal state </test_quality>
<framework_setup>
Test Framework Setup (If None Exists)
When executing a TDD plan but no test framework is configured, set it up as part of the RED phase:
1. Detect project type:
# JavaScript/TypeScript
if [ -f package.json ]; then echo "node"; fi
# Python
if [ -f requirements.txt ] || [ -f pyproject.toml ]; then echo "python"; fi
# Go
if [ -f go.mod ]; then echo "go"; fi
# Rust
if [ -f Cargo.toml ]; then echo "rust"; fi
2. Install minimal framework:
| Project | Framework | Install |
|---|---|---|
| Node.js | Jest | npm install -D jest @types/jest ts-jest |
| Node.js (Vite) | Vitest | npm install -D vitest |
| Python | pytest | pip install pytest |
| Go | testing | Built-in |
| Rust | cargo test | Built-in |
3. Create config if needed:
- Jest:
jest.config.jswith ts-jest preset - Vitest:
vitest.config.tswith test globals - pytest:
pytest.iniorpyproject.tomlsection
4. Verify setup:
# Run empty test suite - should pass with 0 tests
npm test # Node
pytest # Python
go test ./... # Go
cargo test # Rust
5. Create first test file: Follow project conventions for test location:
*.test.ts/*.spec.tsnext to source__tests__/directorytests/directory at root
Framework setup is a one-time cost included in the first TDD plan's RED phase. </framework_setup>
<error_handling>
Error Handling
Test doesn't fail in RED phase:
- Feature may already exist - investigate
- Test may be wrong (not testing what you think)
- Fix before proceeding
Test doesn't pass in GREEN phase:
- Debug implementation
- Don't skip to refactor
- Keep iterating until green
Tests fail in REFACTOR phase:
- Undo refactor
- Commit was premature
- Refactor in smaller steps
Unrelated tests break:
- Stop and investigate
- May indicate coupling issue
- Fix before proceeding </error_handling>
<commit_pattern>
Commit Pattern for TDD Plans
TDD plans produce 2-3 atomic commits (one per phase):
test(08-02): add failing test for email validation
- Tests valid email formats accepted
- Tests invalid formats rejected
- Tests empty input handling
feat(08-02): implement email validation
- Regex pattern matches RFC 5322
- Returns boolean for validity
- Handles edge cases (empty, null)
refactor(08-02): extract regex to constant (optional)
- Moved pattern to EMAIL_REGEX constant
- No behavior changes
- Tests still pass
Comparison with standard plans:
- Standard plans: 1 commit per task, 2-4 commits per plan
- TDD plans: 2-3 commits for single feature
Both follow same format: {type}({phase}-{plan}): {description}
Benefits:
- Each commit independently revertable
- Git bisect works at commit level
- Clear history showing TDD discipline
- Consistent with overall commit strategy </commit_pattern>
<gate_enforcement>
Gate Enforcement Rules
When workflow.tdd_mode is enabled in config, the RED/GREEN/REFACTOR gate sequence is enforced for all type: tdd plans.
Gate Definitions
| Gate | Required | Commit Pattern | Validation |
|---|---|---|---|
| RED | Yes | test({phase}-{plan}): ... |
Test exists AND fails before implementation — intentionally: check tdd-red-evidence returns RED_EVIDENCE_OK (target test failed on an assertion for the behavior; anything else is INVALID_RED) |
| GREEN | Yes | feat({phase}-{plan}): ... |
Test passes after implementation |
| REFACTOR | No | refactor({phase}-{plan}): ... |
Tests still pass after cleanup |
Fail-Fast Rules
- Unexpected GREEN in RED phase: If the test passes before any implementation code is written, STOP. The feature may already exist or the test is wrong. Investigate before proceeding.
- INVALID_RED in RED phase (#3770): A nonzero exit is not RED by itself. Zero-test discovery, fixture/load crashes, nonzero exits with no failing test, unrelated failing tests, and unexpected greens all classify as INVALID_RED (
gsd_run check tdd-red-evidence). STOP and fix the RED phase — do NOT proceed to GREEN. - Missing RED commit: If no
test(...)commit precedes thefeat(...)commit, the TDD discipline was violated. Flag in SUMMARY.md. - REFACTOR breaks tests: Undo the refactor immediately. Commit was premature — refactor in smaller steps.
Executor Gate Validation
After completing a type: tdd plan, the executor validates the git log:
# The commit protocol promises no zero-padding for ${PHASE}/${PLAN} — strip both and
# match the commit-scope position anchored (#4003). #4619: PHASE may be decimal/
# N-segment; zero-strip only the leading integer segment, escape the rest.
PHASE_INT=${PHASE%%.*}; PHASE_FRAC=${PHASE#"$PHASE_INT"}
PHASE_N="$((10#$PHASE_INT))${PHASE_FRAC//./\\.}"
PLAN_N=$((10#${PLAN}))
# Check for RED gate commit
git log --oneline -E --grep="^test\((0*${PHASE_N})-(0*${PLAN_N})\):" | head -1
# Check for GREEN gate commit
git log --oneline -E --grep="^feat\((0*${PHASE_N})-(0*${PLAN_N})\):" | head -1
# Check for optional REFACTOR gate commit
git log --oneline -E --grep="^refactor\((0*${PHASE_N})-(0*${PLAN_N})\):" | head -1
If RED or GREEN gate commits are missing, add a ## TDD Gate Compliance section to SUMMARY.md with the violation details.
</gate_enforcement>
<end_of_phase_review>
End-of-Phase TDD Review Checkpoint
When workflow.tdd_mode is enabled, the execute-phase orchestrator inserts a collaborative review checkpoint after all waves complete but before phase verification.
Review Checkpoint Format
### TDD REVIEW — Phase {X}
TDD Plans: {count} | Gate violations: {count}
| Plan | RED | GREEN | REFACTOR | Status |
|------|-----|-------|----------|--------|
| {id} | ✓ | ✓ | ✓ | Pass |
| {id} | ✓ | ✗ | — | FAIL |
{If violations exist:}
⚠ Gate violations are advisory — review before advancing.
What the Review Checks
- Gate sequence: Each TDD plan has RED → GREEN commits in order
- Test quality: RED phase tests fail for the right reason (not import errors or syntax)
- Minimal GREEN: Implementation is minimal — no premature optimization in GREEN phase
- Refactor discipline: If REFACTOR commit exists, tests still pass
This checkpoint is advisory — it does not block phase completion but surfaces TDD discipline issues for human review. </end_of_phase_review>
<context_budget>
Context Budget
TDD plans target ~40% context usage (lower than standard plans' ~50%).
Why lower:
- RED phase: write test, run test, potentially debug why it didn't fail
- GREEN phase: implement, run test, potentially iterate on failures
- REFACTOR phase: modify code, run tests, verify no regressions
Each phase involves reading files, running commands, analyzing output. The back-and-forth is inherently heavier than linear task execution.
Single feature focus ensures full quality throughout the cycle. </context_budget>