* test(#3884): failing-first coverage for strict argv and absence-signalling --pick ADR-3473 §8.4 says failure is a value. Three families currently encode failure as success, and this commit pins each one RED before the fix lands. Measured on this tree, 2026-08-26: gsd-tools generate-slug "test" --pick nonexistent -> empty stdout, exit 0 (#3365) gsd-tools audit-open --pick nonexistent_field -> dumps the entire human-readable audit report, exit 0 gsd-tools generate-slug "Hello World" --raw --pick bogus -> prints "hello-world", another field's value, exit 0 gsd-tools query state.planned-phase 3 (positional, no --phase) -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted current_phase_name (#3358) tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract ("returns empty string for missing field", success === true). That assertion is replaced by the required behavior rather than deleted. The new parseNamedArgs block calls the spec-object signature that does not exist yet, so it fails today by construction. The 11 existing behavior-lock tests are left untouched here; they are corrected in the implementation commit. C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180 Decision 4(b). A unit assertion on the parser would have passed throughout this defect's life. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3884): failure is a value — strict argv, and --pick that signals absence Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable ways to say "I could not answer". parseNamedArgs (src/command-arg-projection.cts) Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the hub's Result shape instead of a bare Record. Declaring the positional arity is what makes #3358's call site unrepresentable rather than merely detectable: an unrecognized flag or a token past the declared boundary is now InvalidArgs, naming the offending token and listing the accepted flags. The legacy positional-array call shape throws a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale hand-written .cjs call site fails loudly instead of destructuring undefined off a Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a projection over the one parser, not a second parser. Measured before, against a STATE.md with a populated phase-2 block: query state.planned-phase 3 (positional, no --phase) -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted current_phase_name After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical. The flag form is unchanged and still updates STATE.md. --pick <field> (gsd-core/bin/gsd-tools.cjs) extractField returns {found,value}, and the pick block no longer shares one catch between "output was not JSON" and "field was absent". An absent field exits 1 with pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1 with pick_output_not_json instead of dumping the command's entire output. A field that is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer, not a failure, and it is what keeps `--pick count` printing 0 on a fresh project. Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" — a different field's value, confidently, at exit 0. ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0 would demote "could not answer" to "the answer is zero" — the hazard docs/how-to/resolve-unreachable-guard-findings.md already warns against. Guard ledger (ADR-3473 Decision 6) scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls, a nullglob mechanism this change does not touch) is retained in full, as are the shared scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file is not deleted. Call-site audit 45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an if test, && chain, or a pipeline whose status is consumed, and no shell block in workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind a prior found/existence check. No ADR-3409-class "field the command never produces" remains. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows Two review findings, both fixed here rather than recorded as limits. 1. A newline in an untrusted token forged a second stderr line. Before, plain-text mode: $ gsd-tools query state.planned-phase $'foo\nError: forged second line' Error: unexpected positional argument "foo Error: forged second line" After: Error: unexpected positional argument "foo\nError: forged second line" --json-errors mode was never affected — io.error runs that payload through JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the three new InvalidArgs reasons plus the two new --pick diagnostics all interpolate a token that comes straight from argv. Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at every interpolation site — not a copy per site. It is deliberately NOT applied inside error() itself: several callers in this tree emit intentional multi-line diagnostics, and escaping newlines there would mangle them. The available-top-level-keys list needed the same treatment for a reason the review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user document and echoes that document's own keys into the diagnostic. Verified reachable — a frontmatter key containing a newline reaches the key list — so formatKeyForDiagnosticList is guarding a live path, not a hypothetical one. Ordinary keys still render plain and unquoted; a fix that merely dropped the key would also have passed a "one line" assertion, so the test pins the escaped key's presence too. 2. Five behavior-table rows were implemented but nothing pinned them: B7 a dotted path that dies partway B9 bracket syntax on a non-array B10 a negative array index, in and out of range B14 a JSON root that is not an object B17 an @file: payload over 50KB B17 is the load-bearing one. output() writes @file:<path> instead of inline JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no test, a future reordering of those two steps turns every large result into a false pick_output_not_json. The fixture seeds 1200 phase directories and measures the payload at 62474 characters, asserting the spill actually happened rather than assuming it. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): correct the strict-argv surface against a full verification run The first full run came back with 90 failures across 12 files, none in the new tests. They were the argv surface telling me what it actually is. Ten root causes; each classified before anything was changed. I over-implemented, and that is reverted. ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional tokens". It says nothing about a value flag whose value is missing. Making that an error was my design decision, not the rule, and it broke a deliberately recorded contract: `--prd` with no value resolving to null (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5; tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The "requires a value" branch is deleted outright rather than kept behind an option — an unused strictness mode is speculative generality. Unknown-flag and unexpected-positional rejection, which is what §8.4 actually mandates, is unchanged. --wave needed a third flag kind the original design did not anticipate. `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the shipped workflow reconstructs and passes it (execute-phase.md:84), while #2932 records token-PRESENCE semantics: the CLI cares only that the flag appeared, and the value belongs to the workflow layer. That is neither a boolean flag nor a value flag, so `optionalValueFlags` now exists — presence-only in `data`, and the validation cursor consumes a following non-flag token so it is not reported as a stray positional. Every other declared boolean flag was checked against every argument-hint and prose usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one of this shape. Five tests were pinning forms that never worked. tests/adr857-core-without-capabilities.test.cjs passed `init plan-phase --phase 01-stub`, but the documented form is positional (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form is the literal string "--phase". Measured on the pre-fix build against a real .planning/phases/01-stub/ directory: init plan-phase 01-stub -> phase_found=true init plan-phase --phase 01-stub -> phase_found=false The test asserted only exit 0 and key presence, so it had been green while proving nothing about phase resolution. Corrected to the documented form and strengthened to assert phase_found === true. Same class in state.test.cjs (`--plan-count`, a flag that does not exist; the real one is `--plans`), milestone-archive.test.cjs (`init new-milestone --json`, silently ignored), and concurrency-safety.test.cjs (a bare positional field name whose OR-assertion passed because a whole-document dump happens to contain the substring it looked for). Six handlers had no argv validation at all — the same #3358 shape this phase exists to close, found while fixing the rest: init verify-work / phase-op / review / todos / remove-workspace read args[2] with nothing checking the rest, and validate health read --repair/--backfill through a bare args.includes() scan that bypassed the parser entirely. All now go through the seam, so the flag has one owner. tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's local expectation does not override §8, so they are inverted and renamed — a test still called "ignores an unrecognized flag" while asserting rejection would be its own defect. Row C6's point is its PWNED canary; that assertion is kept verbatim and only its exit-status expectation changed, because the hostile token is now rejected rather than absorbed. The blast-radius estimate in 40-design.md is corrected rather than quietly left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was accurate for what the graph can see — parseNamedArgs's callers. It cannot see that those callers' handlers accept argv shapes wider than the code reading args[2] suggests, which is where the real surface was. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert Second full run: 46 failures, down from 90. Four causes, two of them mine. Reverted `validate health` entirely — it was scope creep, and it broke a real flag. ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`. The previous commit routed `validate health` through the parser on the reasoning that a flag should have one owner. That was wrong twice over: §8.4 names parseNamedArgs and count queries, and `validate health` was never a parseNamedArgs call site — it read its flags, just not through the parser, so it had no silent-drop defect to fix. Tightening it omitted `--json`, which the health-diagnostic suites use heavily. The handler is now byte-for-behaviour back to its pre-branch form. `validate context` stays converted: it genuinely was a call site, and its `--json` is now declared rather than read by a second `args.includes` scan. The five handlers that had NO validation at all — init verify-work / phase-op / review / todos / remove-workspace — stay fixed. Those read args[2] with nothing checking the rest, which is the #3358 shape this phase owns. Finished the A2/A3 revert. Three tests still encoded the deleted "a value flag with a missing value is an error" rule, including one added by the previous commit for that rule. All three now assert the reverted null contract, and the ones whose titles said "rejected" are renamed — a test named for a contract it no longer asserts is its own defect. `--wave=` and `--wave --weird` are correctly rejected. Neither is documented in commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and neither is emitted by the shipped prompt layer, so both are unrecognized tokens that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the property it exists for — asserted directly now, at the parser, that `--wave` does not swallow a following flag as its value — and only its exit-status expectation changed. A contradiction inside this branch, surfaced by the audit and resolved the safe way. Two pre-existing #3573 tests call `state begin-phase '2'` and `state planned-phase '2'` with a bare positional, relying on the old permissive parser to ignore it. This branch's own #3358 regression test requires that exact argv to be REJECTED. The two are mutually exclusive. Widening the router to accept a bare positional — mirroring complete-phase — would have silently re-opened #3358, and was verified to do exactly that: with the widened router, `query state.planned-phase 3` returned exit 0 and wrote current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the two #3573 tests move to it. Their assertions were never about the call shape — only that total_phases survives the resync — and both still pass. complete-phase is untouched: its bare positional IS documented, and it keeps the dynamic boundary and the negative-space note that record why. The audit that produced this is in the PR body: for every handler whose declaration changed, the flags it reads anywhere in its body, the flags the shipped surface documents, and the shapes the suite passes, compared. The `--json` miss was a pattern, not an accident — declaring a handler's flags from its parseNamedArgs call alone misses whatever it reads elsewhere. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3884): backfill the changeset PR number Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
687 lines
30 KiB
JavaScript
687 lines
30 KiB
JavaScript
// Reads .md/.json/.yml product files whose deployed text IS what the
|
|
// runtime loads — testing text content tests the deployed contract.
|
|
//
|
|
// docs-guard-exempt: #3884 — the docs/CLI-TOOLS.md:736 citation below is an
|
|
// explanatory comment pointing at documented CLI shape, not a read target;
|
|
// this file performs unrelated filesystem reads (STATE.md, phase/plan
|
|
// fixtures under .planning/) and never reads any docs/ file.
|
|
|
|
/**
|
|
* GSD Tools Tests - Concurrency Safety
|
|
*
|
|
* Tests for fix/concurrency-safety-1473a:
|
|
* - Planning lock integration (withPlanningLock in phase/roadmap operations)
|
|
* - readModifyWriteStateMd (atomic state updates)
|
|
* - normalizeMd behavioral equivalence (O(n) insideFence rewrite)
|
|
* - Warnings (frontmatter parse warning, stateReplaceFieldWithFallback)
|
|
* - Performance benchmarks (normalizeMd O(n) verification)
|
|
* - Snapshot tests for normalizeMd (regression detection)
|
|
* - Multi-process concurrent write tests
|
|
* - Stress tests at scale (50+ phases)
|
|
*/
|
|
|
|
const { test, describe, beforeEach, afterEach } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('fs');
|
|
const path = require('path');
|
|
const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs');
|
|
|
|
const { normalizeContent } = require('../gsd-core/bin/lib/shell-command-projection.cjs');
|
|
// normalizeMd was removed from core.cjs (Phase 4 — issue #3468); the same algorithm now
|
|
// lives in the shell-command-projection seam. Wrap normalizeContent so existing
|
|
// behavioral / snapshot / perf assertions stay point-of-truth.
|
|
const normalizeMd = (input) => normalizeContent('test.md', input).content;
|
|
|
|
// ─── Helpers ────────────────────────────────────────────────────────────────
|
|
|
|
|
|
function writeMinimalStateMd(tmpDir, content) {
|
|
const defaultContent = content || `# Session State\n\n## Current Position\n\nPhase: 1\n`;
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
defaultContent
|
|
);
|
|
}
|
|
|
|
|
|
/**
|
|
* Generate a 50-phase project structure for stress testing.
|
|
*/
|
|
function create50PhaseProject(tmpDir, completedCount = 25) {
|
|
let roadmapContent = '# Roadmap v1.0\n\n';
|
|
for (let i = 1; i <= 50; i++) {
|
|
roadmapContent += `- [${i <= completedCount ? 'x' : ' '}] Phase ${i}: Feature ${i}\n`;
|
|
}
|
|
roadmapContent += '\n';
|
|
for (let i = 1; i <= 50; i++) {
|
|
const pad = String(i).padStart(2, '0');
|
|
roadmapContent += `### Phase ${i}: Feature ${i}\n\n`;
|
|
roadmapContent += `**Goal:** Build feature ${i}\n`;
|
|
roadmapContent += `**Requirements:** REQ-${pad}\n`;
|
|
roadmapContent += `**Plans:** 1 plans\n\n`;
|
|
roadmapContent += `Plans:\n- [${i <= completedCount ? 'x' : ' '}] ${pad}-01-PLAN.md\n\n`;
|
|
}
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'ROADMAP.md'),
|
|
roadmapContent
|
|
);
|
|
|
|
const phasesDir = path.join(tmpDir, '.planning', 'phases');
|
|
for (let i = 1; i <= 50; i++) {
|
|
const pad = String(i).padStart(2, '0');
|
|
const dirName = `${pad}-feature-${i}`;
|
|
const phaseDir = path.join(phasesDir, dirName);
|
|
fs.mkdirSync(phaseDir, { recursive: true });
|
|
fs.writeFileSync(
|
|
path.join(phaseDir, `${pad}-01-PLAN.md`),
|
|
`# Phase ${i} Plan 1\n\nBuild feature ${i}.\n`
|
|
);
|
|
if (i <= completedCount) {
|
|
fs.writeFileSync(
|
|
path.join(phaseDir, `${pad}-01-SUMMARY.md`),
|
|
`# Phase ${i} Plan 1 Summary\n\nFeature ${i} completed.\n`
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// 1. Planning lock integration
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
describe('planning lock integration', () => {
|
|
let tmpDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = createTempProject();
|
|
});
|
|
|
|
afterEach(() => {
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
test('phase add creates and releases .planning/.lock during ROADMAP write', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'ROADMAP.md'),
|
|
`# Roadmap v1.0\n\n### Phase 1: Foundation\n**Goal:** Setup\n\n---\n`
|
|
);
|
|
|
|
const result = runGsdTools('phase add Testing', tmpDir);
|
|
assert.ok(result.success, `Command failed: ${result.error}`);
|
|
|
|
const lockPath = path.join(tmpDir, '.planning', '.lock');
|
|
assert.ok(!fs.existsSync(lockPath), '.lock file should be released after phase add');
|
|
|
|
const output = JSON.parse(result.output);
|
|
assert.strictEqual(output.phase_number, 2, 'should be phase 2');
|
|
});
|
|
|
|
test('phase complete creates and releases .planning/.lock', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'ROADMAP.md'),
|
|
`# Roadmap\n\n- [ ] Phase 1: Foundation\n\n### Phase 1: Foundation\n**Goal:** Setup\n**Plans:** 1 plans\n\n### Phase 2: API\n**Goal:** Build\n`
|
|
);
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
`# State\n\n**Current Phase:** 01\n**Current Phase Name:** Foundation\n**Status:** In progress\n**Current Plan:** 01-01\n**Last Activity:** 2025-01-01\n**Last Activity Description:** Working\n`
|
|
);
|
|
|
|
const p1 = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
|
fs.mkdirSync(p1, { recursive: true });
|
|
fs.writeFileSync(path.join(p1, '01-01-PLAN.md'), '# Plan');
|
|
fs.writeFileSync(path.join(p1, '01-01-SUMMARY.md'), '# Summary');
|
|
fs.writeFileSync(
|
|
path.join(p1, '01-VERIFICATION.md'),
|
|
'---\nstatus: passed\nscore: "1/1"\n---\n# Verification\nPassed.\n',
|
|
);
|
|
fs.mkdirSync(path.join(tmpDir, '.planning', 'phases', '02-api'), { recursive: true });
|
|
|
|
const result = runGsdTools('phase complete 1', tmpDir);
|
|
assert.ok(result.success, `Command failed: ${result.error}`);
|
|
|
|
const lockPath = path.join(tmpDir, '.planning', '.lock');
|
|
assert.ok(!fs.existsSync(lockPath), '.lock file should be released after phase complete');
|
|
|
|
const output = JSON.parse(result.output);
|
|
assert.strictEqual(output.completed_phase, '1', 'phase should be completed');
|
|
});
|
|
|
|
test('roadmap update-plan-progress creates and releases .planning/.lock', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'ROADMAP.md'),
|
|
`# Roadmap\n\n| Phase | Plans | Status | Updated |\n|-------|-------|--------|---------|\n| 1 | 0/0 | Not started | - |\n\n### Phase 1: Foundation\n**Goal:** Setup\n`
|
|
);
|
|
|
|
const phaseDir = path.join(tmpDir, '.planning', 'phases', '01-foundation');
|
|
fs.mkdirSync(phaseDir, { recursive: true });
|
|
fs.writeFileSync(path.join(phaseDir, '01-01-PLAN.md'), '# Plan');
|
|
fs.writeFileSync(path.join(phaseDir, '01-01-SUMMARY.md'), '# Summary');
|
|
|
|
const result = runGsdTools('roadmap update-plan-progress 1', tmpDir);
|
|
assert.ok(result.success, `Command failed: ${result.error}`);
|
|
|
|
const lockPath = path.join(tmpDir, '.planning', '.lock');
|
|
assert.ok(!fs.existsSync(lockPath), '.lock file should be released after roadmap update');
|
|
});
|
|
|
|
test('lock file does NOT persist after successful phase operations', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'ROADMAP.md'),
|
|
`# Roadmap v1.0\n`
|
|
);
|
|
|
|
runGsdTools('phase add First Phase', tmpDir);
|
|
runGsdTools('phase add Second Phase', tmpDir);
|
|
|
|
const lockPath = path.join(tmpDir, '.planning', '.lock');
|
|
assert.ok(!fs.existsSync(lockPath), '.lock file should not persist after multiple operations');
|
|
});
|
|
|
|
test('phase add still works correctly with lock (behavioral regression)', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'ROADMAP.md'),
|
|
`# Roadmap v1.0\n\n### Phase 1: Foundation\n**Goal:** Setup\n\n### Phase 2: API\n**Goal:** Build API\n\n---\n`
|
|
);
|
|
|
|
const result = runGsdTools('phase add User Dashboard', tmpDir);
|
|
assert.ok(result.success, `Command failed: ${result.error}`);
|
|
|
|
const output = JSON.parse(result.output);
|
|
assert.strictEqual(output.phase_number, 3, 'should be phase 3');
|
|
assert.strictEqual(output.slug, 'user-dashboard');
|
|
|
|
assert.ok(
|
|
fs.existsSync(path.join(tmpDir, '.planning', 'phases', '03-user-dashboard')),
|
|
'directory should be created'
|
|
);
|
|
|
|
const roadmap = fs.readFileSync(path.join(tmpDir, '.planning', 'ROADMAP.md'), 'utf-8');
|
|
assert.ok(roadmap.includes('### Phase 3: User Dashboard'), 'roadmap should include new phase');
|
|
});
|
|
});
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// 2. readModifyWriteStateMd (tested via CLI commands that use it)
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
describe('readModifyWriteStateMd (via state patch)', () => {
|
|
let tmpDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = createTempProject();
|
|
});
|
|
|
|
afterEach(() => {
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
test('transforms content atomically (read + modify + write under lock)', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
`# Project State\n\n**Current Phase:** 03\n**Status:** Planning\n**Current Plan:** 03-01\n`
|
|
);
|
|
|
|
const result = runGsdTools('state patch --Status "In progress" --"Current Plan" 03-02', tmpDir);
|
|
assert.ok(result.success, `Command failed: ${result.error}`);
|
|
|
|
const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8');
|
|
assert.ok(content.includes('**Status:** In progress'), 'Status should be updated');
|
|
assert.ok(content.includes('03-02'), 'Current Plan should be updated');
|
|
|
|
const lockPath = path.join(tmpDir, '.planning', 'STATE.md.lock');
|
|
assert.ok(!fs.existsSync(lockPath), 'STATE.md.lock should be released after patch');
|
|
});
|
|
|
|
test('lock file cleaned up after state patch operation', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
`# Project State\n\n**Current Phase:** 01\n**Status:** Ready\n`
|
|
);
|
|
|
|
runGsdTools('state patch --Status "In progress"', tmpDir);
|
|
|
|
const lockPath = path.join(tmpDir, '.planning', 'STATE.md.lock');
|
|
assert.ok(!fs.existsSync(lockPath), 'STATE.md.lock should not persist after operation');
|
|
});
|
|
|
|
test('state patch still works correctly via readModifyWriteStateMd path (behavioral regression)', () => {
|
|
const stateMd = [
|
|
'# Project State',
|
|
'',
|
|
'**Current Phase:** 03',
|
|
'**Status:** Planning',
|
|
'**Current Plan:** 03-01',
|
|
'**Last Activity:** 2024-01-15',
|
|
].join('\n') + '\n';
|
|
|
|
fs.writeFileSync(path.join(tmpDir, '.planning', 'STATE.md'), stateMd);
|
|
|
|
const result = runGsdTools('state patch --Status Complete --"Current Phase" 04', tmpDir);
|
|
assert.ok(result.success, `Command failed: ${result.error}`);
|
|
|
|
const updated = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8');
|
|
assert.ok(updated.includes('**Status:** Complete'), 'Status should be updated to Complete');
|
|
assert.ok(updated.includes('**Last Activity:** 2024-01-15'), 'Last Activity should be unchanged');
|
|
});
|
|
|
|
test('two sequential state patches both persist (patch A then patch B)', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
`# Project State\n\n**Current Phase:** 01\n**Status:** Planning\n**Current Plan:** 01-01\n**Last Activity:** 2024-01-01\n`
|
|
);
|
|
|
|
const resultA = runGsdTools('state patch --Status "In progress"', tmpDir);
|
|
assert.ok(resultA.success, `Patch A failed: ${resultA.error}`);
|
|
|
|
const resultB = runGsdTools('state patch --"Current Plan" 01-02', tmpDir);
|
|
assert.ok(resultB.success, `Patch B failed: ${resultB.error}`);
|
|
|
|
const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8');
|
|
assert.ok(content.includes('**Status:** In progress'), 'Patch A (Status) should persist');
|
|
assert.ok(content.includes('01-02'), 'Patch B (Current Plan) should persist');
|
|
assert.ok(content.includes('**Last Activity:** 2024-01-01'), 'Untouched field should be preserved');
|
|
});
|
|
|
|
test('lock file does not persist after rapid sequential patches', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
`# Project State\n\n**Current Phase:** 01\n**Status:** Planning\n**Current Plan:** 01-01\n`
|
|
);
|
|
|
|
runGsdTools('state patch --Status "In progress"', tmpDir);
|
|
runGsdTools('state patch --"Current Plan" 01-02', tmpDir);
|
|
runGsdTools('state patch --Status Complete', tmpDir);
|
|
|
|
const lockPath = path.join(tmpDir, '.planning', 'STATE.md.lock');
|
|
assert.ok(!fs.existsSync(lockPath), 'STATE.md.lock should not persist after rapid sequential patches');
|
|
});
|
|
});
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// 3. Multi-process concurrent write tests
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
describe('multi-process concurrent write tests', () => {
|
|
let tmpDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = createTempProject();
|
|
});
|
|
|
|
afterEach(() => {
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
// 'two concurrent state patches to DIFFERENT fields both persist' deleted per #453 (clock-seam):
|
|
// plain Promise.all without a barrier is non-deterministic — the weakened OR assertion
|
|
// (aOk || bOk) and the vacuous assert.ok(true, ...) make the test pass even when one write
|
|
// is lost. Deterministic barrier-based coverage lives in locking-bugs-1909-1916-1925-1927:
|
|
// 'state update: both concurrent updates to different fields survive'.
|
|
|
|
// 'lock file does not persist after concurrent operations' deleted per #453 (clock-seam):
|
|
// plain Promise.all without a barrier; lock-cleanup coverage is in clock-seam.test.cjs
|
|
// describe('exit cleanup: STATE.md.lock removed on process exit').
|
|
|
|
test('three rapid sequential patches all persist', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
[
|
|
'# Project State',
|
|
'',
|
|
'**Current Phase:** 01',
|
|
'**Status:** Planning',
|
|
'**Current Plan:** 01-01',
|
|
'**Last Activity:** 2025-01-01',
|
|
'',
|
|
].join('\n')
|
|
);
|
|
|
|
const r1 = runGsdTools('state patch --Status "In progress"', tmpDir);
|
|
assert.ok(r1.success, `Patch 1 failed: ${r1.error}`);
|
|
|
|
const r2 = runGsdTools('state patch --"Current Plan" 01-02', tmpDir);
|
|
assert.ok(r2.success, `Patch 2 failed: ${r2.error}`);
|
|
|
|
const r3 = runGsdTools('state patch --"Last Activity" 2025-06-15', tmpDir);
|
|
assert.ok(r3.success, `Patch 3 failed: ${r3.error}`);
|
|
|
|
const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8');
|
|
assert.ok(content.includes('In progress'), 'Patch 1 (Status) should persist');
|
|
assert.ok(content.includes('01-02'), 'Patch 2 (Current Plan) should persist');
|
|
assert.ok(content.includes('2025-06-15'), 'Patch 3 (Last Activity) should persist');
|
|
assert.ok(content.includes('**Current Phase:** 01'), 'Untouched field should be preserved');
|
|
});
|
|
});
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// 4. normalizeMd behavioral equivalence (O(n) insideFence rewrite)
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
describe('normalizeMd behavioral equivalence', () => {
|
|
test('simple markdown with headings and paragraphs', () => {
|
|
const input = '# Title\nSome text.\n## Section\nMore text.\n';
|
|
const result = normalizeMd(input);
|
|
assert.ok(result.includes('# Title\n\nSome text.'), 'title heading should have blank line after');
|
|
assert.ok(result.includes('\n\n## Section\n\nMore text.'), 'section heading should have blank lines around it');
|
|
assert.ok(result.endsWith('\n'), 'should end with newline');
|
|
assert.ok(!result.endsWith('\n\n'), 'should not end with double newline');
|
|
});
|
|
|
|
test('single fenced code block gets blank lines before/after', () => {
|
|
const input = 'Some text\n```js\nconst x = 1;\n```\nMore text\n';
|
|
const result = normalizeMd(input);
|
|
assert.ok(result.includes('Some text\n\n```js'), 'code block should have blank line before');
|
|
assert.ok(result.includes('```\n\nMore text'), 'code block should have blank line after');
|
|
assert.ok(result.includes('const x = 1;'), 'code content should be preserved');
|
|
});
|
|
|
|
test('multiple fenced code blocks', () => {
|
|
const input = 'Intro\n```js\nfoo();\n```\nMiddle\n```py\nbar()\n```\nEnd\n';
|
|
const result = normalizeMd(input);
|
|
assert.ok(result.includes('Intro\n\n```js'), 'first code block should have blank line before');
|
|
assert.ok(result.includes('```\n\nMiddle'), 'first code block should have blank line after');
|
|
assert.ok(result.includes('Middle\n\n```py'), 'second code block should have blank line before');
|
|
assert.ok(result.includes('```\n\nEnd'), 'second code block should have blank line after');
|
|
});
|
|
|
|
test('unclosed fence at end of file (edge case)', () => {
|
|
const input = 'Some text\n```js\nconst x = 1;\n';
|
|
const result = normalizeMd(input);
|
|
assert.ok(typeof result === 'string', 'should return a string');
|
|
assert.ok(result.includes('```js'), 'fence opener should be preserved');
|
|
assert.ok(result.includes('const x = 1;'), 'content after unclosed fence should be preserved');
|
|
assert.ok(result.endsWith('\n'), 'should end with newline');
|
|
});
|
|
|
|
test('empty string input', () => {
|
|
assert.strictEqual(normalizeMd(''), '', 'empty string should return empty string');
|
|
});
|
|
|
|
test('mixed headings + lists + fences (complex case)', () => {
|
|
const input = [
|
|
'# Title',
|
|
'## Section One',
|
|
'Paragraph text.',
|
|
'- item 1',
|
|
'- item 2',
|
|
'## Section Two',
|
|
'```bash',
|
|
'echo hello',
|
|
'```',
|
|
'After code.',
|
|
'## Section Three',
|
|
'1. First',
|
|
'2. Second',
|
|
'Done.',
|
|
].join('\n') + '\n';
|
|
|
|
const result = normalizeMd(input);
|
|
|
|
assert.ok(result.includes('\n\n## Section One\n\n'), 'Section One heading needs blank lines');
|
|
assert.ok(result.includes('\n\n## Section Two\n\n'), 'Section Two heading needs blank lines');
|
|
assert.ok(result.includes('\n\n## Section Three\n\n'), 'Section Three heading needs blank lines');
|
|
assert.ok(result.includes('Paragraph text.\n\n- item 1'), 'list should have blank line before');
|
|
assert.ok(result.includes('\n\n```bash'), 'code block should have blank line before');
|
|
assert.ok(result.includes('```\n\nAfter code.'), 'code block should have blank line after');
|
|
assert.ok(result.includes('echo hello'), 'code content should be preserved');
|
|
assert.ok(!result.includes('\n\n\n'), 'should not have 3+ consecutive blank lines');
|
|
});
|
|
});
|
|
|
|
// normalizeMd performance benchmark tests deleted per #453 (clock-seam):
|
|
// wall-clock assertions (elapsed < 50ms, elapsed < 200ms) are inherently
|
|
// flaky on loaded CI runners. Correctness is covered by the snapshot tests
|
|
// in describe('normalizeMd snapshot tests') below.
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// 6. normalizeMd snapshot tests
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
describe('normalizeMd snapshot tests', () => {
|
|
test('snapshot - heading spacing', () => {
|
|
const input = '# Title\nParagraph\n## Section\nMore text';
|
|
const expected = '# Title\n\nParagraph\n\n## Section\n\nMore text\n';
|
|
const result = normalizeMd(input);
|
|
assert.strictEqual(result, expected,
|
|
`Heading spacing snapshot mismatch.\nGot: ${JSON.stringify(result)}\nExpected: ${JSON.stringify(expected)}`
|
|
);
|
|
});
|
|
|
|
test('snapshot - code block spacing', () => {
|
|
const input = 'Text before\n```js\nconst x = 1;\n```\nText after\n';
|
|
const expected = 'Text before\n\n```js\nconst x = 1;\n```\n\nText after\n';
|
|
const result = normalizeMd(input);
|
|
assert.strictEqual(result, expected,
|
|
`Code block spacing snapshot mismatch.\nGot: ${JSON.stringify(result)}\nExpected: ${JSON.stringify(expected)}`
|
|
);
|
|
});
|
|
|
|
test('snapshot - list spacing', () => {
|
|
const input = 'Paragraph\n- item 1\n- item 2\nAnother paragraph';
|
|
const expected = 'Paragraph\n\n- item 1\n- item 2\n\nAnother paragraph\n';
|
|
const result = normalizeMd(input);
|
|
assert.strictEqual(result, expected,
|
|
`List spacing snapshot mismatch.\nGot: ${JSON.stringify(result)}\nExpected: ${JSON.stringify(expected)}`
|
|
);
|
|
});
|
|
|
|
test('snapshot - complex mixed document', () => {
|
|
const input = [
|
|
'# Main Title',
|
|
'Intro paragraph.',
|
|
'## Section One',
|
|
'Some text here.',
|
|
'```js',
|
|
'const a = 1;',
|
|
'```',
|
|
'- first item',
|
|
'- second item',
|
|
'## Section Two',
|
|
'Final text.',
|
|
].join('\n');
|
|
|
|
const expected = [
|
|
'# Main Title',
|
|
'',
|
|
'Intro paragraph.',
|
|
'',
|
|
'## Section One',
|
|
'',
|
|
'Some text here.',
|
|
'',
|
|
'```js',
|
|
'const a = 1;',
|
|
'```',
|
|
'',
|
|
'- first item',
|
|
'- second item',
|
|
'',
|
|
'## Section Two',
|
|
'',
|
|
'Final text.',
|
|
'',
|
|
].join('\n');
|
|
|
|
const result = normalizeMd(input);
|
|
assert.strictEqual(result, expected,
|
|
`Complex mixed document snapshot mismatch.\nGot: ${JSON.stringify(result)}\nExpected: ${JSON.stringify(expected)}`
|
|
);
|
|
});
|
|
});
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// 7. Warnings (frontmatter parse, state field miss)
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
describe('warnings', () => {
|
|
let tmpDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = createTempProject();
|
|
});
|
|
|
|
afterEach(() => {
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
test('must_haves parse warning fires for block with content but 0 items', () => {
|
|
const planDir = path.join(tmpDir, '.planning', 'phases', '01-test');
|
|
fs.mkdirSync(planDir, { recursive: true });
|
|
fs.writeFileSync(
|
|
path.join(planDir, '01-01-PLAN.md'),
|
|
`---
|
|
phase: "01"
|
|
plan: "01"
|
|
must_haves:
|
|
acceptance:
|
|
bare content without dash prefix
|
|
another line without dash prefix
|
|
---
|
|
|
|
# Plan 01-01
|
|
`
|
|
);
|
|
|
|
// #3884 (ADR-3473 §8.4): `frontmatter get <file> <field>` is not a real
|
|
// form — the documented shape is `frontmatter get <file> [--field key]`
|
|
// (docs/CLI-TOOLS.md:736). Under the pre-#3884 permissive parser the bare
|
|
// "must_haves" positional was silently dropped, `field` resolved to
|
|
// null, and cmdFrontmatterGet dumped the WHOLE frontmatter object — the
|
|
// test's original `result.output.includes('acceptance')` branch passed
|
|
// only because the full dump happens to contain that substring, not
|
|
// because field selection ever worked. `cmdFrontmatterGet` never calls
|
|
// `parseMustHavesBlock` (that WARNING is only emitted by other
|
|
// consumers), so the WARNING branch of the old assertion could never
|
|
// fire through this command either. Corrected to the real `--field`
|
|
// form and strengthened to assert on the actual, field-scoped payload.
|
|
const result = runGsdTools(
|
|
['frontmatter', 'get', path.join(planDir, '01-01-PLAN.md'), '--field', 'must_haves'],
|
|
tmpDir
|
|
);
|
|
|
|
assert.ok(result.success, `frontmatter get --field must_haves failed: ${result.error}`);
|
|
const parsed = JSON.parse(result.output);
|
|
assert.deepStrictEqual(
|
|
Object.keys(parsed),
|
|
['must_haves'],
|
|
`--field must_haves must scope the output to only that field, got keys: ${Object.keys(parsed).join(', ')}`,
|
|
);
|
|
// Dash-less list items under a YAML mapping key fold into a single plain
|
|
// scalar string, not an array — this is the actual "0 items" parse
|
|
// hazard the test's title names, surfaced directly rather than via a
|
|
// WARNING this command path never emits.
|
|
assert.strictEqual(
|
|
typeof parsed.must_haves.acceptance,
|
|
'string',
|
|
`bare-content (no dash prefix) must_haves.acceptance must parse as a scalar string, not a list, got: ${JSON.stringify(parsed.must_haves.acceptance)}`,
|
|
);
|
|
});
|
|
|
|
test('stateReplaceFieldWithFallback logs warning on miss', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
`# Project State\n\n**Current Phase:** 01\n**Current Plan:** 1\n**Total Plans in Phase:** 3\n`
|
|
);
|
|
|
|
const result = runGsdTools('state advance-plan', tmpDir);
|
|
assert.ok(result.success, `Command failed: ${result.error}`);
|
|
|
|
const output = JSON.parse(result.output);
|
|
assert.ok(output.advanced === true || output.reason === 'last_plan', 'advance should complete');
|
|
});
|
|
});
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// 8. Malformed input resilience
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
describe('malformed input resilience', () => {
|
|
let tmpDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = createTempProject();
|
|
});
|
|
|
|
afterEach(() => {
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
test('STATE.md with invalid bold format -- state patch returns gracefully', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
'# Project State\n\n**Current Phase: 01\n**Status:** Planning\n'
|
|
);
|
|
|
|
const result = runGsdTools('state patch --Status "In progress"', tmpDir);
|
|
const didNotCrash = result.success || (result.output !== undefined);
|
|
assert.ok(didNotCrash, `state patch should not crash on malformed bold format: ${result.error}`);
|
|
|
|
if (result.success) {
|
|
const content = fs.readFileSync(path.join(tmpDir, '.planning', 'STATE.md'), 'utf-8');
|
|
assert.ok(
|
|
content.includes('In progress'),
|
|
'Status field (with valid bold format) should be updated'
|
|
);
|
|
}
|
|
});
|
|
|
|
test('STATE.md with only frontmatter, no body -- state patch handles gracefully', () => {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'STATE.md'),
|
|
'---\nphase: "01"\n---\n'
|
|
);
|
|
|
|
const result = runGsdTools('state patch --Status "In progress"', tmpDir);
|
|
const didNotCrash = result.success || (result.output !== undefined);
|
|
assert.ok(didNotCrash, `state patch should not crash on frontmatter-only STATE.md: ${result.error}`);
|
|
});
|
|
});
|
|
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
// 9. Stress tests with 50+ phases
|
|
// ─────────────────────────────────────────────────────────────────────────────
|
|
|
|
describe('stress tests with 50+ phases', () => {
|
|
let tmpDir;
|
|
|
|
beforeEach(() => {
|
|
tmpDir = createTempProject();
|
|
});
|
|
|
|
afterEach(() => {
|
|
cleanup(tmpDir);
|
|
});
|
|
|
|
// roadmap analyze on 50-phase ROADMAP: behavioral test without timing gate.
|
|
// The elapsed < N wall-clock gate was a persistent flake source (#7, #453).
|
|
// Equivalent behavioral coverage (50 phases, 25 complete) lives in
|
|
// tests/clock-seam.test.cjs describe('roadmap analyze behavioral correctness').
|
|
|
|
|
|
test('phase complete on phase 26 of 50-phase project works correctly', () => {
|
|
create50PhaseProject(tmpDir, 25);
|
|
writeMinimalStateMd(tmpDir, '# Session State\n\n**Current Phase:** 26\n**Status:** In progress\n');
|
|
|
|
const phase26Dir = path.join(tmpDir, '.planning', 'phases', '26-feature-26');
|
|
fs.writeFileSync(
|
|
path.join(phase26Dir, '26-01-SUMMARY.md'),
|
|
'# Phase 26 Plan 1 Summary\n\nFeature 26 completed.\n'
|
|
);
|
|
fs.writeFileSync(
|
|
path.join(phase26Dir, '26-VERIFICATION.md'),
|
|
'---\nstatus: passed\nscore: "1/1"\n---\n# Verification\nPassed.\n',
|
|
);
|
|
|
|
const result = runGsdTools('phase complete 26', tmpDir);
|
|
assert.ok(result.success, `phase complete 26 should succeed: ${result.error}`);
|
|
|
|
const roadmapContent = fs.readFileSync(
|
|
path.join(tmpDir, '.planning', 'ROADMAP.md'),
|
|
'utf-8'
|
|
);
|
|
const phase26Checkbox = roadmapContent.match(/-\s*\[(x| )\]\s*.*Phase\s+26/i);
|
|
assert.ok(phase26Checkbox, 'Should find Phase 26 checkbox in ROADMAP');
|
|
assert.strictEqual(phase26Checkbox[1], 'x', 'Phase 26 should now be marked as complete [x]');
|
|
});
|
|
});
|