Files
msd-core/tests/adr857-core-without-capabilities.test.cjs
Tom Boucher e20744eacb enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick

ADR-3473 §8.4 says failure is a value. Three families currently encode failure as
success, and this commit pins each one RED before the fix lands.

Measured on this tree, 2026-08-26:

  gsd-tools generate-slug "test" --pick nonexistent
    -> empty stdout, exit 0                                     (#3365)

  gsd-tools audit-open --pick nonexistent_field
    -> dumps the entire human-readable audit report, exit 0

  gsd-tools generate-slug "Hello World" --raw --pick bogus
    -> prints "hello-world", another field's value, exit 0

  gsd-tools query state.planned-phase 3        (positional, no --phase)
    -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to
       "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted
       current_phase_name                                        (#3358)

tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract
("returns empty string for missing field", success === true). That assertion is
replaced by the required behavior rather than deleted.

The new parseNamedArgs block calls the spec-object signature that does not exist
yet, so it fails today by construction. The 11 existing behavior-lock tests are
left untouched here; they are corrected in the implementation commit.

C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180
Decision 4(b). A unit assertion on the parser would have passed throughout this
defect's life.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3884): failure is a value — strict argv, and --pick that signals absence

Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable
ways to say "I could not answer".

parseNamedArgs (src/command-arg-projection.cts)
  Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the
  hub's Result shape instead of a bare Record. Declaring the positional arity is what
  makes #3358's call site unrepresentable rather than merely detectable: an unrecognized
  flag or a token past the declared boundary is now InvalidArgs, naming the offending
  token and listing the accepted flags. The legacy positional-array call shape throws
  a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale
  hand-written .cjs call site fails loudly instead of destructuring undefined off a
  Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a
  projection over the one parser, not a second parser.

  Measured before, against a STATE.md with a populated phase-2 block:
    query state.planned-phase 3        (positional, no --phase)
    -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to
       "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted
       current_phase_name
  After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical.
  The flag form is unchanged and still updates STATE.md.

--pick <field> (gsd-core/bin/gsd-tools.cjs)
  extractField returns {found,value}, and the pick block no longer shares one catch
  between "output was not JSON" and "field was absent". An absent field exits 1 with
  pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1
  with pick_output_not_json instead of dumping the command's entire output. A field that
  is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer,
  not a failure, and it is what keeps `--pick count` printing 0 on a fresh project.

  Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable
  audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" —
  a different field's value, confidently, at exit 0.

  ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The
  sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the
  ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0
  would demote "could not answer" to "the answer is zero" — the hazard
  docs/how-to/resolve-unreachable-guard-findings.md already warns against.

Guard ledger (ADR-3473 Decision 6)
  scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a
  `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the
  correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls,
  a nullglob mechanism this change does not touch) is retained in full, as are the shared
  scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file
  is not deleted.

Call-site audit
  45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an
  if test, && chain, or a pipeline whose status is consumed, and no shell block in
  workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the
  prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind
  a prior found/existence check. No ADR-3409-class "field the command never produces"
  remains.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows

Two review findings, both fixed here rather than recorded as limits.

1. A newline in an untrusted token forged a second stderr line.

   Before, plain-text mode:
     $ gsd-tools query state.planned-phase $'foo\nError: forged second line'
     Error: unexpected positional argument "foo
     Error: forged second line"

   After:
     Error: unexpected positional argument "foo\nError: forged second line"

   --json-errors mode was never affected — io.error runs that payload through
   JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the
   three new InvalidArgs reasons plus the two new --pick diagnostics all
   interpolate a token that comes straight from argv.

   Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at
   every interpolation site — not a copy per site. It is deliberately NOT
   applied inside error() itself: several callers in this tree emit intentional
   multi-line diagnostics, and escaping newlines there would mangle them.

   The available-top-level-keys list needed the same treatment for a reason the
   review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user
   document and echoes that document's own keys into the diagnostic. Verified
   reachable — a frontmatter key containing a newline reaches the key list — so
   formatKeyForDiagnosticList is guarding a live path, not a hypothetical one.
   Ordinary keys still render plain and unquoted; a fix that merely dropped the
   key would also have passed a "one line" assertion, so the test pins the
   escaped key's presence too.

2. Five behavior-table rows were implemented but nothing pinned them:
   B7  a dotted path that dies partway
   B9  bracket syntax on a non-array
   B10 a negative array index, in and out of range
   B14 a JSON root that is not an object
   B17 an @file: payload over 50KB

   B17 is the load-bearing one. output() writes @file:<path> instead of inline
   JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no
   test, a future reordering of those two steps turns every large result into a
   false pick_output_not_json. The fixture seeds 1200 phase directories and
   measures the payload at 62474 characters, asserting the spill actually
   happened rather than assuming it.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): correct the strict-argv surface against a full verification run

The first full run came back with 90 failures across 12 files, none in the new
tests. They were the argv surface telling me what it actually is. Ten root
causes; each classified before anything was changed.

I over-implemented, and that is reverted.

  ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional
  tokens". It says nothing about a value flag whose value is missing. Making
  that an error was my design decision, not the rule, and it broke a
  deliberately recorded contract: `--prd` with no value resolving to null
  (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5;
  tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The
  "requires a value" branch is deleted outright rather than kept behind an
  option — an unused strictness mode is speculative generality. Unknown-flag
  and unexpected-positional rejection, which is what §8.4 actually mandates,
  is unchanged.

--wave needed a third flag kind the original design did not anticipate.

  `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the
  shipped workflow reconstructs and passes it (execute-phase.md:84), while
  #2932 records token-PRESENCE semantics: the CLI cares only that the flag
  appeared, and the value belongs to the workflow layer. That is neither a
  boolean flag nor a value flag, so `optionalValueFlags` now exists —
  presence-only in `data`, and the validation cursor consumes a following
  non-flag token so it is not reported as a stray positional. Every other
  declared boolean flag was checked against every argument-hint and prose
  usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one
  of this shape.

Five tests were pinning forms that never worked.

  tests/adr857-core-without-capabilities.test.cjs passed
  `init plan-phase --phase 01-stub`, but the documented form is positional
  (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form
  is the literal string "--phase". Measured on the pre-fix build against a
  real .planning/phases/01-stub/ directory:

    init plan-phase 01-stub          -> phase_found=true
    init plan-phase --phase 01-stub  -> phase_found=false

  The test asserted only exit 0 and key presence, so it had been green while
  proving nothing about phase resolution. Corrected to the documented form and
  strengthened to assert phase_found === true. Same class in state.test.cjs
  (`--plan-count`, a flag that does not exist; the real one is `--plans`),
  milestone-archive.test.cjs (`init new-milestone --json`, silently ignored),
  and concurrency-safety.test.cjs (a bare positional field name whose
  OR-assertion passed because a whole-document dump happens to contain the
  substring it looked for).

Six handlers had no argv validation at all — the same #3358 shape this phase
exists to close, found while fixing the rest: init verify-work / phase-op /
review / todos / remove-workspace read args[2] with nothing checking the rest,
and validate health read --repair/--backfill through a bare args.includes()
scan that bypassed the parser entirely. All now go through the seam, so the
flag has one owner.

tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must
NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's
local expectation does not override §8, so they are inverted and renamed —
a test still called "ignores an unrecognized flag" while asserting rejection
would be its own defect. Row C6's point is its PWNED canary; that assertion is
kept verbatim and only its exit-status expectation changed, because the
hostile token is now rejected rather than absorbed.

The blast-radius estimate in 40-design.md is corrected rather than quietly
left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was
accurate for what the graph can see — parseNamedArgs's callers. It cannot see
that those callers' handlers accept argv shapes wider than the code reading
args[2] suggests, which is where the real surface was.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert

Second full run: 46 failures, down from 90. Four causes, two of them mine.

Reverted `validate health` entirely — it was scope creep, and it broke a real flag.

  ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`.
  The previous commit routed `validate health` through the parser on the
  reasoning that a flag should have one owner. That was wrong twice over:
  §8.4 names parseNamedArgs and count queries, and `validate health` was never
  a parseNamedArgs call site — it read its flags, just not through the parser,
  so it had no silent-drop defect to fix. Tightening it omitted `--json`, which
  the health-diagnostic suites use heavily. The handler is now byte-for-behaviour
  back to its pre-branch form. `validate context` stays converted: it genuinely
  was a call site, and its `--json` is now declared rather than read by a second
  `args.includes` scan.

  The five handlers that had NO validation at all — init verify-work / phase-op /
  review / todos / remove-workspace — stay fixed. Those read args[2] with nothing
  checking the rest, which is the #3358 shape this phase owns.

Finished the A2/A3 revert. Three tests still encoded the deleted
"a value flag with a missing value is an error" rule, including one added by the
previous commit for that rule. All three now assert the reverted null contract,
and the ones whose titles said "rejected" are renamed — a test named for a
contract it no longer asserts is its own defect.

`--wave=` and `--wave --weird` are correctly rejected. Neither is documented in
commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and
neither is emitted by the shipped prompt layer, so both are unrecognized tokens
that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the
property it exists for — asserted directly now, at the parser, that `--wave` does
not swallow a following flag as its value — and only its exit-status expectation
changed.

A contradiction inside this branch, surfaced by the audit and resolved the safe way.

  Two pre-existing #3573 tests call `state begin-phase '2'` and
  `state planned-phase '2'` with a bare positional, relying on the old permissive
  parser to ignore it. This branch's own #3358 regression test requires that exact
  argv to be REJECTED. The two are mutually exclusive.

  Widening the router to accept a bare positional — mirroring complete-phase —
  would have silently re-opened #3358, and was verified to do exactly that: with
  the widened router, `query state.planned-phase 3` returned exit 0 and wrote
  current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and
  docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the
  two #3573 tests move to it. Their assertions were never about the call shape —
  only that total_phases survives the resync — and both still pass.

  complete-phase is untouched: its bare positional IS documented, and it keeps the
  dynamic boundary and the negative-space note that record why.

The audit that produced this is in the PR body: for every handler whose declaration
changed, the flags it reads anywhere in its body, the flags the shipped surface
documents, and the shapes the suite passes, compared. The `--json` miss was a
pattern, not an accident — declaring a handler's flags from its parseNamedArgs call
alone misses whatever it reads elsewhere.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3884): backfill the changeset PR number

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 00:12:13 -04:00

580 lines
23 KiB
JavaScript

'use strict';
/**
* adr857-core-without-capabilities.test.cjs
*
* ADR-857 deliverable B — "the core loop ships and runs without any plug-in"
* (Consequences, §"Positive": "The core loop ships and runs without any plug-in;
* plan-phase.md/execute-phase.md shrink to the irreducible five steps.")
*
* Verified contracts:
* B1. All 12 canonical loop points return activeHooks:[] when every
* capability when-key is explicitly false (real registry, all-caps-off config).
* B2. The CLI `loop render-hooks <point>` exits 0 and emits activeHooks:[],
* placeholder rendered for representative points with all-caps-off config.
* B3. Init bundles for the 5-step loop's entry seam (plan-phase, execute-phase,
* verify-work) resolve with exit 0 and valid JSON when capabilities are off.
* B4. An EMPTY registry (byLoopPoint:{}) at all 12 points → activeHooks:[]
* (loop tolerates a capability-less install).
* B5. [BVA] Exactly one capability ON (tdd_mode) → that capability's points
* non-empty, all OTHER points still empty (caps are additive; core is baseline).
*
* RULESET: no readFileSync + .includes() on source files (source-grep ban).
* All assertions drive real exported functions / subprocess and inspect typed results.
*/
const { describe, test, before, after } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const os = require('node:os');
const path = require('node:path');
const { execFileSync } = require('child_process');
// ── Module under test ─────────────────────────────────────────────────────────
const {
resolveLoopHooks,
renderLoopHooks,
CANONICAL_POINTS,
} = require('../gsd-core/bin/lib/loop-resolver.cjs');
// Real registry (compiled from capabilities/ at build time)
const realRegistry = require('../gsd-core/bin/lib/capability-registry.cjs');
// ── Helpers from test harness ─────────────────────────────────────────────────
const { cleanup } = require('./helpers.cjs');
// ── Paths ─────────────────────────────────────────────────────────────────────
const GSD_TOOLS = path.join(__dirname, '..', 'gsd-core', 'bin', 'gsd-tools.cjs');
// ── Config fixtures ───────────────────────────────────────────────────────────
/**
* All 12 canonical loop points. Derived from the exported constant so the
* assertion set cannot drift from the resolver's own authoritative list.
*/
const ALL_12_POINTS = [...CANONICAL_POINTS];
/**
* All when-keys discovered from the real registry, set to false.
* Built by scanning every hook in every loop point's steps/contributions/gates arrays.
* This gives us a "caps-off" config that passes through the activation resolver as
* explicitly false rather than relying on missing-key default behaviour.
*
* Structure: nested (workflow.* → workflow:{...}, intel.enabled → intel:{enabled:false})
* because _getNestedConfigValue expects a nested object, not a flat dotted key.
*/
function buildAllFalseConfig() {
const workflow = {};
const intel = {};
for (const point of ALL_12_POINTS) {
const entry = realRegistry.byLoopPoint[point];
if (!entry) continue;
for (const kind of ['steps', 'contributions', 'gates']) {
for (const hook of entry[kind] || []) {
const when = hook.when;
if (typeof when !== 'string' || !when) continue;
if (when.startsWith('workflow.')) {
const key = when.slice('workflow.'.length);
workflow[key] = false;
} else if (when === 'intel.enabled') {
intel.enabled = false;
}
// Any future top-level keys would need extending here.
}
}
}
return { workflow, intel };
}
const ALL_FALSE_CONFIG = buildAllFalseConfig();
/**
* All-false config with tdd_mode: true.
* Only workflow.tdd_mode differs from ALL_FALSE_CONFIG.
*/
function buildTddOnlyConfig() {
return {
...ALL_FALSE_CONFIG,
workflow: { ...ALL_FALSE_CONFIG.workflow, tdd_mode: true },
};
}
// ── Helpers ───────────────────────────────────────────────────────────────────
/**
* Run gsd-tools subprocess and return { exitCode, output }.
* Does NOT throw on non-zero exit — let the test assert.
*/
function runCli(args, cwd) {
try {
const stdout = execFileSync(process.execPath, [GSD_TOOLS, ...args], {
cwd,
encoding: 'utf-8',
timeout: 30000,
});
return { exitCode: 0, output: stdout.trim() };
} catch (err) {
return {
exitCode: err.status ?? 1,
output: err.stdout?.toString().trim() ?? '',
error: err.stderr?.toString().trim() ?? '',
};
}
}
/** Create a temp project dir with a .planning/ sub-dir. */
function makeProject(configJson = null) {
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'adr857-b-'));
const planning = path.join(dir, '.planning');
fs.mkdirSync(planning, { recursive: true });
if (configJson !== null) {
fs.writeFileSync(path.join(planning, 'config.json'), JSON.stringify(configJson), 'utf8');
}
return dir;
}
/** Remove a temp dir safely. */
function removeTmp(dir) {
if (dir) cleanup(dir);
}
// ─────────────────────────────────────────────────────────────────────────────
// B1. [happy/aggregate] All 12 points → activeHooks:[] with all-caps-off config
// ─────────────────────────────────────────────────────────────────────────────
describe('B1 — real registry, all-caps-off config: every canonical point resolves to activeHooks:[]', () => {
test('all 12 CANONICAL_POINTS return activeHooks:[] simultaneously when every capability when-key is false', () => {
const failures = [];
for (const point of ALL_12_POINTS) {
const result = resolveLoopHooks({
point,
registry: realRegistry,
config: ALL_FALSE_CONFIG,
});
// Shape guard — result must be an object with an array
assert.ok(result && typeof result === 'object', `${point}: result must be an object`);
assert.ok(Array.isArray(result.activeHooks), `${point}: activeHooks must be an array`);
if (result.activeHooks.length !== 0) {
failures.push({
point,
count: result.activeHooks.length,
capIds: result.activeHooks.map(h => h.capId),
});
}
}
// Genuine assertion: if any point has activeHooks, report them concretely.
// This fails on regression to the specific wrong value, not just "not empty".
assert.deepStrictEqual(
failures,
[],
`Expected zero active hooks at all 12 points with all-caps-off config but got: ${JSON.stringify(failures)}`,
);
});
test('CANONICAL_POINTS exports exactly 12 points', () => {
assert.strictEqual(
ALL_12_POINTS.length,
12,
`CANONICAL_POINTS must have 12 entries (ADR-857 §"Loop Extension Points (the 12)"), got ${ALL_12_POINTS.length}`,
);
});
test('each of the 12 known point names is present in CANONICAL_POINTS', () => {
const expected = [
'discuss:pre', 'discuss:post',
'plan:pre', 'plan:post',
'execute:pre', 'execute:wave:pre', 'execute:wave:post', 'execute:post',
'verify:pre', 'verify:post',
'ship:pre', 'ship:post',
];
for (const p of expected) {
assert.ok(
ALL_12_POINTS.includes(p),
`Expected canonical point "${p}" to be in CANONICAL_POINTS`,
);
}
});
});
// ─────────────────────────────────────────────────────────────────────────────
// B2. [happy] CLI render-hooks E2E: exit 0, activeHooks:[], placeholder rendered
// ─────────────────────────────────────────────────────────────────────────────
describe('B2 — CLI loop render-hooks: exit 0, activeHooks:[], placeholder for all-caps-off project', () => {
let tmpDir;
before(() => {
tmpDir = makeProject(ALL_FALSE_CONFIG);
});
after(() => {
removeTmp(tmpDir);
tmpDir = null;
});
for (const point of ['plan:pre', 'execute:wave:post', 'ship:post']) {
test(`loop render-hooks ${point} → exit 0, activeHooks:[], non-empty rendered placeholder`, () => {
const { exitCode, output, error } = runCli(['loop', 'render-hooks', point], tmpDir);
assert.strictEqual(
exitCode,
0,
`Expected exit 0 for "loop render-hooks ${point}" with all-caps-off config; got ${exitCode}. stderr: ${error ?? ''}`,
);
// Must parse as JSON
let parsed;
try {
parsed = JSON.parse(output);
} catch (e) {
assert.fail(`CLI output for ${point} is not valid JSON: ${output.slice(0, 200)}`);
}
// activeHooks must be present and empty
assert.ok(
Array.isArray(parsed.activeHooks),
`${point}: activeHooks must be an array`,
);
assert.strictEqual(
parsed.activeHooks.length,
0,
`${point}: expected activeHooks:[] with all-caps-off config, got ${JSON.stringify(parsed.activeHooks)}`,
);
// rendered field must be a non-empty placeholder string (loop still renders output)
assert.ok(
typeof parsed.rendered === 'string' && parsed.rendered.length > 0,
`${point}: rendered must be a non-empty string, got ${JSON.stringify(parsed.rendered)}`,
);
// Genuine assertion: the placeholder contains the point name so it doesn't silently
// return a generic empty string detached from the requested point.
assert.ok(
parsed.rendered.includes(point),
`${point}: rendered placeholder must reference the point name "${point}", got: "${parsed.rendered}"`,
);
});
}
});
// ─────────────────────────────────────────────────────────────────────────────
// B3. [happy] Init bundles resolve with exit 0 and valid JSON with caps off
// ─────────────────────────────────────────────────────────────────────────────
describe('B3 — init bundles for 5-step loop entry seam: exit 0 and valid JSON with capabilities off', () => {
// Cases: the 5-step loop's main init entry points (those available without git)
const INIT_CASES = [
// #3884 (ADR-3473 §8.4): these three init subcommands take the phase as a
// bare POSITIONAL (docs/CLI-TOOLS.md:775 `init plan-phase <phase>`, :787
// `init execute-phase <phase>`) — there is no `--phase` flag for any of
// them. Under the pre-#3884 permissive parser, `--phase 01-stub` silently
// resolved to phase_found:false (args[2] literally became the string
// "--phase"), which these tests never caught because they only checked
// exit 0 and key PRESENCE. Corrected to the real, documented positional
// form and strengthened to assert phase_found === true so a future
// regression of this kind fails loudly instead of passing vacuously.
{
label: 'init plan-phase',
args: ['init', 'plan-phase', '01-stub'],
// Required fields that prove the bundle is a real JSON object used by the loop
requiredFields: ['tdd_mode', 'phase_found', 'planning_exists'],
},
{
label: 'init execute-phase',
args: ['init', 'execute-phase', '01-stub'],
requiredFields: ['tdd_mode', 'phase_found', 'config_exists'],
},
{
label: 'init verify-work',
args: ['init', 'verify-work', '01-stub'],
requiredFields: ['phase_found', 'commit_docs'],
},
];
for (const { label, args, requiredFields } of INIT_CASES) {
describe(label, () => {
let tmpDir;
before(() => {
// Bare project with .planning/ but all caps off in config
tmpDir = makeProject(ALL_FALSE_CONFIG);
// A real phase directory matching the "01-stub" argument used by every
// INIT_CASES entry above — without it, phase_found is trivially false
// regardless of whether the CLI call shape is correct, and the
// strengthened phase_found:true assertion below could not distinguish
// a working positional form from the #3358-class silent-drop bug.
fs.mkdirSync(path.join(tmpDir, '.planning', 'phases', '01-stub'), { recursive: true });
});
after(() => {
removeTmp(tmpDir);
tmpDir = null;
});
test(`${label} exits 0 with capabilities off`, () => {
const { exitCode, error } = runCli(args, tmpDir);
assert.strictEqual(
exitCode,
0,
`${label}: expected exit 0 with all-caps-off project, got ${exitCode}. stderr: ${error ?? ''}`,
);
});
test(`${label} returns parseable JSON with expected fields`, () => {
const { output } = runCli(args, tmpDir);
let parsed;
try {
parsed = JSON.parse(output);
} catch (e) {
assert.fail(`${label}: output is not valid JSON: ${output.slice(0, 200)}`);
}
assert.ok(
parsed && typeof parsed === 'object' && !Array.isArray(parsed),
`${label}: JSON must be a plain object`,
);
for (const field of requiredFields) {
assert.ok(
Object.prototype.hasOwnProperty.call(parsed, field),
`${label}: bundle must contain field "${field}", got keys: ${Object.keys(parsed).join(', ')}`,
);
}
// Genuine assertion, not merely key presence (see the fix note on
// INIT_CASES above): the seeded project has a real "01-stub" phase
// directory, so the phase MUST actually resolve. A silently-dropped
// phase argument still has a `phase_found` key (value `false`) — key
// presence alone would pass on that defect.
assert.strictEqual(
parsed.phase_found,
true,
`${label}: phase "01-stub" must actually resolve (phase_found:true), got: ${JSON.stringify(parsed.phase_found)}`,
);
});
test(`${label} returns parseable JSON with bare project (no config at all)`, () => {
const bareDir = makeProject(null); // no config.json
try {
const { exitCode, output, error } = runCli(args, bareDir);
assert.strictEqual(
exitCode,
0,
`${label}: expected exit 0 with bare project (no config), got ${exitCode}. stderr: ${error ?? ''}`,
);
let parsed;
try {
parsed = JSON.parse(output);
} catch (e) {
assert.fail(`${label}: bare project output is not valid JSON: ${output.slice(0, 200)}`);
}
assert.ok(
parsed && typeof parsed === 'object' && !Array.isArray(parsed),
`${label}: bare project JSON must be a plain object`,
);
} finally {
removeTmp(bareDir);
}
});
});
}
});
// ─────────────────────────────────────────────────────────────────────────────
// B4. [negative] Empty registry → activeHooks:[] at all 12 points
// ─────────────────────────────────────────────────────────────────────────────
describe('B4 — empty registry (byLoopPoint:{}) at all 12 points → activeHooks:[]', () => {
const EMPTY_REGISTRY = {
byLoopPoint: {},
capabilities: {},
configKeys: {},
configSchema: {},
commandFamilies: {},
};
test('loop tolerates a capability-less install: all 12 points return activeHooks:[]', () => {
const failures = [];
for (const point of ALL_12_POINTS) {
const result = resolveLoopHooks({
point,
registry: EMPTY_REGISTRY,
config: {},
});
assert.ok(
result && typeof result === 'object',
`${point}: result must be an object`,
);
assert.ok(
Array.isArray(result.activeHooks),
`${point}: activeHooks must be an array`,
);
if (result.activeHooks.length !== 0) {
failures.push({
point,
count: result.activeHooks.length,
capIds: result.activeHooks.map(h => h.capId),
});
}
}
assert.deepStrictEqual(
failures,
[],
`Empty registry: expected zero active hooks at all 12 points, got non-empty at: ${JSON.stringify(failures)}`,
);
});
test('empty registry does not throw for any of the 12 canonical points', () => {
for (const point of ALL_12_POINTS) {
assert.doesNotThrow(
() => resolveLoopHooks({ point, registry: EMPTY_REGISTRY, config: {} }),
`resolveLoopHooks must not throw for empty registry at point "${point}"`,
);
}
});
test('renderLoopHooks with empty activeHooks returns a non-empty placeholder string', () => {
const placeholder = renderLoopHooks({ point: 'plan:pre', activeHooks: [] });
assert.ok(
typeof placeholder === 'string' && placeholder.length > 0,
`renderLoopHooks must return a non-empty string for empty activeHooks, got: ${JSON.stringify(placeholder)}`,
);
// Specific value check — genuineness: this must change if the placeholder format changes
assert.strictEqual(
placeholder,
'_No active hooks at plan:pre._',
`renderLoopHooks placeholder must be "_No active hooks at plan:pre._", got: "${placeholder}"`,
);
});
});
// ─────────────────────────────────────────────────────────────────────────────
// B5. [BVA] One capability ON (tdd_mode) → its 2 points non-empty, 10 others empty
// ─────────────────────────────────────────────────────────────────────────────
describe('B5 — BVA: tdd_mode ON, all other caps OFF → additive: only tdd points active', () => {
// tdd contributes at plan:pre and execute:post (verified from capability registry)
const TDD_ACTIVE_POINTS = ['plan:pre', 'execute:post'];
const TDD_INACTIVE_POINTS = ALL_12_POINTS.filter(p => !TDD_ACTIVE_POINTS.includes(p));
const TDD_ON_CONFIG = buildTddOnlyConfig();
test('plan:pre has exactly 1 active hook and it belongs to tdd', () => {
const result = resolveLoopHooks({
point: 'plan:pre',
registry: realRegistry,
config: TDD_ON_CONFIG,
});
assert.ok(Array.isArray(result.activeHooks), 'activeHooks must be an array');
assert.strictEqual(
result.activeHooks.length,
1,
`plan:pre: expected 1 active hook (tdd), got ${result.activeHooks.length}: ${JSON.stringify(result.activeHooks.map(h => h.capId))}`,
);
assert.strictEqual(
result.activeHooks[0].capId,
'tdd',
`plan:pre: expected activeHooks[0].capId to be "tdd", got "${result.activeHooks[0].capId}"`,
);
});
test('execute:post has exactly 1 active hook and it belongs to tdd', () => {
const result = resolveLoopHooks({
point: 'execute:post',
registry: realRegistry,
config: TDD_ON_CONFIG,
});
assert.ok(Array.isArray(result.activeHooks), 'activeHooks must be an array');
assert.strictEqual(
result.activeHooks.length,
1,
`execute:post: expected 1 active hook (tdd gate), got ${result.activeHooks.length}: ${JSON.stringify(result.activeHooks.map(h => h.capId))}`,
);
assert.strictEqual(
result.activeHooks[0].capId,
'tdd',
`execute:post: expected activeHooks[0].capId to be "tdd", got "${result.activeHooks[0].capId}"`,
);
});
test('all 10 non-tdd points return activeHooks:[] even with tdd_mode ON', () => {
const failures = [];
for (const point of TDD_INACTIVE_POINTS) {
const result = resolveLoopHooks({
point,
registry: realRegistry,
config: TDD_ON_CONFIG,
});
assert.ok(Array.isArray(result.activeHooks), `${point}: activeHooks must be an array`);
if (result.activeHooks.length !== 0) {
failures.push({
point,
count: result.activeHooks.length,
capIds: result.activeHooks.map(h => h.capId),
});
}
}
assert.deepStrictEqual(
failures,
[],
`Expected 10 non-tdd points to be empty with tdd_mode ON, got non-zero at: ${JSON.stringify(failures)}`,
);
});
test('tdd hook at plan:pre is a contribution kind (not a gate or step)', () => {
const result = resolveLoopHooks({
point: 'plan:pre',
registry: realRegistry,
config: TDD_ON_CONFIG,
});
assert.strictEqual(
result.activeHooks.length,
1,
'Expected exactly 1 active hook at plan:pre with tdd ON',
);
assert.strictEqual(
result.activeHooks[0].kind,
'contribution',
`plan:pre tdd hook must be kind "contribution", got "${result.activeHooks[0].kind}"`,
);
});
test('tdd hook at execute:post is a gate kind (not a contribution or step)', () => {
const result = resolveLoopHooks({
point: 'execute:post',
registry: realRegistry,
config: TDD_ON_CONFIG,
});
assert.strictEqual(
result.activeHooks.length,
1,
'Expected exactly 1 active hook at execute:post with tdd ON',
);
assert.strictEqual(
result.activeHooks[0].kind,
'gate',
`execute:post tdd hook must be kind "gate", got "${result.activeHooks[0].kind}"`,
);
});
test('turning tdd_mode OFF restores both tdd points to activeHooks:[]', () => {
// Regression check: tdd_mode OFF → both previously-active points go back to empty
const tddOffConfig = buildAllFalseConfig(); // tdd_mode: false
for (const point of TDD_ACTIVE_POINTS) {
const result = resolveLoopHooks({
point,
registry: realRegistry,
config: tddOffConfig,
});
assert.strictEqual(
result.activeHooks.length,
0,
`${point}: expected activeHooks:[] with tdd_mode OFF, got ${JSON.stringify(result.activeHooks.map(h => h.capId))}`,
);
}
});
});