Files
msd-core/tests/init-debug.test.cjs
Tom Boucher e20744eacb enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick

ADR-3473 §8.4 says failure is a value. Three families currently encode failure as
success, and this commit pins each one RED before the fix lands.

Measured on this tree, 2026-08-26:

  gsd-tools generate-slug "test" --pick nonexistent
    -> empty stdout, exit 0                                     (#3365)

  gsd-tools audit-open --pick nonexistent_field
    -> dumps the entire human-readable audit report, exit 0

  gsd-tools generate-slug "Hello World" --raw --pick bogus
    -> prints "hello-world", another field's value, exit 0

  gsd-tools query state.planned-phase 3        (positional, no --phase)
    -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to
       "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted
       current_phase_name                                        (#3358)

tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract
("returns empty string for missing field", success === true). That assertion is
replaced by the required behavior rather than deleted.

The new parseNamedArgs block calls the spec-object signature that does not exist
yet, so it fails today by construction. The 11 existing behavior-lock tests are
left untouched here; they are corrected in the implementation commit.

C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180
Decision 4(b). A unit assertion on the parser would have passed throughout this
defect's life.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3884): failure is a value — strict argv, and --pick that signals absence

Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable
ways to say "I could not answer".

parseNamedArgs (src/command-arg-projection.cts)
  Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the
  hub's Result shape instead of a bare Record. Declaring the positional arity is what
  makes #3358's call site unrepresentable rather than merely detectable: an unrecognized
  flag or a token past the declared boundary is now InvalidArgs, naming the offending
  token and listing the accepted flags. The legacy positional-array call shape throws
  a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale
  hand-written .cjs call site fails loudly instead of destructuring undefined off a
  Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a
  projection over the one parser, not a second parser.

  Measured before, against a STATE.md with a populated phase-2 block:
    query state.planned-phase 3        (positional, no --phase)
    -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to
       "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted
       current_phase_name
  After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical.
  The flag form is unchanged and still updates STATE.md.

--pick <field> (gsd-core/bin/gsd-tools.cjs)
  extractField returns {found,value}, and the pick block no longer shares one catch
  between "output was not JSON" and "field was absent". An absent field exits 1 with
  pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1
  with pick_output_not_json instead of dumping the command's entire output. A field that
  is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer,
  not a failure, and it is what keeps `--pick count` printing 0 on a fresh project.

  Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable
  audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" —
  a different field's value, confidently, at exit 0.

  ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The
  sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the
  ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0
  would demote "could not answer" to "the answer is zero" — the hazard
  docs/how-to/resolve-unreachable-guard-findings.md already warns against.

Guard ledger (ADR-3473 Decision 6)
  scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a
  `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the
  correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls,
  a nullglob mechanism this change does not touch) is retained in full, as are the shared
  scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file
  is not deleted.

Call-site audit
  45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an
  if test, && chain, or a pipeline whose status is consumed, and no shell block in
  workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the
  prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind
  a prior found/existence check. No ADR-3409-class "field the command never produces"
  remains.

Design:      .gsd/phase/feat-3884-failure-is-a-value/40-design.md
Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows

Two review findings, both fixed here rather than recorded as limits.

1. A newline in an untrusted token forged a second stderr line.

   Before, plain-text mode:
     $ gsd-tools query state.planned-phase $'foo\nError: forged second line'
     Error: unexpected positional argument "foo
     Error: forged second line"

   After:
     Error: unexpected positional argument "foo\nError: forged second line"

   --json-errors mode was never affected — io.error runs that payload through
   JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the
   three new InvalidArgs reasons plus the two new --pick diagnostics all
   interpolate a token that comes straight from argv.

   Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at
   every interpolation site — not a copy per site. It is deliberately NOT
   applied inside error() itself: several callers in this tree emit intentional
   multi-line diagnostics, and escaping newlines there would mangle them.

   The available-top-level-keys list needed the same treatment for a reason the
   review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user
   document and echoes that document's own keys into the diagnostic. Verified
   reachable — a frontmatter key containing a newline reaches the key list — so
   formatKeyForDiagnosticList is guarding a live path, not a hypothetical one.
   Ordinary keys still render plain and unquoted; a fix that merely dropped the
   key would also have passed a "one line" assertion, so the test pins the
   escaped key's presence too.

2. Five behavior-table rows were implemented but nothing pinned them:
   B7  a dotted path that dies partway
   B9  bracket syntax on a non-array
   B10 a negative array index, in and out of range
   B14 a JSON root that is not an object
   B17 an @file: payload over 50KB

   B17 is the load-bearing one. output() writes @file:<path> instead of inline
   JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no
   test, a future reordering of those two steps turns every large result into a
   false pick_output_not_json. The fixture seeds 1200 phase directories and
   measures the payload at 62474 characters, asserting the spill actually
   happened rather than assuming it.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): correct the strict-argv surface against a full verification run

The first full run came back with 90 failures across 12 files, none in the new
tests. They were the argv surface telling me what it actually is. Ten root
causes; each classified before anything was changed.

I over-implemented, and that is reverted.

  ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional
  tokens". It says nothing about a value flag whose value is missing. Making
  that an error was my design decision, not the rule, and it broke a
  deliberately recorded contract: `--prd` with no value resolving to null
  (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5;
  tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The
  "requires a value" branch is deleted outright rather than kept behind an
  option — an unused strictness mode is speculative generality. Unknown-flag
  and unexpected-positional rejection, which is what §8.4 actually mandates,
  is unchanged.

--wave needed a third flag kind the original design did not anticipate.

  `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the
  shipped workflow reconstructs and passes it (execute-phase.md:84), while
  #2932 records token-PRESENCE semantics: the CLI cares only that the flag
  appeared, and the value belongs to the workflow layer. That is neither a
  boolean flag nor a value flag, so `optionalValueFlags` now exists —
  presence-only in `data`, and the validation cursor consumes a following
  non-flag token so it is not reported as a stray positional. Every other
  declared boolean flag was checked against every argument-hint and prose
  usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one
  of this shape.

Five tests were pinning forms that never worked.

  tests/adr857-core-without-capabilities.test.cjs passed
  `init plan-phase --phase 01-stub`, but the documented form is positional
  (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form
  is the literal string "--phase". Measured on the pre-fix build against a
  real .planning/phases/01-stub/ directory:

    init plan-phase 01-stub          -> phase_found=true
    init plan-phase --phase 01-stub  -> phase_found=false

  The test asserted only exit 0 and key presence, so it had been green while
  proving nothing about phase resolution. Corrected to the documented form and
  strengthened to assert phase_found === true. Same class in state.test.cjs
  (`--plan-count`, a flag that does not exist; the real one is `--plans`),
  milestone-archive.test.cjs (`init new-milestone --json`, silently ignored),
  and concurrency-safety.test.cjs (a bare positional field name whose
  OR-assertion passed because a whole-document dump happens to contain the
  substring it looked for).

Six handlers had no argv validation at all — the same #3358 shape this phase
exists to close, found while fixing the rest: init verify-work / phase-op /
review / todos / remove-workspace read args[2] with nothing checking the rest,
and validate health read --repair/--backfill through a bare args.includes()
scan that bypassed the parser entirely. All now go through the seam, so the
flag has one owner.

tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must
NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's
local expectation does not override §8, so they are inverted and renamed —
a test still called "ignores an unrecognized flag" while asserting rejection
would be its own defect. Row C6's point is its PWNED canary; that assertion is
kept verbatim and only its exit-status expectation changed, because the
hostile token is now rejected rather than absorbed.

The blast-radius estimate in 40-design.md is corrected rather than quietly
left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was
accurate for what the graph can see — parseNamedArgs's callers. It cannot see
that those callers' handlers accept argv shapes wider than the code reading
args[2] suggests, which is where the real surface was.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert

Second full run: 46 failures, down from 90. Four causes, two of them mine.

Reverted `validate health` entirely — it was scope creep, and it broke a real flag.

  ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`.
  The previous commit routed `validate health` through the parser on the
  reasoning that a flag should have one owner. That was wrong twice over:
  §8.4 names parseNamedArgs and count queries, and `validate health` was never
  a parseNamedArgs call site — it read its flags, just not through the parser,
  so it had no silent-drop defect to fix. Tightening it omitted `--json`, which
  the health-diagnostic suites use heavily. The handler is now byte-for-behaviour
  back to its pre-branch form. `validate context` stays converted: it genuinely
  was a call site, and its `--json` is now declared rather than read by a second
  `args.includes` scan.

  The five handlers that had NO validation at all — init verify-work / phase-op /
  review / todos / remove-workspace — stay fixed. Those read args[2] with nothing
  checking the rest, which is the #3358 shape this phase owns.

Finished the A2/A3 revert. Three tests still encoded the deleted
"a value flag with a missing value is an error" rule, including one added by the
previous commit for that rule. All three now assert the reverted null contract,
and the ones whose titles said "rejected" are renamed — a test named for a
contract it no longer asserts is its own defect.

`--wave=` and `--wave --weird` are correctly rejected. Neither is documented in
commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and
neither is emitted by the shipped prompt layer, so both are unrecognized tokens
that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the
property it exists for — asserted directly now, at the parser, that `--wave` does
not swallow a following flag as its value — and only its exit-status expectation
changed.

A contradiction inside this branch, surfaced by the audit and resolved the safe way.

  Two pre-existing #3573 tests call `state begin-phase '2'` and
  `state planned-phase '2'` with a bare positional, relying on the old permissive
  parser to ignore it. This branch's own #3358 regression test requires that exact
  argv to be REJECTED. The two are mutually exclusive.

  Widening the router to accept a bare positional — mirroring complete-phase —
  would have silently re-opened #3358, and was verified to do exactly that: with
  the widened router, `query state.planned-phase 3` returned exit 0 and wrote
  current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and
  docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the
  two #3573 tests move to it. Their assertions were never about the call shape —
  only that total_phases survives the resync — and both still pass.

  complete-phase is untouched: its bare positional IS documented, and it keeps the
  dynamic boundary and the negative-space note that record why.

The audit that produced this is in the PR body: for every handler whose declaration
changed, the flags it reads anywhere in its body, the flags the shipped surface
documents, and the shapes the suite passes, compared. The `--json` miss was a
pattern, not an accident — declaring a handler's flags from its parseNamedArgs call
alone misses whatever it reads elsewhere.

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3884): backfill the changeset PR number

Refs #3884

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 00:12:13 -04:00

482 lines
20 KiB
JavaScript

'use strict';
/**
* `init.debug` — the dedicated init entry point for `/gsd:debug` (#3149).
*
* Prerequisite for #3128 condition 1: ADR-1671 admission gate (2), "a fact the
* init seam demonstrably computes at a real entry point"
* (`docs/adr/1671-dynamic-context-management-platform.md:122-131`). Before this,
* `gsd-core/workflows/debug.md` was one of the last workflows with no `cmdInit*`
* of its own, so no debug-scoped fact could ever be computed and any `when=` atom
* naming one would have evaluated FALSE forever — the silent-exclusion bug that
* rule exists to prevent.
*
* Matrix: `.gsd/phase/feat-3149-cmdinitdebug/50-test-matrix.md` groups A-E, G3.
*
* Every test drives the REAL CLI (`runGsdTools` spawns `gsd-tools.cjs`) rather
* than requiring `cmdInitDebug` directly — the handler is not exported, and the
* flag plumbing under test exists only at the `init-command-router.cjs` seam.
* Same rationale recorded in `tests/section-manifest-init-facts.test.cjs:10-14`.
*
* Group A is the load-bearing half: this change's entire claim is "one round-trip
* instead of three, with identical resolved values", so each A-row cross-checks
* `init.debug` against the exact command it replaced.
*/
const { describe, test, beforeEach, afterEach } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('fs');
const path = require('path');
const { runGsdTools, cleanup, createTempDir, createTempProject } = require('./helpers.cjs');
function writeConfig(tmpDir, config, { ws = null } = {}) {
const dir = ws
? path.join(tmpDir, '.planning', 'workstreams', ws)
: path.join(tmpDir, '.planning');
fs.mkdirSync(dir, { recursive: true });
fs.writeFileSync(path.join(dir, 'config.json'), JSON.stringify(config, null, 2));
}
/** Runs a gsd-tools query and parses its JSON, asserting a clean exit first. */
function runJson(argv, cwd, env = {}) {
const result = runGsdTools(argv, cwd, env);
assert.ok(result.success, `Command failed: ${result.error}`);
return JSON.parse(result.output);
}
// ─── Group A: equivalence with the three calls init.debug replaces ──────────
describe('init.debug resolves identically to the three calls it replaces (matrix §A)', () => {
let tmpDir;
beforeEach(() => {
tmpDir = createTempProject('init-debug-a-');
});
afterEach(() => {
cleanup(tmpDir);
});
test('debug_dir matches state.load exactly (row A1)', () => {
const viaInit = runJson(['init', 'debug'], tmpDir);
const viaState = runJson(['query', 'state.load'], tmpDir);
assert.equal(
viaInit.debug_dir,
viaState.debug_dir,
'init.debug must resolve the same debug directory state.load does — debug.md builds ' +
'debug_file_path from it (#2376) and a divergence silently writes sessions elsewhere'
);
});
test('debug_dir agrees with state.load under an active workstream (row A2)', () => {
fs.mkdirSync(path.join(tmpDir, '.planning', 'workstreams', 'ws1'), { recursive: true });
const viaInit = runJson(['init', 'debug'], tmpDir, { GSD_WORKSTREAM: 'ws1' });
const viaState = runJson(['query', 'state.load'], tmpDir, { GSD_WORKSTREAM: 'ws1' });
assert.equal(viaInit.debug_dir, viaState.debug_dir);
assert.match(
viaInit.debug_dir,
/\/workstreams\/ws1\/debug$/,
'an active workstream must scope debug_dir into that workstream, not the project root'
);
});
test('commit_docs matches state.load (row A3)', () => {
writeConfig(tmpDir, { planning: { commit_docs: false } });
const viaInit = runJson(['init', 'debug'], tmpDir);
const viaState = runJson(['query', 'state.load'], tmpDir);
assert.equal(viaInit.commit_docs, viaState.config.commit_docs);
assert.equal(viaInit.commit_docs, false, 'sanity: the configured value, not the default');
});
test('response_language matches state.load config (row A4)', () => {
writeConfig(tmpDir, { response_language: 'es' });
const viaInit = runJson(['init', 'debug'], tmpDir);
const viaState = runJson(['query', 'state.load'], tmpDir);
assert.equal(viaInit.response_language, viaState.config.response_language);
assert.equal(viaInit.response_language, 'es');
});
test('debugger_model matches the resolve-model query (row A5)', () => {
const viaInit = runJson(['init', 'debug'], tmpDir);
const viaResolve = runJson(['query', 'resolve-model', 'gsd-debugger'], tmpDir);
assert.equal(
viaInit.debugger_model,
viaResolve.model,
'debug.md omits the model param when this is empty or "inherit" (#2517) — the value ' +
'must be the same one resolve-model produced, not a re-derived default'
);
});
test('tdd_mode matches config-get when set (row A6)', () => {
writeConfig(tmpDir, { workflow: { tdd_mode: true } });
const viaInit = runJson(['init', 'debug'], tmpDir);
const viaConfigGet = runGsdTools(['query', 'config-get', 'workflow.tdd_mode', '--raw'], tmpDir);
assert.ok(viaConfigGet.success);
assert.equal(viaInit.tdd_mode, true);
assert.equal(String(viaInit.tdd_mode), viaConfigGet.output.trim());
});
test('tdd_mode matches config-get when the key is absent (row A7)', () => {
writeConfig(tmpDir, {});
const viaInit = runJson(['init', 'debug'], tmpDir);
const viaConfigGet = runGsdTools(['query', 'config-get', 'workflow.tdd_mode', '--raw'], tmpDir);
assert.equal(viaInit.tdd_mode, false);
assert.equal(String(viaInit.tdd_mode), viaConfigGet.output.trim());
});
test('tdd_mode matches config-get under workstream inheritance (row A8)', () => {
// The one case where the two resolution paths could genuinely disagree:
// `config-get` inherits from the ROOT config when an active workstream has
// no config.json of its own (#2702, src/config.cts), while the init seam
// reads loadConfig's root+workstream merge (src/config-loader.cts). Both
// must land on the same boolean or the consolidation changes behavior for
// workstream users.
writeConfig(tmpDir, { workflow: { tdd_mode: true } });
fs.mkdirSync(path.join(tmpDir, '.planning', 'workstreams', 'ws1'), { recursive: true });
const viaInit = runJson(['init', 'debug'], tmpDir, { GSD_WORKSTREAM: 'ws1' });
const viaConfigGet = runGsdTools(
['query', 'config-get', 'workflow.tdd_mode', '--raw'],
tmpDir,
{ GSD_WORKSTREAM: 'ws1' }
);
assert.equal(viaInit.tdd_mode, true, 'the root value must be inherited, not lost');
assert.equal(String(viaInit.tdd_mode), viaConfigGet.output.trim());
});
test('honors workflow.tdd_mode, ignores a bare top-level tdd_mode (row A9)', () => {
// The invariant tests/debug-session-management.test.cjs used to guard by
// grepping debug.md for `config-get workflow.tdd_mode`. Asserted here
// behaviorally instead, which is strictly stronger: a bare top-level key
// must NOT be honored, whatever the read mechanism.
writeConfig(tmpDir, { tdd_mode: true });
assert.equal(
runJson(['init', 'debug'], tmpDir).tdd_mode,
false,
'a bare top-level tdd_mode key must be ignored — the canonical key is workflow.tdd_mode'
);
writeConfig(tmpDir, { workflow: { tdd_mode: true } });
assert.equal(
runJson(['init', 'debug'], tmpDir).tdd_mode,
true,
'the canonical workflow.tdd_mode key must be honored'
);
});
});
// ─── Group B: bundle shape ─────────────────────────────────────────────────
describe('init.debug bundle shape (matrix §B)', () => {
let tmpDir;
beforeEach(() => {
tmpDir = createTempProject('init-debug-b-');
});
afterEach(() => {
cleanup(tmpDir);
});
test('emits the documented field set (row B1)', () => {
const output = runJson(['init', 'debug'], tmpDir);
for (const key of ['project_root', 'debug_dir', 'commit_docs', 'debugger_model', 'tdd_mode', 'diagnose']) {
assert.ok(
Object.prototype.hasOwnProperty.call(output, key),
`init.debug must emit "${key}"`
);
}
assert.ok(
Object.prototype.hasOwnProperty.call(output, 'section_manifest'),
'section_manifest must be present even when it degrades to null'
);
});
test('omits response_language entirely when unset (row B2)', () => {
writeConfig(tmpDir, {});
const output = runJson(['init', 'debug'], tmpDir);
assert.equal(
Object.prototype.hasOwnProperty.call(output, 'response_language'),
false,
'withProjectRoot injects response_language ONLY when configured — an absent key means ' +
'"English", and emitting null/"" instead would make absence look like a degraded read'
);
});
test('debug_dir is an absolute POSIX path (row B3)', () => {
const output = runJson(['init', 'debug'], tmpDir);
assert.equal(output.debug_dir.includes('\\'), false, 'no backslash separators (#2376)');
assert.ok(output.debug_dir.endsWith('/debug'), 'points at the debug directory');
assert.notEqual(output.debug_dir, 'debug');
assert.notEqual(output.debug_dir, '.planning/debug', 'must be absolute, never a bare relative literal');
});
test('succeeds with no .planning directory (row B4)', () => {
const bare = createTempDir('init-debug-bare-');
try {
const result = runGsdTools(['init', 'debug'], bare);
assert.ok(result.success, `must not require an initialized project: ${result.error}`);
const output = JSON.parse(result.output);
assert.ok(output.debug_dir.endsWith('/debug'));
} finally {
cleanup(bare);
}
});
test('succeeds on an empty config object (row B5)', () => {
writeConfig(tmpDir, {});
const output = runJson(['init', 'debug'], tmpDir);
assert.equal(output.tdd_mode, false);
});
test('survives valid-JSON-not-an-object config (row B6)', () => {
// Valid JSON that is not an object is the input class nobody enumerates:
// every one of these parses cleanly and then fails on property access.
for (const body of ['0', '"str"', '[]', 'null', 'true']) {
fs.mkdirSync(path.join(tmpDir, '.planning'), { recursive: true });
fs.writeFileSync(path.join(tmpDir, '.planning', 'config.json'), body);
const result = runGsdTools(['init', 'debug'], tmpDir);
assert.ok(result.success, `config.json = ${body} must degrade, not crash: ${result.error}`);
const output = JSON.parse(result.output);
assert.equal(output.tdd_mode, false, `config.json = ${body} must resolve tdd_mode to false`);
assert.ok(output.debug_dir.endsWith('/debug'));
}
});
test('survives a present-but-empty config file (row B7)', () => {
fs.mkdirSync(path.join(tmpDir, '.planning'), { recursive: true });
fs.writeFileSync(path.join(tmpDir, '.planning', 'config.json'), '');
const result = runGsdTools(['init', 'debug'], tmpDir);
assert.ok(result.success, `an empty config file must degrade, not crash: ${result.error}`);
assert.equal(JSON.parse(result.output).tdd_mode, false);
});
test('is insensitive to CRLF in config.json (row B8)', () => {
fs.mkdirSync(path.join(tmpDir, '.planning'), { recursive: true });
fs.writeFileSync(
path.join(tmpDir, '.planning', 'config.json'),
'{\r\n "workflow": {\r\n "tdd_mode": true\r\n }\r\n}\r\n'
);
const output = runJson(['init', 'debug'], tmpDir);
assert.equal(output.tdd_mode, true, 'CRLF must not change how the config parses');
});
});
// ─── Group C: --diagnose forwarding + CLI negative matrix ──────────────────
describe('init.debug --diagnose forwarding and hostile argv (matrix §C)', () => {
let tmpDir;
beforeEach(() => {
tmpDir = createTempProject('init-debug-c-');
});
afterEach(() => {
cleanup(tmpDir);
});
test('--diagnose surfaces as diagnose:true (row C1)', () => {
const output = runJson(['init', 'debug', '--diagnose'], tmpDir);
assert.equal(output.diagnose, true);
});
test('absent --diagnose is false, not undefined (row C2)', () => {
const output = runJson(['init', 'debug'], tmpDir);
assert.equal(output.diagnose, false);
assert.notEqual(output.diagnose, undefined, 'parseNamedArgs materializes false; never leak undefined');
});
test('duplicate --diagnose is idempotent (row C3)', () => {
const output = runJson(['init', 'debug', '--diagnose', '--diagnose'], tmpDir);
assert.equal(output.diagnose, true);
});
test('rejects an unrecognized flag with the flag named (row C4)', () => {
// ADR-3473 §8.4 (Bucket-B correction): "`parseNamedArgs` rejects
// unrecognized ... tokens with a non-zero exit — it is called by agents
// that will drift again." An unrecognized flag is exactly the mandated
// rejection, not a thing to silently absorb.
const result = runGsdTools(['init', 'debug', '--nope'], tmpDir);
assert.equal(result.success, false, 'an unknown flag must now fail the command');
assert.match(result.error, /--nope/, 'the rejection must name the offending flag');
});
test('rejects a flag-shaped trailing token, naming it (row C5)', () => {
const result = runGsdTools(['init', 'debug', '--diagnose', '--weird'], tmpDir);
assert.equal(result.success, false, 'an unrecognized flag-shaped token must now fail the command');
assert.match(result.error, /--weird/, 'the rejection must name the offending flag');
});
test('does not interpolate shell metacharacters even though the hostile positional is now rejected (row C6)', () => {
const canary = path.join(tmpDir, 'PWNED');
const hostile = `; touch ${canary}; $(touch ${canary}) \`touch ${canary}\` && touch ${canary}`;
const result = runGsdTools(['init', 'debug', hostile], tmpDir);
// §8.4 now rejects this as an unexpected positional argument (exit
// non-zero) instead of silently absorbing it — that is at least as safe
// as the old accept-and-ignore behavior. The canary assertion is the
// actual point of this test and is unchanged: no shell ever touches this
// string, whether the token is accepted or rejected.
assert.equal(result.success, false, 'a stray positional argument must now fail the command');
assert.equal(fs.existsSync(canary), false, 'no shell interpolation of an attacker-controlled argument');
assert.equal(result.error.includes(' at '), false, 'no stack trace in non-debug output');
});
test('rejects a very long or unicode positional argument, not just tolerates it (row C7/C8)', () => {
// Classified as the same §8.4 unexpected-positional-argument shape as
// C6: `init debug` declares no positionals, so any bare token here is a
// stray positional and must now be rejected rather than silently
// absorbed.
const long = 'x'.repeat(8192);
const unicode = 'ünïcødé-🐛-测试';
for (const arg of [long, unicode]) {
const result = runGsdTools(['init', 'debug', arg], tmpDir);
assert.equal(result.success, false, `argument of length ${arg.length} must now fail the command, not crash`);
}
});
});
// ─── Group D: section_manifest, null vs [] ─────────────────────────────────
describe('init.debug section_manifest degradation (matrix §D)', () => {
let tmpDir;
let manifestDir;
beforeEach(() => {
tmpDir = createTempProject('init-debug-d-');
manifestDir = createTempDir('init-debug-d-manifest-');
});
afterEach(() => {
cleanup(tmpDir);
cleanup(manifestDir);
});
function withManifest(body) {
const manifestPath = path.join(manifestDir, 'manifest.json');
fs.writeFileSync(manifestPath, typeof body === 'string' ? body : JSON.stringify(body));
return { GSD_SECTION_MANIFEST: manifestPath };
}
test('section_manifest is null while debug has no manifest key (row D1)', () => {
// Drives the SHIPPED artifact deliberately: `debug` carries no gsd:section
// markers until #3128, so the shipped manifest has no `debug` key and the
// field must degrade to null — which debug.md reads as "read everything".
const output = runJson(['init', 'debug'], tmpDir);
assert.equal(output.section_manifest, null);
});
test('an explicit empty debug key computes [], not null (row D2)', () => {
const output = runJson(['init', 'debug'], tmpDir, withManifest({ workflows: { debug: [] } }));
assert.notEqual(output.section_manifest, null, 'a present key must never collapse to the degraded value');
assert.deepEqual(output.section_manifest.included, []);
assert.deepEqual(output.section_manifest.excluded, []);
});
test('selects an always-section for the debug workflow (row D3)', () => {
// Proves the workflow key really is 'debug' — a handler passing the wrong
// name would silently return null forever and look identical to D1.
const output = runJson(['init', 'debug'], tmpDir, withManifest({
workflows: {
debug: [{ id: 'probe-protocol', when: 'always', read: 'gsd-core/workflows/debug/steps/probe-protocol.md' }],
},
}));
assert.notEqual(output.section_manifest, null);
assert.equal(output.section_manifest.workflow, 'debug');
assert.deepEqual(output.section_manifest.included, ['probe-protocol']);
assert.deepEqual(output.section_manifest.read, ['gsd-core/workflows/debug/steps/probe-protocol.md']);
});
test('a missing manifest file degrades to null (row D4)', () => {
const missing = path.join(manifestDir, 'does-not-exist.json');
assert.equal(fs.existsSync(missing), false, 'sanity: file must not exist');
const result = runGsdTools(['init', 'debug'], tmpDir, { GSD_SECTION_MANIFEST: missing });
assert.ok(result.success, `a missing manifest must not crash: ${result.error}`);
assert.equal(JSON.parse(result.output).section_manifest, null);
});
test('a malformed manifest degrades to null (row D5)', () => {
const result = runGsdTools(['init', 'debug'], tmpDir, withManifest('{ not json'));
assert.ok(result.success, `a malformed manifest must not crash: ${result.error}`);
assert.equal(JSON.parse(result.output).section_manifest, null);
});
test('a pre-6.1 flat manifest shape degrades to null (row D6)', () => {
// The pre-#2992 shape had no workflow key at all. Accepting it would
// mis-attribute some other workflow's sections to debug.
const result = runGsdTools(['init', 'debug'], tmpDir, withManifest({ sections: [{ id: 'x', when: 'always' }] }));
assert.ok(result.success);
assert.equal(JSON.parse(result.output).section_manifest, null);
});
});
// ─── Group E: PlanningPaths.debug ──────────────────────────────────────────
describe('planningPaths exposes the debug directory (matrix §E)', () => {
const { planningPaths } = require('../gsd-core/bin/lib/planning-workspace.cjs');
let tmpDir;
beforeEach(() => {
tmpDir = createTempDir('init-debug-e-');
});
afterEach(() => {
cleanup(tmpDir);
});
test('planningPaths exposes debug (row E1)', () => {
assert.equal(planningPaths(tmpDir).debug, path.join(tmpDir, '.planning', 'debug'));
});
test('planningPaths.debug is workstream-scoped (row E2)', () => {
assert.equal(
planningPaths(tmpDir, 'feature-x').debug,
path.join(tmpDir, '.planning', 'workstreams', 'feature-x', 'debug')
);
});
test('planningPaths.debug does not weaken the traversal guard (row E4)', () => {
assert.throws(() => planningPaths(tmpDir, '../../etc'), /invalid path characters/);
assert.throws(() => planningPaths(tmpDir, 'foo/bar'), /invalid path characters/);
});
});
// ─── Group G: regressions this change must not cause ───────────────────────
describe('init.debug does not widen the applicability grammar (matrix §G)', () => {
test('WHEN_VOCABULARY is unchanged at 29 entries (row G3)', () => {
// ADR-1671: the vocabulary is CLOSED and widening it is a coordinated
// amendment. This PR delivers admission gate (2) only — the atom that
// consumes it belongs to #3128, which owns the amendment.
const { WHEN_VOCABULARY } = require('../gsd-core/bin/lib/workflow-fragments.cjs');
assert.equal(WHEN_VOCABULARY.length, 29);
assert.equal(WHEN_VOCABULARY.includes('flag:--diagnose'), false);
assert.equal(WHEN_VOCABULARY.includes('flag:--runtime-probes'), false);
});
});