* test(#3884): failing-first coverage for strict argv and absence-signalling --pick ADR-3473 §8.4 says failure is a value. Three families currently encode failure as success, and this commit pins each one RED before the fix lands. Measured on this tree, 2026-08-26: gsd-tools generate-slug "test" --pick nonexistent -> empty stdout, exit 0 (#3365) gsd-tools audit-open --pick nonexistent_field -> dumps the entire human-readable audit report, exit 0 gsd-tools generate-slug "Hello World" --raw --pick bogus -> prints "hello-world", another field's value, exit 0 gsd-tools query state.planned-phase 3 (positional, no --phase) -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted current_phase_name (#3358) tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract ("returns empty string for missing field", success === true). That assertion is replaced by the required behavior rather than deleted. The new parseNamedArgs block calls the spec-object signature that does not exist yet, so it fails today by construction. The 11 existing behavior-lock tests are left untouched here; they are corrected in the implementation commit. C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180 Decision 4(b). A unit assertion on the parser would have passed throughout this defect's life. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3884): failure is a value — strict argv, and --pick that signals absence Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable ways to say "I could not answer". parseNamedArgs (src/command-arg-projection.cts) Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the hub's Result shape instead of a bare Record. Declaring the positional arity is what makes #3358's call site unrepresentable rather than merely detectable: an unrecognized flag or a token past the declared boundary is now InvalidArgs, naming the offending token and listing the accepted flags. The legacy positional-array call shape throws a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale hand-written .cjs call site fails loudly instead of destructuring undefined off a Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a projection over the one parser, not a second parser. Measured before, against a STATE.md with a populated phase-2 block: query state.planned-phase 3 (positional, no --phase) -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted current_phase_name After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical. The flag form is unchanged and still updates STATE.md. --pick <field> (gsd-core/bin/gsd-tools.cjs) extractField returns {found,value}, and the pick block no longer shares one catch between "output was not JSON" and "field was absent". An absent field exits 1 with pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1 with pick_output_not_json instead of dumping the command's entire output. A field that is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer, not a failure, and it is what keeps `--pick count` printing 0 on a fresh project. Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" — a different field's value, confidently, at exit 0. ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0 would demote "could not answer" to "the answer is zero" — the hazard docs/how-to/resolve-unreachable-guard-findings.md already warns against. Guard ledger (ADR-3473 Decision 6) scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls, a nullglob mechanism this change does not touch) is retained in full, as are the shared scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file is not deleted. Call-site audit 45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an if test, && chain, or a pipeline whose status is consumed, and no shell block in workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind a prior found/existence check. No ADR-3409-class "field the command never produces" remains. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows Two review findings, both fixed here rather than recorded as limits. 1. A newline in an untrusted token forged a second stderr line. Before, plain-text mode: $ gsd-tools query state.planned-phase $'foo\nError: forged second line' Error: unexpected positional argument "foo Error: forged second line" After: Error: unexpected positional argument "foo\nError: forged second line" --json-errors mode was never affected — io.error runs that payload through JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the three new InvalidArgs reasons plus the two new --pick diagnostics all interpolate a token that comes straight from argv. Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at every interpolation site — not a copy per site. It is deliberately NOT applied inside error() itself: several callers in this tree emit intentional multi-line diagnostics, and escaping newlines there would mangle them. The available-top-level-keys list needed the same treatment for a reason the review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user document and echoes that document's own keys into the diagnostic. Verified reachable — a frontmatter key containing a newline reaches the key list — so formatKeyForDiagnosticList is guarding a live path, not a hypothetical one. Ordinary keys still render plain and unquoted; a fix that merely dropped the key would also have passed a "one line" assertion, so the test pins the escaped key's presence too. 2. Five behavior-table rows were implemented but nothing pinned them: B7 a dotted path that dies partway B9 bracket syntax on a non-array B10 a negative array index, in and out of range B14 a JSON root that is not an object B17 an @file: payload over 50KB B17 is the load-bearing one. output() writes @file:<path> instead of inline JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no test, a future reordering of those two steps turns every large result into a false pick_output_not_json. The fixture seeds 1200 phase directories and measures the payload at 62474 characters, asserting the spill actually happened rather than assuming it. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): correct the strict-argv surface against a full verification run The first full run came back with 90 failures across 12 files, none in the new tests. They were the argv surface telling me what it actually is. Ten root causes; each classified before anything was changed. I over-implemented, and that is reverted. ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional tokens". It says nothing about a value flag whose value is missing. Making that an error was my design decision, not the rule, and it broke a deliberately recorded contract: `--prd` with no value resolving to null (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5; tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The "requires a value" branch is deleted outright rather than kept behind an option — an unused strictness mode is speculative generality. Unknown-flag and unexpected-positional rejection, which is what §8.4 actually mandates, is unchanged. --wave needed a third flag kind the original design did not anticipate. `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the shipped workflow reconstructs and passes it (execute-phase.md:84), while #2932 records token-PRESENCE semantics: the CLI cares only that the flag appeared, and the value belongs to the workflow layer. That is neither a boolean flag nor a value flag, so `optionalValueFlags` now exists — presence-only in `data`, and the validation cursor consumes a following non-flag token so it is not reported as a stray positional. Every other declared boolean flag was checked against every argument-hint and prose usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one of this shape. Five tests were pinning forms that never worked. tests/adr857-core-without-capabilities.test.cjs passed `init plan-phase --phase 01-stub`, but the documented form is positional (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form is the literal string "--phase". Measured on the pre-fix build against a real .planning/phases/01-stub/ directory: init plan-phase 01-stub -> phase_found=true init plan-phase --phase 01-stub -> phase_found=false The test asserted only exit 0 and key presence, so it had been green while proving nothing about phase resolution. Corrected to the documented form and strengthened to assert phase_found === true. Same class in state.test.cjs (`--plan-count`, a flag that does not exist; the real one is `--plans`), milestone-archive.test.cjs (`init new-milestone --json`, silently ignored), and concurrency-safety.test.cjs (a bare positional field name whose OR-assertion passed because a whole-document dump happens to contain the substring it looked for). Six handlers had no argv validation at all — the same #3358 shape this phase exists to close, found while fixing the rest: init verify-work / phase-op / review / todos / remove-workspace read args[2] with nothing checking the rest, and validate health read --repair/--backfill through a bare args.includes() scan that bypassed the parser entirely. All now go through the seam, so the flag has one owner. tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's local expectation does not override §8, so they are inverted and renamed — a test still called "ignores an unrecognized flag" while asserting rejection would be its own defect. Row C6's point is its PWNED canary; that assertion is kept verbatim and only its exit-status expectation changed, because the hostile token is now rejected rather than absorbed. The blast-radius estimate in 40-design.md is corrected rather than quietly left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was accurate for what the graph can see — parseNamedArgs's callers. It cannot see that those callers' handlers accept argv shapes wider than the code reading args[2] suggests, which is where the real surface was. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert Second full run: 46 failures, down from 90. Four causes, two of them mine. Reverted `validate health` entirely — it was scope creep, and it broke a real flag. ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`. The previous commit routed `validate health` through the parser on the reasoning that a flag should have one owner. That was wrong twice over: §8.4 names parseNamedArgs and count queries, and `validate health` was never a parseNamedArgs call site — it read its flags, just not through the parser, so it had no silent-drop defect to fix. Tightening it omitted `--json`, which the health-diagnostic suites use heavily. The handler is now byte-for-behaviour back to its pre-branch form. `validate context` stays converted: it genuinely was a call site, and its `--json` is now declared rather than read by a second `args.includes` scan. The five handlers that had NO validation at all — init verify-work / phase-op / review / todos / remove-workspace — stay fixed. Those read args[2] with nothing checking the rest, which is the #3358 shape this phase owns. Finished the A2/A3 revert. Three tests still encoded the deleted "a value flag with a missing value is an error" rule, including one added by the previous commit for that rule. All three now assert the reverted null contract, and the ones whose titles said "rejected" are renamed — a test named for a contract it no longer asserts is its own defect. `--wave=` and `--wave --weird` are correctly rejected. Neither is documented in commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and neither is emitted by the shipped prompt layer, so both are unrecognized tokens that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the property it exists for — asserted directly now, at the parser, that `--wave` does not swallow a following flag as its value — and only its exit-status expectation changed. A contradiction inside this branch, surfaced by the audit and resolved the safe way. Two pre-existing #3573 tests call `state begin-phase '2'` and `state planned-phase '2'` with a bare positional, relying on the old permissive parser to ignore it. This branch's own #3358 regression test requires that exact argv to be REJECTED. The two are mutually exclusive. Widening the router to accept a bare positional — mirroring complete-phase — would have silently re-opened #3358, and was verified to do exactly that: with the widened router, `query state.planned-phase 3` returned exit 0 and wrote current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the two #3573 tests move to it. Their assertions were never about the call shape — only that total_phases survives the resync — and both still pass. complete-phase is untouched: its bare positional IS documented, and it keeps the dynamic boundary and the negative-space note that record why. The audit that produced this is in the PR body: for every handler whose declaration changed, the flags it reads anywhere in its body, the flags the shipped surface documents, and the shapes the suite passes, compared. The `--json` miss was a pattern, not an accident — declaring a handler's flags from its parseNamedArgs call alone misses whatever it reads elsewhere. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3884): backfill the changeset PR number Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
441 lines
20 KiB
JavaScript
441 lines
20 KiB
JavaScript
/**
|
|
* GSD Tools Tests - --pick flag
|
|
*
|
|
* Regression tests for the --pick CLI flag that extracts a single field
|
|
* from JSON output, replacing the need for jq as an external dependency.
|
|
*/
|
|
|
|
const { test, describe } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
const { runGsdTools, createTempProject, cleanup } = require('./helpers.cjs');
|
|
const { seedPhase } = require('./fixtures/index.cjs');
|
|
const { runCli } = require('./helpers/cli-negative.cjs');
|
|
|
|
// ─── --pick flag ─────────────────────────────────────────────────────────────
|
|
|
|
describe('--pick flag', () => {
|
|
test('extracts a top-level field from JSON output', () => {
|
|
const result = runGsdTools('generate-slug "hello world" --pick slug');
|
|
assert.strictEqual(result.success, true);
|
|
assert.strictEqual(result.output, 'hello-world');
|
|
});
|
|
|
|
test('extracts a top-level field using array args', () => {
|
|
const result = runGsdTools(['generate-slug', 'hello world', '--pick', 'slug']);
|
|
assert.strictEqual(result.success, true);
|
|
assert.strictEqual(result.output, 'hello-world');
|
|
});
|
|
|
|
// #3365 / ADR-3473 §8.4 P6: an ABSENT field is a failure ("I could not
|
|
// answer"), never a demotion to the empty answer at exit 0. This inverts
|
|
// the old pinned assertion below (kept as a comment for the historical
|
|
// record — measured on this tree, 2026-08-26, exit 0 + empty stdout):
|
|
// const result = runGsdTools('generate-slug "test" --pick nonexistent');
|
|
// assert.strictEqual(result.success, true);
|
|
// assert.strictEqual(result.output, '');
|
|
test('absentFieldExitsNonZero_3365', () => {
|
|
const result = runGsdTools('generate-slug "test" --pick nonexistent');
|
|
assert.strictEqual(result.success, false, 'an absent --pick field must exit non-zero');
|
|
assert.strictEqual(result.output, '');
|
|
assert.match(result.error, /nonexistent/, 'stderr must name the requested field');
|
|
});
|
|
|
|
// P2 (test matrix): a count of zero is a real value, not absence — this
|
|
// must keep PASSING before and after the fix (the non-change half of #3365).
|
|
test('zeroCountPrintsZeroAtExitZero', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const result = runGsdTools('query phases.list --type summaries --pick count', tmpDir);
|
|
assert.strictEqual(result.success, true);
|
|
assert.strictEqual(result.output, '0');
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
|
|
// P4 (test matrix, negative space N1): a field present with an explicit
|
|
// `null` value is an answer, not a failure — must keep PASSING before and
|
|
// after the fix. Measured: `phases.list --type plans --pick phase_dir` on
|
|
// the enumeration path (no --phase given) returns `phase_dir: null`.
|
|
test('presentButNullIsEmptyAtExitZero', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const result = runGsdTools('query phases.list --type plans --pick phase_dir', tmpDir);
|
|
assert.strictEqual(result.success, true);
|
|
assert.strictEqual(result.output, '');
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
|
|
// P10 (test matrix): required boundary triple over an array field of
|
|
// known length N=3 (three seeded phase directories).
|
|
test('arrayIndexBoundaryAtLenMinus1_Len_LenPlus1', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
seedPhase(tmpDir, '01-alpha');
|
|
seedPhase(tmpDir, '02-beta');
|
|
seedPhase(tmpDir, '03-gamma');
|
|
|
|
const atLenMinus1 = runGsdTools('query phases.list --pick directories[2]', tmpDir);
|
|
assert.strictEqual(atLenMinus1.success, true);
|
|
assert.strictEqual(atLenMinus1.output, '03-gamma');
|
|
|
|
const atLen = runGsdTools('query phases.list --pick directories[3]', tmpDir);
|
|
assert.strictEqual(atLen.success, false, 'index == length is out of range and must exit non-zero');
|
|
|
|
const atLenPlus1 = runGsdTools('query phases.list --pick directories[4]', tmpDir);
|
|
assert.strictEqual(atLenPlus1.success, false, 'index == length + 1 is out of range and must exit non-zero');
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
|
|
// P12 (test matrix): the measured B11 defect — a non-JSON command's
|
|
// `--pick` must never dump the whole document as a coincidental "success".
|
|
// TODAY (measured on this tree, 2026-08-26), against an empty temp project:
|
|
// $ gsd-tools audit-open --pick nonexistent_field
|
|
// ### Milestone Close: Open Artifact Audit
|
|
//
|
|
// All artifact types clear. Safe to proceed.
|
|
//
|
|
// ---
|
|
// exit 0
|
|
test('nonJsonOutputDoesNotDumpWholeDocument', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const result = runGsdTools('audit-open --pick nonexistent_field', tmpDir);
|
|
assert.strictEqual(result.success, false);
|
|
assert.strictEqual(result.output, '');
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
|
|
// P13 (test matrix): `--raw --pick <known field>` withdraws its
|
|
// coincidental "success" — measured today: exit 0, stdout `hello-world`,
|
|
// via the same non-JSON dump `--raw` produces.
|
|
test('rawPlusPickIsRejectedNotCoincidentallyRight', () => {
|
|
const result = runGsdTools(['generate-slug', 'Hello World', '--raw', '--pick', 'slug']);
|
|
assert.strictEqual(result.success, false);
|
|
});
|
|
|
|
// P14 (test matrix): the confidently-wrong case — measured today: exit 0,
|
|
// stdout `hello-world` (the SLUG field's value, not the bogus field asked
|
|
// for), via the same non-JSON dump.
|
|
test('rawPlusPickBogusDoesNotEmitAnotherFieldsValue', () => {
|
|
const result = runGsdTools(['generate-slug', 'Hello World', '--raw', '--pick', 'bogus']);
|
|
assert.strictEqual(result.success, false);
|
|
assert.ok(
|
|
!result.output.includes('hello-world'),
|
|
`must not leak another field's value; got: ${JSON.stringify(result.output)}`,
|
|
);
|
|
});
|
|
|
|
test('errors when --pick has no value', () => {
|
|
const result = runGsdTools('generate-slug "test" --pick');
|
|
assert.strictEqual(result.success, false);
|
|
assert.match(result.error, /Missing value for --pick/);
|
|
});
|
|
|
|
test('errors when --pick value starts with --', () => {
|
|
const result = runGsdTools(['generate-slug', 'test', '--pick', '--raw']);
|
|
assert.strictEqual(result.success, false);
|
|
assert.match(result.error, /Missing value for --pick/);
|
|
});
|
|
|
|
test('does not collide with frontmatter --field flag', () => {
|
|
// frontmatter subcommand uses --field internally; --pick should not interfere
|
|
const result = runGsdTools('generate-slug "test-value" --pick slug');
|
|
assert.strictEqual(result.success, true);
|
|
assert.strictEqual(result.output, 'test-value');
|
|
});
|
|
|
|
test('works with current-timestamp command', () => {
|
|
const result = runGsdTools('current-timestamp --pick timestamp');
|
|
assert.strictEqual(result.success, true);
|
|
assert.ok(result.output.length > 0, 'timestamp should not be empty');
|
|
assert.match(result.output, /^\d{4}-\d{2}-\d{2}T/);
|
|
});
|
|
|
|
// B7 (design 40-design.md; test matrix P7): a dotted path that resolves
|
|
// partway then dies. `count` is a number on `phases.list --type summaries`
|
|
// (a real, always-present field); walking `.missing` off it is not a
|
|
// plain object, so the path dies partway through — the same failure class
|
|
// as B6 (field absent outright), not a crash.
|
|
test('absentDottedPathExitsNonZero', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const result = runCli(
|
|
['query', 'phases.list', '--type', 'summaries', '--pick', 'count.missing'],
|
|
{ cwd: tmpDir },
|
|
);
|
|
assert.notStrictEqual(result.status, 0);
|
|
assert.strictEqual(result.reason, 'pick_field_absent');
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
|
|
// B9 (test matrix P11): bracket syntax applied to a non-array field.
|
|
test('bracketOnNonArrayExitsNonZero', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const result = runCli(
|
|
['query', 'phases.list', '--type', 'summaries', '--pick', 'count[0]'],
|
|
{ cwd: tmpDir },
|
|
);
|
|
assert.notStrictEqual(result.status, 0);
|
|
assert.strictEqual(result.reason, 'pick_field_absent');
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
|
|
// B10 (test matrix): the existing boundary test above only exercises
|
|
// non-negative indices — negative-index normalization
|
|
// (`arr.length + index`) is a separate branch in extractField and was
|
|
// otherwise untested. N=3 seeded phase directories.
|
|
test('negativeArrayIndexInRangeResolves', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
seedPhase(tmpDir, '01-alpha');
|
|
seedPhase(tmpDir, '02-beta');
|
|
seedPhase(tmpDir, '03-gamma');
|
|
|
|
const last = runGsdTools('query phases.list --pick directories[-1]', tmpDir);
|
|
assert.strictEqual(last.success, true);
|
|
assert.strictEqual(last.output, '03-gamma');
|
|
|
|
const first = runGsdTools('query phases.list --pick directories[-3]', tmpDir);
|
|
assert.strictEqual(first.success, true);
|
|
assert.strictEqual(first.output, '01-alpha');
|
|
|
|
const outOfRange = runGsdTools('query phases.list --pick directories[-4]', tmpDir);
|
|
assert.strictEqual(outOfRange.success, false, 'index == -(N+1) is out of range and must exit non-zero');
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
|
|
// B14 (test matrix P15): the JSON root itself is not an object. VERIFIED
|
|
// on this tree: config-get with --default on a missing key emits the bare
|
|
// JSON string "fallback" (not an object), so --pick must fail rather than
|
|
// walk a string as if it had named fields.
|
|
test('nonObjectJsonRootExitsNonZero', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const plain = runGsdTools('config-get nonexistent.key --default fallback', tmpDir);
|
|
assert.strictEqual(plain.success, true);
|
|
assert.strictEqual(plain.output, '"fallback"');
|
|
|
|
const result = runCli(
|
|
['config-get', 'nonexistent.key', '--default', 'fallback', '--pick', 'value'],
|
|
{ cwd: tmpDir },
|
|
);
|
|
assert.notStrictEqual(result.status, 0);
|
|
assert.strictEqual(result.reason, 'pick_field_absent');
|
|
assert.match(result.message, /JSON string, not an object/);
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
|
|
// B17 (test matrix P19/P20, negative space N8): io.cjs's output() writes
|
|
// `@file:<path>` instead of inline JSON once the serialized payload
|
|
// exceeds 50000 characters, and `--pick` MUST resolve that redirection
|
|
// BEFORE parsing — otherwise every large result becomes a false
|
|
// pick_output_not_json. The fixture below seeds exactly PHASE_COUNT real
|
|
// phase directories (skipping phase number 999, which phase.cjs's sentinel
|
|
// predicate — SENTINEL_RANGES [0,999] — excludes from the list regardless
|
|
// of padding width, confirmed empirically; skipping it keeps `count`
|
|
// exactly PHASE_COUNT so the assertions below are deterministic) with
|
|
// padded names long enough that the serialized JSON provably exceeds the
|
|
// threshold. The spill is MEASURED, not assumed: the plain (non --pick)
|
|
// path transparently resolves @file: back to inline JSON (#1891), so its
|
|
// stdout length IS the real serialized payload size.
|
|
test('largeAtFilePayloadStillResolves + largeAtFilePayloadAbsentFieldIsAbsentNotNonJson', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const PHASE_COUNT = 1200;
|
|
let made = 0;
|
|
for (let i = 1; made < PHASE_COUNT; i++) {
|
|
if (i === 999) continue; // sentinel phase id — excluded from the list, would skew `count`
|
|
seedPhase(tmpDir, `${String(i).padStart(5, '0')}-phase-name-padding-to-make-this-longer`);
|
|
made++;
|
|
}
|
|
|
|
const plain = runGsdTools('query phases.list', tmpDir);
|
|
assert.strictEqual(plain.success, true);
|
|
assert.ok(
|
|
plain.output.length > 50000,
|
|
`fixture must exceed the 50000-char @file: spill threshold; measured ${plain.output.length}`,
|
|
);
|
|
|
|
// Present field ("largeAtFilePayloadStillResolves"): resolves at exit
|
|
// 0 through the @file: payload.
|
|
const present = runGsdTools('query phases.list --pick count', tmpDir);
|
|
assert.strictEqual(present.success, true);
|
|
assert.strictEqual(present.output, String(PHASE_COUNT));
|
|
|
|
// Absent field ("largeAtFilePayloadAbsentFieldIsAbsentNotNonJson"):
|
|
// must be pick_field_absent, NOT pick_output_not_json — proving the
|
|
// @file: resolution ran before the JSON.parse/absence check.
|
|
const absent = runCli(
|
|
['query', 'phases.list', '--pick', 'nonexistent_field'],
|
|
{ cwd: tmpDir },
|
|
);
|
|
assert.notStrictEqual(absent.status, 0);
|
|
assert.strictEqual(absent.reason, 'pick_field_absent');
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
});
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// formatDiagnosticToken (io.cjs) — untrusted --pick token escaping.
|
|
//
|
|
// Adversarial review finding (isolated review, verified live): `--pick`'s
|
|
// field value reaches the "field not found" / "output was not JSON"
|
|
// diagnostics verbatim. Repro on this tree BEFORE the fix:
|
|
// $ node gsd-tools.cjs generate-slug x --pick $'a\nError: forged'
|
|
// Error: --pick a
|
|
// Error: forged: field not found; available top-level keys: slug
|
|
// Spawns the real CLI (the vulnerability is about the literal bytes on
|
|
// stderr) and asserts on the RAW stderr string — a trimmed assertion would
|
|
// hide a leading/trailing forged blank line.
|
|
// ---------------------------------------------------------------------------
|
|
|
|
describe('formatDiagnosticToken escapes untrusted --pick field values', () => {
|
|
const HOSTILE_TOKENS = [
|
|
['embedded newline forging a second Error: line', 'a\nError: forged'],
|
|
['embedded double quote', 'a"bogus'],
|
|
['embedded C0 control character', 'a\x07bogus'],
|
|
];
|
|
|
|
for (const [label, token] of HOSTILE_TOKENS) {
|
|
test(`--pick field-not-found diagnostic stays single-line for a token with ${label}`, () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const result = runCli(['generate-slug', 'x', '--pick', token], { cwd: tmpDir, jsonErrors: false });
|
|
assert.notStrictEqual(result.status, 0);
|
|
const rawStderr = result.stderr;
|
|
const nonEmptyLines = rawStderr.split('\n').filter((l) => l.length > 0);
|
|
assert.strictEqual(
|
|
nonEmptyLines.length, 1,
|
|
`expected exactly one non-empty stderr line, got raw stderr: ${JSON.stringify(rawStderr)}`,
|
|
);
|
|
assert.match(nonEmptyLines[0], /^Error: /);
|
|
const errorPrefixedLineCount = rawStderr.split('\n').filter((l) => l.startsWith('Error:')).length;
|
|
assert.strictEqual(errorPrefixedLineCount, 1, 'the hostile token must not forge a second "Error:" line');
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
}
|
|
});
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// formatKeyForDiagnosticList (gsd-tools.cjs) — untrusted frontmatter KEY
|
|
// escaping, as distinct from the untrusted --pick TOKEN escaping above.
|
|
//
|
|
// `frontmatter get <file>` (no --field) reads an arbitrary user-authored
|
|
// markdown file and, on a subsequent --pick miss, echoes that document's own
|
|
// top-level frontmatter keys straight into the "field not found; available
|
|
// top-level keys: ..." diagnostic. A key is therefore untrusted input from a
|
|
// user document in exactly the way a --pick argv token is untrusted input
|
|
// from the shell — formatKeyForDiagnosticList (gsd-tools.cjs) exists to
|
|
// neutralize it the same way formatDiagnosticToken neutralizes the token.
|
|
//
|
|
// Reachable + verified live on this tree, e.g. for a frontmatter key
|
|
// containing a real embedded newline:
|
|
// $ printf '---\n"weird\\nkey": v\nplain: y\n---\n\nbody\n' > f.md
|
|
// $ gsd-tools frontmatter get f.md --pick absent_field
|
|
// Error: --pick "absent_field": field not found; available top-level keys: weird\nkey, plain
|
|
// The `\n` in that stderr is the two-character ESCAPED sequence backslash-n,
|
|
// not a real newline — the raw/untrimmed stderr assertions below pin that.
|
|
// ---------------------------------------------------------------------------
|
|
|
|
describe('formatKeyForDiagnosticList escapes untrusted frontmatter keys', () => {
|
|
test('hostileFrontmatterKeyCannotForgeASecondErrorLine', () => {
|
|
const HOSTILE_KEY_FIXTURES = [
|
|
['embedded newline', '---\n"weird\\nkey": v\nplain: y\n---\n\nbody\n', 'weird\\nkey'],
|
|
['embedded double quote', '---\n"weird\\"quotekey": v\nplain: y\n---\n\nbody\n', 'weird\\"quotekey'],
|
|
['embedded C0 control character', '---\n"weird\\u0007ctrlkey": v\nplain: y\n---\n\nbody\n', 'weird\\u0007ctrlkey'],
|
|
];
|
|
|
|
for (const [label, frontmatterSource, expectedEscapedKey] of HOSTILE_KEY_FIXTURES) {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const f = path.join(tmpDir, 'weird.md');
|
|
fs.writeFileSync(f, frontmatterSource);
|
|
|
|
const result = runCli(
|
|
['frontmatter', 'get', f, '--pick', 'absent_field'],
|
|
{ cwd: tmpDir, jsonErrors: false },
|
|
);
|
|
assert.notStrictEqual(result.status, 0, `[${label}] must exit non-zero`);
|
|
|
|
const rawStderr = result.stderr;
|
|
// Split on the raw, UNTRIMMED stderr and drop exactly one trailing
|
|
// empty element (the newline error() always terminates its message
|
|
// with) — a hostile key that forged a second line would leave MORE
|
|
// than one element after that single drop.
|
|
const lines = rawStderr.split('\n');
|
|
assert.strictEqual(
|
|
lines[lines.length - 1], '',
|
|
`[${label}] expected a single trailing empty element from the terminating newline, got raw stderr: ${JSON.stringify(rawStderr)}`,
|
|
);
|
|
const linesWithoutTrailingEmpty = lines.slice(0, -1);
|
|
assert.strictEqual(
|
|
linesWithoutTrailingEmpty.length, 1,
|
|
`[${label}] expected exactly one line after dropping the trailing empty element, got raw stderr: ${JSON.stringify(rawStderr)}`,
|
|
);
|
|
const errorPrefixedLineCount = linesWithoutTrailingEmpty.filter((l) => l.startsWith('Error:')).length;
|
|
assert.strictEqual(errorPrefixedLineCount, 1, `[${label}] the hostile key must not forge a second "Error:" line`);
|
|
|
|
// The diagnostic must stay USEFUL, not merely safe: a "fix" that
|
|
// dropped the offending key entirely would also pass the one-line
|
|
// assertions above, so pin that the escaped key is still present.
|
|
assert.ok(
|
|
linesWithoutTrailingEmpty[0].includes(`available top-level keys: ${expectedEscapedKey}, plain`),
|
|
`[${label}] expected the escaped key to still name the offending key, got: ${JSON.stringify(linesWithoutTrailingEmpty[0])}`,
|
|
);
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
}
|
|
});
|
|
|
|
// Negative-space control: an ORDINARY frontmatter document (no hostile
|
|
// bytes in any key) must still produce a plain, readable, unquoted key
|
|
// list — proving the escape does not turn every normal diagnostic into
|
|
// JSON-quoted noise.
|
|
test('ordinaryFrontmatterKeysStayPlainAndUnquotedInDiagnostic', () => {
|
|
const tmpDir = createTempProject();
|
|
try {
|
|
const f = path.join(tmpDir, 'plain.md');
|
|
fs.writeFileSync(f, '---\nalpha: v\nbeta: y\n---\n\nbody\n');
|
|
|
|
const result = runCli(
|
|
['frontmatter', 'get', f, '--pick', 'absent_field'],
|
|
{ cwd: tmpDir, jsonErrors: false },
|
|
);
|
|
assert.notStrictEqual(result.status, 0);
|
|
assert.match(result.stderr, /available top-level keys: alpha, beta\n$/);
|
|
// Only the KEY LIST must stay unquoted — `--pick "absent_field"` earlier
|
|
// in the same message is legitimately JSON-quoted by formatDiagnosticToken
|
|
// (a separate escape, for the untrusted argv token, not the frontmatter
|
|
// key), so scope the "no JSON-quoting noise" assertion to the key-list
|
|
// segment rather than the whole stderr string.
|
|
const keyListSegment = result.stderr.slice(result.stderr.indexOf('available top-level keys:'));
|
|
assert.ok(!keyListSegment.includes('"'), `ordinary keys must not be JSON-quoted, got: ${JSON.stringify(keyListSegment)}`);
|
|
} finally {
|
|
cleanup(tmpDir);
|
|
}
|
|
});
|
|
});
|