Files
msd-core/tests/plan-document.test.cjs
Tom Boucher dd4f179672 feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam

Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute
plus a new optional `taskContentResolver` capability-manifest field let a
capability resolve a task's action/verify/acceptance-criteria/read_first/done
content from an external issue tracker instead of PLAN.md's inline body.

- src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId`
- src/task-content-resolution.cts: new leaf module — split/find/build/resolve,
  with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed
  resolution, never a silent fallback to possibly-stale inline text
- src/task-command-router.cts: new `task resolve-content --plan --task-id --raw`
  CLI verb wiring the module into a real process exit code
- gsd-core/bin/lib/capability-validator.cjs: validates the new
  `taskContentResolver` manifest field (feature-role only, cross-capability
  trackerPrefix uniqueness)
- gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md,
  docs/reference/capability-manifest.md: wire the seam into the per-task loop
  and document it as a new `execute:task` point outside the existing
  contribution/step/gate vocabulary (unconditional in autonomous mode)

Closes #3970

* fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard

Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646
Phase 1) found three defects:

1. execute-plan.md's task-content-resolution bullet fired on any
   tracker-id-bearing task with no check that it wasn't type="checkpoint:*",
   contradicting ADR-3646 Decision 1 (a checkpoint task must never enter
   resolve-content). plan-document.cts already parses trackerId: null
   unconditionally for checkpoint tasks; only the workflow prose needed the
   fix, so the bullet now explicitly excludes checkpoint tasks.

2. task-content-resolution.cts's parseResolverDeclaration accepted any
   non-empty trackerPrefix with no grammar check, while capability-
   validator.cjs's KEBAB_RE enforces kebab-case at install time — a
   Generative Fix Divergence gap. Added the same grammar (as a literal
   regex, documented as intentionally not shared across the .cts/.cjs build
   boundary) plus a parity test asserting the two surfaces agree across a
   valid/invalid trackerPrefix table.

3. task-command-router.cts's routeResolveContent path-traversal guard on
   --plan had zero test coverage. Added a test exercising a
   ../../../etc/passwit-shaped path and asserting the USAGE rejection names
   the offending path.

* fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs

Two findings caught by an isolated security-review pass on the task
content resolution seam:

- ResolverFailedError/ResolverMalformedOutputError embedded raw,
  unsanitized subprocess stderr/stdout (attacker/model-influenced via
  the tracker-id argv token) into .message. A hostile or buggy resolver
  could smuggle a newline plus a forged "Error: " line, or terminal
  escape sequences, into a diagnostic io.cjs's error() writes verbatim
  to stderr. Fixed at the constructor (task-content-resolution.cts) via
  io.cjs's existing formatDiagnosticToken(), so every caller of
  resolveTaskContent gets a safe .message by construction.

- capability-validator.cjs's validateTaskContentResolverFields had no
  upper bound on taskContentResolver.invoke.timeoutMs, letting a
  manifest declare an effectively unbounded value and defeat the
  "bounded subprocess" design intent. Added a 120000ms ceiling specific
  to this field, without touching the shared isPositiveIntegerMs()
  helper (still used unbounded by the reviewer lane's timeoutFloorMs
  and probe timeoutMs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion

gsd-test (remote dockerized matrix) came back red with 5 failures on this
PR; all five are real defects, fixed here.

- tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's
  execute-plan.md entry pointed at line 415, which ffc190df4's
  checkpoint-exclusion caveat (added near line 221) shifted down by one
  line. The actual "validated downstream by gsd-tools uat
  classify-coverage" descriptive mention now sits at line 416. Updated
  the allowlist entry's line number to match.

- tests/task-command-router-resolve-content.test.cjs: the path-traversal
  test asserted the outside-project-scope diagnostic against the thrown
  ExitError's own .message. io.cts's error() (ADR-3889) writes its
  human-readable message to fd 2 via writeAllSync and then throws a bare
  `new ExitError(1)` with no message argument — by design, so the
  exception carries no duplicate text and the thrown ExitError's message
  defaults to "process exit 1" (cli-exit.cts's ExitError constructor).
  Root cause was the test, not the source: task-command-router.cjs's
  outside-project-scope rejection already calls error() correctly and the
  diagnostic text is genuinely emitted, just on fd 2, not on the
  exception. Fixed the test to capture fd-2 writes (mirroring
  tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this
  same file's own captureStdout for fd 1) and assert against the captured
  stderr text instead of err.message. This was masked locally because a
  manual `node -e` sanity check that only inspects the caught exception's
  .message cannot see what the real node:test run actually failed on.

Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3970): backfill changeset PR number (pr:0 -> pr:4000)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 13:17:04 -04:00

109 lines
3.3 KiB
JavaScript

'use strict';
/**
* Unit tests for plan-document.cjs
*
* Module: gsd-core/bin/lib/plan-document.cjs
*
* Covers the `tracker-id` attribute (ADR-3646 Phase 1, #3970) added to the
* `<task>` element grammar, plus regression coverage proving the addition
* does not alter pre-existing task-parsing behaviour.
*
* Matrix rows referenced below are from
* .gsd/phase/feat-3970-task-content-resolution-seam/50-test-matrix.md
*/
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const { parsePlanDocument } = require('../gsd-core/bin/lib/plan-document.cjs');
describe('plan-document: tracker-id attribute', () => {
test('row 1 — no tracker-id attribute yields trackerId: null', () => {
const doc = parsePlanDocument(`
<task type="auto">
<name>Do a thing</name>
</task>
`);
assert.equal(doc.tasks.length, 1);
assert.equal(doc.tasks[0].trackerId, null);
});
test('row 2 — tracker-id is read verbatim, never split', () => {
const doc = parsePlanDocument(`
<task type="auto" tracker-id="beads:GSD-42">
<name>Do a thing</name>
</task>
`);
assert.equal(doc.tasks.length, 1);
assert.equal(doc.tasks[0].trackerId, 'beads:GSD-42');
});
test('row 3 — tracker-id="" (empty string) normalises to null', () => {
const doc = parsePlanDocument(`
<task type="auto" tracker-id="">
<name>Do a thing</name>
</task>
`);
assert.equal(doc.tasks.length, 1);
assert.equal(doc.tasks[0].trackerId, null);
});
test('row 4 — checkpoint tasks never read tracker-id, even when present', () => {
const doc = parsePlanDocument(`
<task type="checkpoint:decision" tracker-id="beads:GSD-99">
<decision>Ship it</decision>
</task>
`);
assert.equal(doc.tasks.length, 1);
assert.equal(doc.tasks[0].kind, 'checkpoint');
assert.equal(doc.tasks[0].trackerId, null);
});
});
describe('plan-document: regression — legacy behaviour unchanged', () => {
test('legacy `## Task N` markdown fallback still parses with trackerId: null', () => {
const doc = parsePlanDocument(`
## Task 1: Do a thing
Some body text.
## Task 2: Do another thing
`);
assert.equal(doc.tasks.length, 2);
for (const t of doc.tasks) {
assert.equal(t.kind, 'auto');
assert.equal(t.type, null);
assert.equal(t.trackerId, null);
assert.deepEqual(t.plannedFiles, []);
assert.deepEqual(t.acceptanceCriteria, []);
assert.equal(t.done, null);
}
assert.equal(doc.tasks[0].name, 'Task 1: Do a thing');
assert.equal(doc.tasks[1].name, 'Task 2: Do another thing');
});
test('ordinary task with name/files/acceptance_criteria still parses correctly alongside trackerId', () => {
const doc = parsePlanDocument(`
<task type="auto" tracker-id="beads:GSD-7">
<name>Implement the seam</name>
<files>src/a.cts, src/b.cts</files>
<acceptance_criteria>
- criterion one
- criterion two
</acceptance_criteria>
<done>Merged.</done>
</task>
`);
assert.equal(doc.tasks.length, 1);
const t = doc.tasks[0];
assert.equal(t.kind, 'auto');
assert.equal(t.type, 'auto');
assert.equal(t.name, 'Implement the seam');
assert.deepEqual(t.plannedFiles, ['src/a.cts', 'src/b.cts']);
assert.deepEqual(t.acceptanceCriteria, ['criterion one', 'criterion two']);
assert.equal(t.done, 'Merged.');
assert.equal(t.trackerId, 'beads:GSD-7');
});
});