Files
msd-core/tests/capability-validator-task-content-resolver.test.cjs
Tom Boucher dd4f179672 feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam

Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute
plus a new optional `taskContentResolver` capability-manifest field let a
capability resolve a task's action/verify/acceptance-criteria/read_first/done
content from an external issue tracker instead of PLAN.md's inline body.

- src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId`
- src/task-content-resolution.cts: new leaf module — split/find/build/resolve,
  with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed
  resolution, never a silent fallback to possibly-stale inline text
- src/task-command-router.cts: new `task resolve-content --plan --task-id --raw`
  CLI verb wiring the module into a real process exit code
- gsd-core/bin/lib/capability-validator.cjs: validates the new
  `taskContentResolver` manifest field (feature-role only, cross-capability
  trackerPrefix uniqueness)
- gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md,
  docs/reference/capability-manifest.md: wire the seam into the per-task loop
  and document it as a new `execute:task` point outside the existing
  contribution/step/gate vocabulary (unconditional in autonomous mode)

Closes #3970

* fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard

Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646
Phase 1) found three defects:

1. execute-plan.md's task-content-resolution bullet fired on any
   tracker-id-bearing task with no check that it wasn't type="checkpoint:*",
   contradicting ADR-3646 Decision 1 (a checkpoint task must never enter
   resolve-content). plan-document.cts already parses trackerId: null
   unconditionally for checkpoint tasks; only the workflow prose needed the
   fix, so the bullet now explicitly excludes checkpoint tasks.

2. task-content-resolution.cts's parseResolverDeclaration accepted any
   non-empty trackerPrefix with no grammar check, while capability-
   validator.cjs's KEBAB_RE enforces kebab-case at install time — a
   Generative Fix Divergence gap. Added the same grammar (as a literal
   regex, documented as intentionally not shared across the .cts/.cjs build
   boundary) plus a parity test asserting the two surfaces agree across a
   valid/invalid trackerPrefix table.

3. task-command-router.cts's routeResolveContent path-traversal guard on
   --plan had zero test coverage. Added a test exercising a
   ../../../etc/passwit-shaped path and asserting the USAGE rejection names
   the offending path.

* fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs

Two findings caught by an isolated security-review pass on the task
content resolution seam:

- ResolverFailedError/ResolverMalformedOutputError embedded raw,
  unsanitized subprocess stderr/stdout (attacker/model-influenced via
  the tracker-id argv token) into .message. A hostile or buggy resolver
  could smuggle a newline plus a forged "Error: " line, or terminal
  escape sequences, into a diagnostic io.cjs's error() writes verbatim
  to stderr. Fixed at the constructor (task-content-resolution.cts) via
  io.cjs's existing formatDiagnosticToken(), so every caller of
  resolveTaskContent gets a safe .message by construction.

- capability-validator.cjs's validateTaskContentResolverFields had no
  upper bound on taskContentResolver.invoke.timeoutMs, letting a
  manifest declare an effectively unbounded value and defeat the
  "bounded subprocess" design intent. Added a 120000ms ceiling specific
  to this field, without touching the shared isPositiveIntegerMs()
  helper (still used unbounded by the reviewer lane's timeoutFloorMs
  and probe timeoutMs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion

gsd-test (remote dockerized matrix) came back red with 5 failures on this
PR; all five are real defects, fixed here.

- tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's
  execute-plan.md entry pointed at line 415, which ffc190df4's
  checkpoint-exclusion caveat (added near line 221) shifted down by one
  line. The actual "validated downstream by gsd-tools uat
  classify-coverage" descriptive mention now sits at line 416. Updated
  the allowlist entry's line number to match.

- tests/task-command-router-resolve-content.test.cjs: the path-traversal
  test asserted the outside-project-scope diagnostic against the thrown
  ExitError's own .message. io.cts's error() (ADR-3889) writes its
  human-readable message to fd 2 via writeAllSync and then throws a bare
  `new ExitError(1)` with no message argument — by design, so the
  exception carries no duplicate text and the thrown ExitError's message
  defaults to "process exit 1" (cli-exit.cts's ExitError constructor).
  Root cause was the test, not the source: task-command-router.cjs's
  outside-project-scope rejection already calls error() correctly and the
  diagnostic text is genuinely emitted, just on fd 2, not on the
  exception. Fixed the test to capture fd-2 writes (mirroring
  tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this
  same file's own captureStdout for fd 1) and assert against the captured
  stderr text instead of err.message. This was masked locally because a
  manual `node -e` sanity check that only inspects the caught exception's
  .message cannot see what the real node:test run actually failed on.

Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3970): backfill changeset PR number (pr:0 -> pr:4000)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 13:17:04 -04:00

317 lines
11 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// docs-guard-exempt: no docs/ file reads in this test.
'use strict';
process.env.GSD_TEST_MODE = '1';
/**
* capability-validator-task-content-resolver.test.cjs — behavioral tests for
* the OPTIONAL `taskContentResolver` body on `role: "feature"` capability
* manifests (ADR-3646, #3970).
*
* Implements test-matrix rows 20–24 of
* `.gsd/phase/feat-3970-task-content-resolution-seam/50-test-matrix.md`.
* See `docs/adr/3646-per-task-content-resolution-seam.md` Decision 3 for the
* shape this validates.
*/
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const {
validateCapability,
validateTaskContentResolver,
validateCrossCapability,
} = require('../gsd-core/bin/lib/capability-validator.cjs');
// ─── Fixture builders ──────────────────────────────────────────────────────
// House convention (tests/capability-manifest-version.test.cjs): builder
// functions return a VALID fixture, which each test then mutates. Every call
// returns a FRESH object — no shared mutable state, no execution-order
// dependence.
function validResolver() {
return {
trackerPrefix: 'beads',
invoke: {
binary: 'bd',
args: ['show', '{{id}}', '--json'],
timeoutMs: 10000,
},
};
}
function featureCap(overrides) {
return {
id: 'demo',
role: 'feature',
version: '1.2.3',
title: 'Demo',
description: 'A demo capability.',
tier: 'standard',
requires: [],
engines: { gsd: '>=1.6.0' },
runtimeCompat: { supported: ['*'], unsupported: [] },
skills: [],
agents: [],
hooks: [],
config: {},
steps: [],
contributions: [],
gates: [],
...overrides,
};
}
function runtimeCap(overrides) {
return {
id: 'demo-rt',
role: 'runtime',
version: '1.2.3',
title: 'Demo RT',
description: 'A demo runtime.',
tier: 'standard',
requires: [],
engines: { gsd: '>=1.6.0' },
runtime: {
configHome: { kind: 'dot-home', name: '.demo', env: [] },
localConfigDir: '.demo',
configFormat: 'settings-json',
artifactLayout: { global: [], local: [] },
commandStyle: 'slash-hyphen',
hooksSurface: 'settings-json',
sandboxTier: 'none',
supportTier: 2,
installSurface: 'settings-json',
writesSharedSettings: false,
permissionWriter: null,
extendedHookEvents: [],
hostIntegration: {
embeddingMode: 'imperative',
commandSurface: 'slash-file',
dispatch: { namedDispatch: true, nested: true, maxDepth: -1, background: true, subagentToolkit: 'full', backgroundDispatch: false },
modelMode: 'passive',
hookBus: 'host',
stateIO: 'filesystem',
effortSurface: 'none',
isolation: 'process',
},
},
...overrides,
};
}
function reviewerCap(overrides) {
return {
id: 'demo-reviewer',
role: 'reviewer',
version: '1.2.3',
title: 'Demo Reviewer',
description: 'A demo reviewer lane.',
tier: 'standard',
requires: [],
engines: { gsd: '>=1.6.0' },
reviewer: {
slug: 'demo-reviewer',
flags: ['--demo-reviewer'],
transport: 'spawn',
probe: { kind: 'command-exists', binary: 'demo-reviewer' },
invoke: {
binary: 'demo-reviewer',
args: [],
promptChannel: 'stdin',
outputChannel: 'stdout',
modelArg: null,
effortChannel: 'none',
},
timeoutFloorMs: 5000,
emptyOutput: 'stub-with-stderr',
reviewsSection: 'Demo Reviewer',
evidenceClass: 'source-grounded',
requiresBinaries: [],
promptBudgetKey: null,
handler: null,
},
...overrides,
};
}
// ─── Row 20: valid taskContentResolver on a feature manifest ───────────────
describe('row 20 — valid taskContentResolver body', () => {
test('valid taskContentResolver body passes', () => {
const cap = featureCap({ taskContentResolver: validResolver() });
const errs = validateCapability(cap, cap.id);
assert.deepEqual(errs, [], `expected no errors, got: ${JSON.stringify(errs)}`);
});
test('omitting taskContentResolver entirely on an otherwise-valid feature manifest yields zero errors', () => {
const cap = featureCap();
const errs = validateCapability(cap, cap.id);
assert.deepEqual(errs, [], `expected no errors, got: ${JSON.stringify(errs)}`);
});
});
// ─── Row 21: feature-only field ────────────────────────────────────────────
describe('row 21 — taskContentResolver on non-feature role is rejected', () => {
test('role:runtime declaring taskContentResolver is rejected', () => {
const cap = runtimeCap({ taskContentResolver: validResolver() });
const errs = validateCapability(cap, cap.id);
assert.ok(
errs.some((e) => e.includes('taskContentResolver') && e.includes('feature-only')),
`expected a feature-only rejection, got: ${JSON.stringify(errs)}`,
);
});
test('role:reviewer declaring taskContentResolver is rejected', () => {
const cap = reviewerCap({ taskContentResolver: validResolver() });
const errs = validateCapability(cap, cap.id);
assert.ok(
errs.some((e) => e.includes('taskContentResolver') && e.includes('feature-only')),
`expected a feature-only rejection, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── Row 22: malformed trackerPrefix grammar ───────────────────────────────
describe('row 22 — malformed trackerPrefix is rejected', () => {
test('capital-cased trackerPrefix violates KEBAB_RE', () => {
const cap = featureCap({
taskContentResolver: { ...validResolver(), trackerPrefix: 'Beads' },
});
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('trackerPrefix') && e.includes('kebab-case')),
`expected a grammar error naming trackerPrefix, got: ${JSON.stringify(errs)}`,
);
});
test('empty-string trackerPrefix is rejected', () => {
const cap = featureCap({
taskContentResolver: { ...validResolver(), trackerPrefix: '' },
});
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('trackerPrefix')),
`expected a trackerPrefix error, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── Row 23: invoke.timeoutMs boundary ─────────────────────────────────────
describe('row 23 — invoke.timeoutMs must be a positive integer', () => {
for (const bad of [0, -1, 1.5, undefined]) {
test(`timeoutMs ${JSON.stringify(bad)} is rejected`, () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, timeoutMs: bad };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('timeoutMs')),
`expected a timeoutMs error for ${JSON.stringify(bad)}, got: ${JSON.stringify(errs)}`,
);
});
}
test('a legitimate positive integer timeoutMs (10000) is accepted', () => {
const cap = featureCap({ taskContentResolver: validResolver() });
const errs = validateTaskContentResolver(cap);
assert.ok(
!errs.some((e) => e.includes('timeoutMs')),
`expected no timeoutMs error, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── invoke.timeoutMs upper ceiling (security review finding, #3970) ───────
// A manifest declaring an unbounded-in-practice timeoutMs (e.g.
// Number.MAX_SAFE_INTEGER) would let resolve-content hang near-indefinitely
// on a stuck/malicious resolver, defeating the "bounded subprocess" design
// intent. Boundary-inclusive per CLAUDE.md's limit-1/limit/limit+1 rule.
describe('invoke.timeoutMs upper ceiling (120000ms)', () => {
test('timeoutMs 120000 (exactly at the ceiling) is accepted', () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, timeoutMs: 120000 };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
!errs.some((e) => e.includes('timeoutMs')),
`expected no timeoutMs error at the boundary, got: ${JSON.stringify(errs)}`,
);
});
test('timeoutMs 120001 (one past the ceiling) is rejected', () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, timeoutMs: 120001 };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('timeoutMs') && e.includes('120000')),
`expected a ceiling timeoutMs error, got: ${JSON.stringify(errs)}`,
);
});
test('an unbounded-in-practice timeoutMs (Number.MAX_SAFE_INTEGER) is rejected', () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, timeoutMs: Number.MAX_SAFE_INTEGER };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('timeoutMs')),
`expected a timeoutMs error, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── invoke.args must carry the {{id}} placeholder ─────────────────────────
describe('invoke.args without {{id}} placeholder', () => {
test('args missing the {{id}} placeholder is rejected', () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, args: ['show', 'GSD-42', '--json'] };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('{{id}}')),
`expected a placeholder error, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── Row 24: cross-capability trackerPrefix uniqueness ─────────────────────
describe('row 24 — duplicate trackerPrefix across capabilities fails cross-capability validation', () => {
test('two manifests declaring the same trackerPrefix collide', () => {
const capA = featureCap({ id: 'resolver-a', taskContentResolver: validResolver() });
const capB = featureCap({ id: 'resolver-b', taskContentResolver: validResolver() });
const capMap = new Map([
[capA.id, capA],
[capB.id, capB],
]);
const errs = validateCrossCapability(capMap, new Set());
assert.ok(
errs.some((e) => e.includes('trackerPrefix') && e.includes('resolver-a') && e.includes('resolver-b')),
`expected a collision error naming both ids, got: ${JSON.stringify(errs)}`,
);
});
test('two manifests declaring different trackerPrefix values do not collide', () => {
const capA = featureCap({ id: 'resolver-a', taskContentResolver: validResolver() });
const capB = featureCap({
id: 'resolver-b',
taskContentResolver: { ...validResolver(), trackerPrefix: 'linear' },
});
const capMap = new Map([
[capA.id, capA],
[capB.id, capB],
]);
const errs = validateCrossCapability(capMap, new Set());
assert.ok(
!errs.some((e) => e.includes('trackerPrefix')),
`expected no trackerPrefix collision, got: ${JSON.stringify(errs)}`,
);
});
});