* feat(#3970): per-task external-tracker content-resolution seam Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute plus a new optional `taskContentResolver` capability-manifest field let a capability resolve a task's action/verify/acceptance-criteria/read_first/done content from an external issue tracker instead of PLAN.md's inline body. - src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId` - src/task-content-resolution.cts: new leaf module — split/find/build/resolve, with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed resolution, never a silent fallback to possibly-stale inline text - src/task-command-router.cts: new `task resolve-content --plan --task-id --raw` CLI verb wiring the module into a real process exit code - gsd-core/bin/lib/capability-validator.cjs: validates the new `taskContentResolver` manifest field (feature-role only, cross-capability trackerPrefix uniqueness) - gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md, docs/reference/capability-manifest.md: wire the seam into the per-task loop and document it as a new `execute:task` point outside the existing contribution/step/gate vocabulary (unconditional in autonomous mode) Closes #3970 * fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646 Phase 1) found three defects: 1. execute-plan.md's task-content-resolution bullet fired on any tracker-id-bearing task with no check that it wasn't type="checkpoint:*", contradicting ADR-3646 Decision 1 (a checkpoint task must never enter resolve-content). plan-document.cts already parses trackerId: null unconditionally for checkpoint tasks; only the workflow prose needed the fix, so the bullet now explicitly excludes checkpoint tasks. 2. task-content-resolution.cts's parseResolverDeclaration accepted any non-empty trackerPrefix with no grammar check, while capability- validator.cjs's KEBAB_RE enforces kebab-case at install time — a Generative Fix Divergence gap. Added the same grammar (as a literal regex, documented as intentionally not shared across the .cts/.cjs build boundary) plus a parity test asserting the two surfaces agree across a valid/invalid trackerPrefix table. 3. task-command-router.cts's routeResolveContent path-traversal guard on --plan had zero test coverage. Added a test exercising a ../../../etc/passwit-shaped path and asserting the USAGE rejection names the offending path. * fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs Two findings caught by an isolated security-review pass on the task content resolution seam: - ResolverFailedError/ResolverMalformedOutputError embedded raw, unsanitized subprocess stderr/stdout (attacker/model-influenced via the tracker-id argv token) into .message. A hostile or buggy resolver could smuggle a newline plus a forged "Error: " line, or terminal escape sequences, into a diagnostic io.cjs's error() writes verbatim to stderr. Fixed at the constructor (task-content-resolution.cts) via io.cjs's existing formatDiagnosticToken(), so every caller of resolveTaskContent gets a safe .message by construction. - capability-validator.cjs's validateTaskContentResolverFields had no upper bound on taskContentResolver.invoke.timeoutMs, letting a manifest declare an effectively unbounded value and defeat the "bounded subprocess" design intent. Added a 120000ms ceiling specific to this field, without touching the shared isPositiveIntegerMs() helper (still used unbounded by the reviewer lane's timeoutFloorMs and probe timeoutMs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion gsd-test (remote dockerized matrix) came back red with 5 failures on this PR; all five are real defects, fixed here. - tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's execute-plan.md entry pointed at line 415, which ffc190df4's checkpoint-exclusion caveat (added near line 221) shifted down by one line. The actual "validated downstream by gsd-tools uat classify-coverage" descriptive mention now sits at line 416. Updated the allowlist entry's line number to match. - tests/task-command-router-resolve-content.test.cjs: the path-traversal test asserted the outside-project-scope diagnostic against the thrown ExitError's own .message. io.cts's error() (ADR-3889) writes its human-readable message to fd 2 via writeAllSync and then throws a bare `new ExitError(1)` with no message argument — by design, so the exception carries no duplicate text and the thrown ExitError's message defaults to "process exit 1" (cli-exit.cts's ExitError constructor). Root cause was the test, not the source: task-command-router.cjs's outside-project-scope rejection already calls error() correctly and the diagnostic text is genuinely emitted, just on fd 2, not on the exception. Fixed the test to capture fd-2 writes (mirroring tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this same file's own captureStdout for fd 1) and assert against the captured stderr text instead of err.message. This was masked locally because a manual `node -e` sanity check that only inspects the caught exception's .message cannot see what the real node:test run actually failed on. Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3970): backfill changeset PR number (pr:0 -> pr:4000) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
317 lines
11 KiB
JavaScript
317 lines
11 KiB
JavaScript
// docs-guard-exempt: no docs/ file reads in this test.
|
||
'use strict';
|
||
process.env.GSD_TEST_MODE = '1';
|
||
|
||
/**
|
||
* capability-validator-task-content-resolver.test.cjs — behavioral tests for
|
||
* the OPTIONAL `taskContentResolver` body on `role: "feature"` capability
|
||
* manifests (ADR-3646, #3970).
|
||
*
|
||
* Implements test-matrix rows 20–24 of
|
||
* `.gsd/phase/feat-3970-task-content-resolution-seam/50-test-matrix.md`.
|
||
* See `docs/adr/3646-per-task-content-resolution-seam.md` Decision 3 for the
|
||
* shape this validates.
|
||
*/
|
||
|
||
const { test, describe } = require('node:test');
|
||
const assert = require('node:assert/strict');
|
||
|
||
const {
|
||
validateCapability,
|
||
validateTaskContentResolver,
|
||
validateCrossCapability,
|
||
} = require('../gsd-core/bin/lib/capability-validator.cjs');
|
||
|
||
// ─── Fixture builders ──────────────────────────────────────────────────────
|
||
// House convention (tests/capability-manifest-version.test.cjs): builder
|
||
// functions return a VALID fixture, which each test then mutates. Every call
|
||
// returns a FRESH object — no shared mutable state, no execution-order
|
||
// dependence.
|
||
|
||
function validResolver() {
|
||
return {
|
||
trackerPrefix: 'beads',
|
||
invoke: {
|
||
binary: 'bd',
|
||
args: ['show', '{{id}}', '--json'],
|
||
timeoutMs: 10000,
|
||
},
|
||
};
|
||
}
|
||
|
||
function featureCap(overrides) {
|
||
return {
|
||
id: 'demo',
|
||
role: 'feature',
|
||
version: '1.2.3',
|
||
title: 'Demo',
|
||
description: 'A demo capability.',
|
||
tier: 'standard',
|
||
requires: [],
|
||
engines: { gsd: '>=1.6.0' },
|
||
runtimeCompat: { supported: ['*'], unsupported: [] },
|
||
skills: [],
|
||
agents: [],
|
||
hooks: [],
|
||
config: {},
|
||
steps: [],
|
||
contributions: [],
|
||
gates: [],
|
||
...overrides,
|
||
};
|
||
}
|
||
|
||
function runtimeCap(overrides) {
|
||
return {
|
||
id: 'demo-rt',
|
||
role: 'runtime',
|
||
version: '1.2.3',
|
||
title: 'Demo RT',
|
||
description: 'A demo runtime.',
|
||
tier: 'standard',
|
||
requires: [],
|
||
engines: { gsd: '>=1.6.0' },
|
||
runtime: {
|
||
configHome: { kind: 'dot-home', name: '.demo', env: [] },
|
||
localConfigDir: '.demo',
|
||
configFormat: 'settings-json',
|
||
artifactLayout: { global: [], local: [] },
|
||
commandStyle: 'slash-hyphen',
|
||
hooksSurface: 'settings-json',
|
||
sandboxTier: 'none',
|
||
supportTier: 2,
|
||
installSurface: 'settings-json',
|
||
writesSharedSettings: false,
|
||
permissionWriter: null,
|
||
extendedHookEvents: [],
|
||
hostIntegration: {
|
||
embeddingMode: 'imperative',
|
||
commandSurface: 'slash-file',
|
||
dispatch: { namedDispatch: true, nested: true, maxDepth: -1, background: true, subagentToolkit: 'full', backgroundDispatch: false },
|
||
modelMode: 'passive',
|
||
hookBus: 'host',
|
||
stateIO: 'filesystem',
|
||
effortSurface: 'none',
|
||
isolation: 'process',
|
||
},
|
||
},
|
||
...overrides,
|
||
};
|
||
}
|
||
|
||
function reviewerCap(overrides) {
|
||
return {
|
||
id: 'demo-reviewer',
|
||
role: 'reviewer',
|
||
version: '1.2.3',
|
||
title: 'Demo Reviewer',
|
||
description: 'A demo reviewer lane.',
|
||
tier: 'standard',
|
||
requires: [],
|
||
engines: { gsd: '>=1.6.0' },
|
||
reviewer: {
|
||
slug: 'demo-reviewer',
|
||
flags: ['--demo-reviewer'],
|
||
transport: 'spawn',
|
||
probe: { kind: 'command-exists', binary: 'demo-reviewer' },
|
||
invoke: {
|
||
binary: 'demo-reviewer',
|
||
args: [],
|
||
promptChannel: 'stdin',
|
||
outputChannel: 'stdout',
|
||
modelArg: null,
|
||
effortChannel: 'none',
|
||
},
|
||
timeoutFloorMs: 5000,
|
||
emptyOutput: 'stub-with-stderr',
|
||
reviewsSection: 'Demo Reviewer',
|
||
evidenceClass: 'source-grounded',
|
||
requiresBinaries: [],
|
||
promptBudgetKey: null,
|
||
handler: null,
|
||
},
|
||
...overrides,
|
||
};
|
||
}
|
||
|
||
// ─── Row 20: valid taskContentResolver on a feature manifest ───────────────
|
||
|
||
describe('row 20 — valid taskContentResolver body', () => {
|
||
test('valid taskContentResolver body passes', () => {
|
||
const cap = featureCap({ taskContentResolver: validResolver() });
|
||
const errs = validateCapability(cap, cap.id);
|
||
assert.deepEqual(errs, [], `expected no errors, got: ${JSON.stringify(errs)}`);
|
||
});
|
||
|
||
test('omitting taskContentResolver entirely on an otherwise-valid feature manifest yields zero errors', () => {
|
||
const cap = featureCap();
|
||
const errs = validateCapability(cap, cap.id);
|
||
assert.deepEqual(errs, [], `expected no errors, got: ${JSON.stringify(errs)}`);
|
||
});
|
||
});
|
||
|
||
// ─── Row 21: feature-only field ────────────────────────────────────────────
|
||
|
||
describe('row 21 — taskContentResolver on non-feature role is rejected', () => {
|
||
test('role:runtime declaring taskContentResolver is rejected', () => {
|
||
const cap = runtimeCap({ taskContentResolver: validResolver() });
|
||
const errs = validateCapability(cap, cap.id);
|
||
assert.ok(
|
||
errs.some((e) => e.includes('taskContentResolver') && e.includes('feature-only')),
|
||
`expected a feature-only rejection, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
|
||
test('role:reviewer declaring taskContentResolver is rejected', () => {
|
||
const cap = reviewerCap({ taskContentResolver: validResolver() });
|
||
const errs = validateCapability(cap, cap.id);
|
||
assert.ok(
|
||
errs.some((e) => e.includes('taskContentResolver') && e.includes('feature-only')),
|
||
`expected a feature-only rejection, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
});
|
||
|
||
// ─── Row 22: malformed trackerPrefix grammar ───────────────────────────────
|
||
|
||
describe('row 22 — malformed trackerPrefix is rejected', () => {
|
||
test('capital-cased trackerPrefix violates KEBAB_RE', () => {
|
||
const cap = featureCap({
|
||
taskContentResolver: { ...validResolver(), trackerPrefix: 'Beads' },
|
||
});
|
||
const errs = validateTaskContentResolver(cap);
|
||
assert.ok(
|
||
errs.some((e) => e.includes('trackerPrefix') && e.includes('kebab-case')),
|
||
`expected a grammar error naming trackerPrefix, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
|
||
test('empty-string trackerPrefix is rejected', () => {
|
||
const cap = featureCap({
|
||
taskContentResolver: { ...validResolver(), trackerPrefix: '' },
|
||
});
|
||
const errs = validateTaskContentResolver(cap);
|
||
assert.ok(
|
||
errs.some((e) => e.includes('trackerPrefix')),
|
||
`expected a trackerPrefix error, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
});
|
||
|
||
// ─── Row 23: invoke.timeoutMs boundary ─────────────────────────────────────
|
||
|
||
describe('row 23 — invoke.timeoutMs must be a positive integer', () => {
|
||
for (const bad of [0, -1, 1.5, undefined]) {
|
||
test(`timeoutMs ${JSON.stringify(bad)} is rejected`, () => {
|
||
const resolver = validResolver();
|
||
resolver.invoke = { ...resolver.invoke, timeoutMs: bad };
|
||
const cap = featureCap({ taskContentResolver: resolver });
|
||
const errs = validateTaskContentResolver(cap);
|
||
assert.ok(
|
||
errs.some((e) => e.includes('timeoutMs')),
|
||
`expected a timeoutMs error for ${JSON.stringify(bad)}, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
}
|
||
|
||
test('a legitimate positive integer timeoutMs (10000) is accepted', () => {
|
||
const cap = featureCap({ taskContentResolver: validResolver() });
|
||
const errs = validateTaskContentResolver(cap);
|
||
assert.ok(
|
||
!errs.some((e) => e.includes('timeoutMs')),
|
||
`expected no timeoutMs error, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
});
|
||
|
||
// ─── invoke.timeoutMs upper ceiling (security review finding, #3970) ───────
|
||
// A manifest declaring an unbounded-in-practice timeoutMs (e.g.
|
||
// Number.MAX_SAFE_INTEGER) would let resolve-content hang near-indefinitely
|
||
// on a stuck/malicious resolver, defeating the "bounded subprocess" design
|
||
// intent. Boundary-inclusive per CLAUDE.md's limit-1/limit/limit+1 rule.
|
||
|
||
describe('invoke.timeoutMs upper ceiling (120000ms)', () => {
|
||
test('timeoutMs 120000 (exactly at the ceiling) is accepted', () => {
|
||
const resolver = validResolver();
|
||
resolver.invoke = { ...resolver.invoke, timeoutMs: 120000 };
|
||
const cap = featureCap({ taskContentResolver: resolver });
|
||
const errs = validateTaskContentResolver(cap);
|
||
assert.ok(
|
||
!errs.some((e) => e.includes('timeoutMs')),
|
||
`expected no timeoutMs error at the boundary, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
|
||
test('timeoutMs 120001 (one past the ceiling) is rejected', () => {
|
||
const resolver = validResolver();
|
||
resolver.invoke = { ...resolver.invoke, timeoutMs: 120001 };
|
||
const cap = featureCap({ taskContentResolver: resolver });
|
||
const errs = validateTaskContentResolver(cap);
|
||
assert.ok(
|
||
errs.some((e) => e.includes('timeoutMs') && e.includes('120000')),
|
||
`expected a ceiling timeoutMs error, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
|
||
test('an unbounded-in-practice timeoutMs (Number.MAX_SAFE_INTEGER) is rejected', () => {
|
||
const resolver = validResolver();
|
||
resolver.invoke = { ...resolver.invoke, timeoutMs: Number.MAX_SAFE_INTEGER };
|
||
const cap = featureCap({ taskContentResolver: resolver });
|
||
const errs = validateTaskContentResolver(cap);
|
||
assert.ok(
|
||
errs.some((e) => e.includes('timeoutMs')),
|
||
`expected a timeoutMs error, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
});
|
||
|
||
// ─── invoke.args must carry the {{id}} placeholder ─────────────────────────
|
||
|
||
describe('invoke.args without {{id}} placeholder', () => {
|
||
test('args missing the {{id}} placeholder is rejected', () => {
|
||
const resolver = validResolver();
|
||
resolver.invoke = { ...resolver.invoke, args: ['show', 'GSD-42', '--json'] };
|
||
const cap = featureCap({ taskContentResolver: resolver });
|
||
const errs = validateTaskContentResolver(cap);
|
||
assert.ok(
|
||
errs.some((e) => e.includes('{{id}}')),
|
||
`expected a placeholder error, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
});
|
||
|
||
// ─── Row 24: cross-capability trackerPrefix uniqueness ─────────────────────
|
||
|
||
describe('row 24 — duplicate trackerPrefix across capabilities fails cross-capability validation', () => {
|
||
test('two manifests declaring the same trackerPrefix collide', () => {
|
||
const capA = featureCap({ id: 'resolver-a', taskContentResolver: validResolver() });
|
||
const capB = featureCap({ id: 'resolver-b', taskContentResolver: validResolver() });
|
||
const capMap = new Map([
|
||
[capA.id, capA],
|
||
[capB.id, capB],
|
||
]);
|
||
const errs = validateCrossCapability(capMap, new Set());
|
||
assert.ok(
|
||
errs.some((e) => e.includes('trackerPrefix') && e.includes('resolver-a') && e.includes('resolver-b')),
|
||
`expected a collision error naming both ids, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
|
||
test('two manifests declaring different trackerPrefix values do not collide', () => {
|
||
const capA = featureCap({ id: 'resolver-a', taskContentResolver: validResolver() });
|
||
const capB = featureCap({
|
||
id: 'resolver-b',
|
||
taskContentResolver: { ...validResolver(), trackerPrefix: 'linear' },
|
||
});
|
||
const capMap = new Map([
|
||
[capA.id, capA],
|
||
[capB.id, capB],
|
||
]);
|
||
const errs = validateCrossCapability(capMap, new Set());
|
||
assert.ok(
|
||
!errs.some((e) => e.includes('trackerPrefix')),
|
||
`expected no trackerPrefix collision, got: ${JSON.stringify(errs)}`,
|
||
);
|
||
});
|
||
});
|