* feat(#3970): per-task external-tracker content-resolution seam Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute plus a new optional `taskContentResolver` capability-manifest field let a capability resolve a task's action/verify/acceptance-criteria/read_first/done content from an external issue tracker instead of PLAN.md's inline body. - src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId` - src/task-content-resolution.cts: new leaf module — split/find/build/resolve, with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed resolution, never a silent fallback to possibly-stale inline text - src/task-command-router.cts: new `task resolve-content --plan --task-id --raw` CLI verb wiring the module into a real process exit code - gsd-core/bin/lib/capability-validator.cjs: validates the new `taskContentResolver` manifest field (feature-role only, cross-capability trackerPrefix uniqueness) - gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md, docs/reference/capability-manifest.md: wire the seam into the per-task loop and document it as a new `execute:task` point outside the existing contribution/step/gate vocabulary (unconditional in autonomous mode) Closes #3970 * fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646 Phase 1) found three defects: 1. execute-plan.md's task-content-resolution bullet fired on any tracker-id-bearing task with no check that it wasn't type="checkpoint:*", contradicting ADR-3646 Decision 1 (a checkpoint task must never enter resolve-content). plan-document.cts already parses trackerId: null unconditionally for checkpoint tasks; only the workflow prose needed the fix, so the bullet now explicitly excludes checkpoint tasks. 2. task-content-resolution.cts's parseResolverDeclaration accepted any non-empty trackerPrefix with no grammar check, while capability- validator.cjs's KEBAB_RE enforces kebab-case at install time — a Generative Fix Divergence gap. Added the same grammar (as a literal regex, documented as intentionally not shared across the .cts/.cjs build boundary) plus a parity test asserting the two surfaces agree across a valid/invalid trackerPrefix table. 3. task-command-router.cts's routeResolveContent path-traversal guard on --plan had zero test coverage. Added a test exercising a ../../../etc/passwit-shaped path and asserting the USAGE rejection names the offending path. * fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs Two findings caught by an isolated security-review pass on the task content resolution seam: - ResolverFailedError/ResolverMalformedOutputError embedded raw, unsanitized subprocess stderr/stdout (attacker/model-influenced via the tracker-id argv token) into .message. A hostile or buggy resolver could smuggle a newline plus a forged "Error: " line, or terminal escape sequences, into a diagnostic io.cjs's error() writes verbatim to stderr. Fixed at the constructor (task-content-resolution.cts) via io.cjs's existing formatDiagnosticToken(), so every caller of resolveTaskContent gets a safe .message by construction. - capability-validator.cjs's validateTaskContentResolverFields had no upper bound on taskContentResolver.invoke.timeoutMs, letting a manifest declare an effectively unbounded value and defeat the "bounded subprocess" design intent. Added a 120000ms ceiling specific to this field, without touching the shared isPositiveIntegerMs() helper (still used unbounded by the reviewer lane's timeoutFloorMs and probe timeoutMs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion gsd-test (remote dockerized matrix) came back red with 5 failures on this PR; all five are real defects, fixed here. - tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's execute-plan.md entry pointed at line 415, which ffc190df4's checkpoint-exclusion caveat (added near line 221) shifted down by one line. The actual "validated downstream by gsd-tools uat classify-coverage" descriptive mention now sits at line 416. Updated the allowlist entry's line number to match. - tests/task-command-router-resolve-content.test.cjs: the path-traversal test asserted the outside-project-scope diagnostic against the thrown ExitError's own .message. io.cts's error() (ADR-3889) writes its human-readable message to fd 2 via writeAllSync and then throws a bare `new ExitError(1)` with no message argument — by design, so the exception carries no duplicate text and the thrown ExitError's message defaults to "process exit 1" (cli-exit.cts's ExitError constructor). Root cause was the test, not the source: task-command-router.cjs's outside-project-scope rejection already calls error() correctly and the diagnostic text is genuinely emitted, just on fd 2, not on the exception. Fixed the test to capture fd-2 writes (mirroring tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this same file's own captureStdout for fd 1) and assert against the captured stderr text instead of err.message. This was masked locally because a manual `node -e` sanity check that only inspects the caught exception's .message cannot see what the real node:test run actually failed on. Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3970): backfill changeset PR number (pr:0 -> pr:4000) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
85 lines
3.3 KiB
JavaScript
85 lines
3.3 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* Parity test: `capability-validator.cjs`'s install-time `KEBAB_RE` grammar
|
|
* check on `taskContentResolver.trackerPrefix` (`validateTaskContentResolver`)
|
|
* MUST agree with `task-content-resolution.cts`'s resolve-time re-validation
|
|
* inside `parseResolverDeclaration` (exercised here via `findResolver`) on
|
|
* every `trackerPrefix` value.
|
|
*
|
|
* This is the "Generative Fix Divergence" guard CLAUDE.md requires whenever
|
|
* two surfaces share a rule with no single source of truth: the grammar is
|
|
* duplicated as a literal regex in both files (see `task-content-
|
|
* resolution.cts`'s `TRACKER_PREFIX_RE` docstring for why it is not a shared
|
|
* import), so this test is what actually keeps them from drifting apart. If a
|
|
* future change to either regex loosens or tightens it without mirroring the
|
|
* change in the other file, this test fails.
|
|
*/
|
|
|
|
const { test, describe } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
|
|
const { validateTaskContentResolver } = require('../gsd-core/bin/lib/capability-validator.cjs');
|
|
const { findResolver } = require('../gsd-core/bin/lib/task-content-resolution.cjs');
|
|
|
|
function validInvoke() {
|
|
return { binary: 'bd', args: ['show', '{{id}}', '--json'], timeoutMs: 10000 };
|
|
}
|
|
|
|
function featureCapWithResolver(trackerPrefix) {
|
|
return {
|
|
id: 'demo',
|
|
role: 'feature',
|
|
taskContentResolver: { trackerPrefix, invoke: validInvoke() },
|
|
};
|
|
}
|
|
|
|
/** True when the validator accepts this trackerPrefix (no trackerPrefix-naming error). */
|
|
function validatorAccepts(trackerPrefix) {
|
|
const cap = featureCapWithResolver(trackerPrefix);
|
|
const errs = validateTaskContentResolver(cap);
|
|
return !errs.some((e) => e.includes('trackerPrefix'));
|
|
}
|
|
|
|
/** True when the resolve-time seam accepts this trackerPrefix (finds a match, not `null`). */
|
|
function resolverAccepts(trackerPrefix) {
|
|
const capabilities = [featureCapWithResolver(trackerPrefix)];
|
|
const result = findResolver(trackerPrefix, capabilities);
|
|
return result !== null && result !== 'ambiguous';
|
|
}
|
|
|
|
const TABLE = [
|
|
{ trackerPrefix: 'beads', valid: true },
|
|
{ trackerPrefix: 'my-tracker', valid: true },
|
|
{ trackerPrefix: 'Beads', valid: false },
|
|
{ trackerPrefix: 'has_underscore', valid: false },
|
|
{ trackerPrefix: 'UPPER', valid: false },
|
|
{ trackerPrefix: '', valid: false },
|
|
{ trackerPrefix: '1leading-digit', valid: false },
|
|
];
|
|
|
|
describe('trackerPrefix grammar parity — capability-validator.cjs vs task-content-resolution.cts', () => {
|
|
for (const { trackerPrefix, valid } of TABLE) {
|
|
test(`'${trackerPrefix}' — validator and resolver agree (expected valid: ${valid})`, () => {
|
|
const validatorResult = validatorAccepts(trackerPrefix);
|
|
const resolverResult = resolverAccepts(trackerPrefix);
|
|
assert.strictEqual(
|
|
validatorResult,
|
|
valid,
|
|
`validator disagreed with expected table value for '${trackerPrefix}'`,
|
|
);
|
|
assert.strictEqual(
|
|
resolverResult,
|
|
valid,
|
|
`resolver disagreed with expected table value for '${trackerPrefix}'`,
|
|
);
|
|
assert.strictEqual(
|
|
validatorResult,
|
|
resolverResult,
|
|
`PARITY BREAK: validator and resolver disagree for trackerPrefix '${trackerPrefix}' ` +
|
|
`(validator: ${validatorResult}, resolver: ${resolverResult})`,
|
|
);
|
|
});
|
|
}
|
|
});
|