Files
msd-core/docs/how-to/develop-a-task-content-resolver-capability.md
Tom Boucher dd4f179672 feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam

Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute
plus a new optional `taskContentResolver` capability-manifest field let a
capability resolve a task's action/verify/acceptance-criteria/read_first/done
content from an external issue tracker instead of PLAN.md's inline body.

- src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId`
- src/task-content-resolution.cts: new leaf module — split/find/build/resolve,
  with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed
  resolution, never a silent fallback to possibly-stale inline text
- src/task-command-router.cts: new `task resolve-content --plan --task-id --raw`
  CLI verb wiring the module into a real process exit code
- gsd-core/bin/lib/capability-validator.cjs: validates the new
  `taskContentResolver` manifest field (feature-role only, cross-capability
  trackerPrefix uniqueness)
- gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md,
  docs/reference/capability-manifest.md: wire the seam into the per-task loop
  and document it as a new `execute:task` point outside the existing
  contribution/step/gate vocabulary (unconditional in autonomous mode)

Closes #3970

* fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard

Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646
Phase 1) found three defects:

1. execute-plan.md's task-content-resolution bullet fired on any
   tracker-id-bearing task with no check that it wasn't type="checkpoint:*",
   contradicting ADR-3646 Decision 1 (a checkpoint task must never enter
   resolve-content). plan-document.cts already parses trackerId: null
   unconditionally for checkpoint tasks; only the workflow prose needed the
   fix, so the bullet now explicitly excludes checkpoint tasks.

2. task-content-resolution.cts's parseResolverDeclaration accepted any
   non-empty trackerPrefix with no grammar check, while capability-
   validator.cjs's KEBAB_RE enforces kebab-case at install time — a
   Generative Fix Divergence gap. Added the same grammar (as a literal
   regex, documented as intentionally not shared across the .cts/.cjs build
   boundary) plus a parity test asserting the two surfaces agree across a
   valid/invalid trackerPrefix table.

3. task-command-router.cts's routeResolveContent path-traversal guard on
   --plan had zero test coverage. Added a test exercising a
   ../../../etc/passwit-shaped path and asserting the USAGE rejection names
   the offending path.

* fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs

Two findings caught by an isolated security-review pass on the task
content resolution seam:

- ResolverFailedError/ResolverMalformedOutputError embedded raw,
  unsanitized subprocess stderr/stdout (attacker/model-influenced via
  the tracker-id argv token) into .message. A hostile or buggy resolver
  could smuggle a newline plus a forged "Error: " line, or terminal
  escape sequences, into a diagnostic io.cjs's error() writes verbatim
  to stderr. Fixed at the constructor (task-content-resolution.cts) via
  io.cjs's existing formatDiagnosticToken(), so every caller of
  resolveTaskContent gets a safe .message by construction.

- capability-validator.cjs's validateTaskContentResolverFields had no
  upper bound on taskContentResolver.invoke.timeoutMs, letting a
  manifest declare an effectively unbounded value and defeat the
  "bounded subprocess" design intent. Added a 120000ms ceiling specific
  to this field, without touching the shared isPositiveIntegerMs()
  helper (still used unbounded by the reviewer lane's timeoutFloorMs
  and probe timeoutMs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion

gsd-test (remote dockerized matrix) came back red with 5 failures on this
PR; all five are real defects, fixed here.

- tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's
  execute-plan.md entry pointed at line 415, which ffc190df4's
  checkpoint-exclusion caveat (added near line 221) shifted down by one
  line. The actual "validated downstream by gsd-tools uat
  classify-coverage" descriptive mention now sits at line 416. Updated
  the allowlist entry's line number to match.

- tests/task-command-router-resolve-content.test.cjs: the path-traversal
  test asserted the outside-project-scope diagnostic against the thrown
  ExitError's own .message. io.cts's error() (ADR-3889) writes its
  human-readable message to fd 2 via writeAllSync and then throws a bare
  `new ExitError(1)` with no message argument — by design, so the
  exception carries no duplicate text and the thrown ExitError's message
  defaults to "process exit 1" (cli-exit.cts's ExitError constructor).
  Root cause was the test, not the source: task-command-router.cjs's
  outside-project-scope rejection already calls error() correctly and the
  diagnostic text is genuinely emitted, just on fd 2, not on the
  exception. Fixed the test to capture fd-2 writes (mirroring
  tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this
  same file's own captureStdout for fd 1) and assert against the captured
  stderr text instead of err.message. This was masked locally because a
  manual `node -e` sanity check that only inspects the caught exception's
  .message cannot see what the real node:test run actually failed on.

Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3970): backfill changeset PR number (pr:0 -> pr:4000)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 13:17:04 -04:00

7.0 KiB
Raw Blame History

Develop a task-content resolver capability

Goal: Declare a taskContentResolver in a capability manifest so execute-plan.md resolves a task's <action>/<verify>/<acceptance_criteria>/<read_first>/<done> content from your external issue tracker instead of reading it inline from PLAN.md.

Prerequisites: You already have a capability (capability.json), or you are creating one — see Develop a Capability for GSD 1.5+ first if this is your first one. GSD 1.28 or later (ADR-3646).


Why this exists

A project that wants an external tracker — beads, Linear, Jira, GitHub Issues — to own task content, not just task status, has no seam for that today: execute-plan.md reads every task's instructions directly out of the PLAN.md task block. taskContentResolver adds one, with a hard-halt guarantee: if your tracker is declared as the source of truth and it fails to resolve, execution stops rather than silently falling back to stale PLAN.md text. That guarantee is the entire point of the feature — a silent fallback would require the tracker and PLAN.md to stay in sync forever, defeating the reason to move content out of PLAN.md in the first place.


Declare the resolver

Add a taskContentResolver block to your capability's manifest body (role: "feature" only):

{
  "id": "beads",
  "role": "feature",
  "version": "1.0.0",
  "title": "Beads issue tracker",
  "description": "Resolves task content from the bd issue tracker.",
  "tier": "standard",
  "requires": [],
  "runtimeCompat": { "supported": ["*"], "unsupported": [] },
  "skills": [],
  "agents": [],
  "hooks": [],
  "config": {},
  "steps": [],
  "contributions": [],
  "gates": [],

  "taskContentResolver": {
    "trackerPrefix": "beads",
    "invoke": {
      "binary": "bd",
      "args": ["show", "{{id}}", "--json"],
      "timeoutMs": 10000
    }
  }
}

Two fields decide whether the seam works at all:

  • trackerPrefix must match the prefix of a task's tracker-id attribute — everything before the first :. Given <task tracker-id="beads:GSD-42">, the prefix is beads and the id passed to your resolver is GSD-42. If a tracker's own ids contain colons, that is fine: only the first colon splits prefix from id, so beads:team:GSD-42 resolves to id team:GSD-42.
  • invoke.args must contain the {{id}} placeholder at least once — GSD substitutes it with the task's id before spawning your binary. A declaration whose args never carries the placeholder fails validation at install time, because the id could never reach your resolver.

invoke.timeoutMs is required. An unbounded resolver subprocess is this repo's named Unbounded Subprocesses defect class — declare a bound that matches how long your tracker's lookup actually takes, plus margin.

trackerPrefix must be unique across the merged first-party ∪ overlay capability set; a collision — two installed capabilities both claiming "beads" — is a build-time validation error, not a runtime ambiguity.


What your resolver must output

execute-plan.md invokes your invoke.binary/invoke.args and expects a single JSON object on stdout when the lookup succeeds:

Field Type Required Maps to
description string Yes The task's <action>
verify string No The task's <verify>
acceptance_criteria string[] No The task's <acceptance_criteria>
read_first string[] No The task's <read_first>
done string No The task's <done>

An absent or empty-string description is treated as "nothing resolved" — execute-plan.md falls back to the task's inline PLAN.md content, the one legitimate pre-migration boundary case (for tasks authored before your tracker migration). This is the only silent fallback path; every other failure is a hard halt.

Exit code and stderr matter. Exit 0 with valid JSON on stdout is the only success path. Anything else — a non-zero exit, a timeout past invoke.timeoutMs, or stdout that fails to parse as JSON — is treated as a resolution failure. Write a clear one-line reason to stderr; it is surfaced verbatim to the person watching execution. Stderr on a successful (exit 0) run is not an error — write informational logs there if your CLI already does; only the exit code and the JSON parse outcome decide success.


What happens on failure

A resolver that is declared, invoked, and fails — non-zero exit, timeout, or malformed JSON — makes gsd_run task resolve-content itself exit non-zero. execute-plan.md treats that as a hard halt: it stops before doing any work on the task, surfaces the tracker-id, the tracker prefix, and your resolver's stderr, and never proceeds to read the task's inline PLAN.md content as a substitute. See ADR-3646 for why this is the load-bearing safety property of the whole feature, and loop-hook-dispatch.md for how this call site differs from the twelve prose-dispatched loop extension points.


Worked example: a bd/beads-shaped resolver

Suppose bd show GSD-42 --json already returns:

{
  "id": "GSD-42",
  "title": "Add rate-limit check to login",
  "status": "open",
  "description": "Add a rate-limit check to processLogin using the existing RateLimiter.",
  "acceptance": [
    "Login attempts beyond the configured limit return 429",
    "Existing successful-login tests still pass"
  ],
  "notes": "See src/util/rate.ts for the existing limiter."
}

Your resolver is a thin adapter, not bd itself — it maps bd's field names onto the shape execute-plan.md expects and exits non-zero on anything bd itself reports as a failure:

#!/usr/bin/env bash
set -euo pipefail

id="$1"
raw=$(bd show "$id" --json)

node -e '
  const raw = JSON.parse(process.argv[1]);
  const out = {
    description: raw.description || "",
    acceptance_criteria: raw.acceptance || [],
  };
  process.stdout.write(JSON.stringify(out));
' "$raw"

Declare it as the invoke.binary/invoke.args pair (or point invoke.binary at bd directly if its own --json output already matches the expected field names — no adapter needed in that case). Either way, invoke.args must carry {{id}} so GSD can substitute the task's tracker id before spawning it.