* feat(#3970): per-task external-tracker content-resolution seam Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute plus a new optional `taskContentResolver` capability-manifest field let a capability resolve a task's action/verify/acceptance-criteria/read_first/done content from an external issue tracker instead of PLAN.md's inline body. - src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId` - src/task-content-resolution.cts: new leaf module — split/find/build/resolve, with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed resolution, never a silent fallback to possibly-stale inline text - src/task-command-router.cts: new `task resolve-content --plan --task-id --raw` CLI verb wiring the module into a real process exit code - gsd-core/bin/lib/capability-validator.cjs: validates the new `taskContentResolver` manifest field (feature-role only, cross-capability trackerPrefix uniqueness) - gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md, docs/reference/capability-manifest.md: wire the seam into the per-task loop and document it as a new `execute:task` point outside the existing contribution/step/gate vocabulary (unconditional in autonomous mode) Closes #3970 * fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646 Phase 1) found three defects: 1. execute-plan.md's task-content-resolution bullet fired on any tracker-id-bearing task with no check that it wasn't type="checkpoint:*", contradicting ADR-3646 Decision 1 (a checkpoint task must never enter resolve-content). plan-document.cts already parses trackerId: null unconditionally for checkpoint tasks; only the workflow prose needed the fix, so the bullet now explicitly excludes checkpoint tasks. 2. task-content-resolution.cts's parseResolverDeclaration accepted any non-empty trackerPrefix with no grammar check, while capability- validator.cjs's KEBAB_RE enforces kebab-case at install time — a Generative Fix Divergence gap. Added the same grammar (as a literal regex, documented as intentionally not shared across the .cts/.cjs build boundary) plus a parity test asserting the two surfaces agree across a valid/invalid trackerPrefix table. 3. task-command-router.cts's routeResolveContent path-traversal guard on --plan had zero test coverage. Added a test exercising a ../../../etc/passwit-shaped path and asserting the USAGE rejection names the offending path. * fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs Two findings caught by an isolated security-review pass on the task content resolution seam: - ResolverFailedError/ResolverMalformedOutputError embedded raw, unsanitized subprocess stderr/stdout (attacker/model-influenced via the tracker-id argv token) into .message. A hostile or buggy resolver could smuggle a newline plus a forged "Error: " line, or terminal escape sequences, into a diagnostic io.cjs's error() writes verbatim to stderr. Fixed at the constructor (task-content-resolution.cts) via io.cjs's existing formatDiagnosticToken(), so every caller of resolveTaskContent gets a safe .message by construction. - capability-validator.cjs's validateTaskContentResolverFields had no upper bound on taskContentResolver.invoke.timeoutMs, letting a manifest declare an effectively unbounded value and defeat the "bounded subprocess" design intent. Added a 120000ms ceiling specific to this field, without touching the shared isPositiveIntegerMs() helper (still used unbounded by the reviewer lane's timeoutFloorMs and probe timeoutMs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion gsd-test (remote dockerized matrix) came back red with 5 failures on this PR; all five are real defects, fixed here. - tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's execute-plan.md entry pointed at line 415, which ffc190df4's checkpoint-exclusion caveat (added near line 221) shifted down by one line. The actual "validated downstream by gsd-tools uat classify-coverage" descriptive mention now sits at line 416. Updated the allowlist entry's line number to match. - tests/task-command-router-resolve-content.test.cjs: the path-traversal test asserted the outside-project-scope diagnostic against the thrown ExitError's own .message. io.cts's error() (ADR-3889) writes its human-readable message to fd 2 via writeAllSync and then throws a bare `new ExitError(1)` with no message argument — by design, so the exception carries no duplicate text and the thrown ExitError's message defaults to "process exit 1" (cli-exit.cts's ExitError constructor). Root cause was the test, not the source: task-command-router.cjs's outside-project-scope rejection already calls error() correctly and the diagnostic text is genuinely emitted, just on fd 2, not on the exception. Fixed the test to capture fd-2 writes (mirroring tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this same file's own captureStdout for fd 1) and assert against the captured stderr text instead of err.message. This was masked locally because a manual `node -e` sanity check that only inspects the caught exception's .message cannot see what the real node:test run actually failed on. Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3970): backfill changeset PR number (pr:0 -> pr:4000) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
173 lines
7.0 KiB
Markdown
173 lines
7.0 KiB
Markdown
# Develop a task-content resolver capability
|
||
|
||
**Goal:** Declare a `taskContentResolver` in a capability manifest so `execute-plan.md` resolves
|
||
a task's `<action>`/`<verify>`/`<acceptance_criteria>`/`<read_first>`/`<done>` content from your
|
||
external issue tracker instead of reading it inline from `PLAN.md`.
|
||
|
||
**Prerequisites:** You already have a capability (`capability.json`), or you are creating one —
|
||
see [Develop a Capability for GSD 1.5+](develop-a-capability.md) first if this is your first one.
|
||
GSD 1.28 or later ([ADR-3646](../adr/3646-per-task-content-resolution-seam.md)).
|
||
|
||
---
|
||
|
||
## Why this exists
|
||
|
||
A project that wants an external tracker — beads, Linear, Jira, GitHub Issues — to own task
|
||
*content*, not just task *status*, has no seam for that today: `execute-plan.md` reads every
|
||
task's instructions directly out of the `PLAN.md` task block. `taskContentResolver` adds one, with
|
||
a hard-halt guarantee: if your tracker is declared as the source of truth and it fails to resolve,
|
||
execution stops rather than silently falling back to stale `PLAN.md` text. That guarantee is the
|
||
entire point of the feature — a silent fallback would require the tracker and `PLAN.md` to stay in
|
||
sync forever, defeating the reason to move content out of `PLAN.md` in the first place.
|
||
|
||
---
|
||
|
||
## Declare the resolver
|
||
|
||
Add a `taskContentResolver` block to your capability's manifest body (`role: "feature"` only):
|
||
|
||
```json
|
||
{
|
||
"id": "beads",
|
||
"role": "feature",
|
||
"version": "1.0.0",
|
||
"title": "Beads issue tracker",
|
||
"description": "Resolves task content from the bd issue tracker.",
|
||
"tier": "standard",
|
||
"requires": [],
|
||
"runtimeCompat": { "supported": ["*"], "unsupported": [] },
|
||
"skills": [],
|
||
"agents": [],
|
||
"hooks": [],
|
||
"config": {},
|
||
"steps": [],
|
||
"contributions": [],
|
||
"gates": [],
|
||
|
||
"taskContentResolver": {
|
||
"trackerPrefix": "beads",
|
||
"invoke": {
|
||
"binary": "bd",
|
||
"args": ["show", "{{id}}", "--json"],
|
||
"timeoutMs": 10000
|
||
}
|
||
}
|
||
}
|
||
```
|
||
|
||
Two fields decide whether the seam works at all:
|
||
|
||
- **`trackerPrefix`** must match the prefix of a task's `tracker-id` attribute — everything
|
||
before the **first** `:`. Given `<task tracker-id="beads:GSD-42">`, the prefix is `beads` and
|
||
the id passed to your resolver is `GSD-42`. If a tracker's own ids contain colons, that is fine:
|
||
only the first colon splits prefix from id, so `beads:team:GSD-42` resolves to id `team:GSD-42`.
|
||
- **`invoke.args`** must contain the `{{id}}` placeholder at least once — GSD substitutes it with
|
||
the task's id before spawning your binary. A declaration whose `args` never carries the
|
||
placeholder fails validation at install time, because the id could never reach your resolver.
|
||
|
||
`invoke.timeoutMs` is required. An unbounded resolver subprocess is this repo's named Unbounded
|
||
Subprocesses defect class — declare a bound that matches how long your tracker's lookup actually
|
||
takes, plus margin.
|
||
|
||
`trackerPrefix` must be unique across the merged first-party ∪ overlay capability set; a
|
||
collision — two installed capabilities both claiming `"beads"` — is a build-time validation
|
||
error, not a runtime ambiguity.
|
||
|
||
---
|
||
|
||
## What your resolver must output
|
||
|
||
`execute-plan.md` invokes your `invoke.binary`/`invoke.args` and expects a single JSON object on
|
||
stdout when the lookup succeeds:
|
||
|
||
| Field | Type | Required | Maps to |
|
||
|---|---|---|---|
|
||
| `description` | string | Yes | The task's `<action>` |
|
||
| `verify` | string | No | The task's `<verify>` |
|
||
| `acceptance_criteria` | string[] | No | The task's `<acceptance_criteria>` |
|
||
| `read_first` | string[] | No | The task's `<read_first>` |
|
||
| `done` | string | No | The task's `<done>` |
|
||
|
||
An absent or empty-string `description` is treated as "nothing resolved" — `execute-plan.md`
|
||
falls back to the task's inline `PLAN.md` content, the one legitimate pre-migration boundary case
|
||
(for tasks authored before your tracker migration). This is the *only* silent fallback path; every
|
||
other failure is a hard halt.
|
||
|
||
**Exit code and stderr matter.** Exit `0` with valid JSON on stdout is the only success path.
|
||
Anything else — a non-zero exit, a timeout past `invoke.timeoutMs`, or stdout that fails to parse
|
||
as JSON — is treated as a resolution failure. Write a clear one-line reason to stderr; it is
|
||
surfaced verbatim to the person watching execution. Stderr on a **successful** (exit 0) run is not
|
||
an error — write informational logs there if your CLI already does; only the exit code and the
|
||
JSON parse outcome decide success.
|
||
|
||
---
|
||
|
||
## What happens on failure
|
||
|
||
A resolver that is declared, invoked, and fails — non-zero exit, timeout, or malformed JSON —
|
||
makes `gsd_run task resolve-content` itself exit non-zero. `execute-plan.md` treats that as a
|
||
**hard halt**: it stops before doing any work on the task, surfaces the tracker-id, the tracker
|
||
prefix, and your resolver's stderr, and never proceeds to read the task's inline `PLAN.md` content
|
||
as a substitute. See [ADR-3646](../adr/3646-per-task-content-resolution-seam.md) for why this is
|
||
the load-bearing safety property of the whole feature, and
|
||
[`loop-hook-dispatch.md`](../../gsd-core/references/loop-hook-dispatch.md#the-executetask-point-a-different-shape)
|
||
for how this call site differs from the twelve prose-dispatched loop extension points.
|
||
|
||
---
|
||
|
||
## Worked example: a `bd`/beads-shaped resolver
|
||
|
||
Suppose `bd show GSD-42 --json` already returns:
|
||
|
||
```json
|
||
{
|
||
"id": "GSD-42",
|
||
"title": "Add rate-limit check to login",
|
||
"status": "open",
|
||
"description": "Add a rate-limit check to processLogin using the existing RateLimiter.",
|
||
"acceptance": [
|
||
"Login attempts beyond the configured limit return 429",
|
||
"Existing successful-login tests still pass"
|
||
],
|
||
"notes": "See src/util/rate.ts for the existing limiter."
|
||
}
|
||
```
|
||
|
||
Your resolver is a thin adapter, not `bd` itself — it maps `bd`'s field names onto the shape
|
||
`execute-plan.md` expects and exits non-zero on anything `bd` itself reports as a failure:
|
||
|
||
```bash
|
||
#!/usr/bin/env bash
|
||
set -euo pipefail
|
||
|
||
id="$1"
|
||
raw=$(bd show "$id" --json)
|
||
|
||
node -e '
|
||
const raw = JSON.parse(process.argv[1]);
|
||
const out = {
|
||
description: raw.description || "",
|
||
acceptance_criteria: raw.acceptance || [],
|
||
};
|
||
process.stdout.write(JSON.stringify(out));
|
||
' "$raw"
|
||
```
|
||
|
||
Declare it as the `invoke.binary`/`invoke.args` pair (or point `invoke.binary` at `bd` directly if
|
||
its own `--json` output already matches the expected field names — no adapter needed in that
|
||
case). Either way, `invoke.args` must carry `{{id}}` so GSD can substitute the task's tracker id
|
||
before spawning it.
|
||
|
||
---
|
||
|
||
## Related
|
||
|
||
- [ADR-3646](../adr/3646-per-task-content-resolution-seam.md) — the design decision and rejected
|
||
alternatives
|
||
- [Capability manifest → `taskContentResolver`](../reference/capability-manifest.md#taskcontentresolver) —
|
||
the full field table
|
||
- [`loop-hook-dispatch.md`](../../gsd-core/references/loop-hook-dispatch.md#the-executetask-point-a-different-shape) —
|
||
how `execute:task` differs from the twelve loop extension points
|
||
- [Develop a Capability for GSD 1.5+](develop-a-capability.md) — manifests, registry generation,
|
||
and federated config
|