feat(#3970): per-task external-tracker content-resolution seam (#4000)

* feat(#3970): per-task external-tracker content-resolution seam

Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute
plus a new optional `taskContentResolver` capability-manifest field let a
capability resolve a task's action/verify/acceptance-criteria/read_first/done
content from an external issue tracker instead of PLAN.md's inline body.

- src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId`
- src/task-content-resolution.cts: new leaf module — split/find/build/resolve,
  with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed
  resolution, never a silent fallback to possibly-stale inline text
- src/task-command-router.cts: new `task resolve-content --plan --task-id --raw`
  CLI verb wiring the module into a real process exit code
- gsd-core/bin/lib/capability-validator.cjs: validates the new
  `taskContentResolver` manifest field (feature-role only, cross-capability
  trackerPrefix uniqueness)
- gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md,
  docs/reference/capability-manifest.md: wire the seam into the per-task loop
  and document it as a new `execute:task` point outside the existing
  contribution/step/gate vocabulary (unconditional in autonomous mode)

Closes #3970

* fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard

Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646
Phase 1) found three defects:

1. execute-plan.md's task-content-resolution bullet fired on any
   tracker-id-bearing task with no check that it wasn't type="checkpoint:*",
   contradicting ADR-3646 Decision 1 (a checkpoint task must never enter
   resolve-content). plan-document.cts already parses trackerId: null
   unconditionally for checkpoint tasks; only the workflow prose needed the
   fix, so the bullet now explicitly excludes checkpoint tasks.

2. task-content-resolution.cts's parseResolverDeclaration accepted any
   non-empty trackerPrefix with no grammar check, while capability-
   validator.cjs's KEBAB_RE enforces kebab-case at install time — a
   Generative Fix Divergence gap. Added the same grammar (as a literal
   regex, documented as intentionally not shared across the .cts/.cjs build
   boundary) plus a parity test asserting the two surfaces agree across a
   valid/invalid trackerPrefix table.

3. task-command-router.cts's routeResolveContent path-traversal guard on
   --plan had zero test coverage. Added a test exercising a
   ../../../etc/passwit-shaped path and asserting the USAGE rejection names
   the offending path.

* fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs

Two findings caught by an isolated security-review pass on the task
content resolution seam:

- ResolverFailedError/ResolverMalformedOutputError embedded raw,
  unsanitized subprocess stderr/stdout (attacker/model-influenced via
  the tracker-id argv token) into .message. A hostile or buggy resolver
  could smuggle a newline plus a forged "Error: " line, or terminal
  escape sequences, into a diagnostic io.cjs's error() writes verbatim
  to stderr. Fixed at the constructor (task-content-resolution.cts) via
  io.cjs's existing formatDiagnosticToken(), so every caller of
  resolveTaskContent gets a safe .message by construction.

- capability-validator.cjs's validateTaskContentResolverFields had no
  upper bound on taskContentResolver.invoke.timeoutMs, letting a
  manifest declare an effectively unbounded value and defeat the
  "bounded subprocess" design intent. Added a 120000ms ceiling specific
  to this field, without touching the shared isPositiveIntegerMs()
  helper (still used unbounded by the reviewer lane's timeoutFloorMs
  and probe timeoutMs).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion

gsd-test (remote dockerized matrix) came back red with 5 failures on this
PR; all five are real defects, fixed here.

- tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's
  execute-plan.md entry pointed at line 415, which ffc190df4's
  checkpoint-exclusion caveat (added near line 221) shifted down by one
  line. The actual "validated downstream by gsd-tools uat
  classify-coverage" descriptive mention now sits at line 416. Updated
  the allowlist entry's line number to match.

- tests/task-command-router-resolve-content.test.cjs: the path-traversal
  test asserted the outside-project-scope diagnostic against the thrown
  ExitError's own .message. io.cts's error() (ADR-3889) writes its
  human-readable message to fd 2 via writeAllSync and then throws a bare
  `new ExitError(1)` with no message argument — by design, so the
  exception carries no duplicate text and the thrown ExitError's message
  defaults to "process exit 1" (cli-exit.cts's ExitError constructor).
  Root cause was the test, not the source: task-command-router.cjs's
  outside-project-scope rejection already calls error() correctly and the
  diagnostic text is genuinely emitted, just on fd 2, not on the
  exception. Fixed the test to capture fd-2 writes (mirroring
  tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this
  same file's own captureStdout for fd 1) and assert against the captured
  stderr text instead of err.message. This was masked locally because a
  manual `node -e` sanity check that only inspects the caught exception's
  .message cannot see what the real node:test run actually failed on.

Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#3970): backfill changeset PR number (pr:0 -> pr:4000)

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Tom Boucher
2026-08-28 13:17:04 -04:00
committed by GitHub
parent 12f9d1d9a0
commit dd4f179672
25 changed files with 2273 additions and 10 deletions

View File

@@ -0,0 +1,5 @@
---
type: Added
pr: 4000
---
**Capabilities can now source a task's content from an external issue tracker.** A capability that declares a `taskContentResolver` for a tracker prefix lets a plan's `tracker-id` attribute resolve the task's action, verify, acceptance criteria, and done text from that external tracker at execution time instead of PLAN.md, and any resolution failure — ambiguous match, non-zero exit, timeout, or malformed output — hard-halts rather than silently falling back. (#3970)

1
.gitignore vendored
View File

@@ -302,6 +302,7 @@ build/
/gsd-core/bin/lib/audit.cjs
/gsd-core/bin/lib/git-base-branch.cjs
/gsd-core/bin/lib/host-runtime-detection.cjs
/gsd-core/bin/lib/task-content-resolution.cjs
__pycache__/
*.pyc
.venv/

View File

@@ -136,7 +136,10 @@ Leaf module generalizing the proven `hooks/lib/git-cmd.js` token-walk (#3129) in
Module owning the parsed projection of `.planning/` that a diagnostic rule may read, per ADR-3180 §8.1 (Decision 8, Phase 10, #3308). `buildPlanningSnapshot(cwd) → PlanningSnapshot` is composed EXCLUSIVELY from the already-consolidated §7 owners — `getMilestoneInfo` (Roadmap Parser Module), `listMilestonePhaseDirs` (Phase Locator Module), `isPhaseComplete` (Verification Module), `scanPhasePlans` (Plan Scan Module), `stateFieldValue`/`stateCurrentPositionSlice` (STATE.md Document Module), `planningPaths` (Planning Workspace Module) — and introduces no new semantic derivation of its own. `PlanningSnapshot` exposes `milestone`/`phaseDirs`/`phases`/`currentPhaseLabel`, each a `{value, scope}` pair per the Planning Scope Module's frozen `SCOPE` enum; `phases` additionally carries a `PhaseSnapshot[]` (`dir`, `complete`, `verificationStatus`, `planCount`, `summaryCount`, `scope`). The one new piece of logic this module adds is `worstScope(...scopes) → Scope`, a pure severity-ordered combinator (`UNREADABLE` > `UNSCOPED` > `TRUNCATED` > `COMPLETE`) that folds several independently-scoped owner answers about the same phase directory into one composite signal — NOT a re-derivation of any owner (each owner's own algorithm is untouched; only their already-computed `scope` verdicts are combined), but new coordination logic no single owner has the visibility to express. Every exposed field carries PARSED values only, never raw document text — this is structural, not advisory: a diagnostic rule given only the parsed value cannot re-derive a field's location the way `#3162`'s three inert `Current Phase` literal-search predicates did. Read failures on STATE.md (exists-but-unreadable, distinct from absent) are reported via the Unusable Input Diagnostic Module's `warnUnusableInput(UNUSABLE_REASON.STATE_UNREADABLE)`. Guarded by `scripts/lint-planning-snapshot-bypass-drift.cjs` (ratcheted per Decision 4(e), scoped to `DIAGNOSTIC_RULE_FUNCTIONS` — currently `cmdValidateHealth` in `src/verify.cts` only, acknowledging its existing raw `.planning/` reads as debt owned by Phase 11, #3309, which migrates it onto this snapshot). Source of truth: `gsd-core/bin/lib/planning-snapshot.cjs` (generated from `src/planning-snapshot.cts`). Design: `.gsd/phase/refactor-3308-planning-snapshot-parsed-projection/40-design.md`.
### Plan Document Module
Leaf module owning the parse of a `*-PLAN.md` document BODY: `<objective>` extraction, the `<task>` block grammar (with the legacy `## Task N` heading fallback), per-task `<files>` / `<acceptance_criteria>` / `<done>`, and the frontmatter-derived scheduling metadata (`wave`, `depends_on`, `autonomous`, `agent_hint`, `files_modified`). `parsePlanDocument(content, planPath?) → PlanDocument`; `TASK_KIND` is a frozen `{AUTO, CHECKPOINT}` enum so a `<task type="checkpoint:*">` block — which carries an entirely different element set (`<decision>`/`<what-built>`, no `<name>`/`<files>`) — is reported as its own kind rather than as a malformed auto task. Extracted from `cmdPhasePlanIndex`'s inline pass-1 loop (#2790) because two commands in two families now need it (`phase.plan-index` and `planning.inspect`); leaving it in `phase.cts` would have forced a `planning` → `phase` dependency, and copying it is the `DEFECT.GENERATIVE-FIX` shape. **NOT an ADR-3180 §7 derivation** — §6 puts the document-parsing layer (#2143) explicitly out of that epic's scope; this module answers "what does this plan document say", never "how many plans are outstanding" (`scanPhasePlans`, §7.5) or "is this phase complete" (`isPhaseComplete`, §7.4). Behaviour is byte-for-behaviour identical to the prior inline code, INCLUDING the invariant `taskCount === tasks.length === (xmlTaskCount || mdTaskCount)` and its known fence-blindness (a `## Task 1` inside a fenced block still counts) — characterised, not endorsed: changing it would silently alter `phase.plan-index`'s output for existing projects. Source of truth: `gsd-core/bin/lib/plan-document.cjs` (generated from `src/plan-document.cts`).
Leaf module owning the parse of a `*-PLAN.md` document BODY: `<objective>` extraction, the `<task>` block grammar (with the legacy `## Task N` heading fallback), per-task `<files>` / `<acceptance_criteria>` / `<done>`, and the frontmatter-derived scheduling metadata (`wave`, `depends_on`, `autonomous`, `agent_hint`, `files_modified`). `parsePlanDocument(content, planPath?) → PlanDocument`; `TASK_KIND` is a frozen `{AUTO, CHECKPOINT}` enum so a `<task type="checkpoint:*">` block — which carries an entirely different element set (`<decision>`/`<what-built>`, no `<name>`/`<files>`) — is reported as its own kind rather than as a malformed auto task. Extracted from `cmdPhasePlanIndex`'s inline pass-1 loop (#2790) because two commands in two families now need it (`phase.plan-index` and `planning.inspect`); leaving it in `phase.cts` would have forced a `planning` → `phase` dependency, and copying it is the `DEFECT.GENERATIVE-FIX` shape. **NOT an ADR-3180 §7 derivation** — §6 puts the document-parsing layer (#2143) explicitly out of that epic's scope; this module answers "what does this plan document say", never "how many plans are outstanding" (`scanPhasePlans`, §7.5) or "is this phase complete" (`isPhaseComplete`, §7.4). Behaviour is byte-for-behaviour identical to the prior inline code, INCLUDING the invariant `taskCount === tasks.length === (xmlTaskCount || mdTaskCount)` and its known fence-blindness (a `## Task 1` inside a fenced block still counts) — characterised, not endorsed: changing it would silently alter `phase.plan-index`'s output for existing projects. As of ADR-3646 (#3970) this module also owns the `tracker-id` task attribute: parsed verbatim off the `<task>` block into `PlanTask.trackerId` (string or `null`), never split into prefix/id here — that split is the Task Content Resolution Module's job. Source of truth: `gsd-core/bin/lib/plan-document.cjs` (generated from `src/plan-document.cts`).
### Task Content Resolution Module
Given a task's `tracker-id` attribute (Plan Document Module) and the set of installed capabilities' `taskContentResolver` declarations, resolves that task's content — `action`/`verify`/`acceptanceCriteria`/`readFirst`/`done` — from the matching external tracker via one bounded subprocess call (ADR-3646 Phase 1, #3970), or reports that no resolution applies. Pure/impure split: `splitTrackerId` (first-`:`-only split, colons after the first stay in the id verbatim), `findResolver` (matches a capability's `trackerPrefix`, returns `null`/one match/`'ambiguous'`, never silently picks one), and `buildInvocation` (expands `invoke.args`' `{{id}}` placeholder) are pure and total, never throwing even on hostile third-party-shaped capability input; `resolveTaskContent` is the one impure boundary, spawning through an injectable `execFn` (`spawnSync` with a bounded `timeout`) so tests never spawn a real process or wait a real timeout. **Four non-throwing outcomes** (`TASK_CONTENT_RESULT`: `not-applicable`, `no-resolver`, `resolved`, `empty`) and **HARD-HALT on four throwing outcomes** (`ResolverAmbiguousError`, `ResolverFailedError`, `ResolverTimeoutError`, `ResolverMalformedOutputError`) — ADR-3646 Decision 4: an ambiguous resolver match, a non-zero resolver exit, a timeout, or malformed resolver stdout must never degrade to a silently-empty or silently-picked result. Wired at `task resolve-content --plan <path> --task-id <tracker-id> --raw` (`task-command-router.cts`'s `routeResolveContent`), which turns each of the four throws into the CLI's own non-zero exit rather than swallowing them into a `{resolved: false}` JSON answer; capabilities are loaded through the merged first-party + validated-installed-overlay registry (ADR-1244 D2) so a third-party capability's resolver is honored. Source of truth: `gsd-core/bin/lib/task-content-resolution.cjs` (generated from `src/task-content-resolution.cts`).
### Planning Inspect Module
Module owning the **schema-v1 canonical planning snapshot** emitted by the read-only `planning inspect` query (#2790), for downstream harness UIs that need truthful `.planning/` state without parsing ROADMAP/REQUIREMENTS/PLAN/SUMMARY Markdown a second time. `buildPlanningInspect(cwd) → payload`; `cmdPlanningInspect(cwd, raw)` emits it through `output()` (so the existing >50 KB `@file:` spill seam applies unchanged). `PLANNING_INSPECT_SCHEMA_VERSION = 1` is the wire contract — a consumer MUST reject any other value rather than best-effort-parse an unknown shape. Composes, never re-derives: milestone identity/windowing and phase enumeration via `buildPlanningSnapshot` (Planning Snapshot Module), completion via `isPhaseComplete` (§7.4, disk-strict), live-plan counting via `scanPhasePlans` (§7.5), percent via `clampPercent` (§7.6), STATE fields via `stateFieldValue`/`stateCurrentPositionSlice` (§7.7), plan bodies via `parsePlanDocument`, requirement IDs via `parseRequirements`, UAT items via `parseUatItems`/`selectPhaseUatFiles`. **It deliberately does NOT serialize `PlanningSnapshot`**: that shape is the §8.1 diagnostic-rule subject and is explicitly additive/growing (4 fields at Phase 10, 20+ by Phase 12), so handing it to external consumers would freeze an internal contract by accident (Hyrum's Law) — this module declares its own flat schema and maps into it, and a field added to `PlanningSnapshot` must never change schema-v1 output. Three frozen enums carry every non-answer — `INSPECT_DIAGNOSTIC`, `TASK_STATUS` (`done|pending|unknown`), `PROVENANCE` (`task_scoped|plan_scoped|absent`) and `AGREEMENT` (`agreed|conflicting|unknown`) — because unknown or conflicting evidence serializes as `unknown` plus a diagnostic and is **never inferred, reconciled, or defaulted**; keys are always present, `null` is the explicit non-answer. Roadmap acceptance, verification and UAT are reported **side by side and never folded into one verdict**, and a ROADMAP checkbox is emitted with `authoritative: false` per §7.4. **Not a diagnostic rule**, and deliberately NOT registered in `scripts/lint-planning-snapshot-bypass-drift.cjs`, which is `DIAGNOSTIC_RULE_FUNCTIONS`-scoped and must remain prunable to zero when #3309 lands. Dispatched by the Planning Command Router (`src/planning-command-router.cts`, family `planning`, subcommand `inspect`, no arguments in v1 — a stray positional or unknown flag is a fail-loud `ERROR_REASON.USAGE`). Source of truth: `gsd-core/bin/lib/planning-inspect.cjs` (generated from `src/planning-inspect.cts`). Design: `.gsd/phase/feat-2790-planning-inspect/40-design.md`.

View File

@@ -693,6 +693,43 @@ walkthrough: [Consume the planning snapshot](how-to/consume-the-planning-snapsho
---
### `task resolve-content --plan <path> --task-id <id> --raw`
Resolves one task's content (`action`/`verify`/`acceptance_criteria`/`read_first`/`done`) from an
external issue tracker instead of reading it inline from a task's `PLAN.md` body. Called by
`execute-plan.md`'s per-task loop, once per task carrying a `tracker-id` attribute, before that
task's read_first gate. See [ADR-3646](adr/3646-per-task-content-resolution-seam.md) and
[Develop a task-content resolver capability](how-to/develop-a-task-content-resolver-capability.md).
| Argument | Required | Description |
|----------|----------|-------------|
| `--plan` | **Yes** | Path to the `PLAN.md` the task belongs to |
| `--task-id` | **Yes** | The task's `tracker-id` attribute value, e.g. `beads:GSD-42` |
| `--raw` | No | Machine-readable JSON output |
**Exit codes:**
| Exit | Meaning |
|------|---------|
| `0` | Resolution attempted (or not needed) — see `resolved`/`reason` below |
| non-zero | **Hard halt.** A resolver was found and invoked but failed (tracker unreachable, id not found, timeout, malformed JSON output). stderr names the tracker-id, the tracker prefix, and the resolver's error. Never fall back to inline `PLAN.md` content on this outcome. |
**Output fields (JSON, exit 0 only):**
| Field | Type | Description |
|-------|------|-------------|
| `resolved` | `boolean` | `true` only when a resolver was found, invoked, and returned non-empty content |
| `reason` | `string` | Present when `resolved` is `false`: `"no-resolver"` (task has a `tracker-id` but no installed capability declares a matching `trackerPrefix`) or `"empty"` (the resolver ran successfully but returned empty/absent content — the one legitimate pre-migration fallback case) |
| `content` | `object` | Present when `resolved` is `true`. Supersedes this task's inline `<action>`/`<verify>`/`<acceptance_criteria>`/`<read_first>`/`<done>` for every downstream gate in the execute step |
```bash
node gsd-tools.cjs task resolve-content --plan .planning/phases/03-name/03-1-PLAN.md --task-id beads:GSD-42 --raw
```
`execute-plan.md` only invokes this command when the task carries a `tracker-id` attribute; a task with no `tracker-id` is unaffected.
---
## Navigation Commands
### `/gsd-next`

View File

@@ -204,6 +204,7 @@
- [Hooks Declare Their Crash Policy](#3911-hooks-declare-their-crash-policy)
- [gsd-tools Declares Outcomes, Pinned at v1](#3912-gsd-tools-declares-outcomes-pinned-at-v1)
- [Reachable Lint Rules and a Non-Destructive Quick-Task Append](#3951-reachable-lint-rules-and-a-non-destructive-quick-task-append)
- [Per-Task External-Tracker Content-Resolution Seam](#3970-per-task-external-tracker-content-resolution-seam)
---
@@ -4142,6 +4143,47 @@ regressing to the old seconds-based default had been inert. It is now row-scoped
---
### 3970. Per-Task External-Tracker Content-Resolution Seam
**Purpose:** Let a capability declare that an external issue tracker — beads, Linear, Jira,
GitHub Issues — owns a task's *content* (`<action>`/`<verify>`/`<acceptance_criteria>`/
`<read_first>`/`<done>`), not just its status, so `execute-plan.md` can resolve that content
from the tracker at execution time instead of reading it inline out of `PLAN.md`.
**What changed (ADR-3646, #3970):**
- A new optional feature-body manifest field, `taskContentResolver`, declares a `trackerPrefix`
(matched against a task's `<task tracker-id="beads:GSD-42">` attribute — everything before the
first `:`) and a bounded `invoke` (`binary`, `args` carrying the `{{id}}` placeholder,
`timeoutMs`).
- `execute-plan.md`'s per-task loop gains one new, unconditional call before that task's
`read_first` gate: `gsd_run task resolve-content --plan <path> --task-id <tracker-id> --raw`.
A task with no `tracker-id` attribute is unaffected — the call is only made when the attribute
is present, and resolves instantly to a no-op for every project that declares none.
- **The safety property is a real process exit code, not a prose dispatch.** No capability
registered for the tracker, or resolution succeeds with empty content, exits `0` with
`resolved: false` and falls back to inline `PLAN.md` — the one legitimate pre-migration
boundary case. Resolution succeeding with non-empty content exits `0` with `resolved: true` and
its `content` supersedes the task's inline fields for every downstream gate in the execute step.
A resolver that is declared but fails — tracker unreachable, id not found, timeout, malformed
JSON — makes `task resolve-content` itself **exit non-zero**, which `execute-plan.md` treats as
a **hard halt**: stop, surface the tracker-id/prefix/stderr, never fall back to stale
`PLAN.md` content.
- `execute:task` is a new dispatch shape below wave granularity, deliberately **not** one of the
12 existing loop extension points (`discuss:pre` … `ship:post`) and not routed through
`gsd_run loop render-hooks <point>` / `activeHooks`. It exists because the existing
`step`/`gate` prose-dispatch mechanism cannot deliver a hard-halt guarantee while dispatch
reliability at that layer is an open concern (#3647) — see ADR-3646's Context and Rejected
Alternatives for the full reasoning.
See [Develop a task-content resolver capability](../how-to/develop-a-task-content-resolver-capability.md)
for the authoring walkthrough, [Capability manifest → `taskContentResolver`](../reference/capability-manifest.md#taskcontentresolver)
for the field reference, and
[`loop-hook-dispatch.md`](../../gsd-core/references/loop-hook-dispatch.md#the-executetask-point-a-different-shape)
for how `execute:task` differs from the twelve prose-dispatched points.
---
_Generated by `scripts/gen-features.cjs` — add a fragment under `docs/features/` and run `--write`._
<!-- FEATURES:END -->

View File

@@ -515,6 +515,7 @@
"state.cjs",
"surface.cjs",
"task-command-router.cjs",
"task-content-resolution.cjs",
"teams-status.cjs",
"template.cjs",
"text-lines.cjs",

View File

@@ -646,7 +646,8 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
| `state-document.cjs` | Pure STATE.md field extraction, replacement, status normalization, and progress calculation transforms |
| `milestone-lock.cjs` | Milestone lock (compiled from `src/milestone-lock.cts`, gitignored) — advisory (phase, session id) claim over STATE.md's single Current Position slot: `.planning/milestone.lock` claim IO, liveness (TTL + heartbeat), conflict detection, and the shared stderr warning; consumed by `state.begin-phase` / `state.advance-plan` / `phase.complete` so parallel phases in one working tree get a visible conflict instead of silently overwriting each other (#3311) |
| `surface.cjs` | Runtime surface module — manages the runtime enable/disable surface state independently of the install-time profile marker (ADR-0011 Phase 2) |
| `task-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools task` |
| `task-command-router.cjs` | Thin CJS subcommand router adapter for `gsd-tools task`; `resolve-content` subcommand resolves a task's `tracker-id` via `task-content-resolution.cjs` (ADR-3646, #3970) |
| `task-content-resolution.cjs` | Resolves a task's `tracker-id` to external-tracker content via a capability-declared `taskContentResolver` and a bounded subprocess call; hard-halts on ambiguous/failed/timeout/malformed resolution (compiled from `src/task-content-resolution.cts`, gitignored) (ADR-3646, #3970) |
| `teams-status.cjs` | Detects agent-teams status from environment and runtime; pure core (#1355) |
| `template.cjs` | Template selection and filling with variable substitution |
| `text-lines.cjs` | Line-terminator handling seam — `splitLines`/`normalizeEol`/`detectEol`/`joinLines`, the sole owner of `\r?\n` splitting and CRLF normalization; closes #3360's split-then-match fix in `frontmatter.cjs` (ADR-3212 §3, epic #3212 Phase 2, #3413) |

View File

@@ -64,6 +64,7 @@ Language versions: [English](README.md) · [Português (pt-BR)](pt-BR/README.md)
- [Design a UI phase](how-to/design-a-ui-phase.md) — use the UI phase loop for frontend and visual work
- [Enable live-DOM verification](how-to/enable-live-dom-verification.md) — opt a project into browser-backed UI acceptance checks during execution, handle the browser-profile lock, and tell "nothing to report" apart from "could not look"
- [Develop a Capability for GSD 1.5+](how-to/develop-a-capability.md) — add feature Capabilities, hook fragments, and registry entries
- [Develop a task-content resolver capability](how-to/develop-a-task-content-resolver-capability.md) — declare a `taskContentResolver` so `execute-plan.md` resolves per-task content from your external issue tracker instead of `PLAN.md`
- [Ship a reviewer lane in your capability](how-to/ship-a-reviewer-lane.md) — declare a `reviewer` body so `/gsd-review` discovers, invokes, and renders your external review CLI or model endpoint
- [List your reviewer lane in the registry](how-to/list-your-reviewer-lane.md) — publish a lane you have built to the Reviewer Lane Registry so other people can find and install it
- [Take over a capability or EoS integration](how-to/take-over-a-capability-or-eos.md) — assume maintainership of an existing third-party capability, reviewer lane, or EoS host integration through a handoff, an adoption fork, first-party absorption, or a de-listing

View File

@@ -0,0 +1,42 @@
---
id: 3970
title: Per-Task External-Tracker Content-Resolution Seam
group: v1.7.0 Features
---
**Purpose:** Let a capability declare that an external issue tracker — beads, Linear, Jira,
GitHub Issues — owns a task's *content* (`<action>`/`<verify>`/`<acceptance_criteria>`/
`<read_first>`/`<done>`), not just its status, so `execute-plan.md` can resolve that content
from the tracker at execution time instead of reading it inline out of `PLAN.md`.
**What changed (ADR-3646, #3970):**
- A new optional feature-body manifest field, `taskContentResolver`, declares a `trackerPrefix`
(matched against a task's `<task tracker-id="beads:GSD-42">` attribute — everything before the
first `:`) and a bounded `invoke` (`binary`, `args` carrying the `{{id}}` placeholder,
`timeoutMs`).
- `execute-plan.md`'s per-task loop gains one new, unconditional call before that task's
`read_first` gate: `gsd_run task resolve-content --plan <path> --task-id <tracker-id> --raw`.
A task with no `tracker-id` attribute is unaffected — the call is only made when the attribute
is present, and resolves instantly to a no-op for every project that declares none.
- **The safety property is a real process exit code, not a prose dispatch.** No capability
registered for the tracker, or resolution succeeds with empty content, exits `0` with
`resolved: false` and falls back to inline `PLAN.md` — the one legitimate pre-migration
boundary case. Resolution succeeding with non-empty content exits `0` with `resolved: true` and
its `content` supersedes the task's inline fields for every downstream gate in the execute step.
A resolver that is declared but fails — tracker unreachable, id not found, timeout, malformed
JSON — makes `task resolve-content` itself **exit non-zero**, which `execute-plan.md` treats as
a **hard halt**: stop, surface the tracker-id/prefix/stderr, never fall back to stale
`PLAN.md` content.
- `execute:task` is a new dispatch shape below wave granularity, deliberately **not** one of the
12 existing loop extension points (`discuss:pre` … `ship:post`) and not routed through
`gsd_run loop render-hooks <point>` / `activeHooks`. It exists because the existing
`step`/`gate` prose-dispatch mechanism cannot deliver a hard-halt guarantee while dispatch
reliability at that layer is an open concern (#3647) — see ADR-3646's Context and Rejected
Alternatives for the full reasoning.
See [Develop a task-content resolver capability](../how-to/develop-a-task-content-resolver-capability.md)
for the authoring walkthrough, [Capability manifest → `taskContentResolver`](../reference/capability-manifest.md#taskcontentresolver)
for the field reference, and
[`loop-hook-dispatch.md`](../../gsd-core/references/loop-hook-dispatch.md#the-executetask-point-a-different-shape)
for how `execute:task` differs from the twelve prose-dispatched points.

View File

@@ -0,0 +1,172 @@
# Develop a task-content resolver capability
**Goal:** Declare a `taskContentResolver` in a capability manifest so `execute-plan.md` resolves
a task's `<action>`/`<verify>`/`<acceptance_criteria>`/`<read_first>`/`<done>` content from your
external issue tracker instead of reading it inline from `PLAN.md`.
**Prerequisites:** You already have a capability (`capability.json`), or you are creating one —
see [Develop a Capability for GSD 1.5+](develop-a-capability.md) first if this is your first one.
GSD 1.28 or later ([ADR-3646](../adr/3646-per-task-content-resolution-seam.md)).
---
## Why this exists
A project that wants an external tracker — beads, Linear, Jira, GitHub Issues — to own task
*content*, not just task *status*, has no seam for that today: `execute-plan.md` reads every
task's instructions directly out of the `PLAN.md` task block. `taskContentResolver` adds one, with
a hard-halt guarantee: if your tracker is declared as the source of truth and it fails to resolve,
execution stops rather than silently falling back to stale `PLAN.md` text. That guarantee is the
entire point of the feature — a silent fallback would require the tracker and `PLAN.md` to stay in
sync forever, defeating the reason to move content out of `PLAN.md` in the first place.
---
## Declare the resolver
Add a `taskContentResolver` block to your capability's manifest body (`role: "feature"` only):
```json
{
"id": "beads",
"role": "feature",
"version": "1.0.0",
"title": "Beads issue tracker",
"description": "Resolves task content from the bd issue tracker.",
"tier": "standard",
"requires": [],
"runtimeCompat": { "supported": ["*"], "unsupported": [] },
"skills": [],
"agents": [],
"hooks": [],
"config": {},
"steps": [],
"contributions": [],
"gates": [],
"taskContentResolver": {
"trackerPrefix": "beads",
"invoke": {
"binary": "bd",
"args": ["show", "{{id}}", "--json"],
"timeoutMs": 10000
}
}
}
```
Two fields decide whether the seam works at all:
- **`trackerPrefix`** must match the prefix of a task's `tracker-id` attribute — everything
before the **first** `:`. Given `<task tracker-id="beads:GSD-42">`, the prefix is `beads` and
the id passed to your resolver is `GSD-42`. If a tracker's own ids contain colons, that is fine:
only the first colon splits prefix from id, so `beads:team:GSD-42` resolves to id `team:GSD-42`.
- **`invoke.args`** must contain the `{{id}}` placeholder at least once — GSD substitutes it with
the task's id before spawning your binary. A declaration whose `args` never carries the
placeholder fails validation at install time, because the id could never reach your resolver.
`invoke.timeoutMs` is required. An unbounded resolver subprocess is this repo's named Unbounded
Subprocesses defect class — declare a bound that matches how long your tracker's lookup actually
takes, plus margin.
`trackerPrefix` must be unique across the merged first-party ∪ overlay capability set; a
collision — two installed capabilities both claiming `"beads"` — is a build-time validation
error, not a runtime ambiguity.
---
## What your resolver must output
`execute-plan.md` invokes your `invoke.binary`/`invoke.args` and expects a single JSON object on
stdout when the lookup succeeds:
| Field | Type | Required | Maps to |
|---|---|---|---|
| `description` | string | Yes | The task's `<action>` |
| `verify` | string | No | The task's `<verify>` |
| `acceptance_criteria` | string[] | No | The task's `<acceptance_criteria>` |
| `read_first` | string[] | No | The task's `<read_first>` |
| `done` | string | No | The task's `<done>` |
An absent or empty-string `description` is treated as "nothing resolved" — `execute-plan.md`
falls back to the task's inline `PLAN.md` content, the one legitimate pre-migration boundary case
(for tasks authored before your tracker migration). This is the *only* silent fallback path; every
other failure is a hard halt.
**Exit code and stderr matter.** Exit `0` with valid JSON on stdout is the only success path.
Anything else — a non-zero exit, a timeout past `invoke.timeoutMs`, or stdout that fails to parse
as JSON — is treated as a resolution failure. Write a clear one-line reason to stderr; it is
surfaced verbatim to the person watching execution. Stderr on a **successful** (exit 0) run is not
an error — write informational logs there if your CLI already does; only the exit code and the
JSON parse outcome decide success.
---
## What happens on failure
A resolver that is declared, invoked, and fails — non-zero exit, timeout, or malformed JSON —
makes `gsd_run task resolve-content` itself exit non-zero. `execute-plan.md` treats that as a
**hard halt**: it stops before doing any work on the task, surfaces the tracker-id, the tracker
prefix, and your resolver's stderr, and never proceeds to read the task's inline `PLAN.md` content
as a substitute. See [ADR-3646](../adr/3646-per-task-content-resolution-seam.md) for why this is
the load-bearing safety property of the whole feature, and
[`loop-hook-dispatch.md`](../../gsd-core/references/loop-hook-dispatch.md#the-executetask-point-a-different-shape)
for how this call site differs from the twelve prose-dispatched loop extension points.
---
## Worked example: a `bd`/beads-shaped resolver
Suppose `bd show GSD-42 --json` already returns:
```json
{
"id": "GSD-42",
"title": "Add rate-limit check to login",
"status": "open",
"description": "Add a rate-limit check to processLogin using the existing RateLimiter.",
"acceptance": [
"Login attempts beyond the configured limit return 429",
"Existing successful-login tests still pass"
],
"notes": "See src/util/rate.ts for the existing limiter."
}
```
Your resolver is a thin adapter, not `bd` itself — it maps `bd`'s field names onto the shape
`execute-plan.md` expects and exits non-zero on anything `bd` itself reports as a failure:
```bash
#!/usr/bin/env bash
set -euo pipefail
id="$1"
raw=$(bd show "$id" --json)
node -e '
const raw = JSON.parse(process.argv[1]);
const out = {
description: raw.description || "",
acceptance_criteria: raw.acceptance || [],
};
process.stdout.write(JSON.stringify(out));
' "$raw"
```
Declare it as the `invoke.binary`/`invoke.args` pair (or point `invoke.binary` at `bd` directly if
its own `--json` output already matches the expected field names — no adapter needed in that
case). Either way, `invoke.args` must carry `{{id}}` so GSD can substitute the task's tracker id
before spawning it.
---
## Related
- [ADR-3646](../adr/3646-per-task-content-resolution-seam.md) — the design decision and rejected
alternatives
- [Capability manifest → `taskContentResolver`](../reference/capability-manifest.md#taskcontentresolver) —
the full field table
- [`loop-hook-dispatch.md`](../../gsd-core/references/loop-hook-dispatch.md#the-executetask-point-a-different-shape) —
how `execute:task` differs from the twelve loop extension points
- [Develop a Capability for GSD 1.5+](develop-a-capability.md) — manifests, registry generation,
and federated config

View File

@@ -117,6 +117,28 @@ Gates check a condition at a loop extension point and optionally block progressi
| Predicate | `{ "predicate": { "kind": "artifact-exists" \| "config-equals" \| …, … } }` | Yes | Declarative; no code path. |
| Agent verdict | `{ "agentVerdict": { "ref": …, "prompt": … } }` | No (forced advisory) | LLM evaluation; non-deterministic checks may not halt the loop. |
### `taskContentResolver`
Declares that this capability resolves per-task content (`<action>`/`<verify>`/
`<acceptance_criteria>`/`<read_first>`/`<done>`) from an external issue tracker instead of
`execute-plan.md`'s per-task loop reading it inline from a task's `PLAN.md` body. This is **not**
one of `steps` / `contributions` / `gates`, and it does not use a `point` value from the closed
12-point vocabulary above — it is dispatched directly, once per task, by `execute-plan.md` before
that task's `read_first` gate, documented separately in
[`loop-hook-dispatch.md`](../../gsd-core/references/loop-hook-dispatch.md#the-executetask-point-a-different-shape).
See [ADR-3646](../adr/3646-per-task-content-resolution-seam.md) for the full design and
[Develop a task-content resolver capability](../how-to/develop-a-task-content-resolver-capability.md)
for the authoring walkthrough.
| Sub-field | Type | Required | Description |
|---|---|---|---|
| `trackerPrefix` | string (kebab-case) | Yes | Matches the prefix of a task's `<task tracker-id="beads:GSD-42">` attribute — everything before the **first** `:`. Text after the first colon, including further colons, is passed through verbatim as the id. Must be unique across the merged first-party ∪ overlay capability set. |
| `invoke.binary` | string | Yes | Executable name or path for the resolver subprocess. |
| `invoke.args` | string[] | Yes | Argv passed to `invoke.binary`. Must contain the `{{id}}` placeholder at least once — GSD substitutes it with the task's tracker id (everything after the first `:`); an `args` array that never carries the placeholder fails validation, since the id could never reach the resolver. |
| `invoke.timeoutMs` | number | Yes | Bound on the subprocess invocation. Required — an unbounded resolver subprocess is this repo's named Unbounded Subprocesses defect class. A resolver exceeding this bound is killed and `task resolve-content` exits non-zero. |
`taskContentResolver` is feature-role only (`role: "feature"`); it is not admissible on `role: "runtime"` or `role: "reviewer"` bodies.
---
## Valid `point` values
@@ -289,6 +311,7 @@ The following invariants are enforced at **build time** by `scripts/gen-capabili
- **`engines.gsd` is a hard gate.** A capability whose `engines.gsd` range does not satisfy the installed GSD version is blocked at install and skipped (with a warning) at load time.
- **Path confinement.** Declared module paths may not use parent-directory traversal (`../`); modules are `require()`'d only from the capability's own install root.
- **Reserved namespace.** Capability `id` values beginning with `gsd-`, `gsd-core-`, or `anthropic-` are reserved; third-party capabilities using these prefixes are rejected.
- **`taskContentResolver.trackerPrefix` uniqueness, feature-only.** `trackerPrefix` must be unique across the merged first-party ∪ overlay capability set (mirrors `reviewsSection` uniqueness on the reviewer body above); a collision is a build-time violation. `taskContentResolver` is admissible only on `role: "feature"` bodies.
---

View File

@@ -345,6 +345,11 @@ export default tseslint.config(
// src/pattern.cts — module resolution for a .cts source is relative to
// src/, not the output dir). Same verbatim-third-party exemption.
'src/vendor/**',
// #3970 (ADR-3646 Phase 1): tsc-generated runtime artifact — the
// default `import childProcess from 'node:child_process'` import emits
// tsc's `__importDefault` helper (uses `var`), same class as 007/009/010
// above. Lint the src/task-content-resolution.cts source, not this.
'gsd-core/bin/lib/task-content-resolution.cjs',
// #3904 (ADR-3889 Phase 0): tsc-generated runtime artifact — generated
// by scripts/gen-scripts-cli-exit.cjs from a fresh compile of
// src/cli-exit.cts, and byte-guarded by `npm run lint:generated-sync`

View File

@@ -396,6 +396,11 @@ function validateCapability(cap, folderId) {
errors.push(...validateRuntimeBody(cap));
// A host that is ALSO a reviewer keeps exactly one manifest (ADR-2782 D1).
errors.push(...validateReviewerBody(cap));
// ADR-3646: a runtime capability installs a host CLI, it does not resolve
// task content — taskContentResolver is feature-only.
if (cap.taskContentResolver !== undefined) {
errors.push('role:runtime capability must not have a "taskContentResolver" body (feature-only field)');
}
} else if (cap.role === 'reviewer') {
// ADR-2782 D3 — a lane that is not an install target. No runtime body, no
// install surface, no runtimeCompat (it surfaces through no host runtime).
@@ -718,6 +723,118 @@ function validateFeatureBody(cap) {
}
}
// ADR-3646: optional per-task external-tracker content-resolution seam.
errors.push(...validateTaskContentResolver(cap));
return errors;
}
/**
* ADR-3646 — validate an OPTIONAL `taskContentResolver` body on a `role:
* "feature"` capability. Absence is never an error (most manifests won't
* have one); presence is strictly validated.
*
* Wrapped in try/catch to degrade any unexpected throw (a hostile Proxy, a
* throwing getter, etc.) to a single validation error rather than crashing
* every consumer of loadRegistry, per the #1461 OVL-1 discipline that
* `validateReviewerBody` follows.
*
* @param {object} cap The parsed capability manifest.
* @returns {string[]} Array of error strings; empty = valid or absent.
*/
function validateTaskContentResolver(cap) {
try {
return validateTaskContentResolverFields(cap);
} catch (err) {
return ['capability taskContentResolver body could not be validated: ' + safeErrorMessage(err)];
}
}
/**
* Upper bound for `taskContentResolver.invoke.timeoutMs`, specific to this
* field only. `isPositiveIntegerMs()` has no ceiling and stays that way — it
* is shared with the reviewer lane's `timeoutFloorMs` and probe `timeoutMs`,
* which may legitimately need a longer or unbounded value. Without a ceiling
* here, a manifest could declare `Number.MAX_SAFE_INTEGER` and let
* `resolve-content` hang near-indefinitely on a stuck/malicious resolver,
* defeating the feature's "bounded subprocess" design intent.
*/
const TASK_CONTENT_RESOLVER_TIMEOUT_CEILING_MS = 120000;
function validateTaskContentResolverFields(cap) {
const errors = [];
if (typeof cap !== 'object' || cap === null || Array.isArray(cap)) return errors;
const tcr = cap.taskContentResolver;
if (tcr === undefined) return errors; // optional — never an error to omit
const ctx = 'capability "' + (typeof cap.id === 'string' ? cap.id : '(unknown)') + '"';
if (typeof tcr !== 'object' || tcr === null || Array.isArray(tcr)) {
const got = tcr === null ? 'null' : Array.isArray(tcr) ? 'array' : typeof tcr;
errors.push(
ctx + ' taskContentResolver must be an object (got: ' + got + '). ' +
'Omit the key entirely to declare no resolver — an explicit null is not an omission.',
);
return errors; // cannot validate fields of a non-object
}
// ── trackerPrefix — same grammar as a capability id ─────────────────────
if (typeof tcr.trackerPrefix !== 'string' || tcr.trackerPrefix.length === 0 || !KEBAB_RE.test(tcr.trackerPrefix)) {
errors.push(
ctx + ' taskContentResolver.trackerPrefix must be a non-empty kebab-case string matching ' +
String(KEBAB_RE) + ' (got: ' + describeValue(tcr.trackerPrefix) + ')',
);
}
// ── invoke ────────────────────────────────────────────────────────────
const inv = tcr.invoke;
if (typeof inv !== 'object' || inv === null || Array.isArray(inv)) {
const got = inv === null ? 'null' : Array.isArray(inv) ? 'array' : typeof inv;
errors.push(ctx + ' taskContentResolver.invoke must be an object (got: ' + got + ')');
return errors; // cannot validate sub-fields of a non-object
}
if (typeof inv.binary !== 'string' || inv.binary.length === 0) {
errors.push(ctx + ' taskContentResolver.invoke.binary must be a non-empty string');
}
if (!Array.isArray(inv.args)) {
errors.push(ctx + ' taskContentResolver.invoke.args must be an array of strings');
} else {
let hasPlaceholder = false;
for (const a of inv.args) {
if (typeof a !== 'string') {
errors.push(ctx + ' taskContentResolver.invoke.args entries must be strings (got: ' + describeValue(a) + ')');
} else if (a === '{{id}}') {
hasPlaceholder = true;
}
}
if (!hasPlaceholder) {
errors.push(
ctx + ' taskContentResolver.invoke.args must contain a "{{id}}" placeholder — ' +
'without it the resolved tracker id could never reach the resolver subprocess',
);
}
}
if (!isPositiveIntegerMs(inv.timeoutMs)) {
errors.push(
ctx + ' taskContentResolver.invoke.timeoutMs must be a positive integer of milliseconds — ' +
'an unbounded resolver call could hang task execution indefinitely ' +
'(got: ' + describeValue(inv.timeoutMs) + ')',
);
} else if (inv.timeoutMs > TASK_CONTENT_RESOLVER_TIMEOUT_CEILING_MS) {
// Ceiling specific to this field — `isPositiveIntegerMs()` itself stays
// unbounded because it is shared with the reviewer lane's
// `timeoutFloorMs`/probe `timeoutMs`, which have no such ceiling.
errors.push(
ctx + ' taskContentResolver.invoke.timeoutMs must not exceed ' +
TASK_CONTENT_RESOLVER_TIMEOUT_CEILING_MS + 'ms — an unbounded-in-practice value defeats the ' +
'"bounded subprocess" design intent (got: ' + inv.timeoutMs + ')',
);
}
return errors;
}
@@ -943,7 +1060,7 @@ const HTTP_ONLY_INVOKE_FIELDS = ['hostConfigKey', 'defaultHost', 'path', 'model
// Feature-only fields are as forbidden on a lane-only capability as on a runtime
// one; a `role: "reviewer"` capability owns no artefacts and wires no loop point.
const FEATURE_FIELDS_FORBIDDEN_ON_REVIEWER = ['skills', 'agents', 'steps', 'contributions', 'gates', 'hooks', 'activationKey'];
const FEATURE_FIELDS_FORBIDDEN_ON_REVIEWER = ['skills', 'agents', 'steps', 'contributions', 'gates', 'hooks', 'activationKey', 'taskContentResolver'];
// GATE A: installSurface → allowed hooksSurface values (DEFECT.GENERATIVE-FIX: parity invariant)
// Derived from the actual pairings in the 16 real runtime descriptors.
@@ -3184,6 +3301,11 @@ function validateCrossCapability(capMap, centralKeys, centralPatterns = []) {
const laneSlugClaims = new Map(); // slug → capId[]
const laneFlagClaims = new Map(); // flag → capId[]
const laneSectionClaims = new Map(); // reviewsSection → capId[]
// ADR-3646: task-content resolver tracker-prefix uniqueness across the
// MERGED first-party ∪ overlay set, mirroring the reviewer-lane collision
// pattern above — two resolvers claiming the same prefix would make
// `execute:task` dispatch ambiguous (which capability's resolver runs?).
const trackerPrefixClaims = new Map(); // trackerPrefix → capId[]
// Claims are ACCUMULATED and reported after the sweep, never reported on the
// second claimant. Reporting pairwise-on-collision looks equivalent and is not:
@@ -3204,6 +3326,15 @@ function validateCrossCapability(capMap, centralKeys, centralPatterns = []) {
};
for (const [capId, cap] of capMap) {
// ADR-3646: a MALFORMED taskContentResolver body was already reported by
// validateCapability — do not double-report; only claim well-shaped
// bodies. Independent of the reviewer-lane checks below, so it runs even
// for capabilities that carry no `reviewer` body at all.
const tcr = cap.taskContentResolver;
if (typeof tcr === 'object' && tcr !== null && !Array.isArray(tcr)) {
claim(trackerPrefixClaims, tcr.trackerPrefix, capId);
}
const r = cap.reviewer;
// A capability with no lane contributes to no uniqueness set. A MALFORMED
// body was already reported by validateCapability — do not double-report.
@@ -3225,6 +3356,7 @@ function validateCrossCapability(capMap, centralKeys, centralPatterns = []) {
[laneSlugClaims, 'slug'],
[laneFlagClaims, 'flag'],
[laneSectionClaims, 'reviewsSection'],
[trackerPrefixClaims, 'taskContentResolver.trackerPrefix'],
]) {
for (const [key, claimants] of claims) {
if (claimants.length < 2) continue;
@@ -3734,6 +3866,7 @@ module.exports = {
validateCommandEntry,
validateRuntimeCompat,
validateFeatureBody,
validateTaskContentResolver,
validateConfigHome,
validateArtifactKindEntry,
validateArtifactLayout,

View File

@@ -96,3 +96,25 @@ Honor `onError` if the check itself errors: `skip` means treat as non-blocking a
If `activeHooks` is absent, null, or an empty array, skip silently and continue to the next
step in the workflow. No output to the user is needed.
## The `execute:task` point (a different shape)
`execute:task` exists below wave granularity — it is evaluated once per task, inside the
`execute:wave:pre` / `execute:wave:post` bracket, immediately before that task's `read_first`
gate. It is **not** one of the 12 points documented above, does not appear in `steps` /
`contributions` / `gates`, and is never dispatched through `gsd_run loop render-hooks <point>` or
this file's `activeHooks` envelope.
Instead, a capability declares task-content resolution directly in its manifest body via
`taskContentResolver` (`trackerPrefix` + a bounded `invoke`) — see
[Capability manifest reference](../../docs/reference/capability-manifest.md). `execute-plan.md`'s
per-task loop calls `gsd_run task resolve-content --plan <path> --task-id <tracker-id> --raw`
directly, an unconditional, required subprocess invocation with a real, binding exit code —
never a prose-dispatched `step`/`gate` entry chosen from an `activeHooks` array.
This point always runs — there is no `when` config gate and no autonomous-mode elision. That is
deliberate, not an oversight: the twelve points above are best-effort prose dispatch, which
`execute:task`'s hard-halt safety property cannot be built on top of (a missed dispatch is
indistinguishable from a legitimate resolver-empty fallback). See
[ADR-3646](../../docs/adr/3646-per-task-content-resolution-seam.md) for the full rationale,
including why a `kind: "gate"` shape was rejected outright.

View File

@@ -218,11 +218,12 @@ Deviations are normal — handle via rules below.
1. Read @context files from prompt
2. **MCP tools:** If CLAUDE.md or project instructions reference MCP tools (e.g. jCodeMunch for code navigation), prefer them over Grep/Glob when available. Fall back to Grep/Glob if MCP tools are not accessible.
3. Per task:
- **MANDATORY read_first gate:** If the task has a `<read_first>` field, you MUST read every listed file BEFORE making any edits. This is not optional. Do not skip files because you "already know" what's in them — read them. The read_first files establish ground truth for the task.
- **Task-content resolution:** If this task is NOT `type="checkpoint:*"` AND it carries a `tracker-id` attribute, run `gsd_run task resolve-content --plan "<plan path>" --task-id "<tracker-id>" --raw` BEFORE anything else for this task — before the read_first gate below, since read_first is itself a field this call can resolve. A `type="checkpoint:*"` task NEVER enters this resolution, regardless of whether it carries a `tracker-id` attribute — its interactive structure always stays sourced from PLAN.md (ADR-3646 Decision 1); see the `type="checkpoint:*"` bullet below. A **non-zero exit is a HARD HALT**: surface the tracker-id, the tracker prefix, and the command's stderr, and STOP — do NOT proceed to read this task's inline PLAN.md content as a fallback (that defeats the point of the seam: see ADR-3646). On exit 0 with `resolved: true`, the returned `content` object SUPERSEDES this task's `<read_first>`/`<action>`/`<verify>`/`<acceptance_criteria>`/`<done>` for every remaining bullet in this step — every gate below reads from the resolved content instead of PLAN.md. On exit 0 with `resolved: false` (for any `reason`), proceed exactly as today and read the task's inline PLAN.md content. A task with no `tracker-id` attribute, or a checkpoint task, is unaffected — this step is unconditional but resolves to a no-op instantly for the common case.
- **MANDATORY read_first gate:** If the task has a `<read_first>` field (or the resolved content carries one), you MUST read every listed file BEFORE making any edits. This is not optional. Do not skip files because you "already know" what's in them — read them. The read_first files establish ground truth for the task.
- `type="auto"`: if `tdd="true"` → TDD execution. Implement with deviation rules + auth gates. Verify done criteria. Commit (see task_commit). Track hash for Summary.
- `type="tracer"`: execute like `type="auto"` (production-quality, real `<verify>`, commit), then run the tracer feedback gate BEFORE any expansion task — an early integration checkpoint. Evaluate in order (#3299). First, `gate="blocking-human"` → STOP → return a `checkpoint:human-verify` via checkpoint_protocol — every mode, auto included (golden rule 6, checkpoints.md). Next, Auto mode active (`AUTO_CHAIN` or `AUTO_CFG`): re-run the tracer `<verify>`; on failure HALT and surface (deviation) — do NOT start expansion tasks. Next, `HUMAN_VERIFY_MODE` is `end-of-phase` (default) AND the tracer's `<verify>` carries only `<automated>` (no `<human-check>`) → re-run the tracer `<verify>`; on failure HALT and surface as a deviation exactly as in the auto-mode branch — never a checkpoint; on success log `⚡ Tracer verified end-to-end — expanding` and continue to expansion, do NOT synthesize a checkpoint. Otherwise (`mid-flight`, or the tracer carries genuine human-observable evidence) → STOP → return a `checkpoint:human-verify` for the tracer via checkpoint_protocol before expansion.
- `type="checkpoint:*"`: STOP → checkpoint_protocol → wait for user → continue only after confirmation.
- **HARD GATE — acceptance_criteria verification:** After completing each task, if it has `<acceptance_criteria>`, you MUST run a verification loop before proceeding:
- **HARD GATE — acceptance_criteria verification:** After completing each task, if it has `<acceptance_criteria>` (inline or resolved), you MUST run a verification loop before proceeding:
1. For each criterion: execute the grep, file check, or CLI command that proves it passes
2. Log each result as PASS or FAIL with the command output
3. If ANY criterion fails: fix the implementation immediately, then re-run ALL criteria

View File

@@ -36,6 +36,7 @@ const DOCS_GUARD_EXEMPT_BASELINE = [
'agent-marker-documentation-guard.test.cjs',
'antigravity-upgrades.test.cjs',
'capability-cli.test.cjs',
'capability-validator-task-content-resolver.test.cjs',
'ci-docs-guard-registry.test.cjs',
'ci-test-scope.test.cjs',
'cline-install.test.cjs',
@@ -107,6 +108,10 @@ const DOCS_GUARD_EXEMPT_DOCS_PATHS = {
'agent-marker-documentation-guard.test.cjs': ['docs/reference', 'docs/reference/workflow-fragments.md'],
'antigravity-upgrades.test.cjs': ['docs/cli', 'docs/cli/gcli-migration', 'docs/cli/permissions'],
'capability-cli.test.cjs': ['docs/reference/gsd-capability-command.md'],
// #3970: cites docs/adr/3646-per-task-content-resolution-seam.md in an
// explanatory comment describing ADR-3646's Decision 3; the file never
// reads that (or any) docs/ file.
'capability-validator-task-content-resolver.test.cjs': ['docs/adr/3646-per-task-content-resolution-seam.md'],
'ci-docs-guard-registry.test.cjs': [
'docs/AGENTS.md', 'docs/COMMANDS.md', 'docs/INVENTORY.md', 'docs/a.md', 'docs/adr',
'docs/adr/0001-example.md', 'docs/adrenaline.md', 'docs/bar.md', 'docs/foo.md', 'docs/how-to/foo.md',

View File

@@ -2,9 +2,10 @@
* Plan Document Module — the single parser for a `*-PLAN.md` document BODY.
*
* Owns: objective extraction, the task-block grammar (`<task>` elements, with
* the legacy `## Task N` heading fallback), planned-file extraction, and the
* frontmatter-derived scheduling metadata (`wave`, `depends_on`, `autonomous`,
* `agent_hint`, `files_modified`).
* the legacy `## Task N` heading fallback — including the optional `tracker-id`
* attribute, ADR-3646 Phase 1, read verbatim and never split here), planned-file
* extraction, and the frontmatter-derived scheduling metadata (`wave`,
* `depends_on`, `autonomous`, `agent_hint`, `files_modified`).
*
* WHY THIS IS A LEAF MODULE. This logic was written inline inside
* `cmdPhasePlanIndex` (`src/phase.cts`). Two commands in two different families
@@ -67,6 +68,14 @@ interface PlanTask {
acceptanceCriteria: string[];
/** `<done>` text, or null. */
done: string | null;
/**
* Verbatim `tracker-id` attribute value (e.g. `beads:GSD-42`), or null.
* Never split or parsed here — that belongs to the resolution seam
* (ADR-3646), not this grammar layer. Null for a checkpoint task (never
* read), an absent attribute, or an empty-string value (`tracker-id=""`
* normalises to null, same as every other optional attribute here).
*/
trackerId: string | null;
}
interface PlanDocument {
@@ -196,6 +205,7 @@ function parseXmlTasks(content: string): PlanTask[] {
plannedFiles: [],
acceptanceCriteria: [],
done: null,
trackerId: null,
};
}
@@ -207,6 +217,7 @@ function parseXmlTasks(content: string): PlanTask[] {
plannedFiles: splitFileList(elementBody(block, 'files')),
acceptanceCriteria: splitCriteria(elementBody(block, 'acceptance_criteria')),
done: collapseWhitespace(elementBody(block, 'done')),
trackerId: tagAttribute(openTag, 'tracker-id'),
};
});
}
@@ -225,6 +236,7 @@ function parseMarkdownTasks(content: string): PlanTask[] {
plannedFiles: [],
acceptanceCriteria: [],
done: null,
trackerId: null,
}));
}

View File

@@ -11,6 +11,20 @@ import path from 'node:path';
// eslint-disable-next-line @typescript-eslint/no-require-imports
import ioMod = require('./io.cjs');
const { output, error, ERROR_REASON } = ioMod;
// eslint-disable-next-line @typescript-eslint/no-require-imports
import planDocumentMod = require('./plan-document.cjs');
const { parsePlanDocument } = planDocumentMod;
// eslint-disable-next-line @typescript-eslint/no-require-imports
import capabilityLoaderMod = require('./capability-loader.cjs');
// eslint-disable-next-line @typescript-eslint/no-require-imports
import taskContentResolutionMod = require('./task-content-resolution.cjs');
const {
resolveTaskContent,
ResolverAmbiguousError,
ResolverFailedError,
ResolverTimeoutError,
ResolverMalformedOutputError,
} = taskContentResolutionMod;
// ─── Types ────────────────────────────────────────────────────────────────────
@@ -32,6 +46,26 @@ interface RouteTaskCommandOptions {
raw: boolean;
}
interface PlanTaskLike {
trackerId: string | null;
}
interface CapabilityLike {
id: string;
taskContentResolver?: unknown;
}
/**
* Testability seam for `routeResolveContent` (mirrors this codebase's other
* routers' `_`-prefixed injection convention, e.g.
* `refactor-trigger-command-router.cts`'s `_git`/`_windows`/`_core`).
* Production callers omit both fields.
*/
interface ResolveContentDeps {
loadCapabilities?: (cwd: string) => CapabilityLike[];
resolveTaskContentFn?: typeof resolveTaskContent;
}
// ─── Implementation ───────────────────────────────────────────────────────────
function isBehaviorAddingTaskContent(content: string): BehaviorAddingResult {
@@ -75,10 +109,127 @@ function isBehaviorAddingTaskContent(content: string): BehaviorAddingResult {
};
}
/**
* Default (production) capability loader for `resolve-content`: the merged
* first-party + validated-installed-overlay registry (ADR-1244 D2), the
* established runtime read path for "installed capabilities including
* third-party" — as opposed to `capability-loader.cts`'s heavier build-time
* validation entry points or the static `capability-registry.cjs` alone
* (first-party only, would miss a third-party capability's
* `taskContentResolver` declaration entirely).
*/
function defaultLoadCapabilities(cwd: string): CapabilityLike[] {
const registry = capabilityLoaderMod.loadRegistry({ includeInstalled: true, cwd }) as {
capabilities?: Record<string, CapabilityLike>;
};
return Object.values(registry.capabilities ?? {});
}
function parseResolveContentArgs(args: string[]): { plan: string | null; taskId: string | null } {
let plan: string | null = null;
let taskId: string | null = null;
for (let i = 2; i < args.length; i++) {
if (args[i] === '--plan') {
plan = args[i + 1] ?? null;
i++;
} else if (args[i] === '--task-id') {
taskId = args[i + 1] ?? null;
i++;
}
}
return { plan, taskId };
}
/**
* `task resolve-content --plan <PLAN.md path> --task-id <tracker-id value> --raw`
* (ADR-3646 Decision 2). Resolves one task's content from the external
* tracker its `tracker-id` attribute names, via `task-content-resolution.cts`.
*
* HARD-HALT CONTRACT: a thrown `ResolverAmbiguousError` / `ResolverFailedError`
* / `ResolverTimeoutError` / `ResolverMalformedOutputError` from
* `resolveTaskContent` is turned into this CLI's own non-zero exit via
* `error()` — never swallowed into a `{resolved: false}` JSON answer. Any
* other thrown error is not one of the four documented resolver-error
* classes and is allowed to propagate uncaught.
*/
function routeResolveContent(
{ args, cwd, raw }: RouteTaskCommandOptions,
deps: ResolveContentDeps = {},
): void {
const usage = 'Usage: task resolve-content --plan <path> --task-id <tracker-id> --raw';
const { plan, taskId } = parseResolveContentArgs(args);
if (!plan || !taskId) {
error(usage, ERROR_REASON.USAGE);
return;
}
const projectRoot = path.resolve(cwd || process.cwd());
const resolvedPlanPath = path.resolve(projectRoot, plan);
const rel = path.relative(projectRoot, resolvedPlanPath);
if (rel === '..' || rel.startsWith(`..${path.sep}`)) {
error(`Plan file is outside project scope: ${plan}`, ERROR_REASON.USAGE);
return;
}
if (!fs.existsSync(resolvedPlanPath)) {
error(`Plan file not found: ${plan}`, ERROR_REASON.USAGE);
return;
}
const planContent = fs.readFileSync(resolvedPlanPath, 'utf-8');
const parsedPlan = parsePlanDocument(planContent, resolvedPlanPath) as { tasks?: PlanTaskLike[] };
const task = (parsedPlan.tasks ?? []).find((t) => t.trackerId === taskId);
if (!task) {
error(`No task with tracker-id '${taskId}' found in plan: ${plan}`, ERROR_REASON.USAGE);
return;
}
const loadCapabilities = deps.loadCapabilities ?? defaultLoadCapabilities;
const capabilities = loadCapabilities(projectRoot);
const resolveFn = deps.resolveTaskContentFn ?? resolveTaskContent;
let result;
try {
result = resolveFn({ trackerId: task.trackerId, capabilities });
} catch (err) {
if (
err instanceof ResolverAmbiguousError ||
err instanceof ResolverFailedError ||
err instanceof ResolverTimeoutError ||
err instanceof ResolverMalformedOutputError
) {
error((err as Error).message, ERROR_REASON.UNKNOWN);
return;
}
throw err;
}
switch (result.kind) {
case 'not-applicable':
output({ resolved: false }, raw, undefined);
return;
case 'no-resolver':
output({ resolved: false, reason: 'no-resolver' }, raw, undefined);
return;
case 'empty':
output({ resolved: false, reason: 'empty' }, raw, undefined);
return;
case 'resolved':
output({ resolved: true, content: result.content }, raw, undefined);
return;
}
}
function routeTaskCommand({ args, cwd, raw }: RouteTaskCommandOptions): void {
const subcommand = args[1];
if (subcommand === 'resolve-content') {
routeResolveContent({ args, cwd, raw });
return;
}
if (subcommand !== 'is-behavior-adding') {
error('Unknown task subcommand. Available: is-behavior-adding', ERROR_REASON.SDK_UNKNOWN_COMMAND);
error(
'Unknown task subcommand. Available: is-behavior-adding, resolve-content',
ERROR_REASON.SDK_UNKNOWN_COMMAND,
);
}
let content: string | null = null;
@@ -108,4 +259,5 @@ function routeTaskCommand({ args, cwd, raw }: RouteTaskCommandOptions): void {
export = {
isBehaviorAddingTaskContent,
routeTaskCommand,
routeResolveContent,
};

View File

@@ -0,0 +1,475 @@
/**
* Task Content Resolution Module (ADR-3646 Phase 1, #3970).
*
* Given a task's `tracker-id` attribute value (parsed verbatim by
* `plan-document.cts`, never split there) and the set of installed
* capabilities' `taskContentResolver` declarations, resolves the task's
* content from the matching external tracker via a bounded subprocess call —
* or reports that no resolution applies.
*
* PURE / IMPURE SPLIT, loosely mirroring `review-lane-invocation.cts`'s
* resolve-then-run shape, but deliberately NOT copying its full machinery
* (Gall's Law — see `40-design.md`'s "Laws that apply" section): this problem
* has no probe/model/effort/prompt-channel axes, just one deterministic
* id-lookup. `splitTrackerId`, `findResolver`, and `buildInvocation` are pure
* and total (never throw, even on hostile third-party-shaped input — a
* capability manifest is third-party-authored, and while `capability-
* validator.cjs` validates it at install time, this module re-validates
* defensively rather than trusting that boundary). `resolveTaskContent` is
* the one impure boundary: it spawns exactly one bounded subprocess, through
* an injectable `execFn` so tests never spawn a real process or wait a real
* timeout (CLAUDE.md's clock-seam rule).
*
* HARD-HALT CONTRACT (ADR-3646 Decision 4): an ambiguous resolver match, a
* non-zero resolver exit, a timeout, or malformed resolver stdout all THROW.
* None of these degrade to a silently-empty or silently-picked result — a
* task-content resolution failure must halt the caller (`task-command-
* router.cts`'s `resolve-content` subcommand turns each throw into the CLI's
* own non-zero exit), never fall back to inline PLAN.md content pretending
* nothing happened.
*
* ADR-457 build-at-publish: source in src/task-content-resolution.cts,
* compiled to gsd-core/bin/lib/task-content-resolution.cjs (gitignored).
*/
// Use non-destructured namespace import so test-time mock.method(childProcess,
// 'spawnSync') can intercept calls from this seam — destructured imports
// capture references at load time and become un-mockable (matches the
// convention documented in shell-command-projection.cts).
import childProcess from 'node:child_process';
// eslint-disable-next-line @typescript-eslint/no-require-imports
import ioMod = require('./io.cjs');
const { formatDiagnosticToken } = ioMod;
// ─── Result taxonomy ──────────────────────────────────────────────────────────
/**
* The four non-throwing outcomes of `resolveTaskContent`. Frozen because the
* `kind` discriminant is the product — callers (the CLI seam) switch on it
* directly rather than string-matching prose.
*/
const TASK_CONTENT_RESULT = Object.freeze({
NOT_APPLICABLE: 'not-applicable',
NO_RESOLVER: 'no-resolver',
RESOLVED: 'resolved',
EMPTY: 'empty',
} as const);
type TaskContentResultKind =
(typeof TASK_CONTENT_RESULT)[keyof typeof TASK_CONTENT_RESULT];
interface ResolvedTaskContent {
action: string | null;
verify: string | null;
acceptanceCriteria: string[];
readFirst: string[];
done: string | null;
}
type TaskContentResolution =
| { kind: 'not-applicable' }
| { kind: 'no-resolver' }
| { kind: 'empty' }
| { kind: 'resolved'; content: ResolvedTaskContent };
// ─── Throwable error taxonomy ─────────────────────────────────────────────────
// These four ALWAYS throw — they are configuration/execution defects, never a
// value `resolveTaskContent` returns. See the module docstring's hard-halt
// contract.
/**
* Two or more installed capabilities declare a `taskContentResolver` for the
* same `trackerPrefix`. Structurally impossible in a correctly-validated
* install (`capability-validator.cjs` enforces cross-capability prefix
* uniqueness), but `findResolver` must still refuse to silently pick one if a
* test harness or a validator bug ever produces this shape.
*/
class ResolverAmbiguousError extends Error {
prefix: string;
capabilityIds: string[];
constructor(prefix: string, capabilityIds: string[]) {
super(
`tracker prefix '${prefix}' matches ${capabilityIds.length} installed capability resolvers ` +
`(${capabilityIds.join(', ')}) — ambiguous resolver registration must never silently pick one`,
);
this.name = 'ResolverAmbiguousError';
this.prefix = prefix;
this.capabilityIds = capabilityIds;
}
}
/**
* The resolver subprocess exited non-zero (or failed to spawn at all).
*
* `stderrTail` is UNTRUSTED subprocess-sourced text (the resolver binary is
* declared by a capability manifest and invoked with an argv token derived
* from a PLAN.md `tracker-id` attribute, which is often LLM/agent-authored —
* a hostile or buggy resolver could echo attacker-influenced text back on
* its own stderr). `.message` embeds it through `io.cjs`'s
* `formatDiagnosticToken()` so every caller of `resolveTaskContent` gets a
* `.message` that is already safe to write verbatim to a plain-text
* diagnostic — see that function's docstring for why this must happen here,
* at the point the untrusted substring is embedded, rather than at each
* call site.
*/
class ResolverFailedError extends Error {
exitCode: number | null;
stderrTail: string;
constructor(binary: string, exitCode: number | null, stderrTail: string) {
super(
`resolver command '${binary}' exited ${exitCode === null ? 'with no exit code (spawn failure)' : exitCode}` +
(stderrTail ? `: ${formatDiagnosticToken(stderrTail)}` : ''),
);
this.name = 'ResolverFailedError';
this.exitCode = exitCode;
this.stderrTail = stderrTail;
}
}
/** The resolver subprocess exceeded its declared `invoke.timeoutMs` bound. */
class ResolverTimeoutError extends Error {
timeoutMs: number;
constructor(binary: string, timeoutMs: number) {
super(`resolver command '${binary}' timed out after ${timeoutMs}ms`);
this.name = 'ResolverTimeoutError';
this.timeoutMs = timeoutMs;
}
}
/**
* The resolver's stdout was not valid JSON, or was valid JSON that is not a
* plain object (a `null`, array, string, number, or boolean top-level value
* is rejected — only a plain object can carry the `description`/`verify`/
* `acceptance_criteria`/`read_first`/`done` fields this seam reads).
*/
class ResolverMalformedOutputError extends Error {
stdoutSample: string;
constructor(binary: string, reason: string, stdoutSample: string) {
super(
`resolver command '${binary}' produced malformed output: ${reason}` +
(stdoutSample ? ` (stdout sample: ${formatDiagnosticToken(stdoutSample)})` : ''),
);
this.name = 'ResolverMalformedOutputError';
this.stdoutSample = stdoutSample;
}
}
// ─── Manifest shapes ───────────────────────────────────────────────────────────
interface TaskContentResolverInvoke {
binary: string;
args: string[];
timeoutMs: number;
}
interface TaskContentResolverDeclaration {
capabilityId: string;
trackerPrefix: string;
invoke: TaskContentResolverInvoke;
}
interface CapabilityLike {
id: string;
taskContentResolver?: unknown;
}
// ─── Pure functions ────────────────────────────────────────────────────────────
/**
* The SAME kebab-case grammar `capability-validator.cjs`'s `KEBAB_RE`
* enforces on `taskContentResolver.trackerPrefix` at install time
* (`validateTaskContentResolverFields`). Duplicated as a literal rather than
* imported — `.cts` (build-at-publish, ADR-457) and the hand-written
* `capability-validator.cjs` are genuinely two different build targets with
* no shared-constants module between them today — but `tests/task-content-
* resolver-grammar-parity.test.cjs` asserts both regexes agree on a shared
* table of inputs, so a future edit to either one that silently diverges from
* the other fails a test instead of drifting quietly (CLAUDE.md's Generative
* Fix Divergence rule).
*/
const TRACKER_PREFIX_RE = /^[a-z][a-z0-9-]*$/;
/**
* Split a `tracker-id` attribute value into its prefix and id, on the FIRST
* `:` only — colons after the first stay in the id verbatim (a tracker whose
* native ids contain colons, e.g. `beads:issue:GSD-1` → `{prefix: "beads",
* id: "issue:GSD-1"}`).
*
* PURE. Returns `null` for `null`/empty input, and for a string with no `:`
* at all (nothing to split — there is no prefix to match a resolver against).
*/
function splitTrackerId(trackerId: string | null): { prefix: string; id: string } | null {
if (typeof trackerId !== 'string' || trackerId.length === 0) return null;
const colonIdx = trackerId.indexOf(':');
if (colonIdx === -1) return null;
const prefix = trackerId.slice(0, colonIdx);
const id = trackerId.slice(colonIdx + 1);
if (!prefix || !id) return null;
return { prefix, id };
}
/**
* Defensively re-validate a raw `taskContentResolver` declaration's shape.
* `capability-validator.cjs` already enforces this at install time, but this
* module treats every capability manifest as third-party-authored input and
* never trusts a shape it has not itself checked — a malformed declaration is
* treated as though it does not exist for matching purposes, never thrown on.
*/
function parseResolverDeclaration(
capabilityId: string,
raw: unknown,
): TaskContentResolverDeclaration | null {
if (raw === null || typeof raw !== 'object' || Array.isArray(raw)) return null;
const body = raw as { trackerPrefix?: unknown; invoke?: unknown };
const trackerPrefix = typeof body.trackerPrefix === 'string' ? body.trackerPrefix.trim() : '';
if (!trackerPrefix || !TRACKER_PREFIX_RE.test(trackerPrefix)) return null;
const inv = body.invoke;
if (inv === null || typeof inv !== 'object' || Array.isArray(inv)) return null;
const invBody = inv as { binary?: unknown; args?: unknown; timeoutMs?: unknown };
const binary = typeof invBody.binary === 'string' ? invBody.binary.trim() : '';
if (!binary) return null;
const args = Array.isArray(invBody.args)
? invBody.args.filter((a): a is string => typeof a === 'string')
: null;
if (args === null || args.length !== (invBody.args as unknown[]).length) return null;
if (!args.includes('{{id}}')) return null;
const timeoutMs = invBody.timeoutMs;
if (typeof timeoutMs !== 'number' || !Number.isInteger(timeoutMs) || timeoutMs <= 0) return null;
return {
capabilityId,
trackerPrefix,
invoke: { binary, args, timeoutMs },
};
}
/**
* Find the resolver declared for `prefix` among `capabilities`.
*
* PURE, total. Returns:
* - the single matching declaration when exactly one well-formed resolver
* declares `trackerPrefix === prefix`;
* - `null` when zero capabilities declare a well-formed resolver for it
* (an unrecognized prefix is a data case, not a defect — see design row 11);
* - the literal string `'ambiguous'` when two or more do — this must be
* structurally impossible in a correctly-validated install, but this
* function refuses to silently pick one regardless.
*/
function findResolver(
prefix: string,
capabilities: Array<CapabilityLike>,
): TaskContentResolverDeclaration | 'ambiguous' | null {
const matches: TaskContentResolverDeclaration[] = [];
for (const cap of capabilities ?? []) {
if (!cap || typeof cap !== 'object') continue;
const decl = parseResolverDeclaration(cap.id, cap.taskContentResolver);
if (decl && decl.trackerPrefix === prefix) matches.push(decl);
}
if (matches.length === 0) return null;
if (matches.length > 1) return 'ambiguous';
return matches[0];
}
/**
* Expand a resolver's `invoke.args` template, replacing every `"{{id}}"`
* entry with the literal `id` string. Exact-match token replacement, not
* template-string interpolation — mirrors `review-lane-invocation.cts`'s
* argv-expansion discipline (a placeholder is a whole array element, not a
* substring).
*
* PURE.
*/
function buildInvocation(
resolver: { invoke: TaskContentResolverInvoke },
id: string,
): TaskContentResolverInvoke {
return {
binary: resolver.invoke.binary,
args: resolver.invoke.args.map((a) => (a === '{{id}}' ? id : a)),
timeoutMs: resolver.invoke.timeoutMs,
};
}
// ─── Subprocess boundary ──────────────────────────────────────────────────────
interface ExecResult {
status: number | null;
stdout: string;
stderr: string;
error?: Error;
}
type ExecFn = (binary: string, args: string[], opts: { timeout: number }) => ExecResult;
/**
* Real subprocess execution — the default `execFn`. Uses Node's `spawnSync`
* with the `timeout` option so a hung resolver is killed at the bound rather
* than hanging the caller (CLAUDE.md's Unbounded Subprocesses gauntlet line).
*/
function realExec(binary: string, args: string[], opts: { timeout: number }): ExecResult {
const result = childProcess.spawnSync(binary, args, {
encoding: 'utf-8',
stdio: 'pipe',
timeout: opts.timeout,
windowsHide: true,
});
return {
status: result.status ?? null,
stdout: (result.stdout ?? '').toString(),
stderr: (result.stderr ?? '').toString(),
error: result.error ?? undefined,
};
}
/**
* True when an `ExecResult` indicates the subprocess was killed by the
* `timeout` option, i.e. it never completed and reported a real answer. Only
* `error.code === 'ETIMEDOUT'` is checked — Node.js guarantees this
* cross-platform when `spawnSync`'s `timeout` option fires; pairing it with a
* `signal === 'SIGTERM'` check is platform-fragile (Windows does not
* necessarily report SIGTERM the same way) and risks a false negative. Same
* predicate discipline as `shell-command-projection.cts`'s `isSpawnTimeout`.
*/
function isExecTimeout(result: ExecResult): boolean {
const err: NodeJS.ErrnoException | undefined = result.error;
return err?.code === 'ETIMEDOUT';
}
// ─── Resolver JSON → content mapping ──────────────────────────────────────────
function coerceStringOrNull(value: unknown): string | null {
return typeof value === 'string' ? value : null;
}
function coerceStringArray(value: unknown): string[] {
return Array.isArray(value) ? value.filter((v): v is string => typeof v === 'string') : [];
}
/**
* Map a resolver's validated JSON object onto `ResolvedTaskContent`. Every
* field is coerced defensively — the resolver's JSON is a third-party CLI's
* output, validated for exit code and JSON-object-shape upstream, but never
* trusted field-by-field. A missing or wrong-typed field degrades sanely
* (string/null fields fall back to `null`, array fields fall back to `[]`);
* only the caller's `description`-emptiness check throws no further errors
* here — this function is never the one that decides `resolved` vs `empty`.
*/
function mapResolverOutput(body: Record<string, unknown>): ResolvedTaskContent {
return {
action: coerceStringOrNull(body['description']),
verify: coerceStringOrNull(body['verify']),
acceptanceCriteria: coerceStringArray(body['acceptance_criteria']),
readFirst: coerceStringArray(body['read_first']),
done: coerceStringOrNull(body['done']),
};
}
// ─── Entry point ────────────────────────────────────────────────────────────────
interface ResolveTaskContentInput {
trackerId: string | null;
capabilities: Array<CapabilityLike>;
/** Override the resolver's declared `invoke.timeoutMs`, primarily for tests. */
timeoutOverrideMs?: number;
execFn?: ExecFn;
}
/**
* Orchestrate one task's content resolution: split the `tracker-id`, find the
* matching capability's resolver, invoke it through the bounded subprocess
* boundary, and map its JSON output onto the four documented outcomes.
*
* The only impure boundary is `execFn` (defaults to a real `spawnSync` call).
* Injecting a fake `execFn` lets tests assert every outcome — including a
* timeout — deterministically, without spawning a real process or waiting a
* real `timeoutMs`.
*/
function resolveTaskContent(input: ResolveTaskContentInput): TaskContentResolution {
const split = splitTrackerId(input.trackerId);
if (split === null) return { kind: TASK_CONTENT_RESULT.NOT_APPLICABLE };
const resolver = findResolver(split.prefix, input.capabilities ?? []);
if (resolver === null) return { kind: TASK_CONTENT_RESULT.NO_RESOLVER };
if (resolver === 'ambiguous') {
// findResolver never returns the capability ids for the ambiguous case
// (it discards the losing matches) — re-derive them here for the error.
const ids = (input.capabilities ?? [])
.filter((cap) => {
const decl = parseResolverDeclaration(cap?.id, cap?.taskContentResolver);
return decl !== null && decl.trackerPrefix === split.prefix;
})
.map((cap) => cap.id);
throw new ResolverAmbiguousError(split.prefix, ids);
}
const invocation = buildInvocation(resolver, split.id);
const timeoutMs = input.timeoutOverrideMs ?? invocation.timeoutMs;
const execFn = input.execFn ?? realExec;
const result = execFn(invocation.binary, invocation.args, { timeout: timeoutMs });
if (isExecTimeout(result)) {
throw new ResolverTimeoutError(invocation.binary, timeoutMs);
}
if (result.error || result.status !== 0) {
const stderrTail = (result.stderr ?? '').trim().slice(-2000);
throw new ResolverFailedError(invocation.binary, result.status, stderrTail);
}
let parsed: unknown;
try {
parsed = JSON.parse(result.stdout);
} catch {
throw new ResolverMalformedOutputError(
invocation.binary,
'stdout is not valid JSON',
result.stdout.slice(0, 200),
);
}
if (parsed === null || typeof parsed !== 'object' || Array.isArray(parsed)) {
throw new ResolverMalformedOutputError(
invocation.binary,
`stdout parsed as valid JSON but is not a plain object (got ${Array.isArray(parsed) ? 'array' : typeof parsed})`,
result.stdout.slice(0, 200),
);
}
const content = mapResolverOutput(parsed as Record<string, unknown>);
const description = typeof content.action === 'string' ? content.action.trim() : '';
if (description.length === 0) {
return { kind: TASK_CONTENT_RESULT.EMPTY };
}
return { kind: TASK_CONTENT_RESULT.RESOLVED, content };
}
const taskContentResolution = {
TASK_CONTENT_RESULT,
splitTrackerId,
findResolver,
buildInvocation,
resolveTaskContent,
ResolverAmbiguousError,
ResolverFailedError,
ResolverTimeoutError,
ResolverMalformedOutputError,
};
// eslint-disable-next-line @typescript-eslint/no-namespace
declare namespace taskContentResolution {
export {
TaskContentResultKind,
ResolvedTaskContent,
TaskContentResolution,
TaskContentResolverInvoke,
TaskContentResolverDeclaration,
CapabilityLike,
ExecResult,
ExecFn,
ResolveTaskContentInput,
};
}
export = taskContentResolution;

View File

@@ -0,0 +1,316 @@
// docs-guard-exempt: no docs/ file reads in this test.
'use strict';
process.env.GSD_TEST_MODE = '1';
/**
* capability-validator-task-content-resolver.test.cjs — behavioral tests for
* the OPTIONAL `taskContentResolver` body on `role: "feature"` capability
* manifests (ADR-3646, #3970).
*
* Implements test-matrix rows 20–24 of
* `.gsd/phase/feat-3970-task-content-resolution-seam/50-test-matrix.md`.
* See `docs/adr/3646-per-task-content-resolution-seam.md` Decision 3 for the
* shape this validates.
*/
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const {
validateCapability,
validateTaskContentResolver,
validateCrossCapability,
} = require('../gsd-core/bin/lib/capability-validator.cjs');
// ─── Fixture builders ──────────────────────────────────────────────────────
// House convention (tests/capability-manifest-version.test.cjs): builder
// functions return a VALID fixture, which each test then mutates. Every call
// returns a FRESH object — no shared mutable state, no execution-order
// dependence.
function validResolver() {
return {
trackerPrefix: 'beads',
invoke: {
binary: 'bd',
args: ['show', '{{id}}', '--json'],
timeoutMs: 10000,
},
};
}
function featureCap(overrides) {
return {
id: 'demo',
role: 'feature',
version: '1.2.3',
title: 'Demo',
description: 'A demo capability.',
tier: 'standard',
requires: [],
engines: { gsd: '>=1.6.0' },
runtimeCompat: { supported: ['*'], unsupported: [] },
skills: [],
agents: [],
hooks: [],
config: {},
steps: [],
contributions: [],
gates: [],
...overrides,
};
}
function runtimeCap(overrides) {
return {
id: 'demo-rt',
role: 'runtime',
version: '1.2.3',
title: 'Demo RT',
description: 'A demo runtime.',
tier: 'standard',
requires: [],
engines: { gsd: '>=1.6.0' },
runtime: {
configHome: { kind: 'dot-home', name: '.demo', env: [] },
localConfigDir: '.demo',
configFormat: 'settings-json',
artifactLayout: { global: [], local: [] },
commandStyle: 'slash-hyphen',
hooksSurface: 'settings-json',
sandboxTier: 'none',
supportTier: 2,
installSurface: 'settings-json',
writesSharedSettings: false,
permissionWriter: null,
extendedHookEvents: [],
hostIntegration: {
embeddingMode: 'imperative',
commandSurface: 'slash-file',
dispatch: { namedDispatch: true, nested: true, maxDepth: -1, background: true, subagentToolkit: 'full', backgroundDispatch: false },
modelMode: 'passive',
hookBus: 'host',
stateIO: 'filesystem',
effortSurface: 'none',
isolation: 'process',
},
},
...overrides,
};
}
function reviewerCap(overrides) {
return {
id: 'demo-reviewer',
role: 'reviewer',
version: '1.2.3',
title: 'Demo Reviewer',
description: 'A demo reviewer lane.',
tier: 'standard',
requires: [],
engines: { gsd: '>=1.6.0' },
reviewer: {
slug: 'demo-reviewer',
flags: ['--demo-reviewer'],
transport: 'spawn',
probe: { kind: 'command-exists', binary: 'demo-reviewer' },
invoke: {
binary: 'demo-reviewer',
args: [],
promptChannel: 'stdin',
outputChannel: 'stdout',
modelArg: null,
effortChannel: 'none',
},
timeoutFloorMs: 5000,
emptyOutput: 'stub-with-stderr',
reviewsSection: 'Demo Reviewer',
evidenceClass: 'source-grounded',
requiresBinaries: [],
promptBudgetKey: null,
handler: null,
},
...overrides,
};
}
// ─── Row 20: valid taskContentResolver on a feature manifest ───────────────
describe('row 20 — valid taskContentResolver body', () => {
test('valid taskContentResolver body passes', () => {
const cap = featureCap({ taskContentResolver: validResolver() });
const errs = validateCapability(cap, cap.id);
assert.deepEqual(errs, [], `expected no errors, got: ${JSON.stringify(errs)}`);
});
test('omitting taskContentResolver entirely on an otherwise-valid feature manifest yields zero errors', () => {
const cap = featureCap();
const errs = validateCapability(cap, cap.id);
assert.deepEqual(errs, [], `expected no errors, got: ${JSON.stringify(errs)}`);
});
});
// ─── Row 21: feature-only field ────────────────────────────────────────────
describe('row 21 — taskContentResolver on non-feature role is rejected', () => {
test('role:runtime declaring taskContentResolver is rejected', () => {
const cap = runtimeCap({ taskContentResolver: validResolver() });
const errs = validateCapability(cap, cap.id);
assert.ok(
errs.some((e) => e.includes('taskContentResolver') && e.includes('feature-only')),
`expected a feature-only rejection, got: ${JSON.stringify(errs)}`,
);
});
test('role:reviewer declaring taskContentResolver is rejected', () => {
const cap = reviewerCap({ taskContentResolver: validResolver() });
const errs = validateCapability(cap, cap.id);
assert.ok(
errs.some((e) => e.includes('taskContentResolver') && e.includes('feature-only')),
`expected a feature-only rejection, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── Row 22: malformed trackerPrefix grammar ───────────────────────────────
describe('row 22 — malformed trackerPrefix is rejected', () => {
test('capital-cased trackerPrefix violates KEBAB_RE', () => {
const cap = featureCap({
taskContentResolver: { ...validResolver(), trackerPrefix: 'Beads' },
});
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('trackerPrefix') && e.includes('kebab-case')),
`expected a grammar error naming trackerPrefix, got: ${JSON.stringify(errs)}`,
);
});
test('empty-string trackerPrefix is rejected', () => {
const cap = featureCap({
taskContentResolver: { ...validResolver(), trackerPrefix: '' },
});
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('trackerPrefix')),
`expected a trackerPrefix error, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── Row 23: invoke.timeoutMs boundary ─────────────────────────────────────
describe('row 23 — invoke.timeoutMs must be a positive integer', () => {
for (const bad of [0, -1, 1.5, undefined]) {
test(`timeoutMs ${JSON.stringify(bad)} is rejected`, () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, timeoutMs: bad };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('timeoutMs')),
`expected a timeoutMs error for ${JSON.stringify(bad)}, got: ${JSON.stringify(errs)}`,
);
});
}
test('a legitimate positive integer timeoutMs (10000) is accepted', () => {
const cap = featureCap({ taskContentResolver: validResolver() });
const errs = validateTaskContentResolver(cap);
assert.ok(
!errs.some((e) => e.includes('timeoutMs')),
`expected no timeoutMs error, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── invoke.timeoutMs upper ceiling (security review finding, #3970) ───────
// A manifest declaring an unbounded-in-practice timeoutMs (e.g.
// Number.MAX_SAFE_INTEGER) would let resolve-content hang near-indefinitely
// on a stuck/malicious resolver, defeating the "bounded subprocess" design
// intent. Boundary-inclusive per CLAUDE.md's limit-1/limit/limit+1 rule.
describe('invoke.timeoutMs upper ceiling (120000ms)', () => {
test('timeoutMs 120000 (exactly at the ceiling) is accepted', () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, timeoutMs: 120000 };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
!errs.some((e) => e.includes('timeoutMs')),
`expected no timeoutMs error at the boundary, got: ${JSON.stringify(errs)}`,
);
});
test('timeoutMs 120001 (one past the ceiling) is rejected', () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, timeoutMs: 120001 };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('timeoutMs') && e.includes('120000')),
`expected a ceiling timeoutMs error, got: ${JSON.stringify(errs)}`,
);
});
test('an unbounded-in-practice timeoutMs (Number.MAX_SAFE_INTEGER) is rejected', () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, timeoutMs: Number.MAX_SAFE_INTEGER };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('timeoutMs')),
`expected a timeoutMs error, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── invoke.args must carry the {{id}} placeholder ─────────────────────────
describe('invoke.args without {{id}} placeholder', () => {
test('args missing the {{id}} placeholder is rejected', () => {
const resolver = validResolver();
resolver.invoke = { ...resolver.invoke, args: ['show', 'GSD-42', '--json'] };
const cap = featureCap({ taskContentResolver: resolver });
const errs = validateTaskContentResolver(cap);
assert.ok(
errs.some((e) => e.includes('{{id}}')),
`expected a placeholder error, got: ${JSON.stringify(errs)}`,
);
});
});
// ─── Row 24: cross-capability trackerPrefix uniqueness ─────────────────────
describe('row 24 — duplicate trackerPrefix across capabilities fails cross-capability validation', () => {
test('two manifests declaring the same trackerPrefix collide', () => {
const capA = featureCap({ id: 'resolver-a', taskContentResolver: validResolver() });
const capB = featureCap({ id: 'resolver-b', taskContentResolver: validResolver() });
const capMap = new Map([
[capA.id, capA],
[capB.id, capB],
]);
const errs = validateCrossCapability(capMap, new Set());
assert.ok(
errs.some((e) => e.includes('trackerPrefix') && e.includes('resolver-a') && e.includes('resolver-b')),
`expected a collision error naming both ids, got: ${JSON.stringify(errs)}`,
);
});
test('two manifests declaring different trackerPrefix values do not collide', () => {
const capA = featureCap({ id: 'resolver-a', taskContentResolver: validResolver() });
const capB = featureCap({
id: 'resolver-b',
taskContentResolver: { ...validResolver(), trackerPrefix: 'linear' },
});
const capMap = new Map([
[capA.id, capA],
[capB.id, capB],
]);
const errs = validateCrossCapability(capMap, new Set());
assert.ok(
!errs.some((e) => e.includes('trackerPrefix')),
`expected no trackerPrefix collision, got: ${JSON.stringify(errs)}`,
);
});
});

View File

@@ -107,7 +107,7 @@ const PROSE_ALLOWLIST = [
{ file: 'agents/gsd-phase-researcher.md', line: 33, reason: 'package-legitimacy provenance rule names the command as the source of an OK verdict; descriptive' },
{ file: 'agents/gsd-roadmapper.md', line: 642, reason: 'parenthetical "e.g." naming SDK queries a user *could* run; not an agent instruction' },
{ file: 'agents/gsd-intel-updater.md', line: 40, reason: 'cross-platform note names the `gsd-tools intel <subcommand>` CLI surface descriptively ("CLI invocations go through..."); not an agent instruction' },
{ file: 'gsd-core/workflows/execute-plan.md', line: 415, reason: 'describes the downstream SDK validation step (`validated downstream by ...`); names the mechanism, does not instruct the agent to type it' },
{ file: 'gsd-core/workflows/execute-plan.md', line: 416, reason: 'describes the downstream SDK validation step (`validated downstream by ...`); names the mechanism, does not instruct the agent to type it' },
];
// Resolver-snippet definition lines / probes that must never be flagged. A line

View File

@@ -0,0 +1,108 @@
'use strict';
/**
* Unit tests for plan-document.cjs
*
* Module: gsd-core/bin/lib/plan-document.cjs
*
* Covers the `tracker-id` attribute (ADR-3646 Phase 1, #3970) added to the
* `<task>` element grammar, plus regression coverage proving the addition
* does not alter pre-existing task-parsing behaviour.
*
* Matrix rows referenced below are from
* .gsd/phase/feat-3970-task-content-resolution-seam/50-test-matrix.md
*/
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const { parsePlanDocument } = require('../gsd-core/bin/lib/plan-document.cjs');
describe('plan-document: tracker-id attribute', () => {
test('row 1 — no tracker-id attribute yields trackerId: null', () => {
const doc = parsePlanDocument(`
<task type="auto">
<name>Do a thing</name>
</task>
`);
assert.equal(doc.tasks.length, 1);
assert.equal(doc.tasks[0].trackerId, null);
});
test('row 2 — tracker-id is read verbatim, never split', () => {
const doc = parsePlanDocument(`
<task type="auto" tracker-id="beads:GSD-42">
<name>Do a thing</name>
</task>
`);
assert.equal(doc.tasks.length, 1);
assert.equal(doc.tasks[0].trackerId, 'beads:GSD-42');
});
test('row 3 — tracker-id="" (empty string) normalises to null', () => {
const doc = parsePlanDocument(`
<task type="auto" tracker-id="">
<name>Do a thing</name>
</task>
`);
assert.equal(doc.tasks.length, 1);
assert.equal(doc.tasks[0].trackerId, null);
});
test('row 4 — checkpoint tasks never read tracker-id, even when present', () => {
const doc = parsePlanDocument(`
<task type="checkpoint:decision" tracker-id="beads:GSD-99">
<decision>Ship it</decision>
</task>
`);
assert.equal(doc.tasks.length, 1);
assert.equal(doc.tasks[0].kind, 'checkpoint');
assert.equal(doc.tasks[0].trackerId, null);
});
});
describe('plan-document: regression — legacy behaviour unchanged', () => {
test('legacy `## Task N` markdown fallback still parses with trackerId: null', () => {
const doc = parsePlanDocument(`
## Task 1: Do a thing
Some body text.
## Task 2: Do another thing
`);
assert.equal(doc.tasks.length, 2);
for (const t of doc.tasks) {
assert.equal(t.kind, 'auto');
assert.equal(t.type, null);
assert.equal(t.trackerId, null);
assert.deepEqual(t.plannedFiles, []);
assert.deepEqual(t.acceptanceCriteria, []);
assert.equal(t.done, null);
}
assert.equal(doc.tasks[0].name, 'Task 1: Do a thing');
assert.equal(doc.tasks[1].name, 'Task 2: Do another thing');
});
test('ordinary task with name/files/acceptance_criteria still parses correctly alongside trackerId', () => {
const doc = parsePlanDocument(`
<task type="auto" tracker-id="beads:GSD-7">
<name>Implement the seam</name>
<files>src/a.cts, src/b.cts</files>
<acceptance_criteria>
- criterion one
- criterion two
</acceptance_criteria>
<done>Merged.</done>
</task>
`);
assert.equal(doc.tasks.length, 1);
const t = doc.tasks[0];
assert.equal(t.kind, 'auto');
assert.equal(t.type, 'auto');
assert.equal(t.name, 'Implement the seam');
assert.deepEqual(t.plannedFiles, ['src/a.cts', 'src/b.cts']);
assert.deepEqual(t.acceptanceCriteria, ['criterion one', 'criterion two']);
assert.equal(t.done, 'Merged.');
assert.equal(t.trackerId, 'beads:GSD-7');
});
});

View File

@@ -0,0 +1,279 @@
'use strict';
/**
* Tests for `task resolve-content` (ADR-3646 Decision 2, issue #3970).
* Covers test matrix rows 17-19:
* 17 — plan exists, task-id not found in it -> USAGE, non-zero exit.
* 18 — end-to-end happy path via injected `resolveTaskContentFn`.
* 19 — missing --plan and/or --task-id -> USAGE, before any filesystem access.
* Plus the hard-halt path: a thrown resolver-error class must surface as a
* non-zero CLI exit, never a swallowed `{resolved:false}` JSON answer.
*
* In-process style (routeResolveContent's own `_`-prefixed-equivalent
* `deps` injection seam), mirroring `refactor-trigger-command-router.cts`'s
* injection convention — no subprocess spawn needed for these rows.
*/
const { test, describe, mock } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('node:fs');
const path = require('node:path');
const { createTempDir, cleanup } = require('./helpers.cjs');
const { routeResolveContent } = require('../gsd-core/bin/lib/task-command-router.cjs');
const { ExitError } = require('../gsd-core/bin/lib/cli-exit.cjs');
const {
ResolverFailedError,
} = require('../gsd-core/bin/lib/task-content-resolution.cjs');
function writePlan(dir, taskXml) {
const planPath = path.join(dir, '01-PLAN.md');
fs.writeFileSync(planPath, `# Plan\n\n${taskXml}\n`, 'utf8');
return planPath;
}
/**
* `io.cjs`'s `output()` writes to fd 1 via `fs.writeSync` directly (never
* `process.stdout.write`) — see `writeAllSync` in
* `gsd-core/bin/lib/io.cjs`. Capture fd-1 writes by mocking `fs.writeSync`
* itself; every other fd passes through to the real implementation
* unmocked (so the plan-file reads and any other fs traffic inside a test
* body still work).
*/
function captureStdout(fn) {
const chunks = [];
const origWriteSync = fs.writeSync.bind(fs);
const writeMock = mock.method(fs, 'writeSync', (fd, buffer, ...rest) => {
if (fd === 1) {
chunks.push(Buffer.isBuffer(buffer) ? buffer.toString('utf8') : String(buffer));
return Buffer.isBuffer(buffer) ? buffer.length : Buffer.byteLength(String(buffer));
}
return origWriteSync(fd, buffer, ...rest);
});
try {
fn();
} finally {
writeMock.mock.restore();
}
return chunks.join('');
}
/**
* `io.cjs`'s `error()` (ADR-3889) writes its human-readable message to fd 2
* via `writeAllSync`, then throws a bare `new ExitError(1)` with NO message
* — the stderr write already happened, so the thrown Error's own `.message`
* defaults to `"process exit ${code}"` (see `cli-exit.cjs`'s `ExitError`
* constructor) and never carries the diagnostic text. Asserting against
* `err.message` therefore can never see the "outside project scope" text;
* the diagnostic must be read off the captured fd-2 bytes instead. Mirrors
* the same fd-mock idiom `captureStdout` above uses for fd 1, and the
* established repo pattern in `tests/estimate-calibrate.test.cjs`'s
* `runCalibrateExpectError`.
*/
function captureStderr(fn) {
const chunks = [];
const origWriteSync = fs.writeSync.bind(fs);
const writeMock = mock.method(fs, 'writeSync', (fd, buffer, ...rest) => {
if (fd === 2) {
chunks.push(Buffer.isBuffer(buffer) ? buffer.toString('utf8') : String(buffer));
return Buffer.isBuffer(buffer) ? buffer.length : Buffer.byteLength(String(buffer));
}
return origWriteSync(fd, buffer, ...rest);
});
try {
fn();
} finally {
writeMock.mock.restore();
}
return chunks.join('');
}
const RESOLVABLE_TASK = '<task type="auto" tracker-id="test:1"><name>x</name><action>do the thing</action></task>';
describe('task resolve-content (rows 17-19)', () => {
test('row 19: missing --plan and --task-id -> USAGE, no filesystem access', (t) => {
const dir = createTempDir('gsd-resolve-content-19-');
t.after(() => cleanup(dir));
const readMock = require('node:test').mock.method(fs, 'readFileSync', () => {
throw new Error('must not read the filesystem before usage validation');
});
t.after(() => readMock.mock.restore());
assert.throws(
() => routeResolveContent({ args: ['task', 'resolve-content'], cwd: dir, raw: false }),
ExitError,
);
});
test('row 19: missing --task-id only -> USAGE', (t) => {
const dir = createTempDir('gsd-resolve-content-19b-');
t.after(() => cleanup(dir));
writePlan(dir, RESOLVABLE_TASK);
assert.throws(
() =>
routeResolveContent({
args: ['task', 'resolve-content', '--plan', '01-PLAN.md'],
cwd: dir,
raw: false,
}),
ExitError,
);
});
test('row 17: plan exists, task-id not found -> USAGE, non-zero exit', (t) => {
const dir = createTempDir('gsd-resolve-content-17-');
t.after(() => cleanup(dir));
writePlan(dir, RESOLVABLE_TASK);
assert.throws(
() =>
routeResolveContent({
args: ['task', 'resolve-content', '--plan', '01-PLAN.md', '--task-id', 'nope:999'],
cwd: dir,
raw: false,
}),
(err) => {
assert.ok(err instanceof ExitError, `expected ExitError, got ${err}`);
assert.strictEqual(err.code, 1);
return true;
},
);
});
test('row 18: end-to-end happy path, resolved:true content shape', (t) => {
const dir = createTempDir('gsd-resolve-content-18-');
t.after(() => cleanup(dir));
writePlan(dir, RESOLVABLE_TASK);
const fakeCapabilities = [
{
id: 'fake-tracker',
taskContentResolver: {
trackerPrefix: 'test',
invoke: { binary: 'fake-cli', args: ['show', '{{id}}'], timeoutMs: 5000 },
},
},
];
let capturedInvocation = null;
const resolveTaskContentFn = (input) => {
capturedInvocation = input;
return {
kind: 'resolved',
content: {
action: 'do the resolved thing',
verify: 'run the resolved verify',
acceptanceCriteria: ['criterion a'],
readFirst: ['README.md'],
done: 'done marker',
},
};
};
const stdout = captureStdout(() => {
routeResolveContent(
{
args: ['task', 'resolve-content', '--plan', '01-PLAN.md', '--task-id', 'test:1'],
cwd: dir,
raw: true,
},
{ loadCapabilities: () => fakeCapabilities, resolveTaskContentFn },
);
});
assert.strictEqual(capturedInvocation.trackerId, 'test:1');
assert.deepStrictEqual(capturedInvocation.capabilities, fakeCapabilities);
const printed = JSON.parse(stdout);
assert.strictEqual(printed.resolved, true);
assert.strictEqual(printed.content.action, 'do the resolved thing');
assert.strictEqual(printed.content.verify, 'run the resolved verify');
assert.deepStrictEqual(printed.content.acceptanceCriteria, ['criterion a']);
assert.deepStrictEqual(printed.content.readFirst, ['README.md']);
assert.strictEqual(printed.content.done, 'done marker');
});
test('hard-halt: a thrown ResolverFailedError becomes a non-zero exit, never {resolved:false}', (t) => {
const dir = createTempDir('gsd-resolve-content-hardhalt-');
t.after(() => cleanup(dir));
writePlan(dir, RESOLVABLE_TASK);
const fakeCapabilities = [
{
id: 'fake-tracker',
taskContentResolver: {
trackerPrefix: 'test',
invoke: { binary: 'fake-cli', args: ['show', '{{id}}'], timeoutMs: 5000 },
},
},
];
let thrown = null;
const stdout = captureStdout(() => {
try {
routeResolveContent(
{
args: ['task', 'resolve-content', '--plan', '01-PLAN.md', '--task-id', 'test:1'],
cwd: dir,
raw: true,
},
{
loadCapabilities: () => fakeCapabilities,
resolveTaskContentFn: () => {
throw new ResolverFailedError('fake-cli', 1, 'boom');
},
},
);
} catch (err) {
thrown = err;
}
});
assert.ok(thrown instanceof ExitError, `expected ExitError, got ${thrown}`);
assert.strictEqual(stdout, '', 'resolver failure must never write a JSON {resolved:false} answer to stdout');
});
test('path traversal: --plan escaping the project root -> USAGE, non-zero exit, no filesystem read', (t) => {
const dir = createTempDir('gsd-resolve-content-traversal-');
t.after(() => cleanup(dir));
const escapingPath = '../../../etc/passwit';
let thrown = null;
const stderr = captureStderr(() => {
try {
routeResolveContent({
args: ['task', 'resolve-content', '--plan', escapingPath, '--task-id', 'test:1'],
cwd: dir,
raw: false,
});
} catch (err) {
thrown = err;
}
});
assert.ok(thrown instanceof ExitError, `expected ExitError, got ${thrown}`);
assert.strictEqual(thrown.code, 1);
assert.ok(
stderr.includes('outside project scope') && stderr.includes(escapingPath),
`expected an outside-project-scope rejection naming the path on stderr, got: ${stderr}`,
);
});
test('not-applicable: task-id with no ":" -> {resolved:false}, no reason field', (t) => {
const dir = createTempDir('gsd-resolve-content-na-');
t.after(() => cleanup(dir));
writePlan(dir, '<task type="auto" tracker-id="notrackerprefix"><name>y</name></task>');
const stdout = captureStdout(() => {
routeResolveContent(
{
args: ['task', 'resolve-content', '--plan', '01-PLAN.md', '--task-id', 'notrackerprefix'],
cwd: dir,
raw: true,
},
{ loadCapabilities: () => [] },
);
});
const printed = JSON.parse(stdout);
assert.deepStrictEqual(printed, { resolved: false });
});
});

View File

@@ -0,0 +1,343 @@
'use strict';
// Task Content Resolution Module tests (ADR-3646 Phase 1, #3970).
// Covers test matrix rows 5-16 (.gsd/phase/feat-3970-task-content-resolution-seam/50-test-matrix.md).
// Every subprocess call is a fake execFn — never a real spawn, never a real
// timeout wait (CLAUDE.md's clock-seam rule).
const { test } = require('node:test');
const assert = require('node:assert');
const fc = require('fast-check');
const m = require('../gsd-core/bin/lib/task-content-resolution.cjs');
const {
TASK_CONTENT_RESULT,
splitTrackerId,
findResolver,
buildInvocation,
resolveTaskContent,
ResolverAmbiguousError,
ResolverFailedError,
ResolverTimeoutError,
ResolverMalformedOutputError,
} = m;
function beadsCapability(overrides = {}) {
return {
id: 'beads-capability',
taskContentResolver: {
trackerPrefix: 'beads',
invoke: {
binary: 'bd',
args: ['show', '{{id}}', '--json'],
timeoutMs: 10000,
},
...overrides,
},
};
}
function timeoutError() {
const e = new Error('spawnSync bd ETIMEDOUT');
e.code = 'ETIMEDOUT';
return e;
}
// ─── splitTrackerId — pure ─────────────────────────────────────────────────────
test('splitTrackerId returns null for null and empty input', () => {
assert.strictEqual(splitTrackerId(null), null);
assert.strictEqual(splitTrackerId(''), null);
});
test('splitTrackerId splits on the first colon', () => {
assert.deepStrictEqual(splitTrackerId('beads:GSD-42'), { prefix: 'beads', id: 'GSD-42' });
});
test('splitTrackerId with no colon returns null (row 16 boundary partner)', () => {
assert.strictEqual(splitTrackerId('noprefix'), null);
});
test('id containing colons splits on first colon only (row 16)', () => {
assert.deepStrictEqual(
splitTrackerId('beads:issue:GSD-1'),
{ prefix: 'beads', id: 'issue:GSD-1' },
);
});
// ─── buildInvocation — pure ─────────────────────────────────────────────────────
test('buildInvocation replaces every {{id}} entry with the literal id', () => {
const resolver = { invoke: { binary: 'bd', args: ['show', '{{id}}', '--json'], timeoutMs: 10000 } };
assert.deepStrictEqual(buildInvocation(resolver, 'GSD-42'), {
binary: 'bd',
args: ['show', 'GSD-42', '--json'],
timeoutMs: 10000,
});
});
test('buildInvocation does not touch args that merely contain {{id}} as a substring', () => {
const resolver = { invoke: { binary: 'bd', args: ['prefix-{{id}}-suffix'], timeoutMs: 1000 } };
assert.deepStrictEqual(buildInvocation(resolver, 'X').args, ['prefix-{{id}}-suffix']);
});
// ─── findResolver — pure ─────────────────────────────────────────────────────
test('findResolver returns null when zero capabilities match the prefix', () => {
assert.strictEqual(findResolver('beads', []), null);
assert.strictEqual(findResolver('beads', [{ id: 'other', taskContentResolver: undefined }]), null);
});
test('findResolver returns the single well-formed match', () => {
const cap = beadsCapability();
const result = findResolver('beads', [cap]);
assert.strictEqual(result.capabilityId, 'beads-capability');
assert.strictEqual(result.trackerPrefix, 'beads');
});
test('findResolver returns "ambiguous" for two matching capabilities (row 7)', () => {
const capA = beadsCapability();
const capB = { ...beadsCapability(), id: 'other-capability' };
assert.strictEqual(findResolver('beads', [capA, capB]), 'ambiguous');
});
test('findResolver ignores a declaration missing the {{id}} placeholder (row 15)', () => {
const bad = {
id: 'bad-capability',
taskContentResolver: { trackerPrefix: 'beads', invoke: { binary: 'bd', args: ['show'], timeoutMs: 1000 } },
};
assert.strictEqual(findResolver('beads', [bad]), null);
});
test('findResolver ignores garbage taskContentResolver shapes without throwing', () => {
const garbage = [
{ id: 'a', taskContentResolver: null },
{ id: 'b', taskContentResolver: 'not-an-object' },
{ id: 'c', taskContentResolver: [] },
{ id: 'd', taskContentResolver: { trackerPrefix: 'beads' } }, // no invoke
{ id: 'e', taskContentResolver: { trackerPrefix: 'beads', invoke: { binary: '', args: ['{{id}}'], timeoutMs: 1000 } } },
{ id: 'f', taskContentResolver: { trackerPrefix: 'beads', invoke: { binary: 'bd', args: ['{{id}}'], timeoutMs: 0 } } },
{ id: 'g', taskContentResolver: { trackerPrefix: 'beads', invoke: { binary: 'bd', args: ['{{id}}'], timeoutMs: -5 } } },
{ id: 'h', taskContentResolver: { trackerPrefix: 'beads', invoke: { binary: 'bd', args: ['{{id}}'], timeoutMs: 1.5 } } },
];
assert.strictEqual(findResolver('beads', garbage), null);
});
// ─── resolveTaskContent — orchestration ─────────────────────────────────────────
test('row 5: no tracker-id resolves not-applicable', () => {
const result = resolveTaskContent({ trackerId: null, capabilities: [beadsCapability()] });
assert.deepStrictEqual(result, { kind: TASK_CONTENT_RESULT.NOT_APPLICABLE });
});
test('row 6: unmatched prefix resolves no-resolver', () => {
const result = resolveTaskContent({ trackerId: 'unknownprefix:1', capabilities: [] });
assert.deepStrictEqual(result, { kind: TASK_CONTENT_RESULT.NO_RESOLVER });
});
test('row 7: ambiguous prefix registration throws, never silently picks one', () => {
const capA = beadsCapability();
const capB = { ...beadsCapability(), id: 'other-capability' };
assert.throws(
() => resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [capA, capB], execFn: () => { throw new Error('must not be called'); } }),
ResolverAmbiguousError,
);
});
test('row 8: non-empty description resolves true with mapped content', () => {
const execFn = () => ({
status: 0,
stdout: JSON.stringify({
description: 'do X',
verify: 'run tests',
acceptance_criteria: ['a', 'b'],
read_first: ['docs/x.md'],
done: 'X is done',
}),
stderr: '',
});
const result = resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn });
assert.deepStrictEqual(result, {
kind: TASK_CONTENT_RESULT.RESOLVED,
content: {
action: 'do X',
verify: 'run tests',
acceptanceCriteria: ['a', 'b'],
readFirst: ['docs/x.md'],
done: 'X is done',
},
});
});
test('row 9: empty description string resolves empty', () => {
const execFn = () => ({ status: 0, stdout: JSON.stringify({ description: '' }), stderr: '' });
const result = resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn });
assert.deepStrictEqual(result, { kind: TASK_CONTENT_RESULT.EMPTY });
});
test('row 9b: whitespace-only description resolves empty', () => {
const execFn = () => ({ status: 0, stdout: JSON.stringify({ description: ' ' }), stderr: '' });
const result = resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn });
assert.deepStrictEqual(result, { kind: TASK_CONTENT_RESULT.EMPTY });
});
test('row 10: absent description resolves empty, same as empty string', () => {
const execFn = () => ({ status: 0, stdout: JSON.stringify({}), stderr: '' });
const result = resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn });
assert.deepStrictEqual(result, { kind: TASK_CONTENT_RESULT.EMPTY });
});
test('row 11: non-zero resolver exit throws ResolverFailedError, never falls back silently', () => {
const execFn = () => ({ status: 1, stdout: '', stderr: 'no such issue GSD-42' });
assert.throws(
() => resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn }),
(err) => {
assert.ok(err instanceof ResolverFailedError);
assert.strictEqual(err.exitCode, 1);
assert.match(err.stderrTail, /no such issue/);
return true;
},
);
});
test('row 11b: a spawn error (e.g. ENOENT) also throws ResolverFailedError', () => {
const enoent = new Error('spawnSync bd ENOENT');
enoent.code = 'ENOENT';
const execFn = () => ({ status: null, stdout: '', stderr: '', error: enoent });
assert.throws(
() => resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn }),
ResolverFailedError,
);
});
test('row 12: resolver exceeding timeoutMs throws ResolverTimeoutError (simulated, no real wait)', () => {
const execFn = () => ({ status: null, stdout: '', stderr: '', error: timeoutError() });
assert.throws(
() => resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn }),
(err) => {
assert.ok(err instanceof ResolverTimeoutError);
assert.strictEqual(err.timeoutMs, 10000);
return true;
},
);
});
test('row 13: malformed JSON stdout throws, is not conflated with empty content', () => {
const execFn = () => ({ status: 0, stdout: 'not json', stderr: '' });
assert.throws(
() => resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn }),
ResolverMalformedOutputError,
);
});
test('row 14: valid JSON non-object stdout (null/array/string/number/bool) throws', () => {
const cases = [null, [], 'x', 0, true];
for (const value of cases) {
const execFn = () => ({ status: 0, stdout: JSON.stringify(value), stderr: '' });
assert.throws(
() => resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn }),
ResolverMalformedOutputError,
`expected throw for JSON value ${JSON.stringify(value)}`,
);
}
});
test('row 15: invoke.args without {{id}} placeholder is rejected — resolves no-resolver, never spawns', () => {
const badCapability = {
id: 'bad-capability',
taskContentResolver: { trackerPrefix: 'beads', invoke: { binary: 'bd', args: ['show'], timeoutMs: 1000 } },
};
const execFn = () => { throw new Error('must not spawn a rejected declaration'); };
const result = resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [badCapability], execFn });
assert.deepStrictEqual(result, { kind: TASK_CONTENT_RESULT.NO_RESOLVER });
});
test('row 16: id containing colons splits on first colon only, end to end', () => {
let capturedArgs = null;
const execFn = (binary, args) => {
capturedArgs = args;
return { status: 0, stdout: JSON.stringify({ description: 'do X' }), stderr: '' };
};
const result = resolveTaskContent({ trackerId: 'beads:issue:GSD-1', capabilities: [beadsCapability()], execFn });
assert.deepStrictEqual(capturedArgs, ['show', 'issue:GSD-1', '--json']);
assert.strictEqual(result.kind, TASK_CONTENT_RESULT.RESOLVED);
});
// ─── stderr-on-success is not an error (design.md negative space) ──────────────
test('stderr output on a successful (exit 0) run is not treated as a failure', () => {
const execFn = () => ({ status: 0, stdout: JSON.stringify({ description: 'do X' }), stderr: 'informational log line' });
const result = resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn });
assert.strictEqual(result.kind, TASK_CONTENT_RESULT.RESOLVED);
});
// ─── fast-check property test (row 14, gauntlet enumeration) ───────────────────
test('property: any valid-JSON-but-not-an-object stdout always throws ResolverMalformedOutputError, never resolves silently', () => {
fc.assert(
fc.property(
fc.oneof(
fc.constant(null),
fc.array(fc.anything()),
fc.string(),
fc.double(),
fc.boolean(),
),
(value) => {
const execFn = () => ({ status: 0, stdout: JSON.stringify(value), stderr: '' });
try {
resolveTaskContent({ trackerId: 'beads:GSD-42', capabilities: [beadsCapability()], execFn });
return false;
} catch (err) {
return err instanceof ResolverMalformedOutputError;
}
},
),
{ numRuns: 100 },
);
});
// ─── stderr/stdout sanitization in .message (security review, #3970) ───────
// `stderrTail`/`stdoutSample` are UNTRUSTED subprocess-sourced text (the
// resolver binary and its argv `{{id}}` token come from a PLAN.md
// `tracker-id` attribute, which is often LLM/agent-authored). Embedding them
// raw into `.message` would let a hostile or buggy resolver smuggle its own
// `\n` and forge a second `Error: ` line when a caller (task-command-
// router.cts's routeResolveContent) writes `.message` to stderr via
// io.cjs's error(). `formatDiagnosticToken()` (io.cts) is `JSON.stringify`
// under the hood: it wraps the value in quotes and escapes control
// characters, including `\n` -> the two-character sequence `\n`, so the
// result can never span more than one line.
test('ResolverFailedError.message has no raw literal newline even when stderrTail smuggles one', () => {
const hostileStderr = 'real failure text\nError: fake message forged by a hostile resolver';
const err = new ResolverFailedError('bd', 1, hostileStderr);
assert.ok(
!err.message.includes('\n'),
`expected no raw newline in .message, got: ${JSON.stringify(err.message)}`,
);
// The raw field is left untouched for programmatic callers — only the
// rendered .message is sanitized.
assert.strictEqual(err.stderrTail, hostileStderr);
// formatDiagnosticToken === JSON.stringify: the escaped '\n' (backslash-n,
// two characters) survives inside the JSON-quoted substring.
assert.ok(err.message.includes('\\n'));
assert.ok(err.message.includes(JSON.stringify(hostileStderr)));
});
test('ResolverMalformedOutputError.message has no raw literal newline even when stdoutSample smuggles one', () => {
const hostileStdout = '{"description":"x"\nError: fake message forged by a hostile resolver';
const err = new ResolverMalformedOutputError('bd', 'stdout is not valid JSON', hostileStdout);
assert.ok(
!err.message.includes('\n'),
`expected no raw newline in .message, got: ${JSON.stringify(err.message)}`,
);
assert.strictEqual(err.stdoutSample, hostileStdout);
assert.ok(err.message.includes('\\n'));
assert.ok(err.message.includes(JSON.stringify(hostileStdout)));
});
test('ResolverFailedError.message with an empty stderrTail omits the trailing colon segment', () => {
const err = new ResolverFailedError('bd', 1, '');
assert.strictEqual(err.message, "resolver command 'bd' exited 1");
});

View File

@@ -0,0 +1,84 @@
'use strict';
/**
* Parity test: `capability-validator.cjs`'s install-time `KEBAB_RE` grammar
* check on `taskContentResolver.trackerPrefix` (`validateTaskContentResolver`)
* MUST agree with `task-content-resolution.cts`'s resolve-time re-validation
* inside `parseResolverDeclaration` (exercised here via `findResolver`) on
* every `trackerPrefix` value.
*
* This is the "Generative Fix Divergence" guard CLAUDE.md requires whenever
* two surfaces share a rule with no single source of truth: the grammar is
* duplicated as a literal regex in both files (see `task-content-
* resolution.cts`'s `TRACKER_PREFIX_RE` docstring for why it is not a shared
* import), so this test is what actually keeps them from drifting apart. If a
* future change to either regex loosens or tightens it without mirroring the
* change in the other file, this test fails.
*/
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const { validateTaskContentResolver } = require('../gsd-core/bin/lib/capability-validator.cjs');
const { findResolver } = require('../gsd-core/bin/lib/task-content-resolution.cjs');
function validInvoke() {
return { binary: 'bd', args: ['show', '{{id}}', '--json'], timeoutMs: 10000 };
}
function featureCapWithResolver(trackerPrefix) {
return {
id: 'demo',
role: 'feature',
taskContentResolver: { trackerPrefix, invoke: validInvoke() },
};
}
/** True when the validator accepts this trackerPrefix (no trackerPrefix-naming error). */
function validatorAccepts(trackerPrefix) {
const cap = featureCapWithResolver(trackerPrefix);
const errs = validateTaskContentResolver(cap);
return !errs.some((e) => e.includes('trackerPrefix'));
}
/** True when the resolve-time seam accepts this trackerPrefix (finds a match, not `null`). */
function resolverAccepts(trackerPrefix) {
const capabilities = [featureCapWithResolver(trackerPrefix)];
const result = findResolver(trackerPrefix, capabilities);
return result !== null && result !== 'ambiguous';
}
const TABLE = [
{ trackerPrefix: 'beads', valid: true },
{ trackerPrefix: 'my-tracker', valid: true },
{ trackerPrefix: 'Beads', valid: false },
{ trackerPrefix: 'has_underscore', valid: false },
{ trackerPrefix: 'UPPER', valid: false },
{ trackerPrefix: '', valid: false },
{ trackerPrefix: '1leading-digit', valid: false },
];
describe('trackerPrefix grammar parity — capability-validator.cjs vs task-content-resolution.cts', () => {
for (const { trackerPrefix, valid } of TABLE) {
test(`'${trackerPrefix}' — validator and resolver agree (expected valid: ${valid})`, () => {
const validatorResult = validatorAccepts(trackerPrefix);
const resolverResult = resolverAccepts(trackerPrefix);
assert.strictEqual(
validatorResult,
valid,
`validator disagreed with expected table value for '${trackerPrefix}'`,
);
assert.strictEqual(
resolverResult,
valid,
`resolver disagreed with expected table value for '${trackerPrefix}'`,
);
assert.strictEqual(
validatorResult,
resolverResult,
`PARITY BREAK: validator and resolver disagree for trackerPrefix '${trackerPrefix}' ` +
`(validator: ${validatorResult}, resolver: ${resolverResult})`,
);
});
}
});