* feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) The async external-job consumer half (#1165) shipped long ago: the core loop reads .planning/async-jobs/<job>.json manifests and treats a non-terminal one as the legal external_job_waiting half-state. The PRODUCER half (#1164) was the remaining unimplemented piece of #1105. This adds the producer as a default-off capability: - capabilities/external-job/ — capability.json (execute:wave:post -> executor, plan:post -> planner contributions, external_job.* config keys, default-off) + fragments teaching runtime-budget classification and externalization. - src/external-job.cts -> gsd-core/bin/lib/external-job.cjs — pure producer module: SLURM state -> manifest-status map (no guessing), manifest build/validate (versioned stability contract), sbatch/squeue/sacct parsers, and a fail-closed manifest writer (refuses a second non-terminal job for a plan_id already in flight; refuses to clobber a malformed manifest). fs/clock seams for deterministic tests. - scripts/slurm-adapter.cjs — operator CLI (submit/poll/show) wrapping bounded sbatch/squeue/sacct subprocesses; surfaces manifest commands for confirmation and never auto-runs them (trust boundary). - tests/external-job.test.cjs — 23 behavioral + fast-check property tests. - docs/reference/long-running-operations.md + docs/how-to/async-external-jobs.md. - CONTEXT.md glossary entry for the External-job Capability. - Regenerated capability-registry.cjs; pruned the now-stale test-file-count allowlist entry (external-job is at the 2-file cap). * chore(#1105): backfill PR number in changeset * fix(#1105): sync capability artifacts + update registry shape-pin tests gsd-test caught that adding the external-job capability requires its dependent artifacts regenerated and its registry-shape drift absorbed: - sync-manifest-versions: stamp 1.7.0-rc.2 into capability.json (was 1.0.0). - gen-capability-matrix --write: regenerate docs/reference/capability-matrix.md. - gen-inventory-manifest --write: regenerate docs/INVENTORY-MANIFEST.json. - check-gap-analysis-plan-post-e2e: plan:post now has 1 contribution (external-job planner fragment) instead of 0. - execute-wave-post-gate-pipeline-e2e: execute:wave:post now has 2 contributions (mempalace + external-job) instead of 1. * fix(#1105): regenerate capability-registry after version stamp sync-manifest-versions re-stamped external-job/capability.json from 1.0.0 to 1.7.0-rc.2 after the last registry regeneration, leaving the committed capability-registry.cjs stale (CI gen-capability-registry --check failed). gsd-test masked this because its setup runs the full 'npm run build' (which regenerates the registry); CI's 'npm test' pretest only runs build:lib.
5.4 KiB
Long-running operations and async external jobs
Reference (Diátaxis). The operation-classification policy and the async external-job contract that GSD executors use when a task legitimately exceeds a child-agent timeout. The producer is the default-off
external-jobCapability (#1164, part of #1105); the core loop consumes the manifest (#1165 —external_job_waiting).
1. The problem
Heavy GSD phases (HPC solvers, model training, large simulations) can legitimately exceed short child-agent timeouts. Raising the timeout alone is not sufficient: it prevents legitimate subagents from being killed, but it also lets truly hung commands consume the whole agent budget. GSD must distinguish legitimate heavy work from suspicious hangs, keep safety nets finite, and avoid blocking an agent turn on hours-long compute.
2. Operation classification policy
Every executable task carries a runtime budget. Planners emit it via a
<runtime_budget> element (taught by the external-job Capability's
plan:post fragment); executors branch on it at execute:wave:post.
| Budget | Meaning | Execute behavior |
|---|---|---|
quick |
Under ~2 min | Run normally in the foreground. |
medium |
~2–30 min | Foreground, but with explicit progress expectations. |
unknown |
Runtime not characterized | Run a first-health check and set a soft-review deadline before trusting the child timeout. Define a progress signal, an abort condition, and expected output. A truly hung command must surface before the child timeout is exhausted. |
long_compute |
Over ~30–60 min | Externalize — submit an async external job, record durable state, and return external_job_waiting. Never block the agent turn. |
The classification is advisory metadata; the executor owns the decision at
dispatch time. unknown is the safety-critical class: it is what catches a
hung solver before it burns the budget, without making timeouts infinite.
3. The async external-job half-state
When an executor externalizes a long_compute task it enters a legal
deferred state, not an illegal partial-plan state:
input/code committed
external job submitted
.planning/async-jobs/<job>.json committed
handoff committed
SUMMARY.md deferred until verification
SUMMARY.md is deferred until the job reaches a terminal state and its
expected_artifacts are verified. Until then, the plan is
external_job_waiting, and every resume/pause/dispatch path reconciles
against the manifest — it never re-dispatches the plan (re-dispatching would
duplicate the external job).
4. The manifest — a versioned stability contract
The manifest schema, status enum, trust boundary, matching rules, and the
glob-safe matching probe are the stability contract documented in
planning-artifacts.md.
The core loop depends only on the named fields and ignores any others; the
version field is the evolution escape hatch. Producers MUST write the named
fields and MAY add their own.
Status enum (closed, scheduler-agnostic — producers map backend states onto these):
| Status | Class | Resume action |
|---|---|---|
submitted, running |
non-terminal | Re-check; never re-dispatch. |
completed-unverified |
finished, unverified | Verify expected_artifacts / run verification_command; on success write SUMMARY.md and close. |
failed, cancelled, timeout |
terminal failure | Surface terminal_details; offer recovery (re-run reconciliation, abort, or mark-and-skip). Resubmitting compute is a user action, never automatic. |
5. Trust boundary
The manifest crosses a trust seam: a Capability (or anything that can write
.planning/) produces it; the core loop consumes it. submit_command,
verification_command, and resume_command are therefore untrusted. The
core loop — and the slurm-adapter show subcommand — surface these commands
for explicit operator confirmation; they are never auto-executed. Validate
before trusting a manifest: recognized version, plan_id matches the plan
under reconciliation, and status is one of the closed enum values. On a
malformed manifest or multiple manifests for one plan_id, fail closed.
6. Scheduler pluggability
SLURM is the first backend. The design does not hardcode a cluster, account,
partition, or project layout — per-job artifact directories
(external_job.artifact_dir, default Artifacts/jobs/<jobid>/) avoid fixed
log paths. The backend field on the manifest (opaque to core) and the
external_job.backend config key are the pluggability seams for future
backends (LSF, PBS, Kubernetes batch). The pure producer logic — SLURM
state→manifest-status mapping, manifest build/validate, sbatch/squeue/
sacct parsers, and the fail-closed writer — is backend-aware but lives
behind a single module; a new backend adds a sibling state map and parser
without touching core.
7. Related
- How-To:
../how-to/async-external-jobs.md— using the SLURM adapter. - Contract:
planning-artifacts.md— the manifest stability contract. - Capability manifest:
../../capabilities/external-job/capability.json. - Pure module:
gsd-core/src/external-job.cts→gsd-core/bin/lib/external-job.cjs. - Operator CLI:
scripts/slurm-adapter.cjs.