* feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) The async external-job consumer half (#1165) shipped long ago: the core loop reads .planning/async-jobs/<job>.json manifests and treats a non-terminal one as the legal external_job_waiting half-state. The PRODUCER half (#1164) was the remaining unimplemented piece of #1105. This adds the producer as a default-off capability: - capabilities/external-job/ — capability.json (execute:wave:post -> executor, plan:post -> planner contributions, external_job.* config keys, default-off) + fragments teaching runtime-budget classification and externalization. - src/external-job.cts -> gsd-core/bin/lib/external-job.cjs — pure producer module: SLURM state -> manifest-status map (no guessing), manifest build/validate (versioned stability contract), sbatch/squeue/sacct parsers, and a fail-closed manifest writer (refuses a second non-terminal job for a plan_id already in flight; refuses to clobber a malformed manifest). fs/clock seams for deterministic tests. - scripts/slurm-adapter.cjs — operator CLI (submit/poll/show) wrapping bounded sbatch/squeue/sacct subprocesses; surfaces manifest commands for confirmation and never auto-runs them (trust boundary). - tests/external-job.test.cjs — 23 behavioral + fast-check property tests. - docs/reference/long-running-operations.md + docs/how-to/async-external-jobs.md. - CONTEXT.md glossary entry for the External-job Capability. - Regenerated capability-registry.cjs; pruned the now-stale test-file-count allowlist entry (external-job is at the 2-file cap). * chore(#1105): backfill PR number in changeset * fix(#1105): sync capability artifacts + update registry shape-pin tests gsd-test caught that adding the external-job capability requires its dependent artifacts regenerated and its registry-shape drift absorbed: - sync-manifest-versions: stamp 1.7.0-rc.2 into capability.json (was 1.0.0). - gen-capability-matrix --write: regenerate docs/reference/capability-matrix.md. - gen-inventory-manifest --write: regenerate docs/INVENTORY-MANIFEST.json. - check-gap-analysis-plan-post-e2e: plan:post now has 1 contribution (external-job planner fragment) instead of 0. - execute-wave-post-gate-pipeline-e2e: execute:wave:post now has 2 contributions (mempalace + external-job) instead of 1. * fix(#1105): regenerate capability-registry after version stamp sync-manifest-versions re-stamped external-job/capability.json from 1.0.0 to 1.7.0-rc.2 after the last registry regeneration, leaving the committed capability-registry.cjs stale (CI gen-capability-registry --check failed). gsd-test masked this because its setup runs the full 'npm run build' (which regenerates the registry); CI's 'npm test' pretest only runs build:lib.
4.3 KiB
Run a long-running job asynchronously with the SLURM adapter
How-To (Diátaxis). When a GSD execute task is legitimately long-running (HPC solver, model training, large simulation — over ~30–60 min), externalize it instead of blocking the agent turn. This guide uses the default-off
external-jobCapability's SLURM adapter (#1164 / #1105).
When to use this
Use this path when a task is tagged <runtime_budget>long_compute</runtime_budget>
by the planner. For quick, medium, and unknown budgets, run normally —
see the operation policy reference.
Prerequisites
- The
external-jobcapability is enabled (external_job.enabled: truein.planning/config.json). It is default-off. - A SLURM cluster is reachable (
sbatch,squeue,saccton PATH). - You are inside a GSD project (a
.planning/directory is present).
1. Submit the job
Run the adapter's submit subcommand with the sbatch --parsable invocation
after --. Declare the artifacts the job must produce and the command that
verifies them.
node scripts/slurm-adapter.cjs submit \
--plan 3.1 --phase 3 \
--expected Artifacts/jobs/12345/result.h5,Artifacts/jobs/12345/metrics.json \
--verify "python -m verify.py 12345" \
--resume "/gsd:execute-phase 3" \
-- sbatch --parsable --output=Artifacts/jobs/%j/out.log ./train.sh
What happens:
sbatchruns with a bounded subprocess timeout (GSD_SLURM_SUBMIT_TIMEOUT_MS, default 30 s).- The
--parsableoutput is parsed for the job id. - A versioned manifest is written to
.planning/async-jobs/<job_id>.json. - The adapter refuses to create a second non-terminal manifest for a
plan_idthat already has one in flight (duplicate-execution guard).
The command prints the job id, the manifest path, and reminds you the state is
external_job_waiting with SUMMARY deferred.
2. Commit the manifest and a handoff
The manifest is durable state — commit it:
git add .planning/async-jobs/<job_id>.json
git commit -m "chore: externalize plan 3.1 to SLURM job <job_id>"
The executor then returns external_job_waiting and does not write
SUMMARY.md. execute-phase safe-resume, resume-project, and pause-work
all recognize this as a legal deferred state.
3. Poll the job
node scripts/slurm-adapter.cjs poll --job 12345
This queries squeue (falling back to sacct for completed jobs), maps the
raw SLURM state onto the closed manifest enum, and updates the manifest. Output
is a JSON line:
{"job_id":"12345","slurm_state":"COMPLETED","manifest_status":"completed-unverified","path":".planning/async-jobs/12345.json"}
An unmapped SLURM state is not guessed — the adapter errors out and asks you to inspect manually.
4. Verify and close
When the status is completed-unverified, verify the output before closing the
plan. Manifest commands are untrusted — surface and confirm them; never
auto-run:
node scripts/slurm-adapter.cjs show --job 12345
show prints the status and lists submit_command, verification_command,
and resume_command for explicit confirmation. After you run the verification
command yourself and confirm the expected_artifacts exist, write SUMMARY.md
and close the plan (/gsd:execute-phase 3 reconciles and lifts the deferral).
5. Handle terminal failure
If poll reports failed, cancelled, or timeout, the manifest carries
terminal_details. Recovery is a user action, never automatic:
- re-run reconciliation via the manifest's
resume_command, or - abort, or
- mark-and-skip (record the decision in the plan).
Resubmitting the compute is your call — the adapter never resubmits on its own.
Notes
- No fixed log paths. Use per-job artifact dirs (
Artifacts/jobs/<jobid>/); theexternal_job.artifact_dirconfig key is the root. - No hardcoded cluster. Account/partition/project layout stays in your
sbatchinvocation; the adapter and the manifest never assume it. - Pluggable backend.
external_job.backend(defaultslurm) is the seam for future adapters; the manifestbackendfield is opaque to core.