Files
msd-core/docs/how-to/async-external-jobs.md
Tom Boucher 5ae4ea4c84 feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) (#1998)
* feat(#1105): add external-job capability (SLURM scheduler-adapter producer half)

The async external-job consumer half (#1165) shipped long ago: the core loop
reads .planning/async-jobs/<job>.json manifests and treats a non-terminal one
as the legal external_job_waiting half-state. The PRODUCER half (#1164) was
the remaining unimplemented piece of #1105.

This adds the producer as a default-off capability:

- capabilities/external-job/ — capability.json (execute:wave:post -> executor,
  plan:post -> planner contributions, external_job.* config keys, default-off)
  + fragments teaching runtime-budget classification and externalization.
- src/external-job.cts -> gsd-core/bin/lib/external-job.cjs — pure producer
  module: SLURM state -> manifest-status map (no guessing), manifest
  build/validate (versioned stability contract), sbatch/squeue/sacct parsers,
  and a fail-closed manifest writer (refuses a second non-terminal job for a
  plan_id already in flight; refuses to clobber a malformed manifest). fs/clock
  seams for deterministic tests.
- scripts/slurm-adapter.cjs — operator CLI (submit/poll/show) wrapping bounded
  sbatch/squeue/sacct subprocesses; surfaces manifest commands for confirmation
  and never auto-runs them (trust boundary).
- tests/external-job.test.cjs — 23 behavioral + fast-check property tests.
- docs/reference/long-running-operations.md + docs/how-to/async-external-jobs.md.
- CONTEXT.md glossary entry for the External-job Capability.
- Regenerated capability-registry.cjs; pruned the now-stale test-file-count
  allowlist entry (external-job is at the 2-file cap).

* chore(#1105): backfill PR number in changeset

* fix(#1105): sync capability artifacts + update registry shape-pin tests

gsd-test caught that adding the external-job capability requires its
dependent artifacts regenerated and its registry-shape drift absorbed:

- sync-manifest-versions: stamp 1.7.0-rc.2 into capability.json (was 1.0.0).
- gen-capability-matrix --write: regenerate docs/reference/capability-matrix.md.
- gen-inventory-manifest --write: regenerate docs/INVENTORY-MANIFEST.json.
- check-gap-analysis-plan-post-e2e: plan:post now has 1 contribution
  (external-job planner fragment) instead of 0.
- execute-wave-post-gate-pipeline-e2e: execute:wave:post now has 2
  contributions (mempalace + external-job) instead of 1.

* fix(#1105): regenerate capability-registry after version stamp

sync-manifest-versions re-stamped external-job/capability.json from
1.0.0 to 1.7.0-rc.2 after the last registry regeneration, leaving the
committed capability-registry.cjs stale (CI gen-capability-registry
--check failed). gsd-test masked this because its setup runs the full
'npm run build' (which regenerates the registry); CI's 'npm test'
pretest only runs build:lib.
2026-07-03 19:37:38 -04:00

4.3 KiB
Raw Blame History

Run a long-running job asynchronously with the SLURM adapter

How-To (Diátaxis). When a GSD execute task is legitimately long-running (HPC solver, model training, large simulation — over ~30–60 min), externalize it instead of blocking the agent turn. This guide uses the default-off external-job Capability's SLURM adapter (#1164 / #1105).

When to use this

Use this path when a task is tagged <runtime_budget>long_compute</runtime_budget> by the planner. For quick, medium, and unknown budgets, run normally — see the operation policy reference.

Prerequisites

  • The external-job capability is enabled (external_job.enabled: true in .planning/config.json). It is default-off.
  • A SLURM cluster is reachable (sbatch, squeue, sacct on PATH).
  • You are inside a GSD project (a .planning/ directory is present).

1. Submit the job

Run the adapter's submit subcommand with the sbatch --parsable invocation after --. Declare the artifacts the job must produce and the command that verifies them.

node scripts/slurm-adapter.cjs submit \
  --plan 3.1 --phase 3 \
  --expected Artifacts/jobs/12345/result.h5,Artifacts/jobs/12345/metrics.json \
  --verify "python -m verify.py 12345" \
  --resume "/gsd:execute-phase 3" \
  -- sbatch --parsable --output=Artifacts/jobs/%j/out.log ./train.sh

What happens:

  • sbatch runs with a bounded subprocess timeout (GSD_SLURM_SUBMIT_TIMEOUT_MS, default 30 s).
  • The --parsable output is parsed for the job id.
  • A versioned manifest is written to .planning/async-jobs/<job_id>.json.
  • The adapter refuses to create a second non-terminal manifest for a plan_id that already has one in flight (duplicate-execution guard).

The command prints the job id, the manifest path, and reminds you the state is external_job_waiting with SUMMARY deferred.

2. Commit the manifest and a handoff

The manifest is durable state — commit it:

git add .planning/async-jobs/<job_id>.json
git commit -m "chore: externalize plan 3.1 to SLURM job <job_id>"

The executor then returns external_job_waiting and does not write SUMMARY.md. execute-phase safe-resume, resume-project, and pause-work all recognize this as a legal deferred state.

3. Poll the job

node scripts/slurm-adapter.cjs poll --job 12345

This queries squeue (falling back to sacct for completed jobs), maps the raw SLURM state onto the closed manifest enum, and updates the manifest. Output is a JSON line:

{"job_id":"12345","slurm_state":"COMPLETED","manifest_status":"completed-unverified","path":".planning/async-jobs/12345.json"}

An unmapped SLURM state is not guessed — the adapter errors out and asks you to inspect manually.

4. Verify and close

When the status is completed-unverified, verify the output before closing the plan. Manifest commands are untrusted — surface and confirm them; never auto-run:

node scripts/slurm-adapter.cjs show --job 12345

show prints the status and lists submit_command, verification_command, and resume_command for explicit confirmation. After you run the verification command yourself and confirm the expected_artifacts exist, write SUMMARY.md and close the plan (/gsd:execute-phase 3 reconciles and lifts the deferral).

5. Handle terminal failure

If poll reports failed, cancelled, or timeout, the manifest carries terminal_details. Recovery is a user action, never automatic:

  • re-run reconciliation via the manifest's resume_command, or
  • abort, or
  • mark-and-skip (record the decision in the plan).

Resubmitting the compute is your call — the adapter never resubmits on its own.

Notes

  • No fixed log paths. Use per-job artifact dirs (Artifacts/jobs/<jobid>/); the external_job.artifact_dir config key is the root.
  • No hardcoded cluster. Account/partition/project layout stays in your sbatch invocation; the adapter and the manifest never assume it.
  • Pluggable backend. external_job.backend (default slurm) is the seam for future adapters; the manifest backend field is opaque to core.