Files
msd-core/docs/how-to/adopt-the-v2-exit-contract.md
Tom Boucher d24e22b156 enhance(#3912): gsd-tools declares outcomes, pinned at v1 (#3983)
* enhance(#3912): gsd-tools declares outcomes, pinned at v1

ADR-3889 §4. Phase 6 already moved error()'s terminator onto the seam, so what
remained was the declaration — and the pin that makes it invisible today.

The census corrected two documented figures before any code changed.
ERROR_REASON has exactly 25 members (the ADR and epic were right; an earlier
note of mine claiming 23 was wrong and is corrected). And output({error}) is
**64 sites across 9 files, not the 60 ADR-2980 ratified** — the module shape
holds but the total drifted +4: frontmatter 7 not 6, phase 4 not 2, roadmap 3
not 2. That matters because this phase's criterion demands the pin be asserted
over the enumerated population rather than sampled; asserting over a stale 60
would leave four sites unpinned while claiming full coverage, which is the
shape of failure this epic exists to remove.

The issue does not state the fact that shapes the design: output() never
touches the exit code. Confirmed by reading it — it writes fd 1 and returns.
So a declared outcome for those 64 sites had nowhere to be READ. The mapping
was never the work; wiring somewhere for the declaration to land was.

The seam already existed twice over. cli-exit.cts holds two globalThis-Symbol
cells, each because the module is emitted to three locations and a module-level
`let` would let instances disagree, and runMain already maps a code returned by
main(). A third cell inherits that solution. output() records DEGRADED for any
{error} payload — key-order agnostic, which is exactly why the "42 sites"
figure undercounts — and runMain projects the cell only when main() returns
nothing, so an explicit return still wins.

error() maps its reason through a table over the closed 25-member enum, leaving
all 278 call sites untouched; 226 of them pass no reason at all. The version
gate lives in error(), NOT in projectOutcome: registered names are
version-invariant there, so mapping a reason straight through would make USAGE
project to 64 under v1 and break the pin on its first line. projectOutcome is
left exactly as Phase 2 shipped it, DEGRADED's 0/80 asymmetry included.

Proven rather than asserted. v1 is byte-identical across three real CLI paths —
config-get plain, config-get --json-errors, and an output({error}) path —
matching exit code and exact bytes against the pre-change build. Under
GSD_EXIT_CONTRACT=v2 the same commands now exit 66 (CONFIG_KEY_NOT_FOUND ->
NO_INPUT) and 80 (DEGRADED), both looked up through the registry. An
anti-vacuity test pins that v1 and v2 genuinely differ for at least one reason,
because without it a mapping where everything projects to 1 under both versions
would satisfy every other assertion and the declaration would be theatre.

A1 iterates all 25 enum members and A3 asserts over the measured 64-site
population, so a 26th reason or a 65th site fails until it is given a mapping —
the drift guard this phase needs, given ADR-2980's own count had drifted +4
unnoticed.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the outcome cell must never lower an exit code

The remote run caught a fail-open that this phase introduced, in the phase
whose entire purpose is removing fail-opens.

`state validate --strict` on a missing STATE.md exited **0** where it must exit
1. Mechanism: `runMain` projected the pending outcome whenever `main()` returned
void, and under v1 DEGRADED projects to 0 — so a `process.exitCode` already set
non-zero by the command was clobbered down to success. Confirmed live against a
fixture, before and after.

This refutes a review conclusion recorded earlier in this phase, that the cell
was "fail-closed and can never mask a failure as success". It could, and did.
Recording that plainly so the assumption is not repeated: the cell's danger was
never only that it might add a failure — it was that projecting it
unconditionally overwrites whatever decision came before.

Projection is now guarded: it may set a code only when none is set, and an
already-non-zero exit code always wins. The full precedence — explicit `main()`
return, then an existing non-zero exitCode, then the declared outcome — is
written at the projection site. A regression test drives a void return with a
pre-set non-zero code and a pending DEGRADED, and fails against the pre-fix
build.

The second failure was my test encoding the wrong contract, not a code defect.
It asserted `output({found:false, error: undefined})` records DEGRADED because
the KEY is present. `JSON.stringify` drops undefined, so the payload the user
receives is `{"found":false}` — carrying no error at all, and calling that
degraded would hand back exit 80 under v2 for output that reads as clean. The
discriminator is a serializable error VALUE, not key presence. The test now
pins `{error: undefined}` as explicitly NOT degraded, and the design doc's
wording is tightened to match.

Verification runs on the remote runner.

Refs #3912

* docs(#3912): the versioned exit contract, and a flag defect the docs found

Diataxis pass for Phase 8, plus a real fix that only surfaced because writing
the how-to meant running its own examples.

The docs. ADR-2980's "Revisit if" clause asked for exactly the versioned
projection this phase provides, so it gets an amendment naming #3912 /
ADR-3889 section 4 as that boundary: v1 stays 0 byte-for-byte, v2 projects
DEGRADED to 80. The amendment also records the count drift rather than
restating a stale figure — the ADR ratified 60 output({error}) sites in 9
modules; the AST-measured population is 64 across the same 9 (frontmatter 7
not 6, phase 4 not 2, roadmap 3 not 2). The pin is asserted over the
enumerated 64. json-errors.md gains the outcome-declaration reference,
including the precedence order a review pass got wrong and the suite refuted:
an explicit main() return, then an already-set non-zero process.exitCode, then
the declared outcome. Projection may only ever set a code, never lower one.

A how-to is owed here and is written, not skipped. Under v1 nothing changes,
so the audience is an operator opting into v2 and needing to know what the
codes mean for a CI gate — a migration, which is how-to shaped. It covers
turning v2 on, the code table, why 80 is "ran and reported a condition" rather
than a crash, and how to split a gate that treats any non-zero as fatal. No
tutorial: there is no new entry point to learn, and under the default contract
a reader would be walked through observing nothing.

The defect. Running the how-to's own Step 1 example returned

    $ gsd-tools --exit-contract=v2 state validate --strict
    Error: Unknown command: --exit-contract=v2          (exit 64)

while the same flag trailing the subcommand worked and exited 80. The flag
half-worked, by argv position. resolveContractVersion scans argv
non-destructively, so the token survived into the dispatcher, which treats
argv[2] as the command name. --json-errors had already solved precisely this
at gsd-tools.cjs:4455, under a comment naming the hazard verbatim: "The argv
splice must happen here too, otherwise the dispatcher below sees
--json-errors as an unknown command." The later flag never got the same
treatment.

Fixed rather than documented around: the version is resolved first — which
memoizes the cell and makes an invalid value throw early — and then every
--exit-contract= token is spliced out of the dispatcher's argv copy.
--exit-contract is now listed in TOP_LEVEL_USAGE, where it never was. The
regression test pins leading position, trailing position, agreement between
the two, and a loud failure on v3 rather than a silent fall back to v1.

Neither review engine would have caught this: the defect is invisible in the
diff, because the diff does not touch argv handling. It surfaced only from
running the documentation's own example. Writing a how-to is an execution pass.

Verification runs on the remote runner.

Refs #3912

* fix(#3912): the flag splice has to run before the run-with-timeout return

An isolated review of the previous commit found that the fix did not deliver
what it claimed, and that two of its own tests were weak. All three findings
reproduced by execution before any change was made.

The fix was placed below a return. main() intercepts `run-with-timeout` at
gsd-tools.cjs:4436 and returns from there — above both the --json-errors block
and the --exit-contract splice added in the previous commit. So the flag still
died in leading position for that one command:

    $ gsd-tools --exit-contract=v2 run-with-timeout 5 -- node -e "..."
    Error: Unknown command: run-with-timeout        (exit 64, child never ran)

The previous commit message and the test's describe-block both claimed
position-independence unconditionally. That was an overclaim, not a gap left
open, and it is the part worth naming: the fix was verified by hand on the
commands I happened to think of, and `run-with-timeout` returns before the
code I was verifying.

Both global-flag blocks now run above the interception, with a comment naming
it so a later edit cannot slide them back down. Moving --json-errors up fixes
the identical pre-existing bug for that flag, verified failing beforehand
(exit 1, sdk_unknown_command). Fixing the sibling is deliberate: same defect,
same block, and a known-broken twin next to a fixed one is not a resting state.

Two tests were not pulling their weight. The invalid-value test was vacuous —
it passed against the pre-fix build, because `--exit-contract=v3` already
exited 1 there and already printed the resolve error lazily through
error() -> getContractVersion. Both its assertions held before the fix, so it
pinned nothing. The real discriminator is that the pre-fix build emits BOTH
"Unknown command: --exit-contract=v3" and the resolve error, while the fixed
build emits only the latter; the test now asserts that absence.

The leading-position and leading==trailing tests asserted proxies — "not 64",
"no Unknown command", "the two agree" — none of which pin a value, and all of
which would survive both positions being identically broken. With a .planning
directory and no STATE.md, state-snapshot exits exactly 80 under v2 and 0
under v1 in both positions. Those numbers are pinned now. The multi-token case
the descending splice loop exists for is covered too, and run-with-timeout has
regression tests for both flags.

The lesson is narrower than "test more". Hand-verifying the production
behavior does not verify that the test would have caught its absence. The
pre-fix binary has to be run against the test's own assertions.

Investigated and deliberately not changed: splicing before --cwd parsing
degrades one diagnostic from "Missing value for --cwd" to "Invalid --cwd:
<path>", but that is pre-existing — verified on the pre-fix build via
--json-errors, which already did it. This change joins the pattern rather than
creating it, and both forms exit 64 on malformed input either way.

Verification runs on the remote runner.

Refs #3912

* chore(#3912): backfill changeset pr numbers to 3983

* test(#3912): pin the reason-table invariant as set equality, not a count

A graph-backed review flagged the unchecked lookup in
expectedErrorCode3912. Investigated by execution: the drift guard DOES
hold — for an unmapped reason under v2 the production error() yields 1
while the table yields undefined, so the assertion fails. Not a
correctness defect, and deliberately NOT made tolerant, since a tolerant
lookup would destroy the guard.

Two real problems remained. The guard asserted the wrong invariant: it
counted the TABLE's keys at 25 rather than checking they match the
ENUM's values, so a renamed member keeps the count at 25 and slips past,
and a 26th member leaves the table at 25 and slips past too. Both were
then caught only indirectly, by an undefined mismatch producing 'must
exit undefined'. It is now a sorted set equality, so the failure names
the specific missing or extra reason.

And the comment above it described a '?? FAIL' fallback that does not
exist anywhere in the function. It now states what the code actually
does, verified by running it rather than by reading it.

Refs #3912

---------

Co-authored-by: sim <sim@local>
2026-08-28 08:09:05 -04:00

7.7 KiB

How to adopt the v2 exit contract

gsd-tools runs, by default, under exit-contract v1 — every observable exit code is byte-identical to every prior release. v2 is an opt-in projection, per ADR-3889 §4, that turns a small set of previously same-looking outcomes into distinct, non-zero exit codes a CI gate can branch on without parsing stdout. This page covers how to turn it on, what the new codes mean, and how to migrate a gate that today treats "any non-zero exit" as fatal.

Should you turn this on?

Turn it on if a script or CI gate wraps gsd-tools and needs to tell "ran fine", "you called it wrong", "there was nothing to find", "a prerequisite was missing", "it crashed", and "it ran but is reporting a condition in its payload" apart from the exit code alone, instead of parsing JSON on stdout to find out. If your caller only ever needs pass/fail, v1 already gives you that — there is nothing to adopt.

Step 1 — turn it on

Either a flag or an environment variable activates v2. The flag beats the env var if both are given in the same invocation.

# Per-invocation (preferred in test code and one-off scripts):
node gsd-tools.cjs --exit-contract=v2 state validate --strict

# Process-wide (preferred for CI and shell wrappers):
export GSD_EXIT_CONTRACT=v2
gsd-tools state validate --strict

The flag works in any argv position — before the subcommand or after it — and gsd-tools strips it before dispatch, exactly as it already does for --json-errors.

Only the exact lowercase tokens v1/v2 are accepted. Anything else present — a typo, V2, v3, or an explicitly empty --exit-contract= — throws rather than silently falling back to v1; a selector for a contract whose whole point is "nothing fails with success" must not itself fail open. That throw surfaces as a stack trace at exit 1, not as a tidy USAGE 64, and that is deliberate: the failure is the version being unresolvable, so there is no contract version yet under which to project a code. Loud and coarse beats quiet and wrong. An empty GSD_EXIT_CONTRACT= (nothing after the =) reads as unset, not as an explicit selection, so a shell that exports it empty still gets v1.

Step 2 — read the code table

v2 projects a declared outcome through one generated registry (gsd-core/bin/lib/exit-code-registry.cjs). Every code below is stable and machine-checked; do not hardcode the integers in your own scripts — call exitCodeFor(name) if you are writing Node, or just compare against the number after reading it here once.

Code Name Meaning
0 (PASS) The operation ran and its verdict is affirmative.
1 (FAIL) The operation ran and its verdict is negative — the honest default when nothing more specific applies.
64 USAGE Caller error — bad argv, unknown subcommand, missing required argument.
66 NO_INPUT Ran; zero units were in scope, and that emptiness is known to be genuine.
69 UNAVAILABLE Could not run — a prerequisite was absent, unreadable, or scope was never established.
70 INTERNAL Self-failure — the run itself broke (crash, timeout, killed subprocess), not its inputs.
80 DEGRADED Ran to completion and is reporting a condition through its result payload, not failing as a process.

2 is reserved to the Claude Code hook-protocol deny and is never produced by gsd-tools itself — you will not see it from a CLI invocation.

Step 3 — understand 80 specifically

80 (DEGRADED) is the one code most CI authors get wrong, because it is the one code where "ran to completion" and "found a problem" are the same event. It fires when a command's JSON payload carries a serializable error key — for example gsd-tools state-snapshot in a project with no STATE.md returns {"error": "STATE.md not found"}. Under v1 that exits 0; under v2 it exits 80.

80 is not a crash. The process did its job: it determined, correctly, that the artifact you asked about is absent, unparseable, or that a required argument was missing — see ADR-2980 for the full population this covers (64 call sites across nine modules) and why it stayed a payload-carried signal rather than a thrown fault. 70 (INTERNAL) is the code for an actual crash. Do not conflate the two: a gate that maps 80 to "the tool is broken" will page someone for a condition the tool successfully diagnosed.

Step 4 — migrate a gate that treats any non-zero exit as fatal

The naive shell form,

if ! gsd-tools state-snapshot > snap.json; then
  echo "gsd-tools failed" >&2
  exit 1
fi

is correct as a fail-safe under v2 — every registered code is non-zero, so this can only turn a false green red, never a red green (ADR-3889 §5). What it cannot do on its own is tell you which non-zero condition fired, which matters if your policy is "treat DEGRADED as a soft warning but still hard-fail on USAGE/UNAVAILABLE/INTERNAL":

set +e
gsd-tools --exit-contract=v2 state-snapshot > snap.json
code=$?
set -e

case "$code" in
  0)   ;;                                            # PASS
  80)  echo "degraded result — inspect snap.json" >&2 ;;  # ran, reported a condition — your call whether this gates the pipeline
  64|66|69|70)
       echo "gsd-tools failed (exit $code)" >&2
       exit 1
       ;;
  *)   echo "gsd-tools failed (unrecognized exit $code)" >&2
       exit 1
       ;;
esac

Whether DEGRADED itself should gate your pipeline is a policy decision only you can make — the contract's only promise is that 80 is distinguishable from 64/66/69/70, not that it is always safe to ignore. A gate that wants the old, coarser behavior (any non-zero is fatal, including 80) needs no case statement at all; the naive form above already does that correctly.

Four things that will surprise you

  1. --json-errors is a different, orthogonal switch. It governs whether error()'s stderr envelope is JSON or plain text; it does nothing to output()'s exit code. You can run v2 with or without --json-errors.
  2. Precedence can surprise a caller reading only output()'s contract. An explicit main() return, or a non-zero process.exitCode a command set directly, always wins over a DEGRADED declared earlier in the same invocation — projection can only raise a code, never lower one. See docs/json-errors.md for the full precedence order.
  3. The declaration does not accumulate. If a command calls output() more than once — a diagnostic degraded payload followed by a clean final one — only the last call's declaration is live when the process exits.
  4. v1 and v2 are the only recognized versions today, and the default flips to v2 at the next major (ADR-3889 §4). Pin --exit-contract=v1 explicitly in a script that must keep today's codes indefinitely, rather than relying on the current default staying v1 forever.
  • ADR-3889 — the exit-code registry and the versioned projection this page walks through
  • ADR-2980 — why output({error}) exits 0 under v1, and the population DEGRADED covers under v2
  • docs/json-errors.md — the full reference for both failure channels, the error-code taxonomy, and the outcome-declaration precedence rules
  • Resolve a raw-terminator finding — the sibling page for the lint rule that keeps every termination path routed through this same seam