Files
msd-core/docs/json-errors.md
Adnan bdfc62889b fix(#3784): read the hybrid Current Plan: N of M shape, keep zero-padding, and name the accepted shapes on failure (#3791)
* fix(state): read the hybrid "Current Plan: N of M" shape

advancePlanCore derived the value FORMAT from the field NAME, so it
handled the legacy pair (`Current Plan` + `Total Plans in Phase`) and the
compound `Plan: N of M`, but not the hybrid of the two: the legacy field
name carrying a compound value with no Total Plans sibling. `legacyTotal`
is null so the legacy branch fell through, and the compound branch reads
the `Plan` field through a `^Plan:`-anchored pattern that never matches
`Current Plan:`. Both produced NaN against a file whose plan numbers are
plainly readable.

The shape is not exotic. An agent wrote it unprompted into a project's
STATE.md, believing it was the parseable form, and every subsequent run
in that project inherited the failure and worked around it by hand.

Track the field name and the value shape separately (`planSourceField`,
`planRawValue`) so write-back targets whichever field the value came
from. The legacy pair still takes precedence when both fields exist, so
a stray "of N" inside Current Plan cannot override an explicit Total
Plans — covered by a new test.

Also replace the caller's catch-all error. It reported "Cannot parse
Current Plan or Total Plans" for ANY transition failure, and named no
accepted shape, so a reader learned neither what failed nor what to
write. It now distinguishes "no result" from "unreadable plan position"
and lists all three shapes. The existing test asserted the literal
"cannot parse"; it now asserts the message names the shapes, which is
the property that makes it actionable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* fix(state): keep zero-padding when advancing a compound plan value

The compound write-back rewrote only the leading half of "N of M", so a
padded value drifted lopsided: "04 of 06" advanced to "5 of 06". Cosmetic
on its own, but a plan line that looks wrong is one the next writer tidies
by hand, and hand-tidying this particular line is what produced the hybrid
shape the previous commit had to teach the parser to read.

Pad the incremented number to the width it was written with. padStart never
truncates, so a value that outgrows its padding widens correctly: 09 of 12
advances to 10 of 12. Unpadded values are untouched — 2 of 6 still advances
to 3 of 6.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* fix(state): pass a literal field name to the compound write-back

The previous commit passed `planSourceField` — a variable — as the field-name
argument to `stateReplaceField`, which trips the state-write-path drift guard's
`unstripped_content_write` axis (ADR-3408 §8.3(b)). The guard is right to care:
a Title-Case literal cannot collide with a lowercase or snake_case frontmatter
key, so it is safe whatever the content argument is, while a variable could
hold anything and therefore requires its content to be demonstrably
frontmatter-stripped first.

The content argument here IS stripped — `body` is `stripFrontmatter(content)` —
but the guard does a narrow backward scan rather than dataflow tracking, by
design, and the nearest preceding assignment to `body` is another
`stateReplaceField` result. Rather than baseline a bypass or ask a future
reader to re-derive that the invariant holds, dispatch on the discriminator and
pass the literal.

Guard goes from 1 finding to 0; its own 32 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* chore(3784): add changeset fragment for #3785

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* test(#3784): cover the maintainer's AC1 write-back and AC6 reader-anchoring

Triage published six acceptance criteria; two were only half-covered.

AC1 asks that the hybrid write back to the SAME field with padding preserved.
The existing hybrid test used an unpadded value and asserted only `result.data`,
so it proved the parse but never the write. Now asserts the written content is
`05 of 06` on the original field, and that no separate `Plan:` field appears as
a side effect.

AC6 asks that the shared field reader not be loosened. Reading the hybrid is the
transition's job; `stateExtractField('Plan')` is line-anchored and has 13+
callers, so teaching it to match a name merely ENDING in "Plan" would be the
wrong fix and would silently change what those callers read. This holds by
construction here — the reader is untouched — but nothing locked it in. The new
test fails if anyone later reaches for that shortcut.

Also drops the changeset fragment written against the auto-closed PR number.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* chore(#3784): add changeset fragment for #3791

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8

* fix(#3784): write the advanced plan back to the field it was read from

Review findings 2-6 on #3791 were one defect seen from several angles: the
read path learned the hybrid `Current Plan: N of M` shape, the write path
did not follow it.

- `bumpLeadingNumber` now owns the increment for all three parse branches.
  Only the leading digits belong to this transition; the padding width and
  everything after it (` of M`, and the `\r` of a CRLF file) are the
  author's text and are preserved. The legacy branch wrote `String(newPlan)`,
  which turned `2 of 99` into `3` and `04` into `5`.
- `mutateCurrentPositionForAdvance` takes the plan field NAME. Its plan arm
  only ever looked for `Plan:`, so on a hybrid file the `## Current Position`
  section was never reached; combined with the body-level write being
  single-shot and bold-preferring, a file carrying the field at both sites
  advanced the header and left the section a plan behind. The parameter
  defaults to `Plan`, so the two callers that pass no plan are unchanged.
- Tests: both-sites-advance (fails without the section arm), legacy
  write-back content assertions (the previous test read only `data` and so
  could not see the lossy write), hybrid boundary at limit-1 and limit+1, a
  CRLF fixture, and an fc property pinning the padding-width contract.

Two characterization tests pinned `**Current Plan:** 02` advancing to `3`.
That dropped padding is the defect #3784 reports, so the expectation is
corrected to `03` rather than the fix being narrowed around it.

* fix(#3784): drop the unreachable advance-plan error branch, sync the doc

Findings 1 and 8 on #3791.

The `!resultData` arm could not fire: the transform callback assigns
`resultData` unconditionally, only runs once STATE.md is known to exist (the
missing-file case returns "STATE.md not found" upstream), and every
`advancePlanCore` return path sets `data`. It was a speculative second
failure mode with a message no caller could receive, and the comment beside
it claimed to distinguish two things that were never two. `!resultData`
stays in the condition as a type guard, which is all it ever was.

`docs/json-errors.md:142` quoted the old error literal verbatim and was the
sole occurrence in the tree; it now quotes the emitted one.

* chore(#3784): describe the write-back fix in the changeset

* fix(#3784): anchor the plan grammar and widen the schema row to match

Review round 3 on #3791: B1, B2, M1, M2, M3, M4 and the planSourceField nit.

B1 — `STATE_FIELD_SCHEMA.current_plan.acceptedShapes` widens to
`['N', 'N of M']`, which is what `src/state-md-schema.cts`'s own comment
instructed this PR to do on merge. `'N/M'` stays undeclared so row 23 keeps a
non-vacuous undeclared candidate to probe. Verified `gen-state-md-docs
--check` exits 0 and `--write` rewrites 0 of 6: the generated artifacts do not
surface this row, so there is nothing stale to regenerate.

B2/M4 — the discriminator was `/of\s+(\d+)/`, unanchored, so a total could be
read out of prose. `Current Plan: 4 — blocked on review of 2 PRs` parsed as
`4 of 2`, took the `currentPlan >= totalPlans` branch and WROTE
`Status: Phase complete — ready for verification` into the user's file. Both
shapes are now anchored at the start and every number comes from a capture
group via `planNumberFrom`, which rejects anything past `Number.MAX_SAFE_INTEGER`
rather than letting `data` and the persisted string disagree. Nothing on this
path calls `parseInt` on a raw field value any more.

The grammar keeps a trailing remainder after the total, because
`Plan: 2 of 5 in current phase` is a real tested shape. The refusal comes from
requiring `of <total>` to follow the leading number immediately, not from
forbidding a suffix.

M1 — `bumpLeadingNumber` is total. It returned its input unchanged when there
were no leading digits, so `+2` reported `advanced: true` while writing the
file untouched.

M2 — both section arms use replacer functions. File-derived text was being
spliced into a `String.replace` replacement string, where `$&` / `` $` `` /
`$'` expand: a value of `04 of 06 $&` spliced part of the document into itself.
`stateReplaceField` already used a function; these now agree with it.

M3 — the section arm targets the name the SECTION carries, and the body write
now writes both spellings, each with its own rendering. Keying off the header's
name left the other name stale in both directions: a legacy header beside a
`Current Plan:` section line, and a `**Plan:**` header beside one.

* fix(#3784): derive the shape error from the schema, widen the test coverage

Review round 3 on #3791: B3, m1, m2, m5 and the two test nits.

B3 — the accepted-shape set had two owners: the parser branches and an English
list hand-written beside them in `state.cts`. Nothing coupled them, so adding a
branch left the message stale and removing one left it advertising a shape that
errors, with no test able to see either. The message is now built from
`STATE_FIELD_SCHEMA.current_plan.acceptedShapes`, and the CLI test walks the
schema instead of restating the list. `Plan: N of M` is still spelled out
explicitly because no schema row owns the body-only `Plan` field —
`buildStateFrontmatter` never reads it into frontmatter, so it has no key to
hang a row on.

m1 — the property drove only the pre-existing `**Plan:**` branch, i.e. not the
branch under review. It now drives both compound spellings and ranges past 99
so the width transition is covered by the property rather than one example. A
second property covers the legacy pair's own preservation contract. Both were
mutation-checked: dropping the padStart turns 9 tests red.

m2 — degenerate boundary fixtures around the threshold (`0 of 0` is
phase-complete, not an error; `0 of 3` advances) plus the shapes the anchored
grammar must refuse, including Arabic-Indic digits.

m5 — `docs/json-errors.md` described rather than quoted the message, since it
is now schema-derived and a verbatim quote would be a third owner.

Nits — the CRLF assertion could not see a `\n` at index 0; the
`!/^Plan:/m` presence proxy is now an identity assertion on the whole
`## Current Position` body.

* fix(#3784): give the section plan write its own flag, and stop narrowing what parses

Review round 4 on #3791: Blockers 1-4, Majors 1-2, Minors 1-2.

B1 — the section fallback was guarded by `!mutated`, and `mutated` is
FUNCTION-wide, already set by the phase/status/lastActivity arms that
`advancePlanCore` always populates. A section spelling the field bold or as a
pipe-table row therefore skipped its fallback because an UNRELATED field had
been refreshed, and stayed a plan behind the header — the split-brain document
this arm exists to prevent. The arm now tracks its own `planWritten`.

Worth recording: the reviewer's fixture does not reproduce. The body-level
status write lands on the section's own `Status:` when the document has no
header `Status:`, so `mutated` is still false by the time the plan arm runs and
the fallback fires. The discriminating shape needs a header `Status:` to absorb
that write AND a bold section plan line. The mechanism was right; the example
was not, and the regression test uses the shape that actually fails.

B2 — `fallbackName` chose one name by ternary. In the legacy shape both values
are populated, so it always chose `Current Plan` and a `**Plan:**` section line
— which base did write — got nothing. Each name is now attempted independently
with its own fallback.

B3 — `PLAN_SHAPE_N` was anchored harder than `PLAN_SHAPE_N_OF_M`, so values base
parsed via `parseInt` began to hard-error: `Total Plans in Phase: 5 phases`,
`Current Plan: 3 (blocked)`. #3784's brief puts normalizing plan numbers beyond
this transition's read/write out of scope, so that narrowing was not licensed.
Both grammars now carry the same trailing tolerance. The prose defect stays
closed by the START anchor, not by forbidding suffixes.

Major 1 — the whole-body `Plan` write is scoped to documents that declare a
`Plan` field, instead of firing unconditionally where `stateReplaceField`'s
first match could be prose outside `## Current Position`.

Major 2 — the error message names both `Plan` spellings the parser accepts; it
previously omitted the sibling-paired form, which is the same
message-disagrees-with-parser drift the derivation exists to close.

B4 — the changeset claimed a guarantee B1 broke; it now describes what ships.

Minors — safe-integer boundary coverage at limit-1/limit/limit+1, and the CRLF
comment states the real mechanism (`stateExtractField`'s `(.+)` stops before the
CR; the trailing group is belt-and-braces, not the primary defence).

All three blocker regression tests verified red against the pre-fix source.

* test(#3784): pin the hybrid shape against #3807's ambiguity refusal

#4028 landed `advance-plan`'s multi-`Phase:` refusal on `next` after this
branch's last run, on the same function. The guard sits above the parse, so
a refused document is never parsed and the shape #3784 adds cannot reach the
mutation — but that is a property of source ordering, so assert it as
behaviour instead.

Fail-first proven, not assumed: with `phaseCandidates.length > 1` disabled,
the ambiguous hybrid document advances its FIRST entry's
`Current Plan: 04 of 06` to `05 of 06` and writes it — #3807's exact defect,
reached through #3784's shape. Both tests go red; both go green with the
guard restored.

The control pins the other direction: an unambiguous hybrid section still
advances, and its zero-padding still survives.

* fix(#3784): advance every spelling from its own text, refuse when they disagree

Round 6 review. B1 and M1 are one defect, so they are one fix.

`advancePlanCore` picked one field to parse from, computed `newPlan`, then
wrote BOTH spellings from that field's numbers. Two symptoms:

  B1  With `Plan` as the parse source, `Current Plan` was re-stamped with
      the number just derived from `Plan`. `Current Plan: 7` beside
      `Plan: 2 of 5` silently became `Current Plan: 3` — a value nothing
      derived for that field, no error, no diagnostic.
  M1  With the legacy pair winning, the `Plan:` line was re-rendered from
      a bare `${newPlan} of ${totalPlans}` built out of the sibling field.
      `Plan: 2 of 9` became `3 of 5`; `Plan: 03 of 05` became `4 of 5`.
      The changeset's claim that padding and everything after it survive
      was true only for whichever field happened to be the parse source.

Now: every spelling is advanced from its own raw text via
`bumpLeadingNumber`, so each keeps its own padding, its own total and its
own trailing annotation. Differing TOTALS are preserved, not reconciled —
`Plan: 2 of 9` beside a `Total Plans in Phase: 5` advances to `3 of 9`.

Differing CURRENT numbers are refused, with `reason:
"ambiguous_plan_position"` and both candidates named. Same posture as
#3807's multi-`Phase:` guard one field over: name the conflict, let the
caller resolve it, never pick. The guard sits immediately after the parse,
BEFORE the phase-complete branch — guarding only the normal advance would
let `Current Plan: 7` beside `Plan: 5 of 5` write a terminal "Phase
complete" into a document whose two spellings never agreed.

A field present but unreadable (`Plan: TBD`) is left exactly as authored.
Refusing the whole document because an unrelated line cannot be read would
be a narrowing #3784 does not license; writing a derived number over it is
the fabrication B1 was filed for.

The `planSourceField`/`planRawValue`/`useCompoundFormat` tracking is gone.
It existed only so the write path could ask which field the value came
from, and the write path no longer asks.

M2. The bare `Plan: N` + `Total Plans in Phase: M` shape is dropped. A
revision of this PR added it; base refused it. It cannot be given the
schema-row + forcing-test coupling the other shapes have, because `Plan`
is body-only and `buildStateFrontmatter` never reads it into frontmatter,
so there is no `current_*` key to hang a row on. Parser, the spelling in
`advancePlanShapeError`, and the lockstep test move together — the
invariant is the lockstep, not the length of the list.

N1. The whitespace narrowing (`5phases` no longer parses where `parseInt`
read 5) is documented in the changeset beside the other deliberate
narrowings, rather than loosened. Loosening restores the half-parse this
change exists to remove.

Tests: eight new cases plus a property that crosses the two spellings with
agreeing and disagreeing numbers — the review noted the existing
properties never did. Fail-first proven: restoring the old write path
reddens seven of the eight, both new property arms, and two pre-existing
padding tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

* fix(#3784): report Current Plan as updated only when it was written

The write became conditional in the previous commit — a `Current Plan:`
that is present but unreadable is left as authored — but the `updated`
push stayed unconditional, so `transitionCore` reported a field it had not
touched. `reconcileReportedFields` would have caught it against the
persisted bytes at the `state.cts` caller, but `transitionCore`'s own
`updated` is consumed directly (milestone-lock, the transition tests) and
has to be true on its own.

Covers the mirror of the unreadable-spelling case: `Current Plan: TBD`
beside a readable `Plan: 2 of 5`, where `Plan` is the parse source and the
legacy field is the one that cannot advance. Fail-first proven.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

* test(#3784): account for the new refusal in the output({error}) census

`tests/io.test.cjs`' A3 census asserts the exact population of
`output({error})` call sites in `src/`, per module. The
`ambiguous_plan_position` refusal added a 27th to `state.cts`, so the
census went red at 26/65.

Updated the way #3807 updated it when it added the ambiguous-POSITION
error one line above: bump the count and name the addition inline, so the
next person reads why the number is what it is. The alarm did its job —
it is the only gate that noticed a new user-visible error path had been
introduced.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-01 15:03:20 -04:00

16 KiB

JSON Error Mode — gsd-tools Structured Errors

Overview

gsd-tools supports a JSON error mode that emits most errors as structured JSON objects on stderr instead of free-form text. This is the recommended surface for tests and tooling that need to assert on error types without grepping raw text (see CONTRIBUTING.md — "Prohibited: Raw Text Matching on Test Outputs"). Usage errors are an intentional exception — see the ExitError carve-out below.

This page describes one of two failure channels. A second, equally intentional one reports conditions in the result payload on stdout with exit 0. A caller that branches on exit status alone will not see it. Read Degraded results vs faults before writing anything that consumes gsd-tools output.

Activating

Either flag or env var activates the mode:

# Flag (preferred in test code):
node gsd-tools.cjs --json-errors <command> [args]

# Env var (preferred for shell wrappers and CI):
GSD_JSON_ERRORS=1 node gsd-tools.cjs <command> [args]

Wire format

On any error, exactly one JSON line is written to stderr and the process exits with code 1:

{ "ok": false, "reason": "<error_code>", "message": "<human text>" }

Fields:

Field Type Description
ok false Always false for error objects.
reason string Typed reason code from the taxonomy below.
message string Human-readable description (may change; do not assert on it).

ExitError carve-out (plain text, not JSON)

Usage errors and explicit exit-code signals take a different path: they throw ExitError (src/cli-exit.cts), which runMain catches before the JSON-envelope branch. An ExitError writes its message as plain text to stderr (not a JSON object) and exits with the error's own code (which may differ from 1). This is intentional — usage messages are operator-facing prose, not structured failures.

If you are testing a usage/flag error, do not parse stderr as JSON; assert on the exit code and (if needed) the plain-text message. The "parse stderr as JSON" guidance below applies only to the structured-envelope branch (non-ExitError failures).

Which tools honor this. Both surfaces that run runMain do: the compiled gsd-core/bin/lib/cli-exit.cjs and the scripts/lib/cli-exit.cjs that the repo's own scripts/** tooling requires. Before #3904 the latter was a separate hand-written copy that never gained the structured-envelope branch, so a scripts/-side tool failing unexpectedly printed a raw stack trace even under --json-errors. It is now generated from the same source and byte-compared by npm run lint:generated-sync, so the two cannot answer differently again.

Degraded results vs faults — read this before writing a caller

gsd-tools has two ways of telling you something went wrong, and they use different exit codes. The wire format above describes only one of them. If you write a caller that branches on exit status alone, you will silently miss the other.

Fault Degraded result
Produced by error(message, reason) output({ error: … })
Stream stderr stdout
Exit code 1 0
Shape { "ok": false, "reason": …, "message": … } the command's ordinary result object, with an added error key
Honors --json-errors yes no — it is a payload, not an error envelope
How a caller detects it exit code inspect the payload

A degraded result means: the command ran to completion and is reporting a condition through its result. It is not a process failure. The command succeeded at the job of determining that, for example, the artifact you asked about is absent.

$ gsd-tools state-snapshot          # in a project with no STATE.md
{
  "error": "STATE.md not found"
}
$ echo $?
0

Some verbs return a companion result alongside the key, which is the shape that makes the intent clearest:

$ gsd-tools roadmap get-phase --phase 1      # no ROADMAP.md
{
  "found": false,
  "error": "ROADMAP.md not found"
}
$ echo $?
0

This is a ratified contract, not an accident — see ADR-2980 for the decision and the blast radius that drove it. It applies to 64 call sites across nine modules — state, verify, workstream, frontmatter, commands, template, phase, roadmap, and gsd2-import. (Issues #2966 and #2980 record this as "42 sites"; that figure counts only the sites where error happens to be the object's first key. ADR-2980 itself re-derived the population as "60" by brace-matching; a further AST re-measure for #3912 found the true current count is 64 — the same nine modules, with frontmatter, phase, and roadmap each having grown since. See ADR-2980's amendment for the breakdown.)

Writing a correct caller

The obvious shell form is wrong for a degraded result:

# WRONG — the process exits 0, so this branch never runs
if ! gsd-tools state-snapshot > snap.json; then
  echo "failed"
fi

Check both channels — the exit code for faults, the payload for degraded results:

if ! out=$(gsd-tools state-snapshot); then
  echo "fault (exit non-zero)" >&2      # error() path
  exit 1
fi
if err=$(printf '%s' "$out" | jq -er '.error // empty'); then
  echo "degraded: $err" >&2             # output({error}) path
fi

Four things that will surprise you

  1. --json-errors does nothing here. It governs error() only. A degraded result is byte-identical with and without the flag, and still exits 0.
  2. --raw is not uniform on this path. Most sites pass no raw value, so --raw still yields the JSON object rather than bare text — but eleven sites do pass one and behave differently. Do not infer either behavior from --raw alone; check the verb.
  3. Not every degraded result is an absent artifact. A missing required argument is reported the same way — gsd-tools state add-blocker with no --text returns {"error":"text required"} and exits 0. So is unusable input: gsd-tools state advance-plan against a STATE.md it cannot parse returns an {"error": …} naming the plan-position shapes it accepts (the list is derived from STATE_FIELD_SCHEMA.current_plan.acceptedShapes, so do not quote it verbatim), also exit 0. The exit code does not distinguish absent from malformed from misinvoked — see ADR-2980's Consequences, where this is recorded as a known cost.
  4. message/error text is not stable. Assert on structure and on typed reason codes, never on prose. The rule in "Writing tests" below applies to both paths.

Which one should new code use?

Prefer the fault path, or a result with a named field. ADR-2980 ratifies an existing population; it is not a license to add a 61st output({ error: … }) site. Where a verb needs to report a non-fatal condition in its payload, prefer the shape state update-progress already uses — a named field plus a reason, with no overloaded error key:

$ gsd-tools state update-progress            # STATE.md present, no Progress field
{
  "updated": false,
  "reason": "Progress field not found in STATE.md"
}

Outcome declaration and the versioned exit contract (ADR-3889 §4, #3912)

Both failure channels above now declare an outcome on every terminating path, per ADR-3889. Declaration is unconditional; whether it changes the observed exit code depends on which exit-contract version the process is running under.

Turn on v2 with either --exit-contract=v2 or GSD_EXIT_CONTRACT=v2 (a flag beats the env var if both are given). Absent either, the process runs v1 — today's default and, for every existing caller, byte-identical to pre-#3912 behavior. See Adopt the v2 exit contract for a worked migration.

error(message, reason)

error()'s reason argument now maps onto a declared outcome name (USAGE, NO_INPUT, UNAVAILABLE, INTERNAL, or FAIL) via a fixed table over all 25 ERROR_REASON members.

  • Under v1, the mapping is recorded but never projected. error() still throws ExitError(1) unconditionally, exactly as before — stderr and the exit code are byte-identical to every prior release.
  • Under v2, the mapping is projected through the exit-code registry. error() throws ExitError(exitCodeFor(<mapped outcome>)) instead of a hardcoded 1 — so, for example, a call with ERROR_REASON.SDK_MISSING_ARG or ERROR_REASON.SDK_UNKNOWN_COMMAND exits 64 (USAGE) under v2, and one with ERROR_REASON.CONFIG_KEY_NOT_FOUND exits 66 (NO_INPUT).
  • Most call sites are unaffected either way. 226 of the 278 error() call sites in the repo pass no reason at all, defaulting to ERROR_REASON.UNKNOWN, which maps to the generic FAIL outcome (exit 1) under both versions.

output({ error: … }) — a degraded result is also a declared outcome

The degraded-result idiom above now declares the outcome DEGRADED whenever output()'s payload carries a serializable error value — any key order, and regardless of that value's own truthiness (0/null/'' all count). The one exclusion: { error: undefined } does not declare DEGRADED, because JSON.stringify (the exact serializer output() uses) drops an object property whose value is undefined before it ever reaches the wire — a payload built that way reaches the caller as {}, with nothing to be degraded about.

  • Under v1, DEGRADED projects to 0 — deliberately: this is ADR-2980's compatibility boundary, pinned so all 64 ratified sites keep exiting 0 byte-for-byte.
  • Under v2, DEGRADED projects to 80 — looked up from the exit-code registry, never hardcoded, so a future re-allocation of DEGRADED's number cannot silently desync this doc from the shipped table.

Precedence — what code a void-returning command actually exits with

A command's main() can end up producing a code from more than one source. The order, highest precedence first, is:

  1. An explicit main() return (a number or a registered outcome-name string) — always wins.
  2. A non-zero process.exitCode main() already set directly before returning — wins over anything declared through output(). This is what keeps state validate --strict correct: it sets process.exitCode = 1 itself on a missing STATE.md, and a DEGRADED declared earlier in the same call must not clobber that 1 back down to DEGRADED's v1 projection of 0.
  3. The declared outcome pending from output() — consulted only when neither of the above set anything.
  4. Otherwise the process exits 0.

Projection may only ever set a code, never lower one. A prior review pass concluded the pending declaration was fail-closed by construction; it was not — without rule 2 above, state validate --strict briefly exited 0 on a case that must exit 1. If you add a new call path that sets process.exitCode directly, check it still wins over a later output({error}) in the same invocation.

The declaration does not accumulate across calls. output()'s declaration follows last-write-wins: a clean payload clears a prior DEGRADED declaration in the same invocation, and runMain clears the cell on every exit regardless of which branch produced the final code, so a later runMain call in the same process never inherits a stale declaration.

Error code taxonomy

Codes are frozen constants in gsd-core/bin/lib/core.cjs under ERROR_REASON. Tests must assert on reason values (stable), not message text (unstable).

Dispatch errors (gsd-tools routing layer)

Code When emitted
sdk_unknown_command Unknown top-level command (gsd-tools bogus-cmd)
sdk_unknown_command Unknown dotted command (gsd-tools foo.bar where foo is not a known command)
sdk_unknown_command Unknown subcommand within a domain (e.g. gsd-tools intel bogus-sub)
sdk_missing_arg Required argument omitted by an SDK-level guard
sdk_fail_fast SDK fail-fast policy triggered

Usage / flag errors

Code When emitted
usage --pick flag used without a following value
usage Version flag (--version, -v) which gsd-tools never accepts
usage Top-level no-args invocation (usage text)

--pick <field> errors (ADR-3473 §8.4, #3884)

Code When emitted
pick_field_absent --pick <field> names a field that does not exist in the command's JSON output (missing key, out-of-range index, a partially-missing dotted path, or a non-object JSON root) — see CLI-TOOLS.md's --pick contract
pick_output_not_json --pick <field> is combined with a command whose output is not JSON (including --raw output)

Config errors (config-get, config-set, config-ensure-section)

Code When emitted
config_key_not_found config-get for a key that is absent from the config file
config_no_file Config operation when .planning/config.json does not exist
config_parse_failed Config file exists but is not valid JSON
config_invalid_key config-set for a key outside the allowed whitelist

Phase / workflow errors

Code When emitted
phase_not_found Phase directory lookup returns no match
summary_no_planning Summary operation when no .planning/ directory exists

Estimate errors

Code When emitted
estimate_phases_unreadable estimate-calibrate when .planning/phases/ exists but could not be read (EACCES/EIO) — refused rather than silently rebuilding calibration from a phantom empty sample set (#3882, ADR-3473 §8.5)

Graphify errors

Code When emitted
graphify_no_graph Graphify query or diff when no graph has been built
graphify_invalid_query Graphify query with a malformed query string

Hook / security errors

Code When emitted
hooks_opt_out Hooks are disabled via opt-out config
security_scan_failed Security scan produced a finding that blocks the operation

Fallback

Code When emitted
unknown All other errors without a specific reason code assigned

Writing tests

For non-usage errors (the structured-envelope branch), parse stderr with JSON.parse and assert on typed fields. Never use .includes(), .match(), or regex on the raw error string.

// CORRECT: parse then assert on typed field
const result = runGsdTools(['--json-errors', 'bogus-command'], tmpDir);
assert.strictEqual(result.success, false);
const err = JSON.parse(result.error);
assert.strictEqual(err.ok, false);
assert.strictEqual(err.reason, 'sdk_unknown_command');

// WRONG: text matching (banned by lint-no-source-grep policy)
// assert.ok(result.error.includes('Unknown command'));

Adding a new error code

  1. Add the constant to ERROR_REASON in gsd-core/bin/lib/core.cjs (snake_case, prefixed by subsystem).
  2. Pass it as the second argument to error() at the call site.
  3. Add a row to this document.
  4. Add a test asserting the new reason code via JSON.parse.