Files
msd-core/docs/how-to/change-the-state-md-schema.md
Tom Boucher ddde001af6 enhance(#3873): the STATE.md schema — one owner, generated artifacts (#3880)
* test(#3873): failing-first locale parity, plus tripwires for what must not move

Pins ADR-3473 §8.8 at the artifact a reader actually sees. The English STATE.md
reference carries a Status lifecycle section that is missing from all four
translations — the section documenting the status enum whose clobbering is
#3853. The test derives the heading set rather than hard-coding the missing
one, and names the locale and the heading when it fails.

Two tripwires that must pass today and after. The field-drift guard still
catches a re-derived fallback ladder: §8.8 instructs deleting that script, and
that instruction rests on a wrong premise about what it guards, so the test
stops a future reader from deleting it on the ADR's word. And last_activity's
label resolution is pinned to what ships today, because it is declared in one
of the two tables this phase consolidates and not the other — the
consolidation must not silently pick a side.

The locale test buckets under docs rather than state, which is what it tests;
that bucket is allowlisted with justification rather than folded into an
unrelated docs suite. It reads only markdown, so it carries no allow-test-rule
marker — a marker there would suppress nothing and would grow the unverified
pool against its ceiling.

Refs #3873

* feat(#3873): one schema owns the STATE.md key set, three tables become projections

ADR-3473 §8.8. The key set was declared in four places that had to agree by
hand and already did not: FIELD_CLASSIFICATION, FRONTMATTER_BODY_SOURCE,
FRONTMATTER_KEY_TO_BODY_LABEL and buildStateFrontmatter's emit behavior. One
frozen null-prototype schema now declares each key's type, enum, cardinality,
source, preservation, body source, body label, accepted parse shapes and
whether it is emitted unconditionally; the three tables are derived from it at
module load.

The projections are byte-identical to the literals they replace, key order
included, and the parity tests compare against verbatim copies of today's
tables rather than re-deriving both sides from the schema — a parity test fed
from one source proves nothing, which is how a consolidation ships a changed
policy under a green test.

last_activity was the live disagreement: present in one table, absent from the
other. The schema declares what ships today rather than the tidier answer, and
a test pins it.

The schema is a leaf module and owns the four field-policy types, re-exported
from state-transition so existing importers are untouched — the same split
health-diagnostic-types made to break a CJS require cycle.

Refs #3873

* feat(#3873): generate the schema-derived regions, parity-check the prose tables

ADR-3473 §8.8's generator half. gen-state-md-docs.cjs owns marked regions in
the shipped template and all five reference docs, follows gen-features.cjs's
fail-closed contract, and is wired into regen:derived and lint:generated-sync.

The Status lifecycle section was missing from all four translations — the
section documenting the status enum behind #3853 — and is now generated into
every locale. Field cardinality is a new generated table: pure schema data,
no prose, so nothing to lose.

The Field-reference and Status-values tables are parity-CHECKED rather than
generated. Their Purpose, When-populated and Matched-text columns are
genuinely hand-translated per locale, and §8.8 itself says prose stays
hand-translated; generating them from an English registry would overwrite four
locales' translations on every write. The row set is checked against the schema
instead, so a key added to one and not the other fails, which is what field
drift actually means. Building that check found last_activity_desc
undocumented in all five tables.

Three keys the docs describe are absent from the schema — active_phase,
next_action, next_phases. They are grandfathered by name, not by wildcard, so a
fourth fails: a declared gap with a forcing function rather than a silent one.

Refs #3873

* fix(#3873): declare what the parsers do, and close the shape-parity gap

Two declarations in the new schema described intended behavior rather than
actual — the defect class this epic exists to end, committed inside the epic.
Both were caught by executing the parsers instead of reading their docstrings.

current_plan.acceptedShapes claimed ['N', 'N of M']. Standalone, the hybrid
shape errors; the path that looks like support is parseInt truncating '2 of 5'
to 2 and discarding the rest. Narrowed to ['N']. The parser is deliberately NOT
fixed here: that is #3784 and PR #3791 is already doing it. When #3791 lands
this row must widen, and the shape test will go red until it does — the schema
and the parser cannot drift apart quietly, which is what §8.8's checked-not-
generated rule is for.

STATUS_LIFECYCLE_ENUM claimed to be the closed set status can hold.
normalizeStateStatus passes unrecognized prose through unchanged, so it is not
closed at runtime. The seven members are the canonical values it maps onto; the
docstring now says that and the test asserts the real lenient contract.

Closes the acceptance item that a test asserts the parsers accept exactly the
declared shapes: the check is table-driven over every row carrying
acceptedShapes, guarded against passing vacuously on an empty set, and fails
loudly if a future row has no registered driver. Adds the unwired-label throw
and the fast-check property that every projection agrees with its schema row.

Refs #3873

* fix(#3873): keep the shipped template's frontmatter first, and make row 27 able to fail

The remote matrix caught 12 failures with one cause. Making the template's
frontmatter a generated region wrapped it in its own yaml fence ahead of the
markdown fence, so extractFileTemplate and readShippedStateTemplateBody — which
both match the single markdown block — found the heading first, not the
frontmatter. That breaks the contract every new project's STATE.md is created
from: bug #21 and epic #1969 B8 pin that the File Template block starts with
frontmatter and carries gsd_state_version.

The markers now sit inside the single markdown fence, so the fence opens before
the frontmatter and the region still ends ahead of the heading. Same layout as
before this phase, with markers embedded rather than a second fence.

Row 27 existed to catch exactly this and did not, because it was writer-seeded:
it asserted against the generator's own output shape, so it passed on the broken
template. It now parses the fence the way production does and was verified to
fail against the broken shape before being trusted against the fixed one. A test
that would not have caught the bug it exists to prevent is worse than no test.

The emitted-attribution failure was separate and the fragment was the wrong
remedy: gsd-core/templates/state.md self-attributes under a verbatim-copy
identity rule, so a diff touching it needs no acknowledgment. Fragment deleted
rather than left explaining nothing.

Refs #3873

* docs(#3873): how to change the STATE.md schema

The phase gate was right and my docs artifact was wrong. I listed
lint:generated-sync as the second enablement step, which is a verification
command dressed as one, and then claimed a one-step sequence owed no how-to.

The real sequence is build:lib then regen:derived, and the ordering is a trap:
the generator reads the COMPILED schema, so regenerating before building
regenerates against the previous schema and commits artifacts that look
plausible while disagreeing with the code just written. A reference table
cannot carry an ordering dependency; that is what the how-to test is for.

The page covers adding, changing and removing a key, every reason code the
check emits and what to do about each, what is generated versus hand-translated
and why the two prose-bearing tables are parity-checked instead of generated,
adding a language, and the three grandfathered keys. Indexed from docs/README.md.

Refs #3873

* chore(#3873): backfill changeset PR number

---------

Co-authored-by: sim <sim@local>
2026-08-26 01:57:47 -04:00

6.4 KiB

How to change the STATE.md schema

Every key in .planning/STATE.md's frontmatter is declared once, in src/state-md-schema.cts. The field-classification tables, the shipped template and the five reference documents are all derived from it. This page is how you add, change or remove a key without any of those falling out of step.

If you only want to know what the keys are, read the STATE.md reference instead.

Add a key

1. Declare it in src/state-md-schema.cts.

my_new_key: {
  type: 'string',
  cardinality: 'optional',
  source: 'body',
  preservation: 'preserve-when-unchanged',
  bodySource: ['My New Key'],
  bodyLabel: 'My New Key',
  emitted: 'when-present',
},

Every column is required except enum, guard, mergeStrategy, bodySource, bodyLabel and acceptedShapes. What each means is documented on the type itself — read it there rather than copying a neighbouring row and hoping.

2. Build, then regenerate. In that order.

npm run build:lib
npm run regen:derived

The order matters and getting it wrong fails quietly. The generator reads the compiled gsd-core/bin/lib/state-md-schema.cjs, not the TypeScript source. Regenerating before building regenerates against the previous schema, produces artifacts that look plausible, and commits a document that disagrees with the code you just wrote. If you are ever unsure whether the build is current, run npm run build:lib again — it is cheap and idempotent.

3. Commit the regenerated artifacts. They are generated and committed:

  • gsd-core/templates/state.md
  • docs/reference/state-md.md and its ja-JP, zh-CN, ko-KR, pt-BR siblings

4. Add the human-facing rows by hand. Two tables in the reference documents are deliberately not generated — see What is generated and what is not. Add your key's row to the Field reference table in each locale. The parity check will tell you if you miss one.

Change or remove a key

Same two commands. Removing a key also means removing its row from the Field-reference tables in all five locales, or the parity check fails naming each one.

Before you change a key's preservation or source, read ADR-3408 §8 — those columns drive what survives a STATE.md write, and a change there is a behavior change, not a documentation edit.

What the check is telling you

npm run lint:generated-sync runs node scripts/gen-state-md-docs.cjs --check, and lint:ci runs it for you. It exits non-zero with a reason code and the file and region involved.

Reason What happened What to do
region_stale A generated region does not match what the schema would produce. npm run build:lib && npm run regen:derived, then commit the result.
markers_missing A target file has no STATE-MD-SCHEMA marker pair for a region. Add the marker pair where the region belongs. The generator never invents a location.
marker_unclosed A :START: marker has no matching :END:. Fix the markers. The generator refuses to write rather than guess where the region ends — a wrong guess would eat hand-written prose.
field_reference_drift The Field-reference table's row set disagrees with the schema. Add the missing row, or remove the row for a key that no longer exists.
status_values_drift The Status-values table disagrees with the schema's status enum. Same.

--json gives you the same information structurally if you are scripting against it.

What is generated and what is not

Region Generated?
The template's frontmatter block yes
### Status lifecycle yes, in all five locales
### Field cardinality yes, in all five locales
Field reference table no — row set parity-checked only
Status values table no — row set parity-checked only
All prose outside a marked region no, ever

The last three are the point. The Field-reference and Status-values tables carry per-row prose — Purpose, When populated, Matched text — that is genuinely hand-translated. The Japanese Matched-text column reads `discussing` を含む, not the English. Generating those tables from a single English source would overwrite four languages' translations every time anyone regenerated. So the schema owns the row set, which is what drift actually means, and translators own the prose.

If you edit inside a marked region, the next --write will overwrite you and --check will report it first. If you edit outside one, nothing touches it.

Adding a language

Copy an existing locale's reference/state-md.md, translate the prose, and keep the STATE-MD-SCHEMA marker pairs where they are. Then run the two commands above; the generator fills every marked region for the new locale, and the parity check starts holding it to the same row set as the rest.

Column headers come from a per-locale string table in the generator — add yours there so the generated tables are not headed in English.

Keys the schema does not model

active_phase, next_action and next_phases are real frontmatter keys (#2833) that are documented but sit outside the schema. They are grandfathered by exact name in KNOWN_SCHEMA_GAP_FIELDS, so the parity check tolerates those three and no others — a fourth undocumented key fails, which is what stops the list quietly becoming a wildcard.

If you are adding one of those three to the schema properly, remove its name from that list in the same change.

Why the schema exists

Before it, this key set was declared in four places that had to agree by hand, and they did not: one table carried last_activity, another did not. The five reference documents disagreed too — the section documenting the status values was missing from all four translations, and it is the section that matters for #3853.

ADR-3473 §8.8 is the contract, and its governing idea is worth keeping in mind when you edit the schema: the schema declares what the code does, not what it should do. If you find yourself writing a row that describes intended behavior, you are writing a document that lies — declare today's behavior and fix the code separately.