* test(#3007): failing-first suite for per-model Codex effort capability RED by construction. Binds to behavior renderEffortForRuntime does not yet have: an optional third `model` argument, a per-model advertised-level table, `max` passing through instead of clamping to `xhigh`, `minimal` clamping to `low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/ `reason`) so a downgrade is legible from resolver output rather than silent. Two of these pin defects that exist on next today: - `max` is discarded. Both Codex models whose catalog entries are retrievable (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing. - `minimal` is emitted to a model that refuses it. providerPresets.openai. haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's advertised floor is `low`. GSD is sending a value into a document Codex itself validates. The parity test is what pins that fixed, and it names the offending path/model/effort when it trips. Also corrects tests/model-resolver.test.cjs:351, which asserted renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as though it were a contract. ADR-443 recorded "Codex has no max" as fact and it was true when written; Codex has since added both `max` and `ultra`. That is a stale premise, so the assertion is corrected here rather than worked around. The property test asserts the invariant the whole change exists for: a rendered effort is always a level the target model actually advertises, or an explicit rejection. There is no third outcome. * fix(#3007): resolve Codex effort per model, and make every clamp visible Codex declares supported_reasoning_levels per MODEL and validates against it, so a single per-runtime capability set cannot be right for all of them. GSD's was wrong in both directions at once. `max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped max -> xhigh on that basis; it was accurate when written, and Codex has since added both `max` and `ultra`. Every Codex model whose catalog entry is retrievable advertises `max`, so the clamp was discarding a level the provider supports, silently, on the most-used path. `minimal` stops reaching Codex. No Codex model advertises it -- both retrievable entries floor at `low` -- yet providerPresets.openai.haiku.low paired gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the receiver validates and refuses into a file the receiver reads. Being unconservative in what you send is the half of Postel's rule with no defensible reading, so that preset is corrected and a parity test pins it. `ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum reasoning with automatic task delegation": at ultra, effective_multi_agent_mode returns Proactive and Codex spawns sub-agents on its own initiative, underneath GSD's orchestration rather than inside it (#2167). It is a mode switch, not a reasoning depth, so it is not added to the universal ladder -- which stays provider-agnostic by ADR-443's design -- and it is rejected even for gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered and rejected: that silently discards what the user actually asked for. Clamping is now visible. RenderedEffort carries requested/clamped/reason and resolve-execution surfaces them. The previous table clamped correctly but invisibly, so a user asking for `max` on Codex had no way to find out they were getting `xhigh` -- exactly the failure mode the robustness principle's modern critique warns about, and why "be liberal" has to mean "liberal and loud". Also closes a latent trap found while reviewing the implementation: the clamp-up loop walks the ladder upward, and for a future model advertising `ultra` but not `max` it would have selected `ultra` as the clamp target -- re-entering by the back door the mode the rejection above exists to keep out. A clamp may never produce a value that a direct request for that value would refuse. Unreachable with today's catalog, which is why no test caught it; a test now asserts the invariant directly. Signature stability is preserved: the third `model` argument is optional and the two-argument form still resolves, against the family baseline. That form's BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old answer would have fixed the defect only where a model happened to be threaded through and left it live everywhere else. tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract and is corrected here rather than worked around. * fix(#3007): close every review finding on the Codex effort alignment Two isolated reviewers, correctness and security. Both found the same two blockers, and the per-model work was inert on every surface that matters until this commit. BLOCKER — resolve-execution never passed the model and discarded the clamp. cmdResolveExecution called the two-argument form and emitted only effort_rendered/effort_param/effort_propagation, so the per-model table was unreachable from production code (tests were its only caller) and requested/ clamped/reason were computed and thrown away. Requested outcome 3 names "the effective rendered effort in resolver output" specifically, so the feature was unmet on the exact surface the issue asks for. Now passes the resolved model and emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the existing key convention rather than introducing a nested object. BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed a nested {"effort": ...} sample; the real result is flat and those keys were absent entirely. A reference doc asserting a JSON path a reader can copy is worse than no doc. Corrected against the actual emitted key set. MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex kept minimal in its supported set and still clamped max down to xhigh, so the invocation-time and install-time channels disagreed about the same runtime's capability: --host codex with max emitted xhigh while the generated TOML said max. This is the repo's documented generative-fix-divergence class, so both tables now cross-reference each other and a parity test fails if they ever diverge again. MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null _baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback never fired and every effort rendered as null. And a non-array value made the Set constructor throw at module load — model-catalog.cjs is required across the whole CLI, so one bad JSON value killed every command, not just codex effort. Guarded on size and filtered to array values; both degrade to the hardcoded baseline. MAJOR — value widened to a nullable string with two consumers left behind. runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a null effort key in generated frontmatter); install-effort-resolver still declared a non-nullable return, a structural lie that silently defeated null checking. Both corrected, both omitting the key on null — the same posture as 'inherit', where omission means "follow the host default". MAJOR — the per-model table is inert today, and the docs now say so. All three shipped models advertise the same usable range and ultra (sol's only differentiator) is rejected for every model, so no observable output differs by model. The table stays because Codex declares capability per model and the sets are free to diverge — a single per-runtime assumption is precisely what went stale and produced this issue — but overselling it as a visible per-model feature would have been the same class of error as the doc blocker above. Tests: three passed under a full revert and are strengthened rather than deleted, since each guards a real contract (#3533's inherit rule, the undeclared-host rule, off-ladder handling) — they now also assert the clamp-visibility fields, which only exist after this change. The fast-check property is kept for its shrinking, and a deterministic nested loop over the full cross-product now sits beside it so coverage is exhaustive rather than sampled. Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg form and would have written a literal null reasoning effort on the ultra path; CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT. The installer defect was found by the co-change gate, not by a reviewer — install.js is a historical co-change partner of model-catalog.cts that this diff had not touched. * test(#3007): correct assertions that pinned Codex's stale effort premise Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the shipped commit. Every one is a stale pin, not a defect: each was probed against the built module before its expectation was changed, and none failed for a reason other than this premise correction. Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale by a production change must not ride inside another commit, because the release-sdk hotfix cherry-pick filter routes by subject prefix and a correction buried under the wrong prefix ships a half-state (v1.42.3, #3621). The most valuable one was tests/model-resolver.test.cjs's cross-provider validity invariant, which hardcoded the Codex enum as `minimal|low|medium|high|xhigh` and failed with "real API would 400". That message is now false in both directions: Codex accepts `max`, and rejects `minimal`, which no model advertises. The enum is corrected to `low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the "would the real API refuse this" check worth having, and it was right to fail here. It simply carried the stale fact in its own fixture. Test NAMES were corrected alongside their assertions wherever the name asserted the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal passthrough". A renamed test that still claims the old thing is worse than a failing one, and a green test whose name states a falsehood is how the next reader inherits the wrong premise. Both channels are covered: install-time (renderEffortForRuntime, and the generated .toml in install-runtime-artifacts) and invocation-time argv (effort-surface-axis). They were deliberately brought into agreement in this change, so their assertions had to move together. Each site carries a #3007 comment recording that Codex gained max/ultra and that capability is declared per model, so a future reader can tell this was a deliberate premise correction rather than a test bent to fit an implementation. * test(#3007): separate the effort-precedence case from the clamp case The previous stale-assertion pass over-corrected one test. It saw `effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`, assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to "max passes through" and changed the expectation to `max`. The remote runner disagreed. Reproduced against the real CLI: with that config and `gsd-planner`, the resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`, `effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE — gsd-planner is heavy/opus tier and its routing-tier default outranks `effort.default` — so `max` never reaches the renderer at all. The test says nothing about clamping and never did; it only looked like a clamp pin because both mechanisms happened to yield the same string. Restored to `xhigh` and renamed to say what it actually tests. It now also asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is what makes it impossible to mistake for a clamp pin again: those two fields prove the value is what the resolver produced rather than something the renderer downgraded. Before #3007 there was no way to tell the two apart from the output — which is precisely why the previous pass could not tell them apart either. Added the test that was actually missing: `effort.agent_overrides`, which outranks the tier default, so the requested level genuinely reaches the renderer and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified by probe before asserting. One test now pins the precedence rule and the other pins the #3007 behavior, and neither can be read as the other. That the clamp-visibility fields are what resolved this is a small argument for having added them. * chore(#3007): backfill changeset pr number to 3765 * test(#3007): put model-catalog under the mutation gate The Stryker shard showed as `skipping` on this PR despite the diff rewriting model-catalog's effort logic. That was legitimate, not a detection bug: `model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the whole module — including everything #3007 touches — sat entirely outside mutation scoring with has_work "false". Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no temp dirs. That shape is not stylistic — it is the #2790 precedent this file already documents. Stryker's command runner treats a whole `node --test <file>` invocation as ONE test costing whatever its slowest case costs, and re-runs it per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses runGsdTools throughout) would reproduce exactly the 15-minute shard-cap cancellation #2790 hit. The integration file is unaffected and keeps running in full in the normal test job. Coverage spans the module rather than only the diff, because the score is measured over the whole file: effort rendering across every model and ladder level in both channels, the prototype-chain host guard, the exported enums and maps, isAnthropicFlavoredModel's provider namespacings, the profile projections, nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are worth naming — every uncovered exported function is score given away, and mergeEffortTierDefaults turned out to have a genuinely interesting contract (#3531: a partial override merges over the built-ins rather than replacing them, and isValid gates the VALUE, not the tier name, so an unknown tier key is still merged in). Every expectation was probed against the built module before being asserted. minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment. Floors in this repo are measured, not chosen — the existing entries sit at 94, 75 and 56 — and they can only be measured in CI, because mutation shards run `node --test`, which is hard-blocked locally. The first CI run on this branch reports the real number and the floor gets ratcheted to it before merge. A placeholder of 1 reaching `next` would make the gate decorative: it would pass whether or not a single mutant is ever killed. Note the target is "never regress from measured", not a fixed 80 — planning-inspect sits at 56 and is documented as an accepted ratchet candidate. * test(#3007): bootstrap model-catalog's mutation floor legally The placeholder floor was structurally illegal and the remote run said so. tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module must carry a matching RATCHET_BASELINE entry in the same diff, minScore must EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three. That is the ratchet working exactly as intended — a floor nobody can satisfy accidentally is the point of it. Bootstrapped at 50 in both places. Fifty is not a measured score and the comment says so plainly: it is the minimum the guard permits, and it coincides with Stryker's own configured `break` threshold, so it is the lowest legal starting point for a module that has never been measured. It still must be ratcheted to floor(measured) - 1 before this PR merges. Also corrected a real defect in the file's own instructions. "HOW TO UPDATE" step 1 read "Run the per-module Stryker shard locally" — which cannot be done here, and which the same file contradicts eighty lines further down, where the #2790 scores are recorded as "not a local run; mutation shards run `node --test`, hard-blocked in this repo's local environment". stryker.config.mjs confirms the command runner invokes `node --test` once per mutant, and .claude/hooks/block-local-node-test.sh denies exactly that. So the documented first step sends the next contributor at a wall. Rewritten to describe the path that works — push, read the measured score off the CI shard, then set the floor and its baseline together in one diff — and to say why local measurement is not available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched. The two-step is inherent to the environment rather than a shortcut: a floor cannot be measured before the first CI run exists, and the guard rightly refuses to accept an unmeasured one below its minimum. * test(#3007): ratchet model-catalog's mutation floor to its measured score The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this file's own rule, minScore = floor(measured) - 1, which is the same arithmetic every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94. Both halves moved together, because the ratchet guard asserts minScore equals its RATCHET_BASELINE entry and would reject them drifting apart. The spawn-free unit surface is vindicated by the clock: 57 seconds, against a 15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing the shard at tests/model-resolver.test.cjs — #2790 recorded shards being CANCELLED at that cap when they targeted a runGsdTools-heavy integration file. The registry comment is rewritten rather than deleted. It previously warned that the floor was provisional and must not ship that way; leaving that text next to a measured floor would make the file lie in the other direction. It now records the measurement the way the sibling entries do, including that 59.62 sits below TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 — comfortably clear of its own floor with real room to grow. Raise it as the tests improve; never lower it. Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is an honest floor for a module that had NO mutation coverage at all an hour ago, and it is now pinned so it cannot silently regress. --------- Co-authored-by: sim <sim@local>
Architecture Decision Records
This directory contains Architecture Decision Records (ADRs) for GSD.
Each ADR documents one architectural decision: what was decided, why, and what consequences follow. ADRs are append-only. Amendments extend existing ADRs with a dated section rather than replacing them.
Reading this corpus
Start with the index below, and respect the status. The index is grouped so that the first table — Active decisions — is the set that governs the system as it stands. An ADR in Superseded, Retired, and Legacy is historical: it records what was once decided and names what replaced it. Do not cite it as current architecture.
Two things the index makes explicit, because getting them wrong has actually misled readers here:
- "Read first" on an active ADR points at a broader ADR that now frames it. A decision can be entirely correct and still not be the whole picture. The runtime capability descriptor (ADR-1016) is live and load-bearing, but ADR-1239 (EoS — GSD as an Embeddable Orchestration Engine) subsumes it as the declarative adapter and inverts its direction: GSD is the engine a host embeds, not an installer that projects onto a host. For how GSD meets a host, EoS is the current frame.
Proposedmeans not ratified — and it is kept honest. On 2026-07-17 the corpus was audited against the shipped tree and nine ADRs whose decisions had demonstrably shipped were ratified toAccepted, each carrying a dated Ratification section with the evidence (see ADR-857 for the fullest example). The ADRs that remainProposedareProposedfor a reason recorded in the file — an unmet acceptance criterion, an outstanding phase, or a successor ADR already planned — not through neglect. Trust the label; if you think it is wrong, prove it in a dated section and see Ratifying a staleProposed.
Naming Convention
New ADRs use issue#-prefix slug naming:
docs/adr/<issue#>-<kebab-slug>.md
Examples: 2264-golden-parity-redesign.md, 1239-gsd-embeddable-orchestration-engine.md.
Why
Two developers computing "next ADR number" locally against main will independently pick the same integer and both ship. The collision is already on disk — 0010-* exists twice and 0011-* exists three times. GitHub issue numbers are server-assigned and atomic: the moment you open an issue, that number is reserved globally. Two PRs that both edit the ### Fixed block of CHANGELOG.md always conflict on merge — two PRs that each use a distinct issue# as their ADR prefix never collide. Same shape, same solution.
Legacy naming is not Legacy status
Files 0001-* through 0012-* are preserved as immutable historical record of the old local-compute numbering. The duplicate 0010-* and the three-way 0011-* are documented residue of that convention — not patterns to imitate. Do not renumber them.
This is the single authoritative statement of the legacy range.
docs/contributor-standards.mdreferences it rather than restating it, so the two cannot drift.
Two other zero-padded files look legacy but are not: 0174-retire-gsd-sdk-package-boundary.md (issue #174) and 0656-research-module-seam.md (issue #656) are mis-padded modern ADRs — modern, issue-numbered files whose four-digit padding is a mistake. They are NOT part of the legacy sequential set above and are not "old local-compute numbering" residue.
This is a statement about filenames only. Many of those ADRs are Accepted and load-bearing today (ADR-0002, ADR-0004, ADR-0008, ADR-0009). An old filename says nothing about whether a decision still holds. The Legacy status in the table below is a separate claim — see the vocabulary.
Because 0010-* and 0011-* each resolve to more than one file, a bare cross-reference like "ADR-0011" is genuinely ambiguous. Link the file (see Lifecycle rules).
Full process
See CONTRIBUTING.md — "Proposing an ADR or PRD" for the end-to-end workflow: opening the issue, waiting for approval, naming the file, and submitting the PR.
PRDs live in docs/prd/, not here. (0011-review-default-reviewers-prd.md predates that directory and is kept in place as frozen historical record.)
Lifecycle rules
These are enforced by scripts/gen-adr-index.cjs, which runs in CI via npm run lint:generated-sync. A violation fails the build with the exact file and fix.
1. Every ADR declares one status from the canonical vocabulary
The first word of the Status field must be one of:
| Status | Means | Obligation |
|---|---|---|
Accepted |
Decided and in force. Cite it. | — |
Proposed |
Decided in principle, not ratified. Do not cite as settled. | If the work has demonstrably shipped, ratify it (below) — do not leave the label lying. |
Superseded |
A specific newer ADR replaced this decision. | Must name the successor as a file link. |
Retired |
What this ADR decided no longer exists at all, and no single ADR replaced it. | Say what was removed and when. |
Legacy |
Frozen historical record, kept for provenance; not a pattern to follow. | Say why it is frozen. |
Prose may follow the token (Superseded by [ADR-0174](0174-retire-gsd-sdk-package-boundary.md) (2026-05-23); originally Accepted (2026-05-09)). Both the bullet form (- **Status:** Accepted) and the table form (| **Status** | Accepted |) are accepted.
2. Cross-references to other ADRs are file links, never bare ids
Write [ADR-0011](0011-skill-surface-budget-module.md), not ADR-0011. Bare ids are ambiguous for 0010/0011, and unlinked references cannot be checked.
If you mean an issue, write #857 — not ADR-857. (An ADR and its owning issue often share a number; that is intentional and not a conflict.)
3. Supersession and subsumption are symmetric
These are different relations. Do not conflate them:
Supersedes/Superseded by— the target is replaced. Its status becomesSuperseded.Subsumes/Subsumed by— the target still holds, but a broader ADR now frames it. Its status is unchanged; it becomes a component of the larger decision.
If A declares either relation toward B, B must record the reciprocal. A one-way pointer is the failure this corpus actually suffered: ADR-1239 declared it subsumed four ADRs, none of which said so, and none of which pointed back — so a reader landing on any of them concluded the superseded frame was the way forward.
Only an Accepted ADR is owed the back-link. A Proposed ADR's claim is prospective: it has not taken effect, so its target is not marked. On ratification, the check begins demanding the back-links.
4. The declared id matches the filename
An H1 of # ADR-0175: … in a file named 218-*.md is a rename that never finished. The id in the title must match the filename's prefix.
5. A trailing H1 status bracket must agree with the Status field
Many ADRs restate their status in the H1 — # ADR-1610: … [Accepted]. That bracket is the first thing a reader sees, and the index strips it when rendering the title, so a stale one used to be invisible to everyone but the reader it misled.
If the H1 ends in a bracket holding a status token, it must name the same status as the Status field. Comparison is case-insensitive and against the parsed token, so [Superseded] agrees with Status: Superseded by [ADR-0174](0174-retire-gsd-sdk-package-boundary.md) (2026-05-23).
A trailing bracket that is not a status token — [Draft], [WIP] — is treated as part of the title and left alone. If you want a bracket the gate ignores, do not spell it like a status.
6. Every relative link resolves
A link whose target does not exist on disk fails the check, naming the file, the line, and the unresolved target. This covers every markdown file in this directory, including this README and any file whose name breaks the convention above.
| Written as | Treated as |
|---|---|
[t](900-beta.md), [t](../prd/) |
resolved — a directory counts |
[t](900-beta.md#section) |
the file is resolved; the #fragment is not checked |
[t](https://…), [t](mailto:…), [t](//host/x) |
out of scope — absolute destinations are never fetched |
[t](#lifecycle-rules) |
out of scope — a same-document anchor is not a file reference |
[t](/docs/adr/x.md) |
resolved against the repository root, as GitHub does |
a link inside a ``` fence or `backticks` |
not a link — markdown does not render one there, so it is never resolved |
[text][ref] reference-style, <a href>, bare autolinks |
not supported; write an inline link |
Two consequences worth stating outright:
- Case matters, on every platform.
[t](0001-Alpha.md)pointing at0001-alpha.mdfails even on macOS and Windows, because it 404s on github.com and reds the Linux CI lane. The failure names the entry it found so the fix is obvious. - A link to a generated or ignored path fails. Nothing here consults
.gitignore; the question is only whether a reader following the link lands somewhere. Cite the hand-authored source rather than the build artifact.
If the gate rejects something you wrote
Reproduce it locally first — it is the same command CI runs, and it names the file, the line, and the target:
node scripts/gen-adr-index.cjs --check
Then work from the reason:
| What it says | What to do |
|---|---|
does not resolve — no such file or directory at … |
Fix the path. It is relative to docs/adr/, so a sibling ADR is just 900-slug.md. If the target genuinely does not exist yet, drop the link rather than leaving it pointing nowhere. |
…Did you mean X? — link targets are case-sensitive on github.com |
Match the on-disk name exactly. Your machine may open the file regardless; github.com and the Linux CI lane will not. |
escapes the repository |
The path resolves outside the repo. Link something inside it, or use an absolute URL — those are out of scope and never checked. |
is a symlink that escapes the repository |
An ADR file itself is a symlink pointing outside the repo. Commit a real file. |
H1 status bracket […] contradicts the Status field (…) |
Update whichever of the two is stale so they agree. The Status field is authoritative; the bracket is a restatement for the reader. |
A link that is an example, not a destination, belongs in backticks. The gate skips fenced blocks and inline code entirely, because markdown does not render a link there. That is the escape hatch for illustrative syntax — the table above is written that way, which is why it does not fail this check. An indented code block (four spaces) is not skipped; use backticks.
To consume the result from a script rather than by eye, use --json (below) and branch on each violation's stable reason code.
Ratifying a stale Proposed
A stale Proposed is not cosmetic: it tells contributors and agents that live architecture is an unbuilt idea. Fix it — but on evidence, not vibes.
The bar. All four must hold before flipping to Accepted:
- The decided mechanism demonstrably exists in the tree — name the files, symbols, and tests.
- The owning issue is closed as completed. A closed issue is not proof:
stateReasonof not planned / duplicate means the decision was dropped (that isLegacyorRetired, notAccepted). - No material part is unshipped. If the ADR defines phases and one is outstanding, or states its own bar for acceptance and that bar is unmet, it stays
Proposed. - No later ADR supersedes it, and no approved issue already plans its graduation as separate work.
The procedure. Set the status to Accepted — ratified <date> (originally Proposed <date>), add a dated ## Ratification section holding the evidence, then run node scripts/gen-adr-index.cjs --write. If the ADR claims to supersede or subsume others, the gate will now demand their back-links — that is the point. Ratify deliberately.
Two traps worth knowing, both hit during the 2026-07-17 audit:
- Shipped code is necessary, not sufficient. Eight ADRs had every named module, symbol, and test present and their epics closed — and still failed the bar: ADR-2264's own headline acceptance criterion is unmet in the tree, ADR-230's decided branch protection does not match the live API, ADR-660's namesake mechanism is performed by hand, and ADR-959 has an approved issue planning its graduation as its own ADR. Verify the decision, not just the code.
- "Supersedes" is often "subsumes". Read what the ADR means before the gate makes you act on what it says. ADR-857 said "Supersedes (generalizes)"; taken literally, ratifying it would have stamped two live seams (ADR-0011, ADR-58) as dead. The parenthetical was the truth; the field name was wrong.
Maintaining the index
The index is generated. Do not hand-edit it. Everything between the ADR-INDEX:START / ADR-INDEX:END markers is derived from the ADR files themselves:
node scripts/gen-adr-index.cjs # print the index
node scripts/gen-adr-index.cjs --write # regenerate it into this file
node scripts/gen-adr-index.cjs --check # CI: fail if stale or invalid
node scripts/gen-adr-index.cjs --json # same checks, machine-readable report
After adding an ADR, or changing any ADR's status or relations, run --write and commit the result. npm run lint:generated-sync runs --check in CI, so a missing or stale row fails the build rather than rotting silently.
--json runs the same validation as --check and writes a report to stdout instead of prose to stderr, with the same exit code. Each violation carries a stable reason code, so a tool consuming this never has to pattern-match an error message:
{
"ok": false,
"adrCount": 76,
"indexStale": false,
"violations": [
{ "file": "2704-example.md", "line": 41, "reason": "link_unresolved",
"target": "reference/x.md", "resolved": "docs/adr/reference/x.md" }
]
}
An unrecognized flag is rejected rather than ignored.
This replaces a hand-maintained table that had drifted to 40 of 65 ADRs — the entire capability family and EoS itself were missing from it, which is precisely why the ADRs a reader most needed were the ones they could not find.
Index
Active decisions
These govern the system as it stands. Cite these.
| ADR | Title | Status | Read first |
|---|---|---|---|
| ADR-0001 | Dispatch policy module as single seam for query execution outcomes | Accepted | — |
| ADR-0002 | Command Contract Validation Module | Accepted | — |
| ADR-0003 | Model Catalog Module as single source of truth for agent profiles and runtime tier defaults | Accepted | — |
| ADR-0004 | Planning Workspace Module as single seam for worktree and workstream state | Accepted | — |
| ADR-0006 | Planning Path Projection Module for SDK query handlers | Accepted | — |
| ADR-0008 | Installer Migration Module owns install-time upgrade safety | Accepted | — |
| ADR-0009 | Shell Command Projection Module owns runtime-aware OS command rendering | Accepted | — |
| ADR-0011 | review.default_reviewers config key scopes the no-flag /gsd-review fan-out |
Accepted | — |
| ADR-0011 | Skill Surface Budget Module owns install-time profile staging and runtime surface control | Accepted | ADR-857 |
| ADR-15 | Cross-AI Plan Convergence via Existing Orchestration Commands | Accepted | — |
| ADR-22 | Plan-vs-codebase drift guard: defaults and symbol-resolver seam | Accepted | — |
| ADR-58 | Runtime Install Policy Module owns the typed install-plan projection | Accepted | ADR-1239, ADR-857 |
| ADR-0174 | Retire @opengsd/gsd-sdk package boundary — single-runtime collapse | Accepted | — |
| ADR-218 | Harden release-workflow version validation — reject leading zeros and pre-check npm | Accepted | — |
| ADR-227 | Input validation must check semantic shape, not just type | Accepted | — |
| ADR-415 | Prevent stale-base reintroduction of retired runtime tokens | Accepted | — |
| ADR-443 | Unified cross-provider effort controls and fast-mode-aware routing | Accepted | — |
| ADR-452 | Adopt standard ESLint flat-config lint harness | Accepted | — |
| ADR-456 | Test-rigor architecture — deterministic scheduling, antagonistic tier, typed-surface mandate, and delete-bad-tests policy | Accepted | — |
| ADR-457 | Generation model for bin/lib/*.cjs type safety |
Accepted | — |
| ADR-550 | spec-phase probe pattern and prohibition contract | Accepted | — |
| ADR-0656 | Research Module — L2-hybrid seam for cached, curated-first research | Accepted | — |
| ADR-766 | Claude Code Plugin Manifest Module owns the projection of gsd-core surfaces onto the Claude Code plugin contract | Accepted | — |
| ADR-857 | Capability system — five-step loop as core, features as plug-ins behind Loop Extension Points | Accepted | — |
| ADR-894 | Capability declaration format + registry generation | Accepted | ADR-1239 |
| ADR-959 | Capability Command Contribution | Accepted | — |
| ADR-1016 | Runtime Capability Descriptor | Accepted | ADR-1239 |
| ADR-1235 | Migrate agent conversion to the descriptor-driven install path | Accepted | — |
| ADR-1239 | GSD as an Embeddable Orchestration Engine | Accepted | — |
| ADR-1244 | Capability Ecosystem: third-party authoring, versioned manifests, and URL import/upgrade/remove | Accepted | — |
| ADR-1372 | Canonical markdown-structure parsing — the markdown-sectionizer seam |
Accepted | — |
| ADR-1411 | Resolution must report provenance, not fall open silently | Accepted | — |
| ADR-1508 | Runtime Artifact Conversion Module owns per-runtime content rewriting | Accepted | — |
| ADR-1517 | Reviewer instances — bounded config surface for same-adapter multi-model review | Accepted | — |
| ADR-1577 | Untrusted-input boundary + opt-in injection blocking | Accepted | — |
| ADR-1593 | Skill mapping & converter methodology across runtimes | Accepted | — |
| ADR-1610 | workflow & agent size-budget ratchet (per-file byte baseline + tier hard caps) | Accepted | — |
| ADR-1703 | Cross-platform portability enforcement as AST ESLint rules | Accepted | — |
| ADR-1769 | STATE.md Transition Module — intent-based transitions over scattered RMW callbacks | Accepted | — |
| ADR-1787 | /gsd:next smart-entry front door delegates advancement to /gsd:progress --next |
Accepted | — |
| ADR-1817 | STATE.md rebuild — derivability contract (capstone transition) | Accepted | — |
| ADR-1820 | Spec-Optional Predicate Rail — the Spec-Section Detection Module, the fallback toggle, and the SPEC↔probe precedence contract | Accepted | — |
| ADR-1866 | agent_skills dual injection — orchestrator-side + agent-side self-load | Accepted | — |
| ADR-1990 | Existing Code Onboarding Module owns deterministic repo-state detection and onboarding route selection | Accepted | — |
| ADR-2008 | Generic gate-predicate evaluator | Accepted | — |
| ADR-2121 | Phase-Identifier Parsing Consolidation | Accepted | — |
| ADR-2143 | Markdown Table Model, Bounded Mutation, and Fail-Loud Consolidation (#1372 part 2) | Accepted | — |
| ADR-2164 | Statusline draws its data boundary at local, read-only sources | Accepted | — |
| ADR-2207 | STATE.md Status lifecycle — phase-completion writes an intermediate state; milestone-close owns termination |
Accepted | — |
| ADR-2313 | Codex Adopts the Passive / Session-Only Model Posture | Accepted | — |
| ADR-2346 | Command Dispatch Completion | Accepted | — |
| ADR-2363 | A capability's skill body is an instruction surface — trusted, unscanned, and disclosed | Accepted | — |
| ADR-2619 | Observability and shareable diagnostics — wire the dispatch seam, add the outbound trust boundary | Accepted | — |
| ADR-2629 | Phase effort is estimated against a calibrated smart-zone budget, not a static heuristic | Accepted | — |
| ADR-2719 | Emitted-artifact attribution — replace the committed parity fixtures with a computed conservation law | Accepted | — |
| ADR-2782 | Reviewer Lane — the cross-AI reviewer handoff becomes a declared capability surface | Accepted | — |
| ADR-2866 | Install-surface resolution — the install pipeline resolves (runtime × scope × trigger) as a value |
Accepted | — |
| ADR-2966 | Test the five-step loop as a continuous walk, not isolated points | Accepted | — |
| ADR-2980 | A payload-carried error key is a degraded result, not a fault |
Accepted | — |
| ADR-3180 | Planning Semantic Model — Single Owner per Derivation | Accepted | — |
| ADR-3212 | The Lexical Seam — Safe Pattern Construction, Line-Terminator Normalization, and Tokenizer-First Stateful Grammars | Accepted | — |
| ADR-3408 | STATE.md Write Path — One Declared Policy, One Write Seam | Accepted | — |
| ADR-3409 | Shell Guards Must Observe Their Own Failure Arm | Accepted | — |
| ADR-3574 | Install materialization shares primitives, not one writer | Accepted | — |
| ADR-3625 | The platform seam keeps its own Windows binary resolution rather than adopting a spawn library | Accepted | — |
| ADR-3660 | Runtime Artifact Layout Module owns per-runtime artifact placement | Accepted | ADR-1239 |
Proposed
Decided in principle, not yet ratified. Do not cite as settled architecture.
| ADR | Title | Status | Read first |
|---|---|---|---|
| ADR-230 | Introduce next as a long-lived integration branch |
Proposed | — |
| ADR-612 | Bracket Phase-ID Convention | Proposed | — |
| ADR-660 | Release from the head of next; immutable release tags; @next dist-tag as the RC surface |
Proposed | — |
| ADR-1143 | Claude orchestration capability — Workflow tool (ultracode) as a runtime-gated loop execution backend | Proposed | — |
| ADR-1213 | Capability write side — the Capability State Writer | Proposed | — |
| ADR-1606 | prohibition-enforcement verify-time seam | Proposed | — |
| ADR-1671 | Dynamic context management platform | Proposed | — |
| ADR-1953 | Complexity-triggered refactor — the loop measures the entropy it just added | Proposed | — |
| ADR-3128 | Adaptive runtime evidence for GSD Debug | Proposed | — |
Superseded, Retired, and Legacy
Historical record. Do not follow these — each names what replaced it, or why it was retired.
| ADR | Title | Status | Replaced by |
|---|---|---|---|
| ADR-0005 | SDK Architecture seam map for query/runtime surfaces | Superseded | ADR-0174 |
| ADR-0007 | SDK Package Seam Module owns SDK-to-get-shit-done-redux compatibility | Superseded | ADR-0174 |
| ADR-0010 | File Operation Engine Module owns safe runtime/config file mutations | Superseded | ADR-0009 |
| ADR-0010 | Skill Surface Budget Module owns install-time skill listing curation | Superseded | ADR-0011 |
| ADR-0011 | PRD — review.default_reviewers config key for /gsd-review reviewer selection |
Legacy | — |
| ADR-0012 | CommandRoutingHub as single dispatch seam for CJS command families | Superseded | ADR-0174 |
| ADR-2264 | Redesign golden-install-parity — single-source manifest builder + split invariant | Superseded | ADR-2719 |
| ADR-3524 | CJS↔SDK hard seam — one source of truth per Shared Module | Superseded | ADR-0174 |
Generated by scripts/gen-adr-index.cjs — run --write after adding or restatusing an ADR.
Seam map
Orientation for the module-ownership ADRs. This section is prose and hand-maintained; the index above is the authority on status.
How GSD meets a host — start at ADR-1239 (EoS). It is the current frame and subsumes the descriptor/projection ADRs (ADR-1016, ADR-58, ADR-3660, ADR-894) as adapters beneath it.
The SDK seam map is gone. ADR-0005 was once the entry point for SDK module ownership; it is superseded by ADR-0174, which retired the @opengsd/gsd-sdk package boundary entirely. There is no sdk/ tree. Read ADR-0174 for the single-runtime collapse; the seam-Module vocabulary survives under one src/.
ADR-0006 documents how query handlers project planning paths (cwd → effectiveRoot → .planning/<project>/...). Cross-reference the Planning Workspace Module (ADR-0004) for workstream pointer policy.
ADR-0008 documents the Installer Migration Module for safe install-time moves, removals, config rewrites, and user-data preservation.
ADR-0009 documents the Shell Command Projection Module seam for runtime-aware projection of installer-owned command text and projection IR. Its Phases 3–4 absorbed the File Operation Engine Module (ADR-0010).
ADR-0011 documents the Skill Surface Budget Module for install-time skill/agent profile staging (--profile=<name>, .gsd-profile marker, requires: closure) and the Phase 2 runtime /gsd:surface command.
ADR-1411 establishes the Resolution Provenance principle: context resolution (config loading, project-root anchoring, workstream resolution) must report its provenance rather than fall open silently to defaults. It is the resolution-side analog of ADR-227 (input-validation shape).