Files
msd-core/docs/adr/1610-workflow-agent-size-budget-ratchet.md
Tom Boucher 67a9243cf1 chore(#2356): make the ADR index a generated artifact and enforce ADR lifecycle invariants (#2367)
* chore: rebuild ADR index as a generated artifact and enforce lifecycle invariants

The ADR index in docs/adr/README.md was hand-maintained with nothing checking
it, and had drifted to 40 of 65 ADRs. The absent rows included the entire
capability family (857/894/959/1016/1143/1213/1244) and ADR-1239 (EoS) itself,
so the decisions a reader most needed were the ones they could not find.

Make the index a derived artifact, matching the repo's existing generated-file
idiom (lint:generated-sync), and enforce the corpus' lifecycle invariants:

- scripts/gen-adr-index.cjs generates the index between markers and validates
  the status vocabulary (Accepted/Proposed/Superseded/Legacy/Retired),
  successor links, id/filename agreement, and supersession symmetry.
- Wire --check into lint:generated-sync so drift fails CI.

Correct the lifecycle metadata the gate surfaced, without flipping any status:

- ADR-1239 (EoS) declared it subsumed ADR-1016/58/3660/894; none recorded it.
  Add reciprocal "Subsumed by" pointers + dated amendments. Subsumption keeps
  the target Accepted -- these are live adapters, not dead decisions.
- ADR-857/894 carry dated status caveats: they read Proposed while the
  capability system shipped and epic #857 is closed. Ratification is a
  maintainer act and is deliberately left open.
- Link ADR-0005/0007/0012/3524 -> ADR-0174 and ADR-0010 -> ADR-0009; record
  the reciprocal Supersedes on ADR-0009.
- ADR-218 declared itself "ADR-0175" -- an unfinished rename.
- The 0011 PRD moves from the non-canonical "Draft" to "Legacy".

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: capture stderr via spawnSync; record ADR-0010 draft supersession

Two fixes surfaced by the first gsd-test run and by regenerating the index:

- tests/adr-index-gate.test.cjs used execFileSync, which only surfaces stderr
  through the thrown error on non-zero exit. The `--write` path exits 0 while
  reporting outstanding violations on stderr, so the helper always saw ''.
  spawnSync captures both streams on both outcomes.
- The hand-maintained index recorded 0010-skill-surface-budget-module.md as
  "earlier draft superseded by ADR-0011" while the file itself still said
  Proposed. Deriving the index from the files would have dropped that
  assertion and resurrected a superseded draft as a live decision, so it is
  recorded at its source, with the reciprocal Supersedes on ADR-0011.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: drop the dead sdk/ model-catalog candidate retired by ADR-0174

src/model-catalog.cts resolved model-catalog.json through three candidates, the
second being sdk/shared/model-catalog.json three levels up. That was the legacy
source-repo fallback kept by the #3288 fix ("check the co-located path FIRST,
before the legacy source-repo path").

ADR-0174 then retired the @opengsd/gsd-sdk package boundary and deleted the sdk/
tree (11918dcc3), so the candidate can no longer resolve in any layout: a source
repo has no sdk/, and an install layout points it at ~/.claude/sdk/shared/, which
the installer never writes -- the original #3288 bug. It was dead weight implying
a package boundary this repo no longer has.

No test depends on it: the #3288 regression tests in tests/install.test.cjs write
their own synthetic old-path fixture and assert it throws.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: ratify nine shipped ADRs; record why ten others stay Proposed

The corpus carried 19 Proposed ADRs, most describing architecture that had
already shipped. A Proposed label on live architecture tells contributors and
agents the decision is an unbuilt idea -- the capability system and EoS were
both being misread that way.

Audited all 19 against the shipped tree and GitHub. Each candidate flip then had
to survive two independent reviewers instructed to refute it.

Ratified Proposed -> Accepted, each with a dated Ratification section carrying
the verified evidence (file:line, symbols, tests, issue state):

  857  capability system      894  declaration format   1244 capability ecosystem
  1577 injection boundary     1610 size-budget ratchet  1990 existing-code onboarding
  15   cross-AI convergence   22   plan-drift guard     0011 default reviewers

Held ten, each now carrying a "Why this is still Proposed" section naming the
blocker and its unblock condition, so the audit is not repeated:

  2264 its own headline acceptance criterion is unmet in the tree
  230  live branch protection contradicts the decided spec (1 approval, not 2)
  660  the namesake release/<version> re-cut is manual, not automated
  959  issue #2346 is approved and plans its graduation as its own ADR
  1213 the shipped writer's return shape differs from the decided interface
  443  the orchestrator override path has no live caller
  1143 / 1606 each states its own bar for acceptance; neither is met
  612 / 1671 legitimately open

Shipped code proved necessary but not sufficient: eight ADRs had every named
module, symbol, and test present with their epics closed, and still failed the
bar. That lesson is written into README.md's ratification procedure.

Also corrected ADR-857's "Supersedes (generalizes)" to "Subsumes": taken
literally it would have marked two live seams dead -- ADR-0011 (surface.cts:348)
and ADR-58 (runtime-artifact-install-plan.cts:82). Both keep Accepted status and
gain Subsumed-by pointers.

Index: Active 39->48, Proposed 19->10, Superseded/Legacy 7. 65 total.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: harden gen-adr-index against hostile titles and non-ADR filenames (#2356)

Three findings from the pre-PR orthogonal security review, all confirmed:

- An ADR title containing the literal ADR-INDEX:END marker was emitted verbatim
  into its table cell, relocating the splice boundary so the NEXT --write
  spliced against the wrong marker and truncated README.md. Titles now render
  through cellText(), which escapes pipes and angle brackets -- making an HTML
  comment (and any other HTML) unformable from ADR-authored text.
- A docs/adr/*.md without a numeric prefix crashed on match(...)[1] of null.
  Such a file is also invisible to the index -- the very failure this gate
  exists to prevent -- so it is now reported as a naming-convention violation
  naming the file and the fix.
- Tests leaked their mkdtemp dirs. They now use helpers.createTempDir/cleanup
  via t.after(); helpers.cleanup carries the Windows-EBUSY retry budget that a
  raw fs.rmSync lacks (caught by local/no-raw-rmsync-in-tests).

Adds five regression tests: marker hijack, HTML injection, pipe cell-break,
non-conforming filename, and splice stability across repeated writes.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: close two gate false-passes; read ## Supersedes sections (#2356)

Second round of confirmed findings from the pre-PR orthogonal code review. Both
false-passes matter more than a false-fail: a gate that silently misses a
violation is worse than no gate, because it is trusted.

- A relation field mixing a link with a bare id silently dropped the bare claim:
  the check tested `rel.links.length` (does this field have ANY link?) instead
  of whether THAT id was linked. `Supersedes: [ADR-0001](...), ADR-0011` passed
  clean -- accepting exactly the ambiguous bare reference the rule forbids. Now
  each bare id is checked against the ids actually linked in the same field, so
  a repeat in trailing prose stays quiet while an unlinked claim is flagged.
- The ratification guard (`statusToken !== 'Accepted'`) skipped BOTH relation
  directions, which killed the IN check entirely: `supersedes.in` is only ever
  populated on an ADR whose status IS `Superseded`, so a dangling `Superseded by
  X` where X never claims it always passed. The guard now applies to OUT only --
  a prospective claim must not obligate its target, but an ADR's statement about
  ITSELF is always owed a reciprocal.
- Fixing that surfaced a parser gap: ADR-0174 declares its supersessions in a
  `## Supersedes` table SECTION, not a header field, and headerBlock() stops at
  the first `##`. The repo's best-documented supersession was invisible. Section
  form is now parsed for both relations.
- Replaced a vacuous test: the em-dash negation case passed whether or not
  NEGATED_RELATION_RE matched (a mutation to /$^/ survived). It now carries a
  link that would create a failing asymmetric relation if negation did not fire.

Also removes docs/adr/9401-test-target.md -- a synthetic fixture a reviewer
created in the worktree while reproducing a finding, swept in by `git add -A`.

Adds regression tests for each: mixed link+bare, linked-and-repeated-in-prose,
dangling superseded-by from a non-Accepted ADR, and the ADR-0174 section shape.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: escape backslashes before pipes in the ADR index cell renderer (#2356)

CodeQL js/incomplete-sanitization (high) on scripts/gen-adr-index.cjs: cellText()
escaped `|` -> `\|` without first escaping the backslash. Markdown's escape
character is the backslash, so the input `\|` became `\\|`, which renders as a
literal backslash followed by an UNESCAPED pipe -- re-opening the cell break the
pipe escape exists to prevent. Order is load-bearing: escape the escape
character first, then everything that emits one.

Same class as the index-marker hijack fixed earlier: ADR-authored text breaking
out of the cell it is rendered into.

Adds a regression test asserting a `\|`-bearing title leaves exactly the row's
own 5 unescaped delimiters and cannot forge a Status cell. Uses split(/\r?\n/)
per local/no-crlf-fragile-split -- a literal "\n" split is CRLF-fragile on the
Windows CI leg.

Refs #2356

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 10:51:58 -04:00

9.6 KiB

ADR 1610: workflow & agent size-budget ratchet (per-file byte baseline + tier hard caps) [Accepted]

  • Status: Accepted — ratified 2026-07-17 (originally Proposed 2026-06-22); see "Ratification" below
  • Date: 2026-06-22

Provenance. Drafted 2026-06-22 to give an already-shipped architectural governance decision its first ADR. The decision landed across three PRs — #1089 (additive per-file workflow baseline guard), #1096 (swap enforcement to baseline + loose hard caps), #1097 (agent-size baseline + line→byte rebase) — under epic #1074, building on #717 (bytes rebase) and #683 (LF-normalized byte count) and superseding the #597 tier-max ratchet. Authored by the implementer. Verified against tests/{workflow,agent}-size-budget.test.cjs, tests/{workflow,agent}-size-baseline.json, scripts/workflow-size.cjs, and scripts/lib/allowlist-ratchet.cjs on next. The rationale here is lifted from those tests' own doc comments (the decision was documented in-code but never as an ADR).

Ratification (2026-07-17): Proposed → Accepted

Ratified by explicit maintainer directive after independent re-verification of the evidence below; the ADR had sat in Proposed for 25 days after the decision it documents had already shipped.

Evidence the decision shipped:

  • Owning issue #1074 ("replace tier-max workflow size-budget ratchet with a per-file baseline + loose hard caps") is CLOSED, stateReason: COMPLETED, closed 2026-06-12T14:00:56Z.
  • All three landing PRs are MERGED: #1089 (test(#1074): add additive per-file workflow size baseline guard, 2026-06-12T03:59:57Z), #1096 (test(#1074): swap workflow size enforcement to baseline + loose hard caps, 2026-06-12T13:30:44Z), #1097 (test(#1074): agent-size-budget per-file baseline + line→byte rebase, 2026-06-12T13:58:44Z).
  • scripts/workflow-size.cjs:32-35 — lfByteCount() implements the CRLF→LF-normalized byte count described in Decision point 2 (#683).
  • scripts/workflow-size.cjs:64-72,80-82 — measureMdFiles/measureWorkflows is the single shared measurement path cited in Decision point 5, re-exported for both the guard and scripts/update-size-baseline.cjs.
  • scripts/lib/allowlist-ratchet.cjs:180 exports assertFileBaseline — the per-file baseline assertion named in Decision point 3 and Cross-references.
  • tests/workflow-size-budget.test.cjs:95-97,102 defines XL_CAP = 98304 (96 KiB), LARGE_CAP = 61440 (60 KiB), DEFAULT_CAP = 40960 (40 KiB), NEW_FILE_CAP = 32768 (32 KiB) — the exact numbers quoted in Decision point 3.

Governance: owning issue #1074, stateReason: COMPLETED, closed 2026-06-12T14:00:56Z.

Context

gsd-core/workflows/*.md and agents/*.md are loaded verbatim into agent context every time the corresponding command/agent runs. Unbounded growth is paid on every invocation across every session, and — more importantly — degrades quality: larger context erodes recall and reasoning ("context rot" / finite attention budget). With prompt caching the per-invocation cost premise is weak (cache reads are ~10% of input), so the caching-independent quality argument is the load-bearing one: lean, high-signal instructions produce better plans.

The prior mechanism (#597) was a tier-max tighten-only ratchet: it bound only the single largest file per tier, leaving the other ~85 files free to grow silently. That left the actual risk — broad, quiet creep across many files — unguarded.

No ADR governs how GSD bounds instruction-document size. The decision currently lives only in test doc comments, so it is invisible to anyone browsing docs/adr/ and at risk of being weakened (e.g. "just bump the baseline") without encountering its rationale.

Decision

  1. Measure in BYTES, not lines (#717). Line count is a poor proxy — markdown tables and fenced code are token-dense, so a line budget over-penalizes prose and under-catches dense additions. Bytes are cheap, deterministic, and need no tokenizer; they are also the unit vendors bound on (Codex caps instruction docs at 32,768 bytes, project_doc_max_bytes, and truncates past it). We adopt the unit, not the exact number.

  2. Count LF-normalized bytes (#683). Normalize CRLF→LF (content.replace(/\r\n/g, '\n')) before counting so a CRLF (Windows) checkout yields the same byte count as an LF checkout — a raw on-disk count would add one byte per line on Windows and make the guard platform-dependent.

  3. Two complementary guards, neither a tier-max ceiling (#1074):

    • Per-file baseline (the anti-creep). Every workflow/agent file is pinned to its exact byte size in tests/{workflow,agent}-size-baseline.json. Any growth fails with the file name and delta. A deliberate change is recorded via npm run size:baseline as a one-line reviewable diff. This is the day-to-day guard and it covers every file, not just the largest-per-tier.
    • Tier hard caps (the outer bound). XL / LARGE / DEFAULT are absolute red lines with real headroom (XL_CAP = 98304 / 96 KiB, LARGE_CAP = 61440 / 60 KiB, DEFAULT_CAP = 40960 / 40 KiB), a few KB above the largest file in each tier — the largest XL orchestrators (plan-phase.md, execute-phase.md) currently sit in the low-90s KB, a few KB under XL_CAP. (Exact per-file sizes are not duplicated here: they live in tests/workflow-size-baseline.json, the enforced source of truth, and an absolute byte count quoted in prose drifts with every edit.) Hard caps are never raised in normal work; crossing one is a signal to do lazy extraction, not a +N bump. New workflow files default to the Codex 32 KiB anchor (NEW_FILE_CAP = 32768) unless explicitly tiered in the same PR.
  4. The metric is a proxy for bounded loaded context — do not game it. The real target is total context loaded at runtime. Because @~/.claude/gsd-core/references/... imports are loaded eagerly, moving prose into an eagerly @-imported reference shrinks the measured file while leaving (or growing) total loaded context — that is gaming the proxy (Goodhart's Law: the moment the byte count is the target, the cheapest way to satisfy it is to relocate bytes, not remove them). Legitimate extraction is LAZY: content Read only at the step that needs it (the workflows/discuss-phase/modes/ progressive-disclosure pattern). The per-file baseline + tier-cap pair is itself a Goodhart hedge — two guards pulling in different directions are harder to game than one.

  5. Shared measurement path. The guard and the baseline generator both measure via scripts/workflow-size.cjs (lfByteCount/measureWorkflows, re-exported to scripts/update-size-baseline.cjs and asserted via scripts/lib/allowlist-ratchet.cjs assertFileBaseline), so the recorded baseline and the enforced size can never drift apart (#1074).

Alternatives considered (rejected & deferred)

  • Tier-max tighten-only ratchet (#597) — SUPERSEDED. Bound only the largest file per tier; the other ~85 files could grow silently. Replaced by the per-file baseline, which guards every file. Re-open only if the per-file baseline proves too noisy in practice (it has not).
  • A line-count budget — REJECTED (#717). Token-dense tables/code make lines a poor proxy; bytes are the vendor-bound, tokenizer-free unit. Re-open only if a cheap token count becomes portably available and measurably better than bytes.
  • Eager @-import extraction to "fit" a file — REJECTED as proxy-gaming (#717). Shrinks the measured file without shrinking loaded context. Only lazy (Read-at-step) extraction counts. Not re-openable — it defeats the goal the metric proxies.
  • Codex 32 KiB as a hard ceiling for the grandfathered orchestrators — REJECTED. The XL/ LARGE tiers sit above 32 KiB because they are top-level orchestrators loaded by Claude, not Codex AGENTS.md docs; 32 KiB is the new-file anchor, not a universal cap.

Consequences

  • Positive: broad, quiet size creep is caught per-file across the whole surface; deliberate growth is an explicit, reviewable one-line baseline diff; the byte unit is deterministic and cross-platform (LF-normalized); the "bounded loaded context" goal and its anti-gaming rule are recorded where a contributor will meet them; the quality (context-rot) rationale is no longer caching-dependent hand-waving.
  • Costs: every legitimate edit to a workflow/agent file must regenerate the baseline (npm run size:baseline) or the guard reds — intentional friction that makes growth visible. The baseline JSON is a merge-conflict surface on concurrent edits (resolved by regenerating).
  • Boundary: this governs instruction-document size (workflow/agent .md loaded into context). It is distinct from the install-time skill-surface budget of ADR-0010/0011 (how many skills are eagerly registered) — different lever, different file, do not conflate.

Cross-references

  • Epic #1074 (per-file baseline + hard caps); #717 (bytes rebase + rationale); #683 (LF byte count); #597 (superseded tier-max ratchet); PRs #1089 / #1096 / #1097.
  • Code: scripts/workflow-size.cjs, scripts/update-size-baseline.cjs, scripts/lib/allowlist-ratchet.cjs (assertFileBaseline), tests/{workflow,agent}-size-budget.test.cjs, tests/{workflow,agent}-size-baseline.json.
  • Related but distinct: ADR-0010/0011 (skill-surface budget — install-time, not file size); ADR-456 (test-rigor architecture — the test these guards are authored under).
  • External anchors: Codex project_doc_max_bytes (32 KiB); Anthropic "effective context engineering for AI agents" (the context-rot / attention-budget argument).