Files
msd-core/src/workflow-fragments.cts
Tom Boucher ff4a57b78c chore(#1671): migrate the remaining 13 LARGE/XL workflows to the fragment model — Phase 6.3 (#3030)
* chore(#2994): fragmentize progress.md forensic audit onto the fragment model

Extract the --forensic-gated forensic_audit step to
workflows/progress/steps/forensic-audit.md behind a section marker, and
repair progress.md's init line to forward --forensic so the atom is
actually true in production rather than only under direct CLI tests.

progress.md shrinks 32630 -> 27207 bytes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize the four manifest-wired workflows

new-project, quick, new-milestone and progress each already had a
dedicated cmdInit* entry point but zero marked sections. Extract nine
gated bodies to workflows/<wf>/steps/ behind section markers and repair
each init line to forward its flags.

Fold --full into the discuss/research/validate facts inside cmdInitQuick
so the when= grammar never sees an OR, per the chunked-mode precedent.

Fixes found while working, per the no-defer rule:
- cmdInitProgress passed no phase info to buildSectionManifestField, so
  state:phase-mvp-mode was permanently false — an atom in the vocabulary
  whose fact could never be computed.
- the quick init router folded flag tokens into the free-text
  description, which the new forwarding would have corrupted.
- a #2508 dispatch note was nested inside quick.md's Agent(prompt=)
  fence, leaking orchestrator guidance into the subagent prompt.
- progress.md had a 3-vs-4 backtick outer-fence imbalance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize verify-work.md and admit state:ui-phase-active

Wire cmdInitVerifyWork to buildSectionManifestField — it was a dedicated
entry point that never emitted a manifest — and mark two sections.

state:ui-phase-active folds (plan:pre hooks include an active ui step) OR
(the phase dir holds a *-UI-SPEC.md) into one boolean in init.cts, so the
grammar still sees a single operator-free atom. The inner Playwright-MCP
check stays as prose inside the fragment: it is live session state and no
init seam can precompute it.

The MVP false-branch note is a real fallback, not redundant prose, so it
sits outside the marker — gating it away would delete the text needed
precisely when MVP mode is off.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): follow moved workflow content in drift guards

Retarget every guard that asserted on content this branch moved into
workflows/<wf>/steps/, mirroring 815b3d897. Each retargeted assertion was
verified to still fail when its step file is blanked, so none was
weakened into vacuity.

Three assertions in verify-mvp-uat were genuinely red. Three more were
worse than red — passing for the wrong reason:
- quick-commit-boundary and worktree-cleanup anchored on indexOf('Step
  5.6'), which matched a later cross-reference and sliced 16069 chars
  that coincidentally held the asserted substrings. Replaced with an
  expandWorkflowSections helper that splices step content back in place.
- phase6-review-capabilities lost its end boundary and widened to EOF.
- playwright-ui-verify matched 'UI' in an unrelated bullet and 'fall
  back' in a subagent-dispatch line after the real content moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize code-review and complete-milestone, admit three atoms

Add dedicated cmdInitCodeReview and cmdInitCompleteMilestone entry points
alongside the shared generic ones rather than modifying them — init.phase-op
and init.manager carry a CRITICAL blast radius (179 dependents, 24
processes) and stay byte-identical for their other callers.

Admit flag:--fix, state:fallow-enabled and state:git-create-tag, each with
a consuming section and a fact its own entry point computes.

Both sections had the resolver-in-body hazard: the fallow config-gate and
the git.create_tag check each sat inside the very block being gated, so
gating would have disabled the resolver that decides the gate. Both are
hoisted into init and the bodies now consume the resolved fact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): retarget code-review and milestone drift guards, fix two red tests

Retarget guards that asserted on content moved into steps/, proving
non-vacuity by blanking each step file and confirming failure.

Also fixes two genuinely red tests found while working, per the no-defer
rule:
- workflow-fragments' frozen-vocabulary lock was missing
  state:ui-phase-active, so commit 7ef7f8336 shipped red. Lint and build
  both passed over it, which is why neither is sufficient verification.
- code-review's quick.md capability-hook assertion carried a stale
  delimiter after the 18ff35d20 extraction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize autonomous.md and admit state:plan-strategy-converge

Five sections share one atom, the pattern plan-phase already uses for
flag:--research-phase. The atom folds --converge OR --cross-ai into a
single boolean in cmdInitAutonomous so the grammar stays operator-free.

cmdInitAutonomous is additive; init.milestone-op, init.manager and
init.phase-op are untouched and still consumed. The $PLAN_STRATEGY bash
resolver is deliberately retained — ungated local-planning bullets still
read it, so the init-side fact supplements it rather than replacing it.

converge-fail-fast required splitting one bash fence so the always-run
CONVERGENCE_ARGS construction stays outside the marker. All three
flag-absent fallbacks were left outside their markers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize review and discuss-phase-assumptions

Admit state:reviewer-instances-configured (two peripheral notes share it;
the core reviewer-lane dispatch stays unmarked — it is the workflow's
primary always-evaluated logic, not an optional branch) and
state:auto-advance-active, which folds --auto OR two config keys into one
boolean so the grammar stays operator-free.

discuss-phase-assumptions was the highest-risk edit in this PR. Its
auto_advance step is a full if/elif/else; gating it whole would have
deleted the flag-absent fallback needed exactly when --auto is off. Split
verified exact: resolvers 636-651 and the 'End here' fallback 668-669 both
stay outside the marker; only 653-667 is gated.

Adds emitted-drift acks for the two files that grew — review.md (+55 B)
and autonomous.md (+737 B from 80799211c, which had none and would have
red-gated the push.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize docs-update, update, transition and new-milestone Part A

Completes the 13-workflow rollout. Three of these had no init call at all
and gained a dedicated entry point plus their first gsd_run query line.

Admits state:is-monorepo and adds state:next-channel, state:workstream-active
and state:flat-mode. Vocabulary 26 -> 30 atoms.

Part A of new-milestone applies when NO workstream is active — the negation
of state:workstream-active. Rather than teach the grammar negation, which is
the Greenspun drift the frozen list exists to prevent, it gets a separate
positively-phrased atom whose fact is the inverse. Part B, which always runs,
stays outside the marker.

flag:--verify-only is deliberately NOT admitted: docs-update has no
contiguous purely-additive region for it, and an atom without a consuming
section is dead vocabulary. Evidence recorded in the slice report.

update.md reuses its existing resolved $GSD_TOOLS rather than prepending the
canonical preamble, which would have clobbered it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): stop automated-ui-verification re-resolving its own gate, retire dead vocabulary

Two defects the new tests caught.

The automated-ui-verification step re-ran gsd_run loop render-hooks and
recomputed UI_PHASE_ACTIVE inside a body that is only read when that fact
is already true — the circular self-disabling pattern this design forbids,
introduced by 3c654b168. cmdInitVerifyWork now exposes ui_phase_active and
the step consumes it. Its launcher preamble goes too: no gsd_run remains.
The Playwright-MCP check stays as prose — that is live session state.

Dead vocabulary predating this PR: flag:--full and state:needs-codebase-map
were admitted with a gate-1 claim that never materialized. flag:--full is
removed, redundant once quick folds it into discuss/research/validate.
state:needs-codebase-map gets the real consumer it always lacked, gating
new-project's codebase-map offer. Vocabulary 30 -> 29, and no atom is now
without a consuming section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): add the atom-admission, inversion and resolver-hoist gates

The two existing parity guards prove vocabulary/predicate symmetry but
never that a fact is computed — an atom no cmdInit* assembles evaluates
false forever. These close that hole:

- per-atom satisfiability for all 29 atoms, plus an anti-vacuity assertion
  so the loop cannot silently cover zero atoms
- dead-vocabulary check against the shipped manifest
- inversion guard: the flag-absent fallbacks in discuss-phase-assumptions
  and verify-work must stay outside their markers
- data-driven resolver-hoist guard over the shipped manifest, so a future
  extraction cannot reintroduce the circular class
- compound-fold coverage (--full, --cross-ai, --rc, config-only --auto)
- null-vs-[] degraded/computed distinction, and flag value shapes

Also repairs the frozen-vocabulary lock, which was stale and red for the
seven atoms earlier commits on this branch shipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2994): add changeset for the fragment-model rollout

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): cite the issue on the two new allow-test-rule exemptions

ADR-456 requires an issue ref on the same line as the annotation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2994): correct the atom-count claims after retiring flag:--full

The vocabulary doc comments still said 30 entries; it is 29 since
flag:--full was removed as dead vocabulary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): dedupe the phase-fallback block and harden --ws parsing

Review findings.

MAJOR: the three new init entry points each pasted a verbatim copy of the
guardedFindPhase/guardedGetRoadmapPhase fallback, taking the repo from four
copies to seven — DEFECT.GENERATIVE-FIX. Extracted applyRoadmapFallback and
folded six of the seven; each call site keeps its own field-set via a
closure. Duplication removed rather than papered over with a parity test.
cmdInitPhaseOp stays out: its fallback omits has_reviews, so it is not a
byte-identical copy, and it is CRITICAL-radius.

LOW, pre-existing: GSD_WS captured [^[:space:]]+ and expands unquoted, so a
workstream name holding glob metacharacters would expand against the
filesystem. Narrowed to [A-Za-z0-9._-]+. The unquoted expansion is kept —
it must word-split into two args and vanish when empty.

Also restores the vocabulary ordering convention, and fixes a masked test
bug the mandated run surfaced: the flag-forwarding guard checked only the
first init line per workflow, but new-milestone has two, so a real failure
was reporting exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): drop the stale new-milestone emitted-drift ack

new-milestone.md was acked for a +406 B growth measured against an
intermediate commit. Net against origin/next it SHRANK by 8 bytes, so
nothing needed the ack and it explained nothing — which the differential
attribution check reports as a stale acknowledgment, not a pass.

update.md's entry stays: it genuinely grew +703 B.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): resolve the 15 failures from the full matrix run

All 15 were real and identical on both lanes.

REAL REGRESSION: autonomous.md hit 41479 chars against the #2196 guard's
40960 cap — a CHARS cap distinct from the LARGE tier byte cap, which the
five section stubs pushed it over. Extracted the 3a.5 UI Design Contract
body to references/; now 39968 chars, and the file nets -795 B vs base, so
its growth ack is deleted rather than left stale.

REAL DEFECT: docs referenced /gsd-transition, which is not a live
registered command. Reworded.

STALE FIXTURE: the emission byte-identity test hardcoded two marked
workflows; this branch legitimately marks fifteen. Fixture corrected — the
source was right.

The rest were drift guards over the eight workflows the earlier sweep did
not cover, retargeted at where the content now lives with non-vacuity
proven by blanking each step file and confirming failure. The GSD_WS
forwarding guard was checked as a possible real break and is not one: the
charclass narrowing is intact and forwarding works end to end.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): drop the ack for a newly-added reference file

A new file's emitted ripple is attributable to the diff that adds it, so
the acknowledgment explained nothing and the differential check reports it
as stale. Removing the last entry removes the fragment — an empty one
signals nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): retarget the UI-contract guards and clear two transitive advisories

The §3a.5 extraction that brought autonomous.md under the #2196 char cap
moved its body to references/autonomous-ui-design-contract.md, so ten
guards in autonomous-ui-steps and check-ui-safety-gate were asserting it
against the host. Retargeted via a combined read, each proven non-vacuous
by blanking the reference file and confirming failure.

This class had already bitten twice on this branch because each sweep was
scoped to the workflows touched at that moment, so this one was
exhaustive: ~70 test files across all 13 workflows, zero further broken or
vacuous assertions found.

Also clears two high transitive advisories the matrix flagged on one lane
— fast-uri GHSA-7p8r-x3mc-p8w7 and three ip-address SSRF/trust-boundary
issues. Both pre-date this branch: package-lock.json was untouched until
now, so the production tree was byte-identical to the base. Lockfile-only,
package.json unchanged, verified against a real npm ci install.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): backfill changeset pr number to 3030

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 19:59:58 -04:00

601 lines
27 KiB
TypeScript

/**
* Workflow Fragments — in-file `<!-- gsd:section -->` marker parser/composer
* for GSD workflow markdown files (ADR-1671 epic #1671, Phase 3 / issue #2930,
* `.gsd/phase/chore-2930-fragmentize-xl-workflow/40-design.md`).
*
* Pure module: no I/O, no dependency beyond node built-ins and the shared
* budget-trim seam `context-composer.cjs` (issue #2929). Emission order is
* `parseWorkflowSections` -> `toFragments` -> `composeWithinBudget` ->
* `renderFragments` (= `composeWorkflow`), run BEFORE the per-runtime
* converters so a marker attribute never reaches a path-rewrite regex.
*
* ## Marker grammar (CLOSED)
*
* Open: a line whose only content (after trimming leading/trailing
* whitespace) is `<!-- gsd:section id="<id>" when="<when>" -->`.
* Close: a line whose only content is `<!-- /gsd:section -->`.
*
* Attribute order is free and inner spacing is flexible (Postel on FORMAT);
* `id` and `when` VALUES are validated strictly and fail closed (Postel is
* deliberately NOT applied to semantics — an unrecognized `when` is an
* authoring instruction that must never be silently dropped). `id` matches
* `/^[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$/`; `when` must be `===` exactly one
* entry of the frozen {@link WHEN_VOCABULARY} — no operators, no negation,
* no nesting (Greenspun's Tenth Rule: extending the vocabulary is a
* coordinated ADR amendment, never an organic edit).
*
* ## Partition invariant
*
* `parseWorkflowSections` returns sections that PARTITION the document:
* every byte that is not part of a marker LINE belongs to exactly one
* section, in document order. Text outside any marker pair becomes a
* synthesized gap section (`explicit: false`, id `gap-<n>`, n from 0). A
* marker line is removed IN FULL — text and its original line terminator —
* so an unmarked document (88 of 89 workflows today) parses to exactly one
* implicit gap fragment and composes back byte-identical.
*
* Line splitting is CRLF-aware per line (not `content.split('\n')`, which
* would leave a stray `\r` glued to `.text` and cannot express a mixed
* CRLF-marker/LF-body document): {@link splitLinesPreservingEol} records
* each line's own terminator (`''`, `'\n'`, or `'\r\n'`) so reassembly is
* exact regardless of line-ending mixture.
*
* ## Fence + comment interleaving (the highest-risk code here)
*
* A marker is structural only when it is NOT inside a fenced code block and
* NOT inside an unrelated HTML comment (`<!-- gsd:loop-host ... -->` is a
* different marker family entirely and is left untouched by construction —
* it does not match the `gsd:section` token). Fences and comments are
* scanned in ONE left-to-right interleaved pass with two mutually exclusive
* states (`fence`, `inComment`), copying the discipline documented in
* `src/context-predicates.cts`'s module comment (DEFECT.CONTEXT-PREDICATES-
* COMMENT-FENCE-BLIND, #2928): while a fence is open, only a matching closer
* can end it (a `<!--`/`-->` token on a fenced line is fence content, never
* a comment boundary); while a comment is open, only a `-->` token can end
* it (a fence delimiter inside it is comment content, never a fence
* boundary); when neither is open, a comment opener is checked BEFORE a
* fence opener (HTML comments are lexically outermost). A two-pass design
* (mask one construct, then scan for the other) resolves this wrongly in
* one direction and silently skips to EOF — that is the exact defect this
* module avoids by construction. An unclosed fence at EOF does NOT throw;
* everything after it is simply literal.
*
* One deliberate refinement beyond a naive "does the trimmed line START
* WITH `<!--` and END WITH `-->`" check: whether a comment PERSISTS past
* the current line is decided by `.includes('-->')` (does a close token
* appear anywhere on the line), not by `.endsWith('-->')`. A line like
* `<!-- TODO: fix --> some trailing prose` closes its comment on the same
* line and must not swallow the rest of the document — it is simply not a
* `gsd:section` marker (a marker's grammar requires the comment to be the
* line's ONLY content), and is left as ordinary content in whichever
* section/gap contains it.
*
* Known inherited limitation (shared with `context-predicates.cts`, not a
* regression introduced here): comment-open detection is anchored to the
* start of the trimmed line. An HTML comment that opens *mid-line* (prose
* followed by an unclosed `<!--`) is not tracked, so a `gsd:section`-shaped
* line appearing on a later line inside that comment would be misread as
* real. GSD workflow markers are always authored on their own line, so this
* does not affect the production shape; documented here rather than papered
* over.
*
* ADR-457 build-at-publish: compiled by tsc to
* gsd-core/bin/lib/workflow-fragments.cjs (gitignored).
*/
// eslint-disable-next-line @typescript-eslint/no-require-imports -- context-composer.cjs is a CommonJS module compiled from a sibling .cts source; `import x = require()` reads its module.exports namespace directly.
import contextComposer = require('./context-composer.cjs');
/**
* Frozen, CLOSED applicability vocabulary for the `when=` attribute.
* Extending this list requires an ADR amendment, not an organic edit
* (Greenspun's Tenth Rule — see the module doc comment).
*
* Widened from 4 to 14 entries via the ADR-1671 amendment for #2992 (epic
* #1671 Phase 6.1; see `.gsd/phase/chore-2992-widen-when-vocabulary/
* 40-design.md`), then from 14 to 19 via the ADR-1671 amendment for #2993
* (epic #1671 Phase 6.2; see `.gsd/phase/chore-2993-fragmentize-plan-phase/
* 40-design.md`), then from 19 to 20 via the ADR-1671 amendment for #2994
* (epic #1671 Phase 6.3, `verify-work.md`), then from 20 to 23 via a further
* #2994 amendment fragmentizing `code-review.md` and `complete-milestone.md`,
* then from 23 to 24 via a further #2994 amendment fragmentizing
* `autonomous.md`, then from 24 to 26 via a still further #2994 amendment
* fragmentizing `review.md` and `discuss-phase-assumptions.md`, then from 26
* to 30 (then 29; `flag:--full` retired as dead vocabulary) via the FINAL
* #2994 slice (epic #1671 Phase 6.3) fragmentizing
* `docs-update.md`, `update.md`, `transition.md`, and `new-milestone.md` —
* every one of the 13 workflows targeted by ADR-1671 is now on the fragment
* model. The vocabulary remains CLOSED: no operators, no negation, no nesting.
* Cardinality is not expressiveness — a 29-entry flat list with no
* composition is still not a language.
*
* Held at 14, not wider: an atom whose fact is never computed always
* evaluates FALSE, so a section marked with it would silently never
* include — a silent-exclusion bug, not a feature. One further atom
* (`flag:--verify-only`) was surveyed but is NOT admitted even now that
* `docs-update` has its own `cmdInit*` entry point (`cmdInitDocsUpdate`):
* the flag's control flow is INTERLEAVED across three non-contiguous
* touch-points in `docs-update.md` (an inline early-exit check in
* `init_context` — "If `--verify-only` is present…skip to
* verify_only_report" — a "Skip condition" note embedded in another step's
* body, and the `verify_only_report` step itself) rather than a single
* contiguous, whole-line, purely-additive region — admitting the atom to
* gate only the `verify_only_report` step would leave the other two
* touch-points as un-migrated raw `$ARGUMENTS` checks. `state:is-monorepo`
* (`dispatch-monorepo-packages` section) IS admitted in this slice — see the
* paragraph below.
* `flag:--fix`, `state:fallow-enabled`, and `state:git-create-tag` were
* withheld for the same reason until a further #2994 amendment gave
* `code-review` and `complete-milestone` their own dedicated `cmdInit*`
* entry points (`cmdInitCodeReview`, `cmdInitCompleteMilestone`). The
* originally-surveyed `flag:--converge` never shipped under that name: once
* `autonomous` gained its own dedicated `cmdInitAutonomous` entry point, the
* admitted atom is `state:plan-strategy-converge` instead — `--cross-ai` is
* a documented alias for `--converge` (`autonomous.md`'s own `PLAN_STRATEGY`
* resolver folds both into one value), so a `flag:--converge`-only atom
* would have left `--cross-ai`-only invocations silently excluded from the
* same sections; see the `state:plan-strategy-converge` paragraph below.
* `state:reviewer-instances-configured` and `state:auto-advance-active` were
* withheld the same way until `review` and `discuss-phase-assumptions`
* gained their own dedicated `cmdInit*` entry points (`cmdInitReview`,
* `cmdInitDiscussPhaseAssumptions`).
*
* The #2993 widening adds 5 entries fragmentizing `plan-phase.md`:
* `flag:--ingest`, `flag:--prd`, `flag:--research-phase`, `flag:--reviews`,
* `state:chunked-mode`. `state:chunked-mode` is a disjunction (`--chunked`
* flag OR `.planning/config.json` `workflow.plan_chunked`) resolved to a
* single boolean FACT by the init seam (`src/init.cts`) — the grammar still
* sees exactly one atom with no operator, preserving the same guard.
*
* The #2994 widening adds 1 entry fragmentizing `verify-work.md`:
* `state:ui-phase-active`. Like `state:chunked-mode`, it is a disjunction —
* the phase's active `plan:pre` loop hooks include the `ui-phase` step OR
* the phase directory already contains a `*-UI-SPEC.md` — resolved to a
* single boolean FACT by the init seam before it ever reaches this grammar.
*
* A further #2994 widening (epic #1671 Phase 6.3) adds 3 entries
* fragmentizing `code-review.md` and `complete-milestone.md`: `flag:--fix`
* (`dispatch-fix` section), `state:fallow-enabled` (`structural-pre-pass`
* section — the fallow config-gate resolver, previously re-derived inside
* the section body itself, is hoisted into `cmdInitCodeReview` and exposed
* as top-level `fallow_*` init-bundle fields), and `state:git-create-tag`
* (`git-tag` section — the `git.create_tag` config-gate resolver is hoisted
* into `cmdInitCompleteMilestone`).
*
* A still further #2994 widening (epic #1671 Phase 6.3) adds 1 entry
* fragmentizing `autonomous.md`: `state:plan-strategy-converge`, gating five
* sections (`converge-fail-fast`, `converge-banner`, `converge-dispatch-bg`,
* `converge-dispatch-inline`, `converge-loop`) that all share the same atom
* — legal and precedented (`plan-phase.md`'s `research-only-*` pair already
* shares `flag:--research-phase`). It is a disjunction — `--converge` OR its
* documented alias `--cross-ai` — resolved to a single boolean FACT by the
* new `cmdInitAutonomous` entry point (`flags.has('--converge') ||
* flags.has('--cross-ai')`) before it ever reaches this grammar, same
* discipline as `state:chunked-mode`/`state:ui-phase-active` above.
*
* A still further #2994 widening (epic #1671 Phase 6.3) adds 2 entries
* fragmentizing `review.md` and `discuss-phase-assumptions.md`:
* `state:reviewer-instances-configured` (`reviewer-instances-note-1` and
* `reviewer-instances-note-2` sections — two peripheral notes sharing one
* atom, the same sharing pattern `plan-phase.md`'s `research-only-*` pair
* already established) and `state:auto-advance-active` (`auto-advance-dispatch`
* section — a disjunction, `--auto` flag OR a consolidated auto-mode config
* fact, resolved to a single boolean FACT by the new
* `cmdInitDiscussPhaseAssumptions` entry point before it ever reaches this
* grammar, same discipline as `state:chunked-mode` above).
*
* The FINAL #2994 widening (epic #1671 Phase 6.3) adds 4 entries, closing
* out the last four workflows on ADR-1671's fragmentization list —
* `docs-update.md`, `update.md`, `transition.md`, `new-milestone.md` — none
* of which carried a `gsd_run query init.*` call before this slice:
* `state:is-monorepo` (`dispatch-monorepo-packages` section, new
* `cmdInitDocsUpdate` entry point — reuses `docs.cts`'s own
* `detectMonorepoWorkspaces` detector rather than a second scan);
* `state:next-channel` (`channel-banner` section, new `cmdInitUpdate` entry
* point — `--next` OR its documented alias `--rc`, resolved in PARALLEL
* with, not in place of, `update.md`'s own `TAG="next"` case-statement,
* which issue #815's regression test requires to stay literal in the
* workflow); `state:workstream-active` (`workstream-collision-check`
* section, new `cmdInitTransition` entry point — a workstream is active,
* `GSD_WORKSTREAM` env falling back to the stored active-workstream
* pointer, the same authoritative source `cmdInitProgress` already uses);
* and `state:flat-mode` (`project-md-milestone-write` section,
* `cmdInitNewMilestone` — the positively-phrased INVERSE of
* `state:workstream-active`, introduced because the grammar has no negation
* operator and `new-milestone.md`'s Step 4 Part A is gated on the OPPOSITE
* condition from `transition.md`'s section).
*/
export const WHEN_VOCABULARY: readonly string[] = Object.freeze([
'always',
'flag:--wave',
'state:gap-closure-phase',
'state:has-prior-phases',
'flag:--auto',
'flag:--discuss',
'flag:--fix',
'flag:--forensic',
'flag:--ingest',
'flag:--prd',
'flag:--research',
'flag:--research-phase',
'flag:--reset-phase-numbers',
'flag:--reviews',
'flag:--validate',
'state:auto-advance-active',
'state:chunked-mode',
'state:fallow-enabled',
'state:flat-mode',
'state:git-create-tag',
'state:is-monorepo',
'state:needs-codebase-map',
'state:next-channel',
'state:phase-mvp-mode',
'state:plan-strategy-converge',
'state:reviewer-instances-configured',
'state:ui-phase-active',
'state:workstream-active',
'state:worktrees-enabled',
]);
/**
* Frozen, stable reason codes for every `fail()` throw site in this module.
* Tests assert via `assert.equal(err.reason, REASON.X)` rather than
* regex-/substring-matching the human-readable message (CONTRIBUTING.md
* "Prohibited: Raw Text Matching on Test Outputs"; shape copied from this
* repo's own `gsd-core/bin/verify-reapply-patches.cjs` REASON enum) — a
* message reword must never silently pass a test that exists to catch a
* behavior regression.
*
* Adding a new reason requires updating this map AND the test that locks
* `Object.keys(REASON).sort()` as a coordinated change.
*/
export const REASON = Object.freeze({
UNCLOSED_SECTION: 'unclosed_section',
UNMATCHED_CLOSE: 'unmatched_close',
NESTED_SECTION: 'nested_section',
DUPLICATE_ID: 'duplicate_id',
MISSING_ID: 'missing_id',
MISSING_WHEN: 'missing_when',
MALFORMED_ID: 'malformed_id',
UNKNOWN_WHEN: 'unknown_when',
MALFORMED_ATTRIBUTES: 'malformed_attributes',
UNRECOGNIZED_ATTRIBUTE: 'unrecognized_attribute',
CLOSE_WITH_ATTRIBUTES: 'close_with_attributes',
});
/** One partitioned section of a parsed workflow document. */
export interface WorkflowSection {
readonly id: string;
readonly when: string;
readonly body: string;
/** false for a synthesized gap fragment (unmarked text between/around marker pairs). */
readonly explicit: boolean;
/** 1-based. The `<!-- gsd:section -->` open marker's line for an explicit section, or the first line of the gap for a synthesized one. */
readonly startLine: number;
}
/** One line of source content plus its ORIGINAL terminator, individually. */
interface LineRecord {
readonly text: string;
readonly eol: '' | '\n' | '\r\n';
}
/**
* Split `content` into per-line records that each carry their OWN original
* terminator, so CRLF/LF mixes and a missing trailing terminator reassemble
* byte-for-byte via `record.text + record.eol` concatenation. See the
* module doc comment's "Line splitting is CRLF-aware" note for why a bare
* `content.split('\n')` cannot serve this.
*
* @param content - full source document text
*/
function splitLinesPreservingEol(content: string): LineRecord[] {
const lines: LineRecord[] = [];
let i = 0;
while (i < content.length) {
const nlIdx = content.indexOf('\n', i);
if (nlIdx === -1) {
lines.push({ text: content.slice(i), eol: '' });
break;
}
const hasCr = content[nlIdx - 1] === '\r';
const end = hasCr ? nlIdx - 1 : nlIdx;
lines.push({ text: content.slice(i, end), eol: hasCr ? '\r\n' : '\n' });
i = nlIdx + 1;
}
return lines;
}
// Fence delimiter line matcher — mirrors `context-predicates.cts`'s (itself
// mirroring `markdown-sectionizer.cts`'s `scanFencedBlocks`) exactly: >=3
// backticks/tildes, <=3-space indent tolerance. This is a single-line
// fence-OPENER/CLOSER probe, not a multiline fence-block-strip regex — it
// does not trip `local/no-adhoc-markdown-parsing`'s fenceRegex fingerprint
// (no `[\s\S]` multiline body in the pattern).
const FENCE_DELIM_RE = /^( {0,3})(`{3,}|~{3,})(.*)$/;
const OPEN_TAG_RE = /^gsd:section(?=\s|$)/;
const CLOSE_TAG = '/gsd:section';
const ID_RE = /^[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$/;
/**
* Parse a candidate marker's attribute string (everything after `gsd:section`,
* already trimmed) into a `key -> value` map, or `null` if it does not
* consist entirely of zero-or-more `key="value"` tokens (attribute order is
* free; spacing around `=` and between tokens is flexible — Postel on
* FORMAT). Returns `null` on a duplicate attribute key too.
*
* @param attrsPart - the marker's attribute text, e.g. `id="x" when="always"`
*/
function parseAttrs(attrsPart: string): Map<string, string> | null {
const attrs = new Map<string, string>();
let remaining = attrsPart;
const ATTR_RE = /^\s*([A-Za-z][A-Za-z0-9_-]*)\s*=\s*"([^"]*)"/;
while (remaining.length > 0) {
const m = ATTR_RE.exec(remaining);
if (!m) return null;
const [full, key, value] = m;
if (attrs.has(key)) return null;
attrs.set(key, value);
remaining = remaining.slice(full.length);
}
return attrs;
}
/** A `TypeError` carrying a stable {@link REASON} code alongside the human-readable message. */
export interface WorkflowFragmentsError extends TypeError {
readonly reason: string;
}
/**
* Throws a `TypeError` naming `sourcePath` (when given) and the 1-based
* `line`, carrying `reason` (one of {@link REASON}) as a typed property so
* callers/tests never need to pattern-match the message prose.
*/
function fail(sourcePath: string | undefined, line: number, reason: string, message: string): never {
const loc = sourcePath ? `${sourcePath}:${line}` : `line ${line}`;
const err = new TypeError(`workflow-fragments: ${message} (${loc})`) as TypeError & { reason: string };
err.reason = reason;
throw err;
}
/** Result of classifying a complete one-line HTML comment's inner text. */
type MarkerClassification =
| { readonly kind: 'open'; readonly id: string; readonly when: string }
| { readonly kind: 'close' }
| { readonly kind: 'none' };
/**
* Classify a complete one-line HTML comment's inner text (already stripped
* of `<!--`/`-->` and trimmed) as a `gsd:section` open attempt, a close
* marker, or "not a marker at all" — including `gsd:loop-host` and any
* other unrelated comment, which never match the `gsd:section` token and
* fall through to `{kind: 'none'}` untouched. Throws on any STRUCTURAL
* violation of a recognized open/close attempt (fail-closed grammar).
*
* @param inner - the comment's inner text, e.g. `gsd:section id="x" when="always"`
* @param sourcePath - optional file path named in thrown errors
* @param lineNo - 1-based line number named in thrown errors
*/
function classifyMarker(inner: string, sourcePath: string | undefined, lineNo: number): MarkerClassification {
if (inner === CLOSE_TAG) {
return { kind: 'close' };
}
if (inner.startsWith(CLOSE_TAG) && /^\s/.test(inner.slice(CLOSE_TAG.length))) {
fail(sourcePath, lineNo, REASON.CLOSE_WITH_ATTRIBUTES, 'close marker must not carry attributes');
}
if (!OPEN_TAG_RE.test(inner)) {
return { kind: 'none' };
}
const attrsPart = inner.slice('gsd:section'.length).trim();
const attrs = parseAttrs(attrsPart);
if (attrs === null) {
fail(sourcePath, lineNo, REASON.MALFORMED_ATTRIBUTES, 'malformed section marker attributes');
}
const extraKeys = [...attrs.keys()].filter((k) => k !== 'id' && k !== 'when');
if (extraKeys.length > 0) {
fail(sourcePath, lineNo, REASON.UNRECOGNIZED_ATTRIBUTE, `unrecognized attribute "${extraKeys[0]}" on section marker`);
}
const id = attrs.get('id');
const when = attrs.get('when');
if (id === undefined) {
fail(sourcePath, lineNo, REASON.MISSING_ID, 'section marker missing required "id" attribute');
}
if (when === undefined) {
fail(sourcePath, lineNo, REASON.MISSING_WHEN, 'section marker missing required "when" attribute');
}
if (!ID_RE.test(id)) {
fail(sourcePath, lineNo, REASON.MALFORMED_ID, `section marker "id" value "${id}" does not match ${ID_RE}`);
}
if (!WHEN_VOCABULARY.includes(when)) {
fail(sourcePath, lineNo, REASON.UNKNOWN_WHEN, `section marker "when" value "${when}" is not in the frozen WHEN_VOCABULARY`);
}
return { kind: 'open', id, when };
}
/**
* Parse a workflow document's `<!-- gsd:section -->` markers into a
* document-order partition of {@link WorkflowSection}s. See the module doc
* comment for the full grammar, partition invariant, and fence/comment
* interleaving discipline.
*
* @param content - full workflow markdown source
* @param sourcePath - optional file path named in thrown errors
*/
export function parseWorkflowSections(content: string, sourcePath?: string): WorkflowSection[] {
const lines = splitLinesPreservingEol(content);
const sections: WorkflowSection[] = [];
let fence: { char: '`' | '~'; len: number } | null = null;
let inComment = false;
let currentOpen: { id: string; when: string; startLineIndex: number } | null = null;
const seenIds = new Set<string>();
let gapCounter = 0;
let cursor = 0;
const joinRange = (from: number, to: number): string => {
let out = '';
for (let k = from; k <= to; k++) {
out += lines[k].text + lines[k].eol;
}
return out;
};
const flushGapBefore = (nextIndex: number): void => {
if (nextIndex > cursor) {
sections.push({
id: `gap-${gapCounter}`,
when: 'always',
body: joinRange(cursor, nextIndex - 1),
explicit: false,
startLine: cursor + 1,
});
gapCounter += 1;
}
};
for (let i = 0; i < lines.length; i++) {
const lineNo = i + 1;
const rawText = lines[i].text;
if (fence !== null) {
// Inside a real fence: only a matching closer can end it. Any
// `<!--`/`-->` on this line is fence content, never a comment
// boundary (row 5/6 of 50-test-matrix.md).
const m = FENCE_DELIM_RE.exec(rawText);
if (m) {
const char = m[2][0] as '`' | '~';
const len = m[2].length;
const trailing = m[3];
if (char === fence.char && len >= fence.len && /^\s*$/.test(trailing)) {
fence = null;
}
}
continue;
}
if (inComment) {
// Inside a real (unrelated) comment: only '-->' can end it. Any
// fence delimiter on this line is comment content, never a fence
// boundary (row 7 of 50-test-matrix.md).
if (rawText.includes('-->')) inComment = false;
continue;
}
const trimmed = rawText.trim();
if (trimmed.startsWith('<!--')) {
// Persistence is decided by whether a close token appears ANYWHERE on
// the line, not by whether the line ENDS with one — see the module
// doc comment's "deliberate refinement" note. Marker-hood additionally
// requires the comment to be the line's ENTIRE content.
const hasClose = trimmed.includes('-->');
if (hasClose && trimmed.endsWith('-->')) {
const inner = trimmed.slice(4, trimmed.length - 3).trim();
const classification = classifyMarker(inner, sourcePath, lineNo);
if (classification.kind === 'open') {
if (currentOpen !== null) {
fail(sourcePath, lineNo, REASON.NESTED_SECTION, `nested gsd:section marker (already inside "${currentOpen.id}")`);
}
if (seenIds.has(classification.id)) {
fail(sourcePath, lineNo, REASON.DUPLICATE_ID, `duplicate section id "${classification.id}"`);
}
flushGapBefore(i);
seenIds.add(classification.id);
currentOpen = { id: classification.id, when: classification.when, startLineIndex: i };
cursor = i + 1;
} else if (classification.kind === 'close') {
if (currentOpen === null) {
fail(sourcePath, lineNo, REASON.UNMATCHED_CLOSE, 'unmatched /gsd:section close marker');
}
sections.push({
id: currentOpen.id,
when: currentOpen.when,
body: joinRange(currentOpen.startLineIndex + 1, i - 1),
explicit: true,
startLine: currentOpen.startLineIndex + 1,
});
currentOpen = null;
cursor = i + 1;
}
// classification.kind === 'none': ordinary self-contained comment
// (e.g. a one-line `gsd:loop-host` or unrelated comment) — no state change.
}
if (!hasClose) {
inComment = true; // multi-line: stays open until a later '-->'
}
continue;
}
const fenceMatch = FENCE_DELIM_RE.exec(rawText);
if (fenceMatch) {
const char = fenceMatch[2][0] as '`' | '~';
const trailing = fenceMatch[3];
// CommonMark §4.5: a backtick fence opener's info string must not
// itself contain a backtick.
if (!(char === '`' && trailing.includes('`'))) {
fence = { char, len: fenceMatch[2].length };
}
}
}
if (currentOpen !== null) {
fail(sourcePath, currentOpen.startLineIndex + 1, REASON.UNCLOSED_SECTION, `unclosed gsd:section marker "${currentOpen.id}"`);
}
flushGapBefore(lines.length);
return sections;
}
/**
* Map parsed sections to `context-composer` fragments. Every strategy is
* `{kind: 'verbatim'}` (design row 23 / test matrix rows 26-29): non-
* lossiness in this phase is a STRUCTURAL guarantee of the strategy choice,
* never a large-budget trick.
*
* @param sections - document-order sections from {@link parseWorkflowSections}
*/
export function toFragments(sections: readonly WorkflowSection[]): contextComposer.Fragment[] {
return sections.map((section) => ({
id: section.id,
content: section.body,
strategy: { kind: 'verbatim' as const },
}));
}
/**
* Concatenate a {@link contextComposer.ComposeResult}'s fragment contents,
* in declaration order, back into a document. Every fragment here is
* `verbatim` with an empty wrapper, so this is a plain join.
*
* @param result - the plan returned by `composeWithinBudget`
*/
export function renderFragments(result: contextComposer.ComposeResult): string {
return result.fragments.map((f) => f.content).join('');
}
/**
* THE emission entry point: parse -> toFragments -> composeWithinBudget ->
* render. `budget` defaults to `Number.MAX_SAFE_INTEGER` (no pressure).
* Because every fragment is `verbatim`, the output is identical regardless
* of the budget value (design row 23) — this is never relied upon as the
* source of non-lossiness; the strategy set is.
*
* @param content - full workflow markdown source
* @param opts - `sourcePath` named in thrown parse errors; `budget` in bytes
*/
export function composeWorkflow(content: string, opts: { sourcePath?: string; budget?: number } = {}): string {
const { sourcePath, budget = Number.MAX_SAFE_INTEGER } = opts;
const sections = parseWorkflowSections(content, sourcePath);
const fragments = toFragments(sections);
const composed = contextComposer.composeWithinBudget({
fragments,
budget,
measure: (text: string) => Buffer.byteLength(text, 'utf8'),
options: { charsPerUnit: 1 },
});
return renderFragments(composed);
}