Fold augment's runtime-literal conversion branches onto descriptor-driven
hostBehaviors and delete dead code:
- Site A (_applyRuntimeRewrites case 'augment'): the 4 ~/.augment dot-dir
regexes now derive from getDirName('augment') via escapeRegExp (byte-
identical; getDirName('augment')==='.augment') — no runtime literal.
- Site B (applyRuntimeContentRewritesForCommandsInPlace): the
`if (runtime==='augment')` markdown-converter branch now reads
runtime.hostBehaviors.commandBodyConverter and dispatches through a local
COMMAND_BODY_CONVERTERS map (degrade-closed on unknown/absent name).
- Deleted dead `claudeToAugmentTools` map (zero refs; orphaned by ADR-1508
single-sourcing) and the unreachable `else if (isAugment)` agent-conversion
branch (augment ∈ _DESCRIPTOR_AGENTS_RUNTIMES → gated out upstream).
- Incidental orphan cleanup (no-defer): removed the equally-unreachable
`else if (isTrae)` agent-conversion arm left behind by trae's already-merged
migration #2094 (trae ∈ _DESCRIPTOR_AGENTS_RUNTIMES, same upstream gate).
copilot/windsurf/codebuddy arms are removed by their own pending migrations.
UPGRADE 3 (transport:mcp): register the GSD companion MCP server in Augment's
settings.json under mcpServers.gsd (Augment hosts MCP in settings.json, not a
standalone file). mergeGsdMcpServerIntoSettings mutates the in-memory settings
object finishInstall already writes (gated on hostBehaviors.mcpCompanion===
'settings-json'); non-destructive + idempotent; symmetric uninstall removal.
settings.json is golden-excluded, so no golden change. UPGRADE 1 (named/
background dispatch) + UPGRADE 2 (settings-json hook bus, Claude dialect)
were already live in production — this adds tests exercising both.
Golden: byte-identical for all 16 runtimes (folds preserve regex behavior;
MCP lives in golden-excluded settings.json) — verified by a real double-install
tree diff. Tests: declarative-reference-augment (adapter/axes/fail-closed/
undocumented-sub-axes + source-grep guard scoped to conversion-logic branches)
+ augment-upgrades (dispatch negotiation, hook-bus live install, MCP add/
idempotent/preserve/uninstall). Matrix + connect-gsd-mcp-server + a stale
install-on-your-runtime hook-ownership claim corrected; changeset (Changed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fold all runtime==='kimi'/isKimi logic branches into descriptor-driven
hostBehaviors (localInstallDeferred, verificationStyle, agentManifestStyle,
reapplyCommand, doneBannerStyle) + add 'kimi' to _DESCRIPTOR_AGENTS_RUNTIMES.
Kimi's skills/kimi-agents dispatch was already descriptor-driven (converter-by-
name + kimi-agents kind). Zero isKimi/runtime==='kimi' branches remain.
UPGRADE 1 (native hook bus): new hooksSurface 'kimi-hooks-toml' + a marker-
delimited config.toml [[hooks]] emitter (buildKimiHooksTomlBlock/writeKimiHooksToml
in runtime-hooks-surface.cts; resolveKimiHooksTomlDir in runtime-homes.cts).
GSD's lifecycle hooks now wire into Kimi's native ~/.kimi/config.toml (Context7-
confirmed path) at SessionStart/PreToolUse/Stop/PreCompact/SubagentStart/
SubagentStop — kimi becomes a hooks/ consumer (the 3 && !isKimi exclusion guards
removed). config.toml holds absolute install paths so it's golden-excluded via
an exact relative-path (.kimi/config.toml), not a basename (which would blind
Codex's config.toml). New hooksSurface value added to the closed enum in
capability-validator + runtime-config-adapter-registry.
UPGRADE 2 (background dispatch): flip dispatch.backgroundDispatch true (Kimi's
Agent tool takes run_in_background; root agent already gets the Agent tool), so
negotiation no longer flattens dispatch. subagentToolkit stays 'undocumented'
per AC (coder/explore/plan have distinct tool policies).
MCP transport explicitly deferred (no installer-driven MCP for any runtime).
Golden: only kimi.json changes (hooks/ scripts now installed); all 15 others +
claude-local byte-identical (kilo/zcode keep their own exclusions). Tests:
kimi-imperative-reference (adapter/axes/fail-closed/hostBehaviors + source-grep
guard) + kimi-upgrades (config.toml [[hooks]] SessionStart + marker idempotency
+ backgroundDispatch negotiation). CONTEXT.md glossary + matrix + how-to updated;
changeset (Added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fold trae logic branches into descriptor-driven reads: skipSharedHooksInstall
gates (dropped && !isTrae), and the case 'trae' path-rewrite arm now computes
the self-alias from the descriptor-driven dirName (.trae). Dead isTrae bindings
removed from uninstall/writeManifest/finishInstall. trae's skills dispatch was
already descriptor-driven (converter-by-name). RUNTIME_CONTENT_DISPATCH.trae is
left as a runtime-keyed table registration (its regex/callback rewrites can't be
a byte-identical descriptor map — matches cursor/windsurf/cline). trae stays in
RUNTIME_FLAG_IDS: isTrae still gates the agents-converter selection (agents out
of scope; removal gated on the cross-runtime agents-dispatch migration).
Byte-identical golden parity for all 16 runtimes.
UPGRADE: SOLO stage/trigger metadata — emitted Trae SKILL.md now carries
stage: workflow (descriptor-gated via hostBehaviors.soloStageMetadata) so
Trae's SOLO Agent can auto-invoke GSD skills at the corresponding stage. Field
shape is best-effort/inferred (Trae publishes no formal schema). trae.json
golden regenerated.
Tests: trae-imperative-reference (adapter/axes/fail-closed shouldFlattenDispatch
+ no runtime==='trae' source-grep, isTrae exempted for agents) + trae-upgrades
(stage: workflow on installed SKILL.md, descriptor-gated). Matrix note +
changeset added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Fold remaining isKilo logic branches into descriptor-driven reads:
finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks
copy), and a skills converter-name registry (the artifactLayout.converter
field is now load-bearing, not decorative). frontmatterDialect stays the
documented dispatch key for frontmatter (no descriptor field for it). Dead
isKilo destructure bindings removed. Byte-identical golden parity for all 16
runtimes (opencode, which shares kilo's combined-family path, verified clean).
UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin +
extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus).
UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread
modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped
from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's
mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is
the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented'
per AC so dispatch degrades to 'degraded' by design.
Model-catalog single-source edit ripples the shared model-catalog.json hash
into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale-
bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/
codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the
connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error.
Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/
hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades
(plugin parity+load, model-override converter, agents dispatch surface, MCP
doc). Matrix + how-to + config docs updated; changeset added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Record resolved UI-state considerations in the UI-SPEC, backward-compatibly
(Phase 2, WIRE-02). templates/UI-SPEC.md gains a '## UI Considerations' section
(analog of SPEC '## Edge Coverage') after '## Copywriting Contract' — a
| Category | Element(s) | Status | Resolution / Reason | table with
covered/backstop/unresolved rows in the locked probe-core projectTruths format
the shipped plan-phase lift (plan-phase.md:921) reads. Empty/error COPY stays in
Copywriting; this section covers shape-rooted STATE and references those rows
(de-dup).
Tests: docs-fixtures parsed-heading assertion (UI Considerations present +
distinct from Copywriting Contract; allow-test-rule, parsed structure); typed
backward-compat (projectTruths(undefined/[])===[], old UI-SPEC still plans —
Hyrum), format-match, and idempotency (proposeElements determinism). The
template-structure test was the RED driver.
Install-parity cascade (INVENTORY-MANIFEST + 16 golden fixtures +
agent-size-baseline) deferred to Phase 3 SHIP-01, as planned.
Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS
Add the live ui-phase producer path for the UI-consideration probe (Phase 2,
WIRE-01). Two small exports on the Phase-1 adapter — proposeElements (the
propose-then-confirm view of detected kinds + applicable categories) and
autoResolve (the deterministic --auto floor that never dismisses and never
auto-backstops an unclassified item, #1110) — plus a post-verification
'## 9.5 UI-Consideration Probe' step in ui-phase.md mirroring spec-phase 5.5's
RUNTIME_DIR shim + fatal-invoke/malformed-report/zero-applicable fail-closed
guards, propose-then-confirm (the partial-cue recall mitigation), and the
'## UI Considerations' write-back in the shipped plan-phase.md:921 lift format.
autoResolve is the CODE floor; the covered-upgrade stays workflow prose (the
two-layer --auto). Un-upgraded backstops route to insufficient_spec ->
human_needed at verify, never a silent pass (#1154).
Tests: +8 typed (proposeElements shape/determinism, autoResolve never-dismiss,
partial-cue strict-subset) — structured-value only. ui-phase.md size baseline
ratcheted 15477->24447 (under DEFAULT cap). The plan-phase.md PRE_PHASE6 ceiling
stays RED pending #1852 (unchanged from Phase 1).
Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS
Reference doc mirrors edge-probe.md structure but links rather than re-argues;
states the MIXED-axis boundary (closed compiled shape-rooted subset here; open
UX subset - real-time/offline, a11y depth, i18n/RTL - prose-owned in
domain-probes.md). Docs-parity test pins doc taxonomy ids == code UI_TAXONOMY
ids and asserts disjointness from domain-probes.md topics (ADR-456
runtime-contract exemption, see #1867).
Third probe-core adapter on the UI element/state axis, mirroring edge-probe.
Closed 8-id shape-rooted UI_TAXONOMY + element-cue relevance filter
(UI_CUES -> classifyElement -> applicableCategories); unclassified fail-loud
soft-signal (#1110); fail-closed on invalid authored element kinds. All
lifecycle/merge/validation delegated to probe-core verbatim (no fork). Item
question carried in the shared Item.probe field. LIFT-01 proven at the
probe-core primitive level (projectTruths/dispositionForUnverifiableTruth):
backstop considerations route to insufficient_spec, never a silent pass.
Tests assert typed returns off the built .cjs (no source-grep).
#2056 fixed the foreign-prefix collapse for init plan-phase only.
The identical defect remained in three sibling commands that called
findPhaseInternal/getRoadmapPhaseInternal without the guard:
- cmdInitExecutePhase
- cmdInitVerifyWork
- cmdInitPhaseOp
Extracted the #2056 guard into shared helpers (guardedFindPhase /
guardedGetRoadmapPhase) and routed all four init commands through them.
Deleted the local parsePhasePrefix/isForeignPrefixedPhaseQuery copies in
init.cts — the canonical export from phase-id.cts is now used directly,
eliminating the drift risk flagged by both reviewers.
Added 5 regression tests (3 reject + 2 accept-branch) mirroring the
#2056 plan-phase tests.
Update runtimeTierDefaults.codex and providerPresets.openai in
model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna),
advancing from the superseded GPT-5.4/5.5 generation.
Model IDs verified against OpenAI developer API docs:
- gpt-5.6-sol: flagship, /, reasoning xhigh
- gpt-5.6-terra: balanced, .50/, reasoning medium
- gpt-5.6-luna: fast/cheap, /, reasoning medium
Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast),
so profile semantics are unchanged — only the underlying IDs advance.
Updates: catalog JSON, test assertions (catalog defaults), docs
(CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings),
and changeset.
Closes#2122
Final convergence review found the shared <tag> seam introduced 3 behavior
regressions; fixed all + locked with tests:
- #557 REGRESSION: stripTaggedBlocks's attribute-tolerance stripped `<details open>`
(the ACTIVE-milestone marker) that the old `<details>`-only regex preserved. The
seam now takes `allowAttributes` (default false) — details/decisions strip is
attr-INTOLERANT (preserves `<details open>`); only `<task type="…">` opts in.
Regression test added to roadmap-parser + markdown-sectionizer suites.
- verify.cts actionZones (negative-grep-echo security scan): reverted to a bounded
to-first-close scan `<action>([\s\S]{0,20000}?)</action>` so a grep-echo trick
can't hide behind an unterminated inner <action> (the seam's stop-at-next-open
would drop it). ReDoS-safe via the cap.
- check-command-router HTML-comment strip: `(?:-->|$)` fallback wiped to EOF
(fail-closed spurious gate block) — replaced with stop-at-next-open so an
unclosed `<!--` leaves downstream tags intact.
- Updated the extractTaggedBlocks nested-tag tests to the new (stop-at-next-open)
behavior: `<x><x>inner</x></x>` -> ['inner'].
All vectors still linear; every fix verified in-process.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Review caught that the prior commit bounded only the paren tag clause and left
the SIBLING bracket-prefix `(?:\[[^\]]+\]\s*)?` (same host regexes, before Phase)
UNBOUNDED — the identical quadratic reachable via a `[...]` run (measured ~16s at
1.7MB). Bound `[^\]]+`/`[^\]]*` -> {1,200}/{0,200} across all 19 phase/milestone
heading prefixes. Comprehensive re-measurement now shows EVERY vector linear
(bracket/paren/id/name/milestone all ~2-44ms at 2.45MB; bracket scaling
2k->2ms, 4k->5ms, 8k->10ms). Also: update the #1729 literal-mirror parity test
off its stale unbounded constant, and add limit-1 (199) boundary coverage.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The canonical OPTIONAL_PHASE_TAG_SOURCE tag clause `(?:\s*\([^)\n]*\))?` (and its
inlined literal mirrors across 11 modules) had an UNBOUNDED body, making the
optional-group + /g header scan quadratic on adversarial ROADMAP.md/STATE.md — a
long run of `(` after a header ran ~18.8s at 1.7MB. Bound the body to {0,200} in
the constant AND every mirror in lockstep (the #1729 "both forms change together"
contract), so the scan is linear: the same 1.7MB input now resolves in ~9ms
(measured), while real tags (a handful of chars) still match and a 201-char tag
is rejected. Added a #2128 boundary regression to the #1729 suite.
Pre-existing (byte-identical before/after the Phase 4 migrations); folded in at
maintainer direction rather than deferred.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Re-review found the `// phase-id-owner:` suppression treated a `//` embedded in a
string literal as a comment — help/doc text quoting the sanction syntax (the exact
string the scanner's own main() prints) would silently suppress a real
re-derivation. Require the marker to LEAD its own comment line (`^\s*//…`), so a
`//` inside a string or trailing a code line never counts. All 5 real sanctions
are already dedicated lines (scanRepo stays green); trailing same-line sanctions
are no longer honored — put the comment on the line directly above.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Correctness review of the Phase 4 guard found the allowlist over-broad and the
scanner/guards evadable. Fixed all findings:
- Migrate 9 sites that were wrongly sanctioned: their regex is the PURE canonical
token (`\d+[A-Z]?(?:\.\d+)*`, no variant), byte-identical to already-migrated
siblings. The old justification argued against swapping to the extractPhaseToken()
FUNCTION (behavior-risky) — but the guard only wants the same regex built from
the SOURCE string (byte-equal, zero risk). Coverage is now 32 migrated / 5
sanctioned, not the overstated 23 / 14 (audit.cts x3, uat.cts, init.cts x4,
roadmap-upgrade.cts). Each conversion proven byte-equal (.source + .flags).
- Harden the drift detector: also catch the `[0-9]`-in-place-of-`\d` variant;
document the accepted limits (cross-line split, semantic restructuring —
covered by the identity guard + review, not a text scan).
- Sanction robustness: a `phase-id-owner:` marker now counts only inside a `//`
comment (a bare substring in a string no longer suppresses a real flag), and
the preceding-line window skips blank lines (an auto-formatter's blank line no
longer reactivates the flag).
- roadmap-parser.cts:462 comment: corrected — that regex carries no /i flag, so
its [A-Za-z] class does real case work (matches state.cts:1409's rationale).
- Identity guard: surface require failures instead of silently skipping, and
floor coverage at >75% of consumer modules (inspects 156/157).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Phase 4 of epic #2121 (ADR-2121 Decision 7), closing the recurrence loop that
produced #2111 / #2114 / #2104: no module outside src/phase-id.cts may
re-implement phase-ID parsing without failing CI.
- phase-id.cts: add PHASE_NUMBER_TOKEN_SOURCE — the canonical phase-number-token
grammar (\d+[A-Z]?(?:\.\d+)*) for enumeration/scan call sites, the ANY-phase
counterpart to phaseMarkdownRegexSource(n)'s known-number lookup. Extend-only
(never touches normalizePhaseName; blast radius 79 fns / CRITICAL).
- scripts/lint-phase-id-drift.cjs: pure findPhaseIdRegexDrift(text) + scanRepo(root),
wired to `npm run check:phase-id-drift`. Flags a literal re-derivation of the
canonical token (both /\d/ and new-RegExp `\\d` escaping, plus the [A-Za-z] and
[.-] near-variants) anywhere in src/** outside phase-id.cts, unless sanctioned
with `// phase-id-owner: <reason>`. Narrow by design: bare \d+, digits-only
captures, \w ids, status-message text and pipe-tables are not flagged.
- tests/phase-id-drift-guard.test.cjs: fail-first drift cases (AC1) + live
scanRepo(ROOT) zero-drift (AC3) + identity guard — phase-id.cjs exports the
complete locked surface and no consumer re-exports a divergent copy (AC2).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Re-review found the test comment + changeset prose inaccurately claimed a bare
query "always" surfaced malformed_roadmap. Empirically, on origin/next a
project-code-prefixed checklist entry was a silent {found:false} for BOTH query
forms — the prefixed pass discarded its malformed candidate and the bare regex
could not match the PROJ- prefix at all. The unified 3-source lookup newly grants
the diagnostic to both forms; correct the prose to say so. No logic change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adversarial review of the Phase 3 branch surfaced three verified defects; fix
all three in place (no defer):
- install-runtime-artifacts.test.cjs: finish the fold-triplication dedup started
earlier (only enh-1511 had been collapsed). 11 B1-batch __foldDescribe blocks
were byte-identical triplicates (~5.9k lines, ~49% of the file), tripling the
subprocess-spawning installer suites under --test-concurrency — the same
starvation that produced the temp-dir races this branch fixes. Byte-identity
verified per block before removal; 230 distinct test/it titles preserved
(origin/next: 230 -> 230), interleaved B3/B5/B6 singletons untouched.
- config-get-default.test.cjs: make runExpectError faithful to production. The
throwing process.exit seam was caught by cmdConfigGet's "No config.json"
guard and reclassified into a spurious 2nd error() with the wrong reason
(CONFIG_PARSE_FAILED). Drive io.setJsonErrorMode + carry the original message
on the sentinel so the guard re-throws (single fire), assert exitCount===1,
and strengthen both probes to assert the typed reason (CONFIG_NO_FILE /
CONFIG_KEY_NOT_FOUND).
- roadmap.test.cjs: lock the #2121/#2114 malformed_roadmap parity — a
project-code-prefixed query against a checklist-only roadmap now surfaces the
same diagnostic a bare query always did (fails on prior silent-empty behavior).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The prohibition-enforcement real-runner tests linted src/clock.cts (a .cts) as
their clean target. Under eslint.config.mjs's type-aware block for src/**/*.cts
(recommendedTypeChecked + parserOptions.project: tsconfig.build.json), each eslint
spawn loaded the WHOLE tsconfig.build.json program (~2s, CPU-heavy). The
real-runner tests spawn eslint repeatedly; under --test-concurrency those
full-program type-checks oversubscribed the bench CPU and blew the 60s subprocess
bound -> fail-closed (intermittent, load-dependent — passed 24241/24241 in an
earlier run, failed here).
Root fix (not a retry/timeout bandaid; measured projectService = no faster since
a single-file .cts lint still loads type info): add tests/_ff_lint_clean.cjs, a
KNOWN-CLEAN lint-scoped .cjs companion to _ff_lint_violation.cjs, with a
flat-config block enabling local/no-source-grep so the clean pass stays
non-vacuous. Repoint the 6 src/clock.cts real-runner usages (5 targets + the FF-02
toothless violationFixture) at it. Each spawn is now ~0.8s non-type-aware (no
whole-program load) — starvation removed. All 6 tests' semantics verified
in-process (SF-01 greens; toothless/fail-closed stay unverified); full-repo
`eslint .` green.
Refs #2126, #1259
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Phase 3's gsd-test surfaced 8 pre-existing test-isolation races (in #2090's test
files now on next). Per CLAUDE.md's no-defer rule these are fixed inline in the
current change. Root-caused via /qa-test-architect — all bad-test (the
rewrite-engine production code is race-free):
- install-runtime-artifacts.test.cjs: the "rmSync when readFileSync throws" test
diffed the SHARED os.tmpdir() for gsd-cmd-rewrites-* dirs and force-deleted any
new one with no ownership check. Under --test-concurrency it deleted a sibling
test file's LIVE tempDir mid-copy (the #1575 "ENOENT .../graphify.md") and
misattributed it as its own leak. Fixed: capture the exact tempDir THIS call
creates (fs.mkdtempSync monkeypatch, restored in finally) and assert only on
that — never sweep/delete the shared os.tmpdir(). Also deduped the enh-1511
block the #1969 consolidation folded in 3x byte-identically (#1970/#1974/#1975)
down to 1 copy; 308 unique test titles unchanged (verified).
- issue-1575-agent-descriptor-parity.test.cjs: a missing }); nested the M2
'cursor attribution' test inside the per-runtime loop so it ran 7x (widening
the tempDir window). Fixed the brace -> runs once as a describe sibling.
- config-get-default.test.cjs: local run()/runRaw() spawned node via
execFileSync with a fixed 5s timeout and no retry -> ETIMEDOUT under Docker
load. Redesigned to call cmdConfigGet in-process (fs.writeSync fd-capture +
process.exit sentinel, both restored in finally) — no subprocess, no wall clock.
- runtime-artifact-conversion.cts: fixed the stale "No production caller today"
JSDoc on rewriteStagedCommandBodies (real callers: applySurface,
createRuntimeArtifactInstallPlan) — the false doc invited the bad test.
Refs #2126, #2090
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Phase 3 of epic #2121. cmdRoadmapGetPhase and getRoadmapPhaseWithFallback now
iterate the shared roadmapPhaseLookupSources (exact -> numeric -> prefix-tolerant,
owned by phase-id.cts since Phase 1) instead of a hand-rolled 2-source lookup, so
all three roadmap resolvers share one resolution contract.
Drives #2114: `roadmap get-phase <bare-N>` now resolves a drifted
`### Phase AB-29:` heading (matching getRoadmapPhaseInternal / init.phase-op),
previously EMPTY from the CLI. The malformed_roadmap checklist-fallback and the
milestone-then-full precedence are preserved (a milestone checklist never blocks
a full-roadmap header match).
Behavior reversal (approved in-session): a bare query now also resolves a
*drifted-only* prefixed heading when no bare sibling exists, reversing the #3599
counter-test's expectation. #3599's real anti-steal intent (a bare sibling wins
over a distinct prefixed one) is preserved by the exact->numeric->prefix-tolerant
ordering and re-asserted in the updated test; a new #2114 block covers the
drifted-only case. Fail-first demonstrated.
Closes#2126
Refs #2121
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Orthogonal review surfaced that resolvePhaseIdForCompletePhase (state.cts) and
cmdStateCompletePhase's idempotency check still used an unanchored
/(\d+[A-Z]?(?:\.\d+)*)/i — even more permissive than the parseProsePhaseField
regex this phase fixes. Reachable corruption: after `milestone complete v0.5`,
`state complete-phase` (no --phase) mined "0.5" from the body line
"Phase: Milestone v0.5 complete" and rewrote STATE.md as "Phase 0.5 complete".
Both sites now delegate to phase-id.cts:parsePhaseFromProse (the same anchored
parser), so a milestone-closure line yields no token and the existing
"unable to resolve" guard fires instead of corrupting. Canonical tokens
(3, 03, 3A, 3.3, "3 of 5", "1 — Setup") are preserved unchanged.
Regression (tests/state.test.cjs, complete-phase suite): `state complete-phase`
on a "Milestone v0.5 complete" STATE.md now rejects and does not mine "0.5".
Demonstrated fail-first.
Refs #2125, #2121
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Phase 2 of epic #2121. state.cts:parseProsePhaseField now delegates to the
anchored phase-id.cts:parsePhaseFromProse (built in Phase 1), removing this
module's independent prose phase-id regex.
Drives #2111: `milestone complete vX.Y` no longer corrupts current_phase. The
body line "Phase: Milestone v0.5 complete" previously had "5" mined from it by
the unanchored regex; the anchored parser returns { phase: null }, so
syncStateFrontmatter's #905 guard preserves the real current_phase. This also
fixes the broader family the review surfaced — every milestone completion
(e.g. v1.0 -> "0") was silently corrupting current_phase, not just .5-versions.
Regression (tests/milestone.test.cjs, in the milestone-complete suite, #2111):
`milestone complete v0.5` on a project with current_phase: "19" now preserves
"19". Demonstrated fail-first end-to-end: reverting the migration reproduces
current_phase = "5".
Closes#2125
Refs #2121
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Orthogonal security review of the Phase 1 surface found two issues; both fixed
and regression-tested:
- MEDIUM ReDoS: the name-extraction regexes /\(([^)]+)\)/ and
/—\s*([^(\n]+?).../ backtrack O(n^2) on a crafted STATE.md field value with a
long unterminated "(" / "—" run (reviewer measured ~38s at 320k chars).
Length-bound both quantifiers to {1,200} -> linear (320k now ~100ms). A real
phase name is far shorter than the cap.
- LOW: parsePhaseFromProse threw on non-string truthy input, unlike its three
sibling #2121 functions. Coerce via String(value) up front.
The identical ReDoS regexes are copied verbatim from the pre-existing
state.cts:parseProsePhaseField; per the no-defer rule that surfaced defect is
fixed inline there too (Phase 2 / #2125 later supersedes that function by
delegating to the bounded phase-id.cts parser).
Adds a behavioral bound-guard regression test (a >200-char parenthetical is not
extracted) and a non-string-coercion test.
Refs #2124, #2121
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add the ADR-2121-locked canonical functions to src/phase-id.cts. No consumer
behavior changes — Phases 2-4 migrate the divergent call sites against them.
- parsePhaseFromProse: anchored prose parser. A phase is returned only when the
STATE.md "Phase:" field VALUE begins with a phase token, so
"Milestone v0.5 complete" yields { phase: null } instead of "5" (the #2111
root cause: the old unanchored \b(\d+..)\b mined the minor-version digit).
Name extraction (parenthetical / em-dash tail, minus status words) unchanged.
- stripConfiguredProjectCodePrefix / isForeignPrefixedPhaseQuery: config-aware
prefix policy. A foreign prefix (MEM-01 when the configured code is LKML) is
preserved rather than collapsed to a bare numeric phase — the #2104 fix's
canonical home (consumed later, outside this epic's critical path).
- roadmapPhaseLookupSources: moved from roadmap-parser.cts so phase-id.cts is
the single owner of the exact -> numeric -> prefix-tolerant ordering.
roadmap-parser.cts now imports it (behavior-identical); its two now-unused
imports (phaseMarkdownRegexSourceExact, OPTIONAL_PROJECT_CODE_PREFIX_SOURCE)
are dropped.
Tests: subject-named suites in tests/phase-id.test.cjs covering the ADR
boundary set (v0.5, v1.0, MEM-01, AB-29, bare 29, zero-padded 029) plus two
fast-check properties: the #2111 "Milestone vX.Y complete never yields a phase"
invariant and a parse/normalize property.
Extend-never-mutate: the 12 pre-existing phase-id.cts exports are unchanged.
Closes#2124
Refs #2121
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
MEDIUM fixes (code review):
- Wire resolveManagedHookEvents + resolveHookScripts + buildHookBusEntries
from imperative-hook-bus.cts into writeCursorHooksJson — the install path
is now truly descriptor-driven (reads hostBehaviors.managedHookEvents),
not a hardcoded constant that happens to match the descriptor. bin/install.js
passes the descriptor list via opts.managedHookEvents.
- buildHookBusEntries is now consumed (was dead code); entry-building is no
longer duplicated inline.
- Remove try/finally from cursor-hook-bus-upgrade.test.cjs test bodies
(violated CONTRIBUTING.md L342; redundant with t.after cleanup).
LOW fixes:
- Remove dead require('fs')/require('path') from gsd-cursor-pre-tool.js
- Fix resolveManagedHookEvents docstring (all-invalid fallback behavior)
- Add src/runtime-hooks-surface.cts to the AC2 source-guard file list
Security review: no CRITICAL/HIGH/MEDIUM findings (3 LOW are pre-existing
#777 baseline patterns, not regressions).
Replaces the timing-dependent execFileSync(timeout:5000) approach with a
spawnSync-based runHook that tests the scanner's RESULT (exit code + output
shape), never how long it takes.
Root design flaw in the prior approach: execFileSync's timeout (5000ms)
was identical to the scanner's own internal setTimeout(5000ms), creating a
non-repeatable race (F.I.R.S.T. violation: not Repeatable). Under concurrent
test-chunk load — which #2089's 3 new cursor test files redistribute —
node22's event-loop scheduling let execFileSync's SIGTERM win the race,
producing err.status=null → exitCode=1 → spurious property-test failure.
Redesign (F.I.R.S.T.):
- spawnSync (not execFileSync): non-zero exits return a result object,
not an exception — cleaner for property tests
- Non-serializable payloads (BigInt, circular refs, Symbol) are SKIPPED:
the scanner receives JSON via stdin, so these values are outside its
protocol — JSON.stringify throwing is a test-harness artifact, not a
scanner defect
- 30s safety-net timeout is NOT a test assertion: scanner exits in <100ms;
30s only catches a genuinely hung process (6x the scanner's own 5s
internal timer → no race possible)
- Assertions check exit===0 and output structure, never timing
qa-test-architect pipeline: risk=HIGH (security boundary); automation=
subprocess (real shipped hook); test-cases cover happy/boundary/negative/
independence; verified via gsd-test.
The property test's execFileSync timeout (5000ms) was identical to the
scanner's own internal stdin-timeout (hooks/gsd-read-injection-scanner.js:109,
also 5000ms). Under concurrent test-chunk load on linux-node22 — which #2089's
3 new cursor test files redistribute — the scanner subprocess's stdin 'end'
event can fire late enough that execFileSync's SIGTERM arrives before the
scanner's own process.exit(0), producing err.status=null → exitCode=1 →
spurious property-test failure.
The scanner has no process.exit(N!=0) paths; the only non-zero exit is from
the signal-kill race. Doubling the test ceiling to 10000ms gives the scanner's
5000ms internal exit a 5s buffer to win the race deterministically on every
node version.
- Add /gsd-core/bin/lib/host-integration-adapters/imperative-hook-bus.cjs to
.gitignore (tsc-emitted build artifact per ADR-457 convention; matches the
sibling adapter entries at .gitignore:70-89). The subagent authored the
.cts source but missed this entry, leaving the compiled output untracked.
- Fix cosmetic 'Context3' -> 'Context7' typo in test section header comment.