Commit Graph

59 Commits

Author SHA1 Message Date
sim
2bead6ca1d feat(#3149): add dedicated init.debug entry point for /gsd:debug
/gsd:debug was one of the last workflows with no cmdInit* of its own: its
Step 0 made three separate round-trips (state.load, resolve-model
gsd-debugger, config-get workflow.tdd_mode) to assemble one context. Because
no debug-scoped fact was computed at any entry point, ADR-1671 admission gate
(2) could never be satisfied for debug — an applicability atom naming such a
fact would evaluate FALSE forever and silently exclude its section.

Adds cmdInitDebug (init.debug), registers it in the init router and the
command-alias table, and collapses debug.md Step 0 to one call. Every field
resolves through the same primitive the call it replaces used: loadConfig for
commit_docs, withProjectRoot for response_language (#2402), planningPaths for
debug_dir, resolveModelInternal for debugger_model, and the existing
Boolean(workflow.tdd_mode) idiom for tdd_mode.

PlanningPaths gains a debug field so state.load and init.debug share ONE
debug-directory expression rather than two kept in sync by hand. state.load
keeps emitting debug_dir: it is a shipped query surface with its own test
anchor, so narrowing it would break unseen consumers for no gain.

No WHEN_VOCABULARY atom and no gsd:section marker: gate (1), a consuming
section of at least 400 bytes, belongs to the change that adds the section.

Closes #3149

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-07 09:30:27 -04:00
Tom Boucher
fb3ee56651 fix(#3033): resolve zero-plan split-parent phase as complete when roadmap checkbox is checked (#3114)
* fix(#3033): resolve zero-plan split-parent phase as complete when roadmap checkbox is checked

A phase split into sub-phases (parent kept as shared context, zero plans
by design) was permanently stuck as 'researched' because the roadmap-
checkbox override at line 2266 required completion.phase_complete (derived
from plan/summary counts), which is always false for zero-plan phases.
The parent was permanently eligible for current-phase selection and
re-planning recommendations.

The override now fires when roadmapComplete AND planCount === 0 (the
split-parent shape), treating it as complete regardless of the plan-count
derivation. A zero-plan phase whose checkbox is still unchecked stays
in-progress (researched). Ordinary phases with plans are unchanged.

* chore(#3033): backfill changeset PR number 3114

---------

Co-authored-by: sim <sim@local>
2026-08-06 06:44:28 -04:00
Tom Boucher
53ea8e0664 fix(#3057): make a guard's failure distinguishable from its benign result — Wave 1 (#3088)
* fix(#3057): refuse the write when the duplicate scan cannot complete

writeManifest documents itself as a fail-closed duplicate guard: if any
existing manifest shares plan_id with a different, non-terminal job_id it must
refuse, because dispatching again would duplicate the external job.

It could not honour that. The scan reads every sibling manifest looking for the
duplicate, and an unreadable or unparseable sibling was `continue`d past. If
the corrupt file was the one holding the live duplicate, the scan found nothing
and a duplicate external job dispatched.

The asymmetry is what gives it away: a malformed TARGET refused with
malformed_existing because clobbering is unacceptable, while a malformed
SIBLING was skipped — yet siblings are the only thing the duplicate check
reads.

Adds a scan_incomplete verdict that refuses and names the offending file, so an
operator can quarantine or repair it. Fail-closed alone would let one stale
corrupt manifest wedge every dispatch for that planning dir permanently; naming
the file is what makes refusing survivable. malformed_existing is untouched, so
the target/sibling distinction stays visible. The docstring is updated — it
previously stated a rule the function did not keep.

memFs() gains an optional failReads map so these branches are reachable at all;
they had zero coverage because the fake could not express a per-file read
fault. The signature is additive and every existing caller is unchanged.

The regression is proved by a pair, not a single test. A control writes a
readable sibling holding a genuine non-terminal duplicate and asserts
duplicate_plan_id, establishing the scenario is real; the regression then makes
that same path unreadable and asserts scan_incomplete. A first draft of this
test used a corrupt-JSON fixture containing no plan_id at all while its comment
claimed otherwise — it duplicated the unparseable-sibling case and proved
nothing, which is the defect class this phase exists to remove.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3057): make a guard's failure distinguishable from its benign result

Wave 1 of the negative-space backfill: the branches where a guard that could
not verify something reported the same value it reports when everything is
fine. That indistinguishability is the defect; every fix here makes the two
states tellable apart, and every test proves it with a pair — one for the
failure, one for the benign case. A single test cannot establish that two
states are distinguishable, which is the whole property being fixed.

state.cts phaseInventoryProvider returned null for both a real disk-scan
failure and a genuinely empty phases dir, so `state rebuild` could report
success while phase-table reconciliation never ran. It now returns a
discriminated result and the CLI surfaces phase_inventory_scan_failed plus a
reason. The reason field turned out never to have been wired into the emitted
JSON at all — it existed only as an internal variable — so a test could only
assert on the operator-facing note. It is a real field now.

state.cts treated an unreadable lock body the same as an empty one, applying
the 1-second stealable floor. A lock we cannot read is not a lock we know is
stale; an unreadable body is now held to the deadman ceiling like a live
holder.

verification.cts findStaleVerificationSummary returned null on any fs, scan or
clock failure — meaning "not stale". It now returns a discriminated
StaleCheckResult and the caller records that the check was indeterminate.

git-base-branch resolveBaseBranch returned 'main' both when no candidate branch
existed and when every git tier timed out. A diagnostics variant now reports
whether the answer was verified, and the CLI writes an unverified-fallback note
to stderr. The stdout contract five workflows parse is untouched.

worktree-safety snapshotWorktreeInventory left exists:true when statSync threw,
so a guard that could not check reported the worktree present; exists is now
tri-state and a stat failure surfaces as an 'unverified' finding.
planWorktreePrune reported 'no_worktrees' for a parse failure, which is not the
same as an empty list — and it drives a prune. It now reports 'parse_failed'.

Fixing the inventory change exposed a second fail-open in verify.cts: the
validate-health consumer silently dropped findings whose kind it did not
recognise, so the new kind would have vanished. That is closed too — worth
noting that the survey enumerated producers of degraded verdicts, not consumers
that discard them.

worktree-base-ref and state-transition gain the distinguishing signal without
changing what they do: headAbsenceVerified, and a phase-inventory scan meta.
Whether those guards should ACT differently is a product question this change
does not answer, and both are flagged rather than quietly settled.

rescueSummaryArtifacts is left alone: rescuing on an uncertain cat-file is
deliberate per #2556. It now has tests proving it, and a recorded negative
finding — git cat-file -e returns 128 for both "absent from HEAD" and a fatal
error, so "uncertain" and "certain-and-fine" are not separable at the git
level.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): assert typed values, not rendered text

Ten assertions in the rebuild CLI suite matched substrings of produced output —
STATE.md body fields, a markdown table row, an audit-log heading, and JSON keys
read as text. CONTRIBUTING prohibits that: if the code under test produces
text, the test asserts on its structured surface instead.

No production surface had to be built. Every one already existed and was
already compiled into bin/lib: stateExtractField for body fields,
parseMarkdownTable for the phase table, collectSection for the audit-log
section, and result.data.log — already a typed RebuildLogEntry[]. The tests
were matching rendered text sitting next to the structured data.

One of those assertions was passing for the wrong reason. `stdout.includes
('rebuilt')` matched the JSON KEY name, not a value: the dry-run path emits
`mutated` and the real path emits `rebuilt`, so it would have passed whether
the value was true or false. It now asserts the value.

external-job's refusal already had to name the offending file — that naming is
why the fail-closed variant is survivable rather than a permanent wedge — but
the tests proved it by substring of a prose message. The failure result now
carries offendingPath as its own field and the tests assert it by value. The
human message is unchanged; operators read it.

Array membership is left alone. `phaseIds.includes('99')` and
`result.updated.includes('Completed Phases')` are membership checks on real
arrays, not text matching, and converting them would weaken nothing and clarify
nothing.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): execute acquireStateLock instead of grepping its source

The non-EEXIST lock test asserted on the TEXT of the built .cjs and never
called acquireStateLock. It carried an allow-test-rule: architectural-invariant
exemption to permit that. A source grep proves a literal is present in a file,
not that the behaviour works — it is weaker than a liveness test, which at
least runs the code, and it was the only coverage the fatal-errno path had.

Replaced with tests that inject the errno through fs and assert what actually
happens: a fatal EACCES propagates out of acquireStateLock with zero backoff
sleeps, while EAGAIN/EINTR/EINVAL/EIO/ENOENT/ESTALE/EPERM/EBUSY retry once and
succeed. The exemption is removed and its allowlist entry with it.

One old assertion is deliberately not carried over: it checked the retryable
errnos were expressed as a Set rather than an inline literal. That is a shape
check with no runtime signature; the behavioural tests fail if the code reverts
to the old inline check, which is the regression it was really guarding.

The #3057 lock-body tests move into that same file rather than a new one, which
is what lint-test-file-count asks for and puts every acquireStateLock test in
one place.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3057): surface an indeterminate staleness check to its callers

An isolated review caught an inconsistency inside this wave. Two of the three
"add the distinguishing signal" fixes wire through to something a user sees:
git base-branch writes an unverified-fallback diagnostic to stderr, and an
unverifiable worktree surfaces as a W020 finding. The third set
staleCheckIndeterminate on readVerificationStatus's result and nothing read it.

A signal nobody consumes leaves the fail-open exactly as silent as before: the
staleness check could fail and the operator saw precisely what they would see
if the answer were genuinely "not stale". That is the defect this issue exists
to remove, so it is not defensible as scaffolding when its two siblings in the
same change already wire through.

All five callers now surface it, each through the channel it already had rather
than a mechanism imposed uniformly: phase complete adds it to its existing
warnings array and, on the blocked path, as an additive note on the error text;
init and roadmap carry it as a field on output they already emit; the UAT
report carries it without ever gating passed/blockers; workstream inventory
takes an injectable writeDiagnostic mirroring the git base-branch idiom,
because its return shape had nowhere to hang a per-phase field without
rippling the builder's types.

The routing decision is unchanged everywhere. What changes is only that a
caller and an operator can now tell a failed check from a completed one.

That diagnostic carries structured meta rather than being asserted by regex —
the default still writes only the human message to stderr, but tests assert
phaseDir and reason by value. Two earlier assertions in this branch were
converted the same way; this was the last raw-text assertion left.

Also records a scope correction: the completePhaseCore guards now compare
stateReplaceField's result to the body instead of testing truthiness, so a
field whose substitution produced identical text no longer reports as updated.
That is a real behaviour fix, not the signal-only change this file was
described as carrying, and its tests cover both the changed and unchanged
cases.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): bound two heavy subprocesses for a loaded bench, not an idle one

The remote matrix surfaced three failures unrelated to this branch's changes.
All were bad tests, and a re-run would have hidden every one of them.

The reviewer-flags parse block bounded bash -> node -> a full gsd-tools cold
start at 5 seconds. On a bench running thirty thousand tests in parallel that
is not a hang, it is a busy machine. Raised to 30s, matching the convention
sibling suites already use for script invocations, with a comment saying what
the budget covers so nobody tightens it back. Two further copies of the same
5-second spawn in the same file had the identical defect and are raised too —
they were not in the failure report, but they will be next time.

The fragment-propagation test bounded npm run regen:derived — a full build plus
eight generators, the heaviest subprocess in the suite — at five minutes, and
node22 was killed near the end. The captured output proves it: every generator
had written its files and gen:install-tree had emitted all fifteen runtimes
before the kill. Raised to fifteen minutes.

That failure read as `null !== 0`, which says nothing. status null means killed,
not a non-zero exit, and the two want different responses: one is a timeout to
size correctly, the other is a real build break. The assertion now distinguishes
them and names the signal.

Neither test's assertions were weakened and no retry was added. A retry here
would suppress exactly the signal the timeout exists to produce.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): capture fd 1 through the mock tracker, not a raw reassignment

The phase suite reported zero test results on both lanes while running for five
and a half minutes and exiting 1. No assertion text, no stderr, four events for
the whole file: enqueue, start, dequeue, complete. That shape is not a failing
assertion — it is the runner being unable to read the child at all, because it
parses its event stream from the child's stdout.

The cause was the capture helper reassigning fs.writeSync directly. Proven
rather than assumed: a standalone probe patched fs.writeSync and called
process.stdout.write, and the interception fired only when fd 1 resolved to a
FILE, not when it was a pipe. The remote runner captures the event stream to a
file, so a helper that was invisible against a pipe swallowed the reporter's own
output on the bench. That is also why the two sibling suites wired the same way
in this change pass cleanly — they use the mock tracker, the seam io.test.cjs
established for this exact function.

The helper now uses t.mock.method with an explicit restore after each call, so
teardown belongs to node:test rather than a second hand-rolled implementation,
and the interception cannot outlive the one synchronous call it wraps even if
that call throws. Ten call sites thread the test context through; three test
callbacks gained the parameter they lacked.

The three B3 tests are untouched — same assertions, same fault injection. Only
how the context reaches the helper changed.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3057): capture phase-complete output from a subprocess, not fd 1

Two attempts to make in-process fd-1 interception safe both failed on the
bench. The suite reported zero test results on either lane while exiting 1 —
four events for the whole file — because the runner parses its event stream
from the child's stdout, and process.stdout.write routes through fs.writeSync
whenever fd 1 resolves to a file, which is how the runner captures. Patching
that seam anywhere in a file can therefore destroy the file's own reporting,
and tightening the window only moved the runtime from 326s to 125s without
recovering a single event.

So the interception is gone rather than tuned. The helper now spawns gsd-tools
as a real subprocess and reads stdout the way the OS already gives it to us,
which is what the rest of the suite does. It asserts the command succeeded
before parsing, so a genuine failure can no longer present as a JSON parse
error.

The two fault-injecting tests could not survive that move as written: a
subprocess cannot see a mock installed in the parent. Instead of reinstating
the interception they now produce the fault on disk — the summary artifact is
created as a dangling symlink, so the staleness check's real statSync throws
inside the child. That is a more honest fixture than a mock in any case, since
it is a condition a user's tree can actually be in. Skipped on Windows, matching
the existing symlink precedent in the write-guard suite.

Three further call sites turned out to depend on parent-process writeFileSync
mocks the subprocess could not see. Those call the CJS function directly, which
is what they always wanted — they never needed stdout at all.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3057): one name for one signal, one encoding for one distinction

Standards review found four things this branch introduced, all of them
inconsistencies with itself rather than with the repo.

One upstream bit reached its consumers under three names —
verification_stale_check_indeterminate in two modules, the same value with
"stale" dropped in a third, and stderr only in the fourth. Standardised on the
long name wherever it is a field. The workstream inventory keeps its stderr
channel, since its return shape has nowhere to hang a per-phase field without
rippling the builder's types, but it now says the same word for the same thing.

worktree-safety encoded one three-way distinction two ways in a single file: a
named union for a finding's kind, and boolean|null for an inventory entry's
existence. The second is now a named union too.

Two assertions matched human prose because the blocked and non-blocked
completion paths carried no typed field for the signal. Both now assert typed
values. The first round of this fix added the field but left the regex beside
it, which is the banned pattern sitting next to its own replacement; the second
removed it and added an assertion on the reason enum so nothing was lost.

The remaining two were reasoned away before being fixed, and both reasons were
bad. "No typed surface exists" is the condition CONTRIBUTING says to fix by
adding one — it took three lines. "The file already does this dozens of times"
is not licence to add instance number thirty-one; a convention that violates a
documented rule is debt, not precedent.

Vocabulary differing across DIFFERENT modules is left alone: CONTEXT.md rejects
a single shared result envelope, so per-module shapes are precedented, and a
baseline smell does not outrank a documented standard.

A census of every line this branch adds to a test file now finds no regex or
substring assertion on produced prose: 87 strictEqual, 25 ok (all non-empty or
shape guards), 12 equal, 3 throws (all typed err.code predicates), 3
deepStrictEqual, 2 notStrictEqual.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3057): backfill changeset pr number to 3088

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 16:00:52 -04:00
Tom Boucher
ef823ca9d9 fix(#2830): propagate a halted plan to its transitive dependents (#3038)
* test(#2830): add failing regression tests for halted-plan dependent blocking

Add tests/fix-2830-halted-plan-dependents.test.cjs covering direct,
transitive (2 and 3 hop), and diamond dependents of a halted plan across
both independent "which plans are incomplete" readers (phase-plan-index's
cmdPhasePlanIndex and findPhaseInternal/searchPhaseInDir), the negative
case (an unrelated decoupled plan stays runnable), and a parity check that
the two readers agree. Uses only modules that already exist at this
commit (gsd-tools.cjs via subprocess, the pre-existing phase-locator.cjs)
so the test file loads and runs cleanly on a fresh clone of this exact
commit. These fail against current behavior: neither reader has any
concept of a halted plan or a blocked_by/runnable view yet.

* fix(#2830): a halted plan no longer leaves its dependents on the runnable work list

A plan that reaches a designed stop still writes a SUMMARY, so both
"which plans are incomplete" readers saw it as an ordinary completion and
reported its dependents as ordinary runnable work — never checking
whether an upstream plan had halted rather than finished.

- New `status: halted` frontmatter value, documented in all four SUMMARY
  templates alongside the existing `status: complete`.
- New shared src/plan-dependency-graph.cts: a single computeHaltPropagation
  pass that both phase.cts's cmdPhasePlanIndex (wave-grouping) and
  phase-locator.cts's searchPhaseInDir (the phase-location primitive, ~50
  dependent symbols across 5 command routers) now call, so the
  two-implementation divergence that caused this bug cannot recur. It
  accepts an optional precomputedOrder so cmdPhasePlanIndex — which already
  runs Kahn's algorithm in computeDependencyLevels for wave assignment —
  passes that order straight through instead of a second traversal;
  searchPhaseInDir (no prior traversal) lets the module derive its own.
  The two small duplicated predicates each reader would otherwise carry
  (is this status "halted"?, which summary file matches which plan id?)
  are centralized in the same module as isHaltedStatus/buildSummaryFileIndex.
- Additive fields only: `halted`/`blocked_by`/`runnable` on
  cmdPhasePlanIndex's plans[] and top level, `halted_plans`/`blocked_by`/
  `runnable_plans` on searchPhaseInDir's result. The pre-existing
  `incomplete`/`incomplete_plans` fields are unchanged in meaning and
  membership.
- execute-phase.md's discover_and_group_plans step now also skips any
  plan whose `blocked_by` is non-empty, reporting it by name with its
  blocking chain, in addition to (not instead of) the existing
  has_summary skip rule.

Extends tests/fix-2830-halted-plan-dependents.test.cjs (introduced in the
prior commit) with direct unit coverage of computeHaltPropagation
(including the precomputedOrder call shape) and a fast-check property
test — both only possible once this commit's new module exists.

Closes #2830

* fix(#2830): surface the halt-aware view from init execute-phase

The adopted work made phase-locator compute halted_plans / blocked_by /
runnable_plans, but cmdInitExecutePhase builds its output by explicitly
enumerating fields, so all three were computed and then silently dropped at
the exact consumer the issue names as regressed.

Forwards them additively -- incomplete_plans and incomplete_count keep their
name, type and semantics byte-for-byte -- and adds the same three empty
defaults to the roadmap-only fallback so the shape is consistent in both
branches. Covered by a new test that drives the real CLI end to end rather
than the locator function, since the locator already worked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2830): fail closed on dependency cycles and stop the templates inviting the defect

Three review findings, all fixed:

- BLOCKER (isolated adversarial). Cycle participants never reach indegree 0 in
  the Kahn pass, so they were excluded from the topological order, never visited
  by the forward pass, and vanished from blocked_by entirely -- i.e. reported as
  runnable. The wave-grouping reader hard-fails on a cycle so it never hit this,
  but the phase-location reader does not, so init execute-phase offered a plan
  depending directly on a halted plan. Reproduced, then fixed in the shared
  engine so every consumer is safe regardless of pre-checks: a node absent from
  the order is now blocked with a deterministic, non-empty named cause. A plan
  silently missing from both blocked_by and runnable is the exact disappearance
  this issue exists to prevent.

- MAJOR (isolated adversarial). All four summary templates showed the field as
  an inline comment on the value line. Frontmatter parsing does not strip
  trailing comments, so an executor copying the templates' own presentation
  wrote a halt that parsed as a non-halted string, silently reproducing the
  original bug. Guidance moved off the value line, and the halt predicate now
  tolerates an unquoted trailing comment.

- HARD standards violation. A test regex-matched child-process stderr prose for
  /cycle/i, which CONTRIBUTING bans. Replaced with the structured failure signal
  plus a differential assertion (same fixture without the cycle edge must
  succeed), so it stays cycle-specific without matching prose.

Also folds the duplicated read-summary-and-check-halted wrapper out of both
readers into the shared module -- centralizing only the predicate left the exact
two-copies-that-drift pattern the module exists to prevent -- and commits the
artifact-types documentation for the new status value.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2830): stop the property generator hanging the whole suite

The remote runner did not fail -- it hung. Two containers sat in this file for
31+ minutes, and an earlier attempt ran 9 hours before I killed it. The runner
passes --test-timeout=0, so nothing ever reaps it: this would have hung CI
indefinitely, not reported a failure.

Root cause: the DAG generator built edges by rejection --

  from: fc.integer({ min: 0, max: n - 1 })
  to:   fc.integer({ min: 0, max: n - 1 })
  .filter(({ from, to }) => from < to)

With n === 1 both integers are forced to 0, so the predicate is unsatisfiable
and fast-check retries value generation forever. n is drawn from 1..12 and
fast-check biases toward boundary values, so n === 1 is reached almost at once.

This also explains why the failing-first run completed normally while the fixed
run hung: before the fix the graph module did not exist, so the property test
threw on import and never reached generation. It only starts hanging once the
code under test works.

Generates the DAG by construction instead -- `to` is drawn strictly above
`from`, with the degenerate single-node case short-circuited to an empty edge
list -- so no rejection sampling is involved. Switches the import to the shared
fast-check setup so the seed and run count are pinned per CONTRIBUTING, and adds
a bounded regression guard that samples the arbitrary directly, so a future
reintroduction fails loudly instead of hanging.

Verified: the file now completes in 2 seconds, 29 tests started and 29 finished,
zero failures, against an indefinite hang before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2830): restore the depends_on display contract and acknowledge the workflow growth

Full-suite run surfaced two things the focused harnesses could not.

1. Regression of a pinned pre-existing contract (#3785). A refactor routed the
   EMITTED depends_on field through the new dependency resolver, which also
   consults the canonical-prefix map. The original consulted the plan map only,
   so a short canonical prefix passed through verbatim -- '24-01' stayed
   '24-01' rather than becoming '24-01-auth-hardening'. The emitted field is a
   DISPLAY mapping, not the DAG resolution, and #3785 pins that. Reverted with
   a comment recording why it must not use the resolver; full resolution is
   still used for the wave DAG and halt propagation, which is what needs it.

2. The workflow file grew 518 bytes without an acknowledgment, from the
   halt-aware skip rule and the widened parse contract. Acknowledged.

Note on where the acknowledgment landed: the guidance is to add a NEW fragment,
but execute-phase.md is already named by an existing fragment and the linter
hard-fails when two ack sources name the same path. Appending to the owning
fragment, following its own established multi-PR pattern, was the only
lint-clean option.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2830): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 06:47:46 -04:00
Tom Boucher
ff4a57b78c chore(#1671): migrate the remaining 13 LARGE/XL workflows to the fragment model — Phase 6.3 (#3030)
* chore(#2994): fragmentize progress.md forensic audit onto the fragment model

Extract the --forensic-gated forensic_audit step to
workflows/progress/steps/forensic-audit.md behind a section marker, and
repair progress.md's init line to forward --forensic so the atom is
actually true in production rather than only under direct CLI tests.

progress.md shrinks 32630 -> 27207 bytes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize the four manifest-wired workflows

new-project, quick, new-milestone and progress each already had a
dedicated cmdInit* entry point but zero marked sections. Extract nine
gated bodies to workflows/<wf>/steps/ behind section markers and repair
each init line to forward its flags.

Fold --full into the discuss/research/validate facts inside cmdInitQuick
so the when= grammar never sees an OR, per the chunked-mode precedent.

Fixes found while working, per the no-defer rule:
- cmdInitProgress passed no phase info to buildSectionManifestField, so
  state:phase-mvp-mode was permanently false — an atom in the vocabulary
  whose fact could never be computed.
- the quick init router folded flag tokens into the free-text
  description, which the new forwarding would have corrupted.
- a #2508 dispatch note was nested inside quick.md's Agent(prompt=)
  fence, leaking orchestrator guidance into the subagent prompt.
- progress.md had a 3-vs-4 backtick outer-fence imbalance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize verify-work.md and admit state:ui-phase-active

Wire cmdInitVerifyWork to buildSectionManifestField — it was a dedicated
entry point that never emitted a manifest — and mark two sections.

state:ui-phase-active folds (plan:pre hooks include an active ui step) OR
(the phase dir holds a *-UI-SPEC.md) into one boolean in init.cts, so the
grammar still sees a single operator-free atom. The inner Playwright-MCP
check stays as prose inside the fragment: it is live session state and no
init seam can precompute it.

The MVP false-branch note is a real fallback, not redundant prose, so it
sits outside the marker — gating it away would delete the text needed
precisely when MVP mode is off.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): follow moved workflow content in drift guards

Retarget every guard that asserted on content this branch moved into
workflows/<wf>/steps/, mirroring 815b3d897. Each retargeted assertion was
verified to still fail when its step file is blanked, so none was
weakened into vacuity.

Three assertions in verify-mvp-uat were genuinely red. Three more were
worse than red — passing for the wrong reason:
- quick-commit-boundary and worktree-cleanup anchored on indexOf('Step
  5.6'), which matched a later cross-reference and sliced 16069 chars
  that coincidentally held the asserted substrings. Replaced with an
  expandWorkflowSections helper that splices step content back in place.
- phase6-review-capabilities lost its end boundary and widened to EOF.
- playwright-ui-verify matched 'UI' in an unrelated bullet and 'fall
  back' in a subagent-dispatch line after the real content moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize code-review and complete-milestone, admit three atoms

Add dedicated cmdInitCodeReview and cmdInitCompleteMilestone entry points
alongside the shared generic ones rather than modifying them — init.phase-op
and init.manager carry a CRITICAL blast radius (179 dependents, 24
processes) and stay byte-identical for their other callers.

Admit flag:--fix, state:fallow-enabled and state:git-create-tag, each with
a consuming section and a fact its own entry point computes.

Both sections had the resolver-in-body hazard: the fallow config-gate and
the git.create_tag check each sat inside the very block being gated, so
gating would have disabled the resolver that decides the gate. Both are
hoisted into init and the bodies now consume the resolved fact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): retarget code-review and milestone drift guards, fix two red tests

Retarget guards that asserted on content moved into steps/, proving
non-vacuity by blanking each step file and confirming failure.

Also fixes two genuinely red tests found while working, per the no-defer
rule:
- workflow-fragments' frozen-vocabulary lock was missing
  state:ui-phase-active, so commit 7ef7f8336 shipped red. Lint and build
  both passed over it, which is why neither is sufficient verification.
- code-review's quick.md capability-hook assertion carried a stale
  delimiter after the 18ff35d20 extraction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize autonomous.md and admit state:plan-strategy-converge

Five sections share one atom, the pattern plan-phase already uses for
flag:--research-phase. The atom folds --converge OR --cross-ai into a
single boolean in cmdInitAutonomous so the grammar stays operator-free.

cmdInitAutonomous is additive; init.milestone-op, init.manager and
init.phase-op are untouched and still consumed. The $PLAN_STRATEGY bash
resolver is deliberately retained — ungated local-planning bullets still
read it, so the init-side fact supplements it rather than replacing it.

converge-fail-fast required splitting one bash fence so the always-run
CONVERGENCE_ARGS construction stays outside the marker. All three
flag-absent fallbacks were left outside their markers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize review and discuss-phase-assumptions

Admit state:reviewer-instances-configured (two peripheral notes share it;
the core reviewer-lane dispatch stays unmarked — it is the workflow's
primary always-evaluated logic, not an optional branch) and
state:auto-advance-active, which folds --auto OR two config keys into one
boolean so the grammar stays operator-free.

discuss-phase-assumptions was the highest-risk edit in this PR. Its
auto_advance step is a full if/elif/else; gating it whole would have
deleted the flag-absent fallback needed exactly when --auto is off. Split
verified exact: resolvers 636-651 and the 'End here' fallback 668-669 both
stay outside the marker; only 653-667 is gated.

Adds emitted-drift acks for the two files that grew — review.md (+55 B)
and autonomous.md (+737 B from 80799211c, which had none and would have
red-gated the push.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): fragmentize docs-update, update, transition and new-milestone Part A

Completes the 13-workflow rollout. Three of these had no init call at all
and gained a dedicated entry point plus their first gsd_run query line.

Admits state:is-monorepo and adds state:next-channel, state:workstream-active
and state:flat-mode. Vocabulary 26 -> 30 atoms.

Part A of new-milestone applies when NO workstream is active — the negation
of state:workstream-active. Rather than teach the grammar negation, which is
the Greenspun drift the frozen list exists to prevent, it gets a separate
positively-phrased atom whose fact is the inverse. Part B, which always runs,
stays outside the marker.

flag:--verify-only is deliberately NOT admitted: docs-update has no
contiguous purely-additive region for it, and an atom without a consuming
section is dead vocabulary. Evidence recorded in the slice report.

update.md reuses its existing resolved $GSD_TOOLS rather than prepending the
canonical preamble, which would have clobbered it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): stop automated-ui-verification re-resolving its own gate, retire dead vocabulary

Two defects the new tests caught.

The automated-ui-verification step re-ran gsd_run loop render-hooks and
recomputed UI_PHASE_ACTIVE inside a body that is only read when that fact
is already true — the circular self-disabling pattern this design forbids,
introduced by 3c654b168. cmdInitVerifyWork now exposes ui_phase_active and
the step consumes it. Its launcher preamble goes too: no gsd_run remains.
The Playwright-MCP check stays as prose — that is live session state.

Dead vocabulary predating this PR: flag:--full and state:needs-codebase-map
were admitted with a gate-1 claim that never materialized. flag:--full is
removed, redundant once quick folds it into discuss/research/validate.
state:needs-codebase-map gets the real consumer it always lacked, gating
new-project's codebase-map offer. Vocabulary 30 -> 29, and no atom is now
without a consuming section.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): add the atom-admission, inversion and resolver-hoist gates

The two existing parity guards prove vocabulary/predicate symmetry but
never that a fact is computed — an atom no cmdInit* assembles evaluates
false forever. These close that hole:

- per-atom satisfiability for all 29 atoms, plus an anti-vacuity assertion
  so the loop cannot silently cover zero atoms
- dead-vocabulary check against the shipped manifest
- inversion guard: the flag-absent fallbacks in discuss-phase-assumptions
  and verify-work must stay outside their markers
- data-driven resolver-hoist guard over the shipped manifest, so a future
  extraction cannot reintroduce the circular class
- compound-fold coverage (--full, --cross-ai, --rc, config-only --auto)
- null-vs-[] degraded/computed distinction, and flag value shapes

Also repairs the frozen-vocabulary lock, which was stale and red for the
seven atoms earlier commits on this branch shipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2994): add changeset for the fragment-model rollout

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(#2994): cite the issue on the two new allow-test-rule exemptions

ADR-456 requires an issue ref on the same line as the annotation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(#2994): correct the atom-count claims after retiring flag:--full

The vocabulary doc comments still said 30 entries; it is 29 since
flag:--full was removed as dead vocabulary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): dedupe the phase-fallback block and harden --ws parsing

Review findings.

MAJOR: the three new init entry points each pasted a verbatim copy of the
guardedFindPhase/guardedGetRoadmapPhase fallback, taking the repo from four
copies to seven — DEFECT.GENERATIVE-FIX. Extracted applyRoadmapFallback and
folded six of the seven; each call site keeps its own field-set via a
closure. Duplication removed rather than papered over with a parity test.
cmdInitPhaseOp stays out: its fallback omits has_reviews, so it is not a
byte-identical copy, and it is CRITICAL-radius.

LOW, pre-existing: GSD_WS captured [^[:space:]]+ and expands unquoted, so a
workstream name holding glob metacharacters would expand against the
filesystem. Narrowed to [A-Za-z0-9._-]+. The unquoted expansion is kept —
it must word-split into two args and vanish when empty.

Also restores the vocabulary ordering convention, and fixes a masked test
bug the mandated run surfaced: the flag-forwarding guard checked only the
first init line per workflow, but new-milestone has two, so a real failure
was reporting exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): drop the stale new-milestone emitted-drift ack

new-milestone.md was acked for a +406 B growth measured against an
intermediate commit. Net against origin/next it SHRANK by 8 bytes, so
nothing needed the ack and it explained nothing — which the differential
attribution check reports as a stale acknowledgment, not a pass.

update.md's entry stays: it genuinely grew +703 B.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): resolve the 15 failures from the full matrix run

All 15 were real and identical on both lanes.

REAL REGRESSION: autonomous.md hit 41479 chars against the #2196 guard's
40960 cap — a CHARS cap distinct from the LARGE tier byte cap, which the
five section stubs pushed it over. Extracted the 3a.5 UI Design Contract
body to references/; now 39968 chars, and the file nets -795 B vs base, so
its growth ack is deleted rather than left stale.

REAL DEFECT: docs referenced /gsd-transition, which is not a live
registered command. Reworded.

STALE FIXTURE: the emission byte-identity test hardcoded two marked
workflows; this branch legitimately marks fifteen. Fixture corrected — the
source was right.

The rest were drift guards over the eight workflows the earlier sweep did
not cover, retargeted at where the content now lives with non-vacuity
proven by blanking each step file and confirming failure. The GSD_WS
forwarding guard was checked as a possible real break and is not one: the
charclass narrowing is intact and forwarding works end to end.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): drop the ack for a newly-added reference file

A new file's emitted ripple is attributable to the diff that adds it, so
the acknowledgment explained nothing and the differential check reports it
as stale. Removing the last entry removes the fragment — an empty one
signals nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(#2994): retarget the UI-contract guards and clear two transitive advisories

The §3a.5 extraction that brought autonomous.md under the #2196 char cap
moved its body to references/autonomous-ui-design-contract.md, so ten
guards in autonomous-ui-steps and check-ui-safety-gate were asserting it
against the host. Retargeted via a combined read, each proven non-vacuous
by blanking the reference file and confirming failure.

This class had already bitten twice on this branch because each sweep was
scoped to the workflows touched at that moment, so this one was
exhaustive: ~70 test files across all 13 workflows, zero further broken or
vacuous assertions found.

Also clears two high transitive advisories the matrix flagged on one lane
— fast-uri GHSA-7p8r-x3mc-p8w7 and three ip-address SSRF/trust-boundary
issues. Both pre-date this branch: package-lock.json was untouched until
now, so the production tree was byte-identical to the base. Lockfile-only,
package.json unchanged, verified against a real npm ci install.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(#2994): backfill changeset pr number to 3030

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 19:59:58 -04:00
Tom Boucher
ad3b9ec486 chore(#1671): fragmentize plan-phase.md and repair flag forwarding to the init bundle — Phase 6.2 (#3019)
* chore(#2993): fragmentize plan-phase.md onto the fragment model

Epic #1671 Phase 6.2. plan-phase.md is the largest workflow in the repo and
carried zero markers; it was deferred out of the Phase 3 pilot for two
reasons, both now dead. The 36-byte PRE_PHASE6 headroom was never the
blocker it looked like — fragmentizing is net-negative on host source, so
the trim is what creates the room. The --mvp interleaving was resolved by
measurement in #2992 and no sub-line mechanism is built.

- widen WHEN_VOCABULARY 14 -> 19 via a second coordinated ADR-1671
  amendment: flag:--ingest, flag:--prd, flag:--research-phase,
  flag:--reviews, state:chunked-mode
- state:chunked-mode is `--chunked` OR config workflow.plan_chunked, and
  that disjunction is resolved in the FACT, never in the grammar, so a
  compound condition never becomes an operator
- parse the new flags on the plan-phase route; extract six gated bodies to
  gsd-core/workflows/plan-phase/steps/ behind manifest-gated stubs
- prd-express-path.md was already extracted but read unconditionally; its
  wrapper is now gated, so the existing extraction finally pays off

plan-phase.md 94,483 -> 87,575 bytes (cap 94,519): headroom goes from 36
bytes to 6,944.

Also closes a surfaced docs gap: five real plan-phase flags (--chunked,
--skip-ui, --bounce, --skip-bounce, --granularity) were documented in
neither the argument-hint nor help. Making --chunked load-bearing without
fixing its siblings would leave the defect class half-open.

Refs #2993

* fix(#2993): forward flags to the init bundle so section gating actually fires

Blocker found by the correctness review, confirmed directly, and missed by
both the isolated reviewer and every test in this branch.

Neither workflow forwarded its flags to the init CLI:

  plan-phase.md:71    INIT=$(gsd_run query init.plan-phase "$PHASE" $GRAN_PARAM)
  execute-phase.md:84 INIT=$(gsd_run query init.execute-phase "${PHASE_ARG}")

So every flag: atom was permanently false in production and its section
permanently excluded. For plan-phase that made the PRD express path
UNREACHABLE — a regression, since it was an unconditional read before.
For execute-phase this is PRE-EXISTING: #2932 shipped `flag:--wave` gating
that has never once been true, so `--wave` silently dropped its own
wave-filtering guidance. Fixed here under the no-defer rule.

Why every test missed it: they drive the init CLI directly with flags,
which works. Production goes through the workflow's bash line, which did
not pass them — the exact "assert against the shape production uses" trap
this branch's own test matrix warns about.

- parse and forward --prd/--ingest/--research-phase/--reviews/--chunked
  (plan-phase) and --wave (execute-phase), using the anchored regex idiom
  the neighbouring GRAN_PARAM line already uses
- add a regression guard DERIVED FROM THE MANIFEST: for every flag:--X
  section, the owning workflow's init line must forward --X. It fails
  against the pre-fix files and covers any future atom, rather than
  spot-checking today's six.

Verified through the workflow shape, not the CLI shape: `3 --prd spec.md`
now yields ["prd-express-gate"] (was []), `2 --wave 2` yields
["partial-wave"] (was []).

Refs #2993

* test(#2993): acknowledge the execute-phase ripple and regenerate install-tree fixtures

Remote matrix was red with 46 unique failures, identical on both lanes.
Both causes are mechanical consequences of changing shipped workflow
content, and neither is visible to any local gate.

- emitted-attribution: execute-phase.md grew 163 bytes from the WAVE_PARAM
  forwarding fix and was unacknowledged, while the ack fragment named
  plan-phase.md, which SHRANK and therefore needed no ack at all — a stale
  entry is itself a failure. The reason now names the real ripple.
  The entry had to merge into the existing 2930 fragment: the ack linter
  does unconditional cross-fragment duplicate-key detection with no
  spent/live exception, so a second fragment declaring execute-phase.md
  collides even when the first is already merged and inert. Resolved per
  the linter's own guidance and that file's precedent of appending
  successive ripple reasons to one entry.
- golden-install-tree: tests/fixtures/install-tree/*.json are committed and
  deliberately excluded from the ADR-2719 attribution cutover, so they must
  be regenerated when shipped tree content changes. Regenerated after
  build:lib per the ordering landmine. 19 runtimes each gained exactly the
  six new plan-phase step files; zero paths removed, which is the absolute
  failure shape those fixtures exist to catch.

Refs #2993

* fix(#2993): restore the launcher preamble in an extracted step and follow moved content in its drift guards

Second red run: 26 unique failures, identical on both lanes, in two classes.

RUNTIME BUG (runtime-launcher-parity, 7 failures) — chunked-planning-mode.md
calls gsd_run but carried no canonical launcher preamble, which is what
DEFINES gsd_run(). On any non-Claude runtime that step would fail outright.
The preamble is now copied verbatim from the canonical source of truth,
gsd-core/workflows/_runtime-launcher.snippet.sh, and the fence dedented to
column 0 to match the prd-express-path.md sibling (a list-continuation
indent breaks the byte-equal preamble match). prd-express-path.md already
had a correct one. This is the same defect #2932 hit when it extracted
steps; the parity test caught a real bug, not a stale assertion.

DRIFT GUARDS (plan-phase-drift-guard, issue-2762-plan-reviews-chunked,
skill-frontmatter-contract) — these assert plan-phase.md contains content
this branch moved into step files. Retargeted at where the content now
lives, with the asserted property unchanged; the ALL-RUNTIMES label COUNT
test now reads host + every step file so the count is preserved across the
split rather than reduced. Each retargeted guard was verified to still fail
when its step file is stripped, so none was weakened into vacuity.

No emitted-drift ack was needed: currentSizes() enumerates
gsd-core/workflows/*.md non-recursively, so files under
plan-phase/steps/ are never in the size ratchet's scope.

Refs #2993

* chore(#2993): backfill changeset pr number to 3019

---------

Co-authored-by: sim <sim@local>
2026-08-03 09:38:58 -04:00
Tom Boucher
f1af47766a chore(#1671): widen the when= grammar and key the section manifest per workflow — Phase 6.1 (#3013)
* chore(#2992): widen the when= grammar and key the section manifest per workflow

Epic #1671 Phase 6.1. Two blockers stopped the fragment model reaching any
file beyond execute-phase.md: the when= vocabulary was frozen at 4 atoms
(3 execute-phase-specific), and the section manifest was single-workflow by
construction with 'execute-phase' hardcoded into buildSectionManifestField.

- widen WHEN_VOCABULARY 4 -> 14 via a coordinated ADR-1671 amendment; the
  grammar stays CLOSED (one atom, no operators, negation or nesting) and
  WHEN_PREDICATES stays a hand-written literal map, never deriving a
  predicate from its atom string
- InvocationFacts gains flags: ReadonlySet<string> plus three computed state
  booleans; add the missing reverse vocabulary/predicate parity guard
- key the manifest artifact per workflow; a stale flat {sections:[...]}
  artifact now fails shape validation instead of being misattributed
- wire the field into six init entry points and parse the flags each needs

An atom ships only with both a real consuming section and a fact the init
seam actually computes. Six surveyed atoms are withheld because their
workflows have no dedicated init entry point; an atom without a computed
fact evaluates false forever and silently disables its own section.

Fixes a defect found while wiring: parseNamedArgs always materializes a
boolean flag key, so folding its false into the absent sentinel is required
or every flag reads as present and gating is silently always-on.

Also resolves ADR-1671:194 by measurement: --mvp stays unmarkable, because
its interleaved sites are always-run flag resolution and a ~340 byte block
that already delegates lazily.

Refs #2992

* fix(#2992): treat any falsy option value as an absent flag and reject unsafe manifest read paths

Findings from two orthogonal reviews (Claude /code-review + an isolated
adversarial pass); both independently reproduced the first one.

- MAJOR: the flags-builder treated only `undefined` as absent, but
  parseNamedArgs yields `null` for an absent value-flag and `false` for an
  absent boolean-flag, so `--granularity` read as present on every
  plan-phase invocation. Fixed at the root: a flag is present iff its
  option value is truthy. The six per-handler `|| undefined` folds are now
  redundant and removed, which also closes the duplicate-translation and
  missed-onboard-handler findings.
- MAJOR: state:needs-codebase-map had zero coverage. Added unit, property
  and real-CLI integration tests.
- MINOR: reject absolute, UNC/drive and `..`-traversing `read` paths in the
  manifest, degrading the whole load to null like every other shape
  violation. Verified: `/etc/passwd` previously reached section_manifest.read.
- MINOR: corrected a stale "4 to 20" doc comment; the vocabulary is 14.

Refs #2992

* test(#2992): update the generator suite for the per-workflow manifest shape

The remote matrix went red with 5 unique failures, identical on
linux-node22 and linux-node24, all in tests/gen-section-manifest.test.cjs.
Re-keying the artifact to {workflows:{...}} left this suite asserting the
old flat {sections:[...]} shape; nothing else in the tree still does.

- three tests read manifest.sections.length, now undefined; retargeted at
  workflows.<name> with their original intent preserved (a fenced or
  loop-host marker still asserts NO section is produced, not merely a
  changed count)
- the stale-manifest test wrote its fixture in the OLD shape, so it tripped
  shape validation and stopped exercising staleness at all. Its fixture is
  now valid-but-mismatched so FAIL_STALE is genuinely reached again.
- added the coverage that exposed: a pre-6.1 flat artifact must report
  FAIL_MANIFEST_MALFORMED_SHAPE. That is the real upgrade path for an
  installed tree and nothing covered it.

Refs #2992

* chore(#2992): backfill changeset pr number to 3013

---------

Co-authored-by: sim <sim@local>
2026-08-02 22:36:45 -04:00
Tom Boucher
a987cf2731 chore(#2932): emit a per-invocation section manifest from the init bundle (#2987)
* chore(#2932): emit a per-invocation section manifest from init

Extends the init bundle with a typed per-invocation section manifest so an
invocation loads only the branch guidance it will actually take.

The three flag/state-gated branches in execute-phase.md move into their own
step files; the parent keeps its gsd:section markers wrapping a one-line
on-demand reference, so each section's prose lives in exactly one file and
the parent shrinks 93369 -> 89507 bytes. A new drift-guarded generator
derives the shipped section manifest from those markers, and a new pure
evaluator maps invocation facts to applicable section ids.

The evaluator is a lookup over the frozen WHEN_VOCABULARY, never a parser
(Greenspun's Tenth Rule, ADR-1671:69); a parity test asserts the vocabulary
and the predicate map stay exhaustively in sync.

Closes #2932

* fix(#2932): fail closed on prototype-chain when values

An isolated adversarial review found WHEN_PREDICATES[section.when] was a
bracket lookup on a plain-prototype object, so inherited Object.prototype
members resolved as predicates: "constructor"/"toString"/"valueOf"/
"hasOwnProperty" returned truthy and SILENTLY INCLUDED the section, and
"__proto__" threw an untyped TypeError carrying no .reason. Both violate
the module's documented fail-closed contract, and the manifest is read from
disk at run time so it cannot be assumed trustworthy.

Builds the predicate map on a null prototype and guards the lookup with an
explicit Object.hasOwn check. Adds table-driven coverage for nine
Object.prototype-shaped keys asserting the TYPED reason (asserting only
that it throws would still pass while broken) plus a fast-check property
injecting a hostile value at an arbitrary document position.

* test(#2932): retarget execute-phase step assertions at extracted step files

* fix(#2932): emit typed reasons for generator lib-load and write failures

* fix(#2932): restore launcher preamble in extracted steps and refresh derived fixtures

* chore(#2932): backfill changeset pr number to 2987

---------

Co-authored-by: sim <sim@local>
2026-08-02 12:34:41 -04:00
Tom Boucher
000a322489 fix(#2941): hint at global: prefix when bare skill name matches a global skill (#2973)
* test(#2941): add regression for bare-skill-name global: hint

When a bare skill name matches an existing global skill, the skip warning
must hint at the global: prefix. When no global skill matches, the original
'Skill not found' warning is unchanged.

* fix(#2941): hint at global: prefix when a bare skill name matches a global skill

When buildAgentSkillsBlock skips a bare skill name that doesn't exist as a
project-relative path, check if it matches an existing global skill. If so,
append a hint to the warning: 'a global skill named X exists; use global:X
to reference it'. When no global skill matches, the warning is unchanged.

getGlobalSkillDir and getGlobalSkillsBase are already imported in this module
for the global: branch. The hint is guarded on globalSkillsBase being non-null
(runtimes without a skills directory don't support the prefix).

* chore(#2941): add changeset fragment

* chore(#2941): backfill changeset PR number 2973

---------

Co-authored-by: sim <sim@local>
2026-08-01 10:58:50 -04:00
Daniel Einspanjer
f0ff23635e fix(#2602): discover project-local Codex agents (#2623)
* fix(#2602): discover project-local Codex agents

- Select an existing local Codex agents directory before global fallback
- Prove init reports the canonical local installation through compiled CJS

* test(#2602): lock Codex agent precedence

- Cover override, local authority, global fallback, and runtime compatibility
- Exercise installed state through the compiled resolver

* fix(#2602): resolve local Codex agent skills

- Pass the canonical project root to the non-Claude persona fallback
- Cover nested-Codex fallback and Claude compatibility through the CLI

* test(#2602): cover local Codex validation status

- Assert emitted validate and health commands use the project-local install
- Preserve empty local-directory authority beside complete global agents

* fix(#2602): align validation with local Codex discovery

- Pass the resolved runtime and project root to health W010
- Resolve the validate-agents runtime before checking installation status

* test(#2602): cover local Codex docs status

- Assert docs-init reports an authoritative empty local install as unhealthy

* fix(#2602): align docs with local Codex discovery

- Pass the resolved runtime and canonical project root to the shared agent checker

* fix(#2602): honor agent-skills runtime override

- Resolve agent-skills fallback runtime through the canonical project resolver
- Cover conflicting config and GSD_RUNTIME values through the emitted CLI

* fix(#2602): ignore non-directory local agents paths

- Treat only a local Codex agents directory as authoritative
- Cover regular-file fallback through the emitted install checker

* chore(#2602): add changelog fragment

- record the user-visible local Codex agent discovery fix for PR #2623

* fix(#2602): align local agent discovery with runtime policy

- Resolve Codex's local config directory through the canonical runtime policy
- Use test-managed cleanup for local-agent discovery coverage

* fix(#2602): discover local agents across runtimes

- Prefer manifest-backed project-local installs for non-Claude runtimes
- Respect runtime-specific local install roots and preserve global fallback behavior
- Cover native, partial, cross-runtime, and project-root local discovery

* fix(#2602): preserve agent discovery fallback

- Fall back globally when local-install probes fail
- Document and test symlink rejection
- Align the changeset with repository format

* fix(#2602): reuse local directory policy

- Resolve runtimes without local config through the canonical sentinel
- Document the manifest gate and refresh the context index

---------

Co-authored-by: Daniel E. <daniel.e@teachingstrategies.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:20:46 -04:00
Tom Boucher
9a76ca6783 fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter

extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.

Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.

The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.

Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (3eb1cede2), making it
the only non-text file under src. file(1) reported it as data and text tools
silently skipped it, defeating the audit rule that says to search the authored
source; tsc passed because the bytes sat inside a comment, so no gate caught it.
It is live on next.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): pin unterminated-frontmatter detection and its negative space

Covers the discriminator on both sides. The positive rows are the issue's own
repro (LF and CRLF) plus the key-count boundary 0/1/2 around the ">= 1 parsed
key" threshold. The negative rows are the documents that reach the same branch
and must stay silent -- above all a Markdown thematic break at byte 0, which is
how this class of check has previously shipped a false positive on valid
Markdown.

Deduplication is tested on both halves of the composite key: a repeat of the
same (path, cause) is suppressed, a genuine second failure in a different file
is not, and a Windows and POSIX spelling of one path resolve to a single key.
The reset seam is asserted to actually clear -- #2674 is the precedent where a
reset that silently failed to clear made every later dedup assertion a vacuous
pass, and the cases only passed because each happened to pick an unused key, so
every case here uses a path unique to itself.

Assertions are on typed surfaces throughout -- the frozen reason enum and the
dedup-set size -- never on diagnostic prose. The one CLI-level case asserts a
differential between two runs (whether stderr is empty) rather than matching a
message, and is the wired user-reachable surface for this fix. Stream failure is
injected by overriding process.stderr.write and restoring it, never chmod 0o000,
which root bypasses.

Two properties guard the ~50 call sites of the changed function: the new
optional path argument is inert with respect to the parsed value, and LF/CRLF
spellings of a document still parse identically.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): raise the truncation threshold and repair the dedup key

Isolated adversarial review found the one-key discriminator false-positives on
ordinary Markdown: a thematic break above a single labelled line -- `Note:`,
`Author:`, `TODO:`, `See:` -- parses as exactly one key and was reported as
corruption, which is the precise failure the design claimed to prevent and the
changeset promised was fixed. The threshold is now two keys. A file truncated
after exactly one key becomes a false negative; that is the same
precision-over-recall direction already taken at zero keys, and every GSD
artefact this guards carries two or more frontmatter keys.

Three dedup-key defects, each of which could silently swallow a real diagnostic:

- Backslash normalization is removed. A backslash is a legal filename character
  on Linux and macOS, so folding it to a forward slash made two genuinely
  different files share one key. Two spellings of one Windows path may now
  report twice; two distinct files can never silence each other. Lost signal is
  the worse failure.
- The key namespaces are tagged so a file literally named like the unnamed
  digest fallback can no longer collide with a path-less caller whose content
  hashes to that digest -- computable for any predictable content, no brute
  force needed.
- The source identity is computed once rather than hashed twice per emission.

Corrects the previous commit. The two NUL bytes in src/config-loader.cts were
NOT in a JSDoc comment as that message claimed; they were deliberate separators
in the live dedup key, and stripping them degraded it to bare concatenation.
They are restored as escape sequences -- byte-identical runtime string, and the
file is text again so grep can see it. The diagnostic script that misled me
indexed a character-offset string with a byte offset.

Also threads sourcePath through the STATE.md and PLAN.md readers so the two
artefacts epic #1879 is actually about name their file rather than reporting
under a content digest.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): correct fixtures and assertions left behind by the review fixes

The previous commit changed two behaviours deliberately and the suite still
encoded the old ones, so gsd-test came back red with six failures across both
lanes -- all of them mine.

Fixtures carrying a single frontmatter key no longer clear the two-key
truncation threshold, so the CLI differential and the two path-less dedup cases
were asserting a diagnostic that is now correctly withheld. They now carry two
keys, which is what a real interrupted write of a GSD artefact looks like.

The Windows/POSIX case asserted that two spellings of one path collapse to a
single key -- the exact folding that was removed because it also collapsed
genuinely distinct POSIX files whose names contain a backslash. Inverted to
assert they now report separately, with the reasoning recorded inline so the
trade is not silently reversed later: mild duplicate noise on one Windows path
is acceptable, a swallowed diagnostic is not.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): name the file at every read site, and report each file once

The diagnostic reached only the four frontmatter CLI verbs, so ~47 of 53 call
sites reported a truncated file under an anonymous content digest instead of
naming it. Since naming the file is the whole point -- it is what an operator
can act on -- that was a gap in the deliverable, not a scoping choice. 43 of 53
sites now pass the resolved path.

Closing it surfaced a defect the original design missed. A single truncated
STATE.md is parsed twice in a normal run: once by the read wrapper, which holds
the path, and again by a pure core downstream, which is handed only the string
and cannot know it. Those two parses keyed separately, so one file produced two
diagnostics -- and wiring more sites made the collision more likely, not less.
Every emission now registers both identities the input could be known by and
checks both before writing, so whichever caller arrives first speaks and the
other is suppressed. Distinct files with distinct content still report
separately, which is the property ADR-1411 actually requires; two files whose
truncated content is byte-identical collapse to one report, which stays the
documented limit.

Ten call sites deliberately keep no path. Two are frontmatter's own round-trip
checks during set and merge, where passing a path would report on every write.
The other eight are the state-transition pure cores, which ADR-1769 defines as
(content, intent, deps) -> newContent with injected I/O; threading a path
through them would contradict that recorded decision, so it is surfaced rather
than taken unilaterally. With the widened key they no longer double-report, and
in the normal flow the named parse runs first, so the file is still named.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): inject the STATE.md path into the transition cores

The six state-transition cores parsed STATE.md frontmatter without knowing
which file it came from, so a truncated STATE.md reached the operator as an
anonymous content digest on exactly the artefact epic #1879 is named for.

ADR-1769 section 3 shapes these as (content, intent, deps) -> newContent with
injected deps, and deps is the seam for precisely this: something the core
cannot derive without doing I/O. It already carries roadmapProvider and a
phase-inventory provider on that basis, each documented as injected rather than
imported so the core stays pure and testable without disk access. A resolved
path is data, not I/O, so an optional sourcePath member extends the established
pattern rather than contradicting it, and every existing stub keeps compiling
because the member is optional.

updateCore and reconcileCurrentPosition take no deps and are left alone. With
the widened dedup key they cannot double-report, and in the normal flow the read
wrapper has already named the file by the time they run.

Also regenerates gsd-core/bin/lib/state-transition.cjs. That artifact is tracked
rather than gitignored, unlike most of its siblings, so leaving it stale would
have shipped a runtime without this change to anyone reading the repo without
building. tsc had skipped the re-emit because its incremental build info still
recorded an emit that had since been reverted, so the stale output survived a
clean build; clearing tsconfig.build.tsbuildinfo forced it. The
compiled-artifact-sync gate is what surfaced the drift and now reports all nine
tracked artifacts matching their source.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): stop the widened dedup key from hiding a second file

The previous commit widened the dedup guard so one file parsed twice -- once by
a read wrapper holding the path, once by a pure core holding only the string --
reported once instead of twice. It did that by checking BOTH keys before
emitting, which silently traded one defect for a worse one: two DIFFERENT files
whose truncated content happened to be byte-identical now collided on the shared
content digest, and the second file's diagnostic was swallowed. That is the
over-coarse keying ADR-1411 explicitly forbids, reintroduced while fixing
something else.

The guard now checks only the key matching what the caller actually knows -- a
named read checks its path key, a path-less read checks its digest key -- while
still recording every key the input could later be identified by. The redundant
path-less re-parse of an already-named file stays silent, and two distinct files
always both report.

Verified across all six orderings: same file named-then-anonymous reports once;
two different files with identical content report twice; two different files
with different content report twice; the same path twice reports once; two
path-less parses of identical content report once; two path-less parses of
different content report twice.

The suite caught this -- twenty failures, all in the unusable-input tests that
reuse one truncated fixture across different paths. The local probe written
alongside the broken change did not, because it compared two files with
different content and could therefore only confirm the expected behaviour.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): count diagnostics emitted, not identities interned

The suite measured the size of the dedup set as a stand-in for "how many
diagnostics were emitted". That held only while one emission recorded exactly
one key. Once an emission began recording every identity the input could later
be matched by -- a path key and a content key for the same file -- the set grew
by two per write and twenty assertions read 2 where they expected 1.

The production behaviour was correct throughout; the proxy was not. Set size
counts identities, which is an implementation detail of the guard. The
behavioural claim these tests exist to make is how many diagnostics an operator
actually saw, so the module now exposes that directly as an emission counter and
the suite asserts on it. The set-size accessor stays for assertions genuinely
about key shape.

The local probe written alongside the change did not catch this because it
counted process.stderr.write calls -- the right thing -- while the suite counted
set growth. Verification now asserts both and requires them to agree, so a
future divergence between the counter and real writes fails immediately rather
than being discovered a bench run later.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): retire two assertions that outlived the behaviour they described

Both tests encoded assumptions the dedup fix invalidated, and both were caught
by the suite rather than by the probe written alongside the change.

The forged-path case asserted that a file named like the anonymous digest
fallback must not suppress a later path-less report. That premise is gone: an
emission now records every identity its input could be matched by, so ANY named
report of some content silences the anonymous re-parse of that same content --
which is the same-file guard working as intended, and has nothing to do with the
crafted name. The property still worth defending is that a crafted filename can
never silence a real file reported under its own path, so that is what the test
now asserts, with the deliberate suppression documented beside it.

The reset-seam case ended by reading the size of the dedup set and expecting 1.
Set size counts interned identities, not diagnostics written, and one emission
now interns two. It asserts the emission counter for the event and keeps a
weaker set-size check for the interning.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1882): close the review findings on the discriminator, dry-run and counter

Three orthogonal review passes ran against the final diff. Their findings:

A labelled preamble under a leading rule was still misreported. Raising the key
threshold to two only moved the boundary, because two colon-labelled lines are
as common in ordinary prose as one -- a document opening with a rule over an
Author and a Reviewed-by line, then prose, was called corrupt. Key count alone
cannot separate the two. What does is what follows: a write interrupted part way
through a frontmatter block ends mid-block, so every line of the region is still
frontmatter-shaped, whereas a document merely opening with a rule goes on to
prose. Both conditions are now required, and each closes a false-positive class
the other leaves open. Nested list values and indented continuations stay
frontmatter-shaped, so legitimate truncations are unaffected.

`state rebuild --dry-run` reported a truncated STATE.md anonymously. The write
path is named only because readModifyWriteStateMd parses with the path first;
the dry-run branch reads the file directly and never did. Dry-run is the
read-only mode an operator reaches for first when they suspect corruption, so it
is the one that most needed to name the file. reconcileCurrentPosition takes the
path as an optional argument now and rebuildCore passes it down. That function
was previously left alone on the grounds that a read wrapper always names the
file first -- this is the flow that disproves it.

The emission counter counted write attempts rather than writes, so on a broken
stderr it claimed a diagnostic had reached the operator when nothing had. It is
incremented only after a write that completed, and the broken-stderr test now
asserts the count as well as the return value.

Two documentation defects. The module described a guarantee it does not keep:
one file yields one diagnostic only when the named read comes first. The reverse
ordering emits twice, and that is deliberate -- a path-less caller cannot
identify its file, so suppressing the later named report would also suppress a
genuine second failure in a different file whenever two files share identical
truncated bytes, which ADR-1411 ranks the worse failure. The comment now states
the asymmetric guarantee and a test pins it. Separately, the CONTEXT.md glossary
entry still described backslash normalization that a later commit removed, and
asserted the opposite of what the tests pin; no lint checks prose against code,
so nothing caught it.

Also converts three body-level try/finally blocks to t.after(), per
CONTRIBUTING.md's rule that try/finally belongs only in helpers with no test
context -- the file's own emissionsDuring helper already did this correctly.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(#1882): tell the operator what the truncated-frontmatter warning means

A user who has just seen the new warning is acting, not studying, so this lands
in the How-To quadrant beside the other "if you see X" branches in
debug-a-failed-execution, not in reference or explanation. It gives them what
the warning means for this run, three steps to restore the file, and the fact
that the warning changes no return value or exit code.

It also states the case that matters more than the warning itself: silence does
not prove the file is intact. GSD says nothing when the partial block carries
fewer than two fields or reads as prose, because a Markdown document opening
with a horizontal rule is indistinguishable from one of those. A reader chasing
missing metadata needs to know not to treat quiet as clean. Why that threshold
exists is explanation and deliberately stays out of a how-to.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1882): backfill changeset pr number to 2712

* test(#1882): constrain each branch of the frontmatter-shape check

CI's mutation gate came in at 61.56 against a threshold of 62, and the surviving
mutants were concentrated in isFrontmatterShaped -- the function added last, in
response to review, and the only one never given tests of its own. It was
exercised solely through extractFrontmatter, which covers the composite decision
but leaves each branch of the predicate unconstrained: drop the blank-line
filter, or any one of the three shape alternatives, and every existing assertion
still passed.

Four cases now pin the halves independently. A blank line inside an interrupted
block must not disqualify it, which constrains the filter and its comparison. An
unindented list item and an indented folded-scalar continuation each exercise one
shape alternative that no other case reaches on its own -- the folded line is
neither a key nor a list item, so it is the only input that distinguishes the
indented branch. And two keys followed by prose must stay silent, which is the
negative half: it fails if the predicate is ever mutated to accept everything,
and it is the case that proves key count alone was never sufficient.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1882): register the unusable-input suite with the frontmatter mutation shard

The mutation gate reported an identical 61.56 across two runs whose only
difference was four added tests. That is the tell: the tests were never
executed. The frontmatter shard runs a fixed file list in stryker.config.mjs and
scripts/mutation-matrix.cjs, and tests/unusable-input.test.cjs was in neither, so
the entire suite covering the new unterminated-fence branch was invisible to the
gate while passing perfectly well in the normal run.

So the score was not measuring weak tests, it was measuring absent ones: #1882
added mutants to frontmatter.cjs and no test in the shard covered them. Both
lists gain the file; the config already notes they must stay in sync.

This is a registration ripple a new test file carries when it covers a
mutation-tracked module, alongside the .gitignore, eslint, inventory, glossary
and size-baseline ripples a new module carries. Nothing warned about it, which
is why two runs were spent before the identical score gave it away.

Refs #1879

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-27 16:50:12 -04:00
Tom Boucher
27c2279a39 fix(#2617): project verification next_command onto the runtime's command surface (#2700)
* fix(#2617): project verification next_command onto the runtime's command surface

`src/verification.cts` stored and synthesized hard-coded `/gsd:…` command strings
with no runtime context, and `phase complete` relayed that raw field straight into
its verification-blocked error. On a Codex project the suggested next step was
`/gsd:execute-phase`, a surface Codex does not install — it installs
`$gsd-execute-phase`.

The colon form is wrong twice over: `runtime-slash.cts` documents that "the colon
form is never emitted", so EVERY runtime — not just Codex — was being handed a
deprecated shape.

Fixed at the one routing seam rather than per caller:

- The routing table now stores BARE command names (`execute-phase`), never a
  prefixed literal. A prefixed literal in the table is what leaked.
- A single `projectNextCommand(bare, runtime, tail)` helper runs every return path
  through `formatGsdSlash`, preserving the argument tail (`01 --gaps`) untouched.
  An empty command stays empty, so "no next step" never becomes a bare prefix.
- `readVerificationStatus` accepts `opts.runtime`; `cmdVerificationStatus` and
  `phase complete` pass `resolveRuntime(cwd)`. The default is `claude`, which
  yields the canonical `/gsd-` hyphen form.

All four routed states are covered: missing, unknown, gaps_found, stale.

`init.cts` keeps its own projector deliberately. It already formats correctly, and
its command CONTENT differs from the router's on purpose (it appends the phase
number to `execute-phase`, and routes `human_needed` to `verify-work`).
Consolidating them would silently change `init`'s user-visible output, which this
issue did not ask for — so the divergence is left intact and the new tests instead
pin the property that matters on both surfaces: no raw colon form escapes.

Failing-first record: `origin/next:src/verification.cts` carried the four `/gsd:`
literals (lines 101, 108, 382, 392), and 11 existing assertions in
tests/verification-status.test.cjs asserted the colon form. Those 11 are corrected
in this commit — they passed before the fix and fail after it, which is precisely
the regression this closes.

Tests are folded into the module's primary suite rather than added as a third file
(`lint-test-file-count` caps the `verification` module at two, and consolidating is
its documented remedy — growing the allowlist is not). The `phase complete`
assertion reads `res.error`, not `res.stderr`: `runGsdTools` exposes a clean
non-zero exit's stderr as `error`, and reading the wrong field yields '' and makes
the whole check vacuous — which is how this user-visible path stayed untested.

Closes #2617

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* test(#2617): scope the new hooks to their describes; cover gaps_found through the CLI

Two findings from the orthogonal review of the first commit, both in the tests
this change added.

1. The folded block's `beforeEach`/`afterEach` were declared at MODULE scope.
   node:test applies module-scope hooks to every test in the file, so hooks added
   for the #2617 suites also wrapped the ~40 pre-existing tests in
   verification-status.test.cjs — making an unrelated block a single point of
   failure for them (currently benign, but a throwing hook would have failed
   suites it has nothing to do with). They now install inside their own describes
   via a small `useProjectionPhaseDir()` helper, with a comment recording why.

2. The live-CLI `phase complete` test exercised only the `missing` state, so a
   regression in any other routed branch would have shown up in the router's
   return object but not in the text a user actually reads. Added a `gaps_found`
   case per runtime, asserting the projected `plan-phase <N> --gaps` reaches the
   blocked-completion error.

Whole file verified green: 48 tests, 48 pass — the ~40 pre-existing ones included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* test(#2617): correct the last colon-form assertion in phase.test.cjs

The remote run surfaced one more stale assertion outside
tests/verification-status.test.cjs: the `phase complete` canonical-gate suite
matched the blocked-completion message against `/\/gsd:verify-work 0?1/`.

That project fixture configures no runtime, so it takes the `claude` default,
which now yields the canonical `/gsd-verify-work 01` hyphen form. The colon form
this asserted is exactly the deprecated shape #2617 removes — `runtime-slash.cts`
documents that "the colon form is never emitted".

Like the eleven corrected in the first commit, this assertion passed before the
fix and fails after it, which is the regression record rather than a test being
loosened: the surrounding assertions (failure reason, `stale` wording, and that
neither ROADMAP.md nor STATE.md was mutated) are untouched.

Verified against the real CLI: the emitted message is now
"Phase 1 verification is incomplete: Verification is stale. Re-run verify-work
before transition. Next: /gsd-verify-work 01".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* fix(#2617): collapse the two verification projectors into one seam

The orthogonal review found that `init.cts` carried a second, independently
maintained `verificationNextCommand()` that had drifted from the router's table
in CONTENT, not just formatting:

  state          router (before)          init.cts
  missing        execute-phase            execute-phase <N>
  unknown        execute-phase            execute-phase <N>
  human_needed   ""  (no command)         verify-work <N>

The `human_needed` row is the sharp one: two GSD surfaces disagreed about whether
a next command existed at all, and the router's own next_action told the user to
"re-run the verify step until status is passed" while naming no command to run.

init's answers were the useful ones, so the router adopts them and init now
delegates to it — satisfying the issue's "keep one verification-routing seam"
direction. `verificationNextCommand()` is deleted.

Appending the phase number surfaced a trap the old bare commands hid.
`extractPhaseToken` also returns project-code forms (`PROJ-07`), which are
indistinguishable by shape from an ordinary directory name — `gsd-651-parent`
yields `gsd-651` — so deriving the argument blindly emits
`execute-phase gsd-651`. The number is therefore appended only when it is
unambiguously numeric, or when the caller supplies it explicitly. `init` does
supply it: its `phaseDir` is unresolved in several branches, where the router
could not derive one at all.

  dir `01-example`      -> $gsd-execute-phase 01, $gsd-verify-work 01
  dir `gsd-651-parent`  -> $gsd-execute-phase,    $gsd-verify-work

Suites verified green against the built lib: verification-status 50/50,
phase 268/268, init 143/143, init-manager 40/40. `npm run lint:ci` clean.

User-visible change beyond the reported bug, as agreed: `query verification.status`
and `phase complete` now append the phase number for missing/unknown, and emit
`verify-work <N>` for human_needed where they previously emitted nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf

* chore(#2617): backfill changeset PR number (#2700)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 10:18:10 -04:00
Tom Boucher
936a345381 feat(#2505): Phase 3 — agent-skills fallback for non-dispatchable runtimes (#2521)
* feat(#2454): PR 2 — cmdAgentSkills fallback reads installed agent prompt

When no agent_skills config entry exists for a given agent type (the common
case on AGENTS-native runtimes), cmdAgentSkills previously returned empty
output. Workflows that inject ${AGENT_SKILLS_*} into subagent dispatch
prompts then carried nothing — the persona was lost.

The fallback: resolve the runtime's agents directory via checkAgentsInstalled
and read <agentsDir>/<agentType>.md. The installed agent prompt content
(now present for kimi-code via the flat-skills install layout) flows into
the dispatch prompt so the persona survives even without explicit config
opt-in. This is the reporter's suggested fix #2 from #2454.

The fallback triggers for ALL runtimes (not just kimi-code) when no config
entry exists — it is strictly additive (returns content the previous empty
path could not). If the agent file is not found on disk, the block stays
empty (same as before).

* docs(changeset): Phase 3 agent-skills fallback Added (#2510)

* docs(changeset): backfill PR #2521 for Phase 3 (#2510)
2026-07-22 01:25:41 -04:00
Tom Boucher
b6e6a22fce fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer

Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage
(seven commits, never pushed) onto current origin/next as a single squashed commit.
The original work was substantial and correct; this commit preserves its full scope,
trimmed where rebase conflicts + workflow size budgets required it.

Three independent layers where response_language was being dropped are closed:

Layer 1 — orchestrator-facing directives across workflows. Adds the strong
"All user-facing output in this workflow MUST be presented in {response_language};
technical terms, code, paths, and subagent prompts stay in English" directive
to ~40 workflows that previously either lacked it entirely (verify-work,
new-project, new-milestone, quick, manager, and ~35 more) or carried only
the weak subagent-prompt-only form (plan-phase, execute-phase). The directive
covers narration between tool calls and banner output, not just the
AskUserQuestion prompts.

Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts
an optional responseLanguage parameter and renders the frame strings
("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.")
in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/
Chinese/Korean/Italian) with an alias table covering ~30 input variants
(en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads
config.response_language via loadConfig(cwd) and passes it through, so the
byte-for-byte block verify-work.md reprints verbatim is already localized
when written — preserving the anti-injection hygiene rule at verify-work.md
(the model is forbidden to translate after the fact). CJK display width is
computed by East Asian Width property ranges (W/F) so the right ║ border of
the banner stays aligned for full-width characters. English fallback is
byte-identical to the pre-fix behavior when response_language is unset or
unrecognized.

Layer 3 — literal English report templates in execute-phase. The top-of-
workflow directive covers all template sites (templates are a structural
source, not literal output). Inline render-language notes that previously
sat at each template site were removed during the squash because they
pushed execute-phase.md over its frozen pre-phase-6 byte ceiling
(93600 — ADR-857 Phase 6 capstone). The single top directive covers the
same surface with fewer bytes.

Also extends src/docs.cts and src/init.cts to propagate response_language
into the init JSON bundle of the additional workflows so the directive can
read it.

Tests added:
- tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls
  back to English default; recognized language swaps only the two frame
  strings while structural lines stay untouched; CJK display-width regression
  (independent recomputation of East Asian Width W/F ranges).
- tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language
  wiring through docs.cts/init.cts.

References: #2402; reporter's three-layer triage + Layer-4 follow-up; the
byte-for-byte anti-injection hygiene rule at verify-work.md (the reason
Layer 2 must be renderer-side, not model-translated).

This is a squash of the in-flight bot branch — seven commits representing
the original implementation plus its subsequent fix/CJK-padding/test/
changeset/regen cycles, none of which were ever pushed or PR'd. The squash
captures the final coherent state.

* chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md

* chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451)

Rebase conflicts were entirely in generated artifacts (golden-install-parity
fixtures + workflow-size-baseline.json). After taking theirs during rebase,
regenerated cleanly against the merged source tree.
2026-07-20 14:22:27 -04:00
Tom Boucher
1a46bc068a fix(#2376): emit absolute subagent-facing paths from init/state, convert workflow literals (#2428)
* fix(#2376): emit absolute subagent-facing init/state paths

Make init.* and state.* path fields absolute rather than cwd-relative
so subagent prompts resolve correctly regardless of working directory.
Adds intel_dir/conflicts_path/requirements_path/roadmap_path/state_path
to cmdInitIngestDocs, an absolute debug_dir to cmdStateLoad, and
replaces bare .planning/... literals in 12 workflow Agent() prompt
blocks with the absolute init-JSON path fields. Includes decoy-cwd
regression tests and realpath'd tmpdir fixtures for macOS.

Squashed rebase of the #2376 commit series onto a fresh origin/next
(previous merge ee25543a1 was against a now-stale next).

* chore(#2376): add changeset

* chore(#2376): regenerate golden fixtures + workflow size baseline

Regenerated after rebasing the absolute-path fix onto current next
(picks up #2351's run-with-timeout content in execute-phase.md too).

* fix(#2376): trim execute-phase.md redundancy to stay under the size margin

* chore(#2376): regenerate golden/size baseline after rebase onto next
2026-07-19 15:43:48 -04:00
Tom Boucher
b0f672f88c fix(#2337): capture and surface todo severity (#2381)
add-todo.md gains a confirm-based infer_severity step (infer from the blocker/major/minor/cosmetic taxonomy, confirm via AskUserQuestion with TEXT_MODE fallback, before writing) and a severity frontmatter field. cmdListTodos and cmdInitTodos now surface severity, backward-compatible (key omitted when absent), in parity.

Closes #2337. Admin-merged (self-review bypass) with full green CI.
2026-07-17 14:01:38 -04:00
Tom Boucher
5f609762ed fix(#2237): fail loud on ambiguous bare-number phase directory collision (#2262)
* fix(#2237): fail loud on ambiguous bare-number phase directory collision

When two unrelated projects share a .planning/phases/ tree, a bare phase
number silently resolved to the first 0N-* directory found — risking
cross-project file writes. The fix detects multiple matches for the same
phase number and surfaces an ambiguous_matches result instead of silently
taking the first.

Changes:
- src/phase-locator.cts: searchPhaseInDir uses filter() + ambiguity check
- src/phase.cts: cmdFindPhase same pattern
- src/init.cts: cmdInitPhaseOp surfaces ambiguous_matches in the result
- tests/phase-locator.test.cjs: 3 regression tests

* docs: backfill changeset PR number (#2262)

* merge: keep up to date with next

* fix: regenerate stale capability-registry after next merge
2026-07-14 15:17:14 -04:00
Tom Boucher
e0f969af6a refactor(#2246): centralize cross-platform path-separator handling (toPosixPath / toNativePath / posixNormalize) (#2247)
Replace every open-coded separator translation across the installer/hooks
source with named, tested seams in shell-command-projection.cts (the platform
seam), removing all hardcoded `/`+`\` from path handling:

- toPosixPath(p)   — this machine's native path → POSIX (running-OS relative;
                     for local filesystem paths).
- toNativePath(p)  — POSIX → native (collapses the win32 `/\//g,'\\'` ternary).
- posixNormalize(p)— unconditional `\`→`/`, OS-independent; for emitting paths
                     to a POSIX/bash TARGET (which may differ from the running
                     OS) and for parsing mixed-separator input.

core-utils.toPosixPath now delegates to the seam, so its 20+ existing consumers
resolve to one implementation; no duplicate helper.

- ~47 sites across runtime-hooks-surface, runtime-artifact-conversion,
  runtime-artifact-install-plan, drift, init, worktree-safety,
  installer-migrations, installer-migration-authoring, install-engine, surface,
  verify, runtime-artifact-layout, schema-detect, check-command-router.
- Closes the latent POSIX-literal-backslash corruption class (the regex form
  corrupts a POSIX path containing a literal backslash; split(path.sep) does not).
- New unit + fast-check property tests for all three helpers.

Closes #2246

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 15:36:17 -04:00
Tom Boucher
58b20c31f8 fix(#2136): migrate ALL operator-facing date sites to localToday (anti-pattern elimination)
The UTC-slice anti-pattern (deriving a calendar day a human reads by slicing a
UTC instant) remained in several operator-facing sites beyond the original
seamed last_activity set. Eliminate it everywhere a human reads the value as a
calendar day — do not leave known bad code in place:

- commands.cts cmdTodoComplete + cmdScaffold: completion/scaffold dates.
- gsd2-import.cts: migrated STATE.md 'Last activity' / 'Last session'.
- template.cts: plan frontmatter 'completed:' date.
- verify.cts: health --repair session-log date + '(Backfilled: <date>)' header.
- workstream.cts: workstream-create 'Last Activity' / 'created' + archive dirname.
- init.cts: JSON-bundle 'date' (-> localToday) + 'timestamp' (-> nowIso) at all
  three sites; drop the now-dead 'const now = new Date()' in cmdInitTodos /
  cmdInitMapCodebase (cmdInitQuick keeps it for the local branch-id derivation).
- state.cts: prune-archive '## Pruned <date>' header.
- state-transition.cts: the 7 seamed last_activity writes (prior commit).

No realClock.today() / .clock.today() / raw new Date().toISOString().split('T')
operator-facing sites remain in src/. Rebuilds the tracked
bin/lib/state-transition.cjs artifact to match.
2026-07-12 15:34:59 -04:00
Tom Boucher
f910352ea6 fix(2104): guard foreign-prefix collapse in init execute-phase/verify-work/phase-op
#2056 fixed the foreign-prefix collapse for init plan-phase only.
The identical defect remained in three sibling commands that called
findPhaseInternal/getRoadmapPhaseInternal without the guard:
- cmdInitExecutePhase
- cmdInitVerifyWork
- cmdInitPhaseOp

Extracted the #2056 guard into shared helpers (guardedFindPhase /
guardedGetRoadmapPhase) and routed all four init commands through them.
Deleted the local parsePhasePrefix/isForeignPrefixedPhaseQuery copies in
init.cts — the canonical export from phase-id.cts is now used directly,
eliminating the drift risk flagged by both reviewers.

Added 5 regression tests (3 reject + 2 accept-branch) mirroring the
#2056 plan-phase tests.
2026-07-10 14:24:14 -04:00
Tom Boucher
7866e22457 Merge remote-tracking branch 'origin/next' into fix/2056-plan-phase-foreign-prefix
# Conflicts:
#	src/init.cts
2026-07-10 12:01:32 -04:00
Tom Boucher
c1cd43a39f fix(#2128): bound the phase-tag clause to {0,200} — kill quadratic ReDoS
The canonical OPTIONAL_PHASE_TAG_SOURCE tag clause `(?:\s*\([^)\n]*\))?` (and its
inlined literal mirrors across 11 modules) had an UNBOUNDED body, making the
optional-group + /g header scan quadratic on adversarial ROADMAP.md/STATE.md — a
long run of `(` after a header ran ~18.8s at 1.7MB. Bound the body to {0,200} in
the constant AND every mirror in lockstep (the #1729 "both forms change together"
contract), so the scan is linear: the same 1.7MB input now resolves in ~9ms
(measured), while real tags (a handful of chars) still match and a 201-char tag
is rejected. Added a #2128 boundary regression to the #1729 suite.

Pre-existing (byte-identical before/after the Phase 4 migrations); folded in at
maintainer direction rather than deferred.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 09:31:17 -04:00
Tom Boucher
e2eaa5b046 fix(#2128): address review — migrate 9 mis-allowlisted sites, harden scanner + guards
Correctness review of the Phase 4 guard found the allowlist over-broad and the
scanner/guards evadable. Fixed all findings:

- Migrate 9 sites that were wrongly sanctioned: their regex is the PURE canonical
  token (`\d+[A-Z]?(?:\.\d+)*`, no variant), byte-identical to already-migrated
  siblings. The old justification argued against swapping to the extractPhaseToken()
  FUNCTION (behavior-risky) — but the guard only wants the same regex built from
  the SOURCE string (byte-equal, zero risk). Coverage is now 32 migrated / 5
  sanctioned, not the overstated 23 / 14 (audit.cts x3, uat.cts, init.cts x4,
  roadmap-upgrade.cts). Each conversion proven byte-equal (.source + .flags).
- Harden the drift detector: also catch the `[0-9]`-in-place-of-`\d` variant;
  document the accepted limits (cross-line split, semantic restructuring —
  covered by the identity guard + review, not a text scan).
- Sanction robustness: a `phase-id-owner:` marker now counts only inside a `//`
  comment (a bare substring in a string no longer suppresses a real flag), and
  the preceding-line window skips blank lines (an auto-formatter's blank line no
  longer reactivates the flag).
- roadmap-parser.cts:462 comment: corrected — that regex carries no /i flag, so
  its [A-Za-z] class does real case work (matches state.cts:1409's rationale).
- Identity guard: surface require failures instead of silently skipping, and
  floor coverage at >75% of consumer modules (inspects 156/157).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 09:14:15 -04:00
Tom Boucher
dfad3a7510 refactor(#2128): single-source 23 phase-token re-derivations; sanction 14 context-specific sites
Route 23 literal re-derivations of the canonical phase-number token through
phase-id.cjs `PHASE_NUMBER_TOKEN_SOURCE` (via new RegExp). Each conversion was
proven BYTE-IDENTICAL (old.source === new.source && old.flags === new.flags), so
the runtime regexes are unchanged — zero behavior change by construction.

The remaining 14 phase-token sites are genuine but context-specific and stay
literal with a `// phase-id-owner: <reason>` sanction: dir-name parses whose
dash-continuation semantics differ from extractPhaseToken, and the [A-Za-z]
case-variant / [.-] dot-or-dash separator forms that are not source-byte-equal
to the canonical token.

Scanner (`npm run check:phase-id-drift`) is now green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 08:49:36 -04:00
Tom Boucher
9d9139b8b1 Merge branch 'next' into fix/2056-plan-phase-foreign-prefix 2026-07-09 19:08:08 -04:00
Tom Boucher
6e773d97df feat(#2088): migrate Codex onto the Embeddable Orchestration System (ADR-1239)
Drive Codex install/uninstall through the descriptor-driven Host-Integration
Interface (declarative embedding adapter → engine surface dispatch) and fold
every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven
`runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated
(tests/fixtures/golden-install-parity/codex.json); no other runtime changes.

Three Context7-verified upgrades, each with a test on the user-reachable surface:
- Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override,
  with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install
  and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest,
  and the skill-manifest inventory to honor the override so --skills-root /
  sync-skills / the manifest report the real location.
- Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact,
  PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall;
  extendedHookEvents reconciled [] -> the schema-valid wired subset.
- Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the
  negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a
  known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and
  unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own
  AgentsToml scalars (max_threads etc.) instead of dropping them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 21:42:11 -04:00
Tom Boucher
250c901b5d fix(#2056): guard init plan-phase against foreign-prefix numeric collapse
normalizePhaseName() strips any [A-Z][A-Z0-9_]*- prefix as a project code,
so a foreign-prefixed query like MEM-01 collapsed to 01 and resolved to
the unrelated numeric Phase 01 via the dir/roadmap fallback. Add a guard
in cmdInitPlanPhase: when the query carries a prefix that is not the
configured project_code, require exact prefixed evidence (a phase dir whose
token literally IS the prefixed query, or a roadmap entry literally headed
with it) before accepting a match; otherwise report phase_found:false.
The project_code's own prefixed phases pass through unchanged.
2026-07-08 14:04:52 -04:00
Tom Boucher
8ffb1261d1 Merge branch 'next' into fix/2028-phase-complete-milestone-end-and-workstream-guard 2026-07-07 23:26:39 -04:00
Tom Boucher
4483300253 fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.

Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
  → ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
  → REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
  → FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
  reviewer's own override was ignored); init.quick now resolves `reviewer_model`
  (gsd-code-reviewer) and the spawn threads it.

resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.

Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.

Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.

Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
  models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
  a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
  new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 21:22:48 -04:00
Tom Boucher
51dfa683d4 fix(#2028): phase.complete milestone-end out-of-order + workstream root-fallback guard
Two code-confirmed defects in `gsd-tools phase complete` (re-verified against
next; the three severe corruption paths the issue filed are superseded by the
ADR-1769 Transition Module migration + #2012, so this is the confirmed remainder).

1. Milestone-end mislabel (isLastPhase). The milestone-end determination only
   cleared isLastPhase when a HIGHER-numbered phase existed, so completing the
   numerically-highest phase out of order (e.g. Phase 10 before Phase 9) stamped
   STATE.md `Status: Milestone complete` while a lower phase was still outstanding.
   Added a lower-phase check: after the existing higher-phase scans, if any earlier
   phase in the current milestone has an unchecked roadmap checkbox (`[ ]`),
   isLastPhase becomes false AND next_phase/next_phase_name point at the LOWEST
   outstanding lower phase — so STATE.md advances to the real gap instead of
   parking on the just-completed phase. A completed phase always has `[x]`
   (phase.complete sets it), so all-lower-complete still reports milestone-end;
   heading-only roadmaps (no checkboxes) retain prior behavior. The checkbox regex
   mirrors the sibling phasePattern's anchoring (whitespace/bold + required `:`) so
   unrelated checklist lines mentioning "Phase N" don't match.

2. Workstream root-fallback (no guard). cmdPhaseComplete resolves every path via
   planningDir(cwd); with a `workstreams/` dir present but no active workstream and
   no --ws, that returns root `.planning`, so phase.complete wrote STATE.md/
   ROADMAP.md (and the mislabel) into the shared root other workstreams read.
   Added the same #1912 fail-safe guard init.progress got: refuse (asking for
   `--ws`/active workstream) instead of silently writing root. Resolution itself
   was already wired globally (resolveActiveWorkstream: --ws > GSD_WORKSTREAM >
   pointer, set in bin/gsd-tools.cjs), so only the refusal guard was missing.

The workstream-mode detection (`listAvailableWorkstreams`) is extracted into
planning-workspace.cts as the single source of truth and consumed by BOTH
init.progress and phase.complete, so the two fail-safe paths cannot drift.

Tests (tests/phase.test.cjs, new #2028 describe): out-of-order completion becomes
`Ready to plan` with is_last_phase=false, next_phase pointing at the outstanding
phase and Current Phase advancing to it (not the completed phase); all-lower-
complete still reports milestone-end; workstream-mode-no-active refuses with an
`--ws` hint; `--ws` completes in the workstream leaving root untouched; flat mode
unaffected. Fail-first verified locally via direct gsd-tools invocation.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 14:40:44 -04:00
Jeremy McSpadden
64d9f84cf2 no-mistakes(review): Forward onboard text flag 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
a5298c1fc0 refactor: project onboard routing in init 2026-07-05 19:16:08 +00:00
Jeremy McSpadden
b3555b103d no-mistakes(review): fix onboard fast root anchoring 2026-07-05 19:16:07 +00:00
Cursor Agent
7b3bf9be3f fix: align new-project map gate and guard onboarding summary overwrite
init new-project now uses the same seven-file codebase map completeness
check as init onboard, so partial .planning/codebase/ directories no longer
skip the brownfield mapping offer after onboarding warns about an incomplete
map.

The onboard workflow now branches on onboarding_summary_exists and asks for
confirmation before regenerating SUMMARY.md on repeat runs.
2026-07-05 19:16:06 +00:00
Jeremy McSpadden
66ff521c71 no-mistakes(review): Fix onboard runtime and doc detection 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
f29f981486 no-mistakes(review): Fix onboarding docs gate detection 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
6b0b5f1b2c no-mistakes(review): Fix onboard planning detection 2026-07-05 19:16:06 +00:00
Jeremy McSpadden
896c2740d3 feat(#1990): add onboard command for brownfield setup 2026-07-05 19:16:06 +00:00
Behruz Nassre Esfahani
12b35eeeaf fix(#1729): resolve phase headers with a pre-colon parenthetical tag (#1765)
* fix(#1729): resolve phase headers with a pre-colon parenthetical tag

A phase header may carry a parenthetical tag between the number and the
colon, e.g. `### Phase 26 (Cluster B): Title`. Every phase-header regex
built `Phase\s+<num>` immediately against the colon delimiter, so the
tagged phase was invisible: the resolver returned found:false and, just
as bad, the capture-all enumeration/parse paths (roadmap analyze,
milestone listing + milestone-scope filter, verify, init/import, state
total_phases, validate, the command router, preamble stripping, and the
phase-remove renumbering rewrite) silently dropped, miscounted, or
failed to renumber it — wrong phase_count, progress_percent, next_phase,
or corrupt numbering after a removal.

The fix tolerates the tag at the header seam. Parameterized resolver
sites compose the exported OPTIONAL_PHASE_TAG_SOURCE fragment; literal
enumeration sites inline its character-for-character mirror
`(?:\s*\([^)\n]*\))?`, placed immediately before the colon so it cannot
alter an existing match (optional, single-line, one paren pair, no
capture-group shift). In the renumber-on-removal rewrite the tag is
folded into the re-emitted suffix capture so it survives verbatim. Both
forms are documented to change together and a drift-guard test asserts
their behavioral equivalence over a header corpus.

Deliberately excluded: roadmap-upgrade.cts (legacy one-time migration),
where tolerating the tag would silently drop it on header rewrite — that
needs its own data-preserving treatment. Known boundaries left for
follow-up: checklist/bullet-style phase entries (`- [ ] Phase N (tag):`)
and a malformed space-before-colon variant, both pre-existing.

Validated empirically against the issue's reproduction: `roadmap
get-phase 26` resolves and `roadmap analyze` lists Phase 26 with the tag
excluded from the name (phase_count 2, next 26); an all-tagged versioned
roadmap now scopes correctly instead of falling back to a pass-all
filter. Regression coverage in tests/phase.test.cjs asserts resolver
parity (pre- vs post-colon), padding tolerance (#3537), decimal
sub-phases, no cross-phase false match, the shared seam, enumeration
coherence, renumber-preserves-tag, and seam/mirror drift. Full unit
suite green (7291 pass, 0 fail); eslint + regression-name +
resolution-provenance + changeset lints pass. Reviewed by Codex
(no critical/high; the two enumeration misses it surfaced are folded in).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1729): add changeset for pre-colon phase-tag fix

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-02 11:55:44 -04:00
Tom Boucher
8084f626ba fix(#1912): init.progress fails safe in workstream mode with no active ws (#1918)
cmdInitProgress used planningDir(cwd), which resolves to root .planning
when no active workstream and no --ws/GSD_WORKSTREAM is set — regardless
of mode:workstream. So /gsd-progress confidently reported a stale root
milestone with no signal it was stale.

Fail safe: when .planning/workstreams/ has workstreams AND no active
workstream is resolved, error with an actionable hint naming the available
workstreams and the --ws / `workstream set` fix. Flat mode (no workstreams
dir) and --ws <name> are unchanged.

Regression in tests/init.test.cjs: errors when workstreams exist but none
active (no stale root report); succeeds with --ws; flat mode unchanged.

Closes #1912
2026-07-02 10:24:53 -04:00
Tom Boucher
be54c0dbf5 fix(#1838): exclude backlog 999.x milestone phases (#1843)
* fix(#1838): exclude backlog 999.x milestone phases

* chore(#1838): add changeset fragment
2026-06-30 21:29:35 -04:00
Tom Boucher
88a7500e91 fix(#1836): count prefixed milestone phase dirs (#1844)
* fix(#1836): count prefixed milestone phase dirs

* chore(#1836): add changeset fragment
2026-06-30 20:43:28 -04:00
Tom Boucher
9d52043f50 feat(#1733): normalize-path-in-content production AST rule + fix Windows agent-skills content leak (Phase 5) (#1736)
* feat(#1733): normalize-path-in-content production AST rule (Phase 5)

ADR-1703 Phase 5 — the first production-code rule. local/normalize-path-in-content
(src/**/*.cts, @typescript-eslint/parser): flags a path-returning fn result
(path.basename excluded — returns a separator-less filename) interpolated into an
@-reference / config-dir markdown body without .replace(/\\/g,'/') normalization,
per RULESET.CONTENT-PATH-NORMALIZATION / DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT.

Build-and-assess found the canonical defect site (computePathPrefix) already
compliant and only 1 src/ hit — a false positive (path.basename in a status
message) — eliminated by narrowing (exclude basename; require a real @-ref/
config-dir marker, not bare .md). 0 src/ violations: clean forward-prevention.

The out-of-band disable-ban now scans src/**/*.cts too (typescript-estree) so the
production rule also cannot be eslint-disabled. Registered (error) + PROTECTED_RULES;
CONTEXT.md predicates + how-to doc updated.

- RuleTester suite (26 cases)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1733): add changeset for Windows agent-skills path-leak fix

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: harden mutation-matrix.cjs stdin read against EAGAIN on non-blocking pipe

scripts/mutation-matrix.cjs read piped stdin via readFileSync(process.stdin.fd).
On macOS libuv marks the stdin pipe fd non-blocking, so a synchronous read can
throw EAGAIN before the writer fills the pipe — intermittently, under heavy CI
shard load — aborting the script (status 2) and flaking mutation-matrix-ratchet.
Replace with readStdinSync(): an fs.readSync loop that retries on EAGAIN (1ms
synchronous Atomics.wait yield), stops on 0-byte/EOF, and rethrows other errors.
Deterministic regression test injects EAGAIN via an fs.readSync monkeypatch.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* ci: re-run golden-install-parity on src/lib + installer changes (close drift guard)

golden-install-parity hashes every installed bin/lib/*.cjs per runtime, so it
must re-run whenever the built lib could change. ci-test-scope selected it for
neither src/** nor installer changes, so a source-only edit (e.g. #1691's
milestone.cts/roadmap.cts) recompiled bin/lib and silently drifted the golden
fixtures past the scoped lane. Add golden-install-parity.test.cjs to both the
'TS runtime sources' and 'installer and package layout' selection rules, with
behavioral regression tests for each.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: review-bot <review-bot@gsd>
2026-06-25 21:49:22 -04:00
Jeremy McSpadden
77c7b4fc9d fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition

* no-mistakes(review): Fix canonical verification closeout gates

* no-mistakes(review): Fix verify-work frontmatter promotion command

* no-mistakes(review): Fix stale verification gates

* no-mistakes(review): Fix canonical verification routing gates

* no-mistakes(review): Fix verification dependency and runtime routing gates

* no-mistakes(review): Block stale verification bypasses

* fix: handle large init manager outputs in verification workflows

* chore: update changeset pr number

* fix(verify-work): use fresh verification.status for stale gate

The stale check after UAT used phase_completion.verification_status from
session-start INIT while human_needed promotion already queried fresh
verification.status. Align the stale gate with the canonical query so
mid-session verification refresh is not ignored.

* fix(init): skip roadmap-checked phases when selecting next_phase

Roadmap-only phases without a disk directory were still promoted to
next_phase when their checkbox was already checked. Exclude
checkboxComplete phases so progress routing does not point at work the
roadmap already marks done.

* fix: gaps_found not overridden by stale, transition uses canonical verification

- verification.cts: check gaps_found before stale so gap-closure routing
  is not masked by a newer summary mtime
- phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus
  already handles stale detection
- transition.md: replace raw grep on file content with verification.status query
  to avoid false-positive blocks from body text matching

* ci: retrigger tests after rebase

* fix(transition): replace gsd_run advisory check with awk frontmatter extraction

The runtime launcher is not defined until the update_roadmap_and_state step
bash block (~line 165). The early verify_completion block used gsd_run to
query verification.status, which violated the runtime-launcher-parity test:
'preamble appears AFTER the first gsd_run reference'.

Replace the gsd_run call with an awk-based frontmatter extractor that reads
only the status: field between the two --- fences. This avoids both the
preamble-ordering constraint and the original false-positive grep bug where
body text like 'previous_status: gaps_found' would match a full-text regex.

The phase.complete gate at update_roadmap_and_state is the canonical
enforcement point; this early check is advisory only.

Also update workflow-size-baseline.json for the updated transition.md size.

Fixes: runtime-launcher-parity test (B)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: re-check verification under planning lock in phase complete

Move readVerificationStatus into withPlanningLock so stale verification
cannot slip through when a SUMMARY.md is written between the gate and
the roadmap/state mutation. Return the blocked status from the lock
callback and emit the error after release to avoid leaving .lock behind.

* fix(transition): gate on canonical verification.status including stale

Replace awk frontmatter read with verification.status query so transition
blocks when summaries are newer than VERIFICATION.md, matching phase.complete
and other workflows (autonomous, progress, verify-work).

* Fix workflow verification gates for yolo transition and stale routing

Require VERIFY_STATUS passed before yolo/interactive transition advance.
Route stale verification recovery to verify-work, matching canonical projection.

* fix(transition): use verification.status query for stale-aware advisory check

The awk-based check read raw frontmatter status: passed, which misses the
stale case where summaries are newer than the VERIFICATION.md file even
though the frontmatter still says passed. The stale status is computed from
file modification times, not stored in frontmatter.

Move the preamble to the verify_completion bash block (the first block with
a gsd_run call) so gsd_run query verification.status can be used for the
advisory check. This gives the full readVerificationStatus logic including
mtime-based staleness detection, matching the enforcement gate at phase.complete.

Capture full JSON (VERIFY_JSON) so next_action can be included in the
advisory output alongside the status.

Also update workflow-size-baseline.json for the updated transition.md size.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* ci: trigger test matrix for 525b946

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(transition): restore awk frontmatter extraction for pre-shim verification check

The gsd_run launcher shim is not defined until line ~163 of transition.md,
so the verification debt check at line ~80 cannot use gsd_run. Restore the
awk-based frontmatter extraction that correctly reads status without needing
the runtime, and restore the shim at its proper location before
phase.complete.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(#1522): clarify transition verification gate wording

* fix(#1522): update transition workflow size baseline

* fix(#1522): update workflow-size-baseline after rebase onto next

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review)

Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync
and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections,
so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any
FS error threw uncaught into callers NOT under the planning lock (init.manager /
init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and
thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller)
for parity with readVerificationStatus's no-throw contract and testability. Also adds the
Verification Module glossary entry to CONTEXT.md (review B3).

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-24 13:19:10 -04:00
Behruz Nassre Esfahani
6bbd8919d5 fix(#1400): flush agent-skills block via writeAllSync instead of process.exit(0)
cmdAgentSkills' plain (non-JSON) path wrote the <agent_skills> block with
process.stdout.write then immediately called process.exit(0). stdout.write is
async on pipes/files, so on Windows the process tore down before Node flushed
the buffer — the workflows' `$(gsd_run query agent-skills <type>)` capture
received 0 bytes and every ${AGENT_SKILLS_*} substitution expanded empty,
silently dropping configured per-agent skills.

Route the plain path through the existing synchronous output() helper
(writeAllSync, src/io.cts) + return — the same flush-safe mechanism the --json
branch already uses — instead of write + process.exit(0).

Adds a #1400 regression block to tests/agent-skills.test.cjs that captures
stdout via a real file descriptor and asserts the block is non-empty and
byte-identical to the --json .block.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 16:49:18 -07:00
Tom Boucher
b0c774c2e3 feat(#1416): formalize Resolution convention + agent-skills value envelope (Resolution Provenance P3) (#1425)
Narrows P3 of ADR-1411 (Resolution Provenance, epic #1411) based on an
adversarial fit-analysis that showed a single Resolution<T> envelope adopted
by agent-skills, capability-state, and capability-writer fails the deletion
test: configured/reason are meaningless for capability verbs, and
capability-writer's errors[] (operation-not-applied) cannot fold into
warnings[]. The only genuinely shared seam is warnings: string[].

Changes:

- src/resolution.cts: new pure types+builder leaf — exports Resolution<T>
  {value, configured, reason, warnings}, makeResolution<T>() builder, and
  AgentSkillsValue {block, skills_count}. No other src/ imports.

- src/init.cts: cmdAgentSkills --json IR gains additive value:{block,
  skills_count} field (built via makeResolution). All existing flat fields
  (agent_type, block, skills_count, warnings, configured, reason, source,
  degraded) are retained unchanged for back-compat.

- src/capability-state.cts: doc comment on ResolveCapabilityRuntimeStateResult
  naming it the canonical read-verb envelope. No JSON change.

- src/capability-writer.cts: doc comment on SetCapabilityStateResult naming it
  the canonical mutation-verb result (warnings=advisory, errors=operation-
  not-applied). No JSON change.

- CONTEXT.md: new ### Resolution Convention glossary entry after
  ### Resolution Provenance.

- docs/adr/1411-resolution-provenance.md: P3 narrowing amendment appended.

- tests/resolution.test.cjs: 9 unit tests for makeResolution (new).
- tests/agent-skills.test.cjs: 2 P3 tests for value.block/value.skills_count
  and back-compat of all flat fields.

All 277 tests pass (5 suites). npm run lint clean. All lint checks pass.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:44:52 -04:00
Tom Boucher
484a5b7b86 fix(#1415): loadConfig provenance + agent-skills diagnostic (Resolution Provenance P2) (#1424)
Implements ADR-1411 P2 / #1415, closing #1366.

Part 1 — config-loader.cts:
- Adds `loadConfigResolved(cwd, options) → ConfigResolution { config, source, degraded }`
  with six tagged branch return paths:
  A1 ws+wsconfig → source:'workstream', degraded:false
  A2 no-ws+config → source:'root', degraded:false
  B  ws requested, wsconfig absent → source:'root', degraded:true (intercepts recursive call)
  C  .planning/ exists, no config → source:'builtin-defaults', degraded:false
  D  no .planning/, global defaults readable → source:'global-defaults', degraded:false
  E  no .planning/, no global → source:'builtin-defaults', degraded:false
- `loadConfigResolved` calls `findProjectRoot` at entry to anchor resolution to the
  nearest .planning/ ancestor (cwd-drift fix, heuristic 4 from P1)
- `loadConfig` becomes a one-line delegation: `return loadConfigResolved(cwd, options).config`
- Exports `loadConfigResolved` in the `export =` block

Part 2 — init.cts cmdAgentSkills:
- Imports `findProjectRoot` from `./project-root.cjs`
- Anchors to project root before loading config (fixes #1366 cwd-drift)
- Uses `loadConfigResolved` for provenance; passes projectRoot to buildAgentSkillsBlock
- Computes `configured` + `reason` (AgentSkillsReason enum): 'resolved' |
  'not_configured' | 'configured_empty' | 'configured_unresolved'
- configured_empty and configured_unresolved emit stderr WARNING; not_configured is silent
- --json IR gains: configured, reason, source, degraded (in addition to existing
  agent_type, block, skills_count, warnings)

Tests (TDD):
- tests/config-loader.test.cjs: 8 new provenance tests (RED before impl, GREEN after)
- tests/agent-skills.test.cjs: 7 new diagnostic tests (RED before impl, GREEN after)
- All 306 tests across 4 suites pass (config-loader:33, agent-skills:72, init:105, workstream:96)

Docs:
- CONTEXT.md: Config Loader Module entry updated with loadConfigResolved interface;
  Resolution Provenance entry notes P2 is now implemented
- docs/CLI-TOOLS.md: --json field reference table added for agent-skills

Closes #1366
Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:05:15 -04:00
Tom Boucher
66085d0080 fix(#1374): surface diagnostic when configured agent skills all fail to resolve (#1376)
* fix(#1374): surface diagnostic when configured agent skills all fail to resolve

buildAgentSkillsBlock returned '' (only ad-hoc per-path stderr warnings) when an agent configured via agent_skills had paths that all failed to resolve — missing SKILL.md, unsafe path, invalid global name, OR a malformed (non-string/non-array) value. query agent-skills --json reported skills_count>0 with an empty block and no machine-readable signal, so a fully-dropped configuration was indistinguishable from a resolved one.

Thread an optional diagnostics collector through buildAgentSkillsBlock: route every skip warning through a warn() helper (stderr + collector), flag truthy-but-malformed config values, emit an aggregate warning when configured paths resolve to zero skills, and surface the collected reasons in a new warnings[] field on the query agent-skills --json IR. Empty arrays and falsy values stay silent (skills_count is honestly 0). skills_count semantics unchanged. Docs updated for the new IR field.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1374): backfill changeset PR number (#1376)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 10:16:51 -04:00
Tom Boucher
ec2ecdf28b refactor(#1283): T3 — migrate 9 multi-leaf callers off the core spine (batch 2) (#1285)
Migrate 9 files' entire core surface to the leaf modules directly
(behaviour-identical — leaves are the objects core re-exports by reference):
config, docs, gap-checker, graphify-command-router (namespace core.output ->
io.output), init (17 core symbols -> 8 leaves), profile-output, uat,
verification, workstream.

All 9 now import zero core symbols and are removed from the allowlist
(18 -> 9). Stale core.* docstrings corrected. core.cts re-exports untouched
(still serve the remaining 9 files); teardown is T-final. No behaviour change.

Closes #1283

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 16:40:48 -04:00
Tom Boucher
54420bae9e refactor(#1277): T1 — decouple agent-install-check + git-base-branch from the core spine (#1280)
First leaf-migration tranche of epic #1267 (after T0 #1268). Migrate the
via-core callers of the two leaves T0 created to import from the leaf
modules directly, and stop core re-exporting their symbols:
- checkAgentsInstalled: docs.cts, verify.cts, init.cts -> agent-install-check.cjs
- gitWorktreeInfoInternal: init.cts -> git-base-branch.cjs
- getAgentsDir had no external via-core caller (internal to the leaf)

core no longer re-exports getAgentsDir / checkAgentsInstalled /
gitWorktreeInfoInternal; the now-unused agent-install-check + git-base-branch
requires are dropped from core; the shim-identity assertions for these are
deleted (behaviour tests retained). Convergence lint stays green (0 new).
No behaviour change.

Closes #1277

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:46:41 -04:00