Files
msd-core/gsd-core/workflows/progress.md
Tom Boucher 832dcbb751 fix(#3707): surface UAT rows audit-uat silently dropped, and never report a clean result for a file it could not read (#3887)
* test(#3707): failing-first coverage for the three parseUatItems false negatives

Nine tests that must be red and three controls that must already be green.

The controls are the point of the split. `result: pass` staying unsurfaced is
what stops the fix inverting the filter so eagerly that every passing test
becomes an outstanding item, and the classic single-line shape is the no-churn
control for rewriting the adjacency regex. Both were confirmed green against
the current build before being written down; a control that is red today would
be a second bug, not a control.

Each failing fixture was run through the built parser first and returns []
for its stated cause — the issue row matched then filtered, the block-scalar
and wrapped rows never matched at all, the all-unparseable file vanishing whole.
That evidence is in 50-test-matrix.md rather than asserted.

Tests target ../gsd-core/bin/lib/uat.cjs, the built live module, and drive the
real CLI through runGsdTools. #3706 lost a full RED/GREEN cycle to tests that
imported a different copy of the function under test, so the import target was
verified before anything was written.

* fix(#3707): stop parseUatItems dropping outstanding UAT rows

Three independent false negatives, all in the audit path, plus one the issue
did not mention.

The matcher no longer requires `expected:` and `result:` to be adjacent single
lines. It slices each `### N.` block to the next heading and reads the first
`result:` line within it, taking `expected:` from parseExpectedFromTestBlock —
the seam that already parsed both the block-scalar and inline forms correctly
and was sitting unused two hundred lines away. Two parsers in one module read
the same field with different grammars; now there is one.

The result filter is inverted from an inclusion list of three to an exclusion
of a minimal PASS set. This was the issue's one open design question, which the
reporter explicitly declined to answer for the maintainer; it was asked and
decided deliberately. The fail-safe direction is what parseGapsItems documents
seventy lines below for this same false-negative class (#2286): a token nobody
recognised surfaces rather than vanishing. The trade is a visible, correctable
false positive if a project invents a novel pass-word, against today's silent
and invisible drop.

`issue` also needed a category. It is template-sanctioned with its own `issues:`
counter, but categorizeItem fell through to `unknown` — surfacing it in the
wrong bucket would have been a half-fix.

Finally, a file parsing to zero items no longer vanishes with its frontmatter
`status:`. One with a non-terminal status is reported with `parse_gap: true`,
so the reader gets a cue to look; a `complete` one stays omitted as before.
That is what made the first two defects dangerous rather than merely lossy —
the audit omitted the phase instead of under-counting it.

* fix(#3707): close the review blockers, including a regression I introduced

The remote suite was RED on the previous commit and both reviews found real
defects. Everything below was verified by execution, not by reading.

I introduced a regression against origin/next. The rewritten result matcher was
END-anchored where the old one was not, so `result: pending (blocked on
staging)`, `result: [skipped] # no device` and `result: blocked - waiting` all
returned a row before this branch and returned nothing on it — me reproducing
the exact defect class this issue exists to kill, in the fix for it. The anchor
is gone and each shape has a regression test; trailing text now falls back to
`reason` when the block has none.

`parse_gap` was inferred from the wrong signal. It fired for ANY zero-item file
whose status was not `complete`, which asserted something false about a
perfectly-parsed all-pass file, swept in archived phases left at `testing`, and
is what turned the #2286 Gaps tests red — a control this change was supposed to
keep green. It now derives from headings SEEN BUT UNYIELDED, reported by a new
parseUatItemsWithStats, so an all-pass file and a Gaps-only file are not parse
gaps and a file whose blocks carry no `result:` line is.

The fix was also invisible end to end, which both reviewers caught
independently. parse_gap entries carry no items, and both audit-uat.md and
progress.md gate on `total_items === 0` — so the headline symptom, the phase
vanishing, still reproduced for a user and only the raw JSON had changed. There
is now a `parse_gap_files` counter and both workflows gate and report on it.

Also: categorizeItem compared case-sensitively while the new PASS check
lowercased, so `result: PENDING` surfaced as `unknown`; blocks are bounded at
the next heading of any level, so a trailing `## Gaps` entry no longer bleeds
its `reason` onto the preceding test; dead unreachable fallbacks removed; and
the all-pass control was strengthened, since asserting only `total_items === 0`
let it stay green through the bogus parse_gap entry.

* chore(#3707): acknowledge the workflow growth the fix required

The emitted-attribution guard went red because audit-uat.md and progress.md
grew, and it is right to ask: runtime-loaded workflow prose is the product,
so growth there is a real change to what an executing agent reads.

The growth is not incidental to this fix, it IS the fix reaching a user. Both
reviewers found independently that emitting `parse_gap` in the JSON changed
nothing observable, because both workflows gated their output on
`total_items === 0` and parse-gap entries carry no items — so a phase whose
rows could not be parsed still printed "All Clear" and still vanished from the
progress report. The widened gates and the branches that name the unparsed
files with their phase and path are what close that.

Acks exactly the two paths the guard reported, keyed on the bare filename. The
three spent acknowledgments it also listed are inert by its own description —
the base already absorbs them — so they are left alone rather than swept up
here, where they would just add unrelated churn to this diff.

* fix(#3707): close the mixed-file blocker and the second false-clean surface

The suite was GREEN and the isolated review still found a blocker, which is
the useful part: none of this was covered by a test.

A MIXED file dropped its unparseable rows silently. `parse_gap` sat behind an
`else if` on `items.length > 0`, so one parseable row was enough to discard
`headingsSeen` entirely — a file with one pending row and two unreadable blocks
reported one item and no gap. That is the exact class this issue exists to kill,
reappearing inside its own fix for the third time. The flag is now set
independently of item count and the entry carries `unparsed_blocks`, so the
count is quantified rather than merely flagged.

A `result:` inside a fenced code block was being read as real, so a PASSING
test could be reported as outstanding from a value in a code sample — another
regression against origin/next, whose adjacency regex ignored it. Field scans
now run against a fence-stripped copy while `expected:` still reads the raw
block, since a block scalar may legitimately contain fenced-looking text.

The workflow report was still unreachable whenever anything else was
outstanding: the unparsed table lived in the all-clear branch, so a project with
one pending row in phase 01 and an unreadable phase 02 rendered phase 02
nowhere. It now fires on `parse_gap_files > 0` from the `present` step.

planning-inspect was the second surface making a false-clean claim — for
exactly the files audit-uat now flags it emitted `scope: 'complete'` with an
empty unresolved list, positively asserting completeness over a file it could
not read. It consumes the stats now and reports SCOPE.TRUNCATED with a
`uat_unreadable` diagnostic, reusing the vocabulary already used two lines above
for an unreadable file rather than inventing a token.

Also: headings with no name no longer vanish whole; the trailing-text-to-reason
synthesis I had added is removed, since it was never required by the blocker and
silently changed categorization for `result: [skipped] # no device`; the emitted
`result` token is normalized to lower case so it agrees with `category` in a
published contract; and an O(n^2) indexOf is gone from the heading loop.

* fix(#3707): stop rows stealing each other's fields, on all three surfaces

The suite was green when the security review found these. Two are
blocker-severity and one of them is a direct hit on my own verification.

A `### N.` line indented two spaces inside an `expected: |` scalar is a valid
ATX heading, so it became a phantom row that STOLE the real row's result token
while the real row vanished. I had probed this shape and declared it fixed — my
probe asserted the item COUNT and the result token, both of which the phantom
satisfied, so it passed for exactly the reason it should have failed. Block
scalar bodies are now masked to blank lines (line count preserved, so offsets
still line up) before headings are tokenized, and the tests assert row IDENTITY
— number and name — not presence.

Feeding parseExpectedFromTestBlock the raw slice let one row publish another
row's `expected:` from inside a fence the stripped view had correctly excluded.
Blocks handed to it are now clipped at the first fence opener. This was not
cosmetic on the render-checkpoint path: a checkpoint banner a HUMAN reads and
answers was rendering a different row's expected text.

A balanced fence pair straddling a test block made that heading invisible, so
an outstanding row disappeared with no item, no gap and no count — the exact
false-clean this issue exists to close, and a regression against origin/next.
Suppressed `### N.` lines now count toward headingsSeen so the file is flagged.

An unterminated fence swallowed the rest of the document including `## Gaps`,
producing a whole-file false clean. Such a file is now treated as a parse gap,
following what uat-predicate already does.

Found and fixed inline while there: parseExpectedFromTestBlock's scalar opener
required a bare newline, so a CRLF `expected: |` fell through to the inline arm
and published `expected: "|"`, silently discarding the entire value. The same
fall-through hit `|-` and `|+`.

parseFirstPendingTest had the identical exposure on the render-checkpoint path
and now shares the same masking and clipping. Five legitimate fixtures — inline
expected, a real block scalar, CRLF, bracketed pending, and a first-pending
that is not the first test — are byte-identical before and after.

Also from the code review: the admit condition disagreed with the terminal
status guard, so a `status: complete` file with an unparseable block was
emitted as an empty entry that rendered nowhere but inflated total_files; a
control test was vacuous because its fixture filename did not match its phase
dir, so #3511 scoping meant the file was never opened — and that vacuity is why
the admit regression shipped green; an unterminated fence discarded the flag
that would have caught it; `### 1.2.3` parsed as test 1; and planning-inspect
did not share the terminal-status rule.

* test(#3707): assert what the render-checkpoint fix actually does

The suite went red on three of my own tests and the source was right — the
assertions were wrong, in a way worth naming.

One forbade the rendered checkpoint from containing `### 3. Fake Row`. But in
that fixture the string IS row 1's legitimate `expected:` block-scalar value; a
heading-shaped line inside a scalar is inert text and rendering it is correct.
The test was forbidding correct output. It now asserts row IDENTITY — the
checkpoint is for test 1 named Alpha and never test 3 named Fake Row — which is
the property that actually distinguishes the fix from the bug.

The other expected success where the correct outcome is a clean error: row 1 in
that fixture has no `expected:` of its own and only ever appeared to have one by
stealing row 2's from inside a fence. Depending on the bug to produce a pass is
how a test ends up pinning the defect. The fixture now gives row 1 its own
value and asserts the checkpoint carries it and never the fence-hidden text, and
the error path gets its own test asserting it fails cleanly without leaking.

All three were checked against the real rendered output before the assertion was
written, and each was reasoned through for whether it can fail: the identity
test breaks if a phantom row is parsed, the clipping test breaks if the raw
block is read again, and the error test would pass-not-fail under the old
stealing behavior.

* fix(#3707): correct the scalar masking frame and cover every YAML block opener

Two reviews independently found the same blocker, and it is the sharpest defect
on this issue: maskBlockScalarBodies computed line offsets in UTF-16 units but
spliced them into Array.from(content), a CODE POINT array. One emoji anywhere
earlier in the file shifted every later mask write, so the mask blanked the
wrong characters and spilled past line ends. Measured: at two astral characters
a result token truncated `pending` to `pendi` and recategorized to unknown; at
six the real row vanished; at twelve the FOLLOWING row's `result: blocked`
disappeared and the file reported clean. That is the false-clean class this
issue exists to close, reintroduced by the mitigation written to prevent it, and
defeating both new detectors at once. The mask is rebuilt line by line now,
which is frame-agnostic and length-preserving by construction.

The opener grammar was also incomplete. YAML block scalar headers take an
optional indentation indicator and an optional chomping indicator in either
order, so `|2`, `|2-`, `|-2`, `>2`, `>2+` are all valid — and none were matched.
An unmasked `expected: |2` body meant a `### N.` line inside the value became a
real heading: reproduced, row 1 disappeared and a fabricated row 2 named
"Phantom" took its identity.

Fixing that exposed a third instance of the same family, found by my own probe
rather than by review: the value extractor understood only the `|` openers, so
every `>` folded scalar published the LITERAL OPENER as its value — `expected`
came back as ">" or ">2+" and the whole scalar was discarded. The extractor now
shares the opener grammar and implements real folding, joining paragraph lines
with a space and turning a blank line into a newline, rather than pretending `>`
means `|`.

Also from the reviews: the shortfall counter scanned the masked copy but not a
fence-stripped one, so a `### N.`-shaped line inside a properly closed
documentation fence — the ordinary way to document the row format inside a UAT
file — counted as a suppressed row and flagged the file against nothing; and
clipping at the first fence discarded a legitimate `expected:` that appeared
after a closed fence, which is silent field loss.

Every opener now verified for both row identity and exact extracted value, in
LF and CRLF, alongside the emoji fixtures at 1/2/6/12.

* refactor(#3707): replace the scalar masking with a column-0 heading rule

The fix had grown to five helpers whose only job was undoing one
over-permissive rule: tokenizeHeadings treats a heading indented up to three
spaces as real, so a `### N.` inside an `expected: |` body was parsed as a row
and stole the real row's identity. Every blocker in the last three review
rounds came out of that machinery rather than the reported bug — worst of all
a UTF-16-versus-code-point frame mismatch that corrupted any document
containing an emoji.

A UAT test heading is at column 0. The shipped template puts all of them there,
no `*UAT*.md` in the repo has an indented one, and the only indented `### N.`
lines in the tree are the adversarial fixtures that must not parse. Requiring
column 0 makes a scalar-interior heading a non-heading by construction, so
maskBlockScalarBodies, indentWidthOf and BLOCK_SCALAR_OPENER_RE are gone along
with the mask-invariant test that existed only to guard them. The frame bug is
now structurally unreachable: no code-point array or offset splicing remains.

The premise was incomplete and the reviewer caught it rather than forcing it
through. Masking had been doing double duty — it also hid indented FENCE
delimiters from the tokenizer, so removing it let a two-space fence inside a
scalar body swallow a later column-0 row. The alternative on offer was to
rewrite that test to assert the row is merely counted, which is a behavior
regression dressed as a passing suite. Instead there is a small line-based pass
that blanks only indented fence delimiters — same "column 0 is structure" rule
extended consistently, no YAML knowledge, and line-based by construction so the
frame bug cannot come back. It was proven load-bearing by a negative control:
reverting the wiring reproduces the regression exactly.

Kept, because they fix defects column-0 does not touch: the fence clipping that
stops one row reading another's `expected:` from inside a fence, and the folded
scalar handling that stopped `expected: >` publishing the literal ">".

Also corrects a comment left pointing at a symbol this commit deletes.

* fix(#3707): blank neutralized fence bodies, and give both parse paths one grammar

Reviewing the simplification found two more, and the first is row theft again —
the sixth time this class has surfaced on this issue, and the second time
inside a mitigation written to stop it.

Neutralizing an indented fence blanked only its DELIMITER lines. If the block's
body held a column-0 `### N.`, un-hiding the delimiters made that line a real
heading, which then took the preceding row's fields: an `### 1. Alpha` document
came back as a single row 9 named Phantom, with Alpha gone. Neutralized blocks
are now blanked open-to-close, body included, which is the honest reading of
the intent — an indented fence inside a scalar is content, so nothing in it
should be able to produce structure. The raw block is still what the expected
extractor reads, so a legitimate `expected: |` carrying a fenced code sample
keeps its full text.

The two parse paths also disagreed about what a test row IS.
parseFirstPendingTest filtered on `^\d+\.\s+` while parseUatItemsWithStats used
`^\d+\.(?!\d)`, so `### 3.Foo` was a row when audited and not a row when
resumed. Both now share one predicate and one extractor. The extractor mattered
as much as the filter: the checkpoint path's name-mandatory pattern would have
skipped exactly the shapes the widened filter admits, so fixing the filter alone
would have moved the divergence down a line rather than closing it.

Also from the security pass: parseUatItems had become an export with no callers
and no direct test once both consumers moved to the stats form. It stays, since
deleting an exported symbol from a shipped module is a contract change and not
this issue's business, but it is now documented as the items-only wrapper and
has a test. And the PASS check lowercased a value that extraction had already
lowercased; normalization now happens once.

* fix(#3707): revert the whole-block fence blanking, and pin the rule instead

My previous commit over-corrected and the suite caught it. Blanking a
neutralized fence block open-to-close destroys content legitimately living
between the delimiters, and on an UNTERMINATED opener it blanks to EOF and
deletes every later row. No framing makes that correct, and it is what turned
two earlier tests red.

The "blocker" that prompted it was my own misreading. This change adopted the
rule that column 0 is structure and indentation is content. Under that rule an
indented fence delimiter is not a fence, so a column-0 `### 9.` sitting between
two indented delimiters genuinely IS a heading, and a `result:` after it
genuinely belongs to it. That document is malformed and the parser reading it
that way is consistent, not stealing. Nothing is silently lost either: the row
whose result was taken surfaces as the parse gap.

So the blanking is back to delimiters only, and rather than leaving the
question open, the behavior is now pinned by a test asserting the rows by
identity, with the rule stated at the site — so the next person does not
oscillate the way I just did.

Kept from the reverted commit: the shared row-heading grammar and extractor
across both parse paths, the parseUatItems wrapper documentation and its test,
and the single point of lowercasing.

* fix(#3707): scope the shortfall scan, and stop neutralized content becoming structure

The security review found a HIGH that is the earlier frame-mismatch bug wearing
different clothes. The shortfall scan compared a SECTION-scoped raw line count —
the `## Tests` body — against a DOCUMENT-wide token count. So a single legal
`### N.` row anywhere outside `## Tests` decremented the shortfall and switched
the fence-straddle detector off: an identical `## Tests` section went from
`headingsSeen 1, parse_gap true` to `headingsSeen 0, no gap` purely because a
`## Prior` section existed. A document whose rows render as ordinary blocked rows
in any CommonMark renderer audited as totally clean. Both sides of the
comparison now come from the same surface, by filtering tokens to the scan
span rather than re-tokenizing, so there is no second offset basis to keep in
step.

Neutralizing a fence could also promote its former CONTENT into structure: a
column-0 delimiter run inside an indented pair became an opener once the
enclosing delimiters were blanked, hiding every later heading to EOF. Column-0
delimiter-shaped lines inside a neutralized block are now blanked too — and only
those, so a column-0 heading between neutralized delimiters is still a heading
(the pinned behaviour) and the field lines of a row living between two scalars
still survive. The reviewer corrected my repro while fixing it: an even number of
inner runs re-pairs and hides nothing, so the live shape needs an odd one, and
both are now tests.

Two more from that pass. A legal scalar header carrying a trailing comment
(`expected: | # sample`) failed the end-anchored grammar, publishing the literal
header and raising a false gap. And the indented-row counter walked backwards
per row: 3.6 seconds at sixteen thousand rows, now 15ms, via one forward pass —
though the reviewer also established my example was not the quadratic shape,
which needs an uninterrupted scalar body.

Carried in from the previous round: the indented-row counter keys on any block
scalar rather than only `expected:`, so a template-sanctioned `reported: |`
holding user prose with a heading-shaped line no longer raises a false gap; and
`reason:`/`blocked_by:` read block scalars through the same shared extractor
instead of publishing the literal `"|"`, which also means a multi-line reason
can finally reach categorizeItem — a `reason:` mentioning a server now
categorizes as server_blocked, which was impossible while the value was thrown
away.

* fix(#3707): make the shortfall scan whole-document on both sides

Second HIGH in this area, and the diagnosis is the useful part: I closed the
first one by making the two sides agree, but I did it by NARROWING the token
side to the `## Tests` span while the parse side stayed whole-document. Rows
outside that section are still parsed and surfaced when visible, so when a fence
straddled one it fell through both sides of the comparison — no item, no gap,
file never entered the results at all. A `## Regression Tests` section, or a
second `## Tests` (collectSection takes the first), audited as totally clean
while origin/next surfaced those rows.

Both sides are whole-document now. Symmetry is the property that matters here;
every attempt to be clever about which scope to compare has produced one of
these, twice at HIGH severity.

That reinstates a known over-report, deliberately: a `### N.`-shaped line inside
a closed fence in a `## Notes` section — the ordinary way to document the row
format — counts as a suppressed row and raises a gap on a file with nothing
missing. Noisy, but visible and fail-safe, against two silent false-cleans on
the other side of the trade. This issue exists to eliminate false cleans, so the
trade goes that way, and the reasoning is written at the site so it does not get
optimized back.

Three existing tests encoded the retired scoping and are replaced rather than
worked around: two now assert the accepted over-report, and one asserting a
4-space row is "not counted" was already contradicted by widening the counter to
any indentation — refusing to PARSE a 4-space heading is right, refusing to
COUNT it reopened the hole the counter exists to close.

Also in this commit, from the same review round: the inner-delimiter sweep tested
a column-0-anchored pattern, so an INDENTED delimiter inside a neutralized block
was still promoted to structure and lost a row; it is indent-tolerant now.

A refinement was identified and deliberately not taken — keying the
documentation-sample exemption on the fence info string rather than on section
scope. It is content-based and symmetric, so it would not reintroduce the
asymmetry, but it belongs in its own change rather than riding this one.

* docs(#3707): correct two claims in the over-report justification

Both from review, both comment-only, and both matter because they would
mislead the next person into "fixing" something correct.

The over-report note called the triggering shape "the ordinary way to document
the row format". It is narrower than that: the scan requires literal digits, so
the conventional placeholder `### N. Name` does not trigger it at all — only a
sample written with real numbers does, and no phase UAT file in-tree has one,
only the shipped template, which selectPhaseUatFiles never scans. A maintainer
who tested the documented placeholder form would find no over-report and could
reasonably conclude the pin was stale. That is now stated, and it also makes the
trade look better than I claimed: the real-world frequency is lower.

The attribution guard is described as structural rather than positional. It is
positional in one respect: the walk stops at the nearest column-0 line, so a
block scalar nested inside a `## Gaps` bullet is transparent to it and a
heading-shaped line in that value gets counted. Same accepted over-report,
reached by a path the comment did not mention — recorded so it is not later
mistaken for a new defect.

* fix(#3707): a complete status no longer switches off the parse-gap detector

The security review named this as the last silent-clean path in the change, and
its phrase is the right one: a self-declared kill switch over the very detector
this issue built. A file whose frontmatter said `status: complete` was omitted
unconditionally, so one containing a fence-straddled `result: blocked` computed
headingsSeen = 1 — the detector fired — and then emitted no entry at all. The
audit reported nothing.

The predicate is now status-independent: a file is surfaced when blocks were
seen but yielded nothing, whatever it claims about itself. A terminal status is
an assertion by the author, and an assertion is exactly what must not be allowed
to suppress the signal that would contradict it. What does not change is the
thing the status is actually for — a complete file with nothing to parse, and a
complete file whose rows all parse and all pass, both stay silent, verified
through the real CLI.

I replaced a control test of mine, and it is worth saying why that is not a
weakening: its name was already false. "A zero-item file with a complete status
is still omitted" used a fixture with a `### 1.` block carrying no `result:`
line, so headingsSeen was 1 — it was never a zero-item file, it was the kill
switch itself, pinned. The intent it claimed is now covered by two stricter
tests, one for a file with no blocks and one for a file where every row parses
and passes, each asserting both that no entry exists and that no items are
counted, where the old test asserted only the former.

Everything else is byte-identical: 61 regression cases and the non-complete
equivalents of all four shapes produce exactly the same output as before, with
the delta confined to the two cells this change is meant to move.

* fix(#3707): close the moved kill switch, and keep the archive out of the live gate

Both reviewers independently found that closing the kill switch on one surface
left it standing on the other. cmdAuditUat dropped the terminal-status guard,
but buildUatRows in planning-inspect kept it — and its comment justified that by
claiming to mirror a guard cmdAuditUat no longer had. One byte-identical file
with `status: complete` and a fence-straddled `result: blocked` reported
parse_gap through audit-uat while planning-inspect published
`uat.scope: "complete"` with no diagnostic at all. That is the repo's own
generative-fix-divergence class, and no test pinned that arm, which is why it
survived. The clause and the false comment are gone and the arm now has tests.

Removing it exposed a MAJOR the security pass had not reached: archived phase
dirs are deliberately not milestone-filtered and archived UAT files are
`complete` by definition, so status-independence newly admitted the entire
project archive. One live pending row plus four signed-off milestones produced
parse_gap_files 4 — and since progress.md gates Verification Debt on that
counter, a mature project would have warned on every run, forever, about closed
history no user action can clear. Warning fatigue that buries the next real gap
is the feature defeating itself.

So the counter is split rather than suppressed: `parse_gap_files` counts live
phases only and remains the gate, `archived_parse_gap_files` carries the rest,
and every archived entry stays in `results` with its parse_gap and its
milestone. Nothing became silent; the live signal stayed actionable. Both
workflows report the archived bucket as closed history rather than as something
to act on.

The scope cascade is also decoupled, on the security reviewer's advice that it
is load-bearing here rather than a follow-up: `uat.scope` still reports
TRUNCATED honestly so no completeness is claimed over an unread row, while the
accepted fence-shortfall over-report no longer flips the aggregate fold that
withholds a phase's percentage. A genuinely unreadable file degrades as before.

Also corrects a frequency claim of mine: "no phase UAT file in-tree triggers
this" was true over a sample of zero, since the only UAT file in the tree is the
shipped template. The comment now says the shape is uncommon, which is what I
can actually support.

* fix(#3707): state that the live/archived split does not extend to outstanding_debt

Review MINOR: the split's rationale read as though it governed every counter, but
`summary.total_items` was never split — so a single archived `result: pending` row
re-trips the same Verification Debt warning the split exists to stop. The asymmetry
is deliberate: an archived parse gap is a row nobody can read, so the warning can
never be cleared, whereas an archived pending row is legible work someone can still
pay down by retesting. Debt that can be settled stays counted. The prose now says so
at the point of the claim, instead of leaving the next reader to file it as a miss.

Also rewrites the changeset, which described only the secondary fixes and omitted all
three defects the issue actually reports: `result: issue` dropped, any wrapped or
block-scalar `expected:` never matched at all, and the phase vanishing outright.

* fix(#3707): count every parse gap, dropping the live/archived split

The split had two regressions, both reproduced through the real CLI, and its
premise was false.

uat.cts carries #2766's rationale ~190 lines above the code I added: 'Outstanding
UAT items do not stop mattering when a milestone closes: a deferred human-UAT
scenario or a skipped live-stack test is exactly what gets archived still-open.'
So 'archived UAT files are complete by definition' was never true, and the split
rested on it.

Regression 1: archived-ness was inferred from path shape alone. A phase in the
CURRENT milestone, status in_progress, filed under .planning/milestones/v1.1-phases/
was classified archived and demoted out of the gate — live work reported as closed
history that needs no action.

Regression 2: the split was one-sided. total_items has no archived split, so an
archived outstanding row that PARSES gates Verification Debt while the identical
row that fails to parse was informational. The parse failure was what buried the
debt — the exact bug class this issue exists to fix, re-created one surface over.

parse_gap_files counts every parse_gap entry again, archived or not, so it agrees
with total_items on what archived means. The pre-existing archived_milestone field
and archived-phase scanning are untouched. Regression tests added for both cases.

* fix(#3707): correct the changeset clause left behind by the split revert

The changeset was rewritten before the split was removed, so its final clause
still claimed archived parse gaps are counted separately because signed-off
history is not work anyone can act on. There is one counter now, and that premise
is the one uat.cts refutes and the revert was made over. This text lands in
CHANGELOG.md verbatim, so it would have shipped a description of behavior the
code does not have.

* chore(#3707): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-26 09:16:22 -04:00

25 KiB

Check project progress, summarize recent work and what's ahead, then intelligently route to the next action — either executing an existing plan or creating the next one. Provides situational awareness before continuing work.

<required_reading> Read all files referenced by the invoking prompt's execution_context before starting. </required_reading>

**Load progress context (paths only):**
_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; _gsd_at() { for _p; do if [ -f "$_p" ]; then GSD_TOOLS="$_p"; return 0; fi; done; return 1; }; if _gsd_at "${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif unset -f gsd_run; _G="$(command -v gsd_run)"; then GSD_TOOLS="$_G"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif _gsd_at "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; then gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd_run is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; GSD_IDENTITY_STATUS=unverified; case "$(gsd_run runtime-identity --raw 2>/dev/null || true)" in '{"packageName":"@opengsd/gsd-core"'*'}') GSD_IDENTITY_STATUS=ok;; esac; export GSD_IDENTITY_STATUS; [ "$GSD_IDENTITY_STATUS" = ok ] || echo "WARNING: \"$GSD_TOOLS\" did not prove it is @opengsd/gsd-core - it is either a different package or an @opengsd/gsd-core older than the runtime-identity verb. See docs/how-to/diagnose-a-foreign-gsd-tools.md" >&2; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
FORENSIC_PARAM=""; if [[ "$ARGUMENTS" =~ (^|[[:space:]])--forensic([[:space:]]|$) ]]; then FORENSIC_PARAM="--forensic"; fi
INIT=$(gsd_run query init.progress $FORENSIC_PARAM)
if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi

Extract from init JSON: project_exists, roadmap_exists, state_exists, phases, current_phase, next_phase, milestone_version, completed_count, phase_count, paused_at, state_path, roadmap_path, project_path, config_path, phase_mvp_mode.

DISCUSS_MODE=$(gsd_run query config-get workflow.discuss_mode 2>/dev/null || echo "discuss")

If project_exists is false (no .planning/ directory):

No planning structure found.

Run /gsd:new-project to start a new project.

Exit.

If missing STATE.md: suggest /gsd:new-project.

If ROADMAP.md missing but PROJECT.md exists:

This means a milestone was completed and archived. Go to Route F (between milestones).

If missing both ROADMAP.md and PROJECT.md: suggest /gsd:new-project.

**Use structured extraction from `gsd_run query`:**

Instead of reading full files, use targeted tools to get only the data needed for the report:

  • ROADMAP=$(gsd_run query roadmap.analyze)
  • STATE=$(gsd_run query state-snapshot)

This minimizes orchestrator context usage.

**Get comprehensive roadmap analysis (replaces manual parsing):**
ROADMAP=$(gsd_run query roadmap.analyze)

This returns structured JSON with:

  • All phases with disk status (complete/partial/planned/empty/no_directory)
  • Goal and dependencies per phase
  • Plan and summary counts per phase
  • Aggregated stats: total plans, summaries, progress percent
  • Current and next phase identification

Use this instead of manually reading/parsing ROADMAP.md.

**Gather recent work context:**
  • Find the 2-3 most recent SUMMARY.md files
  • Use summary-extract for efficient parsing:
    gsd_run query summary-extract <path> --fields one_liner
    
  • This shows "what we've been working on"
**Parse current position from init context and roadmap analysis:**
  • Use current_phase and next_phase from $ROADMAP
  • Note paused_at if work was paused (from $STATE)
  • Count pending todos: use init todos or list-todos
  • Check for active debug sessions: (ls .planning/debug/*.md 2>/dev/null || true) | grep -v resolved | wc -l
> ⚠️ Context authority: PROJECT.md, STATE.md, and ROADMAP.md are the authoritative sources > for project name, milestone, current phase, and next-step routing. CLAUDE.md ## Project > blocks are a secondary config aid that may be significantly stale — do NOT use the > CLAUDE.md project description as a source for any progress report field.

Generate progress bar from gsd_run query progress / progress.json, then present rich status report:

# Get formatted progress bar
PROGRESS_BAR=$(gsd_run query progress.bar --raw)

Present:

# [Project Name]

**Progress:** {PROGRESS_BAR}
**Profile:** [quality/balanced/budget/inherit]
**Discuss mode:** {DISCUSS_MODE}

## Recent Work
- [Phase X, Plan Y]: [what was accomplished - 1 line from summary-extract]
- [Phase X, Plan Z]: [what was accomplished - 1 line from summary-extract]

## Current Position
Phase [N] of [total]: [phase-name]
Plan [M] of [phase-total]: [status]
CONTEXT: [✓ if has_context | - if not]

## Key Decisions Made
- [extract from $STATE.decisions[]]
- [e.g. jq -r '.decisions[].decision' from state-snapshot]

## Blockers/Concerns
- [extract from $STATE.blockers[]]
- [e.g. jq -r '.blockers[].text' from state-snapshot]

## Pending Todos
- [count] pending — /gsd:capture --list to review

## Open Windows
- [count] open in `.planning/WINDOWS.md` — /gsd:ship blocks while any remain
(Only show this section if count > 0; suppressed when ledger is empty or absent)

```bash
WINDOWS_STATUS=$(gsd_run windows status --raw 2>/dev/null || echo '')
WINDOWS_OPEN=$(printf '%s' "$WINDOWS_STATUS" | jq -r '.ledger.open_count // 0' 2>/dev/null || echo 0)
WINDOWS_WAIVED=$(printf '%s' "$WINDOWS_STATUS" | jq -r '.ledger.waived_count // 0' 2>/dev/null || echo 0)
```

Render `Open Windows` only when `$WINDOWS_OPEN` is greater than `0` (or `$WINDOWS_WAIVED` is greater than `0`, so an auditable deferral history remains visible). Phrase: `{WINDOWS_OPEN} open, {WINDOWS_WAIVED} waived — resolves with /gsd:ship gate; inspect via gsd_run windows status`. The ledger is cross-phase; the count is the project total, not the current phase's.


## Active Debug Sessions
- [count] active — /gsd:debug to continue
(Only show this section if count > 0)

## What's Next
[Next phase/plan objective from roadmap analyze]

If section_manifest is null or "mvp-display" is in its included list: read and execute gsd-core/workflows/progress/steps/mvp-display.md. Otherwise skip — do not read the file.

**Determine next action based on verified counts.**

Step 0: Resume-incomplete-phase invariant (Route 0)

Before any current-phase-scoped counting, scan ALL phases for incomplete execution. This catches the case where STATE.md's current_phase was advanced past the phase that actually has unfinished work (common after a mid-execution session death from hang, token exhaustion, or API disruption). Without this guard, the current-phase-scoped count in Step 1 would inspect the wrong phase and the routing would skip the unfinished work.

Skip if --no-resume or --force is present in $ARGUMENTS.

Scan all phases via the $ROADMAP JSON already loaded in analyze_roadmap. For each phase entry, compare plans length to summaries length using the same plans-without-summaries predicate as determine_next_action Route 4 (plans.length > summaries.length). Stop at the first (lowest-numbered) phase where the predicate is true. Record its phase number as INCOMPLETE_PHASE.

If $ROADMAP is empty or the query failed, surface a warning rather than silently proceeding:

INCOMPLETE_PHASE=""
if [ -z "$ROADMAP" ]; then
  echo "⚠ WARNING: resume-incomplete-phase scan could not run (\$ROADMAP is empty)." >&2
  echo "  The incomplete-phase invariant (#160) could not be verified." >&2
  echo "  Review project state carefully before continuing." >&2
else
  for PHASE_NUM in $(echo "$ROADMAP" | jq -r '.phases[] | (.number // .phase_number)'); do
    PHASE_DATA=$(echo "$ROADMAP" | jq --arg n "$PHASE_NUM" '.phases[] | select((.number // .phase_number) == ($n | tonumber))')
    # #3218: $PHASE_DATA is a `.phases[]` entry from `roadmap.analyze`, which
    # emits `plan_count`/`summary_count` SCALARS (src/roadmap.cts) — it has
    # never emitted `.plans`/`.summaries` ARRAYS. Reading those absent keys
    # (even with a `// []` fallback) always produced 0, permanently disabling
    # this resume-incomplete-phase check. Read the scalars the producer
    # actually emits.
    PLAN_COUNT=$(echo "$PHASE_DATA" | jq '.plan_count // 0')
    SUMMARY_COUNT=$(echo "$PHASE_DATA" | jq '.summary_count // 0')
    if [ "${PLAN_COUNT:-0}" -gt "${SUMMARY_COUNT:-0}" ]; then
      INCOMPLETE_PHASE="$PHASE_NUM"
      break
    fi
  done
fi

If INCOMPLETE_PHASE is non-empty: emit a one-line resume notice in the routing output and route to /gsd:execute-phase ${INCOMPLETE_PHASE} instead of running Step 1's current-phase routing. The progress report (already displayed by the report step above) gives the user full project status before this routing decision is shown.

---

## ▶ Next Up — Resuming incomplete Phase ${INCOMPLETE_PHASE}

`/clear` then:

`/gsd:execute-phase ${INCOMPLETE_PHASE} ${GSD_WS}`

(plans without summaries detected; use --no-resume to skip this check and route by current_phase instead; --force to skip all gates)

---

Then exit the route step. Do NOT run Steps 1 through Routes A-F.

If INCOMPLETE_PHASE is empty: continue to Step 1.

Step 1: Count plans, summaries, and issues in current phase

Get plan/summary counts for the current phase from the single owner (#3218 — LIVE counts, i.e. status: superseded plans excluded, matching "outstanding work"):

PHASE_COUNTS=$(gsd_run query find-phase "${CURRENT_PHASE}")
X=$(echo "$PHASE_COUNTS" | jq -r '.plan_count // 0')
Y=$(echo "$PHASE_COUNTS" | jq -r '.summary_count // 0')
(ls -1 .planning/phases/[current-phase-dir]/*-UAT.md 2>/dev/null || true) | wc -l

State: "This phase has {X} plans, {Y} summaries."

Step 1.5: Check for unaddressed UAT gaps

Check for UAT.md files with status "diagnosed" (has gaps needing fixes).

# Check for diagnosed UAT with gaps or partial (incomplete) testing
grep -l "status: diagnosed\|status: partial" .planning/phases/[current-phase-dir]/*-UAT.md 2>/dev/null || true

Track:

  • uat_with_gaps: UAT.md files with status "diagnosed" (gaps need fixing)
  • uat_partial: UAT.md files with status "partial" (incomplete testing)

Step 1.6: Cross-phase health check

Scan ALL phases in the current milestone for outstanding verification debt using the CLI (which respects milestone boundaries via getMilestonePhaseFilter):

DEBT=$(gsd_run query audit-uat --raw 2>/dev/null)

Parse JSON for summary.total_items, summary.total_files, and summary.parse_gap_files.

Track: outstanding_debt — summary.total_items from the audit. Track parse_gap_files — summary.parse_gap_files from the audit.

summary.parse_gap_files counts EVERY file with parse_gap: true, archived or not — the same as outstanding_debt (summary.total_items), which has no archived split either. An outstanding item does not stop mattering because its phase belongs to an already-archived milestone: a deferred human-UAT scenario or a skipped live-stack test is exactly what gets archived still-open, so an archived parse gap is exactly as much unread outstanding work as an archived result: pending row.

If outstanding_debt > 0 OR parse_gap_files > 0: Add a warning section to the progress report output (in the report step), placed between "## What's Next" and the route suggestion:

## Verification Debt ({N} files across prior phases)

| Phase | File | Issue |
|-------|------|-------|
| {phase} | {filename} | {pending_count} pending, {skipped_count} skipped, {blocked_count} blocked |
| {phase} | {filename} | human_needed — {count} items |
| {phase} | {filename} | {unresolved_count} deferred items |
| {phase} | {filename} | unparsed — test blocks with no readable `result:` line |

Review: `/gsd:audit-uat ${GSD_WS}` — full cross-phase audit
Resume testing: `/gsd:verify-work {phase} ${GSD_WS}` — retest specific phase

The unparsed row comes from results entries with parse_gap: true (summary.parse_gap_files counts exactly those, archived or not). This is a WARNING, not a blocker — routing proceeds normally. The debt is visible so the user can make an informed choice.

Step 1.7: Check verification status for the current phase

A phase whose verification is missing, unknown, gaps_found, or human_needed is NOT complete, even when every PLAN.md has a matching SUMMARY.md. The count-based status (roadmap.analyze) only sees plans/summaries, so without this check such a phase is reported complete and routing skips straight to the next phase. When the phase appears count-complete (summaries = plans AND plans > 0), consult the verification report (the same verification.status gate ship and execute-phase use, from #651):

PHASE_DIR=".planning/phases/[current-phase-dir]"
VERIFICATION=$(gsd_run query verification.status "${PHASE_DIR}" 2>/dev/null)
VERIFICATION_STATUS=$(printf '%s' "$VERIFICATION" | jq -r '.status' 2>/dev/null || echo "")
VERIFICATION_NEXT_ACTION=$(printf '%s' "$VERIFICATION" | jq -r '.next_action' 2>/dev/null || echo "")

Track: verification_status — the .status field (passed | stale | gaps_found | human_needed | missing | unknown). The query/projection handles a missing VERIFICATION.md (missing), unexpected values, and stale verification (stale, when summaries are newer than verification). Only passed routes as phase complete (Step 3); every other status routes back to close verification debt (Step 2).

Step 2: Route based on counts

Condition Meaning Action
uat_partial > 0 UAT testing incomplete Go to Route E.2
uat_with_gaps > 0 UAT gaps need fix plans Go to Route E
summaries < plans Unexecuted plans exist Go to Route A
summaries = plans AND plans > 0 AND verification_status = missing Phase executed; verification report missing Go to Route V.missing
summaries = plans AND plans > 0 AND verification_status = unknown Phase executed; verification status unknown Go to Route V.unknown
summaries = plans AND plans > 0 AND verification_status = stale Phase executed; verification is stale Go to Route V.stale
summaries = plans AND plans > 0 AND verification_status = gaps_found Phase executed; verification found gaps Go to Route V.gaps
summaries = plans AND plans > 0 AND verification_status = human_needed Phase executed; awaiting human verification Go to Route V.human
summaries = plans AND plans > 0 AND verification_status = passed Phase complete (verification passed) Go to Step 3
plans = 0 Phase not yet planned Go to Route B

Rows are evaluated top to bottom; the first matching row wins. The verification_status rows must precede the passed row so non-passed verification is not reported as complete.


Route A: Unexecuted plan exists

Find the first PLAN.md without matching SUMMARY.md. Read its <objective> section.

---

## ▶ Next Up — [${PROJECT_CODE}] ${PROJECT_TITLE}

**{phase}-{plan}: [Plan Name]** — [objective summary from PLAN.md]

`/clear` then:

`/gsd:execute-phase {phase} ${GSD_WS}`

---

Route B: Phase needs planning

Check if {phase_num}-CONTEXT.md exists in phase directory.

Check if current phase has UI indicators:

PHASE_SECTION=$(gsd_run query roadmap.get-phase "${CURRENT_PHASE}" 2>/dev/null)
PHASE_HAS_UI=$(echo "$PHASE_SECTION" | grep -qi "UI hint.*yes" && echo "true" || echo "false")

If CONTEXT.md exists:

---

## ▶ Next Up — [${PROJECT_CODE}] ${PROJECT_TITLE}

**Phase {N}: {Name}** — {Goal from ROADMAP.md}
<sub>✓ Context gathered, ready to plan</sub>

`/clear` then:

`/gsd:plan-phase {phase-number} ${GSD_WS}`

---

If CONTEXT.md does NOT exist AND phase has UI (PHASE_HAS_UI is true):

---

## ▶ Next Up — [${PROJECT_CODE}] ${PROJECT_TITLE}

**Phase {N}: {Name}** — {Goal from ROADMAP.md}

`/clear` then:

`/gsd:discuss-phase {phase}` — gather context and clarify approach

---

**Also available:**
- `/gsd:ui-phase {phase}` — generate UI design contract (recommended for frontend phases)
- `/gsd:plan-phase {phase}` — skip discussion, plan directly
- `/gsd:discuss-phase {phase}` — include assumptions check before planning

---

If CONTEXT.md does NOT exist AND phase has no UI:

---

## ▶ Next Up — [${PROJECT_CODE}] ${PROJECT_TITLE}

**Phase {N}: {Name}** — {Goal from ROADMAP.md}

`/clear` then:

`/gsd:discuss-phase {phase} ${GSD_WS}` — gather context and clarify approach

---

**Also available:**
- `/gsd:plan-phase {phase} ${GSD_WS}` — skip discussion, plan directly
- `/gsd:discuss-phase {phase} ${GSD_WS}` — include assumptions check before planning

---

Route E: UAT gaps need fix plans

UAT.md exists with gaps (diagnosed issues). User needs to plan fixes.

---

## ⚠ UAT Gaps Found

**{phase_num}-UAT.md** has {N} gaps requiring fixes.

`/clear` then:

`/gsd:plan-phase {phase} --gaps ${GSD_WS}`

---

**Also available:**
- `/gsd:execute-phase {phase} ${GSD_WS}` — execute phase plans
- `/gsd:verify-work {phase} ${GSD_WS}` — run more UAT testing

---

Route E.2: UAT testing incomplete (partial)

UAT.md exists with status: partial — testing session ended before all items resolved.

---

## Incomplete UAT Testing

**{phase_num}-UAT.md** has {N} unresolved tests (pending, blocked, or skipped).

`/clear` then:

`/gsd:verify-work {phase} ${GSD_WS}` — resume testing from where you left off

---

**Also available:**
- `/gsd:audit-uat ${GSD_WS}` — full cross-phase UAT audit
- `/gsd:execute-phase {phase} ${GSD_WS}` — execute phase plans

---

Route V.missing: verification report missing

All plans have summaries, but canonical verification has not passed. The phase is implementation-complete, not phase-complete.

---

## Verification Report Missing

**Phase {phase}** has all plans summarized, but no canonical `*-VERIFICATION.md` exists yet. ${VERIFICATION_NEXT_ACTION}

`/clear` then:

`/gsd:execute-phase {phase} ${GSD_WS}` — resumes at the verification gates

---

Route V.unknown: verification status unknown

VERIFICATION.md has an unexpected status. The phase is implementation-complete, not phase-complete.

---

## Verification Status Unexpected

**Phase {phase}** has all plans summarized, but its `*-VERIFICATION.md` reports an unexpected status. ${VERIFICATION_NEXT_ACTION}

`/clear` then:

`/gsd:execute-phase {phase} ${GSD_WS}` — regenerate verification

---

Route V.stale: verification is stale

VERIFICATION.md has status: passed, but one or more SUMMARY.md files are newer than the verification report. The phase is implementation-complete, not phase-complete.

`/gsd:verify-work {phase} ${GSD_WS}` — re-run verification against the latest summaries

Route V.gaps: verification found gaps (gaps_found)

VERIFICATION.md exists with status: gaps_found — verification identified gaps that need fix plans. The phase is NOT complete.

---

## ⚠ Verification Gaps Found

**{phase_num}-VERIFICATION.md** reports `gaps_found`. ${VERIFICATION_NEXT_ACTION}

`/clear` then:

`/gsd:plan-phase {phase} --gaps ${GSD_WS}`

---

Route V.human: human verification required (human_needed)

VERIFICATION.md exists with status: human_needed — automated checks passed but manual verification items remain. The phase is NOT complete until they are resolved.

---

## Human Verification Required

**{phase_num}-VERIFICATION.md** reports `human_needed`. ${VERIFICATION_NEXT_ACTION}

`/clear` then:

`/gsd:verify-work {phase} ${GSD_WS}` — resume human verification

---

Step 3: Check milestone status (only when phase complete)

Read ROADMAP.md and identify:

  1. Current phase number
  2. All phase numbers in the current milestone section

Count total phases and identify the highest phase number.

State: "Current phase is {X}. Milestone has {N} phases (highest: {Y})."

Route based on milestone status:

Condition Meaning Action
current phase < highest phase More phases remain Go to Route C
current phase = highest phase All phases complete Go to Route D

Route C: Phase complete, more phases remain

Read ROADMAP.md to get the next phase's name and goal.

Check if next phase has UI indicators:

NEXT_PHASE_SECTION=$(gsd_run query roadmap.get-phase "$((Z+1))" 2>/dev/null)
NEXT_HAS_UI=$(echo "$NEXT_PHASE_SECTION" | grep -qi "UI hint.*yes" && echo "true" || echo "false")

If next phase has UI (NEXT_HAS_UI is true):

---

## ✓ Phase {Z} Complete

## ▶ Next Up — [${PROJECT_CODE}] ${PROJECT_TITLE}

**Phase {Z+1}: {Name}** — {Goal from ROADMAP.md}

`/clear` then:

`/gsd:discuss-phase {Z+1}` — gather context and clarify approach

---

**Also available:**
- `/gsd:ui-phase {Z+1}` — generate UI design contract (recommended for frontend phases)
- `/gsd:plan-phase {Z+1}` — skip discussion, plan directly
- `/gsd:verify-work {Z}` — user acceptance test before continuing

---

If next phase has no UI:

---

## ✓ Phase {Z} Complete

## ▶ Next Up — [${PROJECT_CODE}] ${PROJECT_TITLE}

**Phase {Z+1}: {Name}** — {Goal from ROADMAP.md}

`/clear` then:

`/gsd:discuss-phase {Z+1} ${GSD_WS}` — gather context and clarify approach

---

**Also available:**
- `/gsd:plan-phase {Z+1} ${GSD_WS}` — skip discussion, plan directly
- `/gsd:verify-work {Z} ${GSD_WS}` — user acceptance test before continuing

---

Route D: All phases complete (milestone ready to close)

---

## 🎉 Milestone Complete

All {N} phases finished!

## ▶ Next Up — [${PROJECT_CODE}] ${PROJECT_TITLE}

**Complete Milestone** — archive and prepare for next

`/clear` then:

`/gsd:complete-milestone ${GSD_WS}`

---

**Also available:**
- `/gsd:verify-work ${GSD_WS}` — user acceptance test before completing milestone

---

Route F: Between milestones (ROADMAP.md missing, PROJECT.md exists)

A milestone was completed and archived. Ready to start the next milestone cycle.

Read MILESTONES.md to find the last completed milestone version.

---

## ✓ Milestone v{X.Y} Complete

Ready to plan the next milestone.

## ▶ Next Up — [${PROJECT_CODE}] ${PROJECT_TITLE}

**Start Next Milestone** — questioning → research → requirements → roadmap

`/clear` then:

`/gsd:new-milestone ${GSD_WS}`

---
**Handle edge cases:**
  • Phase complete but next phase not planned → offer /gsd:plan-phase [next] ${GSD_WS}
  • All work complete → offer milestone completion
  • Blockers present → highlight before offering to continue
  • Handoff file exists → mention it, offer /gsd:resume-work ${GSD_WS}

If section_manifest is null or "forensic-audit" is in its included list: read and execute gsd-core/workflows/progress/steps/forensic-audit.md. Otherwise skip — do not read the file.

<success_criteria>

  • Rich context provided (recent work, decisions, issues)
  • Current position clear with visual progress
  • What's next clearly explained
  • Smart routing: /gsd:execute-phase if plans exist, /gsd:plan-phase if not
  • User confirms before any action
  • Seamless handoff to appropriate gsd command </success_criteria>