e6d047decc23dfc802a0ffcc032d4a23b863bf7f
645 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b3906c66f6 | chore: sync next package version to 1.13.0 | ||
|
|
f9f72cb54c |
enhance(#3777): opt-in concurrent per-plan planners in chunked mode (#4346)
* test(#3777): add failing-first coverage for concurrent per-plan planner dispatch Extracts and executes the real bash blocks this PR is about to add to plan-phase.md and chunked-planning-mode.md (CHUNKED_PARALLEL resolution and the BATCH_PLAN_IDS dedup guard), plus config-set/config-get coverage for the new planning.chunked_parallel key. Expected RED against the current shipped workflow text — the extraction anchors do not exist yet. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(#3777): dispatch chunked mode's per-plan planners concurrently within a Wave Adds opt-in planning.chunked_parallel (default false, byte-identical to the existing serial loop). When true and the runtime's negotiated dispatch capacity (dispatch-capacity, #3673) is greater than 1, chunked planning's per-plan Tasks that share one outline Wave are issued together instead of one at a time; a later Wave still waits for the current one to be verified on disk and committed. A host with no declared maxConcurrency (most non-Claude runtimes today) stays serial regardless of the setting. Resolution and the Plan-ID dedup guard live in chunked-planning-mode.md itself (gated on the section's own CHUNKED_MODE skip-check) rather than in plan-phase.md, so a non-chunked run pays no extra gsd_run calls. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#3777): repoint extraction at chunked-planning-mode.md after the move CHUNKED_PARALLEL resolution moved out of plan-phase.md into chunked-planning-mode.md itself (see the preceding commit); update the test's extraction path and header comment to match. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3777): relocate the canonical runtime-launcher preamble before its first use The CHUNKED_PARALLEL resolution block's two gsd_run calls landed earlier in the file than the sole existing preamble (in the commit step), which tests/runtime-launcher-parity.test.cjs's (B) check requires to precede every gsd_run call in the file. Move the preamble (not duplicate it) to the top of the resolution block; the commit step's fenced block now just calls gsd_run directly. Caught by the GREEN checkpoint gsd-test run before push. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3777): strip the canonical preamble from the extracted resolution block The CHUNKED_PARALLEL resolution fence now carries the relocated runtime-launcher preamble as its first line (previous commit). Extracting the whole fence and running it after the test's own gsd_run stub let the embedded preamble's own resolver logic `unset -f gsd_run` and exit 1 before reaching the resolution logic, since no real gsd-tools.cjs exists in the temp script dir — every test calling runChunkedParallelResolution() failed. Strip the preamble (sourced from gsd-core/workflows/_runtime-launcher.snippet.sh, the same file scripts/sync-runtime-launcher.cjs treats as canonical) before splicing in the stub, so this suite tests only the resolution logic it is actually about. Caught by the post-rebase gsd-test run before push. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3777): add the How-To page the phase gate requires Enablement is 2 commands (config-set, then --chunked), which this repo's own doc-quadrant gate flags as how-to-owed: a reference table cannot carry a sequence. Covers enablement, the dispatch-capacity gate's honest "most runtimes today: no effect" case, and the two accepted trade-offs. An earlier reasoning pass (recorded in .gsd/phase/.../70-docs.json before this commit) had incorrectly claimed #3034 shipped with no equivalent how-to page, as precedent for skipping one here. That claim was false — docs/how-to/enable-parallel-reviewer-lanes.md exists and is indexed. The phase gate caught the omission before merge; corrected here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3777): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1db726ebbf |
feat(#3806): canonize the Review Dispositions Ledger contract (#4345)
* test(#3806): add parity tests for the Review Dispositions Ledger contract Failing-first: asserts references/planner-reviews.md, workflows/plan-phase.md, and agents/gsd-plan-checker.md agree on a single canonical "Review Dispositions Ledger" heading, its round-scoping, L##@{sha} anchor format, and append-only supersession rule. These fail until the canon and its two references are added. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(#3806): canonize the Review Dispositions Ledger contract Promote the existing planner-reviews.md Step 4 return-payload tables (Review Feedback Addressed/Deferred) into a canonical `## Review Dispositions Ledger` PLAN.md section, stated once in planner-reviews.md and referenced (not restated) from plan-phase.md's <review_incorporation_contract> and gsd-plan-checker.md's Review Incorporation dimension. Adds round-scoping (`### Round {N} — {REVIEWS_sha}`), a `L##@{sha}` line-anchor format so a REVIEWS.md reference survives the file being rewritten each round, and an append-only supersession rule. Scoped to part 1 only per the maintainer's approved-feature verdict — the deterministic lint/check verb (part 2) is explicitly deferred to a follow-up. Also: ADR-3806 recording the decision, a docs/features/ fragment (FEATURES.md is generated), and a changeset fragment. Closes #3806 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): fenced-example count bug and lint findings from review - tests/plan-review-convergence.test.cjs: the "heading exactly once" test counted the canonical heading text globally, so it also matched the illustrative fenced-code example in planner-reviews.md that shows the same heading as sample content, always failing 2 !== 1. Rewritten as a bounded line scanner that skips fenced blocks (found by an isolated adversarial review pass). Also bounded an unbounded regex quantifier over readFileSync content flagged by local/no-unbounded-quantifier. - docs/features/review-dispositions-ledger.md: match house fragment style (bold-lead paragraphs, not #### headings) per the Standards-axis review; regenerated docs/FEATURES.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): fit reference-cite fix within size hard caps; ack growth Trims the plan-phase.md / gsd-plan-checker.md reference-cite text to a single short clause pointing at gsd-core/references/planner-reviews.md (also fixes the bare `references/planner-reviews.md` cite the #3576 shipped-reference-cites gate rejects), bringing both files back under their SIZE hard caps and the plan-phase.md phase6 shrink-only baseline. Both files still grow slightly versus origin/next, acknowledged below per ADR-2719's emitted-drift-ack contract. Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap. Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): correct malformed Emitted-Drift-Ack-Growth trailer block The previous commit's two Emitted-Drift-Ack-Growth trailers were separated from the Co-Authored-By trailer by a blank line, so git's own trailer parser (which tests/helpers/emitted-runtime.cjs reads via `%(trailers:key=...)`) only recognized the last contiguous block (Co-Authored-By) and treated the Ack-Growth lines as ordinary body text — invisible to the emitted-attribution gate, not malformed data. Restating them here as one contiguous trailer block, git log over the PR range aggregates trailers from every commit, so this is additive. Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap. Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3806): isolate the ack-trailer paragraph as its own trailer block Git's trailer parser requires the trailer paragraph to be the message's final paragraph, preceded by a blank line, and to contain nothing but trailer-shaped lines. The prior commit's blank line before the trailer lines was missing, which folded the leading Emitted-Drift-Ack-Growth lines into an ordinary prose paragraph. Emitted-Drift-Ack-Growth: gsd-plan-checker.md — adds a short pointer (in the existing Review Incorporation bullet) to the canonical Review Dispositions Ledger location (#3806); stays within the LARGE hard cap. Emitted-Drift-Ack-Growth: plan-phase.md — adds a short pointer (in the existing review_incorporation_contract bullet) to the canonical Review Dispositions Ledger location (#3806); stays under the XL hard cap and the phase6 shrink-only baseline. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3806): backfill PR #4345 into changeset and ADR Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c20675cc4d |
fix(#3819): widen executor's pre-commit guard beyond worktree mode (#4343)
* fix(#3819): widen executor's pre-commit guard beyond worktree mode The pre-commit protected-branch assertion in the executor agent (#2924) only fired inside a Claude Code worktree and matched a hardcoded five-name branch list. It never ran in an ordinary checkout and never covered this repo's own default branch ("next"), so gsd-executor could commit planning-repo documents directly onto a shared checkout's default branch with no PR ever created. Widen the guard to run in every isolation mode, and resolve the protected branch via the repository's actual default branch (with the existing five-name list retained as a fallback when the resolver itself cannot be invoked) plus any configured git.protected_branches. Add a git.allow_default_branch_commits escape hatch for projects that intentionally execute on their default branch. Also point the separate <final_commit> commit helper back at the same guard, so it cannot be sidestepped by that path. Emitted-Drift-Ack-Growth: gsd-executor.md — widened pre-commit protected-branch guard (#3819); tightened comments to stay under the size cap. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3819): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2f920c5af3 |
fix(#4097): evidence preservation no longer sweeps the run's own input copies into .review-diagnostics/ (#4329)
* test(#4097): failing-first regression — run input copies swept into .review-diagnostics The present_results preserve+cleanup glob `_DIAG_MD=( "$RUN_DIR"/gsd-review-*.md )` matches both lane outputs and the run's own assembled input copies (prompt, instructions, roadmap, per-plan copies, project/context/research/requirements, per-lane trimmed prompts) because both share the gsd-review- prefix. Seed the full input-copy set into the existing runWriteReviewsFlow fixture and assert (a) the diagnostics dir holds exactly the lane report + .err sidecar and no input basenames, and (b) an inputs-only run creates no diagnostics dir at all and still cleans up. Both RED against the shipped block. * fix(#4097): preserve lane output only — exclude the run's input copies from evidence The preserve+cleanup glob treated every gsd-review-*.md in RUN_DIR as lane evidence, but the workflow itself writes the run's assembled INPUTS there under the same prefix (prompt, instructions, roadmap, per-plan copies, project/ context/research/requirements, per-lane trimmed prompts). Filter _DIAG_MD by basename against that closed input set instead. Direct glob iteration + case filter — identical under bash and zsh (#4099/ #4109), nullglob-safe (#2962), and the exclusion list is closed and owned in this step: a future input basename cannot silently rejoin the evidence set. Lane reports, diagnostic stubs and non-empty .err sidecars are unchanged. Emitted-Drift-Ack-Growth: review.md — #4097 narrows the present_results evidence-preservation glob: the closed input-basename exclusion list (case filter) plus its rationale note are deliberate additions so a future input basename cannot silently rejoin the evidence set. * changeset(#4097): fixed — review diagnostics no longer sweep run input copies * changeset(#4097): backfill PR number 4329 --------- Co-authored-by: sim <sim@local> |
||
|
|
1017898cb9 |
fix(#3771): make remediation examples non-binding and surface revision conflicts (#3916)
* fix(#3771): separate the binding property from the advisory remediation
Checker findings fused "what property failed" with "how to fix it" into a
single `fix_hint` and never marked which half binds. The checker rendered
every hint under a "must fix" heading, the orchestrators injected the issues
verbatim and ordered targeted updates, and the shared revision references
mapped each hint to a prescriptive strategy — so a contract-following planner
applied a hint literally even when a smaller mechanism satisfied the same
property, or when the hint contradicted a locked decision. There was no
channel to report that conflict, and every attempt burned a revision
iteration.
Checker side: every issue now carries a binding `required_property` (the
invariant that failed) plus its evidence and severity, and `fix_hint` is
labelled non-binding wherever it appears — including the human-facing blocker
rendering, so "must fix" unambiguously names the property and never the
example.
Planner side: revision re-checks locked decisions, capability guidance and
existing plan constraints before editing; satisfying a blocker through a
smaller valid alternative counts as addressing it; and a hint that conflicts
with any of those returns `REVISION_CONFLICT` carrying the conflict and the
alternatives considered. Orchestrators route that to user choice or the
configured plan-review convergence loop without consuming retry budget.
Also applied to the UI-spec revision loop and the gap-plan hint, and the
generic pattern's stray `suggested_fix` field name is reconciled to the
plan-checker's `fix_hint`.
Nothing legitimately binding is weakened: blockers still block, severity
still gates, iteration caps and stall escalation still fire, and required
task fields and decision coverage still hold.
Refs #3771
* test(#3771): pin the binding/advisory split across the revision chain
Locks the separation at every link that carries it: the checker's issue
schema and blocker rendering, the planner's constraint re-check and
REVISION_CONFLICT return, the generic pattern's reconciled field names, and
each orchestrator's conflict routing without retry-budget consumption. Also
pins what must not have been weakened — blockers, severity gating, iteration
caps and stall escalation.
Red against the pre-fix prose: 32 of 34 assertions fail (the 2 that pass are
the preservation checks, correctly).
Refs #3771
* chore(#3771): add changeset fragment for the remediation-binding fix
* chore(#3771): acknowledge the remediation-binding growth
Five runtime-loaded files grow: the two checkers carry the binding/advisory
split where the model reads it (a `required_property` on every dimension
example, since a schema the examples contradict teaches the examples), and
the three orchestrators carry the REVISION_CONFLICT route, which has to live
with the `iteration_count`/`revision_count` state it declines to spend.
Deletes tests/emitted-drift-acks/3172-stated-failing-direction.json: it is
fully spent on next and still owned plan-phase.md, so it walls off a key it
can no longer clear (#3078). Its removal is the documented remedy for the
duplicate-key collision, not drive-by cleanup.
* fix(#3771): close the review gaps in the conflict contract
Adversarial review (Codex, Antigravity) found four real defects in the first
pass, each confirmed against the source before acting:
- The UI checker's structured return still ordered `Fix: {exact fix required}`
and "list each BLOCK dimension with exact fix required". The dimension
examples had been marked non-binding but the rendering the researcher
actually reads had not — the same omission this issue is about.
- `ui-phase` and the canonical `revision-loop` flow incremented their counter
BEFORE dispatching the reviser, so "do NOT increment on REVISION_CONFLICT"
was unreachable prose: the iteration was already spent. The increment now
sits on the return path in both.
- The conflict gate offered "accept as-is", which is an early exit from a
still-failing blocker — a weakening the brief explicitly forbids. The three
options are now adopt an alternative / override the constraint / amend the
constraint; every one resolves the conflict. Accepting an unaddressed blocker
remains available only at the unchanged iteration-cap escalation.
- The convergence route was declarative: nothing in
plan-review-convergence.md could receive a conflict. plan-phase now records
it in REVIEWS.md — the channel that loop already consumes — convergence
refuses to declare convergence over an open entry, and routing back into a
run convergence itself started is explicitly excluded as a cycle. `quick` has
no REVIEWS.md and no phase, so its convergence branch was dead prose and is
deleted in favour of asking the user.
Also reconciles the last two drifted field names (`finding`, `affected_field`)
to the plan-checker schema, and repairs a silent no-op: the few-shot
`required_property` insertion never applied because those lines are
blockquoted, and the test's own block filter was anchored on indentation only,
so a vacuous loop passed over zero blocks. Both are fixed and the filter now
asserts it found blocks.
Refs #3771
* chore(#3771): extend the growth acknowledgment for the review round
plan-review-convergence.md joins the list: the conflict route needed a
receiving end, and it lands on the seam that loop already reads (REVIEWS.md)
rather than a new mechanism. The plan-phase, ui-phase and gsd-ui-checker
entries gain the second-pass reasoning — an executable convergence branch, the
increment moved onto the return path, and the structured return that still
ordered an exact fix.
* fix(#3771): make the conflict route bounded, ordered, and owned
Round-2 adversarial review found five more defects, each confirmed in the
source before acting:
- The convergence gate sat AFTER `gsd_run state planned-phase` and the success
banner, so a run could write and announce convergence over an unresolved
conflict. OPEN_CONFLICTS is now read from REVIEWS.md and is part of the
converged CONDITION, evaluated before any write.
- plan-phase's cycle-exclusion ("unless this run was invoked by convergence")
was not a question the orchestrator can answer at runtime. plan-phase now
never invokes convergence at all — it records the conflict when a phase
REVIEWS.md exists and resolves it with the user in-place, which removes the
cycle instead of describing it.
- Closure had no owner. plan-phase writes the row, so plan-phase strikes it
resolved; convergence only reads. An open row is a live blocker, never a
stale artifact.
- Declining to increment the counter removed the only bound on the conflict
path: an agent returning the same conflict forever would loop unattended. A
conflict naming the same `required_property` twice in a row is now a stall
and escalates through the existing gate.
- verify-work's gap-plan revision loop hands `<revision_context>` to
gsd-planner and so inherits the whole contract, but stated none of it and
could not handle the conflict return. It is now covered like the others, and
is in the test's orchestrator table.
Refs #3771
* chore(#3771): acknowledge the round-2 growth
verify-work.md joins the list — the flow the second review found missed — and
the plan-phase, ui-phase and plan-review-convergence entries gain the
round-2 reasoning: the gate moved ahead of the state write, the convergence
hand-off replaced with a runtime-checkable record-and-resolve, and the
recurrence bound that replaces the counter the conflict path stopped spending.
* fix(#3771): make the convergence gate countable and stop the conflict fall-through
Third adversarial round (Antigravity) found three defects:
- The OPEN_CONFLICTS pipeline had no `grep -v '~~'` despite its own comment
claiming one, and `grep -c '^| '` also counts a markdown table's header and
separator rows — every resolved conflict would have read as open and
convergence would have deadlocked instead of converging. plan-phase now
records each conflict as a `- [ ]` checklist line and flips it to `- [x]`, so
the gate is an exact fixed-string match with no table parsing.
- "then continue below" fell through to the checker spawn, so a SECOND
REVISION_CONFLICT would have been handed to the checker as though it were a
revised plan. plan-phase, quick and verify-work now re-evaluate the return
from the top of the conflict handler; ui-phase already looped back.
- revision-loop.md still described plan-phase routing a conflict to the
convergence loop instead of asking — the behaviour round 2 removed. Recording
is now stated as being in addition to asking, never instead of it.
Refs #3771
* chore(#3771): bring the changeset in line with what shipped
Two review rounds widened the change after the fragment was written:
verify-work's gap-plan revision and the convergence loop are covered, two
more drifted field names are reconciled, and the conflict path carries an
explicit recurrence bound.
* fix(#3771): declare and emit the REVISION_CONFLICT marker
check:contract-drift on CI caught what local lint never reached: four
workflows dispatch on `## REVISION_CONFLICT`, but no agent declared or emitted
it — an orphan consumer, matching a marker nothing produces. The shared
reference (planner-revision.md Step 7b) described the return; the agent
definitions did not carry it.
gsd-planner and gsd-ui-researcher now emit the marker in-fence alongside their
other return markers, and both registry rows in agent-contracts.md declare it.
gsd-planner's Consumed by gains the two workflows that dispatch on it and were
missing from the row.
The gate is right: a return contract belongs where the agent is defined, not
only in a reference the agent happens to load.
Refs #3771
* chore(#3771): acknowledge the return-marker growth
gsd-planner.md and gsd-ui-researcher.md each gain the REVISION_CONFLICT
marker that check:contract-drift requires them to emit.
* fix(#3771): hoist the shared conflict protocol out of the workflows
Two CI failures, both correct gates:
- tests/few-shot-calibration.test.cjs pins the plan-checker calibration file
at exactly 4 examples (2 positive, 2 negative). The example added in the
first pass broke that balance — and described PLANNER behaviour in the
CHECKER's calibration set, which is the wrong surface for it. Removed; the
smaller-alternative rule is already normative in gsd-plan-checker.md and
planner-revision.md, and pinned by the regression suite.
- tests/phase6-capstone-conformance.test.cjs (ADR-857 phase 6, #1168) requires
plan-phase.md to stay BELOW its pre-phase-6 baseline of 94519 bytes. The
inline conflict block pushed it to 94988.
The fix for the second is the one that should have been made first: the
record/resolve/close protocol and the recurrence bound were identical in four
workflows, and revision-loop.md — which plan-phase already @-imports — is what
a shared contract is for. The protocol now lives there once; plan-phase states
only its bindings (which counter, which artifact, which next step) and points
at it. plan-phase.md: 94988 -> 92739, under the ratchet with headroom, and the
four-way duplication is gone.
quick, ui-phase and verify-work do not import the reference, so they keep their
inline statements. The suite asserts each rule against what the runtime
actually loads for that orchestrator, not against the file in isolation.
Refs #3771
* docs(#3771): state the shared-protocol relationship accurately
Three of the four revision-bearing workflows do not @-import revision-loop.md,
so 'follows it verbatim' overstated the coupling. Only plan-phase defers; the
others restate the rules inline and this section is the authority they must
agree with.
* refactor(#3771): name the authority instead of restating it four times
Self-review finding: the same ~700-byte reviser paragraph was inlined in four
prompts while gsd-planner already loads planner-revision.md whenever
<revision_context> is present (agents/gsd-planner.md:555, :588) — a fifth copy
of the same contract. Four places to edit in lockstep is precisely the drift
class this PR exists to fix; the generic pattern calling the field
suggested_fix while the checker emitted fix_hint is what that looks like after
a year.
Each prompt now carries only the load-bearing clauses and names the authority
it summarises. Deliberately NOT reduced to a bare pointer: these are LLM
prompts, and a contract stated only in a file the reader is supposed to fetch
is the failure mode of this very bug. The saving is modest (~50 bytes each) —
the point is the named source of truth, not the bytes.
Suite: dropped seven assertions that pinned heading text and bold-lead
phrasing a reword would break without changing what the runtime is told; their
neighbours already pin the same contract by content. 519 -> 499 lines, 61
tests. Red gate against origin/next: 56 of 61 fail.
Refs #3771
* fix(#3771): sanitize agent-authored conflict text and bound total conflicts
Cross-AI review (agy/Gemini) raised two security findings. I confirmed the
first by running the gate's own awk against a crafted file and reproduced it
exactly; the second is a real hole in my bound.
1. FAIL-OPEN in the convergence gate. The conflict text is agent-authored and
is appended to REVIEWS.md, which the gate scans with an awk that stops at
the next '## ' heading. One line of agent text beginning '## ' ends that
scan early, so conflicts below it are never counted and convergence declares
success over a live blocker. Measured: 3 open conflicts, awk returned 2.
Fixed at the write boundary, which is the trust boundary: every field has
newlines and tabs collapsed to spaces and a leading '#', '-', '|' or fence
stripped, so one conflict is exactly one line. Both producing agents now
declare their fields single-line plain text, and the reader states the
invariant it depends on so a later edit cannot silently break it. Verified:
3 open + 1 resolved now counts 3; missing file and absent section count 0.
2. The recurrence bound was 'same required_property twice in a row', which an
agent alternating property names never trips, leaving the un-incremented
conflict path unbounded. Now bounded twice: the repeat rule catches the
common case, and the THIRD conflict return of a loop escalates whatever
property it names. A conflict still never consumes a revision iteration;
this cap is separate from and additional to the revision cap.
Rejected from the same review: deleting 'a planner that reaches
required_property by a smaller or different mechanism has addressed the issue
in full' from the CHECKER prompt as misplaced. It is load-bearing exactly
there. A checker that does not know a different mechanism counts will re-flag
the issue on re-check, which is the revision loop that never terminates. The
argument offered for deleting it, that the checker evaluates the new state
independently, describes the failure mode.
Refs #3771
* fix(#3771): fail closed on an unverifiable convergence gate
Second cross-AI pass (agy, this time with the full files rather than the diff)
found two more, both real:
1. The gate read REVIEWS_FILE with `2>/dev/null || echo 0`, so an unreadable or
empty path counted as ZERO open conflicts and converged. That path is
resolved a few lines earlier by a pre-existing unquoted
`ls ${phase_dir}/${padded_phase}-REVIEWS.md` (line 346, not touched by this
PR), which yields an empty string rather than an error when the path
contains a space. Unverifiable is not clean: the gate now tests -z and -r
first and BLOCKS. Verified both branches.
The unquoted ls itself is left alone deliberately — it predates this change
and belongs to the reviews lookup, not the conflict gate. Fixing it at my
own boundary removes its effect on this gate without widening scope.
2. REVIEWS.md is writable by the review agent, which could flip a `- [ ]` to
`- [x]` or delete the section and forge the state of a blocking gate. The
section now declares a single writer: /gsd:plan-phase appends and closes,
every other agent leaves it byte-for-byte alone, readers read.
Also trimmed a clause that explained the increment ordering by reference to
what the file said before this PR. Commit history is not instruction, and
these files are prompts.
Rejected: the claim that quick's conflict gate deadlocks autonomous pipelines
by asking the user. Its existing max-iteration escalation in the same file
already asks the user the same way; this adds no new interaction class.
Noted but out of scope: the per-dimension YAML example blocks and the shim
boilerplate duplicated across agent prompts both predate this change.
Refs #3771
* fix(#3771): count conflicts by line shape, not by section
CodeRabbit review on the rehearsal PR. Five findings, all valid, all applied.
The best one is a deletion. The convergence gate scanned between
'## Plan-Revision Conflicts' and the next '## ' heading, and that scan stops at
the FIRST heading it meets — so one stray '## ' line hid every conflict beneath
it and returned 0, converging over a live blocker. Reproduced: section-scan 0,
shape-scan 1. Sanitizing at the write boundary does not cover a hand-edited,
legacy, or foreign-written REVIEWS.md, so the reader needed its own guarantee.
It now matches the conflict line SHAPE anywhere in the file:
grep -c '^- \[ \] .*required_property:'
No section bookkeeping, nothing a heading can truncate, and it composes with the
writer's sanitization (which strips a leading '-' from agent text, so agent prose
cannot forge the shape). Verified: injected heading -> 1, all resolved -> 0.
The other four:
- Both checkers told the author never to emit a contradictory fix_hint, then
offered an escape hatch that put the forbidden route in the hint anyway. They
now name NO route in that case and state only that the property conflicts with
the constraint. A hint carrying a forbidden route is applied by anyone who
trusts hints.
- The REVISION_CONFLICT marker description in gsd-planner.md was narrower than
planner-revision.md: it covered a contradictory hint but not an unreachable
required_property. A planner reading only the agent file would have burned
retry budget on the case the reference routes to a conflict.
- The few-shot calibration examples used uppercase BLOCKER/INFO while the schema
defines blocker/warning/info. Pre-existing, but it is the same schema-vs-example
disagreement this PR exists to end, and the file was already being edited.
- verify-work's re-entry instruction existed but sat after the Bounded clause, so
the paragraph read "re-spawn ... stop re-spawning ... after re-spawning". The
re-entry now immediately follows the re-spawn, and states that only a
non-conflict return may reach the checker or increment iteration_count.
Refs #3771
* fix(#3771): resolve the contradictory scope_sanity severity examples
Sixth CodeRabbit finding, posted outside the diff range and missed on my first
read — I had claimed all findings were addressed after reading only the five
inline comments. This one was in the review body.
agents/gsd-plan-checker.md carried TWO scope_sanity examples with identical
metrics (5 tasks, 12 files) and OPPOSITE severities: warning in Dimension 5,
blocker in <examples>. Line 872 states "2-3 tasks/plan good, 4 warning, 5+
blocker" and the severity table lists warning as "Scope 4 tasks (borderline)",
so the warning example contradicted both.
ADR-2629 Decision 5's "over budget is a WARNING, never a blocker" does not
excuse it: that rule governs the smart-zone TOKEN estimate (the estimate-check
verb, lines 299-306), which is a different axis from task count. Verified in
source before touching it.
The contradiction is pre-existing but this PR made it binding and visible:
severity is now declared part of the binding payload, and both examples were
given the same required_property, so they now disagree on the severity of an
identical finding about an identical property.
Deviating from the proposed correction, which was warning -> blocker: that
would duplicate the <examples> entry outright (same tasks, files, severity).
The Dimension 5 example is instead made a genuine 4-task borderline warning, so
the file keeps one worked example per severity and the thresholds, the severity
table and both examples finally agree.
Refs #3771
* fix(#3771): stop laundering a grep error into zero open conflicts
Seventh CodeRabbit finding — from a SECOND review round my own CR-4 push
triggered, which I had not looked for. This one is a regression I introduced
while fixing the previous fail-open.
CR-4 replaced the truncatable section scan with:
OPEN_CONFLICTS=$(grep -c '^- \[ \] .*required_property:' "$REVIEWS_FILE" || true)
`|| true` masks every grep failure. grep exits 1 for "no matches" (a legitimate
zero) but 2 for a read error, and `|| true` turns both into an empty capture
that `${OPEN_CONFLICTS:-0}` renders as 0. If REVIEWS.md is removed or becomes
unreadable between the -r check and the scan, the gate reports no conflicts and
convergence proceeds. Proven: unreadable file -> captured empty -> 0.
The status is now inspected, and only exit 1 counts as zero; anything else
blocks.
My first attempt at this fix was itself wrong and my own harness caught it: I
wrote `if ! grep ...; then grep_status=$?`, but `!` inverts the status, so `$?`
in that branch is 0 and every failure reads as success — the clean-file case
printed "BLOCKED (grep exit 0)". The status must be read in the ELSE branch of a
non-negated `if`, which is what CodeRabbit proposed. Both traps are now pinned
by tests.
Verified end to end: all resolved -> 0, no conflicts at all -> 0, injected
heading -> 1, unreadable file -> BLOCKED with grep exit 2.
Refs #3771
* test(#3771): execute the conflict gate instead of reading it
CodeRabbit round three: 0 actionable, 1 nitpick — "these assertions inspect
Markdown source only; they do not prove that grep status 1 produces zero
conflicts or that a scan error exits before convergence." Rated Trivial. It is
the most valuable finding of the three rounds.
This gate has been wrong three times: a section scan a heading could truncate, a
`|| true` that laundered grep's error status into zero, and an `if !` whose `$?`
reported the negation rather than the command. Every one of those passed the
text assertions that existed at the time. I proved each fix by hand in a shell,
and none of that proof lived in the suite.
The gate is one self-contained fenced block, so the test now extracts it from
the workflow — located by content, not line number — writes it to a script and
RUNS it against fixtures: two open plus one resolved counts 2; no matches counts
0 and does not fail; a conflict below an injected `## ` heading still counts; an
unreadable path and an empty path both BLOCK with a non-zero status and no zero
count on stdout.
Non-vacuity proven by mutation rather than asserted. Reverting the gate to each
of its three historical broken forms reds the suite:
section-scan awk -> 7 failures (5 in the gate cases)
|| true -> 4 failures (3 in the gate cases)
if ! (negated $?) -> 4 failures (3 in the gate cases)
restored -> 69 pass, 0 fail
The prose assertions stay: they are the right instrument for a prompt. This
covers the one part of the change that is real shell an orchestrator executes.
Refs #3771
* test(#3771): route the gate harness through the shared test helpers
ESLint's project rules caught three violations in the new harness: an unbounded
execFileSync (DEFECT.UNBOUNDED-SUBPROCESS — an unbounded spawn is an indefinite
hang, and on macOS CI that is how a stuck run stops reporting instead of failing)
and two raw fs.rmSync calls, which skip the Windows-EBUSY retry budget that
helpers.cleanup carries.
Now uses createTempDir/cleanup from tests/helpers.cjs and passes an explicit
30s timeout. Suppressing the rules was available and would have been the wrong
call: both exist because of real CI failure modes on platforms I am not testing
on.
* chore(#3771): backfill the changeset PR number
The pr: field is drift-checked against the PR event payload, so it cannot be
written before the PR exists. Set to 3916.
* fix(#3771): close revision conflict persistence gaps
Use the authoritative review artifact, keep conflict and normal retry paths disjoint, and enforce one writer-reader grammar so malformed state fails closed.
Emitted-Drift-Ack-Growth: diagnose-issues.md — #3771 marks the gap-plan remediation hint non-binding while keeping root_cause authoritative
Emitted-Drift-Ack-Growth: gsd-plan-checker.md — #3771 separates binding required_property evidence from advisory fix_hint examples across the checker contract
Emitted-Drift-Ack-Growth: gsd-planner.md — #3771 declares the REVISION_CONFLICT return used when remediation contradicts governing constraints
Emitted-Drift-Ack-Growth: gsd-ui-checker.md — #3771 applies the same binding-property and advisory-hint split to UI review findings
Emitted-Drift-Ack-Growth: gsd-ui-researcher.md — #3771 defines the UI revision producer's structured REVISION_CONFLICT return
Emitted-Drift-Ack-Growth: plan-phase.md — #3771 routes and persists bounded revision conflicts before spending the normal retry budget
Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3771 adds the fail-closed owned-block parser and prevents convergence over open conflicts
Emitted-Drift-Ack-Growth: review.md — #3771 emits and preserves the canonical writer-owned conflict block across review regeneration
Emitted-Drift-Ack-Growth: ui-phase.md — #3771 routes UI revision conflicts to resolution before consuming revision_count
Emitted-Drift-Ack-Growth: verify-work.md — #3771 gives gap-plan revision the same bounded conflict route before iteration_count
* test(#3916): guard rebases against schema drift
Load the current-base progressive-disclosure examples so every integrated issue remains bound by required_property after branch reconciliation.
* test(#3916): skip the extracted-gate suite's bash spawns on win32
Third review round's sole survivor: runConflictGate()/withReviews() spawn
bash against a Node-native temp path built by createTempDir(), which is
backslash-separated on the Windows CI lane and not a path Git Bash is
guaranteed to accept (DEFECT.WINDOWS-TEST-PORTABILITY, matching the
observed CI failure at revision-remediation-binding.test.cjs:844,
ENOENT on a path Windows read as a directory separator). No eslint rule
catches it since the call has neither a chmod nor a `bash -c` form.
Guards the four call sites with the repo's existing skipOnWin32
convention (describe/test `{ skip: IS_WINDOWS }`) rather than
normalizing the harness path to forward slashes, which would defeat the
one test whose purpose is proving the production gate does NOT rewrite
a literal backslash in a POSIX filename.
* fix(#3916): backfill changeset pr field to the fork validation PR number
* fix(#3771): forbid silently accepting an open plan-revision conflict at max-cycles escalation
The max-cycles escalation prompt only surfaced HIGH_COUNT and ACTIONABLE_COUNT; an open
plan-revision conflict (OPEN_CONFLICTS > 0) was never disclosed there, and "Proceed anyway"
could exit successfully over it — exactly the failure mode this PR exists to close (a success
banner over an unresolved conflict nobody resolved). Blockers still block: withhold "Proceed
anyway" and route to Manual review whenever a conflict is open.
* fix(#3771): do not hard-block REVISION_CONFLICT persistence when no REVIEWS.md exists yet
A phase's first-ever revision cycle can return REVISION_CONFLICT before any REVIEWS.md has been
written — REVIEWS_PATH is then legitimately empty, not a corrupt or deleted file. The persistence
gate's own accompanying prose already says the record channel applies 'when REVIEWS_FILE is
non-empty', but the bash condition never checked that, so it hard-blocked every conflict on a
brand-new phase regardless of whether persistence was even expected to run. Require a non-empty
REVIEWS_FILE before treating a missing file as an error.
* chore(#3771): raise the plan-phase.md ADR-857 host-loop ceiling to 96700
The frozen pre-phase-6 ceiling (94519) collided on rebase: this PR's own
REVISION_CONFLICT persistence/routing gate is core planner control flow, not an
un-extracted optional feature, and landed alongside an unrelated, already-merged
same-file growth (the #4.6 context-drift pre-check) already on next. Same
rationale #1298 already established for execute-phase.md's ceiling.
* chore(#3916): backfill changeset pr field to the upstream PR number
* fix(#3771): make the writer-side REVISION_CONFLICT sanitize step real shell
The Conflict Return record channel sanitized agent-authored fields via a
prose instruction ("Sanitize each agent-authored field before appending")
for the orchestrator LLM to apply by hand, while the reader-side gate in
plan-review-convergence.md parses the same slot with real, executed awk.
Flagged Minor across two review rounds (round 4, round 6) since no code
performed the sanitize anywhere.
plan-phase.md's Conflict Return step now runs a real bash gate: sanitize
each field (collapse newline/tab to space, strip a leading #/-/|/fence),
build the one-line record, skip the append if an identical line already
exists (idempotent), insert before the writer-owned end delimiter, and
fail closed if that delimiter is missing rather than silently dropping
the conflict.
tests/revision-remediation-binding.test.cjs extracts and RUNS the new
fence (matching how the reader gate is already tested), composing it
with the existing reader gate: hostile-field fast-check fuzzing, a
repeated-conflict idempotency check, and a missing-delimiter fail-closed
check that the file is left byte-for-byte unchanged on failure.
Emitted-Drift-Ack-Growth: plan-phase.md — #3916 turns the writer-side
REVISION_CONFLICT sanitize+insert step into real, executed shell instead
of a prose instruction, matching the reader gate's existing rigor
* fix(#3916): backfill changeset pr field to the fork validation PR number
Fork CI's changeset-lint reads the real PR number from its own event
payload; the fragment still carried the upstream number from the last
sync, so the DEFECT.CHANGESET-PR-FIELD-DRIFT check failed on this fork
PR. Re-backfill to the upstream number before the final push.
* fix(#3771): close the awk -v forgery and same-session close gaps agy found
Adversarial review (gemini-3.8-flash-high via the internal agy review
lane) on the full PR found two BLOCKERs against the just-added
writer-side conflict gate:
1. `awk -v line="${LINE}"` decodes escape sequences in its argument, so
a literal two-character `\n` in agent-authored text became a real
newline inside awk, splitting the appended record across two
physical lines. `tr` only strips actual control bytes, so it never
saw this — it defeated the exact forgery the gate exists to
prevent, both the reader's zero-count and the writer's own
idempotency check. Fixed by passing LINE/END through awk's
ENVIRON, which is not escape-decoded.
2. A conflict resolved and re-spawned within the same plan-phase
session was never flipped from `- [ ]` to `- [x]` — the record
channel bullet said "plan-phase closes it," but no step did. Only
a *separate* `--reviews` re-entry (line ~622, still prose-only)
closes conflicts; the in-session resolve path left them open
forever, permanently blocking convergence. Fixed by carrying the
just-written line in `PENDING_CONFLICT` and closing it in the
`Otherwise` branch before the checker re-spawns.
Also fixed a MAJOR: docs/COMMANDS.md described the `--max-cycles`
escalation gate as uniformly offering "proceed or review manually,"
but the code (this PR's own change) withholds "Proceed anyway"
specifically when a plan-revision conflict is open — only manual
review is offered in that case. Docs now say so.
Not applied: the reviewer's `\r` truncated to plain tr from a MINOR
that also asked for temp-file permission preservation across `mktemp`.
Applying `chmod --reference` is not portable to macOS/BSD `chmod`, so
this is left as a documented low-severity tradeoff — the temp file now
sits alongside REVIEWS.md (same filesystem, atomic `mv`), which was
the same finding's more substantive half. Also not applied: a
suggested `gsd_run review record-conflict` CLI subcommand to
deduplicate the two `awk` blocks — a new command plus wiring is out of
scope for a review-remediation fix.
tests/revision-remediation-binding.test.cjs adds regression coverage
for both BLOCKERs: a literal-backslash-n hostile field composed with
the reader gate, and a close-gate extraction that verifies the flip to
`[x]`, the reader's count dropping to 0, and a fail-closed path when
the pending line is missing.
Emitted-Drift-Ack-Growth: plan-phase.md — #3916 fixes an awk -v escape-
decoding forgery and adds the missing same-session conflict-close step
an adversarial review found in the writer-side gate
* chore(#3916): backfill changeset pr field to the upstream PR number
Fork-validation CI needed pr: 1 to pass its own changeset-lint; restore
pr: 3916 before this push reaches open-gsd/gsd-core.
* fix(#3771): trim plan-phase.md prose back under the XL byte cap
Merging origin/next's unrelated growth pushed plan-phase.md 473 bytes
past the workflow-size-budget XL cap and the ADR-857 phase-6 baseline,
both tripped by CI after review approval. Removed an unpinned inert
bash comment and tightened connective prose in three REVISION_CONFLICT
bullets; no executable shell or test-pinned substring changed.
* chore(rehearsal): pin changeset pr field to fork rehearsal PR #25
Scratch-only commit for the rehearsal branch's own CI. Will not be
carried onto the branch backing upstream #3916 — that keeps pr: 3916.
* fix(#3771): address CodeRabbit findings on the REVISION_CONFLICT protocol
Fork rehearsal PR #25's first CodeRabbit pass surfaced 7 findings against
the already-approved #3916 diff; each verified against current code
before fixing (none hallucinated):
- plan-phase.md: writer-side awk gates now strip a trailing \r before
comparing lines, matching the reader gate (plan-review-convergence.md)
-- a CRLF REVIEWS.md previously made both writer gates fail closed.
- plan-phase.md: the close-fence's REVIEWS_FILE/PENDING_CONFLICT/
CONFLICT_RESOLUTION were read without ever being (re)defined in that
fence -- shell state does not survive across separate fenced blocks
(same convention already documented in review.md). Added the explicit
recompute/set instruction.
- revision-loop.md: previous_conflict_property was never reset after a
normal (non-conflict) revision, so a later, unrelated conflict on the
same property could be misread as a repeat and escalate prematurely.
- gsd-plan-checker.md / few-shot-examples/plan-checker.md: two example
required_property strings were unconditionally binding in a way their
own dimension's rules aren't (no-analog RESEARCH.md fallback; tasks
that create no functions), now scoped to match.
- quick/steps/plan-checker-loop.md: added the same disjoint
"Otherwise (not REVISION_CONFLICT)" branch plan-phase.md already had,
closing an ambiguity between the conflict and non-conflict return paths.
- revision-remediation-binding.test.cjs: the REVIEWS_PATH init-order
assertion used indexOf() without checking for -1, so it would pass
vacuously if either anchor were renamed away.
Also restores an "Export the row's CONFLICT_*" instruction I had cut in
the prior byte-budget trim -- checked non-pinned by tests, but it was the
only text telling the agent to set those vars before the awk block reads
them via ENVIRON.
Net growth from these fixes required reclaiming bytes elsewhere in
plan-phase.md (verified against every pinned substring in
revision-remediation-binding.test.cjs) to stay under the XL tier's
hard 98304-byte cap; final size 98245 bytes.
* fix(#3771): resync the #4079 shrink-only mirror to the current PRE_PHASE6 line
tests/plan-phase-background-wait-wakeup.test.cjs (landed on next via an
unrelated #4079 PR, merged in by this branch's next-sync) mirrored
plan-phase.md's phase6 shrink-only ceiling as a hardcoded local constant
(94519) rather than reading tests/phase6-capstone-conformance.test.cjs's
PRE_PHASE6 value. That value has since been legitimately raised twice
during this PR's own review (94519 -> 96700 -> 98300) to accommodate the
REVISION_CONFLICT persistence/routing gate. The two branches' independent
histories left the mirror stale post-merge -- not a textual git conflict,
but the same class of thing. Resynced to 98300.
* fix(#3771): address round-2 CodeRabbit findings on the conflict gates
CodeRabbit's re-review of the previous remediation commit found two real
issues in what it had already flagged:
- Both writer-side awk CRLF fixes used \`sub(/\r$/, "")\` directly on \`\$0\`,
which mutates it in place -- \`{ print }\` then emitted the CR-stripped
copy for every passed-through line, silently rewriting an unrelated
CRLF REVIEWS.md to LF on any insert or close. Now compares against a
separate \`cur\` copy and prints the original, untouched \`\$0\`.
- The close-fence's "recompute REVIEWS_FILE/PENDING_CONFLICT" prose
implied in-fence derivation, but the fence has no such code and the
test harness (\`runCloseGate\`) deliberately supplies all three as
pre-set env vars -- matching how the open fence's "Export the row's
CONFLICT_*" instruction already works. Reworded to "export ... in the
same invocation", matching that established, test-verified pattern
instead of promising logic that isn't there.
Added a regression test proving the CRLF fix no longer touches
passthrough lines (red against the mutate-in-place version, green now).
* fix(#3771): use a CRLF-safe check in the new passthrough regression test
local/no-crlf-fragile-split forbids splitting readFileSync content on a
literal \n (Windows git-autocrlf checkouts yield \r\n). My CRLF
passthrough-preservation test from the previous commit did exactly that
to inspect the first line. Replaced with a direct startsWith() check
against the known CRLF-terminated header, which needs no split.
* test(#3771): assert the record itself is inserted in the CRLF passthrough test
CodeRabbit nitpick (round 3): the passthrough-preservation test checked
gate status and the pre-existing line's CRLF ending, but never asserted
the new REVISION_CONFLICT record was actually written.
* fix(#3771): address agy/gemini-3.8-flash-high adversarial review findings
Full-PR adversarial review (internal /gsd-review antigravity lane,
gemini-3.8-flash-high) surfaced 9 findings; each verified against current
code before fixing (none hallucinated):
HIGH:
- quick-batch/steps/plan-checker-loop.md never received the
required_property/fix_hint binding language or REVISION_CONFLICT
handling this PR added everywhere else -- a genuinely unmigrated
producing context. Migrated to match quick/steps/plan-checker-loop.md,
and added it to the ORCHESTRATORS consistency battery in
revision-remediation-binding.test.cjs so future drift is caught
automatically.
- The close-fence's PENDING_CONFLICT was an agent-supplied env var that
had to exactly reconstruct a five-field sanitized line across a
multi-minute subagent dispatch -- fragile, and a scalar var also meant
a second simultaneous conflict silently dropped the first on overwrite.
Redesigned to match the open conflict by CONFLICT_DIMENSION/
CONFLICT_PLAN identity instead: the agent re-supplies two short,
already-tracked identifiers rather than reconstructing the full
sanitized text, and each conflict resolves independently regardless of
how many are open. Updated the test harness's runCloseGate contract to
match, and added a two-open-conflicts regression test.
MEDIUM:
- plan-phase.md's `--reviews` replanning path told the reader to "flip
the matching line to [x]" in prose only, with no executable path to
it -- pointed it at the same close gate used in step 12.
- plan-review-convergence.md's reader-gate awk tolerated a blank line
before the opening delimiter but not before the heading that follows
it; a formatter or LLM writer inserting one would hard-abort
convergence on an otherwise well-formed REVIEWS.md. Added the same
tolerance already granted above it, with a regression test.
LOW:
- Clarified that the escalation destination for a stalled conflict is
the same iteration/revision-count cap gate already defined in each of
quick, quick-batch, ui-phase, and verify-work, rather than an
undefined "stall" concept.
- Clarified "twice in a row" means no successful revision intervened,
matching revision-loop.md's now-explicit previous_conflict_property
reset.
- Fixed gsd-ui-researcher.md's stale rationale text, copied verbatim
from planner-revision.md: ui-phase presents the conflict table
directly to the user, it does not persist to a shared file scanned by
heading.
Net growth again required reclaiming bytes in plan-phase.md (verified
against every pinned substring in revision-remediation-binding.test.cjs)
to stay under the XL tier's hard 98304-byte cap; removed a now-dead
PENDING_CONFLICT assignment in the process. Final size 98258 bytes.
* fix(#3771): scope row 48's quick/steps guard away from plan-checker-loop.md
tests/gsd-quick-batch-quick-regression.test.cjs's row 48 (#3676) flagged
this branch's quick-batch/steps/plan-checker-loop.md migration (the agy
HIGH finding) as a violation, because it also edits
quick/steps/plan-checker-loop.md for the same underlying #3771 protocol
fix.
Verified against git history before scoping:
|
||
|
|
7ff196c505 |
fix(#4096): honor --dry-run in todo complete and write completion keys inside the frontmatter fence (#4325)
* fix(#4096): honor --dry-run in todo complete and upsert completion keys inside the frontmatter fence * review(#4096): tighten todo complete flag rejection to any dash-prefixed token * chore(#4096): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
2e1ede6d99 |
fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline (#4318)
* test(#4093): regression matrix for advance-plan zero-labeled-fields decline * fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline * refactor(#4093): collapse IIFE to a plain block (review finding) * docs(#4093): document the advance-plan recovery decline + changeset * chore(#4093): backfill PR number in changeset * fix(#4093): budget lint-compiled-artifact-sync's tsc compile as a compile, not a probe --------- Co-authored-by: sim <sim@local> |
||
|
|
b327331747 |
fix(#4086): resolve skills/ manifest keys at the runtime's actual skills root (#4311)
* fix(#4086): resolve skills/ manifest keys at the runtime's actual skills root Codex installs skills to ~/.agents/skills (skills-kind home override), but saveLocalPatches() and verify-reapply-patches.cjs resolved every manifest key config-dir-relative only — every skills/ key missed, so user modifications to Codex skills were never hash-compared, never backed up, and silently overwritten on update; the reapply verifier false-failed the same keys with fail_installed_missing. configDir stays first (non-override runtimes byte-identical); the skills root (same _resolveSkillsRootDir / skillsManifestPrefix seams the write side uses) is a containment-guarded fallback when the config-dir path is absent. * fix(#4086): drop unused test param; add changeset fragment * chore(#4086): backfill PR number in changeset fragment --------- Co-authored-by: agent-4086 <agent-4086@local> |
||
|
|
5869febb16 |
enhance(#4155): invalidate verification results when covered inputs change (#4290)
* enhance(#4155): invalidate verification results when covered inputs change readVerificationStatus() now recomputes a deterministic sha256 fingerprint over a VERIFICATION.md's declared covered_files (phase PLAN/SUMMARY, requirements, implementation files in the verified change set) and returns stale on any mismatch, fail-closed when a covered file is missing, unreadable, or escapes the project root. Legacy reports with no fingerprint metadata keep the prior SUMMARY-mtime staleness check unchanged. The verifier computes covered_digest via the new verification.fingerprint CLI command rather than by hand, since a digest is deterministic math, not an LLM-estimated value. * chore(#4155): backfill fork PR number in changeset * fix(#4155): trim gsd-verifier.md fingerprint instructions to fit LARGE tier byte cap * fix(#4155): address CodeRabbit findings on fingerprint fail-closed behavior Partial fingerprint metadata (one of covered_files/covered_digest present, the other missing or malformed) now fails closed to stale instead of silently downgrading to the legacy mtime-only check. computeCoveredDigest also canonicalizes with realpathSync before re-confining, so an in-root symlink whose target escapes the project root can no longer produce a matching digest. gsd-verifier.md restores the completeness requirement and checklist item trimmed by the earlier size-budget fix, within the LARGE tier byte cap. * chore(#4155): acknowledge gsd-verifier.md growth for the #4155 fingerprint instructions Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the covered-input fingerprint instructions and frontmatter fields the #4155 verification staleness mechanism requires; trimmed to stay within the LARGE tier byte cap * fix(#4155): address gemini adversarial review findings computeCoveredDigest now threads the caller-supplied opts.fs seam through its confinement and read paths instead of always using raw node:fs — a caller like planning-inspect.cts's containmentEnforcingVerificationFs (GAP 2, #2790 follow-up) was silently bypassed for covered-input reads. The project-root anchor itself still canonicalizes through real fs (it is a trusted value the caller derived, not attacker-influenced covered-input data); only per-file candidate reads go through the injected seam. Covered-file paths are now canonicalized (./ prefixes, redundant slashes, internal .. segments) before becoming dedup/sort/hash keys or confinement subjects — closes both a spurious-stale false positive (two spellings of the same file hashing differently) and a confinement gap (an internal .. segment that doesn't start the string). gsd-verifier.md now states covered-file paths are project-root-relative, not phaseDir-relative, closing an ambiguity that would have made a real verifier agent's first fingerprint invocation fail closed. defaultFsImpl's methods now late-bind through fs.<method> rather than capturing function references at module load — the earlier direct-capture form was invisible to existing tests' t.mock.method(fs, 'statSync', ...) seams, a real regression caught by the full suite (not the reviewer). * fix(#4155): catch a plan/summary added to the phase dir after verification but never declared The content digest only recomputes hashes for paths the verifier actually declared in covered_files — it had no way to notice a plan or summary added to the phase directory after verification if that new file was never declared, silently regressing behind the legacy mtime check it replaces (which scans the live directory, not a declared list). findUncoveredCurrentArtifact re-scans the live phase directory for every current *-PLAN.md/*-SUMMARY.md and requires each to be represented in covered_files, closing that gap; a directory scan failure fails closed to stale rather than silently skipping the check. CONTEXT.md's Verification Module entry corrected to describe the fingerprint path's stricter fail-closed FS-error contract (routes to stale) instead of the module's original degrade-to-safe one (missing / not-stale), which only the legacy path still keeps. * refactor(#4155): extract canonicalizeCoveredFiles, add real nested-project e2e test computeCoveredDigest and cmdVerificationFingerprint each normalized/deduped/ sorted covered_files independently — one shared helper now backs both (gemini review's ponytail-lens finding). Adds one CLI-to-readVerificationStatus test against a genuine .planning/phases/NN-x/ project with an implementation file outside .planning/ entirely, closing the review finding that prior #4155 unit fixtures put phaseDir directly under an ownerless tmpdir (findProjectRoot falls back to phaseDir itself there) and never exercised real multi-level path resolution. * fix(#4155): route computeCoveredDigest through real fs, fail closed on unreadable plans/ Two independent review rounds (opus critical-reviewer + opus ponytail + agy, run twice) found two instances of the same fail-open class: - computeCoveredDigest's per-file reads routed through the caller's injected fsImpl. planning-inspect.cts passes a `.planning/`-confined containment fs into readVerificationStatus's opts.fs, so any covered implementation file outside `.planning/` (mandatory per the issue) made the confinement wrapper throw, which was caught and turned into a stale digest -- reporting every fingerprinted phase permanently stale via `planning.inspect`, regardless of actual drift. Per-file reads now always use real node:fs, matching the pre-existing treatment of root canonicalization; the realRel-vs-realRoot check is the real confinement boundary for this data and needs no seam. - allCurrentArtifactsCovered's try/catch never fired (scanPhasePlans reports readdir failures via a `scope` field, it never throws), so an unreadable nested plans/ dir was silently treated as "zero artifacts, all covered" instead of failing closed. Now branches on scope !== SCOPE.COMPLETE. Also, per ponytail's second-round findings: reverted an unwarranted FINGERPRINT_VERSION bump and digest length-prefix from the first fix (no v1 digest has ever existed -- the feature is unreleased -- and the prefix closed a collision that grants no capability beyond what a writer of covered_files already has more cheaply); removed a verifier-facing escape-hatch instruction whose own example was a case that should trigger staleness, not bypass it; corrected CONTEXT.md references to the renamed allCurrentArtifactsCovered and a stale "unconditional" rescan claim; simplified the isStale derivation, removed dead FsLike members, and tightened test coverage. Regression tests for both fail-open bugs are included and were each confirmed to fail against the pre-fix code before the fix landed. full test suite: 2558/2560 pass, 2 skipped, 0 fail * fix(#4155): trim gsd-verifier.md under the LARGE size cap Fork CI caught what my local runs missed: the superseded/nested-plans instruction added earlier pushed gsd-verifier.md to 49299 bytes, 147 over the LARGE tier's 49152-byte hard cap (tests/agent-size-budget.test.cjs). Tightened the #4155 instruction's wording and dropped a redundant inline comment tag; no content lost. * chore(#4155): point changeset at the upstream PR number pr: 19 was the fork PR opened for internal review-lane CI; now that open-gsd/gsd-core#4290 exists, the changeset field must match it per CONTRIBUTING.md's release-notes convention. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
5ad9a36f35 |
fix(#4255): resolve reviewer-lane effort from the lane, not from gsd-plan-checker (#4275)
`review-lane plan` resolved every cross-AI reviewer lane's reasoning effort by spawning `query resolve-execution gsd-plan-checker --host <slug>`. The agent id was a hardcoded literal, so `--host` chose only the argv RENDERING while the LEVEL always came from the installed plan-checker's frontmatter — `low` under every shipped model profile. Every prompt-fed lane therefore ran at a fast structural verifier's effort, and because the rendered argument is a CLI config override it silently beat the effort the operator had configured for that CLI. At `low` a large source-grounded prompt makes a model end its turn with no final message, so the lane came back empty and its stub read as a crash. Effort is a property of the review, so the lane declares it. Two new fields on ReviewerLane — `effortConfigKey` (`review.effort.<slug>`) and `defaultEffort` — carried through each capability manifest and the generated registry, set on the three lanes with an argv effort channel and null on the other nine. A new pure `resolveLaneEffort()` resolves config key -> lane default -> nothing, where "nothing" emits no effort argument at all and the reviewer CLI's own configuration decides; `inherit` selects that path explicitly and an unrecognized level falls back to the lane default rather than being forwarded to a CLI that would reject it. The host's negotiated effortSurface still gates the rendering, so ADR-1239/#2481's trust boundary holds on this path too. Resolving in-process also removes up to twelve subprocess spawns per review. The empty-output stub now names the effort the lane ran at and distinguishes a clean exit from a timeout kill, a non-zero exit, and a process that never ran — `status` is null for both a timeout and a signal, so those were indistinguishable before. The hint is hedged: a clean empty exit is most often a model stopping short, but it is also consistent with a CLI writing its output elsewhere. Also: the capability validator now knows both fields, rejects a malformed key or an out-of-vocabulary default, and rejects a default declared without a config key (a level the operator could never override). An existing end-to-end row in tests/effort-surface-axis.test.cjs asserted the old coupling; it now configures the lane's own key and pins the decoupling in the same real spawn, with the agent execution tier set to a level that must not appear. Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort; leaving the new key undocumented there is the same invisibility that made the plan-checker coupling survive this long. Emitted-Drift-Ack-Growth: review.md — the effort/model resolution-order table this fix adds. The workflow is where an operator looks to find out which knob set a lane's model and effort, so leaving the new key undocumented there is the same invisibility that let the plan-checker coupling survive. Claude-Session: https://claude.ai/code/session_01CRMEuzNMWn3gs5uUW2ghcF Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
925a363879 |
enhance(#4032): apply configured agent tool grants (#4238)
* test(4032): add failing installed-agent grants contract Cover global and project agent_tools precedence at the real Claude installer seam before adding implementation. * feat(4032): apply configured agent tool grants during staging Resolve selector-level global and project config once per staging call, then append validated grants before runtime conversion. * test(4032): cover host grant and quoted MCP contracts Exercise installed host artifacts and prove ZCode must treat quoted MCP scalars like plain MCP grants. * feat(4032): apply configured agent tool grants across runtimes Move augmentation and scalar identity into the converter seam so every staged artifact preserves host policy. * fix(4032): register agent tool grants in configuration Accept documented agent_tools config without unknown-key warnings.\n\nKeep installer fixtures on the shared temporary-directory helper. * fix(4032): translate configured MCP grants for Kilo Reuse the converter-owned scalar decoder so quoted canonical grants reach Kilo's native permission keys without altering other host policies. * fix(4032): decode YAML-escaped tool grants * fix(4032): emit valid inline agent tool grants * fix(4032): reject invalid trailing-colon grants * test(#4032): cover cross-review remediation gaps * fix(#4032): close cross-runtime grant gaps * test(#4032): expose Kimi global project context * fix(#4032): preserve Kimi project config context * chore(#4032): add release note * test(#4032): expose fork review regressions * fix(#4032): address fork review findings * test(#4032): make byte-stability assertion portable Compare repeat installs at one root so platform-specific path rendering cannot masquerade as an agent_tools behavior change. * chore(#4032): bind changeset to upstream PR 4238 * fix(#4032): address trek-e review findings (2,3,4,5,6,7,8) Fixes fail-closed decode-failure handling in ZCode's mcp__ stripper, a comment-only `tools:` header mis-parse that silently dropped configured grants, and a naive comma-split that could tear a quoted scalar containing a literal comma. Documents Kilo's inherent `{server}_{tool}` MCP-permission-key collision (external, fixed format — not ours to widen) and locks the existing first-seen-wins resolution in with a regression test. Opts kimi/kimi-code out of the ADR-1235 pre-converter path-rewrite step: routing Kimi through that pipeline (needed so project-scoped agent_tools selectors reach it) was short-circuiting Kimi's own neutralizeKimiAgentPrompt, which expects the original ~/.claude/gsd-core text rather than a pre-rewritten Kimi path. Extends the fast-check token pool and per-runtime install coverage with the missing comment/comma/broad-runtime cases the prior review flagged as untested. * docs(#4032): add CONTEXT.md glossary entries for agent_tools resolver + pre-converter step Documents readGsdEffectiveAgentTools (Install Model Override Resolver Module) and the appendAgentTools pre-converter pipeline step (Runtime Artifact Conversion Module), per contributor-standards.md's new-seam glossary requirement (finding 1). * fix(#4032): address agy adversarial review findings An agy (gemini-3.8-flash-high) adversarial pass over the prior review-fix commit found the fixes for findings 3, 4, 6 and 8 had unfixed sibling gaps, plus a genuine new regression and two CONTEXT.md inaccuracies: - ZCode's comment-only `tools: # note` header matched the inline-value branch instead of falling through to the block-list scan, so a following mcp__* item leaked through unstripped — the exact defect finding 4 fixed in appendAgentTools, unfixed in this sibling function. - Reverted capabilities/kimi-code/capability.json's noPathRewrite: true. kimi-code uses the standard 'agents' kind with converter: null (not kimi-agents — confirmed by reading the descriptor, not its prose description), so it never went through the pipeline change finding 5 fixed, and disabling its path rewrite broke every ~/.claude/ embed in its shipped agents instead. - decodeToolScalar never stripped a trailing ` # comment` from a bare (unquoted) scalar, so a comment after a block-list item, or after an appended grant on an inline line, became part of the "tool name" — fixed at the source (one call site fixes every consumer). - appendAgentTools's comment-index scan wasn't quote-aware, so a `#` inside a quoted scalar (`"mcp__server #1"`) was mistaken for a comment start and corrupted the quote. - parseFrontmatterTools (Kimi/Qwen's tool-list reader, downstream of appendAgentTools's own output) had the same naive comma-split and comment-only-header gaps as findings 4 and 6, unpatched. - The all-runtime smoke test's presence assertion was built on a guessed omit-list; empirically only 7 of 17 runtimes keep an arbitrary mcp__ grant recognizable, replaced with a verified allowlist. - CONTEXT.md claimed a `project:<agent>` selector prefix that does not exist (project override is a same-key merge across two config files) and mislabeled stageAgentsForRuntimeWithConverter's module. * fix(#4032): address full-PR review (Opus critical/ponytail + agy) A whole-PR pass (critical-code-reviewer + ponytail-review on Opus, plus a second agy full-source adversarial pass) surfaced defects the earlier finding-scoped passes couldn't reach: - appendAgentTools corrupted a `tools:` line whose ENTIRE value is a leading quoted scalar (`tools: "Read"` -> `tools: "Read", Write`, invalid YAML) — there is no safe line-surgical rewrite here, so it now refuses to touch that shape instead of emitting broken frontmatter. - decodeToolScalar's malformed-trailing-quote check ran BEFORE comment stripping, so a bare tool name with a quote inside its own trailing comment (`Bash # note: "internal"`) was wrongly rejected. Reordered. - findUnquotedCommentIndex (added in the prior remediation commit) was built on a wrong model of YAML: a `#` after whitespace starts a real comment in a plain scalar regardless of nearby quote characters — verified against the actual parser. The one case that DOES need protection (a leading quoted scalar) is now refused outright above, so the quote-tracking scan was dead weight solving a problem that no longer reaches it. Removed; reverted to the plain `[ \t]#` scan. - Kilo has a SEPARATE agent-frontmatter parser (convertClaudeToKiloFrontmatter, distinct from the buildKiloAgentPermissionBlock fixed earlier) with the same comment-only-header and naive-comma-split gaps as findings 4 and 6 — unfixed in both its src/ and bin/install.js copies. Fixed in both, exporting splitToolScalars for bin/install.js to reuse rather than reimplementing it. - Pipeline docstring in stageAgentsForRuntimeWithConverter still listed 5 steps, omitting appendAgentTools (now step 3 of 6). - docs/CONFIGURATION.md didn't state that a --global install still discovers agent_tools from the cwd's .planning/config.json (confirmed intentional and already covered by a dedicated test, not a bug). - Removed install-engine.cts's deps.cwd injection seam: zero callers or tests ever populated it. Two claims from this round were verified and rejected, not fixed: prototype pollution via a `__proto__` selector key (empirically confirmed `Object.prototype` is never touched — only reassigns the resolver's own local object's prototype, with no observable effect), and a `*` grant value crashing YAML parsing as an alias reference (empirically confirmed it parses as plain scalar text, no crash). A pre-existing, unrelated defect (extractFrontmatterField returns null for block-list `tools:` on Copilot/Antigravity/Cursor/Codex/Qwen, affecting two shipped agents today) was filed as a follow-up rather than fixed here — it predates #4032 and isn't caused or worsened by this PR. * fix(#4032): update stale slug-derivation-drift-guard fixture line normalizeKimiSkillName's real closing brace moved from line 616 to 635 as a side effect of this PR's edits to runtime-artifact-conversion.cts; the MAJOR-1 fixture's hardcoded realEndLine had gone stale. * fix(#4032): address CodeRabbit findings on projectDir threading and flow-sequence tools bin/install.js's installAgentsKindStandalone call site omitted the projectDir argument the function already supports, so a global install through this legacy branch silently fell back to the runtime config dir instead of process.cwd() when resolving project-scoped agent_tools grants — inconsistent with the sibling installOpencodeFamilyArtifacts call site, which already threads it correctly. appendAgentTools' leading-quoted-scalar bailout did not cover a YAML flow sequence (`tools: [Bash, Read]`): splitToolScalars tore it apart on the in-sequence commas and appended past its closing bracket, producing invalid frontmatter. Extended the bailout regex to also refuse a value starting with `[`, matching the same "whole node, nothing may follow" reasoning already applied to quoted scalars. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
e8800287d5 |
enhance(#4153): fail closed unresolved update targets (#4237)
* test(#4153): cover unresolved update target * fix(#4153): fail closed unresolved update target * test(#4153): require a concrete recovery installer * fix(#4153): use concrete unresolved recovery command * chore(#4153): bind changeset to fork PR * test(#4153): cover portable update diagnostics * fix(#4153): keep update diagnostics portable * fix(#4153): harden update version diagnostics * test(#4153): reject jq in update version checks * test(#4153): expose step-local parser gap * fix(#4153): keep JSON parsing step-local * docs(#4153): align update target guidance * test(#4153): expose workflow runtime fallback * test(#4153): expose resolver runtime fallback * fix(#4153): leave unknown workflow runtime empty * fix(#4153): stop inferring Claude for unknown targets * test(#4153): preserve Claude workflow targeting * test(#4153): preserve known runtime directory identity * fix(#4153): recognize Claude workflow paths * fix(#4153): reuse known runtime directory identities * chore(#4153): acknowledge emitted workflow growth The fail-closed diagnostic and known-runtime preservation deliberately add 48 emitted bytes. Emitted-Drift-Ack-Growth: update.md — explicit unresolved-target diagnostics and known-runtime preservation * test(#4153): expose missing Windsurf workflow contract * docs(#4153): document Windsurf update targets * chore(#4153): bind changeset to upstream PR * fix(#4153): gate unresolved-target exit before the VERSION-missing fallback The VERSION-missing bullet in get_installed_version sat before the UPDATE_TARGET_UNRESOLVED exit and shared its trigger condition (version 0.0.0). An LLM agent reading the workflow top-to-bottom could satisfy "proceed to install" without ever reaching the fail-closed exit this PR adds, reopening the ill-defined mutating path #4153 closes. Reorder so the unresolved-target gate runs first and scope the VERSION-missing bullet to require an already-resolved target. Also drop two vacuous mutationSpies entries: they checked '--sync'/ '--reapply' (commands/gsd/update.md content) against `step`, a slice of workflows/update.md — always -1 regardless of correctness. Those routes bypass get_installed_version entirely and are already covered by install.test.cjs, reapply-patches.test.cjs, and skill-frontmatter-contract.test.cjs. * chore(#4153): point changeset pr field at fork PR #10 for fork CI * test(#4153): guard RUNTIME_DIRS/update.md table parity, confirm narrowing intent Nit 1: update.md's PREFERRED_RUNTIME prose and RUNTIME_DIRS (src/update-context.cts) are two independently maintained copies of the same runtime->dir mapping with no parity check; add one so a future edit to either surface without the other fails loudly instead of silently drifting. Nit 2: call out in the changeset that a custom --config-dir matching no known runtime, marker file, or env var now resolves unresolved instead of silently defaulting to claude -- this narrowing is intentional, it's the fail-closed behavior #4153 asks for. * fix(#4153): drop dead $UC fallback in check_latest_version's uc_field, cover unresolved-runtime fast path agy (gemini-3.8-flash-high) adversarial review of the full PR: 1. check_latest_version's uc_field() copy-pasted get_installed_version's `${2:-$UC}` fallback, but every call site here passes $2 explicitly and $UC does not exist in this step's scope -- dead, misleading reference. Use $2 directly. 2. No unit test covered resolveUpdateContext's preferredConfigDir fast path returning runtime: '' for a custom --config-dir matching no RUNTIME_DIRS suffix, marker file, or env var (the exact fail-closed case #4153 adds). Added. A third finding (update.md:90 using /gsd:update vs docs using /gsd-update) was investigated and rejected: /gsd:update is the actual registered Claude Code command name (commands/gsd/update.md name: gsd:update) and is locked by this PR's own test (tests/update-workflow.test.cjs); /gsd-update is a separate, pre-existing, intentional prose convention used in audience-facing docs (README/INVENTORY/FEATURES). Not a defect. * chore(#4153): backfill changeset pr field to upstream PR #4237 --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
d0d542e478 |
fix(#4183): resolve root phase Fallow base (#4215)
* test(#4183): reproduce root phase Fallow base failure * fix(#4183): resolve root phase Fallow base Fall back to the phase root commit when its parent is unresolvable. * chore(#4183): add release note * chore(#4183): bind changeset to fork PR * docs(#4183): describe root fallback precisely * chore(#4183): bind changeset to upstream PR --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
eedb6b5431 |
enhance(#4107): sequence external review after internal fixes (#4206)
* enhance(#4107): sequence external review after internal fixes Teach the planner to finish internal review and accepted fixes before opening a PR known to trigger automatic external review. If an open-time property exists, re-check it immediately before opening with nothing intervening; post-open CI, review, changeset, and tracking work may follow. Emitted-Drift-Ack-Growth: gsd-planner.md — issue #4107 adds the review-before-publish ordering rule * chore(#4107): add PR #11 changeset * chore(changeset): link upstream PR 4206 * fix(#4107): ground external-review terms and tighten ordering test Addresses trek-e review on PR #4206: - Ground 'known automatic external review' and 'open-time property' with concrete anchors (CodeRabbit App / .coderabbit.yaml, not-behind-base). - Suffix the antipatterns heading with (#4107), matching sibling sections. - Replace vacuous negative assertion with inverted-order fixtures that prove the ordering regexes reject bad phrasing, not just co-occurrence. * fix(#4107): make directionality fixtures genuinely adversarial agy (gemini-3.8-flash-high) adversarial review found the two negative fixtures added in 584ec1cda were vacuous: they proved the ordering regexes require certain keywords, not that they reject inverted order — the bad strings simply omitted required tokens rather than reordering them. - Rebuild both fixtures to contain every required token, reordered/negated, so a real reordering would still slip past a weaker regex. - Drop the unsupported 'changeset' mention from the Wave 4+ antipatterns example — gsd-core/workflows/ship.md never references changeset work, so naming it here implied a step this rule doesn't actually govern. * fix(#4107): make the full review-then-fix-then-open sequence explicit CodeRabbit (fork PR #11) flagged that the planner prose only ordered accepted fixes before PR open, without explicitly naming 'run internal review' as its own earlier step, and that no fixture tested the planner text's own wording for inversion (only the antipatterns example had one). - Prose now reads 'run internal review and apply the accepted internal-review fixes before the final open'. - Added a planner-text-specific inverted-order fixture alongside the existing antipatterns-example one. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a262ad6b61 |
fix(#4148): dispatch wave-pre step hooks (#4185)
* fix(#4148): dispatch wave-pre step hooks External capabilities can render step hooks before a wave, but the execute workflow consumed only contributions and silently skipped every step. Reuse the shared dispatch contract before executor spawning and pin the capability-validator boundary with a red-first regression. Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning * test(#4148): pin wave-pre dispatch ordering * test(#4148): pin wave-pre dispatch contract * chore(#4148): bind upstream changeset PR * chore(#4148): restore fork changeset identity * fix(#4148): align wave-pre dispatch contract Mirror the sibling wave-post all-shapes clarification while pruning redundant prose so the rebased workflow remains below its frozen byte ceiling. Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning * chore(#4148): restore upstream changeset identity * fix(#4148): align wave-pre capability guidance * docs(#4148): identify wave-pre manifest input Name the third-party manifest trust origin at the wave-pre dispatch boundary so the reviewer-requested validation guidance matches wave-post. Emitted-Drift-Ack-Growth: execute-phase.md — wave-pre now carries the missing generic step-dispatch contract before executor spawning * docs(#4148): preserve execute-phase byte budget Remove a redundant advisory label while retaining the non-blocking contract, keeping the reviewer-required trust-boundary wording at the enforced 93,400-byte ceiling. * fix(#4148): mark wave-pre manifest-input validation as security-relevant Reviewer nit on PR #4185: wave-pre's step-dispatch sentence had the (third-party manifest input) parenthetical but dropped the ⚠ marker that wave-post's parallel sentence (execute-phase.md:1044) carries, losing the visual flag that this validation is security-motivated. Trims the redundant "of one" from "not one shape of one" to reclaim the 4 bytes the marker adds — the ADR-857 byte-margin gate (tests/claude-orchestration.test.cjs) leaves zero slack at the 93,400-byte ceiling. * fix(#4148): trim wave-pre step-dispatch prose to clear ADR-857 byte ceiling Merging next's unrelated growth (#3990's TDD_APPLICABLE conditional) pushed execute-phase.md 116 bytes past the 93,400-byte ceiling, failing CI on all three platforms. The security-relevant ⚠ marker and ref.command validation call-out (added per prior reviewer nit) are preserved verbatim per the pinned regression test in capability-registry.test.cjs; only the non-pinned connective prose is trimmed. * fix(#4148): recalibrate execute-phase.md self-imposed margin, restore security marker next grew execute-phase.md by ~230 bytes across two unrelated merges during this fix (#3990's TDD_APPLICABLE conditional, then a further step-extraction commit), consuming this test's own self-imposed 93,400 safety buffer under ADR-857's actual, unmodified 93,600 ceiling (docs/adr/857-capability-system.md:22). The wave-pre step-dispatch sentence cannot shrink further without dropping one of the pinned substrings this same test file asserts on (kind=="step", loop-hook-dispatch, never blocks or redirects executor spawning, Validate `ref.command`). Raises the self-imposed margin to 93,550 (still 50 bytes under the real, untouched ADR ceiling) and restores the ⚠ marker the prior reviewer round required for the ref.command validation call-out, which byte pressure had dropped. * fix(#4148): restore full ref.command validation wording, drop self-imposed margin Adversarial review (agy/gemini-3.8-flash-high) flagged two issues in the prior CI-recovery commit: 1. Trimming "in-context before any shell use" from the step-dispatch warning weakened the inline operational instruction (the reader is told WHAT to validate but not the specific in-context-not-shell mechanism the referenced loop-hook-dispatch.md:45-51 threat model requires). Restored it - the merge with next since the last commit freed enough real margin (77 bytes under the untouched 93,600 ADR-857 ceiling) to afford it without any margin change. 2. The prior commit self-imposed margin bump (93400 to 93550) was, on reflection, the wrong lever: it is a number this PR invented, not an ADR value, and re-bumping it every time next grows execute-phase.md is a losing pattern (already needed twice in one session). Removed the redundant assertion; the same line existing bytes-under-93600 check against the real, frozen ADR-857 ceiling (docs/adr/857-capability-system.md:22) is the actual invariant and is untouched. workflow-size-budget.test.cjs tier hard cap (98304 bytes, extract-not-bump by design) remains the correct backstop for runaway growth. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
c27da5e993 |
fix(#4079): forbid ScheduleWakeup wake-tool escape hatch at background-wait sites (#4299)
* test(#4079): guard every background-wait site against ScheduleWakeup wake-tool escape hatch Content-contract tests over plan-phase.md, stall-detection-helpers.md, and manager.md: every wait-instruction site must forbid ScheduleWakeup (the host /loop tool whose partial-args call surfaces the red validation error in #4079), while the sanctioned wait mechanisms stay byte-preserved. * fix(#4079): forbid ScheduleWakeup wake-tool escape hatch at every background-wait site The orchestrator could literalize 'I'll wait' by calling the host's ScheduleWakeup tool (/loop pacing surface) with partial arguments while a background subagent was in flight, surfacing the red validation error '`prompt` is required when `stop` is not true.' GSD never sanctioned a wake call; now every wait-instruction site says so explicitly: the two synchronous stop-and-wait ORCHESTRATOR RULEs in plan-phase.md, the stall-detection-helpers.md fragment (loaded before every gsd_stall_watch wait), and both manager.md background-dispatch rules. The sanctioned wait mechanisms (blocking Agent() return, gsd_stall_watch polling, dashboard loop) are unchanged. The regression test uses the shared readFileNormalized helper (code-review finding). Emitted-Drift-Ack-Growth: plan-phase.md — #4079 guard sentence at the two stop-and-wait rules (+396 B, within the phase6 shrink-only line) Emitted-Drift-Ack-Growth: manager.md — #4079 guard sentence at both background-dispatch rules * chore(#4079): add changeset fragment (pr placeholder, backfilled after PR open) * chore(#4079): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
0fca71eaae |
enhance(#2529): cover every workflow with response-language directives + CI lint (#2558)
* enhance(#2529): cover every workflow with response-language directives + CI lint Every workflow now carries response-language coverage in one of three forms, and a CI lint keeps it that way. - 43 workflows load the new shared reference, `gsd-core/references/response-language-directive.md`, by eager `@`-import. - Lazy-loaded modes/steps/templates, which cannot rely on an eager import, carry an exact inline directive; 35 such paths are pinned by exact path. - Fragments dispatched by a covered parent inherit coverage, proven per file rather than granted per directory. The 45 workflows whose directive covered only "questions, prompts, and explanations" now name inter-tool narration, which is the defect #2529 reports: the running commentary between tool calls stayed English while the answers around it were translated. `scripts/lint-response-language-coverage.cjs` enforces it and fails closed on three independent discovery failures (unreadable catalog, empty catalog, unfollowed symlink). It resolves which reference a workflow imports and applies the same four-predicate test to that file, so a weakened shared reference uncovers its importers instead of passing silently, reported once as a systemic failure rather than 43 times. The walk follows symlinked subtrees with a realpath cycle bound. `lint:ci` invokes it by name. REQ-LANG-03 and REQ-LANG-04 state the contract in docs/FEATURES.md; REQ-LANG-04 names the two forms that satisfy it ("narration", "between tool calls") rather than enumerating class members an author cannot use verbatim, and a test pins that text to what the matcher accepts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): register the coverage test in the docs-guard lane `107eb8c1` (#3787) landed the docs-guard lane on `next` while this PR was open: a test that reads a `docs/` path must be named in `scripts/docs-guard-registry.cjs` or carry a `docs-guard-exempt` marker, so the guards that read a doc run on the PR that changes it. `tests/response-language-coverage.test.cjs` reads `docs/FEATURES.md` -- it extracts every form REQ-LANG-04 offers an author and runs each through the matcher that enforces it. Registration, not exemption, is the correct side of that gate: a reword of the requirement with no code change is precisely the diff this test exists to catch, and it is the diff the lane would otherwise skip. Registered narrowly (`['docs/FEATURES.md']`) rather than with the `'*'` sentinel, so an unrelated docs change does not pull this test into the lane. Verified: lint-docs-guard-registration 0 violations, tests/ci-docs-guard-registry.test.cjs 51/51, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): consolidate this PR's emitted-growth acks into its own fragment This PR ripples emitted bytes across 85 workflow paths. Until now each ripple was acknowledged by appending to whichever live fragment owned that path, because two ack sources may never name the same path. `a84f7563` (#3078) swept all 45 fully-spent fragments off `next`. Forty-two of the paths this PR grows were owned by swept fragments, so those keys are now unowned and this PR's own fragment declares them directly -- one path, one source, and no dependence on a fragment that no longer exists. Each adopted entry keeps its measurement and records where it came from. Two paths are handled differently, because the sweep did not free them: - `review.md` is now owned by `3034-parallel-reviewer-lanes.json`, which landed on `next` after the sweep. Its entry is live, so the old route still applies: this PR's note is appended to that entry rather than declared a second time. - `plan-review-convergence.md` keeps the arrangement made in round 24. Result: 3 fragments in the directory, 85 keys in this PR's own, 0 cross-source duplicates. `lint-emitted-drift-ack` exit 0, `tests/emitted-attribution.test.cjs` green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): move REQ-LANG-03/04 into the feature fragment that now generates them `36375513` (#3845) made docs/FEATURES.md a generated projection of docs/features/*.md, marked "do not edit by hand". This PR wrote REQ-LANG-03 and REQ-LANG-04 straight into the generated file, so the rebase left the requirement present in the projection and absent from its source -- the next regeneration would have deleted both, and `tests/features-index-gate.test.cjs` was already red on the mismatch. Both requirements now live in docs/features/response-language-config.md alongside REQ-LANG-01 and -02. Regenerating produces a docs/FEATURES.md that is byte-identical to the committed one, so the text this PR shipped is unchanged -- only its source of truth moved to where #3840 put it. The docs-guard registration is widened to name the fragment as well as the projection. The requirement's source is the fragment now, and an edit there that skips regeneration would otherwise reach this guard through neither path. Verified: features-index-gate 68/68, lint-docs-guard-registration 0 violations, ci-docs-guard-registry + response-language-coverage 142/142. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): hand the plan-phase ack back to its new live owner `c933184b` (#3825) landed `3172-stated-failing-direction.json` on `next` after fragment had adopted that path when the sweep left it unowned, so the merged tree named it from two sources -- a hard failure in `scripts/lint-emitted-drift-ack.cjs`. The path has a live owner again, so the append route applies: this PR's note joins that entry, carrying its own measurement, and the key is dropped from this PR's fragment (84 keys left, the others untouched). The provenance sentence written for the swept-fragment case is removed rather than reused -- this path was never orphaned, so that account of it would be false. Same shape as `review.md` and `plan-review-convergence.md`: ownership is a property of the merged tree, and a fragment landing upstream after a push can reclaim a key no local check would have flagged. Verified: lint-emitted-drift-ack exit 0, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): state byte figures that are true against the tree The reference claimed `execute-phase.md` has "2 bytes of headroom under the ceiling named below". That was true when the sentence was written -- the file sat at 93398 against the 93400 comfort assert -- and upstream has since shrunk it to 91493 against a 93600 hard ceiling, so the figure now understates the headroom by three orders of magnitude. The rationale the sentence supports does not depend on the number, so the number is gone rather than refreshed: a restated figure would go stale again on the next upstream edit, and nothing parses it. Audited every other numeric claim this PR ships the same way, mechanically against the merge base: all 82 FILE-delta claims in the ack fragment match the real per-file delta exactly, and the 1,629-byte reference and 63-byte import line check out. One class was imprecise: the 41 notes for workflows whose inline directive was rewritten in place quoted the conversion counterfactual as "+1,692 bytes more loaded context", which is the reference form's whole weight, not the increase over the inline directive those files already carry. Each now names both quantities and the net (+1,605 / +1,609 / +1,584). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): one rule for pinned vs inherited coverage, and the docs to pick it Review measured that 14 of the 35 pinned fragments would pass by inheritance anyway, and that the PR asserted both readings at once: inheritance is real coverage (so those 14 pins are noise) or it is not (so 30 inheriting fragments are green-but-uncovered). Only one can be true. Inheritance is real: the predicate proves it per file -- the parent must dispatch this exact path from a read/execute context AND be covered itself -- so the parent's directive is in the loaded context by the time the fragment is read. The 14 pins are therefore removed along with the directive lines they pinned, and those files inherit like the 30 structurally identical ones. The rule is now stated where the set is declared, and enforced from the other side by a test: no member of the pinned set may be one that would have inherited. That is what decides the form for the next fragment. - pinned set 35 -> 21; 14 workflow files revert to their base content - `findViolations` no longer returns early on a pinned path: a file that becomes eagerly loaded and takes the shared reference is strictly better off, and the gate must not red that. The reference form is admitted because its own wording is validated in turn; an arbitrary reworded inline line still fails. - the reference-directive cache is keyed by size and mtime, not by path alone, so a rewritten reference re-asked in one process no longer returns the stale verdict - `carriesInlineDirective` names its negation blindness: four independent hits read vocabulary, not polarity - the real-tree scan asserts each source produced files instead of `> 152`, a constant that read as the workflow count and would have passed a scan that lost one of its two directories - the pinned-set size assertion goes the same way: the size follows from the rule, so the rule is what the suite asserts Docs, for the gate that now governs every future workflow: - `docs/contributing/response-language-coverage.md` -- why the narration class is the discriminator, the four coverage forms, the decision order that picks one, the pinned line, and what each failure message means - a row in CONTRIBUTING.md's CI checks table, matching the docs-guard row - `docs/CONFIGURATION.md` points at it from the `response_language` entry Also: the changeset said 45 reworded workflows; it is 44 (42 @-reference + 21 pinned + 44 rewritten = 107 touched). That text ships to CHANGELOG.md. `3707-parse-gap-reporting.json` landed on `next` reclaiming `audit-uat.md` and `progress.md`; both handed back by the append route, leaving 82 keys here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): correct the reference-taker count, 43 -> 42 The ack notes said the import line is byte-identical "in each of the 43 workflows that take the reference" and that the alternative would be "43 inline copies". The shared reference has 42 importers; the 43rd file in review's table is `execute-phase.md`, which imports the OTHER reference. Corrected in all 41 notes that carry the sentence, across this PR's fragment and the two it appends to. Found by re-running the numeric audit from the previous round after the rebase, which also re-verified all 84 FILE-delta claims against the new base -- all exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): migrate the emitted-drift ack from a fragment to commit trailers ADR-3942 (#3954) landed while this PR was open: the acknowledgment is now a commit trailer and tests/emitted-drift-acks/ no longer exists. The fragment is deleted and each key it declared becomes one trailer, reasons unchanged. The four keys this PR had handed to 3034-*, 3172-* and 3707-* under the one-source rule come home here. That rule was the whole reason for the hand-backs, and the trailer model has no shared namespace to collide in -- five of this PR's rounds were spent on exactly those collisions. Emitted-Drift-Ack-Growth: add-backlog.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-tests.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: add-todo.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: ai-integration-phase.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by |
||
|
|
5214ad5802 |
fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication (#4295)
* fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication execute-plan.md and gsd-executor.md each cited "Red-Green-Refactor Cycle" for three facts (commit-scope contract, fail-fast rule, error handling), but only the commit-scope contract lives there. Fail-fast is in tdd.md's "Fail-Fast Rules" subsection (under "Gate Enforcement Rules") and error handling is in tdd.md's "Error Handling" section — cite each correctly. gsd-executor.md's "Plan-Level TDD Gate Enforcement" section also fully restated the gate-sequence rules tdd.md's "Gate Enforcement Rules" already owns (and covers more thoroughly, including the actual git-log validation script). Collapse it to a short pointer, matching the treatment already used by the cycle-steps pointer immediately above it. Adds tests/tdd-reference-correctness.test.cjs asserting the pointer text cites the correct section names, that those sections actually carry the guidance, and that the old gate-sequence restatement is gone from gsd-executor.md. Closes #4267 Closes #4269 * docs(#4267): add changeset for tdd.md pointer correctness fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4267): acknowledge execute-plan.md growth execute-plan.md grew 95 bytes (39766 -> 39861) from the corrected three-section citation in the #3990/#4267 cycle-steps pointer. Emitted-Drift-Ack-Growth: execute-plan.md — net +95 bytes from citing the "Fail-Fast Rules" and "Error Handling" sections by name instead of a single mis-scoped section (#4267). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4267): update PROSE_ALLOWLIST line numbers shifted by the pointer-citation edit This branch's edits to agents/gsd-executor.md and gsd-core/workflows/execute-plan.md shifted line numbers, leaving tests/no-bare-gsd-tools-command-position.test.cjs's PROSE_ALLOWLIST pointing at stale lines. Update both entries to their new correct lines (811 and 419 respectively) without changing the underlying prose. * fix(#4267): restore INVALID_RED citation, fix allowlist line shift after #3770 rebase The rebase onto next picked up #3770's already-merged fail-fast update to gsd-executor.md's plan-level gate section, which this branch's own commit collapses into a pointer. The conflict resolution kept the pointer but dropped the literal "INVALID_RED" term that tests/tdd-red-evidence.test.cjs requires gsd-executor.md to name — restored it. Also updates no-bare-gsd-tools-command-position.test.cjs's PROSE_ALLOWLIST line number for gsd-executor.md, shifted again by the rebase. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4267): fix non-matching regression-guard regex in tdd-reference-correctness The fail-fast regression guard asserted gsd-executor.md no longer contains "If a test passes unexpectedly during the RED phase" — but the actual old prose (removed by this branch's pointer-collapse) read "If a test passes unexpectedly during RED, STOP". The regex never matched the real old text, so the assertion would have passed even against the unmodified pre-change file. Caught by an isolated orthogonal review pass. Fixed to match the actual removed wording, and confirmed (via a direct grep) it is genuinely absent from the current file. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4267): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d29b50d696 |
fix(#4051): route specific intents first and confirm before dispatch in --do (#4289)
* test(#4051): pin freeform routing specificity contract in do.md * fix(#4051): order freeform routing specific-first, confirm before dispatch, argument-aware forwarding * fix(#4051): regenerate FEATURES.md, satisfy docs-guard on new routing test Emitted-Drift-Ack-Growth: do.md — deliberate growth: specific-first routing table (code-review, plan review, ui-review, secure-phase, audit, docs-update, phase CRUD rows), a REQ-DO-03 confirm step, and argument-hint-aware dispatch. * chore(#4051): fold regression into non-bug-prefixed test filename per lint-regression-test-names * fix(#4051): review fixes — em-dash description style, split audit-fix route * chore(#4051): sync skill mirrors of execute-phase/phase descriptions * chore(#4051): add changeset (pr backfill to follow) * chore(#4051): backfill PR 4289 in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
8249ebcf6e |
fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do not exist yet; every row fails on require. Per #3770 only an intentional target-test failure may authorize GREEN; zero-test discovery, fixture crashes, unrelated failures, and unexpected green are INVALID_RED. * fix(3770): require intentional RED evidence before GREEN Only an intentional failure of the TARGET test (distinctly named, TAP-reported assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test discovery, fixture/load crashes (file-named failures), nonzero exits without a failing test, unrelated failures, unexpected greens, and malformed/missing records are INVALID_RED and block GREEN. - src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses the prohibition-enforcement TAP primitives; fail-closed, never throws) - check tdd-red-evidence <record.json>: validates the persisted record (command, exit code, failing test, expected, actual) - gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now requires the evidence record + gate verdict, not a nonzero exit or a RED: tag * chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs * fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib - gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B < 49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid) - tests: the row-6 fixture used String.replace (first-occurrence), so the `not ok` line still named the target test and the classifier was right to accept it; replaceAll makes the failure genuinely unrelated - eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs (lint the src/*.cts source, per ADR-457 migration rule) Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line * chore(3770): add changeset * chore(3770): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
2f4f7538e9 | fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) | ||
|
|
580059251a |
fix(#4040): route partially-created .planning to initialization recovery (#4283)
* test(#4040): add failing-first regression tests for partial-init routing Red: init.progress/init.resume/init.new-project payloads carry no partial-init discriminator, and progress.md/resume-project.md/ new-project.md route an interrupted bootstrap (.planning/PROJECT.md + config.json only) to Route F / STATE reconstruction / a hard error. * fix(#4040): route partially-created .planning to initialization recovery A bootstrap interrupted after .planning/PROJECT.md (but before REQUIREMENTS.md/ROADMAP.md/STATE.md) was mis-routed three ways: progress.md read it as between-milestones (Route F) or 'no planning structure', resume-project.md offered STATE.md reconstruction, and new-project.md errored 'already initialized' — a routing loop with no recovery exit. Add a shared buildInitCompletenessFields discriminator (planning_exists / requirements_exists / milestones_exists / init_incomplete) to the init.progress, init.resume and init.new-project payloads, and branch on init_incomplete in progress.md, resume-project.md and new-project.md BEFORE the legacy branches. MILESTONES.md presence excludes the archival between-milestones state, so Route F and the STATE-reconstruction path keep working. Emitted-Drift-Ack-Growth: progress.md — deliberate #4040 growth: new init_incomplete recovery branch (routing text + guard on the no-planning and Route F branches) added ahead of the legacy init_context routes. Emitted-Drift-Ack-Growth: resume-project.md — deliberate #4040 growth: new init_incomplete branch routing an interrupted bootstrap to initialization recovery before the STATE.md-reconstruction branch. Emitted-Drift-Ack-Growth: new-project.md — deliberate #4040 growth: project_exists gate split on init_incomplete so a partial bootstrap resumes initialization instead of erroring. * chore(#4040): add changeset fragment * chore(#4040): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
97ce61dee2 |
fix(#3990): state the RED/GREEN/REFACTOR cycle once, embed tdd.md conditionally (#4228)
* test(#3990): the RED/GREEN/REFACTOR cycle is stated once, embedded conditionally * fix(#3990): state the cycle once — pointers in consumers, conditional tdd.md embeds Emitted-Drift-Ack-Growth: execute-phase.md — #3990 conditions the tdd.md embed on the dispatch being TDD * chore(#3990): changeset for the single-statement TDD cycle * chore(#3990): backfill changeset pr number * fix(#4228): linear cycle check — the lazy-span regex pinned a Windows core for the whole job cap * test(#3990): allowlist pin tracks the rebased line * fix(tests): npm-integrity gate names an empty audit output explicitly — empty stdout crashed the parse as a bare SyntaxError * fix: name an empty npm-audit stdout explicitly — it crashed the parse as a bare SyntaxError Observed on CI (several branches, all lanes): spawnSync npm ETIMEDOUT with empty stdout; the empty string survived the recovery path and surfaced as 'SyntaxError: Unexpected end of JSON input', hiding the captured error. The recovery path now requires non-empty stdout, and an empty result throws with the captured stdout/stderr/message so the actual error is on the record. Root cause of the ETIMEDOUT itself is NOT diagnosed here — this change only stops masking it. --------- Co-authored-by: sim <sim@local> |
||
|
|
515191f07d |
feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only) * test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED) Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/ merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and confirms the prior research pass's Open Question 1: a coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) leaves BATCH.json at "pending" with no STATE.md row yet (only written in Step 9), so --resume's eligibility re-derivation would dispatch a second executor into a new worktree for the same item, orphaning the first. This test asserts worktree-dispatch.md's Step 6 excludes an item whose SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's existing PLAN.md-existence check one layer earlier. Fails against the current worktree-dispatch.md, which has no such guard. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the full trace and fix-location rationale. * fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN) worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round via the same quick-batch resume call resume-mode.md uses, but had no check for "did this item already finish executing" the way planner-wave.md already checks "did this item already get planned" (PLAN.md existence) before re-planning. A coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) left the item eligible for a second dispatch on --resume, orphaning the first worktree's real, already- committed work and silently losing it once the second executor's SUMMARY.md write clobbered the first at the same item_dir path. Adds a SUMMARY.md-existence exclusion before spawn-plan is computed, symmetric to planner-wave.md's PLAN.md check. The excluded item is not lost: merge-wave.md's own mergeable-wave criterion (status=pending, SUMMARY.md on disk, not yet merged) already picks it up independently of this eligible/spawn list. Workflow-prose-only fix — touches no already-merged/reviewed .cts module. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the fix-location rationale (why not resumeBatch itself). * test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677, epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership attempts", "scope drift", "submodules"): - Arbitrary-worktree ownership tampering: a manifest entry naming a non-agent branch is silently dropped at normalization before any git subprocess runs; a manifest entry naming a plausible agent-branch that was never actually created by this repo's own worktree.create (a genuinely foreign repo/branch) is blocked via base_mismatch. Both leave the foreign location and repoRoot's HEAD provably untouched. - Advisory scope drift: a committed path outside declared files_modified still merges successfully (advisory, never blocking) while surfacing a scope_out_of_declared warning naming the drifted path; an exact declared-scope match produces zero warnings (boundary case). - Real .gitmodules submodule integration: a repo containing a real local git submodule merges cleanly through executeWorktreeWaveCleanupPlan for an unrelated plan; a real gitlink pointer bump (declared) merges cleanly with the superproject tree reflecting the new pinned commit; an undeclared bump is advisory-only and surfaces a scope warning naming vendor/sub, same as any other undeclared modification. No src/*.cts changes — all three gaps were coverage-only; the underlying primitives already behaved correctly (independently verified against real git subprocess output before writing each assertion). * docs(#3677): document how to diagnose a preserved quick-batch worktree Extends the one-sentence "worktree is preserved (never deleted)" mention into a concrete diagnosis procedure: where the preserved directory is, how to read the executor's real commits/diff against the plan's declared files_modified, how to read the item's own SUMMARY.md independent of merge outcome, how to manually merge-and-clean-up or discard, and how to re-run --resume afterward. Also documents that a SUMMARY.md-written-but-still- pending item (the crash-window case fixed in this same PR) needs no manual intervention — --resume routes it straight to the merge step. * chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only) * fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable Orthogonal review (Spec finding): the crash-window regression test added earlier this phase only asserted readStep('worktree-dispatch.md') + regex matches against the markdown prose — proving the DOCUMENTATION says the right thing, never that the runtime condition (pending status + on-disk SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own "Alternatives considered" explicitly rejects "document recovery without fault injection" for exactly this reason. Extracts the filtering decision into a pure, independently testable function, filterAlreadyExecuted(eligibleIds, executedIds) in src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed` CLI verb (src/quick-batch-command-router.cts) — the same pure-decision- then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already establish. worktree-dispatch.md now calls this verb explicitly instead of only describing the decision in prose. A genuine fixture-based test in tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch), writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the REAL resumeBatch, and proves both that resumeBatch alone still reports the item eligible AND that filterAlreadyExecuted (fed a real filesystem check) correctly excludes it. The prior prose-assertion tests are kept — they now prove the workflow markdown is correctly WIRED to the verb — but are no longer the only proof. Self-discovered defect while building that fixture (fixed inline, not deferred): tracing merge-wave.md against /gsd:quick's own prior art (QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed $QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed coordinator correctly does not re-dispatch an already-executed item (this fix), but nothing durably recorded that item's worktree_path/branch/base either — Step 7 in the resumed process would have had no data to build its cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/ dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT a reuse of the pre-existing `worktree` field, whose loadBatch validation requires the path to exist on disk (verified empirically: reusing it made the batch permanently unloadable the moment a legitimately-merged worktree was removed). worktree-dispatch.md persists the triple once a worktree is created; merge-wave.md falls back to it when the ephemeral manifest lacks an entry, clears it after a successful merge, and fails closed rather than guessing if no record exists anywhere. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1 and §9.3 for the full trace, empirical verification notes, and rejected alternatives (reusing `worktree` directly). * test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees Orthogonal review (Security finding): the two existing ownership-tampering tests didn't test ownership — one was trivially rejected by WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch- NAME filtering, not ownership), the other pointed at a wholly separate, never-linked foreign repo, so merge-base failed immediately because the branch didn't exist as a ref at all. Neither exercised the real scenario: a manifest entry whose worktree_path/branch are swapped to point at a DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot, with a branch name passing the shape check and a base in allowed_bases. Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts) directly: this is NOT a reachable gap. Git enforces branch-per-worktree uniqueness, so a swapped-in entry.branch can only match worktree_path's ACTUAL checked-out branch if it names that sibling's own real, uniquely- generated branch name — which manifest tampering confined to one batch's own record has no way to know (branch names are agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is collision-checked GLOBALLY across every existing quick task and batch, not merely within one batch). Adds a stronger test that empirically proves this: two REAL, concurrently- alive sibling worktrees of the same repo (both via real `git worktree add`, both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with worktree_path/branch swapped between them in both directions. Both attempts are blocked via branch_mismatch; both real worktrees, their branches, and one sibling's real uncommitted-to-main commit survive completely untouched. Supplements (does not replace) the original two tests, which still prove distinct, real boundaries. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2 for the full trace, including the one explicitly-documented (not fixed) trust boundary this investigation surfaced: the primitive defends against fabricated data, not a caller bug that misattributes a real-but-wrong item's own triple to a different item. * chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only) * docs(#3677): add changeset for PR 4240 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3639ab0431 |
fix(#3968): measure commit claims at all three surfaces — ledger, verifier BLOCKER, porcelain HANDOFF (#4230)
* test(#3968): commit claims must be measured against git, never narrated * fix(#3968): measure commit claims at all three surfaces Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument Emitted-Drift-Ack-Growth: pause-work.md — #3968 uncommitted_files from git status --porcelain * fix(#3968): retired slash syntax, allowlist line pin, git-compare test pin * fix(#3968): persist the ledger on disk and reconcile with the same rev-list instrument Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument * fix(#3968): hold the gsd-executor size cap with a compact ledger contract Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) * fix(#3968): allowlist pin and HALT regex track the final prose * chore(#3968): changeset for measured commit claims * chore(#3968): backfill changeset pr number * fix(#3968): quote the BASE expansion (SC2086) * ci: raise the test-lane budget 21 to 32 minutes (measured cost grew past the cap) --------- Co-authored-by: sim <sim@local> |
||
|
|
d5f8191f66 |
fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick (#4216)
* test(#3730): a legacy Quick Tasks table must be migratable to canonical * fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick * fix(#3730): review fixes — usage parity, contiguous table span, collision-safe bucket, template-width delimiter Emitted-Drift-Ack-Growth: fast.md — #3730 runs quick-tasks-migrate before the first append (auto-migration on first quick run) Emitted-Drift-Ack-Growth: quick.md — #3730 replaces the match-any-format note with the migration instruction * chore(#3730): backfill changeset pr number * fix(#3730): scope the quick-batch row-48 guard to branches touching quick-batch --------- Co-authored-by: sim <sim@local> |
||
|
|
2f64e6230a |
feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local> |
||
|
|
7c52344284 |
fix(#4003): anchor the safe-resume gate's plan-scope greps to the milestone (#4194)
* test(#4003): safe_resume_gate must grep an anchored padding-tolerant scope * fix(#4003): anchor the resume-gate scope greps and bound them to the milestone tag Emitted-Drift-Ack-Growth: execute-phase.md — #4003 rewrites three commit-scope greps (safe_resume_gate, TDD RED, completion spot-check) to anchored zero-pad-tolerant regexes with a milestone tag bound; growth is the fix itself * test(#4003): align shape assertions with the implemented gate text * fix: bump fast-uri past GHSA-jqff-g426-hqxp (transitive, advisory reddened next) * fix(#4003): bound the TDD RED grep to the milestone and fix tdd.md's example greps * test(#4003): the gate pin tracks the anchored scope grep * fix(#4003): trim the gate rationale to hold the 93400 margin ceiling * test(#4003): the RED-grep pin tracks the milestone-bounded invocation * chore(#4003): changeset for the anchored resume-gate scope * chore(#4003): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
acb903c2e8 |
enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable Add `workflow.code_review_point` (`execute:post` default, or `execute:wave:post`) so a multi-wave phase can run code review once per wave instead of once at the end, scoped to what changed since the phase's prior review. The code-review capability now declares its step at both loop points via a new generic `pointFrom` step field: `pointFrom` names an enum config key, and the step is only active at its own `point` when that key resolves to a matching value. `_resolvePointGate` (capability-activation.cts) is the single shared implementation consumed identically by loop-resolver.cts and capability-state.cts, and capability-validator.cjs enforces that `pointFrom` references an enum key whose values cover the declaring step's own point. code-review.md's manual-invocation gate now reads `workflow.code_review` directly instead of probing registry presence at the hardcoded execute:post point (so manual `/gsd-code-review` keeps working regardless of which automatic point is configured), and its file-scope tiers narrow to what changed since the phase's last review commit when one exists. execute-phase.md's wave-post step dispatch gets a small, precedented carve-out so the code-review skill still receives its required phase argument when dispatched generically (caught by the isolated spec review). Closes #3661 Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers. Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill. * docs: backfill changeset PR number for #3661 (#4159) * fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd Five fault-injection mocks in the "bug #1008" describe blocks intercepted every fs.writeSync call regardless of file descriptor, and several threw or truncated unconditionally on the first call. This surfaced as an intermittent macOS CI failure: node:test's own IPC channel back to the parent process (which also goes through fs.writeSync internally) could get a bogus injected error or truncated write if node's internal machinery called it while one of these mocks was active, corrupting the message frame the parent tried to deserialize ("Unable to deserialize cloned data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file IPC crash, not a test assertion failure). Root cause confirmed by a working counter-example already in the same file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection and were never implicated. Applied the same fd-scoped pattern to the five unscoped mocks (four output()-targeting tests gate on fd 1, one error()-targeting test gates on fd 2), and added a regression test proving an unrelated fd passes through untouched while the fault-injection mock is active. Found while verifying #3661; unrelated to that change's own diff. --------- Co-authored-by: sim <sim@local> |
||
|
|
5c7243e54b |
fix(#3995): derive the review diff base from the phase directory (#4181)
* test(#3995): diff base keys on the phase directory, not commit subjects All three derivation sites (Tier 3, spawn_reviewer, fallow pre-pass) must anchor on the phase directory's first commit; the milestone-blind repro (an archived milestone's same-numbered phase commit capturing the base) is the failing-first row. #3191/#3503 rows reworked to the directory-anchor contract; T6 docs-parity forbids any remaining phase-scope message-grep site. * fix(#3995): derive the review diff base from the phase directory A phase number is unique within a milestone, not a repository; the message grep had no milestone bound and tail -1 deliberately selected the oldest same-numbered subject, dragging archived milestones phases into the scope (7 files to 3388 plus the >50 depth downgrade). All three lockstep sites now anchor on the first commit that added anything under the phase own directory — the same anchor class git-base-branch phaseStartCommit uses. ShellCheck baseline gains the escaped fragment shifted parse signature. Emitted-Drift-Ack-Growth: code-review.md — phase-directory anchor replaces the message-grep derivation at both sites (#3995) * chore(#3995): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
647365faf1 |
fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as a gate condition, the end-of-phase escalation must not require MVP, the executor agent's gate section triggers on TDD_MODE alone, and the gate semantics reference loads without MVP_MODE. * fix(#4011): key the TDD runtime gate on TDD_MODE alone The RED-commit gate shipped as #76's MVP slice kept the paired invocation's conjunct, so workflow.tdd_mode=true was silently inert on every non-MVP phase, contradicting references/tdd.md's own contract. Drops the MVP conjunct from the per-task gate and the end-of-phase review escalation; rescopes execute-mvp-tdd.md's load condition, gsd-executor's gate section, and mvp-concepts' intersection claim. MVP remains free to imply TDD; the file is not renamed (stated assumption in the PR body). * test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing Review follow-ups: the detector now only inspects if/[ condition lines so explanatory prose mentioning both flags cannot trip it; remaining 'under/outside MVP+TDD' phrases in execute-phase.md, the gate reference, and docs/INVENTORY.md now describe TDD-mode semantics. Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011) Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011) * chore(#4011): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
7960374d15 |
fix(#3962): rename the TDD-Audit trailer token to gate-status (#4174)
* test(#3962): shipped TDD-Audit trailer token must round-trip through git Behavioral coverage: extract the trailers:key token from ship.md and prove a real git commit carrying that trailer reads back via %(trailers:key=token,valueonly). gate_status contains an underscore, which git's trailer machinery cannot tokenize, so the audit read was structurally empty. * fix(#3962): rename the TDD-Audit trailer token to gate-status Underscore is not a valid git trailer token character, so %(trailers:key=gate_status,...) could never match. Renames the token at the read, the documented aggregate write, and the section's prose/table header. Self-suppression semantics (#2431) unchanged. * test(#3962): compare outcome against the seam's literal, fix doc token Review follow-ups: the round-trip test compared outcome to 'EXITED' but the process seam emits 'exited'; and the trailer rename is propagated to docs/ship-pr-body-sections.md's read/write examples. * chore(#3962): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
10593e2cf2 |
fix(#3959): give review-prompt plans citable path anchors (#4170)
* test(#3959): plan copies carry source names, prompt carries path anchors Regression coverage: budget copies named gsd-review-plan-<plan-id>.md (not a bare padded index), no bare-index copy survives, and the Plans to Review template instructs a per-plan repo-relative #### path header. Also corrects the stale -00 assertion to the source-named copy. * fix(#3959): give review-prompt plans citable path anchors The Plans to Review template now instructs a per-plan repo-relative #### path header, and the budget copies are named gsd-review-plan-<plan-id>.md instead of a bare padded index, restoring provenance for prompt-budget's per-plan headers while keeping the trim glob intact. Emitted-Drift-Ack-Growth: review.md — per-plan path-header instruction in Plans to Review + provenance-named copy loop (#3959) * chore(#3959): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
ff0361071d |
feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane (#4160)
* feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane cursor-agent exposes --model (204 selectable models) but the cursor lane declared modelArg: null / modelConfigKey: null, so review.models.cursor was rejected as an unknown config key and the #1517 reviewer-instances escape hatch silently discarded a configured model at modelExpansion. Wire the lane the same way codex already is: inject {{model}} into args right after -p, set modelArg to --model, and declare modelConfigKey as review.models.cursor plus its config schema entry. An unconfigured lane still invokes byte-identically to today. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3653): add changeset for review.models.cursor Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3653): update co-change surfaces that assumed cursor has no model key gsd-test surfaced three surfaces still hardcoding "cursor declares no modelConfigKey", broken by wiring review.models.cursor: - tests/reviewer-config-federation.test.cjs: the #3691-narrows-#2797 invariant test listed cursor among lanes that must own no model key. - tests/settings-integrations.test.cjs: the #3651 keyless-lane test listed cursor as keyless, including a live config-set assertion that now correctly succeeds instead of failing (swapped to qwen). - gsd-core/workflows/settings-integrations.md: the settable-keys enumeration and two prose call-outs still named cursor as keyless. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3653): acknowledge deliberate growth of settings-integrations.md settings-integrations.md grew 4 bytes because it now enumerates review.models.cursor as a settable key alongside the other reviewer lanes, matching the modelConfigKey wired for cursor in this PR. Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3653): fix malformed Emitted-Drift-Ack-Growth trailer The previous commit's trailer was separated from Co-Authored-By by a blank line, splitting it into an earlier, non-trailer paragraph — git's trailer parser only recognizes the last contiguous block. Restating it here immediately adjacent to Co-Authored-By so both parse as trailers. Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3653): backfill changeset pr number pr:0 -> pr:4160 now that https://github.com/open-gsd/gsd-core/pull/4160 exists. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
bf4485ada2 |
enhance(#3717): make the edge probe's shape cues language-aware via an optional text_en field (#4156)
* test(#3717): add failing-first coverage for text_en language-aware classification Adds unit tests for the not-yet-implemented text_en field on Requirement (fallback selection, empty/whitespace/non-string rejection, shapes-override precedence), a SHAPE_CUES/VALID_SHAPES parity guard (RULESET.GENERATIVE-FIX), and workflow-prose contract tests asserting spec-phase.md Step 5.5 documents populating text_en for response_language projects. All new tests are RED until src/edge-probe.cts and the workflow docs are updated. * feat(#3717): make edge-probe shape classification read an optional text_en field Requirement gains an optional text_en; classifyShape's own signature stays untouched (a locked, directly-tested export), and the text_en ?? text selection is pushed to proposeEdges' single call site instead. text_en is validated fail-closed: an empty or whitespace-only value throws rather than silently winning the ?? fallback and degrading classification to zero shapes. This makes the #2773 doc-only translation convention an explicit, validatable field instead of an invisible instruction, per the approved Form-1 scope on #3717. * docs(#3717): document the text_en field across spec-phase, reference and how-to docs Updates Step 5.5's response_language instructions, the edge-probe reference Inputs contract, the FEATURES.md fragment, and the non-English how-to guide to describe the new text_en field: text keeps the requirement's own wording in all cases, text_en (when populated) is the engine-only English rendering the classifier prefers. * docs(#3717): record the text_en locked-surface change in CONTEXT.md and ADR-550 Updates the Edge Probe Module glossary entry to describe the text_en field and its fail-closed validation, and appends an ADR-550 amendment recording why this is additive and does not re-open the #652 LLM-classifier rejection (text_en is a plain field read by the existing deterministic regex classifier, not a new model-dependent surface). * docs(#3717): add changeset fragment and regenerate FEATURES.md pr:0 placeholder — backfilled with the real PR number after the PR opens. * docs(#3717): attribute the text_en machine check to engine-level validation, not prose tests Code-review (Spec axis) finding: the workflow-prose contract tests and the ADR-550 amendment overclaimed themselves as "the machine check the #2773 doc-only stopgap lacked." That check is actually engine-level (validateRequirement/classifyShape, covered in tests/edge-probe.test.cjs) — the prose tests are the same style of assertion #2773 already used. Reworded both to attribute the claim correctly. * fix(#3717): rewrap spec-phase.md so the id-unchanged sentence stays on one line The #3717 rewrite of Step 5.5's response_language paragraph moved a line break so "requirement `id`s" ended one physical line and "are never translated" started the next. The pre-existing #2773 regression test (tests/edge-probe-spec-phase-contract.test.cjs) asserts id + "never translated" on the SAME line (no \n in between, matching git's own line-oriented prose), so the reflow silently broke it. Rewrapped so the sentence lands on one line again, verified against every #2773/#3717 regex assertion in that test file. Emitted-Drift-Ack-Growth: spec-phase.md — #3717 adds text_en documentation to Step 5.5 (response_language paragraph + REQS_JSON heredoc comment); this growth is this PR's own diff, not incidental drift. * chore(#3717): backfill changeset PR number pr:0 -> pr:4156 now that the PR exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
4dfc46bbe7 |
enhance(#3348): add a context-drift pre-check gate to plan-phase (#4147)
* test(#3348): add failing-first coverage for the context-drift gate * feat(#3348): add context-drift pre-check gate for plan-phase Compares each phase's *-RESEARCH.md/*-PATTERNS.md/*-VALIDATION.md/*-SPEC.md effective last-changed time (git commit time, falling back to mtime for uncommitted edits) against *-CONTEXT.md's, so plan-phase no longer silently reuses an upstream artifact that predates a decision added to CONTEXT.md after that artifact was derived from it. Deterministic, no model call. New `gsd_run verify context-drift <phase>` command, sibling to the existing verify.codebase-drift/verify.schema-drift gates in the drift capability. Warn-only by default (workflow.context_drift_precheck), with an opt-in workflow.context_drift_action: block escape hatch. Wired at plan:pre in plan-phase.md, before both the RESEARCH.md and PATTERNS.md reuse decisions. * fix(#3348): address code-review findings — raw-text-match, stale comment, import placement, duplicated phase resolution * fix(#3859): pin the real commit's diff.ignoreSubmodules to match the empty-diff probe The #3859 empty-diff guard decides whether a submodule bump would land using `--ignore-submodules=dirty`, overriding the caller's `diff.ignoreSubmodules` config. The real `git commit -- <paths>` that follows was never given the same override, so under a bare `diff.ignoreSubmodules=all` repo config the two calculations disagree: driven on git 2.39.5 (Debian bookworm, the linux-node24 test-matrix image), the guard correctly stands aside but the scoped commit itself then silently fails (exit 1, no error text) for a gitlink bump it had just confirmed would be recorded, surfacing as commit_failed instead of committed:true. Pin `-c diff.ignoreSubmodules=dirty` onto the scoped commit call too, so the probe and the commit it protects can never diverge. Harmless when no submodule path is involved (driven: identical outcome on an ordinary scoped file, with and without the flag). * fix(#3348): guard resolvePhaseDirByToken's exact-match fallback against path traversal * fix(#3348): retarget phase-enumeration-drift exemption to the consolidated resolvePhaseDirByToken helper cmdVerifySchemaDrift's inline readdirSync was already function-scoped-exempt in lint-phase-enumeration-drift.cjs as a single-phase LOOKUP (not a current-milestone enumeration). This PR's refactor pass lifted that block into a shared helper, resolvePhaseDirByToken, also used by the new cmdVerifyContextDrift — the guard tracks exemptions by enclosing function name, so the readdirSync now lives in an unexempted function and started firing. Move the exemption to resolvePhaseDirByToken (same written reason, now covering both callers) instead of migrating to listAllPhaseDirs, which would introduce two real behavior deltas here: it catches readdirSync failures internally (old code let them throw) and sorts results by phase number before matchPhaseDirs picks matches[0] (old code used raw, OS-dependent readdirSync order). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): satisfy lint:ci — slash form, capability registry regen - docs/features/context-drift-gate.md used the deprecated /gsd: colon form; docs are never passed through the install-time slash-form converters, so lint-docs-command-form requires the hyphen form. Regenerated docs/FEATURES.md from the corrected fragment. - Regenerated gsd-core/bin/lib/capability-registry.cjs after editing capabilities/drift/capability.json (lint:generated-sync). * fix(#3859): pin the real commit's diff.ignoreSubmodules via env, not argv -c The prior fix pinned `-c diff.ignoreSubmodules=dirty` onto the scoped commit's argv via `commitArgs.unshift(...)`. `-c key=val` must precede the `commit` subcommand, so this shifted `commitArgs[0]` from `'commit'` to `'-c'` for every scoped commit call, breaking 17 position-based assertions in the commit-files pathspec regression suite that read `a[0] === 'commit'` to find the commit invocation among recorded git calls. `execGit` already accepts an `env` option merged onto `process.env` before spawning. Git honors `GIT_CONFIG_COUNT`/`GIT_CONFIG_KEY_0`/`GIT_CONFIG_VALUE_0` as a per-invocation config override functionally identical to `-c key=val`, expressed via env instead of argv. Passing that env alongside the existing commitArgs (still `['commit', ..., '--', ...stagedPaths]`, argv unchanged) fixes the real commit's effective diff.ignoreSubmodules to match the empty-diff guard's probe without moving anything in argv position 0. Scoped to exactly the canScope branch, matching the probe's own preconditions and leaving no behavior change for commits the probe never evaluated. No test file changes needed — the 17 previously-failing assertions test argv[0] against the array passed into execGit, which never changes. * fix(#3348): register verify-context-drift in the check subcommand router The drift capability's new plan:pre gate declares check.query "verify.context-drift", which normalizes to `check verify-context-drift`, but no such subcommand was routed — phase6-capstone-conformance's uniform-block-field test failed with "Unknown check subcommand" for every declared gate query. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): extend #1592's exact-key-list snapshot for the new context-drift config keys tests/capability-registry.test.cjs asserted an exact, hardcoded snapshot of the drift capability's config keys. #3348 legitimately adds two new keys (workflow.context_drift_precheck, workflow.context_drift_action) for its own plan:pre context-drift gate — extend the expected set (and clarify the assertion message) without weakening the test's exactness. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): reconcile E2's exemption-migration pin with the resolvePhaseDirByToken extraction #3348 (an earlier commit on this branch, e4b80ad81) extracted cmdVerifySchemaDrift's inline phasesDir readdirSync/matchPhaseDirs block into the shared resolvePhaseDirByToken helper (also used by the new cmdVerifyContextDrift), and retargeted lint-phase-enumeration-drift.cjs's function-scoped exemption from cmdVerifySchemaDrift to resolvePhaseDirByToken accordingly — cmdVerifySchemaDrift no longer contains a line the guard's detectors match, so it needs no exemption. tests/phase-locator.test.cjs's E2 test still pinned the exemption to the old name (cmdVerifySchemaDrift), unaware of the migration. Update E2 to match the same "migrated call site's exemption must move, not duplicate" pattern the test already applies to cmdRoadmapAnalyze and cmdInitMilestoneOp just below it: drop cmdVerifySchemaDrift from the still-exempt list and add symmetric assertions that it no longer carries the exemption while resolvePhaseDirByToken now does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): fix two self-contradicting/nondeterministic tests in context-drift.test.cjs 'always exits 0 (query command contract)' included the no-phase-arg case, which contradicts the file's own earlier 'errors with usage message on missing phase arg' test (that case legitimately exits 1 via the Usage error) — drop it from the always-exits-0 cases. 'degrades to mtime comparison outside a git repo' and '...in a repo with no commits' relied on real wall-clock ordering between two back-to-back writeFileSync calls to prove CONTEXT.md is newer than RESEARCH.md; on a fast filesystem both can land in the same mtime tick, producing a tie that computeContextDrift's strict `<` correctly treats as not-stale, so stale_artifacts comes back empty. Make both tests deterministic via explicit fs.utimesSync instead of relying on timing (CONTRIBUTING.md: never assert elapsed wall-clock time). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): add context_drift_precheck:false to the plan:pre all-off fixture The "all plan:pre when-keys false" fixture explicitly disables every known workflow.* plan:pre toggle, but didn't yet know about the new workflow.context_drift_precheck key (defaults to true), so the new drift context-drift gate stayed active and broke the empty-activeHooks assertion. Emitted-Drift-Ack-Growth: plan-phase.md — adds the #3348 context-drift plan:pre pre-check section (new ## 4.6); this PR's own diff, not incidental drift. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3348): backfill changeset PR number (pr:0 -> 4147) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
c0fd2e3f4c |
feat(#3673): add dispatch.maxConcurrency axis and dispatch-capacity query (#4162)
* test(#3673): add failing tests for dispatch.maxConcurrency axis and dispatch-capacity CLI route Extends tests/host-integration.test.cjs with negotiateHostCapabilities maxConcurrency negotiation coverage (test matrix rows 1-15, including a fast-check property test) and a new #3673 dispatch-capacity CLI route describe block spawning the real gsd-tools.cjs (rows 16-25). Extends tests/host-integration-validator-parity.test.cjs with an all-19-descriptor maxConcurrency presence/validity sweep (row 26) and adds a hostile-input validator test to host-integration.test.cjs (row 27). The dispatch.maxConcurrency field does not exist yet, so these tests fail. * feat(#3673): add dispatch.maxConcurrency axis, negotiation, validator parity, and the dispatch-capacity query Adds a numeric dispatch.maxConcurrency sub-field to the Host-Integration Interface (ADR-1239 Phase 1), following the existing dispatch.isolation sub-field pattern: DispatchCapability interface, SAFE_DEFAULTS/PROFILE_BASELINES floors, and a negotiateHostCapabilities branch that passes through a positive safe integer and fails closed to 1 otherwise (no engine-side reduction, per the design doc's explicit rejection of a min(host,engine) rule). capability-validator.cjs gains parity validation for the new field (optional, positive safe integer or the "undocumented" sentinel — mirroring isolation's "added after existing descriptors" treatment). gsd-tools.cjs gains a new `query dispatch-capacity` route, a pure-read sibling of `dispatch-isolation` with no side effects: live env (GSD_DISPATCH_MAX_CONCURRENCY) > descriptor > fallback-to-1 precedence. All 19 capabilities/*/capability.json descriptors now declare dispatch.maxConcurrency: claude carries the one cited value (20, per code.claude.com/docs/en/sub-agents); the other 18 carry "undocumented" (not yet researched for this axis). * docs(#3673): document dispatch.maxConcurrency and add its citation row to the capability matrix Updates docs/reference/host-integration-interface.md's dispatch struct entry (also backfilling the previously-undocumented isolation/backgroundDispatch sub-fields found stale in the same table) and adds fail-closed/live-transport precedence prose for the new maxConcurrency field. Adds a dispatch.maxConcurrency row (with citation) to all 19 host sections in docs/reference/host-integration-capability-matrix.md — required for tests/host-integration-descriptors.test.cjs's kimi-code matrix-parity check, which asserts every declared dispatch sub-axis is documented there. Updates docs/how-to/add-or-update-a-host-integration.md's dispatch checklist and example descriptor block to mention maxConcurrency (and, likewise backfilling a stale gap, isolation/backgroundDispatch). * fix(#3673): extract shared maxConcurrency validator, drop dead reserved-name check Exports isPositiveSafeInteger from src/host-integration.cts as the single source of truth for the dispatch.maxConcurrency positive-safe-integer contract; negotiateHostCapabilities and gsd-tools.cjs's routeDispatchCapacity now both call it instead of independently reimplementing the same predicate. Also removes the __proto__/constructor/prototype reserved-name branch from capability-validator.cjs's maxConcurrency check — copy-pasted from the string-enum fields above it, but unreachable for a numeric field (the generic positive-safe-integer branch already rejects any string) and absent from maxDepth, the field the code's own comment claims to mirror. --------- Co-authored-by: sim <sim@local> |
||
|
|
b848b23861 |
feat(#3778): dispatch plan:pre planner contributions before quick planning (#3934)
* feat(#3778): dispatch plan:pre planner contributions in quick.md - Add plan:pre capability gate to quick.md Step 5, mirroring plan-phase.md's existing render + generic contribution dispatch pattern - Inject planner-targeted contribution fragments into the planner prompt, after AGENT_SKILLS_PLANNER, matching D-08 ordering - Add tests/quick-plan-pre-capabilities.test.cjs proving the dispatch is generic (D-01) via real scanWiredKinds/coveredKindsInRegion functions - Record quick.md's byte-growth rationale in this commit trailer for every gsd-core-verbatim runtime Emitted-Drift-Ack-Growth: quick.md — #3778: Step 5 (Spawn planner, quick mode) gains a `plan:pre` capability gate, mirroring `plan-phase.md:420-424` and `:797`. This is shipped shell and prose read by an agent at runtime, not compiled, so the reasoning has to travel with the feature rather than being deferred to a reference doc: (1) the dispatch paragraph phrases role routing possessively ("the role each entry's `into` names") rather than as an `into ==` equality, because `coveredKindsInRegion` (scripts/gen-loop-host-contract.cjs) voids a segment's `kind == "contribution"` coverage credit when a role or capability equality shares that same segment — an equality phrasing here would silently fail the generic-dispatch proof required by D-01; (2) `activeHooks` is read directly in-context from `PLAN_PRE_HOOKS_JSON`/`HOOKS_JSON` and the unfiltered `rendered` digest is explicitly forbidden from being pasted, because `rendered` carries every kind and role — including non-planner-targeted contributions such as a `into: "checker"` twin — and pasting it would leak checker-scoped guidance into the planner's prompt (T-01-02 in the threat model); (3) the injection block sits inside `<planning_context>` AFTER `${AGENT_SKILLS_PLANNER}` and after the Project skills line, matching plan-phase's `:741` -> `:797` ordering (D-08), so agent-skills content is never shadowed by capability-contributed prose. No prose was moved into an eagerly `@`-imported reference to shrink the measured file — @gsd-core/references/loop-hook-dispatch.md already existed before this change and is deferred to for the generic contract only, exactly as plan-phase.md already does. * test(#3778): expand quick.md plan:pre dispatch coverage to all nine locked conditions Extend tests/quick-plan-pre-capabilities.test.cjs with D-02 (silent omit-when-empty), D-03 (single shared planner spawn), D-06 (array-order dispatch phrasing), D-07 (planner-only into filter), and D-08 (render call < agent-skills placeholder < injection block < spawn ordering) assertions, all extracted via a brace-bounded slice anchored on the literal injection instruction rather than a naive first-brace scan (${AGENT_SKILLS_PLANNER} and the surrounding prompt's ${VALIDATE_MODE ? ...} ternaries also contain brace pairs). Add a capability-registry.test.cjs describe block proving the registry-wide D-07 exclusion is meaningful: at least one plan:pre contribution exists, every plan:pre contribution has a non-empty into/fragment.inline, and the registry as a whole carries at least one non-planner-into contribution. Add a loop-host-contract.test.cjs regression pin for D-09: quick.md stays absent from STEP_WORKFLOWS, parseLoopHostBlock still throws on quick.md's real content, and buildContract() still yields exactly 5 entries. Verified red-without-Task-1 by temporarily reverting quick.md to its pre-f30de9cc content and re-running these three suites (D-08 failed as expected), then restored via git checkout and re-confirmed green. * docs(#3778): note quick planning also renders plan:pre in the tutorial The tutorial's Step 6 named only /gsd-plan-phase as the trigger for the plan:pre hook set. Since quick.md now dispatches the same hook set (f30de9cc), the sentence understated the capability's real reach. * feat(#3778): add changeset fragment * chore(#3778): reference the upstream issue in the changeset fragment The fragment was the only one of 81 in .changeset/ without a trailing (#NNNN) reference or a bold lead-in. serializeChangelog auto-appends only the pr: field, so the rendered CHANGELOG entry carried no link back to issue #3778. * test(#3778): scope the D-07 registry assertion to what it actually proves The registry-wide non-planner check was named "D-07 exclusion is meaningful", which overclaims: it proves only that `into` takes non-planner values somewhere in the registry, not that anything is excluded at plan:pre. Every plan:pre contribution is currently into: "planner", so the filter is a forward-looking safeguard there. Narrowing the assertion to plan:pre (as review suggested) would fail today. Asserting plan:pre is all-planner would be brittle — it would break the day a legitimate non-planner plan:pre contribution lands, which is exactly when the safeguard starts doing work. So the assertion is unchanged and only the name and comment are corrected. * chore(#3778): point the changeset fragment at the upstream PR The fragment carried pr: 3, the fork staging PR. changeset lint derives the real PR number from GITHUB_EVENT_PATH, so on the upstream PR that would read as pr-field drift. Point it at open-gsd/gsd-core#3934. * test(#3778): require contributions in Quick revision prompts * test(loop-host): require Quick auxiliary registration * fix(#3778): preserve contributions in Quick plan revisions * fix(#3778): validate Quick as a planner contribution host * fix(#3778): tighten Quick contribution contract * test(#3778): drop unnecessary source-contract exemption * fix(#3778): require Quick planner target coverage * docs(#3778): describe targeted auxiliary coverage --------- Co-authored-by: davdittrich <davdittrich@gmail.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
c9bff8d2b9 |
fix(#3899): resolve REVIEWS.md by quoted assignment, not an unquoted ls (#3928)
* fix(#3899): resolve REVIEWS.md by quoted assignment, not an unquoted ls The convergence loop resolved the phase REVIEWS.md through `$(ls ${phase_dir}/... 2>/dev/null)`. The unquoted ${phase_dir} word-splits, so a project path containing a space handed `ls` arguments that do not exist, the discarded stderr hid the error, and the empty result was reported as "review agent did not produce REVIEWS.md" — blaming the agent for a quoting defect. A glob metacharacter fails worse: the pattern expands to whatever sibling directory matches, so the loop silently reads a different phase's REVIEWS.md and `ls` still exits 0. Assign the path directly, quoted, and fail closed with an error that names the expected location. The regression block extracts the bash fence from the shipped workflow and RUNS it against space-, glob- and missing-file fixtures, because every text assertion in that suite passed against the broken line. * chore(#3899): add the changelog fragment * fix(#3899): reject a non-file reviews path and anchor the gate harness Adversarial review (codex, gemini) on the first cut, three findings taken: - `[ -r ]` alone is true for a readable DIRECTORY, so a directory standing where the reviews file belongs passed the gate and reached the consumers that read it. Require a regular file as well. - the harness picked its fence by scanning the WHOLE document for a bash block assigning REVIEWS_FILE. If the real fence ever stopped assigning it and an unrelated one started, the harness would execute the wrong block and report green. Anchor the span to the verification step, between its opening sentence and the next heading. - the no-subshell assertion matched only `$(`; a backtick rewrite would reintroduce identical word-splitting unseen. The discarded-stderr assertion is scoped to lines naming REVIEWS_FILE rather than banning the redirect across the whole fence. Four findings declined: `return 1 || exit 1` (the document's own idiom is a bare `exit 1`, including the guard immediately below), `${phase_dir:-}` for `set -u` (unchanged exposure from the old line, and an unset phase_dir is a bug worth surfacing), an explicit empty-phase_dir branch (the guard already fails closed and prints the truncated path), and a downstream `[ -n "$REVIEWS_FILE" ]` hazard (no such consumer exists — the guard exits before any of them run). * chore(#3899): backfill the changeset PR number * fix(#3899): address required review feedback Use the fragment-aware workflow reader, remove the inert exemption marker, document the unreadable-arm coverage limit, align errors with the workflow convention, and name an empty phase_dir directly. Migrate the obsolete emitted-drift fragment to the current commit-trailer contract. Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3899: replace unquoted $(ls ... 2>/dev/null) resolution with a quoted direct path, fail-closed file/readability checks, and an explicit phase_dir diagnostic; surrounding prose records why quoting is load-bearing. --------- Co-authored-by: davdittrich <davdittrich@gmail.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
b6dd4e2e74 |
fix(#3936): dispatch quick research with researcher persona and model (#4158)
* test(#3936): quick research dispatch must use researcher persona/model Regression tests for the planner-persona leak in quick's Step 4.75 research dispatch: init quick emits researcher_model, the workflow resolves AGENT_SKILLS_RESEARCHER and parses researcher_model, and the researcher dispatch uses the researcher persona/tier while planner and executor dispatches keep their own. * fix(#3936): dispatch quick research with researcher persona and model cmdInitQuick now emits researcher_model (parity with the plan-phase init), quick.md resolves AGENT_SKILLS_RESEARCHER and parses the new field, and Step 4.75's dispatch swaps ${AGENT_SKILLS_PLANNER}/ {planner_model} for the researcher's own persona and tier. The #2517 model-omission list gains researcher_model. Emitted-Drift-Ack-Growth: quick.md — researcher persona resolution line + researcher_model parse-list entry (#3936) * chore(#3936): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
eca9c2b590 |
fix(#4112): portable grep in workflow markdown + submodule commit fix under diff.ignoreSubmodules (#4149)
* fix(#4112): ban GNU-only grep -P in workflow markdown to prevent macOS regressions Add scripts/lint-portable-grep.cjs, wire it into lint:ci, and add tests/lint-portable-grep.test.cjs. The #4112 shell-syntax fix newly exposed a pre-existing grep -oP invocation that silently resolves to "" on stock macOS's BSD grep (no -P support). This ratchet catches the same class of GNU-coreutils assumption before it merges, mirroring lint-portable-timeout.cjs. * fix(#4112): drop unneeded ls -l long-format that broke basename/dirname extraction The previous commit on this branch replaced grep -oP 'phases/\K[^/]+' (GNU-only, silently fails on macOS's BSD grep) with basename "$(dirname "$phase")"), but left ls -lt (long format, -l) in place. -l output is a full detail line (permissions, owner, size, date, path), not a bare path, so dirname on that string throws "illegal option -- r" (BSD) / errors under GNU coreutils too -- the extraction never produced a usable value, on any platform. -l was never needed here; only mtime-sort plus the bare path mattered. Dropping it to ls -t restores one-bare-path-per-line output, which basename/dirname actually requires. Confirmed on gsd-test's Linux bench: the #4112 regression test (tests/pause-work-context-detection.test.cjs) now passes phase, spike, and sketch resolution. Emitted-Drift-Ack-Growth: pause-work.md — portability fix (#4112): dropping grep -oP for a portable basename/dirname extraction is a few characters longer per line; no functional growth beyond the fix. Emitted-Drift-Ack-Growth: sync-skills.md — portability fix (#4112): replacing grep -oP '(?<=--from )\S+' with a portable sed -E capture-group equivalent (no PCRE lookbehind available) is a longer expression; growth is the direct cost of the fix, not new functionality. * fix(#3859): pin git commit itself against diff.ignoreSubmodules=all, not just the probe git 2.39.x (the exact version on the CI Linux bench) resolves a pathspec-scoped `git commit -- <path>` through the same diff.ignoreSubmodules-gated machinery as `git diff`, so a genuinely bumped submodule gitlink was silently dropped by the real commit even though the #3859 empty-diff probe correctly saw the change. The commit was misclassified as generic commit_failed because git's refusal text never says "nothing to commit". Prefix the pathspec-scoped commit invocation with -c diff.ignoreSubmodules=dirty, the same override the probe already carries, so the two can no longer disagree. Confirmed on git 2.50.1 (no-op) and git 2.39.5 via the actual gsd-tester-linux v1.8.0-node24 bench image (turns the silent refusal into a commit). * fix(#3859): carry the diff.ignoreSubmodules override via env, not argv The prior commit prepended `-c diff.ignoreSubmodules=dirty` to the real commit's argv, which shifted `commitArgs[0]` off `'commit'` and broke several pre-existing #3859 regression tests (tests/commit-files-pathspec.test.cjs) that assert on the raw argv captured at the execGit seam, e.g. `gitCalls.some((a) => a[0] === 'commit')`. Carry the same override via `GIT_CONFIG_COUNT` / `GIT_CONFIG_KEY_0` / `GIT_CONFIG_VALUE_0` env vars instead, which git honours identically and leaves argv untouched. Re-confirmed on the gsd-tester-linux v1.8.0-node24 bench (git 2.39.5): the submodule commit still succeeds. * test(#3859): pin the commitEnv GIT_CONFIG_* override to canScope directly The commitEnv/GIT_CONFIG_* override added in e935694fc/b3d37b929 was only exercised indirectly through pre-existing submodule integration tests. Add two dedicated regression tests that pin the canScope=false side of the scoping decision: a whole-index commit (no --files) and an --amend commit must both keep recording a bumped submodule gitlink under diff.ignoreSubmodules=all with no override applied, since neither shape carries a pathspec for that git internal check to consult. * fix(#4112): match grep invocations after then/do/else/elif shell keywords lint-portable-grep.cjs's GREP_INVOCATION_RE only anchored to line-start or right after `| & ; ( \` {`, so a grep call positioned right after `then`, `do`, `else`, or `elif` (e.g. `if x; then grep -oP '...'; fi`) was never flagged, letting the GNU-only grep -P defect this lint exists to catch reappear undetected in that shape. * fix(#3859): apply the diff.ignoreSubmodules override to every commit, not just scoped ones The commitEnv GIT_CONFIG_* override landed in e935694fc/b3d37b929 only when canScope was true, on the assumption that git 2.39.5's silent refusal of a bumped submodule gitlink under diff.ignoreSubmodules=all only affects a pathspec-limited `git commit -- <paths>`. Reproduced directly against the pinned CI tester image (ghcr.io/open-gsd/gsd-tester-linux:v1.8.0-node24, git 2.39.5): a bare whole-index `git commit -m ...` with no pathspec at all is refused identically when the only staged change is a submodule gitlink, since git's "nothing to commit" check is a real diff against HEAD that honours diff.ignoreSubmodules regardless of whether a pathspec narrows it. Apply the override unconditionally instead of gating it on canScope; it remains a confirmed no-op for --amend, which never hits this refusal at all. Caught by the new regression test added in 579ac9f8f, which failed against the real bench git version before this change. * fix(#3859): apply diff.ignoreSubmodules override to commit-to-subrepo and pr-subrepo cmdCommit already carries the GIT_CONFIG_* override that forces diff.ignoreSubmodules=dirty on the actual git commit call so a bumped submodule gitlink is not spuriously refused under git 2.39.5. The two sibling multi-repo commit paths, cmdCommitToSubrepo and cmdPrSubrepo, build the identical canScope-branched commit invocation but never carried the override, so the same refusal there surfaces as a generic commit_failed/error instead of a recorded gitlink. Apply the override unconditionally to both, matching the corrected cmdCommit shape. * fix(#3859): pin pr-subrepo's change detection against diff.ignoreSubmodules cmdPrSubrepo discovers what to commit via `git status --porcelain`, which (like the empty-diff probe fixed for cmdCommit) honors a local diff.ignoreSubmodules=all config. Under that config, a genuinely bumped submodule gitlink is invisible to the status scan, so changedFiles comes back empty and the function reports nothing_to_commit before ever reaching the commit call this same issue already fixed. Pin the status probe with --ignore-submodules=dirty, mirroring the flag cmdCommit's diff probe already uses, so a real gitlink bump is detected regardless of local config. Found while adding a regression test for the previous #3859 follow-up fix: the test failed not on the commit step but on this earlier detection step. * docs(#3859): update changeset to cover all three fixed commit sites * docs(#4112): backfill changeset PR number (pr:0 -> pr:4149) --------- Co-authored-by: sim <sim@local> |
||
|
|
80d2236ca5 |
fix(#4112): drop redundant $(( subshell wrapping in pause-work.md Context Detection (#4140)
* test(#4112): add failing-first regression coverage for pause-work.md Context Detection Extracts and executes the shipped Context Detection bash block to prove the $(( arithmetic-context misparse (SC1102/SC1106/SC2205) as a hard syntax error, plus boundary/independence coverage for the phase/spike/sketch/ deliberation resolution the fix must preserve. * test(#4112): fix false-green regression test — assert dash/sh failure, not bash -n The prior version asserted bash -n syntax validity and wrapped bash -c execution, neither of which detects this bug: bash/zsh silently fall back to a working command-substitution reading of the ambiguous $(( construct, and bash -n never evaluates arithmetic-context content at parse time. A gsd-test run against the still-broken file passed 41594/41594 with this test in place — a false green. POSIX sh/dash does not implement that fallback and throws a real "Syntax error: Missing '))'", which is what this version asserts against. * fix(#4112): remove leaked tool-call tags corrupting the regression test file A subagent's Write introduced trailing </content>/</invoke> markup at the end of tests/pause-work-context-detection.test.cjs. gsd-test's own tests/portability-rule-disable-ban.test.cjs (which parses every test file) correctly caught this as a parse error. Strips the garbage lines; no behavioral change. * fix(#4112): drop redundant $(( subshell wrapping in pause-work.md Context Detection phase=/spike=/sketch= opened with $((, which POSIX sh/dash parses as arithmetic expansion and rejects (Syntax error: Missing '))') since the enclosed text isn't valid arithmetic. bash/zsh silently retry it as command substitution, which masked the defect there. Dropping the inner grouping parens (matching the deliberation= line already in the same block) makes the construct valid under bash, zsh, and dash alike, with identical fallback-to-empty-string behavior when nothing matches. Once the surrounding syntax parses, ShellCheck can now also analyze the inner ls -lt calls it previously couldn't see past the parse failure, newly surfacing the same pre-existing SC2012 ("use find instead of ls") suggestion already baselined for the file's other ls usage. Baseline updated to reflect the true current count; no ls-vs-find behavior change made, as that is a separate, pre-existing, out-of-scope question. * fix(#4112): verify shell discrimination instead of trusting the binary name Code review finding: falling back from dash to sh could silently produce a non-discriminating test on a host where /bin/sh is bash-compatible (e.g. macOS) — such a shell never throws on the broken construct either, so the test would pass whether the bug were present or not. resolvePosixShell now verifies the candidate actually rejects a known-ambiguous $(( snippet before using it, and skips with an explicit reason when none does. The gsd-test Linux bench (dash as /bin/sh) is unaffected either way. * fix(#4112): route regression test through process-seam, splitLines, cleanup npm run lint:ci flagged 7 violations gsd-test's node:test run doesn't check: local/no-adhoc-markdown-parsing (single fence-spanning regex), local/no-crlf-fragile-split (bare \n split), local/no-unbounded-spawn (x4, hand-rolled execFileSync with no timeout), and local/no-raw-rmsync-in-tests. Rewrites the test to route every subprocess call through tests/helpers/process-seam.cjs's runHook() (bounded by construction), tests/helpers.cjs's cleanup() for temp-dir removal, and a line-by-line fence scan using the text-lines.cts splitLines() seam, matching this repo's established extraction idiom (tests/no-hardcoded-home-gsd-tools.test.cjs). No behavioral change to what is asserted. * docs(#4112): add changeset for the pause-work Context Detection fix * docs(#4112): backfill changeset PR number (pr:0 -> pr:4140) --------- Co-authored-by: sim <sim@local> |
||
|
|
05092ff369 |
fix(#4113): add no-op to empty then-body in chunked-planning-mode.md's outline resume-check (#4125)
* fix(#4113): add no-op to empty then-body in chunked-planning-mode.md's outline resume-check The `### 8.5.1 Outline Phase` resume-detection `if` block only contained a comment in its `then` clause, which is a syntax error under both bash and zsh if executed literally. The workflow ShellCheck lint added by #4109 already flags this (SC1009/1048/1072/1073/2105) but the findings were masked by matching entries in the pre-existing-findings baseline; those five entries are removed here so a regression re-surfaces as a new failure instead of being silently re-absorbed. * chore: backfill changeset PR number for #4125 * fix(#4113): fix second empty-then/continue-outside-loop defect surfaced by baseline cleanup Removing the 5 stale baseline entries for chunked-planning-mode.md in the prior commit unmasked a second, genuinely separate ShellCheck finding (SC2105) in a different fenced block of the same file: a bare `continue` outside any bash loop in the ### 8.5.2 per-plan resume-check. Same root cause class as the original bug (pseudocode fenced as literal bash) — fixed the same way, with a `:` no-op preserving the documentation. --------- Co-authored-by: sim <sim@local> |
||
|
|
8c9265d4e5 |
fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a blocker" but tagged severity: warning — the tier plan-phase's revision loop counts as must-fix — and the planner is never taught the rule, so every multi-wave phase touching shared mutable state replans at least once, and intentionally coupled plans re-flag identically every iteration to the stall prompt. Three coordinated changes: - gsd-plan-checker: retag 3b to severity: info, the tier references/revision-loop.md already exempts by design; recognize a coupling_justified frontmatter declaration in the Do-NOT-flag list so deliberate pairs converge. Additions are offset by trimming 3b motivation prose — the checker sits 45 bytes under its LARGE hard cap. - plan-phase step 12: INFO-only accept — an issues block with zero BLOCKER/WARNING entries accepts the plan and surfaces the advisories instead of re-entering the revision loop. Real blockers and warnings still gate unconditionally. - gsd-planner: slim pointer in assign_waves to the new progressive-disclosure reference gsd-core/references/planner-coupling.md (the planner sits 19 chars under its own cap), which carries the shared-mutable-state rule and the coupling_justified escape hatch so first-pass plans avoid the finding when the coupling is unintentional. Documented the coupling_justified field in docs/reference/plan-md.md. Growth acks per #2914; inventory manifest and install-tree fixtures regenerated for the new reference file. Closes #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): pin Dimension 3b at severity: info The severity retag makes the old assertion (severity: warning) stale; lock the advisory tier from both directions — info must be present, warning must not — so a future edit cannot silently re-arm the revision-loop trigger. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * chore(#3724): changeset fragment for PR #3758 Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * docs(#3724): roster planner-coupling.md in docs/INVENTORY.md The new reference was enumerated in the manifest and all 19 install-tree fixtures but missing its row in the Modular Planner Decomposition table — the roster half the manifest-sync test cannot check. (Review Blocker.) Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): cover all four acceptance criteria (review round 1) - plan-checker-coupling: the 3b severity assertion is now a PARITY check deriving the exempt tier from revision-loop.md's flow instead of hardcoding info — editing either side alone reds the suite. New describe pins the other three criteria: plan-phase's INFO-only accept clause (proven failing-first), the BLOCKER + WARNING count staying intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and the planner pointer + planner-coupling.md content. - ack fragment: $comment's plan-phase figure corrected to +79B; the 2775 pin note carried forward into the gsd-planner.md entry, updated for upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves verbatim). The parallel-dependent-plans re-anchor this commit originally carried was superseded by upstream #3764 during review; this branch no longer touches that file. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 2 — align the stance enumeration, complete the template contract MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it. Funded by extracting the inline <examples> block to the new progressive- disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined from the same spot; #1949 precedent), which also restores the 3b motivation clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base (49107 -> 48486) — the extraction the byte pressure was owed. MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and the field's shape becomes one 'plan-id: reason' string per coupled peer so a plan justified against two peers can express it; docs/reference/plan-md.md's Type column names the shape. NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow ratchet an 'XL tier'. Acks and derived artifacts updated accordingly (checker entry removed — a shrink needs no ack; INVENTORY roster row + regen:derived for the new file). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): derive the 3b negative severity assertion (review round 2) Every severity token in the 3b span must BE the tier revision-loop.md exempts, replacing the hardcoded severity:warning negative — if the loop's exemption ever moves, the failure names the real conflict instead of blaming the agent file with a mutually-unsatisfiable pair. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): refit the planner coupling pointer under the char cap Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the base, leaving 5 chars of headroom where the +16-char pointer was measured against 13 more. The pointer prose shortens to 'Non-file coupling:' — 49150 chars, back under the strict 49152-char cap — and the ack figures follow. The @-path the tests pin is unchanged. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep Upstream #3078/#3823 deleted all fully-spent ack fragments, including 3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md +79B append. Per the collision remedy that sweep added: take the deletion and home the still-live entry in this PR's own fragment. Figures re-measured at this merge base (90871 -> 90950 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack Upstream #3825 shipped 3172-stated-failing-direction.json naming only plan-phase.md, now spent at the base — colliding with this PR's live plan-phase entry. Per the #3003 pattern the fully-spent single-path fragment is deleted and this fragment stays the path's one source; figures re-measured at this base (93073 -> 93152 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 3 — true up the ack figures, restore the wave comment The fragment's absolute sizes are re-measured and anchored to base |
||
|
|
e242d85c92 |
fix(#4109): rewrap unquoted DISPATCH_SLUGS-shaped consumption to survive zsh (#4116)
* test(#4109): add zsh regression coverage for gate-check and invoke_reviewers dispatch Extends tests/review-plan-coverage-manifest.test.cjs to cover the two DISPATCH_SLUGS consumption sites the existing #3301 coverage-check tests don't reach: the write_reviews gate-check block (ALL_LANES_SKIPPED / TOTAL_LANE_FAILURE counting) and invoke_reviewers' dispatch + join loops. Both extract the real shipped bash from review.md and run it under bash and zsh, same convention as the existing coverage-check rows. Expected RED under zsh pre-fix, GREEN post-fix. * fix(#4109): rewrap unquoted DISPATCH_SLUGS-shaped consumption to survive zsh Bash word-splits an unquoted scalar on IFS by default; zsh does not (no setopt SH_WORD_SPLIT anywhere in these files), so a scalar accumulated from multiple space-separated tokens and then consumed via bare `for x in $VAR` collapses onto one bogus iteration under zsh whenever it holds 2+ tokens. Fixes all four review.md sites reported by #4108/#4109 (invoke_reviewers dispatch + join loops, the gate-check block, and the #3301 coverage-check block), plus the identical pattern found by a repo-wide sweep in complete-milestone.md, code-review.md, pr-branch.md, sync-skills.md, and execute-phase/steps/per-plan-worktree-gate.md. Each site is rewrapped in unquoted command substitution (`$(printf '%s' "$VAR")`), which re-splits identically under both shells regardless of SH_WORD_SPLIT — the same mechanism that already made the accumulator-building loops in these files shell-safe. * test(#4109): warn loudly when the zsh probe fails instead of skipping silently detectShells() drops the zsh test lane whenever a live zsh probe fails (e.g. on a CI runner without zsh installed), and a dropped lane reads identically to a passing one in the suite's own output — exactly the blind spot that let #4109's bug class ship undetected. Both copies of this helper (this file and its byte-identical duplicate in review-build-prompt-optional-sections.test.cjs) now emit a greppable console.warn when the probe fails, so a run without zsh reads as "zsh coverage unknown" rather than "all lanes green". * ci(#4109): gate workflows/*.md's embedded bash blocks in lint:ci Adds scripts/lint-workflow-shellcheck.cjs, wired into npm run lint:ci, so this bug class can't land undetected a third time. Two independent checks run over every ```bash block in gsd-core/workflows/**/*.md: - ShellCheck (new `shellcheck` devDependency, downloads and caches the real koalaman/shellcheck binary) catches the general unquoted-expansion/ word-splitting family in argument position. 212 pre-existing findings across the tree are absorbed into scripts/lint-workflow-shellcheck-baseline.json (matched on {file, code, message}, not line number, so unrelated edits elsewhere in a file can't spuriously un-baseline anything) — fixing all of them is out of scope for this issue; only NEW findings fail the build. - A custom structural check specifically for #4109's own shape: ShellCheck does not flag a bare `for x in $VAR` word-list — it treats that as an intentional idiom under any ruleset (confirmed empirically). This check does, and gates the build on any occurrence, with zero tolerance (no baseline) since every known site was already swept and fixed on this branch. * test(#4109): update pr-branch cherry-pick loop test anchor for the zsh fix extractPickLoop()'s PICK_LOOP_MARKER located the create_pr_branch cherry-pick loop by the literal substring "for HASH in $INCLUDED_COMMITS", which no longer appears verbatim after this issue's fix rewrapped that loop in unquoted command substitution. Updates the anchor (and its error message, now derived from the same constant instead of duplicating stale text) to the new literal form. No behavior change to the extraction logic itself. Emitted-Drift-Ack-Growth: review.md — #4109's fix adds explanatory comments at 4 sites; net code is functionally equivalent, comment expansion grows the file Emitted-Drift-Ack-Growth: complete-milestone.md — #4109's fix adds explanatory comments at 3 sites documenting the zsh word-splitting divergence Emitted-Drift-Ack-Growth: code-review.md — #4109's fix adds an explanatory comment documenting the zsh word-splitting divergence Emitted-Drift-Ack-Growth: pr-branch.md — #4109's fix adds explanatory comments at 3 sites (TRANSIENT_DIRS, INCLUDED_COMMITS, FILTER_PATHS) Emitted-Drift-Ack-Growth: sync-skills.md — #4109's fix adds explanatory comments at 2 sites (CREATE_LIST/UPDATE_LIST, REMOVE_LIST) * test(#4109): cover lint-workflow-shellcheck parsers + add timeout Adds tests/lint-workflow-shellcheck.test.cjs covering the hand-rolled parser/logic functions exported by scripts/lint-workflow-shellcheck.cjs (stripCommandSubstitutions, stripShellComments, substitutePlaceholders, extractForLoops, findBareForLoopSplits, findingKey, partitionAgainstBaseline) that shipped with zero coverage — CLAUDE.md requires at least one fast-check property test for parsers, included here (bare/braced forms always flagged naming the variable; quoted/substituted/literal forms never are). Also bounds runShellcheck's subprocess with a 60s timeout: the `shellcheck` npm package's own API has no timeout option and internally blocks on a synchronous spawnSync, so this reimplements the binary resolve/download step via the package's own exported config/download and calls spawnSync directly with a native timeout. And guards the CLI entry point with `if (require.main === module)`, matching this repo's sibling dual-purpose lint scripts — without it, requiring the module for its exported functions (as the new test file does) also triggered a live ShellCheck run as a side effect. * docs(#4109): add changeset for the zsh word-splitting fix * chore(#4109): backfill changeset PR number (#4116) --------- Co-authored-by: sim <sim@local> |
||
|
|
ef9ce3e598 |
fix(#3702): count asterisk, plus and ordered markers as deferred-items entries (#3739)
* fix(#3702): deferred-items counts `*`, `+` and ordered markers as list items
`deferred-items.md` has no template and no mandated shape, but its parser
recognised only the `- ` hyphen marker. Asterisk bullets, plus bullets and
dot-terminated ordered lists — all lists in CommonMark and GFM — contributed
ZERO entries on both the headless and the heading-delimited path, and a mixed
file dropped its non-hyphen entries while keeping their hyphenated siblings,
under-reporting without ever looking empty.
The restriction was a regex literal inherited from the Gaps seam, where the
template genuinely mandates the hyphen YAML-lite form; nothing in the module's
stated rationale distinguishes `*` from `-`.
Widened on the deferred path only:
- `splitGapsEntriesCore`'s entry opener, `extractGapEntryFields`' line-0 strip
and `rawGapEntryText`'s line-0 strip take a `BulletMarkers` parameter that
DEFAULTS to the hyphen-only set, so `## Gaps` keeps its template-mandated
grammar byte-for-byte and the module still has exactly one grouping pass.
- `splitDeferredHeadingEntries`' body-bullet test, `stripLeadingBulletMarker`
and `acknowledgeDeferredItem`'s status-field regexes move in lockstep —
widening what OPENS an entry without widening what is STRIPPED before field
extraction would surface an entry that can never resolve.
Unchanged, and pinned by tests: prose-only and bare headings still contribute
nothing ("prose is not an item"); a table under a leaf heading still yields
exactly its rows, since table lines are skipped before the body-marker flag can
be set and a `|` row is not a list marker; the paren-terminated ordered form
`1)` is out of this fix's scope.
* docs(#3702): changeset fragment (pr: 0 placeholder pre-create)
* fix(#3702): widen the forensic-audit prose entry rule to match the parser
Sibling site of the same defect class, found by a defect-class sweep of the
deferred-items consumers. `/gsd-progress` check 7 does NOT go through
`gsd-tools query` — it globs `deferred-items.md` and has the model read entries
by a prose rule that mandated "one entry per top-level `- ` line". Left as-is,
the marker widening would hold on the CLI path while the one consumer that
bypasses the parser kept reporting "No unresolved deferred items" for a file
written with `*`, `+` or an ordered marker: the same false negative, surviving
in the only place the fix could not reach by code.
Also pass DEFERRED_BULLET_MARKERS explicitly where the heading path extracts
fields. It was already correct — stripLeadingBulletMarker pre-strips the widened
set from every line, so the default hyphen strip is a no-op there — but relying
on that leaves a detection site and a strip site nominally on different marker
sets, which is exactly the asymmetry the BulletMarkers doc comment warns about.
Explicit is local; inferred is a trap for whoever edits the strip next.
Out of scope, noted rather than fixed: forensic-audit.md globs only
`.planning/phases/*/` and so misses archived milestone phases that
`scanDeferredItems` covers. Pre-existing, a different defect, and not this
issue's ruling.
* docs(#3702): note the milestone-close halt for heading-shape non-hyphen files in the changeset
A heading-delimited deferred-items.md written with */+/ordered markers
previously parsed to zero and closed silently; it now yields entries whose
heading shape acknowledgeDeferredItem refuses, halting complete-milestone
until hand-edited. User-visible, so the fragment states it.
* chore(#3702): set changeset fragment pr to 3739
* fix(#3702): CR-normalise the heading path and the acknowledge writer (review B1, M4, m2)
B1 — `splitDeferredHeadingEntries` stored RAW lines; on a CRLF file every
body line but the last still carried its `\r`, the `$`-anchored marker
strip failed on it, the marker survived into field extraction and the
field was lost — a `**Status:** resolved` that was not the file's final
line resurfaced its entry as open. The heading path now stores CR-stripped
lines like the headless path already did, and the strip regex tolerates a
trailing CR on its own. Round 1's CRLF test put `**Status:**` on the last
line, the one position `collectSection`'s `.trimEnd()` had already
de-CR'd; the new tests put it first and mid-body.
M4 (pre-existing on `next`) — `acknowledgeDeferredItem` found the status
line on a CR-stripped copy but rewrote the raw line with a `$`-anchored
`.*`, which cannot consume `\r`; `replace` returned its input, and the
writer reported `ok` over byte-identical content. The rewrite now runs on
a CR-stripped line. The comment that claimed `.*$` consumed the `\r` is
corrected — it was the bug, stated as the design.
m2 — the indent probe for an inserted `status:` line ran on the raw line
and fell back to indent 0 on CRLF; it is CR-stripped too.
* fix(#3702): derive every deferred-items marker regex from one source (review M3, N1, N2)
M3 — round 1 carried the marker alternation in FOUR places: the
`BulletMarkers` pair and two inline literals inside
`acknowledgeDeferredItem`, under a doc comment saying the interface
existed so a detection site and its strip site could not drift. All four
now derive from `DEFERRED_MARKER_ALT`; drift is impossible rather than
discouraged. A parity test pins the vocabulary against
`markdown-sectionizer`'s `iterateBullets` on everything the two grammars
are meant to agree on, and names the two points they deliberately differ.
N1 — the ordered marker is `\d{1,9}\.` (CommonMark §5.2), not `\d+\.`.
N2 — the marker is followed by `[ \t]`, not `\s`, which also accepted
`\r`; the tab remains accepted (CommonMark-legal) and the divergence from
`iterateBullets`' literal space is pinned rather than papered over.
The four regexes are exported for the parity test only.
* fix(#3702): an ordered marker opens an entry only from `1.` or inside a run (review B2, m1)
B2 — `\d+\.` alone read ordinary prose as a list: "2026. was a bad year
for this module" and, under `### Notes`, "3. is the number of retries we
settled on." both opened an entry on round 1, the second straight through
the "prose is not an item" contract that round's AC4 claimed to preserve.
CommonMark §5.3 faces the same ambiguity when an ordered list would
interrupt a paragraph and resolves it by requiring the list to start with
1; `matchListOpener` applies that rule wherever an ordered marker is seen,
with the run carried per list (headless) or per leaf-heading body. Numbers
after the first are ignored, as CommonMark ignores them. Stated cost,
pinned: a hand-numbered list starting at 2 reads as prose — every ordered
record in the #3702 scan starts at 1.
Both reviewer cases are pinned as prose; the ruling's `1. alpha / 2. beta`
shape still counts.
m1 — the 9-digit boundary is pinned at both sides (`999999999.` opens,
ten digits is not a marker), and the 3-vs-4-space indentation cliff is
pinned as deliberately NOT applied: the parser is indent-lenient because
surfacing a questionable hand-written entry beats dropping a real one.
* fix(#3702): thematic breaks close the list and fenced code never opens an entry (review M1, M2)
M1 — `- - -` was a phantom `"- -"` entry on base; widening the marker set
added `* * *` and `+ + +` to the class, and `* * *` is the separator an
author writing in the `*` style is most likely to use. A CommonMark §4.1
thematic break (plus the `+ + +` gesture, which is the same garbage as an
entry name) now closes the open entry on the headless path and is dropped
from the body on the heading path — neither an item nor a continuation.
M2 — neither splitter was fence-aware, so `+ `-prefixed diff lines and
`1.`-numbered repro steps inside a code block counted as entries; #3702's
wild records carry exactly those blocks. Both splitters now classify lines
by the sectionizer's own `scanFencedBlocks` (so `~~~`, indented and
unterminated fences behave as `stripFencedCode` would): fence content
never opens an entry, is continuation inside an open one — keeping the
span invariant `acknowledgeDeferredItem` re-verifies — and is discarded
before the first.
* test(#3702): range the #2287 deferred-items property over marker × shape × line ending (review B3)
The `#2287` property hard-coded `- ` and filtered `\r\n` out of its
arbitraries, so the widened marker set — an enumerated domain, exactly
what a property is for — was never under it. It now ranges over
`{-, *, +, ordered}` × `{headless, heading}` × `{LF, CRLF}`, with the
heading shape placing `**Status:**` first or last: the review's
prescription (markers × line endings) would not have reached B1, which
lives on the heading path only, so the shape axis is the load-bearing
addition. Ordered entries are numbered from 1, so the B2 run rule is
under the property too.
A second property drives `acknowledgeDeferredItem` over every unresolved
headless entry across the same marker × line-ending grid — the one that
reaches M4 (a CRLF rewrite that reported `ok` and wrote nothing) and m2.
* test(#3702): pin the milestone-close halt on a heading-delimited `*`/`+`/`1.` file (review m3)
A heading-delimited `deferred-items.md` written with a non-hyphen marker
previously parsed to zero entries and let `complete-milestone` close
silently; it now yields entries whose heading shape `acknowledgeDeferredItem`
refuses, which the milestone loop turns into `record_ack_failure` → exit 1.
The loop is prose in a workflow, so the test drives the two CLI calls it
makes: `audit-open --json` must list the entry, and
`audit-open acknowledge --text <the audit's own text>` must refuse with the
heading-delimited message and write nothing.
* docs(#3702): changeset and forensic-audit prose carry the round-2 grammar
The changeset names the CRLF fixes, the ordered start-at-1 rule, thematic
breaks and fences. The `/gsd-progress` forensic-audit step is the one
prose parser of this file and must state the same grammar the code has.
* fix(#3702): round-review refinements — run ends at a paragraph, rejected ordinals unstripped, breaks at any indent, fenced fields, `## Gaps` scope
Findings from the pre-push adversarial review of round 2, each pinned:
- An ordered run ENDS at a paragraph that follows a blank line (CommonMark
§5.3); a non-indented line with no blank before it is lazy continuation
and keeps the run open. `1. a` / blank / `paragraph` / blank / `5. x` is
one entry, not two.
- The heading path strips the marker off every body line before field
extraction (#3457); a line whose ordinal `matchListOpener` REJECTED must
not be stripped, or "3. status: resolved" as prose loses its `3. ` and
reads as a resolved field. `splitDeferredHeadingEntriesDetailed` now
carries a per-line opener flag and only accepted openers are stripped —
in headless regions of a heading-shaped file too.
- A thematic break is recognised at any indent, matching the parser's
indent-lenient reading of items; ` * * *` was a phantom `* *`.
- Fenced lines carry no FIELDS either: a `status: resolved` quoted inside a
code block no longer resolves its entry on either path.
- Block structure (breaks, fences) is a property of the GRAMMAR, carried as
`BulletMarkers.blockStructure`: the deferred set opts in, the Gaps set
does not, so `## Gaps` is byte-for-byte on its `next` behaviour — the
round-2 M1/M2 change had reached it through the shared splitter.
* test(#3702): the property exercises the rejected-ordinal branch; the N2 control is independent
Round review: the widened #2287 property numbered every ordered run from 1
and so never generated an ordinal the start-at-1 rule rejects — it could
not tell round 1 from round 2 on B2. Each entry may now carry a decoy prose
line beginning with a non-1 ordinal, placed where it cannot end a run
(before the first headless entry; first in a heading body), followed by a
`status: resolved` that must never become a field; and a decoy-only
heading body must yield no entry.
The N2 assertion accepted a tab, which round 1's `\s` accepted too, so a
`[ \t]` → `\s` revert alone stayed green. NBSP, form-feed and vertical-tab
are now asserted refused — the assertion that fails on that revert on its
own, and the disclosure that `[ \t]` narrows what round 1 accepted.
* fix(#3702): the splitter records its own opener flags; an opener clears the blank-line memory
Round-review continuation, two state defects in the ordered-run logic:
- `blankSeen` survived the headless splitter's opener branch, so an opener
followed by a lazy continuation line read as "paragraph after a blank" and
ended the run — `1. a` / blank / `2. b` / lazy / `3. c` folded `c` into `b`.
The opener branch now clears it.
- The heading path re-derived per-line opener flags for headless regions
without the paragraph reset, re-accepting a rejected `3. status: resolved`
under a stale run and stripping it into a field. `GapsEntrySpan` now
carries the flags the splitter itself computed, and the heading path reads
them; the re-derivation is deleted.
* fix(#3702): ordered-run memory is per indent — nested runs resolve, nested ordinals never inherit the top-level run
Round-review continuation 2: nested openers consulted the TOP-LEVEL run
flag and never wrote their own, so a nested `1. / 2.` run under a hyphen
entry rejected its `2. status: resolved` (round 1 resolved it), while a
nested `3. status: resolved` under a nested `- ` bullet inherited an open
top-level run and was stripped into a false field.
`OrderedRuns` keys the memory by indent: a new opener at indent d resets
every deeper level, a paragraph after a blank at indent d ends the runs at
d and deeper, a thematic break or a heading clears all. Both splitters use
it; the top level still decides entry boundaries, nested levels decide
only which continuation lines are accepted openers for field stripping.
Pinned for LF and CRLF.
* fix(#3702): run levels — one top level at or above the base, CommonMark column indents, a fence ends its level's runs
Round-review continuation 3:
- A dedenting top-level list (` 1.` / ` 2.` / `3.`) lost its entry
boundaries: the exact-indent run lookup rejected the shallower ordinals
before the boundary check ran. Every indent at or shallower than the
list's base is now ONE level, in both splitters.
- `indentOf` counted characters, so a tab and a space aliased to one level
and `\t1. nested` / ` 2. status: resolved` resolved falsely. Indent is now
measured in CommonMark columns (§2.2: a tab advances to the next multiple
of 4), for the run level and the entry-boundary check alike.
- A nested run survived a fenced block. A fence is a non-list block: its
opening delimiter ends the runs at its level and deeper, exactly as a
paragraph after a blank does.
* fix(#3702): the indent measure is grammar-scoped — Gaps keeps next's character count
`blockStructure: false` promised the Gaps grammar byte-for-byte parity with
`next`, but the CommonMark-column indent measure added for the deferred
grammar was shared by the whole splitter core, so tab-indented Gaps input
changed entry boundaries in BOTH directions:
`\t- a` / ` - b` — next folded into one entry, HEAD split into two
` - a` / `\t- b` — next split into two, HEAD folded into one
`indentWidth` now keys the measure on the grammar: columns for the deferred
set, raw character count for Gaps. The opt-out covers indent semantics, not
only fences and thematic breaks.
Four cases pin both halves — the two flipped Gaps pairs, the two Gaps pairs
that never moved, and the same tab/space pairs on the deferred path returning
the opposite (column-measured) verdict by design.
* fix(#3702): the acknowledge path reads and writes through one classifier
Round 3, Blockers 1 and 3, and Minors 7 and 8 — one mechanism, so one commit.
Every consumer of an entry's lines now reads the splitter's own per-line
verdict instead of a re-derivation of it.
B1. Round 2 widened the WRITER's status-line finder to the deferred marker set
while `extractGapEntryFields` still de-bulleted line 0 only. A nested
` * status: pending` was therefore selectable by the writer and invisible to
the reader: acknowledge rewrote it in place, returned `ok`, and the item stayed
outstanding on every later audit. Measured against a `next` build, `*`, `+` and
`1.` each resolved on base and stopped resolving at round 2's head — a
regression, not a gap in new behaviour. The hyphen form of the same shape was
already broken on `next` and is fixed here too: one classifier cannot be right
for three markers and wrong for the fourth.
`parseGapEntryFieldLine` is now the single place a line is classified as a
field, and it reports the offset at which the VALUE begins. The rewrite happens
at that offset rather than through a second regex, so a line the classifier can
select is one whose rewrite it has already located — the selection and the
rewrite cannot disagree. Both `DEFERRED_STATUS_FIELD_RE` and
`DEFERRED_STATUS_REWRITE_RE` are deleted rather than widened. A read-back guard
returns `rewrite_not_readable` rather than `ok`; it is unreachable by
construction today and is the fail-loud floor under the next divergence.
B3. This is the end state the round-3 review prescribed on both #3739 and
#3773: #3773's shared classifier, parameterised by this PR's marker set, with
this PR's two status regexes deleted. #3773 lands first. Its hyphen-only strip
is consistent with `next`'s hyphen-only splitter today, so the writer/reader
divergence is created by THIS merge, which is why widening every consumer
belongs to the PR that widens the domain.
m7. The heading path marker-stripped its lines before calling the reader, so
the reader's fence scan ran over text the splitter never saw: `- ```sh` is an
ordinary bullet to the splitter but strips to a fence opener, and a
`**Status:** resolved` after it was suppressed as fence content — a resolved
entry resurfaced as open. Stripping now happens inside the reader, after the
fence scan.
m8. `rawGapEntryText` stripped a marker off line 0 unconditionally, but on the
heading shape line 0 is the heading TEXT: `### 1. Race in the writer` was
silently renamed to `Race in the writer`, and the name is the key acknowledge
matches on. Line 0 is stripped only when the splitter accepted it as an opener.
Also removed: `splitDeferredHeadingEntries`, whose sole caller only null-checked
it (round 3, M4 — the claim was zero callers, which was wrong; the wrapper's
`.map` was waste at the one call site), and `stripLeadingBulletMarker`, which
this change leaves with no callers at all. The export surface narrows to the two
splitter regexes the behavioural parity test reads (M6).
[PEER-ASK pr-order-12d5]
q: Reviewer blocked both on merge order. I'm declaring #3773 lands first and
building the end-state shape into #3739 now (both my status regexes
deleted). Does that match your plan?
reply: CONFIRMED - same order, derived independently. #3773 cannot carry the
fold: `DEFERRED_BULLET_MARKERS`/`BulletMarkers` have zero occurrences at
`next` (verified), so the prescribed end state is not executable inside
#3773 without absorbing this PR's work.
deadline: 03:55 UTC (answered before it)
fallback: declare #3773 first, adopt end-state shape in #3739, push+comment
decision: proceeded as stated; #3773 lands first, this PR carries the widening
of every consumer.
Refs #3740
* test(#3702): pin the detect/strip symmetry, and drop a white-box test that could not reach it
Round 3, Blocker 2 and Minors 6 and 9.
B2. The regression shipped green because no fixture put a marker on a nested
status line. Four markers x {nested status line}, each asserting the entry
READS BACK as acknowledged rather than that acknowledge merely reported `ok` —
reporting `ok` over a line the reader skips is the whole defect. Plus the bare
capitalised `Status:` case (the reader stores it case-sensitively, so the
writer must not select it), and an idempotence test, which is the failure the
defect actually produced: the item resurfaces, is acknowledged again, and never
settles.
Each of these was run against the pre-fix build first: all five fail there and
pass here. Two further assertions in the block are labelled CONTROL because
they held pre-fix — they guard the new offset-based rewrite and the opener-flag
threading against regressing, and calling them regression tests for a reported
defect would overclaim.
M6. The round-2 parity test asserted that four writer-side regexes embedded the
same source string. That is true of a detect/read asymmetry too, so it could
not have caught B1 — and two of the four regexes were widened into `export =`
purely to let it read them. Replaced with a behavioural test that drives the
real seam: every marker that opens an entry must also resolve it through
acknowledge. The structural assertion is kept for the two splitter regexes,
which really are two copies of one alternation.
m9. `expectedResolved` was computed and immediately voided; the loop beneath it
already asserts both polarities.
m7/m8 coverage lands here too: a bullet whose content is a fence opener must
not suppress the entry's fields, and a heading beginning with a list marker
must keep it in the entry name.
* docs(#3702): document the deferred-items entry shape where the file is written
Round 3, Major 5, and #3702's own item 2. The widened grammar was documented in
the reader (`forensic-audit.md`) but not at the write site, where
`executor-examples.md` still said only "log to deferred-items.md" — so the
question the issue actually raised, which shapes count, remained unanswered
anywhere a human writes the file.
States what opens an entry (`-`, `*`, `+`, and `1.` when the list starts at
`1.`), that `1)` is not a marker here, that a separator closes the list and
fenced content is never an entry or a field, and that an entry without an
explicit `status: resolved` stays open by design.
* chore(#3702): regenerate the changeset through the generator
Round 3, Minor 10. The fragment was hand-named against 64 generated names on
`next`, and its body ran ~250 words against CONTRIBUTING's one-sentence form.
Regenerated via `npm run changeset`, which is also what the random three-word
name is for: concurrent PRs never collide.
* fix(#3702): the fence gate lives on the seam both sides call, not just the reader
Found by the pre-push adversarial review of this round, and it is a regression
this round introduced rather than a pre-existing one.
`extractGapEntryFields` applied `fencedLineSet` before classifying; the
acknowledge writer's status-line search did not. So a `status:` line inside a
fenced block was SELECTED by the writer and SKIPPED by the reader — the write
produced a line nothing reads, the read-back guard refused it, and the entry
became impossible to acknowledge at all: `audit acknowledge` raised an internal
error and `complete-milestone` halted on it.
Measured, `- alpha` / fence / ` status: pending` / fence:
next ack=ok -> reads back "acknowledged"
round-2 head ack=ok -> reads back "" (the B1 defect)
before this ack=rewrite_not_readable -> refuses entirely (worse than next)
`entryFieldLines` is now the seam — per line of an entry, the field it declares
or `null`, fences included — and the reader and the writer both go through it.
That makes "the writer cannot select a line the reader will not read back"
structural rather than asserted, which is what the previous commit's message
claimed while a second read-side filter still lived outside the classifier.
Two comments corrected with it. The read-back guard is NOT "unreachable by
construction": this round shipped a reachable path to it, which is precisely
what an invariant asserted in a comment is worth. And the M6 replacement test
put its marker only on the entry opener, so it passed against the defective
build — the exact weakness it was introduced to fix in round 2's test. It now
marks the nested status line too, and fails pre-fix like the rest.
Round-3 tests against the pre-fix build: 10 of 12 fail there, and the 2 that
hold are labelled CONTROL because they guard this round's new code rather than
pin a reported defect.
* fix(#3702): one end-of-file CRLF algorithm, adopting #3773's with its B4 closed
Round-4 M1. Two open PRs shipped two different answers to "what line ending
does an entry that ENDS THE FILE get?", and the review's ruling was that the
disagreement needs one answer, not two. Neither shipped answer was that one.
Measured on builds of both heads:
case #3739 r3 #3773 here
undelimited single entry, CRLF preamble pass FAIL pass
LF-dominant list, one stray CRLF at EOF FAIL pass pass
(the other five) pass pass pass
This PR's content.endsWith('\r\n', matchIndexInContent) reads the terminator of
the PREVIOUS line, so it propagated an isolated CRLF into an LF-dominant list --
refuted by #3773's own LF-dominant fixture, ported here. Withdrawn.
#3773's crlfAtEof asks the right question -- does anything before the entry,
within scope, contradict CRLF -- and fails closed. But its scope goes EMPTY for
an undelimited single-entry list, because the entry-list region runs from the
first entry's start to the insertion point and those coincide; crlfAtEof('') is
false by its own before.length > 0 guard, so 'preamble\r\n\r\n- alpha' gained a
bare \n in a CRLF document. That is #3773's B4, verified by driving its head.
Adopted here with the scope widened to everything preceding the insertion point
where the preferred region is empty, rather than asserting LF from no evidence.
That only ever loosens a scope carrying zero information, and the predicate
stays fail-closed over the wider one. An entry at offset 0 of an undelimited
document has no evidence under either scope and stays LF.
Tests: 10 added. Negative control, driven -- 1 of the 10 fails against this
branch's own pre-fix head (the stray-CRLF fixture); B4 fails against #3773's
head; the remaining 8 are the scope counterexamples ported with the function,
which were regression pins in #3773 and are guards here. Each still kills a
simpler algorithm: drop any one and a refuted scope passes again.
Four deferred-items suites 450/450, 0 skipped. npm run lint:ci exit 0.
* fix(#3702): drop the unreachable rewrite_not_readable guard (B3)
Round-4 B3: the status had zero test coverage in either file. The review
offered two branches -- drive it from a test, or delete it and stop carrying an
untested terminal status. Taking the second, with the reason stated rather than
assumed.
Why it cannot be driven. Round 3 added the guard after a fenced `status:` line
proved the writer could select a line the reader would not read back. Round 3
then closed that divergence STRUCTURALLY, by routing the writer's line selection
and the reader's field extraction through one entryFieldLines seam. The guard
now detects a state construction prevents: 21 document shapes were driven
against it -- fence openers on the bullet line for every marker in the widened
set, duplicate and triplicate status lines, bolded and nested variants, fences
between duplicates -- and none reached it. The only seam that would is routing
the internal call through the module's exports so a test could stub it, which
reshapes production surface for a test.
Why leaving it undriven is not free. RULESET.TESTS.mutation-score runs Stryker
incrementally over changed files at an 80% threshold and says to treat a
surviving mutant as a failing test specification. An undriven `if` on a changed
file is exactly that, on both the condition and the .toLowerCase() comparison.
What this gives up, stated rather than hidden: if a future change re-splits the
writer's selection from the reader's extraction, acknowledgeDeferredItem returns
ok over an item that stays outstanding -- the original #3702 defect class. One
correction to the review's framing: match_verification_failed does NOT backfill
it. That check runs BEFORE the write and compares the matched span to the
target, so it cannot see a post-write read-back failure. The protection against
re-splitting is the shared seam and the round-3 tests that pin it, not a runtime
assertion. A comment at the removal site records all of this.
Removing it also drops the union member from both files, which resolves the PR
body's internal contradiction (it claimed no type-signature changes while adding
one) and the duplicate-status surface #3773 collides on.
No test changed behaviour: 450/450 across the four deferred-items suites, 149/149
across the audit suites, npm run lint:ci exit 0 -- the same figures as before the
removal, which is itself the evidence that nothing exercised the branch.
* fix(#3702): the deferred fence gate is indent-unbounded, like the rest of the grammar (M2)
Round-4 M2. scanFencedBlocks is CommonMark, which caps a fence delimiter's
indent at three spaces -- a fourth makes it an indented code block instead. This
grammar had already opted out of that cliff for entry openers ([ \t]*) and for
THEMATIC_BREAK_RE (^[ \t]*), but not for fences. So a fence at four spaces was
not a fence to the gate, and a `status: resolved` line inside it RESOLVED the
entry containing it.
That is not an exotic shape. A fenced block written under a NESTED bullet sits
at four spaces, so ordinary hand-written deferred-items.md files reach it.
Driven before the fix at indents 4, 5, 8 and a leading tab: all four silently
resolved. It is the #3702 silent-resolution defect class in a new place.
gsd-core/references/executor-examples.md, added by this PR, states flatly that
"nothing inside a fenced code block is an entry or a field". The review offered
fixing the parser or bounding that claim in three places. Fixing it -- the claim
is the one users will rely on, and the grammar had already chosen unbounded
indent everywhere else.
NO second fence dialect (the rule blankIndentedFenceDelimiters states). The
classification is still done by scanFencedBlocks, the one exported CommonMark
state machine, over a de-indented VIEW of the same lines. Run lengths, backtick
vs tilde, closer-must-match-and-not-trail, info-string rules and the
unterminated-at-EOF case remain that engine's answers. Indent is the only
dimension hidden from it, and it is exactly the dimension this grammar has
already declared it does not measure. Index alignment is 1:1 -- map preserves
length -- so every returned line index still addresses the original line.
Scope is the deferred grammar only. Both marker-parameterised call sites gate on
markers.blockStructure, which the Gaps set does not set, so Gaps reaches an empty
set. Verified, not asserted: the 47-fixture Gaps differential (marker x
line-ending x separator x fence x break x key-shape x list-shape) is
BYTE-IDENTICAL across this change, 8033 bytes both sides.
Tests: 14 added, of which 8 fail against the pre-fix source and pass here; the
other 6 are the deliberate controls -- indents 0 through 3, which must NOT move,
and the Gaps opt-out guard.
Four deferred-items suites green; the 58 suites touching uat/deferred/sectionizer
run 6045 tests with an IDENTICAL failing set before and after this change (17
pre-existing environment failures -- installs and an unpinned GSD_EMITTED_BASE;
emitted-attribution passes 259/259 in isolation with its base pinned). lint:ci
exit 0.
* fix(#3702): changeset, both prose parsers, and the minors (M3, M4, m1-m3, m5, n1-n2)
M3 -- the changeset omitted a user-BREAKING change. Measured against next: a
heading-delimited deferred-items.md written with `*`, `+` or `1.` went from
"0 entries, so complete-milestone has nothing to acknowledge and closes" to
"1 entry, the CLI writer refuses the heading shape, ACK_FAILURES accumulates,
exit 1". The `-` form already halted and is unchanged. That is release-note
material: a close that used to succeed now fails, and the correct response is to
fix the file, not revert. Also names the fence-indent fix below, and adds #3740
so #3773's issue is attributed here as it is absorbed.
M4 -- gsd-core/workflows/progress/steps/forensic-audit.md is a SECOND,
model-executed parser of the same grammar, and prose cannot carry a parity test.
Its widened text stated the start-at-1 rule, fences and separators but not the
`1)` exclusion nor the nine-digit ordinal cap, both enforced in code with pinned
tests. Both stated now, along with the round-4 fence-indent rule. (No ack
fragment: the size ratchet's currentSizes does a NON-recursive readdirSync of
gsd-core/workflows and agents, so a file under workflows/progress/steps/ is
outside its scope -- verified by reading the helper, not by the green.)
n1 -- executor-examples.md documented that the BOLDED status key is matched
case-insensitively and left the bare key's rule to inference. Driven: bare
`Status: resolved` is NOT read, so the entry stays open with no warning, while
`**Status:**` is. Stated explicitly, with the digit cap and the any-indent fence
rule (n2).
m1 -- boundary coverage was 2/3. limit (999999999.) and limit+1 (1234567890.)
were pinned; limit-1 (12345678.) added, per RULESET.TESTS.boundary-coverage.
m2 -- THEMATIC_BREAK_RE and the tab-expanding indent counter are hand-rolled
CommonMark rules with no in-repo peer to compare against, so the parity
assertion is against the SPEC: eight positive and five negative fixtures, plus
the two DELIBERATE divergences pinned as deliberate (`+` is a separator here but
not in CommonMark, because `+` is a list marker in this grammar and `+ + +`
would otherwise be a phantom entry; indent is unbounded). One fixture was
initially wrong -- `-- -` IS a CommonMark break, since the spec allows free
spacing between the three characters -- and the parser was right.
m3 -- the result union is hand-duplicated in audit.cts as part of a deliberate
structural view of uat.cjs, so the fix is not to delete a copy but to make drift
observable. Every REACHABLE status is now driven from a fixture; four of the six
(ambiguous, unsupported_heading_shape, already_resolved, match_verification_failed)
had no assertion anywhere in the suite before this. match_verification_failed is
still undriven and the test says so rather than omitting it.
m5 -- DECLINED, with the measurement. The review is right that `(\s*)` in the
opener and `/^[ \t]*/` in the reader disagree about \f, \v and NBSP, but its
prescribed narrowing was implemented, driven and REVERTED: as shipped, an entry
indented with any of those surfaces, parses its status field, acknowledges, and
reads back acknowledged -- a complete round-trip. Narrowing turns all three into
SILENTLY DROPPED entries, which is the #3702 defect class itself and the opposite
of this file's stated fail-safe rule. A latent inconsistency in the safe
direction is not worth a live regression in the unsafe one. Pinned by three
round-trip tests so the prescription cannot be re-applied silently; if it is ever
closed, the direction is to make the readers agree with the opener, not to make
the opener reject lines it accepts today.
Four deferred-items suites 475/475, 0 skipped. lint:ci and lint:changeset exit 0.
The 47-fixture Gaps differential is byte-identical at 8033 bytes.
* fix(#3702): the pinned `## Gaps` phantom now cites its issue (m4)
Round-4 m4. The second assertion in the Gaps byte-for-byte test pins a real
defect as expected output: a spaced hyphen thematic break in `## Gaps` is read
as an ITEM, so `- - -` surfaces a phantom open gap named `- -`. Reproduced on
pristine next at
|
||
|
|
996196fe02 |
fix(#4099): stop word-splitting an unquoted PLAN_IDS string (bash vs zsh diverge) (#4102)
* fix(review): stop word-splitting an unquoted PLAN_IDS string (bash vs zsh diverge) #3301's plan-coverage-manifest block built a space-separated PLAN_IDS accumulator and re-split it via unquoted `for x in $PLAN_IDS`. Bash word-splits an unquoted scalar on IFS by default; zsh does not, so every plan id collapsed onto one iteration under zsh whenever a review had 2+ plans, undercounting the manifest and merging bullets onto one line. This is exactly what was failing origin/next's CI on every push. Derive PLAN_COUNT and the bullet list directly from the same *-PLAN.md glob loop used for copying — glob iteration, unlike re-splitting a stored string, is identical under bash and zsh, so it never needs word-splitting. Closes #4099 Emitted-Drift-Ack-Growth: review.md — fix comment documents the bash/zsh word-splitting divergence (gsd-core#4099); net code is smaller, comment expansion grows the file * docs(#4099): add changeset for the plan-coverage-manifest zsh fix * docs(#4099): fix changeset body to match the bold-dash-explanation format Code review (standards axis) flagged the changeset body as not matching CONTRIBUTING.md's canonical `**<Bold>** — <explanation>.` format — the bold clause had its own trailing period instead of being followed directly by the em-dash separator. * chore(#4099): backfill changeset PR number (#4102) --------- Co-authored-by: sim <sim@local> |
||
|
|
6beaa66b25 |
enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate Content-assertion suite for the Step 7 re-verification evidence gate (agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md). Committed before the implementation to prove RED via gsd-test. * enhance(#3304): gate re-verification blockers on deterministic evidence Step 7's anti-pattern scan re-runs at full, unbounded scope on every re-verification pass, independent of the must-haves established in Step 2. A blocker it finds — other than the self-evidencing debt-marker check — previously reverted a completed gap-closure round and started another --gaps cycle on nothing more than the verifier's own new judgment call, with no bound on how many times that could repeat. A Step 7 blocker now blocks unconditionally in re-verification mode only if it is a carried-forward gap (present in the prior VERIFICATION.md's gaps: list) or the flagged file was git-modified since the prior pass (a regression; fails closed toward blocking when history is unresolvable). Otherwise it predates the gap-closure round unflagged and needs deterministic evidence — a named test run red, or another concrete reproducible artifact — to stay blocking. Unevidenced, it downgrades to a new advisory: frontmatter list and report section instead of setting status: gaps_found, and never reverts a completed must-have. Maintainer approval was narrowed to this evidence condition only, explicitly rejecting the broader "advisory whenever untraceable to a requirement/decision/prior-gap" proposal — implemented and pinned by tests/verifier-evidence-gate.test.cjs and documented as rejected in gsd-core/references/verifier-evidence-gate.md so it can't silently re-expand. Closes #3304 * fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself (not the production prose): a {0,600} match window was shorter than the 724-char paragraph it was scanning (the "exclude from Step 9 Rule 1" phrase starts at offset 662), and two regexes assumed no indentation after a markdown list-continuation line break. All three phrases are confirmed unique across agents/gsd-verifier.md, so the windowed submatches are replaced with direct whole-string assertions instead of just widening the window. Also acknowledges the deliberate byte growth in agents/gsd-verifier.md that the differential-attribution check (ADR-2719) correctly flagged. Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152). * docs(#3304): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
62b0d939b6 |
feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey Add an optional `timeoutConfigKey` field to the reviewer lane descriptor, resolved in `resolveLanePlan` at invocation time and falling back to the frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes declare `review.timeouts.<slug>` on both surfaces (the descriptor and their capability.json manifest), validated by capability-validator.cjs. For the antigravity lane, the native `agy --print-timeout` flag — previously a second hardcoded literal (`540s`) independent of the outer cap — is now derived from the same resolved outer timeout in `antigravityArgv`, preserving the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is declared data, the inner one is handler-owned). The antigravity default timeoutFloorMs stays at 600s per the maintainer's disposition; users raise it through the new config key instead. * docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper Address code-review findings on the timeoutConfigKey change: extract the inline timeout-resolution logic into a named, exported, directly-tested resolveTimeoutMs helper (matching the file's existing configString/ normalizeHost convention); document the new review.timeouts.* federated config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md, and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment. * fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner gsd-test caught two design mistakes in the prior commits: 1. SpawnPlan.argv is documented and tested as fully resolved by resolveLanePlan (model/effort/output/prompt already folded in) — leaving the antigravity '{{nativeTimeout}}' marker unresolved until the runner's antigravityArgv violated that contract and broke tests that read plan.argv directly (tests/antigravity-reviewer.test.cjs, tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}' is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself via the new nativeTimeoutToken() helper, exactly like the other four. antigravityArgv reverts to its pre-#3274 four-argument form. Also missed updating capabilities/antigravity/capability.json's invoke.args to match the descriptor, which broke the manifest/descriptor parity test. 2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three lanes with neither a model flag nor a host — may own no config key beyond their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes violated it. Fix: those three keep timeoutConfigKey: null and own no review.timeouts.* key, matching their existing modelConfigKey: null. The other 9 lanes are unaffected. * chore(#3274): backfill changeset PR number (pr:0 -> 4083) --------- Co-authored-by: sim <sim@local> |