* fix(#1520): randomize mktemp temp paths on BSD/macOS (XXXXXX must be path-final)
BSD/macOS mktemp only substitutes the XXXXXX template when it is the final
path component. Templates like `...-XXXXXX.json` / `gsd-pr-body.XXXXXX.md`
return a LITERAL `XXXXXX` path (no randomization) on macOS, so concurrent
workflow runs collide on the same temp manifest/body file — one run can
overwrite or consume another's. Reproduced on macOS: the second call to the
suffixed template fails `mkstemp: File exists`.
Fix: use a suffixless `XXXXXX` template (so it IS the final component), then
rename to add the intended extension — portable across BSD + GNU userlands,
no GNU-only `--suffix` flag. Empty-file-then-write semantics are preserved at
every site.
Affected workflow temp files:
- execute-phase.md: gsd-worktree-wave-*.json (wave worktree manifest)
- quick.md: gsd-quick-worktree-*.json
- spec-phase.md: edge-probe-reqs-*.json
- ship.md: gsd-pr-body-*.md
- profile-user.md: gsd-profile-answers-*.json, gsd-profile-analysis-*.json
The execute-phase.md edit uses a compact intermediate var + trailing comment
to stay under the ADR-857 phase-6 size ceiling (93166); regenerated the
workflow size baseline accordingly. Validated on macOS: 20 concurrent calls
yield 20 unique randomized paths.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1520): add changeset fragment (Fixed)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1520): add fail-first workflow-prose guard for mktemp XXXXXX suffix
Repo-wide scan of gsd-core/workflows/**/*.md that fails on any mktemp
template whose XXXXXX run is followed by a filename suffix (the BSD/macOS
non-randomizing form). Fails on the six pre-fix instances and passes on
the fix, and locks the copy-paste-prone idiom out of future workflows.
Mirrors the bug-637 hardcoded-$HOME workflow guard.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1520): rename regression test to fix- prefix (regression-test-names lint)
New tests/bug-NNNN-*.test.cjs files are banned by the lint-regression-test-names
ratchet; use the fix- prefix (matches the fix-1445 precedent).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1520): add issue ref to allow-test-rule exemption (ADR-456 lint)
lint-allow-test-rule-refs requires every new `allow-test-rule:` comment to
carry a #NNN reference (don't allowlist). Add (#1520) to the source-text
exemption.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1520): abort touched mktemp chains on failure (|| exit 1)
Per review: the VAR=$(mktemp …) && mv … && VAR=… chains dropped the issue's
suggested failure guard. If mktemp fails, $VAR is empty and the subsequent
mv/write lands on an unintended relative path. Add `|| exit 1` to all six
touched chains so a mktemp failure aborts the snippet. Regenerated the
workflow size baseline for the slightly longer lines.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1520): rebase onto next — regen size baseline + describe rename
Resolve the workflow-size-baseline.json conflict from next advancing by
regenerating from the current workflow sizes. Also rename the test describe
from `bug #1520` to `#1520` (the file uses the fix- prefix) per review nit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1520): harmonize profile-user mktemp to ${TMPDIR:-/tmp} (review nit)
The two profile-user.md temp sites this PR already rewrites kept a hardcoded
/tmp while the four sibling workflows use ${TMPDIR:-/tmp}. Harmonize for
consistency and macOS-correctness (some sandboxes have no writable /tmp).
Regenerated the workflow size baseline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1520): regen size baseline after rebase onto next
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* feat(#323): fish-shell support in post-install PATH suggestion
Two additive changes to the post-install PATH-suggestion seam, both scoped
to existing functions.
A. Projection: add a fish entry to the persist-mode shell-action list in
projectPathActionProjection() (src/shell-command-projection.cts). fish has
no `export`/`$PATH`-list syntax, so the existing zsh/bash `export PATH=...`
commands are inert when pasted. The new entry emits the fish-native
`fish_add_path '<dir>'` (fish 3.2+, persists via the universal-variable
store, de-duplicating). The directory is single-quoted with the same POSIX
literal escaping as the zsh/bash siblings; verified round-tripping through
real fish 3.7.0 for paths containing quotes, spaces, `$`, `*`, backticks
and unicode.
B. Detection: add homePathCoveredByFishConfig() in bin/install.js, called
from maybeSuggestPathExport() alongside homePathCoveredByRc(). fish does
not use sh-style `export PATH=` rc files, so a fish user whose
fish_user_paths already covers the global bin would otherwise get a
false-positive "not on your PATH" warning on every install. Two
side-effect-free detection routes (no fish subprocess):
1. The universal-variable store (~/.config/fish/fish_variables). fish
serializes this with `full_escape`: every byte outside [A-Za-z0-9/_]
becomes `\xHH` (space -> \x20, `-` -> \x2d, `.` -> \x2e, `$` -> \x24,
unicode -> \uXXXX) and list elements are joined by the literal 4-char
token `\x1e` (NOT a raw 0x1e byte). The detector splits on `\x1e`,
decodes the escapes, then compares each as an absolute literal — a
decoded `$` is part of the directory name, not an unexpanded variable.
Verified against real fish 3.7.0 output.
2. config.fish (`fish_add_path`, `set -gx PATH`, `set -Ux fish_user_paths`)
— plain shell tokens: HOME forms ($HOME/${HOME}/~) are expanded and a
token still holding `$` (e.g. `$PATH`, `$fish_user_paths`) is skipped.
Honours $XDG_CONFIG_HOME and always also checks ~/.config/fish.
No behaviour change for bash/zsh/PowerShell/cmd/Git-Bash users: their entries
and command strings are unchanged; the fish entry is additive and the fish
detector only narrows the set of cases that warn.
Tests: update the projection length assertion (2 -> 3) and fish escaping in
bug-3441; add fish detection + suppression cases in install-path-detection
(uvar store with real fish escaping, dot/hyphen/space/$-literal decode
regressions, config.fish routes, commented-out, relative-segment guard,
unreadable-file fault injection, suppression and emission via
maybeSuggestPathExport).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(changeset): add Changed fragment for #323 fish PATH support (#727)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#323): address review — action-only fish docs, decoder property test, win32 guard
Addresses @trek-e's review on #727:
- docs (blocker): keep the how-to action-only (Diátaxis). Drop the
`# fish — persists via …` comment and the internal-mechanism clause
naming fish_variables/config.fish; leave one command + the exec-fish
directive.
- tests (minor): extract decodeFishUniversalValue to a pure, exported
module function and add fast-check round-trip properties
(decode(fishEscape(p)) === p over arbitrary unicode, abs-path variant,
totality). Consolidated into install-path-detection.test.cjs to respect
the install test-file-count ratchet.
- tests (follow-up): port #721's win32 negative-projection test (no fish
action on win32; persist projection is PowerShell/cmd.exe/Git Bash).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#323): address review — drop unused 'after' import, clarify escaping comment
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1577): isolate WebFetch/WebSearch ingress + opt-in injection blocking
Split A of #1573 (security-critical). Scans WebFetch/WebSearch output (the
largest untrusted channel) in gsd-read-injection-scanner; shared
untrusted-input-boundary reference @-included by the 8 ingest agents
(randomized per-wrap delimiters, in-prompt self-scan guard, task-anchoring);
opt-in security.injection_blocking (default advisory — non-breaking).
arXiv: 2506.05739 (PPA), 2507.15219 (PromptArmor), 2504.20472 (Referencing), 2503.00061 (defense-in-depth).
* fix(#1577): address review — honest blocking docs, config key, ADR, property test, revert localized
- A1: rewrote the opt-in-blocking doc + Security changeset honestly — the PostToolUse hook is a
circuit-breaker (halts the agent's next step), NOT a redactor; it does not scrub content already
in the transcript. The prompt-level data/instruction boundary is the primary control.
- A2: registered security.injection_blocking in the config schema + defaults manifests (default
false) + an e2e config-roundtrip test; the dotted setter writes the nested shape the hook reads.
- A3: reverted the 4 hand-edited localized security-model.md (canonical EN only, per convention).
- A5: ADR-1577 (untrusted-input boundary + opt-in blocking; redaction-vs-circuit-breaker rationale).
- A6: property test — scanner never crashes / only emits valid JSON on unicode/large/malformed input.
- Also: inventory (untrusted-input-boundary.md) + agent-size baseline (8 ingest agents) +
drift-guard matcher update (Read -> Read|WebFetch|WebSearch). A7 (content<20 early-exit) left as
the noted pre-existing follow-up.
* fix(#1577): allowlist untrusted-input-boundary.md in injection-scan CI gate
The new reference quotes injection phrases ('ignore previous instructions',
'you are now…') as examples agents must NOT comply with, tripping the repo's
own prompt-injection-scan.sh diff gate (the standalone 'security' CI job, red
on HEAD). Allowlist it alongside the other security docs (security-model.md,
TEST-EXAMPLES.md) that legitimately demonstrate injection patterns. The JS
scanner test doesn't scan references/, so only the shell gate needed it.
Verified: scan --diff origin/next -> 0 findings; scanner JS test 15/15.
* fix(#1577): cover AC #2's gsd-ui-researcher + gsd-assumptions-analyzer
trek-e Major 1: the @-included set dropped two AC #2 agents. Restore them so
no named web-ingress agent is uncovered, keeping the two justified additions
(gsd-ai-researcher, gsd-domain-researcher). Final set = AC's 8 + 2 = 10.
- gsd-ui-researcher carries the full WebSearch/WebFetch + MCP-fetch toolset.
- gsd-assumptions-analyzer reads 5-15 codebase source files (external/source-
document ingress per the boundary), though it has no web tools.
INGEST_AGENTS in the isolation test now asserts all 10; size baselines
regenerated (+60 bytes each, both well under the DEFAULT cap); changeset
reworded 8 -> 10.
Verified: untrusted-input-isolation 14/14; agent-size-budget 39/39.
* docs(#1577): document security.injection_blocking + boundary seam
trek-e Major 2 + Minor:
- docs/CONFIGURATION.md: add the top-level security.injection_blocking key to
the Full Schema and a Security Settings subsection, distinguishing it from
the workflow.security_* namespace; honest circuit-breaker-not-redactor
framing matching ADR-1577 / security-model.
- CONTEXT.md: add the 'Untrusted-input boundary' seam glossary entry.
Verified: lint:docs ok; config-field-docs + contributor-standards green.
* test(#1577): make read-injection property test git-text, not binary
trek-e nit (and more): the file embedded a raw U+FFFF AND a raw NUL byte as
degenerate-edge inputs. The NUL is what actually made git classify it binary
(git binary = NUL in first 8K). Replace both with text-safe escapes that keep
the identical runtime values: '\\x00' and String.fromCodePoint(0xFFFF). File
now diffs/blames line-by-line.
Verified: property test 2/2; no NUL/raw-noncharacter bytes remain.
* docs(#1577): align untrusted boundary docs
Name all 10 ingress agents in INVENTORY/security-model and allowlist the intentional read-injection property corpus for the prompt-injection scanner.
* docs(#1577): align ADR ingest agent count
Update ADR-1577 from 8 to 10 ingest agents so it matches the actual boundary include set and the rest of the docs.
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* fix(#1571): resolve schema-drift phase by token, not substring
verify schema-drift <phase> resolved the phase directory with a naive
entry.name.includes(phaseArg) test, so a non-existent phase could
silently match a different phase whose directory name merely contained
the requested token (e.g. "1" matched "11-expansion"), running the drift
gate against the wrong phase. Use the canonical phaseTokenMatches +
normalizePhaseName, matching find-phase, verify phase-completeness, and
this file's own unstarted-phase check.
Regression coverage folded into tests/schema-drift.test.cjs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1571): add changeset for schema-drift token-match fix
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(#1516): expose adaptive model_profile in /gsd-new-project AI Models prompt
Both onboarding paths (Step 2a auto-mode + Step 5 interactive) enumerated only
4 profiles (Balanced/Quality/Budget/Inherit), omitting 'adaptive' even though the
model catalog (model-catalog.json profiles) and docs/CONFIGURATION.md register 5.
Mirrors the proven /gsd:settings two-question split (#3784): Q1 routes between
Adaptive/Standard-tier/Inherit; Q2 (conditional on Q1=Standard) picks
Quality/Balanced/Budget — keeping every AskUserQuestion within the 4-option cap.
Both config-new-project example payloads now list adaptive. Regression cases
folded into the owning tests/new-project-mvp-prompt.test.cjs (per the
lint-regression-test-names ban on new top-level bug-NNNN files): each AI Models
prompt makes adaptive reachable, all 5 profiles reachable, 4-option cap honored,
both example enums include adaptive, brace balance. Workflow size baseline bumped
(new-project.md 62324 -> 66138 bytes; still well under the XL hard cap).
* chore(#1516): backfill changeset pr ref to 1654
* fix(#1659): dedup By-Phase rows across padded/unpadded phase numbers
phaseRowPattern matched the phase number literally (escapeRegex(String(phaseNum))),
so a seeded zero-padded row '| 05 |' was not matched by 'phase complete 5' (pattern '| 5 |'),
producing a duplicate row that double-counted the phase. Canonicalize a numeric phase to
its integer form (Number('05')===Number('5')===5) and match with a 0* prefix so 5/05/005
all collapse to the same row in either direction. Regression folded into state.test.cjs:
seeded '| 05 |' + 'phase complete 5' yields exactly one phase-5 row. Non-numeric phase IDs
retain the literal escapeRegex match.
* chore(#1659): backfill changeset pr ref to 1663
* fix(#1659): add verification fixture to padded-dedup test under #1522 gate
* fix(#1582): derive phase-complete velocity from By-Phase table (idempotent)
updatePerformanceMetricsSection blind-added summaryCount onto the prior velocity
total on every phase complete, so re-running phase complete on an already-complete
phase incremented the total each time (the sibling of #4, which fixed the Completed
Phases counter the same way). The velocity total is now derived as the sum of the
By-Phase table's Plans column AFTER the row upsert — re-completing a phase upserts
the same row, so the sum is stable; a hand-edited inflated total self-heals downward
to the true sum on the next completion. When the By-Phase table is absent the total
is left unchanged (no crash). Strengthens the misnamed 'idempotent' test (its comment
explicitly declined to assert velocity idempotency — the latent gap) and adds a
self-heal regression; corrects the #320 behavior-lock velocity assertion which had
encoded the blind-add (3 = 1+2 double-count) — the derived value is 2.
* chore(#1582): backfill changeset pr ref to 1655
* fix(#1582): velocity sum tolerates indented By-Phase rows (codex review)
Adversarial review (codex, gpt-5.5/high) flagged that byPhaseTablePattern's
data-row capture allows leading whitespace ([ \t]*\|), but the derive sum was
anchored at ^\| and would skip indented hand-edited/legacy rows — capturing them
in the table but silently undercounting. Align the sum regex (^\s*\|) with the
table capture's tolerance. Adds an indented-row regression. Two other codex
findings are pre-existing and out of scope: padded/unpadded phase dedup
(phaseRowPattern, identical in old code — derive yields the same value as the old
blind-add) and CRLF tables (the shared byPhaseTablePattern header requires bare
\n, so the upsert was already broken on CRLF; the fix changes stale-vs-
double-count, does not worsen it).
* fix(#1582): add verification fixtures to velocity tests under #1522 gate
Post-rebase onto next+#1548, the #1582 velocity tests (self-heal, indented-row) use
phase complete, which now fail-closes under #1522's canonical verification gate without a
passed *-VERIFICATION.md. Add writePassedVerification(tmpDir,'02-next','02') to both.
* fix(#1658): make byPhaseTablePattern CRLF-tolerant on STATE.md tables
byPhaseTablePattern required a bare \n after the header and separator rows, so a
STATE.md with CRLF (\r\n) line endings (Windows, or hand-edited) had its By-Phase
table treated as absent: phase complete never upserted the row (and the velocity-from-
table derivation went stale). Make the header/separator terminators and the closing
lookahead CRLF-tolerant ([ \t]*\r?\n, (?=\r?\n|$)). Backward-compatible with LF.
Regression folded into tests/state.test.cjs: phase complete on a CRLF STATE.md upserts
the row and removes the placeholder. CONTRIBUTING's QA matrix lists Mixed CRLF/LF as a
required parser case.
* chore(#1658): backfill changeset pr ref to 1662
* fix(#1657): recover malformed (non-object) ~/.gsd/defaults.json in finishInstall
JSON.parse of defaults.json succeeds for valid-JSON-but-non-object values (null, [],
42, "str"), which then bypassed the parse catch: null threw a TypeError on property
access (swallowed by the outer try/catch), and array/number/string had resolve_model_ids
set on a non-object whose JSON.stringify round-trip kept the broken shape. The non-Claude
finishInstall step now resets any non-object (null, non-object, or array) parse result to
{} before reading/writing, so the file is repaired and resolve_model_ids defaults normally.
Regression folded into the owning tests/bug-410-install-defaults-test-mode-guard.test.cjs
(parameterized over null/[]/42/"str").
* chore(#1657): backfill changeset pr ref to 1661
* fix(#1660): fail-closed frontmatter set of object-list fields instead of silent no-op
cmdFrontmatterSet reported {updated:true} even when spliceFrontmatter returned the
content unchanged, which happened whenever the new value's extractFrontmatter projection
equalled the original's — notably for object-list fields like must_haves, whose
{path,provides} items flatten to scalar strings under the lossy parser. Detect a no-op
(newContent === content) for a dict-valued field and surface an error directing the user
to edit the file directly, instead of silently accepting a no-op set. Scalars and scalar
arrays round-trip faithfully, so idempotent sets of those are intentionally NOT flagged
(two precision regression tests lock this). Folded into frontmatter-cli.test.cjs.
* chore(#1660): backfill changeset pr ref to 1664
* refactor(#1660): extract noOpObjectListSetError as pure tested helper (Stryker coverage)
cmdFrontmatterSet is not in Stryker's property/unit test set, so the inline no-op
detection added survivors that dropped the frontmatter module below its 62% mutation
threshold. Extract the detection into a pure exported helper noOpObjectListSetError and
unit-test every branch directly (changed content, scalar, scalar-array, null, dict
no-op). cmdFrontmatterSet now calls the helper. Same pattern as the #1572 spliceFrontmatter
coverage fix.
* fix(#1572): preserve must_haves object-lists across frontmatter set/merge
spliceFrontmatter round-tripped the WHOLE frontmatter through extractFrontmatter
(a scalar-only parser) then reconstructFrontmatter (a lossy serializer), so any
must_haves object-list — artifacts {path, provides}, prohibitions {statement,
status} — was flattened to scalar strings and re-emitted as a malformed inline
array whenever an UNRELATED field changed, silently dropping every provides:/
status: value. The write now preserves the original raw text for any top-level
key whose value is structurally unchanged between the original parse and the new
object (generalizing the existing whole-document no-op guard to per-key
fidelity), and regenerates only the key that actually changed. The key set is
still defined by newObj (the cmdSet/cmdMerge flow always passes the full merged
object). spliceFrontmatter's only callers are cmdFrontmatterSet/Merge — the
STATE.md read-modify-write family calls reconstructFrontmatter directly and is
unaffected. Regression cases folded into tests/frontmatter-cli.test.cjs:
artifacts/prohibitions object-lists survive set and merge; idempotent on repeat
sets. Asserted via parseMustHavesBlock (the structure-preserving parser).
* chore(#1572): backfill changeset pr ref to 1656
* fix(#1572): fail-closed when set/merge would emit [object Object] (codex review)
Adversarial review (codex, gpt-5.5/high) flagged that directly setting a must_haves
object-list (a CHANGED key) still routed through the lossy reconstructFrontmatter,
emitting literal "[object Object]" and destroying the data. The reported case
(mutating an UNRELATED field) was already fixed by per-key raw-text preservation,
but the changed-object-list path was still silently lossy. Add fail-closed: when a
regenerated key's text contains the "[object Object]" sentinel, spliceFrontmatter
throws — cmdFrontmatterSet/Merge error out WITHOUT writing, directing the user to
edit the file directly. The no-frontmatter (generate-from-scratch) path is guarded
the same way. Adds a test that a refused set leaves the file unchanged and the
original object-list intact. Codex finding #2 (a contrived flattened-projection
no-op) is a deeper limitation noted in the PR — non-destructive, and the fail-closed
message already directs users to edit object-list blocks directly.
* test(#1572): add spliceFrontmatter per-key preservation + fail-closed unit coverage
Stryker mutates gsd-core/bin/lib/frontmatter.cjs against tests/frontmatter.{property,unit}.test.cjs
(MinScore 62). The #1572 regression cases live in frontmatter-cli.test.cjs, which is NOT in
Stryker's test set, so the new functions (sliceTopLevelFrontmatterSegments, the per-key
preserve/regenerate/drop/append loop, regenerateFrontmatterKey's [object Object] fail-closed)
had surviving mutants that dropped the module below threshold. Add unit-level coverage in
frontmatter.unit.test.cjs exercising every new branch directly via spliceFrontmatter:
unchanged object-list preserved (provides survives) when a scalar sibling changes; changed
scalar regenerates only that key; orphan keys dropped; new keys appended; indented nested
block stays attached to its parent key; whole-document no-op returns input verbatim; both
fail-closed paths (changed object-list + no-frontmatter) throw.
* fix(#1639): parseDecisions handles titled-colon bullet form
bulletColonRe anchors on ':**' (colon immediately before close-bold) and bulletEmDashRe
requires an em-dash, so the titled-colon form '- **D-NN: Title.** body' (title between the
colon and the closing **) matched neither and was dropped by the parse-miss guard. When all
decisions used the titled convention, parseDecisions returned 0 and check.decision-coverage-
plan passed vacuously — the same false-coverage failure mode as #1343/#1364/#1365. Add a
third per-form regex bulletTitledColonRe, checked LAST (strict superset of bulletColonRe,
so it only catches bullets the other two miss — minimal blast radius); id + [tags]
trackability honored. Regression folded into decisions.test.cjs: titled-colon parses,
coexists with colon/em-dash, tags, all-titled-13 no longer vacuously 0.
* fix(#1639): tighten titled-colon title to [^:*]* so malformed pre-colon-run bullets still reject
The first cut's title run [^*]* was too permissive: it matched a genuinely-malformed
bullet with a colon in the pre-separator freeform run (e.g. 'D-07 ratio 3:1:**') by
treating the 3:1 colon as the separator, regressing the #1343 parse-miss guard tests.
Tighten the title to [^:*]* (no colon, no star) so the separator colon remains the only
colon permitted before ** — matching bulletColonRe's existing [^:*]* discipline. Valid
titled forms (colon-free titles) still parse; the malformed colon-in-freeform case still
falls through to the parse-miss guard.
* chore(#1639): backfill changeset pr ref to 1665
* fix(#1569): preserve explicit resolve_model_ids in non-Claude installs
The non-Claude finishInstall step keyed its resolve_model_ids:"omit" write on
!== "omit", so an explicit true opt-in (resolveModelInternal returns full model
IDs) was silently clobbered on every install/upgrade across all 14 non-Claude
runtimes, making generated agent manifests inherit the active chat model instead
of pinning the resolved model. Now only absent/falsy is defaulted to "omit"; an
explicit true (and an existing "omit") is preserved. Regression test
parameterizes across codex/opencode/gemini and covers the absent/false/idempotent/
claude/malformed boundaries.
* chore(#1569): backfill changeset pr ref to 1653
* fix(#1569): default non-canonical resolve_model_ids values to omit (codex review)
Adversarial review (codex, gpt-5.5/high) flagged that the original allowlist-by-
enumeration condition (undefined/null/false -> omit) preserved malformed values
(0, "", "yes", {}) instead of defaulting them to omit, letting them leak Claude
aliases a non-Claude runtime cannot resolve. Switch to an allowlist condition
(existing !== true && existing !== 'omit') so only an explicit canonical true
opt-in and an existing omit are preserved; everything else defaults to the safe
non-Claude omit. Adds a parameterized test over [0, "", "yes", {}].
* fix: require fresh phase verification before transition
* no-mistakes(review): Fix canonical verification closeout gates
* no-mistakes(review): Fix verify-work frontmatter promotion command
* no-mistakes(review): Fix stale verification gates
* no-mistakes(review): Fix canonical verification routing gates
* no-mistakes(review): Fix verification dependency and runtime routing gates
* no-mistakes(review): Block stale verification bypasses
* fix: handle large init manager outputs in verification workflows
* chore: update changeset pr number
* fix(verify-work): use fresh verification.status for stale gate
The stale check after UAT used phase_completion.verification_status from
session-start INIT while human_needed promotion already queried fresh
verification.status. Align the stale gate with the canonical query so
mid-session verification refresh is not ignored.
* fix(init): skip roadmap-checked phases when selecting next_phase
Roadmap-only phases without a disk directory were still promoted to
next_phase when their checkbox was already checked. Exclude
checkboxComplete phases so progress routing does not point at work the
roadmap already marks done.
* fix: gaps_found not overridden by stale, transition uses canonical verification
- verification.cts: check gaps_found before stale so gap-closure routing
is not masked by a newer summary mtime
- phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus
already handles stale detection
- transition.md: replace raw grep on file content with verification.status query
to avoid false-positive blocks from body text matching
* ci: retrigger tests after rebase
* fix(transition): replace gsd_run advisory check with awk frontmatter extraction
The runtime launcher is not defined until the update_roadmap_and_state step
bash block (~line 165). The early verify_completion block used gsd_run to
query verification.status, which violated the runtime-launcher-parity test:
'preamble appears AFTER the first gsd_run reference'.
Replace the gsd_run call with an awk-based frontmatter extractor that reads
only the status: field between the two --- fences. This avoids both the
preamble-ordering constraint and the original false-positive grep bug where
body text like 'previous_status: gaps_found' would match a full-text regex.
The phase.complete gate at update_roadmap_and_state is the canonical
enforcement point; this early check is advisory only.
Also update workflow-size-baseline.json for the updated transition.md size.
Fixes: runtime-launcher-parity test (B)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: re-check verification under planning lock in phase complete
Move readVerificationStatus into withPlanningLock so stale verification
cannot slip through when a SUMMARY.md is written between the gate and
the roadmap/state mutation. Return the blocked status from the lock
callback and emit the error after release to avoid leaving .lock behind.
* fix(transition): gate on canonical verification.status including stale
Replace awk frontmatter read with verification.status query so transition
blocks when summaries are newer than VERIFICATION.md, matching phase.complete
and other workflows (autonomous, progress, verify-work).
* Fix workflow verification gates for yolo transition and stale routing
Require VERIFY_STATUS passed before yolo/interactive transition advance.
Route stale verification recovery to verify-work, matching canonical projection.
* fix(transition): use verification.status query for stale-aware advisory check
The awk-based check read raw frontmatter status: passed, which misses the
stale case where summaries are newer than the VERIFICATION.md file even
though the frontmatter still says passed. The stale status is computed from
file modification times, not stored in frontmatter.
Move the preamble to the verify_completion bash block (the first block with
a gsd_run call) so gsd_run query verification.status can be used for the
advisory check. This gives the full readVerificationStatus logic including
mtime-based staleness detection, matching the enforcement gate at phase.complete.
Capture full JSON (VERIFY_JSON) so next_action can be included in the
advisory output alongside the status.
Also update workflow-size-baseline.json for the updated transition.md size.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* ci: trigger test matrix for 525b946
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix(transition): restore awk frontmatter extraction for pre-shim verification check
The gsd_run launcher shim is not defined until line ~163 of transition.md,
so the verification debt check at line ~80 cannot use gsd_run. Restore the
awk-based frontmatter extraction that correctly reads status without needing
the runtime, and restore the shim at its proper location before
phase.complete.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(#1522): clarify transition verification gate wording
* fix(#1522): update transition workflow size baseline
* fix(#1522): update workflow-size-baseline after rebase onto next
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review)
Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync
and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections,
so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any
FS error threw uncaught into callers NOT under the planning lock (init.manager /
init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and
thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller)
for parity with readVerificationStatus's no-throw contract and testability. Also adds the
Verification Module glossary entry to CONTEXT.md (review B3).
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* refactor(#1646): route capability routers through Command Routing Hub per ADR-959
Phase 2 of parent #1641. Converts graphify, intel, and audit command
routers from hand-rolled if/else dispatch to routeHubCommandFamily,
implementing the ADR-959 §III(B) line 75 mandate. The three routers
now share the uniform dispatch shape with the 14 host routers.
src/cjs-command-router-adapter.cts
* Imported ERROR_REASON from io.cjs.
* UnknownCommand translation now passes ERROR_REASON.SDK_UNKNOWN_COMMAND
as the second arg to error() — additive for host routers (their
existing one-arg error callbacks ignore the second arg), required
for capability routers whose tests assert reason === 'sdk_unknown_command'
on the JSON-error envelope.
src/graphify-command-router.cts
* Replaced 4-branch if/else with routeHubCommandFamily + handlers map.
* Validation handlers (missing term, missing/invalid --budget) now
return makeInvalidArgs(arg, reason, ERROR_REASON.USAGE) Results
instead of calling error() directly (Q2=C, Q4=ii from grilling).
* Success handlers keep direct output() calls.
* Subcommands array is alphabetical for byte-identical 'Available:'
text in the unknown-subcommand message.
* The unknown-subcommand path is now owned by the Hub's manifest
check (the adapter passes SDK_UNKNOWN_COMMAND).
src/intel-command-router.cts
* Replaced 9-branch if/else with routeHubCommandFamily + handlers map.
* Validation handlers (missing term, missing filePath for patch-meta
and extract-exports) return makeInvalidArgs Results.
* Preserved the timeAgo mutation in the non-raw status handler.
* Preserved the lazy require('./intel.cjs') inside the route function.
src/audit-command-router.cts
* routeAuditUat: routes through the Hub with a synthetic 'run'
defaultSubcommand (no real subcommands). Gives uniform observability.
* routeAuditOpen: captures --json in a closure, strips it from args
before Hub dispatch (so it isn't mistaken for a subcommand by the
manifest check), then branches on wantJson inside the handler to
preserve the formatAuditReport success-path quirk.
docs/CONFIGURATION.md
* Observability section: noted capability commands (graphify, intel,
audit-uat, audit-open) now emit DispatchEvent records since #1646.
.changeset/capability-routers-via-hub.md
* Changed fragment describing the user-visible audit-trail expansion.
pr:0 placeholder will be backfilled after gh pr create returns the
real PR number (DEFECT.CHANGESET-PR-FIELD-DRIFT).
Verification
* graphify cutover tests: 119/119 pass (all unit, dispatch, behavior,
error path, JSON-errors, and registry assertions)
* intel cutover tests: 39/39 pass
* audit cutover tests: 24/24 pass
* bug-974-graphify-budget-missing-value regression test: pass
* npm run test:unit (full suite): 2384 tests, 0 fail
* gsd-test-summary on docker: outcome=passed, 0 failures
(RULESET.PR-FLOW.docker-before-push)
JSON-error envelope parity verified byte-identical: reason values
('usage', 'sdk_unknown_command') and message texts are preserved
across all three routers' error paths.
* chore(#1646): backfill changeset pr: 1647 (DEFECT.CHANGESET-PR-FIELD-DRIFT)
* fix(#1634): honor capability hook matcher and node-prefix command
Capability hook install (applyCapabilitySharedEdits) wrote each settings.json
hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a
fail-closed guard could then block the whole session), and emitted a bare
single-quoted script path so a .js-family hook from a git/tarball source
without +x failed with Permission denied on every matching call.
- Pass through an optional declared `matcher` (entry-level sibling of `hooks`);
absent => omitted (match-all), so existing shipped capabilities are unchanged.
- Validate `matcher` in the declaration (non-empty string, no control chars).
- Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party);
.sh and others keep the bare quoted path (unchanged).
Root cause: the manifest hook schema (validator rule C4) was {event, script}
only with no matcher, and applyCapabilitySharedEdits never read or wrote one;
the command used shellSingleQuote(absScript) with no node prefix.
Regression tests fail-first on both defects (matcher dropped; bare path) and
pass after the fix; #1460 command assertions updated for the node prefix.
* chore(#1634): backfill changeset pr:1638
* fix(#1634): resolve lint and windows CI failures
- validator: replace the control-character range regex with a char-code loop.
The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char
codes are equally precise and lint-clean. Behavior unchanged (still rejects
matchers containing ASCII control characters incl. DEL).
- test: gate the executable-bit precondition on POSIX. Windows fs does not
honor POSIX write modes (a 0o644 write reads back as 0o666), so the
precondition is meaningless there and failed the windows-latest lane. The
node-prefix assertion — the actual fix — is platform-independent and still
runs everywhere.
* docs(#1634): amend ADR-894 for optional lifecycle hook matcher
The `role: "feature"` `hooks[]` entry now carries an optional `matcher`
(settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document
the field in the §2 schema table and record a Grilling-amendments entry:
the install path projects a declared matcher onto the emitted settings.json
hook entry (absent = match-all, so shipped capabilities are unchanged), and
per-runtime matcher projection (ADR-857 D8) stays a separate concern. This
amendment ships with the fix that introduced the field rather than as a
follow-up.
* docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md
Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a
test that writes a file with a POSIX mode and then asserts statSync().mode
& 0o777 === <octal> passes on macOS/Linux but fails on windows-latest
(Windows fs does not honor POSIX write modes — reads back 0o666). Added as a
machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/
prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward:
gate the mode-bit precondition on process.platform !== 'win32' and keep the
platform-independent behavioral assertion running everywhere.
* fix(#1619): normalize pruned mise node execPath to the stable shim
resolveNodeRunner() bakes process.execPath into managed .js hook commands.
Node realpaths execPath, so under mise it resolves to a concrete
<data>/installs/node/<ver>/bin/node that mise prunes on `mise up`, after
which every managed hook 404s — the same ephemeral-path failure #977 fixed
for fnm and #3181 for Homebrew. normalizeNodePath now rewrites a mise
versioned install path to the stable sibling shim <data>/shims/node when it
exists (deriving <data> from execPath so a custom MISE_DATA_DIR works),
falling back to the raw execPath otherwise. Tests folded into
install.test.cjs per the regression test-name lint.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(changeset): set pr number to 1621
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Joe Seymour <joese@iarx.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
workflow.security_block_on was documented as the minimum threat severity
that blocks advancement, but threats carried no severity and the auditor's
threats_open count (the SECURITY.md gate field) counted every open threat
regardless of severity — so the threshold had no effect, and the auditor's
block_on vocabulary (open/unregistered/none) did not even match the config
enum (critical/high/medium/low/none).
- planner: add a Severity column to the STRIDE threat register; assign
severity per threat.
- auditor: read severity; reconcile the <config> block_on domain to the
severity enum; redefine threats_open as the count of OPEN threats whose
severity is at or above block_on (none => 0). Below-threshold opens are
reported as non-blocking and excluded from threats_open.
- SECURITY.md template + planning-config.md reconciled.
No gate-check site changed: threats_open == 0 stays the gate everywhere;
only its computation is now severity-filtered.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
workflow.security_asvs_level was display-only — the planner hardcoded
'mitigate if ASVS L1 requires it' and the auditor only echoed the level,
so L2/L3 behaved identically to L1.
- New reference gsd-core/references/security-asvs-levels.md defines L1
(opportunistic), L2 (standard), L3 (comprehensive) for both planner
threat disposition and auditor verification depth (higher = superset).
- planner: disposition now scales with the configured ASVS level (no
hardcoded L1) + @-pointer to the reference.
- auditor: verification depth scales with asvs_level (L1 grep-presence,
L2 boundary/vector check, L3 end-to-end trace + bypass check).
- planning-config.md + INVENTORY updated; planner kept under its 48K cap
by extracting the goal-backward worked example to planner-guidance.md.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Three related defects in cmdConfigSet, all 'config-set stores invalid values silently':
1. Missing guards: workflow.security_block_on (enum) and
workflow.security_asvs_level (integer 1-3) had no store-time validation.
2. Systemic JSON-coercion bypass: every string-enum guard used
VALID_X.includes(String(parsedValue)). Because the value is JSON-parsed
before validation, String(["member"]) === "member" let a JSON array
slip through and an array was stored in a scalar key. Reproduced on
human_verify_mode, statusline.context_position, context_guard_mode,
fallow.scope/profile, source_grounding_authority, drift_action, context.
3. Unvalidated capability keys: 32 capability-registry-owned keys (4 enum,
25 boolean, 2 number, 1 string) had no hardcoded guard, so any value —
including coerced arrays/objects and out-of-enum strings like
code_review_depth=garbage — was stored silently.
Fix: a type-safe assertEnumValue() helper (requires typeof === 'string'
before membership), routed through all nine central string-enum guards
(messages preserved byte-for-byte); plus a generic capability-registry
validation block that validates every capability key against its declared
type/values (enum via the registry's values — single source of truth —
boolean, number, string). Behavioral regression tests cover every central
enum key and representative capability keys (array + object coercion
rejected, out-of-enum rejected, valid accepted) with boundary coverage for
the security keys.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Pre-#1615 Windsurf installs wrote skills under .devin/skills/gsd-*/ (Devin Desktop preferred dir, #1085). PR #1615 moved Windsurf to .windsurf/workflows/ but never cleaned up the old layout. Users upgrading from a pre-#1615 install were left with dead .devin/skills/gsd-* directories that nothing reads anymore.
Fix: added cleanupWindsurfLegacyDevinSkills() which mirrors the Codex cleanupCodexSkillMetadataSidecars() pattern. Runs on Windsurf local install, removes GSD-managed .devin/skills/gsd-* dirs, preserves user content (non-gsd- dirs, gsd-dev-preferences per #2973, symlinks). Empty .devin/ and .devin/skills/ containers are pruned; non-empty ones are left intact.
5 regression tests: removes gsd-* dirs; preserves user content; skips symlinks (escape guard); no-op when absent; end-to-end install removes pre-staged legacy artifacts.
Refs #1629 (Finding B; Finding A addressed in #1630).
Add a non-blocking, warn-only codebase-drift gate at plan:pre so a stale
STRUCTURE.md is surfaced before /gsd:plan-phase spawns the planner, instead
of being discovered mid-execution by the existing execute:wave:post gate.
Gated on a dedicated workflow.plan_drift_precheck toggle (default true),
independent of schema_drift_gate. Never blocks planning, never spawns the
mapper agent at plan time.
Review feedback (#1595):
- Use a documented conventional-commit type (feat, not enhance) per
CONTRIBUTING.md / gsd-validate-commit.sh.
- Normalize the plan_drift_precheck command references to the colon prose
form (/gsd:plan-phase, /gsd:map-codebase) to match plan-phase.md §5.65;
registry regenerated from capability.json.
- Make the test temp dirs hermetic: drain mkdtemp dirs in an after() hook
via the helpers.cleanup() budget (local/no-raw-rmsync-in-tests-compliant).
Closes#1592
Claude-Session: https://claude.ai/code/session_016JBiXEAofvB3prJim29XMS
Phase B-provide of epic #1258. Adds a build-generated skills/ dir +
a skills manifest field so plugin-installed GSD exposes gsd-core:<skill>
the native Claude Code way. Closes the gap where plugin-only installs
lacked the skill surface because bin/install.js never ran.
- scripts/gen-plugin-skills.cjs: build step converting commands/gsd/*.md
to skills/gsd-<stem>/SKILL.md via convertClaudeCommandToClaudeSkill
- .claude-plugin/plugin.json: add "skills": "./skills/"
- package.json: add skills to files, gen:plugin-skills to build chain
- tests/issue-766-plugin-manifest.test.cjs: Section H conformance
(manifest field + dir + frontmatter + count parity) + C2 skills symlink
- docs/adr/766-*.md: dated amendment adding skills surface row
- .changeset/rapid-bears-hum.md: type Added
- skills/: 69 generated gsd-<stem>/SKILL.md files (build-committed)
Closes#1596
GitHub Copilot reads repository-wide instructions only from
.github/copilot-instructions.md (confirmed via GitHub Docs), not a root
copilot-instructions.md. Aligns getProjectInstructionFile with the installer
(runtime-config-adapter-registry installSurface 'copilot-instructions') and
cites the docs source in the doc-comment.
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit
Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed),
enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way
to browse or audit parked seeds on demand. This adds a read-only listing,
following the established --list → workflow pattern (per the approved scope on
- gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the
seeds dir, returns { count, seeds[], summary } JSON with each seed's id,
slug, status, scope, trigger_when, planted, title. Optional case-insensitive
status filter. User-controlled content is sanitized (sanitizeForDisplay) and
every path validated (requireSafePath); read-only. Independent of
audit.scanSeeds, which only returns unimplemented seeds for the milestone surface.
- /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that
renders the seed table.
Closes#441
* chore(#441): point changeset fragment at PR #722
* test(#441): allowlist list-seeds test in prompt-injection scan
The test asserts that list-seeds neutralizes injection payloads
(<system>, [INST]) embedded in seed content, so the fixtures legitimately
contain those patterns — same as the sibling security tests already on the
allowlist.
* fix(#441): use canonical /gsd:capture colon form in list-seeds workflow
Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use
the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is
retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new
list-seeds workflow used the hyphen form.
* docs(#441): sync help full.md + INVENTORY for --list-seeds
Adds the --list-seeds entry to the help reference (help/modes/full.md, per
bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow
in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json.
* docs(#441): add --list-seeds how-to + drop phantom statuses
Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers):
- USER-GUIDE.md Seeds section (how-to): extend the task to cover
auditing parked seeds on demand via --list-seeds, including the
status filter — kept task-oriented per Diataxis how-to mode.
- CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected
from the list-seeds filter vocabulary; the system only produces
dormant|active|triggered (src/audit.cts scanSeeds). Reference must
be factually accurate and complete.
* fix(#441): guard non-scalar status frontmatter in cmdListSeeds
A seed with a bare `status:` line (extractFrontmatter yields {}) or a
`status: [a, b]` value (yields an array) crashed the whole audit list:
`(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string.
Coerce every frontmatter read through a `fmStr` helper (mirrors the existing
`typeof fm.id === 'string'` guard), so a non-scalar status falls back to
dormant and non-scalar scope/trigger_when/title can no longer leak a raw
array/object into the JSON contract. Title is now capped symmetrically.
Adds regression coverage for empty and array `status:` and non-scalar fields.
Refs #441
* docs(#441): align list-seeds workflow status vocabulary
The load_seeds step listed `implemented` as an example status filter, but the
real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds);
`implemented` has no producer. Matches the earlier CLI-TOOLS.md correction.
Refs #441
* refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds
Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported
deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested
in-process (review minor #1). No behavior change.
Filter comparison now matches the raw lowercased status (both sides already
normalized) instead of sanitizeForDisplay(status); sanitization is for output,
not matching (review nit #3).
* test(#441): add fast-check property coverage and count=1 boundary for list-seeds
Adds tests/list-seeds.property.test.cjs with four fast-check properties over
deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug
invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing
(review minor #1).
Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2).
* chore(#441): sync runtime launcher snippet into list-seeds workflow
Propagate the current _runtime-launcher.snippet.sh (with non-Claude
runtime home probes) into the new list-seeds.md workflow via
scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation.
* test(#441): record list-seeds.md in workflow size baseline (#1074)
---------
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
* feat(#1318): require external reviewers to verify plan claims against source
/gsd-review built its external-reviewer prompt from plan text only and never
asked reviewers to open the repo and verify claims, so a grounded HIGH could be
outvoted by ungrounded LOWs. Add a concise, generic source-grounding block to
build_prompt's Review Instructions: treat yourself as running in the working
tree, open referenced files, cite path:line + mechanism, trace asserted
mechanisms, downgrade to an open question if you have no file access, and know
that grounded findings are weighted more heavily.
Also clarify that CodeRabbit (a diff-only reviewer that never receives the
prompt) must not be weighted as a grounded plan-level verdict in consensus
synthesis. Workflow stays under its size cap (baseline bumped deliberately).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1318): add changeset for reviewer source-grounding
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore(#1318): mark changeset docs-exempt (internal reviewer-prompt wording)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(#1318): document reviewer source-grounding in COMMANDS.md; drop docs-exempt
Review: a user-visible behavioral Changed warrants a docs touch, not a
docs-exempt. Add a sentence to the /gsd-review entry in docs/COMMANDS.md
(reviewers verify against source, cite file:line, grounded findings weighted
higher) and remove the changeset docs-exempt marker so lint:docs passes via
docs-updated. Also note the literal build_prompt test anchor.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(#1318): harden build_prompt fence extraction to be fence-run-aware
Addresses maintainer review on PR #1421 (required-before-merge).
The buildPromptReviewInstructions() test helper located the closing fence with
`src.indexOf('\n```')`, which terminates at the FIRST triple-backtick line — so
a build_prompt ```markdown block whose body embeds a fenced code example would
truncate mid-content (dropping the `## Review Instructions` section) and give a
spurious failure or false pass. Since this feature feeds source/plan content
(which routinely contains code fences) to reviewers, that is a live fragility.
Rewrite the extraction to be fence-run-aware, mirroring the CommonMark close
rule in src/markdown-sectionizer.cts stripFencedCode: parse the opener's
backtick run length, then close on the first line with >= that many backticks
and only trailing whitespace — so a shorter nested fence is treated as content.
Add a fail-first regression test (a 4-backtick outer fence wrapping a nested
```bash block) asserting the trailing `## Review Instructions` still extracts.
Test-only change; no production .cts touched. Verified: test file 7/7,
empirical fail-first proof the old indexOf logic truncated, full suite
4236/4236, eslint clean. Codex review: approve.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>