05b170e4489bf4b3cabc0f276d24f4cb78f4f7bc
4875 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
05b170e448 |
chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam Productionizes the ADR-1671 Option-E reference example as a real module: src/context-predicates.cts (parser + selector + index builder) compiled to gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's --check/--write drift-guard idiom and wired into lint:generated-sync. Parser behavior is deliberately prototype-equivalent in this commit so the next commit's regression matrix binds to the real defects rather than to a missing module. Two locked design deviations from the prototype: - duplicates carry a count, not line numbers - the committed index carries no line field at all, resolving ADR-1671 open question 4: an artifact without line numbers cannot drift on a line shift, so promoting --check to a CI gate does not make it routinely red Also reconciles the one remaining duplicate predicate ID (RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording is removed) so the gate can land fail-closed on duplicates. Refs #1671 * test(#2928): failing-first matrix for the predicate fact-store Adds the regression matrix from the phase test plan: parser declaration forms, fence and comment regions, ID/value grammar boundaries at limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard CLI, the selector query surface, and four document-shaped fast-check properties. Seven rows are RED for behavioral reasons against the ported parser: indented-bare, star-list, plus-list and numbered-list declaration forms are dropped; a tilde fence and a four-backtick fence containing a shorter fence are not skipped; and a multi-line HTML comment is parsed as live. Eleven selector rows are RED because the query surface is not wired yet. Negative fixtures come from real repo documents that predate the grammar (CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the fixture-provenance rule, and the property generators are document-shaped rather than seeded from our own serializer. Refs #1671 * fix(#2928): consume the shared fence scanner, relocate the index, wire the selector Drives the failing-first matrix green. Parser: replaces the ported naive triple-backtick toggle with the shared markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord gain an export keyword — the only change to that module, which has 71 upstream dependents — because it already returns line-indexed spans, which is exactly what a line-reporting parser needs. It also already documents itself as the second copy of the fence state machine pending consolidation; adding a third copy here would have been the generative-fix divergence this repo warns about. A parity suite now pins predicate fence-skipping against that scanner across eight fence shapes. HTML-comment skipping stays local because the sectionizer has no comment scanner. Declaration forms widen to indented-bare, star, plus and numbered list items. Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The remote matrix run caught the original choice — a committed .cjs there ships ~120KB of CONTEXT.md prose into a runtime module, and two content guards fired truthfully on it (a leaked .claude install path, and four hardcoded package-name literals). Neither guard was allowlisted; the artifact moved instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to require it — it is a drift-detection artifact, so the selector parses CONTEXT.md live and is always current. Generator: adds a frozen REASON enum and --check --json so the gate's outcome is asserted structurally instead of by matching prose, and --context-path/--index-path so tests drive the real CLI against a temp tree with no filesystem monkeypatching. Selector: gsd_run query context-predicates with --class/--prefix/--contains, structured output carrying a matched count, own-property guards, and no project-root resolution. Registering it exposed that the query dispatch table and the usage string had drifted: a new parity test found 20 routed commands missing from the usage list, all added here rather than deferred. Refs #1671 * test(#2928): lock the newly-public scanFencedBlocks contract Exporting scanFencedBlocks made it public API for the first time, so it needs its own contract test independent of the consumer that motivated the export. Memtrace's co-change analysis flagged the gap: this suite changes together with markdown-sectionizer.cts 8 times in 90 days and was absent from the diff. Covers the documented rules: 0-based indices, -1 for an unterminated fence, the same-char/>=length/no-trailing-text closer rule, a shorter fence inside a longer one staying content, CommonMark 4.5 backtick-in-info-string, and <=3-space indent tolerance. Refs #1671 * fix(#2928): address both isolated review passes Two independent reviewers (correctness axis and security axis, neither the author) found seven findings. All are fixed here with regression tests; none deferred. BLOCKER — comment-blind fence scanning caused silent, permanent predicate loss. The HTML-comment scan and the fence scan ran as two independent passes, and the fence scanner is comment-blind, so a fence delimiter inside an HTML comment with no later close read as an unterminated fence and skipped every remaining line to EOF. Worse, the drift-guard could not catch it: it diffs against a baseline produced by the same corrupted parse. The two constructs now interleave in a single pass so each suppresses the other's boundary detection while active, covered in both directions. The parity suite still binds this scanner to markdown-sectionizer's for comment-free documents, so the two cannot diverge unnoticed. BLOCKER — the selector was not consumed anywhere, leaving the phase's acceptance criterion unmet. Now wired into the pre-work predicate-citation step in contributor-standards, which is the repo's actual brief-assembly path; no code-level brief assembler exists to wire into. MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id regex nested a dot-containing character class inside a dot-prefixed repeat, so N consecutive dots had exponentially many partitions: 40 dots took 565ms and growth was exponential. CI runs this parser over a pull request's own CONTEXT.md, so any contributor could have hung a shared runner with one line. Replaced with linear per-segment validation. Doubled-dot ids are now rejected; the real document contains none. MAJOR — the duplicate-id gate had only ever been proven on synthetic fixtures. A test now re-inserts the exact line this branch removed and asserts the real generator names it. MAJOR — --check together with --write silently let write win, turning the gate into a writer; a missing path value resolved to the cwd and leaked an EISDIR stack trace. Both are now clean usage errors. MINOR — the hoisted skip-list was exported as a live mutable Set; replaced with a read-only predicate. MINOR — flag-shaped selector values were unmatchable; the inline --flag=value form now provides the escape hatch. Refs #1671 * chore(#2928): backfill changeset PR number 2938 --------- Co-authored-by: sim <sim@local> |
||
|
|
c043f2946c |
fix(#2914): per-PR ack fragments instead of one shared mutable file (#2923)
* fix(#2914): never persist a spent emitted-drift ack on next tests/emitted-drift-ack.json held 34 spent #2834 entries merged via #2900. Every entry is scoped to the diff that introduced it (#2789), so once merged to next it is at the base by definition -- spent and inert. Its presence is still load-bearing though: each PR rewrites the paths map wholesale, making a persistent base copy a shared cell. Five of six conflicting PRs in the open queue collided on this file and nothing else. Deletes the stale document and adds a push-to-next guard asserting it stays absent. The guard is deliberately NOT wired into lint:ci -- a PR-lane check against the base is the #2768 shape #2789 exists to end. Closes #2914 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2914): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2914): per-PR ack fragments instead of one shared mutable file The emitted-drift acknowledgment lived in a single tests/emitted-drift-ack.json whose paths map every PR rewrote wholesale. That is a shared mutable cell: any two PRs needing an ack edit the same lines and conflict. Five of six conflicting PRs in the open queue collided on this file and nothing else. Acks now live as per-PR fragments under tests/emitted-drift-acks/, the same shape .changeset/ already uses to solve this exact problem. Two PRs pick different filenames, so they cannot collide, and fragments lingering on next are harmless rather than toxic. The legacy file's 35 entries are MIGRATED into a fragment, not deleted. An earlier delete-only attempt failed verification twice: the ratchet lost the spec-phase.md acknowledgment from #2779 and reported a 10-byte growth with no ack. Relocating preserves every acknowledgment. The legacy single file is still READ (unioned with the fragments) because five open PRs carry it; dropping support would break all of them. A duplicate path key across sources is a hard error, never last-wins. The push-to-next guard is retargeted accordingly: it now asserts only that the legacy SHARED file never reappears on next. Fragments may persist harmlessly. Closes #2914 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d2d2f7c088 |
fix(#2848): non-Latin titles no longer produce empty slugs (Cyrillic transliteration) (#2934)
* test(#2848): add failing-first regression for non-Latin slug transliteration generateSlugInternal and slugify both strip non-ASCII chars with no transliteration step, so an all-Cyrillic title reduces to an empty slug. 12-row matrix: Cyrillic regression (both impls), Latin negative control, multi-letter mappings, soft/hard sign drops, Ukrainian extras, null contract, CJK unaffected, mixed scripts, truncation parity, slugify's distinct no-truncate contract. * fix(#2848): transliterate Cyrillic titles to ASCII before slug strip Both generateSlugInternal (src/core-utils.cts) and slugify (src/gsd2-import.cts) stripped non-ASCII with no transliteration, so an all-Cyrillic title reduced to an empty slug, producing unnamed phase directories (01-) and empty milestone_slug init JSON. Add a shared transliterateForSlug primitive (core-utils) covering Russian + the reported Ukrainian/Belarusian extras (і ї є ґ ў), with multi-letter mappings (ж→zh ч→ch ш→sh щ→sch ю→yu я→ya) and dropped soft/hard signs (ъ ь). It runs BEFORE the existing ASCII filter, so Latin-script text hits zero map entries and is byte-for-byte unchanged (negative control). slugify consumes the shared primitive, preserving its distinct single hyphen-strip + no-truncation contract. CJK/unmapped scripts keep the existing strip-to-ASCII behavior. Also corrects two test assertions to match the chosen й→y mapping and the б→b (not bie) transliteration. * changeset(#2848): Fixed — non-Latin slug transliteration * changeset(#2848): backfill PR number 2934 --------- Co-authored-by: sim <sim@local> |
||
|
|
81eeb8a53a |
docs(#2926): refresh ADR-1671 with findings re-verified on next (#2936)
Re-measured the Option-E prototype's reported figures against CONTEXT.md on next (2026-07-31) and recorded the delta rather than overwriting the June numbers: - index counts 393/18 (2026-06-24) -> 416/20 today; CONTEXT.md gained the PROBE (11) and PROHIB (10) classes - of the 3 duplicate predicate IDs, only RULESET.WORKFLOW_MARKDOWN.FENCES remains; the two RULESET.GEMINI.* went with the Gemini runtime removal - gen-context-index.cjs --check exits 1 on next, so Phase 0's "--check green in CI" criterion is unmet (invisible to CI: the example sits outside tests/) Adds Open question 4 (index keyed on baked line numbers re-drifts on any CONTEXT.md line shift, which matters once Phase 1 promotes --check to a CI gate), and records the reviewer-proposed eval-gate question as resolved by the PROBE.*/PROHIB.* predicate classes (ADR-550 D4/D7, ADR-1606). Also names both surfaces of the Windsurf 12 KB throw in Decision 2, since it is duplicated byte-identically in bin/install.js and src/runtime-artifact-conversion.cts. Docs-only. No code, no runtime-loaded text, no behavior change. Co-authored-by: sim <sim@local> |
||
|
|
76b7d73039 |
fix(#2733): route gate-passed spec-phase paths into the probe steps (#2779)
* fix(#2733): route gate-passed spec-phase paths into the probe steps All four gate-passed transitions in spec-phase.md said "Jump to Step 6", textually bypassing the mandatory Step 5.5 edge-completeness and Step 5.6 prohibition-completeness probes. Steps 5.5/5.6 were spliced between Step 5 and Step 6 by two later feature commits and the pre-existing jumps were never re-pointed, so no jump instruction in the file reached Step 5.5 at all and both probes were unreachable dead prose. Re-point the four gate-passed jumps (lines 129, 162, 168, 170) to Step 5.5. Control then flows 5.5 -> 5.6 -> 6 as the probes' own preconditions prescribe. The max-rounds "write anyway" bypasses and the probes' own "proceed to Step 6" exits are deliberately unchanged. Add tests/spec-phase-probe-reachability.test.cjs, which derives the mandatory probe steps from the file's own headings rather than hardcoding 5.5/5.6, so a future spliced-in probe step is covered without editing the test. It also locks the two coupled constraints: the max-rounds bypass must not be redirected into a probe, and each probe must keep its own onward exit. The existing probe contract tests are untouched and still pass; both scope from the "## Step 5.5"/"## Step 5.6" heading onward and were structurally incapable of observing the upstream jump text. * chore(changeset): Fixed fragment for #2779 (spec-phase probe reachability) * fix(#2733): route Step 5.5's own soft gate into Step 5.6 Round-1 review blocker. The four upstream gate-passed jumps were re-pointed to Step 5.5, but Step 5.5's own terminal soft gate at :305 still read "proceed to Step 6" - so the COMMON path (all applicable edges resolved) skipped the prohibition-completeness probe outright. Same defect class as the four this PR already fixed, on the success path of the very step being fixed: the SPEC shipped with an empty Prohibitions section instead of an empty Edge Coverage one. Its sibling at :393 is byte-identical yet correct, because Step 6 genuinely follows Step 5.6. Position, not phrasing, is the discriminator. The guard could not see it: the transition matcher keyed only on the literal "Jump to Step", and :305 says "proceed to Step". Widened it to a verb alternation (jump/proceed/continue/go/return/skip + "to Step N", case-insensitive) and renamed it TRANSITION_RE to match what it now models. This makes the file's own docstring promise - that a future spliced-in probe is covered without editing the test - true for a step whose exit is worded differently. Verified no false positives: the two pre-existing "continue to Step 3/4" transitions are upstream of both probes but target pre-probe steps, and the max-rounds bypass block contains no step transitions at all. Fail-first verified before fixing :305 - with the widened matcher against the unfixed workflow the guard fails naming exactly "spec-phase.md:305 jumps to Step 6, skipping mandatory Step 5.6", 4 pass / 1 fail; after the fix, 5/5. The two sibling probe contract tests stay 16/16. Also from review: - STEP_HEADING_RE gains an explicit \r? before $. Without it, on a CRLF checkout `.` stops before the \r and the unanchored $ fails to match, yielding ZERO steps and vacuously passing every assertion in the file. Not live today (.gitattributes forces eol=lf) but this repo has a recurring CRLF-regex bug class, so the guard no longer leans on it. - allow-test-rule category corrected to source-text-is-the-product; the previous runtime-contract-is-the-product is not one of the six recognized categories (CONTRIBUTING.md:609-619). - changeset body given the documented bold-lead-in form. - emitted-drift ack reason updated: +8 -> +10 bytes across five transitions (31987 -> 31997), DEFAULT tier, cap 40960. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
d49a7d0c4d |
fix(#2853): roadmap.update-plan-progress preserves hand-written annotations (#2916)
* test(#2853): add failing-first regression for plan-progress annotation preservation The count-bump regex's trailing [^\n]+ swallowed the whole Plans line and the replacement wrote back only the regenerated count, deleting any hand-written annotation after it. 8-row matrix covers bold/plain forms, bare template form, executed path, CRLF, and idempotency. * fix(#2853): preserve hand-written annotations in roadmap plan-progress bump The count-bump regex's trailing [^\n]+ swallowed the entire Plans line and the replacement wrote back only the regenerated count, deleting any hand-written prose after the count (e.g. a gap-closure annotation). The verb owns the count token only. Capture the existing count token ($2) and the trailing line text ($3), and rebuild the line as <label><new count><surviving text>. Trailing text is preserved ONLY when a real count token preceded it, so the fresh-template bracketed placeholder (`[Number of plans…]`) is still replaced cleanly rather than glued after the count (pre-#2853 behaviour on the template path preserved). CRLF \r is preserved via [^\r\n]. Widens replaceInCurrentMilestone to accept a replacement callback (needed to branch on whether the count group matched). The bare Plans: checklist header is still skipped — the lazy match lands on the summary line first and a count-less bare header yields no count to anchor preservation to. * changeset(#2853): backfill PR number 2916 --------- Co-authored-by: Test <test@example.com> Co-authored-by: sim <sim@local> |
||
|
|
8635cc447a |
chore(#2913): prune changeset fragments already promoted in the v1.9.1 CHANGELOG (#2922)
The v1.9.1 finalize consumed these 8 fragments on hotfix/1.9.1 and that
deletion reached main, but the back-merge did not propagate it to next
(
|
||
|
|
9bd0dbf0dd |
docs(#2534): rewrite your-first-project tutorial for beginners (#2569)
* docs(#2534): rewrite your-first-project tutorial for beginners Adds a loop mental-model primer (Mermaid), per-step "what just happened" callouts, a prerequisites flow, a glossary and a troubleshooting table. Same commands, same .planning artefacts, same to-do CLI example. Closes #2534 * docs(#2534): make the tutorial runtime-agnostic (all IDEs) Adds a "Pick your runtime" section (Cursor, Claude Code, OpenCode, Codex, Gemini CLI, Copilot, Windsurf, Kilo, Cline, Qwen, Antigravity, ...) with the installer flag and command syntax per runtime (/gsd-*, /gsd:* colon form, and Cline rules). Keeps the same guaranteed worked example and .planning artefacts. Closes #2534 * docs(#2534): address review - drop gsd-cursor aside + dead hero comment - Remove the '(pair with the gsd-cursor EoS ...)' parenthetical from the Cursor row. - Remove the commented-out reference to a non-existent hero asset. (Gemini CLI references retained: --gemini is still live in bin/install.js on next.) Closes #2534 * docs(#2534): fix review defects (keep multi-runtime) - Replace dead Gemini CLI / --gemini with its live successor Antigravity (#1928); remove the invalid --gemini row/flag everywhere. - Replace fabricated Step 1 output with realistic installer lines (71 skills/commands + destination suffix; exact lines vary by runtime). - Fix 'Skip research' -> choose 'No' on the real Research prompt. - behaviours -> behaviors (2x). Multi-runtime 'Pick your runtime' section retained per author intent; scope re-approval on #2534 still pending. * docs(#2534): scope tutorial back to single-runtime (Claude Code) Per trek-e's 2026-07-27 review, resolve the multi-runtime blockers by returning to the approved scope: - Remove the 'Pick your runtime' table + per-runtime notes; leave a one- line pointer to docs/how-to/install-on-your-runtime.md (which already documents all runtimes) rather than duplicate it (avoids the drift). This kills Blocker 1 (Antigravity is slash-hyphen, not colon) and Blocker 2 (Codex is $gsd-*) at the source. - Step 1 uses --claude concretely; config-dir prose is Claude-local. - Step 5: fix singular researcher (plan-phase spawns one gsd-phase- researcher), and make the research choice consistent with Step 3 (choose 'Skip research'); drop the RESEARCH.md artifact line. - Glossary/troubleshooting/prereqs/Step 2 de-multi-runtimed. Returns the PR to #2534's approved 'docs-only, same commands' scope. * docs(#2534): correct tutorial prerequisites and outputs * docs(#2534): match tutorial research prompts to workflow * docs(#2534): complete tutorial step guidance --------- Co-authored-by: clezcoding <clezcoding@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
2f66788cd3 |
docs(#2619): add ADR-2619 observability and shareable diagnostics (#2862)
Completes ADR-0174 §6's observability rollout and adds the outbound trust boundary that ADR-1577's inbound boundary has no counterpart for. D1 (wire the seam behind the existing opt-in gate) shipped via #2620 / PR #2621. D1b records that the unconditional stderr-on-error rule at 0174:105 is the target state, deferred behind an explicit --json-errors envelope version plus migration note -- disclosed as a partial supersede rather than retconned. D2-D5 are Directional; D6's non-goals are binding. Regenerates docs/adr/README.md via scripts/gen-adr-index.cjs --write. Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
49793465d7 |
docs(#2915): how-to for listing a reviewer lane, and correct the stale listing section (#2917)
* docs(#2904): how-to for listing a reviewer lane in the registry #2912 shipped the Reviewer Lane Registry, which lands in 1.9.1. Two docs consequences. New: docs/how-to/list-your-reviewer-lane.md. A Diataxis how-to for the publish task -- which of the three catalogs applies (and why a runtime carrying a reviewer body lists under its primary install shape instead), opening the required discussion thread BEFORE the PR, the three fields that reject entries most often (slug grammar differs from id, flags stay kebab when the slug is snake, install/uninstall must be copy-pasteable), regenerate-don't-hand-edit, and register-once-then-Releases. Links the registry README for the field table rather than duplicating it -- the spec is reference, this is the task flow. Corrected: ship-a-reviewer-lane.md said "Listing your lane is not wired yet" and pointed at #2904 as future work. #2906 merged at 11:25Z and #2912 at 12:11Z, so that section shipped false the moment the registry landed. Replaced with the publish pointer. Indexed the new guide and the generated catalog in docs/README.md, and added the guide to develop-a-capability.md's ecosystem list. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2904): caveat credential-bearing configKeys in both the guide and the spec Isolated security review found the worked entry's `configKeys: ["acme.api_key"]` modelled storing a live credential with no note on where that value ends up. Verified: config values are written in plaintext to .planning/config.json (docs/CONFIGURATION.md:227 -- masking is display-only, "that file is the security boundary"), and planning.commit_docs defaults to true (:466). So a credential declared that way lands in the installing user's git repository unless they have gitignored .planning/. None of the twelve first-party lanes does this -- they own only review.models.*, host, and prompt-budget keys. The pattern originates in docs/registries/README.md:221, shipped by #2912, so the caveat goes on BOTH surfaces rather than only on the copy that inherited it -- the spec's example is what future authors will read first. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2915): backfill changeset PR number pr: 0 -> 2917. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a9aba61f89 |
Merge pull request #2920 from open-gsd/chore/backmerge-main-to-next-4f1cce98
chore: back-merge main → next (
|
||
|
|
c3c6566ac2 |
chore: back-merge main into next (4f1cce98)
|
||
|
|
932f99907c |
Merge pull request #2919 from open-gsd/chore/sync-next-version-1.9.1
chore: sync next package version to 1.9.1 |
||
|
|
854c93533c | chore: sync next package version to 1.9.1 | ||
|
|
4f1cce9875 |
Merge pull request #2918 from open-gsd/hotfix/1.9.1
chore: merge release v1.9.1 to main |
||
|
|
957ebd8e6c | chore: promote CHANGELOG for v1.9.1 | ||
|
|
538cb0fc1d |
enh(#2904): add a reviewer entry type so third-party reviewer lanes are discoverable (#2912)
* feat(#2904): add a `reviewer` entry type so third-party reviewer lanes are discoverable
ADR-2782 made a reviewer lane installable by a third party, but neither
discoverability catalog could hold one. The Community Capability Registry
requires a non-empty `loopExtensionPoints` and forbids a lane from declaring
any hook kind, so a `role: "reviewer"` entry is unsatisfiable by construction;
the EoS Registry is for ADR-1239 host integrations, which a lane is not.
Adds a third catalog — `docs/registries/reviewers.json` →
`docs/registries/reviewer-registry.md` — whose `interactions` describes the
lane: slug, flags, transport, evidenceClass, reviewsSection, requiresBinaries,
configKeys, runtimeCompat.
The lane vocabulary is a hand-written mirror of `capability-validator.cjs`
(the same pattern as `AXES` mirroring `HOST_INTEGRATION_AXES`), with parity
enforced by tests/registry-reviewer-parity.test.cjs. `slug` deliberately uses
the runtime `LANE_SLUG_RE` grammar rather than the registry's kebab-only `id`
rule, so real lanes (`lm_studio`, `4o-mini`) are not rejected.
Two binary type branches became three-way Map dispatch. Both now fail loudly
on an unrecognized type instead of silently treating it as a capability —
`renderMarkdown` in particular writes a committed catalog file, so a silent
wrong-title render was the worst failure mode available.
Also fixed while here: `gen-registry.cjs` parsed source JSON with no error
handling, so a malformed or non-array `capabilities.json` surfaced as a raw
SyntaxError/TypeError instead of an actionable CLI error.
Closes #2904
* fix(#2904): bound and sanitize untrusted registry `interactions` strings
Review findings from the pre-PR passes.
Security (isolated pass): `interactions` string fields reached the generated,
committed Markdown catalog with no control-character check and no length
bound. A `reviewsSection` carrying ESC and a `requiresBinaries` element
carrying NUL plus 5000 characters validated clean and landed verbatim in the
rendered page — `mdInline` escapes Markdown metacharacters and collapses CRLF,
but nothing else. The identical gap already existed on the capability type's
`configKeys`/`requires`/`runtimeCompat`/`produces`/`consumes`, so it is fixed
there too rather than inherited into a third type.
`hasDisallowedControlChar` is lifted to module scope so exactly one
implementation exists, and a shared `validateStringArrayField` enforces
control-character rejection, a 200-character element cap and a 50-element
array cap for both types.
Correctness (standards pass): `renderMarkdown`'s per-entry summary builder was
still an if/else-if chain whose final `else` was the capability branch — the
one per-type dispatch point this change had not converted, and the same silent
fallthrough it removes elsewhere. It now lives in `RENDER_META` alongside the
title, so a fourth type cannot silently inherit capability's rendering. All
three types' rendered output is byte-identical to before the refactor.
Also corrects a test comment that still claimed the reviewer suites were
failing-first against an unmodified module.
* chore(#2904): backfill changeset PR number (#2912)
(cherry picked from commit
|
||
|
|
f72f70ad39 |
docs(#2782): how-to for declaring a reviewer lane in a capability (#2906)
* docs(#2782): how-to for declaring a reviewer lane in a capability
ADR-2782 shipped the Reviewer Lane capability surface in 1.9.0, but the
only documentation was reference (capability-manifest.md) and rationale
(ADR-2782). A capability author had no task-oriented path from "I have a
review CLI" to "/gsd:review invokes it".
Adds docs/how-to/ship-a-reviewer-lane.md: role selection, the spawn and
openai-http worked examples, federated config ownership, the
review-lane query surface as the verification step, what the install
disclosure and egress-host re-verification mean for the author, and the
data-only boundary with the two named CLIs that do not fit today.
Also corrects the manifest reference's `invoke` row, which understated
three enums against the shipped validator: promptChannel omitted `argv`,
effortChannel omitted `env`, and the openai-http sub-shape plus the
required-with-file-arg `outputArg` were undocumented entirely.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(#2782): correct four claims found by the two orthogonal reviews
Security review (1 major, 2 minor):
- "a reserved slug" implied the gsd-/anthropic- namespace rule, which
guards the capability id, not reviewer.slug. The slug guard is
isReservedName (__proto__/constructor/prototype), a prototype-pollution
barrier. Both rules are now stated and kept apart.
- Added ADR-2782 D5's own caveat verbatim: disclosure and host pinning
make the channel visible, pinned and revocable, not safe, and
consent-at-install is a weaker gate for a standing egress channel than
for a hook.
- Named integrity/SHA pinning and engines.gsd as the controls that make
the disclosure tamper-evident and the version range enforceable.
Correctness review (1 major, 1 minor):
- Claimed a name collision is "a hard failure at install". It is not.
installCapability never runs validateCrossCapability; the check runs in
loadRegistry, and a colliding overlay is dropped from acceptedMap with a
warning while the install reports success. Documented as the quiet
failure mode it is, with the symptom to look for.
- The feature-only field list omitted hooks and activationKey, both of
which FEATURE_FIELDS_FORBIDDEN_ON_REVIEWER rejects.
Both worked examples re-validated against validateCapability() -> [].
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(#2782): restore Kimi Code to the cross-AI reviewer list
set-up-cross-ai-review.md named eleven reviewers; twelve lanes ship. The
kimi-code lane (added by #2718, declared as manifest data by #2798) was
never added here — the same roster-drift class as #2781, which #2800's
parity gate covers for COMMANDS.md and FEATURES.md but not for how-to
prose.
Also points readers at the declared-lane model rather than a static list:
the roster is now generated, a capability can ship its own lane, and
`gsd-tools review-lane sections` answers "what do I actually have".
Verified against the twelve declared bodies in capabilities/*/capability.json.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(#2782): use the invocation form the runtime descriptors actually declare
The new guide used /gsd:review. Nothing in this repo produces that form.
- capability-validator VALID_COMMAND_STYLES is {slash-hyphen, shell-var};
there is no colon/namespaced style in the vocabulary at all.
- 18 of 19 runtime capabilities declare commandStyle "slash-hyphen",
claude included; codex is "shell-var". Every artifactLayout prefix is
"gsd-".
- A plain-file command install never namespaces, so .claude/commands/
gsd-review.md is typed /gsd-review.
- The Claude Code plugin surface would namespace on plugin.json "name",
which is "gsd-core" -- so the plugin form would be /gsd-core:review.
The commands/gsd/ subdirectory is cosmetic and contributes nothing to
the invoked name.
So /gsd:review is neither the installed form nor the plugin form. Uses
/gsd-review, matching set-up-cross-ai-review.md and the 18 descriptors.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(#2782): state that lane listing is not wired yet
The guide could describe authoring, validating, and installing a lane but
implied a publishing path that does not exist. Neither discoverability
catalog can accept one: a Community Capability Registry entry requires a
non-empty loopExtensionPoints plus hookKinds, and a role:"reviewer"
capability is forbidden from declaring steps/contributions/gates, so both
fields are unsatisfiable rather than merely unset. The EoS Registry is
ADR-1239 host integrations, which a lane is not.
The registry schema predates the reviewer role by 17 days (#2182 Jul 11,
ADR-2782 Jul 28) and registry-schema.cjs has zero occurrences of
"reviewer". Tracked for a 1.9.x point release by #2904.
Says so explicitly, and tells authors NOT to file a loop extension point
they do not use to get past validation -- a schema satisfiable only by
lying is one that will be lied to.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(#2782): backfill PR number and fix the changeset invocation form
pr: 0 -> 2906.
Also corrects /gsd:review -> /gsd-review in the fragment body. The
fragment renders into CHANGELOG.md, which is a reader-facing docs surface
and is never passed through the install-time converter -- so the colon
form would ship the #2903 drift into a permanent release artifact.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(#2907): propagate the invocation-form fix and restore lane order
Two findings from an isolated review of the post-review delta.
- docs/README.md and develop-a-capability.md still said /gsd:review in
the cross-links added for the new guide. The form was corrected in the
guide itself but not in the two entries pointing at it, leaving three
docs making the same claim in two different forms.
- set-up-cross-ai-review.md inserted Kimi Code between Antigravity and
Ollama. REVIEWER_LANES is ordered by write_reviews order and kimi-code
is 12th, appended after llama_cpp; the prose list mirrored declaration
order before this change and now does again.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Test <test@example.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
(cherry picked from commit
|
||
|
|
7112c6ca47 |
fix(#2844): verify-summary ignores future/prose path mentions; resolves project root (#2910)
* fix(#2844): verify-summary binds file-claim extraction to a creation-claim context
verify-summary's Pattern 1 matched any backticked path-like token with no
context check, so a prose mention of a future deliverable (`shared/types.ts`
in a 'next phase will add…' sentence) was checked for existence and its absence
failed the verdict on a healthy phase. #2685 added shape filtering but no
context check.
- src/verify.cts: both extraction patterns now require a claim label on the line
(Created/Modified/Added/Updated/Edited/key-files). A bare prose mention no
longer matches; genuine labeled claims still do.
- gsd-core/bin/gsd-tools.cjs: remove 'verify-summary' from SKIP_ROOT_RESOLUTION
so relative claim paths resolve against the project root, not the raw cwd
(subdirectory invocation no longer manufactures missing files).
Regression tests: prose mention not treated as a claim; prose-only SUMMARY
passes; absent claimed file still fails.
* chore(#2844): backfill changeset PR 2910
---------
Co-authored-by: Test <test@example.com>
(cherry picked from commit
|
||
|
|
e1b275766d |
fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH (#2909)
* fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH
findProjectRoot's heuristic (3) used isInsideGitRepo(parent), which only checked
'does SOME .git exist between start and the ancestor' — it never verified the
.git was co-located with / bounded the trusted .planning/. A nested child repo
(own .git, no .planning) under an ancestor GSD project satisfied the check, so
resolution silently crossed into the ancestor project (wrong identity, exit 0).
Add nearestGitRoot(from, upTo) (fs-walk, no spawn) and use it in heuristics (3)
and (4): if the caller is inside its own nested repo whose root is strictly below
the candidate ancestor, do not return that ancestor. The plain-descendant (#1414),
co-located .git+.planning, and sub_repos/multiRepo cases are unchanged.
Regression test: a nested child .git under an ancestor .planning no longer
resolves to the ancestor; the co-located single-repo case still resolves.
* chore(#2843): backfill changeset PR 2909
---------
Co-authored-by: Test <test@example.com>
(cherry picked from commit
|
||
|
|
3aabb0f441 |
fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point (#2905)
* fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point
gsd-code-fixer was the only writer that hand-rolled a git worktree inside the
agent prompt and the only one that never read workflow.use_worktrees. With the
setting explicitly false, --fix still created worktrees; the fresh worktree had
no node_modules, so runs improvised a teardown whose `rm -rf` followed a Windows
junction into the REAL node_modules (silent data loss, 3x observed).
Defect 1: gate setup_worktree + its cleanup tail on workflow.use_worktrees
(same gsd_run query config-get read the four sibling workflows use). When false:
edit/commit in the main checkout (wt='.', no temp branch, no sentinel, no cleanup).
Defect 2: forbid rm -rf on a possible reparse point in the spec — never fall
through to a destructive remove; on failure, stop and surface the error.
Defect 3: REVIEW-FIX records where verification ran (main checkout vs worktree).
The transactional worktree path (#2839/#2990/#2686) is unchanged when worktrees
are enabled. Docs-parity guards in tests/code-review.test.cjs bind the spec to
the fix.
* chore(#2825): backfill changeset PR 2905
* fix(#2825): drop stale agents/gsd-code-fixer.toml ack + spent entries (emitted-attribution)
gsd-test failed: agents/gsd-code-fixer.toml is a STALE ack here — that TOML was
changed by #2834 (now in next), not this PR. The 4 base-carried acks
(autonomous/discuss-phase-assumptions/next/plan-phase) are spent/inert.
---------
Co-authored-by: Test <test@example.com>
(cherry picked from commit
|
||
|
|
dc73680532 |
fix(#2834): write defaults.json before agent TOML generation on clean Codex install (#2900)
* fix(#2834): write defaults.json before agent TOML generation on clean Codex install
Extracted writeNonClaudeDefaults(runtime) and called it BEFORE installCodexConfig
so the runtime-aware model resolver has resolve_model_ids=omit + runtime=codex in
~/.gsd/defaults.json before agent TOMLs are generated. Pre-fix, a clean first Codex
install generated TOMLs with no model fields (the resolver didn't know the runtime);
a second run fixed it. The original inline defaults-write block (which ran AFTER agent
generation) is replaced by the earlier function call (idempotent).
* chore(#2834): changeset fragment
* fix+test(#2834): acknowledge codex TOML drift (emitted-attribution) + fix test comment window
The emitted-attribution gate flags 19 codex agent TOMLs that now carry model-routing
fields (the fix's correct effect) but can't link them to a .md or src/ change (the fix
is in bin/install.js ordering). Acknowledge the drift in emitted-drift-ack.json. Fix the
test's comment-detection window (300 chars to capture the #2834 rationale).
* fix(#2834): ack remaining 15 codex TOML drift paths
* chore(#2834): backfill changeset PR number (2900)
* chore(#2834): ack code-review.md growth from concurrent merge (rebase pickup)
* fix(#2834): remove stale code-review.md ack (emitted-attribution failure)
CI failed: 'differential attribution over the real tree' — the code-review.md
ack added in 6e0b4b3b8 ('ack code-review.md growth from concurrent merge') is
STALE: this PR's diff does not touch code-review.md (only bin/install.js + tests
+ changeset), so the ack explains growth that isn't here. The base already
absorbed the concurrent code-review.md growth; the ack is inert here and the
gate flags it as stale. Remove it.
---------
Co-authored-by: Test <test@example.com>
(cherry picked from commit
|
||
|
|
a38e4d080d |
fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row (#2902)
* fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row
Two coupled defects in the requirement traceability state machine:
Defect 1 (terminal state): requirements revert-phase (the gaps_found response)
left a row at 'Gaps Found' with no inverse — neither mark-complete's /^pending$/i
guard nor the phase-complete reconcile's /^(?:pending|in progress)$/i accepted it,
so a single failed verification stranded every requirement permanently and blocked
the milestone. Widen both guards to accept 'gaps found' so a genuinely-satisfied
stranded row reaches Complete again.
Defect 2 (false success): mark-complete ORed checkboxHit || tableHit for 'updated',
so on a Gaps Found row it flipped the checkbox but could not move the row, yet
reported updated:true. When a traceability table has a row for an ID, gate 'updated'
on the row moving (tableHit) — a checkbox-only flip on a table-bearing file no
longer lies. The #2140 table_unmatched path (no row for the ID) is preserved.
* chore(#2788): backfill changeset PR 2902
---------
Co-authored-by: Test <test@example.com>
(cherry picked from commit
|
||
|
|
14cfbdaad0 |
fix(#2667): run-with-timeout mediates .cmd/.bat spawns on Windows (CVE-2024-27980); fallow pre-pass names failure kind (#2897)
* fix(#2667): mediate .cmd/.bat/.exe spawns on Windows; split fallow pre-pass failure diagnostic
run-with-timeout spawned .cmd/.bat/.exe commands without shell:true on Windows,
tripping Node's CVE-2024-27980 EINVAL (April 2024 security hardening). The fallow
structural pre-pass then no-op'd silently — a hard execution failure read the same
as 'optional dependency absent'.
(A) gsd-core/bin/gsd-tools.cjs runWithTimeout: gate shell:true on
(win32 && command ends in .cmd/.bat/.exe). Narrow by design — never fires for
the 7 `bash -c` callers (command is `bash`, no such suffix), so the recorded
no-shell-for-argv-array security contract (DEFECT.UNBOUNDED-SUBPROCESS) is
preserved; cmdArgs stays an array. POSIX untouched.
(B) code-review.md fallow pre-pass: name the failure KIND (timeout / spawn failure
/ crash / not-found) so a Windows .cmd spawn failure is not mistaken for an
absent binary.
Regression test in tests/run-with-timeout.test.cjs gated to win32 (.cmd/.bat/.exe
shims run with exit 0 + non-empty stdout; pre-fix EINVAL → exit 125/empty). POSIX
negative-space test guards the unchanged bash -c callers.
* chore(#2667): changeset fragment
* chore(#2667): backfill changeset PR 2897 + correct body (cmd.exe array, not shell:true)
* fix(#2667): exclude .exe from the win32 spawn-mediation gate; ack code-review.md growth
CI caught two failures on the first push:
1. windows-24: 'exits 124 when the wall-clock budget is exceeded' regressed. The
gate matched .exe, so the HANG command (node.exe -e 'setTimeout(...)') was
wrapped in 'cmd.exe /c node.exe ...' — the wrapped child escaped the timeout
cap's process-group reap (exit 124 never fired; hit the 30s harness backstop)
AND cmd.exe risked mis-parsing the -e script arg. .exe is INTENTIONALLY
excluded now: real PE executables spawn fine directly; only .cmd/.bat are the
CVE-2024-27980 EINVAL cases. The .exe test becomes a negative-space test
(node.exe spawned directly, exit 0).
2. ubuntu-22: emitted-attribution — code-review.md grew 1177 bytes from the
#2667 fallow pre-pass failure-KIND case statement; acknowledge it.
---------
Co-authored-by: Test <test@example.com>
(cherry picked from commit
|
||
|
|
13fcbe45a6 | chore: bump version to 1.9.1 for hotfix | ||
|
|
90771ddf02 |
enh(#2904): add a reviewer entry type so third-party reviewer lanes are discoverable (#2912)
* feat(#2904): add a `reviewer` entry type so third-party reviewer lanes are discoverable ADR-2782 made a reviewer lane installable by a third party, but neither discoverability catalog could hold one. The Community Capability Registry requires a non-empty `loopExtensionPoints` and forbids a lane from declaring any hook kind, so a `role: "reviewer"` entry is unsatisfiable by construction; the EoS Registry is for ADR-1239 host integrations, which a lane is not. Adds a third catalog — `docs/registries/reviewers.json` → `docs/registries/reviewer-registry.md` — whose `interactions` describes the lane: slug, flags, transport, evidenceClass, reviewsSection, requiresBinaries, configKeys, runtimeCompat. The lane vocabulary is a hand-written mirror of `capability-validator.cjs` (the same pattern as `AXES` mirroring `HOST_INTEGRATION_AXES`), with parity enforced by tests/registry-reviewer-parity.test.cjs. `slug` deliberately uses the runtime `LANE_SLUG_RE` grammar rather than the registry's kebab-only `id` rule, so real lanes (`lm_studio`, `4o-mini`) are not rejected. Two binary type branches became three-way Map dispatch. Both now fail loudly on an unrecognized type instead of silently treating it as a capability — `renderMarkdown` in particular writes a committed catalog file, so a silent wrong-title render was the worst failure mode available. Also fixed while here: `gen-registry.cjs` parsed source JSON with no error handling, so a malformed or non-array `capabilities.json` surfaced as a raw SyntaxError/TypeError instead of an actionable CLI error. Closes #2904 * fix(#2904): bound and sanitize untrusted registry `interactions` strings Review findings from the pre-PR passes. Security (isolated pass): `interactions` string fields reached the generated, committed Markdown catalog with no control-character check and no length bound. A `reviewsSection` carrying ESC and a `requiresBinaries` element carrying NUL plus 5000 characters validated clean and landed verbatim in the rendered page — `mdInline` escapes Markdown metacharacters and collapses CRLF, but nothing else. The identical gap already existed on the capability type's `configKeys`/`requires`/`runtimeCompat`/`produces`/`consumes`, so it is fixed there too rather than inherited into a third type. `hasDisallowedControlChar` is lifted to module scope so exactly one implementation exists, and a shared `validateStringArrayField` enforces control-character rejection, a 200-character element cap and a 50-element array cap for both types. Correctness (standards pass): `renderMarkdown`'s per-entry summary builder was still an if/else-if chain whose final `else` was the capability branch — the one per-type dispatch point this change had not converted, and the same silent fallthrough it removes elsewhere. It now lives in `RENDER_META` alongside the title, so a fourth type cannot silently inherit capability's rendering. All three types' rendered output is byte-identical to before the refactor. Also corrects a test comment that still claimed the reviewer suites were failing-first against an unmodified module. * chore(#2904): backfill changeset PR number (#2912) |
||
|
|
557d46984e |
docs(#2782): how-to for declaring a reviewer lane in a capability (#2906)
* docs(#2782): how-to for declaring a reviewer lane in a capability ADR-2782 shipped the Reviewer Lane capability surface in 1.9.0, but the only documentation was reference (capability-manifest.md) and rationale (ADR-2782). A capability author had no task-oriented path from "I have a review CLI" to "/gsd:review invokes it". Adds docs/how-to/ship-a-reviewer-lane.md: role selection, the spawn and openai-http worked examples, federated config ownership, the review-lane query surface as the verification step, what the install disclosure and egress-host re-verification mean for the author, and the data-only boundary with the two named CLIs that do not fit today. Also corrects the manifest reference's `invoke` row, which understated three enums against the shipped validator: promptChannel omitted `argv`, effortChannel omitted `env`, and the openai-http sub-shape plus the required-with-file-arg `outputArg` were undocumented entirely. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2782): correct four claims found by the two orthogonal reviews Security review (1 major, 2 minor): - "a reserved slug" implied the gsd-/anthropic- namespace rule, which guards the capability id, not reviewer.slug. The slug guard is isReservedName (__proto__/constructor/prototype), a prototype-pollution barrier. Both rules are now stated and kept apart. - Added ADR-2782 D5's own caveat verbatim: disclosure and host pinning make the channel visible, pinned and revocable, not safe, and consent-at-install is a weaker gate for a standing egress channel than for a hook. - Named integrity/SHA pinning and engines.gsd as the controls that make the disclosure tamper-evident and the version range enforceable. Correctness review (1 major, 1 minor): - Claimed a name collision is "a hard failure at install". It is not. installCapability never runs validateCrossCapability; the check runs in loadRegistry, and a colliding overlay is dropped from acceptedMap with a warning while the install reports success. Documented as the quiet failure mode it is, with the symptom to look for. - The feature-only field list omitted hooks and activationKey, both of which FEATURE_FIELDS_FORBIDDEN_ON_REVIEWER rejects. Both worked examples re-validated against validateCapability() -> []. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2782): restore Kimi Code to the cross-AI reviewer list set-up-cross-ai-review.md named eleven reviewers; twelve lanes ship. The kimi-code lane (added by #2718, declared as manifest data by #2798) was never added here — the same roster-drift class as #2781, which #2800's parity gate covers for COMMANDS.md and FEATURES.md but not for how-to prose. Also points readers at the declared-lane model rather than a static list: the roster is now generated, a capability can ship its own lane, and `gsd-tools review-lane sections` answers "what do I actually have". Verified against the twelve declared bodies in capabilities/*/capability.json. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2782): use the invocation form the runtime descriptors actually declare The new guide used /gsd:review. Nothing in this repo produces that form. - capability-validator VALID_COMMAND_STYLES is {slash-hyphen, shell-var}; there is no colon/namespaced style in the vocabulary at all. - 18 of 19 runtime capabilities declare commandStyle "slash-hyphen", claude included; codex is "shell-var". Every artifactLayout prefix is "gsd-". - A plain-file command install never namespaces, so .claude/commands/ gsd-review.md is typed /gsd-review. - The Claude Code plugin surface would namespace on plugin.json "name", which is "gsd-core" -- so the plugin form would be /gsd-core:review. The commands/gsd/ subdirectory is cosmetic and contributes nothing to the invoked name. So /gsd:review is neither the installed form nor the plugin form. Uses /gsd-review, matching set-up-cross-ai-review.md and the 18 descriptors. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2782): state that lane listing is not wired yet The guide could describe authoring, validating, and installing a lane but implied a publishing path that does not exist. Neither discoverability catalog can accept one: a Community Capability Registry entry requires a non-empty loopExtensionPoints plus hookKinds, and a role:"reviewer" capability is forbidden from declaring steps/contributions/gates, so both fields are unsatisfiable rather than merely unset. The EoS Registry is ADR-1239 host integrations, which a lane is not. The registry schema predates the reviewer role by 17 days (#2182 Jul 11, ADR-2782 Jul 28) and registry-schema.cjs has zero occurrences of "reviewer". Tracked for a 1.9.x point release by #2904. Says so explicitly, and tells authors NOT to file a loop extension point they do not use to get past validation -- a schema satisfiable only by lying is one that will be lied to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2782): backfill PR number and fix the changeset invocation form pr: 0 -> 2906. Also corrects /gsd:review -> /gsd-review in the fragment body. The fragment renders into CHANGELOG.md, which is a reader-facing docs surface and is never passed through the install-time converter -- so the colon form would ship the #2903 drift into a permanent release artifact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2907): propagate the invocation-form fix and restore lane order Two findings from an isolated review of the post-review delta. - docs/README.md and develop-a-capability.md still said /gsd:review in the cross-links added for the new guide. The form was corrected in the guide itself but not in the two entries pointing at it, leaving three docs making the same claim in two different forms. - set-up-cross-ai-review.md inserted Kimi Code between Antigravity and Ollama. REVIEWER_LANES is ordered by write_reviews order and kimi-code is 12th, appended after llama_cpp; the prose list mirrored declaration order before this change and now does again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Test <test@example.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
42f4f184c0 |
fix(#2844): verify-summary ignores future/prose path mentions; resolves project root (#2910)
* fix(#2844): verify-summary binds file-claim extraction to a creation-claim context verify-summary's Pattern 1 matched any backticked path-like token with no context check, so a prose mention of a future deliverable (`shared/types.ts` in a 'next phase will add…' sentence) was checked for existence and its absence failed the verdict on a healthy phase. #2685 added shape filtering but no context check. - src/verify.cts: both extraction patterns now require a claim label on the line (Created/Modified/Added/Updated/Edited/key-files). A bare prose mention no longer matches; genuine labeled claims still do. - gsd-core/bin/gsd-tools.cjs: remove 'verify-summary' from SKIP_ROOT_RESOLUTION so relative claim paths resolve against the project root, not the raw cwd (subdirectory invocation no longer manufactures missing files). Regression tests: prose mention not treated as a claim; prose-only SUMMARY passes; absent claimed file still fails. * chore(#2844): backfill changeset PR 2910 --------- Co-authored-by: Test <test@example.com> |
||
|
|
39dbe5e0f5 |
fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH (#2909)
* fix(#2843): findProjectRoot stops at git-repo boundaries, not just MAX_DEPTH findProjectRoot's heuristic (3) used isInsideGitRepo(parent), which only checked 'does SOME .git exist between start and the ancestor' — it never verified the .git was co-located with / bounded the trusted .planning/. A nested child repo (own .git, no .planning) under an ancestor GSD project satisfied the check, so resolution silently crossed into the ancestor project (wrong identity, exit 0). Add nearestGitRoot(from, upTo) (fs-walk, no spawn) and use it in heuristics (3) and (4): if the caller is inside its own nested repo whose root is strictly below the candidate ancestor, do not return that ancestor. The plain-descendant (#1414), co-located .git+.planning, and sub_repos/multiRepo cases are unchanged. Regression test: a nested child .git under an ancestor .planning no longer resolves to the ancestor; the co-located single-repo case still resolves. * chore(#2843): backfill changeset PR 2909 --------- Co-authored-by: Test <test@example.com> |
||
|
|
5d0fd4dc53 |
fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point (#2905)
* fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point gsd-code-fixer was the only writer that hand-rolled a git worktree inside the agent prompt and the only one that never read workflow.use_worktrees. With the setting explicitly false, --fix still created worktrees; the fresh worktree had no node_modules, so runs improvised a teardown whose `rm -rf` followed a Windows junction into the REAL node_modules (silent data loss, 3x observed). Defect 1: gate setup_worktree + its cleanup tail on workflow.use_worktrees (same gsd_run query config-get read the four sibling workflows use). When false: edit/commit in the main checkout (wt='.', no temp branch, no sentinel, no cleanup). Defect 2: forbid rm -rf on a possible reparse point in the spec — never fall through to a destructive remove; on failure, stop and surface the error. Defect 3: REVIEW-FIX records where verification ran (main checkout vs worktree). The transactional worktree path (#2839/#2990/#2686) is unchanged when worktrees are enabled. Docs-parity guards in tests/code-review.test.cjs bind the spec to the fix. * chore(#2825): backfill changeset PR 2905 * fix(#2825): drop stale agents/gsd-code-fixer.toml ack + spent entries (emitted-attribution) gsd-test failed: agents/gsd-code-fixer.toml is a STALE ack here — that TOML was changed by #2834 (now in next), not this PR. The 4 base-carried acks (autonomous/discuss-phase-assumptions/next/plan-phase) are spent/inert. --------- Co-authored-by: Test <test@example.com> |
||
|
|
00c859fae5 |
fix(#2834): write defaults.json before agent TOML generation on clean Codex install (#2900)
* fix(#2834): write defaults.json before agent TOML generation on clean Codex install Extracted writeNonClaudeDefaults(runtime) and called it BEFORE installCodexConfig so the runtime-aware model resolver has resolve_model_ids=omit + runtime=codex in ~/.gsd/defaults.json before agent TOMLs are generated. Pre-fix, a clean first Codex install generated TOMLs with no model fields (the resolver didn't know the runtime); a second run fixed it. The original inline defaults-write block (which ran AFTER agent generation) is replaced by the earlier function call (idempotent). * chore(#2834): changeset fragment * fix+test(#2834): acknowledge codex TOML drift (emitted-attribution) + fix test comment window The emitted-attribution gate flags 19 codex agent TOMLs that now carry model-routing fields (the fix's correct effect) but can't link them to a .md or src/ change (the fix is in bin/install.js ordering). Acknowledge the drift in emitted-drift-ack.json. Fix the test's comment-detection window (300 chars to capture the #2834 rationale). * fix(#2834): ack remaining 15 codex TOML drift paths * chore(#2834): backfill changeset PR number (2900) * chore(#2834): ack code-review.md growth from concurrent merge (rebase pickup) * fix(#2834): remove stale code-review.md ack (emitted-attribution failure) CI failed: 'differential attribution over the real tree' — the code-review.md ack added in 6e0b4b3b8 ('ack code-review.md growth from concurrent merge') is STALE: this PR's diff does not touch code-review.md (only bin/install.js + tests + changeset), so the ack explains growth that isn't here. The base already absorbed the concurrent code-review.md growth; the ack is inert here and the gate flags it as stale. Remove it. --------- Co-authored-by: Test <test@example.com> |
||
|
|
9f567a1627 |
fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row (#2902)
* fix(#2788): recover Gaps Found rows; mark-complete no longer false-succeeds on a rejected row Two coupled defects in the requirement traceability state machine: Defect 1 (terminal state): requirements revert-phase (the gaps_found response) left a row at 'Gaps Found' with no inverse — neither mark-complete's /^pending$/i guard nor the phase-complete reconcile's /^(?:pending|in progress)$/i accepted it, so a single failed verification stranded every requirement permanently and blocked the milestone. Widen both guards to accept 'gaps found' so a genuinely-satisfied stranded row reaches Complete again. Defect 2 (false success): mark-complete ORed checkboxHit || tableHit for 'updated', so on a Gaps Found row it flipped the checkbox but could not move the row, yet reported updated:true. When a traceability table has a row for an ID, gate 'updated' on the row moving (tableHit) — a checkbox-only flip on a table-bearing file no longer lies. The #2140 table_unmatched path (no row for the ID) is preserved. * chore(#2788): backfill changeset PR 2902 --------- Co-authored-by: Test <test@example.com> |
||
|
|
9f73703a4c |
Merge pull request #2901 from open-gsd/chore/backmerge-main-to-next-efbbcc35
chore: back-merge main → next (
|
||
|
|
ee519e883c |
chore: back-merge main into next (efbbcc35)
|
||
|
|
79ed181ec0 |
fix(#2667): run-with-timeout mediates .cmd/.bat spawns on Windows (CVE-2024-27980); fallow pre-pass names failure kind (#2897)
* fix(#2667): mediate .cmd/.bat/.exe spawns on Windows; split fallow pre-pass failure diagnostic run-with-timeout spawned .cmd/.bat/.exe commands without shell:true on Windows, tripping Node's CVE-2024-27980 EINVAL (April 2024 security hardening). The fallow structural pre-pass then no-op'd silently — a hard execution failure read the same as 'optional dependency absent'. (A) gsd-core/bin/gsd-tools.cjs runWithTimeout: gate shell:true on (win32 && command ends in .cmd/.bat/.exe). Narrow by design — never fires for the 7 `bash -c` callers (command is `bash`, no such suffix), so the recorded no-shell-for-argv-array security contract (DEFECT.UNBOUNDED-SUBPROCESS) is preserved; cmdArgs stays an array. POSIX untouched. (B) code-review.md fallow pre-pass: name the failure KIND (timeout / spawn failure / crash / not-found) so a Windows .cmd spawn failure is not mistaken for an absent binary. Regression test in tests/run-with-timeout.test.cjs gated to win32 (.cmd/.bat/.exe shims run with exit 0 + non-empty stdout; pre-fix EINVAL → exit 125/empty). POSIX negative-space test guards the unchanged bash -c callers. * chore(#2667): changeset fragment * chore(#2667): backfill changeset PR 2897 + correct body (cmd.exe array, not shell:true) * fix(#2667): exclude .exe from the win32 spawn-mediation gate; ack code-review.md growth CI caught two failures on the first push: 1. windows-24: 'exits 124 when the wall-clock budget is exceeded' regressed. The gate matched .exe, so the HANG command (node.exe -e 'setTimeout(...)') was wrapped in 'cmd.exe /c node.exe ...' — the wrapped child escaped the timeout cap's process-group reap (exit 124 never fired; hit the 30s harness backstop) AND cmd.exe risked mis-parsing the -e script arg. .exe is INTENTIONALLY excluded now: real PE executables spawn fine directly; only .cmd/.bat are the CVE-2024-27980 EINVAL cases. The .exe test becomes a negative-space test (node.exe spawned directly, exit 0). 2. ubuntu-22: emitted-attribution — code-review.md grew 1177 bytes from the #2667 fallow pre-pass failure-KIND case statement; acknowledge it. --------- Co-authored-by: Test <test@example.com> |
||
|
|
b54e023f05 |
Merge pull request #2899 from open-gsd/chore/sync-next-version-1.9.0
chore: sync next package version to 1.9.0 |
||
|
|
4232a79396 | chore: sync next package version to 1.9.0 | ||
|
|
efbbcc359e |
Merge pull request #2898 from open-gsd/release/1.9.0
chore: merge release v1.9.0 to main |
||
|
|
7d270a205c | chore: promote CHANGELOG for v1.9.0 | ||
|
|
63503b2fc1 | chore: finalize v1.9.0 | ||
|
|
6e0bc50142 |
fix(#2666): code-review scopes root-level + extensionless build files, cross-checks against git diff (#2895)
* test(#2666): add regression + docs-parity guards for code-review file scoper The Tier-2 SUMMARY.md extractor dropped every repository-root file (no `/`) and every extensionless build file (Dockerfile/Makefile/etc.) via an AND-joined predicate. Adds behavioral tests against the pure-function mirror plus docs-parity structural guards that bind the shipped workflow .md to the fix. RED: the docs-parity guards fail against the pre-fix shipped predicate. * fix(#2666): accept root-level + extensionless build files in code-review scope; intersect-and-warn Two coordinated edits to gsd-core/workflows/code-review.md compute_file_scope: (A) Tier-2 SUMMARY extractor: replace the AND-joined predicate `/\\//.test(raw) && /\\.[A-Za-z0-9]+$/.test(raw)` (which required BOTH a directory separator AND a trailing extension, silently dropping every root-level file and every extensionless build file) with a relaxed predicate that accepts any path with a trailing extension OR a known extensionless build basename (Dockerfile/Containerfile/Makefile/Justfile/Procfile). (B) Tier-3: convert the eq-zero git-diff gate into an intersect-and-warn — whenever a reliable diff base is available, cross-check the SUMMARY scope against `git diff --name-only` and warn about (then add) any changed files the SUMMARY extractor did not surface. Portable (bash 3.2, no associative arrays) so a partial SUMMARY result can no longer silently ship an incomplete review scope. * chore(#2666): changeset fragment * fix(#2666): use exact whole-line matching (grep -Fxq) in Tier-3 cross-check Adversarial review found the unanchored `case "$IN_SCOPE" in *"$file"$\\n*` substring membership test would false-match: a root-level `Dockerfile` in the diff substring-matches an already-scoped `docker/Dockerfile`, silently skipping it — reintroducing the exact class of silent-scope-loss bug this PR fixes. Switch to `grep -Fxq` (exact whole-line match). Add docs-parity guard for exact matching + the basename-collision regression. * fix(#2666): resolve gsd-test failures — paraphrase predicate in comment, ack code-review.md growth gsd-test caught 3 issues on 7469c3f22: 1. The docs-parity guard fired on the .md COMMENT which restated the buggy predicate verbatim — paraphrase the comment so it no longer contains the exact string the guard detects. 2. Cascade subtest failure from #1. 3. emitted-attribution: code-review.md grew 2814 bytes — acknowledge the deliberate #2666 growth in tests/emitted-drift-ack.json. * chore(#2666): backfill changeset PR number 2895 --------- Co-authored-by: Test <test@example.com> |
||
|
|
f093738412 |
fix(#2891): normalize emitted version against the measured tree, not the measuring repo (#2894)
* fix(#2891): normalize emitted version against the measured tree, not the measuring repo buildParityManifest normalized the install-time {{GSD_VERSION}} stamp using PKG_VERSION, bound at module load from the MEASURING repo's package.json. Since #2767, currentManifests({repoRoot}) measures a DIFFERENT checkout, so during a release cut the baseline worktree (origin/next, 1.8.0) was normalized with the current tree's version (1.9.0) and its literal 1.8.0 stamp survived into the hash. All 364 emitted hook paths diverged and the differential attribution gate hard-failed every finalize/rc run. Normalize against the version of the tree that PRODUCED the emitted output: buildParityManifest takes an explicit pkgVersion, and currentManifests resolves it from the measured tree via a new fail-closed measuredPackageVersion(). * chore(#2891): backfill changeset pr number (#2894) --------- Co-authored-by: Test <test@example.com> |
||
|
|
cd5b8643ae |
fix(#2828): state sync reports correct total_phases on a flat unmilestoned roadmap (#2892)
* fix+test(#2828): total_phases uses roadmap count on flat unmilestoned roadmap The read-path disk-scan cache fell back to phaseDirs.length (1) when milestoneBounded was false, even though roadmapPhaseCount (6) was correct for a flat roadmap (no sibling milestones to conflate). Use roadmapPhaseCount as the floor when > 0, matching the write-path (cmdStateSync) which already did this. The milestoneBounded flag still flows to milestoneUnbounded for the percent-skip (#1761 guard preserved). Regression test asserts state-sync writes progress.total_phases:6 for a flat 6-phase roadmap + 1 phase dir. * chore(#2828): changeset fragment * test(#2828): add negative-space coverage (Math.max floor mutant + no-roadmap fallback) — review findings The 6-phase test alone couldn't kill a Math.max-dropping mutant (1<6). Add: a 3-phase-dir/2-roadmap-phase case proving Math.max(dirs,count) floor; a no-roadmap case proving phaseDirs.length fallback. * fix(#2828): refine — distinguish flat unmilestoned from milestoned-unbounded (preserve #1761) The first-pass fix (roadmapPhaseCount > 0 always) re-broke #1761: a milestoned- unbounded roadmap (asserted milestone not among existing version headings) conflated sibling milestones (8 = 4+4). Refine with a hasMilestoneSectioning discriminator: ^#{2,3}(?!Phase) detects non-Phase h2/h3 milestone section headings. A FLAT roadmap (only ### Phase headings + a # title) has none → safe to use roadmapPhaseCount; a SECTIONED-but-unbounded roadmap has them → fall back to phaseDirs.length (#1761). Verified both cases locally (flat→6, sectioned-unbounded→1). * test(#2828): remove two fragile negative-space tests (phase-dir scanner internals) The Math.max-floor and no-roadmap tests made assumptions about the phase-dir scanner's internals (which dirs count as 'realized') that didn't hold. The core regression test (6-phase flat → total_phases:6) plus the existing #1761 conflation tests (which the refined fix preserves) provide sufficient coverage. * chore(#2828): backfill changeset PR number 2892 * fix(#2828): replace ReDoS-prone regex in regression test with line-by-line parse CodeQL flagged the nested-quantifier regex (`(?:[ \t]+\w+:.+\r?\n?)*?`) in tests/issue-2828-flat-roadmap-total-phases.test.cjs as a high-severity catastrophic-backtracking risk. Rewrite the STATE.md progress.total_phases extraction as a ReDoS-safe line-by-line block walk. --------- Co-authored-by: Test <test@example.com> |
||
|
|
54cb4145bf |
fix(#2765): bump brace-expansion to patched 1.1.18/5.0.9 (high-severity devDep advisory) (#2888)
* fix(#2765): bump brace-expansion to patched 1.1.18/5.0.9 (high-severity devDep advisory) npm audit fix (non-breaking) bumps the lockfile: brace-expansion 1.1.15→1.1.18 (eslint-nested via minimatch@3.x) and 5.0.6→5.0.9 (stryker-nested). Both 1.1.18 and 5.0.9 were published 2026-07-30 as the patch backports for GHSA-3jxr-9vmj-r5cp / GHSA-mh99-v99m-4gvg (range <=5.0.7). No overrides needed (in-range bump), no major bumps, no --force. Production (npm audit --omit=dev) unaffected (devDep only). Add a structural test pinning the installed versions so the bump can't silently regress. * chore(#2765): changeset fragment * fix(#2765): correct changeset issue ref + parse patch version as number (review findings) - changeset cited #2762 (typo) — fix to (#2765). - test compared v.split('.')[2] as a string (false-pass for 1.1.9) — parse all segments as Number. * chore(#2765): backfill changeset PR number (2888) --------- Co-authored-by: Test <test@example.com> |
||
|
|
4bd6fb066b |
chore(#2880): close ADR-2143 deployment misses — table-regex fingerprint + state-document seam migration (#2889)
* chore(#2880): close ADR-2143 seam misses — widen table-regex fingerprint, migrate state-document onto the seam The no-adhoc-markdown-parsing rule matched only a negated class whose sole member was a pipe ([^|]), so the stricter and more common [^|\n] spelling evaded it entirely -- src/state-document.cts hand-rolled exactly that shape and linted clean. Widen the fingerprint to any negated class excluding a pipe, which is the ADR-2143 section 7 prohibition as written. With the rule fixed, state-document.cts goes red. Replace tableRowPattern with locateFieldRow: a line scan using the markdown-table seam's splitTableRow for cell semantics, returning the value cell's byte range, and splice that range instead of running a whole-document content.replace. An edit now physically cannot cross a row boundary (section 4). Behavior is frozen -- stateReplaceField has 79 dependents across 5 command processes. Characterization tests lock all 14 table-branch rows plus CRLF, extract round-trip and the withFallback caller shape; a fast-check property asserts every non-target line stays byte-identical. Refs #2880, epic #2143 * fix(#2880): address adversarial review — lone-CR rows, field-name padding, quadratic scan, over-broad fingerprint Isolated adversarial review found four defects in the first commit. 1. locateFieldRow split lines on \n only. JS treats a lone \r as a line terminator, so the regex it replaced matched rows separated by bare CR. "| Phase | 3 |\r| Other | 9 |" returned 3 before and null after. Now CR, LF and CRLF are all terminators, byte offsets unchanged. 2. The field name was normalised with trim().toLowerCase(). The old regex embedded it verbatim, so its whitespace had to be absorbed by the row's own padding -- and because the group is a literal-character match rather than a whitespace class, a tab-padded cell does not accept a space-padded name. Replaced with an offset-aligned search reproducing the original backtracking exactly. 3. The widened fingerprint regex had two unbounded [^\]]* around an optional and ran quadratically over every regex source in every linted file: 256000 chars took 23 seconds. Replaced with a single-pass scanner that never rescans; the same input is now ~1ms. 4. The fingerprint also matched non-table idioms such as [^\s|] and [^"|]. Narrowed to a class excluding the pipe plus only \n, \r or \t. Differential fuzz against origin/next: 20000 cases, 0 mismatches. Refs #2880, epic #2143 * test(#2880): drop wall-clock assertion from the ReDoS regression guard local/no-elapsed-assertion flagged the elapsed-time check, and CLAUDE.md bans timing assertions outright as flaky. The 256000-char input stays as the regression guard for the quadratic scan; correctness of the verdict is what is asserted. If the quadratic path returns, the test stops completing and surfaces as a suite timeout rather than a silent pass. Also adds the changeset fragment for #2880. Refs #2880 * fix(#2880): spec-correct case folding, property tests, naming Code-review findings. The field-name comparison used toLowerCase(). The regex it replaced used /i WITHOUT /u, and ECMAScript Canonicalize deliberately does not fold a non-ASCII character onto an ASCII one -- KELVIN SIGN U+212A matched ASCII K where the old code returned null. Replaced with spec-correct Canonicalize, including the multi-character uppercase case (eszett -> SS), which a naive uppercase comparison also gets wrong. Added the fast-check property tests CLAUDE.md requires for parsers: one for the negated-class scanner, one for the field-name fold semantics, each against an independent reference implementation. Both reference impls failed on first run against real bugs, so neither property is vacuous. Renamed p2/p3 to name the exactly-three-pipes invariant, and reduced a duplicated comment to a cross-reference. Differential fuzz vs origin/next: 20000 runs, 0 mismatches, with the harness proven to discriminate the KELVIN case. Refs #2880 * chore(#2880): backfill changeset PR number (#2889) * docs(#2890): correct the local ESLint plugin path in CONTEXT.md CONTEXT.md named the local AST-rule plugin directory as scripts/eslint-rules/, which does not exist. The real location is eslint-rules/ at the repo root -- what eslint.config.mjs actually imports -- and CONTEXT.md's own later entry already says so explicitly, so the file disagreed with itself. Found by a line-by-line audit of all 1036 lines against the live graph; this was the only confirmed inaccuracy. Closes #2890 --------- Co-authored-by: Test <test@example.com> |
||
|
|
7372d99a26 |
enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales The reviewer lane roster was hand-enumerated across five documentation surfaces and three workflow files that had drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md. Adds checkReviewerDocsParity, a second pure gate deliberately separate from checkReviewerLaneParity so a stale doc cannot make the runtime checker look red. Workflows now derive their flag lists from a new review-lane flags query instead of hand-enumerating them, which also retires the unanchored grep that matched --agy inside --antigravity. Documents the previously absent reviewer body and hostBehaviors field in the capability manifest reference. Closes #2800 Closes #2781 Closes #2272 * fix(#2800): key the docs parity table arm on first-cell position Review found the flag arm was file-scoped, so the forwarding row that lists every flag in its third cell satisfied it on its own. Deleting a lane's own reviewer-table row -- the #2781 regression this gate exists to prevent -- therefore passed undetected. Arm 4 keys on the FIRST table cell, which separates a lane row from the forwarding row structurally and in every locale. Regression test included. * fix(#2800): shape-filter the flags subcommand output All three consumers read review-lane flags through an unquoted command substitution so the output word-splits into loop items. Phase 2 admits third-party overlay lanes, so an overlay flag containing whitespace would inject a second loop item and one containing a glob would expand against the cwd. Emit only well-formed flags so neither reaches the shell. * fix(#2800): remove the regex length ceiling and count only prose mentions Review found two real defects in the docs parity gate. The never-throws contract was false: building a RegExp from a declared flag or section title throws SyntaxError past ~100k chars, and Phase 2 admits overlay lanes whose declared strings are untrusted in length. Every one of these matches is literal, so String.includes replaces the regex outright, which also deletes escapeLiteral and the llama.cpp escaping it existed for. Arm 1 was context-blind: a flag mentioned only inside a fenced example or a commented-out row counted as documented. Both are stripped before matching. Also advertises all 13 lane flags in the argument-hint and corrects a stale eleven-lane count in the slug grammar note. * test(#2800): repoint the convergence suite off deleted workflow text The derived flag loop deleted the literal per-flag grep lines four tests matched on. Two of those failed loudly. The behavioral and property tests failed SILENTLY instead: their end marker no longer resolved, so the parse block extracted empty and both passed vacuously, and the property test's gsd_run stub had a no-op default that hid it. All now share one extractor and execute the real deployed block through a gsd_run shim backed by the actual binary. The whitelist assertions become an anti-parity check: re-adding a hand-written flag list must fail. Also repairs two vacuous cases in the docs parity suite. The unreadable-doc test called its own mock rather than the reader, and the integration test bounded nothing, so a doc losing its marker would have been silently skipped and still passed green. * fix(#2800): run the derived flag loop after the launcher preamble The remote matrix caught a real runtime bug, not a test artifact. In autonomous.md and plan-review-convergence.md the launcher preamble that defines gsd_run lives in a separate, LATER bash fence than the derived loop. Each fence is its own shell, so gsd_run was undefined where the loop ran: the command substitution yielded nothing and zero reviewer flags would have been forwarded. Worse than the drift this epic fixes, and silent. The whole CONVERGENCE_ARGS construction moves as one unit, because the --max-cycles append sits between the loop and the preamble and would otherwise have run against an uninitialized variable and then been dropped by the relocated initializer. Also documents all 13 lane flags in help/modes/full.md, which the repo gates bidirectionally against each command's argument-hint. * test(#2800): repoint the two converge suites off deleted flag literals Both asserted workflow.includes('--codex') against the hand-enumerated list the derived loop removed. They now assert the derivation itself, keep --all and --text (convergence controls, still literal), and add an anti-parity guard so re-adding a hardcoded list fails. The lost pass-through proof is replaced with a real one: every flag the tests used to hardcode is asserted present in the actual roster emitted by the binary, which is the property the old assertion was protecting. * test(#2800): acknowledge the workflow byte growth from the derived flag loop * chore(#2800): backfill changeset pr number to 2882 * fix(#2800): strip HTML comments to a fixed point in the parity gate CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick construction smuggles a commented-out row past the gate and it counts as documented. Not an injection risk here since nothing is rendered, but it is the exact false pass this helper exists to prevent. Strips to a fixed point, then treats any surviving opener as unterminated so the multi-line branch closes it on a later line. Terminates because every pass strictly shortens the string. * test(#2800): pin the comment-smuggling regression with a real reproducer The obvious fixture for this class does not reproduce it: <!--<!---->--> leaves a dangling --> rather than a live <!--, and is caught either way, so it would have passed with and without the fix. The join-trick construction (<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely regresses on the single-pass strip and is what the test now uses. --------- Co-authored-by: Test <test@example.com> |
||
|
|
b7b5c3712c |
fix(#2762): chunked --reviews replans instead of no-op + outline resume marker written to file (#2887)
* test(#2762): chunked --reviews must replan, not no-op (outline marker + per-plan --reviews exception) * fix(#2762): chunked --reviews replans plans instead of skipping 100% + outline resume marker written to file Defect A: §8.5.1 outline resume-check greped for a marker the agent only RETURNED (never wrote to the file) → outline always re-ran (broke crash-resume). Fix: the outline agent writes ## OUTLINE COMPLETE into the file. Defect B: §8.5.2 per-plan resume-check skipped any plan with frontmatter, no --reviews exception → --reviews skipped 100% of plans (contradicted §6 'go straight to replanning'). Fix: gate the skip on --reviews being ABSENT. Crash-resume (non-reviews) still skips. Condensed adjacent §8.5 prose to keep plan-phase.md under the 94519B cap (net -33B). * chore(#2762): changeset fragment * chore(#2762): backfill changeset PR number (2887) --------- Co-authored-by: Test <test@example.com> |
||
|
|
3af1941948 |
fix(#2772): resolve four discuss-phase text inconsistencies (dead MAX_PASSES read, gate-prompts drift, circular auto_advance, answer_validation drift) (#2886)
* test(#2772): structural guards for the four discuss-phase text inconsistencies * fix(#2772): resolve four discuss-phase text inconsistencies 1. auto.md: remove the dead MAX_PASSES/max_discuss_passes config read (contradicted the mandated single-pass rule + wasted a shim invocation per auto run). 2. gate-prompts.md: context-handling options now match the actual check_existing flow (Update it | View it | Skip, not Overwrite|Append|Cancel); gray-area-option no longer mandates 'Let Claude decide' (contradicts discuss-phase.md's no-cop-out rule). 3. discuss-phase.md: auto_advance fallback ends the workflow instead of routing back to the already-run confirm_creation step (circular). 4. discuss-phase-assumptions.md: re-sync answer_validation to the parent canonical block (had drifted — lost the 'Other' empty-text branch). * chore(#2772): changeset fragment * fix+test(#2772): also fix the assumptions auto_advance circularity (review minor 1) + add positive test anchors (review minor 2) The sibling discuss-phase-assumptions.md had the identical auto_advance→confirm_creation circularity; fix it the same way (end the workflow). Add positive anchors to both auto_advance tests so a re-phrased regression can't slip past. File #2885 for the dead max_discuss_passes config still advertised in settings/registry/docs (review minor 3). * fix(#2772): keep discuss-phase.md under the 32000B #717 cap + ack assumptions growth The auto_advance fixes + the assumptions answer_validation re-sync grew both files past the emitted-attribution gate (and discuss-phase.md past the #717 32000B cap). Condense the auto_advance prose in both files (discuss-phase.md now net -11, under cap; auto.md already net -4650 from the MAX_PASSES shim removal). Add discuss-phase-assumptions.md to tests/emitted-drift-ack.json for its residual +220 (answer_validation re-sync + auto_advance fix). * chore(#2772): backfill changeset PR number (2886) --------- Co-authored-by: Test <test@example.com> |
||
|
|
dbb0a653be |
fix(#2771): advisor mode spawns registered gsd-advisor-researcher subagent instead of general-purpose (#2884)
* test(#2771): advisor mode must spawn gsd-advisor-researcher, not general-purpose * fix(#2771): spawn registered gsd-advisor-researcher subagent instead of general-purpose in advisor mode universal-anti-patterns rule 10 (injected into discuss-phase via <required_reading>) says NEVER use non-GSD agent types. The advisor mode spawned general-purpose and manually told the agent to read the def — but gsd-advisor-researcher IS registered, so spawning by type auto-loads it. Drop the manual-read prompt line (re-specifying the def is a drift risk) and use the registered type. * chore(#2771): changeset fragment (mentions follow-up #2883) * test(#2771): widen manual-read-line regex to deny phrasing variants (review minor) /read\s+@.*gsd-advisor-researcher\.md/i (case-insensitive, any 'read @' lead-in) so a drift variant like 'Read @' or 'Load @' can't sneak the manual-def-read back in. * chore(#2771): backfill changeset PR number (2884) --------- Co-authored-by: Test <test@example.com> |
||
|
|
185da024cb |
fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block (#2881)
* test(#2770): empty contextPath argument must fail closed, not green-skip the decision-coverage gate The handler conflated empty-arg (caller error) with file-missing (legitimate skip), returning passed:true/skipped on an empty argument. Add: empty arg → passed:false; real-path-to-absent-file → legitimate green skip preserved; omitted arg → fail closed. * fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block Handler (check-command-router.cts): split the guard — empty/missing contextPath argument is a caller error (fail closed, passed:false, mirrors #1365); a real path whose file genuinely does not exist keeps the legitimate green skip. Workflow (plan-phase.md): recompute CONTEXT_PATH inside the consuming Bash block (it was set in the step-1 init block, which does not survive into the separately- spawned gate block — so the gate ran with an empty arg and silently green-skipped). * chore(#2770): changeset fragment * fix(#2770): guard workflow empty-glob case (review blocker) + update drift-guard test The handler now fails closed on an empty contextPath arg, so the workflow's unguarded glob (empty when a phase genuinely has no CONTEXT.md) would invoke the gate with an empty arg → passed:false → exit 1, hard-halting the legitimate 'Continue without context' plan-phase path. Guard the empty-glob case: only run the gate when a CONTEXT.md actually exists. Update the F1 drift-guard test (which gave false coverage — it only checked for the ${CONTEXT_PATH} token) to assert the in-block recompute AND the empty-glob guard. * fix(#2770): keep plan-phase.md under ADR-857 size cap + ack emitted drift + fix drift-guard window The workflow fix grew plan-phase.md past the ADR-857 phase-6 size cap (94519B) and triggered emitted-attribution. Condense adjacent §13a prose/JSON to offset (net +89B, under cap). Add tests/emitted-drift-ack.json acknowledging the residual growth. Widen the drift-guard test window (the gate invocation is now nested in the empty-glob guard, so the old 400-char window missed the glob recompute). * chore(#2770): backfill changeset PR number (2881) --------- Co-authored-by: Test <test@example.com> |