Commit Graph

104 Commits

Author SHA1 Message Date
Tom Boucher
779f67cb11 fix(#4378): mint collision-free SEED-YYMMDD-xxx seed ids instead of a shared count (#4754)
* test(#4378): regression tests for collision-free seed ids

* fix(#4378): mint collision-free SEED-YYMMDD-xxx ids, not a shared count

plant-seed derived the next seed id from 'ls .planning/seeds/SEED-*.md | wc -l'.
.planning/seeds/ is shared but each worktree only sees what has merged, so two
workstreams planting before either merges computed the same id and git merged
both files silently.

The id is now the local date plus a 3-char random base36 suffix -- the shape
.planning/quick/ already uses -- computed from knowledge one worktree has alone,
with a same-day regen guard. deriveSeedIdentity learns the new canonical grammar
alongside legacy SEED-NNN (whose parsing never changes), the --enrich parser and
the filename-prefix fallback keep the full new-format id, and the docs that
state the filename shape move to it.

The prefix fallback previously truncated any non-pure-numeric id at
'SEED-<digits>' -- the same one-id-two-answers ambiguity the issue reports,
reproduced one level down.

* fix(#4378): harden seed id generation per adversarial review

- parse-idea: anchor the --enrich extractor to the flag and capture the
  complete id, uppercase-tolerant; a leftmost 'SEED-[0-9]+' truncated an
  uppercase or malformed suffix to its date and enriched an arbitrary
  same-day seed via head -1. Ambiguous and unmatched targets now fail
  closed instead.
- generate-seed-id: tolerate the expected SIGPIPE under pipefail, abort
  loudly when the suffix cannot be drawn (an empty suffix would collapse
  every seed's id to the bare date), and run the same-day regen guard as
  a find existence test (the 'ls <glob>' shape trips the #3409 drift
  guard and degenerates under a stray nullglob).
- deriveSeedIdentity: document the theoretical legacy/new grammar
  ambiguity (6-digit counter + 3-char base36 slug, no frontmatter).
- changeset: state the residual same-day collision bound instead of
  implying zero.

Emitted-Drift-Ack-Growth: plant-seed.md — the counting step became hardened date+random generation with explicit failure modes; growth is the failure handling, not duplicated logic

* fix(#4378): address standards and spec review findings

- tests: move the allow-test-rule marker to its suppression site (the
  file-header placement was inert per CONTRIBUTING site-scoping); add
  width-boundary coverage (5/7-digit dates, 2/4-char suffixes pin the
  documented branch behavior); add a writer-to-reader parity property
  that parses the mint widths out of the shipped workflow so the two
  grammar owners cannot drift; cover uppercase ids end-to-end in the
  reader.
- plant-seed.md: draw/retry restructured as one loop with a loud
  terminal failure; SEED_SUFX renamed SEED_SUFFIX; regen guard drops
  the redundant head -1; the ambiguity error no longer advises an
  impossible 'complete id' for duplicate legacy ids.
- commands.cts: refresh the cmdListSeeds comment still describing
  SEED-NNN as the only canonical form.
- changeset: drop the audit claim the spec axis showed to be an
  overstatement (audit's id display is filename-derived, pre-existing).
- remove a stray untracked artifact file swept into the tree.

* test(#4378): correct boundary expectations to the module's real branch behavior

The first matrix run on the boundary tests caught my hand-trace of the
regex branches, not a module defect: the slug regex's alternation
backtracks to the legacy branch whenever the canonical branch cannot
complete (so the slug is the remainder after the legacy numeric
prefix), and the 7-digit case fails the canonical branch at its 7th
digit before the dash. Pin the verified values.

* docs(#4378): backfill changeset PR number

* fix(#4378): audit seed identity uses the canonical grammar

Review of this PR found the audit surface publishing a fused filename
stem (SEED-081-region for SEED-081-region.md) where list-seeds reports
the canonical id -- one id, two answers across surfaces, the same
ambiguity class the issue files. scanSeeds now derives identity through
the SAME deriveSeedIdentity the list-seeds gate uses (frontmatter id,
then filename id-prefix, then stem), and audit-open acknowledge
resolves --seed-id by scanning for the derived identity, falling back
to the literal stem so callers scripted against pre-canonical output
keep working. Roll-in per the fix-inline rule: found during this PR's
review, same seed-identity seam.

RED probe: pre-fix audit published seed_id SEED-081-region-becomes /
slug 081-region-becomes for a legacy seeded file; post-fix SEED-081 /
region-becomes, matching list-seeds.

* test(#4378): probe timeout uses the class norm after windows-lane timeout

The windows conformance shard failed its bounded sh -c probes at the
local 5000ms bound (cold sh.exe spawn under shard load) while the
identical code passed this PR's two earlier windows waves. The probe
now uses PROBE_TIMEOUT_MS from the class-norm module instead of a local
override, per the helpers/timeouts.cjs convention.

---------

Co-authored-by: sim <sim@local>
2026-09-15 11:41:29 -04:00
Tom Boucher
eb49ff98df fix(#4728): stop presenting the retired Gemini CLI as a supported runtime (#4743)
* fix(#4728): stop presenting the retired Gemini CLI as a supported runtime

#1928 removed the Gemini CLI runtime after Google sunset it on 2026-06-18, and
updated the ENGLISH docs. The locale mirrors and the runtime-loaded workflow
prose were not updated in the same change, and no gate asserts the ABSENCE of a
retired runtime, so both drifted quietly for a year.

The finding that shaped this change: English is already correct. docs/
ARCHITECTURE.md, CONFIGURATION.md, USER-GUIDE.md, how-to/install-on-your-runtime.md
and CLI-TOOLS.md carry zero runtime-axis Gemini references; the only English hits
anywhere are a Gemini 2.5 Pro MODEL line, the GEMINI_API_KEY row, and prose that
correctly documents the retirement. So the docs half of this is translation lag,
not a content decision, and every locale edit here is parity with an existing
English line rather than new wording:

  - install-on-your-runtime.md  English has NO `### Gemini CLI` section  -> deleted
  - USER-GUIDE.md :843          "…, Antigravity CLI, Kilo)"              -> substituted
  - ARCHITECTURE.md             English has NO Gemini CLI table row      -> row deleted
  - ARCHITECTURE.md :24         English holds `Kimi CLI` in that slot    -> Kimi CLI
  - context-monitor.md :3       "`AfterTool` for Antigravity CLI"        -> substituted
  - spike-and-sketch.md :93     "(Codex, Antigravity CLI, etc.)"         -> substituted
  - configure-model-profiles    "Codex, OpenCode, Antigravity CLI, or Kilo" -> substituted
  - COMMANDS.md                 English keeps only hyphen + Codex bullets -> colon bullet deleted
  - FEATURES.md                 source docs/features/multi-runtime-support.md:10
                                lists no Gemini CLI                       -> name removed

ARCHITECTURE.md:24 is the clearest case for reading English rather than
substituting blind: Antigravity ALREADY appears later in that list, so replacing
Gemini CLI with Antigravity would have named it twice. English holds Kimi CLI
there, so that is what the locales get.

The largest single class was hand-duplicated boilerplate. A "Text mode" paragraph
repeated across 34 runtime-loaded workflow files ends "…required for non-Claude
runtimes (OpenAI Codex, Gemini CLI, etc.)". No lint enforces that sentence and no
script syncs it, so every copy was edited. These files are read by the agent at
runtime, so they steer behavior rather than only informing a reader — which is why
this class matters more than its word count suggests.

The slash-command-form section is restructured in all four languages to match
English, which had already dropped its colon-form bullet. That bullet claimed the
colon form is "Gemini CLI only", which was false on its own terms independent of
the retirement: `/gsd:…` is GSD's canonical AUTHORING token, rewritten per runtime
at install time, and NO runtime registers it — VALID_COMMAND_STYLES is
{slash-hyphen, shell-var} and 18 of 19 runtimes declare slash-hyphen. Substituting
the runtime name would have left the claim false with Antigravity's name in it, so
the claim is gone, matching English.

Two anchor regressions were caught and fixed while doing that. zh-CN lost its
explicit {#slash-command-forms-hyphen-vs-colon} anchor while its TOC still linked
it; the anchor is restored. ko-KR and pt-BR never had an explicit anchor and rely
on the slug generated from the heading text, so shortening the heading broke their
own TOC links; those links now point at the new slugs. English's heading lost its
anchor while its TOC still links the old one — that latent English bug is
deliberately NOT copied.

Preserved, because `gemini` is not one thing here and a blanket sweep breaks the
product: ~/.gemini/antigravity{,-ide,-cli} and ~/.gemini as their parent;
~/.gemini/config (#3738); GEMINI.md; hookEvents "gemini"; GEMINI_API_KEY in all
four locales; every gemini-* model id and the Gemini 2.5 Pro references in
ko-KR/pt-BR/zh-CN (ja-JP genuinely lacks that line — the locales have diverged, so
a uniform patch would be wrong); the hook-event dialect notes, which are
RE-ATTRIBUTED rather than deleted because Antigravity inherits that dialect;
reapply-patches.md:93's legacy-install note; host-integration-capability-matrix.md
:27 and :342, which correctly record the sunset and Antigravity's contract;
whats-new-1.7.0.md and FEATURES.md:3506, which document the retirement itself; and
the generated launcher preamble, which belongs to epic #4632 — zero
_GSD_SHIM_NAME lines appear in this diff.

Coverage: a #4728 block in tests/gemini-runtime-removed.test.cjs asserts the
retired name is gone from STRUCTURAL POSITIONS (a level-3 heading, a table row's
first cell, a runtime-example parenthetical) rather than asserting the string is
absent, which would be wrong. It pairs those with positive PRESERVE assertions
over the same files — Antigravity's heading, ~/.gemini/antigravity, GEMINI_API_KEY,
AfterTool — so a patch that deletes too much fails as loudly as one that deletes
too little. The model-axis test pins both the presence in three locales and the
absence in ja-JP, so a later uniform patch that "helpfully" adds it back fails.
The new docs/ reads tripped lint-docs-guard-registration for the first time in
this file, so the test is registered in scripts/docs-guard-registry.cjs.

Not covered here, by design: nothing above would catch a Gemini-as-runtime
reference appearing in a NEW file tomorrow. That is the repo-wide drift guard,
#4729, which must land last — written now it would red on the very references this
change removes.

Fixes #4728

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4728): fix four review blockers, including a vacuous test and my own duplicate

A full matrix run on 31f12d7943 FAILED with 3 real failures, and an isolated
adversarial review returned BLOCK on four blockers. All of it was correct.

1. I committed the exact error I claimed to have avoided. The commit message
   boasted that ARCHITECTURE.md:24 proved the value of reading English rather
   than substituting blind, because Antigravity already appeared later in that
   list. Five hundred lines further down the SAME four files, my
   `Gemini:` -> `Antigravity:` substitution produced TWO consecutive
   `- Antigravity:` bullets, because an Antigravity bullet was already there.
   English (ARCHITECTURE.md:827) merges them into one. Now merged in all four
   locales, reusing each locale's existing words.

2. `--gemini` survived in the runtime-detection CLI flag list in all four
   locale ARCHITECTURE.md files. English:817 holds `--kimi` in that slot and
   already lists `--antigravity` later, so this is another place where
   substituting Antigravity would have duplicated it. Now `--kimi`.

3. Two runtime-loaded workflow files still enumerated Gemini one line ABOVE the
   line I had already corrected -- the "Adaptive (Recommended)" option in
   settings.md:192 and new-project/steps/auto-mode-config.md:95.

4. THE NEW TEST WAS VACUOUS for two of its five files. It matched only
   `non-Claude runtimes (` and `(e.g. `, and neither regex could reach the two
   lines the change actually fixed: health.md:52 reads `non-Claude (Codex, ...)`
   without the word "runtimes", and execute-phase.md:1028 has no parenthetical
   at all. The reviewer proved it by re-introducing Gemini at both lines and
   watching the assertion stay GREEN. That same blind spot is what hid finding 3.

   Replaced with a case-sensitive `/\bGemini\b/` walk over every
   `gsd-core/workflows/**/*.md`, which works because every LEGITIMATE gemini
   reference in that tree is spelled differently and cannot match: Antigravity's
   paths are lowercase with a slash (`~/.gemini/antigravity`), Google's model ids
   are lowercase and hyphenated (`gemini-3.1-pro-preview`), and the env vars are
   uppercase (`GEMINI_CONFIG_DIR`, `GEMINI_SESSION_ID`). A bare capitalised
   `Gemini` there means the retired RUNTIME is being named. The walk asserts it
   found at least 50 files so an empty walk cannot pass vacuously, and it now
   covers the nested `new-project/steps/` directory where finding 3 lived.

   Two allowlist entries, both by line CONTENT and both justified:
   reapply-patches.md's `Legacy: ... pre-#1928` note, and settings-advanced.md's
   `Known provider` menu. The second was escalated by the agent rather than
   decided: Section 8 of that file says model policy is defined "independently"
   of the runtime, so `(Claude / OpenAI / Gemini / Qwen)` is the PROVIDER axis --
   the same axis as the lowercase model ids -- and must keep working.

   Proven to fail, not just asserted: the predicate reports 0 offenders on the
   real tree and exactly 2 on a /tmp copy with Gemini re-injected at
   health.md:52 and execute-phase.md:1028.

Also from the review: a `| Gemini |` COLUMN survived in the locale FEATURES.md
comparison tables (English has none) -- removed from all three, with header,
separator and every body row kept aligned; two ENGLISH runtime-axis sites were
missed by my own parity standard (how-to/execute-a-phase.md:88 and
how-to/verify-and-ship.md:89, the latter doubly stale since #4716 retired the
Gemini reviewer lane); docs/USER-GUIDE.md:12 linked a dead anchor, which I had
found and deliberately left -- record-and-proceed on a known defect is exactly
what the rules forbid, so it is fixed; docs/COMMANDS.md:12 and all four mirrors
still claimed "the hyphen and colon forms are runtime-specific spellings" with
no colon form documented anywhere, so that false sentence is deleted; and ko-KR
had the installer rather than the user doing the targeting.

The other two matrix failures were the compact-content benchmark baseline, which
drifted because this PR changes byte counts, refreshed via the script's own
`--write` path rather than by hand; and this commit's emitted-drift-ack trailers.

Method note on the acks: the failing run measured growth against
origin/next@1110c3b4ee, which is the STALE LOCAL `next` ref -- gsd-test merges
into the local base branch, and this machine's `next` is seven commits behind
origin/next, which is checked out in the main worktree and so cannot be
fast-forwarded from here. The 32 trailers below are computed against the REAL
base (origin/next @ ca8d9d4459) by comparing each tracked file's blob size, which
is one more file than that run reported -- the extra is settings.md, grown again
by fix 3. docs-update.md and map-codebase.md are deliberately NOT acked: they
SHRANK, since there the fix deleted ", Gemini CLI" rather than substituting, and
acking a file no delta consumed is itself an error.

Refs #4728

Emitted-Drift-Ack-Growth: add-tests.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: add-todo.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ai-integration-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: check-todos.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: cleanup.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: complete-milestone.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: do.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: eval-review.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: execute-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: execute-plan.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: health.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: import.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: inbox.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: manager.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: new-milestone.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: new-workspace.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: note.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: onboard.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: plant-seed.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: profile-user.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: quick.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: remove-workspace.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: secure-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: settings.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ship.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: smart-entry.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ui-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: ui-review.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: undo.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: update.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: validate-phase.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Emitted-Drift-Ack-Growth: verify-work.md — retiring the Gemini CLI runtime name; Antigravity is one byte longer
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4728): add the changeset fragment

The PR body claimed one was present and it was not — caught by
scripts/changeset/lint.cjs reporting fail_missing_fragment, not by the
checklist, which is exactly why the lint exists.

Type Fixed: the diff is prose, and a docs-only fix uses Fixed since there is
no Documentation type.

Refs #4728

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 16:49:52 -04:00
Tom Boucher
241646a43a fix(#4651): classify .env names by final extension, and close the trailing-dot alias bypass — Phase 1 of #4636 (#4659)
* test(#4651): failing-first coverage for final-extension classification

Phase 1 of epic #4636, absorbing #4580. Tests only; no fix. These MUST fail.

The guard classifies a name by comparing everything after `.env.` as one
token against a set whose members are FINAL EXTENSIONS. So `.env.local.example`
yields suffix `local.example`, which is not a member, and a committed
secret-free template is refused. That is a category error, not strictness.

Two arms are covered because the same classification is hand-rolled twice in
one file: `isSecretBasename` for Read/Bash, and `globAltSelectsSecret`
(`lit.startsWith('.env.')`) for Grep globs. Fixing one alone would ship a
guard that allows `cat .env.local.example` while refusing
`Grep --glob '.env.local.example'` — the same file, the same hook, opposite
answers. A cross-arm parity loop over one shared list asserts the two cannot
drift.

Rows that exist because they are the ones nobody enumerates:

- `.env.example.local` must stay BLOCKED. Final extension is `local`; this is
  dotenv's documented local-override convention and a real secret. Any fix
  shaped as "contains example" admits it.
- `.env.local.` must stay BLOCKED — empty final extension is not a member.
- `.env.` must stay ALLOWED. Note #4580's proposed patch adds
  `if (suffix === '') return true;`, which flips it to blocked; that breaks the
  existing `allows` assertion in this suite and broadens the protected set,
  which epic #4636's non-goals forbid. Not applied.
- `.env.local.exam*` (partial glob literal) must stay BLOCKED — it can select
  `.env.local`, and a partial literal cannot be classified.
- `*.example` and `*` must stay ALLOWED — regression protection on the arm
  that already works.

Local behavioral repro of the current guard, confirming the tests fail for the
right reason rather than by construction:

  .env.local.example  rc=2 (blocked)   <- the defect
  .env.example        rc=0 (allowed)
  .env.local          rc=2 (blocked)
  .env.example.local  rc=2 (blocked)
  .env.               rc=0 (allowed)
  glob .env.local.example  rc=2        <- the second arm

Regressions are folded into the owning module's suite rather than a new
tests/fix-NNNN-*.test.cjs file, per scripts/lint-regression-test-names.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): classify by final extension so .env.<name>.example is readable

Phase 1 of epic #4636, absorbing #4580. Implements ADR-4650 decision 5.

The guard compared everything after `.env.` as ONE token against a set whose
members are FINAL EXTENSIONS. `.env.local.example` yielded `local.example`,
which is not a member, so a committed, secret-free template was refused — the
guard blocked the one file that exists so nobody has to open the real `.env`.

That is a category error, not strictness. The fix is not "add local.example to
the set"; it is to compare the right token. hooks/lib/filename-classification.js
now owns that distinction and is the only place it is expressed.

Both arms are fixed, because the same classification was hand-rolled twice in
this one file:

  - isSecretBasename (Read/Bash) now tests finalExtension(suffix).
  - globAltSelectsSecret (Grep --glob) split its first branch. With no
    wildcard the alternative IS a whole filename, so it is classified exactly
    via isSecretBasename. With a wildcard present the literal is only a
    PARTIAL prefix (`.env.local.exam*` can still select `.env.local`) and
    cannot be classified, so the original conservative rule stays.

Fixing only the first would have shipped a self-contradicting guard: `cat
.env.local.example` allowed while `Grep --glob '.env.local.example'` refused —
same file, same hook, opposite answers. A cross-arm parity loop over one shared
list now asserts the two cannot drift.

Two deliberate departures from #4580's suggested patch, both verified:

  - Its `if (suffix === '') return true;` is NOT applied. That flips `.env.`
    from allowed to blocked, breaking an existing assertion in this suite and
    broadening the protected set, which epic #4636's non-goals forbid.
  - `fullSuffix` was drafted alongside finalExtension and removed before
    commit: zero production consumers, and none planned (Phases 2-4 are
    containment, duplicate draining and the path-join ratchet, none of which
    classify filenames). A zero-caller export is dead code. The distinction is
    pinned instead by a test asserting finalExtension('local.example') is
    'example' and explicitly NOT 'local.example'.

The protected set is unchanged. `.env.example.local` stays BLOCKED — its final
extension is `local`, dotenv's local-override convention and a real secret;
any fix shaped as "contains example" admits it.

Scoped out by measurement, not assumption: src/validate.cts:395 and
src/phase.cts:1674 also hand-roll lastIndexOf('.'), but both parse phase
identifiers (`3.2` -> parent `3`), owned by the phase-id.cts seam. Folding
them in would repeat this same category error in the opposite direction.

Checkpoint 1 (prove RED) on the tests-only commit 91d3d6e1: outcome=failed,
26 failures / 45330, all 26 in the two new test files, zero pre-existing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): document the widened template exemption and cover the Bash arm

Two findings from the isolated adversarial review, both fixed in place.

1. The header's "Stated cost" passage named only the four literal template
   names, but since this change the exemption keys on the FINAL EXTENSION, so
   the trusted set is `.env.<anything>.{example,sample,template,dist}` — an
   unbounded family. The reviewer demonstrated it: `.env.prod-real-secrets.example`
   is allowed. That is the deliberate and necessary cost of fixing #4580, but
   it was materially larger than what the header disclosed, and a silent
   expansion of a security guard's trusted set is not acceptable. The passage
   now states the family, the concrete bypass, and that it applies across
   Read, Grep and Bash alike.

2. The cross-arm parity loop asserted Read and the exact-literal Grep glob but
   not Bash, whose `namesSecret` -> `isSecretBasename` path is genuinely
   distinct. The Bash arm was covered only by two one-off tests outside the
   shared table, so the table could not have caught a drift there. The loop now
   drives all three arms from the same TEMPLATES/SECRETS arrays.

No classification logic changed. The Read-arm behavioral table is byte-identical
before and after: rc=0 for .env.local.example / .env.example / .env. ; rc=2 for
.env.local / .env.example.local / .env / .secrets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): treat trailing dots and spaces as aliases of the protected file

Closes a Windows path-alias bypass surfaced by the isolated adversarial review
of this phase. Maintainer-approved as in scope.

Win32 strips trailing dots and spaces from every path component, so `.env.`,
`.env..`, `.env `, `.env. `, `.env .`, `.secrets.` and `.secrets ` all resolve
to the real `.env` / `.secrets` on Windows. The guard allowed every one of them
— a bypass of a file it already protects, reachable from Read, Grep and Bash
alike. `isSecretBasename` now normalizes the basename before classifying.

The whole class is fixed, not the reported name. `.env.` alone would have left
`.secrets.` and the trailing-space forms open, which is the same
one-cause-explains-every-failure trap this epic exists to close.

Two consequences, both measured rather than assumed:

  - `.env.example.` flips blocked -> ALLOWED. It aliases the already-trusted
    `.env.example` template, so this is correct; it was previously blocked only
    because the trailing dot broke final-extension parsing.
  - A Bash token that is exactly `.env` plus trailing whitespace flips
    allowed -> BLOCKED. Verified this is CONSISTENCY, not a new false-positive
    class: the bare `.env` token was ALREADY blocked as an operand in the same
    position before this change, so the alias now simply behaves like the thing
    it aliases.

The header's "No whitespace trimming" guarantee is preserved and now stated
precisely: leading and interior whitespace is still never trimmed, so prose
like a commit message mentioning `.env` in a sentence stays prose and stays
allowed. Only TRAILING dots and spaces are stripped. Two tests pin that.

This lands at the same behavior #4580's proposed `if (suffix === '') return
true;` would have produced for `.env.`, which this phase earlier rejected. The
rejection was correct on its stated grounds — that line broadens the protected
set, which epic #4636's non-goals forbid. The Windows framing is different:
normalizing an alias of an already-protected file is not a broadening, and the
fix is reached by normalization rather than by special-casing an empty suffix,
so it generalizes to `.secrets.` and the space forms.

Cannot be reproduced on this host — the remote matrix is Linux-only and Windows
coverage arrives from CI — so this ships on the Win32 path-normalization
contract plus the CI lane, and that limitation is stated rather than implied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4651): one owner for path segmentation, closing a Read/Grep divergence

Four findings from the two-axis review, all fixed in place.

The real one: the guard had TWO path-segmentation rules. `lastSegment` (used
by Read and Bash via `namesSecret`) splits on both `/` and `\`, while
`classifyGrepGlob` hand-rolled its own on `/` only. Measured:

  Read  of `config\.env`      rc=2  BLOCKED
  Grep  --glob 'config\.env'  rc=0  ALLOWED

Same logical file, opposite answers — precisely the divergence this epic
exists to remove, sitting inside the file this phase was already fixing.
`lastSegment` now lives in hooks/lib/filename-classification.js and both arms
call it. All five path-bearing cases (both separators) now agree.

Note on how this was nearly missed: the first measurement of it reported
"both allow", which looked like the reviewer was wrong. That reading was a
measurement artifact — `config\.env` inside a printf'd JSON payload is an
invalid escape, so the hook fails open at rc=0 and the test was observing
JSON breakage rather than the predicate. Re-measured with correct escaping,
the divergence is real. The tests added here use properly escaped literals
and were verified by running, not by reasoning about the escaping.

Also fixed:

  - Both fast-check properties were satisfied by a degenerate
    always-return-'' implementation: "never contains a dot / is a suffix" and
    "never ends with dot-or-space / is a prefix" are both trivially true of
    the empty string. They now additionally pin content preservation — the
    removed tail must match /^[. ]*$/, and a name with nothing to strip must
    come back unchanged.
  - The cross-arm parity loop used only bare basenames, so it could not have
    caught the divergence above. It now covers path-bearing names with both
    separators.
  - That loop's description overclaimed: Read and Bash BOTH route through
    `namesSecret`, so they are not independent paths; only the Grep glob arm
    is genuinely separate. The description now says so rather than implying
    three-way independence.
  - `normalizeWindowsBasename` runs on every platform, not only Windows. Its
    doc now states that explicitly: the guard must answer identically
    everywhere, and a name is judged by what Win32 would resolve it to.

No classification logic changed; the 12-name regression sweep is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): regenerate install-tree goldens, correct the guard's user-facing docs

Three things, all consequences of the fix rather than new behavior.

1. Install-tree goldens. `hooks/lib/filename-classification.js` is a SHIPPED
   file — package.json `files` includes `hooks` — so every per-runtime install
   tree gains a path. Checkpoint 2 failed on exactly this: 11 failures, all in
   tests/golden-install-tree.test.cjs, against 45356 passing. Regenerated via
   scripts/gen-install-tree-fixtures.cjs; 11 goldens changed, matching the 11
   failures one-for-one.

   This ripple was identified at design time and then not acted on. Fleet's
   impact preview named golden-install-tree.test.cjs before any code was
   written, and 40-design.md records it under "Ripples identified". Writing a
   risk down is not the same as discharging it, and a full matrix run was spent
   discovering something already known.

2. docs/USER-GUIDE.md made a precise and now-false claim about the guard's
   protected set: it named `.env.example` / `.sample` / `.template` / `.dist`
   as the four exempt names. The exemption keys on the FINAL EXTENSION, so the
   exempt set is the unbounded family `.env.<anything>.{example,sample,template,dist}`.
   The page now states that family, the widened residual, that order matters
   and only the last segment counts (`.env.example.local` is a secret), and
   that trailing dots and spaces are stripped because Windows resolves them to
   the protected file. A wrong user-facing model of what a security guard
   protects is worth correcting even though Fixed/Security changesets are
   exempt from the required-docs rule.

   docs/ARCHITECTURE.md and docs/INVENTORY.md say "templates such as
   `.env.example` exempt" — non-exhaustive, still true, deliberately left
   alone. Same for the ja-JP / zh-CN / ko-KR / pt-BR rows, which carry the same
   hedged phrasing; hand-translating a security description unreviewed is not
   something to do silently.

3. Two changeset fragments, not one. A refusal corrected is `Fixed`; a bypass
   closed is `Security`. Folding the second into the first would under-report
   it in the release notes. Both carry `pr: 0` for backfill once the PR exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4651): backfill changeset PR number to 4659

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 08:42:43 -04:00
Tom Boucher
2cefa5a5ac enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes (#4587)
* enhance(#4139): Phase 8 — the toggle becomes discoverable, and the ledger closes

ADR-4139's final phase. workflow.compact_content already defaulted to false
(Phase 1's buildNewProjectConfig hardcoded default), but nothing surfaced it:
/gsd-new-project never asked, and /gsd-settings/config had no toggle path for
an already-initialized project — config-set/config-get were the only route.

new-project.md gains a fourth question in the existing Round 2 AskUserQuestion
array (grouped with the other general-workflow-behavior toggles, not the
per-agent capability questions above it) and threads compact_content into the
config-new-project CLI JSON literal. settings.md mirrors the exact pattern
every other non-capability workflow.* key already follows: read_current bullet,
question block, update_config write, the safe-merge non-capability-keys list,
save_as_defaults, and the confirm summary table — seven edits, zero new
src/*.cts code, since Phase 1's merge logic is a generic passthrough. Its
success_criteria question-count ("24 settings") is bumped to 25 to match the
now-25-entry main AskUserQuestion batch.

settings-advanced.md deliberately does NOT get a duplicate question: no other
boolean toggle in this repo is asked in both settings.md and
settings-advanced.md, and there's no reason to start with this one.

docs/CONFIGURATION.md, docs/USER-GUIDE.md, and a new docs/features/4139-compact-
content.md fragment (regenerated into docs/FEATURES.md) document the toggle.

ADR-4139 itself: Status flips Proposed -> Accepted, the acceptance-criteria
section becomes a guard ledger — a 13-row table covering all 12 of #4139's
original checkboxes plus the shipped-content guard criterion, each with real
evidence (the merged PR that satisfied it, fetched via `gh issue view
--json closedByPullRequestsReferences` rather than asserted from phase
numbers) — and both "Open questions for the implementation phases" are
resolved rather than left dangling: discuss-phase was never converted to
spine+detail shape (verified: no detail/ subdir exists) — a genuine gap, not a
reasoned decline; the disjointness check is confirmed line-based by reading
compact-content-split.cjs's normalizeNonTrivialLines directly.

Orthogonal review (isolated Standards/Spec code-review + security-review
sub-agents) found and this fixes two real defects: the changeset fragment's
body didn't match CONTRIBUTING.md's single em-dash-sentence format (was
multi-sentence prose naming implementation file paths); and settings.md's own
success_criteria still said "24 settings" after the new question pushed the
main batch to 25. Also fixed, found by the Spec pass while confirming
commands/gsd/settings.md correctly needed no sync edit: that file and its
skills/gsd-settings/SKILL.md twin both still described "Interactive 5-question
prompt (model, research, plan_check, verifier, branching)", stale since long
before this phase (the batch has had far more than 5 questions for a while) —
replaced with a description that names the current set without hardcoding a
count that will drift again.

gsd-test (real run, sha 1da78fe2) caught a third real regression the local
sweep missed: new-project.md is a registered spine+detail split for Phase 4's
token-reduction benchmark (scripts/benchmark-compact-content.cjs), and the new
question's +167 tokens drifted the committed baseline
(tests/fixtures/compact-content-benchmark-baseline.json). The benchmark itself
is designed never to fail CI on drift, but the test asserting the COMMITTED
baseline is currently non-drifted correctly caught it. Regenerated via
`node scripts/benchmark-compact-content.cjs --write`; re-verified --check now
reports "up to date" and the test file passes 27/27.

Closes #4408.
Closes #4139.

Emitted-Drift-Ack-Growth: new-project.md — new 4th Round-2 AskUserQuestion entry (Compact Content, #4139) plus the config-new-project CLI JSON field and explanatory sentence; a new opt-in toggle needs new prose.
Emitted-Drift-Ack-Growth: settings.md — new workflow.compact_content read_current bullet, question block, update_config write, safe-merge key, save_as_defaults field, and confirm summary row (the same seven-edit pattern every other non-capability workflow.* toggle already follows), plus the 24->25 success_criteria count fix found in review.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* chore(#4408): backfill changeset PR number

pr:0 -> pr:4587 now that gh pr create has returned the real number.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-10 00:04:51 -04:00
Dennis Alexis Valin Dittrich
18c899def5 enhance(#4209): optional external source reviewer lanes for /gsd:code-review (#4323)
* test(01-01): define reviewer-support trait contract

Add failing coverage for step.supportsReviewerLanes (#4209 DISP-02):
validator rejects non-boolean values with an exact field path, accepts
missing/true/false, and the real code-review capability.json steps
must declare supportsReviewerLanes: true. Add loop-resolver projection
coverage proving the trait reaches activeHooks verbatim for a
provider-neutral synthetic step (not code-review-specific), and that
omitted/false values stay inert (no key on the active hook).

All 8 new assertions fail today: the validator has no such field, and
loop-resolver has nothing to project. RED before GREEN.

* feat(01-01): declare reviewer-capable steps

Add step.supportsReviewerLanes (#4209 DISP-02): a strict optional
boolean opt-in trait, step-scoped (not capability-wide). Only a
literal true validates and projects; false/omitted stay inert (no
key on the projected active hook), and every non-boolean type fails
capability-validator.cjs with an exact field-path error.

Opt both existing code-review steps (execute:post, execute:wave:post)
into the trait in capabilities/code-review/capability.json. Project
the validated field through src/loop-resolver.cts into activeHooks
so a provider-neutral generic interpreter can read it without any
code-review-specific knowledge. Document the field in
docs/reference/capability-manifest.md and regenerate
gsd-core/bin/lib/capability-registry.cjs via the generator (never
hand-edited).

Makes all 8 RED assertions from the prior commit pass.

* test(01-02): define shared reviewer dispatch

- Add tests/reviewer-step-dispatch.test.cjs covering dispatchReviewerLanes:
  inert when the supportsReviewerLanes trait is off or nothing is selected,
  exactly-once plan/invoke per selected lane, duplicate-alias dedup, the
  bounded metadata-only source-review prompt (repo root, paths+baseSha,
  depth, four fixed prohibitions), and capability-neutral reuse via a
  second synthetic step context.
- RED: module under test (src/reviewer-step-dispatch.cts) does not exist
  yet, so require() fails and every assertion is unreached.

* feat(01-02): dispatch reviewers for opted-in steps

- Add src/reviewer-step-dispatch.cts: dispatchReviewerLanes(input, deps),
  ONE interpreter for a step's supportsReviewerLanes trait. Reuses
  resolveReviewerSelection for selection and resolveLanePlan for planning
  (both already-existing, pure building blocks); invocation is the one
  required, caller-injected seam (deps.invoke) since runLane needs
  OS-aware spawn plumbing this module does not own.
- trait !== true, or a selection resolving to zero lanes, dispatches
  nothing (zero plan/invoke calls). Each selected lane is planned and
  invoked exactly once, in the selector's deduped/sorted order.
- buildSourceReviewPrompt assembles a metadata-only bounded prompt
  (repo root, canonical paths + base SHA, depth, four fixed
  prohibitions) — never file contents — written once per dispatch and
  shared across every invoked lane.
- GREEN: tests/reviewer-step-dispatch.test.cjs now passes.

* test(01-02): define reviewer dispatch failures

- Extend tests/reviewer-step-dispatch.test.cjs with the fail-closed
  matrix: an explicitly requested lane the selector could not resolve
  still lets the OTHER resolved lane run, but the aggregate result must
  never read as a clean success (and 'every explicit lane unavailable'
  must be distinguishable from the plain no-flags-passed inert case);
  request-level validation (path traversal, absolute paths outside
  repoRoot, empty/non-string paths, missing depth/base SHA) halts the
  whole dispatch before any lane is planned or invoked; a per-lane
  prompt-budget overflow hard-fails only that lane before invoke while
  its sibling still runs.
- RED: src/reviewer-step-dispatch.cts does not yet implement any of
  these guards, so 9 of the new assertions fail against the current
  (Task 1) implementation.

* fix(01-02): fail closed in reviewer dispatch

- src/reviewer-step-dispatch.cts: add the fail-closed guards the prior
  commit deliberately left out. An explicitly requested lane the
  selector could not resolve no longer lets the aggregate read as a
  clean success — lanes that DID resolve still run and keep their
  results (never narrow the requested set), but selection.errors now
  flips the aggregate ok to false, and 'every explicit lane
  unavailable' is now distinguishable (SELECTION_FAILED) from the
  plain no-flags-passed inert case (NO_LANES_SELECTED).
- Add request-level validation (validatePaths, depth/baseSha presence)
  that halts the WHOLE dispatch before any lane is planned or invoked:
  path traversal, absolute paths outside repoRoot, empty/non-string
  paths, and missing provenance are all rejected up front.
- Add per-lane prompt-budget enforcement (resolveBudget, mirroring
  gsd-tools.cjs's budgetFor convention including budget 0 = unbounded):
  a lane whose resolved budget the prompt exceeds hard-fails before
  invoke runs for it, without cancelling a sibling lane already
  planned.
- Document the supportsReviewerLanes trait and its dispatch-step
  interpreter in gsd-core/references/loop-hook-dispatch.md.
- GREEN: all 19 tests in tests/reviewer-step-dispatch.test.cjs pass;
  no regressions in the review-lane/reviewer-selection/prompt-budget
  suites (356 passing).

* test(01-03): define optional source reviewer flow

RED: assert code-review.md dispatches roster-derived reviewer-lane flags
through a single review-lane dispatch-step call (DISP-01..05), that the
no-flag path stays byte-for-behavior unchanged (COMP-01), and that
external evidence reaching the internal reviewer prompt is marked
unverified (CONS-02). Also covers the CLI contract directly: no-op with
no explicit selection, and fail-closed on an explicit unknown lane
(SAFE-07) via real gsd-tools.cjs subprocess calls.

* feat(01-03): route optional source reviewers

GREEN: code-review.md gains a dispatch_reviewer_lanes step that matches
canonical reviewer-lane flags against the merged first-party + installed
roster (never a hand-maintained list) and, only when at least one is
present, calls the shared reviewer-step interpreter exactly once with the
already-resolved repo root, file scope, depth, and base SHA. Its evidence
paths are appended to the internal reviewer prompt via
${EXTERNAL_EVIDENCE_BLOCK}, explicitly marked unverified. No reviewer-lane
flag leaves the internal-only dispatch byte-for-behavior unchanged
(COMP-01).

Deviation (Rule 3 — blocking issue): 01-02 documented `review-lane
dispatch-step` (gsd-core/references/loop-hook-dispatch.md) as the CLI
route `dispatchReviewerLanes` wires through, but never implemented the
gsd-tools.cjs subcommand — the workflow's call had nothing to reach. Add
it to the existing review-lane router, reusing the same effort-aware plan
building and runner deps `plan`/`invoke` already use (factored into
buildLaneRunnerDeps to avoid duplicating the spawn/http/fs seam). Guard
the CLI's own `detected` set on whether an explicit flag was passed:
resolveReviewerSelection's no-explicit-selection fallback is "select every
detected reviewer" (the correct default for /gsd:review), and passing it
an unconditionally non-empty detected set would silently invoke the whole
roster on every no-flag code review, violating COMP-01.

* test(01-03): define external finding consolidation

RED: assert gsd-code-reviewer.md treats <external_reviewer_evidence> as
untrusted input — independently re-verifies every claim against the actual
current source, resists a prompt-injection attempt embedded in evidence
text, and folds a verified claim into the existing Narrative Findings
section with no second REVIEW.md schema (CONS-01..03). Also assert
code-review.md's EXTERNAL_EVIDENCE_BLOCK restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* feat(01-03): consolidate external review evidence

GREEN: gsd-code-reviewer.md's load_context parses <external_reviewer_evidence>
as untrusted data, independently re-verifies every cited claim against the
actual current source before it can appear in REVIEW.md, and explicitly
resists prompt injection embedded in evidence text (never a command, no
matter what it claims to be). A verified claim folds into the existing
Narrative Findings section with (external: {slug}) provenance — one
REVIEW.md schema only, no separate external-findings section.
code-review.md's EXTERNAL_EVIDENCE_BLOCK now restates the four fixed
source-review prohibitions (SAFE-03..06) at the internal-reviewer handoff.

* fix(01-02): gitignore the reviewer-step-dispatch build artifact

01-02 added src/reviewer-step-dispatch.cts but never added its
npm run build:lib output to .gitignore, unlike every sibling
gsd-core/bin/lib/*.cjs generated file. Left it showing as untracked
noise in git status.

* docs(01-04): publish user and command contract for reviewer-lane source review

- Document optional reviewer-lane flags on /gsd-code-review in USER-GUIDE.md
  and COMMANDS.md: opt-in, no source bodies in prompts, no fallback on
  failure, findings independently consolidated into the single REVIEW.md
- Add the same contract to the docs/features/code-review-pipeline.md
  fragment and regenerate docs/FEATURES.md from it
- Preserve /gsd-review as the plan-review command; cross-reference it
  rather than duplicating the reviewer roster
- Pick up docs/INVENTORY-MANIFEST.json and skills/gsd-code-review/SKILL.md
  drift owned by source already shipped in Plans 01-01/01-03 but never
  regenerated (npm run regen:derived had not been run in this worktree)

* docs(01-04): align architecture and agent ownership docs for reviewer-lane trait

- ARCHITECTURE.md: trace the #4209 capability trait (supportsReviewerLanes)
  through the shared dispatchReviewerLanes interpreter to the existing
  review-lane plan/invoke machinery, ending at gsd-code-reviewer as the
  sole REVIEW.md consolidator
- AGENTS.md: document gsd-code-reviewer's full-context verification scope
  and its treatment of external reviewer evidence as unverified input
- No new diagram, abstraction, or config key; docs/CONFIGURATION.md is
  unchanged since the feature adds no setting or default

* fix(01-02): eslint-ignore the reviewer-step-dispatch build artifact

Same gap as the earlier .gitignore fix: 01-02 added
src/reviewer-step-dispatch.cts but never added its generated
gsd-core/bin/lib/reviewer-step-dispatch.cjs output to
eslint.config.mjs's ignore list like every sibling generated file,
so tsc's emitted __importDefault CommonJS-interop var tripped
no-var.

* fix(01-04): add the reviewer-step-dispatch.cjs roster row to docs/INVENTORY.md

01-04 regenerated docs/INVENTORY-MANIFEST.json (which now lists
cli_modules/reviewer-step-dispatch.cjs) but the hand-written roster
row in docs/INVENTORY.md — required by design, since a role sentence
cannot be generated — was never added.

* fix(01-01): update the code-review capability-step fixture for supportsReviewerLanes

refactor-trigger-cli.test.cjs's preservesCodeReviewHookShapeAlongsideRefactorHook
strict-deep-equals the code-review step's exact shape at execute:post; 01-01 added
supportsReviewerLanes: true to that step and this fixture was not updated.

* chore(01-03): acknowledge emitted-doc growth for code-review.md and gsd-code-reviewer.md

Both files grew as a direct, intended consequence of wiring optional
reviewer lanes into /gsd:code-review (the new dispatch_reviewer_lanes
step and the untrusted-evidence consolidation contract) — not
incidental drift.

Emitted-Drift-Ack-Growth: code-review.md — new dispatch_reviewer_lanes step and EXTERNAL_EVIDENCE_BLOCK wiring for optional reviewer lanes (#4209)
Emitted-Drift-Ack-Growth: gsd-code-reviewer.md — untrusted external-evidence consolidation contract for optional reviewer lanes (#4209)

* test(01-05): define WR-01/WR-02 reliability contract for dispatchReviewerLanes

From internal code review: dispatched must be false when zero lanes
actually reached plan(), and a throwing plan()/invoke() for one lane
must not discard results already collected for a sibling lane —
matching the fail-closed pattern gsd-tools.cjs already uses for the
same resolveLanePlan call (#2494/#2605/#1698/#1936/#2073/#2176/#2589/#2794).

Refs: gsd-core-dks.16, gsd-core-dks.17

* fix(01-05): close WR-01/WR-02/IN-01/IN-02 from internal review

- WR-01: dispatched now tracks whether any lane actually reached
  plan(), not results.length — an unresolvable selected slug no
  longer reports dispatched:true.
- WR-02: plan()/writePromptFile()/invoke() wrapped per-lane so a
  throw for one lane can never discard results already collected
  for a sibling lane, matching the same guard gsd-tools.cjs already
  has around the identical resolveLanePlan call.
- IN-01: documents the intentional budget===0-is-unbounded
  convention (#2797) the caller already relies on.
- IN-02: review-lane dispatch-step no longer blocks indefinitely on
  an un-piped interactive TTY; fails closed to empty paths instead.

Refs: gsd-core-dks.16, gsd-core-dks.17

* docs(01-05): add changeset fragment for PR #17

* fix(01-03): allowlist prompt-injection-scan false positive on the untrusted-evidence contract

agents/gsd-code-reviewer.md's untrusted-evidence section and its
pinning regression test both quote injection phrases as the exact
attack they defend against/detect — same
DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as the existing
allowlist entries, not an actual injection vector.

* test(01-05): extend WR-02 coverage to writePromptFile/invoke throws; DIFF_BASE-empty skip

From CodeRabbit review: WR-02's earlier fix only wrapped plan() —
writePromptFile()/deps.invoke() still ran unguarded, so a throw
there still aborted every later selected lane. Also covers the
dispatch_reviewer_lanes DIFF_BASE-empty-provenance gap (explicit
lanes silently not running when no prior review and no phase-start
commit exist).

* fix(01-05): skip dispatch_reviewer_lanes with a clear warning when DIFF_BASE cannot be resolved

Previously an explicit reviewer-lane request with no prior review and
no resolvable phase-start commit reached dispatch-step with an empty
--base-sha, which fails closed via missing_provenance — correct, but
silent about why explicitly requested lanes didn't run. Now skip
dispatch entirely in that case with a stderr warning naming the
actual cause.

* fix(01-05): wrap writePromptFile/invoke in the same per-lane try/catch as plan()

WR-02's original fix only guarded plan() — a throw from
writePromptFile() or deps.invoke() still aborted the whole dispatch,
discarding results already collected for lanes processed earlier in
the loop. CodeRabbit caught the gap; WR-02b/WR-02c pin it.

* fix(01-05): WR-02b mock must throw only on the first writePromptFile() call

The committed mock threw unconditionally, so codex's retry also threw and
failed for the same reason as claude's — the test could not distinguish
'sibling still runs' from 'sibling also breaks'. Gate the throw to the
first call, matching WR-02/WR-02c's single-failure intent.

* fix(#4209): close review findings from adversarial + critical-code-reviewer pass

Two independent reviews (agy adversarial review, Opus critical-code-reviewer +
ponytail) found 6 Blocking and 7 Required issues in the reviewer-lane dispatch
wiring around dispatchReviewerLanes. All 13 tracked in gsd-core-dks.18-30 and
fixed here:

- dispatch-step's reducer silently swallowed whole-dispatch rejections
  (invalid paths, missing provenance, etc); it now checks parsed.ok/reason.
- spawn_reviewer recomputed its own stale DIFF_BASE, diverging from the
  LAST_REVIEW_COMMIT-aware value dispatch_reviewer_lanes uses on re-review;
  now shares the single compute_file_scope derivation.
- the external reviewer prompt had no actual review request or citation
  requirement, only prohibitions; added both.
- removed the supportsReviewerLanes trait plumbing (capability registry,
  validator, loop-resolver, docs, tests) — it was never consulted by the
  real dispatch path, which gates on explicit CLI flags instead.
- flag-resolution require() was a fragile cwd-relative literal that failed
  silently on non-vendored installs; now resolves via GSD_TOOLS's own
  directory and warns instead of swallowing failure.
- reducer didn't unwrap the @file: overflow protocol for large payloads.
- deduplicated resolveBudget/budgetFor into one resolveLaneBudget.
- lane artifacts now write to a mktemp run dir instead of $PHASE_DIR, so a
  second dispatch can't overwrite prior evidence.
- validatePaths rejects control characters, closing a markdown-injection
  vector into the external prompt via crafted filenames.
- reworded the one line that tripped prompt-injection-scan.sh instead of
  allowlisting the whole production prompt file.
- fixed a stale docstring range and a dispatched-field ordering bug.
- added 3 integration tests executing the actual reducer against synthetic
  dispatch-step JSON, replacing markdown-substring-only assertions.

771/771 tests pass across every touched suite; tsc --noEmit clean.

* fix(#4209): wire supportsReviewerLanes as the maintainer's required reusable trait

The maintainer's approval on issue #4209 explicitly redirected implementation
shape: reviewer-lane dispatch must be a reusable capability/step-dispatch
trait ("supportsReviewerLanes"), not code-review.md hand-wiring the call
itself. My previous commit (e2558326) deleted that trait entirely after
finding it declared-but-never-consulted, which was backwards — the fix was to
wire it, not remove it.

Restores the trait (capability.json, generated registry, validator,
loop-resolver.cts, docs, tests) and wires it for real: dispatch_reviewer_lanes
now resolves its own active hook via `gsd_run loop render-hooks` for the
configured workflow.code_review_point and only proceeds to CLI-flag matching
when supportsReviewerLanes reads true. Explicit flags no longer bypass the
trait; a matching flag with the trait false resolves zero slugs (proven by a
new integration test executing the real fence with both trait states).

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — the
dispatch_reviewer_lanes step grows a trait-resolution fence (#4209 maintainer
redirect requires the capability layer, not the workflow, own the opt-in
decision).

* fix(#4209): dispatch-step self-verifies the reviewer-lane trait via --cap-id/--point

Both an agy adversarial review and an Opus critical-code-reviewer pass
independently found the same gap in my previous commit (9b2c3773d): the trait
check I wired into code-review.md only protected code-review's OWN
invocation — gsd-tools.cjs's dispatch-step handler still hardcoded
`trait: true` unconditionally, so a second capability declaring
supportsReviewerLanes would get zero enforcement from the shared CLI unless
it correctly re-implemented the ~15-line render-hooks scrape itself. That is
exactly the "each workflow.md hand-wiring the call" the maintainer's redirect
said to eliminate.

Moves the trait check into dispatch-step itself: given --cap-id/--point, it
self-invokes `loop render-hooks <point>` (relocating the one subprocess
code-review.md used to spawn for this, not adding a new one) and derives the
real trait from that capId's active hook, rather than trusting a
caller-passed boolean. code-review.md now only passes
--cap-id code-review --point "$CODE_REVIEW_POINT" and no longer resolves or
gates on the trait itself — the ~20-line scrape it previously carried is
gone. Any other capability opts into the identical enforcement by declaring
the trait and passing the same two flags.

Replaced the two tests that stipulated SUPPORTS_REVIEWER_LANES as an input
variable (they proved a bash branch honors a variable, not that the variable
reflects the real capability manifest) with three integration tests that
invoke the real dispatch-step CLI against the real first-party capability
registry: the real code-review trait resolves true, an unknown --cap-id
resolves false (trait_not_enabled, fail-closed), and omitting
--cap-id/--point entirely resolves false (no context means no opt-in).

Also: reject \x7f/U+2028/U+2029 in validatePaths' control-character check
(agy-F1 was incomplete), and delete the promptWritten per-lane coupling
flag — the prompt write is idempotent, so writing it once per lane instead
of gating on "did any lane write it yet" removes a latent bug where a
deps.plan override that ever varies promptPath per lane would silently skip
writing for a later lane.

Emitted-Drift-Ack-Growth: gsd-core/workflows/code-review.md — net line count
drops (the trait scrape moved into dispatch-step), but the file still grew
this session across multiple commits; acknowledging per the growth-tracking
convention.

* fix(#4209): remove per-run token waste from the shipped prompts

Runtime prompt content, not session tokens: two real, per-invocation token
costs in the code that ships.

1. agents/gsd-code-reviewer.md's critical_rules restated nearly all of
   load_context step 5's ~180-word untrusted-evidence contract in ~90 more
   words, breaking this section's own established terse one-liner style
   (every other rule here is 1-2 sentences). This prompt loads fresh on
   every /gsd:code-review invocation. Shrunk to a one-line cross-reference,
   matching how write_review's own reference to step 5 already does it.

2. buildSourceReviewPrompt repeated the base SHA on every single file line
   even though it is identical for every file and already stated once at
   the top of the prompt — O(files) wasted tokens on every dispatched lane
   for a 50-file review, for zero information gain. File lines are now bare
   paths.

* fix(#4209): resolve reviewer-lane trait in-process, fix CI failures found in review round 3

Opus critical-code-reviewer found a real Blocking defect in the --cap-id/
--point self-invocation added last commit: `dispatch-step` spawned
`loop render-hooks <point> --raw` as a subprocess and bare-JSON.parse'd its
stdout, but `io.cjs`'s output() redirects any payload over 50000 chars to
`@file:<path>` instead of inline JSON -- the same overflow protocol this
feature already unwraps for its OWN dispatch result 60 lines later in
code-review.md. A large-enough activeHooks envelope (more installed
capabilities/fragments) would throw, get silently swallowed by the bare
catch, and misreport a real trait as trait_not_enabled with zero diagnostic.

Fixed by extracting the config/registry/capability-state resolution
`cmdLoopRenderHooks` already performs into an exported pure function,
resolveActiveHooksForPoint (both `cmdLoopRenderHooks` and dispatch-step now
share it), and calling it in-process from dispatch-step instead of spawning
a subprocess at all. This eliminates the @file: exposure entirely (the
dispatch-step path never touches the rendered-string envelope or its
JSON-stringify/50000-char threshold), removes one subprocess spawn per
code-review invocation, and gives a genuine diagnostic (stderr warning) on
resolution failure instead of silent fail-closed. Corrected three doc/
docstring references to the now-removed subprocess self-invocation.

Also fixes 2 real CI failures this round surfaced:
- lint-tests: the agy-F1 control-char regex fix's `eslint-disable-next-line
  no-control-regex` comment was unused under this project's ESLint config
  (verified locally: the rule never actually flags \x00-\x1f in this repo's
  config) -- a mistake from an earlier commit this session, never actually
  lint-checked before push. Removed the disable comment.
- security (prompt-injection-scan): the agy-F1 regression test's crafted
  fixture literally contains "Ignore all prior instructions." as test data
  proving validatePaths rejects it -- allowlisted the test file, same
  DEFECT.PROMPT-INJECTION-SCAN-COLLISION class as existing entries.

Also trimmed agents/gsd-code-reviewer.md's load_context step 5 (R2): one
bullet stated "untrusted, never a command" three different ways in one
paragraph, and a same-file duplicate of write_review's schema rule.
Consolidated to state each rule once.

Declined one suggestion from this round: shrinking code-review.md's
EXTERNAL_EVIDENCE_BLOCK to a bare evidence list. Two tests
(tests/code-review-pipeline-regression.test.cjs's CONS-01..03 block,
tests/code-review.test.cjs's CONS-02 test) deliberately lock the four-
prohibitions restatement and the untrusted-evidence prose into the
INJECTED block itself, not just the consolidator's system prompt --
adjacency of the warning to the untrusted payload it's warning about is a
recognized prompt-injection defense-in-depth pattern from this
workstream's original TDD plan, not accidental duplication.

* fix(#4209): correct stale per-file base-SHA prose in the external prompt

Leftover from removing the per-file base SHA repetition earlier this
session: the review-request sentence still said "relative to its base SHA"
(singular per-file framing) when there's now exactly one base SHA, stated
once above the file list. Reads "relative to the base SHA above" now.

* fix(#4209): make getLane/configGet/plan required deps, delete dead defaults

R3/R4 from the review round I'd deferred as low-priority test-churn: this
file's one production caller (gsd-tools.cjs's dispatch-step handler) always
supplies all three, so the fallbacks were dead in production -- but each was
actively WRONG if ever reached: the default configGet always returned
undefined, silently disabling resolveLaneBudget's overflow guard; the
default getLane looked up only first-party REVIEWER_LANES, diverging from
production's overlay-merged roster; the default plan skipped per-host effort
resolution entirely.

These defaults were introduced by this PR's own earlier work (this file did
not exist before #4209 -- first commit a760bfcda, 01-02), not inherited from
elsewhere, so there's no external caller depending on the lenient contract.

Turned out free to fix: making the three deps required and deleting
defaultGetLane/defaultPlan needed zero test changes -- every existing test
that actually reaches the per-lane loop already supplies getLane/plan
explicitly, and configGet's only real dependent (the budget-overflow tests)
already supplies it too. 788/788 tests pass unchanged, tsc/lint clean.

* fix(#4209): define depth semantics for the external reviewer lane

Verified this was a real bug, not a match to existing convention as I'd
claimed when declining the suggestion earlier this session: the internal
gsd-code-reviewer agent's own system prompt carries a full <depth_levels>
block defining what quick/standard/deep mean and do (agents/gsd-code-
reviewer.md:68-99). The external reviewer lane has no access to that
persona at all -- it only ever sees buildSourceReviewPrompt's bounded text,
which sent the bare depth label with zero definition to a third-party CLI
with no other source of truth for what "standard" means.

Added depthMeaning(), condensed from the internal reviewer's own
<depth_levels> definitions so the two stay consistent, and interpolated it
into the review-request sentence. 150/150 tests pass, tsc/lint clean.

* fix(#4209): merge dispatch_reviewer_lanes' split fences into one shell invocation

CR-01 (Opus critical-code-reviewer, confirmed by direct execution): the
roster-matching fence set EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS, and a
SEPARATE later fence read them via ${#EXPLICIT_REVIEWER_SLUGS[@]} to decide
whether to dispatch at all. This file's own documented rule (its
depth-resolution guard, stated explicitly a few hundred lines earlier) is
that a guard and the extraction it protects must run as one shell
control-flow decision, because markdown-fenced blocks do not share shell
state -- this step violated its own file's rule for the entire feature's
gating condition.

Merged the roster-resolution fence and the dispatch-decision fence into one
continuous bash block, removing the intervening prose that split them.
Fixed the stderr-based failure detection in the same edit (RQ-01: checking
whether stderr is non-empty misfires on any benign Node warning; now checks
the actual exit status of the roster-resolution command).

Verified by extracting the merged fence and executing it standalone, driving
both branches: --codex resolves EXPLICIT_JOINED=codex, SLUGS_COUNT=1, and a
real dispatch-step call succeeds; no flags resolves EXPLICIT_JOINED empty,
SLUGS_COUNT=0, dispatch-step never invoked (COMP-01). 141/141 workflow tests
pass, tsc/lint clean.

* fix(#4209): depthMeaning accuracy, injection defense on all embedded fields, hoisted prompt write

Batch of Required/Suggestion fixes from the Opus critical-code-reviewer +
writing-for-agents pass:

- CR-02/CR-03: depthMeaning() dropped real categories from quick (empty catch
  blocks, commented-out code) and deep (error propagation, state mutation
  consistency, circular dependencies) relative to the real <depth_levels>
  block, and had zero test coverage. Restored full accuracy and added tests
  that read the real agents/gsd-code-reviewer.md file directly, so drift
  between the two can't recur silently. Unrecognised depth now normalizes to
  standard's definition, matching that agent's own documented rule, instead
  of rendering an undefined bare label.

- RQ-04: depth/baseSha/repoRoot/runDir land in the same markdown prompt
  `paths` does, but weren't checked for control characters like paths were
  (agy-F1's original finding). Hoisted CONTROL_CHAR to module scope and
  applied it to all four fields at the same provenance-check boundary.
  runDir previously had zero validation at all.

- S1: deleted the dead `identity` parameter on `invoke` -- the one production
  caller already ignores it, no test read it by name.

- S2: hoisted the shared prompt write above the per-lane loop -- promptPath
  is derived from runDir alone (constant across lanes by construction), so
  writing it once is both correct and cheaper than the per-lane write R1
  introduced earlier this session. Discovered and fixed a real regression
  from the naive version of this hoist: an unguarded throw would have
  escaped dispatchReviewerLanes as an uncaught exception instead of a clean
  per-lane failure. Added a new PROMPT_WRITE_FAILED whole-dispatch reason,
  matching the existing validatePaths/MISSING_PROVENANCE halt pattern, with
  a dedicated regression test.

- S3: moved `planned = true` past the budget-overflow gate, so `dispatched`
  only reports true once a lane has cleared BOTH plan and budget checks.

- S5: relayed gsd-code-reviewer.md's own "performance issues are out of
  scope unless also correctness issues" policy into the external-lane
  prompt, which previously had no such guidance and could return findings
  the internal reviewer's own contract excludes.

- RQ-05 (partial): shrunk this file's own header docstring's restatement of
  the trait-reuse architecture to a pointer at
  gsd-core/references/loop-hook-dispatch.md, the canonical home.

234/234 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): dedupe roster-merge logic, consolidate trait architecture prose, add step completion criterion

RQ-02: added a `review-lane explicit-from-argv` subcommand that reuses the
SAME merged-roster logic (`laneBySlug`) `dispatch-step`/`plan`/`invoke`
already share. code-review.md's ~18-line inline `node -e` reimplementing
`loadRegistry`+`mergeReviewerLanes` (a rename-only copy of the block in
gsd-tools.cjs) is now a single call to this subcommand -- the exact
violation code-review-flags.cjs's own header warns against ("this is the
canonical flag-parsing surface -- do not replicate inline bash parsing").

RQ-03: an empty --cap-id XOR --point now warns distinctly from the
legitimate no-context opt-out (both absent) -- a caller that named a
capability without its point was silently indistinguishable from a correct
opt-out. Also hardened the CODE_REVIEW_POINT config-get fallback: it only
ever fires when the config-get COMMAND ITSELF fails (config-get already
resolves the manifest's own schema default in the normal case), but that
failure was previously silent.

RQ-05/W-01/W-12/W-13: the "supportsReviewerLanes is a reusable trait
resolved inside dispatch-step" explanation was restated in full in 5
places across this session's own review cycles. Consolidated to ONE
canonical statement in gsd-core/references/loop-hook-dispatch.md; the other
4 (this file's own header, gsd-tools.cjs's comment, docs/ARCHITECTURE.md,
code-review.md's step-opening comment) now point at it instead.

W-05/W-06: loop-hook-dispatch.md described "false or non-boolean" as two
inert cases when capability-validator.cjs already rejects non-boolean at
load -- restated as the two cases that actually reach this code. Removed a
"do not hand-roll trait resolution" prohibition whose target no longer
exists once the positive description precedes it.

W-04: deleted a no-op sentence in agents/gsd-code-reviewer.md ("missing
block means proceed as normal") -- an absent optional block already means
proceed as normal without being told.

W-08/W-09: replaced longhand "zero selection/plan/invoke calls" and the
made-up compound "byte-for-behavior [un]changed" with the token this
session's own docs already coined for this concept (inert) and the word
that means what byte-for-behavior was reaching for (unchanged).

W-10: dispatch_reviewer_lanes had no completion criterion -- added one
sentence naming the checkable end state (EXTERNAL_EVIDENCE_BLOCK is set,
either populated or empty). This exact sentence would have caught the
cross-fence bug fixed two commits ago at authoring time.

Declined from this round, with reasoning: W-02/W-03 (trim the
untrusted-evidence restatement in EXTERNAL_EVIDENCE_BLOCK/critical_rules) --
two tests deliberately lock this as intentional adjacency-based
prompt-injection defense-in-depth, not accidental duplication (see this
branch's own earlier commit). S4 (wrap LANE_RUN_DIR in a creation-site
`trap ... EXIT`) -- would fire at the end of the CREATING fence, before
spawn_reviewer's agent ever reads the evidence files, given this file's own
documented fenced-block execution model; the existing named cross-reference
between creation and cleanup already satisfies the co-location concern
without introducing that regression.

853/853 tests pass across the full reviewer-lane test suite, tsc/lint clean.

* fix(#4209): merge CODE_REVIEW_POINT into dispatch_reviewer_lanes' one fence, stop test from spawning real codex

Round-5 review (agy) found the same cross-fence-split bug CR-01 already fixed
for EXPLICIT_JOINED/EXPLICIT_REVIEWER_SLUGS: CODE_REVIEW_POINT's config-get
fallback lived in an earlier, separate fence from the fence that consumes it
via --point, split only by prose (not a guard, per this step's own documented
rule). Merged into the single continuous fence and added a structural test
asserting exactly one bash fence in the step.

The new end-to-end regression test for this used --codex, which drives the
fence's real `review-lane dispatch-step` call and, with the codex binary
present on PATH, spawns the real external CLI — which then blocks on
interactive auth with no stdin (BL-01). Stubbed gsd_run for
`review-lane dispatch-step` only (captures argv instead of executing),
keeping the real config-get/explicit-from-argv calls the test is actually
about.

* fix(#4209): split control-char vs missing provenance reason, realpath-check path escapes, stale comment

Round-5 review (Opus) warning-tier findings:

- WR-04: MISSING_PROVENANCE covered both "field absent" and "field present but
  a control-character injection attempt" — a caller distinguishing a config
  problem from a security event couldn't tell them apart. Split into
  MISSING_PROVENANCE (absent) and INVALID_PROVENANCE (present but invalid).
- WR-05: validatePaths' containment check was lexical only (path.resolve),
  so a symlink whose own path sits inside repoRoot could still point outside
  it. Added an fs.realpathSync check (ENOENT-tolerant — a git-diff path can
  legitimately name a file already deleted in a stale worktree), realpathing
  repoRoot itself too so a symlinked repoRoot (e.g. /tmp on macOS) doesn't
  false-positive-reject its own real children.
- WR-08: a comment in the per-lane loop still said a throwing writePromptFile()
  was caught there — stale since the prompt write was hoisted above the loop
  in an earlier round.

WR-03 (validate depth against the quick/standard/deep enum) was considered
and declined: this dispatcher is deliberately capability-neutral (see the
existing "synthetic step context" test, which passes a non-code-review depth
label on purpose to prove no code-review-specific special-casing exists).
WR-01 (double registry load), WR-02 (trim-vs-hard-fail budget semantics), and
WR-07 (reason omitted on the aggregate return) were verified against source
and are not bugs — see review notes.

* docs(#4209): document LANE_RUN_DIR's early-exit trade-off as accepted, not a gap

Round-5 review (Opus, BL-03) flagged that an early exit between
dispatch_reviewer_lanes and commit_review leaks the run-scoped temp dir. A
trap-based cleanup was considered and rejected: if a step genuinely runs as
a separate process, a trap set at creation time would fire at the end of
that SAME fence, deleting the directory before spawn_reviewer/commit_review
ever read it — worse than the leak it would fix.

review.md's own gather_context/cleanup pair for the identical resource class
(a run-scoped reviewer temp dir) already makes and documents this exact
trade-off: cleanup runs only on a documented success path, and a leftover
$TMPDIR entry is explicitly called cheaper than destroyed evidence. Recording
that precedent here so this isn't re-raised as a live gap in a future review.

* fix(#4209): register the WR-05 symlink-escape test's synthetic docs/ path

reviewer-step-dispatch.test.cjs's "capability-neutral reuse" fixture passes
paths: ['docs/spec.md'] as a synthetic, never-read path proving the
dispatcher has no code-review-specific special-casing. lint-docs-guard-
registration correctly flagged this as an unregistered docs/ path reference —
add the docs-guard-exempt marker and its pinned baseline entry, the same
pattern every other synthetic docs/ literal in this test suite already uses.

* fix(#4209): backfill changeset pr: field with the real upstream PR number

changeset-lint's fail_pr_field_drift caught the fragment still pointing at
the fork PR (17) instead of the upstream one (open-gsd/gsd-core#4323) this
branch is now also open against.

* docs(#4209): amend ADR-2782 for the supportsReviewerLanes step-trait seam

trek-e's review (2026-09-07, gsd-core#4323) found a real ADR gap: every
decision in ADR-2782 (D1-D9) and every prior dated amendment governs the
`role: "reviewer"` capability body and its one consumer, /gsd:review. This
PR's actual new seam - a `supportsReviewerLanes: true` trait on an ordinary
feature capability's `steps[]` entry, projected through loop-resolver.cts
and resolved in-process via resolveActiveHooksForPoint - is a different
capability axis (steps/gates/contributions) that the ADR's own scope note
explicitly places out of reach. Per docs/contributor-standards.md's
"Amending an accepted ADR", an in-place dated section is the established,
lighter-weight path for an addition that stays within the ADR's existing
decisions - used twice already in this same file - so this appends a third
dated entry documenting the new seam, its consumer, and why it reuses the
existing D1-D9-governed plan/invoke machinery rather than adding a second
one. No decision is reversed; no new Amends/Amended-by pair is needed since
the steps/gates/contributions axis already carries reciprocal links to
ADR-857 and ADR-894.

* fix(#4209): close two test-quality gaps trek-e's review found

Minor 1: validatePaths (a path-shape parser guarding the prompt-
injection/path-traversal trust boundary) had only example-based coverage,
violating ADR-456's rule that parsers/budget limits carry at least one
fast-check property test. Adds three: safe-segment paths are never
rejected, a single leading "../" always escapes the one-segment repoRoot,
and a control character anywhere is always rejected - one property per
rejection reason validatePaths owns.

Minor 2: the budget-overflow check (`estimatedTokens > budget`) was only
ever exercised far below budget or at budget:0 (unbounded), never at the
exact threshold crossing where a `>` vs `>=` off-by-one would hide. Adds
three exact-boundary tests using the real estimateTokens/
buildSourceReviewPrompt the module calls internally, so the resolved
token count is exact rather than approximated: budget == estimate (must
pass), budget == estimate - 1 (must fail), budget == estimate + 1 (must
pass).

Also extracts okPlan()'s fixture timeoutMs into a named constant -
local/no-adhoc-timeout-literal (#4446) landed on next after this branch
was authored and flagged the pre-existing literal on rebase; it is fixture
data for a synthetic plan object dispatchReviewerLanes never waits on, a
distinct class from tests/helpers/timeouts.cjs's real subprocess norms.

* fix(#4209): update docs-guard-registration baseline for the new ADR citation

reviewer-step-dispatch.test.cjs's new fast-check property tests cite
docs/adr/456-test-rigor-architecture.md in a justifying comment (never a
real read). lint-docs-guard-registration fingerprints every docs/ path
string an exempted test file mentions and fails on drift so a human
re-confirms the exemption still holds - re-confirmed, and the baseline is
updated to match.

* fix(#4209): point changeset pr: field at the fork PR for CI validation

changeset-lint's fail_pr_field_drift check compares the fragment's pr:
field against the PR the CI run is actually attached to (GITHUB_EVENT_PATH),
not a fixed target. Rehearsing this branch on fork PR
davdittrich/gsd-core#17 needs pr: 17 to pass that check; the prior commit's
pr: 4323 (the real open-gsd upstream PR number) is correct for that PR but
fails here. Backfill to 4323 happens again, as the last commit, immediately
before the approved push to open-gsd#4323 - never leaving pr: 17 on the
branch that ships upstream.

* fix(#4209): reject promptChannel:none lanes from source-review dispatch

CodeRabbit found a real scope mismatch: coderabbit's lane declares
promptChannel: 'none' and reviews the working tree on its own terms,
fed nothing (review.md:367). Silently dispatching it through
dispatchReviewerLanes would ignore the bounded paths/depth/baseSha scope
buildSourceReviewPrompt promises and let the lane review whatever it
independently sees fit, violating this interpreter's own scoped,
metadata-only contract. Reject before plan()/invoke(), same as an
unresolved slug.

* fix(#4209): scope CONS-02 test to the evidence-block line, not the whole file

CodeRabbit found the whole-file match on workflowContent would still
pass if UNVERIFIED and re-open/reopen appeared in two unrelated parts
of this 1000+-line workflow, proving nothing about the actual evidence
block's contract. Line-filtered via splitLines (not a bare-\n regex
spanning readFileSync content) so this stays CRLF-portable and passes
local/no-unbounded-quantifier and local/no-crlf-fragile-split.

* fix(#4209): guard DISPATCH_JSON substitution and capture its stderr

CodeRabbit found the dispatch-step command substitution unguarded: a
non-zero exit could leave DISPATCH_JSON empty (or halt the step under
errexit with no warning), and the downstream reducer would only ever
report the generic unparseable_dispatch_output reason, discarding the
command's own diagnostic. Guarded like the existing CODE_REVIEW_POINT/
EXPLICIT_JOINED calls above it: capture stderr to a temp file, surface
it in a warning on failure, and fall back to a parseable dispatch_
command_failed JSON stub so the reducer's existing reason-reporting
path still fires.

* docs(#4209): fix byte-for-behavior wording and missing colon, regenerate

CodeRabbit found "byte-for-behavior" should read "byte-for-byte" (the
established repo term for output-identical unchanged behavior) and a
missing colon after the bold "Optional external reviewer lanes (#4209)"
lead-in in docs/features/code-review-pipeline.md. Fixed in the two
hand-authored sources (commands/gsd/code-review.md, docs/features/
code-review-pipeline.md) and regenerated the two derived projections
(skills/gsd-code-review/SKILL.md via gen-plugin-skills.cjs, docs/
FEATURES.md via gen-features.cjs) so they stay in sync.

* fix(#4209): drop the fabricated DISPATCH_JSON fallback stub (Windows CI)

The prior fix's fallback `DISPATCH_JSON='{"ok":false,...}'` embeds
double-quoted JSON keys inside a single-quoted shell literal. That
extra quote density, inside an already quote-heavy ~8KB driver string,
passed bash -n and the full local suite on Linux but broke Windows
Git-Bash: `dispatch_reviewer_lanes computes CODE_REVIEW_POINT ... end
to end (#4209 round 5)` failed on two Windows CI shards with `bash -c:
unexpected EOF while looking for matching '''` — a Windows argv-to-
command-line re-quoting edge case, reproducible on rerun, not a flake.
Root-caused via gh api job logs plus a byte-identical local
reconstruction of the test's own driver script.

Fix: drop the fabricated stub. The downstream node -e reducer already
falls back to reason `unparseable_dispatch_output` on any JSON.parse
failure, so an empty/partial DISPATCH_JSON on command failure is still
handled correctly, with zero new quoting risk.

* revert(#4209): drop the DISPATCH_JSON stderr-guard nitpick (Windows CI)

Two materially different mechanisms for the same CodeRabbit Nitpick
("Trivial | Quick win") both broke Windows Git-Bash reproducibly:
a single-quoted JSON-literal fallback ("bash -c: unexpected EOF ...
matching '''") and, after removing that, a plain `head -1 "$VAR"`
inside a nested command substitution ("unexpected EOF ... matching
'"'"). Both passed bash -n and the full local suite on Linux every
time; both failed the SAME test deterministically on Windows CI. Two
attempts at the same class of fix (nested-quote construction near
this exact step) is the retry limit - reverting to the original,
already-shipped, Windows-verified unguarded form rather than
continuing to guess at a third quoting mechanism for a Trivial-
severity nitpick. Logged as bug-221/bug-222 in .wolf/buglog.json for
anyone attempting this again: the fix belongs outside this specific
markdown-fence-driver test harness (e.g., a real .sh helper script)
if it's worth doing at all.

* fix(#4209): backfill changeset pr: field to the real upstream PR before push

Fork validation (davdittrich/gsd-core#17) needed pr: 17 to satisfy
changeset-lint's PR-number check while rehearsing there; this is the
last commit before the approved push to the real upstream PR
(open-gsd/gsd-core#4323), so the field points at that PR number again.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-07 22:52:33 -04:00
Cody Anderson
77e2472ca0 enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration

Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on
Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets
(the .env.example/.sample/.template/.dist templates stay readable).
Read checks file_path; Grep checks an explicit path and judges the glob
per brace alternative; Bash runs a two-pass token scan (quotes, comments,
redirects with fd digits, separators, $( )/backtick/<( ) recursion,
heredoc bodies never scanned as commands, nested bash -c/eval rescans,
git <ref>:<path> shapes) with a closed non-reading exemption set for
existence checks. Fail-open crash policy; 1 MiB commands are denied as
command-too-large; more than 64 glob alternatives as glob-too-complex.

Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt
for approval whenever any Read() deny rule exists, even in auto mode. A
hook denial is not a permission rule and never arms that check. The
installer-written deny rules are retired in the follow-up commit.

Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks
HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking
guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell),
shell-command-projection managed sets, installer-migration-report,
OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch),
docs tables in five locales, ADR-766 always-on list, regen:derived
fixtures, and a new table-driven unit suite.

* test(#4221): pin the secret-read guard in existing hook gates

Register gsd-secret-read-guard.js in every existing hook gate: the
hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest
REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity
EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal-
hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS,
kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and
typed-payload floors, the OpenCode adapter (grep mapping, include ->
glob, three dispatch tests) and a Kimi TOML matcher assertion.

* fix(#4221): retire installer Read() deny rules (legacy filter)

Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS
and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets)
strings. mergeClaudePermissions now only filters them out of an existing
permissions.deny: an absent deny key stays absent, a malformed one is
still repaired to [], and an array emptied by the filter is deleted so
no `"deny": []` residue is left. Uninstall filters the same legacy list
and, symmetric with the Antigravity branch, drops an emptied allow or
deny key and an emptied permissions object.

Unlike the #2278 allow-side migration there is no surviving current
deny list, so the constant is renamed rather than mirrored. Removal is
byte-exact: a hand-written identical rule is indistinguishable from the
installer's and is removed too (the manifest never recorded permission
strings). USER-GUIDE and CONTEXT.md updated.

* test(#4221): flip install-regressions deny-rule assertions to the retired shape

The fresh-merge, non-destructive merge, idempotency, end-to-end install,
reinstall and uninstall assertions now expect no Read(.env*) deny rules
and no permissions.deny key on a fresh install; the deny:null repair case
is kept. A new describe block covers the legacy filter: retired strings
removed with a user entry kept, partial sets, near-miss strings
untouched, idempotency, GSD-only deny array deleted, a pre-existing
empty deny preserved, and uninstall symmetry for allow/deny/permissions.

* chore(#4221): add changeset fragment for PR #4236

* fix(#4221): case-fold names; scan shell stdin and xargs pipes

Review round 1 (trek-e):

- Blocker: secret-name matching is now case-insensitive in the Read,
  Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a
  case-insensitive filesystem are recognized as the same secret file.
- Major: a shell interpreter's script is now scanned wherever it comes
  from. The tokenizer keeps heredoc bodies as per-segment tokens and
  records separator operators; pass 2 groups by segment id and resolves
  bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined
  `-lc`) scans the script operand, a file operand is checked as a file
  (a `<( )` operand's echo/printf output is reconstructed), otherwise
  stdin is the script and heredocs, here-strings and a piped echo/printf
  source are scanned. `eval` joins all its operands; `source`/`.` handle
  process substitution. Data heredocs (`cat <<EOF`, the commit-message
  shape) stay unscanned.
- Major: `… | xargs <cmd>` checks the upstream segment's operands as
  file names when the sub-command reads (`echo .env | xargs cat`,
  `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the
  inference; a shell sub-command's `-c` script is scanned.

Header, USER-GUIDE bullet and changeset updated; documented gaps now
include piped scripts from non-echo sources and `exec`/`timeout`
wrappers. 60 new suite cases pin the block and allow shapes.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-05 04:00:08 -04:00
Tom Boucher
cf15682d1c enhance(#3028): responsive Markdown separators instead of fixed-width rules (#3789)
* feat(#3028): responsive Markdown separators instead of fixed-width rules

Stage banners, checkpoints, completion and error panels used fixed-width
runs of box-drawing characters -- a 53-column heavy rule and a 62-column
double-line box. Those runs are ordinary text to a Markdown-rendering
host, so in a narrower pane they wrap and the border comes apart from
the heading it framed.

Shipped content now emits an ATX heading for a titled section and a
blank-line-delimited --- for a break between sections, both of which
adapt to the available width. The same convention is applied to the
three code sites that built these strings at runtime: the UAT
checkpoint renderer, the milestone-close audit report, and the TDD
review checkpoint table.

Removing the box also removes its only reason to exist -- the
east-asian-width padding helpers that kept its right border aligned
(checkpointBoxLine, displayWidth, isWideCodePoint, ZERO_WIDTH_MARK_RE,
CHECKPOINT_BOX_WIDTH). RTL directional isolation is unchanged.

The convention is specified in gsd-core/references/ui-brand.md and
enforced across all shipped content by tests/responsive-separators.test.cjs.

Refs #3028

* test(#3028): pin the heading form in checkpoint and audit-report assertions

These suites asserted the exact box borders and the 62-column padded
banner interior. With the box gone they assert the ### heading form,
the --- break and the bolded instruction line, and each now carries a
positive assertion that no box character remains -- which is what pins
the fix rather than merely tolerating it.

Language coverage is converted, not dropped: Japanese, Chinese, Korean,
Hindi and Arabic all still assert their rendered banner, and the Arabic
case still asserts the RTL directional isolates the box removal must
not disturb. Adds a case for a banner longer than the old inner width,
which previously produced a ragged border and now has none.

Refs #3028

* chore(#3028): acknowledge execute-plan.md growth from the checkpoint display spec

The checkpoint_protocol display spec described the drawn box; it now
describes the heading, the --- break and the bolded action prompt,
which costs 22 bytes (40111 -> 40133, 827 under the cap).

Appended to the existing #3370 fragment rather than filed as a new one:
a growth ack keys on the bare filename and #3370 already declares
execute-plan.md, so a second source naming it would be a hard
duplicate-key error. Same supersede-by-append route #3370 took for the
spent #2652 fragment.

Refs #3028

* docs(#3028): state the load-bearing half of the separator rule, and amend the zh-CN reference

Review found three things.

The rule as first written demanded a blank line above AND below every
---. Only the one above is load-bearing: it is what stops CommonMark
reading the rule as a setext underline for the line above. The one below
is cosmetic, because a thematic break is a leaf block. The rule now says
that, with the reason, instead of asserting a stricter form the content
does not keep.

The zh-CN reference had received the mechanical box-to-heading swap but
none of the prose behind it: it still claimed a 62-character checkpoint
width and still listed --- among forbidden mixed banner styles, so it
contradicted the convention it was translating. It now carries the
separator section, the setext reasoning, the unconditional-vs-per-runtime
rationale and a corrected anti-pattern list, in Chinese.

The user guide asserted that a heading is not a degradation anywhere.
That is an assertion, not a demonstration. It now says what was actually
traded away in a plain terminal, points at the recorded rationale, and
invites the report that would justify the capability flag instead.

Refs #3028

* chore(#3028): backfill changeset PR number

Refs #3028

---------

Co-authored-by: sim <sim@local>
2026-08-23 22:38:12 -04:00
Tom Boucher
693f12ad56 refactor(#3187): give state field extraction one canonical owner (#3283)
* refactor(#3187): give state field extraction one canonical owner

stateFieldValue in state-document.cts becomes the single owner of the #1760
frontmatter-then-body fallback chain. The new whole-repo guard found 14
independent re-derivations where the epic scoped 5, all now routed through it:
cmdStateSnapshot (11), cmdStatePrune (2) and smart-entry fmScalar (1).

state validate was a gate that could not fail. Every warning it could emit sat
behind a phase resolved without the frontmatter tier, so a STATE.md whose phase
lives only in frontmatter skipped the drift scan entirely and returned
valid:true. It also read unstripped content, letting a frontmatter status: key
shadow the body field (#1255 class). Both fixed; output gains a scope field so
could-not-look stops being output-identical to looked-and-clean.

Verified on the remote runner.

* docs(#3187): document the state validate scope field and its reason codes

Adds docs/how-to/interpret-state-validate-results.md so a reader can tell
nothing-to-report from could-not-look, updates the COMMANDS.md and USER-GUIDE.md
entries, corrects the CONTEXT.md glossary overstatement about Current Position
sole ownership, and drops the changeset fragment.

* fix(#3187): close three drift-guard evasion shapes and test the refuse path

The isolated adversarial review found the ladder detector was evadable by
ordinary reformatting, not just deliberately: a member or computed operand
(fm.key / fm[key]) missed the bare-identifier backreference, a swapped tier
order missed a hardcoded number-then-boolean sequence, and a ladder wrapped
across lines missed single-line detection. All three now caught, each with its
own test plus a proven boundary control.

The frontmatter-parse refuse path on the destructive complete-phase route was
unreachable and therefore untested. It is now driven by an injected parse
failure and asserts STATE.md is byte-identical after the refusal, rather than
shipping untested defensive code on a path that rewrites user state.

Verified on the remote runner.

* fix(#3187): widen the drift guard to the prompt layer and disclose tier-2 changes

The code-review spec axis found the guard's scan surface was src/ only, which is
Decision 4(d)'s forbidden allowlist one directory wide - and it had a live miss:
gsd-core/workflows/smart-entry.md tells an agent to read status from frontmatter
or the body, a prose expression of this same chain. The surface now covers the
prompt layer. That one site carries a permanent written exemption rather than a
ratchet: it is the gsd-tools-is-down fallback, so it cannot call the owner by
construction, and a ratchet would imply removable debt that does not exist.

Two tier-2 output changes shipped undisclosed and are now named in the changeset
and docs: complete-phase's idempotency guard consulting frontmatter, and the
workstream inventory resolving frontmatter-only fields. docs/COMMANDS.md gains a
state complete-phase entry, which it never had.

Also records Amendment 5 on ADR-3180, extracts the duplicated frontmatter-parse
block the epic's own thesis forbids, and re-points two assertions from free-form
warning prose onto the structured drift object.

Verified on the remote runner.

* chore(#3187): backfill changeset PR number

pr:0 placeholder replaced with the real PR number now that #3283 exists.

---------

Co-authored-by: sim <sim@local>
2026-08-09 22:49:39 -04:00
Tom Boucher
b7431a9259 feat(#1956): flag cross-artifact fact drift in the plan drift guard (#3259)
* test(#1956): failing-first contract for cross-artifact fact-drift pass

* feat(#1956): flag cross-artifact fact drift in the plan drift guard

* fix(#1956): correct config-key assertion and bidirectional lifecycle-lag exemption

* docs(#1956): document the cross-artifact axis in the architecture reference

* feat(#1956): decide the phase-status drift axis deterministically

* fix(#1956): scope the progress-table lookup, abstain without a position section, rank deferred

* docs(#1956): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-09 14:23:35 -04:00
Tom Boucher
636ec92107 refactor(#3185): phase enumeration has one owner and a decidable scope (#3222)
* test(#3185): failing-first phase-enumeration single-owner suite

Covers the enumeration rows with direct code evidence: 999.* backlog dirs
listed by progress/stats, the phase-0 sentinel divergence, the #1324
letter-prefixed-decimal negative space, and the destructive-path find —
cmdPhasesClear carries a fifth sentinel copy (/^999(?:\.|$)/) that excludes
999 but not 0, so a 0-* directory roadmap.analyze preserves is deleted there.

Also covers the pass-all degrade, which is where the defect actually lives:
when the milestone window declares no phases the filter becomes a literal
() => true and its heading-side sentinel exclusion is unreachable. A fixture
carrying phase headings keeps the filter active and never reaches that path.

Named for the derivation, not a module: the suite drives commands, phase,
milestone, workstream-inventory and state, and both the phase and
phase-locator buckets are already at the per-module test-file cap.

Committed alone so the remote runner records the failure before the fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): phase enumeration has one owner and a decidable scope

Adds phase-locator.cts::listMilestonePhaseDirs as the single canonical owner
of "which phase directories belong to the current milestone". It applies the
milestone window AND the sentinel filter and returns a ScopedResult, so a
caller can tell a genuinely-empty milestone from an enumeration that could
not be scoped.

The sentinel test now runs against DIRECTORY NAMES and is unconditional.
getMilestonePhaseFilter excludes sentinels from its ROADMAP heading set, but
degrades to a literal () => true pass-all predicate when that set is empty --
at which point the heading set is never consulted and its sentinel exclusion
is unreachable exactly when it is needed. That degrade is the #3167 path, and
it is why stats already used the filter and still listed backlog directories.
The narrowing is sentinel-only: pass-all stays over-inclusive otherwise.

Sentinel copies deleted, canonical isSentinelPhaseId adopted:
  - cmdRoadmapAnalyze's local closure (parseInt === 0 || === 999), 2 call sites
  - cmdPhasesClear's /^999(?:\.|$)/ -- the DESTRUCTIVE path, which excluded
    999 but not 0, so a 0-* directory roadmap.analyze preserves was deleted

cmdStats also seeded rows from ROADMAP headings with no sentinel filter, so a
999 heading produced a row with no directory; that seed is filtered now.

cmdPhasesList routes only its ENUMERATION. --phase lookup searches the
physical set (scoping it would report an out-of-window phase as not found) and
--include-archived still merges archived dirs (they are by definition from
other milestones). Both exempt by documented reason, never a file allowlist.

Fixed inline, found while building: isDirInMilestone could not match a #1324
letter-prefixed-decimal directory (P0.0-foundation) to its own Phase P0.0
heading, so stats reported the phase with plans: 0 while its directory held
plan files. Defers to phase-id's extractPhaseToken rather than widening a
fourth bespoke regex; additive, so it can only admit directories.

Refs #3180. Closes #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): route the last two enumeration re-derivations

workstream-inventory countRoadmapPhases counted every `Phase` heading across
the whole ROADMAP -- no window, no sentinel filter -- so it counted 999.*
backlog and Phase 0 and spanned every milestone the document ever had. Its own
caller already resolved a currentVersion and passed it to getMilestonePhaseFilter
elsewhere in the same file; this was the sibling copy that never got the fix.

state.cts phaseInventoryProvider enumerated phase dirs with its own
/^(\d+)-(.+)$/ convention regex and neither filter, so a rebuilt STATE.md
inventory carried backlog and sentinel directories as current-milestone phases.
A non-COMPLETE enumeration scope now throws to the outer catch as a real scan
failure rather than reporting a confident undercount, mirroring the per-phase
scanPhasePlans contract beside it.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): consolidate 23 sentinel re-derivations onto one predicate

The whole-repo drift guard (ADR-3180 Decision 4a, no file allowlist) found the
sentinel rule re-implemented 23 times across 8 modules, in three regex variants
plus four integer-comparison forms. Most tested 999 only, so Phase 0 slipped
through them while roadmap.analyze and the engine-wide convention (#1580) both
treat 0 and 999 alike. That disagreement is the defect class this epic removes.

All 23 now call phase-id's isSentinelPhaseId (SENTINEL_RANGES [0,999]). Sites:
init recommended-actions and backlog counts, milestone phase scan, the
phase-lifecycle progress table, phase.cts used-number collection and the four
renumber-on-remove guards, roadmap-parser's heading and bullet milestone
counts, roadmap get-phase fallbacks, and state's heading denominator.

Excluding Phase 0 at these sites is a deliberate behavior change and the point
of the consolidation — several carried comments already saying 0 should be
excluded while the literal beside them caught only 999.

Adds scripts/lint-phase-enumeration-drift.cjs, wired into lint:ci. It scans the
whole src/ tree with no file allowlist and reports both shapes: an independent
phases-dir enumeration, and an independent sentinel literal. Exemptions are
function-scoped with a written reason. The guard is comment-aware — its first
pass flagged JSDoc and a comment documenting that the code below uses the
canonical owner, which would have trained readers to exempt prose.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* refactor(#3185): resolve every phases-dir enumeration; drift guard reports zero

Per-site triage of the 31 remaining whole-repo guard hits, applying the rule
generalized from #3183's Amendment 1: a LOOKUP, DIAGNOSTIC, ARCHIVAL or
MUTATION pass wants the physical set; only "which phases belong to this
milestone" wants the scoped set.

Routed (10): init new-milestone phase_dir_count, init milestone-op fallback
count, init manager, init progress, milestone complete stats/dry-run/archive
move, phase complete's next-phase scan, state update-progress, state
frontmatter stats, and uat audit's active set.

Exempt with a written function-scoped reason (never a file allowlist): the
audit/UAT/verification sweeps that deliberately scan every directory to report
gaps, phase create/insert/rename/renumber mutations, single-phase lookups,
roadmap-upgrade's cross-milestone migration, cmdPhasesClear's whole-tree
destructive pass, and the reads that list a phase dir's FILES rather than
enumerating the phases dir at all.

Latent defects fixed by the routing: sentinel directories leaked into
cmdInitNewMilestone's phase_dir_count, cmdMilestoneComplete's stats, dry-run
AND ARCHIVE MOVE, cmdStateUpdateProgress, buildStateFrontmatter and
cmdAuditUat's active set — every one of those hand-rolled an isDirInMilestone
filter with no sentinel exclusion, so `milestone complete` was archiving
backlog directories.

scripts/lint-phase-enumeration-drift.cjs now reports 0 re-derivations and
npm run lint:ci is green.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* docs(#3185): document milestone-scoped enumeration and record ADR Amendment 3

Changeset fragment (Changed), CLI-TOOLS/COMMANDS/USER-GUIDE updates for the
scoped output of progress, stats, phases list, phases clear and milestone
complete, the CONTEXT.md Phase Locator glossary entry naming
listMilestonePhaseDirs, and ADR-3180 Amendment 3.

Amendment 3 records: the SCOPE contract held unchanged; the declared deviation
from Decision 1's provisional signature (the window needs cwd/ws, which the
locked roadmapContent parameter cannot supply); the copy count being a lower
bound for the third consecutive phase (4 scoped vs 54 found); the load-bearing
finding that the sentinel exclusion sat on the heading set and was unreachable
under the pass-all degrade; the two destructive-path defects; and the
generalized exemption rule.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* fix(#3185): wire scope to consumers; revert two wrong routings the suite caught

Review + remote runner findings, all fixed:

The three consumers computed the enumeration scope and threw it away, so
TRUNCATED/UNSCOPED/UNREADABLE collapsed into the same output as COMPLETE --
reproducing this epic's own output-identical-failure defect one layer up.
progress, stats and phases list now emit phase_scope (null on the phases list
--phase lookup path, which performs no enumeration).

Two routings were wrong and the suite proved it:

roadmap-parser's two milestone phase-count scans are reverted to the 999-only
literal. isSentinelPhaseId is BROADER than what it replaced: its legacy branch
runs /^0*(\d+)/ over "00.1", which backtracks to capture 0, so it read #2554's
decimal phase ids as sentinel milestone 0 and stopped counting them.

state.cts phaseInventoryProvider is reverted to the physical disk scan.
`state rebuild` is a RECONCILIATION pass -- scoping it made it throw on healthy
trees whose fixture resolves no window, swallowed the raw readdirSync fault
message #3057 B1 requires verbatim, and stopped it dropping orphan STATE.md
rows, which is the job.

Both are now function-scoped guard exemptions with written reasons, not
silent reverts. This is the consolidation trap named in the epic: a canonical
rule can cover MORE than the copy it replaces, and only real inputs show it.

Adds phases list coverage, a scope-branch test, and a drift-guard unit suite;
backports comment-awareness to the milestone-window and plan-count guards so
all three siblings share one false-positive profile; names #3161 alongside
#3167 in Amendment 3's subsumption record.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* fix(#3185): correct isSentinelPhaseId's decimal-zero misclassification

An isolated security review caught this branch committing the epic's own sin:
the over-broad predicate was worked around at ONE call site and left live at
the destructive ones.

isSentinelPhaseId's legacy branch ran /^0*(\d+)/, which backtracks so any id
whose leading digit run is all zeros before a non-digit captures 0 -- "0.1",
"00.1" and "0.2554" all read as sentinel milestone 0. Two pinned contracts
disagree with that: #2554 requires "00.1" to be counted as a real phase, and
the 999 icebox is a whole reserved milestone so "999.1" must stay sentinel.

The rule is asymmetric and now says so explicitly: 999 is sentinel with or
without a decimal part; 0 is sentinel only when bare. A decimal phase under
either is a real phase for 0 and reserved for 999, because 999 reserves a
MILESTONE while 0 reserves a PHASE.

Fixing the owner lets the earlier workaround go: getMilestonePhaseFilter's two
scans route through isSentinelPhaseId again and the guard exemption that
existed only to accommodate the defect is deleted. The state.cts cmdStateRebuild
exemption stays -- that one is a genuine reconciliation-wants-the-physical-set
case.

Also corrects tests/adr-612-bracket-grammar.test.cjs, which asserted
isSentinelPhaseId('0.1') === true and so had encoded the defect as expected
behavior.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* fix(#3185): keep isSentinelPhaseId's semantics — 0.x is layered, not wrong

Reverts the previous commit. The remote suite failed six tests proving it
wrong, and the reason is the sharpest finding of this phase.

An isolated security review observed that isSentinelPhaseId reads 0.1 and 00.1
as sentinel milestone 0 and judged that a defect against #2554. Correcting the
canonical predicate broke #2949. Both contracts are pinned and both are right,
because they ask different questions:

  #2554  is this dir part of the current milestone's phase SET?  -> count 00.1
  #2949  must this phase COMPLETE before the milestone closes?   -> 0.x sentinel

No single global predicate answers both. isSentinelPhaseId keeps its semantics
(0.x IS a sentinel, #2949), and the milestone-window layer keeps a narrower
999-only rule (#2554) as a function-scoped guard exemption with a written
reason — not a second silent copy.

That corrects how Decision 1 reads: "one owner per derivation" governs who
computes an answer, not how many questions share it. An over-broad canonical
rule is as much a defect as a divergent copy and fails worse, because it looks
like consolidation. Recorded in Amendment 3 as the lesson for Phases 4 and 5.

Where a review's inference about intent conflicts with a pinned contract, the
pinned contract wins; the finding is adjudicated, not fixed.

The boundary tables in the enumeration suite are corrected to assert 0.x IS a
sentinel, with the layering explained.

Refs #3180 #3185.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

* chore(#3185): set changeset fragment pr to 3222

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 14:22:10 -04:00
Tom Boucher
610ebdebe8 docs(#3043): add caution blocks for --dangerously-skip-permissions (#3121)
* docs(#3043): add caution blocks for --dangerously-skip-permissions

The flag was presented without a caveat in docs/USER-GUIDE.md,
docs/tutorials/onboarding-an-existing-codebase.md, and all four translated
locales. Only the English first-project tutorial carried a proper [!CAUTION]
block. All 10 uncaveated occurrences now carry the same caution block
(optional flag, throwaway/low-stakes use, how to keep confirmations, link
to security model).

* chore(#3043): backfill changeset PR number 3121

---------

Co-authored-by: sim <sim@local>
2026-08-06 10:56:53 -04:00
Tom Boucher
ffd5370464 fix(#2903): use the command form that actually works in reader-facing docs (#3047)
* fix(#2903): use the command form that actually works in reader-facing docs

Docs told readers to type the colon form, which no runtime registers -- 18 of
19 runtimes use slash-hyphen and the 19th uses shell-var -- so anyone copying an
example got an unrecognized command. Swept 178 occurrences across 53 files,
locale mirrors included so they do not re-diverge from English.

The colon form is a source-authoring token, not a user-facing one: install-time
converters key on it to produce the hyphen form runtimes actually register. So
the sweep is scoped, and three things are deliberately left alone:

- ADRs, which are a historical record; editing their prose falsifies what was
  written at the time.
- The legacy release-notes archive, pending a maintainer decision on whether it
  follows the same historical carve-out. Excluding it keeps a later reversal
  additive rather than a revert.
- Source artifacts under commands, workflows and agents, where the colon form is
  load-bearing. Rewriting those would break the installed-skill guarantee across
  every runtime -- the single largest hazard here.

The plugin namespace form is a real, separate token and survives untouched.

Adds a lint enforcing exactly that boundary, since the correct form genuinely
differs by directory and nothing previously caught the drift.

Also fixes a hardcoded colon form in the capability-matrix generator. The sweep
alone would have left the generated matrix disagreeing with the template that
produces it, so the fix is at the source and the output regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2903): stop the sweep misquoting source frontmatter

Adversarial review caught three lines where the sweep rewrote a citation of the
literal YAML name: key from a source command file. That key genuinely is the
colon form -- this change's own carve-out logic says source-authoring tokens keep
it -- so the docs ended up misquoting the real files. One of the three is an
acceptance-checklist assertion, which the sweep turned into a false statement.

Restored the three citations to match their sources verbatim, surgically: where a
line carried both a name: citation and a real reader-facing slash command, only
the citation reverted and the command stayed corrected.

The guard needed the same distinction, or it would have flagged the restoration
and reddened the build: a gsd:<cmd> token preceded by name: is a citation of a
source token and is now permitted. The exemption is deliberately narrow -- a bare
gsd:<cmd> anywhere else still fails -- with a test pinning that narrowness.

Also makes the detection case-insensitive. Review found /GSD:next slipped through
silently; no such casing exists in the tree today, so this closes a latent gap
rather than fixing a live one.

Swept the whole tree for further corrupted citations: none beyond the three.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2903): retire the stale-next invariant and sweep next like every other command

Maintainer decision on a genuine conflict between two contracts.

Invariant #3054 banned the literal /gsd-next from user-facing docs because it
named a retired workflow-advance command. But commands/gsd/next.md is a live
command -- the state-aware smart-entry launcher -- and this issue requires docs
to use the hyphen form every runtime actually registers. Both could not hold for
this one command, so docs had been sidestepping the ban by keeping the colon
form, which is exactly the defect this issue exists to remove.

FEATURES.md already recorded the reassignment: the hyphen form "is not the
retired workflow-advance command; it is reserved for the state-aware smart-entry
launcher. Workflow advancement remains under /gsd-progress --next." With that
reassignment the invariant's premise is obsolete and the guard now contradicts
the documented command form, so it is retired with a comment recording why
rather than deleted silently.

next is now swept like every other command, and the earlier exemption added to
the new guard is removed so nothing is special-cased.

Four citations of the literal name: frontmatter key stay in colon form, because
the source file really does carry name: gsd:next and a doc quoting it must
reproduce it verbatim. Two of those lines were reworded to say which side is the
frontmatter key and which is the slash command, since they previously conflated
the two.

Verified the retired scan would now genuinely fail against this tree -- the
conflict was real and resolved, not dodged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2903): backfill changeset pr number

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 13:23:44 -04:00
Tom Boucher
de78f2eef2 docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate (#3010)
* docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate

security-model.md, USER-GUIDE.md, ARCHITECTURE.md, COMMANDS.md,
FEATURES.md, and gsd-planner.md's STRIDE template (+ ja-JP mirrors)
described the pre-ADR-0656 design: slopcheck as the install-or-degrade
gate, with unavailability degrading every package to [ASSUMED].
ADR-0656 inverted this months ago — registry-API verdicts (npm/PyPI/
crates.io) are the gate; slopcheck is an optional escalate-only adapter
that no shipped configuration wires. Verified every replacement claim
against src/package-legitimacy.cts (checkPackages, classifyPackage,
lookupNpm/lookupPypi/lookupCrates) via Memtrace before writing it, so
the corrected prose matches the live implementation rather than
restating the ADR from memory.

Restored docs/explanation/security-model.md:79-84 (and its ja-JP
mirror) to original wording after an orthogonal spec review caught
that an earlier draft had edited the "Why WebSearch packages are
always [ASSUMED]" paragraph — inside the range issue #2775 explicitly
named as correct and to leave alone.

The ja-JP mirror was missing the closing clause present in the
corrected English original ("its absence leaves registry-API verdicts
intact rather than downgrading everything to [ASSUMED]") — added for
parity. This completes the ja-JP mirror the issue's acceptance
criteria named explicitly.

zh-CN/ko-KR/pt-BR (not named by #2775, but carrying the same stale
design) get the mechanical portion of the same fix: command-string
swaps, table headers, ARCHITECTURE.md diagram labels, and technical-
term swaps that reuse a word already attested elsewhere in the same
file (合法性/적법성/legitimidade for "legitimacy") — surrounding prose
untouched. The remainder in those three locales — full-paragraph
rewrites of the corrected degrade-path mechanism, deleted "External
dependency" bullets, and "manually install slopcheck" code blocks —
needs prose composed by a fluent speaker of each language and is filed
as open-gsd/gsd-core#3002 with an exact file:line inventory.

* test(#2775): acknowledge gsd-planner.md byte growth from the STRIDE-row fix

agents/gsd-planner.md grew 14 bytes (49309 -> 49323) from the STRIDE
supply-chain row correction (slopcheck -> package-legitimacy gate).
Emitted agent/workflow files are byte-tracked; this fragment
acknowledges the growth per tests/emitted-attribution.test.cjs's
"differential attribution over the real tree" check.

* docs(#2775): close ja-JP FEATURES.md gap; fix a ko-KR transliterated heading

docs/ja-JP/FEATURES.md:2808 still read the katakana transliteration
"スロップチェック verdict" in REQ-PKG-GATE-01 — invisible to a literal
"slopcheck" grep, so it was missed when ja-JP parity was checked and
declared complete. Corrected to "正当性判定" (legitimacy verdict),
matching the term already established in ja-JP/explanation/
security-model.md and ja-JP/USER-GUIDE.md. This was the only
remaining ja-JP gap; a full sweep for the transliterated form across
docs/ja-JP/ now returns zero hits, and the ja-JP mirror is genuinely
at parity.

docs/ko-KR/USER-GUIDE.md:398's heading "슬롭체크 판정:" had the same
transliteration problem. Fixed inline to "적법성 판정:", reusing the
적법성/legitimacy word already attested two lines below in the same
table. A parallel sweep of zh-CN and pt-BR found no transliterated
forms of "slopcheck" in either locale. The remaining transliterated
occurrence in ko-KR (USER-GUIDE.md:406, the lead-in to the
pip-install code block) needs prose composition like the rest of that
block and is added to open-gsd/gsd-core#3002's inventory.

* chore(#2775): backfill changeset PR number to 3010

---------

Co-authored-by: sim <sim@local>
2026-08-02 20:24:42 -04:00
0xdhx
c61dd49d95 enhance(#2255): blocking catastrophic-shrink guard for curated .planning/ writes (#2301)
* feat(#2255): blocking catastrophic-shrink guard for .planning writes

Adds hooks/gsd-write-guard.js, a PreToolUse hook that hard-blocks
(decision: 'block', exit 2) a whole-file Write collapsing a curated
.planning/ artifact (ROADMAP.md, .planning/milestones/*-ROADMAP.md,
STATE.md) below 40% of its on-disk line count. Files under 40 lines
are exempt; GSD_ALLOW_PLANNING_SHRINK=1 (named in the block message)
bypasses for legitimate milestone resets.

Fix 3 of #973 — the only defense independent of per-agent tool config.
Registered on the Claude plugin surface (hooks.json), settings-json
runtimes (runtime-hooks-surface.cts, self-contained pattern), Kimi
spec, and the OpenCode/Kilo plugin buses. Golden install fixtures and
INVENTORY regenerated; regression tests negative-controlled (16/16
RED with the hook absent, 16/16 GREEN with it present).

* chore(#2255): backfill changeset pr number to 2301

* enhance(#2255): address review — fail-closed reads, typed block output, registration, property test

Review fixes for trek-e's CHANGES_REQUESTED on PR #2301:

- Blocker 2: register gsd-write-guard.js in BUNDLED_GSD_HOOK_FILES
  (no-shipping-drift test).
- Blocker 3: update the always-on hook enumerations in ADR-766 and
  CONTEXT.md from six to seven.
- Major 4: fail CLOSED on non-ENOENT read errors — only a missing file
  (new-file Write) passes; EACCES/EISDIR/ELOOP/etc now block, with a
  typed readError field and the override still honored. Tested, with a
  negative control against the pre-fix hook.
- Major 5: fast-check property test for the SHRINK_RATIO/FLOOR_LINES
  budget contract (blocked ⟺ newLines < oldLines*SHRINK_RATIO above the
  floor; sub-floor always exempt), boundary examples pinned.
- Major 6: block output now carries typed oldLines/newLines/
  overrideEnvVar fields; tests assert on those instead of regexing the
  free-form reason string.
- Minor: CURATED_PATTERNS are case-insensitive (case-insensitive-FS
  bypass on macOS/Windows); limit+1 boundary tests added for both the
  floor and the ratio.

* enhance(#2255): engage the write guard on Kimi's native payload shape

The guard shipped with Claude-vocabulary checks (tool_name 'Write',
tool_input.file_path), which #2304 showed leaves a guard dormant on
Kimi: the [[hooks]] matcher is registered pre-translated but kimi-cli
forwards its native payload verbatim — tool_name 'WriteFile' (bare or
module-qualified) and tool_input.path per its tool schemas
(src/kimi_cli/tools/file/write.py). The guard matched, saw an unknown
name, and exited 0.

Apply the same per-guard normalization PR #2326 gives the three
sibling guards (name + field mapping, inlined — hook scripts stage as
standalone files), and write the block reason to stderr as well as
stdout JSON: Kimi feeds stderr, not stdout, back to the model on
exit 2, so a stdout-only reason blocks without telling the model why
or naming the documented override.

Regression tests pipe Kimi-shaped payloads (engage, qualified-name,
stderr-reason) plus exemption pins (StrReplaceFile stays out of scope
by design; non-curated paths pass) — verified red against the pre-fix
guard, green after.

* enhance(#2255): rebase onto next; regenerate golden-parity fixtures

* enhance(#2255): wire the escape hatch into complete-milestone's reorganize step

Review Blocker 1: the guard hard-blocked /gsd:complete-milestone's ROADMAP
reorganize — the tree's only legitimate milestone reset and the exact caller
GSD_ALLOW_PLANNING_SHRINK was built for. The reorganize step now performs the
rewrite through a shell write with the hatch set on the command (a hook
inherits the runtime env, so a bare Write cannot carry a per-step override),
and a binding test derives the env var name from the guard's typed output and
asserts (a) the workflow step sets it and (b) the guard passes the identical
catastrophic payload under it — so the next complete-milestone.md edit cannot
silently re-break the wiring.

* enhance(#2255): drop dead Edit-class mapping from normalizeKimiPayload

Review Major 1: StrReplaceFile -> 'Edit' and the old_string/new_string
reconstruction were unreachable-by-effect — the guard exits 0 for any
tool_name !== 'Write', so nothing ever read the fields they set, leaving
guaranteed-surviving mutants against the Stryker bar. The map now carries
only WriteFile -> 'Write'; the StrReplaceFile exemption test message states
the fall-through it actually exercises.

* enhance(#2255): review minors — American spellings; writeSync before exit(2)

Minor 1: normalised/normalise -> American house style. Minor 2: the two
block paths wrote stdout+stderr via async pipe writes then exit(2) —
async-on-Windows, unflushed at exit; fs.writeSync(1/2, ...) makes the block
payload durable.

* enhance(#2255): assert stderr equals the typed reason, not raw prose

Minor 3: the last raw-text match in the suite pinned override-name prose on
stderr. The contract is "stderr carries the reason Kimi feeds back" — now
asserted as stderr non-empty and byte-equal to the parsed stdout.reason.

* enhance(#2255): bind the write-guard's Kimi normalization into the parity test

Review Major 2: the guard's normalizeKimiPayload is a 4th inlined copy with
nothing binding it. This extends PR #2326's kimi-guard-normalization-parity
test (same path and helpers, authored as a superset so either merge order
resolves cleanly): sibling byte-parity is existence-gated zero-or-all —
trivially green until #2326 lands, full-strength after — and the write-guard
copy is bound semantically (map is the value-inverse of convertKimiToolName;
the Kimi name for Write must map, or the guard is dormant on Kimi; the
path -> file_path half must be present). Byte-parity is deliberately not
asserted for this copy: it legitimately omits the Edit-class mapping
(Major 1 — dead code in a Write-only guard).

* enhance(#2255): refresh golden-parity fixtures for revised guard + workflow

* chore(#2255): regenerate golden fixtures after rebase onto next

The committed fixture hashes were generated against a tree predating
next's latest 11 commits, which independently modified the same
install-parity surface. Rebased onto next and regenerated with
`npm run gen:golden`.

Verified: against upstream/next the regenerated fixtures differ by
exactly this PR's own entries -- hooks/gsd-write-guard.js (new),
hooks/managed-hooks-registry.cjs, plugins/gsd-core.js, and
gsd-core/workflows/complete-milestone.md. No unrelated drift.

* fix(#2255): regenerate workflow size baseline for complete-milestone

`complete-milestone.md` grew 31071 -> 32061 (+990) when the round-2
review fix bound GSD_ALLOW_PLANNING_SHRINK=1 into the reorganize step,
but tests/workflow-size-baseline.json was never regenerated. The
per-file workflow baseline test (issue #1074) failed on
ubuntu-latest/22 and both macOS shard 1/3 jobs.

The growth is justified: it is the escape-hatch binding requested in
review round 2 (the guard must not hard-block the tree's only
legitimate milestone reset), not incidental bloat.

Regenerated via `npm run size:baseline`; the diff is exactly the one
entry.

* chore(#2255): regenerate golden fixtures and size baseline after rebase onto next

* enhance(#2255): bind the shrink escape hatch mechanically — single-use sentinel the guard consumes

Round-5 M1: the per-step `GSD_ALLOW_PLANNING_SHRINK=1 tee` prefix was inert
(no PreToolUse hook exists on Bash in this family; the write succeeded by
dodging the guard, not by the override firing) and the protection was prose.
The hatch is now a transport code consults: complete-milestone's reorganize
step arms `.planning/.gsd-allow-shrink` with the target's path, keeps the
Write tool as the sanctioned path, and the guard — at the block point only —
verifies the sentinel is fresh (15 min) and names the pending target, then
CONSUMES it and allows that one write. Path-bound + single-use + freshness
keep it from becoming a standing unlock. The env var remains as the
interactive transport, where it can actually reach the hook.

Regression tests written first (negative control: 3 failed pre-fix): the
armed-sentinel Write passes and consumes; stale does not exempt; a token for
a different file neither exempts nor is consumed; the binding test now takes
the sentinel name from the guard's typed output (overrideSentinel), asserts
the step arms it, and asserts the step no longer routes the rewrite around
Write via a shell pipe.

Also in this commit, same file:
- m2: block emission is exception-safe — emitBlock() wraps both writeSync
  sites in their own try/catch that still exits 2, so an EPIPE can no longer
  convert fail-closed into the outer catch's fail-open.
- Header discloses the two reviewed design limits (cumulative sequential
  shrink; lexical match vs symlinked paths) per round-5 scoping.

* docs(#2255): document the sentinel transport across guard surfaces; changeset ends with the (#2255) parenthetical (m4)

USER-GUIDE bullet, INVENTORY row (en + ja/ko/pt/zh), the
runtime-hooks-surface registration comment, and the changeset now describe
both hatches — the single-use sentinel for workflow steps and the env var
for interactive use — instead of implying a per-step env can reach a hook.
The changeset's trailing `Resolves #2255.` prose becomes the `(#2255)`
parenthetical the repo's fragments use (round-5 m4).

* chore(#2255): regenerate derived families on the rebased tree (full sweep)

Full generator sweep after rebasing onto next @ the body-parser-patched
lockfile: build, gen-inventory-manifest, gen:golden, size:baseline. Every
regen delta verified to be either a PR-owned entry (gsd-write-guard.js,
complete-milestone.md, INVENTORY/USER-GUIDE) or exact convergence to next's
committed value for entries our arbitrary-side conflict resolution had left
stale (all 18 runtime fixtures checked mechanically).

* test(#2255): use helpers.cleanup for sentinel teardown, not raw fs.rmSync

The repo's local/no-raw-rmsync-in-tests rule exists for the Windows-EBUSY
retry budget; the sentinel disarm now rides it like every other teardown.

* chore(#2255): regenerate derived families after rebase onto next

Full sweep on the rebased tree (build -> gen-inventory-manifest ->
gen:golden -> size:baseline). Every delta is either a PR-owned entry
(hooks/gsd-write-guard.js, its registration surfaces
hooks/managed-hooks-registry.cjs and the two plugin buses,
gsd-core/workflows/complete-milestone.md) or exact convergence to
next's committed value across all 18 runtime fixtures.

* chore(#2255): regenerate derived families after rebase onto next @ a5180d96

Rebase onto current `next` (a5180d96) resolved 12 conflicting
golden-install-parity fixtures; all regenerated via the full generator
sweep (build, gen:golden, size:baseline) rather than a single generator.

`lint:generated-sync` reports every generated artifact in sync. All 45
differing fixture keys and the single workflow-size-baseline entry map
to files this PR actually touches; no foreign drift.

* fix(#2255): remove the stale unguarded reorganize_roadmap step (round-8 blocker)

complete-milestone.md carried a second ROADMAP-collapsing step,
`reorganize_roadmap`, distinct from the sentinel-armed
`reorganize_roadmap_and_delete_originals` this PR wired. It is a vestige
of the pre-archive-then-reorganize design: it sits BEFORE
archive_milestone, so executing it as written would collapse ROADMAP.md
before the archive snapshots the full phase detail — and its Write is
exactly the shape gsd-write-guard hard-blocks, with no hatch armed. The
file's own success criteria describe only one reorganize outcome
(Backlog-preserving, overwrite-in-place — the later step's properties),
and archive_milestone points forward to "the reorganize step".

Removed rather than wired, per the round-8 review's confirm-and-remove
option. A new binding test asserts the sentinel-armed step is the ONLY
reorganize step in the workflow, so an unguarded collapse step cannot be
silently reintroduced (negative-controlled: fails against the pre-fix
tree). Golden-parity fixtures and the size baseline regenerate for the
shrunk file; every changed fixture key is complete-milestone.md's own.

* test(#2255): document why the read-error injection is a path collision, not an fs monkeypatch

Round-8 nit: the non-ENOENT tests inject via a directory-at-target-path
collision instead of the repo's fs-method monkeypatch pattern. That is
deliberate, not drift — runHook exercises the hook as a spawnSync child
process, so an in-process fs.readFileSync patch (the pattern the cited
siblings use on require'd, in-process code) can never reach the code
under test. Record the reasoning at the injection site.

* chore(#2255): regenerate derived families after rebase onto next @ 0d08c320

Rebase onto current next (0d08c320) for the CONFLICTING/DIRTY state. All 32
conflicts were generated artifacts (19 golden-install-parity, 12 install-tree,
workflow-size-baseline); resolved arbitrarily and regenerated via a full
generator sweep (build, gen:golden, size:baseline, gen-inventory-manifest)
rather than hand-merged. No source conflicts.

Regen diff verified against the PR's changed-file set: 7 distinct differing
keys, all PR-owned (gsd-write-guard.js, managed-hooks-registry.cjs,
plugins/gsd-core.js, complete-milestone.md, and their .kimi mirrors).
lint:generated-sync clean.

* chore(#2255): regenerate derived families after rebase onto next @ 9138271b

Conflict set was 20 paths, every one a generated artifact, zero source
conflicts — resolved arbitrarily during the replay and regenerated here,
per the maintainer's round-9 recipe (never hand-merged).

Generator sweep (not just gen:golden): npm run build, gen:golden,
size:baseline, gen-inventory-manifest, gen:registry. INVENTORY-MANIFEST
came back byte-identical, so the merged value was already correct.

Regen diff verified == PR-touched entries: every differing leaf key
attributes to a file this PR changes (complete-milestone.md, the guard,
managed-hooks-registry.cjs, both plugin buses) — derived from the
changed-file set, not a feature-name filter.

* fix(#2255): port the #2595 Kimi-payload hardening the base landed for the sibling guards

Self-found via the round-10 base sweep: a8b40fa5 (#2547/PR #2595) hardened
gsd-prompt/read/worktree-path-guard's normalizeKimiPayload after this PR was
authored, and this guard carried all three pre-fix shapes:

- `path` now wins outright over a model-supplied `file_path` — kimi-cli's
  WriteFile schema carries no `file_path`, so a payload pairing a curated
  `path` with a spurious `file_path: ""` had the guard reading '' and
  exiting 0 while kimi-cli wrote to `path` (one-key bypass, no crash needed).
- null/primitive payloads pass through deliberately instead of throwing into
  the fail-open catch.
- the `file_path` read is typed (`typeof === 'string'`), closing the
  crash-to-allow on `[]`/`{}` — this is what the base-added
  kimi-guard-typed-payload-reads scan flags.

Negative-controlled: the shadowing payload exits 0 against the pre-fix guard
and 2 against this one.

* fix(#2255): round 9 Major 2 — stop printing the sentinel recipe in the agent-facing denial

The block message ended by telling the blocked agent exactly how to arm the
single-use sentinel — handing the #973 actor a mechanical self-authorization
in the moment it is blocked. The recipe is gone from the prose; the
corrective advice and the env override's name stay (the latter is a #2255
acceptance criterion, and a per-step env prefix cannot reach a hook anyway),
and the typed overrideSentinel field stays for the binding tests. The hatch
remains documented in USER-GUIDE.md and complete-milestone.md, where humans
and the workflow engine read.

* fix(#2255): round 9 Minors 1-2 — realpath-resolve the target before the curated match; disclose the /i Linux cost

Minor 1: a Write to a non-curated path that symlinks into a curated file was
not matched while writeFileSync followed the link — the target is now
realpath-resolved before the curated match (ENOENT keeps the lexical
resolution so new-file Writes still pass; any other realpath error falls
through to the read, which fails closed). Negative-controlled: the symlink
payload exits 0 against the pre-fix guard, 2 against this one. Test skips on
win32, where symlink creation needs privilege.

Minor 2: the header's design-limits block now names the unconditional /i
cost on case-sensitive Linux (a genuinely distinct .planning/roadmap.md is
also treated as curated) next to the stateless limit, and drops the closed
symlink limit.

* test(#2255): round 9 Minors 3-4 — CRLF counting pin + a passing Write leaves a fresh sentinel unburned

Minor 3: countLines' split('\n') is CRLF-safe for a count (the \r rides
along), confirmed by trace in the review — this pins it against this repo's
recurring CRLF regressions, on both sides of the compare and at the 40%
boundary.

Minor 4: consumeSentinelFor runs only after the ratio check would block, so
a within-tolerance Write never burns the workflow's token — true by
construction, previously un-asserted.

* fix(#2255): round 9 Major 3 — correct the stale env-var line in archive_milestone's summary

complete-milestone.md's "After archival" bullet still said the reorganize
happens "under GSD_ALLOW_PLANNING_SHRINK=1" — the wording from the round-2
design this PR's own history rejected in round 5 (a per-step env var cannot
reach a hook; setting it in a Bash step silently does nothing). It now points
at the sentinel mechanics the reorganize step actually documents, matching
that step and USER-GUIDE.md.

* docs(#2255): round 9 Major 1 — user-facing docs state the stateless per-Write limit

The changeset and USER-GUIDE described the guard as covering "catastrophically
shrinks" with no caveat, while the stateless design was disclosed only in the
hook header — an operator reading the shipped docs would conclude iterative
erosion is covered. Both surfaces now state the per-Write comparison and the
erosion non-goal explicitly, in line with what the guard does.

* fix(#2255): round 10 Blocker 1 — arm the hatch in the roadmapper's curated writes

gsd-roadmapper Step 7 Writes both .planning/ROADMAP.md and .planning/STATE.md
wholesale, and /gsd:new-milestone spawns it against the OUTGOING milestone's
files. new-milestone's `phases.clear` archives phase DIRECTORIES, never
ROADMAP.md, so nothing compacts it first, and no ordering rule forces
/gsd:complete-milestone to run before /gsd:new-milestone.

Measured against the shipped hook at the #973 file size (292 lines): a new
4-phase roadmap lands at 18.2% and an 8-phase one at 31.8% — both blocked; only
a 12-phase replacement (45.5%) clears. So the collapse is reachable on a
first-party flow, which is the failure mode #2255 para 4 says the guard must not
have.

Arms the single-use sentinel immediately before each Write, one per target
(the token is path-bound and single-use, so one arming cannot cover both), and
gates each arming on `[ -f ]` so the /gsd:new-project path — where the guard
already exempts via ENOENT — strands no unconsumed token.

Binding test derives the sentinel name from the guard's own typed output, so a
rename on either side fails the suite instead of silently unwiring the hatch.
Negative-controlled: 2 assertions fail against the pre-fix roadmapper.

* docs(#2255): round 10 Major 2 — scope the claim to the guarantee that ships

The sentinel is a plain file, so an agent that would reason past an advisory
can arm one with a single Bash call it is already permitted to make. The
"only defense independent of per-agent tool config" framing outran that, and
the changeset was on its way into CHANGELOG.md.

Retitles the claim on all three surfaces (changeset, guard header, USER-GUIDE)
to what the guard actually delivers: it blocks accidental and single-shot
collapse and is not a defense against a determined agent; what it converts is
"ignore a sentence" into "take one deliberate, path-bound, single-use,
auditable action".

Pinned by test on the DURABLE surfaces only — the guard header and USER-GUIDE.
The changeset fragment is deliberately not pinned: it is consumed at release,
so a test reading it would start failing the moment the release lands. The
bound-statement assertion normalizes comment markers and whitespace first, so
it pins the claim rather than the paragraph's line wrapping.

Negative-controlled: both assertions fail against the pre-fix surfaces.

* test(#2255): acknowledge the roadmapper growth from the round 10 Blocker 1 wiring

The emitted-attribution gate (#2719/#2767) flags gsd-roadmapper.md growing 1130
bytes without an acknowledgment. The growth is the Blocker 1 sentinel wiring
plus the rationale a future editor needs to keep it, so it gets an ack fragment
rather than a silencing regen — the gate's own message is explicit that there is
nothing left to regenerate.

Fragment is PR-scoped (2301-…) per the gate's naming instruction, and uses the
plain-string reason form the shipped fragments use.

Verified against the TRUE upstream tip, not the fork's origin/next: a stale
origin made this same gate report unrelated phantom drift (1 emitted path + 6
grown files + 5 stale acks) that vanishes when GSD_EMITTED_BASE is pinned.

* test(#2255): renumber the roadmapper PROSE_ALLOWLIST pin after the Step 7 wiring

CI red on shard 2/3, all four platforms. The #2751 gate keys PROSE_ALLOWLIST on
{file, line}; the Blocker 1 wiring added 18 lines above the allowlisted
parenthetical in agents/gsd-roadmapper.md, moving it 624 -> 642. Both halves of
the gate then fired: the moved line reads as a new offender, and the stale
entry no longer matches anything.

Line content at 642 is byte-identical to what the entry describes — a
descriptive "e.g." naming SDK queries a user could run — so this is a
renumber, not a re-classification.

Swept the defect class rather than the instance: agents/gsd-roadmapper.md is
the only line-pinned reference to any file this round changed.

Negative-controlled: both assertions fail against the un-renumbered allowlist.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:19:49 -04:00
Tom Boucher
c87f6f358e enhance(#1854): offer restore for user-added files backed up on update (#2679)
* test(#1854): failing-first coverage for user-files-backup restore

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(#1854): offer restore for user-added files backed up on update

Adds a restore-custom-files gsd-tools verb and wires it into update.md as a
restore_custom_files step: plan, compatibility-check against the newly
installed release, then restore only on explicit opt-in. The backup is never
deleted, a shipped path is never overwritten, and a single unwritable entry
does not abort the rest.

Also drops the jq pipe from update-context field extraction (#2589 class,
missed by that sweep) and repairs a broken code fence in docs/CLI-TOOLS.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): reject symlinked restore destinations and backup roots

Self-review of the restore path found two write-through holes: copyFileSync
follows a symlinked destination, so a link planted at the restore target wrote
outside the config dir with every ancestor still a real directory; and statSync
on the backup root followed a link, letting the walk read arbitrary files and
present them as the user's own backup. Both now lstat.

Also marks the report's path/detail strings as untrusted data in update.md so
the rendered step cannot carry instructions into the runtime model.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1854): move the update-context jq guard into the #2589 sweep

update.md joins the AUDITED list rather than carrying a duplicate assertion in
the backup-restore suite, and the guard gains a negative-proof companion so
'no jq pipe' cannot pass by the fields simply no longer being read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): validate manifest files map shape before trusting it

Security review flagged that Object.keys on a non-plain-object files field
yields numeric-index keys matching nothing, so the managed-path check dies
silently while manifest_found still reports true. Shape, not just type
(ADR-227): an array or scalar files map is now an unusable manifest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): size the restore prompt by eligible_count

Spec review found the prompt was driven by entries.length, so a backup holding
only blocked entries asked "Restore 1 file(s)?" when accepting would restore
zero. The question now reads eligible_count, and an all-blocked backup reports
its reasons instead of offering a choice that cannot be honored. The decline
path names the resolved backup_dir rather than the bare directory name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#1854): use t.skip on hosts without symlink support

A bare return in a node:test body registers as a PASS, so the four symlink
guards silently reported green on unprivileged Windows instead of skipping.
Adds the dangling-link destination case the security review called out, and
moves outside-dir teardown to t.after so a failing assert cannot leak it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#1854): unfence restore hint, regen goldens, widen install timeout

Three gate failures from the c99d612a5 run, all root-caused:

1. capability-registry (3): update.md's decline message put an instructional
   'gsd-tools ...' line in an UNTAGGED fence, and the guard treats untagged
   fences as shell blocks. Retagged both display blocks as text and switched
   the hint to the resolved 'node <config-dir>/.../gsd-tools.cjs' form users
   can actually paste.

2. golden-install-parity (19): update.md and gsd-tools.cjs ship, so every
   runtime fixture moved. Regenerated; the diff is exactly those two hashes
   per fixture, no other drift.

3. install.test.cjs (5): one real failure, four cascades. The Cursor suite's
   before hook died on 'spawnSync ETIMEDOUT' at the 60s cap while the node22
   lane passed the SAME commit in 12.7s. A full install measures 13-30s idle,
   so 60s was under 2x headroom and shrinks with every file added to the
   payload. Raised to 120s, matching the heavy case already in this file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#1854): backfill changeset pr number to 2679

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-07-26 19:50:56 -04:00
Tom Boucher
57b2bd8368 fix(#2491): finish todos/done -> todos/completed rename (14 stale refs) + guard (#2626)
* test(#2491): add todos/done rename under-sweep guard

* fix(#2491): finish todos/done -> todos/completed rename (14 stale refs)

* chore(#2491): backfill changeset pr to 2626

* test(#2491): fix lint-legacy-dir-name + allow-test-rule-refs (split legacy token, add issue ref)
2026-07-24 21:36:33 -04:00
Tom Boucher
517bae8d6d fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message

Bug: check.decision-coverage-plan's remediation message told the user to
cite decisions "(or body)" but extractPlanDesignatedSections only scanned
<objective>/<tasks>/<task>/<action>. A decision cited in <read_first>,
<behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to
the gate — false BLOCKING coverage gap, plus the message's own fix-hint
sent the user to "the body" where re-citing still failed.

Two-part fix (must change together — that drift was the bug):

1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also
   match <read_first>, <behavior>, <verify>, <acceptance_criteria>,
   <done>. These are all planner-canonical tags the planner is told to
   use (plan-phase.md:830-862, plan-phase.md:772). The body negative-
   lookahead mirrors the opening-tag set so each tag's body is captured
   independently.

2. Correct buildPlanMessage to name ONLY the surfaces the extractor
   actually scans (front-matter must_haves/truths/objective,
   designated markdown headings, and the nine planner-canonical tag
   bodies). The misleading "(or body)" clause is gone.

Also updates the planner's documented contract (agents/gsd-planner.md:69)
and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to
reflect the wider scan.

Regression tests in tests/decisions.test.cjs cover each newly-scanned
tag body, a control (no citation still uncovered), and a message/extractor
parity assertion that names every scanned surface — so the two cannot
drift apart again.

Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage
is a separate command (decision-coverage-verify) checking shipped
artifacts, not plan citations — untouched.

* chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures

gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision-
coverage contract (5 new scanned tag names + heading clarification).
Growth is justified: the contract surface is itself the fix — the prior
text under-described what the gate scans, which was the bug.

Updates:
- tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294)
- 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime)
- tests/fixtures/install-tree/*.json (regenerated by gen:golden)

* fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting

Code review (subagent) flagged a Medium edge-case regression from the
single-alternation regex: when a newly-scanned tag nests inside another
scanned tag, the alternation's negative lookahead halts the outer tag's
body at the inner tag — losing any D-NN citation in the outer tag's
prefix prose. Concretely:

  <action>per D-05 <verify>npm test</verify></action>

  → 3-tag alternation (old):  captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught
  → 9-tag alternation (bug):  captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST

Switches extractXmlTagBodies to per-tag matching: each tag gets its own
regex whose negative-lookahead tempers only against the SAME tag's
reopening. So <verify> inside <action> is absorbed into <action>'s body
(D-05 caught) AND <verify> is matched separately on its own pass.

Per-tag preserves both:
- the reporter's case (sibling tags inside <read_first>)
- nested-tag citations in outer-tag prefix prose
- ReDoS safety (each per-tag regex keeps the #2128 body tempering)

Also adds the reviewer's other requested edge-case tests:
- non-scanned tag (<name>) bearing D-NN must NOT count
- self-closing form <read_first /> safely ignored
- attribute form <verify type="...">D-NN</verify> (canonical planner shape)
- CRLF newlines in tag body do not break capture

* chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md
2026-07-19 23:07:13 -04:00
Tom Boucher
4a9833d3e3 fix(#2278): use Edit() not Write() for Claude allow-permissions + migrate legacy (#2302)
GSD_CLAUDE_ALLOW_PERMISSIONS pre-populated Claude Code settings.json
with Write(.planning/*) and Write(STATE.md). Claude Code has no
standalone Write permission gate — file-editing tools are gated
collectively via Edit(pattern) — so those rules never matched, fresh
installs still hit first-run approval prompts for .planning/* and
STATE.md, and Claude Code emitted a session-start warning about the
unmatched rules.

Swap the two entries to Edit(.planning/*) / Edit(STATE.md). Add a
GSD_CLAUDE_LEGACY_ALLOW_PERMISSIONS list of the retired Write(...) forms,
consulted by mergeClaudePermissions (actively remove stale entries when
adding current ones, idempotent, user entries preserved) and by the
uninstall cleanup filter (still removes the legacy form). Sample
settings.json in docs/USER-GUIDE.md corrected to match.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 13:46:26 -04:00
Tom Boucher
60e3c4988a feat(#2182): scaffold community capability + EoS registry (tests + stubbed core)
Adds the discoverability-registry surface for issue #2182: JSON-sourced
capability/eos catalogs, a pure schema/vocab module (registry-schema.cjs)
with the ADR-857 loop points + ADR-1239 axes, thin validate/gen CLIs,
the registry-entry PR template, README spec, and CONTEXT.md glossary terms.

The three pure functions (isValidGsdRange/validateEntries/renderMarkdown)
are stubbed here so the comprehensive test suite fails first (red), per the
feature-implementation red-first directive; the next commit implements them.

Refs #2182

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 14:13:02 -04:00
Jeremy McSpadden
a3cca0704d no-mistakes(document): Sync onboard documentation 2026-07-05 19:16:06 +00:00
Tom Boucher
8f2ebbe9bf feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity

Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap).

--gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): backfill changeset PR number (#1996)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1928): drop Gemini CLI from issue templates (review nit)

Removes the sunset Gemini CLI runtime from the two GitHub issue-template
runtime lists that the removal PR missed, per @davesienkowski's review nit:
- feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise
  request a feature for a runtime GSD no longer supports)
- bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json
  retrieval-help line

Leaves the post-removal templates fully consistent with the Antigravity redirect.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 13:32:51 -04:00
Tom Boucher
e32eac56a6 docs(#1618): correct Claude skill layout from nested to flat 2026-06-23 14:11:22 -04:00
Tom Boucher
08dfcad5f1 docs(#1617): correct Antigravity skill layout from nested to flat 2026-06-23 14:05:49 -04:00
Tom Boucher
ba96c70b14 feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a
deterministic classifier that `verify-work` consumes to route deliverables to
auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic.

- New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage
  block (extractFrontmatter can't — its `-` items are scalars-only; this is a
  focused parser, sibling of parseMustHavesBlock), validates each entry, and
  classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE
  typed-IR surface. Exposed via `uat classify-coverage --summary <f>`.
- Auto-pass is the narrow proven case only: strict-boolean human_judgment:false
  AND non-empty all-`pass` verification AND zero validation errors. Everything
  else — judgment, empty/failing verification, malformed entry — routes to the
  human (fail-safe). A malformed block falls back to legacy prose extraction and
  surfaces an error; an absent block is byte-identical to pre-#1602.
- execute-plan create_summary populates the block (fail-safe default
  human_judgment:true); verify-work extract_tests consumes it; create_uat_file
  marks auto-passed entries `source: automated`.
- Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY,
  eslint/gitignore registration, and Diataxis docs (COMMANDS reference +
  USER-GUIDE explanation) updated.
- Behavioral tests via the CLI (no source-grep); parser-robustness regressions
  for the null-entry/comment-header/mis-indent cases found in adversarial review.

Closes #1602

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 23:43:22 -04:00
Joe
e12a2abfd8 feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit

Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed),
enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way
to browse or audit parked seeds on demand. This adds a read-only listing,
following the established --list → workflow pattern (per the approved scope on

- gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the
  seeds dir, returns { count, seeds[], summary } JSON with each seed's id,
  slug, status, scope, trigger_when, planted, title. Optional case-insensitive
  status filter. User-controlled content is sanitized (sanitizeForDisplay) and
  every path validated (requireSafePath); read-only. Independent of
  audit.scanSeeds, which only returns unimplemented seeds for the milestone surface.
- /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that
  renders the seed table.

Closes #441

* chore(#441): point changeset fragment at PR #722

* test(#441): allowlist list-seeds test in prompt-injection scan

The test asserts that list-seeds neutralizes injection payloads
(<system>, [INST]) embedded in seed content, so the fixtures legitimately
contain those patterns — same as the sibling security tests already on the
allowlist.

* fix(#441): use canonical /gsd:capture colon form in list-seeds workflow

Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use
the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is
retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new
list-seeds workflow used the hyphen form.

* docs(#441): sync help full.md + INVENTORY for --list-seeds

Adds the --list-seeds entry to the help reference (help/modes/full.md, per
bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow
in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json.

* docs(#441): add --list-seeds how-to + drop phantom statuses

Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers):

- USER-GUIDE.md Seeds section (how-to): extend the task to cover
  auditing parked seeds on demand via --list-seeds, including the
  status filter — kept task-oriented per Diataxis how-to mode.
- CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected
  from the list-seeds filter vocabulary; the system only produces
  dormant|active|triggered (src/audit.cts scanSeeds). Reference must
  be factually accurate and complete.

* fix(#441): guard non-scalar status frontmatter in cmdListSeeds

A seed with a bare `status:` line (extractFrontmatter yields {}) or a
`status: [a, b]` value (yields an array) crashed the whole audit list:
`(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string.
Coerce every frontmatter read through a `fmStr` helper (mirrors the existing
`typeof fm.id === 'string'` guard), so a non-scalar status falls back to
dormant and non-scalar scope/trigger_when/title can no longer leak a raw
array/object into the JSON contract. Title is now capped symmetrically.

Adds regression coverage for empty and array `status:` and non-scalar fields.

Refs #441

* docs(#441): align list-seeds workflow status vocabulary

The load_seeds step listed `implemented` as an example status filter, but the
real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds);
`implemented` has no producer. Matches the earlier CLI-TOOLS.md correction.

Refs #441

* refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds

Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported
deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested
in-process (review minor #1). No behavior change.

Filter comparison now matches the raw lowercased status (both sides already
normalized) instead of sanitizeForDisplay(status); sanitization is for output,
not matching (review nit #3).

* test(#441): add fast-check property coverage and count=1 boundary for list-seeds

Adds tests/list-seeds.property.test.cjs with four fast-check properties over
deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug
invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing
(review minor #1).

Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2).

* chore(#441): sync runtime launcher snippet into list-seeds workflow

Propagate the current _runtime-launcher.snippet.sh (with non-Claude
runtime home probes) into the new list-seeds.md workflow via
scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation.

* test(#441): record list-seeds.md in workflow size baseline (#1074)

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-06-22 00:59:41 -04:00
Tom Boucher
f3c06f59df fix(#1326): stop emitting Codex agents/openai.yaml sidecars; clean up stale ones (#1360)
* fix(#1326): stop emitting Codex agents/openai.yaml sidecars; clean up stale ones

Codex installs wrote an agents/openai.yaml sidecar under every managed gsd-*
skill dir. Recent Codex builds index both SKILL.md and the sidecar, so each
GSD skill appeared twice in autocomplete (canonical gsd-* name + humanized
display_name).

- Replace writeCodexSkillMetadataFiles / generateCodexSkillMetadataYaml with
  cleanupCodexSkillMetadataSidecars: Codex-only (if isCodex), removes stale
  managed gsd-*/agents/openai.yaml and prunes the now-empty agents/ dir.
- Preserve user-owned dirs (gsd-dev-preferences), non-empty agents/ dirs, and
  non-gsd dirs; lstat-guard against symlinked agents/ so a delete can never
  escape the skills tree; fail-open per directory.
- Codex relies on SKILL.md alone for /skills discovery.
- Update USER-GUIDE/FEATURES docs and rewrite the #774 emission tests into
  cleanup tests.

Scope: the sidecar duplicate only. The separate multi-root (~/.agents/skills
shared-skills) duplicate facet is a distinct concern, not addressed here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1326): add changeset for Codex sidecar cleanup

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 21:55:49 -04:00
Tom Boucher
14e0709eed refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active (#1315)
* refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active

Part A: intel's command gate moves from config-only isIntelEnabled to the
shared isCapabilityActive('intel', cwd) — a consistency cutover (intel has
skills:[] so its tri-state collapses to the intel.enabled config leg; the gate
now flows through the resolver's precedence + runtime-aware resolution).
Part B: loop-resolver hook rendering now gates on capability state.active
(=== true, fail-closed) instead of state.enabled, so the capability config
gate is honored by the hook consumer, not just per-hook 'when'. active is now
required in the loop-resolver input types. Regression test proves a config-
disabled (active=false) capability's unconditional hook is not rendered.
Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1307): add changeset for intel + loop-resolver active gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 00:12:54 -04:00
Tom Boucher
e3b829e765 refactor(#1306): gate graphify on isCapabilityActive (tri-state, runtime-aware) (#1313)
* refactor(#1306): gate graphify on isCapabilityActive (tri-state), not config-only

graphify's command gate moves from the config-only isGraphifyEnabled to the
shared isCapabilityActive('graphify', cwd) — so graphify is off unless installed
AND surfaced AND graphify.enabled. Fixes a latent claude-hardcoding in the
resolver: resolveCapabilityRuntimeState now detects the active runtime via
resolveRuntime(cwd) (GSD_RUNTIME -> config.runtime -> 'claude') so non-Claude
runtimes (Codex/Cursor) read their own surface, not ~/.claude. Hermetic
regression test proves config-on+unsurfaced -> disabled; cross-runtime test
proves GSD_RUNTIME=codex honors CODEX_HOME. Gate fails closed on every error
path. Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1306): add changeset for graphify tri-state gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 23:19:21 -04:00
Tom Boucher
91bd82f9a6 enhance(execute): isolated-executor rejected/over-reaching run fails safe — never default recovery to main (#1292) (#1303)
* enhance(execute): isolated-executor rejected/over-reaching run fails safe (#1292)

When an isolated (worktree) executor run is rejected — the user declines to
merge it, the orchestrator surfaces recovery for a blocked/halted plan, or the
run over-reached the requested scope — the orchestrator must no longer
default/propose recovery by editing the primary checkout (`main`). Absent an
explicit guardrail, the LLM orchestrator could improvise "continue on main",
inverting the isolation contract at the moment it matters most.

Added an ISOLATED-RUN RECOVERY — FAIL SAFE policy: default to a safe halt that
offers a fresh, narrowly-scoped worktree or inspect/discard; editing the primary
checkout requires explicit, clearly-labeled confirmation and is never the
default/proposed option.

To respect the ADR-857 phase-6 host-loop size cap on execute-phase.md (it sits
just under the pre-phase-6 baseline), the policy is delivered as an extracted
reference fragment rather than inline:
- New `execute-phase/steps/worktree-recovery-policy.md` holds the recovery policy
  (the existing FAIL-CLOSED rule #48 for base/HEAD mismatches + the #1292
  fail-safe guardrail). No #48 behavior change — moved verbatim.
- execute-phase.md references the fragment at the worktree-spawn recovery point,
  the step-5.5 merge decision, and the stalled-agent "switch to inline execution"
  menu (which for an isolated run now follows the fail-safe policy). Net effect:
  execute-phase.md shrinks below its cap.
- quick.md carries the fail-safe guardrail inline at its post-return merge/discard
  decision (quick.md is not size-capped).

Scoped to the recovery offer only — no automatic scope-overreach detection
(explicitly out of scope per the issue) and no new config key. Adds content
regression tests, a USER-GUIDE note, and a workflow size-baseline update.

Closes #1292

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(changeset): Changed fragment for #1292 isolated-executor fail-safe recovery

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 20:53:38 -04:00
Tom Boucher
93c5ecd645 feat(#792): add devin-desktop runtime alias for windsurf (#1086)
windsurf now also answers to devin-desktop (CLI --devin-desktop) for the Windsurf→Devin Desktop rebrand; all paths unchanged. The .devin/skills/ workspace migration is split to #1085. Closes #792.
2026-06-11 22:32:49 -04:00
Colin
cd5db1f8db test(suites): seed security/slow/integration suites via measured retags
Renames (git mv) with all references updated (ci-test-scope RULES,
windows-parity allowlist, test-file-count allowlist, docs in 6 locales):

- 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step
  ran zero files since the suite taxonomy landed; it is now honest.
- graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite;
  e2e gsd-tools spawns) — runs on full-matrix lanes and push to next.
- installer-migration-install-integration -> *.integration.test.cjs
  (13s; an integration test by its own name).

Coverage gate measured after retags: 88.55% lines (gate 70%).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-09 23:50:41 -04:00
Tom Boucher
0a11d361ca feat(#69): nest concrete skills under namespace routers at install (#883)
Emit the 6 gsd-ns-* routers as the only top-level skill bundles and nest
the ~61 concrete skills under <router>/skills/<name>/SKILL.md on runtimes
with confirmed non-recursive skill loaders (claude global, cline, qwen,
hermes, augment, trae, antigravity). Router bodies rewrite their routing
tables from Skill-tool dispatch to a Read skills/<name>/SKILL.md pattern.
Recursive/unconfirmed loaders (cursor, codex, copilot, windsurf, codebuddy,
opencode, kilo) keep the flat layout. Completes the v1.40 namespace
architecture (#2792) so the eager skill listing drops to ~6 entries.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 15:09:34 -04:00
Tom Boucher
40d48c0508 feat(#815): add /gsd-update --next to install the @next RC channel (#839)
Adds an opt-in --next (alias --rc) flag to /gsd-update targeting the @next RC dist-tag (ADR #660), with a {latest,next} allowlist enforced at three layers, channel-aware version check + banner, and byte-for-byte unchanged default @latest behavior.

Closes #815
2026-06-07 20:22:07 -04:00
Tom Boucher
cb284962bf feat(#789): elevate CodeBuddy — slash commands (#830)
* feat(#789): elevate CodeBuddy — emit slash commands (+ document subagent/MCP scope)

Emit a CodeBuddy slash-command surface so GSD workflows appear in the
'/' menu, reaching parity with other elevated runtimes.

- Add convertClaudeCommandToCodebuddyCommand and register a commands/
  artifact kind for the codebuddy runtime (commands/gsd-<name>.md),
  consistent with the Cursor (#785) and Augment (#790) commands surfaces.
- Mark emitted skills user-invocable:false so the commands surface is the
  sole '/' entry point (no duplicate /gsd-* entries); skills stay
  model-invocable. CodeBuddy's SKILL.md supports this field.
- Normalize $HOME/.codebuddy (bare + slash) path forms in runtime
  rewrites so --config-dir/local installs don't leak the default home.
- Report installed commands/ count on install; uninstall prunes gsd-*
  commands while preserving user-owned commands.

Scope: subagents (~/.codebuddy/agents/) are already emitted by the
generic agents block (unchanged); no mcp.json is written (gsd ships no
MCP server, and CodeBuddy's mcp.json registers only external servers).

Closes #789

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#789): set changeset pr number to 830

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:47 -04:00
Tom Boucher
28c8d524fe feat(#778): cross-runtime command enrichment (Gemini {{args}}/!{}, Qwen priority) (#825)
* feat(#778): cross-runtime command enrichment (Gemini {{args}}/!{}, Qwen priority)

Enrich the installer's per-runtime command/skill generators with native,
verified, additive fields:

- Gemini CLI: map Claude's $ARGUMENTS -> Gemini's {{args}} in generated TOML
  commands so typed arguments interpolate; inject live .planning/STATE.md into
  /gsd:progress via a fixed, injection-safe !{cat .planning/STATE.md 2>/dev/null}
  shell block (no interpolated input).
- Qwen Code: emit the optional numeric `priority` field on main-loop skills so
  the most-used workflows sort first in the /skills list (higher = earlier per
  the Qwen skills spec; the issue's inverse numbering was corrected).

OpenCode per-command model/agent/subtask/variant enrichment was evaluated and
intentionally not implemented: `model` reintroduces the #1156
ProviderModelNotFoundError regression for non-Anthropic providers (the converter
deliberately strips model:), `subtask`/`agent` change execution semantics for
GSD's interactive commands, and `variant` is not in the OpenCode command schema.

Schemas verified against primary docs (Gemini custom-commands, Qwen skills,
OpenCode commands/skills). Adds tests/enh-778-* and how-to + USER-GUIDE docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#778): set changeset PR number to 825

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:42 -04:00
Tom Boucher
19280510fe feat(#774): emit service_tier/model_verbosity in Codex agent TOML + agents/openai.yaml skill chip (#828)
* feat(#774): emit service_tier/model_verbosity in Codex agent TOML + agents/openai.yaml skill chip

- Add service_tier = "flex" and model_verbosity = "low" to the Codex
  ConfigProfile TOML for light-tier agents (gsd-research-synthesizer,
  gsd-codebase-mapper, gsd-plan-checker, and 8 others identified via
  AGENT_DEFAULT_TIERS). Field names/values verified against Codex schema
  (profile_toml.rs / config_types.rs Verbosity enum). Non-light agents
  are unaffected.

- Add generateCodexSkillMetadataYaml() and writeCodexSkillMetadataFiles():
  after installRuntimeArtifacts, iterate every gsd-* skill directory,
  read the short-description already emitted in the SKILL.md frontmatter
  by convertClaudeCommandToCodexSkill, and write agents/openai.yaml with
  interface.display_name and interface.short_description for the Codex
  TUI skill picker chip.
  - yamlQuote (JSON.stringify) handles all YAML-unsafe chars.
  - User-owned gsd-dev-preferences dir is never overwritten.
  - Errors per-skill are swallowed so a bad SKILL.md can't abort install.
  - agents/openai.yaml is covered by the snapshot/rollback system and
    manifest hash (writeManifest hashes skill dirs recursively).
  - Uninstall symmetry: _removeGsdEntries removes whole gsd-* dirs.

- 21 new tests in codex-config.test.cjs covering service_tier/verbosity
  TOML emission, YAML generation (round-trip via js-yaml), and
  writeCodexSkillMetadataFiles including an e2e integration test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#774): correct docs-lint coverage — proper changeset format + USER-GUIDE entry

Rewrite the changeset fragment from old @opengsd/gsd-core:patch format to the
required type:/pr: schema so the docs-lint parser can consume it.  Add a new
"Codex skill picker and agent scheduling (#774)" section to docs/USER-GUIDE.md
describing the flex-tier scheduling and /skills TUI chip enrichments — both are
user-visible and belong in docs rather than behind a docs-exempt marker.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:25 -04:00
Tom Boucher
dc7f1557c6 feat(#768): pre-populate settings.json permissions.allow/deny for Claude Code (#819)
* feat(#768): pre-populate settings.json permissions.allow/deny for Claude Code

Adds mergeClaudePermissions() to bin/install.js which non-destructively
appends GSD's known-safe tool-call patterns to permissions.allow and
defense-in-depth credential-file patterns to permissions.deny during
Claude Code installs. Merge is idempotent (no duplicates on reinstall)
and additive (existing user entries preserved). Uninstall removes only
the exact GSD-owned entries.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: update changeset pr number to 819

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:56:20 -04:00
Tom Boucher
a3aa0ae142 feat(#775): ship a gemini-extension.json extension package (#818)
Add a Gemini CLI extension package so users can install, update, and
remove GSD through Gemini's own extension lifecycle and have it appear in
`gemini extensions list`:

  gemini extensions install https://github.com/open-gsd/gsd-core
  gemini extensions update gsd-core
  gemini extensions uninstall gsd-core
  gemini extensions link /path/to/gsd-core   # dev

This mirrors the additive Claude Code plugin manifest (#766): a thin,
version-stamped manifest enforced by an in-repo drift test. The extension
ships the context-file payload (GEMINI.md), loaded into every Gemini
session; slash-command/agent/hook TOML projection into the extension is a
documented follow-up. The manual `npx gsd-core --gemini` installer (which
provides the /gsd:* commands) is unchanged — purely additive, no breaking
change.

- gemini-extension.json: name=binName, version tracks package.json,
  description, contextFileName=GEMINI.md (minimal; no mcpServers — gsd
  ships no MCP server)
- GEMINI.md: Gemini-session context payload
- package.json: add both artifacts to files[] so they publish
- CONTEXT.md: add "Gemini Extension Package" glossary entry
- docs: USER-GUIDE + install-on-your-runtime how-to
- tests/issue-775-gemini-extension.test.cjs: manifest validity, version
  parity with package.json, contextFileName existence, files[] publication

Closes #775

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 18:54:08 -04:00
Tom Boucher
3025a6846e fix(#812): honor COPILOT_HOME in Copilot global config-dir resolution (#814)
* fix(#812): honor COPILOT_HOME in Copilot global config-dir resolution

getGlobalConfigDir('copilot') resolved the global config directory using
only --config-dir > COPILOT_CONFIG_DIR > ~/.copilot, ignoring the
COPILOT_HOME env var. Per GitHub's Copilot CLI docs, COPILOT_HOME
overrides the default ~/.copilot location (and user-level hooks are read
from $COPILOT_HOME/hooks/), so a global --copilot install wrote all
artifacts (skills, agents, copilot-instructions.md, the gsd-session.json
hook) to ~/.copilot even when the user relocated their Copilot home,
making them undiscoverable by Copilot CLI.

Mirror the codex/CODEX_HOME branch: precedence is now
--config-dir > COPILOT_CONFIG_DIR > COPILOT_HOME > ~/.copilot. Uninstall
uses the same resolver, so it stays symmetric.

Also: document COPILOT_HOME in the installer --help notes, the
USER-GUIDE env-var table, and the installer-migrations Copilot row; and
clear COPILOT_HOME in the two default-path test suites so they stay
hermetic now that the resolver honors it.

Closes #812

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#812): add changeset for PR #814

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 15:27:46 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
3bb2f8f1c5 docs: rebrand to GSD Core and restructure docs with Diataxis (#605)
* chore: wire docs/agents config into AGENTS.md Agent skills section

Add the `## Agent skills` discovery block pointing the engineering
skills at the existing docs/agents/{issue-tracker,triage-labels,domain}.md
files (issue tracker, triage label mapping, single-context domain docs).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: rebrand to GSD Core and restructure docs with Diataxis

Reorganise the root README and docs/ around the Diataxis framework
(tutorials, how-to guides, reference, explanation), add new how-to
guides and schema references (STATE.md / CONTEXT.md / PLAN.md /
planning artifacts), and cross-link the whole set. Update the lone
legacy gsd-build reference to open-gsd; keep internal get-shit-done/
filesystem paths unchanged (directory rename tracked separately in
open-gsd/gsd-core#604). Regenerate the ja-JP, ko-KR, pt-BR and zh-CN
localised trees to mirror the new structure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: backfill changeset PR number (#605)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 08:13:09 -04:00
Tom Boucher
e3bc53f835 enhancement(#558): add liveness hints to all GSD spawn announcements (#566)
* enhancement(#558): add liveness hints to all GSD spawn announcements

Append '(runs in a subagent — no output until it returns, ~1–5 min; expected,
not a freeze)' inline to every ◆ Spawning… banner and subagent dispatch
instruction across 26 workflows. Silent subagents look identical to frozen
sessions — this note sets the expectation so users wait instead of killing
healthy in-progress work.

Changes:
- references/ui-brand.md: document liveness convention under Spawning Indicators
- 10 banner workflows: append liveness note to ◆ Spawning… lines in-place
- 18 subagent-only workflows: add print instruction with liveness phrase
- tests/spawn-liveness-banner.test.cjs: new test; fails if any workflow with
  subagent_type omits 'runs in a subagent'
- docs/USER-GUIDE.md: troubleshooting entry for frozen-looking spawns
- .changeset/558-spawn-liveness-banner.md: changeset fragment

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#558): add missing pr field to changeset fragment

The changeset lint requires pr: <NNN> in frontmatter; the fragment was
written without it, causing parse.cjs to reject it.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#558): address codex review — missed spawns and tighten test

- plan-phase.md: add liveness note to chunked outline planner and
  per-plan chunked planner banners (two missed ◆ Spawning… lines)
- quick.md: add liveness note to research banner and add missing
  display line before planner spawn in Step 5
- plan-review-convergence.md: add liveness note to initial planning
  and review-agent spawn Display lines
- docs-update.md: add Print instructions with liveness note before
  gsd-doc-verifier spawns in Phase 1 and Phase 2
- autonomous.md: add Print instruction with liveness note before
  background plan-phase agent dispatch in step 3b
- tests/spawn-liveness-banner.test.cjs: replace single file-level
  check with two assertions:
  (1) every ◆ Spawning… banner line carries the phrase on that line
  (2) every file with subagent_type contains the phrase somewhere
  The tighter test would have caught all five missed spawns.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#558): tighten spawn-liveness test regex to catch spawn-word-anywhere variants

Previous SPAWN_BANNER_RE only matched ◆ immediately followed by Spawning|spawning.
Replace with /◆[^\n]*\bspawning?\b/i which matches the spawn word anywhere on the
◆ line — catching "◆ Chunked mode: spawning outline planner..." and similar.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#558): rename changeset to PR number 566 and correct pr field

Changeset was filed as 558-spawn-liveness-banner.md (issue#) but the
convention is the PR number. Renamed to 566-spawn-liveness-banner.md
and updated pr: 558 → pr: 566 so release notes link to the right PR.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-31 23:33:49 -04:00
Tom Boucher
0fbe1d899e chore(#191): retire the gsd-sdk shim — route everything at gsd-tools (#522)
* chore(#191): migrate gsd-sdk query call sites to gsd-tools query

Retiring the gsd-sdk shim. gsd-tools.cjs already accepts `query` as a
meta-prefix (gsd-tools query <command>), so this is a behavior-preserving 1:1
swap across the runtime reference prompts, the graphify hook's commit-detection
gate, and two bin/lib comment/message references.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#191): remove vestigial gsd-sdk shim code from installer + projection

The gsd-sdk shim was already not wired up (no gsd-sdk bin in package.json;
buildWindowsShimTriple had zero call sites). Remove the dead code:
- shell-command-projection.cjs: buildWindowsShimTriple + formatSdkPathDiagnostic
  (+ their now-unused PACKAGE_NAME import) and exports
- install.js: the re-export wrappers + imports, the #3406 stale-standalone-sdk
  detection (detectStaleStandaloneSdk/formatStaleStandaloneSdkWarning + its
  global-install call site), and the exports

Preserved (retained, not gsd-sdk): buildCodexHookWindowsShimIR (#3426) — only
its comments referenced the gsd-sdk pattern; reworded. Also kept the
homePathCoveredByRc 'reopen your shell' branch in maybeSuggestPathExport — its
logic is bin-dir-agnostic, only the message mentioned gsd-sdk; reworded to use
the actual bin dir.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#191): update tests for retired gsd-sdk shim

- bug-3441/bug-3442: drop the formatSdkPathDiagnostic / buildWindowsShimTriple
  assertions (functions removed); retained PATH-action + drift-guard tests stay
- bug-505: remove the 'still exported' assertions for detectStaleStandaloneSdk /
  formatStaleStandaloneSdkWarning / the shim contract surface (#505 kept them;
  #191 removes them)
- graphify-auto-update: migrate the hook-dispatch inputs gsd-sdk query commit ->
  gsd-tools query commit to match the migrated commit hook

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#191): point active docs at gsd-tools query (gsd-sdk shim retired)

Update the user/agent-facing docs (AGENTS, COMMANDS, CONFIGURATION, USER-GUIDE,
ship-pr-body-sections) that presented gsd-sdk query as a current command to
gsd-tools query. Historical docs (ADRs, PRDs, release notes) left untouched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#191): correct state.load vs state.json description for gsd-tools query

Adversarial-review (codex) finding: the migrated USER-GUIDE line claimed both
'gsd-tools query state.json' and 'state.load' resolve to the frontmatter-rebuild
handler. Verified they don't — state.load returns the CJS load shape
(config + state_raw + flags), state.json returns the frontmatter shape. Both are
available via gsd-tools query; corrected the text to say so.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#191): add changeset for gsd-sdk shim retirement

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 19:19:14 -04:00
Tom Boucher
79002a00cb chore(#518): rename npm package + bin to @opengsd/gsd-core (#519)
* chore: rename npm package + bin to @opengsd/gsd-core (functional)

- package.json: name @opengsd/get-shit-done-redux → @opengsd/gsd-core,
  bin key get-shit-done-redux → gsd-core, repository/homepage/bugs URLs
- package-lock.json: regenerated (npm install --package-lock-only)
- tests/**, scripts/**, bin/**, .github/**, agents/**, commands/**,
  get-shit-done/bin/**, get-shit-done/workflows/**:
  applied the 4-rule replacement (scoped npm ref, GitHub repo path,
  bin/clone invocations) per #505 single-source refactor

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: sweep live references to @opengsd/gsd-core

Update all live documentation (README.md + translations, docs/**,
CONTRIBUTING.md, VERSIONING.md, SECURITY.md, CONTEXT.md,
docs/CANARY.md) to reflect the renamed package and repository.

Rules applied:
- @opengsd/get-shit-done-redux → @opengsd/gsd-core (scoped npm name)
- open-gsd/get-shit-done-redux → open-gsd/gsd-core (GitHub repo)
- GSD-redux/get-shit-done-redux → open-gsd/gsd-core (stale badge org)
- bare bin/clone refs → gsd-core

CHANGELOG.md, docs/adr/**, docs/RELEASE-*.md, docs/research/**,
and .changeset/** are preserved byte-identical.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: add negative lookbehind to slash-command regex in bug-2954 test

The extractSlashReferences regex matched /gsd-core inside npm package
URLs (@opengsd/gsd-core), producing a false /gsd:core command reference.
Adding a negative lookbehind (?<![a-z]) excludes matches preceded by a
letter, so only standalone /gsd-<cmd> and /gsd:<cmd> tokens are found.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#518): add changeset for package rename

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#518): update package-identity expectations to the renamed coordinates

The rebase regenerated the seam to @opengsd/gsd-core (bin gsd-core, repo
open-gsd/gsd-core). The #498 seam tests assert deriveIdentity against the REAL
package.json, so their expected literals must follow the rename. The drift-lint
unit test is left as-is — its SEAM is a self-consistent fixture and its
stale-literal detection cases would shift if altered; the live-repo scan in it
already passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-30 17:25:02 -04:00
Tom Boucher
05cdec5f47 feat(#22): plan-vs-codebase drift guard (source-grounded reviewer + intel surface) (#487)
* feat(#22): add plan_review.source_grounding + _authority config keys

Two additive opt-out keys for the drift guard: source_grounding (bool,
default true) gates the source-grounded reviewer pass; _authority (enum
grep|intel|treesitter|lsp|scip, default grep) selects the resolver rung.
No existing default changed.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): add intel api-surface renderer + CLI subcommand

Renders .planning/intel/api-map.json into a human-readable API-SURFACE.md
for planner injection. Empty/missing map still writes a surface that
announces itself incomplete (absence = unknown, not 'does not exist').
Gated on intel.enabled like all intel functions.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): add source-grounding pass to plan-review-convergence

Default-on reviewer pass (plan_review.source_grounding) that enumerates
every symbol a plan cites, excludes declared new artifacts, resolves each
against source via the configured authority adapter, and records
three-valued verdicts. rung-0/1 MISSING is needs-acknowledgement, not a
hard block; UNCHECKABLE is logged in a REVIEWS.md coverage section.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): inject API-SURFACE.md into planner + require Artifacts section

When intel.enabled, plan-phase regenerates API-SURFACE.md and injects it
as a HINT (prefer, may be incomplete, absence = unknown), never a hard
rule. Every plan must now emit an 'Artifacts this phase produces' section
so the source-grounding reviewer can separate new symbols from references
to existing code.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(#22): surface drift-guard in setup + settings, add docs

/gsd:new-project asks to enable plan_review.source_grounding (default Y);
/gsd:settings exposes the toggle and authority knob. Documents both config
keys in CONFIGURATION.md, the intel api-surface command in COMMANDS.md,
the drift guard in USER-GUIDE.md, and links ADR 22 from ARCHITECTURE.md.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#22): respect AskUserQuestion 4-option cap and plan-phase XL line budget

settings drift-guard toggle moved to its own 2-option question; #22
plan-phase additions condensed to bring the file back under the 1810-line
XL budget without dropping the intel gate, the incomplete-surface hint, or
the Artifacts-section requirement.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#22): use live slash-command forms in drift-guard docs

Doc-parity gate requires every slash-command token in docs/*.md to resolve
to a registered command. Corrected the command form(s) referenced in the
#22 drift-guard / api-surface documentation.

The unresolved token was /gsd-core, matched from the GitHub repo reference
"open-gsd/gsd-core#22" in docs/adr/22-plan-drift-guard.md. This is the
same pattern as the existing 'test-runner' exemption (open-gsd/gsd-test-runner).
Added 'core' to INTERNAL_COMPONENT_SLUGS with a matching explanatory comment.

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#22): add changeset fragment for drift guard (PR #487)

Refs #22

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-30 17:08:11 -04:00
Tom Boucher
5b3646d6c8 fix(#213): support antigravity 2.x config directory split (#217)
* fix(#213): support antigravity 2.x config directory split

* fix(#213): harden antigravity tests for windows parity
2026-05-24 14:17:03 -04:00
Tom Boucher
c439890e26 docs(2): clarify installer is required for cross-runtime compatibility (#144)
* docs(readme): add cross-runtime compatibility note for installer requirement

Source files in agents/ and commands/ are Claude Code-format frontmatter.
The installer (bin/install.js:5208 convertClaudeToOpencodeFrontmatter) is
the mandatory conversion layer for non-Claude-Code runtimes. Manually
copying source files to ~/.config/opencode/agents (or equivalent) bypasses
conversion and produces schema validation errors.

README lines 37, 165, 273 all claim OpenCode as a first-class runtime
without flagging the installer as mandatory. This adds an explicit note
immediately after the Getting Started section (README.md line ~175).

Refs #2

* docs(user-guide): add manual install / no-Node.js setup section for non-Claude runtimes

Users without Node.js (Windows + OpenCode being the most common case per
issue #2) cannot run the installer and may attempt to copy agents/ source
files directly. This section documents what manual conversion is required
for OpenCode (remove tools:, convert color: to hex), references the
installer function at bin/install.js:5208, links to the OpenCode schema
docs, and covers the Docker/WSL alternative.

USER-GUIDE.md insertion after line 1258 (after the Codex/non-Claude
runtime section, before Installing for Cline).

Refs #2

* ci: retrigger checks after transient git-auth runner failure

The original run for this PR had a single CI job fail with:
"fatal: could not read Username for 'https://github.com': terminal prompts disabled"
That is a hosted-runner infrastructure flake — no code defect. The run
cannot be retried via gh CLI (too old). This empty commit kicks a fresh
full CI cycle.
2026-05-23 15:50:20 -04:00
Tom Boucher
334a64168e chore(npm): rebrand packages to @opengsd scope (#127)
* chore(npm): rebrand packages to @opengsd scope

Rename:
- get-shit-done-redux → @opengsd/get-shit-done-redux
- @gsd-redux/sdk → @opengsd/gsd-sdk

Add publishConfig.access=public for first-time scoped publish.
CLI binary names (get-shit-done-redux, gsd-sdk, gsd-tools) unchanged.

Sweeps install commands, npx invocations, CI publish/version-check
workflows, tests, docs, READMEs (all translations), and the
PACKAGE_NAME constant in check-latest-version.

Bumps qs 6.15.1 → 6.15.2 to clear a moderate advisory surfaced by
the audit-clean test (GHSA-q8mj-m7cp-5q26).

Closes #126

* chore: pin 2.0.0 release + remove canary workflow

- Bump both packages 1.50.0-canary.0 → 2.0.0 for first @opengsd publish
- Remove .github/workflows/canary.yml and canary dist-tag handling in
  release.yml / release-sdk.yml
- Drop canary section from VERSIONING.md

Refs #126

* chore: address review findings + harden tarball-smoke timeout

- .changeset/opengsd-org-rename.md: match project's custom
  parse.cjs frontmatter (type: Changed / pr: 127); the scoped
  @changesets/cli keys were silently rejected.
- CONTEXT.md: drop two canary-stream policy lines and a dangling
  DEFECT.CANARY-VERSION-LEAK.cross-ref now that canary.yml is gone.
- tests/release-tarball-smoke.install.test.cjs: pass
  timeout: 600_000 for npm pack + global install; the 3-minute
  runNpm default was timing out on slower Docker hosts (cartographer).

Refs #126

* fix(sdk): add missing type/runtime devDependencies for build

prepublishOnly invokes tsc which couldn't resolve @types/node,
@types/ws, or synckit. They had been hoisted from root but were
not declared in sdk/'s own package.json — first publish from a
clean SDK tree failed.

Refs #126

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(ci): use npm pack stdout instead of glob to find tarball

`npm pack --silent` for a scoped package (@opengsd/get-shit-done-redux)
produces `opengsd-get-shit-done-redux-*.tgz`, not `get-shit-done-redux-*.tgz`.
Capture the filename from stdout instead of a hardcoded glob so the step
works regardless of package name format.

Fixes smoke (ubuntu-latest, 22, false) CI failure.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* ci: treat workflow-file changes as test-skip eligible

`.github/workflows/install-smoke.yml` (and other workflow files)
were in neither `test.yml` paths nor `test-skip.yml` paths-ignore,
so neither workflow ran on a workflow-only commit — leaving the
required test-skip check perpetually missing.

Refs #126

* chore: reset version to 1.0.0 for first @opengsd publish

Nothing has been published yet under the @opengsd scope, so the
inaugural release uses 1.0.0 rather than 2.0.0. The "major bump"
in the changeset reflects the breaking install-command change for
users migrating from the prior unscoped `get-shit-done-redux`, not
a numeric continuation from a 1.x line under the new identity.

Refs #126

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-22 16:22:41 -04:00
Tom Boucher
2a915c1b82 chore: migrate references from gsd-build to open-gsd/get-shit-done-redux (#120) (#121)
Security-motivated migration of all stale repository and npm-scope references.

Three categories of changes (58 files, 174 substitutions):

1. gsd-build → open-gsd (security-critical):
   - .github/workflows/release-sdk.yml — npm token comment, tarball filename pattern
   - .github/workflows/hotfix.yml — same
   - .changeset/fix-3406-detect-stale-sdk-shadow.md — @gsd-build/sdk → @open-gsd/sdk
   - .changeset/sharp-quails-leap.md — same
   - get-shit-done/workflows/update.md — CHANGELOG raw GitHub URL

2. GSD-redux org slug → open-gsd (canonical rename):
   - package.json + sdk/package.json — repository/homepage/bugs metadata
   - All README.*.md — live badge and link sections
   - CONTRIBUTING.md, CONTEXT.md, QUICK-WINS-CONFIRMED-BUGS.md
   - .coderabbit.yaml, .release-monitor.sh, scripts/sync-rulesets.sh
   - docs/** — all live agent/ADR/user-facing documentation
   - tests/** — repo slug assertions and test fixtures
   - scripts/changeset/cli.cjs + github-release-notes.cjs
   - .github/ISSUE_TEMPLATE/*, .github/pull_request_template.md
   - bin/install.js, get-shit-done/bin/lib/model-catalog.cjs
   - sdk/HANDOVER-*.md, sdk/src/*.test.ts

3. CLAUDE.md (gitignored local file — not in this commit):
   Updated separately outside git: --repo gsd-build/get-shit-done →
   --repo open-gsd/get-shit-done-redux with security warning.

Intentionally unchanged: CHANGELOG.md, docs/RELEASE-*.md,
.changeset/README.md, .changeset/build-hooks-atomic-write.md,
README.md migration table (historical fork record),
tests/changeset-serialize.test.cjs line 78 (serialization fixture).

The gsd-build/get-shit-done repo is compromised (rug-pull documented in
README.md). Do not push to or interact with that repo.

Closes #120
2026-05-22 12:28:16 -04:00
Tom Boucher
dff176bfd2 chore: rebrand to GSD-redux/get-shit-done-redux
Mirror of code, issues, and PRs from the upstream gsd-build/get-shit-done,
which appears compromised or abandoned (maintainer unreachable since
2026-04-01; $GSD token linked to rug-pull).

- Adds rebrand notice block at top of English README
- Removes $GSD token badge and @gsd_foundation X badge (keeps Discord)
- Renames npm packages: get-shit-done-cc -> get-shit-done-redux,
  @gsd-build/sdk -> @gsd-redux/sdk
- Updates all repo URLs across docs, workflows, package.json, bin/
- Updates ci@gsd-build -> ci@gsd-redux in workflow git identities
- Leaves CHANGELOG and .changeset/* alone (historical, time-stamped)
2026-05-22 08:27:07 -04:00