next
15 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6cfa0c55d2 |
refactor: drop 12 runtimes, keep Claude, Codex, OpenCode, Cursor, ZCode, Antigravity
Removes kilo, kimi, kimi-code, copilot, windsurf, augment, trae, qwen, hermes, cline, codebuddy and pi end to end: capability descriptors, installer branches and converters (bin/install.js 14.9k -> 11.2k lines), TypeScript converters, hook surfaces and runtime homes, review lanes qwen/kimi-code, the two pi migrations, Kimi payload normalization in the hook guards, dead hostBehaviors vocabulary, launcher home probes, fixtures, runtime-specific tests and the prose that presented them as supported. Installer output for the six kept runtimes is byte-identical to before the prune. The Kimi tool-vocabulary tests in workflow-guard, read-guard and read-injection-scanner are left in place pending a decision. |
||
|
|
a9a7a328e6 |
refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted. |
||
|
|
eea9247c93 |
enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select (#4912)
* enhance(#4095): checkpoint:decision auto-selection is opt-in via auto_select Auto-mode used to auto-select a checkpoint:decision's first <option> unconditionally, making a decision checkpoint's safety depend on option presentation order rather than an authored choice. Add an optional auto_select="<option-id>" attribute on the <task> tag: absent, auto-mode now escalates to a human exactly like gate="blocking-human" does; present, it names the option auto-mode selects; naming an id with no matching <option id> is a hard structural-validation error at plan-parse time rather than a silent fallback to the first option. gate="blocking-human" continues to win over everything, unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): anchor auto_select/id attribute regexes past hyphenated decoys An isolated adversarial review of the auto_select work found that both new attribute regexes used \b as their left anchor, which is a word boundary, not a "start of attribute name" boundary. A decoy attribute ending in the same word (e.g. data-id="...") sitting before the real id="..." on the same <option> tag matched first, silently corrupting the extracted option id. Anchor on (?:^|\s) instead so only the real attribute name can match. Adds a regression test reproducing the exact decoy-attribute shape, plus a Unicode option-id test from the same review pass. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4095): register auto-select-attribute.test.cjs in the docs-guard lane lint-docs-guard-registration failed: the new test reads docs/reference/ plan-md.md but was not registered, so a future edit to that doc could silently desync from the test without the guard catching it on the PR that changed the doc. Registered alongside its direct precedents (precondition-element.test.cjs, reversibility-tagging.test.cjs), which read the same file for the same reason. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4095): fit the decision bullet under execute-phase.md's frozen byte ceiling The remote gsd-test run caught what local checks missed: execute-phase.md carries a frozen ADR-857 Phase-6 byte ceiling (93600) with only 36 bytes of headroom before this change, and the original checkpoint:decision wording pushed it to 93772 (over the ceiling). Cascaded into failures in phase6-capstone-conformance, execute-phase-completion-reconciliation, claude-orchestration, and the compact-content drift-report test. Also caught: tests/package-legitimacy-gate.test.cjs anchors a "decision is conditional, not unconditional" safety check on the literal phrase "first option" in the decision bullet — which #4095 deliberately removes, since there is no more unconditional first-option pick. The test was asserting an assumption this change intentionally makes obsolete; re-anchored on tokens that still identify the bullet ('decision', 'auto-spawn') without weakening what the test actually verifies (the bullet must still carry a blocking-human carve-out). Also fixed a word-order mismatch between my own new test's regex and the actual doc text it was asserting against (tests/auto-select-attribute.test.cjs). Regenerated the compact-content benchmark baseline (tests/fixtures/compact-content-benchmark-baseline.json) to match the new byte counts. Emitted-Drift-Ack-Growth: gsd-executor.md — +7 bytes (49138 -> 49145), from the auto_select carve-out added to the checkpoint:decision auto-mode bullet; already trimmed once to fit the 49152 hard cap. Emitted-Drift-Ack-Growth: execute-phase.md — +12 bytes (93564 -> 93576), from the same carve-out in the orchestrator's decision bullet; kept 24 bytes under the frozen 93600 ADR-857 ceiling after two rounds of trimming for clarity vs. margin. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4095): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
88b5775dc8 |
enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI (#4477)
* enhance(#4223): default-off interaction capture for gsd-ui-auditor via the chrome-devtools CLI gsd-ui-auditor is chartered to audit interaction and handed a capture driver with no interaction verb: `npx playwright screenshot` cannot click, fill, hover, press or snapshot, so a hover state, an open menu, a focus ring or a form's validation state never appears in its evidence and every Experience Design finding degrades to code reading. Implements the shape approved at triage, not a new capability: - capabilities/ui/capability.json declares `workflow.ui_interaction_capture` (boolean, default false) on the capability that already owns the auditor (ADR-894 one-owner invariant); capability-registry.cjs regenerated. - gsd-core/workflows/ui-review.md reads the key through gsd_run and hands it to the auditor as `interaction_capture:` in the spawn <config> block — the auditor carries no gsd_run resolver, so the key travels by value. - agents/gsd-ui-auditor.md gains an anchored interaction-capture section AFTER the static block. With the key on and a Chrome binary resolved it starts the `chrome-devtools` CLI (chrome-devtools-mcp, floor ^1.8.0) on an --isolated profile, opens the dev URL the static block reached, takes the a11y snapshot for element uids, captures the baseline and a Tab focus-ring state, drives the UI-SPEC's interactive components, saves console output, and stops the daemon unconditionally. Key off, no dev server, or no Chrome: one status line, and the Playwright-only static path runs exactly as before — the static fence is untouched. Needs only Bash: no MCP server, no tools: change. Chromium-only by nature; Firefox/WebKit stay on Playwright. `wait_for` is MCP-only, so readiness is polled through evaluate_script. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): bind the interaction-capture shape and containment - manifest, generated registry, config schema and config-set/loadConfig all know workflow.ui_interaction_capture as a default-off boolean, and hand-written non-booleans fall to the slice default - the orchestrator reads the key and hands it down; the auditor never grows a gsd_run dependency - the static fence stays Playwright-only and the interaction fence chrome-devtools-only, so key-off is today's path - the interaction fence runs under bash with a stub driver on PATH: key off / absent / no dev server / no Chrome invoke nothing; the happy path starts first and stops last on the [selected] pageId with the documented flags; a failed capture is removed and not counted; new_page and start failures still honour the stop-only-if-started rule; CHROME_BIN and CHROME_DEVTOOLS_MCP_VERSION overrides flow through - docs/CONFIGURATION.md row shape; registered in the docs-guard lane Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * docs(#4223): document workflow.ui_interaction_capture and its how-to - docs/CONFIGURATION.md: one row in the workflow.* table, default-off - docs/AGENTS.md: the gsd-ui-auditor entry names the key and what the interaction-capture section adds, skips and never claims - docs/how-to/enable-ui-interaction-capture.md: turn it on, read the `**Interaction captures:**` outcomes, what it does not do, turn it off - docs/README.md: index the how-to beside live-DOM verification Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): add changeset Added-type fragment; pr: carries the issue number until the PR exists. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): use the /gsd:ui-review namespace form in the auditor's prose Claude-facing source (agents/, workflows/) uses the /gsd:<cmd> namespace; the hyphen form is retired there and the slash-command-namespace guard rejects it. docs/ keep the hyphen form by convention. Emitted-Drift-Ack-Growth: gsd-ui-auditor.md — #4223: the anchored default-off interaction-capture section (prose + one bash fence) appended after the static Playwright block inside <screenshot_approach>, plus one `**Interaction captures:**` line in each of the two report templates, one completion-checklist line and one Step-3 sentence. The static fence is byte-identical to next; nothing was removed or reordered. Emitted-Drift-Ack-Growth: ui-review.md — #4223: a two-line config-get read + true/false normalisation in step 0 and one `interaction_capture:` line in the spawn <config> block with a three-line note on why the value travels by prompt. No step, gate, or dispatch shape changed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): per-run daemon session, bounded navigation, and step failures that count Three findings from the pre-file adversarial review of the interaction fence, folded in: - `--sessionId <epoch>-<pid>` on every driver call. `start` restarts whatever daemon shares its session and --isolated isolates only the browser profile, so two concurrent audits — or an audit beside the operator's own CLI daemon — would otherwise stop each other. The CLI accepts hex and dashes only; the id is validated by the test stub. - `new_page --timeout 30000`: the one verb that takes a bound, placed before every verb that does not, so a hung page is caught first. - a failed take_snapshot or press_key now increments the failure count and is named on stdout; two clean screenshots can no longer read as `0 failed` after the step that gives the interactions their uids failed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): subshell-unique session id, CRLF-safe page-id parse, stale-snapshot removal Second review round, both reviewers: - session id is `<epoch>-<BASHPID>-<RANDOM>`: `$$` is inherited by a subshell, so two audits forked from one parent in the same second shared an id and could stop each other's daemon (driven by the reviewer) - `tr -d '\r'` before the `[selected]` parse so a CRLF-emitting driver under Git Bash still matches the `$` anchor, and `|| true` on the assignment so a failed new_page cannot abort the block under `set -e -o pipefail` before the unconditional stop - a failed take_snapshot removes any snapshot.txt it left or inherited from a reused directory, so stale uids never drive the interactions - `<config>` placeholder is `{interaction_capture}`, lowercase like its `{phase_dir}` / `{padded_phase}` siblings — the block is a prompt template, not a bash heredoc - how-to: the `not captured` row no longer claims the daemon started Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): check new_page's exit status before parsing its output; regression cases for the edges Third review round: - a new_page that prints a page line and then exits non-zero is a failed navigation, not a page id: the exit status is checked in an `if` before the output is parsed (driven by the reviewer against the previous `|| true`, which masked exactly that) - regression cases for what the last two rounds added: CRLF driver output, a stale snapshot removed on failure, partial-output new_page failure, and the whole fence under `set -e -o pipefail` (both the failed-navigation path and the happy path) - the harness whitelist gains `date`; the session-id assertion now requires all three parts, so a silently empty epoch cannot hide again Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * enhance(#4223): keep gsd-ui-auditor under the DEFAULT-tier size cap; changeset pr placeholder - the three review folds pushed agents/gsd-ui-auditor.md to 25179 bytes, over the 24576-byte hard cap tests/agent-size-budget.test.cjs enforces; the interaction section's comments are tightened to the same content in fewer bytes (23559 now). No bash changed — the fence's own tests and the real-browser run are unchanged. - .changeset/vivid-yaks-fly.md carries the policy placeholder `pr: 0`, which the post-create backfill rewrites to the PR's own number. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * chore(#4223): set changeset fragment pr to 4477 * test(#4223): compare the fence's status path with the separator the fence uses On the windows-latest lane the happy-path case failed on `\interaction` vs `/interaction` alone: the fence joins "$SCREENSHOT_DIR/interaction" with a literal slash, and the assertion built its expectation with path.join. Every other case in the file passed on that lane, including the CRLF and errexit/pipefail ones. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PanfY8KaLb4RVVubcUoGP6 * test(#4223): drop the inert file-header allow-test-rule marker Review round 1 on #4477: the `source-text-is-the-product` marker sat at line 2, outside no-source-grep's 8-line lookahead of every readFileSync site (the first is ~60 lines down), so it suppressed nothing. It was also unnecessary: every read in this file is a .md/.json path, which the rule does not trigger on. Deleted rather than relocated — there is no site to relocate it to. Negative control: `eslint` on the file is clean without it. * chore(#4223): regenerate the platform-conformance-tier lists for the new test Review round 3 on #4477. `next` gained chore(#4591)'s platform-conformance-tier gate after this branch opened; its two committed lists must name every file under tests/, and this PR's tests/ui-interaction-capture.test.cjs had never been in them. Once the branch was updated against next the lists were stale and three jobs went red on head 575667dd: lint-tests (gen-platform-conformance-tier --check), conformance test (macos-latest) at 546 !== 547, and shard 1/3's fragment-single-edit-propagation, which sees the same staleness as regen:derived touching files beyond the fragment edit under test. Regenerated with the repo's own generators, no hand-editing. The general tier goes 546 -> 547 and the macOS tier 196 -> 197, each by exactly this one entry; both --check arms are clean. Verified the red is this PR's own file and not base drift: at upstream/next both generators report "list matches" (546 / 196), and our committed copies were byte-identical to next's before this commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CBQTeGX1JYHF5DRWp4wvZ * fix(#4223): bound, confine and trap the chrome-devtools driver fence Round 4 — three findings in one fence, interleaved on the same lines, so one commit: - Every driver call is time-bounded. `cdt <ceiling> <verb>` runs the client as a background job in its own process group (`set -m`) under a watchdog that kills the whole group at the ceiling — TERM, then KILL two seconds later. One pid is not enough: npm forwards SIGTERM only to its direct child, so killing `npx` alone leaves the client holding the fence's stdout and a `$(cdt … new_page)` capture blocked past the ceiling (driven against a real npx tree by the round's adversarial review; the pid-only first cut of this commit had exactly that hole). The watchdog is an exec'd bash (`"$BASH" -c`), never a `( … )` subshell: a subshell inherits bash's saved copies of the caller's stdio (the fds ≥10 a function-level `>/dev/null` redirect leaves behind) and holds them open, so a runner waiting for EOF waited out the whole 60 s ceiling whenever a watchdog outlived its kill — measured as the intermittent 30 s test run the review flagged; 0/60 after. It polls the job's process GROUP (`kill -0 -- -pgid`, every 0.1 s) and stands down by itself once the group is empty; nothing ever signals it. The group, not the leader pid: a child can outlive the leader while holding the `$(cdt … new_page)` pipe, and a leader-pid poll stood down at once and left the substitution open-ended (driven by the round's adversarial review at 6× the ceiling; a pgid cannot be reused while any member lives, which a bare pid can). The daemon `start` launches is spawned detached (its own session) and never in that group. Two platforms forced the never-signalled shape. Under bash 3.2.57 the earlier `trap … TERM; sleep & wait $!` form ignored its TERM in 3 of 300 fast calls and slept out the whole ceiling — CI's macos job hanging 30 s right after `start`. On Git Bash a signal to a watchdog still starting up hung the fence's `wait` for it: 18 of 20 fence tests at the harness's 30 s cap in 3 of 3 full-file runs, while a fence slowed by xtrace, or three tests run alone, never hit it (a startup race; the mechanism is not pinned further). Polling: 0/300 slow calls and 0 orphaned sleeps under 3.2.57 and 5.2, the fence suite 20/20 in 3 of 3 full-file runs on Git Bash 5.2.37 (fractional `sleep 0.1`: driven on GNU, msys and busybox sleep; BSD sleep documents it). A clock that cannot launch (`sleep … || exit 0`) stands the watchdog down rather than firing at once and killing a healthy call — by design that leaves a hung call unbounded, the pre-round-4 behaviour, instead of failing a healthy one. A hung call returns once its group is gone: at the ceiling, plus up to the 2 s TERM-to-KILL grace. The KILL after the grace is sent only to a group that is still alive: a pgid freed during the grace can be reused, and an unconditional KILL could hit an unrelated group (the round's review). `start` (npx fetch + Chrome launch) gets CHROME_DEVTOOLS_START_TIMEOUT (180 s), every verb CHROME_DEVTOOLS_STEP_TIMEOUT (60 s). timeout(1) is absent on macOS and this agent carries no gsd-tools resolver, hence a bash watchdog rather than either. - --allowUnrestrictedPaths -> --workspace "$INTERACTION_DIR": the driver may write under the run's interaction/ directory and nowhere else. Relative, like every --filePath (unchanged from rounds 1-3): the daemon resolves both against one cwd (chrome-devtools-mcp 1.9.0 spawns it with cwd: process.cwd() and path.resolve()s both), and a relative path needs no dialect translation — an absolute `pwd -P` path is an msys path on Git Bash, which a Windows-native daemon cannot resolve (CI's windows conformance shard caught the first cut). --workspace is a 1.9.0 flag (absent from 1.8.0's `start --help`, verified), so the documented floor moves from ^1.8.0 to ^1.9.0, where --allowUnrestrictedPaths is deprecated. - `stop` is owed by an EXIT trap after a successful `start`, not by position (it replaces any earlier EXIT trap — none exists in this file); the explicit call keeps it in order, a flag makes the trap a no-op afterwards, and only the shell that installed the trap may act: a subshell copy of the fence state carries CDT_STARTED=1 and, under a timing race CI's ubuntu job hit (reproduced locally at 3/40 under load: the second `stop` came from a subshell pid, never main), issued a second `stop`. The identity is `$(exec /bin/sh -c 'echo "$PPID"')`, not $BASHPID — macOS ships bash 3.2, where BASHPID does not exist and CI's macos conformance job showed the guard comparing empty to empty. The fence was driven under bash 3.2.57 for the injected-subshell, errexit failed-new_page, errexit failed-resize, hung-start, hung-new_page and happy paths. A failed resize_page is a counted failed step now, not the one bare command an errexit runner could abort on. Prose in the section is tightened to pay for the mechanism: 23559 -> 24517 bytes against the 24576 DEFAULT-tier cap. Tests: the stub driver hangs as a real child tree (sh waiting on a child that holds stdout — never an exec), so a pid-only kill fails the new aHungNewPageWhoseChildHoldsStdoutIsStillCutOffAtTheCeiling test (negative- controlled: it blocks for the harness's whole cap on the old wrapper). A hung start and a hung capture are cut off within ceiling + grace + slack and still reach stop; an injected bare failure under errexit reaches stop through the trap, exactly once; an injected subshell call of cdt_stop issues nothing; the happy path issues exactly one stop; every driver call site names a ceiling and the only bare $CDT is the wrapper's own spawn; the start line carries --workspace with the capture directory, every --filePath lies under it, and no code line carries --allowUnrestrictedPaths. A driver whose leader exits at once while a child keeps holding the capture pipe is still cut off at the ceiling (negative-controlled: a leader-pid poll blocks for the harness's whole cap). A watchdog whose clock cannot launch leaves a 300 ms driver call alone (negative-controlled: the trap form kills `start` in under 20 ms). The harness EXPORTS its stub-only PATH — unexported, the exec'd watchdog fell through to bash's compiled-in default PATH and never saw the stub dir — and ships `sleep` there as an exec-wrapper script (portable to Git Bash, pid-preserving). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * fix(#4223): gitignore gate covers the capture directory, and upgrades an existing file Round 4 Blocker. The gate enumerated image extensions, so snapshot.txt (the accessibility tree, with entered form values) and console.txt (which can carry tokens) were committable by `git add .`. The gate now ignores `interaction/` as a directory — the next artifact type is covered by construction — and it appends whatever an existing .gitignore lacks instead of writing once. The write-once form was the same defect one step later: every project that had already run an audit would never have received the new pattern at all. Tests run the gate fence under bash: a fresh file carries every pattern; an image-only file from an earlier audit gains interaction/ and keeps its own header without duplicating present lines; a second run appends nothing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj * test(#4223): declare the interaction-capture anchor as a comment marker The #4324 colon-token gate (slash-command-namespace) landed on next after this branch was opened and reads `<!-- gsd:ui-interaction-capture -->` as an unconvertible /gsd: command token. It is a section anchor of the same family as gsd:live-dom-families and gsd:write-continue, so it is declared in COMMENT_MARKER_TOKENS rather than renamed. Found by running the base-added gates against the merged tree; CI at ca8d2508 predates the gate. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LgNMb2G67rAJfFQRHEBTAj --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
eb49ff98df |
fix(#4728): stop presenting the retired Gemini CLI as a supported runtime (#4743)
* fix(#4728): stop presenting the retired Gemini CLI as a supported runtime
#1928 removed the Gemini CLI runtime after Google sunset it on 2026-06-18, and
updated the ENGLISH docs. The locale mirrors and the runtime-loaded workflow
prose were not updated in the same change, and no gate asserts the ABSENCE of a
retired runtime, so both drifted quietly for a year.
The finding that shaped this change: English is already correct. docs/
ARCHITECTURE.md, CONFIGURATION.md, USER-GUIDE.md, how-to/install-on-your-runtime.md
and CLI-TOOLS.md carry zero runtime-axis Gemini references; the only English hits
anywhere are a Gemini 2.5 Pro MODEL line, the GEMINI_API_KEY row, and prose that
correctly documents the retirement. So the docs half of this is translation lag,
not a content decision, and every locale edit here is parity with an existing
English line rather than new wording:
- install-on-your-runtime.md English has NO `### Gemini CLI` section -> deleted
- USER-GUIDE.md :843 "…, Antigravity CLI, Kilo)" -> substituted
- ARCHITECTURE.md English has NO Gemini CLI table row -> row deleted
- ARCHITECTURE.md :24 English holds `Kimi CLI` in that slot -> Kimi CLI
- context-monitor.md :3 "`AfterTool` for Antigravity CLI" -> substituted
- spike-and-sketch.md :93 "(Codex, Antigravity CLI, etc.)" -> substituted
- configure-model-profiles "Codex, OpenCode, Antigravity CLI, or Kilo" -> substituted
- COMMANDS.md English keeps only hyphen + Codex bullets -> colon bullet deleted
- FEATURES.md source docs/features/multi-runtime-support.md:10
lists no Gemini CLI -> name removed
ARCHITECTURE.md:24 is the clearest case for reading English rather than
substituting blind: Antigravity ALREADY appears later in that list, so replacing
Gemini CLI with Antigravity would have named it twice. English holds Kimi CLI
there, so that is what the locales get.
The largest single class was hand-duplicated boilerplate. A "Text mode" paragraph
repeated across 34 runtime-loaded workflow files ends "…required for non-Claude
runtimes (OpenAI Codex, Gemini CLI, etc.)". No lint enforces that sentence and no
script syncs it, so every copy was edited. These files are read by the agent at
runtime, so they steer behavior rather than only informing a reader — which is why
this class matters more than its word count suggests.
The slash-command-form section is restructured in all four languages to match
English, which had already dropped its colon-form bullet. That bullet claimed the
colon form is "Gemini CLI only", which was false on its own terms independent of
the retirement: `/gsd:…` is GSD's canonical AUTHORING token, rewritten per runtime
at install time, and NO runtime registers it — VALID_COMMAND_STYLES is
{slash-hyphen, shell-var} and 18 of 19 runtimes declare slash-hyphen. Substituting
the runtime name would have left the claim false with Antigravity's name in it, so
the claim is gone, matching English.
Two anchor regressions were caught and fixed while doing that. zh-CN lost its
explicit {#slash-command-forms-hyphen-vs-colon} anchor while its TOC still linked
it; the anchor is restored. ko-KR and pt-BR never had an explicit anchor and rely
on the slug generated from the heading text, so shortening the heading broke their
own TOC links; those links now point at the new slugs. English's heading lost its
anchor while its TOC still links the old one — that latent English bug is
deliberately NOT copied.
Preserved, because `gemini` is not one thing here and a blanket sweep breaks the
product: ~/.gemini/antigravity{,-ide,-cli} and ~/.gemini as their parent;
~/.gemini/config (#3738); GEMINI.md; hookEvents "gemini"; GEMINI_API_KEY in all
four locales; every gemini-* model id and the Gemini 2.5 Pro references in
ko-KR/pt-BR/zh-CN (ja-JP genuinely lacks that line — the locales have diverged, so
a uniform patch would be wrong); the hook-event dialect notes, which are
RE-ATTRIBUTED rather than deleted because Antigravity inherits that dialect;
reapply-patches.md:93's legacy-install note; host-integration-capability-matrix.md
:27 and :342, which correctly record the sunset and Antigravity's contract;
whats-new-1.7.0.md and FEATURES.md:3506, which document the retirement itself; and
the generated launcher preamble, which belongs to epic #4632 — zero
_GSD_SHIM_NAME lines appear in this diff.
Coverage: a #4728 block in tests/gemini-runtime-removed.test.cjs asserts the
retired name is gone from STRUCTURAL POSITIONS (a level-3 heading, a table row's
first cell, a runtime-example parenthetical) rather than asserting the string is
absent, which would be wrong. It pairs those with positive PRESERVE assertions
over the same files — Antigravity's heading, ~/.gemini/antigravity, GEMINI_API_KEY,
AfterTool — so a patch that deletes too much fails as loudly as one that deletes
too little. The model-axis test pins both the presence in three locales and the
absence in ja-JP, so a later uniform patch that "helpfully" adds it back fails.
The new docs/ reads tripped lint-docs-guard-registration for the first time in
this file, so the test is registered in scripts/docs-guard-registry.cjs.
Not covered here, by design: nothing above would catch a Gemini-as-runtime
reference appearing in a NEW file tomorrow. That is the repo-wide drift guard,
#4729, which must land last — written now it would red on the very references this
change removes.
Fixes #4728
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#4728): fix four review blockers, including a vacuous test and my own duplicate
A full matrix run on 31f12d7943 FAILED with 3 real failures, and an isolated
adversarial review returned BLOCK on four blockers. All of it was correct.
1. I committed the exact error I claimed to have avoided. The commit message
boasted that ARCHITECTURE.md:24 proved the value of reading English rather
than substituting blind, because Antigravity already appeared later in that
list. Five hundred lines further down the SAME four files, my
`Gemini:` -> `Antigravity:` substitution produced TWO consecutive
`- Antigravity:` bullets, because an Antigravity bullet was already there.
English (ARCHITECTURE.md:827) merges them into one. Now merged in all four
locales, reusing each locale's existing words.
2. `--gemini` survived in the runtime-detection CLI flag list in all four
locale ARCHITECTURE.md files. English:817 holds `--kimi` in that slot and
already lists `--antigravity` later, so this is another place where
substituting Antigravity would have duplicated it. Now `--kimi`.
3. Two runtime-loaded workflow files still enumerated Gemini one line ABOVE the
line I had already corrected -- the "Adaptive (Recommended)" option in
settings.md:192 and new-project/steps/auto-mode-config.md:95.
4. THE NEW TEST WAS VACUOUS for two of its five files. It matched only
`non-Claude runtimes (` and `(e.g. `, and neither regex could reach the two
lines the change actually fixed: health.md:52 reads `non-Claude (Codex, ...)`
without the word "runtimes", and execute-phase.md:1028 has no parenthetical
at all. The reviewer proved it by re-introducing Gemini at both lines and
watching the assertion stay GREEN. That same blind spot is what hid finding 3.
Replaced with a case-sensitive `/\bGemini\b/` walk over every
`gsd-core/workflows/**/*.md`, which works because every LEGITIMATE gemini
reference in that tree is spelled differently and cannot match: Antigravity's
paths are lowercase with a slash (`~/.gemini/antigravity`), Google's model ids
are lowercase and hyphenated (`gemini-3.1-pro-preview`), and the env vars are
uppercase (`GEMINI_CONFIG_DIR`, `GEMINI_SESSION_ID`). A bare capitalised
`Gemini` there means the retired RUNTIME is being named. The walk asserts it
found at least 50 files so an empty walk cannot pass vacuously, and it now
covers the nested `new-project/steps/` directory where finding 3 lived.
Two allowlist entries, both by line CONTENT and both justified:
reapply-patches.md's `Legacy: ... pre-#1928` note, and settings-advanced.md's
`Known provider` menu. The second was escalated by the agent rather than
decided: Section 8 of that file says model policy is defined "independently"
of the runtime, so `(Claude / OpenAI / Gemini / Qwen)` is the PROVIDER axis --
the same axis as the lowercase model ids -- and must keep working.
Proven to fail, not just asserted: the predicate reports 0 offenders on the
real tree and exactly 2 on a /tmp copy with Gemini re-injected at
health.md:52 and execute-phase.md:1028.
Also from the review: a `| Gemini |` COLUMN survived in the locale FEATURES.md
comparison tables (English has none) -- removed from all three, with header,
separator and every body row kept aligned; two ENGLISH runtime-axis sites were
missed by my own parity standard (how-to/execute-a-phase.md:88 and
how-to/verify-and-ship.md:89, the latter doubly stale since #4716 retired the
Gemini reviewer lane); docs/USER-GUIDE.md:12 linked a dead anchor, which I had
found and deliberately left -- record-and-proceed on a known defect is exactly
what the rules forbid, so it is fixed; docs/COMMANDS.md:12 and all four mirrors
still claimed "the hyphen and colon forms are runtime-specific spellings" with
no colon form documented anywhere, so that false sentence is deleted; and ko-KR
had the installer rather than the user doing the targeting.
The other two matrix failures were the compact-content benchmark baseline, which
drifted because this PR changes byte counts, refreshed via the script's own
`--write` path rather than by hand; and this commit's emitted-drift-ack trailers.
Method note on the acks: the failing run measured growth against
origin/next@1110c3b4ee, which is the STALE LOCAL `next` ref -- gsd-test merges
into the local base branch, and this machine's `next` is seven commits behind
origin/next, which is checked out in the main worktree and so cannot be
fast-forwarded from here. The 32 trailers below are computed against the REAL
base (origin/next @
|
||
|
|
4d65c248e5 |
fix(#4641): make test-conformance the sole Windows selector and narrow the tier to 28.5% (#4643)
* test(#4641): failing-first tests for the tier ceiling and a single Windows selector Tests only, committed ahead of the implementation so the RED run is real. - tests/platform-conformance-tier.test.cjs: tier-size ceiling asserted as a ratio against a live denominator (Windows 33%, macOS 25%); per-helper negative cases proving seam calls and path-call-plus-slash-literal are not platform signals; positive pins that genuine platform content, seam-bypassing spawns, chmod and symlink still classify in; macOS signal set and generated list unchanged. - tests/ci-full-lane-sharding.test.cjs: the test job has zero windows-latest rows and test-conformance still has 3 windows + 1 macOS. - tests/ci-test-scope.test.cjs: windows_tests is absent rather than empty, a non-tier test file no longer forces full_matrix, a RULE-pulled windows-hint test does, and resolveSelection rejects the retired windows scope. Refs #4589, #4591, #4592, #4593, #4603 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): delete the second Windows selector and narrow the conformance tier Epic #4589's goal — the OS-agnostic bulk on Linux, a small explicitly-scoped conformance tier on real Windows/macOS — was not met. Measured on PR #4640 (run 34618834118): 7 non-Linux jobs, a 546/930 (58.7%) "tier", and 5 of 7 changed test files running on a real Windows runner twice. Two selectors, only one in the epic's scope. The test job's three scope:windows shards predate the epic (#494, sharded #3057) and gate on product_changed, not full_matrix, so they fire on every product PR whatever Phase 3's classifier decides. They are deleted; test-conformance becomes the sole Windows selector, as it already was for macOS. Non-Linux jobs 7 -> 4. Gating the lane instead was rejected as provably redundant: for a test file reachesConformanceTierOrSeam is literally CONFORMANCE_TIER_FILES.includes(file), and that same predicate sets full_matrix, which turns test-conformance on. Every file a gated lane would run is already covered in the same run. The lane's one non-redundant residue -- RULE-pulled tests matched by the isWindowsHint filename heuristic -- is ported into reachesConformanceTierOrSeam so it sets full_matrix instead of feeding a parallel lane. Two detectors matched the repo's own test idiom rather than any platform signal and carried 226 of the tier's sole-signal membership against 41 for the other eight: process-seam-subprocess (335 files, 118 unique) matches the tests/helpers.cjs entry points nearly every CLI test uses, and going through the seam is the opposite of a platform signal since shell-command-projection takes platform as an injected parameter; hardcoded-path-vs-path-call (328, 108) needs only a path call anywhere plus a slash literal anywhere, and that class is already enforced by ADR-1703's Linux-runnable ESLint rules. Both are removed. Tier 546 -> 254 (27.3%). src/ reachability is unchanged at 28 files, measured. Adds the size gate Phase 2 never had, as a ratio against a live denominator so it cannot stop binding as the suite grows. 292 files leave real-OS Windows execution. The drop-out set was audited: 14 have a platform-suggestive filename and all 14 are static source-text analyses or seam-mediated CLI tests. raw-child-process was investigated as a suspected false negative and left unchanged -- relaxing it adds 13 files, all false positives. macOS is untouched: MACOS_CATEGORIES is a separate array and the regenerated macos-conformance-tier.generated.cjs is byte-identical at 196 files. Fixes #4641 Refs #4589, #4591, #4592, #4593, #4603 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): register the new ADR path in the docs-guard exempt baseline tests/ci-test-scope.test.cjs references docs/adr/4641-windows-selector-consolidation.md in a comment justifying the retired windows scope; lint-docs-guard-registration tracks that reference set, so the baseline needs the new path. Verified the exemption still holds: the path is prose, not a filesystem read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): make the escalation tier-backed and drop every hardcoded count Three follow-ups from measuring the first pass rather than trusting it. The windows-hint escalation now requires tier membership as well as the filename hint. Setting full_matrix runs test-conformance, which runs only the tier; escalating on a test that is NOT in the tier costs four jobs and still never runs that test on Windows. Measured over the 16 RULES entries the narrowed predicate fires on exactly the same rules today, so this is correct-by-construction rather than a behavior change. The broader variant -- escalate on any tier member a rule pulls in, ignoring the hint -- was measured at 14/16 rules and rejected as over-broad. Removes the hardcoded counts. A hardcoded macOS tier length of 196 broke as soon as the rebase pulled in one new test file from #4253, which is the whole argument against them: the ceilings are ratios against a live denominator, the committed lists are pinned by comparison against a fresh classification of the live tree, and the three named probe files now assert on their SIGNAL rather than on membership in a literal list -- asserting by filename is the exact error this PR fixes in the classifier. Regenerates both lists against the rebased tree. Same-tree figures are now 547 -> 255 of 931 eligible (58.8% -> 27.4%), 292 entries removed and none added; macOS is unchanged at 197 with a zero-line diff. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): restore real-shell-spawn coverage and repair assertions the narrowing broke An isolated adversarial review found a real false negative. Removing the blanket process-seam-subprocess detector also removed the only coverage for tests that spawn a REAL shell: tests/helpers/process-seam.cjs's runHook spawns options.interpreter via real spawnSync, so runHook('-c', [script], { interpreter: 'bash' }) runs a real bash binary executing a shell script extracted from workflow markdown. The seam argument holds for src/shell-command-projection.cts, which takes platform as an injected parameter; it does NOT hold for the test helpers, which spawn real binaries. Conflating the two is what made the blanket detector look purely noisy -- it was 99% noise wrapping a real signal. Adds a narrow shell-interpreter-spawn category keyed on a real interpreter option. Measured 2026-09-11: 33 files match, 9 were outside the tier and are added back, taking it 255 -> 264 of 931 (27.4% -> 28.4%), still under the 33% ceiling. All 9 confirmed by reading the matching source line, zero comment or fixture matches. runGit-alone and non-node-spawnSeam alternatives were measured and rejected -- each adds 9 files but misses the counterexample entirely. Fixes a real bug the suite caught: jobs.test is ubuntu-only now that its scope:windows rows are gone, so it must wire GSD_STRICT_LIVE_CONFIG_GUARD strictly rather than carrying the Windows report-only carve-out. The carve-out now lives solely on test-conformance, whose matrix does include windows. Repairs seven pre-existing assertions the category removal invalidated, preserving each case's purpose rather than deleting coverage, and converts the last hardcoded tier bounds to live-derived ratios -- including the macOS sanity range that was still a magic [100, 350]. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): keep the confinement test on a real OS via a documented allowlist A security review found tests/external-descriptor-confinement.test.cjs had dropped out of the Windows tier. It must stay in, and no content signal can express why: it exercises isPathConfined (src/external-descriptor-trust.cts), which uses the AMBIENT path module -- path.resolve(root, target) and path.sep -- with no injection. Its win32 semantics (drive letters, UNC, separator) are only reachable by actually running on Windows, and it is a security-relevant write-confinement gate. A content classifier cannot see 'this module reads the ambient path module', so no regex belongs here. Adds ALWAYS_REAL_OS, a Map of path -> recorded reason, unioned into the Windows tier only. A Map rather than a list so an entry without a reason is impossible by construction, and tests assert every entry names a file that exists on disk so a stale entry fails loudly instead of rotting. This is the centrally- enumerated single source of truth epic #4589 Phase 2 asked for and ADR-1703's portability-vocab.cjs already models -- deliberately not a heuristic. Windows tier 264 -> 265 of 931 (28.5%), still under the 33% ceiling. macOS is untouched and byte-identical: the win32 concern does not apply to a POSIX runner, and a test asserts the allowlist does not leak into that tier. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4641): inject the path impl into isPathConfined and correct the ADR count Two review findings, both fixed rather than dispositioned. A security review found tests/external-descriptor-confinement.test.cjs had left real-OS execution. The allowlist pinned it back, but that only restored INCIDENTAL coverage: isPathConfined used the ambient path module, and its test carried POSIX-only literals, so a win32 confinement escape was unverified on every platform including Windows. isPathConfined now takes an optional third parameter carrying the path implementation, defaulting to the ambient module. Blast radius is CRITICAL -- 53 affected symbols across 19 files -- so the change is purely additive and every existing two-argument caller is byte-identical. Tests now inject path.win32 and path.posix, covering a different drive letter, a cross-drive absolute, backslash and forward-slash traversal, UNC, and the startsWith prefix-boundary bug (.gsdEVIL against root .gsd) on both separators. Proved load-bearing: dropping the + p.sep from the prefix check fails exactly the two boundary cases and nothing else. Callers' suites 149/149. The spec review caught an off-by-one: the ADR narrated a 264-file tier while the committed list holds 265. The ADR now records the full chain 547 -> 255 -> 264 -> 265 (28.5%). Also corrects a stale comment in scripts/docs-guard-registry.cjs that narrated classify() as zeroing windows_tests, a key this change removes -- kept as historical narration but labelled as such. Refs #4641 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#131): make the unwritable-HOME test actually test something Found by sweeping for the root-bypass class after fixing commit-files-deletion. This one is the silent variant, and it was broken twice over. First, the condition: the test made a fake HOME unwritable with chmod 0o500. The gsd-test Docker bench runs as root, root bypasses mode bits, so HOME stayed writable and the hostile condition never existed. Replaced with a HOME whose PARENT is a regular file, so every write under it fails ENOTDIR at the VFS layer for every uid -- no permission check is involved at all. Second, and more fundamental: the probe was npm --version, which on npm 11.19.0 performs zero filesystem I/O against HOME. Proven rather than assumed -- neutralizing runNpm()'s isolation turned the sibling test red while this one stayed green, so its assertion could never detect the regression it guards, on any uid, with or without the condition fix. npm config get cache was tried next and proved vacuous the same way (it only string-resolves the path). The probe is now npm cache verify, which really does mkdir _cacache under HOME. Re-proved load-bearing after the change: with isolation neutralized the test now fails with ENOTDIR on <blocker>/home/.npm/_cacache. tests/helpers.cjs was restored and verified diff-clean; suite 13/13. Refs #4641 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): correct the net drop-out figure in ADR-4641 The Consequences section still said 292 files leave real-OS Windows execution. That was the count before the narrow shell-interpreter-spawn replacement restored 9 and ALWAYS_REAL_OS pinned 1. Net is 282. Also names both real-binary categories rather than only raw-child-process, and clarifies that the 14-file filename audit was against the 292 initially dropped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record the rejected concentration ceiling and its measurement Applying Goodhart's own question to the new ceiling -- how would you make this metric look good without improving what it represents -- surfaces a real weakness: a ratio can be satisfied by inflating the denominator, so adding OS-agnostic tests loosens it without narrowing the tier. The obvious companion gate was a sole-signal concentration ceiling, since the original defect was one detector carrying half the tier. Measured and rejected: peak concentration post-fix is raw-child-process at 53/265 = 20.0%, against the historic offenders at 21.6% and 19.8%. Any threshold above 20% misses the original defect; any threshold below it fails on a legitimate category. The discriminator is whether a signal is platform-meaningful, which no threshold encodes. Weakness disclosed rather than covered by a gate that does not bind. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4641): add the changeset fragment for the confinement-check change changeset-lint failed on PR #4643: the PR touches user-facing paths and carried no fragment. The earlier no-changeset call matched #4604's CI-only precedent and was correct then; it was not revisited once the PR grew a src/ change, which is my miss. The fragment describes the real user-visible improvement: the external-descriptor write-confinement check's Windows semantics are now verified deterministically rather than only when the suite happened to run on Windows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): correct the tier count in TESTING-SUITES.md Said the tier narrowed from 546 to 254. The final committed list is 265 of 931 eligible (58.8% -> 28.5%) after the shell-interpreter-spawn replacement restored 9 files and ALWAYS_REAL_OS pinned 1. Same error class the spec review caught in the ADR, in a live reference page rather than a dated record, so it states the current truth rather than carrying an amendment note. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record the measured aggregate from real CI job lists Epic #4589's closeout asserted its reduction from a static count; #4641's acceptance criterion asks for a figure read off a real run. Recorded here: test.yml job count 21 -> 15 and non-Linux 7 -> 4, comparing PR #4640's run against this PR's own. Against the true pre-epic baseline of 9, that is 9 -> 4. Also states the caveat that a PR's total CHECK count is not a clean before/after comparison, since many gates are path-scoped and this change touches a broader path set -- the like-for-like figure is the test.yml job count. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): compare job totals the same way on both sides The measured-aggregate table put #4640's COMPLETED run total (21) against this run's count at matrix-expansion time (15). Those are not the same measurement: the completed total includes the post-test Coverage gate and baseline-publisher jobs. Counted identically, it is 21 -> 17. The load-bearing figure, non-Linux jobs 7 -> 4, was correct and is unchanged. Called out in the table rather than silently corrected -- comparing two differently-derived numbers is exactly the error class this ADR is about. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): record measured conformance wall-clock and date the stale counterfactual Adds the per-job durations from both runs. The honest read is that this is a correctness win more than a speed one: file count fell 52% but wall-clock only 9-29%, because what was removed were the cheap static tests and what remains is concentrated in expensive spawn-heavy work. Stated explicitly so nobody expects a future narrowing to buy time proportional to file count. The load-bearing figure is windows shard 3/3: 40m24s against a 45-minute cap on the 547-file tier -- 90% of the cliff #869 and #3057 were both filed about -- pulled back to 31m27s. macOS moved the wrong way (17m48s -> 21m02s) while its tier was UNCHANGED at 197 files, which fixes that as runner variance and is noted as a caution against reading a single duration as signal. Also dates the symlink-keyword counterfactual, which cited a 254-file tier from before the replacement category and allowlist took it to its final 265. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#4641): re-measure against the rebased tree and disclose the allowlist's zero next gained #4644 mid-flight, so every absolute count shifted. Re-measured on the tree this actually ships against (932 eligible): 548 -> 257 by detector removal, 257 -> 266 once shell-interpreter-spawn restores 9. Net 282 removed, 9 restored. macOS 198, unchanged by this PR. The percentages did not move across three rebases (58.8% -> 28.5%), which is the whole argument for expressing the ceilings as ratios rather than counts -- noted in the ADR since it is now evidence rather than assertion. Also discloses that ALWAYS_REAL_OS now contributes ZERO files: this PR's own win32 test cases introduced the literal win32 into the pinned file, so it classifies in on content via win32-darwin-literal. The entry stays and the reason is written down, because the file's real-OS need is a property of the code under test (isPathConfined reads the ambient path module), not of the test's text -- the text that currently saves it is incidental and could be refactored away silently. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e54d3aa159 |
enhance(#4401): register workflow.compact_content as a validated config key (#4441)
* feat(#4401): register workflow.compact_content as a validated config key - Add compact_content: false to the nested workflow object in gsd-core/bin/shared/config-defaults.manifest.json - Add 'workflow.compact_content': false to SCHEMA_DEFAULTS in src/config.cts so an absent key resolves to false via config-get --raw - validKeys entry in config-schema.manifest.json already present Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(#4401): behavioral and boundary tests for workflow.compact_content - 19 behavioral tests covering config-set/config-get round trip, invalid-shape rejection (banana, 42, empty string), the corrected null-unset semantics (#2046), absent-key resolution against config-defaults.manifest.json, config-new-project wiring, and doc-row shape assertions - Drops the install-tree fixture-parity block (and its docstring item) that asserted gsd-core/references/compact-content-gate.md and gsd-core/workflows/compact/map-codebase.md fixture entries — those paths belong to #4402 and do not exist on this filtered branch Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(#4401): document workflow.compact_content in both config references - One 4-cell row in docs/CONFIGURATION.md (workflow.* run) - One 5-cell row under Workflow Fields in gsd-core/references/planning-config.md - Both cross-reference ADR-4139 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(#4401): add changeset - Added-type fragment, pr: 4401 (issue number; backfill to the real PR number is a required follow-up once the PR is opened, per D-08 and CHANGESET-PR- FIELD-DRIFT) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(#4401): backfill changeset pr field to #4441 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(#4401): derive workflow.compact_content default from CONFIG_DEFAULTS SCHEMA_DEFAULTS['workflow.compact_content'] hardcoded the literal false instead of deriving it from CONFIG_DEFAULTS the way 3 of its 8 sibling entries do (smart_zone_tokens, pr_strict, inline_plan_threshold), leaving a single-source-of-truth drift risk: a future manifest-only edit to the default could silently diverge from this literal, only caught later by the D-03 test if it ever happened to manifest. Adds compact_content to CONFIG_DEFAULTS in src/config-loader.cts and derives SCHEMA_DEFAULTS from it in src/config.cts, matching the majority sibling pattern. Found during maintainer review (review-open-prs) of this PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4401): map compact_content in config-field-docs NAMESPACE_MAP The previous commit added compact_content to CONFIG_DEFAULTS in src/config-loader.cts but missed the matching entry in tests/config-field-docs.test.cjs's NAMESPACE_MAP, which maps flat CONFIG_DEFAULTS keys to their namespaced doc form before checking gsd-core/references/planning-config.md for a match. Without it, the test looked for a bare `compact_content` doc reference instead of the actual `workflow.compact_content` row, and failed: "CONFIG_DEFAULTS keys missing from planning-config.md: compact_content". Found by actually running gsd-test against the branch rather than trusting the plausible-looking fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4401): register compact-content-4139 test in the docs-guard lane tests/compact-content-4139.test.cjs's D-06 tests read docs/CONFIGURATION.md directly (fs.readFileSync) to assert the workflow.compact_content doc row's shape, which makes it a doc-reading test file under the #3753 docs-guard lane. It was never added to scripts/docs-guard-registry.cjs's DOCS_GUARD_TESTS map and carries no docs-guard-exempt marker, so tests/ci-docs-guard-registry.test.cjs's registration lint correctly failed: "compact-content-4139.test.cjs reads a docs/ path but is not registered in the docs-guard lane and carries no docs-guard-exempt marker". Registers it with ['docs/CONFIGURATION.md'] (the only real docs/-prefixed path it reads; gsd-core/references/planning-config.md is outside this registry's docs/ scope, matching the sibling config-field-docs.test.cjs entry's existing convention). Found by actually running gsd-test against the branch — this gap predates the maintainer's config-loader.cts fix and was already present in the original PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: sim <sim@local> |
||
|
|
0fca71eaae |
enhance(#2529): cover every workflow with response-language directives + CI lint (#2558)
* enhance(#2529): cover every workflow with response-language directives + CI lint Every workflow now carries response-language coverage in one of three forms, and a CI lint keeps it that way. - 43 workflows load the new shared reference, `gsd-core/references/response-language-directive.md`, by eager `@`-import. - Lazy-loaded modes/steps/templates, which cannot rely on an eager import, carry an exact inline directive; 35 such paths are pinned by exact path. - Fragments dispatched by a covered parent inherit coverage, proven per file rather than granted per directory. The 45 workflows whose directive covered only "questions, prompts, and explanations" now name inter-tool narration, which is the defect #2529 reports: the running commentary between tool calls stayed English while the answers around it were translated. `scripts/lint-response-language-coverage.cjs` enforces it and fails closed on three independent discovery failures (unreadable catalog, empty catalog, unfollowed symlink). It resolves which reference a workflow imports and applies the same four-predicate test to that file, so a weakened shared reference uncovers its importers instead of passing silently, reported once as a systemic failure rather than 43 times. The walk follows symlinked subtrees with a realpath cycle bound. `lint:ci` invokes it by name. REQ-LANG-03 and REQ-LANG-04 state the contract in docs/FEATURES.md; REQ-LANG-04 names the two forms that satisfy it ("narration", "between tool calls") rather than enumerating class members an author cannot use verbatim, and a test pins that text to what the matcher accepts. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): register the coverage test in the docs-guard lane `107eb8c1` (#3787) landed the docs-guard lane on `next` while this PR was open: a test that reads a `docs/` path must be named in `scripts/docs-guard-registry.cjs` or carry a `docs-guard-exempt` marker, so the guards that read a doc run on the PR that changes it. `tests/response-language-coverage.test.cjs` reads `docs/FEATURES.md` -- it extracts every form REQ-LANG-04 offers an author and runs each through the matcher that enforces it. Registration, not exemption, is the correct side of that gate: a reword of the requirement with no code change is precisely the diff this test exists to catch, and it is the diff the lane would otherwise skip. Registered narrowly (`['docs/FEATURES.md']`) rather than with the `'*'` sentinel, so an unrelated docs change does not pull this test into the lane. Verified: lint-docs-guard-registration 0 violations, tests/ci-docs-guard-registry.test.cjs 51/51, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): consolidate this PR's emitted-growth acks into its own fragment This PR ripples emitted bytes across 85 workflow paths. Until now each ripple was acknowledged by appending to whichever live fragment owned that path, because two ack sources may never name the same path. `a84f7563` (#3078) swept all 45 fully-spent fragments off `next`. Forty-two of the paths this PR grows were owned by swept fragments, so those keys are now unowned and this PR's own fragment declares them directly -- one path, one source, and no dependence on a fragment that no longer exists. Each adopted entry keeps its measurement and records where it came from. Two paths are handled differently, because the sweep did not free them: - `review.md` is now owned by `3034-parallel-reviewer-lanes.json`, which landed on `next` after the sweep. Its entry is live, so the old route still applies: this PR's note is appended to that entry rather than declared a second time. - `plan-review-convergence.md` keeps the arrangement made in round 24. Result: 3 fragments in the directory, 85 keys in this PR's own, 0 cross-source duplicates. `lint-emitted-drift-ack` exit 0, `tests/emitted-attribution.test.cjs` green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): move REQ-LANG-03/04 into the feature fragment that now generates them `36375513` (#3845) made docs/FEATURES.md a generated projection of docs/features/*.md, marked "do not edit by hand". This PR wrote REQ-LANG-03 and REQ-LANG-04 straight into the generated file, so the rebase left the requirement present in the projection and absent from its source -- the next regeneration would have deleted both, and `tests/features-index-gate.test.cjs` was already red on the mismatch. Both requirements now live in docs/features/response-language-config.md alongside REQ-LANG-01 and -02. Regenerating produces a docs/FEATURES.md that is byte-identical to the committed one, so the text this PR shipped is unchanged -- only its source of truth moved to where #3840 put it. The docs-guard registration is widened to name the fragment as well as the projection. The requirement's source is the fragment now, and an edit there that skips regeneration would otherwise reach this guard through neither path. Verified: features-index-gate 68/68, lint-docs-guard-registration 0 violations, ci-docs-guard-registry + response-language-coverage 142/142. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): hand the plan-phase ack back to its new live owner `c933184b` (#3825) landed `3172-stated-failing-direction.json` on `next` after fragment had adopted that path when the sweep left it unowned, so the merged tree named it from two sources -- a hard failure in `scripts/lint-emitted-drift-ack.cjs`. The path has a live owner again, so the append route applies: this PR's note joins that entry, carrying its own measurement, and the key is dropped from this PR's fragment (84 keys left, the others untouched). The provenance sentence written for the swept-fragment case is removed rather than reused -- this path was never orphaned, so that account of it would be false. Same shape as `review.md` and `plan-review-convergence.md`: ownership is a property of the merged tree, and a fragment landing upstream after a push can reclaim a key no local check would have flagged. Verified: lint-emitted-drift-ack exit 0, lint:ci exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): state byte figures that are true against the tree The reference claimed `execute-phase.md` has "2 bytes of headroom under the ceiling named below". That was true when the sentence was written -- the file sat at 93398 against the 93400 comfort assert -- and upstream has since shrunk it to 91493 against a 93600 hard ceiling, so the figure now understates the headroom by three orders of magnitude. The rationale the sentence supports does not depend on the number, so the number is gone rather than refreshed: a restated figure would go stale again on the next upstream edit, and nothing parses it. Audited every other numeric claim this PR ships the same way, mechanically against the merge base: all 82 FILE-delta claims in the ack fragment match the real per-file delta exactly, and the 1,629-byte reference and 63-byte import line check out. One class was imprecise: the 41 notes for workflows whose inline directive was rewritten in place quoted the conversion counterfactual as "+1,692 bytes more loaded context", which is the reference form's whole weight, not the increase over the inline directive those files already carry. Each now names both quantities and the net (+1,605 / +1,609 / +1,584). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): one rule for pinned vs inherited coverage, and the docs to pick it Review measured that 14 of the 35 pinned fragments would pass by inheritance anyway, and that the PR asserted both readings at once: inheritance is real coverage (so those 14 pins are noise) or it is not (so 30 inheriting fragments are green-but-uncovered). Only one can be true. Inheritance is real: the predicate proves it per file -- the parent must dispatch this exact path from a read/execute context AND be covered itself -- so the parent's directive is in the loaded context by the time the fragment is read. The 14 pins are therefore removed along with the directive lines they pinned, and those files inherit like the 30 structurally identical ones. The rule is now stated where the set is declared, and enforced from the other side by a test: no member of the pinned set may be one that would have inherited. That is what decides the form for the next fragment. - pinned set 35 -> 21; 14 workflow files revert to their base content - `findViolations` no longer returns early on a pinned path: a file that becomes eagerly loaded and takes the shared reference is strictly better off, and the gate must not red that. The reference form is admitted because its own wording is validated in turn; an arbitrary reworded inline line still fails. - the reference-directive cache is keyed by size and mtime, not by path alone, so a rewritten reference re-asked in one process no longer returns the stale verdict - `carriesInlineDirective` names its negation blindness: four independent hits read vocabulary, not polarity - the real-tree scan asserts each source produced files instead of `> 152`, a constant that read as the workflow count and would have passed a scan that lost one of its two directories - the pinned-set size assertion goes the same way: the size follows from the rule, so the rule is what the suite asserts Docs, for the gate that now governs every future workflow: - `docs/contributing/response-language-coverage.md` -- why the narration class is the discriminator, the four coverage forms, the decision order that picks one, the pinned line, and what each failure message means - a row in CONTRIBUTING.md's CI checks table, matching the docs-guard row - `docs/CONFIGURATION.md` points at it from the `response_language` entry Also: the changeset said 45 reworded workflows; it is 44 (42 @-reference + 21 pinned + 44 rewritten = 107 touched). That text ships to CHANGELOG.md. `3707-parse-gap-reporting.json` landed on `next` reclaiming `audit-uat.md` and `progress.md`; both handed back by the append route, leaving 82 keys here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2529): correct the reference-taker count, 43 -> 42 The ack notes said the import line is byte-identical "in each of the 43 workflows that take the reference" and that the alternative would be "43 inline copies". The shared reference has 42 importers; the 43rd file in review's table is `execute-phase.md`, which imports the OTHER reference. Corrected in all 41 notes that carry the sentence, across this PR's fragment and the two it appends to. Found by re-running the numeric audit from the previous round after the rebase, which also re-verified all 84 FILE-delta claims against the new base -- all exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2529): migrate the emitted-drift ack from a fragment to commit trailers ADR-3942 (#3954) landed while this PR was open: the acknowledgment is now a commit trailer and tests/emitted-drift-acks/ no longer exists. The fragment is deleted and each key it declared becomes one trailer, reasons unchanged. The four keys this PR had handed to 3034-*, 3172-* and 3707-* under the one-source rule come home here. That rule was the whole reason for the hand-backs, and the trailer model has no shared namespace to collide in -- five of this PR's rounds were spent on exactly those collisions. Emitted-Drift-Ack-Growth: add-backlog.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-phase.md — #2529 — RESTATED in round 10, superseding this PR's earlier "+63 bytes, prose only" wording, which reported a file delta as if it were the whole cost. The workflow gains the shared response-language directive as a single `@`-reference line. FILE delta: +63 bytes, byte-identical in each of the 42 workflows that take the reference. LOADED-CONTEXT delta: +1,692 bytes per workflow — the 63-byte import line plus the 1,629 bytes of `gsd-core/references/response-language-directive.md`, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"). The repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so they see 63 of those 1,692 bytes; the remaining 1,629 are declared here because no gate reads them. The eager import is accepted on its merits, not hidden: 42 inline copies would be 43 places for the wording to drift, and the reference is the one place it is maintained. Prose only: no step, gate, tool invocation, or subagent dispatch shape changed. Emitted-Drift-Ack-Growth: add-tests.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: add-todo.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Emitted-Drift-Ack-Growth: ai-integration-phase.md — #2529 MAJOR 1 (round 10): the workflow's pre-existing inline response-language directive is rewritten IN PLACE so the sentence names inter-tool NARRATION explicitly — narration between tool calls, status updates, progress notes, findings — instead of only "questions, prompts, and explanations". That older wording is the defect #2529 reports (it leaves the running commentary between tool calls in English while the answers around it are translated), and `scripts/lint-response-language-coverage.cjs` had been certifying it as coverage, so the gate legitimised the bug. +87 bytes, prose only: no step, gate, tool invocation, or subagent dispatch shape changed. FILE delta and LOADED-CONTEXT delta are both +87 here, and that identity is the point — the directive was deliberately NOT converted to an `@`-reference, because an `@`-import in this repo is EAGER (ADR-1610 Decision point 4; docs/ARCHITECTURE.md: moving prose into a file that is still eagerly `@`-imported "shrinks the measured file without shrinking loaded context"), so the conversion would have bought a smaller measured file at a cost of 1,692 bytes of loaded context per workflow — the 63-byte import line plus the 1,629-byte reference — against the 87 bytes this inline directive costs, a net +1,605. Stated plainly because the gates cannot state it: the repo's size gates — the tier hard caps in `tests/workflow-size-budget.test.cjs` and this size ratchet — measure the FILE, not the transitive inline, so an `@`-reference conversion would have READ as a smaller change to every gate in the repo while costing 1,605 bytes more loaded context per invocation. Re-homed in round 16: the fragment that carried this sentence (`3423-required-reading.json`) was retired on `next` by |
||
|
|
472f585f7c |
fix(#3726)!: require --confirm before milestone complete mutates (#3774)
* fix(#3726): require --confirm before milestone complete mutates `milestone complete <version>` is a one-way door — ROADMAP.md and REQUIREMENTS.md archived, every phase directory in the milestone MOVED, STATE.md rewritten — and ran unconditionally on first invocation through every invocation path, including `query milestone.complete <version>`, whose `query` meta-prefix reads as a read-only namespace but performs no filtering (#167's invocation-compatibility shim + #3243's dotted-form normalization). The gate lives on the destructive command itself, not on the `query` prefix (the prefix is an intentional invocation mechanism, not a permission boundary — restricting it would break dozens of shipped workflow callers). Without --confirm and without --dry-run the command now refuses via error() before reading anything beyond its arg checks, so an unconfirmed invocation is a guaranteed no-op on disk. --dry-run still previews with no confirmation needed and is now documented in the usage block (it was only documented for the sibling archive-quick). --force keeps its narrow meaning — bypassing the TRUNCATED-scope and unstarted-phase guards — and does not double as the mutation opt-in. --confirm follows the existing `phases clear --confirm` idiom in the same module. complete-milestone.md's two invocations pass --confirm (the workflow has gathered explicit user intent by that step). Existing tests get --confirm appended — pre-change behavior is exactly confirmed behavior — and a #3726 regression block covers: refusal + full-tree byte-identity on both invocation forms, --force not satisfying the gate, --dry-run still passing without confirmation, and --confirm proceeding. The refusal tests fail against pre-fix code (negative control run). Fixes #3726 * docs(#3726): document the --confirm requirement in CLI-TOOLS and COMMANDS Cross-AI review of the fix diff (codex, pre-create) caught three shipped doc sites still instructing the now-refused bare invocation: the CLI-TOOLS.md milestone-complete synopsis + flag table, and COMMANDS.md's two guard-override instructions (`--force` alone now refuses without --confirm). Localized CLI-TOOLS copies already lag the English synopsis (no --force/--dry-run either) and follow the translation pipeline, not this fix. * chore(#3726): set changeset fragment pr to 3774 * test(#3726): confirm-gate CI repairs — QA scenario caller + growth ack Two CI reds from the --confirm gate, both this branch's own misses: - tests/qa/scenarios/milestone-rollover.json invoked `milestone complete 1.0 --force` as a JSON arg-array fixture — a caller shape the test sweep (which grepped runGsdTools/runSdkQuery in tests/*.cjs) never enumerated. Adds --confirm; the scenario's boundary-crossing contract is otherwise untouched. - complete-milestone.md's +420-byte --confirm note trips the emitted-attribution growth ratchet. Acknowledged as a #3726 append to the existing complete-milestone.md entry in 3409-unreachable-guard-arms.json (two ack sources may never name the same path, per that fragment's own precedent). Local: lint-emitted-drift-ack ok; loop-walk.qa 115/115 green sandboxed. * docs(#3726): CLI-TOOLS.md guard-override sentences say --force --confirm Review Major 1: the truncated-window and unstarted-phase guard paragraphs still told the reader to "Pass `--force` to override", which now refuses (--force alone does not satisfy the confirmation gate), while the flag table 470 lines later said the opposite. Mirror the docs/COMMANDS.md pair so the file no longer contradicts itself. * docs(#3726): synopsis renders --confirm and --dry-run as alternatives Review Nit 1: `milestone complete <version> --confirm [--dry-run]` read as "a dry run still needs --confirm", the opposite of AC 3. Render the pair as `(--confirm | --dry-run)` in the CLI-TOOLS.md synopsis and the usage docblock, and let the flag rows carry the rule. * test(#3726): pass --confirm in base-added milestone fixtures; re-file the growth ack Rebase onto next (26 commits) surfaced three tests the gate now refuses: the #3685 write-flag contract pair in tests/milestone.test.cjs and the `milestone complete` boundary fixture in tests/state-contract.test.cjs all invoke the command bare. Each now passes --confirm (a mutating run is exactly what they assert on). The +420 byte complete-milestone.md growth ack rode on 3409-unreachable-guard-arms.json, which #3078 swept from next as fully spent — hence the modify/delete conflict. Re-filed under a fresh fragment named for this issue, never resurrecting the swept one. * test(#3726): pin the present-but-falsy arm of the confirmation gate Review Minor 1: the boundary triple covered absent and present but not present-but-falsy. The gate is an exact-token match, so --confirm=false and --confirm=0 refuse today — pinned (canonical + query forms, whole .planning/ tree byte-identical) so a future `=`-aware or prefix-matching parser cannot silently turn --confirm=false into a confirmed run of an irreversible command. * test(#3726): drop --confirm from dry-run-only invocations Review Nit 2: --confirm was mass-appended to 14 pre-existing --dry-run invocations that never needed it, so each stopped standing as incidental proof that a preview needs no confirmation. Reverted to the pre-PR form; the dedicated AC-3 test carries the explicit assertion. * docs(#3726): sync the localized CLI-TOOLS synopsis with the confirm gate REQ-I18N-02 (docs/features/internationalized-documentation.md) requires translations to stay synchronized with the English source. The four localized CLI-TOOLS.md guides still advertised a bare `milestone complete <version>`, which now exits 1. Render the English synopsis verbatim — `(--confirm | --dry-run)` plus the `[--force]` and `[--archive-quick]` flags the translations had also fallen behind on. * test(#3726): drop --confirm from the remaining preview-only invocations Round 2 reverted the --confirm appends on --dry-run-only invocations in tests/milestone.test.cjs, but four more sat in two files the sweep missed: tests/milestone-archive.test.cjs (three) and tests/milestone-window-single-owner.test.cjs (one). Each is a preview run whose whole purpose is to document that a preview mutates nothing, so `--dry-run ... --confirm` contradicted the semantics the test exists to pin. Dropping the token restores each as incidental proof that a preview needs no confirmation; the dedicated AC-3 test keeps the explicit assertion. No assertion added, relaxed, or removed — the change is four tokens. * chore(#3726): migrate the emitted-drift ack from a fragment to a commit trailer #3954 (ADR-3942) moved emitted-drift acknowledgments out of tests/emitted-drift-acks/ and into git commit trailers, and the fragment directory no longer exists on next. The reason this PR's fragment carried moves verbatim into the Emitted-Drift-Ack-Growth trailer on this commit; the fragment file is removed rather than resurrected. Emitted-Drift-Ack-Growth: complete-milestone.md — #3726: +420 bytes (40186 -> 40606). The archive_milestone step's two `milestone complete` invocations now pass the required --confirm flag (the command refuses to mutate without it — the archive is irreversible), with a note explaining the flag and pointing at --dry-run for previews. Deliberate runtime-loaded workflow text for the new gate, not converter drift. * fix(#3726): name --confirm in the version-required refusal The documented arg-discovery path (gsd-tools.cjs top-level usage: invoke the command without args and the error lists what is required) stopped at `version required for milestone complete (e.g., v1.0)` — one required argument short. Discovering --confirm took a second round trip through the gate. The refusal now reads `… — and --confirm to mutate`, pinned by a test that also asserts the version-less invocation leaves .planning/ untouched. * test(#3726): pin the milestone complete docs against a silent regression The changeset is `type: Fixed`, which the docs-required lint exempts, so nothing in CI would notice a later edit that reinstated the bare-`--force` override prose or dropped `--confirm` from the synopsis. Four tests in tests/milestone.test.cjs now pin: the synopsis line in docs/CLI-TOOLS.md and its four localized mirrors; the `--confirm` flag row; both guard-override instructions in docs/CLI-TOOLS.md and docs/COMMANDS.md, by guard name (a substring match on each instruction's `--force --confirm` text); and — as an identity ratchet over the milestone-complete sections — every `--force` sentence or clause that lacks `--confirm`, so a new bare instruction in its own sentence or clause fails whatever its wording. Named residual: a bare instruction spliced into the same clause as a compliant one coalesces with it and passes the ratchet; the by-name pins are what keep the four known instructions from losing the pairing that way. The file is registered in scripts/docs-guard-registry.cjs so the pin runs on the PR that changes those docs, not only after merge. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
519ac23ebb |
fix(#3839): hook tables say PreToolUse (validate-commit) and SessionStart (session-state) (#4041)
* test(#3839): docs hook tables must match surface registrations (failing first) * docs(#3839): hook tables say PreToolUse for validate-commit, SessionStart for session-state gsd-validate-commit.sh is registered PreToolUse (src/runtime-hooks-surface.cts; its exit-2 block IS the contract — a post-tool hook cannot prevent a commit) and gsd-session-state.sh is registered SessionStart (session orientation, not post-tool tracking). Both rows said PostToolUse in ARCHITECTURE.md and the three INVENTORY locales; the issue asked for a neighbouring-row scan, which is how the session-state row was found. All other rows in the four tables verify against the surface. * fix(#3839): review fold-ins — 10 more wrong rows in ko-KR/pt-BR/zh-CN, parser authority + drift pins Adversarial review found the same two wrong rows shipped in five more files the issue's table missed (ko-KR ARCHITECTURE+INVENTORY, pt-BR ARCHITECTURE+INVENTORY, zh-CN ARCHITECTURE) — all fixed; DOC_TABLES now covers all ten shipped tables. The parity parser unioned only the Kimi mirror list, silently exempting agent-isolation-guard (registered via the dynamic preToolEvent push): probes are now parsed too, with bare hook names resolved against hooks/ ground truth and dynamic event variables resolved to their canonical (non-Gemini) events; an exact-set pin replaces the loose size guard. allow-test-rule marker carries the issue ref; unverified-ceiling 280→281 (audited: the new marker is legitimate — the suite reads product docs whose text is the contract). * fix(#3839): register the hook-table parity suite in the docs-guard lane The new suite reads ten docs/ paths, so lint-docs-guard-registration requires it in the docs-guard registry — the first GREEN bench run caught the omission (the RED run's docs-guard failures were the same signal, previously misread as marker fallout). * chore(#3839): changeset fragment (pr number backfilled after PR creation) * chore(#3839): backfill changeset PR number (4041) --------- Co-authored-by: sim <sim@local> |
||
|
|
12f9d1d9a0 |
enhance(#3913): docs, and the guards come down (#3994)
ADR-3889 terminal phase. Generated docs/reference/exit-codes.md from the exit-code declaration with a --check drift arm; deleted the inert soft-error-exit-zero oracle; promoted untyped-success from SMELL to VIOLATION so it can fail a build; pruned all 5 smell-baseline entries. Fixed inline: two mis-scoped oracles (routing-validity, value-hygiene), a second source behind the band table, unescaped declaration strings reaching Markdown, and a pre-existing Windows 8.3 short-name path-comparison defect. Guard ledger corrected from a claimed net -4 to a measured net -1. Closes #3913 |
||
|
|
ddde001af6 |
enhance(#3873): the STATE.md schema — one owner, generated artifacts (#3880)
* test(#3873): failing-first locale parity, plus tripwires for what must not move Pins ADR-3473 §8.8 at the artifact a reader actually sees. The English STATE.md reference carries a Status lifecycle section that is missing from all four translations — the section documenting the status enum whose clobbering is #3853. The test derives the heading set rather than hard-coding the missing one, and names the locale and the heading when it fails. Two tripwires that must pass today and after. The field-drift guard still catches a re-derived fallback ladder: §8.8 instructs deleting that script, and that instruction rests on a wrong premise about what it guards, so the test stops a future reader from deleting it on the ADR's word. And last_activity's label resolution is pinned to what ships today, because it is declared in one of the two tables this phase consolidates and not the other — the consolidation must not silently pick a side. The locale test buckets under docs rather than state, which is what it tests; that bucket is allowlisted with justification rather than folded into an unrelated docs suite. It reads only markdown, so it carries no allow-test-rule marker — a marker there would suppress nothing and would grow the unverified pool against its ceiling. Refs #3873 * feat(#3873): one schema owns the STATE.md key set, three tables become projections ADR-3473 §8.8. The key set was declared in four places that had to agree by hand and already did not: FIELD_CLASSIFICATION, FRONTMATTER_BODY_SOURCE, FRONTMATTER_KEY_TO_BODY_LABEL and buildStateFrontmatter's emit behavior. One frozen null-prototype schema now declares each key's type, enum, cardinality, source, preservation, body source, body label, accepted parse shapes and whether it is emitted unconditionally; the three tables are derived from it at module load. The projections are byte-identical to the literals they replace, key order included, and the parity tests compare against verbatim copies of today's tables rather than re-deriving both sides from the schema — a parity test fed from one source proves nothing, which is how a consolidation ships a changed policy under a green test. last_activity was the live disagreement: present in one table, absent from the other. The schema declares what ships today rather than the tidier answer, and a test pins it. The schema is a leaf module and owns the four field-policy types, re-exported from state-transition so existing importers are untouched — the same split health-diagnostic-types made to break a CJS require cycle. Refs #3873 * feat(#3873): generate the schema-derived regions, parity-check the prose tables ADR-3473 §8.8's generator half. gen-state-md-docs.cjs owns marked regions in the shipped template and all five reference docs, follows gen-features.cjs's fail-closed contract, and is wired into regen:derived and lint:generated-sync. The Status lifecycle section was missing from all four translations — the section documenting the status enum behind #3853 — and is now generated into every locale. Field cardinality is a new generated table: pure schema data, no prose, so nothing to lose. The Field-reference and Status-values tables are parity-CHECKED rather than generated. Their Purpose, When-populated and Matched-text columns are genuinely hand-translated per locale, and §8.8 itself says prose stays hand-translated; generating them from an English registry would overwrite four locales' translations on every write. The row set is checked against the schema instead, so a key added to one and not the other fails, which is what field drift actually means. Building that check found last_activity_desc undocumented in all five tables. Three keys the docs describe are absent from the schema — active_phase, next_action, next_phases. They are grandfathered by name, not by wildcard, so a fourth fails: a declared gap with a forcing function rather than a silent one. Refs #3873 * fix(#3873): declare what the parsers do, and close the shape-parity gap Two declarations in the new schema described intended behavior rather than actual — the defect class this epic exists to end, committed inside the epic. Both were caught by executing the parsers instead of reading their docstrings. current_plan.acceptedShapes claimed ['N', 'N of M']. Standalone, the hybrid shape errors; the path that looks like support is parseInt truncating '2 of 5' to 2 and discarding the rest. Narrowed to ['N']. The parser is deliberately NOT fixed here: that is #3784 and PR #3791 is already doing it. When #3791 lands this row must widen, and the shape test will go red until it does — the schema and the parser cannot drift apart quietly, which is what §8.8's checked-not- generated rule is for. STATUS_LIFECYCLE_ENUM claimed to be the closed set status can hold. normalizeStateStatus passes unrecognized prose through unchanged, so it is not closed at runtime. The seven members are the canonical values it maps onto; the docstring now says that and the test asserts the real lenient contract. Closes the acceptance item that a test asserts the parsers accept exactly the declared shapes: the check is table-driven over every row carrying acceptedShapes, guarded against passing vacuously on an empty set, and fails loudly if a future row has no registered driver. Adds the unwired-label throw and the fast-check property that every projection agrees with its schema row. Refs #3873 * fix(#3873): keep the shipped template's frontmatter first, and make row 27 able to fail The remote matrix caught 12 failures with one cause. Making the template's frontmatter a generated region wrapped it in its own yaml fence ahead of the markdown fence, so extractFileTemplate and readShippedStateTemplateBody — which both match the single markdown block — found the heading first, not the frontmatter. That breaks the contract every new project's STATE.md is created from: bug #21 and epic #1969 B8 pin that the File Template block starts with frontmatter and carries gsd_state_version. The markers now sit inside the single markdown fence, so the fence opens before the frontmatter and the region still ends ahead of the heading. Same layout as before this phase, with markers embedded rather than a second fence. Row 27 existed to catch exactly this and did not, because it was writer-seeded: it asserted against the generator's own output shape, so it passed on the broken template. It now parses the fence the way production does and was verified to fail against the broken shape before being trusted against the fixed one. A test that would not have caught the bug it exists to prevent is worse than no test. The emitted-attribution failure was separate and the fragment was the wrong remedy: gsd-core/templates/state.md self-attributes under a verbatim-copy identity rule, so a diff touching it needs no acknowledgment. Fragment deleted rather than left explaining nothing. Refs #3873 * docs(#3873): how to change the STATE.md schema The phase gate was right and my docs artifact was wrong. I listed lint:generated-sync as the second enablement step, which is a verification command dressed as one, and then claimed a one-step sequence owed no how-to. The real sequence is build:lib then regen:derived, and the ordering is a trap: the generator reads the COMPILED schema, so regenerating before building regenerates against the previous schema and commits artifacts that look plausible while disagreeing with the code just written. A reference table cannot carry an ordering dependency; that is what the how-to test is for. The page covers adding, changing and removing a key, every reason code the check emits and what to do about each, what is generated versus hand-translated and why the two prose-bearing tables are parity-checked instead of generated, adding a language, and the three grandfathered keys. Indexed from docs/README.md. Refs #3873 * chore(#3873): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
36375513b9 |
feat(#3840): generate docs/FEATURES.md from per-feature fragments (#3845)
* feat(#3840): generate docs/FEATURES.md from per-feature fragments docs/FEATURES.md was hand-maintained, and every feature PR wrote into two shared mutable cells: the '### N.' heading whose integer was hand-allocated at authoring time, and the hand-maintained table of contents. Concurrent PRs all picked the same next integer, and two PRs adding differently numbered features still collided on the TOC. #3831 was renumbered 165 -> 166 -> 167 -> 168 across successive rebases, each collision also costing a full matrix verification run because the sha-keyed pass marker dies with the rebase. Mechanism: one fragment per feature at docs/features/<slug>.md carrying id/title/group (and an optional order) in frontmatter, consolidated by scripts/gen-features.cjs --write|--check into a marker-delimited region of docs/FEATURES.md that holds BOTH the TOC and every section body. Group headings and their order are derived too - a group sorts by its lowest-ordered member - so there is no shared registry to edit either; optional per-group prose lives in docs/features/_groups/<slug>.md. A contributor adds exactly one new file. Wired into regen:derived and lint:generated-sync alongside the eight existing generators, matching gen-adr-index.cjs's CLI shape and typed-REASON reporting. Migration froze all 168 existing numbers verbatim: identical section set, identical order, identical bodies. Two defects found in the tree are fixed inline rather than carried forward - the '## Related' block had been spliced into the middle of the document, orphaning §142's Reference line, and four inbound anchors were already broken on next (FEATURES.md#runtime-identity in two files, and #143-spec-phase-edge-completeness-probe off by one). Since the repo has no link checker, --check now validates every inbound FEATURES.md#anchor by resolved target, so that class cannot ship silently again; locale FEATURES.md files resolve elsewhere and stay out of scope. Refs #3840 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3840): carry upstream §69 delta into its fragment and harden the generator Review found section 69 missing '[--strict]' and REQ-STATE-05/06 versus origin/next. Root cause was a stale base, not extraction loss: those lines landed in |
||
|
|
4b84be1da4 |
fix(#3683): wire gated learnings extraction into completion, align copy path (#3810)
* test(#3683): failing-first rows for learnings source resolution and wiring pins * fix(#3683): wire gated learnings extraction into completion, align copy path * test(#3683): register the learnings suite in the docs-guard lane, drop unverified markers * fix(#3683): close review findings — per-item parsing, readdir guards, docs paths * fix(#3683): route phase enumeration through the locator seam, fix assertion targets * fix(#3683): merge execute-phase ack into the 3003 fragment, fix fidelity targets * fix(#3663): replace the spent execute-phase ack entry with the 3683 re-arm * chore(#3683): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
107eb8c1d9 |
feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on
|