af822a8024f732c167b163f163a0487bdbcfebdf
61 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7d0c6339d0 |
fix(#4705): emit Antigravity-native tool names as a YAML sequence (#4822)
* test(#4705): add failing-first coverage for native Antigravity tool sequences * fix(#4705): emit Antigravity-native tool names as a YAML sequence convertClaudeAgentToAntigravityAgent and the installer's twin emitted Gemini CLI tool names as a comma-separated scalar. Antigravity's documented subagent contract (antigravity.google/docs/subagents) wants a YAML sequence of native names — view_file, grep_search, run_command, replace_file_content are the documented examples, and wrong or malformed grants can hang the subagent per Antigravity's own warning. Map values move to the native vocabulary where documented (Read -> view_file, Edit -> replace_file_content, Bash -> run_command, Grep -> grep_search); undocumented entries keep their best-known grant rather than being dropped (dropping would silently remove a restriction). The emitter writes one '- name' item per line; an agent whose every tool was filtered emits an explicit tools: [] instead of an empty scalar. Pre-existing pins updated to the native vocabulary. * test(#4705): update the #4727 map-value pin to the Antigravity-native vocabulary The #4727-era pin held the map VALUES at the Gemini CLI dialect on the belief that Antigravity speaks it; the confirmed bug #4705 (with Antigravity's own documented subagent contract) supersedes that for the four documented names. Key/shape pinning is preserved; only the values move. * docs(#4705): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
2f0e99f9e0 |
fix(#4377): opt in to project-relative includes for local installs (#4425)
* enhance(#4377): opt-in project-relative includes for local installs A local install wrote the includes that point at GSD's own files as absolute paths — whatever the installer resolved at install time. For one checkout that is invisible. Across git worktrees it is not: each worktree gets its own .claude/ copy, but all of them point back at the checkout that ran the installer, so a worktree runs its own gsd-tools.cjs while reading workflow prose from a different checkout. Update that one checkout and every other worktree is running new instructions against an old engine, with nothing to stage the update with. --relative-includes (or GSD_RELATIVE_INCLUDES=1) makes a local install emit `@.claude/gsd-core/...`. Opt-in, and staying opt-in: absolute works for a single checkout, which is most people, and flipping the default would change every existing local install to solve a problem those users do not have. The prefix is the runtime's own localConfigDir descriptor value, never a literal — the same value resolveScope joins onto the cwd to produce the install target, and the same one the rewrite engine already uses for its ./.claude/ -> ./<dir>/ substitutions. Copilot and Antigravity have shipped this shape for local installs since they were added, with hardcoded .github/ and .agents/; this is that behavior, derived rather than written down. Six seams compute a path prefix and all six had to be threaded, which is why the opt-in travels through the environment the way --portable-hooks already does: one variable they all read cannot fall out of sync the way six signatures can. The launcher shim deliberately keeps its ABSOLUTE fallbacks. It probes gsd-tools through ${CLAUDE_CONFIG_DIR:-$HOME/.claude} and one such default per runtime; those are shell word expansions, not includes, and a relative value there resolves against the shell's cwd rather than the project. Trading an include that points at the wrong checkout for a path that points at nothing is not a fix. All three rewrite paths mask ${VAR:-default} spans before substituting and restore them after, and the mask only runs when the prefix is relative, so an absolute install is byte-for-byte unchanged. Every unexpressible case falls back to absolute: no opt-in, a global install, a missing dir name, the configHome.kind === 'none' sentinel, an absolute descriptor value, or one climbing out of the project with '..'. * chore(#4377): add changeset for project-relative local includes * fix(#4377): compare against POSIX-normalized roots in the install e2e arms The emitted prefix is POSIX-normalized by design — it is substituted into markdown @-references, which use forward slashes universally, so a backslash would leak into shipped content (#1615). The e2e arms compared against the raw temp root, which on Windows is `D:\a\...` and appears in no emitted file. That reddened the control arm on the windows shard, and it was worse than a red: the negative arm ("nothing references the checkout") was passing VACUOUSLY there, because a string that cannot occur is trivially absent. Both now go through the same normalization, so the Windows lane asserts what the Linux lane does. * fix(#4377): tolerate a resolved temp root, and make the e2e diff self-diagnosing Two changes, one confirmed and one to stop guessing. Confirmed: the emitted content carries the RESOLVED root, not the spelling mkdtemp handed back. Reproduced on Linux with a symlinked install root — 236 emitted files carry the realpath, zero carry the link path. macOS has this structurally, since /var is a symlink to /private/var. Comparisons now go through both spellings, or the negative arms pass vacuously: "nothing references the checkout" is trivially true when the string being searched for cannot occur. Not confirmed: the macOS shard reported ~every workflow file differing in the "differ ONLY" arm while the five arms around it passed, and the assertion printed a list of filenames — which says a difference exists somewhere across 236 files and leaves the reader to guess which bytes. I cannot reproduce that platform locally, and guessing turns one CI round-trip into four. The assertion now reports the first divergence as text: the file, the byte offset, and a bounded window of both sides. * fix(#4377): strip the longest root spelling first in the install e2e diff The macOS failure was my test corrupting its own comparison, not a product defect. /var/folders/…/X is a SUBSTRING of /private/var/folders/…/X, so stripping the unresolved spelling first matched inside the resolved one and left the /private prefix glued to what followed: @/private/var/…/X/.claude/gsd-core/… -> @/private.claude/gsd-core/… a string present in neither install, which is why all 236 files "differed". Sorting the spellings longest-first consumes the whole occurrence, and the short form then has nothing left to match. Proven in isolation on the exact macOS shapes: short-first yields @/private.claude/…, longest-first yields @.claude/…. The self-diagnosing assertion added in the previous commit is what found this — it named the file, the byte offset, and printed both sides, so the corrupted string was visible rather than inferred from a list of 236 filenames. Keeping it. * fix(#4377): address review findings * test(#4377): scan nested shell defaults without regex backtracking * fix(#4377): close relative include review gaps * fix(#4377): preserve root-target runtime includes * fix(#4377): guard project-root relative includes * test(#4377): normalize Cline fallback roots * fix(#4377): persist relative include style --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
b647e28313 |
chore(#4727): name the tool-conversion helpers for the runtime that uses them (#4732)
* chore(#4727): name the tool-conversion helpers for the runtime that uses them
GSD has had no Gemini runtime since #1928 removed it (Google sunset Gemini CLI
on 2026-06-18, shipped 1.8.0), yet two helpers were still named for it:
claudeToGeminiTools -> claudeToAntigravityTools
convertGeminiToolName -> convertAntigravityToolName
The sole consumer is convertClaudeAgentToAntigravityAgent, whose own comment
read "Map tools to Gemini equivalents (reuse existing convertGeminiToolName)".
Nothing named Gemini consumes them, because nothing named Gemini exists. The
new names follow the convention the file already sets with its neighbouring
Copilot pair, claudeToCopilotTools / convertCopilotToolName.
Zero behavior change. Every mapped VALUE is byte-identical, deliberately:
read_file, write_file, replace, run_shell_command, glob,
search_file_content, google_web_search, web_fetch, write_todos
Those are Gemini's built-in tool dialect and Antigravity genuinely speaks it.
This rename covers only the identifiers, which are the one part of the surface
that was GSD's choice rather than Google's contract.
Renamed in BOTH copies. CLAUDE.md labels bin/install.js "(generated)", but no
script emits it -- build:lib is tsc -p tsconfig.build.json and writes only
gsd-core/bin/lib/**. These converters are the #1099/#1173/#1182 situation: they
were extracted into src/runtime-artifact-conversion.cts while bin/install.js
kept its own working inline copies, so each symbol existed twice in two
independently hand-maintained files. Renaming one would have left two names for
one concept. Verified first that no capability descriptor resolves either by
name -- antigravity's descriptor names only convertClaudeCommandToAntigravitySkill
and convertClaudeAgentToAntigravityAgent, neither of which moved.
Comments keep their reasoning and their issue refs (#3362 AskUserQuestion,
#1394 Skill/SlashCommand); only the subject is corrected, from "Gemini CLI" to
Antigravity speaking the Gemini dialect. Those describe the dialect's behavior,
which is still Antigravity's behavior, so deleting them would destroy the record
of two real bugs.
docs/research/gemini-to-antigravity-migration.md is left unedited and carries a
dated addendum instead: it is pinned to
|
||
|
|
f334f277dd |
fix(#4324): stop the retired /gsd: prefix reaching users (#4712)
* test(#4324): prove colon tokens the installer cannot convert leak Failing-first regression coverage for #4324. The install rewrite (transformContentToHyphen) is gated on an exact match against the commands/gsd stem list, so any /gsd:<token> whose token is not a registered stem survives the install and reaches the user as the deprecated colon form. The gate is load-bearing -- it is the only thing protecting the workflow DSL marker family (gsd:section, gsd:protected, gsd:loop-host, gsd:guard, gsd:dispatch, gsd:plan-revision-conflicts), which workflow-fragments parses as a literal. So this suite asserts the shipped text is convertible rather than asserting the transform is broad, and pins the marker family as explicit negative space. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4324): stop unconvertible colon tokens reaching the user The install rewrite is gated on an exact match against the commands/gsd stem list, so a /gsd:<token> whose token is not a registered stem survives the install and reaches the user as the deprecated colon form. That gate is load-bearing -- it protects the gsd:section / gsd:protected / gsd:loop-host marker family -- so the fix is in the shipped text, and the source stays colon per CONTEXT.md's two-tier rule. - quick-batch command + skill description: close the command token at a boundary so `/gsd:quick`-shaped converts instead of being skipped. - gsd-code-fixer (both variants): execute-plan and diagnose-issues are workflows, not commands, so they never converted and rendered beside two hyphenated siblings on the same line. Name them as workflows. - help topic-mode: the extraction rule hard-coded a colon prefix that the converted full.md never ships, so --brief could never match a signature line and silently fell back on every topic. Describe the signature line without a literal prefix. - update.md: drop the prefix from prose describing a stale command. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4324): add changeset fragment pr:0 placeholder is backfilled with the real number once the PR exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4324): locate the help summary per reference variant Adversarial review finding. Restoring the signature-line match (the #4324 fix) activated a latent defect in the clause next to it: compact scope emitted "the single non-blank line immediately after" the signature, and that clause is only correct for full.md. full.compact.md puts the summary on the signature line itself, after an em-dash, and its next non-blank line is an unrelated "Usage:" line. Both variants ship and both are served, so before this commit the compact variant would have emitted the wrong line as the summary. It was masked until now only because the stale colon prefix meant no signature line ever matched at all. Name the two placements and pick per line, and say explicitly that a Usage: line is never a summary. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4324): de-vacuum the help parity check, narrow the marker waiver Two adversarial review findings against the #4324 coverage. The help-parity assertion went vacuous the moment the fix landed: once topic.md stops spelling a literal prefix, the matched set is empty and the assertion holds for any rewording, correct or not. It now also asserts across BOTH served reference variants that each ships signature lines under the hyphen prefix, that the two genuinely disagree about where the summary sits, and that topic.md still names both placements and the Usage: guard. The marker waiver keyed on "sits inside an HTML comment", which waves through a real broken reference that happens to be commented out -- `<!-- see /gsd:typo-cmd -->` scored clean. Enumerate the six marker families instead. Verified the narrowed rule catches that probe and still passes over the tree; it also surfaced a seventh family, write-continue, that the broad rule was hiding. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#4324): normalize the namespace in skill descriptions Both hyphen-namespace skill converters ran the hyphen transform over the body but rebuilt the frontmatter description from the raw field, so a /gsd:<cmd> mention in a command description survived into the installed SKILL.md -- the exact field the host's skill picker renders, which is the surface this issue was filed about. The local flat-command path was already correct because it rewrites the whole file; only the skills path, used by a global install, was affected. Confirmed by installing into a fake HOME before and after. Fixed in both copies: bin/install.js and the src/ source of truth that compiles into gsd-core/bin/lib. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4324): assert descriptions through the real converters The previous version of this check called transformContentToHyphen on the description line itself and passed, while a real install still shipped the colon form -- the converter never calls that transform on the description. It asserted a proxy for the behaviour instead of the behaviour. Drive convertClaudeCommandToClaudeSkill and convertClaudeCommandToClineSkill over every registered command and assert on the emitted description. Verified it fails against the pre-fix converters and passes against the fixed ones. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4324): regenerate skills after the description change skills/<name>/SKILL.md is generated by gen-plugin-skills, not hand-maintained, and lint:generated-sync caught the hand edit. The regenerated file emits the hyphen form, which also corrects the assumption behind the scan comment in the namespace test: skills/ is runtime-emitter output, not colon source. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#4324): re-sanction normalizeKimiSkillName's real end line The description-normalisation fix inserted five lines above normalizeKimiSkillName in src/runtime-artifact-conversion.cts, moving its closing brace from 635 to 640. MAJOR-1 pins that line deliberately, so the planted violation landed INSIDE the exempted body and went unflagged -- 0 !== 1. Re-sanction the value rather than derive it: the array is named sanctionedRealEndLines, and a pinned line that fails loudly on drift is the design. Deriving it would remove the human check the name asks for. Verified by executing all four MAJOR-1 rows against the real tree: each planted violation is flagged at realEndLine+1 and each unmodified file stays exempt. Emitted-Drift-Ack-Growth: gsd-code-fixer.md — names execute-plan and diagnose-issues as workflows rather than as slash commands that do not exist Emitted-Drift-Ack-Growth: gsd-code-fixer.compact.md — same rewording as its full sibling, kept byte-consistent with it Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#4324): backfill the changeset PR number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
93ba63aeff |
fix(#4482): strip Copilot notes from OpenCode artifacts (#4532)
* fix(#4482): strip Copilot notes from OpenCode artifacts * chore: add changeset for #4532 * fix(#4482): filter runtime notes across emitted surfaces * fix(#4482): make runtime note filtering runtime-neutral * test(#4482): keep ack fixture outside registered provenance --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
86b745b48b |
fix(#4270): forward Codex spawn model routing (#4281)
Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
925a363879 |
enhance(#4032): apply configured agent tool grants (#4238)
* test(4032): add failing installed-agent grants contract Cover global and project agent_tools precedence at the real Claude installer seam before adding implementation. * feat(4032): apply configured agent tool grants during staging Resolve selector-level global and project config once per staging call, then append validated grants before runtime conversion. * test(4032): cover host grant and quoted MCP contracts Exercise installed host artifacts and prove ZCode must treat quoted MCP scalars like plain MCP grants. * feat(4032): apply configured agent tool grants across runtimes Move augmentation and scalar identity into the converter seam so every staged artifact preserves host policy. * fix(4032): register agent tool grants in configuration Accept documented agent_tools config without unknown-key warnings.\n\nKeep installer fixtures on the shared temporary-directory helper. * fix(4032): translate configured MCP grants for Kilo Reuse the converter-owned scalar decoder so quoted canonical grants reach Kilo's native permission keys without altering other host policies. * fix(4032): decode YAML-escaped tool grants * fix(4032): emit valid inline agent tool grants * fix(4032): reject invalid trailing-colon grants * test(#4032): cover cross-review remediation gaps * fix(#4032): close cross-runtime grant gaps * test(#4032): expose Kimi global project context * fix(#4032): preserve Kimi project config context * chore(#4032): add release note * test(#4032): expose fork review regressions * fix(#4032): address fork review findings * test(#4032): make byte-stability assertion portable Compare repeat installs at one root so platform-specific path rendering cannot masquerade as an agent_tools behavior change. * chore(#4032): bind changeset to upstream PR 4238 * fix(#4032): address trek-e review findings (2,3,4,5,6,7,8) Fixes fail-closed decode-failure handling in ZCode's mcp__ stripper, a comment-only `tools:` header mis-parse that silently dropped configured grants, and a naive comma-split that could tear a quoted scalar containing a literal comma. Documents Kilo's inherent `{server}_{tool}` MCP-permission-key collision (external, fixed format — not ours to widen) and locks the existing first-seen-wins resolution in with a regression test. Opts kimi/kimi-code out of the ADR-1235 pre-converter path-rewrite step: routing Kimi through that pipeline (needed so project-scoped agent_tools selectors reach it) was short-circuiting Kimi's own neutralizeKimiAgentPrompt, which expects the original ~/.claude/gsd-core text rather than a pre-rewritten Kimi path. Extends the fast-check token pool and per-runtime install coverage with the missing comment/comma/broad-runtime cases the prior review flagged as untested. * docs(#4032): add CONTEXT.md glossary entries for agent_tools resolver + pre-converter step Documents readGsdEffectiveAgentTools (Install Model Override Resolver Module) and the appendAgentTools pre-converter pipeline step (Runtime Artifact Conversion Module), per contributor-standards.md's new-seam glossary requirement (finding 1). * fix(#4032): address agy adversarial review findings An agy (gemini-3.8-flash-high) adversarial pass over the prior review-fix commit found the fixes for findings 3, 4, 6 and 8 had unfixed sibling gaps, plus a genuine new regression and two CONTEXT.md inaccuracies: - ZCode's comment-only `tools: # note` header matched the inline-value branch instead of falling through to the block-list scan, so a following mcp__* item leaked through unstripped — the exact defect finding 4 fixed in appendAgentTools, unfixed in this sibling function. - Reverted capabilities/kimi-code/capability.json's noPathRewrite: true. kimi-code uses the standard 'agents' kind with converter: null (not kimi-agents — confirmed by reading the descriptor, not its prose description), so it never went through the pipeline change finding 5 fixed, and disabling its path rewrite broke every ~/.claude/ embed in its shipped agents instead. - decodeToolScalar never stripped a trailing ` # comment` from a bare (unquoted) scalar, so a comment after a block-list item, or after an appended grant on an inline line, became part of the "tool name" — fixed at the source (one call site fixes every consumer). - appendAgentTools's comment-index scan wasn't quote-aware, so a `#` inside a quoted scalar (`"mcp__server #1"`) was mistaken for a comment start and corrupted the quote. - parseFrontmatterTools (Kimi/Qwen's tool-list reader, downstream of appendAgentTools's own output) had the same naive comma-split and comment-only-header gaps as findings 4 and 6, unpatched. - The all-runtime smoke test's presence assertion was built on a guessed omit-list; empirically only 7 of 17 runtimes keep an arbitrary mcp__ grant recognizable, replaced with a verified allowlist. - CONTEXT.md claimed a `project:<agent>` selector prefix that does not exist (project override is a same-key merge across two config files) and mislabeled stageAgentsForRuntimeWithConverter's module. * fix(#4032): address full-PR review (Opus critical/ponytail + agy) A whole-PR pass (critical-code-reviewer + ponytail-review on Opus, plus a second agy full-source adversarial pass) surfaced defects the earlier finding-scoped passes couldn't reach: - appendAgentTools corrupted a `tools:` line whose ENTIRE value is a leading quoted scalar (`tools: "Read"` -> `tools: "Read", Write`, invalid YAML) — there is no safe line-surgical rewrite here, so it now refuses to touch that shape instead of emitting broken frontmatter. - decodeToolScalar's malformed-trailing-quote check ran BEFORE comment stripping, so a bare tool name with a quote inside its own trailing comment (`Bash # note: "internal"`) was wrongly rejected. Reordered. - findUnquotedCommentIndex (added in the prior remediation commit) was built on a wrong model of YAML: a `#` after whitespace starts a real comment in a plain scalar regardless of nearby quote characters — verified against the actual parser. The one case that DOES need protection (a leading quoted scalar) is now refused outright above, so the quote-tracking scan was dead weight solving a problem that no longer reaches it. Removed; reverted to the plain `[ \t]#` scan. - Kilo has a SEPARATE agent-frontmatter parser (convertClaudeToKiloFrontmatter, distinct from the buildKiloAgentPermissionBlock fixed earlier) with the same comment-only-header and naive-comma-split gaps as findings 4 and 6 — unfixed in both its src/ and bin/install.js copies. Fixed in both, exporting splitToolScalars for bin/install.js to reuse rather than reimplementing it. - Pipeline docstring in stageAgentsForRuntimeWithConverter still listed 5 steps, omitting appendAgentTools (now step 3 of 6). - docs/CONFIGURATION.md didn't state that a --global install still discovers agent_tools from the cwd's .planning/config.json (confirmed intentional and already covered by a dedicated test, not a bug). - Removed install-engine.cts's deps.cwd injection seam: zero callers or tests ever populated it. Two claims from this round were verified and rejected, not fixed: prototype pollution via a `__proto__` selector key (empirically confirmed `Object.prototype` is never touched — only reassigns the resolver's own local object's prototype, with no observable effect), and a `*` grant value crashing YAML parsing as an alias reference (empirically confirmed it parses as plain scalar text, no crash). A pre-existing, unrelated defect (extractFrontmatterField returns null for block-list `tools:` on Copilot/Antigravity/Cursor/Codex/Qwen, affecting two shipped agents today) was filed as a follow-up rather than fixed here — it predates #4032 and isn't caused or worsened by this PR. * fix(#4032): update stale slug-derivation-drift-guard fixture line normalizeKimiSkillName's real closing brace moved from line 616 to 635 as a side effect of this PR's edits to runtime-artifact-conversion.cts; the MAJOR-1 fixture's hardcoded realEndLine had gone stale. * fix(#4032): address CodeRabbit findings on projectDir threading and flow-sequence tools bin/install.js's installAgentsKindStandalone call site omitted the projectDir argument the function already supports, so a global install through this legacy branch silently fell back to the runtime config dir instead of process.cwd() when resolving project-scoped agent_tools grants — inconsistent with the sibling installOpencodeFamilyArtifacts call site, which already threads it correctly. appendAgentTools' leading-quoted-scalar bailout did not cover a YAML flow sequence (`tools: [Bash, Read]`): splitToolScalars tore it apart on the in-sequence commas and appended past its closing bracket, producing invalid frontmatter. Extended the bailout regex to also refuse a value starting with `[`, matching the same "whole node, nothing may follow" reasoning already applied to quoted scalars. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
dce40eeb6e |
fix(#4002): rewrite zcode command @-refs to the zcode runtime home (#4188)
* test(#4002): zcode commands must rewrite at-refs to the zcode home * fix(#4002): add the missing zcode case to the runtime rewrite engine * chore(#4002): add ZCode to the bug-report runtime dropdown and drop the changeset * fix(#4002): attribute zcode command and skill ripples to the rewrite engine * fix(#4002): attribute zcode nested-skill ripples to the rewrite engine * chore(#4002): backfill changeset pr number * fix: bump qs past GHSA-x5fp-wj9c-mxmx (transitive, advisory reddened next) --------- Co-authored-by: sim <sim@local> |
||
|
|
39673ae9ff |
fix(#3738): antigravity global skills/agents install to ~/.gemini/config (#3921)
* test(#3738): antigravity global skills/agents must resolve under ~/.gemini/config Regression tests (RED first): --skills-root and gsd-tools query surfaces, install-plan dest dirs, and converter skills-path rewrite. * fix(#3738): antigravity global skills/agents install to ~/.gemini/config Antigravity's machine-local discovery scans ~/.gemini/config/{skills,agents}; the configHome (~/.gemini/antigravity) is deprecated for artifacts. Declare the ADR-1239 skills/agents 'home' override on the antigravity global layout — the same mechanism codex uses (.agents) — and divert ~/.claude/skills/ references in converted global content to ~/.gemini/config/skills/. configHome, settings, probe/migration semantics, and the local .agents layout are unchanged. * fix(#3738): retire deprecated configHome artifacts via installer migration 010 Next install converges an existing antigravity install: manifest-managed skills/gsd-*/ and agents/gsd-*.md under the configHome (a location AGY does not scan) are removed — modified files backed up first, unmanifested and non-gsd entries preserved — and now-empty containers retired. Global scope only; the local .agents surface is live. Docs + inventory updated. * fix(#3738): converter sync in bin/install.js, harness emit-root coverage, migration baseline - bin/install.js converter gains the same ~/.claude/skills → ~/.gemini/config/ rewrite as src (ADR-1508 dual copy must stay in sync). - Parity-manifest walk covers home-override emit roots (extraEmitRootsFor) so antigravity's emitted skills/agents stay differential-visible at their new install root; install-tree fixture regen confirms an unchanged key set. - skills-from-commands rule declares the antigravity converter as a runtime-scoped transform; one ack fragment covers the identity-classed workflow whose antigravity copy embeds the old skills path. - Migration 010 checksum baseline + home-override set doc updated; existing tests updated to the #3738 contract (global dest, golden parity via layout dest, integration expectations). * fix(#3738): tolerate an absent extra emit root on baseline-side measurement The base tree's installer predates the home override, so <HOME>/.gemini/config does not exist there; walk() threw ENOENT and the in-job baseline build failed. An absent extra root is the legitimate pre-override shape — skip it. * fix(#3738): review findings — manifest agents root, bare skills-path rewrite, guard comment - writeManifest resolves the agents-kind home override (_kindDestDirSafe), so the manifest records agents at their actual install root and drift detection keeps working (isolated review finding 1, major). - Converter bare forms ~/.claude/skills and $HOME/.claude/skills (no trailing slash) divert to ~/.gemini/config/skills instead of falling through to the retired configHome path (finding 2). - real-home-guard comment updated: antigravity's global agents kind is the first agents-kind home override (finding 3, doc-only). - Regression tests for both behavioral findings. * chore(#3738): changeset fragment (pr number backfilled after PR creation) * chore(#3738): backfill changeset PR number (3921) * fix(#3738): sandbox HOME in tests that install antigravity global artifacts antigravity is the first home-override runtime in the golden-parity and skills-wrapper suites (codex is not in their runtime lists), so those tests never needed HOME sandboxing — the real-home guard now (correctly) refuses their un-sandboxed global installs on CI, where HOME is the passwd home. * fix(#3738): stop the K3 sequential-sandbox env leak; sandbox L2's home-override plans K3's two back-to-back sandboxHome calls leave HOME pointing at the first sandbox once the after-hooks restore (each call saves the env as it found it, so the second saves the first's sandbox as 'original'). On the windows matrix that leaked gsd-k3-qwen-* home into the L2 property, whose antigravity/global run then (correctly) refused via the #3712 real-home guard — antigravity is the runtime that made L2's plan escape into os.homedir(). K3 now manages the env with a single restore; L2 sandboxes HOME per run, mirroring L1. * fix(#3738): L2 property's HOME sandbox must exist on disk The #3712 guard's sandbox exemption fails closed when identify(effectiveHome) is 'absent' — L2 never created its configDir, so on the windows matrix (tmpdir under the real home) the antigravity/global run refused even with HOME sandboxed. Create the per-run sandbox dir and clean it up. --------- Co-authored-by: sim <sim@local> |
||
|
|
8641d0a468 |
fix(#3719): restore @-includes on the agents emit path, so global Claude installs load their guidance (#3918)
* test(#3719): failing-first coverage for the agents emit path's missing tilde restore applyAgentPathRewrites performs its four tilde and HOME substitutions and never calls restoreClaudeGlobalAtRefTilde, which has exactly one call site in the module and it is not this one. So every agents/gsd-*.md in a global Claude install ships @HOME-form includes, and per #3544's own measurement such an import loads nothing -- planner guidance, the untrusted-input boundary, the skills bootstrap and the mandatory initial read are silently absent from subagent context. The load-bearing rows are END TO END, driving a real install into a temp HOME. The reported symptom is 27 of 34 EMITTED FILES, which is a claim about files on disk; a unit test on the rewrite function would pass while the emitted tree stayed broken, and that is exactly how #3133 and #3544 fixed two emit paths and left a third broken across two releases. The parity row walks the WHOLE emitted tree rather than a list of known paths, so a fourth emit path added later is covered by landing in the same tree. Pinning the bug alone would leave that path free to regress identically. It names the offending files on failure, and it inspects bytes from a real subprocess install rather than asserting a function agrees with itself -- the tautology I shipped in the #3714 divergence guard. Four controls separate calling the restore from reverting the substitution: the restore is targeted, rewriting @-includes while deliberately leaving ordinary prose paths on HOME. Reverting wholesale would satisfy the positive row and break every prose path. One boundary row is deliberately red beyond the obvious fix. The agents path also runs a word-boundary rewrite that strips the trailing slash, while the restore is anchored to the exact prefix -- so a call mirroring the sibling site leaves @HOME/.claude with no trailing slash broken. Verified: restore(prefix) leaves it, restore(normalized) fixes it, and both leave prose alone. * fix(#3719): restore @-refs to tilde on the agents emit path The third emit path that needed this. #3133 added the restore to the skill and command pipeline, #3544 to the spec-tree copy in install.js, and the agents pipeline never got it -- so every @~/.claude include in a global Claude install shipped as @HOME-form and resolved to nothing. Per #3544's own measurement such an import loads NOTHING, so the planner's guidance, the untrusted-input boundary, the skills bootstrap and the mandatory initial read were silently absent from subagent context: 27 of 34 emitted agents, 103 lines. Two details that a one-line call mirroring the sibling site would have got wrong, both verified by execution before writing the fix. It passes the NORMALIZED prefix rather than pathPrefix. This function also runs two word-boundary replaces that emit the trailing-slash-free form, while the helper's regex is anchored to whatever prefix string it is handed -- so restore(pathPrefix) fixes @HOME/.claude/x and leaves a bare @HOME/.claude broken. The normalized form is a prefix of both, so one call covers both. And it is guarded on claude. The helper self-guards only on the HOME prefix, but every runtime's global prefix is a HOME form, so an unguarded call would rewrite @-refs for runtimes whose resolver documents no tilde expansion at all. Verified: claude restores, cursor and kilo do not. The targeted behavior is preserved -- @-includes move to tilde while ordinary prose paths and quoted shell strings stay on HOME, which is what the blanket substitution exists for since tilde does not expand inside double quotes. * fix(#3719): stop the word-boundary replaces corrupting a non-default config dir Review found that my fix MASKED a pre-existing bug, which is worse than leaving it. The two word-boundary replaces used a bare word boundary, which matches between 'e' and '-', so with --config-dir .claude-work they turned .claude-work/ into .claude-work-work/ -- 119 dead paths in a real install. #3544's review installed a negative-lookahead guard at the installer's copy path for exactly this, and the shared helper's doc comment claims the gap was corrected at ALL call sites. It was not corrected here. That much is pre-existing; base emits the same 119. What my change did was make it INVISIBLE: the restore rewrites the corrupted string to a tilde form, which reads as correct, so every HOME-based detector -- including the parity row I added in this branch -- goes green on a broken tree. A fix that hides the evidence of a neighbouring bug is not a fix. Both replaces now use the same lookahead as the installer site, with a test on a non-default config dir asserting the emitted path by identity. Three test weaknesses from the same review, all of which would have passed while guarding nothing: The parity row anchored its detector at line start, so it could not see the 48 mid-line refs the helper deliberately supports -- half-blind while being billed as the future-proof row. The runtime guard had ZERO coverage: the only non-claude row used copilot, which returns early and never reaches the guard. Cursor and kilo rows now exercise it. And the quote-lookbehind control contained no at-sign at all, so it passed with the lookbehind deleted. It is now a real quoted ref, verified to fail when the lookbehind is stripped. * test(#3719): distinguish a live @-include from prose describing one My own MINOR fix introduced a BLOCKER. Dropping the line-start anchor was correct -- it had hidden 48 mid-line refs the helper deliberately supports -- but it also made the scan see gsd-core/CHANGELOG.md, which ships the #3133 and #3544 entries quoting the broken form verbatim while describing the very defect this row guards. Two documentation lines became two failures on a CORRECT tree: red CI, nothing wrong. Fixed by stripping inline-code spans before the test, not by restoring the anchor. Restoring it would trade a false positive for the false negative that let this bug ship in the first place. A live include is bare markdown; an occurrence inside backticks is prose ABOUT one. Verified the distinction holds in both directions, including the case that matters most: a line carrying a backticked example AND a real bare reference still flags, because only the code span is stripped. * fix(#3719): escape the replacement pattern, and pin both fixes that shipped unproven Security's remaining item, landed here on its recommendation: the restore used a STRING replacement, so a config dir containing the ampersand or backtick dollar forms was treated as a special pattern. Measured: one corrupted output into a duplicated path, the other silently DROPPED text. Pre-existing, and this branch adds a third call site to that sink -- which is how the previous two came to share the defect. It is a function replacement now. Two fixes on this branch were shipping UNPROVEN and both are now pinned: The word-boundary fix had no regression test at all. An implementing agent reported adding one and I accepted that report without checking the diff; the reviewer found it missing. Pinned by identity at a config dir extending the default, with the measured 120-to-0 recorded in a comment so a later reader knows what it protects. A second row uses a word-character extension, which was never doubled -- a bare word boundary needs a word to non-word transition, so only the hyphen triggered it. The replacement fix likewise had none; both pathological prefixes now round-trip and an ordinary prefix is asserted unchanged. Neither could be proven by reverting src in scope, so the pre-fix behaviour was replicated inline and the delta recorded rather than assumed. * test(#3719): state the trust model accurately in the replacement-pattern note The comment described the config dir as attacker- or operator-controlled. Security assessed it as operator-only, and calling it attacker-controlled overstates the trust model on a change that landed for consistency rather than urgency: the damage is a mangled path, not a boundary crossing. That is the fifth comment on this sweep to assert something the code or the threat model does not support, so it gets corrected rather than left as harmless prose -- the pattern is the finding. * chore(#3719): backfill changeset pr number Doing this immediately after PR creation this time: the same omission was the only red CI on the previous PR tonight. --------- Co-authored-by: sim <sim@local> |
||
|
|
6b7df61938 |
enhance(#3881): one YAML parser — vendored js-yaml replaces the hand-rolled dialect (#3888)
* docs(#3881): answer §8.1's open question and correct three wrong premises ADR-3473 §8.1 carries a blocking open question with a forcing function: it must be answered before any implementation PR for the rule opens. Answered here as (a), a string-coercing adapter, with the measurement that settles it. The sequencing note bet that §8.8's schema would make (b) tractable. Measured against merged reality it does not: only 33 of extractFrontmatter's 78 non-test call sites read STATE.md, and two of the five compensating mechanisms §8.1 lists survive real types, leaving ~31 lines across 3 call sites as the actual prize. Also corrects three claims verified false while answering it. §8.1's justifying sentence names #3349 and #3360 as defects a real parser would fix; both are already fixed on next, confirmed by executing the compiled parser rather than reading it. The guard roster calls lint-frontmatter-scalar-broad-grep.cjs an expected casualty of this rule, but it guards shell grep idioms in workflow bash fences and never touches our parser. The same roster calls lint-vendored-deps.cjs reusable as-is; it is hardcoded to re2js throughout. The last two were caught by applying the rule this amendment records -- a factual claim in this ADR is a hypothesis until the implementing phase executes it -- on its first use. Refs #3881 * docs(#3881): record that §8.1's fork is ill-posed and (a) is not implementable An adversarial pass on the Phase 4 design established by execution that extractFrontmatter is not a YAML parser but a line-oriented scanner whose output is a function of raw source text. Four spellings of the same value collapse to one js-yaml tree but produce four distinct legacy strings, one of them mangled. No adapter over a tree can choose among outputs the tree does not distinguish, so fork (a) -- keep a string-coercing adapter so the existing contract holds -- cannot be built. For any document with a non-scalar value, (a) collapses into (b); about 26 percent of frontmatter-carrying documents have one. Also records three design defects and one new attack surface, all confirmed by execution: catching a parse failure and returning {} would delete the frontmatter block on the next write at eight call sites that conflate empty with unparseable; an empty value yields null where legacy yields {}, and reconstructFrontmatter omits null-valued keys, so the shipped state template's empty progress key would vanish; the #1882 truncation probe is parseYamlRegion itself rather than a pre-parse heuristic, so it cannot both stay unchanged and survive that deletion; and FAILSAFE_SCHEMA still resolves aliases, expanding seven lines to 22.8 MB. The rule is not deferred. The measurement is the deliverable and the re-scoping is recorded as an open question with a forcing function, per section 8's own rule. Refs #3881 * test(#3881): failing-first rows for block scalars, unicode keys and the missing #3594 matrix Creates tests/feat-3594-parser-adversarial-frontmatter.test.cjs, the file the fixture README instructs contributors to register fixtures in but which never existed. Section C: table-driven ownership check over tests/fixtures/adversarial/frontmatter/ so a fixture with no matrix entry fails loudly; six existing fixtures (duplicate-keys, crlf-mixed, unclosed-block, unicode-keys-and-values, null-byte-value, huge-bounded) each get the invariant its README states. B1 blockScalarValueIsNotTheBlockIndicator: parsing commands/gsd/add-tests.md must give argument-instructions the instruction text, not the literal '|'. RED today. B2 blockScalarDoesNotInventATopLevelKey: same parse must not produce a top-level Example key scraped from inside the block body. RED today. B3 unicodeKeyRoundTripsAsIs: the 相 key in unicode-keys-and-values.md must survive parsing; today it is silently dropped. RED today. Refs #3881 * chore(#3881): vendor js-yaml and generalize the vendored-deps guard to a manifest Packaging step for ADR-3473 §8.1: makes js-yaml available to gsd-core/bin/** without promoting it out of devDependencies (promoting broke every installed tree, #3496). gsd-core/bin/lib/vendor/js-yaml.cjs is a verbatim copy of node_modules/js-yaml/dist/js-yaml.js (the self-contained UMD dist bundle, not index.js), exposing load/dump/FAILSAFE_SCHEMA/YAMLException with zero require() calls of its own. src/vendor/js-yaml.d.cts is hand-authored, not copied, because js-yaml ships no upstream .d.ts and @types/js-yaml is not installed. It is deliberately narrow, declaring only the four symbols in use, so anchors/aliases/custom types/loadAll are unreachable from typed code -- a compile-time enforcement of ADR-3473 §8.1's refusal to expand alias resolution for security reasons. Because it has no upstream counterpart it is excluded from the byte-compare. scripts/lint-vendored-deps.cjs is refactored from a script hardcoded to re2js into a table-driven VENDORED manifest (one row per package: upstream/vendored .cjs paths, optional .d.cts paths, twin kind upstream-verbatim vs hand-authored) so a second vendored package does not require a second hardcoded check block, per ADR-3473 §8.3 'one implementation per rule'. The four existing re2js checks (vendored .cjs vs node_modules, vendored .d.cts vs node_modules, src/vendor twin vs bin-side twin, devDependency version pin vs installed version) are preserved unchanged; verified pass/fail identical before and after the refactor, and the guard's ability to fail was re-proven with a deliberate one-byte append to both re2js.cjs and js-yaml.cjs, then restored. docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (via gen-inventory-manifest.cjs --write, run after build:lib) register vendor/js-yaml.cjs. gsd-core/bin/lib/vendor/README.md documents both vendored packages and the two twin kinds. Refs #3881 * feat(#3881): parse .planning frontmatter with the vendored js-yaml ADR-3473 §8.1: extractFrontmatter's read path is no longer a hand-rolled line scanner. parseYamlRegion, escapeDoubleQuoted, unescapeDoubleQuoted and parseQuotedScalar are deleted (not patched); parsing now goes through the vendored js-yaml (./vendor/js-yaml.cjs) under { schema: FAILSAFE_SCHEMA, json: true }. Everything js-yaml does not do is layered on top, in one place, carrying the seven design-doc consequences: 1. Empty value: a null js-yaml value is coerced to {} (matching legacy's own empty-value contract) so reconstructFrontmatter — which omits null-valued keys — still round-trips a bare `key:` line instead of deleting it. Verified live: progress: with no value survives parse -> reconstruct -> re-parse. 2. Unparseable no longer collapses to a bare {}: a new FRONTMATTER_UNPARSEABLE Symbol (exported), keyed exactly like the existing #3257 FULL_LINE_COMMENTS channel, is carried on the {} returned for malformed/refused YAML. Invisible to Object.keys/entries/JSON.stringify/for-in, so the 70 call sites that never inspect it are unaffected; wiring the 8 hasFrontmatter sites to consult it is a separate change, not done here. 3. Non-scalar object-list items (the four spellings of `- test: a b` that js-yaml collapses into one tree shape) are rendered as a canonical `key: value[, key2: value2]` string per item, keeping the existing array-of-strings value SHAPE. A full corpus differential over all 1702 tracked markdown files found 11 residual divergences from the legacy parser (enumerated in the PR/report), most of them the parser now being MORE correct (a dropped quoted top-level key, the block-scalar/phantom-key defect, a dropped Unicode key). 4. The #1882 truncation probe still runs the one real parser, but derives its key count from js-yaml's own thrown error and mark.line when the whole region doesn't parse cleanly (the dominant real truncation shape: fence opened, well-formed keys, no closing fence). Verified against both the clean-parse and the exception-fallback path. 5. The #3257 comment channel now attributes each pending column-0 comment against js-yaml's own parsed top-level key list (matched by literal key text, in document order) instead of the legacy ASCII-only key regex, so a comment above a Unicode key attaches correctly. 6. Anchors, aliases and merge keys are refused outright (a raw-text pre-scan, since FAILSAFE_SCHEMA still resolves them) — corpus occurrences today: zero. A 7-line billion-laughs fixture is verified refused rather than expanded. 7. A literal U+0000 is swapped for a private-use sentinel before the parse and restored in every resulting string afterward, since js-yaml rejects NUL unconditionally under every schema. escapeDoubleQuoted is deleted and reimplemented via js-yaml's dump() (forced double-quoted style), with control-char hex escapes lowercased to keep serialized output byte-stable (#1779 emitted lowercase); it keeps its exported name and signature for its two other call sites (commands.cts, runtime-artifact-conversion.cts), which need no change. frontmatterDeepEqual, the comment channel, sliceTopLevelFrontmatterSegments, regenerateFrontmatterKey's guard, noOpObjectListSetError and parseMustHavesBlock are all unchanged — retiring them is fork (b) and is not this phase. Refs #3881 * fix(#3881): quote template placeholders and preserve unparseable frontmatter SECURITY.md/UI-SPEC.md/VALIDATION.md wrote frontmatter placeholders as bare {N}/{phase-slug}/{date}, which is valid YAML flow-mapping syntax under the vendored js-yaml parser, not the literal placeholder text intended. Quote them so they parse as strings. Wire the FRONTMATTER_UNPARSEABLE Symbol (exported but unused) at the 8 call sites in state.cts/state-transition.cts that compute hasFrontmatter via Object.keys(extractFrontmatter(...)).length > 0 and reassemble the document without a frontmatter block when false. That check conflated 'no frontmatter' with 'unparseable frontmatter' (both parse to {}), so a document with a merge-conflict marker or refused alias in its frontmatter had that block silently dropped on write. Each site now preserves the exact raw bytes stripFrontmatter removed when the marker is set, leaving the genuinely-empty case unchanged. Refs #3881 * test(#3881): consequence and boundary coverage for the js-yaml migration Rows: A1 emptyValuedKeySurvivesAWrite, A2 unparseableDocumentKeepsItsFrontmatterBlock, A3 unparseableIsDistinguishableFromEmpty, A4 nonScalarValuesCanonicalize, A5 truncationProbeStillFiresOnAnOpenFence, A6 commentsStayOnTheirOwnKey, A7 anchorsAndAliasesAreRefused, A8 aliasExpansionCannotExhaustMemory, F1 UNTERMINATED_KEY_THRESHOLD boundary, F2 alias/nesting refusal bound, F3 frontmatter size boundary (huge-bounded.md + larger). Adds tests/fixtures/adversarial/frontmatter/anchor-alias-bomb.md and its entry in the feat-3594 fixture matrix. Refs #3881 * docs(#3881): document the vendored parser, correct a stale rationale, add a vendoring how-to Refs #3881 * docs(#3881): correct the frontmatter glossary entry Two errors in the entry as first written: it named parseYamlRegion as part of the read path when that function is deleted, and it recorded the eight hasFrontmatter call sites as unwired follow-on work when they were wired in e35ac2a2c. Also records the scope caveat that the CLI write path rebuilds the frontmatter block independently, so the marker binds at the transform layer. Refs #3881 * docs(#3881): record the semantic-migration decision and the counted guard ledger The maintainer chose the full semantic migration over splitting the rule into its own epic or patching the scanner, so section 8.1 is answered as "the fork was ill-posed and the migration is semantic" rather than as (a) or (b). Also replaces the pre-implementation guess that this phase would shrink the guard surface with the counted result: excluding vendored third-party lines the hand-maintained surface is net +307, and frontmatter.cts grew by 68 lines despite four functions being deleted, because the compatibility layer over js-yaml is larger than the scanner it replaced. Section 8.1's stated benefit is therefore not delivered as written; what improved is the kind of code maintained, not the amount. Decision 6 requires recording that rather than netting it away. Refs #3881 * chore(#3881): changeset for the vendored YAML parser migration Refs #3881 * test(#3881): golden parity, round-trip property and packaging coverage Refs #3881 * fix(#3881): refuse anchors structurally and fold in review findings ADR-3473 §8.1 review findings, addressed inline: Finding 1 (BLOCKER): refuseAnchorsAndAliases was a raw-line regex that matched only the bare-key spelling (key: &x). A quoted key ("a": &x), a flow mapping ({b: &x}) and a flow sequence ([&x, *x]) all define/use the SAME anchor mechanics while never matching that line shape, so the exact expansion the guard exists to stop went straight through unrefused (a 303-byte quoted-key bomb expanded to ~35.8MB). Replaced with js-yaml's own `load` `listener` callback, which reports `state.anchor` for every event belonging to an anchored node in every spelling, and throws from inside the callback to abort before any expansion (~1-2ms vs full expand-then-discard). A merge key with an alias is still refused (merge always requires a previously anchored node, so the alias itself trips the listener); a bare merge key with NO alias is no longer separately refused, documented as intentional: FAILSAFE_SCHEMA never resolves `!!merge`, so it carries no expansion risk. Table-driven tests added for all four bypass spellings + merge key, plus a quoted-key-spelled billion-laughs fixture registered in the adversarial matrix and README. Finding 2: src/vendor/js-yaml.d.cts's docblock falsely claimed anchors/ aliases were "simply UNREACHABLE from typed code" through the twin. Corrected to state the truth: anchor/alias resolution is document-level `load` mechanics reachable through exactly the declared surface, and refusal is enforced at RUNTIME (Finding 1's listener), not by the type surface. Finding 3 (MAJOR): the null-byte sentinel (U+E000) round-trip was non-injective — restoreNullBytesDeep rewrote every U+E000 in the parsed tree back to NUL, including one the document author legitimately wrote, silently corrupting it. Now refuses outright whenever the raw region already contains U+E000 (consistent with the existing anchor/merge-key refusal path), making the substitution provably injective. Tests added for a real NUL alone (preserved), a pre-existing U+E000 alone (refused, not corrupted), and both together (refused, not merged into one byte). Finding 4 (MAJOR): scripts/lint-vendored-deps.cjs's `srcTwin` field was dead for a hand-authored row (only read inside the upstream-verbatim branch) — exactly how Finding 2's stale docblock drifted unnoticed. Added checkHandAuthoredTwin: every value-level export the twin DECLARES must be an actual own property of the vendored runtime module at require-time. Tests added, including a sensor that a declared-but-nonexistent export IS caught. Finding 5: the existingFm/hasFrontmatter/stripFrontmatter/fmPrefix/ unparseableFm/reassemble preamble, copy-pasted at 7 sites in state-transition.cts plus a sixth hand-inlined copy in state.cts's cmdStateCompletePhase, is now one exported helper (beginFrontmatterReassembly) every site routes through, including the hand-inlined one. Three call sites (beginPhaseCore, patchCore, updateCore) keep a literal `body = stripFrontmatter(content)` assignment alongside the helper call so scripts/lint-state-write-path-drift.cjs's single-hop backward scan (which does not chase aliases) still sees the strip; stripFrontmatter is pure/idempotent so the extra call changes nothing observable. Finding 6: corrected the frontmatter.cts docblock's stale "wiring is a separate change" claim (the 8 call sites are wired on this branch) and the changeset's backlink from (#3473) to (#3881). Finding 7: fixed the lint:ci failures blocking the gate — an @typescript-eslint/only-throw-error violation from throwing a bare Symbol as the anchor-detected signal (now a real Error subclass), unused-var warnings left over from the Finding 5 refactor, a lint-test-file-count cap exceeded by two migration-specific test files (allowlisted with justification), and the lint-state-write-path-drift false positive from Finding 5's helper (fixed above). tests/frontmatter-golden-parity.test.cjs:117's execFileSync already carried an explicit timeout; no change was needed there. Golden fixture: added a golden entry for the new anchor-alias-bomb-quoted.md fixture ({} — matches what the legacy line scanner would also produce, since it independently dropped every quoted top-level key). No other corpus document diverges: real .planning/ documents carry zero anchors/aliases/merge keys/U+E000 today. Refs #3881 * fix(#3881): fold in second-round review findings Finding 1 (BLOCKER): tests/frontmatter.test.cjs pinned the pre-migration ASCII-only key regex for the Unicode fixture; updated to require the 相 key's value now that js-yaml has no such restriction. Audited the rest of the file for other pre-migration pins (block scalars, quoted keys, flattened values, empty values, duplicate keys, unclosed blocks, null bytes) by execution against real fixtures; found none regressed. Finding 2: parseYamlRegion and escapeDoubleQuoted renamed to parseGuardedYamlRegion and escapeDoubleQuotedScalar in src/frontmatter.cts so no function still answers to the deleted hand-rolled scanner's name (ADR-3473 §8.1 "deleted, not patched"). escapeDoubleQuotedScalar's three external call sites (src/commands.cts, src/runtime-artifact-conversion.cts) updated in the same change — a mechanical rename, not an ADR-amendment matter. Finding 3 (BLOCKER): fixed a real crash and a silent data-loss bug found by execution. A top-level key named constructor/__proto__/toString/ valueOf/hasOwnProperty crashed reconstructFrontmatter (bracket read resolving an inherited Object.prototype member); a key literally named __proto__ was silently DROPPED entirely (bracket assignment on an ordinary {} invoked the inherited __proto__ setter instead of creating a data property). Fixed by building every parsed Frontmatter object with Object.create(null), and replacing an `in` check with hasOwnProperty.call in propagateCommentChannel. Added round-trip tests for all five hostile keys, each with its own leading comment. Finding 4 (MAJOR): escapeDoubleQuotedScalar's docstring falsely claimed full byte-stability across the migration. Verified by execution: BEL/NUL/ NEL/NBSP/LS/PS/BOM now emit YAML-named escapes instead of the old hex/raw- literal forms. Proved round-trip equivalence (each escape re-parses to the exact source codepoint) and corrected the docstring. Found and fixed a related real defect while verifying: a lone UTF-16 surrogate was emitted BARE (scalarNeedsDoubleQuoting didn't trigger), producing genuinely unparseable YAML that silently collapsed to {} on re-read — extended scalarNeedsDoubleQuoting to route surrogates through the quoted+escaped path. Finding 5 (MAJOR): countKeysBeforeTruncation went silent on 4 real truncation shapes (unquoted colon, open flow collection, mis-indented sibling key, refused anchor). Root cause: the mark-based prefix recovery excluded the very line whose key needed counting, and a mark-less refusal never entered the recovery branch at all. Fixed by taking the max of two lower bounds: the longest parser-verified line-prefix, and a raw-text count of key-shaped lines (reusing the same key-shape pattern this file already uses for isFrontmatterShaped). Extended test-matrix row A5 table-driven over all 4 regressed shapes. Finding 6: the design doc's claim that no test owned the #3594 adversarial fixture corpus was false — consolidation epic #1969 had already folded it into tests/frontmatter.test.cjs. An earlier commit on this branch re-created a standalone duplicate under that false premise; folded its genuinely-new coverage (fixture-ownership check, anchor-bomb fixtures, block-scalar B1/B2 rows) into frontmatter.test.cjs and deleted the duplicate file. Corrected the false claims in 40-design.md §3.3.1 and the ADR's §8.1 note, including the roadmap-sibling claim (no such file exists). Finding 7: the golden serializer sorted object keys, making it structurally blind to the key-order-parity invariant ADR-3473 §8.1 actually claims. Made it order-preserving and regenerated the golden fixture from a standalone compile of the legacy (pre-#3881) parser at ddde001af; the current parser matches it with zero undocumented divergences, confirming key-order parity genuinely holds. Extended row A2 table-driven across 6 of the remaining 7 transitionCore kinds (all pass) plus documented, by execution, a newly-discovered 8th-site regression: state.cts's cmdStateCompletePhase calls the same preservation helper but its result is clobbered by a later unconditional resync — filed as a distinct finding rather than fixed here (touches syncAndPreserveStateMd, outside this change's verified scope). Refs #3881 * fix(#3881): preserve unparseable frontmatter through the CLI write path Characterization (executed, before/after shown): case (b), not (a). The frontmatter FENCE survives — `state complete-phase` on a conflict-marked STATE.md returns success and a well-formed, freshly-derived frontmatter block, not a document with no frontmatter at all. But the block's actual content (the merge-conflict markers, and with them any signal to a human that the document was in conflict) is silently discarded and replaced. Root cause was two clobber sites, not one: 1. syncStateFrontmatter (src/state.cts) re-parses the already-preserved `transformedContent` from readModifyWriteStateMd, finds {} + the FRONTMATTER_UNPARSEABLE marker, and unconditionally rebuilt a fresh frontmatter block from the body anyway. 2. Even after (1) is fixed, applyPostSyncPreservation's own postFm/applyStatePreservation/authoritativeFm-reassertion machinery re-extracts frontmatter from syncedContent, restores curated fields from the pre-write snapshot, and reconstructs a NEW block again — confirmed live via `state begin-phase`, which still lost the markers after fixing (1) alone. Both are now guarded by the same predicate (isUnparseableFrontmatter, checking FRONTMATTER_UNPARSEABLE): when the ORIGINAL frontmatter did not parse and the caller is not on ADR-3408 §8.3's closed "body wins" list, both functions return their input content unchanged rather than re-deriving over it. The closed list (cmdStateSync #905, /gsd-health --repair's REGENERATE_STATE, both routed only through writeStateMd, which never reaches applyPostSyncPreservation and passes sanctionedPermanentEmptyFallback=true to syncStateFrontmatter) is untouched — neither widened nor narrowed; verified by execution that `state sync` still overwrites the conflict-marked block exactly as before. Other verbs sharing the same readModifyWriteStateMd path were checked and were equally affected before this fix: state update, query state.patch, and state begin-phase all lost the conflict markers (RED, shown by execution), and all three now preserve them (GREEN). Covered table-driven in tests/feat-3881-yaml-parser-consequences.test.cjs's new A2b describe block, which drives the real CLI verbs via runGsdTools — not just the pure transitionCore layer the earlier A2 rows exercised — plus a control asserting state sync's body-wins contract is unchanged. Refs #3881 * fix(#3881): restore the parse surface's prototype and fix remote-runner failures Root cause of the bulk of the 88 remote-runner failures: extractFrontmatter/parseGuardedYamlRegion handed back Object.create(null) trees for prototype-pollution safety, but assert.deepStrictEqual compares prototypes, so every assertion against a plain object literal failed (57 frontmatter.unit.test.cjs + 5 frontmatter.test.cjs + others). Fixed by keeping the internal construction null-prototype (unchanged) and converting to a plain-prototype tree via Object.defineProperty (never bracket assignment, so __proto__/constructor/toString keys stay safe) at the parseGuardedYamlRegion/unparseableResult return boundary only; the internal FULL_LINE_COMMENTS Symbol channel is copied by reference, not recursed, so its own __proto__-safety is untouched. Per-class fixes: (1) bomAcrossArtifactTypes was the same prototype bug, no separate code change needed. (2) frontmatter-cli #1660: added objectListFieldWouldLoseData, a broader lossy-field detector alongside the existing byte-identical noOpObjectListSetError -- js-yaml's flattenObjectListItem now correctly includes every sub-key of an object-list item (a real bug fix over the legacy scanner, which silently dropped every field but the first), so a set that drops that now-included data is no longer byte-identical to the original and needs its own guard. (3) uat.test.cjs: updated the pinned expectation for the human_verification quote-stripping artifact -- js-yaml resolves quoting correctly where the legacy regex left an unbalanced quote; documented as an intentional, non-lossy behavior change. (4) smart-entry: added a fallback-only loadWithAmbiguousColonRepair so a column-0 key: value line whose value itself contains an unquoted colon (the #2571 hand-edited-STATE.md shape) round-trips instead of failing the whole frontmatter block closed. (5) frontmatter.unit.test.cjs bracket-array leniency: added a second fallback, repairMalformedInlineArrays, restoring the legacy scanner's tolerant inline-array handling (consecutive/blank commas, unclosed bracket) -- both repairs run ONLY after the primary parse already threw, so well-formed documents are unaffected. (6) prompt-injection-scan: src/frontmatter.cts had a literal U+FEFF BOM embedded in a comment illustrating the #2977 fix; replaced with the U+FEFF text escape. (7) eslint-glob-coverage: allowlisted the new src/vendor/js-yaml.d.cts vendored type declaration, same precedent as the existing re2js.d.cts entry. (8) frontmatter-golden-parity: git ls-files *.md now runs with -c safe.directory=* (process-scoped) so it survives the remote runner's dubious-ownership check without a persistent git config write. Refs #3881 * chore(#3881): backfill changeset PR number Refs #3881 * test(#3881): make golden parity resistant to unrelated tree churn A corpus-wide snapshot keyed to every tracked *.md file was coupled to mutable-by-design files: .changeset/*.md's pr:0 -> real-PR-number backfill is a required workflow step, not a parser change, yet it turned this suite red. Training people to 'just regenerate the golden' on that kind of failure defeats the point of the snapshot. Exclude .changeset/** from the golden corpus entirely, tolerate tracked *.md files with no golden entry (they postdate the capture) instead of failing on them, keep hard failures for a golden entry whose file has vanished from the tree and for any real parity divergence, and add a coverage floor so the enumeration cannot quietly degrade to comparing a handful of files. Golden regenerated by recompiling the legacy pre-migration parser (git show ddde001af:src/frontmatter.cts) standalone, independent of the current parser, over the same non-changeset corpus. Refs #3881 * test(#3881): make the parser golden hermetic instead of tree-keyed This repo merges ~21 commits/day; a 14-day sample measured 937 touches of the exact files (commands/gsd/*.md, gsd-core/workflows/*.md, agents/*.md, docs/*.md) the prior golden pinned by tracked path. Any PR editing one of those files' frontmatter for reasons unrelated to the parser (an argument-hint addition, an allowed-tools tweak) turned the suite red, and the reflex fix -- "regenerate the golden" -- overwrote the very snapshot meant to catch a real regression. Excluding .changeset/** was not enough; the design itself was wrong: a regression fixture must not be keyed to mutable repo paths, and a single 376-entry JSON every such PR touches is also a guaranteed merge-conflict surface. Rebuilt the fixture to carry its own documents: each of 51 entries stores a stable id, literal documentText (shrunk from a real ddde001af-era corpus document), and an expectedParse captured independently from the pre-migration legacy parser (git show ddde001af:src/frontmatter.cts, compiled standalone against its byte-identical sibling modules). The test reads no tracked path, shells out to no git command, and enumerates no tree -- a PR editing commands/gsd/help.md cannot affect it. Every entry's reconstruction was verified at capture time to reproduce both the current and legacy parser's output on the original document; 0 of 51 candidates were dropped by that check (1, the deliberately-unterminated unclosed-block.md adversarial fixture, has no closing fence to truncate at and is stored unshrunk). Kept the 5 documented DIVERGENCES rows (now diverges:true entries) and the D2 order-preserving structural serializer that keeps the comparison from passing vacuously; dropped the tree-enumeration helpers, the coverage floor, the post-capture-skip logic, and the vanished-file check -- all artifacts of the path-keyed design. Refs #3881 * fix(#3881): resolve vendored-deps paths independently of cwd shape Five rows in tests/lint-vendored-deps-manifest.test.cjs failed on windows-latest CI: the test passed absolute scratch-file paths into compareFiles()/checkRow(), whose helpers joined every input onto ROOT via path.join(ROOT, rel), producing garbage when the input was already absolute. It surfaced on windows-latest specifically because GitHub's Windows runners checkout the repo on a different drive than TEMP, so path.relative(REPO_ROOT, tmpFile) returned the absolute path unchanged (no relative traversal is representable across drives) rather than the relative form the test assumed. The remote gsd-test runner this repo gates pushes on is Linux-only and could never have caught this; GitHub CI's windows-latest job is the only signal that does, and it did. Fixed the helper itself (scripts/lint-vendored-deps.cjs's new resolvePath()) to treat an already-absolute input as absolute-in, absolute-out instead of silently mis-joining it, and updated the test to pass the scratch file's absolute path directly rather than relying on a relative conversion that is not always representable. Kept every mutation-sensor assertion intact and added coverage proving resolvePath is a no-op for relative inputs and correctly passes absolute ones through unchanged. Refs #3881 * fix(#3881): warn when state sync regenerates over unparseable frontmatter state sync (ADR-3408 §8.3's sanctioned regenerate path) correctly overwrites an unparseable frontmatter block per its 'body wins' contract — that overwrite behavior is unchanged here. The defect was the silence: synced:true/exit 0 gave no signal that the existing block (including git merge-conflict markers) could not be parsed and was destroyed, per ADR-3473 §8.5 ('a derived conclusion may not be reported as authoritative when the derivation dropped input it could not resolve') and §8.4 ('failure is a value'). Adds a gsd: warning — ... (#3881) line on stderr, matching the existing #3573 precedent, and surfaces the same disclosure in the JSON result's existing changes[] array so a machine consumer sees it too. Exit code and synced:true are left unchanged — sync did what its contract says. REGENERATE_STATE (/gsd-health --repair's sibling on the same sanctioned-regenerate list) is DESTRUCTIVE-risk and unconditionally refused by applyRepairs's dispatcher before runRepairAction ever runs (src/health-diagnostic.cts), so it is not a live path today and is not in scope for this fix. Refs #3881 * fix(#3881): exit non-zero when a state command returns an error Refs #3881 * chore(#3881): changeset for the state exit-code fix Refs #3881 * fix(#3881): honor the documented --project-dir flag Refs #3881 * revert(#3881): restore exit-0 result envelopes for state errors Reverts 9638f2936 and its changeset. The change was wrong and the revert is the correction. This repo distinguishes two error mechanisms deliberately. error() in src/io.cts writes to stderr and calls process.exit(1) -- the hard-failure path. output({error: ...}) writes a JSON result envelope to stdout and returns normally with exit 0. The reverted commit converted 23 result-envelope sites into hard failures, which is a different contract, not a bug fix. tests/state-contract.test.cjs's errorPathDoesNotPublish asserts the envelope contract directly -- a failing command exits 0 with a JSON error envelope and must not publish state.json -- and the remote matrix run caught it along with four cases in the QA scenario walk. Thirteen tests in tests/state.test.cjs that the original commit rewrote were encoding that real contract, not the bug it claimed; they are restored. Whether an error envelope on stdout with exit 0 is the right CLI design is a genuine question, and it is section 8.4's rule ('failure is a value') with its own phase. It is not something to flip inside this PR. Refs #3881 * chore(#3881): backfill changeset PR number for the project-dir fix Refs #3881 * test(#3881): keep the frontmatter mutation shard inside its time budget The Stryker (frontmatter) shard hit the documented 15-minute (900s) shard cap. Root cause is NOT row-level spawn overhead (contrast the #2790/ core-utils precedent): the three shard test files' own logic runs in ~413ms total (356+30+27ms) with all 392 assertions passing. Instead, src/frontmatter.cts grew from ~825 to 1496 lines (+671/-187) migrating to the vendored YAML parser, proportionally growing the mutant count Stryker generates for gsd-core/bin/lib/frontmatter.cjs. Stryker's command runner bills the full 'node --test <3 files>' invocation once per mutant, and node:test's default per-file process isolation forks a child process for each of the three files on every one of those invocations — pure fork overhead multiplied by a much larger mutant population. Fix: scripts/mutation-matrix.cjs COVERED.frontmatter now declares isolation: 'none', and .github/workflows/mutation.yml passes --test-isolation=${{ matrix.isolation }} (defaulting to 'process' — i.e. unchanged behavior — for the other 8 shards, which were not individually audited for cross-file state leakage under shared-process execution). Measured locally via node:test's run() API on the exact 3-file set: isolation:'process' took ~593ms vs isolation:'none' ~478ms for the same 392 passing assertions. The true CI-shard number can only be confirmed on the GitHub Actions run (Stryker cannot run locally, and 'node --test' is hard-blocked in this environment). Refs #3881 * test(#3881): register the vendored-parser tests in the frontmatter mutation shard stryker.config.mjs's own rule ("Keep this list in sync with the tests arrays in scripts/mutation-matrix.cjs COVERED") was violated: #3881 grew src/frontmatter.cts from ~825 to 1496 lines but its new tests (tests/feat-3881-yaml-parser-consequences.test.cjs, tests/frontmatter-golden-parity.test.cjs, tests/frontmatter-roundtrip.property.test.cjs, and +167 lines in tests/frontmatter.test.cjs) were never added to the frontmatter shard's tests array, so Stryker's mutants in the new vendored-js-yaml adapter had nothing constraining them. PR #3888 measured 55.8% against the 65 floor (748 killed / 593 survived / 17 timeout) and the shard was separately cancelled at 15m04s against the 15-minute per-shard cap. Registers all four files (each earns its slot on evidence of a unique constraining assertion, documented inline), gives the shard a measured/projected 180-minute budget via a new per-module timeoutMinutes field threaded through mutation.yml's job-level timeout-minutes the same way isolation is threaded, and removes the prior isolation:'none' override (re-measured at this file-set size, its savings are within run-to-run noise, not worth the unaudited cross-file-state-leakage risk). Refs #3881 * feat(#3881): derive the mutation test list and ratchet the score floor Refs #3881 * test(#3881): ratchet five stale mutation floors and close the frontmatter gap Raised five module minScore floors per CI run 33012034388 (floor(achieved)-1): config-schema 75.51%->74, prompt-budget 88.95%->87, context-composer 79.92%->78, context-utilization 92.31%->91, active-workstream-store 87.42%->86. Updated both scripts/mutation-matrix.cjs COVERED entries and tests/mutation-matrix-ratchet.test.cjs RATCHET_BASELINE in the same diff per the ratchet's own contract. Closed the frontmatter shard's 63.03%-vs-65 gap with new behavioral tests in tests/feat-3881-yaml-parser-consequences.test.cjs, each paired with a documented near-miss: frontmatterDeepEqual's array-order/length/type-mismatch/key-order semantics (via spliceFrontmatter's no-op guard), scalarNeedsDoubleQuoting's leading/trailing-whitespace and dash/surrogate triggers (via reconstructFrontmatter), repairAmbiguousColonValues' already-quoted vs ambiguous-colon repair paths (via extractFrontmatter), and the null-byte sentinel round-trip surviving at region offset 1. Did not lower minScore. Refs #3881 * test(#3881): decouple the ratchet test from real module floors The CLI end-to-end rows in tests/mutation-score-ratchet.test.cjs hardcoded config-schema's real floor (52), which commit 973321541 legitimately ratcheted to 74 -- breaking a test pinned to the exact value the mechanism under test exists to change. Add an injectable --matrix seam to scripts/check-mutation-score-ratchet.cjs and point the CLI rows at a synthetic module + synthetic floor built via a temp fixture, so the rows are indifferent to any real module's floor moving while still exercising the same fail/pass behaviour. Refs #3881 * refactor(#3881): parse must_haves with the vendored parser and drop re-implemented leniency Refs #3881 * fix(#3881): restore the ambiguous-colon repair its hand-edited-STATE.md contract needs A tracked-document sweep of 910 *.md files cannot see this dependent: repairAmbiguousColonValues's one real caller is user hand-edited STATE.md content that never lives in this repo's tree, only on end users' machines, and is pinned by tests/smart-entry.unit.test.cjs. Restores the function plus its post-throw fallback path (loadWithAmbiguousColonRepair) only; repairMalformedInlineArrays and splitLegacyInlineArrayItems stay deleted, reverified against the full frontmatter test shard. Adds a frontmatter-level regression row in tests/feat-3881-yaml-parser-consequences.test.cjs so the dependency is visible where the function lives. Closes #2571 Refs #3881 --------- Co-authored-by: sim <sim@local> |
||
|
|
382bf7c423 |
fix(#3706): deliver the resolved reasoning effort to OpenCode subagents (#3867)
* test(#3706): failing-first coverage for OpenCode variant emission and frontmatter escaping * fix(#3706): emit the resolved reasoning effort as OpenCode's variant key `query resolve-execution` resolved an effort level for every agent, but the OpenCode bake wrote only `model:` — the effort never reached the generated agent, so subagents ran at whatever the runtime defaulted the model to. This is the effort-side twin of the model-side defect fixed in #3705. The key is written only when an `effort` block is actually configured. `resolveInstallTimeEffort` always returns a level (the catalog default is `high`), so gating on its return value would stamp `variant: high` into every existing OpenCode install — and OpenCode resolves a variant name against a `variants` map in the user's `opencode.jsonc`, so a value nobody declared is not a safe default. Gating on `readGsdEffectiveEffortConfig` keeps installs that never asked for effort routing byte-identical. Kilo does not receive the key: `EFFORT_ARGV` declares surfaces for claude, opencode and codex and has no kilo entry. This is deliberately asymmetric with the model side, where #2794 J8 requires the two runtimes to resolve alike. Both frontmatter sinks now route through `frontmatterScalar`, which quotes and escapes any value that is not a plain scalar. The raw interpolation predates this change, but it was already shown by execution during the #3705 security review to let a config value containing a newline inject additional top-level keys (`tools:`, `permission:`) into a generated agent file. This change adds a second write to that sink, so it is closed here rather than doubled. * fix(#3706): quote frontmatter values YAML would not read back verbatim Self-review of the predicate added in the previous commit. Treating /^[A-Za-z0-9._:/@+-]+$/ as 'safe to emit bare' answers the wrong question: a value can match it and still not round-trip. - A leading '@' is a YAML *reserved* indicator and may not open a plain scalar at all, so a scoped ID like '@org/model' emitted bare is a parse error, not an ambiguity — the whole agent file becomes unreadable. - 'no' / 'y' / 'off' / 'null' resolve to booleans and null, so a variant with one of those names would match no entry in the user's variants map. - '12:30' resolves to 750 under YAML 1.1 sexagesimal, and ':' is legal mid-identifier here, so the form is reachable rather than contrived. Real model IDs pass every clause and stay bare, so already-generated files remain byte-identical. * fix(#3706): route variant through the declared effort seam and cover the live path Addresses six findings from the isolated review, all confirmed by execution. The tests were the serious one: they required `../bin/install.js` while the fix landed in src/, which compiles to gsd-core/bin/lib/. They exercised a different copy of the converter than the one the bake actually uses, so the whole suite was green-by-construction against unchanged code and the remote run failed all 13. Every case now runs against BOTH copies from one table, which doubles as the parity assertion the generative-fix note in runtime-artifact-conversion.cts asks for, and bin/install.js carries the mirrored change. Emission no longer hand-rolls the value. It goes through `renderEffortArgv`, the declared OpenCode effort seam (EFFORT_ARGV.opencode: its own supported set and clamp). That is what rejects a level that is not a wire value — above all `inherit`, which per #3533 (10d) means "omit the key and follow the host default" and was previously written literally, naming a variant that cannot resolve. Reachable two ways, both now pinned: an agent_overrides entry and a routing_tier_defaults entry. A bare effort.default does NOT reach a tiered agent (the #3531 tier ladder answers first), so a test written against `default` alone asserts nothing — that is pinned too. The plain-scalar decision moved into frontmatter.cts beside `scalarNeedsDoubleQuoting` rather than sitting next to it as a second, weaker predicate. `agentScalarNeedsDoubleQuoting` is a documented superset: it adds a trailing `:` (read as a nested mapping key, which fails the whole frontmatter), boolean/null words, and numeric-looking values including YAML 1.1 sexagesimal. Docs now state the cascade plainly: the gate is on effort being configured at all, not on the individual agent being named, so every generated OpenCode agent gets a variant line once any effort block exists. * test(#3706): assert the two frontmatterScalar copies cannot diverge A hand-picked adversarial corpus plus a fast-check property over YAML-significant strings, both run against bin/install.js and the live src copy. Verified the property can actually fail: mutating one copy's quoting rule is killed well inside the run budget. * fix(#3706): close the review findings — predicate, seam, and dead mirror Third review round; every item below was confirmed by execution. The scalar predicate was wrong in two families, both found by a round-trip property test rather than by reading. Basing it on scalarNeedsDoubleQuoting dropped the "first character must be alphanumeric" clause, so `~`, `.inf`, `.nan`, `+1`, `-0` and `.5` went out bare and came back as null/floats/ints; and that base predicate only inspects the FIRST character, so an embedded `: ` (a nested mapping, i.e. a parse error) or ` #` (a comment, i.e. silent truncation) also passed. Dates round out the set: `2026-08-25` opens alphanumeric, survives every other clause, and YAML resolves it to a Date. The property now asserts the contract directly over generated values instead of trusting an enumerated character list. The bin/install.js mirror is gone. Its premise was false — install.js already requires bin/lib at :65 — and it was unreachable besides: install.js's convertClaudeToOpencodeFrontmatter has no `isAgent: true` call site, because its agents path resolves converters from the compiled module. It was a third copy of the YAML rules serving a test rather than a caller, so the file is back to origin/next and the tests target the live copy only. Effort clamping moved to `clampEffortForHost`, which renderEffortArgv now delegates to. The layout was calling renderEffortArgv with a hardcoded 'argv' to borrow its clamp, which read as if the frontmatter key were gated on the invocation-time axis. It is not: claude declares effortSurface "argv" and independently bakes an effort: key. One capability table, one clamp, two channels that no longer pretend to be each other. Also corrects an earlier claim of mine: adding EFFORT_RENDERING.opencode would NOT have made `effort sync` write the wrong key, because it guards on the runtime name before it ever renders. The seam choice stands on other grounds. `effort sync` still skips OpenCode, but its stated reason claimed OpenCode "does not use effort: frontmatter", which this change makes false — so the message now says what is actually true. * docs(#3706): restate the changeset around the round-trip contract * fix(#3706): restore the changeset fragment belonging to #3809 An earlier commit in this branch picked the first file in .changeset/ by glob order instead of the fragment created for this issue, and overwrote agile-geese-squeak.md (PR 3815 / #3809) with this change's body. Restored verbatim from origin/next; this change's text now lives in its own patient-cranes-parade.md, where it was created. * feat(#3706): maintain the OpenCode variant key from effort sync Install bakes the resolved effort into OpenCode agent frontmatter as `variant:`, so `effort sync` has to maintain it or a config change only takes effect on reinstall — and its skip message claimed OpenCode does not use frontmatter effort at all, which this issue made false. cmdEffortSyncOpencode mirrors the codex branch: resolve per agent, clamp through the declared OpenCode capability, then write, strip, or skip. A null target means the key must not exist, which covers both "no effort configured" and "resolved to inherit or to an unsupported level" — the same states under which install writes nothing, so sync and install agree by construction. The frontmatter line-editors are key-parameterised rather than copied: setEffortFrontmatter / removeEffortFrontmatter are now thin wrappers over the same internals the variant path uses, and a test pins that the claude `effort:` behavior did not move. The child-process test harness fixes both HOME and USERPROFILE, so the hermetic-config assertions cannot pass vacuously on Windows. * fix(#3706): scope the frontmatter line editors to the matched block Found by the security review of the sync path, reported as correctness rather than vulnerability, and reproduced against pre-fix code before being fixed. Both editors matched the frontmatter with a regex that can match a block after a preamble, then derived the EOL and the opening-fence length from the START OF THE FILE. On a CRLF document with a preamble those disagree, the offsets shift by one byte, and the reassembled document comes back with a mangled fence (`---\rname: x`). Both now take the EOL from the matched block. `setFrontmatterKeyLine` additionally did a whole-file `/m` replace when the key already existed, gated only on the key being present in the frontmatter body — so a preamble line starting with the same key was rewritten instead of the frontmatter one. It now replaces inside the frontmatter span only, which is the hazard `removeFrontmatterKeyLine` already documented and guarded against. Neither is reachable from an install-written `gsd-*.md` (those begin at byte 0 with `---`), and both predate this change — but the editors are in this diff because #3706 key-parameterised them, so they are fixed here rather than left for the next caller to trip over. Three regression tests, each confirmed to fail against the pre-fix build. * fix(#3706): treat a present-but-empty key as present, and pin the real seam Fourth review round. The MAJOR one: both sync branches read the current value with `(.+?)`, which needs at least one character, so a key present with an EMPTY value read as "key absent". When the target was also null the code concluded "already correct" and skipped — leaving the key in the file, where it reads back as YAML `null`: exactly the unresolvable-variant state this change exists to prevent. Whitespace decided whether it fired, since `variant: ` matched and `variant:` did not. Presence and value are now separate questions at both the opencode and the claude branch. The OpenCode writer now follows the codex branch rather than the claude one: tmp file plus retryRenameSync with orphan cleanup, and a write failure skips that agent and is reported instead of aborting the sweep. Same granularity, same transient-Windows-lock exposure, so the hardened sibling was the right precedent. Also: the generic line-editors escape their interpolated key, the JSDoc stranded by the clampEffortForHost extraction is back on renderEffortArgv, and a cast that declared a nullable function as non-nullable is corrected. Tests close the gaps the review listed — empty value (both spellings), CRLF round-trip through write and strip, the symlink guard, a body line starting `variant:`, a file with no frontmatter, and the YAML classes that actually broke the predicate. The new layout-seam test drives the real stage() path and was verified to FAIL when `variant` is removed from the converter call; a seam test that survives cutting the seam is worse than none. * fix(#3706): clear the round-five review findings No blockers or majors this round; the repo's review gate is zero-tolerance, so the minors are cleared too. A duplicated key was only half-stripped: the strip regex had no `g` flag, so a frontmatter carrying the key twice lost one occurrence, reported success, and left the "a null target means the key must not exist" invariant false on disk — converging only on a second run. Such a document is already invalid YAML, so this is robustness rather than a live corruption path, but a successful sync has to leave the invariant true. A run in which every write failed still summarised as `ok`, so a caller could not tell "nothing to do" from "everything failed". The OpenCode branch now reports `failed` when any write failed. The write-failure path was also the newest code in the change with no coverage at all; it now has a test that injects the failure by monkeypatching the write, per CLAUDE.md §4, rather than by chmod — mode bits do not bite under root in CI. `CodexEffortSyncWriteFailure` is renamed `EffortSyncWriteFailure` now that two branches share it. Removed a guard on the claude concrete path that was provably unreachable — no member of EFFORT_SET renders null there, so it read as protection that did not exist. The claude inherit path's presence check is load-bearing and untouched. Three stale statements corrected: the OpenCode result shape matches codex's, not claude's, now that it emits write_failures; the `thread()` test helper now calls `clampEffortForHost` so it genuinely mirrors the layout instead of merely claiming to; and a test helper restored `USERPROFILE` by assignment, writing the literal string "undefined" into the environment on POSIX — it deletes now. * fix(#3706): converge the set path, degrade on unreadable files, preserve mode Rounds five and six of review. No blockers or majors; the review gate is zero-tolerance, so the minors are cleared too. `setFrontmatterKeyLine` was the mirror of a defect already fixed in its sibling: `remove` was made global, `set` was not, so on a frontmatter carrying the key twice it rewrote the first and left a stale second. Last-wins YAML readers honour the stale value while the sync's own first-occurrence read reports "in sync" — permanently non-converging. It now collapses to exactly one occurrence, in the position of the first, so ordinary single-occurrence documents stay byte-identical (verified across seven shapes before and after). An unreadable agent file used to throw and abort the entire sweep, while a failed WRITE in the same loop degraded into a report. The OpenCode branch now reports read failures alongside write failures; the claude branch degrades to a skip without a new result field, because its shape is long-standing and widely consumed and one bad file aborting the sweep is the actual defect. The tmp+rename publish dropped the original file's mode — a plain writeFileSync preserves it, a rename does not — so a 0600 agent came back 0644. Both the OpenCode and the codex branch now carry the original's permission bits across the publish, masked with 0o7777: the raw stat mode includes the file-type bits, and POSIX leaves those unspecified for chmod. Linux is the only OS the remote matrix runs, so relying on Darwin's tolerance would have been untestable here. Also documents the `from` contract on EffortSyncChange (null means the key was absent, '' means present with an empty value — a distinction earlier rounds introduced and then collapsed in the output), adds OpenCode to the docs paragraph enumerating where the key is omitted under inherit, and records in a comment that the 'failed' summary reaches only raw mode and does not change the exit code, which is a CLI-contract change affecting all three branches and is deliberately not made here. * fix(#3706): guard the codex read, close the tmp permission window, rename the failure type Round seven, plus one thing I found myself. `cmdEffortSyncCodex` still had an unguarded `fs.readFileSync` — a read fault on one agent exited 1 and aborted the whole sweep. The claude and opencode branches were both guarded earlier this round and codex was missed, with the unguarded read sitting ten lines above the chmod block the previous commit did edit. It now reports read failures the way the OpenCode branch does, and a read failure flips its summary to `failed` — which write failures did not do there either, so both are corrected for consistency. The tmp file was created at the default mode and only tightened afterwards, so a 0600 agent's contents sat in a 0644 file for the length of the publish. I measured the window rather than assuming it, then closed it by passing the mode at creation. The chmod after the write is deliberately RETAINED and commented: the `mode` option only applies when the file is actually created, so a leftover tmp from an earlier crashed run would be truncated and reused at its old mode, and the chmod is what corrects that. `EffortSyncWriteFailure` is renamed `EffortSyncFileFailure` — it was typing a `read_failures` array, the same naming-lie the `Codex…` prefix had last round. Also pins the codex mode preservation with a test. It only writes on a path that genuinely rewrites the file, so the fixture is an Anthropic-flavoured model pin the sync strips, and the test asserts the content changed before checking the mode — otherwise it would pass on a sync that did nothing. * fix(#3706): guard the claude writes and share one escaping rule The security sign-off caught a comment of mine that was factually wrong: the new claude read guard said the failure is folded in "like the write path in this same loop does", and there was no write guard in that loop. Rather than correct the sentence, both claude write sites are now guarded the way the read is — a failed file is skipped, the sweep continues, and the raw summary token flips to `failed`. The JSON shape stays frozen deliberately, because it is long-standing and widely consumed; the token is the channel that can carry the signal without a compatibility risk, which is the reviewer's own suggestion. That makes all three branches consistent: reads and writes guarded everywhere, per-file failures degrade instead of aborting, and every branch reports `failed` rather than `ok` when something did not sync. `setFrontmatterKeyLine` interpolated its value raw while the install-side writer quoted through the shared helpers — two writers of the same frontmatter key disagreeing on escaping, the divergence class this repo requires closed. They now share one rule. Verified no churn: all six effort levels are plain scalars and emit byte-identically, with claude's documented minimal-to-low clamp the only difference in the table, exactly as before. * fix(#3706): publish claude agent writes atomically too Both reviewers found this independently, and it is data loss rather than a reporting gap. The claude branch wrote in place, so `fs.writeFileSync`'s O_TRUNC meant a post-open fault left the agent file truncated or half-written: an injected ENOSPC produced an empty file, and under `ulimit -f` a 60000-byte agent came back as 512 bytes of wrong content. The guard added earlier this round then counted that destroyed file as `skipped`, which in JSON mode is indistinguishable from "already in sync" — so a caller would have read the sweep as clean while an agent on disk was corrupt. It now publishes the way the codex and opencode branches already do: write to a tmp file created at the original's masked mode, chmod, then retryRenameSync, with the tmp unlinked and the agent skipped on any failure. The corrupting case is gone rather than merely reported, which matters because this branch deliberately takes no new result key. I had claimed all three branches were consistent after the previous commit. That was true for degradation and reporting and not for atomicity; the reviewer caught the overclaim. It is true now. Also sorts the claude file list, which the other two branches already did — readdir order is platform-dependent, so leaving it unsorted made the reported `changes` ordering differ across machines for identical inputs. * chore(#3706): backfill the changeset PR number pr:0 placeholder replaced with the real PR now that gh api returned it. * test(#3706): kill the frontmatter mutants this change introduced CI's Stryker frontmatter shard scored 60.58 against a break floor of 62. The cause is documented in the lane's own config, from #1882: this PR added a multi-clause predicate to frontmatter.cts and exported the escaper, but the tests constraining them live in tests/runtime-converters.test.cjs, which that shard does not run — so every mutant in the new code was uncovered there even though the behaviour is tested elsewhere. The fix is assertions that kill real mutants, per the repo's own instruction, not a lowered floor and not a Stryker disable: scripts/mutation-matrix.cjs is untouched. Each clause of agentScalarNeedsDoubleQuoting now has a true case AND a near-miss that must answer the opposite way, so flipping the clause fails a specific named test — alnum-first against `a-b`, trailing `:` against `foo:bar`, embedded `: ` against `a:b`, embedded ` #` against `a#b`, the word list against `yes1`/`nullish`, the numeric forms against `1a`/`0xzz`, the timestamp against `2026-08-25x`, plus the case-insensitive spellings that pin the `i` flag. escapeDoubleQuoted is pinned on exact output, including a case constructed so that escaping in the wrong ORDER yields a different string. Two of my expectations were wrong and are asserted as the code actually behaves: `12:99` is NOT quoted, because the sexagesimal alternative never range-checks minutes and so does not match — which is right, since YAML would not read it as sexagesimal either; and `20260825` is quoted by the numeric clause rather than the timestamp one, being a bare integer. * chore(#3706): ratchet the frontmatter mutation floor to 65 The lane measured 66.67 on PR 3867 after the mutant-killing unit tests landed — above its pre-change 63.35 baseline, not merely recovered. Step 3 of this file's own HOW TO UPDATE procedure says to set minScore = floor(measured) - 1 in the same diff, so 62 becomes 65 and the improvement is locked in rather than left free to slide back. The ledger of measured scores now records the new measurement, why the shard broke in the first place (logic added to frontmatter.cts whose only tests lived in a file this lane does not run — the same trap the #1882 note describes), and one discrepancy: step 3 also says to update "the matching RATCHET_BASELINE entry", but no such declaration exists in this file. The name appears only in that comment, so minScore and the ledger are all there is to update. * fix(#3706): update RATCHET_BASELINE alongside the raised floor The ratchet test caught the previous commit: it raised COVERED['frontmatter'] .minScore to 65 without updating the baseline that mirrors it, which is exactly the mismatch that guard exists to make visible in review. I had claimed RATCHET_BASELINE did not exist. It does — in tests/mutation-matrix-ratchet.test.cjs, not in scripts/mutation-matrix.cjs, which is the only file I searched before concluding it was a stale reference. The ledger comment is corrected to say where it lives and to record that the guard caught the error rather than leaving my wrong claim on the record. * docs(#3706): put the mutation ledger entries back under their own dates The 2026-08-25 measurement was spliced into the middle of the 2026-06-14 list, so adr-parser, config-schema, active-workstream-store and core-utils ended up sitting under the wrong heading and misattributing their measurement dates. That ledger is what a future change reads to calibrate a floor, so a wrong date there is not cosmetic. Each measurement is now under the date it was taken. Also drops the first-person account of my own mistake from the entry — the factual half (where RATCHET_BASELINE lives, and that it is updated in the same diff) is what a reader needs; the confession is not. --------- Co-authored-by: sim <sim@local> |
||
|
|
004e9dd741 |
fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765)
* test(#3007): failing-first suite for per-model Codex effort capability RED by construction. Binds to behavior renderEffortForRuntime does not yet have: an optional third `model` argument, a per-model advertised-level table, `max` passing through instead of clamping to `xhigh`, `minimal` clamping to `low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/ `reason`) so a downgrade is legible from resolver output rather than silent. Two of these pin defects that exist on next today: - `max` is discarded. Both Codex models whose catalog entries are retrievable (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing. - `minimal` is emitted to a model that refuses it. providerPresets.openai. haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's advertised floor is `low`. GSD is sending a value into a document Codex itself validates. The parity test is what pins that fixed, and it names the offending path/model/effort when it trips. Also corrects tests/model-resolver.test.cjs:351, which asserted renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as though it were a contract. ADR-443 recorded "Codex has no max" as fact and it was true when written; Codex has since added both `max` and `ultra`. That is a stale premise, so the assertion is corrected here rather than worked around. The property test asserts the invariant the whole change exists for: a rendered effort is always a level the target model actually advertises, or an explicit rejection. There is no third outcome. * fix(#3007): resolve Codex effort per model, and make every clamp visible Codex declares supported_reasoning_levels per MODEL and validates against it, so a single per-runtime capability set cannot be right for all of them. GSD's was wrong in both directions at once. `max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped max -> xhigh on that basis; it was accurate when written, and Codex has since added both `max` and `ultra`. Every Codex model whose catalog entry is retrievable advertises `max`, so the clamp was discarding a level the provider supports, silently, on the most-used path. `minimal` stops reaching Codex. No Codex model advertises it -- both retrievable entries floor at `low` -- yet providerPresets.openai.haiku.low paired gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the receiver validates and refuses into a file the receiver reads. Being unconservative in what you send is the half of Postel's rule with no defensible reading, so that preset is corrected and a parity test pins it. `ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum reasoning with automatic task delegation": at ultra, effective_multi_agent_mode returns Proactive and Codex spawns sub-agents on its own initiative, underneath GSD's orchestration rather than inside it (#2167). It is a mode switch, not a reasoning depth, so it is not added to the universal ladder -- which stays provider-agnostic by ADR-443's design -- and it is rejected even for gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered and rejected: that silently discards what the user actually asked for. Clamping is now visible. RenderedEffort carries requested/clamped/reason and resolve-execution surfaces them. The previous table clamped correctly but invisibly, so a user asking for `max` on Codex had no way to find out they were getting `xhigh` -- exactly the failure mode the robustness principle's modern critique warns about, and why "be liberal" has to mean "liberal and loud". Also closes a latent trap found while reviewing the implementation: the clamp-up loop walks the ladder upward, and for a future model advertising `ultra` but not `max` it would have selected `ultra` as the clamp target -- re-entering by the back door the mode the rejection above exists to keep out. A clamp may never produce a value that a direct request for that value would refuse. Unreachable with today's catalog, which is why no test caught it; a test now asserts the invariant directly. Signature stability is preserved: the third `model` argument is optional and the two-argument form still resolves, against the family baseline. That form's BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old answer would have fixed the defect only where a model happened to be threaded through and left it live everywhere else. tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract and is corrected here rather than worked around. * fix(#3007): close every review finding on the Codex effort alignment Two isolated reviewers, correctness and security. Both found the same two blockers, and the per-model work was inert on every surface that matters until this commit. BLOCKER — resolve-execution never passed the model and discarded the clamp. cmdResolveExecution called the two-argument form and emitted only effort_rendered/effort_param/effort_propagation, so the per-model table was unreachable from production code (tests were its only caller) and requested/ clamped/reason were computed and thrown away. Requested outcome 3 names "the effective rendered effort in resolver output" specifically, so the feature was unmet on the exact surface the issue asks for. Now passes the resolved model and emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the existing key convention rather than introducing a nested object. BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed a nested {"effort": ...} sample; the real result is flat and those keys were absent entirely. A reference doc asserting a JSON path a reader can copy is worse than no doc. Corrected against the actual emitted key set. MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex kept minimal in its supported set and still clamped max down to xhigh, so the invocation-time and install-time channels disagreed about the same runtime's capability: --host codex with max emitted xhigh while the generated TOML said max. This is the repo's documented generative-fix-divergence class, so both tables now cross-reference each other and a parity test fails if they ever diverge again. MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null _baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback never fired and every effort rendered as null. And a non-array value made the Set constructor throw at module load — model-catalog.cjs is required across the whole CLI, so one bad JSON value killed every command, not just codex effort. Guarded on size and filtered to array values; both degrade to the hardcoded baseline. MAJOR — value widened to a nullable string with two consumers left behind. runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a null effort key in generated frontmatter); install-effort-resolver still declared a non-nullable return, a structural lie that silently defeated null checking. Both corrected, both omitting the key on null — the same posture as 'inherit', where omission means "follow the host default". MAJOR — the per-model table is inert today, and the docs now say so. All three shipped models advertise the same usable range and ultra (sol's only differentiator) is rejected for every model, so no observable output differs by model. The table stays because Codex declares capability per model and the sets are free to diverge — a single per-runtime assumption is precisely what went stale and produced this issue — but overselling it as a visible per-model feature would have been the same class of error as the doc blocker above. Tests: three passed under a full revert and are strengthened rather than deleted, since each guards a real contract (#3533's inherit rule, the undeclared-host rule, off-ladder handling) — they now also assert the clamp-visibility fields, which only exist after this change. The fast-check property is kept for its shrinking, and a deterministic nested loop over the full cross-product now sits beside it so coverage is exhaustive rather than sampled. Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg form and would have written a literal null reasoning effort on the ultra path; CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT. The installer defect was found by the co-change gate, not by a reviewer — install.js is a historical co-change partner of model-catalog.cts that this diff had not touched. * test(#3007): correct assertions that pinned Codex's stale effort premise Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the shipped commit. Every one is a stale pin, not a defect: each was probed against the built module before its expectation was changed, and none failed for a reason other than this premise correction. Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale by a production change must not ride inside another commit, because the release-sdk hotfix cherry-pick filter routes by subject prefix and a correction buried under the wrong prefix ships a half-state (v1.42.3, #3621). The most valuable one was tests/model-resolver.test.cjs's cross-provider validity invariant, which hardcoded the Codex enum as `minimal|low|medium|high|xhigh` and failed with "real API would 400". That message is now false in both directions: Codex accepts `max`, and rejects `minimal`, which no model advertises. The enum is corrected to `low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the "would the real API refuse this" check worth having, and it was right to fail here. It simply carried the stale fact in its own fixture. Test NAMES were corrected alongside their assertions wherever the name asserted the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal passthrough". A renamed test that still claims the old thing is worse than a failing one, and a green test whose name states a falsehood is how the next reader inherits the wrong premise. Both channels are covered: install-time (renderEffortForRuntime, and the generated .toml in install-runtime-artifacts) and invocation-time argv (effort-surface-axis). They were deliberately brought into agreement in this change, so their assertions had to move together. Each site carries a #3007 comment recording that Codex gained max/ultra and that capability is declared per model, so a future reader can tell this was a deliberate premise correction rather than a test bent to fit an implementation. * test(#3007): separate the effort-precedence case from the clamp case The previous stale-assertion pass over-corrected one test. It saw `effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`, assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to "max passes through" and changed the expectation to `max`. The remote runner disagreed. Reproduced against the real CLI: with that config and `gsd-planner`, the resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`, `effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE — gsd-planner is heavy/opus tier and its routing-tier default outranks `effort.default` — so `max` never reaches the renderer at all. The test says nothing about clamping and never did; it only looked like a clamp pin because both mechanisms happened to yield the same string. Restored to `xhigh` and renamed to say what it actually tests. It now also asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is what makes it impossible to mistake for a clamp pin again: those two fields prove the value is what the resolver produced rather than something the renderer downgraded. Before #3007 there was no way to tell the two apart from the output — which is precisely why the previous pass could not tell them apart either. Added the test that was actually missing: `effort.agent_overrides`, which outranks the tier default, so the requested level genuinely reaches the renderer and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified by probe before asserting. One test now pins the precedence rule and the other pins the #3007 behavior, and neither can be read as the other. That the clamp-visibility fields are what resolved this is a small argument for having added them. * chore(#3007): backfill changeset pr number to 3765 * test(#3007): put model-catalog under the mutation gate The Stryker shard showed as `skipping` on this PR despite the diff rewriting model-catalog's effort logic. That was legitimate, not a detection bug: `model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the whole module — including everything #3007 touches — sat entirely outside mutation scoring with has_work "false". Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no temp dirs. That shape is not stylistic — it is the #2790 precedent this file already documents. Stryker's command runner treats a whole `node --test <file>` invocation as ONE test costing whatever its slowest case costs, and re-runs it per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses runGsdTools throughout) would reproduce exactly the 15-minute shard-cap cancellation #2790 hit. The integration file is unaffected and keeps running in full in the normal test job. Coverage spans the module rather than only the diff, because the score is measured over the whole file: effort rendering across every model and ladder level in both channels, the prototype-chain host guard, the exported enums and maps, isAnthropicFlavoredModel's provider namespacings, the profile projections, nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are worth naming — every uncovered exported function is score given away, and mergeEffortTierDefaults turned out to have a genuinely interesting contract (#3531: a partial override merges over the built-ins rather than replacing them, and isValid gates the VALUE, not the tier name, so an unknown tier key is still merged in). Every expectation was probed against the built module before being asserted. minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment. Floors in this repo are measured, not chosen — the existing entries sit at 94, 75 and 56 — and they can only be measured in CI, because mutation shards run `node --test`, which is hard-blocked locally. The first CI run on this branch reports the real number and the floor gets ratcheted to it before merge. A placeholder of 1 reaching `next` would make the gate decorative: it would pass whether or not a single mutant is ever killed. Note the target is "never regress from measured", not a fixed 80 — planning-inspect sits at 56 and is documented as an accepted ratchet candidate. * test(#3007): bootstrap model-catalog's mutation floor legally The placeholder floor was structurally illegal and the remote run said so. tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module must carry a matching RATCHET_BASELINE entry in the same diff, minScore must EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three. That is the ratchet working exactly as intended — a floor nobody can satisfy accidentally is the point of it. Bootstrapped at 50 in both places. Fifty is not a measured score and the comment says so plainly: it is the minimum the guard permits, and it coincides with Stryker's own configured `break` threshold, so it is the lowest legal starting point for a module that has never been measured. It still must be ratcheted to floor(measured) - 1 before this PR merges. Also corrected a real defect in the file's own instructions. "HOW TO UPDATE" step 1 read "Run the per-module Stryker shard locally" — which cannot be done here, and which the same file contradicts eighty lines further down, where the #2790 scores are recorded as "not a local run; mutation shards run `node --test`, hard-blocked in this repo's local environment". stryker.config.mjs confirms the command runner invokes `node --test` once per mutant, and .claude/hooks/block-local-node-test.sh denies exactly that. So the documented first step sends the next contributor at a wall. Rewritten to describe the path that works — push, read the measured score off the CI shard, then set the floor and its baseline together in one diff — and to say why local measurement is not available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched. The two-step is inherent to the environment rather than a shortcut: a floor cannot be measured before the first CI run exists, and the guard rightly refuses to accept an unmeasured one below its minimum. * test(#3007): ratchet model-catalog's mutation floor to its measured score The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this file's own rule, minScore = floor(measured) - 1, which is the same arithmetic every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94. Both halves moved together, because the ratchet guard asserts minScore equals its RATCHET_BASELINE entry and would reject them drifting apart. The spawn-free unit surface is vindicated by the clock: 57 seconds, against a 15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing the shard at tests/model-resolver.test.cjs — #2790 recorded shards being CANCELLED at that cap when they targeted a runGsdTools-heavy integration file. The registry comment is rewritten rather than deleted. It previously warned that the floor was provisional and must not ship that way; leaving that text next to a measured floor would make the file lie in the other direction. It now records the measurement the way the sibling entries do, including that 59.62 sits below TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 — comfortably clear of its own floor with real room to grow. Raise it as the tests improve; never lower it. Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is an honest floor for a module that had NO mutation coverage at all an hour ago, and it is now pinned so it cannot silently regress. --------- Co-authored-by: sim <sim@local> |
||
|
|
3ab0007164 |
enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19) preserveUserArtifacts held user files only in an in-memory Map across the wipe, so any process death between preserve and restore lost them outright. Seven call sites, not the four the issue records. Three of them never called the helper at all - they open-coded the same read/wipe/write - so searching for callers under-counted by construction; the extra sites were found by sweeping for the pattern instead. The worst is the mainline install path, where the crash window spans the entire gsd-core tree copy rather than a single rmSync. Adds src/user-artifact-staging.cts: durable on-disk staging with a record written after the copies land as the commit point, plus recovery of orphaned batches on the next run - without recovery the staged bytes survive but the user's file is still gone, which would pass its own test while delivering nothing. Routes copyPreservingSymlink through installFs() so staging cannot bypass the install fs seam, and reunites its symlink-safety docblock with the function it documents. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-3574 with four claims disproved by implementation Implementing Phase 6 disproved four statements the ADR rests on. The central decision - no single materializer - is unaffected and stands. Corrected: decision 3 was already satisfied, so nothing was extracted; the agents-bypass runtime set omitted claude, kilo and opencode, and closing it needed three new pieces of descriptor contract rather than proceeding on its own terms; three of the four blockers the layout comment names were already stale; and F19 is seven call sites, not four. Records the generalizable lesson: the defect is the pattern of holding user data in memory across a wipe, not the helper, so searching for callers of the helper under-counts by construction. Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and notes that copyPreservingSymlink needed routing through the install fs seam before it could be reused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close dangling-symlink blind spot and harden staging recovery An adversarial review found the F19 staging work shipped red and unsafe. Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween missed dangling symlinks in both its root check and its per-segment walk, because it probed with existsSync, which is false for a link whose target does not exist. Fixing only the new module would have reused a guard that was itself blind. This guard protects the whole install tree. Recovery no longer throws: it degrades per entry and per file, so one bad batch cannot block the others. Previously an unrecoverable entry propagated out of the first statement of install and uninstall, before the cleanup that would have removed it - wedging the installer permanently. Partial fs adapters now throw on any omitted method instead of silently reaching the real filesystem, closing the trap that let a test poison list pass while real IO happened. Staged names must be flat, recovery refuses a dangling destination symlink, and a batch whose recovery genuinely failed is no longer swept - it was discarding the only durable copy of the file it had just failed to restore. Replaces three tests that could not fail, including the one labelled negative proof. Known limitation, documented not closed: concurrent installs sharing a staging key can still lose a batch. A real fix needs a cross-process lock. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enh(#2875): make the descriptor authoritative for the agents kind Deletes the inline agent-staging loop in bin/install.js and the _DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from its capability descriptor instead of an inline hostBehaviors dispatch. Closing it needed three pieces of contract the descriptor pipeline never had, all reducible to one missing input - per-agent resolution context: a frontmatter-extensions step for claude's effort and disallowedTools, per-agent model-override resolution for kilo and opencode, and a named branding converter for hermes, whose rewrite data was already declared. Seven runtimes were on the loop, not the six the design recorded - kimi-code was found by a golden fixture, not by analysis. claude-local and kimi-code both silently lost their agents mid-change; the fixtures caught both and the cause was fixed rather than the fixtures regenerated. A parity harness gates the migration: both pipelines over identical inputs, byte-identical output including filenames, per runtime. It is demonstrated red before being trusted. Surface and install paths converge for all seven, which also fixes surface previously writing no agents for these runtimes. Codex's config.toml strip stays put - it mutates host config, which no descriptor kind models. Also routes install-model-override-resolver and install-effort-resolver through the install fs seam. Both leaked real filesystem IO from the install call tree; the stricter adapter is what exposed them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): record the agents-descriptor migration and correct the ADR count The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host integration guide told readers to join a set that is gone. Replaces that with what is now true - declare an agents entry and it installs, on the surface path as well as install - and points anyone needing a per-agent transform at the three extension points rather than at a new inline branch. Corrects the ADR amendment: seven runtimes were on the inline loop, not six. kimi-code was found by a golden fixture going red, not by reading. That is the third short count this phase, all from enumerating by symbol or set membership when the thing that matters is a behavior. Adds the Changed changeset for the surface-path convergence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-2866 - claude global always wrote agents on disk The claude row's global=[skills] described what capability.json declared, not what the installer wrote. bin/install.js's inline agent-staging loop was never scope-gated and never consulted the descriptor, so a claude --global install has always written agents/gsd-*.md. Phase 6 closes the gap by deleting that loop and declaring agents on claude's descriptor at global scope. On-disk bytes are unchanged - the golden fixtures did not move, which is the evidence that the descriptor, not the installer, was incomplete. #2218 is unaffected: agents are not trigger-bearing, so the wider row does not introduce a new shadowing case. Records the warning that an incomplete descriptor is invisible while a second code path silently does its work, and only surfaces when the two are forced into agreement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close review findings across staging, agents and the parity harness Two independent reviews of this branch found defects the local gates missed. Security: a dangling symlink at a migration destination allowed writing outside configDir - the same class this change claimed to close, missed at the terminal write of the flow being added. The staging-root resolver threw as the first statement of install and uninstall, so a hostile symlink bricked both, and symlinked-configDir users lost uninstall as well as install; it now degrades instead of aborting. Recovery gained a source-side symlink check and now refuses a relative destDir, which resolved against cwd. Converter dispatch gained a runtime allowlist - lint-time validation stopped mattering once this branch promoted that dispatch from the surface path to real installs. Correctness: claude --local --minimal exited 1 because the minimal profile legitimately yields zero agents and the new path treated that as a failure. cline --local silently lost its agents - its descriptor declared none while the deleted loop wrote them unconditionally. The agents prune was widened to any gsd-* entry and destroyed user files it never owned. The parity harness, on which the migration's safety argument rested, drove a synthetic registry and never byte-compared the shipped descriptors; two of its trap rows could not fail. It now drives the real registry across 13 runtime-scope rows including kimi-code and cline-local, and its red-proof is demonstrated by corrupting a live capability.json. Three goldens that had encoded the cline regression as expected behavior were corrected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close findings from both mandated review engines /security-review found the staging source-side walk honouring GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write destination. A symlinked files/ component dereferenced because copyPreservingSymlink lstats the leaf only, so an intermediate link is followed. The source walk no longer honours the opt-in; the destination check still does. /code-review spec axis found this branch had reintroduced its own bug: migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after the legacy dir was wiped and before the staged batch was restored, so a planted symlink bricked uninstall permanently and orphaned the batch. Refusal kept, abort removed. kimi-code local silently lost its agents, the same class as the cline bug, and the parity harness recorded that exclusion as intentional - the third test in this branch to pin a regression as correct. --minimal now creates an empty agents/ dir that never existed. Behaviour restored rather than softening the changeset, so its byte-identical claim stays true. Standards axis: try/finally removed from twelve test bodies, fast-check properties added for parseOwnerPid, boundary coverage at the grace window and the ancestor-probe depth, a parity assertion for the staging-root helper duplicated across two files, and the 8-deep config walk deduplicated. Records 60-review.json with every finding and disposition from five passes, including the smells left unfixed and why. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): prune stale agents unconditionally in minimal mode The previous round stopped an empty agents/ directory being created when the resolved profile yields no agents. That was implemented by skipping the agents kind entirely, which also skipped its stale-agent prune - so a full to minimal downgrade left stale gsd-* agents behind. The deleted inline loop pruned unconditionally and only skipped writing. Those are three separate conditions, not one: prune always, write only when there is something to write, create the directory only when writing. Both call sites now run _removeGsdEntries before the empty-staged early exit. The symlink-escape guard moved with it, since the prune also touches dest. Codex .toml agents and the config.toml stanzas are cleaned again, and user-owned agents are still preserved. The agents/ directory is left in place after a prune empties it, matching every sibling kind - none of them remove the destination directory itself. Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh install, so fixture generation is unaffected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): document interrupted-install recovery for user-owned files The durable-staging fix is invisible to the user it protects. Someone whose install died mid-flight has no way to know USER-PROFILE.md was staged before the delete, that the next run restores it, or that recovery happens at the start of that run rather than in the background. Written as the task the user has - finish the interrupted command - rather than as a description of the mechanism, and states what it will not do: overwrite a file already present, or touch staging belonging to another install still running. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2875): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2875): assert the J8 model override without building a regex CodeQL flagged incomplete string escaping: the assertion interpolated the override value into a RegExp while escaping only forward slashes, which is meaningless in a constructor, leaving real metacharacters unescaped. The failure direction was the dangerous one - a metacharacter would have made the match more permissive, so the row would pass when it should fail. That matters here because J8 exists precisely because an earlier revision was a tautology; the rewrite reintroduced a different way for the same assertion to stop discriminating. Replaced with a line-wise exact match, so no regex is constructed at all. Swept the other test files this branch adds; no sibling instances. lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full metachar-escape copy, so a single slash replace slipped under it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a4a02a7a01 |
enhance(#2874): return the executed plan and route install IO through a seam (#3568)
* test(#2874): add failing-first gate for the executed-plan return Four rows from the matrix's red-first order. E3 pins the one early return, for the opencode family, where a void-shaped hole would otherwise survive unnoticed. E13 sweeps every runtime in the registry - enumerated from the registry rather than hardcoded, so a runtime added later cannot slip past. F2 proves absence of real filesystem contact rather than merely that the happy path ran, which is the difference between a complete seam and a partial one. G1 and G3 are the additive guard and must be green before and after. G3 deliberately leaves the two existing adapter test doubles untouched: if this change required editing them it would not be additive, and the acceptance criterion would be unmet. No production code. All 19 runtimes install without throwing today, so E3 and E13 fail on the undefined comparison alone. Refs #2874 * feat(#2874): return the executed plan and route install IO through a seam installRuntimeArtifacts returned void, so its correctness was observable only by re-reading disk. It now returns what it executed - per kind, per scope - including on the combinedFamilyInstall path, which was the one early return where a void-shaped hole would have survived unnoticed. Failure still throws rather than becoming an ok:false return, so control flow is unchanged for both existing callers. A best-effort cleanup that fails is still swallowed, but is now visible in the returned value rather than silently absent. The fs seam is ambient rather than threaded. Explicit deps through install-profiles and the 3000-line conversion module was impractical; the tradeoff, the synchronous-only re-entrancy assumption, the restore guarantee and the partial-adapter fallback trap are all documented at the seam. findInstallSourceRoot and its sibling stay unrouted by design - they locate the package's own source, not the install destination. readCmdNames keeps a second implementation because the standalone CLI that owns the original cannot require the compiled adapter without a build-order dependency on its own output. A parity test fails if the two ever disagree. Refs #2874 * chore(#2874): gitignore the new build artifact install-fs-adapter.cjs is tsc output from src/install-fs-adapter.cts, not a tracked source file. It was added to eslint's ignore list but not to .gitignore, so it landed as a tracked file - the third time this step of the new-.cts ripple has been missed on this epic. Refs #2874 * fix(#2874): close two seam leaks and correct a false comment A correctness review found the seam still leaked in two places, both subtler than the three already closed. readGsdCommandNames was routed when it should not have been: it reads the package's own commands directory, which a destination-fake is never seeded with, so under a fake adapter it returned an empty or wrong roster instead of failing loudly. It now reads real fs, matching the precedent already documented for findInstallSourceRoot. cleanupStagedSkills ran raw rmSync from a process exit handler, which is real filesystem work deferred past the point where withInstallFs has restored - the one thing the synchronous-only contract exists to exclude. Staging now captures the adapter that created each directory and cleanup replays it, so a real install cleans up exactly as before and a fake-staged path never reaches the real filesystem. Also corrected a comment claiming the migration reads were an unrouted, untested residual gap. They are routed and exercised; a comment understating the seam is as corrosive as one overstating it in a module whose trust rests on being honestly documented. Refs #2874 * test(#2874): migrate the exemplar group and cover the matrix AC3's exemplar migration lands in place: the qwen install group now asserts skills and agents destinations from the returned plan in one deepStrictEqual instead of probing the filesystem for each. Nine facts the old probes established were enumerated first. Two moved to the value assertion; seven were retained deliberately - per-file SKILL.md existence, the VERSION file written outside this function, the manifest content, and the post-uninstall absence checks all sit outside the plan's per-kind contract. A migration that quietly asserts less looks like a win and is a regression, so the enumeration is the guard rather than the line count. Also implements the rest of the matrix: the executed-plan shape, adapter failure modes, the security-boundary rows including a fake that cannot certify an install the real filesystem would refuse, cleanup visibility, and two seeded property tests. Only the two external CI gates are left unticked, because self-certifying them would be a claim rather than a check. Refs #2874 * fix(#2874): restore streaming hashes and derive F2 from the boundary rule The checkpoint found three things reasoning had missed. sha256File had been converted from raw-fd streaming to a single readFileSync on the assumption that GSD artifacts are never large. A test named for exactly that contract already existed and went red. Streaming is restored, now routed through the adapter, which gains openSync, readSync and closeSync. The contract was the specification; the assumption was not. Three existing tests inject faults by monkeypatching real fs. They broke because mkInstallTempDir stopped calling real mkdtempSync, not because of any binding subtlety - the real adapter was already late-bound. It now calls the real function when no fake is injected, so a monkeypatch applied after import is still seen and the additive contract holds. F2 poisoned real fs by method, so a deliberately unrouted package-source read failed a correct design. It now poisons by path: destination IO is forbidden, package-source IO is allowed and positively asserted. The claim was always zero real destination IO, and the test now derives from that rule instead of coincidentally matching it. Refs #2874 * docs(#2874): add the contributor how-to for plan-based test migration The phase gate caught a real gap. The docs plan was Reference plus Explanation only, and every CI check would have passed, because the docs-required lint only verifies that some file under docs/ moved. But this phase exists to demonstrate a pattern for follow-on work, and that work is other contributors migrating probing test groups. The sequence has two live traps - a partial fake silently falls back to real fs, and the seam is ambient and synchronous-only - plus one discipline nobody infers: enumerate the facts before converting, or you assert less and call it a win. The page carries the qwen migration's arithmetic, nine facts enumerated and only two converted, because a reader seeing only the diff would reasonably conclude the pattern is to replace probes wholesale. No locale mirrors: none of the four carries any contributor-only how-to, so a single translated file would manufacture parity rather than provide it. Refs #2874 * chore(#2874): backfill changeset pr number * test(#2874): normalize both sides of the G1 tree comparison G1 failed on Windows only, deterministically on both shards. The defect was in the test helper, not production. _computePathPrefix posix-normalizes the resolved config dir unconditionally, so on Windows the path embedded in every emitted SKILL.md body is forward-slash form. hashDirTree stripped against the raw backslash path from mkdtempSync, so the substring never matched and each install's unique temp suffix stayed baked into every file - all fifteen skill bodies hashed differently for two runs that had written identical bytes. Both sides are now normalized unconditionally rather than gated on path.sep, matching the rule this repo already records: backslash paths arrive on Linux too. Production code is untouched and was verified correct. Normalizing this away on the production side would have hidden a real portability bug if one had existed. Refs #2874 --------- Co-authored-by: sim <sim@local> |
||
|
|
fd2b97a52a |
fix(#3544): restore tilde form for at-refs in the global spec tree (#3551)
* fix(#3544): restore tilde form for at-refs in the global spec tree A global claude install emitted @$HOME/.claude/gsd-core/references/*.md in its workflows and references. $HOME does not expand in a Claude Code @-import - only relative, absolute and ~ are documented, and a controlled /context test confirmed a $HOME import loads nothing - so 54 includes across 22 files silently resolved to nothing on a live install. This is a divergence, not a new bug. #3133 already applies exactly this correction to skill and command bodies through _applyRuntimeRewrites's claude case; copyWithPathReplacement, the spec-tree emit path, never had it. Both now call one exported helper, so the two surfaces cannot drift apart again. Deliberately narrower than changing computePathPrefix's return value: shipped markdown also carries double-quoted "$HOME/.claude/..." shell invocations, and ~ does not expand inside double quotes, so rewriting the prefix wholesale would regress #1284. Only @-prefixed references move. Refs #3544 * fix(#3544): derive the tilde restore from the resolved prefix Three review findings, one batch. The restore was hardcoded to the literal .claude directory, so a global install with --config-dir pointing anywhere else silently no-opped and reproduced the very defect this fixes. It now derives the tilde form from the resolved prefix, which also closes the same latent gap in #3133's original path since both call sites share the helper. The @-anchor is quote-aware, so a double-quoted shell path is never rewritten into a form the shell does not expand. Deliberately a lookbehind rather than a line-start anchor: @-references are documented to work mid-line, and anchoring would have traded a theoretical bug for a real one. Found while testing the above: the bare-form rewrites re-matched their own output whenever a config dir name extends .claude, emitting .claude-work-work. Guarded with the same negative-lookahead convention this file already uses to preserve .claude-plugin. The tests prove the emitted form, never that the host resolves it - no CI test can - and both the helper and the suite now say so, because an undocumented verification boundary is how this defect stayed green for its whole life. Refs #3544 * test(#3544): acknowledge the tilde-restore emitted drift The converter change moves 94 emitted paths that no source-file diff can explain, which is exactly the case the per-PR ack fragment exists for. Verified before acknowledging rather than after: both trees were built from real installs and every one of the 211 changed lines across all 94 paths is @$HOME becoming @~, with nothing outside that single kind. Nine spent entries were pruned from the #3151 and #2658 fragments. Those paths moved again here, and two ack sources naming one path is a hard duplicate error rather than last-wins, so the inert entries had to go before this one could land. Both fragments retain their remaining entries. Refs #3544 * chore(#3544): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
147856040b |
fix(#2873): close review findings across fences, sanitizer and docs
Isolated security review found resolveSpecRootReference's fence tracker toggled on any delimiter, so a backtick fence could be closed by a tilde one and an include in the gap was rewritten inside a code block. Fixed by reusing scanFencedBlocks - the canonical engine already behind stripFencedCode and extractFencedBlock - rather than carrying a fourth copy of fence detection, which also closes the duplication the standards review flagged. sanitizeForRender now strips combining marks and zero-width characters alongside the ANSI, control and bidi classes it already handled. Adds the C, E and F matrix rows the spec review found missing, including installer-level coverage that spawns the real install rather than calling the report builder. Ships the how-to, the reference and command docs in five locales, the changeset, the inventory and glossary entries, and regenerates health.md for the new W028 rule. Refs #2873 |
||
|
|
2641e6cb67 |
feat(#2873): detect cross-scope shadowing and reach the local spec tree
4a - the detection floor. A shadowed install now reports which triggers are shadowed and which scope wins, at install time and through a new W028 /gsd-health diagnostic. Exit codes are untouched: a shadowed install is a warning, not a failure. Only triggers whose stem exists at BOTH scopes are reported, so a global full profile beside a local core profile no longer names local artifacts the user does not have. 4b - spec-root reachability, claude runtime and global scope only. The winning global skill stops carrying a static workflow @-include and instead resolves its spec at runtime: prefer the project-local copy, fall back to the global one, stop if neither exists. Every other @-include stays static, and the local emission is byte-identical. It runs after the staged-skills rewrite pass, whose claude branch would otherwise mangle the literal tilde path into an undocumented $HOME form. Also fixed inline: readInstallManifest classified a top-level JSON array as an installed v1 manifest, because typeof [] is object. Refs #2873 |
||
|
|
71180983a0 |
fix(#3423): standardize on <required_reading>, retire the files_to_read emit tag (#3432)
* fix(#3423): standardize on required_reading, retire files_to_read emit tag * test(#3423): flip tag assertions, extend consistency guard to spawner surfaces * fix(#3423): sweep capabilities fragments, regen registry+skills, anchor executor test * chore(#3423): acknowledge tag-rename emitted ripples and workflow growth * chore(#3423): broaden emitted-ripple acknowledgment to all embedders * chore(#3423): settle emitted-drift acks post-rebase (merge 3004/1689-owned keys) * chore(#3423): drop stale ripple acks, ack execute-phase growth * chore(#3423): restore pristine 3004 fragment, keep only consumed appends * chore(#3423): backfill changeset pr number * chore(#3423): settle emitted-drift acks post-merge (move code-review-fix ripple into 3190, tag-rename ripples into 3191/3297) * chore(#3423): re-arm 3324 ack for execute-phase.md tag-rename ripple * fix(#3423): trim 8 bytes from execute-phase model note to hold ADR-857 margin, re-arm 3370 ack for net +4 growth --------- Co-authored-by: sim <sim@local> |
||
|
|
967bddba37 |
fix(#3384): strip mcp__* tool grants from zcode-installed subagents (#3483)
* fix(#3384): strip mcp__* tool grants from zcode-installed subagents ZCode's dispatcher treats every mcp__<server>__* entry in an agent's tools: frontmatter as a required MCP server and hard-fails the subagent spawn (CONFIGURATION_ERROR) when it is not connected, whereas Claude Code treats the same grants as an optional allowlist. ZCode shared Claude's verbatim agents copy (converter: null), so all 8 MCP-granted agents failed to spawn out of the box with zero MCP servers configured. Add convertClaudeAgentToZcodeAgent — a line-surgical converter that filters mcp__* entries out of the frontmatter tools: grant list (both inline comma and YAML block-list shapes) and preserves every other byte. Declare it on both of zcode's capability.json agents entries and cut zcode over to the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES) so the legacy inline loop stops deleting+re-copying the converted agents raw. Claude Code, Kimi, and Gemini install behavior is unchanged. * chore(#3384): link changeset fragment to pr 3483 --------- Co-authored-by: sim <sim@local> |
||
|
|
b9adedbc86 |
fix(#3151): stop emitting effort: into skill frontmatter (cache invalidation) (#3425)
Claude Code applies SKILL.md effort: as output_config.effort; any change from the session baseline invalidates the prompt cache at BOTH scope boundaries (entry + exit, the latter often machine-fired via subagent-completion notification). The reporter's owned measurement confirms it: /gsd-progress (effort:low) in a medium session → cache_creation 63,404 (entry) + 18,589 (exit), while a no-effort skill shows none. ~76% of invocations paid in full. Fix (trek-e AC#2/AC#4): convertClaudeCommandToClaudeSkill no longer emits effort: into Claude-runtime skill frontmatter (src/runtime-artifact-conversion.cts + duplicated bin/install.js). normalizeClaudeSkillEffort removed (dead). The six declaring skills (plan-phase/execute-phase/autonomous/next/progress/stats) no longer carry effort. Source command files keep effort (input, used elsewhere); the separate agent-effort surface (#3160) is untouched. Tests: install-runtime-artifacts #769 block flipped to assert effort is ABSENT from installed SKILL.md + converter output (the AC#4 behavioral coverage). Co-authored-by: sim <sim@local> |
||
|
|
dc3c81e93d |
chore(#3212): src/pattern.cts is the sole owner of runtime-value regex construction — Phase 1 (#3416)
* test(#3412): failing-first suite for the pattern-construction seam Phase 1 of epic #3212 (ADR-3212 §1/§2/§7). Tests only — src/pattern.cts and eslint-rules/no-adhoc-regex-escape.cjs do not exist yet, so both suites fail with MODULE_NOT_FOUND, which is the intended RED. Locks the measured behavior rather than the assumed behavior: RegExp.escape hex-escapes the leading character of nearly every string ("abc" -> "\x61bc"), so the suite asserts match-equivalence against an inlined historical oracle (the implementation being deleted) rather than byte-equivalence of pattern text — 200 seeded fast-check runs plus a fixed corpus, 0 mismatches. Also locks the latent character-class range bug this phase fixes as a side effect: a hyphen-bearing value interpolated into [...] currently forms a real range and matches an unintended character; post-migration it must not. * chore(#3412): src/pattern.cts owns runtime-value regex construction Phase 1 of epic #3212 (ADR-3212 §1/§2/§6/§7). Adds the pattern seam delegating to the built-in RegExp.escape, deletes every hand-rolled copy, and raises the Node floor to the Active LTS line. The census was low, three times over. ADR-3212 counted 10 copies; a graph query found 12; the new lint rule — once live — found 27 more. The difference is that the census counted named helper FUNCTIONS while the rule counts the escape SHAPE, so inline .replace(<class>, '\$&') copies were never in scope. ADR §1's actual requirement is that no module outside the seam escapes a value for regex use, so all of them are, and CLAUDE.md's no-defer rule makes them this change's work. Fourth consecutive epic here whose copy count was low — the argument for ADR-3180 Amendment 3's "state N found by the guard" rule. Also corrected mid-implementation: the survey reported phase-id.cts's escapeRegex had 0 external importers. It had 8 production importers, making its removal a public-surface change to an ADR-2121-owned module and requiring an update to that ADR's locked-surface test. Blast radius revised Medium-High -> High. RegExp.escape is match-equivalent but NOT text-equivalent: it hex-escapes the leading char of nearly every string ("abc" -> "\x61bc"). Equivalence is proven by a seeded fast-check property test against the deleted implementation as oracle. It also fixes a latent bug: a hyphen-bearing value interpolated into a character class previously formed a real range and matched an unintended character. Node floor 22 -> 24 (RegExp.escape is Node 24+), across engines, .nvmrc, package-lock, 9 CI matrix entries, and 5 docs. The aggregate `required-tests` context is unchanged and no job was added or removed, so branch protection cannot be orphaned by the dropped lanes. Enforced by eslint-rules/no-adhoc-regex-escape.cjs (shape-matched, with structural provenance for reviewed pattern-fragment constants rather than a name heuristic) plus a whole-tree companion guard covering the directories ESLint's globs miss. * fix(#3412): close the _SOURCE guard evasion, correct two false claims Three findings from the orthogonal review pass, all fixed. 1. The ESLint rule's `_SOURCE` provenance fallback was pure identifier- name matching with no binding check, so `new RegExp(userInput_SOURCE)` — a function parameter — sailed past the guard. That is the same rename-evasion class issue #3410 documents, reopened by the very fallback meant to complement the structural check. Now bound to the identifier's actual binding kind: import, require-derived const, or module-scope const; parameters, `let`/`var`, and unresolvable bindings fail closed. Four RuleTester cases cover the evasion and prove the legitimate cross-module case still passes. 2. src/pattern.cts's own header carried the stale pre-correction counts (12 copies / 17 call sites) while CONTEXT.md and the design doc carried the corrected ones (~39 / ~44) — a self-contradiction inside the PR whose entire purpose is deleting divergent copies. Rewritten, preserving the durable lesson: a named-function census cannot see inline copies; only a shape-matching guard can. 3. The claim that all deleted copies threw TypeError on non-string was false. phase-id.cts's copy — the one with 8 external importers — did String(value).replace(...) and never threw. The seam's locked signature does not coerce, so this is a real, now-disclosed behavior change rather than the pure preservation the tests asserted. Audited all 32 invocations across the 8 importers and 6 in-file callers: every one is safe by construction (upstream truthy guard or a string-producing derivation), verified by runtime probe against the compiled modules rather than by TS compilation, which cannot see a runtime undefined. Corrected the false claim in both the test comment and the design doc, and added it to Known limits. * docs(#3412): add Changed changeset for the Node 24 floor The only user-visible break in this phase. The escape-behavior change is internal and match-equivalent, so it carries no user-facing note. * fix(#3412): resolve the seam's require graph in script fixtures and packaging Checkpoint 2 came back red with 90 failures on the node24 lane. Three distinct defects, all introduced by routing scripts/ through the new pattern seam, none reproducible by any local gate: 1. ~82 failures — tests/adr-index-gate.test.cjs and tests/removed-but-needed-lint.test.cjs copy a scripts/*.cjs into an mkdtemp fixture and spawn it there (necessary: those scripts resolve their scan root from __dirname/.., so running the real script would scan the real repo). Each harness hand-listed the dependencies to copy alongside. Adding require('../gsd-core/bin/lib/pattern.cjs') to gen-adr-index.cjs made both lists silently incomplete -> MODULE_NOT_FOUND, plus 17 downstream 'did not emit parseable JSON' failures from the same crash. Fixed as a class, not an instance: new tests/helpers/copy-script- fixture.cjs walks a script's transitive static relative-require graph and copies it, so dependencies are derived and never re-declared. It throws (naming the unbuilt artifact) instead of letting the child die with a bare MODULE_NOT_FOUND. Verified for all four seam-consuming scripts: gen-adr-index, lint-removed-but-needed, gen-loop-host- contract, sync-runtime-launcher. 2. 2 failures — scripts/ ships wholesale but eslint-rules/ does not, so the new scripts/lint-no-adhoc-regex-escape.cjs would be MODULE_NOT_FOUND in a published install (#2858 guard). Excluded from the tarball, matching the existing precedent for gen-emitted- baseline.cjs, which is excluded for the identical reason, and locked with a test modeled on that one. Confirmed against a real npm pack: 890 files, 0 from eslint-rules/, and gsd-core/bin/lib/pattern.cjs present (so the other four scripts' requires are legitimate). 3. 6 failures — tests/phase-id.test.cjs asserted the literal escaped source text ('0*29', 'PROJ-42'). RegExp.escape is match-equivalent to the retired hand-rolled escaper but NOT text-equivalent: it hex- escapes the leading character and all hyphens ('0*\x329', '\x50ROJ\x2d42'). Verified NOT a behavior change — 576 match decisions across all three real interpolation prefixes, zero divergence. Those tests now compile each source into the same heading regex src/roadmap.cts's searchPhaseInContent builds and assert what matches and what does not, including the 'i'-flag canonicalization the hex escape has to preserve. Re-pinning the new literals would have rebuilt the same brittleness one layer down. Adds a test for the property the escape exists for: a dot in '1.2' must not act as a wildcard. Also shares one definition of 'a require' between the packaging guard and the fixture copier, so the two cannot disagree about what they scan. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): refuse to copy a fixture dependency outside the fixture root copyScriptWithDeps resolved each relative require and joined the repo-relative result onto fixtureRoot. A require resolving OUTSIDE the repo yields a '../'-prefixed relative path, so path.join climbed out of the fixture and wrote into the surrounding temp dir (verified: repoRoot=/repo + depAbs=/etc/passwd wrote /tmp/etc/passwd). No script in the tree does this today, so this closes an available escape rather than an active one. Refuses via the existing unresolved- require path so the failure names the offending specifier. Covered by a negative proof that the guard fires and that nothing lands outside the fixture. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): parse requires instead of pattern-matching them; restore the foreign-prefix contract Applies all findings from the second orthogonal review round, re-run because real code changed after round 1. HIGH (security) — extractRequires stripped BLOCK comments before LINE comments, so a '//' comment containing '/*' opened a phantom block comment, and a '//' inside a string literal truncated the line. Both hid real requires: 'const u="http://x"; require("./real.cjs")' returned [], and four real requires in gsd-core/bin/gsd-tools.cjs were invisible. Replaced with a real AST parse via espree. This is ADR-3212's own Decision 4 — tokenizer-first for stateful grammars — applied to the case it describes; comment/string/regex nesting is exactly such a grammar, which is why the regex version was wrong. The function was moved byte-identical out of the #2858 packaging guard, so the bug PRE-DATES this branch and has been a live blind spot there: a shipped script could have required an unshipped path undetected. Fixing it makes that guard strictly stronger than on next. espree is promoted from a transitive eslint dependency to an explicit devDependency rather than relying on hoisting. The script parse attempt sets ecmaFeatures.globalReturn because Node wraps CommonJS bodies in a function, making a top-level return legal — scripts/check-coverage-gate .cjs relies on it, and without the flag the guard throws on a file it is supposed to scan. Verified 0 unparseable across all 324 .cjs/.js under scripts/, bin/, and gsd-core/bin/, and 0 new violations against a real npm pack, so the exact extractor does not newly fail the guard. MEDIUM (security) — the repo-containment check guarded dependencies but not the entry path. One escapesContainment predicate now guards both. LOW (security) — containment was lexical while fs follows symlinks, and a directory symlink could mint a fresh dedupe key per level. realpath now resolves both repoRoot and each dependency before the decision, and the realpath-derived path is the dedupe key. Destination layout still uses the original repo-relative path, so copied trees are unchanged. MAJOR (standards) — the round-1 behavioral rewrite of phase-id tests lost the foreign-prefix contract: every assertion was satisfied by an impl returning [A-Z]+\x2d42, i.e. ANY project code — the exact #3599 bug class the exact-source prevents. The literal assertions it replaced were catching this. Now asserts the compiled regex REJECTS a different prefix with the same number. MAJOR (standards) — the test hand-duplicated production's heading regex with no parity guard (CLAUDE.md's 'Generative Fix Divergence'). Removed the parallel surface instead of policing it: src/roadmap.cts exports buildPhaseHeadingRegex, searchPhaseInContent calls it, the test imports it. Byte-identical .source and .flags verified for both escaped forms. MINOR — '..foo' no longer false-flagged as an escape; the inverted spurious-vs-missing doc claim corrected; the dead allow-test-rule header removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3412): backfill changeset pr number to 3416 * fix(#3412): make the escape guard's own regex linear, reword an injection-scan collision Two CI failures on PR #3416, both in code this branch added. CodeQL js/redos (high) — REPLACE_CALL_RE's outer alternation let a bracket run be consumed EITHER by the character-class branch OR one character at a time by the trailing catch-all, so a failing match explored both parses of every pair. Measured on the real regex: n=26 -> 204ms, n=28 -> 791ms, n=30 -> 3475ms, a clean 2^n. This script scans repo source, so a file with a long bracket run after '.replace(/' would hang CI outright — a guard against undisciplined pattern construction was itself the worst pattern in the diff. Fixed the way ADR-3212 already prescribes: the catch-all branch now excludes '[' and ']' so a bracket can only be consumed by the class branch (this is what makes it linear), and every quantifier is bounded (the locked bounded-quantifiers decision) as a second line of defense. Now 0ms at n=2000. Disclosed coverage tradeoff, recorded at the constant: a regex literal with a BARE unescaped ']' outside a class is no longer matched by this backstop. No census shape has that form, and the AST rule remains the primary detector. Verified the guard did not go blind doing it: a real census-shape violation is still reported, and an allow-adhoc-regex-escape suppression comment is still honored. Regression test drives the exported findViolations on a 2000-repetition adversarial input and asserts the RESULT. It makes no wall-clock assertion — elapsed-time tests are forbidden — so a regression surfaces as a harness timeout, which is the correct signal. Prompt injection scan — 'must not act as a regex wildcard' in a test comment matched the scanner's jailbreak pattern act\s+as\s+(a|an|if| my). Reworded to 'behave as'. Deliberately NOT allowlisted: silencing a whole test file over one phrase would blunt the scanner permanently, and the comment has nothing to do with injection. Neither failure was reachable from the remote runner — CodeQL and the injection scan are not in that matrix, so the sha it passed was green and still wrong. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
86101ee612 |
fix(#3133): keep global Claude @-references on tilde, not $HOME (#3393)
* test(#3133): global Claude @-references must resolve on tilde, not $HOME Regression for #3133: on a global Claude install, _applyRuntimeRewrites rewrites @~/.claude/... -> @$HOME/.claude/..., but Claude Code does not expand $HOME in @-file references, so the include silently resolves to nothing and the skill loads an empty execution_context. Rows 1/2/6/9 fail RED on next; rows 3-5/7-8 guard #1284 (shell $HOME), local redirect, #3503 (no homedir leak), and non-Claude isolation. * fix(#3133): keep global Claude @-references on tilde, not $HOME On a global Claude install, _applyRuntimeRewrites rewrote every @~/.claude/ include to @$HOME/.claude/ because computePathPrefix returns the $HOME form for shell-context correctness (#1284: ~ does not expand inside double quotes). Claude Code does not expand $HOME in @-file references, so the rewritten include silently resolved to nothing and skills loaded an empty execution_context. Add a normalization pass in the claude case: when the prefix is the $HOME (global) form, restore @$HOME/.claude/ -> @~/.claude/ (the form Claude expands and the shipped tarball uses). Shell contexts keep $HOME; local installs (absolute prefix) are unaffected; computePathPrefix is untouched (all its tests stay green). * test(#3133): fix row 9 assertion — drop end-of-string anchor Row 9's regex used $ (end-of-string) but a realistic @-ref line ends with a newline, so the anchor could not match. The fix under test produces the correct @~/.claude/gsd-core/references/ui-brand.md tail; only the assertion was wrong. Drop the anchor and assert the tail + no @$HOME. * docs(#3133): add changeset * docs(#3133): backfill changeset PR number (3393) --------- Co-authored-by: sim <sim@local> |
||
|
|
2076d450d7 |
fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not the runtime name (#2728)
* fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not runtime name quick.md and diagnose-issues.md kept the pre-#2584 `RUNTIME != "claude"` worktree gate, so every non-Claude runtime failed closed regardless of the capability it negotiated — including Codex, which declares orchestrator-worktree. Route both through the negotiated dispatch.isolation seam via a new shared reference, and migrate the two execute-phase reference fragments that carried the same runtime-name gate. - new gsd-core/references/dispatch-isolation-gate.md: canonical ISOLATION resolution, harness-flag resolution, single-agent degrade rule - quick.md / diagnose-issues.md read the gate; dispatch uses the {harnessFlag} placeholder rather than a hardcoded isolation="worktree" - execute-phase-wave-guard.md / execute-phase-between-wave-reset.md: migrate [ "$RUNTIME" = "claude" ] -> [ "$ISOLATION" = "harness-worktree" ] - every degrade site now clears BOTH USE_WORKTREES and ISOLATION; clearing one dispatched an isolated agent with no base guard and no manifest - parity guard in host-integration.test.cjs scans workflows AND references and matches six reintroduction shapes - migrate four tests that pinned the pre-#2584 runtime-name contract Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): use the /gsd:<cmd> namespace in the isolation degrade messages The degrade warnings cited /gsd-execute-phase, the retired hyphen form that slash-command-namespace.test.cjs rejects in Claude-facing source. Same length, so the quick.md size budget is unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2652): add changeset for PR #2728 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): normalize dispatch-site paths to forward slashes for Windows path.relative() returns backslash-separated paths on Windows, so the #2652 dispatch-site parity test compared "gsd-core\workflows\quick.md" against the hardcoded forward-slash literal "gsd-core/workflows/quick.md" and failed on every windows-latest CI lane. Normalize with .replace(/\\/g, '/'), matching the existing convention used elsewhere in this suite (e.g. tests/branch-no-track-guard.test.cjs:37). * test(#2652): restore the size-growth acknowledgment The rebase dropped tests/emitted-drift-ack.json. #2757/#2758 fixed the ATTRIBUTION axis, but the SIZE-GROWTH axis is independent: diagnose-issues.md (+2086) and quick.md (+230) still need an ack naming them and saying why. Verified: 65/66 without it (both files named), 66/66 with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): convert execute-plan.md Pattern A onto the dispatch-isolation gate Pattern A hardcoded `isolation="worktree"` — Claude Code's own literal — gated only on `workflow.use_worktrees`, with no capability negotiation at all. It is the same defect #2652 fixes at the other four sites, just a different shape: the file contains no RUNTIME variable, so the new detector correctly does not flag it. Concrete break: a Codex user who follows this PR's own newly-documented pattern and sets `workflow.use_worktrees: true` to get isolated dispatch via /gsd:quick then runs a plan through /gsd-execute-plan Pattern A, and hits an unconverted path — either an Agent() call erroring on an unrecognized parameter or silent unisolated execution, depending on host tolerance. Pattern A is a single-agent dispatch site through the host's own subagent tool, so it takes the same treatment as quick.md and diagnose-issues.md: resolve ISOLATION/HARNESS_FLAG through the canonical reference, degrade to sequential on orchestrator-worktree hosts, and substitute the host's declared {harnessFlag} instead of Claude Code's literal. while the area was open. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2652): add the INVENTORY row for dispatch-isolation-gate.md, refresh CONTEXT Two bookkeeping gaps flagged in review: INVENTORY.md had no row for the new gsd-core/references/dispatch-isolation-gate.md. INVENTORY-MANIFEST.json was regenerated correctly and its --check only diffs a live directory scan against the committed manifest, so CI passed regardless — but gen-inventory-manifest.cjs's own stderr guidance says to add the matching INVENTORY.md row. This is the repo's named "Inventory Drift" pattern. Placed with the dispatch/isolation cluster (worktree-branch-check, runtime-aware-dispatch) rather than alphabetically, matching how that table is grouped. CONTEXT.md's Host-Integration Interface entry still described dispatch.isolation as "declared and negotiated but not yet consumed by any scheduler — Phase 1 of #2584". That was already stale before this PR (execute-phase graduated in Phase 3) and more so now with three single-agent dispatch sites consuming it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): detect reversed-operand runtime gates; add a permutation property All five reintroduction regexes assumed $RUNTIME on the LEFT of the comparison, so `[ "claude" != "$RUNTIME" ]` — the same gate written backwards — evaded every one of them. Verified against the old patterns before fixing: all four reversed shapes (single bracket, double bracket, test builtin, JS template) scored EVADED. Each comparison shape is now generated in both operand orders from a single template, so a shape cannot be added in one order and forgotten in the other. The mutation table gains the four reversed cases. Also adds the fast-check property review suggested in place of the hand-rolled cases: it generates the cross product of the axes an author actually varies — bracket form, operator, operand order, quoting, spacing, runtime id — so a permutation the hand-written patterns miss surfaces here rather than in production. The 11 explicit cases stay as named regression anchors. execute-plan.md joins the scan's required-identities list now that it is a converted dispatch site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): acknowledge the execute-plan.md size growth The Pattern A conversion adds 811 bytes to an emitted workflow. Per #2719 the size axis needs its own acknowledgment, independent of attribution. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): repin the execute-plan.md PROSE_ALLOWLIST line after the rebase The #2751 command-position gate pins its prose exemptions by line number. This branch inserts the dispatch-isolation resolution above the `validated downstream by gsd-tools uat classify-coverage` sentence, moving it from execute-plan.md:387 to :397 — which fired the gate twice for one displacement (an un-allowlisted mention at 397, a stale entry at 387). The prose itself is unchanged from next; only the pin moves. Fixes #2652 * fix(#2652): gate the #2649 base-check on ISOLATION in diagnose-issues.md The rebase onto next merged #2649's pre-dispatch base-check textually, but its degrade flipped USE_WORKTREES after ISOLATION was already resolved, so the degrade never reached the dispatch decision. Gate the block on ISOLATION = "harness-worktree" and degrade ISOLATION itself, the same pairing quick.md already uses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): key quick.md post-dispatch bookkeeping on ISOLATION, not the Claude literal Review Blocker: the manifest append (l.822), worktree merge-back (l.825), and its skip clause (l.839) all conditioned on the literal isolation="worktree" — Claude Code's own rendering of {harnessFlag}. Cursor renders --worktree, so a newly-unblocked isolated Cursor run created a worktree whose committed work was never merged back and never cleaned up, silently. All three now key on ISOLATION = "harness-worktree" at dispatch. The existing parity detector cannot catch this class (its ISOLATION_TOKEN treats the literal as a legitimate marker), so this adds a dedicated literal-condition detector with a discrimination proof against both pre-fix sentences, a benign-mention control, and a positive pin on all three re-keyed conditions. Verified fail-first against the pre-fix quick.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): scope the use_worktrees=false install stamp to isolation=none runtimes `_stampNonClaudeRuntimeDefaults` rewrote every non-Claude runtime's `workflow.use_worktrees` read to `--default false`. That default resolved before `gsd_run query dispatch-isolation` was ever consulted, so the five runtimes that declare worktree support — cursor (harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree) — got ISOLATION=none regardless of what they negotiated. The gate this PR migrates dispatch onto was therefore still deciding isolation by runtime name, one layer down. The stamp's #1521 premise was that worktree isolation *was* Claude Code's isolation="worktree" spawn parameter, which no other host honored. #2584 replaced that premise with the negotiated capability. The stamp is now scoped to runtimes whose negotiated isolation really is `none`, where the default it writes is the outcome the resolver reaches anyway. `_negotiatedDispatchIsolation` mirrors routeDispatchIsolation's resolution against the same registry — closed vocabulary, a harness-worktree host must declare its flag, an orchestrator-worktree host must carry a descriptor that resolves — and fails closed to `none` on anything else, so an undeclared or unknown runtime keeps today's behavior. Two #1515 tests pinned the superseded premise for codex and are re-pointed at the new contract rather than deleted: the safety property they protect is now held by the isolation gate's fail-closed resolution, not by a name-scoped install-time default. Verified fail-first — all five assertions red against the pre-fix source, green after. * test(#2652): acknowledge the emitted ripple and re-point the end-to-end stamp proof Scoping the use_worktrees stamp changes emitted output, and two gates caught it. `gsd-core/workflows/execute-phase.md` now differs at emit time for the five hosts that declare worktree support (cursor harness-worktree; codex, opencode, kimi, kimi-code orchestrator-worktree) — the source file is byte-identical, only the stamp is gone. Acknowledged in this PR's fragment. `tests/install.test.cjs`'s real-install assertion pinned the superseded premise end-to-end, asserting codex receives `--default false`. Re-pointed rather than deleted, matching the two unit tests: it now proves codex keeps the unstamped `true` read. A second arm installs windsurf — which declares isolation `none` — and asserts the false stamp is still applied there, so the change cannot silently degrade into "never stamp" without a test noticing. The ack entry collides with `2658-trae-instruction-file-path.json`, which is fully spent (merged via #2925, so all 25 of its entries are present at base and gate nothing) and is pruned for the same reason and by the same rule as the spent `2649-*` fragment this PR already removed. #2566 prunes the same file for the same collision on `new-project.md`; a delete/delete merges cleanly either way, and the base-side cleanup would make both unnecessary. * fix(#2652): re-record the sentinel when a dispatch site degrades isolation Review Blocker B1/B2/B3. Every isolation degrade in a dispatch site is decided in shell, where routeDispatchIsolation cannot see it. That resolver persists whatever it resolved to the run-scoped sentinel as an unconditional side effect (#3045), so a degrade that only reassigns $ISOLATION leaves the sentinel asserting harness-worktree while the dispatch correctly omits the harness flag. The shipped PreToolUse guard reads the sentinel at the instant of the Agent() call and denies that mismatch with exit 2 — the work does not run unisolated, it does not run at all. Latent on this branch and lands on rebase, since |
||
|
|
c2f24265f2 |
feat(#2870): resolve install scope as a value (#3278)
* test(#2870): failing-first suite for the Install Scope Module 19 tests over the 50-test-matrix rows 1-19. RED by construction: the module under test does not exist yet, so the suite fails at require with MODULE_NOT_FOUND until src/install-scope.cts lands. Every row asserts a returned value with injected env/home/existsSync -- no filesystem, per the issue's acceptance criterion that tests assert the resolved value directly. Row 7 asserts the RELATION rank(global) > rank(local) rather than a literal, so Phase 2 (#2871) can re-base the numbers without a fixture edit. Row 4 iterates the real runtime registry rather than a hardcoded list, excluding vscode, which declares configHome.kind none and is never CLI-installed. * feat(#2870): add the Install Scope Module Scope becomes one resolved value instead of a bare string re-derived at every layer. resolveScope({id, runtime, ...}) returns {id, configHome, settingsFile, consentRequired, hostPrecedenceRank}. It COMPOSES resolveConfigHomeFromDescriptor rather than extending it. That function has 60 dependents across 13 files and 2 process flows -- a CRITICAL blast radius -- so adding a scope parameter to it, which the issue's framing invites, would ripple through all of them. Composing costs nothing and leaves every existing caller byte-identical. The module owns the InstallScope type name, which previously lived privately in runtime-artifact-install-plan.cts; that module now imports it. A fifth spelling of the same concept would have defeated the phase. settingsFile is null for the 18 runtimes that declare no settingsFileByScope -- absence is a value, not an error, and inventing a Claude-shaped default would leak that host's shape onto every other one. hostPrecedenceRank ships unread: Phase 2 (#2871) is its first consumer. It is carried as data only, per this issue's out-of-scope note that precedence semantics belong to that phase. Vocabulary: the install axis standardizes on local. ConsentRecord.scope keeps project deliberately -- that literal is serialized into consent records in the user's home, and renaming it would silently deactivate every project-scoped capability on the machine. CONTEXT.md records the boundary mapping instead. Every environmental input is injectable (env, home, existsSync, cwd), so the resolved value is assertable with no filesystem at all. Registration ripple: .gitignore, eslint.config.mjs, CONTEXT.md glossary, docs/INVENTORY.md, and the inventory manifest (regenerated after build:lib, never before). Verified via the remote runner. * refactor(#2870): route scope re-derivations through the module bin/install.js resolves scope once per function instead of inline at each of its 12 sites, and the settingsFileByScope consumer reads it through resolveScope(). Seven downstream boolean re-derivations now call the module's isGlobalScope() instead of comparing the literal independently: runtime-artifact-install-plan, both runtime-artifact-layout kind builders, dispatchKindEntry, surface, and two install-engine sites. The fifth through seventh were not named in the issue -- they are the same re-derivation class, and leaving them would have made the acceptance criterion false. _computePathPrefix keeps its isGlobal boolean API, so the projection is centralized rather than eliminated. resolveScope and isGlobalScope share one validator, so the two surfaces cannot drift. TWO SITES DELIBERATELY NOT ROUTED: runtime-artifact-conversion's rewriteStagedSkillBodies and rewriteStagedCommandBodies. ADR-1508 fixes the direction as installer/layout -> conversion, never upward, and install-scope composes runtime-homes, so importing it into the conversion module would invert that direction. Left as-is on purpose. Behavior-preserving throughout. Each step was proven by capturing full layout and plan output -- including every kind's home field and the hashed contents of emitted files -- before and after, across both scopes for claude, codex, opencode, hermes, kimi and kilo. Byte-identical. surface.cts keeps a scope ?? 'global' default before the call because Layout.scope is optional there; isGlobalScope throws where the old inline compare returned false, and that difference would have been a placement regression. Verified via the remote runner. * fix(#2870): cover the no-config-home throw and document the strictness Two findings from the isolated adversarial review. The vscode case was implemented but untested. resolveScope throws for a runtime whose descriptor declares configHome.kind 'none', which is the design's own behavior-table row 13, but the registry sweep excluded vscode rather than asserting the throw -- so the behavior shipped with no test. The exclusion is now legitimate because the case has its own test naming the runtime in the assertion. isGlobalScope throws where the inline compare it replaced returned false. No reachable caller can deliver an out-of-union value today, but the types are not enforced at runtime, so a future caller passing an optional Layout.scope would crash rather than silently misroute. That is the better failure -- misrouting writes artifacts to the wrong place -- but it was undocumented, so the reason is now on the function. Adds the changeset the acceptance criteria require. * refactor(#2870): route the last two sites; correct the ADR-1508 claim The previous commit declined to route runtime-artifact-conversion's rewriteStagedSkillBodies and rewriteStagedCommandBodies, claiming ADR-1508's dependency direction forbade the import. That reasoning was wrong, and this commit corrects it. Two independent reviewers checked the actual import graph: runtime-artifact-conversion already imports capability-registry, command-roster, runtime-name-policy and shell-command-projection -- it depends on leaf-tier siblings today. install-scope imports only runtime-homes plus node builtins, and runtime-homes imports only node builtins, so there is no cycle at any depth. ADR-1508 governs the installer/layout to conversion boundary, not a leaf-to-leaf sibling import of the same shape conversion already makes. With those two routed, every isGlobal re-derivation in the tree now goes through one owner and acceptance criterion 1 is fully met rather than partially. Nine sites, not the four the issue enumerated. Also from the review: Tests were falling through to the real process.cwd() at five local-scope call sites, which contradicts the acceptance criterion that the resolved value be assertable with no filesystem. Every one now injects a cwd. One of the five was a site the review had not spotted. bin/install.js carried two near-identical copies of the guarded resolveScope block, one in install() and one in uninstall() -- duplicated scope logic in the phase whose purpose is removing it. Extracted to one helper, and the new sites use the file's existing ternary idiom rather than the if/else that replaced it. Equivalence re-proven across both scopes for claude, codex, opencode, kilo and hermes, now including the staged skill and command body rewrites hashed per file, since those decide the literal spec-root path baked into every emitted artifact. Byte-identical. Verified via the remote runner. * fix(#2870): assert configHome portably instead of with a native separator The windows-latest node24 shard failed on two install-scope assertions. The module was right and the tests were wrong: they built their expected value with path.join, which emits \fake\home\.claude on Windows, while resolveScope normalizes separators unconditionally to /fake/home/.claude. That unconditional normalization is deliberate -- backslash paths arrive on Linux too, so normalizing via path.sep is the documented defect this repo guards against. Weakening it to make the assertion pass would have inverted the fix. Every path.join-built expectation in the suite now goes through toPosixPath from tests/helpers.cjs, which is the pattern the no-path-literal-in-assert rule's own valid-case list sanctions. It splits on the running platform's path.sep and rejoins with forward slashes, so it reverses whatever path.join produced on that same platform and the expectation is invariant everywhere. Two more call sites had the same latent problem and passed on Linux and macOS by luck; they are fixed too. This is the class of defect the remote runner structurally cannot catch -- its matrix is Linux-only, so a green pass there is not evidence of portability, and CI's Windows lane is the only place it surfaces. Verified via the remote runner. --------- Co-authored-by: sim <sim@local> |
||
|
|
b181c2f8c3 |
fix(#3039): clamp max/xhigh effort to high for Claude-runtime skills (#3119)
* fix(#3039): clamp max/xhigh effort to high for Claude-runtime skills effort: max in plan-phase, execute-phase, and autonomous SKILL.md frontmatter was passed through as output_config.effort, which the Anthropic API rejects when extended thinking is disabled (400: effort 'max' is not supported when thinking is disabled on this model). The frontmatter is static at install time and the installer cannot know whether thinking will be on or off at invocation. normalizeClaudeSkillEffort now clamps both 'max' and 'xhigh' to 'high' — the maximum value that works in both thinking states on all supported models. Applied in both src/runtime-artifact-conversion.cts and bin/install.js. * chore(#3039): backfill changeset PR number 3119 * fix(#3039): regenerate skills with clamped effort: high --------- Co-authored-by: sim <sim@local> |
||
|
|
4926c2e904 |
fix(#3004): update Codex adapter collaboration-tool vocabulary (#3104)
* fix(#3004): update Codex adapter collaboration-tool vocabulary The generated Codex skill adapter documented stale tool vocabulary: - wait(ids) → collaboration.wait_agent(timeout_ms=...) (the real tool), with explicit disambiguation from the unrelated exec-cell functions.wait - close_agent(id) unconditional → gated on tool visibility (same schema- detection pattern already used for spawn_agent's agent_type field) - Missing required task_name field and fork_turns parameter → added alongside the existing fork_context guidance (coexist, not replace) Updated the regression test to assert the new vocabulary (wait_agent not wait(ids), functions.wait disambiguation, task_name, fork_turns, tool_search gate on close_agent). * chore(#3004): backfill changeset PR number 3104 --------- Co-authored-by: sim <sim@local> |
||
|
|
88f6d9bd1b |
fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu * fix: preserve installer executable mode * chore: add changeset for PR #2812 * test(#2644): acknowledge Cursor emission changes * test(#2644): drop spent emitted drift acknowledgments * fix(#2644): remove retired Cursor command converter --------- Co-authored-by: clezcoding <clezcoding@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
51f32d2d40 |
fix(#2658): detect trae runtime and resolve its instruction file to a concrete rules file (#3006)
* test(#2658): add failing-first regression for trae runtime detection and instruction path Covers all three collided defects reported in #2658 plus a fourth instance of defect 1 (ingest-docs.md) found while diagnosing it: missing trae detection in workflow runtime-detection blocks, the CLAUDE.md path-mutilation bug in both the js/cjs and md install-time converters, and the missing projectInstructionFile capability declaration. Fails against current source; the next commit fixes it. * fix(#2658): detect trae runtime and resolve its instruction file to a concrete rules file Three defects collided to produce the reported ".claude/.trae/rules/" path: 1. new-project.md and ingest-docs.md's runtime-detection blocks only recognized codex/gemini/opencode and fell through to RUNTIME=claude for trae. Both now recognize the /.trae/ execution-context path and the TRAE_CONFIG_DIR env var before the claude fallback. 2. RUNTIME_CONTENT_DISPATCH.trae.js (bin/install.js) replaced bare "CLAUDE.md" before the ".claude/" prefix was handled, mutilating ".claude/CLAUDE.md" into ".claude/.trae/rules/". Now replaces the full ".claude/CLAUDE.md" path first, and targets a concrete file. 3. convertClaudeToTraeMarkdown (mirrored in bin/install.js and src/runtime-artifact-conversion.cts per the #2094 output-parity test) had the same class of bug with a different wrong output (".trae/.trae/rules/", from its generic ".claude/" rewrite firing after the bare CLAUDE.md rewrite). Both mirrors now match full-path forms before the bare/generic patterns, converging on the same concrete file as the js/cjs converter. 4. capabilities/trae/capability.json didn't declare hostBehaviors.projectInstructionFile, so getProjectInstructionFile fell through to the generic AGENTS.md default even when RUNTIME=trae was resolved correctly. Now declares ".trae/rules/rules.md", regenerated into gsd-core/bin/lib/capability-registry.cjs via npm run gen:capability-registry. Closes #2658 * chore(#2658): add changeset * fix(#2658): preserve arbitrary runtime-dir prefixes in the trae path rewrite Found by the end-to-end --trae install regression test (not by static trace) across two verification runs: 1. copyWithPathReplacement runs a generic ~/.claude/, $HOME/.claude/, and ./.claude/ -> runtime-dir rewrite on every .md file BEFORE calling convertClaudeToTraeMarkdown. The prior fix's .claude/CLAUDE.md-specific patterns never fire on that already-rewritten text, and the bare fallback still doubled whatever prefix the generic pass substituted. A first attempt handled only the fixed "./.trae/" shape and missed the $HOME/.claude/ and ~/.claude/ forms gsd-core/workflows/profile-user.md actually uses, which post-rewrite become an arbitrary absolute local-install-root path, not the fixed relative shape. Fixed with a prefix-preserving pattern that captures whatever precedes a ".trae/" tail and fixes only the filename suffix, instead of assuming one fixed shape. 2. The fix's own explanatory comments literally spelled out the malformed strings and the instruction filename as contiguous text. Since these two files ship verbatim into local --trae installs, where they are themselves run through the same find/replace, the comments got "fixed" right along with the real code, leaking the malformed string into the installed tree. Rewrote every comment in both mirror copies to never spell either the instruction filename or a malformed shape as one contiguous token. Adds an emitted-drift-ack fragment: the corrected replacement target for every CLAUDE.md mention (bare directory -> concrete file) changes trae-emitted output for every repo file that mentions CLAUDE.md, not only the ones that hit the originally reported bug. * test(#2658): extend parity test with arbitrary-prefix .trae/ inputs The bin/install.js vs runtime-artifact-conversion.cjs parity assertion for convertClaudeToTraeMarkdown only fed the pre-existing bare .claude/CLAUDE.md input through both implementations. Feed the prefix-preserving cases (relative, nested-absolute, tilde, $HOME, backtick-wrapped) plus a property-based check through both, so a future edit to only one copy of the .trae/-tail regex fails this test instead of silently diverging. * chore(#2658): backfill changeset PR number to 3006 --------- Co-authored-by: sim <sim@local> |
||
|
|
cc3ee301a7 |
fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root (#2593)
* fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root installSharedHooksBundle wrote `{"type":"commonjs"}` over <configRoot>/package.json unconditionally — no existence check, no merge, no backup — on every install and every /gsd-update re-install. On the 11 affected runtimes that file is often user-owned; on OpenCode and Kilo it is the documented place to declare local-plugin npm dependencies, so a user's name/type/dependencies/scripts were destroyed on each run. The uninstall path already read the file and unlinked it only on an exact content match. That asymmetry was the defect: the discipline existed in the codebase, it just was not applied on the write side. Move the marker into the directories GSD creates and fills with its own .js files — hooks/ (all shared-hooks runtimes, incl. Kimi's own root) and the nativePlugin dir (plugins/ for OpenCode+Kilo, extensions/ for pi) — and stop writing the config root entirely. New src/commonjs-marker.cts owns the marker string plus one ownership predicate (absent / gsd-owned / foreign, fail-closed on an unreadable file) shared by ensureCommonJsMarker and removeCommonJsMarker, so install and uninstall cannot drift apart again. Nothing else depended on the config-root marker: package identity is baked at build time (#378/#498) and version resolution prefers gsd-core/VERSION and already tolerates a missing root package.json (#1383) — Codex has installed without one all along. A package.json in plugins/ or extensions/ is inert to plugin discovery, which globs *.{ts,js} only (see installer-migration 006). Uninstall retires the pre-fix config-root marker, so upgrading users are cleaned up on removal, and still never touches a file it did not write. * fix(#2544): point the changeset fragment at the filed PR The fragment's `pr:` field is only knowable after `gh pr create` returns. * fix(#2544): register commonjs-marker.cjs in the tsc-generated ESLint ignore set bin/lib/commonjs-marker.cjs is tsc output (src/commonjs-marker.cts is the linted source), so it belongs in the ADR-457 ignore list like its siblings. Clears the lint-tests no-var failure and the repo-invariants "linted xor ignored" migration-state test. * fix(#2544): pin the kimi CommonJS marker to hooks/, not the ~/.kimi root The UPGRADE 1 test still asserted the pre-#2544 marker location (~/.kimi/package.json). The marker now lives inside ~/.kimi/hooks — the directory GSD itself creates — matching the updated golden-install-parity and install-tree fixtures. Also asserts the root marker is NOT written. * fix(#2544): make the CommonJS marker write path non-fatal Review round 2, Major 3 + Minor 1 + the stagedHooks nit. ensureCommonJsMarker rethrew any non-EEXIST write error and neither call site caught it, so EACCES on a read-only hooks/, EROFS, or ENOSPC aborted the whole install with a raw stack trace. Every other marker interaction in the module is best-effort — removeCommonJsMarker swallows unlink failures, classifyMarker swallows read failures — and this was the write path, i.e. the one most likely to fail on a locked-down config dir. It now returns a new 'failed' outcome and both call sites warn and continue. Sibling found while sweeping for the same defect class: fs.mkdirSync sat OUTSIDE the try block, so an unwritable parent threw past the guard entirely. Creating the directory is the same environmental hazard as writing into it, so it moved inside. Also in this file: - The hooks marker is now gated on `stagedHooks && hooksOk`, not stagedHooks alone. stagedHooks is computed from the SOURCE listing before the copy loop, so it stays true when the copies land but verifyInstalled() then fails — marking a hooks/ GSD did not successfully populate claims an ownership the install did not earn. - The uninstall rmdir of the native plugin dir is gated on GSD having actually removed something from it. Hoisting it out of the adapter-exists guard (so the marker-only case could prune) had silently widened it into deleting a user-created but empty plugins/ or extensions/ dir — the same "don't touch territory GSD didn't fill" principle this issue is about, inverted. - Kimi's pre-#2544 marker at its native hook root (~/.kimi) is retired at the same call site that writes its replacement. That path is outside kimi's configDir, so installer-migration 007 structurally cannot reach it. * fix(#2544): retire the stale config-root marker via installer-migration 007 Review round 2, Major 1 — the PR's headline claim was false for existing installs. Upgraders kept BOTH markers: the new one under hooks/ and the stale {"type":"commonjs"} at the config root, so their config root stayed pinned to CommonJS and their dependency manifest stayed gone until they uninstalled. The migration is unusual in one way, and it is the part worth reviewing: the config-root marker was never recorded in gsd-file-manifest.json (writeManifest records hooks/, agents/, commands/, scripts/ and the native plugin, never a root package.json), so classifyArtifact answers 'unknown' for it and the planner's own guard downgrades a remove-managed on an 'unknown' classification to preserve-user. 007 therefore supplies the "purpose-built detector for an old GSD-owned shape" that docs/installer-migrations.md#remove-managed sanctions — exact content match, the same predicate removeCommonJsMarker has always used — and declares the resulting classification on the action. A package.json with any other content is left untouched, and there is deliberately no backup-and-remove branch: a non-matching file here is not a patched GSD artifact, it is somebody else's file. Scope is all runtimes. The `runtimes` field is OMITTED rather than `[]`: validateStringArray requires the field to be non-empty WHEN PRESENT, while the runtime filter treats an empty array as "all" — so `runtimes: []` throws at plan time and the migration never runs. The metadata test pins this. Kimi is a deliberate carve-out, named in the migration's own header: its marker lived at ~/.kimi, outside kimi's configDir, and migration relPaths are structurally confined to configDir. It is retired by the installer instead. Registration: shipped-migrations table, .gitignore for the emitted .cjs, the EXPECTED_CHECKSUMS baseline, and the ESLint ignore set. That last one is not copied from migration 006 by rote — 006 needs no entry because it imports nothing, while 007 imports node builtins, so tsc emits its __importDefault helper and the `var` in it trips no-var. This is the same lint gate that made round 1 red. * test(#2544): fault-injection and multi-runtime marker coverage Review round 2, Major 2 + Minors 4 and 5. Major 2 — CONTRIBUTING.md:514-531 is mandatory for install/uninstall flows and the suite had no fs monkeypatching at all. Every branch now covered is one whose doc comment claims it as the module's safety posture: - classifyMarker non-ENOENT lstat error -> 'foreign' (the fail-closed rule), with an ENOENT control alongside it so the test discriminates rather than just asserting one side - classifyMarker readFileSync throw -> 'foreign' (present-but-unreadable never downgrades to the permissive answer) — the fixture's bytes are exactly GSD's marker, so the test fails if the code ever answers on content it could not read - a DIRECTORY at the marker path (CONTRIBUTING:521; the symlink case was already covered with a real symlink, the directory case needs no injection at all) - the ensureCommonJsMarker TOCTOU EEXIST branch — the entire reason for flag:'wx' - the new 'failed' outcome, for both writeFileSync (EACCES/EROFS/ENOSPC) and the mkdirSync that used to sit outside the guard - removeCommonJsMarker unlink throw -> false These save and restore fs methods in `finally` rather than using chmod 0o000, which does not fault under root and would pass vacuously in root Docker and CI. Minor 4 — uninstall was driven for opencode only. pi's extensions/ and both kimi locations now have behavioral coverage, install and uninstall, each paired with a user-authored-file case proving GSD leaves it alone. Minor 5 — the stagedHooks gate had no assertion behind its stated reason. A pre-existing, GSD-untouched hooks/ directory is now driven through a runtime that declares skipSharedHooksInstall and asserted to stay marker-free, with its user content intact. Also regression-tests the uninstall rmdir gate from the previous commit: an empty plugin dir GSD removed nothing from must survive. * docs(#2544): correct stale marker prose, register the module, document the trade-off Review round 2, Minors 2, 3 and 6. Minor 2 — six files asserted the installed ROOT ships the synthetic marker. None was load-bearing (all three walk-up consumers are VERSION-first with try/catch and the marker never carried a `version`), but ADR-457:52 is the rationale for keeping a generated module, so a future reader would mis-derive the constraint from it. Each site is corrected to what is now true: the installed tree carries no package.json with a .name at all, because the only ones GSD stages are {"type":"commonjs"} markers and they now live in GSD's own directories. Two of the six needed more than a location swap. hooks/gsd-check-update-worker.js and the platform-gate test both described `require('../package.json').name` resolving to undefined; post-#2544 that require does not resolve at all, so the history is kept accurate and the present-tense claim corrected rather than just moved. And src/runtime-artifact-conversion.cts described the no-root-package.json case as Codex-only — it is now every runtime, which strengthens that comment's own argument for lazy resolution. The generated .cjs sibling needs no edit: it is gitignored build output, not a tracked file. Minor 3 — src/commonjs-marker.cts had no CONTEXT.md entry, unlike every peer module, and CONTEXT.md is the #2 co-change partner of bin/install.js. Added, including the fail-closed posture and the never-throws contract. Minor 6 — the plugins//extensions/ marker shadows the config root for all .js siblings, so an OpenCode/Kilo user's ESM plugin/*.js stays broken. That is exactly what #2544's Fix section prescribed and it is disclosed in the PR body, but the PR body is not documentation. It now lives in the OpenCode section of docs/how-to/install-on-your-runtime.md, stated as a real constraint rather than a pure improvement, with the .ts mitigation and a fallback for ESM plugins. * test(#2544): attribute the CommonJS marker in the emitted-provenance rules The differential emitted-attribution gate (#2723, landed on `next` after this branch was cut) went red on the macOS shards once this PR rebased onto it. Two distinct causes, both real gaps rather than noise: 1. `plugins/package.json` and `extensions/package.json` matched NO rule — the `native-plugin` rule covers `*.{js,cjs,mjs}` only, so the marker read as an unattributed emitted family. 2. `hooks/package.json` fell through to `hooks-built`, which attributes an emitted `hooks/<X>` to a repo source `hooks/<X>`. There is no `hooks/package.json` in the repo, so it resolved to a nonexistent path. Cause 2 is exactly the failure already documented three lines above it for Copilot's `gsd-session.json` — "a code literal, not a built script" — so the fix follows that precedent rather than inventing one: `package.json` is excluded from `hooks-built` the same way, and a dedicated `commonjs-marker` rule attributes the family across all four roots it can appear in (both hooks roots plus `plugins`/`extensions`) to the sources that actually emit it. Deliberately a RULE, not an entry in tests/emitted-drift-ack.json. An ack is for a one-off ripple and goes stale by design — the gate fails a stale ack precisely so it cannot pre-clear the next change on that path. These markers are a permanent part of the emitted tree from #2544 onward, so they need standing attribution. Verified by reproducing the CI failure locally with GSD_EMITTED_BASE: 3 provenance errors + 12 unattributed paths before, 35/35 green after. * fix(#2544): route the #2717 hooks-surface marker helpers through commonjs-marker #2717 landed a second copy of ensureCommonJsMarker/removeCommonJsMarkerIfGsdOwned in src/runtime-hooks-surface.cts for the runtimes that stage .js hooks via dedicated paths (cursor/windsurf/codex). That copy had drifted from this PR's module on the two properties that matter: - ownership probe: `fs.existsSync` FOLLOWS symlinks and reports false for a DANGLING one, so a dangling package.json symlink classified as absent and the write went straight through it. Demonstrated: against the pre-fix copy, ensureCommonJsMarker() on a hooks/ dir holding a dangling package.json symlink returns true and creates {"type":"commonjs"} OUTSIDE that directory. - create: a plain writeFileSync leaves the classify->write window open, where commonjs-marker creates with flag:'wx' (O_EXCL). Both helpers now delegate to src/commonjs-marker.cts, which is what this PR's own docstring already claimed was the single place these rules are enforced. Exported signatures are unchanged (still boolean), so bin/install.js and the #2717 tests are unaffected. The new subtest is the only coverage that fails if the duplicate is ever reintroduced — the two implementations agree on every non-adversarial input, so the existing suites pass against both. * test(#2544): pin the stagedHooks gate on zcode, not windsurf The Minor-5 coverage picked windsurf because hostBehaviors.skipSharedHooksInstall kept it out of the shared hooks bundle, so GSD staged nothing into hooks/ and the marker was correctly absent. #2717 changed that premise: cursor/windsurf/codex now stage their .js hooks via dedicated paths and get the marker beside those scripts. Measured on this tree, windsurf stages 2 .js hooks and receives a marker — so the assertion was pinning behaviour that is now wrong, not the gate it was written for. ZCode is the durable choice: per #1821 it has hooksSurface:'none' AND no plugin surface to spawn hooks, so GSD stages no .js there by either route (measured: 0 staged, no marker). The property under test is unchanged — a user-created hooks/ directory GSD never fills stays marker-free. * test(#2544): use the shared cleanup helper in the migration test Addresses the review's Major 1. The suppression's stated reason — "no helpers import available" — was not correct: tests/helpers.cjs exports cleanup, and the other test file added in this same PR imports it (tests/commonjs-marker.test.cjs). The local reimplementation dropped two protections that are live on this repo's windows-latest lane: the CWD guard (Windows cannot remove a directory that is the current working directory) and the 20 x 250ms retry budget that absorbs the deferred-scan handle Windows Defender holds on newly-written files. Local function and suppression both removed; local/no-raw-rmsync-in-tests now passes without one. * test(#2544): expect hooks/package.json for the #2717 runtimes The fresh-install contract table predates #2717, which stages cursor/windsurf/ codex .js hooks via dedicated paths and writes the CommonJS marker beside them. All three therefore now receive hooks/package.json legitimately. Measured on this tree: codex stages 3 .js hooks, cursor 6, windsurf 2 — each with the marker; cline/copilot/trae/zcode stage none and get none, so their contracts are unchanged. * fix(#2544): gate the #2717 marker writes on having staged something The three dedicated marker writers #2717 added ran unconditionally. Each one mkdirs hooks/ up front and stages its scripts conditionally on the source existing, so with an absent or empty hook source they created a directory, filled it with nothing, and marked it as GSD's anyway. That is the same write-into-someone-else's-territory this issue is about, and installSharedHooksBundle already guards the identical case with `stagedHooks`. The dedicated paths now carry the matching gate: - cursor / windsurf: `installedScripts.size > 0` - codex: a new `codexStagedHooks` flag. The enclosing guard only proves that hooks/dist EXISTS; it says nothing about whether any CODEX_HOOKS_TO_COPY entry landed. Covered for cursor and windsurf by driving each writer against a src tree whose hooks/ dir is empty. The codex leg is defensive and deliberately uncovered: its trigger state needs a package tree where hooks/dist exists but holds none of the allowlist, which is not constructible from a real checkout. * test(#2544): scope the commonjs-marker sources per root The rule declared one flat source list for every marker root, so `extensions/package.json` was attributed to runtime-hooks-surface.cts (which never writes there) and `.kimi/hooks/package.json` to install-engine.cts. That is not merely untidy. emitted-diff.cjs accepts the FIRST satisfied source, so a flat list containing bin/install.js let any change anywhere in that 13k-line file authorise marker drift for every root — the blanket escape hatch this file's own agents-verbatim comment refuses for exactly the same reason. Sources are now derived per root from ctx.rel. Note the rule ctx is `{ rel, runtime }` and carries no `root`, so keying on ctx.root would have sent every path down one branch silently. * test(#2544): state precisely what the zcode assertion pins The comment claimed the test pinned installSharedHooksBundle's `stagedHooks` gate. It does not, and neither did the windsurf version it replaced: zcode declares skipSharedHooksInstall, so the outer guard skips that helper entirely and the gate is never evaluated. The test passes on the runtime exclusion. What it does pin — the outcome a pre-existing, GSD-untouched hooks/ stays marker-free — is still worth having, and is what the review asked for. The two `staging zero hook scripts` tests are the ones that pin a real staged-nothing gate. Comment corrected rather than left implying coverage that is not there. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
628648d63a |
chore(#2931): cap emitted per-runtime bytes and single-source windsurf (#2984)
* fix(#2931): preserve protected regions and cap emitted per-runtime bytes Route every runtime brand swap through applyClaudeCodeBrandSwap so "Claude Code" survives verbatim inside <runtime_compatibility> regions (#2284b). The fix existed only in bin/install.js's local copies; the src/*.cts exports still used a naive replace, so binding install.js to the single source -- as this phase does for the Windsurf family -- would have silently regressed those runtimes. A table-driven parity guard now covers all nine brand-swapping converters. De-duplicate the Windsurf converter family: delete the six local copies in bin/install.js and bind the four exported ones by reference, guarded by reference-identity assertions (the ADR-1508/#1675 pattern). The two unexported helpers and an unused tool table go with them. Replace the Windsurf 12,000-byte throw with description truncation, matching the bound its sibling skill converter already applied. The throw could only fire on an ~11.7 KB frontmatter description: the largest emitted workflow is 311 bytes. Truncation makes the cap unreachable by construction and leaves 12,000 in exactly one place, eliminating the dual-surface duplication rather than testing for it. Add the emitted-byte cap gate: buildEmittedSizes captures LF- and <HOME>-normalized bytes from the walk buildParityManifest already performs, and evaluateEmittedCaps asserts them against a per-runtime cap table with dead-rule detection. buildParityManifest's return shape is deliberately unchanged -- diffEmitted compares its values with ===, so making them objects would report all 8,529 emitted paths as moved. A regression test pins the values as strings. Add a deterministic trim-safety gate over composeWithinBudget's omitted/shrunk/floored/isolatePrefix metadata, with an anti-vacuity rule, replacing the model-graded eval gate the issue described. * docs(#2931): correct ADR-1671 windsurf premise and trim-safety contract * fix(#2931): bound the windsurf command name and single-source the brand swap Review findings from the orthogonal passes, all fixed inline. The claim that removing the 12,000-byte throw left total emission "bounded by construction" was false. The #1615 regex constrains the character class but not the length, and commandName is interpolated three times into the emitted workflow: a 20,000-character name emitted 60,162 bytes silently. Add WINDSURF_COMMAND_NAME_MAX=128 as a separate, clearly-labelled size control that THROWS -- commandName is the @-ref path target, so truncating it would point the workflow at a file that does not exist (DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED). The #1615 security regex is untouched and still runs first. 128 is generous: the longest shipped name is gsd-plan-review-convergence at 27. Harmonize convertClaudeCommandToWindsurfSkill onto the code-point-safe truncation helper. It still used a UTF-16 slice(0,177) -- the exact surrogate-splitting bug the helper was written to avoid, in the very sibling the helper's comment cites as its model. Bounds are unchanged, so output is byte-identical for every shipped command (descriptions max out at 99 chars). Export applyClaudeCodeBrandSwap and bind it in bin/install.js, deleting the local copy. Adding it to the .cts left two unlinked implementations of identical logic -- the drift class this change exists to remove. Verified byte-identical across eight fixtures and five sequential calls before merging, and guarded by a reference-identity assertion. Convert three try/finally test bodies to t.after (CONTRIBUTING.md:344), add fast-check property coverage for the trim-safety contract, and use fc.pre instead of a bare return in a property callback. * test(#2931): fix three test-authoring bugs the remote matrix caught The remote runner returned 8 unique failures on 6f15cdeb8. All three causes were in the test files, not the modules under test -- local harnesses exercise the modules directly, so nothing executed the test bodies until the matrix did. `{ __proto__: [...] }` in an object literal sets the prototype instead of an own key, so the JSON round-trip erased it and the cap table never saw a reserved runtime key. The production rejection was already correct; the test could not reach it. Use a computed key. Two cap fixtures tripped orthogonal error paths rather than the paths they name: one declared windsurf in the cap table but omitted it from sizes (UNKNOWN_RUNTIME), the other left the sole windsurf pattern matching nothing (a genuine dead rule). Both now include a compliant artifact so the intended branch is what is asserted. The dead-rule and unknown-runtime contracts are deliberate and unchanged. `const { root } = makeSyntheticConfig({ ... `${root}` })` referenced `root` from inside its own initializer -- a temporal dead zone error. makeSyntheticConfig now optionally takes a (root) => files factory. Also raise the npm pack --dry-run bound 60s -> 120s in the shipped- scripts packaging test. That failure is NOT from this branch: the file is byte-identical to next, a fresh tsc measures 1.98s there vs 2.14s here, and the run recorded 60,637ms against a 60,000ms bound -- a timeout under 28,948-test parallel contention, not a slowdown. Fixed rather than deferred because a bound that tight is fragile regardless of which branch trips it. * chore(#2931): backfill changeset pr number to 2984 --------- Co-authored-by: sim <sim@local> |
||
|
|
6ad30f74b6 |
feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler. harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run. Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard. Closes #2627 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c2a305c44d |
feat(#2505): Phase 2 — kimi-code Agent Skills install layout (#2520)
* feat(#2454): PR 2 — kimi-code Agent Skills converter + install layout PR 1 registered the kimi-code EoS descriptor with empty artifactLayout (SKIP_INSTALL_CONTRACT excluded it from the end-to-end install test). PR 2 fills in the install surface: - src/runtime-artifact-conversion.cts: new convertClaudeCommandToKimiCodeSkill function. Today it delegates to convertClaudeCommandToKimiSkill (Python kimi-cli) because Kimi Code uses the same Agent Skills format + /skill: invocation per official docs. The distinct function name lets a future divergence land cleanly if Kimi Code's skill format evolves independently. - gsd-core/bin/lib/capability-validator.cjs: add to ALLOWED_SKILLS_CONVERTERS. - capabilities/kimi-code/capability.json: artifactLayout.global now declares the skills kind with converter='convertClaudeCommandToKimiCodeSkill' + home='.kimi-code' (auto-discovered at ~/.kimi-code/skills/ per Kimi Code docs: merge_all_available_skills = true default). - tests/installer-migration-install.integration.test.cjs: REMOVE the SKIP_INSTALL_CONTRACT exclusion — kimi-code now has a full install surface. - Regenerated capability-registry + capability-matrix + golden install parity + install tree fixtures for kimi-code. * fix(#2454): wire kimi-code converter into SKILLS_CONVERTER_REGISTRY + count bump - src/install-engine.cts: add convertClaudeCommandToKimiCodeSkill to SKILLS_CONVERTER_REGISTRY so the layout-driven skills install path can dispatch off the descriptor's converter string. - tests/capability-registry.test.cjs: bump VALID_CONVERTER_NAMES count 26 → 27 (added convertClaudeCommandToKimiCodeSkill). * fix(#2454): remove home override from kimi-code skills (inherit configDir) The home:'.kimi-code' override made the install plan resolve skills dest to ~/.kimi-code/skills instead of <configDir>/skills, causing the test's temp configDir to miss the install. Removing it lets skills inherit configDir like most runtimes. * fix(#2454): kimi-code install contract surface is flat-skills (no agents) Kimi Code has NO custom named subagents (per official docs: 3 built-in coder/explore/plan only). The kimi-skills-agents surface expects agents/ gsd.yaml + subagents/*.yaml which kimi-code does not produce. Changed to flat-skills which only checks for skills/gsd-* dirs. * docs(changeset): Phase 2 kimi-code install layout Added (#2509) * docs(changeset): backfill PR #2520 for Phase 2 (#2509) |
||
|
|
58028eaf56 |
fix(#2341): de-dup Cursor / menu by marking skills user-invocable:false (#2386)
Cursor installs both a skills and a commands surface and shows both in '/', duplicating every /gsd-*. Extend the #789 CodeBuddy de-dup to Cursor: convertClaudeCommandToCursorSkill (in both src and the live bin/install.js) now emits user-invocable:false, so the skill stays model-invocable while the commands surface is the single '/' entry point. Closes #2341. Admin-merged (self-review bypass) with full green CI. |
||
|
|
e0f969af6a |
refactor(#2246): centralize cross-platform path-separator handling (toPosixPath / toNativePath / posixNormalize) (#2247)
Replace every open-coded separator translation across the installer/hooks
source with named, tested seams in shell-command-projection.cts (the platform
seam), removing all hardcoded `/`+`\` from path handling:
- toPosixPath(p) — this machine's native path → POSIX (running-OS relative;
for local filesystem paths).
- toNativePath(p) — POSIX → native (collapses the win32 `/\//g,'\\'` ternary).
- posixNormalize(p)— unconditional `\`→`/`, OS-independent; for emitting paths
to a POSIX/bash TARGET (which may differ from the running
OS) and for parsing mixed-separator input.
core-utils.toPosixPath now delegates to the seam, so its 20+ existing consumers
resolve to one implementation; no duplicate helper.
- ~47 sites across runtime-hooks-surface, runtime-artifact-conversion,
runtime-artifact-install-plan, drift, init, worktree-safety,
installer-migrations, installer-migration-authoring, install-engine, surface,
verify, runtime-artifact-layout, schema-detect, check-command-router.
- Closes the latent POSIX-literal-backslash corruption class (the regex form
corrupts a POSIX path containing a literal backslash; split(path.sep) does not).
- New unit + fast-check property tests for all three helpers.
Closes #2246
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
d1e9491fef |
feat(#2099): drive GitHub Copilot through the EoS descriptor + multi-event hook bus (ADR-1239)
Fold Copilot's residual runtime-literal branches onto descriptor-driven hostBehaviors. Several issue premises were inaccurate (verified via research) and deliberately NOT followed: writesSharedSettings/"legacy exclusion list" (per-runtime descriptor data, copilot's false is correct); RUNTIME_CONTENT_DISPATCH.copilot + installSurface==='copilot-instructions' (already descriptor-driven); extendedHookEvents (closed Claude/Gemini enum — hooks extended in code instead); reapply ternary (already folded, kimi #2095). Real folds (all byte-parity — golden byte-identical for every runtime): - src/install-engine.cts + src/surface.cts: the two `.agent.md` filename cutovers (_copyStaged + _syncGsdDir) unified onto hostBehaviors.agentFileExtension via a new exported agentFileExtensionFor() accessor (kills the two-mechanism divergence). - src/runtime-artifact-conversion.cts: applyAgentPathRewrites' copilot skip → hostBehaviors.noPathRewrite:true (antigravity #2096 precedent). - bin/install.js uninstall: the two isCopilot cleanup branches → installSurface=== 'copilot-instructions' gate (symmetric with install-time). - bin/install.js: `!isCopilot` in the two skipSharedHooksInstall checks → hostBehaviors.skipSharedHooksInstall:true (copilot has no shared gsd-*.js hooks). - bin/install.js: three dead legacy inline-agent-loop isCopilot refs removed (copilot ∈ _DESCRIPTOR_AGENTS_RUNTIMES → unreachable; byte-parity proven by clean golden + real reachable-runtime install diffs). isCopilot dropped from 4 destructures. Zero live `runtime==='copilot'`/`isCopilot` branches remain in bin/install.js, install-engine.cts, surface.cts, or runtime-artifact-conversion.cts (AC2 guard scans all four). UPGRADE 1 (multi-event hook bus): buildCopilotHookConfig() now emits preToolUse/ postToolUse/userPromptSubmitted/sessionEnd advisory handlers alongside sessionStart (static inline bash/powershell — deterministic, golden-trackable). Only copilot.json's gsd-session.json hash changes. UPGRADE 2 (background dispatch): surfaced via the negotiated contract only — dispatch.background:true exceeds the declarative-cli baseline and survives negotiation with no downgrade warning. NO .agent.md frontmatter field (copilot has none). MCP companion out of scope (AC4 names only 2 upgrades). Tests: declarative-reference-copilot (adapter/axes/fail-closed + AC2 4-file source-grep guard) + copilot-upgrades (live 5-event hook wiring; dispatch.background negotiation). Matrix EoS note + how-to; changeset (Changed). capability-registry regenerated. Incidental flaky-test RE-ARCHITECTURE (no-defer, maintainer-directed): tests/opencode-review-reconstruction.property.test.cjs spawned ~600 synchronous execFileSync('jq') subprocesses (numRuns:200 × 3 fast-check properties, one jq per generated stream); a single jq freezing on a contended macos-22 CI runner hung the whole unit-test chunk to its 600s kill (this PR's CI). --test-force-exit can't interrupt a synchronous execFileSync, so the cure is to stop spawning per case, not just time-bound it. Re-architected to run the SHIPPED jq program over the whole fast-check corpus in ONE jq process: each generated stream is one compact-JSON array per line in a temp file, `jq -c <PROGRAM>` (no -s) applies PROGRAM to each array (`.` == the array, exactly what production's `jq -rs <file>` sees after slurping) and emits one result per line — empirically byte-identical to the per-stream form across embedded-newline/empty/quote/ unicode/null-drop cases, and file-input (like production) so there's no stdin pipe to deadlock on large I/O. ~600 spawns → 6; coverage unchanged (200-case corpus per property, deterministic seeds) plus explicit boundary/diagnostic example batches. Still property- tests the real shipped jq (no JS reimplementation). Per-call jq timeout retained as a belt-and-suspenders bound. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
17aa35fba7 |
feat(#2097): migrate Augment onto EoS declarative adapter + settings.json MCP companion (ADR-1239)
Fold augment's runtime-literal conversion branches onto descriptor-driven
hostBehaviors and delete dead code:
- Site A (_applyRuntimeRewrites case 'augment'): the 4 ~/.augment dot-dir
regexes now derive from getDirName('augment') via escapeRegExp (byte-
identical; getDirName('augment')==='.augment') — no runtime literal.
- Site B (applyRuntimeContentRewritesForCommandsInPlace): the
`if (runtime==='augment')` markdown-converter branch now reads
runtime.hostBehaviors.commandBodyConverter and dispatches through a local
COMMAND_BODY_CONVERTERS map (degrade-closed on unknown/absent name).
- Deleted dead `claudeToAugmentTools` map (zero refs; orphaned by ADR-1508
single-sourcing) and the unreachable `else if (isAugment)` agent-conversion
branch (augment ∈ _DESCRIPTOR_AGENTS_RUNTIMES → gated out upstream).
- Incidental orphan cleanup (no-defer): removed the equally-unreachable
`else if (isTrae)` agent-conversion arm left behind by trae's already-merged
migration #2094 (trae ∈ _DESCRIPTOR_AGENTS_RUNTIMES, same upstream gate).
copilot/windsurf/codebuddy arms are removed by their own pending migrations.
UPGRADE 3 (transport:mcp): register the GSD companion MCP server in Augment's
settings.json under mcpServers.gsd (Augment hosts MCP in settings.json, not a
standalone file). mergeGsdMcpServerIntoSettings mutates the in-memory settings
object finishInstall already writes (gated on hostBehaviors.mcpCompanion===
'settings-json'); non-destructive + idempotent; symmetric uninstall removal.
settings.json is golden-excluded, so no golden change. UPGRADE 1 (named/
background dispatch) + UPGRADE 2 (settings-json hook bus, Claude dialect)
were already live in production — this adds tests exercising both.
Golden: byte-identical for all 16 runtimes (folds preserve regex behavior;
MCP lives in golden-excluded settings.json) — verified by a real double-install
tree diff. Tests: declarative-reference-augment (adapter/axes/fail-closed/
undocumented-sub-axes + source-grep guard scoped to conversion-logic branches)
+ augment-upgrades (dispatch negotiation, hook-bus live install, MCP add/
idempotent/preserve/uninstall). Matrix + connect-gsd-mcp-server + a stale
install-on-your-runtime hook-ownership claim corrected; changeset (Changed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
5695522d5f |
feat(#2096): migrate Antigravity onto EoS declarative adapter + permission-writer + MCP companion (ADR-1239)
Fold all antigravity literal branches into descriptor-driven reads: getConfigDirFromHome (→ configHome.kind 'dot-home-nested'), projectLocalHookPrefix (→ hostBehaviors.hookPathStyle 'raw'), applyAgentPathRewrites (→ noPathRewrite), getProjectInstructionFile (→ projectInstructionFile 'GEMINI.md'); removed the dead inline convertClaudeAgentToAntigravityAgent branch + dead isAntigravity destructures (antigravity is already on the descriptor-agents path). subagentToolkit flipped undocumented→full (Context7: antigravity.google/docs/cli/features); namedDispatch/nested/maxDepth/backgroundDispatch stay undocumented. Byte-identical golden parity for all 16 runtimes. UPGRADE 1 (permission-writer): permissionWriter 'antigravity' + configureAntigravityPermissions merges a scoped permissions.allow block (GSD's own tree + hooks) into Antigravity's settings.json — non-destructive, idempotent, symmetric uninstall. Added to VALID_PERMISSION_WRITERS + the FinishPermissionWriter union. UPGRADE 2 (MCP companion): configureAntigravityMcpConfig writes mcp_config.json registering the gsd-core companion MCP server (Gemini-successor mcpServers schema, best-effort — raw schema unpublished). Both writers dispatch from finishInstall. settings.json is golden-excluded (HOOK_CONFIG_FILES); mcp_config.json (portable, no absolute paths) is golden-tracked → only antigravity.json changes. Tests: declarative-reference-antigravity extended (source-grep guard across 4 modules, fail-closed for the 4 undocumented sub-axes, validator acceptance) + antigravity-upgrades (permission-writer + mcp_config live-install, idempotency, user-preservation). Matrix + ADR-1016 + capability-manifest + CONTEXT.md + connect-gsd-mcp-server docs updated; changeset (Changed). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ab04916682 |
feat(#2095): migrate Kimi CLI onto EoS imperative adapter + native hook-bus + background dispatch (ADR-1239)
Fold all runtime==='kimi'/isKimi logic branches into descriptor-driven hostBehaviors (localInstallDeferred, verificationStyle, agentManifestStyle, reapplyCommand, doneBannerStyle) + add 'kimi' to _DESCRIPTOR_AGENTS_RUNTIMES. Kimi's skills/kimi-agents dispatch was already descriptor-driven (converter-by- name + kimi-agents kind). Zero isKimi/runtime==='kimi' branches remain. UPGRADE 1 (native hook bus): new hooksSurface 'kimi-hooks-toml' + a marker- delimited config.toml [[hooks]] emitter (buildKimiHooksTomlBlock/writeKimiHooksToml in runtime-hooks-surface.cts; resolveKimiHooksTomlDir in runtime-homes.cts). GSD's lifecycle hooks now wire into Kimi's native ~/.kimi/config.toml (Context7- confirmed path) at SessionStart/PreToolUse/Stop/PreCompact/SubagentStart/ SubagentStop — kimi becomes a hooks/ consumer (the 3 && !isKimi exclusion guards removed). config.toml holds absolute install paths so it's golden-excluded via an exact relative-path (.kimi/config.toml), not a basename (which would blind Codex's config.toml). New hooksSurface value added to the closed enum in capability-validator + runtime-config-adapter-registry. UPGRADE 2 (background dispatch): flip dispatch.backgroundDispatch true (Kimi's Agent tool takes run_in_background; root agent already gets the Agent tool), so negotiation no longer flattens dispatch. subagentToolkit stays 'undocumented' per AC (coder/explore/plan have distinct tool policies). MCP transport explicitly deferred (no installer-driven MCP for any runtime). Golden: only kimi.json changes (hooks/ scripts now installed); all 15 others + claude-local byte-identical (kilo/zcode keep their own exclusions). Tests: kimi-imperative-reference (adapter/axes/fail-closed/hostBehaviors + source-grep guard) + kimi-upgrades (config.toml [[hooks]] SessionStart + marker idempotency + backgroundDispatch negotiation). CONTEXT.md glossary + matrix + how-to updated; changeset (Added). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f744635b3d |
feat(#2094): migrate Trae onto EoS imperative adapter + SOLO stage-metadata upgrade (ADR-1239)
Fold trae logic branches into descriptor-driven reads: skipSharedHooksInstall gates (dropped && !isTrae), and the case 'trae' path-rewrite arm now computes the self-alias from the descriptor-driven dirName (.trae). Dead isTrae bindings removed from uninstall/writeManifest/finishInstall. trae's skills dispatch was already descriptor-driven (converter-by-name). RUNTIME_CONTENT_DISPATCH.trae is left as a runtime-keyed table registration (its regex/callback rewrites can't be a byte-identical descriptor map — matches cursor/windsurf/cline). trae stays in RUNTIME_FLAG_IDS: isTrae still gates the agents-converter selection (agents out of scope; removal gated on the cross-runtime agents-dispatch migration). Byte-identical golden parity for all 16 runtimes. UPGRADE: SOLO stage/trigger metadata — emitted Trae SKILL.md now carries stage: workflow (descriptor-gated via hostBehaviors.soloStageMetadata) so Trae's SOLO Agent can auto-invoke GSD skills at the corresponding stage. Field shape is best-effort/inferred (Trae publishes no formal schema). trae.json golden regenerated. Tests: trae-imperative-reference (adapter/axes/fail-closed shouldFlattenDispatch + no runtime==='trae' source-grep, isTrae exempted for agents) + trae-upgrades (stage: workflow on installed SKILL.md, descriptor-gated). Matrix note + changeset added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f014ec83bd |
feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads: finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks copy), and a skills converter-name registry (the artifactLayout.converter field is now load-bearing, not decorative). frontmatterDialect stays the documented dispatch key for frontmatter (no descriptor field for it). Dead isKilo destructure bindings removed. Byte-identical golden parity for all 16 runtimes (opencode, which shares kilo's combined-family path, verified clean). UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin + extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus). UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented' per AC so dispatch degrades to 'degraded' by design. Model-catalog single-source edit ripples the shared model-catalog.json hash into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale- bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/ codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error. Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/ hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades (plugin parity+load, model-override converter, agents dispatch surface, MCP doc). Matrix + how-to + config docs updated; changeset added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c6ce110efa |
feat(#2092): migrate Qwen Code onto EoS imperative adapter + native subagents + SubagentStart (ADR-1239)
Fold all runtime==='qwen'/isQwen logic branches (skill-priority frontmatter, branding/path rewrites, legacy commands/gsd cleanup, hyphen-namespace normalization, RUNTIME_CONTENT_DISPATCH, hooks-surface label) into descriptor-driven runtime.hostBehaviors on capabilities/qwen/capability.json, read via _hostBehaviors(). Shared claude/qwen/hermes legacy-migration branches in install-engine.cts folded to descriptor flags (claude+hermes descriptors updated; FALLBACK_HOST_BEHAVIORS.claude floored). Byte-identical golden parity for qwen/hermes/claude(global+local). UPGRADE 1: native .qwen/agents/*.md subagent projection — new agents artifact-layout kind + convertClaudeAgentToQwenAgent converter (name + description + tools YAML block list; color/model dropped). qwen routed onto the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES). UPGRADE 2: SubagentStart hook wired into extendedHookEvents + the descriptor-gated hook-writer loop (activates only for qwen). Tests: qwen-imperative-reference (adapter/axes/fail-closed/hostBehaviors + no runtime==='qwen' source-grep across 4 files) + qwen-upgrades (agents file validity + SubagentStart mirrors SubagentStop, descriptor-gated). Docs matrix + how-to updated; changeset added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
119702ff29 |
test(#2126): fix os.tmpdir() cross-file race + dedup folds surfaced by gsd-test (no-defer)
Phase 3's gsd-test surfaced 8 pre-existing test-isolation races (in #2090's test files now on next). Per CLAUDE.md's no-defer rule these are fixed inline in the current change. Root-caused via /qa-test-architect — all bad-test (the rewrite-engine production code is race-free): - install-runtime-artifacts.test.cjs: the "rmSync when readFileSync throws" test diffed the SHARED os.tmpdir() for gsd-cmd-rewrites-* dirs and force-deleted any new one with no ownership check. Under --test-concurrency it deleted a sibling test file's LIVE tempDir mid-copy (the #1575 "ENOENT .../graphify.md") and misattributed it as its own leak. Fixed: capture the exact tempDir THIS call creates (fs.mkdtempSync monkeypatch, restored in finally) and assert only on that — never sweep/delete the shared os.tmpdir(). Also deduped the enh-1511 block the #1969 consolidation folded in 3x byte-identically (#1970/#1974/#1975) down to 1 copy; 308 unique test titles unchanged (verified). - issue-1575-agent-descriptor-parity.test.cjs: a missing }); nested the M2 'cursor attribution' test inside the per-runtime loop so it ran 7x (widening the tempDir window). Fixed the brace -> runs once as a describe sibling. - config-get-default.test.cjs: local run()/runRaw() spawned node via execFileSync with a fixed 5s timeout and no retry -> ETIMEDOUT under Docker load. Redesigned to call cmdConfigGet in-process (fs.writeSync fd-capture + process.exit sentinel, both restored in finally) — no subprocess, no wall clock. - runtime-artifact-conversion.cts: fixed the stale "No production caller today" JSDoc on rewriteStagedCommandBodies (real callers: applySurface, createRuntimeArtifactInstallPlan) — the false doc invited the bad test. Refs #2126, #2090 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
396f44bd0b |
feat(architecture): [EoS/opencode] Migrate OpenCode onto the Embeddable Orchestration System (ADR-1239, #2087)
Route OpenCode (and its Kilo sibling) through the public Host-Integration Interface and land two Context7-verified capability upgrades. Byte-identical install output for all 16 runtimes (golden parity asserted). Through the interface (AC2): - OpenCode/Kilo's bespoke commands+skills+plugin install (the inline `else if (isOpencode || isKilo)` block) moves into the engine (installOpencodeFamilyCommands/Artifacts in src/install-engine.cts), dispatched by installRuntimeArtifacts when the descriptor declares hostBehaviors.combinedFamilyInstall. opencode/kilo now flow CLI -> _runtimeAdapter -> installRuntimeArtifacts like the skills runtimes. _isSkillsRuntime no longer excludes them; the bespoke block + dead copyFlattenedCommands are removed. - Every hardcoded `runtime === 'opencode'`/`isOpencode` branch is folded into descriptor-driven runtime.hostBehaviors. ZERO `runtime === 'opencode'`/`'kilo'` string-equality remain in bin/install.js / install-engine.cts / runtime-artifact-conversion.cts. Upgrades (AC4): - Background dispatch: OpenCode shipped experimental background subagents in v1.15 and made them default-on in v1.17 -> dispatch.background/backgroundDispatch flip to true; shouldFlattenDispatch(opencode) now returns false (behavioral change; type: Changed). - Expanded event surface: the OpenCode plugin subscribes permission.asked/replied + session.error. Tests: opencode-imperative-reference (adapter/profile, shouldFlattenDispatch pin, fail-closed negotiate, hostBehaviors, AC2 source-guard) + extended plugin surface test. Docs: capability matrix v1.15/v1.17 citations. Changeset (Changed). gitignore .memdb//.memtrace/. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ce62f2b68d |
refactor(#1763): ADR-1235 agent migration — cut over the trivial-converter group to the descriptor path (#1764)
ADR-1235 step 1: route the trivial-converter runtime group (cursor, windsurf, augment, trae, codebuddy) off the inline install() agent loop onto the descriptor-driven installRuntimeArtifacts path. Establishes the converter-context foundation (pre-converter cross-cutting + no agent-stamp). Agent install output is byte-identical for all 16 runtimes (golden-parity, global + verified local). cline deliberately excluded (local rules-only). Closes #1763. |
||
|
|
2990305f78 |
refactor(#1727): derive NON_CLAUDE_RUNTIMES from the capability registry (ADR-1239 Phase B) (#1728)
* refactor(#1727): derive NON_CLAUDE_RUNTIMES from the capability registry ADR-1239 Phase B (parent #1679). NON_CLAUDE_RUNTIMES was a hand-maintained 15-element literal whose own doc-comment said "keep in sync with bin/install.js and getDirName()" — a parallel source of truth that can drift from the capability registry. Derive it instead: Object.keys(capabilityRegistry.runtimes).filter(id => id !== 'claude').sort() The exported value is byte-identical to the old literal (the registry's runtimes key set minus claude is exactly the 15 entries), so there is no observable behavior change; the list can no longer drift from the registry. capability-registry.cjs is a committed, dependency-free data module (no cycle). Drift-guard test: golden-oracle deepEqual (non-circular) + a role-based cross-check from the registry metadata + every member must have a non-'.claude' getDirName branch (a registry runtime missing a getDirName branch now fails CI). Closes #1727 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1727): backfill changeset PR number (#1728) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1727): put docs-exempt marker on its own line so parse.cjs extracts it DOCS_EXEMPT_RE is line-anchored (^...$ + m flag); the marker only counts on its own line. It was appended to the end of the body text, so it was never extracted and docs-lint failed in CI (fail_docs_missing). Verified via direct parse.cjs extraction (docsExempt now non-empty, marker stripped from body). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4ed208e74b |
fix(#1615): validate commandName to prevent workflow prompt injection
Codex peer review of PR #1622 surfaced that convertClaudeCommandToWindsurfWorkflow interpolated commandName unsanitized into a markdown body that Windsurf loads as an LLM-readable workflow. A plugin author who controls a commands/gsd/*.md filename could inject newlines, markdown structure, or path components (..) to manipulate the workflow body. Validate commandName at function entry against /^(?:gsd-)?[a-z0-9](?:[a-z0-9-]*[a-z0-9])?$/ — rejects slashes, backslashes, spaces, dots, control chars, trailing dash. Pattern requires alphanumeric ending so gsd- alone (which would slice to empty stem) is also rejected. Throws with a JSON.stringify-escaped preview (no literal newlines in the error message). Applied to both bin/install.js (where tests import from) and src/runtime-artifact-conversion.cts (production source). 18 positive + 22 negative test cases lock in the validation. |
||
|
|
527142ad2e |
fix(#1615): normalize Windows backslash paths in workflow content
computePathPrefix returned a Windows-style path (with backslashes from path.join) into markdown @-references. Workflow file content on Windows ended up with mixed separators, breaking substring checks in install/install-runtime-artifacts tests on windows-latest CI only. Normalize resolvedTarget and homeDir to forward slashes inside computePathPrefix. The prefix is always substituted into markdown body text, which uses POSIX paths universally. Idempotent on POSIX. Also normalizes the two test assertions to forward-slash form so they pass on Windows. Adds a regression test for backslash-style input. Documents DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT + RULESET.CONTENT-PATH-NORMALIZATION in CONTEXT.md so this anti-pattern stops recurring. |
||
|
|
fc2a7c0555 | fix(#1615): install Windsurf slash workflows | ||
|
|
793fab0fd6 | Merge branch 'next' into fix/1394-gemini-skill-tool-exclusion |