1110c3b4eef96a533ea0094850347b22d689e536
23 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
54085516c1 |
fix(#4211): materialize Kimi's agent tree recursively during surface apply (#4371)
* fix(#4211): materialize Kimi's agent tree recursively during surface apply kimiAgentsKind stages `gsd.yaml` + `gsd.md` + `subagents/gsd-*.{yaml,md}`, and install copies that tree recursively (_copyStaged). Surface apply fell through to _syncGsdDir's flat command/agent branch, which reads only top-level `*.md`: the YAML half and the whole subagents/ subtree were ignored, and `gsd.md` was written as `gsdgsd.md` because the flat branch re-applies kind.prefix to a name that already carries it. A surface change could therefore corrupt Kimi's installed artifacts while still reporting success. Three divergences from the install path, all in src/surface.cts: - _syncGsdDir gains a kimi-agents branch: recursive copy, then a prune scoped to exactly what install's _removeGsdEntries owns for this kind (the two root files, and gsd-*.{yaml,md} under subagents/). Everything else is user-owned and preserved. - applySurface stages kimi-agents WITH agentCtx and the `skills: '*'` rule for an unmodified full profile, as it already does for the agents kind and as createRuntimeArtifactInstallPlan does for every kind — without it Kimi's generated subagents lost their path-prefix rewrites and attribution trailer, and an unmodified full profile staged only the skill-referenced subset. - applySurface runs rewriteStagedSkillBodies for kimi-agents, which the install plan routes through it alongside skills. * chore: add changeset for #4211 --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
4c60879b5d |
fix(#4132): verify durable runtime surface sources (#4182)
* fix(#4132): verify durable runtime surface sources * chore(#4132): record PR number in changeset * test(#4132): cover rejected commands source alias * fix(#4132): reject aliased package fallback * test(#4132): cover rejected agents source alias * test(#4132): cover partially aliased marker provider * fix(#4132): reject partially aliased source providers * test(#4132): cover routed source identity probes * fix(#4132): route installed source identity probes * refactor(#4132): tighten installer source metadata * test(#4132): cover corpus trust boundary attacks * fix(#4132): close installed corpus trust gaps * refactor(#4132): keep installer authority private * fix(#4132): preserve private installer fallback * test(#4132): preserve fixture source authority * fix(#4132): reject overlapping source fallback * fix(#4132): avoid redundant installed corpus reads * refactor(#4132): simplify provider resolution * test(#4132): sync install tree fixtures after rebase --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
107eb8c1d9 |
feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on
|
||
|
|
a44d513566 |
fix(#3712): confine in-process installs to a sandboxed HOME (#3725)
* fix(#3712): confine in-process installs to a sandboxed HOME
A runtime kind may declare a global `home` override resolved from os.homedir()
rather than from the caller's configDir — codex's skills kind (`home: ".agents"`,
ADR-1239 / #2088) is the only live case. Sandboxing configDir/targetDir does not
contain it, and assertDestWithinConfigHome cannot see the class: that gate
confines a destSubpath to whatever root it is handed, and here the root IS the
escaped home. So an in-process caller that forgot to sandbox HOME wrote to, and
pruned gsd-* entries from, the developer's REAL ~/.agents/skills.
tests/agent-descriptor-parity.install.test.cjs's K1 loop did exactly that: it
iterates every agents-kind runtime (codex included) with a sandboxed targetDir
and an un-sandboxed HOME. Reproduced against a canary home on next @
|
||
|
|
3ab0007164 |
enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19) preserveUserArtifacts held user files only in an in-memory Map across the wipe, so any process death between preserve and restore lost them outright. Seven call sites, not the four the issue records. Three of them never called the helper at all - they open-coded the same read/wipe/write - so searching for callers under-counted by construction; the extra sites were found by sweeping for the pattern instead. The worst is the mainline install path, where the crash window spans the entire gsd-core tree copy rather than a single rmSync. Adds src/user-artifact-staging.cts: durable on-disk staging with a record written after the copies land as the commit point, plus recovery of orphaned batches on the next run - without recovery the staged bytes survive but the user's file is still gone, which would pass its own test while delivering nothing. Routes copyPreservingSymlink through installFs() so staging cannot bypass the install fs seam, and reunites its symlink-safety docblock with the function it documents. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-3574 with four claims disproved by implementation Implementing Phase 6 disproved four statements the ADR rests on. The central decision - no single materializer - is unaffected and stands. Corrected: decision 3 was already satisfied, so nothing was extracted; the agents-bypass runtime set omitted claude, kilo and opencode, and closing it needed three new pieces of descriptor contract rather than proceeding on its own terms; three of the four blockers the layout comment names were already stale; and F19 is seven call sites, not four. Records the generalizable lesson: the defect is the pattern of holding user data in memory across a wipe, not the helper, so searching for callers of the helper under-counts by construction. Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and notes that copyPreservingSymlink needed routing through the install fs seam before it could be reused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close dangling-symlink blind spot and harden staging recovery An adversarial review found the F19 staging work shipped red and unsafe. Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween missed dangling symlinks in both its root check and its per-segment walk, because it probed with existsSync, which is false for a link whose target does not exist. Fixing only the new module would have reused a guard that was itself blind. This guard protects the whole install tree. Recovery no longer throws: it degrades per entry and per file, so one bad batch cannot block the others. Previously an unrecoverable entry propagated out of the first statement of install and uninstall, before the cleanup that would have removed it - wedging the installer permanently. Partial fs adapters now throw on any omitted method instead of silently reaching the real filesystem, closing the trap that let a test poison list pass while real IO happened. Staged names must be flat, recovery refuses a dangling destination symlink, and a batch whose recovery genuinely failed is no longer swept - it was discarding the only durable copy of the file it had just failed to restore. Replaces three tests that could not fail, including the one labelled negative proof. Known limitation, documented not closed: concurrent installs sharing a staging key can still lose a batch. A real fix needs a cross-process lock. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enh(#2875): make the descriptor authoritative for the agents kind Deletes the inline agent-staging loop in bin/install.js and the _DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from its capability descriptor instead of an inline hostBehaviors dispatch. Closing it needed three pieces of contract the descriptor pipeline never had, all reducible to one missing input - per-agent resolution context: a frontmatter-extensions step for claude's effort and disallowedTools, per-agent model-override resolution for kilo and opencode, and a named branding converter for hermes, whose rewrite data was already declared. Seven runtimes were on the loop, not the six the design recorded - kimi-code was found by a golden fixture, not by analysis. claude-local and kimi-code both silently lost their agents mid-change; the fixtures caught both and the cause was fixed rather than the fixtures regenerated. A parity harness gates the migration: both pipelines over identical inputs, byte-identical output including filenames, per runtime. It is demonstrated red before being trusted. Surface and install paths converge for all seven, which also fixes surface previously writing no agents for these runtimes. Codex's config.toml strip stays put - it mutates host config, which no descriptor kind models. Also routes install-model-override-resolver and install-effort-resolver through the install fs seam. Both leaked real filesystem IO from the install call tree; the stricter adapter is what exposed them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): record the agents-descriptor migration and correct the ADR count The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host integration guide told readers to join a set that is gone. Replaces that with what is now true - declare an agents entry and it installs, on the surface path as well as install - and points anyone needing a per-agent transform at the three extension points rather than at a new inline branch. Corrects the ADR amendment: seven runtimes were on the inline loop, not six. kimi-code was found by a golden fixture going red, not by reading. That is the third short count this phase, all from enumerating by symbol or set membership when the thing that matters is a behavior. Adds the Changed changeset for the surface-path convergence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-2866 - claude global always wrote agents on disk The claude row's global=[skills] described what capability.json declared, not what the installer wrote. bin/install.js's inline agent-staging loop was never scope-gated and never consulted the descriptor, so a claude --global install has always written agents/gsd-*.md. Phase 6 closes the gap by deleting that loop and declaring agents on claude's descriptor at global scope. On-disk bytes are unchanged - the golden fixtures did not move, which is the evidence that the descriptor, not the installer, was incomplete. #2218 is unaffected: agents are not trigger-bearing, so the wider row does not introduce a new shadowing case. Records the warning that an incomplete descriptor is invisible while a second code path silently does its work, and only surfaces when the two are forced into agreement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close review findings across staging, agents and the parity harness Two independent reviews of this branch found defects the local gates missed. Security: a dangling symlink at a migration destination allowed writing outside configDir - the same class this change claimed to close, missed at the terminal write of the flow being added. The staging-root resolver threw as the first statement of install and uninstall, so a hostile symlink bricked both, and symlinked-configDir users lost uninstall as well as install; it now degrades instead of aborting. Recovery gained a source-side symlink check and now refuses a relative destDir, which resolved against cwd. Converter dispatch gained a runtime allowlist - lint-time validation stopped mattering once this branch promoted that dispatch from the surface path to real installs. Correctness: claude --local --minimal exited 1 because the minimal profile legitimately yields zero agents and the new path treated that as a failure. cline --local silently lost its agents - its descriptor declared none while the deleted loop wrote them unconditionally. The agents prune was widened to any gsd-* entry and destroyed user files it never owned. The parity harness, on which the migration's safety argument rested, drove a synthetic registry and never byte-compared the shipped descriptors; two of its trap rows could not fail. It now drives the real registry across 13 runtime-scope rows including kimi-code and cline-local, and its red-proof is demonstrated by corrupting a live capability.json. Three goldens that had encoded the cline regression as expected behavior were corrected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close findings from both mandated review engines /security-review found the staging source-side walk honouring GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write destination. A symlinked files/ component dereferenced because copyPreservingSymlink lstats the leaf only, so an intermediate link is followed. The source walk no longer honours the opt-in; the destination check still does. /code-review spec axis found this branch had reintroduced its own bug: migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after the legacy dir was wiped and before the staged batch was restored, so a planted symlink bricked uninstall permanently and orphaned the batch. Refusal kept, abort removed. kimi-code local silently lost its agents, the same class as the cline bug, and the parity harness recorded that exclusion as intentional - the third test in this branch to pin a regression as correct. --minimal now creates an empty agents/ dir that never existed. Behaviour restored rather than softening the changeset, so its byte-identical claim stays true. Standards axis: try/finally removed from twelve test bodies, fast-check properties added for parseOwnerPid, boundary coverage at the grace window and the ancestor-probe depth, a parity assertion for the staging-root helper duplicated across two files, and the 8-deep config walk deduplicated. Records 60-review.json with every finding and disposition from five passes, including the smells left unfixed and why. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): prune stale agents unconditionally in minimal mode The previous round stopped an empty agents/ directory being created when the resolved profile yields no agents. That was implemented by skipping the agents kind entirely, which also skipped its stale-agent prune - so a full to minimal downgrade left stale gsd-* agents behind. The deleted inline loop pruned unconditionally and only skipped writing. Those are three separate conditions, not one: prune always, write only when there is something to write, create the directory only when writing. Both call sites now run _removeGsdEntries before the empty-staged early exit. The symlink-escape guard moved with it, since the prune also touches dest. Codex .toml agents and the config.toml stanzas are cleaned again, and user-owned agents are still preserved. The agents/ directory is left in place after a prune empties it, matching every sibling kind - none of them remove the destination directory itself. Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh install, so fixture generation is unaffected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): document interrupted-install recovery for user-owned files The durable-staging fix is invisible to the user it protects. Someone whose install died mid-flight has no way to know USER-PROFILE.md was staged before the delete, that the next run restores it, or that recovery happens at the start of that run rather than in the background. Written as the task the user has - finish the interrupted command - rather than as a description of the mechanism, and states what it will not do: overwrite a file already present, or touch staging belonging to another install still running. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2875): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2875): assert the J8 model override without building a regex CodeQL flagged incomplete string escaping: the assertion interpolated the override value into a RegExp while escaping only forward slashes, which is meaningless in a constructor, leaving real metacharacters unescaped. The failure direction was the dangerous one - a metacharacter would have made the match more permissive, so the row would pass when it should fail. That matters here because J8 exists precisely because an earlier revision was a tautology; the rewrite reintroduced a different way for the same assertion to stop discriminating. Replaced with a line-wise exact match, so no regex is constructed at all. Swept the other test files this branch adds; no sibling instances. lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full metachar-escape copy, so a single slash replace slipped under it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cf6de5e1c0 |
feat(#2871): resolve triggers and host precedence, not just placement (#3291)
* test(#2871): failing-first suite for trigger-surface resolution 23 tests over the 50-test-matrix rows. RED by construction: resolveTriggerSurface and DEFAULT_TRIGGER_PRECEDENCE do not exist yet, and the validator silently ignores triggerPrecedence today. Written in the per-runtime describe idiom the other four runtime-artifact-layout suites use, not a table. The rows that carry the weight: windsurf must NOT report a shadow it does not have, since its global scope emits only agents and agents are not trigger-bearing; agents and kimi-agents must be absent from the output for every runtime; and reordering a runtime's triggerPrecedence must flip the winner, which is the only assertion that proves the axis is read rather than decorative. Stems are injected, never scanned, so the surface is assertable with no filesystem. * feat(#2871): resolve triggers and host precedence, not just placement resolveTriggerSurface(runtime, scopes) returns every /gsd-<name> trigger a runtime emits, with the scope and kind that produced it, whether the host registers it directly or only through a router, and which artifact shadows it. resolveRuntimeArtifactLayout is untouched -- its 7 callers need placement only and the issue requires them unchanged. AGENTS ARE NOT TRIGGER-BEARING, and ADR-2866 said they were. The host-integration matrix models command and dispatch as separate interface points: an agent is invoked through the Agent tool's subagent_type, not by typing a slash trigger, and _copyStaged never applies the kind prefix to an agents entry. So agents and kimi-agents are excluded from the surface entirely, and this commit amends ADR-2866 with a dated correction. #2218's conclusion is unchanged -- the collision is strictly commands-vs-skills, and claude's local /gsd-* trigger surface is still fully shadowed -- but the ADR implied the local agents surface was lost too, and it is not. That correction is what makes windsurf come out right. Its global scope emits only agents, so it has no global trigger and its local commands are unshadowed. Model agents as trigger-bearing and windsurf falsely reports a full shadow. The triggerPrecedence axis lands on all 19 descriptors as an ordered kind list, one value with one owner, rather than a numeric rank spread across N kind entries with nothing keeping them consistent. Validation uses a required-with-default shape that has no precedent in this validator -- every existing axis is hard-required -- so a third-party capability.json omitting the field still validates, which is what ADR-894's additive-only contract promises. Winner resolution reads Phase 1's scope rank first, then the kind ordering. A test reorders the axis and asserts the winner flips, since an axis that is added, validated and never consulted would pass every other assertion. shadowedBy ships unread. Phase 4 (#2873) is its first consumer, per this issue's out-of-scope note. Verified via the remote runner. * fix(#2871): single-source namespacedByDir and close two test gaps Four findings from the isolated adversarial review. The namespacedByDir rule had reached three copies -- install-engine, surface, and the new trigger resolver -- one of which carried a hand-written keep-in-sync comment and no assertion. That is this repo's generative-fix-divergence class. Extracted to one exported predicate all three now call. Verified by diverging one copy deliberately: the existing #816 parity test failed, and passes again on revert. The omission test was vacuous. Row 16 asserted that a descriptor without triggerPrecedence still validates, but built its fixture from claude's shipped descriptor -- which this PR had just added the axis to. It now clones and deletes the key, following the shippedDescriptorWithout pattern, and asserts both that validation passes and that the resolver still picks the right winner from the default. The second half is what makes it prove anything. resolveTriggerSurface silently dropped an unrecognized scope while every sibling in this epic throws. Two phases of one epic should not disagree about whether an invalid scope is an error, so it now rejects through the same shared validator; an empty scope list still returns empty rather than throwing. The ADR amendment had been spliced into the middle of the References list, orphaning its last bullet. Moved to the top, after the header block, which is where ADR-3660 and ADR-1016 both put dated amendments. No lint checks markdown structure, so this was green while malformed. * fix(#2871): single-source the command filename composition too The earlier fix shared the namespacedByDir boolean but left the filename composition around it written twice -- once in _copyStaged as what actually gets written, once in resolveTriggerSurface as what gets predicted. The predictor could go stale silently. One exported helper now composes it for both. The entry.name asymmetry that looked like it would block extraction does not: entry.name is filtered to end in .md and stem is entry.name minus those three characters, so the two branches are the same string by construction. Divergence proven to fail: injecting a marker into the helper broke the trigger-surface suite; reverting restored 25/25. The four sibling layout suites hold at 227 unchanged. * docs(#2871): correct the ADR timing notes that this phase makes stale The Amended by back-links on ADR-3660 and ADR-1016 were written in Phase 0, when the widenings they describe had not shipped. Each carried a forward-looking clause -- "the module changes at Phase 2, not before, until then this module resolves placement only" -- which becomes false the moment this PR merges. ADR-2866's own Amends header and its reciprocal-notes section carried the same tense. All four now describe what shipped. This is a tense and status correction on Accepted ADRs, not a change to any decision. Worth stating because it is the failure mode this epic keeps meeting: gen-adr-index.cjs tracks only Supersedes and Subsumes, so nothing in CI would have caught either the missing back-link in Phase 0 or these stale clauses now. They stay correct only because someone checks. * chore(#2871): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
c2f24265f2 |
feat(#2870): resolve install scope as a value (#3278)
* test(#2870): failing-first suite for the Install Scope Module 19 tests over the 50-test-matrix rows 1-19. RED by construction: the module under test does not exist yet, so the suite fails at require with MODULE_NOT_FOUND until src/install-scope.cts lands. Every row asserts a returned value with injected env/home/existsSync -- no filesystem, per the issue's acceptance criterion that tests assert the resolved value directly. Row 7 asserts the RELATION rank(global) > rank(local) rather than a literal, so Phase 2 (#2871) can re-base the numbers without a fixture edit. Row 4 iterates the real runtime registry rather than a hardcoded list, excluding vscode, which declares configHome.kind none and is never CLI-installed. * feat(#2870): add the Install Scope Module Scope becomes one resolved value instead of a bare string re-derived at every layer. resolveScope({id, runtime, ...}) returns {id, configHome, settingsFile, consentRequired, hostPrecedenceRank}. It COMPOSES resolveConfigHomeFromDescriptor rather than extending it. That function has 60 dependents across 13 files and 2 process flows -- a CRITICAL blast radius -- so adding a scope parameter to it, which the issue's framing invites, would ripple through all of them. Composing costs nothing and leaves every existing caller byte-identical. The module owns the InstallScope type name, which previously lived privately in runtime-artifact-install-plan.cts; that module now imports it. A fifth spelling of the same concept would have defeated the phase. settingsFile is null for the 18 runtimes that declare no settingsFileByScope -- absence is a value, not an error, and inventing a Claude-shaped default would leak that host's shape onto every other one. hostPrecedenceRank ships unread: Phase 2 (#2871) is its first consumer. It is carried as data only, per this issue's out-of-scope note that precedence semantics belong to that phase. Vocabulary: the install axis standardizes on local. ConsentRecord.scope keeps project deliberately -- that literal is serialized into consent records in the user's home, and renaming it would silently deactivate every project-scoped capability on the machine. CONTEXT.md records the boundary mapping instead. Every environmental input is injectable (env, home, existsSync, cwd), so the resolved value is assertable with no filesystem at all. Registration ripple: .gitignore, eslint.config.mjs, CONTEXT.md glossary, docs/INVENTORY.md, and the inventory manifest (regenerated after build:lib, never before). Verified via the remote runner. * refactor(#2870): route scope re-derivations through the module bin/install.js resolves scope once per function instead of inline at each of its 12 sites, and the settingsFileByScope consumer reads it through resolveScope(). Seven downstream boolean re-derivations now call the module's isGlobalScope() instead of comparing the literal independently: runtime-artifact-install-plan, both runtime-artifact-layout kind builders, dispatchKindEntry, surface, and two install-engine sites. The fifth through seventh were not named in the issue -- they are the same re-derivation class, and leaving them would have made the acceptance criterion false. _computePathPrefix keeps its isGlobal boolean API, so the projection is centralized rather than eliminated. resolveScope and isGlobalScope share one validator, so the two surfaces cannot drift. TWO SITES DELIBERATELY NOT ROUTED: runtime-artifact-conversion's rewriteStagedSkillBodies and rewriteStagedCommandBodies. ADR-1508 fixes the direction as installer/layout -> conversion, never upward, and install-scope composes runtime-homes, so importing it into the conversion module would invert that direction. Left as-is on purpose. Behavior-preserving throughout. Each step was proven by capturing full layout and plan output -- including every kind's home field and the hashed contents of emitted files -- before and after, across both scopes for claude, codex, opencode, hermes, kimi and kilo. Byte-identical. surface.cts keeps a scope ?? 'global' default before the call because Layout.scope is optional there; isGlobalScope throws where the old inline compare returned false, and that difference would have been a placement regression. Verified via the remote runner. * fix(#2870): cover the no-config-home throw and document the strictness Two findings from the isolated adversarial review. The vscode case was implemented but untested. resolveScope throws for a runtime whose descriptor declares configHome.kind 'none', which is the design's own behavior-table row 13, but the registry sweep excluded vscode rather than asserting the throw -- so the behavior shipped with no test. The exclusion is now legitimate because the case has its own test naming the runtime in the assertion. isGlobalScope throws where the inline compare it replaced returned false. No reachable caller can deliver an out-of-union value today, but the types are not enforced at runtime, so a future caller passing an optional Layout.scope would crash rather than silently misroute. That is the better failure -- misrouting writes artifacts to the wrong place -- but it was undocumented, so the reason is now on the function. Adds the changeset the acceptance criteria require. * refactor(#2870): route the last two sites; correct the ADR-1508 claim The previous commit declined to route runtime-artifact-conversion's rewriteStagedSkillBodies and rewriteStagedCommandBodies, claiming ADR-1508's dependency direction forbade the import. That reasoning was wrong, and this commit corrects it. Two independent reviewers checked the actual import graph: runtime-artifact-conversion already imports capability-registry, command-roster, runtime-name-policy and shell-command-projection -- it depends on leaf-tier siblings today. install-scope imports only runtime-homes plus node builtins, and runtime-homes imports only node builtins, so there is no cycle at any depth. ADR-1508 governs the installer/layout to conversion boundary, not a leaf-to-leaf sibling import of the same shape conversion already makes. With those two routed, every isGlobal re-derivation in the tree now goes through one owner and acceptance criterion 1 is fully met rather than partially. Nine sites, not the four the issue enumerated. Also from the review: Tests were falling through to the real process.cwd() at five local-scope call sites, which contradicts the acceptance criterion that the resolved value be assertable with no filesystem. Every one now injects a cwd. One of the five was a site the review had not spotted. bin/install.js carried two near-identical copies of the guarded resolveScope block, one in install() and one in uninstall() -- duplicated scope logic in the phase whose purpose is removing it. Extracted to one helper, and the new sites use the file's existing ternary idiom rather than the if/else that replaced it. Equivalence re-proven across both scopes for claude, codex, opencode, kilo and hermes, now including the staged skill and command body rewrites hashed per file, since those decide the literal spec-root path baked into every emitted artifact. Byte-identical. Verified via the remote runner. * fix(#2870): assert configHome portably instead of with a native separator The windows-latest node24 shard failed on two install-scope assertions. The module was right and the tests were wrong: they built their expected value with path.join, which emits \fake\home\.claude on Windows, while resolveScope normalizes separators unconditionally to /fake/home/.claude. That unconditional normalization is deliberate -- backslash paths arrive on Linux too, so normalizing via path.sep is the documented defect this repo guards against. Weakening it to make the assertion pass would have inverted the fix. Every path.join-built expectation in the suite now goes through toPosixPath from tests/helpers.cjs, which is the pattern the no-path-literal-in-assert rule's own valid-case list sanctions. It splits on the running platform's path.sep and rejoins with forward slashes, so it reverses whatever path.join produced on that same platform and the expectation is invariant everywhere. Two more call sites had the same latent problem and passed on Linux and macOS by luck; they are fixed too. This is the class of defect the remote runner structurally cannot catch -- its matrix is Linux-only, so a green pass there is not evidence of portability, and CI's Windows lane is the only place it surfaces. Verified via the remote runner. --------- Co-authored-by: sim <sim@local> |
||
|
|
8c1962200d |
fix(#2911): resolve surface re-stage destinations the way the installer does (#3049)
* fix(#2911): resolve surface re-stage destinations the way the installer does Two writers computed the same destination differently. The installer honors a skills-kind home override; the surface re-stage ignored it and always resolved against configDir. For a global Codex install that override points at $HOME/.agents, so every re-stage built a second GSD-managed skill tree under $CODEX_HOME alongside the correct one, with nothing indicating which was live. Honors the override as a fallback, never a replacement -- runtimes without one still resolve against configDir, which is most of them. The real deliverable is the parity test, not the one-expression fix: it walks every runtime in the registry across both scopes, computes the installer and surface destinations, and fails naming the runtime if they ever disagree. Today only Codex global carries an override, so it discriminates on exactly one runtime -- stating that plainly rather than implying broader coverage -- but it is derived from the registry, so a newly-added runtime is covered without anyone remembering to add it. Two further defects fixed rather than deferred: - The legacy dev-preferences migration carried the identical defect, which the issue flagged as a latent instance of the same shape. - Fixing it exposed a symlink-escape guard confined against the wrong root: it checked the span between configDir and the skill dir, but a home override moves the skill dir outside configDir entirely, so the span was meaningless and threw a false-positive escape. Now confined against the install root the destination actually resolves under. The guard is unchanged in strength and still honors its opt-in; only the root it measures from is corrected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2911): honor the home override in the fourth destination writer too Adversarial review found a writer the fix had missed: the opencode-family skills installer resolved its destination and its symlink guard against targetDir, never consulting the skills-kind home override, while its three siblings all already honored it. Pre-existing and currently dormant -- it is reachable only for the combined-family runtimes, and none of them declares an override today, so no user is affected right now. Fixed anyway rather than left as a latent instance of the same shape, which is exactly what this issue asked for in the case of the legacy migration. Mirrors the shape used for the other three: a single installRoot local that both the destination and the guard derive from, so the two cannot drift apart. The guard's message now names the root it actually confined against. Coverage extended to this writer and proven non-theatre: reverting the change in a scratch build makes it fail for both combined-family runtimes. Enumerated every remaining site that computes a destination from destSubpath or calls the confinement helper -- install and uninstall paths, the surface module, the read-side skills-root reporter. All honor the override or structurally cannot express one. No fifth defect. The one adjacent shape, the flat command directory, reads a different descriptor field that no kind declares an override for in the current schema; noted rather than papered over with a fallback for a field that cannot exist. Verified no behavior change for the affected runtimes today: normalized file-tree hashes before and after are identical. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2911): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
88f6d9bd1b |
fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu * fix: preserve installer executable mode * chore: add changeset for PR #2812 * test(#2644): acknowledge Cursor emission changes * test(#2644): drop spent emitted drift acknowledgments * fix(#2644): remove retired Cursor command converter --------- Co-authored-by: clezcoding <clezcoding@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
23a65c4a3d |
fix(#2322): materialize installed third-party capability skills (#2340)
* test(#2322): fail-first tests for third-party capability skill materialization Red phase: tests (1) and (6) fail — resolveSurface reports the third-party stem surfaced (#2045) but no SKILL.md is ever written to disk. The other four are controls that must keep holding: first-party-wins collision, profile-tier filter, nested-router layout unperturbed, and absent/malformed capability must not throw. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): materialize installed third-party capability skills A capability could report installed:true, surfaced:true, active:true and still never exist as an invocable command. #2045 fixed the registry layer — resolveSurface unions registry.capabilityClusters into the resolved skill set — but the materialization layer never got the matching fix. stageSkillsForRuntimeAsSkills only ever read gsd-core's own bundled commands/gsd/*.md and silently skipped any stem it couldn't find there, so a third-party skill living at <GSD_HOME>/.gsd/capabilities/<id>/skills/<stem>/ was never copied. Registry said surfaced; disk had nothing. Installed capability skills are now staged alongside the first-party ones, copied verbatim (they are authored complete for their target runtime and need no converter). First-party stems always win a collision, the profile filter still applies, and an absent or malformed capability degrades rather than throwing. Security: capability.json's skills[] entries are validated only as non-empty non-reserved strings (capability-validator.cjs:503-514) — no path shape is enforced upstream — so stems are sanitized (rejecting separators, '..', absolute paths, NUL) with an independent isPathConfined check on both the read and write paths. A '../../evil' stem writes nothing outside the capability's own dir. Also fixes a defect this surfaced in pruneSkillDirs: a materialized capability skill dir has no first-party manifest entry, so every apply logged "preserving (user-owned or unknown)" for a live GSD-managed dir. The retained check now precedes the manifest gate; no deletion outcome changes, and genuinely unknown gsd-* dirs still warn and are preserved. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2322): address security review — bind skills to declaring capability, fix full profile An independent security review BLOCKED the first pass. Both blockers were mine. BLOCKER 1 (security): readInstalledCapabilitySkill scanned every capability dir and returned the first sorted match, never checking that a capability DECLARES the stem — ownership was inferred from attacker-controlled filesystem layout. Since install copies the whole bundle and the validator only checks DECLARED entries, a capability declaring `skills: []` could ship an undeclared skills/deploy/SKILL.md and win the `deploy` stem on sort order, supplying the agent-invocable instructions the user believed came from the registered capability. Stems are now bound to their owning capId via registry.capabilityClusters, and only that capability's dir is read. BLOCKER 2: the fill-in pass was gated `skills !== '*'` on the premise that applySurface materializes `full` into a concrete Set. True for applySurface — false for the installer, which is the default path: resolveProfile returns the '*' sentinel and bin/install.js passes it straight to staging. So #2322 survived on the default `full` profile, i.e. the fix didn't fix the reported bug. The registry is now plumbed to staging, and '*' stages all capability-cluster stems. Wiring this surfaced a second gap: the ADR-1239 imperative adapter (the primary install path) never threaded its registry either, which would have silently defeated the fix on the real default install. HIGH: staged capability skills were never prunable — pruneSkillDirs gates on the first-party manifest, so uninstalling a capability left its instructions live in the agent's context forever. Staged skills now carry a marker making them GSD-owned and prunable; genuinely unknown gsd-* dirs still warn and are preserved. MEDIUM: the "staged verbatim" claim was false — applySurface rewrites bodies over the whole stage dir. The tests asserted byte-equality and passed only because their fixtures contained no rewrite triggers. Claim dropped; tests now assert the rewrite against triggering content. LOW: isPathConfined is lexical, not realpath (symlink-defeatable, currently unreachable because install rejects symlinks) — comment corrected. The validator does not enforce non-empty, so isSafeCapabilitySkillStem is the sole defense, not a second layer — comment corrected and it now has traversal/NUL/absolute/empty test coverage (previously mutating it to `return true` left every test green). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2322): pin that the imperative adapter forwards a capability registry The delegation-args test deep-equalled the exact argv to installRuntimeArtifacts, so threading the composed capability registry through the ADR-1239 imperative adapter (required for #2322 — without it the default `full` install path never materializes third-party capability skills) failed it. The contract legitimately gained a parameter, so this is a stale-test correction, not a regression. Rather than deep-equalling the whole composed registry (brittle — it embeds the full agent/profile map), the test pins the leading args exactly and asserts only that a registry-shaped value is forwarded. That still fails if the adapter stops threading it, which is the regression the test exists to catch. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2322): backfill PR number 2340 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e0f969af6a |
refactor(#2246): centralize cross-platform path-separator handling (toPosixPath / toNativePath / posixNormalize) (#2247)
Replace every open-coded separator translation across the installer/hooks
source with named, tested seams in shell-command-projection.cts (the platform
seam), removing all hardcoded `/`+`\` from path handling:
- toPosixPath(p) — this machine's native path → POSIX (running-OS relative;
for local filesystem paths).
- toNativePath(p) — POSIX → native (collapses the win32 `/\//g,'\\'` ternary).
- posixNormalize(p)— unconditional `\`→`/`, OS-independent; for emitting paths
to a POSIX/bash TARGET (which may differ from the running
OS) and for parsing mixed-separator input.
core-utils.toPosixPath now delegates to the seam, so its 20+ existing consumers
resolve to one implementation; no duplicate helper.
- ~47 sites across runtime-hooks-surface, runtime-artifact-conversion,
runtime-artifact-install-plan, drift, init, worktree-safety,
installer-migrations, installer-migration-authoring, install-engine, surface,
verify, runtime-artifact-layout, schema-detect, check-command-router.
- Closes the latent POSIX-literal-backslash corruption class (the regex form
corrupts a POSIX path containing a literal backslash; split(path.sep) does not).
- New unit + fast-check property tests for all three helpers.
Closes #2246
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
d1e9491fef |
feat(#2099): drive GitHub Copilot through the EoS descriptor + multi-event hook bus (ADR-1239)
Fold Copilot's residual runtime-literal branches onto descriptor-driven hostBehaviors. Several issue premises were inaccurate (verified via research) and deliberately NOT followed: writesSharedSettings/"legacy exclusion list" (per-runtime descriptor data, copilot's false is correct); RUNTIME_CONTENT_DISPATCH.copilot + installSurface==='copilot-instructions' (already descriptor-driven); extendedHookEvents (closed Claude/Gemini enum — hooks extended in code instead); reapply ternary (already folded, kimi #2095). Real folds (all byte-parity — golden byte-identical for every runtime): - src/install-engine.cts + src/surface.cts: the two `.agent.md` filename cutovers (_copyStaged + _syncGsdDir) unified onto hostBehaviors.agentFileExtension via a new exported agentFileExtensionFor() accessor (kills the two-mechanism divergence). - src/runtime-artifact-conversion.cts: applyAgentPathRewrites' copilot skip → hostBehaviors.noPathRewrite:true (antigravity #2096 precedent). - bin/install.js uninstall: the two isCopilot cleanup branches → installSurface=== 'copilot-instructions' gate (symmetric with install-time). - bin/install.js: `!isCopilot` in the two skipSharedHooksInstall checks → hostBehaviors.skipSharedHooksInstall:true (copilot has no shared gsd-*.js hooks). - bin/install.js: three dead legacy inline-agent-loop isCopilot refs removed (copilot ∈ _DESCRIPTOR_AGENTS_RUNTIMES → unreachable; byte-parity proven by clean golden + real reachable-runtime install diffs). isCopilot dropped from 4 destructures. Zero live `runtime==='copilot'`/`isCopilot` branches remain in bin/install.js, install-engine.cts, surface.cts, or runtime-artifact-conversion.cts (AC2 guard scans all four). UPGRADE 1 (multi-event hook bus): buildCopilotHookConfig() now emits preToolUse/ postToolUse/userPromptSubmitted/sessionEnd advisory handlers alongside sessionStart (static inline bash/powershell — deterministic, golden-trackable). Only copilot.json's gsd-session.json hash changes. UPGRADE 2 (background dispatch): surfaced via the negotiated contract only — dispatch.background:true exceeds the declarative-cli baseline and survives negotiation with no downgrade warning. NO .agent.md frontmatter field (copilot has none). MCP companion out of scope (AC4 names only 2 upgrades). Tests: declarative-reference-copilot (adapter/axes/fail-closed + AC2 4-file source-grep guard) + copilot-upgrades (live 5-event hook wiring; dispatch.background negotiation). Matrix EoS note + how-to; changeset (Changed). capability-registry regenerated. Incidental flaky-test RE-ARCHITECTURE (no-defer, maintainer-directed): tests/opencode-review-reconstruction.property.test.cjs spawned ~600 synchronous execFileSync('jq') subprocesses (numRuns:200 × 3 fast-check properties, one jq per generated stream); a single jq freezing on a contended macos-22 CI runner hung the whole unit-test chunk to its 600s kill (this PR's CI). --test-force-exit can't interrupt a synchronous execFileSync, so the cure is to stop spawning per case, not just time-bound it. Re-architected to run the SHIPPED jq program over the whole fast-check corpus in ONE jq process: each generated stream is one compact-JSON array per line in a temp file, `jq -c <PROGRAM>` (no -s) applies PROGRAM to each array (`.` == the array, exactly what production's `jq -rs <file>` sees after slurping) and emits one result per line — empirically byte-identical to the per-stream form across embedded-newline/empty/quote/ unicode/null-drop cases, and file-input (like production) so there's no stdin pipe to deadlock on large I/O. ~600 spawns → 6; coverage unchanged (200-case corpus per property, deterministic seeds) plus explicit boundary/diagnostic example batches. Still property- tests the real shipped jq (no JS reimplementation). Per-call jq timeout retained as a belt-and-suspenders bound. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
847de596b8 |
fix: third-party capability skills surface correctly (#2045)
A skills-only role:feature third-party capability installed 'active' but its skills never reached the runtime surface, capability enable/set rejected it as 'unknown capability', and capability list disagreed with capability state. Three defects, fix shape 1b (teach resolveSurface, no on-disk linking):
D1 (materialization): resolveSurface built the surfaced skills Set only from the on-disk manifest; third-party cap skills live at ~/.gsd/capabilities/<id>/skills/ and never entered the Set -> surfaced:false. Fix: union registry.capabilityClusters values into the Set in the full-profile branch (idempotent for first-party, additive for third-party; prototype-pollution guarded).
D2 (enable/set unknown): setCapabilityState validated the capId against the static first-party registry instead of the composed overlay-aware registry. Fix: validate against loadRegistry({includeInstalled}) like capability-state.cts does.
D3 (list vs state): capability list derived status purely from ledger existence, never surface composition. Fix: add a 'surfaced' field to list rows sourced from the same resolver capability state uses, so list and state agree.
Regression coverage: tests/issue-2045-third-party-skills-surface.test.cjs asserts all five acceptance criteria (resolveSurface union, state surfaced:true, enable/set not-unknown, list/state agreement, first-party + unknown-id regression). gsd-test: 23916/23916 green (linux-node22 + linux-node24).
|
||
|
|
dc746ce328 |
fix(#1575): address code review M1+M2+L1
M1: gate skills:'*' sentinel on unmodified-full profile (base profile must be
'full' AND no surface mods) so tiered profiles (core/standard) don't over-stage.
M2: thread resolveAttribution through capability-writer materialize opts; add
parity test variant with non-undefined Co-Authored-By attribution.
L1: remove redundant .agent.md filter condition (.endsWith('.md') already covers it).
|
||
|
|
d671171698 |
feat(#1575): complete agent-converter descriptor cutover for copilot/antigravity + surface path parity
- Teach applySurface to build agentCtx (pathPrefix + attribution) and pass it to kind.stage() for agents kind, mirroring createRuntimeArtifactInstallPlan (ADR-1235 §1). Surface-path agents now receive path-rewrite + attribution + converter + normalize, matching install output byte-for-byte. - Pass skills:'*' sentinel for agents staging when no surface state modifications exist, so ALL agents are staged (not just those referenced by _calls_agents_). - Declare converted agents kind in copilot and antigravity capability.json; add to _DESCRIPTOR_AGENTS_RUNTIMES in bin/install.js. - Handle copilot .agent.md filename rename in both _copyStaged (install path) and _syncGsdDir (surface path). - Ship golden-parity harness (ADR-1235 §0): tests/issue-1575-agent-descriptor- parity.test.cjs asserts applySurface output is byte-identical to installRuntimeArtifacts for all 7 descriptor-driven runtimes, plus stale- cleanup convergence and prune data-loss coverage. - Update ADR-1235 with cutover progress. Cline remains deferred (rules-only local branch + local/global complication). |
||
|
|
460956bfed |
fix(#2018): applySurface empty manifest no longer deletes gsd-* agents (#2031)
* fix(#2018): applySurface empty manifest no longer deletes gsd-* agents The agent-prune loop in _syncGsdDir deleted any gsd-*.md not in the staged set. When the manifest resolved empty (null/non-object, no entries, no files key, or unresolvable install source root), the staged set was empty → every gsd-* agent was deleted. Skills were guarded by pruneSkillDirs' manifest- membership check; agents had no equivalent. - src/surface.cts: skip the agent-prune loop when the manifest is empty/absent (copy still runs — genuinely new agents are added). - tests/surface-empty-manifest-agents.test.cjs: boundary matrix — empty manifest preserves agents (Map() + undefined); non-empty manifest + empty staged still prunes (boundary); empty manifest + new staged copies new + preserves existing. Closes #2018 * docs(#2018): backfill changeset pr 2031 |
||
|
|
fb5f89db10 |
feat(#1704): destSubpath write-confinement (ADR-1239 Phase B) (#1706)
* feat(#1679): confine install writes within configHome ADR-1239 Phase B write-confinement: a pure assertDestWithinConfigHome(configDir, destSubpath) rejects a destSubpath that escapes configHome (path traversal / NUL byte) at plan-build time on BOTH the install and uninstall plan paths; surface.applySurface and installOpencodeFamilySkills route through it, and _copyStaged carries a defense-in-depth containment check. Security-load-bearing for the Phase C third-party-descriptor loader. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1704): add changeset for destSubpath write-confinement Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1704): fix windows path-portability in confinement test The N1 'accepts a true child subpath' assertion compared against path.join (no drive resolution) while the helper uses path.resolve — on Windows that mismatches the C: drive prefix. Compute the expected via path.resolve to mirror the helper. Windows-CI-only failure (local gsd-test is Mac+Linux). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
658ea33cb6 |
fix(#1615): applySurface rewrites commands kind, not just skills
Codex adversarial orthogonal review of PR #1622 surfaced that applySurface (src/surface.cts) only called rewriteStagedSkillBodies for kind='skills', skipping kind='commands'. The gap meant /gsd-surface profile changes on any runtime with commands kinds (windsurf, opencode, kilo, cursor, augment, codebuddy, gemini) wrote raw @~/.claude/... references into synced command/workflow bodies, which fail at invocation time on non-Claude runtimes. For Windsurf specifically, this left workflow files containing @~/.claude/gsd-core/commands/gsd/X.md after a profile change — paths that don't exist on a Windsurf install. Verified by the new regression test which fails before the fix (workflow bodies contained @~/.claude/) and passes after (workflow bodies reference the install target). Captures the return value of rewriteStagedCommandBodies (temp dir path — commands rewrite uses copy-then-rewrite to avoid mutating the package source), syncs from the temp dir, then cleans up. Type annotations satisfy typescript-eslint strict mode. Findings 2 (install ordering) and 3 (legacy .devin cleanup) from the same review are tracked in #1629 — both real but out of scope for #1615. |
||
|
|
eb81faaeab |
refactor(#1511): move content-rewrite engine to conversion module, delete the install.js relay (#1513)
* refactor(#1511): move content-rewrite engine to conversion module, delete the install.js relay Phase 2 of epic #1507 (ADR-1508). Behavior-preserving: makes the Runtime Artifact Conversion Module the single owner of per-runtime content rewriting and removes the last upward .cts -> bin/install.js dependency. - src/runtime-artifact-conversion.cts now owns the engine (_applyRuntimeRewrites, 5-arg with INJECTED attribution), the staged-content walkers (applyRuntimeContentRewritesInPlace / ...ForCommandsInPlace), computePathPrefix (private, exported as _computePathPrefix for tests), and the deep public seam rewriteStagedSkillBodies / rewriteStagedCommandBodies({runtime, configDir, scope, homedir?, platform?, resolveAttribution?}). - src/surface.cts:applySurface calls rewriteStagedSkillBodies directly (no resolveAttribution -> undefined). Co-Authored-By is absent from ALL rewritten content, so processAttribution is vacuous there and undefined is provably behavior-identical. surface no longer imports getInstallExports. - src/runtime-artifact-layout.cts: deleted getInstallExports / loadInstallExports / InstallExports + the GSD_TEST_MODE require('bin/install.js') relay. - bin/install.js: binds computePathPrefix / the two walkers / _applyRuntimeRewrites from the conversion module (single implementation, exports preserved for Hyrum); install callsites pass getCommitAttribution(runtime) as the injected attribution. getCommitAttribution stays here (impure install-time config I/O). - DEFECT.GENERATIVE-FIX guard: tests assert install.X === conversion.X reference identity for computePathPrefix + both walkers (no drift). New tests/enh-1511-*.test.cjs (engine, attribution injection, deep seam, prefix, layout-no-relay guard, reference-identity). 316 affected-suite tests green; lint clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD * test(#1511): make rewrite-engine path assertions Windows-robust The deep seam normalizes paths as path.resolve(configDir).replace(/\\/g,'/') and compares homedir().replace(/\\/g,'/'). Three assertions in the new test rebuilt expected paths without that normalization, so they passed on Mac/Linux but failed on Windows CI (PR #1513): - two absolute-branch asserts rebuilt resolvedTarget via path.resolve(configDir) without the backslash→slash replace → mismatch on Windows. - the $HOME-branch test fed a POSIX-literal /home/testuser, which Windows path.resolve re-roots onto the cwd drive (D:/home/...), so the resolvedTarget.startsWith(homeDir) check failed and the $HOME shorthand was never produced. Fix is test-only (engine unchanged, still behavior-preserving): mirror the engine's .replace(/\\/g,'/') in the two absolute-branch asserts, and use a real absolute path (path.resolve(os.tmpdir(), ...)) + platform: process.platform for the $HOME-branch test so the comparison holds on all platforms. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0187qgypdy1wkWRpdaf2hRuD --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
dc139b38e0 |
feat(#949): install/surface consume derived profiles/clusters (ADR-857 phase 4c) (#954)
Make install + surface read the registry's derived profileMembership/ capabilityClusters so a capability's tier drives what installs + surfaces. resolveProfile (when given the registry) unions capability skills for the profiles its tier implies before the requires: closure; resolveSurface merges capabilityClusters into the cluster map. bin/install.js, /gsd:surface, and the capability-state resolver all thread the registry. Shipped as a proven no-op: the UI capability is reconciled to tier:full (its skills were full-only in the hand-authored profiles), so it contributes only to the full profile (already the '*' sentinel) and core/standard are unchanged. Equivalence tests prove resolveProfile/resolveSurface/listSurface/staging/ capability-state are identical with vs without the registry; the core-alias staging path is verified equivalent (empty manifest → raw PROFILES.core). Closes #949 Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1373f9bfb4 |
fix(#816): mirror install command-prefix handling in _syncGsdDir (#822)
* fix(#816): mirror install command-prefix handling in _syncGsdDir applySurface() via _syncGsdDir handled command artifacts differently from a fresh install. For flat command dirs (cursor/augment/opencode/kilo) install's _copyStaged adds kind.prefix (gsd-<stem>.md) and _removeGsdEntries prunes prefix-scoped, but _syncGsdDir copied staged files verbatim (unprefixed) and pruned by exact name. So every /gsd:surface toggle wrote wrong filenames, orphaned the installed gsd-*.md, and deleted user-authored command files. _syncGsdDir's commands/agents branch now mirrors install: - flat command dirs get the gsd- prefix on copy; namespaced dirs (commands/gsd) and agents keep staged names, using install's namespacedByDir rule - prune is prefix-scoped so user files in shared flat dirs are preserved; namespaced commands/gsd stays membership-pruned so superseded commands are still removed on profile shrink The naming rule is intentionally re-implemented (not via require('bin/install.js') to avoid its module-load banner side-effect); a strict parity test asserts applySurface and installRuntimeArtifacts produce identical command filenames for opencode/kilo/cursor/augment/gemini, guarding against drift. Closes #816 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#816): add changeset fragment for PR #822 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
907cd01aa3 |
fix(#813): apply per-runtime skill path rewrites in applySurface (#817)
* fix(#813): apply per-runtime skill path rewrites in applySurface applySurface() re-staged skill artifacts but, unlike installRuntimeArtifacts(), never applied the per-runtime path rewrites. So /gsd:surface (profile/enable/disable/reset) overwrote installed SKILL.md bodies with the converter's default ~/.claude paths instead of the install target (pathPrefix), silently regressing skill path references for every skillsKind runtime until the next reinstall. applySurface now mirrors installRuntimeArtifacts: for kind.kind === 'skills' it derives pathPrefix the same way and applies applyRuntimeContentRewritesInPlace on the staged dir before syncing. - bin/install.js: export applyRuntimeContentRewritesInPlace - runtime-artifact-layout.cts: carry resolved scope on Layout; export getInstallExports; type computePathPrefix/applyRuntimeContentRewritesInPlace on InstallExports - surface.cts: lazily derive pathPrefix (only when a skills kind exists) and apply the rewrite via the shared getInstallExports accessor — single source of truth with install, only skills kinds rewritten (matches install) - tests: regression test parameterized over cursor + codex asserting post-applySurface bodies carry the install pathPrefix, not ~/.claude - CONTEXT.md: glossary updated for the applySurface rewrite parity + scope seam Closes #813 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#813): add changeset fragment for PR #817 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#813): normalize configDir prefix to forward slashes for Windows CI The #813 regression assertion compared skill bodies against a raw ${configDir}/ prefix, but production derives pathPrefix via path.resolve(configDir).replace(/\\/g, '/'). On Windows, mkdtempSync returns backslash paths while the rewritten body uses forward slashes, so the assertion would fail Windows-only (not covered by local gsd-test). Normalize the expected prefix the same way production does. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
df04aae5e4 |
enhancement(#537): migrate all hand-written bin/lib/*.cjs to TypeScript source of truth (ADR-457) (#602)
* enhancement(#537): migrate code-review-flags to TS source of truth Collapse the hand-written get-shit-done/bin/lib/code-review-flags.cjs to a TypeScript source of truth (src/code-review-flags.cts), compiled by tsc to a gitignored .cjs build artifact at the same path, per ADR-457 (build-at-publish). Second module after the semver-compare pilot (#541). Behaviour is preserved byte-for-behaviour (characterization test added in tests/code-review-flags.test.cjs locks the parser quirks). Adds compile-time type checking: CodeReviewFlags interface + CodeReviewWorkflow literal union. The require() path is unchanged, so code-review.md and the bug-3727 test keep working. The emitted .cjs is gitignored and eslint-ignored, mirroring the pilot. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 9 leaf bin/lib modules to TS source of truth ADR-457 build-at-publish, batch 1 (pure leaf modules, 0 sibling-deps): 001-legacy-orphan-files, context-utilization, redaction, artifacts, command-arg-projection, clock, ui-safety-gate, review-reviewer-selection, clusters. Each moves to src/*.cts (strict TS, typed), compiled by tsc to a gitignored .cjs at the same require() path; behaviour preserved byte-for- behaviour. Adds src/node-globals.d.ts (minimal ambient shim; "types":[]). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#537): add @types/node, drop hand-rolled node-globals shim ADR-457 migration infra: replace the temporary src/node-globals.d.ts ambient shim with @types/node@22 + "types":["node"] in tsconfig.build.json. Unblocks migrating the ~49 remaining bin/lib modules that use node:fs/path/os/ child_process. Build + full suite (3030 pass) + lint all green; no .cts type changes were needed (real Node types matched the shim). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 9 more bin/lib modules to TS (batch 2) ADR-457 build-at-publish. Clean leaves: installer-migration-report, prompt-budget. Type-error-prone leaves (were tsconfig.lint-excluded; now strict-typed and removed from that exclude list): secrets, phase-lifecycle, workstream-name-policy, decisions, validate, schema-detect. Plus runtime-name-policy. Strict type fixes narrow unknown->concrete domain types (no any/ts-ignore); behaviour preserved. Full suite green, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate runtime-slash to TS (cross-import proof) ADR-457. First cross-module TS->TS import: src/runtime-slash.cts imports ./runtime-name-policy.cjs and tsc resolves the sibling .cts types under strict (no declaration files; NodeNext .cjs->.cts mapping), emitting a correct require("./runtime-name-policy.cjs"). Confirms the recipe for coupled modules, which must be migrated in dependency order (leaves-up). Suite green, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 10 more bin/lib modules to TS (batch 3) ADR-457 build-at-publish, Wave-1 leaves: event, workstream-inventory-builder, plan-scan, fallow-runner, project-root, installer-migration-authoring, update-context, 000-first-time-baseline, runtime-homes, model-catalog. Strict typing fixed real issues (narrowing unknown, qualified fs/path calls, removed unnecessary casts); plan-scan/project-root/workstream-inventory-builder dropped from tsconfig.lint exclude. Behaviour preserved; suite green, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 5 large Wave-1 leaves to TS (batch 4) ADR-457 build-at-publish: configuration, state-document, shell-command- projection (42 dependents), security, command-aliases. shell-command- projection keeps a namespace child_process import for mock-intercept testability. loadConfig/migrateOnDisk emit synchronously (every caller uses them sync; the one awaited migrateOnDisk caller tolerates a non-Promise) — full suite (3030 pass) confirms behaviour preserved. configuration/ state-document/command-aliases dropped from tsconfig.lint exclude. Also fixes the malformed batch-3 changeset frontmatter (type/pr) that failed lint:docs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 6 Wave-2 modules to TS (batch 5) ADR-457 build-at-publish: config-schema, model-profiles, 002-codex-legacy-hooks-json, logger, active-workstream-store, adr-parser. First batch importing already-migrated siblings (configuration, model-catalog, shell-command-projection, redaction, security) via ./sibling.cjs specifiers. Strict type narrowing (typeof guards over String(unknown)); behaviour preserved; suite 3030 pass, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 5 large Wave-2 modules to TS (batch 6) ADR-457 build-at-publish: graphify, install-profiles, intel, installer-migrations, worktree-safety. installer-migrations preserves its dynamic require() loader for numbered migration modules (scoped lint suppressions). Strict typing (typeof guards over String(unknown)); behaviour preserved; suite 3030 pass, lint 0 errors. Wave 2 complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate Wave-3 modules to TS (batch 7) ADR-457 build-at-publish: planning-workspace, runtime-artifact-layout, command-routing-hub, drift. Uses `import x = require()` for export= siblings; drift's lazy require of runtime-slash hoisted to a top-level import (verified non-circular). Behaviour preserved; suite 3030 pass, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate small Wave-4 modules to TS (batch 8) ADR-457 build-at-publish: cjs-command-router-adapter, phase-command-router, surface, roadmap-upgrade. Typed the hub router handler results as the HubResult discriminated union; surface drops 4 genuinely-unused imports. Behaviour preserved; suite 3030 pass, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate core hub (2.5k LOC, 68 dependents) to TS (batch 9) ADR-457 build-at-publish: get-shit-done/bin/lib/core.cjs -> src/core.cts, preserving all 63 exports via export=. All sibling deps already migrated (shell-command-projection, model-profiles, model-catalog, worktree-safety, planning-workspace, project-root, configuration, config-schema). Strict types, no any/ts-ignore; config-schema lazy require hoisted (non-circular). Behaviour preserved (independently verified: core's shard 3030 pass / 0 fail). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537): make ESLint-coverage + test-sprawl checks migration-aware #551 test hardcoded 12 now-migrated modules as "hand-written, must be linted"; that invariant is obsoleted by the ADR-457 migration. Rewrite it to a filesystem-driven invariant that holds at every stage: a bin/lib/*.cjs must be eslint-ignored IFF it has a src/*.cts source (tsc-generated), else linted (covers package-identity, which has no TS source). Also eslint-ignore config-types.cjs (has a src counterpart) and drop the redundant tests/clock.test.cjs (clock already covered by clock-seam + bug-474 tests), which tripped the lint-test-file-count ratchet. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 9 Wave-5 router/inventory modules to TS (batch 10) ADR-457 build-at-publish: phases/verify/init/agent/task/validate/roadmap/state command routers + workstream-inventory. Router handler results typed against core's exported shapes; behaviour preserved (caught+fixed a --verify boolean flag regression mid-migration). Full suite green across all shards (only the 4 local gpg-env changeset-notes failures remain; CI passes them). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 7 Wave-5 modules to TS (batch 11) ADR-457 build-at-publish: gap-checker, docs, check-command-router, frontmatter, learnings, gsd2-import, profile-pipeline. Behaviour preserved; full suite green across all shards (only the 4 local gpg-env failures remain). Also broadens atomic-write-coverage.test.cjs to accept the tsc-compiled namespace-import form while still asserting platformWriteSync is called (safety guard intact). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate config + profile-output to TS (batch 12) ADR-457 build-at-publish: config (729 LOC), profile-output (1142 LOC). All exports preserved; cmdMigrateConfig de-asynced (migrateOnDisk is sync, awaited caller tolerates it). Behaviour preserved; suite green across all shards (only the 4 local gpg-env failures). Wave 5 complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 5 Wave-6 modules to TS (batch 13) ADR-457 build-at-publish: template, uat, workstream, roadmap, audit. Behaviour preserved (dead toPosixPath import dropped from audit; inline requires hoisted). Suite green across all shards (only the 4 local gpg-env failures). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate commands + state hubs to TS (batch 14) ADR-457 build-at-publish: commands (1305 LOC), state (2074 LOC, 17 dependents). All exports preserved; inner requires kept non-hoisted where load-order matters (install.js, per-call security); acquireStateLock cast inlined to preserve the err.code source token a structural test inspects. Behaviour preserved; suite green across all shards (only the 4 local gpg-env failures). Wave 6 complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate milestone to TS (batch 15a, hand-authored) ADR-457 build-at-publish: milestone -> src/milestone.cts. Authored directly (subagent capacity was unavailable). Also relaxes core.output()'s 3rd param to optional, matching its real always-optional call contract (unblocks remaining 2-arg output callers). Behaviour preserved; suite green across all shards (only the 4 local gpg-env failures). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate phase, verify, init to TS (batch 15, final modules) ADR-457 build-at-publish, Wave 7 (the last hubs): phase (1608 LOC), verify (1615), init (2113). Adds src/package-identity.d.cts so verify can import the permanently value-baked package-identity.cjs under strict TS. Fixes two regressions the migration introduced in verify: restore cmdValidateHealth's `return result` (callers/tests read result.warnings — it is NOT side-effect-only), and make the bug-3384 source-pattern test tolerant of the tsc-compiled bracket-notation form of the git_list_failed->W020 branch (behaviour intact). Full suite green across all shards (only the 4 local gpg-env failures); lint 0 errors. All 86 migratable bin/lib modules are now TypeScript sources. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#537): finalize ADR-457 migration — retire tsconfig.lint.json All hand-written bin/lib/*.cjs are now src/*.cts sources, so the checkJs stopgap tsconfig.lint.json (unused; not wired into eslint, scripts, or CI) is deleted per ADR-457's final step. Also gitignore the tsc-generated config-types.cjs (was still committed) for consistency with every other emitted artifact. package-identity.cjs stays value-baked (declared via src/package-identity.d.cts). Suite green; #551 ESLint-coverage test green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): add prepare script so unpacked/git installs build bin/lib artifacts ADR-457 build-at-publish: bin/lib/*.cjs are now gitignored, built by tsc. The prepack/prepublishOnly hooks cover `npm pack`/publish, but `npm install -g <dir>` and git installs run the `prepare` lifecycle — which was missing — so the unpacked install shipped without the compiled .cjs and failed at startup with "Cannot find module './lib/core.cjs'" (caught by the smoke-unpacked CI job). Add `prepare` mirroring prepublishOnly (build:lib + build:hooks). prepare does NOT run for registry consumers (they get the pre-built tarball), only for source/local/pack installs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): make CI build/lockfile checks work with gitignored bin/lib artifacts ADR-457 build-at-publish exposed two CI assumptions that bin/lib/*.cjs are always present on disk: - check:env's lockfile-sync ran `npm ci --dry-run`, which now triggers the `prepare` build (tsc) — but it runs before deps are installed, so tsc is absent and it misreported the lockfile as out of sync. Add --ignore-scripts (a lockfile check must not build). - the lint-tests job installs with --ignore-scripts (no prepare build), but lint:skill-deps require()s the built install-profiles.cjs. Add an explicit `npm run build:lib` step after install. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): narrow prepare to build:lib only (unbreak packed-smoke pack step) prepare running build:hooks emitted "✓ Copying ..." stdout during `npm pack`, which the install-smoke "Pack root tarball" step captures into $GITHUB_OUTPUT — breaking it with "Invalid format". build:lib (tsc) is silent on success and is all the unpacked/source install needs (the smoke-unpacked assertions exercise gsd-tools, i.e. bin/lib, and tolerate hook setup with `|| true`). Matches prepack. build:hooks still runs on prepublishOnly for real publishes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): wire Stryker mutation gate to build-at-publish layout The gate scored 0.00 because it mutated changed bin/lib/*.cjs that (a) were generated artifacts and (b) included modules with no coverage in the command's test set. Rework: mutation.yml now derives changed COVERED modules from src/*.cts and maps them to their built bin/lib/*.cjs; Stryker mutates those built artifacts with a no-rebuild command (mutating src/*.cts + per-mutant tsc was ~3x over the 30-min CI budget). NOTE: with the gate now correctly measuring the covered modules, their actual mutation score is 42.94% (< break 50) — a pre-existing test-coverage gap (adr-parser/prompt-budget/etc.), not introduced by this behaviour-preserving migration. Reaching 50 needs more tests, a threshold/scope change, or a waiver — a maintainer decision. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537): raise mutation coverage of covered modules above the 50 gate Adds focused example-based unit tests that kill surviving mutants in the two lowest-scoring covered modules: - tests/prompt-budget.unit.test.cjs (112 tests): 17.9% -> 97.9% - tests/adr-parser.unit.test.cjs (205 tests): 44.7% -> 89.4% Both wired into stryker.config.mjs's command. Fresh full run over the 6 covered modules now scores 82.25% (>= break 50); every covered module is >= 68%. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537,#609): parallelize mutation gate via dynamic per-module matrix The serial Stryker run timed out at 30 min once the migration's added tests made every mutant re-run ~300 tests. Replace it with a dynamic matrix so the gate completes well under budget — folded into this PR (was tracked as #609) because it's a prerequisite for this PR's mutation gate to pass. - scripts/mutation-matrix.cjs: single source of truth (covered-module -> test files) computing changed covered modules from git diff -> {has_work, matrix}. - mutation.yml: detect -> dynamic `matrix: fromJSON(...)` mutate job (one parallel shard per changed module, scoped via MUTATION_TEST_CMD to only that module's tests, 15-min/shard) -> summary job that KEEPS the legacy check name "Stryker mutation score (changed files only)" so branch protection is unchanged. Per-shard jobs report as "Stryker (<module>)". - stryker.config.mjs: commandRunner.command reads MUTATION_TEST_CMD (falls back to the full command locally). Closes #609. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537,#609): give each mutation shard ≥50% on its own tests; drop blacksmith note Per-module sharding revealed that active-workstream-store (46.5%) and frontmatter (7.4%) only cleared 50% in the old serial run via timeout-noise from the bloated 300-test command; on their own tests they were below the gate. Add focused unit tests: - tests/active-workstream-store.unit.test.cjs (115 tests): 46.5% -> 81.9% - tests/frontmatter.unit.test.cjs (165 tests): 7.4% -> 63.4% Both wired into scripts/mutation-matrix.cjs (per-module test map) and stryker.config.mjs DEFAULT_TEST_CMD. All 6 covered modules now clear break:50 with only their own tests (config-schema/context-utilization/prompt-budget/ adr-parser already did). Also removes the leftover blacksmith TODO comment — GitHub-hosted runners only; speed comes from parallel per-module shards. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537,#609): strengthen prompt-budget tests to clear the gate on its own tests prompt-budget scored 39.58% when mutation-tested with ONLY its own tests (the way the per-module CI shard runs it) — an earlier ~98% reading was inflated by accidentally running the full multi-module command. Add 96 targeted tests to tests/prompt-budget.unit.test.cjs (exact note-template text, plan-truncation arithmetic/percentages, drop-block strings, noteInjected/hardFailed booleans): scoped score 39.58% -> 68.75% (>= break 50). All 6 covered modules now clear the gate on their own tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |