* fix(#3712): confine in-process installs to a sandboxed HOME
A runtime kind may declare a global `home` override resolved from os.homedir()
rather than from the caller's configDir — codex's skills kind (`home: ".agents"`,
ADR-1239 / #2088) is the only live case. Sandboxing configDir/targetDir does not
contain it, and assertDestWithinConfigHome cannot see the class: that gate
confines a destSubpath to whatever root it is handed, and here the root IS the
escaped home. So an in-process caller that forgot to sandbox HOME wrote to, and
pruned gsd-* entries from, the developer's REAL ~/.agents/skills.
tests/agent-descriptor-parity.install.test.cjs's K1 loop did exactly that: it
iterates every agents-kind runtime (codex included) with a sandboxed targetDir
and an un-sandboxed HOME. Reproduced against a canary home on next @ adb46cdd8 —
71 gsd-* skill dirs deleted, a foreign `cloudflare` skill surviving, suite still
exit 0. It is silent because the runtime's own config home is untouched, so the
manifest keeps reporting a healthy install.
FIVE writers resolve a kind `home` and then destroy under it. Three are reachable
today — installRuntimeArtifacts, uninstallRuntimeArtifacts (install-engine.cts)
and applySurface (surface.cts). Two are descriptor-dependent and guarded against a
future descriptor change rather than a present escape: installOpencodeFamilySkills
(behind the combined-family early return) and installAgentsKindStandalone. Those
two are scoped to the single kind each destroys — passing the whole layout made
codex's unrelated skills override trip a writer that never touches it.
- src/test-home-guard.cts: refuse when a run under a test runner cannot be shown
to have sandboxed HOME. NODE_TEST_CONTEXT (set by `node --test`) gates it, so
installs outside a Node test context are untouched; GSD_TEST_MODE is unusable,
as several candidate files including the offender never set it. Homes are
compared by FILESYSTEM IDENTITY (st_dev + st_ino), not by pathname:
path.resolve() resolves neither symlinks nor case, and realpath returns a
canonical pathname that two routes to one directory can still disagree on (bind
mounts). Verified on macOS/APFS — HOME=/users/<name> made the strings differ
while naming the same directory, and the lexical form ALLOWED a write into the
physical real home. FAILS CLOSED: a pair is "different" only when both identify,
or one is definitively absent (ENOENT/ENOTDIR) while the other identifies; every
other errno is "cannot tell" and refuses. Only when neither home identifies is a
marker consulted, and it carries the sandbox PATH and must equal the home in
effect — a boolean checked first let an ambient or stale value disarm the guard.
- helpers: promote sandboxHome() out of its two byte-identical private copies,
which is also what makes them record the sandbox; the three withFakeHome()
helpers record it too. The marker NAME is duplicated as a bare string rather
than required from the compiled guard, keeping helpers.cjs's documented
no-built-lib-at-import-time contract; a test pins the two together.
- agent-descriptor-parity: sandbox HOME across the K1 loop.
- helpers-process-isolation: #3156's canary asserts on <home>/.gsd only, and its
`--cursor --local` spawn cannot reach `.agents` at all, so an assertion added
there would pass with all confinement removed. Add a discriminating row — a
`--codex --global` spawn against a seeded ambient home — which also asserts the
runtime still declares the override. Its check is a sampled inventory (dir names
+ each SKILL.md), not a tree compare.
- install-write-confinement: predicate rows through the deps seam, covering the
symlinked HOME, ambient and stale markers, and each sameDirectory branch
(both-identify, one-absent, neither-identifiable), plus wiring rows that drive
the REAL entrypoints so deleting a guard call site is red.
Verified: guard fires end-to-end against a real un-sandboxed HOME (exit 1, zero
deletions); the case-variant fail-open reproduced on APFS before the fix and
refuses after; K1 file 29/29 green with skills intact; mutation-tested — each of
the three reachable call sites, lexical-only comparison, and treating an unknown
errno as "absent" each take exactly one row red, with every mutation echoed back;
the process-isolation row negative-controlled by reverting installerEnv to its
pre-#3156 leak (16/0 -> 13/3); a full npm test leaves ~/.agents/skills at 71.
Stated residual: the two descriptor-dependent writers have no wiring test, because
no runtime declares a `home` override on those kinds and neither can be exercised
without inventing a descriptor.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore(#3712): add changeset
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#3712): let a sandbox nested inside the real home through the guard
All six Windows shards of #3725 failed on legitimately sandboxed
destinations. On Windows os.tmpdir() is %LOCALAPPDATA%\Temp — inside the
user's home — so every sandbox a test creates is a descendant of the real
home, and "does this land inside the real home?" answers yes for the safe
case and the dangerous one alike. POSIX conceals this: /tmp and
/var/folders both sit outside $HOME.
Add the missing conjunct: a destination inside the real home is allowed
only when it also sits beneath a HOME that was sandboxed away from the
passwd home. Both halves are required — dropping the first re-admits a
plain un-sandboxed install, and dropping the second decays into the
"is HOME sandboxed?" check the module rejects, which a layout resolved
before the sandbox walks straight through. Each is mutation-proven by a
row that goes red without it.
Also covers the two fail-closed branches of the new exemption, which
survived mutation to `true` with the suite green, and avoids `<user>` in
a docblock — the prompt-injection scanner reads it as a delimiter tag.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3712): close three writer/rollback gaps found reviewing the whole PR
Cross-AI review of the full PR (not just the round's delta) surfaced
three ways the guard could still be defeated:
- The nested-sandbox exemption trusted the SPELLING of a destination.
With HOME sandboxed to a directory inside the real home — legitimate on
Windows — an aliased `.agents` (symlink, junction, subordinate bind
mount) beneath it redirected an allowed path into the real home. Decide
containment on the path the write RESOLVES to: walk up to the nearest
existing ancestor, canonicalize, re-append the tail.
- `migrateLegacyDevPreferencesToSkill` is a SIXTH writer that resolves a
skills-kind `home` override. It creates rather than prunes, which is
why it was missed, and `_runLegacyInstallMigrations` runs it before
`installRuntimeArtifacts`' own assertion. Guarded, scoped to that kind.
- Worst of the three: `bin/install.js` snapshots the resolved skills root
before installing, and its outer catch rolls back by deleting and
recreating every snapshotted `gsd-*` directory there. The guard's own
throw landed in that catch, so refusing an un-sandboxed codex install
provoked exactly the mutation the guard exists to prevent. Refusals are
now marked and rethrown without rollback — nothing was written, so
there is no partial install to undo. Every other error still rolls back.
Also carries the sandbox marker into `installSpawnEnv`, so spawned
installers are not refused on passwd-less CI images, and corrects three
claims that no longer hold: "every writer" (six, and named), the
unconditional "fails CLOSED" (the passwd-less marker branch is a
deliberate weakening, and TOCTOU is out of scope), and the assertion that
Windows os.tmpdir() is always %LOCALAPPDATA%\Temp (Node honors TEMP/TMP).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(#3712): name the guard's two limits instead of overclaiming
Round-2 review found the prose had drifted ahead of the code. Corrected,
with no behavior change:
- The module still said FIVE writers; there are six, and the sixth is
now named along with why it was missed (it creates rather than prunes)
and why it carries its own assertion (it runs before the main one).
- The canonicalization docblock listed subordinate bind mounts among the
aliases it closes. It does not close them: a bind mount is not a link,
so realpath keeps the mount-point spelling. `sameDirectory` already
recorded that limit; the new helper now inherits it explicitly rather
than contradicting it. Closing it needs mount-table introspection.
- "FAILS CLOSED" was unqualified while the passwd-less marker branch is
a deliberate weakening — with no passwd entry, nothing can contradict a
marker naming the real home.
- "Refuses BEFORE any write" was too broad: legacy install migrations run
ahead of the layout-driven ones, which is exactly why the two
rollbackInstallerMigrations() calls still execute before the rethrow.
Only the codex skills-root rollback is skipped, and that is the only
_codexPreConfigRollback() call site — applySurface is never called from
bin/install.js and uninstall cannot reach it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3712): sandbox HOME in the opencode-family home-override parity rows
The last two Windows failures, and the same platform asymmetry in a
different disguise. This row drives a skills-kind `home` override on
purpose — precisely what the guard polices — but relied on the override
temp dir happening to sit outside the real home. It does on POSIX
(/tmp, /var/folders); on Windows os.tmpdir() is under %USERPROFILE%, so
the guard correctly refused and only Windows went red.
Declare the sandbox instead of depending on the platform: HOME becomes
the override itself, which is the home the call writes under. This is the
fix the guard's own message prescribes, applied to the test rather than
to the guard.
Both failing Windows shards fail on exactly these two rows and nothing
else; every other shard is green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3712): make sameDirectory answer NO when it cannot tell
Review Major 1. sameDirectory()'s only caller is the passwd-less marker
branch, which reads a `true` as permission to PROCEED:
if (marker && sameDirectory(marker, osMod.homedir())) return;
The fallthrough returned `true` whenever neither side identified — two
absent paths, or two stats failing EACCES/EPERM/EIO on a locked-down
host — on the reasoning that "cannot tell" should make the caller refuse.
That reasoning was inverted with respect to this caller: it turned the
passwd-less escape hatch into an unconditional bypass for any marker
value at all, on precisely the hosts the fallback exists to serve. Only
two things now answer yes: one resolved pathname, or two readable
identities that match. Restoring the old fallthrough takes the new row
red.
Also from review:
- Major 2 asked whether st_dev/st_ino discriminate directories on
Windows, where Node derives them from BY_HANDLE_FILE_INFORMATION. The
whole guard rests on that primitive, so assert it rather than argue it:
a row comparing two distinct temp directories, and one directory
reached by two spellings. It runs on every platform in the matrix, so
Windows answers the question itself.
- Minor 1: the refusal now names the real home it compared against, not
just the destination it refused. That is the one fact needed to tell a
true positive from a false one, and its absence is what made the
Windows case a CI-log dig rather than a glance.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3712): refuse a HOME that merely spells the real home more widely
Review round: one Blocker, four Minors, a Nit.
N2 (the one with teeth) — a destination's ancestor chain is linear, so
"inside the real home AND inside the effective HOME" admits two
arrangements, not one. The intended `effectiveHome ⊂ realHome` is the
Windows temp shape; `realHome ⊂ effectiveHome` — HOME at /Users, /home,
C:\Users — is not a sandbox at all, it is the real home reached by a
wider spelling, and it was exempting a stale destination pointing
straight at ~/.agents. Third conjunct added; the docblock no longer
claims two conditions suffice. Removing the conjunct reds the new row
and nothing else.
N4 — the migration guard resolved its OWN layout, and without
capabilityRegistry, so a registry-dependent descriptor could make it
vouch for a path the migration does not write: a guard reporting safe
while the unsafe write proceeds. It now guards the destination already
resolved by _resolveDevPreferencesSkillTarget, keyed on
`installRoot !== targetDir` — which is exactly the condition under which
a `home` override was declared, read off that same result.
N1 — CONTEXT.md gains the Test Home Guard Module glossary entry that
contributor-standards.md requires of a new Module. Not CI-enforced, so
green CI was never evidence it was met.
N3 — the docs/INVENTORY.md row was misfiled between install-fs-adapter
and install-model-override-resolver; the table is alphabetical and the
manifest already had it right. Moved, and its text now names six writers
and the third conjunct.
N5 — applySurface's signature docblock was two parameters stale; this PR
added the second of them.
N6 — the duplicated rollbackInstallerMigrations() adjacent to the new
rethrow: two identical consecutive calls, not two phases.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3712): guard the sixth writer, and close two false-ALLOW paths
Review round 3 (NEW-1, NEW-2) plus three defects Codex found in the
whole-PR pass, each reproduced before it was fixed.
NEW-1 — migrateLegacyDevPreferencesToSkill called the guard with no
`deps`, so it bound real os/process.env and could not be wiring-tested
the way the other three reachable writers were. It now takes the same
optional `deps: { os?, env? }` tail parameter. The wiring block gains
the missing fourth row, and a fifth pinning the ALLOW half; the test
file's header docblock said "FIVE writers ... the three reachable
today", contradicting the six/four statement this PR already put in
src/test-home-guard.cts, CONTEXT.md, docs/INVENTORY.md and the
changeset. Both directions of the guard's condition now fail a row
when broken — previously neither did.
NEW-2 — derivesFromSandboxedHome's docblock claimed "THREE conditions
are required, and no two of them suffice". False for {2,3}: isInside is
reflexive, so whenever conjunct 1 fires conjunct 2 already returns
false on its own. Reworded as a fast path, which is what it is.
Codex 1 (false ALLOW) — on a host with no readable passwd entry the
marker branch returned as soon as the marker matched the effective
HOME. That attests a caller sandboxed HOME and says nothing about where
an already-resolved destination points, so a layout captured before
sandboxHome() — still naming the real ~/.agents — was waved straight
through: the same stale-layout shape the primary branch refuses by
design. The marker must now identify AND contain every destination.
Codex 2 (false ALLOW) — `installRoot !== targetDir` was the stand-in
for "the skills kind declared a home override". The two are not
equivalent: the inequality is false when the override resolves onto
targetDir itself, which is exactly a configDir of $HOME/.agents. The
guard was skipped and SKILL.md written into the real home under a test
runner. _resolveDevPreferencesSkillTarget now reports hasHomeOverride
off the same resolution instead of inferring it from two paths.
Codex 3 (prose) — the shared refusal message claimed every guarded
writer prunes; the migrate writer only creates. The changeset headline
claimed in-process installer calls can no longer reach the real home,
which is wider than the guard: writeNonClaudeDefaults still writes
~/.gsd/defaults.json through os.homedir(). INVENTORY's and CONTEXT's
fail-closed sentences omitted sameDirectory's pathname-equality
shortcut. All four narrowed to what the code does.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3712): let the sandbox marker follow an overridden HOME
Found by Codex in the whole-PR pass. installSpawnEnv spreads
`overrides` last so an explicit HOME wins — deliberate, and its
docblock tells callers needing per-spawn isolation to pass their own
{ HOME, USERPROFILE }. But the #3712 marker was set before that spread,
so such a caller got HOME=<theirs> and marker=<helper default>. On a
host with no readable passwd entry the guard compares the two and
refuses a legitimately sandboxed spawn — tests/install.test.cjs:7143
and install-shared.cjs's own runInstaller both take that path.
The marker is now derived from the final HOME unless the caller
supplied one explicitly. The contract test asserted HOME after an
override but not the marker, which is why it stayed green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(#3712): name the shipped guard condition, not the deleted one
Review round 4 of #3725. Two artifacts this PR adds still described
`target.installRoot !== targetDir` in the PRESENT tense as the live guard
condition on `migrateLegacyDevPreferencesToSkill`. The shipped condition is
`runtime && target.hasHomeOverride` (src/install-engine.cts:510).
This is not ordinary doc drift. The named condition is the exact false-ALLOW
the previous round closed: a `home` override resolving onto `targetDir` — a
configDir of `$HOME/.agents`, which is where codex's override points — makes
the inequality FALSE while the override is declared, so the guard was skipped.
A maintainer reading CONTEXT.md:290 as authoritative would believe the guard
still skips that case.
- tests/install-write-confinement.test.cjs — the ALLOW-half row's comment.
Its "teeth" rationale is unchanged and still correct as written.
- CONTEXT.md:290 — the Test Home Guard Module glossary entry, a documented
PR gate. Now states the condition and names the inequality only as what it
is NOT, with the reason.
The three surviving mentions of the inequality are all past-tense or negated
(src/install-engine.cts:448, :508 and the sibling test comment at :3698) and
are correct as they stand.
Verified: `npm run lint:ci` exit 0; full `npm test` 31327 tests / 31312 pass /
0 fail / 14 skipped, run with TMPDIR unset.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3712): canonicalization fails closed, matching identify's errno split
Codex full-PR review of #3725, run against the round-4 head.
`resolveThroughLinks` caught EVERY realpathSync error and fell back to
`path.resolve(dest)` — the lexical spelling. That inverts the function's own
purpose. An aliased `<sandbox>/.agents` that cannot be canonicalized keeps its
sandbox spelling, satisfies the nested-sandbox exemption at :227, and the write
is ALLOWED into the real home — the exact escape this walk exists to close. The
module documents that it fails CLOSED with ONE named exception (the marker
branch); this was a second, unnamed one.
Split by errno, and deliberately by the SAME split `identify` already draws
rather than a second policy in one module — both answer "does this path exist
as named?", so they must not disagree:
ENOENT / ENOTDIR -> walk up. The ordinary case: a fresh install resolves a
destination nothing has created yet, so realpath fails on the leaf and on
every not-yet-created ancestor. Refusing here rejects every install.
anything else (EACCES, EPERM, ELOOP, EIO) -> refuse. The component exists but
cannot be resolved, so the guard cannot tell where the write lands.
Three rows in the predicate block, beside the other aliasing rows:
- a symlink CYCLE in the destination path (ELOOP) -> REFUSE
- a destination that does not exist yet (ENOENT) -> ALLOW
- a component behind a regular file (ENOTDIR) -> ALLOW
Teeth checked against the artifact the test loads, not the source: reverting
the condition to the swallow-everything shape in the compiled
test-home-guard.cjs turns row 1 — and only row 1 — red. The ENOTDIR row caught
a stale build during development, which is the point of asserting on the
compiled file.
CONTEXT.md and the resolveThroughLinks docblock both record the new behaviour,
so this does not repeat the prose-vs-code drift the round-4 finding was about.
The changeset's existing scope sentence now bounds "six writers" to the
`installRuntimeArtifacts` call tree and names `cmdGenerateDevPreferences` —
which resolves the same codex `home` override through `getGlobalSkillsBase` and
writes SKILL.md beneath it unguarded. It has no in-process caller today (its
only direct require-and-call is a spawnSync with HOME sandboxed), so it is
latent rather than live, and whether it belongs in this PR is raised with the
maintainer rather than decided here.
Verified: `npm run lint:ci` exit 0; full `npm test` 31330 tests / 31315 pass /
0 fail / 14 skipped, TMPDIR unset.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
657 lines
36 KiB
JavaScript
657 lines
36 KiB
JavaScript
'use strict';
|
||
|
||
// allow-test-rule: source-text-is-the-product #2875 — the J-row assertions
|
||
// pattern-match rendered agent-.md frontmatter (`effort:`, `model:`,
|
||
// disallowedTools, branding text). That frontmatter IS the deployed artifact
|
||
// each runtime loads at dispatch time (no typed IR exists between
|
||
// applyAgentFrontmatterExtensions/injectEffortFrontmatter and the file a
|
||
// runtime reads) — matching CONTRIBUTING.md's `source-text-is-the-product`
|
||
// exemption, not a workaroundable "hide the grep in a parser" case.
|
||
|
||
/**
|
||
* agent-descriptor-parity.install.test.cjs — #2875 Part 2 (the agents-bypass
|
||
* closure), 50-test-matrix.md sections H, I, J (rows H1-H8, I1-I3, J1-J10).
|
||
*
|
||
* HONEST BASELINE (rewritten — a prior revision of this file was found to
|
||
* test the wrong thing on every axis; see the fixed defects below):
|
||
*
|
||
* bin/install.js's inline agent-staging loop is GONE (deleted in the same
|
||
* commit that made every runtime descriptor-driven for agents). There is no
|
||
* second, independently-maintained agent-staging pipeline left in this
|
||
* codebase to diff against — an "old pipeline vs new pipeline" comparison is
|
||
* therefore IMPOSSIBLE post-deletion, and a prior revision of this file's
|
||
* header claiming to compare against "bin/install.js's inline agent-staging
|
||
* loop" was false the moment that loop was deleted.
|
||
*
|
||
* What THIS revision actually proves instead, and how:
|
||
*
|
||
* 1. The REAL production entry point (`installAgentsKindStandalone`,
|
||
* `install-engine.cjs`) is driven directly, against the REAL, BUILT
|
||
* `capability-registry.cjs` — no synthetic registry override. Every H
|
||
* row therefore byte-compares actual output written by a real
|
||
* `capabilities/<runtime>/capability.json` edit, not a hand-rolled
|
||
* stand-in for one. A wrong-but-syntactically-valid `converter` name
|
||
* landing in a real descriptor changes the ACTUAL side's output and is
|
||
* caught (see H8's red-proof, which demonstrates this directly).
|
||
*
|
||
* 2. `computeExpectedOutput` (the "oracle") independently assembles the
|
||
* EXPECTED bytes by calling the individual conversion PRIMITIVES
|
||
* directly: `composeWorkflow`, `applyAgentPathRewrites`,
|
||
* `processAttribution`, the runtime's converter function (dispatched
|
||
* off `EXPECTED_CONVERTER_NAME_BY_RUNTIME` — a hand-verified,
|
||
* independent map, NOT read from the descriptor under test),
|
||
* `applyAgentFrontmatterExtensions`, `normalizeAgentBodyForRuntime`.
|
||
* These primitives are — by construction, not accident —
|
||
* single-sourced: there is no live duplicate of `composeWorkflow` or
|
||
* `applyAgentPathRewrites` to diff against either, because #2875 Part 2
|
||
* collapsed the duplication into these shared functions. Reusing them
|
||
* here does not defeat the test: the property under test in every H row
|
||
* is "does capability.json's declared `converter` name resolve to the
|
||
* CORRECT converter and get invoked in the correct position of the
|
||
* pipeline" — which the independent `EXPECTED_CONVERTER_NAME_BY_RUNTIME`
|
||
* map exists specifically to keep decoupled from the descriptor.
|
||
*
|
||
* 3. Both scopes are exercised for every runtime that declares a
|
||
* per-scope `agents` kind (claude, cline, codex, hermes, kilo, opencode,
|
||
* kimi-code × global+local). The prior revision was global-only, which
|
||
* is exactly the class of gap that let two separate agents-drop
|
||
* regressions reach `next` undetected: cline-local (fixed alongside
|
||
* this rewrite — capabilities/cline/capability.json's `local`
|
||
* artifactLayout now declares an `agents` kind) and kimi-code-local
|
||
* (fixed the same way — its `local` artifactLayout previously declared
|
||
* no `agents` kind at all, so the deleted inline loop's implicit
|
||
* scope-gate was silently replaced with NO gate, dropping every
|
||
* kimi-code local install's agents/gsd-*.md entirely).
|
||
*
|
||
* H8 is mandatory, not optional: a parity harness never demonstrated failing
|
||
* is decoration. It feeds the oracle a DELIBERATELY WRONG (but real,
|
||
* allowlisted) converter name and asserts the comparison goes red against
|
||
* the REAL (correctly-configured) production output — proving that if
|
||
* capability.json's declared converter ever regressed, this exact harness
|
||
* would catch it.
|
||
*
|
||
* WHAT THIS FILE DOES NOT PROVE: H8's red-proof demonstrates exactly one
|
||
* failure class — the oracle and the real descriptor path resolving to a
|
||
* DIFFERENT converter for the same runtime. It says nothing about a bug
|
||
* INSIDE a shared primitive (`composeWorkflow`, `applyAgentPathRewrites`,
|
||
* `processAttribution`, `applyAgentFrontmatterExtensions`,
|
||
* `normalizeAgentBodyForRuntime`, or a named converter itself): both the
|
||
* oracle (point 2 above) and the real descriptor path call the identical
|
||
* function, so a regression inside one of those functions changes BOTH
|
||
* sides identically and every H row stays green. That is a deliberate,
|
||
* unavoidable consequence of point 2's single-sourcing (there is no second,
|
||
* independently-implemented copy of those primitives left to diff against
|
||
* post-#2875-Part-2) — this file is a converter-WIRING parity gate, not a
|
||
* substitute for direct unit coverage of the primitives themselves (which
|
||
* live in their own owning test files, e.g. `runtime-artifact-conversion`'s
|
||
* suite).
|
||
*/
|
||
|
||
const { test } = require('node:test');
|
||
const assert = require('node:assert/strict');
|
||
const fs = require('node:fs');
|
||
const os = require('node:os');
|
||
const path = require('node:path');
|
||
const { cleanup, createTempDir, sandboxHome } = require('./helpers.cjs');
|
||
|
||
const REPO_ROOT = path.join(__dirname, '..');
|
||
const LIB_DIR = path.join(REPO_ROOT, 'gsd-core', 'bin', 'lib');
|
||
|
||
const runtimeArtifactConversion = require(path.join(LIB_DIR, 'runtime-artifact-conversion.cjs'));
|
||
const installModelOverrideResolver = require(path.join(LIB_DIR, 'install-model-override-resolver.cjs'));
|
||
const installEngine = require(path.join(LIB_DIR, 'install-engine.cjs'));
|
||
const capabilityRegistry = require(path.join(LIB_DIR, 'capability-registry.cjs'));
|
||
const { composeWorkflow } = require(path.join(LIB_DIR, 'workflow-fragments.cjs'));
|
||
const installBin = require(path.join(REPO_ROOT, 'bin', 'install.js'));
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// Fixture builders
|
||
// ---------------------------------------------------------------------------
|
||
|
||
/** Deterministic sample agent sources — NOT the real agents/ tree (the real
|
||
* tree's exact roster can change independently of this suite). Covers:
|
||
* ~/.claude/ + $HOME/.claude/ (anchored + bare) path forms, a Co-Authored-By
|
||
* trailer, and one row-J4 "disallowedTools hit" agent name plus one "miss". */
|
||
const SAMPLE_AGENTS = {
|
||
'gsd-planner.md': [
|
||
'---',
|
||
'name: gsd-planner',
|
||
'description: Plans phases for GSD workflows.',
|
||
'tools: Read, Write, Edit, Bash',
|
||
'---',
|
||
'',
|
||
'Reads @~/.claude/gsd-core/commands/gsd/plan-phase.md and $HOME/.claude/CLAUDE.md.',
|
||
'Bare forms too: ~/.claude and $HOME/.claude.',
|
||
'References Claude Code and CLAUDE.md and .claude/settings.json.',
|
||
'',
|
||
'Co-Authored-By: Claude <noreply@anthropic.com>',
|
||
'',
|
||
].join('\n'),
|
||
'gsd-plan-checker.md': [
|
||
'---',
|
||
'name: gsd-plan-checker',
|
||
'description: Checks plans for GSD workflows.',
|
||
'tools: Read, Grep, Glob',
|
||
'---',
|
||
'',
|
||
'A read-only checker agent (row J4 "hit" — declared in READONLY_AGENT_DISALLOWED_TOOLS).',
|
||
'',
|
||
].join('\n'),
|
||
};
|
||
|
||
function buildSourceTree(agentFiles) {
|
||
const root = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-agent-parity-src-'));
|
||
const commandsGsd = path.join(root, 'commands', 'gsd');
|
||
fs.mkdirSync(commandsGsd, { recursive: true });
|
||
const agentsDir = path.join(root, 'agents');
|
||
fs.mkdirSync(agentsDir, { recursive: true });
|
||
for (const [name, content] of Object.entries(agentFiles)) {
|
||
fs.writeFileSync(path.join(agentsDir, name), content);
|
||
}
|
||
return { root, commandsGsd, agentsDir };
|
||
}
|
||
|
||
/** A fresh "install destination" dir with a `.gsd-source` marker pointing at
|
||
* `commandsGsd` — the same marker findInstallSourceRoot/findAgentsSourceRoot
|
||
* read (runtime-artifact-layout.cts). */
|
||
function buildTargetDir(commandsGsd) {
|
||
const targetDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-agent-parity-dest-'));
|
||
fs.writeFileSync(path.join(targetDir, '.gsd-source'), commandsGsd);
|
||
return targetDir;
|
||
}
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// Oracle — independently-assembled EXPECTED output (see header doc)
|
||
// ---------------------------------------------------------------------------
|
||
|
||
/** Hand-verified, independent of any capability.json — this is the thing
|
||
* every H row's ACTUAL side (a real capability.json) is checked against.
|
||
* `null` means converter:null (identity — claude, kimi-code). */
|
||
const EXPECTED_CONVERTER_NAME_BY_RUNTIME = {
|
||
claude: null,
|
||
'kimi-code': null,
|
||
cline: 'convertClaudeAgentToClineAgent',
|
||
codex: 'convertClaudeAgentToCodexAgent',
|
||
hermes: 'convertClaudeAgentToHermesAgent',
|
||
kilo: 'convertClaudeToKiloFrontmatter',
|
||
opencode: 'convertClaudeToOpencodeFrontmatter',
|
||
};
|
||
|
||
/** Converter functions callable by name — bin/install.js still owns
|
||
* cline/codex/kilo/opencode's (never migrated to runtime-artifact-conversion.cjs);
|
||
* hermes's is descriptor-native (#2875 Part 2 / J9-J10). */
|
||
const NAMED_CONVERTERS = {
|
||
convertClaudeAgentToClineAgent: installBin.convertClaudeAgentToClineAgent,
|
||
convertClaudeAgentToCodexAgent: installBin.convertClaudeAgentToCodexAgent,
|
||
convertClaudeAgentToHermesAgent: runtimeArtifactConversion.convertClaudeAgentToHermesAgent,
|
||
convertClaudeToKiloFrontmatter: installBin.convertClaudeToKiloFrontmatter,
|
||
convertClaudeToOpencodeFrontmatter: installBin.convertClaudeToOpencodeFrontmatter,
|
||
};
|
||
|
||
/** kilo/opencode take an options bag (`{isAgent, modelOverride}`), resolved
|
||
* ONCE per call via the single shared precedence resolver (J5-J8) — every
|
||
* other converter here takes only `content`. */
|
||
const MODEL_OVERRIDE_CONVERTER_NAMES = new Set(['convertClaudeToKiloFrontmatter', 'convertClaudeToOpencodeFrontmatter']);
|
||
|
||
/**
|
||
* Assemble the EXPECTED per-file output for `runtime` from `agentsDir`,
|
||
* calling the shared conversion primitives directly (see header doc for why
|
||
* this is not circular). `converterNameOverride`, when passed, replaces
|
||
* `EXPECTED_CONVERTER_NAME_BY_RUNTIME[runtime]` — used ONLY by H8's
|
||
* red-proof to inject a deliberately wrong converter.
|
||
*/
|
||
function computeExpectedOutput(runtime, agentsDir, ctx, converterNameOverride) {
|
||
const converterName = converterNameOverride !== undefined ? converterNameOverride : EXPECTED_CONVERTER_NAME_BY_RUNTIME[runtime];
|
||
const { pathPrefix, attribution, targetDir } = ctx;
|
||
const out = new Map();
|
||
const entries = fs.readdirSync(agentsDir, { withFileTypes: true });
|
||
for (const entry of entries) {
|
||
if (!entry.isFile() || !entry.name.endsWith('.md')) continue;
|
||
const agentSourcePath = path.join(agentsDir, entry.name);
|
||
let content = fs.readFileSync(agentSourcePath, 'utf8');
|
||
// Step 0 (#2995): strip gsd:section markers — same order the real
|
||
// pipeline uses (stageAgentsForRuntimeWithConverter, install-profiles.cts).
|
||
content = composeWorkflow(content, { sourcePath: agentSourcePath });
|
||
const agentName = runtimeArtifactConversion.deriveAgentName(entry.name);
|
||
// Step 1: path rewrites
|
||
content = runtimeArtifactConversion.applyAgentPathRewrites(content, runtime, pathPrefix);
|
||
// Step 2: attribution
|
||
content = runtimeArtifactConversion.processAttribution(content, attribution);
|
||
// Step 3: converter — dispatched off the INDEPENDENT expected-name map,
|
||
// never off the descriptor under test.
|
||
if (converterName) {
|
||
const fn = NAMED_CONVERTERS[converterName];
|
||
assert.ok(typeof fn === 'function', `oracle: no converter function registered for "${converterName}"`);
|
||
if (MODEL_OVERRIDE_CONVERTER_NAMES.has(converterName)) {
|
||
const modelOverride = installModelOverrideResolver.resolveAgentModelOverride(
|
||
agentName,
|
||
installModelOverrideResolver.readGsdEffectiveModelOverrides(targetDir),
|
||
installModelOverrideResolver.readGsdRuntimeProfileResolver(targetDir),
|
||
);
|
||
content = fn(content, { isAgent: true, modelOverride });
|
||
} else {
|
||
content = fn(content);
|
||
}
|
||
}
|
||
// converter:null (claude, kimi-code) — content unchanged by this step.
|
||
// Step 4: frontmatter extensions (effort/disallowedTools)
|
||
content = runtimeArtifactConversion.applyAgentFrontmatterExtensions(content, { runtime, agentName, targetDir });
|
||
// Step 5: normalize colon->hyphen refs
|
||
content = runtimeArtifactConversion.normalizeAgentBodyForRuntime(
|
||
content,
|
||
runtime,
|
||
runtimeArtifactConversion.readGsdCommandNames(),
|
||
);
|
||
out.set(entry.name, content);
|
||
}
|
||
return out;
|
||
}
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// Real descriptor path — the REAL production entry point, REAL registry
|
||
// ---------------------------------------------------------------------------
|
||
|
||
/**
|
||
* Drive the ACTUAL production entry point (`installAgentsKindStandalone`)
|
||
* against the REAL, built `capability-registry.cjs` — no override. This is
|
||
* exactly what `bin/install.js`'s `install()` calls for every runtime whose
|
||
* layout is not otherwise reached by the generic `installRuntimeArtifacts`
|
||
* loop (and, for the runtimes reached by that loop, produces identical
|
||
* output to it — both route through the SAME `convertedAgentsKind` /
|
||
* `_resolveNamedConverter` dispatch in runtime-artifact-layout.cts).
|
||
*/
|
||
function runRealDescriptorPath(runtime, scope, targetDir, ctx) {
|
||
const resolvedProfile = { name: 'full', skills: '*', agents: new Set() };
|
||
const result = installEngine.installAgentsKindStandalone(
|
||
runtime,
|
||
targetDir,
|
||
scope,
|
||
resolvedProfile,
|
||
ctx.pathPrefix,
|
||
() => ctx.attribution,
|
||
);
|
||
const out = new Map();
|
||
if (!result) return out;
|
||
for (const entry of fs.readdirSync(result.destDir, { withFileTypes: true })) {
|
||
if (entry.isFile()) out.set(entry.name, fs.readFileSync(path.join(result.destDir, entry.name), 'utf8'));
|
||
}
|
||
return out;
|
||
}
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// Comparison helper
|
||
// ---------------------------------------------------------------------------
|
||
|
||
/** Every (runtime, scope) pair the REAL capabilities/<runtime>/capability.json
|
||
* today declares an `agents` kind for (measured 2026-08-17). K1 below is the
|
||
* machine-checked guarantee that this list cannot silently go stale — it
|
||
* sweeps the real registry and fails if a declarant is missing here. */
|
||
const RUNTIME_SCOPE_PAIRS = [
|
||
['claude', 'global'], ['claude', 'local'],
|
||
['cline', 'global'], ['cline', 'local'],
|
||
['codex', 'global'], ['codex', 'local'],
|
||
['hermes', 'global'], ['hermes', 'local'],
|
||
['kilo', 'global'], ['kilo', 'local'],
|
||
['opencode', 'global'], ['opencode', 'local'],
|
||
['kimi-code', 'global'], ['kimi-code', 'local'],
|
||
];
|
||
|
||
/** Model override literal shared by J8's two `_stageWithModelOverride` calls. */
|
||
const J8_OVERRIDE_MODEL = 'shared/explicit-model';
|
||
|
||
/**
|
||
* Standalone helper (module scope, no test-context access — the
|
||
* CONTRIBUTING.md "Never use try/finally inside test bodies" exemption) for
|
||
* J8: stage a single-agent source tree with a real `.planning/config.json`
|
||
* model_overrides block for `runtime` through the real descriptor path.
|
||
*/
|
||
function _stageWithModelOverride(runtime, overrideModel) {
|
||
const agentFiles = { 'gsd-planner.md': SAMPLE_AGENTS['gsd-planner.md'] };
|
||
const { commandsGsd, root } = buildSourceTree(agentFiles);
|
||
const targetDir = buildTargetDir(commandsGsd);
|
||
fs.mkdirSync(path.join(targetDir, '.planning'), { recursive: true });
|
||
fs.writeFileSync(
|
||
path.join(targetDir, '.planning', 'config.json'),
|
||
JSON.stringify({ model_overrides: { 'gsd-planner': overrideModel } }),
|
||
);
|
||
try {
|
||
const ctx = { pathPrefix: `${targetDir}/`, attribution: undefined, targetDir };
|
||
return runRealDescriptorPath(runtime, 'global', targetDir, ctx);
|
||
} finally {
|
||
cleanup(root);
|
||
cleanup(targetDir);
|
||
}
|
||
}
|
||
|
||
function comparePipelines(runtime, scope, converterNameOverride) {
|
||
const { commandsGsd, agentsDir, root } = buildSourceTree(SAMPLE_AGENTS);
|
||
const targetDir = buildTargetDir(commandsGsd);
|
||
try {
|
||
const ctx = { pathPrefix: `${targetDir}/`, attribution: undefined, targetDir };
|
||
const expected = computeExpectedOutput(runtime, agentsDir, ctx, converterNameOverride);
|
||
const actual = runRealDescriptorPath(runtime, scope, targetDir, ctx);
|
||
return { expected, actual };
|
||
} finally {
|
||
cleanup(root);
|
||
cleanup(targetDir);
|
||
}
|
||
}
|
||
|
||
/** Asserts H1-H6 + H7 in one shot: same filenames (Set equality, order-free)
|
||
* AND byte-identical content per filename. */
|
||
function assertMapsIdentical(expected, actual) {
|
||
assert.deepEqual(
|
||
[...expected.keys()].sort(),
|
||
[...actual.keys()].sort(),
|
||
'filenames diverged between the oracle and the real descriptor path (row H7)',
|
||
);
|
||
for (const [name, expectedContent] of expected) {
|
||
assert.equal(
|
||
actual.get(name),
|
||
expectedContent,
|
||
`content diverged for ${name} between the oracle and the real descriptor path`,
|
||
);
|
||
}
|
||
}
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// H rows — the parity gate, one row per (runtime, scope)
|
||
// ---------------------------------------------------------------------------
|
||
|
||
for (const [runtime, scope] of RUNTIME_SCOPE_PAIRS) {
|
||
test(`agent-descriptor-parity: H row — ${runtime} (${scope}) real descriptor output matches the independent oracle`, () => {
|
||
const { expected, actual } = comparePipelines(runtime, scope);
|
||
assert.ok(expected.size > 0, 'fixture produced no oracle output — test is vacuous');
|
||
assertMapsIdentical(expected, actual);
|
||
});
|
||
}
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// H7 — the harness compares filenames, not only content (meta)
|
||
// ---------------------------------------------------------------------------
|
||
|
||
test('agent-descriptor-parity: H7 — a filename-only divergence fails the harness', (t) => {
|
||
const { commandsGsd, agentsDir, root } = buildSourceTree(SAMPLE_AGENTS);
|
||
const targetDir = buildTargetDir(commandsGsd);
|
||
t.after(() => {
|
||
cleanup(root);
|
||
cleanup(targetDir);
|
||
});
|
||
const ctx = { pathPrefix: `${targetDir}/`, attribution: undefined, targetDir };
|
||
const expected = computeExpectedOutput('claude', agentsDir, ctx);
|
||
const actual = runRealDescriptorPath('claude', 'global', targetDir, ctx);
|
||
// Deliberately rename one actual-side entry — same bytes, different name.
|
||
const [firstName, firstContent] = [...actual.entries()][0];
|
||
actual.delete(firstName);
|
||
actual.set(`RENAMED-${firstName}`, firstContent);
|
||
assert.throws(
|
||
() => assertMapsIdentical(expected, actual),
|
||
/filenames diverged/,
|
||
'a renamed output file must fail the harness',
|
||
);
|
||
});
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// H8 — the harness can actually FAIL (mandatory, not optional)
|
||
// ---------------------------------------------------------------------------
|
||
|
||
test('agent-descriptor-parity: H8 — a deliberately-wrong (but real, allowlisted) expected converter turns the harness RED', (t) => {
|
||
// Hermes's REAL capability.json declares convertClaudeAgentToHermesAgent.
|
||
// Feed the ORACLE a different, real, allowlisted converter name
|
||
// (convertClaudeAgentToCodexAgent) instead — simulating exactly the failure
|
||
// mode row H exists to catch: capability.json's declared converter silently
|
||
// diverging from the correct one. The REAL descriptor path is untouched and
|
||
// still uses hermes's real (correct) converter, so this proves: IF
|
||
// capability.json ever regressed to the wrong name, THIS harness's H row
|
||
// for hermes would go red exactly like this.
|
||
const { commandsGsd, agentsDir, root } = buildSourceTree(SAMPLE_AGENTS);
|
||
const targetDir = buildTargetDir(commandsGsd);
|
||
t.after(() => {
|
||
cleanup(root);
|
||
cleanup(targetDir);
|
||
});
|
||
const ctx = { pathPrefix: `${targetDir}/`, attribution: undefined, targetDir };
|
||
const wrongExpected = computeExpectedOutput('hermes', agentsDir, ctx, 'convertClaudeAgentToCodexAgent');
|
||
const realActual = runRealDescriptorPath('hermes', 'global', targetDir, ctx);
|
||
let threw = false;
|
||
let observedDiff = null;
|
||
try {
|
||
assertMapsIdentical(wrongExpected, realActual);
|
||
} catch (err) {
|
||
threw = true;
|
||
observedDiff = err.message;
|
||
}
|
||
assert.equal(threw, true, 'H8 FAILED: the harness did not go red for a deliberately-wrong expected converter');
|
||
assert.match(observedDiff, /content diverged for gsd-(planner|plan-checker)\.md/, 'expected the content-diverged assertion to name the mismatched file');
|
||
// Verbatim red-proof output for the record (see CHANGES report):
|
||
console.log(`H8 red-proof observed: ${observedDiff}`);
|
||
});
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// I1-I3 — per-agent resolution context
|
||
// ---------------------------------------------------------------------------
|
||
|
||
test('agent-descriptor-parity: I3 — deriveAgentName matches the pipeline exactly, including a no-.md-suffix boundary', () => {
|
||
assert.equal(runtimeArtifactConversion.deriveAgentName('gsd-planner.md'), 'gsd-planner');
|
||
// Boundary: a filename with no trailing .md is returned unchanged (the
|
||
// regex has nothing to match) — matches `entry.name.replace(/\.md$/, '')`.
|
||
assert.equal(runtimeArtifactConversion.deriveAgentName('gsd-planner'), 'gsd-planner');
|
||
assert.equal(runtimeArtifactConversion.deriveAgentName('gsd-planner.MD'), 'gsd-planner.MD');
|
||
});
|
||
|
||
test('agent-descriptor-parity: I2 — real descriptor path with no agentCtx is unaffected (converter-only)', (t) => {
|
||
const { commandsGsd, agentsDir, root } = buildSourceTree(SAMPLE_AGENTS);
|
||
const targetDir = buildTargetDir(commandsGsd);
|
||
t.after(() => {
|
||
cleanup(root);
|
||
cleanup(targetDir);
|
||
});
|
||
const runtimeArtifactLayout = require(path.join(LIB_DIR, 'runtime-artifact-layout.cjs'));
|
||
const realLayout = runtimeArtifactLayout.resolveRuntimeArtifactLayout('claude', targetDir, 'global');
|
||
const agentsKindEntry = realLayout.kinds.find((k) => k.kind === 'agents');
|
||
const stagedDir = agentsKindEntry.stage({ name: 'full', skills: '*', agents: new Set() }); // no agentCtx
|
||
t.after(() => cleanup(stagedDir));
|
||
const planner = fs.readFileSync(path.join(stagedDir, 'gsd-planner.md'), 'utf8');
|
||
const original = fs.readFileSync(path.join(agentsDir, 'gsd-planner.md'), 'utf8');
|
||
assert.equal(planner, original, 'no agentCtx must leave content byte-identical to source (converter:null == identity)');
|
||
});
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// J1-J4 — frontmatter extensions
|
||
// ---------------------------------------------------------------------------
|
||
|
||
test('agent-descriptor-parity: J1 — claude effort is injected via applyAgentFrontmatterExtensions', () => {
|
||
const content = '---\nname: gsd-planner\ndescription: x\n---\n\nBody.\n';
|
||
const viaShared = runtimeArtifactConversion.applyAgentFrontmatterExtensions(content, { runtime: 'claude', agentName: 'gsd-planner', targetDir: null });
|
||
assert.match(viaShared, /^effort: /m, 'expected an effort: key to be injected for claude');
|
||
});
|
||
|
||
test('agent-descriptor-parity: J2 — effort resolving to inherit writes NO effort: key at all, exercised via the REAL guard in applyAgentFrontmatterExtensions (#3533 trap row)', (t) => {
|
||
// #2875 Part 2 defect fix: a prior revision of this row asserted the DUMB
|
||
// half (injectEffortFrontmatter DOES emit the literal 'inherit' if called
|
||
// with it) and never called applyAgentFrontmatterExtensions at all — so
|
||
// deleting the `universalEffort !== 'inherit'` guard at
|
||
// runtime-artifact-conversion.cts:3504 left this row green. This revision
|
||
// writes a REAL .planning/config.json under targetDir (readGsdEffectiveEffortConfig
|
||
// walks up from targetDir looking for it — install-effort-resolver.cts)
|
||
// and calls the REAL applyAgentFrontmatterExtensions end to end, so removing
|
||
// that guard makes THIS assertion fail (verified: red with the guard
|
||
// removed, green with it restored — see CHANGES report).
|
||
const targetDir = fs.mkdtempSync(path.join(os.tmpdir(), 'gsd-agent-parity-j2-'));
|
||
t.after(() => cleanup(targetDir));
|
||
fs.mkdirSync(path.join(targetDir, '.planning'), { recursive: true });
|
||
fs.writeFileSync(
|
||
path.join(targetDir, '.planning', 'config.json'),
|
||
JSON.stringify({ effort: { agent_overrides: { 'gsd-inherit-agent': 'inherit' } } }),
|
||
);
|
||
const content = '---\nname: gsd-inherit-agent\ndescription: x\n---\n\nBody.\n';
|
||
const out = runtimeArtifactConversion.applyAgentFrontmatterExtensions(content, { runtime: 'claude', agentName: 'gsd-inherit-agent', targetDir });
|
||
assert.doesNotMatch(out, /^effort:/m, 'an agent resolving to "inherit" must get no effort: key at all — not even effort: inherit');
|
||
// Sanity: injectEffortFrontmatter itself is dumb and WOULD write the
|
||
// literal if called with it — the guard is applyAgentFrontmatterExtensions
|
||
// never calling it for 'inherit', which is exactly what the assertion above proves.
|
||
assert.match(runtimeArtifactConversion.injectEffortFrontmatter(content, 'inherit'), /^effort: inherit$/m);
|
||
});
|
||
|
||
test('agent-descriptor-parity: J3 — a runtime NOT declaring agentFrontmatterExtensions gets nothing injected', () => {
|
||
const content = '---\nname: gsd-plan-checker\ndescription: x\n---\n\nBody.\n';
|
||
const out = runtimeArtifactConversion.applyAgentFrontmatterExtensions(content, { runtime: 'opencode', agentName: 'gsd-plan-checker', targetDir: null });
|
||
assert.equal(out, content, 'opencode declares no agentFrontmatterExtensions — output must be byte-identical to input');
|
||
});
|
||
|
||
test('agent-descriptor-parity: J4 — disallowedTools injected only on a READONLY_AGENT_DISALLOWED_TOOLS hit', () => {
|
||
const content = '---\nname: gsd-plan-checker\ndescription: x\n---\n\nBody.\n';
|
||
const hit = runtimeArtifactConversion.applyAgentFrontmatterExtensions(content, { runtime: 'claude', agentName: 'gsd-plan-checker', targetDir: null });
|
||
assert.match(hit, /^disallowedTools: /m, 'gsd-plan-checker is a declared read-only agent — expected a disallowedTools hit');
|
||
|
||
const missContent = '---\nname: gsd-not-a-readonly-agent\ndescription: x\n---\n\nBody.\n';
|
||
const miss = runtimeArtifactConversion.applyAgentFrontmatterExtensions(missContent, { runtime: 'claude', agentName: 'gsd-not-a-readonly-agent', targetDir: null });
|
||
assert.doesNotMatch(miss, /^disallowedTools:/m, 'an agent absent from READONLY_AGENT_DISALLOWED_TOOLS must get no disallowedTools key');
|
||
});
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// J5-J8 — model-override resolution (kilo/opencode), single-sourced
|
||
// ---------------------------------------------------------------------------
|
||
|
||
test('agent-descriptor-parity: J5 — explicit model_overrides[agent] wins (highest precedence)', () => {
|
||
const modelOverrides = { 'gsd-planner': 'anthropic/explicit-model' };
|
||
const runtimeResolver = { resolve: () => ({ model: 'anthropic/tier-model' }) }; // would win if precedence were wrong
|
||
const result = installModelOverrideResolver.resolveAgentModelOverride('gsd-planner', modelOverrides, runtimeResolver);
|
||
assert.equal(result, 'anthropic/explicit-model');
|
||
});
|
||
|
||
test('agent-descriptor-parity: J6 — falls back to the runtime tier resolver when no explicit override exists', () => {
|
||
const runtimeResolver = { resolve: (agentName) => (agentName === 'gsd-planner' ? { model: 'anthropic/tier-model' } : null) };
|
||
const result = installModelOverrideResolver.resolveAgentModelOverride('gsd-planner', null, runtimeResolver);
|
||
assert.equal(result, 'anthropic/tier-model');
|
||
});
|
||
|
||
test('agent-descriptor-parity: J7 — neither configured resolves to null (omit), never "" or the string "null"', () => {
|
||
const result = installModelOverrideResolver.resolveAgentModelOverride('gsd-planner', null, null);
|
||
assert.equal(result, null);
|
||
assert.notEqual(result, '');
|
||
const contentWithoutModelOverride = installBin.convertClaudeToOpencodeFrontmatter(
|
||
'---\nname: gsd-planner\ndescription: x\ntools: Read\n---\n\nBody.\n',
|
||
{ isAgent: true, modelOverride: result },
|
||
);
|
||
assert.doesNotMatch(contentWithoutModelOverride, /^model:/m, 'an omitted override must not appear as a model: key at all');
|
||
});
|
||
|
||
test('agent-descriptor-parity: J8 — kilo and opencode both resolve model overrides through the REAL descriptor path (installAgentsKindStandalone), proving ONE shared resolution path, not a tautology', () => {
|
||
// #2875 Part 2 defect fix: a prior revision of this row called
|
||
// resolveAgentModelOverride TWICE with identical arguments and asserted
|
||
// equality — true of ANY pure function regardless of whether kilo/opencode
|
||
// are actually wired to it. This revision drives BOTH runtimes through the
|
||
// REAL production entry point with a REAL .planning/config.json
|
||
// model_overrides block and asserts BOTH staged outputs carry the SAME
|
||
// resolved model — which only happens if both are genuinely wired to the
|
||
// one shared installModelOverrideResolver.resolveAgentModelOverride.
|
||
const kiloOut = _stageWithModelOverride('kilo', J8_OVERRIDE_MODEL);
|
||
const opencodeOut = _stageWithModelOverride('opencode', J8_OVERRIDE_MODEL);
|
||
// J8_OVERRIDE_MODEL is a fixed constant: assert the exact expected
|
||
// frontmatter line, matched line-wise, rather than building a RegExp from
|
||
// an interpolated value (CodeQL js/incomplete-sanitization — a
|
||
// metachar-bearing value would silently widen the match).
|
||
const expectedModelLine = `model: ${J8_OVERRIDE_MODEL}`;
|
||
assert.ok(
|
||
kiloOut.get('gsd-planner.md').split('\n').some((line) => line.trim() === expectedModelLine),
|
||
'kilo must apply the shared model override via the real descriptor path',
|
||
);
|
||
assert.ok(
|
||
opencodeOut.get('gsd-planner.md').split('\n').some((line) => line.trim() === expectedModelLine),
|
||
'opencode must apply the SAME shared model override via the real descriptor path',
|
||
);
|
||
});
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// J9-J10 — hermes branding converter
|
||
// ---------------------------------------------------------------------------
|
||
|
||
test('agent-descriptor-parity: J9 — hermes branding converter is byte-identical to the shared generic branding-rewrite function, including \\bClaude Code\\b word-boundary semantics', () => {
|
||
const content = 'Claude Code and ClaudeCodeExtra and CLAUDE.md and .claude/foo and reClaude Code.\n';
|
||
const viaSharedFn = runtimeArtifactConversion.applyAgentBrandingRewrites(content, 'hermes');
|
||
const viaNamedConverter = runtimeArtifactConversion.convertClaudeAgentToHermesAgent(content);
|
||
assert.equal(viaSharedFn, viaNamedConverter, 'the named converter must be a pure delegate to the generic branding-rewrite function');
|
||
// \bClaude Code\b: neither "ClaudeCodeExtra" (no space/boundary between
|
||
// "Claude" and "Code") nor "reClaude Code" (no boundary between the 'e' of
|
||
// "re" and the 'C' of "Claude" — both word chars) satisfy the word-boundary
|
||
// requirement, so BOTH are left untouched; only the standalone occurrence is
|
||
// rewritten.
|
||
assert.match(viaSharedFn, /Hermes Agent and ClaudeCodeExtra and HERMES\.md and \.hermes\/foo and reClaude Code\./);
|
||
});
|
||
|
||
test('agent-descriptor-parity: J10 — the branding converter is descriptor-data-driven, not hardcoded to hermes strings', () => {
|
||
const content = 'Claude Code uses CLAUDE.md under .claude/.\n';
|
||
// A runtime with NO brandingRewrites declared gets nothing rewritten.
|
||
assert.equal(runtimeArtifactConversion.applyAgentBrandingRewrites(content, 'claude'), content);
|
||
// hermes (the only runtime with brandingRewrites AND no dedicated converter
|
||
// pre-#2875) gets ITS OWN declared rewrite table applied — proving the
|
||
// function reads the runtime's descriptor rather than a hermes-hardcoded literal.
|
||
const hermesOut = runtimeArtifactConversion.applyAgentBrandingRewrites(content, 'hermes');
|
||
assert.notEqual(hermesOut, content);
|
||
assert.match(hermesOut, /Hermes Agent uses HERMES\.md under \.hermes\/\./);
|
||
});
|
||
|
||
// ---------------------------------------------------------------------------
|
||
// K1 — migration completeness: every registry runtime with an `agents` kind
|
||
// is reachable from the REAL production entry point, not the (now-deleted)
|
||
// inline loop.
|
||
// ---------------------------------------------------------------------------
|
||
|
||
/**
|
||
* bin/install.js's inline agent-staging loop and its `_DESCRIPTOR_AGENTS_RUNTIMES`
|
||
* gate were DELETED in #2875 Part 2 Task C — there is no longer a symbol to
|
||
* assert absent (a source-text check would violate `local/no-source-grep` and
|
||
* would prove nothing about runtime behavior anyway, per CLAUDE.md's
|
||
* "Behavioral tests are required"). Row K1 is instead proven the only way
|
||
* that is actually meaningful once the code is gone: for EVERY `role:
|
||
* "runtime"` capability in the REAL capability-registry that declares an
|
||
* `agents` kind (either scope), the REAL production entry point
|
||
* (`installRuntimeArtifacts` — which internally routes combinedFamilyInstall
|
||
* runtimes like kilo/opencode through `installOpencodeFamilyAgents`, #2875
|
||
* Part 2 Task A) actually materializes agents/ on disk. If any runtime were
|
||
* still silently depending on the deleted inline loop, this call would write
|
||
* nothing to agents/ for it (the deleted code was the ONLY thing that used to
|
||
* write it for the seven runtimes migrated in this change) and the assertion
|
||
* below would fail.
|
||
*/
|
||
test('agent-descriptor-parity: K1 — every registry runtime declaring an agents kind is reachable from installRuntimeArtifacts (the inline loop is gone)', (t) => {
|
||
const runtimesWithAgentsKind = Object.entries(capabilityRegistry.runtimes || {})
|
||
.filter(([, cap]) => {
|
||
const layout = cap.runtime && cap.runtime.artifactLayout;
|
||
if (!layout) return false;
|
||
const entries = [...(layout.global || []), ...(layout.local || [])];
|
||
return entries.some((e) => e.kind === 'agents');
|
||
})
|
||
.map(([id]) => id);
|
||
|
||
assert.ok(runtimesWithAgentsKind.length >= 7, 'sanity: expected at least the seven #2875 Part 2 runtimes to declare an agents kind');
|
||
assert.ok(runtimesWithAgentsKind.includes('kimi-code'), 'kimi-code (found via golden fixture, not analysis) must be covered here');
|
||
assert.ok(runtimesWithAgentsKind.includes('cline'), 'cline must be covered here');
|
||
|
||
// #3712: this loop reaches `codex`, whose global skills kind declares a `home`
|
||
// override (`.agents`) resolved from os.homedir() rather than from targetDir.
|
||
// Sandboxing targetDir alone does NOT contain it — before this line the call
|
||
// below pruned every gsd-* skill from the developer's REAL ~/.agents/skills
|
||
// (71 -> 0) while the suite still exited 0. Sandbox HOME for the whole loop so
|
||
// codex's skills land in <homeDir>/.agents/skills inside a temp dir we clean.
|
||
const homeDir = createTempDir('gsd-adp-home-');
|
||
t.after(() => cleanup(homeDir));
|
||
sandboxHome(t, homeDir);
|
||
|
||
for (const runtime of runtimesWithAgentsKind) {
|
||
const { commandsGsd, root } = buildSourceTree(SAMPLE_AGENTS);
|
||
const targetDir = buildTargetDir(commandsGsd);
|
||
t.after(() => {
|
||
cleanup(root);
|
||
cleanup(targetDir);
|
||
});
|
||
const resolvedProfile = { name: 'full', skills: '*', agents: new Set() };
|
||
const result = installEngine.installRuntimeArtifacts(runtime, targetDir, 'global', resolvedProfile, () => undefined, undefined);
|
||
const agentsKindEntry = result.kinds.find((k) => k.kind === 'agents');
|
||
assert.ok(agentsKindEntry, `${runtime}: installRuntimeArtifacts reported no agents kind in the executed plan — it did not go through the descriptor path`);
|
||
const writtenFiles = fs.readdirSync(agentsKindEntry.destDir).filter((f) => f.endsWith('.md'));
|
||
assert.ok(writtenFiles.length > 0, `${runtime}: agents kind reported but nothing was actually written to ${agentsKindEntry.destDir}`);
|
||
}
|
||
});
|