f7df920681f233ae0fe064ee659550bdf41ff708
34 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ea594300d9 |
fix(#3606): validate hook-kind coverage at call sites and dispatch generically (#3687)
* test(#3606): pin hook-kind coverage in the wired guard * fix(#3606): validate hook-kind coverage at call sites and dispatch generically * fix(#3606): address review - segment-granular narrowing, zero-coverage diagnosis, quick.md, fragment extraction * fix(#3606): drop stale shrink-ack, export HOOK_GROUP_KINDS, dedupe scanner regex * chore(#3606): regenerate install-tree fixtures for new wave-post fragment * chore(#3606): sync canonical launcher preamble into new fragment * fix(#3606): keep fragment preamble ahead of first gsd_run mention * fix(#3606): revert sync script's preamble move in explore.md * chore(#3606): regenerate derived manifests post-rebase * chore(#3606): allowlist peer test files - base was red on the count lane * chore(#3606): regenerate inventory for peer's verify-command-grounding doc * chore(#3606): grounding test maps to its own module by longest prefix * chore(#3606): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
9e4f0e99ad |
fix(#3631): exclude only __pycache__-resident bytecode from the consent digest (#3650)
* test(3631): failing-first coverage for bytecode-cache in the consent hash
bundleContentHash digests a walk with no exclusion, so a routine 'python3 -m unittest'
inside a Python-backed capability bundle writes __pycache__ under the bundle, the
recomputed hash stops matching the consent record, and the capability silently goes
inactive — no error, no warning, and loop render-hooks then omits its step and gate.
Two distinct triggers, and the second is the sharper one: collectBundleEntries pushes a
{kind:'dir'} entry for EVERY directory and the digest emits a TAG_DIR marker for it, so an
EMPTY __pycache__/ flips the hash before a single .pyc is written. A fix filtering only
*.pyc would leave that live. Verified by execution against the built lib: 5 of 7 probe
rows diverge from intent today, including the empty-directory row.
The anti-regression rows are the point of the shape: editing a real scripts/m.py and
adding node_modules/pkg/index.js must BOTH still change the hash. node_modules is
deliberately not excludable — its contents are required at runtime, so dropping it from
the digest would stop consent binding executable content. The symlink row pins ordering:
exclusion must apply after the lstat fail-closed rejection, never before.
Refs #3631
* fix(3631): exclude derived bytecode caches from the consent digest
RED proven at e5ba8f1fe on the remote runner: 8 failures, exactly the rows predicted to
fail, with the four anti-regression rows already green.
collectBundleEntries now skips a hardcoded, gitignore-independent set from the DIGEST:
basenames __pycache__, .pytest_cache, .DS_Store, and any .pyc/.pyo file. Matching is
byte-exact on the raw Buffer name (the walk never utf8-decodes) and case-sensitive, so the
digest does not vary with how a name happens to be spelled on a case-insensitive volume.
Three properties were preserved deliberately, each pinned by a test:
- The filter runs AFTER the lstat symlink/non-regular fail-closed rejection. Filtering
first would have turned the exclusion into a way to smuggle a symlink past the check;
a symlink named x.pyc still throws.
- Excluded entries still count toward BUNDLE_MAX_FILES and BUNDLE_MAX_TOTAL_BYTES. The
caps guard the WALK; the digest answers a different question, and exclusion must not
become an unbounded-bytes hole.
- An excluded DIRECTORY is neither emitted as a TAG_DIR marker nor recursed into. The
directory marker was the sharper half of this bug: an empty __pycache__ flipped the
hash before any .pyc existed, so a *.pyc-only filter would have left it live.
The issue proposed either a gitignore-aware walk or a list including node_modules. Both
are rejected. A consent binding must not delegate its scope to a .gitignore the bundle
author does not control — one line there would drop arbitrary executable content out of
the hash. And node_modules holds code that is required at runtime; excluding it would stop
consent binding executable content, turning a usability bug into a supply-chain hole. What
makes __pycache__ different is that CPython validates each .pyc against its sibling
source, which remains hashed, so a real code change still invalidates consent.
Docs: CONTEXT.md's 'EVERY regular file AND directory' claim is corrected in place.
ADR-2363's residual-gap section said the walk had 'no exclusions' — per
docs/adr/README.md ('ADRs are append-only') that is corrected by a dated amendment rather
than an in-place edit. Its D4 argument is unaffected: skill bodies are .md and stay bound.
Fixes #3631
* fix(3631): narrow the digest exclusion after two isolated security reviews
The first cut of this fix passed the full suite and was still wrong. Both orthogonal
reviews rejected it, and the second one found a hole that has nothing to do with Python.
HIGH — an excluded DIRECTORY was 'continue'd before recursion, so its whole subtree was
permanently outside the digest. Declared hook script paths allow '_', '.' and '/' with no
directory or extension rule, so hooks:[{script:'__pycache__/run.js'}] installed, executed
via node, and its bytes could be rewritten forever without moving the hash. Ship benign
v1, collect consent, then own the machine. No Python involved.
FALSE RATIONALE — the justification I wrote into the code, CONTEXT.md, the ADR amendment
and the changeset claimed CPython validates a cached .pyc against its sibling source, so
the source staying hashed kept consent honest. That is not true, and I proved it by
execution rather than argument: default timestamp invalidation compares only the source's
mtime and size, both settable by anyone who can write the bundle. A forged pyc ran while
the .py was byte-identical.
Also wrong: '*.pyc' matched anywhere, but a legacy sourceless scripts/x.pyc IS importable,
so excluding it was a live vector.
Narrowed to what is actually defensible:
- a DIRECTORY named __pycache__/.pytest_cache has only its TAG_DIR marker suppressed;
the walk still recurses and hashes every non-excluded child.
- .pyc/.pyo are excluded ONLY when the parent basename is exactly __pycache__.
- a regular FILE named __pycache__, and a DIRECTORY named x.pyc, stay bound.
- declared hook paths containing a __pycache__/.pytest_cache segment or a .pyc/.pyo
basename are now rejected in both validator copies — a file named .pyc can contain
perfectly valid JavaScript, so the exclusion must not be reachable from a declared
surface.
Accepted residual risk, stated plainly in ADR-2363 and CONTEXT.md instead of explained
away: a forged __pycache__/mod.pyc matching an unmodified, still-hashed mod.py executes
without moving the digest. Before this change that write was detected. It is accepted to
stop routine bytecode caching from silently deactivating capabilities, and it is bounded —
the attacker needs post-consent write access, everything outside __pycache__/*.pyc stays
hashed, and no declared surface can point into the excluded space.
Known limitation, not papered over: .pytest_cache CONTENTS still move the digest. Only the
directory marker is suppressed. Excluding that subtree would reopen the HIGH finding.
Refs #3631
* fix(3631): drop the .DS_Store exclusion and pin what the caps actually bind
Second round of isolated review findings. The hardening closed the two original holes —
both re-reviews confirmed that by execution — but it introduced a new one of the same
shape, and left three claims unbacked.
HIGH, self-inflicted: .DS_Store was excluded from the digest at any depth, but the hook
path validator was hardened only for __pycache__/.pytest_cache/.pyc/.pyo. So
script:'hooks/.DS_Store' was ACCEPTED, runnableHookCommand emits the bare quoted path for
a non-.js name (the branch .sh hooks already use), and capability-source copies it with
its mode bit intact. Ship it +x with a benign shebang, take consent, then rewrite it
forever — the digest never moves. Fixed by DELETING the .DS_Store exclusion rather than
teaching the validator about it: .DS_Store has nothing to do with this issue's Python
bytecode symptom, and an excluded filename is a permanently unhashed name. The narrower
the exclusion, the smaller the hole.
The residual-risk bound in ADR-2363 and CONTEXT.md claimed declared surfaces cannot reach
excluded space. That is false and is now stated correctly: node resolves an unregistered
extension through the default .js handler, so a hashed, consent-covered hooks/run.js that
requires '../__pycache__/mod.pyc' reaches it in one hop. The validator guard raises the
bar for DECLARED surfaces; it does not contain the risk. The two bounds that are real —
post-consent write access required, everything outside __pycache__/*.pyc still hashed —
are kept.
The BUNDLE_MAX_FILES boundary test had gone vacuous: it padded with root-level *.pyc,
which the hardening made non-excluded, so it no longer proved anything about excluded
entries while the ADR claimed the caps were test-pinned. It now pads __pycache__/f{i}.pyc,
with the arithmetic re-derived by execution (capability.json + the still-counted
__pycache__ dir + N). BUNDLE_MAX_TOTAL_BYTES had zero coverage at all and is now pinned by
a sparse 32 MiB __pycache__/big.pyc that must still trip the size cap — the test that
proves exclusion did not become an unbounded-bytes hole.
Added the parity assertion CLAUDE.md's Generative Fix Divergence rule requires for the two
isSafeHookScriptPath copies, and proved it can fail: mutating one BUILT copy to drop .pyo
made the parity check report the divergence. Also pinned semantics that were correct but
untested and would have survived mutation — __pycache__/sub/x.pyc stays hashed (the parent
resets to sub, which is the recursion threading itself), .pytest_cache/y.pyc stays hashed,
and .pyo in both directions, which was a free surviving mutant.
Changeset rewritten: it still described the rejected wholesale-exclusion semantics.
Refs #3631
* chore(3631): backfill changeset PR number (#3650)
---------
Co-authored-by: sim <sim@local>
|
||
|
|
3ab0007164 |
enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19) preserveUserArtifacts held user files only in an in-memory Map across the wipe, so any process death between preserve and restore lost them outright. Seven call sites, not the four the issue records. Three of them never called the helper at all - they open-coded the same read/wipe/write - so searching for callers under-counted by construction; the extra sites were found by sweeping for the pattern instead. The worst is the mainline install path, where the crash window spans the entire gsd-core tree copy rather than a single rmSync. Adds src/user-artifact-staging.cts: durable on-disk staging with a record written after the copies land as the commit point, plus recovery of orphaned batches on the next run - without recovery the staged bytes survive but the user's file is still gone, which would pass its own test while delivering nothing. Routes copyPreservingSymlink through installFs() so staging cannot bypass the install fs seam, and reunites its symlink-safety docblock with the function it documents. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-3574 with four claims disproved by implementation Implementing Phase 6 disproved four statements the ADR rests on. The central decision - no single materializer - is unaffected and stands. Corrected: decision 3 was already satisfied, so nothing was extracted; the agents-bypass runtime set omitted claude, kilo and opencode, and closing it needed three new pieces of descriptor contract rather than proceeding on its own terms; three of the four blockers the layout comment names were already stale; and F19 is seven call sites, not four. Records the generalizable lesson: the defect is the pattern of holding user data in memory across a wipe, not the helper, so searching for callers of the helper under-counts by construction. Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and notes that copyPreservingSymlink needed routing through the install fs seam before it could be reused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close dangling-symlink blind spot and harden staging recovery An adversarial review found the F19 staging work shipped red and unsafe. Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween missed dangling symlinks in both its root check and its per-segment walk, because it probed with existsSync, which is false for a link whose target does not exist. Fixing only the new module would have reused a guard that was itself blind. This guard protects the whole install tree. Recovery no longer throws: it degrades per entry and per file, so one bad batch cannot block the others. Previously an unrecoverable entry propagated out of the first statement of install and uninstall, before the cleanup that would have removed it - wedging the installer permanently. Partial fs adapters now throw on any omitted method instead of silently reaching the real filesystem, closing the trap that let a test poison list pass while real IO happened. Staged names must be flat, recovery refuses a dangling destination symlink, and a batch whose recovery genuinely failed is no longer swept - it was discarding the only durable copy of the file it had just failed to restore. Replaces three tests that could not fail, including the one labelled negative proof. Known limitation, documented not closed: concurrent installs sharing a staging key can still lose a batch. A real fix needs a cross-process lock. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enh(#2875): make the descriptor authoritative for the agents kind Deletes the inline agent-staging loop in bin/install.js and the _DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from its capability descriptor instead of an inline hostBehaviors dispatch. Closing it needed three pieces of contract the descriptor pipeline never had, all reducible to one missing input - per-agent resolution context: a frontmatter-extensions step for claude's effort and disallowedTools, per-agent model-override resolution for kilo and opencode, and a named branding converter for hermes, whose rewrite data was already declared. Seven runtimes were on the loop, not the six the design recorded - kimi-code was found by a golden fixture, not by analysis. claude-local and kimi-code both silently lost their agents mid-change; the fixtures caught both and the cause was fixed rather than the fixtures regenerated. A parity harness gates the migration: both pipelines over identical inputs, byte-identical output including filenames, per runtime. It is demonstrated red before being trusted. Surface and install paths converge for all seven, which also fixes surface previously writing no agents for these runtimes. Codex's config.toml strip stays put - it mutates host config, which no descriptor kind models. Also routes install-model-override-resolver and install-effort-resolver through the install fs seam. Both leaked real filesystem IO from the install call tree; the stricter adapter is what exposed them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): record the agents-descriptor migration and correct the ADR count The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host integration guide told readers to join a set that is gone. Replaces that with what is now true - declare an agents entry and it installs, on the surface path as well as install - and points anyone needing a per-agent transform at the three extension points rather than at a new inline branch. Corrects the ADR amendment: seven runtimes were on the inline loop, not six. kimi-code was found by a golden fixture going red, not by reading. That is the third short count this phase, all from enumerating by symbol or set membership when the thing that matters is a behavior. Adds the Changed changeset for the surface-path convergence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-2866 - claude global always wrote agents on disk The claude row's global=[skills] described what capability.json declared, not what the installer wrote. bin/install.js's inline agent-staging loop was never scope-gated and never consulted the descriptor, so a claude --global install has always written agents/gsd-*.md. Phase 6 closes the gap by deleting that loop and declaring agents on claude's descriptor at global scope. On-disk bytes are unchanged - the golden fixtures did not move, which is the evidence that the descriptor, not the installer, was incomplete. #2218 is unaffected: agents are not trigger-bearing, so the wider row does not introduce a new shadowing case. Records the warning that an incomplete descriptor is invisible while a second code path silently does its work, and only surfaces when the two are forced into agreement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close review findings across staging, agents and the parity harness Two independent reviews of this branch found defects the local gates missed. Security: a dangling symlink at a migration destination allowed writing outside configDir - the same class this change claimed to close, missed at the terminal write of the flow being added. The staging-root resolver threw as the first statement of install and uninstall, so a hostile symlink bricked both, and symlinked-configDir users lost uninstall as well as install; it now degrades instead of aborting. Recovery gained a source-side symlink check and now refuses a relative destDir, which resolved against cwd. Converter dispatch gained a runtime allowlist - lint-time validation stopped mattering once this branch promoted that dispatch from the surface path to real installs. Correctness: claude --local --minimal exited 1 because the minimal profile legitimately yields zero agents and the new path treated that as a failure. cline --local silently lost its agents - its descriptor declared none while the deleted loop wrote them unconditionally. The agents prune was widened to any gsd-* entry and destroyed user files it never owned. The parity harness, on which the migration's safety argument rested, drove a synthetic registry and never byte-compared the shipped descriptors; two of its trap rows could not fail. It now drives the real registry across 13 runtime-scope rows including kimi-code and cline-local, and its red-proof is demonstrated by corrupting a live capability.json. Three goldens that had encoded the cline regression as expected behavior were corrected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close findings from both mandated review engines /security-review found the staging source-side walk honouring GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write destination. A symlinked files/ component dereferenced because copyPreservingSymlink lstats the leaf only, so an intermediate link is followed. The source walk no longer honours the opt-in; the destination check still does. /code-review spec axis found this branch had reintroduced its own bug: migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after the legacy dir was wiped and before the staged batch was restored, so a planted symlink bricked uninstall permanently and orphaned the batch. Refusal kept, abort removed. kimi-code local silently lost its agents, the same class as the cline bug, and the parity harness recorded that exclusion as intentional - the third test in this branch to pin a regression as correct. --minimal now creates an empty agents/ dir that never existed. Behaviour restored rather than softening the changeset, so its byte-identical claim stays true. Standards axis: try/finally removed from twelve test bodies, fast-check properties added for parseOwnerPid, boundary coverage at the grace window and the ancestor-probe depth, a parity assertion for the staging-root helper duplicated across two files, and the 8-deep config walk deduplicated. Records 60-review.json with every finding and disposition from five passes, including the smells left unfixed and why. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): prune stale agents unconditionally in minimal mode The previous round stopped an empty agents/ directory being created when the resolved profile yields no agents. That was implemented by skipping the agents kind entirely, which also skipped its stale-agent prune - so a full to minimal downgrade left stale gsd-* agents behind. The deleted inline loop pruned unconditionally and only skipped writing. Those are three separate conditions, not one: prune always, write only when there is something to write, create the directory only when writing. Both call sites now run _removeGsdEntries before the empty-staged early exit. The symlink-escape guard moved with it, since the prune also touches dest. Codex .toml agents and the config.toml stanzas are cleaned again, and user-owned agents are still preserved. The agents/ directory is left in place after a prune empties it, matching every sibling kind - none of them remove the destination directory itself. Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh install, so fixture generation is unaffected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): document interrupted-install recovery for user-owned files The durable-staging fix is invisible to the user it protects. Someone whose install died mid-flight has no way to know USER-PROFILE.md was staged before the delete, that the next run restores it, or that recovery happens at the start of that run rather than in the background. Written as the task the user has - finish the interrupted command - rather than as a description of the mechanism, and states what it will not do: overwrite a file already present, or touch staging belonging to another install still running. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2875): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2875): assert the J8 model override without building a regex CodeQL flagged incomplete string escaping: the assertion interpolated the override value into a RegExp while escaping only forward slashes, which is meaningless in a constructor, leaving real metacharacters unescaped. The failure direction was the dangerous one - a metacharacter would have made the match more permissive, so the row would pass when it should fail. That matters here because J8 exists precisely because an earlier revision was a tautology; the rewrite reintroduced a different way for the same assertion to stop discriminating. Replaced with a line-wise exact match, so no regex is constructed at all. Swept the other test files this branch adds; no sibling instances. lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full metachar-escape copy, so a single slash replace slipped under it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
967bddba37 |
fix(#3384): strip mcp__* tool grants from zcode-installed subagents (#3483)
* fix(#3384): strip mcp__* tool grants from zcode-installed subagents ZCode's dispatcher treats every mcp__<server>__* entry in an agent's tools: frontmatter as a required MCP server and hard-fails the subagent spawn (CONFIGURATION_ERROR) when it is not connected, whereas Claude Code treats the same grants as an optional allowlist. ZCode shared Claude's verbatim agents copy (converter: null), so all 8 MCP-granted agents failed to spawn out of the box with zero MCP servers configured. Add convertClaudeAgentToZcodeAgent — a line-surgical converter that filters mcp__* entries out of the frontmatter tools: grant list (both inline comma and YAML block-list shapes) and preserves every other byte. Declare it on both of zcode's capability.json agents entries and cut zcode over to the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES) so the legacy inline loop stops deleting+re-copying the converted agents raw. Claude Code, Kimi, and Gemini install behavior is unchanged. * chore(#3384): link changeset fragment to pr 3483 --------- Co-authored-by: sim <sim@local> |
||
|
|
0396d9cab1 |
enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection
The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.
That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.
Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.
review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.
* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines
Two CI failures from the first push, both mine:
1. lint-tests: the new regression test split readFileSync content on a
literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
"\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).
2. golden-install-parity / workflow-size-budget / workflow-compat: editing
gsd-core/workflows/review.md changes its content hash and byte size, and
both are pinned in committed baselines. Regenerated via the repo's own
generators (npm run size:baseline, npm run gen:golden).
The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.
Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).
* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape
The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.
* enhance(#2483): also guard the claude leg against auto-memory injection
CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).
* enhance(#2483): match the claude binary in command position, not argument position
The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as
|
||
|
|
cf6de5e1c0 |
feat(#2871): resolve triggers and host precedence, not just placement (#3291)
* test(#2871): failing-first suite for trigger-surface resolution 23 tests over the 50-test-matrix rows. RED by construction: resolveTriggerSurface and DEFAULT_TRIGGER_PRECEDENCE do not exist yet, and the validator silently ignores triggerPrecedence today. Written in the per-runtime describe idiom the other four runtime-artifact-layout suites use, not a table. The rows that carry the weight: windsurf must NOT report a shadow it does not have, since its global scope emits only agents and agents are not trigger-bearing; agents and kimi-agents must be absent from the output for every runtime; and reordering a runtime's triggerPrecedence must flip the winner, which is the only assertion that proves the axis is read rather than decorative. Stems are injected, never scanned, so the surface is assertable with no filesystem. * feat(#2871): resolve triggers and host precedence, not just placement resolveTriggerSurface(runtime, scopes) returns every /gsd-<name> trigger a runtime emits, with the scope and kind that produced it, whether the host registers it directly or only through a router, and which artifact shadows it. resolveRuntimeArtifactLayout is untouched -- its 7 callers need placement only and the issue requires them unchanged. AGENTS ARE NOT TRIGGER-BEARING, and ADR-2866 said they were. The host-integration matrix models command and dispatch as separate interface points: an agent is invoked through the Agent tool's subagent_type, not by typing a slash trigger, and _copyStaged never applies the kind prefix to an agents entry. So agents and kimi-agents are excluded from the surface entirely, and this commit amends ADR-2866 with a dated correction. #2218's conclusion is unchanged -- the collision is strictly commands-vs-skills, and claude's local /gsd-* trigger surface is still fully shadowed -- but the ADR implied the local agents surface was lost too, and it is not. That correction is what makes windsurf come out right. Its global scope emits only agents, so it has no global trigger and its local commands are unshadowed. Model agents as trigger-bearing and windsurf falsely reports a full shadow. The triggerPrecedence axis lands on all 19 descriptors as an ordered kind list, one value with one owner, rather than a numeric rank spread across N kind entries with nothing keeping them consistent. Validation uses a required-with-default shape that has no precedent in this validator -- every existing axis is hard-required -- so a third-party capability.json omitting the field still validates, which is what ADR-894's additive-only contract promises. Winner resolution reads Phase 1's scope rank first, then the kind ordering. A test reorders the axis and asserts the winner flips, since an axis that is added, validated and never consulted would pass every other assertion. shadowedBy ships unread. Phase 4 (#2873) is its first consumer, per this issue's out-of-scope note. Verified via the remote runner. * fix(#2871): single-source namespacedByDir and close two test gaps Four findings from the isolated adversarial review. The namespacedByDir rule had reached three copies -- install-engine, surface, and the new trigger resolver -- one of which carried a hand-written keep-in-sync comment and no assertion. That is this repo's generative-fix-divergence class. Extracted to one exported predicate all three now call. Verified by diverging one copy deliberately: the existing #816 parity test failed, and passes again on revert. The omission test was vacuous. Row 16 asserted that a descriptor without triggerPrecedence still validates, but built its fixture from claude's shipped descriptor -- which this PR had just added the axis to. It now clones and deletes the key, following the shippedDescriptorWithout pattern, and asserts both that validation passes and that the resolver still picks the right winner from the default. The second half is what makes it prove anything. resolveTriggerSurface silently dropped an unrecognized scope while every sibling in this epic throws. Two phases of one epic should not disagree about whether an invalid scope is an error, so it now rejects through the same shared validator; an empty scope list still returns empty rather than throwing. The ADR amendment had been spliced into the middle of the References list, orphaning its last bullet. Moved to the top, after the header block, which is where ADR-3660 and ADR-1016 both put dated amendments. No lint checks markdown structure, so this was green while malformed. * fix(#2871): single-source the command filename composition too The earlier fix shared the namespacedByDir boolean but left the filename composition around it written twice -- once in _copyStaged as what actually gets written, once in resolveTriggerSurface as what gets predicted. The predictor could go stale silently. One exported helper now composes it for both. The entry.name asymmetry that looked like it would block extraction does not: entry.name is filtered to end in .md and stem is entry.name minus those three characters, so the two branches are the same string by construction. Divergence proven to fail: injecting a marker into the helper broke the trigger-surface suite; reverting restored 25/25. The four sibling layout suites hold at 227 unchanged. * docs(#2871): correct the ADR timing notes that this phase makes stale The Amended by back-links on ADR-3660 and ADR-1016 were written in Phase 0, when the widenings they describe had not shipped. Each carried a forward-looking clause -- "the module changes at Phase 2, not before, until then this module resolves placement only" -- which becomes false the moment this PR merges. ADR-2866's own Amends header and its reciprocal-notes section carried the same tense. All four now describe what shipped. This is a tense and status correction on Accepted ADRs, not a change to any decision. Worth stating because it is the failure mode this epic keeps meeting: gen-adr-index.cjs tracks only Supersedes and Subsumes, so nothing in CI would have caught either the missing back-link in Phase 0 or these stale clauses now. They stay correct only because someone checks. * chore(#2871): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
653f95e39f |
chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias (#3272)
* test(#2801): failing-first suite for the hostBehaviors.reviewerCli alias removal Inverts the Phase 5a rows that assert the derived legacy alias still contributes a reviewer slug, and adds the removal-warning coverage the alias's exit needs (ADR-2782 D9). RED against unmodified production code, by design: the six shipped manifests still declare the key and collectReviewerWarnings emits nothing for hostBehaviors. Refs #2801 * chore(#2801): remove the hostBehaviors.reviewerCli deprecated alias ADR-2782 D9, Phase 7 — the final phase of epic #2782. The derived legacy alias survived one release (Phase 5a shipped in 1.9.0; 1.9.1 and 1.10.0 have since gone out), so it goes. A declared reviewer body is now the only route onto the reviewer roster. - deriveReviewerSlugs no longer reads runtime.hostBehaviors.reviewerCli - the key is stripped from the six manifests that carried it; each already declares a reviewer body whose slug equals its capability id, so the derived roster is unchanged at the same twelve slugs - collectReviewerWarnings emits a presence-based, non-fatal removal notice for any manifest still declaring the key, reaching both the build-time registry generation and the third-party overlay load path. The check runs before the reviewer-body early-return, because the manifest it exists for is the alias-only one that has no body. - hostBehaviors stays an open, unvalidated bag for its other 59 keys; this adds one keyed removal notice, not general validation Refs #2801 * refactor(#2801): give the reviewer-warning channel a typed IR Review finding: the new tests asserted with String#includes() on the warning prose, which CONTRIBUTING.md's 'Prohibited: Raw Text Matching on Test Outputs' bans in favor of a typed intermediate representation. Adds the IR beside the renderer rather than replacing it, which is the shape that section prescribes and bin/verify-reapply-patches.cjs already models: - REVIEWER_WARNING, a frozen code enum - REMOVED_REVIEWER_CLI_FIELD, so the emitting site and its test share one symbol instead of duplicating a literal - collectReviewerWarningRecords(cap), returning typed records collectReviewerWarnings(cap) keeps its exact string[] contract as a thin map over the records, so both production consumers are untouched. Every section-K row now asserts on record.code/field/capId and none on the rendered message. Locks the code surface, asserts the renderer stays one-to-one with the records, and migrates the pre-existing Phase 2 test on the same channel off prose matching. Refs #2801 * test(#2801): invert the section F alias fall-through regression row Caught by the remote runner: 2 unique failures on both Node lanes out of 31,692. tests/reviewer-lane-declarations.test.cjs section F — Phase 5a's isolated-security-review regressions — asserted that a blank reviewer.slug falls through to the hostBehaviors.reviewerCli alias rather than dropping the lane. That is the direct inverse of this phase's contract. The original rationale held only while the alias existed. With it gone there is nothing to fall through to: a blank body is not a declaration, and a declaration is the only route onto the roster. Inverted rather than deleted — the row carries the adversarial-review provenance for the slug trim, and removing a security regression guard to make a change pass is backwards. The duplicate row added earlier in section C is dropped instead; section F is its canonical home. Also corrects two count strings Phase 5b left at eleven while asserting twelve, which would misreport on failure. Refs #2801 * docs(#2801): give the removed reviewerCli flag a migration path The Reference edit alone satisfied CI — a file under docs/ moved, so lint-docs-required.cjs was green — while the task-oriented quadrant said nothing about the removal. A maintainer whose lane had just gone silent would have found the field documented as removed and no page telling them what to do about it. Adds a migration section to the how-to: the symptom, the verbatim warning they will see, the before/after manifest, and the note to keep the reviewer slug equal to the capability id so existing review.default_reviewers entries and --<slug> flags survive. Refs #2801 * chore(#2801): backfill changeset pr number to 3272 * feat(#2801): close the runtime.hostBehaviors vocabulary ADR-1016 closes twelve descriptor axes and rejects an open escape hatch in the descriptor. It never mentioned runtime.hostBehaviors, and that silence was read as permission: 59 keys across 18 manifests, 39 of them set by a single capability, validated by nothing. The reference docs went further and attributed the open seam to ADR-1016, which does not mention the field at all. KNOWN_HOST_BEHAVIORS enumerates the vocabulary. An undeclared key yields a non-fatal UNKNOWN_HOST_BEHAVIOR record on the same D4.3 channel as the alias removal notice, reaching both build-time generation and overlay install. Warning, never error, for the reason this phase exists: an error would hard-break an out-of-tree descriptor carrying a bespoke key with no deprecation window, which is what reviewerCli was given a release to avoid. Escalation is a separate decision. reviewerCli is excluded from the unknown-key sweep so it keeps its own notice with the migration pointer rather than drawing two records. A parity test binds the vocabulary to the shipped manifests in both directions, and a second asserts no shipped capability draws a notice, so the closure is provably inert in-tree. Records the decision and the miscitation as an ADR-1016 amendment. Refs #2801 * fix(#2801): bound and sanitize the unknown-key diagnostics Two findings from an isolated adversarial review of the closure commit, both proven by execution rather than asserted. MAJOR, introduced by the closure: the new Object.keys(hostBehaviors) sweep had no ceiling. An installed third-party manifest is bounded only by MANIFEST_MAX_BYTES, and an 8.69MB manifest with 800,000 keys produced 800,000 records and ~139MB of message text, retained for the registry's lifetime in OverlayMeta.diagnostics. Now capped at ten records plus a summary carrying omittedCount, mirroring capability-loader's existing slice(0,3) idiom. The same manifest now yields 11 records and 1748 chars. MINOR, newly reachable: manifest-supplied key names were interpolated raw. Unlike cap.id, which validateCapability gates on KEBAB_RE before these diagnostics run, hostBehaviors keys have no grammar check anywhere, so ANSI escapes and CRLF reached stderr and OverlayMeta.warnings intact. New describeKey replaces C0/C1 controls and clips at 80 chars. The file already had describeValue for this and applied it only to values. Both fixes land on the pre-existing reviewer.* sweep too — it carried the identical pair, and fixing only the new copy would leave the same defect one screen from its own fix. Refs #2801 --------- Co-authored-by: sim <sim@local> |
||
|
|
88f6d9bd1b |
fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu * fix: preserve installer executable mode * chore: add changeset for PR #2812 * test(#2644): acknowledge Cursor emission changes * test(#2644): drop spent emitted drift acknowledgments * fix(#2644): remove retired Cursor command converter --------- Co-authored-by: clezcoding <clezcoding@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ce38d44811 |
fix(#2777): remove stale codex local home metadata (#2831)
* fix(#2777): remove stale codex local home metadata * chore(#2777): add changeset for codex local layout metadata --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
0408276791 |
chore(#2797): federate reviewer config keys off the central schema (#2841)
* chore(#2797): federate reviewer config keys off the central schema Phase 4 of epic #2782 (ADR-2782 D9, config half). Runs AFTER 5a per the ADR's swap amendment: a federated config slice lives inside a capabilities/<id>/capability.json, and three of the five key families had no capability directory until 5a created them. Four key families move to the lanes that use them; the central-schema removal and the federated addition land in this one commit because the exclusivity invariant fails the build on a key present in both. review.max_prompt_tokens, review.default_reviewers and review.reviewer_instances describe policy ACROSS lanes and stay central. Two things the issue did not name, both found while building it: 1. THE EXCLUSIVITY GATE WAS BLIND TO PATTERNS. It compared federated keys against manifest.validKeys only, and two of the four families (review.models.<slug>, review.max_prompt_tokens_per_reviewer.<slug>) were pattern-backed. That is not cosmetic: isCentralConfigKey consults those patterns and mergeFederatedConfig skips every key for which it returns true, so declaring a slice while the pattern survived would have shipped an INERT slice behind a green gate — the exact half-migrated shape the invariant exists to prevent. The gate now loads the patterns from the same manifest the runtime reads. 2. AN UNSET PER-LANE BUDGET NOW RESOLVES TO 0, NOT NOT-FOUND, because a federated key always resolves to its declared default. The three fallback guards in review.md checked only empty-or-"null", so a user who set the GLOBAL review.max_prompt_tokens would have silently lost trimming on the HTTP lanes. The guards now treat 0 as unset. D9 says review.models.<slug> is owned by "the lane whose slug it names". That is false for one lane: the shipped key is review.models.agy while the slug is antigravity. Ownership follows the lane; the key name is preserved, because renaming would break every config that sets it. Existing tests updated rather than left asserting the old world: config-get on a cleared federated key yields empty instead of not-found (what #2046 actually protects — never persisting the literal "null" — is unchanged and still asserted); the config-schema dynamic pattern representative moves to reviewer_instances; the prototype-pollution guard case moves to a surviving dynamic prefix so alert #26 keeps its coverage, with a new case asserting the old key is now rejected earlier; and Phase 2's harvest-widening inertness assertion becomes an ownership assertion, since Phase 4 is what consumes it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): use a -1 sentinel so an explicit per-lane budget of 0 survives A federated config key always resolves to its declared default, so an unset per-lane prompt budget needed a value the workflow could treat as 'not configured'. The first cut used 0 — which is wrong: 0 is already a LEGITIMATE per-lane budget meaning 'do not trim this lane' (the early-return guard in prepare_trimmed_prompt_for_reviewer). Treating it as unset would have silently switched a user who deliberately disabled trimming for one lane onto the global budget. The sentinel is now -1, which is not a valid token budget, so all three states stay distinguishable: unset falls back to global, an explicit 0 disables trimming for that lane, and an explicit N is used. Locked by three CLI round-trip tests. Surfaced by the isolated security reviewer before it crashed mid-run; verified independently against the shipped trim guard rather than taken on trust. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): update central-registration assertions and stay under the review.md cap The remote runner caught both; my local sweep missed the files. 1. tests/plan-review-convergence.test.cjs asserted the three local-server host keys are in VALID_CONFIG_KEYS. They are federated to their lane capabilities now, and the exclusivity invariant forbids a key living in both places. What #2306-local actually protects is that config-set ACCEPTS them, so that is what is asserted — via isValidConfigKey, the predicate config-set itself uses, which spans central and federated. A second assertion pins federated ownership, so a silent reversion back to the central schema fails too. 2. review.md exceeded the LARGE tier hard cap (62583 > 61440). That cap is a red line, not a budget to raise. The three per-lane budget guard comments were near-identical; condensed to one terse line each. 61371 bytes, 69 to spare. Real extraction to workflows/review/modes/ is Phase 5b/6 work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): fail closed on a broken config-schema manifest; reconcile stale docs Isolated security review findings. MAJOR — loadCentralConfigPatterns failed OPEN. It swallowed a JSON parse error and returned [], while its sibling loadCentralConfigKeys, reading the SAME file, writes to stderr and throws ExitError(1) on that identical failure class. Fail-open here defeats the gate this function exists to feed: with zero patterns, validateCrossCapability's pattern-collision check silently passes and an inert federated slice ships green. It was masked in the one production call site only because loadCentralConfigKeys runs first against the same path — a coincidence of ordering, not a guarantee, and this function is exported and called standalone. The two now share a contract: ENOENT is the legitimate absent case, anything else throws loudly. A single unparseable PATTERN is still skipped, which degrades to "checked less" rather than blocking every build. The branch had zero coverage; it now has two tests (malformed JSON, EISDIR). MINOR — docs/CONFIGURATION.md still listed review.models.qwen and review.models.cursor as settable, ~770 lines below this PR's own new Ownership section. Those lanes take no model flag, so they declare no model key and config-set now rejects them. Rows removed; the missing review.models.agy row added; the per-reviewer budget row corrected to name only the lanes that own a budget key, and to document that a per-lane 0 disables trimming for that lane. Also fixes a shadowed "raw" binding introduced by the fail-closed change, which made the generator unrequirable — caught immediately by its own --check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2797): backfill changeset pr number to 2841 * fix(#2452): make the base-ref mutation test hermetic against leaked GIT_* env tests/mutation-workflow-base-ref.test.cjs fails on PR branches while next stays green, and it is currently blocking at least three unrelated PRs (#2841, #2832, #2827) with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees The existing loop comment attributes this to `git add .` rehashing O(n^2) blobs "before the object write had landed" and works around it by staging one path per iteration. That is not the cause: sequential execFileSync calls cannot race each other's object writes, and the failure persisted after that change — it simply moved to a lower commit index. The cause is that the git() helper inherited the runner's environment. A leaked GIT_INDEX_FILE makes `git add` write into a DIFFERENT repository's index; GIT_OBJECT_DIRECTORY / GIT_ALTERNATE_OBJECT_DIRECTORIES send the blob to another object store; GIT_DIR / GIT_WORK_TREE redirect the whole operation. In every case `git commit` then cannot resolve a blob it just staged, which is precisely the error above. Verified by negative control: with GIT_DIR exported, this test fails on the unfixed helper (the git commands operate on the wrong repository entirely); with the helper stripping GIT_* it passes. The single-path staging is kept — it is genuinely less work — but it is no longer load bearing. Found while shipping #2797. Fixed in place rather than deferred: it is a defect surfaced during the work, and it is blocking other contributors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2452): build the base-advance commits empty, removing the lost-object class The base-ref guard has been failing in CI with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees It is currently red on at least three unrelated PRs (#2841, #2832, #2827) while next stays green. Two theories have now been tried and neither held. #1881 blamed `git add .` rehashing O(n^2) blobs and switched to staging one path per iteration; the failure moved from commit 32 to commit 25 and carried on. The preceding commit here made the git helper hermetic against leaked GIT_* environment — that IS a real vulnerability (with GIT_DIR exported the helper operates on the wrong repository entirely, proven by negative control) but it produces a different error than CI reports, so it is not demonstrably the cause either. Neither trigger reproduces off-CI, so this stops guessing at the trigger and removes the failure CLASS instead. The loop needs base-branch DEPTH and nothing else: no assertion reads these commits' contents, and base-side files cannot appear in `origin/base...HEAD` regardless. `--allow-empty` writes no blob and no tree, so there is no object for the index to reference and lose. It is also far less work than 60 write+hash+index cycles. The guard still proves its mechanism: the test asserts that a --depth=1 base fetch FAILS and a full fetch resolves, so a broken topology would surface immediately rather than passing vacuously. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Test <test@example.com> |
||
|
|
46e84d5e39 |
chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat (#2823)
* chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat Phase 2 of epic #2782 under ADR-2782. Delivers D1, D2, D3, D7, D8 and the four Phase-1 vocabulary amendments (A1-A4). - VALID_ROLES gains "reviewer"; the reviewer body is admissible on role:runtime (a host that is also a reviewer keeps one manifest) and on the new role:reviewer (a lane that is not an install target). A reviewer body on role:feature is an error: declaring one is an assertion of lane-ness. - validateReviewerBody + validateLaneProbe + validateLaneInvoke: nine closed enums, a transport discriminator selecting mutually-exclusive invoke sub-shapes, bounded probes (D7), and outputArg required-iff outputChannel is file-arg and forbidden otherwise. - Absent-safe (D4.1): only `undefined` is absent. null/{}/[]/false/0 are malformed assertions and error. 39 of 39 shipped capabilities depend on this. - collectReviewerWarnings: an unknown field inside the body warns, never errors, so a forward-built manifest degrades visibly instead of failing the build. - D8 uniqueness (slug / flags / reviewsSection) lives in validateCrossCapability, so it is enforced at build time over first-party AND at load time over the merged first-party union overlay set, with first-party-wins falling out of the loader's existing ordering rather than a new provenance check. - Config harvest widened past the role==="feature" branch in both the generator and the ownership loop. The often-cited cause of the stranded reviewer config keys -- the runtime body forbidding feature-only fields -- is not the mechanism: `config` is not in FEATURE_FIELDS_FORBIDDEN_ON_RUNTIME. The cause is two harvest sites that never read it. Verified inert: no shipped capability declares config on a non-feature role, and the generated registry is unchanged. Three ADR corrections are folded in (Phase 1 set the precedent of amending in-phase): the misattributed config-stranding cause, D3's inverted profile-membership claim, and the specified capability folder names for lm_studio / llama_cpp, which would have failed the id kebab-case invariant. Closes #2795 * chore(#2795): collapse nine enum checks into one validateEnumField helper Standards-axis review findings, both applied: - Duplicated Code: the enum-membership + enumerate-the-members error shape repeated near-verbatim at nine call sites. Routing them through one helper makes "the error names its valid members" structural rather than a convention repeated nine times, where it would drift. That property is load-bearing until Phase 6 ships the prose reference, because these errors are currently the only documentation of the vocabulary. - Speculative Generality: the isReservedName() pre-check on every enum field was inert. A VALID_* set never contains __proto__/constructor/prototype, so membership alone already rejects them, and "must be one of: ..." is more actionable than "is a reserved name". The literal guards remain where they do real work -- the key-derived write sites in the registry generator and the claim() accumulator. The reserved-name test now asserts all three reserved names are rejected via enum membership, rather than one name via a branch that no longer exists. * fix(#2795): align lane slug grammar with Phase 1 and wire the load-time diagnostic channel Spec-axis review findings, both applied. (1) The slug grammar had diverged from Phase 1's core descriptor. Phase 1 exports LANE_SLUG_RE = /^[a-z0-9][a-z0-9_-]*$/ (leading digit permitted); the manifest validator required a leading LETTER. A slug the core descriptor accepts -- a model-named lane such as 4o-mini -- would have been rejected by the manifest validator, which is exactly the translation layer ADR-2782 exists to delete. It was inert only because all eleven shipped slugs begin with a letter, so nothing else would have caught it until a third party shipped such a lane. The grammar cannot be reduced to one definition: Phase 1's module compiles to gitignored build output, and capability-validator.cjs is a committed plain .cjs that must load on a fresh worktree before build:lib has ever run. That makes this the repo's DEFECT.GENERATIVE-FIX class, so the duplication now carries a parity assertion -- laneSlugGrammarMatchesPhase1Descriptor -- which compares both the source grammar and the accept/reject verdict for a shared input set, and fails if the two ever drift again. (2) collectReviewerWarnings had exactly one caller: the build-time generator, which only ever sees first-party in-repo manifests. The real third-party overlay loader never called it and ValidatorModule did not declare it, so ADR-2782 D4.3 -- an unknown field inside a reviewer body is ignored WITH A WARNING -- surfaced nowhere at runtime, which is precisely the case D4.3 exists for. loadRegistry now collects those diagnostics on the accept path, behind a typeof-guard (an older built validator without the function still loads) and a try/catch (ADR-1244 D2's never-crash contract outranks a diagnostic). They land in a NEW OverlayMeta.diagnostics field rather than OverlayMeta.warnings, because warnings records capabilities that were SKIPPED and a consumer treating every entry as inactive would mislabel a working lane. Covered end-to-end by overlayLaneWithUnknownFieldIsAcceptedAndDiagnosed, which drives a real global-scope overlay through loadRegistry and asserts the lane is accepted, produces no skip warning, and yields a diagnostic naming the field. * fix(#2795): make the reviewer validators honour their documented totality contract Isolated adversarial review finding (MAJOR), reproduced by execution. validateReviewerBody documents itself as "TOTAL: returns an array of error strings for ANY input and never throws", and the overlay loader contracts every validator to RETURN errors -- #1461 OVL-1 records a validator that THREW and would have crashed every consumer of loadRegistry. The contract was false at ten sites: JSON.stringify throws on a BigInt and on a circular structure, and every enum/scalar rejection path interpolated the rejected value into its own rejection message. Reading the value could throw too, before any message was built, via a throwing getter or a Proxy get/ownKeys trap. Not reachable through a capability.json today -- every ingestion path is a plain JSON.parse of file text, which cannot express any of those shapes. Fixed anyway: the contract is stated on an EXPORTED function, and a caller must not have to re-derive today's reachability analysis to know whether it holds. Two layers, because serialization safety alone is insufficient: - describeValue() renders any value without throwing, so messages stay useful (a BigInt now reads "got: 10n" rather than degrading to a generic fallback). - A structural try/catch around validateReviewerBody and collectReviewerWarnings makes the guarantee absolute rather than argued, covering read-time throws that fire before any message exists. The same review found the property test guarding this contract was FALSE CONFIDENCE, which is the more important half. fc.anything() at default constraints emits no BigInt, no circular reference, no getter and no Proxy -- 20,000 sampled draws produced zero of each -- so the test was named for a contract its generator could not reach. Even withBigInt is insufficient under whole-value fuzzing, because the defect needs an exotic value in a specifically NAMED field and random key names never land on one. The property is now field-targeted across all twelve reviewer fields, and a companion test enumerates the shapes fast-check cannot generate at all (BigInt, circular, throwing getter, symbol, function, null-prototype) across scalar positions, array-element positions, and read-time traps. Verified red-before-green: with the fix reverted both property tests fail; with it restored all 119 pass. * chore(#2795): backfill changeset pr number to 2823 |
||
|
|
6ad30f74b6 |
feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler. harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run. Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard. Closes #2627 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4a66d62d10 |
feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver (#2625)
* feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver
Phase 2 of the negotiated executor-isolation feature (ADR-1239 Codex-binding amendment). Two building blocks for `dispatch.isolation: orchestrator-worktree` hosts, both unconsumed — no scheduler wires them yet (that is Phase 3), so no runtime behavior changes.
worktree create verb (planWorktreeCreate / executeWorktreeCreatePlan / cmdWorktreeCreate in worktree-safety.cts, routed via routeWorktree in gsd-tools.cjs): validates the wave base, creates a bounded branch+worktree, records it in the run manifest reusing record-agent 4-field entry shape, returns the executor working directory. Bounded git (10s timeout, degrade-not-throw); all manifest read/parse/validate/dedupe precedes the single git side effect (no unmanifested-orphan on a bad manifest); timeout-only best-effort partial rollback (a clean collision-exit never removes a live peer worktree); fail-closed on bad base, unsafe leading-dash / .. inputs, and malformed/mis-shaped manifest.
resolveOrchestratorExec (host-integration.cts): pure descriptor->argv resolver reading the new runtime.orchestratorExec descriptor field (codex/opencode/kimi/kimi-code), fail-closed on missing/invalid shape. Validator (capability-validator.cjs) + a parity guard asserting every orchestrator-worktree host declares a resolvable orchestratorExec.
Adding the create route edits the installed gsd-core/bin/gsd-tools.cjs, so the golden-install-parity fixtures for all 19 runtimes are regenerated (npm run gen:golden) — the only changed hash is gsd-tools.cjs. CONTEXT.md glossary updated; capability-registry regenerated. Behavioral tests (worktree-safety + host-integration) incl. a fast-check property test and the parity sweep.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: rebuild tracked state-transition.cjs to match #2400 source
The tracked compiled artifact drifted from src/state-transition.cts: #2400 (commit
|
||
|
|
ec7978c0b4 | feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) | ||
|
|
d579daa3ed |
docs(#2505): Phase 6 — migration guide + built-in-only subagent-toolkit enum (#2538)
* docs(#2512): Phase 6 — migration guide + built-in-only subagent-toolkit enum * fix #2512: update CONTRACT-PIN for built-in-only subagentToolkit value * docs(changeset): backfill PR #2538 for Phase 6 (#2512) |
||
|
|
c2a305c44d |
feat(#2505): Phase 2 — kimi-code Agent Skills install layout (#2520)
* feat(#2454): PR 2 — kimi-code Agent Skills converter + install layout PR 1 registered the kimi-code EoS descriptor with empty artifactLayout (SKIP_INSTALL_CONTRACT excluded it from the end-to-end install test). PR 2 fills in the install surface: - src/runtime-artifact-conversion.cts: new convertClaudeCommandToKimiCodeSkill function. Today it delegates to convertClaudeCommandToKimiSkill (Python kimi-cli) because Kimi Code uses the same Agent Skills format + /skill: invocation per official docs. The distinct function name lets a future divergence land cleanly if Kimi Code's skill format evolves independently. - gsd-core/bin/lib/capability-validator.cjs: add to ALLOWED_SKILLS_CONVERTERS. - capabilities/kimi-code/capability.json: artifactLayout.global now declares the skills kind with converter='convertClaudeCommandToKimiCodeSkill' + home='.kimi-code' (auto-discovered at ~/.kimi-code/skills/ per Kimi Code docs: merge_all_available_skills = true default). - tests/installer-migration-install.integration.test.cjs: REMOVE the SKIP_INSTALL_CONTRACT exclusion — kimi-code now has a full install surface. - Regenerated capability-registry + capability-matrix + golden install parity + install tree fixtures for kimi-code. * fix(#2454): wire kimi-code converter into SKILLS_CONVERTER_REGISTRY + count bump - src/install-engine.cts: add convertClaudeCommandToKimiCodeSkill to SKILLS_CONVERTER_REGISTRY so the layout-driven skills install path can dispatch off the descriptor's converter string. - tests/capability-registry.test.cjs: bump VALID_CONVERTER_NAMES count 26 → 27 (added convertClaudeCommandToKimiCodeSkill). * fix(#2454): remove home override from kimi-code skills (inherit configDir) The home:'.kimi-code' override made the install plan resolve skills dest to ~/.kimi-code/skills instead of <configDir>/skills, causing the test's temp configDir to miss the install. Removing it lets skills inherit configDir like most runtimes. * fix(#2454): kimi-code install contract surface is flat-skills (no agents) Kimi Code has NO custom named subagents (per official docs: 3 built-in coder/explore/plan only). The kimi-skills-agents surface expects agents/ gsd.yaml + subagents/*.yaml which kimi-code does not produce. Changed to flat-skills which only checks for skills/gsd-* dirs. * docs(changeset): Phase 2 kimi-code install layout Added (#2509) * docs(changeset): backfill PR #2520 for Phase 2 (#2509) |
||
|
|
09b535ac00 |
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.
Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.
Every per-host value is documentation-sourced, never inferred:
- claude argv -- verified via 'claude --help' (--effort <level>)
- opencode argv -- verified via 'opencode run --help' (--variant)
- codex argv -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
flag (config.toml key only), so the global -c override is the
only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
fails closed rather than inheriting a profile baseline
No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by
|
||
|
|
a0fafedfa0 |
feat(#2103): drive VS Code through the Embeddable Orchestration System (ADR-1239)
VS Code is a net-new EoS runtime that — unlike every prior migration — is NOT CLI-installed (Marketplace/VSIX extension). It has zero runtime==='vscode' branches in bin/install.js and stays that way (regression-guarded); it is driven entirely through the negotiated imperative Host-Integration adapter. Registry + validator (the hard part): - capabilities/vscode/capability.json (role:runtime): full hostIntegration block (imperative / palette / active vscode.lm model / engine hook bus / sandboxed-storage / mcp transport / sandboxed-web runtime; dispatch nested, maxDepth 5 per VS Code's documented subagent depth). - capability-validator.cjs extended so a role:runtime capability can legitimately declare "extension-distributed, no config directory": new configHome.kind:'none' + installSurface:'none' (+ GATE-A pairing + the parity maps), with localConfigDir and configHome.name made conditional on kind!=='none'. All 18 runtimes still validate; getDirName returns a distinct sentinel (not '.claude') for a no-config runtime. - The add-a-registry-runtime tax: NON_INSTALLABLE_RUNTIMES exemption in the runtime-flags drift guard, vscode added to global-config-home SPECIAL_CASED, EXPECTED_PROFILES.vscode='ide', and the config-adapter/derivation/pin-count guards updated. No golden-install fixture, model-catalog, or CONFIGURATION rows (vscode never enters allRuntimes). Dispatch + extension surface: - Fixed vscode/extension.js's createHub()-no-args bug (every dispatch was UnknownCommand, masked by a vacuous reachability test) — now reuses the shared dispatchGsdCommand subprocess-shim (Node/desktop); the reachability test is tightened to assert real dispatch. - Promoted the #1933 host binding to a shipped vscode/host-binding.js; activate() now composes the model/hookBus/stateIO seams through it. Corrected the model seam to VS Code's real API (vscode.lm.selectChatModels() -> model.sendRequest(); vscode.lm.sendRequest does not exist) so the binding actually composes on real desktop VS Code instead of throwing. - New vscode/browser.js Web Extension entry with ZERO Node APIs (the engine's config/capability loading is Node-bound, so the web entry registers the surface and directs full dispatch to the native MCP server — honestly documented). - UPGRADE 1: GSD skills as native Language Model Tools (contributes.languageModelTools + vscode.lm.registerTool), invoke() dispatching through the hub. - UPGRADE 2: native subagent dispatch wired onto #runSubagent / chat.subagents.allowInvocationsFromSubagents (fail-soft on API availability, maxDepth 5 enforced). - vscode/package.json: browser entry, engines.vscode ^1.105, chatParticipants + languageModelTools contributions; fixed a stale activationPoints->activationEvents manifest key. Added "vscode" to the package files array. Docs (## vscode matrix section) + changeset (Added). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bd613566cb |
feat(#2100): drive Windsurf through the EoS descriptor + wire Cascade's blocking hook bus (ADR-1239)
Fold all 10 residual isWindsurf branches in bin/install.js onto descriptor-driven hostBehaviors (byte-parity — no fold changes any install output): - 2 dead destructures dropped (uninstall, finishInstall); the dead `else if (isWindsurf)` legacy agent-loop arm removed (windsurf ∈ _DESCRIPTOR_AGENTS_RUNTIMES → unreachable). - skipSharedHooksInstall:true folds the two `!isWindsurf` shared-hooks exclusions. - legacyDevinSkillsCleanup:true folds the `.devin`→`.windsurf` one-time cleanup gate. - installsCommandBodiesForWorkflowDelegation:true folds the #1629 command-body copy (workflow-delegation target — load-bearing; local-install verified intact). - verificationStyle:"windsurf-workflows" folds the workflow-count report. - Corrected stale _LEGACY_SCAN_SUBDIR_NAMES + hooks-json manifest comments (cursor + windsurf). Zero live runtime==='windsurf'/isWindsurf branches remain across bin/install.js, install-engine.cts, surface.cts, runtime-artifact-conversion.cts (AC2 guard scans all four). UPGRADE (Cascade hook bus): wire GSD's write/command safety guards into Windsurf's native hook bus. New hooksSurface 'windsurf-hooks-json' (VALID_HOOKS_SURFACES 7→8, GATE A profile-marker-only allowlist, the HooksSurface union) + writeWindsurfHooksJson (Cursor-templated, Cascade's flat {hooks:{<event>:[{command}]}} shape) writing .windsurf/hooks.json with two BLOCKING pre-hooks: - pre_write_code → gsd-windsurf-pre-write.js: blocks writes to a file outside the active git worktree / into .git internals. - pre_run_command → gsd-windsurf-pre-command.js: conservative destructive-command deny-list (rm -rf of root/home incl. sudo/env/path-prefixed forms; fork bombs; force-push refspec forms — HEAD:main, +main, --force/-f — to main/master/next). Both use Cascade's protocol (stdin JSON, exit 2 + stderr to block, exit 0 to allow, fail-open on error/timeout). Tokenize-based classifier (no catastrophic-backtracking regex; 4096-char cap) with the fail-closed false-positives fixed post-review. The 4 advisory GSD guards + pre_mcp_tool_use + 5 post_* logging events are deliberately NOT wired: Cascade has no context-injection channel for advisory hooks and GSD has no MCP guard — porting them would be non-functional padding (documented; codebuddy #2098 / copilot #2099 faithful-subset precedent). extendedHookEvents stays []. Golden: the 2 guard scripts ship in the shared hook bundle (HOOKS_TO_COPY + the shared managed-hooks-registry), exactly like cursor's 6 gsd-cursor-*.js scripts — so the 8 shared-bundle runtimes' fixtures gain the 2 inert windsurf scripts + the registry hash (functionally inert for non-windsurf; the established cursor pattern). No install-output change beyond that (the folds are byte-parity; skip-bundle runtimes untouched). New scripts registered in managed-hooks-registry + build-hooks + INVENTORY. Tests: declarative-reference- windsurf (adapter/axes/fail-closed + AC2 guard) + windsurf-hooks-bridge (live exit-2 blocking + allow/fail-open + ReDoS-bound + writer/reconcile/remove idempotency); VALID_HOOKS_SURFACES pin updated to 8. Matrix hookBus delta + changeset (Changed). capability-registry regenerated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5695522d5f |
feat(#2096): migrate Antigravity onto EoS declarative adapter + permission-writer + MCP companion (ADR-1239)
Fold all antigravity literal branches into descriptor-driven reads: getConfigDirFromHome (→ configHome.kind 'dot-home-nested'), projectLocalHookPrefix (→ hostBehaviors.hookPathStyle 'raw'), applyAgentPathRewrites (→ noPathRewrite), getProjectInstructionFile (→ projectInstructionFile 'GEMINI.md'); removed the dead inline convertClaudeAgentToAntigravityAgent branch + dead isAntigravity destructures (antigravity is already on the descriptor-agents path). subagentToolkit flipped undocumented→full (Context7: antigravity.google/docs/cli/features); namedDispatch/nested/maxDepth/backgroundDispatch stay undocumented. Byte-identical golden parity for all 16 runtimes. UPGRADE 1 (permission-writer): permissionWriter 'antigravity' + configureAntigravityPermissions merges a scoped permissions.allow block (GSD's own tree + hooks) into Antigravity's settings.json — non-destructive, idempotent, symmetric uninstall. Added to VALID_PERMISSION_WRITERS + the FinishPermissionWriter union. UPGRADE 2 (MCP companion): configureAntigravityMcpConfig writes mcp_config.json registering the gsd-core companion MCP server (Gemini-successor mcpServers schema, best-effort — raw schema unpublished). Both writers dispatch from finishInstall. settings.json is golden-excluded (HOOK_CONFIG_FILES); mcp_config.json (portable, no absolute paths) is golden-tracked → only antigravity.json changes. Tests: declarative-reference-antigravity extended (source-grep guard across 4 modules, fail-closed for the 4 undocumented sub-axes, validator acceptance) + antigravity-upgrades (permission-writer + mcp_config live-install, idempotency, user-preservation). Matrix + ADR-1016 + capability-manifest + CONTEXT.md + connect-gsd-mcp-server docs updated; changeset (Changed). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ab04916682 |
feat(#2095): migrate Kimi CLI onto EoS imperative adapter + native hook-bus + background dispatch (ADR-1239)
Fold all runtime==='kimi'/isKimi logic branches into descriptor-driven hostBehaviors (localInstallDeferred, verificationStyle, agentManifestStyle, reapplyCommand, doneBannerStyle) + add 'kimi' to _DESCRIPTOR_AGENTS_RUNTIMES. Kimi's skills/kimi-agents dispatch was already descriptor-driven (converter-by- name + kimi-agents kind). Zero isKimi/runtime==='kimi' branches remain. UPGRADE 1 (native hook bus): new hooksSurface 'kimi-hooks-toml' + a marker- delimited config.toml [[hooks]] emitter (buildKimiHooksTomlBlock/writeKimiHooksToml in runtime-hooks-surface.cts; resolveKimiHooksTomlDir in runtime-homes.cts). GSD's lifecycle hooks now wire into Kimi's native ~/.kimi/config.toml (Context7- confirmed path) at SessionStart/PreToolUse/Stop/PreCompact/SubagentStart/ SubagentStop — kimi becomes a hooks/ consumer (the 3 && !isKimi exclusion guards removed). config.toml holds absolute install paths so it's golden-excluded via an exact relative-path (.kimi/config.toml), not a basename (which would blind Codex's config.toml). New hooksSurface value added to the closed enum in capability-validator + runtime-config-adapter-registry. UPGRADE 2 (background dispatch): flip dispatch.backgroundDispatch true (Kimi's Agent tool takes run_in_background; root agent already gets the Agent tool), so negotiation no longer flattens dispatch. subagentToolkit stays 'undocumented' per AC (coder/explore/plan have distinct tool policies). MCP transport explicitly deferred (no installer-driven MCP for any runtime). Golden: only kimi.json changes (hooks/ scripts now installed); all 15 others + claude-local byte-identical (kilo/zcode keep their own exclusions). Tests: kimi-imperative-reference (adapter/axes/fail-closed/hostBehaviors + source-grep guard) + kimi-upgrades (config.toml [[hooks]] SessionStart + marker idempotency + backgroundDispatch negotiation). CONTEXT.md glossary + matrix + how-to updated; changeset (Added). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f014ec83bd |
feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads: finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks copy), and a skills converter-name registry (the artifactLayout.converter field is now load-bearing, not decorative). frontmatterDialect stays the documented dispatch key for frontmatter (no descriptor field for it). Dead isKilo destructure bindings removed. Byte-identical golden parity for all 16 runtimes (opencode, which shares kilo's combined-family path, verified clean). UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin + extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus). UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented' per AC so dispatch degrades to 'degraded' by design. Model-catalog single-source edit ripples the shared model-catalog.json hash into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale- bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/ codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error. Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/ hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades (plugin parity+load, model-override converter, agents dispatch surface, MCP doc). Matrix + how-to + config docs updated; changeset added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c6ce110efa |
feat(#2092): migrate Qwen Code onto EoS imperative adapter + native subagents + SubagentStart (ADR-1239)
Fold all runtime==='qwen'/isQwen logic branches (skill-priority frontmatter, branding/path rewrites, legacy commands/gsd cleanup, hyphen-namespace normalization, RUNTIME_CONTENT_DISPATCH, hooks-surface label) into descriptor-driven runtime.hostBehaviors on capabilities/qwen/capability.json, read via _hostBehaviors(). Shared claude/qwen/hermes legacy-migration branches in install-engine.cts folded to descriptor flags (claude+hermes descriptors updated; FALLBACK_HOST_BEHAVIORS.claude floored). Byte-identical golden parity for qwen/hermes/claude(global+local). UPGRADE 1: native .qwen/agents/*.md subagent projection — new agents artifact-layout kind + convertClaudeAgentToQwenAgent converter (name + description + tools YAML block list; color/model dropped). qwen routed onto the descriptor-driven agents path (_DESCRIPTOR_AGENTS_RUNTIMES). UPGRADE 2: SubagentStart hook wired into extendedHookEvents + the descriptor-gated hook-writer loop (activates only for qwen). Tests: qwen-imperative-reference (adapter/axes/fail-closed/hostBehaviors + no runtime==='qwen' source-grep across 4 files) + qwen-upgrades (agents file validity + SubagentStart mirrors SubagentStop, descriptor-gated). Docs matrix + how-to updated; changeset added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f152e9a069 |
feat(#2091): migrate hermes onto EoS imperative adapter + extensionEvents dialect (ADR-1239)
- Fold 7 hardcoded isHermes/runtime === 'hermes' branches in bin/install.js into descriptor-driven _hostBehaviors lookups (skillFrontmatterVersion, skillsManifestPrefix, trackCategoryDescription, writeCategoryDescription, reportSkillsCount, legacyCommandsGsdCleanup, brandingRewrites) - Add runtime.hostBehaviors block to capabilities/hermes/capability.json - Register EXTENSION_EVENT_SURFACES.hermes (13 real plugin hook events) in src/host-integration.cts — replaces the borrowed hookEvents:'claude' 6-event surface that silently never fired on Hermes - Add extensionEvents:'hermes' to the descriptor - Add 'hermes' to VALID_EXTENSION_EVENTS in capability-validator.cjs - Regenerate capability-registry.cjs - Tests: hermes-imperative-reference (negotiation, axes, fail-closed, source-guard), hermes-dispatch-upgrade (dispatch posture, degradation, fail-closed) - Changeset + docs update |
||
|
|
b51cbf96cf |
feat(#1943): extensionEvents vocabulary — separate from hookEvents (extension-system surface) (#1946)
* feat(#1943): extensionEvents vocabulary — separate from hookEvents (extension-system surface) * fix(#1943): re-export VALID_EXTENSION_EVENTS from gen-capability-registry (test import path) * fix(#1943): regenerate capability-registry.cjs + add changeset fragment * fix(#1943): import VALID_EXTENSION_EVENTS from validator, not gen-capability-registry (golden parity) |
||
|
|
a3d3c2a445 |
refactor(#1756): derive getDirName from a documented runtime.localConfigDir descriptor axis (#1757)
ADR-1239 Phase B (parent #1679). getDirName was a hand-maintained 15-branch if-chain mapping each runtime to its local content-rewrite dot-dir. Relocate those values into a documented runtime.localConfigDir descriptor field; derive getDirName from registry.runtimes[id].runtime.localConfigDir (fallback .claude). - 16 capability.json gain runtime.localConfigDir (byte-identical values) - capability-validator.cjs requires it (non-empty dot-dir); registry regenerated - docs/reference/capability-manifest.md documents the field + the three divergent values (copilot=.github, antigravity=.agents, kimi=.kimi-code) - drift-guard test: golden value map + key-set equality both ways Byte-identical install output for all 16 runtimes (golden-parity harness #1730). Closes #1756 Co-authored-by: review-bot <review-bot@gsd> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
cf2e66b39e |
feat(#1708): typed documentation-sourced #853 dispatch-flatten (ADR-1239 Phase B) (#1719)
* feat(#1708): typed documentation-sourced #853 dispatch-flatten Graduate the #853 orchestrator-backgrounding decision from a scattered RUNTIME==='codex' prose check to a typed, documentation-sourced engine decision. Adds a backgroundDispatch dispatch sub-axis (sourced per host: codex+cursor documented true, 9 documented false, 5 undocumented), shouldFlattenDispatch(dispatch) (inline UNLESS background && backgroundDispatch, fail-closed), and a gsd_run query dispatch-should-flatten the plan/execute workflows call. Cursor is newly background-eligible per its docs (inline->background) — a documentation-justified behavior change. No RUNTIME-name residue for this decision. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1708): backgroundDispatch citations in matrix + CONTEXT note Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1708): address review findings on typed dispatch-flatten Code/adversarial review: convert the manager.md/autonomous.md Compound Action preamble from hardcoded 'On Codex' to FLATTEN-based branching (the handlers already use the query; the preamble contradicted them and was wrong for cursor); make shouldFlattenDispatch null-safe + type-honest (accepts raw 'undocumented' registry values); make backgroundDispatch a required descriptor field (matching its siblings, all 16 carry it); strengthen the config.runtime behavioral test; update the bug-853 prose-pin test + comment. Security review clean; Codex confirmed no fail-open. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): backfill backgroundDispatch in role:runtime test fixtures Making backgroundDispatch a required descriptor field broke role:runtime fixtures in capability-manifest-version/capability-registry/host-integration-descriptors tests that build a dispatch object without it (caught by full gsd-test, not scoped npm test). Backfill backgroundDispatch:false into the well-formed fixtures; the deliberately-malformed 'required-field' test fixture is left malformed by design. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): update fix-1521 dispatch-gating assertion to the FLATTEN gate fix-1521 pinned the codex-specific run_in_background prose that #1708 graduated to the typed dispatch-should-flatten/FLATTEN gate. Update its assertions to verify FLATTEN=false gating (not a runtime name) + that the old RUNTIME===codex gate is gone. Caught by full gsd-test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1708): add changeset for typed dispatch-flatten Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1708): remove stray temp PR-body file Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1708): add issue ref to bug-853 allow-test-rule annotations ADR-456 requires every allow-test-rule exemption to carry a see #NNN reference; the source-text-is-the-product annotations added when migrating the prose assertions lacked it (lint-tests CI gate). Add (see #1708). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
30d4b85de5 |
feat(#1684): negotiated host-integration interface (ADR-1239 Phase A) (#1690)
* feat(#1684): add negotiated host-integration interface module ADR-1239 Phase A: a pure, additive, no-I/O module exposing PROTOCOL_VERSION, the 8-axis HOST_INTEGRATION_AXES closed vocabulary, the UNDOCUMENTED fail-closed sentinel, negotiateHostCapabilities (effective subset of host-declared and engine-known), a typed degradation ladder, and host-capability profiles. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1684): validate and document host-integration axes (16 runtimes) Extend validateRuntimeBody to validate the 8 hostIntegration axes (closed enums + undocumented sentinel + dispatch struct + reserved-key guards) and the widened runtime vocabulary; author a documentation-sourced hostIntegration block in all 16 runtime descriptors; regenerate the registry. Every per-CLI value is documented (cited) or the explicit undocumented sentinel. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add host-integration capability matrix and adr amendment New per-CLI, per-axis citation reference (value/source/evidence for all 16 CLIs); ADR-1239 Phase-A-implemented amendment; CONTEXT.md glossary seam entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1684): harden dispatch negotiation edge cases Code-review hardening: treat NaN/Infinity maxDepth as missing (fail-closed, +warning); reset nested/background when namedDispatch collapses to false (struct consistency); SAFE_DEFAULTS dispatch floor to read-only; warn on non-finite protocolVersion; symmetric undocumented warnings for dispatch fields. Pure module — no consumers; behaviour fail-closed throughout. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1684): register host-integration.cjs in lint-ignore and inventory New tsc-generated bin/lib artifact: add to the eslint ignore list (ADR-457 — lint the .cts source), regenerate docs/INVENTORY-MANIFEST.json, and add the docs/INVENTORY.md CLI-modules row. Fixes the 3 gsd-test failures (551-eslint-bin-lib-coverage x2 + inventory-manifest-sync). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add changeset fragment for host-integration interface Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1684): add how-to for sourcing a host's integration axes Diataxis how-to guide for adding/updating a host's runtime.hostIntegration axes from authoritative docs, the undocumented-sentinel rule, validation, and extending the closed vocabulary. Completes the Step-5 doc quadrants (reference + explanation + how-to). Indexed in docs/README.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bcc5a6d1ba |
fix(#1634): honor capability hook matcher and node-prefix command (#1638)
* fix(#1634): honor capability hook matcher and node-prefix command Capability hook install (applyCapabilitySharedEdits) wrote each settings.json hook entry with no `matcher`, so a tool-scoped hook fired on every tool (a fail-closed guard could then block the whole session), and emitted a bare single-quoted script path so a .js-family hook from a git/tarball source without +x failed with Permission denied on every matching call. - Pass through an optional declared `matcher` (entry-level sibling of `hooks`); absent => omitted (match-all), so existing shipped capabilities are unchanged. - Validate `matcher` in the declaration (non-empty string, no control chars). - Emit `node <quoted-path>` for .js/.cjs/.mjs hooks (mirrors first-party); .sh and others keep the bare quoted path (unchanged). Root cause: the manifest hook schema (validator rule C4) was {event, script} only with no matcher, and applyCapabilitySharedEdits never read or wrote one; the command used shellSingleQuote(absScript) with no node prefix. Regression tests fail-first on both defects (matcher dropped; bare path) and pass after the fix; #1460 command assertions updated for the node prefix. * chore(#1634): backfill changeset pr:1638 * fix(#1634): resolve lint and windows CI failures - validator: replace the control-character range regex with a char-code loop. The literal /[\x00-\x1f\x7f]/ tripped ESLint's no-control-regex rule; char codes are equally precise and lint-clean. Behavior unchanged (still rejects matchers containing ASCII control characters incl. DEL). - test: gate the executable-bit precondition on POSIX. Windows fs does not honor POSIX write modes (a 0o644 write reads back as 0o666), so the precondition is meaningless there and failed the windows-latest lane. The node-prefix assertion — the actual fix — is platform-independent and still runs everywhere. * docs(#1634): amend ADR-894 for optional lifecycle hook matcher The `role: "feature"` `hooks[]` entry now carries an optional `matcher` (settings.json tool-scoping pattern: exact/pipe/wildcard/regex). Document the field in the §2 schema table and record a Grilling-amendments entry: the install path projects a declared matcher onto the emitted settings.json hook entry (absent = match-all, so shipped capabilities are unchanged), and per-runtime matcher projection (ADR-857 D8) stays a separate concern. This amendment ships with the fix that introduced the field rather than as a follow-up. * docs(#1634): record WINDOWS-POSIX-MODE-BIT-ASSERT defect in CONTEXT.md Capture the CI failure pattern from #1634/PR #1638 so it is not repeated: a test that writes a file with a POSIX mode and then asserts statSync().mode & 0o777 === <octal> passes on macOS/Linux but fails on windows-latest (Windows fs does not honor POSIX write modes — reads back 0o666). Added as a machine-greppable DEFECT predicate (symptom/examples/detect/fix-forward/ prevention) next to DEFECT.WINDOWS-TEST-PORTABILITY, with the fix-forward: gate the mode-bit precondition on process.platform !== 'win32' and keep the platform-independent behavioral assertion running everywhere. |
||
|
|
fc2a7c0555 | fix(#1615): install Windsurf slash workflows | ||
|
|
08d1c57d6e |
fix(#1460): verify-or-reject capability --integrity per source; confine hook commands to the bundle
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e7855bc217 | fix(#1459): user-owned consent store gates third-party capability activation; env/cwd in disclosure; loader validator parity (#1473) | ||
|
|
353f63d170 |
feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) Promote the registry from a frozen data file to loadRegistry({includeInstalled}), composing the first-party registry with a validated installed overlay (ADR-1244 D2): - Extract the conformance validator to a shared runtime-callable module (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it verbatim, guarded by a generative-parity test (no build-time/runtime drift). - capability-loader.cts: loadRegistry({includeInstalled}) composes first-party ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic- prefixes); full merged-set cross-capability validation; engines.gsd load-time re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path escapes rejected. - semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed. - Wire surface/state + loop to the overlay; loop injects a blocking gate for each skipped gate-kind overlay (fail-closed). - cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd) + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/ config-set call (never eager at module load, never wrong-cwd); first-party path unchanged with no cwd. - run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact); capability-validator.cjs stays linted (#551 migration coverage). Closes #1431 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1431): add changeset for runtime capability registry overlay Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52) The cwd-aware overlay config-key federation added to config-schema.cts (_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd threading) introduced mutable surface uncovered by config-schema's mutation test set, dropping its score to 39.58% (below the 52 break threshold). Add a real-overlay-fixture describe block exercising every branch (cwd guard, overlay loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker score 39.58% -> 77.08%. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |