bbdf7e8e84b965d3a625bcc79b5667fcb59befc5
69 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
30468f16fe |
test(#4515): migrate installer-runs batch to named timeout constants (#4575)
* test(#4515): migrate installer-runs batch to named timeout constants Batch 4 of the ad hoc timeout literal migration (epic #4445). Replaces every bare numeric timeout/timeoutMs object-literal property across tests/install.test.cjs, tests/install-minimal-hooks.test.cjs, tests/fragment-single-edit-propagation.install.test.cjs, tests/install-regressions.test.cjs, tests/install-runtime-artifacts.test.cjs, tests/npm-integrity-gate.test.cjs, tests/faulty-deps.test.cjs, tests/release-tarball-smoke.install.test.cjs, tests/install-write-confinement.test.cjs, tests/plugin-manifest.test.cjs, and tests/packaging-shipped-scripts-require-only-shipped.test.cjs with a named constant, per eslint-rules/no-adhoc-timeout-literal.cjs. Removes these 11 files from the rule's allowlist. Adds one new shared constant to tests/helpers/timeouts.cjs for a fixture-JSON value (in seconds, not ms) mimicking a Claude Code settings.json hook-entry's own timeout field, shared across two files in this batch. Every other new constant is file-local, each carrying a comment explaining why its call site is a distinct operational class from the existing shared norms even where its digits numerically coincide with one. No src/bin file touched, no numeric timeout value changed anywhere. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: combine gsd-worktree-path-guard's git rev-parse calls to cut subprocess count under CI load Discovered while verifying PR #4575 (a full test (windows-latest) failure in an unrelated test, tests/kilo-upgrades.test.cjs's worktree-path-guard rejection test). Root cause: this hook's up-to-4 sequential git spawns (git-dir, branch, worktree-toplevel, file-toplevel), invoked inside a native-plugin's own 8000ms-capped subprocess wrapper, left zero margin for node/git startup overhead — the guard is deliberately fail-open on any timeout, so CI-load-induced latency in this chain silently downgrades a security block into an allow. Combines the first three git rev-parse calls (git-dir, branch via --abbrev-ref, worktree-toplevel) into ONE spawn instead of three, using git's documented multi-flag rev-parse output (one line per flag, in order) — verified empirically including the detached-HEAD edge case. No timeout value changed; this reduces subprocess COUNT, not any tolerance. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs: add changeset fragment for the worktree-path-guard fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f09e7ed08c |
fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath (#4375)
* test(#4137): keg-only Homebrew Cellar path falls back to raw execPath Regression tests for the Homebrew branch of normalizeNodePath: the rewrite to <prefix>/bin/node must be existsSync-guarded like the mise/volta branches, falling through to the raw execPath when the keg-only formula was never linked into <prefix>/bin. Also makes the existing #3181/#2185 Cellar assertions hermetic by injecting existsSync stubs (granting existence to exactly the one candidate each asserts) so they no longer depend on the runner machine's real /usr/local/bin/node. * fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath The Homebrew branch of normalizeNodePath returned <prefix>/bin/node unconditionally — the only one of five runtime branches that never probed its rewrite candidate. On a keg-only or versioned Homebrew install (node@24 never brew-linked) that path does not exist, so every managed hook command baked by resolveNodeRunner/buildBakedNodeToken/ buildNodeRunnerChainToken failed at invocation with exit 127, /bin/sh: <prefix>/bin/node: No such file or directory. Guard the rewrite with the already-injected existsSync exactly like the mise and volta branches: when <prefix>/bin/node exists (linked formula) the rewrite is byte-identical to today; when it does not, fall through to the raw execPath — a working keg path instead of an immediately broken one. Also drops two now-unused constants from the regression tests. * test(#4137): make the #977 non-fnm Cellar assertions hermetic too The Bug #977 folded block's two 'still maps to stable symlink' assertions called normalizeNodePath without an existsSync stub, silently depending on the runner machine's real /usr/local/bin/node (present on the Linux bench image, absent for /opt/homebrew). With the #4137 guard these become environment-dependent; grant each exactly the one candidate it asserts. * chore(#4137): add changeset fragment * chore(#4137): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
86b745b48b |
fix(#4270): forward Codex spawn model routing (#4281)
Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
77e2472ca0 |
enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets (the .env.example/.sample/.template/.dist templates stay readable). Read checks file_path; Grep checks an explicit path and judges the glob per brace alternative; Bash runs a two-pass token scan (quotes, comments, redirects with fd digits, separators, $( )/backtick/<( ) recursion, heredoc bodies never scanned as commands, nested bash -c/eval rescans, git <ref>:<path> shapes) with a closed non-reading exemption set for existence checks. Fail-open crash policy; 1 MiB commands are denied as command-too-large; more than 64 glob alternatives as glob-too-complex. Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt for approval whenever any Read() deny rule exists, even in auto mode. A hook denial is not a permission rule and never arms that check. The installer-written deny rules are retired in the follow-up commit. Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell), shell-command-projection managed sets, installer-migration-report, OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch), docs tables in five locales, ADR-766 always-on list, regen:derived fixtures, and a new table-driven unit suite. * test(#4221): pin the secret-read guard in existing hook gates Register gsd-secret-read-guard.js in every existing hook gate: the hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal- hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS, kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and typed-payload floors, the OpenCode adapter (grep mapping, include -> glob, three dispatch tests) and a Kimi TOML matcher assertion. * fix(#4221): retire installer Read() deny rules (legacy filter) Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets) strings. mergeClaudePermissions now only filters them out of an existing permissions.deny: an absent deny key stays absent, a malformed one is still repaired to [], and an array emptied by the filter is deleted so no `"deny": []` residue is left. Uninstall filters the same legacy list and, symmetric with the Antigravity branch, drops an emptied allow or deny key and an emptied permissions object. Unlike the #2278 allow-side migration there is no surviving current deny list, so the constant is renamed rather than mirrored. Removal is byte-exact: a hand-written identical rule is indistinguishable from the installer's and is removed too (the manifest never recorded permission strings). USER-GUIDE and CONTEXT.md updated. * test(#4221): flip install-regressions deny-rule assertions to the retired shape The fresh-merge, non-destructive merge, idempotency, end-to-end install, reinstall and uninstall assertions now expect no Read(.env*) deny rules and no permissions.deny key on a fresh install; the deny:null repair case is kept. A new describe block covers the legacy filter: retired strings removed with a user entry kept, partial sets, near-miss strings untouched, idempotency, GSD-only deny array deleted, a pre-existing empty deny preserved, and uninstall symmetry for allow/deny/permissions. * chore(#4221): add changeset fragment for PR #4236 * fix(#4221): case-fold names; scan shell stdin and xargs pipes Review round 1 (trek-e): - Blocker: secret-name matching is now case-insensitive in the Read, Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a case-insensitive filesystem are recognized as the same secret file. - Major: a shell interpreter's script is now scanned wherever it comes from. The tokenizer keeps heredoc bodies as per-segment tokens and records separator operators; pass 2 groups by segment id and resolves bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined `-lc`) scans the script operand, a file operand is checked as a file (a `<( )` operand's echo/printf output is reconstructed), otherwise stdin is the script and heredocs, here-strings and a piped echo/printf source are scanned. `eval` joins all its operands; `source`/`.` handle process substitution. Data heredocs (`cat <<EOF`, the commit-message shape) stay unscanned. - Major: `… | xargs <cmd>` checks the upstream segment's operands as file names when the sub-command reads (`echo .env | xargs cat`, `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the inference; a shell sub-command's `-c` script is scanned. Header, USER-GUIDE bullet and changeset updated; documented gaps now include piped scripts from non-echo sources and `exec`/`timeout` wrappers. 60 new suite cases pin the block and allow shapes. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
15af0f5536 |
enhance(#3951): B6+B7 — widen two unreachable lint rules and make the guard ledger true (#3965)
* fix(#3951): two lint rules that could not reach the code they govern B6 names two widenings. Measuring them first turned up a defect the criterion did not know about, and refuted the reason it gave for one of them. 1. no-adhoc-markdown-parsing self-gates on its own filename. Lines 107-110 short-circuit create() to {} unless the path matches /(?:^|\/)src\/[^/]+\.cts$/. B6 says to widen the files: glob in eslint.config.mjs - but doing only that ships an INERT rule, because the gate still returns {} for every new path. Both halves have to change, and the gate is the load-bearing one. That same regex hides a live hole: [^/]+ is FLAT-ONLY, so it requires the file to sit directly in src/. The registered glob is src/**/*.cts, which includes subdirectories. 28 .cts files - health-diagnostic-rules/ (10), installer-migrations/ (11), observability/ (3), host-integration-adapters/ (2), vendor/ (2) - are inside the registered glob and silently skipped. Measured with the gate neutralized: 0 violations there today. The hole is hiding nothing right now, and is fixed anyway, because "no violations today" is not a property that keeps holding. The fix is not invented: require-subprocess-timeout.cjs:196 already carries the correct form of this guard, /(?:^|\/)src\/.*\.cts$/ with .*, one directory over. Checked the other 21 rules for the same bug - no-adhoc-regex-escape and no-private-binary-resolution short-circuit only to exempt their own seam file, which is the right shape, and no-crlf-fragile-split has no filename gate at all. This bug is unique to the one rule. 2. no-adhoc-regex-escape could not see the shape that actually occurs. Line 396 gated the whole UNSAFE-NEW-REGEXP arm on arg.type === 'Identifier'. Every check below it - the _SOURCE provenance check, the isSoleReturnOfOwnParameter shape - lives inside that branch, so new RegExp(obj['key']) and new RegExp(cfg.pattern) were never examined at all. Runtime data arrives as a property access far more often than as a bare identifier, which is exactly why this rule never fired on the #3477 ReDoS. Widened to MemberExpression, measured by AST walk across all five registered blocks rather than by grep. 27 sites, zero TSAsExpression: 18 safe new RegExp(X.source, flags) -> exempted, keyed strictly on the PROPERTY being `source`, never on the object. Keying on the object would wave through X.anything and buy nothing. B6 estimated ~10; that was an undercount. 3 _SOURCE-suffixed constants reached through a required module namespace (phaseId.BRACKET_PHASE_TOKEN_SOURCE) -> the same provenance-exempt class the rule already recognizes for bare identifiers, extended to reach them. Without this the widening produces 3 false flags. 6 real findings -> marked, each a test extracting a pattern from a shipped file at test time, where the runtime contract IS the product. Deliberately the NARROW MemberExpression form. The rule's own isSoleReturnOfOwnParameter doc comment records that an earlier broad "any non-literal identifier" heuristic produced ~25 false positives and was rejected; a re-run of the census after this change flags exactly the 6 above and nothing else. Verified by execution, not by reading: the gate now accepts src/<subdir>/x.cts, still accepts flat src/x.cts, and still exempts paths outside src/ - each pinned by a test proven to fail against the old regex. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): give no-adhoc-markdown-parsing its reach, and fix the 80 parses it finds The rule self-gates on filename AND is registered on one glob, so widening either half alone is inert. Both move here: the gate now accepts tests/**/*.cjs and scripts/**/*.cjs alongside src/**/*.cts, and eslint.config.mjs registers it on the same two. A test pins that the gate and the registration AGREE, in both directions. The original defect was a gate narrower than its registration; the failure mode of this fix is a gate wider than its registration. Both are silent, so the test asserts the pair rather than either half. 80 violations across 43 files, all in tests/, zero in scripts/. 70 are routed through the existing seams - scanFencedBlocks, collectSection, stripFencedCode, tokenizeHeadings from markdown-sectionizer; splitTableRow, parseMarkdownTable, findTableWithColumns from markdown-table. Headerless STATE.md tables use splitTableRow per line, because parseMarkdownTable needs a real delimiter row. 10 are suppressed, 12.5%, well under the third that would have meant the rule is mis-scoped for tests/ rather than the tests carrying debt. Each names its reason: three regression guards (#3873 / bug-#21) are deliberately independent of the generator's own fence handling, and routing them through the seam would have them test the generator against itself; one is a negative-text probe that extracts nothing; six are a shell-pipe-to-jq detector whose regex coincidentally matches the table fingerprint and is not markdown parsing at all. All ten sit in tests whose subject is .md content, which is normally a reason to prefer the seam. The marker used is allow-adhoc-markdown, distinct from no-source-grep's allow-test-rule, and lint:ci's lint-allow-test-rule-refs reports the same 280/280 unverified count as before - checked rather than assumed, because those two markers are easy to conflate. The widening earned its keep immediately: it found a test that passed for the wrong reason. tests/config-field-docs.test.cjs asserted notEqual(<cell>, '600') against the TYPE column instead of the DEFAULT column. notEqual('number', '600') is true forever, so the guard against workflow.subagent_timeout regressing to the old seconds default could never fire. docs/CONFIGURATION.md:434 is `| workflow.subagent_timeout | number | 300000 | ... |`, so the default is cell index 2; the assertion is now row-scoped through splitTableRow and reads 300000. That is the argument for the widening in one case: the violation was invisible to lint, the suite was green, and the assertion was vacuous. A rule that cannot reach a file cannot tell you the file is lying. Not fixed here, and recorded rather than assumed: #3426/#3239 are NOT reachable by this widening. tests/package-legitimacy-gate.test.cjs yields zero violations even with the gate bypassed - its hand-rolled scans are real, but built from line filters and split('|') rather than the regex-literal fingerprints this rule detects. They need new detectors. The epic assumed a wider glob would catch them. build:lib, lint and lint:ci all exit 0; the post-fix census across tests/** and scripts/** is 0 violations. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): B7 — and #3356's defects were still live in the code B7 asks that each closed child be driven fail-first with a behavioral identity test at the CONSUMER's output. Four of eleven children had no test citing their issue number. Auditing them by BEHAVIOR rather than by number-grep changed the answer for three of the four. #3364 and #2540 — traceability only. Both were implemented by #3941 and their consumer-output tests exist and were shown failing-first; neither cited its originating issue, so an audit that greps for the number reports them uncovered. Tagged the specific asserting test in each file, following the citation form those files already use. #3372 — covered, but only at helper level, and the triage narrowed it. Of the four commands the issue names, only estimate-cli's collectCalibrationSamples actually enumerates phase dirs from disk; smart-entry, audit and roadmap-upgrade derive from ROADMAP/body text and never reach the sentinel path, so they are benign by construction and were left alone rather than "fixed" into churn. The existing #3882 rows asserted the helper's return value. Added a consumer-output test driving `query estimate-calibrate` and asserting sample_count and the persisted document. RED proof: reverted collectCalibrationSamples to a raw readdirSync and ran the real CLI - sample_count 3, sentinel leaked; restored - sample_count 2. #3356 — NOT covered, and BOTH halves of the defect were still live in source. The issue is closed; the bug was not fixed. Fixed here rather than writing tests that document a bug as correct. Defect 1, the contradicted row. quick.md:627 claimed `quick-tasks-append` performs "the equivalent write" to the Step 7c row. It did not: the `#` cell was a positional ordinal and `Directory` read `—`, because the route had no way to receive a quick id or task directory. Added OPTIONAL `--quick-id` / `--slug` / `--directory`. A caller with neither - fast.md, the original #2133 caller - omits them and gets the byte-identical prior row, so nothing existing changes. A caller that HAS a real id and directory now gets the canonical row quick.md:632 renders. The false-equivalence sentence itself is corrected rather than left to mislead the next reader. Defect 2, the forced re-derive. The route called readModifyWriteStateMd with no options, so a body-only append to the Quick Tasks table triggered a full re-derive of the disk-derived progress.* frontmatter. Every other body-only writer passes { resync: false } - src/state.cts's own docstring prescribes it - and this route was the lone outlier. RED proof: reverted the option, seeded a project with 2 real phase dirs and a curated total_phases of 25, ran quick-tasks-append; total_phases collapsed to 2. Restored; it stayed 25. That second one is the shape this epic exists to close: a silent write that replaces curated state with a re-derivation nobody asked for, exit 0 throughout. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3951): amend B6's ledger to what was measured, and document the new flags The ADR gains a ledger amendment in its own correction style - the sixth wrong premise it records, found the same way as the other five, by measuring before building. B6 says the net guard count must fall. It rose: 62 -> 69, +7, measured from the epic's filing commit to origin/next. The attribution is the point, though. Five of the seven came from PRs unrelated to this epic, one was added by a phase of it, and the epic did retire something sub-file - #3884 removed a detector with an explicit "net: -1 detector, 0 added" ledger. Every named casualty is load-bearing, two already carry retractions in this same document, and a sweep of all 22 rules plus every scripts/lint-* found no provably dead guard. There is no honest way to make the count fall; forcing it would trade coverage for a number, which is the Goodhart outcome Decision 6 exists to prevent. The amendment also records that B6's own prescribed fix for one widening was inert. no-adhoc-markdown-parsing self-gates on its filename, so widening only the files: glob - which is what the criterion says to do - ships a rule that still returns {} for every new path. And #3426/#3239 are not reachable by that widening at all; their scans use line filters and split('|'), not the regex fingerprints the rule detects. The roster row tracked them against the wrong mechanism. Three roster rows updated from aspiration to fact: the two widenings are DONE with their measured counts, and lint-phase-enumeration-drift is marked RETAINED rather than "expected casualty - verify before retiring", because Phase 5 verified it and kept it. The rule Decision 6 should carry forward is stated plainly: a guard ledger is a claim about COVERAGE, not about COUNT. "Net count must fall" is measurable and wrong. "Every guard is reachable, and each retirement names what makes its defect unrepresentable" is the property that was actually wanted. CLI-TOOLS.md documents the optional --quick-id/--slug/--directory flags and says plainly that omitting them keeps the pre-#3356 row byte-identical, plus that the append no longer re-derives progress frontmatter. New features fragment (id 3951); FEATURES.md regenerated rather than hand-edited. Changeset is Changed, pr:0 pending backfill. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): correct four rows that pinned the lint rule's old narrow reach The remote suite came back RED with 5 failures, all in tests/eslint-rules.test.cjs. They are stale tests, not a regression: four rows assert that no-adhoc-markdown-parsing is inert outside src/*.cts, which is exactly the contract this deliverable changes. Confirmed by reading rather than inferred from the names - the row at :1981 used filename: 'tests/some.test.cjs' and filename: 'scripts/helper.cjs', the two roots the rule now covers on purpose. Worth recording WHY local gates missed this. npm run lint and lint:ci were green, and the touched test files passed standalone. Lint only reports violations in real files; these rows assert the rule's REACH using synthetic RuleTester filenames, so nothing but the full suite could see them. Local green on a rule change says nothing about the rule's own tests. Each row is rewritten with BOTH halves rather than flipped from valid to invalid: - the same fingerprint under tests/ or scripts/ is now flagged, with the right messageId - the negative space is preserved - the same fingerprint under a path outside all three roots (gsd-core/bin/lib/foo.cjs) is still NOT flagged The second half is the one that matters. Without it the rule has no boundary and nothing would catch an over-wide gate later, which is the mirror image of the bug this deliverable just fixed. Each row is renamed to state the current contract; the old names said "non-src/*.cts ... is not flagged" and would have been actively misleading once the bodies changed. Proven to test the widening rather than restate it: every flagged half was run against HEAD~2's pre-widening rule and does NOT fire there, then against the current rule and does. 12/12 on that probe; the full file is 178/178. Swept for the same staleness elsewhere and found none. require-subprocess-timeout's own "inert outside src/*.cts" row is untouched - that rule's gate was not widened here - and no-adhoc-regex-escape's test file already carries correctly-targeted rows. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): acknowledge the quick.md growth the attribution guard reported The full suite came back RED with one failure, and it is mine: 1 file(s) grew without an acknowledgment: quick.md grew 364 bytes gsd-core/workflows/quick.md is runtime-loaded emitted content, so correcting its false 'performs the equivalent write' claim trips emitted-attribution by construction. This is the acknowledgment, not a workaround - there is nothing to regenerate. The fragment names ONE path, which is the only one the guard reported. The four spent acknowledgments it also listed (audit-uat, plan-phase, progress, review) belong to other fragments whose ripple the base already absorbs; they are inert, not failures, and are deliberately NOT copied here - naming paths I did not change would make this record false in the other direction. Byte figure corrected before committing: the guard reported 37220 -> 37584 (+364), but origin/next has since moved and quick.md is 37232 there now, so the measured delta is +352. The reason text says so and names the base as a moving figure rather than pinning a number that is already stale. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): move the quick.md growth ack to a trailer, delete the obsolete fragment The acknowledgment mechanism changed under this branch. Merging next brought in the redesign - it also deleted .github/workflows/ack-fragment-sweep.yml, which was in the merge status and which I did not register at the time - and the guard now says so directly: Add a trailer to a commit in this PR (never a new file). Emitted-Drift-Ack-Growth: quick.md - <why this growth is deliberate> So tests/emitted-drift-acks/3951-quick-append-equivalence.json is obsolete on arrival. A fragment file is no longer read by anything, and leaving it would be a dead record that looks like an active one. It is deleted here rather than kept "just in case". The byte figure moved again with the merge: 37232 -> 37596, +364. The earlier fragment said +352, measured before the merge auto-merged quick.md itself. The trailer carries no number, which is the better design - the figure was stale twice in two attempts. Refs #3951 Emitted-Drift-Ack-Growth: quick.md — #3356/#3951 replaces a false claim with an accurate one. Line 627 said the `quick-tasks-append` shortcut "performs the equivalent write" to the Step 7c row rendered above it; it did not, and that was the documented half of #3356 — with no quick id or task directory the route emitted a positional ordinal in `#` and an em-dash in `Directory`, a visibly different row. The corrected sentence has to carry three facts the original elided: what the shortcut actually writes when it has neither input, that this is honest behavior for its real caller (`fast.md`, which has neither), and how a caller with both now gets the byte-identical canonical row via the new optional `--quick-id`/`--slug`/`--directory` flags. Prose is the product here — an executing agent reads this line to decide whether the shortcut is safe for its case, and a shorter correction would either drop the flags (leaving the reader unable to act on the fix) or drop the limitation (recreating the false claim in gentler words). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3951): backfill changeset pr number Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
39673ae9ff |
fix(#3738): antigravity global skills/agents install to ~/.gemini/config (#3921)
* test(#3738): antigravity global skills/agents must resolve under ~/.gemini/config Regression tests (RED first): --skills-root and gsd-tools query surfaces, install-plan dest dirs, and converter skills-path rewrite. * fix(#3738): antigravity global skills/agents install to ~/.gemini/config Antigravity's machine-local discovery scans ~/.gemini/config/{skills,agents}; the configHome (~/.gemini/antigravity) is deprecated for artifacts. Declare the ADR-1239 skills/agents 'home' override on the antigravity global layout — the same mechanism codex uses (.agents) — and divert ~/.claude/skills/ references in converted global content to ~/.gemini/config/skills/. configHome, settings, probe/migration semantics, and the local .agents layout are unchanged. * fix(#3738): retire deprecated configHome artifacts via installer migration 010 Next install converges an existing antigravity install: manifest-managed skills/gsd-*/ and agents/gsd-*.md under the configHome (a location AGY does not scan) are removed — modified files backed up first, unmanifested and non-gsd entries preserved — and now-empty containers retired. Global scope only; the local .agents surface is live. Docs + inventory updated. * fix(#3738): converter sync in bin/install.js, harness emit-root coverage, migration baseline - bin/install.js converter gains the same ~/.claude/skills → ~/.gemini/config/ rewrite as src (ADR-1508 dual copy must stay in sync). - Parity-manifest walk covers home-override emit roots (extraEmitRootsFor) so antigravity's emitted skills/agents stay differential-visible at their new install root; install-tree fixture regen confirms an unchanged key set. - skills-from-commands rule declares the antigravity converter as a runtime-scoped transform; one ack fragment covers the identity-classed workflow whose antigravity copy embeds the old skills path. - Migration 010 checksum baseline + home-override set doc updated; existing tests updated to the #3738 contract (global dest, golden parity via layout dest, integration expectations). * fix(#3738): tolerate an absent extra emit root on baseline-side measurement The base tree's installer predates the home override, so <HOME>/.gemini/config does not exist there; walk() threw ENOENT and the in-job baseline build failed. An absent extra root is the legitimate pre-override shape — skip it. * fix(#3738): review findings — manifest agents root, bare skills-path rewrite, guard comment - writeManifest resolves the agents-kind home override (_kindDestDirSafe), so the manifest records agents at their actual install root and drift detection keeps working (isolated review finding 1, major). - Converter bare forms ~/.claude/skills and $HOME/.claude/skills (no trailing slash) divert to ~/.gemini/config/skills instead of falling through to the retired configHome path (finding 2). - real-home-guard comment updated: antigravity's global agents kind is the first agents-kind home override (finding 3, doc-only). - Regression tests for both behavioral findings. * chore(#3738): changeset fragment (pr number backfilled after PR creation) * chore(#3738): backfill changeset PR number (3921) * fix(#3738): sandbox HOME in tests that install antigravity global artifacts antigravity is the first home-override runtime in the golden-parity and skills-wrapper suites (codex is not in their runtime lists), so those tests never needed HOME sandboxing — the real-home guard now (correctly) refuses their un-sandboxed global installs on CI, where HOME is the passwd home. * fix(#3738): stop the K3 sequential-sandbox env leak; sandbox L2's home-override plans K3's two back-to-back sandboxHome calls leave HOME pointing at the first sandbox once the after-hooks restore (each call saves the env as it found it, so the second saves the first's sandbox as 'original'). On the windows matrix that leaked gsd-k3-qwen-* home into the L2 property, whose antigravity/global run then (correctly) refused via the #3712 real-home guard — antigravity is the runtime that made L2's plan escape into os.homedir(). K3 now manages the env with a single restore; L2 sandboxes HOME per run, mirroring L1. * fix(#3738): L2 property's HOME sandbox must exist on disk The #3712 guard's sandbox exemption fails closed when identify(effectiveHome) is 'absent' — L2 never created its configDir, so on the windows matrix (tmpdir under the real home) the antigravity/global run refused even with HOME sandboxed. Create the per-run sandbox dir and clean it up. --------- Co-authored-by: sim <sim@local> |
||
|
|
27318ecec3 |
fix(#3704): rewrite fnm versioned node paths to the stable alias on macOS and Linux (#3856)
* test(#3704): failing-first coverage for fnm versioned-path normalization on POSIX * fix(#3704): rewrite fnm versioned node paths to the stable alias on POSIX * fix(#3704): share one root normalizer across both fnm branches so no baked path carries a doubled separator * chore(#3704): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
107eb8c1d9 |
feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on
|
||
|
|
682eaae3f0 |
enh(#2876): retire the dead and pass-through exports from bin/install.js (#3615)
* enh(#2876): retire the dead and pass-through exports from bin/install.js The installer exported 197 names and had zero production consumers - every non-test require of it repo-wide sits inside a comment. Its interface was shaped by test access, not by callers. Removes 9 dead exports and 61 pass-throughs, repointing their tests onto the extracted modules' own interfaces. 197 down to 127. Every count in the issue was wrong: 197 exports not 188, 9 dead not 12, 61 pass-throughs not 49, 44 test files not 42 - and the audit itself then missed 7 more consumer files. restoreUserArtifacts was on the dead list but ceased to exist in phase 6, and two _GSD_EFFORT_MANIFEST_* names listed as dead are now genuinely asserted, so acting on that list would have deleted live exports. 7 of the 9 dead names collide with an independent declaration that install.js delegates TO. Each removal was justified by which declaration a reference resolves to, never by whether the name appears somewhere. Coverage parity was the gate rather than test greenness: per-file counts were captured before any edit and diffed after. 44 of 45 files are byte-identical; the single delta is one added assertion, not a loss. The sweep for scattered require sites found two forms static grep misses - require(VARIABLE) and multi-line require() - plus tests asserting that install.js re-exports the SAME object, which now assert retirement instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2876): close review findings — restore the duplicate-body guard, sweep orphaned code Both review engines found real defects in the first cut. The DEFECT.GENERATIVE-FIX single-owner guard from #1511 had been repointed from a reference-identity check to install.X === undefined. Those are not equivalent: the guard exists to catch a duplicate function body reintroduced into install.js, and the replacement passes cleanly if that duplicate is used internally and never exported. It now walks bin/install.js's real top-level bindings, so it catches a duplicate under either shape, exported or not - strictly stronger than the check it replaced. Proved by injecting a duplicate and watching it go red. That weakening survived the coverage-parity gate because the assertion count never moved. The gate compares counts, so an assertion that changes meaning rather than number is invisible to it. Removing the exports had orphaned their wrapper bodies: 14 dead wrappers, 9 consts and 9 destructure entries, several pre-existing and found by the same sweep. Dead code left in the file this phase exists to shrink. Three more comments claimed re-exports this phase removed, and tests were reading Cursor and Windsurf hook constants from install.js's local copy while calling functions from the hooks surface - equal today, with nothing holding them equal. The local consts now reference the owning module. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2876): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3ab0007164 |
enh(#2875): materialization primitives — durable user-artifact staging and descriptor-authoritative agents (#3600)
* fix(#2875): stage user artifacts durably across install wipes (#1874-F19) preserveUserArtifacts held user files only in an in-memory Map across the wipe, so any process death between preserve and restore lost them outright. Seven call sites, not the four the issue records. Three of them never called the helper at all - they open-coded the same read/wipe/write - so searching for callers under-counted by construction; the extra sites were found by sweeping for the pattern instead. The worst is the mainline install path, where the crash window spans the entire gsd-core tree copy rather than a single rmSync. Adds src/user-artifact-staging.cts: durable on-disk staging with a record written after the copies land as the commit point, plus recovery of orphaned batches on the next run - without recovery the staged bytes survive but the user's file is still gone, which would pass its own test while delivering nothing. Routes copyPreservingSymlink through installFs() so staging cannot bypass the install fs seam, and reunites its symlink-safety docblock with the function it documents. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-3574 with four claims disproved by implementation Implementing Phase 6 disproved four statements the ADR rests on. The central decision - no single materializer - is unaffected and stands. Corrected: decision 3 was already satisfied, so nothing was extracted; the agents-bypass runtime set omitted claude, kilo and opencode, and closing it needed three new pieces of descriptor contract rather than proceeding on its own terms; three of the four blockers the layout comment names were already stale; and F19 is seven call sites, not four. Records the generalizable lesson: the defect is the pattern of holding user data in memory across a wipe, not the helper, so searching for callers of the helper under-counts by construction. Also resolves the ADR's open question on USER_OWNED_ARTIFACTS membership, and notes that copyPreservingSymlink needed routing through the install fs seam before it could be reused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close dangling-symlink blind spot and harden staging recovery An adversarial review found the F19 staging work shipped red and unsafe. Root cause, shared by two arbitrary-write findings: hasExistingSymlinkBetween missed dangling symlinks in both its root check and its per-segment walk, because it probed with existsSync, which is false for a link whose target does not exist. Fixing only the new module would have reused a guard that was itself blind. This guard protects the whole install tree. Recovery no longer throws: it degrades per entry and per file, so one bad batch cannot block the others. Previously an unrecoverable entry propagated out of the first statement of install and uninstall, before the cleanup that would have removed it - wedging the installer permanently. Partial fs adapters now throw on any omitted method instead of silently reaching the real filesystem, closing the trap that let a test poison list pass while real IO happened. Staged names must be flat, recovery refuses a dangling destination symlink, and a batch whose recovery genuinely failed is no longer swept - it was discarding the only durable copy of the file it had just failed to restore. Replaces three tests that could not fail, including the one labelled negative proof. Known limitation, documented not closed: concurrent installs sharing a staging key can still lose a batch. A real fix needs a cross-process lock. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enh(#2875): make the descriptor authoritative for the agents kind Deletes the inline agent-staging loop in bin/install.js and the _DESCRIPTOR_AGENTS_RUNTIMES set, so every runtime materializes agents from its capability descriptor instead of an inline hostBehaviors dispatch. Closing it needed three pieces of contract the descriptor pipeline never had, all reducible to one missing input - per-agent resolution context: a frontmatter-extensions step for claude's effort and disallowedTools, per-agent model-override resolution for kilo and opencode, and a named branding converter for hermes, whose rewrite data was already declared. Seven runtimes were on the loop, not the six the design recorded - kimi-code was found by a golden fixture, not by analysis. claude-local and kimi-code both silently lost their agents mid-change; the fixtures caught both and the cause was fixed rather than the fixtures regenerated. A parity harness gates the migration: both pipelines over identical inputs, byte-identical output including filenames, per runtime. It is demonstrated red before being trusted. Surface and install paths converge for all seven, which also fixes surface previously writing no agents for these runtimes. Codex's config.toml strip stays put - it mutates host config, which no descriptor kind models. Also routes install-model-override-resolver and install-effort-resolver through the install fs seam. Both leaked real filesystem IO from the install call tree; the stricter adapter is what exposed them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): record the agents-descriptor migration and correct the ADR count The _DESCRIPTOR_AGENTS_RUNTIMES allow-list no longer exists, so the host integration guide told readers to join a set that is gone. Replaces that with what is now true - declare an agents entry and it installs, on the surface path as well as install - and points anyone needing a per-agent transform at the three extension points rather than at a new inline branch. Corrects the ADR amendment: seven runtimes were on the inline loop, not six. kimi-code was found by a golden fixture going red, not by reading. That is the third short count this phase, all from enumerating by symbol or set membership when the thing that matters is a behavior. Adds the Changed changeset for the surface-path convergence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): amend ADR-2866 - claude global always wrote agents on disk The claude row's global=[skills] described what capability.json declared, not what the installer wrote. bin/install.js's inline agent-staging loop was never scope-gated and never consulted the descriptor, so a claude --global install has always written agents/gsd-*.md. Phase 6 closes the gap by deleting that loop and declaring agents on claude's descriptor at global scope. On-disk bytes are unchanged - the golden fixtures did not move, which is the evidence that the descriptor, not the installer, was incomplete. #2218 is unaffected: agents are not trigger-bearing, so the wider row does not introduce a new shadowing case. Records the warning that an incomplete descriptor is invisible while a second code path silently does its work, and only surfaces when the two are forced into agreement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close review findings across staging, agents and the parity harness Two independent reviews of this branch found defects the local gates missed. Security: a dangling symlink at a migration destination allowed writing outside configDir - the same class this change claimed to close, missed at the terminal write of the flow being added. The staging-root resolver threw as the first statement of install and uninstall, so a hostile symlink bricked both, and symlinked-configDir users lost uninstall as well as install; it now degrades instead of aborting. Recovery gained a source-side symlink check and now refuses a relative destDir, which resolved against cwd. Converter dispatch gained a runtime allowlist - lint-time validation stopped mattering once this branch promoted that dispatch from the surface path to real installs. Correctness: claude --local --minimal exited 1 because the minimal profile legitimately yields zero agents and the new path treated that as a failure. cline --local silently lost its agents - its descriptor declared none while the deleted loop wrote them unconditionally. The agents prune was widened to any gsd-* entry and destroyed user files it never owned. The parity harness, on which the migration's safety argument rested, drove a synthetic registry and never byte-compared the shipped descriptors; two of its trap rows could not fail. It now drives the real registry across 13 runtime-scope rows including kimi-code and cline-local, and its red-proof is demonstrated by corrupting a live capability.json. Three goldens that had encoded the cline regression as expected behavior were corrected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): close findings from both mandated review engines /security-review found the staging source-side walk honouring GSD_ALLOW_SYMLINKED_DEST, an opt-in documented as relaxing only the write destination. A symlinked files/ component dereferenced because copyPreservingSymlink lstats the leaf only, so an intermediate link is followed. The source walk no longer honours the opt-in; the destination check still does. /code-review spec axis found this branch had reintroduced its own bug: migrateLegacyDevPreferencesToSkill's new symlink refusal threw unguarded after the legacy dir was wiped and before the staged batch was restored, so a planted symlink bricked uninstall permanently and orphaned the batch. Refusal kept, abort removed. kimi-code local silently lost its agents, the same class as the cline bug, and the parity harness recorded that exclusion as intentional - the third test in this branch to pin a regression as correct. --minimal now creates an empty agents/ dir that never existed. Behaviour restored rather than softening the changeset, so its byte-identical claim stays true. Standards axis: try/finally removed from twelve test bodies, fast-check properties added for parseOwnerPid, boundary coverage at the grace window and the ancestor-probe depth, a parity assertion for the staging-root helper duplicated across two files, and the 8-deep config walk deduplicated. Records 60-review.json with every finding and disposition from five passes, including the smells left unfixed and why. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2875): prune stale agents unconditionally in minimal mode The previous round stopped an empty agents/ directory being created when the resolved profile yields no agents. That was implemented by skipping the agents kind entirely, which also skipped its stale-agent prune - so a full to minimal downgrade left stale gsd-* agents behind. The deleted inline loop pruned unconditionally and only skipped writing. Those are three separate conditions, not one: prune always, write only when there is something to write, create the directory only when writing. Both call sites now run _removeGsdEntries before the empty-staged early exit. The symlink-escape guard moved with it, since the prune also touches dest. Codex .toml agents and the config.toml stanzas are cleaned again, and user-owned agents are still preserved. The agents/ directory is left in place after a prune empties it, matching every sibling kind - none of them remove the destination directory itself. Golden fixtures confirmed byte-identical: the prune is a no-op on a fresh install, so fixture generation is unaffected. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2875): document interrupted-install recovery for user-owned files The durable-staging fix is invisible to the user it protects. Someone whose install died mid-flight has no way to know USER-PROFILE.md was staged before the delete, that the next run restores it, or that recovery happens at the start of that run rather than in the background. Written as the task the user has - finish the interrupted command - rather than as a description of the mechanism, and states what it will not do: overwrite a file already present, or touch staging belonging to another install still running. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2875): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2875): assert the J8 model override without building a regex CodeQL flagged incomplete string escaping: the assertion interpolated the override value into a RegExp while escaping only forward slashes, which is meaningless in a constructor, leaving real metacharacters unescaped. The failure direction was the dangerous one - a metacharacter would have made the match more permissive, so the row would pass when it should fail. That matters here because J8 exists precisely because an earlier revision was a tautology; the rewrite reintroduced a different way for the same assertion to stop discriminating. Replaced with a line-wise exact match, so no regex is constructed at all. Swept the other test files this branch adds; no sibling instances. lint:ci passed on the original - lint-no-adhoc-regex-escape matches a full metachar-escape copy, so a single slash replace slipped under it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a4a02a7a01 |
enhance(#2874): return the executed plan and route install IO through a seam (#3568)
* test(#2874): add failing-first gate for the executed-plan return Four rows from the matrix's red-first order. E3 pins the one early return, for the opencode family, where a void-shaped hole would otherwise survive unnoticed. E13 sweeps every runtime in the registry - enumerated from the registry rather than hardcoded, so a runtime added later cannot slip past. F2 proves absence of real filesystem contact rather than merely that the happy path ran, which is the difference between a complete seam and a partial one. G1 and G3 are the additive guard and must be green before and after. G3 deliberately leaves the two existing adapter test doubles untouched: if this change required editing them it would not be additive, and the acceptance criterion would be unmet. No production code. All 19 runtimes install without throwing today, so E3 and E13 fail on the undefined comparison alone. Refs #2874 * feat(#2874): return the executed plan and route install IO through a seam installRuntimeArtifacts returned void, so its correctness was observable only by re-reading disk. It now returns what it executed - per kind, per scope - including on the combinedFamilyInstall path, which was the one early return where a void-shaped hole would have survived unnoticed. Failure still throws rather than becoming an ok:false return, so control flow is unchanged for both existing callers. A best-effort cleanup that fails is still swallowed, but is now visible in the returned value rather than silently absent. The fs seam is ambient rather than threaded. Explicit deps through install-profiles and the 3000-line conversion module was impractical; the tradeoff, the synchronous-only re-entrancy assumption, the restore guarantee and the partial-adapter fallback trap are all documented at the seam. findInstallSourceRoot and its sibling stay unrouted by design - they locate the package's own source, not the install destination. readCmdNames keeps a second implementation because the standalone CLI that owns the original cannot require the compiled adapter without a build-order dependency on its own output. A parity test fails if the two ever disagree. Refs #2874 * chore(#2874): gitignore the new build artifact install-fs-adapter.cjs is tsc output from src/install-fs-adapter.cts, not a tracked source file. It was added to eslint's ignore list but not to .gitignore, so it landed as a tracked file - the third time this step of the new-.cts ripple has been missed on this epic. Refs #2874 * fix(#2874): close two seam leaks and correct a false comment A correctness review found the seam still leaked in two places, both subtler than the three already closed. readGsdCommandNames was routed when it should not have been: it reads the package's own commands directory, which a destination-fake is never seeded with, so under a fake adapter it returned an empty or wrong roster instead of failing loudly. It now reads real fs, matching the precedent already documented for findInstallSourceRoot. cleanupStagedSkills ran raw rmSync from a process exit handler, which is real filesystem work deferred past the point where withInstallFs has restored - the one thing the synchronous-only contract exists to exclude. Staging now captures the adapter that created each directory and cleanup replays it, so a real install cleans up exactly as before and a fake-staged path never reaches the real filesystem. Also corrected a comment claiming the migration reads were an unrouted, untested residual gap. They are routed and exercised; a comment understating the seam is as corrosive as one overstating it in a module whose trust rests on being honestly documented. Refs #2874 * test(#2874): migrate the exemplar group and cover the matrix AC3's exemplar migration lands in place: the qwen install group now asserts skills and agents destinations from the returned plan in one deepStrictEqual instead of probing the filesystem for each. Nine facts the old probes established were enumerated first. Two moved to the value assertion; seven were retained deliberately - per-file SKILL.md existence, the VERSION file written outside this function, the manifest content, and the post-uninstall absence checks all sit outside the plan's per-kind contract. A migration that quietly asserts less looks like a win and is a regression, so the enumeration is the guard rather than the line count. Also implements the rest of the matrix: the executed-plan shape, adapter failure modes, the security-boundary rows including a fake that cannot certify an install the real filesystem would refuse, cleanup visibility, and two seeded property tests. Only the two external CI gates are left unticked, because self-certifying them would be a claim rather than a check. Refs #2874 * fix(#2874): restore streaming hashes and derive F2 from the boundary rule The checkpoint found three things reasoning had missed. sha256File had been converted from raw-fd streaming to a single readFileSync on the assumption that GSD artifacts are never large. A test named for exactly that contract already existed and went red. Streaming is restored, now routed through the adapter, which gains openSync, readSync and closeSync. The contract was the specification; the assumption was not. Three existing tests inject faults by monkeypatching real fs. They broke because mkInstallTempDir stopped calling real mkdtempSync, not because of any binding subtlety - the real adapter was already late-bound. It now calls the real function when no fake is injected, so a monkeypatch applied after import is still seen and the additive contract holds. F2 poisoned real fs by method, so a deliberately unrouted package-source read failed a correct design. It now poisons by path: destination IO is forbidden, package-source IO is allowed and positively asserted. The claim was always zero real destination IO, and the test now derives from that rule instead of coincidentally matching it. Refs #2874 * docs(#2874): add the contributor how-to for plan-based test migration The phase gate caught a real gap. The docs plan was Reference plus Explanation only, and every CI check would have passed, because the docs-required lint only verifies that some file under docs/ moved. But this phase exists to demonstrate a pattern for follow-on work, and that work is other contributors migrating probing test groups. The sequence has two live traps - a partial fake silently falls back to real fs, and the seam is ambient and synchronous-only - plus one discipline nobody infers: enumerate the facts before converting, or you assert less and call it a win. The page carries the qwen migration's arithmetic, nine facts enumerated and only two converted, because a reader seeing only the diff would reasonably conclude the pattern is to replace probes wholesale. No locale mirrors: none of the four carries any contributor-only how-to, so a single translated file would manufacture parity rather than provide it. Refs #2874 * chore(#2874): backfill changeset pr number * test(#2874): normalize both sides of the G1 tree comparison G1 failed on Windows only, deterministically on both shards. The defect was in the test helper, not production. _computePathPrefix posix-normalizes the resolved config dir unconditionally, so on Windows the path embedded in every emitted SKILL.md body is forward-slash form. hashDirTree stripped against the raw backslash path from mkdtempSync, so the substring never matched and each install's unique temp suffix stayed baked into every file - all fifteen skill bodies hashed differently for two runs that had written identical bytes. Both sides are now normalized unconditionally rather than gated on path.sep, matching the rule this repo already records: backslash paths arrive on Linux too. Production code is untouched and was verified correct. Normalizing this away on the production side would have hidden a real portability bug if one had existed. Refs #2874 --------- Co-authored-by: sim <sim@local> |
||
|
|
8a6d87538f |
test(#3466): replace 8 source-grep assertions with behavioral tests (#3500)
Phase 2 of #3464, following #3465. Rewrites every assertion that read a shipped .cjs/.js file and text-searched it, so the file no longer needs an allow-test-rule exemption. 6 of the 8 files are now marker-free; the ceiling drops 285 -> 278 against a measured 277. A source-grep passes when a STRING is present, not when the code WORKS. It survives a refactor that keeps the string but breaks the behavior, and breaks on a refactor that keeps the behavior but renames the string. Both failure modes are silent about the thing the test claims to protect. That is the anti-pattern ADR-456 and local/no-source-grep exist to prevent. Rewrites, each against the real exported seam: - discuss-mode: calls cmdInitPlanPhase() against a fixture whose config sets workflow.text_mode, asserts the value propagates to its emitted JSON. - effort-surface-axis: runs review-lane invoke against a real project with a fake `claude` PATH shim, asserts the resolved --effort actually lands in the shim's captured argv. - install-minimal-hooks: calls applySettingsJsonHooks() with a hook source missing, asserts it is neither registered nor silently registers wholesale, with sibling present hooks as the positive control. - install: calls the exported inferPreferredRuntime({fs, env, preferredConfigDir}) via its injected fs seam, asserting 'kilo' from both the config-marker and env-var paths. - opencode-permissions: spawns the real installer with a custom config dir, asserts the written opencode.json permission paths are anchored on it. - repo-layout: spawns the installer for copilot local vs global, asserting AGENTS.md is written only in the local case. - runtime-config-adapter-registry: stubs resolveInstallPlan and force-reloads install.js, asserting the runtime's artifact stops being written -- proving install.js genuinely routes through the registry. - runtime-homes-descriptor-drive: calls buildAgentSkillsBlock() for cursor and claude against real fixture SKILL.md files, asserting each runtime's refs land under its own config dir and never the other's. Every rewrite was mutation-checked before its marker was dropped. Because node --test is not runnable locally in this repo, each assertion was replicated in a standalone probe that requires the same module: run green against the real file, then red against a deliberately broken one (text_mode propagation deleted, existsSync guard removed, config dir hardcoded, !isGlobal guard dropped, kilo branch removed, effort resolution bypassed, skills base hardcoded back to .claude), then the production file restored and confirmed byte-identical. An assertion that could not be made to fail would not have shipped -- a behavioral test that passes regardless of correctness is strictly worse than the source-grep it replaces, because it looks rigorous while asserting nothing. Three claims were checked rather than trusted. All three were wrong: - repo-layout's own marker cited #1188 asserting the `!isGlobal` lexical scope was "unprovable at runtime". It is provable: the guard decides whether AGENTS.md is written, which is directly observable. Both directions verified. - The triage for runtime-config-adapter-registry claimed its source-grep was redundant with the file's EXPECTED_TABLE tests, so deletion would be safe. Those tests only exercise resolveInstallPlan() directly and never load bin/install.js, so they do not cover it. A real behavioral assertion was written instead of deleting coverage. - An earlier revision of this change dropped runtime-config-adapter-registry's two markers on the grounds that ESLint stayed silent without them. An adversarial review caught that this was wrong, and it is restored here. The file still genuinely source-greps bin/install.js at two sites; ESLint is silent only because no-source-grep's TEXT_METHODS omits matchAll. Dropping a marker because the linter cannot see the violation is exploiting the blind spot, not resolving it -- and it would go red the moment the rule is widened. Those two assertions are also irreducible: they assert that EVERY inline `runtime === '...'` branch in install.js names a registry-known runtime, and a branch naming an unregistered runtime would simply never execute, so no runtime observation can prove its absence. The markers now say so explicitly. Two coverage gaps in no-source-grep surfaced while doing this, recorded on #3464 rather than fixed here, since widening the rule is its own change with its own blast radius: - TEXT_METHODS omits matchAll, so a matchAll source-grep never trips the rule (the case above). - The path test requires a literal quoted bin/lib/gsd-core/src segment and tracks the binding one hop, so a read through a dynamic path or an intermediate variable evades it. install-minimal-hooks' genuinely load-bearing read at line 975 is itself unmarked for a related reason. install-minimal-hooks therefore keeps its markers too: its remaining real source-grep of bin/install.js (the Codex legacy gsd-update-check migration check, line 975) is outside this issue's 8 sites. It is the last blocker for that file and is a clean follow-up. Closes #3466 Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
69e7afd0c7 |
chore(#3212): bounded quantifiers over document content — prohibition with teeth — Phase 4 (#3441)
* feat(#3415): ship local/no-unbounded-quantifier, burn down ReDoS class Phase 4 of epic #3212 (ADR-3212 §5/§7, the final phase). New rule flags an unbounded */+/{n,} quantifier over a broad character class ([\s\S], dotAll ., or a 1-2-unit negated class like [^\n]/[^)\n] — the exact #2128-fixed shape) applied to a regex whose match target is data-flow-traced to readFileSync content. eslint-rules/lib/readfilesync-trace.cjs extracts the data-flow tracer shared with no-crlf-fragile-split (Phase 2) rather than a second copy — no-crlf-fragile-split refactored onto it with zero behavior change, parity-tested. Real triage, not 798 mechanical edits: the ADR's census (2026-08-08) screened every unbounded quantifier in the tree unscoped. Correctly scoped to readFileSync-derived content (matching Phase 2's own G2/G3 scoping), the rule found 162 real hits across two detection waves — the second wave (93) surfaced only after a genuine off-by-one bug in this rule's own first draft was caught while writing its RuleTester tests and fixed (the bug silently missed every directly-quantified [\s\S]* with no gap before the quantifier — exactly the class this rule exists to catch). 3 hits landed in production src/ (commands.cts, milestone.cts, roadmap.cts) and were each empirically timed against adversarial input (matching #2128's own measured-not-assumed precedent) — all confirmed linear-time/benign, left unbounded with a measured-evidence comment rather than mechanically bounded. The remaining 159 are test-file fixture parsing (test-author-controlled, fixed-size content, not adversarial input) — each suppressed with a specific, non-generic reason. Zero functional behavior changed anywhere in this diff. tests/no-pending-3212-markers.test.cjs locks the epic's own closing invariant (ADR §7: "assert zero pending #3212 markers remain") — ground truth confirmed trivially true today (no phase left any such marker behind), now regression-locked going forward. Design: .gsd/phase/chore-3415-prohibition-with-teeth/40-design.md Test matrix: .gsd/phase/chore-3415-prohibition-with-teeth/50-test-matrix.md Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): correct rule category mislabel, add CI test-scope entry An orthogonal Standards-axis review found eslint-rules/no-unbounded-quantifier.cjs mistakenly carried meta.docs.category: 'Portability', copied from a sibling rule without realizing what that implied: docs/contributing/cross-platform- portability-rules.md governs an ADR-1703 rule family under a hard "zero escape hatches" contract (tests/portability-rule-disable-ban.test.cjs's PROTECTED_RULES bans eslint-disable for those rules entirely). This rule is not part of that family — it's ADR-3212 (ReDoS/CWE-1333), a different epic — and its eslint-disable-next-line suppressions (159 of them, added earlier this same phase after empirical benign-verification) are an intentional, correct design, not a bypass. Corrected to category: 'Best Practices', matching the actual precedent (no-adhoc-regex-escape.cjs, Phase 1 of the same epic, which is also correctly outside PROTECTED_RULES), and the rule's own docstring now states this explicitly so a future reader doesn't have to re-derive it. Also registers a new scripts/ci-test-scope.cjs bucket so editing this rule or the shared eslint-rules/lib/readfilesync-trace.cjs helper re-runs their own test suites under targeted CI selection — was previously unregistered and invisible to that fast-path (this PR's own gsd-test checkpoint runs the full suite regardless, so this only affects future narrowly-scoped PRs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): bound no-unbounded-quantifier's own scanner (CWE-1333, ironic) Security review found the rule meant to catch algorithmic-complexity bugs had one of its own: hasUnboundedBroadQuantifier's negated-class inner scan walked from each `[^` occurrence to the next `]` (or EOF) with no bound, while the outer loop only ever advanced by one character — O(n²) total work on a pattern with many unclosed `[^` runs. Runs unconditionally inside checkPattern on any `new RegExp('literal string')` argument in any linted file, before the (cheap) readFileSync data-flow gate — so a single crafted string literal, no valid regex syntax required, could make `npm run lint` / CI hang. Empirically confirmed both the bug and the fix: pre-fix, n=4000/8000/ 16000/32000 chars took 30.8/115.6/463.8/1874.3ms (~4x work per 2x n, quadratic); extrapolated, the 300000-char repro from the finding would run ~165s. Post-fix (bail the inner scan once units exceeds the rule's own 1-2-unit scope, rather than continuing to hunt for a closing `]`), the same 300000-char input runs in 8.7ms via the real rule module, independently reconfirmed at 18ms via a fresh Linter.verify() call. New regression row in tests/no-unbounded-quantifier.rule.test.cjs asserts the RuleTester run on a 50000-char adversarial pattern completes and returns a defined result — no wall-clock assertion (CLAUDE.md Clock Seams / local/no-elapsed-assertion). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): triage 3 new sites, re-raise ceiling after upstream batch next merged 12 more PRs during this PR's review. Two consequences: - tests/edit-phase.test.cjs (fix #3262, unrelated) added 3 new content.match(/<tag>([\s\S]*?)<\/tag>/) reads of this repo's own workflow .md content — the same Class A pattern as the ~159 sites already triaged elsewhere in this PR. Suppressed with the same established reason. - lint-allow-test-rule-refs' ratchet ceiling needed re-raising again (301 -> 303) for the same reason as the two prior bumps: organic growth from unrelated, already-reviewed PRs landing concurrently, not a defect in this branch's own diff. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ad07f76a31 |
test(#3336): fold the installer & runtime surface issue-* cluster — Wave 4 (#3376)
* test(#3336): fold the installer & runtime surface issue-* cluster — Wave 4 Folds 10 legacy issue-*.test.cjs regression files (79 test() blocks) into their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053). First of 4 issue-* waves (following the 3 fix-* waves, all merged). - 1 file with no prior target coverage: renamed (git mv) into legacy-cleanup.test.cjs (sole comprehensive suite for that module). - 9 files merged into 6 pre-existing suites: golden-parity-single-source, runtime-artifact-layout-surface, codex-config (4 sources merged jointly in one pass per the issue's own instruction, to catch overlap between the 4 sources themselves, not just against the pre-existing target — zero overlap found, all 20 blocks additive), runtime-config-adapter-registry (1 of 10 source blocks dropped as a proven subset of existing coverage), cline-install, install.test.cjs. Incidental fixes required to keep this wave's own ratchets green: - Fixed a stale ADR doc reference (docs/adr/1235) to a folded-away filename. - scripts/lint-allow-test-rule-refs: pruned 4 stale allowlist entries for renamed/merged-away files, cited 2 previously-uncited allow-test-rule comments that surfaced as "new" only because their file path changed, added 1 fresh allowlist entry for a pre-existing uncited comment that predates this PR, and tightened the exemption-file ceiling 309 -> 305 to match the real post-fold high-water mark. Zero net test-coverage loss. No production code changed. * test(#3336): fix orthogonal-review findings — Wave 4 fold Standards-axis review + Memtrace graph pass found real issues in the just-folded suites, all fixed here: - Standardized the fold-wrapper convention (block-scoped __foldDescribe) across golden-parity-single-source.test.cjs, runtime-artifact-layout- surface.test.cjs, runtime-config-adapter-registry.test.cjs, and cline-install.test.cjs to match the pattern already used by codex-config.test.cjs and install.test.cjs in this same wave (and by earlier folds elsewhere in the epic) — repeats the exact inconsistency Wave 3 (#3335) already fixed once in this epic. - Fixed a stale allowlist entry's alphabetical position (cosmetic, not tool-gated, caught by review anyway). - Fixed two stale test-filename references in PRODUCTION code comments (src/capability-writer.cts, src/runtime-config-adapter-registry.cts) caught by lint-removed-but-needed — a class of stale reference this wave's fold agents didn't check for, since they were scoped to docs/ and gsd-core/references/ only, not src/. First fix attempt wrongly edited the gitignored gsd-core/bin/lib/*.cjs BUILD OUTPUT instead of the tracked .cts source; caught and corrected before commit. - Fixed one remaining stale doc reference in docs/adr/1235 (a prior partial fix in this same wave missed it). No test() count changed in any file. No production code BEHAVIOR changed — comment-only fixes in src/. --------- Co-authored-by: sim <sim@local> |
||
|
|
2076d450d7 |
fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not the runtime name (#2728)
* fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not runtime name quick.md and diagnose-issues.md kept the pre-#2584 `RUNTIME != "claude"` worktree gate, so every non-Claude runtime failed closed regardless of the capability it negotiated — including Codex, which declares orchestrator-worktree. Route both through the negotiated dispatch.isolation seam via a new shared reference, and migrate the two execute-phase reference fragments that carried the same runtime-name gate. - new gsd-core/references/dispatch-isolation-gate.md: canonical ISOLATION resolution, harness-flag resolution, single-agent degrade rule - quick.md / diagnose-issues.md read the gate; dispatch uses the {harnessFlag} placeholder rather than a hardcoded isolation="worktree" - execute-phase-wave-guard.md / execute-phase-between-wave-reset.md: migrate [ "$RUNTIME" = "claude" ] -> [ "$ISOLATION" = "harness-worktree" ] - every degrade site now clears BOTH USE_WORKTREES and ISOLATION; clearing one dispatched an isolated agent with no base guard and no manifest - parity guard in host-integration.test.cjs scans workflows AND references and matches six reintroduction shapes - migrate four tests that pinned the pre-#2584 runtime-name contract Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): use the /gsd:<cmd> namespace in the isolation degrade messages The degrade warnings cited /gsd-execute-phase, the retired hyphen form that slash-command-namespace.test.cjs rejects in Claude-facing source. Same length, so the quick.md size budget is unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2652): add changeset for PR #2728 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): normalize dispatch-site paths to forward slashes for Windows path.relative() returns backslash-separated paths on Windows, so the #2652 dispatch-site parity test compared "gsd-core\workflows\quick.md" against the hardcoded forward-slash literal "gsd-core/workflows/quick.md" and failed on every windows-latest CI lane. Normalize with .replace(/\\/g, '/'), matching the existing convention used elsewhere in this suite (e.g. tests/branch-no-track-guard.test.cjs:37). * test(#2652): restore the size-growth acknowledgment The rebase dropped tests/emitted-drift-ack.json. #2757/#2758 fixed the ATTRIBUTION axis, but the SIZE-GROWTH axis is independent: diagnose-issues.md (+2086) and quick.md (+230) still need an ack naming them and saying why. Verified: 65/66 without it (both files named), 66/66 with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): convert execute-plan.md Pattern A onto the dispatch-isolation gate Pattern A hardcoded `isolation="worktree"` — Claude Code's own literal — gated only on `workflow.use_worktrees`, with no capability negotiation at all. It is the same defect #2652 fixes at the other four sites, just a different shape: the file contains no RUNTIME variable, so the new detector correctly does not flag it. Concrete break: a Codex user who follows this PR's own newly-documented pattern and sets `workflow.use_worktrees: true` to get isolated dispatch via /gsd:quick then runs a plan through /gsd-execute-plan Pattern A, and hits an unconverted path — either an Agent() call erroring on an unrecognized parameter or silent unisolated execution, depending on host tolerance. Pattern A is a single-agent dispatch site through the host's own subagent tool, so it takes the same treatment as quick.md and diagnose-issues.md: resolve ISOLATION/HARNESS_FLAG through the canonical reference, degrade to sequential on orchestrator-worktree hosts, and substitute the host's declared {harnessFlag} instead of Claude Code's literal. while the area was open. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2652): add the INVENTORY row for dispatch-isolation-gate.md, refresh CONTEXT Two bookkeeping gaps flagged in review: INVENTORY.md had no row for the new gsd-core/references/dispatch-isolation-gate.md. INVENTORY-MANIFEST.json was regenerated correctly and its --check only diffs a live directory scan against the committed manifest, so CI passed regardless — but gen-inventory-manifest.cjs's own stderr guidance says to add the matching INVENTORY.md row. This is the repo's named "Inventory Drift" pattern. Placed with the dispatch/isolation cluster (worktree-branch-check, runtime-aware-dispatch) rather than alphabetically, matching how that table is grouped. CONTEXT.md's Host-Integration Interface entry still described dispatch.isolation as "declared and negotiated but not yet consumed by any scheduler — Phase 1 of #2584". That was already stale before this PR (execute-phase graduated in Phase 3) and more so now with three single-agent dispatch sites consuming it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): detect reversed-operand runtime gates; add a permutation property All five reintroduction regexes assumed $RUNTIME on the LEFT of the comparison, so `[ "claude" != "$RUNTIME" ]` — the same gate written backwards — evaded every one of them. Verified against the old patterns before fixing: all four reversed shapes (single bracket, double bracket, test builtin, JS template) scored EVADED. Each comparison shape is now generated in both operand orders from a single template, so a shape cannot be added in one order and forgotten in the other. The mutation table gains the four reversed cases. Also adds the fast-check property review suggested in place of the hand-rolled cases: it generates the cross product of the axes an author actually varies — bracket form, operator, operand order, quoting, spacing, runtime id — so a permutation the hand-written patterns miss surfaces here rather than in production. The 11 explicit cases stay as named regression anchors. execute-plan.md joins the scan's required-identities list now that it is a converted dispatch site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): acknowledge the execute-plan.md size growth The Pattern A conversion adds 811 bytes to an emitted workflow. Per #2719 the size axis needs its own acknowledgment, independent of attribution. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): repin the execute-plan.md PROSE_ALLOWLIST line after the rebase The #2751 command-position gate pins its prose exemptions by line number. This branch inserts the dispatch-isolation resolution above the `validated downstream by gsd-tools uat classify-coverage` sentence, moving it from execute-plan.md:387 to :397 — which fired the gate twice for one displacement (an un-allowlisted mention at 397, a stale entry at 387). The prose itself is unchanged from next; only the pin moves. Fixes #2652 * fix(#2652): gate the #2649 base-check on ISOLATION in diagnose-issues.md The rebase onto next merged #2649's pre-dispatch base-check textually, but its degrade flipped USE_WORKTREES after ISOLATION was already resolved, so the degrade never reached the dispatch decision. Gate the block on ISOLATION = "harness-worktree" and degrade ISOLATION itself, the same pairing quick.md already uses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): key quick.md post-dispatch bookkeeping on ISOLATION, not the Claude literal Review Blocker: the manifest append (l.822), worktree merge-back (l.825), and its skip clause (l.839) all conditioned on the literal isolation="worktree" — Claude Code's own rendering of {harnessFlag}. Cursor renders --worktree, so a newly-unblocked isolated Cursor run created a worktree whose committed work was never merged back and never cleaned up, silently. All three now key on ISOLATION = "harness-worktree" at dispatch. The existing parity detector cannot catch this class (its ISOLATION_TOKEN treats the literal as a legitimate marker), so this adds a dedicated literal-condition detector with a discrimination proof against both pre-fix sentences, a benign-mention control, and a positive pin on all three re-keyed conditions. Verified fail-first against the pre-fix quick.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): scope the use_worktrees=false install stamp to isolation=none runtimes `_stampNonClaudeRuntimeDefaults` rewrote every non-Claude runtime's `workflow.use_worktrees` read to `--default false`. That default resolved before `gsd_run query dispatch-isolation` was ever consulted, so the five runtimes that declare worktree support — cursor (harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree) — got ISOLATION=none regardless of what they negotiated. The gate this PR migrates dispatch onto was therefore still deciding isolation by runtime name, one layer down. The stamp's #1521 premise was that worktree isolation *was* Claude Code's isolation="worktree" spawn parameter, which no other host honored. #2584 replaced that premise with the negotiated capability. The stamp is now scoped to runtimes whose negotiated isolation really is `none`, where the default it writes is the outcome the resolver reaches anyway. `_negotiatedDispatchIsolation` mirrors routeDispatchIsolation's resolution against the same registry — closed vocabulary, a harness-worktree host must declare its flag, an orchestrator-worktree host must carry a descriptor that resolves — and fails closed to `none` on anything else, so an undeclared or unknown runtime keeps today's behavior. Two #1515 tests pinned the superseded premise for codex and are re-pointed at the new contract rather than deleted: the safety property they protect is now held by the isolation gate's fail-closed resolution, not by a name-scoped install-time default. Verified fail-first — all five assertions red against the pre-fix source, green after. * test(#2652): acknowledge the emitted ripple and re-point the end-to-end stamp proof Scoping the use_worktrees stamp changes emitted output, and two gates caught it. `gsd-core/workflows/execute-phase.md` now differs at emit time for the five hosts that declare worktree support (cursor harness-worktree; codex, opencode, kimi, kimi-code orchestrator-worktree) — the source file is byte-identical, only the stamp is gone. Acknowledged in this PR's fragment. `tests/install.test.cjs`'s real-install assertion pinned the superseded premise end-to-end, asserting codex receives `--default false`. Re-pointed rather than deleted, matching the two unit tests: it now proves codex keeps the unstamped `true` read. A second arm installs windsurf — which declares isolation `none` — and asserts the false stamp is still applied there, so the change cannot silently degrade into "never stamp" without a test noticing. The ack entry collides with `2658-trae-instruction-file-path.json`, which is fully spent (merged via #2925, so all 25 of its entries are present at base and gate nothing) and is pruned for the same reason and by the same rule as the spent `2649-*` fragment this PR already removed. #2566 prunes the same file for the same collision on `new-project.md`; a delete/delete merges cleanly either way, and the base-side cleanup would make both unnecessary. * fix(#2652): re-record the sentinel when a dispatch site degrades isolation Review Blocker B1/B2/B3. Every isolation degrade in a dispatch site is decided in shell, where routeDispatchIsolation cannot see it. That resolver persists whatever it resolved to the run-scoped sentinel as an unconditional side effect (#3045), so a degrade that only reassigns $ISOLATION leaves the sentinel asserting harness-worktree while the dispatch correctly omits the harness flag. The shipped PreToolUse guard reads the sentinel at the instant of the Agent() call and denies that mismatch with exit 2 — the work does not run unisolated, it does not run at all. Latent on this branch and lands on rebase, since |
||
|
|
c28134ab39 |
fix(#3271): delete 25 duplicated folded test suites and fix three runner defects found doing it (#3285)
* test(#3271): guard against a folded suite appearing twice in one host Adds local/no-duplicate-fold-marker, an AST rule that reports the second and every subsequent `folded:<name>` marker in a host file, plus RuleTester cases and a tree-wide regression assertion. Failing-first on purpose: the rule is registered at error and the 25 duplicated regions are still present, so eslint and the new tree-wide test are RED. The deletions land in the next commit. The marker key is the whitespace-delimited token after `folded:` — not the issue's `[a-z0-9-]*` slice, which truncates at `.` and false-positives on tests/model-resolver.test.cjs where feat-443-effort-fast-mode.integration and feat-443-effort-fast-mode are two distinct folded suites. Refs #3271 * fix(#3271): delete 25 duplicated folded suites from three install hosts Three consolidated install suites each carried a verbatim second copy of a contiguous run of #1969 B1 folded blocks. Byte-identical, constant offset, and green — each duplicated block registered and ran twice on every lane. tests/install.test.cjs 5981-9937 (3957 lines, 18 blocks) tests/install-minimal-hooks.test.cjs 2734-4015 (1282 lines, 5 blocks) tests/install-write-confinement.test.cjs 1754-2321 ( 568 lines, 2 blocks) Introduced by |
||
|
|
2e2b8ba4a7 |
enhance(#2704): resolve documentation links and compare H1 status brackets in the ADR gate (#3266)
* test(#2704): failing-first coverage for ADR link resolution and H1 status brackets Binds the gate to two assertions it does not yet make: every relative markdown link under docs/adr/ must resolve, and an H1 trailing status bracket must agree with the Status: field instead of being silently stripped. Covers all 51 rows of the phase test matrix across two altitudes - the pure extractLinks/maskCode IR for fence and inline-code-span boundaries, hostile input and the fast-check totality properties, and the real CLI verdict for the end-to-end classes. Includes the DEFECT.GENERATIVE-FIX parity test that iterates the exported STATUSES array so a sixth status is covered the day it is added. * feat(#2704): resolve ADR documentation links and compare H1 status brackets The ADR gate validated naming, relation symmetry and index freshness but never resolved a link target, and it stripped an ADR's trailing H1 status bracket for display rather than comparing it against that ADR's own Status: field. Both classes were structurally invisible: #2691 found five dangling references by manual audit roughly a year after they were introduced, one of which reached the published npm payload, while CI reported green throughout. Both are now assertions on the same --check path, using only node:fs and node:path - no dependency and no subprocess. Fenced blocks and inline code spans are masked before scanning, because markdown does not render a link inside code. That is not a policy choice: the corpus contains exactly two such sequences today and both are ordinary JavaScript. Masking preserves length and column positions so findings still name a real line. Resolution is case-exact on every platform - a link that resolves only through macOS or Windows case-folding still 404s on github.com and still fails the Linux lane - and a destination resolving outside the repository is reported before any filesystem call is made. Also single-sources two duplicated surfaces this change would otherwise have extended: the H1 bracket vocabulary (a second hand-written copy of STATUSES with nothing asserting agreement, a DEFECT.GENERATIVE-FIX instance) and the docs/adr directory traversal. Two tests added by #2691 that reimplemented link resolution and bracket comparison inside the test file are removed for the same reason; the corpus assertion is now made by running the real gate against the real corpus. * fix(#2704): reject symlinks that leave the repository and linearize code masking Four defects from the isolated adversarial security review, plus one it noted. BLOCKER - a symlink defeated path containment. path.relative(ROOT, abs) is purely lexical, but the case-exact walk then calls readdirSync, which follows symlinks at the OS level: a contributor-committed docs/adr/x -> /etc together with a link through it passed containment and listed the real external directory, and a wrong-case probe echoed a real external filename through the "Did you mean" hint into publicly-readable fork-PR logs. Every segment is now lstat'd before descent; a symlink is realpathed and re-checked against realpath(ROOT) - realpath on both sides, so a root under /var does not produce false escapes - and an escape emits no hint and reads nothing further. The same rule now governs which FILES are read: an ADR entry that is a symlink out of the repository is excluded and reported rather than parsed, closing the vector this change had widened by newly reading README.md, naming-violation files, and full bodies rather than only header fields. MAJOR - inline-span masking rescanned the line remainder per backtick run, roughly O(n^1.6) on adversarial input: 1.76s for an 800KB line. Rewritten as a single linear pass pairing runs through forward-only per-length cursors. Same input now takes 3.31ms, with behavior unchanged. MINOR - an unreadable or broken entry threw, and the generic handler wrote a raw stack trace carrying absolute CI paths to stderr. The scan is now fault-tolerant and reports excluded entries as ordinary violations. The status vocabulary is escaped before being interpolated into a dynamic RegExp - defence-in-depth, not a live bug. The containment predicate had reached three hand-written copies while fixing this; it is now the single escapesRoot() helper used by all four call sites. * feat(#2704): add a --json report so the gate's tests assert on typed values Maintainer-directed addition. CONTRIBUTING.md's "Prohibited: Raw Text Matching on Test Outputs" requires that a system under test producing text also expose a structured intermediate representation, and that tests assert on that IR rather than on rendered prose. This gate had no such surface, so its verdict tests matched on stderr. --json runs exactly the same validation as --check and writes a report to stdout with the same exit code, following the frozen-REASON-enum pattern already used by verify-reapply-patches.cjs. Every violation carries a stable reason code plus the fields a consumer needs, so nothing has to pattern-match an error message. Adding a reason stays three coordinated changes - the enum, the emitting site, and the test locking Object.keys(REASON).sort(). The human output is unchanged, deliberately: a large pre-existing suite asserts on it and migrating that is not this PR's concern. Verified by running the pre-change and post-change scripts against an identical violating corpus and diffing their stderr - character-for-character identical. This PR's own verdict tests now assert on parsed --json. Absence checks improve the most: "no bracket violation" is now a reason-code predicate rather than a negative regex over prose, which could pass for the wrong reason. The security assertions were strengthened rather than translated - no leaked filename may appear in ANY field of the serialized report. Unknown flags are now rejected instead of silently falling through to printing the index. * test(#2704): fix the status-parity fixture and guard hooks/dist before overlay builds Two failures from the matrix run of 79b29909. The status-parity fixture was mine. It built, per status token, an ADR whose H1 bracket and Status field both carried that token - but Superseded carries an obligation beyond the bracket: it must name its successor as a file link and be symmetric with it. The fixture declared a bare Superseded, tripped that unrelated invariant, and the test reported a bracket-parity failure for a reason that had nothing to do with bracket parity. The fixture now satisfies each token's own obligations in both the agreeing and contradicting corpora, derived from the status actually declared rather than special-cased on one name, so a future token carrying obligations is handled rather than silently skipped. The second failure was not mine but is fixed here rather than deferred. mcp-catalog-parity.install.test.cjs hardlinks hooks/dist/* while building its overlay, but hooks/dist is a gitignored build artifact produced only by build:hooks. The suite had no guard, so it passed only when some other suite happened to build it first - an execution-order dependency, which is why it failed on node22 and passed on node24 for identical code. install.test.cjs already documents this exact hazard and guards it. Six behaviorally identical copies of that guard existed across three files. Rather than add a seventh, they are now one canonical tests/helpers/hooks-dist.cjs - idempotent and bounded by the shared BUILD_TIMEOUT_MS class norm - which is the same single-sourcing this PR applies to the ADR gate itself. * docs(#2704): add a how-to for contributors the ADR gate rejects The reference and explanation quadrants were covered by Lifecycle rules 5 and 6, but the task-oriented one was thin: a contributor meets this gate because it failed on their PR, under pressure, and the rules told them what is checked without telling them what to do about it. Adds the command to reproduce the CI failure locally and a message-to-remedy table covering every reason code that can be hit - unresolved target, wrong case with the did-you-mean hint, repository escape, symlinked ADR file, bracket contradiction - plus the backtick escape hatch for illustrative links and the caveat that indented code blocks are not skipped. The table is itself written in backticked inline code, so the gate skips it: the escape hatch demonstrated on the page that documents it. * chore(#2704): backfill changeset PR number pr:0 placeholder replaced with the real PR number now that #3266 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
66a4940d6f | Merge branch 'next' into fix/2665-test-env-base-config-location-vars | ||
|
|
342590c70e |
refactor(#3184): milestone windowing has one owner and a decidable failure signal (#3209)
* test(#3184): failing-first milestone-window single-owner suite Covers the 50 input classes in the phase test matrix: scope classification (genuinely-empty vs truncated vs unscoped vs unreadable), the section-end owner's level boundaries, consumer-output identity per ADR-3180 Decision 4(c), the milestone.complete refusal with negative proof that no directory moved, the version-token boundary defect, drift-guard behavior, and three fast-check properties over document-shaped generators. Committed alone so the remote runner records the failure before the fix lands. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * refactor(#3184): milestone windowing routes through one owner Three copies of the milestone section-end walk lived in roadmap-parser.cts — two distinct computeSectionEnd function nodes plus an inline third in getMilestonePhaseFilter's versionOverride branch. computeMilestoneSectionEnd is now the sole owner and the other two are deleted, not kept in sync by comment. The whole-repo drift guard found what the epic did not: state.cts held three more re-derivations of the same vocabulary — two byte-identical milestone bounding checks carrying a defect neither reported copy has (no boundary after the version token, so v2.0 matched inside v2.0.1), and a milestone-sectioning predicate. All three route through the owner now. A composition-level duplicate appeared inside this change's own first pass: getMilestonePhaseFilter and cmdMilestoneComplete each re-assembled a window out of the owner's primitives, and had already diverged on whether to skip a closed milestone heading. sliceMilestoneWindow is the one composition. Windows now carry the ADR-3180 SCOPE discriminator, so a truncated window is distinguishable from a genuinely empty milestone — those were output-identical, which is the whole failure class. roadmap analyze emits it (#3165), and milestone complete refuses to archive on anything but COMPLETE rather than pass-all moving every phase directory on disk (#3166). The pass-all degrade is preserved where its premise holds: making the filter deny-all would trade a silent over-inclusive answer for a silent under-inclusive one on the read paths that count with it. extractCurrentMilestone keeps its signature — 200+ affected symbols across 41 files and 25 process flows — and is a one-line wrapper over the scoped owner. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): fence-aware phase detection and one heading-selection owner Review fixes from the two orthogonal passes. The blocker: hasPhaseEntries matched ATX phase headings fence-aware via tokenizeHeadings but tested the #2199 bullet form against un-stripped markdown, so a fenced EXAMPLE of the bullet syntax counted as a real phase. A genuinely empty milestone then classified TRUNCATED and milestone complete refused a legitimate archive — a false positive in the destructive direction, worse than the defect this phase set out to fix. Both that path and getMilestonePhaseFilter own pre-existing bullet scan now run on stripFencedCode, since leaving one meant the owner file gave two different answers to the same question. The selection rule — locate, prefer the non-closed heading, else the first — had been written three more times inside the file whose thesis is single ownership. selectMilestoneHeading owns it; all three sites route through it. The copies were behaviorally identical, so this is de-duplication with no observable change, verified by probing that all three paths select the same heading. roadmap analyze emitting a scope no consumer read left #3165's actual symptom alive, so Route 0 in next.md now treats a non-complete scope as scan-failed rather than as a clean empty scan, and the ADR amendment no longer overstates what shipped. Also: the scope refusal moved above the archive-directory create, so a refusal leaves nothing on disk; the versionOverride comment names all four consumers; COMMANDS.md documents the new guard beside its sibling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#2658): exclude the changelog from the malformed-path scan The gate walks every emitted .md/.js/.cjs file in an installed tree and asserts none contains `.claude/.trae/rules` or `.trae/.trae/rules`. CHANGELOG.md ships into that tree, and its #2658 entry quotes both malformed paths while describing the fix that removed them — so the release note documenting the fix trips the fix's own regression test. Red on next before this branch. The installer is correct: a probe over a real --trae --local install found 621 emitted files, exactly one hit, and it was gsd-core/CHANGELOG.md. The scan scope was the defect, not the product. Excluded by exact relative path rather than by loosening the patterns or skipping all markdown — the emitted agent and command markdown is precisely what #2658 was about, so the gate stays strong everywhere it matters. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): regenerate install-tree fixtures for the shared drift scanner scripts/lib/ ships in the npm package and installer, so extracting the shared tree-walk into scripts/lib/drift-scan.cjs adds one path to every runtime's install tree. Regenerated via npm run gen:install-tree; the delta is exactly that one path per fixture. The two drift guards themselves do not ship (scripts/lint-*.cjs is excluded), so only the extracted library moves. This matches the existing scripts/lib/allowlist-ratchet.cjs precedent, which is likewise a lint-only helper carried in the shipped tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): restore the #730 sub-milestone boundary and narrow the refusal The remote runner caught two regressions this branch introduced. Both were mine, and neither review pass found them — only running the existing suite did. The version-token boundary. I replaced locateMilestoneHeadings' \b with (?![\w.-]), reasoning that v2.0 matching inside v2.0.1 was the same defect #2562 fixed in isMilestoneShippedInRoadmap. It is not the same question. A milestone state of v8.0 legitimately selects the '## v8.0-B' sub-milestone section over a closed v8.0-A sibling (#730), and \b is what allows it while the stricter boundary forbids it — nine tests in roadmap-phase-fallback said so. Reverted to \b; the state.cts consolidation is now a straight merge with no behavior change, and the v2.0/v2.0.1 ambiguity is left exactly as it was. The ADR amendment and the design doc no longer claim otherwise. The refusal scope. I refused whenever the window was not COMPLETE, but #3166 is about the TRUNCATED window specifically — the heading is found and the section closes before the phase region, so pass-all archives everything. UNREADABLE and UNSCOPED are pre-existing, legitimately handled states, and refusing on them broke 'handles missing ROADMAP.md gracefully' and three archive tests. Narrowed to TRUNCATED; docs corrected to match. One of the new tests was also wrong: its fixture gave the shipped and current milestones' phases the same numeric id, and the filter matches on that id, so it could not have distinguished the two windows. Fixture corrected to exercise what it claims to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): enumerate drift-scan.cjs for uninstall The installer copies scripts/lib/ wholesale, but uninstall removes an explicit set — deliberately, so a user's own helpers in that directory survive. The extracted drift-scan.cjs was copied in and never enumerated, so it outlived uninstall, left the directory non-empty, and the rmdir that follows failed. Added to GSD_SCRIPTS_LIB_FILES, following allowlist-ratchet.cjs, which is likewise a lint-only helper that ships there and is enumerated. Verified with a real install-then-uninstall into a temp target: scripts/lib/ held exactly the three GSD files and was gone afterwards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): assert install and uninstall agree on scripts/lib and scripts/changeset Found while shipping this phase, and fixed here rather than noted. install() copies scripts/lib/ and scripts/changeset/ into the target WHOLESALE — the comment at the copy site literally says "and any future lib helpers". uninstall() removes them by hardcoded enumeration, deliberately, so a user's own helpers in those directories survive. A wholesale writer paired with an enumerated remover cannot stay in sync by construction: any file added to either directory ships to every user and is then orphaned in their repo forever, since it survives uninstall, leaves the directory non-empty, and the rmdir that follows fails. Nothing reported this. 31,225 tests were green over it. That is the same divergence class this epic exists to delete, sitting in the installer, so it gets the same remedy CLAUDE.md prescribes for it: a parity assertion that fails the moment the two surfaces disagree. The test compares each directory's real contents against its enumeration and names the offending file plus the constant to add it to. Both enumerations are hoisted to module scope and exported, so the test asserts on the actual arrays rather than pattern-matching the installer's source — no allow-test-rule annotation needed. Proven non-vacuous both ways: empty diff on the current tree, correct report when an unenumerated file is injected. scripts/changeset/ turned out to carry the identical defect and is covered too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * chore(#3184): backfill changeset PR number Also narrows the wording to match the shipped behavior: the refusal fires on a truncated window specifically, not on any non-complete scope. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
35df0891af |
fix(#3156): sandbox HOME on raw installer spawns — the one leak no scrub reaches
The strict live-config guard this PR ships went red on CI: the suite creates $HOME/.gsd/defaults.json. Diagnosed rather than suppressed, because the guard is right — this is #2665's class arriving through the one door the scrub set is structurally unable to close. bin/install.js writeNonClaudeDefaults() (#2834) writes path.join(os.homedir(), '.gsd', 'defaults.json') for every non-Claude runtime. os.homedir() consults NO GSD variable, so: - no entry in CONFIG_LOCATION_ENV_KEYS can reach it, however the set is derived; and - blanking GSD_HOME does not reach it either -- a blank GSD_HOME falls back to exactly that homedir(). Only a sandboxed HOME contains it, and HOME is deliberately excluded from TEST_ENV_BASE because blanking it would break far more than it fixed. So the containment belongs per-spawn, which is the discipline the suite already applies by hand -- install-minimal-hooks.test.cjs:576 carries a comment naming this exact hazard for the --codex spawn, while the parameterized --${runtime} spawn 130 lines below it does not. Instance fixed, class open. Rather than add a fifth hand-synced env shape to a PR whose subject is that hand-synced copies drift, this adds ONE export -- installSpawnEnv() in tests/helpers.cjs -- and routes every raw installer spawn through it, including the shared tests/helpers/install-shared.cjs installerEnv(), which every install suite already consumes. Callers passing an explicit { HOME, USERPROFILE } are unaffected: overrides spread last. Census (measured, not reasoned): of the 119 test files that spawn bin/install.js, exactly four wrote into a sandboxed $HOME before this commit -- install, copilot-install, install-minimal-hooks, opencode-plugin-adapter -- and zero do after. Only install.test.cjs sits in CI's targeted lane, which is why ubuntu went red on one file while the macOS full lanes went red on four. Attribution: the leak reproduces unchanged at upstream/next itself, so the defect is base-owned and pre-existing; only the detector is new. The guard found a real leak on next within one run. No new failures: the surviving names under a sandboxed HOME (folded:enh-2380-sync-skills, getGlobalConfigDir (Copilot)) fail at base too, and base additionally fails folded:bug-3288-model-catalog-install-path, which this tree does not. |
||
|
|
466db2f1a3 |
fix(#2665): scrub config-location env on the parent for in-process install()
The blocker the review said decides this PR. These tests call the real installer IN-PROCESS with only HOME/USERPROFILE sandboxed. getGlobalConfigDir is env-first, so an ambient CLAUDE_CONFIG_DIR beats the sandbox and a complete global install -- agents/, commands/, skills/, gsd-core/, manifest, settings -- lands in the developer's live config dir. No child-env scrub can reach it; only clearing the parent's env can. The review named two describe blocks in install.test.cjs. A sweep for the shape found four there (bug #3571 and bug #3288 each appear twice in the file), and the post-suite guard added later in this series found a fifth in runtime-artifact-layout.test.cjs, which the review did not name. All five now save/clear/restore via scrubConfigLocationEnv(). Verified: codebuddy-install, cline-install and codex-config already guard their own config-location var around in-process install(), so the class is closed. Addresses review finding: Blocker 1. |
||
|
|
3f349e551d |
fix(#3024): route sync-skills through the shipped gsd-tools instead of an unshipped install.js (#3195)
* fix(#3024): sync-skills workflow uses gsd-tools query skills-root instead of unshipped install.js The sync-skills workflow Step 2 shelled out to gsd-core/bin/install.js --skills-root, but install.js is not shipped in installed trees (only in the npm tarball root bin/). Every /gsd-update --sync invocation failed with MODULE_NOT_FOUND. Fix: added 'gsd-tools query skills-root <runtime>' subcommand (gsd-tools IS shipped) that calls the same getGlobalSkillsBase function install.js used. Updated the workflow to call gsd_run query skills-root instead of the dead install.js path. Also documented the #3025 verbatim-cp limitation in Step 5 with a workaround. * test(#3024): failing-first guards for the three defects in the adopted fix The cherry-picked commit came from an aborted run that never executed its own tests. Its raw-path assertion fails as written, which is the clearest evidence the work never reached verification. Covers: - --raw must emit a bare path, not JSON (output() takes a third rawValue arg that routeSkillsRoot omits, so the raw branch never fires) - an unknown, empty, whitespace, traversing, or metacharacter-bearing runtime must be rejected, not silently resolved to claude's skills root - sync-skills.md must contain zero references to the unshipped install.js, including the guard's remediation text — the issue's second reported defect - parity across every runtime in the registry, not three hardcoded ones, so the two entry points cannot drift Also converts the adopted tests off a hand-rolled spawnSync onto the bounded process seam, per CONTRIBUTING. Fails before the fix. Verified via the remote runner. * fix(#3024): make the skills-root query actually work and reach non-Claude runtimes The cherry-picked commit never ran its own tests. Six defects, all fixed here. --raw was ignored: output() is output(result, raw, rawValue) and the third argument was omitted, so the raw branch never fired and the workflow captured a JSON blob as SRC_SKILLS_ROOT. Every downstream cp -r then resolved against a nonexistent path — the command would have shipped still broken. An unknown runtime silently resolved to claude's skills root, because getGlobalSkillsBase falls back rather than returning null, leaving the existing === null guard dead. The runtime id is now validated at the CLI boundary against the shipped registry, so a typo'd --from/--to fails instead of reading from or writing into the wrong runtime's tree. getGlobalSkillsBase('vscode') threw a raw TypeError. vscode is non-installable by descriptor, so it has no skills root — null is the answer, not a crash. The resolver now short-circuits configHome.kind 'none', which also fixes the same latent crash in install.js --skills-root vscode. Every caller already gates on === null. sync-skills.md used gsd_run WITHOUT the canonical launcher preamble, so gsd_run was undefined on non-Claude runtimes — the fix would have been dead in exactly the place the original bug bit. Preamble propagated via sync-runtime-launcher. Also registers skills-root in TOP_LEVEL_USAGE (the help/dispatch parity guard caught it), removes the last two install.js references including the guard's remediation text (the issue's second reported defect), and updates the stale assertion that still described the removed contract. Verified on the remote runner. * fix(#3024): align the documented runtime list with the registry and gate both entry points Isolated review returned BLOCK on two findings. The workflow's Supported-runtimes list and its --to all expansion named grok and gemini, neither of which is a registered runtime. Once this branch added validation, --to all — a documented first-class feature — aborted. The list was hand-copied prose shadowing the registry, so correcting it alone would drift again; a parity assertion now fails in BOTH directions if the doc and the registry disagree. vscode is excluded by name: it is installSurface 'none', so syncing skills to it is meaningless and would abort. bin/install.js --skills-root reached getGlobalSkillsBase with no own-property gate, so --skills-root __proto__ silently resolved to claude's skills root. This branch had just hardened the OTHER entry point to the same function; leaving one of two parallel surfaces open is the same divergence class as the first finding. Both now call one shared isRegisteredRuntimeId() rather than a copied check, and the parity test covers the hostile ids so the two can never disagree again. Also guards the workflow's root resolution: neither command substitution checked its exit status and only the source had an existence guard, so a failed destination resolution left DEST_ROOT empty and turned rm -rf "$DEST_ROOT/$SKILL" into an absolute path at filesystem root. Both resolutions are now checked, and Step 5 requires both roots to be non-empty and absolute before any destructive command. Verified on the remote runner. * test(#3024): anchor the runtime-list parity extractor to the list span The extractor captured (.+) to end of line, so it swallowed the em-dash prose that explains the vscode exclusion — and that sentence contains backticked `runtimes` and `null`, which is where the three phantom ids came from. The documented list was correct; the test was reading its own explanation back as data. Anchored to the id-list span. Both directions still fail as intended: proven by injecting a bogus id and by removing a registered one. * test(#3024): anchor the --to all extractor and fail loudly on empty captures The workflow has three TO_RUNTIMES= assignments and the regex matched the first one — an empty array initializer at line 28 — so the extractor captured nothing and the assertion diffed [] against 18 ids as if that were data. That is the same failure twice, so the fix is the general one: every extractor in this test now asserts it captured a plausible list before comparing, naming which extractor found nothing and what it was looking for. An extractor that silently yields [] is a confident wrong answer, and a parity guard that reports it as a data mismatch teaches the reader to loosen the assertion. Verified against the real workflow and against doctored copies with each target construct removed, plus both teeth directions. * fix(#3024): merge duplicate process-seam import after rebase The rebase applied cleanly but left runNode declared twice: next had gained its own import of the seam while this branch added one carrying OUTCOME. A clean rebase is not a correct one — the file no longer parsed. Merged into a single import providing both. * fix(#3024): bind DEST_ROOT per destination instead of a dangling map Step 2 stored each destination's root into DEST_SKILLS_ROOTS, which nothing ever read, while Steps 3 and 5 used a scalar DEST_ROOT that nothing ever assigned. The array was also never declare -A'd, so on bash 3.2 — macOS system bash, which this repo supports — every destination collapsed onto index 0. The absolute-path guard added earlier was the only thing standing between that and rm -rf "/$SKILL"; it turned a silent disaster into a hard stop, but the feature still could not complete. Each destination now binds its own DEST_ROOT where it is used, and the unread map is gone rather than replaced. Step 2 keeps eager validation, so a bad runtime id in a multi-destination --to aborts before any destination is written rather than after some already have been. Verified on bash 3.2 with a two-destination run binding distinct roots, and with a bad id aborting before any destructive call. * fix(#3024): restore grok support broken by the registry gate The registry gate added earlier rejected grok, and that was my error. I confirmed grok was absent from the capability registry and concluded the hardcoded branch was dead — without checking what it resolved to. It resolves to ~/.agents/skills, a real grok-specific path, exactly as the pre-fix workflow documented ('grok uses the ~/.agents layout'), and there is a support discussion doc for it. So a working, documented runtime silently lost --skills-root and sync-skills support as a side effect of prototype-pollution hardening — and the parity test I added locked that in as correct. gemini is the one that really was dead: it fell through to CLAUDE's skills root, so rejecting it is right and it stays rejected, as do bogus ids, __proto__, empty, whitespace and traversal. The validator's real question is 'does this id have a genuine runtime-specific resolution', not 'is it in the registry map'. Registry membership was a proxy that happened to miss grok. Legacy non-registry runtimes with dedicated resolution branches are now a named, documented set; enumerating every hardcoded branch in getGlobalConfigDir against the registry confirms grok is the only one. The new tests assert grok resolves UNDER .agents and specifically not to claude's root. Allow-listing an id proves nothing about whether it resolves correctly — that assertion is what would have caught my mistake. Also uses the shared PROBE_TIMEOUT_MS instead of a duplicate literal, and guards Step 3's DEST_ROOT re-resolution, which contradicted the file's own stated guarantee. Verified on the remote runner. * test(#3024): guard against LEGACY_NON_REGISTRY_RUNTIME_IDS drifting The named legacy set is a second hand-maintained proxy for the same predicate the registry check got wrong — 'does this id resolve runtime-specifically'. Nothing stopped a third hardcoded branch being added to getGlobalConfigDir without updating the Set, reproducing the exact class of bug that broke grok. Production stays explicit and greppable; the test derives the truth instead. It resolves a sentinel id to learn the generic fallback, classifies every candidate against it, and fails in both directions — an id resolving runtime-specifically that is in neither the registry nor the Set, or a Set entry that no longer earns its exemption. The failure message names the remedy. Confirms grok resolves runtime-specifically and gemini does not, which is the distinction the original registry check could not see. Also reverts the shared-timeout swap: SKILLS_ROOT_PROBE_TIMEOUT_MS is pre-existing on next and arrived by rebase, so changing it here was scope creep into another issue's territory. Verified on the remote runner. * chore(#3024): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
3fac6e629f |
test(#3145): bound the installer/runtime cluster onto the process seam (#3176)
* test(#3145): bound the installer/runtime cluster onto the process seam Migrates 156 unbounded sync spawn sites across 47 files. Allowlist 120 to 73. Timeouts are sized from evidence already in the tree rather than a house default, because this wave spawns installers rather than git plumbing and an undersized bound does not catch a hang -- it manufactures CI flake, which is worse, since a flake gets re-run instead of investigated. install.test.cjs records a real spawnSync ETIMEDOUT at a 60000ms cap on a loaded bench while another lane passed the same commit in 12.7s, so full installs are bound at 120000ms against that recorded incident. Also adds an auditable escape to the guard's timeout ceiling. The 600000ms cap was set in #3143 from partial evidence, but fragment-single-edit- propagation carries a documented, load-tested 900000ms bound on a run that chains a full build plus eight generators -- the guard would have rejected a correct timeout the moment that file left the allowlist. A value above the ceiling is now permitted only with an inline allow-spawn-timeout-ceiling marker carrying a non-empty reason. It raises the ceiling; it never waives the requirement for a bound, which is asserted directly. install-shared.cjs keeps its hand-rolled assert rather than routing through throwIfFailed: its message embeds both streams, and throwIfFailed carries only a trimmed stderr. The message now also names the outcome, so a bounded timeout reads as such across its 38 importers instead of as expected null to equal 0. * test(#3145): extract class-norm timeouts and correct the build-hooks sizing A pre-PR review found 52 copies of four class-norm timeout constants across this wave. These are not per-suite fixture bindings -- they are shared facts about how long a class of subprocess takes, derived from a recorded bench incident. That norm already moved once (60000 to 120000 after a real ETIMEDOUT), and 52 copies would have drifted the next time it moved. Extracts tests/helpers/timeouts.cjs, where each norm is justified once, and converts the copies. A site that genuinely differs -- a real tsc compile, or regen:derived -- keeps its own local constant with its own justification. Also corrects a misclassification: scripts/build-hooks.js was sized as a build at 120000 in twelve places and 60000 in another, but it compiles and bundles nothing. Its own header says no bundling needed; it copies pre-built files and syntax-checks them with vm. Three different values bounded one script; now there is one. * test(#3145): fix red CI — lint self-match and a Windows chunk overrun Two failures on PR 3176. lint-allow-test-rule-refs read a RuleTester fixture as a real exemption. The fixture exists to prove an unrelated marker does NOT suppress the rule, so it carries that marker's literal text as test data. Split via concatenation, the same idiom no-unbounded-spawn-allowlist.test.cjs already uses for its own self-match problem. The explanatory comment needed the same treatment. The Windows shard 3/3 chunk was killed at its 600000ms budget. Output stopped seven minutes before the kill, so this was an overrun rather than a slow chunk: regenDerivedPropagatesSingleFragmentEditWithNoSecondSourceSurface runs regen:derived bounded at 900000ms, which is larger than the whole chunk budget, so the chunk killer always fires first and it can never complete there. Both the test and that bound predate this change; modifying the file pulled it into the Windows targeted set and exposed it. Skipped on Windows with the reason recorded; the Linux lanes cover it. The 900000 bound and its ceiling marker are unchanged -- they are correct. * test(#3145): refresh the stale test-timings cost table The Windows shard was killed at its 600000ms per-chunk budget. run-tests.cjs packs chunks by measured duration from tests/test-timings.json, and an unknown file falls back to the table's median weight -- advisory by design, but it silently underweights exactly the files that matter. Four of the failing chunk's 22 files were absent from the table, including the two heaviest: fragment-single-edit-propagation.install.test.cjs at 230s (it runs regen:derived) and agent-fragments-emission.install.test.cjs at 79s. Both were weighted as average, so the chunk's total weight read 53.68 against a budget of 60 and the packer produced a single chunk. Regenerated from a passing full-suite run, per the remedy the script itself documents. 700 to 770 entries, 70 added, 0 dropped -- verified, since gen-test-timings.cjs replaces the table wholesale rather than merging. Proven against the real packer: the same 22 files now weigh 103.91 and split into two chunks. No logic, budget, or timeout was changed; raising a budget to make a red gate pass is not a fix. --------- Co-authored-by: sim <sim@local> |
||
|
|
60cf18999b |
fix(#3026): document --pi and --gemini in installer --help (#3112)
* fix(#3026): document --pi and --gemini in installer --help --help documented 16 runtime flags but the installer accepts 18: --pi (named only in banner prose) and --gemini (invisible everywhere). Both install correctly when passed but are undiscoverable via --help. Added both to the Options section. Added a parity test that asserts every accepted runtime flag appears in --help output, with an exclusion set for legacy aliases (--both, --kimi-code). * chore(#3026): backfill changeset PR number 3112 --------- Co-authored-by: sim <sim@local> |
||
|
|
7dd9e59f6b |
test(#3090): stop exempting violations under categories that do not fit
An allow-test-rule annotation citing a category that does not apply is worse than no annotation, because it reads as reviewed. Eight were confirmed by reading the assertions each one covered, and auditing the rest found five more plus one refutation — a converter test whose wording described the wrong mechanism while the covered assertion genuinely was deployed-text. The instructive one used the CANONICAL string for the same mistake: STATE.md command output labelled as a deployed artifact. A canonical string is not evidence the category fits, which is why normalising strings alone would have laundered the problem rather than fixed it. Every mapping the audit had inferred rather than code-verified was spot-checked before rewriting, and the ones that turned out not to fit were re-annotated rather than relabelled. Fourteen STATE.md assertions had a typed extractor available all along and now use it; their annotations came out because nothing needs exempting. Eight assertions genuinely need a production change first — CLI stdout and stderr with no structured mode — and are tagged pending-migration-to-typed-ir citing #3090, which is what that category is for. It had zero real uses before this, while one file carried a real citation to migration issue #2974 under a non-canonical tag. Six annotations covered assertions that do no text matching at all. An exemption for a violation that does not exist is noise that makes the real ones harder to audit; those are removed. atomic-write-coverage gains the annotation it always warranted — its own docstring describes a structural-regression-guard while the file carried none. Fifty-nine non-canonical strings across roughly thirty files are normalised, and the allow-test-rule allowlist is regenerated to match. 472 annotations became 463: every one now uses a canonical category, and the two remaining non-canonical strings are ESLint RuleTester fixtures, not annotations. Refs #3057 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c87f6f358e |
enhance(#1854): offer restore for user-added files backed up on update (#2679)
* test(#1854): failing-first coverage for user-files-backup restore Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#1854): offer restore for user-added files backed up on update Adds a restore-custom-files gsd-tools verb and wires it into update.md as a restore_custom_files step: plan, compatibility-check against the newly installed release, then restore only on explicit opt-in. The backup is never deleted, a shipped path is never overwritten, and a single unwritable entry does not abort the rest. Also drops the jq pipe from update-context field extraction (#2589 class, missed by that sweep) and repairs a broken code fence in docs/CLI-TOOLS.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): reject symlinked restore destinations and backup roots Self-review of the restore path found two write-through holes: copyFileSync follows a symlinked destination, so a link planted at the restore target wrote outside the config dir with every ancestor still a real directory; and statSync on the backup root followed a link, letting the walk read arbitrary files and present them as the user's own backup. Both now lstat. Also marks the report's path/detail strings as untrusted data in update.md so the rendered step cannot carry instructions into the runtime model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): move the update-context jq guard into the #2589 sweep update.md joins the AUDITED list rather than carrying a duplicate assertion in the backup-restore suite, and the guard gains a negative-proof companion so 'no jq pipe' cannot pass by the fields simply no longer being read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): validate manifest files map shape before trusting it Security review flagged that Object.keys on a non-plain-object files field yields numeric-index keys matching nothing, so the managed-path check dies silently while manifest_found still reports true. Shape, not just type (ADR-227): an array or scalar files map is now an unusable manifest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): size the restore prompt by eligible_count Spec review found the prompt was driven by entries.length, so a backup holding only blocked entries asked "Restore 1 file(s)?" when accepting would restore zero. The question now reads eligible_count, and an all-blocked backup reports its reasons instead of offering a choice that cannot be honored. The decline path names the resolved backup_dir rather than the bare directory name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): use t.skip on hosts without symlink support A bare return in a node:test body registers as a PASS, so the four symlink guards silently reported green on unprivileged Windows instead of skipping. Adds the dangling-link destination case the security review called out, and moves outside-dir teardown to t.after so a failing assert cannot leak it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): unfence restore hint, regen goldens, widen install timeout Three gate failures from the c99d612a5 run, all root-caused: 1. capability-registry (3): update.md's decline message put an instructional 'gsd-tools ...' line in an UNTAGGED fence, and the guard treats untagged fences as shell blocks. Retagged both display blocks as text and switched the hint to the resolved 'node <config-dir>/.../gsd-tools.cjs' form users can actually paste. 2. golden-install-parity (19): update.md and gsd-tools.cjs ship, so every runtime fixture moved. Regenerated; the diff is exactly those two hashes per fixture, no other drift. 3. install.test.cjs (5): one real failure, four cascades. The Cursor suite's before hook died on 'spawnSync ETIMEDOUT' at the 60s cap while the node22 lane passed the SAME commit in 12.7s. A full install measures 13-30s idle, so 60s was under 2x headroom and shrinks with every file added to the payload. Raised to 120s, matching the heavy case already in this file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1854): backfill changeset pr number to 2679 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bf8f320083 |
feat(#2505): Phase 1 — EoS descriptor split (kimi-code capability.json + drift-guard registration) (#2519)
* feat(#2454): add kimi-code as an EoS capability (Node Kimi Code CLI) PR 1 of N for #2454. Establishes the EoS descriptor foundation for splitting GSD's kimi support into two distinct products per the user's directive: - kimi (existing): Moonshot's Python kimi-cli (~/.kimi, runtime: python) - kimi-code (new): Moonshot's Node Kimi Code CLI (~/.kimi-code, runtime: node, KIMI_CODE_HOME env) Per ADR-1239 EoS, runtime behavior is driven by capabilities/<id>/capability.json descriptors, not hardcoded branches in install.js. The new descriptor uses the existing primitives (dot-home configHome, skills artifactLayout, kimi-hooks-toml hooksSurface — same TOML [[hooks]] format Kimi Code reads per its docs). Critical Kimi Code constraint reflected in the descriptor: hostIntegration.dispatch.namedDispatch: false hostIntegration.dispatch.builtInSubagents: ['coder', 'explore', 'plan'] hostBehaviors.namedSubagentsSupported: false Kimi Code's official docs confirm only 3 built-in subagents with NO custom- subagent registration (the [subagent] table only has timeout_ms). The kimi-agents YAML layout (used by Python kimi-cli) is therefore NOT in kimi-code's artifactLayout. Schema adjustments: - subagentToolkit set to 'undocumented' (the existing escape hatch); the schema enum (full/read-only) lacks a 'limited'/'built-in-only' value. A follow-up PR can extend the schema enum to add 'built-in-only' as a first-class axis value reflecting Kimi Code's documented model. Registration: - capabilities/kimi-code/capability.json (new descriptor, modeled on codex) - bin/install.js: allRuntimes array + --all list + --kimi-code flag - gsd-core/bin/shared/runtime-aliases.manifest.json: kimi-code aliases (kimi-code, kimicode, kimi_code) - src/runtime-name-policy.cts: FALLBACK_ALIASES map - gsd-core/bin/lib/capability-registry.cjs: regenerated via scripts/gen-capability-registry.cjs --write Tests: - tests/multi-runtime-select.test.cjs updated for the new runtime count (18) + new --kimi-code flag test + 'All' shortcut renumbered 18 → 19. Out of scope for PR 1 (follow-up PRs in the sequence): - Install-time decision logic (kimi vs kimi-code detection / prompt) - agent-install-check semantics for kimi-code (verify Agent Skills presence) - cmdAgentSkills fallback returning subagent prompt content - Workflow template mapping (named agents → built-in coder/explore/plan) - Migration guidance for users currently on 'kimi' who are actually on Kimi Code - Schema enum extension for subagentToolkit: 'built-in-only' Refs #2454, #2095 (EoS/kimi migration epic), ADR-1239 (EoS). * fix(#2454): complete drift-guard registrations for kimi-code runtime The drift guards caught every surface that pins runtime enumeration. Each update is mechanical, driven by the guard's named failure mode: - src/runtime-name-policy.cts RUNTIME_LABELS: 'Kimi Code' label for kimi-code - src/runtime-name-policy.cts RUNTIME_FLAG_IDS: add kimi-code to the isKimiCode predicate generator - bin/install.js runtimeMap: option '11' → 'kimi-code', renumber downstream entries (11..17 → 12..18), ALL_RUNTIMES_OPTION 18 → 19 - gsd-core/bin/shared/model-catalog.json runtimeTierDefaults: kimi-code entry (null/null/null — same as kimi, no model tier defaults until configured) - docs/reference/capability-matrix.md: regenerated via scripts/gen-capability-matrix.cjs --write (kimi-code row added) - tests/global-config-home-fragment.test.cjs GOLDEN_FRAGMENT_MAP: kimi-code → '.kimi-code' - tests/fixtures/golden-install-parity/*.json: regenerated via npm run gen:golden (the runtime-aliases.manifest.json hash changed; all 17 runtime fixtures updated) The capability-registry is already regenerated from the prior commit. * test(#2454): update drift-guard tests for kimi-code runtime registration Multiple drift guards pin runtime enumeration counts and option numbering. Each update is mechanical, driven by the guard's named failure mode: - tests/runtime-flags.test.cjs: EXPECTED_FLAGS gains isKimiCode (16 → 17); 'all 16 flags' → 'all 17 flags' in test names + messages. - tests/multi-runtime-select.test.cjs: parseRuntimeInput option renumbering cascade — kilo moves 11→12, opencode 12→13, pi 13→14, qwen 14→15, trae 15→16, windsurf 16→17, zcode 17→18, All 18→19. New single-choice test for kimi-code (option 11). Prompt test updated for new numbering. - tests/host-integration-descriptors.test.cjs: EXPECTED_PROFILES gains kimi-code → 'programmatic-cli' (terminal CLI per Kimi Code docs); EXPECTED_FLATTEN gains kimi-code → false (backgroundDispatch:true per docs, same as Python kimi/opencode). - tests/global-config-home-fragment.test.cjs: table-count test renamed 13 → 14 table runtimes (kimi-code added to GOLDEN_FRAGMENT_MAP earlier). * fix(#2454): empty artifactLayout for kimi-code (PR 1 scope) The skills kind requires a converter (existing converters are per-runtime like convertClaudeCommandToKimiSkill). PR 1 of this multi-PR sequence only registers the descriptor; the actual Agent Skills converter (and a new 'convertClaudeCommandToKimiCodeSkill' function) lands in PR 2 alongside the install-time decision logic. Empty artifactLayout.global is valid and means 'nothing to install yet via the layout seam'. Also: added kimi-code to RUNTIME_META in tests/helpers/install-shared.cjs (localDir .kimi-code, globalSuffix .kimi-code), and added Kimi Code as option 11 in install.js's buildRuntimePromptText (renumbered downstream options 11..17 → 12..18, All 18 → 19). * fix(#2454): camelCase runtimeFlags for hyphenated ids (kimi-code → isKimiCode) The runtimeFlags generator previously produced 'isKimi-code' (hyphen preserved) for the new kimi-code runtime id. Property names with hyphens are awkward for consumers (flags['isKimi-code'] instead of flags.isKimiCode). The new runtimeIdToFlagName helper folds -[a-z] boundaries to uppercase, producing the conventional PascalCase flag name. The 16 prior single-word runtime ids are unaffected (the regex finds no hyphens). * fix(#2454): update remaining drift-guard tests + gen kimi-code fixtures - tests/runtime-flags.test.cjs drift guard: use proper kebab-case conversion (isKimiCode → kimi-code, not 'kimicode') so the registry comparison doesn't false-positive on hyphenated runtime ids. - tests/multi-runtime-select.test.cjs: fix kilo/opencode/pi/qwen/trae single-choice tests for the renumbered options (kilo 11→12, opencode 12→13, pi 13→14, qwen 14→15, trae 15→16). - tests/install.test.cjs: Kilo integration option 11→12, prompt test regex updated. - tests/fixtures/golden-install-parity/kimi-code.json + install-tree/ kimi-code.json: generated via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1. The kimi-code install produces the standard GSD install layout (skills, contexts, references, etc.) — 436 paths, same shape as other runtimes that have no custom converter yet. * fix(#2454): add kimi-code install contract + global config home fragment - src/runtime-name-policy.cts GLOBAL_CONFIG_HOME_FRAGMENTS: add kimi-code → '.kimi-code' so getGlobalConfigHomeFragment returns the correct path instead of falling through to the default '.claude'. - tests/installer-migration-install.integration.test.cjs RUNTIME_INSTALL_CONTRACTS: kimi-code entry (same surface as kimi for PR 1; PR 2 will specialize once the Agent Skills converter lands). - tests/multi-runtime-select.test.cjs: fix space-separated-choices test for the renumbered kilo option (11 → 12). - tests/fixtures/golden-install-parity/kimi-code.json + install-tree/ kimi-code.json: regenerated after rebasing onto current next (new planner-reversibility.md from #2471 etc. now included). * test(#2454): skip kimi-code install contract until PR 2 ships install layout The end-to-end install test (tests/installer-migration-install.integration .test.cjs) asserts every allRuntimes entry installs a runtime-specific artifact surface. PR 1 of #2454 registers kimi-code in allRuntimes + the capability descriptor + flags + labels, but the install LAYOUT (Agent Skills converter + global AGENTS.md at $KIMI_CODE_HOME/AGENTS.md) lands in PR 2. The SKIP_INSTALL_CONTRACT set marks this exclusion explicit and self-removing — PR 2 removes the entry alongside adding the install surface, restoring the contract loop to full coverage. * fix(#2454): restore compact model-catalog.json format (M1 review) Per code-review M1: my prior 'fix(#2454): complete drift-guard registrations' commit used python json.dump(indent=2) which inflated the file from 165→607 lines (every nested entry got expanded) and lost the trailing newline. The semantic change was just a 3-line kimi-code entry. Restored the original hybrid format (top-level indent=2 + inner entries' one-line style) and added kimi-code in matching form. Regenerated golden install parity + install tree fixtures since the model-catalog.json hash changed. * fix(#2454): update CONTEXT.md allRuntimes glossary (17 → 18, add kimi-code) CI lint-tests job failed on the glossary drift guard (scripts/check-glossary-refs.cjs --check): ✗ CONTEXT.md's allRuntimes enum-count sentence claims 17 values but bin/install.js's allRuntimes array has 18. ✗ CONTEXT.md's allRuntimes member list has drifted from bin/install.js (missing from CONTEXT.md's list: kimi-code). Missed in the prior commits because gsd-test does not run the glossary check (it's a CI lint-tests-only check). Updating CONTEXT.md's two claims to 18 values + kimi-code in the member list. * chore(#2505): regen capability-registry + stamp kimi-code version 1.8.0 (#2511) * docs(changeset): Phase 1 kimi-code runtime Added (#2511) * test(#2511): regen kimi-code golden parity fixture after Phase 0 guard normalization lands * docs(changeset): backfill PR #2519 for Phase 1 (#2511) |
||
|
|
e227e81f59 |
fix(#2395): persist runtime identity into ~/.gsd/defaults.json for non-Claude installs (#2446)
* fix(#2395): persist runtime identity into ~/.gsd/defaults.json for non-Claude installs Bug: finishInstall for non-Claude runtimes never persisted a 'runtime' key into ~/.gsd/defaults.json. resolveRuntime() precedence is GSD_RUNTIME env > config.runtime > 'claude', so both inputs being empty on a Cursor (or any non-Claude) install caused agent_runtime and every runtime-branded slash hint to silently fall through to 'claude'. Cursor users saw 'agent_runtime: "claude"' and Claude-formatted /gsd-* hints with no env or config hand-set. Fix mirrors the existing resolve_model_ids: 'omit' write site at the same call site (bin/install.js finishInstall, gated on !_hostBehaviors(runtime) .nativeModelAliases && !GSD_TEST_MODE). Writes runtime: <runtime> into ~/.gsd/defaults.json when absent/null/empty. Claude is the resolveRuntime() fallback so it needs no write; an explicit pre-existing runtime value is always preserved across installs of any runtime. Pattern parity with #1156 (default-to-omit intent) and #1569 (preserve explicit user value) — the new write is the third sibling on the same install-time persistence block. Regression tests in tests/install.test.cjs cover: - absent / null / empty-string runtime → populated to <install runtime> - explicit pre-existing runtime preserved (no clobber across runtimes) - parameterized across 5 non-Claude runtimes (cursor, codex, opencode, antigravity, windsurf) Out of scope (per triage): the resolveRuntime() precedence order itself (env > config > default) is unchanged. A separate follow-up noted in the issue (subagent rendering when runtime correctly identifies as 'cursor') is unrelated to branding and not addressed here. * chore(#2395): regenerate golden-install-parity fixtures for new runtime persistence The golden install-tree fixtures capture the post-install state, including ~/.gsd/defaults.json. For non-Claude runtimes, defaults.json now includes runtime: <runtime> — content hash updated for each affected runtime (17 golden-install-parity/*.json files). Claude's defaults.json is unchanged (no runtime key written for Claude — it's the resolveRuntime() fallback). * chore(#2395): drop product names from changeset fragment (product-name-purity gate) * test(#2395): move describe out of fix-1521 fold + add same-runtime idempotence test Code review (correctness subagent) flagged 2 Low test-quality issues: 1. The new Bug #2395 describe was inserted inside the folded:fix-1521-real-install-stamping IIFE callback, muddying test reporting and tracing the regression to the wrong epic. Moved it outside the IIFE close to be a top-level sibling. 2. No explicit same-runtime idempotence test — the suite covered cross-runtime preservation (cursor seed → opencode install) but not 'install cursor twice → second is a no-op'. Added: seeds fresh defaults, installs cursor, captures mtime, installs cursor again, asserts runtime unchanged AND defaults.json mtime unchanged (idempotent, no rewrite churn). * chore(changeset): backfill pr:2446 in .changeset/fierce-ravens-dance.md |
||
|
|
dcb4954131 |
fix(#2335): normalize volta node image paths to the stable shim (#2375)
Adds a volta branch to normalizeNodePath() that rewrites the version-pinned node image path to volta's stable shim, so managed hooks survive a volta node prune (the fnm/Homebrew/mise class, now covered for volta). Also unifies the release-smoke install timeout into one shared 600s constant across before() and runSmoke() so slow benches no longer spuriously time out. Closes #2335. Admin-merged (self-review bypass) with full green CI. |
||
|
|
75bedf16fd |
fix(#2284): project Hermes named dispatch onto delegate_task; protect comparison tables (#2309)
Hermes installs brand-swapped "Claude Code" -> "Hermes Agent" in shipped workflows/*.md but never projected the Agent(...) dispatch calls onto Hermes's delegate_task contract, so installed workflows kept literal Agent(...) syntax and falsely asserted "The Agent tool IS available" (Hermes exposes delegate_task, not Agent). Dispatch projection: a generic named-dispatch engine (projectNamedDispatchToStructuralDelegate) wired into the per-runtime RUNTIME_CONTENT_DISPATCH.hermes.md converter, branching entirely on the documentation-sourced hostIntegration.dispatch facts read via _hostIntegrationDispatch (capability.json unchanged): namedDispatch:false -> resolve the gsd-* role and embed a load-its-prompt instruction in the payload; background:true -> map onto delegate_task background; read-only / maxDepth:1 -> no nested delegation to leaf roles; per-call model dropped. Span detection uses literal Agent( scanning + local balanced paren/quote matching (immune to upstream document quote imbalance) and handles all three corpus call forms (multi-line, object-literal, single-line compact). An independent, mask-free post-projection guard fails the install loud on any residual Agent(/subagent_type/leaked model. Fail-closed: install throws if a literal gsd-* role reference cannot be resolved. commands -> skill path untouched. No literal Agent( survives in installed Hermes workflows. Folded in (maintainer-directed) a pre-existing cross-cutting branding defect: the "Claude Code" -> brand swap corrupted <runtime_compatibility> comparison tables (where "Claude Code" is a compared-runtime label) for every branding runtime. New shared applyClaudeCodeBrandSwap helper protects <runtime_compatibility> regions via split-and-rejoin (no sentinel token) while still rebranding genuine self-references; adopted by all six branding .md converters. Also a surgical prose-consistency fix so plan-review-convergence.md's dispatch-adjacent terminology is coherent post-projection (no broad bare-word rename). Golden install-parity regenerated for the six branding runtimes (dispatch/branding scope only); other runtimes unchanged. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
dae8e70a5a |
test(#2185): cover Linuxbrew + custom-prefix Cellar paths in normalizeNodePath
Adds the Linuxbrew Cellar layout (/home/linuxbrew/.linuxbrew/Cellar/...), a post-version-bump path, a node@version formula, and a custom HOMEBREW_PREFIX — all mapping to the stable <prefix>/bin/node symlink. Mirrored in both consolidated describe blocks. |
||
|
|
6e773d97df |
feat(#2088): migrate Codex onto the Embeddable Orchestration System (ADR-1239)
Drive Codex install/uninstall through the descriptor-driven Host-Integration Interface (declarative embedding adapter → engine surface dispatch) and fold every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven `runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated (tests/fixtures/golden-install-parity/codex.json); no other runtime changes. Three Context7-verified upgrades, each with a test on the user-reachable surface: - Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override, with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest, and the skill-manifest inventory to honor the override so --skills-root / sync-skills / the manifest report the real location. - Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact, PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall; extendedHookEvents reconciled [] -> the schema-valid wired subset. - Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own AgentsToml scalars (max_threads etc.) instead of dropping them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8f2ebbe9bf |
feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap). --gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): backfill changeset PR number (#1996) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): drop Gemini CLI from issue templates (review nit) Removes the sunset Gemini CLI runtime from the two GitHub issue-template runtime lists that the removal PR missed, per @davesienkowski's review nit: - feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise request a feature for a runtime GSD no longer supports) - bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json retrieval-help line Leaves the post-removal templates fully consistent with the Antigravity redirect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
6d072435d0 |
test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks, read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim block-scoped describe wrappers; 427 subtests conserved 1:1. Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec, so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard. Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6, docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md (EN + ja/ko/pt/zh) and ADR-0002. lint:ci green. Part of epic #1969. Closes #1975. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0cc7a1a426 |
test(#1974): consolidate 27 installer/hooks remainder tests into module suites
Fold 27 issue-named installer/hooks/statusline/migration/reapply regression files into their canonical module suites (installer-migrations, installer-migration-report, gsd-statusline, reapply-verify-hunks, install-*, gsd-check-update-worker-platform-gate, etc.). Verbatim block-scoped describe wrappers; 276 subtests conserved 1:1. No new files. The one subdir origin (tests/installer-migrations/001-legacy-orphan-files) moved up one level into installer-migrations.test.cjs; its single ../../ module require corrected to ../ so it resolves from tests/ root (verified). Host-env pre-check: no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. Regenerates regression-name allowlist (222->205), ratchets file-count allowlist (verify 11->8, validate entry removed), makes 16 relocated allow-test-rule exemptions issue-ref- compliant (ADR-456; prunes stale ids). Repoints 15 tests/ references across state-md.md (EN + ja/ko/pt/zh). lint:ci green. Part of epic #1969. Closes #1974. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c0a8f383db |
test(#1970): scope HOME/USERPROFILE in folded bug-130/bug-410 blocks
Adversarial review (Codex) surfaced a latent cross-suite leak: the folded bug-130 and bug-410 install.test.cjs blocks set process.env.HOME/USERPROFILE to a temp home at collection time and never restored — harmless when each ran as its own process, but after consolidation it leaked into sibling folded suites in the same process. Scope both mutations to before()/after() hooks that restore the prior values, matching the save/restore pattern used elsewhere in the file. Assertions unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4f779eda43 |
test(#1970): consolidate 61 install-suite regression tests into function suites
Fold 61 issue-named install/codex/runtime regression files into the canonical test file that owns each subject-under-test, preserving every assertion and its origin issue number as provenance (block-scoped describe wrappers, zero assertion loss — 652 subtests conserved 1:1). Routes: - codex-config.test.cjs +19 (codex config/toml/hooks/adapter/skill surface) - install.test.cjs +18 (node-runner norm, manifest, arg parse, finishInstall) - install-runtime-artifacts +13 (per-runtime conversion + emission) - install-minimal-hooks +6 (hook-event dialects + guards) - path-replacement +2 (opencode absolute pathPrefix) - install-write-confinement +2 (pristine dir writes) - install-regressions +1 (user-artifact preservation) Removes 61 tests/ files → 61 fewer CI processes. Regenerates the regression-name allowlist (271→222), ratchets the file-count allowlist (config 10→9, install 12→9), and makes 13 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes 13 stale allowlist ids). Repoints ADR-0009's moved-test list. lint:ci green. Part of epic #1969. Closes #1970. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a72dbfa58c |
feat(#1707): no-path-literal-in-assert AST rule + portability foundation (#1710)
Phase 1 of epic #1702. Closes #1707. |
||
|
|
fc5ca178a2 |
test(#1178): consolidate duplicated agent-roster helper into tests/helpers (#1420)
* test(#1178): consolidate duplicated agent-roster helper into tests/helpers The "list gsd-*.md agent files, strip .md, sort" derivation was hand-duplicated across the suite (two listAgentFiles(), an identical agentFilesOnDisk(), and inline readdir blocks). Add tests/helpers/agent-roster.cjs exporting listAgentFiles(agentsDir?) and route the genuinely-identical source-roster sites through it. Semantically-different sites (installed-dest dirs, absolute-path returns, .toml-inclusive Codex rosters, full-.md-filename readers, the uniform multi-family inventory table) are left intact, each with a one-line comment. Test-only; no production code touched. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1178): note AGENTS_DIR export is for future call sites Review nit: clarify that the currently-unused AGENTS_DIR export is intentional — available for future tests needing the canonical source agents path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
9d12725e4e |
fix(#1619): normalize pruned mise node execPath to the stable shim in normalizeNodePath (#1621)
* fix(#1619): normalize pruned mise node execPath to the stable shim resolveNodeRunner() bakes process.execPath into managed .js hook commands. Node realpaths execPath, so under mise it resolves to a concrete <data>/installs/node/<ver>/bin/node that mise prunes on `mise up`, after which every managed hook 404s — the same ephemeral-path failure #977 fixed for fnm and #3181 for Homebrew. normalizeNodePath now rewrites a mise versioned install path to the stable sibling shim <data>/shims/node when it exists (deriving <data> from execPath so a custom MISE_DATA_DIR works), falling back to the raw execPath otherwise. Tests folded into install.test.cjs per the regression test-name lint. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(changeset): set pr number to 1621 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Joe Seymour <joese@iarx.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
db2b4d326a |
fix(#1629): cleanup legacy .devin/skills/gsd- dirs on Windsurf reinstall
Pre-#1615 Windsurf installs wrote skills under .devin/skills/gsd-*/ (Devin Desktop preferred dir, #1085). PR #1615 moved Windsurf to .windsurf/workflows/ but never cleaned up the old layout. Users upgrading from a pre-#1615 install were left with dead .devin/skills/gsd-* directories that nothing reads anymore. Fix: added cleanupWindsurfLegacyDevinSkills() which mirrors the Codex cleanupCodexSkillMetadataSidecars() pattern. Runs on Windsurf local install, removes GSD-managed .devin/skills/gsd-* dirs, preserves user content (non-gsd- dirs, gsd-dev-preferences per #2973, symlinks). Empty .devin/ and .devin/skills/ containers are pruned; non-empty ones are left intact. 5 regression tests: removes gsd-* dirs; preserves user content; skips symlinks (escape guard); no-op when absent; end-to-end install removes pre-staged legacy artifacts. Refs #1629 (Finding B; Finding A addressed in #1630). |
||
|
|
b65939cdff |
fix(#1629): copy Windsurf command bodies so workflow delegation targets exist
PR #1622 (issue #1615) shipped Windsurf /gsd-* workflow wrappers that delegate to command bodies at <targetDir>/.windsurf/gsd-core/commands/gsd/X.md via a hardcoded @~/.claude/gsd-core/commands/gsd/ path. The path-rewrite pipeline correctly substitutes ~/.claude/ to the install target. But the source gsd-core/ dir does not ship with commands/ — the canonical command source lives at the package root (commands/gsd/). Without this copy, every /gsd-* workflow in Cascade references a file that does not exist. The slash commands appear in the / menu but silently fail when invoked because the LLM is told to read a missing file. None of the original reviews caught this: not the security review, not Codex's adversarial orthogonal review (gpt-5.5/high), not Memtrace's graph-backed review. It was surfaced by a #1629 regression test that verifies 'every workflow @-reference target exists on disk after install' — the test failed, revealing the bug. Fix: for Windsurf local installs, copy commands/gsd/*.md into <targetDir>/gsd-core/commands/gsd/ via copyWithPathReplacement (applies the same path+brand rewrites as the rest of the install). Guarded on isWindsurf && !isGlobal since global Windsurf workflow install is an explicit no-op. Documented as DEFECT.WORKFLOW-DELEGATION-TARGET-NOT-INSTALLED in CONTEXT.md so the pattern is locked in: any new converter emitting a wrapper that delegates to another file MUST verify the delegation target is actually installed. |
||
|
|
527142ad2e |
fix(#1615): normalize Windows backslash paths in workflow content
computePathPrefix returned a Windows-style path (with backslashes from path.join) into markdown @-references. Workflow file content on Windows ended up with mixed separators, breaking substring checks in install/install-runtime-artifacts tests on windows-latest CI only. Normalize resolvedTarget and homeDir to forward slashes inside computePathPrefix. The prefix is always substituted into markdown body text, which uses POSIX paths universally. Idempotent on POSIX. Also normalizes the two test assertions to forward-slash form so they pass on Windows. Adds a regression test for backslash-style input. Documents DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT + RULESET.CONTENT-PATH-NORMALIZATION in CONTEXT.md so this anti-pattern stops recurring. |
||
|
|
fc2a7c0555 | fix(#1615): install Windsurf slash workflows | ||
|
|
cbd21092a9 | fix(#1614): install Antigravity skills flat | ||
|
|
dcceb1a004 |
fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing (#1442)
* fix(#1441): antigravity resolver prefers GSD-owned dir over first-existing resolveConfigHomeFromDescriptor's dot-home-nested probe returned the first bare-existing candidate, so a CLI user (~/.gemini/antigravity-cli) who also had the IDE's ~/.gemini/antigravity dir was silently shadowed to the legacy dir (probed first). Regression from #217 — pre-#217 returned ~/.gemini/antigravity unconditionally. Add an optional probeMarker (gsd-core/VERSION) to the dot-home-nested descriptor: a two-pass probe prefers the candidate GSD installed into, then bare existence, then probe[0]. Behavior is byte-identical when probeMarker is absent (windsurf etc. unaffected). Adds detectAntigravityDirAmbiguity() for installer/operator guidance on already-misinstalled users (auto-relocation ruled out per ADR-0008's single-configDir migration bound). Regression tests fail before / pass after: coexistence + marker-priority cases, end-to-end through the registry descriptor. Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB * chore(#1441): add changeset for antigravity resolver fix Claude-Session: https://claude.ai/code/session_01JU2WB23JLE3QTysPXuFJjB |
||
|
|
00acbc8868 |
fix(#1223): install scripts/fix-slash-commands.cjs so gsd-tools loads (#1240)
* fix(#1223): install scripts/fix-slash-commands.cjs so gsd-tools loads Before this fix, bin/install.js copied scripts/changeset/ and scripts/lib/ into the runtime config dir but omitted scripts/fix-slash-commands.cjs. gsd-core/bin/lib/command-roster.cjs requires this file at module load via require('../../../scripts/fix-slash-commands.cjs'), so every gsd-tools command crashed with MODULE_NOT_FOUND on every installed runtime. Four changes: - bin/install.js copy step: copy fix-slash-commands.cjs into <configDir>/scripts/ with source-missing hard-fail and verifyFileInstalled smoke check - bin/install.js writeManifest: track scripts/fix-slash-commands.cjs (not covered by the changeset/lib subdir loops) - bin/install.js uninstall: best-effort unlinkSync before scripts/ rmdir - scripts/fix-slash-commands.cjs readCmdNames(): wrap readdirSync in try/catch returning [] so skill-based/global installs without a local commands/gsd/ directory do not throw ENOENT Tests added to tests/install.test.cjs (6 new tests): - smoke: install() copies fix-slash-commands.cjs - e2e: spawned gsd-tools.cjs does not crash with MODULE_NOT_FOUND - manifest: writeManifest() tracks the file - uninstall: uninstall() removes the file - readCmdNames unit: export returns an array - readCmdNames spawn: absent COMMANDS_DIR returns exit 0 (no throw) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1223): backfill changeset PR number (#1240) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5fa4dcd78c |
fix: recover silently-excluded test dirs + test-architecture audit hardening (#1195)
* fix: recurse test discovery so subdir test suites actually run
scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.
Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: retire 5 verified-worthless tests
Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
and bug-782-cline-skills-emission)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: add ADR-218 release version-validation coverage
ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: redesign weak tests into behavioral, deterministic assertions
Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
collision) and unconditional plugin.json schema validation (issue-766)
Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add no-tautological-assert lint rule, error in test suite
New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).
Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gate new allow-test-rule exemptions to require an issue ref
ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: add ADR test-audit evidence report (#1192)
Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture
feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at
|
||
|
|
b4a7eabaae |
feat(#1085): migrate windsurf workspace skills to .devin/ + fix global content refs (#1093)
Fresh windsurf/devin-desktop workspace installs write skills under .devin/ (legacy .windsurf/ recognized); global ~/.codeium/windsurf/ unchanged. Also threads real isGlobal through _applyRuntimeRewrites so global skill content references the codeium path. Closes #1085. |
||
|
|
77a671ec53 |
feat(#791): migrate antigravity workspace base dir .agent → .agents (#1090)
Fresh antigravity workspace installs write under the canonical .agents/ (plural) base; legacy .agent/ stays recognized (dual-read). Global ~/.gemini/antigravity/ path unchanged. Closes #791. |