af822a8024f732c167b163f163a0487bdbcfebdf
30 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8d0b6868ae |
fix(#4725): write normalization preserves tight paragraph-list shape (#4842)
* test(#4725): write normalization must not reflow untouched prose (failing first) * fix(#4725): stop write normalization injecting a blank before a list after prose * test(#4725): repair ordered-list fixture and list-spacing snapshot * test(#4725): assert whole-file prose stability, fix heading-list comment * chore(#4725): backfill changeset PR number (4842) --------- Co-authored-by: sim <sim@local> |
||
|
|
ad1477d659 |
enhance(#4154): validate configured entrypoints before reporting install success (#4249)
* test(260903-m7p): expose configured-entrypoint validation gap * enhance(260903-m7p): validate configured entrypoints before success * test(260903-m7p): require pre-success entrypoint validation * enhance(260903-m7p): gate install success on entrypoints * test(260903-m7p): cover configured entrypoints across runtimes * enhance(260903-m7p): cover emitted runtime entrypoints * fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number - finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process before the new configured-entrypoint assertion throws. Without a HOME + config-location-env sandbox that write resolved through the ambient environment and landed in the developer's live ~/.gsd (confirmed absent on origin/next baseline, present only on this branch — full-suite HERMETICITY WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the duration of the test, matching the existing in-process finishInstall/ install() pattern in tests/install.test.cjs (#2665). - .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the fork PR number until the upstream PR number is known. * fix(260903-m7p): repair cross-platform and pre-existing shape fallout - tests/configured-entrypoint-validation.test.cjs: the win32 branch of ensureCodexHooksJsonSessionStart writes a .cmd shim under <codexRoot>/hooks/; create that dir in the test (the real installer only calls this once hooks/gsd-check-update.js already exists) and assert the platform-common entrypoint shape instead of a fixed non-Windows array, since win32 legitimately emits two entries (cmd shim + script). - tests/install.test.cjs: finishInstall's shared settings-json return now carries configuredEntrypoints/rollbackInstallerMigrations for every runtime on that path (trae included, not just Claude/Cursor/Windsurf); update the trae install() exact-shape assertion to match. * fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash configuredEntrypointsForHook's shell branch dropped interpreterCandidates entirely when resolveBashExecutable returned null, unlike the sibling portableHooks runner entry a few lines below (which correctly falls back to the literal 'bash' token). Found via agy adversarial review; verified unreachable through the current call graph (buildHookCommand's own resolveBashRunner==null gate already short-circuits before recordConfiguredHookCommand runs), so this is a defensive consistency fix, not a live-bug patch — kept for the next caller that does not share that gate. * chore(260903-m7p): backfill changeset pr number to the opened upstream PR .changeset/quick-wasps-sing.md carried the fork PR number (16) as a placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now open, so record its real number per CONTRIBUTING.md's changeset pr-field convention. * fix(#4154): track already-registered hooks for entrypoint validation on update applySettingsJsonHooks registers each guard hook only if absent, so a hook already present from a prior install keeps its stale on-disk command. The new entrypoint tracker always records the freshly-computed command for it, which never matches what is actually persisted, so the exact-string filter in finishInstall silently dropped it from validation — the Blocker case this feature exists to catch (an already-installed entrypoint going stale between installs) was exactly the case it never validated. Match on the managed script's basename instead, which the persisted command carries either way, so an already-registered hook stays in the validated set. Regression test forces this path by mutating a freshly-installed hook's persisted command before a second install. * fix(#4154): distinguish an unreadable script from a missing one validateConfiguredEntrypoints folded an EACCES statSync failure into the same 'missing' reason as ENOENT, misreporting a real permission problem as an absent file. Check the error code and report 'unreadable' instead. * docs(#4154): document entrypoint validation's rollback and PATH scope CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite bin/install.js x CONTEXT.md being this repo's strongest co-change pairing. The update-gsd.md how-to overstated what a validation failure undoes: for Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/ config.toml inside install() before the aggregate validation call runs, so there is no rollback path for that write regardless of "where available" phrasing. Also note that interpreter resolution checks the installer's own PATH, not necessarily the PATH a hook fires under later (#2979 launchers). * chore(#4154): point changeset pr field at the fork PR while CI runs there Mirrors the branch's own prior backfill commit: pr: matches whichever PR number changeset-lint is currently validating against (fork PR #16 during the fork-first CI/review loop), flipped back to the upstream PR number right before the final push to open-gsd/gsd-core. * fix(#4249): address adversarial-review findings in entrypoint validation An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR found several real gaps beyond the human reviewer's Blocker, verified against source before fixing: - Codex's install() result bound rollbackInstallerMigrations to the narrow installer-migrations-only rollback instead of restoreCodexSnapshot (#3245), the full pre-install snapshot/restore Codex already owns for exactly this case — a validation failure discovered outside install() reverted nothing of the config.toml/hooks.json that call had already written. - The register-only-if-absent basename match from the prior fix used a bare substring, which an unrelated user command mentioning the same filename could false-positive into GSD's validated set — anchored on the `/hooks/<basename>` path segment instead. - nodeCandidates checked raw process.execPath (always true — we're running in that process) instead of normalizeNodePath's stable version-manager alias, the same one buildNodeRunnerChainToken bakes as its first choice — a false green regardless of whether that alias itself still resolves. - An entry with no interpreterCandidates (Cline's PreToolUse hook, or a Windows-Claude .sh hook invoked without a bash runner) runs via its own shebang; validateConfiguredEntrypoints checked only file-type, never the execute bit. Cline's writer also never reported an entrypoint at all. - Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor hook registered across several events) were validated once per duplicate. Each fix is covered by a new or extended test; the Codex one required inlining runCodexInstall's env sandboxing so the rollback closure — which re-resolves the $HOME-relative skills root live — runs before the sandbox is torn down, matching how installAllRuntimes' real aggregate gate calls it. * docs(#4249): document the round-2 entrypoint-validation fixes Runtime Hooks Surface Module and Installer Module entries now name ConfiguredEntrypoint's not-executable reason, the normalizeNodePath alignment, Cline's tracked hook, and which install() result the finishInstall/installAllRuntimes rollback path actually reverts per runtime (Codex's full snapshot vs. the others' narrow migrations-only rollback). * chore(#4249): point changeset pr field at the upstream PR now that fork CI is green * fix(#4249): address agy adversarial-review findings - validateConfiguredEntrypoints: statSync alone never detects a chmod-000 script (it only needs parent-dir search permission), so an interpreter-invoked entry with an unreadable script passed validation. Add an explicit R_OK check for the interpreterCandidates branch only — the candidate-less/shebang branch already has its own X_OK gate. - docs/how-to/update-gsd.md: the blanket "does not revert" claim was false for Codex, which reverts config.toml/hooks.json via its full pre-install snapshot; qualify it per runtime. - tests/codex-config.test.cjs: the #4249 rollback regression test asserted skills/ and VERSION were reverted but never asserted config.toml/hooks.json were too, despite the test's own stated intent. - CONTEXT.md: qualify which interpreterCandidates entries get normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's Windows .cmd shim as a third candidate-less case that relies on extension dispatch, not a shebang. * fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node', so it needs the execute bit (like any shebang-invoked entry), but its interpreter is looked up on PATH by 'env' at hook-fire time (unlike every other GSD JS hook, which bakes an absolute node path specifically to avoid that dependency). The candidate-less/interpreterCandidates fork treated these as mutually exclusive, so Cline's entry silently skipped interpreter resolution entirely — a completely missing 'node' on PATH would still validate successfully. Add an orthogonal selfExecutable flag so both checks run for entries that need them. (CodeRabbit finding on the fork rehearsal PR.) * fix(#4249): address second-round adversarial review findings (opus + agy) - validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry, not just interpreterCandidates ones — a self-executable shebang script is still opened and read by its kernel-invoked interpreter, so X_OK alone never proved it was readable. - selfExecutable is now the sole, explicit source of truth for the execute-bit check (every producer that needs it sets the flag) instead of being partly inferred from an absent interpreterCandidates, which Cline's hybrid entry also carries. - The execute-bit check now skips explicitly on win32 (matching resolveExecutableBinary's own carve-out) instead of relying on Node's accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows machine and not a test that simulates win32 on a POSIX runner. - bin/install.js: fixed a stale comment claiming no runtime's install()-time writes have a rollback path — Codex's does (restoreCodexSnapshot) — and added the omitted Cline to both that comment and CONTEXT.md's equivalent lists. - CONTEXT.md: fixed the Cline description left stale by the previous commit's selfExecutable addition, and rewrote the validation-mechanism paragraph for clarity (writing-for-agents pass). - docs/how-to/update-gsd.md: split an overloaded 4-clause sentence. - Removed a fault-injection integration test that could not reliably exercise the real installAllRuntimes -> finalize -> rollback wiring without fighting the installer's own pre-registration existence guards; the constituent pieces remain covered individually. * fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner X_OK is a POSIX-only concept, skipped entirely when an entry's platform is win32 (matching production). Two test entries omitted platform, defaulting to process.platform — on an actual windows-latest CI runner that silently skipped the very check they were meant to exercise, turning 'not-executable' into a false pass. Pin platform: 'linux' so these are deterministic regardless of which OS runs the suite. * fix(#4249): classify EPERM the same as EACCES in statSync error handling Windows raises EPERM (not EACCES) for a parent directory that couldn't be traversed into — was falling through to 'missing', misreporting a genuine permission problem as a nonexistent path. * docs(#4249): address final CodeRabbit doc-completeness findings - CONTEXT.md: install()'s documented result shape omitted configuredEntrypoints; the ConfiguredEntrypoint shape omitted selfExecutable. - docs/how-to/update-gsd.md: the failure-mode sentence omitted unreadable and lacks-execute-permission, which the installer also rejects. * fix(#4249): stop double-validating every configured entrypoint on install/update installAllRuntimes' finalize() already runs assertConfiguredEntrypoints once over the aggregate set; finishInstall then re-ran the identical check per runtime in the printSummaries loop right after, so every entrypoint paid its statSync/accessSync/interpreter-resolution cost twice on every install and update. Add entrypointsAlreadyValidated to skip the redundant pass specifically on that path, while leaving the check intact for any caller that invokes finishInstall directly. * chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there * perf(#4249): memoize interpreter candidate resolution across entrypoints resolveExecutableBinary walked PATH once per (entry, candidate) pair; a typical install has a dozen-plus entries sharing the same few candidate lists (process.execPath for JS hooks, bash for shell hooks). Cache by (platform, candidate) so each distinct pair resolves once per validation call instead of once per entry. * chore(#4249): point changeset pr field at the rebased rehearsal fork PR * fix(#4249): drop entrypoint tracking from the now-dead Codex event writer #2586 (landed on next after this branch forked) removed install.js's CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent no longer runs during install or update. The ConfiguredEntrypoint records this branch added inside it were therefore unreachable and untested. Restore the function to its upstream shape; the entrypoints it used to report were never collected by any caller. * refactor(#4249): drop the revalidation bypass flag and the candidate cache Both were this PR's own micro-optimisations over a set of roughly a dozen entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall gate off to save one statSync/accessSync pass; `resolvedCandidateCache` memoised resolveExecutableBinary across entries that are already deduped by (configPath, scriptPath). Neither is measurable, and the flag was the only way to reach finishInstall with validation disabled. finishInstall now always validates what it is given. * chore(#4249): point the changeset pr field back at the upstream PR * refactor(#4249): track settings.json entrypoints without the hooksSurface gate The install-surface writer only tracked configured entrypoints when the runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing asserts that axis agrees with `installSurface`, so a descriptor that broke the coupling would silently pass `configuredEntrypoints: undefined` and drop that runtime out of the validation this PR adds — reintroducing the exact 'reports Done! over a broken entrypoint' failure #4154 exists to close. Remove the dependence rather than test it: everything recorded on this path lands in settings.json by construction, and the registered-command filter already discards entries no persisted hook references. * chore(#4249): put the changeset body in the documented two-part format CONTRIBUTING.md and .changeset/README.md both show `**<bold change>** — <symptom-led explanation>.`; the fragment was a single unbolded sentence. * chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there * fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback #3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-* and gsd-core/VERSION. The install overwrites every other GSD-owned file too — hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest itself — before the entrypoint-validation gate runs, so a validation failure left the new payload sitting on top of the restored old config. Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims, before runInstallerMigrations so the bytes are the true pre-install state, and restore it from both Codex rollback closures ahead of the per-surface restores. Files only the failed install introduced are removed, read from the manifest now on disk. The manifest is already the authoritative record of what GSD owns, so no second hand-written list can drift out of sync, and user-owned files are never snapshotted or removed. Every path is confined through resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into an arbitrary-path write. Non-Codex runtimes are unaffected: the snapshot is gated on the same tomlConfigInstall + non-minimal condition as #3245's. * fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest Two follow-on defects in the previous commit's snapshot: - The capture was gated on `!isMinimalMode`, copied from #3245. A core/ --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the snapshot came back empty while the rollback still ran — and its removal pass would have deleted every file the new manifest lists. Gate on tomlConfigInstall alone, matching where the rollback actually reaches. - An unreadable or unparseable prior manifest was caught alongside ENOENT and treated as a fresh install. That is the same empty-snapshot state, so a failed update over a real install with a corrupt manifest could delete its prior payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means known-empty; any other read error or a parse failure means unknown, and the restore closure returns without touching anything, degrading to #3245's narrower rollback. Deliberately not fatal — a corrupt manifest has to stay repairable by reinstalling over it. Both paths are covered by red-checked regression tests. * fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too commit removed from the manifest snapshot. restoreCodexSnapshot is reachable for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-* skill dir and gsd-* agent file the snapshot does not claim — so with an empty minimal-mode snapshot a rollback deleted the whole skills/agents surface with nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via the ADR-1239 skills-kind home override, so this is also the reason manifest `skills/` keys do not resolve under configDir: that surface belongs to this snapshot, not to the manifest-driven one. Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal mode — doing nothing on an early failure is the non-destructive side. Covered by a red-checked regression test that plants bytes in an alternate-home skill file, reinstalls under the core profile marker, and asserts the rollback restores it. * fix(#4249): never remove on rollback unless a prior manifest proves what predates the install Three defects in the manifest-driven Codex rollback, all in its removal half: - ENOENT marked the snapshot usable, arming the removal pass on a FIRST install. GSD may have overwritten a user's file at a manifest-tracked path there, and no prior manifest records the difference — so rollback deleted it where before it merely left it overwritten. Absent, unreadable and malformed manifests now all leave the prior set UNKNOWN and skip removal entirely. - Membership was tested against the map of files whose pre-install read SUCCEEDED, so a tracked file that existed but was unreadable read as introduced-by-this-install and was removed. Track the prior manifest's paths in their own Set and test against that. - The unreachable "delete the manifest when there was no prior one" branch is gone: usable now implies a parsed prior manifest. Also adds the end-to-end test the aggregate gate was missing — the four Codex rollback tests drove the closure directly, proving the restore but not the wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes Cline's `env node` entry fail validation for real, and asserts Codex's payload comes back. Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper instead of six copies. Both new tests are red-checked. * test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its Windows-EBUSY retry budget. That budget is for directory trees; this removes a single file, which unlinkSync says more precisely and the rule does not flag. * chore(#4249): point the changeset pr field back at the upstream PR * fix(#4249): use an unambiguous dedup key and surface partial-restore failures trek-e's 2026-09-08 adversarial pass flagged two findings in the new entrypoint-validation/rollback code: - assertConfiguredEntrypoints' dedup key already used a raw NUL separator (introduced in ceebb65f2d), but git/Read render NUL as a space, so the key looked like a plain-space join to every reviewer that read the diff. Replace it with JSON.stringify([configPath, scriptPath]) so the separator is visible and unambiguous. - restoreManagedFileSnapshot's per-file restore catch block claimed to 'surface the original error' but only swallowed it, matching (and widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add an actual console.warn using the existing best-effort-warning convention, scoped to just this PR's new function. * fix(#4249): treat a files-less prior manifest as unknown, not known-empty agy's gemini-3.8-flash-high adversarial pass (round 5) found and I reproduced empirically: a structurally-valid manifest missing the files key (e.g. {"version":1}) parses without throwing, so Object.keys(undefined || {}) silently read as 'zero files predate this install' instead of the UNKNOWN state the malformed-manifest guard exists to produce. Rollback's removal pass then deleted every GSD-owned file the failed install's own manifest listed, including ones that predated it — the exact data loss the #4249 CodeRabbit malformed-manifest fix was supposed to prevent, reachable through a JSON.parse success instead of a failure. Route the shapeless case into the same catch-all UNKNOWN path via an explicit shape check. Regression test reproduces the deletion before the fix and confirms the file survives after it. Also extend restoreManagedFileSnapshot's removal-pass rmSync and final manifest-rewrite catches with the same real console.warn trek-e's round-4 review asked for on the per-file restore catch — same rollback function, same operator-facing-signal gap. * docs(#4249): correct which runtimes actually leave a written config on rollback agy's completeness audit (round 5, holistic pass) caught this new paragraph claiming 'for every other runtime, the configuration file(s) already written during that update are left in place' — false for Claude Code and other settings.json-based runtimes, whose write never happens on failure (assertConfiguredEntrypoints runs before finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline actually match that description, since they persist their config file inside install() ahead of the gate. Split the one sentence into the three actual outcomes; matches the PR body's own accurate Before/After wording, which this doc addition had drifted from. * fix(#4249): clean up doc/comment mismatches and dead fields from opus review Opus critical-code-reviewer + ponytail-review pass on the final diff: - assertConfiguredEntrypoints carried finishInstall's old docblock ("Apply statusline config, then print completion message") from before this function was inserted between comment and callee. finishInstall already has its own accurate #4249 comment, so the stale docblock is removed rather than moved. - checked: number on ConfiguredEntrypointValidationResult and error.configuredEntrypointValidation on the thrown error: the first had zero consumers anywhere in the repo, including its own defining file, and is removed. The second matches an existing repo convention (bin/install.js's installerMigrationRollbackFailures, #4249 predates this PR) of attaching structured diagnostic context to a re-thrown Error even before a consumer exists, so it's kept. - finishInstall's own assertConfiguredEntrypoints call is a redundant backstop on the real production path (installAllRuntimes's aggregate call already validates the superset first), but its comment read as though this call alone provided the before-the-write guarantee. Clarified rather than removed — it's the only gate for a caller that invokes finishInstall directly. * chore(#4249): split the manifest-driven rollback engine out into #4544 Issue #4154 asked the installer to consume a validation failure "through the existing rollback mechanism, without a second transaction mechanism". The manifest-driven rollback widening added during review (capture every path the prior gsd-file-manifest.json claims, restore those bytes, remove what only the failed install introduced) is that second mechanism on a plain reading. It is a real fix for a #3245-era gap, but an independent one, so it moves to its own bug report and PR. Removed here: - bin/install.js: the pre-install managed-file capture block and restoreManagedFileSnapshot, plus its call sites in _codexPreConfigRollback and restoreCodexSnapshot (99 lines). - tests/configured-entrypoint-validation.test.cjs: the five tests that exercise the manifest engine. - CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the widened restore. update-gsd.md again documents the #3245 surfaces only. Kept, because it is #4154's own scope: - the entrypoint-validation gate itself; - Codex's install() result binding rollbackInstallerMigrations to restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*, agents/gsd-*, gsd-core/VERSION); - the !isMinimalMode gate removal on that snapshot. Binding the closure to the result made it reachable for a core/--minimal install, where its pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot does not claim; an empty minimal-mode snapshot therefore deleted the whole surface with nothing to restore. The surviving aggregate-failure test now asserts on config.toml, a surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which only the manifest engine restored. Refs #4544 * test(#4249): cover configured entrypoints through the packed install path #4154's scope lists install smoke coverage alongside the installer gate — "assert representative configured entrypoints resolve for supported runtime profiles". The gate itself (assertConfiguredEntrypoints / validateConfiguredEntrypoints) is unit-covered by in-process install() calls; nothing proved the property survives npm pack -> npm install -g -> install.js. Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct config surfaces GSD writes launch paths into (settings.json, and hooks.json + config.toml) — run the tarball-installed installer into a throwaway HOME, then re-read that runtime's own written config and return the new ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check becomes a release gate on every matrix host without workflow changes. The scan re-derives paths from the written config instead of reusing the installer's own entrypoint list, and test I shows why that matters: a registration the installer never touched during a run is invisible to the in-process gate, so the install exits 0 and only reading the config back off disk catches the dangling launch path. * ci(#4249): pack a publish-shaped tarball in the install smoke lane `npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs after `npm ci` carries no hook scripts at all — the lane has been smoking a package that differs from the published one in exactly the artifacts the lifecycle smoke is supposed to launch. That went unnoticed because the lane's init runs `--local`, which registers no statusline and therefore registers no hook whose target is missing. A `--global` install on the same tarball exits 1 on #4249's own gate (`gsd-statusline.js (missing)`), which is what the new configured-entrypoint cycle performs, so without this step the cycle would report INIT_FAILED instead of checking anything. Build hooks before packing so the smoked tarball matches prepublishOnly. The CLI now reports 16 configured entrypoints for claude and 1 for codex instead of zero. * fix(#4249): scope Codex's full snapshot restore to entrypoint failures Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage exception un-install a Codex install that had already succeeded and already printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints` call. Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration` scopes finalize-stage rollback to installer *migrations* ("the executor uses the journal to restore modified paths"), and this PR's own operator-facing paragraph in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/ agents/VERSION revert to entrypoint-validation failures specifically ("If a script is missing, unreadable, ... For Codex, this reverts ..."). The wide behaviour is also incoherent as a transaction abort: the same doc says Cursor, Windsurf, Kimi and Cline keep the config they wrote inside install(). Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install bytes — on an update, silently downgrading a working Codex install to the previous version while the user had just been told it was Done. Select the rollback by error kind instead. `assertConfiguredEntrypoints` already tags its error with `configuredEntrypointValidation`, so the full snapshot restore runs for that error (and anything downstream of it, including finishInstall's per-runtime backstop) and the installer-migrations-only closure runs for everything else. The codex result now also exposes that narrow closure as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning the full restore, so the direct-call contract asserted by tests/codex-config.test.cjs is unchanged. Adds a regression test that installs codex+kilo together, injects EACCES on the Kilo permission write by monkeypatching node:fs (restored in a finally — never chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps the bytes the successful install wrote. Verified red against the pre-fix unconditional path. Cline cannot host this test: its plan is writesSharedSettings:false + finishPermissionWriter:null, so its finishInstall performs no write and has no non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally (unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded fs.writeFileSync. * docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing CONTEXT.md still described Codex's rollback as an unconditional bind to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint- validation failures via rollbackInstallerMigrationsOnly and the configuredEntrypointValidation error tag. Caught during the round-6 PR body pass. * fix(#4249): stop rollbackInstallerMigrations meaning its own opposite Codex's install() result bound `rollbackInstallerMigrations` to restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the actual installer-migrations-only closure behind `rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name meant the opposite of what it says, and CONTEXT.md had to concede as much in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure for every runtime, matching both its name and the meaning it already has on next, and the snapshot restore gets its own Codex-only field, `rollbackPreInstallSnapshot`. The selection in rollbackFinalizedInstallerMigrations collapses to one line and no longer needs a fallback chain. Also in this commit, all against the same rollback path: - Correct the rollbackFinalizedInstallerMigrations comment. It read as if the round-6 narrowing prevented any sibling-triggered revert of a Codex install the user has already seen "Done!" for. It does not, and is not meant to: `wide` is true for ANY entrypoint-validation error from ANY runtime, because the aggregate gate is all-or-nothing — an invalid Cline entrypoint reverts Codex's snapshot, which tests/configured-entrypoint-validation.test.cjs's 'an aggregate entrypoint validation failure rolls the Codex install back (#4249)' asserts directly. The discriminator is the error's KIND, not which runtime owns the failing path. Comment and CONTEXT.md now say that. - Name the runtime in the "Configured entrypoint validation failed" error. ConfiguredEntrypointInvalid already carries `runtime`; the message threw it away, leaving an operator of a multi-runtime install unable to tell whose entrypoint broke — which matters precisely because the failure can revert a runtime that was itself fine. - Set `configuredEntrypoints: []` explicitly on the copilot-instructions early return. Every other branch states the key; this one relied on installAllRuntimes' `(result.configuredEntrypoints || [])` defence. `[]` is correct, not a workaround: every Copilot hook is an inline printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no GSD-managed script or interpreter to resolve. No behaviour change beyond the error-message text. * docs(#4249): narrow the smoke scan's config-surface claim to what it checks RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is told to launch is registered in one of settings.json / hooks.json / config.toml, and that nothing else in a config dir is runtime configuration. Both halves are false as stated. Cline registers its hook at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native [[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi), a directory separate from Kimi's own GSD configDir — the same gap installer-migration 007 already documents as structurally unreachable. The scan is in fact correct for what it runs against: entrypointRuntimes defaults to claude + codex, whose launch paths do all live in those three top-level files. Restate the docstring at that scope, name the two known out-of-scope surfaces, and warn that adding either runtime to entrypointRuntimes without teaching scanConfiguredEntrypoints about its surface yields a scan that finds zero entrypoints and proves nothing. The entrypointRuntimes default comment carried the same overgeneralization ("every other runtime reuses one of them") and is corrected with it. Documentation only; no code change. * fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found that install()'s settings-json early return for an unparseable settings.local.json omitted configuredEntrypoints and rollbackInstallerMigrations from its result, unlike every other branch. rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations unconditionally, so this branch silently dropped its own installer-migration rollback on a later finalize-stage failure. Completed the return shape: configuredEntrypoints: [] (matching Copilot's equally-early no-entrypoints-yet return) and rollbackInstallerMigrations (already in closure scope). Red-then-green regression test added. * test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs Same internal adversarial review (round 8): this suite's before() packed the tarball directly, without the ensureHooksDist() guard every sibling install-test suite (install.test.cjs, install-minimal-hooks.test.cjs, mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run in isolation ahead of a suite that builds hooks/dist itself, this suite's pack would ship a tarball with no hook scripts and fail closed on SMOKE.INIT_FAILED instead of testing anything. * fix(#4249): refresh stale test-timings weight for the codex-config split next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the CI shard packer's weight table (tests/test-timings.json) was never updated: codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block this PR extends with its own #4249 install()-pipeline test — had no entry at all, so the packer would silently underestimate it at the table's median weight (roughly a 9x underestimate against its real cost). trek-e's most recent review flagged a Windows shard timeout in-flight on codex-config.test.cjs, plausibly aggravated by this PR's own addition to that file before the rebase moved it. Re-measured both files locally (node --test --test-reporter=tap, max of 3 runs, matching the table's own max-across-streams methodology) and patched just these two entries — not a full regeneration, which would need real multi-lane CI data this session doesn't have access to. * fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists next's platform-conformance-tier classifier (#4591/#4598) landed after this branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's gen-platform-conformance-tier --check and both the Linux and macOS conformance suites. * fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation next's #4733 (landed after this branch's last rebase) replaced the static ISOLATED_HEAVY_FILES set with a threshold derived live from tests/test-timings.json, and pins the current derived result in EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still named codex-config.test.cjs, whose own weight this PR already dropped from 127783ms to 189ms (after splitting its heavy install()-pipeline blocks into codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The live-computed set correctly no longer includes it; the pinned expectation is updated to match. * fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it trek-e's review flagged two Major gaps: the thrown error read identically regardless of which of three real outcomes a runtime hit (nothing persisted / snapshot reverted / config left broken on disk), and no test proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/ Kimi/Cline — only Codex's revert path was ever asserted. assertConfiguredEntrypoints now tags each invalid entry with its actual consequence, mirrored from docs/how-to/update-gsd.md's existing rollback-matrix disclosure. A new test drives the same aggregate failure through Cline (whose own entrypoint is the one that fails) and asserts its hook file is still on disk afterward. * fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR One review pass (gemini-3.8-flash-high via the antigravity review lane) against this PR's full diff against next, findings independently verified against source before fixing: - Copilot's install() return object was the only one of 6 runtime branches missing rollbackInstallerMigrations — reachable now that this PR's own aggregate gate runs rollback across every result on any runtime's entrypoint failure, not just Copilot's own. - buildHookCommand's unresolved-bash early return skipped track() entirely, so a win32 install with no Git Bash silently produced an unregistered .sh hook instead of the 'unresolved-interpreter' validation failure configuredEntrypointsForHook's own comment said it would. - release-tarball-smoke.cjs reported a Cycle 4 install failure under SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing SMOKE.INSTALL_FAILED. - SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's trailing args, which also truncated any configDir containing a space (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan. Anchored the match on the already-known configDir prefix instead of a generic absolute-path guess: removes the ambiguity outright rather than patching the character class, and stays a raw-text scan on purpose (it catches a writer that emits a path without registering it — a JSON.parse of the expected schema would miss exactly that case). One suggested finding (test-timings.json "missing" the new test file) was verified false — that table only holds measured CI timings, populated after a file's first real run — and one Ponytail suggestion (a JSON.stringify dedup key) was rejected as it would reintroduce a real, if narrow, key-collision risk for no benefit. * fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment changeset-lint requires pr: to match the PR it runs on (16 on the fork, not the eventual upstream number) — rehearsal-branch convention already established earlier in this PR's history. lint-docs-guard-registration's quote-pairing heuristic doesn't require the docs/ path itself to be quoted — it flags a file once ANY quote-delimited span containing "docs/" appears anywhere in it, alongside any real fs read call. A comment ending "...update-gsd.md's rollback-matrix paragraph" supplied the closing quote character (the possessive apostrophe) the heuristic paired with an unrelated single-quoted string earlier in the file. Reworded to avoid the unquoted apostrophe next to the path. * chore(#4249): point the changeset pr field back at the upstream PR Fork rehearsal (PR #16) is green; the real target for this changeset is upstream PR #4249. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
77e2472ca0 |
enhance(#4221): replace installer Read() deny rules with a managed secret-read guard hook (#4236)
* feat(#4221): gsd-secret-read-guard PreToolUse hook + registration Add hooks/gsd-secret-read-guard.js, a blocking PreToolUse guard on Read|Grep|Bash that denies reads of .env, .env.<suffix> and .secrets (the .env.example/.sample/.template/.dist templates stay readable). Read checks file_path; Grep checks an explicit path and judges the glob per brace alternative; Bash runs a two-pass token scan (quotes, comments, redirects with fd digits, separators, $( )/backtick/<( ) recursion, heredoc bodies never scanned as commands, nested bash -c/eval rescans, git <ref>:<path> shapes) with a closed non-reading exemption set for existence checks. Fail-open crash policy; 1 MiB commands are denied as command-too-large; more than 64 glob alternatives as glob-too-complex. Why: Claude Code 2.1.259 makes every `cd DIR && grep …` compound prompt for approval whenever any Read() deny rule exists, even in auto mode. A hook denial is not a permission rule and never arms that check. The installer-written deny rules are retired in the follow-up commit. Registration: hooks.json (Read|Grep|Bash, timeout 5), build-hooks HOOKS_TO_COPY, managed-hooks-registry, runtime-hooks-surface (blocking guard with BLOCKING_GUARD_TIMEOUT_S; Kimi ReadFile|Grep|Shell), shell-command-projection managed sets, installer-migration-report, OpenCode/Kilo plugin (grep tool mapping, include -> glob, dispatch), docs tables in five locales, ADR-766 always-on list, regen:derived fixtures, and a new table-driven unit suite. * test(#4221): pin the secret-read guard in existing hook gates Register gsd-secret-read-guard.js in every existing hook gate: the hooks-crash-policy table (deny row; 6 -> 7 deny cases), plugin-manifest REQUIRED_HOOKS and its Read|Grep|Bash group, docs-hooks-table-parity EXPECTED_SURFACE_HOOKS, install.test MANAGED_JS_HOOKS, install-minimal- hooks JS_HOOKS/BLOCKING_GUARDS, portable-node-runner GUARD_HOOKS, kilo-upgrades PLUGIN_GUARD_HOOKS, the Kimi normalization-parity and typed-payload floors, the OpenCode adapter (grep mapping, include -> glob, three dispatch tests) and a Kimi TOML matcher assertion. * fix(#4221): retire installer Read() deny rules (legacy filter) Rename GSD_CLAUDE_DENY_PERMISSIONS to GSD_CLAUDE_LEGACY_DENY_PERMISSIONS and stop adding the three Read(.env) / Read(.env.*) / Read(.secrets) strings. mergeClaudePermissions now only filters them out of an existing permissions.deny: an absent deny key stays absent, a malformed one is still repaired to [], and an array emptied by the filter is deleted so no `"deny": []` residue is left. Uninstall filters the same legacy list and, symmetric with the Antigravity branch, drops an emptied allow or deny key and an emptied permissions object. Unlike the #2278 allow-side migration there is no surviving current deny list, so the constant is renamed rather than mirrored. Removal is byte-exact: a hand-written identical rule is indistinguishable from the installer's and is removed too (the manifest never recorded permission strings). USER-GUIDE and CONTEXT.md updated. * test(#4221): flip install-regressions deny-rule assertions to the retired shape The fresh-merge, non-destructive merge, idempotency, end-to-end install, reinstall and uninstall assertions now expect no Read(.env*) deny rules and no permissions.deny key on a fresh install; the deny:null repair case is kept. A new describe block covers the legacy filter: retired strings removed with a user entry kept, partial sets, near-miss strings untouched, idempotency, GSD-only deny array deleted, a pre-existing empty deny preserved, and uninstall symmetry for allow/deny/permissions. * chore(#4221): add changeset fragment for PR #4236 * fix(#4221): case-fold names; scan shell stdin and xargs pipes Review round 1 (trek-e): - Blocker: secret-name matching is now case-insensitive in the Read, Grep (path and glob) and Bash paths, so `.ENV` / `.Secrets` on a case-insensitive filesystem are recognized as the same secret file. - Major: a shell interpreter's script is now scanned wherever it comes from. The tokenizer keeps heredoc bodies as per-segment tokens and records separator operators; pass 2 groups by segment id and resolves bash/sh/zsh/dash/ksh/su invocation mode: `-c` (including combined `-lc`) scans the script operand, a file operand is checked as a file (a `<( )` operand's echo/printf output is reconstructed), otherwise stdin is the script and heredocs, here-strings and a piped echo/printf source are scanned. `eval` joins all its operands; `source`/`.` handle process substitution. Data heredocs (`cat <<EOF`, the commit-message shape) stay unscanned. - Major: `… | xargs <cmd>` checks the upstream segment's operands as file names when the sub-command reads (`echo .env | xargs cat`, `find . -name .env | xargs cat`); `-a`/`--arg-file` suppresses the inference; a shell sub-command's `-c` script is scanned. Header, USER-GUIDE bullet and changeset updated; documented gaps now include piped scripts from non-echo sources and `exec`/`timeout` wrappers. 60 new suite cases pin the block and allow shapes. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
fd63889d1f |
fix(#3854): write normalization preserves tight multi-line lists (#4049)
* test(#3854): write normalization must preserve tight multi-line lists (failing first) * fix(#3854): no blank before a bullet whose previous line is an indented continuation _normalizeMd's 'separate a list from a preceding paragraph' rule inserted a blank before any bullet whose previous line wasn't a bullet — but an INDENTED CONTINUATION of the previous multi-line item also isn't a bullet. Every .md write (phase.complete in the report, but any write through platformWriteSync) therefore converted tight lists to loose ones: +61 blank lines on the reporter's 1015-line ROADMAP, one before each bullet following a wrapped item. Tight and loose lists render differently, so this was a rendering change plus misleading diff noise; one-shot (idempotent afterwards), which is why integrity checks on headings/content passed. The guard is the mirror image of the after-a-bullet rule two lines below, which already excludes indented next lines. Paragraph→list and heading→list separations — the rule's purpose — are pinned unchanged by the new suite. * fix(#3854): review fold-ins — ceiling tracks next's 281 + this branch's marker (282), header/require nits The ceiling is not ratcheted but must track the tree: origin/next raised it to 281 (sibling branch's marker file); this tree adds one more (shell-command-projection-md-normalize), so 282/282. Also fixes the test header's stale pre-rename filename and hoists the inline require to the file's single import. * chore(#3854): changeset fragment (pr number backfilled after PR creation) * chore(#3854): backfill changeset PR number (4049) --------- Co-authored-by: sim <sim@local> |
||
|
|
3a6c0412a9 |
enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) (#3976)
* enhance(#3624): local/no-exact-case-env-access — ratchet ADR-1703 onto production env reads (epic #3411 Phase 4) Extends ADR-1703's portability rule catalog with a second production-runtime rule: it flags an exact-case read of a Windows case-varying environment variable (PATH, PATHEXT, ComSpec, USERPROFILE, TEMP, TMP, APPDATA) off any receiver that is not process.env itself, matched via an env-shaped-receiver check to avoid colliding with ordinary `.path`-named properties elsewhere in the tree. Exports the seam's private `_envGet` as `envGet` so the rule's remediation message names a real helper, and fixes the one pre-existing violation the tightened rule found (`src/runtime-hooks-surface.cts`'s `env.APPDATA` read). Closes #3624 * fix(#3624): extractStaticName recognizes non-computed Literal destructuring keys; add missing accessor-call test case Review findings from the code-review + isolated-adversarial passes: - extractStaticName only matched non-computed Identifier keys, so a destructuring like `const { 'PATH': v } = opts.env;` (the issue's own I8 acceptance case) silently evaded the rule. Widened to accept a Literal key regardless of computed, which is safe for MemberExpression too (its non-computed property is always an Identifier by grammar). - Added the missing RuleTester valid case for "a case-insensitive accessor call" (envGet(env, 'PATH')) from the issue's Done-when checklist. * docs: backfill changeset PR number for #3624 (PR #3976) --------- Co-authored-by: sim <sim@local> |
||
|
|
aaf47c5fc2 |
fix(#3691): let every reviewer lane take a prompt cap, and make the documented global resolve (#3832)
* test(#3691): failing-first coverage for the reviewer prompt budget No prompt cap can reach any CLI reviewer lane, by any configuration. Two independent defects compound: all nine `transport: spawn` lanes declare `promptBudgetKey: null`, so `budgetFor` returns on its first line; and the documented global `review.max_prompt_tokens` is advertised in the schema manifest but declared nowhere, so the resolver never materializes it and `budgetFor`'s fallback is dead code. Adds to tests/reviewer-config-federation.test.cjs, which already owns the per-reviewer budget config-set/config-get idiom: - a CLI lane inherits the global cap (RED: reports null) - an http lane with the -1 sentinel inherits the global cap (RED: reports null) - the resolved review surface carries max_prompt_tokens at all (RED: absent) - per-lane overrides the global on a CLI lane - the sentinel boundary: -1 inherits, 0 means do-not-trim and must NOT read as unset, 1 is the smallest real budget — the regression budgetFor's own comment warns about - anti-tightening pins that must stay green: an empty config leaves every lane null, the three existing budgeted lanes are unchanged, and config-set still rejects a per-reviewer key naming something that is not a declared lane - a fast-check property over the resolution contract itself, with -1, 0 and non-finite inputs generated explicitly rather than left to chance Every row was reproduced by hand against the real CLI before being written, so the RED/GREEN split is observed rather than predicted. Refs #3691 * fix(#3691): let every reviewer lane take a prompt cap, and make the global resolve No prompt cap could reach any CLI reviewer lane, by any configuration. Two independent defects compounded. The nine spawn-transport lanes — claude, coderabbit, antigravity, cursor, gemini, codex, kimi-code, opencode, qwen — declared `promptBudgetKey: null`, so `budgetFor` returned on its first line and `review-lane plan` reported `promptBudget: null` no matter what was configured. Each now declares `review.max_prompt_tokens_per_reviewer.<slug>` with the same `-1`-is-unset sentinel the three local-server lanes already use. Separately, the central `review.max_prompt_tokens` was listed in the schema manifest's validKeys and documented as a supported setting, but declared nowhere — the resolved surface is built from capability declarations plus the defaults manifest, and neither carried it. `configGet` returned undefined and `budgetFor`'s documented fallback was dead code. It is now declared with a `null` default, exactly as docs/CONFIGURATION.md already specified, so the default behavior is unchanged: nothing configured means nothing trims. Two things the diagnosis had not predicted, found and fixed while implementing: - `REVIEWER_LANES` in src/review-lane-descriptor.cts is a second, hardcoded registration site that `mergeReviewerLanes` prefers over the capability registry on a slug collision. Editing only the capability files left every CLI lane still null. Both sites now agree. - The generated `gsd-core/bin/lib/capability-registry.cjs` was stale and masked the capability edits; regenerated with `npm run gen:capability-registry` rather than hand-edited. docs/CONFIGURATION.md said "Only lanes that declare a budget key accept one — today ollama, lm_studio and llama_cpp". That is false as of this change and is corrected rather than left to rot. The trim-versus-refuse question the issue raises is deliberately not taken up here: the refusal path already exists for the case that matters — a reviewer whose minimum set exceeds its budget is skipped rather than sent a misleading prompt — and trimming above that floor is the documented, shipped design of the feature. Changing it would alter behavior for the three lanes that already work, which is not what the issue asks for. Fixes #3691 * fix(#3691): document the new global and narrow an invariant this change obsoleted The full suite surfaced two consequences of giving every CLI lane a budget key. `review.max_prompt_tokens` entered CONFIG_DEFAULTS without a matching entry in the planning-config reference, which config-field-docs guards. Documented, including the sentinel semantics a reader needs: a per-lane value overrides the global, `-1` means unset and inherits it, and `0` means "do not trim that lane" and is not unset. The #2797 federation guard asserted that "a lane with no model flag and no host owns no config keys". That held only because budget keys existed solely on the three local-server lanes, all of which have hosts. A lane can now legitimately own a config key for a third reason, so qwen tripped it. The assertion is narrowed rather than weakened: such a lane must still own no model key and no host key, and may own at most its own `review.max_prompt_tokens_per_reviewer.<slug>` — never another lane's. That is strictly more specific in the dimensions that still matter. Proven to still bite: hypothetically giving qwen a `review.models.qwen` key fails it with `model/host: review.models.qwen`. The name and comment cite #3691 for why the premise changed, so a reader sees a deliberate narrowing, not erosion. Checked the sibling assertions in that describe block; the other three do not rest on the obsolete premise and are untouched. Refs #3691 * fix(#3685): port the write-flag content-change contract to its three sibling sites #3685 fixed `phase complete`'s `roadmap_updated` / `state_updated`, which reported `fs.existsSync(path)` rather than whether the transaction wrote anything. Three sibling sites carried the identical defect and are ported here. - `cmdPhaseRemove` reported `roadmap_updated: true`, hardcoded. `updateRoadmapAfterPhaseRemoval` now returns whether the content changed and the flag reports it. #2640/#2974 already fixed `state_updated` at this same call site and left this one behind, so the correct shape was adjacent. - `cmdMilestoneComplete` reported `state_updated: fs.existsSync(statePath)` — byte-identical to #3685's bug in a different command. - `cmdMilestoneComplete` reported `milestones_updated: true`, hardcoded, never consulting the MILESTONES.md write. `gsd-core/workflows/remove-phase.md:100` extracts `roadmap_updated` for display and never branches on it, so the flip from always-true to content-based changes no workflow behavior. Verified by reading the step, not assumed. One trap found while implementing: the obvious in-memory `finalContent !== originalStateContent` comparison — copying `cmdPhaseComplete`'s shipped shape verbatim — gives a FALSE POSITIVE for milestone completion. `platformWriteSync` normalizes Markdown at write time, and the milestone-closure transform regenerates `## Current Position` fresh on every call, so its pre-normalize output always differs from the already-normalized file on disk even when the persisted bytes are identical. The comparison is therefore made against the post-write on-disk content. `cmdPhaseComplete`'s own comparisons are left untouched — their repeat-no-op tests pass, so they are not exposed to this artifact. `milestones_updated` has no reachable no-op: the MILESTONES.md write unconditionally appends an entry every call. Only the true direction is pinned, documented inline rather than faked with a passing test. Refs #3685 * fix(#3685): compare write-flag content through the writer's own normalizer An independent reviewer disproved a claim made while porting #3685's contract to its sibling sites: that `cmdPhaseComplete`'s comparisons were not exposed to the Markdown-normalization artifact already diagnosed in `cmdMilestoneComplete`. `platformWriteSync` normalizes on write — CRLF stripped, blank-line runs collapsed, a blank line inserted after a heading, a single trailing newline enforced. Every flag that compares the PRE-normalization in-memory string against the on-disk pre-image can therefore report a change when the persisted bytes are identical. `cmdMilestoneComplete` had been worked around by re-reading the file after the write; the other sites compared raw strings. All of them now go through one exported seam, `contentChangedAfterNormalize(filePath, before, after)`, which normalizes both sides exactly as the writer does. That removes the extra disk read the milestone workaround needed, and makes the sites agree by construction rather than by four independent implementations of one rule — the divergence the repo names as an anti-pattern. Reachability, stated precisely rather than uniformly: the seam is load-bearing at `cmdPhaseComplete`'s `roadmapUpdated`, `requirementsUpdated` and `stateUpdated`, where section-rewrite logic genuinely regenerates content into a different-but-normalization-equivalent shape. At `updateRoadmapAfterPhaseRemoval` it is defense-in-depth: the no-match branch never reassigns `content`, so the raw comparison was already correct there. The first analysis claimed the reverse; this is the corrected finding. Also fixes an unsound test premise the remote suite caught. The byte-identity precondition in `roadmap_updated is false when ROADMAP.md comes out byte-identical` asserted against a hand-authored, un-normalized fixture — so the very first write reformatted it and the file could not come back identical. The fixture is now written already-normalized, so the assertion compares a normalized pre-image against a normalized post-image and still fails if the flag regresses to a hardcoded `true`. Not platform-specific; it reproduces on macOS too, and the earlier local check simply never exercised it. The sibling true-direction and milestone tests were checked for the same premise and do not share it — they assert `notEqual`, or compare two post-write states produced through the same normalizing seam. Refs #3685 * chore(changeset): backfill PR number for #3691 fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
314ea20fa4 |
fix(#3663): fold path casing only on win32 in the w027 active-worktree check (#3793)
* test(#3663): failing-first rows for w027 path-casing normalization * fix(#3663): fold path casing only on win32 in the w027 active-worktree check * fix(#3663): close review findings — seam-owned compare key, deterministic case pin * chore(#3663): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
4af59f8dd3 |
fix(#3662): resolve managed hook node runners at hook-fire time (#3790)
* test(#3662): failing-first suite for runtime-resolving hook runners * fix(#3662): resolve managed hook node runners at hook-fire time * fix(#3662): close review findings and document the resolver * fix(#3662): close adversarial and security review findings * chore(#3662): backfill changeset pr number * test(#3662): honor win32 skip return and platform-aware sh runner pin * test(#3662): pin the bare win32-claude sh-hook shape omitting the bash runner --------- Co-authored-by: sim <sim@local> |
||
|
|
2972da4c9d |
enhance(#3619): ratchet the platform seam with local/no-private-binary-resolution (epic #3411 Phase 3) (#3636)
* chore(#3619): ratchet the platform seam with local/no-private-binary-resolution
Epic #3411 Phase 3, the ratchet. Scope revised with maintainer approval and
recorded on the issue: the epic's literal ask was a rule rejecting a bare-name
spawn outside the seam. Surveyed at
|
||
|
|
ac1b6d679f |
enhance(#3618): fold fallow-runner onto the canonical binary resolver (epic #3411 Phase 2) (#3633)
* chore(#3618): fold fallow-runner onto the canonical binary resolver Epic #3411 Phase 2. src/fallow-runner.cts was the fourth divergent implementation of Windows binary resolution the epic enumerated — candidateNames, isExecutableFile, findInPath, findInNodeModules, 40 lines. All four are deleted; resolveFallowBinary is one seam call. Two OPT-IN options were added to resolveExecutableBinary to make the fold behavior-preserving, both defaulting off so Phase 1's callers are byte-identical: prependPaths dirs searched before env.PATH, in order, through the identical per-directory candidate logic. This expresses node_modules/.bin-first precedence without env surgery — the rejected alternative re-introduced the spread-loses-the-proxy hazard the Windows lane caught in Phase 1, at every future call site instead of once. requireExecutable POSIX-only accessSync(X_OK); a no-op on win32 where mode bits do not mean execute. Opt-in rather than default because unconditional X_OK breaks #3445's suite, which stages candidates with plain writeFileSync and never sets an exec bit — the repo bans chmod in tests — so every one would resolve to null on POSIX. Deliberate behavior change on Windows: fallow's prior candidate list ended in a BARE fallow. The seam never tries a bare name there, so an extensionless file beside fallow.cmd is no longer resolved. That is the fix, not a regression — the extensionless file is npm's POSIX sh shim, which CreateProcess cannot run (#3275). Rows 7 and 8 of the design record it. Defect found while working, fixed inline: the resolution order was documented BACKWARDS as PATH-then-.bin in structural-pre-pass.md, docs/INVENTORY.md and four INVENTORY translations. The code has always been .bin first, and .bin first is correct — a project-local tool should beat a global one. The archived changeset is left alone as a historical record. fallow-runner had no test file at all. tests/fallow-runner.test.cjs is new (F1-F15) and the seam options are pinned by S1-S12 folded into the existing dispatch suite. RED proven by execution: with both source files stashed and build:lib re-run, 7 of 27 probe cases failed. Refs #3411 * chore(#3618): backfill changeset pr number 3633 * fix(#3618): assert both platform contracts in F4 instead of a POSIX-only premise Windows CI on #3633 failed F4. The test monkeypatched accessSync to throw and asserted resolveFallowBinary returned null — but that premise, that the X_OK check is consulted at all, is POSIX-only by design. requireExecutable is a deliberate no-op on win32 because Windows mode bits do not mean execute, so the staged fixture correctly resolved there. 40-design.md's negative-space section already states this carve-out verbatim. The test contradicted the design it was written from: fixtures were made platform-adaptive in the previous commit, and this assertion was left platform-blind. F4 now asserts BOTH contracts — null on POSIX, resolves on win32 — rather than skipping either. A t.skip on one lane would have been green and would have left the win32 carve-out unpinned by fallow's own entry point. Audited every other row for the same class. F1-F3, F5, F6, F11-F15 hold on both platforms; F7-F10 and S1-S12 inject platform explicitly and are unaffected. F4 was the only row with a single-platform premise. The local probe runs on one platform and structurally cannot catch this, which is why it was green — that limitation is now stated at the top of the probe so a green probe is not mistaken for platform coverage. The win32 branch was proven by injecting platform:'win32' with accessSync throwing and asserting it still resolves. Refs #3411 --------- Co-authored-by: sim <sim@local> |
||
|
|
bf87dd4156 |
enhance(#3617): one canonical Windows binary resolver in the platform seam (epic #3411 Phase 1) (#3621)
* feat(#3411): one canonical Windows binary resolver in the platform seam CONTEXT.md declares src/shell-command-projection.cts the single OS-facing seam, but Windows binary resolution had grown four divergent implementations outside it. #3445 folded two of them together — inside gsd-core/bin/gsd-tools.cjs, not the seam — so the declaration stayed untrue and execTool still had no handling at all. Lift the resolver into the seam as resolveExecutableBinary, and export the half that actually executes as projectSpawnInvocation: CreateProcess cannot run a .cmd/.bat, so the cmd.exe mediation is inseparable from the lookup and splitting them is how the copies accumulated. cmd.exe is invoked with an explicit argv array, never shell:true — CVE-2024-27980's vector and Node 26's DEP0190. execTool now resolves on win32. POSIX is a strict no-op by construction, which matters: execTool rates CRITICAL blast radius (167 symbols, 53 files). gsd-tools.cjs deletes its private scan and its private mediation and delegates. Two semantics grown beyond #3445's resolver, both additive: a name already carrying a PATHEXT-listed extension is tried as-is before the append loop, and a suffix outside PATHEXT is not treated as an extension. Refs #3411 * fix(#3411): keep mediating a declared .cmd that PATH resolution misses Standards review caught a narrowing against the code this replaces. gsd-tools.cjs computed `target = resolveSpawnBinary(binary) || binary` and keyed the shim test on `target`, so a declared .cmd mediated whether or not PATH resolution found it. That is load-bearing: resolveExecutableBinary scans PATH only, while `cmd.exe /c` also finds a batch file in the current directory. Mediation now keys on the target — resolved path, else declared name. The ENOENT contract still holds for BARE unresolved names, which is the case it was written for. P9/P10 pin both halves. Spec review found E1/E2/E3/E5 promised by 50-test-matrix.md but never written; added. E3 is the integration proof that the CVE-relevant mediation fires through execTool, not only through projectSpawnInvocation in isolation. Also adds the CONTEXT.md glossary entry for the seam's new resolution ownership (a PR gate) and the changeset fragment. Refs #3411 * fix(#3617): pass mediated cmd.exe arguments verbatim so metacharacters cannot inject The isolated security pass found the mediation shape carried an argument-injection surface. libuv's quote_cmd_arg force-quotes an argv element only when it contains a space, tab, or quote — never for a cmd metacharacter — and cmd.exe re-parses everything after /c. So an arg of a&calc arrived unquoted and cmd ran calc. Node's own CVE-2024-27980 escaping cannot help: it fires only when the spawned FILE is the .bat/.cmd, and here the file is cmd.exe. Caret-escaping is not a fix. It is correct only when libuv does not quote, and libuv quotes whenever the arg also contains a space — no per-arg transform is right in both cases. So build the command line and pass it through verbatim, the shape Rust's std adopted for the sibling CVE-2024-24576: one outer quote pair that cmd /c strips, every token inside force-quoted, embedded quotes doubled. An argument containing CR or LF is refused rather than mediated — a newline cannot be represented in a Windows command line, so mediating would silently truncate. Failing visibly is correct. Known limit, documented at the seam: %VAR% still expands inside a /c string and has no escape outside a batch file. That is information disclosure, not arbitrary execution, and is the same limit Rust's std documents. This was byte-for-byte the shape #3445 shipped, so the fix closes it for the reviewer-lane spawn path too, not only for execTool's newly reachable route. Refs #3411 * docs(#3617): document the subprocess-execution security posture Adds Layer 4 to the security model: why GSD never uses shell:true for binary invocation (CVE-2024-27980, Node 26 DEP0190), why resolution is explicit and never tries the bare name on Windows (the npm extensionless-shim trap behind #3275), and why .cmd/.bat mediation builds a verbatim force-quoted command line rather than relying on default escaping — Node's own CVE protection cannot fire once the started program is cmd.exe. The residual %VAR% expansion limit is stated plainly under Trade-offs rather than left implicit: it is information disclosure, not arbitrary execution, and callers passing untrusted text to a Windows .cmd should not assume the value arrives byte-identical. Docs-only; no code change. Refs #3411 * chore(#3617): backfill changeset pr number 3621 * fix(#3617): read PATH, PATHEXT and ComSpec case-insensitively The Windows CI lane on #3621 failed E5, and the root cause was a defect in the implementation, not the assertion. Windows names the variable Path, not PATH. process.env is a case-insensitive proxy, so process.env.PATH works — but execTool builds { ...process.env, ...opts.env } whenever a caller supplies opts.env, and spreading discards the proxy while keeping the OS's actual casing. The exact-case env['PATH'] lookup then returned undefined, the PATH scan saw zero segments, resolution returned null, and the change degraded to precisely the spawn ENOENT it exists to fix. ComSpec and PATHEXT had the same exposure. #3445's tests never caught it because they pass uppercase keys explicitly, and neither did the Linux remote runner — this is a defect only the Windows lane could see. _envGet resolves a variable by exact match first (so a canonical caller pays no scan) and falls back to a case-insensitive sweep. R23 and P16 pin it and were proven RED by execution: with the fix stashed and build:lib re-run, R23 returned null and P16 returned the cmd.exe default. R24 was rewritten because the first version was vacuous — it staged foo.CMD, so the default PATHEXT already contained .CMD and it passed against the broken code for the wrong reason. It now stages foo.XYZ, an extension absent from the default, and carries a negative control asserting that dropping the Pathext key yields null. Re-proven RED the same way. E5's assertion was corrected alongside the fix: 'PATH' in options.env expressed the wrong contract. It now checks case-insensitively for the key. Refs #3411 * fix(#3617): execTool spawns the declared name unless mediation is required The Windows full-test lane on #3621 failed tests/graphify.test.cjs — the python3 identity check asserted 'python3' and got the absolute resolved path C:\hostedtoolcache\windows\Python\3.12.10\x64\python3.EXE instead. Those tests are correct and the change was wrong. They pin a long-standing contract — execTool spawns the program name it was given — by spying on spawnSync's first argument, and routing every win32 call through the projected invocation broke it. Resolving a .exe buys nothing. libuv's CreateProcess path already performs PATH + PATHEXT search, which is why spawning a bare 'node' has always worked on Windows. The only case the OS genuinely cannot spawn is a .cmd/.bat. So execTool now adopts the projection only when mediation actually happened — windowsVerbatimArguments is exactly that flag — and otherwise passes the declared program and args through untouched. 40-design.md already rejected gratuitous change for this reason: symmetry is not worth a behavior change to 53 files that fixes nothing. That reasoning was applied to POSIX and missed the win32 non-batch case. Rows 5 and 20 now record it, and the CONTEXT.md glossary states the caller-choice rule. deps.spawn deliberately still adopts the resolved path: its hasBinary probe answers from the same resolver, so probe and spawn must agree on the exact file (#3445). The asymmetry is now documented at both call sites rather than latent. E7 pins the restored contract and was verified by executing execTool against a monkeypatched spawnSync: python3 in, python3 spawned. Refs #3411 --------- Co-authored-by: sim <sim@local> |
||
|
|
dc3c81e93d |
chore(#3212): src/pattern.cts is the sole owner of runtime-value regex construction — Phase 1 (#3416)
* test(#3412): failing-first suite for the pattern-construction seam Phase 1 of epic #3212 (ADR-3212 §1/§2/§7). Tests only — src/pattern.cts and eslint-rules/no-adhoc-regex-escape.cjs do not exist yet, so both suites fail with MODULE_NOT_FOUND, which is the intended RED. Locks the measured behavior rather than the assumed behavior: RegExp.escape hex-escapes the leading character of nearly every string ("abc" -> "\x61bc"), so the suite asserts match-equivalence against an inlined historical oracle (the implementation being deleted) rather than byte-equivalence of pattern text — 200 seeded fast-check runs plus a fixed corpus, 0 mismatches. Also locks the latent character-class range bug this phase fixes as a side effect: a hyphen-bearing value interpolated into [...] currently forms a real range and matches an unintended character; post-migration it must not. * chore(#3412): src/pattern.cts owns runtime-value regex construction Phase 1 of epic #3212 (ADR-3212 §1/§2/§6/§7). Adds the pattern seam delegating to the built-in RegExp.escape, deletes every hand-rolled copy, and raises the Node floor to the Active LTS line. The census was low, three times over. ADR-3212 counted 10 copies; a graph query found 12; the new lint rule — once live — found 27 more. The difference is that the census counted named helper FUNCTIONS while the rule counts the escape SHAPE, so inline .replace(<class>, '\$&') copies were never in scope. ADR §1's actual requirement is that no module outside the seam escapes a value for regex use, so all of them are, and CLAUDE.md's no-defer rule makes them this change's work. Fourth consecutive epic here whose copy count was low — the argument for ADR-3180 Amendment 3's "state N found by the guard" rule. Also corrected mid-implementation: the survey reported phase-id.cts's escapeRegex had 0 external importers. It had 8 production importers, making its removal a public-surface change to an ADR-2121-owned module and requiring an update to that ADR's locked-surface test. Blast radius revised Medium-High -> High. RegExp.escape is match-equivalent but NOT text-equivalent: it hex-escapes the leading char of nearly every string ("abc" -> "\x61bc"). Equivalence is proven by a seeded fast-check property test against the deleted implementation as oracle. It also fixes a latent bug: a hyphen-bearing value interpolated into a character class previously formed a real range and matched an unintended character. Node floor 22 -> 24 (RegExp.escape is Node 24+), across engines, .nvmrc, package-lock, 9 CI matrix entries, and 5 docs. The aggregate `required-tests` context is unchanged and no job was added or removed, so branch protection cannot be orphaned by the dropped lanes. Enforced by eslint-rules/no-adhoc-regex-escape.cjs (shape-matched, with structural provenance for reviewed pattern-fragment constants rather than a name heuristic) plus a whole-tree companion guard covering the directories ESLint's globs miss. * fix(#3412): close the _SOURCE guard evasion, correct two false claims Three findings from the orthogonal review pass, all fixed. 1. The ESLint rule's `_SOURCE` provenance fallback was pure identifier- name matching with no binding check, so `new RegExp(userInput_SOURCE)` — a function parameter — sailed past the guard. That is the same rename-evasion class issue #3410 documents, reopened by the very fallback meant to complement the structural check. Now bound to the identifier's actual binding kind: import, require-derived const, or module-scope const; parameters, `let`/`var`, and unresolvable bindings fail closed. Four RuleTester cases cover the evasion and prove the legitimate cross-module case still passes. 2. src/pattern.cts's own header carried the stale pre-correction counts (12 copies / 17 call sites) while CONTEXT.md and the design doc carried the corrected ones (~39 / ~44) — a self-contradiction inside the PR whose entire purpose is deleting divergent copies. Rewritten, preserving the durable lesson: a named-function census cannot see inline copies; only a shape-matching guard can. 3. The claim that all deleted copies threw TypeError on non-string was false. phase-id.cts's copy — the one with 8 external importers — did String(value).replace(...) and never threw. The seam's locked signature does not coerce, so this is a real, now-disclosed behavior change rather than the pure preservation the tests asserted. Audited all 32 invocations across the 8 importers and 6 in-file callers: every one is safe by construction (upstream truthy guard or a string-producing derivation), verified by runtime probe against the compiled modules rather than by TS compilation, which cannot see a runtime undefined. Corrected the false claim in both the test comment and the design doc, and added it to Known limits. * docs(#3412): add Changed changeset for the Node 24 floor The only user-visible break in this phase. The escape-behavior change is internal and match-equivalent, so it carries no user-facing note. * fix(#3412): resolve the seam's require graph in script fixtures and packaging Checkpoint 2 came back red with 90 failures on the node24 lane. Three distinct defects, all introduced by routing scripts/ through the new pattern seam, none reproducible by any local gate: 1. ~82 failures — tests/adr-index-gate.test.cjs and tests/removed-but-needed-lint.test.cjs copy a scripts/*.cjs into an mkdtemp fixture and spawn it there (necessary: those scripts resolve their scan root from __dirname/.., so running the real script would scan the real repo). Each harness hand-listed the dependencies to copy alongside. Adding require('../gsd-core/bin/lib/pattern.cjs') to gen-adr-index.cjs made both lists silently incomplete -> MODULE_NOT_FOUND, plus 17 downstream 'did not emit parseable JSON' failures from the same crash. Fixed as a class, not an instance: new tests/helpers/copy-script- fixture.cjs walks a script's transitive static relative-require graph and copies it, so dependencies are derived and never re-declared. It throws (naming the unbuilt artifact) instead of letting the child die with a bare MODULE_NOT_FOUND. Verified for all four seam-consuming scripts: gen-adr-index, lint-removed-but-needed, gen-loop-host- contract, sync-runtime-launcher. 2. 2 failures — scripts/ ships wholesale but eslint-rules/ does not, so the new scripts/lint-no-adhoc-regex-escape.cjs would be MODULE_NOT_FOUND in a published install (#2858 guard). Excluded from the tarball, matching the existing precedent for gen-emitted- baseline.cjs, which is excluded for the identical reason, and locked with a test modeled on that one. Confirmed against a real npm pack: 890 files, 0 from eslint-rules/, and gsd-core/bin/lib/pattern.cjs present (so the other four scripts' requires are legitimate). 3. 6 failures — tests/phase-id.test.cjs asserted the literal escaped source text ('0*29', 'PROJ-42'). RegExp.escape is match-equivalent to the retired hand-rolled escaper but NOT text-equivalent: it hex- escapes the leading character and all hyphens ('0*\x329', '\x50ROJ\x2d42'). Verified NOT a behavior change — 576 match decisions across all three real interpolation prefixes, zero divergence. Those tests now compile each source into the same heading regex src/roadmap.cts's searchPhaseInContent builds and assert what matches and what does not, including the 'i'-flag canonicalization the hex escape has to preserve. Re-pinning the new literals would have rebuilt the same brittleness one layer down. Adds a test for the property the escape exists for: a dot in '1.2' must not act as a wildcard. Also shares one definition of 'a require' between the packaging guard and the fixture copier, so the two cannot disagree about what they scan. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): refuse to copy a fixture dependency outside the fixture root copyScriptWithDeps resolved each relative require and joined the repo-relative result onto fixtureRoot. A require resolving OUTSIDE the repo yields a '../'-prefixed relative path, so path.join climbed out of the fixture and wrote into the surrounding temp dir (verified: repoRoot=/repo + depAbs=/etc/passwd wrote /tmp/etc/passwd). No script in the tree does this today, so this closes an available escape rather than an active one. Refuses via the existing unresolved- require path so the failure names the offending specifier. Covered by a negative proof that the guard fires and that nothing lands outside the fixture. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): parse requires instead of pattern-matching them; restore the foreign-prefix contract Applies all findings from the second orthogonal review round, re-run because real code changed after round 1. HIGH (security) — extractRequires stripped BLOCK comments before LINE comments, so a '//' comment containing '/*' opened a phantom block comment, and a '//' inside a string literal truncated the line. Both hid real requires: 'const u="http://x"; require("./real.cjs")' returned [], and four real requires in gsd-core/bin/gsd-tools.cjs were invisible. Replaced with a real AST parse via espree. This is ADR-3212's own Decision 4 — tokenizer-first for stateful grammars — applied to the case it describes; comment/string/regex nesting is exactly such a grammar, which is why the regex version was wrong. The function was moved byte-identical out of the #2858 packaging guard, so the bug PRE-DATES this branch and has been a live blind spot there: a shipped script could have required an unshipped path undetected. Fixing it makes that guard strictly stronger than on next. espree is promoted from a transitive eslint dependency to an explicit devDependency rather than relying on hoisting. The script parse attempt sets ecmaFeatures.globalReturn because Node wraps CommonJS bodies in a function, making a top-level return legal — scripts/check-coverage-gate .cjs relies on it, and without the flag the guard throws on a file it is supposed to scan. Verified 0 unparseable across all 324 .cjs/.js under scripts/, bin/, and gsd-core/bin/, and 0 new violations against a real npm pack, so the exact extractor does not newly fail the guard. MEDIUM (security) — the repo-containment check guarded dependencies but not the entry path. One escapesContainment predicate now guards both. LOW (security) — containment was lexical while fs follows symlinks, and a directory symlink could mint a fresh dedupe key per level. realpath now resolves both repoRoot and each dependency before the decision, and the realpath-derived path is the dedupe key. Destination layout still uses the original repo-relative path, so copied trees are unchanged. MAJOR (standards) — the round-1 behavioral rewrite of phase-id tests lost the foreign-prefix contract: every assertion was satisfied by an impl returning [A-Z]+\x2d42, i.e. ANY project code — the exact #3599 bug class the exact-source prevents. The literal assertions it replaced were catching this. Now asserts the compiled regex REJECTS a different prefix with the same number. MAJOR (standards) — the test hand-duplicated production's heading regex with no parity guard (CLAUDE.md's 'Generative Fix Divergence'). Removed the parallel surface instead of policing it: src/roadmap.cts exports buildPhaseHeadingRegex, searchPhaseInContent calls it, the test imports it. Byte-identical .source and .flags verified for both escaped forms. MINOR — '..foo' no longer false-flagged as an escape; the inverted spurious-vs-missing doc claim corrected; the dead allow-test-rule header removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3412): backfill changeset pr number to 3416 * fix(#3412): make the escape guard's own regex linear, reword an injection-scan collision Two CI failures on PR #3416, both in code this branch added. CodeQL js/redos (high) — REPLACE_CALL_RE's outer alternation let a bracket run be consumed EITHER by the character-class branch OR one character at a time by the trailing catch-all, so a failing match explored both parses of every pair. Measured on the real regex: n=26 -> 204ms, n=28 -> 791ms, n=30 -> 3475ms, a clean 2^n. This script scans repo source, so a file with a long bracket run after '.replace(/' would hang CI outright — a guard against undisciplined pattern construction was itself the worst pattern in the diff. Fixed the way ADR-3212 already prescribes: the catch-all branch now excludes '[' and ']' so a bracket can only be consumed by the class branch (this is what makes it linear), and every quantifier is bounded (the locked bounded-quantifiers decision) as a second line of defense. Now 0ms at n=2000. Disclosed coverage tradeoff, recorded at the constant: a regex literal with a BARE unescaped ']' outside a class is no longer matched by this backstop. No census shape has that form, and the AST rule remains the primary detector. Verified the guard did not go blind doing it: a real census-shape violation is still reported, and an allow-adhoc-regex-escape suppression comment is still honored. Regression test drives the exported findViolations on a 2000-repetition adversarial input and asserts the RESULT. It makes no wall-clock assertion — elapsed-time tests are forbidden — so a regression surfaces as a harness timeout, which is the correct signal. Prompt injection scan — 'must not act as a regex wildcard' in a test comment matched the scanner's jailbreak pattern act\s+as\s+(a|an|if| my). Reworded to 'behave as'. Deliberately NOT allowlisted: silencing a whole test file over one phrase would blunt the scanner permanently, and the comment has nothing to do with injection. Neither failure was reachable from the remote runner — CodeQL and the injection scan are not in that matrix, so the sha it passed was green and still wrong. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bcf7b04864 |
chore(#2896): convert CONTEXT.md prose defect registry into enforced gates (#3325)
* chore(#2896): convert CONTEXT.md prose defect registry into enforced gates Squashes the prior 4-commit sequence and fixes defects found while resuming this branch: 5 orphaned/corrupted DEFECT fragment lines left by an earlier botched edit, 17 "Source of truth: Memtrace `find_symbol`" placeholders that had destroyed real file-path citations, and 3 DEFECT.GENERATIVE-* entries merged into one RULESET.GENERATIVE-FIX predicate (policy, not an unenforced defect) to satisfy the zero DEFECT.<NAME>.<field>= acceptance criterion. Six mechanizable defects get real gates: DEFECT.UNBOUNDED-SUBPROCESS (eslint-rules/require-subprocess-timeout.cjs), DEFECT.CANARY-VERSION-LEAK (scripts/lint-canary-version-leak.cjs + version-gate.yml), DEFECT.CHANGESET-PR-FIELD-DRIFT (findPrFieldDrift in changeset/lint.cjs), DEFECT.FRONTMATTER-SCALAR-BROAD-GREP, DEFECT.REMOVED-BUT-NEEDED, and DEFECT.DEFAULT-FLIP-DOCUMENTATION (new lint scripts, wired into lint:ci). Already-enforced and unenforceable prose entries are deleted; the gate is the record. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#2896): route the new lint tests' subprocess calls through the bounded process-seam helper The 4 new test files for this PR's lint checks called cp.spawnSync/ execFileSync directly with no timeout, tripping this repo's own existing local/no-unbounded-spawn ESLint rule. Route every one through runNode/gitOrThrow (tests/helpers/process-seam.cjs, tests/helpers/git-fixture.cjs) instead, matching the pattern already used elsewhere in the suite (e.g. tests/changeset-lint.test.cjs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: register claude-orchestration.cjs and regenerate stale generated indexes Pre-existing drift on next, unrelated to this PR's own change, surfaced by running lint:ci as part of verifying #2896: two cli_modules (claude-orchestration.cjs, write-set.cjs) landed without a manifest regen, and CONTEXT.md's own edits in this PR staled its two generated indexes. Adds the missing docs/INVENTORY.md row for claude-orchestration.cjs (write-set.cjs already had one — only its manifest entry was stale) and regenerates docs/INVENTORY-MANIFEST.json, docs/CONTEXT-INDEX.json, and examples/dynamic-context-management/CONTEXT-INDEX.json. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#2896): default-flip-documentation lint's local fallback base was main, not next Found in review: every other base-ref fallback in this repo (see scripts/changeset/lint.cjs's DEFAULT_BASE, #2988) defaults to `next`, the integration branch every PR actually targets — `main` is the release branch. This script's local fallback (used only when GITHUB_BASE_REF is unset, i.e. never in CI, but potentially on a local or direct invocation) diffed against the wrong ref. No test exercised the unset-env-var path, so it shipped unnoticed; every e2e test sets GITHUB_BASE_REF explicitly and is unaffected by this fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#2896): stale eslint comment, overclaiming CONTEXT.md wording, and an incompletely-regenerated manifest Found by the isolated Standards code-review pass: - eslint.config.mjs's require-subprocess-timeout comment said "'warn' for now... flip to 'error' once migrated" while the rule already shipped as 'error' with all 8 sites migrated in the same commit — described a state that never existed. - The CONTEXT.md pointer block claimed the rule's bounded call sites "never throw", but roadmap-upgrade.cts's pre-mutation clean-tree check correctly still throws on failure (it gates a destructive real-run migration; degrading to "assume clean" would risk clobbering uncommitted work) — softened the claim to describe both shapes accurately instead of overclaiming one. - docs/INVENTORY-MANIFEST.json's claude-orchestration.cjs/write-set.cjs entries from the prior "fix: register claude-orchestration.cjs..." commit didn't actually land — re-running the generator now includes them; lint:generated-sync is green. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#2896): backfill changeset pr field with the real PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#2896): normalize buildCorpus file paths to POSIX in lint-removed-but-needed Windows CI caught it: path.relative(root, abs) returns backslash- separated paths on Windows, but findSurvivingReferences's package-lock special case does file.startsWith('.github/workflows') — a forward- slash literal. On Windows the check silently never matched, so tests/removed-but-needed-lint.test.cjs's real-defect-shape fixture got exit 0 instead of the expected exit 1. Normalize at the production source (RULESET.CONTENT-PATH-NORMALIZATION) rather than the test side. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
0e6fa2e2cf |
enhance(#3118): close the dead injectables and the shell projection follow-on — Wave 4 (#3124)
* test(#3118): failing-first coverage for the dead injectables and the shell projection Adds the counter-tests Wave 4 closes against, before any fix: - antigravityWatermark had zero test references. The four existing tests that look like watermark coverage hand the fallback a literal mark and never call the producer, so nothing pinned whether a real run's mark is correct. Covers all six branches plus the non-object cache classes. - Pins the fail-open: a transcript read that throws reports lines:0, indistinguishable from a genuinely empty transcript, and the consumer then replays a previous run's review as this run's. - Pins the export-line escaping across the repair, persist and win32 bash lanes, including the parity assertion that they must not diverge. - sliceCurrentPositionSection: empty-vs-absent, fenced heading, second occurrence, H3, CRLF. - Proves deps.progressProvider is inert by supplying a throwing stub to all ten transition intents. Verification through the remote runner only. Refs #3118 * fix(#3118): distinguish an unreadable transcript from an empty one antigravityWatermark's final read can throw on a transcript that indisputably exists. It returned lines:0, which is the same value a genuinely empty transcript produces, so the caller could not tell the two apart. antigravityTranscriptFallback derives its skip from that count. A mark of {convId:'c1', lines:0} for a conversation that pre-dates the run makes it skip nothing and return the last PLANNER_RESPONSE in a transcript written before this run started — a previous review presented as this one's, which is exactly what the function's own 'never stale' docstring promises cannot happen. The unreadable case now sets unreadable:true and the fallback declines for a same-conv-id unreadable mark. An absent or empty transcript is untouched: those genuinely have zero prior lines. * fix(#3118): escape the export line for the file it lands in, not the echo Three lanes emit export PATH="<dir>:$PATH". repair escaped it with escapePosixDoubleQuoted; persist and the win32 Git Bash lane escaped it with escapeSingleQuotedShellLiteral instead. The single-quoting is correct for the echo, so nothing runs when the user pastes the command. But the bytes appended to ~/.bashrc are the export line itself, and inside double quotes in an rc file a $(...) or a backtick in the directory name is command substitution that runs on every new shell. Those characters are legal in a path on both POSIX and Windows, so the path was reachable. projectPathExportLine is now the single source of that line and escapes for its final rc-file context; each lane still applies its own transport escaping on top. fish keeps the single-quote escaper — its value really does stay single-quoted. The cmd.exe lane interpolated into a cmd double-quoted string with no cmd-level escaping, so a quote closed the region and &cmd& ran. A quote is reserved on Windows and cannot appear in a real path, so there is no correct command to suggest: the win32 lanes now fail closed for one. Metacharacter-free paths render byte-identically on every lane. * fix(#3118): drop a stray carriage return and a deps field nobody reads locateCurrentPosition subtracted a fixed one byte to exclude the newline before the next heading, which assumes LF. On a CRLF document the slice kept an unpaired trailing carriage return. It now walks back over the newline and over a preceding carriage return if there is one. StateTransitionDeps also required a progressProvider that 33 sites supplied and no site ever called. A required field nothing reads widens the module's interface without changing its implementation, which is the shape epic #3051 cites as its reason for refusing blanket injection. Removed along with the ProgressRecord alias that existed only as its return type; state-document.cts's unrelated interface of the same name is untouched. * fix(#3118): stop an empty span duplicating bytes, and name the empty results Three findings from the isolated review pass. locateCurrentPosition could return end < start when the section was empty and the next heading followed with no blank line between. Every mutator splices with slice(0,start) + body + slice(end), so an inverted span duplicated the region between them — a blank line silently inserted into STATE.md on every transition, two bytes on CRLF. The span is now clamped, and an empty section is a zero-length span, which is what it always meant. The win32 fail-closed path left the installer printing 'Add it with one of:' with nothing under it. An empty shellActions folded two different facts together, so projectPathActionProjection now carries a frozen PATH_ACTION_REASON and the installer branches on it. Two empty results with different causes staying distinguishable is the subject of the epic this belongs to. fish_add_path parses a leading dash as an option, so a directory named -v printed 'No paths to add' instead of being added. Verified against fish 4.8.1: the end-of-options separator fixes it. Replaces the console-prose test the second fix first arrived with — a regex over captured stdout is what CONTRIBUTING prohibits, and the typed reason is the surface it asks for instead. * fix(#3118): escape TOML control characters, and stop a test name overstating Five findings from the two review axes. escapeTomlDoubleQuotedString escaped only backslash and quote. TOML basic strings also require U+0000-U+0008, U+000A-U+001F and U+007F to be escaped, so a value carrying a raw newline or NUL wrote a config.toml no parser accepts — rejecting the whole file, not just that value. Four of its call sites write real config. Tab stays raw; the grammar exempts it. The byte-identity test claimed every lane was unchanged for an ordinary path, which is false: fish now takes the end-of-options separator on every path, not only hostile ones. Renamed, and the one intended delta now has its own named test instead of hiding inside a claim that read as broader than it was. Also: exact-equality assertions in place of substring checks that could pass on a subtly wrong escape, newline and null-byte cases for all five quoting primitives, and a temp dir registered with t.after so it is removed when an assertion fails. * docs(#3118): add the changeset fragments * fix(#3118): degrade instead of throwing on a null conversation cache A cache file whose whole content is the literal null — what a truncated or zeroed write leaves behind — made both antigravityWatermark and antigravityTranscriptFallback throw. JSON.parse('null') succeeds, so the try/catch wrapping the parse never fired, and resolveConvId then called hasOwnProperty on null. Both functions advertise the opposite; the existing test next to them is named 'a missing cache or transcript degrades to empty, never throws'. Parsing successfully is not the same fact as the payload being usable, and a guard that only wraps the parse cannot tell them apart. resolveConvId is now total for any non-object input, so one guard covers both callers. Caught by the null case in this wave's own cache matrix. * test(#3118): correct a stale fish expectation and a parity comparison The pre-existing 'POSIX persist mode escapes single quotes' test pinned fish_add_path without the end-of-options separator this wave adds, so it asserted behavior that is no longer correct. A repo-wide scan found one such hardcoded expectation; every other site derives its expectation from the projection. The new parity test compared the token from a POSIX path against the win32 lane, which posix-normalizes its input first — two different inputs, so the tokens differed for a reason that had nothing to do with the parity it claims to check. It now derives the win32 expectation from the same input the lane receives. * docs(#3118): reword a comment the injection scanner reads as an instruction The scanner pattern act\s+as\s+(?:a|an|the)\s+ carries no word boundary, so 'the same fact as the payload' matched on the tail of 'fact'. Reworded per the documented remedy for this collision. The missing boundary is a scanner defect rather than a prose problem — any contributor writing 'fact as the' trips it — but the pattern is gate plumbing, which the sibling epic owns, so it is surfaced rather than changed here. * chore(#3118): backfill changeset pr number to 3124 * chore(#3118): backfill changeset pr number to 3124 * fix(#2784): make the negation scan single-pass and index it correctly Three defects in the negation suppression added by #3127, all in one block, none of which had a test. The pair scan was verbs.some(nouns.some(...)) with a slice and a split per pair, so it grew cubically with clause length: 1.1ms before that PR and 8462ms after, on 800 verb+noun pairs in one clause. api-coverage's property test generates documents large enough to reach the runner's 600s file cap, which is why it hangs as 'fail 0, cancelled 1' rather than failing an assertion. Every (verb, noun) window is a subset of the single widest one, so one scan of that window answers the same question in a linear pass. Verified equivalent against the old predicate over 20,000 generated clauses. Both checks also subtracted clause.start from offsets that collectTerm- Matches already returns clause-local. The first clause on a line has start 0 so it worked there and nowhere else: later clauses went negative, and slice reads a negative index from the end, so suppression silently examined unrelated text. The comment claimed 'without any API integration' was suppressed. It is not — the qualifier sits outside the two-word lookback and the noun precedes the verb. Widening the window would trade a false positive that costs one declaration line for a false negative that slips a real integration past a blocking gate, so the behavior stands and the comment now says so. Pinned by a test. The qualifier sets were also rebuilt for every line of every document. |
||
|
|
7e1004d89e |
test(#3056): add the in-process fault-injection adapter, and normalize execGit's shape (#3077)
* refactor(#3071): normalize execGit's call and result shape, unify ExecGitFn ExecGitFn was declared four times. Three were hand-copies of one signature and two of those were wrong: they typed exitCode as number|null when _spawnResult returns `result.status ?? 1` and can never yield null, weakened signal from NodeJS.Signals to string, and widened error from Error to unknown. Only verification.cts got it right, via `typeof execGit`. The root cause was a missing export: SpawnResultOutput was declared without `export`, so no other module could name the return type of execGit. Three authors independently hand-copied it instead. Exported now. Normalizing the type alone would have left the pressure that caused the divergence, so the function is normalized on both sides. It now ACCEPTS every call its consumers make — worktree-safety's declaration could not express an env-carrying call at all — and RETURNS every result code they need: timedOut moves into _spawnResult, so execGit, execNpm and execTool all carry it and the one extension that justified a separate type disappears. All four sites are now `typeof execGit` with nothing left to restate. timedOut reuses the existing isSpawnTimeout predicate introduced by #3050 rather than re-deriving it. That predicate checks error.code === 'ETIMEDOUT' only; the signal === 'SIGTERM' conjunct was deliberately dropped there because Windows does not reliably report SIGTERM and requiring it risks a false negative. There is no false-positive risk, and a test proves it: an externally-delivered SIGTERM leaves error null, so it is still not reported as a timeout. No dead null-checks surfaced. Every exitCode comparison in the two affected modules is === 0, !== 0 or === 128 — never a null guard — so the nullable declaration had never been written against. Closes #3071 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3056): add the in-process fault-injection adapter Adds tests/helpers/faulty-deps.cjs — makeFaultyGit() and withFaultyFs() — so a module's error branch can be driven deterministically and its degraded verdict asserted, instead of a counter-test that only proves the call did not throw. makeFaultyGit returns a value structurally assignable to `typeof execGit`, so one stub satisfies all four seams that the #3071 normalization collapsed into that single shape. A parity test drives the same stub through a real injectable entry point in each of worktree-safety, git-base-branch, worktree-base-ref and verification; it fails the moment any of them re-grows its own shape. Faults are scoped rather than global — by argv predicate and by call ordinal — because a fault adapter that faults everything looks like it works and proves nothing, and because verification.cts's two-call error handling needs to fault the second call only. Invocations are recorded so a test can assert an exact call count. The timeout fault carries error.code === 'ETIMEDOUT', and a test asserts the real isSpawnTimeout predicate matches it, so the fixture cannot drift from the production definition of a timeout. A companion test asserts an externally delivered SIGTERM with a null error is still NOT reported as a timeout — the false-positive guard for #3050's dropped conjunct. withFaultyFs restores in a finally so a throwing body still restores, patches only the named methods, and nests without clobbering an outer saved original. It never uses chmod: that no-ops under root, so the test would pass with zero coverage in root Docker/CI. The adapter is in-process via deps only. The Phase 1 process seam is documented as deliberately not a fault-injection surface — it cannot distinguish an injected timeout from a genuine bench OOM and would retry it. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3071): add changeset fragment for the execGit normalization Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3071): route the last two timeout checks through the shared predicate An isolated review found this branch had normalised the timeout verdict but left two callers still hand-rolling the fragile version of it. check-command-router's runBoundedShell computed `timedOut: r.signal === 'SIGTERM'` while the correctly-derived `r.timedOut` sat on the same result object. graphify's execGraphify branched on `result.signal === 'SIGTERM'`, with a comment asserting the very premise isSpawnTimeout exists to reject. Both fail in both directions. On Windows a genuine timeout is not reliably reported as SIGTERM, so the guard silently fails to fire — the false negative #3050 was raised for. And an externally-delivered SIGTERM is not a timeout at all, so the check also fires when it should not; isSpawnTimeout avoids that because `error` is null in that case and it keys on error.code. Both now read the derived `timedOut`, and graphify's comment states the actual rule instead of the fragile assumption. Also replaces a vacuous test: "execGitDefault now accepts env" never called execGitDefault (it is unexported), called execGit — whose signature already accepted env before this branch — and asserted only that exitCode was a number, which would pass whether or not the change under test existed. It now proves env reaches the child by asserting `git var GIT_EDITOR` returns the injected sentinel. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3071): make the graphify timeout fixture faithful to a real timeout The remote matrix failed on both Linux lanes: graphify's "returns exitCode 124 on timeout" got 1 instead of 124. The fixture was wrong, not the production change. It stubbed spawnSync as { status: null, signal: 'SIGTERM', error: undefined }. That is not a timeout. A real spawnSync timeout also sets error.code 'ETIMEDOUT'; a SIGTERM with no error is an externally delivered signal — a kill. The old `result.signal === 'SIGTERM'` check accepted it as a timeout, which is the false positive the shared predicate exists to reject, so this test was locking that bug in rather than guarding against it. The fixture now carries a real ETIMEDOUT error and all three original assertions pass unchanged. A counter-test is added alongside it: an externally delivered SIGTERM with no error must NOT be reported as a timeout. That is the assertion whose absence let the false positive live. Swept every SIGTERM/SIGKILL stub under tests/ for the same unfaithful shape. No other instance: the worktree-safety, worktree-base-ref and commit-staging fixtures already set ETIMEDOUT, and the remaining hits are either deliberate external-kill tests or feed code that never consults timedOut. Two sites keep their own signal check deliberately and are NOT changed: capability-source.cts:1301,1386 fail closed on ANY abnormal termination, which is correct — reading timedOut there would stop it failing closed on a kill and let it parse stdout from a killed process. Their reason strings, and check-latest-version.cjs:115, label any signal as "timed out", which is imprecise wording over a correct verdict, not a silent failure. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3071): backfill changeset pr number to 3077 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4eb8e3648c |
fix(#3050): consolidate the spawn-timeout predicate and propagate the unresolved-root reason (#3060)
* chore(#3050): changeset and review artifacts for the follow-up Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3050): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e35534c3ec |
fix(#2236): add hookShell parameter for PowerShell call operator (#2261)
* fix(#2236): add hookShell parameter for PowerShell call operator Windows Claude Code with a PowerShell hook runner failed on every hook with 'Unexpected token' — the installer emitted bare quoted paths that PowerShell parses as string literals, not commands. The root cause was that hookCommandNeedsPowerShellCallOperator returned false unconditionally — the &/no-& decision was keyed on (platform, runtime) only, with no signal for the effective hook-execution shell. One runtime (Claude Code) hosts either Git Bash or PowerShell on Windows. Fix: thread a new hookShell parameter through the projection chain. When hookShell='powershell', the & call operator is prepended. Default (Git Bash, no prefix) is unchanged and regression-locked. * docs: backfill changeset PR number (#2261) * merge: keep up to date with next * fix: regenerate stale capability-registry after next merge |
||
|
|
e0f969af6a |
refactor(#2246): centralize cross-platform path-separator handling (toPosixPath / toNativePath / posixNormalize) (#2247)
Replace every open-coded separator translation across the installer/hooks
source with named, tested seams in shell-command-projection.cts (the platform
seam), removing all hardcoded `/`+`\` from path handling:
- toPosixPath(p) — this machine's native path → POSIX (running-OS relative;
for local filesystem paths).
- toNativePath(p) — POSIX → native (collapses the win32 `/\//g,'\\'` ternary).
- posixNormalize(p)— unconditional `\`→`/`, OS-independent; for emitting paths
to a POSIX/bash TARGET (which may differ from the running
OS) and for parsing mixed-separator input.
core-utils.toPosixPath now delegates to the seam, so its 20+ existing consumers
resolve to one implementation; no duplicate helper.
- ~47 sites across runtime-hooks-surface, runtime-artifact-conversion,
runtime-artifact-install-plan, drift, init, worktree-safety,
installer-migrations, installer-migration-authoring, install-engine, surface,
verify, runtime-artifact-layout, schema-detect, check-command-router.
- Closes the latent POSIX-literal-backslash corruption class (the regex form
corrupts a POSIX path containing a literal backslash; split(path.sep) does not).
- New unit + fast-check property tests for all three helpers.
Closes #2246
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
79d7657eff |
feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.
Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
slash-programmatic / active-model / native-extension / bun) + hostBehaviors
{nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
(global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
markdown); the 16 other fixtures + claude-local change only by the shared
model-catalog hash line.
Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
(every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
consumes params; getArgumentCompletions; before_provider_request active-model
steering (fail-open on null resolution); functional session_start /
before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
vocabulary.
Docs (host-integration matrix + how-to) + changeset (Added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
5695522d5f |
feat(#2096): migrate Antigravity onto EoS declarative adapter + permission-writer + MCP companion (ADR-1239)
Fold all antigravity literal branches into descriptor-driven reads: getConfigDirFromHome (→ configHome.kind 'dot-home-nested'), projectLocalHookPrefix (→ hostBehaviors.hookPathStyle 'raw'), applyAgentPathRewrites (→ noPathRewrite), getProjectInstructionFile (→ projectInstructionFile 'GEMINI.md'); removed the dead inline convertClaudeAgentToAntigravityAgent branch + dead isAntigravity destructures (antigravity is already on the descriptor-agents path). subagentToolkit flipped undocumented→full (Context7: antigravity.google/docs/cli/features); namedDispatch/nested/maxDepth/backgroundDispatch stay undocumented. Byte-identical golden parity for all 16 runtimes. UPGRADE 1 (permission-writer): permissionWriter 'antigravity' + configureAntigravityPermissions merges a scoped permissions.allow block (GSD's own tree + hooks) into Antigravity's settings.json — non-destructive, idempotent, symmetric uninstall. Added to VALID_PERMISSION_WRITERS + the FinishPermissionWriter union. UPGRADE 2 (MCP companion): configureAntigravityMcpConfig writes mcp_config.json registering the gsd-core companion MCP server (Gemini-successor mcpServers schema, best-effort — raw schema unpublished). Both writers dispatch from finishInstall. settings.json is golden-excluded (HOOK_CONFIG_FILES); mcp_config.json (portable, no absolute paths) is golden-tracked → only antigravity.json changes. Tests: declarative-reference-antigravity extended (source-grep guard across 4 modules, fail-closed for the 4 undocumented sub-axes, validator acceptance) + antigravity-upgrades (permission-writer + mcp_config live-install, idempotency, user-preservation). Matrix + ADR-1016 + capability-manifest + CONTEXT.md + connect-gsd-mcp-server docs updated; changeset (Changed). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8f2ebbe9bf |
feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap). --gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): backfill changeset PR number (#1996) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): drop Gemini CLI from issue templates (review nit) Removes the sunset Gemini CLI runtime from the two GitHub issue-template runtime lists that the removal PR missed, per @davesienkowski's review nit: - feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise request a feature for a runtime GSD no longer supports) - bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json retrieval-help line Leaves the post-removal templates fully consistent with the Antigravity redirect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
dcd1d7f973 |
fix(#1693): don't double-quote $CLAUDE_PROJECT_DIR-anchored hook paths on Windows (#1746)
* fix(#1693): don't double-quote $CLAUDE_PROJECT_DIR-anchored hook paths on Windows The installer's #2979 legacy-node rewrite ran every managed node hook path through JSON.stringify on Windows. For local installs the path already carries a "$CLAUDE_PROJECT_DIR"-anchored quoted prefix, so stringifying produced "\"$CLAUDE_PROJECT_DIR\"/...". Node then received an argument starting with a literal " , treated it as relative, and failed MODULE_NOT_FOUND — breaking every node managed hook at once (a self-locking PreToolUse-guard deadlock). projectLegacySettingsHookCommand now emits an already-anchored token verbatim and only JSON.stringify-quotes bare absolute paths (which may contain spaces). Scoped to win32 so non-Windows token-shape behavior is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1693): add changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
871621c3c8 |
feat(#1740): require-fs-op-fallback production AST rule + Windows transient-lock retry (Phase 6) (#1742)
* feat(#1740): require-fs-op-fallback production AST rule + Windows transient-lock retry (Phase 6) ADR-1703 Phase 6 of the cross-platform portability epic (#1702). Adds the second production-code portability AST rule + the ADR-mandated glob expansion to bin/install.js and scripts/build-hooks.js. - eslint-rules/require-fs-op-fallback.cjs: flags an unguarded fs.rename / fs.renameSync (the atomic-publish primitive named first in DEFECT.WINDOWS-FS-OPS.symptom) that is NOT inside a try/catch whose handler references a transient errno ('EPERM'/'EBUSY'/'EACCES' or a *RETRY_ERRNOS set) AND NOT behind a Windows platform guard. A catch that silently swallows or cleans-up-and-rethrows without an errno check does NOT satisfy the defect's 'never silently swallow' clause. copyFile/unlink are deliberately not flagged (they are the fallback primitives per the defect's own fix-forward). Scope narrowed to rename per Phase 5's precision discipline; documented on #1740. - src/shell-command-projection.cts: export retryRenameSync(from, to) — the drop-in bounded-retry helper over the existing atomicRenameWithRetry. - 27 bare fs.renameSync sites across 11 modules routed through retryRenameSync (capability-lifecycle/lock/source, installer-migrations, milestone, phase, planning-workspace, roadmap-upgrade, runtime-hooks-surface, state, workstream). Idempotent on POSIX; resilient to AV/indexer transient locks on Windows. - eslint.config.mjs: register rule at error on src/**/*.cts; new focused portability-rules block covering bin/install.js + scripts/build-hooks.js (ADR-1703 L124-126 glob expansion — both files are compliant: zero rename violations). - tests: 15-case RuleTester suite; portability-rule-disable-ban extended (PROTECTED_RULES + scans bin/install.js/build-hooks.js with shebang handling); ci-test-scope portability-lint selection rule. - CONTEXT.md DEFECT.WINDOWS-FS-OPS predicate rewritten to point at the rule; docs/contributing/cross-platform-portability-rules.md reference + how-to. Closes #1740 * chore(#1740): backfill changeset pr:1742 * fix(#1740): tighten require-fs-op-fallback precision (codex review HIGH-1/HIGH-2) Addresses two false-negative findings from the codex (gpt-5.5/high) adversarial review of PR #1742: HIGH-1 — a catch that REFERENCES a transient errno but only rethrows (no retry/fallback) was marked compliant. The DEFECT.WINDOWS-FS-OPS fix-forward requires retry, not just recognition. Fix: catchHandlerHasRetrySignal now requires a loop `continue` backedge OR a `return <call>` delegation; a bare rethrow is flagged. The misleading `/* retry logic */` valid test is replaced with a real retry loop, and the rethrow-only shape is added as invalid. HIGH-2 — the nested-try ancestor walk treated an OUTER errno-catch as protecting the rename even when an INNER catch intercepted/swallowed the error (the outer catch is unreachable). Fix: isInsideTransientErrnoTryCatch now stops at the NEAREST enclosing TryStatement WITH A CATCH HANDLER whose block contains the rename (try-finally is skipped — it doesn't catch); outer catches are no longer consulted. The unsound nested-try valid test is converted to invalid, and a try-finally-skipped valid case is added. Verified: 17 RuleTester cases pass; zero new production violations (the 27 fixed sites use retryRenameSync; the real retry loops — atomicRenameWithRetry, capability-ledger/consent, build-hooks — remain compliant via continue/errno); lint:ci green; disable-ban + vocab-drift green. --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
cbf7c82841 |
feat(#323): fish-shell support in post-install PATH suggestion (#727)
* feat(#323): fish-shell support in post-install PATH suggestion Two additive changes to the post-install PATH-suggestion seam, both scoped to existing functions. A. Projection: add a fish entry to the persist-mode shell-action list in projectPathActionProjection() (src/shell-command-projection.cts). fish has no `export`/`$PATH`-list syntax, so the existing zsh/bash `export PATH=...` commands are inert when pasted. The new entry emits the fish-native `fish_add_path '<dir>'` (fish 3.2+, persists via the universal-variable store, de-duplicating). The directory is single-quoted with the same POSIX literal escaping as the zsh/bash siblings; verified round-tripping through real fish 3.7.0 for paths containing quotes, spaces, `$`, `*`, backticks and unicode. B. Detection: add homePathCoveredByFishConfig() in bin/install.js, called from maybeSuggestPathExport() alongside homePathCoveredByRc(). fish does not use sh-style `export PATH=` rc files, so a fish user whose fish_user_paths already covers the global bin would otherwise get a false-positive "not on your PATH" warning on every install. Two side-effect-free detection routes (no fish subprocess): 1. The universal-variable store (~/.config/fish/fish_variables). fish serializes this with `full_escape`: every byte outside [A-Za-z0-9/_] becomes `\xHH` (space -> \x20, `-` -> \x2d, `.` -> \x2e, `$` -> \x24, unicode -> \uXXXX) and list elements are joined by the literal 4-char token `\x1e` (NOT a raw 0x1e byte). The detector splits on `\x1e`, decodes the escapes, then compares each as an absolute literal — a decoded `$` is part of the directory name, not an unexpanded variable. Verified against real fish 3.7.0 output. 2. config.fish (`fish_add_path`, `set -gx PATH`, `set -Ux fish_user_paths`) — plain shell tokens: HOME forms ($HOME/${HOME}/~) are expanded and a token still holding `$` (e.g. `$PATH`, `$fish_user_paths`) is skipped. Honours $XDG_CONFIG_HOME and always also checks ~/.config/fish. No behaviour change for bash/zsh/PowerShell/cmd/Git-Bash users: their entries and command strings are unchanged; the fish entry is additive and the fish detector only narrows the set of cases that warn. Tests: update the projection length assertion (2 -> 3) and fish escaping in bug-3441; add fish detection + suppression cases in install-path-detection (uvar store with real fish escaping, dot/hyphen/space/$-literal decode regressions, config.fish routes, commented-out, relative-segment guard, unreadable-file fault injection, suppression and emission via maybeSuggestPathExport). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(changeset): add Changed fragment for #323 fish PATH support (#727) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#323): address review — action-only fish docs, decoder property test, win32 guard Addresses @trek-e's review on #727: - docs (blocker): keep the how-to action-only (Diátaxis). Drop the `# fish — persists via …` comment and the internal-mechanism clause naming fish_variables/config.fish; leave one command + the exec-fish directive. - tests (minor): extract decodeFishUniversalValue to a pure, exported module function and add fast-check round-trip properties (decode(fishEscape(p)) === p over arbitrary unicode, abs-path variant, totality). Consolidated into install-path-detection.test.cjs to respect the install test-file-count ratchet. - tests (follow-up): port #721's win32 negative-projection test (no fish action on win32; persist projection is PowerShell/cmd.exe/Git Bash). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#323): address review — drop unused 'after' import, clarify escaping comment Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
3857912ff6 |
fix(#1540): platformWriteSync retries transient rename locks instead of truncating readers (#1541)
* fix(core): platformWriteSync must retry transient rename locks, not truncate readers (#1540) platformWriteSync fell back to a non-atomic fs.writeFileSync(filePath) on ANY error from the temp+rename path. On Windows, renameSync onto a target a reader holds open throws EPERM/EBUSY/EACCES (the common case for hot files like STATE.md), so the fallback fired and a concurrent reader saw the file mid-truncation. Mirror the capability-ledger rename-retry idiom: retry transient lock errnos (EPERM/EBUSY/EACCES) with a bounded Atomics.wait backoff; on a persistent lock, surface the error rather than do the truncating non-atomic write. Genuinely unrenameable cases (EXDEV cross-device) and tmp-write failures still fall back. Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz * chore(changeset): Fixed fragment for #1541 (platformWriteSync rename retry) Claude-Session: https://claude.ai/code/session_01R88n7Q54bAaVHFkDbbH1yz --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
898d55788e |
fix(#976): detect command+args (wrapped) hook registrations in installer presence checks (#994)
* fix(#976): detect command+args (wrapped) hook registrations in installer presence checks Add referencesHook() helper that inspects both h.command (standard form) and h.args[] (args-form / wrapped-launcher form) when checking whether a managed hook is already registered. Rewrite all has*Hook predicates and the alreadyHas* guards to use it so args-form registrations suppress the duplicate stock string-command entry that was previously appended on every install/update. Also add an explicit args-form skip to rewriteLegacyManagedNodeHookCommands so entries with a non-empty args[] are left untouched (they are intentional user wrappers, not legacy bare-node commands to migrate). Extend isManagedHookCommand() in shell-command-projection.cts with an optional args: unknown[] parameter that checks whether any arg's basename matches the managed hook surface set — backward compatible; existing callers are unaffected. Regression test added to tests/install-regressions.test.cjs: - two-pass install with an args-form SessionStart entry pre-written to settings.local.json asserts exactly 1 hook entry remains after reinstall (previously 2 — the original args-form + a new stock string-command duplicate) - rewriteLegacyManagedNodeHookCommands test asserts args-form entries unchanged Closes #976 * chore(#976): backfill changeset pr number (994) --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> |
||
|
|
1b6bd66f2c |
feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged) (#821)
* feat(#770): register Claude Code lifecycle hooks (SubagentStop/Stop/PreCompact/FileChanged) Wire three new context-tracking events (SubagentStop, Stop, PreCompact) to gsd-context-monitor so context-headroom warnings surface at model-stop and subagent-finalisation moments — not just on PostToolUse. Add a new FileChanged hook (gsd-config-reload.js) that hot-reloads .planning/config.json context mid-session when the user edits it, injecting a config summary as hookSpecificOutput.additionalContext. Updates plugin manifest hooks.json, managed-hooks-registry, installer-migration-report allowlist, and shell-command-projection cleanup tables. Tests: 21 new assertions in enh-770-claude-hook-events.test.cjs; enh-788 and issue-766 test suites updated. Closes #770 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#770): document newly-registered Claude Code lifecycle hooks Add a Hook coverage table to the Claude Code npm installer section of docs/how-to/install-on-your-runtime.md describing SubagentStop, Stop, PreCompact, and the new FileChanged (gsd-config-reload.js) hook that hot-reloads .planning/config.json mid-session. Also fixes the changeset frontmatter (adds type: Added + pr: 821) so docs-lint can consume the fragment. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): add gsd-config-reload.js to INVENTORY.md and regenerate manifest The feat commit added hooks/gsd-config-reload.js but did not bump the Hooks count in docs/INVENTORY.md (14→15) or add the new row, and did not regenerate docs/INVENTORY-MANIFEST.json. Both inventory-counts and inventory-manifest-sync tests failed across the full CI matrix. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): make lifecycle-hook tests deterministic on scoped runner Replace the shared hooks/dist/ ensemble setup (ensureHooksDist / teardownHooksDist) in the Claude hook tests with per-test isolation: pre-populate each test's own tmpDir/.claude/hooks/ with stub files and pass installerMigrations:[] to install() so the first-time-baseline migration does not remove the stubs before the copy step can run. Root cause: hooks/dist/ is gitignored and absent on a fresh npm ci. ensureHooksDist() created it and teardownHooksDist() deleted it, but with --test-concurrency=4 both test files ran concurrently as separate Node.js worker processes sharing the same filesystem. One file's afterEach teardown deleted hooks/dist/ while the other file's install() was copying from it, producing an ENOENT (reproduced 2/10 runs locally). The additional issue: even with pre-placed stubs surviving the copy race, the 000-first-time-baseline migration classified hooks/gsd-*.js as bundled-gsd-hook artifacts, auto-removed them, and the copy step never re-ran (hooks/dist/ absent) — leaving contextMonitorFile missing and all hook registrations silently skipped (the 'got: []' symptom). Fix: pre-populate targetDir/hooks/ per-test (isolated temp dir) AND pass installerMigrations:[] so the baseline scan is skipped. The Qwen suites already used this pattern correctly; the Claude suites are aligned to it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): ship gsd-config-reload.js by adding it to build-hooks HOOKS_TO_COPY The #770 feature added hooks/gsd-config-reload.js and registered it in MANAGED_HOOKS, the installer, INVENTORY, and the test EXPECTED_ALL_HOOKS list — but never added it to scripts/build-hooks.js HOOKS_TO_COPY. As a result the hook was never copied into hooks/dist/ during the build, so: - the hook would never ship to users (real production bug — the FileChanged config-reload feature was dead-on-arrival), and - install-minimal-hooks.test.cjs #1755 ("all expected hooks are copied from hooks/dist/ to target", ".js hooks are executable after copy", "manifest contains .js hook entries") failed on any environment with a clean checkout (no pre-existing hooks/dist/): coverage, full test macos-22/macos-24, test ubuntu-24. The failures were masked locally only by a stale hooks/dist/ left from a prior build (build-hooks copies into dist without clearing it). On CI's fresh `npm ci` there is no dist, so the omission surfaced. Fix: add 'gsd-config-reload.js' to HOOKS_TO_COPY so build-hooks stages it into hooks/dist/ alongside the other JS hooks. Verified by removing hooks/dist/ and rerunning the full suite green (0 fail). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#770): make config prototype-pollution beforeEach deterministic on scoped runner Root cause: the #663 and alert-#26 prototype-pollution describe blocks seeded .planning/config.json in beforeEach via a bare runGsdTools('config-ensure-section') whose result was discarded. That command runs in a spawned gsd-tools child; on the scoped CI lane (--test-concurrency=4, config.test.cjs scheduled alongside the heavy install/tarball suites that #770 pulled into the targeted set) the child can be transiently killed under resource pressure (non-zero exit, empty stderr — an OS-level kill, not an app error). The swallowed failure left config.json absent, so the first subtest's readConfig() threw ENOENT opening <tmp>/.planning/config.json. Only 1 of 4 subtests failed, confirming a per-invocation transient, not a deterministic miss; the full suite schedules files differently so config.test.cjs did not collide with those heavy neighbors → passed there. Fix: add ensureConfigReady(tmpDir) which retries config-ensure-section on ANY failure or missing file and throws a clear diagnostic if it still cannot create config.json, then use it in both prototype-pollution beforeEach blocks. Setup is now deterministic under load; the #663/alert-#26 security assertions are unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
67703d1586 |
feat(#772): adopt stable Codex hook events + commandWindows for Windows parity (#827)
* feat(#772): adopt stable Codex hook events + commandWindows for Windows parity Register three new stable Codex hook events (SubagentStart, Stop, PostToolUse) wired to gsd-context-monitor.js so Codex installs get the same context-headroom tracking at subagent and session boundaries that Claude/Qwen already have. Add commandWindows field to the SessionStart hook entry on Windows so Codex uses the .cmd shim directly (Git Bash/MSYS cannot POSIX-exec node.exe). commandWindows is only emitted on win32; POSIX is unchanged. Refactor reconcileCodexHooksJsonSessionStart into a generic reconcileCodexHooksJsonEvent so any event name can be reconciled with the same dedup/preserve-user-entries logic. Add gsd-context-monitor.js and .cmd to MANAGED_HOOK_COMMAND_BASENAMES _BY_SURFACE so idempotent re-runs de-duplicate entries correctly. 30 new tests covering: export surface, event registration for each of the three events, commandWindows parity (POSIX vs win32), idempotency, uninstall, and user-entry preservation. Closes #772 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#772): windows path normalization + docs-lint - Normalize scriptPath backslashes to forward slashes in ensureCodexHooksJsonEvent and ensureCodexHooksJsonSessionStart so that isManagedHookCommand can match stored commands against configDir on Windows CI runners. path.resolve returns backslash paths on Windows, but when platform is not 'win32' (e.g. platform:'linux' in tests), projectManagedHookCommand skips normalization — producing a mismatch that breaks idempotency deduplication (the same hook entry appended twice on re-register). Forward-slash paths are always valid in both Node.js and Codex, so the normalization is safe for all platforms. - Fix changeset pr: 0 → 827 to resolve fail_malformed_fragment. - Add Codex hook coverage table to docs/how-to/install-on-your-runtime.md documenting the SubagentStart/Stop/PostToolUse events + commandWindows Windows-parity field added by this enhancement. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
10f6981743 |
fix(#685): set windowsHide on all Windows child-process spawns (#688)
* fix(#685): set windowsHide on all Windows child-process spawns A visible "gsd-core" console window flashed on Windows whenever a gsd-core child process spawned without `windowsHide: true`. The most visible offenders fire on every SessionStart / `/clear` (execNpm's `shell:true` npm view via the update-check worker) and on every Edit/Write/MultiEdit in a worktree (the worktree-path guard's git probe). Add `windowsHide: true` to every external-binary spawn in the runtime source: - hooks/gsd-context-monitor.js (record-session spawn) - hooks/gsd-worktree-path-guard.js (SPAWNOPT) - hooks/gsd-workflow-guard.js (git branch --show-current) - src/shell-command-projection.cts (execGit / execNpm / execTool) - src/check-command-router.cts (git log execFileSync) - src/roadmap-upgrade.cts (git status/rev-parse/reset/clean execSync) gsd-check-update.js already had it (the precedent). probeTty's tty call is POSIX-only and intentionally untouched. Adds a regression test that asserts each site plus a repo-wide completeness guard so a future external-binary spawn that omits windowsHide fails CI. No behavior change off-Windows. Closes #685 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#685): set changeset pr number to 688 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
df04aae5e4 |
enhancement(#537): migrate all hand-written bin/lib/*.cjs to TypeScript source of truth (ADR-457) (#602)
* enhancement(#537): migrate code-review-flags to TS source of truth Collapse the hand-written get-shit-done/bin/lib/code-review-flags.cjs to a TypeScript source of truth (src/code-review-flags.cts), compiled by tsc to a gitignored .cjs build artifact at the same path, per ADR-457 (build-at-publish). Second module after the semver-compare pilot (#541). Behaviour is preserved byte-for-behaviour (characterization test added in tests/code-review-flags.test.cjs locks the parser quirks). Adds compile-time type checking: CodeReviewFlags interface + CodeReviewWorkflow literal union. The require() path is unchanged, so code-review.md and the bug-3727 test keep working. The emitted .cjs is gitignored and eslint-ignored, mirroring the pilot. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 9 leaf bin/lib modules to TS source of truth ADR-457 build-at-publish, batch 1 (pure leaf modules, 0 sibling-deps): 001-legacy-orphan-files, context-utilization, redaction, artifacts, command-arg-projection, clock, ui-safety-gate, review-reviewer-selection, clusters. Each moves to src/*.cts (strict TS, typed), compiled by tsc to a gitignored .cjs at the same require() path; behaviour preserved byte-for- behaviour. Adds src/node-globals.d.ts (minimal ambient shim; "types":[]). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#537): add @types/node, drop hand-rolled node-globals shim ADR-457 migration infra: replace the temporary src/node-globals.d.ts ambient shim with @types/node@22 + "types":["node"] in tsconfig.build.json. Unblocks migrating the ~49 remaining bin/lib modules that use node:fs/path/os/ child_process. Build + full suite (3030 pass) + lint all green; no .cts type changes were needed (real Node types matched the shim). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 9 more bin/lib modules to TS (batch 2) ADR-457 build-at-publish. Clean leaves: installer-migration-report, prompt-budget. Type-error-prone leaves (were tsconfig.lint-excluded; now strict-typed and removed from that exclude list): secrets, phase-lifecycle, workstream-name-policy, decisions, validate, schema-detect. Plus runtime-name-policy. Strict type fixes narrow unknown->concrete domain types (no any/ts-ignore); behaviour preserved. Full suite green, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate runtime-slash to TS (cross-import proof) ADR-457. First cross-module TS->TS import: src/runtime-slash.cts imports ./runtime-name-policy.cjs and tsc resolves the sibling .cts types under strict (no declaration files; NodeNext .cjs->.cts mapping), emitting a correct require("./runtime-name-policy.cjs"). Confirms the recipe for coupled modules, which must be migrated in dependency order (leaves-up). Suite green, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 10 more bin/lib modules to TS (batch 3) ADR-457 build-at-publish, Wave-1 leaves: event, workstream-inventory-builder, plan-scan, fallow-runner, project-root, installer-migration-authoring, update-context, 000-first-time-baseline, runtime-homes, model-catalog. Strict typing fixed real issues (narrowing unknown, qualified fs/path calls, removed unnecessary casts); plan-scan/project-root/workstream-inventory-builder dropped from tsconfig.lint exclude. Behaviour preserved; suite green, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 5 large Wave-1 leaves to TS (batch 4) ADR-457 build-at-publish: configuration, state-document, shell-command- projection (42 dependents), security, command-aliases. shell-command- projection keeps a namespace child_process import for mock-intercept testability. loadConfig/migrateOnDisk emit synchronously (every caller uses them sync; the one awaited migrateOnDisk caller tolerates a non-Promise) — full suite (3030 pass) confirms behaviour preserved. configuration/ state-document/command-aliases dropped from tsconfig.lint exclude. Also fixes the malformed batch-3 changeset frontmatter (type/pr) that failed lint:docs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 6 Wave-2 modules to TS (batch 5) ADR-457 build-at-publish: config-schema, model-profiles, 002-codex-legacy-hooks-json, logger, active-workstream-store, adr-parser. First batch importing already-migrated siblings (configuration, model-catalog, shell-command-projection, redaction, security) via ./sibling.cjs specifiers. Strict type narrowing (typeof guards over String(unknown)); behaviour preserved; suite 3030 pass, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 5 large Wave-2 modules to TS (batch 6) ADR-457 build-at-publish: graphify, install-profiles, intel, installer-migrations, worktree-safety. installer-migrations preserves its dynamic require() loader for numbered migration modules (scoped lint suppressions). Strict typing (typeof guards over String(unknown)); behaviour preserved; suite 3030 pass, lint 0 errors. Wave 2 complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate Wave-3 modules to TS (batch 7) ADR-457 build-at-publish: planning-workspace, runtime-artifact-layout, command-routing-hub, drift. Uses `import x = require()` for export= siblings; drift's lazy require of runtime-slash hoisted to a top-level import (verified non-circular). Behaviour preserved; suite 3030 pass, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate small Wave-4 modules to TS (batch 8) ADR-457 build-at-publish: cjs-command-router-adapter, phase-command-router, surface, roadmap-upgrade. Typed the hub router handler results as the HubResult discriminated union; surface drops 4 genuinely-unused imports. Behaviour preserved; suite 3030 pass, lint 0 errors. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate core hub (2.5k LOC, 68 dependents) to TS (batch 9) ADR-457 build-at-publish: get-shit-done/bin/lib/core.cjs -> src/core.cts, preserving all 63 exports via export=. All sibling deps already migrated (shell-command-projection, model-profiles, model-catalog, worktree-safety, planning-workspace, project-root, configuration, config-schema). Strict types, no any/ts-ignore; config-schema lazy require hoisted (non-circular). Behaviour preserved (independently verified: core's shard 3030 pass / 0 fail). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537): make ESLint-coverage + test-sprawl checks migration-aware #551 test hardcoded 12 now-migrated modules as "hand-written, must be linted"; that invariant is obsoleted by the ADR-457 migration. Rewrite it to a filesystem-driven invariant that holds at every stage: a bin/lib/*.cjs must be eslint-ignored IFF it has a src/*.cts source (tsc-generated), else linted (covers package-identity, which has no TS source). Also eslint-ignore config-types.cjs (has a src counterpart) and drop the redundant tests/clock.test.cjs (clock already covered by clock-seam + bug-474 tests), which tripped the lint-test-file-count ratchet. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 9 Wave-5 router/inventory modules to TS (batch 10) ADR-457 build-at-publish: phases/verify/init/agent/task/validate/roadmap/state command routers + workstream-inventory. Router handler results typed against core's exported shapes; behaviour preserved (caught+fixed a --verify boolean flag regression mid-migration). Full suite green across all shards (only the 4 local gpg-env changeset-notes failures remain; CI passes them). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 7 Wave-5 modules to TS (batch 11) ADR-457 build-at-publish: gap-checker, docs, check-command-router, frontmatter, learnings, gsd2-import, profile-pipeline. Behaviour preserved; full suite green across all shards (only the 4 local gpg-env failures remain). Also broadens atomic-write-coverage.test.cjs to accept the tsc-compiled namespace-import form while still asserting platformWriteSync is called (safety guard intact). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate config + profile-output to TS (batch 12) ADR-457 build-at-publish: config (729 LOC), profile-output (1142 LOC). All exports preserved; cmdMigrateConfig de-asynced (migrateOnDisk is sync, awaited caller tolerates it). Behaviour preserved; suite green across all shards (only the 4 local gpg-env failures). Wave 5 complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate 5 Wave-6 modules to TS (batch 13) ADR-457 build-at-publish: template, uat, workstream, roadmap, audit. Behaviour preserved (dead toPosixPath import dropped from audit; inline requires hoisted). Suite green across all shards (only the 4 local gpg-env failures). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate commands + state hubs to TS (batch 14) ADR-457 build-at-publish: commands (1305 LOC), state (2074 LOC, 17 dependents). All exports preserved; inner requires kept non-hoisted where load-order matters (install.js, per-call security); acquireStateLock cast inlined to preserve the err.code source token a structural test inspects. Behaviour preserved; suite green across all shards (only the 4 local gpg-env failures). Wave 6 complete. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate milestone to TS (batch 15a, hand-authored) ADR-457 build-at-publish: milestone -> src/milestone.cts. Authored directly (subagent capacity was unavailable). Also relaxes core.output()'s 3rd param to optional, matching its real always-optional call contract (unblocks remaining 2-arg output callers). Behaviour preserved; suite green across all shards (only the 4 local gpg-env failures). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537): migrate phase, verify, init to TS (batch 15, final modules) ADR-457 build-at-publish, Wave 7 (the last hubs): phase (1608 LOC), verify (1615), init (2113). Adds src/package-identity.d.cts so verify can import the permanently value-baked package-identity.cjs under strict TS. Fixes two regressions the migration introduced in verify: restore cmdValidateHealth's `return result` (callers/tests read result.warnings — it is NOT side-effect-only), and make the bug-3384 source-pattern test tolerant of the tsc-compiled bracket-notation form of the git_list_failed->W020 branch (behaviour intact). Full suite green across all shards (only the 4 local gpg-env failures); lint 0 errors. All 86 migratable bin/lib modules are now TypeScript sources. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#537): finalize ADR-457 migration — retire tsconfig.lint.json All hand-written bin/lib/*.cjs are now src/*.cts sources, so the checkJs stopgap tsconfig.lint.json (unused; not wired into eslint, scripts, or CI) is deleted per ADR-457's final step. Also gitignore the tsc-generated config-types.cjs (was still committed) for consistency with every other emitted artifact. package-identity.cjs stays value-baked (declared via src/package-identity.d.cts). Suite green; #551 ESLint-coverage test green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): add prepare script so unpacked/git installs build bin/lib artifacts ADR-457 build-at-publish: bin/lib/*.cjs are now gitignored, built by tsc. The prepack/prepublishOnly hooks cover `npm pack`/publish, but `npm install -g <dir>` and git installs run the `prepare` lifecycle — which was missing — so the unpacked install shipped without the compiled .cjs and failed at startup with "Cannot find module './lib/core.cjs'" (caught by the smoke-unpacked CI job). Add `prepare` mirroring prepublishOnly (build:lib + build:hooks). prepare does NOT run for registry consumers (they get the pre-built tarball), only for source/local/pack installs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): make CI build/lockfile checks work with gitignored bin/lib artifacts ADR-457 build-at-publish exposed two CI assumptions that bin/lib/*.cjs are always present on disk: - check:env's lockfile-sync ran `npm ci --dry-run`, which now triggers the `prepare` build (tsc) — but it runs before deps are installed, so tsc is absent and it misreported the lockfile as out of sync. Add --ignore-scripts (a lockfile check must not build). - the lint-tests job installs with --ignore-scripts (no prepare build), but lint:skill-deps require()s the built install-profiles.cjs. Add an explicit `npm run build:lib` step after install. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): narrow prepare to build:lib only (unbreak packed-smoke pack step) prepare running build:hooks emitted "✓ Copying ..." stdout during `npm pack`, which the install-smoke "Pack root tarball" step captures into $GITHUB_OUTPUT — breaking it with "Invalid format". build:lib (tsc) is silent on success and is all the unpacked/source install needs (the smoke-unpacked assertions exercise gsd-tools, i.e. bin/lib, and tolerate hook setup with `|| true`). Matches prepack. build:hooks still runs on prepublishOnly for real publishes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#537): wire Stryker mutation gate to build-at-publish layout The gate scored 0.00 because it mutated changed bin/lib/*.cjs that (a) were generated artifacts and (b) included modules with no coverage in the command's test set. Rework: mutation.yml now derives changed COVERED modules from src/*.cts and maps them to their built bin/lib/*.cjs; Stryker mutates those built artifacts with a no-rebuild command (mutating src/*.cts + per-mutant tsc was ~3x over the 30-min CI budget). NOTE: with the gate now correctly measuring the covered modules, their actual mutation score is 42.94% (< break 50) — a pre-existing test-coverage gap (adr-parser/prompt-budget/etc.), not introduced by this behaviour-preserving migration. Reaching 50 needs more tests, a threshold/scope change, or a waiver — a maintainer decision. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537): raise mutation coverage of covered modules above the 50 gate Adds focused example-based unit tests that kill surviving mutants in the two lowest-scoring covered modules: - tests/prompt-budget.unit.test.cjs (112 tests): 17.9% -> 97.9% - tests/adr-parser.unit.test.cjs (205 tests): 44.7% -> 89.4% Both wired into stryker.config.mjs's command. Fresh full run over the 6 covered modules now scores 82.25% (>= break 50); every covered module is >= 68%. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * enhancement(#537,#609): parallelize mutation gate via dynamic per-module matrix The serial Stryker run timed out at 30 min once the migration's added tests made every mutant re-run ~300 tests. Replace it with a dynamic matrix so the gate completes well under budget — folded into this PR (was tracked as #609) because it's a prerequisite for this PR's mutation gate to pass. - scripts/mutation-matrix.cjs: single source of truth (covered-module -> test files) computing changed covered modules from git diff -> {has_work, matrix}. - mutation.yml: detect -> dynamic `matrix: fromJSON(...)` mutate job (one parallel shard per changed module, scoped via MUTATION_TEST_CMD to only that module's tests, 15-min/shard) -> summary job that KEEPS the legacy check name "Stryker mutation score (changed files only)" so branch protection is unchanged. Per-shard jobs report as "Stryker (<module>)". - stryker.config.mjs: commandRunner.command reads MUTATION_TEST_CMD (falls back to the full command locally). Closes #609. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537,#609): give each mutation shard ≥50% on its own tests; drop blacksmith note Per-module sharding revealed that active-workstream-store (46.5%) and frontmatter (7.4%) only cleared 50% in the old serial run via timeout-noise from the bloated 300-test command; on their own tests they were below the gate. Add focused unit tests: - tests/active-workstream-store.unit.test.cjs (115 tests): 46.5% -> 81.9% - tests/frontmatter.unit.test.cjs (165 tests): 7.4% -> 63.4% Both wired into scripts/mutation-matrix.cjs (per-module test map) and stryker.config.mjs DEFAULT_TEST_CMD. All 6 covered modules now clear break:50 with only their own tests (config-schema/context-utilization/prompt-budget/ adr-parser already did). Also removes the leftover blacksmith TODO comment — GitHub-hosted runners only; speed comes from parallel per-module shards. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#537,#609): strengthen prompt-budget tests to clear the gate on its own tests prompt-budget scored 39.58% when mutation-tested with ONLY its own tests (the way the per-module CI shard runs it) — an earlier ~98% reading was inflated by accidentally running the full multi-module command. Add 96 targeted tests to tests/prompt-budget.unit.test.cjs (exact note-template text, plan-truncation arithmetic/percentages, drop-block strings, noteInjected/hardFailed booleans): scoped score 39.58% -> 68.75% (>= break 50). All 6 covered modules now clear the gate on their own tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |