* test(260903-m7p): expose configured-entrypoint validation gap * enhance(260903-m7p): validate configured entrypoints before success * test(260903-m7p): require pre-success entrypoint validation * enhance(260903-m7p): gate install success on entrypoints * test(260903-m7p): cover configured entrypoints across runtimes * enhance(260903-m7p): cover emitted runtime entrypoints * fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number - finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process before the new configured-entrypoint assertion throws. Without a HOME + config-location-env sandbox that write resolved through the ambient environment and landed in the developer's live ~/.gsd (confirmed absent on origin/next baseline, present only on this branch — full-suite HERMETICITY WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the duration of the test, matching the existing in-process finishInstall/ install() pattern in tests/install.test.cjs (#2665). - .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the fork PR number until the upstream PR number is known. * fix(260903-m7p): repair cross-platform and pre-existing shape fallout - tests/configured-entrypoint-validation.test.cjs: the win32 branch of ensureCodexHooksJsonSessionStart writes a .cmd shim under <codexRoot>/hooks/; create that dir in the test (the real installer only calls this once hooks/gsd-check-update.js already exists) and assert the platform-common entrypoint shape instead of a fixed non-Windows array, since win32 legitimately emits two entries (cmd shim + script). - tests/install.test.cjs: finishInstall's shared settings-json return now carries configuredEntrypoints/rollbackInstallerMigrations for every runtime on that path (trae included, not just Claude/Cursor/Windsurf); update the trae install() exact-shape assertion to match. * fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash configuredEntrypointsForHook's shell branch dropped interpreterCandidates entirely when resolveBashExecutable returned null, unlike the sibling portableHooks runner entry a few lines below (which correctly falls back to the literal 'bash' token). Found via agy adversarial review; verified unreachable through the current call graph (buildHookCommand's own resolveBashRunner==null gate already short-circuits before recordConfiguredHookCommand runs), so this is a defensive consistency fix, not a live-bug patch — kept for the next caller that does not share that gate. * chore(260903-m7p): backfill changeset pr number to the opened upstream PR .changeset/quick-wasps-sing.md carried the fork PR number (16) as a placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now open, so record its real number per CONTRIBUTING.md's changeset pr-field convention. * fix(#4154): track already-registered hooks for entrypoint validation on update applySettingsJsonHooks registers each guard hook only if absent, so a hook already present from a prior install keeps its stale on-disk command. The new entrypoint tracker always records the freshly-computed command for it, which never matches what is actually persisted, so the exact-string filter in finishInstall silently dropped it from validation — the Blocker case this feature exists to catch (an already-installed entrypoint going stale between installs) was exactly the case it never validated. Match on the managed script's basename instead, which the persisted command carries either way, so an already-registered hook stays in the validated set. Regression test forces this path by mutating a freshly-installed hook's persisted command before a second install. * fix(#4154): distinguish an unreadable script from a missing one validateConfiguredEntrypoints folded an EACCES statSync failure into the same 'missing' reason as ENOENT, misreporting a real permission problem as an absent file. Check the error code and report 'unreadable' instead. * docs(#4154): document entrypoint validation's rollback and PATH scope CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite bin/install.js x CONTEXT.md being this repo's strongest co-change pairing. The update-gsd.md how-to overstated what a validation failure undoes: for Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/ config.toml inside install() before the aggregate validation call runs, so there is no rollback path for that write regardless of "where available" phrasing. Also note that interpreter resolution checks the installer's own PATH, not necessarily the PATH a hook fires under later (#2979 launchers). * chore(#4154): point changeset pr field at the fork PR while CI runs there Mirrors the branch's own prior backfill commit: pr: matches whichever PR number changeset-lint is currently validating against (fork PR #16 during the fork-first CI/review loop), flipped back to the upstream PR number right before the final push to open-gsd/gsd-core. * fix(#4249): address adversarial-review findings in entrypoint validation An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR found several real gaps beyond the human reviewer's Blocker, verified against source before fixing: - Codex's install() result bound rollbackInstallerMigrations to the narrow installer-migrations-only rollback instead of restoreCodexSnapshot (#3245), the full pre-install snapshot/restore Codex already owns for exactly this case — a validation failure discovered outside install() reverted nothing of the config.toml/hooks.json that call had already written. - The register-only-if-absent basename match from the prior fix used a bare substring, which an unrelated user command mentioning the same filename could false-positive into GSD's validated set — anchored on the `/hooks/<basename>` path segment instead. - nodeCandidates checked raw process.execPath (always true — we're running in that process) instead of normalizeNodePath's stable version-manager alias, the same one buildNodeRunnerChainToken bakes as its first choice — a false green regardless of whether that alias itself still resolves. - An entry with no interpreterCandidates (Cline's PreToolUse hook, or a Windows-Claude .sh hook invoked without a bash runner) runs via its own shebang; validateConfiguredEntrypoints checked only file-type, never the execute bit. Cline's writer also never reported an entrypoint at all. - Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor hook registered across several events) were validated once per duplicate. Each fix is covered by a new or extended test; the Codex one required inlining runCodexInstall's env sandboxing so the rollback closure — which re-resolves the $HOME-relative skills root live — runs before the sandbox is torn down, matching how installAllRuntimes' real aggregate gate calls it. * docs(#4249): document the round-2 entrypoint-validation fixes Runtime Hooks Surface Module and Installer Module entries now name ConfiguredEntrypoint's not-executable reason, the normalizeNodePath alignment, Cline's tracked hook, and which install() result the finishInstall/installAllRuntimes rollback path actually reverts per runtime (Codex's full snapshot vs. the others' narrow migrations-only rollback). * chore(#4249): point changeset pr field at the upstream PR now that fork CI is green * fix(#4249): address agy adversarial-review findings - validateConfiguredEntrypoints: statSync alone never detects a chmod-000 script (it only needs parent-dir search permission), so an interpreter-invoked entry with an unreadable script passed validation. Add an explicit R_OK check for the interpreterCandidates branch only — the candidate-less/shebang branch already has its own X_OK gate. - docs/how-to/update-gsd.md: the blanket "does not revert" claim was false for Codex, which reverts config.toml/hooks.json via its full pre-install snapshot; qualify it per runtime. - tests/codex-config.test.cjs: the #4249 rollback regression test asserted skills/ and VERSION were reverted but never asserted config.toml/hooks.json were too, despite the test's own stated intent. - CONTEXT.md: qualify which interpreterCandidates entries get normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's Windows .cmd shim as a third candidate-less case that relies on extension dispatch, not a shebang. * fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node', so it needs the execute bit (like any shebang-invoked entry), but its interpreter is looked up on PATH by 'env' at hook-fire time (unlike every other GSD JS hook, which bakes an absolute node path specifically to avoid that dependency). The candidate-less/interpreterCandidates fork treated these as mutually exclusive, so Cline's entry silently skipped interpreter resolution entirely — a completely missing 'node' on PATH would still validate successfully. Add an orthogonal selfExecutable flag so both checks run for entries that need them. (CodeRabbit finding on the fork rehearsal PR.) * fix(#4249): address second-round adversarial review findings (opus + agy) - validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry, not just interpreterCandidates ones — a self-executable shebang script is still opened and read by its kernel-invoked interpreter, so X_OK alone never proved it was readable. - selfExecutable is now the sole, explicit source of truth for the execute-bit check (every producer that needs it sets the flag) instead of being partly inferred from an absent interpreterCandidates, which Cline's hybrid entry also carries. - The execute-bit check now skips explicitly on win32 (matching resolveExecutableBinary's own carve-out) instead of relying on Node's accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows machine and not a test that simulates win32 on a POSIX runner. - bin/install.js: fixed a stale comment claiming no runtime's install()-time writes have a rollback path — Codex's does (restoreCodexSnapshot) — and added the omitted Cline to both that comment and CONTEXT.md's equivalent lists. - CONTEXT.md: fixed the Cline description left stale by the previous commit's selfExecutable addition, and rewrote the validation-mechanism paragraph for clarity (writing-for-agents pass). - docs/how-to/update-gsd.md: split an overloaded 4-clause sentence. - Removed a fault-injection integration test that could not reliably exercise the real installAllRuntimes -> finalize -> rollback wiring without fighting the installer's own pre-registration existence guards; the constituent pieces remain covered individually. * fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner X_OK is a POSIX-only concept, skipped entirely when an entry's platform is win32 (matching production). Two test entries omitted platform, defaulting to process.platform — on an actual windows-latest CI runner that silently skipped the very check they were meant to exercise, turning 'not-executable' into a false pass. Pin platform: 'linux' so these are deterministic regardless of which OS runs the suite. * fix(#4249): classify EPERM the same as EACCES in statSync error handling Windows raises EPERM (not EACCES) for a parent directory that couldn't be traversed into — was falling through to 'missing', misreporting a genuine permission problem as a nonexistent path. * docs(#4249): address final CodeRabbit doc-completeness findings - CONTEXT.md: install()'s documented result shape omitted configuredEntrypoints; the ConfiguredEntrypoint shape omitted selfExecutable. - docs/how-to/update-gsd.md: the failure-mode sentence omitted unreadable and lacks-execute-permission, which the installer also rejects. * fix(#4249): stop double-validating every configured entrypoint on install/update installAllRuntimes' finalize() already runs assertConfiguredEntrypoints once over the aggregate set; finishInstall then re-ran the identical check per runtime in the printSummaries loop right after, so every entrypoint paid its statSync/accessSync/interpreter-resolution cost twice on every install and update. Add entrypointsAlreadyValidated to skip the redundant pass specifically on that path, while leaving the check intact for any caller that invokes finishInstall directly. * chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there * perf(#4249): memoize interpreter candidate resolution across entrypoints resolveExecutableBinary walked PATH once per (entry, candidate) pair; a typical install has a dozen-plus entries sharing the same few candidate lists (process.execPath for JS hooks, bash for shell hooks). Cache by (platform, candidate) so each distinct pair resolves once per validation call instead of once per entry. * chore(#4249): point changeset pr field at the rebased rehearsal fork PR * fix(#4249): drop entrypoint tracking from the now-dead Codex event writer #2586 (landed on next after this branch forked) removed install.js's CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent no longer runs during install or update. The ConfiguredEntrypoint records this branch added inside it were therefore unreachable and untested. Restore the function to its upstream shape; the entrypoints it used to report were never collected by any caller. * refactor(#4249): drop the revalidation bypass flag and the candidate cache Both were this PR's own micro-optimisations over a set of roughly a dozen entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall gate off to save one statSync/accessSync pass; `resolvedCandidateCache` memoised resolveExecutableBinary across entries that are already deduped by (configPath, scriptPath). Neither is measurable, and the flag was the only way to reach finishInstall with validation disabled. finishInstall now always validates what it is given. * chore(#4249): point the changeset pr field back at the upstream PR * refactor(#4249): track settings.json entrypoints without the hooksSurface gate The install-surface writer only tracked configured entrypoints when the runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing asserts that axis agrees with `installSurface`, so a descriptor that broke the coupling would silently pass `configuredEntrypoints: undefined` and drop that runtime out of the validation this PR adds — reintroducing the exact 'reports Done! over a broken entrypoint' failure #4154 exists to close. Remove the dependence rather than test it: everything recorded on this path lands in settings.json by construction, and the registered-command filter already discards entries no persisted hook references. * chore(#4249): put the changeset body in the documented two-part format CONTRIBUTING.md and .changeset/README.md both show `**<bold change>** — <symptom-led explanation>.`; the fragment was a single unbolded sentence. * chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there * fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback #3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-* and gsd-core/VERSION. The install overwrites every other GSD-owned file too — hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest itself — before the entrypoint-validation gate runs, so a validation failure left the new payload sitting on top of the restored old config. Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims, before runInstallerMigrations so the bytes are the true pre-install state, and restore it from both Codex rollback closures ahead of the per-surface restores. Files only the failed install introduced are removed, read from the manifest now on disk. The manifest is already the authoritative record of what GSD owns, so no second hand-written list can drift out of sync, and user-owned files are never snapshotted or removed. Every path is confined through resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into an arbitrary-path write. Non-Codex runtimes are unaffected: the snapshot is gated on the same tomlConfigInstall + non-minimal condition as #3245's. * fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest Two follow-on defects in the previous commit's snapshot: - The capture was gated on `!isMinimalMode`, copied from #3245. A core/ --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the snapshot came back empty while the rollback still ran — and its removal pass would have deleted every file the new manifest lists. Gate on tomlConfigInstall alone, matching where the rollback actually reaches. - An unreadable or unparseable prior manifest was caught alongside ENOENT and treated as a fresh install. That is the same empty-snapshot state, so a failed update over a real install with a corrupt manifest could delete its prior payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means known-empty; any other read error or a parse failure means unknown, and the restore closure returns without touching anything, degrading to #3245's narrower rollback. Deliberately not fatal — a corrupt manifest has to stay repairable by reinstalling over it. Both paths are covered by red-checked regression tests. * fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too commit removed from the manifest snapshot. restoreCodexSnapshot is reachable for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-* skill dir and gsd-* agent file the snapshot does not claim — so with an empty minimal-mode snapshot a rollback deleted the whole skills/agents surface with nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via the ADR-1239 skills-kind home override, so this is also the reason manifest `skills/` keys do not resolve under configDir: that surface belongs to this snapshot, not to the manifest-driven one. Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal mode — doing nothing on an early failure is the non-destructive side. Covered by a red-checked regression test that plants bytes in an alternate-home skill file, reinstalls under the core profile marker, and asserts the rollback restores it. * fix(#4249): never remove on rollback unless a prior manifest proves what predates the install Three defects in the manifest-driven Codex rollback, all in its removal half: - ENOENT marked the snapshot usable, arming the removal pass on a FIRST install. GSD may have overwritten a user's file at a manifest-tracked path there, and no prior manifest records the difference — so rollback deleted it where before it merely left it overwritten. Absent, unreadable and malformed manifests now all leave the prior set UNKNOWN and skip removal entirely. - Membership was tested against the map of files whose pre-install read SUCCEEDED, so a tracked file that existed but was unreadable read as introduced-by-this-install and was removed. Track the prior manifest's paths in their own Set and test against that. - The unreachable "delete the manifest when there was no prior one" branch is gone: usable now implies a parsed prior manifest. Also adds the end-to-end test the aggregate gate was missing — the four Codex rollback tests drove the closure directly, proving the restore but not the wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes Cline's `env node` entry fail validation for real, and asserts Codex's payload comes back. Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper instead of six copies. Both new tests are red-checked. * test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its Windows-EBUSY retry budget. That budget is for directory trees; this removes a single file, which unlinkSync says more precisely and the rule does not flag. * chore(#4249): point the changeset pr field back at the upstream PR * fix(#4249): use an unambiguous dedup key and surface partial-restore failures trek-e's 2026-09-08 adversarial pass flagged two findings in the new entrypoint-validation/rollback code: - assertConfiguredEntrypoints' dedup key already used a raw NUL separator (introduced in ceebb65f2d), but git/Read render NUL as a space, so the key looked like a plain-space join to every reviewer that read the diff. Replace it with JSON.stringify([configPath, scriptPath]) so the separator is visible and unambiguous. - restoreManagedFileSnapshot's per-file restore catch block claimed to 'surface the original error' but only swallowed it, matching (and widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add an actual console.warn using the existing best-effort-warning convention, scoped to just this PR's new function. * fix(#4249): treat a files-less prior manifest as unknown, not known-empty agy's gemini-3.8-flash-high adversarial pass (round 5) found and I reproduced empirically: a structurally-valid manifest missing the files key (e.g. {"version":1}) parses without throwing, so Object.keys(undefined || {}) silently read as 'zero files predate this install' instead of the UNKNOWN state the malformed-manifest guard exists to produce. Rollback's removal pass then deleted every GSD-owned file the failed install's own manifest listed, including ones that predated it — the exact data loss the #4249 CodeRabbit malformed-manifest fix was supposed to prevent, reachable through a JSON.parse success instead of a failure. Route the shapeless case into the same catch-all UNKNOWN path via an explicit shape check. Regression test reproduces the deletion before the fix and confirms the file survives after it. Also extend restoreManagedFileSnapshot's removal-pass rmSync and final manifest-rewrite catches with the same real console.warn trek-e's round-4 review asked for on the per-file restore catch — same rollback function, same operator-facing-signal gap. * docs(#4249): correct which runtimes actually leave a written config on rollback agy's completeness audit (round 5, holistic pass) caught this new paragraph claiming 'for every other runtime, the configuration file(s) already written during that update are left in place' — false for Claude Code and other settings.json-based runtimes, whose write never happens on failure (assertConfiguredEntrypoints runs before finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline actually match that description, since they persist their config file inside install() ahead of the gate. Split the one sentence into the three actual outcomes; matches the PR body's own accurate Before/After wording, which this doc addition had drifted from. * fix(#4249): clean up doc/comment mismatches and dead fields from opus review Opus critical-code-reviewer + ponytail-review pass on the final diff: - assertConfiguredEntrypoints carried finishInstall's old docblock ("Apply statusline config, then print completion message") from before this function was inserted between comment and callee. finishInstall already has its own accurate #4249 comment, so the stale docblock is removed rather than moved. - checked: number on ConfiguredEntrypointValidationResult and error.configuredEntrypointValidation on the thrown error: the first had zero consumers anywhere in the repo, including its own defining file, and is removed. The second matches an existing repo convention (bin/install.js's installerMigrationRollbackFailures, #4249 predates this PR) of attaching structured diagnostic context to a re-thrown Error even before a consumer exists, so it's kept. - finishInstall's own assertConfiguredEntrypoints call is a redundant backstop on the real production path (installAllRuntimes's aggregate call already validates the superset first), but its comment read as though this call alone provided the before-the-write guarantee. Clarified rather than removed — it's the only gate for a caller that invokes finishInstall directly. * chore(#4249): split the manifest-driven rollback engine out into #4544 Issue #4154 asked the installer to consume a validation failure "through the existing rollback mechanism, without a second transaction mechanism". The manifest-driven rollback widening added during review (capture every path the prior gsd-file-manifest.json claims, restore those bytes, remove what only the failed install introduced) is that second mechanism on a plain reading. It is a real fix for a #3245-era gap, but an independent one, so it moves to its own bug report and PR. Removed here: - bin/install.js: the pre-install managed-file capture block and restoreManagedFileSnapshot, plus its call sites in _codexPreConfigRollback and restoreCodexSnapshot (99 lines). - tests/configured-entrypoint-validation.test.cjs: the five tests that exercise the manifest engine. - CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the widened restore. update-gsd.md again documents the #3245 surfaces only. Kept, because it is #4154's own scope: - the entrypoint-validation gate itself; - Codex's install() result binding rollbackInstallerMigrations to restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*, agents/gsd-*, gsd-core/VERSION); - the !isMinimalMode gate removal on that snapshot. Binding the closure to the result made it reachable for a core/--minimal install, where its pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot does not claim; an empty minimal-mode snapshot therefore deleted the whole surface with nothing to restore. The surviving aggregate-failure test now asserts on config.toml, a surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which only the manifest engine restored. Refs #4544 * test(#4249): cover configured entrypoints through the packed install path #4154's scope lists install smoke coverage alongside the installer gate — "assert representative configured entrypoints resolve for supported runtime profiles". The gate itself (assertConfiguredEntrypoints / validateConfiguredEntrypoints) is unit-covered by in-process install() calls; nothing proved the property survives npm pack -> npm install -g -> install.js. Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct config surfaces GSD writes launch paths into (settings.json, and hooks.json + config.toml) — run the tarball-installed installer into a throwaway HOME, then re-read that runtime's own written config and return the new ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check becomes a release gate on every matrix host without workflow changes. The scan re-derives paths from the written config instead of reusing the installer's own entrypoint list, and test I shows why that matters: a registration the installer never touched during a run is invisible to the in-process gate, so the install exits 0 and only reading the config back off disk catches the dangling launch path. * ci(#4249): pack a publish-shaped tarball in the install smoke lane `npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs after `npm ci` carries no hook scripts at all — the lane has been smoking a package that differs from the published one in exactly the artifacts the lifecycle smoke is supposed to launch. That went unnoticed because the lane's init runs `--local`, which registers no statusline and therefore registers no hook whose target is missing. A `--global` install on the same tarball exits 1 on #4249's own gate (`gsd-statusline.js (missing)`), which is what the new configured-entrypoint cycle performs, so without this step the cycle would report INIT_FAILED instead of checking anything. Build hooks before packing so the smoked tarball matches prepublishOnly. The CLI now reports 16 configured entrypoints for claude and 1 for codex instead of zero. * fix(#4249): scope Codex's full snapshot restore to entrypoint failures Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage exception un-install a Codex install that had already succeeded and already printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints` call. Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration` scopes finalize-stage rollback to installer *migrations* ("the executor uses the journal to restore modified paths"), and this PR's own operator-facing paragraph in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/ agents/VERSION revert to entrypoint-validation failures specifically ("If a script is missing, unreadable, ... For Codex, this reverts ..."). The wide behaviour is also incoherent as a transaction abort: the same doc says Cursor, Windsurf, Kimi and Cline keep the config they wrote inside install(). Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install bytes — on an update, silently downgrading a working Codex install to the previous version while the user had just been told it was Done. Select the rollback by error kind instead. `assertConfiguredEntrypoints` already tags its error with `configuredEntrypointValidation`, so the full snapshot restore runs for that error (and anything downstream of it, including finishInstall's per-runtime backstop) and the installer-migrations-only closure runs for everything else. The codex result now also exposes that narrow closure as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning the full restore, so the direct-call contract asserted by tests/codex-config.test.cjs is unchanged. Adds a regression test that installs codex+kilo together, injects EACCES on the Kilo permission write by monkeypatching node:fs (restored in a finally — never chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps the bytes the successful install wrote. Verified red against the pre-fix unconditional path. Cline cannot host this test: its plan is writesSharedSettings:false + finishPermissionWriter:null, so its finishInstall performs no write and has no non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally (unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded fs.writeFileSync. * docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing CONTEXT.md still described Codex's rollback as an unconditional bind to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint- validation failures via rollbackInstallerMigrationsOnly and the configuredEntrypointValidation error tag. Caught during the round-6 PR body pass. * fix(#4249): stop rollbackInstallerMigrations meaning its own opposite Codex's install() result bound `rollbackInstallerMigrations` to restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the actual installer-migrations-only closure behind `rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name meant the opposite of what it says, and CONTEXT.md had to concede as much in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure for every runtime, matching both its name and the meaning it already has on next, and the snapshot restore gets its own Codex-only field, `rollbackPreInstallSnapshot`. The selection in rollbackFinalizedInstallerMigrations collapses to one line and no longer needs a fallback chain. Also in this commit, all against the same rollback path: - Correct the rollbackFinalizedInstallerMigrations comment. It read as if the round-6 narrowing prevented any sibling-triggered revert of a Codex install the user has already seen "Done!" for. It does not, and is not meant to: `wide` is true for ANY entrypoint-validation error from ANY runtime, because the aggregate gate is all-or-nothing — an invalid Cline entrypoint reverts Codex's snapshot, which tests/configured-entrypoint-validation.test.cjs's 'an aggregate entrypoint validation failure rolls the Codex install back (#4249)' asserts directly. The discriminator is the error's KIND, not which runtime owns the failing path. Comment and CONTEXT.md now say that. - Name the runtime in the "Configured entrypoint validation failed" error. ConfiguredEntrypointInvalid already carries `runtime`; the message threw it away, leaving an operator of a multi-runtime install unable to tell whose entrypoint broke — which matters precisely because the failure can revert a runtime that was itself fine. - Set `configuredEntrypoints: []` explicitly on the copilot-instructions early return. Every other branch states the key; this one relied on installAllRuntimes' `(result.configuredEntrypoints || [])` defence. `[]` is correct, not a workaround: every Copilot hook is an inline printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no GSD-managed script or interpreter to resolve. No behaviour change beyond the error-message text. * docs(#4249): narrow the smoke scan's config-surface claim to what it checks RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is told to launch is registered in one of settings.json / hooks.json / config.toml, and that nothing else in a config dir is runtime configuration. Both halves are false as stated. Cline registers its hook at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native [[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi), a directory separate from Kimi's own GSD configDir — the same gap installer-migration 007 already documents as structurally unreachable. The scan is in fact correct for what it runs against: entrypointRuntimes defaults to claude + codex, whose launch paths do all live in those three top-level files. Restate the docstring at that scope, name the two known out-of-scope surfaces, and warn that adding either runtime to entrypointRuntimes without teaching scanConfiguredEntrypoints about its surface yields a scan that finds zero entrypoints and proves nothing. The entrypointRuntimes default comment carried the same overgeneralization ("every other runtime reuses one of them") and is corrected with it. Documentation only; no code change. * fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found that install()'s settings-json early return for an unparseable settings.local.json omitted configuredEntrypoints and rollbackInstallerMigrations from its result, unlike every other branch. rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations unconditionally, so this branch silently dropped its own installer-migration rollback on a later finalize-stage failure. Completed the return shape: configuredEntrypoints: [] (matching Copilot's equally-early no-entrypoints-yet return) and rollbackInstallerMigrations (already in closure scope). Red-then-green regression test added. * test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs Same internal adversarial review (round 8): this suite's before() packed the tarball directly, without the ensureHooksDist() guard every sibling install-test suite (install.test.cjs, install-minimal-hooks.test.cjs, mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run in isolation ahead of a suite that builds hooks/dist itself, this suite's pack would ship a tarball with no hook scripts and fail closed on SMOKE.INIT_FAILED instead of testing anything. * fix(#4249): refresh stale test-timings weight for the codex-config split next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the CI shard packer's weight table (tests/test-timings.json) was never updated: codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block this PR extends with its own #4249 install()-pipeline test — had no entry at all, so the packer would silently underestimate it at the table's median weight (roughly a 9x underestimate against its real cost). trek-e's most recent review flagged a Windows shard timeout in-flight on codex-config.test.cjs, plausibly aggravated by this PR's own addition to that file before the rebase moved it. Re-measured both files locally (node --test --test-reporter=tap, max of 3 runs, matching the table's own max-across-streams methodology) and patched just these two entries — not a full regeneration, which would need real multi-lane CI data this session doesn't have access to. * fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists next's platform-conformance-tier classifier (#4591/#4598) landed after this branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's gen-platform-conformance-tier --check and both the Linux and macOS conformance suites. * fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation next's #4733 (landed after this branch's last rebase) replaced the static ISOLATED_HEAVY_FILES set with a threshold derived live from tests/test-timings.json, and pins the current derived result in EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still named codex-config.test.cjs, whose own weight this PR already dropped from 127783ms to 189ms (after splitting its heavy install()-pipeline blocks into codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The live-computed set correctly no longer includes it; the pinned expectation is updated to match. * fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it trek-e's review flagged two Major gaps: the thrown error read identically regardless of which of three real outcomes a runtime hit (nothing persisted / snapshot reverted / config left broken on disk), and no test proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/ Kimi/Cline — only Codex's revert path was ever asserted. assertConfiguredEntrypoints now tags each invalid entry with its actual consequence, mirrored from docs/how-to/update-gsd.md's existing rollback-matrix disclosure. A new test drives the same aggregate failure through Cline (whose own entrypoint is the one that fails) and asserts its hook file is still on disk afterward. * fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR One review pass (gemini-3.8-flash-high via the antigravity review lane) against this PR's full diff against next, findings independently verified against source before fixing: - Copilot's install() return object was the only one of 6 runtime branches missing rollbackInstallerMigrations — reachable now that this PR's own aggregate gate runs rollback across every result on any runtime's entrypoint failure, not just Copilot's own. - buildHookCommand's unresolved-bash early return skipped track() entirely, so a win32 install with no Git Bash silently produced an unregistered .sh hook instead of the 'unresolved-interpreter' validation failure configuredEntrypointsForHook's own comment said it would. - release-tarball-smoke.cjs reported a Cycle 4 install failure under SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing SMOKE.INSTALL_FAILED. - SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's trailing args, which also truncated any configDir containing a space (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan. Anchored the match on the already-known configDir prefix instead of a generic absolute-path guess: removes the ambiguity outright rather than patching the character class, and stays a raw-text scan on purpose (it catches a writer that emits a path without registering it — a JSON.parse of the expected schema would miss exactly that case). One suggested finding (test-timings.json "missing" the new test file) was verified false — that table only holds measured CI timings, populated after a file's first real run — and one Ponytail suggestion (a JSON.stringify dedup key) was rejected as it would reintroduce a real, if narrow, key-collision risk for no benefit. * fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment changeset-lint requires pr: to match the PR it runs on (16 on the fork, not the eventual upstream number) — rehearsal-branch convention already established earlier in this PR's history. lint-docs-guard-registration's quote-pairing heuristic doesn't require the docs/ path itself to be quoted — it flags a file once ANY quote-delimited span containing "docs/" appears anywhere in it, alongside any real fs read call. A comment ending "...update-gsd.md's rollback-matrix paragraph" supplied the closing quote character (the possessive apostrophe) the heuristic paired with an unrelated single-quoted string earlier in the file. Reworded to avoid the unquoted apostrophe next to the path. * chore(#4249): point the changeset pr field back at the upstream PR Fork rehearsal (PR #16) is green; the real target for this changeset is upstream PR #4249. --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
81 KiB
GSD Core Architecture
System architecture for contributors and advanced users. For user-facing documentation, see Feature Reference or User Guide.
Table of Contents
- System Overview
- Design Principles
- Component Architecture
- Agent Model
- Data Flow
- File System Layout
- Installer Architecture
- Hook System
- CLI Tools Layer
- Runtime Abstraction
System Overview
GSD Core is a meta-prompting framework that sits between the user and AI coding agents (Claude Code, Kimi CLI, OpenCode, Kilo, Codex, Copilot, Antigravity, Trae, Cline, Augment Code). It provides:
- Context engineering — Structured artifacts that give the AI everything it needs per task (see Context engineering)
- Multi-agent orchestration — Thin orchestrators that spawn specialized agents with fresh context windows (see Multi-agent orchestration)
- Spec-driven development — Requirements → research → plans → execution → verification pipeline
- State management — Persistent project memory across sessions and context resets
┌──────────────────────────────────────────────────────┐
│ USER │
│ /gsd-command [args] │
└─────────────────────┬────────────────────────────────┘
│
┌─────────────────────▼────────────────────────────────┐
│ COMMAND LAYER │
│ commands/gsd/*.md — Prompt-based command files │
│ (Claude Code custom commands / Codex skills) │
└─────────────────────┬────────────────────────────────┘
│
┌─────────────────────▼────────────────────────────────┐
│ WORKFLOW LAYER │
│ gsd-core/workflows/*.md — Orchestration logic │
│ (Reads references, spawns agents, manages state) │
└──────┬──────────────┬─────────────────┬──────────────┘
│ │ │
┌──────▼──────┐ ┌─────▼─────┐ ┌────────▼───────┐
│ AGENT │ │ AGENT │ │ AGENT │
│ (fresh │ │ (fresh │ │ (fresh │
│ context) │ │ context)│ │ context) │
└──────┬──────┘ └─────┬─────┘ └────────┬───────┘
│ │ │
┌──────▼──────────────▼─────────────────▼──────────────┐
│ CLI TOOLS LAYER │
│ gsd-tools.cjs command families + domain modules │
│ command-routing-hub + observability seams │
└──────────────────────┬───────────────────────────────┘
│
┌──────────────────────▼───────────────────────────────┐
│ FILE SYSTEM (.planning/) │
│ PROJECT.md | REQUIREMENTS.md | ROADMAP.md │
│ STATE.md | config.json | phases/ | research/ │
└──────────────────────────────────────────────────────┘
Design Principles
1. Fresh Context Per Agent
Every agent spawned by an orchestrator gets a clean context window (up to 200K tokens). This eliminates context rot — the quality degradation that happens as an AI fills its context window with accumulated conversation.
2. Thin Orchestrators
Workflow files (gsd-core/workflows/*.md) never do heavy lifting. They:
- Load context via
gsd-tools.cjs init <workflow> - Spawn specialized agents with focused prompts
- Collect results and route to the next step
- Update state between steps
3. File-Based State
All state lives in .planning/ as human-readable Markdown and JSON. No database, no server, no external dependencies. This means:
- State survives context resets (
/clear) - State is inspectable by both humans and agents
- State can be committed to git for team visibility
4. Absent = Enabled
Workflow feature flags follow the absent = enabled pattern. If a key is missing from config.json, it defaults to true. Users explicitly disable features; they don't need to enable defaults.
5. Defense in Depth
Multiple layers prevent common failure modes:
- Plans are verified before execution (plan-checker agent)
- Execution produces atomic commits per task
- Post-execution verification checks against phase goals
- UAT provides human verification as final gate
Component Architecture
Commands (commands/gsd/*.md)
User-facing entry points. Each file contains YAML frontmatter (name, description, allowed-tools) and a prompt body that bootstraps the workflow. Commands are installed as:
- Claude Code: Custom slash commands (hyphen form,
/gsd-command-name) - OpenCode / Kilo: Slash commands (hyphen form,
/gsd-command-name) - Codex: Skills (
$gsd-command-name) - Copilot: Slash commands (hyphen form,
/gsd-command-name) - Kimi CLI: Agent Skills (
/skill:gsd-command-name) plus an explicit custom agent launch withkimi --agent-file - Antigravity: Skills
Total commands: see docs/INVENTORY.md for the authoritative count and full roster.
Two-stage hierarchical routing (v1.40, #2792)
To keep the eager skill-listing token cost low, v1.40 introduces six namespace meta-skills (gsd-workflow, gsd-project, gsd-quality, gsd-context, gsd-manage, gsd-ideate — sourced from commands/gsd/ns-*.md, but the invocable name: is the bare form shown here) layered above the concrete sub-skills. On runtimes with non-recursive skill loaders (cline, qwen, hermes, augment, trae) the installer now realizes this fully: it emits only the 6 namespace router bundles as top-level skills and nests the ~61 concrete skills under <router>/skills/<name>/SKILL.md, so the eager listing is ≈6 entries instead of ≈67. The model selects a namespace router, which instructs it to read the nested concrete skill file via a routing table embedded in the router body. On these runtimes concrete skills are not directly invocable by bare name via the Skill tool; they are reachable through the router. Slash commands (/gsd-*, via the separate commands surface) are unaffected where the runtime has one. On runtimes with recursive or unconfirmed skill loaders (claude global, cursor, codex, copilot, windsurf, codebuddy, opencode, kilo, antigravity) the layout remains flat — all skills emitted at the top level as before. Antigravity moved from nested to flat in #1614: agy scans only skills/<name>/SKILL.md, so nested sub-skills were unreachable. Claude was reverted to flat in #924: the Skill tool hard-errors on unknown names rather than re-routing via the router, so nested concrete skills were uninvokable.
The router descriptions use pipe-separated keyword tags (≤ 60 chars) per the Tool Attention research showing keyword-dense tags outperform prose for routing at ~40 % the token cost.
MCP token-budget interaction
The eager skill listing is one of two recurring per-turn token costs. The other is the MCP tool schema injected by every enabled MCP server in .claude/settings.json. Heavyweight MCP servers (browser/playwright, Mac-tools, Windows-tools) can each cost 20 k+ tokens per turn — often dwarfing what model_profile tuning saves. The toggle lives in the Claude Code harness (enabledMcpjsonServers / disabledMcpjsonServers in .claude/settings.json) and is not a GSD concern. Together, the two-stage routing layer (#2792) and disciplined MCP enablement are the largest cost levers per turn. See docs/USER-GUIDE.md and references/context-budget.md for the audit checklist.
Workflows (gsd-core/workflows/*.md)
Orchestration logic that commands reference. Contains the step-by-step process including:
- Context loading via
gsd-tools.cjs inithandlers - Agent spawn instructions with model resolution
- Gate/checkpoint definitions
- State update patterns
- Error handling and recovery
Total workflows: see docs/INVENTORY.md for the authoritative count and full roster.
Progressive disclosure for workflows
Workflow files are loaded verbatim into Claude's context every time the
corresponding /gsd-* command is invoked. The workflow size budget enforced by
tests/workflow-size-budget.test.cjs keeps each file bounded, mirroring the
the agent size-budget convention. The budget is measured in bytes (#717), not lines:
line count over-penalizes prose and under-catches token-dense tables and code
blocks, whereas bytes are deterministic and match the unit our vendors bound on
— Codex truncates instruction docs past 32,768 bytes (project_doc_max_bytes).
We adopt that unit, not that exact number: the XL/LARGE ceilings below sit above
32,768 because these are grandfathered top-level orchestrators loaded by Claude,
not Codex AGENTS.md docs.
| Tier | Per-file byte limit |
|---|---|
XL |
90,000 — top-level orchestrators (execute-phase, plan-phase, new-project) |
LARGE |
54,000 — multi-step planners and large feature workflows |
DEFAULT |
38,000 — focused single-purpose workflows (the target tier) |
Ceilings are not fixed forever: under the tighten-only ratchet (#597) each one tracks its tier's current high-water mark within a small grace band, so budgets may only decrease over time.
Why the budget exists. With prompt caching the per-invocation cost of a large workflow is modest (cache reads run ~10% of input). The stronger, caching-independent reason is quality: as context grows, recall and reasoning degrade ("context rot" / attention budget), so leaner, higher-signal instructions produce better plans. The ceiling protects the agent's attention, not just the token bill.
Because the budget measures one file, it is a proxy for the real goal —
bounded loaded context. Extraction only helps when the extracted content is
loaded lazily (Read at the step that needs it). Moving prose into a file
that is still eagerly @-imported shrinks the measured file without shrinking
loaded context, which games the proxy rather than serving the goal.
workflows/discuss-phase.md is held to a stricter <30,000-byte ceiling per
the discuss-phase byte budget (#717; the discuss-phase/modes split keeps it ≈32000 bytes). When a workflow grows
beyond its tier, extract per-mode bodies into
workflows/<workflow>/modes/<mode>.md, templates into
workflows/<workflow>/templates/, and shared knowledge into
gsd-core/references/. The parent file becomes a thin dispatcher that
Reads only the mode and template files needed for the current invocation.
workflows/discuss-phase/ is the canonical example of this pattern —
parent dispatches, modes/ holds per-flag behavior (power.md, all.md,
auto.md, chain.md, text.md, batch.md, analyze.md, default.md,
advisor.md), and templates/ holds CONTEXT.md, DISCUSSION-LOG.md, and
checkpoint.json schemas that are read only when the corresponding output
file is being written.
workflows/plan-phase.md, workflows/execute-phase.md, and the
gsd-planner / gsd-executor agent definitions apply the same discipline
to their MVP-only reference bodies — planner-mvp-mode.md,
user-story-template.md, skeleton-template.md, and execute-mvp-tdd.md
are referenced for the planner/executor to Read only on MVP,
Walking-Skeleton, or MVP+TDD paths, rather than eagerly @-imported, so
non-MVP runs do not pay their context cost (guards against the "@-import
behind a conditional still loads eagerly" leak; see #720). The dedicated
mvp-phase workflow keeps its eager imports, since it is always MVP.
Agents (agents/*.md)
Specialized agent definitions with frontmatter specifying:
name— Agent identifierdescription— Role and purposetools— Allowed tool access (Read, Write, Edit, Bash, Grep, Glob, WebSearch, etc.)color— Terminal output color for visual distinction
Total agents: 33
References (gsd-core/references/*.md)
Shared knowledge documents that workflows and agents @-reference (see docs/INVENTORY.md for the authoritative full roster):
Core references:
checkpoints.md— Checkpoint type definitions and interaction patternsgates.md— 4 canonical gate types (Confirm, Quality, Safety, Transition) wired into plan-checker and verifiermodel-profiles.md— Per-agent model tier assignmentsmodel-profile-resolution.md— Model resolution algorithm documentationverification-patterns.md— How to verify different artifact typesverification-overrides.md— Per-artifact verification override rulesplanning-config.md— Full config schema and behaviorgit-integration.md— Git commit, branching, and history patternsgit-planning-commit.md— Planning directory commit conventionsquestioning.md— Dream extraction philosophy for project initializationtdd.md— Test-driven development integration patternsui-brand.md— Visual output formatting patternscommon-bug-patterns.md— Common bug patterns for code review and verification
Workflow references:
agent-contracts.md— Formal interface between orchestrators and agentscontext-budget.md— Context window budget allocation rulescontinuation-format.md— Session continuation/resume formatdomain-probes.md— Domain-specific probing questions for discuss-phasegate-prompts.md— Gate/checkpoint prompt templatesrevision-loop.md— Plan revision iteration patternsuniversal-anti-patterns.md— Common anti-patterns to detect and avoidartifact-types.md— Planning artifact type definitionsphase-argument-parsing.md— Phase argument parsing conventionsdecimal-phase-calculation.md— Decimal sub-phase numbering rulesworkstream-flag.md— Workstream active pointer conventionsuser-profiling.md— User behavioral profiling methodologythinking-partner.md— Conditional thinking partner activation at decision points
Thinking model references:
References for integrating thinking-class models (o3, o4-mini, Gemini 2.5 Pro) into GSD workflows:
thinking-models-debug.md— Thinking model patterns for debugging workflowsthinking-models-execution.md— Thinking model patterns for execution agentsthinking-models-planning.md— Thinking model patterns for planning agentsthinking-models-research.md— Thinking model patterns for research agentsthinking-models-verification.md— Thinking model patterns for verification agents
Modular planner decomposition:
The planner agent (agents/gsd-planner.md) was decomposed from a single monolithic file into a core agent plus reference modules to stay under the 50K character limit imposed by some runtimes:
planner-gap-closure.md— Gap closure mode behavior (reads VERIFICATION.md, targeted replanning)planner-reviews.md— Cross-AI review integration (reads REVIEWS.md from/gsd-review)planner-revision.md— Plan revision patterns for iterative refinement
Templates (gsd-core/templates/)
Markdown templates for all planning artifacts. Used by gsd-tools.cjs template fill / phase.scaffold (and top-level scaffold) to create pre-structured files:
project.md,requirements.md,roadmap.md,state.md— Core project filesphase-prompt.md— Phase execution prompt templatesummary.md(+summary-minimal.md,summary-standard.md,summary-complex.md) — Granularity-aware summary templatesDEBUG.md— Debug session tracking templateUI-SPEC.md,UAT.md,VALIDATION.md— Specialized verification templatesdiscussion-log.md— Discussion audit trail templatecodebase/— Brownfield mapping templates (architecture, stack)research-project/— Research output templates (SUMMARY, STACK, FEATURES, ARCHITECTURE, PITFALLS)
Hooks (hooks/)
Runtime hooks that integrate with the host AI agent:
| Hook | Event | Purpose |
|---|---|---|
gsd-statusline.js |
statusLine |
Displays model (long-context suffixes like (1M context) collapse to a compact (1M) badge), task, directory, and context usage bar |
gsd-context-monitor.js |
PostToolUse / AfterTool |
Injects agent-facing context warnings at 35%/25% remaining by default (configurable — see CONFIGURATION.md) |
gsd-check-update.js |
SessionStart |
Foreground trigger for the background update check |
gsd-ensure-canonical-path.js |
SessionStart |
For Claude Code plugin installs, symlinks ~/.claude/gsd-core/{bin,contexts,references,templates,workflows} to the plugin's bundled tree so @~/.claude/gsd-core/... includes resolve; runs first in SessionStart, no-op in classic installs, self-heals after claude plugin update (#997) |
gsd-check-update-worker.js |
(helper) | Background worker spawned by gsd-check-update.js; no direct event registration |
gsd-prompt-guard.js |
PreToolUse |
Scans .planning/ writes for prompt injection patterns (advisory) |
gsd-read-injection-scanner.js |
PostToolUse |
Scans Read tool output for injected instructions in untrusted content |
gsd-workflow-guard.js |
PreToolUse |
Detects file edits outside GSD workflow context (advisory, opt-in via hooks.workflow_guard) |
gsd-secret-read-guard.js |
PreToolUse |
Hard-blocks Read / Grep / Bash reads of .env, .env.<suffix> (templates such as .env.example exempt) and .secrets; replaces the installer-written Read(.env*) permission deny rules, which made every cd DIR && grep … compound prompt for approval on Claude Code ≥ 2.1.259 (#4221) |
gsd-read-guard.js |
PreToolUse |
Advisory guard preventing Edit/Write on files not yet read in the session |
gsd-session-state.sh |
SessionStart |
Session state tracking for shell-based runtimes |
gsd-validate-commit.sh |
PreToolUse |
Commit validation for conventional commit enforcement |
gsd-phase-boundary.sh |
PostToolUse |
Phase boundary detection for workflow transitions |
See docs/INVENTORY.md for the authoritative hook roster.
Crash policy (ADR-3889 Phase 7, #3911). Every enforcement hook terminates
through hooks/lib/hook-exit.js's allow(payload) (exit 0), deny(payload, stderrPayload?) (exit 2), or crash(onCrash, payload) — the last dispatching
per a HOOK_ON_CRASH policy the hook must declare explicitly (ALLOW or
DENY, no default), so a hook's fail-open/fail-closed stance is a visible
declaration rather than an inference from a bare process.exit(N). Two hooks
are deliberate exceptions — gsd-read-injection-scanner.js (PostToolUse) and
gsd-cursor-subagent-start.js (Cursor) — whose harnesses read the block
decision from the JSON response body at exit 0, not from the exit code, so
they never call deny(). See
Declare a hook's crash policy.
Command Routing Hub (gsd-core/bin/lib/command-routing-hub.cjs)
CJS command family routers dispatch through CommandRoutingHub. The hub owns the no-throw pure-result contract (hub.dispatch() catches internal exceptions and returns { ok: false, kind, ...typedPayload }) and the closed runtime error taxonomy (UnknownCommand, InvalidArgs, HandlerRefusal, HandlerFailure). Router adapters remain thin CLI translators — they build the hub, call dispatch, then map the Result to output()/error() calls. The runtime is single-path (no dual-runtime mode selection). See docs/adr/0174-retire-gsd-sdk-package-boundary.md.
Planned (ADR-2346 / epic #2345): the
runCommand73-case switch is being dissolved into a two-layer dispatch — families via thecommandFamiliesregistry (ADR-959 mechanism, completed) and single-purpose leaf verbs via a table filling the prepared_dispatchNonFamilyseam — collapsingrunCommandto a ~15-line dispatcher. Behavior-preserving; tracked phase-by-phase under epic #2345. The current-state description above holds until each phase lands.
Capability Command Dispatch (gsd-core/bin/gsd-tools.cjs, ADR-1244 D7)
Command families declared by capabilities (commands: [{ family, module, router }]) are dispatched from the registry rather than a hardcoded switch. The runCommand default arm tries, in order:
- First-party —
dispatchCapabilityCommandagainst the frozencapability-registry.cjscommandFamilies, loading the router frombin/lib/. The in-tree families (graphify,intel,audit) reach their routers this way (the legacy hardcoded switch is retired). - Third-party (installed overlay) —
dispatchOverlayCapabilityCommandcallsloadRegistry({ includeInstalled })and dispatches a family only when itscapIdappears in_overlay.commandRoots. The loader lists a command root only for an accepted overlay capability with a committed ledger entry (consent gate), and the router module isrequire()'d from that capability's install root, confined by basename validation +realpathcontainment (rejecting..traversal and symlink escape). This is the one point where third-party capability code executes; see the capability trust model for the consent + confinement + project-scope trust boundary.
Both paths share the same guards: prototype-pollution-safe command keys, an own-property router check, and synchronous-only routers (an async router is a fail-fast error).
Reviewer-Lane Capability Trait (#4209, ADR-2782)
/gsd-code-review optionally corroborates its internal review with external reviewer lanes (--codex, --agy, ...), gated by the reusable supportsReviewerLanes capability-step trait and dispatched through the single dispatchReviewerLanes interpreter — see gsd-core/references/loop-hook-dispatch.md for the trait and src/reviewer-step-dispatch.cts for the interpreter's fail-closed contract. gsd-code-reviewer is the sole consolidator: it independently re-verifies every external claim against the actual source before writing anything to REVIEW.md, so a lane's evidence is corroborating input, never a second output schema.
Research Module (src/research-{store,provider}.cts, src/package-legitimacy.cts)
The Research Module implements an L2-hybrid seam: code owns the cache, provider policy, and package legitimacy verdicts; MCP owns the actual network fetch.
Three compiled modules (generated to gsd-core/bin/lib/*.cjs per ADR-457) are reachable via gsd-tools query research-plan | research-store | package-legitimacy:
- Research Store — content-addressed cache (
sha256(ecosystem+library+version+query+kind)) with per-source TTL (curated-doc: 30 d, medium: 7 d, web/synthesis: 1 d) and two storage tiers:~/.gsd/research-cachefor cross-project curated-doc hits,.planning/research/.cachefor project-local web/synthesis results. - Research Provider — single
PROVIDER_WATERFALL(Context7→Ref→Jina→websearchfor docs;Exa→Tavily→Perplexity→Brave→websearchfor web;Firecrawl→Jinafor scrape-only).planResearch()returns cache hits plus a fetch plan;classifyConfidence()stampsHIGH|MEDIUM|LOWby provider tier. - Package Legitimacy — registry-API verdicts (npm/PyPI/crates.io injectable adapters) producing
OK|SUS|SLOPper package.slopcheckis an optional escalate-only adapter; absence leaves registry verdicts intact rather than downgrading everything to[ASSUMED].
Data flow:
agent
│
▼
gsd-tools query research-plan ← Research Provider: check cache, build fetch plan
│
├── [cache hits] ──────────────────► RESEARCH.md (digest only, no raw content)
│
└── [fetch plan] ──────────────────► MCP fetch (agent calls MCP tools with the plan)
│
▼
gsd-tools query research-store (put)
│
▼
RESEARCH.md path returned to orchestrator
Agents always return a RESEARCH.md path, never raw fetched content. Context discipline is enforced through subagent isolation, compact provider output, and fetch-to-disk. See ADR-0656.
Context Predicate Fact-Store (src/context-predicates.cts, ADR-1671)
The CONTEXT.md predicate fact-store — every backtick-wrapped CLASS.subkey=value declaration in the repo-root CONTEXT.md — has a compiled parser/selector seam (generated to gsd-core/bin/lib/context-predicates.cjs per ADR-457) reachable live via gsd-tools query context-predicates --class|--prefix|--contains. Fence-aware line skipping mirrors markdown-sectionizer.cts's exported scanFencedBlocks delimiter-matching rule exactly (proven by a fence-skip parity test suite), but is scanned by a LOCAL, interleaved single pass rather than a call into that seam directly: fences and HTML comments must mutually suppress each other's open/close detection while either is active (a fence delimiter inside a real comment, or a comment token inside a real fence, must not falsely toggle the other construct), and that precedence cannot be resolved by two independent passes over scanFencedBlocks's comment-blind output — see src/context-predicates.cts's module doc comment.
scripts/gen-context-index.cjs --check is the CI drift-guard for the committed docs/CONTEXT-INDEX.json artifact: it fails on staleness between a fresh parse of CONTEXT.md and the committed file, and on any duplicate predicate ID. It is wired into lint:generated-sync (so lint:ci, so CI). docs/CONTEXT-INDEX.json is generated — never hand-edit it; regenerate with gen-context-index.cjs --write (also wired into build, after build:lib, and into regen:derived). The generator require()s the compiled context-predicates.cjs, so it must run after build:lib in any pipeline; .github/workflows/test.yml does this.
The committed index intentionally carries no line field for any predicate (ADR-1671 open question 4, resolved by #2928) — committed-but-uncompared metadata goes silently stale, the same defect class the drift-guard exists to catch, with the alarm removed. The live gsd-tools query context-predicates parse still returns line/section for callers that want to cite a source location. See ADR-1671 and CLI Tools Reference.
Workflow Fragmentization and Emission (src/workflow-fragments.cts, ADR-1671)
Workflow markdown under gsd-core/workflows/*.md can mark one or more sections with an
in-file <!-- gsd:section id="<id>" when="<when>" --> / <!-- /gsd:section --> pair. A
compiled parser/composer seam (generated to gsd-core/bin/lib/workflow-fragments.cjs per
ADR-457) partitions a marked document into fragments and recomposes them through the shared
context-composer.cjs budget seam (ADR-1671, #2929) before any per-runtime converter sees the
text — so a marker attribute can never be corrupted by a .claude/ → .windsurf/-style
path-rewrite regex. bin/install.js's copyWithPathReplacement calls composeWorkflow on
every workflow file at emit time; an unmarked file (88 of the 89 shipped workflows today)
parses to a single implicit fragment and round-trips byte-identical, so this is a no-op for
every workflow that hasn't opted in yet.
Every fragment in this phase carries the verbatim strategy, so composition is structurally
non-lossy — nothing is trimmed regardless of budget. Fence and HTML-comment interleaving
reuses the same LOCAL, single-pass, mutually-suppressing scan discipline as
context-predicates.cts (see above), so a marker-shaped line inside a fenced code block or an
unrelated comment is never misread as structural. Markers are stripped at emit — the
installed artifact carries no build metadata and is smaller than the source by exactly the
stripped marker bytes.
See Reference: Workflow fragments for the full marker
grammar, the frozen when= vocabulary, and fail-closed authoring rules, and
ADR-1671 (open questions 1 and 2) for why
in-file markers were chosen over separate fragment files or a sidecar manifest.
Section Manifest (src/section-manifest.cts, ADR-1671 Phases 5 and 6.1)
Two seams turn a workflow's gsd:section markers into per-invocation applicability data.
scripts/gen-section-manifest.cjs --write (wired into build after build:lib, and into
lint:generated-sync) scans gsd-core/workflows/*.md and writes the committed
gsd-core/workflows/section-manifest.json, keyed per workflow —
{workflows: {"<name>": [{id, when, read}]}} — where read is the path of the step file the
section's body was extracted to. A workflow with no marked sections contributes no key at
all: an absent key means degraded/unknown (the caller reads every section, the safe superset),
while a key present with an empty array means "computed, nothing applies". The generator reuses
parseWorkflowSections unchanged rather than re-implementing marker parsing, and fails closed
(--check) on a marker naming a step file that does not exist, a step file no marker
references, or a committed artifact still carrying the pre-6.1 flat {sections: [...]} shape.
A separate pure evaluator, src/section-manifest.cts (compiled to
gsd-core/bin/lib/section-manifest.cjs per ADR-457), maps one invocation's facts —
{flags, phaseNumber, hasPriorPhases} plus the optional needsCodebaseMap, phaseMvpMode and
worktreesEnabled booleans — to an included/excluded partition of section ids via
selectSections. flags is a ReadonlySet<string> of flag tokens; because parseNamedArgs
always materializes a boolean flag key (false when the token was absent, never undefined),
presence is truthiness, and the init router folds a boolean flag's own false into the
absent sentinel before the facts are built. Per Greenspun's Tenth Rule, the evaluator is a total
lookup over the frozen 14-atom when= vocabulary, never a parser: WHEN_PREDICATES is a
hand-written literal map that never derives a predicate from its atom string, and an
unrecognized value fails closed rather than being silently excluded. An atom is admitted only
when it has both a real consuming section and a fact the init seam actually computes — an atom
without the latter would evaluate false forever and silently disable its own section.
execute-phase.md's partial-wave and gap-closure-artifacts sections — previously inlined
directly per #2930's pilot — now delegate to dedicated step files under
gsd-core/workflows/execute-phase/steps/, the same pattern the pre-existing regression-gate
section already used.
CLI Tools (gsd-core/bin/)
Node.js CLI utility (gsd-tools.cjs) with domain modules split across gsd-core/bin/lib/ (see docs/INVENTORY.md for the authoritative roster):
| Module | Responsibility |
|---|---|
config-loader.cjs |
Project config loading — defaults merge, legacy-key migration, workstream overlay, unknown-key/profile-override validation, and federated config overlay (ADR-857 phase 3b) (extracted from core.cjs, ADR-857) |
federated-config.cjs |
Defensive merge of capability-declared config slices (ADR-857 phase 3b); exports mergeFederatedConfig; live for migrated Capability keys that are absent from the central config schema |
core-utils.cjs |
Shared low-level utility primitives — POSIX path normalization, sub-repo/subdirectory scanning, phase file stats, slug/one-liner/plan-id helpers, time-ago (extracted from core.cjs, ADR-857) |
core.cjs |
Shared utilities; compatibility re-exports for planning, I/O (io.cjs), and phase-id helpers |
io.cjs |
CLI I/O primitives — output/error emission, JSON-error mode, large-payload temp-file spillover |
phase-id.cjs |
Pure phase-id parsing/matching helpers — normalize, token match, regex builders (extracted from core.cjs, ADR-857) |
phase-locator.cjs |
Phase-directory search and location — active-phase discovery (searchPhaseInDir, findPhaseInternal) and archived-phase-dir enumeration (getArchivedPhaseDirs), matching phase ids/tokens against the filesystem (extracted from core.cjs, ADR-857) |
roadmap-parser.cjs |
ROADMAP.md parsing — milestone slicing, current-milestone extraction, phase/milestone lookups, milestone-phase filter (extracted from core.cjs, ADR-857) |
planning-workspace.cjs |
Planning seam (planningDir, planningPaths, active workstream routing, .planning/.lock) |
state.cjs |
STATE.md parsing, updating, progression, metrics |
phase.cjs |
Phase directory operations, decimal numbering, plan indexing |
roadmap.cjs |
ROADMAP.md parsing, phase extraction, plan progress |
config.cjs |
config.json read/write, section initialization |
verify.cjs |
Plan structure, phase completeness, reference, commit validation |
template.cjs |
Template selection and filling with variable substitution |
frontmatter.cjs |
YAML frontmatter CRUD operations |
init.cjs |
Compound context loading for each workflow type |
milestone.cjs |
Milestone archival, requirements marking |
commands.cjs |
Misc commands (slug, timestamp, todos, scaffolding, stats) |
model-profiles.cjs |
Model profile resolution table |
model-resolver.cjs |
Model and effort resolution policy — resolves model, tier, granularity, effort, and fast-mode for a given agent from project config and model profiles/catalog (extracted from core.cjs, ADR-857) |
security.cjs |
Path traversal prevention, prompt injection detection, safe JSON parsing, shell argument validation |
uat.cjs |
UAT file parsing, verification debt tracking, audit-uat support |
docs.cjs |
Docs-update workflow init, Markdown scanning, monorepo detection |
workstream.cjs |
Workstream CRUD, migration, session-scoped active pointer |
schema-detect.cjs |
Schema-drift detection for ORM patterns (Prisma, Drizzle, etc.) |
profile-pipeline.cjs |
User behavioral profiling data pipeline, session file scanning |
profile-output.cjs |
Profile rendering, USER-PROFILE.md and dev-preferences.md generation |
context-predicates.cjs |
CONTEXT.md predicate fact-store parser/selector (ADR-1671, #2928); backs query context-predicates and scripts/gen-context-index.cjs's docs/CONTEXT-INDEX.json drift guard; compiled from src/context-predicates.cts |
loop-host-contract.cjs |
Generated Loop Host Contract — 12 loop points, per-step agent roles, and core artifacts; emitted by scripts/gen-loop-host-contract.cjs from workflow markers (ADR-894 §3); consumed by gen-capability-registry.cjs |
capability-loader.cjs |
Runtime registry overlay loader (ADR-1244 D2) — loadRegistry({ includeInstalled }) composes the frozen first-party registry with a validated installed overlay of third-party capability manifests read from global $GSD_HOME/.gsd/capabilities/ and project <projectRoot>/.gsd/capabilities/; first-party always wins; load-time engines.gsd re-gate skips incompatible overlays with a warning; gate-kind hooks on skipped capabilities fail OPEN — no gate is injected; a loud warning (stderr + envelope warnings) names the load failure and the gsd capability remove <id> remediation (#2009) |
capability-registry.cjs |
Generated central Capability Registry — role-partitioned index of all co-located capability declarations; emitted by scripts/gen-capability-registry.cjs (ADR-894 §5) |
loop-resolver.cjs |
Loop Extension Point resolver — ADR-857 phase 3c registry-consuming query; consumes resolved Capability State, filters byLoopPoint by capability enablement plus config activation, renders active hooks as markdown, emits { point, activeHooks, rendered } envelope; gsd-tools loop render-hooks <point> [--config-dir <path>] |
capability-state.cjs |
Unified capability-state resolver — ADR-857 phase 4b/6; composes install profile, runtime surface, and config activation into one per-capability view consumed by workflow hook rendering; pure resolveCapabilityState, reusable resolveCapabilityRuntimeState, I/O cmdCapabilityState, and convenience predicate isCapabilityActive(capId, cwd); gsd-tools capability state [--config-dir <path>] emits { runtimeConfigDir, capabilities[] } where each entry carries enabled (installed && surfaced) and active (enabled && configActivation via the capability's activationKey; absent key → active===enabled) |
capability-validator.cjs |
Shared capability conformance validator (ADR-1244 D2) — extracted from scripts/gen-capability-registry.cjs so the build-time generator and the runtime overlay loader share one validateCapability(manifest) implementation; generative-parity is CI-guarded |
graphify-command-router.cjs |
ADR-959 capability command router — first real capability command cutover (phase 4d-impl-2); extracted from the case 'graphify': arm in gsd-tools.cjs; dispatches build/query/status/diff subcommands; discovered via commandFamilies in the capability registry |
audit-command-router.cjs |
ADR-959 capability command router (phase 4d-impl-3); extracted from the case 'audit-uat': and case 'audit-open': arms in gsd-tools.cjs; routeAuditUat → uat.cjs:cmdAuditUat, routeAuditOpen → audit.cjs:{auditOpenArtifacts,formatAuditReport}; discovered via commandFamilies in the capability registry |
intel-command-router.cjs |
ADR-959 capability command router (phase 4d-impl-4, last first-party cutover); extracted from the case 'intel': arm in gsd-tools.cjs; routeIntelCommand → all 9 intel subcommands via lazy require('./intel.cjs'); preserves non-raw timeAgo transform on status.files[*].updated_at; discovered via commandFamilies in the capability registry |
runtime-hooks-surface.cjs |
Hook-surface writer and configured-entrypoint validation module (ADR-857 phase 5f-1); owns Cline rules/agents-md/pre-tool-use hook generation, Cursor/Windsurf hooks.json, Kimi hook TOML, Copilot session-hook config, Codex hook-block management, and construction-time records for paths emitted into runtime configuration. Installer success requires those records to pass non-executing file-type and interpreter-resolution checks. |
Agent Model
Orchestrator → Agent Pattern
Orchestrator (workflow .md)
│
├── Load context: gsd-tools.cjs init <workflow> <phase>
│ Returns JSON with: project info, config, state, phase details
│
├── Resolve model: gsd-tools.cjs resolve-model <agent-name>
│ Returns: opus | sonnet | haiku | inherit
│
├── Spawn Agent (Task/SubAgent call)
│ ├── Agent prompt (agents/*.md)
│ ├── Context payload (init JSON)
│ ├── Model assignment
│ └── Tool permissions
│
├── Collect result
│
└── Update state: gsd-tools.cjs state update / state patch / state advance-plan
Primary Agent Spawn Categories
Conceptual spawn-pattern taxonomy for the primary agents. For the authoritative agent roster (including the advanced/specialized agents such as gsd-pattern-mapper, gsd-code-reviewer, gsd-code-fixer, gsd-ai-researcher, gsd-domain-researcher, gsd-eval-planner, gsd-eval-auditor, gsd-framework-selector, gsd-debug-session-manager, gsd-intel-updater), see docs/INVENTORY.md.
| Category | Agents | Parallelism |
|---|---|---|
| Researchers | gsd-project-researcher, gsd-phase-researcher, gsd-ui-researcher, gsd-advisor-researcher | 4 parallel (stack, features, architecture, pitfalls); advisor spawns during discuss-phase |
| Synthesizers | gsd-research-synthesizer | Sequential (after researchers complete) |
| Planners | gsd-planner, gsd-roadmapper | Sequential |
| Checkers | gsd-plan-checker, gsd-integration-checker, gsd-ui-checker, gsd-nyquist-auditor | Sequential (verification loop, max 3 iterations) |
| Executors | gsd-executor | Parallel within waves, sequential across waves |
| Verifiers | gsd-verifier | Sequential (after all executors complete) |
| Mappers | gsd-codebase-mapper | 4 parallel (tech, arch, quality, concerns) |
| Debuggers | gsd-debugger | Sequential (interactive) |
| Auditors | gsd-ui-auditor, gsd-security-auditor | Sequential |
| Doc Writers | gsd-doc-writer, gsd-doc-verifier | Sequential (writer then verifier) |
| Profilers | gsd-user-profiler | Sequential |
| Analyzers | gsd-assumptions-analyzer | Sequential (during discuss-phase) |
Wave Execution Model
During execute-phase, plans are grouped into dependency waves:
Wave Analysis:
Plan 01 (no deps) ─┐
Plan 02 (no deps) ─┤── Wave 1 (parallel)
Plan 03 (depends: 01) ─┤── Wave 2 (waits for Wave 1)
Plan 04 (depends: 02) ─┘
Plan 05 (depends: 03,04) ── Wave 3 (waits for Wave 2)
Each executor gets:
- Fresh 200K context window (or up to 1M for models that support it)
- The specific PLAN.md to execute
- Project context (PROJECT.md, STATE.md)
- Phase context (CONTEXT.md, RESEARCH.md if available)
Adaptive Context Enrichment (1M Models)
When the context window is 500K+ tokens (1M-class models like Opus 4.6, Sonnet 4.6), subagent prompts are automatically enriched with additional context that would not fit in standard 200K windows:
- Executor agents receive prior wave SUMMARY.md files and the phase CONTEXT.md/RESEARCH.md, enabling cross-plan awareness within a phase
- Verifier agents receive all PLAN.md, SUMMARY.md, CONTEXT.md files plus REQUIREMENTS.md, enabling history-aware verification
The orchestrator reads context_window from config (gsd-tools.cjs config-get context_window) and conditionally includes richer context when the value is >= 500,000. For standard 200K windows, prompts use truncated versions with cache-friendly ordering to maximize context efficiency.
Parallel Commit Safety
When multiple executors run within the same wave, two mechanisms prevent conflicts:
--no-verifycommits — Parallel agents skip pre-commit hooks (which can cause build lock contention, e.g., cargo lock fights in Rust projects). The orchestrator runsgit hook run pre-commitonce after each wave completes.- STATE.md file locking — All
writeStateMd()calls use lockfile-based mutual exclusion (STATE.md.lockwithO_EXCLatomic creation). This prevents the read-modify-write race condition where two agents read STATE.md, modify different fields, and the last writer overwrites the other's changes. Includes stale lock detection (10s timeout) and spin-wait with jitter.
The STATE.md Write Path
Locking decides who writes. A separate contract decides what survives the write.
STATE.md carries the same fact in two places — YAML frontmatter and the document body — and the body is authoritative. Every write therefore re-derives frontmatter from the body, which raises the question the write path exists to answer: when a re-derived value disagrees with the one already in frontmatter, which wins?
FIELD_CLASSIFICATION (src/state-transition.cts) answers it per field, declaring a preservation policy — preserve-when-unchanged, preserve-always, preserve-if-placeholder, derive — that applyStatePreservation executes after syncStateFrontmatter re-derives. (A fifth policy, clear, was listed here until ADR-3408 §8.6's amendment removed it: no row used it and no executor existed for it.)
The pipeline's precondition is a type, not a convention (ADR-3473 §8.6). A policy row can only be honored if the pre-write frontmatter snapshot it compares against is actually present. That snapshot now travels as a StateTransaction, built by openStateTransaction() — preservation applies — or rebuildStateTransaction() — it does not. Both carry the snapshot, and a transaction cannot be constructed without one: an absent snapshot is a construction failure, not a runtime skip. That distinction is the whole point. Previously the snapshot was nulled to signal "re-derive from disk", so a declared preserve-always row and a silently-skipped one were indistinguishable at runtime, which is how a curated progress: block was erased by verbs that had nothing to do with progress.
rebuildStateTransaction() is the typed form of ADR-3408 §8.3's closed exception list: state sync, which exists to let the body win, and /gsd-health --repair's factory reset. Both are deliberate and permanent, not debt — and because the type names them, the write-path drift guard no longer has to track them as strings in a ratcheted baseline.
What a command reports it wrote is the same snapshot, read back (ADR-3473 §8.7). Every state.* command returns an updated array. That array is now derived by comparing what was actually persisted against the transaction's pre-write snapshot: a field appears if, and only if, its persisted value changed. One comparison answers both of the questions that used to need separate machinery — a field the caller asked for that the pipeline then discarded is persisted-equals-snapshot and drops out, and a field nobody asked for that the write moved anyway is different and appears. Nothing is filtered by its preservation policy.
Reporting is at leaf granularity: when a single counter moves you are told progress.total_plans, not progress. The leaves are the ones the field-classification table already declares, so the report is bounded by a schema rather than by walking the document.
One field is excluded, and it is excluded for its provenance rather than its policy: last_updated is stamped on every save regardless of what you changed, so admitting it would make state patch's success signal — which is simply whether updated is non-empty — permanently true, and a patch in which every field failed would report success. state_head is deliberately not excluded: it is recomputed on every save but only changes when the commit it records actually moved, so reporting it tells you something true.
A practical consequence worth knowing: these arrays are now longer than they used to be, because they used to under-report. If you compare one exactly, expect more entries — and expect them to be the ones that really changed.
ADR-3408 is the normative contract for that path: one executor per declared policy, one write seam, and reports computed from what was actually persisted rather than from what the caller intended to write. Where the contract and the code disagree, the code is the defect. It is the write-side counterpart of ADR-3180, which gave each read-side derivation a single owner.
ADR-3473 owns the invariants that sit outside that contract. ADR-3408 governs what survives a write; it does not govern the pipeline's precondition, what a command reports it wrote, or where the set of STATE.md keys, types and enums is declared. ADR-3473 owns those, alongside document parsing, enumeration, and the return contract of every routine that can fail. It is the third application of ADR-3180's mechanism and the first whose success metric requires the guard surface to shrink as each seam lands.
Data Flow
New Project Flow
User input (idea description)
│
▼
Questions (questioning.md philosophy)
│
▼
4x Project Researchers (parallel)
├── Stack → STACK.md
├── Features → FEATURES.md
├── Architecture → ARCHITECTURE.md
└── Pitfalls → PITFALLS.md
│
▼
Research Synthesizer → SUMMARY.md
│
▼
Requirements extraction → REQUIREMENTS.md
│
▼
Roadmapper → ROADMAP.md
│
▼
User approval → STATE.md initialized
Phase Execution Flow
discuss-phase → CONTEXT.md (user preferences)
│
▼
ui-phase → UI-SPEC.md (design contract, optional)
│
▼
plan-phase
├── Research gate (blocks if RESEARCH.md has unresolved open questions)
├── Phase Researcher → RESEARCH.md
│ └── Package Legitimacy Gate: registry-API verdict on every package; [SLOP] removed,
│ [SUS]/[ASSUMED] flagged; Audit table written to RESEARCH.md
├── Planner (with reachability check) → PLAN.md files
│ └── checkpoint:human-verify injected before [ASSUMED]/[SUS] installs;
│ T-{phase}-SC STRIDE row added for install-bearing plans
├── Plan Checker → Verify loop (max 3x)
├── Requirements coverage gate (REQ-IDs → plans)
└── Decision coverage gate (CONTEXT.md `<decisions>` → plans, BLOCKING — #2492)
│
▼
state planned-phase → STATE.md (Planned/Ready to execute)
│
▼
execute-phase (context reduction: truncated prompts, cache-friendly ordering)
├── Wave analysis (dependency grouping)
├── Executor per plan → code + atomic commits
├── SUMMARY.md per plan
└── Verifier → VERIFICATION.md
└── Decision coverage gate (CONTEXT.md decisions → shipped artifacts, NON-BLOCKING — #2492)
│
▼
verify-work → UAT.md (user acceptance testing)
│
▼
ui-review → UI-REVIEW.md (visual audit, optional)
Context Propagation
Each workflow stage produces artifacts that feed into subsequent stages:
PROJECT.md ────────────────────────────────────────────► All agents
REQUIREMENTS.md ───────────────────────────────────────► Planner, Verifier, Auditor
ROADMAP.md ────────────────────────────────────────────► Orchestrators
STATE.md ──────────────────────────────────────────────► All agents (decisions, blockers)
CONTEXT.md (per phase) ────────────────────────────────► Researcher, Planner, Executor
RESEARCH.md (per phase) ───────────────────────────────► Planner, Plan Checker
PLAN.md (per plan) ────────────────────────────────────► Executor, Plan Checker
SUMMARY.md (per plan) ─────────────────────────────────► Verifier, State tracking
UI-SPEC.md (per phase) ────────────────────────────────► Executor, UI Auditor
File System Layout
Installation Files
~/.claude/ # Claude Code (global install)
├── skills/gsd-ns-*/SKILL.md # Global skills — nesting runtimes: 6 namespace routers (authoritative roster: docs/INVENTORY.md)
│ └── skills/<name>/SKILL.md # concrete skills nested under each router
│ (flat runtimes: skills/gsd-*/SKILL.md — all ~67 skills at top level)
├── commands/gsd/*.md # Local Claude installs use slash commands instead of global skills
├── gsd-core/
│ ├── bin/gsd-tools.cjs # CLI utility
│ ├── bin/lib/*.cjs # Domain modules (authoritative roster: docs/INVENTORY.md)
│ ├── workflows/*.md # Workflow definitions (authoritative roster: docs/INVENTORY.md)
│ ├── references/*.md # Shared reference docs (authoritative roster: docs/INVENTORY.md)
│ └── templates/ # Planning artifact templates
├── agents/*.md # Agent definitions (authoritative roster: docs/INVENTORY.md)
├── hooks/*.js # Node.js hooks (statusline, guards, monitors, update check)
├── hooks/*.sh # Shell hooks (session state, commit validation, phase boundary)
├── settings.json # Hook registrations
└── VERSION # Installed version number
Equivalent paths for other runtimes:
- OpenCode:
~/.config/opencode/global or./.opencode/local - Kilo:
~/.config/kilo/global or./.kilo/local - Kimi CLI: first-existing generic global root (
~/.config/agents/recommended, then~/.agents/if itsskills/directory already exists); local install is deferred and guarded - Codex:
~/.codex/global or./.codex/local - Copilot:
~/.copilot/global or./.github/local - Antigravity: auto-detected global root (
~/.gemini/antigravity/,~/.gemini/antigravity-ide/, or~/.gemini/antigravity-cli/) for settings and runtime files; global skills/agents under~/.gemini/config/(the machine-local discovery dir, #3738) or./.agent/local - Cursor:
~/.cursor/global or./.cursor/local - Windsurf/Devin Desktop:
~/.codeium/windsurf/global config or./.windsurf/local workflows - Augment Code:
~/.augment/global or./.augment/local - Trae:
~/.trae/global or./.trae/local - Qwen Code:
~/.qwen/global or./.qwen/local - Hermes Agent:
~/.hermes/global or./.hermes/local - CodeBuddy:
~/.codebuddy/global or./.codebuddy/local - Cline:
~/.cline/global or project-root.clineruleslocal
Project Files (.planning/)
.planning/
├── PROJECT.md # Project vision, constraints, decisions, evolution rules
├── REQUIREMENTS.md # Scoped requirements (v1/v2/out-of-scope)
├── ROADMAP.md # Phase breakdown with status tracking
├── STATE.md # Living memory: position, decisions, blockers, metrics
├── config.json # Workflow configuration
├── MILESTONES.md # Completed milestone archive
├── research/ # Domain research from /gsd-new-project
│ ├── SUMMARY.md
│ ├── STACK.md
│ ├── FEATURES.md
│ ├── ARCHITECTURE.md
│ └── PITFALLS.md
├── codebase/ # Brownfield mapping (from /gsd-map-codebase or /gsd-onboard)
├── onboarding/ # Brownfield onboarding summary (from /gsd-onboard)
│ ├── STACK.md # YAML frontmatter carries `last_mapped_commit`
│ ├── ARCHITECTURE.md # for the post-execute drift gate (#2003)
│ ├── CONVENTIONS.md
│ ├── CONCERNS.md
│ ├── STRUCTURE.md
│ ├── TESTING.md
│ └── INTEGRATIONS.md
├── phases/
│ └── XX-phase-name/
│ ├── XX-CONTEXT.md # User preferences (from discuss-phase)
│ ├── XX-RESEARCH.md # Ecosystem research (from plan-phase)
│ ├── XX-YY-PLAN.md # Execution plans
│ ├── XX-YY-SUMMARY.md # Execution outcomes
│ ├── XX-VERIFICATION.md # Post-execution verification
│ ├── XX-VALIDATION.md # Nyquist test coverage mapping
│ ├── XX-UI-SPEC.md # UI design contract (from ui-phase)
│ ├── XX-UI-REVIEW.md # Visual audit scores (from ui-review)
│ └── XX-UAT.md # User acceptance test results
├── quick/ # Quick task tracking
│ └── YYMMDD-xxx-slug/
│ ├── PLAN.md
│ └── SUMMARY.md
├── todos/
│ ├── pending/ # Captured ideas
│ └── completed/ # Completed todos
├── threads/ # Persistent context threads (from /gsd-thread)
├── seeds/ # Forward-looking ideas (from /gsd-capture --seed)
├── debug/ # Active debug sessions
│ ├── *.md # Active sessions
│ ├── resolved/ # Archived sessions
│ └── knowledge-base.md # Persistent debug learnings
├── ui-reviews/ # Screenshots from /gsd-ui-review (gitignored)
└── continue-here.md # Context handoff (from pause-work)
Post-Execute Codebase Drift Gate (#2003)
After the last wave of /gsd-execute-phase commits, the workflow runs a
non-blocking codebase_drift_gate step (between schema_drift_gate and
verify_phase_goal). It compares the diff last_mapped_commit..HEAD
against .planning/codebase/STRUCTURE.md and counts four kinds of
structural elements:
- New directories outside mapped paths
- New barrel exports at
(packages|apps)/<name>/src/index.* - New migration files
- New route modules under
routes/orapi/
If the count meets workflow.drift_threshold (default 3), the gate either
warns (default) with the suggested /gsd-map-codebase --paths … command,
or auto-remaps (workflow.drift_action = auto-remap) by spawning
gsd-codebase-mapper scoped to the affected paths. Any error in detection
or remap is logged and the phase continues — drift detection cannot fail
verification.
last_mapped_commit lives in YAML frontmatter at the top of each
.planning/codebase/*.md file; bin/lib/drift.cjs provides
readMappedCommit and writeMappedCommit round-trip helpers.
The baseline is written by gsd-tools stamp-codebase-map, a shell step in the
map-codebase workflow, not by the mapper agent. The mapper's own freshness
markers (**Analysis Date:**, <!-- refreshed: ... -->) are restamped
unconditionally on an Update run, so an agent that rewrites only the dates still
looks current to a reader; the machine-readable stamp is the one marker that
cannot be satisfied by a date-only rewrite, which is exactly why it is not the
agent's to write. --files a.md,b.md narrows the stamp to the documents a
caller actually refreshed, as the auto-remap path does.
An absent or unresolvable baseline is reported as skipped with reason
no-mapped-commit or unresolvable-mapped-commit, never as drift. Diffing
HEAD against the empty tree would report every tracked file as newly added,
which makes a stale map indistinguishable from a fresh one. Files under
.planning/ are excluded from the diff: the map's own commit is a planning
artifact, not codebase structure.
Installer Architecture
The installer (bin/install.js, ~10,700 lines) handles:
- Runtime detection — Interactive prompt or CLI flags (
--claude,--opencode,--kimi,--kilo,--codex,--copilot,--antigravity,--cursor,--windsurf,--augment,--trae,--qwen,--hermes,--codebuddy,--cline,--all) - Location selection — Global (
--global) or local (--local) - File deployment — Copies commands, skills, workflows, references, templates, agents, and hooks
- Runtime adaptation — Transforms file content per runtime:
- Claude Code: Uses as-is
- OpenCode: Converts commands/agents to OpenCode-compatible flat command + subagent format
- Kilo: Reuses the OpenCode conversion pipeline with Kilo config paths
- Codex: Generates TOML config + skills from commands
- Kimi CLI: Generates Agent Skills under
skills/gsd-*/SKILL.md, custom agent YAML/prompt files, and explicitkimi_cli.tools.*module paths - Copilot: Maps tool names (Read→read, Bash→execute, etc.)
- Antigravity: Skills-first with Google model equivalents; adjusts hook event names (
AfterToolinstead ofPostToolUse) - Cursor: Skills-first with Cursor rule references
- Windsurf: Skills-first with Windsurf rule references
- Trae: Skills-first install to
~/.trae/./.traewith nosettings.jsonor hook integration - Qwen Code: Skills-first with Qwen-branded path and prompt rewrites
- Hermes Agent: Category-based skills under
skills/gsd/ - CodeBuddy: Skills-first with CodeBuddy path and prompt rewrites
- Cline: Writes
.clinerulesfor rule-based integration - Augment Code: Skills-first with full skill conversion and config management
- Path normalization — Replaces
~/.claude/paths with runtime-specific paths - Settings integration — Registers hooks in runtime's
settings.json - Patch backup — Since v1.17, backs up locally modified files to
gsd-local-patches/for/gsd-update --reapply - Manifest tracking — Writes
gsd-file-manifest.jsonfor clean uninstall. The manifest also records whichruntimeand whichscope(global/local) wrote it, under amanifestVersionschema field, so a reader can answer "which surfaces are installed, at which scopes" without inferring it from the directory the file sits in (ADR 2866, #2872). Manifests written before that carry no such fields and are read without error — no reinstall is required. See Installer Migrations → File Manifest - Uninstall mode —
--uninstallremoves all GSD files, hooks, and settings
installRuntimeArtifacts (install-engine.cjs) returns the executed plan it ran — per kind, per
scope, including on the combined OpenCode/Kilo family path, which previously early-returned void —
rather than being observable only by re-reading disk afterward. Its destination-writing IO (copies,
removals, snapshot/restore, best-effort cleanup) now routes through an injectable fs seam,
install-fs-adapter.cjs, so a full install can be exercised against a fake adapter with zero real
destination IO; locating this package's own source tree remains real by design (a destination-fake
is never seeded with the repo's own paths). Writes stay byte-identical and existing void-ignoring
callers are unaffected. This completes ADR 58's
registry → adapter → helpers → cleanup rollout — the cleanup step had not previously landed
(#2874, epic #2866 Phase 5).
Install-time file moves, stale-artifact cleanup, config rewrites, and user-data preservation are governed by the Installer Migration Module. See Installer Migrations and ADR 0008. The migration module also owns the gated first-time baseline scan for legacy installs, classifying known runtime install surfaces before later migrations remove or rewrite anything.
The plan drift guard (plan_review.source_grounding) — which verifies symbol references in generated plans against live source before execution — is specified in ADR 22.
The same switch gates a second, cross-artifact axis: a fact-drift pass that compares the same fact as stated in ROADMAP.md, PLAN.md, STATE.md and CONTEXT.md and reports contradictions (a phase status, a success criterion, a requirement ID, a glossary term) with both locations and the authoritative side named. Where the source-grounding axis grounds a plan against code, this one grounds the planning artifacts against each other. It keys on contradicting knowledge rather than similar-looking text, and is advisory only — it never sets hardBlock and never contributes to the convergence counts.
Platform Handling
- Windows:
windowsHideon child processes, EPERM/EACCES protection on protected directories, path separator normalization - WSL: Detects Windows Node.js running on WSL and warns about path mismatches
- Docker/CI: Supports
CLAUDE_CONFIG_DIRenv var for custom config directory locations
Hook System
Architecture
Runtime Engine (Claude Code / Antigravity CLI)
│
├── statusLine event ──► gsd-statusline.js
│ Reads: stdin (session JSON)
│ Writes: stdout (formatted status), /tmp/claude-ctx-{session}.json (bridge)
│
├── PostToolUse/AfterTool event ──► gsd-context-monitor.js
│ Reads: stdin (tool event JSON), /tmp/claude-ctx-{session}.json (bridge)
│ Writes: stdout (hookSpecificOutput with additionalContext warning)
│
└── SessionStart event
├──► gsd-ensure-canonical-path.js (runs first)
│ Reads: ${CLAUDE_PLUGIN_ROOT}/gsd-core/ (plugin installs only)
│ Writes: ~/.claude/gsd-core/{bin,contexts,references,templates,workflows} symlinks
│ (no-op in classic installs; preserves user files; self-heals)
└──► gsd-check-update.js
Reads: VERSION file
Writes: ~/.claude/cache/gsd-update-check.json (spawns background process)
Context Monitor Thresholds
| Remaining Context | Level | Agent Behavior |
|---|---|---|
| > 35% | Normal | No warning injected |
| ≤ 35% | WARNING | "Avoid starting new complex work" |
| ≤ 25% | CRITICAL | "Context nearly exhausted, inform user" |
The two fire-points are defaults. hooks.context_warning_threshold and
hooks.context_critical_threshold in .planning/config.json move them per
project; see context-monitor.md for the resolution and
fallback rules.
Debounce: 5 tool uses between repeated warnings. Severity escalation (WARNING→CRITICAL) bypasses debounce.
Safety Properties
- All hooks wrap in try/catch, exit silently on error
- stdin timeout guard (3s) prevents hanging on pipe issues
- Stale metrics (>60s old) are ignored
- Missing bridge files handled gracefully (subagents, fresh sessions)
- Context monitor is advisory — never issues imperative commands that override user preferences
Package Legitimacy Gate (v1.42.1)
The researcher → planner → executor pipeline includes a supply-chain gate against slopsquatting (AI-hallucinated package names pre-registered with malicious post-install scripts).
Threat model: GSD automates the full path from "researcher names a package" to "executor runs npm install". A hallucinated name that passes npm view (proving only registration, not legitimacy) would previously flow through undetected. ~20% of AI-generated package references are hallucinated; ~43% of those names recur consistently across prompts, making pre-registration economically viable for attackers.
Gate layers:
| Layer | Component | Action |
|---|---|---|
| Research | gsd-phase-researcher |
Runs gsd-tools query package-legitimacy check --ecosystem <npm|pypi|crates> <pkgs>; writes ## Package Legitimacy Audit table to RESEARCH.md; strips [SLOP] packages before RESEARCH.md is written |
| Planning | gsd-planner |
Reads Audit table; inserts checkpoint:human-verify before any [ASSUMED] or [SUS] install task; adds T-{phase}-SC STRIDE supply-chain row to <threat_model> |
| Execution | gsd-executor |
RULE 3 excludes package installation from auto-fix scope; failed installs surface as checkpoints, never silent substitutions |
Claim provenance integration: Package names discovered via WebSearch are tagged [ASSUMED] (not [VERIFIED]) regardless of the registry-API verdict. This extends the existing [ASSUMED] / [VERIFIED] / [CITED] provenance system by enforcing the provenance tag as a hard gate at the install boundary — [ASSUMED] always generates a checkpoint:human-verify in PLAN.md.
Ecosystem coverage: The gate resolves signals directly from each ecosystem's registry API rather than a single generic check — registry.npmjs.org + api.npmjs.org/downloads (Node), pypi.org/pypi/<pkg>/json (Python), the crates.io API (Rust). This catches cross-ecosystem hallucination (~9% rate documented in 2025 USENIX research).
Graceful degradation: Each registry adapter degrades to null signals (never throws) on a failed lookup; missing signals push a package to [SUS], which is gated behind the same checkpoint:human-verify checkpoint as [ASSUMED]. Research and planning proceed; the system never hard-fails on a network or tool outage. slopcheck is an optional escalate-only adapter — it can only raise a verdict, never lower it, and is not the install-or-degrade gate. No shipped configuration wires it.
Security Hooks (v1.27)
For a conceptual overview of how the hook and guard layers fit into the broader security approach, see Security model.
Prompt Guard (gsd-prompt-guard.js):
- Triggers on Write/Edit to
.planning/files - Scans content for prompt injection patterns (role override, instruction bypass, system tag injection)
- Advisory-only — logs detection, does not block
- Patterns are inlined (subset of
security.cjs) for hook independence - Output contract:
hookSpecificOutputcarries bothadditionalContextandfindings— an array of{ ruleId, match }records (INJECTION-PATTERNorINVISIBLE-UNICODE), module-local to this hook (not shared withgsd-read-injection-scanner.js's ownRULE_IDS). The advisory is rendered fromfindingsvia a single mapper, so the two cannot disagree. Consumers should readfindingsrather than parsing the advisory text.
Read Injection Scanner (gsd-read-injection-scanner.js):
- Triggers on
Read/WebFetch/WebSearchPostToolUse events - Advisory by default; blocks only
HIGHseverity, and only whensecurity.injection_blockingistrue - Severity is
LOWfor 1-2 matched patterns,HIGHfor 3 or more - Skips content shorter than 20 characters, and skips excluded paths (
.planning/,REVIEW.md,CHECKPOINT*, security/injection docs, and GSD's own staged hook bundle) - Rule ids: the
MD-LINK-*markdown-link rules mirrored fromsecurity.cjs'sMARKDOWN_LINK_PATTERNS, plusINJECTION-PATTERN,INVISIBLE-UNICODE, andUNICODE-TAG-BLOCK - Patterns are shared with
gsd-prompt-guard.jsviahooks/lib/injection-patterns.js(#3504); the markdown-link list is inlined for hook independence - Output contract:
hookSpecificOutputcarriesadditionalContext(the human-readable advisory sentence),findings— an array of{ ruleId, match }records naming each rule that fired — plusseverity(LOWfor 1-2 matches,HIGHfor 3+) andsource(the scanned file path, URL, orsearch: <query>string).findingsandseverityare the structured surface; the advisory is rendered from them, so the three cannot disagree.matchisnullfor rules with no captured text (INVISIBLE-UNICODE,UNICODE-TAG-BLOCK). Consumers should readfindings/severity/sourcerather than parsing the advisory text.
Workflow Guard (gsd-workflow-guard.js):
- Triggers on Write/Edit to non-
.planning/files - Detects edits outside GSD workflow context (no active
/gsd-command or Task subagent) - Advises using
/gsd-quickor/gsd-fastfor state-tracked changes - Opt-in via
hooks.workflow_guard: true(default: false) - Output contract: the advisory leg's
hookSpecificOutputcarriescode: 'WORKFLOW_ADVISORY'alongsideadditionalContext. This is distinct from the hook's separate force-add block leg (code: 'WORKTREE_AGENT_FORCE_ADD_FORBIDDEN',decision: 'block') — the two are disambiguated bycode, never by presence.
Runtime Abstraction
GSD supports multiple AI coding runtimes through a unified command/workflow architecture:
Runtime Install Contract Matrix
This matrix describes the runtime surfaces the installer materializes today. The migration-specific ownership and source snapshots live in Installer Migrations.
| Runtime | Global root | Local root | Invocation surface | Agent surface | Config and hooks |
|---|---|---|---|---|---|
| Claude Code | ~/.claude |
./.claude |
Global skills/gsd-*/SKILL.md (flat, #924); local commands/gsd/*.md |
agents/gsd-*.md |
settings.json hook and statusLine entries |
| OpenCode | ~/.config/opencode |
./.opencode |
commands/gsd-*.md |
agents/gsd-*.md |
opencode.json or opencode.jsonc; no GSD hooks |
| Kilo | ~/.config/kilo |
./.kilo |
command/gsd-*.md |
agents/gsd-*.md |
kilo.json or kilo.jsonc; no GSD hooks |
| Kimi CLI | First-existing generic root: ~/.config/agents recommended, then ~/.agents when ~/.agents/skills exists and ~/.config/agents/skills does not |
Deferred and guarded | skills/gsd-*/SKILL.md (flat) invoked as /skill:gsd-* |
agents/gsd.yaml, agents/gsd.md, and agents/subagents/gsd-* YAML/prompt pairs |
Explicit kimi --agent-file <configRoot>/agents/gsd.yaml; no GSD hooks or statusline |
| Codex | ~/.codex |
./.codex |
skills/gsd-*/SKILL.md (flat) |
agents/ source markdown plus per-agent TOML (Codex auto-discovers each agents/gsd-*.toml; this is the sole canonical role registration, #2406) |
config.toml bare [agents] dispatch-tuning scalar (max_depth, no per-role [agents.gsd-*] tables), [features].hooks (canonical; legacy alias codex_hooks is recognized and migrated forward on reinstall, #3566), and hook tables |
| GitHub Copilot | ~/.copilot |
./.github |
skills/gsd-*/SKILL.md (flat), copilot-instructions.md, and AGENTS.md (repo root, local) |
.agent.md files |
Self-contained sessionStart hook (hooks/gsd-session.json, inline command type); no statusline |
| Antigravity | auto-detected: ~/.gemini/antigravity, ~/.gemini/antigravity-ide, or ~/.gemini/antigravity-cli |
./.agent |
~/.gemini/config/skills/gsd-*/SKILL.md (flat, #1614; global home override #3738) |
~/.gemini/config/agents/gsd-*.md (#3738) |
Gemini-style settings.json hook entries when installed by GSD |
| Cursor | ~/.cursor |
./.cursor |
skills/gsd-*/SKILL.md (flat) |
agents/gsd-*.md |
Rule references under rules/; hooks.json with sessionStart context injection and postToolUse STATE.md monitor (#777) |
| Windsurf | ~/.codeium/windsurf config |
./.windsurf |
workflows/gsd-*.md slash-command workflows |
No custom-agent artifact surface | No GSD hooks |
| Augment Code | ~/.augment |
./.augment |
skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) |
agents/gsd-*.md |
No GSD hooks or statusline |
| Trae | ~/.trae |
./.trae |
skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) |
agents/gsd-*.md |
Rule references under rules/; no GSD hooks |
| Qwen Code | ~/.qwen |
./.qwen |
skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) |
agents/gsd-*.md |
Common GSD settings and hook entries where supported |
| Hermes Agent | ~/.hermes |
./.hermes |
skills/gsd/ns-*/SKILL.md (6 routers, prefix='') + skills/gsd/ns-*/skills/<name>/SKILL.md (nested concretes) |
agents/gsd-*.md |
Common GSD settings and hook entries where supported |
| CodeBuddy | ~/.codebuddy |
./.codebuddy |
skills/gsd-*/SKILL.md (flat, user-invocable: false) |
agents/gsd-*.md |
/gsd-* slash commands under commands/; common GSD settings and hook entries where supported |
| Cline | ~/.cline |
project root | skills/gsd-ns-*/SKILL.md (6 routers) + skills/gsd-ns-*/skills/<name>/SKILL.md (nested concretes) + .clinerules |
Rules only | No GSD hooks or statusline |
Upstream Contract Sources
Runtime install expectations are checked against primary documentation where available. The current source snapshot is 2026-05-11, with Kimi CLI rechecked on 2026-06-07:
- Claude Code: Anthropic slash commands, settings, hooks, and subagents docs.
- OpenCode and Kilo: OpenCode config docs and Kilo custom subagent docs.
- Qwen Code: command/config docs; Qwen command docs were last updated 2026-05-06.
- Kimi CLI: Agent Skills docs for user-level brand roots and first-existing
generic roots (
~/.config/agents/skills/recommended, then~/.agents/skills/), plus Agents docs for YAML files,system_prompt_path,kimi_cli.tools.*module paths, and explicitkimi --agent-filelaunch. - Codex: OpenAI Codex docs and
config-schema.json; the installer also carries Codex 0.124.0 compatibility for agent table shape. - Copilot, Cursor, Cline, Augment, Hermes, and CodeBuddy: vendor docs for custom instructions, rules, skills, or config.
- Antigravity, Windsurf, and Trae: source-limited rows. The installer documents current compatibility shims, and migrations must refresh those sources before rewriting their config.
Abstraction Points
- Tool name mapping — Each runtime has its own tool names (e.g., Claude's
Bash→ Copilot'sexecute) - Hook event names — Claude uses
PostToolUse, Antigravity usesAfterTool - Agent frontmatter — Each runtime has its own agent definition format
- Path conventions — Each runtime stores config in different directories
- Model references —
inheritprofile lets GSD defer to runtime's model selection
The installer handles all translation at install time. Workflows and agents are written in Claude Code's native format and transformed during deployment.