Files
msd-core/docs/ARCHITECTURE.md
Dennis Alexis Valin Dittrich ad1477d659 enhance(#4154): validate configured entrypoints before reporting install success (#4249)
* test(260903-m7p): expose configured-entrypoint validation gap

* enhance(260903-m7p): validate configured entrypoints before success

* test(260903-m7p): require pre-success entrypoint validation

* enhance(260903-m7p): gate install success on entrypoints

* test(260903-m7p): cover configured entrypoints across runtimes

* enhance(260903-m7p): cover emitted runtime entrypoints

* fix(260903-m7p): sandbox HOME in finishInstall test and fix changeset pr number

- finishInstall(...'cline'...) calls writeNonClaudeDefaults(runtime) in-process
  before the new configured-entrypoint assertion throws. Without a HOME +
  config-location-env sandbox that write resolved through the ambient
  environment and landed in the developer's live ~/.gsd (confirmed absent on
  origin/next baseline, present only on this branch — full-suite HERMETICITY
  WARNING). Sandbox HOME/USERPROFILE and scrub config-location env for the
  duration of the test, matching the existing in-process finishInstall/
  install() pattern in tests/install.test.cjs (#2665).
- .changeset/quick-wasps-sing.md: pr: 0 is a never-backfilled placeholder
  (CONTRIBUTING.md) that fails changeset-lint's invalid_pr check; set to the
  fork PR number until the upstream PR number is known.

* fix(260903-m7p): repair cross-platform and pre-existing shape fallout

- tests/configured-entrypoint-validation.test.cjs: the win32 branch of
  ensureCodexHooksJsonSessionStart writes a .cmd shim under
  <codexRoot>/hooks/; create that dir in the test (the real installer only
  calls this once hooks/gsd-check-update.js already exists) and assert the
  platform-common entrypoint shape instead of a fixed non-Windows array,
  since win32 legitimately emits two entries (cmd shim + script).
- tests/install.test.cjs: finishInstall's shared settings-json return now
  carries configuredEntrypoints/rollbackInstallerMigrations for every
  runtime on that path (trae included, not just Claude/Cursor/Windsurf);
  update the trae install() exact-shape assertion to match.

* fix(260903-m7p): keep .sh interpreter tracking consistent with unresolved bash

configuredEntrypointsForHook's shell branch dropped interpreterCandidates
entirely when resolveBashExecutable returned null, unlike the sibling
portableHooks runner entry a few lines below (which correctly falls back to
the literal 'bash' token). Found via agy adversarial review; verified
unreachable through the current call graph (buildHookCommand's own
resolveBashRunner==null gate already short-circuits before
recordConfiguredHookCommand runs), so this is a defensive consistency fix,
not a live-bug patch — kept for the next caller that does not share that
gate.

* chore(260903-m7p): backfill changeset pr number to the opened upstream PR

.changeset/quick-wasps-sing.md carried the fork PR number (16) as a
placeholder until the upstream PR existed; open-gsd/gsd-core#4249 is now
open, so record its real number per CONTRIBUTING.md's changeset pr-field
convention.

* fix(#4154): track already-registered hooks for entrypoint validation on update

applySettingsJsonHooks registers each guard hook only if absent, so a hook
already present from a prior install keeps its stale on-disk command. The
new entrypoint tracker always records the freshly-computed command for it,
which never matches what is actually persisted, so the exact-string filter
in finishInstall silently dropped it from validation — the Blocker case
this feature exists to catch (an already-installed entrypoint going stale
between installs) was exactly the case it never validated.

Match on the managed script's basename instead, which the persisted
command carries either way, so an already-registered hook stays in the
validated set. Regression test forces this path by mutating a
freshly-installed hook's persisted command before a second install.

* fix(#4154): distinguish an unreadable script from a missing one

validateConfiguredEntrypoints folded an EACCES statSync failure into the
same 'missing' reason as ENOENT, misreporting a real permission problem as
an absent file. Check the error code and report 'unreadable' instead.

* docs(#4154): document entrypoint validation's rollback and PATH scope

CONTEXT.md's Runtime Hooks Surface Module / Installer Module entries had
no mention of ConfiguredEntrypoint/validateConfiguredEntrypoints, despite
bin/install.js x CONTEXT.md being this repo's strongest co-change pairing.

The update-gsd.md how-to overstated what a validation failure undoes: for
Codex/Cursor/Windsurf/Kimi, their own writer already persisted hooks.json/
config.toml inside install() before the aggregate validation call runs, so
there is no rollback path for that write regardless of "where available"
phrasing. Also note that interpreter resolution checks the installer's own
PATH, not necessarily the PATH a hook fires under later (#2979 launchers).

* chore(#4154): point changeset pr field at the fork PR while CI runs there

Mirrors the branch's own prior backfill commit: pr: matches whichever PR
number changeset-lint is currently validating against (fork PR #16 during
the fork-first CI/review loop), flipped back to the upstream PR number
right before the final push to open-gsd/gsd-core.

* fix(#4249): address adversarial-review findings in entrypoint validation

An internal adversarial review (agy/gemini-3.8-flash-high) of the whole PR
found several real gaps beyond the human reviewer's Blocker, verified
against source before fixing:

- Codex's install() result bound rollbackInstallerMigrations to the narrow
  installer-migrations-only rollback instead of restoreCodexSnapshot (#3245),
  the full pre-install snapshot/restore Codex already owns for exactly this
  case — a validation failure discovered outside install() reverted nothing
  of the config.toml/hooks.json that call had already written.
- The register-only-if-absent basename match from the prior fix used a bare
  substring, which an unrelated user command mentioning the same filename
  could false-positive into GSD's validated set — anchored on the
  `/hooks/<basename>` path segment instead.
- nodeCandidates checked raw process.execPath (always true — we're running
  in that process) instead of normalizeNodePath's stable version-manager
  alias, the same one buildNodeRunnerChainToken bakes as its first choice —
  a false green regardless of whether that alias itself still resolves.
- An entry with no interpreterCandidates (Cline's PreToolUse hook, or a
  Windows-Claude .sh hook invoked without a bash runner) runs via its own
  shebang; validateConfiguredEntrypoints checked only file-type, never the
  execute bit. Cline's writer also never reported an entrypoint at all.
- Duplicate (configPath, scriptPath) entries (e.g. Kimi's context-monitor
  hook registered across several events) were validated once per duplicate.

Each fix is covered by a new or extended test; the Codex one required
inlining runCodexInstall's env sandboxing so the rollback closure — which
re-resolves the $HOME-relative skills root live — runs before the sandbox
is torn down, matching how installAllRuntimes' real aggregate gate calls it.

* docs(#4249): document the round-2 entrypoint-validation fixes

Runtime Hooks Surface Module and Installer Module entries now name
ConfiguredEntrypoint's not-executable reason, the normalizeNodePath
alignment, Cline's tracked hook, and which install() result the
finishInstall/installAllRuntimes rollback path actually reverts per
runtime (Codex's full snapshot vs. the others' narrow migrations-only
rollback).

* chore(#4249): point changeset pr field at the upstream PR now that fork CI is green

* fix(#4249): address agy adversarial-review findings

- validateConfiguredEntrypoints: statSync alone never detects a
  chmod-000 script (it only needs parent-dir search permission), so an
  interpreter-invoked entry with an unreadable script passed validation.
  Add an explicit R_OK check for the interpreterCandidates branch only —
  the candidate-less/shebang branch already has its own X_OK gate.
- docs/how-to/update-gsd.md: the blanket "does not revert" claim was
  false for Codex, which reverts config.toml/hooks.json via its full
  pre-install snapshot; qualify it per runtime.
- tests/codex-config.test.cjs: the #4249 rollback regression test
  asserted skills/ and VERSION were reverted but never asserted
  config.toml/hooks.json were too, despite the test's own stated intent.
- CONTEXT.md: qualify which interpreterCandidates entries get
  normalizeNodePath'd (Node hooks only, not .sh/bash) and note Codex's
  Windows .cmd shim as a third candidate-less case that relies on
  extension dispatch, not a shebang.

* fix(#4249): validate Cline's PATH-dependent interpreter, not just its execute bit

Cline's hook is a hybrid: it self-executes via '#!/usr/bin/env node',
so it needs the execute bit (like any shebang-invoked entry), but its
interpreter is looked up on PATH by 'env' at hook-fire time (unlike
every other GSD JS hook, which bakes an absolute node path specifically
to avoid that dependency). The candidate-less/interpreterCandidates
fork treated these as mutually exclusive, so Cline's entry silently
skipped interpreter resolution entirely — a completely missing 'node'
on PATH would still validate successfully.

Add an orthogonal selfExecutable flag so both checks run for entries
that need them. (CodeRabbit finding on the fork rehearsal PR.)

* fix(#4249): address second-round adversarial review findings (opus + agy)

- validateConfiguredEntrypoints: R_OK now runs for every scriptOk entry,
  not just interpreterCandidates ones — a self-executable shebang script
  is still opened and read by its kernel-invoked interpreter, so X_OK
  alone never proved it was readable.
- selfExecutable is now the sole, explicit source of truth for the
  execute-bit check (every producer that needs it sets the flag) instead
  of being partly inferred from an absent interpreterCandidates, which
  Cline's hybrid entry also carries.
- The execute-bit check now skips explicitly on win32 (matching
  resolveExecutableBinary's own carve-out) instead of relying on Node's
  accessSync(X_OK)-as-F_OK no-op, which only protects a real Windows
  machine and not a test that simulates win32 on a POSIX runner.
- bin/install.js: fixed a stale comment claiming no runtime's
  install()-time writes have a rollback path — Codex's does
  (restoreCodexSnapshot) — and added the omitted Cline to both that
  comment and CONTEXT.md's equivalent lists.
- CONTEXT.md: fixed the Cline description left stale by the previous
  commit's selfExecutable addition, and rewrote the validation-mechanism
  paragraph for clarity (writing-for-agents pass).
- docs/how-to/update-gsd.md: split an overloaded 4-clause sentence.
- Removed a fault-injection integration test that could not reliably
  exercise the real installAllRuntimes -> finalize -> rollback wiring
  without fighting the installer's own pre-registration existence
  guards; the constituent pieces remain covered individually.

* fix(#4249): pin platform in X_OK-testing entries so they're deterministic cross-CI-runner

X_OK is a POSIX-only concept, skipped entirely when an entry's platform
is win32 (matching production). Two test entries omitted platform,
defaulting to process.platform — on an actual windows-latest CI runner
that silently skipped the very check they were meant to exercise,
turning 'not-executable' into a false pass. Pin platform: 'linux' so
these are deterministic regardless of which OS runs the suite.

* fix(#4249): classify EPERM the same as EACCES in statSync error handling

Windows raises EPERM (not EACCES) for a parent directory that couldn't
be traversed into — was falling through to 'missing', misreporting a
genuine permission problem as a nonexistent path.

* docs(#4249): address final CodeRabbit doc-completeness findings

- CONTEXT.md: install()'s documented result shape omitted
  configuredEntrypoints; the ConfiguredEntrypoint shape omitted
  selfExecutable.
- docs/how-to/update-gsd.md: the failure-mode sentence omitted
  unreadable and lacks-execute-permission, which the installer also
  rejects.

* fix(#4249): stop double-validating every configured entrypoint on install/update

installAllRuntimes' finalize() already runs assertConfiguredEntrypoints
once over the aggregate set; finishInstall then re-ran the identical
check per runtime in the printSummaries loop right after, so every
entrypoint paid its statSync/accessSync/interpreter-resolution cost
twice on every install and update. Add entrypointsAlreadyValidated to
skip the redundant pass specifically on that path, while leaving the
check intact for any caller that invokes finishInstall directly.

* chore(#4154): point changeset pr field at rehearsal fork PR while CI runs there

* perf(#4249): memoize interpreter candidate resolution across entrypoints

resolveExecutableBinary walked PATH once per (entry, candidate) pair; a
typical install has a dozen-plus entries sharing the same few candidate
lists (process.execPath for JS hooks, bash for shell hooks). Cache by
(platform, candidate) so each distinct pair resolves once per validation
call instead of once per entry.

* chore(#4249): point changeset pr field at the rebased rehearsal fork PR

* fix(#4249): drop entrypoint tracking from the now-dead Codex event writer

#2586 (landed on next after this branch forked) removed install.js's
CODEX_EXTENDED_HOOK_EVENTS registration loop, so ensureCodexHooksJsonEvent
no longer runs during install or update. The ConfiguredEntrypoint records
this branch added inside it were therefore unreachable and untested. Restore
the function to its upstream shape; the entrypoints it used to report were
never collected by any caller.

* refactor(#4249): drop the revalidation bypass flag and the candidate cache

Both were this PR's own micro-optimisations over a set of roughly a dozen
entries. `entrypointsAlreadyValidated` let a caller turn the finishInstall
gate off to save one statSync/accessSync pass; `resolvedCandidateCache`
memoised resolveExecutableBinary across entries that are already deduped by
(configPath, scriptPath). Neither is measurable, and the flag was the only
way to reach finishInstall with validation disabled. finishInstall now
always validates what it is given.

* chore(#4249): point the changeset pr field back at the upstream PR

* refactor(#4249): track settings.json entrypoints without the hooksSurface gate

The install-surface writer only tracked configured entrypoints when the
runtime's descriptor also declared `hooksSurface: 'settings-json'`. Nothing
asserts that axis agrees with `installSurface`, so a descriptor that broke
the coupling would silently pass `configuredEntrypoints: undefined` and drop
that runtime out of the validation this PR adds — reintroducing the exact
'reports Done! over a broken entrypoint' failure #4154 exists to close.

Remove the dependence rather than test it: everything recorded on this path
lands in settings.json by construction, and the registered-command filter
already discards entries no persisted hook references.

* chore(#4249): put the changeset body in the documented two-part format

CONTRIBUTING.md and .changeset/README.md both show
`**<bold change>** — <symptom-led explanation>.`; the fragment was a single
unbolded sentence.

* chore(#4249): point the changeset pr field at the rehearsal fork PR while CI runs there

* fix(#4249): restore the whole manifest-tracked GSD file set on Codex rollback

#3245's snapshot covers config.toml, hooks.json, skills/gsd-*, agents/gsd-*
and gsd-core/VERSION. The install overwrites every other GSD-owned file too —
hooks/, gsd-core/CHANGELOG.md, scripts/, gsd-core/.gsd-runtime, the manifest
itself — before the entrypoint-validation gate runs, so a validation failure
left the new payload sitting on top of the restored old config.

Snapshot the file set the PREVIOUS install's gsd-file-manifest.json claims,
before runInstallerMigrations so the bytes are the true pre-install state, and
restore it from both Codex rollback closures ahead of the per-surface restores.
Files only the failed install introduced are removed, read from the manifest
now on disk. The manifest is already the authoritative record of what GSD owns,
so no second hand-written list can drift out of sync, and user-owned files are
never snapshotted or removed. Every path is confined through
resolveInstallRelativePath, so a hand-edited manifest cannot turn rollback into
an arbitrary-path write.

Non-Codex runtimes are unaffected: the snapshot is gated on the same
tomlConfigInstall + non-minimal condition as #3245's.

* fix(#4249): keep the managed-file snapshot honest in minimal mode and on a bad manifest

Two follow-on defects in the previous commit's snapshot:

- The capture was gated on `!isMinimalMode`, copied from #3245. A core/
  --minimal Codex install still writes gsd-core/, hooks/, scripts/ and the
  manifest, and restoreCodexSnapshot is reachable in that mode (#2695), so the
  snapshot came back empty while the rollback still ran — and its removal pass
  would have deleted every file the new manifest lists. Gate on
  tomlConfigInstall alone, matching where the rollback actually reaches.

- An unreadable or unparseable prior manifest was caught alongside ENOENT and
  treated as a fresh install. That is the same empty-snapshot state, so a failed
  update over a real install with a corrupt manifest could delete its prior
  payload. Track whether the pre-install GSD-owned set is KNOWN: ENOENT means
  known-empty; any other read error or a parse failure means unknown, and the
  restore closure returns without touching anything, degrading to #3245's
  narrower rollback. Deliberately not fatal — a corrupt manifest has to stay
  repairable by reinstalling over it.

Both paths are covered by red-checked regression tests.

* fix(#4249): snapshot Codex skills, agents and VERSION in minimal mode too

commit removed from the manifest snapshot. restoreCodexSnapshot is reachable
for a core/--minimal install (#2695), and its pass-2 sweeps remove every gsd-*
skill dir and gsd-* agent file the snapshot does not claim — so with an empty
minimal-mode snapshot a rollback deleted the whole skills/agents surface with
nothing to restore it from. Codex resolves skills to $HOME/.agents/skills via
the ADR-1239 skills-kind home override, so this is also the reason manifest
`skills/` keys do not resolve under configDir: that surface belongs to this
snapshot, not to the manifest-driven one.

Gate on tomlConfigInstall alone. _codexPreConfigRollback stays null in minimal
mode — doing nothing on an early failure is the non-destructive side.

Covered by a red-checked regression test that plants bytes in an alternate-home
skill file, reinstalls under the core profile marker, and asserts the rollback
restores it.

* fix(#4249): never remove on rollback unless a prior manifest proves what predates the install

Three defects in the manifest-driven Codex rollback, all in its removal half:

- ENOENT marked the snapshot usable, arming the removal pass on a FIRST
  install. GSD may have overwritten a user's file at a manifest-tracked path
  there, and no prior manifest records the difference — so rollback deleted it
  where before it merely left it overwritten. Absent, unreadable and malformed
  manifests now all leave the prior set UNKNOWN and skip removal entirely.

- Membership was tested against the map of files whose pre-install read
  SUCCEEDED, so a tracked file that existed but was unreadable read as
  introduced-by-this-install and was removed. Track the prior manifest's paths
  in their own Set and test against that.

- The unreachable "delete the manifest when there was no prior one" branch is
  gone: usable now implies a parsed prior manifest.

Also adds the end-to-end test the aggregate gate was missing — the four Codex
rollback tests drove the closure directly, proving the restore but not the
wiring. installAllRuntimes(['codex','cline']) under an emptied PATH makes
Cline's `env node` entry fail validation for real, and asserts Codex's payload
comes back.

Test preamble (HOME/USERPROFILE sandbox + config-env scrub) is now one helper
instead of six copies. Both new tests are red-checked.

* test(#4249): use unlinkSync, not rmSync, to drop the manifest in a test

lint:ci's raw-fs.rmSync rule points tests at helpers.cleanup for its
Windows-EBUSY retry budget. That budget is for directory trees; this removes a
single file, which unlinkSync says more precisely and the rule does not flag.

* chore(#4249): point the changeset pr field back at the upstream PR

* fix(#4249): use an unambiguous dedup key and surface partial-restore failures

trek-e's 2026-09-08 adversarial pass flagged two findings in the new
entrypoint-validation/rollback code:
- assertConfiguredEntrypoints' dedup key already used a raw NUL
  separator (introduced in ceebb65f2d), but git/Read render NUL as a
  space, so the key looked like a plain-space join to every reviewer
  that read the diff. Replace it with JSON.stringify([configPath,
  scriptPath]) so the separator is visible and unambiguous.
- restoreManagedFileSnapshot's per-file restore catch block claimed to
  'surface the original error' but only swallowed it, matching (and
  widening) the pre-existing #3245 restoreCodexSnapshot pattern. Add
  an actual console.warn using the existing best-effort-warning
  convention, scoped to just this PR's new function.

* fix(#4249): treat a files-less prior manifest as unknown, not known-empty

agy's gemini-3.8-flash-high adversarial pass (round 5) found and I
reproduced empirically: a structurally-valid manifest missing the
files key (e.g. {"version":1}) parses without throwing, so
Object.keys(undefined || {}) silently read as 'zero files predate
this install' instead of the UNKNOWN state the malformed-manifest
guard exists to produce. Rollback's removal pass then deleted every
GSD-owned file the failed install's own manifest listed, including
ones that predated it — the exact data loss the #4249 CodeRabbit
malformed-manifest fix was supposed to prevent, reachable through a
JSON.parse success instead of a failure. Route the shapeless case
into the same catch-all UNKNOWN path via an explicit shape check.
Regression test reproduces the deletion before the fix and confirms
the file survives after it.

Also extend restoreManagedFileSnapshot's removal-pass rmSync and
final manifest-rewrite catches with the same real console.warn
trek-e's round-4 review asked for on the per-file restore catch —
same rollback function, same operator-facing-signal gap.

* docs(#4249): correct which runtimes actually leave a written config on rollback

agy's completeness audit (round 5, holistic pass) caught this new
paragraph claiming 'for every other runtime, the configuration file(s)
already written during that update are left in place' — false for
Claude Code and other settings.json-based runtimes, whose write never
happens on failure (assertConfiguredEntrypoints runs before
finishInstall's writeSettings). Only Cursor/Windsurf/Kimi/Cline
actually match that description, since they persist their config file
inside install() ahead of the gate. Split the one sentence into the
three actual outcomes; matches the PR body's own accurate Before/After
wording, which this doc addition had drifted from.

* fix(#4249): clean up doc/comment mismatches and dead fields from opus review

Opus critical-code-reviewer + ponytail-review pass on the final diff:
- assertConfiguredEntrypoints carried finishInstall's old docblock
  ("Apply statusline config, then print completion message") from
  before this function was inserted between comment and callee.
  finishInstall already has its own accurate #4249 comment, so the
  stale docblock is removed rather than moved.
- checked: number on ConfiguredEntrypointValidationResult and
  error.configuredEntrypointValidation on the thrown error: the first
  had zero consumers anywhere in the repo, including its own defining
  file, and is removed. The second matches an existing repo
  convention (bin/install.js's installerMigrationRollbackFailures,
  #4249 predates this PR) of attaching structured diagnostic context
  to a re-thrown Error even before a consumer exists, so it's kept.
- finishInstall's own assertConfiguredEntrypoints call is a redundant
  backstop on the real production path (installAllRuntimes's aggregate
  call already validates the superset first), but its comment read as
  though this call alone provided the before-the-write guarantee.
  Clarified rather than removed — it's the only gate for a caller that
  invokes finishInstall directly.

* chore(#4249): split the manifest-driven rollback engine out into #4544

Issue #4154 asked the installer to consume a validation failure "through
the existing rollback mechanism, without a second transaction mechanism".
The manifest-driven rollback widening added during review (capture every
path the prior gsd-file-manifest.json claims, restore those bytes, remove
what only the failed install introduced) is that second mechanism on a
plain reading. It is a real fix for a #3245-era gap, but an independent
one, so it moves to its own bug report and PR.

Removed here:
- bin/install.js: the pre-install managed-file capture block and
  restoreManagedFileSnapshot, plus its call sites in
  _codexPreConfigRollback and restoreCodexSnapshot (99 lines).
- tests/configured-entrypoint-validation.test.cjs: the five tests that
  exercise the manifest engine.
- CONTEXT.md and docs/how-to/update-gsd.md: the sentences describing the
  widened restore. update-gsd.md again documents the #3245 surfaces only.

Kept, because it is #4154's own scope:
- the entrypoint-validation gate itself;
- Codex's install() result binding rollbackInstallerMigrations to
  restoreCodexSnapshot (config.toml, hooks.json, skills/gsd-*,
  agents/gsd-*, gsd-core/VERSION);
- the !isMinimalMode gate removal on that snapshot. Binding the closure
  to the result made it reachable for a core/--minimal install, where its
  pass-2 sweeps delete every gsd-* skill dir and agent file the snapshot
  does not claim; an empty minimal-mode snapshot therefore deleted the
  whole surface with nothing to restore.

The surviving aggregate-failure test now asserts on config.toml, a
surface the #3245 snapshot owns, instead of gsd-core/CHANGELOG.md, which
only the manifest engine restored.

Refs #4544

* test(#4249): cover configured entrypoints through the packed install path

#4154's scope lists install smoke coverage alongside the installer gate —
"assert representative configured entrypoints resolve for supported runtime
profiles". The gate itself (assertConfiguredEntrypoints /
validateConfiguredEntrypoints) is unit-covered by in-process install() calls;
nothing proved the property survives npm pack -> npm install -g -> install.js.

Add Cycle 4 to runSmoke. For each of claude and codex — the two distinct
config surfaces GSD writes launch paths into (settings.json, and hooks.json +
config.toml) — run the tarball-installed installer into a throwaway HOME, then
re-read that runtime's own written config and return the new
ENTRYPOINT_UNRESOLVED code when a script path it names does not resolve to a
file. install-smoke.yml already asserts .code == "ok" on the CLI, so the check
becomes a release gate on every matrix host without workflow changes.

The scan re-derives paths from the written config instead of reusing the
installer's own entrypoint list, and test I shows why that matters: a
registration the installer never touched during a run is invisible to the
in-process gate, so the install exits 0 and only reading the config back off
disk catches the dangling launch path.

* ci(#4249): pack a publish-shaped tarball in the install smoke lane

`npm pack` runs prepack/prepare (build:lib); only prepublishOnly runs
build:hooks. hooks/dist is gitignored, so the tarball install-smoke.yml packs
after `npm ci` carries no hook scripts at all — the lane has been smoking a
package that differs from the published one in exactly the artifacts the
lifecycle smoke is supposed to launch.

That went unnoticed because the lane's init runs `--local`, which registers no
statusline and therefore registers no hook whose target is missing. A
`--global` install on the same tarball exits 1 on #4249's own gate
(`gsd-statusline.js (missing)`), which is what the new configured-entrypoint
cycle performs, so without this step the cycle would report INIT_FAILED
instead of checking anything.

Build hooks before packing so the smoked tarball matches prepublishOnly. The
CLI now reports 16 configured entrypoints for claude and 1 for codex instead
of zero.

* fix(#4249): scope Codex's full snapshot restore to entrypoint failures

Binding Codex's result to `restoreCodexSnapshot` made ANY finalize-stage
exception un-install a Codex install that had already succeeded and already
printed its own "Done!" summary — `rollbackFinalizedInstallerMigrations` wraps
the whole `finalize()` body, not just the aggregate `assertConfiguredEntrypoints`
call.

Nothing documents that. `docs/installer-migrations.md#phase-4-installupdate-integration`
scopes finalize-stage rollback to installer *migrations* ("the executor uses the
journal to restore modified paths"), and this PR's own operator-facing paragraph
in `docs/how-to/update-gsd.md` scopes the Codex config.toml/hooks.json/skills/
agents/VERSION revert to entrypoint-validation failures specifically ("If a
script is missing, unreadable, ... For Codex, this reverts ..."). The wide
behaviour is also incoherent as a transaction abort: the same doc says Cursor,
Windsurf, Kimi and Cline keep the config they wrote inside install().

Concretely: `installAllRuntimes(['codex', 'kilo'])` where Kilo's finishInstall
hits EACCES writing kilo.json rolled Codex's config.toml back to its pre-install
bytes — on an update, silently downgrading a working Codex install to the
previous version while the user had just been told it was Done.

Select the rollback by error kind instead. `assertConfiguredEntrypoints` already
tags its error with `configuredEntrypointValidation`, so the full snapshot
restore runs for that error (and anything downstream of it, including
finishInstall's per-runtime backstop) and the installer-migrations-only closure
runs for everything else. The codex result now also exposes that narrow closure
as `rollbackInstallerMigrationsOnly`; `rollbackInstallerMigrations` keeps meaning
the full restore, so the direct-call contract asserted by
tests/codex-config.test.cjs is unchanged.

Adds a regression test that installs codex+kilo together, injects EACCES on the
Kilo permission write by monkeypatching node:fs (restored in a finally — never
chmod 0o000, which root bypasses in CI), and asserts Codex's config.toml keeps
the bytes the successful install wrote. Verified red against the pre-fix
unconditional path.

Cline cannot host this test: its plan is writesSharedSettings:false +
finishPermissionWriter:null, so its finishInstall performs no write and has no
non-entrypoint failure path. Kilo's configureKiloPermissions runs unconditionally
(unlike OpenCode's, it is not GSD_TEST_MODE-gated) and ends in an unguarded
fs.writeFileSync.

* docs(#4249): sync CONTEXT.md's rollback description with the round-6 narrowing

CONTEXT.md still described Codex's rollback as an unconditional bind
to restoreCodexSnapshot after ff13adc00 scoped it to entrypoint-
validation failures via rollbackInstallerMigrationsOnly and the
configuredEntrypointValidation error tag. Caught during the round-6
PR body pass.

* fix(#4249): stop rollbackInstallerMigrations meaning its own opposite

Codex's install() result bound `rollbackInstallerMigrations` to
restoreCodexSnapshot (the FULL pre-install snapshot restore) and put the
actual installer-migrations-only closure behind
`rollbackInstallerMigrationsOnly` — so for one runtime the unsuffixed name
meant the opposite of what it says, and CONTEXT.md had to concede as much
in prose. Invert it: `rollbackInstallerMigrations` is the narrow closure
for every runtime, matching both its name and the meaning it already has
on next, and the snapshot restore gets its own Codex-only field,
`rollbackPreInstallSnapshot`. The selection in
rollbackFinalizedInstallerMigrations collapses to one line and no longer
needs a fallback chain.

Also in this commit, all against the same rollback path:

- Correct the rollbackFinalizedInstallerMigrations comment. It read as if
  the round-6 narrowing prevented any sibling-triggered revert of a Codex
  install the user has already seen "Done!" for. It does not, and is not
  meant to: `wide` is true for ANY entrypoint-validation error from ANY
  runtime, because the aggregate gate is all-or-nothing — an invalid Cline
  entrypoint reverts Codex's snapshot, which
  tests/configured-entrypoint-validation.test.cjs's 'an aggregate
  entrypoint validation failure rolls the Codex install back (#4249)'
  asserts directly. The discriminator is the error's KIND, not which
  runtime owns the failing path. Comment and CONTEXT.md now say that.

- Name the runtime in the "Configured entrypoint validation failed" error.
  ConfiguredEntrypointInvalid already carries `runtime`; the message threw
  it away, leaving an operator of a multi-runtime install unable to tell
  whose entrypoint broke — which matters precisely because the failure can
  revert a runtime that was itself fine.

- Set `configuredEntrypoints: []` explicitly on the copilot-instructions
  early return. Every other branch states the key; this one relied on
  installAllRuntimes' `(result.configuredEntrypoints || [])` defence.
  `[]` is correct, not a workaround: every Copilot hook is an inline
  printf one-liner (GSD_COPILOT_*_HOOK_BASH/PWSH), so there is no
  GSD-managed script or interpreter to resolve.

No behaviour change beyond the error-message text.

* docs(#4249): narrow the smoke scan's config-surface claim to what it checks

RUNTIME_CONFIG_FILES claimed every GSD-managed executable a runtime is
told to launch is registered in one of settings.json / hooks.json /
config.toml, and that nothing else in a config dir is runtime
configuration. Both halves are false as stated. Cline registers its hook
at .clinerules/hooks/PreToolUse — a subdirectory, and not one of those
names (writeClineArtifacts, src/runtime-hooks-surface.cts). Kimi's native
[[hooks]] config.toml lives under resolveKimiHooksTomlDir() (~/.kimi),
a directory separate from Kimi's own GSD configDir — the same gap
installer-migration 007 already documents as structurally unreachable.

The scan is in fact correct for what it runs against: entrypointRuntimes
defaults to claude + codex, whose launch paths do all live in those three
top-level files. Restate the docstring at that scope, name the two known
out-of-scope surfaces, and warn that adding either runtime to
entrypointRuntimes without teaching scanConfiguredEntrypoints about its
surface yields a scan that finds zero entrypoints and proves nothing.
The entrypointRuntimes default comment carried the same overgeneralization
("every other runtime reuses one of them") and is corrected with it.

Documentation only; no code change.

* fix(#4249): complete configuredEntrypoints/rollback shape on unparseable settings.local.json

An internal adversarial review (agy/gemini-3.8-flash-medium, round 8) found
that install()'s settings-json early return for an unparseable
settings.local.json omitted configuredEntrypoints and
rollbackInstallerMigrations from its result, unlike every other branch.
rollbackFinalizedInstallerMigrations reads result.rollbackInstallerMigrations
unconditionally, so this branch silently dropped its own installer-migration
rollback on a later finalize-stage failure.

Completed the return shape: configuredEntrypoints: [] (matching Copilot's
equally-early no-entrypoints-yet return) and rollbackInstallerMigrations
(already in closure scope). Red-then-green regression test added.

* test(#4249): ensure hooks/dist before packing in release-tarball-smoke.install.test.cjs

Same internal adversarial review (round 8): this suite's before() packed the
tarball directly, without the ensureHooksDist() guard every sibling
install-test suite (install.test.cjs, install-minimal-hooks.test.cjs,
mcp-catalog-parity.install.test.cjs) already uses. On a clean tree, or run
in isolation ahead of a suite that builds hooks/dist itself, this suite's
pack would ship a tarball with no hook scripts and fail closed on
SMOKE.INIT_FAILED instead of testing anything.

* fix(#4249): refresh stale test-timings weight for the codex-config split

next's own consolidation split (#4139/#4540) moved tests/codex-config.test.cjs's
heavy install()-pipeline blocks into tests/codex-config-hooks.test.cjs, but the
CI shard packer's weight table (tests/test-timings.json) was never updated:
codex-config.test.cjs still carried its pre-split weight (127783ms, ~18x the
suite mean), and codex-config-hooks.test.cjs — which now holds the #3245 block
this PR extends with its own #4249 install()-pipeline test — had no entry at
all, so the packer would silently underestimate it at the table's median
weight (roughly a 9x underestimate against its real cost).

trek-e's most recent review flagged a Windows shard timeout in-flight on
codex-config.test.cjs, plausibly aggravated by this PR's own addition to that
file before the rebase moved it. Re-measured both files locally (node --test
--test-reporter=tap, max of 3 runs, matching the table's own max-across-streams
methodology) and patched just these two entries — not a full regeneration,
which would need real multi-lane CI data this session doesn't have access to.

* fix(#4249): register configured-entrypoint-validation tests in the conformance-tier lists

next's platform-conformance-tier classifier (#4591/#4598) landed after this
branch's last rebase, so tests/configured-entrypoint-validation.test.cjs and
tests/codex-config-hooks.test.cjs were never classified, failing lint:ci's
gen-platform-conformance-tier --check and both the Linux and macOS conformance
suites.

* fix(#4249): drop codex-config.test.cjs from the #4733 pinned isolated-set expectation

next's #4733 (landed after this branch's last rebase) replaced the static
ISOLATED_HEAVY_FILES set with a threshold derived live from
tests/test-timings.json, and pins the current derived result in
EXPECTED_ISOLATED_UNIT_FILES for regression coverage. That pinned list still
named codex-config.test.cjs, whose own weight this PR already dropped from
127783ms to 189ms (after splitting its heavy install()-pipeline blocks into
codex-config-hooks.test.cjs) — well under #4733's derived 120000ms bar. The
live-computed set correctly no longer includes it; the pinned expectation is
updated to match.

* fix(#4249): name the rollback consequence in the entrypoint-validation error, and prove Cline's file survives it

trek-e's review flagged two Major gaps: the thrown error read identically
regardless of which of three real outcomes a runtime hit (nothing
persisted / snapshot reverted / config left broken on disk), and no test
proved the disclosed "left on disk, unreverted" case for Cursor/Windsurf/
Kimi/Cline — only Codex's revert path was ever asserted.

assertConfiguredEntrypoints now tags each invalid entry with its actual
consequence, mirrored from docs/how-to/update-gsd.md's existing
rollback-matrix disclosure. A new test drives the same aggregate failure
through Cline (whose own entrypoint is the one that fails) and asserts
its hook file is still on disk afterward.

* fix(#4249): close 4 gaps antigravity's adversarial review found in the entrypoint-validation PR

One review pass (gemini-3.8-flash-high via the antigravity review lane)
against this PR's full diff against next, findings independently verified
against source before fixing:

- Copilot's install() return object was the only one of 6 runtime branches
  missing rollbackInstallerMigrations — reachable now that this PR's own
  aggregate gate runs rollback across every result on any runtime's
  entrypoint failure, not just Copilot's own.
- buildHookCommand's unresolved-bash early return skipped track() entirely,
  so a win32 install with no Git Bash silently produced an unregistered
  .sh hook instead of the 'unresolved-interpreter' validation failure
  configuredEntrypointsForHook's own comment said it would.
- release-tarball-smoke.cjs reported a Cycle 4 install failure under
  SMOKE.INIT_FAILED (Cycle 1's code) instead of the already-existing
  SMOKE.INSTALL_FAILED.
- SCRIPT_PATH_RE excluded whitespace to avoid swallowing a shell command's
  trailing args, which also truncated any configDir containing a space
  (e.g. a real "/Users/John Doe/.claude"), silently zeroing the scan.
  Anchored the match on the already-known configDir prefix instead of a
  generic absolute-path guess: removes the ambiguity outright rather than
  patching the character class, and stays a raw-text scan on purpose (it
  catches a writer that emits a path without registering it — a
  JSON.parse of the expected schema would miss exactly that case).

One suggested finding (test-timings.json "missing" the new test file) was
verified false — that table only holds measured CI timings, populated
after a file's first real run — and one Ponytail suggestion (a
JSON.stringify dedup key) was rejected as it would reintroduce a real, if
narrow, key-collision risk for no benefit.

* fix(#4249): fix fork CI red from a stale changeset pr field and an unquoted docs/ comment

changeset-lint requires pr: to match the PR it runs on (16 on the fork,
not the eventual upstream number) — rehearsal-branch convention already
established earlier in this PR's history.

lint-docs-guard-registration's quote-pairing heuristic doesn't require the
docs/ path itself to be quoted — it flags a file once ANY quote-delimited
span containing "docs/" appears anywhere in it, alongside any real fs read
call. A comment ending "...update-gsd.md's rollback-matrix paragraph"
supplied the closing quote character (the possessive apostrophe) the
heuristic paired with an unrelated single-quoted string earlier in the
file. Reworded to avoid the unquoted apostrophe next to the path.

* chore(#4249): point the changeset pr field back at the upstream PR

Fork rehearsal (PR #16) is green; the real target for this changeset is
upstream PR #4249.

---------

Co-authored-by: Test <test@test.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-09-16 04:23:53 -04:00

1045 lines
81 KiB
Markdown

# GSD Core Architecture
> System architecture for contributors and advanced users. For user-facing documentation, see [Feature Reference](FEATURES.md) or [User Guide](USER-GUIDE.md).
---
## Table of Contents
- [System Overview](#system-overview)
- [Design Principles](#design-principles)
- [Component Architecture](#component-architecture)
- [Agent Model](#agent-model)
- [Data Flow](#data-flow)
- [File System Layout](#file-system-layout)
- [Installer Architecture](#installer-architecture)
- [Hook System](#hook-system)
- [CLI Tools Layer](#cli-tools-layer)
- [Runtime Abstraction](#runtime-abstraction)
---
## System Overview
GSD Core is a **meta-prompting framework** that sits between the user and AI coding agents (Claude Code, Kimi CLI, OpenCode, Kilo, Codex, Copilot, Antigravity, Trae, Cline, Augment Code). It provides:
1. **Context engineering** — Structured artifacts that give the AI everything it needs per task (see [Context engineering](explanation/context-engineering.md))
2. **Multi-agent orchestration** — Thin orchestrators that spawn specialized agents with fresh context windows (see [Multi-agent orchestration](explanation/multi-agent-orchestration.md))
3. **Spec-driven development** — Requirements → research → plans → execution → verification pipeline
4. **State management** — Persistent project memory across sessions and context resets
```
┌──────────────────────────────────────────────────────┐
│ USER │
│ /gsd-command [args] │
└─────────────────────┬────────────────────────────────┘
│
┌─────────────────────▼────────────────────────────────┐
│ COMMAND LAYER │
│ commands/gsd/*.md — Prompt-based command files │
│ (Claude Code custom commands / Codex skills) │
└─────────────────────┬────────────────────────────────┘
│
┌─────────────────────▼────────────────────────────────┐
│ WORKFLOW LAYER │
│ gsd-core/workflows/*.md — Orchestration logic │
│ (Reads references, spawns agents, manages state) │
└──────┬──────────────┬─────────────────┬──────────────┘
│ │ │
┌──────▼──────┐ ┌─────▼─────┐ ┌────────▼───────┐
│ AGENT │ │ AGENT │ │ AGENT │
│ (fresh │ │ (fresh │ │ (fresh │
│ context) │ │ context)│ │ context) │
└──────┬──────┘ └─────┬─────┘ └────────┬───────┘
│ │ │
┌──────▼──────────────▼─────────────────▼──────────────┐
│ CLI TOOLS LAYER │
│ gsd-tools.cjs command families + domain modules │
│ command-routing-hub + observability seams │
└──────────────────────┬───────────────────────────────┘
│
┌──────────────────────▼───────────────────────────────┐
│ FILE SYSTEM (.planning/) │
│ PROJECT.md | REQUIREMENTS.md | ROADMAP.md │
│ STATE.md | config.json | phases/ | research/ │
└──────────────────────────────────────────────────────┘
```
---
## Design Principles
### 1. Fresh Context Per Agent
Every agent spawned by an orchestrator gets a clean context window (up to 200K tokens). This eliminates context rot — the quality degradation that happens as an AI fills its context window with accumulated conversation.
### 2. Thin Orchestrators
Workflow files (`gsd-core/workflows/*.md`) never do heavy lifting. They:
- Load context via `gsd-tools.cjs init <workflow>`
- Spawn specialized agents with focused prompts
- Collect results and route to the next step
- Update state between steps
### 3. File-Based State
All state lives in `.planning/` as human-readable Markdown and JSON. No database, no server, no external dependencies. This means:
- State survives context resets (`/clear`)
- State is inspectable by both humans and agents
- State can be committed to git for team visibility
### 4. Absent = Enabled
Workflow feature flags follow the **absent = enabled** pattern. If a key is missing from `config.json`, it defaults to `true`. Users explicitly disable features; they don't need to enable defaults.
### 5. Defense in Depth
Multiple layers prevent common failure modes:
- Plans are verified before execution (plan-checker agent)
- Execution produces atomic commits per task
- Post-execution verification checks against phase goals
- UAT provides human verification as final gate
---
## Component Architecture
### Commands (`commands/gsd/*.md`)
User-facing entry points. Each file contains YAML frontmatter (name, description, allowed-tools) and a prompt body that bootstraps the workflow. Commands are installed as:
- **Claude Code:** Custom slash commands (hyphen form, `/gsd-command-name`)
- **OpenCode / Kilo:** Slash commands (hyphen form, `/gsd-command-name`)
- **Codex:** Skills (`$gsd-command-name`)
- **Copilot:** Slash commands (hyphen form, `/gsd-command-name`)
- **Kimi CLI:** Agent Skills (`/skill:gsd-command-name`) plus an explicit custom agent launch with `kimi --agent-file`
- **Antigravity:** Skills
**Total commands:** see [`docs/INVENTORY.md`](INVENTORY.md#commands) for the authoritative count and full roster.
#### Two-stage hierarchical routing (v1.40, [#2792](https://github.com/open-gsd/gsd-core/issues/2792))
To keep the eager skill-listing token cost low, v1.40 introduces six namespace **meta-skills** (`gsd-workflow`, `gsd-project`, `gsd-quality`, `gsd-context`, `gsd-manage`, `gsd-ideate` — sourced from `commands/gsd/ns-*.md`, but the invocable `name:` is the bare form shown here) layered above the concrete sub-skills. On runtimes with non-recursive skill loaders (cline, qwen, hermes, augment, trae) the installer now realizes this fully: it emits only the 6 namespace router bundles as top-level skills and nests the ~61 concrete skills under `<router>/skills/<name>/SKILL.md`, so the eager listing is ≈6 entries instead of ≈67. The model selects a namespace router, which instructs it to read the nested concrete skill file via a routing table embedded in the router body. On these runtimes concrete skills are **not** directly invocable by bare name via the Skill tool; they are reachable through the router. Slash commands (`/gsd-*`, via the separate commands surface) are unaffected where the runtime has one. On runtimes with recursive or unconfirmed skill loaders (claude global, cursor, codex, copilot, windsurf, codebuddy, opencode, kilo, antigravity) the layout remains flat — all skills emitted at the top level as before. Antigravity moved from nested to flat in #1614: `agy` scans only `skills/<name>/SKILL.md`, so nested sub-skills were unreachable. Claude was reverted to flat in #924: the Skill tool hard-errors on unknown names rather than re-routing via the router, so nested concrete skills were uninvokable.
The router descriptions use pipe-separated keyword tags (≤ 60 chars) per the Tool Attention research showing keyword-dense tags outperform prose for routing at ~40 % the token cost.
#### MCP token-budget interaction
The eager skill listing is one of two recurring per-turn token costs. The other is the MCP tool schema injected by every enabled MCP server in `.claude/settings.json`. Heavyweight MCP servers (browser/playwright, Mac-tools, Windows-tools) can each cost 20 k+ tokens per turn — often dwarfing what `model_profile` tuning saves. The toggle lives in the Claude Code harness (`enabledMcpjsonServers` / `disabledMcpjsonServers` in `.claude/settings.json`) and is **not** a GSD concern. Together, the two-stage routing layer (#2792) and disciplined MCP enablement are the largest cost levers per turn. See [`docs/USER-GUIDE.md`](USER-GUIDE.md) and `references/context-budget.md` for the audit checklist.
### Workflows (`gsd-core/workflows/*.md`)
Orchestration logic that commands reference. Contains the step-by-step process including:
- Context loading via `gsd-tools.cjs init` handlers
- Agent spawn instructions with model resolution
- Gate/checkpoint definitions
- State update patterns
- Error handling and recovery
**Total workflows:** see [`docs/INVENTORY.md`](INVENTORY.md#workflows) for the authoritative count and full roster.
#### Progressive disclosure for workflows
Workflow files are loaded verbatim into Claude's context every time the
corresponding `/gsd-*` command is invoked. The workflow size budget enforced by
`tests/workflow-size-budget.test.cjs` keeps each file bounded, mirroring the
the agent size-budget convention. The budget is measured in **bytes** (#717), not lines:
line count over-penalizes prose and under-catches token-dense tables and code
blocks, whereas bytes are deterministic and match the unit our vendors bound on
— Codex truncates instruction docs past 32,768 bytes (`project_doc_max_bytes`).
We adopt that unit, not that exact number: the XL/LARGE ceilings below sit above
32,768 because these are grandfathered top-level orchestrators loaded by Claude,
not Codex AGENTS.md docs.
| Tier | Per-file byte limit |
|-----------|---------------------|
| `XL` | 90,000 — top-level orchestrators (`execute-phase`, `plan-phase`, `new-project`) |
| `LARGE` | 54,000 — multi-step planners and large feature workflows |
| `DEFAULT` | 38,000 — focused single-purpose workflows (the target tier) |
Ceilings are not fixed forever: under the tighten-only ratchet (#597) each one
tracks its tier's current high-water mark within a small grace band, so budgets
may only decrease over time.
**Why the budget exists.** With prompt caching the per-invocation *cost* of a
large workflow is modest (cache reads run ~10% of input). The stronger,
caching-independent reason is **quality**: as context grows, recall and
reasoning degrade ("context rot" / attention budget), so leaner, higher-signal
instructions produce better plans. The ceiling protects the agent's attention,
not just the token bill.
Because the budget measures one file, it is a proxy for the real goal —
*bounded loaded context*. Extraction only helps when the extracted content is
loaded **lazily** (Read at the step that needs it). Moving prose into a file
that is still eagerly `@`-imported shrinks the measured file without shrinking
loaded context, which games the proxy rather than serving the goal.
`workflows/discuss-phase.md` is held to a stricter <30,000-byte ceiling per
the discuss-phase byte budget (#717; the discuss-phase/modes split keeps it ≈32000 bytes). When a workflow grows
beyond its tier, extract per-mode bodies into
`workflows/<workflow>/modes/<mode>.md`, templates into
`workflows/<workflow>/templates/`, and shared knowledge into
`gsd-core/references/`. The parent file becomes a thin dispatcher that
Reads only the mode and template files needed for the current invocation.
`workflows/discuss-phase/` is the canonical example of this pattern —
parent dispatches, modes/ holds per-flag behavior (`power.md`, `all.md`,
`auto.md`, `chain.md`, `text.md`, `batch.md`, `analyze.md`, `default.md`,
`advisor.md`), and templates/ holds CONTEXT.md, DISCUSSION-LOG.md, and
checkpoint.json schemas that are read only when the corresponding output
file is being written.
`workflows/plan-phase.md`, `workflows/execute-phase.md`, and the
`gsd-planner` / `gsd-executor` agent definitions apply the same discipline
to their MVP-only reference bodies — `planner-mvp-mode.md`,
`user-story-template.md`, `skeleton-template.md`, and `execute-mvp-tdd.md`
are referenced for the planner/executor to Read only on MVP,
Walking-Skeleton, or MVP+TDD paths, rather than eagerly `@`-imported, so
non-MVP runs do not pay their context cost (guards against the "`@`-import
behind a conditional still loads eagerly" leak; see #720). The dedicated
`mvp-phase` workflow keeps its eager imports, since it is always MVP.
### Agents (`agents/*.md`)
Specialized agent definitions with frontmatter specifying:
- `name` — Agent identifier
- `description` — Role and purpose
- `tools` — Allowed tool access (Read, Write, Edit, Bash, Grep, Glob, WebSearch, etc.)
- `color` — Terminal output color for visual distinction
**Total agents:** 33
### References (`gsd-core/references/*.md`)
Shared knowledge documents that workflows and agents `@-reference` (see [`docs/INVENTORY.md`](INVENTORY.md#references) for the authoritative full roster):
**Core references:**
- `checkpoints.md` — Checkpoint type definitions and interaction patterns
- `gates.md` — 4 canonical gate types (Confirm, Quality, Safety, Transition) wired into plan-checker and verifier
- `model-profiles.md` — Per-agent model tier assignments
- `model-profile-resolution.md` — Model resolution algorithm documentation
- `verification-patterns.md` — How to verify different artifact types
- `verification-overrides.md` — Per-artifact verification override rules
- `planning-config.md` — Full config schema and behavior
- `git-integration.md` — Git commit, branching, and history patterns
- `git-planning-commit.md` — Planning directory commit conventions
- `questioning.md` — Dream extraction philosophy for project initialization
- `tdd.md` — Test-driven development integration patterns
- `ui-brand.md` — Visual output formatting patterns
- `common-bug-patterns.md` — Common bug patterns for code review and verification
**Workflow references:**
- `agent-contracts.md` — Formal interface between orchestrators and agents
- `context-budget.md` — Context window budget allocation rules
- `continuation-format.md` — Session continuation/resume format
- `domain-probes.md` — Domain-specific probing questions for discuss-phase
- `gate-prompts.md` — Gate/checkpoint prompt templates
- `revision-loop.md` — Plan revision iteration patterns
- `universal-anti-patterns.md` — Common anti-patterns to detect and avoid
- `artifact-types.md` — Planning artifact type definitions
- `phase-argument-parsing.md` — Phase argument parsing conventions
- `decimal-phase-calculation.md` — Decimal sub-phase numbering rules
- `workstream-flag.md` — Workstream active pointer conventions
- `user-profiling.md` — User behavioral profiling methodology
- `thinking-partner.md` — Conditional thinking partner activation at decision points
**Thinking model references:**
References for integrating thinking-class models (o3, o4-mini, Gemini 2.5 Pro) into GSD workflows:
- `thinking-models-debug.md` — Thinking model patterns for debugging workflows
- `thinking-models-execution.md` — Thinking model patterns for execution agents
- `thinking-models-planning.md` — Thinking model patterns for planning agents
- `thinking-models-research.md` — Thinking model patterns for research agents
- `thinking-models-verification.md` — Thinking model patterns for verification agents
**Modular planner decomposition:**
The planner agent (`agents/gsd-planner.md`) was decomposed from a single monolithic file into a core agent plus reference modules to stay under the 50K character limit imposed by some runtimes:
- `planner-gap-closure.md` — Gap closure mode behavior (reads VERIFICATION.md, targeted replanning)
- `planner-reviews.md` — Cross-AI review integration (reads REVIEWS.md from `/gsd-review`)
- `planner-revision.md` — Plan revision patterns for iterative refinement
### Templates (`gsd-core/templates/`)
Markdown templates for all planning artifacts. Used by `gsd-tools.cjs template fill` / `phase.scaffold` (and top-level `scaffold`) to create pre-structured files:
- `project.md`, `requirements.md`, `roadmap.md`, `state.md` — Core project files
- `phase-prompt.md` — Phase execution prompt template
- `summary.md` (+ `summary-minimal.md`, `summary-standard.md`, `summary-complex.md`) — Granularity-aware summary templates
- `DEBUG.md` — Debug session tracking template
- `UI-SPEC.md`, `UAT.md`, `VALIDATION.md` — Specialized verification templates
- `discussion-log.md` — Discussion audit trail template
- `codebase/` — Brownfield mapping templates (architecture, stack)
- `research-project/` — Research output templates (SUMMARY, STACK, FEATURES, ARCHITECTURE, PITFALLS)
### Hooks (`hooks/`)
Runtime hooks that integrate with the host AI agent:
| Hook | Event | Purpose |
|------|-------|---------|
| `gsd-statusline.js` | `statusLine` | Displays model (long-context suffixes like `(1M context)` collapse to a compact `(1M)` badge), task, directory, and context usage bar |
| `gsd-context-monitor.js` | `PostToolUse` / `AfterTool` | Injects agent-facing context warnings at 35%/25% remaining by default (configurable — see [CONFIGURATION.md](CONFIGURATION.md)) |
| `gsd-check-update.js` | `SessionStart` | Foreground trigger for the background update check |
| `gsd-ensure-canonical-path.js` | `SessionStart` | For Claude Code plugin installs, symlinks `~/.claude/gsd-core/{bin,contexts,references,templates,workflows}` to the plugin's bundled tree so `@~/.claude/gsd-core/...` includes resolve; runs first in `SessionStart`, no-op in classic installs, self-heals after `claude plugin update` (#997) |
| `gsd-check-update-worker.js` | (helper) | Background worker spawned by `gsd-check-update.js`; no direct event registration |
| `gsd-prompt-guard.js` | `PreToolUse` | Scans `.planning/` writes for prompt injection patterns (advisory) |
| `gsd-read-injection-scanner.js` | `PostToolUse` | Scans Read tool output for injected instructions in untrusted content |
| `gsd-workflow-guard.js` | `PreToolUse` | Detects file edits outside GSD workflow context (advisory, opt-in via `hooks.workflow_guard`) |
| `gsd-secret-read-guard.js` | `PreToolUse` | Hard-blocks Read / Grep / Bash reads of `.env`, `.env.<suffix>` (templates such as `.env.example` exempt) and `.secrets`; replaces the installer-written `Read(.env*)` permission deny rules, which made every `cd DIR && grep …` compound prompt for approval on Claude Code ≥ 2.1.259 (#4221) |
| `gsd-read-guard.js` | `PreToolUse` | Advisory guard preventing Edit/Write on files not yet read in the session |
| `gsd-session-state.sh` | `SessionStart` | Session state tracking for shell-based runtimes |
| `gsd-validate-commit.sh` | `PreToolUse` | Commit validation for conventional commit enforcement |
| `gsd-phase-boundary.sh` | `PostToolUse` | Phase boundary detection for workflow transitions |
See [`docs/INVENTORY.md`](INVENTORY.md#hooks) for the authoritative hook roster.
**Crash policy (ADR-3889 Phase 7, #3911).** Every enforcement hook terminates
through `hooks/lib/hook-exit.js`'s `allow(payload)` (exit 0), `deny(payload,
stderrPayload?)` (exit 2), or `crash(onCrash, payload)` — the last dispatching
per a `HOOK_ON_CRASH` policy the hook must declare explicitly (`ALLOW` or
`DENY`, no default), so a hook's fail-open/fail-closed stance is a visible
declaration rather than an inference from a bare `process.exit(N)`. Two hooks
are deliberate exceptions — `gsd-read-injection-scanner.js` (PostToolUse) and
`gsd-cursor-subagent-start.js` (Cursor) — whose harnesses read the block
decision from the JSON response body at exit 0, not from the exit code, so
they never call `deny()`. See
[Declare a hook's crash policy](how-to/declare-a-hook-crash-policy.md).
### Command Routing Hub (`gsd-core/bin/lib/command-routing-hub.cjs`)
CJS command family routers dispatch through `CommandRoutingHub`. The hub owns the no-throw pure-result contract (`hub.dispatch()` catches internal exceptions and returns `{ ok: false, kind, ...typedPayload }`) and the closed runtime error taxonomy (`UnknownCommand`, `InvalidArgs`, `HandlerRefusal`, `HandlerFailure`). Router adapters remain thin CLI translators — they build the hub, call `dispatch`, then map the Result to `output()`/`error()` calls. The runtime is single-path (no dual-runtime mode selection). See `docs/adr/0174-retire-gsd-sdk-package-boundary.md`.
> **Planned (ADR-2346 / epic #2345):** the `runCommand` 73-case switch is being dissolved into a two-layer dispatch — families via the `commandFamilies` registry (ADR-959 mechanism, completed) and single-purpose leaf verbs via a table filling the prepared `_dispatchNonFamily` seam — collapsing `runCommand` to a ~15-line dispatcher. Behavior-preserving; tracked phase-by-phase under epic #2345. The current-state description above holds until each phase lands.
### Capability Command Dispatch (`gsd-core/bin/gsd-tools.cjs`, ADR-1244 D7)
Command families declared by capabilities (`commands: [{ family, module, router }]`) are dispatched from the registry rather than a hardcoded switch. The `runCommand` default arm tries, in order:
1. **First-party** — `dispatchCapabilityCommand` against the frozen `capability-registry.cjs` `commandFamilies`, loading the router from `bin/lib/`. The in-tree families (`graphify`, `intel`, `audit`) reach their routers this way (the legacy hardcoded switch is retired).
2. **Third-party (installed overlay)** — `dispatchOverlayCapabilityCommand` calls `loadRegistry({ includeInstalled })` and dispatches a family only when its `capId` appears in `_overlay.commandRoots`. The loader lists a command root **only** for an accepted overlay capability with a **committed** ledger entry (consent gate), and the router module is `require()`'d **from that capability's install root**, confined by basename validation + `realpath` containment (rejecting `..` traversal and symlink escape). This is the one point where third-party capability code executes; see [the capability trust model](explanation/capability-trust-model.md) for the consent + confinement + project-scope trust boundary.
Both paths share the same guards: prototype-pollution-safe command keys, an own-property router check, and synchronous-only routers (an async router is a fail-fast error).
### Reviewer-Lane Capability Trait (#4209, ADR-2782)
`/gsd-code-review` optionally corroborates its internal review with external reviewer lanes (`--codex`, `--agy`, ...), gated by the reusable `supportsReviewerLanes` capability-step trait and dispatched through the single `dispatchReviewerLanes` interpreter — see `gsd-core/references/loop-hook-dispatch.md` for the trait and `src/reviewer-step-dispatch.cts` for the interpreter's fail-closed contract. `gsd-code-reviewer` is the sole consolidator: it independently re-verifies every external claim against the actual source before writing anything to `REVIEW.md`, so a lane's evidence is corroborating input, never a second output schema.
### Research Module (`src/research-{store,provider}.cts`, `src/package-legitimacy.cts`)
The Research Module implements an **L2-hybrid seam**: code owns the cache, provider policy, and package legitimacy verdicts; MCP owns the actual network fetch.
Three compiled modules (generated to `gsd-core/bin/lib/*.cjs` per ADR-457) are reachable via `gsd-tools query research-plan | research-store | package-legitimacy`:
- **Research Store** — content-addressed cache (`sha256(ecosystem+library+version+query+kind)`) with per-source TTL (curated-doc: 30 d, medium: 7 d, web/synthesis: 1 d) and two storage tiers: `~/.gsd/research-cache` for cross-project curated-doc hits, `.planning/research/.cache` for project-local web/synthesis results.
- **Research Provider** — single `PROVIDER_WATERFALL` (`Context7→Ref→Jina→websearch` for docs; `Exa→Tavily→Perplexity→Brave→websearch` for web; `Firecrawl→Jina` for scrape-only). `planResearch()` returns cache hits plus a fetch plan; `classifyConfidence()` stamps `HIGH|MEDIUM|LOW` by provider tier.
- **Package Legitimacy** — registry-API verdicts (npm/PyPI/crates.io injectable adapters) producing `OK|SUS|SLOP` per package. `slopcheck` is an optional escalate-only adapter; absence leaves registry verdicts intact rather than downgrading everything to `[ASSUMED]`.
**Data flow:**
```
agent
│
▼
gsd-tools query research-plan ← Research Provider: check cache, build fetch plan
│
├── [cache hits] ──────────────────► RESEARCH.md (digest only, no raw content)
│
└── [fetch plan] ──────────────────► MCP fetch (agent calls MCP tools with the plan)
│
▼
gsd-tools query research-store (put)
│
▼
RESEARCH.md path returned to orchestrator
```
Agents always return a `RESEARCH.md` path, never raw fetched content. Context discipline is enforced through subagent isolation, compact provider output, and fetch-to-disk. See [ADR-0656](adr/0656-research-module-seam.md).
### Context Predicate Fact-Store (`src/context-predicates.cts`, ADR-1671)
The `CONTEXT.md` predicate fact-store — every backtick-wrapped `CLASS.subkey=value` declaration in the repo-root `CONTEXT.md` — has a compiled parser/selector seam (generated to `gsd-core/bin/lib/context-predicates.cjs` per ADR-457) reachable live via `gsd-tools query context-predicates --class|--prefix|--contains`. Fence-aware line skipping mirrors `markdown-sectionizer.cts`'s exported `scanFencedBlocks` delimiter-matching rule exactly (proven by a fence-skip parity test suite), but is scanned by a LOCAL, interleaved single pass rather than a call into that seam directly: fences and HTML comments must mutually suppress each other's open/close detection while either is active (a fence delimiter inside a real comment, or a comment token inside a real fence, must not falsely toggle the other construct), and that precedence cannot be resolved by two independent passes over `scanFencedBlocks`'s comment-blind output — see `src/context-predicates.cts`'s module doc comment.
`scripts/gen-context-index.cjs --check` is the CI drift-guard for the committed `docs/CONTEXT-INDEX.json` artifact: it fails on staleness between a fresh parse of `CONTEXT.md` and the committed file, and on any duplicate predicate ID. It is wired into `lint:generated-sync` (so `lint:ci`, so CI). `docs/CONTEXT-INDEX.json` is **generated — never hand-edit it**; regenerate with `gen-context-index.cjs --write` (also wired into `build`, after `build:lib`, and into `regen:derived`). The generator `require()`s the compiled `context-predicates.cjs`, so it must run after `build:lib` in any pipeline; `.github/workflows/test.yml` does this.
The committed index intentionally carries **no `line` field** for any predicate (ADR-1671 open question 4, resolved by #2928) — committed-but-uncompared metadata goes silently stale, the same defect class the drift-guard exists to catch, with the alarm removed. The live `gsd-tools query context-predicates` parse still returns `line`/`section` for callers that want to cite a source location. See [ADR-1671](adr/1671-dynamic-context-management-platform.md) and [CLI Tools Reference](CLI-TOOLS.md#query-context-predicates).
### Workflow Fragmentization and Emission (`src/workflow-fragments.cts`, ADR-1671)
Workflow markdown under `gsd-core/workflows/*.md` can mark one or more sections with an
in-file `<!-- gsd:section id="<id>" when="<when>" -->` / `<!-- /gsd:section -->` pair. A
compiled parser/composer seam (generated to `gsd-core/bin/lib/workflow-fragments.cjs` per
ADR-457) partitions a marked document into fragments and recomposes them through the shared
`context-composer.cjs` budget seam (ADR-1671, #2929) before any per-runtime converter sees the
text — so a marker attribute can never be corrupted by a `.claude/` → `.windsurf/`-style
path-rewrite regex. `bin/install.js`'s `copyWithPathReplacement` calls `composeWorkflow` on
every workflow file at emit time; an unmarked file (88 of the 89 shipped workflows today)
parses to a single implicit fragment and round-trips byte-identical, so this is a no-op for
every workflow that hasn't opted in yet.
Every fragment in this phase carries the `verbatim` strategy, so composition is structurally
non-lossy — nothing is trimmed regardless of budget. Fence and HTML-comment interleaving
reuses the same LOCAL, single-pass, mutually-suppressing scan discipline as
`context-predicates.cts` (see above), so a marker-shaped line inside a fenced code block or an
unrelated comment is never misread as structural. Markers are **stripped at emit** — the
installed artifact carries no build metadata and is smaller than the source by exactly the
stripped marker bytes.
See [Reference: Workflow fragments](reference/workflow-fragments.md) for the full marker
grammar, the frozen `when=` vocabulary, and fail-closed authoring rules, and
[ADR-1671](adr/1671-dynamic-context-management-platform.md) (open questions 1 and 2) for why
in-file markers were chosen over separate fragment files or a sidecar manifest.
### Section Manifest (`src/section-manifest.cts`, ADR-1671 Phases 5 and 6.1)
Two seams turn a workflow's `gsd:section` markers into per-invocation applicability data.
`scripts/gen-section-manifest.cjs --write` (wired into `build` after `build:lib`, and into
`lint:generated-sync`) scans `gsd-core/workflows/*.md` and writes the committed
`gsd-core/workflows/section-manifest.json`, keyed **per workflow** —
`{workflows: {"<name>": [{id, when, read}]}}` — where `read` is the path of the step file the
section's body was extracted to. A workflow with no marked sections contributes **no key at
all**: an absent key means degraded/unknown (the caller reads every section, the safe superset),
while a key present with an empty array means "computed, nothing applies". The generator reuses
`parseWorkflowSections` unchanged rather than re-implementing marker parsing, and fails closed
(`--check`) on a marker naming a step file that does not exist, a step file no marker
references, or a committed artifact still carrying the pre-6.1 flat `{sections: [...]}` shape.
A separate pure evaluator, `src/section-manifest.cts` (compiled to
`gsd-core/bin/lib/section-manifest.cjs` per ADR-457), maps one invocation's facts —
`{flags, phaseNumber, hasPriorPhases}` plus the optional `needsCodebaseMap`, `phaseMvpMode` and
`worktreesEnabled` booleans — to an included/excluded partition of section ids via
`selectSections`. `flags` is a `ReadonlySet<string>` of flag tokens; because `parseNamedArgs`
always materializes a boolean flag key (`false` when the token was absent, never `undefined`),
presence is **truthiness**, and the init router folds a boolean flag's own `false` into the
absent sentinel before the facts are built. Per Greenspun's Tenth Rule, the evaluator is a total
lookup over the frozen 14-atom `when=` vocabulary, never a parser: `WHEN_PREDICATES` is a
hand-written literal map that never derives a predicate from its atom string, and an
unrecognized value fails closed rather than being silently excluded. An atom is admitted only
when it has both a real consuming section and a fact the init seam actually computes — an atom
without the latter would evaluate `false` forever and silently disable its own section.
`execute-phase.md`'s `partial-wave` and `gap-closure-artifacts` sections — previously inlined
directly per #2930's pilot — now delegate to dedicated step files under
`gsd-core/workflows/execute-phase/steps/`, the same pattern the pre-existing `regression-gate`
section already used.
### CLI Tools (`gsd-core/bin/`)
Node.js CLI utility (`gsd-tools.cjs`) with domain modules split across `gsd-core/bin/lib/` (see [`docs/INVENTORY.md`](INVENTORY.md#cli-modules) for the authoritative roster):
| Module | Responsibility |
| ---------------------- | --------------------------------------------------------------------------------------------------- |
| `config-loader.cjs` | Project config loading — defaults merge, legacy-key migration, workstream overlay, unknown-key/profile-override validation, and federated config overlay (ADR-857 phase 3b) (extracted from `core.cjs`, ADR-857) |
| `federated-config.cjs` | Defensive merge of capability-declared config slices (ADR-857 phase 3b); exports `mergeFederatedConfig`; live for migrated Capability keys that are absent from the central config schema |
| `core-utils.cjs` | Shared low-level utility primitives — POSIX path normalization, sub-repo/subdirectory scanning, phase file stats, slug/one-liner/plan-id helpers, time-ago (extracted from `core.cjs`, ADR-857) |
| `core.cjs` | Shared utilities; compatibility re-exports for planning, I/O (`io.cjs`), and phase-id helpers |
| `io.cjs` | CLI I/O primitives — output/error emission, JSON-error mode, large-payload temp-file spillover |
| `phase-id.cjs` | Pure phase-id parsing/matching helpers — normalize, token match, regex builders (extracted from `core.cjs`, ADR-857) |
| `phase-locator.cjs` | Phase-directory search and location — active-phase discovery (`searchPhaseInDir`, `findPhaseInternal`) and archived-phase-dir enumeration (`getArchivedPhaseDirs`), matching phase ids/tokens against the filesystem (extracted from `core.cjs`, ADR-857) |
| `roadmap-parser.cjs` | ROADMAP.md parsing — milestone slicing, current-milestone extraction, phase/milestone lookups, milestone-phase filter (extracted from `core.cjs`, ADR-857) |
| `planning-workspace.cjs` | Planning seam (`planningDir`, `planningPaths`, active workstream routing, `.planning/.lock`) |
| `state.cjs` | STATE.md parsing, updating, progression, metrics |
| `phase.cjs` | Phase directory operations, decimal numbering, plan indexing |
| `roadmap.cjs` | ROADMAP.md parsing, phase extraction, plan progress |
| `config.cjs` | config.json read/write, section initialization |
| `verify.cjs` | Plan structure, phase completeness, reference, commit validation |
| `template.cjs` | Template selection and filling with variable substitution |
| `frontmatter.cjs` | YAML frontmatter CRUD operations |
| `init.cjs` | Compound context loading for each workflow type |
| `milestone.cjs` | Milestone archival, requirements marking |
| `commands.cjs` | Misc commands (slug, timestamp, todos, scaffolding, stats) |
| `model-profiles.cjs` | Model profile resolution table |
| `model-resolver.cjs` | Model and effort resolution policy — resolves model, tier, granularity, effort, and fast-mode for a given agent from project config and model profiles/catalog (extracted from `core.cjs`, ADR-857) |
| `security.cjs` | Path traversal prevention, prompt injection detection, safe JSON parsing, shell argument validation |
| `uat.cjs` | UAT file parsing, verification debt tracking, audit-uat support |
| `docs.cjs` | Docs-update workflow init, Markdown scanning, monorepo detection |
| `workstream.cjs` | Workstream CRUD, migration, session-scoped active pointer |
| `schema-detect.cjs` | Schema-drift detection for ORM patterns (Prisma, Drizzle, etc.) |
| `profile-pipeline.cjs` | User behavioral profiling data pipeline, session file scanning |
| `profile-output.cjs` | Profile rendering, USER-PROFILE.md and dev-preferences.md generation |
| `context-predicates.cjs` | `CONTEXT.md` predicate fact-store parser/selector (ADR-1671, #2928); backs `query context-predicates` and `scripts/gen-context-index.cjs`'s `docs/CONTEXT-INDEX.json` drift guard; compiled from `src/context-predicates.cts` |
| `loop-host-contract.cjs` | Generated Loop Host Contract — 12 loop points, per-step agent roles, and core artifacts; emitted by `scripts/gen-loop-host-contract.cjs` from workflow markers (ADR-894 §3); consumed by `gen-capability-registry.cjs` |
| `capability-loader.cjs` | Runtime registry overlay loader (ADR-1244 D2) — `loadRegistry({ includeInstalled })` composes the frozen first-party registry with a validated installed overlay of third-party capability manifests read from global `$GSD_HOME/.gsd/capabilities/` and project `<projectRoot>/.gsd/capabilities/`; first-party always wins; load-time `engines.gsd` re-gate skips incompatible overlays with a warning; gate-kind hooks on skipped capabilities fail OPEN — no gate is injected; a loud warning (stderr + envelope `warnings`) names the load failure and the `gsd capability remove <id>` remediation (#2009) |
| `capability-registry.cjs` | Generated central Capability Registry — role-partitioned index of all co-located capability declarations; emitted by `scripts/gen-capability-registry.cjs` (ADR-894 §5) |
| `loop-resolver.cjs` | Loop Extension Point resolver — ADR-857 phase 3c registry-consuming query; consumes resolved Capability State, filters `byLoopPoint` by capability enablement plus config activation, renders active hooks as markdown, emits `{ point, activeHooks, rendered }` envelope; `gsd-tools loop render-hooks <point> [--config-dir <path>]` |
| `capability-state.cjs` | Unified capability-state resolver — ADR-857 phase 4b/6; composes install profile, runtime surface, and config activation into one per-capability view consumed by workflow hook rendering; pure `resolveCapabilityState`, reusable `resolveCapabilityRuntimeState`, I/O `cmdCapabilityState`, and convenience predicate `isCapabilityActive(capId, cwd)`; `gsd-tools capability state [--config-dir <path>]` emits `{ runtimeConfigDir, capabilities[] }` where each entry carries `enabled` (installed && surfaced) and `active` (enabled && configActivation via the capability's `activationKey`; absent key → active===enabled) |
| `capability-validator.cjs` | Shared capability conformance validator (ADR-1244 D2) — extracted from `scripts/gen-capability-registry.cjs` so the build-time generator and the runtime overlay loader share one `validateCapability(manifest)` implementation; generative-parity is CI-guarded |
| `graphify-command-router.cjs` | ADR-959 capability command router — first real capability command cutover (phase 4d-impl-2); extracted from the `case 'graphify':` arm in `gsd-tools.cjs`; dispatches build/query/status/diff subcommands; discovered via `commandFamilies` in the capability registry |
| `audit-command-router.cjs` | ADR-959 capability command router (phase 4d-impl-3); extracted from the `case 'audit-uat':` and `case 'audit-open':` arms in `gsd-tools.cjs`; `routeAuditUat` → `uat.cjs:cmdAuditUat`, `routeAuditOpen` → `audit.cjs:{auditOpenArtifacts,formatAuditReport}`; discovered via `commandFamilies` in the capability registry |
| `intel-command-router.cjs` | ADR-959 capability command router (phase 4d-impl-4, last first-party cutover); extracted from the `case 'intel':` arm in `gsd-tools.cjs`; `routeIntelCommand` → all 9 intel subcommands via lazy `require('./intel.cjs')`; preserves non-raw `timeAgo` transform on `status.files[*].updated_at`; discovered via `commandFamilies` in the capability registry |
| `runtime-hooks-surface.cjs` | Hook-surface writer and configured-entrypoint validation module (ADR-857 phase 5f-1); owns Cline rules/agents-md/pre-tool-use hook generation, Cursor/Windsurf `hooks.json`, Kimi hook TOML, Copilot session-hook config, Codex hook-block management, and construction-time records for paths emitted into runtime configuration. Installer success requires those records to pass non-executing file-type and interpreter-resolution checks. |
---
## Agent Model
### Orchestrator → Agent Pattern
```
Orchestrator (workflow .md)
│
├── Load context: gsd-tools.cjs init <workflow> <phase>
│ Returns JSON with: project info, config, state, phase details
│
├── Resolve model: gsd-tools.cjs resolve-model <agent-name>
│ Returns: opus | sonnet | haiku | inherit
│
├── Spawn Agent (Task/SubAgent call)
│ ├── Agent prompt (agents/*.md)
│ ├── Context payload (init JSON)
│ ├── Model assignment
│ └── Tool permissions
│
├── Collect result
│
└── Update state: gsd-tools.cjs state update / state patch / state advance-plan
```
### Primary Agent Spawn Categories
Conceptual spawn-pattern taxonomy for the primary agents. For the authoritative agent roster (including the advanced/specialized agents such as `gsd-pattern-mapper`, `gsd-code-reviewer`, `gsd-code-fixer`, `gsd-ai-researcher`, `gsd-domain-researcher`, `gsd-eval-planner`, `gsd-eval-auditor`, `gsd-framework-selector`, `gsd-debug-session-manager`, `gsd-intel-updater`), see [`docs/INVENTORY.md`](INVENTORY.md#agents).
| Category | Agents | Parallelism |
| ---------------- | --------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- |
| **Researchers** | gsd-project-researcher, gsd-phase-researcher, gsd-ui-researcher, gsd-advisor-researcher | 4 parallel (stack, features, architecture, pitfalls); advisor spawns during discuss-phase |
| **Synthesizers** | gsd-research-synthesizer | Sequential (after researchers complete) |
| **Planners** | gsd-planner, gsd-roadmapper | Sequential |
| **Checkers** | gsd-plan-checker, gsd-integration-checker, gsd-ui-checker, gsd-nyquist-auditor | Sequential (verification loop, max 3 iterations) |
| **Executors** | gsd-executor | Parallel within waves, sequential across waves |
| **Verifiers** | gsd-verifier | Sequential (after all executors complete) |
| **Mappers** | gsd-codebase-mapper | 4 parallel (tech, arch, quality, concerns) |
| **Debuggers** | gsd-debugger | Sequential (interactive) |
| **Auditors** | gsd-ui-auditor, gsd-security-auditor | Sequential |
| **Doc Writers** | gsd-doc-writer, gsd-doc-verifier | Sequential (writer then verifier) |
| **Profilers** | gsd-user-profiler | Sequential |
| **Analyzers** | gsd-assumptions-analyzer | Sequential (during discuss-phase) |
### Wave Execution Model
During `execute-phase`, plans are grouped into dependency waves:
```
Wave Analysis:
Plan 01 (no deps) ─┐
Plan 02 (no deps) ─┤── Wave 1 (parallel)
Plan 03 (depends: 01) ─┤── Wave 2 (waits for Wave 1)
Plan 04 (depends: 02) ─┘
Plan 05 (depends: 03,04) ── Wave 3 (waits for Wave 2)
```
Each executor gets:
- Fresh 200K context window (or up to 1M for models that support it)
- The specific PLAN.md to execute
- Project context (PROJECT.md, STATE.md)
- Phase context (CONTEXT.md, RESEARCH.md if available)
### Adaptive Context Enrichment (1M Models)
When the context window is 500K+ tokens (1M-class models like Opus 4.6, Sonnet 4.6), subagent prompts are automatically enriched with additional context that would not fit in standard 200K windows:
- **Executor agents** receive prior wave SUMMARY.md files and the phase CONTEXT.md/RESEARCH.md, enabling cross-plan awareness within a phase
- **Verifier agents** receive all PLAN.md, SUMMARY.md, CONTEXT.md files plus REQUIREMENTS.md, enabling history-aware verification
The orchestrator reads `context_window` from config (`gsd-tools.cjs config-get context_window`) and conditionally includes richer context when the value is >= 500,000. For standard 200K windows, prompts use truncated versions with cache-friendly ordering to maximize context efficiency.
#### Parallel Commit Safety
When multiple executors run within the same wave, two mechanisms prevent conflicts:
1. `--no-verify` commits — Parallel agents skip pre-commit hooks (which can cause build lock contention, e.g., cargo lock fights in Rust projects). The orchestrator runs `git hook run pre-commit` once after each wave completes.
2. **STATE.md file locking** — All `writeStateMd()` calls use lockfile-based mutual exclusion (`STATE.md.lock` with `O_EXCL` atomic creation). This prevents the read-modify-write race condition where two agents read STATE.md, modify different fields, and the last writer overwrites the other's changes. Includes stale lock detection (10s timeout) and spin-wait with jitter.
#### The STATE.md Write Path
Locking decides *who* writes. A separate contract decides *what survives the write*.
STATE.md carries the same fact in two places — YAML frontmatter and the document body — and the body is authoritative. Every write therefore re-derives frontmatter from the body, which raises the question the write path exists to answer: when a re-derived value disagrees with the one already in frontmatter, which wins?
`FIELD_CLASSIFICATION` (`src/state-transition.cts`) answers it per field, declaring a `preservation` policy — `preserve-when-unchanged`, `preserve-always`, `preserve-if-placeholder`, `derive` — that `applyStatePreservation` executes after `syncStateFrontmatter` re-derives. (A fifth policy, `clear`, was listed here until ADR-3408 §8.6's amendment removed it: no row used it and no executor existed for it.)
**The pipeline's precondition is a type, not a convention (ADR-3473 §8.6).** A policy row can only be honored if the pre-write frontmatter snapshot it compares against is actually present. That snapshot now travels as a `StateTransaction`, built by `openStateTransaction()` — preservation applies — or `rebuildStateTransaction()` — it does not. Both carry the snapshot, and a transaction cannot be constructed without one: an absent snapshot is a *construction failure*, not a runtime skip. That distinction is the whole point. Previously the snapshot was nulled to signal "re-derive from disk", so a declared `preserve-always` row and a silently-skipped one were indistinguishable at runtime, which is how a curated `progress:` block was erased by verbs that had nothing to do with progress.
`rebuildStateTransaction()` is the typed form of ADR-3408 §8.3's closed exception list: `state sync`, which exists to let the body win, and `/gsd-health --repair`'s factory reset. Both are deliberate and permanent, not debt — and because the type names them, the write-path drift guard no longer has to track them as strings in a ratcheted baseline.
**What a command reports it wrote is the same snapshot, read back (ADR-3473 §8.7).** Every `state.*` command returns an `updated` array. That array is now derived by comparing what was actually persisted against the transaction's pre-write snapshot: **a field appears if, and only if, its persisted value changed.** One comparison answers both of the questions that used to need separate machinery — a field the caller asked for that the pipeline then discarded is persisted-equals-snapshot and drops out, and a field nobody asked for that the write moved anyway is different and appears. Nothing is filtered by its preservation policy.
Reporting is at **leaf granularity**: when a single counter moves you are told `progress.total_plans`, not `progress`. The leaves are the ones the field-classification table already declares, so the report is bounded by a schema rather than by walking the document.
One field is excluded, and it is excluded for its *provenance* rather than its policy: `last_updated` is stamped on every save regardless of what you changed, so admitting it would make `state patch`'s success signal — which is simply whether `updated` is non-empty — permanently true, and a patch in which every field failed would report success. `state_head` is deliberately **not** excluded: it is recomputed on every save but only *changes* when the commit it records actually moved, so reporting it tells you something true.
A practical consequence worth knowing: these arrays are now longer than they used to be, because they used to under-report. If you compare one exactly, expect more entries — and expect them to be the ones that really changed.
**[ADR-3408](adr/3408-state-write-path-preservation.md) is the normative contract** for that path: one executor per declared policy, one write seam, and reports computed from what was actually persisted rather than from what the caller intended to write. Where the contract and the code disagree, the code is the defect. It is the write-side counterpart of [ADR-3180](adr/3180-planning-semantic-model-single-owner.md), which gave each read-side derivation a single owner.
**[ADR-3473](adr/3473-enforcement-by-construction.md) owns the invariants that sit outside that contract.** ADR-3408 governs what survives a write; it does not govern the pipeline's precondition, what a command reports it wrote, or where the set of STATE.md keys, types and enums is declared. ADR-3473 owns those, alongside document parsing, enumeration, and the return contract of every routine that can fail. It is the third application of ADR-3180's mechanism and the first whose success metric requires the guard surface to *shrink* as each seam lands.
---
## Data Flow
### New Project Flow
```
User input (idea description)
│
▼
Questions (questioning.md philosophy)
│
▼
4x Project Researchers (parallel)
├── Stack → STACK.md
├── Features → FEATURES.md
├── Architecture → ARCHITECTURE.md
└── Pitfalls → PITFALLS.md
│
▼
Research Synthesizer → SUMMARY.md
│
▼
Requirements extraction → REQUIREMENTS.md
│
▼
Roadmapper → ROADMAP.md
│
▼
User approval → STATE.md initialized
```
### Phase Execution Flow
```
discuss-phase → CONTEXT.md (user preferences)
│
▼
ui-phase → UI-SPEC.md (design contract, optional)
│
▼
plan-phase
├── Research gate (blocks if RESEARCH.md has unresolved open questions)
├── Phase Researcher → RESEARCH.md
│ └── Package Legitimacy Gate: registry-API verdict on every package; [SLOP] removed,
│ [SUS]/[ASSUMED] flagged; Audit table written to RESEARCH.md
├── Planner (with reachability check) → PLAN.md files
│ └── checkpoint:human-verify injected before [ASSUMED]/[SUS] installs;
│ T-{phase}-SC STRIDE row added for install-bearing plans
├── Plan Checker → Verify loop (max 3x)
├── Requirements coverage gate (REQ-IDs → plans)
└── Decision coverage gate (CONTEXT.md `<decisions>` → plans, BLOCKING — #2492)
│
▼
state planned-phase → STATE.md (Planned/Ready to execute)
│
▼
execute-phase (context reduction: truncated prompts, cache-friendly ordering)
├── Wave analysis (dependency grouping)
├── Executor per plan → code + atomic commits
├── SUMMARY.md per plan
└── Verifier → VERIFICATION.md
└── Decision coverage gate (CONTEXT.md decisions → shipped artifacts, NON-BLOCKING — #2492)
│
▼
verify-work → UAT.md (user acceptance testing)
│
▼
ui-review → UI-REVIEW.md (visual audit, optional)
```
### Context Propagation
Each workflow stage produces artifacts that feed into subsequent stages:
```
PROJECT.md ────────────────────────────────────────────► All agents
REQUIREMENTS.md ───────────────────────────────────────► Planner, Verifier, Auditor
ROADMAP.md ────────────────────────────────────────────► Orchestrators
STATE.md ──────────────────────────────────────────────► All agents (decisions, blockers)
CONTEXT.md (per phase) ────────────────────────────────► Researcher, Planner, Executor
RESEARCH.md (per phase) ───────────────────────────────► Planner, Plan Checker
PLAN.md (per plan) ────────────────────────────────────► Executor, Plan Checker
SUMMARY.md (per plan) ─────────────────────────────────► Verifier, State tracking
UI-SPEC.md (per phase) ────────────────────────────────► Executor, UI Auditor
```
---
## File System Layout
### Installation Files
```
~/.claude/ # Claude Code (global install)
├── skills/gsd-ns-*/SKILL.md # Global skills — nesting runtimes: 6 namespace routers (authoritative roster: docs/INVENTORY.md)
│ └── skills/<name>/SKILL.md # concrete skills nested under each router
│ (flat runtimes: skills/gsd-*/SKILL.md — all ~67 skills at top level)
├── commands/gsd/*.md # Local Claude installs use slash commands instead of global skills
├── gsd-core/
│ ├── bin/gsd-tools.cjs # CLI utility
│ ├── bin/lib/*.cjs # Domain modules (authoritative roster: docs/INVENTORY.md)
│ ├── workflows/*.md # Workflow definitions (authoritative roster: docs/INVENTORY.md)
│ ├── references/*.md # Shared reference docs (authoritative roster: docs/INVENTORY.md)
│ └── templates/ # Planning artifact templates
├── agents/*.md # Agent definitions (authoritative roster: docs/INVENTORY.md)
├── hooks/*.js # Node.js hooks (statusline, guards, monitors, update check)
├── hooks/*.sh # Shell hooks (session state, commit validation, phase boundary)
├── settings.json # Hook registrations
└── VERSION # Installed version number
```
Equivalent paths for other runtimes:
- **OpenCode:** `~/.config/opencode/` global or `./.opencode/` local
- **Kilo:** `~/.config/kilo/` global or `./.kilo/` local
- **Kimi CLI:** first-existing generic global root (`~/.config/agents/` recommended, then `~/.agents/` if its `skills/` directory already exists); local install is deferred and guarded
- **Codex:** `~/.codex/` global or `./.codex/` local
- **Copilot:** `~/.copilot/` global or `./.github/` local
- **Antigravity:** auto-detected global root (`~/.gemini/antigravity/`, `~/.gemini/antigravity-ide/`, or `~/.gemini/antigravity-cli/`) for settings and runtime files; global skills/agents under `~/.gemini/config/` (the machine-local discovery dir, #3738) or `./.agent/` local
- **Cursor:** `~/.cursor/` global or `./.cursor/` local
- **Windsurf/Devin Desktop:** `~/.codeium/windsurf/` global config or `./.windsurf/` local workflows
- **Augment Code:** `~/.augment/` global or `./.augment/` local
- **Trae:** `~/.trae/` global or `./.trae/` local
- **Qwen Code:** `~/.qwen/` global or `./.qwen/` local
- **Hermes Agent:** `~/.hermes/` global or `./.hermes/` local
- **CodeBuddy:** `~/.codebuddy/` global or `./.codebuddy/` local
- **Cline:** `~/.cline/` global or project-root `.clinerules` local
### Project Files (`.planning/`)
```
.planning/
├── PROJECT.md # Project vision, constraints, decisions, evolution rules
├── REQUIREMENTS.md # Scoped requirements (v1/v2/out-of-scope)
├── ROADMAP.md # Phase breakdown with status tracking
├── STATE.md # Living memory: position, decisions, blockers, metrics
├── config.json # Workflow configuration
├── MILESTONES.md # Completed milestone archive
├── research/ # Domain research from /gsd-new-project
│ ├── SUMMARY.md
│ ├── STACK.md
│ ├── FEATURES.md
│ ├── ARCHITECTURE.md
│ └── PITFALLS.md
├── codebase/ # Brownfield mapping (from /gsd-map-codebase or /gsd-onboard)
├── onboarding/ # Brownfield onboarding summary (from /gsd-onboard)
│ ├── STACK.md # YAML frontmatter carries `last_mapped_commit`
│ ├── ARCHITECTURE.md # for the post-execute drift gate (#2003)
│ ├── CONVENTIONS.md
│ ├── CONCERNS.md
│ ├── STRUCTURE.md
│ ├── TESTING.md
│ └── INTEGRATIONS.md
├── phases/
│ └── XX-phase-name/
│ ├── XX-CONTEXT.md # User preferences (from discuss-phase)
│ ├── XX-RESEARCH.md # Ecosystem research (from plan-phase)
│ ├── XX-YY-PLAN.md # Execution plans
│ ├── XX-YY-SUMMARY.md # Execution outcomes
│ ├── XX-VERIFICATION.md # Post-execution verification
│ ├── XX-VALIDATION.md # Nyquist test coverage mapping
│ ├── XX-UI-SPEC.md # UI design contract (from ui-phase)
│ ├── XX-UI-REVIEW.md # Visual audit scores (from ui-review)
│ └── XX-UAT.md # User acceptance test results
├── quick/ # Quick task tracking
│ └── YYMMDD-xxx-slug/
│ ├── PLAN.md
│ └── SUMMARY.md
├── todos/
│ ├── pending/ # Captured ideas
│ └── completed/ # Completed todos
├── threads/ # Persistent context threads (from /gsd-thread)
├── seeds/ # Forward-looking ideas (from /gsd-capture --seed)
├── debug/ # Active debug sessions
│ ├── *.md # Active sessions
│ ├── resolved/ # Archived sessions
│ └── knowledge-base.md # Persistent debug learnings
├── ui-reviews/ # Screenshots from /gsd-ui-review (gitignored)
└── continue-here.md # Context handoff (from pause-work)
```
### Post-Execute Codebase Drift Gate (#2003)
After the last wave of `/gsd-execute-phase` commits, the workflow runs a
non-blocking `codebase_drift_gate` step (between `schema_drift_gate` and
`verify_phase_goal`). It compares the diff `last_mapped_commit..HEAD`
against `.planning/codebase/STRUCTURE.md` and counts four kinds of
structural elements:
1. New directories outside mapped paths
2. New barrel exports at `(packages|apps)/<name>/src/index.*`
3. New migration files
4. New route modules under `routes/` or `api/`
If the count meets `workflow.drift_threshold` (default 3), the gate either
**warns** (default) with the suggested `/gsd-map-codebase --paths …` command,
or **auto-remaps** (`workflow.drift_action = auto-remap`) by spawning
`gsd-codebase-mapper` scoped to the affected paths. Any error in detection
or remap is logged and the phase continues — drift detection cannot fail
verification.
`last_mapped_commit` lives in YAML frontmatter at the top of each
`.planning/codebase/*.md` file; `bin/lib/drift.cjs` provides
`readMappedCommit` and `writeMappedCommit` round-trip helpers.
The baseline is written by `gsd-tools stamp-codebase-map`, a shell step in the
map-codebase workflow, not by the mapper agent. The mapper's own freshness
markers (`**Analysis Date:**`, `<!-- refreshed: ... -->`) are restamped
unconditionally on an Update run, so an agent that rewrites only the dates still
looks current to a reader; the machine-readable stamp is the one marker that
cannot be satisfied by a date-only rewrite, which is exactly why it is not the
agent's to write. `--files a.md,b.md` narrows the stamp to the documents a
caller actually refreshed, as the auto-remap path does.
An absent or unresolvable baseline is reported as `skipped` with reason
`no-mapped-commit` or `unresolvable-mapped-commit`, never as drift. Diffing
HEAD against the empty tree would report every tracked file as newly added,
which makes a stale map indistinguishable from a fresh one. Files under
`.planning/` are excluded from the diff: the map's own commit is a planning
artifact, not codebase structure.
---
## Installer Architecture
The installer (`bin/install.js`, ~10,700 lines) handles:
1. **Runtime detection** — Interactive prompt or CLI flags (`--claude`, `--opencode`, `--kimi`, `--kilo`, `--codex`, `--copilot`, `--antigravity`, `--cursor`, `--windsurf`, `--augment`, `--trae`, `--qwen`, `--hermes`, `--codebuddy`, `--cline`, `--all`)
2. **Location selection** — Global (`--global`) or local (`--local`)
3. **File deployment** — Copies commands, skills, workflows, references, templates, agents, and hooks
4. **Runtime adaptation** — Transforms file content per runtime:
- Claude Code: Uses as-is
- OpenCode: Converts commands/agents to OpenCode-compatible flat command + subagent format
- Kilo: Reuses the OpenCode conversion pipeline with Kilo config paths
- Codex: Generates TOML config + skills from commands
- Kimi CLI: Generates Agent Skills under `skills/gsd-*/SKILL.md`, custom agent YAML/prompt files, and explicit `kimi_cli.tools.*` module paths
- Copilot: Maps tool names (Read→read, Bash→execute, etc.)
- Antigravity: Skills-first with Google model equivalents; adjusts hook event names (`AfterTool` instead of `PostToolUse`)
- Cursor: Skills-first with Cursor rule references
- Windsurf: Skills-first with Windsurf rule references
- Trae: Skills-first install to `~/.trae` / `./.trae` with no `settings.json` or hook integration
- Qwen Code: Skills-first with Qwen-branded path and prompt rewrites
- Hermes Agent: Category-based skills under `skills/gsd/`
- CodeBuddy: Skills-first with CodeBuddy path and prompt rewrites
- Cline: Writes `.clinerules` for rule-based integration
- Augment Code: Skills-first with full skill conversion and config management
5. **Path normalization** — Replaces `~/.claude/` paths with runtime-specific paths
6. **Settings integration** — Registers hooks in runtime's `settings.json`
7. **Patch backup** — Since v1.17, backs up locally modified files to `gsd-local-patches/` for `/gsd-update --reapply`
8. **Manifest tracking** — Writes `gsd-file-manifest.json` for clean uninstall. The manifest also records which `runtime` and which `scope` (`global`/`local`) wrote it, under a `manifestVersion` schema field, so a reader can answer "which surfaces are installed, at which scopes" without inferring it from the directory the file sits in ([ADR 2866](adr/2866-install-surface-resolution.md), #2872). Manifests written before that carry no such fields and are read without error — no reinstall is required. See [Installer Migrations → File Manifest](installer-migrations.md#file-manifest)
9. **Uninstall mode** — `--uninstall` removes all GSD files, hooks, and settings
`installRuntimeArtifacts` (`install-engine.cjs`) returns the executed plan it ran — per kind, per
scope, including on the combined OpenCode/Kilo family path, which previously early-returned `void` —
rather than being observable only by re-reading disk afterward. Its destination-writing IO (copies,
removals, snapshot/restore, best-effort cleanup) now routes through an injectable fs seam,
`install-fs-adapter.cjs`, so a full install can be exercised against a fake adapter with zero real
destination IO; locating this package's own source tree remains real by design (a destination-fake
is never seeded with the repo's own paths). Writes stay byte-identical and existing `void`-ignoring
callers are unaffected. This completes [ADR 58](adr/58-runtime-install-policy-module.md)'s
`registry → adapter → helpers → cleanup` rollout — the `cleanup` step had not previously landed
(#2874, epic #2866 Phase 5).
Install-time file moves, stale-artifact cleanup, config rewrites, and user-data
preservation are governed by the Installer Migration Module. See
[Installer Migrations](installer-migrations.md) and
[ADR 0008](adr/0008-installer-migration-module.md).
The migration module also owns the gated first-time baseline scan for legacy
installs, classifying known runtime install surfaces before later migrations
remove or rewrite anything.
The plan drift guard (`plan_review.source_grounding`) — which verifies symbol references in generated plans against live source before execution — is specified in [ADR 22](adr/22-plan-drift-guard.md).
The same switch gates a second, cross-artifact axis: a fact-drift pass that compares the *same* fact as stated in `ROADMAP.md`, `PLAN.md`, `STATE.md` and `CONTEXT.md` and reports contradictions (a phase status, a success criterion, a requirement ID, a glossary term) with both locations and the authoritative side named. Where the source-grounding axis grounds a plan against code, this one grounds the planning artifacts against each other. It keys on contradicting knowledge rather than similar-looking text, and is advisory only — it never sets `hardBlock` and never contributes to the convergence counts.
### Platform Handling
- **Windows:** `windowsHide` on child processes, EPERM/EACCES protection on protected directories, path separator normalization
- **WSL:** Detects Windows Node.js running on WSL and warns about path mismatches
- **Docker/CI:** Supports `CLAUDE_CONFIG_DIR` env var for custom config directory locations
---
## Hook System
### Architecture
```
Runtime Engine (Claude Code / Antigravity CLI)
│
├── statusLine event ──► gsd-statusline.js
│ Reads: stdin (session JSON)
│ Writes: stdout (formatted status), /tmp/claude-ctx-{session}.json (bridge)
│
├── PostToolUse/AfterTool event ──► gsd-context-monitor.js
│ Reads: stdin (tool event JSON), /tmp/claude-ctx-{session}.json (bridge)
│ Writes: stdout (hookSpecificOutput with additionalContext warning)
│
└── SessionStart event
├──► gsd-ensure-canonical-path.js (runs first)
│ Reads: ${CLAUDE_PLUGIN_ROOT}/gsd-core/ (plugin installs only)
│ Writes: ~/.claude/gsd-core/{bin,contexts,references,templates,workflows} symlinks
│ (no-op in classic installs; preserves user files; self-heals)
└──► gsd-check-update.js
Reads: VERSION file
Writes: ~/.claude/cache/gsd-update-check.json (spawns background process)
```
### Context Monitor Thresholds
| Remaining Context | Level | Agent Behavior |
| ----------------- | -------- | --------------------------------------- |
| > 35% | Normal | No warning injected |
| ≤ 35% | WARNING | "Avoid starting new complex work" |
| ≤ 25% | CRITICAL | "Context nearly exhausted, inform user" |
The two fire-points are defaults. `hooks.context_warning_threshold` and
`hooks.context_critical_threshold` in `.planning/config.json` move them per
project; see [context-monitor.md](context-monitor.md) for the resolution and
fallback rules.
Debounce: 5 tool uses between repeated warnings. Severity escalation (WARNING→CRITICAL) bypasses debounce.
### Safety Properties
- All hooks wrap in try/catch, exit silently on error
- stdin timeout guard (3s) prevents hanging on pipe issues
- Stale metrics (>60s old) are ignored
- Missing bridge files handled gracefully (subagents, fresh sessions)
- Context monitor is advisory — never issues imperative commands that override user preferences
### Package Legitimacy Gate (v1.42.1)
The researcher → planner → executor pipeline includes a supply-chain gate against slopsquatting (AI-hallucinated package names pre-registered with malicious post-install scripts).
**Threat model:** GSD automates the full path from "researcher names a package" to "executor runs `npm install`". A hallucinated name that passes `npm view` (proving only registration, not legitimacy) would previously flow through undetected. ~20% of AI-generated package references are hallucinated; ~43% of those names recur consistently across prompts, making pre-registration economically viable for attackers.
**Gate layers:**
| Layer | Component | Action |
|-------|-----------|--------|
| Research | `gsd-phase-researcher` | Runs `gsd-tools query package-legitimacy check --ecosystem <npm\|pypi\|crates> <pkgs>`; writes `## Package Legitimacy Audit` table to RESEARCH.md; strips `[SLOP]` packages before RESEARCH.md is written |
| Planning | `gsd-planner` | Reads Audit table; inserts `checkpoint:human-verify` before any `[ASSUMED]` or `[SUS]` install task; adds `T-{phase}-SC` STRIDE supply-chain row to `<threat_model>` |
| Execution | `gsd-executor` | RULE 3 excludes package installation from auto-fix scope; failed installs surface as checkpoints, never silent substitutions |
**Claim provenance integration:** Package names discovered via WebSearch are tagged `[ASSUMED]` (not `[VERIFIED]`) regardless of the registry-API verdict. This extends the existing `[ASSUMED]` / `[VERIFIED]` / `[CITED]` provenance system by enforcing the provenance tag as a hard gate at the install boundary — `[ASSUMED]` always generates a `checkpoint:human-verify` in PLAN.md.
**Ecosystem coverage:** The gate resolves signals directly from each ecosystem's registry API rather than a single generic check — `registry.npmjs.org` + `api.npmjs.org/downloads` (Node), `pypi.org/pypi/<pkg>/json` (Python), the crates.io API (Rust). This catches cross-ecosystem hallucination (~9% rate documented in 2025 USENIX research).
**Graceful degradation:** Each registry adapter degrades to null signals (never throws) on a failed lookup; missing signals push a package to `[SUS]`, which is gated behind the same `checkpoint:human-verify` checkpoint as `[ASSUMED]`. Research and planning proceed; the system never hard-fails on a network or tool outage. `slopcheck` is an optional escalate-only adapter — it can only raise a verdict, never lower it, and is not the install-or-degrade gate. No shipped configuration wires it.
---
### Security Hooks (v1.27)
For a conceptual overview of how the hook and guard layers fit into the broader security approach, see [Security model](explanation/security-model.md).
**Prompt Guard** (`gsd-prompt-guard.js`):
- Triggers on Write/Edit to `.planning/` files
- Scans content for prompt injection patterns (role override, instruction bypass, system tag injection)
- Advisory-only — logs detection, does not block
- Patterns are inlined (subset of `security.cjs`) for hook independence
- **Output contract:** `hookSpecificOutput` carries both `additionalContext` and `findings` — an array of `{ ruleId, match }` records (`INJECTION-PATTERN` or `INVISIBLE-UNICODE`), module-local to this hook (not shared with `gsd-read-injection-scanner.js`'s own `RULE_IDS`). The advisory is rendered from `findings` via a single mapper, so the two cannot disagree. Consumers should read `findings` rather than parsing the advisory text.
**Read Injection Scanner** (`gsd-read-injection-scanner.js`):
- Triggers on `Read` / `WebFetch` / `WebSearch` PostToolUse events
- Advisory by default; blocks only `HIGH` severity, and only when `security.injection_blocking` is `true`
- Severity is `LOW` for 1-2 matched patterns, `HIGH` for 3 or more
- Skips content shorter than 20 characters, and skips excluded paths (`.planning/`, `REVIEW.md`, `CHECKPOINT*`, security/injection docs, and GSD's own staged hook bundle)
- Rule ids: the `MD-LINK-*` markdown-link rules mirrored from `security.cjs`'s `MARKDOWN_LINK_PATTERNS`, plus `INJECTION-PATTERN`, `INVISIBLE-UNICODE`, and `UNICODE-TAG-BLOCK`
- Patterns are shared with `gsd-prompt-guard.js` via `hooks/lib/injection-patterns.js` (#3504); the markdown-link list is inlined for hook independence
- **Output contract:** `hookSpecificOutput` carries `additionalContext` (the human-readable advisory sentence), `findings` — an array of `{ ruleId, match }` records naming each rule that fired — plus `severity` (`LOW` for 1-2 matches, `HIGH` for 3+) and `source` (the scanned file path, URL, or `search: <query>` string). `findings` and `severity` are the structured surface; the advisory is rendered from them, so the three cannot disagree. `match` is `null` for rules with no captured text (`INVISIBLE-UNICODE`, `UNICODE-TAG-BLOCK`). Consumers should read `findings`/`severity`/`source` rather than parsing the advisory text.
**Workflow Guard** (`gsd-workflow-guard.js`):
- Triggers on Write/Edit to non-`.planning/` files
- Detects edits outside GSD workflow context (no active `/gsd-` command or Task subagent)
- Advises using `/gsd-quick` or `/gsd-fast` for state-tracked changes
- Opt-in via `hooks.workflow_guard: true` (default: false)
- **Output contract:** the advisory leg's `hookSpecificOutput` carries `code: 'WORKFLOW_ADVISORY'` alongside `additionalContext`. This is distinct from the hook's separate force-add block leg (`code: 'WORKTREE_AGENT_FORCE_ADD_FORBIDDEN'`, `decision: 'block'`) — the two are disambiguated by `code`, never by presence.
---
## Runtime Abstraction
GSD supports multiple AI coding runtimes through a unified command/workflow architecture:
### Runtime Install Contract Matrix
This matrix describes the runtime surfaces the installer materializes today.
The migration-specific ownership and source snapshots live in
[Installer Migrations](installer-migrations.md#runtime-configuration-contract-registry).
| Runtime | Global root | Local root | Invocation surface | Agent surface | Config and hooks |
| --- | --- | --- | --- | --- | --- |
| Claude Code | `~/.claude` | `./.claude` | Global `skills/gsd-*/SKILL.md` (flat, #924); local `commands/gsd/*.md` | `agents/gsd-*.md` | `settings.json` hook and statusLine entries |
| OpenCode | `~/.config/opencode` | `./.opencode` | `commands/gsd-*.md` | `agents/gsd-*.md` | `opencode.json` or `opencode.jsonc`; no GSD hooks |
| Kilo | `~/.config/kilo` | `./.kilo` | `command/gsd-*.md` | `agents/gsd-*.md` | `kilo.json` or `kilo.jsonc`; no GSD hooks |
| Kimi CLI | First-existing generic root: `~/.config/agents` recommended, then `~/.agents` when `~/.agents/skills` exists and `~/.config/agents/skills` does not | Deferred and guarded | `skills/gsd-*/SKILL.md` (flat) invoked as `/skill:gsd-*` | `agents/gsd.yaml`, `agents/gsd.md`, and `agents/subagents/gsd-*` YAML/prompt pairs | Explicit `kimi --agent-file <configRoot>/agents/gsd.yaml`; no GSD hooks or statusline |
| Codex | `~/.codex` | `./.codex` | `skills/gsd-*/SKILL.md` (flat) | `agents/` source markdown plus per-agent TOML (Codex auto-discovers each `agents/gsd-*.toml`; this is the sole canonical role registration, #2406) | `config.toml` bare `[agents]` dispatch-tuning scalar (`max_depth`, no per-role `[agents.gsd-*]` tables), `[features].hooks` (canonical; legacy alias `codex_hooks` is recognized and migrated forward on reinstall, #3566), and hook tables |
| GitHub Copilot | `~/.copilot` | `./.github` | `skills/gsd-*/SKILL.md` (flat), `copilot-instructions.md`, and `AGENTS.md` (repo root, local) | `.agent.md` files | Self-contained `sessionStart` hook (`hooks/gsd-session.json`, inline `command` type); no statusline |
| Antigravity | auto-detected: `~/.gemini/antigravity`, `~/.gemini/antigravity-ide`, or `~/.gemini/antigravity-cli` | `./.agent` | `~/.gemini/config/skills/gsd-*/SKILL.md` (flat, #1614; global home override #3738) | `~/.gemini/config/agents/gsd-*.md` (#3738) | Gemini-style `settings.json` hook entries when installed by GSD |
| Cursor | `~/.cursor` | `./.cursor` | `skills/gsd-*/SKILL.md` (flat) | `agents/gsd-*.md` | Rule references under `rules/`; `hooks.json` with sessionStart context injection and postToolUse STATE.md monitor (#777) |
| Windsurf | `~/.codeium/windsurf` config | `./.windsurf` | `workflows/gsd-*.md` slash-command workflows | No custom-agent artifact surface | No GSD hooks |
| Augment Code | `~/.augment` | `./.augment` | `skills/gsd-ns-*/SKILL.md` (6 routers) + `skills/gsd-ns-*/skills/<name>/SKILL.md` (nested concretes) | `agents/gsd-*.md` | No GSD hooks or statusline |
| Trae | `~/.trae` | `./.trae` | `skills/gsd-ns-*/SKILL.md` (6 routers) + `skills/gsd-ns-*/skills/<name>/SKILL.md` (nested concretes) | `agents/gsd-*.md` | Rule references under `rules/`; no GSD hooks |
| Qwen Code | `~/.qwen` | `./.qwen` | `skills/gsd-ns-*/SKILL.md` (6 routers) + `skills/gsd-ns-*/skills/<name>/SKILL.md` (nested concretes) | `agents/gsd-*.md` | Common GSD settings and hook entries where supported |
| Hermes Agent | `~/.hermes` | `./.hermes` | `skills/gsd/ns-*/SKILL.md` (6 routers, prefix='') + `skills/gsd/ns-*/skills/<name>/SKILL.md` (nested concretes) | `agents/gsd-*.md` | Common GSD settings and hook entries where supported |
| CodeBuddy | `~/.codebuddy` | `./.codebuddy` | `skills/gsd-*/SKILL.md` (flat, `user-invocable: false`) | `agents/gsd-*.md` | `/gsd-*` slash commands under `commands/`; common GSD settings and hook entries where supported |
| Cline | `~/.cline` | project root | `skills/gsd-ns-*/SKILL.md` (6 routers) + `skills/gsd-ns-*/skills/<name>/SKILL.md` (nested concretes) + `.clinerules` | Rules only | No GSD hooks or statusline |
### Upstream Contract Sources
Runtime install expectations are checked against primary documentation where
available. The current source snapshot is 2026-05-11, with Kimi CLI rechecked
on 2026-06-07:
- Claude Code: Anthropic slash commands, settings, hooks, and subagents docs.
- OpenCode and Kilo: OpenCode config docs and Kilo custom subagent docs.
- Qwen Code: command/config docs; Qwen command docs were last
updated 2026-05-06.
- Kimi CLI: Agent Skills docs for user-level brand roots and first-existing
generic roots (`~/.config/agents/skills/` recommended, then
`~/.agents/skills/`), plus Agents docs for YAML files, `system_prompt_path`,
`kimi_cli.tools.*` module paths, and explicit `kimi --agent-file` launch.
- Codex: OpenAI Codex docs and `config-schema.json`; the installer also carries
Codex 0.124.0 compatibility for agent table shape.
- Copilot, Cursor, Cline, Augment, Hermes, and CodeBuddy: vendor docs for
custom instructions, rules, skills, or config.
- Antigravity, Windsurf, and Trae: source-limited rows. The installer documents
current compatibility shims, and migrations must refresh those sources before
rewriting their config.
### Abstraction Points
1. **Tool name mapping** — Each runtime has its own tool names (e.g., Claude's `Bash` → Copilot's `execute`)
2. **Hook event names** — Claude uses `PostToolUse`, Antigravity uses `AfterTool`
3. **Agent frontmatter** — Each runtime has its own agent definition format
4. **Path conventions** — Each runtime stores config in different directories
5. **Model references** — `inherit` profile lets GSD defer to runtime's model selection
The installer handles all translation at install time. Workflows and agents are written in Claude Code's native format and transformed during deployment.
---
## Related
- [Multi-agent orchestration](explanation/multi-agent-orchestration.md)
- [Security model](explanation/security-model.md)
- [CLI tools](CLI-TOOLS.md)
- [docs index](README.md)