A GSD verification gate resolves a project's test command and runs it. vitest
defaults to WATCH mode in an interactive TTY — exactly where a user runs
`gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never
exited and the orchestrator waited indefinitely. Recovery needed the user to
manually prompt "something blocking?".
Fix — one shared helper + a bounded, surfacing timeout on the test-command gates:
- New pure module src/normalize-test-command.cts + `gsd-tools query
normalize-test-command` verb: rewrites a resolved command to a best-effort
one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`;
a package-manager `test` script whose package.json runner is watch-vitest →
`CI=true` prefix; handles `--dir`; already-one-shot commands unchanged —
never double-flagged). Named `normalize-test-command` (not `test-*`) so the
file does not match node --test's default `test-*` discovery glob.
- The three gates that HUNG or silently-continued route through that ONE helper
and bound execution with `timeout $(config-get workflow.test_gate_timeout)`
(new config key, default 600s): the regression gate (extracted to
execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen —
it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the
audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124.
- verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so
it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint
on 124, staying under its frozen 40960-byte tier cap.
Security hardening (review): the normalizer only rewrites a runner named as a
standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never
mangled), is length-capped and uses only linear-time split-based scanning (no
super-linear backtracking on an adversarial `workflow.test_command`), and reads
package.json only when it is a regular file (never blocks on a FIFO via `--dir`).
Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the
schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors
workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores,
inventory manifest/index. All 16 golden-install-parity fixtures + workflow size
baseline regenerated for the changed shipped files; bin/lib is excluded from the
parity manifest.
Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security
hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route
through the shared helper + configured timeout + exit-124 watch-mode hint;
verify-phase asserted as normalize-only/already-bounded).
tests/execute-phase-active-flags.test.cjs repointed at the extracted step;
tests/planner-language-regression.test.cjs allowlist comment updated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A skills-only role:feature third-party capability installed 'active' but its skills never reached the runtime surface, capability enable/set rejected it as 'unknown capability', and capability list disagreed with capability state. Three defects, fix shape 1b (teach resolveSurface, no on-disk linking):
D1 (materialization): resolveSurface built the surfaced skills Set only from the on-disk manifest; third-party cap skills live at ~/.gsd/capabilities/<id>/skills/ and never entered the Set -> surfaced:false. Fix: union registry.capabilityClusters values into the Set in the full-profile branch (idempotent for first-party, additive for third-party; prototype-pollution guarded).
D2 (enable/set unknown): setCapabilityState validated the capId against the static first-party registry instead of the composed overlay-aware registry. Fix: validate against loadRegistry({includeInstalled}) like capability-state.cts does.
D3 (list vs state): capability list derived status purely from ledger existence, never surface composition. Fix: add a 'surfaced' field to list rows sourced from the same resolver capability state uses, so list and state agree.
Regression coverage: tests/issue-2045-third-party-skills-surface.test.cjs asserts all five acceptance criteria (resolveSurface union, state surfaced:true, enable/set not-unknown, list/state agreement, first-party + unknown-id regression). gsd-test: 23916/23916 green (linux-node22 + linux-node24).
- warn (don't silently ignore) when --runtime is an unknown runtime that
canonicalizeRuntimeName rejects; the warning surfaces via warnings[] so a
typo like --runtime cluade or a runtime known to runtime-homes but not the
alias manifest (e.g. grok) no longer silently resolves to the persisted
runtime's config dir on this diagnostic command [M-1]
- add end-to-end CLI test for loop render-hooks --runtime (the exact command
the bug report calls out as silently no-op'ing) [L-2]
- add closed-vocabulary rejection test: crafted --runtime values
(../../etc/passwd, __proto__, --config-dir, garbage) are rejected, warn,
and fall through to the persisted runtime — pins the security-load-bearing
contract [NIT-01]
- add boundary tests: --config-dir wins over --runtime (precedence); missing
--runtime value errors with USAGE [N-1]
Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed --runtime cannot coerce getGlobalConfigDir into an
arbitrary path (closed-vocabulary Map lookup + registry hash-key gate) and
does not expand the trust surface beyond the existing operator-controlled
--config-dir flag.
The --runtime parsing + help-text edit to gsd-core/bin/gsd-tools.cjs changes
the installed file's content (gsd-tools.cjs is installed and compared by the
golden snapshot, unlike gsd-core/bin/lib/ which is excluded). Regenerated via
UPDATE_GOLDEN=1; every runtime's manifest updates exactly one line (the
gsd-tools.cjs hash).
resolveCapabilityRuntimeState derived the config dir from resolveRuntime(cwd)
(GSD_RUNTIME -> config.runtime -> 'claude') when no --config-dir was passed,
so a repo with persisted runtime:'codex' resolved the config dir to ~/.codex
where the Claude skill isn't installed -> surfaced:false / hooks silently
no-op when the operator drove from Claude Code. capability state and loop
render-hooks parsed only --config-dir, never --runtime, so there was no way
to assert the actually-active runtime.
Add a runtimeOverride param to resolveCapabilityRuntimeState (canonicalized
via runtime-name-policy so aliases like codex-app work); when present it
short-circuits the persisted-runtime fallback and resolves getGlobalConfigDir
for the explicit runtime. Thread --runtime through cmdCapabilityState and
cmdLoopRenderHooks, and parse it in gsd-tools.cjs for both commands (dual
--runtime X / --runtime=X form, mirroring --config-dir and the existing
capability-set --runtime precedent). Help text updated.
Without the override, behavior is byte-identical to today (regression-guarded).
Mirrors the #1160 installed-layout block for the runtime auto-detection gap.
Covers: runtimeOverride='claude' bypasses persisted config.runtime:'codex';
no override still honours persisted runtime (regression guard); alias
canonicalization (codex-app -> codex); and an end-to-end CLI test proving
'capability state --runtime claude' resolves the Claude config dir despite a
persisted runtime:'codex'.
Expected RED against unfixed resolveCapabilityRuntimeState (no runtimeOverride
param) and unfixed gsd-tools.cjs (no --runtime parsing for capability state /
loop render-hooks).
- restructure _loadFlatCommandsGsdManifest try/catch to wrap read+parse+set
together (mirrors loadSkillsManifest exactly), so a thrown parser degrades
both keys to [] — closes the latent catch-scope parity drift [Nit-1]
- add boundary test: gsd-.md (empty stem) skipped, gsd-x.md single-char stem
kept (slice(4,-3) boundary) [Low-1]
- add unreadable-file test (POSIX-gated): both keys degrade to [] [Low-2]
- strengthen parity test to also compare requires + _calls_agents_ VALUES,
not just the stem set [Nit-2]
Both orthogonal reviews returned APPROVE with no Critical/High findings.
Security review confirmed no new trust-boundary crossing, prototype-pollution
immune (Map/Set throughout), and symlink/path-traversal surface identical to
the pre-existing nested loader (not a regression).
_resolveManifest only recognized the nested source layout (commands/gsd/*.md)
and the installed-runtime skills layout (skills/gsd-<stem>/SKILL.md). A flat
source install (Claude local project shape: commands/gsd-<stem>.md, no
commands/gsd/ subdir) matched neither branch, so the manifest came back empty
and resolveSurface materialized the full profile to an empty Set — silently
reporting every skill-bearing capability as surfaced:false / enabled:false /
active:false. The nyquist/code-review/security/ui verify:post and execute:post
hooks never fired even with their workflow.* toggles on.
Add a third branch: when commandsGsdDir is absent, scan dirname(commandsGsdDir)
for gsd-<stem>.md files, strip the gsd- prefix, and build the same Map shape
the nested loader produces (requires via shared parseRequires, companion
_calls_agents_<stem> via shared parseCallsAgents). Falls through to the
installed-skills branch when the flat dir has no gsd-*.md files (precedence:
nested > flat-source > installed).
Also export parseCallsAgents from install-profiles so capability-state reuses
the SAME parser the nested loader uses (no drift; mirrors the existing
parseRequires export+reuse pattern).
Mirrors the #1160 installed-layout tests for the flat source layout
(<repo>/commands/gsd-<stem>.md, no commands/gsd/ subdir). Covers stem
extraction (strip gsd- prefix), requires parsing via shared parseRequires,
companion _calls_agents_ key parity, _resolveManifest flat-branch detection,
precedence (flat-empty falls through to installed), and a generative-parity
assertion that the flat loader and nested loader produce identical stem sets
for the real command tree.
Expected RED against unfixed capability-state.cts (_resolveManifest has no
flat branch; _loadFlatCommandsGsdManifest not exported).
- add typeof guard so a non-string override passes through verbatim instead of
crashing on .startsWith (preserves pre-fix no-crash behaviour) [LOW-1]
- use Object.hasOwn() for the alias lookup so __proto__/constructor cannot
return a truthy non-string from the plain object literal [LOW-D3]
- cap the unmappable-override stderr warning at 64 chars so an oversized or
secret-shaped value cannot leak in full to stderr/logs [LOW-D4]
- remove the unused mapClaudeOverrideForRuntime export (helpers are covered
behaviourally via resolveModelInternal/resolveModelForTier) [NIT]
- add resolveModelForTier unmappable-override fall-through test (closes the
mutation-score gap) [MEDIUM-1]
- add case-sensitivity contract test (Claude-Sonnet-5 passes through verbatim) [LOW-2]
Both orthogonal reviews returned APPROVE with no Critical/High findings.
model_overrides values that are full Claude model IDs (claude-sonnet-5,
claude-opus-4-8, claude-haiku-4-5, claude-fable-5) were returned verbatim on
the claude runtime and handed to the Claude Agent tool, whose typed model
parameter documents only tier aliases (opus/sonnet/haiku/fable). The
model_policy path already mapped full IDs -> aliases via
CLAUDE_POLICY_ID_TO_ALIAS (#1144); model_overrides skipped that mapping, so
the two resolver paths produced different shapes for the same underlying
Claude model. The fix mirrors #1144 on the override path via a shared
mapClaudeOverrideForRuntime helper used by both resolveModelInternal and
resolveModelForTier. Bare aliases pass through verbatim; non-Claude runtimes
and non-Claude custom/vendor values keep full IDs verbatim (parity). An
unmappable Claude ID (e.g. claude-opus-4-5) warns once to stderr and falls
through to tier resolution, exactly as the model_policy path already does.
Alias mapping is also the documented best practice (prevents staleness when
new model versions ship).
Mirrors the #1133 model_policy alias-mapping tests for the model_overrides
path. Covers AC1-AC6: mappable Claude full IDs (claude-sonnet-5/opus-4-8/
haiku-4-5/fable-5) resolve to aliases on runtime:claude; bare aliases pass
through; non-claude runtimes keep full IDs verbatim; unmappable Claude IDs
warn-once + fall through; resolveModelForTier escalation path also maps;
non-Claude custom/vendor values pass through verbatim (regression guards).
Expected RED against unfixed model-resolver.cts (override short-circuit at
lines 162-167 / 288-290 returns override verbatim with no alias mapping).
- compareSemver: implement full SemVer 2.0.0 §11 pre-release identifier
comparison (two pre-releases of the same triple now order correctly; was 0).
- capability description + fragment: scope the plan-checker/verifier claim
(this capability delivers the parallel-execution backend; those gates remain
inline until separately wired). Correct the 'each wave is one barrier' prose
(a wave splits into multiple sequential parallel() barriers on files_modified
overlap). Frame detect-backend CLI as a simulation harness; the pure function
with the live host descriptor is the real detection seam.
- partitionStages docstring: 'near-minimal via greedy first-fit' (not 'fewest');
document empty-files_modified behavior.
Documents the new architectural surface #1820 introduces, per the
contributor-standards ADR requirement (a new Module seam that other code
will depend on):
- The spec-section detection Module seam (src/spec-section.cts) and its
locked exported surface, supply rule, suffix-tolerant header invariant,
and ownership boundary (detection only).
- The workflow.specless_probe_fallback toggle as a policy decision
(default-on, disableable cost-gate over the fallback INVOCATION path, not
the verifier<->predicate contract) — records the maintainer 857:66 ruling
rather than amending it.
- The SPEC-supplied <-> probe-derived precedence & authoring contract:
section-level precedence (a SPEC-supplied section is never re-run), one
projectProhibitions serializer (no second producer), descriptor-less
fallback predicates flag/abstain (never green, never auto-dismissed),
no-silent-drop equality.
Does not restate ADR-857/550/1606; cross-references them. Resolves the
sole remaining review blocker on #1835.
Refs #1820
Claude-Session: https://claude.ai/code/session_017vYn26e3nkDNxcpty1ciPJ
Addresses trek-e's R1 blocker: src/spec-section.cts is a new named module
but had no ## Domain terms entry in CONTEXT.md (INVENTORY.md had one).
Adds the ### Spec-Section Helper Module paragraph covering the owned
surface (SpecSectionKey, SECTION_HEADERS, SectionStatus, countSectionDataRows,
specSectionStatus), the suffix-tolerant header matching invariant, the
supplied = present AND dataRows > 0 rule, and the source-of-truth note
(gitignored per ADR-457). Maintainer owns final wording.