Commit Graph

18 Commits

Author SHA1 Message Date
0xdhx
f0ef9063d5 test(#2665): restore the three agent-skills tests this PR deleted
Commit 2bed9fd8 ("replace the hand-synced TEST_ENV_BASE copies with the
canonical import") also removed markLocalGsdInstall and three behavioural tests
from tests/agent-skills.test.cjs -- 79 lines, no replacement, and no mention in
the commit message or the PR body:

  - unconfigured Codex reads its local companion agent from a descendant cwd
  - workstream runtime selects the local Codex companion when root config differs
  - unconfigured Claude remains empty when a local Codex companion exists

RULESET.TESTS.delete-bad-tests permits deleting a bad test only when it is
replaced with compliant tests in the same PR. Nothing was replaced, and these
were not bad tests -- they were collateral in a mechanical edit. A silent net
loss of behavioural coverage inside a PR whose subject is test hygiene is the
one thing that should not pass here, and the reviewer was right to block on it.

Restored verbatim. They need no adaptation to the canonical TEST_ENV_BASE
import: they pass their env explicitly, and an explicit env still spreads last
over the base. The file goes 80 -> 83 tests, all green.

Non-vacuity checked rather than assumed: dropping the local-install marker the
first two depend on fails both. The third is a negative assertion and correctly
stays green, which is why it is named here rather than counted as covered.

Addresses review finding: Blocker 1.
2026-08-08 05:50:49 -05:00
0xdhx
bc5c362610 fix(#2665): replace the hand-synced TEST_ENV_BASE copies with the canonical import
Eight declarations were kept in sync by hand with no parity assertion. The drift
was already in the tree: api-coverage-gate-e2e and representative-corpus declared
TERM_SESSION, but the real variable is TERM_SESSION_ID, so both scrubbed nothing
for that slot and leaked TERM_SESSION_ID into every child. Deleting the copies
removes the dead key with them -- there is no longer a second place to get wrong.

run-tests-harness keeps its local session-identity literal: that helper mirrors
the production runGsdTools to prove its CONTRACT, so it must not re-import what
it is testing. The config-LOCATION keys are a safety scrub rather than part of
that contract, so it spreads the canonical derived set and keeps the rest local.

Addresses review findings: Blocker 2, Blocker 3.
2026-08-08 05:50:17 -05:00
0xdhx
08021b02c0 fix(#2665): scrub config-location env vars in every TEST_ENV_BASE declaration
TEST_ENV_BASE blanks session-identity variables but none of the three that
decide WHERE a child process writes: CLAUDE_CONFIG_DIR, GSD_RUNTIME and
CODEX_HOME. The config-home resolver is env-first (runtime-homes.cts, the
dot-home case consults the env var before the home-derived fallback), so an
ambient CLAUDE_CONFIG_DIR in the developer's shell beats a call site that
sandboxes only HOME. The suite then writes into the developer's real config
directory -- including a registered skill under <configDir>/skills/ whose
body carries behavioural directives that load into later sessions.

Blank all three alongside the session-identity vars. `...env` still spreads
last, so the five call sites that already constrain these locally keep
winning with their explicit values.

TEST_ENV_BASE is re-declared in nine files, so the three lines are added
nine times rather than once. Consolidating the nine into a single exported
constant -- and fixing the TERM_SESSION / TERM_SESSION_ID drift between the
copies -- is deliberately left out of this change; see the PR body.

One call site needed adjusting. capability-state.test.cjs's
`capability state --runtime claude` CLI test passed no env at all and
compared the CHILD's resolved config dir against the PARENT process's
getGlobalConfigDir('claude'). That agreed only because the child inherited
the developer's ambient CLAUDE_CONFIG_DIR -- i.e. it passed *because of*
the leak. It now redirects both runtime homes into the sandbox and asserts
against values the test controls, so it is hermetic with the variable set
or unset.

Regression case folded into the owning module's test file rather than a new
bug-NNNN file, per scripts/lint-regression-test-names.cjs. It sets the
variable on the PARENT process, which is the actual vector; setting it in
the per-call env argument would exercise a path that was never broken.
2026-08-08 05:49:33 -05:00
Tom Boucher
9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00
Tom Boucher
000a322489 fix(#2941): hint at global: prefix when bare skill name matches a global skill (#2973)
* test(#2941): add regression for bare-skill-name global: hint

When a bare skill name matches an existing global skill, the skip warning
must hint at the global: prefix. When no global skill matches, the original
'Skill not found' warning is unchanged.

* fix(#2941): hint at global: prefix when a bare skill name matches a global skill

When buildAgentSkillsBlock skips a bare skill name that doesn't exist as a
project-relative path, check if it matches an existing global skill. If so,
append a hint to the warning: 'a global skill named X exists; use global:X
to reference it'. When no global skill matches, the warning is unchanged.

getGlobalSkillDir and getGlobalSkillsBase are already imported in this module
for the global: branch. The hint is guarded on globalSkillsBase being non-null
(runtimes without a skills directory don't support the prefix).

* chore(#2941): add changeset fragment

* chore(#2941): backfill changeset PR number 2973

---------

Co-authored-by: sim <sim@local>
2026-08-01 10:58:50 -04:00
Daniel Einspanjer
f0ff23635e fix(#2602): discover project-local Codex agents (#2623)
* fix(#2602): discover project-local Codex agents

- Select an existing local Codex agents directory before global fallback
- Prove init reports the canonical local installation through compiled CJS

* test(#2602): lock Codex agent precedence

- Cover override, local authority, global fallback, and runtime compatibility
- Exercise installed state through the compiled resolver

* fix(#2602): resolve local Codex agent skills

- Pass the canonical project root to the non-Claude persona fallback
- Cover nested-Codex fallback and Claude compatibility through the CLI

* test(#2602): cover local Codex validation status

- Assert emitted validate and health commands use the project-local install
- Preserve empty local-directory authority beside complete global agents

* fix(#2602): align validation with local Codex discovery

- Pass the resolved runtime and project root to health W010
- Resolve the validate-agents runtime before checking installation status

* test(#2602): cover local Codex docs status

- Assert docs-init reports an authoritative empty local install as unhealthy

* fix(#2602): align docs with local Codex discovery

- Pass the resolved runtime and canonical project root to the shared agent checker

* fix(#2602): honor agent-skills runtime override

- Resolve agent-skills fallback runtime through the canonical project resolver
- Cover conflicting config and GSD_RUNTIME values through the emitted CLI

* fix(#2602): ignore non-directory local agents paths

- Treat only a local Codex agents directory as authoritative
- Cover regular-file fallback through the emitted install checker

* chore(#2602): add changelog fragment

- record the user-visible local Codex agent discovery fix for PR #2623

* fix(#2602): align local agent discovery with runtime policy

- Resolve Codex's local config directory through the canonical runtime policy
- Use test-managed cleanup for local-agent discovery coverage

* fix(#2602): discover local agents across runtimes

- Prefer manifest-backed project-local installs for non-Claude runtimes
- Respect runtime-specific local install roots and preserve global fallback behavior
- Cover native, partial, cross-runtime, and project-root local discovery

* fix(#2602): preserve agent discovery fallback

- Fall back globally when local-install probes fail
- Document and test symlink rejection
- Align the changeset with repository format

* fix(#2602): reuse local directory policy

- Resolve runtimes without local config through the canonical sentinel
- Document the manifest gate and refresh the context index

---------

Co-authored-by: Daniel E. <daniel.e@teachingstrategies.com>
Co-authored-by: Rezolv <dave@sienkowski.com>
2026-07-31 21:20:46 -04:00
Behruz Nassre Esfahani
6bbd8919d5 fix(#1400): flush agent-skills block via writeAllSync instead of process.exit(0)
cmdAgentSkills' plain (non-JSON) path wrote the <agent_skills> block with
process.stdout.write then immediately called process.exit(0). stdout.write is
async on pipes/files, so on Windows the process tore down before Node flushed
the buffer — the workflows' `$(gsd_run query agent-skills <type>)` capture
received 0 bytes and every ${AGENT_SKILLS_*} substitution expanded empty,
silently dropping configured per-agent skills.

Route the plain path through the existing synchronous output() helper
(writeAllSync, src/io.cts) + return — the same flush-safe mechanism the --json
branch already uses — instead of write + process.exit(0).

Adds a #1400 regression block to tests/agent-skills.test.cjs that captures
stdout via a real file descriptor and asserts the block is non-empty and
byte-identical to the --json .block.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-20 16:49:18 -07:00
Tom Boucher
b0c774c2e3 feat(#1416): formalize Resolution convention + agent-skills value envelope (Resolution Provenance P3) (#1425)
Narrows P3 of ADR-1411 (Resolution Provenance, epic #1411) based on an
adversarial fit-analysis that showed a single Resolution<T> envelope adopted
by agent-skills, capability-state, and capability-writer fails the deletion
test: configured/reason are meaningless for capability verbs, and
capability-writer's errors[] (operation-not-applied) cannot fold into
warnings[]. The only genuinely shared seam is warnings: string[].

Changes:

- src/resolution.cts: new pure types+builder leaf — exports Resolution<T>
  {value, configured, reason, warnings}, makeResolution<T>() builder, and
  AgentSkillsValue {block, skills_count}. No other src/ imports.

- src/init.cts: cmdAgentSkills --json IR gains additive value:{block,
  skills_count} field (built via makeResolution). All existing flat fields
  (agent_type, block, skills_count, warnings, configured, reason, source,
  degraded) are retained unchanged for back-compat.

- src/capability-state.cts: doc comment on ResolveCapabilityRuntimeStateResult
  naming it the canonical read-verb envelope. No JSON change.

- src/capability-writer.cts: doc comment on SetCapabilityStateResult naming it
  the canonical mutation-verb result (warnings=advisory, errors=operation-
  not-applied). No JSON change.

- CONTEXT.md: new ### Resolution Convention glossary entry after
  ### Resolution Provenance.

- docs/adr/1411-resolution-provenance.md: P3 narrowing amendment appended.

- tests/resolution.test.cjs: 9 unit tests for makeResolution (new).
- tests/agent-skills.test.cjs: 2 P3 tests for value.block/value.skills_count
  and back-compat of all flat fields.

All 277 tests pass (5 suites). npm run lint clean. All lint checks pass.

Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:44:52 -04:00
Tom Boucher
484a5b7b86 fix(#1415): loadConfig provenance + agent-skills diagnostic (Resolution Provenance P2) (#1424)
Implements ADR-1411 P2 / #1415, closing #1366.

Part 1 — config-loader.cts:
- Adds `loadConfigResolved(cwd, options) → ConfigResolution { config, source, degraded }`
  with six tagged branch return paths:
  A1 ws+wsconfig → source:'workstream', degraded:false
  A2 no-ws+config → source:'root', degraded:false
  B  ws requested, wsconfig absent → source:'root', degraded:true (intercepts recursive call)
  C  .planning/ exists, no config → source:'builtin-defaults', degraded:false
  D  no .planning/, global defaults readable → source:'global-defaults', degraded:false
  E  no .planning/, no global → source:'builtin-defaults', degraded:false
- `loadConfigResolved` calls `findProjectRoot` at entry to anchor resolution to the
  nearest .planning/ ancestor (cwd-drift fix, heuristic 4 from P1)
- `loadConfig` becomes a one-line delegation: `return loadConfigResolved(cwd, options).config`
- Exports `loadConfigResolved` in the `export =` block

Part 2 — init.cts cmdAgentSkills:
- Imports `findProjectRoot` from `./project-root.cjs`
- Anchors to project root before loading config (fixes #1366 cwd-drift)
- Uses `loadConfigResolved` for provenance; passes projectRoot to buildAgentSkillsBlock
- Computes `configured` + `reason` (AgentSkillsReason enum): 'resolved' |
  'not_configured' | 'configured_empty' | 'configured_unresolved'
- configured_empty and configured_unresolved emit stderr WARNING; not_configured is silent
- --json IR gains: configured, reason, source, degraded (in addition to existing
  agent_type, block, skills_count, warnings)

Tests (TDD):
- tests/config-loader.test.cjs: 8 new provenance tests (RED before impl, GREEN after)
- tests/agent-skills.test.cjs: 7 new diagnostic tests (RED before impl, GREEN after)
- All 306 tests across 4 suites pass (config-loader:33, agent-skills:72, init:105, workstream:96)

Docs:
- CONTEXT.md: Config Loader Module entry updated with loadConfigResolved interface;
  Resolution Provenance entry notes P2 is now implemented
- docs/CLI-TOOLS.md: --json field reference table added for agent-skills

Closes #1366
Part of #1411

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 07:05:15 -04:00
Tom Boucher
6210621494 test(#1377): pin real agent-skills IR contracts in two vacuous tests (#1380)
The empty-IR test asserted `parsed === '' || typeof parsed === 'object'`,
which always passes (typeof null === 'object'). Pin the real contract:
no agent type → `output('', raw, '')` → the --json IR is the empty string.

The nonexistent-skill-path test asserted only `block === ''`. Since #1376
added a warnings[] field to the --json IR, also assert warnings[] names the
skipped path so the test guards the silent-drop regression it is named for.

Test-only; no product behavior change.

Closes #1377

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 11:20:21 -04:00
Tom Boucher
66085d0080 fix(#1374): surface diagnostic when configured agent skills all fail to resolve (#1376)
* fix(#1374): surface diagnostic when configured agent skills all fail to resolve

buildAgentSkillsBlock returned '' (only ad-hoc per-path stderr warnings) when an agent configured via agent_skills had paths that all failed to resolve — missing SKILL.md, unsafe path, invalid global name, OR a malformed (non-string/non-array) value. query agent-skills --json reported skills_count>0 with an empty block and no machine-readable signal, so a fully-dropped configuration was indistinguishable from a resolved one.

Thread an optional diagnostics collector through buildAgentSkillsBlock: route every skip warning through a warn() helper (stderr + collector), flag truthy-but-malformed config values, emit an aggregate warning when configured paths resolve to zero skills, and surface the collected reasons in a new warnings[] field on the query agent-skills --json IR. Empty arrays and falsy values stay silent (skills_count is honestly 0). skills_count semantics unchanged. Docs updated for the new IR field.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1374): backfill changeset PR number (#1376)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 10:16:51 -04:00
Tom Boucher
cf68841220 enh(#1243): consume Claude plugin-provided skills in agent_skills (epic #1258 Phase B) (#1261)
* feat(#1243): consume Claude plugin-provided skills via native Skill-tool directive + grant Skill to agent_skills-consumer agents

- Relax global skill name validation to accept namespaced form `^[A-Za-z0-9_-]+(:[A-Za-z0-9_-]+)*$`
- Namespaced names (containing colon) on claude runtime emit a Skill-tool load directive instead of a @-include line
- Namespaced names on non-claude runtimes are skipped with a warning
- Bare unresolved names retain existing warn-and-skip behavior (no promotion to directive)
- Grant `Skill` tool to all 22 agent_skills consumer agents; 5 generated agents updated via research-profiles.cjs + regen, 17 hand-authored agents edited directly
- Add 16 TDD tests in describe('bug #1243') covering happy/mixed/precedence/negative/cross-runtime/regression/grant cases

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* docs(#1243): document plugin-provided skills in agent_skills

Update the Agent Skills Injection reference in CONFIGURATION.md with
the three entry forms (project-relative, global:<name>,
global:<plugin>:<skill>), the Claude-only runtime behaviour of the
namespaced form and the warn-skip on other runtimes, the plugin
pre-install prerequisite, and the consumer-agent Skill tool grant.

Add docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md with a
step-by-step guide for installing the plugin, locating the namespaced
skill name, wiring it into agent_skills, and verifying injection.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#1243): align agent_skills docs with emitted block format + mixed-block regression test (code-review)

- Replace two-section mixed-block example (bogus "Load these plugin-provided skills using the Skill tool:" header) with the actual single-section inline format in CONFIGURATION.md and docs/how-to/attach-a-plugin-skill-to-a-gsd-agent.md
- Fix quoted warning text in how-to doc to exactly match the emitted string: [agent-skills] WARNING: Plugin-namespaced skill "global:<name>" requires a Skill-tool-capable runtime (claude) — skipping on runtime "<runtime>"
- Replace phantom agent slugs (gsd-checker, gsd-researcher, gsd-advisor, gsd-synthesizer) in CONFIGURATION.md Supported Agent Types with real agents/gsd-*.md examples (gsd-plan-checker, gsd-phase-researcher, gsd-code-reviewer, gsd-ui-auditor, gsd-research-synthesizer)
- Add byte-identical mixed-block regression test: one path-resolvable global skill + one plugin-namespaced skill on claude runtime → asserts r.ir.block === single-section interleaved block, no secondary header

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(#1243): regenerate agent-size baseline for the Skill-tool grant

The 22 agent_skills-consumer agents each grew +7 bytes from adding `Skill`
to their tools list; refresh the committed per-agent size baseline (#1074 guard).

* chore(#1243): add Added changeset fragment

* fix(#1243): traceable allow-test-rule ref + separator-agnostic byte-identical tests (CI)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-15 00:49:12 -04:00
Tom Boucher
f7e902f1cf feat(#52): add agent_skills_security.trusted_global_roots allowlist for global skills (#754)
* feat(#52): add agent_skills_security.trusted_global_roots allowlist

Opt-in allowlist so a global: agent skill whose SKILL.md realpath resolves
outside the default global skills base (e.g. ~/.claude/skills) is accepted
when its real target lies under a user-declared trusted root. Default [] is
byte-identical to prior behavior; the symlink-escape guard is preserved and
simply re-applied against each declared root.

- src/security.cts: loadTrustedGlobalRoots — tilde-expand (~ and ~/), reject
  project-relative and dangerously broad roots (filesystem/UNC root, homedir),
  realpath-canonicalize each root every run and drop non-existent ones.
- src/init.cts: on base-check failure the guard consults the trusted roots
  (hoisted out of the loop); emits a stderr NOTE when a skill is accepted via
  a trusted root so the widened boundary is visible.
- src/core.cts: thread agent_skills_security through loadConfig.
- config-schema.manifest.json: allow the new key path.
- docs/CONFIGURATION.md: document the option and its security model.
- tests/agent-skills.test.cjs: unit + end-to-end CLI coverage (regression,
  feature, negative, broad-root hardening, stderr NOTE).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#52): add changeset fragment for trusted_global_roots (#754)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-07 01:06:41 -04:00
Tom Boucher
a28dcec981 chore(#597): replace count-based ratchet guards with AST lint + named-set allowlists (#603)
The windows-test-parity ratchet greps test source for fs.rmSync-without-
maxRetries (and six other Windows-portability anti-patterns), failing when an
integer offender COUNT exceeds a frozen baseline (rmSync: 95). A count ratchet
is a Goodhart metric: fixing one offender and adding another keeps the count
constant, so a new defect slips through green. Replace it — and every other
count ratchet in the repo — with a layered, masking-proof design.

Behavioral seam test
- tests/helpers-cleanup.test.cjs proves helpers.cleanup() carries the Windows
  EBUSY retry budget. cleanup() delegates retries to Node's fs.rmSync via
  maxRetries (it owns no loop), so the test asserts the option contract
  (recursive/force/maxRetries>0/retryDelay>0) + real-FS removal + the cwd-guard,
  rather than a loop that does not exist. The EBUSY risk is now tested ONCE at
  the helper, not approximated textually at every call site.

Write-time ESLint rule (AST-accurate, replaces the grep)
- eslint-rules/no-raw-rmsync-in-tests.cjs (error in tests/**/*.test.cjs) bans
  raw fs.rmSync, steering to cleanup(). Catches member, computed (fs['rmSync']),
  destructured and aliased forms; escape hatch is inline
  `// eslint-disable-next-line local/no-raw-rmsync-in-tests -- <reason>` only.
- Migrated 336 raw fs.rmSync teardown calls across ~116 test files to cleanup().
  ~18 genuinely load-bearing sites (mid-test SUT/fault-injection removals,
  error-swallowing or name-colliding local teardown helpers) keep the raw call
  with an inline eslint-disable + reason.

Shared anti-ratchet primitive
- scripts/lib/allowlist-ratchet.cjs:
  - assertWithinAllowlist: fails on NOVEL ids (new offender introduced) AND on
    STALE ids (a known offender was fixed but not pruned) — identity, not count,
    and a ratchet DOWN toward zero.
  - assertTightCeiling: a size/length budget whose ceiling must stay within a
    grace band of the high-water mark, so budgets may only tighten, never creep.

Ratchets converted onto the primitive
- windows-test-parity-guard.test.cjs: rmSync rule deleted (now ESLint-enforced);
  the remaining six patterns moved from integer baselines to named-set
  allowlists with ratchet-down.
- scripts/lint-test-file-count.{cjs,allowlist.json}: per-module integer counts →
  named filename sets (closes the swap-a-file-keep-the-count blind spot); a
  module dropping under cap now FAILS to force pruning its allowlist entry.
- enh-2790 skill-count `<= 63` → named skill allowlist (ratchets toward ~58).

Size budgets hardened (tighten-only)
- agent-size / workflow-size / feat-3039 help-tiered: ceilings lowered to the
  current high-water mark and an assertTightCeiling anti-creep check added per
  tier. Fixed external-contract limits (description ≤100 chars, agent ≤100 KB)
  are intentionally left as-is — they are not grandfathered creeping budgets.

No user-facing behavior change (tests + tooling only); no USER_FACING_PREFIXES
touched, so no changeset fragment is required.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 22:43:49 -04:00
Tom Boucher
b8c33647d8 refactor(tests): retire output-grep & source-grep via typed surfaces (finish #2974) (#462)
* refactor(#455): implement typed surfaces to retire grep tests

Production surfaces added:
- hooks/managed-hooks-registry.cjs: new CJS module exporting MANAGED_HOOKS
  as a typed array; gsd-check-update-worker.js now requires it instead of
  declaring an inline array
- bin/install.js: elevate inline gsdHooks to module-level GSD_UNINSTALL_HOOKS,
  export it alongside runtimeMap/allRuntimes (already exported)
- scripts/build-hooks.js: export HOOKS_TO_COPY; guard build() behind
  require.main===module so tests can require the file without triggering a build
- get-shit-done/bin/lib/init.cjs: add --json mode to agent-skills command,
  emitting typed IR { agent_type, block, skills_count } for test assertions
- get-shit-done/bin/gsd-tools.cjs: wire --json flag for agent-skills dispatch

Category-B source-grep migrations:
- tests/managed-hooks.test.cjs: require MANAGED_HOOKS from registry, drop fs.readFileSync+regex
- tests/orphaned-hooks.test.cjs: require MANAGED_HOOKS+HOOKS_TO_COPY as typed exports
- tests/hooks-opt-in.test.cjs: replace gsdHooks regex-parse with GSD_UNINSTALL_HOOKS import
- tests/install-minimal-hooks.test.cjs: replace gsdHooks regex-parse with GSD_UNINSTALL_HOOKS
- tests/copilot-install.test.cjs: replace src.includes() checks with typed
  assertions on runtimeMap, allRuntimes, parseRuntimeInput, buildRuntimePromptText
- tests/agent-skills.test.cjs: migrate to --json typed IR assertions

pending-migration-to-typed-ir token cleared (87 of 87 files):
- 78 files already had source-text-is-the-product; removed duplicate token
- 5 files already used typed assertions; reclassified or annotated
- 4 files required individual reclassification to source-text-is-the-product
  or architectural-invariant

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): update workflow-guard test to typed GSD_UNINSTALL_HOOKS import; isolate HOME in runtime-launcher (D) test

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): guard install.js main() behind require.main===module so the typed export is require-safe

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#455): document --json typed surfaces for agent-skills, progress, validate context

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(#455): add changeset fragment for new --json surfaces

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(#455): complete grep migration for files flagged by lint-tests

The branch commit 4e630d99 stripped `allow-test-rule: pending-migration-to-typed-ir`
from ~80 test files without replacing their assertions or adding the correct
exemption annotation. The files were NOT source-grep tests — they read .md
workflow/agent/command/reference files (source-text-is-the-product) or hook
source files for structural invariants (structural-regression-guard). No
assertion logic was changed; only the correct allow-test-rule annotation was
added to each file per CONTRIBUTING.md exception matrix.

73 files: `source-text-is-the-product` — workflow/agent/command/reference .md
7 files:  `structural-regression-guard` — hook .js / bin/install.js structural checks

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: CI Rebase Check <ci@gsd-redux>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-05-29 11:39:52 -04:00
Tibsfox
0a18fc3464 fix(config): global skill symlink guard, tests, and empty-name handling (#1992)
Review feedback from @trek-e — three blocking issues and one style fix:

1. **Symlink escape guard** — Added validatePath() call on the resolved
   global skill path with allowAbsolute: true. This routes the path
   through the existing symlink-resolution and containment logic in
   security.cjs, preventing a skill directory symlinked to an arbitrary
   location from being injected. The name regex alone prevented
   traversal in the literal name but not in the underlying directory.

2. **5 new tests** covering the global: code path:
   - global:valid-skill resolves and appears in output
   - global:invalid!name rejected by regex, skipped without crash
   - global:missing-skill (directory absent) skipped gracefully
   - Mix of global: and project-relative paths both resolve
   - global: with empty name produces clear warning and skips

3. **Explicit empty-name guard** — Added before the regex check so
   "global:" produces "empty skill name" instead of the confusing
   'Invalid global skill name ""'.

4. **Style fix** — Hoisted require('os') and globalSkillsBase
   calculation out of the loop, alongside the existing validatePath
   import at the top of buildAgentSkillsBlock.

All 16 agent-skills tests pass.
2026-04-11 03:39:29 -07:00
Tom Boucher
2703422be8 refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md (#1675)
* refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md

- Replace require('node:assert') with require('node:assert/strict') across
  all 73 test files to enforce strict equality (no type coercion)
- Replace try/finally cleanup blocks with t.after() hooks in core.test.cjs
  and hooks-opt-in.test.cjs per the test lifecycle standards
- Utility functions in codex-config and security-scan retain try/finally
  as that is appropriate for per-function resource guards, not lifecycle hooks

Closes #1674

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(tests): add --test-concurrency=4 to test runner for parallel file execution

Node.js --test-concurrency controls how many test files run as parallel child
processes. Set to 4 by default, configurable via TEST_CONCURRENCY env var.
Fixes tests at a known level rather than inheriting os.availableParallelism()
which varies across CI environments.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): allowlist verify.test.cjs in prompt-injection scanner

tests/verify.test.cjs uses <human>...</human> as GSD phase task-type
XML (meaning "a human should verify this step"), which matches the
scanner's fake-message-boundary pattern for LLM APIs. This is a
false positive — add it to the allowlist alongside the other test files
that legitimately contain injection-adjacent patterns.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-04 14:29:03 -04:00
Tom Boucher
db3eeb8fe4 feat: agent skill injection via config (#1355)
Add agent_skills config section that maps agent types to skill directory
paths. At spawn time, workflows load configured skills and inject them
as <agent_skills> blocks in Task() prompts, giving subagents access to
project-specific skill files.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-24 16:47:40 -04:00