Commit Graph

33 Commits

Author SHA1 Message Date
0xdhx
a294ec2a2b test(#2665): widen the hermeticity guard to its two blind surfaces, and cover its budget
Round 2, both Majors. They are one defect seen twice: the recurrence guard did
not cover the surface it exists to guard.

Blind surfaces. resolveLiveConfigRoots enumerates getGlobalConfigDir per registry
runtime plus a hardcoded grok branch, so it can only ever see runtime config
ROOTS. Two live write surfaces are not roots and passed through silently:

  $GSD_HOME/.gsd     — GSD's user-owned store. Watched WHOLESALE: unlike ~/.claude
                       this root is exclusively ours, so the shared-root
                       false-positive trap the module documents does not apply.
  <kimi>/config.toml — the file GSD writes its native [[hooks]] block into. The
                       INVERSE case: ~/.kimi belongs to Kimi CLI, so only the one
                       file GSD writes is watched, never the root.

That asymmetry is why this is not a two-line "add two roots" patch — one target
needs the whole tree, the other needs exactly one file, and collapsing them
either under-watches the store or trips the guard's own documented
false-positive trap on a third party's directory.

Extras are passed to snapshotLiveConfig explicitly rather than resolved inside
it, so a caller snapshotting a fixture root cannot silently pull the developer's
real ~/.gsd into its own assertions. run-tests.cjs now snapshots when EITHER the
roots or the extras are non-empty — previously an unbuilt tree yielding zero
roots disabled the entire guard without saying so.

Budget coverage. The MAX_ENTRIES/MAX_DEPTH bound and the truncated -> 'unverified'
branch had zero tests, despite this module's own docstring naming "a truncated
scan reading as clean" as the safety-critical case. Added per
RULESET.TESTS.boundary-coverage (N in {limit-1, limit, limit+1}, exercised
through newestMtime's injected budget so the boundary is real without
materialising 20000 files) and RULESET.TESTS.property-based-testing (fast-check:
truncation is monotone in the budget; reported newest never exceeds the true
maximum). A regression flipping `truncated` to false on an exhausted budget now
breaks the property for every budget below the tree size.

Negative-controlled: neutering the extras wiring fails exactly the two
new-surface tests and nothing else. 21/21 green with it restored.
2026-08-08 05:50:36 -05:00
0xdhx
e2eed1c58a test(#2665): ship the hermeticity guard at report level, not fatal
Its first CI run found PRE-EXISTING leaks on the Windows lane —
C:\Users\runneradmin\.claude\gsd-core and skills\gsd-dev-preferences — with all
1196 Windows tests otherwise passing. os.homedir() reads USERPROFILE on Windows,
and ~190 test sites across 31 files sandbox HOME alone, so the suite has been
installing GSD into the runner's real home directory invisibly. That is exactly
the class the guard exists to surface, and exactly the class this PR's review
said CI could never catch.

It is also a different defect from the one #2665 closes, and too large to fold in
here. A brand-new gate that immediately reds an unrelated lane gets bypassed or
reverted rather than obeyed, so the guard reports by default and fails only under
GSD_STRICT_LIVE_CONFIG_GUARD=1.

This is the repo's own established ratchet, not a hedge: the local/no-source-grep
ESLint rule shipped at `warn` and was promoted to `error` after its cleanup sweep
(ADR 452). Promote this the same way once the USERPROFILE sweep lands.
2026-08-08 05:50:17 -05:00
0xdhx
a02462e050 test(#2665): fail the suite when it writes into a live config dir
The recurrence guard, and #2665's own "Optional hardening". This class is silent
by construction: TEST_ENV_BASE cannot see an in-process caller, and CI cannot see
the class at all because CI never has these env vars set. It damages the
developer's machine and reports nothing -- which is how two prior authors each
diagnosed it and fixed only the instance in front of them.

run-tests.cjs snapshots GSD's install footprint in every live runtime config dir
before the suite and re-checks it after, failing the run on a create or a modify.
Roots come from the product's own getGlobalConfigDir, so the guard watches
wherever the product actually points, including through an ambient var.

Scope is ownership-based, not whole-root: the top-level install footprint plus
gsd-prefixed children of dirs GSD shares with the host agent. A config root like
~/.claude is shared, and watching it wholesale would false-positive on the host's
own history.jsonl or settings.json -- a guard that cries wolf gets disabled, and
then catches nothing. The prefix test is load-bearing: the first version watched
only the three top-level entries and MISSED a real leak into skills/gsd-*.

It earned its place immediately -- it is what found the fifth in-process leak in
runtime-artifact-layout.test.cjs, which no amount of reading the review would have
surfaced. Known gap documented in the module: a write to a file GSD does not own
is out of scope by construction.

Lives in scripts/, deliberately NOT scripts/lib/ -- the installer copies that dir
into every user's config dir wholesale while uninstall removes only an allowlist,
so a test-only module there would ship to users and survive uninstall.

Addresses review finding: Minor 8.
2026-08-08 05:50:17 -05:00
Tom Boucher
33fd203ccd test(#2966): loop QA walk — drive real scenarios across all five loop steps (#2976)
* test(#2966): loop QA walk — drive real scenarios across all five loop steps

Adds a headless walk that carries accumulating project state across
discuss -> plan -> execute -> verify -> ship against one temp project,
layered over the existing tests/helpers.cjs runGsdTools substrate.

Findings carry severity. A violation breaks a stated contract and fails
the build; a smell is legal under today's implementation but structurally
questionable, is recorded, and never reddens CI. Without that split an
oracle set derived from current behavior can only ever confirm current
behavior -- the harness could not say "this works and is still wrong".

The end-to-end test asserts the walk produces at least one smell: a QA
harness that reports nothing on a first run against a real engine is far
more likely mis-specified than the engine is perfect. It deliberately does
not pin smell ids or counts, which would re-freeze current behavior.

First run against the real engine: 0 violations, 3 smell classes --
init returns agents_dir outside the project tree; smart-entry emits prose
unconditionally so routing cannot be asserted; state-snapshot reports a
missing STATE.md through a payload key with exit 0.

Also fixes tests/fixtures/index.cjs: createFixture with git:true and
planning:false staged nothing, so the commit failed with "nothing to
commit". That combination was unreachable until greenfield needed it.

Extends RULESET.TESTS.feedback-loop-convergence from estimation to the
loop itself. Design lock: docs/adr/2966-loop-qa-walk.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): wire fault injection, make perturbations discriminating

Independent review found tests/qa/mutations.cjs entirely unwired: 462
lines exercised only by their own unit tests, with no mutation hook in
the scenario DSL and no scenario applying one, while the module header
and the ADR described fault injection in the present tense. Dead code
documented as live.

Adds a `mutate` step field, three perturbation scenarios, and a wiring
detector: a self-test scenario whose expectations are known-false and
which MUST fail. The previous anti-vacuity check asserted only that the
walk produced a smell, which passes on well-known engine behavior
regardless of whether the harness wiring works.

First perturbation attempt produced zero signal -- progress does not
structurally parse ROADMAP.md, so a corrupted roadmap sailed through. A
perturbation that cannot fail is the same defect in a new costume.
Probes now target roadmap get-phase, and each mutated step runs a clean
baseline first so `mutationObserved` records whether the corruption
changed anything at all.

Also clears four review findings: classify() returned PROSE for exit-0
with empty stdout; `warnings` was structurally unpopulatable on the
success path (execFileSync discards it) and is now documented as
error-path-only; read-only-idempotence passed vacuously when asked to
check idempotence without the data to check it; the ADR miscounted the
oracles.

Discrimination matrix across 8 mutations x 6 commands: bom,
duplicate-phase-id and escaped-pipes are absorbed silently by every
probed surface, and progress / smart-entry / roadmap validate never
reacted to any mutation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): add path-containment guard for scenario-supplied targets

Security review found scenario-supplied paths joined to the temp project
with no containment check. step.mutate.target and agent.write keys were
validated only as non-empty strings, so a target of ../../../../etc/hosts
reached fs.unlinkSync / fs.writeFileSync / fs.symlinkSync outside the
project. The symlink mutation was worst: it read the traversed file, wrote
a sibling copy, deleted the original and symlinked it back.

Not exploitable today -- all shipped scenarios target .planning/ROADMAP.md
and scenarios are repo-committed, not runtime input. Fixed anyway: it is a
live primitive any future scenario or copied helper can reach.

Adds tests/qa/paths.cjs with resolveWithin(): rejects absolute paths, NUL
bytes and empty input, normalizes separators unconditionally, and requires
containment by path segment so a sibling like <base>-evil is not treated as
inside. Non-existent targets resolve via nearest existing ancestor rather
than falling back to a lexical compare. Scenario load now rejects traversing
or absolute targets up front.

oracles.cjs previously carried its own copy of the containment logic; both
now share paths.cjs, since a duplicated containment check is exactly the
divergence class this repo calls out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): complete trajectory corpus, report emission, boundary-aware oracle

Adds the remaining trajectories and drives all 11 mutations end-to-end.
20 scenarios, 72 steps, 0 violations, 25 smells.

Adds qa-report.json with per-step verdicts and a copy-pasteable repro
command, plus --keep / GSD_QA_KEEP=1 to preserve a failing tree. A repro
line for a tree that was not preserved is marked NOT RUNNABLE rather than
emitting a command pointing at a deleted directory.

monotonic-progress is now boundary-aware. Two scenarios had been trimmed
to stop the oracle complaining at a milestone rollover, which destroys the
signal the trajectory exists to produce. Evidence: counters legitimately
reset to zero at milestone complete, but the payload milestone_version
lags until a new ROADMAP.md is written. So the oracle now scopes by
milestone plus workstream, keeps a same-scope decrease as a violation, and
records a boundary crossing as a smell. Both scenarios walk the real
boundary again.

Standards review fixes: oracle findings now carry a structured subject so
tests assert on typed fields instead of substring-matching the free-form
detail string, resolveWithin throws a typed EPATHESCAPE error, and the
absolute-path predicate scenario.cjs had re-implemented now comes from
paths.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): fix silently-vacuous fixtures and guard the class

Every fixture carried its #2371 provenance comment BEFORE the frontmatter
block, and extractFrontmatter returns {} when anything precedes the opening
---. So every scenario reading status/phase/name was operating on an empty
object and reporting green. Nine fixtures repositioned; the comment stays,
it just moves below the closing ---.

Both UAT fixtures lacked a parser-recognized result block, so
evaluateUatPassed saw checks.length===0 and could never return passed:true.
The uat-fail-then-remediate scenario could not have proven a remediation.
Its expect block only inspected blockers, which is empty before AND after,
which is why the corpus never noticed. Both fixtures now carry real result
blocks and the scenario asserts passed and no_uat_artifacts on each side of
the flip.

The actual deliverable is the guard: a fixture-integrity block asserting
every fixture with a frontmatter shape parses to a non-empty object, that
every fixture carries its provenance marker, and that the two UAT fixtures
produce opposite verdicts through the real evaluateUatPassed. The first
guard written required --- at byte 0, which would never have fired on the
regression it exists to prevent; it was rewritten and proven by deliberately
re-breaking a fixture.

No engine defect here. no_uat_artifacts means no parsed check items, not no
UAT files, and it was reporting correctly on fixtures that had none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): make the walk report — smell ratchet, baseline, CI job

The harness computed smells into a gitignored qa-report.json that nothing
read. In CI it surfaced nothing at all: violations failed the build, but the
half of the tool that says "this works and is still wrong" was inert. A QA
tool nobody hears is decoration.

Adds a ratchet on the same idiom this repo already uses three times over
(the regression-test-name allowlist, the emitted-drift acks, the size
baseline): a committed smell-baseline.json, per-PR acknowledgment fragments
under tests/qa/smell-acks/, and a ratchet script wired into CI.

The design invariant is preserved exactly. A smell still never fails a build
on its own merits. What fails is an UNACKNOWLEDGED NEW smell -- the absence
of a decision -- leaving an author two honest exits: fix it, or record a
fragment with a real reason. An empty reason is rejected. The baseline is
shrink-only, so a fixed smell must prune its entry. Violations remain
unacknowledgeable.

Fingerprints are composed only from stable fields (oracle id, scenario,
argv, subject discriminator) -- never temp paths, timestamps or counts.
Verified byte-identical across two runs in separate temp dirs; an unstable
fingerprint would have false-positived every CI run.

CI gains a qa-loop-walk job that runs the suite and the ratchet, uploads the
report with `if: always()` (it matters most when it failed), and renders a
summary a reviewer reads without downloading anything.

Also fixes the report runner invoking main() unconditionally on require, so
importing it double-ran every scenario and clobbered its own output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): every smell terminates in a defect or a fixed detector

The baseline accepted a smell with a free-text reason. That is a mechanism
for designing smells in -- an allowlist nobody revisits. The harness is
brand new, so nothing it found is inherited legacy; every finding is a
FIRST finding. Each must now terminate in exactly one of two states:

  REAL           -> an assigned defect, entry carries the issue number
  FALSE POSITIVE -> the detector is wrong and gets fixed, never baselined

There is no third "accepted with a good explanation" state, so the ratchet
now requires a positive-integer `issue` on every entry. A reason may remain
as a human note but can never substitute. `--update` refuses to invent
issue numbers: a new smell is written with `issue: null` and a TODO, and
the next plain run rejects it, forcing triage rather than accumulation.

Working the 21 existing entries through that rule found 16 were my own
detectors being wrong:

value-hygiene (10) flagged $.agents_dir, a field whose entire contract is
to point at the install tree outside any project. Fixed with a leaf-key
allowlist of contractually-external fields, verified as the only such key
in the init payload. Genuinely unexpected out-of-project paths still smell.

monotonic-progress (6) fired on legitimate boundary crossings -- milestone
v1.0 to v2.0, workstream beta to alpha -- and on one payload carrying no
scope fields at all, where a change cannot even be known. Scope changes now
reset silently and scope-less observations are skipped. The same-scope
decrease remains a violation; that is the real invariant and is regression-
guarded.

The five survivors are real and now tracked: soft-error-exit-zero (#2980),
untyped-success (#2979). Baseline 25 -> 5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): keep the ratchet out of the tarball, unpin the qa CI job

The remote matrix returned failed -- 3 unique failures, identical on
node22 and node24, both root causes in this branch's own diff.

The ratchet lives under scripts/, which ships in the npm tarball, and it
requires three modules under tests/, which does not. In a published
install it is MODULE_NOT_FOUND at load. This is exactly the class the
#2858 guard was added to catch, and it caught it. Fixed the way #2858
fixed the same shape for its own repo-only CI script: a targeted files[]
negation, so the ratchet stays in the repo for CI and out of the tarball.
Not solved by moving or inlining the required modules -- the ratchet must
keep using the same code the harness uses, or the two drift.

Verified both directions: the script is no longer in the pack list, and
build-hooks.js, fix-slash-commands.cjs and gen-capability-registry.cjs are
all still shipped. Over-negating there would have broken installs, since
bin/install.js requires them.

The qa-loop-walk job also carried CI_REBASE_BASE_SHA copied from a
neighbouring job without the paired GSD_EMITTED_BASE, which the #2854
invariant forbids by name: diverging them makes the differential compare a
tree against a baseline from a different commit. The job runs only the qa
suite and the ratchet and invokes no emitted-attribution test, so it needs
no rebase-pinned base at all -- the step was removed rather than paired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): stop monotonic-progress going blind on scope-less payloads

The full remote suite caught a false NEGATIVE I introduced while fixing a
false positive. Silencing the boundary-crossing noise had made the oracle
skip ANY observation lacking milestone fields -- so a minimal payload like
{total_summaries: n} produced no violation at all, and the oracle stopped
catching the exact defect it exists to catch. For a QA tool that is
strictly worse than the noise it replaced.

Scope is only indeterminate when the two observations DISAGREE about
having it:

  both scoped, same scope, decrease -> VIOLATION
  both scoped, different scope      -> reset silently
  NEITHER scoped, decrease          -> VIOLATION   (the regression)
  mixed                             -> skip the comparison

Implementing the mixed case surfaced a second blind spot: advancing the
reference point on a skipped pair lets a scope-less observation sitting
between two same-scope ones mask a real decrease. Mixed now leaves the
reference untouched. All four branches carry explicit coverage; only one
did before, which is why this shipped.

The self-test that failed was right and the code was wrong, so the code
moved. Corpus behavior is unchanged: still 5 smells, 0 new, 0 stale, 0
violations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 16:13:28 -04:00
Tom Boucher
bf0d715733 fix(#2472): cost-balanced test sharding and pinned CI base commit (#2480)
* fix(#2472): weight-aware shard partition

Windows shard 1/3 hit the 20-minute job cap with no failing assertion. Root
cause is the shard layer, not the chunk layer: selectShard partitioned by
sorted ARRAY INDEX (k % n, #1212), which balances file COUNTS and ignores
file COST. On the real unit suite that produced 12.4m / 19.2m / 15.2m — a
1.23x max/ideal ratio leaving the heaviest shard 5% under the cap. Because
assignment keyed off position, inserting one test file re-indexed every file
after it and could tip that shard over; deterministic, so a re-run reproduced
it exactly.

This is NOT the chunk packer (#2456/#2463). That fix works and applies one
level down, WITHIN a shard. The across-shard partition predated it and never
consumed the cost table. Both layers now share one cost model.

selectShard takes an optional weightOf and, when given one, partitions by LPT
(longest-processing-time-first) — the same algorithm packChunks uses. Omitting
it keeps the legacy round-robin byte-identical, so every existing test above
still exercises that path unchanged and callers without timing data lose
nothing. A missing timings table yields uniform weight 1, under which LPT
degenerates to the equal-count split.

Projected on the real suite: 16.4/17.3/13.0 -> 15.6/15.6/15.6 (worst shard
17.3m -> 15.6m).

Tests: a skewed-cost regression (round-robin clusters all four heavy files
onto one shard at 2.98x ideal; LPT does not), back-compat equivalence,
determinism, tie-breaking, order preservation, and two fast-check properties
— the partition is exhaustive and disjoint (getting this wrong silently DROPS
tests from CI, the worst failure mode for a harness), and no shard exceeds
average + heaviest file.

Two assertions were corrected during authoring rather than shipped wrong:
- an initial "LPT within 4/3 of ideal" bound was false. The 4/3 figure is
  relative to the OPTIMAL makespan, not the average, and the two differ when
  item sizes force a pairing. Replaced with Graham's average+max bound, which
  is what is actually provable.
- "weighted is never worse than round-robin" is also false; fast-check
  falsified it with [19316,10190,1,9128,29353,20227] over 2 shards (rr 48670,
  lpt 48671). Round-robin can win by luck on a specific input. Dropped, with
  the counterexample recorded in place so it is not re-asserted later.

Closes #2472

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): rotate tied bins; restore #1212 test block; lazy cost table

Isolated-review findings, all fixed.

HIGH — zero weights collapsed the whole partition onto shard 1. The
lightest-bin scan compared weight only, and adding a zero-weight file leaves
its bin's weight unchanged, so bin 0 stayed tied-minimum forever and every
such file landed on it. Verified: all-zero weights gave shard1=[a..f],
shard2=[], shard3=[] — two of three CI runners idle while one ran everything.
Reachable through safeWeight's own clamp (a NaN/negative/Infinity entry in a
corrupted or hand-edited timings table) and through any genuine 0ms
measurement, so the clamp reproduced the exact failure its comment claimed to
prevent. Ties now break on file COUNT after weight, which rotates. Pinned by
two regression tests (all-zero, and clamped NaN/negative/Infinity) plus a
property over list size x shard count. The live table has no 0ms entries
(min 19ms), so production was not affected — but nothing prevented it.

MEDIUM — the new describe block had swallowed #1212's pre-existing property
test, which is why a test under a "weight-aware" heading never passed a
weigher. That was a bad block boundary in the previous commit, not a bad
test: the #2472 describe was opened before #1212's last test instead of
after. Moved back where it belongs; #1212 is 762-879 and #2472 is 894-1082.

LOW — that relocated property test ran unseeded. Seeded (12120) per the
repo's property-test convention so a failure reproduces. Verified passing
under the new seed.

LOW — hoisting the timings load above the shard block charged a readFileSync
+ JSON.parse to invocations that exit before needing it (empty selection,
--files matching nothing). Now lazily memoized, so neither consumer reads the
table unless it is used and it is still read at most once.

Real-suite projection unchanged at 15.6m / 15.6m / 15.6m.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#2472): correct stale round-robin sharding descriptions

The partition is now cost-balanced, so the header block in run-tests.cjs
and the two comments in test.yml describing '--shard' as a round-robin over
sorted file index were actively wrong. Updated to describe LPT over measured
duration, and to state the degenerate case explicitly: with no timing data
every file weighs the same and the partition collapses back to k % n, which
is why the pre-existing #1212 CLI tests still pass unchanged (their nine
synthetic files are absent from the timings table, so all take the identical
median weight).

Remaining 'round-robin' mentions are correct — they describe the unweighted
fallback path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): shard diagnostics, cost-routing E2E test, table validation

Second orthogonal review (operational lens) findings, all fixed.

HIGH — cross-runner partition divergence. Each of the up-to-12 CI jobs runs
its own 'merge base into head' and computes its own partition, so if the
inputs differ between jobs (the file list, or the timings table) two jobs can
place the same file in different shards or in none. Every job stays
internally exhaustive and disjoint, so nothing errors: a test simply never
runs and CI stays green.

The risk class is pre-existing — round-robin diverges identically when the
file set differs between jobs, which is literally this issue's insertion
instability — but weighting adds tests/test-timings.json as a second input
that must match, so it widens the hole. Properly closing it means pinning the
partition inputs per run, a workflow change beyond this fix.

What IS closed here is the silence. Each shard now prints an input
fingerprint over the FULL pre-partition list and the weight assigned to each
file — deliberately not this shard's slice, which would differ by design and
be useless for comparison. All shard jobs of one run must print an identical
sig; a mismatch is direct proof the runners disagreed about the input.
Verified: three independent computations agree, and the sig changes when the
input drifts by one file.

MEDIUM — nothing proved main() actually threads fileWeightOf() into
selectShard. Every pre-existing --shard E2E test uses synthetic filenames
absent from the real table, so all collapse to a uniform median weight, under
which LPT is mathematically identical to k % n — a typo on that one wiring
line would have passed the whole suite. Added an E2E test that injects a
table via RUN_TESTS_TIMINGS_FILE with differing costs, placing the heavy
files at exactly the indices round-robin hands to shard 1, and asserts shard 1
does NOT receive all three. Plus a test that all three shards emit the same
sig.

MEDIUM/LOW — no observability. The diagnostic line now reports files,
weighed count, aggregate weight, and whether the table loaded, so a table
that silently failed to parse shows table=absent/weighed=0 instead of being
indistinguishable from a healthy load. (The reviewer confirmed the advisory
fallback is already live on next: feat-2296-provider-escalation.test.cjs is
missing from the table.)

LOW — typeof [] === 'object', so a hand-edit turning the map into a list was
accepted as a valid table. Now rejected via Array.isArray, falling back to
uniform weight like any other malformed table.

LOW — stale round-robin wording in ci-test-scope.test.cjs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2472): pin every CI job to one base commit

Closes the cross-runner divergence at its source instead of only making it
visible.

Each job of a run executes the rebase-check step independently, minutes apart
across a 12-job matrix, and merged the MOVING origin/<branch> ref. If the base
advanced mid-run, different jobs merged different trees. That was survivable
when jobs only had to agree on pass/fail; it is not once they must agree on a
PARTITION. Each shard job computes the whole split and keeps its own slice, so
jobs working from different trees can place a file in two shards or in none —
and every job still looks internally consistent, so nothing errors. A test
silently never runs and CI stays green.

ci-rebase-check.cjs now accepts CI_REBASE_BASE_SHA and pins BOTH the fetch and
the merge to that one commit, so the two can never disagree. test.yml passes
github.event.pull_request.base.sha on all three rebase-check steps; that value
is fixed for the life of a run, so all jobs merge the identical base.

This also closes the PRE-EXISTING half of the divergence. Round-robin had the
same exposure whenever the test-file set differed between jobs — that is this
issue's insertion instability — so the pin fixes the older hole too, not just
the timings-table input weighting added.

Only a full 40-hex sha is accepted; empty (push/workflow_dispatch), malformed,
or injected values fall back to the branch ref rather than handing an arbitrary
string to git fetch as a refspec. resolveBaseRefs is extracted pure and
exported, and runMain is guarded behind require.main === module, so the pin
contract is testable without spawning git.

Tests (tests/ci-test-scope.test.cjs): every rebase-check step must carry the
pin; a valid sha pins both refs; absence falls back correctly; and five hostile
values — short sha, uppercase, --upload-pack= injection, ref expression, empty
— are each rejected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 09:15:59 -04:00
Tom Boucher
448e148058 fix(#2455): abort remaining chunks when one hits the per-chunk timeout (#2465)
The per-chunk timeout exists so a bad chunk fails loudly 'rather than
silently burn the job's wall-clock budget until the CI runner cancels the
whole job' (run-tests.cjs:633-637). The control flow defeated that: after a
timeout kill the loop fell through to the next chunk. Since the timeout
(600000ms) is half the 20m job cap and a healthy Windows full pass is
~11m42s, continuing after a timeout can essentially never finish.

Observed on run 29749380190 (windows shard 2/3): chunk 1/5 was killed at
exactly 600s, the loop pressed on through chunks 2-4, and the job was
cancelled mid-chunk-5 at the 20m wall. The failure surfaced as
'##[error]The operation was canceled.' — the timeout diagnostic ended up
~38,000 log lines from the end and 'gh run view --log-failed' returned
nothing, making the real cause very hard to find.

Abort the remaining chunks on a timeout so the diagnostic survives as the
visible failure. Ordinary test failures still run every chunk, so the
operator keeps seeing all failures in one pass.

Fixes #2455

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 16:58:13 -04:00
Tom Boucher
953b8043ea fix(#2456): weight test chunks by measured cost and pack with LPT (#2463)
* fix(#2456): weight test chunks by measured cost and pack with LPT

scripts/run-tests.cjs guessed each test file's cost from its filename
(basename matching /^(?:install|codex-)/ scored 12, everything else 1).
Measured durations show that guess is wrong in both directions:
installer-migration-authoring.test.cjs scored 12 while running ~0.1s, and
the two most expensive files in the suite both scored 1 —
run-tests-harness.test.cjs never matched the prefix, and
release-tarball-smoke.install.test.cjs was missed because the regex is
anchored to the START of the basename.

Chunks were therefore balanced by file COUNT, not cost. On the real
shard 2/3 the two heaviest files packed into the SAME chunk, leaving the
slowest chunk 2.8x the lightest and sitting near the 600s per-chunk
timeout while other chunks idled.

Weight each file by its measured duration from a checked-in, regenerable
timings table and pack with LPT (heaviest first, into the lightest
chunk). On the same shard this drops the slowest chunk from 383s to 238s
and the imbalance from 2.79x to 1.00x, and separates the two heavy files.

Timings are advisory, never gated: an unknown file falls back to the
table's median weight, a missing or corrupt table falls back to uniform
weight, and a count-based floor guarantees the packer never produces
fewer chunks than plain count-based packing would.

Closes #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2456): harden chunk packing against degenerate knobs and table keys

Follow-up hardening found while reviewing the packer, fixed inline.

The chunk knobs are read from the environment with Number(), so a typo
(RUN_TESTS_MAX_FILES_PER_CHUNK=abc) yields NaN and an explicit 0 yields
0. Both flow into the new chunk-count arithmetic: NaN made Math.ceil
return NaN, Array.from({length: NaN}) produce zero bins, and packChunks'
retry loop spin forever — a hung CI job with no output. Zero made the
count Infinity and threw RangeError: Invalid array length. The previous
count-based packer degraded to a single chunk instead, so this was a
regression introduced by the LPT rewrite.

Normalize the knobs at the environment boundary (positiveNumberEnv:
anything not a positive finite number falls back to the default) and
guard packChunks itself, since it is exported and cannot assume its
caller normalized. Non-finite weights from an arbitrary weightOf are
clamped too. RUN_TESTS_CHUNK_TIMEOUT_MS gets the same treatment.

Also resolve timing-table lookups with Object.hasOwn: the table is
JSON-parsed, so a bare index would walk the prototype chain and return a
function for a file named constructor.test.cjs or toString.test.cjs.
The typeof guard already rejected that, but the lookup now resolves
correctly rather than relying on the downstream check.

Refs #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2456): correct prototype-lookup rationale and guard generator keys

Two findings from independent security review, fixed inline.

The makeFileWeigher comment claimed a bare table lookup "would return a
FUNCTION for a file named constructor.test.cjs". That premise is false:
basename('constructor.test.cjs') is 'constructor.test.cjs', which is not
an Object.prototype key, and walkTestFiles only ever collects *.test.cjs.
The prototype chain was never reachable from a real selection, and the
existing typeof guard already rejected the function it would return, so
Object.hasOwn is defense-in-depth rather than a behavior change. The
comment now says that instead of asserting something untrue.

The accompanying test inherited the same false premise: it fed
constructor.test.cjs and asserted a median fallback that would have held
with or without the guard, so it passed for a reason unrelated to what
it claimed to prove. It now uses BARE keys (constructor, toString,
valueOf, hasOwnProperty, __proto__) — the only inputs that actually
resolve on Object.prototype — and asserts the real exported contract:
any key absent from the table weighs the median, never a function.

gen-test-timings.cjs built its output object by computed-key assignment
from basenames taken out of a reporter stream it does not control — the
js/prototype-polluting-assignment shape, and this repo has a CodeQL
barrier for exactly that pattern. It was not exploitable (the value is
always a rounded number, so the __proto__ setter is a silent no-op), but
it silently DROPPED such an entry rather than reporting it. Validate every
key against a test-basename pattern and fail loudly instead, and build
the table with a null prototype.

Refs #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2456): replace tautological chunking tests and clamp chunk count

Six findings from independent correctness review, all reproduced and
fixed inline.

The two subprocess tests written to carry the #2088 guarantee forward
were tautological: every seeded file weighed exactly 1, so both passed
under the OLD prefix-heuristic packer and with the timings file deleted
entirely. Neither could fail for the reason it existed. Both are rebuilt
so the old algorithm produces a different packing and the assertion goes
red: the spread test now uses three expensive files named so the old
heuristic scored them 1 alongside three trivial `install-`-prefixed
files it scored 12 — inverted from real cost, giving {2,2,1,1} under the
old packer versus {2,2,2} under measured weights. The companion test
covers the other direction: four trivial `install-` files the old
heuristic split into four single-file chunks now stay in one.

packChunks clamped the chunk count from below but not above, so a
legitimate but tiny budget (RUN_TESTS_MAX_FILES_PER_CHUNK=1e-9, which
positiveNumberEnv accepts) asked for 637,000,000,000 bins and threw
RangeError. More chunks than files is never useful; the count now clamps
at one file per chunk.

The generator's basename-collision guard compared full dirnames, so two
OS lanes reporting the same file under different container roots
(/work/tests vs C:/work/tests) flagged every shared basename as a
collision — on the script's own documented multi-lane usage. Detection is
now scoped per stream, where the root is constant; a genuine same-lane
collision is still caught.

Also: the LPT tie-break compared raw paths, so a path separator (0x2F vs
0x5C) could order a subdir file differently per platform, contradicting
the documented byte-identical guarantee — it now normalizes separators.
loadTestTimings now honors schema_version instead of writing it and
never reading it, falling back to uniform weight on an unknown version.
A comment claiming an all-uniform suite "chunks exactly as it did
before" was false and contradicted by this PR's own test: the chunk
count is preserved, the composition is not. And the missing-table test
created a temp dir it never cleaned up, for a path that only needed to
not exist.

Refs #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 16:51:35 -04:00
Tom Boucher
8b99f4f3c3 fix(#2088): weight install-heavy test files so they spread across chunks
The targeted CI lane runs changed files UNSHARDED; #2088 touched 13 install-heavy
test files that all landed in one chunk, blowing the 600s per-chunk backstop on
the slow Windows runner (pure slowness, not a leak — per run-tests.cjs's own
comment). Weight install*/codex-* files (~10x a unit file) toward the per-chunk
budget so they spread across chunks instead of clustering; light-file chunking is
unchanged (weight 1). Adds harness regression tests (heavy split vs light control).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 22:49:25 -04:00
Tom Boucher
97972ca7fd fix(#1575): lower MAX_FILES_PER_CHUNK from 90 to 60 to fix macOS Node 22 timeout
Shard 2/3 chunk 2 (~80 files including state.test.cjs, perf-*, worktree-cleanup)
exceeded the 600s per-chunk timeout on macOS Node 22. Reducing the cap from 90
to 60 splits this into two ~40-file chunks, each well within the 600s budget.
Three chunks at ~5 min each = ~15 min, safely under the 20m job cap.
2026-07-06 12:11:08 -04:00
Jeremy McSpadden
e5ef323b15 feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command

Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.

* feat: add /gsd-start smart-entry command

State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.

- src/smart-entry.cts: deterministic situation classifier (no-project,
  paused, blocked, verify-failed, needs-first-phase, planning, executing,
  verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
  ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire  case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
  dispatcher presenting an AskUserQuestion menu (with --text fallback for
  non-Claude runtimes) and dispatching to existing commands. Falls back
  to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
  situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
  (markdown-layer invariants + every emitted command resolves to a real
  slash command).

Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.

* refactor: rename smart-entry command to /gsd:next

Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.

All affected tests (188) pass; lint:ci clean.

* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)

Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.

- detectSignals now reads total_phases + percent from nested progress{}
  first, then scalar fm, then body; current_phase falls back to the
  body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
  body Phase field) covering verify-pending + executing situations.

Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.

* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)

Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.

Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
   layer is down — read .planning/STATE.md directly with the Read tool
   and synthesize a minimal situation + actions menu so /gsd:next stays
   useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
   /gsd:progress as before.

Matches the direct-read resilience the live agent already did by hand.

* docs: add gsd-next skill surface

* chore: trigger no-mistakes validation

* no-mistakes(review): Fix smart-entry phase ordering

* no-mistakes(review): Fix decimal smart-entry phase ordering

* no-mistakes(test): Fix smart-entry next test contracts

* no-mistakes(document): Docs synced for smart entry

* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: shorten next.md description and update golden install parity fixtures

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: trigger no-mistakes validation

* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files

Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.

* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: regenerate golden install parity fixtures for /gsd:next

Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.

* Fix smart-entry verify-failed phase scoping and empty resolve shim step

Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.

* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* fix: recapture all 16 golden fixtures with updated smart-entry.md hash

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* chore: regenerate fixtures + inventory manifest after rebase onto next

Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>

* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next

Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).

Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)

The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs

Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo

Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
  recommended action, 1-4 unique-id /gsd:* actions (previously the
  one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout

Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:

- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
  zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
  kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
  consolidation file doing dozens of real installs — 41s even on a fast Mac
  (much worse on the slow Windows I/O path), plus an install-heavy cluster.

Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.

Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 12:18:25 -04:00
Tom Boucher
b188c0d085 fix(#1967): build hooks/dist once upfront in run-tests to close scoped-CI empty-dir race (#1968)
* fix(#1967): build hooks/dist once upfront in run-tests to close scoped-CI empty-dir race

hooks/dist/ is gitignored and not built by prepare (build:lib only), so the
scoped CI lane starts with it absent. The first install test's before() hook
triggers build-hooks.js, which creates DIST_DIR empty then fills it file-by-
file; a concurrently-spawned install.js reader can observe the empty window and
fail with 'Failed to install hooks: directory is empty' (intermittently failing
e.g. bug-3683-workflow-colon-namespace-leak on scoped legs).

Add ensureBuiltHooks() to scripts/run-tests.cjs — the same upfront chokepoint as
ensureBuiltArtifacts — to build hooks/dist once, single-process, before any
concurrent test spawns install.js. Completeness is checked against
build-hooks.js HOOKS_TO_COPY (absent/empty/partial/zero-byte -> rebuild; complete
-> no-op). Folds regression coverage into bug-969-test-infra-flake-hardening
(Part C), proven fail-first (ensureBuiltHooks undefined on next).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1967): set changeset pr to 1968

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-02 23:07:14 -04:00
Tom Boucher
353f63d170 feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2) (#1440)
* feat(#1431): runtime capability registry overlay (ADR-1244 Phase 2)

Promote the registry from a frozen data file to loadRegistry({includeInstalled}),
composing the first-party registry with a validated installed overlay (ADR-1244 D2):

- Extract the conformance validator to a shared runtime-callable module
  (gsd-core/bin/lib/capability-validator.cjs); the generator re-exports it
  verbatim, guarded by a generative-parity test (no build-time/runtime drift).
- capability-loader.cts: loadRegistry({includeInstalled}) composes first-party
  ∪ validated overlay from $GSD_HOME/.gsd/capabilities (global) and
  <root>/.gsd/capabilities (project) via the canonical buildRegistry. First-party
  always wins (id/skill/agent/config/command-family + reserved gsd-/anthropic-
  prefixes); full merged-set cross-capability validation; engines.gsd load-time
  re-gate (skip-with-warning); gate-kind capabilities FAIL CLOSED; fragment-path
  escapes rejected.
- semverSatisfies (hand-written, no dep) for the engines.gsd gate, fail-closed.
- Wire surface/state + loop to the overlay; loop injects a blocking gate for each
  skipped gate-kind overlay (fail-closed).
- cwd-aware overlay config-key federation: config-loader _federatedConfigSchema(cwd)
  + config-schema isValidConfigKey(key, cwd) compose the overlay per loadConfig/
  config-set call (never eager at module load, never wrong-cwd); first-party path
  unchanged with no cwd.
- run-tests.cjs sandboxes GSD_HOME (idempotent — nested spawns reuse it) for test
  hermeticity; capability-loader.cjs git+eslint-ignored (tsc artifact);
  capability-validator.cjs stays linted (#551 migration coverage).

Closes #1431

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(#1431): add changeset for runtime capability registry overlay

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1431): kill config-schema cwd-aware federation mutants (Stryker ≥52)

The cwd-aware overlay config-key federation added to config-schema.cts
(_capabilityConfigSchema(cwd) + isCapabilityConfigKey/isValidConfigKey cwd
threading) introduced mutable surface uncovered by config-schema's mutation
test set, dropping its score to 39.58% (below the 52 break threshold). Add a
real-overlay-fixture describe block exercising every branch (cwd guard, overlay
loadRegistry, found-branch, first-party fallback, cwd threading); local Stryker
score 39.58% -> 77.08%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-18 14:41:20 -04:00
Tom Boucher
120f85164b feat(#1355): detect-and-warn guard for claude-code agent-teams (#1371)
* feat(#1355): detect-and-warn guard for claude-code agent-teams

GSD's multi-agent orchestration can stall under claude-code's experimental
agent-teams (a subagent's completion fails to route to the orchestrator). Per
the maintainer decision, the accepted scope is a read-only detector + one
non-fatal warning — NOT the declined run_in_background/TaskOutput conversion.

- New Teams Status Module (src/teams-status.cts → gsd-core/bin/lib/teams-status.cjs):
  pure resolveTeamsStatus({runtime, env}) + thin CLI cmdTeamsStatus reusing
  resolveRuntime. active = strictly-truthy env flag AND runtime === 'claude'.
- Wire `gsd-tools query teams-status [--active]` (read-only; no capability
  registration needed — conformance gates govern features, not query commands).
- One non-fatal warning in plan-phase.md before the first Agent spawn, gated on
  `query teams-status --active`; zero behavior change on non-claude/teams-off.
- Hermeticity: clear CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS in run-tests.cjs +
  SESSION_ENV_KEYS. Docs reference + CONTEXT.md glossary. Built lib gitignored.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): add changeset for teams-detect guard

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#1355): bump plan-phase.md workflow size baseline (+407B for teams warning)

The non-fatal agent-teams warning block added to plan-phase.md grew it
92759 → 93166 bytes, past its committed per-file baseline ratchet. The growth
is small, deliberate, and still well under the workflow tier hard cap. Regenerate
the baseline via `npm run size:baseline` (only plan-phase.md changed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#1355): register teams-status.cjs in the inventory manifest

The new teams-status CLI module is a tracked surface; regenerate
docs/INVENTORY-MANIFEST.json (cli_modules family) via
gen-inventory-manifest.cjs --write so the inventory-manifest-sync gate passes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-17 08:52:23 -04:00
Tom Boucher
22f56f4431 ci(#1212): shard windows full-test lane to remove timeout cliff (#1222)
The `full test (windows-latest, *)` lane ran the entire unit suite (~740+
files) in one job whose wall-clock crept against the 20m cap and intermittently
CANCELLED (false-negative gate, observed on PR #1207). Prior tactical fixes
#869 (15→20m bump) and #1051 (handle-leak) deferred the cliff structurally.

Shard the unit suite across 3 parallel runners per OS/node leg so per-job
wall-clock is O(total/3) and stays under the cap as the suite grows.

- scripts/run-tests.cjs: add `--shard <i>/<n>` — a deterministic, balanced
  round-robin partition (fileIndex % n === i-1) over the SORTED selected file
  list. parseShardArg strictly validates i∈1..n, n≥1, integer-only; n=1 is a
  pure no-op. The 28K Windows argv chunking is preserved within each shard. A
  legitimately-empty shard (n > file count) exits 0; a selection empty BEFORE
  sharding still hits the discovery hard error. Composes with --suite and is
  order-independent (sorted before partition). Exports selectShard/parseShardArg.
- .github/workflows/test.yml: test-full becomes the 3 legs × 3 shards = 9-job
  cross-product (explicit include rows — a base shard dim does not cross-product
  with include legs, and a nested matrix.leg.os is unresolvable by the H1
  shell-policy linter). Unit suite runs sharded; integration/security run once
  per leg (shard 1). The Required tests fan-in is unchanged: it already needs
  test-full and checks the matrix-aggregate result, so a failed/cancelled shard
  fails the gate; the branch-protection check name is preserved.
- tests: partition/CLI + pure selectShard contract (completeness, disjointness,
  balance, determinism, boundaries, fast-check property) + parseShardArg
  validation, in run-tests-harness.test.cjs; a DEFECT.GENERATIVE-FIX parity
  guard (per-row shard values 1..N, every leg runs all shards, N == --shard /N
  denominator) + Required-tests name/needs pin, in ci-test-scope.test.cjs.

Closes #1212

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 12:25:29 -04:00
Tom Boucher
5fa4dcd78c fix: recover silently-excluded test dirs + test-architecture audit hardening (#1195)
* fix: recurse test discovery so subdir test suites actually run

scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.

Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: retire 5 verified-worthless tests

Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
  covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
  and bug-782-cline-skills-emission)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: add ADR-218 release version-validation coverage

ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: redesign weak tests into behavioral, deterministic assertions

Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
  bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
  plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
  context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
  core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
  collision) and unconditional plugin.json schema validation (issue-766)

Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: add no-tautological-assert lint rule, error in test suite

New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).

Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: gate new allow-test-rule exemptions to require an issue ref

ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs: add ADR test-audit evidence report (#1192)

Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture

feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at b10e5681 — confirmed: base blob has 1 NUL, this fix has 0), which
made git treat the file as binary and would break grep/editors. Switch to the
\x00 escape; the runtime string value (a real NUL in the parser input) is
unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: address adversarial-review findings

Codex adversarial pass over the branch:
- capability-registry drift test no longer mutates the committed generated
  capability-registry.cjs in place (concurrency hazard) — uses in-memory
  checkPipeline comparison instead.
- allow-test-rule ratchet now detects exemptions in ALL comment forms (block
  /* */ too, matching no-source-grep) so a block comment can't bypass it;
  one newly-surfaced pre-existing offender grandfathered (323->324).
- install.test Kilo case asserts on what install(false,'kilo') actually writes
  rather than manually calling configureKiloPermissions (masked the call site).
- issue-766 drops the undeclared transitive ajv dep for explicit structural
  assertions from the schema fixture.
- adr-218 test notes the hotfix leading-zero gap is tracked in #1186.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: address code-review findings (subdir discovery, rule + test gaps)

xhigh code review surfaced 15 confirmed issues, all fixed:
- run-tests.cjs --files now resolves subdir tests by bare basename + handles
  Windows backslash paths (ambiguous basenames error clearly).
- affected-tests-lib.cjs listTestFiles made recursive — the targeted CI lane was
  silently dropping changed subdir tests (same false-green class the audit fixed).
- no-tautological-assert now catches 'true || cond' and empty []/{}  equality.
- verify-test-quality: restore provenance-classification coverage, tighten the
  writeFile circular-detection check, guard the module-level file read.
- sh-hook-paths: cover the global-install .sh delegation branch (#2045 guard).
- active-workstream null-guard runs deterministically (no longer skipped on TTY).
- adr-218 structural guards tightened (major/minor leading-zero; needs: membership).
- repo-layout AGENTS.md guard no longer false-alarms on equivalent refactors.
- cross-ai ordering guard fails red when the step is missing.
- issue-766 parses required fields from the schema fixture (auto-enforced).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test: stub USERPROFILE alongside HOME in feat-488 (Windows parity)

The feat-488 redesign stubbed process.env.HOME but not USERPROFILE; os.homedir()
resolves from USERPROFILE on Windows, so the home stub was not hermetic there —
caught by windows-test-parity-guard (stubsHomeNoUserProfile). Save/set/restore
USERPROFILE symmetrically with HOME (delete-if-originally-undefined).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: reconcile allow-test-rule allowlist after rebase onto next

Rebasing onto current next pulled in merged PR #1170, which added
inventory-headings-countfree.test.cjs (a baseline allow-test-rule exemption) and
deleted inventory-counts.test.cjs. Grandfather the former and prune the latter so
the ratchet matches the merged tree. No new debt from this PR.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 23:35:08 -04:00
Tom Boucher
fd01e7a12e feat(#1132): complete contribution hook prerequisite
Closes #1132
2026-06-12 18:39:57 -04:00
Tom Boucher
4698b3e349 fix(#1051): force-exit + per-chunk timeout for the windows full-test lane; close leaked test handles (#1054)
The `full test (windows-latest, 22)` job intermittently got CANCELLED at its
20m wall-clock cap with no failed test step — a false-negative gate (recurrence
of #869). Root cause: a unit test leaves an open event-loop handle, so the
chunk's `node --test` child hangs ~150s on Windows after its last test prints;
two such stalls push the already-~13m job past 20m.

Fix (defense in depth):
- run-tests.cjs: pass --test-force-exit (Node >=22; engines requires >=22.0.0)
  so the runner exits once all tests finish regardless of lingering handles —
  the durable backstop. Account for the flag in the argv-length ceiling.
- run-tests.cjs: add a per-chunk execFileSync timeout (default 600000ms, env
  RUN_TESTS_CHUNK_TIMEOUT_MS) that fails loudly with a diagnostic naming the
  chunk's files, so a hung chunk can never silently eat the job budget.
- perf-316 test: terminate both Worker threads on all paths (afterEach +
  finally) so they cannot outlive the test.
- locking-bugs test: kill spawned children in a finally that wraps the whole
  spawn -> waitFor -> barrier-release -> Promise.all sequence, so a barrier
  timeout no longer leaks live child processes.
- Refresh the stale synckit comment (synckit/SDK bridge was removed).

Regression tests in run-tests-harness: a hung chunk hits the per-chunk timeout
and fails with a clear message; force-exit lets a chunk with a leaked handle
exit cleanly.

Closes #1051

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 14:11:50 -04:00
Tom Boucher
adaf3e17d8 fix(#1001): make bug-969 hardening tests hermetic + move build tsbuildinfo out of shipped tree (regression from #996) (#1002)
* fix(#969): make bug-969 hardening tests hermetic and move build tsbuildinfo out of shipped tree (regression from #996)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(#969): self-heal legacy bin-local tsbuildinfo and make sentinel test hermetic (adversarial-review follow-ups)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(changeset): set pr number to 1002

* docs(#1001): record DEFECT.SHARED-ARTIFACT-MUTATION-IN-CONCURRENT-TEST anti-pattern in CONTEXT.md

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-10 13:58:56 -04:00
Tom Boucher
88e30d5342 test(#969): fix stale-build flake (incremental + re-emit-on-missing) and make runGsdTools retry-once before surfacing subprocess kills (#996)
Closes #969

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-10 12:04:59 -04:00
Tom Boucher
a480510f54 fix(#872): make roadmap-phase-fallback tests hermetic against ambient GSD env (#873)
extractCurrentMilestone reads STATE.md via planningDir(cwd), which is
workstream-aware (honours GSD_PROJECT/GSD_WORKSTREAM). The fixtures write
STATE.md to the plain <tmp>/.planning/STATE.md, so a developer shell inside a
GSD workstream (GSD_WORKSTREAM exported) redirected the read to a non-existent
workstream subdir -> version=null -> closed milestone sections leaked into the
slice and assertions failed. Clean CI/Docker env never hit it. Not a Node-26
regex bug; reproduces identically on any Node with GSD_WORKSTREAM set.

- scripts/run-tests.cjs: strip GSD_PROJECT/GSD_WORKSTREAM before spawning test
  children so the local runner env matches clean CI/Docker.
- tests/roadmap-phase-fallback.test.cjs: file-level beforeEach/afterEach
  save/delete/restore of both vars; new regression test pinning workstream-aware
  STATE.md resolution.
- tests/run-tests-harness.test.cjs: guard asserting the runner strips both vars
  (so removing the deletion fails clean CI).

Closes #872

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 12:04:26 -04:00
Tom Boucher
f729101eec refactor(scripts): replace process.exit() with ExitError + runMain handler (#739) (#740)
Part 1 of 2 of the n/no-process-exit cleanup (umbrella #738): convert every
process.exit() call in standalone scripts/** CLIs to the rule-compliant pattern.

- New shared helper scripts/lib/cli-exit.cjs: ExitError(code,message) + runMain()
  which translates a thrown ExitError / returned number into process.exitCode
  (never process.exit()), flushing output and still firing process.on('exit').
- main()-based entrypoints: throw new ExitError(code) for errors, return <code>
  for verdicts; invoked via runMain(main). Child exit codes preserved via return.
- top-level-only scripts: imperative body extracted into main() so mid-flow
  aborts (throw ExitError) actually halt; pure consts/helpers stay at module scope.
- diff-touches-shipped-paths.cjs: stdin event handling restructured to an async
  read so the whole flow runs under runMain; uncaughtException/unhandledRejection
  nets replaced by an in-band catch that preserves EXIT_ERROR=2.

Exit codes verified unchanged for every converted script (success/error/help and
the 0/1/2 semantic codes in diff-touches). Rule stays warn here; flipped to error
in part 2 (#738) once gsd-core/bin/** is also clean.

Refs #739

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 16:13:13 -04:00
Tom Boucher
a41d8d0cf6 fix(#641): teach --files-from to expand bare suite tokens (#647)
When ci-test-scope falls back to the 'unit' sentinel (#408 intent) and
ci-prepare-test-scope writes it verbatim, run-tests --files-from received
a bare 'unit' token that was not a filename, causing exit 2 with
"requested test file(s) not found: unit".

selectExplicitFiles() now recognises any SUITES member and delegates to
the existing selectFiles() resolver before the path-existence check,
reusing the suite expansion logic rather than reimplementing it.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-03 10:11:52 -04:00
Tom Boucher
463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00
Tom Boucher
5cd52eb151 enhancement(#537): pilot TS build-at-publish for bin/lib (semver-compare) (#541)
* docs(#457): rewrite ADR-457 to ground truth and accept build-at-publish

The prior draft asserted a codebase state that never existed (13 tsc-generated
files, src/ trees, a tests/cjs-ts-parity.test.cjs). Corrected to verified ground
truth (84 bin/lib .cjs, 1 value-baked package-identity.cjs, no tsc pipeline),
distinguished value-baking from transpilation so package-identity stops being
miscited as precedent, made check-in-the-artifact vs build-at-publish the central
decision, and flipped status to Accepted (build-at-publish).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* build(#537): pilot TS build-at-publish for bin/lib (semver-compare)

First hand-written module collapsed to a TypeScript source of truth per ADR-457.
src/semver-compare.cts compiles (tsc, strict, noEmitOnError) to a gitignored
get-shit-done/bin/lib/semver-compare.cjs. build:lib is wired into build, pretest,
pretest:coverage, and prepublishOnly so the artifact is built before test and
shipped on publish. Type-aware ESLint on src/**/*.cts immediately caught the
params were over-typed as `unknown` (no-base-to-string); narrowed to a honest
VersionInput domain type. Behavior preserved: semver-compare.test.cjs (14) and
bug-10 (4) pass against the generated output; runtime consumer changeset/cli.cjs
unaffected.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): make build-at-publish robust across all CI paths (codex review)

Adversarial review found the pilot's generated artifact would be missing on
clean CI checkouts. `pretest`/`pretest:coverage` only fire for `npm test`, but
CI runs `test:unit`/`test:integration`/`test:install` and `node run-tests.cjs`
directly — none of which built the artifact, so any suite requiring
semver-compare.cjs would hit module-not-found on a clean checkout, and
install-smoke's `npm pack` could ship without it.

- Add a `prepare` script (`npm run build:lib`). `npm ci` runs it automatically,
  so every CI test job and install-smoke's pack emit the artifact before use.
  This is the idiomatic npm mechanism for compiled-output-not-in-git and fixes
  both the test and pack paths in one place.
- Add `src/` + `tsconfig.build.json` to ci-test-scope and the install-smoke /
  mutation path filters, so a source-only edit to a migrated module still
  triggers its tests and mutation coverage (prevents silent CI skips).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): map src/*.cts to built artifact in mutation changed-files detection

Follow-up to the codex re-review. The prior commit added src/**/*.cts to the
mutation workflow's path trigger but left its "compute changed core lib files"
step diffing only get-shit-done/bin/lib/**/*.cjs — which are now gitignored and
never appear in a diff. A source-only edit would trigger the workflow then
early-exit ("no core lib files changed"), silently skipping mutation testing.

Map each changed src/*.cts to its built get-shit-done/bin/lib/*.cjs path (the
on-disk artifact Stryker mutates after prepare/build:lib), merge with the
hand-written .cjs diff, and apply the test/excluded-module filters once.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): use 'src/' pathspec in mutation diff (git glob doesn't match top-level)

Codex review caught that `git diff -- 'src/**/*.cts'` returns empty for a
top-level file like src/semver-compare.cts — git's default pathspec glob does
not match `**` across zero directories (verified on git 2.50.1). The prior
commit's src-detection therefore never fired, so source-only changes still
skipped mutation. Switch to the dir-scoped pathspec 'src/' + a `.cts` grep
(robust for flat and nested layouts), and broaden the workflow path trigger to
'src/**' to match install-smoke. Verified end-to-end: a change to
src/semver-compare.cts now resolves to get-shit-done/bin/lib/semver-compare.cjs
in the --mutate list.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#537): add changeset fragment for build-at-publish pilot

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): replace prepare with prepack + build-if-missing; defer mutation wiring

CI surfaced three real issues the local run and codex review missed:

1. lockfile-sync failed on every platform. Root cause: `npm ci --dry-run` (the
   repo's lockfile health check) RUNS the `prepare` script, but in dry-run the
   devDependencies aren't installed, so `tsc` is not found (exit 127) and the
   check reports a misleading "out of sync". `prepare` is the wrong hook for a
   build needing a devDep. Replace it with `prepack` (runs only on pack/publish,
   when node_modules exists) for the tarball path, and build the artifact inside
   scripts/run-tests.cjs (build-if-missing) for the test path — the universal
   chokepoint every CI test invocation funnels through, including the direct
   `node run-tests.cjs --files-from` step that bypasses npm lifecycle hooks. The
   guard is a no-op once built, so the run-tests harness test is unaffected.

2. The Stryker mutation gate ran only 1 test against semver-compare (~0% score,
   71/71 mutants surviving) — a Stryker test-selection problem orthogonal to the
   build migration, and raising the score needs property tests (ADR-456). Revert
   the mutation.yml src wiring; mutation coverage for src-authored modules is a
   separate follow-up tracked in #537. (The deletion of the gitignored top-level
   .cjs does not match the workflow's `bin/lib/**/*.cjs` git pathspec, so the
   gate skips cleanly.)

Verified: clean-room `npm ci --dry-run` exits 0; deleting the artifact then
running a suite rebuilds it; run-tests harness 22/22 green; `npm pack` includes
the built artifact via prepack.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-05-31 14:47:48 -04:00
Tom Boucher
8c8887f00e fix(#370): scope affected-tests runner to PR suites, exclude push-only install/slow (#395)
Root cause: the affected-tests runner called runAllSuites() on critical-path
changes (running every suite including install/slow on all matrix cells including
Windows), and pickAffectedTests injected DEFAULT_SMOKE_TESTS (an install test)
as the empty-selection fallback — causing install suite tests to run on PR lanes
where they are push-only per docs/TESTING-SUITES.md.

Fix: PR_EXCLUDED_SUITES filter at the pickAffectedTests chokepoint strips
install/slow from every selection path (direct-change, reverse-index, stem-match).
Empty selection now returns [] and the caller runs the unit suite as smoke.
Critical-path fallback replaces runAllSuites with PR_FULL_SUITES
(unit, integration, security). suiteOf exported from run-tests.cjs
(with require.main guard) so affected-tests-lib reuses canonical detection.

Fixes #370

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:43 -04:00
Tom Boucher
9dd2c87c6c ci: supersede #367 with combined tiered PR pipeline redesign (#369)
* ci: streamline PR pipeline gates

* test: annotate ci-test-scope stderr assertions for lint rule

---------

Co-authored-by: Colin <colin@solvely.net>
2026-05-26 19:11:58 -04:00
Tom Boucher
3169d5cda6 fix(ci): reduce Windows test concurrency 4→2 to prevent synckit worker exhaustion on Node 24 (#173)
Under Node 24 on Windows, running node --test with --test-concurrency=4
causes 4 concurrent gsd-tools subprocesses to each spawn a synckit
worker_threads worker for the SDK bridge. The 4 workers simultaneously
contend on SharedArrayBuffer + Atomics.wait under Windows Defender
scanning and NTFS latency, triggering OS-level resource exhaustion that
kills worker processes with empty stderr before any output is flushed.

The symptom: intermittent exit 1 with 0 test failures, varying affected
test files per run, all sharing the pattern of invoking gsd-tools as a
subprocess. Empty stderr distinguishes OS crash from gsd-tools app error
(the error() path writes to stderr before exiting).

Fix: platform-aware concurrency default — 2 on win32, 4 on Linux/macOS.
The existing TEST_CONCURRENCY env-var override is preserved. Also adds
a [stderr: (empty) exit:N] diagnostic note in helpers.cjs runGsdTools
catch block so future empty-stderr crashes are visible in CI logs.

Fixes gsd-build/get-shit-done#3869

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-23 23:10:06 -04:00
Tom Boucher
7fa5eb7e63 fix(3597): resolve CR threads — drop substring-assertion guidance, use $testDir
Two CodeRabbit threads from PR #3649 review:

- docs/TESTING-SUITES.md:75 — removed the "stable message substring"
  fallback from the error-assertion guidance. Project rule (per the
  no-source-grep lint and lint-no-source-grep.cjs) is structured/typed
  checks only — err.code, JSON fields, enums. Substring matching
  re-introduces the exact prose-coupling we banned.

- scripts/run-tests.cjs:106 — the "no test files found" error now
  reports the resolved testDir variable instead of the hardcoded
  'tests/' string, so when GSD_TEST_DIR points elsewhere the message
  names the actual directory the harness searched.

Local: docker gsd-test-summary 11224/0 on holodeck.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-16 10:44:51 -04:00
Tom Boucher
52f23ac0a0 fix(3597): chunk node --test spawn to survive Windows CreateProcess limit
Windows CreateProcess caps lpCommandLine at 32,767 chars. The original
`execFileSync(node, ['--test', ...546 paths])` exceeded that on every
Windows runner and exited within ~70ms with no test output. Linux/macOS
allow ~2 MB ARG_MAX so the same call worked there.

`scripts/run-tests.cjs` now splits selected files into chunks that keep
each spawn's argv under 28,000 chars (operator-overridable via
RUN_TESTS_MAX_CMDLINE_CHARS), runs them sequentially, and reports the
first non-zero exit. Cross-platform regression test forces chunking with
a low ceiling and asserts the `run-tests: chunk N/M …` stderr marker.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-16 09:25:36 -04:00
Tom Boucher
e82876fe45 feat(3597): split test suites and add Node 22/24/26 OS matrix
scripts/run-tests.cjs gains `--suite <name>` filtering using a filename
suffix convention (`*.security.test.cjs`, `*.integration.test.cjs`, …).
Files with no marker are `unit` (the default fast lane); files with a
marker land in the matching suite. No `--suite` flag preserves the prior
behavior of running every test (backcompat for `npm test` and
`npm run test:coverage`).

New package scripts wire the suites to stable entrypoints:
test:unit, test:integration, test:install, test:security, test:slow,
test:coverage:unit, test:coverage:all. Unknown suite → exit 2 with the
list of valid suites; empty suite → exit 0 with a stderr notice so empty
lanes (e.g. `security` before adversarial tests land) don't gate CI.

CI matrix grows from `ubuntu × {22,24}` + a single macOS lane to
`{ubuntu, macos, windows} × {22, 24, 26}`. `fail-fast: false` so one
lane failure doesn't cancel siblings. Node 26 is `continue-on-error`
until actions/setup-node stabilises that image. PR CI runs unit +
integration + security on every cell; `install` and `slow` only on
`main` push. A dedicated `coverage` job runs `test:coverage:unit` on
ubuntu/Node 24 and uploads the report.

Grouping policy lives in docs/TESTING-SUITES.md with a pointer from
CONTRIBUTING.md. New harness test covers arg parsing, filter selection,
empty-suite behavior, and failure propagation.

Closes #3597.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-16 08:34:09 -04:00
Tom Boucher
2703422be8 refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md (#1675)
* refactor(tests): standardize to node:assert/strict and t.after() per CONTRIBUTING.md

- Replace require('node:assert') with require('node:assert/strict') across
  all 73 test files to enforce strict equality (no type coercion)
- Replace try/finally cleanup blocks with t.after() hooks in core.test.cjs
  and hooks-opt-in.test.cjs per the test lifecycle standards
- Utility functions in codex-config and security-scan retain try/finally
  as that is appropriate for per-function resource guards, not lifecycle hooks

Closes #1674

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* perf(tests): add --test-concurrency=4 to test runner for parallel file execution

Node.js --test-concurrency controls how many test files run as parallel child
processes. Set to 4 by default, configurable via TEST_CONCURRENCY env var.
Fixes tests at a known level rather than inheriting os.availableParallelism()
which varies across CI environments.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(security): allowlist verify.test.cjs in prompt-injection scanner

tests/verify.test.cjs uses <human>...</human> as GSD phase task-type
XML (meaning "a human should verify this step"), which matches the
scanner's fake-message-boundary pattern for LLM APIs. This is a
false positive — add it to the allowlist alongside the other test files
that legitimately contain injection-adjacent patterns.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-04 14:29:03 -04:00
Lex Christopherson
02a5319777 fix(ci): propagate coverage env in cross-platform test runner
The run-tests.cjs child process now inherits NODE_V8_COVERAGE from the
parent so c8 collects coverage data. Also restores npm scripts to use
the cross-platform runner for both test and test:coverage commands.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 10:07:02 -06:00
Lex Christopherson
ccb8ae1d18 fix(ci): cross-platform test runner for Windows glob expansion
npm scripts pass `tests/*.test.cjs` to node/c8 as a literal string on
Windows (PowerShell/cmd don't expand globs). Adding `shell: bash` to CI
steps doesn't help because c8 spawns node as a child process using the
system shell. Use a Node script to enumerate test files cross-platform.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-27 10:00:26 -06:00