Tom Boucher 4e70b245e8 fix(#3582): stop the cold-tree fixture racing concurrent hook builds (#3656)
* fix(3582): stop the cold-tree fixture racing concurrent hook builds

Two tests in tests/gsd-check-update-worker-platform-gate.test.cjs failed a verification run
with `ENOENT: no such file or directory, lstat '/work/hooks/.dist-staging-20836'`. This is
a race I introduced in #3582, not a flake, and it passed when #3582 merged because it only
fires when the timing lines up.

buildColdInstallTree() copied the LIVE repo hooks/ directory with a filter that excluded
only the basename 'dist'. scripts/build-hooks.js writes atomically through a per-PID
staging dir, hooks/.dist-staging-<pid>, and removes it when finished — and the archived
build-hooks-atomic-write changeset records that NINE test files invoke build-hooks.js from
their before() hooks. So several test processes create and delete staging directories
inside hooks/ while other tests are reading it. cpSync enumerated one, and the owning
process removed it before cpSync got to it.

The helper's own header already states the rule it needed: hooks/dist is excluded because
it "is not present in a raw marketplace checkout either". hooks/.dist-staging-* is
gitignored (.gitignore:21) and equally absent from a raw checkout — it was simply missed.

Fixed by enumerating hooks/ explicitly and skipping 'dist' and any '.dist-staging' prefix
BY NAME, before anything stats or copies the entry, then copying each surviving entry
individually. A name-first skip means a vanishing staging dir is never touched at all.

Worth recording because it corrects the assumption this fix was written under: cpSync's
filter IS invoked before the entry is lstat'd, and returning false leaves it untouched
(verified by deleting inside the callback and returning false — no throw). So merely adding
'.dist-staging' to the old filter would also have closed the race. The explicit enumeration
was kept anyway so correctness does not depend on that Node implementation detail.

Proven by execution both ways: with a staging dir planted in hooks/, the OLD
cpSync-with-filter form copied it straight through into the fixture, while the new form
succeeds and produces no .dist-staging entry with the real hook set intact.

Regression test added beside the existing cold-tree tests: it plants a real
hooks/.dist-staging-test-<random>, asserts the fixture builds clean without it, and removes
only the directory it created.

Repo swept for the same exposure: this helper is the only place doing a bulk enumeration of
the whole live hooks/ tree. The other hooks/-touching tests reference specific named files
or hooks/dist/ and are not exposed. scripts/build-hooks.js is deliberately untouched — its
per-PID staging is what makes its own writes atomic and is correct.

Refs #3582

* fix(3582): make the race regression test hermetic instead of mutating the live tree

The regression test added in the previous commit failed the runner with "failed running
after hook", and it was wrong in two ways — the second one worse than the first.

cleanup() (tests/helpers.cjs:452-487) deliberately THROWS for any path outside the known
temp roots. The test planted hooks/.dist-staging-test-<random> inside the repo and then
asked cleanup() to remove it, so the after-hook threw. That guard is correct and is left
alone.

The real problem is that the test mutated the LIVE hooks/ directory while other test files
concurrently read it — the exact shared-state hazard this change exists to remove. A
regression test for a race must not introduce one.

buildColdInstallTree now takes an optional opts.repoRoot (defaulting to the real REPO_ROOT
and used for both copies it performs), so the test builds a fake repo root under the temp
dir, plants representative hooks plus dist/ and .dist-staging-99999/ THERE, and asserts the
fixture excludes both. All six pre-existing callers pass no arguments and are unaffected.
The test also asserts the real hooks/ listing is identical before and after, so a future
edit that reintroduces live-tree mutation fails loudly.

The name rule is now pinned directly rather than only through the copy. shouldCopyHookEntry
is exported and asserted, including the two cases a sloppier implementation would get
wrong: 'dist-staging-no-dot' and 'distant.js' must both be KEPT. Anything matching on a
loose 'dist' substring or startsWith passes every other case and fails those two.

Also corrected the issue number on the tests introduced here: they were labelled #3631,
which is the unrelated capability-consent bytecode work. This is #3582.

Verified by execution: the predicate rule holds on all nine cases; a fake-root fixture
yields exactly the representative hooks with dist and .dist-staging excluded; the no-arg
default still copies the real tree (29 entries); and the real hooks/ listing is byte-identical
before and after.

Refs #3582

* chore(3582): re-trigger CI after an orphaned Validate Branch Name run

The Validate Branch Name run for this branch (32211622051) sat queued from 03:17 and was
never picked up — updatedAt never advanced past createdAt while the same workflow completed
normally for other branches. `gh run rerun` refused it ("already running") and
`gh run cancel` returned HTTP 500, so the run is orphaned on the GitHub side.

Closing and reopening the PR re-fired the other pull_request workflows but not that one,
whose triggers evidently do not include reopened. An empty commit is the remaining way to
get a fresh run.

No file changes: the tree is identical to dc71534b6, whose remote-runner pass carries
forward unchanged.

Recording this rather than admin-merging past the pending check. Everything else was green
(24 pass, 0 fail), but admin merge is sanctioned only for the missing-secondary-reviewer
case, never to skip a gate that has not actually run.

Refs #3582

---------

Co-authored-by: sim <sim@local>
2026-08-19 01:33:40 -04:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%