Tom Boucher 953b8043ea fix(#2456): weight test chunks by measured cost and pack with LPT (#2463)
* fix(#2456): weight test chunks by measured cost and pack with LPT

scripts/run-tests.cjs guessed each test file's cost from its filename
(basename matching /^(?:install|codex-)/ scored 12, everything else 1).
Measured durations show that guess is wrong in both directions:
installer-migration-authoring.test.cjs scored 12 while running ~0.1s, and
the two most expensive files in the suite both scored 1 —
run-tests-harness.test.cjs never matched the prefix, and
release-tarball-smoke.install.test.cjs was missed because the regex is
anchored to the START of the basename.

Chunks were therefore balanced by file COUNT, not cost. On the real
shard 2/3 the two heaviest files packed into the SAME chunk, leaving the
slowest chunk 2.8x the lightest and sitting near the 600s per-chunk
timeout while other chunks idled.

Weight each file by its measured duration from a checked-in, regenerable
timings table and pack with LPT (heaviest first, into the lightest
chunk). On the same shard this drops the slowest chunk from 383s to 238s
and the imbalance from 2.79x to 1.00x, and separates the two heavy files.

Timings are advisory, never gated: an unknown file falls back to the
table's median weight, a missing or corrupt table falls back to uniform
weight, and a count-based floor guarantees the packer never produces
fewer chunks than plain count-based packing would.

Closes #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2456): harden chunk packing against degenerate knobs and table keys

Follow-up hardening found while reviewing the packer, fixed inline.

The chunk knobs are read from the environment with Number(), so a typo
(RUN_TESTS_MAX_FILES_PER_CHUNK=abc) yields NaN and an explicit 0 yields
0. Both flow into the new chunk-count arithmetic: NaN made Math.ceil
return NaN, Array.from({length: NaN}) produce zero bins, and packChunks'
retry loop spin forever — a hung CI job with no output. Zero made the
count Infinity and threw RangeError: Invalid array length. The previous
count-based packer degraded to a single chunk instead, so this was a
regression introduced by the LPT rewrite.

Normalize the knobs at the environment boundary (positiveNumberEnv:
anything not a positive finite number falls back to the default) and
guard packChunks itself, since it is exported and cannot assume its
caller normalized. Non-finite weights from an arbitrary weightOf are
clamped too. RUN_TESTS_CHUNK_TIMEOUT_MS gets the same treatment.

Also resolve timing-table lookups with Object.hasOwn: the table is
JSON-parsed, so a bare index would walk the prototype chain and return a
function for a file named constructor.test.cjs or toString.test.cjs.
The typeof guard already rejected that, but the lookup now resolves
correctly rather than relying on the downstream check.

Refs #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2456): correct prototype-lookup rationale and guard generator keys

Two findings from independent security review, fixed inline.

The makeFileWeigher comment claimed a bare table lookup "would return a
FUNCTION for a file named constructor.test.cjs". That premise is false:
basename('constructor.test.cjs') is 'constructor.test.cjs', which is not
an Object.prototype key, and walkTestFiles only ever collects *.test.cjs.
The prototype chain was never reachable from a real selection, and the
existing typeof guard already rejected the function it would return, so
Object.hasOwn is defense-in-depth rather than a behavior change. The
comment now says that instead of asserting something untrue.

The accompanying test inherited the same false premise: it fed
constructor.test.cjs and asserted a median fallback that would have held
with or without the guard, so it passed for a reason unrelated to what
it claimed to prove. It now uses BARE keys (constructor, toString,
valueOf, hasOwnProperty, __proto__) — the only inputs that actually
resolve on Object.prototype — and asserts the real exported contract:
any key absent from the table weighs the median, never a function.

gen-test-timings.cjs built its output object by computed-key assignment
from basenames taken out of a reporter stream it does not control — the
js/prototype-polluting-assignment shape, and this repo has a CodeQL
barrier for exactly that pattern. It was not exploitable (the value is
always a rounded number, so the __proto__ setter is a silent no-op), but
it silently DROPPED such an entry rather than reporting it. Validate every
key against a test-basename pattern and fail loudly instead, and build
the table with a null prototype.

Refs #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#2456): replace tautological chunking tests and clamp chunk count

Six findings from independent correctness review, all reproduced and
fixed inline.

The two subprocess tests written to carry the #2088 guarantee forward
were tautological: every seeded file weighed exactly 1, so both passed
under the OLD prefix-heuristic packer and with the timings file deleted
entirely. Neither could fail for the reason it existed. Both are rebuilt
so the old algorithm produces a different packing and the assertion goes
red: the spread test now uses three expensive files named so the old
heuristic scored them 1 alongside three trivial `install-`-prefixed
files it scored 12 — inverted from real cost, giving {2,2,1,1} under the
old packer versus {2,2,2} under measured weights. The companion test
covers the other direction: four trivial `install-` files the old
heuristic split into four single-file chunks now stay in one.

packChunks clamped the chunk count from below but not above, so a
legitimate but tiny budget (RUN_TESTS_MAX_FILES_PER_CHUNK=1e-9, which
positiveNumberEnv accepts) asked for 637,000,000,000 bins and threw
RangeError. More chunks than files is never useful; the count now clamps
at one file per chunk.

The generator's basename-collision guard compared full dirnames, so two
OS lanes reporting the same file under different container roots
(/work/tests vs C:/work/tests) flagged every shared basename as a
collision — on the script's own documented multi-lane usage. Detection is
now scoped per stream, where the root is constant; a genuine same-lane
collision is still caught.

Also: the LPT tie-break compared raw paths, so a path separator (0x2F vs
0x5C) could order a subdir file differently per platform, contradicting
the documented byte-identical guarantee — it now normalizes separators.
loadTestTimings now honors schema_version instead of writing it and
never reading it, falling back to uniform weight on an unknown version.
A comment claiming an all-uniform suite "chunks exactly as it did
before" was false and contradicted by this PR's own test: the chunk
count is preserved, the composition is not. And the missing-table test
created a temp dir it never cleaned up, for a path that only needed to
not exist.

Refs #2456

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 16:51:35 -04:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%