Tom Boucher a0f8f956c4 enhance(#4139): Phase 3 — partition rules + the five checks (#4497)
* enhance(#4139): Phase 3 — partition rules + the five checks

ADR-4139 Decision 5, epic #4139 Phase 3. Issue #4403's own "Proposed behavior"
section lists four checks; the ADR's Decision 5 and its own phase table ("partition
rules + the five checks") list five — the same four plus "boundary moves are
declared, ongoing". Same issue-vs-ADR drift Phase 2 hit on the detail.md vs
detail/*.md layout: the ADR is the locked, reviewed document, so it wins. This PR
implements all five.

docs/PARTITION-RULES.md (new) is the partition-rules document: the partition rule
itself, the protected-content list and <!-- gsd:protected --> sentinel syntax
(relocated unchanged from gsd-core/references/compact-content-protected-content.md,
now deleted — it was never referenced by any runtime workflow Read, only by the
predecessor test as documentation, so nothing at runtime regresses, and removing it
from gsd-core/references/ also drops it from all 19 installed-project shipped-content
trees for a file nothing ever read), and the five checks explained for a human
reader. Referenced from a new CONTRIBUTING.md subsection under "Editing shipped
content".

tests/helpers/compact-content-split.cjs (new) is the shared mechanics: split
discovery (any gsd-core/workflows/<name>/detail/*.md paired with <name>.md — no
registry file, a pair is registered by existing on disk), line normalization
(carries forward Phase 2's bare-label-line isTrivial fix and the canonical
gsd_run-launcher-preamble exclusion), sentinel extraction, and a
Boundary-Move-Declared commit-trailer reader that is a direct structural port of
tests/helpers/emitted-runtime.cjs's Emitted-Drift-Ack-Hash/-Growth trailer reader
(ADR-3942) — same merge-base range, same fail-closed throw on an uncomputable range,
same dedupe/conflict rules.

tests/compact-content-partition-guard.test.cjs (new) is the actual guard, superseding
tests/plan-phase-compact-split.test.cjs (deleted — its per-pair checks are now the
general guard's job for plan-phase specifically). Checks 2 (disjointness) and 3
(registration + size cap) run unconditionally against every registered split. Checks
1 (completeness, fires once per split on the PR that introduces a new detail/ path),
4 (protected content — no trailer can ever excuse this one, unlike check 5) and 5
(boundary moves declared) are PR-diff-scoped against the resolved base ref and skip
cleanly when there's nothing to compare (a fresh clone, no PR in flight) — a
deliberate asymmetry from check 5's trailer reader, which must throw rather than
silently pass when ITS range is uncomputable, since that function is answering "did
this PR declare its moves" rather than "is there even a diff to look at". Each of the
five checks carries a RED (deliberately broken fixture) / GREEN (fixed) test pair,
built against synthetic temp files or real throwaway git repos, per this repo's rule
that a guard nobody has seen go red is not yet a guard. Building the real fixtures
caught and fixed one real bug before it shipped: check 4's line-presence test was
using the trivial-line-filtered normalizer, so a byte-identical spine falsely
reported its own protected code-fence line as "deleted" — fixed with a
non-filtering membership check.

Extending docs/INVENTORY.md's "Workflow Sub-Files" table for `detail` surfaced a
pre-existing, unrelated gap in the SAME area: gsd-core/workflows/<name>/templates/*.md
is a fourth workflow sub-file kind that already existed on disk and was already known
to lint-response-language-coverage.cjs's FRAGMENT_DIRS, but was invisible to
gen-inventory-manifest.cjs and undocumented in that table. Fixed alongside it, same
pattern, same PR, rather than deferred.

Also, mechanically required by the new fourth sub-file kind:
- scripts/lint-response-language-coverage.cjs: `detail` added to FRAGMENT_DIRS
  alongside modes/steps/templates — a detail/<part>.md inherits its parent's
  response_language coverage through the same per-file proof, not a parallel one.
- tests/workflow-size-budget.test.cjs: explicit regression test locking that
  detail/ files are governed solely by the hard, non-waivable NEW_FILE_CAP
  (tests/helpers/emitted-diff.cjs) and never by the XL/LARGE/DEFAULT spine tiers —
  true by construction (measureWorkflows/listWorkflowStems don't recurse), made
  explicit per the issue's own Done-when item rather than left true-by-omission.
- scripts/gen-inventory-manifest.cjs: `workflow_detail` and `workflow_templates`
  NESTED_FAMILIES entries; docs/INVENTORY-MANIFEST.json regenerated
  (plan-phase/detail/elaboration.md, discuss-phase/templates/*.md now tracked);
  docs/INVENTORY.md's table updated to four kinds.

Verified: `npm run lint:ci` clean with the eslint cache cleared.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): review findings + a real gsd-test failure in the new guard

Two orthogonal review passes (Standards + Spec, isolated sub-agents) plus a
separate security review ran against the prior commit. Fixed everything each
surfaced:

- Security (Low, path-traversal existence oracle): checkRegistration's
  dangling-reference check extracted detail-path-shaped substrings from spine
  PROSE via a regex that permits `.`/`/` freely, then joined them onto repoRoot
  and probed fs.existsSync with no containment check — a spine file containing
  `../../../etc/detail/passwd.md`-shaped text could make the guard test file
  existence outside the repo. Added a path.relative-based containment check
  before the fs.existsSync call; anything that resolves outside repoRoot is now
  reported as a dangling reference directly, never probed on disk.
- Standards (Boundary Coverage): the size-cap fixtures covered NEW_FILE_CAP and
  NEW_FILE_CAP-1 but not NEW_FILE_CAP+1 — added the third boundary-point case
  CLAUDE.md's TEST RULES require (limit-1/limit/limit+1).
- Standards (Property-Based Testing): extractProtectedBlocks (a sentinel
  parser) and the new parseBoundaryMoveTrailerValues (a declare/dedupe/conflict
  parser, bijective-shaped) had no fast-check property test. Added three: a
  render/parse bijectivity property for the trailer parser (mirroring the exact
  ADR-3942 sibling test's alphabet/idiom), a dedupe-is-idempotent property for
  the same parser, and a well-formed-sentinel-round-trips property for
  extractProtectedBlocks.

Then dispatched gsd-test on the resulting commit. It found a real bug the
reviews couldn't have caught (none of them can run inside gsd-test's sandbox):
checks 4/5's real-repo assertion failed against plan-phase's own split,
reporting DISK_PLANS/#3218-comment lines as "undeclared boundary moves" —
content Phase 2 (#4402) legitimately moved into detail/elaboration.md months
before this PR's Boundary-Move-Declared mechanism existed to require a
trailer for it. Root cause: `resolveBase()`'s own doc comment already documents
that no `origin/*` remote-tracking ref exists inside the gsd-test sandbox
container, and its fallback candidate (a bare `next` branch) can resolve to a
point in history that predates an already-merged, already-reviewed split —
making that split look "newly introduced" from the sandbox's vantage point.
Check 1 (completeness) already scopes itself correctly to only genuinely-new
detail paths (git diff status 'A'); checks 4 and 5 did not share that scoping,
so a stale base made them re-litigate a settled split retroactively. Fixed by
having checks 4/5 skip any split name check 1 already counted as newly-split —
their own premise ("did an EXISTING split shed/undeclare something") does not
apply to a split that is, from the resolved base's vantage point, brand new;
that is check 1's domain alone. Verified locally (25/25 tests pass via a
direct `node -e` require, since `node --test` is blocked in this repo) and via
re-reasoning through the exact real-repo scenario the gsd-test failure showed.

Also regenerated all 19 tests/fixtures/install-tree/*.json goldens — the
prior commit's deletion of gsd-core/references/compact-content-protected-content.md
was never reflected there, which is what golden-install-tree.test.cjs's other
19 failures in the same gsd-test run were.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* docs(#4403): backfill changeset pr number to 4497

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* fix(#4403): isolate codex-config.test.cjs into its own chunk, root-causing the Windows CI failure

PR #4497's "full test (windows-latest, 24, shard 2/3)" job failed: run-tests
killed chunk 3/8 at the 600s per-chunk backstop, with codex-config.test.cjs
(weight 17.87, by far the chunk's dominant cost) packed alongside 39 other
files. Traced, not assumed:

- scripts/run-tests.cjs's own timeout-headroom comment for the OUTER
  per-shard timeout documents that "adding one test file reshuffled 115 of
  268 unit files between shards" — shard/chunk composition is architecturally
  known to be unstable to single-file additions, which is exactly what this
  PR's own new tests/compact-content-partition-guard.test.cjs is.
- A second comment, dated 2026-09-06 (one day before this PR, PR #4428's own
  CI), already documents the SAME chunk hitting the SAME 600s backstop with
  the SAME file (codex-config.test.cjs, "a genuinely MEASURED weight of
  17.87 — not a stale-table miss") dominating it — the fix then was cutting
  the Windows per-chunk budget from 60 to 40. That cut clearly was not
  enough: two documented incidents in two days, at two different budget
  settings, both centered on one file that alone consumes ~45% of even the
  reduced Windows budget.
- tests/test-timings.json's own header confirms its source data
  (test-events-linux-node22/24.jsonl) is Linux-only, and run-tests.cjs's own
  chunk-timeout diagnostic already prints "real Windows cost runs ~2.2x the
  recorded figure" — the packer's weight-balancing is working off data that
  is both stale (table last regenerated 2026-08-07) and known to
  underestimate the platform where the failure occurs.

Given codex-config.test.cjs is disproportionately heavy AND every companion
sharing its chunk is decided by a packing algorithm already documented as
reshuffling unpredictably on any new file, tuning the shared budget a third
time only moves the marginal line to wherever the next new file happens to
land — it does not remove the gamble. Isolating codex-config.test.cjs into
its own dedicated single-file chunk, unconditionally and on every platform,
removes it at the source: the file never enters the pool packChunks balances,
so no other file's packing changes, and no future single-file addition
(mine or anyone else's) can silently reintroduce this exact failure by
landing in its chunk.

Extracted as a small pure function, partitionIsolatedFiles (mirroring this
file's existing pattern of pulling packing/analysis logic out of main() for
in-process unit coverage — see computeSweepProtectSet, analyzeChunkEvents),
with 6 new tests in tests/run-tests-harness.test.cjs covering basename
matching across path separators, near-miss non-matches, the empty-list case,
and the isolated-set contents.

Root cause is now closed rather than papered over with a retry: this failure
is a property of one specific heavy file's chunk placement, not something
that recurs randomly. If codex-config.test.cjs itself is ever genuinely sped
up, this isolation can be revisited — this is a packing-side mitigation for
a known file's cost, not a claim the cost is irreducible.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-07 16:44:53 -04:00
2026-09-06 02:09:28 +00:00
2026-09-06 02:09:28 +00:00
2026-09-06 02:09:28 +00:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 79 MiB
Languages
JavaScript 82.2%
TypeScript 17.4%
Shell 0.4%