Tom Boucher 98866a0c69 feat(#1187): per-module Stryker mutation-score ratchet (ADR-456 80% floor) (#1200)
* feat(#1187): per-module mutation-score ratchet + graduate core-utils

ADR-456's 80% mutation floor was unenforceable as a single global break=50:
4 of 6 covered modules sit at 63-79% and forcing them to 80 would require
brittle exact-string assertions on equivalent string-literal mutants (a
Goodhart's-Law trap). Instead, each covered module declares a minScore floor
(locked at its measured score, TARGET 80) enforced per CI shard via
stryker --break, ratcheting up over time without brittle tests.
- mutation-matrix.cjs: minScore per module + TARGET_MUTATION_SCORE=80, emitted
  in the matrix; require.main guard + exports for testability.
- mutation.yml: per-shard --break <minScore>.
- stryker.config.mjs: global break 50->60 as a local backstop (CI uses minScore).
- Graduated core-utils (measured 77.5%, floor 75).
- context-utilization 79.5->92.3% via behavioral killers (state classification
  outputs + error-value contract, not exact-string matches) -> minScore 80 (TARGET).
- ratchet-integrity guard test (28 cases).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1187): pass mutation break via MUTATION_BREAK env (no stryker --break flag)

Adversarial review caught that Stryker 9.x has no --break CLI flag, so the
per-shard 'stryker run --break <minScore>' errored out every mutation shard.
Read the per-module floor from process.env.MUTATION_BREAK in stryker.config.mjs
and set it per shard via env in mutation.yml. Red-green verified: MUTATION_BREAK=99
exits 1, =80 exits 0. Also make the ratchet guard monotonic (RATCHET_BASELINE
floors; lowering a floor now fails the guard unless the baseline is edited).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1187): fail closed on bad MUTATION_BREAK + monotonic ratchet baseline

Code review: Number(env)||60 failed OPEN — an empty/invalid MUTATION_BREAK
(e.g. a future module missing minScore -> matrix expands to '') silently
degraded the shard to break 60, letting a high-floor module regress undetected.
resolveMutationBreak() now returns 60 only when the env is truly unset (local
backstop) and THROWS on present-but-empty/non-numeric/out-of-range (fail closed);
stryker.config.mjs imports it via createRequire. Also make RATCHET_BASELINE an
equality mirror (=== not >=) so any floor change is explicit in review and no
floor can be silently lowered. Tests: 46 (incl resolveMutationBreak cases).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#1187): recalibrate config-schema/prompt-budget floors to CI scores

First CI mutation run failed two shards: the floors were set from local Stryker
runs whose TIMEOUTS were counted as kills (env-variable), inflating scores. CI
runs with timeout~0, so the real deterministic scores are lower:
- config-schema: local 69.7% -> CI 54.55% (5 local timeouts vanished) -> floor 52
- prompt-budget: local 99.6% -> CI 68.33% (239 local timeouts vanished) -> floor 66
Calibrate floors from CI (the documented source of truth) and record the lesson
in the comment so future floors aren't set from timeout-inflated local runs.
Baseline updated to match. The other 5 shards passed (deterministic CI scores
above their floors).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-14 08:25:41 -04:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Gemini CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Gemini CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Gemini CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start your first project:

/gsd-new-project

New here? Follow Your first project for a guided walkthrough from install to first shipped phase.


Documentation

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 79 MiB
Languages
JavaScript 82.2%
TypeScript 17.4%
Shell 0.4%