Tom Boucher 8a6d87538f test(#3466): replace 8 source-grep assertions with behavioral tests (#3500)
Phase 2 of #3464, following #3465. Rewrites every assertion that read a
shipped .cjs/.js file and text-searched it, so the file no longer needs an
allow-test-rule exemption. 6 of the 8 files are now marker-free; the ceiling
drops 285 -> 278 against a measured 277.

A source-grep passes when a STRING is present, not when the code WORKS. It
survives a refactor that keeps the string but breaks the behavior, and breaks
on a refactor that keeps the behavior but renames the string. Both failure
modes are silent about the thing the test claims to protect. That is the
anti-pattern ADR-456 and local/no-source-grep exist to prevent.

Rewrites, each against the real exported seam:

- discuss-mode: calls cmdInitPlanPhase() against a fixture whose config sets
  workflow.text_mode, asserts the value propagates to its emitted JSON.
- effort-surface-axis: runs review-lane invoke against a real project with a
  fake `claude` PATH shim, asserts the resolved --effort actually lands in the
  shim's captured argv.
- install-minimal-hooks: calls applySettingsJsonHooks() with a hook source
  missing, asserts it is neither registered nor silently registers wholesale,
  with sibling present hooks as the positive control.
- install: calls the exported inferPreferredRuntime({fs, env,
  preferredConfigDir}) via its injected fs seam, asserting 'kilo' from both
  the config-marker and env-var paths.
- opencode-permissions: spawns the real installer with a custom config dir,
  asserts the written opencode.json permission paths are anchored on it.
- repo-layout: spawns the installer for copilot local vs global, asserting
  AGENTS.md is written only in the local case.
- runtime-config-adapter-registry: stubs resolveInstallPlan and force-reloads
  install.js, asserting the runtime's artifact stops being written -- proving
  install.js genuinely routes through the registry.
- runtime-homes-descriptor-drive: calls buildAgentSkillsBlock() for cursor and
  claude against real fixture SKILL.md files, asserting each runtime's refs
  land under its own config dir and never the other's.

Every rewrite was mutation-checked before its marker was dropped. Because
node --test is not runnable locally in this repo, each assertion was
replicated in a standalone probe that requires the same module: run green
against the real file, then red against a deliberately broken one (text_mode
propagation deleted, existsSync guard removed, config dir hardcoded, !isGlobal
guard dropped, kilo branch removed, effort resolution bypassed, skills base
hardcoded back to .claude), then the production file restored and confirmed
byte-identical. An assertion that could not be made to fail would not have
shipped -- a behavioral test that passes regardless of correctness is strictly
worse than the source-grep it replaces, because it looks rigorous while
asserting nothing.

Three claims were checked rather than trusted. All three were wrong:

- repo-layout's own marker cited #1188 asserting the `!isGlobal` lexical scope
  was "unprovable at runtime". It is provable: the guard decides whether
  AGENTS.md is written, which is directly observable. Both directions verified.
- The triage for runtime-config-adapter-registry claimed its source-grep was
  redundant with the file's EXPECTED_TABLE tests, so deletion would be safe.
  Those tests only exercise resolveInstallPlan() directly and never load
  bin/install.js, so they do not cover it. A real behavioral assertion was
  written instead of deleting coverage.
- An earlier revision of this change dropped runtime-config-adapter-registry's
  two markers on the grounds that ESLint stayed silent without them. An
  adversarial review caught that this was wrong, and it is restored here. The
  file still genuinely source-greps bin/install.js at two sites; ESLint is
  silent only because no-source-grep's TEXT_METHODS omits matchAll. Dropping a
  marker because the linter cannot see the violation is exploiting the blind
  spot, not resolving it -- and it would go red the moment the rule is
  widened. Those two assertions are also irreducible: they assert that EVERY
  inline `runtime === '...'` branch in install.js names a registry-known
  runtime, and a branch naming an unregistered runtime would simply never
  execute, so no runtime observation can prove its absence. The markers now
  say so explicitly.

Two coverage gaps in no-source-grep surfaced while doing this, recorded on
#3464 rather than fixed here, since widening the rule is its own change with
its own blast radius:

- TEXT_METHODS omits matchAll, so a matchAll source-grep never trips the rule
  (the case above).
- The path test requires a literal quoted bin/lib/gsd-core/src segment and
  tracks the binding one hop, so a read through a dynamic path or an
  intermediate variable evades it. install-minimal-hooks' genuinely
  load-bearing read at line 975 is itself unmarked for a related reason.

install-minimal-hooks therefore keeps its markers too: its remaining real
source-grep of bin/install.js (the Codex legacy gsd-update-check migration
check, line 975) is outside this issue's 8 sites. It is the last blocker for
that file and is a clean follow-up.

Closes #3466

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:57:25 -04:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%