sim 3925839f2a test(#4652): failing-first coverage for the four unconfined boundaries
Phase 2 of epic #4636, absorbing #4327 and #4354. Tests only; no fix. These
MUST fail.

Four CLI boundaries join externally-supplied input to a managed root with no
containment validation. Each was driven through the real CLI and confirmed
unconfined before the assertions were written:

  todo complete <name>                     src/commands.cts cmdTodoComplete
  check predicate --phase-dir <dir>        check-command-router cmdCheckPredicate
  check decision-coverage-plan <dir>       check-command-router resolvePath
  check gap-analysis.plan-post <dir>       check-command-router

Boundary 1 is worse than the issue describes. #4327 reports that a traversal
name "resolves outside the todos root", which reads as an information leak.
Measured, it is destructive: `todo complete ../../../../b1out/leak.md` exited
0, MOVED the outside file into completed/, and unlinked the original. The file
was gone. cmdTodoComplete ends in fs.unlinkSync(sourcePath), so an unconfined
name does not merely read across the boundary, it consumes across it.

Boundary 2 reproduces #4354 exactly: a BLOCKING gate returned
{"block":false,"details":{"match":true}} sourced entirely from a SECURITY.md
in a directory the caller chose, outside the project.

Boundaries 3 and 4 are not named in the epic. Both accepted an outside phase
dir and exited 0.

Rows that exist because they are the ones nobody enumerates:

- ORDERING. A real file is created outside the todos root, then the traversal
  name targeting it is asserted rejected AND the outside file asserted still
  present and unmoved. #4327 notes the existence check and the move target
  BOTH follow the unvalidated join, so a rejection that lands after the read
  has already leaked — and, per the finding above, after the unlink has
  already destroyed.
- `a/../../b.md` — looks balanced, resolves outside.
- --dry-run must reject too; a preview must not leak a resolved outside path.
- ${PHASE_DIR} interpolation into a command-exit-zero predicate is the SECOND
  predicate kind, which a fix inside gate-predicate-evaluator.cts would miss.
- An absolute path INSIDE the project must still be accepted at every
  boundary — absolute is not a synonym for escaping.

Cross-boundary rows loop over one shared list of escaping inputs and assert
all four reject with the same shape, so four sites adopting one predicate
cannot drift into four rejection contracts.

Property tests cover BOTH directions — outside is always rejected, inside is
always accepted. A property asserting only rejection is satisfied by a
predicate that rejects everything, which is the degenerate-implementation trap
found in Phase 1's review. Both are seeded.

Regressions fold into the owning module suites rather than a new
tests/fix-NNNN-*.test.cjs, per scripts/lint-regression-test-names.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 09:43:50 -04:00
2026-09-06 02:09:28 +00:00
2026-09-06 02:09:28 +00:00

GSD Core

Git. Ship. Done.

English · Português · 简体中文 · 日本語 · 한국어

A light-weight meta-prompting, context engineering, and spec-driven development system for Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more.

npm version npm downloads Tests Discord GitHub stars License


What is GSD Core

GSD Core is a context-engineering and spec-driven development framework that drives AI coding agents (Claude Code, Codex, Antigravity CLI, Kimi CLI, Copilot, Cursor, and more) through a disciplined phase loop. It solves context rot — the quality degradation that accumulates as an AI fills its context window — by running all heavy research, planning, and execution work in fresh-context subagents while keeping your main session lean.


How it works

Each milestone repeats the same five-step loop, one phase at a time:

  1. Discuss — capture implementation decisions before anything is planned
  2. Plan — research, decompose, and verify the plan fits a fresh context window
  3. Execute — run plans in parallel waves; each executor starts with a clean 200k-token context
  4. Verify — walk through what was built; diagnose and fix before declaring done
  5. Ship — create the PR, archive the phase, repeat for the next one

Quickstart

npx @opengsd/gsd-core@latest

The installer prompts for your runtime (Claude Code, OpenCode, Antigravity CLI, Kimi CLI, Kilo, Codex, Copilot, Cursor, Windsurf, and more) and whether to install globally or locally. The installer is required for cross-runtime compatibility — do not copy files from agents/ or commands/ directly.

On another runtime or without Node.js? See Install on your runtime.

Once installed, start a new project or onboard an existing repo:

/gsd-new-project   # greenfield project
/gsd-onboard       # existing codebase

New here? Follow Your first project for a guided walkthrough from install to first shipped phase, or Onboarding an existing codebase for brownfield setup.


Documentation

What's new in 1.7.0 → docs/whats-new-1.7.0.md

Tutorials — learning by doing:

How-to guides — task-focused recipes:

Reference — authoritative facts:

Explanation — concepts and design decisions:

Full index: docs/README.md. Other languages: 日本語 · 한국어 · Português · 简体中文.


Why it works

Most AI-coding setups fail at scale because context bloat silently degrades output quality, there is no shared memory between sessions, and nothing verifies that code actually works. GSD Core solves all three: heavy work runs in fresh subagents, structured artifacts like STATE.md and CONTEXT.md survive session boundaries, and the verify step walks through what was built and generates fix plans before a phase is declared done. See docs/explanation/context-engineering.md for the full reasoning.

Troubleshooting? See docs/how-to/recover-and-troubleshoot.md.


Community

Project Platform
gsd-opencode Original OpenCode port
Discord Community support

Star History

Star History Chart

License

MIT License. See LICENSE for details.


Claude Code is powerful. GSD Core makes it reliable.

Description
No description provided
Readme MIT 77 MiB
Languages
JavaScript 82.3%
TypeScript 17.4%
Shell 0.3%