Files
msd-core/docs/FEATURES.md
Jakub Zych 6cfa0c55d2 refactor: drop 12 runtimes, keep Claude, Codex, OpenCode, Cursor, ZCode, Antigravity
Removes kilo, kimi, kimi-code, copilot, windsurf, augment, trae, qwen, hermes,
cline, codebuddy and pi end to end: capability descriptors, installer branches
and converters (bin/install.js 14.9k -> 11.2k lines), TypeScript converters,
hook surfaces and runtime homes, review lanes qwen/kimi-code, the two pi
migrations, Kimi payload normalization in the hook guards, dead hostBehaviors
vocabulary, launcher home probes, fixtures, runtime-specific tests and the
prose that presented them as supported.

Installer output for the six kept runtimes is byte-identical to before the
prune. The Kimi tool-vocabulary tests in workflow-guard, read-guard and
read-injection-scanner are left in place pending a decision.
2026-10-06 20:02:40 +02:00

280 KiB
Raw Blame History

MSD Feature Reference

Feature index and reference for MSD Core. For architecture details, see Architecture. For command syntax, see Command Reference. Return to docs index.


Table of Contents


Core Features

1. Project Initialization

Command: /msd-new-project [--auto @file.md]

Purpose: Transform a user's idea into a fully structured project with research, scoped requirements, and a phased roadmap.

Requirements:

  • REQ-INIT-01: System MUST conduct adaptive questioning until project scope is fully understood
  • REQ-INIT-02: System MUST spawn parallel research agents to investigate the domain ecosystem
  • REQ-INIT-03: System MUST extract requirements into v1 (must-have), v2 (future), and out-of-scope categories
  • REQ-INIT-04: System MUST generate a phased roadmap with requirement traceability
  • REQ-INIT-05: System MUST require user approval of the roadmap before proceeding
  • REQ-INIT-06: System MUST prevent re-initialization when .planning/PROJECT.md already exists
  • REQ-INIT-07: System MUST support --auto @file.md flag to skip interactive questions and extract from a document

Produces:

Artifact Description
PROJECT.md Project vision, constraints, technical decisions, evolution rules
REQUIREMENTS.md Scoped requirements with unique IDs (REQ-XX)
ROADMAP.md Phase breakdown with status tracking and requirement mapping
STATE.md Initial project state with position, decisions, metrics
config.json Workflow configuration
research/SUMMARY.md Synthesized domain research
research/STACK.md Technology stack investigation
research/FEATURES.md Feature implementation patterns
research/ARCHITECTURE.md Architecture patterns and trade-offs
research/PITFALLS.md Common failure modes and mitigations

Process:

  1. Questions — Adaptive questioning guided by the "dream extraction" philosophy (not requirements gathering)
  2. Research — 4 parallel researcher agents investigate stack, features, architecture, and pitfalls
  3. Synthesis — Research synthesizer combines findings into SUMMARY.md
  4. Requirements — Extracted from user responses + research, categorized by scope
  5. Roadmap — Phase breakdown mapped to requirements, with granularity setting controlling phase count

Functional Requirements:

  • Questions adapt based on detected project type (web app, CLI, mobile, API, etc.)
  • Research agents have web search capability for current ecosystem information
  • Granularity setting controls phase count: coarse (2-4), standard (4-6), fine (6-10)
  • --auto mode extracts all information from the provided document without interactive questioning
  • Existing codebase context (from /msd-map-codebase) is loaded if present

2. Phase Discussion

Command: /msd-discuss-phase [N] [--auto] [--batch]

Purpose: Capture user's implementation preferences and decisions before research and planning begin. Eliminates the gray areas that cause AI to guess.

Requirements:

  • REQ-DISC-01: System MUST analyze the phase scope and identify decision areas (gray areas)
  • REQ-DISC-02: System MUST categorize gray areas by type (visual, API, content, organization, etc.)
  • REQ-DISC-03: System MUST ask only questions not already answered in prior CONTEXT.md files
  • REQ-DISC-04: System MUST persist decisions in {phase}-CONTEXT.md with canonical references
  • REQ-DISC-05: System MUST support --auto flag to auto-select recommended defaults
  • REQ-DISC-06: System MUST support --batch flag for grouped question intake
  • REQ-DISC-07: System MUST scout relevant source files before identifying gray areas (code-aware discussion)
  • REQ-DISC-08: System MUST adapt gray area language to product-outcome terms when USER-PROFILE.md indicates a non-technical owner (learning_style: guided, jargon in frustration_triggers, or high-level explanation depth)
  • REQ-DISC-09: When REQ-DISC-08 applies, advisor_research rationale paragraphs MUST be rewritten in plain language — same decisions, translated framing

Produces: {padded_phase}-CONTEXT.md — User preferences that feed into research and planning

Gray Area Categories:

Category Example Decisions
Visual features Layout, density, interactions, empty states
APIs/CLIs Response format, flags, error handling, verbosity
Content systems Structure, tone, depth, flow
Organization Grouping criteria, naming, duplicates, exceptions

3. UI Design Contract

Command: /msd-ui-phase [N]

Purpose: Lock design decisions before planning so that all components in a phase share consistent visual standards.

Requirements:

  • REQ-UI-01: System MUST detect existing design system state (shadcn components.json, Tailwind config, tokens)
  • REQ-UI-02: System MUST ask only unanswered design contract questions
  • REQ-UI-03: System MUST validate against 7 dimensions (Copywriting, Visuals, Color, Typography, Spacing, Registry Safety, Inventory Provenance)
  • REQ-UI-04: System MUST enter revision loop if validation returns BLOCKED (max 2 iterations)
  • REQ-UI-05: System MUST offer shadcn initialization for React/Next.js/Vite projects without components.json
  • REQ-UI-06: System MUST enforce registry safety gate for third-party shadcn registries

Produces: {padded_phase}-UI-SPEC.md — Design contract consumed by executors

7 Validation Dimensions:

  1. Copywriting — CTA labels, empty states, error messages
  2. Visuals — Focal points, visual hierarchy, icon accessibility
  3. Color — Accent usage discipline, 60/30/10 compliance
  4. Typography — Font size/weight constraint adherence
  5. Spacing — Grid alignment, token consistency
  6. Registry Safety — Third-party component inspection requirements
  7. Inventory Provenance — Component inventory enumerated from the installed design system, not recalled

shadcn Integration:

  • Detects missing components.json in React/Next.js/Vite projects
  • Guides user through ui.shadcn.com/create preset configuration
  • Preset string becomes a planning artifact reproducible across phases
  • Safety gate requires npx shadcn view and npx shadcn diff before third-party components

4. Phase Planning

Command: /msd-plan-phase [N] [--auto] [--skip-research] [--skip-verify]

Purpose: Research the implementation domain and produce verified, atomic execution plans.

Requirements:

  • REQ-PLAN-01: System MUST spawn a phase researcher to investigate implementation approaches
  • REQ-PLAN-02: System MUST produce plans with 2-3 tasks each, sized for a single context window
  • REQ-PLAN-03: System MUST structure plans as XML with <task> elements containing name, files, action, verify, and done fields
  • REQ-PLAN-04: System MUST include read_first and acceptance_criteria sections in every plan
  • REQ-PLAN-05: System MUST run plan checker verification loop (up to 3 iterations) unless --skip-verify is set
  • REQ-PLAN-06: System MUST support --skip-research flag to bypass research phase
  • REQ-PLAN-07: System MUST prompt user to run /msd-ui-phase if frontend phase detected and no UI-SPEC.md exists (UI safety gate)
  • REQ-PLAN-08: System MUST include Nyquist validation mapping when workflow.nyquist_validation is enabled
  • REQ-PLAN-09: System MUST verify all phase requirements are covered by at least one plan before planning completes (requirements coverage gate)
  • REQ-PLAN-10: System MUST support an optional <reversibility rating="reversible|costly|one-way"> element recording how costly a decision would be to undo, and MUST insert a checkpoint:decision before the task implementing a one-way decision unless --no-reversibility-gates is set (costly is flagged without blocking; reversible and unrated flow normally)

Produces:

Artifact Description
{phase}-RESEARCH.md Ecosystem research findings
{phase}-{N}-PLAN.md Atomic execution plans (2-3 tasks each)
{phase}-VALIDATION.md Test coverage mapping (Nyquist layer)

Plan Structure (XML):

<task type="auto">
  <name>Create login endpoint</name>
  <files>src/app/api/auth/login/route.ts</files>
  <action>
    Use jose for JWT. Validate credentials against users table.
    Return httpOnly cookie on success.
  </action>
  <verify>curl -X POST localhost:3000/api/auth/login returns 200 + Set-Cookie</verify>
  <done>Valid credentials return cookie, invalid return 401</done>
</task>

Plan Checker Verification (8 Dimensions):

  1. Requirement coverage — Plans address all phase requirements
  2. Task atomicity — Each task is independently committable
  3. Dependency ordering — Tasks sequence correctly
  4. File scope — No excessive file overlap between plans
  5. Verification commands — Each task has testable done criteria
  6. Context fit — Tasks fit within a single context window
  7. Gap detection — No missing implementation steps
  8. Nyquist compliance — Tasks have automated verify commands (when enabled)

5. Phase Execution

Command: /msd-execute-phase <N>

Purpose: Execute all plans in a phase using wave-based parallelization with fresh context windows per executor.

Requirements:

  • REQ-EXEC-01: System MUST analyze plan dependencies and group into execution waves
  • REQ-EXEC-02: System MUST spawn independent plans in parallel within each wave
  • REQ-EXEC-03: System MUST give each executor a fresh context window (200K tokens)
  • REQ-EXEC-04: System MUST produce atomic git commits per task
  • REQ-EXEC-05: System MUST produce a SUMMARY.md for each completed plan
  • REQ-EXEC-06: System MUST run post-execution verifier to check phase goals were met
  • REQ-EXEC-07: System MUST support git branching strategies (none, phase, milestone)
  • REQ-EXEC-08: System MUST invoke node repair operator on task verification failure (when enabled)
  • REQ-EXEC-09: System MUST run prior phases' test suites before verification to catch cross-phase regressions

Produces:

Artifact Description
{phase}-{N}-SUMMARY.md Execution outcomes per plan
{phase}-VERIFICATION.md Post-execution verification report
Git commits Atomic commits per task

Wave Execution:

  • Plans with no dependencies → Wave 1 (parallel)
  • Plans depending on Wave 1 → Wave 2 (parallel, waits for Wave 1)
  • Continues until all plans complete
  • File conflicts force sequential execution within same wave

Executor Capabilities:

  • Reads PLAN.md with full task instructions
  • Has access to PROJECT.md, STATE.md, CONTEXT.md, RESEARCH.md
  • Commits each task atomically with structured commit messages
  • Uses --no-verify on commits during parallel execution to avoid build lock contention
  • Handles checkpoint types: auto, checkpoint:human-verify, checkpoint:decision, checkpoint:human-action
  • Reports deviations from plan in SUMMARY.md

Parallel Safety:

  • Pre-commit hooks: Skipped by parallel agents (--no-verify), run once by orchestrator after each wave
  • STATE.md locking: File-level lockfile prevents concurrent write corruption across agents

6. Work Verification

Command: /msd-verify-work [N]

Purpose: User acceptance testing — walk the user through testing each deliverable and auto-diagnose failures.

Requirements:

  • REQ-VERIFY-01: System MUST extract testable deliverables from the phase
  • REQ-VERIFY-02: System MUST present deliverables one at a time for user confirmation
  • REQ-VERIFY-03: System MUST spawn debug agents to diagnose failures automatically
  • REQ-VERIFY-04: System MUST create fix plans for identified issues
  • REQ-VERIFY-05: System MUST inject cold-start smoke test for phases modifying server/database/seed/startup files
  • REQ-VERIFY-06: System MUST produce UAT.md with pass/fail results

Produces: {phase}-UAT.md — User acceptance test results, plus fix plans if issues found


6.5. Ship

Command: /msd-ship [N] [--draft]

Purpose: Bridge local completion → merged PR. After verification passes, push branch, create PR with auto-generated body from planning artifacts, optionally trigger review, and track in STATE.md.

Requirements:

  • REQ-SHIP-01: System MUST verify phase has passed verification before shipping
  • REQ-SHIP-02: System MUST push branch and create PR via gh CLI
  • REQ-SHIP-03: System MUST auto-generate PR body from SUMMARY.md, VERIFICATION.md, and REQUIREMENTS.md
  • REQ-SHIP-04: System MUST update STATE.md with shipping status and PR number
  • REQ-SHIP-05: System MUST support --draft flag for draft PRs
  • REQ-SHIP-06: System MUST support append-only project PR body sections configured with ship.pr_body_sections

Prerequisites: Phase verified, gh CLI installed and authenticated, work on feature branch

Produces: GitHub PR with rich body, optional configured PRD-style sections, STATE.md updated

User documentation: Custom PR Body Sections


7. UI Review

Command: /msd-ui-review [N]

Purpose: Retroactive 6-pillar visual audit of implemented frontend code. Works standalone on any project.

Requirements:

  • REQ-UIREVIEW-01: System MUST score each of the 6 pillars on a 1-4 scale
  • REQ-UIREVIEW-02: System MUST capture screenshots via Playwright CLI to .planning/ui-reviews/
  • REQ-UIREVIEW-03: System MUST create .gitignore for screenshot directory
  • REQ-UIREVIEW-04: System MUST identify top 3 priority fixes
  • REQ-UIREVIEW-05: System MUST work standalone (without UI-SPEC.md) using abstract quality standards

6 Audit Pillars (scored 1-4):

  1. Copywriting — CTA labels, empty states, error states
  2. Visuals — Focal points, visual hierarchy, icon accessibility
  3. Color — Accent usage discipline, 60/30/10 compliance
  4. Typography — Font size/weight constraint adherence
  5. Spacing — Grid alignment, token consistency
  6. Experience Design — Loading/error/empty state coverage

Produces: {padded_phase}-UI-REVIEW.md — Scores and prioritized fixes


8. Milestone Management

Commands: /msd-audit-milestone, /msd-complete-milestone, /msd-new-milestone [name]

Purpose: Verify milestone completion, archive, tag release, and start the next development cycle.

Requirements:

  • REQ-MILE-01: Audit MUST verify all milestone requirements are met
  • REQ-MILE-02: Audit MUST detect stubs, placeholder implementations, and untested code
  • REQ-MILE-03: Audit MUST check Nyquist validation compliance across phases
  • REQ-MILE-04: Complete MUST archive milestone data to MILESTONES.md
  • REQ-MILE-05: Complete MUST offer git tag creation for the release
  • REQ-MILE-06: Complete MUST offer squash merge or merge with history for branching strategies
  • REQ-MILE-07: Complete MUST clean up UI review screenshots
  • REQ-MILE-08: New milestone MUST follow same flow as new-project (questions → research → requirements → roadmap)
  • REQ-MILE-09: New milestone MUST NOT reset existing workflow configuration

Planning Features

9. Phase Management

Commands: /msd-phase, /msd-phase --insert [N], /msd-phase --remove [N]

Purpose: Dynamic roadmap modification during development.

Requirements:

  • REQ-PHASE-01: Add MUST append a new phase to the end of the current roadmap
  • REQ-PHASE-02: Insert MUST use decimal numbering (e.g., 3.1) between existing phases
  • REQ-PHASE-03: Remove MUST renumber all subsequent phases
  • REQ-PHASE-04: Remove MUST prevent removing phases that have been executed
  • REQ-PHASE-05: All operations MUST update ROADMAP.md and create/remove phase directories
  • REQ-PHASE-06: Bare-number phase lookup MUST resolve digit-leading slug names consistently across phase verbs, preserve project-code-prefixed result shaping, and fail loudly when multiple directories match

10. Quick Mode

Command: /msd-quick [--full] [--discuss] [--research]

Purpose: Ad-hoc task execution with MSD guarantees but a faster path.

Requirements:

  • REQ-QUICK-01: System MUST accept freeform task description
  • REQ-QUICK-02: System MUST use same planner + executor agents as full workflow
  • REQ-QUICK-03: System MUST skip research, plan checker, and verifier by default
  • REQ-QUICK-04: --full flag MUST enable plan checking (max 2 iterations) and post-execution verification
  • REQ-QUICK-05: --discuss flag MUST run lightweight pre-planning discussion
  • REQ-QUICK-06: --research flag MUST spawn focused research agent before planning
  • REQ-QUICK-07: Flags MUST be composable (--discuss --research --full)
  • REQ-QUICK-08: System MUST track quick tasks in .planning/quick/YYMMDD-xxx-slug/
  • REQ-QUICK-09: System MUST produce atomic commits for quick task execution

11. Autonomous Mode

Command: /msd-autonomous [--from N]

Purpose: Run all remaining phases autonomously — discuss → plan → execute per phase.

Requirements:

  • REQ-AUTO-01: System MUST iterate through all incomplete phases in roadmap order
  • REQ-AUTO-02: System MUST run discuss → plan → execute for each phase
  • REQ-AUTO-03: System MUST pause for explicit user decisions (gray area acceptance, blockers, validation)
  • REQ-AUTO-04: System MUST re-read ROADMAP.md after each phase to catch dynamically inserted phases
  • REQ-AUTO-05: --from N flag MUST start from a specific phase number

12. Freeform Routing

Command: /msd-progress --do (see also /msd-manager for interactive routing)

Purpose: Analyze freeform text and route to the appropriate MSD command.

Requirements:

  • REQ-DO-01: System MUST parse user intent from natural language input
  • REQ-DO-02: System MUST map intent to the best matching MSD command
  • REQ-DO-03: System MUST confirm the routing with the user before executing
  • REQ-DO-04: System MUST handle project-exists vs no-project contexts differently
  • REQ-DO-05: Routing rules MUST order specific operations before the generic keyword rules they shadow (specific-before-generic)
  • REQ-DO-06: Dispatch MUST forward only arguments the selected command accepts; the freeform sentence is forwarded only when that command explicitly accepts a freeform task description

13. Note Capture

Command: /msd-capture

Purpose: Zero-friction idea capture without interrupting workflow. Append timestamped notes, list all notes, or promote notes to structured todos.

Requirements:

  • REQ-NOTE-01: System MUST save timestamped note files with a single Write call
  • REQ-NOTE-02: System MUST support list subcommand to show all notes from project and global scopes
  • REQ-NOTE-03: System MUST support promote N subcommand to convert a note into a structured todo
  • REQ-NOTE-04: System MUST support --global flag for global scope operations
  • REQ-NOTE-05: System MUST NOT use Task, AskUserQuestion, or Bash — runs inline only

14. Auto-Advance (Next)

Command: /msd-progress --next

Purpose: Automatically detect current project state and advance to the next logical workflow step, eliminating the need to remember which phase/step you're on.

Requirements:

  • REQ-NEXT-01: System MUST read STATE.md, ROADMAP.md, and phase directories to determine current position
  • REQ-NEXT-02: System MUST detect whether discuss, plan, execute, or verify is needed
  • REQ-NEXT-03: System MUST invoke the correct command automatically
  • REQ-NEXT-04: System MUST suggest /msd-new-project if no project exists
  • REQ-NEXT-05: System MUST suggest /msd-complete-milestone when all phases are complete

State Detection Logic:

State Action
No .planning/ directory Suggest /msd-new-project
Phase has no CONTEXT.md Run /msd-discuss-phase
Phase has no PLAN.md files Run /msd-plan-phase
Phase has plans but no SUMMARY.md Run /msd-execute-phase
Phase executed but no VERIFICATION.md Run /msd-verify-work
All phases complete Suggest /msd-complete-milestone

3806. Review Dispositions Ledger

Purpose: Reviews-mode planning (/msd-plan-phase {N} --reviews) has required every current actionable REVIEWS.md finding to be incorporated into PLAN.md or explicitly deferred/rejected there since v1.5.0 (#724/#728). Nothing canonized where in PLAN.md, what shape, or how a REVIEWS.md line reference survives the next round rewriting the file wholesale. Two independently-invented, mutually incompatible disposition formats were observed across two consecutive rounds of the same phase, each written by a different planner subagent instance improvising from prose alone.

Behavior: The existing return-payload tables from references/planner-reviews.md Step 4 — ### Review Feedback Addressed / ### Review Feedback Deferred — are now the canonical Review Dispositions Ledger, promoted verbatim in shape into the affected PLAN.md itself under a ## Review Dispositions Ledger heading. Each reviews-mode round gets its own ### Round {N} — {REVIEWS_sha} subsection, where {REVIEWS_sha} is the commit that wrote that round's REVIEWS.md snapshot (workflows/review.md already commits REVIEWS.md as its own commit). A REVIEWS.md line reference cites L##@{REVIEWS_sha}; a bare line number is non-conforming. The ledger is append-only — a later round adds a new row naming what it supersedes rather than editing or deleting an earlier round's tables.

The contract is stated once, in references/planner-reviews.md; workflows/plan-phase.md's <review_incorporation_contract> and agents/msd-plan-checker.md's Review Incorporation dimension both reference it by name rather than restating it, guarded by a parity test (tests/plan-review-convergence.test.cjs) that fails if the three drift apart.

{Concern}/{Reason} stay free text — the reviewer roster is capability-owned and open to third-party additions, so no closed reviewer/severity enum is introduced.

Known limits: No lint or check verb enforces this shape yet — a follow-up (tracked as part 2 of #3806) will add deterministic enforcement once a migration story for the two pre-existing ad-hoc formats already in the wild is decided. Legacy PLAN.md content written before this convention is not migrated or flagged.

Reference: ADR-3806 · Cross-AI Peer Review


4015. Quick Batch Mode

Command: /msd-quick-batch [--file <path>] [--jobs auto|N] [--validate] [--research] [--resume <batch-id>]

Purpose: Batch several /msd-quick-shaped tasks together — one coordinator plans, dispatches, and merges them as a single run, with per-item leaves and deterministic merge ordering (ADR-1239 "Quick-batch binding").

Requirements:

  • REQ-QB-01: System MUST accept an inline task list (≥2 items) or --file <path>
  • REQ-QB-02: System MUST reject --discuss and --full with a usage error before any dispatch
  • REQ-QB-03: System MUST reject a malformed --jobs value before any dispatch
  • REQ-QB-04: System MUST resolve effective concurrency as min(task count, jobsN, capacity) for --jobs N, or capacity alone for --jobs auto
  • REQ-QB-05: System MUST force a mutating (worktree/executor) wave's concurrency to 1 when isolation is none, without capping a non-mutating (research/planning-only) wave
  • REQ-QB-06: System MUST dispatch a planner per eligible item per DAG layer, providing the full batch task catalog and always requiring depends_on/files_modified frontmatter
  • REQ-QB-07: System MUST recompute execution waves after each planning layer from the planners' declared dependencies/files
  • REQ-QB-08: System MUST serialize worktree create/merge/cleanup while allowing already-created worktrees to run concurrently
  • REQ-QB-09: System MUST merge items strictly in the deterministic wave order, never completion order
  • REQ-QB-10: System MUST NOT call the STATE.md completion primitive for an item routed to human_needed
  • REQ-QB-11: System MUST fail an item routed to gaps_found/merge_failed/scope_violation without rollback, without an automatic retry, and with its worktree preserved
  • REQ-QB-12: System MUST support --resume <batch-id> to re-derive eligibility and dispatch only still-runnable items, refusing closed on an unknown batch id or a diverged base revision

Quality Assurance Features

15. Nyquist Validation

Purpose: Map automated test coverage to phase requirements before any code is written. Named after the Nyquist sampling theorem — ensures a feedback signal exists for every requirement.

Requirements:

  • REQ-NYQ-01: System MUST detect existing test infrastructure during plan-phase research
  • REQ-NYQ-02: System MUST map each requirement to a specific test command
  • REQ-NYQ-03: System MUST identify Wave 0 tasks (test scaffolding needed before implementation)
  • REQ-NYQ-04: Plan checker MUST enforce Nyquist compliance as 8th verification dimension
  • REQ-NYQ-05: System MUST support retroactive validation via /msd-validate-phase
  • REQ-NYQ-06: System MUST be disableable via workflow.nyquist_validation: false

Produces: {phase}-VALIDATION.md — Test coverage contract

Retroactive Validation (/msd-validate-phase [N]):

  • Scans implementation and maps requirements to tests
  • Identifies gaps where requirements lack automated verification
  • Spawns auditor to generate tests (max 3 attempts)
  • Never modifies implementation code — only test files and VALIDATION.md
  • Flags implementation bugs as escalations for user to address

16. Plan Checking

Purpose: Goal-backward verification that plans will achieve phase objectives before execution.

Requirements:

  • REQ-PLANCK-01: System MUST verify plans against 8 quality dimensions
  • REQ-PLANCK-02: System MUST loop up to 3 iterations until plans pass
  • REQ-PLANCK-03: System MUST produce specific, actionable feedback on failures
  • REQ-PLANCK-04: System MUST be disableable via workflow.plan_check: false

17. Post-Execution Verification

Purpose: Automated check that the codebase delivers what the phase promised.

Requirements:

  • REQ-POSTVER-01: System MUST check against phase goals, not just task completion
  • REQ-POSTVER-02: System MUST produce VERIFICATION.md with pass/fail analysis
  • REQ-POSTVER-03: System MUST log issues for /msd-verify-work to address
  • REQ-POSTVER-04: System MUST be disableable via workflow.verifier: false

18. Node Repair

Purpose: Autonomous recovery when task verification fails during execution.

Requirements:

  • REQ-REPAIR-01: System MUST analyze failure and choose one strategy: RETRY, DECOMPOSE, or PRUNE
  • REQ-REPAIR-02: RETRY MUST attempt with a concrete adjustment
  • REQ-REPAIR-03: DECOMPOSE MUST break task into smaller verifiable sub-steps
  • REQ-REPAIR-04: PRUNE MUST remove unachievable tasks and escalate to user
  • REQ-REPAIR-05: System MUST respect repair budget (default: 2 attempts per task)
  • REQ-REPAIR-06: System MUST be configurable via workflow.node_repair_budget and workflow.node_repair

19. Health Validation

Command: /msd-health [--repair] [--backfill]

Purpose: Validate .planning/ directory integrity and auto-repair issues.

Requirements:

  • REQ-HEALTH-01: System MUST check for missing required files
  • REQ-HEALTH-02: System MUST validate configuration consistency
  • REQ-HEALTH-03: System MUST detect orphaned plans without summaries
  • REQ-HEALTH-04: System MUST check phase numbering and roadmap sync
  • REQ-HEALTH-05: --repair flag MUST auto-fix recoverable issues except DESTRUCTIVE-risk ones, which it MUST report but never auto-apply
  • REQ-HEALTH-06: --backfill flag MUST synthesize missing MILESTONES.md entries from archived milestone snapshots

20. Cross-Phase Regression Gate

Purpose: Prevent regressions from compounding across phases by running prior phases' test suites after execution.

Requirements:

  • REQ-REGR-01: System MUST run test suites from all completed prior phases after phase execution
  • REQ-REGR-02: System MUST report any test failures as cross-phase regressions
  • REQ-REGR-03: Regressions MUST be surfaced before post-execution verification
  • REQ-REGR-04: System MUST identify which prior phase's tests were broken

When: Runs automatically during /msd-execute-phase before the verifier step.


21. Requirements Coverage Gate

Purpose: Ensure all phase requirements are covered by at least one plan before planning completes.

Requirements:

  • REQ-COVGATE-01: System MUST extract all requirement IDs assigned to the phase from ROADMAP.md
  • REQ-COVGATE-02: System MUST verify each requirement appears in at least one PLAN.md
  • REQ-COVGATE-03: Uncovered requirements MUST block planning completion
  • REQ-COVGATE-04: System MUST report which specific requirements lack plan coverage

When: Runs automatically at the end of /msd-plan-phase after the plan checker loop.


4273. TDD-Applicability Predicate

Purpose: Give the workflow engine one code-owned computation for whether TDD's RED/GREEN/REFACTOR procedure applies to a given plan, instead of restating the same precedence logic as hand-written prose in each dispatch backend — a restatement that had already drifted between two backends (#4264, #4265). This is Phase 1 of epic #4272 (ADR-3473's fourth application of the single-owner-predicate pattern): it ships the isolated phase.tdd-applicable query verb only. Wiring execute-phase.md and its executor-isolation-dispatch step to consume the verb instead of their own inline predicates is a later phase of the same epic.

Command: msd-tools query phase.tdd-applicable <plan-file> [--cli-flag]

Requirements:

  • REQ-TDDA-01: System MUST resolve applicability via a fixed precedence: --cli-flag (explicit override) > plan frontmatter type: tdd > any task in the plan carrying tdd="true" > project config workflow.tdd_mode
  • REQ-TDDA-02: System MUST report which precedence tier decided the outcome (cli_flag, plan_frontmatter, task_attribute, config, or none) alongside the boolean result
  • REQ-TDDA-03: System MUST emit JSON (applicable, source, plan_type, config_tdd_mode, cli_flag_present) so callers can consume the decision without re-deriving it

Context Engineering Features

22. Context Window Monitoring

Purpose: Prevent context rot by alerting both user and agent when context is running low.

Requirements:

  • REQ-CTX-01: Statusline MUST display context usage percentage to user
  • REQ-CTX-02: Context monitor MUST inject agent-facing warnings at the WARNING fire-point — ≤35% remaining by default, overridable per project via hooks.context_warning_threshold
  • REQ-CTX-03: Context monitor MUST inject agent-facing warnings at the CRITICAL fire-point — ≤25% remaining by default, overridable per project via hooks.context_critical_threshold
  • REQ-CTX-04: Warnings MUST debounce (5 tool uses between repeated warnings)
  • REQ-CTX-05: Severity escalation (WARNING→CRITICAL) MUST bypass debounce
  • REQ-CTX-06: Context monitor MUST differentiate MSD-active vs non-MSD-active projects
  • REQ-CTX-07: Warnings MUST be advisory, never imperative commands that override user preferences
  • REQ-CTX-08: All hooks MUST fail silently and never block tool execution

Architecture: Two-part bridge system:

  1. Statusline writes metrics to /tmp/claude-ctx-{session}.json
  2. Context monitor reads metrics and injects additionalContext warnings

23. Session Management

Commands: /msd-pause-work, /msd-resume-work, /msd-progress

Purpose: Maintain project continuity across context resets and sessions.

Requirements:

  • REQ-SESSION-01: Pause MUST save current position and next steps to continue-here.md and structured HANDOFF.json
  • REQ-SESSION-02: Resume MUST restore full project context from HANDOFF.json (preferred) or state files (fallback)
  • REQ-SESSION-03: Progress MUST show current position, next action, and overall completion
  • REQ-SESSION-04: Progress MUST read all state files (STATE.md, ROADMAP.md, phase directories)
  • REQ-SESSION-05: All session operations MUST work after /clear (context reset)
  • REQ-SESSION-06: HANDOFF.json MUST include blockers, human actions pending, and in-progress task state
  • REQ-SESSION-07: Resume MUST surface human actions and blockers immediately on session start

24. Session Reporting

Command: /msd-pause-work --report

Purpose: Generate a structured post-session summary document capturing work performed, outcomes achieved, and estimated resource usage.

Requirements:

  • REQ-REPORT-01: System MUST gather data from STATE.md, git log, and plan/summary files
  • REQ-REPORT-02: System MUST include commits made, plans executed, and phases progressed
  • REQ-REPORT-03: System MUST estimate token usage and cost based on session activity
  • REQ-REPORT-04: System MUST include active blockers and decisions made
  • REQ-REPORT-05: System MUST recommend next steps

Produces: .planning/reports/SESSION_REPORT.md

Report Sections:

  • Session overview (duration, milestone, phase)
  • Work performed (commits, plans, phases)
  • Outcomes and deliverables
  • Blockers and decisions
  • Resource estimates (tokens, cost)
  • Next steps recommendation

25. Multi-Agent Orchestration

Purpose: Coordinate specialized agents with fresh context windows for each task.

Requirements:

  • REQ-ORCH-01: Each agent MUST receive a fresh context window
  • REQ-ORCH-02: Orchestrators MUST be thin — spawn agents, collect results, route next
  • REQ-ORCH-03: Context payload MUST include all relevant project artifacts
  • REQ-ORCH-04: Parallel agents MUST be truly independent (no shared mutable state)
  • REQ-ORCH-05: Agent results MUST be written to disk before orchestrator processes them
  • REQ-ORCH-06: Failed agents MUST be detected (spot-check actual output vs reported failure)

26. Model Profiles

Command: /msd-config --profile <quality|balanced|budget|adaptive|inherit>

Purpose: Control which AI model each agent uses, balancing quality vs cost.

Requirements:

  • REQ-MODEL-01: System MUST support 4 profiles: quality, balanced, budget, inherit
  • REQ-MODEL-02: Each profile MUST define model tier per agent (see profile table)
  • REQ-MODEL-03: Per-agent overrides MUST take precedence over profile
  • REQ-MODEL-04: inherit profile MUST defer to runtime's current model selection
  • REQ-MODEL-04a: inherit profile MUST be used when running non-Anthropic providers (OpenRouter, local models) to avoid unexpected API costs
  • REQ-MODEL-05: Profile switch MUST be programmatic (script, not LLM-driven)
  • REQ-MODEL-06: Model resolution MUST happen once per orchestration, not per spawn

Profile Assignments:

Agent quality balanced budget inherit
msd-planner Opus Opus Sonnet Inherit
msd-roadmapper Opus Sonnet Sonnet Inherit
msd-executor Opus Sonnet Sonnet Inherit
msd-phase-researcher Opus Sonnet Haiku Inherit
msd-project-researcher Opus Sonnet Haiku Inherit
msd-research-synthesizer Sonnet Sonnet Haiku Inherit
msd-debugger Opus Sonnet Sonnet Inherit
msd-codebase-mapper Sonnet Haiku Haiku Inherit
msd-verifier Sonnet Sonnet Haiku Inherit
msd-plan-checker Sonnet Sonnet Haiku Inherit
msd-integration-checker Sonnet Sonnet Haiku Inherit
msd-nyquist-auditor Sonnet Sonnet Haiku Inherit

4139. Compact Content Mode

Config: workflow.compact_content: false

Purpose: Per-project opt-in to token-minimized variants of MSD's own shipped prompt content — workflow instructions, planning-artifact templates, and non-Claude agent-persona payloads — so the always-loaded instruction window leaves more of the model's attention on the developer's own code (ADR-4139 Decision 2: finite attention, not per-invocation price, since prompt caching already discounts the latter).

Nothing is compressed at runtime. Compact variants are hand-authored, reviewed files sitting beside their canonical siblings; the config key only chooses which one gets read. With the key off (the default), every covered workflow, template, and agent persona behaves exactly as it did before this feature existed.

Requirements:

  • REQ-COMPACT-01: System MUST default workflow.compact_content to false — off costs nothing and changes no existing behavior
  • REQ-COMPACT-02: Eagerly @-included workflow files MUST keep their host-guaranteed load; compactness on this stream comes from a spine + deferred detail/*.md elaboration, never from converting the @-include itself
  • REQ-COMPACT-03: A missed runtime Read of a deferred elaboration or compact variant MUST degrade to a complete, correct, terser state — never to a state with no instructions
  • REQ-COMPACT-04: No compact variant MAY weaken or remove protected content (guardrails, output-format contracts, few-shot examples, security language, structural headings)
  • REQ-COMPACT-05: An agent with no compact persona variant registered MUST fall back to its canonical persona and disclose the fallback inside the served payload, never fail or serve nothing
  • REQ-COMPACT-06: /msd-new-project MUST ask the question and persist the answer; /msd-settings and /msd-config MUST toggle it on an already-initialized project

Config:

Setting Type Default Description
workflow.compact_content boolean false When true, loads token-minimized instruction/template/agent-persona variants wherever one is registered; falls back to canonical content everywhere else

See also: ADR-4139, CONFIGURATION.md, USER-GUIDE.md


Brownfield Features

27. Codebase Mapping

Command: /msd-map-codebase [area]

Purpose: Analyze an existing codebase before starting a new project or as the mapping handoff from /msd-onboard, so MSD understands what exists.

Requirements:

  • REQ-MAP-01: System MUST spawn parallel mapper agents for each analysis area
  • REQ-MAP-02: System MUST produce structured documents in .planning/codebase/
  • REQ-MAP-03: System MUST detect: tech stack, architecture patterns, coding conventions, concerns
  • REQ-MAP-04: Subsequent /msd-new-project MUST load codebase mapping and focus questions on what's being added
  • REQ-MAP-05: Optional [area] argument MUST scope mapping to a specific area

Produces:

Document Content
STACK.md Languages, frameworks, databases, infrastructure
ARCHITECTURE.md Patterns, layers, data flow, boundaries
CONVENTIONS.md Naming, file organization, code style, testing patterns
CONCERNS.md Technical debt, security issues, performance bottlenecks
STRUCTURE.md Directory layout and file organization
TESTING.md Test infrastructure, coverage, patterns
INTEGRATIONS.md External services, APIs, third-party dependencies

Incremental remap — --paths (#2003): The mapper accepts an optional --paths <p1,p2,...> scope hint. When provided, it restricts exploration to the listed repo-relative prefixes instead of scanning the whole tree. This is the pathway used by the post-execute codebase-drift gate to refresh only the subtrees the phase actually changed. Each produced document carries last_mapped_commit in its YAML frontmatter so drift can be measured against the mapping point, not HEAD.


27b. Existing Codebase Onboarding

Command: /msd-onboard [--fast] [--text]

Purpose: Guide first-time setup for an existing repository by checking brownfield state, routing through codebase mapping and docs ingest, then handing off to project initialization without silently overwriting planning artifacts.

Requirements:

  • REQ-ONBOARD-01: System MUST detect existing code, package manifests, planning documents, partial .planning/ state, and complete or missing codebase-map files.
  • REQ-ONBOARD-02: System MUST hand off to /msd-map-codebase or /msd-map-codebase --fast when brownfield code lacks the required .planning/codebase/ map files; fast-map readiness is partial and MUST NOT be treated as sufficient for /msd-new-project.
  • REQ-ONBOARD-03: System MUST offer /msd-ingest-docs before /msd-new-project when ADR/PRD/SPEC/RFC candidates exist and no project exists.
  • REQ-ONBOARD-04: System MUST refuse to report onboarding complete until PROJECT.md, REQUIREMENTS.md, ROADMAP.md, and STATE.md all exist.
  • REQ-ONBOARD-05: System MUST create or confirm .planning/onboarding/SUMMARY.md only after project setup exists.
  • REQ-ONBOARD-06: System MUST support --text for numbered plain-text gates on runtimes without interactive menus.

Produces:

Artifact Description
.planning/codebase/ Codebase map produced by the /msd-map-codebase handoff
.planning/PROJECT.md, REQUIREMENTS.md, ROADMAP.md, STATE.md Planning setup produced by /msd-new-project or /msd-ingest-docs
.planning/onboarding/SUMMARY.md Onboarding status, artifact index, and next-command summary

27a. Post-Execute Codebase Drift Detection

Introduced by: #2003 Trigger: Runs automatically at the end of every /msd-execute-phase Configuration:

  • workflow.drift_threshold (integer, default 3) — minimum new structural elements before the gate acts.
  • workflow.drift_action (warn | auto-remap, default warn) — warn-only or spawn msd-codebase-mapper with --paths scoped to affected subtrees.

What counts as drift:

  • New directory outside mapped paths
  • New barrel export at (packages|apps)/*/src/index.*
  • New migration file (supabase/prisma/drizzle/src/migrations/…)
  • New route module under routes/ or api/

Non-blocking guarantee: any internal failure (missing STRUCTURE.md, git errors, mapper spawn failure) logs a single line and the phase continues. Drift detection cannot fail verification.

Requirements:

  • REQ-DRIFT-01: System MUST detect the four drift categories from git diff --name-status last_mapped_commit..HEAD
  • REQ-DRIFT-02: Action fires only when element count ≥ workflow.drift_threshold
  • REQ-DRIFT-03: warn action MUST NOT spawn any agent
  • REQ-DRIFT-04: auto-remap action MUST pass sanitized --paths to the mapper
  • REQ-DRIFT-05: Detection/remap failure MUST be non-blocking for /msd-execute-phase
  • REQ-DRIFT-06: last_mapped_commit round-trip through YAML frontmatter on each .planning/codebase/*.md file

Utility Features

28. Debug System

Command: /msd-debug [description]

Purpose: Systematic debugging with persistent state across context resets.

Requirements:

  • REQ-DEBUG-01: System MUST create debug session file in .planning/debug/
  • REQ-DEBUG-02: System MUST track hypotheses, evidence, and eliminated theories
  • REQ-DEBUG-03: System MUST persist state so debugging survives context resets
  • REQ-DEBUG-04: System MUST require human verification before marking resolved
  • REQ-DEBUG-05: Resolved sessions MUST append to .planning/debug/knowledge-base.md
  • REQ-DEBUG-06: Knowledge base MUST be consulted on new debug sessions to prevent re-investigation

Debug Session States: gathering → investigating → fixing → verifying → awaiting_human_verify → resolved


29. Todo Management

Commands: /msd-capture [desc], /msd-capture --list

Purpose: Capture ideas and tasks during sessions for later work.

Requirements:

  • REQ-TODO-01: System MUST capture todo from current conversation context
  • REQ-TODO-02: Todos MUST be stored in .planning/todos/pending/
  • REQ-TODO-03: Completed todos MUST move to .planning/todos/completed/
  • REQ-TODO-04: Check-todos MUST list all pending items with selection to work on one

30. Statistics Dashboard

Command: /msd-stats

Purpose: Display project metrics — phases, plans, requirements, git history, and timeline.

Requirements:

  • REQ-STATS-01: System MUST show phase/plan completion counts
  • REQ-STATS-02: System MUST show requirement coverage
  • REQ-STATS-03: System MUST show git commit metrics
  • REQ-STATS-04: System MUST support multiple output formats (json, table, bar)

31. Update System

Command: /msd-update

Purpose: Update MSD to the latest version with changelog preview.

Requirements:

  • REQ-UPDATE-01: System MUST check for new versions via npm
  • REQ-UPDATE-02: System MUST display changelog for new version before updating
  • REQ-UPDATE-03: System MUST be runtime-aware and target the correct directory
  • REQ-UPDATE-04: System MUST back up locally modified files to msd-local-patches/
  • REQ-UPDATE-05: /msd-update --reapply MUST restore local modifications after update
  • REQ-UPDATE-06: /msd-update --next (alias --rc) MUST target the @next RC dist-tag for version check and install; omitting the flag MUST keep @latest behavior unchanged (ADR #660)
  • REQ-UPDATE-07: System MUST back up user-added files found inside MSD-managed directories to msd-user-files-backup/ before the clean install
  • REQ-UPDATE-08: When that backup is non-empty, the update MUST offer an explicit restore choice before finishing, and MUST leave the backup intact whichever way the user answers
  • REQ-UPDATE-09: A restore MUST NOT overwrite a path the newly installed release ships, MUST NOT overwrite a different file already on disk, and MUST report best-effort compatibility warnings for restored files without blocking on them

32. Settings Management

Command: /msd-settings

Purpose: Interactive configuration of workflow toggles and model profile.

Requirements:

  • REQ-SETTINGS-01: System MUST present current settings with toggle options
  • REQ-SETTINGS-02: System MUST update .planning/config.json
  • REQ-SETTINGS-03: System MUST support saving as global defaults (~/.msd/defaults.json)

Configurable Settings:

Setting Type Default Description
mode enum interactive interactive or yolo (auto-approve)
granularity enum standard coarse, standard, or fine
model_profile enum balanced quality, balanced, budget, or inherit
models.<phase_type> enum (none) Per-phase-type tier override (planning, discuss, research, execution, verification, completion). Values: opus, sonnet, haiku, inherit. Coarse phase-level tuning that wins over model_profile but loses to per-agent model_overrides. See CONFIGURATION.md. Added in v1.40
granularities.<phase_type> enum (none) Per-phase-type granularity override (planning, discuss, research, execution, verification, completion). Values: coarse, standard, fine. Mirrors models.<phase_type> for granularity. See CONFIGURATION.md. Added in v1.43 (#68). /msd-plan-phase --granularity <coarse|standard|fine> overrides all config-based granularity for a single invocation (takes precedence over granularities.planning, top-level granularity, and planning.granularity). (#703)
dynamic_routing.enabled boolean false Master switch for failure-tier escalation. When true, agents resolve to tier_models[default_tier] and escalate one tier on orchestrator-detected soft failure. Capped by max_escalations. See CONFIGURATION.md. Added in v1.40
workflow.research boolean true Domain research before planning
workflow.plan_check boolean true Plan verification loop
workflow.verifier boolean true Post-execution verification
workflow.auto_advance boolean false Auto-chain discuss→plan→execute
workflow.nyquist_validation boolean true Nyquist test coverage mapping
workflow.ui_phase boolean true UI design contract generation
workflow.ui_safety_gate boolean true Prompt for ui-phase on frontend phases
workflow.node_repair boolean true Autonomous task repair
workflow.node_repair_budget number 2 Max repair attempts per task
planning.commit_docs boolean true Commit .planning/ files to git
planning.search_gitignored boolean false Include gitignored files in searches
parallelization.enabled boolean true Run independent plans simultaneously
git.branching_strategy enum none none, phase, or milestone

33. Test Generation

Command: /msd-add-tests [N]

Purpose: Generate tests for a completed phase based on UAT criteria and implementation.

Requirements:

  • REQ-TEST-01: System MUST analyze completed phase implementation
  • REQ-TEST-02: System MUST generate tests based on UAT criteria and acceptance criteria
  • REQ-TEST-03: System MUST use existing test infrastructure patterns

Infrastructure Features

Looking for a third-party add-on instead? See the MSD Community Capability Registry & EoS Registry — non-endorsing discoverability catalogs for community-contributed Capabilities and EoS host integrations.

34. Git Integration

Purpose: Atomic commits, branching strategies, and clean history management.

Requirements:

  • REQ-GIT-01: Each task MUST get its own atomic commit
  • REQ-GIT-02: Commit messages MUST follow structured format: type(scope): description
  • REQ-GIT-03: System MUST support 3 branching strategies: none, phase, milestone
  • REQ-GIT-04: Phase strategy MUST create one branch per phase
  • REQ-GIT-05: Milestone strategy MUST create one branch per milestone
  • REQ-GIT-06: Complete-milestone MUST offer squash merge (recommended) or merge with history
  • REQ-GIT-07: System MUST respect commit_docs setting for .planning/ files
  • REQ-GIT-08: System MUST auto-detect .planning/ in .gitignore and skip commits

Commit Format:

type(phase-plan): description

# Examples:
docs(08-02): complete user registration plan
feat(08-02): add email confirmation flow
fix(03-01): correct auth token expiry

35. CLI Tools

Purpose: Programmatic utilities for workflows and agents, replacing repetitive inline bash patterns.

Requirements:

  • REQ-CLI-01: System MUST provide atomic commands for state, config, phase, roadmap operations
  • REQ-CLI-02: System MUST provide compound init commands that load all context for each workflow
  • REQ-CLI-03: System MUST support --raw flag for machine-readable output
  • REQ-CLI-04: System MUST support --cwd flag for sandboxed subagent operation
  • REQ-CLI-05: All operations MUST use forward-slash paths on Windows

Command Categories: State (11 subcommands), Phase (5), Roadmap (3), Verify (8), Template (2), Frontmatter (4), Scaffold (4), Init (12), Validate (2), Progress, Stats, Todo


36. Multi-Runtime Support

Purpose: Run MSD across multiple AI coding agent runtimes.

Requirements:

  • REQ-RUNTIME-01: System MUST support Claude Code, OpenCode, Codex, Antigravity, Cursor
  • REQ-RUNTIME-02: Installer MUST transform content per runtime (tool names, paths, frontmatter)
  • REQ-RUNTIME-03: Installer MUST support interactive and non-interactive (--claude --global) modes
  • REQ-RUNTIME-04: Installer MUST support both global and local installation
  • REQ-RUNTIME-05: Uninstall MUST cleanly remove all MSD files without affecting other configurations
  • REQ-RUNTIME-06: Installer MUST handle platform differences (Windows, macOS, Linux, WSL, Docker)
  • REQ-RUNTIME-07: Runtimes with lifecycle hook support MUST register per-turn context-headroom tracking events at install time
  • REQ-RUNTIME-08: Native packaging manifests MUST be version-stamped and enable runtime-native install/update/uninstall flows

Runtime Transformations:

Aspect Claude Code OpenCode Codex Antigravity Cursor
Commands Slash commands Slash commands Skills (TOML) Skills Skills + Slash commands
Agent format Claude native mode: subagent Skills Skills Skills
Skills emission N/A On-demand SKILL.md (1.4.0) /skills picker (1.4.0) N/A SKILL.md
Hook events SessionStart, PreToolUse, PostToolUse, SubagentStop, Stop, PreCompact, FileChanged N/A SessionStart, SubagentStart, Stop, PostToolUse N/A sessionStart, postToolUse
Config settings.json opencode.json(c) TOML Config Config

Cursor artifact surfaces: msd install --cursor writes two artifact kinds:

  • ~/.cursor/skills/msd-<name>/SKILL.md — rich skills with YAML frontmatter, Cursor tool-name mapping, and adapter context header (existing surface)
  • ~/.cursor/commands/msd-<name>.md — plain markdown slash commands (no frontmatter) invocable via / in the Agent input (Cursor 1.6+)

Native skills emission (1.4.0): One runtime now emits MSD as on-demand native skills (skills/<name>/SKILL.md) at install time, in addition to their existing command and agent surfaces. Skills respect the active install profile and are removed on uninstall.

  • OpenCode — emits skills alongside its existing surfaces

New slash-command surfaces (1.4.0):

  • Cursor (Cursor >= 1.6) — .cursor/commands/msd-<name>.md so MSD appears in the / command menu

Cross-runtime lifecycle hooks (1.4.0): Each supported runtime registers lifecycle hook events for per-turn context-headroom tracking and workflow state management. Notable registrations:

  • Claude Code: SubagentStop, Stop, PreCompact (context-headroom warnings), FileChanged (hot-reloads .planning/config.json mid-session)
  • Codex: SubagentStart, Stop, PostToolUse (new in 1.4.0); on Windows the SessionStart hook entry gains a commandWindows field so the .cmd shim is used for native execution
  • Cursor: sessionStart (injects workflow state), postToolUse (nudges .planning updates)

Runtime-specific enrichments (1.4.0):

  • Codex emits service_tier: flex for light-tier agents; MSD skills appear in the Codex /skills picker via SKILL.md (no agents/openai.yaml sidecar is emitted — doing so caused duplicate autocomplete entries, #1326)

Native packaging:

  • Claude Code: MSD Core ships a .claude-plugin/plugin.json manifest, enabling installation and lifecycle management via claude plugin install|enable|disable|update msd-core. Commands load under the /msd-core: namespace (e.g. /msd-core:plan-phase), avoiding slash-command collisions with the classic npm installer which uses /msd:. Always-on guard and update hooks are wired automatically via hooks/hooks.json. The plugin path is additive — the npm installer (npx @golem15/msd-core) remains fully supported.

37. Hook System

Purpose: Runtime event hooks for context monitoring, status display, and update checking.

Requirements:

  • REQ-HOOK-01: Statusline MUST display model, current task, directory, and context usage
  • REQ-HOOK-02: Context monitor MUST inject agent-facing warnings at threshold levels
  • REQ-HOOK-03: Update checker MUST run in background on session start
  • REQ-HOOK-04: All hooks MUST respect CLAUDE_CONFIG_DIR env var
  • REQ-HOOK-05: All hooks MUST include 3-second stdin timeout guard
  • REQ-HOOK-06: All hooks MUST fail silently on any error
  • REQ-HOOK-07: Context usage MUST normalize for autocompact buffer (16.5% reserved)
  • REQ-HOOK-08: Update banner MUST be opt-in and silent unless an update is available (PR #2795)

Statusline Display:

[⬆ /msd-update │] model │ [current task │] directory [█████░░░░░ 50%]

Color coding: <50% green, <65% yellow, <80% orange, ≥80% red with skull emoji

Update Banner (opt-in, when MSD statusline isn't used):

When the user declines (or keeps a non-MSD) statusline, the installer offers a SessionStart banner that surfaces update availability without occupying statusline real estate. The banner reads ~/.cache/msd/msd-update-check.json (written by msd-check-update-worker.js) and emits one line only when an update is available:

MSD update available: 1.39.0 → 1.40.0. Run /msd-update.

The banner is silent when up-to-date and rate-limits "check failed" diagnostics to once per 24 hours. Removed cleanly by npx @golem15/msd-core --uninstall or by deleting the SessionStart entry that references msd-update-banner.js.


38. Developer Profiling

Command: /msd-profile-user [--questionnaire] [--refresh]

Purpose: Analyze Claude Code session history to build behavioral profiles across 8 dimensions, generating artifacts that personalize Claude's responses to the developer's style.

Dimensions:

  1. Communication style (terse vs verbose, formal vs casual)
  2. Decision patterns (rapid vs deliberate, risk tolerance)
  3. Debugging approach (systematic vs intuitive, log preference)
  4. UX preferences (design sensibility, accessibility awareness)
  5. Vendor/technology choices (framework preferences, ecosystem familiarity)
  6. Frustration triggers (what causes friction in workflows)
  7. Learning style (documentation vs examples, depth preference)
  8. Explanation depth (high-level vs implementation detail)

Generated Artifacts:

  • USER-PROFILE.md — Full behavioral profile with evidence citations
  • CLAUDE.md profile section — Auto-discovered by Claude Code

Flags:

  • --questionnaire — Interactive questionnaire fallback when session history is unavailable
  • --refresh — Re-analyze sessions and regenerate profile

Pipeline Modules:

  • profile-pipeline.cjs — Session scanning, message extraction, sampling
  • profile-output.cjs — Profile rendering, questionnaire, artifact generation
  • msd-user-profiler agent — Behavioral analysis from session data

Requirements:

  • REQ-PROF-01: Session analysis MUST cover at least 8 behavioral dimensions
  • REQ-PROF-02: Profile MUST cite evidence from actual session messages
  • REQ-PROF-03: Questionnaire MUST be available as fallback when no session history exists
  • REQ-PROF-04: Generated artifacts MUST be discoverable by Claude Code (CLAUDE.md integration)

39. Execution Hardening

Purpose: Three additive quality improvements to the execution pipeline that catch cross-plan failures before they cascade.

Components:

1. Pre-Wave Dependency Check (execute-phase) Before spawning wave N+1, verify key-links from prior wave artifacts exist and are wired correctly. Catches cross-plan dependency gaps before they cascade into downstream failures.

2. Cross-Plan Data Contracts — Dimension 9 (plan-checker) New analysis dimension that checks plans sharing data pipelines have compatible transformations. Flags when one plan strips data that another plan needs in its original form.

3. Export-Level Spot Check (verify-phase) After Level 3 wiring verification passes, spot-check individual exports for actual usage. Catches dead stores that exist in wired files but are never called.

Requirements:

  • REQ-HARD-01: Pre-wave check MUST verify key-links from all prior wave artifacts before spawning next wave
  • REQ-HARD-02: Cross-plan contract check MUST detect incompatible data transformations between plans
  • REQ-HARD-03: Export spot-check MUST identify dead stores in wired files

40. Verification Debt Tracking

Command: /msd-audit-uat

Purpose: Prevent silent loss of UAT/verification items when projects advance past phases with outstanding tests. Surfaces verification debt across all prior phases so items are never forgotten.

Components:

1. Cross-Phase Health Check (progress.md Step 1.6) Every /msd-progress call scans ALL phases in the current milestone for outstanding items (pending, skipped, blocked, human_needed, gaps_found). Displays a non-blocking warning section with actionable links.

A verification report counts as outstanding under EITHER terminal non-passing status: human_needed contributes its human_verification: entries, and gaps_found contributes both its human_verification: and its gaps: entries, excluding any already closed. What counts as closed is per key: a gaps: entry closes on status: resolved and nothing else — the same rule the ## Gaps markdown reader applies, so one authored entry cannot read closed in one reader and open in the other — while a human_verification: entry also closes on a bare resolution: field, provided no status: contradicts it (#3850).

2. status: partial (verify-work.md, UAT.md) New UAT status that distinguishes between "session ended" and "all tests resolved". Prevents status: complete when tests are still pending, blocked, or skipped without reason.

3. result: blocked with blocked_by tag (verify-work.md, UAT.md) New test result type for tests blocked by external dependencies (server, physical device, release build, third-party services). Categorized separately from skipped tests.

4. HUMAN-UAT.md Persistence (execute-phase.md) When verification returns human_needed, items are persisted as a trackable HUMAN-UAT.md file with status: partial. Feeds into the cross-phase health check and audit systems.

5. Phase Completion Warnings (phase.cjs, transition.md) phase complete CLI returns verification debt warnings in its JSON output. Transition workflow surfaces outstanding items before confirmation.

Requirements:

  • REQ-DEBT-01: System MUST surface outstanding UAT/verification items from ALL prior phases in /msd-progress
  • REQ-DEBT-02: System MUST distinguish incomplete testing (partial) from completed testing (complete)
  • REQ-DEBT-03: System MUST categorize blocked tests with blocked_by tags
  • REQ-DEBT-04: System MUST persist human_needed verification items as trackable UAT files
  • REQ-DEBT-05: System MUST warn (non-blocking) during phase completion and transition when verification debt exists
  • REQ-DEBT-06: /msd-audit-uat MUST scan all phases, categorize items by testability, and produce a human test plan

v1.27 Features

41. Fast Mode

Command: /msd-fast [task description]

Purpose: Execute trivial tasks inline without spawning subagents or generating PLAN.md files. For tasks too small to justify planning overhead: typo fixes, config changes, small refactors, forgotten commits, simple additions.

Requirements:

  • REQ-FAST-01: System MUST execute the task directly in the current context without subagents
  • REQ-FAST-02: System MUST produce an atomic git commit for the change
  • REQ-FAST-03: System MUST track the task in .planning/quick/ for state consistency
  • REQ-FAST-04: System MUST NOT be used for tasks requiring research, multi-step planning, or verification

When to use vs /msd-quick:

  • /msd-fast — One-sentence tasks executable in under 2 minutes (typo, config change, small addition)
  • /msd-quick — Anything needing research, multi-step planning, or verification

42. Cross-AI Peer Review

Command: /msd-review --phase N [--claude] [--codex] [--coderabbit] [--opencode] [--cursor] [--agy] [--antigravity] [--ollama] [--lm-studio] [--llama-cpp] [--all]

Purpose: Invoke external AI CLIs (Claude, Codex, CodeRabbit, OpenCode, Cursor, Antigravity) and local OpenAI-compatible servers (Ollama, LM Studio, llama.cpp) to independently review phase plans. Produces structured REVIEWS.md with per-reviewer feedback.

Each reviewer is a declared lane: its binary, prompt and output channels, timeout, availability probe, and empty-output policy come from a capability manifest rather than hand-written per-CLI logic, so a reviewer can be shipped as an installable capability instead of a core change.

Requirements:

  • REQ-REVIEW-01: System MUST detect available AI CLIs on the system
  • REQ-REVIEW-02: System MUST build a structured review prompt from phase plans
  • REQ-REVIEW-03: System MUST invoke each selected CLI independently
  • REQ-REVIEW-04: System MUST collect responses and produce REVIEWS.md
  • REQ-REVIEW-05: Reviews MUST be consumable by /msd-plan-phase --reviews
  • REQ-REVIEW-06: System MUST support project-level no-flag defaults via review.default_reviewers
  • REQ-REVIEW-07: Reviewer precedence MUST be explicit flags > --all > review.default_reviewers > all detected reviewers

Produces: {phase}-REVIEWS.md — Per-reviewer structured feedback

User configuration note:

  • Set review.default_reviewers in .planning/config.json (or via msd config-set) to control no-flag /msd-review fan-out.
  • review.default_reviewers may include configured review.reviewer_instances names; each instance runs as an independent reviewer identity backed by its configured adapter/model. Instance names are not CLI flags.
  • Use --all for a full pre-merge sweep without changing project defaults.
  • For local model servers with small context windows, set review.max_prompt_tokens_per_reviewer to auto-trim prompts per reviewer — see Prompt budgets for small-context reviewers in CONFIGURATION.md.

Why record which model produced a review (#2295): reviewers: in the frontmatter recorded which CLIs ran, but not which model each one resolved to. Without a pin, the model is whatever the CLI's own config or internal default happens to pick, so a "Codex vs Antigravity" comparison could quietly be a frontier model against a cheap-tier default with nothing in the record to say so — and a CLI update, or an unrelated config edit, could silently make past and future reviews incomparable.

The fix records the model and its provenance. Provenance is what makes the value trustworthy: pinned (from review.models.<slug>) is certain, while banner and transcript are recovered from third-party CLI output this project does not own — a startup banner or an undocumented session log.

That third-party dependence is a real trade-off, held honestly rather than papered over: the banner and transcript arms read formats MSD does not control, so they are best-effort by design and degrade to unknown rather than guessing or failing the run. A recorded unknown is a real answer — a wrong model name attributed to a review would be worse than none.


43. Backlog Parking Lot

Commands: /msd-capture --backlog <description>, /msd-review-backlog, /msd-capture --seed <idea>, /msd-capture --list-seeds [status]

Purpose: Capture ideas that aren't ready for active planning. Backlog items use 999.x numbering to stay outside the active phase sequence. Seeds are forward-looking ideas with trigger conditions that surface automatically at the right milestone. --list-seeds provides a read-only audit of all parked seeds (with optional status filter) without waiting for the next milestone.

Requirements:

  • REQ-BACKLOG-01: Backlog items MUST use 999.x numbering to stay outside active phase sequence
  • REQ-BACKLOG-02: Phase directories MUST be created immediately so /msd-discuss-phase and /msd-plan-phase work on them
  • REQ-BACKLOG-03: /msd-review-backlog MUST support promote, keep, and remove actions per item
  • REQ-BACKLOG-04: Promoted items MUST be renumbered into the active milestone sequence
  • REQ-SEED-01: Seeds MUST capture the full WHY and WHEN to surface conditions
  • REQ-SEED-02: /msd-new-milestone MUST scan seeds and present matches
  • REQ-SEED-03: /msd-capture --list-seeds MUST list seeds with status, scope, and trigger for audit, with optional status filtering

Produces:

Artifact Description
.planning/phases/999.x-slug/ Backlog item directory
.planning/seeds/SEED-YYMMDD-xxx-slug.md Seed with trigger conditions

44. Persistent Context Threads

Command: /msd-thread [name | description]

Purpose: Lightweight cross-session knowledge stores for work that spans multiple sessions but doesn't belong to any specific phase. Lighter weight than /msd-pause-work — no phase state, no plan context.

Requirements:

  • REQ-THREAD-01: System MUST support create, list, and resume modes
  • REQ-THREAD-02: Threads MUST be stored in .planning/threads/ as markdown files
  • REQ-THREAD-03: Thread files MUST include Goal, Context, References, and Next Steps sections
  • REQ-THREAD-04: Resuming a thread MUST load its full context into the current session
  • REQ-THREAD-05: Threads MUST be promotable to phases or backlog items

Produces: .planning/threads/{slug}.md — Persistent context thread


45. PR Branch Filtering

Command: /msd-pr-branch [target branch]

Purpose: Create a clean branch suitable for pull requests by filtering out .planning/ commits. Reviewers see only code changes, not MSD planning artifacts.

Requirements:

  • REQ-PRBRANCH-01: System MUST identify commits that only modify .planning/ files
  • REQ-PRBRANCH-02: System MUST create a new branch with planning commits filtered out
  • REQ-PRBRANCH-03: Code changes MUST be preserved exactly as committed
  • REQ-PRBRANCH-04: System MUST NOT delete a .planning/ path the target branch already tracks
  • REQ-PRBRANCH-05: Verification MUST assert against the active filter mode's contract, not an unconditional zero

Filter modes. planning.pr_strict selects what "filtered" means. The default mode treats .planning/ as two populations: structural state that belongs in review (STATE.md, ROADMAP.md, MILESTONES.md, PROJECT.md, REQUIREMENTS.md, milestones/**) and transient per-phase artifacts that do not (phases/, quick/, research/, threads/, todos/, debug/, seeds/, codebase/, ui-reviews/). Strict mode collapses that distinction: nothing under .planning/ reaches the PR branch, and a commit is carried over only when it touches at least one file outside .planning/.

Strict mode exists because the two ways to keep planning private are not equivalent. Turning off planning.commit_docs keeps .planning/ out of git, which also takes parallel executor worktrees with it — a worktree is checked out from a commit, so an untracked planning tree is simply absent inside it and the executor has no PLAN.md to read. Strict mode leaves planning committed, so worktrees and revert paths keep working, and moves the guarantee to the publication boundary instead. See Publish PRs without planning artifacts.

Both modes filter by forcing the excluded paths back to whatever the target branch already tracks, in the index and the working tree. Un-staging alone would record a deletion of any planning file the target branch carries, and would leave the picked file untracked on disk, where it makes a later cherry-pick of the same path abort.


46. Security Hardening

Purpose: Defense-in-depth security for MSD's planning artifacts. Because MSD generates markdown files that become LLM system prompts, user-controlled text flowing into these files is a potential indirect prompt injection vector.

Components:

1. Centralized Security Module (security.cjs)

  • Path traversal prevention — validates file paths resolve within the project directory
  • Prompt injection detection — scans for known injection patterns in user-supplied text
  • Safe JSON parsing — catches malformed input before state corruption
  • Field name validation — prevents injection through config field names
  • Shell argument validation — sanitizes user text before shell interpolation

2. Prompt Injection Guard Hook (msd-prompt-guard.js) PreToolUse hook that scans Write/Edit calls targeting .planning/ for injection patterns. Advisory-only — logs detection for awareness without blocking legitimate operations.

3. Workflow Guard Hook (msd-workflow-guard.js) PreToolUse hook that detects when Claude attempts file edits outside a MSD workflow context. Advises using /msd-quick or /msd-fast instead of direct edits. Configurable via hooks.workflow_guard (default: false).

4. CI-Ready Injection Scanner (prompt-injection-scan.security.test.cjs) Test suite that scans all agent, workflow, and command files for embedded injection vectors.

Requirements:

  • REQ-SEC-01: All user-supplied file paths MUST be validated against the project directory
  • REQ-SEC-02: Prompt injection patterns MUST be detected before text enters planning artifacts
  • REQ-SEC-03: Security hooks MUST be advisory-only (never block legitimate operations)
  • REQ-SEC-04: JSON parsing of user input MUST catch malformed data gracefully
  • REQ-SEC-05: macOS /var → /private/var symlink resolution MUST be handled in path validation

47. Multi-Repo Workspace Support

Purpose: Auto-detection and project root resolution for monorepos and multi-repo setups. Supports workspaces where .planning/ may need to resolve across repository boundaries.

Requirements:

  • REQ-MULTIREPO-01: System MUST auto-detect multi-repo workspace configuration
  • REQ-MULTIREPO-02: System MUST resolve project root across repository boundaries
  • REQ-MULTIREPO-03: Executor MUST record per-repo commit hashes in multi-repo mode

48. Discussion Audit Trail

Purpose: Auto-generate DISCUSSION-LOG.md during /msd-discuss-phase for full audit trail of decisions made during discussion.

Requirements:

  • REQ-DISCLOG-01: System MUST auto-generate DISCUSSION-LOG.md during discuss-phase
  • REQ-DISCLOG-02: Log MUST capture questions asked, options presented, and decisions made
  • REQ-DISCLOG-03: Decision IDs MUST enable traceability from discuss-phase to plan-phase

v1.28 Features

49. Forensics

Command: /msd-forensics [description]

Purpose: Post-mortem investigation of failed or stuck MSD workflows.

Requirements:

  • REQ-FORENSICS-01: System MUST analyze git history for anomalies (stuck loops, long gaps, repeated commits)
  • REQ-FORENSICS-02: System MUST check artifact integrity (completed phases have expected files)
  • REQ-FORENSICS-03: System MUST generate a markdown report saved to .planning/forensics/
  • REQ-FORENSICS-04: System MUST offer to create a GitHub issue with findings
  • REQ-FORENSICS-05: System MUST NOT modify project files (read-only investigation)

Produces:

Artifact Description
.planning/forensics/report-{timestamp}.md Post-mortem investigation report

Process:

  1. Scan — Analyze git history for anomalies: stuck loops, long gaps between commits, repeated identical commits
  2. Integrity Check — Verify completed phases have expected artifact files
  3. Report — Generate markdown report with findings, saved to .planning/forensics/
  4. Issue — Offer to create a GitHub issue with findings for team visibility

50. Milestone Summary

Command: /msd-milestone-summary [version]

Purpose: Generate comprehensive project summary from milestone artifacts for team onboarding.

Requirements:

  • REQ-SUMMARY-01: System MUST aggregate phase plans, summaries, and verification results
  • REQ-SUMMARY-02: System MUST work for both current and archived milestones
  • REQ-SUMMARY-03: System MUST produce a single navigable document

Produces:

Artifact Description
MILESTONE-SUMMARY.md Comprehensive navigable summary of milestone artifacts

Process:

  1. Collect — Aggregate phase plans, summaries, and verification results from the target milestone
  2. Synthesize — Combine artifacts into a single navigable document with cross-references
  3. Output — Write MILESTONE-SUMMARY.md suitable for team onboarding and stakeholder review

51. Workstream Namespacing

Command: /msd-workstreams

Purpose: Parallel workstreams for concurrent work on different milestone areas.

Requirements:

  • REQ-WS-01: System MUST isolate workstream state in separate .planning/workstreams/{name}/ directories
  • REQ-WS-02: System MUST validate workstream names (alphanumeric + hyphens only, no path traversal)
  • REQ-WS-03: System MUST support list, create, switch, status, progress, complete, resume subcommands

Produces:

Artifact Description
.planning/workstreams/{name}/ Isolated workstream directory structure

Process:

  1. Create — Initialize a named workstream with isolated .planning/workstreams/{name}/ directory
  2. Switch — Change active workstream context for subsequent MSD commands
  3. Manage — List, check status, track progress, complete, or resume workstreams

52. Manager Dashboard

Command: /msd-manager

Purpose: Interactive command center for managing multiple phases from one terminal.

Requirements:

  • REQ-MGR-01: System MUST show overview of all phases with status
  • REQ-MGR-02: System MUST filter to current milestone scope
  • REQ-MGR-03: System MUST show phase dependencies and conflicts

Produces: Interactive terminal output

Process:

  1. Scan — Load all phases in the current milestone with their statuses
  2. Display — Render overview showing phase dependencies, conflicts, and progress
  3. Interact — Accept commands to navigate, inspect, or act on individual phases

53. Assumptions Discussion Mode

Command: /msd-discuss-phase with workflow.discuss_mode: 'assumptions'

Purpose: Replace interview-style questioning with codebase-first assumption analysis.

Requirements:

  • REQ-ASSUME-01: System MUST analyze codebase to generate structured assumptions before asking questions
  • REQ-ASSUME-02: System MUST classify assumptions by confidence level (Confident/Likely/Unclear)
  • REQ-ASSUME-03: System MUST produce identical CONTEXT.md format as default discuss mode
  • REQ-ASSUME-04: System MUST support confidence-based skip gate (all HIGH = no questions)

Produces:

Artifact Description
{phase}-CONTEXT.md Same format as default discuss mode

Process:

  1. Analyze — Scan codebase to generate structured assumptions about implementation approach
  2. Classify — Categorize assumptions by confidence level: Confident, Likely, Unclear
  3. Gate — If all assumptions are HIGH confidence, skip questioning entirely
  4. Confirm — Present unclear assumptions as targeted questions to the user
  5. Output — Produce {phase}-CONTEXT.md in identical format to default discuss mode

54. UI Phase Auto-Detection

Part of: /msd-new-project and /msd-progress

Purpose: Automatically detect UI-heavy projects and surface /msd-ui-phase recommendation.

Requirements:

  • REQ-UI-DETECT-01: System MUST detect UI signals in project description (keywords, framework references)
  • REQ-UI-DETECT-02: System MUST annotate ROADMAP.md phases with ui_hint when applicable
  • REQ-UI-DETECT-03: System MUST suggest /msd-ui-phase in next steps for UI-heavy phases
  • REQ-UI-DETECT-04: System MUST NOT make /msd-ui-phase mandatory

Process:

  1. Detect — Scan project description and tech stack for UI signals (keywords, framework references)
  2. Annotate — Add ui_hint markers to applicable phases in ROADMAP.md
  3. Surface — Include /msd-ui-phase recommendation in next steps for UI-heavy phases

55. Multi-Runtime Installer Selection

Part of: npx @golem15/msd-core

Purpose: Select multiple runtimes in a single interactive install session.

Requirements:

  • REQ-MULTI-RT-01: Interactive prompt MUST support multi-select (e.g., Claude Code + Antigravity)
  • REQ-MULTI-RT-02: CLI flags MUST continue to work for non-interactive installs

Process:

  1. Detect — Identify available AI CLI runtimes on the system
  2. Prompt — Present multi-select interface for runtime selection
  3. Install — Configure MSD for all selected runtimes in a single session

v1.29 Features

57. Internationalized Documentation

Part of: docs/

Purpose: Provide MSD documentation in Portuguese, Korean, and Japanese.

Requirements:

  • REQ-I18N-01: Documentation MUST be available in Portuguese (pt), Korean (ko), and Japanese (ja)
  • REQ-I18N-02: Translations MUST stay synchronized with English source documents

Process:

  1. Translate — Convert core documentation into target languages
  2. Publish — Make translated documentation accessible alongside English originals

v1.31 Features

59. Schema Drift Detection

Command: Automatic during /msd-execute-phase

Purpose: Detect when ORM schema files are modified without corresponding migration or push commands, preventing false-positive verification.

Requirements:

  • REQ-SCHEMA-01: System MUST detect modifications to ORM schema files (Prisma, Drizzle, Payload, Sanity, Mongoose)
  • REQ-SCHEMA-02: System MUST verify corresponding migration/push commands exist when schema changes are detected
  • REQ-SCHEMA-03: System MUST implement two-layer defense: plan-time injection and execute-time gate
  • REQ-SCHEMA-04: System MUST support MSD_SKIP_SCHEMA_CHECK env var to override detection
  • REQ-SCHEMA-05: System MUST prevent false-positive verification when schema is modified without migration

Process:

  1. Detect — Monitor ORM schema file modifications during plan execution
  2. Verify — Check that corresponding migration/push commands are present in the plan
  3. Gate — Block execution if schema drift is detected without migration (execute-time gate)
  4. Inject — Add migration reminders during plan generation (plan-time injection)

Config: MSD_SKIP_SCHEMA_CHECK environment variable to bypass detection.


60. Security Enforcement

Command: /msd-secure-phase <N>

Purpose: Threat-model-anchored security verification for phase implementations.

Requirements:

  • REQ-SEC-01: System MUST perform threat-model-anchored verification (not blind scanning)
  • REQ-SEC-02: System MUST support configurable OWASP ASVS verification levels (1-3)
  • REQ-SEC-03: System MUST block phase advancement based on configurable severity threshold
  • REQ-SEC-04: System MUST spawn msd-security-auditor agent for analysis

Produces:

Artifact Description
Security audit report Threat-model-anchored findings with severity classification

Process:

  1. Model — Build threat model from phase implementation context
  2. Audit — Spawn msd-security-auditor to verify against threat model
  3. Gate — Block phase advancement if findings meet or exceed security_block_on severity

Config:

Setting Type Default Description
security_enforcement boolean true Enable threat-model security verification
security_asvs_level number (1-3) 1 OWASP ASVS verification level
security_block_on string "high" Minimum severity to block phase advancement

61. Documentation Generation

Command: /msd-docs-update

Purpose: Generate and verify project documentation with accuracy checks.

Requirements:

  • REQ-DOCS-01: System MUST spawn msd-doc-writer agent to generate documentation
  • REQ-DOCS-02: System MUST spawn msd-doc-verifier agent to check accuracy
  • REQ-DOCS-03: System MUST verify generated documentation against actual implementation

Produces:

Artifact Description
Updated project documentation Generated and verified documentation files

Process:

  1. Generate — Spawn msd-doc-writer to create or update documentation from implementation
  2. Verify — Spawn msd-doc-verifier to check documentation accuracy against codebase
  3. Output — Produce verified documentation with accuracy annotations

62. Discuss Chain Mode

Flag: /msd-discuss-phase <N> --chain

Purpose: Auto-chain discuss, plan, and execute phases in one flow to reduce manual command sequencing.

Requirements:

  • REQ-CHAIN-01: System MUST auto-chain discuss → plan → execute when --chain flag is provided
  • REQ-CHAIN-02: System MUST respect all gate settings between chained phases
  • REQ-CHAIN-03: System MUST halt the chain if any phase fails

Process:

  1. Discuss — Run discuss-phase to gather context
  2. Plan — Automatically invoke plan-phase with gathered context
  3. Execute — Automatically invoke execute-phase with generated plan

63. Single-Phase Autonomous

Flag: /msd-autonomous --only N

Purpose: Execute just one phase autonomously instead of all remaining phases.

Requirements:

  • REQ-ONLY-01: System MUST execute only the specified phase number when --only N is provided
  • REQ-ONLY-02: System MUST follow the same discuss → plan → execute flow as full autonomous mode
  • REQ-ONLY-03: System MUST stop after the specified phase completes

Process:

  1. Select — Identify the target phase from --only N argument
  2. Execute — Run full autonomous flow (discuss → plan → execute) for that single phase
  3. Stop — Halt after the phase completes instead of advancing to the next

64. Scope Reduction Detection

Part of: /msd-plan-phase

Purpose: Prevent silent requirement dropping during plan generation with three-layer defense.

Requirements:

  • REQ-SCOPE-01: System MUST prohibit planners from reducing scope without explicit justification
  • REQ-SCOPE-02: System MUST have plan-checker verify requirement dimension coverage
  • REQ-SCOPE-03: System MUST have orchestrator recover dropped requirements and re-inject them
  • REQ-SCOPE-04: System MUST implement three-layer defense: planner prohibition, checker dimension, orchestrator recovery

Process:

  1. Prohibit — Planner instructions explicitly forbid scope reduction
  2. Check — Plan-checker verifies all phase requirements are covered in the plan
  3. Recover — Orchestrator detects dropped requirements and re-injects them into the planning loop

65. Claim Provenance Tagging

Part of: /msd-plan-phase --research-phase <N>

Purpose: Ensure research claims are tagged with source evidence and assumptions are logged separately.

Requirements:

  • REQ-PROVENANCE-01: Researcher MUST mark claims with source evidence references
  • REQ-PROVENANCE-02: Assumptions MUST be logged separately from sourced claims
  • REQ-PROVENANCE-03: System MUST distinguish between evidenced facts and inferred assumptions

Process:

  1. Research — Researcher gathers information from codebase and domain sources
  2. Tag — Each claim is annotated with its source (file path, documentation, API response)
  3. Separate — Assumptions without direct evidence are logged in a distinct section

66. Worktree Toggle

Config: workflow.use_worktrees: false

Purpose: Disable git worktree isolation for users who prefer sequential execution.

Requirements:

  • REQ-WORKTREE-01: System MUST respect workflow.use_worktrees setting when deciding isolation strategy
  • REQ-WORKTREE-02: System MUST default to true (worktrees enabled) for backward compatibility
  • REQ-WORKTREE-03: System MUST fall back to sequential execution when worktrees are disabled

Config:

Setting Type Default Description
workflow.use_worktrees boolean true When false, disables git worktree isolation

67. Project Code Prefixing

Config: project_code: "ABC"

Purpose: Prefix phase directory names with a project code for multi-project disambiguation.

Requirements:

  • REQ-PREFIX-01: System MUST prefix phase directories with project code when configured (e.g., ABC-01-setup/)
  • REQ-PREFIX-02: System MUST use standard naming when project_code is not set
  • REQ-PREFIX-03: System MUST apply prefix consistently across all phase operations

Config:

Setting Type Default Description
project_code string (none) Prefix for phase directory names

68. Claude Code Skills Migration

Part of: npx @golem15/msd-core

Purpose: Migrate MSD commands to Claude Code 2.1.88+ skills format with backward compatibility.

Requirements:

  • REQ-SKILLS-01: Installer MUST write skills/msd-*/SKILL.md for Claude Code 2.1.88+
  • REQ-SKILLS-02: Installer MUST auto-clean legacy commands/msd/ directory
  • REQ-SKILLS-03: Installer MUST maintain backward compatibility with older Claude Code versions via the legacy commands/msd/ path

Process:

  1. Detect — Check Claude Code version to determine skills support
  2. Migrate — Write skills/msd-*/SKILL.md files for each MSD command
  3. Clean — Remove legacy commands/msd/ directory if skills are installed
  4. Fallback — Maintain legacy commands/msd/ path compatibility for older Claude Code versions

v1.32 Features

69. STATE.md Consistency Gates

Commands: state validate [--strict], state sync [--verify], state planned-phase --phase N --plans N

Purpose: Detect and repair drift between STATE.md and the actual filesystem, preventing cascading errors from stale state.

Requirements:

  • REQ-STATE-01: state validate MUST detect drift between STATE.md fields and filesystem reality
  • REQ-STATE-02: state sync MUST reconstruct STATE.md from actual project state on disk
  • REQ-STATE-03: state sync --verify MUST perform a dry-run showing proposed changes without writing
  • REQ-STATE-04: state planned-phase MUST record the state transition after plan-phase completes (Planned/Ready to execute)
  • REQ-STATE-05: state validate MUST report a Last activity value that no reader can parse, rather than validating clean
  • REQ-STATE-06: state validate --strict MUST reflect valid in the process exit status, leaving the default exit status unchanged

Produces:

Artifact Description
Updated STATE.md Corrected state reflecting filesystem reality

Process:

  1. Validate — Compare STATE.md fields against filesystem (phase directories, plan files, summaries)
  2. Sync — Reconstruct STATE.md from disk when drift is detected
  3. Transition — Record post-planning state with plan count for execute-phase readiness

70. Autonomous --to N Flag

Flag: /msd-autonomous --to N

Purpose: Stop autonomous execution after completing a specific phase, allowing partial autonomous runs.

Requirements:

  • REQ-TO-01: System MUST stop execution after the specified phase number completes
  • REQ-TO-02: System MUST follow the same discuss -> plan -> execute flow for each phase up to N
  • REQ-TO-03: --to N MUST be combinable with --from N for bounded autonomous ranges

Process:

  1. Bound — Set the upper phase limit from --to N argument
  2. Execute — Run autonomous flow for each phase up to and including phase N
  3. Stop — Halt after phase N completes

71. Research Gate

Part of: /msd-plan-phase

Purpose: Block planning when RESEARCH.md has unresolved open questions, preventing plans built on incomplete information.

Requirements:

  • REQ-RESGATE-01: System MUST scan RESEARCH.md for unresolved open questions before planning begins
  • REQ-RESGATE-02: System MUST block plan-phase entry when open questions exist
  • REQ-RESGATE-03: System MUST surface the specific unresolved questions to the user

Process:

  1. Scan — Check RESEARCH.md for open questions section with unresolved items
  2. Gate — Block planning if unresolved questions are found
  3. Surface — Display the specific open questions requiring resolution

72. Verifier Milestone Scope Filtering

Part of: /msd-execute-phase (verifier step)

Purpose: Distinguish between genuine gaps and items deferred to later phases, reducing false negatives in verification.

Requirements:

  • REQ-VSCOPE-01: Verifier MUST check whether a gap is addressed in a later milestone phase
  • REQ-VSCOPE-02: Gaps addressed in later phases MUST be marked as "deferred", not "gap"
  • REQ-VSCOPE-03: Only genuine gaps (not covered by any future phase) MUST be reported as failures

Process:

  1. Verify — Run standard goal-backward verification
  2. Filter — Cross-reference detected gaps against later milestone phases
  3. Classify — Mark deferred items separately from genuine gaps

73. Read-Before-Edit Guard Hook

Part of: Hooks (PreToolUse)

Purpose: Prevent infinite retry loops in non-Claude runtimes by ensuring files are read before editing.

Requirements:

  • REQ-RBE-01: Hook MUST detect Edit/Write tool calls that target files not previously read in the session
  • REQ-RBE-02: Hook MUST advise reading the file first (advisory, non-blocking)
  • REQ-RBE-03: Hook MUST prevent infinite retry loops common in runtimes without built-in read-before-edit enforcement

74. Context Reduction

Part of: prompt assembly pipeline

Purpose: Reduce context prompt sizes through markdown truncation and cache-friendly prompt ordering.

Requirements:

  • REQ-CTXRED-01: System MUST truncate oversized markdown artifacts to fit within context budgets
  • REQ-CTXRED-02: System MUST order prompts for cache-friendly assembly (stable prefixes first)
  • REQ-CTXRED-03: Reduction MUST preserve essential information (headings, requirements, task structure)
  • REQ-CTXRED-04: Skill description: fields MUST be ≤ 100 chars; enforced by npm run lint:descriptions (see scripts/lint-descriptions.cjs and tests/skill-frontmatter-contract.test.cjs)

Process:

  1. Measure — Calculate total prompt size for the workflow
  2. Truncate — Apply markdown-aware truncation to oversized artifacts
  3. Order — Arrange prompt sections for optimal KV-cache reuse

75. Discuss-Phase --power Flag

Flag: /msd-discuss-phase --power

Purpose: File-based bulk question answering for discuss-phase, enabling batch input from a prepared answers file.

Requirements:

  • REQ-POWER-01: System MUST accept a file containing pre-written answers to discussion questions
  • REQ-POWER-02: System MUST map answers to the corresponding gray area questions
  • REQ-POWER-03: System MUST produce CONTEXT.md identical to interactive discuss-phase

76. Debug --diagnose Flag

Flag: /msd-debug --diagnose

Purpose: Diagnosis-only mode that investigates without attempting fixes.

Requirements:

  • REQ-DIAG-01: System MUST perform full debug investigation (hypotheses, evidence, root cause)
  • REQ-DIAG-02: System MUST NOT attempt any code modifications
  • REQ-DIAG-03: System MUST produce a diagnostic report with findings and recommended fixes

77. Phase Dependency Analysis

Command: /msd-manager --analyze-deps

Purpose: Detect phase dependencies and suggest Depends on entries for ROADMAP.md before running /msd-manager.

Requirements:

  • REQ-DEP-01: System MUST detect file overlap between phases
  • REQ-DEP-02: System MUST detect semantic dependencies (API/schema producers and consumers)
  • REQ-DEP-03: System MUST detect data flow dependencies (output producers and readers)
  • REQ-DEP-04: System MUST suggest dependency entries with user confirmation before writing

Produces: Dependency suggestion table; optionally updates ROADMAP.md Depends on fields


78. Anti-Pattern Severity Levels

Part of: /msd-resume-work

Purpose: Mandatory understanding checks at resume with severity-based anti-pattern enforcement.

Requirements:

  • REQ-ANTI-01: System MUST classify anti-patterns by severity level
  • REQ-ANTI-02: System MUST enforce mandatory understanding checks at session resume
  • REQ-ANTI-03: Higher severity anti-patterns MUST block workflow progression until acknowledged

79. Methodology Artifact Type

Part of: Planning artifacts

Purpose: Define consumption mechanisms for methodology documents, ensuring they are consumed correctly by agents.

Requirements:

  • REQ-METHOD-01: System MUST support methodology as a distinct artifact type
  • REQ-METHOD-02: Methodology artifacts MUST have defined consumption mechanisms for agents

80. Planner Reachability Check

Part of: /msd-plan-phase

Purpose: Validate that plan steps are achievable before committing to execution.

Requirements:

  • REQ-REACH-01: Planner MUST validate that each plan step references reachable files and APIs
  • REQ-REACH-02: Unreachable steps MUST be flagged during planning, not discovered during execution

81. Playwright-MCP UI Verification

Part of: /msd-verify-work (optional)

Purpose: Automated visual verification using Playwright-MCP during verify-phase.

Requirements:

  • REQ-PLAY-01: System MUST support optional Playwright-MCP visual verification during verify-phase
  • REQ-PLAY-02: Visual verification MUST be opt-in, not mandatory
  • REQ-PLAY-03: System MUST capture and compare visual state against UI-SPEC.md expectations

82. Pause-Work Expansion

Part of: /msd-pause-work

Purpose: Support non-phase contexts with richer handoff data for broader pause-work applicability.

Requirements:

  • REQ-PAUSE-01: System MUST support pausing in non-phase contexts (quick tasks, debug sessions, threads)
  • REQ-PAUSE-02: Handoff data MUST include richer context appropriate to the current work type

83. Response Language Config

Config: response_language

Purpose: Cross-phase language consistency for non-English users.

Requirements:

  • REQ-LANG-01: System MUST respect response_language setting across all phases and agents
  • REQ-LANG-02: Setting MUST propagate to all spawned agents for consistent language output
  • REQ-LANG-03: Every workflow MUST carry response-language coverage — through an exact inline directive, a shared @-referenced directive (msd-core/references/response-language-directive.md), or inheritance from the parent workflow that dispatches it; enforced in CI by scripts/lint-response-language-coverage.cjs (#2529)
  • REQ-LANG-04: A covering directive MUST name inter-tool narration, not only the question/prompt surface. A directive names it by using the word "narration" or the phrase "between tool calls"; the class it denotes is the model's running commentary between tool calls, status updates, progress notes and findings included, and enumerating those items without naming the class does not satisfy the rule. A directive worded around questions and prompts alone leaves the model's running commentary in English beside translated answers, which is the defect #2529 reports; scripts/lint-response-language-coverage.cjs rejects it (#2529)

Config:

Setting Type Default Description
response_language string (none) Language code for agent responses (e.g., "pt", "ko", "ja")

84. Manual Update Procedure

Part of: docs/manual-update.md

Purpose: Document a manual update path for environments where npx is unavailable or npm publish is experiencing outages.

Requirements:

  • REQ-MANUAL-01: Documentation MUST describe step-by-step manual update procedure
  • REQ-MANUAL-02: Procedure MUST work without npm access

86. Autonomous --interactive Flag

Flag: /msd-autonomous --interactive

Purpose: Lean-context autonomous mode that keeps discuss-phase interactive (user answers questions) while dispatching plan and execute as background agents on runtimes that support nested background dispatch; on Claude Code, plan and execute run inline to preserve worktree isolation and independent verification.

Requirements:

  • REQ-INTERACT-01: --interactive MUST run discuss-phase inline with interactive questions (not auto-answered)
  • REQ-INTERACT-02: --interactive MUST dispatch plan-phase and execute-phase as background agents for context isolation on runtimes where a backgrounded agent can spawn subagents; on Claude Code, plan and execute run inline
  • REQ-INTERACT-03: --interactive MUST enable pipeline parallelism — discuss Phase N+1 while Phase N builds (applies on runtimes that support nested background dispatch; on Claude Code, discuss does not overlap planning/execution)
  • REQ-INTERACT-04: Main context MUST only accumulate discuss conversations (lean context) on runtimes that support nested background dispatch; on Claude Code, inline plan/execute also accumulate in the main context

Process:

  1. Discuss inline — Run discuss-phase in the main context with user interaction
  2. Dispatch — On runtimes that support nested background dispatch: send plan and execute to background agents with fresh context windows. On Claude Code: run plan and execute inline.
  3. Pipeline — On runtimes with background dispatch: while background agents build Phase N, begin discussing Phase N+1. On Claude Code: phases run sequentially.

87. Commit-Docs Guard Hook

Hook: msd-commit-docs.js

Purpose: PreToolUse hook that enforces the commit_docs configuration, preventing .planning/ files from being committed when planning.commit_docs is false.

Requirements:

  • REQ-COMMITDOCS-01: Hook MUST intercept git commit commands that stage .planning/ files
  • REQ-COMMITDOCS-02: Hook MUST block commits containing .planning/ files when commit_docs is false
  • REQ-COMMITDOCS-03: Hook MUST be advisory — does not block when commit_docs is true or absent

88. Community Hooks Opt-In

Hooks: msd-validate-commit.sh, msd-session-state.sh, msd-phase-boundary.sh

Purpose: Optional git and session hooks for MSD projects, gated behind hooks.community: true in config.

Requirements:

  • REQ-COMMUNITY-01: All community hooks MUST be no-ops unless hooks.community is true in .planning/config.json
  • REQ-COMMUNITY-02: msd-validate-commit.sh MUST enforce Conventional Commits format on git commit messages
  • REQ-COMMUNITY-03: msd-session-state.sh MUST track session state transitions
  • REQ-COMMUNITY-04: msd-phase-boundary.sh MUST enforce phase boundary checks

Config:

Setting Type Default Description
hooks.community boolean false Enable optional community hooks for commit validation, session state, and phase boundaries
hooks.commit_types array of strings [] Extra Conventional Commits types msd-validate-commit.sh accepts, in addition to the built-in feat, fix, docs, style, refactor, perf, test, build, ci, chore — never replaces them. Each entry must match ^[a-z][a-z0-9-]*$ (lowercase letters, digits, hyphens); non-conforming or non-string entries are dropped. Example: { "hooks": { "community": true, "commit_types": ["enhance", "enh", "revert"] } }.

v1.34.0 Features

89. Global Learnings Store

Commands: Auto-triggered at phase completion; consumed by planner Config: features.global_learnings

Purpose: Persist cross-session, cross-project learnings in a global store so the planner agent can learn from patterns across the entire project history — not just the current session.

Requirements:

  • REQ-LEARN-01: Learnings MUST be auto-copied from .planning/ to the global store at phase completion
  • REQ-LEARN-02: The planner agent MUST receive relevant learnings at spawn time via injection
  • REQ-LEARN-03: Injection MUST be capped by learnings.max_inject to avoid context bloat
  • REQ-LEARN-04: Feature MUST be opt-in via features.global_learnings: true

Config:

Setting Type Default Description
features.global_learnings boolean false Enable cross-project learnings pipeline
learnings.max_inject number (system default) Maximum learnings entries injected into planner

90. Queryable Codebase Intelligence

Command: /msd-map-codebase --query [<term>|status|diff|refresh] Config: intel.enabled

Purpose: Maintain a queryable JSON index of codebase structure, API surface, dependency graph, file roles, and architecture decisions in .planning/intel/. Enables targeted lookups without reading the entire codebase.

Requirements:

  • REQ-INTEL-01: Intel files MUST be stored as JSON in .planning/intel/
  • REQ-INTEL-02: query mode MUST search across all intel files for a term and group results by file
  • REQ-INTEL-03: status mode MUST report freshness (FRESH/STALE, stale threshold: 24 hours)
  • REQ-INTEL-04: diff mode MUST compare current intel state to the last snapshot
  • REQ-INTEL-05: refresh mode MUST spawn the intel-updater agent to rebuild all files
  • REQ-INTEL-06: Feature MUST be opt-in via intel.enabled: true

Intel files produced:

File Contents
stack.json Technology stack and dependencies
api-map.json Exported functions and API surface
dependency-graph.json Inter-module dependency relationships
file-roles.json Role classification for each source file
arch-decisions.json Detected architecture decisions

91. Execution Context Profiles

Config: context_profile

Purpose: Select a pre-configured execution context (mode, model, workflow settings) tuned for a specific type of work without manually adjusting individual settings.

Requirements:

  • REQ-CTX-01: dev profile MUST optimize for iterative development (balanced model, plan_check enabled)
  • REQ-CTX-02: research profile MUST optimize for research-heavy work (higher model tier, research enabled)
  • REQ-CTX-03: review profile MUST optimize for code review work (verifier and code_review enabled)

Available profiles: dev, research, review

Config:

Setting Type Default Description
context_profile string (none) Execution context preset: dev, research, or review

92. Gates Taxonomy

References: msd-core/references/gates.md Agents: plan-checker, verifier

Purpose: Define 4 canonical gate types that structure all workflow decision points, enabling plan-checker and verifier agents to apply consistent gate logic.

Gate types:

Type Description
Confirm User approves before proceeding (e.g., roadmap review)
Quality Automated quality check must pass (e.g., plan verification loop)
Safety Hard stop on detected risk or policy violation
Transition Phase or milestone boundary acknowledgment

Requirements:

  • REQ-GATES-01: plan-checker MUST classify each checkpoint as one of the 4 gate types
  • REQ-GATES-02: verifier MUST apply gate logic appropriate to the gate type
  • REQ-GATES-03: Hard stop safety gates MUST never be bypassed by --auto flags

93. Code Review Pipeline

Commands: /msd-code-review, /msd-code-review --fix

Purpose: Structured review of source files changed during a phase, with a separate auto-fix pass that commits each fix atomically.

Requirements:

  • REQ-REVIEW-01: msd-code-review MUST scope files to the phase using SUMMARY.md and git diff fallback
  • REQ-REVIEW-02: Review MUST support three depth levels: quick, standard, deep
  • REQ-REVIEW-03: Findings MUST be severity-classified: Critical, Warning, Info
  • REQ-REVIEW-04: msd-code-review --fix MUST read REVIEW.md and fix Critical + Warning findings by default
  • REQ-REVIEW-05: Each fix MUST be committed atomically with a descriptive message
  • REQ-REVIEW-06: --auto flag MUST enable fix + re-review iteration loop, capped at 3 iterations
  • REQ-REVIEW-07: Feature MUST be gated by workflow.code_review config flag
  • REQ-REVIEW-08: workflow.code_review_point MUST select which loop point the automatic review step registers at (execute:post default, or execute:wave:post), independent of the workflow.code_review on/off gate and of manual /msd-code-review invocation (#3661)
  • REQ-REVIEW-09: The in-phase code_review_gate MUST report the per-severity counts it parses from REVIEW.md, so a review with one info finding is distinguishable from a review with a Critical
  • REQ-REVIEW-10: Each finding MUST carry a recorded disposition, so a triaged finding is distinguishable from a forgotten one

Config:

Setting Type Default Description
workflow.code_review boolean true Enable code review commands
workflow.code_review_point string execute:post Loop point for the automatic review: execute:post (once per phase, default) or execute:wave:post (once per completed wave, scoped to what changed since the phase's prior review). See below.
workflow.code_review_depth string standard Default review depth: quick, standard, or deep
workflow.code_review_depth_overrides array [] Ordered { paths, depth } rules that escalate depth for directories matched by path prefix against the changed-file set (#2554). See below.

Reviewing per wave instead of per phase (#3661)

Setting workflow.code_review_point to execute:wave:post moves the automatic review from "once, after the whole phase's waves have all landed" to "once per completed wave." Each wave's review scopes to what changed since the phase's previous review — the whole phase's diff on the first wave, then just that wave's diff on every wave after — so review batches stay small instead of growing with the phase. A finding introduced early is caught after the wave that introduced it, not after the last wave of the phase.

This only affects the automatic dispatch inside a wave-based phase execution. Manual /msd-code-review <phase> runs are gated by workflow.code_review alone and are unaffected by this key. /msd-autonomous and /msd-quick have no wave granularity of their own, so setting this to execute:wave:post means automatic review does not run inside those two flows — the same way every other wave-scoped capability step already behaves for them.

Path-scoped code review depth overrides

workflow.code_review_depth_overrides matches rules against the review's changed-file set by whole-segment directory-path prefix — src/auth matches src/auth/token.ts and src/auth itself, never src/authfoo/x.ts or docs/src/auth/x.ts — and is case-sensitive, following git.

Escalation is whole-review, not per-file: depth is a single scalar handed to the reviewer agent, not a per-file setting, so the strongest matching tier across the whole rule set applies to every file in the review — a sensitive file is never reviewed shallowly because it shared a review with an unrelated one.

v1 supports directory-prefix matching only, not glob syntax: no glob engine (minimatch, picomatch, fast-glob) exists in this project and none was added for this feature. A path containing * or ? (e.g. src/auth/**) is a configuration error rather than a silent near-miss, because accepting it as sugar for a prefix would make unsupported patterns look armed when they match nothing. Every use case in the issue is expressible as a directory prefix. See Scope code review depth by path for the resolution order, error table, and a worked example.

Optional external reviewer lanes (#4209): /msd-code-review accepts the same reviewer-lane flags as /msd-review — any flag the roster declares (run msd_run review-lane flags to list them for your installation, e.g. --codex, --agy). No reviewer-lane flag is the default and is byte-for-byte unchanged from before #4209: zero lane selection, plan, or invoke calls, and only the internal msd-code-reviewer agent runs. Passing one or more flags asks those lanes to independently review the same already-resolved file scope alongside the internal agent, through the same shared capability-trait interpreter and review-lane plan/invoke machinery /msd-review uses — no second implementation. Each lane's prompt carries only the repository root, canonical file paths, review depth, and base SHA, never source file contents, under four fixed prohibitions (no source mutation, no test execution, no background processes, no polling). External findings are unverified corroborating evidence: msd-code-reviewer independently re-verifies every claim against the actual source before writing it to REVIEW.md, so there remains exactly one REVIEW.md schema regardless of how many lanes ran. An explicitly requested lane that is unavailable or fails is reported as a warning, never silently dropped and never a raw-CLI fallback. This is separate from /msd-review, which reviews PLAN.md files before execution, not source code. In-phase review reporting and disposition

/msd-execute-phase's code_review_gate runs code review, then reports what it found:

Code review: 23 findings — 1 critical, 9 warning, 8 info.
Consider running: /msd-code-review 1 --fix

Both critical: and its documented tier-equivalent blocker: are accepted. A REVIEW.md written without a findings: block has no counts to report, and the gate falls back to the countless form rather than printing a half-filled line.

The gate then writes <NN>-REVIEW-DISPOSITION.md beside the review — one row per finding ID, defaulting to open:

Finding Severity Disposition Source
CR-01 critical open -
WR-01 warning fixed 01-REVIEW-FIX.md

open means recorded but not yet triaged, and it is the only value the gate assigns on its own.

Where the reconciliation happens, which is not where you might expect. The in-phase gate runs immediately after review, and at that moment <NN>-REVIEW-FIX.md does not exist — the gate invokes review with neither --fix nor --auto — so every row it writes is open. The fixed and skipped outcomes are reconciled by /msd-code-review <N> --fix, which records them once the fix report is on disk. Running the gate alone therefore tells you what was found; running --fix is what records what happened to it.

--auto's iterations are reconciled too, and that takes reading more than the final report. The loop overwrites REVIEW-FIX.md on every pass and the re-review stops reporting a finding once it is fixed, so a finding closed in iteration 1 appears in neither final artifact. The gate therefore also reads the per-iteration backups the loop writes (<NN>-REVIEW-FIX.iterN.md), newest first, so the most recent statement about an ID wins; the backups are removed after the ledger has read them, not before. Without this a fully successful multi-iteration run recorded its early fixes as open (not in the current review) — indistinguishable from a finding that vanished for an unrelated reason, which is the one distinction this ledger exists to make. A finding a fix report decided but the current review no longer reports gets a row of its own, carrying that decision and marked (not in the current review).

A converged --auto run also reaches the gate with a clean review and, on a direct /msd-code-review invocation, no ledger from the in-phase gate. A fix report on disk is reason enough to record: without that, a run in which every finding was fixed and committed produced no disposition record at all.

A decision belongs to a finding, not to an ID. IDs are reused across re-reviews — --auto renumbers — so the ledger records each finding's title alongside its disposition, and a recorded decision is carried forward only while the ID still names the same finding. Without that, a prior CR-01 fixed would be inherited by a brand-new CR-01, which is a false decision in the one artifact whose purpose is telling triaged from forgotten.

The limitation that follows, stated rather than hidden. When an ID is reused, the new finding is recorded open (nobody decided anything about it) and the earlier decision loses its row — the ledger keys rows on the finding ID, and two rows under one ID is an ambiguity, not a record. The drop is reported on the console, naming the ID and what had been decided — for a recorded decision. A prior row still sitting at open is replaced silently, and deliberately: open means nobody had decided anything, so there is no decision to lose. Preserving a recorded decision in the file was tried and withdrawn: it needed a second identity scheme and produced a fresh defect on each of three review passes. Recovery is the ledger's own git history where commit_docs is on — which is why the console reports the drop rather than pointing at a commit that may not exist.

Two residuals, since (ID, title) is not proof of identity. A ledger written before titles were recorded carries none, so its decisions are inherited on the ID alone — refusing there would reset every decision in every existing ledger, which is the loss the guard exists to prevent. And two genuinely distinct findings that share both an ID and a title are indistinguishable to this key. Separating them needs a second identity scheme, which is the thing that was just withdrawn.

Reconciliation applies an outcome only when the fix report names the same finding, because finding IDs are reused across re-reviews and a stale report would otherwise declare a brand-new CR-01 already fixed. Titles are compared with runs of whitespace collapsed, so a fixer that re-spaces a title still reconciles. A title that differs otherwise — including one wrapped across lines, since a ### heading is one line and the continuation is a separate paragraph — leaves the row open and is reported, naming the ids it could not reconcile, rather than passing silently. The report does not claim to know whether such a report is stale or merely re-titled, because it cannot tell.

The disposition column is a closed vocabulary — open, fixed, skipped, deferred. A hand-edited value outside it is not treated as a decision: the row falls back to open, so a typo cannot quietly mark a phase as triaged.

Severity comes from the section a finding sits under (## Critical Issues, ## Warnings, ## Info) when the review uses those headings, and from the ID prefix otherwise. The section is the reviewer's own statement of severity, so a Critical filed under ## Critical Issues is recorded critical even if its ID was mis-numbered WR-04. That severity is then remembered: a row the current review no longer reports, or reports under no recognized heading, keeps the severity the ledger recorded rather than having it re-inferred from the ID prefix — so the mis-numbered WR-04 above stays critical on the second run, deferred or not. The precedence is the current review's section, then the recorded value, then the prefix; the recorded value is inherited only while the ID still names the same finding, by the same title check the disposition uses, so a reused ID starts from its own review. That check has the same compatibility arm the disposition has: a ledger written before titles were recorded carries no title to compare, so its severities — like its decisions — inherit on the ID alone.

If the review's total: exceeds the number of findings whose headings the gate could parse, the shortfall is stated — on the console and as an unparsed: key in the ledger's frontmatter. A finding the gate cannot record is the one a human most needs to see, so it is never dropped silently. The one input that produces no shortfall is a findings: block that disagrees with itself: where critical (or blocker), warning and info are all present, numeric and do not sum to total, the gate has no trustworthy number to reconcile against and withholds the key rather than reporting a figure derived from one — the same input on which the console line already withholds the severity breakdown. Counts that are merely absent, partial or non-numeric are not a disagreement, and the shortfall is still reported from total alone. deferred is the one disposition the gate never writes: it is recorded by hand, and the reason recorded beside it in the Source cell is preserved across re-runs, a literal | included once escaped. One exception, because it cannot be resolved: a reason ending in the literal phrase (not in the current review) loses that trailing phrase, since it is indistinguishable from the carried marker the gate appends. The alternative is worse — a stored marker never leaves, so a carried finding that later reappears would keep claiming it is absent from the review reporting it.

Re-running the gate keeps every row it can, so a decision recorded here is not overwritten by a later pass. A finding the current review no longer reports — --auto re-reviews and rewrites REVIEW.md, so this happens routinely — is carried rather than dropped, marked (not in the current review). That holds whether or not it was triaged: losing a decided row would erase the record that the finding was seen, and losing an untriaged one would erase the record that it was never answered, which is the trace this ledger exists to keep. The cost is that a renumbered finding shows under both IDs until the old row is decided; the marker makes that legible. A run that changes nothing rewrites nothing, so a re-executed phase does not produce a docs commit with no content.

One residual is concurrency: the ledger is read, rebuilt and written whole, with no lock. Two dispatchers can run this step (code_review_gate and code-review-fix's record_disposition) and a human is invited to hand-edit the file, so two writers overlapping would lose one's update. This is the shape #3780 reported for WINDOWS.md under parallel executors, closed there by a cross-process lock (#4681); the disposition step does not take that lock. No lost update has been reproduced; the window is stated so it is not mistaken for a guarantee.

Not to be confused with the Review Dispositions Ledger of reviews-mode planning (ADR-3806, docs/features/review-dispositions-ledger.md): that one is a ## Review Dispositions Ledger section inside PLAN.md, append-only per round, over REVIEWS.md findings. This artifact is a sibling file beside REVIEW.md, rewritten idempotently with rows carried. Same word, different inputs, writers, files and durability rules; neither governs the other.

The record is a sibling artifact rather than a section inside REVIEW.md because --auto's re-review loop rewrites REVIEW.md on every iteration — a ledger kept inside it would not survive the next pass — and because REVIEW.md has a single writer (msd-code-reviewer) that the gate is not. The gate remains advisory throughout: it reports and records, and never blocks phase completion.


94. Socratic Exploration

Command: /msd-explore [topic]

Purpose: Guide a developer through exploring an idea via Socratic probing questions before committing to a plan. Routes outputs to the appropriate MSD artifact: notes, todos, seeds, research questions, requirements updates, or a new phase.

Requirements:

  • REQ-EXPLORE-01: Exploration MUST use Socratic probing — ask questions before proposing solutions
  • REQ-EXPLORE-02: Session MUST offer to route outputs to the appropriate MSD artifact
  • REQ-EXPLORE-03: An optional topic argument MUST prime the first question
  • REQ-EXPLORE-04: Exploration MUST optionally spawn a research agent for technical feasibility
  • REQ-EXPLORE-05: A research pass MUST disposition each surfaced claim (admit / refute / abstain) and route every abstention to a visible Unresolved Ledger — never smoothing an ungrounded claim into the narrative as confident prose

95. Safe Undo

Command: /msd-undo --last N | --phase NN | --plan NN-MM

Purpose: Roll back MSD phase or plan commits safely using the phase manifest and git log, with dependency checks and a hard confirmation gate before any revert is applied.

Requirements:

  • REQ-UNDO-01: --phase mode MUST identify all commits for the phase via manifest and git log fallback
  • REQ-UNDO-02: --plan mode MUST identify all commits for a specific plan
  • REQ-UNDO-03: --last N mode MUST display recent MSD commits for interactive selection
  • REQ-UNDO-04: System MUST check for dependent phases/plans before reverting
  • REQ-UNDO-05: A confirmation gate MUST be shown before any git revert is executed

96. Plan Import

Command: /msd-import --from <filepath>

Purpose: Ingest an external plan file into the MSD planning system with conflict detection against PROJECT.md decisions, converting it to a valid MSD PLAN.md and validating it through the plan-checker.

Requirements:

  • REQ-IMPORT-01: Importer MUST detect conflicts between the external plan and existing PROJECT.md decisions
  • REQ-IMPORT-02: All detected conflicts MUST be presented to the user for resolution before writing
  • REQ-IMPORT-03: Imported plan MUST be written as a valid MSD PLAN.md format
  • REQ-IMPORT-04: Written plan MUST pass msd-plan-checker validation

97. Rapid Codebase Scan

Command: /msd-map-codebase --fast [--focus tech|arch|quality|concerns]

Purpose: Lightweight alternative to /msd-map-codebase that spawns a single mapper agent for one or two combined focus areas, producing targeted output in .planning/codebase/ without the overhead of 4 parallel agents.

Requirements:

  • REQ-SCAN-01: Scan MUST spawn exactly one mapper agent (not four parallel agents)
  • REQ-SCAN-02: Focus area MUST be one of: tech, arch, quality, concerns, or the combined tech+arch shorthand (default: tech+arch); combined focus runs as a single agent covering both areas in one pass
  • REQ-SCAN-03: Output MUST be written to .planning/codebase/ in the same format as /msd-map-codebase

98. Autonomous Audit-to-Fix

Command: /msd-audit-fix [--source <audit>] [--severity high|medium|all] [--max N] [--dry-run]

Purpose: End-to-end pipeline that runs an audit, classifies findings as auto-fixable vs. manual-only, then autonomously fixes auto-fixable issues with test verification and atomic commits.

Requirements:

  • REQ-AUDITFIX-01: Findings MUST be classified as auto-fixable or manual-only before any changes
  • REQ-AUDITFIX-02: Each fix MUST be verified with tests before committing
  • REQ-AUDITFIX-03: Each fix MUST be committed atomically
  • REQ-AUDITFIX-04: --dry-run MUST show classification table without applying any fixes
  • REQ-AUDITFIX-05: --max N MUST limit the number of fixes applied in one run (default: 5)

99. Improved Prompt Injection Scanner

Hook: msd-prompt-guard.js, msd-read-injection-scanner.js Script: scripts/prompt-injection-scan.sh, scripts/base64-scan.sh

Purpose: Defense-in-depth detection of prompt injection attempts in planning artifacts and ingested content. Live hooks inline their own pattern subsets for hook independence (they do not import from security.cts). The CI scanner (scanForInjection in security.cts) provides a centralized engine for codebase-wide scanning in tests.

Requirements:

  • REQ-SCAN-INJ-01: Live hooks MUST detect invisible Unicode characters (zero-width spaces, soft hyphens, Unicode tag block U+E0000–E007F)

  • REQ-SCAN-INJ-02: Live hooks MUST detect known injection patterns (instruction override, role manipulation, system-prompt extraction, fake message boundaries). Base64-decode scanning is a CI-time control (scripts/base64-scan.sh), not a live hook — live hooks match a base64-exfiltration phrase regex only, they do not decode.

  • REQ-SCAN-INJ-03: Scanner MUST apply entropy analysis — Entropy analysis (scanEntropyAnomalies) was removed in #2198 as dead code (zero production callers; live hooks do not perform entropy analysis). This requirement is deferred pending a maintainable live implementation.

  • REQ-SCAN-INJ-04: Scanner MUST remain advisory-only — detection is logged, not blocking

  • REQ-SCAN-INJ-05: A scanner that could not establish its file list MUST NOT report clean (#3908). The CI scanners (prompt-injection-scan.sh, base64-scan.sh, secret-scan.sh) distinguish four outcomes rather than collapsing them into exit 0:

    Outcome Exit Meaning
    scanned, no findings 0 files were in scope and none matched
    findings 1 the scan's own verdict
    nothing in scope NO_INPUT the diff resolved and was genuinely empty — e.g. a docs-only PR
    could not scan UNAVAILABLE the file list was never established: a bad ref, no repository, or a repository with no commits

    Codes come from the exit-code registry (ADR-3889), sourced from msd-core/bin/shared/exit-codes.sh, never written into the scripts. Every one is non-zero, so a caller written if ! scanner; then behaves identically for a clean scan and trips for everything else — this can turn a false green red, never a red green. .github/workflows/security-scan.yml treats nothing in scope as a pass and could not scan as a failure; previously the latter passed silently, having scanned nothing.


100. Stall Detection in Plan-Phase

Command: /msd-plan-phase

Purpose: Detect when the planner revision loop has stalled — producing the same output across multiple iterations — and break the cycle by escalating to a different strategy or exiting with a clear diagnostic.

Requirements:

  • REQ-STALL-01: Revision loop MUST detect identical plan output across consecutive iterations
  • REQ-STALL-02: On stall detection, system MUST escalate strategy before retrying
  • REQ-STALL-03: Maximum stall retries MUST be bounded (capped at the existing max 3 iterations)

101. Hard Stop Safety Gates in /msd-progress --next

Command: /msd-progress --next

Purpose: Prevent /msd-progress --next from entering runaway loops by adding hard stop safety gates and a consecutive-call guard that interrupts autonomous chaining when repeated identical steps are detected.

Requirements:

  • REQ-NEXT-GATE-01: /msd-progress --next MUST track consecutive same-step calls
  • REQ-NEXT-GATE-02: On repeated same-step, system MUST present a hard stop gate to the user
  • REQ-NEXT-GATE-03: User MUST explicitly confirm to continue past a hard stop gate

102. Adaptive Model Preset

Config: model_profile: "adaptive"

Purpose: Role-based model assignment that automatically selects the appropriate model tier based on the current agent's role, rather than applying a single tier to all agents.

Requirements:

  • REQ-ADAPTIVE-01: adaptive preset MUST assign model tiers based on agent role (planner → quality tier, executor → balanced tier, etc.)
  • REQ-ADAPTIVE-02: adaptive MUST be selectable via /msd-config --profile adaptive

103. Post-Merge Hunk Verification

Command: /msd-update --reapply

Purpose: After applying local patches post-update, verify that all hunks were actually applied by comparing the expected patch content against the live filesystem. Surface any dropped or partial hunks immediately rather than silently accepting incomplete merges.

Requirements:

  • REQ-PATCH-VERIFY-01: Reapply-patches MUST verify each hunk was applied after the merge
  • REQ-PATCH-VERIFY-02: Dropped or partial hunks MUST be reported to the user with file and line context
  • REQ-PATCH-VERIFY-03: Verification MUST run after all patches are applied, not per-patch

v1.35.0 Features

105. GSD-2 Reverse Migration

Command: /msd-import --from-gsd2 [--dry-run] [--force] [--path <dir>]

Purpose: Migrate a project from GSD-2 format (.gsd/ directory with Milestone→Slice→Task hierarchy) back to the v1 .planning/ format, restoring full compatibility with all MSD v1 commands.

Requirements:

  • REQ-FROM-MSD2-01: Importer MUST read .gsd/ from the specified or current directory
  • REQ-FROM-MSD2-02: Milestone→Slice hierarchy MUST be flattened to sequential phase numbers (M001/S01→phase 01, M001/S02→phase 02, M002/S01→phase 03, etc.)
  • REQ-FROM-MSD2-03: System MUST guard against overwriting an existing .planning/ directory without --force
  • REQ-FROM-MSD2-04: --dry-run MUST preview all changes without writing any files
  • REQ-FROM-MSD2-05: Migration MUST produce PROJECT.md, REQUIREMENTS.md, ROADMAP.md, STATE.md, and sequential phase directories

Flags:

Flag Description
--dry-run Preview migration output without writing files
--force Overwrite an existing .planning/ directory
--path <dir> Specify the GSD-2 root directory

106. AI Integration Phase Wizard

Command: /msd-ai-integration-phase [N]

Purpose: Guide developers through selecting, integrating, and planning evaluation for AI/LLM capabilities in a project phase. Produces a structured AI-SPEC.md that feeds into planning and verification.

Requirements:

  • REQ-AISPEC-01: Wizard MUST present an interactive decision matrix covering framework selection, model choice, and integration approach
  • REQ-AISPEC-02: System MUST surface domain-specific failure modes and eval criteria relevant to the project type
  • REQ-AISPEC-03: System MUST spawn 3 parallel specialist agents: domain-researcher, framework-selector, and eval-planner
  • REQ-AISPEC-04: Output MUST produce {phase}-AI-SPEC.md with framework recommendation, implementation guidance, and evaluation strategy

Produces: {phase}-AI-SPEC.md in the phase directory


107. AI Eval Review

Command: /msd-eval-review [N]

Purpose: Retroactively audit an executed AI phase's evaluation coverage against the AI-SPEC.md plan. Identifies gaps between planned and implemented evaluation before the phase is closed.

Requirements:

  • REQ-EVALREVIEW-01: Review MUST read AI-SPEC.md from the specified phase
  • REQ-EVALREVIEW-02: Each eval dimension MUST be scored as COVERED, PARTIAL, or MISSING
  • REQ-EVALREVIEW-03: Output MUST include findings, gap descriptions, and remediation guidance
  • REQ-EVALREVIEW-04: EVAL-REVIEW.md MUST be written to the phase directory

Produces: {phase}-EVAL-REVIEW.md with scored eval dimensions, gap analysis, and remediation steps


v1.36.0 Features

108. Plan Bounce

Command: /msd-plan-phase N --bounce

Purpose: After plans pass the checker, optionally refine them through an external script (a second AI, a linter, a custom validator). The bounce step backs up each plan, runs the script, validates YAML frontmatter integrity on the result, re-runs the plan checker, and restores the original if anything fails.

Requirements:

  • REQ-BOUNCE-01: --bounce flag or workflow.plan_bounce: true activates the step; --skip-bounce always disables it
  • REQ-BOUNCE-02: workflow.plan_bounce_script must point to a valid executable; missing script produces a warning and skips
  • REQ-BOUNCE-03: Each plan is backed up to *-PLAN.pre-bounce.md before the script runs
  • REQ-BOUNCE-04: Bounced plans with broken YAML frontmatter or that fail the plan checker are restored from backup
  • REQ-BOUNCE-05: workflow.plan_bounce_passes (default: 2) controls how many refinement passes the script receives

Configuration: workflow.plan_bounce, workflow.plan_bounce_script, workflow.plan_bounce_passes


109. External Code Review Command

Command: /msd-ship (enhanced)

Purpose: Before the manual review step in /msd-ship, automatically run an external code review command if configured. The command receives the diff and phase context via stdin and returns a JSON verdict (APPROVED or REVISE). Falls through to the existing manual review flow regardless of outcome.

Requirements:

  • REQ-EXTREVIEW-01: workflow.code_review_command must be set to a command string; null means skip
  • REQ-EXTREVIEW-02: Diff is generated against BASE_BRANCH with --stat summary included
  • REQ-EXTREVIEW-03: Review prompt is piped via stdin (never shell-interpolated)
  • REQ-EXTREVIEW-04: 120-second timeout; stderr captured on failure
  • REQ-EXTREVIEW-05: JSON output parsed for verdict, confidence, summary, issues fields

Configuration: workflow.code_review_command


110. Cross-AI Execution Delegation

Command: /msd-execute-phase N --cross-ai

Purpose: Delegate individual plans to an external AI runtime for execution. Plans with cross_ai: true in their frontmatter (or all plans when --cross-ai is used) are sent to the configured command via stdin. Successfully handled plans are removed from the normal executor queue.

Requirements:

  • REQ-CROSSAI-01: --cross-ai forces all plans through cross-AI; --no-cross-ai disables it
  • REQ-CROSSAI-02: workflow.cross_ai_execution: true and plan frontmatter cross_ai: true required for per-plan activation
  • REQ-CROSSAI-03: Task prompt is piped via stdin to prevent injection
  • REQ-CROSSAI-04: Dirty working tree produces a warning before execution
  • REQ-CROSSAI-05: On failure, user chooses: retry, skip (fall back to normal executor), or abort

Configuration: workflow.cross_ai_execution, workflow.cross_ai_command, workflow.cross_ai_timeout


111. Architectural Responsibility Mapping

Command: /msd-plan-phase (enhanced research step)

Purpose: During phase research, the phase-researcher now maps each capability to its architectural tier owner (browser, frontend server, API, CDN/static, database). The planner cross-references tasks against this map, and the plan-checker enforces tier compliance as Dimension 7c.

Requirements:

  • REQ-ARM-01: Phase researcher produces an Architectural Responsibility Map table in RESEARCH.md (Step 1.5)
  • REQ-ARM-02: Planner sanity-checks task-to-tier assignments against the map
  • REQ-ARM-03: Plan checker validates tier compliance as Dimension 7c (WARNING for general mismatches, BLOCKER for security-sensitive ones)

Produces: ## Architectural Responsibility Map section in {phase}-RESEARCH.md


112. Extract Learnings

Command: /msd-extract-learnings N

Purpose: Extract structured knowledge from completed phase artifacts. Reads PLAN.md and SUMMARY.md (required) plus VERIFICATION.md, UAT.md, and STATE.md (optional) to produce four categories of learnings: decisions, lessons, patterns, and surprises. Optionally captures each item to an external knowledge base via capture_thought tool.

Requirements:

  • REQ-LEARN-01: Requires PLAN.md and SUMMARY.md; exits with clear error if missing
  • REQ-LEARN-02: Each extracted item includes source attribution (artifact and section)
  • REQ-LEARN-03: If capture_thought tool is available, captures items with source, project, and phase metadata
  • REQ-LEARN-04: If capture_thought is unavailable, completes successfully and logs that external capture was skipped
  • REQ-LEARN-05: Running twice overwrites the previous LEARNINGS.md

Produces: {phase}-LEARNINGS.md with YAML frontmatter (phase, project, counts per category, missing_artifacts)

Optional integration — capture_thought: capture_thought is a convention, not a bundled tool. MSD does not ship one and does not require one. The workflow checks whether any MCP server in the current session exposes a tool named capture_thought and, if so, calls it once per extracted learning with the signature below. If no such tool is present, the step is skipped silently and LEARNINGS.md remains the primary output.

Expected tool signature:

capture_thought({
  category: "decision" | "lesson" | "pattern" | "surprise",
  phase: <phase_number>,
  content: <learning_text>,
  source: <artifact_name>
})

Users who run a memory / knowledge-base MCP server (for example, ExoCortex-style servers, claude-mem, or mem0-style servers) can implement this tool name to have learnings routed into their knowledge base automatically with project, phase, and source metadata. Everyone else can use /msd-extract-learnings without any extra setup — the LEARNINGS.md artifact is the feature.

With features.global_learnings: true, phase completion runs the extraction for the just-completed phase automatically and copies the artifact to the global store at ~/.msd/knowledge/ (#3683) — extraction and copy failures never block completion. With the gate off (the default), extraction stays fully manual.


114. Context-Window-Aware Prompt Thinning

Purpose: Reduce static prompt overhead by ~40% for models with context windows under 200K tokens. Extended examples and anti-pattern lists are extracted from agent definitions into reference files loaded on demand via @ required_reading.

Requirements:

  • REQ-THIN-01: When CONTEXT_WINDOW < 200000, executor and planner agent prompts omit inline examples
  • REQ-THIN-02: Extracted content lives in references/executor-examples.md and references/planner-antipatterns.md
  • REQ-THIN-03: Standard (200K-500K) and enriched (500K+) tiers are unaffected
  • REQ-THIN-04: Core rules and decision logic remain inline; only verbose examples are extracted

Reference files: executor-examples.md, planner-antipatterns.md


115. Configurable CLAUDE.md Path

Purpose: Allow projects to store their CLAUDE.md in a non-root location. The claude_md_path config key controls where /msd-profile-user and related commands write the generated CLAUDE.md file.

Requirements:

  • REQ-CMDPATH-01: claude_md_path defaults to ./.claude/CLAUDE.md (a valid project-scoped memory location; changed from ./CLAUDE.md in v1.5 per #1098 so generated content does not pollute a hand-crafted repo-root CLAUDE.md)
  • REQ-CMDPATH-02: Profile generation commands read the path from config and write to the specified location
  • REQ-CMDPATH-03: Relative paths are resolved from the project root
  • REQ-CMDPATH-04: generate-claude-md never overwrites an existing instruction file that lacks MSD section markers (a hand-crafted file) unless --force is passed

Configuration: claude_md_path


116. TDD Pipeline Mode

Purpose: Opt-in TDD (red-green-refactor) as a first-class phase execution mode. When enabled, the planner aggressively selects type: tdd for eligible tasks and the executor enforces RED/GREEN/REFACTOR gate sequence with fail-fast on unexpected GREEN before RED.

Requirements:

  • REQ-TDD-01: workflow.tdd_mode config key (boolean, default false)
  • REQ-TDD-02: When enabled, planner applies TDD heuristics from references/tdd.md to all eligible tasks (business logic, APIs, validations, algorithms, state machines)
  • REQ-TDD-03: Executor enforces gate sequence for type: tdd plans — RED commit (test(...)) must precede GREEN commit (feat(...))
  • REQ-TDD-04: Executor fails fast if tests pass unexpectedly during RED phase (feature already exists or test is wrong)
  • REQ-TDD-05: End-of-phase collaborative review checkpoint verifies gate compliance across all TDD plans (advisory, non-blocking)
  • REQ-TDD-06: Gate violations surfaced in SUMMARY.md under ## TDD Gate Compliance section

Configuration: workflow.tdd_mode Reference files: tdd.md, checkpoints.md


v1.37.0 Features

117. Spike Command

Command: /msd-spike [idea] [--quick]

Purpose: Run 2–5 focused feasibility experiments before committing to an implementation approach. Each experiment uses Given/When/Then framing, produces executable code, and returns a VALIDATED / INVALIDATED / PARTIAL verdict. Companion /msd-spike --wrap-up packages findings into a project-local skill.

Requirements:

  • REQ-SPIKE-01: Each experiment MUST produce a Given/When/Then hypothesis before any code is written
  • REQ-SPIKE-02: Each experiment MUST include working code or a minimal reproduction
  • REQ-SPIKE-03: Each experiment MUST return one of: VALIDATED, INVALIDATED, or PARTIAL verdict with evidence
  • REQ-SPIKE-04: Results MUST be stored in .planning/spikes/NNN-experiment-name/ with a README and MANIFEST.md
  • REQ-SPIKE-05: --quick flag skips intake conversation and uses the argument text as the experiment direction
  • REQ-SPIKE-06: /msd-spike --wrap-up MUST package findings into .claude/skills/spike-findings-[project]/

Produces:

Artifact Description
.planning/spikes/NNN-name/README.md Hypothesis, experiment code, verdict, and evidence
.planning/spikes/MANIFEST.md Index of all spikes with verdicts
.claude/skills/spike-findings-[project]/ Packaged findings (via /msd-spike --wrap-up)

118. Sketch Command

Command: /msd-sketch [idea] [--quick] [--text]

Purpose: Explore design directions through throwaway HTML mockups before committing to implementation. Produces 2–3 interactive variants per design question, all viewable directly in a browser with no build step. Companion /msd-sketch --wrap-up packages winning decisions into a project-local skill.

Requirements:

  • REQ-SKETCH-01: Each sketch MUST answer one specific visual design question
  • REQ-SKETCH-02: Each sketch MUST include 2–3 meaningfully different variants in a single index.html with tab navigation
  • REQ-SKETCH-03: All interactive elements (hover, click, transitions) MUST be functional
  • REQ-SKETCH-04: Sketches MUST use real-ish content, not lorem ipsum
  • REQ-SKETCH-05: A shared themes/default.css MUST provide CSS variables adapted to the agreed aesthetic
  • REQ-SKETCH-06: --quick flag skips mood intake; --text flag replaces AskUserQuestion with numbered lists for non-Claude runtimes
  • REQ-SKETCH-07: The winning variant MUST be marked in the README frontmatter and with a ★ in the HTML tab
  • REQ-SKETCH-08: /msd-sketch --wrap-up MUST package winning decisions into .claude/skills/sketch-findings-[project]/

Produces:

Artifact Description
.planning/sketches/NNN-name/index.html 2–3 interactive HTML variants
.planning/sketches/NNN-name/README.md Design question, variants, winner, what to look for
.planning/sketches/themes/default.css Shared CSS theme variables
.planning/sketches/MANIFEST.md Index of all sketches with winners
.claude/skills/sketch-findings-[project]/ Packaged decisions (via /msd-sketch --wrap-up)

119. Agent Size-Budget Enforcement

Purpose: Keep agent prompt files lean with tiered line-count limits enforced in CI. Oversized agents are caught before they bloat context windows in production.

Requirements:

  • REQ-BUDGET-01: agents/msd-*.md files are classified into three tiers: XL (≤ 1 600 lines), Large (≤ 1 000 lines), Default (≤ 500 lines)
  • REQ-BUDGET-02: Tier assignment is declared in the file's YAML frontmatter (size: xl | large | default)
  • REQ-BUDGET-03: tests/agent-size-budget.test.cjs enforces limits and fails CI on violation
  • REQ-BUDGET-04: Files without a size frontmatter key default to the Default (500-line) limit

Test file: tests/agent-size-budget.test.cjs


120. Shared Boilerplate Extraction

Purpose: Reduce duplication across agents by extracting two common boilerplate blocks into shared reference files loaded on demand. Keeps agent files within size budget and makes boilerplate updates a single-file change.

Requirements:

  • REQ-BOILER-01: Mandatory-initial-read instructions extracted to references/mandatory-initial-read.md
  • REQ-BOILER-02: Project-skills-discovery instructions extracted to references/project-skills-discovery.md
  • REQ-BOILER-03: Agents that previously inlined these blocks MUST now reference them via @ required_reading

Reference files: references/mandatory-initial-read.md, references/project-skills-discovery.md


121. Knowledge Graph Integration

Purpose: Build, query, and inspect a lightweight knowledge graph of the project in .planning/graphs/. Opt-in per project. Exposed as the /msd-graphify user-facing command and the msd-tools.cjs graphify … programmatic verb family. Complements /msd-map-codebase --query (snapshot-oriented) with a graph-oriented view of nodes and edges across commands, agents, workflows, and phases.

Requirements:

  • REQ-GRAPH-01: Opt-in via graphify.enabled: true in .planning/config.json. When disabled, /msd-graphify prints an activation hint and stops without writing.
  • REQ-GRAPH-02: Slash-command /msd-graphify exposes subcommands build, query <term>, status, diff. The programmatic CLI node msd-tools.cjs graphify … additionally exposes snapshot, which is also invoked automatically as the final step of graphify build.
  • REQ-GRAPH-03: Build runs within the configurable graphify.build_timeout (seconds); exceeding the timeout aborts cleanly without leaving a partial graph.
  • REQ-GRAPH-04: graphify.cjs falls back to graph.links when graph.edges is absent so older graph artifacts keep rendering.
  • REQ-GRAPH-05: Graphify is invoked through msd-tools.cjs graphify ... command handlers.
  • REQ-GRAPH-06: The knowledge-graph location is configurable via graphify.graph_path (issue #1825) so one umbrella-level cross-repo graph can serve multiple sibling projects; query/status/diff read the configured graph (relative to project root), with a byte-identical .planning/graphs/ default when unset.
  • REQ-GRAPH-07: status reports the resolved graph location as graph_path — the same absolute path query/diff read, after the graphify.graph_path override is applied — so a caller shelling out to the graphify CLI passes it as --graph instead of re-deriving the default location (issue #4836).

Configuration: graphify.enabled, graphify.build_timeout, graphify.graph_path Reference files: commands/msd/graphify.md, bin/lib/graphify.cjs


v1.40.0 Features

122. Skill Surface Consolidation

Purpose: Cut the eager skill-listing overhead by folding 31 micro-skills into 4 new grouped parents and 6 existing parents that absorb sub-operations as flags. Zero functional loss — every removed micro-skill's behavior survives via a flag on a consolidated parent. After consolidation, commands/msd/*.md ships 60 sub-skills (plus 6 namespace meta-skills, see #123).

Requirements:

  • REQ-CONSOLIDATE-01: Four new grouped skills replace clusters of micro-skills:
    • /msd-capture — folds add-todo (default), note (--note), add-backlog (--backlog), plant-seed (--seed), check-todos (--list)
    • /msd-phase — folds add-phase (default), insert-phase (--insert), remove-phase (--remove), edit-phase (--edit)
    • /msd-config — folds settings-advanced (--advanced), settings-integrations (--integrations), set-profile (--profile)
    • /msd-workspace — folds new-workspace (--new), list-workspaces (--list), remove-workspace (--remove)
  • REQ-CONSOLIDATE-02: Six existing parents absorb wrap-up / sub-operations as flags: /msd-update --sync, /msd-update --reapply, /msd-sketch --wrap-up, /msd-spike --wrap-up, /msd-map-codebase --fast, /msd-map-codebase --query, /msd-code-review --fix, /msd-progress --do, /msd-progress --next.
  • REQ-CONSOLIDATE-03: /msd-next is not the retired workflow-advance command; it is reserved for the state-aware smart-entry launcher. Workflow advancement remains under /msd-progress --next.
  • REQ-CONSOLIDATE-04: Deleted micro-skill slash forms (the bare msd-add-todo, msd-add-backlog, msd-plant-seed, msd-check-todos, msd-add-phase, msd-insert-phase, msd-remove-phase, msd-edit-phase, msd-new-workspace, msd-list-workspaces, msd-remove-workspace, msd-settings-advanced, msd-settings-integrations, msd-set-profile, msd-sketch-wrap-up, msd-spike-wrap-up, msd-reapply-patches, msd-code-review-fix, …) MUST resolve to "Unknown command" — no shadow stubs.
  • REQ-CONSOLIDATE-05: autonomous.md invokes /msd-code-review --fix (was previously calling the deleted msd-code-review-fix).

Reference issue: #2790


123. Namespace Meta-Skills (Two-Stage Routing)

Purpose: Replace the flat eager skill listing with a two-stage hierarchical routing layer. The model sees 6 namespace routers instead of 86 entries, selects a namespace, then routes to the sub-skill. Descriptions use pipe-separated keyword tags (≤ 60 chars) for routing density.

Commands:

  • /msd-workflow — phase pipeline router (discuss / plan / execute / verify / phase / progress / next)
  • /msd-project — project lifecycle (milestones, audits, summary)
  • /msd-quality — quality gates (code review, debug, audit, security, eval, ui)
  • /msd-context — codebase intelligence (map, graphify, docs, learnings)
  • /msd-manage — config / workspace / workstreams / thread / update / ship / inbox
  • /msd-ideate — exploration & capture (explore, sketch, spike, spec, capture)

Token cost:

Entries Approx tokens
Pre-1.40 full install 86 ~2,150
Namespace meta-skills 6 ~120

Requirements:

  • REQ-NS-01: Six commands/msd/ns-*.md namespace routers ship with pipe-separated keyword-tag descriptions (≤ 60 chars).
  • REQ-NS-02: Existing sub-skills are unchanged and still invocable directly — namespace skills are additive, not a replacement for direct slash forms.
  • REQ-NS-03: The body of each namespace router contains a routing table that maps user intent to the correct concrete sub-skill on the post-#2790 consolidated surface.
  • REQ-NS-04: Tests validate namespace files exist, include matching command requires, and reference only existing sub-skill files.

Reference issue: #2792


124. Context-Window Utilization Guard

Command: /msd-health --context

Purpose: Quality guard against context-window saturation. Two thresholds: 60 % utilization warns ("consider /msd-thread"), 70 % is critical ("reasoning quality may degrade"; matches the fracture-point per recent context-attention research).

Requirements:

  • REQ-CTX-GUARD-01: /msd-health --context prints a structured status line with current utilization, threshold tier (ok / warn / critical), and a remediation suggestion.
  • REQ-CTX-GUARD-02: The same triage is exposed as msd-tools.cjs validate context --tokens-used <int> --context-window <int> — a structured envelope for status-line and hook callers (#125). Both flags are required; the handler returns the same { percent, state } envelope as the pure classifier in REQ-CTX-GUARD-03.
  • REQ-CTX-GUARD-03: The classifier (bin/lib/context-utilization.cjs) is pure: input (tokensUsed, contextWindow), output { percent, state }. Easy to unit-test, easy to reuse from any caller.

Reference issue: #2792


125. Phase-Lifecycle Status-Line Read-Side

Purpose: Surface phase orchestration state on the status-line. parseStateMd() reads four new STATE.md frontmatter fields and formatMsdState() renders in-flight, idle, and progress scenes. Write-side wiring follows in a later RC.

Requirements:

  • REQ-LIFECYCLE-01: parseStateMd() reads four optional fields:
    • active_phase — phase number when an orchestrator is in flight
    • next_action — recommended next command when idle
    • next_phases — YAML flow array of next phase numbers
    • progress — nested total_phases / completed_phases / percent block
  • REQ-LIFECYCLE-02: formatMsdState() checks the lifecycle fields in priority order and emits the first matching scene (Phase active → Idle next-recommended → Milestone complete → Default fallback).
  • REQ-LIFECYCLE-03: All four fields default to undefined; existing STATE.md files render byte-for-byte identically.

Reference issue: #2833 — see docs/STATE-MD-LIFECYCLE.md for the full field reference and rendering rules.


v1.41.0 Features

126. Per-Phase-Type Model Selection

Purpose: Express model tuning at the phase level (planning, research, execution, verification) without learning the full agent taxonomy. Sits between per-agent model_overrides (precise, verbose) and the global model_profile tier (coarse, uniform).

Config key: models in .planning/config.json

Phase-type slots:

Slot Agents assigned
planning msd-planner, msd-roadmapper, msd-pattern-mapper
discuss msd-assumptions-analyzer
research msd-phase-researcher, msd-project-researcher, msd-research-synthesizer, msd-codebase-mapper, msd-ui-researcher
execution msd-executor, msd-debugger, msd-doc-writer
verification msd-verifier, msd-plan-checker, msd-integration-checker, msd-nyquist-auditor, msd-ui-checker, msd-ui-auditor, msd-doc-verifier, msd-code-reviewer
completion (reserved for future subagent)

Accepted values: "opus" / "sonnet" / "haiku" / "inherit"

Resolution precedence (highest → lowest):

1. model_overrides[<agent>]
2. dynamic_routing.tier_models[<tier>]   (when enabled)
3. models[<phase_type>]                  (this feature)
4. model_profile
5. Runtime default

Requirements:

  • REQ-PHASE-MODELS-01: Six named models.* slots accepted by config-schema.cjs and config-schema.ts; config-set rejects unknown phase-types.
  • REQ-PHASE-MODELS-02: Configs without a models block behave byte-for-byte identically to pre-v1.41 behavior.
  • REQ-PHASE-MODELS-03: discuss and completion are accepted by the schema for forward compatibility; setting them today is a no-op until a subagent maps to each.

Reference issue: #3023


127. Dynamic Routing with Failure-Tier Escalation

Purpose: Pay for the cheap tier by default; escalate to a more capable model automatically when the orchestrator detects a soft failure (verification inconclusive, plan-check FLAG, etc.).

Config key: dynamic_routing in .planning/config.json

Behavior:

  • enabled: false (default) — feature is off; all agents use the precedence chain unchanged.
  • enabled: true — the resolver picks tier_models[default_tier] for the first spawn and escalates one tier up on orchestrator-detected soft failure, capped by max_escalations.

Composition: model_overrides always wins; dynamic_routing.tier_models[<tier>] resolves above models.<phase_type> and model_profile.

Requirements:

  • REQ-DYNROUTE-01: dynamic_routing.enabled acts as a master switch; when false or block is absent, zero behavior change.
  • REQ-DYNROUTE-02: New resolver resolveModelForTier(cwd, agent, attempt) in core.cjs is the single call-site for orchestrator integration.
  • REQ-DYNROUTE-03: max_escalations caps the escalation chain to prevent runaway cost.

Reference issue: #3024


128. Update Banner Opt-In

Purpose: Surface update availability to users who have declined or bypassed the MSD statusline, without requiring the statusline.

Behavior:

  • At install time, if the installer detects no MSD statusline, it offers an opt-in SessionStart hook.
  • The hook reads the existing ~/.cache/msd/msd-update-check.json cache — the same cache used by the statusline — and prints a banner only when an update is available.
  • Silent when up-to-date.
  • Failure diagnostics rate-limited to once per 24 h.
  • Cleanly removed by npx @golem15/msd-core --uninstall.

Requirements:

  • REQ-BANNER-01: Banner does not install without explicit opt-in.
  • REQ-BANNER-02: No additional network requests — reuses the existing background update-check cache.
  • REQ-BANNER-03: Uninstall path removes the banner hook.

Reference issue: #2795


129. Issue-Driven Orchestration Guide

Purpose: Document a recipe for driving the full MSD workflow from a GitHub / Linear / Jira issue, mapping tracker-centric concepts onto existing MSD primitives.

Document: docs/issue-driven-orchestration.md

Covered workflow:

  1. Create an isolated workspace per issue (/msd-workspace --new)
  2. Run the manager dashboard to get oriented (/msd-manager)
  3. Execute autonomously (/msd-autonomous)
  4. Verify and review (/msd-verify-work, /msd-review)
  5. Ship and close the issue (/msd-ship)

No new commands or daemon process — purely a documentation artifact that maps existing primitives onto a tracker-driven workflow.

Reference issue: #2840


130. Graphify Commit-Based Staleness

Purpose: Surface whether the architecture graph was built from the current commit or an older one, complementing the existing mtime-based stale signal.

Command: /msd-graphify status

New fields returned (graphify v0.7+ graphs):

Field Type Description
built_at_commit string Commit SHA the graph was built from
current_commit string Current git HEAD
commits_behind number How many commits behind HEAD the graph is
commit_stale boolean | null true=stale, false=current, null=unavailable (pre-v0.7, non-git)

Rendered output (when signal is available):

Source commit: abc1234 (3 commits behind HEAD)

Security: built_at_commit validated as 4–40 hex chars before reaching git — a hostile graph.json cannot inject dashed options into argv.

Fallback: pre-v0.7 graphs and non-git checkouts return commit_stale: null; callers fall back to the existing mtime-based stale flag. No behavior change for existing users.

Reference issue: #3170


v1.42.1 Features

132. Package Legitimacy Gate

Purpose: Stop hallucinated, suspicious, or slopsquatting package names before they reach a shell install command.

Behavior:

  • Phase research writes a ## Package Legitimacy Audit table for recommended packages.
  • Packages verified only through search are treated as [ASSUMED], not trusted.
  • [SLOP] packages are removed from recommendations.
  • Plans that need [ASSUMED] or suspicious packages add a human verification checkpoint.
  • Executor install failures stop for human verification instead of auto-trying similarly named packages.

Requirements:

  • REQ-PKG-GATE-01: Research MUST record package registry, age, download/source signals, legitimacy verdict, and disposition.
  • REQ-PKG-GATE-02: Planner MUST gate unverified or suspicious package installs before execution.
  • REQ-PKG-GATE-03: Executor MUST NOT auto-substitute package names after failed package-manager installs.

Reference: v1.42.1 Release Notes


133. Skill Surface Budgeting

Purpose: Let users reduce installed skill and agent surface area when context budget matters.

Install profiles:

Profile Purpose
core Minimal main-loop surface
standard Core plus common phase-management commands
full Complete surface; default

Runtime control: /msd-surface lists profile state and enables, disables, or resets skill clusters without reinstalling.

Requirements:

  • REQ-SURFACE-01: Installer MUST resolve --profile=<name> and persist the active profile in .msd-profile.
  • REQ-SURFACE-02: --minimal and --core-only MUST remain aliases for --profile=core.
  • REQ-SURFACE-03: Runtime surface state MUST persist outside the install profile marker.

Reference: ADR-0011


134. Installer Migrations

Purpose: Make runtime config cleanup explicit, auditable, and rollback-aware during installs and updates.

Capabilities:

  • First-time baseline migration records managed files.
  • Legacy stale-file cleanup uses ownership evidence before deleting or rewriting.
  • User-owned artifacts are preserved.
  • Ambiguous MSD-looking files block with a clear report instead of being silently overwritten.
  • Migration plans support dry-run reporting and rollback protection.

Requirements:

  • REQ-INSTALL-MIGRATION-01: Migration records MUST include metadata, install scope, and ownership evidence.
  • REQ-INSTALL-MIGRATION-02: Destructive actions MUST fail closed when ownership is ambiguous.
  • REQ-INSTALL-MIGRATION-03: Install failures MUST restore the pre-install state when rollback data exists.

Reference: Installer Migrations


135. Custom Ship PR Body Sections

Command: /msd-ship

Config key: ship.pr_body_sections

Purpose: Add project-specific PRD-style sections to generated PR bodies without editing MSD workflow files.

Behavior: Configured sections append after the required Summary, Changes, Requirements Addressed, Verification, and Key Decisions sections. They can copy from artifact headings, render templates, or fall back to static text.

Requirements:

  • REQ-SHIP-SECTIONS-01: Custom sections MUST NOT replace, remove, or reorder required PR sections.
  • REQ-SHIP-SECTIONS-02: Unknown template tokens MUST be rejected by config validation.
  • REQ-SHIP-SECTIONS-03: Disabled sections MUST stay in config without appearing in PR output.

Reference: Custom PR Body Sections


136. Review Default Reviewers

Command: /msd-review

Config key: review.default_reviewers

Purpose: Let teams choose the default reviewer subset for no-flag /msd-review runs.

Precedence:

explicit reviewer flags -> --all -> review.default_reviewers -> all detected reviewers

Requirements:

  • REQ-REVIEW-DEFAULTS-01: Missing review.default_reviewers MUST preserve the previous all-detected behavior.
  • REQ-REVIEW-DEFAULTS-02: Empty arrays MUST be rejected; remove the key to restore all-detected behavior.
  • REQ-REVIEW-DEFAULTS-03: Known but unavailable reviewers MUST be skipped with diagnostics rather than hard-failing the run.

Reference: Configuration Reference


137. Fallow Structural Review Pre-Pass

Command: /msd-code-review

Config keys: code_quality.fallow.*

Purpose: Add an optional structural analysis pass before the agent review.

Behavior: When enabled, MSD resolves a fallow binary, runs a bounded audit, writes FALLOW.json, and embeds structural findings in REVIEW.md.

Requirements:

  • REQ-FALLOW-01: Fallow MUST be opt-in and disabled by default.
  • REQ-FALLOW-02: Missing or failing fallow runs MUST produce clear diagnostics.
  • REQ-FALLOW-03: Findings larger than the embed budget MUST be skipped with a warning, preserving the raw JSON artifact.

Reference: Configuration Reference


138. End-of-Phase Human Verification Mode

Config key: workflow.human_verify_mode

Purpose: Reduce mid-flight human checkpoint interruptions while preserving human verification requirements.

Behavior: The default "end-of-phase" mode embeds human checks into <verify><human-check> blocks for phase review. "mid-flight" restores blocking checkpoint:human-verify tasks.

Requirements:

  • REQ-HUMAN-VERIFY-01: checkpoint:decision and checkpoint:human-action MUST remain blocking regardless of mode.
  • REQ-HUMAN-VERIFY-02: Human-needed verification MUST remain pending until the end-of-phase review resolves it.
  • REQ-HUMAN-VERIFY-03: Configs without the key MUST use "end-of-phase".

Reference: Checkpoints Reference


139. Quota and Rate-Limit Failure Classification

Command: /msd-execute-phase

Purpose: Treat provider quota and rate-limit failures as wait-and-resume conditions, not normal executor failures.

Behavior: Agent output is classified for signals such as 429, rate limit, usage limit, RESOURCE_EXHAUSTED, and usage_limit_reached. Matching failures present a wait-for-reset recovery path.

Requirements:

  • REQ-QUOTA-01: Quota failures MUST NOT offer immediate retry as the primary recovery.
  • REQ-QUOTA-02: Classification MUST cover Claude, Codex, and generic provider sentinels.
  • REQ-QUOTA-03: Non-quota failures MUST continue through the normal execution failure path.

Reference: Provider Rate Limit Signals


140. Statusline Context Position

Config key: statusline.context_position

Purpose: Keep the context meter visible in narrow terminals.

Options:

Value Behavior
"end" Default; render context meter near the line tail
"front" Render context meter immediately after the model name

Requirements:

  • REQ-STATUSLINE-POS-01: Invalid values MUST be rejected by config validation.
  • REQ-STATUSLINE-POS-02: Missing config MUST preserve existing end-position rendering.

Reference: Configuration Reference


141. Milestone Tag Creation Toggle

Command: /msd-complete-milestone

Config key: git.create_tag

Purpose: Let projects with external release automation complete milestones without creating local git tags.

Behavior: git.create_tag: false skips milestone tag creation. The workflow still updates milestone artifacts and state.

Requirements:

  • REQ-MILESTONE-TAG-01: Missing config MUST preserve automatic tag creation.
  • REQ-MILESTONE-TAG-02: Existing tag collisions MUST fail clearly instead of overwriting tags.
  • REQ-MILESTONE-TAG-03: Disabling tag creation MUST NOT skip milestone archival.

Reference: Configuration Reference


142. Structured JSON Error Mode

CLI: msd-tools --json-errors

Purpose: Give automation callers stable machine-readable error envelopes.

Behavior: Commands that fail under --json-errors return structured ok: false payloads with error kind, message, command context, and exit mapping instead of prose-only stderr.

Requirements:

  • REQ-JSON-ERRORS-01: Unknown commands, validation errors, timeouts, native failures, fallback failures, and internal errors MUST map to canonical error kinds.
  • REQ-JSON-ERRORS-02: CLI exit code mapping MUST remain stable for automation callers.
  • REQ-JSON-ERRORS-03: Human-readable output MUST remain the default when --json-errors is absent.

Reference: JSON Error Mode


143. UAT-Passed Predicate

CLI: node msd-tools.cjs phase uat-passed <N> [--require-verification]

Purpose: Provide a runtime-neutral, automatable predicate that evaluates HUMAN-UAT results for a phase and returns a structured pass/fail verdict with full diagnostic detail.

Behavior: Locates *-UAT.md and optionally *-VERIFICATION.md files for the given phase, parses UAT test blocks (heading-block parser, column-0 result lines) with a markdown-aware stripper that removes false-positive contexts (YAML frontmatter, fenced code blocks, HTML comments, and blockquotes). Returns passed: true only when at least one check exists AND all checks pass AND no blockers — fail-closed, no vacuous pass. The --require-verification flag requires at least one *-VERIFICATION.md with an allowlisted passing status; the command fails without one.

Output envelope: { passed, uat_files[], verification_files[], checks[], blockers[], no_uat_artifacts, policy: { require_verification } }

Field Type Description
passed boolean true only when ≥1 check exists AND all passing AND no blockers
uat_files string[] Filenames of *-UAT.md files evaluated
verification_files string[] Filenames of *-VERIFICATION.md files evaluated
checks[] { file, test, name, result, passing }[] Per-item results from heading blocks
blockers[] string[] Human-readable failure reasons (frontmatter, failing/missing items, policy, malformed markdown)
no_uat_artifacts boolean true when no test items were parsed; passed is always false when true
policy.require_verification boolean Whether --require-verification was active

Requirements:

  • REQ-UAT-PRED-01: The predicate MUST ignore result lines inside YAML frontmatter, fenced code blocks, HTML comments, and blockquotes.
  • REQ-UAT-PRED-02: passed: true MUST require at least one check AND all checks passing AND no blockers (fail-closed, no vacuous pass).
  • REQ-UAT-PRED-03: --require-verification MUST cause the command to fail when no *-VERIFICATION.md file with an allowlisted passing status is found.
  • REQ-UAT-PRED-04: blockers[] contains all human-readable failure reasons including frontmatter issues, policy violations, and malformed markdown — NOT limited to a subset of checks[].
  • REQ-UAT-PRED-05: The module MUST be runtime-neutral (no runtime-specific env checks or exit shortcuts).
  • REQ-UAT-PRED-06: A heading block with no column-0 result: line emits result:'missing' (blocker); test items are never silently dropped.

Reference: Phase Management Commands


144. Spec-Phase Edge-Completeness Probe

Command: /msd-spec-phase

Purpose: Surface the omitted domain-boundary edges that silently invalidate a requirement — touching intervals, empty inputs, rounding ties, grapheme truncation — before they become production defects. Runs as Step 5.5 of spec-phase, after the ambiguity gate.

Behavior: For each SPEC requirement the probe classifies its data/behavior shape, then raises only the applicable categories from a closed 8-category taxonomy (boundary, adjacency, empty, encoding, ordering, precision, idempotency, concurrency) via a relevance filter. Each raised category proposes one concrete candidate edge, which the author resolves to exactly one of four states:

State Meaning Downstream effect
covered An acceptance criterion handles the edge Pass/fail line written into the SPEC Acceptance Criteria block; lifted into plan-phase must_haves.truths
dismissed The edge cannot occur (requires a non-empty reason) Recorded with its reason; empty dismissals are rejected
backstop Intent recorded, needs a held-out/property-based test Lifted into must_haves.truths as a non-inferable check
unresolved Deferred Soft-gates the spec; row stamped ⚠ Edge unresolved — planner must treat as assumption

When a requirement's prose matches no shape cue, the probe does not silently drop it (#1110): it emits a single unclassified — review manually candidate so the zero-cue requirement is surfaced for the author to resolve like any other (specify / dismiss-with-reason / defer) — a manual-review nudge, not a hard block.

Non-English projects: the probe reads English via text_en, the SPEC does not have to (#2773, durable fix #3717). The shape cues are English word-boundary patterns, so a project running with response_language set would otherwise have every requirement match nothing, classify to zero shapes, and land in unclassified — the taxonomy silently contributing nothing to exactly the kind of spec it exists to harden. spec-phase Step 5.5 therefore populates an optional text_en field alongside each requirement's text with a faithful English translation: text_en is engine input, never user-facing output, so it is translated while text keeps the requirement's own wording (the SPEC stays in the original language), requirement ids are left untouched, and any acceptance criteria written back from the resolved edges return to response_language. Translation makes the classifier applicable; it does not make it omniscient. A requirement carrying no shape cue in any language still classifies to zero — that is the classifier's recorded recall gap (ADR-857 §98), not a translation failure — and the remedy there is the same one an English project uses: author an explicit shapes array on the requirement instead of relying on prose classification.

The resolved edges populate a ## Edge Coverage section in SPEC.md. Unresolved applicable edges trigger a soft gate (Resolve / Write-anyway-flagged / Keep-probing) rather than a hard block. Under --auto, the probe never auto-dismisses — it auto-covers where a defensible criterion exists, otherwise auto-backstops, and logs [auto] edge coverage: C covered, B backstop, U unresolved. The one exception is an unclassified candidate: --auto leaves it unresolved (surfaced as a flagged assumption), never auto-backstop — a missing shape is not evidence an edge exists, so minting a held-out edge obligation would be a false claim.

The load-bearing wire is the plan-phase lift: covered and backstop edges become must_haves.truths the verifier can check, so the section is not merely documentation. A backstop edge is lifted as a structured non-inferable marker ({ statement, verification: backstop }, a flat scalar — not a prose note), which the honest verifier then consumes (see below) — closing the loop the edge-probe opened.

Honest verifier — abstention on non-inferable checks (#1154). A non-inferable (backstop) truth is one whose correct behavior is not derivable from the spec alone, so the verifier cannot self-detect the gap and would confidently false-pass it (~100% of the time). At verify time, a backstop truth the verifier cannot confirm with explicit evidence (a passing wired held-out/property-based test, or a directly-observed behavior) abstains → human_needed with reason insufficient_spec (reported as unverified — held-out test recommended), never a silent passed. This is the verify-time, truth-axis mirror of the prohibition judgment-tier disposition (ADR-550 D4): exogenous (driven by the backstop tag, never a self-judged "abstain if unsure"), routing-not-diagnosis (the held-out test carries the omitted rule), and capable-tier dependent (reliable on sonnet+; the budget haiku tier degrades toward current behavior). An inferable truth is never abstained (the over-abstention guard). Reference: Honest Verifier.

Requirements:

  • REQ-EDGE-01: The edge pass MUST run after the ambiguity gate and emit a ## Edge Coverage SPEC section.
  • REQ-EDGE-02: The relevance filter MUST raise only applicable categories; each raised edge resolves to exactly one of covered/dismissed/backstop/unresolved.
  • REQ-EDGE-03: A dismissed resolution MUST require a non-empty reason.
  • REQ-EDGE-04: An unresolved applicable edge MUST trigger the soft gate; write-anyway stamps the row as a planner assumption.
  • REQ-EDGE-05: --auto MUST never auto-dismiss — auto-cover or auto-backstop only.
  • REQ-EDGE-06: plan-phase MUST lift covered criteria and backstop notes into must_haves.truths.
  • REQ-EDGE-07: A requirement whose prose matches no shape cue MUST surface an unclassified — review manually candidate (never silently dropped); --auto MUST leave it unresolved, never auto-backstop.
  • REQ-EDGE-08: plan-phase MUST lift a backstop edge into must_haves.truths as a structured flat-scalar marker ({ statement, verification: backstop }), never a prose parenthetical.
  • REQ-HONEST-01: At verify time a backstop truth that cannot be confirmed with explicit evidence MUST abstain → human_needed (reason insufficient_spec), never passed; an inferable truth MUST never be abstained (over-abstention guard); abstention MUST be exogenous (driven by the backstop tag, not self-judgment).

Reference: Edge Probe


v1.43.0 Features

145. MemPalace Memory Capability

Purpose: Opt-in cross-session and cross-project memory via the MemPalace external service (local-first, MCP + CLI). Wires deliberate recall before discuss/plan and verbatim capture + temporal-KG sync at phase boundaries through the ADR-857 capability mechanism. Default-resilient: disabled by default, every hook is onError: skip, and an absent MemPalace installation leaves the loop unchanged.

Commands: /msd-mempalace-recall, /msd-mempalace-capture

Requirements:

  • REQ-MP-01: Opt-in via mempalace.enabled: true. Default false — the loop is unchanged when unset.
  • REQ-MP-02: At plan:pre, skill mempalace-recall produces MEMORY-RECALL.md from prior decisions, patterns, and surprises retrieved via wake-up + semantic search + KG timeline. When MemPalace is unreachable, writes an "unavailable" stub and continues.
  • REQ-MP-03: At discuss:post, plan:post, and verify:post, skill mempalace-capture files the phase artifact verbatim into the appropriate MemPalace room (decisions, planning, milestones). Capture is idempotent via mempalace_check_duplicate.
  • REQ-MP-04: At ship:post, agent msd-mempalace-curator writes a diary entry, proposes cross-project tunnels (when mempalace.cross_project_tunnels: true), and runs wing-scoped sync pruning.
  • REQ-MP-05: mempalace.memory_mode has three wired values: augment (default — palace is an additive recall layer alongside MSD native memory, which stays authoritative), kg_backend (knowledge-graph queries resolve against the palace's temporal KG as the primary source, .planning/graphs/ as fallback; non-KG drawer recall stays additive), replace (recall resolves through the palace as the source of truth, native memory as fallback). Every mode is onError:skip and default-resilient — an unreachable palace degrades to native memory and MSD keeps writing .planning/graphs/, so no mode loses memory. Cross-mode migration of existing .planning/graphs/ into the palace is out of scope (not yet implemented).
  • REQ-MP-06: Every hook is onError: skip. No hook carries blocking: true. Memory never halts or fails a phase.
  • REQ-MP-07: Interactive runs prefer MCP tools; headless/cron runs prefer the MemPalace CLI (mempalace wake-up, mempalace search, mempalace mine, mempalace sync).
  • REQ-MP-08: mempalace.auto_capture_hooks is forward-declared and not yet functional. No native Claude Code hooks (stop, precompact, session-start) are installed by this key; the capability's hooks array is empty. This key is reserved for the future "Connected Capability" phase. Default false.

Configuration: mempalace.enabled, mempalace.memory_mode, mempalace.wing, mempalace.recall_on_discuss, mempalace.recall_on_plan, mempalace.capture_artifacts, mempalace.mirror_kg, mempalace.cross_project_tunnels, mempalace.diary_journal, mempalace.auto_capture_hooks

See Configuration Reference for full schema and How to enable cross-session memory with MemPalace for a setup walkthrough.


146. Spec-Phase Prohibition Probe

Command: /msd-spec-phase

Purpose: Surface the unwritten must-NOT constraints — the values/safety/ethics interpretations a feature could silently become that the author would never want but the spec does not forbid — before any code is written. The edge probe reaches data-shape edges; it structurally cannot reach prohibitions. This is the missing instrument, running as Step 5.6 of spec-phase, after the edge probe.

Behavior: A two-stage, prose-orchestrated pass per requirement (no compiled recall engine — recall is inherently model-driven, ADR-550 D7b):

  1. Recall (adversarial probe): "What could this feature silently become that the author would NOT want, but the spec does not forbid?" — model-robust open-vocabulary elicitation across values/safety/ethics.
  2. Precision (one-pass classifier): drop routine-engineering items, keep genuine values/safety/ethics prohibitions — collapses the raw list to the load-bearing few.

Each surfaced prohibition is resolved to exactly one of three states:

State Meaning Downstream effect
resolved Confirmed a real must-NOT NEGATIVE acceptance criterion written into the SPEC ## Prohibitions (must-NOT) section; lifted into plan-phase must_haves.prohibitions (its own sibling block, never truths)
dismissed Not a genuine prohibition (requires a non-empty reason) Recorded with its reason; empty dismissals are rejected
unresolved Deferred Soft-gates the spec; surfaced as a planner assumption

Each resolved prohibition carries a verification tier — test (a negative test can enforce it) or judgment (only human/LLM judgment can). At verify time, judgment-tier prohibitions route to a never-silent / never-hard-halt soft gate (autonomous emits an unverified-prohibition — human review recommended flag); test-tier prohibitions are enforced via the deterministic check prohibition-enforcement gate — green when the wired negative test / lint rule passes, hard-gate (flagged, non-green) when missing or failing, in both interactive and autonomous modes (#1259, ADR-550 D5d). Under --auto, the probe never auto-dismisses. Canon-bound concerns (OWASP / GDPR / fairness) are referred to /msd-secure-phase rather than minting SPEC prohibitions (ADR-550 D6).

The load-bearing wire is the plan-phase lift into must_haves.prohibitions, so the section is not merely documentation.

Deterministic prohibition-check descriptor source (#1278). A resolved test-tier prohibition MAY carry an optional check descriptor — the flat-scalar keys check_kind (node-test | lint-rule), check_target, and check_rule (lint-rule only) — authored at spec-phase. projectProhibitions projects these scalars deterministically and verify-phase reads them back to locate the check handed to check prohibition-enforcement, so a wired, passing test closes the gap with zero manual descriptor authoring (previously the verify-phase LLM had to invent {kind, target, rule} each run, #1259). The descriptor is optional and backward-compatible — a descriptor-less prohibition parses and disposes byte-identically to today — and fail-closed: a partial, invalid, or absent descriptor falls through to the producer's existing fail-closed locate, never a silent green. failFirst stays a verify-time caller attestation (machine-proven fail-first is tracked in #1279).

Requirements:

  • REQ-PROHIB-01: The prohibition pass MUST run after the edge probe and emit a ## Prohibitions (must-NOT) SPEC section.
  • REQ-PROHIB-02: Stage 1 MUST ask the adversarial recall question; Stage 2 MUST drop routine-engineering items and keep values/safety/ethics prohibitions.
  • REQ-PROHIB-03: A dismissed resolution MUST require a non-empty reason.
  • REQ-PROHIB-04: --auto MUST never auto-dismiss.
  • REQ-PROHIB-05: plan-phase MUST lift resolved prohibitions into must_haves.prohibitions (never truths).
  • REQ-PROHIB-06: A well-formed but unwired test-tier prohibition MUST fail closed at verify time — never a silent pass.
  • REQ-PROHIB-07: A test-tier prohibition with a machine-proven-fail-first, genuinely-passing (non-vacuous) wired mechanical check (a node --test negative test OR a lint/AST rule) MUST dispose green and be satisfiable; a missing, un-provable, or non-passing check MUST hard-gate (flagged, non-green) in both interactive and autonomous modes. Fail-first is machine-proven, not caller-attested (#1279, ADR-550 D5d): before a clean pass greens, the producer independently runs the wired check against a known violation (the descriptor's violationFixture) and confirms it goes RED — a lint rule via the violating fixture, a node test via the violating subject injected through the MSD_PROHIB_SUBJECT convention; absent a violation source it fails closed, never falling back to attestation. (Enforcement half shipped #1259; deterministic descriptor auto-locate in #1278.)

Reference: Prohibition Probe


147. Capability Management Command

Command: msd capability install | update | remove | list | outdated | disable | enable

Purpose: The user-facing CLI for the ADR-1244 capability ecosystem — install, upgrade, remove, list, check for updates, and toggle MSD capabilities (first-party and third-party overlays) from a registry / git / npm / tarball / local source. Wires the Phase-3/4 lifecycle library (source resolver, install ledger, trust gate) to a command users actually run.

Behavior:

  • install <spec> [--integrity sha512-…] [--scope global|project] [--yes] [--shared-file <rel>]… — resolve (copy-only) → verify integrity / SHA pin → engines.msd gate → disclose executable surfaces → consent (--yes grants; without it an executable install aborts after printing the disclosure and writes nothing) → validate → extract → record the ledger.
  • update [<id> | --all] [--scope] [--yes] — re-resolve the capability's recorded source and upgrade via atomic stage-then-swap; re-consent when the executable set changed; --all reports a per-capability outcome and exits non-zero on any partial failure.
  • remove <id> [--purge-data] [--scope] — strip the ledger-recorded files + marker-isolated shared edits; first-party capabilities are rejected (use the product uninstaller).
  • list [--json] — first-party + installed overlay capabilities (both scopes) as a JSON array.
  • outdated [--json] [--scope] — light remote peek of each installed overlay's recorded source (ADR-1244 D6 per-source matrix: git ls-remote --tags, npm view … version resolving the highest version matching the recorded range, local re-read; tarball → manual, registry → unknown) reporting outdated / current / pinned / manual / unknown per capability. A source pinned to an immutable ref (git #sha: or #tag:, or an exact npm version) is reported pinned. A bare git #<ref> is classified at the remote: if it resolves exclusively under refs/tags/ it is an immutable tag → pinned; if it resolves to a mutable branch (or is ambiguous) it is unknown. Bounded subprocesses (git ≤30s, npm ≤60s) and a failing peek degrades that row to unknown without crashing the command. --json for machine output, default for a table.
  • disable | enable <id> — toggle activation state (equivalent to msd capability set <id> --off / --on).

Trust boundary: install never executes capability code (copy-only staging); executable surfaces require explicit consent; sources are gated by the project-scoped capabilities.strict_known_registries policy (fail-closed on a malformed/unparseable value); every shared-config write/delete is realpath-confined to the scope root, and a name collision with a user's mcpServers entry is never clobbered.

Reference: msd capability command reference · ADR-1244


148. Smart Entry Launcher

Command: /msd-next

Tool: msd-tools smart-entry [--json]

Purpose: Provide a state-aware front door that reads project/workflow state, classifies the user's situation, presents a short menu, and dispatches exactly one existing MSD command.

Requirements:

  • REQ-SMART-ENTRY-01: Detection MUST be read-only and deterministic; classification lives in msd-tools smart-entry.
  • REQ-SMART-ENTRY-02: The launcher MUST never perform project work directly; it only displays a menu and dispatches one command.
  • REQ-SMART-ENTRY-03: The workflow MUST fall back to /msd-progress if detection fails.
  • REQ-SMART-ENTRY-04: Each classified situation MUST provide exactly one recommended action and valid slash commands.
  • REQ-SMART-ENTRY-05: Text-mode runtimes MUST receive a numbered-list fallback instead of being stranded by interactive UI assumptions.

Situations: no project, paused, blocked, verify failed, needs first phase, planning, executing, verify pending, idle stranded, complete, unknown.

Reference: Smart Entry Design


v1.7.0 Features

These are features new to @golem15/msd-core 1.7.0 (the current release line: 1.0.0 → 1.2.0 → … → 1.6.1 → 1.7.0). The preceding v1.27–v1.43.0 sections use the retired get-shit-done-cc / get-shit-done-redux feature numbering and are not msd-core releases — see Legacy Release Notes.

149. Embeddable Orchestration System (Host-Integration Interface)

Purpose: Express every host integration against one public, versioned contract (ADR-1239 Phase A, #1690) instead of bespoke per-host wiring, so onboarding a new host becomes additive descriptor work.

Behavior: The interface exposes six interface points (command, dispatch, model, hooks, state, artifact), eight negotiated axes, and a PROTOCOL_VERSION handshake that negotiates down to min(host, engine). In 1.7.0, 14 runtimes were migrated onto the interface via imperative adapters (OpenCode #2087, Cursor #2089, Cline #2090, Hermes #2091, Qwen #2092, Kilo #2093, Trae #2094, Kimi #2095, Antigravity #2096, Augment #2097), a declarative adapter (Codex #2088), plus full lifecycle-hook wiring for CodeBuddy (#2098), GitHub Copilot (#2099), and Windsurf (#2100). Descriptors gained an extensionEvents vocabulary (#1946), and /msd-surface now reproduces a runtime's agent output byte-for-byte from the installer's descriptors (#1575).

New runtimes: ZCode (Z.ai — Agentic Development Environment for GLM-5.2, #1925) and a repo-local VS Code extension driven through the adapter (#2103). The retired Gemini CLI now redirects to Antigravity CLI, its official successor (#1928).

Reference: The Embeddable Orchestration System · Host-Integration Interface · Interface versioning policy


150. Discoverability Registries

Purpose: Two non-endorsing catalogs for third-party extensions (#2182).

Behavior: The Community Capability Registry (#2188) lists third-party Feature Capabilities installed with msd capability install; the EoS Registry (#2193) lists third-party host integrations built on the ADR-1239 interface. Every entry embeds a live release badge and links to a GitHub Discussion. Registration is a documentation PR, regenerated with npm run gen:registry.

Reference: MSD Registries


151. Companion MCP Server

Command: msd-mcp-server

Purpose: A companion MCP server exposing MSD over stdio JSON-RPC 2.0, covering interface points 1 and 5 (#1681).

Behavior: OpenCode installs auto-register it as mcp.msd (#1682). OpenCode also gained the opencode-subset hook dialect plus session.idle handling (#1682) and now runs MSD's lifecycle safety hooks — prompt-injection guard, read-before-edit guard, and injection scanner (#1923).


152. Statusline Token Count & Git Segment

Purpose: Opt-in statusline additions surfacing more session context.

Behavior: An absolute token count on the context meter (#2161) and a git branch + working-state segment (#2163), both opt-in. A companion opt-in compact MSD-state format condenses the MSD state segment (#2162).

Configuration: statusline.*


153. Model Catalog Advances

Purpose: Refresh the default model tiers and how models are surfaced.

Behavior: Codex/OpenAI defaults advance to the GPT-5.6 family (Sol / Terra / Luna) (#2122); the verbose (1M context) model suffix collapses to a compact (1M) badge (#2160). MSD warns when model config changes without re-running the installer on static-frontmatter runtimes such as Codex and OpenCode (#1688).

Reference: Configuration · Configure model profiles


154. Claude Orchestration Capability (BETA)

Purpose: A default-off, BETA, Claude-only capability that adopts Claude Code's Workflow tool for parallel sub-agent orchestration (#1143).

Reference: The Claude orchestration capability


155. External-Job Capability

Purpose: A default-off capability that externalizes long-running compute as asynchronous external jobs, e.g. SLURM submission (#1165).

Configuration: external_job.submit_timeout_ms, external_job.poll_timeout_ms, external_job.artifact_dir (#1164)


156. API-Coverage Gate

Command: /msd-verify-work

Purpose: A phase that integrates an external API, SDK, or service can no longer seal verification without a decided coverage matrix (#1562).

Behavior: At seal time the gate reads the phase scope — the plan bodies, falling back to this phase's ROADMAP section — and runs the deterministic detector over it. An integration signal without a COVERAGE.md matrix blocks the seal; no signal passes.

Unestablished scope is not a negative verdict (#3909). A phase with no plan body and no roadmap section gives the detector nothing to examine. The gate used to run detection over zero bytes and pass, certifying "no external-API integration" from a probe that never looked. It now holds the seal instead, reporting scope_unavailable: true. A phase whose plans are real and simply contain no API vocabulary is unaffected — the discriminator is bytes examined, never signals found.

Breaking change: a phase that previously sealed because its detector could not establish a scope is now correctly held. Add the phase plan, or record a reasoned No external API integration: <reason> declaration in COVERAGE.md. See Resolve a skipped capability probe.


157. State Rebuild & Configurable Graph Path

Behavior: A new msd-tools state rebuild subcommand re-derives STATE.md from source (#1830). The new graphify.graph_path setting makes the knowledge-graph location configurable, so a single umbrella graph can serve several projects (#1825).


158. Broken-Windows Ledger

Behavior: A cross-phase defect register at .planning/WINDOWS.md accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths (#1950). /msd-ship blocks while any entry is open; an entry can be waived only with a recorded reason (auditable) or marked fixed (removed from the blocking set). /msd-progress surfaces the open + waived counts.

Commands: msd-tools windows status | append | waive | fixed.

Config: workflow.windows_enforce (gate active, default false — opt-in enforcement). Enable with msd config-set workflow.windows_enforce true. Tracking (the ledger itself, populated by the executor) is always on; only the ship gate is opt-in.

Backward compatibility: A project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly; the gate only activates once windows are recorded.

Milestone attribution (#4487): each entry carries a milestone field, stamped at record time from the workstream's resolved milestone version (STATE.md milestone: frontmatter, or the ROADMAP.md in-progress marker as a fallback). Phase numbers are unique only within one active phases/ directory — milestone complete frees them for reuse — so this is what lets an entry be attributed to the milestone it was actually recorded under, even after that milestone is archived and its phase numbers reused. null when no milestone could be resolved, including every entry recorded before this field existed.

Configuration: graphify.graph_path


159. Complexity-Triggered Refactor

Behavior: An execute:post step measures the complexity of the files a phase touched (decision-point counting over comment- and literal-stripped source, no external dependency) and surfaces a scoped refactor proposal at .planning/phases/<N>/<NN>-REFACTOR.md when a function's score exceeds refactor.complexity_threshold or its growth over its recorded anchor exceeds refactor.complexity_jump_delta — whichever trips first, both reported. Trigger semantics are strictly greater (ESLint's complexity: {max: N} convention), so a score exactly equal to the threshold does not trigger. The anchor is set the first time a function is observed and moves only when the proposal is dispositioned via refactor accept or refactor decline — never when the score alone improves — so the jump delta is cumulative growth since the last conscious decision about that function, not the change made in a single phase. Advisory by default: the proposal is informational only, never edits code, and never blocks. Opt-in refactor.trigger_strict records an untriaged proposal as an open deviation entry in the broken-windows ledger (#1950) instead — it does not block on its own; ship-blocking is broken-windows' existing ship:pre gate, enabled separately with workflow.windows_enforce. Without broken-windows installed, strict mode still records the proposal locally and says so. Enabling refactor.trigger_strict without workflow.windows_enforce also on (or with broken-windows absent) surfaces a typed refactor_strict_not_enforcing warning on every triggering evaluate, naming the exact remediation, so this enforcement gap is never silent. A declined proposal resolves its ledger entry as waived with the recorded reason; an accepted one resolves as fixed. The metric is approximate by construction: biased against a flat switch, blind to nesting depth, JS/TS-family only, and a renamed function loses its anchor (issue #1953).

Commands: msd-tools refactor evaluate | status | accept | decline.

Config: refactor.trigger_enabled (master gate, default false), refactor.complexity_threshold (default 15), refactor.complexity_jump_delta (default 5), refactor.trigger_strict (default false). See Configuration Reference.

Backward compatibility: Off by default. When refactor.trigger_enabled is false the hook never runs and writes nothing; a project that never enables it is completely unaffected.


160. Archive Quick Tasks at Milestone Close

Command: /msd-complete-milestone (forward path), /msd-cleanup (retroactive path), msd-tools milestone complete --archive-quick / msd-tools milestone archive-quick <version> (#2142)

Behavior: .planning/quick/ otherwise accumulates one directory per /msd-quick task forever. /msd-complete-milestone now offers a Yes/Skip prompt — when accepted, it moves every directory under .planning/quick/ into .planning/milestones/<version>-quick/, (re)writes that archive directory's README.md (an index built by scanning the archive directory, one entry per task, linked to its SUMMARY.md when one exists), and clears the data rows of STATE.md's ### Quick Tasks Completed table while preserving its header and detected column variant. /msd-cleanup offers the same archival retroactively, for milestones that were already closed before their quick tasks were swept, via the narrower milestone archive-quick <version> command — identical move/index/reset behavior, but without touching ROADMAP.md, REQUIREMENTS.md, MILESTONES.md, or milestone-completion guards, so it can be re-run safely against an already-completed milestone.

Why opt-in. Phase-directory archival is default-ON (#1871) — omitting a phase directory from an archive would silently leave stale execution history in the way of the next milestone's roadmap. Quick tasks carry no such downstream conflict, so archival here defaults OFF: a user who never passes --archive-quick sees zero behavior change. This is a deliberate asymmetry with phase archival, not an oversight.

Why bucket-all, not per-milestone. .planning/quick/ is a flat directory with no on-disk record of which milestone a given task belongs to. Splitting tasks per milestone was considered and rejected — inferring provenance from dates (creation time vs. a milestone's shipped date) is a proxy, not a fact, and a wrong inference on a one-way mv is silently irreversible. Archival instead buckets everything currently in .planning/quick/ into the one milestone being completed (or, on the retroactive path, the one milestone chosen), and says so in the confirmation prompt.

Why the index is built from disk, not from STATE.md's table. The ### Quick Tasks Completed table is a running log a workflow step appends to — it demonstrably drifts from what's actually in .planning/quick/ (the motivating case: 53 rows against 49 directories, ~22 rows pointing at directories that no longer existed, 18 directories with no row at all). Building the archive's README.md index by scanning the archive directory itself, rather than trusting the table, means the index can never inherit that drift; a re-run's index also naturally includes entries a prior run already archived, since it's re-derived from what's physically present.

Known limits:

  • No per-milestone provenance — bucket-all is the only option (see above).
  • A ### Quick Tasks Completed table whose columns match neither registered variant (with/without a Status column) is left untouched with a warning rather than reset, since clearing it would risk destroying rows under a schema MSD doesn't recognize.
  • A STATE.md with no ### Quick Tasks Completed section at all is a normal, silent no-op for the reset step — the section is created lazily by /msd-quick, not present in the project template.

See Archiving quick tasks for the full walkthrough.


161. Verify-Command Path Grounding

Command: /msd-plan-phase (automatic), msd-tools check verify-command-paths <N> (#2401)

Behavior: A planner authoring a per-task <automated> verify command has no line of sight to whether the path it just wrote actually resolves, and msd-plan-checker had no deterministic way to check — so it hand-reasoned the filesystem and, in the motivating case, prescribed two successively-wrong replacement paths (the second citing a package.json that did not exist). Two changes close that:

  1. Prior-command inheritance. The nearest prior phase's <automated> commands are surfaced to the planner as prior_verify_commands, at every context window. Cross-phase enrichment was previously gated on context_window >= 500000; at 200k the planner re-invented the command and got it wrong. This payload is a handful of one-liners, so it is never gated.
  2. A deterministic probe. msd-tools check verify-command-paths <N> resolves each <automated> command's target directory and reports whether it exists and holds the manifest the command needs. /msd-plan-phase runs it before the plan-check pass and hands the JSON to the checker, which acts on severity instead of guessing.

It never executes command text. PLAN.md is model-authored, so running it from the checker would be arbitrary code execution — and would trigger the real lint/build as a side effect. The probe only resolves paths and stats directories; a package.json it finds is read for script names only.

Why a recognizer, not a shell parser. Interpreting shell would mean maintaining a bad shell. Exactly two forms are grounded — a leading cd <literal> chain and npm --prefix <literal> — and any path carrying a variable, glob, substitution, or ~ returns unresolvable, which is a warning and never a blocker. The parser's incompleteness is the specification: it degrades to "cannot prove" rather than growing features. Refusing to guess is the fix, not a limitation of it.

It reports, it never prescribes. The payload carries the target that failed and what was missing; there is deliberately no suggestion field. Choosing the replacement is the planner's job — and the planner now has the prior phase's proven command to reach for.

Not findings: a target an earlier task in this phase creates (pending_creation), a command with no cd/--prefix at all, and the Nyquist MISSING — Wave 0 … sentinel, which Dimension 8 owns.

Known limits:

  • Only cd <literal> and npm --prefix <literal> are recognized. pushd, make -C, yarn --cwd, pnpm -C, and cargo --manifest-path report unresolvable.
  • Verdicts are relative to the checker's project root. Under parallel worktree execution the executor's root differs, so a bare ancestor climb (cd ../..) is reported outside_root as a warning rather than asserted about.
  • script_missing is advisory only — this phase may be adding the script — so a genuinely mistyped npm script still reaches the executor.

See Resolve verify-command path findings and msd-tools check verify-command-paths.


162. Statusline STATE.md Freshness Marker

Config key: statusline.show_state_freshness (default false)

Purpose: A solo developer returning to a project after time away reads "Phase 4, executing" in STATE.md and acts on it — without noticing the codebase has moved 40 commits since that line was written. /msd-health reports this as W024, but only if the user thinks to run it. The statusline is the one surface seen continuously without asking (#2734).

Behavior: Renders state ~N commits back inside the MSD-state segment when STATE.md carries a state_head stamp (#2573) and HEAD is at least STATE_HEAD_ADVISORY_COMMITS (20) commits past it. Both statusline formats carry it — the default renderer and the compact statusline.state_format one.

The threshold is 20, deliberately not 1. With commit_docs: true (the default) the commit carrying a STATE.md sync advances HEAD by one, so a > 0 threshold would render state ~1 commits back permanently on a project that is by construction fresh — alarm fatigue on the one always-visible surface.

It degrades to silence rather than to a wrong answer. The marker is absent — never "fresh" — when the stamp is malformed, when the project root does not own its .git (an enclosing unrelated repo would otherwise answer), in a planning.sub_repos workspace (the outer HEAD never advances when code lands in children), when history was rewound past the stamp, and when git is unavailable or slow. A freshness claim the project cannot substantiate degrades to unknown.

Cost: exactly one bounded git rev-list call per render, and only when enabled and a stamp is present — rev-list --left-right --count answers ancestry and distance together, and repo pinning is a filesystem check rather than a subprocess. Disabled (the default) it adds none.

A proxy, never a drift measurement. The count includes commits that touched nothing STATE.md describes, and the stamp restamps on every state write — so a low count means "something wrote STATE recently", not "STATE is accurate". Rendered with a ~; never gate on it.

Reference: Configuration · Read the statusline freshness marker · ADR-2164


163. Read-Only Planning Snapshot (planning inspect)

Command: msd-tools query planning inspect

Purpose: Give downstream consumers — harness UIs, mission-control surfaces, dashboards, bots — one schema-versioned JSON document describing everything .planning/ knows, so nothing outside msd-core has to parse ROADMAP.md / REQUIREMENTS.md / *-PLAN.md / *-SUMMARY.md a second time. msd-core is the single source of .planning/ truth; a second parser is a second answer.

Requirements:

  • REQ-INSP-01: PLANNING_INSPECT_SCHEMA_VERSION = 1 is emitted as schema_version. Consumers MUST reject any other value rather than best-effort-parse an unknown shape.
  • REQ-INSP-02: Read-only. The command mutates no planning state, and mutates nothing on disk, under any input.
  • REQ-INSP-03: Unknown or conflicting evidence serializes as null / "unknown" with a coded entry in diagnostics[] — never inferred, reconciled, or defaulted. Every key is always present; a key is never omitted to signal absence.
  • REQ-INSP-04: Argument errors fail loud (non-zero exit, typed ERROR_REASON); data gaps do not. v1 takes no arguments, and a stray positional or unknown flag is a usage error rather than a silently-ignored one.
  • REQ-INSP-05: Roadmap acceptance, verification status, and UAT items are reported side by side per phase and are never folded into a single verdict. A ROADMAP checkbox carries authoritative: false — completion is derived from disk state.
  • REQ-INSP-06: accepted_phases and completed_plans are independent fractions. percent is null whenever the scope is not complete, per the same rule the roadmap and progress surfaces follow.
  • REQ-INSP-07: Payloads over ~50 KB use the existing @file: spill channel, resolved transparently before stdout.

Why it does not simply serialize the internal snapshot. PlanningSnapshot (the diagnostic-rule subject introduced by ADR-3180 §8.1) is deliberately additive and still growing — four fields at Phase 10, twenty-plus by Phase 12. Handing that shape to external consumers would freeze an internal contract by accident. planning inspect declares its own flat schema and maps into it, so a field added to PlanningSnapshot never changes what this command emits.

Composed, never re-derived. Milestone identity and phase enumeration arrive via buildPlanningSnapshot; completion from isPhaseComplete (disk-strict); live-plan counting from scanPhasePlans; the percentage arithmetic from clampPercent; STATE fields from stateFieldValue; plan bodies from the Plan Document Module; requirement IDs from parseRequirements; UAT items from parseUatItems. Markdown structure is read through the Markdown Sectionizer and Markdown Table Model seams, so the Traceability table is resolved by column name against its registered schema rather than by a position-anchored regex.

Known limit — task-scoped file provenance. A <task> declares the files it plans to touch, but SUMMARY.md's ## Files Created/Modified describes the whole plan. Spreading that list across a plan's tasks would be inference, so a task's changed_files is populated only where the summary attributes files to that specific task; otherwise it is null with provenance: "plan_scoped". Closing this needs a change to the SUMMARY format, not to the reader.

Reference: CLI Tools · Consume the planning snapshot


164. Live-DOM UAT Capability

Config key: workflow.live_dom_uat (default false)

Purpose: A phase with a live-UI acceptance criterion could not be finished by the agent that executed it. msd-executor carries no browser tools, so it correctly returned a checkpoint:human-action — even though the work was not human-only, just tool-less. Every such phase quietly degraded from executed by the executor to executed, then finished by hand in the orchestrator, and the plan's autonomous: false marker could not distinguish "a human must judge this" from "the executor lacks the tool" (#2856).

Behavior: A default-off capability owns one boolean key, one agent, and one additive step. When the key is on, msd-dom-verifier runs at execute:wave:post and writes {phase}-DOM-VERIFY.md; the orchestrator's automated_ui_verification step additionally considers mcp__chrome-devtools__* / mcp__claude-in-chrome__* when present.

The executor's tool surface is unchanged in every configuration. Widening it was the reported proposal and was refused: for a first-party agent the static tools: list is the only control that exists — no capability can grant tools to one (ADR-1244 D2), no hook kind grants tool permissions (ADR-857 D4), and there is no per-dispatch override. Browser reach lives in one purpose-built agent that carries no Bash.

Two independent gates, both fail-closed. The capability's activationKey makes it resolve inactive when the key is off — resolveLoopHooks renders a hook only on state.active === true — and the step carries its own when guard. Tool presence alone never activates it: a browser MCP configured for unrelated work is not driven by default.

The pre-existing Playwright path is untouched. mcp__playwright__* keeps the gating it already had (presence plus an active UI phase). Pulling it behind a new default-off key would have silently removed working behavior from current users on upgrade; the key gates only the newly added families.

It tolerates the browser-profile lock rather than coordinating it. chrome-devtools-mcp holds an exclusive lock on its profile, so parallel waves collide. --isolated is a flag on the operator's own MCP server registration — MSD neither launches that server nor passes its arguments — so the verifier reports could_not_look / profile_locked, names the flag, and stops. No retry, no held-up wave.

nothing_to_report is never conflated with could_not_look. A report claiming no issues when it never opened a browser is worse than no report; the artifact carries a closed reason enum so the two are always distinguishable.

Known limits: no sandbox — once enabled, nothing constrains which origins are reached (ADR-1244 D5); DOM observation only, no screenshot diffing, accessibility audit, or performance tracing.

Reference: Configuration · Enable live-DOM verification · Explanation · Agents


165. Opt-In Parallel Reviewer Lanes

Command: /msd-review, /msd-plan-review-convergence

Config key: review.parallel_lanes (default false)

Purpose: Reviewer lanes within one review pass have no data dependency on each other — they all inspect the same immutable plan snapshot — but were dispatched strictly one at a time, so a pass with Codex, Antigravity and Claude cost roughly the sum of three long reviewer calls. The serialization was a deliberate, unconditional protection against provider rate limits, which made it a global policy imposed on users whose providers could comfortably take concurrent requests, or who run local model servers with no limits at all (#3034).

Behavior: With the key enabled, the invoke_reviewers step dispatches each selected lane as a background job and joins all of them before REVIEWS.md and consensus are rendered. Wall-clock cost falls toward the slowest lane rather than the sum. Default remains false, preserving the existing sequential dispatch and its rate-limit protection.

The guard is strict equality, and it fails safe. Only the exact value true opts in — "1", "yes" and "TRUE" all stay sequential, so a mistyped config gets the conservative behavior rather than concurrent requests at a rate-limited provider. A failure to read the config falls back to sequential too. This polarity is deliberately the opposite of the commit_docs guard, which fails open: there, failing open preserves user intent; here it would fire the very requests the default exists to prevent.

Result ordering is unchanged in both modes. Per-lane results are written to slug-scoped files and concatenated in reviewer-selection order after the join, so msd-review-lane-results.jsonl reads identically whether lanes ran sequentially or concurrently. Completion order never reaches the artifact. This also means concurrent lanes never share an append handle — a lane result larger than the pipe-atomicity bound cannot interleave and corrupt the models: / model_sources: frontmatter that write_reviews renders from that file.

Per-lane semantics are untouched. Timeouts, prompt budgets, the diagnostic stub for an empty or failed lane, explicit-lane failure (ADR-2782 D4), trust/egress checks and result-file layout all behave exactly as they do sequentially. A failing lane does not abort its siblings.

Known limits: convergence cycles stay sequential by design (review → replan → re-review has a genuine data dependency), so this speeds up each pass rather than reducing the number of passes; there is no concurrency bound, so every selected lane dispatches at once; and reviewer instances sharing one adapter dispatch concurrently against that single provider, which is the most likely way to hit a limit.

Reference: Configuration · Enable parallel reviewer lanes · Commands


166. Machine-Readable State Contract (.planning/state.json)

Purpose: External tools that display MSD project state — a workbench, a dashboard, an editor extension — had to parse STATE.md and ROADMAP.md heuristically. Those are human surfaces: their shape drifts as the templates evolve, and every consumer ends up carrying a brittle second parser that silently reports wrong numbers after an upgrade. MSD now publishes a small, versioned JSON snapshot instead, so the reader binds to a contract rather than to markdown (#3227).

Behavior: At every step boundary, MSD writes .planning/state.json — contract, flavor, milestone, phases[], next, updated_at. The boundaries are state begin-phase / planned-phase / advance-plan / complete-phase / milestone-switch, phase add / add-batch / insert / remove / complete, and milestone complete. The write is best-effort and completely invisible to the command that triggered it: it cannot change an exit code, cannot change stdout, and cannot fail a workflow. Readers prefer the file when it is present and fall back to markdown when it is not.

Requirements:

  • REQ-SC-01: contract is semver, 1.0.0 at introduction. Consumers gate on the MAJOR version; 1.x changes are additive only. Every key is ALWAYS present — an unknown value is null, never an omitted key, because an omitted key is itself an observable a consumer would bind to.
  • REQ-SC-02: phases[] carries {number, name, status} per phase, status drawn from exactly complete | in_progress | pending. number is a string ("01" and "2.1" are both real ids and neither survives a number cast); name is null when the roadmap gives a phase no name, never a fabricated placeholder.
  • REQ-SC-03: next is the same recommended action the /msd front door routes, derived from the smart-entry classifier itself rather than from a second copy of its routing table.
  • REQ-SC-04: A missing ROADMAP.md, a missing or unreadable .planning/, an unwritable target, or any other failure NEVER errors the parent command. A directory that is not a MSD project stays untouched — the publisher will not create .planning/ in order to publish into it.
  • REQ-SC-05: The skills own the file; readers never write it. It is a derived cache — safe to delete, regenerated at the next boundary.

Composed, never re-derived. Milestone identity comes from getMilestoneInfo; phase rows from locateProgressTable, the same ## Progress locator the progress counters use, so state.json can never disagree with the rest of MSD about which phases are complete; the recommended action from classifyProject. This module introduces no second answer to any question MSD already answers.

Why it does not reuse planning inspect's schema. The two surfaces answer different questions and have opposite shapes. planning inspect is a rich, diagnostic-carrying pull query a consumer runs; this is a small push artifact a consumer watches. Publishing planning inspect's payload at every phase add would mean opening every plan, summary and requirements document on a hot path, and freezing a much larger surface as a contract.

It costs up to three bounded git calls per boundary. Deriving next from the smart-entry classifier means inheriting its git signals — git status --porcelain, and git log @{u}..HEAD. Each is timeout-bounded and swallows every error, so nothing can hang or fail because of it, but a command like phase add did not previously touch git at all. "Invisible to the parent command" is exact about exit code and output; it is not a claim about latency.

Known limits: an empty phases: [] cannot be told apart from "no ROADMAP.md" or "roadmap unreadable" — the 1.0 schema carries no diagnostic channel, and planning inspect is the surface that does. A roadmap phase marked Deferred is reported as pending, because the roadmap vocabulary has four values and this contract has three; inventing a fourth wire value would break every existing reader. phases[] is not milestone-scoped, so a long-running project lists every phase it has ever had.

Reference: Consume the state contract · Consume the planning snapshot


167. Stated Failing Direction

Command: /msd-plan-phase (automatic), msd-tools check verify-failure-directions <N> (#3172)

Behavior: A plan's <automated> block is the thing that decides whether work is done, and nothing checked that the command inside it could fail. In the motivating case six plans shipped 21 commands that could not run at all — cargo test -p <pkg> --lib against a package with no library target. They read as rigour and were not falsifiable, so three separate executors each rediscovered the defect and improvised a substitute at execution time. Every runnable <automated> command now needs a <fails_when> sibling naming what output constitutes failure:

<verify>
  <automated>npm --prefix apps/api test -- auth.spec.ts</automated>
  <fails_when>non-zero exit, or "0 passed" in the summary line</fails_when>
</verify>

msd-planner emits it; msd-tools check verify-failure-directions <N> verifies it deterministically; /msd-plan-phase runs the probe before the plan-check pass and hands the JSON to msd-plan-checker, whose check 8f blocks on severity.

Why this shape and not the two obvious alternatives. Validating command shape — teaching the checker Cargo's --lib/--bin target resolution, then pytest's node-ids, then the next one — always trails the newest toolchain. Executing each command at plan time is the strongest signal but means running planner-invented commands, with whatever side effects they carry, during planning. Requiring a stated failing direction needs no toolchain knowledge at all, and it is the only one of the three that catches the dangerous case: the motivating command exited non-zero, so it failed loudly, but the same class of error with a command that exits 0 on a no-op passes green and silently. Naming the failure signal is what makes that visible.

Presence, not quality — deliberately split. The probe is deterministic and owns the blockers: a statement is missing, blank, or a whole-value placeholder (TBD, TODO, N/A, NA, none, unknown, TBA, ?, -). Whether the statement names the right signal is prose judgment, so msd-plan-checker raises a vacuous statement ("the command fails") as a WARNING only. Every BLOCKER stays reproducible; judgment stays advisory.

It reports, it never prescribes. The payload names the command with no stated failure mode and stops there. A prescribed statement would be copied verbatim and carry zero information — reproducing the original defect one level up.

Not findings: the Nyquist MISSING — Wave 0 … sentinel (not runnable, so it has no failure mode to state), an empty <automated> body (check 8a owns command presence), and a <verify> with no <automated> at all.

Known limits:

  • Presence only. A statement that is present and specific can still name the wrong signal; that is caught, if at all, by judgment rather than by the probe.
  • Breaking for plans authored before this shipped. A phase planned earlier has no <fails_when> anywhere and blocks on re-check until statements are added or the phase is re-planned.
  • The adjacent vacuous pass — a command that runs successfully and asserts nothing, such as a test-name filter matching zero tests and exiting 0 — is a distinct problem and is explicitly out of scope.

See State a failing direction and msd-tools check verify-failure-directions.


168. Runtime Identity

Purpose: The predecessor package get-shit-done-cc publishes a binary named msd-tools, and so does this one. They answer some of the same verb names with different semantics. #3129 is the worked example: phases.clear archives here and deletes there. Both print success-shaped output, and .planning/ is gitignored by default, so a user lost 43 phase directories with no error, no warning, and nothing recoverable from git. The failure was silent in both directions — the workflow could not tell it had reached the wrong handler, and the handler could not tell it had been called by a workflow written for a different contract (#3146).

Behavior: two independent defenses, one structural and one asserted.

Structural — the PATH branch. The launcher's PATH resolution branch looks for msd_run instead of msd-tools. Only this package publishes msd_run; the predecessor publishes msd-tools and msd-sdk. Our msd_run follows its own symlink chain and executes the msd-tools.cjs sitting beside it, so resolving it cannot land on a foreign handler.

Asserted — every other branch. The path-based branches (a project-local install, a runtime config directory) have no such guarantee: they trust their configured location. So once resolution finishes, and before any verb runs, the preamble probes the tool it picked with runtime-identity --raw and matches the answer anchored against the compact payload. It exports the result as a two-valued MSD_IDENTITY_STATUS (ok / unverified) and, when it is unverified, prints one actionable line naming both plausible causes. The same msd-tools runtime-identity verb remains available by hand, so a human or a support thread can settle "which tool am I actually running?" in one command.

The match is anchored at both ends, not a substring. A substring search for @golem15/msd-core accepts the decoy {"packageName":"get-shit-done-cc","note":"@golem15/msd-core"}, which any colliding package could publish. The preamble instead requires the payload to begin with {"packageName":"@golem15/msd-core" and to end with a closing brace, so a truncated answer fails as well. Closing on } costs nothing in future-proofing: a JSON object's own brace is always the last character, whatever type the last value has.

The status is a value, not prose. MSD_IDENTITY_STATUS exists so the gate can be tested — and read by a later step — without anyone parsing the warning text.

The byte budget is why the assertion arrived second. The preamble is inlined into 113 shipped files, several of which sat within single-digit bytes of frozen size ceilings — agents/msd-verifier.md had 2 bytes of headroom — and those caps are red lines, not budgets. A first attempt to inline an assertion broke five of them. What made it fit was collapsing the resolver's twenty near-identical elif [ -f … ] arms into a single candidate-list helper, which is worth far more bytes than the assertion costs: the preamble is now 1,876 bytes smaller than the version that carried no assertion at all, so every one of the 113 files moved away from its ceiling.

It fails closed. If no msd_run is reachable, the resolver falls through its remaining path-based branches and finally errors with an install command. It does not fall back to executing whatever msd-tools happens to be on PATH — that fallback was the vulnerability.

A doubly-sourced preamble cannot build a recursive launcher. command -v msd_run finds the shell function on a second source and would return the bare string msd_run, defining the function in terms of itself. unset -f msd_run leads that branch, so the second source resolves exactly as the first did. (An executability guard was tried here instead and removed: it rejected the bare name, fell through every branch, and hit the resolver's exit 1 — which, in a sourced script, kills the caller's shell.)

Known limits:

  • The assertion warns; it does not yet stop the run. The rollout is warn-then-fail. It cannot hard-fail yet because an @golem15/msd-core older than the runtime-identity verb answers exactly as a foreign package does — neither answers — and at rollout the old-version case is the common one. The warning therefore names both causes. A later release turns unverified into a refusal.
  • An installation old enough to predate bin/msd_run (#381) is not reachable through the PATH branch and must be upgraded or invoked through one of the path-based branches.
  • The probe costs one extra process launch per preamble source. It is a pure local read of baked coordinates, deliberately kept off the SDK bridge for that reason.

Reference: runtime-identity · Diagnose which msd-tools is running


Generated by scripts/gen-features.cjs — add a fragment under docs/features/ and run --write.


3348. Context Drift Gate

Purpose: Warns (or optionally blocks) before /msd-plan-phase reuses an existing RESEARCH.md, PATTERNS.md, VALIDATION.md, or SPEC.md that predates a decision added to the phase's CONTEXT.md after that artifact was derived from it. Deterministic — compares git commit time (falling back to mtime for uncommitted edits), no model call. Sibling to the existing codebase-drift and schema-drift gates in the drift capability. Configure with workflow.context_drift_precheck (on/off) and workflow.context_drift_action (warn/block).


3884. "Failure Is a Value" — Strict Argv Rejection and the --pick Absence Contract

Purpose: ADR-3473 §8.4 states the rule directly: absence, emptiness, and failure are three different things, and a routine that cannot tell them apart eventually reports the wrong one. Before this change, msd-tools had two silent instances of exactly that collapse.

Half one — a stray positional corrupted state, silently (#3358). parseNamedArgs read only the flags it recognized and dropped everything else — an unrecognized --flag or an extra positional argument (for example, a stray phase number appended after state.planned-phase) was silently discarded rather than rejected. The caller's own positional read (args[2], etc.) still worked, so the command ran anyway, on the wrong phase, and overwrote the previously-current phase block with no error at all. The fix makes parseNamedArgs(args, spec) return the command-routing hub's own Result shape ({ok:true,data} | {ok:false,kind:'InvalidArgs',...}) and requires every call site to declare positionals: number | 'rest' — the count of leading argv slots the caller itself reads directly. An unknown flag or an unexpected positional past that boundary is now a loud, non-zero-exit InvalidArgs failure instead of a token quietly falling on the floor. A duplicate flag, a negative-number value (--plans -1), and a documented free-text tail (init quick <description>) are deliberately left alone — none of them are the defect this closes, and forbidding them would just break working call sites for no gain.

Half two — --pick on an absent field answered '' at exit 0, exactly like a present-but-empty one (#3365). --pick <field> extracted one field from a command's JSON output, but a missing key, an out-of-range array index, a partially-missing dotted path, or non-JSON output (including a --raw command's output) all rendered the same way: empty stdout, exit 0. That is indistinguishable from a field that genuinely holds null or '' — a real answer. The shell idiom X=$(… --pick F) || X=default could therefore never observe the failure it was written to react to; only a typo in the verb name would ever make it exit non-zero. --pick now exits 1 with a diagnostic on stderr (pick_field_absent naming the field and the available top-level keys, or pick_output_not_json when the output could not be parsed as JSON at all) whenever the field cannot be resolved. A present field's value — including 0, false, null, and '' — is unchanged: those are answers, not failures, and remain exit 0. See CLI-TOOLS.md's --pick <field> contract for the full outcome table and json-errors.md for the two new reason codes.

Why not just default to zero for an absent count. The sub-issue's own Done-when checkbox suggested treating an absent field the same as a zero-valued one. That is rejected on the merits: it demotes "I could not answer" to "the answer is zero", which would make a count-gated shell guard fire unconditionally on the very projects that could never resolve the count in the first place — the opposite of what a gate is for.

Consequence for scripts/lint-unreachable-guard-drift.cjs. That guard's Detector A existed specifically because the old --pick behavior made a --pick … || echo <default> line's fallback arm permanently unreachable. Once --pick exits non-zero on absence, that premise is false and the shape the detector forbade becomes the correct idiom — so Detector A was retired rather than kept. Detector B (the unrelated cat/ls-over-a-glob nullglob hazard) is untouched. See Resolve unreachable-guard findings, Shape A.

Known limits:

  • A value token beginning with -- still cannot be passed to a declared value flag (--summary "--force is now default" now fails loudly instead of silently dropping the value) — strictly better, but no --flag=value escape was added.
  • --pick still cannot distinguish an absent field from a null one on stdout alone — the distinction is carried entirely by exit code.
  • The ~10 msd-tools.cjs call sites of parseNamedArgs get no compile-time check (that file is hand-written JavaScript, not .cts); enforcement there is the runtime throw on a stale legacy call shape plus behavioral tests.
  • This phase does not sweep every routine in msd-core for Result conformance — it applies the rule to the argument-projection seam and the --pick extractor it names, not the whole codebase.

3885. No Silent Swallow, No Verdict From Dropped Data

Purpose: ADR-3473 §8.5 states the rule directly: a failure or a gap in the input must not be absorbed into an output that reads as authoritative. A routine that drops data it could not read or could not resolve, and then reports a clean result anyway, turns a diagnosable gap into a confidently wrong answer. This closes four instances of that collapse found across msd-tools.

intel query no longer crashes past ~12000 levels of nesting (#3427). searchJsonEntries / matchesInValue recursed with no depth bound at all — an intel JSON file nested deeply enough overflowed the call stack with an uncaught RangeError instead of a diagnosis. The original SDK-era bound (MAX_JSON_SEARCH_DEPTH = 48, lost in the ADR-0174 consolidation) is restored, paired with a truncated result field: a match at or above the ceiling is not returned, and the result says so rather than reporting a bare "not found" that is indistinguishable from a genuine miss. A match at depth 48 (inclusive) or shallower is unaffected; the bound is on nesting depth, not breadth or total node count, so a shallow object with many siblings still works unchanged.

phase-plan-index no longer blames the author for an edge the tool itself dropped (#3427). A depends_on: token that resolves to no plan in the phase (typo, or a stale cross-phase reference) silently dropped that edge, making the dependent plan a DAG root — its own docstring recorded the intent as "ignore this edge, never a throw." The tool then compared the resulting degraded wave against the plan's declared wave: and reported the author's correct declaration as a mismatch. The unresolved token is now named in its own warnings[] entry (plan and token together), and the wave-mismatch warning is suppressed for that plan only — a plan with no dropped edges and a genuinely wrong wave: still warns as before. The token is escaped (quoted, control characters and embedded newlines backslash-escaped) before it is embedded in the warning text, so a depends_on value crafted to contain a newline or a quote cannot forge a second, fabricated warning entry.

A code-review run where every lane failed no longer writes REVIEWS.md from nothing (#3352). review.md's aggregation step wrote REVIEWS.md regardless of whether any lane actually produced results — a run where every lane failed still emitted a completed-looking review artifact, and the per-lane outputs and .err files that would have explained the failure were then destroyed by the run's own cleanup. REVIEWS.md is now withheld when the aggregate has zero lines (every lane failed, not merely skipped under a lower budget), the run reports the failure instead, and per-lane outputs and non-empty .err files are preserved beside the phase's artifacts before cleanup runs.

Unreadable directories are distinguished from absent ones (#3473 B5). Four call sites collapsed an EACCES/EIO on a phase directory into the same "nothing here" result as a directory that genuinely does not exist — countPhasePlansAndSummaries (hasContext:false), runGapAnalysis, and two guarded blocks in init.cts (context_path absent). Each now distinguishes "could not read" from "does not exist" and names the discarded path and error in a dedicated field (context_read_error / phase_dir_read_error) rather than silently reading as absent.

Audited, no defect found: every retry-set / swallowed-catch call site this phase's rule covers that had not already been fixed by a prior PR (withPlanningLock, acquireStateLock, atomicRenameWithRetry, renameWithRetry) was reviewed and found to already fail loudly on a fatal errno rather than folding it into a retry.

Known limits:

  • A match deeper than 48 levels is still not surfaced by intel query — it is reported as truncated rather than as absent, but the value itself is not returned. Raising the ceiling is a separate decision.
  • phase-plan-index's waves / wave fields remain computed from the degraded DAG when an edge is dropped — this phase stops the tool from manufacturing a false verdict about it, but does not invent the missing edge. A consumer that schedules work from wave (e.g. --wave N filtering) is still working from the degraded assignment.
  • review.md's evidence preservation is bounded by what a lane actually wrote — a lane that produced no output at all leaves nothing to preserve.

3897. Runtime Marker Resolution, Derived Codex Sandbox, and In-Phase Short-Form Dependencies

Purpose: ADR-3473 §8.3 states the rule directly — one implementation per invariant, not a hand-maintained copy that quietly drifts from the rule it stands in for. This closes three instances across msd-tools: a resolver that never read the install-time signal it was documented to read, a Codex sandbox map that was fully redundant with the tool contract it stood in for, and a dependency-resolution tier lost when the SDK lineage was retired.

A non-Claude install now resolves its own runtime with no config needed (#3897). resolveRuntime's ladder was MSD_RUNTIME env var → project config.runtime → 'claude' — the per-install .msd-runtime marker the installer has written beside VERSION since #2297 was read by four separate hand-rolled copies (model-resolver.cts, and two more inside msd-cursor-subagent-start.js), but never by resolveRuntime itself. A Codex, Cursor, or other non-Claude install with no MSD_RUNTIME set and no runtime key in .planning/config.json therefore still resolved claude everywhere resolveRuntime is consulted (slash-command style, query teams-status, validate agents's agent-directory selection, and 19 other call sites). The marker is now the third rung — MSD_RUNTIME → config.runtime → install marker → 'claude' — so those installs resolve their own runtime by default. The marker's contents are never trusted verbatim: they are routed through the same name-normalization the env rung already uses, so a marker holding an unexpected or hostile value degrades exactly like an unexpected MSD_RUNTIME value would.

Codex sandbox permissions are derived from each agent's own tool contract, not a hand-maintained map (#3897). generateCodexAgentToml looked up sandbox_mode in an 11-entry CODEX_AGENT_SANDBOX map, falling back to read-only — silently — for every role the map didn't name. Measured against all 35 shipped roles, the map's 11 entries agree with deriving sandbox_mode from each role's declared tools: frontmatter (workspace-write when it declares Write or Edit, read-only otherwise) with zero disagreements, so the map is deleted rather than clamped. The fallback, however, was under-granting: 16 of the 24 roles that hit it declare Write/Edit and would derive workspace-write. Pending a decision on whether Codex actually enforces sandbox_mode (a question the derivation can't answer on its own), those 16 are held at read-only by an explicit, self-invalidating hold list — a hold whose role no longer derives broader, or that names a role that no longer exists, fails loudly instead of being silently honored. Every one of the 35 emitted .toml files is byte-identical to before this change — the fix is in provenance (an explicit, reviewable rule instead of a silent default), not in any installed agent's actual permissions today.

validate agents now reports Codex sandbox drift (#3897). A new sandbox_posture field — report-only, exit 0, same shape as the existing codex_posture — flags any installed Codex .toml whose sandbox_mode disagrees with what its role's tool contract derives. Populated only when the active runtime is codex.

depends_on accepts the bare plan number (#3897). A plan's frontmatter could already reference a dependency by its full id ("03-01-auth-hardening") or its canonical phase-plan prefix ("03-01"). A third form — the bare plan number alone ("01") — existed in the retired SDK lineage but was lost when that lineage was consolidated; a plan written with it silently dropped the edge entirely, collapsing into wave 1 regardless of its declared dependency. That form is restored, scoped to the same phase only: "01" resolves to the sibling plan whose canonical id ends -01. This is an observable behavior change — a phase whose plans used the bare form and had silently collapsed into a single wave will now execute in its actual declared waves. Two plans in the same phase sharing a bare form resolve first-write-wins, by sorted plan-file order — deterministic, but arbitrary where the collision happens, matching the retired behavior exactly.

Known limits:

  • The 17 held Codex roles are pinned at read-only, not widened. A faithful derivation from the tool contract would widen them, because they declare Write or Edit; the previous hand-maintained map never listed them and they fell through a silent || 'read-only' default instead. Deriving and holding keeps emitted TOML byte-identical for all 35 roles today while the derivation becomes the single owner of the rule. Widening them is a follow-up once Codex's actual enforcement of sandbox_mode is confirmed — until then a hold is reversible and a widened sandbox is not. The hold list is self-invalidating: an entry naming a role that no longer derives broader, or that has no file in the shipped roster, fails rather than rotting into the subset map this change deletes.
  • The bare plan-number form is ambiguous by construction across two plans in the same phase that share a short form; first-write-wins is deterministic but not a conflict warning. Prefer the full or canonical id when a phase's plan numbering risks a short-form collision.
  • The install marker never feeds model-tier resolution (model_profile_overrides, model_policy.runtime_tiers) — that still reads config.runtime alone, and reporting-only host detection (agent_runtime) is a separate, pre-existing ladder this change does not touch.

3910. The Raw Terminator Is Banned by Construction

Purpose: Make a bare process.exit(...) a lint error everywhere it matters, so the "nothing fails with success" defect class ADR-3889 exists to close cannot silently reopen through a new call site.

What changed (ADR-3889 Phase 6, #3910):

  • New rule local/require-registered-exit (eslint-rules/require-registered-exit.cjs) flags any CallExpression shaped exactly like process.exit(...). It does not flag process.exitCode = N — that assignment is the correct drain-then-exit pattern runMain itself uses, and the two are structurally distinct (an assignment target is never a CallExpression).
  • Registered on four globs: src/**/*.cts, scripts/**/*.cjs, hooks/**/*.js, msd-core/bin/**/*.cjs (eslint.config.mjs:420-426,545-547,574-576,601). Registering on src/**/*.cts — not only the emitted msd-core/bin/lib/*.cjs mirrors, which are globally eslint-ignored (ADR-457) — is load-bearing: a rule registered only on the emitted surface is blind to the real sources, the same way n/no-process-exit went invisible (#3496).
  • The dead n/no-process-exit: 'off' carve-out for hooks/** is deleted: Phase 7 (#3911) migrated every enforcement hook onto terminateNow, so it protected nothing.
  • Exactly two allowlist entries, repo-wide:
    1. The body of terminateNow in src/cli-exit.cts — detected structurally (any process.exit() lexically nested inside a function named terminateNow, and the file's basename is cli-exit.cts), not by path+line, so it does not rot when the function moves.
    2. msd-core/bin/msd-tools.cjs's ensureRuntimeBuild bootstrap-failure path, via an inline // eslint-disable-next-line local/require-registered-exit with a stated reason — it runs before ./lib/cli-exit.cjs is even required, so the registered-exit seam does not exist yet at that point in the process's lifetime.
  • The last raw terminators in src/**/*.cts were migrated onto the seam, most notably src/io.cts's error(): it changed from an uncatchable process.exit(1) to a catchable throw new ExitError(1) (stderr output is byte-identical; runMain projects the exit code). terminateNow could not serve this site — ADR-3889 §1 makes exit codes 0 and 1 unallocatable, so nameForExitCode(1) throws. That control-flow change required three interceptor fixes so an ExitError reaches runMain: command-routing-hub's dispatch() now rethrows it, and the profile-pipeline router's detached .catch() no longer calls error() — it writes stderr and sets exitCode in place.
  • Known limits (documented and test-pinned, not endorsed): the rule matches the literal process.exit(...) shape only, with no scope/flow analysis. It does not catch process['exit'](0) (computed member access), const e = process.exit; e(1) (aliasing to a local binding before calling), or process.exit.call(...)/.apply(...) (indirect invocation). Catching these needs binding/scope-aware analysis, out of scope for this issue; pinning tests in tests/eslint-rules.test.cjs assert today's non-detection so a future widening is a visible choice, not a silent one.

See Resolve a raw-terminator finding for what to do when this rule fires, and ADR-3889 for the exit-code registry the seam is layered over.


3911. Hooks Declare Their Crash Policy

Purpose: Give every shipped enforcement hook (hooks/*.js, hooks/*.sh) a named, auditable termination vocabulary instead of a bare process.exit(N) scattered per file — and make a hook's fail-open/fail-closed choice a declaration a reviewer can see, rather than an inference from which literal integer follows process.exit( in its outer catch.

What changed (ADR-3889 Phase 7, #3911):

  • hooks/lib/hook-exit.js (hand-written) exposes allow(payload) → exit 0, deny(payload, stderrPayload?) → exit 2, and crash(onCrash, payload), which dispatches to allow/deny per a HOOK_ON_CRASH policy the caller must supply — crash() has no default policy, so a hook cannot fail open by omission.
  • Every one of the 19 enforcement hooks under hooks/*.js now declares const ON_CRASH = HOOK_ON_CRASH.ALLOW (or DENY) once, with a hook-specific comment naming why, and calls crash(ON_CRASH, payload) from its outer catch instead of a bare process.exit(0) / process.exit(2). No hook's effective exit code changed — this is a naming-and-declaration migration, not a behavior change.
  • hooks/lib/cli-exit.js and hooks/lib/exit-code-registry.js are new, generated, git-tracked copies of the exit-code seam (src/cli-exit.cts / msd-core/bin/shared/exit-codes.json), so a shipped hook can terminate correctly on a raw, unbuilt clone without depending on msd-core/bin/lib/ tsc output. Generated by scripts/gen-hooks-cli-exit.cjs and scripts/gen-exit-code-registry.cjs, both --checked by npm run lint:generated-sync.
  • terminateNow gained an optional third argument, stderrPayload, so a deny can send a full JSON body to stdout and a distinct plain-text reason to stderr — needed because msd-write-guard.js (a host's native hook bus may read stderr verbatim back to the model) always sent only the bare reason string on fd 2. The two streams are now written in independent try/catch blocks: previously a payload that failed to serialize on fd 1 aborted before fd 2 ever wrote, producing a deny with an empty stderr reason.
  • Two hooks are deliberately not migrated to deny(): msd-read-injection-scanner.js (PostToolUse — its harness reads the block decision from the JSON response body, not the exit code) and msd-cursor-subagent-start.js (follows Cursor's own subagentStart protocol, which reads permission: "deny" from the JSON body at exit 0). Both still use allow()/crash() for their no-op and crash paths.
  • msd-phase-boundary.sh, msd-session-state.sh, and msd-validate-commit.sh gained set -euo pipefail, and msd-validate-commit.sh's three swallow-and-pass sites (the opt-in config read, JSON command extraction, and the isGitSubcommand classifier) now distinguish a genuine negative from "could not run" — on the latter they emit a stderr diagnostic and exit 0 instead of silently allowing every commit (#3838).

See Declare a hook's crash policy for the full how-to, and ADR-3889 for the exit-code registry this vocabulary is layered over.


3912. msd-tools Declares Outcomes, Pinned at v1

Purpose: Give every msd-tools terminating path a declared outcome name, and project that declaration through the versioned exit contract (ADR-3889 §4) — without changing a single exit code for a caller that has not opted in.

Reference — what changed (ADR-3889 Phase 8, #3912):

  • error(message, reason) now maps its reason argument onto a declared outcome name (USAGE, NO_INPUT, UNAVAILABLE, INTERNAL, FAIL) via a fixed table closed over all 25 ERROR_REASON members (src/io.cts's REASON_TO_OUTCOME). Under the default contract version v1, the declaration is recorded but error() still throws ExitError(1) unconditionally, byte-identical to every prior release. Under v2 (--exit-contract=v2 / MSD_EXIT_CONTRACT=v2), it throws ExitError(projectOutcome(outcome, 'v2')) instead — e.g. SDK_MISSING_ARG/SDK_UNKNOWN_COMMAND project to 64 (USAGE), CONFIG_KEY_NOT_FOUND to 66 (NO_INPUT). All 278 call sites are untouched; 226 pass no reason and default to UNKNOWN -> FAIL -> exit 1 under both versions.
  • output() now declares DEGRADED whenever its payload carries a serializable error value (any key order — {found:false, error} counts the same as {error, found:false}). The discriminator is survives-JSON.stringify, not mere key presence: { error: undefined } does not declare DEGRADED, because JSON.stringify drops an undefined-valued property before the payload reaches the wire.
  • A third globalThis cell (src/cli-exit.cts's PENDING_OUTCOME_KEY) holds the pending declared outcome between output() and runMain. Semantics: last declaration wins, cleared on consumption — a later clean output() call in the same invocation undoes an earlier degraded one, and runMain clears the cell on every exit so a second runMain in the same process never inherits a stale declaration.
  • Precedence for the code a void-returning main() ends up with, highest first: (1) an explicit main() return, (2) a non-zero process.exitCode main() already set directly, (3) the pending declared outcome, (4) otherwise 0. Projection may only ever set a code, never lower one — a review pass wrongly concluded the cell was fail-closed by construction; without rule 2, state validate --strict briefly exited 0 where it must exit 1.
  • v1 is byte-identical. DEGRADED projects to 0 under v1 and to 80 (exitCodeFor('DEGRADED')) under v2 — that asymmetry is ADR-2980's compatibility boundary, deliberately preserved, not a bug to reconcile.

Explanation — why this is the shape it is:

ADR-2980 ratified output({error})'s exit-0 population on measured blast radius (output has 170 direct callers) and Hyrum's Law grounds — a CLI exit code has no /v2/ of its own, so normalizing it in place would have broken every caller already treating exit 0 as a soft signal. Its own "Revisit if" clause named the missing piece: "a future msd-tools major version provides a compatibility boundary that a CLI exit code otherwise lacks." ADR-3889 §4 built exactly that boundary — a versioned projection selected by flag or env var, defaulting to today's behavior — and this phase is what wires error() and output() onto it. Declaring an outcome is unconditional and immediate; only its projection onto an integer is deferred behind the version switch, so the population ADR-2980 ratified keeps exiting 0 until a caller explicitly asks for something else.

The count matters here too: an AST re-measure for this phase found 64 output({error}) call sites across the same nine modules ADR-2980 named — not the 60 that ADR itself recorded, the drift concentrated in frontmatter.cts, phase.cts, and roadmap.cts. The v2 projection is asserted over the enumerated 64, not a restated 60; see ADR-2980's amendment for the module-by-module breakdown.

See Adopt the v2 exit contract for how to opt in and what it means for a CI gate, docs/json-errors.md for the full reference, and ADR-2980 / ADR-3889 for the decisions.


3951. Reachable Lint Rules and a Non-Destructive Quick-Task Append

Purpose: Make two ESLint rules cover the code they were written to govern, and stop quick-tasks-append from overwriting curated progress.* values on a body-only write.

What changed:

  • local/no-adhoc-markdown-parsing reaches its whole registered surface. The rule short-circuited unless a file's path matched a flat src/*.cts pattern, so it self-gated on its own filename. Two consequences: 28 .cts files in src/ subdirectories sat inside the src/**/*.cts glob it was registered on and were silently skipped, and the rule could not be extended by configuration at all — widening the glob alone left it inert. Both halves now move together, and a test pins that the gate and the registration agree in both directions.
  • The rule now also covers tests/** and scripts/**, which surfaced 80 hand-rolled markdown parses across 43 test files. Seventy are routed through the existing markdown-sectionizer and markdown-table seams; ten are suppressed with a stated reason (six of those are a shell-pipe detector whose regex merely resembles a table).
  • local/no-adhoc-regex-escape sees property access. Its unsafe-new RegExp arm examined only bare identifiers, so new RegExp(obj['key']) — the shape runtime data actually arrives in — was invisible. That is why it never fired on a known ReDoS. It now inspects MemberExpression, with an exemption keyed strictly on the property being source (18 safe sites), plus provenance exemptions for _SOURCE constants reached through a required module (3 sites).
  • quick-tasks-append can write the canonical row. Optional --quick-id, --slug and --directory let a caller that has a real quick task emit the same row /msd-quick renders. Omit them — as fast.md does, having neither an id nor a directory — and the row is byte-identical to before.
  • A body-only append no longer re-derives progress. The route was the only body-only STATE.md writer not passing { resync: false }, so appending one row triggered a full re-derive of the disk-derived progress.* frontmatter and replaced curated values. Reproduced: a project with two real phase directories and a curated total_phases: 25 collapsed to 2 on append.

Found by the widening: tests/config-field-docs.test.cjs asserted that workflow.subagent_timeout's documented default is not 600 — but read the Type column instead of Default, so it compared 'number' against '600' and could never fail. The guard against regressing to the old seconds-based default had been inert. It is now row-scoped and real.

Known limits:

  • #3426 and #3239 are not closed by this. Their hand-rolled scans in tests/package-legitimacy-gate.test.cjs are built from line filters and split('|'), not the regex-literal fingerprints this rule detects — measured at zero violations even with the gate bypassed. They need new detectors, which is a separate design.
  • The 10 suppressions are suppressions, not fixes. Each names why the raw markdown text is the subject of that assertion.
  • The src/ subdirectory hole was latent — zero violations existed there when it was fixed. It is closed because "no violations today" is not a property that keeps holding, not because it was hiding anything.

3970. Per-Task External-Tracker Content-Resolution Seam

Purpose: Let a capability declare that an external issue tracker — beads, Linear, Jira, GitHub Issues — owns a task's content (<action>/<verify>/<acceptance_criteria>/ <read_first>/<done>), not just its status, so execute-plan.md can resolve that content from the tracker at execution time instead of reading it inline out of PLAN.md.

What changed (ADR-3646, #3970):

  • A new optional feature-body manifest field, taskContentResolver, declares a trackerPrefix (matched against a task's <task tracker-id="beads:MSD-42"> attribute — everything before the first :) and a bounded invoke (binary, args carrying the {{id}} placeholder, timeoutMs).
  • execute-plan.md's per-task loop gains one new, unconditional call before that task's read_first gate: msd_run task resolve-content --plan <path> --task-id <tracker-id> --raw. A task with no tracker-id attribute is unaffected — the call is only made when the attribute is present, and resolves instantly to a no-op for every project that declares none.
  • The safety property is a real process exit code, not a prose dispatch. No capability registered for the tracker, or resolution succeeds with empty content, exits 0 with resolved: false and falls back to inline PLAN.md — the one legitimate pre-migration boundary case. Resolution succeeding with non-empty content exits 0 with resolved: true and its content supersedes the task's inline fields for every downstream gate in the execute step. A resolver that is declared but fails — tracker unreachable, id not found, timeout, malformed JSON — makes task resolve-content itself exit non-zero, which execute-plan.md treats as a hard halt: stop, surface the tracker-id/prefix/stderr, never fall back to stale PLAN.md content.
  • execute:task is a new dispatch shape below wave granularity, deliberately not one of the 12 existing loop extension points (discuss:pre … ship:post) and not routed through msd_run loop render-hooks <point> / activeHooks. It exists because the existing step/gate prose-dispatch mechanism cannot deliver a hard-halt guarantee while dispatch reliability at that layer is an open concern (#3647) — see ADR-3646's Context and Rejected Alternatives for the full reasoning.

See Develop a task-content resolver capability for the authoring walkthrough, Capability manifest → taskContentResolver for the field reference, and loop-hook-dispatch.md for how execute:task differs from the twelve prose-dispatched points.


4014. Unreadable-Directory Scope Signal

Purpose: ADR-3473 §8.4 ("failure is a value") applies to filesystem listings, not only command argv. #3885 (B5) gave roadmap analyze, gap-checker, and init's JSON bundles a context_read_error / phase_dir_read_error string naming an unreadable phase directory — but the underlying has_context / hasContext boolean stayed false either way, so a consumer branching on that boolean alone still cannot tell "genuinely no context file" from "could not read the directory at all." This closes that gap with a typed signal, reusing ADR-3180's existing frozen SCOPE enum rather than a new vocabulary.

findContextMdIn (src/planning-workspace.cts) now reports its own scope. Called with a directory path, it returns { file, files, scope } instead of a bare filename-or-null, and never throws — an unreadable directory reports scope: 'unreadable' (previously it threw, forcing every caller to hand-roll its own try/catch); a genuinely absent directory (ENOENT) reports scope: 'complete', the same "real empty" answer as today. The array-input call form (an already-read listing) is unchanged.

Five downstream call sites gain an additive scope field, none renamed or removed: roadmap analyze's AnalyzePhase.context_scope, gap-checker's phase_dir_scope, and init's context_scope on all three JSON bundles (init plan-phase, init phase-op, init manager) — including cmdInitManager, whose own read failure previously vanished into a bare empty catch {} with no signal of any kind. getPhaseFileStats (src/core-utils.cts) — the shared listing owner behind roadmap analyze and init's has_context — no longer lets its own failed read get masked by an unrelated, already-successful scanPhasePlans scope on the same phase directory.

Known limits:

  • context_read_error / phase_dir_read_error's message text is now a fixed "Could not read phase directory <path>" rather than embedding the underlying OS errno text — findContextMdIn's directory-string form reports only the SCOPE discriminator, not the raw caught error. The field's presence and type are unchanged; only its message detail is coarser than before #4014.
  • init.cts's three call sites call findContextMdIn for the scope signal and then still run their own, pre-existing fs.readdirSync on the same path for the rest of their output — an intentional, additive-only choice to avoid altering already-complex failure control-flow at those sites, not a performance optimization.

4836. Graphify CLI Preferred for Planner and Researcher Graph Queries

Purpose: msd-planner gets one knowledge-graph query per phase and msd-phase-researcher gets two or three. That single shot decides which modules the plan treats as related, and therefore how tasks are ordered into waves. It was spent on the built-in reader, which seeds by case-insensitive substring match over a node's label and description and then expands a hardcoded two hops — so the phase "User Authentication" seeds on author, authoring, and unauthorized with exactly the same weight as authenticate, and when the inflated result exceeds --budget the trimmer drops edges by confidence tier. The graphify CLI, already a hard dependency of /msd-graphify build, ranks seeds (IDF weighting, trigram fuzzy matching) and applies context filters before traversal.

Both prompts now prefer the CLI and fall back to the built-in reader. The branch is command -v graphify, the same degradation shape the repo already uses for Context7 → ctx7 in references/research-documentation-lookup.md. No new config key: a graph.json can only exist if graphify update . ran, which requires the binary, so binary presence is a self-satisfying gate. The fallback covers edge cases — a CI checkout with a committed graph, a binary since removed — not the common path. No new tool grant either: both agents already have Bash.

The planner additionally runs graphify affected. The reference states its own goal as "which subsystems may be affected by changes in this phase", which is literally reverse traversal by relation. The built-in reader only approximates it with undirected two-hop expansion, and has no equivalent verb, so affected is skipped on the fallback path.

msd-tools graphify status now returns graph_path. The CLI takes the graph location as --graph, and the prompts must not re-derive .planning/graphs/graph.json for it — that would point the CLI at a non-existent local mirror in exactly the umbrella multi-repo setup graphify.graph_path (#1825) exists to serve. status already resolves the override, so it now reports the absolute path it resolved, on both the graph-present and the graph-missing branch. For the same reason the presence gate in both prompts is now the status call itself rather than a bare ls of the default location.

Known limits:

  • The two paths return different shapes. graphify query emits prose and has no --json flag; msd-tools graphify query emits JSON with per-edge confidence tiers and budget_met/budget_estimate. Both are consumed by a model, and nothing machine-parses this block, but the prompts now say so explicitly instead of implying a stable shape.
  • --budget means different things on the two paths — rendered output on the CLI, estimated payload bytes in the built-in reader (#2738). Same flag name, different unit.
  • With graphify absent from PATH the fallback runs and the injected graph context is byte-identical to before.

Generated by scripts/gen-features.cjs — add a fragment under docs/features/ and run --write.