* test(#2751): regression guard — no command-position bare gsd-tools calls Agents/workflows instructed bare `gsd-tools <verb>` invocations that fail with 'command not found' on a shim-only install (#725 fixed only the Codex conversion pipeline; the Claude-facing source shipped them verbatim). Adds a source-text guard (allow-test-rule: source-text-is-the-product) scanning agents/*.md + gsd-core/workflows/*.md for the operative shape `gsd-tools <verb> <arg>`, excluding command -v probes / resolver definitions, with a documented PROSE_ALLOWLIST for descriptive mentions that name the command without instructing literal invocation. A stale-allowlist check ensures entries stay real. RED first; source fix lands next commit. * fix(#2751): normalize bare gsd-tools calls to gsd_run in Claude-facing source The 12 command-position bare `gsd-tools <verb>` instructions across agents/ and gsd-core/workflows/ failed with 'command not found' on a shim-only install (no gsd-tools binary on PATH). #725 fixed this only for the Codex install-conversion pipeline; the Claude-facing SOURCE shipped the bare calls verbatim, and new ones kept accumulating (new-project.md:114 landed 12 days AFTER #725 closed). Rewrite each operative site to the portable `gsd_run` resolver that the same files already define (3-18x each) — a pure command-position token swap preserving all arguments, flags, --files, and surrounding prose. Every runtime now benefits from one source change instead of each needing its own converter. Touched sites (12): gsd-intel-updater (validate/snapshot/extract-exports), gsd-code-fixer (query commit), gsd-planner (learnings.query), gsd-project- researcher (websearch/research-plan/classify-confidence), gsd-phase-researcher (websearch/research-plan/classify-confidence), new-project (project-instruction- file / commit --files), new-milestone (commit --files). Preserves command -v gsd-tools probes, resolver-snippet definitions, and the 4 descriptive prose mentions that NAME the command without instructing invocation. RED @ 1bb12ba1. * chore(#2751): allow-test-rule issue ref + changeset fragment Add the (#2751) tracking ref to the source-text-is-the-product annotation per ADR-456, and the .changeset Fixed fragment (pr:0, backfilled post-PR). * fix(#2751): convert remaining command-position bare gsd-tools calls (verify-summary, windows, worktree, smart-entry, quick-tasks-append) Isolated adversarial review (Step 4) found the first pass missed genuine command-position bare calls because the regression test's hand-maintained 6-verb list silently false-passed verify-summary (the 'verify' branch matched the prefix then died on the hyphen) and omitted windows/worktree/smart-entry/ quick-tasks-append entirely. Convert these 8 additional operative sites across new-project.md, new-milestone.md, ship.md, execute-phase.md, progress.md, smart-entry.md, quick.md. * chore(#2751): backfill changeset PR number (2851) * fix(#2751): normalize allowlist paths to forward slashes — Windows path-separator false-flag The PROSE_ALLOWLIST is keyed by file:line using forward-slash paths, but path.relative() returns backslash separators on Windows, so the allowlist lookup failed and the 6 descriptive mentions were flagged as offenders on the windows-latest CI lane. Normalize rel to forward slashes before the lookup so the allowlist matches identically on every OS. --------- Co-authored-by: Test <test@example.com>
This commit is contained in:
5
.changeset/lively-finches-bark.md
Normal file
5
.changeset/lively-finches-bark.md
Normal file
@@ -0,0 +1,5 @@
|
||||
---
|
||||
type: Fixed
|
||||
pr: 2851
|
||||
---
|
||||
**Agents and workflows no longer instruct a bare `gsd-tools` that fails on a shim-only install** — command-position `gsd-tools` invocations in the shipped agent/workflow source are now the portable `gsd_run` resolver (already defined in those files), so they resolve the runtime-local shim on installs with no `gsd-tools` binary on PATH. Previously only the Codex install-conversion pipeline rewrote these; the Claude-facing source shipped them verbatim and failed with `command not found`. (#2751)
|
||||
@@ -456,7 +456,7 @@ For each finding in sorted order:
|
||||
|
||||
**If verification passed:**
|
||||
|
||||
Use `gsd-tools query commit` with conventional format (message first, then every staged file path):
|
||||
Use `gsd_run query commit` with conventional format (message first, then every staged file path):
|
||||
```bash
|
||||
_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif command -v gsd-tools >/dev/null 2>&1; then GSD_TOOLS="$(command -v gsd-tools)"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLAUDE_CONFIG_DIR:-$HOME/.claude}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd-tools is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
|
||||
gsd_run query commit \
|
||||
|
||||
@@ -123,7 +123,7 @@ All JSON files include a `_meta` object with `updated_at` (ISO timestamp) and `v
|
||||
}
|
||||
```
|
||||
|
||||
**exports constraint:** Array of ACTUAL exported symbol names extracted from `module.exports` or `export` statements. MUST be real identifiers (e.g., `"configLoad"`, `"stateUpdate"`), NOT descriptions (e.g., `"config operations"`). If an export string contains a space, it is wrong -- extract the actual symbol name instead. Use `gsd-tools intel extract-exports <file>` to get accurate exports.
|
||||
**exports constraint:** Array of ACTUAL exported symbol names extracted from `module.exports` or `export` statements. MUST be real identifiers (e.g., `"configLoad"`, `"stateUpdate"`), NOT descriptions (e.g., `"config operations"`). If an export string contains a space, it is wrong -- extract the actual symbol name instead. Use `gsd_run intel extract-exports <file>` to get accurate exports.
|
||||
|
||||
Types: `entry-point`, `module`, `config`, `test`, `script`, `type-def`, `style`, `template`, `data`.
|
||||
|
||||
@@ -255,7 +255,7 @@ gsd_run intel patch-meta .planning/intel/arch-decisions.json
|
||||
|
||||
### Step 6.5: Self-Check
|
||||
|
||||
Run: `gsd-tools intel validate`
|
||||
Run: `gsd_run intel validate`
|
||||
|
||||
Review the output:
|
||||
|
||||
@@ -267,7 +267,7 @@ This step is MANDATORY -- do not skip it.
|
||||
|
||||
### Step 7: Snapshot
|
||||
|
||||
Run: `gsd-tools intel snapshot`
|
||||
Run: `gsd_run intel snapshot`
|
||||
|
||||
This writes `.last-refresh.json` with accurate timestamps and hashes. Do NOT write `.last-refresh.json` manually.
|
||||
</execution_flow>
|
||||
|
||||
@@ -138,7 +138,7 @@ For each item where `fetch` is present, invoke the MCP tool matching `fetch.prov
|
||||
| `exa` | `mcp__exa__web_search_exa` with `fetch.query` |
|
||||
| `tavily` | `mcp__tavily__search` with `fetch.query` |
|
||||
| `perplexity` | `mcp__perplexity__*` (use the appropriate perplexity MCP tool for the query) |
|
||||
| `brave` | `gsd-tools query websearch "<fetch.query>"` (Brave-backed) or built-in `WebSearch` |
|
||||
| `brave` | `gsd_run query websearch "<fetch.query>"` (Brave-backed) or built-in `WebSearch` |
|
||||
| `firecrawl` | `mcp__firecrawl__scrape` with url (scrape kind) or `mcp__firecrawl__search` |
|
||||
| `websearch` | built-in `WebSearch` tool |
|
||||
| `webfetch` | built-in `WebFetch` tool |
|
||||
@@ -709,7 +709,7 @@ docker info 2>/dev/null | head -3
|
||||
|
||||
## Step 3: Execute Research Protocol
|
||||
|
||||
For each domain, use the `<tool_strategy>` seam (Steps A–D): build questions JSON, call `gsd-tools query research-plan`, run the indicated provider per item, then cache each digest. Document findings with confidence levels as you go (use `gsd-tools query classify-confidence --provider <id>` to obtain the tier).
|
||||
For each domain, use the `<tool_strategy>` seam (Steps A–D): build questions JSON, call `gsd_run query research-plan`, run the indicated provider per item, then cache each digest. Document findings with confidence levels as you go (use `gsd_run query classify-confidence --provider <id>` to obtain the tier).
|
||||
|
||||
## Step 4: Validation Architecture Research (if nyquist_validation enabled)
|
||||
|
||||
|
||||
@@ -726,7 +726,7 @@ Read the most recent milestone retrospective and cross-milestone trends. Extract
|
||||
</step>
|
||||
|
||||
<step name="inject_global_learnings">
|
||||
If `features.global_learnings` is `true`: run `gsd-tools query learnings.query --tag <tag> --limit 5` once per tag from PLAN.md frontmatter `tags` (or use the single most specific keyword). The handler matches one `--tag` at a time. Prefix matches with `[Prior learning from <project>]` as weak priors. Project-local decisions take precedence. Skip silently if disabled or no matches.
|
||||
If `features.global_learnings` is `true`: run `gsd_run query learnings.query --tag <tag> --limit 5` once per tag from PLAN.md frontmatter `tags` (or use the single most specific keyword). The handler matches one `--tag` at a time. Prefix matches with `[Prior learning from <project>]` as weak priors. Project-local decisions take precedence. Skip silently if disabled or no matches.
|
||||
</step>
|
||||
|
||||
<step name="gather_phase_context">
|
||||
|
||||
@@ -102,7 +102,7 @@ For each item where `fetch` is present, invoke the MCP tool matching `fetch.prov
|
||||
| `exa` | `mcp__exa__web_search_exa` with `fetch.query` |
|
||||
| `tavily` | `mcp__tavily__search` with `fetch.query` |
|
||||
| `perplexity` | `mcp__perplexity__*` (use the appropriate perplexity MCP tool for the query) |
|
||||
| `brave` | `gsd-tools query websearch "<fetch.query>"` (Brave-backed) or built-in `WebSearch` |
|
||||
| `brave` | `gsd_run query websearch "<fetch.query>"` (Brave-backed) or built-in `WebSearch` |
|
||||
| `firecrawl` | `mcp__firecrawl__scrape` with url (scrape kind) or `mcp__firecrawl__search` |
|
||||
| `websearch` | built-in `WebSearch` tool |
|
||||
| `webfetch` | built-in `WebFetch` tool |
|
||||
@@ -490,7 +490,7 @@ Orchestrator provides: project name/description, research mode, project context,
|
||||
|
||||
## Step 3: Execute Research
|
||||
|
||||
For each domain, use the `<tool_strategy>` seam (Steps A–D): build questions JSON, call `gsd-tools query research-plan`, run the indicated provider per item, then cache each digest. Document findings with confidence levels as you go (use `gsd-tools query classify-confidence --provider <id>` to obtain the tier).
|
||||
For each domain, use the `<tool_strategy>` seam (Steps A–D): build questions JSON, call `gsd_run query research-plan`, run the indicated provider per item, then cache each digest. Document findings with confidence levels as you go (use `gsd_run query classify-confidence --provider <id>` to obtain the tier).
|
||||
|
||||
## Step 4: Quality Check
|
||||
|
||||
|
||||
@@ -125,7 +125,7 @@ fi
|
||||
|
||||
When `USE_WORKTREES` is `false`, `ISOLATION` is forced to `none`: executors run sequentially on the main working tree. The per-plan decision below has no effect when worktrees are project-disabled.
|
||||
|
||||
`USE_WORKTREES` and `ISOLATION` are also reset for the run when `worktree base-check` detects the orchestrator HEAD has diverged from the worktree fork base (#683 — e.g. an unmerged milestone branch). This runs for **any** isolated run, not only Claude: fork-base divergence is a property of the repository, so it degrades a GSD-created worktree exactly as a harness-created one. The auto-degrade prints a one-line warning to stderr and falls through to the sequential path so executors do not hit the exit-42 worktree-branch-check halt. To restore parallel worktree execution, set `worktree.baseRef:"head"` in `.claude/settings.local.json` (or run `gsd-tools worktree set-baseref`) — this makes the fork base track the live HEAD instead of a fixed remote ref. The `worktree-branch-check` exit-42 guard inside each executor remains in place as a backstop.
|
||||
`USE_WORKTREES` and `ISOLATION` are also reset for the run when `worktree base-check` detects the orchestrator HEAD has diverged from the worktree fork base (#683 — e.g. an unmerged milestone branch). This runs for **any** isolated run, not only Claude: fork-base divergence is a property of the repository, so it degrades a GSD-created worktree exactly as a harness-created one. The auto-degrade prints a one-line warning to stderr and falls through to the sequential path so executors do not hit the exit-42 worktree-branch-check halt. To restore parallel worktree execution, set `worktree.baseRef:"head"` in `.claude/settings.local.json` (or run `gsd_run worktree set-baseref`) — this makes the fork base track the live HEAD instead of a fixed remote ref. The `worktree-branch-check` exit-42 guard inside each executor remains in place as a backstop.
|
||||
|
||||
Read context window size for adaptive prompt enrichment:
|
||||
|
||||
|
||||
@@ -426,8 +426,8 @@ Commit after writing.
|
||||
|
||||
**Synthesizer output self-heal (#222) — verify SUMMARY.md materialized:** The synthesizer's canonical output is `.planning/research/SUMMARY.md` on disk; its brief structured return (`## SYNTHESIS COMPLETE` plus a few `###` confirmation lines) is NOT the file content. A known LLM false-refusal (issue #222) sometimes makes the agent return the full SUMMARY.md document inline — fabricating a write restriction (e.g. "the runtime is blocking file writes") — instead of writing the file. Prompt hardening alone does not fully eliminate it, so the orchestrator MUST absorb the failure deterministically before spawning `gsd-roadmapper`:
|
||||
|
||||
1. Verify `.planning/research/SUMMARY.md` exists AND is substantive — non-empty, and free of any leftover `<!-- gsd:write-continue -->` continuation sentinel (which marks a truncated/incomplete write). You may validate with `gsd-tools verify-summary .planning/research/SUMMARY.md` — it exits 0 regardless, so check its JSON `passed` field (`"passed": false` means missing or invalid), not the process exit code. If it passes, continue normally.
|
||||
2. If it is MISSING or invalid AND the synthesizer's return message contains the FULL SUMMARY.md document — recognizable by the template's top-level markers `# Project Research Summary`, `## Key Findings`, `## Implications for Roadmap`, and `## Sources`, not merely the brief `## SYNTHESIS COMPLETE` confirmation — the false-refusal fired: write that returned document to `.planning/research/SUMMARY.md` with the Write tool, then commit ALL research artifacts the synthesizer owns (it commits on behalf of the four researchers) with `gsd-tools query commit "docs: complete project research" --files .planning/research/` unless they are already committed. Log `⚠ #222 self-heal: synthesizer returned SUMMARY.md inline without writing it; orchestrator persisted the file.`
|
||||
1. Verify `.planning/research/SUMMARY.md` exists AND is substantive — non-empty, and free of any leftover `<!-- gsd:write-continue -->` continuation sentinel (which marks a truncated/incomplete write). You may validate with `gsd_run verify-summary .planning/research/SUMMARY.md` — it exits 0 regardless, so check its JSON `passed` field (`"passed": false` means missing or invalid), not the process exit code. If it passes, continue normally.
|
||||
2. If it is MISSING or invalid AND the synthesizer's return message contains the FULL SUMMARY.md document — recognizable by the template's top-level markers `# Project Research Summary`, `## Key Findings`, `## Implications for Roadmap`, and `## Sources`, not merely the brief `## SYNTHESIS COMPLETE` confirmation — the false-refusal fired: write that returned document to `.planning/research/SUMMARY.md` with the Write tool, then commit ALL research artifacts the synthesizer owns (it commits on behalf of the four researchers) with `gsd_run query commit "docs: complete project research" --files .planning/research/` unless they are already committed. Log `⚠ #222 self-heal: synthesizer returned SUMMARY.md inline without writing it; orchestrator persisted the file.`
|
||||
3. If it is MISSING or invalid AND the return is only a brief confirmation (no full SUMMARY document to recover), the synthesizer genuinely failed — surface the error and stop; do NOT spawn `gsd-roadmapper` against a missing or incomplete SUMMARY.md.
|
||||
|
||||
This guarantees `gsd-roadmapper` (which lists SUMMARY.md as required reading) never runs against a missing or truncated SUMMARY.md.
|
||||
|
||||
@@ -111,7 +111,7 @@ elif [ -n "$OPENCODE_CONFIG_DIR" ] || [ -n "$OPENCODE_CONFIG" ]; then RUNTIME="o
|
||||
else RUNTIME="claude"; fi
|
||||
```
|
||||
|
||||
Set the instruction file variable via the shared runtime-name policy adapter (`gsd-tools query project-instruction-file`, backed by `getProjectInstructionFile` in `runtime-name-policy.cjs` — the single source of truth shared with `profile-output.cjs`):
|
||||
Set the instruction file variable via the shared runtime-name policy adapter (`gsd_run query project-instruction-file`, backed by `getProjectInstructionFile` in `runtime-name-policy.cjs` — the single source of truth shared with `profile-output.cjs`):
|
||||
```bash
|
||||
INSTRUCTION_FILE=$(gsd_run query project-instruction-file --runtime "$RUNTIME")
|
||||
```
|
||||
@@ -1146,8 +1146,8 @@ Commit after writing.
|
||||
|
||||
**Synthesizer output self-heal (#222) — verify SUMMARY.md materialized:** The synthesizer's canonical output is `.planning/research/SUMMARY.md` on disk; its brief structured return (`## SYNTHESIS COMPLETE` plus a few `###` confirmation lines) is NOT the file content. A known LLM false-refusal (issue #222) sometimes makes the agent return the full SUMMARY.md document inline — fabricating a write restriction (e.g. "the runtime is blocking file writes") — instead of writing the file. Prompt hardening alone does not fully eliminate it, so the orchestrator MUST absorb the failure deterministically before spawning `gsd-roadmapper`:
|
||||
|
||||
1. Verify `.planning/research/SUMMARY.md` exists AND is substantive — non-empty, and free of any leftover `<!-- gsd:write-continue -->` continuation sentinel (which marks a truncated/incomplete write). You may validate with `gsd-tools verify-summary .planning/research/SUMMARY.md` — it exits 0 regardless, so check its JSON `passed` field (`"passed": false` means missing or invalid), not the process exit code. If it passes, continue normally.
|
||||
2. If it is MISSING or invalid AND the synthesizer's return message contains the FULL SUMMARY.md document — recognizable by the template's top-level markers `# Project Research Summary`, `## Key Findings`, `## Implications for Roadmap`, and `## Sources`, not merely the brief `## SYNTHESIS COMPLETE` confirmation — the false-refusal fired: write that returned document to `.planning/research/SUMMARY.md` with the Write tool, then commit ALL research artifacts the synthesizer owns (it commits on behalf of the four researchers) with `gsd-tools query commit "docs: complete project research" --files .planning/research/` unless they are already committed. Log `⚠ #222 self-heal: synthesizer returned SUMMARY.md inline without writing it; orchestrator persisted the file.`
|
||||
1. Verify `.planning/research/SUMMARY.md` exists AND is substantive — non-empty, and free of any leftover `<!-- gsd:write-continue -->` continuation sentinel (which marks a truncated/incomplete write). You may validate with `gsd_run verify-summary .planning/research/SUMMARY.md` — it exits 0 regardless, so check its JSON `passed` field (`"passed": false` means missing or invalid), not the process exit code. If it passes, continue normally.
|
||||
2. If it is MISSING or invalid AND the synthesizer's return message contains the FULL SUMMARY.md document — recognizable by the template's top-level markers `# Project Research Summary`, `## Key Findings`, `## Implications for Roadmap`, and `## Sources`, not merely the brief `## SYNTHESIS COMPLETE` confirmation — the false-refusal fired: write that returned document to `.planning/research/SUMMARY.md` with the Write tool, then commit ALL research artifacts the synthesizer owns (it commits on behalf of the four researchers) with `gsd_run query commit "docs: complete project research" --files .planning/research/` unless they are already committed. Log `⚠ #222 self-heal: synthesizer returned SUMMARY.md inline without writing it; orchestrator persisted the file.`
|
||||
3. If it is MISSING or invalid AND the return is only a brief confirmation (no full SUMMARY document to recover), the synthesizer genuinely failed — surface the error and stop; do NOT spawn `gsd-roadmapper` against a missing or incomplete SUMMARY.md.
|
||||
|
||||
This guarantees `gsd-roadmapper` (which lists SUMMARY.md as required reading) never runs against a missing or truncated SUMMARY.md.
|
||||
|
||||
@@ -141,7 +141,7 @@ WINDOWS_OPEN=$(printf '%s' "$WINDOWS_STATUS" | jq -r '.ledger.open_count // 0' 2
|
||||
WINDOWS_WAIVED=$(printf '%s' "$WINDOWS_STATUS" | jq -r '.ledger.waived_count // 0' 2>/dev/null || echo 0)
|
||||
```
|
||||
|
||||
Render `Open Windows` only when `$WINDOWS_OPEN` is greater than `0` (or `$WINDOWS_WAIVED` is greater than `0`, so an auditable deferral history remains visible). Phrase: `{WINDOWS_OPEN} open, {WINDOWS_WAIVED} waived — resolves with /gsd:ship gate; inspect via gsd-tools windows status`. The ledger is cross-phase; the count is the project total, not the current phase's.
|
||||
Render `Open Windows` only when `$WINDOWS_OPEN` is greater than `0` (or `$WINDOWS_WAIVED` is greater than `0`, so an auditable deferral history remains visible). Phrase: `{WINDOWS_OPEN} open, {WINDOWS_WAIVED} waived — resolves with /gsd:ship gate; inspect via gsd_run windows status`. The ledger is cross-phase; the count is the project total, not the current phase's.
|
||||
|
||||
|
||||
## Active Debug Sessions
|
||||
|
||||
@@ -994,7 +994,7 @@ Use `date` from init:
|
||||
| ${quick_id} | ${DESCRIPTION} | ${date} | ${commit_hash} | [${quick_id}-${slug}](./quick/${quick_id}-${slug}/) |
|
||||
```
|
||||
|
||||
For a schema-safe append outside this workflow (e.g. from fast.md), `gsd-tools quick-tasks-append --task <text>` performs the equivalent write via the shared, schema-backed `appendQuickTaskRow` helper (#2133, ADR-2143 §3/§7).
|
||||
For a schema-safe append outside this workflow (e.g. from fast.md), `gsd_run quick-tasks-append --task <text>` performs the equivalent write via the shared, schema-backed `appendQuickTaskRow` helper (#2133, ADR-2143 §3/§7).
|
||||
|
||||
**7d. Update "Last activity" line:**
|
||||
|
||||
|
||||
@@ -139,14 +139,14 @@ Verify the work is ready to ship:
|
||||
```
|
||||
⚠ Broken-windows ship gate: WINDOWS.md has {WINDOWS_OPEN_COUNT} open window(s).
|
||||
Resolve each entry before shipping, or explicitly waive with a recorded reason:
|
||||
gsd-tools windows fixed <id> # defect resolved
|
||||
gsd-tools windows waive <id> "<reason>" # justified deferral (reason required)
|
||||
gsd_run windows fixed <id> # defect resolved
|
||||
gsd_run windows waive <id> "<reason>" # justified deferral (reason required)
|
||||
Then re-run /gsd:ship.
|
||||
```
|
||||
- **`WINDOWS_OPEN_COUNT` is `"?"`, empty, or non-numeric** → **fail closed and block** with `WINDOWS_SHIP_GATE_READ_FAILED` (the gate is strict equality to `0`; never ship on an unreadable ledger):
|
||||
```
|
||||
⚠ Broken-windows ship gate: could not read open_count from .planning/WINDOWS.md.
|
||||
Inspect the file or run `gsd-tools windows status --raw` to diagnose. The ledger
|
||||
Inspect the file or run `gsd_run windows status --raw` to diagnose. The ledger
|
||||
may be malformed; fix it before shipping (an unparseable ledger is a broken window).
|
||||
```
|
||||
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
<purpose>
|
||||
GSD smart entry — the state-aware front door. Detect the current project situation via `gsd-tools smart-entry --json`, present a short menu of the right next actions, and dispatch to exactly one existing GSD command. This is a launcher/router only; it never does the work itself.
|
||||
GSD smart entry — the state-aware front door. Detect the current project situation via `gsd_run smart-entry --json`, present a short menu of the right next actions, and dispatch to exactly one existing GSD command. This is a launcher/router only; it never does the work itself.
|
||||
|
||||
This is a *menu* front door, not a second router. For in-project forward motion (planning → executing → verify-pending) the recommended action is `/gsd:progress --next`, which delegates to the single gated advancement engine (`workflows/next.md`: Route 0 resume-incomplete-phase + Gates 1-3). smart-entry adds value only where `--next` cannot reach: pre-project, remediation (paused/blocked/verify-failed), and lifecycle exits (idle-stranded/complete). See `docs/adr/1787-gsd-next-smart-entry.md`.
|
||||
</purpose>
|
||||
|
||||
194
tests/no-bare-gsd-tools-command-position.test.cjs
Normal file
194
tests/no-bare-gsd-tools-command-position.test.cjs
Normal file
@@ -0,0 +1,194 @@
|
||||
// allow-test-rule: source-text-is-the-product (#2751)
|
||||
'use strict';
|
||||
|
||||
// Regression guard for #2751: agents/*.md and gsd-core/workflows/*.md must not
|
||||
// instruct an agent to invoke a BARE `gsd-tools <verb> <args>` command. The bare
|
||||
// word fails with "command not found" on a shim-only install (no `gsd-tools`
|
||||
// binary on PATH) — every such instruction must use the `gsd_run` resolver the
|
||||
// same files already define. #725 fixed this only for the Codex install-
|
||||
// conversion pipeline; the Claude-facing SOURCE shipped the bare calls verbatim
|
||||
// until #2751 normalized them.
|
||||
//
|
||||
// A pure regex cannot perfectly distinguish an imperative ("Use `gsd-tools query
|
||||
// commit` to commit") from a descriptive mention ("`gsd-tools query commit`
|
||||
// returns an envelope") — both contain the same command phrase. So this guard
|
||||
// scans for `gsd-tools <verb> <arg>` (the operative shape — a verb followed by
|
||||
// its arguments) and subtracts a documented ALLOWLIST of known descriptive
|
||||
// mentions that NAME the command without instructing literal invocation. Any
|
||||
// site NOT in the allowlist is a new operative bare call and fails the gate.
|
||||
//
|
||||
// Verb coverage is derived LIVE from the gsd-tools host-command router
|
||||
// (`gsd-core/bin/gsd-tools.cjs`) by reading the `'verb': routeHandler` / `verb:
|
||||
// routeHandler` entries. This is deliberate — the original #2751 guard used a
|
||||
// hand-maintained 6-verb list that silently false-passed `verify-summary` (the
|
||||
// `verify` branch matched the prefix, then died on the hyphen), and missed
|
||||
// `windows`, `worktree`, `smart-entry`, and `quick-tasks-append` entirely.
|
||||
// Deriving the set from the router means a new verb registered there is covered
|
||||
// here the moment it lands — no second list to keep in sync.
|
||||
//
|
||||
// Source-text guard: the deployed contract IS the markdown text the runtime
|
||||
// loads. Scans FULL file text (the real hits lived in inline prose/table cells,
|
||||
// not fenced bash blocks).
|
||||
|
||||
const { test } = require('node:test');
|
||||
const assert = require('node:assert/strict');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
const ROOT = path.resolve(__dirname, '..');
|
||||
const ROUTER_PATH = path.join(ROOT, 'gsd-core', 'bin', 'gsd-tools.cjs');
|
||||
const SCAN_DIRS = ['agents', path.join('gsd-core', 'workflows')];
|
||||
|
||||
// Derive the verb set the bare-call guard matches against. Most top-level
|
||||
// verbs live in the host-command router table as `'verb': routeHandler` entries
|
||||
// (~70); this reads those dynamically so new router verbs are covered the moment
|
||||
// they land. A handful of verbs are dispatched as FAMILIES (their own
|
||||
// `command === 'verb'` arm, not a route-table entry): `query` (line ~2876),
|
||||
// `intel`, `verify`, and `graphify`. These are stable, documented families, so
|
||||
// they are supplemented explicitly here rather than parsed from the help string
|
||||
// (whose prose mixes real verbs with English words like "for"/"output"/"working",
|
||||
// producing noise). If a family verb is ever promoted into the route table the
|
||||
// union dedupes harmlessly; if a NEW family verb is added it must be added here.
|
||||
//
|
||||
// Sorted longest-first so a hyphenated verb (`verify-summary`) is preferred over
|
||||
// its prefix (`verify`) — the exact ordering bug that let `verify-summary` slip
|
||||
// past a fixed 6-verb list during the first #2751 pass.
|
||||
const FAMILY_DISPATCHED_VERBS = ['query', 'intel', 'verify', 'graphify'];
|
||||
function readRouterVerbs() {
|
||||
const src = fs.readFileSync(ROUTER_PATH, 'utf8');
|
||||
const re = /(?:'([a-z][a-z-]*)'|([a-z][a-z-]*))\s*:\s*route[A-Z]\w*/g;
|
||||
const verbs = new Set(FAMILY_DISPATCHED_VERBS);
|
||||
let m;
|
||||
while ((m = re.exec(src)) !== null) verbs.add(m[1] || m[2]);
|
||||
return [...verbs].sort((a, b) => b.length - a.length);
|
||||
}
|
||||
|
||||
const VERBS = readRouterVerbs();
|
||||
|
||||
// Operative shape: `gsd-tools <known-verb>` followed by whitespace and a real
|
||||
// argument (the verb is NOT the close of a code span — there is a real arg
|
||||
// after it). Restricting to KNOWN verbs (vs any `[a-z-]+` token) avoids
|
||||
// false-flagging English prose like "gsd-tools through the ...".
|
||||
const BARE_COMMAND_RE = new RegExp(
|
||||
String.raw`(?:^|[^./A-Za-z0-9_-])gsd-tools\s+(` + VERBS.join('|') + String.raw`)\s+[^\s` + '`' + String.raw`]`
|
||||
);
|
||||
|
||||
// Lines that legitimately embed `gsd-tools <verb> <arg>` but are descriptive, not
|
||||
// command-position: they NAME the command (in prose / parenthetical examples /
|
||||
// return-envelope descriptions) rather than instructing an agent to run the bare
|
||||
// word. Keyed `file:line` so a rewording that moves the mention forces a conscious
|
||||
// allowlist update rather than silently passing.
|
||||
//
|
||||
// Each entry MUST carry a one-line reason; the test prints the allowlist on
|
||||
// failure so a reviewer can see exactly what is sanctioned.
|
||||
const PROSE_ALLOWLIST = [
|
||||
{ file: 'agents/gsd-executor.md', line: 791, reason: 'describes the SDK return envelope of `gsd-tools query commit`; not an instruction to run the bare word' },
|
||||
{ file: 'agents/gsd-phase-researcher.md', line: 33, reason: 'package-legitimacy provenance rule names the command as the source of an OK verdict; descriptive' },
|
||||
{ file: 'agents/gsd-roadmapper.md', line: 624, reason: 'parenthetical "e.g." naming SDK queries a user *could* run; not an agent instruction' },
|
||||
{ file: 'agents/gsd-intel-updater.md', line: 40, reason: 'cross-platform note names the `gsd-tools intel <subcommand>` CLI surface descriptively ("CLI invocations go through..."); not an agent instruction' },
|
||||
{ file: 'gsd-core/workflows/execute-plan.md', line: 387, reason: 'describes the downstream SDK validation step (`validated downstream by ...`); names the mechanism, does not instruct the agent to type it' },
|
||||
{ file: 'agents/gsd-research-synthesizer.md', line: 65, reason: 'a code comment inside a fenced block explaining what the commit step loads (`# Planning config loaded via gsd-tools query ...`); descriptive, not an invocation — and explicitly names gsd-tools.cjs as the alternative' },
|
||||
];
|
||||
|
||||
// Resolver-snippet definition lines / probes that must never be flagged. A line
|
||||
// is a resolver/probe definition when it assigns the shim name, probes PATH, or
|
||||
// assigns GSD_TOOLS / defines the gsd_run function — NOT merely when it contains
|
||||
// the substring `gsd-tools.cjs` (which would blanket-exempt any prose that
|
||||
// happens to name the file).
|
||||
const EXCLUSION_RE = /_GSD_SHIM_NAME\s*=|command -v gsd-tools|\bGSD_TOOLS\s*=|gsd_run\s*\(\)\s*\{/;
|
||||
|
||||
function collectMdFiles(dir) {
|
||||
const results = [];
|
||||
let entries;
|
||||
try { entries = fs.readdirSync(dir, { withFileTypes: true }); } catch { return results; }
|
||||
for (const entry of entries) {
|
||||
const full = path.join(dir, entry.name);
|
||||
if (entry.isDirectory()) {
|
||||
results.push(...collectMdFiles(full));
|
||||
} else if (entry.isFile() && entry.name.endsWith('.md')) {
|
||||
results.push(full);
|
||||
}
|
||||
}
|
||||
return results;
|
||||
}
|
||||
|
||||
function findBareCommandPositionCalls() {
|
||||
const offenders = [];
|
||||
for (const scanDir of SCAN_DIRS) {
|
||||
const abs = path.join(ROOT, scanDir);
|
||||
for (const file of collectMdFiles(abs)) {
|
||||
// Normalize to forward slashes so the file:line PROSE_ALLOWLIST keys match
|
||||
// identically on every OS — path.relative() returns backslash separators on
|
||||
// Windows, which would defeat the allowlist lookup and false-flag the 6
|
||||
// descriptive mentions (#2751 Windows CI failure).
|
||||
const rel = path.relative(ROOT, file).split(path.sep).join('/');
|
||||
const lines = fs.readFileSync(file, 'utf8').split(/\r?\n/);
|
||||
for (let i = 0; i < lines.length; i += 1) {
|
||||
const line = lines[i];
|
||||
if (EXCLUSION_RE.test(line)) continue;
|
||||
const match = line.match(BARE_COMMAND_RE);
|
||||
if (!match) continue;
|
||||
const loc = `${rel}:${i + 1}`;
|
||||
const allowed = PROSE_ALLOWLIST.find((a) => a.file === rel && a.line === i + 1);
|
||||
if (allowed) continue;
|
||||
offenders.push({ loc, verb: match[1], text: line.trim() });
|
||||
}
|
||||
}
|
||||
}
|
||||
return offenders;
|
||||
}
|
||||
|
||||
test('verb set was derived from the router (guards against a silent extraction regression)', () => {
|
||||
// If the router is refactored so the `routeHandler` pattern no longer matches,
|
||||
// VERBS would be empty and the bare-call guard below would false-pass
|
||||
// everything. Sanity-bound it: a healthy router exposes many host verbs.
|
||||
assert.ok(
|
||||
VERBS.length > 40,
|
||||
`Expected to derive >40 verbs from ${path.relative(ROOT, ROUTER_PATH)}, got ${VERBS.length}. ` +
|
||||
'If the router table changed shape, update readRouterVerbs() so the guard keeps coverage.'
|
||||
);
|
||||
// Longest-first ordering is what lets verify-summary win over verify.
|
||||
const sample = ['verify-summary', 'verify', 'query', 'intel', 'graphify', 'windows', 'worktree', 'smart-entry', 'commit', 'check'];
|
||||
for (const v of sample) {
|
||||
assert.ok(VERBS.includes(v), `expected verb '${v}' in derived set (router/family drift?)`);
|
||||
}
|
||||
});
|
||||
|
||||
test('no command-position bare gsd-tools <verb> survives in agents/ or workflows/ (#2751)', () => {
|
||||
const offenders = findBareCommandPositionCalls();
|
||||
assert.strictEqual(
|
||||
offenders.length,
|
||||
0,
|
||||
'Bare `gsd-tools <verb> <args>` command-position calls must use the `gsd_run` ' +
|
||||
'resolver (they fail with "command not found" on a shim-only install — #2751). ' +
|
||||
`Found ${offenders.length} offender(s):\n` +
|
||||
offenders.map((o) => ` ${o.loc} [${o.verb}] ${o.text}`).join('\n') +
|
||||
'\n\nIf a hit is a descriptive prose mention (not an instruction to run the bare ' +
|
||||
'word), add it to PROSE_ALLOWLIST in this test with a reason.'
|
||||
);
|
||||
});
|
||||
|
||||
test('every PROSE_ALLOWLIST entry still matches a real gsd-tools mention (no stale allowlist entries)', () => {
|
||||
// An allowlist entry that no longer matches anything is stale — it was either
|
||||
// fixed (remove it) or the line moved (update it). Either way it must not linger.
|
||||
const stale = [];
|
||||
for (const entry of PROSE_ALLOWLIST) {
|
||||
const file = path.join(ROOT, entry.file);
|
||||
let lines;
|
||||
try { lines = fs.readFileSync(file, 'utf8').split(/\r?\n/); } catch (e) {
|
||||
stale.push({ ...entry, problem: 'file missing' });
|
||||
continue;
|
||||
}
|
||||
const line = lines[entry.line - 1];
|
||||
if (!line || !BARE_COMMAND_RE.test(line) || EXCLUSION_RE.test(line)) {
|
||||
stale.push({ ...entry, problem: 'line no longer matches a bare gsd-tools mention', actual: line ? line.trim() : '(line absent)' });
|
||||
}
|
||||
}
|
||||
assert.strictEqual(
|
||||
stale.length,
|
||||
0,
|
||||
'PROSE_ALLOWLIST has stale entries (the mentioned line no longer carries a bare ' +
|
||||
'`gsd-tools <verb> <arg>` mention). Remove or update them:\n' +
|
||||
stale.map((s) => ` ${s.file}:${s.line} — ${s.problem}${s.actual ? ` (actual: ${s.actual.slice(0, 80)})` : ''}`).join('\n')
|
||||
);
|
||||
});
|
||||
Reference in New Issue
Block a user