* enhance(#4836): prefer the graphify CLI for planner and researcher graph queries The planner gets one knowledge-graph query per phase and the researcher two or three, and that single shot decides which modules the plan treats as related — and therefore how tasks are ordered into waves. It was spent on the built-in reader, which seeds by case-insensitive substring match over a node's label and description and then expands a hardcoded two hops. The phase "User Authentication" seeds on `author`, `authoring` and `unauthorized` with the same weight as `authenticate`, and when the inflated payload exceeds `--budget` the trimmer drops edges by confidence tier — so the highest-confidence tier can be discarded to fit a payload that bad seeding inflated in the first place. The graphify CLI is already a hard dependency of /gsd-graphify build, and it ranks seeds (IDF weighting, trigram fuzzy matching) and applies context filters before traversal. Both prompts now prefer it and fall back to the built-in reader, branching on `command -v graphify` — the same degradation shape the repo already uses for Context7 to ctx7. Binary presence is a self-satisfying gate: a graph can only exist if the binary built it, so the fallback covers edge cases (a CI checkout with a committed graph, a binary since removed), not the common path. No new config key and no new tool grant — both agents already have Bash. The planner additionally runs `graphify affected`. The reference states its own goal as "which subsystems may be affected by changes in this phase", which is literally reverse traversal by relation; the built-in reader only approximates it with undirected two-hop expansion and has no equivalent verb, so `affected` is skipped on the fallback path. `graphify status` now reports `graph_path`, the resolved absolute graph location, on both the present and the missing branch. The CLI takes the graph location as `--graph`, and the prompts must not re-derive `.planning/graphs/graph.json` for it: that would point the CLI at a non-existent local mirror in exactly the umbrella multi-repo setup `graphify.graph_path` (#1825) exists to serve. For the same reason the presence gate in both prompts is now the `status` call itself rather than a bare `ls` of the default location, which was already blind to the override. Known limit, stated in both prompts rather than implied: the two paths return different shapes. `graphify query` emits prose and has no `--json` flag; the built-in emits JSON with per-edge confidence tiers and budget_met/budget_estimate. `--budget` also counts rendered output on one and estimated payload bytes on the other (#2738) — same flag name, different unit. Both are read by a model and nothing machine-parses the injected block. With graphify absent from PATH the injected context is byte-identical to before. Closes #4836 Emitted-Drift-Ack-Growth: gsd-phase-researcher.md — the CLI-first branch, the reason it is preferred, and the output-shape warning are the deliverable; a pointer to a part would not be read at the decision point. Emitted-Drift-Ack-Growth: gsd-planner.md — one sentence in the load_graph_context step pointer, so it stops naming the default graph path the reference no longer assumes. * docs(#4836): record the CLI-first graph query in the planner and researcher entries * chore(#4836): add changeset fragment * enhance(#4836): name the full domain word in the planner's query-term examples The reference's own example — phase "User Authentication" → term "auth" — is the exact collision the CLI-first path exists to avoid, and it stays a collision whenever the fallback path runs, since that path matches the term as a substring of label and description. * fix(#4836): surface graph_path on the unparseable-graph status branch graphifyStatus() returned graph_path on the exists:true and exists:false outcomes but not on the third, error, outcome (graph.json present but unparseable). The planner/researcher prompts gate CLI-first dispatch on exists, not on this outcome, so a corrupt graph file made them fall through to the CLI-first branch with the literal <graph> placeholder and no real path to substitute. * docs(#4836): note graph_path's trust boundary at the --graph interpolation graph_path is reflected verbatim into a double-quoted --graph argument the agent executes via Bash. It comes from graphify.graph_path, a config surface already trusted elsewhere, so this isn't a new trust boundary -- but it is a new injection site (no --graph flag existed on this call before). One-line caution for anyone hardening this later. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
48 lines
3.9 KiB
Markdown
48 lines
3.9 KiB
Markdown
# Planner — Load Graph Context
|
|
|
|
> Loaded by `gsd-planner` at the `load_graph_context` step.
|
|
|
|
Check for a knowledge graph and read its freshness in one call. `status` resolves the
|
|
graph through `graphify.graph_path`, so it is also the presence gate — a bare
|
|
`ls` of the default location misses an umbrella graph shared across sibling repos:
|
|
|
|
```bash
|
|
_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif command -v gsd-tools >/dev/null 2>&1; then GSD_TOOLS="$(command -v gsd-tools)"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd-tools is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi
|
|
gsd_run graphify status
|
|
```
|
|
|
|
If `exists` is `false`, continue without graph context — skip the rest of this step.
|
|
|
|
If the status response has `stale: true`, note for later: "Graph is {age_hours}h old -- treat semantic relationships as approximate." Include this annotation inline with any graph context injected below.
|
|
|
|
The same response carries `graph_path` — the resolved graph location. Substitute it for `<graph>` below. `graph_path` comes from `graphify.graph_path` in `.planning/config.json`, a config surface already trusted elsewhere; if it ever carried attacker-controlled content, the literal double-quoted substitution below would need escaping.
|
|
|
|
Query the graph for phase-relevant dependency context (single query per D-06). Prefer the `graphify` CLI when it is on PATH; fall back to the built-in reader otherwise:
|
|
|
|
```bash
|
|
if command -v graphify >/dev/null 2>&1; then
|
|
graphify query "<phase-goal-keyword>" --graph "<graph>" --budget 2000
|
|
graphify affected "<phase-goal-keyword>" --graph "<graph>" --depth 2
|
|
else
|
|
gsd_run graphify query "<phase-goal-keyword>" --budget 2000
|
|
fi
|
|
```
|
|
|
|
Why the CLI is preferred: it ranks seeds (IDF weighting, fuzzy matching) and applies context filters before traversal, where the built-in reader seeds by case-insensitive substring over label and description — so a term like "auth" seeds equally on `author` and `authorize` — and then expands a fixed two hops. `affected` answers "which subsystems may be affected by changes in this phase" directly, by reverse traversal; it has no built-in equivalent, so the fallback path runs the query alone.
|
|
|
|
The two paths return **different shapes**: the CLI emits prose, the built-in emits JSON with per-edge confidence tiers and `budget_met`/`budget_estimate`. `--budget` caps rendered output on the CLI and estimated payload bytes in the built-in — same flag name, different unit. Read whichever you get; do not assume a stable shape and do not paste raw output into PLAN.md.
|
|
|
|
Use the keyword that best captures the phase goal. Prefer the full domain word over a
|
|
prefix of it — on the fallback path a prefix is matched as a substring, so "auth" also
|
|
seeds on `author` and `authoring`. Examples:
|
|
- Phase "User Authentication" -> query term "authentication"
|
|
- Phase "Payment Integration" -> query term "payment"
|
|
- Phase "Database Migration" -> query term "migration"
|
|
|
|
If the query returns related nodes, incorporate as dependency context for planning:
|
|
- Which modules/files are semantically related to this phase's domain
|
|
- Which subsystems may be affected by changes in this phase
|
|
- Cross-document relationships that inform task ordering and wave structure
|
|
|
|
If nothing comes back, continue without graph context.
|