Files
msd-core/src/intel.cts
Tom Boucher 929e02cb2c enhance(#3885): no silent swallow, and no verdict manufactured from dropped data (#3925)
* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict

ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking
answer. Three families do exactly that today; this commit pins each one RED.

Measured on this tree, 2026-08-27:

  intel query, .planning/intel/file-roles.json nested 12000 deep
    -> exit 1, "Error: Maximum call stack size exceeded"
       searchJsonEntries / matchesInValue carry no depth parameter at all.
       The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage
       (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage
       never received it.

  same fixture nested 48 and 49 deep
    -> both return total=1 at exit 0, truncated=undefined
       Nothing distinguishes "searched to the bottom" from "stopped looking".

  query phase-plan-index, a plan whose depends_on names an unresolvable token
    -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it
                  in wave 1"]
       The token is never mentioned. computeDependencyLevels drops the edge
       with `if (!resolvedDep) continue;`, every plan becomes a root, and the
       tool then reports the author's correct wave: as the thing that is wrong.

  countPhasePlansAndSummaries with fs.readdirSync throwing EACCES
    -> hasContext:false, indistinguishable from a phase that simply has no
       CONTEXT.md. context_read_error is undefined.

The shapes these tests assert against, chosen here so the implementation has a
target rather than inventing one later: `truncated: boolean` on the intel query
result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and
`context_read_error: string | null` per analyzed phase.

Deliberately green, and they must stay that way — each stops the fix from
over-firing:

  depth 48 is found and NOT flagged truncated (the ceiling is inclusive)
  a shallow miss reports no truncation           (noise control, N1)
  10,000 siblings at depth 2 are unaffected      (the bound is DEPTH, N2)
  a genuine wave: mismatch on a fully-resolved DAG still warns (N3)
  a genuinely missing directory is absent, not an error
  the emitted depends_on display mapping still passes an unresolved token
    through verbatim — already pinned by the existing #3785 test, so no
    duplicate was added

T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the
real CLI and reads the emitted JSON, because a unit assertion on
computeDependencyLevels would have passed throughout #3427's life.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3885): no silent swallow, and no verdict manufactured from dropped data

Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed
into an output that reads as authoritative.

The recursion bound, restored but NOT verbatim (src/intel.cts)

  MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and
  matchesInValue, which carried no depth parameter at all. The bound existed in
  the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the
  surviving .cts lineage never received it — §8.3's "a consolidation may not
  delete an invariant along with the surface that held it", demonstrated.

  Measured before: a .planning/intel file nested 12000 deep exits 1 with
  "Error: Maximum call stack size exceeded". Reachable from a project document.

  The original returned a bare `false` at the ceiling. Restoring that verbatim
  would trade a crash for a silent "no match" when the truth is "I stopped
  looking" — the same class this epic exists to close, and ADR-3473 Decision 4
  forbids it. So the bound carries a truncation signal:

    nesting 47 -> found,     truncated false
    nesting 48 -> found,     truncated false      (the ceiling is inclusive)
    nesting 49 -> not found, truncated TRUE
    nesting 12000 -> exit 0, truncated TRUE, no RangeError

  A shallow document that simply has no match reports truncated FALSE — the
  flag means "I stopped early", never "I found nothing", or it would be noise.
  The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected.

The dropped edge is named, and stops being blamed on the author (src/phase.cts)

  computeDependencyLevels dropped every unresolvable depends_on token with a
  bare `continue`. Each drop makes a plan a root, so the whole phase collapses
  to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as
  the thing that was wrong.

  Before:
    warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in
                wave 1"]
  After:
    warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not
                resolve to any plan in this phase — edge dropped, wave placement
                for this plan may be unreliable"]

  The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG
  and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId
  stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is
  deliberately not built here. The emitted depends_on display mapping still
  passes an unresolved token through verbatim (#3785).

No artifact from failed inputs (gsd-core/workflows/review.md, #3352)

  A failed lane leaves no result file, so "every lane failed" is exactly "the
  aggregate JSONL has zero lines" — the gate condition already existed as a
  byproduct. REVIEWS.md is no longer written in that case, and the commit step
  is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT
  counted as a failure. Per-lane output and non-empty .err are preserved to
  .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that
  the lanes failed at all; the commit step names one file, never a glob, so the
  diagnostics are not swept in.

Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2)

  Four callers collapsed an EACCES on a phase directory into [] and reported
  hasContext:false — byte-identical to a phase that simply has no CONTEXT.md.
  Each now names the directory it could not read. A genuinely missing directory
  stays absent rather than becoming an error, which is what keeps the fix from
  over-firing.

Fatal errno folded into a retry set: audited, no defect found

  Reported as a verified negative rather than padded with a change.
  withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776;
  atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by
  construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they
  return or rethrow the final error rather than swallowing it. estimate-cli's
  sole caller surfaces that rethrow as write_error in its JSON output.
  Manufacturing a diff to make the checkbox look worked-on is the Goodhart
  outcome Decision 6 exists to prevent.

Disclosed: R46 (the commit step names one file, never a glob) is a real
regression guard but is NOT independently failing-first — the commit fence is
byte-identical pre- and post-fix, so it only fails pre-fix through its shared
extraction dependency. Recorded rather than claimed as fail-first.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence

Two review findings, both real, both in my own change.

An isolated adversarial review found the evidence-preservation block never
checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in
a SEPARATE fenced block. A disk-full or unwritable phase directory therefore
still destroyed the only copy of the failed lanes' output — reintroducing the
exact #3352 data loss this item exists to stop, inside the fix for it.

Preservation and cleanup are now one block, because each fenced block is a
separate execution and a shell variable cannot carry between them. mkdir -p and
each cp are exit-checked; cleanup runs only when preservation succeeded, and a
failure warns naming the intact run directory. "Nothing to preserve" is not a
failure and still cleans up. Driven three ways: success removes run_dir, failure
leaves it intact with the warning, nothing-to-preserve removes it. The failure is
induced by a file-vs-directory conflict rather than chmod 0o000, which root
bypasses.

The new unresolved-depends_on warning embedded a user-authored token verbatim:

  warnings: ["Plan 03-02: depends_on token \"evil
  Plan 03-01: FORGED WARNING\" does not resolve ..."]

The JSON wire form is safe, and the security reviewer judged it non-exploitable
for that reason. It is escaped anyway through formatDiagnosticToken — the helper
#3884 added one phase earlier for exactly this class. warnings[] is an array a
consumer naturally prints line by line, and not reusing the sibling fix is the
generative-fix-divergence shape this epic exists to close. The same treatment is
applied to context_read_error / phase_dir_read_error, which embed a phase
directory path a repository can choose, and to the fs error message, which
echoes the raw path itself.

Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow
array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per
negative space N2, disclosed rather than left to be discovered.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot"

Blocker from the round-2 isolated review, and it is my own inconsistency:
this phase applied "unreadable is not absent" to phase directories and left it
broken in the file it was already editing.

  chmod 000 .planning/intel/file-roles.json
  gsd-tools intel query <term>
  -> {"matches":[],"total":0,"truncated":false}  exit 0

safeReadJson swallowed every read failure and returned null, so an EACCES was
byte-indistinguishable from an absent file AND from a genuine no-match. Now it
separates three states: ENOENT stays silently absent, because not every project
has every intel file and intelQuery loops over all of them expecting misses;
EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel
file previously read as "no matches" too — same defect, same fix.

Threading that outcome through the other three callers found something worse
than the reported case. intelDiff returned no_baseline:true for a corrupt or
unreadable snapshot — not a silent failure but an actively FALSE verdict, telling
the caller they never took a snapshot when they did. That is §8.5's headline
case, so it is fixed and tested rather than noted. intelStatus and
intelApiSurface collapsed the same way; intelApiSurface additionally printed a
"not yet populated" banner that was simply untrue.

Every row is failing-first, including the absent-file ones — the field is new,
so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin
that the fix does not OVER-fire on the ordinary absent case, which is what would
turn this into noise on every project lacking an intel file. IO failure is
injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root
bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as
a test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): build the pathological intel fixture as text, not by stringifying a nested object

The remote runner came back red on Linux with two failures, both
T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on
macOS. The product was never at fault.

writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then
JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed
the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned.
Linux's container stack is smaller than macOS's, which is the whole of the
platform difference.

Measured, with the same document built as JSON TEXT so nothing in the building
process recurses:

  depth=100    rc=0 truncated=true
  depth=5000   rc=0 truncated=true
  depth=12000  rc=0 truncated=true
  depth=60000  rc=0 truncated=true

V8 parses this shape iteratively; only stringify recurses. The bound works at
every depth tried.

The fixture is now built by string concatenation. That is also the more faithful
input — a real deeply nested JSON document on disk is exactly what the bound
guards, where a stringified object was only ever a way to produce one.

The depth stays 12000. Lowering it would have made the test pass by weakening it
to accommodate a fixture bug, and 12000 is a legitimate pathological input the
product handles. T4 remains a genuine fail-first: rebuilt against the parent of
the commit that added the bound, the string-built depth-12000 fixture still
drives the CLI to rc=1 with "Error: Maximum call stack size exceeded".

A comment records why the fixture is text, so it is not "simplified" back into a
macOS-green / Linux-red test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3885): backfill the changeset PR number

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): normalize path separators before splicing into the workflow's bash

CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the
remote runner were all green.

  AssertionError: commit must name the single REVIEWS.md file; got:
    --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md

Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The
harness spliced an OS-native temp path into the extracted bash, and bash consumes
\U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke
RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory
survived — which is the other two assertions.

This is a fixture defect, not a product one, and that was checked rather than
assumed. In production the phase directory is toPosixPath-normalized at every
call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and
the run directory is created by "mktemp -d" running inside the bash block itself
(gsd-core/workflows/review.md:163), which emits POSIX-style output even under
Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it.

The file's pre-existing #3034 harness splices raw native paths too, but only ever
inside double-quoted assignments, so it never tripped this — my new harness
followed that convention faithfully into the one place where it does not hold.
Both now splice through toPosixPath from shell-command-projection, the
established seam, which is a no-op on POSIX and mirrors what production does.

No assertion was weakened. "commit must name the single REVIEWS.md file" and
"the run dir must still be destroyed" still assert exactly that; only how the
fixture supplies its path changed. Nothing is skipped on Windows — a t.skip()
here would have hidden the question of whether the exposure was real, which is
the question that mattered.

Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI
string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input
produces a byte-identical shape, proving the normalization is idempotent.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): stop the harness making the deleted run dir its own cwd

Windows shard 3/3 stayed red after the separator fix, on two assertions the
separator fix never touched:

  AssertionError: the run dir must still be destroyed
  AssertionError: nothing to preserve is not a failure — run dir must still be removed

The separators were a real bug and fixing them fixed the --files assertion. They
were not this bug, and two CI cycles went into the wrong axis before I stopped
converting path forms and looked at what the harness actually does.

runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's
working directory WAS the directory the block under test then removes with
rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified
locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live
process's working directory cannot be removed. So on Windows the directory
survived and both assertions failed, on macOS and Linux it vanished and they
passed. Nothing to do with slashes.

Harness-only. Production never cd's into the run directory; every reference is by
absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself
(gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged.

Fix: the child now runs with its cwd in an unrelated temp directory that the
block under test never deletes. Neither assertion was weakened, and nothing is
skipped on Windows — the tests in this file carry no platform guard and run
there unconditionally, which is how this surfaced at all.

Honest limit: the Windows failure mode cannot be reproduced on macOS, because
POSIX permits the very thing Windows refuses. The diagnosis is grounded in that
documented divergence and in the fact that only the Windows lane failed, but the
green outcome on windows-latest is unverified until CI runs it.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 04:12:47 -04:00

838 lines
29 KiB
TypeScript

/**
* lib/intel.cts -- Intel storage and query operations for GSD.
*
* Provides a persistent, queryable intelligence system for project metadata.
* Intel files live in .planning/intel/ and store structured data about
* the project's files, APIs, dependencies, architecture, and tech stack.
*
* All public functions gate on isCapabilityActive('intel', cwd) — the shared
* tri-state resolver (installed + surfaced + intel.enabled config key).
*
* ADR-457 build-at-publish: the hand-written bin/lib/intel.cjs collapsed
* to a TypeScript source of truth. Behaviour is preserved byte-for-behaviour
* from the prior hand-written .cjs; only types are added.
*/
import fs from 'node:fs';
import path from 'node:path';
import crypto from 'node:crypto';
import { platformWriteSync, platformReadSync, platformEnsureDir } from './shell-command-projection.cjs';
// eslint-disable-next-line @typescript-eslint/no-require-imports
import capabilityStateMod = require('./capability-state.cjs');
const { isCapabilityActive } = capabilityStateMod;
// eslint-disable-next-line @typescript-eslint/no-require-imports
import ioMod = require('./io.cjs');
const { formatDiagnosticToken } = ioMod;
// ─── Constants ───────────────────────────────────────────────────────────────
const INTEL_DIR = '.planning/intel';
const INTEL_FILES: Record<string, string> = {
files: 'file-roles.json',
apis: 'api-map.json',
deps: 'dependency-graph.json',
arch: 'arch-decisions.json',
stack: 'stack.json',
};
/**
* ADR-3473 §8.5 / #3885: recursion bound for the intel JSON search walk.
* Restored from the retired SDK lineage (`sdk/src/query/intel.ts` at `11918dcc3^`),
* lost in the ADR-0174 consolidation. Unlike the original, hitting the ceiling is
* NOT reported as a silent "no match" — the walk that stops early sets a
* `truncated` flag threaded back up to `intelQuery`'s result (Decision 4: a
* routine that discards an input says so). The bound is on DEPTH only; breadth
* (sibling count at any given depth) is unaffected.
*/
const MAX_JSON_SEARCH_DEPTH = 48;
// ─── Internal helpers ────────────────────────────────────────────────────────
/**
* Ensure the intel directory exists under the given planning dir.
*/
function ensureIntelDir(planningDir: string): string {
const intelPath = path.join(planningDir, 'intel');
platformEnsureDir(intelPath);
return intelPath;
}
/**
* Check whether intel is active (installed, surfaced, and config-enabled) for the project at cwd.
* Delegates to the shared tri-state capability resolver (isCapabilityActive) which honours the
* install profile, runtime surface, and activationKey (intel.enabled config gate).
*
* NOTE: planningDir is the legacy entry-point; cwd is derived as path.dirname(planningDir).
* Callers that have cwd directly may call isCapabilityActive('intel', cwd) themselves.
*
* INVARIANT: planningDir is always `<cwd>/.planning` (i.e. path.join(cwd, '.planning')).
* The intel-command-router always constructs planningDir as path.join(cwd, '.planning'),
* so path.dirname(planningDir) === cwd is guaranteed. If a workstream-aware planningDir
* were ever passed here, the dirname would be wrong — but no caller does that.
*/
function isIntelCapabilityActive(planningDir: string): boolean {
return isCapabilityActive('intel', path.dirname(planningDir));
}
interface DisabledResponse {
disabled: true;
message: string;
}
/**
* Return the standard disabled response object.
*/
function disabledResponse(): DisabledResponse {
return { disabled: true, message: 'Intel system disabled. Set intel.enabled=true in config.json to activate.' };
}
/**
* Resolve full path to an intel file.
*/
function intelFilePath(planningDir: string, filename: string): string {
return path.join(planningDir, 'intel', filename);
}
interface IntelData {
_meta?: {
updated_at?: string;
version?: number;
[key: string]: unknown;
};
entries?: Record<string, unknown>;
[key: string]: unknown;
}
/**
* Safely read and parse a JSON intel file.
*
* Returns null for THREE distinct on-disk states, only one of which is a
* "quiet" case (#3885, ADR-3473 §8.5):
* - ABSENT (ENOENT, via platformReadSync returning null): silent — not
* every project has every intel file, and callers already loop over
* the full INTEL_FILES set expecting misses. Never pushed to `errors`.
* - UNREADABLE (EACCES/EIO/... — platformReadSync rethrows anything that
* isn't ENOENT): surfaced, naming the file, when `errors` is supplied.
* - MALFORMED (JSON.parse throws on a file that WAS read successfully):
* also surfaced, naming the file — a corrupt intel file used to read
* identically to "no matches", which is the same defect one layer down.
*
* `errors` is an optional accumulator so callers that want to thread the
* outcome into their own result shape can pass an array and read it back
* after the call; callers that omit it keep the prior fold-to-null shape.
*/
function safeReadJson(filePath: string, errors?: string[]): IntelData | null {
let raw: string | null;
try {
raw = platformReadSync(filePath);
} catch (err) {
if (errors) {
errors.push(`Could not read ${formatDiagnosticToken(filePath)}: ${formatDiagnosticToken((err as Error)?.message ?? String(err))}`);
}
return null;
}
if (raw === null) return null; // ENOENT — genuinely absent, not an error.
try {
return JSON.parse(raw) as IntelData;
} catch (err) {
if (errors) {
errors.push(`Could not parse ${formatDiagnosticToken(filePath)}: ${formatDiagnosticToken((err as Error)?.message ?? String(err))}`);
}
return null;
}
}
/**
* Compute SHA-256 hash of a file's contents.
* Returns null if the file doesn't exist.
*/
function hashFile(filePath: string): string | null {
try {
const content = platformReadSync(filePath);
if (content === null) return null;
return crypto.createHash('sha256').update(content).digest('hex');
} catch {
return null;
}
}
interface SearchMatch {
key: string;
value: unknown;
}
interface SearchWalkResult {
matches: SearchMatch[];
truncated: boolean;
}
/**
* Mutable walk state shared across one searchJsonEntries invocation's
* recursive matchesInValue calls. Set to true the moment ANY branch of the
* walk actually hits MAX_JSON_SEARCH_DEPTH and stops recursing further —
* never inferred from an empty result (a shallow miss must not set this).
*/
interface SearchWalkState {
truncated: boolean;
}
/**
* Search for a term (case-insensitive) in a JSON object's keys and string values.
* Returns matching entries plus whether the walk hit MAX_JSON_SEARCH_DEPTH.
*/
function searchJsonEntries(data: IntelData, term: string): SearchWalkResult {
if (!data || typeof data !== 'object') return { matches: [], truncated: false };
const entries = data.entries || data;
if (!entries || typeof entries !== 'object') return { matches: [], truncated: false };
const lowerTerm = term.toLowerCase();
const matches: SearchMatch[] = [];
const state: SearchWalkState = { truncated: false };
for (const [key, value] of Object.entries(entries)) {
if (key === '_meta') continue;
// Check key match
if (key.toLowerCase().includes(lowerTerm)) {
matches.push({ key, value });
continue;
}
// Check string value match (recursive for objects/arrays, bounded by depth)
if (matchesInValue(value, lowerTerm, 0, state)) {
matches.push({ key, value });
}
}
return { matches, truncated: state.truncated };
}
/**
* Recursively check if a term appears in any string value.
*
* `depth` counts container unwraps already performed (starts at 0 for the
* entry's own value). Strings never fail the ceiling check themselves — only
* a container (object/array) refuses to recurse one level deeper once
* `depth > MAX_JSON_SEARCH_DEPTH`, at which point `state.truncated` is set so
* the caller can report "I stopped looking" rather than a bare "no match".
* The bound is on nesting depth only, never on sibling breadth.
*/
function matchesInValue(value: unknown, lowerTerm: string, depth: number, state: SearchWalkState): boolean {
if (typeof value === 'string') {
return value.toLowerCase().includes(lowerTerm);
}
if (Array.isArray(value)) {
if (depth > MAX_JSON_SEARCH_DEPTH) {
state.truncated = true;
return false;
}
return value.some(v => matchesInValue(v, lowerTerm, depth + 1, state));
}
if (value && typeof value === 'object') {
if (depth > MAX_JSON_SEARCH_DEPTH) {
state.truncated = true;
return false;
}
return Object.values(value).some(v => matchesInValue(v, lowerTerm, depth + 1, state));
}
return false;
}
// ─── Public API ──────────────────────────────────────────────────────────────
interface IntelQueryResult {
matches: Array<{ source: string; entries: SearchMatch[] }>;
term: string;
total: number;
/**
* True iff ANY searched intel file's walk hit MAX_JSON_SEARCH_DEPTH and
* stopped early. Never true merely because nothing was found (a shallow
* miss is not a truncation) — always present as a boolean, never undefined.
*/
truncated: boolean;
/**
* #3885 (ADR-3473 §8.5): one diagnostic string per intel file that was
* present on disk but could not be read (EACCES/EIO/...) or could not be
* parsed (malformed JSON) — each naming the file via formatDiagnosticToken.
* A file that is simply ABSENT (ENOENT) never contributes an entry here;
* that is the normal, expected case (not every project has every intel
* file). Always present as an array, never undefined — empty when every
* searched file was either absent or read cleanly, mirroring `truncated`'s
* always-present convention. Plural (unlike `context_read_error`'s
* singular nullable-string shape in roadmap.cts/init.cts) because a single
* query fans out across every file in INTEL_FILES and more than one can
* independently fail.
*/
read_errors: string[];
}
/**
* Query intel files for a search term.
* Searches across all JSON intel files in INTEL_FILES (keys and values), including arch-decisions.json (parsed as JSON, not as text).
*/
function intelQuery(term: string, planningDir: string): IntelQueryResult | DisabledResponse {
if (!isIntelCapabilityActive(planningDir)) return disabledResponse();
const matches: Array<{ source: string; entries: SearchMatch[] }> = [];
let total = 0;
let truncated = false;
const readErrors: string[] = [];
// Search all JSON intel files
for (const [_key, filename] of Object.entries(INTEL_FILES)) {
const filePath = intelFilePath(planningDir, filename);
const data = safeReadJson(filePath, readErrors);
if (!data) continue;
const { matches: found, truncated: fileTruncated } = searchJsonEntries(data, term);
if (fileTruncated) truncated = true;
if (found.length > 0) {
matches.push({ source: filename, entries: found });
total += found.length;
}
}
return { matches, term, total, truncated, read_errors: readErrors };
}
interface IntelStatusFileEntry {
exists: boolean;
updated_at: string | null;
stale: boolean;
/**
* #3885 (ADR-3473 §8.5): non-null iff the file exists but could not be
* read or parsed — naming the file. `stale` still defaults to true in
* that case (an unknown-freshness file is conservatively treated as
* stale, unchanged from prior behaviour), but this field says WHY rather
* than leaving the reader to assume the file was simply never refreshed.
*/
read_error: string | null;
}
interface IntelStatusResult {
files: Record<string, IntelStatusFileEntry>;
overall_stale: boolean;
}
/**
* Report status and staleness of each intel file.
* A file is considered stale if its updated_at is older than 24 hours.
*/
function intelStatus(planningDir: string): IntelStatusResult | DisabledResponse {
if (!isIntelCapabilityActive(planningDir)) return disabledResponse();
const STALE_MS = 24 * 60 * 60 * 1000; // 24 hours
const now = Date.now();
const files: Record<string, IntelStatusFileEntry> = {};
let overallStale = false;
for (const [_key, filename] of Object.entries(INTEL_FILES)) {
const filePath = intelFilePath(planningDir, filename);
const exists = fs.existsSync(filePath);
if (!exists) {
files[filename] = { exists: false, updated_at: null, stale: true, read_error: null };
overallStale = true;
continue;
}
let updatedAt: string | null = null;
// All intel files are JSON — read _meta.updated_at
const readErrors: string[] = [];
const data = safeReadJson(filePath, readErrors);
if (data && data._meta && data._meta.updated_at) {
updatedAt = data._meta.updated_at;
}
let stale = true;
if (updatedAt) {
const age = now - new Date(updatedAt).getTime();
stale = age > STALE_MS;
}
if (stale) overallStale = true;
files[filename] = { exists: true, updated_at: updatedAt, stale, read_error: readErrors[0] ?? null };
}
return { files, overall_stale: overallStale };
}
interface IntelDiffResult {
changed: string[];
added: string[];
removed: string[];
}
/**
* Show changes since the last full refresh by comparing file hashes.
*/
function intelDiff(planningDir: string): IntelDiffResult | { no_baseline: true; read_error: string | null } | DisabledResponse {
if (!isIntelCapabilityActive(planningDir)) return disabledResponse();
const snapshotPath = intelFilePath(planningDir, '.last-refresh.json');
const readErrors: string[] = [];
const snapshot = safeReadJson(snapshotPath, readErrors);
if (!snapshot) {
// #3885 (ADR-3473 §8.5): `no_baseline: true` stays true for BOTH a
// genuinely-never-snapshotted project (ENOENT, read_error stays null)
// AND a present-but-unreadable/malformed snapshot — collapsing those
// two into a bare `no_baseline: true` would manufacture "you've never
// run a refresh" out of a read failure. `read_error` names which one.
return { no_baseline: true, read_error: readErrors[0] ?? null };
}
const prevHashes = (snapshot.hashes as Record<string, string> | undefined) || {};
const changed: string[] = [];
const added: string[] = [];
const removed: string[] = [];
// Check current files against snapshot
for (const [_key, filename] of Object.entries(INTEL_FILES)) {
const filePath = intelFilePath(planningDir, filename);
const currentHash = hashFile(filePath);
if (currentHash && !prevHashes[filename]) {
added.push(filename);
} else if (currentHash && prevHashes[filename] && currentHash !== prevHashes[filename]) {
changed.push(filename);
} else if (!currentHash && prevHashes[filename]) {
removed.push(filename);
}
}
return { changed, added, removed };
}
/**
* Stub for triggering an intel update.
* The actual update is performed by the intel-updater agent (PLAN-02).
*/
function intelUpdate(planningDir: string): { action: string; message: string } | DisabledResponse {
if (!isIntelCapabilityActive(planningDir)) return disabledResponse();
return {
action: 'spawn_agent',
message: 'Run gsd-tools intel update or spawn gsd-intel-updater agent for full refresh',
};
}
interface SaveRefreshResult {
saved: boolean;
timestamp: string;
files: number;
}
/**
* Save a refresh snapshot with hashes of all current intel files.
* Called by the intel-updater agent after completing a refresh.
*/
function saveRefreshSnapshot(planningDir: string): SaveRefreshResult {
const intelPath = ensureIntelDir(planningDir);
const hashes: Record<string, string> = {};
let fileCount = 0;
for (const [_key, filename] of Object.entries(INTEL_FILES)) {
const filePath = path.join(intelPath, filename);
const hash = hashFile(filePath);
if (hash) {
hashes[filename] = hash;
fileCount++;
}
}
const timestamp = new Date().toISOString();
const snapshotPath = path.join(intelPath, '.last-refresh.json');
platformWriteSync(snapshotPath, JSON.stringify({
hashes,
timestamp,
version: 1,
}, null, 2));
return { saved: true, timestamp, files: fileCount };
}
// ─── CLI Subcommands ─────────────────────────────────────────────────────────
/**
* Thin wrapper around saveRefreshSnapshot for CLI dispatch.
* Writes .last-refresh.json with accurate timestamps and hashes.
*/
function intelSnapshot(planningDir: string): SaveRefreshResult | DisabledResponse {
if (!isIntelCapabilityActive(planningDir)) return disabledResponse();
return saveRefreshSnapshot(planningDir);
}
interface IntelValidateResult {
valid: boolean;
errors: string[];
warnings: string[];
}
/**
* Validate all intel files for correctness and freshness.
*/
function intelValidate(planningDir: string): IntelValidateResult | DisabledResponse {
if (!isIntelCapabilityActive(planningDir)) return disabledResponse();
const errors: string[] = [];
const warnings: string[] = [];
const STALE_MS = 24 * 60 * 60 * 1000;
const now = Date.now();
for (const [key, filename] of Object.entries(INTEL_FILES)) {
const filePath = intelFilePath(planningDir, filename);
// Check existence
if (!fs.existsSync(filePath)) {
errors.push(`${filename}: file does not exist`);
continue;
}
// All intel files are JSON — validate _meta and entries structure
// Parse JSON
const raw = platformReadSync(filePath);
if (raw === null) {
errors.push(`${filename}: file missing`);
continue;
}
let data: IntelData;
try {
data = JSON.parse(raw) as IntelData;
} catch (e) {
errors.push(`${filename}: invalid JSON — ${(e as Error).message}`);
continue;
}
// Check _meta.updated_at recency
if (data._meta && data._meta.updated_at) {
const age = now - new Date(data._meta.updated_at).getTime();
if (age > STALE_MS) {
warnings.push(`${filename}: _meta.updated_at is ${Math.round(age / 3600000)} hours old (>24 hr)`);
}
} else {
warnings.push(`${filename}: missing _meta.updated_at`);
}
// Validate entries are objects with expected fields
if (data.entries && typeof data.entries === 'object') {
// file-roles.json (INTEL_FILES key 'files'): check exports are actual symbol names (no spaces)
if (key === 'files') {
for (const [entryPath, entry] of Object.entries(data.entries)) {
const entryObj = entry as Record<string, unknown>;
if (entryObj.exports && Array.isArray(entryObj.exports)) {
for (const exp of entryObj.exports as unknown[]) {
if (typeof exp === 'string' && exp.includes(' ')) {
warnings.push(`${filename}: "${entryPath}" export "${exp}" looks like a description (contains space)`);
}
}
}
}
// Spot-check first 5 file paths exist on disk
const entryPaths = Object.keys(data.entries).slice(0, 5);
for (const ep of entryPaths) {
if (!fs.existsSync(ep)) {
warnings.push(`${filename}: entry path "${ep}" does not exist on disk`);
}
}
}
// dependency-graph.json (INTEL_FILES key 'deps'): check entries have version, type, used_by
if (key === 'deps') {
for (const [depName, entry] of Object.entries(data.entries)) {
const entryObj = entry as Record<string, unknown>;
const missing: string[] = [];
if (!entryObj.version) missing.push('version');
if (!entryObj.type) missing.push('type');
if (!entryObj.used_by) missing.push('used_by');
if (missing.length > 0) {
warnings.push(`${filename}: "${depName}" missing fields: ${missing.join(', ')}`);
}
}
}
}
}
return { valid: errors.length === 0, errors, warnings };
}
interface IntelApiSurfaceResult {
written: string;
symbolCount: number;
stale: boolean;
/**
* #3885 (ADR-3473 §8.5): non-null iff api-map.json exists but could not
* be read or parsed — naming the file. Null when api-map.json is simply
* absent (the normal "not yet populated" case) or was read cleanly.
*/
read_error: string | null;
}
/**
* Render .planning/intel/api-map.json into a human-readable API-SURFACE.md.
* Always writes the file — even when api-map.json is absent or empty, the
* surface will contain an explicit "incomplete" banner so consumers never
* mistake silence for "nothing exists".
*/
function intelApiSurface(planningDir: string): IntelApiSurfaceResult | DisabledResponse {
if (!isIntelCapabilityActive(planningDir)) return disabledResponse();
const intelPath = ensureIntelDir(planningDir);
const apiMapPath = path.join(intelPath, INTEL_FILES.apis);
const outputPath = path.join(intelPath, 'API-SURFACE.md');
const readErrors: string[] = [];
const data = safeReadJson(apiMapPath, readErrors);
const readError = readErrors[0] ?? null;
const entries = (data && data.entries && typeof data.entries === 'object')
? Object.entries(data.entries)
: [];
const symbolCount = entries.length;
// Staleness: reuse the _meta.updated_at field if present
const STALE_MS = 24 * 60 * 60 * 1000;
let stale = true;
if (data && data._meta && data._meta.updated_at) {
const age = Date.now() - new Date(data._meta.updated_at).getTime();
stale = age > STALE_MS;
}
const lines: string[] = [];
lines.push('# API Surface');
lines.push('');
lines.push('> Generated from `.planning/intel/api-map.json`. Do not edit by hand.');
lines.push('');
if (symbolCount === 0) {
if (readError) {
// #3885: a read/parse failure is NOT "not yet populated" — say so,
// rather than manufacturing the wrong reason for the empty surface.
lines.push(`> **Incomplete:** ${readError}`);
} else {
lines.push('> **Incomplete:** api-map.json has no entries (intel extraction is regex/JS-only or not yet populated).');
}
lines.push('> Treat absence here as "unknown", not "does not exist".');
lines.push('');
} else {
if (stale) {
lines.push('> **Warning:** api-map.json is stale (>24 hours old). Data below may be out of date.');
lines.push('');
}
for (const [symbol, info] of entries) {
lines.push(`## \`${symbol}\``);
lines.push('');
if (info && typeof info === 'object') {
for (const [field, val] of Object.entries(info as Record<string, unknown>)) {
const display = Array.isArray(val) ? val.join(', ') : String(val);
lines.push(`- **${field}:** ${display}`);
}
}
lines.push('');
}
}
platformWriteSync(outputPath, lines.join('\n'));
return { written: outputPath, symbolCount, stale, read_error: readError };
}
interface IntelPatchMetaResult {
patched: boolean;
file?: string;
timestamp?: string;
error?: string;
}
/**
* Patch _meta.updated_at in a JSON intel file to the current timestamp.
* Reads the file, updates _meta.updated_at, increments version, writes back.
*
* NOTE: Does not gate on isCapabilityActive — operates on arbitrary file paths
* for use by agents patching individual files outside the intel store.
*/
function intelPatchMeta(filePath: string): IntelPatchMetaResult {
try {
const content = platformReadSync(filePath);
if (content === null) {
return { patched: false, error: `File not found: ${filePath}` };
}
let data: IntelData;
try {
data = JSON.parse(content) as IntelData;
} catch (e) {
return { patched: false, error: `Invalid JSON: ${(e as Error).message}` };
}
if (!data._meta) {
data._meta = {};
}
const timestamp = new Date().toISOString();
data._meta.updated_at = timestamp;
data._meta.version = (data._meta.version || 0) + 1;
platformWriteSync(filePath, JSON.stringify(data, null, 2) + '\n');
return { patched: true, file: filePath, timestamp };
} catch (e) {
return { patched: false, error: (e as Error).message };
}
}
interface IntelExtractExportsResult {
file: string;
exports: string[];
method: string;
}
/**
* Extract exports from a JS/CJS file by parsing module.exports or exports.X patterns.
*
* NOTE: Does not gate on isCapabilityActive — operates on arbitrary source files
* for use by agents building intel data from project files.
*/
function intelExtractExports(filePath: string): IntelExtractExportsResult {
const content = platformReadSync(filePath);
if (content === null) {
return { file: filePath, exports: [], method: 'none' };
}
const exports = new Set<string>();
let method = 'none';
// Try module.exports = { ... } pattern (handle multi-line)
// Find the LAST module.exports assignment (the actual one, not references in code)
const allMatches = [...content.matchAll(/module\.exports\s*=\s*\{/g)];
if (allMatches.length > 0) {
const lastMatch = allMatches[allMatches.length - 1];
const startIdx = lastMatch.index + lastMatch[0].length;
// Find matching closing brace by counting braces
let depth = 1;
let endIdx = startIdx;
while (endIdx < content.length && depth > 0) {
if (content[endIdx] === '{') depth++;
else if (content[endIdx] === '}') depth--;
if (depth > 0) endIdx++;
}
const block = content.substring(startIdx, endIdx);
method = 'module.exports';
// Extract key names from lines like " keyName," or " keyName: value,"
const lines = block.split('\n');
for (const line of lines) {
const trimmed = line.trim();
// Skip comments and empty lines
if (!trimmed || trimmed.startsWith('//') || trimmed.startsWith('*')) continue;
// Match identifier at start of line (before comma, colon, end of line)
const keyMatch = trimmed.match(/^(\w+)\s*[,}:]/) || trimmed.match(/^(\w+)$/);
if (keyMatch) {
exports.add(keyMatch[1]);
}
}
}
// Also try individual exports.X = patterns (only at start of line, not inside strings/regex)
const individualPattern = /^exports\.(\w+)\s*=/gm;
let im: RegExpExecArray | null;
while ((im = individualPattern.exec(content)) !== null) {
if (!exports.has(im[1])) {
exports.add(im[1]);
if (method === 'none') method = 'exports.X';
}
}
const hadCjs = exports.size > 0;
// ESM patterns
const esmExports = new Set<string>();
// export default function X / export default class X
const defaultNamedPattern = /^export\s+default\s+(?:function|class)\s+(\w+)/gm;
let em: RegExpExecArray | null;
while ((em = defaultNamedPattern.exec(content)) !== null) {
esmExports.add(em[1]);
}
// export default (without named function/class)
const defaultAnonPattern = /^export\s+default\s+(?!function\s|class\s)/gm;
if (defaultAnonPattern.test(content) && esmExports.size === 0) {
esmExports.add('default');
}
// export function X( / export async function X(
const exportFnPattern = /^export\s+(?:async\s+)?function\s+(\w+)\s*\(/gm;
while ((em = exportFnPattern.exec(content)) !== null) {
esmExports.add(em[1]);
}
// export const X = / export let X = / export var X =
const exportVarPattern = /^export\s+(?:const|let|var)\s+(\w+)\s*=/gm;
while ((em = exportVarPattern.exec(content)) !== null) {
esmExports.add(em[1]);
}
// export class X
const exportClassPattern = /^export\s+class\s+(\w+)/gm;
while ((em = exportClassPattern.exec(content)) !== null) {
esmExports.add(em[1]);
}
// export { X, Y, Z } — strip "as alias" parts
const exportBlockPattern = /^export\s*\{([^}]+)\}/gm;
while ((em = exportBlockPattern.exec(content)) !== null) {
const items = em[1].split(',');
for (const item of items) {
const trimmed = item.trim();
if (!trimmed) continue;
// "foo as bar" -> extract "foo"
const name = trimmed.split(/\s+as\s+/)[0].trim();
if (name) esmExports.add(name);
}
}
// Merge ESM exports into the result
for (const e of esmExports) {
exports.add(e);
}
// Determine method
const hadEsm = esmExports.size > 0;
if (hadCjs && hadEsm) {
method = 'mixed';
} else if (hadEsm && !hadCjs) {
method = 'esm';
}
return { file: filePath, exports: [...exports], method };
}
// ─── Exports ─────────────────────────────────────────────────────────────────
export = {
// Public API
intelQuery,
intelUpdate,
intelStatus,
intelDiff,
saveRefreshSnapshot,
// CLI subcommands
intelSnapshot,
intelValidate,
intelExtractExports,
intelPatchMeta,
intelApiSurface,
// Utilities
ensureIntelDir,
isIntelCapabilityActive,
// Constants
INTEL_FILES,
INTEL_DIR,
};