* feat(#656): add Research Store module (content-addressed cache, TTL staleness) Content-addressed research cache behind a clock seam: researchKey (sha256, deterministic), putResearch/getResearch ({hit,stale}, never throws), ttlForSource (curated HIGH 30d / MED 7d / web LOW 1d), two-tier resolveStorePath (curated -> ~/.gsd/research-cache, web/synthesis -> project .planning/research/.cache). 28 behavioral + property tests; boundary coverage at ttl-1/ttl/ttl+1. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): add Research Provider module (waterfall + confidence + plan) Single source of truth for the Balanced provider waterfall (docs Context7->Ref->Jina, web Exa+Tavily, fallback Perplexity/Brave, Firecrawl scrape-only). classifyConfidence stamps HIGH|MEDIUM|LOW by provider (never throws). providerAvailability maps config flags to usable providers. planResearch checks the Research Store (injected seam) and returns cache-hits + a per-question fetch plan, falling through the waterfall to the always-available websearch terminal. 22 behavioral + property tests. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): add Package Legitimacy module (registry-API verdicts, slopcheck optional) Replaces the pip-install-or-degrade slopcheck prose gate with code: classifyPackage (pure, never throws) computes OK|SUS|SLOP from tunable thresholds (minAgeDays 30, minWeeklyDownloads 1000, requireRepo). checkPackages queries injectable npm/PyPI/crates registry adapters (real https with 5s timeout, degraded-not-thrown on failure); slopcheck is one optional adapter that can only escalate severity, never degrade to [ASSUMED]. 34 behavioral + property tests; boundary coverage on age and downloads (limit-1/limit/limit+1). Known follow-up: real npm adapter must add api.npmjs.org last-week downloads fetch (currently null -> unknown-downloads). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): detect Tavily/Ref/Perplexity/Jina provider keys; complete npm downloads adapter config: add tavily_search/ref_search/perplexity/jina availability flags (env var or ~/.gsd/<x>_api_key), mirroring brave_search/exa_search/firecrawl, so the Research Provider waterfall can gate them. package-legitimacy: real npm adapter now fetches api.npmjs.org last-week downloads (bounded, degraded-not-thrown) so weeklyDownloads is populated. +12 config tests; 34 legitimacy tests unchanged. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#656): expose Research seam via gsd-tools query (research-plan, research-store, package-legitimacy) Routes the L2-hybrid surface so agents reach it as CLI: 'query research-store get/put' (cache, HOME-sandboxable), 'query research-plan --input' (cache-hits + fetch plan from planResearch), 'query package-legitimacy check --ecosystem' (async registry verdicts). Commands skip .planning root resolution and appear in top-level usage. 5 behavioral runGsdTools tests; command-contract unchanged (335). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#656): document Research module (CONTEXT predicates, ADR-0656, architecture, changeset) Adds GSD-RESEARCH.* + DEFECT.RESEARCH-PROVIDER-PROSE-DRIFT predicates to CONTEXT.md, ADR-0656 recording the L2-hybrid seam decision, a docs/ARCHITECTURE.md Research Module subsection, and an Added changeset fragment (pr:0, backfill on PR). Notes the #657 deferrals (agent collapse + install.js MCP mapping). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): sync inventory for research modules Regenerate INVENTORY-MANIFEST.json and bump docs/INVENTORY.md CLI Modules count 82->85 with rows for research-store/research-provider/package-legitimacy (DEFECT.INVENTORY-DRIFT). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): eslint-ignore generated research .cjs artifacts (ADR-457) research-store/research-provider/package-legitimacy .cjs are tsc-generated from src/*.cts, so they belong in the ESLint ignore block (lint the .cts source, not the emitted .cjs). Fixes tests/551-eslint-bin-lib-coverage. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): backfill changeset pr number to #664 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#656): satisfy eslint lint-tests gate Fix 20 eslint errors in the new research files: use helpers.cleanup() instead of raw fs.rmSync() in tests (local/no-raw-rmsync-in-tests, Windows-EBUSY retry budget); drop redundant '| string' union members and unnecessary type assertions; deterministic object normalization in researchKey (no-base-to-string). Logic unchanged; 6180 tests still green. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): harden package legitimacy per review (W1/W2/I3/I4) W1: httpsGet now reads statusCode; npm/PyPI/crates map 404 -> exists:false -> SLOP (registry-existence is the #1 slopsquatting defense; previously only npm caught it). Transport made injectable (_setHttpGet) for hermetic 404 tests. W2: suspicious-postinstall is now terminal SLOP independent of the optional slopcheck adapter, and the regex drops the bare https?:// arm (over-fired on esbuild/sharp/node-gyp) for shell-exec/download-exec signatures only. I3: checkPackages now threads version to registry.lookup and adapters verify that specific version exists. I4: moreServerVerdict -> moreSevereVerdict. +11 regression tests (all RED-first); 45 total green. Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): research-store tier coherence + freshness + version TTL (W4/I1/I2/I4) I1: tier now derives from source (curated -> user ~/.gsd, else -> project .planning), not kind, so put-tier and get-tier can't diverge; kind is a key component only. W4: getResearch searches both tiers and returns the freshest (non-stale preferred), never letting a stale curated entry shadow a fresh web one; blank version caps TTL at 1 day (no 30d on version-blind keys). I2: atomic platformWriteSync instead of raw fs.writeFileSync on the shared global path. I4: dropped the dead ttlForSource arm. CLI get now searches both tiers. +5 RED-first regression tests; 38 green. Addresses review by @davesienkowski on #664. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): expose classifyConfidence as a CLI route, killing dead code (W3) Adds 'gsd-tools query classify-confidence --provider X [--verified]' so research agents get the confidence tier FROM CODE (provider waterfall + verification lever) instead of asserting it in prose. classifyConfidence previously had no runtime caller. HIGH means 'trusted provider'; --verified raises web results to MEDIUM (verification semantics documented in ADR-0656). +4 behavioral tests. Addresses review by @davesienkowski on #664 (W3). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): close Codex adversarial-review findings (path-traversal, version-age, malformed-cache) HIGH: research key must be 64-hex sha256 (isValidResearchKey) + resolved-path containment check in put/get + CLI validation -> blocks '../../x' arbitrary-file-write. HIGH: package legitimacy now derives publishedAt from the REQUESTED version (npm time[version], PyPI releases[version] upload_time, crates versions[].created_at) so a new malicious version of an old package can't inherit old age and evade 'too-new'. MEDIUM: getResearch validates entry shape (finite fetched_at + positive ttl + required fields) -> malformed cache entry is a miss, not fresh-forever. +regression tests (RED-first); 111 green. Codex adversarial review (required pre-PR gate). Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): close code-review correctness findings (1) package-legitimacy CLI now rejects unknown --flags instead of silently consuming the following package as a flag value; only --ecosystem takes a value. (2) crates recent_downloads (90-day) normalized to a weekly figure before the minWeeklyDownloads threshold (was ~13x too lenient). (3) research-plan --input validates parsed JSON is an object with an Array questions before destructuring -> clean usage error instead of an uncaught TypeError on null/bad input. (4) research-store put rejects a flag value that is itself a --flag (no more storing '--source' as content). (5) planResearch skips questions whose text is not a non-empty string instead of emitting question:undefined. +13 RED-first regression tests; 143 green. Code-review gate. Issue #656. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(#657): extract researcher documentation_lookup to shared @-reference 6 researcher agents carried a near-duplicate <documentation_lookup> block; consolidate into gsd-core/references/research-documentation-lookup.md (@-included). Unifies the ctx7 CLI fallback to the safer 'command -v ctx7' guard (drops silent 'npx --yes ctx7@latest' execution in 5 agents). Behavior-preserving dedup; inventory 63->64 references. Phase A of the agent collapse. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(#657): extract researcher philosophy + verification-protocol to shared @-references philosophy and the pitfalls+pre-submission-checklist common-core were near-duplicated in project/phase researchers; consolidate into gsd-core/references/research-{philosophy,verification-protocol}.md (@-included). phase-researcher keeps its 3 extra checklist items inline. Pre-submission domains checklist made agent-agnostic so project-researcher doesn't lose features/architecture coverage. Write-contract intentionally left inline (bug-214 tests assert it verbatim). Inventory 64->66 refs. Behavior-preserving. Phase A. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#657): wire gsd-phase-researcher to the Research seam (Phase B / S1) The phase researcher now CALLS the code seam instead of carrying inline mechanics: provider waterfall -> 'gsd-tools query research-plan' (+ research-store put to cache digests); confidence-tier prose -> 'gsd-tools query classify-confidence'; slopcheck pip-install protocol -> 'gsd-tools query package-legitimacy check'. This makes the Research module a real runtime consumer (validates the seam end-to-end, addresses reviewer S1) and removes the duplicated waterfall/confidence/slopcheck prose. RESEARCH.md output contract, commit step, structured returns, and Phase-A @-includes unchanged. package-legitimacy-gate.test.cjs rewritten prose-grep -> behavioral (asserts the seam invocation). Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#657): wire gsd-project-researcher to the seam + add tavily/ref/jina MCP tools (Phase C.1) project-researcher now calls gsd-tools query research-plan / classify-confidence (+ research-store put) instead of the inline provider waterfall + confidence-tier prose (mirrors the phase-researcher rewire; no package-legitimacy — phase-only). Output contract (STACK/FEATURES/ARCHITECTURE/PITFALLS/SUMMARY.md + sections, no-commit, structured returns, Phase-A @-includes) unchanged. Adds mcp__tavily/ref/jina__* to the project/phase/ui researcher tools frontmatter (Balanced provider set) so install.js MCP mapping (C.2) has a consumer. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#657): cover tavily/ref/jina MCP install handling + frontmatter parity guard (Phase C.2) Investigation: exa/firecrawl have no explicit per-runtime tool-mapping — every mcp__<server>__* except context7 rides the generic passthrough (Copilot lowercases; OpenCode/Cursor/Windsurf/Augment keep as-is; Gemini auto-discovers). tavily/ref/jina are handled identically, no install path broken. Added 12 copilot-install passthrough tests + a mcp-tool-inheritance parity guard (tavily co-declared with exa, jina with firecrawl, ref present across the 3 web researchers) so the MCP set can't drift. No io.github registry ids invented (none sourceable in-repo); documented as a follow-up. 488 tests green. Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(#657): profiles as source of truth for researcher agents + drift-guard (Phase C.3) scripts/research-profiles.cjs declares each of the 7 researcher agents' identity + contract (name, description, color, tools, required @-includes, required gsd-tools seam calls, output-contract markers). scripts/gen-research-agents.cjs --check validates every committed agent against its profile; --write regenerates ONLY the frontmatter from profiles (body untouched) and is a verified no-op against the current agents (zero diff = fidelity). tests/research-agent-profiles.test.cjs is the DEFECT.GENERATIVE-FIX drift guard. Design note: profiles govern the generatable/contract surface rather than destructively regenerating the disparate operational prose bodies (those were deduped via @-includes in Phase A). scripts/ is not inventoried (no inventory change). Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#657): complete agent provider-dispatch + parity guard; align legitimacy field; validate profiles Adversarial-review findings: (HIGH) the seam-wired agents' Step-C dispatch only mapped 6 providers, so a planResearch result of jina/ref/perplexity/brave (reachable via the waterfall fallbacks) had no handling -> agent stall; completed both agents' dispatch to all 9 PROVIDER_WATERFALL ids + a catch-all, and added a parity test asserting agent dispatch stays in sync with research-provider PROVIDER_WATERFALL (DEFECT.GENERATIVE-FIX). (MEDIUM) phase-researcher package-legitimacy JSON example used 'package' but the module returns 'name' -> aligned. (LOW) gen-research-agents checkAgent now returns a clear failure for a malformed profile instead of throwing. +parity/validation tests (RED-first). Issue #657. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(#656): make classifyConfidence verification-evidence-driven (W3) Confidence conflated provider authority with claim verification — context7/ref stamped HIGH purely by provider identity, and the only verification lever was a self-set --verified flag. Split into two axes: provider authority (static) + verification evidence (code-computed). HIGH now requires ground-truth corroboration (legitimacyVerdict OK), independent of provider; authority alone caps at MEDIUM; SLOP caps at LOW; the self-reported --verified is demoted to a MEDIUM-only web lever. HIGH = corroborated-against-authoritative-source, not a correctness guarantee. Adds --legitimacy-verdict to the classify-confidence CLI; updates CONTEXT.md predicate + ADR-0656 (tier set unchanged, ADR-consistent). Addresses davesienkowski's W3 review on #664. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#656): bind classify-confidence verdict to code, closing CLI self-grading Adversarial review found the new --legitimacy-verdict flag was caller-supplied, so an agent could self-assert OK->HIGH without any real legitimacy check — reintroducing the exact self-grading hole W3 closes. Remove the free flag; the CLI now computes the verdict via checkPackages only when --package/--ecosystem is given (code-computed, not agent-asserted). Update the stale CLI test (context7 alone -> MEDIUM) and extend the property test to vary legitimacyVerdict. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
189 lines
7.0 KiB
JavaScript
189 lines
7.0 KiB
JavaScript
// allow-test-rule: <runtime-contract-is-the-product> research agent .md content is the governed surface
|
|
// The 7 researcher agent .md files are the deployed AI agent definitions — their
|
|
// frontmatter and @-includes ARE what the runtime loads. Asserting on their content
|
|
// is asserting on the deployed contract, not the test author's source code.
|
|
|
|
'use strict';
|
|
|
|
/**
|
|
* research-agent-profiles.test.cjs — drift guard for the 7 researcher agents.
|
|
*
|
|
* Behavioral contract (DEFECT.GENERATIVE-FIX):
|
|
* 1. The profiles table covers exactly the 7 researcher agents (no missing, no extra).
|
|
* 2. Every agent passes the profile check (frontmatter + includes + seam-calls +
|
|
* output-contract markers all match the profile).
|
|
* 3. (DEFECT.GENERATIVE-FIX parity guard) Every provider id in PROVIDER_WATERFALL
|
|
* has a dispatch mapping in the Step-C section of BOTH seam-wired researcher agents.
|
|
* 4. checkAgent returns a clear failure string for malformed profiles (not a thrown TypeError).
|
|
*
|
|
* If an agent's frontmatter/includes/seam-calls drift from its profile, this test fails.
|
|
*/
|
|
|
|
const { describe, test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const path = require('node:path');
|
|
const fs = require('node:fs');
|
|
|
|
const { PROFILES, checkAgent } = require('../scripts/gen-research-agents.cjs');
|
|
|
|
const ROOT = path.resolve(__dirname, '..');
|
|
|
|
// The canonical set of 7 researcher agent names
|
|
const EXPECTED_AGENT_NAMES = new Set([
|
|
'gsd-project-researcher',
|
|
'gsd-phase-researcher',
|
|
'gsd-advisor-researcher',
|
|
'gsd-ai-researcher',
|
|
'gsd-domain-researcher',
|
|
'gsd-ui-researcher',
|
|
'gsd-research-synthesizer',
|
|
]);
|
|
|
|
// ─── Profile coverage ─────────────────────────────────────────────────────────
|
|
|
|
describe('research-agent-profiles: coverage', () => {
|
|
test('profiles covers exactly the 7 researcher agents — no missing agents', () => {
|
|
const profileNames = new Set(PROFILES.map((p) => p.name));
|
|
const missing = [];
|
|
for (const name of EXPECTED_AGENT_NAMES) {
|
|
if (!profileNames.has(name)) missing.push(name);
|
|
}
|
|
assert.deepEqual(
|
|
missing,
|
|
[],
|
|
'These researcher agents are missing from PROFILES: ' + missing.join(', '),
|
|
);
|
|
});
|
|
|
|
test('profiles covers exactly the 7 researcher agents — no extra agents', () => {
|
|
const profileNames = PROFILES.map((p) => p.name);
|
|
const extra = profileNames.filter((n) => !EXPECTED_AGENT_NAMES.has(n));
|
|
assert.deepEqual(
|
|
extra,
|
|
[],
|
|
'PROFILES contains unexpected agent names: ' + extra.join(', '),
|
|
);
|
|
});
|
|
|
|
test('profiles contains exactly 7 entries', () => {
|
|
assert.equal(
|
|
PROFILES.length,
|
|
7,
|
|
'PROFILES should have 7 entries, got ' + PROFILES.length,
|
|
);
|
|
});
|
|
});
|
|
|
|
// ─── Per-agent parity check ───────────────────────────────────────────────────
|
|
|
|
describe('research-agent-profiles: parity', () => {
|
|
for (const profile of PROFILES) {
|
|
test(profile.name + ' matches its profile', () => {
|
|
const agentPath = path.join(ROOT, 'agents', profile.name + '.md');
|
|
assert.ok(
|
|
fs.existsSync(agentPath),
|
|
'Agent file not found: ' + agentPath,
|
|
);
|
|
|
|
const failures = checkAgent(profile);
|
|
assert.deepEqual(
|
|
failures,
|
|
[],
|
|
profile.name + ' has profile mismatches:\n' + failures.join('\n'),
|
|
);
|
|
});
|
|
}
|
|
});
|
|
|
|
// ─── Provider dispatch parity (DEFECT.GENERATIVE-FIX) ────────────────────────
|
|
//
|
|
// Every provider id in PROVIDER_WATERFALL must have a dispatch mapping in the
|
|
// Step-C section of gsd-phase-researcher.md and gsd-project-researcher.md.
|
|
// This guard fails when code adds a new provider without updating the agents.
|
|
|
|
describe('research-agent-profiles: provider dispatch parity', () => {
|
|
// The two seam-wired researcher agents that contain a Step-C dispatch table.
|
|
const SEAM_AGENTS = ['gsd-phase-researcher', 'gsd-project-researcher'];
|
|
|
|
// Load PROVIDER_WATERFALL from the compiled seam module.
|
|
const { PROVIDER_WATERFALL } = require('../gsd-core/bin/lib/research-provider.cjs');
|
|
|
|
// Compute the union of all provider ids across all waterfall kinds.
|
|
const allProviderIds = new Set();
|
|
for (const ids of Object.values(PROVIDER_WATERFALL)) {
|
|
for (const id of ids) {
|
|
allProviderIds.add(id);
|
|
}
|
|
}
|
|
|
|
// Extract the Step-C section from an agent file.
|
|
// We look for the section between "### Step C" and "### Step D".
|
|
function extractStepC(agentPath) {
|
|
const content = fs.readFileSync(agentPath, 'utf8');
|
|
const stepCStart = content.indexOf('### Step C');
|
|
if (stepCStart === -1) return '';
|
|
const stepDStart = content.indexOf('### Step D', stepCStart);
|
|
if (stepDStart === -1) return content.slice(stepCStart);
|
|
return content.slice(stepCStart, stepDStart);
|
|
}
|
|
|
|
for (const agentName of SEAM_AGENTS) {
|
|
for (const providerId of allProviderIds) {
|
|
test(agentName + ' Step-C dispatch table covers provider: ' + providerId, () => {
|
|
const agentPath = path.join(ROOT, 'agents', agentName + '.md');
|
|
assert.ok(
|
|
fs.existsSync(agentPath),
|
|
'Agent file not found: ' + agentPath,
|
|
);
|
|
const stepC = extractStepC(agentPath);
|
|
assert.ok(
|
|
stepC.includes('`' + providerId + '`') || stepC.includes('"' + providerId + '"'),
|
|
agentName + ' Step-C dispatch table is missing provider "' + providerId + '".\n' +
|
|
'Add a row for this provider in the Step-C dispatch table.\n' +
|
|
'Step-C section content:\n' + stepC,
|
|
);
|
|
});
|
|
}
|
|
}
|
|
});
|
|
|
|
// ─── checkAgent handles malformed profiles without throwing ──────────────────
|
|
|
|
describe('research-agent-profiles: checkAgent malformed profile', () => {
|
|
test('checkAgent returns clear failure string when requiredSeamCalls is missing (not a thrown TypeError)', () => {
|
|
const malformedProfile = {
|
|
name: 'gsd-phase-researcher',
|
|
description: 'some description',
|
|
color: 'cyan',
|
|
tools: 'Read',
|
|
requiredIncludes: [],
|
|
// requiredSeamCalls intentionally omitted
|
|
outputContract: [],
|
|
};
|
|
|
|
let result;
|
|
let threw = false;
|
|
try {
|
|
result = checkAgent(malformedProfile);
|
|
} catch (err) {
|
|
threw = true;
|
|
}
|
|
|
|
assert.ok(
|
|
!threw,
|
|
'checkAgent threw a TypeError instead of returning a failure string. ' +
|
|
'Add array validation at the top of checkAgent().',
|
|
);
|
|
assert.ok(
|
|
Array.isArray(result),
|
|
'checkAgent should return an array, got: ' + typeof result,
|
|
);
|
|
// Should contain a clear failure message about the missing field
|
|
const combined = result.join('\n');
|
|
assert.ok(
|
|
combined.includes('requiredSeamCalls') || combined.includes('missing required array field'),
|
|
'checkAgent should return a message mentioning the missing field "requiredSeamCalls", got: ' + combined,
|
|
);
|
|
});
|
|
});
|