diff --git a/.changeset/eager-moles-wake.md b/.changeset/eager-moles-wake.md new file mode 100644 index 000000000..6b2623de4 --- /dev/null +++ b/.changeset/eager-moles-wake.md @@ -0,0 +1,5 @@ +--- +type: Changed +pr: 1583 +--- +**eval-auditor scoring moved into a deterministic `eval.score` query verb (LLM-playbook principle 10)** — coverage/infra/overall arithmetic and verdict banding are computed in code (`gsd-tools query eval.score`) instead of by the model. Based on arXiv 2601.15130 (Plausibility Trap / DPDM), 2508.15754 (Tool-Integrated Reasoning), 2507.10281 (Table Agent); 2504.00406 / 2510.15955 supporting. diff --git a/.gitignore b/.gitignore index 2a62dfc99..12009ded1 100644 --- a/.gitignore +++ b/.gitignore @@ -174,6 +174,8 @@ build/ /gsd-core/bin/lib/roadmap-upgrade.cjs /gsd-core/bin/lib/phases-command-router.cjs /gsd-core/bin/lib/verify-command-router.cjs +/gsd-core/bin/lib/eval.cjs +/gsd-core/bin/lib/eval-command-router.cjs /gsd-core/bin/lib/init-command-router.cjs /gsd-core/bin/lib/agent-command-router.cjs /gsd-core/bin/lib/agent-install-check.cjs diff --git a/CONTEXT.md b/CONTEXT.md index da8cb1a6e..59e0a691a 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -290,6 +290,9 @@ Runtime-neutral predicate evaluating `*-UAT.md` / `*-VERIFICATION.md` result fie ### Coverage Metadata Module Deterministic classifier for the per-deliverable coverage RTM on SUMMARY.md (#1602). Parses the optional `coverage:` frontmatter block (a list-of-maps-with-nested-list-of-maps that `extractFrontmatter` cannot represent — so a dedicated indentation parser, sibling of `parseMustHavesBlock`), validates each entry's schema, and classifies each into `auto_passed` (deterministically covered) vs `present` (human UAT required). Output envelope: `{ mode, summary_file, total, all_auto_covered, auto_passed[], present[], errors[] }` with frozen `MODE`/`PRESENT_REASON`/`ERROR_CODE` enums. Auto-pass requires the narrow proven case (strict-boolean `human_judgment:false` AND non-empty all-`pass` verification AND zero errors); everything else, including a malformed entry, routes to `present` (fail-safe — never drops a deliverable, never false-auto-passes). `mode:legacy` (absent block) ⇒ caller falls back to prose `## Accomplishments` extraction, byte-identical for un-migrated phases. Source: `gsd-core/bin/lib/coverage.cjs` (generated from `src/coverage.cts`). Wired via `uat classify-coverage --summary ` → `cmdClassify`; authored by `execute-plan` create_summary, consumed by `verify-work` extract_tests. See `RULESET.WORKFLOW.COVERAGE-METADATA`. +### Eval Scoring Module +Deterministic eval-scoring projection (#10 / #1579) that moves the `gsd-eval-auditor`'s weighted arithmetic out of the prompt into code. `computeEvalScore(covered, total, infra[])` returns `{ coverage_score, infra_score, overall_score, verdict }` — coverage = `covered/total*100`, infra = mean of per-item weights (`ok`=1, `partial`=0.5, `missing`=0) over exactly 5 items, `overall = coverage*0.6 + infra*0.4` (2-dp rounding), verdict banded at 80/60/40 (`PRODUCTION READY` / `NEEDS WORK` / `SIGNIFICANT GAPS` / `NOT IMPLEMENTED`). `cmdEvalScore` is the CLI guard: rejects empty/NaN flags, `infra.length !== 5`, and out-of-domain counts (requires `0 <= covered <= total`). Pure arithmetic — no `.planning/` access (it is in `SKIP_ROOT_RESOLUTION`), no `Date.now`/`Math.random`. Wired via the `eval.score` verb (and the `eval score` spaced alias) → `eval-command-router` → `cmdEvalScore`; consumed by `gsd-eval-auditor`. Source of truth: `gsd-core/bin/lib/eval.cjs` (generated from `src/eval.cts`, gitignored per ADR-457). Tests: `tests/eval.test.cjs`, `tests/eval.property.test.cjs`. + ### Probe Core Module Generic spec-phase probe resolution model — the shared seam underlying spec-completeness probes (ADR-550 Decision 7). Owns the `status × verification` model (`status: resolved | dismissed | unresolved` × a per-probe `verification` tier), structural validation (`validateResolution`, `validateRequirement` — fail-closed: `verification` must be null unless status is `resolved`, and an out-of-enum status, a `dismissed`-without-`reason`, or an `unresolved` carrying a `resolution`/`reason`/tier payload all throw rather than silently miscount), the `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject pipeline, the `byVerification` per-tier rollup, and the `runProbeCli` I/O scaffold (parse → validate → analyze → emit, structurally guarding the report shape before write — a malformed report fails closed with stderr + exit 2 instead of stringifying as green). Adapter-agnostic: consumed by the Edge Probe Module today and the Prohibition Probe Module (#644) next. Exports (generic surface): `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` — the prohibition adapter exports that also ship from this module (`projectProhibitions`, `PROHIBITION_VALIDATORS`, `validateProhibitionResolution`, `dispositionForProhibition`) are documented under the Prohibition Probe Module's own locked-surface line. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (generated from `src/probe-core.cts`, gitignored per ADR-457). Tests: `tests/probe-core.test.cjs`. See ADR-550 and Edge Probe Module. Under ADR-857 (phase-6 boundary, settled 2026-06-12) this seam is classified **core verification substrate** on the *contract* side: its deterministic validators are the verifier↔predicate contract's CI-testable surface (ADR-550 Decision 5) — core and non-toggleable, never an off-by-default Feature Capability. (The recall-gapped *generator* is the probe adapters that propose predicates, not this resolution engine — see Edge Probe Module and Verification substrate (predicate boundary).) diff --git a/agents/gsd-eval-auditor.md b/agents/gsd-eval-auditor.md index b0608810c..4b0f96282 100644 --- a/agents/gsd-eval-auditor.md +++ b/agents/gsd-eval-auditor.md @@ -109,17 +109,14 @@ Score 5 components (ok / partial / missing): -``` -coverage_score = covered_count / total_dimensions × 100 -infra_score = (tooling + dataset + cicd + guardrails + tracing) / 5 × 100 -overall_score = (coverage_score × 0.6) + (infra_score × 0.4) +Do NOT compute scores by hand. Call the deterministic verb with your audited inputs: + +```bash +_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif command -v gsd-tools >/dev/null 2>&1; then GSD_TOOLS="$(command -v gsd-tools)"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd-tools is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi +gsd_run query eval.score --covered --total --infra ,,,, --raw ``` -Verdict: -- 80-100: **PRODUCTION READY** — deploy with monitoring -- 60-79: **NEEDS WORK** — address CRITICAL gaps before production -- 40-59: **SIGNIFICANT GAPS** — do not deploy -- 0-39: **NOT IMPLEMENTED** — review AI-SPEC.md and implement +where each infra component is `ok`, `partial`, or `missing` (from the audit_infrastructure step). Parse the JSON result — it returns `coverage_score`, `infra_score`, `overall_score`, and `verdict` (PRODUCTION READY / NEEDS WORK / SIGNIFICANT GAPS / NOT IMPLEMENTED). Use those values verbatim in EVAL-REVIEW.md; never recompute or override them. diff --git a/docs/CLI-TOOLS.md b/docs/CLI-TOOLS.md index 13eeb009d..c698a986f 100644 --- a/docs/CLI-TOOLS.md +++ b/docs/CLI-TOOLS.md @@ -248,6 +248,42 @@ This command is strictly read-only — no config writes, no disk mutation. --- +### `query eval.score` + +```bash +node gsd-tools.cjs query eval.score --covered --total --infra ,,,, +``` + +Deterministic scorer for eval-auditor results. Computes coverage, infrastructure, and overall scores from audited inputs. Called by `gsd-eval-auditor` in its `calculate_scores` step — agents must not recompute these values by hand. + +**Inputs:** + +| Flag | Type | Description | +|---|---|---| +| `--covered` | integer | Number of eval dimensions scored COVERED | +| `--total` | integer | Total planned eval dimensions | +| `--infra` | string | Comma-separated list of 5 infra component statuses (order: tooling, dataset, cicd, guardrails, tracing); each value is `ok`, `partial`, or `missing` | + +**Output JSON:** + +| Field | Type | Description | +|---|---|---| +| `coverage_score` | number | `covered / total × 100` | +| `infra_score` | number | `(sum of component weights) / 5 × 100` (`ok`=1, `partial`=0.5, `missing`=0) | +| `overall_score` | number | `(coverage_score × 0.6) + (infra_score × 0.4)` | +| `verdict` | string | `PRODUCTION READY` (80–100) / `NEEDS WORK` (60–<80) / `SIGNIFICANT GAPS` (40–<60) / `NOT IMPLEMENTED` (0–<40) | + +**Example:** + +```bash +node gsd-tools.cjs query eval.score --covered 3 --total 5 --infra ok,partial,missing,ok,ok +# → {"coverage_score":60,"infra_score":70,"overall_score":64,"verdict":"NEEDS WORK"} +``` + +This command is strictly read-only — no config writes, no disk mutation. + +--- + ## Model Resolution ```bash diff --git a/docs/INVENTORY-MANIFEST.json b/docs/INVENTORY-MANIFEST.json index f10b71a78..8b26886a3 100644 --- a/docs/INVENTORY-MANIFEST.json +++ b/docs/INVENTORY-MANIFEST.json @@ -319,6 +319,8 @@ "docs.cjs", "drift.cjs", "edge-probe.cjs", + "eval-command-router.cjs", + "eval.cjs", "fallow-runner.cjs", "federated-config.cjs", "frontmatter.cjs", diff --git a/docs/INVENTORY.md b/docs/INVENTORY.md index 5be55d91a..1474fb47d 100644 --- a/docs/INVENTORY.md +++ b/docs/INVENTORY.md @@ -429,6 +429,8 @@ Full listing: `gsd-core/bin/lib/*.cjs`. | `docs.cjs` | Docs-update workflow init, Markdown scanning, monorepo detection | | `drift.cjs` | Post-execute codebase structural drift detector (#2003): classifies file changes into new-dir/barrel/migration/route categories and round-trips `last_mapped_commit` frontmatter | | `edge-probe.cjs` | Spec-completeness edge probe (compiled from `src/edge-probe.cts`, gitignored) — the first adapter of the `probe-core` resolution model (ADR-550 Decision 7): shape classification, applicable-category relevance filter, edge proposal, and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`; exports `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `TAXONOMY` (#550) | +| `eval-command-router.cjs` | Routes the `eval.score` verb (compiled from `src/eval-command-router.cts`, gitignored) — thin dispatcher into the eval scoring module (#1579) | +| `eval.cjs` | Deterministic eval scoring (compiled from `src/eval.cts`, gitignored) — `computeEvalScore` (coverage*0.6 + infra*0.4, bands 80/60/40) + `cmdEvalScore` CLI domain guard; moves the gsd-eval-auditor's weighted arithmetic out of the prompt into code (#10 / #1579) | | `fallow-runner.cjs` | Fallow audit adapter for `/gsd-code-review`: binary resolution (`PATH` then `node_modules/.bin`), actionable missing-binary errors, and structural findings normalization | | `federated-config.cjs` | Defensive merge of capability-declared config slices into the loadConfig return value — ADR-857 phase 3b; exports `mergeFederatedConfig({ configSchema, isCentralKey, userConfig })` → `{ values, validKeys, warnings }`; live for migrated Capability keys that are atomically removed from the central config schema | | `frontmatter.cjs` | YAML frontmatter CRUD operations | diff --git a/eslint.config.mjs b/eslint.config.mjs index cf9833ba3..20ec0157b 100644 --- a/eslint.config.mjs +++ b/eslint.config.mjs @@ -133,6 +133,8 @@ export default tseslint.config( 'gsd-core/bin/lib/verify-command-router.cjs', 'gsd-core/bin/lib/verification.cjs', 'gsd-core/bin/lib/verification-command-router.cjs', + 'gsd-core/bin/lib/eval.cjs', + 'gsd-core/bin/lib/eval-command-router.cjs', 'gsd-core/bin/lib/init-command-router.cjs', 'gsd-core/bin/lib/agent-command-router.cjs', 'gsd-core/bin/lib/agent-install-check.cjs', diff --git a/gsd-core/bin/gsd-tools.cjs b/gsd-core/bin/gsd-tools.cjs index 893d64f46..83eadeb7b 100755 --- a/gsd-core/bin/gsd-tools.cjs +++ b/gsd-core/bin/gsd-tools.cjs @@ -222,6 +222,8 @@ const learnings = require('./lib/learnings.cjs'); const gapChecker = require('./lib/gap-checker.cjs'); const { routeStateCommand } = require('./lib/state-command-router.cjs'); const { routeVerifyCommand } = require('./lib/verify-command-router.cjs'); +const { routeEvalCommand } = require('./lib/eval-command-router.cjs'); +const evalMod = require('./lib/eval.cjs'); const { routeVerificationCommand } = require('./lib/verification-command-router.cjs'); const verification = require('./lib/verification.cjs'); const { routeInitCommand } = require('./lib/init-command-router.cjs'); @@ -640,7 +642,7 @@ async function main() { 'generate-dev-preferences, generate-slug, graphify, history-digest, init, intel, ' + 'capability, classify-confidence, git, learnings, list-seeds, list-todos, loop, milestone, package-legitimacy, phase, phase-plan-index, phases, profile-questionnaire, ' + 'profile-sample, progress, project-instruction-file, prompt-budget, requirements, research-plan, research-store, resolve-granularity, resolve-model, roadmap, scaffold, state, ' + - 'task, template, user-story, validate, verify, verify-path-exists, verify-summary, workstream, worktree\n\n' + + 'task, template, user-story, validate, verify, verify-path-exists, verify-summary, eval, workstream, worktree\n\n' + 'Global flags:\n' + ' --raw Emit raw output without post-processing\n' + ' --pick Extract a single field from JSON output (dot/bracket notation)\n' + @@ -694,6 +696,9 @@ async function main() { // .planning/ access needed, and resolving project root would break workflow // invocations that run before .planning/ exists (new-project Step 1). 'project-instruction-file', + // #1579: eval.score is pure arithmetic (covered/total + infra weights); it + // needs no .planning/ access, so skip the findProjectRoot traversal. + 'eval', ]); if (!SKIP_ROOT_RESOLUTION.has(command)) { cwd = findProjectRoot(cwd); @@ -1062,6 +1067,11 @@ async function runCommand(command, args, cwd, raw, defaultValue, originalCommand break; } + case 'eval': { + routeEvalCommand({ evalMod, args, cwd, raw, error }); + break; + } + // ─── Verification Status ─────────────────────────────────────────────── // // verification status diff --git a/scripts/check-alias-drift.cjs b/scripts/check-alias-drift.cjs index d5c3e2784..abaca0be8 100644 --- a/scripts/check-alias-drift.cjs +++ b/scripts/check-alias-drift.cjs @@ -72,6 +72,11 @@ function main() { subcommands: 'ROADMAP_SUBCOMMANDS', routerPath: path.join(ROOT, 'gsd-core', 'bin', 'lib', 'roadmap-command-router.cjs'), }, + { + commandAliases: 'EVAL_COMMAND_ALIASES', + subcommands: 'EVAL_SUBCOMMANDS', + routerPath: path.join(ROOT, 'gsd-core', 'bin', 'lib', 'eval-command-router.cjs'), + }, ]; for (const family of families) { diff --git a/src/command-aliases.cts b/src/command-aliases.cts index 9269ce019..a864aa701 100644 --- a/src/command-aliases.cts +++ b/src/command-aliases.cts @@ -834,3 +834,14 @@ export const PHASE_SUBCOMMANDS: string[] = PHASE_COMMAND_ALIASES.map((entry) => export const PHASES_SUBCOMMANDS: string[] = PHASES_COMMAND_ALIASES.map((entry) => entry.subcommand); export const VALIDATE_SUBCOMMANDS: string[] = VALIDATE_COMMAND_ALIASES.map((entry) => entry.subcommand); export const ROADMAP_SUBCOMMANDS: string[] = ROADMAP_COMMAND_ALIASES.map((entry) => entry.subcommand); + +export const EVAL_COMMAND_ALIASES: CommandAlias[] = [ + { + "canonical": "eval.score", + "aliases": ["eval score"], + "subcommand": "score", + "mutation": false + } +]; + +export const EVAL_SUBCOMMANDS: string[] = EVAL_COMMAND_ALIASES.map((entry) => entry.subcommand); diff --git a/src/eval-command-router.cts b/src/eval-command-router.cts new file mode 100644 index 000000000..b9fb32734 --- /dev/null +++ b/src/eval-command-router.cts @@ -0,0 +1,35 @@ +/** + * Manifest-backed eval subcommand router (#10). + */ + +import { EVAL_SUBCOMMANDS } from './command-aliases.cjs'; +// eslint-disable-next-line @typescript-eslint/no-require-imports +import cjsCommandRouterAdapter = require('./cjs-command-router-adapter.cjs'); +const { routeCjsCommandFamily } = cjsCommandRouterAdapter; + +interface EvalModule { + cmdEvalScore(cwd: string, args: string[], raw: boolean): void; +} + +interface RouteEvalCommandOptions { + evalMod: EvalModule; + args: string[]; + cwd: string; + raw: boolean; + error: (message: string) => void; +} + +function routeEvalCommand({ evalMod, args, cwd, raw, error }: RouteEvalCommandOptions): void { + routeCjsCommandFamily({ + args, + subcommands: EVAL_SUBCOMMANDS, + unsupported: {}, + error, + unknownMessage: (_s: string, available: string[]) => `Unknown eval subcommand. Available: ${available.join(', ')}`, + handlers: { + score: () => evalMod.cmdEvalScore(cwd, args, raw), + }, + }); +} + +export = { routeEvalCommand }; diff --git a/src/eval.cts b/src/eval.cts new file mode 100644 index 000000000..67a71d455 --- /dev/null +++ b/src/eval.cts @@ -0,0 +1,74 @@ +/** + * Deterministic eval scoring verb (#10). + * Moves coverage/infra/overall arithmetic out of the gsd-eval-auditor prompt + * into code, per the framework's code-delegation discipline. + */ + +interface EvalScoreResult { + coverage_score: number; + infra_score: number; + overall_score: number; + verdict: string; +} + +function parseFlag(args: string[], flag: string): string | undefined { + const i = args.indexOf(flag); + return i >= 0 && i + 1 < args.length ? args[i + 1] : undefined; +} + +const INFRA_VALUE: Record = { ok: 1, partial: 0.5, missing: 0 }; +const INFRA_TOKENS = new Set(Object.keys(INFRA_VALUE)); + +function computeEvalScore(covered: number, total: number, infra: string[]): EvalScoreResult { + const coverage = total > 0 ? (covered / total) * 100 : 0; + // unknown/typo tokens are treated as `missing` (score 0) by design — upstream agent only passes ok|partial|missing + const infraSum = infra.reduce((acc, s) => acc + (INFRA_VALUE[s.trim().toLowerCase()] ?? 0), 0); + const infraScore = (infraSum / 5) * 100; + const overall = coverage * 0.6 + infraScore * 0.4; + const round = (n: number) => Math.round(n * 100) / 100; + const o = round(overall); + const verdict = + o >= 80 ? 'PRODUCTION READY' : + o >= 60 ? 'NEEDS WORK' : + o >= 40 ? 'SIGNIFICANT GAPS' : 'NOT IMPLEMENTED'; + return { coverage_score: round(coverage), infra_score: round(infraScore), overall_score: o, verdict }; +} + +function cmdEvalScore(_cwd: string, args: string[], raw: boolean): void { + const coveredRaw = parseFlag(args, '--covered'); + const totalRaw = parseFlag(args, '--total'); + const infraRaw = parseFlag(args, '--infra') || ''; + const infra = infraRaw ? infraRaw.split(',').map((s) => s.trim().toLowerCase()) : []; + const covered = Number(coveredRaw); + const total = Number(totalRaw); + if ( + coveredRaw === undefined || coveredRaw.trim() === '' || + totalRaw === undefined || totalRaw.trim() === '' || + !Number.isFinite(covered) || !Number.isFinite(total) || + infra.length !== 5 + ) { + process.stderr.write('Usage: gsd-tools query eval.score --covered N --total N --infra a,b,c,d,e (each ok|partial|missing)\n'); + process.exitCode = 1; + return; + } + // Domain validation: this is a public CLI verb, so reject out-of-domain inputs + // rather than emit nonsense (covered>total -> coverage_score>100; negatives -> + // negative scores). Counts must be non-negative integers and covered cannot + // exceed total; infra tokens must match the documented ok|partial|missing set. + if (!Number.isInteger(covered) || !Number.isInteger(total) || covered < 0 || total < 0 || covered > total) { + process.stderr.write('Invalid eval.score domain: require integer counts with 0 <= covered <= total.\n'); + process.exitCode = 1; + return; + } + const invalidInfra = infra.find((s) => !INFRA_TOKENS.has(s)); + if (invalidInfra !== undefined) { + process.stderr.write(`Invalid eval.score infra token: ${invalidInfra || ''}. Expected ok|partial|missing.\n`); + process.exitCode = 1; + return; + } + const result = computeEvalScore(covered, total, infra); + process.stdout.write(raw ? JSON.stringify(result) : JSON.stringify(result, null, 2)); + process.stdout.write('\n'); +} + +export = { cmdEvalScore, computeEvalScore }; diff --git a/tests/agent-size-baseline.json b/tests/agent-size-baseline.json index c247f93d6..782aaa5dc 100644 --- a/tests/agent-size-baseline.json +++ b/tests/agent-size-baseline.json @@ -12,7 +12,7 @@ "gsd-doc-verifier.md": 12403, "gsd-doc-writer.md": 38834, "gsd-domain-researcher.md": 6998, - "gsd-eval-auditor.md": 7761, + "gsd-eval-auditor.md": 12362, "gsd-eval-planner.md": 7008, "gsd-executor.md": 43343, "gsd-framework-selector.md": 6778, diff --git a/tests/eval.property.test.cjs b/tests/eval.property.test.cjs new file mode 100644 index 000000000..798ae1488 --- /dev/null +++ b/tests/eval.property.test.cjs @@ -0,0 +1,101 @@ +'use strict'; + +/** + * Property-based tests for the eval scoring module (#10 / #1579). + * + * Module: gsd-core/bin/lib/eval.cjs + * Exported: computeEvalScore(covered, total, infra), cmdEvalScore(cwd, args, raw) + * + * Properties tested: + * (a) determinism — computeEvalScore is pure: identical inputs deep-equal across calls + * (b) output shape — always { coverage_score, infra_score, overall_score, verdict }; + * scores finite; verdict is exactly the band implied by overall_score + * (c) overall_score derivation — equals round(coverage*0.6 + infra*0.4) within rounding + * (d) band monotonicity — a higher overall_score never maps to a lower-quality verdict + * (e) valid-domain bounds — for 0<=covered<=total and infra in {ok,partial,missing}, + * every score lands in [0,100] + * (f) never throws — tolerates arbitrary infra tokens/lengths and numeric inputs + */ + +const { describe, test } = require('node:test'); +const assert = require('node:assert/strict'); +const fc = require('./helpers/fast-check-setup.cjs'); +const { computeEvalScore } = require('../gsd-core/bin/lib/eval.cjs'); + +const INFRA_TOKENS = ['ok', 'partial', 'missing']; +const VERDICTS = ['PRODUCTION READY', 'NEEDS WORK', 'SIGNIFICANT GAPS', 'NOT IMPLEMENTED']; +const RANK = { 'NOT IMPLEMENTED': 0, 'SIGNIFICANT GAPS': 1, 'NEEDS WORK': 2, 'PRODUCTION READY': 3 }; +const band = (o) => + o >= 80 ? 'PRODUCTION READY' : + o >= 60 ? 'NEEDS WORK' : + o >= 40 ? 'SIGNIFICANT GAPS' : 'NOT IMPLEMENTED'; + +// Valid-domain generator: 0 <= covered <= total, exactly 5 infra tokens. +const validDomain = fc.record({ + total: fc.nat({ max: 1000 }), + infra: fc.array(fc.constantFrom(...INFRA_TOKENS), { minLength: 5, maxLength: 5 }), +}).chain(({ total, infra }) => + fc.nat({ max: total }).map((covered) => ({ covered, total, infra }))); + +describe('computeEvalScore — properties', () => { + test('(a) deterministic / pure', () => { + fc.assert(fc.property(validDomain, ({ covered, total, infra }) => { + assert.deepEqual( + computeEvalScore(covered, total, infra), + computeEvalScore(covered, total, infra), + ); + })); + }); + + test('(b) output shape + verdict matches band', () => { + fc.assert(fc.property(validDomain, ({ covered, total, infra }) => { + const r = computeEvalScore(covered, total, infra); + for (const k of ['coverage_score', 'infra_score', 'overall_score']) { + assert.ok(Number.isFinite(r[k]), `${k} must be finite`); + } + assert.ok(VERDICTS.includes(r.verdict), `verdict must be one of the four bands`); + assert.equal(r.verdict, band(r.overall_score)); + })); + }); + + test('(c) overall_score = round(coverage*0.6 + infra*0.4)', () => { + fc.assert(fc.property(validDomain, ({ covered, total, infra }) => { + const r = computeEvalScore(covered, total, infra); + const expected = Math.round((r.coverage_score * 0.6 + r.infra_score * 0.4) * 100) / 100; + // coverage_score/infra_score are pre-rounded to 2dp; allow compounded-rounding slack. + assert.ok(Math.abs(r.overall_score - expected) <= 0.05, + `overall_score ${r.overall_score} should equal ${expected} within rounding`); + })); + }); + + test('(d) verdict band monotonic in overall_score', () => { + fc.assert(fc.property(validDomain, validDomain, (a, b) => { + const ra = computeEvalScore(a.covered, a.total, a.infra); + const rb = computeEvalScore(b.covered, b.total, b.infra); + if (ra.overall_score <= rb.overall_score) { + assert.ok(RANK[ra.verdict] <= RANK[rb.verdict], + `score ${ra.overall_score}<=${rb.overall_score} but verdict rank ${ra.verdict}>${rb.verdict}`); + } + })); + }); + + test('(e) valid-domain scores stay within [0,100]', () => { + fc.assert(fc.property(validDomain, ({ covered, total, infra }) => { + const r = computeEvalScore(covered, total, infra); + for (const k of ['coverage_score', 'infra_score', 'overall_score']) { + assert.ok(r[k] >= 0 && r[k] <= 100, `${k}=${r[k]} must be in [0,100]`); + } + })); + }); + + test('(f) never throws on arbitrary infra tokens / lengths / numbers', () => { + fc.assert(fc.property( + fc.integer({ min: -1000, max: 1000 }), + fc.integer({ min: -1000, max: 1000 }), + fc.array(fc.string(), { maxLength: 12 }), + (covered, total, infra) => { + assert.doesNotThrow(() => computeEvalScore(covered, total, infra)); + }, + )); + }); +}); diff --git a/tests/eval.test.cjs b/tests/eval.test.cjs new file mode 100644 index 000000000..615a6e5a7 --- /dev/null +++ b/tests/eval.test.cjs @@ -0,0 +1,112 @@ +'use strict'; +const { test, describe } = require('node:test'); +const assert = require('node:assert/strict'); +const evalMod = require('../gsd-core/bin/lib/eval.cjs'); + +function capture(fn) { + const orig = process.stdout.write; + let buf = ''; + process.stdout.write = (s) => { buf += s; return true; }; + try { fn(); } finally { process.stdout.write = orig; } + return buf.trim(); +} + +function runCmd(args) { + const origOut = process.stdout.write; + const origErr = process.stderr.write; + const origExitCode = process.exitCode; + let stdout = ''; + let stderr = ''; + process.exitCode = 0; + process.stdout.write = (s) => { stdout += s; return true; }; + process.stderr.write = (s) => { stderr += s; return true; }; + try { + evalMod.cmdEvalScore(process.cwd(), args, true); + return { stdout: stdout.trim(), stderr: stderr.trim(), exitCode: process.exitCode || 0 }; + } finally { + process.stdout.write = origOut; + process.stderr.write = origErr; + process.exitCode = origExitCode; + } +} + +describe('eval.score (#10)', () => { + test('computes coverage/infra/overall + band', () => { + const out = JSON.parse(capture(() => + evalMod.cmdEvalScore(process.cwd(), ['eval', 'score', '--covered', '5', '--total', '5', '--infra', 'ok,ok,ok,ok,ok'], true))); + assert.equal(out.coverage_score, 100); + assert.equal(out.infra_score, 100); + assert.equal(out.overall_score, 100); + assert.equal(out.verdict, 'PRODUCTION READY'); + }); + + test('partial/missing infra weighted correctly', () => { + // coverage 3/5=60; infra (ok,ok,partial,missing,ok)=3.5/5=70; overall=60*.6+70*.4=64 ⇒ NEEDS WORK + const out = JSON.parse(capture(() => + evalMod.cmdEvalScore(process.cwd(), ['eval', 'score', '--covered', '3', '--total', '5', '--infra', 'ok,ok,partial,missing,ok'], true))); + assert.equal(out.coverage_score, 60); + assert.equal(out.infra_score, 70); + assert.equal(out.overall_score, 64); + assert.equal(out.verdict, 'NEEDS WORK'); + }); + + test('band boundary: overall exactly 60 ⇒ NEEDS WORK; 59 ⇒ SIGNIFICANT GAPS', () => { + // 60: coverage 60 (3/5), infra 60 (3/5 ok) ⇒ 60 + const at60 = JSON.parse(capture(() => + evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','3','--total','5','--infra','ok,ok,ok,missing,missing'], true))); + assert.equal(at60.overall_score, 60); + assert.equal(at60.verdict, 'NEEDS WORK'); + // 40: coverage 40 (2/5), infra 40 (2/5 ok) ⇒ 40 SIGNIFICANT GAPS; under ⇒ NOT IMPLEMENTED + const at40 = JSON.parse(capture(() => + evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','2','--total','5','--infra','ok,ok,missing,missing,missing'], true))); + assert.equal(at40.overall_score, 40); + assert.equal(at40.verdict, 'SIGNIFICANT GAPS'); + }); + + test('band boundary: overall exactly 80 ⇒ PRODUCTION READY; 79 ⇒ NEEDS WORK', () => { + const at80 = JSON.parse(capture(() => + evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','4','--total','5','--infra','ok,ok,ok,ok,missing'], true))); + assert.equal(at80.overall_score, 80); + assert.equal(at80.verdict, 'PRODUCTION READY'); + + const at79 = JSON.parse(capture(() => + evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','13','--total','20','--infra','ok,ok,ok,ok,ok'], true))); + assert.equal(at79.overall_score, 79); + assert.equal(at79.verdict, 'NEEDS WORK'); + }); + + test('rounding before banding can promote just-below-80 to PRODUCTION READY', () => { + const out = JSON.parse(capture(() => + evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','159999','--total','200000','--infra','ok,ok,ok,ok,missing'], true))); + assert.equal(out.overall_score, 80); + assert.equal(out.verdict, 'PRODUCTION READY'); + }); + + test('missing --covered value errors: non-zero exitCode, no score JSON on stdout', () => { + const { stdout, exitCode } = runCmd(['eval', 'score', '--total', '5', '--infra', 'ok,ok,ok,ok,ok']); + assert.equal(exitCode, 1); + let parsed; + try { parsed = JSON.parse(stdout); } catch (_) { parsed = null; } + assert.ok(parsed === null || parsed.overall_score === undefined, 'stdout must not be a valid score object'); + }); + + test('unknown infra token errors instead of silently scoring as missing', () => { + const { stdout, stderr, exitCode } = runCmd( + ['eval', 'score', '--covered', '5', '--total', '5', '--infra', 'ok,ok,ok,ok,typo']); + assert.equal(exitCode, 1); + assert.match(stderr, /Invalid eval\.score infra token/i); + assert.equal(stdout, ''); + }); + + test('fractional covered/total counts error instead of smuggling partial credit', () => { + const covered = runCmd(['eval', 'score', '--covered', '0.5', '--total', '1', '--infra', 'ok,ok,ok,ok,ok']); + assert.equal(covered.exitCode, 1); + assert.match(covered.stderr, /integer counts/i); + assert.equal(covered.stdout, ''); + + const total = runCmd(['eval', 'score', '--covered', '1', '--total', '1.5', '--infra', 'ok,ok,ok,ok,ok']); + assert.equal(total.exitCode, 1); + assert.match(total.stderr, /integer counts/i); + assert.equal(total.stdout, ''); + }); +}); diff --git a/tests/feat-3251-command-aliases-manifest-coverage.test.cjs b/tests/feat-3251-command-aliases-manifest-coverage.test.cjs index 7208944c0..e920ce790 100644 --- a/tests/feat-3251-command-aliases-manifest-coverage.test.cjs +++ b/tests/feat-3251-command-aliases-manifest-coverage.test.cjs @@ -78,6 +78,7 @@ describe('feat-3251: command-aliases.cjs manifest coverage', () => { 'PHASES_COMMAND_ALIASES', 'VALIDATE_COMMAND_ALIASES', 'ROADMAP_COMMAND_ALIASES', + 'EVAL_COMMAND_ALIASES', ]; for (const key of familyArrayKeys) { const arr = manifest[key];