enhance(#1579): deterministic gsd-tools query eval.score verb (#1583)

* feat(#1579): deterministic gsd-tools query eval.score verb

Split C of #1573 (pure code, lowest risk). Adds an eval.score query verb
(coverage*0.6 + infra*0.4; bands 80/60/40) mirroring the verify.* chain;
gsd-eval-auditor consumes it instead of doing weighted arithmetic in-prompt.
Non-breaking — additive only.

arXiv: 2601.15130 (Plausibility Trap/DPDM), 2507.10281 (Table Agent), 2508.15754 (TIR).

* fix(#1579): address review — domain guard, property test, glossary, SKIP_ROOT, inventory/baseline

- C3 input-domain: reject out-of-domain eval.score (require 0<=covered<=total; was emitting overall_score>100 / negatives)
- C1 property test: add tests/eval.property.test.cjs (fast-check) — determinism, band monotonicity, [0,100] bounds, never-throws
- C2 glossary: CONTEXT.md "Eval Scoring Module" entry (source-of-truth path + interface)
- C4: add `eval` to SKIP_ROOT_RESOLUTION (pure arithmetic; no .planning/ access)
- inventory: register generated eval.cjs/eval-command-router.cjs (INVENTORY-MANIFEST.json + INVENTORY.md rows)
- size: regen agent-size baseline for gsd-eval-auditor (reused gsd_run shim + eval.score step)
- eslint: ignore generated eval*.cjs (ADR-457 bin/lib migration coverage)

* fix(#1579): register eval family in alias-drift gates

Add EVAL_COMMAND_ALIASES/EVAL_SUBCOMMANDS to scripts/check-alias-drift.cjs
families and to familyArrayKeys in the manifest-coverage test, so the eval
family lands under the same drift guard as every sibling family
(state/verify/init/phase/phases/validate/roadmap). Addresses trek-e review.

check:alias-drift ok; feat-3251 coverage 9/9; eval suites 10/10.

* docs(#1579): use half-open verdict band ranges in CLI-TOOLS

overall_score is fractional and thresholds are >=80/>=60/>=40, so a score
in [79,80) is correctly NEEDS WORK despite the old '60-79' label. Relabel
bands as 60-<80 / 40-<60 / 0-<40 to match the code. Addresses trek-e nit.

* fix(#1579): validate eval.score CLI inputs

Reject unknown infra tokens and fractional counts, and pin the 80-point verdict boundary including rounding-before-banding behavior.
This commit is contained in:
Alex V.
2026-06-25 00:16:13 +03:00
committed by GitHub
parent a63684c222
commit 1a46109b97
17 changed files with 409 additions and 11 deletions

View File

@@ -0,0 +1,5 @@
---
type: Changed
pr: 1583
---
**eval-auditor scoring moved into a deterministic `eval.score` query verb (LLM-playbook principle 10)** — coverage/infra/overall arithmetic and verdict banding are computed in code (`gsd-tools query eval.score`) instead of by the model. Based on arXiv 2601.15130 (Plausibility Trap / DPDM), 2508.15754 (Tool-Integrated Reasoning), 2507.10281 (Table Agent); 2504.00406 / 2510.15955 supporting.

2
.gitignore vendored
View File

@@ -174,6 +174,8 @@ build/
/gsd-core/bin/lib/roadmap-upgrade.cjs
/gsd-core/bin/lib/phases-command-router.cjs
/gsd-core/bin/lib/verify-command-router.cjs
/gsd-core/bin/lib/eval.cjs
/gsd-core/bin/lib/eval-command-router.cjs
/gsd-core/bin/lib/init-command-router.cjs
/gsd-core/bin/lib/agent-command-router.cjs
/gsd-core/bin/lib/agent-install-check.cjs

View File

@@ -290,6 +290,9 @@ Runtime-neutral predicate evaluating `*-UAT.md` / `*-VERIFICATION.md` result fie
### Coverage Metadata Module
Deterministic classifier for the per-deliverable coverage RTM on SUMMARY.md (#1602). Parses the optional `coverage:` frontmatter block (a list-of-maps-with-nested-list-of-maps that `extractFrontmatter` cannot represent — so a dedicated indentation parser, sibling of `parseMustHavesBlock`), validates each entry's schema, and classifies each into `auto_passed` (deterministically covered) vs `present` (human UAT required). Output envelope: `{ mode, summary_file, total, all_auto_covered, auto_passed[], present[], errors[] }` with frozen `MODE`/`PRESENT_REASON`/`ERROR_CODE` enums. Auto-pass requires the narrow proven case (strict-boolean `human_judgment:false` AND non-empty all-`pass` verification AND zero errors); everything else, including a malformed entry, routes to `present` (fail-safe — never drops a deliverable, never false-auto-passes). `mode:legacy` (absent block) ⇒ caller falls back to prose `## Accomplishments` extraction, byte-identical for un-migrated phases. Source: `gsd-core/bin/lib/coverage.cjs` (generated from `src/coverage.cts`). Wired via `uat classify-coverage --summary <f>` → `cmdClassify`; authored by `execute-plan` create_summary, consumed by `verify-work` extract_tests. See `RULESET.WORKFLOW.COVERAGE-METADATA`.
### Eval Scoring Module
Deterministic eval-scoring projection (#10 / #1579) that moves the `gsd-eval-auditor`'s weighted arithmetic out of the prompt into code. `computeEvalScore(covered, total, infra[])` returns `{ coverage_score, infra_score, overall_score, verdict }` — coverage = `covered/total*100`, infra = mean of per-item weights (`ok`=1, `partial`=0.5, `missing`=0) over exactly 5 items, `overall = coverage*0.6 + infra*0.4` (2-dp rounding), verdict banded at 80/60/40 (`PRODUCTION READY` / `NEEDS WORK` / `SIGNIFICANT GAPS` / `NOT IMPLEMENTED`). `cmdEvalScore` is the CLI guard: rejects empty/NaN flags, `infra.length !== 5`, and out-of-domain counts (requires `0 <= covered <= total`). Pure arithmetic — no `.planning/` access (it is in `SKIP_ROOT_RESOLUTION`), no `Date.now`/`Math.random`. Wired via the `eval.score` verb (and the `eval score` spaced alias) → `eval-command-router` → `cmdEvalScore`; consumed by `gsd-eval-auditor`. Source of truth: `gsd-core/bin/lib/eval.cjs` (generated from `src/eval.cts`, gitignored per ADR-457). Tests: `tests/eval.test.cjs`, `tests/eval.property.test.cjs`.
### Probe Core Module
Generic spec-phase probe resolution model — the shared seam underlying spec-completeness probes (ADR-550 Decision 7). Owns the `status × verification` model (`status: resolved | dismissed | unresolved` × a per-probe `verification` tier), structural validation (`validateResolution`, `validateRequirement` — fail-closed: `verification` must be null unless status is `resolved`, and an out-of-enum status, a `dismissed`-without-`reason`, or an `unresolved` carrying a `resolution`/`reason`/tier payload all throw rather than silently miscount), the `analyzeCoverage(items, resolutions?, validators)` merge/rollup/orphan-reject pipeline, the `byVerification` per-tier rollup, and the `runProbeCli` I/O scaffold (parse → validate → analyze → emit, structurally guarding the report shape before write — a malformed report fails closed with stderr + exit 2 instead of stringifying as green). Adapter-agnostic: consumed by the Edge Probe Module today and the Prohibition Probe Module (#644) next. Exports (generic surface): `VALID_STATUS`, `validateResolution`, `validateRequirement`, `analyzeCoverage`, `runProbeCli` — the prohibition adapter exports that also ship from this module (`projectProhibitions`, `PROHIBITION_VALIDATORS`, `validateProhibitionResolution`, `dispositionForProhibition`) are documented under the Prohibition Probe Module's own locked-surface line. Source of truth: `gsd-core/bin/lib/probe-core.cjs` (generated from `src/probe-core.cts`, gitignored per ADR-457). Tests: `tests/probe-core.test.cjs`. See ADR-550 and Edge Probe Module. Under ADR-857 (phase-6 boundary, settled 2026-06-12) this seam is classified **core verification substrate** on the *contract* side: its deterministic validators are the verifier↔predicate contract's CI-testable surface (ADR-550 Decision 5) — core and non-toggleable, never an off-by-default Feature Capability. (The recall-gapped *generator* is the probe adapters that propose predicates, not this resolution engine — see Edge Probe Module and Verification substrate (predicate boundary).)

View File

@@ -109,17 +109,14 @@ Score 5 components (ok / partial / missing):
</step>
<step name="calculate_scores">
```
coverage_score = covered_count / total_dimensions × 100
infra_score = (tooling + dataset + cicd + guardrails + tracing) / 5 × 100
overall_score = (coverage_score × 0.6) + (infra_score × 0.4)
Do NOT compute scores by hand. Call the deterministic verb with your audited inputs:
```bash
_GSD_SHIM_NAME="gsd-tools.cjs"; _GSD_RUNTIME_ROOT="${RUNTIME_DIR:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"; GSD_TOOLS="${_GSD_RUNTIME_ROOT}/gsd-core/bin/${_GSD_SHIM_NAME}"; if [ -f "$GSD_TOOLS" ]; then gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${_GSD_RUNTIME_ROOT}/.codex/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif command -v gsd-tools >/dev/null 2>&1; then GSD_TOOLS="$(command -v gsd-tools)"; gsd_run() { "$GSD_TOOLS" "$@"; }; elif [ -f "$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="$HOME/.claude/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${HERMES_HOME:-$HOME/.hermes}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CURSOR_CONFIG_DIR:-$HOME/.cursor}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEX_HOME:-$HOME/.codex}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GEMINI_CONFIG_DIR:-$HOME/.gemini}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${COPILOT_CONFIG_DIR:-$HOME/.copilot}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${WINDSURF_CONFIG_DIR:-$HOME/.codeium/windsurf}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${AUGMENT_CONFIG_DIR:-$HOME/.augment}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${TRAE_CONFIG_DIR:-$HOME/.trae}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${QWEN_CONFIG_DIR:-$HOME/.qwen}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CODEBUDDY_CONFIG_DIR:-$HOME/.codebuddy}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${CLINE_CONFIG_DIR:-$HOME/.cline}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${GROK_AGENTS_HOME:-$HOME/.agents}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${ANTIGRAVITY_CONFIG_DIR:-$HOME/.gemini/antigravity}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; elif [ -f "${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}" ]; then GSD_TOOLS="${KILO_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/kilo}/gsd-core/bin/${_GSD_SHIM_NAME}"; gsd_run() { node "$GSD_TOOLS" "$@"; }; else echo "ERROR: gsd-tools.cjs not found at $GSD_TOOLS and gsd-tools is not on PATH. Run: npx -y @opengsd/gsd-core@latest --claude --local" >&2; exit 1; fi; if [ -n "${CLAUDE_ENV_FILE:-}" ] && [ -n "${GSD_TOOLS:-}" ]; then printf "export PATH='%s':\"\$PATH\"\n" "${GSD_TOOLS%/*}" >> "$CLAUDE_ENV_FILE" 2>/dev/null || true; fi
gsd_run query eval.score --covered <covered_count> --total <total_dimensions> --infra <tooling>,<dataset>,<cicd>,<guardrails>,<tracing> --raw
```
Verdict:
- 80-100: **PRODUCTION READY** — deploy with monitoring
- 60-79: **NEEDS WORK** — address CRITICAL gaps before production
- 40-59: **SIGNIFICANT GAPS** — do not deploy
- 0-39: **NOT IMPLEMENTED** — review AI-SPEC.md and implement
where each infra component is `ok`, `partial`, or `missing` (from the audit_infrastructure step). Parse the JSON result — it returns `coverage_score`, `infra_score`, `overall_score`, and `verdict` (PRODUCTION READY / NEEDS WORK / SIGNIFICANT GAPS / NOT IMPLEMENTED). Use those values verbatim in EVAL-REVIEW.md; never recompute or override them.
</step>
<step name="write_eval_review">

View File

@@ -248,6 +248,42 @@ This command is strictly read-only — no config writes, no disk mutation.
---
### `query eval.score`
```bash
node gsd-tools.cjs query eval.score --covered <N> --total <N> --infra <tooling>,<dataset>,<cicd>,<guardrails>,<tracing>
```
Deterministic scorer for eval-auditor results. Computes coverage, infrastructure, and overall scores from audited inputs. Called by `gsd-eval-auditor` in its `calculate_scores` step — agents must not recompute these values by hand.
**Inputs:**
| Flag | Type | Description |
|---|---|---|
| `--covered` | integer | Number of eval dimensions scored COVERED |
| `--total` | integer | Total planned eval dimensions |
| `--infra` | string | Comma-separated list of 5 infra component statuses (order: tooling, dataset, cicd, guardrails, tracing); each value is `ok`, `partial`, or `missing` |
**Output JSON:**
| Field | Type | Description |
|---|---|---|
| `coverage_score` | number | `covered / total × 100` |
| `infra_score` | number | `(sum of component weights) / 5 × 100` (`ok`=1, `partial`=0.5, `missing`=0) |
| `overall_score` | number | `(coverage_score × 0.6) + (infra_score × 0.4)` |
| `verdict` | string | `PRODUCTION READY` (80–100) / `NEEDS WORK` (60–<80) / `SIGNIFICANT GAPS` (40–<60) / `NOT IMPLEMENTED` (0–<40) |
**Example:**
```bash
node gsd-tools.cjs query eval.score --covered 3 --total 5 --infra ok,partial,missing,ok,ok
# → {"coverage_score":60,"infra_score":70,"overall_score":64,"verdict":"NEEDS WORK"}
```
This command is strictly read-only — no config writes, no disk mutation.
---
## Model Resolution
```bash

View File

@@ -319,6 +319,8 @@
"docs.cjs",
"drift.cjs",
"edge-probe.cjs",
"eval-command-router.cjs",
"eval.cjs",
"fallow-runner.cjs",
"federated-config.cjs",
"frontmatter.cjs",

View File

@@ -429,6 +429,8 @@ Full listing: `gsd-core/bin/lib/*.cjs`.
| `docs.cjs` | Docs-update workflow init, Markdown scanning, monorepo detection |
| `drift.cjs` | Post-execute codebase structural drift detector (#2003): classifies file changes into new-dir/barrel/migration/route categories and round-trips `last_mapped_commit` frontmatter |
| `edge-probe.cjs` | Spec-completeness edge probe (compiled from `src/edge-probe.cts`, gitignored) — the first adapter of the `probe-core` resolution model (ADR-550 Decision 7): shape classification, applicable-category relevance filter, edge proposal, and the `{explicit, backstop}` verification validators; delegates merge/rollup/CLI to `probe-core`; exports `classifyShape`, `applicableCategories`, `proposeEdges`, `analyzeCoverage`, `validateResolution`, `TAXONOMY` (#550) |
| `eval-command-router.cjs` | Routes the `eval.score` verb (compiled from `src/eval-command-router.cts`, gitignored) — thin dispatcher into the eval scoring module (#1579) |
| `eval.cjs` | Deterministic eval scoring (compiled from `src/eval.cts`, gitignored) — `computeEvalScore` (coverage*0.6 + infra*0.4, bands 80/60/40) + `cmdEvalScore` CLI domain guard; moves the gsd-eval-auditor's weighted arithmetic out of the prompt into code (#10 / #1579) |
| `fallow-runner.cjs` | Fallow audit adapter for `/gsd-code-review`: binary resolution (`PATH` then `node_modules/.bin`), actionable missing-binary errors, and structural findings normalization |
| `federated-config.cjs` | Defensive merge of capability-declared config slices into the loadConfig return value — ADR-857 phase 3b; exports `mergeFederatedConfig({ configSchema, isCentralKey, userConfig })` → `{ values, validKeys, warnings }`; live for migrated Capability keys that are atomically removed from the central config schema |
| `frontmatter.cjs` | YAML frontmatter CRUD operations |

View File

@@ -133,6 +133,8 @@ export default tseslint.config(
'gsd-core/bin/lib/verify-command-router.cjs',
'gsd-core/bin/lib/verification.cjs',
'gsd-core/bin/lib/verification-command-router.cjs',
'gsd-core/bin/lib/eval.cjs',
'gsd-core/bin/lib/eval-command-router.cjs',
'gsd-core/bin/lib/init-command-router.cjs',
'gsd-core/bin/lib/agent-command-router.cjs',
'gsd-core/bin/lib/agent-install-check.cjs',

View File

@@ -222,6 +222,8 @@ const learnings = require('./lib/learnings.cjs');
const gapChecker = require('./lib/gap-checker.cjs');
const { routeStateCommand } = require('./lib/state-command-router.cjs');
const { routeVerifyCommand } = require('./lib/verify-command-router.cjs');
const { routeEvalCommand } = require('./lib/eval-command-router.cjs');
const evalMod = require('./lib/eval.cjs');
const { routeVerificationCommand } = require('./lib/verification-command-router.cjs');
const verification = require('./lib/verification.cjs');
const { routeInitCommand } = require('./lib/init-command-router.cjs');
@@ -640,7 +642,7 @@ async function main() {
'generate-dev-preferences, generate-slug, graphify, history-digest, init, intel, ' +
'capability, classify-confidence, git, learnings, list-seeds, list-todos, loop, milestone, package-legitimacy, phase, phase-plan-index, phases, profile-questionnaire, ' +
'profile-sample, progress, project-instruction-file, prompt-budget, requirements, research-plan, research-store, resolve-granularity, resolve-model, roadmap, scaffold, state, ' +
'task, template, user-story, validate, verify, verify-path-exists, verify-summary, workstream, worktree\n\n' +
'task, template, user-story, validate, verify, verify-path-exists, verify-summary, eval, workstream, worktree\n\n' +
'Global flags:\n' +
' --raw Emit raw output without post-processing\n' +
' --pick <field> Extract a single field from JSON output (dot/bracket notation)\n' +
@@ -694,6 +696,9 @@ async function main() {
// .planning/ access needed, and resolving project root would break workflow
// invocations that run before .planning/ exists (new-project Step 1).
'project-instruction-file',
// #1579: eval.score is pure arithmetic (covered/total + infra weights); it
// needs no .planning/ access, so skip the findProjectRoot traversal.
'eval',
]);
if (!SKIP_ROOT_RESOLUTION.has(command)) {
cwd = findProjectRoot(cwd);
@@ -1062,6 +1067,11 @@ async function runCommand(command, args, cwd, raw, defaultValue, originalCommand
break;
}
case 'eval': {
routeEvalCommand({ evalMod, args, cwd, raw, error });
break;
}
// ─── Verification Status ───────────────────────────────────────────────
//
// verification status <phaseDir>

View File

@@ -72,6 +72,11 @@ function main() {
subcommands: 'ROADMAP_SUBCOMMANDS',
routerPath: path.join(ROOT, 'gsd-core', 'bin', 'lib', 'roadmap-command-router.cjs'),
},
{
commandAliases: 'EVAL_COMMAND_ALIASES',
subcommands: 'EVAL_SUBCOMMANDS',
routerPath: path.join(ROOT, 'gsd-core', 'bin', 'lib', 'eval-command-router.cjs'),
},
];
for (const family of families) {

View File

@@ -834,3 +834,14 @@ export const PHASE_SUBCOMMANDS: string[] = PHASE_COMMAND_ALIASES.map((entry) =>
export const PHASES_SUBCOMMANDS: string[] = PHASES_COMMAND_ALIASES.map((entry) => entry.subcommand);
export const VALIDATE_SUBCOMMANDS: string[] = VALIDATE_COMMAND_ALIASES.map((entry) => entry.subcommand);
export const ROADMAP_SUBCOMMANDS: string[] = ROADMAP_COMMAND_ALIASES.map((entry) => entry.subcommand);
export const EVAL_COMMAND_ALIASES: CommandAlias[] = [
{
"canonical": "eval.score",
"aliases": ["eval score"],
"subcommand": "score",
"mutation": false
}
];
export const EVAL_SUBCOMMANDS: string[] = EVAL_COMMAND_ALIASES.map((entry) => entry.subcommand);

View File

@@ -0,0 +1,35 @@
/**
* Manifest-backed eval subcommand router (#10).
*/
import { EVAL_SUBCOMMANDS } from './command-aliases.cjs';
// eslint-disable-next-line @typescript-eslint/no-require-imports
import cjsCommandRouterAdapter = require('./cjs-command-router-adapter.cjs');
const { routeCjsCommandFamily } = cjsCommandRouterAdapter;
interface EvalModule {
cmdEvalScore(cwd: string, args: string[], raw: boolean): void;
}
interface RouteEvalCommandOptions {
evalMod: EvalModule;
args: string[];
cwd: string;
raw: boolean;
error: (message: string) => void;
}
function routeEvalCommand({ evalMod, args, cwd, raw, error }: RouteEvalCommandOptions): void {
routeCjsCommandFamily({
args,
subcommands: EVAL_SUBCOMMANDS,
unsupported: {},
error,
unknownMessage: (_s: string, available: string[]) => `Unknown eval subcommand. Available: ${available.join(', ')}`,
handlers: {
score: () => evalMod.cmdEvalScore(cwd, args, raw),
},
});
}
export = { routeEvalCommand };

74
src/eval.cts Normal file
View File

@@ -0,0 +1,74 @@
/**
* Deterministic eval scoring verb (#10).
* Moves coverage/infra/overall arithmetic out of the gsd-eval-auditor prompt
* into code, per the framework's code-delegation discipline.
*/
interface EvalScoreResult {
coverage_score: number;
infra_score: number;
overall_score: number;
verdict: string;
}
function parseFlag(args: string[], flag: string): string | undefined {
const i = args.indexOf(flag);
return i >= 0 && i + 1 < args.length ? args[i + 1] : undefined;
}
const INFRA_VALUE: Record<string, number> = { ok: 1, partial: 0.5, missing: 0 };
const INFRA_TOKENS = new Set(Object.keys(INFRA_VALUE));
function computeEvalScore(covered: number, total: number, infra: string[]): EvalScoreResult {
const coverage = total > 0 ? (covered / total) * 100 : 0;
// unknown/typo tokens are treated as `missing` (score 0) by design — upstream agent only passes ok|partial|missing
const infraSum = infra.reduce((acc, s) => acc + (INFRA_VALUE[s.trim().toLowerCase()] ?? 0), 0);
const infraScore = (infraSum / 5) * 100;
const overall = coverage * 0.6 + infraScore * 0.4;
const round = (n: number) => Math.round(n * 100) / 100;
const o = round(overall);
const verdict =
o >= 80 ? 'PRODUCTION READY' :
o >= 60 ? 'NEEDS WORK' :
o >= 40 ? 'SIGNIFICANT GAPS' : 'NOT IMPLEMENTED';
return { coverage_score: round(coverage), infra_score: round(infraScore), overall_score: o, verdict };
}
function cmdEvalScore(_cwd: string, args: string[], raw: boolean): void {
const coveredRaw = parseFlag(args, '--covered');
const totalRaw = parseFlag(args, '--total');
const infraRaw = parseFlag(args, '--infra') || '';
const infra = infraRaw ? infraRaw.split(',').map((s) => s.trim().toLowerCase()) : [];
const covered = Number(coveredRaw);
const total = Number(totalRaw);
if (
coveredRaw === undefined || coveredRaw.trim() === '' ||
totalRaw === undefined || totalRaw.trim() === '' ||
!Number.isFinite(covered) || !Number.isFinite(total) ||
infra.length !== 5
) {
process.stderr.write('Usage: gsd-tools query eval.score --covered N --total N --infra a,b,c,d,e (each ok|partial|missing)\n');
process.exitCode = 1;
return;
}
// Domain validation: this is a public CLI verb, so reject out-of-domain inputs
// rather than emit nonsense (covered>total -> coverage_score>100; negatives ->
// negative scores). Counts must be non-negative integers and covered cannot
// exceed total; infra tokens must match the documented ok|partial|missing set.
if (!Number.isInteger(covered) || !Number.isInteger(total) || covered < 0 || total < 0 || covered > total) {
process.stderr.write('Invalid eval.score domain: require integer counts with 0 <= covered <= total.\n');
process.exitCode = 1;
return;
}
const invalidInfra = infra.find((s) => !INFRA_TOKENS.has(s));
if (invalidInfra !== undefined) {
process.stderr.write(`Invalid eval.score infra token: ${invalidInfra || '<empty>'}. Expected ok|partial|missing.\n`);
process.exitCode = 1;
return;
}
const result = computeEvalScore(covered, total, infra);
process.stdout.write(raw ? JSON.stringify(result) : JSON.stringify(result, null, 2));
process.stdout.write('\n');
}
export = { cmdEvalScore, computeEvalScore };

View File

@@ -12,7 +12,7 @@
"gsd-doc-verifier.md": 12403,
"gsd-doc-writer.md": 38834,
"gsd-domain-researcher.md": 6998,
"gsd-eval-auditor.md": 7761,
"gsd-eval-auditor.md": 12362,
"gsd-eval-planner.md": 7008,
"gsd-executor.md": 43343,
"gsd-framework-selector.md": 6778,

View File

@@ -0,0 +1,101 @@
'use strict';
/**
* Property-based tests for the eval scoring module (#10 / #1579).
*
* Module: gsd-core/bin/lib/eval.cjs
* Exported: computeEvalScore(covered, total, infra), cmdEvalScore(cwd, args, raw)
*
* Properties tested:
* (a) determinism — computeEvalScore is pure: identical inputs deep-equal across calls
* (b) output shape — always { coverage_score, infra_score, overall_score, verdict };
* scores finite; verdict is exactly the band implied by overall_score
* (c) overall_score derivation — equals round(coverage*0.6 + infra*0.4) within rounding
* (d) band monotonicity — a higher overall_score never maps to a lower-quality verdict
* (e) valid-domain bounds — for 0<=covered<=total and infra in {ok,partial,missing},
* every score lands in [0,100]
* (f) never throws — tolerates arbitrary infra tokens/lengths and numeric inputs
*/
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const fc = require('./helpers/fast-check-setup.cjs');
const { computeEvalScore } = require('../gsd-core/bin/lib/eval.cjs');
const INFRA_TOKENS = ['ok', 'partial', 'missing'];
const VERDICTS = ['PRODUCTION READY', 'NEEDS WORK', 'SIGNIFICANT GAPS', 'NOT IMPLEMENTED'];
const RANK = { 'NOT IMPLEMENTED': 0, 'SIGNIFICANT GAPS': 1, 'NEEDS WORK': 2, 'PRODUCTION READY': 3 };
const band = (o) =>
o >= 80 ? 'PRODUCTION READY' :
o >= 60 ? 'NEEDS WORK' :
o >= 40 ? 'SIGNIFICANT GAPS' : 'NOT IMPLEMENTED';
// Valid-domain generator: 0 <= covered <= total, exactly 5 infra tokens.
const validDomain = fc.record({
total: fc.nat({ max: 1000 }),
infra: fc.array(fc.constantFrom(...INFRA_TOKENS), { minLength: 5, maxLength: 5 }),
}).chain(({ total, infra }) =>
fc.nat({ max: total }).map((covered) => ({ covered, total, infra })));
describe('computeEvalScore — properties', () => {
test('(a) deterministic / pure', () => {
fc.assert(fc.property(validDomain, ({ covered, total, infra }) => {
assert.deepEqual(
computeEvalScore(covered, total, infra),
computeEvalScore(covered, total, infra),
);
}));
});
test('(b) output shape + verdict matches band', () => {
fc.assert(fc.property(validDomain, ({ covered, total, infra }) => {
const r = computeEvalScore(covered, total, infra);
for (const k of ['coverage_score', 'infra_score', 'overall_score']) {
assert.ok(Number.isFinite(r[k]), `${k} must be finite`);
}
assert.ok(VERDICTS.includes(r.verdict), `verdict must be one of the four bands`);
assert.equal(r.verdict, band(r.overall_score));
}));
});
test('(c) overall_score = round(coverage*0.6 + infra*0.4)', () => {
fc.assert(fc.property(validDomain, ({ covered, total, infra }) => {
const r = computeEvalScore(covered, total, infra);
const expected = Math.round((r.coverage_score * 0.6 + r.infra_score * 0.4) * 100) / 100;
// coverage_score/infra_score are pre-rounded to 2dp; allow compounded-rounding slack.
assert.ok(Math.abs(r.overall_score - expected) <= 0.05,
`overall_score ${r.overall_score} should equal ${expected} within rounding`);
}));
});
test('(d) verdict band monotonic in overall_score', () => {
fc.assert(fc.property(validDomain, validDomain, (a, b) => {
const ra = computeEvalScore(a.covered, a.total, a.infra);
const rb = computeEvalScore(b.covered, b.total, b.infra);
if (ra.overall_score <= rb.overall_score) {
assert.ok(RANK[ra.verdict] <= RANK[rb.verdict],
`score ${ra.overall_score}<=${rb.overall_score} but verdict rank ${ra.verdict}>${rb.verdict}`);
}
}));
});
test('(e) valid-domain scores stay within [0,100]', () => {
fc.assert(fc.property(validDomain, ({ covered, total, infra }) => {
const r = computeEvalScore(covered, total, infra);
for (const k of ['coverage_score', 'infra_score', 'overall_score']) {
assert.ok(r[k] >= 0 && r[k] <= 100, `${k}=${r[k]} must be in [0,100]`);
}
}));
});
test('(f) never throws on arbitrary infra tokens / lengths / numbers', () => {
fc.assert(fc.property(
fc.integer({ min: -1000, max: 1000 }),
fc.integer({ min: -1000, max: 1000 }),
fc.array(fc.string(), { maxLength: 12 }),
(covered, total, infra) => {
assert.doesNotThrow(() => computeEvalScore(covered, total, infra));
},
));
});
});

112
tests/eval.test.cjs Normal file
View File

@@ -0,0 +1,112 @@
'use strict';
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const evalMod = require('../gsd-core/bin/lib/eval.cjs');
function capture(fn) {
const orig = process.stdout.write;
let buf = '';
process.stdout.write = (s) => { buf += s; return true; };
try { fn(); } finally { process.stdout.write = orig; }
return buf.trim();
}
function runCmd(args) {
const origOut = process.stdout.write;
const origErr = process.stderr.write;
const origExitCode = process.exitCode;
let stdout = '';
let stderr = '';
process.exitCode = 0;
process.stdout.write = (s) => { stdout += s; return true; };
process.stderr.write = (s) => { stderr += s; return true; };
try {
evalMod.cmdEvalScore(process.cwd(), args, true);
return { stdout: stdout.trim(), stderr: stderr.trim(), exitCode: process.exitCode || 0 };
} finally {
process.stdout.write = origOut;
process.stderr.write = origErr;
process.exitCode = origExitCode;
}
}
describe('eval.score (#10)', () => {
test('computes coverage/infra/overall + band', () => {
const out = JSON.parse(capture(() =>
evalMod.cmdEvalScore(process.cwd(), ['eval', 'score', '--covered', '5', '--total', '5', '--infra', 'ok,ok,ok,ok,ok'], true)));
assert.equal(out.coverage_score, 100);
assert.equal(out.infra_score, 100);
assert.equal(out.overall_score, 100);
assert.equal(out.verdict, 'PRODUCTION READY');
});
test('partial/missing infra weighted correctly', () => {
// coverage 3/5=60; infra (ok,ok,partial,missing,ok)=3.5/5=70; overall=60*.6+70*.4=64 ⇒ NEEDS WORK
const out = JSON.parse(capture(() =>
evalMod.cmdEvalScore(process.cwd(), ['eval', 'score', '--covered', '3', '--total', '5', '--infra', 'ok,ok,partial,missing,ok'], true)));
assert.equal(out.coverage_score, 60);
assert.equal(out.infra_score, 70);
assert.equal(out.overall_score, 64);
assert.equal(out.verdict, 'NEEDS WORK');
});
test('band boundary: overall exactly 60 ⇒ NEEDS WORK; 59 ⇒ SIGNIFICANT GAPS', () => {
// 60: coverage 60 (3/5), infra 60 (3/5 ok) ⇒ 60
const at60 = JSON.parse(capture(() =>
evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','3','--total','5','--infra','ok,ok,ok,missing,missing'], true)));
assert.equal(at60.overall_score, 60);
assert.equal(at60.verdict, 'NEEDS WORK');
// 40: coverage 40 (2/5), infra 40 (2/5 ok) ⇒ 40 SIGNIFICANT GAPS; under ⇒ NOT IMPLEMENTED
const at40 = JSON.parse(capture(() =>
evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','2','--total','5','--infra','ok,ok,missing,missing,missing'], true)));
assert.equal(at40.overall_score, 40);
assert.equal(at40.verdict, 'SIGNIFICANT GAPS');
});
test('band boundary: overall exactly 80 ⇒ PRODUCTION READY; 79 ⇒ NEEDS WORK', () => {
const at80 = JSON.parse(capture(() =>
evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','4','--total','5','--infra','ok,ok,ok,ok,missing'], true)));
assert.equal(at80.overall_score, 80);
assert.equal(at80.verdict, 'PRODUCTION READY');
const at79 = JSON.parse(capture(() =>
evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','13','--total','20','--infra','ok,ok,ok,ok,ok'], true)));
assert.equal(at79.overall_score, 79);
assert.equal(at79.verdict, 'NEEDS WORK');
});
test('rounding before banding can promote just-below-80 to PRODUCTION READY', () => {
const out = JSON.parse(capture(() =>
evalMod.cmdEvalScore(process.cwd(), ['eval','score','--covered','159999','--total','200000','--infra','ok,ok,ok,ok,missing'], true)));
assert.equal(out.overall_score, 80);
assert.equal(out.verdict, 'PRODUCTION READY');
});
test('missing --covered value errors: non-zero exitCode, no score JSON on stdout', () => {
const { stdout, exitCode } = runCmd(['eval', 'score', '--total', '5', '--infra', 'ok,ok,ok,ok,ok']);
assert.equal(exitCode, 1);
let parsed;
try { parsed = JSON.parse(stdout); } catch (_) { parsed = null; }
assert.ok(parsed === null || parsed.overall_score === undefined, 'stdout must not be a valid score object');
});
test('unknown infra token errors instead of silently scoring as missing', () => {
const { stdout, stderr, exitCode } = runCmd(
['eval', 'score', '--covered', '5', '--total', '5', '--infra', 'ok,ok,ok,ok,typo']);
assert.equal(exitCode, 1);
assert.match(stderr, /Invalid eval\.score infra token/i);
assert.equal(stdout, '');
});
test('fractional covered/total counts error instead of smuggling partial credit', () => {
const covered = runCmd(['eval', 'score', '--covered', '0.5', '--total', '1', '--infra', 'ok,ok,ok,ok,ok']);
assert.equal(covered.exitCode, 1);
assert.match(covered.stderr, /integer counts/i);
assert.equal(covered.stdout, '');
const total = runCmd(['eval', 'score', '--covered', '1', '--total', '1.5', '--infra', 'ok,ok,ok,ok,ok']);
assert.equal(total.exitCode, 1);
assert.match(total.stderr, /integer counts/i);
assert.equal(total.stdout, '');
});
});

View File

@@ -78,6 +78,7 @@ describe('feat-3251: command-aliases.cjs manifest coverage', () => {
'PHASES_COMMAND_ALIASES',
'VALIDATE_COMMAND_ALIASES',
'ROADMAP_COMMAND_ALIASES',
'EVAL_COMMAND_ALIASES',
];
for (const key of familyArrayKeys) {
const arr = manifest[key];