Files
msd-core/tests/qa/scenario.cjs
Tom Boucher 33fd203ccd test(#2966): loop QA walk — drive real scenarios across all five loop steps (#2976)
* test(#2966): loop QA walk — drive real scenarios across all five loop steps

Adds a headless walk that carries accumulating project state across
discuss -> plan -> execute -> verify -> ship against one temp project,
layered over the existing tests/helpers.cjs runGsdTools substrate.

Findings carry severity. A violation breaks a stated contract and fails
the build; a smell is legal under today's implementation but structurally
questionable, is recorded, and never reddens CI. Without that split an
oracle set derived from current behavior can only ever confirm current
behavior -- the harness could not say "this works and is still wrong".

The end-to-end test asserts the walk produces at least one smell: a QA
harness that reports nothing on a first run against a real engine is far
more likely mis-specified than the engine is perfect. It deliberately does
not pin smell ids or counts, which would re-freeze current behavior.

First run against the real engine: 0 violations, 3 smell classes --
init returns agents_dir outside the project tree; smart-entry emits prose
unconditionally so routing cannot be asserted; state-snapshot reports a
missing STATE.md through a payload key with exit 0.

Also fixes tests/fixtures/index.cjs: createFixture with git:true and
planning:false staged nothing, so the commit failed with "nothing to
commit". That combination was unreachable until greenfield needed it.

Extends RULESET.TESTS.feedback-loop-convergence from estimation to the
loop itself. Design lock: docs/adr/2966-loop-qa-walk.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): wire fault injection, make perturbations discriminating

Independent review found tests/qa/mutations.cjs entirely unwired: 462
lines exercised only by their own unit tests, with no mutation hook in
the scenario DSL and no scenario applying one, while the module header
and the ADR described fault injection in the present tense. Dead code
documented as live.

Adds a `mutate` step field, three perturbation scenarios, and a wiring
detector: a self-test scenario whose expectations are known-false and
which MUST fail. The previous anti-vacuity check asserted only that the
walk produced a smell, which passes on well-known engine behavior
regardless of whether the harness wiring works.

First perturbation attempt produced zero signal -- progress does not
structurally parse ROADMAP.md, so a corrupted roadmap sailed through. A
perturbation that cannot fail is the same defect in a new costume.
Probes now target roadmap get-phase, and each mutated step runs a clean
baseline first so `mutationObserved` records whether the corruption
changed anything at all.

Also clears four review findings: classify() returned PROSE for exit-0
with empty stdout; `warnings` was structurally unpopulatable on the
success path (execFileSync discards it) and is now documented as
error-path-only; read-only-idempotence passed vacuously when asked to
check idempotence without the data to check it; the ADR miscounted the
oracles.

Discrimination matrix across 8 mutations x 6 commands: bom,
duplicate-phase-id and escaped-pipes are absorbed silently by every
probed surface, and progress / smart-entry / roadmap validate never
reacted to any mutation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): add path-containment guard for scenario-supplied targets

Security review found scenario-supplied paths joined to the temp project
with no containment check. step.mutate.target and agent.write keys were
validated only as non-empty strings, so a target of ../../../../etc/hosts
reached fs.unlinkSync / fs.writeFileSync / fs.symlinkSync outside the
project. The symlink mutation was worst: it read the traversed file, wrote
a sibling copy, deleted the original and symlinked it back.

Not exploitable today -- all shipped scenarios target .planning/ROADMAP.md
and scenarios are repo-committed, not runtime input. Fixed anyway: it is a
live primitive any future scenario or copied helper can reach.

Adds tests/qa/paths.cjs with resolveWithin(): rejects absolute paths, NUL
bytes and empty input, normalizes separators unconditionally, and requires
containment by path segment so a sibling like <base>-evil is not treated as
inside. Non-existent targets resolve via nearest existing ancestor rather
than falling back to a lexical compare. Scenario load now rejects traversing
or absolute targets up front.

oracles.cjs previously carried its own copy of the containment logic; both
now share paths.cjs, since a duplicated containment check is exactly the
divergence class this repo calls out.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): complete trajectory corpus, report emission, boundary-aware oracle

Adds the remaining trajectories and drives all 11 mutations end-to-end.
20 scenarios, 72 steps, 0 violations, 25 smells.

Adds qa-report.json with per-step verdicts and a copy-pasteable repro
command, plus --keep / GSD_QA_KEEP=1 to preserve a failing tree. A repro
line for a tree that was not preserved is marked NOT RUNNABLE rather than
emitting a command pointing at a deleted directory.

monotonic-progress is now boundary-aware. Two scenarios had been trimmed
to stop the oracle complaining at a milestone rollover, which destroys the
signal the trajectory exists to produce. Evidence: counters legitimately
reset to zero at milestone complete, but the payload milestone_version
lags until a new ROADMAP.md is written. So the oracle now scopes by
milestone plus workstream, keeps a same-scope decrease as a violation, and
records a boundary crossing as a smell. Both scenarios walk the real
boundary again.

Standards review fixes: oracle findings now carry a structured subject so
tests assert on typed fields instead of substring-matching the free-form
detail string, resolveWithin throws a typed EPATHESCAPE error, and the
absolute-path predicate scenario.cjs had re-implemented now comes from
paths.cjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): fix silently-vacuous fixtures and guard the class

Every fixture carried its #2371 provenance comment BEFORE the frontmatter
block, and extractFrontmatter returns {} when anything precedes the opening
---. So every scenario reading status/phase/name was operating on an empty
object and reporting green. Nine fixtures repositioned; the comment stays,
it just moves below the closing ---.

Both UAT fixtures lacked a parser-recognized result block, so
evaluateUatPassed saw checks.length===0 and could never return passed:true.
The uat-fail-then-remediate scenario could not have proven a remediation.
Its expect block only inspected blockers, which is empty before AND after,
which is why the corpus never noticed. Both fixtures now carry real result
blocks and the scenario asserts passed and no_uat_artifacts on each side of
the flip.

The actual deliverable is the guard: a fixture-integrity block asserting
every fixture with a frontmatter shape parses to a non-empty object, that
every fixture carries its provenance marker, and that the two UAT fixtures
produce opposite verdicts through the real evaluateUatPassed. The first
guard written required --- at byte 0, which would never have fired on the
regression it exists to prevent; it was rewritten and proven by deliberately
re-breaking a fixture.

No engine defect here. no_uat_artifacts means no parsed check items, not no
UAT files, and it was reporting correctly on fixtures that had none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): make the walk report — smell ratchet, baseline, CI job

The harness computed smells into a gitignored qa-report.json that nothing
read. In CI it surfaced nothing at all: violations failed the build, but the
half of the tool that says "this works and is still wrong" was inert. A QA
tool nobody hears is decoration.

Adds a ratchet on the same idiom this repo already uses three times over
(the regression-test-name allowlist, the emitted-drift acks, the size
baseline): a committed smell-baseline.json, per-PR acknowledgment fragments
under tests/qa/smell-acks/, and a ratchet script wired into CI.

The design invariant is preserved exactly. A smell still never fails a build
on its own merits. What fails is an UNACKNOWLEDGED NEW smell -- the absence
of a decision -- leaving an author two honest exits: fix it, or record a
fragment with a real reason. An empty reason is rejected. The baseline is
shrink-only, so a fixed smell must prune its entry. Violations remain
unacknowledgeable.

Fingerprints are composed only from stable fields (oracle id, scenario,
argv, subject discriminator) -- never temp paths, timestamps or counts.
Verified byte-identical across two runs in separate temp dirs; an unstable
fingerprint would have false-positived every CI run.

CI gains a qa-loop-walk job that runs the suite and the ratchet, uploads the
report with `if: always()` (it matters most when it failed), and renders a
summary a reviewer reads without downloading anything.

Also fixes the report runner invoking main() unconditionally on require, so
importing it double-ran every scenario and clobbered its own output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): every smell terminates in a defect or a fixed detector

The baseline accepted a smell with a free-text reason. That is a mechanism
for designing smells in -- an allowlist nobody revisits. The harness is
brand new, so nothing it found is inherited legacy; every finding is a
FIRST finding. Each must now terminate in exactly one of two states:

  REAL           -> an assigned defect, entry carries the issue number
  FALSE POSITIVE -> the detector is wrong and gets fixed, never baselined

There is no third "accepted with a good explanation" state, so the ratchet
now requires a positive-integer `issue` on every entry. A reason may remain
as a human note but can never substitute. `--update` refuses to invent
issue numbers: a new smell is written with `issue: null` and a TODO, and
the next plain run rejects it, forcing triage rather than accumulation.

Working the 21 existing entries through that rule found 16 were my own
detectors being wrong:

value-hygiene (10) flagged $.agents_dir, a field whose entire contract is
to point at the install tree outside any project. Fixed with a leaf-key
allowlist of contractually-external fields, verified as the only such key
in the init payload. Genuinely unexpected out-of-project paths still smell.

monotonic-progress (6) fired on legitimate boundary crossings -- milestone
v1.0 to v2.0, workstream beta to alpha -- and on one payload carrying no
scope fields at all, where a change cannot even be known. Scope changes now
reset silently and scope-less observations are skipped. The same-scope
decrease remains a violation; that is the real invariant and is regression-
guarded.

The five survivors are real and now tracked: soft-error-exit-zero (#2980),
untyped-success (#2979). Baseline 25 -> 5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): keep the ratchet out of the tarball, unpin the qa CI job

The remote matrix returned failed -- 3 unique failures, identical on
node22 and node24, both root causes in this branch's own diff.

The ratchet lives under scripts/, which ships in the npm tarball, and it
requires three modules under tests/, which does not. In a published
install it is MODULE_NOT_FOUND at load. This is exactly the class the
#2858 guard was added to catch, and it caught it. Fixed the way #2858
fixed the same shape for its own repo-only CI script: a targeted files[]
negation, so the ratchet stays in the repo for CI and out of the tarball.
Not solved by moving or inlining the required modules -- the ratchet must
keep using the same code the harness uses, or the two drift.

Verified both directions: the script is no longer in the pack list, and
build-hooks.js, fix-slash-commands.cjs and gen-capability-registry.cjs are
all still shipped. Over-negating there would have broken installs, since
bin/install.js requires them.

The qa-loop-walk job also carried CI_REBASE_BASE_SHA copied from a
neighbouring job without the paired GSD_EMITTED_BASE, which the #2854
invariant forbids by name: diverging them makes the differential compare a
tree against a baseline from a different commit. The job runs only the qa
suite and the ratchet and invokes no emitted-attribution test, so it needs
no rebase-pinned base at all -- the step was removed rather than paired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#2966): stop monotonic-progress going blind on scope-less payloads

The full remote suite caught a false NEGATIVE I introduced while fixing a
false positive. Silencing the boundary-crossing noise had made the oracle
skip ANY observation lacking milestone fields -- so a minimal payload like
{total_summaries: n} produced no violation at all, and the oracle stopped
catching the exact defect it exists to catch. For a QA tool that is
strictly worse than the noise it replaced.

Scope is only indeterminate when the two observations DISAGREE about
having it:

  both scoped, same scope, decrease -> VIOLATION
  both scoped, different scope      -> reset silently
  NEITHER scoped, decrease          -> VIOLATION   (the regression)
  mixed                             -> skip the comparison

Implementing the mixed case surfaced a second blind spot: advancing the
reference point on a skipped pair lets a scope-less observation sitting
between two same-scope ones mask a real decrease. Mixed now leaves the
reference untouched. All four branches carry explicit coverage; only one
did before, which is why this shipped.

The self-test that failed was right and the code was wrong, so the code
moved. Corpus behavior is unchanged: still 5 smells, 0 new, 0 stale, 0
violations.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 16:13:28 -04:00

568 lines
24 KiB
JavaScript

'use strict';
/**
* scenario.cjs — the QA-walk scenario DSL interpreter.
*
* WHY THIS FILE EXISTS
* ────────────────────
* A scenario is a small JSON document describing a sequence of loop-host
* points, artifacts an "agent" would write at each point, and `gsd-tools`
* invocations to run there. `loadScenario` validates the JSON shape so a
* malformed scenario fails fast and loud (never a silent no-op walk);
* `runScenario` drives a `LoopWalk` through the steps, evaluating both
* declared `expect` assertions and the shared `oracles.cjs` checks at every
* step, and NEVER throws on a step failure — a single bad step is recorded
* and the walk continues, so one broken step can never hide the rest of the
* walk's results.
*/
const fs = require('node:fs');
const path = require('node:path');
const { resolveRef } = require('./fixtures/index.cjs');
const { LOOP_HOST_CONTRACT } = require('../../gsd-core/bin/lib/loop-host-contract.cjs');
const { MUTATIONS, apply, NOOP } = require('./mutations.cjs');
const { resolveWithin, isAbsoluteLike, hasTraversalSegment } = require('./paths.cjs');
/**
* True when `relPath` is absolute or contains a `..` path segment. Delegates to
* `paths.cjs`'s `isAbsoluteLike` / `hasTraversalSegment` — the single source of truth for
* both predicates — rather than keeping a second copy here. Used at `validateScenario` LOAD
* time, ahead of `resolveWithin`'s own (equally strict) runtime check: failing a malformed
* scenario fast and loud at load time is better than failing at apply time — the scenario is
* malformed, not the run.
*
* @param {string} relPath
* @returns {boolean}
*/
function isTraversalOrAbsolute(relPath) {
if (typeof relPath !== 'string' || relPath === '') return true;
return isAbsoluteLike(relPath) || hasTraversalSegment(relPath);
}
/** Scenario `fixture` must be one of these — see `loop-walk.cjs` `FIXTURE_BUILDERS`. */
const VALID_FIXTURES = new Set(['greenfield', 'planning', 'seeded']);
/** Every known mutation id, derived from `mutations.cjs` (never hardcoded — see that module's `MUTATIONS`). */
const VALID_MUTATION_IDS = new Set(MUTATIONS.map((m) => m.id));
/** `id -> {id, kind, describe}` lookup, derived once from `MUTATIONS`. */
const MUTATION_BY_ID = new Map(MUTATIONS.map((m) => [m.id, m]));
/**
* The legal set of `step.at` values, derived from the generated loop-host
* contract rather than hardcoded — see `loop-walk.cjs`'s `LOOP_STEPS` header
* comment for why a hardcoded copy would silently drift from the generator.
*
* @returns {Set<string>}
*/
function getLegalPoints() {
const points = new Set();
for (const entry of LOOP_HOST_CONTRACT) {
for (const point of entry.points) points.add(point);
}
return points;
}
/**
* Structural (deep) equality for JSON-ish values. Cycle-tolerant via a
* `WeakMap` pairing "already compared" object references.
*
* @param {unknown} a
* @param {unknown} b
* @param {WeakMap<object, unknown>} [seen]
* @returns {boolean}
*/
function deepEqual(a, b, seen = new WeakMap()) {
if (Object.is(a, b)) return true;
if (a === null || b === null) return false;
if (typeof a !== 'object' || typeof b !== 'object') return false;
if (Array.isArray(a) !== Array.isArray(b)) return false;
if (seen.get(a) === b) return true;
seen.set(a, b);
const aKeys = Object.keys(a);
const bKeys = Object.keys(b);
if (aKeys.length !== bKeys.length) return false;
for (const key of aKeys) {
if (!Object.prototype.hasOwnProperty.call(b, key)) return false;
if (!deepEqual(a[key], b[key], seen)) return false;
}
return true;
}
/**
* Look up a dot-path (e.g. `"a.b.c"` or `"a.b[0].c"`) inside a JSON-ish value.
*
* @param {unknown} value
* @param {string} dotPath
* @returns {{ found: boolean, value: unknown }}
*/
function dotGet(value, dotPath) {
const segments = dotPath
.split('.')
.flatMap((seg) => {
const parts = [];
const re = /^([^[\]]*)((?:\[\d+\])*)$/;
const m = re.exec(seg);
if (!m) return [seg];
if (m[1] !== '') parts.push(m[1]);
const indices = m[2].match(/\[\d+\]/g) || [];
for (const idx of indices) parts.push(Number(idx.slice(1, -1)));
return parts;
});
let current = value;
for (const segment of segments) {
if (current === null || current === undefined) return { found: false, value: undefined };
if (typeof current !== 'object') return { found: false, value: undefined };
const key = segment;
if (Array.isArray(current)) {
if (typeof key !== 'number' || key < 0 || key >= current.length) {
return { found: false, value: undefined };
}
current = current[key];
continue;
}
if (!Object.prototype.hasOwnProperty.call(current, key)) return { found: false, value: undefined };
current = current[key];
}
return { found: true, value: current };
}
/**
* Validate a parsed scenario object, throwing on the first violation with a
* message naming the offending field/step.
*
* @param {unknown} scenario
* @param {string} [sourceLabel] e.g. an absolute file path, for error context.
* @returns {object} `scenario`, unmodified, once fully validated.
*/
function validateScenario(scenario, sourceLabel) {
const label = sourceLabel ? ` (from ${sourceLabel})` : '';
if (!scenario || typeof scenario !== 'object' || Array.isArray(scenario)) {
throw new Error(`loadScenario: scenario${label} must be a JSON object, got ${JSON.stringify(scenario)}`);
}
if (typeof scenario.name !== 'string' || scenario.name.trim() === '') {
throw new Error(`loadScenario: "name"${label} must be a non-empty string, got ${JSON.stringify(scenario.name)}`);
}
if (!VALID_FIXTURES.has(scenario.fixture)) {
throw new Error(
`loadScenario: "fixture"${label} must be one of ${[...VALID_FIXTURES].join(', ')}, got ${JSON.stringify(scenario.fixture)}`,
);
}
if (!Array.isArray(scenario.steps) || scenario.steps.length === 0) {
throw new Error(`loadScenario: "steps"${label} must be a non-empty array — an empty scenario is an error, not a silent pass`);
}
if (scenario.selfTest !== undefined && typeof scenario.selfTest !== 'boolean') {
throw new Error(`loadScenario: "selfTest"${label} must be a boolean, got ${JSON.stringify(scenario.selfTest)}`);
}
const legalPoints = getLegalPoints();
scenario.steps.forEach((step, index) => {
const where = `steps[${index}]${label}`;
if (!step || typeof step !== 'object' || Array.isArray(step)) {
throw new Error(`loadScenario: ${where} must be an object, got ${JSON.stringify(step)}`);
}
if (typeof step.at !== 'string' || !legalPoints.has(step.at)) {
throw new Error(
`loadScenario: ${where}.at is ${JSON.stringify(step.at)}, which is not a legal loop-host-contract point `
+ `(legal points: ${[...legalPoints].sort().join(', ')})`,
);
}
if (step.run !== undefined) {
if (!Array.isArray(step.run)) {
throw new Error(`loadScenario: ${where}.run must be an array of argv arrays, got ${JSON.stringify(step.run)}`);
}
step.run.forEach((argv, ri) => {
if (!Array.isArray(argv) || argv.length === 0 || !argv.every((t) => typeof t === 'string')) {
throw new Error(`loadScenario: ${where}.run[${ri}] must be a non-empty array of strings, got ${JSON.stringify(argv)}`);
}
});
}
if (step.expect !== undefined) {
if (!Array.isArray(step.expect)) {
throw new Error(`loadScenario: ${where}.expect must be an array, got ${JSON.stringify(step.expect)}`);
}
step.expect.forEach((exp, ei) => {
const isValid = exp && typeof exp === 'object' && !Array.isArray(exp)
&& typeof exp.path === 'string' && exp.path !== ''
&& Object.prototype.hasOwnProperty.call(exp, 'is');
if (!isValid) {
throw new Error(
`loadScenario: ${where}.expect[${ei}] must be {path: <non-empty dot-path string>, is: <value>}, got ${JSON.stringify(exp)}`,
);
}
});
}
if (step.jsonErrors !== undefined && typeof step.jsonErrors !== 'boolean') {
throw new Error(`loadScenario: ${where}.jsonErrors must be a boolean, got ${JSON.stringify(step.jsonErrors)}`);
}
if (step.mutate !== undefined) {
if (!step.mutate || typeof step.mutate !== 'object' || Array.isArray(step.mutate)) {
throw new Error(`loadScenario: ${where}.mutate must be an object, got ${JSON.stringify(step.mutate)}`);
}
if (typeof step.mutate.id !== 'string' || !VALID_MUTATION_IDS.has(step.mutate.id)) {
throw new Error(
`loadScenario: ${where}.mutate.id is ${JSON.stringify(step.mutate.id)}, which is not a known mutation id `
+ `(valid ids: ${[...VALID_MUTATION_IDS].sort().join(', ')})`,
);
}
if (typeof step.mutate.target !== 'string' || step.mutate.target === '') {
throw new Error(`loadScenario: ${where}.mutate.target must be a non-empty project-relative path string, got ${JSON.stringify(step.mutate.target)}`);
}
if (isTraversalOrAbsolute(step.mutate.target)) {
throw new Error(
`loadScenario: ${where}.mutate.target ${JSON.stringify(step.mutate.target)} must be project-relative `
+ '— absolute paths and ".." segments are rejected at load time',
);
}
if (step.mutate.targetBytes !== undefined) {
const { targetBytes } = step.mutate;
if (typeof targetBytes !== 'number' || !Number.isFinite(targetBytes) || targetBytes < 0) {
throw new Error(`loadScenario: ${where}.mutate.targetBytes must be a non-negative finite number, got ${JSON.stringify(targetBytes)}`);
}
}
}
if (step.agent !== undefined) {
if (!step.agent || typeof step.agent !== 'object' || Array.isArray(step.agent)) {
throw new Error(`loadScenario: ${where}.agent must be an object, got ${JSON.stringify(step.agent)}`);
}
if (step.agent.write !== undefined) {
if (!step.agent.write || typeof step.agent.write !== 'object' || Array.isArray(step.agent.write)) {
throw new Error(`loadScenario: ${where}.agent.write must be an object, got ${JSON.stringify(step.agent.write)}`);
}
for (const [relPath, ref] of Object.entries(step.agent.write)) {
if (typeof ref !== 'string' || ref === '') {
throw new Error(`loadScenario: ${where}.agent.write["${relPath}"] must be a non-empty ref string, got ${JSON.stringify(ref)}`);
}
if (isTraversalOrAbsolute(relPath)) {
throw new Error(
`loadScenario: ${where}.agent.write key ${JSON.stringify(relPath)} must be project-relative `
+ '— absolute paths and ".." segments are rejected at load time',
);
}
}
}
}
});
return scenario;
}
/**
* Load and validate a scenario JSON file.
*
* @param {string} absPathToJson
* @returns {object} the validated scenario object.
* @throws {Error} on missing/unparsable file or any validation violation
* (message names the offending field).
*/
function loadScenario(absPathToJson) {
let text;
try {
text = fs.readFileSync(absPathToJson, 'utf-8');
} catch (err) {
throw new Error(`loadScenario: cannot read "${absPathToJson}": ${err && err.message}`);
}
let parsed;
try {
parsed = JSON.parse(text);
} catch (err) {
throw new Error(`loadScenario: "${absPathToJson}" is not valid JSON: ${err && err.message}`);
}
return validateScenario(parsed, path.resolve(absPathToJson));
}
/**
* Evaluate a step's `expect` array against a `RunResult`.
*
* @param {Array<{path: string, is: unknown}>} expectations
* @param {{ json: unknown }} result
* @returns {string[]} human-readable failure descriptions, empty when all pass.
*/
function evaluateExpectations(expectations, result) {
const failures = [];
for (const exp of expectations || []) {
const { found, value } = dotGet(result ? result.json : undefined, exp.path);
if (!found) {
failures.push(`path "${exp.path}": not found in result.json`);
continue;
}
if (!deepEqual(value, exp.is)) {
failures.push(`path "${exp.path}": expected ${JSON.stringify(exp.is)}, got ${JSON.stringify(value)}`);
}
}
return failures;
}
/**
* Drive a `LoopWalk` through every step of `scenario`, evaluating `expect`
* assertions and the shared oracle set at each step. Never throws on a step
* failure — failures are recorded on that step's report entry and the walk
* continues, so one bad step never hides the rest.
*
* @param {object} scenario a scenario already validated by `loadScenario`.
* @param {{
* LoopWalk: { create(opts: object): object },
* runOracles: (ctx: object) => { passed: string[], violations: {id:string,detail:string}[], smells: {id:string,detail:string}[], failed: {id:string,detail:string}[] },
* liveCommands?: string[],
* keep?: boolean,
* }} opts `keep` (default `false`) preserves the walk's temp project instead
* of removing it in `finally` — see `LoopWalk#cleanup`. Also honored via
* the `GSD_QA_KEEP=1` environment variable (an `||`, not an override: either
* one being truthy keeps the tree), since a CI operator invoking this
* through a shell cannot pass a JS option.
* @returns {{
* name: string,
* steps: Array<{ at: string, argv: string[], kind: string|null, expectFailures: string[], oracleFailures: {id:string,detail:string}[], smells: {id:string,detail:string}[], mutation: {id:string,target:string}|null, mutationNoop: boolean, mutationObserved: boolean }>,
* ok: boolean,
* smellSummary: Array<{ id: string, count: number, examples: string[] }>,
* preservedDir?: string,
* }}
*/
function runScenario(scenario, opts) {
const { LoopWalk, runOracles, liveCommands = [], keep = false } = opts || {};
if (!LoopWalk || typeof LoopWalk.create !== 'function') {
throw new Error('runScenario: opts.LoopWalk (with a create() factory) is required');
}
if (typeof runOracles !== 'function') {
throw new Error('runScenario: opts.runOracles (function) is required');
}
const shouldKeep = keep || process.env.GSD_QA_KEEP === '1';
const walk = LoopWalk.create({ fixture: scenario.fixture });
/** @type {object[]} */
const history = [];
/** @type {Array<{at:string, argv:string[], kind:string|null, expectFailures:string[], oracleFailures:{id:string,detail:string}[], smells:{id:string,detail:string}[]}>} */
const steps = [];
let preservedDir;
try {
for (const step of scenario.steps) {
try {
if (step.agent && step.agent.write) {
for (const [relPath, ref] of Object.entries(step.agent.write)) {
walk.writeArtifact(relPath, resolveRef(ref));
}
}
// Anti-vacuity for perturbations: BEFORE the mutation is applied,
// run this step's own `run` sequence once against the CLEAN (as-yet
// unmutated) world and keep only the last result's `kind`/`json` —
// never pushed to `history`, never fed to oracles, never counted in
// `statsBefore/After`. This baseline exists solely so that, once the
// mutated run happens below, the two can be compared: a mutation
// that changes nothing observable is indistinguishable from a
// mutation that was never applied, so `mutationObserved` gives that
// distinction a name instead of leaving it implicit in a diff nobody
// looks at.
let cleanBaseline = null;
if (step.mutate) {
const jsonErrorModeForBaseline = step.jsonErrors !== false;
const baselineOptions = { jsonErrors: jsonErrorModeForBaseline };
const baselineRuns = Array.isArray(step.run) ? step.run : [];
let baselineResult = null;
for (const argv of baselineRuns) {
baselineResult = walk.run(...argv, baselineOptions);
}
if (baselineResult) {
cleanBaseline = { kind: baselineResult.kind, json: baselineResult.json };
}
}
// Mutation is applied AFTER `agent.write` and BEFORE `run` — a step
// can write a valid artifact and then corrupt it, so `run` observes
// the corrupted world exactly as a real engine invocation would.
let mutationRecord = null;
let mutationNoop = false;
if (step.mutate) {
const { id, target, targetBytes } = step.mutate;
const entry = MUTATION_BY_ID.get(id);
const absTarget = resolveWithin(walk.dir, target);
if (entry.kind === 'content') {
// This `readFileSync` is harness plumbing, not a test assertion —
// it reads a planning ARTIFACT the walk itself just wrote so the
// mutation catalog's pure string transforms have input to work
// on. It is never string-matched/asserted against; the mutated
// bytes are written straight back to disk for `run` to react to.
// Do NOT "fix" this into a stat-only check — see `mutations.cjs`
// and this file's header for why oracles must never read SUT
// output, which does not apply to this harness-owned write path.
const before = fs.readFileSync(absTarget, 'utf-8');
const mutateOpts = targetBytes !== undefined ? { targetBytes } : undefined;
const after = apply(id, before, mutateOpts);
if (after === NOOP) {
mutationNoop = true;
} else {
fs.writeFileSync(absTarget, after, 'utf-8');
}
} else {
apply(id, { dir: walk.dir, relPath: target });
}
mutationRecord = { id, target };
}
const readOnly = !!step.readOnly;
const runs = Array.isArray(step.run) ? step.run : [];
// `step.jsonErrors === false` opts a step into the human-invocation
// path (`gsd_run <cmd>` without `--json-errors`); any other value
// (including undefined) keeps `LoopWalk#run`'s own default of true.
const jsonErrorMode = step.jsonErrors !== false;
const runOptions = { jsonErrors: jsonErrorMode };
const statsBefore = walk.statSnapshot();
let result = null;
let lastArgv = [];
for (const argv of runs) {
result = walk.run(...argv, runOptions);
lastArgv = argv;
}
const statsAfter = walk.statSnapshot();
let repeatResult = null;
if (readOnly && lastArgv.length > 0) {
repeatResult = walk.run(...lastArgv, runOptions);
}
const expectFailures = evaluateExpectations(step.expect, result);
const { violations: oracleFailures, smells } = runOracles({
result,
repeatResult,
statsBefore,
statsAfter,
history,
liveCommands,
readOnly,
projectDir: walk.dir,
jsonErrorMode,
});
if (result) history.push(result);
// `mutationObserved` is only meaningful for a mutated step: it is
// `true` when the mutated run's result differs from the clean
// baseline captured above, by `kind` OR by deep-inequality of
// `json` — a `kind`-only comparison would miss a mutation that
// keeps `kind: "json"` but silently changes the payload (e.g.
// `found: true` -> `found: false`), which is exactly the class of
// corruption a perturbation scenario exists to catch.
const mutationObserved = step.mutate && cleanBaseline
? (cleanBaseline.kind !== (result ? result.kind : null) || !deepEqual(cleanBaseline.json, result ? result.json : null))
: false;
steps.push({
at: step.at,
argv: lastArgv,
kind: result ? result.kind : null,
expectFailures,
oracleFailures,
smells,
mutation: mutationRecord,
mutationNoop,
mutationObserved,
});
} catch (err) {
steps.push({
at: step && step.at,
argv: [],
kind: null,
expectFailures: [],
oracleFailures: [{ id: 'step-exception', detail: `${err && err.message}` }],
smells: [],
mutation: null,
mutationNoop: false,
mutationObserved: false,
});
}
}
} finally {
preservedDir = walk.cleanup({ keep: shouldKeep });
}
// A SMELL is evidence, never a build break: `ok` is derived from
// `expectFailures` + `oracleFailures` (violations) ONLY. A step whose sole
// findings are smells still counts as `ok`.
const ok = steps.every((s) => s.expectFailures.length === 0 && s.oracleFailures.length === 0);
const smellSummary = summarizeSmells(steps);
return {
name: scenario.name,
steps,
ok,
smellSummary,
...(preservedDir ? { preservedDir } : {}),
};
}
/**
* Group every step's `smells` by oracle `id` across the whole walk, so a
* reader sees "this oracle fired N times, here are up to 3 examples" rather
* than a flat wall of per-step repeats.
*
* @param {Array<{smells: {id:string, detail:string}[]}>} steps
* @returns {Array<{id: string, count: number, examples: string[]}>}
*/
function summarizeSmells(steps) {
/** @type {Map<string, string[]>} */
const byId = new Map();
for (const step of steps) {
for (const smell of step.smells || []) {
const list = byId.get(smell.id) || [];
list.push(smell.detail);
byId.set(smell.id, list);
}
}
return [...byId.entries()].map(([id, details]) => ({
id,
count: details.length,
examples: details.slice(0, 3),
}));
}
/** Absolute path to the wiring self-test scenario — see `assertWiringIsLive`. */
const SELF_TEST_SCENARIO_PATH = path.join(__dirname, 'scenarios', '_selftest-must-fail.json');
/**
* The REAL wiring detector for this harness's assertion machinery.
*
* `totalSmells > 0` on a happy-path walk proves almost nothing: it leans on
* well-known engine behaviors that fire regardless of whether `expect` /
* oracle plumbing actually works. This function instead loads the
* deliberately-broken `scenarios/_selftest-must-fail.json` scenario — whose
* `expect` block asserts something KNOWN FALSE about a real command — and
* runs it for real. If the assertion machinery is wired correctly, the run
* MUST fail (`ok === false`, non-empty `expectFailures`); if it silently
* passes, `expect` is not actually being evaluated and this function throws
* naming that fact, rather than letting a broken harness report a clean bill
* of health.
*
* @param {{ LoopWalk: object, runOracles: Function, liveCommands?: string[] }} opts
* @returns {ReturnType<typeof runScenario>} the self-test's own report, for
* callers that want to inspect or log it.
* @throws {Error} when the self-test scenario is missing `"selfTest": true`,
* or when it does NOT fail — either case means the expect/oracle assertion
* machinery is not provably wired.
*/
function assertWiringIsLive(opts) {
const scenario = loadScenario(SELF_TEST_SCENARIO_PATH);
if (scenario.selfTest !== true) {
throw new Error(`assertWiringIsLive: "${SELF_TEST_SCENARIO_PATH}" is missing "selfTest": true`);
}
const report = runScenario(scenario, opts);
if (report.ok !== false) {
throw new Error(
'assertWiringIsLive: the self-test scenario (a KNOWN-FALSE expectation) did not fail — '
+ 'the expect/oracle assertion machinery is not wired',
);
}
const hasExpectFailure = report.steps.some((s) => s.expectFailures.length > 0);
if (!hasExpectFailure) {
throw new Error(
'assertWiringIsLive: the self-test scenario failed via oracleFailures but recorded no expectFailures — '
+ 'the "expect" assertion machinery specifically is not provably wired',
);
}
return report;
}
module.exports = { loadScenario, runScenario, deepEqual, dotGet, assertWiringIsLive };