* test(#2966): loop QA walk — drive real scenarios across all five loop steps Adds a headless walk that carries accumulating project state across discuss -> plan -> execute -> verify -> ship against one temp project, layered over the existing tests/helpers.cjs runGsdTools substrate. Findings carry severity. A violation breaks a stated contract and fails the build; a smell is legal under today's implementation but structurally questionable, is recorded, and never reddens CI. Without that split an oracle set derived from current behavior can only ever confirm current behavior -- the harness could not say "this works and is still wrong". The end-to-end test asserts the walk produces at least one smell: a QA harness that reports nothing on a first run against a real engine is far more likely mis-specified than the engine is perfect. It deliberately does not pin smell ids or counts, which would re-freeze current behavior. First run against the real engine: 0 violations, 3 smell classes -- init returns agents_dir outside the project tree; smart-entry emits prose unconditionally so routing cannot be asserted; state-snapshot reports a missing STATE.md through a payload key with exit 0. Also fixes tests/fixtures/index.cjs: createFixture with git:true and planning:false staged nothing, so the commit failed with "nothing to commit". That combination was unreachable until greenfield needed it. Extends RULESET.TESTS.feedback-loop-convergence from estimation to the loop itself. Design lock: docs/adr/2966-loop-qa-walk.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2966): wire fault injection, make perturbations discriminating Independent review found tests/qa/mutations.cjs entirely unwired: 462 lines exercised only by their own unit tests, with no mutation hook in the scenario DSL and no scenario applying one, while the module header and the ADR described fault injection in the present tense. Dead code documented as live. Adds a `mutate` step field, three perturbation scenarios, and a wiring detector: a self-test scenario whose expectations are known-false and which MUST fail. The previous anti-vacuity check asserted only that the walk produced a smell, which passes on well-known engine behavior regardless of whether the harness wiring works. First perturbation attempt produced zero signal -- progress does not structurally parse ROADMAP.md, so a corrupted roadmap sailed through. A perturbation that cannot fail is the same defect in a new costume. Probes now target roadmap get-phase, and each mutated step runs a clean baseline first so `mutationObserved` records whether the corruption changed anything at all. Also clears four review findings: classify() returned PROSE for exit-0 with empty stdout; `warnings` was structurally unpopulatable on the success path (execFileSync discards it) and is now documented as error-path-only; read-only-idempotence passed vacuously when asked to check idempotence without the data to check it; the ADR miscounted the oracles. Discrimination matrix across 8 mutations x 6 commands: bom, duplicate-phase-id and escaped-pipes are absorbed silently by every probed surface, and progress / smart-entry / roadmap validate never reacted to any mutation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2966): add path-containment guard for scenario-supplied targets Security review found scenario-supplied paths joined to the temp project with no containment check. step.mutate.target and agent.write keys were validated only as non-empty strings, so a target of ../../../../etc/hosts reached fs.unlinkSync / fs.writeFileSync / fs.symlinkSync outside the project. The symlink mutation was worst: it read the traversed file, wrote a sibling copy, deleted the original and symlinked it back. Not exploitable today -- all shipped scenarios target .planning/ROADMAP.md and scenarios are repo-committed, not runtime input. Fixed anyway: it is a live primitive any future scenario or copied helper can reach. Adds tests/qa/paths.cjs with resolveWithin(): rejects absolute paths, NUL bytes and empty input, normalizes separators unconditionally, and requires containment by path segment so a sibling like <base>-evil is not treated as inside. Non-existent targets resolve via nearest existing ancestor rather than falling back to a lexical compare. Scenario load now rejects traversing or absolute targets up front. oracles.cjs previously carried its own copy of the containment logic; both now share paths.cjs, since a duplicated containment check is exactly the divergence class this repo calls out. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2966): complete trajectory corpus, report emission, boundary-aware oracle Adds the remaining trajectories and drives all 11 mutations end-to-end. 20 scenarios, 72 steps, 0 violations, 25 smells. Adds qa-report.json with per-step verdicts and a copy-pasteable repro command, plus --keep / GSD_QA_KEEP=1 to preserve a failing tree. A repro line for a tree that was not preserved is marked NOT RUNNABLE rather than emitting a command pointing at a deleted directory. monotonic-progress is now boundary-aware. Two scenarios had been trimmed to stop the oracle complaining at a milestone rollover, which destroys the signal the trajectory exists to produce. Evidence: counters legitimately reset to zero at milestone complete, but the payload milestone_version lags until a new ROADMAP.md is written. So the oracle now scopes by milestone plus workstream, keeps a same-scope decrease as a violation, and records a boundary crossing as a smell. Both scenarios walk the real boundary again. Standards review fixes: oracle findings now carry a structured subject so tests assert on typed fields instead of substring-matching the free-form detail string, resolveWithin throws a typed EPATHESCAPE error, and the absolute-path predicate scenario.cjs had re-implemented now comes from paths.cjs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2966): fix silently-vacuous fixtures and guard the class Every fixture carried its #2371 provenance comment BEFORE the frontmatter block, and extractFrontmatter returns {} when anything precedes the opening ---. So every scenario reading status/phase/name was operating on an empty object and reporting green. Nine fixtures repositioned; the comment stays, it just moves below the closing ---. Both UAT fixtures lacked a parser-recognized result block, so evaluateUatPassed saw checks.length===0 and could never return passed:true. The uat-fail-then-remediate scenario could not have proven a remediation. Its expect block only inspected blockers, which is empty before AND after, which is why the corpus never noticed. Both fixtures now carry real result blocks and the scenario asserts passed and no_uat_artifacts on each side of the flip. The actual deliverable is the guard: a fixture-integrity block asserting every fixture with a frontmatter shape parses to a non-empty object, that every fixture carries its provenance marker, and that the two UAT fixtures produce opposite verdicts through the real evaluateUatPassed. The first guard written required --- at byte 0, which would never have fired on the regression it exists to prevent; it was rewritten and proven by deliberately re-breaking a fixture. No engine defect here. no_uat_artifacts means no parsed check items, not no UAT files, and it was reporting correctly on fixtures that had none. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2966): make the walk report — smell ratchet, baseline, CI job The harness computed smells into a gitignored qa-report.json that nothing read. In CI it surfaced nothing at all: violations failed the build, but the half of the tool that says "this works and is still wrong" was inert. A QA tool nobody hears is decoration. Adds a ratchet on the same idiom this repo already uses three times over (the regression-test-name allowlist, the emitted-drift acks, the size baseline): a committed smell-baseline.json, per-PR acknowledgment fragments under tests/qa/smell-acks/, and a ratchet script wired into CI. The design invariant is preserved exactly. A smell still never fails a build on its own merits. What fails is an UNACKNOWLEDGED NEW smell -- the absence of a decision -- leaving an author two honest exits: fix it, or record a fragment with a real reason. An empty reason is rejected. The baseline is shrink-only, so a fixed smell must prune its entry. Violations remain unacknowledgeable. Fingerprints are composed only from stable fields (oracle id, scenario, argv, subject discriminator) -- never temp paths, timestamps or counts. Verified byte-identical across two runs in separate temp dirs; an unstable fingerprint would have false-positived every CI run. CI gains a qa-loop-walk job that runs the suite and the ratchet, uploads the report with `if: always()` (it matters most when it failed), and renders a summary a reviewer reads without downloading anything. Also fixes the report runner invoking main() unconditionally on require, so importing it double-ran every scenario and clobbered its own output. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2966): every smell terminates in a defect or a fixed detector The baseline accepted a smell with a free-text reason. That is a mechanism for designing smells in -- an allowlist nobody revisits. The harness is brand new, so nothing it found is inherited legacy; every finding is a FIRST finding. Each must now terminate in exactly one of two states: REAL -> an assigned defect, entry carries the issue number FALSE POSITIVE -> the detector is wrong and gets fixed, never baselined There is no third "accepted with a good explanation" state, so the ratchet now requires a positive-integer `issue` on every entry. A reason may remain as a human note but can never substitute. `--update` refuses to invent issue numbers: a new smell is written with `issue: null` and a TODO, and the next plain run rejects it, forcing triage rather than accumulation. Working the 21 existing entries through that rule found 16 were my own detectors being wrong: value-hygiene (10) flagged $.agents_dir, a field whose entire contract is to point at the install tree outside any project. Fixed with a leaf-key allowlist of contractually-external fields, verified as the only such key in the init payload. Genuinely unexpected out-of-project paths still smell. monotonic-progress (6) fired on legitimate boundary crossings -- milestone v1.0 to v2.0, workstream beta to alpha -- and on one payload carrying no scope fields at all, where a change cannot even be known. Scope changes now reset silently and scope-less observations are skipped. The same-scope decrease remains a violation; that is the real invariant and is regression- guarded. The five survivors are real and now tracked: soft-error-exit-zero (#2980), untyped-success (#2979). Baseline 25 -> 5. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2966): keep the ratchet out of the tarball, unpin the qa CI job The remote matrix returned failed -- 3 unique failures, identical on node22 and node24, both root causes in this branch's own diff. The ratchet lives under scripts/, which ships in the npm tarball, and it requires three modules under tests/, which does not. In a published install it is MODULE_NOT_FOUND at load. This is exactly the class the #2858 guard was added to catch, and it caught it. Fixed the way #2858 fixed the same shape for its own repo-only CI script: a targeted files[] negation, so the ratchet stays in the repo for CI and out of the tarball. Not solved by moving or inlining the required modules -- the ratchet must keep using the same code the harness uses, or the two drift. Verified both directions: the script is no longer in the pack list, and build-hooks.js, fix-slash-commands.cjs and gen-capability-registry.cjs are all still shipped. Over-negating there would have broken installs, since bin/install.js requires them. The qa-loop-walk job also carried CI_REBASE_BASE_SHA copied from a neighbouring job without the paired GSD_EMITTED_BASE, which the #2854 invariant forbids by name: diverging them makes the differential compare a tree against a baseline from a different commit. The job runs only the qa suite and the ratchet and invokes no emitted-attribution test, so it needs no rebase-pinned base at all -- the step was removed rather than paired. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2966): stop monotonic-progress going blind on scope-less payloads The full remote suite caught a false NEGATIVE I introduced while fixing a false positive. Silencing the boundary-crossing noise had made the oracle skip ANY observation lacking milestone fields -- so a minimal payload like {total_summaries: n} produced no violation at all, and the oracle stopped catching the exact defect it exists to catch. For a QA tool that is strictly worse than the noise it replaced. Scope is only indeterminate when the two observations DISAGREE about having it: both scoped, same scope, decrease -> VIOLATION both scoped, different scope -> reset silently NEITHER scoped, decrease -> VIOLATION (the regression) mixed -> skip the comparison Implementing the mixed case surfaced a second blind spot: advancing the reference point on a skipped pair lets a scope-less observation sitting between two same-scope ones mask a real decrease. Mixed now leaves the reference untouched. All four branches carry explicit coverage; only one did before, which is why this shipped. The self-test that failed was right and the code was wrong, so the code moved. Corpus behavior is unchanged: still 5 smells, 0 new, 0 stale, 0 violations. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
678 lines
27 KiB
JavaScript
678 lines
27 KiB
JavaScript
#!/usr/bin/env node
|
|
'use strict';
|
|
|
|
/**
|
|
* qa-smell-ratchet.cjs — turn a QA-walk "smell" into a decision (#2966).
|
|
*
|
|
* WHY THIS FILE EXISTS
|
|
* ────────────────────
|
|
* `tests/qa/run-report.cjs` computes "smells" — legal-but-questionable engine
|
|
* behavior (see `oracles.cjs`'s `SEVERITY.SMELL`) — and writes them into a
|
|
* gitignored `qa-report.json` that nothing reads. In CI, that means every
|
|
* smell is invisible: a NEW one can appear silently and nobody notices. This
|
|
* script is the pipeline that turns a smell into a decision.
|
|
*
|
|
* ══════════════════════════════════════════════════════════════════════════
|
|
* THE DESIGN INVARIANT (read this before touching anything below)
|
|
* ══════════════════════════════════════════════════════════════════════════
|
|
* A smell must NEVER fail a build on its own merits. What fails is an
|
|
* UNACKNOWLEDGED NEW smell — i.e. the absence of a human decision.
|
|
* Existing/known smells stay green forever.
|
|
*
|
|
* Concretely, that means:
|
|
* - A smell whose fingerprint (`tests/qa/smell-fingerprint.cjs`) is already
|
|
* recorded in `tests/qa/smell-baseline.json` OR in any fragment under
|
|
* `tests/qa/smell-acks/` is KNOWN and never fails the build, no matter
|
|
* how many times it fires or how bad it sounds.
|
|
* - A smell whose fingerprint has never been seen before is NEW, and fails
|
|
* the build — not because the behavior is wrong (it may be perfectly
|
|
* fine), but because nobody has looked at it and said so in writing.
|
|
* - The baseline is SHRINK-ONLY: an entry that stops firing (the engine
|
|
* was fixed, or the scenario changed) becomes STALE and ALSO fails the
|
|
* build, forcing `--update` to prune it. A baseline that only ever grows
|
|
* would let acknowledgments outlive the behavior they describe.
|
|
* - A VIOLATION (`SEVERITY.VIOLATION` — the engine broke a documented
|
|
* contract) is a completely different thing and is NEVER acknowledgeable
|
|
* through this mechanism: it always fails, baseline or no baseline. This
|
|
* script's whole ratchet apparatus applies to smells alone.
|
|
*
|
|
* WHY A BASELINE FILE *AND* A FRAGMENTS DIRECTORY (not just one)
|
|
* ──────────────────────────────────────────────────────────────
|
|
* This follows the exact idiom `tests/emitted-drift-acks/` and `.changeset/`
|
|
* already use in this repo, for the exact same reason: `smell-baseline.json`
|
|
* is a single shared file every PR that acknowledges a smell would otherwise
|
|
* have to rewrite, guaranteeing merge conflicts between any two such PRs in
|
|
* flight at once. A fragment per PR under `tests/qa/smell-acks/` — uniquely
|
|
* named (its own issue/PR number) — means two PRs can never conflict on this
|
|
* seam. A maintainer periodically folds spent fragments into the committed
|
|
* baseline via `--update` and deletes them (see that directory's README).
|
|
*
|
|
* USAGE
|
|
* ─────
|
|
* node scripts/qa-smell-ratchet.cjs # check (CI entry point)
|
|
* node scripts/qa-smell-ratchet.cjs --update # regenerate the baseline
|
|
* node scripts/qa-smell-ratchet.cjs --json <path> # also write the full qa-report
|
|
* node scripts/qa-smell-ratchet.cjs --keep # preserve scenario temp dirs
|
|
* # (real repro commands; see
|
|
* # `report.cjs`'s buildRepro)
|
|
*
|
|
* Exit code 0 only when: zero violations, zero NEW smells, zero STALE
|
|
* baseline/fragment entries. Exit code 1 otherwise.
|
|
*/
|
|
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
const { runAllScenarios } = require('../tests/qa/run-report.cjs');
|
|
const { buildReport } = require('../tests/qa/report.cjs');
|
|
const { fingerprint } = require('../tests/qa/smell-fingerprint.cjs');
|
|
const { ExitError, runMain } = require('./lib/cli-exit.cjs');
|
|
|
|
const REPO_ROOT = path.join(__dirname, '..');
|
|
const BASELINE_REL_PATH = 'tests/qa/smell-baseline.json';
|
|
const ACKS_DIR_REL_PATH = 'tests/qa/smell-acks';
|
|
const BASELINE_PATH = path.join(REPO_ROOT, ...BASELINE_REL_PATH.split('/'));
|
|
const ACKS_DIR = path.join(REPO_ROOT, ...ACKS_DIR_REL_PATH.split('/'));
|
|
|
|
/** Bump when `smell-baseline.json` / fragment shape changes incompatibly. */
|
|
const BASELINE_VERSION = 1;
|
|
|
|
/**
|
|
* Upper bound on fragment files read in one pass — mirrors the identical cap
|
|
* in `scripts/lint-emitted-drift-ack.cjs` (`MAX_ACK_FRAGMENTS`). Exceeding it
|
|
* throws rather than silently truncating the listing, which would silently
|
|
* drop acknowledgments from consideration — exactly the class of silent
|
|
* failure this whole seam exists to prevent.
|
|
*/
|
|
const MAX_ACK_FRAGMENTS = 500;
|
|
|
|
/**
|
|
* `--update` writes this into a newly-discovered entry's OPTIONAL `reason`
|
|
* field (alongside `issue: null`) as a note-to-self, never as a substitute for
|
|
* `issue` — see the header's "THE DESIGN INVARIANT" and #2966 FIX 3. A plain
|
|
* (non-`--update`) run rejects any entry whose `issue` is not a positive
|
|
* integer regardless of what `reason` says, and additionally rejects a
|
|
* `reason` still carrying this placeholder prefix (see `isPlaceholderReason`),
|
|
* so the baseline can never silently ship with a smell nobody has triaged.
|
|
*/
|
|
const PLACEHOLDER_REASON_PREFIX = 'TODO(qa-smell-ratchet):';
|
|
const PLACEHOLDER_REASON =
|
|
`${PLACEHOLDER_REASON_PREFIX} triage this smell — either file a defect and set "issue" to its number (REAL), ` +
|
|
'or fix the oracle so it stops firing (FALSE POSITIVE). A "reason" alone, with no "issue", is never accepted.';
|
|
|
|
const isPlainObject = (v) => v !== null && typeof v === 'object' && !Array.isArray(v);
|
|
const isPlaceholderReason = (reason) => typeof reason === 'string' && reason.startsWith(PLACEHOLDER_REASON_PREFIX);
|
|
|
|
/**
|
|
* Parse CLI argv into `{ update, jsonOut, keep }`.
|
|
*
|
|
* @param {string[]} argv
|
|
* @returns {{ update: boolean, jsonOut: string|null, keep: boolean }}
|
|
*/
|
|
function parseArgs(argv) {
|
|
let update = false;
|
|
let jsonOut = null;
|
|
let keep = false;
|
|
for (let i = 0; i < argv.length; i += 1) {
|
|
const arg = argv[i];
|
|
if (arg === '--update') {
|
|
update = true;
|
|
} else if (arg === '--json') {
|
|
const value = argv[i + 1];
|
|
if (typeof value !== 'string' || value === '') {
|
|
throw new ExitError(2, 'qa-smell-ratchet: --json requires a path argument');
|
|
}
|
|
jsonOut = path.resolve(value);
|
|
i += 1;
|
|
} else if (arg === '--keep') {
|
|
keep = true;
|
|
} else {
|
|
throw new ExitError(
|
|
2,
|
|
`qa-smell-ratchet: unrecognized argument "${arg}" (expected --update, --json <path>, and/or --keep)`,
|
|
);
|
|
}
|
|
}
|
|
return { update, jsonOut, keep };
|
|
}
|
|
|
|
/**
|
|
* Validate one baseline/fragment entry, pushing a message per problem onto
|
|
* `errors`. Does not mutate `entry`.
|
|
*
|
|
* Every entry MUST carry the three string fields (`key`, `id`, `scenario`)
|
|
* AND a positive-integer `issue` — the ONLY two terminal states for a smell
|
|
* are REAL (an assigned defect, cited by its issue number) or FALSE POSITIVE
|
|
* (the oracle gets fixed and the entry is never baselined at all); there is
|
|
* no third "accepted with a good explanation" state, so a free-text `reason`
|
|
* can NEVER substitute for `issue` (#2966 FIX 3). `reason` remains an OPTIONAL
|
|
* human note: when present it must be a non-empty, non-placeholder string,
|
|
* but its absence is never itself an error.
|
|
*
|
|
* @param {unknown} entry
|
|
* @param {string} where human-readable location for error messages
|
|
* (e.g. `"tests/qa/smell-baseline.json.smells[3]"` or a fragment's own
|
|
* relative path).
|
|
* @param {string[]} errors
|
|
* @returns {boolean} true when `entry` has all required fields, a valid
|
|
* `issue`, and (if present) a real (non-placeholder) `reason`.
|
|
*/
|
|
function validateEntryFields(entry, where, errors) {
|
|
if (!isPlainObject(entry)) {
|
|
errors.push(`${where} must be an object, got ${JSON.stringify(entry)}`);
|
|
return false;
|
|
}
|
|
let ok = true;
|
|
for (const field of ['key', 'id', 'scenario']) {
|
|
if (typeof entry[field] !== 'string' || entry[field] === '') {
|
|
errors.push(`${where}.${field} must be a non-empty string, got ${JSON.stringify(entry[field])}`);
|
|
ok = false;
|
|
}
|
|
}
|
|
if (!Number.isInteger(entry.issue) || entry.issue <= 0) {
|
|
errors.push(
|
|
`${where}.issue must be a positive integer, got ${JSON.stringify(entry.issue)} — every acknowledged smell ` +
|
|
'must be REAL (an assigned defect, cited by issue number) or a FALSE POSITIVE (the oracle is fixed, never ' +
|
|
'baselined); a free-text "reason" can never substitute for a tracked issue number',
|
|
);
|
|
ok = false;
|
|
}
|
|
if (entry.reason !== undefined) {
|
|
if (typeof entry.reason !== 'string' || entry.reason === '') {
|
|
errors.push(`${where}.reason, when present, must be a non-empty string, got ${JSON.stringify(entry.reason)}`);
|
|
ok = false;
|
|
} else if (isPlaceholderReason(entry.reason)) {
|
|
errors.push(
|
|
`${where}.reason is still the placeholder ("${entry.reason}") — either remove it or replace it with a `
|
|
+ 'real human note; either way, "issue" (not "reason") is what makes this entry valid',
|
|
);
|
|
ok = false;
|
|
}
|
|
}
|
|
return ok;
|
|
}
|
|
|
|
/**
|
|
* Read and validate `tests/qa/smell-baseline.json`.
|
|
*
|
|
* @param {{ allowMissing: boolean }} opts `allowMissing: true` is used only
|
|
* by `--update`'s bootstrap path — a not-yet-existing baseline is the
|
|
* expected first-run state there, never an error. In check mode a missing
|
|
* baseline is always an error (there is nothing to ratchet against).
|
|
* @returns {{ entries: Array<{key:string,id:string,scenario:string,issue:number,reason?:string}>, errors: string[], existed: boolean }}
|
|
*/
|
|
function readBaseline({ allowMissing }) {
|
|
const existed = fs.existsSync(BASELINE_PATH);
|
|
if (!existed) {
|
|
if (allowMissing) return { entries: [], errors: [], existed };
|
|
return {
|
|
entries: [],
|
|
errors: [`${BASELINE_REL_PATH} is missing — run \`node scripts/qa-smell-ratchet.cjs --update\` to generate it`],
|
|
existed,
|
|
};
|
|
}
|
|
|
|
const raw = fs.readFileSync(BASELINE_PATH, 'utf8');
|
|
if (raw.trim() === '') {
|
|
return { entries: [], errors: [`${BASELINE_REL_PATH} is present but empty`], existed };
|
|
}
|
|
let doc;
|
|
try {
|
|
doc = JSON.parse(raw);
|
|
} catch (err) {
|
|
return { entries: [], errors: [`${BASELINE_REL_PATH} is not valid JSON: ${err.message}`], existed };
|
|
}
|
|
const errors = [];
|
|
if (!isPlainObject(doc)) {
|
|
errors.push(`${BASELINE_REL_PATH} must be a JSON object, got ${Array.isArray(doc) ? 'array' : typeof doc}`);
|
|
return { entries: [], errors, existed };
|
|
}
|
|
if (doc.version !== BASELINE_VERSION) {
|
|
errors.push(`${BASELINE_REL_PATH}: unsupported version ${JSON.stringify(doc.version)} (expected ${BASELINE_VERSION})`);
|
|
}
|
|
if (!Array.isArray(doc.smells)) {
|
|
errors.push(`${BASELINE_REL_PATH}: "smells" must be an array, got ${JSON.stringify(doc.smells)}`);
|
|
return { entries: [], errors, existed };
|
|
}
|
|
|
|
const entries = [];
|
|
doc.smells.forEach((entry, i) => {
|
|
const where = `${BASELINE_REL_PATH}.smells[${i}]`;
|
|
if (validateEntryFields(entry, where, errors)) entries.push(entry);
|
|
});
|
|
return { entries, errors, existed };
|
|
}
|
|
|
|
/**
|
|
* Fragment filenames under `tests/qa/smell-acks/`, sorted. Absent directory
|
|
* == zero fragments. Throws (naming the dir, cap, and actual count) rather
|
|
* than silently truncating when the cap is exceeded.
|
|
*
|
|
* @returns {string[]}
|
|
*/
|
|
function listFragmentFiles() {
|
|
if (!fs.existsSync(ACKS_DIR)) return [];
|
|
const names = fs.readdirSync(ACKS_DIR).filter((n) => n.endsWith('.json')).sort();
|
|
if (names.length > MAX_ACK_FRAGMENTS) {
|
|
throw new ExitError(
|
|
1,
|
|
`qa-smell-ratchet: ${ACKS_DIR_REL_PATH} contains ${names.length} ack fragments, exceeding the cap of `
|
|
+ `${MAX_ACK_FRAGMENTS}. Refusing to read only some of them — a truncated read would silently drop `
|
|
+ 'acknowledgments. Prune spent fragments from this directory.',
|
|
);
|
|
}
|
|
return names;
|
|
}
|
|
|
|
/**
|
|
* Read and validate every fragment under `tests/qa/smell-acks/`. Each
|
|
* fragment is ONE acknowledgment: the same shape as a baseline entry — `key`,
|
|
* `id`, `scenario`, a positive-integer `issue`, and an OPTIONAL `reason` —
|
|
* validated identically via `validateEntryFields` (#2966 FIX 3: there is no
|
|
* separate "acknowledge via PR number" path; every acknowledgment cites the
|
|
* issue tracking the underlying defect).
|
|
*
|
|
* @returns {{ entries: Array<{key:string,id:string,scenario:string,issue:number,reason?:string,_source:string}>, errors: string[] }}
|
|
*/
|
|
function readAckFragments() {
|
|
const errors = [];
|
|
const entries = [];
|
|
for (const name of listFragmentFiles()) {
|
|
const label = `${ACKS_DIR_REL_PATH}/${name}`;
|
|
const raw = fs.readFileSync(path.join(ACKS_DIR, name), 'utf8');
|
|
if (raw.trim() === '') {
|
|
errors.push(`${label} is present but empty`);
|
|
continue;
|
|
}
|
|
let doc;
|
|
try {
|
|
doc = JSON.parse(raw);
|
|
} catch (err) {
|
|
errors.push(`${label} is not valid JSON: ${err.message}`);
|
|
continue;
|
|
}
|
|
if (!validateEntryFields(doc, label, errors)) continue;
|
|
entries.push({ ...doc, _source: label });
|
|
}
|
|
return { entries, errors };
|
|
}
|
|
|
|
/**
|
|
* Merge baseline entries and ack fragments into one `key -> entry` map (the
|
|
* full set of KNOWN smells this run is ratcheted against), plus the list of
|
|
* fragments that are now redundant because the baseline already carries
|
|
* their key (an advisory, not a failure — see this file's header on why
|
|
* cross-source duplication is not hard-blocked here).
|
|
*
|
|
* @param {Array<{key:string}>} baselineEntries
|
|
* @param {Array<{key:string,_source:string}>} fragmentEntries
|
|
* @returns {{ byKey: Map<string, object>, redundantFragments: Array<{key:string, source:string}> }}
|
|
*/
|
|
function mergeKnown(baselineEntries, fragmentEntries) {
|
|
const byKey = new Map();
|
|
for (const e of baselineEntries) byKey.set(e.key, { ...e, source: BASELINE_REL_PATH });
|
|
|
|
const redundantFragments = [];
|
|
for (const e of fragmentEntries) {
|
|
if (byKey.has(e.key)) {
|
|
redundantFragments.push({ key: e.key, source: e._source });
|
|
continue;
|
|
}
|
|
byKey.set(e.key, { ...e, source: e._source });
|
|
}
|
|
return { byKey, redundantFragments };
|
|
}
|
|
|
|
/**
|
|
* Walk `reportObject.scenarios[].steps[]` and split every finding into
|
|
* `smells` (fingerprinted) and `violations` (never acknowledgeable — see
|
|
* this file's header). Both carry the step's `repro` command for later use
|
|
* in failure messages / the GitHub step summary.
|
|
*
|
|
* @param {ReturnType<import('../tests/qa/report.cjs').buildReport>} reportObject
|
|
* @returns {{
|
|
* smells: Array<{key:string,id:string,scenario:string,argv:string[],detail:string,at:string,repro:string}>,
|
|
* violations: Array<{id:string,scenario:string,argv:string[],detail:string,at:string,repro:string}>,
|
|
* }}
|
|
*/
|
|
function collectFindings(reportObject) {
|
|
const smells = [];
|
|
const violations = [];
|
|
for (const scenario of reportObject.scenarios) {
|
|
for (const step of scenario.steps) {
|
|
for (const v of step.violations || []) {
|
|
violations.push({
|
|
id: v.id, scenario: scenario.name, argv: step.argv, detail: v.detail, at: step.at, repro: step.repro,
|
|
});
|
|
}
|
|
for (const smell of step.smells || []) {
|
|
const key = fingerprint(scenario.name, { id: smell.id, subject: smell.subject, argv: step.argv });
|
|
smells.push({
|
|
key,
|
|
id: smell.id,
|
|
scenario: scenario.name,
|
|
argv: step.argv,
|
|
detail: smell.detail,
|
|
at: step.at,
|
|
repro: step.repro,
|
|
});
|
|
}
|
|
}
|
|
}
|
|
return { smells, violations };
|
|
}
|
|
|
|
/** Lowercase, hyphenate, and strip anything that isn't `[a-z0-9-]`, for a fragment-filename skeleton. */
|
|
function slugify(value) {
|
|
return value
|
|
.toLowerCase()
|
|
.replace(/[^a-z0-9]+/g, '-')
|
|
.replace(/^-+|-+$/g, '')
|
|
.slice(0, 60);
|
|
}
|
|
|
|
/**
|
|
* Render the paste-ready fragment skeleton for one NEW smell finding.
|
|
*
|
|
* @param {{key:string,id:string,scenario:string}} finding
|
|
* @returns {string}
|
|
*/
|
|
function fragmentSkeleton(finding) {
|
|
const doc = {
|
|
version: 1,
|
|
key: finding.key,
|
|
id: finding.id,
|
|
scenario: finding.scenario,
|
|
issue: '<YOUR ISSUE NUMBER>',
|
|
};
|
|
const suggestedName = `${ACKS_DIR_REL_PATH}/<issue>-${slugify(finding.id)}-${slugify(finding.scenario)}.json`;
|
|
return `${suggestedName}:\n${JSON.stringify(doc, null, 2)}`;
|
|
}
|
|
|
|
/**
|
|
* Build the markdown block appended to `GITHUB_STEP_SUMMARY`, when set —
|
|
* kept intentionally compact (a PR reviewer's first read, not a log dump).
|
|
*
|
|
* @param {{
|
|
* smells: ReturnType<typeof collectFindings>['smells'],
|
|
* violations: ReturnType<typeof collectFindings>['violations'],
|
|
* newKeys: string[],
|
|
* staleEntries: Array<{key:string,id:string,scenario:string,source:string}>,
|
|
* smellSummary: Array<{id:string,count:number,examples:string[]}>,
|
|
* }} data
|
|
* @returns {string}
|
|
*/
|
|
function buildStepSummaryMarkdown({ smells, violations, newKeys, staleEntries, smellSummary }) {
|
|
const lines = [];
|
|
lines.push('## QA smell ratchet');
|
|
lines.push('');
|
|
lines.push(
|
|
`**${smells.length} smells** (${newKeys.length} new, ${staleEntries.length} stale) · `
|
|
+ `**${violations.length} violations**`,
|
|
);
|
|
lines.push('');
|
|
|
|
if (newKeys.length) {
|
|
lines.push('### 🚨 NEW (unacknowledged) smells');
|
|
lines.push('');
|
|
for (const key of newKeys) {
|
|
const f = smells.find((s) => s.key === key);
|
|
lines.push(`- \`${f.id}\` in **${f.scenario}** — ${f.detail}`);
|
|
}
|
|
lines.push('');
|
|
}
|
|
|
|
if (staleEntries.length) {
|
|
lines.push('### Stale baseline/fragment entries (no longer produced)');
|
|
lines.push('');
|
|
for (const e of staleEntries) {
|
|
lines.push(`- \`${e.id}\` in **${e.scenario}** (${e.source})`);
|
|
}
|
|
lines.push('');
|
|
}
|
|
|
|
if (smellSummary.length) {
|
|
lines.push('### Smells by oracle');
|
|
lines.push('');
|
|
lines.push('| oracle id | count |');
|
|
lines.push('|---|---|');
|
|
for (const entry of smellSummary) {
|
|
lines.push(`| \`${entry.id}\` | ${entry.count} |`);
|
|
}
|
|
lines.push('');
|
|
}
|
|
|
|
const firstFailingRepro = (violations[0] && violations[0].repro)
|
|
|| (newKeys.length && smells.find((s) => s.key === newKeys[0]).repro);
|
|
if (firstFailingRepro) {
|
|
lines.push('### Repro (first failing step)');
|
|
lines.push('');
|
|
lines.push('```sh');
|
|
lines.push(firstFailingRepro);
|
|
lines.push('```');
|
|
lines.push('');
|
|
}
|
|
|
|
return lines.join('\n');
|
|
}
|
|
|
|
function main() {
|
|
const { update, jsonOut, keep } = parseArgs(process.argv.slice(2));
|
|
|
|
const scenarioReports = runAllScenarios({ keep });
|
|
const meta = {
|
|
nodeVersion: process.version,
|
|
platform: process.platform,
|
|
// Only ever used for report METADATA (and, when --json is passed, the
|
|
// written artifact's meta.generatedAt) — never fed into a fingerprint or
|
|
// into smell-baseline.json, which is what keeps this script's fingerprint
|
|
// and baseline output deterministic despite this one real clock read.
|
|
generatedAt: new Date().toISOString(),
|
|
};
|
|
const reportObject = buildReport(scenarioReports, meta);
|
|
|
|
if (jsonOut) {
|
|
fs.mkdirSync(path.dirname(jsonOut), { recursive: true });
|
|
fs.writeFileSync(jsonOut, `${JSON.stringify(reportObject, null, 2)}\n`, 'utf8');
|
|
}
|
|
|
|
const { smells, violations } = collectFindings(reportObject);
|
|
const runKeys = new Set(smells.map((s) => s.key));
|
|
|
|
const baseline = readBaseline({ allowMissing: update });
|
|
const fragments = readAckFragments();
|
|
const sourceErrors = [...baseline.errors, ...fragments.errors];
|
|
|
|
if (update) {
|
|
if (sourceErrors.length) {
|
|
for (const e of sourceErrors) console.error(` - ${e}`);
|
|
throw new ExitError(
|
|
1,
|
|
`qa-smell-ratchet --update: ${sourceErrors.length} problem(s) in existing baseline/fragment source(s) `
|
|
+ '(printed above) — fix or delete the offending source(s) by hand before regenerating.',
|
|
);
|
|
}
|
|
|
|
const { byKey: knownBeforeUpdate } = mergeKnown(baseline.entries, fragments.entries);
|
|
const oldBaselineKeys = new Set(baseline.entries.map((e) => e.key));
|
|
|
|
// `--update` NEVER invents an issue number (#2966 FIX 3). A key already
|
|
// carrying a real `issue` (from the committed baseline or a fragment) keeps
|
|
// it, along with its `reason` if any. A genuinely NEW smell — no prior
|
|
// acknowledgment exists — gets `issue: null` and a TODO `reason`; the very
|
|
// next plain (non-`--update`) run REJECTS that entry, forcing a human to
|
|
// triage it as REAL (cite the issue) or FALSE POSITIVE (fix the oracle).
|
|
const newBaselineEntries = [...runKeys].sort().map((key) => {
|
|
const representative = smells.find((s) => s.key === key);
|
|
const carried = knownBeforeUpdate.get(key);
|
|
const hasKnownIssue = !!carried && Number.isInteger(carried.issue) && carried.issue > 0;
|
|
const entry = {
|
|
key,
|
|
id: representative.id,
|
|
scenario: representative.scenario,
|
|
issue: hasKnownIssue ? carried.issue : null,
|
|
};
|
|
if (hasKnownIssue && typeof carried.reason === 'string' && !isPlaceholderReason(carried.reason)) {
|
|
entry.reason = carried.reason;
|
|
} else if (!hasKnownIssue) {
|
|
entry.reason = PLACEHOLDER_REASON;
|
|
}
|
|
return entry;
|
|
});
|
|
const newBaselineKeys = new Set(newBaselineEntries.map((e) => e.key));
|
|
|
|
const added = [...newBaselineKeys].filter((k) => !oldBaselineKeys.has(k)).sort();
|
|
const removed = [...oldBaselineKeys].filter((k) => !newBaselineKeys.has(k)).sort();
|
|
|
|
fs.mkdirSync(path.dirname(BASELINE_PATH), { recursive: true });
|
|
fs.writeFileSync(
|
|
BASELINE_PATH,
|
|
`${JSON.stringify({ version: BASELINE_VERSION, smells: newBaselineEntries }, null, 2)}\n`,
|
|
'utf8',
|
|
);
|
|
|
|
console.log(
|
|
`qa-smell-ratchet --update: ${oldBaselineKeys.size} -> ${newBaselineKeys.size} baseline entries`
|
|
+ (added.length ? ` | added: ${added.length}` : '')
|
|
+ (removed.length ? ` | removed: ${removed.length}` : ''),
|
|
);
|
|
for (const key of added) {
|
|
const e = newBaselineEntries.find((x) => x.key === key);
|
|
const placeholderNote = e.issue === null ? ' [issue: null — TODO, needs triage before the next check run]' : '';
|
|
console.log(` + ${key}${placeholderNote}`);
|
|
}
|
|
for (const key of removed) console.log(` - ${key}`);
|
|
|
|
const redundant = fragments.entries.filter((e) => newBaselineKeys.has(e.key));
|
|
if (redundant.length) {
|
|
console.log(
|
|
`\n${redundant.length} fragment(s) are now redundant — their key is already in the regenerated baseline. `
|
|
+ 'Delete them (CONTRIBUTING.md fragment idiom: fold, then delete):',
|
|
);
|
|
for (const e of redundant) console.log(` - ${e._source}`);
|
|
}
|
|
|
|
if (process.env.GITHUB_STEP_SUMMARY) {
|
|
const md = buildStepSummaryMarkdown({
|
|
smells,
|
|
violations,
|
|
newKeys: added,
|
|
staleEntries: removed.map((key) => ({ key, id: '(pruned)', scenario: '(pruned)', source: BASELINE_REL_PATH })),
|
|
smellSummary: reportObject.smellSummary,
|
|
});
|
|
fs.appendFileSync(process.env.GITHUB_STEP_SUMMARY, `${md}\n`);
|
|
}
|
|
|
|
console.log(
|
|
`\nqa-smell-ratchet: ${smells.length} smells (${added.length} new, ${removed.length} stale), `
|
|
+ `${violations.length} violations`,
|
|
);
|
|
|
|
if (violations.length) {
|
|
printViolations(violations);
|
|
throw new ExitError(1, 'qa-smell-ratchet --update: baseline regenerated, but VIOLATIONS remain (never acknowledgeable — see above)');
|
|
}
|
|
return;
|
|
}
|
|
|
|
// ── check mode ──────────────────────────────────────────────────────────
|
|
const { byKey: known, redundantFragments } = mergeKnown(baseline.entries, fragments.entries);
|
|
const newKeys = [...runKeys].filter((k) => !known.has(k)).sort();
|
|
const staleKeys = [...known.keys()].filter((k) => !runKeys.has(k)).sort();
|
|
const staleEntries = staleKeys.map((key) => known.get(key));
|
|
|
|
if (sourceErrors.length) {
|
|
console.error(`qa-smell-ratchet: ${sourceErrors.length} problem(s) in baseline/fragment source(s):\n`);
|
|
for (const e of sourceErrors) console.error(` - ${e}`);
|
|
}
|
|
|
|
if (violations.length) {
|
|
printViolations(violations);
|
|
}
|
|
|
|
if (newKeys.length) {
|
|
console.error(`\nqa-smell-ratchet: ${newKeys.length} NEW (unacknowledged) smell(s):\n`);
|
|
for (const key of newKeys) {
|
|
const f = smells.find((s) => s.key === key);
|
|
console.error(`NEW smell: ${f.key}`);
|
|
console.error(` oracle: ${f.id}`);
|
|
console.error(` scenario: ${f.scenario}`);
|
|
console.error(` detail: ${f.detail}`);
|
|
console.error(' remedy: exactly two options — no third "accepted with an explanation" state:');
|
|
console.error(' 1. fix the detector if this is a FALSE POSITIVE (the oracle is wrong; make it stop firing);');
|
|
console.error(' 2. file a defect and add an entry citing its issue number (REAL) — a fragment:\n');
|
|
console.error(`${fragmentSkeleton(f).split('\n').map((l) => ` ${l}`).join('\n')}\n`);
|
|
}
|
|
}
|
|
|
|
if (staleKeys.length) {
|
|
console.error(`\nqa-smell-ratchet: ${staleKeys.length} STALE baseline/fragment entr${staleKeys.length === 1 ? 'y' : 'ies'} (no longer produced by the run):\n`);
|
|
for (const e of staleEntries) {
|
|
console.error(`STALE entry: ${e.key}`);
|
|
console.error(` source: ${e.source}`);
|
|
console.error(` oracle: ${e.id}`);
|
|
console.error(` scenario: ${e.scenario}`);
|
|
console.error(` issue: ${e.issue}`);
|
|
if (e.reason !== undefined) console.error(` reason: ${e.reason}`);
|
|
}
|
|
console.error('\n remedy: node scripts/qa-smell-ratchet.cjs --update');
|
|
}
|
|
|
|
if (redundantFragments.length) {
|
|
console.log(
|
|
`\n${redundantFragments.length} fragment(s) are already covered by the baseline and can be deleted:`,
|
|
);
|
|
for (const e of redundantFragments) console.log(` - ${e.source} (key ${e.key})`);
|
|
}
|
|
|
|
if (process.env.GITHUB_STEP_SUMMARY) {
|
|
const md = buildStepSummaryMarkdown({
|
|
smells,
|
|
violations,
|
|
newKeys,
|
|
staleEntries,
|
|
smellSummary: reportObject.smellSummary,
|
|
});
|
|
fs.appendFileSync(process.env.GITHUB_STEP_SUMMARY, `${md}\n`);
|
|
}
|
|
|
|
console.log(
|
|
`\nqa-smell-ratchet: ${smells.length} smells (${newKeys.length} new, ${staleKeys.length} stale), `
|
|
+ `${violations.length} violations`,
|
|
);
|
|
|
|
if (sourceErrors.length || violations.length || newKeys.length || staleKeys.length) {
|
|
throw new ExitError(1);
|
|
}
|
|
}
|
|
|
|
/**
|
|
* @param {ReturnType<typeof collectFindings>['violations']} violations
|
|
*/
|
|
function printViolations(violations) {
|
|
console.error(`qa-smell-ratchet: ${violations.length} VIOLATION(s) — never acknowledgeable, always fail:\n`);
|
|
for (const v of violations) {
|
|
console.error(`VIOLATION: ${v.id}`);
|
|
console.error(` scenario: ${v.scenario}`);
|
|
console.error(` argv: ${v.argv.join(' ')}`);
|
|
console.error(` detail: ${v.detail}`);
|
|
console.error(` repro: ${v.repro}`);
|
|
}
|
|
}
|
|
|
|
runMain(main);
|
|
|
|
module.exports = {
|
|
parseArgs,
|
|
readBaseline,
|
|
readAckFragments,
|
|
mergeKnown,
|
|
collectFindings,
|
|
fragmentSkeleton,
|
|
slugify,
|
|
isPlaceholderReason,
|
|
PLACEHOLDER_REASON_PREFIX,
|
|
BASELINE_REL_PATH,
|
|
ACKS_DIR_REL_PATH,
|
|
MAX_ACK_FRAGMENTS,
|
|
};
|