* fix: recurse test discovery so subdir test suites actually run
scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.
Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: retire 5 verified-worthless tests
Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
and bug-782-cline-skills-emission)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: add ADR-218 release version-validation coverage
ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: redesign weak tests into behavioral, deterministic assertions
Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
collision) and unconditional plugin.json schema validation (issue-766)
Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add no-tautological-assert lint rule, error in test suite
New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).
Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gate new allow-test-rule exemptions to require an issue ref
ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: add ADR test-audit evidence report (#1192)
Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture
feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at b10e5681 — confirmed: base blob has 1 NUL, this fix has 0), which
made git treat the file as binary and would break grep/editors. Switch to the
\x00 escape; the runtime string value (a real NUL in the parser input) is
unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: address adversarial-review findings
Codex adversarial pass over the branch:
- capability-registry drift test no longer mutates the committed generated
capability-registry.cjs in place (concurrency hazard) — uses in-memory
checkPipeline comparison instead.
- allow-test-rule ratchet now detects exemptions in ALL comment forms (block
/* */ too, matching no-source-grep) so a block comment can't bypass it;
one newly-surfaced pre-existing offender grandfathered (323->324).
- install.test Kilo case asserts on what install(false,'kilo') actually writes
rather than manually calling configureKiloPermissions (masked the call site).
- issue-766 drops the undeclared transitive ajv dep for explicit structural
assertions from the schema fixture.
- adr-218 test notes the hotfix leading-zero gap is tracked in #1186.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: address code-review findings (subdir discovery, rule + test gaps)
xhigh code review surfaced 15 confirmed issues, all fixed:
- run-tests.cjs --files now resolves subdir tests by bare basename + handles
Windows backslash paths (ambiguous basenames error clearly).
- affected-tests-lib.cjs listTestFiles made recursive — the targeted CI lane was
silently dropping changed subdir tests (same false-green class the audit fixed).
- no-tautological-assert now catches 'true || cond' and empty []/{} equality.
- verify-test-quality: restore provenance-classification coverage, tighten the
writeFile circular-detection check, guard the module-level file read.
- sh-hook-paths: cover the global-install .sh delegation branch (#2045 guard).
- active-workstream null-guard runs deterministically (no longer skipped on TTY).
- adr-218 structural guards tightened (major/minor leading-zero; needs: membership).
- repo-layout AGENTS.md guard no longer false-alarms on equivalent refactors.
- cross-ai ordering guard fails red when the step is missing.
- issue-766 parses required fields from the schema fixture (auto-enforced).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: stub USERPROFILE alongside HOME in feat-488 (Windows parity)
The feat-488 redesign stubbed process.env.HOME but not USERPROFILE; os.homedir()
resolves from USERPROFILE on Windows, so the home stub was not hermetic there —
caught by windows-test-parity-guard (stubsHomeNoUserProfile). Save/set/restore
USERPROFILE symmetrically with HOME (delete-if-originally-undefined).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: reconcile allow-test-rule allowlist after rebase onto next
Rebasing onto current next pulled in merged PR #1170, which added
inventory-headings-countfree.test.cjs (a baseline allow-test-rule exemption) and
deleted inventory-counts.test.cjs. Grandfather the former and prune the latter so
the ratchet matches the merged tree. No new debt from this PR.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
179 lines
6.9 KiB
JavaScript
179 lines
6.9 KiB
JavaScript
/**
|
|
* Deterministic property-style parser tests (#3594).
|
|
*
|
|
* Follows TEST-EXAMPLES.md §"Deterministic Property-Style Parser Test":
|
|
* a bounded, seeded loop generates many malformed inputs and asserts a
|
|
* single invariant against each. On failure the seed and case index
|
|
* are printed so the failing input can be reproduced exactly.
|
|
*
|
|
* The generator is a small mulberry32 PRNG so this file has zero
|
|
* external dependencies and is fully reproducible across Node versions.
|
|
* Each test pins its own seed and case count; bumping either is a
|
|
* deliberate test change, not a flake source.
|
|
*
|
|
* Invariant tested (frontmatter): for any random text the parser must
|
|
* either return a plain object or throw — never return null/undefined,
|
|
* never hang, never propagate "Cannot read properties of …" V8 prose.
|
|
*/
|
|
|
|
'use strict';
|
|
|
|
const { test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
|
|
const { extractFrontmatter } = require('../gsd-core/bin/lib/frontmatter.cjs');
|
|
|
|
/**
|
|
* mulberry32 — small fast deterministic PRNG. Seed in, [0,1) out.
|
|
* Same input always produces the same sequence across Node versions.
|
|
*/
|
|
function mulberry32(seed) {
|
|
let a = seed >>> 0;
|
|
return function next() {
|
|
a = (a + 0x6D2B79F5) | 0;
|
|
let t = a;
|
|
t = Math.imul(t ^ (t >>> 15), t | 1);
|
|
t ^= t + Math.imul(t ^ (t >>> 7), t | 61);
|
|
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
|
|
};
|
|
}
|
|
|
|
/**
|
|
* Build a single malformed-ish frontmatter input. Components are mixed
|
|
* deterministically by the supplied PRNG.
|
|
*/
|
|
function makeInput(rng) {
|
|
const fragments = [
|
|
'---\n',
|
|
'title: Generated\n',
|
|
'phase: 99\n',
|
|
'plans:\n - a\n - b\n',
|
|
'extra: \xff\xfe\xfd\n', // invalid UTF-8 bytes
|
|
'unicode: 日本語\n',
|
|
'crlf: ends\r\nin\rcr\n',
|
|
' indented_key: value\n',
|
|
'duplicate: first\nduplicate: second\n',
|
|
'sparse:\n\n\n',
|
|
'malformed_array: [a, "b", c\n', // unclosed inline array
|
|
'null_byte: before\x00after\n',
|
|
];
|
|
// Pick a random subset of fragments in random order. Always include
|
|
// the opening `---`. Closing `---` is included by 50% of cases so we
|
|
// exercise both well-formed and unclosed shapes.
|
|
const head = fragments[0];
|
|
const rest = shuffle(fragments.slice(1), rng).slice(0, 1 + Math.floor(rng() * 6));
|
|
const closing = rng() < 0.5 ? '---\n' : '';
|
|
return head + rest.join('') + closing + '\nBody.\n';
|
|
}
|
|
|
|
/**
|
|
* Fisher-Yates shuffle driven by the supplied PRNG. Returns a new
|
|
* array; does not mutate the input. Replaces the previous
|
|
* `arr.sort(() => rng() - 0.5)` which was non-transitive — the
|
|
* resulting order depended on V8's sort implementation, not only on
|
|
* the seed, so failing cases were unreproducible across Node versions.
|
|
* Fisher-Yates is O(n), transitive (no comparator), and depends only
|
|
* on the RNG output. Codex review on PR #3633 / #3594.
|
|
*/
|
|
function shuffle(arr, rng) {
|
|
const out = arr.slice();
|
|
for (let i = out.length - 1; i > 0; i--) {
|
|
const j = Math.floor(rng() * (i + 1));
|
|
const tmp = out[i];
|
|
out[i] = out[j];
|
|
out[j] = tmp;
|
|
}
|
|
return out;
|
|
}
|
|
|
|
test('extractFrontmatter is total over 500 deterministic random inputs (seed=1234)', () => {
|
|
const seed = 1234;
|
|
const rng = mulberry32(seed);
|
|
const count = 500;
|
|
for (let i = 0; i < count; i++) {
|
|
const input = makeInput(rng);
|
|
let result;
|
|
try {
|
|
result = extractFrontmatter(input);
|
|
} catch (err) {
|
|
// If the parser throws, the failure must be a controlled one —
|
|
// not a V8 "Cannot read properties of undefined" that signals a
|
|
// null-deref bug. Print the seed and case index so the input
|
|
// can be reproduced exactly.
|
|
const msg = String((err && err.message) || err);
|
|
assert.doesNotMatch(
|
|
msg,
|
|
/Cannot read propert/i,
|
|
`seed=${seed} case=${i}: parser must not propagate null-deref TypeError; input=${JSON.stringify(input)}`,
|
|
);
|
|
continue;
|
|
}
|
|
// No throw: result MUST be a plain object (not null, not array, not
|
|
// primitive). Print enough on failure to reproduce.
|
|
assert.equal(typeof result, 'object', `seed=${seed} case=${i}: result must be object, got ${typeof result}`);
|
|
assert.notEqual(result, null, `seed=${seed} case=${i}: result must not be null`);
|
|
assert.equal(Array.isArray(result), false, `seed=${seed} case=${i}: result must not be an array`);
|
|
}
|
|
});
|
|
|
|
test('extractFrontmatter scales sub-quadratically (complexity ratio guard)', () => {
|
|
// Rationale: an absolute wall-clock bound (e.g. < 2000 ms) is flaky —
|
|
// it fails on slow CI machines and passes on a fast local box even when
|
|
// a quadratic regression has been introduced. A *ratio* test is
|
|
// self-calibrating: we measure how much longer the parser takes on a
|
|
// 10x-larger input (by line count). For an O(n) parser the ratio should
|
|
// be near 10; for an O(n^2) parser it would be near 100. We tolerate
|
|
// up to 60x to give ample room for JIT, GC, constant-factor differences,
|
|
// and measurement noise — yet a true quadratic regression (ratio ~100)
|
|
// will still be caught.
|
|
//
|
|
// Input shape: pure key:value lines so the line count directly controls
|
|
// the amount of work the parser does per call. No randomness needed here
|
|
// — the property being tested is complexity, not totality.
|
|
|
|
/** Build a frontmatter string with exactly `lineCount` key:value lines. */
|
|
function buildScaleInput(lineCount) {
|
|
let s = '---\n';
|
|
for (let i = 0; i < lineCount; i++) {
|
|
s += `key${i}: value${i}\n`;
|
|
}
|
|
return s + '---\nBody.\n';
|
|
}
|
|
|
|
const SMALL_LINES = 20;
|
|
const LARGE_LINES = 200; // 10x more lines than SMALL_LINES
|
|
const SIZE_RATIO = LARGE_LINES / SMALL_LINES; // 10
|
|
const REPS = 3000; // enough iterations for hrtime to produce stable ns totals
|
|
const MAX_RATIO = SIZE_RATIO * 6; // 60 — well above O(n) (10) but well below O(n^2) (100)
|
|
|
|
const smallInput = buildScaleInput(SMALL_LINES);
|
|
const largeInput = buildScaleInput(LARGE_LINES);
|
|
|
|
// Warmup: let V8 JIT-compile the hot path before we measure.
|
|
for (let i = 0; i < 300; i++) {
|
|
extractFrontmatter(smallInput);
|
|
extractFrontmatter(largeInput);
|
|
}
|
|
|
|
const t1 = process.hrtime.bigint();
|
|
for (let i = 0; i < REPS; i++) extractFrontmatter(smallInput);
|
|
const dSmall = Number(process.hrtime.bigint() - t1);
|
|
|
|
const t2 = process.hrtime.bigint();
|
|
for (let i = 0; i < REPS; i++) extractFrontmatter(largeInput);
|
|
const dLarge = Number(process.hrtime.bigint() - t2);
|
|
|
|
// Guard against a degenerate measurement (< 1 µs total) that would
|
|
// make the ratio meaningless. If the machine is this fast, the parser
|
|
// is trivially fine and we skip the ratio check.
|
|
if (dSmall < 1000 /* 1 µs */) return;
|
|
|
|
const ratio = dLarge / dSmall;
|
|
assert.ok(
|
|
ratio < MAX_RATIO,
|
|
`complexity ratio ${ratio.toFixed(1)} exceeds ${MAX_RATIO} ` +
|
|
`(${LARGE_LINES}-line input took ${(ratio).toFixed(1)}x longer than ${SMALL_LINES}-line input; ` +
|
|
`expected ≤ ${MAX_RATIO}x for sub-quadratic behaviour — possible O(n²) regression)`,
|
|
);
|
|
});
|