* fix: recurse test discovery so subdir test suites actually run
scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.
Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: retire 5 verified-worthless tests
Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
and bug-782-cline-skills-emission)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: add ADR-218 release version-validation coverage
ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: redesign weak tests into behavioral, deterministic assertions
Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
collision) and unconditional plugin.json schema validation (issue-766)
Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add no-tautological-assert lint rule, error in test suite
New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).
Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gate new allow-test-rule exemptions to require an issue ref
ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: add ADR test-audit evidence report (#1192)
Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture
feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at b10e5681 — confirmed: base blob has 1 NUL, this fix has 0), which
made git treat the file as binary and would break grep/editors. Switch to the
\x00 escape; the runtime string value (a real NUL in the parser input) is
unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: address adversarial-review findings
Codex adversarial pass over the branch:
- capability-registry drift test no longer mutates the committed generated
capability-registry.cjs in place (concurrency hazard) — uses in-memory
checkPipeline comparison instead.
- allow-test-rule ratchet now detects exemptions in ALL comment forms (block
/* */ too, matching no-source-grep) so a block comment can't bypass it;
one newly-surfaced pre-existing offender grandfathered (323->324).
- install.test Kilo case asserts on what install(false,'kilo') actually writes
rather than manually calling configureKiloPermissions (masked the call site).
- issue-766 drops the undeclared transitive ajv dep for explicit structural
assertions from the schema fixture.
- adr-218 test notes the hotfix leading-zero gap is tracked in #1186.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: address code-review findings (subdir discovery, rule + test gaps)
xhigh code review surfaced 15 confirmed issues, all fixed:
- run-tests.cjs --files now resolves subdir tests by bare basename + handles
Windows backslash paths (ambiguous basenames error clearly).
- affected-tests-lib.cjs listTestFiles made recursive — the targeted CI lane was
silently dropping changed subdir tests (same false-green class the audit fixed).
- no-tautological-assert now catches 'true || cond' and empty []/{} equality.
- verify-test-quality: restore provenance-classification coverage, tighten the
writeFile circular-detection check, guard the module-level file read.
- sh-hook-paths: cover the global-install .sh delegation branch (#2045 guard).
- active-workstream null-guard runs deterministically (no longer skipped on TTY).
- adr-218 structural guards tightened (major/minor leading-zero; needs: membership).
- repo-layout AGENTS.md guard no longer false-alarms on equivalent refactors.
- cross-ai ordering guard fails red when the step is missing.
- issue-766 parses required fields from the schema fixture (auto-enforced).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: stub USERPROFILE alongside HOME in feat-488 (Windows parity)
The feat-488 redesign stubbed process.env.HOME but not USERPROFILE; os.homedir()
resolves from USERPROFILE on Windows, so the home stub was not hermetic there —
caught by windows-test-parity-guard (stubsHomeNoUserProfile). Save/set/restore
USERPROFILE symmetrically with HOME (delete-if-originally-undefined).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: reconcile allow-test-rule allowlist after rebase onto next
Rebasing onto current next pulled in merged PR #1170, which added
inventory-headings-countfree.test.cjs (a baseline allow-test-rule exemption) and
deleted inventory-counts.test.cjs. Grandfather the former and prune the latter so
the ratchet matches the merged tree. No new debt from this PR.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
253 lines
10 KiB
JavaScript
253 lines
10 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* Property-based tests for context-utilization.cjs
|
|
*
|
|
* Module: gsd-core/bin/lib/context-utilization.cjs
|
|
* Exported: classifyContextUtilization(tokensUsed, contextWindow) -> { percent, state }
|
|
*
|
|
* Thresholds (from module source):
|
|
* ratio < 0.60 → healthy
|
|
* 0.60 <= ratio < 0.70 → warning
|
|
* ratio >= 0.70 → critical
|
|
*
|
|
* Properties tested:
|
|
* (a) Boundary: across the 60% and 70% thresholds the classify flips correctly
|
|
* (b) Robustness: hostile inputs (null/undefined/NaN/Infinity/negative/wrong-type)
|
|
* always throw TypeError (documented contract) and never throw non-TypeError
|
|
* (c) Return shape: all valid inputs return an object with typed { percent, state }
|
|
*/
|
|
|
|
const { describe, test } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
const fc = require('./helpers/fast-check-setup.cjs');
|
|
|
|
const { classifyContextUtilization, STATES } = require('../gsd-core/bin/lib/context-utilization.cjs');
|
|
|
|
// ─── Boundary constants ───────────────────────────────────────────────────────
|
|
const WARNING_THRESHOLD = 0.60; // ratio < this → healthy
|
|
const CRITICAL_THRESHOLD = 0.70; // ratio < this → warning, else critical
|
|
const WINDOW = 10000; // A fixed contextWindow that gives us clean ratio math
|
|
const W = 50; // boundary exploration width: ±50 tokens around the threshold
|
|
|
|
describe('context-utilization property tests', () => {
|
|
// ─── (a) Boundary property: classify flips at exactly 60% and 70% ────────────
|
|
test('property: classify is healthy below 60%, warning in [60%,70%), critical at ≥70%', () => {
|
|
// Test across a range of context windows.
|
|
// The boundary is at the exact RATIO — not at Math.floor(window * ratio).
|
|
// e.g. with window=1001: floor(1001 * 0.60) = 600, but 600/1001 = 0.5994 < 0.60 → healthy.
|
|
// So we compute the exact first token count that meets or exceeds the threshold.
|
|
fc.assert(
|
|
fc.property(
|
|
// contextWindow: 1000..200000 to keep ratios well-defined
|
|
fc.integer({ min: 1000, max: 200_000 }),
|
|
(contextWindow) => {
|
|
// First integer where tokensUsed/contextWindow >= WARNING_THRESHOLD
|
|
const firstWarning = Math.ceil(contextWindow * WARNING_THRESHOLD);
|
|
// First integer where tokensUsed/contextWindow >= CRITICAL_THRESHOLD
|
|
const firstCritical = Math.ceil(contextWindow * CRITICAL_THRESHOLD);
|
|
|
|
// Just below warning threshold → healthy
|
|
if (firstWarning > 0) {
|
|
const below = firstWarning - 1;
|
|
const r = classifyContextUtilization(below, contextWindow);
|
|
assert.equal(
|
|
r.state,
|
|
STATES.HEALTHY,
|
|
`tokensUsed=${below} contextWindow=${contextWindow} ratio=${(below / contextWindow).toFixed(6)} expected healthy got ${r.state}`
|
|
);
|
|
}
|
|
|
|
// At exact warning boundary → warning (unless critical collapses to same point)
|
|
if (firstWarning < firstCritical && firstWarning <= contextWindow) {
|
|
const r2 = classifyContextUtilization(firstWarning, contextWindow);
|
|
assert.equal(
|
|
r2.state,
|
|
STATES.WARNING,
|
|
`tokensUsed=${firstWarning} ratio=${(firstWarning / contextWindow).toFixed(6)} expected warning got ${r2.state}`
|
|
);
|
|
}
|
|
|
|
// At or above critical threshold → critical
|
|
if (firstCritical <= contextWindow) {
|
|
const r3 = classifyContextUtilization(firstCritical, contextWindow);
|
|
assert.equal(
|
|
r3.state,
|
|
STATES.CRITICAL,
|
|
`tokensUsed=${firstCritical} ratio=${(firstCritical / contextWindow).toFixed(6)} expected critical got ${r3.state}`
|
|
);
|
|
}
|
|
}
|
|
)
|
|
);
|
|
});
|
|
|
|
test('property: near 60% boundary the state is always healthy (never warning/critical)', () => {
|
|
// Tokens strictly below ceil(WINDOW * 0.60) must classify as healthy.
|
|
// The boundary is the FIRST integer where ratio >= 0.60 (Math.ceil).
|
|
// We sample from [firstWarning - W, firstWarning - 1] to probe just below it.
|
|
const firstWarning = Math.ceil(WINDOW * WARNING_THRESHOLD); // = 6000
|
|
const rangeMin = Math.max(0, firstWarning - W); // = 5950
|
|
const rangeMax = firstWarning - 1; // = 5999
|
|
fc.assert(
|
|
fc.property(
|
|
fc.integer({ min: rangeMin, max: rangeMax }),
|
|
(tokensUsed) => {
|
|
const r = classifyContextUtilization(tokensUsed, WINDOW);
|
|
assert.equal(
|
|
r.state,
|
|
STATES.HEALTHY,
|
|
`tokensUsed=${tokensUsed}/${WINDOW}=${(tokensUsed / WINDOW * 100).toFixed(2)}% expected healthy got ${r.state}`
|
|
);
|
|
}
|
|
)
|
|
);
|
|
});
|
|
|
|
test('property: near 70% boundary tokens at/above critical threshold must be critical', () => {
|
|
const criticalFloor = Math.ceil(WINDOW * CRITICAL_THRESHOLD);
|
|
fc.assert(
|
|
fc.property(
|
|
fc.integer({ min: criticalFloor, max: WINDOW }),
|
|
(tokensUsed) => {
|
|
const r = classifyContextUtilization(tokensUsed, WINDOW);
|
|
assert.equal(
|
|
r.state,
|
|
STATES.CRITICAL,
|
|
`tokensUsed=${tokensUsed}/${WINDOW}=${(tokensUsed / WINDOW * 100).toFixed(2)}% expected critical got ${r.state}`
|
|
);
|
|
}
|
|
)
|
|
);
|
|
});
|
|
|
|
// ─── (b) Robustness: hostile inputs always throw TypeError ────────────────────
|
|
test('property: non-integer/negative tokensUsed always throws TypeError', () => {
|
|
const invalidTokensUsed = fc.oneof(
|
|
fc.constant(null),
|
|
fc.constant(undefined),
|
|
fc.constant(NaN),
|
|
fc.constant(Infinity),
|
|
fc.constant(-Infinity),
|
|
fc.constant(-1),
|
|
fc.integer({ min: -10000, max: -1 }),
|
|
fc.double({ min: 0.1, max: 0.9 }),
|
|
fc.string(),
|
|
fc.boolean(),
|
|
fc.constant([]),
|
|
fc.constant({})
|
|
);
|
|
fc.assert(
|
|
fc.property(invalidTokensUsed, (bad) => {
|
|
assert.throws(
|
|
() => classifyContextUtilization(bad, 10000),
|
|
(err) => {
|
|
assert.ok(err instanceof TypeError, `Expected TypeError but got ${err.constructor.name}: ${err.message}`);
|
|
return true;
|
|
}
|
|
);
|
|
})
|
|
);
|
|
});
|
|
|
|
test('property: non-integer/non-positive contextWindow always throws TypeError', () => {
|
|
const invalidWindows = fc.oneof(
|
|
fc.constant(null),
|
|
fc.constant(undefined),
|
|
fc.constant(NaN),
|
|
fc.constant(Infinity),
|
|
fc.constant(-Infinity),
|
|
fc.constant(0),
|
|
fc.constant(-1),
|
|
fc.integer({ min: -10000, max: 0 }),
|
|
fc.double({ min: 0.1, max: 0.9 }),
|
|
fc.string(),
|
|
fc.boolean()
|
|
);
|
|
fc.assert(
|
|
fc.property(invalidWindows, (bad) => {
|
|
assert.throws(
|
|
() => classifyContextUtilization(1000, bad),
|
|
(err) => {
|
|
assert.ok(err instanceof TypeError, `Expected TypeError for contextWindow=${bad} but got ${err.constructor.name}: ${err.message}`);
|
|
return true;
|
|
}
|
|
);
|
|
})
|
|
);
|
|
});
|
|
|
|
// ─── (c) Return shape: all valid inputs produce typed { percent, state } ──────
|
|
//
|
|
// Previously used Math.random() inside fc.property which broke reproducibility
|
|
// under the pinned seed (seed=42). Fixed: tokensUsed is now a seeded fc.integer
|
|
// arbitrary, making both inputs part of the shrinkable, reproducible input tuple.
|
|
//
|
|
// Split into two sub-properties:
|
|
// (c1) shape-only — result is an object with the right field types and ranges
|
|
// (c2) value-correctness — percent value matches the expected ratio arithmetic
|
|
// at three known representative ratios (0%, 50%, 100%)
|
|
|
|
test('property: valid inputs always return { percent: number[0..100], state: string }', () => {
|
|
fc.assert(
|
|
fc.property(
|
|
fc.integer({ min: 1, max: 1_000_000 }), // contextWindow
|
|
fc.integer({ min: 0, max: 1_000_000 }), // tokensUsed (upper-bound clamped below)
|
|
(contextWindow, rawTokens) => {
|
|
// Clamp so tokensUsed is always in [0, contextWindow] — same domain as
|
|
// the former Math.random() draw but now seeded and shrinkable.
|
|
const tokensUsed = rawTokens % (contextWindow + 1);
|
|
const r = classifyContextUtilization(tokensUsed, contextWindow);
|
|
|
|
assert.ok(typeof r === 'object' && r !== null, 'result must be object');
|
|
assert.ok(typeof r.percent === 'number', `percent must be number got ${typeof r.percent}`);
|
|
assert.ok(typeof r.state === 'string', `state must be string got ${typeof r.state}`);
|
|
assert.ok(r.percent >= 0 && r.percent <= 100, `percent ${r.percent} out of [0,100]`);
|
|
assert.ok(
|
|
[STATES.HEALTHY, STATES.WARNING, STATES.CRITICAL].includes(r.state),
|
|
`state must be one of the STATES enum, got ${r.state}`
|
|
);
|
|
}
|
|
)
|
|
);
|
|
});
|
|
|
|
test('property: percent value matches ratio arithmetic at known representative ratios', () => {
|
|
// Use a fixed contextWindow of 10000 so exact percent values are predictable.
|
|
// Three known points: 0% (healthy), 50% (healthy), 100% (critical).
|
|
const knownCases = [
|
|
{ tokensUsed: 0, expectedPercent: 0, expectedState: STATES.HEALTHY },
|
|
{ tokensUsed: 5000, expectedPercent: 50, expectedState: STATES.HEALTHY },
|
|
{ tokensUsed: 10000, expectedPercent: 100, expectedState: STATES.CRITICAL },
|
|
];
|
|
for (const { tokensUsed, expectedPercent, expectedState } of knownCases) {
|
|
const r = classifyContextUtilization(tokensUsed, WINDOW);
|
|
assert.equal(
|
|
r.percent,
|
|
expectedPercent,
|
|
`tokensUsed=${tokensUsed}/${WINDOW}: expected percent=${expectedPercent} got ${r.percent}`
|
|
);
|
|
assert.equal(
|
|
r.state,
|
|
expectedState,
|
|
`tokensUsed=${tokensUsed}/${WINDOW}: expected state=${expectedState} got ${r.state}`
|
|
);
|
|
}
|
|
});
|
|
|
|
test('property: tokensUsed exceeding contextWindow clamps to 100% critical', () => {
|
|
fc.assert(
|
|
fc.property(
|
|
fc.integer({ min: 1, max: 100_000 }),
|
|
fc.integer({ min: 1, max: 100_000 }),
|
|
(contextWindow, extra) => {
|
|
const tokensUsed = contextWindow + extra; // always exceeds window
|
|
const r = classifyContextUtilization(tokensUsed, contextWindow);
|
|
assert.equal(r.state, STATES.CRITICAL, `overflow ${tokensUsed}/${contextWindow} must be critical`);
|
|
assert.equal(r.percent, 100, `overflow percent must clamp to 100, got ${r.percent}`);
|
|
}
|
|
)
|
|
);
|
|
});
|
|
});
|