* fix: recurse test discovery so subdir test suites actually run
scripts/run-tests.cjs discovered tests with a flat readdirSync(testDir),
silently excluding tests/observability/ (4 files), tests/dispatch/ (1) and
tests/installer-migrations/ (1) — 94 passing tests — from `npm test` and all
CI lanes. Walk the tree recursively (relative subpaths preserved), classify
suites by basename, and add a fail-on-zero-executed guard for suite/default
runs (escape hatch GSD_ALLOW_EMPTY_SUITE=1) while preserving the empty
--files/--files-from path the CI inert lane relies on.
Unit suite 735 -> 741 files; surfaces ADR-227's observability/dispatch seam.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: retire 5 verified-worthless tests
Adversarial verification confirmed these 5 prove nothing — their coverage is
provided more strictly elsewhere:
- enh-2790 'has a name: field' spot-checks (command-contract enforces /^gsd[:-]/)
- command-routing-hub duplicate construct + duplicate ERROR_KINDS assertions
- no-cjs-sdk-handsync-tooling (guarded files that never existed on main; bug-190
covers the real retired SDK artifacts)
- runtime-artifact-layout cline edge case (subsumed by the explicit-global test
and bug-782-cline-skills-emission)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: add ADR-218 release version-validation coverage
ADR-218 (reject leading-zero versions like 1.01.0; npm duplicate pre-check) had
zero tests — the logic lived only in release.yml bash. Add a test that extracts
the actual rejection regexes from the workflow and exercises them against a
boundary table (leading-zero/malformed rejected, valid accepted) plus structural
wiring assertions. Goes red if the regex is reverted to [0-9]+.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: redesign weak tests into behavioral, deterministic assertions
Per the ADR test audit, rewrite 27 weak test files (test-only, no source
changes) so each can go red for the defect it guards:
- kill pass-always assert.ok(true) placeholders (research-cli, worktree-baseref,
bug-260 security guard, eslint-rules x24, clusters '|| true')
- replace source-text grep with behavioral calls (install Kilo, sh-hook-paths,
plan-review-convergence) and add a repo-layout governance test
- de-flake real-clock/Math.random coupling (phase last_updated, bug-3707 mtime,
context-utilization property, feat-3594)
- fix independence/shared-state violations (bug-492 singleton, issue-844 tmpRoot,
core reapStaleTempFiles, active-workstream TTY, feat-488 GSD_HOME)
- strengthen property/shape-only tests (research-provider/store classification +
collision) and unconditional plugin.json schema validation (issue-766)
Verified: all 28 files run together 1220 pass / 0 fail / 1 skip.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: add no-tautological-assert lint rule, error in test suite
New custom ESLint rule (eslint-rules/no-tautological-assert.cjs) bans asserts
that can never fail: assert(true)/assert.ok(<always-truthy literal>),
'cond || true' inside an assert, and equality asserts comparing two identical
literals. Wired as error on tests/**; full sweep confirmed zero existing
violations so the suite stays green. Prevents the placeholder-assert regressions
the audit redesigns just removed. RuleTester coverage added (6 valid, 8 invalid).
Note: no-only-tests was already enforced via eslint-plugin-no-only-tests, so no
duplicate rule was added.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gate new allow-test-rule exemptions to require an issue ref
ADR-456 requires any allow-test-rule exemption added after the ADR to carry a
tracking issue number, but nothing enforced it. New ratchet gate
(scripts/lint-allow-test-rule-refs.cjs, wired into lint:ci) fails when a NEW
allow-test-rule comment lacks a #NNN/URL reference; the 323 existing untracked
exemptions are grandfathered in an allowlist that ratchets down as they gain
refs. Red-green verified (novel untracked offender fails; compliant passes).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs: add ADR test-audit evidence report (#1192)
Full risk-first qa-test-architect audit of the ADR portfolio (37 ADRs + 4
platform lenses, adversarial verification of retire verdicts) that drove the
P0 discovery fix, ADR-218 coverage, 5 retires, 27 redesigns, and the two new
lint gates. Filed as point-in-time evidence under docs/issueevidence/, named
for tracking issue #1192.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: replace pre-existing raw NUL byte with escape in feat-3594 fixture
feat-3594's null-byte parser fixture contained a literal NUL byte (pre-existing
on next at b10e5681 — confirmed: base blob has 1 NUL, this fix has 0), which
made git treat the file as binary and would break grep/editors. Switch to the
\x00 escape; the runtime string value (a real NUL in the parser input) is
unchanged.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: address adversarial-review findings
Codex adversarial pass over the branch:
- capability-registry drift test no longer mutates the committed generated
capability-registry.cjs in place (concurrency hazard) — uses in-memory
checkPipeline comparison instead.
- allow-test-rule ratchet now detects exemptions in ALL comment forms (block
/* */ too, matching no-source-grep) so a block comment can't bypass it;
one newly-surfaced pre-existing offender grandfathered (323->324).
- install.test Kilo case asserts on what install(false,'kilo') actually writes
rather than manually calling configureKiloPermissions (masked the call site).
- issue-766 drops the undeclared transitive ajv dep for explicit structural
assertions from the schema fixture.
- adr-218 test notes the hotfix leading-zero gap is tracked in #1186.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: address code-review findings (subdir discovery, rule + test gaps)
xhigh code review surfaced 15 confirmed issues, all fixed:
- run-tests.cjs --files now resolves subdir tests by bare basename + handles
Windows backslash paths (ambiguous basenames error clearly).
- affected-tests-lib.cjs listTestFiles made recursive — the targeted CI lane was
silently dropping changed subdir tests (same false-green class the audit fixed).
- no-tautological-assert now catches 'true || cond' and empty []/{} equality.
- verify-test-quality: restore provenance-classification coverage, tighten the
writeFile circular-detection check, guard the module-level file read.
- sh-hook-paths: cover the global-install .sh delegation branch (#2045 guard).
- active-workstream null-guard runs deterministically (no longer skipped on TTY).
- adr-218 structural guards tightened (major/minor leading-zero; needs: membership).
- repo-layout AGENTS.md guard no longer false-alarms on equivalent refactors.
- cross-ai ordering guard fails red when the step is missing.
- issue-766 parses required fields from the schema fixture (auto-enforced).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: stub USERPROFILE alongside HOME in feat-488 (Windows parity)
The feat-488 redesign stubbed process.env.HOME but not USERPROFILE; os.homedir()
resolves from USERPROFILE on Windows, so the home stub was not hermetic there —
caught by windows-test-parity-guard (stubsHomeNoUserProfile). Save/set/restore
USERPROFILE symmetrically with HOME (delete-if-originally-undefined).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: reconcile allow-test-rule allowlist after rebase onto next
Rebasing onto current next pulled in merged PR #1170, which added
inventory-headings-countfree.test.cjs (a baseline allow-test-rule exemption) and
deleted inventory-counts.test.cjs. Grandfather the former and prune the latter so
the ratchet matches the merged tree. No new debt from this PR.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
226 lines
9.5 KiB
JavaScript
226 lines
9.5 KiB
JavaScript
/**
|
|
* Model Profiles Tests
|
|
*
|
|
* Tests for MODEL_PROFILES data structure, VALID_PROFILES list,
|
|
* formatAgentToModelMapAsTable, getAgentToModelMapForProfile,
|
|
* and resolveModelInternal precedence (override > profile > default).
|
|
*/
|
|
|
|
const { test, describe, beforeEach, afterEach } = require('node:test');
|
|
const assert = require('node:assert/strict');
|
|
|
|
const fs = require('node:fs');
|
|
const path = require('node:path');
|
|
|
|
const {
|
|
MODEL_PROFILES,
|
|
VALID_PROFILES,
|
|
formatAgentToModelMapAsTable,
|
|
getAgentToModelMapForProfile,
|
|
} = require('../gsd-core/bin/lib/model-profiles.cjs');
|
|
|
|
const { resolveModelInternal } = require('../gsd-core/bin/lib/model-resolver.cjs');
|
|
const { createTempProject, cleanup } = require('./helpers.cjs');
|
|
|
|
// ─── temp-project helpers ──────────────────────────────────────────────────────
|
|
|
|
function writeConfig(tmpDir, obj) {
|
|
fs.writeFileSync(
|
|
path.join(tmpDir, '.planning', 'config.json'),
|
|
JSON.stringify(obj, null, 2),
|
|
'utf-8'
|
|
);
|
|
}
|
|
|
|
function agentFilesOnDisk() {
|
|
return fs.readdirSync(path.join(__dirname, '..', 'agents'))
|
|
.filter((f) => /^gsd-.*\.md$/.test(f))
|
|
.map((f) => f.replace(/\.md$/, ''))
|
|
.sort();
|
|
}
|
|
|
|
// ─── MODEL_PROFILES data integrity ────────────────────────────────────────────
|
|
|
|
describe('MODEL_PROFILES', () => {
|
|
test('contains every shipped gsd agent file on disk (#3229)', () => {
|
|
const expectedAgents = agentFilesOnDisk();
|
|
const actualAgents = Object.keys(MODEL_PROFILES).sort();
|
|
assert.deepStrictEqual(actualAgents, expectedAgents);
|
|
});
|
|
|
|
test('every agent has quality, balanced, budget, and adaptive profiles', () => {
|
|
for (const [agent, profiles] of Object.entries(MODEL_PROFILES)) {
|
|
assert.ok(profiles.quality, `${agent} missing quality profile`);
|
|
assert.ok(profiles.balanced, `${agent} missing balanced profile`);
|
|
assert.ok(profiles.budget, `${agent} missing budget profile`);
|
|
assert.ok(profiles.adaptive, `${agent} missing adaptive profile`);
|
|
}
|
|
});
|
|
|
|
test('all profile values are valid model aliases', () => {
|
|
const validModels = ['opus', 'sonnet', 'haiku'];
|
|
for (const [agent, profiles] of Object.entries(MODEL_PROFILES)) {
|
|
for (const [profile, model] of Object.entries(profiles)) {
|
|
assert.ok(
|
|
validModels.includes(model),
|
|
`${agent}.${profile} has invalid model "${model}" — expected one of ${validModels.join(', ')}`
|
|
);
|
|
}
|
|
}
|
|
});
|
|
|
|
test('quality profile never uses haiku', () => {
|
|
for (const [agent, profiles] of Object.entries(MODEL_PROFILES)) {
|
|
assert.notStrictEqual(
|
|
profiles.quality, 'haiku',
|
|
`${agent} quality profile should not use haiku`
|
|
);
|
|
}
|
|
});
|
|
});
|
|
|
|
// ─── VALID_PROFILES ───────────────────────────────────────────────────────────
|
|
|
|
describe('VALID_PROFILES', () => {
|
|
test('contains quality, balanced, budget, adaptive, and inherit', () => {
|
|
assert.deepStrictEqual(VALID_PROFILES.sort(), ['adaptive', 'balanced', 'budget', 'inherit', 'quality']);
|
|
});
|
|
|
|
test('includes all MODEL_PROFILES keys plus inherit', () => {
|
|
const fromData = Object.keys(MODEL_PROFILES['gsd-planner']);
|
|
for (const profile of fromData) {
|
|
assert.ok(VALID_PROFILES.includes(profile), `VALID_PROFILES should include ${profile}`);
|
|
}
|
|
assert.ok(VALID_PROFILES.includes('inherit'), 'VALID_PROFILES should include inherit');
|
|
});
|
|
});
|
|
|
|
// ─── getAgentToModelMapForProfile ─────────────────────────────────────────────
|
|
|
|
describe('getAgentToModelMapForProfile', () => {
|
|
test('returns correct models for balanced profile', () => {
|
|
const map = getAgentToModelMapForProfile('balanced');
|
|
assert.strictEqual(map['gsd-planner'], 'opus');
|
|
assert.strictEqual(map['gsd-codebase-mapper'], 'haiku');
|
|
assert.strictEqual(map['gsd-verifier'], 'sonnet');
|
|
});
|
|
|
|
test('returns correct models for budget profile', () => {
|
|
const map = getAgentToModelMapForProfile('budget');
|
|
assert.strictEqual(map['gsd-planner'], 'sonnet');
|
|
assert.strictEqual(map['gsd-phase-researcher'], 'haiku');
|
|
});
|
|
|
|
test('returns correct models for quality profile', () => {
|
|
const map = getAgentToModelMapForProfile('quality');
|
|
assert.strictEqual(map['gsd-planner'], 'opus');
|
|
assert.strictEqual(map['gsd-executor'], 'opus');
|
|
});
|
|
|
|
test('returns correct models for adaptive profile', () => {
|
|
const map = getAgentToModelMapForProfile('adaptive');
|
|
assert.strictEqual(map['gsd-planner'], 'opus', 'planner should use opus in adaptive');
|
|
assert.strictEqual(map['gsd-debugger'], 'opus', 'debugger should use opus in adaptive');
|
|
assert.strictEqual(map['gsd-executor'], 'sonnet', 'executor should use sonnet in adaptive');
|
|
assert.strictEqual(map['gsd-codebase-mapper'], 'haiku', 'mapper should use haiku in adaptive');
|
|
assert.strictEqual(map['gsd-plan-checker'], 'haiku', 'checker should use haiku in adaptive');
|
|
});
|
|
|
|
// ─── resolution order: override > profile > default ─────────────────────────
|
|
// Uses gsd-phase-researcher because it has visibly distinct values at every
|
|
// level: balanced (default) = sonnet, budget (profile) = haiku, override = opus.
|
|
// Each tier must beat the one below it; the test goes RED if resolveModelInternal
|
|
// ignores model_overrides (returns 'haiku') or conflates default with profile
|
|
// (returns 'sonnet' instead of 'haiku' for budget).
|
|
describe('resolution order: override > profile > default', () => {
|
|
// agent under test — must have three distinct model values across tiers
|
|
const AGENT = 'gsd-phase-researcher';
|
|
const EXPECTED_DEFAULT = 'sonnet'; // balanced profile (no config)
|
|
const EXPECTED_PROFILE = 'haiku'; // budget profile
|
|
const EXPECTED_OVERRIDE = 'opus'; // explicit model_overrides entry
|
|
|
|
let tmpDir;
|
|
beforeEach(() => { tmpDir = createTempProject(); });
|
|
afterEach(() => { cleanup(tmpDir); tmpDir = null; });
|
|
|
|
test('default (no config) resolves to balanced profile model', () => {
|
|
// Sanity-check: balanced is the profile tier when no config is present.
|
|
assert.strictEqual(
|
|
resolveModelInternal(tmpDir, AGENT),
|
|
EXPECTED_DEFAULT,
|
|
`expected balanced-profile default "${EXPECTED_DEFAULT}" but got a different model`
|
|
);
|
|
});
|
|
|
|
test('profile setting (budget) beats the balanced default', () => {
|
|
writeConfig(tmpDir, { model_profile: 'budget' });
|
|
assert.strictEqual(
|
|
resolveModelInternal(tmpDir, AGENT),
|
|
EXPECTED_PROFILE,
|
|
`expected budget-profile model "${EXPECTED_PROFILE}" but got a different model`
|
|
);
|
|
});
|
|
|
|
test('model_overrides entry beats the active profile', () => {
|
|
// budget profile would give haiku; override must win with opus
|
|
writeConfig(tmpDir, {
|
|
model_profile: 'budget',
|
|
model_overrides: { [AGENT]: EXPECTED_OVERRIDE },
|
|
});
|
|
assert.strictEqual(
|
|
resolveModelInternal(tmpDir, AGENT),
|
|
EXPECTED_OVERRIDE,
|
|
`expected override "${EXPECTED_OVERRIDE}" to beat budget-profile model "${EXPECTED_PROFILE}"`
|
|
);
|
|
});
|
|
|
|
test('model_overrides beats the default profile too (no explicit profile key)', () => {
|
|
// Even without an explicit model_profile, override still wins over default
|
|
writeConfig(tmpDir, {
|
|
model_overrides: { [AGENT]: EXPECTED_OVERRIDE },
|
|
});
|
|
assert.strictEqual(
|
|
resolveModelInternal(tmpDir, AGENT),
|
|
EXPECTED_OVERRIDE,
|
|
`expected override "${EXPECTED_OVERRIDE}" to beat balanced default "${EXPECTED_DEFAULT}"`
|
|
);
|
|
});
|
|
});
|
|
|
|
test('returns all agents in the map', () => {
|
|
const map = getAgentToModelMapForProfile('balanced');
|
|
const agentCount = Object.keys(MODEL_PROFILES).length;
|
|
assert.strictEqual(Object.keys(map).length, agentCount);
|
|
});
|
|
});
|
|
|
|
// ─── formatAgentToModelMapAsTable ─────────────────────────────────────────────
|
|
|
|
describe('formatAgentToModelMapAsTable', () => {
|
|
test('produces a table with header and separator', () => {
|
|
const map = { 'gsd-planner': 'opus', 'gsd-executor': 'sonnet' };
|
|
const table = formatAgentToModelMapAsTable(map);
|
|
assert.ok(table.includes('Agent'), 'should have Agent header');
|
|
assert.ok(table.includes('Model'), 'should have Model header');
|
|
assert.ok(table.includes('─'), 'should have separator line');
|
|
assert.ok(table.includes('gsd-planner'), 'should list agent');
|
|
assert.ok(table.includes('opus'), 'should list model');
|
|
});
|
|
|
|
test('pads columns correctly', () => {
|
|
const map = { 'a': 'opus', 'very-long-agent-name': 'haiku' };
|
|
const table = formatAgentToModelMapAsTable(map);
|
|
const lines = table.split('\n').filter(l => l.trim());
|
|
// Separator line uses ┼, data/header lines use │
|
|
const dataLines = lines.filter(l => l.includes('│'));
|
|
const pipePositions = dataLines.map(l => l.indexOf('│'));
|
|
const unique = [...new Set(pipePositions)];
|
|
assert.strictEqual(unique.length, 1, 'all data lines should align on │');
|
|
});
|
|
|
|
test('handles empty map', () => {
|
|
const table = formatAgentToModelMapAsTable({});
|
|
assert.ok(table.includes('Agent'), 'should still have header');
|
|
});
|
|
});
|