* feat(#25): scope gsd-verifier Step 7b to enumerate-or-single-test; forbid full-suite re-runs Step 7b's lone test example (`npm test -- --grep "$PHASE_TEST_PATTERN"`) is mocha/vitest/jest-specific, where `--grep` filters which tests *execute*. Models generalized it to `cargo test --workspace 2>&1 | grep X` (runs the whole suite, filters only *output*) and repeated it once per must-have, adding minutes per verification with no new evidence after the first run. Replace the example with language-agnostic guidance: prove a test EXISTS via enumeration (`cargo test -- --list` / `pytest --collect-only` / `npx vitest list` / `go test -list`), and prove it PASSES via a single named test (`cargo test <name> -- --exact` / `pytest -k` / `npx vitest run -t`). Add a Spot-check constraint forbidding more than one full-suite run per verification or piping a full run through grep per must-have, while still permitting one saved run + grep when a full run is genuinely required. docs/AGENTS.md gains a one-line Key-behaviors note, and a new test asserts the Step 7b content. Scoped per the maintainer decision on the issue: folded into Step 7b (no new top-level Step 7a) with no VERIFICATION.md label changes. Closes #25 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#25): add Changed changeset fragment for PR #753 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
5
.changeset/gentle-orcas-click.md
Normal file
5
.changeset/gentle-orcas-click.md
Normal file
@@ -0,0 +1,5 @@
|
||||
---
|
||||
type: Changed
|
||||
pr: 753
|
||||
---
|
||||
The `gsd-verifier` agent no longer re-runs the full workspace test suite once per must-have during Step 7b spot-checks — it enumerates tests to prove existence and runs a single named test to prove a pass, invoking the full suite at most once per verification.
|
||||
@@ -466,8 +466,11 @@ ls $BUILD_OUTPUT_DIR/*.{js,css} 2>/dev/null | wc -l
|
||||
# Module exports expected functions
|
||||
node -e "const m = require('$MODULE_PATH'); console.log(typeof m.$FUNCTION_NAME)" 2>/dev/null | grep -q "function"
|
||||
|
||||
# Test suite passes (if tests exist for this phase's code)
|
||||
npm test -- --grep "$PHASE_TEST_PATTERN" 2>&1 | grep -q "passing"
|
||||
# A test EXISTS (existence proof — enumerate, do NOT run the suite)
|
||||
cargo test -- --list 2>/dev/null | grep -q "$PHASE_TEST_PATTERN" # pytest --collect-only -q · npx vitest list · go test -list '.*'
|
||||
|
||||
# A specific test PASSES (run ONE named test, never the whole suite)
|
||||
cargo test "$TEST_NAME" -- --exact # pytest -k "$TEST_NAME" · npx vitest run -t "$TEST_NAME"
|
||||
```
|
||||
|
||||
2. **Run each check** and record pass/fail:
|
||||
@@ -487,6 +490,7 @@ npm test -- --grep "$PHASE_TEST_PATTERN" 2>&1 | grep -q "passing"
|
||||
- Each check must complete in under 10 seconds
|
||||
- Do not start servers or services — only test what's already runnable
|
||||
- Do not modify state (no writes, no mutations, no side effects)
|
||||
- **Run the full workspace test command at most once per verification.** Never filter a full run per must-have (`<full-suite> 2>&1 | grep X` repeated per truth) — it re-runs everything and yields no new evidence. Prove a test exists by enumeration (`--list` / `--collect-only`); prove one passes via a single named test. If a full run is genuinely required, run it once and `grep` the saved output.
|
||||
- If the project has no runnable entry points yet, skip with: "Step 7b: SKIPPED (no runnable entry points)"
|
||||
|
||||
## Step 7c: Probe Execution
|
||||
|
||||
@@ -292,6 +292,7 @@ GSD uses a multi-agent architecture where thin orchestrators (workflow files) sp
|
||||
- Logs issues for `/gsd-verify-work` to address
|
||||
- Milestone scope filtering: gaps addressed in later phases are marked as "deferred", not reported as failures (v1.32)
|
||||
- **Test quality audit** (v1.32): verifies that tests prove what they claim by checking for disabled/skipped tests on requirements, circular test patterns (system generating its own expected values), assertion strength (existence vs. value vs. behavioral), and expected value provenance. Blockers from test quality audit override an otherwise passing verification
|
||||
- Runs the full workspace test suite at most once per verification — proves a test *exists* by enumeration and that it *passes* via a single named test, never re-running the whole suite per must-have.
|
||||
|
||||
---
|
||||
|
||||
|
||||
57
tests/verifier-spotcheck-test-discipline.test.cjs
Normal file
57
tests/verifier-spotcheck-test-discipline.test.cjs
Normal file
@@ -0,0 +1,57 @@
|
||||
'use strict';
|
||||
const test = require('node:test');
|
||||
const assert = require('node:assert');
|
||||
const fs = require('node:fs');
|
||||
const path = require('node:path');
|
||||
|
||||
// Issue #25: gsd-verifier Step 7b must not present the runner-specific
|
||||
// full-suite "| grep" anti-pattern, and must steer toward
|
||||
// enumerate-for-existence / single-named-test-for-pass.
|
||||
const verifierPath = path.join(__dirname, '..', 'agents', 'gsd-verifier.md');
|
||||
const content = fs.readFileSync(verifierPath, 'utf8');
|
||||
|
||||
// Scope assertions to the Step 7b section so the guard tracks the right place.
|
||||
const start = content.indexOf('## Step 7b');
|
||||
const end = content.indexOf('## Step 7c', start);
|
||||
assert.ok(start !== -1 && end !== -1 && end > start, 'Step 7b section not found');
|
||||
const step7b = content.slice(start, end);
|
||||
|
||||
test('Step 7b drops the misleading full-suite grep example', () => {
|
||||
assert.ok(
|
||||
!step7b.includes('npm test -- --grep "$PHASE_TEST_PATTERN" 2>&1 | grep -q "passing"'),
|
||||
'the runner-specific `npm test --grep` example should be removed (it mis-generalizes to `<full-suite> | grep`)'
|
||||
);
|
||||
});
|
||||
|
||||
test('Step 7b steers existence proofs to test enumeration', () => {
|
||||
assert.ok(
|
||||
/cargo test -- --list/.test(step7b),
|
||||
'Step 7b should show the correct `cargo test -- --list` enumeration form'
|
||||
);
|
||||
assert.ok(
|
||||
/pytest --collect-only/.test(step7b) &&
|
||||
/vitest list/.test(step7b) &&
|
||||
/go test -list/.test(step7b),
|
||||
'Step 7b should list cross-ecosystem enumeration commands (pytest/vitest/go)'
|
||||
);
|
||||
});
|
||||
|
||||
test('Step 7b steers pass-checks to a single named test', () => {
|
||||
assert.ok(
|
||||
/-- --exact/.test(step7b) &&
|
||||
/pytest -k/.test(step7b) &&
|
||||
/vitest run -t/.test(step7b),
|
||||
'Step 7b should show single-named-test commands across ecosystems'
|
||||
);
|
||||
});
|
||||
|
||||
test('Step 7b forbids re-running the full suite per must-have', () => {
|
||||
assert.ok(
|
||||
/at most once per verification/i.test(step7b),
|
||||
'Step 7b should forbid invoking the full workspace test command more than once per verification'
|
||||
);
|
||||
assert.ok(
|
||||
/grep/i.test(step7b) && /per must-have|per truth/i.test(step7b),
|
||||
'Step 7b should explicitly call out the per-must-have full-suite grep anti-pattern'
|
||||
);
|
||||
});
|
||||
Reference in New Issue
Block a user