* fix(#3951): two lint rules that could not reach the code they govern B6 names two widenings. Measuring them first turned up a defect the criterion did not know about, and refuted the reason it gave for one of them. 1. no-adhoc-markdown-parsing self-gates on its own filename. Lines 107-110 short-circuit create() to {} unless the path matches /(?:^|\/)src\/[^/]+\.cts$/. B6 says to widen the files: glob in eslint.config.mjs - but doing only that ships an INERT rule, because the gate still returns {} for every new path. Both halves have to change, and the gate is the load-bearing one. That same regex hides a live hole: [^/]+ is FLAT-ONLY, so it requires the file to sit directly in src/. The registered glob is src/**/*.cts, which includes subdirectories. 28 .cts files - health-diagnostic-rules/ (10), installer-migrations/ (11), observability/ (3), host-integration-adapters/ (2), vendor/ (2) - are inside the registered glob and silently skipped. Measured with the gate neutralized: 0 violations there today. The hole is hiding nothing right now, and is fixed anyway, because "no violations today" is not a property that keeps holding. The fix is not invented: require-subprocess-timeout.cjs:196 already carries the correct form of this guard, /(?:^|\/)src\/.*\.cts$/ with .*, one directory over. Checked the other 21 rules for the same bug - no-adhoc-regex-escape and no-private-binary-resolution short-circuit only to exempt their own seam file, which is the right shape, and no-crlf-fragile-split has no filename gate at all. This bug is unique to the one rule. 2. no-adhoc-regex-escape could not see the shape that actually occurs. Line 396 gated the whole UNSAFE-NEW-REGEXP arm on arg.type === 'Identifier'. Every check below it - the _SOURCE provenance check, the isSoleReturnOfOwnParameter shape - lives inside that branch, so new RegExp(obj['key']) and new RegExp(cfg.pattern) were never examined at all. Runtime data arrives as a property access far more often than as a bare identifier, which is exactly why this rule never fired on the #3477 ReDoS. Widened to MemberExpression, measured by AST walk across all five registered blocks rather than by grep. 27 sites, zero TSAsExpression: 18 safe new RegExp(X.source, flags) -> exempted, keyed strictly on the PROPERTY being `source`, never on the object. Keying on the object would wave through X.anything and buy nothing. B6 estimated ~10; that was an undercount. 3 _SOURCE-suffixed constants reached through a required module namespace (phaseId.BRACKET_PHASE_TOKEN_SOURCE) -> the same provenance-exempt class the rule already recognizes for bare identifiers, extended to reach them. Without this the widening produces 3 false flags. 6 real findings -> marked, each a test extracting a pattern from a shipped file at test time, where the runtime contract IS the product. Deliberately the NARROW MemberExpression form. The rule's own isSoleReturnOfOwnParameter doc comment records that an earlier broad "any non-literal identifier" heuristic produced ~25 false positives and was rejected; a re-run of the census after this change flags exactly the 6 above and nothing else. Verified by execution, not by reading: the gate now accepts src/<subdir>/x.cts, still accepts flat src/x.cts, and still exempts paths outside src/ - each pinned by a test proven to fail against the old regex. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): give no-adhoc-markdown-parsing its reach, and fix the 80 parses it finds The rule self-gates on filename AND is registered on one glob, so widening either half alone is inert. Both move here: the gate now accepts tests/**/*.cjs and scripts/**/*.cjs alongside src/**/*.cts, and eslint.config.mjs registers it on the same two. A test pins that the gate and the registration AGREE, in both directions. The original defect was a gate narrower than its registration; the failure mode of this fix is a gate wider than its registration. Both are silent, so the test asserts the pair rather than either half. 80 violations across 43 files, all in tests/, zero in scripts/. 70 are routed through the existing seams - scanFencedBlocks, collectSection, stripFencedCode, tokenizeHeadings from markdown-sectionizer; splitTableRow, parseMarkdownTable, findTableWithColumns from markdown-table. Headerless STATE.md tables use splitTableRow per line, because parseMarkdownTable needs a real delimiter row. 10 are suppressed, 12.5%, well under the third that would have meant the rule is mis-scoped for tests/ rather than the tests carrying debt. Each names its reason: three regression guards (#3873 / bug-#21) are deliberately independent of the generator's own fence handling, and routing them through the seam would have them test the generator against itself; one is a negative-text probe that extracts nothing; six are a shell-pipe-to-jq detector whose regex coincidentally matches the table fingerprint and is not markdown parsing at all. All ten sit in tests whose subject is .md content, which is normally a reason to prefer the seam. The marker used is allow-adhoc-markdown, distinct from no-source-grep's allow-test-rule, and lint:ci's lint-allow-test-rule-refs reports the same 280/280 unverified count as before - checked rather than assumed, because those two markers are easy to conflate. The widening earned its keep immediately: it found a test that passed for the wrong reason. tests/config-field-docs.test.cjs asserted notEqual(<cell>, '600') against the TYPE column instead of the DEFAULT column. notEqual('number', '600') is true forever, so the guard against workflow.subagent_timeout regressing to the old seconds default could never fire. docs/CONFIGURATION.md:434 is `| workflow.subagent_timeout | number | 300000 | ... |`, so the default is cell index 2; the assertion is now row-scoped through splitTableRow and reads 300000. That is the argument for the widening in one case: the violation was invisible to lint, the suite was green, and the assertion was vacuous. A rule that cannot reach a file cannot tell you the file is lying. Not fixed here, and recorded rather than assumed: #3426/#3239 are NOT reachable by this widening. tests/package-legitimacy-gate.test.cjs yields zero violations even with the gate bypassed - its hand-rolled scans are real, but built from line filters and split('|') rather than the regex-literal fingerprints this rule detects. They need new detectors. The epic assumed a wider glob would catch them. build:lib, lint and lint:ci all exit 0; the post-fix census across tests/** and scripts/** is 0 violations. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): B7 — and #3356's defects were still live in the code B7 asks that each closed child be driven fail-first with a behavioral identity test at the CONSUMER's output. Four of eleven children had no test citing their issue number. Auditing them by BEHAVIOR rather than by number-grep changed the answer for three of the four. #3364 and #2540 — traceability only. Both were implemented by #3941 and their consumer-output tests exist and were shown failing-first; neither cited its originating issue, so an audit that greps for the number reports them uncovered. Tagged the specific asserting test in each file, following the citation form those files already use. #3372 — covered, but only at helper level, and the triage narrowed it. Of the four commands the issue names, only estimate-cli's collectCalibrationSamples actually enumerates phase dirs from disk; smart-entry, audit and roadmap-upgrade derive from ROADMAP/body text and never reach the sentinel path, so they are benign by construction and were left alone rather than "fixed" into churn. The existing #3882 rows asserted the helper's return value. Added a consumer-output test driving `query estimate-calibrate` and asserting sample_count and the persisted document. RED proof: reverted collectCalibrationSamples to a raw readdirSync and ran the real CLI - sample_count 3, sentinel leaked; restored - sample_count 2. #3356 — NOT covered, and BOTH halves of the defect were still live in source. The issue is closed; the bug was not fixed. Fixed here rather than writing tests that document a bug as correct. Defect 1, the contradicted row. quick.md:627 claimed `quick-tasks-append` performs "the equivalent write" to the Step 7c row. It did not: the `#` cell was a positional ordinal and `Directory` read `—`, because the route had no way to receive a quick id or task directory. Added OPTIONAL `--quick-id` / `--slug` / `--directory`. A caller with neither - fast.md, the original #2133 caller - omits them and gets the byte-identical prior row, so nothing existing changes. A caller that HAS a real id and directory now gets the canonical row quick.md:632 renders. The false-equivalence sentence itself is corrected rather than left to mislead the next reader. Defect 2, the forced re-derive. The route called readModifyWriteStateMd with no options, so a body-only append to the Quick Tasks table triggered a full re-derive of the disk-derived progress.* frontmatter. Every other body-only writer passes { resync: false } - src/state.cts's own docstring prescribes it - and this route was the lone outlier. RED proof: reverted the option, seeded a project with 2 real phase dirs and a curated total_phases of 25, ran quick-tasks-append; total_phases collapsed to 2. Restored; it stayed 25. That second one is the shape this epic exists to close: a silent write that replaces curated state with a re-derivation nobody asked for, exit 0 throughout. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3951): amend B6's ledger to what was measured, and document the new flags The ADR gains a ledger amendment in its own correction style - the sixth wrong premise it records, found the same way as the other five, by measuring before building. B6 says the net guard count must fall. It rose: 62 -> 69, +7, measured from the epic's filing commit to origin/next. The attribution is the point, though. Five of the seven came from PRs unrelated to this epic, one was added by a phase of it, and the epic did retire something sub-file - #3884 removed a detector with an explicit "net: -1 detector, 0 added" ledger. Every named casualty is load-bearing, two already carry retractions in this same document, and a sweep of all 22 rules plus every scripts/lint-* found no provably dead guard. There is no honest way to make the count fall; forcing it would trade coverage for a number, which is the Goodhart outcome Decision 6 exists to prevent. The amendment also records that B6's own prescribed fix for one widening was inert. no-adhoc-markdown-parsing self-gates on its filename, so widening only the files: glob - which is what the criterion says to do - ships a rule that still returns {} for every new path. And #3426/#3239 are not reachable by that widening at all; their scans use line filters and split('|'), not the regex fingerprints the rule detects. The roster row tracked them against the wrong mechanism. Three roster rows updated from aspiration to fact: the two widenings are DONE with their measured counts, and lint-phase-enumeration-drift is marked RETAINED rather than "expected casualty - verify before retiring", because Phase 5 verified it and kept it. The rule Decision 6 should carry forward is stated plainly: a guard ledger is a claim about COVERAGE, not about COUNT. "Net count must fall" is measurable and wrong. "Every guard is reachable, and each retirement names what makes its defect unrepresentable" is the property that was actually wanted. CLI-TOOLS.md documents the optional --quick-id/--slug/--directory flags and says plainly that omitting them keeps the pre-#3356 row byte-identical, plus that the append no longer re-derives progress frontmatter. New features fragment (id 3951); FEATURES.md regenerated rather than hand-edited. Changeset is Changed, pr:0 pending backfill. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): correct four rows that pinned the lint rule's old narrow reach The remote suite came back RED with 5 failures, all in tests/eslint-rules.test.cjs. They are stale tests, not a regression: four rows assert that no-adhoc-markdown-parsing is inert outside src/*.cts, which is exactly the contract this deliverable changes. Confirmed by reading rather than inferred from the names - the row at :1981 used filename: 'tests/some.test.cjs' and filename: 'scripts/helper.cjs', the two roots the rule now covers on purpose. Worth recording WHY local gates missed this. npm run lint and lint:ci were green, and the touched test files passed standalone. Lint only reports violations in real files; these rows assert the rule's REACH using synthetic RuleTester filenames, so nothing but the full suite could see them. Local green on a rule change says nothing about the rule's own tests. Each row is rewritten with BOTH halves rather than flipped from valid to invalid: - the same fingerprint under tests/ or scripts/ is now flagged, with the right messageId - the negative space is preserved - the same fingerprint under a path outside all three roots (gsd-core/bin/lib/foo.cjs) is still NOT flagged The second half is the one that matters. Without it the rule has no boundary and nothing would catch an over-wide gate later, which is the mirror image of the bug this deliverable just fixed. Each row is renamed to state the current contract; the old names said "non-src/*.cts ... is not flagged" and would have been actively misleading once the bodies changed. Proven to test the widening rather than restate it: every flagged half was run against HEAD~2's pre-widening rule and does NOT fire there, then against the current rule and does. 12/12 on that probe; the full file is 178/178. Swept for the same staleness elsewhere and found none. require-subprocess-timeout's own "inert outside src/*.cts" row is untouched - that rule's gate was not widened here - and no-adhoc-regex-escape's test file already carries correctly-targeted rows. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): acknowledge the quick.md growth the attribution guard reported The full suite came back RED with one failure, and it is mine: 1 file(s) grew without an acknowledgment: quick.md grew 364 bytes gsd-core/workflows/quick.md is runtime-loaded emitted content, so correcting its false 'performs the equivalent write' claim trips emitted-attribution by construction. This is the acknowledgment, not a workaround - there is nothing to regenerate. The fragment names ONE path, which is the only one the guard reported. The four spent acknowledgments it also listed (audit-uat, plan-phase, progress, review) belong to other fragments whose ripple the base already absorbs; they are inert, not failures, and are deliberately NOT copied here - naming paths I did not change would make this record false in the other direction. Byte figure corrected before committing: the guard reported 37220 -> 37584 (+364), but origin/next has since moved and quick.md is 37232 there now, so the measured delta is +352. The reason text says so and names the base as a moving figure rather than pinning a number that is already stale. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): move the quick.md growth ack to a trailer, delete the obsolete fragment The acknowledgment mechanism changed under this branch. Merging next brought in the redesign - it also deleted .github/workflows/ack-fragment-sweep.yml, which was in the merge status and which I did not register at the time - and the guard now says so directly: Add a trailer to a commit in this PR (never a new file). Emitted-Drift-Ack-Growth: quick.md - <why this growth is deliberate> So tests/emitted-drift-acks/3951-quick-append-equivalence.json is obsolete on arrival. A fragment file is no longer read by anything, and leaving it would be a dead record that looks like an active one. It is deleted here rather than kept "just in case". The byte figure moved again with the merge: 37232 -> 37596, +364. The earlier fragment said +352, measured before the merge auto-merged quick.md itself. The trailer carries no number, which is the better design - the figure was stale twice in two attempts. Refs #3951 Emitted-Drift-Ack-Growth: quick.md — #3356/#3951 replaces a false claim with an accurate one. Line 627 said the `quick-tasks-append` shortcut "performs the equivalent write" to the Step 7c row rendered above it; it did not, and that was the documented half of #3356 — with no quick id or task directory the route emitted a positional ordinal in `#` and an em-dash in `Directory`, a visibly different row. The corrected sentence has to carry three facts the original elided: what the shortcut actually writes when it has neither input, that this is honest behavior for its real caller (`fast.md`, which has neither), and how a caller with both now gets the byte-identical canonical row via the new optional `--quick-id`/`--slug`/`--directory` flags. Prose is the product here — an executing agent reads this line to decide whether the shortcut is safe for its case, and a shorter correction would either drop the flags (leaving the reader unable to act on the fix) or drop the limitation (recreating the false claim in gentler words). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3951): backfill changeset pr number Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
684 lines
30 KiB
JavaScript
684 lines
30 KiB
JavaScript
// allow-test-rule: source-text-is-the-product
|
||
// Workflow .md / agent .md / command .md / reference .md files — their text
|
||
// IS what the runtime loads. Testing text content tests the deployed contract.
|
||
// Per CONTRIBUTING.md exception matrix.
|
||
'use strict';
|
||
|
||
|
||
/**
|
||
* Planner Language Regression Tests (#2091, #2092)
|
||
*
|
||
* Prevents time-based reasoning and complexity-as-scope-justification
|
||
* from leaking back into planning artifacts via future PRs.
|
||
*
|
||
* These tests scan agent definitions, workflow files, and references
|
||
* for prohibited patterns that import human-world constraints into
|
||
* an AI execution context where those constraints do not exist.
|
||
*/
|
||
|
||
const { test, describe } = require('node:test');
|
||
const assert = require('node:assert/strict');
|
||
const fs = require('fs');
|
||
const path = require('path');
|
||
const { collectSection } = require('../gsd-core/bin/lib/markdown-sectionizer.cjs');
|
||
|
||
const ROOT = path.join(__dirname, '..');
|
||
const AGENTS_DIR = path.join(ROOT, 'agents');
|
||
const WORKFLOWS_DIR = path.join(ROOT, 'gsd-core', 'workflows');
|
||
const REFERENCES_DIR = path.join(ROOT, 'gsd-core', 'references');
|
||
const TEMPLATES_DIR = path.join(ROOT, 'gsd-core', 'templates');
|
||
|
||
/**
|
||
* Collect all .md files from a directory (non-recursive).
|
||
*/
|
||
function mdFiles(dir) {
|
||
if (!fs.existsSync(dir)) return [];
|
||
return fs.readdirSync(dir)
|
||
.filter(f => f.endsWith('.md'))
|
||
.map(f => ({ name: f, path: path.join(dir, f) }));
|
||
}
|
||
|
||
/**
|
||
* Collect all .md files recursively.
|
||
*/
|
||
function mdFilesRecursive(dir) {
|
||
if (!fs.existsSync(dir)) return [];
|
||
const results = [];
|
||
for (const entry of fs.readdirSync(dir, { withFileTypes: true })) {
|
||
const full = path.join(dir, entry.name);
|
||
if (entry.isDirectory()) {
|
||
results.push(...mdFilesRecursive(full));
|
||
} else if (entry.name.endsWith('.md')) {
|
||
results.push({ name: entry.name, path: full });
|
||
}
|
||
}
|
||
return results;
|
||
}
|
||
|
||
/**
|
||
* Files that define planning behavior — agents, workflows, references.
|
||
* These are the files where time-based and complexity-based scope
|
||
* reasoning must never appear.
|
||
*/
|
||
const PLANNING_FILES = [
|
||
...mdFiles(AGENTS_DIR),
|
||
...mdFiles(WORKFLOWS_DIR),
|
||
...mdFiles(REFERENCES_DIR),
|
||
...mdFilesRecursive(TEMPLATES_DIR),
|
||
];
|
||
|
||
// -- Prohibited patterns --
|
||
|
||
/**
|
||
* Time-based task sizing patterns.
|
||
* Matches "15-60 minutes", "X minutes Claude execution time", etc.
|
||
* Does NOT match operational timeouts ("timeout: 5 minutes"),
|
||
* API docs examples ("100 requests per 15 minutes"),
|
||
* or human-readable timeout descriptions in workflow execution steps.
|
||
*/
|
||
const TIME_SIZING_PATTERNS = [
|
||
// "N-M minutes" in task sizing context (not timeout context)
|
||
/each task[:\s]*\*?\*?\d+[-–]\d+\s*min/i,
|
||
// "minutes Claude execution time" or "minutes execution time"
|
||
/minutes?\s+(claude\s+)?execution\s+time/i,
|
||
// Duration-based sizing table rows: "< 15 min", "15-60 min", "> 60 min"
|
||
/[<>]\s*\d+\s*min\s*\|/i,
|
||
];
|
||
|
||
/**
|
||
* Complexity-as-scope-justification patterns.
|
||
* Matches "too complex to implement", "challenging feature", etc.
|
||
* Does NOT match legitimate uses like:
|
||
* - "complex domains" in research/discovery context (describing what to research)
|
||
* - "non-trivial" in verification context (confirming substantive code exists)
|
||
* - "challenging" in user-profiling context (quoting user reactions)
|
||
*/
|
||
const COMPLEXITY_SCOPE_PATTERNS = [
|
||
// "too complex to" — always a scope-reduction justification
|
||
/too\s+complex\s+to/i,
|
||
// "too difficult" — always a scope-reduction justification
|
||
/too\s+difficult/i,
|
||
// "is too complex for" — scope justification (e.g. "Phase X is too complex for")
|
||
/is\s+too\s+complex\s+for/i,
|
||
];
|
||
|
||
/**
|
||
* Files allowed to contain certain patterns because they document
|
||
* the prohibition itself, or use the terms in non-scope-reduction context.
|
||
*/
|
||
const ALLOWLIST = {
|
||
// Plan-checker scans FOR these patterns — it's a detection list, not usage
|
||
'gsd-plan-checker.md': ['complexity_scope', 'time_sizing'],
|
||
// Planner defines the prohibition and the authority limits — uses terms to explain what NOT to do
|
||
'gsd-planner.md': ['complexity_scope'],
|
||
// Debugger uses "30+ minutes" as anti-pattern detection, not task sizing
|
||
'gsd-debugger.md': ['time_sizing'],
|
||
// Doc-writer uses "15 minutes" in API rate limit example, "2 minutes" for doc quality
|
||
'gsd-doc-writer.md': ['time_sizing'],
|
||
// Explore uses "~30 seconds" as operational estimate
|
||
'explore.md': ['time_sizing'],
|
||
// Review uses "up to 5 minutes" for CodeRabbit timeout
|
||
'review.md': ['time_sizing'],
|
||
// Fast uses "under 2 minutes wall time" as operational constraint
|
||
'fast.md': ['time_sizing'],
|
||
// Execute-phase uses a configurable test-gate timeout (workflow.test_gate_timeout, #1857)
|
||
'execute-phase.md': ['time_sizing'],
|
||
// Map-codebase documents subagent_timeout
|
||
'map-codebase.md': ['time_sizing'],
|
||
// Help documents CodeRabbit timing
|
||
'help.md': ['time_sizing'],
|
||
};
|
||
|
||
function isAllowlisted(fileName, category) {
|
||
const entry = ALLOWLIST[fileName];
|
||
return entry && entry.includes(category);
|
||
}
|
||
|
||
// -- Tests --
|
||
|
||
describe('Planner language regression — time-based task sizing (#2092)', () => {
|
||
for (const file of PLANNING_FILES) {
|
||
test(`${file.name} must not use time-based task sizing`, () => {
|
||
if (isAllowlisted(file.name, 'time_sizing')) return;
|
||
|
||
const content = fs.readFileSync(file.path, 'utf-8');
|
||
for (const pattern of TIME_SIZING_PATTERNS) {
|
||
const match = content.match(pattern);
|
||
assert.ok(
|
||
!match,
|
||
[
|
||
`${file.name} contains time-based task sizing: "${match?.[0]}"`,
|
||
'Task sizing must use context-window percentage, not time units.',
|
||
'See issue #2092 for rationale.',
|
||
].join('\n')
|
||
);
|
||
}
|
||
});
|
||
}
|
||
});
|
||
|
||
describe('Planner language regression — complexity-as-scope-justification (#2092)', () => {
|
||
for (const file of PLANNING_FILES) {
|
||
test(`${file.name} must not use complexity to justify scope reduction`, () => {
|
||
if (isAllowlisted(file.name, 'complexity_scope')) return;
|
||
|
||
const content = fs.readFileSync(file.path, 'utf-8');
|
||
for (const pattern of COMPLEXITY_SCOPE_PATTERNS) {
|
||
const match = content.match(pattern);
|
||
assert.ok(
|
||
!match,
|
||
[
|
||
`${file.name} contains complexity-as-scope-justification: "${match?.[0]}"`,
|
||
'Scope decisions must be based on context cost, missing information,',
|
||
'or dependency conflicts — not perceived difficulty.',
|
||
'See issue #2092 for rationale.',
|
||
].join('\n')
|
||
);
|
||
}
|
||
});
|
||
}
|
||
});
|
||
|
||
describe('gsd-planner.md — required structural sections (#2091, #2092)', () => {
|
||
let plannerContent;
|
||
|
||
test('planner file exists and is readable', () => {
|
||
const plannerPath = path.join(AGENTS_DIR, 'gsd-planner.md');
|
||
assert.ok(fs.existsSync(plannerPath), 'agents/gsd-planner.md must exist');
|
||
plannerContent = fs.readFileSync(plannerPath, 'utf-8');
|
||
});
|
||
|
||
test('contains <planner_authority_limits> section', () => {
|
||
assert.ok(
|
||
plannerContent.includes('<planner_authority_limits>'),
|
||
'gsd-planner.md must contain a <planner_authority_limits> section defining what the planner cannot decide'
|
||
);
|
||
});
|
||
|
||
test('authority limits prohibit difficulty-based scope decisions', () => {
|
||
assert.ok(
|
||
plannerContent.includes('The planner has no authority to'),
|
||
'planner_authority_limits must explicitly state what the planner cannot decide'
|
||
);
|
||
});
|
||
|
||
test('authority limits list three legitimate split reasons: context cost, missing info, dependency', () => {
|
||
assert.ok(
|
||
plannerContent.includes('Context cost') || plannerContent.includes('context cost'),
|
||
'authority limits must list context cost as a legitimate split reason'
|
||
);
|
||
assert.ok(
|
||
plannerContent.includes('Missing information') || plannerContent.includes('missing information'),
|
||
'authority limits must list missing information as a legitimate split reason'
|
||
);
|
||
assert.ok(
|
||
plannerContent.includes('Dependency conflict') || plannerContent.includes('dependency conflict'),
|
||
'authority limits must list dependency conflict as a legitimate split reason'
|
||
);
|
||
});
|
||
|
||
test('task sizing uses context percentage, not time units', () => {
|
||
assert.ok(
|
||
plannerContent.includes('context consumption') || plannerContent.includes('context cost'),
|
||
'task sizing must reference context consumption, not time'
|
||
);
|
||
assert.ok(
|
||
!(/each task[:\s]*\*?\*?\d+[-–]\d+\s*min/i.test(plannerContent)),
|
||
'task sizing must not use minutes as sizing unit'
|
||
);
|
||
});
|
||
|
||
test('contains multi-source coverage audit (not just D-XX decisions)', () => {
|
||
assert.ok(
|
||
plannerContent.includes('Multi-Source Coverage Audit') ||
|
||
plannerContent.includes('multi-source coverage audit'),
|
||
'gsd-planner.md must contain a multi-source coverage audit, not just D-XX decision matrix'
|
||
);
|
||
});
|
||
|
||
test('coverage audit includes all four source types: GOAL, REQ, RESEARCH, CONTEXT', () => {
|
||
// The planner file or its referenced planner-source-audit.md must define all four types.
|
||
// The inline compact version uses **GOAL**, **REQ**, **RESEARCH**, **CONTEXT**.
|
||
const refPath = path.join(ROOT, 'gsd-core', 'references', 'planner-source-audit.md');
|
||
const combined = plannerContent + (fs.existsSync(refPath) ? fs.readFileSync(refPath, 'utf-8') : '');
|
||
|
||
const hasGoal = combined.includes('**GOAL**');
|
||
const hasReq = combined.includes('**REQ**');
|
||
const hasResearch = combined.includes('**RESEARCH**');
|
||
const hasContext = combined.includes('**CONTEXT**');
|
||
|
||
assert.ok(hasGoal, 'coverage audit must include GOAL source type (ROADMAP.md phase goal)');
|
||
assert.ok(hasReq, 'coverage audit must include REQ source type (REQUIREMENTS.md)');
|
||
assert.ok(hasResearch, 'coverage audit must include RESEARCH source type (RESEARCH.md)');
|
||
assert.ok(hasContext, 'coverage audit must include CONTEXT source type (CONTEXT.md decisions)');
|
||
});
|
||
|
||
test('coverage audit defines MISSING item handling with developer escalation', () => {
|
||
assert.ok(
|
||
plannerContent.includes('Source Audit: Unplanned Items Found') ||
|
||
plannerContent.includes('MISSING'),
|
||
'coverage audit must define handling for MISSING items'
|
||
);
|
||
assert.ok(
|
||
plannerContent.includes('Awaiting developer decision') ||
|
||
plannerContent.includes('developer confirmation'),
|
||
'MISSING items must escalate to developer, not be silently dropped'
|
||
);
|
||
});
|
||
});
|
||
|
||
describe('plan-phase.md — source audit orchestration (#2091)', () => {
|
||
let workflowContent;
|
||
|
||
test('plan-phase workflow exists and is readable', () => {
|
||
const workflowPath = path.join(WORKFLOWS_DIR, 'plan-phase.md');
|
||
assert.ok(fs.existsSync(workflowPath), 'workflows/plan-phase.md must exist');
|
||
workflowContent = fs.readFileSync(workflowPath, 'utf-8');
|
||
});
|
||
|
||
test('step 9 handles Source Audit return from planner', () => {
|
||
assert.ok(
|
||
workflowContent.includes('Source Audit: Unplanned Items Found'),
|
||
'plan-phase.md step 9 must handle the Source Audit return from the planner'
|
||
);
|
||
});
|
||
|
||
test('step 9c exists for source audit gap handling', () => {
|
||
assert.ok(
|
||
workflowContent.includes('9c') && workflowContent.includes('Source Audit'),
|
||
'plan-phase.md must have a step 9c for handling source audit gaps'
|
||
);
|
||
});
|
||
|
||
test('step 9b does not use "too complex" language', () => {
|
||
// Extract just step 9b content (between "## 9b" and "## 9c" or "## 10")
|
||
const step9bSection = collectSection(workflowContent, (h) => h.text.startsWith('9b.'));
|
||
if (step9bSection) {
|
||
const step9b = step9bSection.body;
|
||
assert.ok(
|
||
!step9b.includes('too complex'),
|
||
'step 9b must not use "too complex" — use context budget language instead'
|
||
);
|
||
}
|
||
});
|
||
|
||
test('phase split recommendation uses context budget framing', () => {
|
||
assert.ok(
|
||
workflowContent.includes('context budget') || workflowContent.includes('context cost'),
|
||
'phase split recommendation must be framed in terms of context budget, not complexity'
|
||
);
|
||
});
|
||
});
|
||
|
||
describe('gsd-plan-checker.md — scope reduction detection includes time/complexity (#2092)', () => {
|
||
let checkerContent;
|
||
|
||
test('plan-checker exists and is readable', () => {
|
||
const checkerPath = path.join(AGENTS_DIR, 'gsd-plan-checker.md');
|
||
assert.ok(fs.existsSync(checkerPath), 'agents/gsd-plan-checker.md must exist');
|
||
checkerContent = fs.readFileSync(checkerPath, 'utf-8');
|
||
});
|
||
|
||
test('scope reduction scan includes complexity-based justification patterns', () => {
|
||
assert.ok(
|
||
checkerContent.includes('too complex') || checkerContent.includes('too difficult'),
|
||
'plan-checker scope reduction scan must detect complexity-based justification language'
|
||
);
|
||
});
|
||
|
||
test('scope reduction scan includes time-based justification patterns', () => {
|
||
assert.ok(
|
||
checkerContent.includes('would take') || checkerContent.includes('hours') || checkerContent.includes('minutes'),
|
||
'plan-checker scope reduction scan must detect time-based justification language'
|
||
);
|
||
});
|
||
});
|
||
|
||
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
// Folded from tests/bug-3805-fast-md-log-to-state-schema.test.cjs — consolidation epic #1969 (B4 #1973)
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
{
|
||
const { describe: __foldDescribe } = require('node:test');
|
||
__foldDescribe("folded:bug-3805-fast-md-log-to-state-schema (consolidation epic #1969 B4 #1973)", () => {
|
||
'use strict';
|
||
|
||
// allow-test-rule: source-text-is-the-product (see #3805)
|
||
// Reads gsd-core/workflows/fast.md whose deployed text IS the product —
|
||
// the workflow markdown is executed verbatim by LLM runtimes.
|
||
|
||
/**
|
||
* #3805 — fast.md log_to_state appends a schema-blind 4-column row to the
|
||
* 5-column "Quick Tasks Completed" table created by quick.md Step 7.
|
||
*
|
||
* quick.md Step 7 creates the table with 5 columns:
|
||
* | # | Description | Date | Commit | Directory |
|
||
*
|
||
* Before this fix, fast.md's log_to_state step appended a hardcoded 4-cell
|
||
* row unconditionally:
|
||
* echo "| $(date +%Y-%m-%d) | fast | $TASK | ✅ |" >> .planning/STATE.md
|
||
*
|
||
* This produces malformed Markdown when the existing table has a different
|
||
* column count.
|
||
*
|
||
* Covers:
|
||
* - fast.md does NOT contain the hardcoded 4-cell echo template
|
||
* - fast.md log_to_state step reads/introspects the existing table header
|
||
* before appending (schema-aware insertion)
|
||
* - fast.md log_to_state step matches the 5-column schema from quick.md
|
||
* Step 7 when that table is present
|
||
* - fast.md log_to_state step skips (does not corrupt) the STATE.md write
|
||
* when the table schema is unrecognized
|
||
*/
|
||
|
||
const { test, describe } = require('node:test');
|
||
const assert = require('node:assert/strict');
|
||
const fs = require('node:fs');
|
||
const path = require('node:path');
|
||
|
||
const REPO_ROOT = path.join(__dirname, '..');
|
||
const FAST_MD_PATH = path.join(REPO_ROOT, 'gsd-core', 'workflows', 'fast.md');
|
||
|
||
// The 5-column schema defined in quick.md Step 7 (non-validate mode).
|
||
// Column count is 5: # | Description | Date | Commit | Directory
|
||
// Named constant for traceability — mirrors quick.md Step 7's table header.
|
||
const QUICK_MD_STEP7_COL_COUNT = 5;
|
||
const QUICK_MD_STEP7_COLUMNS = ['#', 'Description', 'Date', 'Commit', 'Directory'];
|
||
|
||
describe('bug #3805: fast.md log_to_state must be schema-aware', () => {
|
||
let fastMdContent;
|
||
|
||
test('fast.md workflow file exists and is readable', () => {
|
||
assert.ok(fs.existsSync(FAST_MD_PATH), `fast.md not found at ${FAST_MD_PATH}`);
|
||
fastMdContent = fs.readFileSync(FAST_MD_PATH, 'utf-8');
|
||
});
|
||
|
||
test('fast.md log_to_state step does NOT hardcode a 4-cell row template', () => {
|
||
// The old broken template: | date | fast | task | ✅ |
|
||
// This regex matches the exact hardcoded pattern that ignores table schema.
|
||
// A 4-cell row has exactly 4 pipe-delimited fields plus the surrounding pipes.
|
||
const hardcoded4CellPattern = /echo\s+["'][|][^|]*[|][^|]*[|][^|]*[|][^|]*[|]\s*["']/;
|
||
const match = fastMdContent.match(hardcoded4CellPattern);
|
||
assert.ok(
|
||
!match,
|
||
[
|
||
'fast.md log_to_state still contains the hardcoded 4-cell row template:',
|
||
` "${match?.[0]}"`,
|
||
'This appends a malformed row to the 5-column Quick Tasks Completed table',
|
||
'created by quick.md Step 7.',
|
||
`Expected table schema (${QUICK_MD_STEP7_COL_COUNT} cols): | ${QUICK_MD_STEP7_COLUMNS.join(' | ')} |`,
|
||
].join('\n')
|
||
);
|
||
});
|
||
|
||
test('fast.md log_to_state step reads the existing table header (schema introspection)', () => {
|
||
// The fix must inspect the existing STATE.md table header before appending.
|
||
// Acceptable signals: reading STATE.md content, grepping for the header line,
|
||
// or parsing columns from the header row.
|
||
const hasHeaderRead =
|
||
// Reads STATE.md to inspect it (awk/sed/grep on the file for header detection)
|
||
/grep.*Quick Tasks Completed.*STATE\.md/.test(fastMdContent) ||
|
||
/awk.*Quick Tasks Completed/.test(fastMdContent) ||
|
||
/sed.*Quick Tasks Completed/.test(fastMdContent) ||
|
||
// Reads the header line explicitly (head -n, sed -n, awk NR==)
|
||
/head\s+-n/.test(fastMdContent) && /STATE\.md/.test(fastMdContent) ||
|
||
// Counts pipe separators / columns from existing header
|
||
/col.*count|column.*count|count.*col|NF|awk.*\|/.test(fastMdContent) ||
|
||
// References schema detection in prose
|
||
/schema|header|column\s+count|existing.*table/.test(fastMdContent);
|
||
|
||
assert.ok(
|
||
hasHeaderRead,
|
||
[
|
||
'fast.md log_to_state step does not appear to introspect the existing table schema.',
|
||
'The step must read the STATE.md table header to detect column count before appending.',
|
||
'quick.md Step 7 uses schema-aware matching — fast.md must follow the same discipline.',
|
||
].join('\n')
|
||
);
|
||
});
|
||
|
||
test('fast.md log_to_state step references the 5-column quick.md schema', () => {
|
||
// Schema-awareness no longer lives inline in fast.md as hardcoded column
|
||
// names — it was moved to the schema-backed `quick-tasks-append` helper
|
||
// (`appendQuickTaskRow` in markdown-table.cjs; #2133, ADR-2143 §3/§7),
|
||
// which introspects the existing table's schema (5-col or 6-col) itself.
|
||
// Assert the step delegates to that helper instead of requiring the old
|
||
// literal column names.
|
||
const logToStateMatch = fastMdContent.match(/<step name="log_to_state">([\s\S]*?)<\/step>/);
|
||
assert.ok(logToStateMatch, 'fast.md must contain a <step name="log_to_state"> element');
|
||
|
||
const stepContent = logToStateMatch[1];
|
||
|
||
assert.ok(
|
||
stepContent.includes('quick-tasks-append'),
|
||
[
|
||
'fast.md log_to_state step does not invoke the schema-aware quick-tasks-append helper.',
|
||
'Schema-awareness now lives in the gsd-tools quick-tasks-append subcommand',
|
||
`(appendQuickTaskRow, handling both the ${QUICK_MD_STEP7_COL_COUNT}-column quick.md Step 7 schema`,
|
||
`${QUICK_MD_STEP7_COLUMNS.join(', ')} and the 6-column with-status variant) —`,
|
||
'the step must delegate to it rather than hardcoding column names inline.',
|
||
].join('\n')
|
||
);
|
||
});
|
||
|
||
test('fast.md log_to_state step skips STATE.md write on unrecognized schema', () => {
|
||
// The fix must not blindly append when the table schema is unknown.
|
||
// Check for a guard that skips or logs rather than corrupting the file.
|
||
const logToStateMatch = fastMdContent.match(/<step name="log_to_state">([\s\S]*?)<\/step>/);
|
||
assert.ok(logToStateMatch, 'fast.md must contain a <step name="log_to_state"> element');
|
||
|
||
const stepContent = logToStateMatch[1];
|
||
|
||
// Must have a skip/guard path for unrecognized schemas
|
||
const hasSkipGuard =
|
||
/skip/i.test(stepContent) ||
|
||
/unrecognized|unknown|mismatch/i.test(stepContent) ||
|
||
/else\b/.test(stepContent) ||
|
||
/warn|log/i.test(stepContent);
|
||
|
||
assert.ok(
|
||
hasSkipGuard,
|
||
[
|
||
'fast.md log_to_state step does not appear to guard against unrecognized table schemas.',
|
||
'When the existing STATE.md table does not match an expected schema,',
|
||
'the step must skip the write (with a brief log) rather than append a malformed row.',
|
||
].join('\n')
|
||
);
|
||
});
|
||
});
|
||
});
|
||
}
|
||
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
// Folded from tests/bug-2421-planner-grep-gate-hygiene.test.cjs — consolidation epic #1969 (B7 #1976)
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
{
|
||
const { describe: __foldDescribe } = require('node:test');
|
||
__foldDescribe("folded:bug-2421-planner-grep-gate-hygiene (consolidation epic #1969 B7 #1976)", () => {
|
||
// allow-test-rule: source-text-is-the-product (see #2421)
|
||
// Workflow .md / agent .md / command .md / reference .md files — their text
|
||
// IS what the runtime loads. Testing text content tests the deployed contract.
|
||
// Per CONTRIBUTING.md exception matrix.
|
||
|
||
/**
|
||
* Bug #2421: gsd-planner emits grep-count acceptance gates that count comment text
|
||
*
|
||
* The planner must instruct agents to use comment-aware grep patterns in
|
||
* <automated> verify blocks. Without this, descriptive comments in file
|
||
* headers count against the gate and force authors to reword them — the
|
||
* "self-invalidating grep gate" anti-pattern.
|
||
*/
|
||
|
||
const { describe, test } = require('node:test');
|
||
const assert = require('node:assert/strict');
|
||
const fs = require('fs');
|
||
const path = require('path');
|
||
|
||
const PLANNER_PATH = path.join(__dirname, '..', 'agents', 'gsd-planner.md');
|
||
|
||
describe('gsd-planner grep gate hygiene (#2421)', () => {
|
||
test('gsd-planner.md exists in agents source dir', () => {
|
||
assert.ok(fs.existsSync(PLANNER_PATH), 'agents/gsd-planner.md must exist');
|
||
});
|
||
|
||
test('gsd-planner.md contains Grep gate hygiene rule', () => {
|
||
const content = fs.readFileSync(PLANNER_PATH, 'utf-8');
|
||
assert.ok(
|
||
content.includes('Grep gate hygiene') || content.includes('grep gate hygiene'),
|
||
'gsd-planner.md must contain a "Grep gate hygiene" rule to prevent self-invalidating grep gates'
|
||
);
|
||
});
|
||
|
||
test('gsd-planner.md explains self-invalidating grep gate anti-pattern', () => {
|
||
const content = fs.readFileSync(PLANNER_PATH, 'utf-8');
|
||
assert.ok(
|
||
content.includes('self-invalidating'),
|
||
'gsd-planner.md must describe the "self-invalidating" grep gate anti-pattern'
|
||
);
|
||
});
|
||
|
||
test('gsd-planner.md provides comment-stripping grep example', () => {
|
||
const content = fs.readFileSync(PLANNER_PATH, 'utf-8');
|
||
// Must show a pattern that excludes comment lines (grep -v or grep -vE)
|
||
assert.ok(
|
||
content.includes('grep -v') || content.includes('grep -vE') || content.includes('-v '),
|
||
'gsd-planner.md must provide a comment-stripping grep example (grep -v or grep -vE)'
|
||
);
|
||
});
|
||
|
||
test('gsd-planner.md warns against bare zero-count grep gates on whole files', () => {
|
||
const content = fs.readFileSync(PLANNER_PATH, 'utf-8');
|
||
assert.ok(
|
||
content.includes('== 0') || content.includes('zero-count') || content.includes('zero count'),
|
||
'gsd-planner.md must warn against bare zero-count grep gates without comment exclusion'
|
||
);
|
||
});
|
||
|
||
test('gsd-planner.md grep gate hygiene rule appears after Nyquist Rule', () => {
|
||
const content = fs.readFileSync(PLANNER_PATH, 'utf-8');
|
||
const nyquistIdx = content.indexOf('Nyquist Rule');
|
||
const grepGateIdx = content.indexOf('grep gate hygiene') !== -1
|
||
? content.indexOf('grep gate hygiene')
|
||
: content.indexOf('Grep gate hygiene');
|
||
|
||
assert.ok(nyquistIdx !== -1, 'Nyquist Rule must be present in gsd-planner.md');
|
||
assert.ok(grepGateIdx !== -1, 'Grep gate hygiene must be present in gsd-planner.md');
|
||
assert.ok(
|
||
grepGateIdx > nyquistIdx,
|
||
`Grep gate hygiene rule (at ${grepGateIdx}) must appear after Nyquist Rule (at ${nyquistIdx})`
|
||
);
|
||
});
|
||
});
|
||
});
|
||
}
|
||
|
||
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
// Folded from tests/bug-3087-planner-directive-language.test.cjs — consolidation epic #1969 (B7 #1976)
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
{
|
||
const { describe: __foldDescribe } = require('node:test');
|
||
__foldDescribe("folded:bug-3087-planner-directive-language (consolidation epic #1969 B7 #1976)", () => {
|
||
'use strict';
|
||
|
||
// Regression guard for bug #3087.
|
||
//
|
||
// Between v1.38.3 and v1.38.4, agents/gsd-planner.md had 10 instances of
|
||
// CRITICAL/MANDATORY/ALWAYS/MUST directive emphasis systematically removed.
|
||
// The change was undocumented and conflicts with the stated intent of PR #2489
|
||
// (the sycophancy-hardening pass that shipped in the same release). This test
|
||
// enforces the restored directive language so the demotion cannot recur silently.
|
||
|
||
const { test } = require('node:test');
|
||
const assert = require('node:assert/strict');
|
||
const fs = require('node:fs');
|
||
const path = require('node:path');
|
||
|
||
const ROOT = path.join(__dirname, '..');
|
||
let src;
|
||
try {
|
||
src = fs.readFileSync(path.join(ROOT, 'agents', 'gsd-planner.md'), 'utf8');
|
||
} catch (err) {
|
||
throw new Error(`agents/gsd-planner.md not found — was the file renamed? (${err.message})`);
|
||
}
|
||
|
||
const directives = [
|
||
{ desc: 'User Decision Fidelity heading is CRITICAL', pattern: /## CRITICAL: User Decision Fidelity/ },
|
||
{ desc: 'Never Simplify heading is CRITICAL', pattern: /## CRITICAL: Never Simplify User Decisions/ },
|
||
{ desc: 'Multi-Source Audit heading is MANDATORY', pattern: /## Multi-Source Coverage Audit \(MANDATORY in every plan set\)/ },
|
||
{ desc: 'Source audit uses "Audit ALL" imperative', pattern: /Audit ALL four source types before finalizing/ },
|
||
{ desc: 'Discovery is MANDATORY', pattern: /Discovery is MANDATORY unless/ },
|
||
{ desc: 'Split signals use ALWAYS', pattern: /\*\*ALWAYS split if:\*\*/ },
|
||
{ desc: 'requirements field doc uses MUST', pattern: /\*\*MUST\*\* list requirement IDs from ROADMAP/ },
|
||
{ desc: 'Step 0 has CRITICAL requirement ID directive', pattern: /\*\*CRITICAL:\*\* Every requirement ID MUST appear/ },
|
||
{ desc: 'Write tool directive uses ALWAYS', pattern: /\*\*ALWAYS use the Write tool to create files\*\*/ },
|
||
{ desc: 'File naming convention heading is CRITICAL', pattern: /\*\*CRITICAL — File naming convention \(enforced\):\*\*/ },
|
||
];
|
||
|
||
for (const { desc, pattern } of directives) {
|
||
test(`gsd-planner.md: ${desc}`, () => {
|
||
assert.ok(
|
||
pattern.test(src),
|
||
`Directive enforcement missing from gsd-planner.md: "${desc}" — pattern ${pattern} not found. ` +
|
||
`This language was demoted in v1.38.4 (PR #2489) without documentation, conflicting with ` +
|
||
`the sycophancy-hardening intent of that release. See bug #3087.`,
|
||
);
|
||
});
|
||
}
|
||
});
|
||
}
|
||
|
||
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
// Folded from tests/bug-3430-planner-phase-contract.test.cjs — consolidation epic #1969 (B7 #1976)
|
||
// ────────────────────────────────────────────────────────────────────────
|
||
{
|
||
const { describe: __foldDescribe } = require('node:test');
|
||
__foldDescribe("folded:bug-3430-planner-phase-contract (consolidation epic #1969 B7 #1976)", () => {
|
||
// allow-test-rule: source-text-is-the-product (see #3430)
|
||
// Planner markdown is the deployed planning contract; these checks lock the
|
||
// exact canonical forms that downstream phase-plan-index accepts.
|
||
|
||
'use strict';
|
||
|
||
const { test } = require('node:test');
|
||
const assert = require('node:assert/strict');
|
||
const fs = require('node:fs');
|
||
const path = require('node:path');
|
||
|
||
const PLANNER_PATH = path.join(__dirname, '..', 'agents', 'gsd-planner.md');
|
||
|
||
function readPlanner() {
|
||
return fs.readFileSync(PLANNER_PATH, 'utf8');
|
||
}
|
||
|
||
test('#3430: planner SUMMARY instruction uses canonical padded phase/plan form', () => {
|
||
const content = readPlanner();
|
||
assert.match(
|
||
content,
|
||
/Create `\.planning\/phases\/XX-name\/\{padded_phase\}-\{plan\}-SUMMARY\.md` when done/,
|
||
'planner must instruct executors to write SUMMARY files in canonical padded-phase form'
|
||
);
|
||
assert.doesNotMatch(
|
||
content,
|
||
/After completion, create `\.planning\/phases\/XX-name\/\{phase\}-\{plan\}-SUMMARY\.md`/,
|
||
'planner must not instruct the broken {phase}-{plan}-SUMMARY.md form'
|
||
);
|
||
});
|
||
|
||
test('#3430: planner depends_on docs show canonical in-phase plan ids', () => {
|
||
const content = readPlanner();
|
||
assert.match(
|
||
content,
|
||
/depends_on:[^\n]*Use `01-01`\/`01-01-auth-hardening`/,
|
||
'planner must document canonical depends_on examples that phase-plan-index resolves'
|
||
);
|
||
assert.doesNotMatch(
|
||
content,
|
||
/depends_on:[^\n]*01-trust\/01/,
|
||
'planner must not document phase-slug/plan-number depends_on examples as canonical'
|
||
);
|
||
});
|
||
});
|
||
}
|