Files
msd-core/tests/planner-language-regression.test.cjs
Tom Boucher 463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00

339 lines
13 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
// allow-test-rule: source-text-is-the-product
// Workflow .md / agent .md / command .md / reference .md files — their text
// IS what the runtime loads. Testing text content tests the deployed contract.
// Per CONTRIBUTING.md exception matrix.
'use strict';
/**
* Planner Language Regression Tests (#2091, #2092)
*
* Prevents time-based reasoning and complexity-as-scope-justification
* from leaking back into planning artifacts via future PRs.
*
* These tests scan agent definitions, workflow files, and references
* for prohibited patterns that import human-world constraints into
* an AI execution context where those constraints do not exist.
*/
const { test, describe } = require('node:test');
const assert = require('node:assert/strict');
const fs = require('fs');
const path = require('path');
const ROOT = path.join(__dirname, '..');
const AGENTS_DIR = path.join(ROOT, 'agents');
const WORKFLOWS_DIR = path.join(ROOT, 'gsd-core', 'workflows');
const REFERENCES_DIR = path.join(ROOT, 'gsd-core', 'references');
const TEMPLATES_DIR = path.join(ROOT, 'gsd-core', 'templates');
/**
* Collect all .md files from a directory (non-recursive).
*/
function mdFiles(dir) {
if (!fs.existsSync(dir)) return [];
return fs.readdirSync(dir)
.filter(f => f.endsWith('.md'))
.map(f => ({ name: f, path: path.join(dir, f) }));
}
/**
* Collect all .md files recursively.
*/
function mdFilesRecursive(dir) {
if (!fs.existsSync(dir)) return [];
const results = [];
for (const entry of fs.readdirSync(dir, { withFileTypes: true })) {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) {
results.push(...mdFilesRecursive(full));
} else if (entry.name.endsWith('.md')) {
results.push({ name: entry.name, path: full });
}
}
return results;
}
/**
* Files that define planning behavior — agents, workflows, references.
* These are the files where time-based and complexity-based scope
* reasoning must never appear.
*/
const PLANNING_FILES = [
...mdFiles(AGENTS_DIR),
...mdFiles(WORKFLOWS_DIR),
...mdFiles(REFERENCES_DIR),
...mdFilesRecursive(TEMPLATES_DIR),
];
// -- Prohibited patterns --
/**
* Time-based task sizing patterns.
* Matches "15-60 minutes", "X minutes Claude execution time", etc.
* Does NOT match operational timeouts ("timeout: 5 minutes"),
* API docs examples ("100 requests per 15 minutes"),
* or human-readable timeout descriptions in workflow execution steps.
*/
const TIME_SIZING_PATTERNS = [
// "N-M minutes" in task sizing context (not timeout context)
/each task[:\s]*\*?\*?\d+[-–]\d+\s*min/i,
// "minutes Claude execution time" or "minutes execution time"
/minutes?\s+(claude\s+)?execution\s+time/i,
// Duration-based sizing table rows: "< 15 min", "15-60 min", "> 60 min"
/[<>]\s*\d+\s*min\s*\|/i,
];
/**
* Complexity-as-scope-justification patterns.
* Matches "too complex to implement", "challenging feature", etc.
* Does NOT match legitimate uses like:
* - "complex domains" in research/discovery context (describing what to research)
* - "non-trivial" in verification context (confirming substantive code exists)
* - "challenging" in user-profiling context (quoting user reactions)
*/
const COMPLEXITY_SCOPE_PATTERNS = [
// "too complex to" — always a scope-reduction justification
/too\s+complex\s+to/i,
// "too difficult" — always a scope-reduction justification
/too\s+difficult/i,
// "is too complex for" — scope justification (e.g. "Phase X is too complex for")
/is\s+too\s+complex\s+for/i,
];
/**
* Files allowed to contain certain patterns because they document
* the prohibition itself, or use the terms in non-scope-reduction context.
*/
const ALLOWLIST = {
// Plan-checker scans FOR these patterns — it's a detection list, not usage
'gsd-plan-checker.md': ['complexity_scope', 'time_sizing'],
// Planner defines the prohibition and the authority limits — uses terms to explain what NOT to do
'gsd-planner.md': ['complexity_scope'],
// Debugger uses "30+ minutes" as anti-pattern detection, not task sizing
'gsd-debugger.md': ['time_sizing'],
// Doc-writer uses "15 minutes" in API rate limit example, "2 minutes" for doc quality
'gsd-doc-writer.md': ['time_sizing'],
// Discovery-phase uses time for level descriptions (operational, not scope)
'discovery-phase.md': ['time_sizing'],
// Explore uses "~30 seconds" as operational estimate
'explore.md': ['time_sizing'],
// Review uses "up to 5 minutes" for CodeRabbit timeout
'review.md': ['time_sizing'],
// Fast uses "under 2 minutes wall time" as operational constraint
'fast.md': ['time_sizing'],
// Execute-phase uses "timeout: 5 minutes" for test runner
'execute-phase.md': ['time_sizing'],
// Verify-phase uses "timeout: 5 minutes" for test runner
'verify-phase.md': ['time_sizing'],
// Map-codebase documents subagent_timeout
'map-codebase.md': ['time_sizing'],
// Help documents CodeRabbit timing
'help.md': ['time_sizing'],
};
function isAllowlisted(fileName, category) {
const entry = ALLOWLIST[fileName];
return entry && entry.includes(category);
}
// -- Tests --
describe('Planner language regression — time-based task sizing (#2092)', () => {
for (const file of PLANNING_FILES) {
test(`${file.name} must not use time-based task sizing`, () => {
if (isAllowlisted(file.name, 'time_sizing')) return;
const content = fs.readFileSync(file.path, 'utf-8');
for (const pattern of TIME_SIZING_PATTERNS) {
const match = content.match(pattern);
assert.ok(
!match,
[
`${file.name} contains time-based task sizing: "${match?.[0]}"`,
'Task sizing must use context-window percentage, not time units.',
'See issue #2092 for rationale.',
].join('\n')
);
}
});
}
});
describe('Planner language regression — complexity-as-scope-justification (#2092)', () => {
for (const file of PLANNING_FILES) {
test(`${file.name} must not use complexity to justify scope reduction`, () => {
if (isAllowlisted(file.name, 'complexity_scope')) return;
const content = fs.readFileSync(file.path, 'utf-8');
for (const pattern of COMPLEXITY_SCOPE_PATTERNS) {
const match = content.match(pattern);
assert.ok(
!match,
[
`${file.name} contains complexity-as-scope-justification: "${match?.[0]}"`,
'Scope decisions must be based on context cost, missing information,',
'or dependency conflicts — not perceived difficulty.',
'See issue #2092 for rationale.',
].join('\n')
);
}
});
}
});
describe('gsd-planner.md — required structural sections (#2091, #2092)', () => {
let plannerContent;
test('planner file exists and is readable', () => {
const plannerPath = path.join(AGENTS_DIR, 'gsd-planner.md');
assert.ok(fs.existsSync(plannerPath), 'agents/gsd-planner.md must exist');
plannerContent = fs.readFileSync(plannerPath, 'utf-8');
});
test('contains <planner_authority_limits> section', () => {
assert.ok(
plannerContent.includes('<planner_authority_limits>'),
'gsd-planner.md must contain a <planner_authority_limits> section defining what the planner cannot decide'
);
});
test('authority limits prohibit difficulty-based scope decisions', () => {
assert.ok(
plannerContent.includes('The planner has no authority to'),
'planner_authority_limits must explicitly state what the planner cannot decide'
);
});
test('authority limits list three legitimate split reasons: context cost, missing info, dependency', () => {
assert.ok(
plannerContent.includes('Context cost') || plannerContent.includes('context cost'),
'authority limits must list context cost as a legitimate split reason'
);
assert.ok(
plannerContent.includes('Missing information') || plannerContent.includes('missing information'),
'authority limits must list missing information as a legitimate split reason'
);
assert.ok(
plannerContent.includes('Dependency conflict') || plannerContent.includes('dependency conflict'),
'authority limits must list dependency conflict as a legitimate split reason'
);
});
test('task sizing uses context percentage, not time units', () => {
assert.ok(
plannerContent.includes('context consumption') || plannerContent.includes('context cost'),
'task sizing must reference context consumption, not time'
);
assert.ok(
!(/each task[:\s]*\*?\*?\d+[-–]\d+\s*min/i.test(plannerContent)),
'task sizing must not use minutes as sizing unit'
);
});
test('contains multi-source coverage audit (not just D-XX decisions)', () => {
assert.ok(
plannerContent.includes('Multi-Source Coverage Audit') ||
plannerContent.includes('multi-source coverage audit'),
'gsd-planner.md must contain a multi-source coverage audit, not just D-XX decision matrix'
);
});
test('coverage audit includes all four source types: GOAL, REQ, RESEARCH, CONTEXT', () => {
// The planner file or its referenced planner-source-audit.md must define all four types.
// The inline compact version uses **GOAL**, **REQ**, **RESEARCH**, **CONTEXT**.
const refPath = path.join(ROOT, 'gsd-core', 'references', 'planner-source-audit.md');
const combined = plannerContent + (fs.existsSync(refPath) ? fs.readFileSync(refPath, 'utf-8') : '');
const hasGoal = combined.includes('**GOAL**');
const hasReq = combined.includes('**REQ**');
const hasResearch = combined.includes('**RESEARCH**');
const hasContext = combined.includes('**CONTEXT**');
assert.ok(hasGoal, 'coverage audit must include GOAL source type (ROADMAP.md phase goal)');
assert.ok(hasReq, 'coverage audit must include REQ source type (REQUIREMENTS.md)');
assert.ok(hasResearch, 'coverage audit must include RESEARCH source type (RESEARCH.md)');
assert.ok(hasContext, 'coverage audit must include CONTEXT source type (CONTEXT.md decisions)');
});
test('coverage audit defines MISSING item handling with developer escalation', () => {
assert.ok(
plannerContent.includes('Source Audit: Unplanned Items Found') ||
plannerContent.includes('MISSING'),
'coverage audit must define handling for MISSING items'
);
assert.ok(
plannerContent.includes('Awaiting developer decision') ||
plannerContent.includes('developer confirmation'),
'MISSING items must escalate to developer, not be silently dropped'
);
});
});
describe('plan-phase.md — source audit orchestration (#2091)', () => {
let workflowContent;
test('plan-phase workflow exists and is readable', () => {
const workflowPath = path.join(WORKFLOWS_DIR, 'plan-phase.md');
assert.ok(fs.existsSync(workflowPath), 'workflows/plan-phase.md must exist');
workflowContent = fs.readFileSync(workflowPath, 'utf-8');
});
test('step 9 handles Source Audit return from planner', () => {
assert.ok(
workflowContent.includes('Source Audit: Unplanned Items Found'),
'plan-phase.md step 9 must handle the Source Audit return from the planner'
);
});
test('step 9c exists for source audit gap handling', () => {
assert.ok(
workflowContent.includes('9c') && workflowContent.includes('Source Audit'),
'plan-phase.md must have a step 9c for handling source audit gaps'
);
});
test('step 9b does not use "too complex" language', () => {
// Extract just step 9b content (between "## 9b" and "## 9c" or "## 10")
const step9bMatch = workflowContent.match(/## 9b\.([\s\S]*?)(?=## 9c|## 10)/);
if (step9bMatch) {
const step9b = step9bMatch[1];
assert.ok(
!step9b.includes('too complex'),
'step 9b must not use "too complex" — use context budget language instead'
);
}
});
test('phase split recommendation uses context budget framing', () => {
assert.ok(
workflowContent.includes('context budget') || workflowContent.includes('context cost'),
'phase split recommendation must be framed in terms of context budget, not complexity'
);
});
});
describe('gsd-plan-checker.md — scope reduction detection includes time/complexity (#2092)', () => {
let checkerContent;
test('plan-checker exists and is readable', () => {
const checkerPath = path.join(AGENTS_DIR, 'gsd-plan-checker.md');
assert.ok(fs.existsSync(checkerPath), 'agents/gsd-plan-checker.md must exist');
checkerContent = fs.readFileSync(checkerPath, 'utf-8');
});
test('scope reduction scan includes complexity-based justification patterns', () => {
assert.ok(
checkerContent.includes('too complex') || checkerContent.includes('too difficult'),
'plan-checker scope reduction scan must detect complexity-based justification language'
);
});
test('scope reduction scan includes time-based justification patterns', () => {
assert.ok(
checkerContent.includes('would take') || checkerContent.includes('hours') || checkerContent.includes('minutes'),
'plan-checker scope reduction scan must detect time-based justification language'
);
});
});