Files
msd-core/tests/prompt-budget.property.test.cjs
Tom Boucher 463cffd894 chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/

Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the
on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary
(`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers
are unaffected.

Mechanical (bulk, ~90% of the diff):
- `git mv get-shit-done gsd-core`
- Swept path/identifier references across the repo via
  `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead
  preserves the five legitimate slug variants that are NOT the directory:
  get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names).
- Build/manifest wiring: package.json (bin, files, coverage globs),
  tsconfig.build.json (outDir), ~86 .gitignore build-output entries,
  stryker.config.mjs, scan-ignore files, install.js path strings.
- Frozen (not rewritten): CHANGELOG.md history; translated docs
  (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/).

New logic (review here):
- src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper
  ADR-0008 installer migration. On upgrade it walks the legacy
  `~/.claude/get-shit-done/` tree, classifies each file via the prior install
  manifest, and emits remove-managed / backup-and-remove for managed files
  while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked
  root and symlinked entries; bounds-checks every path under configDir). The
  framework rolls back on install failure. Emptied dirs may remain (framework
  has no recursive dir-removal primitive) — documented.
- scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare
  `get-shit-done` directory token (split token to avoid self-match; case-
  insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists
  CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines).
  Wired into the lint-tests CI job.
- Restored scripts/lint-package-identity-drift.cjs detection regexes (the
  mechanical sweep had wrongly rewritten the old-name patterns it exists to
  detect) and marked them as intentional legacy references.
- TDD tests for the migration and the guard; do.md slash-command guard regex
  tightened so a `/gsd-core/bin` path segment is not mistaken for a command;
  changeset + docs/installer-migrations.md row added.

Breaking: the installed runtime path moves `~/.claude/get-shit-done/` ->
`~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed
files (preserving user files) on upgrade. Users with custom hooks/configs
hardcoding the old path must update them.

Closes #604

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unsweep pending changesets + allowlist injection-example docs

CI fixes for the rename PR:
- Do not sweep pending .changeset/*.md (ephemeral release-note fragments,
  like CHANGELOG); reverted those body edits so 5 pre-existing malformed
  fragments (missing type/pr) no longer enter the PR diff and trip docs-lint.
  Allowlisted .changeset/ in the legacy-name guard accordingly.
- Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in
  prompt-injection-scan.sh: they contain intentional injection examples /
  security-model prose; the path-reference rewrites are kept.

CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR;
none in the new migration/guard) and are out of scope for the rename.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): resolve CodeQL alerts surfaced on this PR

The rename diff touched files carrying pre-existing CodeQL findings; per the
no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving
them off. All behavior-preserving:

- scripts/ci-test-scope.cjs: build the config-path match from string
  .includes() instead of a RegExp over an arg-derived value (js/regex-injection).
- src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName
  so the table-cell escape is complete (js/incomplete-sanitization).
- tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment
  strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization).
- tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace,
  keep the meaningful POSIX-class conversion (js/identity-replacement).

Verified: build:lib green; the touched test files + ci-test-scope + profile-output
suites pass; lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization)

The prior commit's fixes for two alerts were ineffective:
- ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file`
  reaching static regex `.test(file)` calls (not the config rule). Removed ALL
  regex over file/t — startsWith/includes/=== string checks + an isWindowsHint
  helper — so there is no regex sink for the tainted value.
- js/incomplete-multi-character-sanitization (3 test files): a single
  `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint
  loop (replace until stable) plus a final bare-opener strip.

Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass;
lint:legacy-name clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL

CodeQL flags the regex PATTERNS syntactically (regex-injection on the
--files arg split; incomplete-multi-character-sanitization on the <!--...-->
replace), so loop fixes do not satisfy it. Made these paths regex-free:
- ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/).
- 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)).
Behavior preserved; ci-test-scope + the 3 suites pass; guard clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): unblock security base64 scan on the large rename diff

The security job hit its 10m timeout: base64-scan.sh choked on the binary
test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/
non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings),
and the ~800-file rename diff is slow to scan regardless.

- scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they
  can't carry base64-obfuscated *text* and feeding NUL bytes through the
  per-line scanner is pathologically slow. collect_files already filtered
  binary *extensions*; this catches binary *content* in text extensions.
- .github/workflows/security-scan.yml: raise the security job timeout 10m->30m
  to accommodate very large diffs (the scan itself is unchanged).

Verified locally: scan skips the fixture, 0 "ignored null byte" warnings,
0 findings, exit 0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): sweep get-shit-done refs introduced by merging next

The branch was updated with next (#614/#384/#618 etc.), which reference the
get-shit-done/ dir (still named that on next). Swept the stale references in
the merged files to gsd-core so the rename stays consistent and lint:legacy-name
passes:
- commands/gsd/discuss-phase.md (runtime-launcher shim paths)
- src/core.cts (getAgentsDir layout comments)
- tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib)

Verified: guard 0 violations; build green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant

The #614 runtime-launcher shim added to discuss-phase.md references
`${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y
refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it
mis-read the directory path as a dangling `/gsd-core` command ref (same class as
the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path
segments are not treated as slash-command references.

Verified locally on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22 image) full suite: 0 failures
- bug-3683 + bug-2954 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI)

CI intermittently failed state.test's gsd-tools subprocess with
"findProjectRoot is not a function" (flip-flopping across legs; not reproducible
on mac full suite, gsd-test linux full suite, test:unit, or state.test x8).
findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs);
binding it via destructure at module-load can be undefined under a load-ordering
edge. Resolve it lazily at call time via a small wrapper so the lookup happens
after core.cjs is fully initialized.

Verified green on BOTH platforms before pushing:
- mac (node 26) full suite: 0 failures
- gsd-test-runner (linux, node22) full suite: 0 failures
- state.test.cjs: 106/106; gsd-tools loads cleanly.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#604): allowlist verification-patterns.md placeholder examples in secret scan

The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling
it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var
examples (illustrative Stripe test-key / database-URL / API-key placeholders) —
not real credentials. Added it to .secretscanignore with the strict annotation,
mirroring the existing gsd-core/workflows/plan-phase.md exception.

Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next
exits 0 with 0 findings.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 18:35:29 -04:00

243 lines
9.4 KiB
JavaScript

'use strict';
/**
* Property-based tests for prompt-budget.cjs
*
* Module: gsd-core/bin/lib/prompt-budget.cjs
* Exported: estimateTokens(text), applyBudget({ sections, budget, options })
*
* Key invariants:
* - estimateTokens: always >= 0, monotonically related to string length
* - applyBudget: when hardFailed=false, estimatedTokens <= effectiveBudget
* - applyBudget: instructions and roadmap are ALWAYS kept verbatim (never trimmed)
* - applyBudget: below the minSet the call returns hardFailed=true with prompt=''
* - Budget boundary: a budget just at the effective floor triggers hard-fail;
* a budget just above it passes
*/
const { describe, test } = require('node:test');
const assert = require('node:assert/strict');
const fc = require('./helpers/fast-check-setup.cjs');
const { estimateTokens, applyBudget } = require('../gsd-core/bin/lib/prompt-budget.cjs');
// ─── Helpers ─────────────────────────────────────────────────────────────────
function minimalSections(overrides = {}) {
return {
instructions: 'Instructions text.',
roadmap: 'Roadmap text.',
plans: [{ file: 'plan.md', content: 'Plan content.' }],
projectMd: null,
context: null,
research: null,
requirements: null,
...overrides,
};
}
// ─── estimateTokens property tests ───────────────────────────────────────────
describe('prompt-budget: estimateTokens properties', () => {
test('property: estimateTokens is always >= 0', () => {
fc.assert(
fc.property(
fc.oneof(
fc.string(),
fc.constant(null),
fc.constant(undefined),
fc.constant(''),
fc.string({ unit: 'binary', maxLength: 200 }),
fc.string({ unit: 'grapheme-composite', maxLength: 200 })
),
(input) => {
const result = estimateTokens(input);
assert.ok(typeof result === 'number', `estimateTokens must return number, got ${typeof result}`);
assert.ok(result >= 0, `estimateTokens(${JSON.stringify(input)}) must be >= 0, got ${result}`);
assert.ok(Number.isInteger(result), `estimateTokens must return integer, got ${result}`);
}
)
);
});
test('property: estimateTokens is monotonically non-decreasing as text grows', () => {
fc.assert(
fc.property(
fc.string({ maxLength: 500 }),
fc.string({ minLength: 1, maxLength: 100 }),
(base, suffix) => {
const short = estimateTokens(base);
const long = estimateTokens(base + suffix);
assert.ok(long >= short, `tokens('${base}' + suffix)=${long} < tokens('${base}')=${short}`);
}
)
);
});
test('property: estimateTokens(null/undefined) returns 0', () => {
assert.equal(estimateTokens(null), 0);
assert.equal(estimateTokens(undefined), 0);
assert.equal(estimateTokens(''), 0);
});
test('property: estimateTokens approximation is ceil(len/4)', () => {
fc.assert(
fc.property(fc.string({ minLength: 1, maxLength: 1000 }), (text) => {
const expected = Math.ceil(text.length / 4);
assert.equal(estimateTokens(text), expected);
})
);
});
});
// ─── applyBudget property tests ───────────────────────────────────────────────
describe('prompt-budget: applyBudget properties', () => {
// (a) Boundary property: budget near the hardFail threshold
test('property: when minSet > effectiveBudget, applyBudget returns hardFailed=true and prompt=""', () => {
fc.assert(
fc.property(
// Use a very small budget to force hard-fail
fc.integer({ min: 1, max: 50 }),
(tinyBudget) => {
const sections = minimalSections({
instructions: 'A'.repeat(200), // ~50 tokens
roadmap: 'B'.repeat(200), // ~50 tokens
});
const result = applyBudget({ sections, budget: tinyBudget });
if (result.metadata.hardFailed) {
assert.equal(result.prompt, '', 'hardFailed must return empty prompt');
assert.equal(result.metadata.hardFailed, true);
}
// If not hard-failed, that is also valid — result is consistent
}
)
);
});
test('property: when budget is adequate, estimatedTokens never exceeds effectiveBudget', () => {
// Use a large budget: instructions ~10 tokens + roadmap ~10 tokens + plan ~10 tokens
// With margin 10%, effectiveBudget = floor(budget * 0.9)
fc.assert(
fc.property(
fc.integer({ min: 500, max: 10_000 }),
(budget) => {
const sections = minimalSections();
const result = applyBudget({ sections, budget });
if (!result.metadata.hardFailed) {
assert.ok(
result.metadata.estimatedTokens <= result.metadata.effectiveBudget,
`estimatedTokens ${result.metadata.estimatedTokens} > effectiveBudget ${result.metadata.effectiveBudget} at budget=${budget}`
);
assert.ok(result.prompt.length > 0, 'non-hardFailed must return non-empty prompt');
}
}
)
);
});
// (b) Robustness: hostile section inputs — applyBudget should either work or throw
// clearly — it must NEVER silently return a broken shape
test('property: applyBudget always returns typed { prompt, metadata } shape on valid budget', () => {
fc.assert(
fc.property(
fc.integer({ min: 100, max: 50_000 }),
fc.string({ maxLength: 200 }),
fc.string({ maxLength: 200 }),
(budget, instructions, roadmap) => {
const sections = minimalSections({ instructions, roadmap });
const result = applyBudget({ sections, budget });
assert.ok(typeof result === 'object' && result !== null);
assert.ok(typeof result.prompt === 'string', 'prompt must be string');
assert.ok(typeof result.metadata === 'object' && result.metadata !== null);
assert.ok(typeof result.metadata.hardFailed === 'boolean');
assert.ok(typeof result.metadata.budget === 'number');
assert.ok(typeof result.metadata.effectiveBudget === 'number');
assert.ok(Array.isArray(result.metadata.omitted));
}
)
);
});
test('property: instructions are always present verbatim in the output prompt', () => {
fc.assert(
fc.property(
fc.string({ minLength: 1, maxLength: 100 }),
(instructions) => {
const sections = minimalSections({ instructions });
const result = applyBudget({ sections, budget: 100_000 });
if (!result.metadata.hardFailed) {
assert.ok(
result.prompt.includes(instructions),
`Instructions not found verbatim in prompt. Instructions: "${instructions.slice(0, 50)}"`
);
}
}
)
);
});
test('property: roadmap is always present verbatim in the output prompt', () => {
fc.assert(
fc.property(
fc.string({ minLength: 1, maxLength: 100 }),
(roadmap) => {
const sections = minimalSections({ roadmap });
const result = applyBudget({ sections, budget: 100_000 });
if (!result.metadata.hardFailed) {
assert.ok(
result.prompt.includes(roadmap),
`Roadmap not found verbatim in prompt. Roadmap: "${roadmap.slice(0, 50)}"`
);
}
}
)
);
});
test('property: safetyMarginPct in [0,50] always produces effectiveBudget <= budget', () => {
fc.assert(
fc.property(
fc.integer({ min: 1000, max: 100_000 }),
fc.integer({ min: 0, max: 50 }),
(budget, safetyMarginPct) => {
const sections = minimalSections();
const result = applyBudget({ sections, budget, options: { safetyMarginPct } });
assert.ok(
result.metadata.effectiveBudget <= budget,
`effectiveBudget ${result.metadata.effectiveBudget} > budget ${budget} at margin ${safetyMarginPct}%`
);
}
)
);
});
test('property: context/research/requirements omission is tracked in metadata.omitted', () => {
// Build a sections object where extras push it over a tight budget
fc.assert(
fc.property(
fc.string({ minLength: 400, maxLength: 800 }), // ~100-200 tokens context
(contextText) => {
const sections = minimalSections({ context: contextText });
// Use a very tight budget that forces trimming
const baseTokens = estimateTokens('Instructions text.') +
estimateTokens('Roadmap text.') +
estimateTokens('Plan content.') + 20; // overhead
const tightBudget = Math.ceil(baseTokens / 0.9) + 1; // just barely fits without context
const result = applyBudget({ sections, budget: tightBudget });
if (!result.metadata.hardFailed && result.metadata.omitted.includes('context')) {
// The note must have been injected if context was dropped
assert.equal(result.metadata.noteInjected, true,
'noteInjected should be true when context was omitted');
}
// Whether or not context was dropped, omitted is always an array
assert.ok(Array.isArray(result.metadata.omitted));
}
)
);
});
});