f7df920681f233ae0fe064ee659550bdf41ff708
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
325fc25c01 |
fix(#3569): require a digit-bearing phase id in the stats heading scan (#3591)
* test(#3569): pin stats phase-id shape — inline-code mentions produce no phantom row Failing-first regression for #3569: cmdStats' heading scan accepted any word as a phase id, so prose mentioning ### Phase N: inside inline code inflated phases_total and disagreed with roadmap analyze. New adversarial fixture phase-heading-inside-inline-code.md (blockquote + bare mention), parity assertion against roadmap analyze, and over-narrowing guards for decimal / milestone-prefixed / letter-prefixed ids. * fix(#3569): require a digit-bearing phase id in the stats heading scan cmdStats' hand-rolled heading pattern captured any word as a phase id, so a ### Phase N: token inside an inline code span (the issue's blockquote) produced a phantom Not-Started row that could never complete, inflating phases_total and deflating percent forever. The id capture is now the canonical #3036 shape roadmap.cts uses (digit required; letter-prefixed, decimal, and milestone-prefixed ids keep counting), so stats and roadmap analyze agree. * fix(#3569): sanction the stats id-shape literal; correct zero-padded expectation Review findings: the phase-id drift guard requires the // phase-id-owner: comment directly above the regex (same form as roadmap.cts); the milestone-prefixed over-narrowing guard must expect normalizePhaseName's zero-padded 02-01 form, not the raw 2-01 token. * chore(#3569): add changeset fragment * chore(#3569): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
4a1ed2531f |
enhance(#3242): validate codex .toml model posture, not just presence (#3290)
* test(#3242): failing-first suite for the codex posture health-check Specifies ADR-2313 D6 before the implementation exists, so the tests bind to the contract rather than to whatever the code happens to do. RED is established by construction, not by a remote run: checkCodexModelPosture and POSTURE_REASON are absent from the compiled lib today, so every row fails on the missing export. A remote checkpoint here would prove only that the function is missing, which is already known — so the run is deliberately deferred to the combined green checkpoint rather than spent proving a tautology. That makes the NEGATIVE PROOFS the rows that carry real signal. Every positive row passes even for a naive implementation that greps /model\s*=/ over the whole file. Six rows fail it: light-tier service_tier/model_verbosity decoupling (#774), hand-added keys, a commented pin, the model_verbosity key-prefix collision, the runtime no-op ordering, and the headline case — a literal `model = "sonnet"` inside the developer_instructions ''' block, which the emitter fills with agent prompts that discuss models constantly. Row 14's fixture was verified to discriminate before being written: a whole-file scan matches it and a header-slice scan does not. Without that check the test would pass trivially and prove nothing, which is the vacuous-test failure this epic has already hit repeatedly. Adversarial TOML fixtures are hand-authored against the real Codex shape rather than generated by generateCodexAgentToml, per #2371 — a fixture from the writer can only confirm what the writer already believed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3242): validate codex .toml model posture, not just presence Implements ADR-2313 D6. checkCodexModelPosture is a new sibling export, not a branch inside checkAgentsInstalled — that function carries 33 upstream dependents, cyclomatic 25, and sits in two traced process flows, so it is deliberately left untouched. It imports isAnthropicFlavoredModel from model-catalog, a genuine leaf. That is what Phase 1's constant move bought: agent-install-check is documented as pure read/verify and imports only leaves, so reaching the rule through model-resolver would have dragged config-loader into it. Reads liberally, judges strictly, and never guesses. Tolerates comments, key order, whitespace, CRLF, and a BOM; anchors on full key names so model_verbosity does not satisfy a `model` probe; treats extra hand-added keys as none of its business, since the check is a predicate on the two fields the posture owns rather than a whitelist over the document. An unreadable file becomes a named violation and the loop keeps going. The scan covers only the header slice — the lines before the developer_instructions ''' marker. The emitter writes agent prompts into that block and GSD's prompts discuss models constantly, so a whole-file scan reports violations for prose. This is the highest-risk defect in the phase and the reason its fixture was verified to discriminate before being written. The non-codex short-circuit runs before any filesystem call, so a stray .toml under another runtime is never inspected. Wired through cmdValidateAgents as an additive codex_posture key, so a violating install is visible from a command a user actually runs rather than only from a library nothing calls. Also fixes a test defect found while implementing: .gitattributes forces `* text=auto eol=lf` repo-wide, so the committed CRLF fixture was normalized to LF in the index — `git ls-files --eol` reported `i/lf w/crlf`, the working copy being stale pre-normalization bytes. The CRLF row was asserting against a file that could not survive a fresh clone. CRLF is now derived at runtime, which puts it under the test's control rather than git's, instead of adding a .gitattributes exception that fights a deliberate repo-wide policy and that anyone could re-normalize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3242): document the codex posture check where a user will look Three quadrants, filed by where the reader actually arrives. How-to (recover-and-troubleshoot.md, under Install and update problems) is titled by the SYMPTOM — "If Codex agents fail to spawn with a 400 about an unsupported model" — and opens with the verbatim error string. Someone hitting this does not know the words "posture" or "ADR-2313"; they have a 400 in their terminal and will search for that. Reference (COMMANDS.md) had no `validate agents` entry at all, though sibling gsd-tools subcommands are documented. Adding user-visible output to an undocumented command and then linking to it from the new how-to would have left a dangling reference. The entry carries the violation-reason table, since the frozen POSTURE_REASON enum is the machine-readable contract a reader needs rather than the prose. Both surfaces state that presence and posture are separate verdicts — a missing agent lands in `missing`, never as a posture violation. That is a deliberate design decision and would otherwise be invisible to someone watching one command emit both. Explanation stays in ADR-2313, which already covers D6 and the liberal-parse/strict-judge boundary. Pointing at it beats duplicating it into COMMANDS.md and creating two copies to drift. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3242): close two false negatives in the posture scan Both found by an isolated reviewer and reproduced before fixing. Both made the check report clean when it was not — the worst direction for this function, since the how-to tells users an empty violations list means the install is posture-clean. A quoted TOML key was never matched. `"model" = "sonnet"` is legal TOML, but the key pattern required a bare identifier, so the pin was silently invisible. Bare, "double" and 'single' quoted forms now normalize to the same key name. The block marker was found by unanchored whole-content search and used to truncate the header. A `description` value merely containing the literal text `developer_instructions = '''` truncated the scan before a real pin, and a user who hand-reordered `model` to sit after the block — still legal TOML — was never scanned at all. Fixed by changing the strategy rather than the regex: find the block's range, anchored at line start, and scan every line OUTSIDE it. That covers both failures and is strictly more correct than truncation, while still never reading prompt prose. An unterminated block excludes the rest of the file, which fails toward a false positive — the safe direction, since misreading prose as a pin wastes a user's time while the alternative hides a real one. Also corrects two overclaims of mine. The how-to named "v1.11", a version that does not exist — package.json is 1.10.0 and unreleased — so it now describes the boundary by behavior and links the ADR. And the test matrix asserted that a naive whole-file scan "fails exactly rows 12,13,14,15,16,25"; the reviewer computed that rows 12, 13, 15 and 16 produce the correct result against that baseline too. They guard real but *different* mistakes, and the matrix now says which one each catches instead of attributing them all to the header-slice defect. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3242): skip symlinked agent files instead of following them Security review finding. The scan listed entries with readdirSync and read them with readFileSync, which follows symlinks — so a symlink in the agents directory pointing anywhere would have its contents read, and any line matching the model pattern echoed into cmdValidateAgents' output through the `value` field. A read-and-echo primitive on an arbitrary path. It needs write access to the agents directory, so it crosses no new trust boundary today. Fixed anyway, for two reasons. This repo already does it correctly next door: cmdEffortSync filters with lstatSync().isFile() and the comment "Skip symlinks — only write regular files to avoid clobbering symlink targets." Being inconsistent with a sibling in the same subsystem IS the defect. And Phase 3 (#3243) extends that same cmdEffortSync to WRITE these files. Establishing symlink-following as the house pattern for Codex .toml handling here would hand Phase 3 a worse starting point while it writes rather than reads. Skipped silently rather than reported, matching the sibling: a symlinked agent file is a structural install choice, which checkAgentsInstalled owns, not a model-content posture defect. An lstat that itself throws excludes the file rather than crashing the scan. That does narrow the guarantee slightly, so the how-to now says an empty list means every REGULAR .toml is clean, and tells anyone symlinking their configs to check the targets by hand. Claiming a clean bill of health over files the check declined to open would be the same kind of false confidence the two false negatives above produced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3242): backfill changeset pr number (#3290) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cd5db1f8db |
test(suites): seed security/slow/integration suites via measured retags
Renames (git mv) with all references updated (ci-test-scope RULES, windows-parity allowlist, test-file-count allowlist, docs in 6 locales): - 5 scanner tests -> *.security.test.cjs — the 'Run security tests' CI step ran zero files since the suite taxonomy landed; it is now honest. - graphify-auto-update -> *.slow.test.cjs (36s, slowest file in the suite; e2e gsd-tools spawns) — runs on full-matrix lanes and push to next. - installer-migration-install-integration -> *.integration.test.cjs (13s; an integration test by its own name). Coverage gate measured after retags: 88.55% lines (gate 70%). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
463cffd894 |
chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/ Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary (`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers are unaffected. Mechanical (bulk, ~90% of the diff): - `git mv get-shit-done gsd-core` - Swept path/identifier references across the repo via `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead preserves the five legitimate slug variants that are NOT the directory: get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names). - Build/manifest wiring: package.json (bin, files, coverage globs), tsconfig.build.json (outDir), ~86 .gitignore build-output entries, stryker.config.mjs, scan-ignore files, install.js path strings. - Frozen (not rewritten): CHANGELOG.md history; translated docs (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/). New logic (review here): - src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper ADR-0008 installer migration. On upgrade it walks the legacy `~/.claude/get-shit-done/` tree, classifies each file via the prior install manifest, and emits remove-managed / backup-and-remove for managed files while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked root and symlinked entries; bounds-checks every path under configDir). The framework rolls back on install failure. Emptied dirs may remain (framework has no recursive dir-removal primitive) — documented. - scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare `get-shit-done` directory token (split token to avoid self-match; case- insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines). Wired into the lint-tests CI job. - Restored scripts/lint-package-identity-drift.cjs detection regexes (the mechanical sweep had wrongly rewritten the old-name patterns it exists to detect) and marked them as intentional legacy references. - TDD tests for the migration and the guard; do.md slash-command guard regex tightened so a `/gsd-core/bin` path segment is not mistaken for a command; changeset + docs/installer-migrations.md row added. Breaking: the installed runtime path moves `~/.claude/get-shit-done/` -> `~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed files (preserving user files) on upgrade. Users with custom hooks/configs hardcoding the old path must update them. Closes #604 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unsweep pending changesets + allowlist injection-example docs CI fixes for the rename PR: - Do not sweep pending .changeset/*.md (ephemeral release-note fragments, like CHANGELOG); reverted those body edits so 5 pre-existing malformed fragments (missing type/pr) no longer enter the PR diff and trip docs-lint. Allowlisted .changeset/ in the legacy-name guard accordingly. - Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in prompt-injection-scan.sh: they contain intentional injection examples / security-model prose; the path-reference rewrites are kept. CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR; none in the new migration/guard) and are out of scope for the rename. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): resolve CodeQL alerts surfaced on this PR The rename diff touched files carrying pre-existing CodeQL findings; per the no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving them off. All behavior-preserving: - scripts/ci-test-scope.cjs: build the config-path match from string .includes() instead of a RegExp over an arg-derived value (js/regex-injection). - src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName so the table-cell escape is complete (js/incomplete-sanitization). - tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization). - tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace, keep the meaningful POSIX-class conversion (js/identity-replacement). Verified: build:lib green; the touched test files + ci-test-scope + profile-output suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization) The prior commit's fixes for two alerts were ineffective: - ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file` reaching static regex `.test(file)` calls (not the config rule). Removed ALL regex over file/t — startsWith/includes/=== string checks + an isWindowsHint helper — so there is no regex sink for the tainted value. - js/incomplete-multi-character-sanitization (3 test files): a single `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint loop (replace until stable) plus a final bare-opener strip. Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL CodeQL flags the regex PATTERNS syntactically (regex-injection on the --files arg split; incomplete-multi-character-sanitization on the <!--...--> replace), so loop fixes do not satisfy it. Made these paths regex-free: - ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/). - 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)). Behavior preserved; ci-test-scope + the 3 suites pass; guard clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unblock security base64 scan on the large rename diff The security job hit its 10m timeout: base64-scan.sh choked on the binary test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/ non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings), and the ~800-file rename diff is slow to scan regardless. - scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they can't carry base64-obfuscated *text* and feeding NUL bytes through the per-line scanner is pathologically slow. collect_files already filtered binary *extensions*; this catches binary *content* in text extensions. - .github/workflows/security-scan.yml: raise the security job timeout 10m->30m to accommodate very large diffs (the scan itself is unchanged). Verified locally: scan skips the fixture, 0 "ignored null byte" warnings, 0 findings, exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): sweep get-shit-done refs introduced by merging next The branch was updated with next (#614/#384/#618 etc.), which reference the get-shit-done/ dir (still named that on next). Swept the stale references in the merged files to gsd-core so the rename stays consistent and lint:legacy-name passes: - commands/gsd/discuss-phase.md (runtime-launcher shim paths) - src/core.cts (getAgentsDir layout comments) - tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib) Verified: guard 0 violations; build green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant The #614 runtime-launcher shim added to discuss-phase.md references `${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it mis-read the directory path as a dangling `/gsd-core` command ref (same class as the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path segments are not treated as slash-command references. Verified locally on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22 image) full suite: 0 failures - bug-3683 + bug-2954 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI) CI intermittently failed state.test's gsd-tools subprocess with "findProjectRoot is not a function" (flip-flopping across legs; not reproducible on mac full suite, gsd-test linux full suite, test:unit, or state.test x8). findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs); binding it via destructure at module-load can be undefined under a load-ordering edge. Resolve it lazily at call time via a small wrapper so the lookup happens after core.cjs is fully initialized. Verified green on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22) full suite: 0 failures - state.test.cjs: 106/106; gsd-tools loads cleanly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): allowlist verification-patterns.md placeholder examples in secret scan The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var examples (illustrative Stripe test-key / database-URL / API-key placeholders) — not real credentials. Added it to .secretscanignore with the strict annotation, mirroring the existing gsd-core/workflows/plan-phase.md exception. Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next exits 0 with 0 findings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e32a53b974 |
feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (#133)
* test(113): add per-rule failing tests + hostile fixture for markdown link payloads RED phase for issue #113 — scanForInjection() currently returns { clean: true } for markdown links containing javascript:, data:text/html, userinfo credentials, and token-in-query payloads. Changes: - tests/fixtures/adversarial/security/context-malicious-markdown-link.md: Extended to contain one hostile example per rule class (MD-LINK-JS-SCHEME, MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY) plus benign negative controls (data:image/png, mailto:, https://github.com, port-only URL). - tests/security-prompt-injection.test.cjs: - Flipped PINNED "malicious-markdown-link fixture is NOT flagged" assertion to "malicious-markdown-link fixture is flagged by scanner" (forward-looking). - Added 4×positive + 4×negative per-rule unit tests asserting structuredFindings with ruleId, file, line, match fields. - Added parity guard: every MARKDOWN_LINK_PATTERNS source string from security.cjs must appear in gsd-read-injection-scanner.js hook source. D3 false-positive grep: 0 legitimate matches — no allowlist entries needed. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(113): detect javascript:/data:/userinfo/token-in-query in markdown links (security.cjs + hook) GREEN phase for issue #113. Rule details (all with primary source citations): MD-LINK-JS-SCHEME Flags ](javascript:...) regardless of case. Source: OWASP XSS Prevention Cheat Sheet https://cheatsheetseries.owasp.org/cheatsheets/Cross_Site_Scripting_Prevention_Cheat_Sheet.html MD-LINK-DATA-SCHEME Flags data: URIs NOT in the explicit safe-list. Safe-list: image/(png|jpeg|gif|webp|bmp|ico|avif|heic) and font/(woff2?|otf|ttf). data:image/svg+xml is intentionally BLOCKED — SVG can host <script>. Source: OWASP File Upload Cheat Sheet — SVG Files https://cheatsheetseries.owasp.org/cheatsheets/File_Upload_Cheat_Sheet.html#svg-files MD-LINK-USERINFO Flags https?://user:pass@host in markdown link targets. Does NOT fire on: mailto:user@host (no :// before user) or https://host:443/path (port, not userinfo). Source: RFC 3986 §3.2.1 (userinfo syntax) https://www.rfc-editor.org/rfc/rfc3986#section-3.2.1 RFC 9110 §4.2.4 (HTTP deprecates userinfo) https://www.rfc-editor.org/rfc/rfc9110#section-4.2.4 MD-LINK-TOKEN-IN-QUERY Flags key NAMES: token, access_token, id_token, refresh_token, api_key, apikey, secret, password, client_secret, code — regardless of value. Source: RFC 9700 OAuth 2.0 Security BCP §4.3.1 https://www.rfc-editor.org/rfc/rfc9700#section-4.3.1 D3 false-positive grep: 0 legitimate matches in codebase — no allowlist needed. Architecture: - scripts/security.cjs: canonical MARKDOWN_LINK_PATTERNS export, scanForInjection() extended with structuredFindings (ruleId, file, line, match) via opts.file. - hooks/gsd-read-injection-scanner.js: patterns inlined for hook independence (same pattern sources, verified by parity test). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test(113): flip PINNED malicious-markdown-link assertion and add parity guard REFACTOR phase — tightening test rigor after test-rigor skill review: 1. Fixture assertion now enumerates all 4 expected ruleIds explicitly: [MD-LINK-JS-SCHEME, MD-LINK-DATA-SCHEME, MD-LINK-USERINFO, MD-LINK-TOKEN-IN-QUERY]. Previously findings.length > 0 would pass even if 3 of 4 rules were broken. 2. line field assertions tightened: `f.line >= 1` (meaningful lower bound for 1-based line numbers) instead of `typeof f.line === 'number'` (vacuous). 3. match field assertions tightened to check the hostile content is present: - MD-LINK-JS-SCHEME: /javascript:/i in match - MD-LINK-DATA-SCHEME: /data:/i in match - MD-LINK-USERINFO: /@/ in match (the @ character is the definitive userinfo marker) - MD-LINK-TOKEN-IN-QUERY: /token=/i in match 4. Parity test checks actual RegExp .source strings (not just lengths), verifying the hook contains the exact canonical pattern sources character-for-character. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(#113): add changeset fragment + Windows/Node 24 state.test compatibility 1. .changeset/113-malicious-markdown-links.md — required Security fragment for the user-facing markdown-link scanner changes in this PR (changeset-lint was failing with FAIL_MISSING_FRAGMENT). 2. get-shit-done/bin/lib/state-command-router.cjs — add OUTPUT_ON_SDK_ERROR set for mutation state subcommands whose CJS contract is always exit-0. On Windows/Node 24 the SDK bridge returns result.ok===false for validation failures (e.g. state record-metric --phase 1 with no --plan/--duration), causing dispatchViaSdk() to call error() (exit 1) instead of output({error}) (exit 0). The fix maps SDK non-ok results to JSON output for the affected mutation commands (record-metric, advance-plan, record-session, add-decision, add-blocker, resolve-blocker, update-progress), restoring the exit-0 CJS contract on all platforms. tests/state.test.cjs:1161 "returns error when required fields missing" passes locally (104/104 pass). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
835dd6ab44 |
test(3596): adversarial security/prompt-injection abuse suite (#3654)
* test(3596): adversarial security/prompt-injection abuse suite Adds `tests/security-prompt-injection.test.cjs` and a fixtures directory at `tests/fixtures/adversarial/security/` covering the attack classes enumerated in #3596: - Command substitution / backticks / heredoc payloads in workstream names — sentinel-file probes prove no shell is spawned, slugifier neutralises the input. - Path traversal through `--ws` and slash-bearing workstream names — rejected with structured `--json-errors` payload, no stack trace, no filesystem mutation outside the project root. - Fake `<system>` / `[SYSTEM]` / `<<SYS>>` / `[INST]` boundary tags — sanitizeForPrompt neutralises every form; structural negative property locked across all six styles in one place. - Zero-width / bidi-override codepoints — stripped per the documented codepoint set; asserted via codePoint inspection, not regex literals. - Hostile read of CONTEXT.md / PLAN.md / ROADMAP.md fixtures — `gsd-read-injection-scanner.js` surfaces the advisory; excluded paths and non-Read tools stay silent; malformed JSON does not crash the hook. - Hostile write of `.planning/` files — `gsd-prompt-guard.js` emits a `PreToolUse` advisory; non-Write/Edit tools stay silent. - Fake `ghp_*` / `sk-*` env tokens — never echoed in CLI stdout or stderr under hostile inputs; covered under `// allow-test-rule: structural-regression-guard` because the only way to assert byte-level absence is `.includes(token)` against the captured streams. - `validatePath`, `validateShellArg`, `validatePhaseNumber`, `validateFieldName` — focused negative-input contract pins. Pinned behavior gaps (documented, NOT fixed in this PR): - `<instructions>` is intentionally whitelisted by both the scanner and the sanitiser (GSD's own prompt scaffolding). Two REGRESSION GUARD tests lock that contract. - The current `scanForInjection` does NOT flag malicious markdown links (javascript:/data:/embedded-credentials URLs). PINNED with negative-proof so any future scope extension fails the assertion and forces a deliberate update to the acceptance map. - `prompt-builder.ts` does not yet wrap plan/context markdown in an "untrusted data" envelope. That seam lives on the TS side and is covered by `sdk/src/prompt-builder.test.ts`; out of scope for a CJS test file. Mentioned in the file header. Verification: - `node --test tests/security-prompt-injection.test.cjs` → 73 tests pass. - `node scripts/lint-no-source-grep.cjs` → 0 violations across 546 test files (one `allow-test-rule: structural-regression-guard` annotation on this file for the token-absence assertions). - `node scripts/run-tests.cjs` → 9730 tests pass, 0 fail. Refs #3596 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(3596): allow adversarial fixtures in scan + harden graphify status parse * fix(3596): skip adversarial security fixtures in secret scan --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
b1317633db |
test(3594): adversarial parser fixtures + frontmatter/roadmap matrix + property-style suite (#3633)
* test(3594): adversarial parser fixtures + frontmatter/roadmap matrix + property-style suite
Lands the adversarial parser-input corpus that CONTRIBUTING.md
§"QA Matrix Requirements / Parser and project-file inputs" and
TEST-EXAMPLES.md §"Parser Adversarial Fixtures" describe.
New tests/fixtures/adversarial/ layout:
frontmatter/
duplicate-keys.md — same key twice (collapses last-wins)
crlf-mixed.md — CRLF endings throughout
unclosed-block.md — `---` open with no close
unicode-keys-and-values.md — non-ASCII + emoji + Greek
null-byte-value.md — U+0000 in a value
huge-bounded.md — 2000-item array, ~30KB
roadmap/
phase-heading-inside-fenced-code.md — #2787 fence shadowing
nested-fenced-code.md — outer + inner ``` blocks
unicode-phase-titles.md — JP / Greek / emoji titles
repeated-phase-ids.md — phase 1 declared twice
decimal-phase-mixed.md — 2 vs 2.1 vs 2.10 vs 21
markdown-headings-inside-html-comment.md — comment shadowing
Test files (all node:test, no try/finally in test bodies, no source-grep,
no raw-text matching on stdout/file content):
tests/feat-3594-parser-adversarial-frontmatter.test.cjs (12 tests)
Loads each fixture, pins parser invariants on extractFrontmatter()
return shape. Cross-corpus "does not throw on any fixture" sweep.
tests/feat-3594-parser-adversarial-roadmap.test.cjs (18 tests)
Loads each fixture into a temp project's .planning/ROADMAP.md and
drives `gsd-tools roadmap get-phase <N>` via the runCli harness
introduced by #3593. Asserts on the typed JSON payload.
tests/feat-3594-parser-property-style.test.cjs (2 tests)
Deterministic mulberry32 PRNG generates 500 malformed-ish
frontmatter inputs per test. Pins (a) extractFrontmatter is total
over the corpus (no null-deref TypeError, always returns a plain
object on success), (b) the suite completes well under 2 seconds
(quadratic-regression guard).
Known-open bugs surfaced and pinned (intentionally NOT fixed in this
PR — separate issues warranted):
- CJS roadmap parser matches `## Phase N:` headings inside fenced
code blocks (the SDK parser tracks fences per the #2787 comment in
sdk/src/query/roadmap.ts but the CJS path has not caught up).
- CJS roadmap parser matches `## Phase N:` headings inside HTML
comments.
Both are documented in-test with the "currently STILL matches it
(open: needs <fix>)" naming pattern so the day the production fix
lands, flipping the assertion from `found: true` to `found: false` is
the regression guard.
Test totals:
- 32 new feat-3594-* tests (12 frontmatter + 18 roadmap + 2 property)
- 108/108 pass when running together with the pre-existing
frontmatter.test.cjs + roadmap.test.cjs suites (76 of theirs).
Closes #3594
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(3594): use Fisher-Yates shuffle for deterministic seeded inputs
Replaces `arr.sort(() => rng() - 0.5)` with a Fisher-Yates shuffle
driven by the supplied PRNG. The sort-based shuffle is non-transitive:
V8's TimSort behavior on non-transitive comparators is engine-defined,
so the same seed produced different orderings across Node versions —
undermining the test's stated reproducibility guarantee.
Fisher-Yates is O(n), transitive (no comparator at all), and consumes
exactly n-1 RNG values in a fixed order. The mulberry32 seed now
determines the input sequence end-to-end.
Codex review on PR #3633.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|