* chore(#604): rename get-shit-done/ runtime directory to gsd-core/ Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary (`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers are unaffected. Mechanical (bulk, ~90% of the diff): - `git mv get-shit-done gsd-core` - Swept path/identifier references across the repo via `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead preserves the five legitimate slug variants that are NOT the directory: get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names). - Build/manifest wiring: package.json (bin, files, coverage globs), tsconfig.build.json (outDir), ~86 .gitignore build-output entries, stryker.config.mjs, scan-ignore files, install.js path strings. - Frozen (not rewritten): CHANGELOG.md history; translated docs (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/). New logic (review here): - src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper ADR-0008 installer migration. On upgrade it walks the legacy `~/.claude/get-shit-done/` tree, classifies each file via the prior install manifest, and emits remove-managed / backup-and-remove for managed files while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked root and symlinked entries; bounds-checks every path under configDir). The framework rolls back on install failure. Emptied dirs may remain (framework has no recursive dir-removal primitive) — documented. - scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare `get-shit-done` directory token (split token to avoid self-match; case- insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines). Wired into the lint-tests CI job. - Restored scripts/lint-package-identity-drift.cjs detection regexes (the mechanical sweep had wrongly rewritten the old-name patterns it exists to detect) and marked them as intentional legacy references. - TDD tests for the migration and the guard; do.md slash-command guard regex tightened so a `/gsd-core/bin` path segment is not mistaken for a command; changeset + docs/installer-migrations.md row added. Breaking: the installed runtime path moves `~/.claude/get-shit-done/` -> `~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed files (preserving user files) on upgrade. Users with custom hooks/configs hardcoding the old path must update them. Closes #604 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unsweep pending changesets + allowlist injection-example docs CI fixes for the rename PR: - Do not sweep pending .changeset/*.md (ephemeral release-note fragments, like CHANGELOG); reverted those body edits so 5 pre-existing malformed fragments (missing type/pr) no longer enter the PR diff and trip docs-lint. Allowlisted .changeset/ in the legacy-name guard accordingly. - Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in prompt-injection-scan.sh: they contain intentional injection examples / security-model prose; the path-reference rewrites are kept. CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR; none in the new migration/guard) and are out of scope for the rename. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): resolve CodeQL alerts surfaced on this PR The rename diff touched files carrying pre-existing CodeQL findings; per the no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving them off. All behavior-preserving: - scripts/ci-test-scope.cjs: build the config-path match from string .includes() instead of a RegExp over an arg-derived value (js/regex-injection). - src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName so the table-cell escape is complete (js/incomplete-sanitization). - tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization). - tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace, keep the meaningful POSIX-class conversion (js/identity-replacement). Verified: build:lib green; the touched test files + ci-test-scope + profile-output suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization) The prior commit's fixes for two alerts were ineffective: - ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file` reaching static regex `.test(file)` calls (not the config rule). Removed ALL regex over file/t — startsWith/includes/=== string checks + an isWindowsHint helper — so there is no regex sink for the tainted value. - js/incomplete-multi-character-sanitization (3 test files): a single `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint loop (replace until stable) plus a final bare-opener strip. Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL CodeQL flags the regex PATTERNS syntactically (regex-injection on the --files arg split; incomplete-multi-character-sanitization on the <!--...--> replace), so loop fixes do not satisfy it. Made these paths regex-free: - ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/). - 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)). Behavior preserved; ci-test-scope + the 3 suites pass; guard clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unblock security base64 scan on the large rename diff The security job hit its 10m timeout: base64-scan.sh choked on the binary test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/ non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings), and the ~800-file rename diff is slow to scan regardless. - scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they can't carry base64-obfuscated *text* and feeding NUL bytes through the per-line scanner is pathologically slow. collect_files already filtered binary *extensions*; this catches binary *content* in text extensions. - .github/workflows/security-scan.yml: raise the security job timeout 10m->30m to accommodate very large diffs (the scan itself is unchanged). Verified locally: scan skips the fixture, 0 "ignored null byte" warnings, 0 findings, exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): sweep get-shit-done refs introduced by merging next The branch was updated with next (#614/#384/#618 etc.), which reference the get-shit-done/ dir (still named that on next). Swept the stale references in the merged files to gsd-core so the rename stays consistent and lint:legacy-name passes: - commands/gsd/discuss-phase.md (runtime-launcher shim paths) - src/core.cts (getAgentsDir layout comments) - tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib) Verified: guard 0 violations; build green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant The #614 runtime-launcher shim added to discuss-phase.md references `${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it mis-read the directory path as a dangling `/gsd-core` command ref (same class as the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path segments are not treated as slash-command references. Verified locally on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22 image) full suite: 0 failures - bug-3683 + bug-2954 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI) CI intermittently failed state.test's gsd-tools subprocess with "findProjectRoot is not a function" (flip-flopping across legs; not reproducible on mac full suite, gsd-test linux full suite, test:unit, or state.test x8). findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs); binding it via destructure at module-load can be undefined under a load-ordering edge. Resolve it lazily at call time via a small wrapper so the lookup happens after core.cjs is fully initialized. Verified green on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22) full suite: 0 failures - state.test.cjs: 106/106; gsd-tools loads cleanly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): allowlist verification-patterns.md placeholder examples in secret scan The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var examples (illustrative Stripe test-key / database-URL / API-key placeholders) — not real credentials. Added it to .secretscanignore with the strict annotation, mirroring the existing gsd-core/workflows/plan-phase.md exception. Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next exits 0 with 0 findings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
6.9 KiB
component, version, last_calibrated, calibration_source
| component | version | last_calibrated | calibration_source |
|---|---|---|---|
| verifier | 2 | 2026-03-25 | Adapted from thinking-partner verification corpus |
Verifier Few-Shot Examples
Positive Examples
Example 1: Thorough multi-level artifact verification with real gap discovery
Input: Must-have artifact: gsd-core/workflows/context-bridge.md: >50 lines
Output:
L1 (Existence): EXISTS --
gsd-core/workflows/context-bridge.mdfound, 187 lines. L2 (Substantive): SUBSTANTIVE -- No TODO/FIXME markers, 4 defined sections (updateContextTier,getContextBudget,injectContextWarning,resetTierState), non-trivial logic with config reads and threshold calculations. L3 (Wired): WIRED -- Referenced bygsd-core/workflows/execute-phase.mdat step 3. Called byagents/gsd-executor.mdvia @-reference include. Arguments verified: receives planning directory path matching the expected parameter. Status: PASSED
Why this is good: The verifier checked all three levels independently with specific evidence at each level -- file path, line count, exported symbols, import location, and argument verification. It did not stop at "file exists" but traced wiring through to actual usage with correct arguments.
Example 2: Correct FAILED classification with grep evidence
Input: Must-have truth: validate-commit.js rejects commits missing scope parentheses
Output:
Method: Ran
grep -n "scope" gsd-core/bin/validate-commit.jsto locate scope validation logic. Evidence: Lines 34-41 show regex/^(feat|fix|refactor|test|docs|chore|wip)\(/which requires opening paren after type. However, no test for MISSING closing paren exists -- the regex acceptsfeat(: description(missing scope content and closing paren). Ran:echo "feat(: test" | node gsd-core/bin/validate-commit.js-- exit code 0 (accepted). Status: FAILED -- Commit messages with empty scopefeat(: desc)are accepted. The truth "rejects commits missing scope parentheses" does not hold for empty-scope edge case.
Why this is good: The verifier did not trust the code's intent -- it ran an actual test with an edge case input and reported the specific failure. Evidence includes the regex, the test command, the exit code, and why the truth does not hold.
Example 3: Catching missing wiring (calibration-derived gap pattern)
Input: Must-have key_link: check-subagent-output.js imports and calls feedback-loop.js extractFeedback()
Output:
Method: Ran
grep -n "feedback-loop\|extractFeedback" gsd-core/bin/check-subagent-output.js. Evidence: Zero matches. The filefeedback-loop.jsexists (172 lines, substantive), andtests/feedback-loop.test.jspasses (229 lines). Butcheck-subagent-output.jsnever imports or calls it. Cross-check:grep -rn "require.*feedback-loop" gsd-core/bin/returns zero results across all scripts. Status: FAILED -- The artifact exists and is tested, but is completely orphaned. The key_link "imports and calls" does not hold. This is a missing wiring gap.
Why this is good: The verifier caught a common calibration pattern: code that exists, is tested, but is never actually wired into the system. L1 and L2 pass, but L3 (wiring) fails. This pattern accounts for 37% of all gaps found in calibration corpus.
Example 4: Catching missing config schema entries (calibration-derived gap pattern)
Input: Must-have truth: Config toggles features.architecture_graph and features.architecture_guard exist with defaults
Output:
Method: Ran
grep "architecture_graph\|architecture_guard" gsd-core/bin/lib/config-schema.json. Evidence: Zero matches in schema. However,grep "architecture" gsd-core/bin/lib/config.cjsshows both flags in CONFIG_DEFAULTS (lines 426, 428). Gap: Defaults exist in code but schema validation doesn't know about them.configValidate()would strip these fields on load, making the defaults unreachable. Status: FAILED -- Config flags work in code but are invisible to schema validation. This is a missing wiring gap between config layers.
Why this is good: The verifier checked multiple config layers (schema vs defaults vs code) rather than just one. Config wiring requires schema + defaults + code to agree -- checking only one layer misses layer-mismatch bugs.
Example 5: Distinguishing planning gaps from execution failures
Input: Must-have truth: All MSYS regex patterns consolidated into shared module
Output:
Method: Ran
grep -rn "\/\[a-zA-Z\]\/" gsd-core/bin/to find inline MSYS regex. Evidence: 2 matches incommands/gsd/misc.md(lines 26, 487). These were NOT in the original plan scope -- Phase 148 extracted code into misc.md AFTER Phase 149's plan was written. Assessment: The executor correctly completed all planned tasks. The gap is a planning ordering issue, not an execution failure. Status: PASS_WITH_NOTES -- Executor did what was planned. Unconsolidated regex noted as follow-up, not as a failure of this phase.
Why this is good: The verifier distinguished between what was planned vs what exists now. Code that changed AFTER plan creation is a planning gap, not an executor failure. This prevents false FAIL verdicts on correctly-executed work.
Calibration-Derived Gap Patterns
Based on analysis of verification corpus (80% pass rate, 8 gaps):
| Pattern | Count | % of Gaps | What to Check |
|---|---|---|---|
| Missing wiring | 3 | 37% | File exists + tested but never imported/called by consumer |
| Missing tests | 2 | 25% | Feature works but test file excluded from CI or never runs |
| Other | 3 | 38% | Config layer mismatches, planning ordering issues, schema gaps |
Negative Examples
Example 1: Blanket pass with no per-criterion evidence
Input: Must-haves: 5 truths, 3 artifacts, 2 key_links from plan frontmatter.
Output:
All must-haves have been verified. The implementation matches the plan requirements. Files exist and are properly connected. Status: PASSED. must_haves_checked: 10, must_haves_passed: 10
Why this is bad: No individual must-have is listed or checked. No file paths, no line numbers, no grep output, no level-by-level breakdown. The verifier restated the must-haves as passing without showing any evidence. This would miss stubs, orphaned exports, and broken wiring.
Example 2: File existence check without substance or wiring verification
Input: Must-have artifact: gsd-core/bin/lib/audit-checks/index.cjs: >100 lines
Output:
L1 (Existence): EXISTS -- File found at
gsd-core/bin/lib/audit-checks/index.cjs, 142 lines. Status: PASSED
Why this is bad: The verifier stopped at Level 1. The file has 142 lines but could contain // TODO: implement all checks with stub functions returning empty objects. Level 2 (substantive) and Level 3 (wired) were skipped entirely. A file that exists but is never imported or contains only placeholder code should not pass.