5e997de5f0f8afed5db2ad009f264df343f29656
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f4d6747d21 |
fix(#2738): report graphify query budget outcome and stop between-tier over-trimming (#2819)
* fix(#2738): report budget outcome from graphify query and stop over-trimming between tiers applyBudget retains seed nodes unconditionally, so the seed set is a floor the edge-tier reduction cannot go below — a --budget 500 request could return the full seed payload (~119k tokens measured) with no signal of the miss. Add budget_met + budget_estimate to the budget result and surface them through graphifyQuery when a budget was requested. Secondary: the tier loop estimated against the full pre-filter node set, so a tier removal that already satisfied the budget (once its orphaned nodes were excluded) still triggered the next, higher-confidence tier drop. Recompute reachability and the estimate after each tier and break as soon as the pruned result fits. Adjacent same-class instance from pre-submit review: the CLI forwards --budget 0 but truthiness checks silently treated it as no budget and returned the unbounded result. Test budget presence with != null so a zero budget is honored and reported as an unmeetable miss. * docs(changeset): backfill PR number for #2738 fragment * fix(#2738): estimate the payload as emitted, not a private compact form budget_met measured a different payload than the caller receives. The estimator serialized a compact `{nodes, edges}`, while output() emits the whole response pretty-printed (2-space indent, plus the term/total_*/trimmed wrapper keys). Measured on the repo's own SAMPLE_GRAPH fixture: reported 183 tokens against 302 actually emitted — 1.65x — so `--budget 200` returned budget_met: true while handing back 302 tokens. That is worse than the old silent miss: an automated consumer stops checking a signal that is confidently wrong. Fix the basis rather than the number: - io.cts gains serializeForOutput(), the single definition of the wire form. output() now calls it, so the estimator and the emitter cannot drift on indentation or shape. Pure extraction; output()'s behaviour is unchanged. - graphify builds its response through one buildQueryResponse() used by both the emitter and the estimator, so the estimate describes exactly the bytes returned. - The tier loop estimates on that same basis, so it keeps trimming until the real payload fits instead of stopping at a smaller internal measure. This makes budget_met === (budget_estimate <= budget) true by construction. - Drop the module-private chars/4 helper for prompt-budget's estimateTokens — the repo's single token scale, per the rule phase-estimation.cts documents. budget_estimate is self-referential (its own digits are part of the emitted bytes), resolved by iterating to a fixed point; the sequence only ever grows, so it settles in a couple of passes and errs toward over-reporting. Tests pin estimator to emitter so this cannot silently re-diverge if output() ever changes its indentation. Both new tests fail against the pre-fix source. * test(#2738): property-test the budget-limit and reporting contract RULESET.TESTS.property-based-testing names budget-limit contracts, and this module is the literal case: #2819 turns it into a *reporting* contract, which is what properties express well. Five invariants over arbitrary small graphs and any budget >= 0: - budget_met === (budget_estimate <= budget) - budget_estimate === the tokens actually emitted - the seed set is a floor the reduction never goes below (the changeset's "seeds are a floor" claim, previously asserted only for one hand-built fixture) - total_nodes/total_edges match the returned arrays - a larger budget never yields a smaller payload Two notes on what these do and do not prove. The emitted-payload property fails against the pre-fix source; the budget_met/budget_estimate agreement property does NOT — pre-fix both derived from the same wrong number, so it is a contract guard, not a regression proof. Monotonicity is asserted over the payload (node/edge counts), not over budget_estimate: the estimate measures emitted bytes exactly, and budget_met renders as "false" (5 chars) or "true" (4), so an identical payload can measure one token larger when the budget is missed. That is the estimate being honest, not a monotonicity break. The generator seeds on `label` — seedAndExpand matches label/description, never id/name, and a fixture that gets this wrong expands to nothing and passes vacuously. * test(#2738): pin the budget boundary and label the forward-guard test Two test-coverage gaps from review. RULESET.TESTS.boundary-coverage wants limit-1 / limit / limit+1. The added tests used 1, 0, 200, 100000, 50 — all far from the decision point. The branch that matters is `estimate <= budgetTokens`, so the input that decides it is budget === estimate exactly: that is the one value where an off-by-one in the comparison flips budget_met, and nothing else in the suite would catch it. Pinned at the limit and either side of it. The "omits budget fields when no budget was requested" test passes unchanged on next — graphifyQuery never set those keys before the fix, so both assertions already held. It has value as a forward guard against the spread leaking budget fields, but it is not failing-first and should not be counted toward RULESET.TESTS.regression-must-fail-first. Said so in a comment, so a later reader does not mistake it for the regression proof. * fix(#2738): keep a non-finite budget out of the budget path Switching `!budgetTokens` to `budgetTokens == null` widened the internal contract to admit NaN, where every `estimate <= NaN` is false: the loop strips all three tiers and returns a seeds-only payload that is indistinguishable from a legitimate aggressive trim. Unreachable through the CLI — graphify-command-router rejects a non-numeric --budget with makeInvalidArgs before graphifyQuery is called — so this is hardening, not a live defect. It is still worth guarding: graphifyQuery and applyBudget are module-level entry points a future caller could reach without the router's validation, and the failure mode is silent. Number.isFinite also routes Infinity to the no-budget path, deliberately: an unbounded budget is not a budget, and parseInt cannot produce one anyway. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
e3b829e765 |
refactor(#1306): gate graphify on isCapabilityActive (tri-state, runtime-aware) (#1313)
* refactor(#1306): gate graphify on isCapabilityActive (tri-state), not config-only graphify's command gate moves from the config-only isGraphifyEnabled to the shared isCapabilityActive('graphify', cwd) — so graphify is off unless installed AND surfaced AND graphify.enabled. Fixes a latent claude-hardcoding in the resolver: resolveCapabilityRuntimeState now detects the active runtime via resolveRuntime(cwd) (GSD_RUNTIME -> config.runtime -> 'claude') so non-Claude runtimes (Codex/Cursor) read their own surface, not ~/.claude. Hermetic regression test proves config-on+unsurfaced -> disabled; cross-runtime test proves GSD_RUNTIME=codex honors CODEX_HOME. Gate fails closed on every error path. Part of #1302. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1306): add changeset for graphify tri-state gate Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
463cffd894 |
chore(#604): rename get-shit-done/ runtime directory to gsd-core/ (#615)
* chore(#604): rename get-shit-done/ runtime directory to gsd-core/ Renames the installed runtime directory `get-shit-done/` to `gsd-core/` so the on-disk name matches the package (`@opengsd/gsd-core`), repo, and binary (`gsd-tools`). The npm package name and binary are unchanged; npx/npm consumers are unaffected. Mechanical (bulk, ~90% of the diff): - `git mv get-shit-done gsd-core` - Swept path/identifier references across the repo via `perl -pe 's/get-shit-done(?!-\w)/gsd-core/g'`. The negative lookahead preserves the five legitimate slug variants that are NOT the directory: get-shit-done-{OLD,cc,classic,cli,redux} (old package/repo names). - Build/manifest wiring: package.json (bin, files, coverage globs), tsconfig.build.json (outDir), ~86 .gitignore build-output entries, stryker.config.mjs, scan-ignore files, install.js path strings. - Frozen (not rewritten): CHANGELOG.md history; translated docs (README.<locale>.md and docs/{ja-JP,ko-KR,pt-BR,zh-CN}/). New logic (review here): - src/installer-migrations/003-rename-get-shit-done-to-gsd-core.cts: a proper ADR-0008 installer migration. On upgrade it walks the legacy `~/.claude/get-shit-done/` tree, classifies each file via the prior install manifest, and emits remove-managed / backup-and-remove for managed files while PRESERVING unknown user-added files. Symlink-safe (skips a symlinked root and symlinked entries; bounds-checks every path under configDir). The framework rolls back on install failure. Emptied dirs may remain (framework has no recursive dir-removal primitive) — documented. - scripts/lint-legacy-dir-name.cjs: CI regression guard forbidding the bare `get-shit-done` directory token (split token to avoid self-match; case- insensitive; `(?!-\w)` lookahead allows the slug variants; allowlists CHANGELOG, translated docs, and `gsd-allow-legacy-name` marker lines). Wired into the lint-tests CI job. - Restored scripts/lint-package-identity-drift.cjs detection regexes (the mechanical sweep had wrongly rewritten the old-name patterns it exists to detect) and marked them as intentional legacy references. - TDD tests for the migration and the guard; do.md slash-command guard regex tightened so a `/gsd-core/bin` path segment is not mistaken for a command; changeset + docs/installer-migrations.md row added. Breaking: the installed runtime path moves `~/.claude/get-shit-done/` -> `~/.claude/gsd-core/`. Migration 003 removes the stale legacy dir's managed files (preserving user files) on upgrade. Users with custom hooks/configs hardcoding the old path must update them. Closes #604 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unsweep pending changesets + allowlist injection-example docs CI fixes for the rename PR: - Do not sweep pending .changeset/*.md (ephemeral release-note fragments, like CHANGELOG); reverted those body edits so 5 pre-existing malformed fragments (missing type/pr) no longer enter the PR diff and trip docs-lint. Allowlisted .changeset/ in the legacy-name guard accordingly. - Allowlisted TEST-EXAMPLES.md and docs/explanation/security-model.md in prompt-injection-scan.sh: they contain intentional injection examples / security-model prose; the path-reference rewrites are kept. CodeQL alerts on this PR are pre-existing (alert lines unchanged by this PR; none in the new migration/guard) and are out of scope for the rename. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): resolve CodeQL alerts surfaced on this PR The rename diff touched files carrying pre-existing CodeQL findings; per the no-pre-existing-dismissal rule, fixing every surfaced alert rather than waving them off. All behavior-preserving: - scripts/ci-test-scope.cjs: build the config-path match from string .includes() instead of a RegExp over an arg-derived value (js/regex-injection). - src/profile-output.cts: escape backslashes before pipe-escaping desc/safeName so the table-cell escape is complete (js/incomplete-sanitization). - tests/{bug-2643,bug-2808,docs-parity-live-registry}: two-pass HTML-comment strip so a bare/unclosed `<!--` cannot survive (js/incomplete-multi-character-sanitization). - tests/inline-plan-threshold: drop the no-op `\s`->`\s` identity replace, keep the meaningful POSIX-class conversion (js/identity-replacement). Verified: build:lib green; the touched test files + ci-test-scope + profile-output suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): correctly resolve remaining CodeQL alerts (regex-injection + sanitization) The prior commit's fixes for two alerts were ineffective: - ci-test-scope.cjs js/regex-injection: the alert is the CLI-arg-derived `file` reaching static regex `.test(file)` calls (not the config rule). Removed ALL regex over file/t — startsWith/includes/=== string checks + an isWindowsHint helper — so there is no regex sink for the tainted value. - js/incomplete-multi-character-sanitization (3 test files): a single `.replace(/<!--...-->/g,'')` can let `<!--` re-form. Replaced with a fixpoint loop (replace until stable) plus a final bare-opener strip. Verified: no regex over file/t remains; ci-test-scope + the 3 test suites pass; lint:legacy-name clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): make ci-test-scope + comment-strippers regex-free to clear CodeQL CodeQL flags the regex PATTERNS syntactically (regex-injection on the --files arg split; incomplete-multi-character-sanitization on the <!--...--> replace), so loop fixes do not satisfy it. Made these paths regex-free: - ci-test-scope.cjs splitFiles: char-by-char separator tokenizer (no /[,\\s]+/). - 3 test files: indexOf/slice HTML-comment stripper (no .replace(/<!--/)). Behavior preserved; ci-test-scope + the 3 suites pass; guard clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): unblock security base64 scan on the large rename diff The security job hit its 10m timeout: base64-scan.sh choked on the binary test fixture tests/feat-3594-parser-property-style.test.cjs (embedded NUL/ non-UTF8 bytes -> thousands of bogus blobs + "ignored null byte" warnings), and the ~800-file rename diff is slow to scan regardless. - scripts/base64-scan.sh: skip binary-by-content files (grep -Iq .) — they can't carry base64-obfuscated *text* and feeding NUL bytes through the per-line scanner is pathologically slow. collect_files already filtered binary *extensions*; this catches binary *content* in text extensions. - .github/workflows/security-scan.yml: raise the security job timeout 10m->30m to accommodate very large diffs (the scan itself is unchanged). Verified locally: scan skips the fixture, 0 "ignored null byte" warnings, 0 findings, exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): sweep get-shit-done refs introduced by merging next The branch was updated with next (#614/#384/#618 etc.), which reference the get-shit-done/ dir (still named that on next). Swept the stale references in the merged files to gsd-core so the rename stays consistent and lint:legacy-name passes: - commands/gsd/discuss-phase.md (runtime-launcher shim paths) - src/core.cts (getAgentsDir layout comments) - tests/bug-384-agents-runtime-aware.test.cjs (require path to runtime lib) Verified: guard 0 violations; build green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): exclude gsd-core/ path segments from bug-3683 command cross-ref invariant The #614 runtime-launcher shim added to discuss-phase.md references `${_GSD_RUNTIME_ROOT}/gsd-core/bin/...`. bug-3683's REF_PATTERN excluded path-y refs only via lookbehind, but `}` precedes `/gsd-core/` in the shim, so it mis-read the directory path as a dangling `/gsd-core` command ref (same class as the #604 bug-2954 fix). Added a trailing `(?![\w-]*\/)` so `/gsd-<x>/...` path segments are not treated as slash-command references. Verified locally on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22 image) full suite: 0 failures - bug-3683 + bug-2954 pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): lazily resolve findProjectRoot in gsd-tools (harden flaky CI) CI intermittently failed state.test's gsd-tools subprocess with "findProjectRoot is not a function" (flip-flopping across legs; not reproducible on mac full suite, gsd-test linux full suite, test:unit, or state.test x8). findProjectRoot is a re-export from core.cjs (sourced from project-root.cjs); binding it via destructure at module-load can be undefined under a load-ordering edge. Resolve it lazily at call time via a small wrapper so the lookup happens after core.cjs is fully initialized. Verified green on BOTH platforms before pushing: - mac (node 26) full suite: 0 failures - gsd-test-runner (linux, node22) full suite: 0 failures - state.test.cjs: 106/106; gsd-tools loads cleanly. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#604): allowlist verification-patterns.md placeholder examples in secret scan The rename git-mv'd references/verification-patterns.md into gsd-core/, pulling it into the secret-scan diff. It documents stub/placeholder RED-FLAG env-var examples (illustrative Stripe test-key / database-URL / API-key placeholders) — not real credentials. Added it to .secretscanignore with the strict annotation, mirroring the existing gsd-core/workflows/plan-phase.md exception. Verified locally: secret-scan-lint --strict OK; secret-scan --diff origin/next exits 0 with 0 findings. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c4e39bc23a |
refactor(tests): consolidate graphify Module — 7 files → 1 (#3769)
* refactor(tests): consolidate graphify Module — 7 files → 1 Closes #3761 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(tests): split graphify.test.cjs along describe-block seams — keep files ≤ 800 LOC - tests/graphify.test.cjs (653 LOC): status + build - tests/graphify-query.test.cjs (447 LOC): query - tests/graphify-visualization.test.cjs (577 LOC): staleness + mvp-viz + regressions - tests/graphify-auto-update.test.cjs (625 LOC): auto-update hook - tests/helpers/graphify.cjs (112 LOC): shared helpers extracted Total: 132 tests, 0 failures. Refs #3761. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |