b6e6a22fce822b30bbc27fcde60deaa1b18c29e5
12 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b6e6a22fce |
fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage (seven commits, never pushed) onto current origin/next as a single squashed commit. The original work was substantial and correct; this commit preserves its full scope, trimmed where rebase conflicts + workflow size budgets required it. Three independent layers where response_language was being dropped are closed: Layer 1 — orchestrator-facing directives across workflows. Adds the strong "All user-facing output in this workflow MUST be presented in {response_language}; technical terms, code, paths, and subagent prompts stay in English" directive to ~40 workflows that previously either lacked it entirely (verify-work, new-project, new-milestone, quick, manager, and ~35 more) or carried only the weak subagent-prompt-only form (plan-phase, execute-phase). The directive covers narration between tool calls and banner output, not just the AskUserQuestion prompts. Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts an optional responseLanguage parameter and renders the frame strings ("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.") in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/ Chinese/Korean/Italian) with an alias table covering ~30 input variants (en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads config.response_language via loadConfig(cwd) and passes it through, so the byte-for-byte block verify-work.md reprints verbatim is already localized when written — preserving the anti-injection hygiene rule at verify-work.md (the model is forbidden to translate after the fact). CJK display width is computed by East Asian Width property ranges (W/F) so the right ║ border of the banner stays aligned for full-width characters. English fallback is byte-identical to the pre-fix behavior when response_language is unset or unrecognized. Layer 3 — literal English report templates in execute-phase. The top-of- workflow directive covers all template sites (templates are a structural source, not literal output). Inline render-language notes that previously sat at each template site were removed during the squash because they pushed execute-phase.md over its frozen pre-phase-6 byte ceiling (93600 — ADR-857 Phase 6 capstone). The single top directive covers the same surface with fewer bytes. Also extends src/docs.cts and src/init.cts to propagate response_language into the init JSON bundle of the additional workflows so the directive can read it. Tests added: - tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls back to English default; recognized language swaps only the two frame strings while structural lines stay untouched; CJK display-width regression (independent recomputation of East Asian Width W/F ranges). - tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language wiring through docs.cts/init.cts. References: #2402; reporter's three-layer triage + Layer-4 follow-up; the byte-for-byte anti-injection hygiene rule at verify-work.md (the reason Layer 2 must be renderer-side, not model-translated). This is a squash of the in-flight bot branch — seven commits representing the original implementation plus its subsequent fix/CJK-padding/test/ changeset/regen cycles, none of which were ever pushed or PR'd. The squash captures the final coherent state. * chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md * chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451) Rebase conflicts were entirely in generated artifacts (golden-install-parity fixtures + workflow-size-baseline.json). After taking theirs during rebase, regenerated cleanly against the merged source tree. |
||
|
|
1720aacf0c |
feat(#1949): <precondition> task element — Design by Contract (#2422)
* test(#1949): add failing-first tests for <precondition> element Red phase for issue #1949 (Design by Contract: <precondition> element asserted before task execution). Tests assert: - docs/reference/plan-md.md documents the new <precondition> element - agents/gsd-planner.md @-references planner-preconditions.md and stays under the 49152-char cap (progressive-disclosure requirement) - gsd-core/references/planner-preconditions.md exists and documents the three emission cases mandated by the issue (user_setup / prior-phase artifact / env-var) and the contract triad mapping - agents/gsd-executor.md asserts <precondition> before task execution and routes unmet preconditions through existing checkpoint machinery - cmdVerifyPlanStructure (behavioral via runGsdTools) accepts plans both with and without <precondition> — the additive-validation guarantee - Parity assertion: plan-md.md and planner-preconditions.md agree on the canonical tag spelling (DEFECT.GENERATIVE-FIX-DIVERGENCE guard) Most prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately (regression guards proving the validator already accepts unknown optional tags). * feat(#1949): <precondition> task element — Design by Contract Add an optional <precondition> element to <task> in PLAN.md (issue #1949, The Pragmatic Programmer Topic 23). The front-of-task side of the plan contract — preconditions (before) ↔ postconditions (<verify>/<done>/ <acceptance_criteria>, after) ↔ invariants (must_haves.truths, across the whole plan). Together with the tracer-bullet proposal (#1945), this closes both ends of the 'outrunning your headlights' failure mode for an autonomous AI executor. Acceptance criteria met: - <precondition> is an optional element on <task>; plans that omit it validate unchanged (cmdVerifyPlanStructure checks for presence of required tags, does not reject unknown optional tags). - gsd-executor evaluates the precondition before any other task work. Unmet halts execution with a checkpoint:human-verify and no partial commit; met or absent produces no visible change to execution flow. Unmet is never auto-approved under AUTO_CFG=true — a missing prerequisite is a fact the executor cannot establish on its own. - gsd-planner emits <precondition> in exactly the three cases the issue mandates: user_setup consumption, prior-phase artifact dependency, and env-var/runtime-config dependency. - Tests cover met, unmet, and absent preconditions plus the additive- validator guarantee. Files: - gsd-core/references/planner-preconditions.md (NEW): full emission rules, the three cases with worked examples, format guidance, anti-patterns, the contract triad mapping, and the executor assertion contract. Progressive disclosure. - agents/gsd-planner.md: slim <precondition> note in Task Anatomy with @-reference to the new file. To stay under the 49152-char agent-file cap (27-char headroom before this change), the inline <comment_text_discipline> and <region_scoped_negative_gate> summaries are compressed to one-line pointers — their full rules already live in planner-antipatterns.md, so no content is lost. - agents/gsd-executor.md: new step 0 'Precondition check' in the execute_tasks loop, before the type dispatch, routing unmet through checkpoint_return_format. - docs/reference/plan-md.md: new Preconditions section in the schema reference, with the canonical example and the three emission cases. - CONTEXT.md: Precondition glossary entry as a sibling of Tracer Bullet. - docs/INVENTORY.md + INVENTORY-MANIFEST.json: row for the new references/planner-preconditions.md (regen via gen-inventory-manifest). - tests/precondition-element.test.cjs: failing-first tests covering schema docs, planner emission contract, executor assertion contract, reference-file presence + the three cases, behavioral additive- validator guarantee, and a parity assertion (DEFECT.GENERATIVE-FIX- DIVERGENCE guard). - .changeset/quick-hawks-bark.md: Added fragment. Companion to #1945 (tracer bullets). * chore(#1949): regen agent-size baseline + install-tree goldens Documented baseline regenerations required by the feat(#1949) prose changes (RULESET.AGENT_SIZE_BUDGET + golden-install-parity): - npm run size:baseline — locks in the new gsd-executor.md size (+1050 bytes: the precondition-check step 0 block). gsd-planner.md is net smaller (-142 bytes: compressed two inline summary blocks whose full rules already lived in planner-antipatterns.md to make room for the slim <precondition> pointer). No hard-cap breach. - npm run gen:golden — pick up the new references/planner-preconditions.md + the two changed agent files across all 18 runtime install trees. Both regens are CI-mandated after intentional agent/reference changes; see CLAUDE.md 'RULESET.AGENT_SIZE_BUDGET' and the comments in tests/golden-install-parity.test.cjs. * fix(#1949): bound <precondition> checks to read-only (security review) Apply the security-review finding (LOW, isolated /security-review subagent): the executor's 'run the cheapest check' phrasing for a plan-author-controlled prose line was broader than ideal — a hostile plan author could craft a <precondition> whose 'cheapest check' is side-effecting (curl to an attacker host under the guise of verification, rm -rf before checking, secret emission). The risk is inherited from GSD's existing plan-trust model (<verify>, <action>, <done> already direct the executor to run arbitrary shell), so <precondition> does not materially expand it. But the new prose actively directs execution ('run the check') rather than passively consuming the element, so the bound is worth making explicit. Tightened across all four surfaces that describe the check shape: - agents/gsd-executor.md step 0: 'Verify with read-only checks only — file existence, env var presence (no value output), idempotent GET /health-style pings. Do NOT run commands with side effects (writes, network POSTs, secret emission) as the check; if a side-effecting check seems required, halt and surface via checkpoint instead.' - gsd-core/references/planner-preconditions.md Format section: same bound, plus the halt-and-surface escape hatch. - docs/reference/plan-md.md Preconditions section: mirrored. - CONTEXT.md Precondition glossary entry: mirrored. Regenerated agent-size baseline (executor grew 46186 -> 46440; still under the 49152 cap) and install-tree goldens. * chore(#1949): backfill changeset pr number 2422 Per CONTRIBUTING.md changeset workflow + feature-builder directive Step 8.7: backfill the placeholder pr:0 with the real PR number immediately after gh pr create returns. Avoids the fail_invalid_fragment gate. * fix(#1949): cite [#1949] on allow-test-rule exemption (ADR-456) CI's lint:ci runs lint-allow-test-rule-refs which per ADR-456 requires every // allow-test-rule: exemption on a NEW test file to carry an issue reference (#NNN or URL). My earlier push omitted it. Local 'npm run lint' (eslint) does NOT run this check — only 'npm run lint:ci' does. CLAUDE.md explicitly warns: 'lint:ci ≠ lint — CI runs lint:ci; a local pass is not the gate.' I should have run lint:ci before pushing; correcting now. Pattern matches the companion feature's test file: tests/tracer-bullet.test.cjs:1 // allow-test-rule: source-text-is-the-product [#1945] |
||
|
|
8d2f8bcb23 |
fix(#2388): gate shared requirement completion on sibling plans, revert on gaps (#2424)
* fix(#2388): gate shared-ID requirement marking and revert on gaps_found Adds requirements.ready-ids (execute-plan.md's update_requirements step) so a requirement ID declared by multiple plans in a phase only marks Complete once every declaring plan has produced a SUMMARY.md, and requirements.revert-phase (execute-phase.md's gaps_found branch) so a gaps_found verdict reverts the phase's own prematurely-Complete IDs before the gap report renders. Single-plan IDs still mark immediately. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2388): regenerate fixtures + lint gate-prep * fix(#2388): repair failing tests after gate verification * chore(#2388): add changeset (#2424) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dd5a2211c9 |
enhance(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) (#2416)
* test(#1964): add failing-first semantic-recall contract tests Epic #1957 Phase 3C (final). Source-text-is-the-product contract tests: semantic recall via MemPalace (top-k meaning-similar prior resolutions, catches same-root-cause/different-wording cases), indexing resolved sessions at archive, graceful degradation to keyword matching when MemPalace is absent, knowledge-base.md stays the durable plain-text source of truth, agent Phase 0 / Matching Logic is semantic-first (the stale 'keyword overlap, not semantic similarity' claim must go), and no new embedding/vector infra (reuse MemPalace). Failing-first: reference, the Matching Logic reframe, the Phase 0 consolidation, and the archive indexing step do not yet exist. * feat(#1964): semantic knowledge-base recall via MemPalace (keyword fallback) Epic #1957 Phase 3C (FINAL). Replaces keyword-overlap matching with semantic recall: at Phase 0 the debugger queries MemPalace with the current symptoms and surfaces the top-k meaning-similar prior resolutions, catching the same-root-cause/different-wording cases keyword overlap missed (the self-noted 'keyword overlap, not semantic similarity' limitation). Resolved sessions are indexed into MemPalace at archive (symptoms + root_cause(s) + fix + recurrence guard). knowledge-base.md remains the durable plain-text source of truth; when MemPalace is absent the debugger falls back to keyword-overlap matching (logged, never a silent skip). No new embedding/vector infrastructure — MemPalace is reused. Size-neutral agent edits: the Matching Logic section reframed (keyword-only -> semantic-first + keyword-fallback + @-include); Phase 0's three keyword bullets consolidated into one semantic-first bullet; one MemPalace-indexing step added at archive. Agent at 57222 B (122 B headroom — final phase). Full rules in gsd-core/references/debugger-semantic-recall.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1964): address orthogonal review (invocation mechanism, index Resolution-not-symptoms + redaction, fallback detail) - HIGH: the 'query MemPalace' instruction was WHAT-level only; the agent has no MCP tools. Added an Invocation section naming the Bash CLI (mempalace search --wing <wing>) + MCP-when-registered + wing resolution (config.mempalace.wing -> project_code -> project dir), matching every other MemPalace integration. Without this the feature silently degraded to keyword matching even when MemPalace was present. - MEDIUM (security x2): index the agent-authored Resolution summary (root_cause + fix + recurrence_guard), NOT raw user-supplied Symptoms — excludes attacker-controlled prose from the cross-session index AND reduces secret/PII leakage. Redact secret-shaped values before indexing. Stated the write order (KB append + commit MUST succeed before indexing). - LOW: restored 'identifiers' + 'case-insensitive' to the keyword fallback; added a test asserting the fallback mechanics survived the Phase 0 consolidation (Error patterns field, 2+ token overlap, identifiers, case-insensitive). * chore(#1964): ratchet agent-size baseline downward (leaner archive bullet shrank gsd-debugger.md 57222->57197) * chore(#1964): backfill changeset pr number (PR #2416) |
||
|
|
c67f301867 |
feat(#1963): emit blameless-postmortem Prevention block at resolution (#2410)
* test(#1963): add failing-first prevention/postmortem contract tests Epic #1957 Phase 3B. Source-text-is-the-product contract tests: blameless 5-Whys that BRANCHES per Phase 2A RCA (not a single-cause chain; treats agent error as 'why was that possible?'), the 'why wasn't this caught?' question, the recurrence-guard taxonomy (regression test / assertion / lint rule / KB pattern), the KB-entry why_not_caught + recurrence_guard fields with backward compat, the session-manager prevention summary line, and the Zawinski scope-boundary (a block, not a subsystem). Failing-first: reference, archive_session edit, KB schema extension, and session-manager summary do not yet exist. * feat(#1963): emit blameless-postmortem Prevention block at resolution Epic #1957 Phase 3B. At archive_session the debugger now produces a Prevention block with three blame-free components: a branching 5-Whys causal chain (branches per Phase 2A RCA, not a single chain; 'agent error' prompts 'why was that possible?', never blame), a 'why wasn't this caught?' answer naming the missed gate (test/typecheck/lint/review/verify), and a concrete recurrence guard (regression test / assertion / lint rule / KB pattern). The knowledge-base entry gains two structured fields (why_not_caught + recurrence_guard) so future Phase-0 recall surfaces the prior prevention, not just the prior fix. Additive: old entries without the fields still load. The session-manager compact summary surfaces a one-line prevention summary. Full rules extracted to gsd-core/references/debugger-prevention.md (slim archive_session step + 2 KB fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1963): address orthogonal review (CRITICAL append-template drift + Phase-0 consumption + parity test) - CRITICAL: the archive_session KB append template omitted Why not caught + Recurrence guard (only the Entry Format had them) — the feature's core deliverable silently did not happen. Added both fields to the append template the agent actually follows (nearest-instruction wins). - HIGH: Phase 0 (KB read) only surfaced root_cause + fix; the new fields were dead data. Extended the Phase 0 Evidence line to consume why_not_caught + recurrence_guard when present (absent on old entries — backward compat holds). - MEDIUM: added a cross-section parity test (every Entry-Format field must also appear in the append template — the guard that would have caught the Critical) + a Phase-0-consumption assertion. - MEDIUM: the 'branches per Phase 2A' claim is now wired — reuses reasoning_checkpoint.candidate_causes across the four categories. - MEDIUM: recurrence-guard taxonomy gains type refinement + config-default change; LOW: added 'build' gate to both surfaces for parity. - NIT: compact-summary fallback shape ('no gate existed'); verify the guard artifact exists before recording it. * test(#1963): anchor Phase-0 consumption test on the specific heading The regex /Phase 0[\s\S]{0,1200}/ matched the first 'Phase 0' in the file (in knowledge_base_protocol prose), not the Phase 0 block in investigation_loop. Anchor on '**Phase 0: Check knowledge base**' and widen to 1500 chars. * chore(#1963): backfill changeset pr number (PR #2410) |
||
|
|
36a311c5bb |
enhance(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) (#2409)
* test(#1962): add failing-first repro-hardening contract tests Epic #1957 Phase 3A. Source-text-is-the-product contract tests: PBT shrinking (fast-check/Hypothesis, minimized seed, manual-minimization degradation), the four oracle types (specified/derived/metamorphic/implicit with implicit flagged weakest), boundary neighbors (off-by-one/min-max/empty-singleton tied to the equivalence class), oracle_type in DEBUG Resolution, and the Phase 1A tie-in (minimized seed + real oracle => the mutation guardrail bites). Failing-first: reference, agent cross-refs, and template field do not yet exist. * feat(#1962): harden regression tests (PBT shrinking + oracle classification + boundaries) Epic #1957 Phase 3A. Extends Minimal Reproduction (shrinking) and Test-First Debugging (oracle classification + boundary neighbors): - Shrinking: wrap an input-space failing input in a property (fast-check JS/TS, Hypothesis Python) and store the MINIMIZED counterexample as the regression seed; degrade to manual minimization when no PBT framework is present. - Oracle classification: state specified / derived (contract/model) / metamorphic / implicit (crash, weakest) before writing the assertion; record under Resolution.oracle_type; never default to implicit silently. - Boundary neighbors: off-by-one, min/max, empty/singleton around the fixed defect's equivalence class. Together they turn the regression test into a root-cause check — what the Phase 1A mutation guardrail needs to bite. Full rules extracted to gsd-core/references/ debugger-repro-hardening.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1962): address orthogonal review (bounding, provenance, oracle scope, sufficient-triple) - HIGH: added a 'Bound the property/shrink run' section (60s timeout, degrade- to-manual on timeout, do-not-raise-default-run-limits, argv-not-shell) — the gauntlet violation the sibling references already honored. - Medium: test-provenance caveat (the failing input often comes from the bug report — author the generator from a sanitized description, cross-ref debugger-fix-acceptance.md). - Medium: oracle scope note — the 4 types cover deterministic bugs; non- deterministic failures re-route to stability-stress per bug-taxonomy. - Medium: Phase 1A tie-in corrected — seed+oracle is necessary not sufficient; boundary neighbors close the adjacent-input escape; the sufficient triple is seed+oracle+neighbors. - Low: preserve the original noisy repro as a secondary reference; operationalize 'equivalence class' (the predicate the fix draws). Nit: degradation reworded. * chore(#1962): backfill changeset pr number (PR #2409) --------- Co-authored-by: sim <sim@local> |
||
|
|
6baa2a8182 |
feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger (#2407)
* test(#1961): add failing-first bug-taxonomy routing contract tests Epic #1957 Phase 2B. Source-text-is-the-product contract tests (3 taxonomy classes, explicit class->technique routing table, Bohrbug->repro+SBFL+bisect, Heisenbug->record-replay/stability+SKIP-SBFL, Concurrency->atomicity/order/ deadlock checklist, bug_class in DEBUG Current Focus, supersede-not-append) plus a routing-table specification object pinning the documented decisions (SBFL forbidden on Heisenbug is the load-bearing 1B/2B seam). Failing-first: reference, Phase 1.75, and routing-table reframe do not yet exist. * feat(#1961): add bug-taxonomy classification + strategy routing to gsd-debugger Epic #1957 Phase 2B (reliability-critical). Adds Phase 1.75: classify the failure as Bohrbug / Heisenbug-Mandelbug / Concurrency, then route the investigation technique via an explicit class->technique table (Kernighan: no opaque heuristic). Bohrbug -> reproduction + SBFL (Phase 1.25) + git bisect; Heisenbug/Mandelbug -> record-replay (rr) + stability-stress + statistical sampling, with SBFL explicitly SKIPPED (a flaky spectrum poisons the Ochiai ranking — the load-bearing 1B/2B seam); Concurrency -> the atomicity/order/deadlock checklist first. Reframes (supersedes, not appends — Zawinski) the flat 'Technique Selection by situation' table into a class-routed table; the 11 techniques remain as routed targets. bug_class recorded in Current Focus (DEBUG template); common-bug- patterns catalog cross-referenced to the taxonomy. Full rules extracted to gsd-core/references/debugger-bug-taxonomy.md. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * fix(#1961): address orthogonal review (phase-name drift, General lane, revoke framing, row-scoped tests, bounding) - HIGH: reference said 'Phase 1B' (epic shorthand); corrected to the deployed 'Phase 1.25' (matches the agent + SBFL reference). - HIGH: 6 of 11 techniques (Rubber duck, Delta, Working backwards, Differential, Comment-out, Follow-the-indirection) were orphaned by the situation-table reframe. Added a 'General (any class, situation-cued)' lane to BOTH the reference routing table and the agent's Technique Selection table that re-homes them — supersede-not-append now holds. - MEDIUM: the SBFL-skip is structurally retroactive (Phase 1.25 runs before Phase 1.75 classification), so reframed the table column from 'Do NOT use' to 'Revoke if already run' + an explicit 'retroactive revocation, not proactive skip' note stating the ordering honestly. - MEDIUM: contract tests are now row-scoped (parse the table by class, assert per-row) instead of presence-only; added a guard that the previously- orphaned techniques now have a General-lane route. - LOW: pinned the canonical bug_class value form (lowercase-kebab: bohrbug|heisenbug-mandelbug|concurrency; prose may use title-case). - NIT: added a 'Bound the Heisenbug-chase runs' note (rr/stability/sampling timeouts) per the unbounded-subprocess gauntlet. * chore(#1961): backfill changeset pr number (PR #2407) |
||
|
|
f8b16d1874 |
enhance(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger (#2405)
* test(#1960): add failing-first RCA-branching contract + schema-invariant tests Epic #1957 Phase 2A. Source-text-is-the-product contract tests (fishbone >=2 categories, AND-gate, multi-cause root_cause, backward compat, reasoning checkpoint candidate_causes+and_gate fields, debugger-philosophy single-cause note, DEBUG template) plus behavioral schema-invariant checks on two fixtures: two contributing causes (AND-gate yes) -> both recorded; single-cause (AND-gate no) -> one root_cause, identical to today. Failing-first: reference, agent edits, and template note do not yet exist. * feat(#1960): add RCA branching (fishbone + AND-gate) to gsd-debugger Epic #1957 Phase 2A. Guards against 5-Whys single-cause bias: before committing root_cause, the debugger enumerates candidate causes across >=2 Ishikawa categories (code/config/environment/data) and explicitly answers an AND-gate question. When the AND-gate fires, every contributing cause is recorded, so a multi-cause fix no longer recurs via the unaddressed second cause. Resolution.root_cause may hold one OR a small set (additive; single-cause sessions are byte-identical to today). The Structured Reasoning Checkpoint gains candidate_causes + and_gate fields; debugger-philosophy.md adds the single-cause-bias trap. Full rules extracted to gsd-core/references/debugger-rca-branching.md (slim Phase 2 routing + 2 checkpoint fields kept in the agent). INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md + DEBUG template updated. * fix(#1960): address orthogonal review (AND-gate self-consistency, parity guard, narrowed claim, ripples) - Reference: the collapse rule now enforces AND-gate self-consistency — and_gate=yes with a single confirmed cause is flagged as incomplete (return to Phase 3); a race/timing note clarifies such bugs bridge categories; the 'byte-identical' backward-compat claim narrowed to 'root_cause shape unchanged; reasoning_checkpoint gains 2 fields in every session'. - DEBUG.md: stale 'five-field' mirror prose -> seven-field (parallel-surface drift the reviewer flagged); new debug-session-management parity test pins the field-count claim to the gsd-debugger.md YAML keys (CRLF-safe). - Scalar-assuming consumers of set-valued root_cause updated: session-manager compact summaries (319/332), diagnose-only return (1062), archive entry (1216), ROOT CAUSE FOUND return (1322). - Test: added the AND-gate-yes/single-cause invariant + fixture; rephrased the fixture describe block honestly as a schema-invariant specification. - Phase 2 bullet phrasing clarified ('at hypothesis formation, before the Phase 4 commit'). * test(#1960): parity regex accepts word-form count ('seven-field' or '7-field') * test(#1960): parity regex counts array-valued YAML keys (no inline value) * chore(#1960): backfill changeset pr number (PR #2405) |
||
|
|
56a5c6404c |
feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger (#2403)
* test(#1959): add failing-first SBFL contract + Ochiai correctness tests Epic #1957 Phase 1B. Source-text-is-the-product contract tests (Ochiai formula documented, Tarantula fallback, top-N seeding, no-coverage skip logged, ranking->Evidence, Bohrbug gating) plus a behavioral Ochiai formula-correctness section: bound [0,1], max-score invariant, a known-fault fixture proving the fault ranks #1 (criterion 2), clean degradation on zero failing tests, and two fast-check properties. Failing-first: reference file and agent routing do not yet exist. * feat(#1959): add spectrum-based fault localization (Ochiai) pre-filter to gsd-debugger Epic #1957 Phase 1B. When a runnable test suite with per-test coverage exists (>=1 failing AND >=1 passing test), the debugger computes an Ochiai suspiciousness ranking over the coverage spectrum and seeds the top-N suspicious locations into Evidence as first-class hypothesis candidates, narrowing the search space deterministically before LLM reasoning. Tarantula documented as fallback. Degrades cleanly (logged, never silent) when there is no test suite, no failing tests, or no per-test coverage, and is explicitly not trusted on flaky/Heisenbug spectra (pairs with Phase 2B bug-taxonomy). Full rules extracted to gsd-core/references/debugger-sbfl.md (slim Phase 1.25 routing kept in the agent to respect the size cap). No new coverage framework — reuses the project's existing test/coverage runner. INVENTORY + manifest + agent-size baseline + install-parity goldens + AGENTS.md updated. * test(#1959): bound property generators to valid coverage counts The [0,1] property generated failedExec independently of totalFailed, but Ochiai's score is only bounded by 1 under the coverage invariant failedExec <= totalFailed (a failing test that executed s is one of the totalFailed failing tests). Out-of-domain inputs (failedExec=100, totalFailed=5) make the formula correctly return >1. Bound failedExec by totalFailed via fc.chain so the property tests the real domain. Also cleaned up the ranking property (removed dead code). * fix(#1959): address orthogonal review (monotonicity property, degradation row, coverage bounding) - Replace vacuous ranking property (true-by-sort-construction) with a non-trivial monotonicity property: holding totalFailed + passedExec fixed, ochiai is non-decreasing in failedExec. An inverted formula would fail it. - Add the missing 'no passing tests' degradation row (preconditions require >=1 passing test; Tarantula would divide by totalPassed=0). - Bound the coverage subprocess (CLAUDE.md gauntlet): cap the coverage run, degrade-to-skip on timeout, never hang the debug session. - Reword 'discard the ranking' -> 'mark the Evidence entry as revoked (do not delete)' per Kernighan auditability. * test(#1959): bound monotonicity-property generator to valid coverage (failedExecA <= totalFailed) * chore(#1959): backfill changeset pr number (PR #2403) |
||
|
|
5e52350736 |
feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger (#2396)
* test(#1958): add failing-first guardrail contract tests Epic #1957 Phase 1A. Adds source-text-is-the-product tests asserting the 5-signal fix-acceptance guardrail contract (target test, mutation check, no-op/deletion detector, adjacent tests, revert-and-reconfirm), graceful degradation, FIX REJECTED BY GUARDRAIL return path, per-signal debug-file recording, and subprocess bounding. Failing-first: reference file and agent sections do not yet exist. * feat(#1958): add multi-signal fix-acceptance guardrail to gsd-debugger Epic #1957 Phase 1A. Prevents accepting a fix that merely greens the test (Goodhart defense / APR overfitting). Adds a 5-signal gate run before fix acceptance: target test, mutation check (Stryker), no-op/behavior-deleting detector, adjacent/held-out tests, revert-and-reconfirm. Degrades gracefully when Stryker or a test suite is absent (each skip logged, never a silent pass), records per-signal results under Resolution.verification, and returns a FIX REJECTED BY GUARDRAIL outcome the session-manager surfaces for revise / accept-as-debt / abandon. Full rules extracted to gsd-core/references/debugger-fix-acceptance.md (slim routing kept in the agent to respect the agent-size cap). Debug template + INVENTORY + manifest + agent-size baseline + AGENTS.md updated. * test(#1958): correct newline-tolerant assertion + regen install-parity goldens The revert-and-reconfirm assertion collapsed whitespace before matching so markdown line-wrapping does not break it. Regenerated the golden-install-parity and install-tree fixtures (npm run gen:golden) to absorb the intentional gsd-debugger.md / gsd-debug-session-manager.md / DEBUG.md / new reference-file changes to the installed artifact tree. * fix(#1958): tighten guardrail per orthogonal review Addresses the isolated reviewer's findings: - signal 5 now states its recorded-repro dependency and routes the no-repro case to the degradation row; revert mechanism specified (git stash / git revert -n); minimality flag tied to diff structure, not revert-ability. - bounded-subprocesses section now bounds the git subprocess (5-30s) too, requires argv-array argument passing, and scopes Stryker to the driving regression test (a mutant killed only by a non-driving test is a finding). - new test-provenance (security) clause: the driving test must be agent-authored; bug-report repro scripts are DATA, never executed verbatim. - tightened 3 contract assertions to bind to specific clauses (guardrail_verdict field, deletion-reject-unless-RCA, 60s+git bounding). - Goodhart framing softened to 'partially-independent'; DEBUG.md template verification field notes the nested map shape. * chore(#1958): backfill changeset pr number (PR #2396) * fix(#1958): add issue ref to allow-test-rule annotation (ADR-456) CI lint-allow-test-rule-refs requires every allow-test-rule exemption to carry a 'see #NNN' issue ref per ADR-456. The new test file's annotation lacked it; this adds (see #1958). |
||
|
|
9ad2bab4be |
fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime (#2332)
* fix(#2297): scope resolve_model_ids:"omit" to the resolving runtime The installer writes resolve_model_ids:"omit" for non-alias runtimes into the machine-wide ~/.gsd/defaults.json (#1156); any runtime read it back, so install order silently flipped Claude's adaptive tier aliases (executor->sonnet, planner->opus) to '' in no-project sessions. Resolution is now scoped to the runtime actually resolving, identified by a new per-install <install>/gsd-core/.gsd-runtime marker (installer writes it beside VERSION). The "omit" branch returns '' only when the PROJECT explicitly set omit (honored for all runtimes, #2517 finding #4) OR the active runtime lacks native aliases. Claude ignores a global-defaults-only omit and keeps its aliases; the active runtime is canonicalized (GSD_RUNTIME -> config.runtime -> marker -> claude) so alias/case spellings can't defeat the check; explicit project omit is workstream/project-scope aware; explicit true still materializes IDs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#2297): backfill PR number 2332 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
89b1bef881 |
refactor(#2267): golden-parity file-set snapshot + anti-staleness CI selection (#2274)
Phase 2 of golden-parity redesign (epic #2264). Adds an install file-set snapshot (golden-install-tree) and a ci-test-scope rule selecting golden-parity whenever any installed-source path changes, closing the silent-staleness hole behind the #2266 red. ADR-2264 amended (the copy/transform split premise was unsound). Closes #2267. |