aa0f7dee99e7b4e893e0dc41bfe0beced10ff6f2
4722 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
aa0f7dee99 |
fix(#2460): pi before_provider_request fail-opens without explicit model_profile_overrides (#2499)
* fix(#2460): pi before_provider_request fail-opens without explicit override pi/gsd.cjs's buildBeforeProviderRequestHandler unconditionally rewrote payload.model to the built-in pi/sonnet tier default (claude-sonnet-5) via resolveTierEntry's catalog fallback. For any pi user on a non-Anthropic provider (kimi-coding, zai, openrouter, openai-codex, minimax, ...), this silently broke every request: pi's chosen model was replaced with one the active provider did not know. The fix inspects model_profile_overrides.pi[tier] explicitly BEFORE calling resolveTierEntry (which falls back to the built-in catalog and would mask the 'user did not opt in' signal). When the user has not set an override (or set it to null), the handler returns undefined — fail-open — and pi's chosen model flows through untouched. Only an explicit opt-in via model_profile_overrides.pi[tier] steers. Tests: - the ACTUALLY-REGISTERED handler fail-opens when no override configured (was: 'steers to default-tier model-catalog pi id' — encoded the bug). - new test reproducing the reporter's exact repro (model: 'k3' → undefined). - override path: explicit model_profile_overrides.pi.sonnet config steers to the user-configured model id. - defensive: explicit null override also fail-opens. Per the reporter's suggested fix #1 of #2460. * test(#2460): regen pi golden parity + install tree fixtures pi/gsd.cjs changed → pi install hash changed → regenerate the parity + install-tree fixtures via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1. * fix(#2460): treat empty-string override as fail-open (M1 review) Per code-review M1 + security M1: an explicit empty-string override (`{ pi: { sonnet: "" } }`) silently bypassed the fail-open guard because the check was `=== undefined || === null` only. resolveTierEntry's falsy `if (userRaw)` then fell back to the built-in catalog and rewrote payload.model to claude-sonnet-5 — re-introducing the exact bug this PR fixes, via a degenerate config shape. Fix: widen the guard to also reject `''`. The test now exercises both null and '' in a loop, asserting fail-open for both. * fix(#2460): clear hono/@hono/node-server moderate advisories via npm override GHSA-v422-hmwv-36x6-class advisories (3 moderate) appeared during this PR's session: - @hono/node-server <2.0.5 (path traversal on Windows via encoded paths) - hono 4.3.3 - 4.12.26 (API Gateway v1 adapter drops distinct repeated request header values) - both transitively via @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol /sdk@1.29.0 The npm audit fix re-resolved hono to 4.12.31 (within the existing ^4.11.4 range declared by MCP SDK), clearing the hono advisory without an override. The @hono/node-server advisory cannot be re-resolved the same way: MCP SDK pins @hono/node-server@^1.19.9, and the fix requires 2.0.5+. There is no MCP SDK release that allows @hono/node-server@2.x (latest 1.29.0 is the most recent), and bumping @anthropic-ai/claude-agent-sdk to 0.3.x does not help (its peerDependency is still @modelcontextprotocol/sdk@^1.29.0). The override (sibling to the existing 'qs' and 'body-parser' entries) is therefore the only available tool — distinct from the body-parser case in re-resolution. Verified: npm audit --omit=dev reports 0/0/0/0 advisories. Also: regenerated the pi golden-install-parity fixture (the pi/gsd.cjs change in this PR altered the pi install hash). * docs(changeset): add Fixed fragment for #2460 PR The two-otters-jog.md changeset was created earlier but lost during the cherry-pick detour to fix #2454 PR 1's npm advisory cascade. Recreating here with PR number 2499 backfilled (no placeholder cycle needed). |
||
|
|
d579daa3ed |
docs(#2505): Phase 6 — migration guide + built-in-only subagent-toolkit enum (#2538)
* docs(#2512): Phase 6 — migration guide + built-in-only subagent-toolkit enum * fix #2512: update CONTRACT-PIN for built-in-only subagentToolkit value * docs(changeset): backfill PR #2538 for Phase 6 (#2512) |
||
|
|
28f79e6e22 |
feat(#2505): Phase 5 — Kimi variant install-time disambiguation (#2535)
* feat(#2513): Phase 5 — Kimi variant install-time disambiguation (descriptions + mismatch warning) * fix #2513: use helpers.cleanup (Windows-EBUSY retry budget) not raw fs.rmSync * fix #2513: spawnSync captures stderr (execFileSync drops it on exit-0) * docs(changeset): backfill PR #2535 for Phase 5 (#2513) |
||
|
|
f654c24a3e |
feat(#2505): Phase 4 — runtime-aware subagent dispatch (Option A; resolve-dispatch-type query) (#2525)
* feat(#2508): Phase 4 Option A — runtime-aware subagent dispatch via resolve-dispatch-type query (#2505) * fix(#2508): prose-variant preamble (avoid scanner-tripping literals) + namedDispatch===false-only mapping * fix(#2508): remove leftover old-preamble lines (keep prose variant only) * fix #2508: prose-only reference file * test #2508: regen golden install parity after workflow preamble additions * fix #2508: remove preamble from plan-phase.md (Phase 6 capstone ceiling); regen size+golden baselines * docs(changeset): backfill PR #2525 for Phase 4 (#2508) |
||
|
|
936a345381 |
feat(#2505): Phase 3 — agent-skills fallback for non-dispatchable runtimes (#2521)
* feat(#2454): PR 2 — cmdAgentSkills fallback reads installed agent prompt When no agent_skills config entry exists for a given agent type (the common case on AGENTS-native runtimes), cmdAgentSkills previously returned empty output. Workflows that inject ${AGENT_SKILLS_*} into subagent dispatch prompts then carried nothing — the persona was lost. The fallback: resolve the runtime's agents directory via checkAgentsInstalled and read <agentsDir>/<agentType>.md. The installed agent prompt content (now present for kimi-code via the flat-skills install layout) flows into the dispatch prompt so the persona survives even without explicit config opt-in. This is the reporter's suggested fix #2 from #2454. The fallback triggers for ALL runtimes (not just kimi-code) when no config entry exists — it is strictly additive (returns content the previous empty path could not). If the agent file is not found on disk, the block stays empty (same as before). * docs(changeset): Phase 3 agent-skills fallback Added (#2510) * docs(changeset): backfill PR #2521 for Phase 3 (#2510) |
||
|
|
c2a305c44d |
feat(#2505): Phase 2 — kimi-code Agent Skills install layout (#2520)
* feat(#2454): PR 2 — kimi-code Agent Skills converter + install layout PR 1 registered the kimi-code EoS descriptor with empty artifactLayout (SKIP_INSTALL_CONTRACT excluded it from the end-to-end install test). PR 2 fills in the install surface: - src/runtime-artifact-conversion.cts: new convertClaudeCommandToKimiCodeSkill function. Today it delegates to convertClaudeCommandToKimiSkill (Python kimi-cli) because Kimi Code uses the same Agent Skills format + /skill: invocation per official docs. The distinct function name lets a future divergence land cleanly if Kimi Code's skill format evolves independently. - gsd-core/bin/lib/capability-validator.cjs: add to ALLOWED_SKILLS_CONVERTERS. - capabilities/kimi-code/capability.json: artifactLayout.global now declares the skills kind with converter='convertClaudeCommandToKimiCodeSkill' + home='.kimi-code' (auto-discovered at ~/.kimi-code/skills/ per Kimi Code docs: merge_all_available_skills = true default). - tests/installer-migration-install.integration.test.cjs: REMOVE the SKIP_INSTALL_CONTRACT exclusion — kimi-code now has a full install surface. - Regenerated capability-registry + capability-matrix + golden install parity + install tree fixtures for kimi-code. * fix(#2454): wire kimi-code converter into SKILLS_CONVERTER_REGISTRY + count bump - src/install-engine.cts: add convertClaudeCommandToKimiCodeSkill to SKILLS_CONVERTER_REGISTRY so the layout-driven skills install path can dispatch off the descriptor's converter string. - tests/capability-registry.test.cjs: bump VALID_CONVERTER_NAMES count 26 → 27 (added convertClaudeCommandToKimiCodeSkill). * fix(#2454): remove home override from kimi-code skills (inherit configDir) The home:'.kimi-code' override made the install plan resolve skills dest to ~/.kimi-code/skills instead of <configDir>/skills, causing the test's temp configDir to miss the install. Removing it lets skills inherit configDir like most runtimes. * fix(#2454): kimi-code install contract surface is flat-skills (no agents) Kimi Code has NO custom named subagents (per official docs: 3 built-in coder/explore/plan only). The kimi-skills-agents surface expects agents/ gsd.yaml + subagents/*.yaml which kimi-code does not produce. Changed to flat-skills which only checks for skills/gsd-* dirs. * docs(changeset): Phase 2 kimi-code install layout Added (#2509) * docs(changeset): backfill PR #2520 for Phase 2 (#2509) |
||
|
|
bf8f320083 |
feat(#2505): Phase 1 — EoS descriptor split (kimi-code capability.json + drift-guard registration) (#2519)
* feat(#2454): add kimi-code as an EoS capability (Node Kimi Code CLI) PR 1 of N for #2454. Establishes the EoS descriptor foundation for splitting GSD's kimi support into two distinct products per the user's directive: - kimi (existing): Moonshot's Python kimi-cli (~/.kimi, runtime: python) - kimi-code (new): Moonshot's Node Kimi Code CLI (~/.kimi-code, runtime: node, KIMI_CODE_HOME env) Per ADR-1239 EoS, runtime behavior is driven by capabilities/<id>/capability.json descriptors, not hardcoded branches in install.js. The new descriptor uses the existing primitives (dot-home configHome, skills artifactLayout, kimi-hooks-toml hooksSurface — same TOML [[hooks]] format Kimi Code reads per its docs). Critical Kimi Code constraint reflected in the descriptor: hostIntegration.dispatch.namedDispatch: false hostIntegration.dispatch.builtInSubagents: ['coder', 'explore', 'plan'] hostBehaviors.namedSubagentsSupported: false Kimi Code's official docs confirm only 3 built-in subagents with NO custom- subagent registration (the [subagent] table only has timeout_ms). The kimi-agents YAML layout (used by Python kimi-cli) is therefore NOT in kimi-code's artifactLayout. Schema adjustments: - subagentToolkit set to 'undocumented' (the existing escape hatch); the schema enum (full/read-only) lacks a 'limited'/'built-in-only' value. A follow-up PR can extend the schema enum to add 'built-in-only' as a first-class axis value reflecting Kimi Code's documented model. Registration: - capabilities/kimi-code/capability.json (new descriptor, modeled on codex) - bin/install.js: allRuntimes array + --all list + --kimi-code flag - gsd-core/bin/shared/runtime-aliases.manifest.json: kimi-code aliases (kimi-code, kimicode, kimi_code) - src/runtime-name-policy.cts: FALLBACK_ALIASES map - gsd-core/bin/lib/capability-registry.cjs: regenerated via scripts/gen-capability-registry.cjs --write Tests: - tests/multi-runtime-select.test.cjs updated for the new runtime count (18) + new --kimi-code flag test + 'All' shortcut renumbered 18 → 19. Out of scope for PR 1 (follow-up PRs in the sequence): - Install-time decision logic (kimi vs kimi-code detection / prompt) - agent-install-check semantics for kimi-code (verify Agent Skills presence) - cmdAgentSkills fallback returning subagent prompt content - Workflow template mapping (named agents → built-in coder/explore/plan) - Migration guidance for users currently on 'kimi' who are actually on Kimi Code - Schema enum extension for subagentToolkit: 'built-in-only' Refs #2454, #2095 (EoS/kimi migration epic), ADR-1239 (EoS). * fix(#2454): complete drift-guard registrations for kimi-code runtime The drift guards caught every surface that pins runtime enumeration. Each update is mechanical, driven by the guard's named failure mode: - src/runtime-name-policy.cts RUNTIME_LABELS: 'Kimi Code' label for kimi-code - src/runtime-name-policy.cts RUNTIME_FLAG_IDS: add kimi-code to the isKimiCode predicate generator - bin/install.js runtimeMap: option '11' → 'kimi-code', renumber downstream entries (11..17 → 12..18), ALL_RUNTIMES_OPTION 18 → 19 - gsd-core/bin/shared/model-catalog.json runtimeTierDefaults: kimi-code entry (null/null/null — same as kimi, no model tier defaults until configured) - docs/reference/capability-matrix.md: regenerated via scripts/gen-capability-matrix.cjs --write (kimi-code row added) - tests/global-config-home-fragment.test.cjs GOLDEN_FRAGMENT_MAP: kimi-code → '.kimi-code' - tests/fixtures/golden-install-parity/*.json: regenerated via npm run gen:golden (the runtime-aliases.manifest.json hash changed; all 17 runtime fixtures updated) The capability-registry is already regenerated from the prior commit. * test(#2454): update drift-guard tests for kimi-code runtime registration Multiple drift guards pin runtime enumeration counts and option numbering. Each update is mechanical, driven by the guard's named failure mode: - tests/runtime-flags.test.cjs: EXPECTED_FLAGS gains isKimiCode (16 → 17); 'all 16 flags' → 'all 17 flags' in test names + messages. - tests/multi-runtime-select.test.cjs: parseRuntimeInput option renumbering cascade — kilo moves 11→12, opencode 12→13, pi 13→14, qwen 14→15, trae 15→16, windsurf 16→17, zcode 17→18, All 18→19. New single-choice test for kimi-code (option 11). Prompt test updated for new numbering. - tests/host-integration-descriptors.test.cjs: EXPECTED_PROFILES gains kimi-code → 'programmatic-cli' (terminal CLI per Kimi Code docs); EXPECTED_FLATTEN gains kimi-code → false (backgroundDispatch:true per docs, same as Python kimi/opencode). - tests/global-config-home-fragment.test.cjs: table-count test renamed 13 → 14 table runtimes (kimi-code added to GOLDEN_FRAGMENT_MAP earlier). * fix(#2454): empty artifactLayout for kimi-code (PR 1 scope) The skills kind requires a converter (existing converters are per-runtime like convertClaudeCommandToKimiSkill). PR 1 of this multi-PR sequence only registers the descriptor; the actual Agent Skills converter (and a new 'convertClaudeCommandToKimiCodeSkill' function) lands in PR 2 alongside the install-time decision logic. Empty artifactLayout.global is valid and means 'nothing to install yet via the layout seam'. Also: added kimi-code to RUNTIME_META in tests/helpers/install-shared.cjs (localDir .kimi-code, globalSuffix .kimi-code), and added Kimi Code as option 11 in install.js's buildRuntimePromptText (renumbered downstream options 11..17 → 12..18, All 18 → 19). * fix(#2454): camelCase runtimeFlags for hyphenated ids (kimi-code → isKimiCode) The runtimeFlags generator previously produced 'isKimi-code' (hyphen preserved) for the new kimi-code runtime id. Property names with hyphens are awkward for consumers (flags['isKimi-code'] instead of flags.isKimiCode). The new runtimeIdToFlagName helper folds -[a-z] boundaries to uppercase, producing the conventional PascalCase flag name. The 16 prior single-word runtime ids are unaffected (the regex finds no hyphens). * fix(#2454): update remaining drift-guard tests + gen kimi-code fixtures - tests/runtime-flags.test.cjs drift guard: use proper kebab-case conversion (isKimiCode → kimi-code, not 'kimicode') so the registry comparison doesn't false-positive on hyphenated runtime ids. - tests/multi-runtime-select.test.cjs: fix kilo/opencode/pi/qwen/trae single-choice tests for the renumbered options (kilo 11→12, opencode 12→13, pi 13→14, qwen 14→15, trae 15→16). - tests/install.test.cjs: Kilo integration option 11→12, prompt test regex updated. - tests/fixtures/golden-install-parity/kimi-code.json + install-tree/ kimi-code.json: generated via UPDATE_GOLDEN=1 + UPDATE_INSTALL_TREE=1. The kimi-code install produces the standard GSD install layout (skills, contexts, references, etc.) — 436 paths, same shape as other runtimes that have no custom converter yet. * fix(#2454): add kimi-code install contract + global config home fragment - src/runtime-name-policy.cts GLOBAL_CONFIG_HOME_FRAGMENTS: add kimi-code → '.kimi-code' so getGlobalConfigHomeFragment returns the correct path instead of falling through to the default '.claude'. - tests/installer-migration-install.integration.test.cjs RUNTIME_INSTALL_CONTRACTS: kimi-code entry (same surface as kimi for PR 1; PR 2 will specialize once the Agent Skills converter lands). - tests/multi-runtime-select.test.cjs: fix space-separated-choices test for the renumbered kilo option (11 → 12). - tests/fixtures/golden-install-parity/kimi-code.json + install-tree/ kimi-code.json: regenerated after rebasing onto current next (new planner-reversibility.md from #2471 etc. now included). * test(#2454): skip kimi-code install contract until PR 2 ships install layout The end-to-end install test (tests/installer-migration-install.integration .test.cjs) asserts every allRuntimes entry installs a runtime-specific artifact surface. PR 1 of #2454 registers kimi-code in allRuntimes + the capability descriptor + flags + labels, but the install LAYOUT (Agent Skills converter + global AGENTS.md at $KIMI_CODE_HOME/AGENTS.md) lands in PR 2. The SKIP_INSTALL_CONTRACT set marks this exclusion explicit and self-removing — PR 2 removes the entry alongside adding the install surface, restoring the contract loop to full coverage. * fix(#2454): restore compact model-catalog.json format (M1 review) Per code-review M1: my prior 'fix(#2454): complete drift-guard registrations' commit used python json.dump(indent=2) which inflated the file from 165→607 lines (every nested entry got expanded) and lost the trailing newline. The semantic change was just a 3-line kimi-code entry. Restored the original hybrid format (top-level indent=2 + inner entries' one-line style) and added kimi-code in matching form. Regenerated golden install parity + install tree fixtures since the model-catalog.json hash changed. * fix(#2454): update CONTEXT.md allRuntimes glossary (17 → 18, add kimi-code) CI lint-tests job failed on the glossary drift guard (scripts/check-glossary-refs.cjs --check): ✗ CONTEXT.md's allRuntimes enum-count sentence claims 17 values but bin/install.js's allRuntimes array has 18. ✗ CONTEXT.md's allRuntimes member list has drifted from bin/install.js (missing from CONTEXT.md's list: kimi-code). Missed in the prior commits because gsd-test does not run the glossary check (it's a CI lint-tests-only check). Updating CONTEXT.md's two claims to 18 values + kimi-code in the member list. * chore(#2505): regen capability-registry + stamp kimi-code version 1.8.0 (#2511) * docs(changeset): Phase 1 kimi-code runtime Added (#2511) * test(#2511): regen kimi-code golden parity fixture after Phase 0 guard normalization lands * docs(changeset): backfill PR #2519 for Phase 1 (#2511) |
||
|
|
7e905aa137 |
feat(#2505): Phase 0 — Kimi PreToolUse guard vocabulary normalization (precondition; carries PR #2326 forward) (#2518)
* fix(#2304): normalize Kimi tool vocabulary in PreToolUse guard payload checks
The Kimi [[hooks]] registrations translate the matcher to Kimi's tool
vocabulary (WriteFile|StrReplaceFile) but the guard scripts early-exit
unless the payload's tool_name is a Claude name (Write/Edit/MultiEdit),
so every guard was dormant on Kimi: the matcher fired, the script saw
WriteFile, and exit(0)'d.
Normalize the payload's tool_name at the top of each guard
(WriteFile -> Write, StrReplaceFile -> Edit; bare or module-qualified
kimi_cli.tools.file:* forms) before the check. Inlined per guard rather
than a hooks/lib/ helper because hook scripts are staged as standalone
files on every hook surface, and a sibling require is a staging
dependency that can fail silently.
Regression tests pipe Kimi-vocabulary payloads at each guard and assert
it engages (typed fields: exit status, decision, hookSpecificOutput) —
verified red against the pre-fix scripts, green after.
* fix(#2304): normalize Kimi tool_input fields and route block reasons to stderr
Cross-AI review of the initial fix, verified against kimi-cli source,
found the tool_name normalization alone leaves the guards dormant on a
real Kimi runtime: kimi-cli forwards tool_input verbatim
(src/kimi_cli/hooks/events.py), and its tool schemas
(src/kimi_cli/tools/file/{write,replace}.py) use path/content and
edit.old/edit.new (single Edit or list) — not Claude's
file_path/old_string/new_string. The guards read file_path, got '',
and exited 0 past the now-open tool_name gate.
Extend the per-guard normalization to the payload fields
(path -> file_path, edit -> old_string/new_string with list flattening),
and write the worktree guard's block reason to stderr as well as the
stdout JSON — Kimi feeds stderr, not stdout, back to the model on
exit 2 (docs/en/customization/hooks.md exit-code table).
Regression tests rewritten to Kimi's actual payload shapes (plus an
edit-list case and a stderr-reason assertion) — verified red against
the name-only fix, green after.
* fix(#2304): join all edit[] entries into old_string, matching new_string
Review nit on #2326: old_string took only edits[0].old while new_string
joined the whole list. Symmetric join removes the latent trap for any
future consumer sizing before/after content (e.g. the #2255 write guard).
* fix(#2304): normalize Kimi ReadFile vocabulary in read-injection scanner
Review Major 2 on #2326: gsd-read-injection-scanner.js had the identical
dormancy — its Kimi matcher fires on 'ReadFile' but the SCANNED_TOOLS
check only knew 'Read', so injected content in read files was never
flagged on Kimi installs.
Folds the same inlined normalization block into the scanner and extends
the shared KIMI_TOOL_NAMES map with ReadFile:'Read' in all four copies so
they stay byte-identical. Harmless in the three write guards: a
normalized 'Read' falls out of their Write/Edit allowlist exactly as the
unmapped name did. Field mapping verified against kimi-cli upstream
(src/kimi_cli/tools/file/read.py Params.path); the existing
path->file_path copy covers the scanner's file_path read.
* test(#2304): parity test binding the four inlined Kimi normalization copies
Review Major 1 on #2326: KIMI_TOOL_NAMES + normalizeKimiPayload is
deliberately inlined in four hook scripts (staging-dependency rationale,
unchanged), with the inverse table in bin/install.js — five
hand-maintained surfaces and nothing binding them.
Static binding, zero runtime coupling:
- the four inlined blocks must be byte-identical;
- each guard-map entry must be the value-inverse of
convertKimiToolName() for its Claude name;
- every guard-relevant Claude tool (Write/Edit/MultiEdit/Read) must have
a reverse entry — a vocabulary rename or extension that updates the
installer without updating the guards now fails in CI instead of
leaving a guard silently dormant (the #2304 recurrence door).
Negative-controlled: diverging one copy or dropping a map entry fails
the suite against the fixed code.
* test(#2304): regenerate golden parity fixtures for guard hook changes
CI red on #2326: all 10 golden-parity failures were the staged guard
hooks drifting from their fixtures. Regenerated with npm run gen:golden
(after npm run build) under throwaway HOME/CLAUDE_CONFIG_DIR; diff
verified to change exactly the four PR-touched guard entries per
surface, nothing else.
* test(#2304): regression tests for Kimi ReadFile engaging the scanner
Mirrors the per-guard Kimi vocabulary tests the PR added for the three
write guards: bare and module-qualified ReadFile produce the advisory,
path exclusions still apply post-normalization, unknown Kimi names stay
fail-open. Negative-controlled against the pre-fold scanner (the two
positive cases fail there; exclusion/fall-through correctly pass on
both sides).
* fix(#2304): normalize Kimi Shell vocabulary in workflow guard
Withdraws the disclosed out-of-scope split: verification showed the
Bash->Shell case needs NO different mapping — kimi-cli's Shell.Params
names its field `command` (src/kimi_cli/tools/shell/__init__.py), same
as Claude's Bash — and the guard's write branch (Write/Edit/MultiEdit
allowlist) was ALSO dormant on Kimi under its Shell|WriteFile|
StrReplaceFile matcher. Same defect class as the other four hooks.
Folds the identical inlined block into gsd-workflow-guard.js and
extends the shared map with Shell:'Bash' in all five copies (harmless
outside the workflow guard: a normalized Bash falls out of the other
guards' checks as before). Parity test now binds five copies and adds
Bash to the dormancy alarm. New workflow-guard test file exercises the
observable block (force-add on a worktree-agent branch): Shell bare and
module-qualified block with WORKTREE_AGENT_FORCE_ADD_FORBIDDEN, benign
Shell passes, Claude Bash unchanged — negative-controlled against the
pre-fold guard (the two Kimi cases fail there). Golden parity fixtures
regenerated; diff verified to change exactly the five guard entries per
surface.
* fix(#2304): map Kimi tool_output and route workflow-guard block to stderr
Third-party review (cross-AI verifier) caught two gaps in the revision:
1. Kimi PostToolUse events carry `tool_output`, not `tool_response`
(kimi-cli src/kimi_cli/hooks/events.py post_tool_use()), so the
read-injection scanner — which reads data.tool_response — was STILL
dormant on real Kimi payloads; the earlier tests passed because they
sent Claude-shaped payloads. The shared normalization block now maps
tool_output -> tool_response (inert in PreToolUse guards, where the
field is absent), and the scanner's Kimi tests send the real shape.
2. The workflow guard's force-add block wrote its reason to stdout only.
Kimi's exit-2 protocol feeds stderr back to the model — the exact
fix this PR already applied to the other blocking guard — so the
newly-awakened block would have been a silent denial. Reason now
also routed to stderr, asserted in the test.
Also: the scanner's "unknown name" test now uses a genuinely unmapped
name (FetchURL) — Shell stopped qualifying when it entered the map —
and the workflow guard's write branch (WriteFile advisory,
StrReplaceFile .planning pass) gains behavioral coverage. All five
copies stay byte-identical (parity test green); golden fixtures
regenerated, diff verified to the five guard entries per surface.
Negative-controlled: 3 new assertions fail against the pre-fix hooks.
* docs(#2304): update changeset to cover the full five-guard fix
Review round 2 (2026-07-18) flagged the changeset as stale: it was
written for the first commit and still described only the three guards
named in the issue. The shipped diff grew to five guards plus two
payload dimensions the original body never mentioned. The body now
names gsd-read-injection-scanner and gsd-workflow-guard, the ReadFile
and Shell vocabulary entries, the tool_output -> tool_response mapping,
and the workflow guard's stderr block-reason routing.
* test(#2304): regenerate kilo golden fixture after #2305 landed on next
The branch's fixture sweep predates
|
||
|
|
9181c77df5 |
fix(#2515): auto-merge the release/hotfix -> main PR when clean (#2516)
The finalize job created the release/hotfix -> main merge-back PR but never merged it, so every release and hotfix needed a manual merge click on main. The opposite direction (main -> next) is already admin-merged by auto-backmerge.yml; this direction was the asymmetric manual step where the recurring divergence got hand-reconciled. Adds a finalize step (after "Verify publish", so main only absorbs a confirmed-published release) that finds the open merge-back PR, polls until GitHub settles its mergeability, and admin-merges it with a merge commit ONLY when MERGEABLE. A CONFLICTING PR is left open for manual resolution rather than force-merged. Non-fatal (continue-on-error): the tag + npm publish already happened, so a merge-back that can't complete (org PR policy, token) must not fail the release. With the main-is-ancestor-of-next invariant restored (#2504), this merge is clean every release -- verified: a simulated 1.8.1 hotfix merges to main producing [1.8.1]->[1.8.0]->[1.7.0]->[1.6.1] and version 1.8.1 with no conflict, and main's hardened auto-backmerge.yml survives the merge. Guarded by three assertions in release-backmerge-invariants.test.cjs (step present + admin-merge, gates on MERGEABLE, continue-on-error); verified they fail when --admin or the MERGEABLE guard or the continue-on-error is removed. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
21c22d8cb0 |
Merge pull request #2514 from open-gsd/chore/backmerge-main-to-next-3a4df4d8
chore: back-merge main → next (
|
||
|
|
60184c6d22 |
chore: back-merge main into next (3a4df4d8)
|
||
|
|
3a4df4d8c4 |
ci(#2504): converge main's auto-backmerge.yml with next (continue-on-error hardening)
Delivers the #2506 blast-radius fix to main NOW instead of deferring to the next release. Deferring specifically fails for a hotfix: 1.8.x hotfix branches are cut from the immutable v1.8.0 tag (which lacks this hardening), so a hotfix would carry the un-hardened workflow to main and never converge it. Converging now makes main's backmerge robust regardless of whether the next release is a minor or a hotfix. The only functional change is `continue-on-error: true` on the two version-sync steps (verified byte-diff vs next). main and next are now identical on auto-backmerge.yml. |
||
|
|
bcdfd21c61 |
fix(#2504): make auto-backmerge survive a broken workflow copy + gate the invariants (#2506)
The main->next auto-backmerge fails after nearly every release, leaving
main not an ancestor of next, so the following release->main merge-back
conflicts. Root cause is a copy-shuffling loop: auto-backmerge.yml must be
identical on main and next, but `-s ours` (main->next) and the release-tree
merge-back (release->main) each overwrite one copy wholesale, so a fix
applied to one copy is repeatedly overwritten by the copy that lacks it.
The build:lib step proves it: added to main (
|
||
|
|
807edc3559 |
Merge pull request #2503 from open-gsd/chore/backmerge-main-to-next-fe573118
chore: back-merge main → next (
|
||
|
|
756369a23e |
chore: back-merge main into next (fe573118)
|
||
|
|
fe5731185e |
chore: merge release v1.8.0 to main (#2500)
Reconciles main to the v1.8.0 tree while preserving the [1.7.0] CHANGELOG section that release/1.8.0 omitted (root-caused in #2502: the main->next auto-backmerge failed after 1.7.0, so next never received 1.7.0's promoted CHANGELOG). Both parents kept so the v1.7.0 and v1.6.1 tags remain in main's ancestry. This also delivers the fixed auto-backmerge.yml (build:lib step, #2281) to main so the main->next backmerge stops failing. |
||
|
|
d39ca0c004 |
Merge pull request #2490 from open-gsd/docs/2481-adr-1239-effort-axis
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort |
||
|
|
3878dbdf02 |
Merge pull request #2501 from open-gsd/chore/sync-next-version-1.8.0
chore: sync next package version to 1.8.0 |
||
|
|
8cf747724f | chore: sync next package version to 1.8.0 | ||
|
|
e4df05126d | chore: promote CHANGELOG for v1.8.0 | ||
|
|
53e3028005 | chore: finalize v1.8.0 | ||
|
|
9fe9da9830 |
fix(#2488): strip leading terminators so changeset bullets survive re-parse (#2492)
* fix(#2488): strip leading terminators so changeset bullets survive re-parse A fragment body beginning with a line terminator rendered as an empty `- ` bullet followed by an orphaned paragraph. `parseChangelog` treats a non-indented line as terminating a bullet, so `github-release-notes.cjs` silently dropped the entry when re-parsing CHANGELOG.md to build the GitHub Release body. Two independent causes, both in scripts/changeset/parse.cjs: 1. `extractDocsExempt` stripped trailing terminators but not leading ones. `DOCS_EXEMPT_RE` is `^...$` under /m, so removing a first-line `<!-- docs-exempt -->` marker left the `\n` that `$` does not consume. 2. `parseFragment` preserved the post-frontmatter body verbatim, so a blank line between the closing `---` and the first content line produced the same leading `\n` with no marker involved. 8 of 256 pending fragments were affected, split 4/4 across the two causes — including the OpenCode MCP binding, the pi extension, and the EoS adapters, all of which would have vanished from the v1.8.0 release notes. Regression tests cover both causes in LF and CRLF form, plus an end-to-end serializeChangelog -> parseChangelog round-trip that pins the user-visible defect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2488): regenerate golden install fixtures for parse.cjs scripts/ ships in the npm package and the installer, so the golden install-parity fixtures record a content hash for every shipped file. Editing scripts/changeset/parse.cjs drifts that hash and fails all 18 per-runtime parity tests. Regenerated via `npm run gen:golden`. The diff is exactly one line per fixture — the scripts/changeset/parse.cjs hash — with no unrelated drift. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ecef1e1670 |
test(#2481): add tracking ref to allow-test-rule exemption
lint-allow-test-rule-refs (ADR-456) requires a #NNN issue ref on the SAME line as allow-test-rule:. CI-only lint (lint:ci), so it passed the local lint gate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5c487bdec5 | chore(#2481): backfill changeset PR number (pr:0 -> 2490) | ||
|
|
09b535ac00 |
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.
Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.
Every per-host value is documentation-sourced, never inferred:
- claude argv -- verified via 'claude --help' (--effort <level>)
- opencode argv -- verified via 'opencode run --help' (--variant)
- codex argv -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
flag (config.toml key only), so the global -c override is the
only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
fails closed rather than inheriting a profile baseline
No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by
|
||
|
|
c19d3d7bda |
chore(#2453): resolve the uniform-router-param conflict and clear the warning floor (#2489)
npx eslint . reported 0 errors, 53 warnings. The floor eroded the signal: a
genuinely new warning had to be spotted against noise, so 'lint is clean' was not
usable as a check. This takes it to zero.
Fifty of the 53 were one category in gsd-core/bin/gsd-tools.cjs — all unused
ARGUMENTS (error x30, cwd x10, raw x7, args x3), never unused variables. They are
the Command Routing Hub's uniform handler signature, function routeX({ args, cwd,
raw, error }), declared identically whether or not a handler uses all four
members. argsIgnorePattern: '^_' is structurally in conflict with that convention:
satisfying it would mean _-prefixing ~50 parameters and making the signature
non-uniform across the table. Option 1 of #2453: disable args checking for that
file only, keeping varsIgnorePattern intact so genuinely dead variables (the #2379
class) still surface — verified by injecting an unused variable, which still warns.
This is the config decision #732 explicitly deferred ('Severities stay warn (no
config change in this pass)').
The remaining three predate the issue's count of 51 and are real defects, not
suppressions:
- tests/workflow-compat.test.cjs — the step-9 lookahead used (?=\*\*10\.|\z).
\z is a Perl/Ruby end-of-input anchor with NO meaning in JavaScript; it matched
a literal 'z', so the lazy span silently stopped at the first z whenever **10.
was absent, truncating the captured step and letting the assertion pass against
a partial block. Corrected to $.
- tests/debugger-prevention.test.cjs — an unused RegExp built one line above the
one actually used. Removed.
- tests/installer-migrations.test.cjs — try/finally inside a test body, which
CONTRIBUTING prohibits ('verbose, masks test failures, not an approved
pattern'); the unused t was the symptom. Converted to t.after().
Closes #2453
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
360b3ebc5a |
chore(#2496): clear five newly-disclosed production advisories (#2497)
* chore(#2496): clear five newly-disclosed production advisories The `#3588: npm audit --omit=dev reports zero advisories` gate began failing mid-release. The v1.8.0 finalize dry run was green at 17:27:27Z; GHSA-frvp-7c67-39w9 and GHSA-xgm2-5f3f-mvvc published at 18:17:25Z and 18:18:13Z, with fast-uri and two further hono advisories in the same window. Nothing in the tree changed — the advisory database did. All five arrive transitively through the one declared dependency @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol/sdk. Two-part fix, both following existing repo precedent: 1. `npm audit fix --omit=dev` (no --force) re-resolves fast-uri and hono inside their already-declared ranges. package.json untouched — the same approach as .changeset/witty-badgers-hum.md (body-parser) and .changeset/archived/fix-3588-npm-audit-clean.md. Clears the only high. 2. overrides["@hono/node-server"] = ">=2.0.5" for the remaining chain, which cannot resolve in-range (^1.19.9 cannot reach 2.0.5) because @modelcontextprotocol/sdk@1.29.0 is already latest and still declares the vulnerable range. Extends the block that already pins qs and body-parser. Resolves to 2.0.11. Bumping @anthropic-ai/claude-agent-sdk to ^0.3.x was tested and REJECTED: 0.3.216 moves @modelcontextprotocol/sdk to peerDependencies, which npm auto-installs, so the chain survives and resolution pulls extra advisories — 5 vulnerabilities including a high, versus 4 moderate. Forced major sits under a dependency no tracked source imports (see src/mcp-server.cts:22 — the JSON-RPC loop is hand-rolled precisely to avoid the MCP SDK), so runtime risk is minimal. Revisit once upstream ships a release depending on patched @hono/node-server. npm audit --omit=dev: found 0 vulnerabilities. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2496): add changeset fragment for the advisory clearance --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
3435218089 |
fix(#2452): stop shallow-fetching the base ref in three-dot-diff CI gates (#2485)
* fix(#2452): stop shallow-fetching the base ref in three-dot-diff CI gates `mutation.yml` re-fetched the base branch with `--depth=1` after checking out with `fetch-depth: 0`. The shallow re-fetch truncates the base ref's ancestry, so `git diff --name-only origin/<base>...HEAD` in scripts/mutation-matrix.cjs can no longer compute a merge base and aborts with `fatal: origin/next...HEAD: no merge base` (exit 2). The `detect` job then fails and the `mutate` shards never run — so the 80% mutation-score threshold went UNVERIFIED rather than enforced. The failure is branch-position dependent, which is why it went unnoticed: a branch already level with the base incidentally passes (its merge base IS the single fetched commit), while a branch that is BEHIND fails. Observed on PRs #2436 and #2005. `changeset-required.yml` and `docs-required.yml` shallow-fetched the base ref too (`--depth=50`), shrinking the same window further. All three now fetch the BASE REF unshallowed. Their shallow *checkout* depth is left at 50: that is a separate, deliberate cost control with fail-closed semantics, owned by tests/policy-lint-shallow-checkout.test.cjs. Only the base-ref fetch changes. The two `${{ }}` interpolations in mutation.yml's run: blocks now pass the base name through `env:`, matching the sibling workflows. Regression coverage in tests/mutation-workflow-base-ref.test.cjs: - a per-workflow contract guard asserting the base fetch carries no --depth (RED on origin/next for all three files, GREEN here). The YAML step parser handles block scalars and skips commented-out steps, so a future refactor to a multi-line `run:` cannot silently degrade the guard. - a real-git mechanism proof with boundary coverage at the shallow edge: with the base advanced 60 commits past the branch point, --depth=1 and --depth=60 both fail with `no merge base`, --depth=61 (merge base exactly at the boundary) succeeds, and an unbounded fetch succeeds. Each variant uses an independent clone, because a plain fetch does not un-shallow a repo that already carries a .git/shallow boundary. Also fixes a startup race in tests/run-with-timeout.test.cjs surfaced by this branch's gsd-test run (C1, linux-node24). The heartbeat file only appeared ~100ms after the grandchild's runtime was up, but the window is 1s spanning two cold node starts, so on a loaded runner the timeout fired before any heartbeat existed and the precondition failed for reasons unrelated to reaping. The child now writes its heartbeat once synchronously at startup; the frozen-vs-ticking comparison that actually proves reaping is unchanged. Closes #2452 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(changeset): backfill PR number to 2485 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c5e0371775 |
feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging Red phase for issue #1951 (reversibility tagging: classify decisions by undo cost, gate one-way doors behind a checkpoint:decision). Tests assert, per the issue's acceptance criteria: - discuss-phase CONTEXT.md template records a **Reversibility:** field with a rationale on captured decisions, and states it is optional - gsd-planner @-references planner-reversibility.md and stays under the 49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs) - a one-way rating inserts a checkpoint:decision before the dependent task; reversible inserts none; costly is flagged but never blocks - the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard) and inserting a checkpoint implies autonomous: false - docs/reference/plan-md.md documents <reversibility> as optional with all three ratings - --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected into the planner prompt, and is advertised in the command argument-hint and help full mode (argument-hint parity) - the override suppresses the gate but still persists the rating - cmdVerifyPlanStructure accepts every rating and the absent case (additive-validator guarantee, behavioral via runGsdTools) - parity: thinking-models-planning.md #4 adopts the canonical three-level taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone - no content loss from the planner extraction made to fit under the cap Prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately — regression guards proving the validator already accepts unknown optional tags. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1951): reversibility tagging — gate one-way-door decisions Classify planning decisions by what undoing them would cost, and give a one-way door a human beat before the agent walks through it (issue #1951, The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way door framing). Acceptance criteria met: - discuss-phase records an optional reversibility rating with a rationale on <decisions> entries in the phase CONTEXT.md template. Unrated decisions are treated as reversible, so existing phases are unaffected. - a one-way rating makes gsd-planner insert a checkpoint:decision before the task that implements the decision, reusing the existing checkpoint mechanism -- no new checkpoint machinery. - reversible ratings trigger no checkpoint; costly ratings are flagged in the plan but never block. - the rating persists on the task as the optional <reversibility rating=> element. cmdVerifyPlanStructure accepts every rating and the absent case; the structural validator does not reject unknown optional tags. - --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings -- the override changes what stops the run, not what the plan remembers. Single taxonomy, not two: references/thinking-models-planning.md #4 already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the canonical three-level vocabulary and now points at planner-reversibility.md as the taxonomy owner, with a parity test that fails if the surfaces diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE). agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the checkpoint DO/DON'T guidance was relocated verbatim into planner-antipatterns.md -- already @-referenced from the same section for the same topic, so the planner still loads it and nothing was dropped. A test guards the relocation against content loss. Files: gsd-core/references/planner-reversibility.md (NEW, canonical taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase workflow/command/help (flag wiring + parity), plan-md.md schema, discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest, size baselines, install goldens, plugin skills regen, changeset. Closes #1951 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1951): address orthogonal review findings Two isolated reviewers (correctness + security), neither of which authored the change. Every finding fixed: Security — the rationale is untrusted input (ADR-1577). It originates in conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each hop an LLM reading the previous hop's output, with no validation on the path. planner-reversibility.md and the discuss-phase template now state it is data and never instructions, and name the </reversibility> early-termination hazard explicitly -- a rationale that closes its own element injects sibling structure the executor reads as real tasks. Four tests guard it. Correctness 1 — nothing machine-enforced the feature's own promise: a task rated one-way with no preceding checkpoint:decision validated as fully clean, so a planner error silently reopened the gap this feature exists to close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A warning, not an error: <reversibility> stays additive and the plan stays valid. Four tests cover ungated (warns), gated (silent), still-valid, and reversible/costly never flagged. Correctness 2 — pass-always test. The --no-reversibility-gates parse test substring-matched the whole workflow file, and plan-phase.md prose mentions both tokens in one sentence, so it passed with the bash conditional deleted: it was testing the documentation, not the parser. Now scoped to the fenced bash blocks and matched as one physical line, with a negative control confirming prose alone cannot satisfy it. Correctness 3 — costly had no itemized emission rule, only one-way did, so two agents could diverge on whether to tag costly at all. Correctness 4 — template convention break: the example ratings were bare while every sibling field uses [...] to signal substitution, inviting an LLM to copy one-way/costly forward as boilerplate. Now bracketed. Correctness 5 — latent false-green: .includes('reversible') also matches inside irreversible/irreversibility, which appear in anti-pattern prose, so a surface that dropped the real taxonomy entry would still pass. Now word-boundary matched. ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md 1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on next). The ceiling may only rise for privileged host machinery, and reversibility gating is optional-feature logic, so the wiring was slimmed to its minimum and the explanatory prose moved to the reference files the planner already loads. plan-phase.md is now 94400 bytes -- 119 under the ceiling and 70 bytes SMALLER than on next, so the host loop shrank while gaining the feature, which is what phase 6 ratchets toward. The tracer contract (tests/tracer-bullet.test.cjs) is unchanged. Lint — fixed an unnecessary non-null assertion in verify.cts and a CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST- PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture must carry the common task elements The gated-one-way fixture built a checkpoint:decision task from the abbreviated skeleton in gsd-planner.md, which shows only the checkpoint-specific elements (<decision>/<context>/<resume-signal>). cmdVerifyPlanStructure requires <name> and <action> on EVERY task regardless of type, so the fixture failed validation for reasons that had nothing to do with reversibility: errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"] Caught by gsd-test on 14d14a39 (2 failures, both this fixture). The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task carries <name>/<files>/<action>/<verify> like any other. Fixture corrected to match. Verified behaviorally against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Not a product defect: the validator's every-task contract is intentional and pre-existing, and docs/reference/plan-md.md scopes its required-element list to type=auto/tracer only because those are the elements a planner must author, not because checkpoints are exempt from <name>. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1951): backfill changeset pr number to 2471 * fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision Both CI failures were real defects in code this PR added, not false positives. CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 — the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`, which escapes the hyphen but not backslash, so the escape was incomplete. It was also unnecessary: `-` carries no special meaning outside a character class. Replaced with a complete metacharacter escape (backslash included). Word-boundary behavior verified unchanged across all three ratings — notably that "irreversible" prose still does not satisfy a "reversible" match, which is the false-green this helper exists to prevent. Prompt injection scan — the checkpoint fixture used the human-verification child element inside <verify>. That tag name is a fake-instruction-boundary pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed files, so copying the shape from tests/verify.test.cjs (unflagged only because it is not in this diff) tripped the gate. Switched to the documented plain-prose <verify> form. The first attempt at that fix failed the same gate a second time: the comment explaining the collision quoted the offending tag literally. The comment now names it in prose instead — the scanner does not care whether a match is code or commentary, which is the whole point of the DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md. Verified locally before push: scan reports 0 findings across 57 changed files, eslint clean, and both fixtures still validate as designed (gated one-way silent, ungated one-way warns, neither errors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): record measured cost and halve gsd-tools spawns The Windows shard 1/3 job timeout was traced to the sharding layer, not to this PR's assertions — see #2472. Two contributing factors were this file's own, and are fixed here. 1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs, so scripts/run-tests.cjs weighted it at the table's median fallback (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x under-weight. Recorded the measured value from the green gsd-test run (max across the node22/node24 lanes, per gen-test-timings.cjs's convention). Only this one entry: a full regen churns 634 entries of run-to-run drift, and the table is explicitly advisory and un-gated, so a 637-line diff does not belong in a feature PR. 2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost. Spawns cut from 9 to 6 with no coverage lost: - the ungated-one-way warning and its stays-valid assertion now share one plan instead of building the same plan twice; - the reversible/costly never-flagged-as-ungated test was strictly subsumed by the additive suite, which already runs those two ratings ungated and asserts no /reversibilit/ warning at all — and the gate warning's text contains both "reversibility" and "one-way", so the broader assertion catches it. It only re-spawned gsd-tools twice to prove the same thing. Both are symptom fixes. The shard imbalance itself (19/11/10 minutes against a 20-minute cap, from a cost-blind round-robin partition that also reshuffles downstream files whenever one is inserted) is tracked in #2472. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture adopts the #2444 type-branched contract Surfaced by rebasing onto next, which gained #2444 (branch plan-structure validation on task type=checkpoint:*) while this PR was in review. cmdVerifyPlanStructure no longer applies one required-element set to every task. A checkpoint:decision now requires <name> + <resume-signal> + <decision> + <options>, and is exempt from the <action>/<verify>/<done>/ <files> set that auto and tracer tasks carry. The gated-one-way fixture predated that split and failed on the new requirement: errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"] Fixture rewritten to mirror the checkpoint:decision contract exactly — real <options> with two <option> children — rather than padding it with fields checkpoints no longer need. That also drops the plain-prose <verify> the earlier revision carried purely to dodge the prompt-injection scan; a checkpoint task has no <verify> requirement at all, so the workaround is moot. Verified against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
04a0eb8d63 |
fix(#2450): CRLF-tolerant session-section rewrite + no-op-detection guard (#2482)
* fix(#2444): re-resolve body-parser to 2.3.0 in lockfile (GHSA-v422-hmwv-36x6) GHSA-v422-hmwv-36x6 (body-parser DoS via invalid limit value, low severity, published 2026-07-20T23:23:26Z) made tests/npm-integrity-gate.test.cjs (#3588: root workspace production tree has no advisories) fail any subsequent npm audit --omit=dev. The advisory affects body-parser >=2.0.0 <2.3.0 pulled transitively via @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol/sdk -> express -> body-parser@2.2.2. express@5.2.1 already declares body-parser as ^2.2.1, so 2.3.0 is a valid re-resolution within express's own compatibility range — no override needed. Regenerated the lockfile via 'npm audit fix --omit=dev' which re-resolves transitive deps within their declared ranges; package.json is unchanged. Verified: npm audit --omit=dev reports 0/0/0/0/0 advisories; body-parser now reads as 2.3.0 in 'npm ls body-parser --omit=dev'. * test(#2450): failing-first CRLF regression for record-session insert path cmdStateRecordSession's section-rewrite regexes (src/state.cts:1166, :1195) used literal \\n which cannot match CRLF STATE.md delimiters. The detector regex (CRLF-tolerant via $ under /m) entered the rewrite branch, the writer regex silently no-op'd, but updated.push(...)/ sessionCreated=true ran unconditionally. Result: caller reported recorded:true with 'Resume File' in updated, but the field was never written to disk. With core.autocrlf=input, the CRLF working-tree file produces no git diff, so the bug was invisible. Adds three regression tests covering all three rewrite paths: - CRLF STATE.md with ## Session and Resume file absent - CRLF STATE.md with ## Session and Stopped at absent - CRLF STATE.md with ## Session Continuity (bootstrap shape) Each asserts the field IS on disk (the bug discriminator: the pre-fix command's JSON output looked identical to a successful write). * fix(#2450): CRLF-tolerant session-section rewrite + no-op-detection guard Two regexes in cmdStateRecordSession used literal \\n which cannot match CRLF STATE.md, silently no-op'ing the section rewrite while the CRLF- tolerant detector above entered the branch. The reporter's exact repro: on a CRLF STATE.md with one canonical session field absent, the command returned recorded:true + updated:['Resume File'] but the field was never written to disk. Three changes: 1. src/state.cts:1166 (canonical ## Session rewrite regex): \\n -> \\r?\\n 2. src/state.cts:1195 (## Session Continuity insert regex): \\n -> \\r?\\n 3. Defensive invariant (#2450 class fix per reporter's suggestion): track whether the chosen branch's replace actually matched via callback flag. Only set sessionCreated=true and push to updated when rewriteMatched. Unreachable post-fix, but fail-loud is the right posture for a silent- success gate. If a future drift between the detector and writer regexes reintroduces the asymmetry, the caller will not see false updated entries. Same canonical CRLF-tolerant form already in use at check-command-router.cts :205 (extractPlanDesignatedSections). Same bug class previously fixed in #1658, #1668, #2206, #2449. * fix(#2450): address review followups + add changeset Code-review + security-review both flagged the unreachable else at the Session Continuity branch (defaulted rewriteMatched=true in dead code, re-arming the bug class for future drift). Removed the else; the remaining code path leaves rewriteMatched=false if linesToInsert is empty, preserving the fail-loud posture. Added scope-limitation doc to the rewriteMatched gate: it covers the INSERT path only, not the earlier in-place stateReplaceField successes (which DID land on disk and correctly push to updated unconditionally). Tests: - Normalized STATE_CRLF_SESSION_MISSING_RESUME fixture to match the canonical 6-key frontmatter of STATE_WITH_SESSION (code-review I2). - Added mixed-ending test (LF frontmatter + CRLF body) to close CONTRIBUTING.md:490 'Mixed CRLF/LF newlines' requirement (I1). Added Fixed changeset (code-review H1). * docs(changeset): backfill PR number to 2482 |
||
|
|
bf0d715733 |
fix(#2472): cost-balanced test sharding and pinned CI base commit (#2480)
* fix(#2472): weight-aware shard partition Windows shard 1/3 hit the 20-minute job cap with no failing assertion. Root cause is the shard layer, not the chunk layer: selectShard partitioned by sorted ARRAY INDEX (k % n, #1212), which balances file COUNTS and ignores file COST. On the real unit suite that produced 12.4m / 19.2m / 15.2m — a 1.23x max/ideal ratio leaving the heaviest shard 5% under the cap. Because assignment keyed off position, inserting one test file re-indexed every file after it and could tip that shard over; deterministic, so a re-run reproduced it exactly. This is NOT the chunk packer (#2456/#2463). That fix works and applies one level down, WITHIN a shard. The across-shard partition predated it and never consumed the cost table. Both layers now share one cost model. selectShard takes an optional weightOf and, when given one, partitions by LPT (longest-processing-time-first) — the same algorithm packChunks uses. Omitting it keeps the legacy round-robin byte-identical, so every existing test above still exercises that path unchanged and callers without timing data lose nothing. A missing timings table yields uniform weight 1, under which LPT degenerates to the equal-count split. Projected on the real suite: 16.4/17.3/13.0 -> 15.6/15.6/15.6 (worst shard 17.3m -> 15.6m). Tests: a skewed-cost regression (round-robin clusters all four heavy files onto one shard at 2.98x ideal; LPT does not), back-compat equivalence, determinism, tie-breaking, order preservation, and two fast-check properties — the partition is exhaustive and disjoint (getting this wrong silently DROPS tests from CI, the worst failure mode for a harness), and no shard exceeds average + heaviest file. Two assertions were corrected during authoring rather than shipped wrong: - an initial "LPT within 4/3 of ideal" bound was false. The 4/3 figure is relative to the OPTIMAL makespan, not the average, and the two differ when item sizes force a pairing. Replaced with Graham's average+max bound, which is what is actually provable. - "weighted is never worse than round-robin" is also false; fast-check falsified it with [19316,10190,1,9128,29353,20227] over 2 shards (rr 48670, lpt 48671). Round-robin can win by luck on a specific input. Dropped, with the counterexample recorded in place so it is not re-asserted later. Closes #2472 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2472): rotate tied bins; restore #1212 test block; lazy cost table Isolated-review findings, all fixed. HIGH — zero weights collapsed the whole partition onto shard 1. The lightest-bin scan compared weight only, and adding a zero-weight file leaves its bin's weight unchanged, so bin 0 stayed tied-minimum forever and every such file landed on it. Verified: all-zero weights gave shard1=[a..f], shard2=[], shard3=[] — two of three CI runners idle while one ran everything. Reachable through safeWeight's own clamp (a NaN/negative/Infinity entry in a corrupted or hand-edited timings table) and through any genuine 0ms measurement, so the clamp reproduced the exact failure its comment claimed to prevent. Ties now break on file COUNT after weight, which rotates. Pinned by two regression tests (all-zero, and clamped NaN/negative/Infinity) plus a property over list size x shard count. The live table has no 0ms entries (min 19ms), so production was not affected — but nothing prevented it. MEDIUM — the new describe block had swallowed #1212's pre-existing property test, which is why a test under a "weight-aware" heading never passed a weigher. That was a bad block boundary in the previous commit, not a bad test: the #2472 describe was opened before #1212's last test instead of after. Moved back where it belongs; #1212 is 762-879 and #2472 is 894-1082. LOW — that relocated property test ran unseeded. Seeded (12120) per the repo's property-test convention so a failure reproduces. Verified passing under the new seed. LOW — hoisting the timings load above the shard block charged a readFileSync + JSON.parse to invocations that exit before needing it (empty selection, --files matching nothing). Now lazily memoized, so neither consumer reads the table unless it is used and it is still read at most once. Real-suite projection unchanged at 15.6m / 15.6m / 15.6m. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#2472): correct stale round-robin sharding descriptions The partition is now cost-balanced, so the header block in run-tests.cjs and the two comments in test.yml describing '--shard' as a round-robin over sorted file index were actively wrong. Updated to describe LPT over measured duration, and to state the degenerate case explicitly: with no timing data every file weighs the same and the partition collapses back to k % n, which is why the pre-existing #1212 CLI tests still pass unchanged (their nine synthetic files are absent from the timings table, so all take the identical median weight). Remaining 'round-robin' mentions are correct — they describe the unweighted fallback path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2472): shard diagnostics, cost-routing E2E test, table validation Second orthogonal review (operational lens) findings, all fixed. HIGH — cross-runner partition divergence. Each of the up-to-12 CI jobs runs its own 'merge base into head' and computes its own partition, so if the inputs differ between jobs (the file list, or the timings table) two jobs can place the same file in different shards or in none. Every job stays internally exhaustive and disjoint, so nothing errors: a test simply never runs and CI stays green. The risk class is pre-existing — round-robin diverges identically when the file set differs between jobs, which is literally this issue's insertion instability — but weighting adds tests/test-timings.json as a second input that must match, so it widens the hole. Properly closing it means pinning the partition inputs per run, a workflow change beyond this fix. What IS closed here is the silence. Each shard now prints an input fingerprint over the FULL pre-partition list and the weight assigned to each file — deliberately not this shard's slice, which would differ by design and be useless for comparison. All shard jobs of one run must print an identical sig; a mismatch is direct proof the runners disagreed about the input. Verified: three independent computations agree, and the sig changes when the input drifts by one file. MEDIUM — nothing proved main() actually threads fileWeightOf() into selectShard. Every pre-existing --shard E2E test uses synthetic filenames absent from the real table, so all collapse to a uniform median weight, under which LPT is mathematically identical to k % n — a typo on that one wiring line would have passed the whole suite. Added an E2E test that injects a table via RUN_TESTS_TIMINGS_FILE with differing costs, placing the heavy files at exactly the indices round-robin hands to shard 1, and asserts shard 1 does NOT receive all three. Plus a test that all three shards emit the same sig. MEDIUM/LOW — no observability. The diagnostic line now reports files, weighed count, aggregate weight, and whether the table loaded, so a table that silently failed to parse shows table=absent/weighed=0 instead of being indistinguishable from a healthy load. (The reviewer confirmed the advisory fallback is already live on next: feat-2296-provider-escalation.test.cjs is missing from the table.) LOW — typeof [] === 'object', so a hand-edit turning the map into a list was accepted as a valid table. Now rejected via Array.isArray, falling back to uniform weight like any other malformed table. LOW — stale round-robin wording in ci-test-scope.test.cjs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2472): pin every CI job to one base commit Closes the cross-runner divergence at its source instead of only making it visible. Each job of a run executes the rebase-check step independently, minutes apart across a 12-job matrix, and merged the MOVING origin/<branch> ref. If the base advanced mid-run, different jobs merged different trees. That was survivable when jobs only had to agree on pass/fail; it is not once they must agree on a PARTITION. Each shard job computes the whole split and keeps its own slice, so jobs working from different trees can place a file in two shards or in none — and every job still looks internally consistent, so nothing errors. A test silently never runs and CI stays green. ci-rebase-check.cjs now accepts CI_REBASE_BASE_SHA and pins BOTH the fetch and the merge to that one commit, so the two can never disagree. test.yml passes github.event.pull_request.base.sha on all three rebase-check steps; that value is fixed for the life of a run, so all jobs merge the identical base. This also closes the PRE-EXISTING half of the divergence. Round-robin had the same exposure whenever the test-file set differed between jobs — that is this issue's insertion instability — so the pin fixes the older hole too, not just the timings-table input weighting added. Only a full 40-hex sha is accepted; empty (push/workflow_dispatch), malformed, or injected values fall back to the branch ref rather than handing an arbitrary string to git fetch as a refspec. resolveBaseRefs is extracted pure and exported, and runMain is guarded behind require.main === module, so the pin contract is testable without spawning git. Tests (tests/ci-test-scope.test.cjs): every rebase-check step must carry the pin; a valid sha pins both refs; absence falls back correctly; and five hostile values — short sha, uppercase, --upload-pack= injection, ref expression, empty — are each rejected. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
909a3b180b |
fix(#2470): install pi's extension as gsd.js so pi actually discovers it (#2478)
* test(#2470): failing-first — pi extension must satisfy pi's auto-discovery filter pi auto-discovers extensions/ entries through isExtensionFile(), which accepts only .ts and .js. GSD installs its extension as gsd.cjs, so pi silently skips it: no /gsd command, no error, no log line. Encodes pi's discovery PREDICATE rather than a literal filename, so the contract keeps holding across future renames, and adds the migration-006 test matrix for retiring the stale gsd.cjs left in pre-fix installs. Red until the fix lands. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2470): install pi's extension as gsd.js so pi actually discovers it pi auto-discovers extensions/ entries via isExtensionFile(), which accepts only .ts and .js and skips everything else silently. capabilities/pi declared the dest as gsd.cjs, so the extension installed correctly and was then ignored forever: no /gsd command, no error, no log line. Install it as gsd.js. The in-repo source stays pi/gsd.cjs — tests require() it directly and .cjs is unambiguous CommonJS; only the installed name has to satisfy pi, and pi loads accepted files through jiti, which handles CJS and ESM alike. (The reporter's premise that ~/.pi/agent/package.json declares "type":"commonjs" does not hold — pi never writes that file.) Renaming an installed artifact requires a migration record, so add 006 to retire the stale gsd.cjs from pre-fix installs; without it the old path drops out of the manifest and uninstall can never remove it. The migration plans nothing for an unmanifested gsd.cjs: emitting remove-managed there would have the executor downgrade it to preserve-user and mark it blocked, failing the install for anyone who hand-placed their own file. Also pins body-parser >=2.3.0 (GHSA-v422-hmwv-36x6). The advisory reaches the production tree transitively via the Claude Agent SDK and fails the npm-integrity gate, blocking any PR; pinned via the existing overrides idiom. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2470): address orthogonal review findings + register migration checksum Code review: - pi/gsd.cjs's install docstring still told readers to copy the file to extensions/gsd.cjs — the exact silently-broken state this PR fixes. Anyone following it recreated the bug. - Two stale extensions/gsd.cjs comments in install-minimal-hooks.test.cjs. Security review: - _installNativePluginIfDeclared confined nativePlugin.dir but joined nativePlugin.file onto the validated directory unchecked, so a descriptor whose file carried .., an absolute path, or a NUL byte would have written outside configHome. Not reachable in a shipped build (descriptors are first-party and compiled into the capability registry), but file is exactly the field this PR changes. Confine the full dest path instead; for a well-formed descriptor this resolves identically to the previous mkdir(dir) + join(dir, file). Covered by four new write-confinement tests. Also register migration 006 in the #670 EXPECTED_CHECKSUMS baseline — shipped migration bodies are locked to a committed checksum and a new migration fails CI until it is listed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2470): never dereference a symlinked managed path when snapshotting fs.copyFileSync follows symlinks, so a managed path replaced by a link had the REFERENT's bytes copied into the migration journal's rollback and backup trees — a gsd.cjs symlinked at a private key would land that key's contents under gsd-migration-journal/. Deletion was already safe (fs.rmSync unlinks the link, never the target); the copy was not. Nothing GSD installs is ever a symlink, so the faithful snapshot of a symlinked managed path is the link itself. copyPreservingSymlink recreates it, which keeps rollback fidelity (restore re-creates the same link) while never reading the referent. Scoped the pre-delete to the symlink branch only, so the regular-file path keeps copyFileSync's overwrite-in-place and a mid-restore failure cannot destroy the destination. The restore-side existence check moves to lstat, since existsSync follows a link whose target is gone and would silently skip the restore. This lives in the engine all six migrations share, so 000-005 are hardened too. Also regenerates the pi golden-parity hash: correcting pi/gsd.cjs's own install docstring changes the extension's content, which the golden suite caught. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2470): symlink-preserve the in-apply failure-recovery restore too The previous commit routed three copy sites through copyPreservingSymlink but missed a fourth: the catch block inside applyInstallerMigrationPlan, which replays rollback snapshots taken earlier in the SAME apply attempt. Those snapshots are symlinks precisely because of that commit, so the raw copyFileSync there dereferenced them and wrote the referent's bytes to the LIVE install path — worse than the journal-tree leak it was meant to fix, since it is user-visible and at a predictable location. Verified by experiment rather than assertion: with the pre-fix line restored, the managed path comes back as a REGULAR FILE containing the referent's bytes; with the fix it comes back as a symlink and the bytes appear nowhere. The accompanying test injects the failure by letting the delete succeed and then throwing once, modelling a later step failing after the delete. That ordering is load-bearing — an earlier draft injected before the delete, which leaves the live path in place, so the pre-fix copyFileSync hit a same-file collision and threw instead of leaking. That draft passed against the bug it was written to catch; this one fails against it. Adds the missing rollback() coverage as well: a restored symlinked managed path must come back as a link pointing at its original target. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#2470): read the backup location from the journal, not the plan The new backup-content assertion read backupRelPath off result.plan.actions, where it is always null: the planner reserves the field and apply chooses the concrete location, recording it in the journal. The assertion therefore failed on "backup path must be recorded for the user" rather than on anything about the behavior it was written to check. Read it from the journal, which is the authoritative record. Verified by executing all four new test bodies in-process against the built engine — the backup file exists and holds the locally patched content. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2470): backfill changeset pr number to 2478 * chore(#2470): backfill changeset pr number to 2478 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b6ebf9d4b4 |
fix(#2449): make tdd.review-checkpoint frontmatter regex CRLF-tolerant (#2477)
* test(#2449): failing-first CRLF regression for tdd.review-checkpoint cmdTddReviewCheckpoint's frontmatter regex /^---\\n...\\n---/ (line 751) cannot match CRLF PLAN.md, so a type:tdd plan with Windows line endings was silently classified as 'no type:tdd plans' — output indistinguishable from a phase that genuinely contains no TDD plans. The advisory gate then short-circuited to a confident pass with no violations table. Adds tddPlanCrlf() fixture helper (CRLF twin of tddPlan) and a [crlf] test case that asserts tddPlans===1 for a CRLF type:tdd plan (would be 0 before the fix), plus block:true/violations:1 since the fixture has no RED/GREEN commits. * fix(#2449): make tdd-review-checkpoint frontmatter regex CRLF-tolerant The frontmatter delimiter regex at src/check-command-router.cts:751 used literal \\n which cannot match a CRLF PLAN.md delimiter (---\\r\\n). frontmatterMatch was null, the plan was never classified type:tdd, tddPlanFiles stayed empty, and the advisory gate short-circuited to a confident pass with no violations table. The fix replaces /^---\\n([\\s\\S]*?)\\n---/ with /^---\\r?\\n([\\s\\S]*?)\\r?\\n---/ — the same CRLF-tolerant form already used elsewhere in the same file at line 205 (extractPlanDesignatedSections). Same canonical pattern; one-line change. Same bug class previously fixed in #1658, #1668, #2206; this instance is in a different file and is not a duplicate of any of them. * fix(#2444): re-resolve body-parser to 2.3.0 in lockfile (GHSA-v422-hmwv-36x6) GHSA-v422-hmwv-36x6 (body-parser DoS via invalid limit value, low severity, published 2026-07-20T23:23:26Z) made tests/npm-integrity-gate.test.cjs (#3588: root workspace production tree has no advisories) fail any subsequent npm audit --omit=dev. The advisory affects body-parser >=2.0.0 <2.3.0 pulled transitively via @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol/sdk -> express -> body-parser@2.2.2. express@5.2.1 already declares body-parser as ^2.2.1, so 2.3.0 is a valid re-resolution within express's own compatibility range — no override needed. Regenerated the lockfile via 'npm audit fix --omit=dev' which re-resolves transitive deps within their declared ranges; package.json is unchanged. Verified: npm audit --omit=dev reports 0/0/0/0/0 advisories; body-parser now reads as 2.3.0 in 'npm ls body-parser --omit=dev'. * test(#2449): add mixed-endings (CRLF frontmatter + LF body) variant Per code-review Low finding: CONTRIBUTING.md:490 names 'Mixed CRLF/LF newlines' as a required adversarial fixture class. The pure-CRLF [crlf] test covers the reported bug (Windows editor + autocrlf=input). This [crlf-mixed] variant covers the more adversarial case where frontmatter delimiters are CRLF but the body is LF (editor that normalizes body text, or a toolchain concatenating CRLF + LF fragments). The classifier only reads frontmatter, so detection is unaffected — but the test locks the behavior. * docs(changeset): add Fixed fragment for #2449 PR * docs(changeset): backfill PR number to 2477 |
||
|
|
46ba9ed464 |
fix(#2444): branch plan-structure validation on task type=checkpoint:* (#2473)
* test(#2444): failing-first regression for checkpoint:* plan-structure validation Add acceptance-criteria tests covering the three canonical checkpoint task types (human-verify, decision, human-action) plus an unknown-subtype forward-compat case. Each canonical type must pass verify plan-structure when it carries its type-specific required fields (per gsd-core/references/checkpoints.md), and must be flagged when those fields are missing. Non-checkpoint tasks keep the existing <action>/<verify>/<done>/<files> requirements unchanged (AC3 regression guards). The existing 'errors when checkpoint task but autonomous is true' fixture is updated to use the canonical checkpoint:human-verify triple (<what-built>/<how-to-verify>/<resume-signal>) so it does not collide with the new per-type validator; the assertion (autonomous is not false) is unchanged. * fix(#2444): branch plan-structure validation on task type=checkpoint:* cmdVerifyPlanStructure unconditionally required <action>/<verify>/<done>/ <files> on every task, so every checkpoint:* task — which uses the checkpoint convention's type-specific fields instead — was reported as a structural error. Checkpoint-heavy phases produced walls of false findings. The fix introduces two pure helpers in verify.cts: - extractPlanTaskInfos(content): single ReDoS-safe pass over <task ...>...</task> blocks that captures BOTH the opening-tag attribute string (so the type= selector is not lost, as it is with extractTaggedBlocks) and the body, returning a typed PlanTaskInfo. - validatePlanTaskStructure(task): branches on the task's type. checkpoint:human-verify requires <what-built>/<how-to-verify>/ <resume-signal> (the canonical triple). checkpoint:decision requires <decision>/<options>/<resume-signal>. checkpoint:human-action requires <action>/<instructions>/ <verification>/<resume-signal>. Unknown checkpoint:* subtypes require only the universal <resume-signal> (forward-compat). All other types keep the historical <action>/<verify>/<done>/<files> requirements unchanged. Canonical reference: gsd-core/references/checkpoints.md. Per-type field sets validated against the documented templates in agents/gsd-planner.md and gsd-core/templates/phase-prompt.md. * fix(#2444): re-resolve body-parser to 2.3.0 in lockfile (GHSA-v422-hmwv-36x6) GHSA-v422-hmwv-36x6 (body-parser DoS via invalid limit value, low severity, published 2026-07-20T23:23:26Z) made tests/npm-integrity-gate.test.cjs (#3588: root workspace production tree has no advisories) fail any subsequent npm audit --omit=dev. The advisory affects body-parser >=2.0.0 <2.3.0 pulled transitively via @anthropic-ai/claude-agent-sdk -> @modelcontextprotocol/sdk -> express -> body-parser@2.2.2. express@5.2.1 already declares body-parser as ^2.2.1, so 2.3.0 is a valid re-resolution within express's own compatibility range — no override needed. Regenerated the lockfile via 'npm audit fix --omit=dev' which re-resolves transitive deps within their declared ranges; package.json is unchanged. Verified: npm audit --omit=dev reports 0/0/0/0/0 advisories; body-parser now reads as 2.3.0 in 'npm ls body-parser --omit=dev'. * test(#2444): close review gap-closure tests + harden type-attr charset Orthogonal review (code-review + security-review subagents) returned APPROVE on Standards and Spec. Per the playbook's zero-tolerance policy, address every Low finding: Spec gap-closures: - AC3 verbatim: add explicit <done> and <files> regression tests for non-checkpoint tasks (pre-existing tests only covered <action> and <verify>). - AC2: add checkpoint:decision missing <decision>, checkpoint:human-action missing <action>, checkpoint:human-action missing <verification> cases (the implementation enforces all of these; only one missing-field case per type was previously tested). - Remove the duplicate 'returns error for nonexistent file' test that leaked into the new describe block from the insertion edit. Security hardening (Low-sev, defense-in-depth): - Tighten the task type= attribute extractor in src/verify.cts from [^"'>\s]+ to [\w:-]+ so a hostile type= attribute cannot carry markup fragments (e.g. type=evil<fragment) into the verifier's typed JSON output. All legitimate type values (auto, tracer, manual, checkpoint:human-verify, checkpoint:decision, checkpoint:human-action, checkpoint:tdd-review) match the tighter charset. - Add adversarial regression test asserting type=evil<fragment surfaces as 'evil' (capture stops at '<'), with no markup chars (< > ( ) &) in the surfaced type field. * docs(changeset): add Fixed fragments for #2444 PR Two fragments: - sturdy-jays-tumble.md: the verify plan-structure checkpoint fix - witty-badgers-hum.md: the body-parser 2.3.0 re-resolution PR number backfilled to 0 placeholder per CLAUDE.md 'PR Number Handling'; will backfill to the real PR number immediately after gh pr create returns. * docs(changeset): backfill PR number to 2473 Per CLAUDE.md 'PR Number Handling': backfill the placeholder pr:0 with the real PR number returned by gh pr create. |
||
|
|
a54feb4216 |
fix(#2440): per-counter progress ratchet — total_plans always takes derived value (#2468)
* fix(#2440): per-counter progress ratchet — total_plans always takes derived value Two sites fixed (targeted — existing body-only write tests preserved): Site A — read path: shouldPreserveExistingProgress (state-document.cts:167) removed total_plans from the all-or-nothing ratchet check. It now joins total_phases as an always-derived counter. Only completed_phases and completed_plans keep ratchet behaviour (they are monotonic). This fixes gsd-tools query state.json reporting stale total_plans when a curated completed_plans triggers the ratchet. Site B — write path: applyStatePreservation (state-transition.cts:162) gained a deriveProgressKeys opt-in flag. When true (passed by cmdStatePlannedPhase only), total_plans and total_phases take the derived (post-sync) value instead of the wholesale curated restore. When false (the default — state.update, state.patch), the existing #3242 wholesale protection stays fully in force. This fixes the state planned-phase verb writing a stale total_plans. The opt-in approach preserves all 8 existing #3242/#1264/#500 body-only write tests that assert wholesale progress preservation during non- progress updates. Tests: - tests/state.test.cjs: 4 unit tests for shouldPreserveExistingProgress (total_plans upward/downward/equality + completed_plans ratchet active). - tests/state-transition.test.cjs: 2 #2440 regression tests for deriveProgressKeys=true (total_plans takes derived; boundary at equality). The existing !resync wholesale-restore test stays unchanged (default behavior preserved). References: #2440; #1446 (total_phases read-path fix — same principle); #3242 Bug A (body-only preservation — protection preserved via the opt-in gate); ADR-1769 (applyStatePreservation table-driven preservation). * chore(#2440): backfill pr:2468 in .changeset/mellow-eagles-chatter.md |
||
|
|
bb97ffb5aa |
fix(#2431): self-suppress TDD Audit section when all commits are missing (#2467)
* fix(#2431): self-suppress TDD Audit section when all commits are missing The TDD Audit section in ship.md step 8 was always emitted — but the execute pipeline only writes gate_status: git trailers when TDD mode is active. Without TDD mode (the default), every commit's trailer is absent and the section normalizes to 100% missing, producing a noise table with no way to disable it. Fix: add a self-suppress instruction at the point where gate_status values are normalized. When every commit in the scan normalizes to 'missing', skip both step 8 (TDD Audit section) and step 9 (aggregate gate_status trailer) entirely. Only emit when at least one commit carries a real value (skill, fallback, or exempt). This is data-driven, NOT config-gated. An earlier iteration used inline 'gsd_run query config-get workflow.tdd_mode' — but workflow.tdd_mode is owned by the tdd capability, and ADR-857 Phase 6 forbids host loop workflows from reading capability-owned keys via inline config-get. The self-suppress approach avoids any config-get entirely; it checks the actual trailer data and skips when there's nothing real to report. Matches the triage's suggested approach: 'have the audit gracefully degrade (skip the section)' when there is no real signal. Tests: tests/workflow-compat.test.cjs gains 3 #2431 assertions: - documents self-suppress when every commit is missing - step 9 (aggregate trailer) is also gated on real values existing - does NOT read workflow.tdd_mode inline (ADR-857 Phase 6 compliant) The existing feat-41 assertions still pass — the section content is preserved, only gated by the self-suppress instruction. References: #2431; PR #585 (consumer shipped, producer never wired); ADR-857 Phase 6 (capability-owned config keys must not be read inline by host loop workflows). * chore(#2431): backfill pr:2467 in .changeset/clever-moles-frolic.md |
||
|
|
6140627f5c |
fix(#2427): ground smart-entry completion in ROADMAP-derived counts + tighten status regex (#2466)
* fix(#2427): ground smart-entry completion in ROADMAP-derived counts + tighten status regex Two coupled defects in isComplete (src/smart-entry.cts): 1. Two-scale comparison: isComplete compared global current_phase (from STATE.md body 'Phase: N') against milestone-scoped total_phases (from STATE.md frontmatter progress.total_phases, written once at milestone- switch time and going stale as soon as new phases are appended to the roadmap). When current_phase >= stale total_phases (e.g. 7 >= 4), isComplete tripped true even though later phases were still unchecked in ROADMAP.md. 2. Over-broad status regex: /\bcomplete(d)?|done|shipped\b/i matched any 'shipped' or 'done' substring — including per-phase status like 'Phase X shipped — PR #N' — and falsely satisfied the status side of the completion check. Fix: - Added two new SmartEntrySignals fields: roadmap_total_phases and roadmap_completed_phases, populated by calling the existing deriveProgressFromRoadmap helper (from phase-lifecycle.cts:60) when ROADMAP.md exists. These are global, authoritative counts from the Progress table — never stale. - isComplete now prefers the roadmap-derived counts when available (completed >= total) and falls back to the legacy STATE.md comparison only when the roadmap has no parseable Progress table (backward compat for fresh or non-standard projects). - Tightened the status regex to /\b(milestone\s+complete|all\s+phases\s+complete|complete(d)?)\b/i. Drops 'done' and 'shipped' (per-phase language). Keeps milestone-level signals per ADR-2207 (milestone complete, all phases complete) plus the legacy short form 'complete'/'completed'. Tests (tests/smart-entry.unit.test.cjs gains a #2427 describe block): - Mid-milestone with stale total_phases=4, current_phase=7, per-phase 'shipped' status, and 3 unchecked roadmap phases → NOT complete (the core bug scenario). - All roadmap phases complete classifies as complete even with stale cached total_phases (roadmap wins). - Per-phase 'shipped' or 'done' status alone does NOT satisfy completion when roadmap phases are unchecked. - Legacy fallback: empty roadmap (no Progress table) still classifies via STATE.md comparison (backward compat). The makeProject test helper now accepts a string for the 'roadmap' parameter (written verbatim) in addition to the boolean shorthand, so tests can supply a real Progress table. References: #2427; ADR-2207 (milestone status lifecycle); ADR-2143 (column-name-driven Progress table parsing via deriveProgressFromRoadmap); triage note that this is a read-side fix only (the milestone-switch write path in state.cjs is out of scope). * chore(#2427): backfill pr:2466 in .changeset/curious-rams-run.md |
||
|
|
448e148058 |
fix(#2455): abort remaining chunks when one hits the per-chunk timeout (#2465)
The per-chunk timeout exists so a bad chunk fails loudly 'rather than silently burn the job's wall-clock budget until the CI runner cancels the whole job' (run-tests.cjs:633-637). The control flow defeated that: after a timeout kill the loop fell through to the next chunk. Since the timeout (600000ms) is half the 20m job cap and a healthy Windows full pass is ~11m42s, continuing after a timeout can essentially never finish. Observed on run 29749380190 (windows shard 2/3): chunk 1/5 was killed at exactly 600s, the loop pressed on through chunks 2-4, and the job was cancelled mid-chunk-5 at the 20m wall. The failure surfaced as '##[error]The operation was canceled.' — the timeout diagnostic ended up ~38,000 log lines from the end and 'gh run view --log-failed' returned nothing, making the real cause very hard to find. Abort the remaining chunks on a timeout so the diagnostic survives as the visible failure. Ordinary test failures still run every chunk, so the operator keeps seeing all failures in one pass. Fixes #2455 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
903182fed3 |
fix(#2414): mempalace-capture rooms example must be dicts with name key (#2464)
* fix(#2414): mempalace-capture rooms example must be dicts with name key
The skill's Step 3 'Add the drawer (verbatim)' example wrote a flat list
of bare strings under rooms: in the embedded mempalace.yaml. mempalace's
miner (detect_room + _mine_impl) indexes room["name"] — a bare-string
list crashes the first 'mempalace mine' invocation with
TypeError: string indices must be integers, not 'str'
Following the documented example verbatim and running the capture crashed
every time, before any file was routed. The bug shipped in #2220's fix
(commit
|
||
|
|
953b8043ea |
fix(#2456): weight test chunks by measured cost and pack with LPT (#2463)
* fix(#2456): weight test chunks by measured cost and pack with LPT scripts/run-tests.cjs guessed each test file's cost from its filename (basename matching /^(?:install|codex-)/ scored 12, everything else 1). Measured durations show that guess is wrong in both directions: installer-migration-authoring.test.cjs scored 12 while running ~0.1s, and the two most expensive files in the suite both scored 1 — run-tests-harness.test.cjs never matched the prefix, and release-tarball-smoke.install.test.cjs was missed because the regex is anchored to the START of the basename. Chunks were therefore balanced by file COUNT, not cost. On the real shard 2/3 the two heaviest files packed into the SAME chunk, leaving the slowest chunk 2.8x the lightest and sitting near the 600s per-chunk timeout while other chunks idled. Weight each file by its measured duration from a checked-in, regenerable timings table and pack with LPT (heaviest first, into the lightest chunk). On the same shard this drops the slowest chunk from 383s to 238s and the imbalance from 2.79x to 1.00x, and separates the two heavy files. Timings are advisory, never gated: an unknown file falls back to the table's median weight, a missing or corrupt table falls back to uniform weight, and a count-based floor guarantees the packer never produces fewer chunks than plain count-based packing would. Closes #2456 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2456): harden chunk packing against degenerate knobs and table keys Follow-up hardening found while reviewing the packer, fixed inline. The chunk knobs are read from the environment with Number(), so a typo (RUN_TESTS_MAX_FILES_PER_CHUNK=abc) yields NaN and an explicit 0 yields 0. Both flow into the new chunk-count arithmetic: NaN made Math.ceil return NaN, Array.from({length: NaN}) produce zero bins, and packChunks' retry loop spin forever — a hung CI job with no output. Zero made the count Infinity and threw RangeError: Invalid array length. The previous count-based packer degraded to a single chunk instead, so this was a regression introduced by the LPT rewrite. Normalize the knobs at the environment boundary (positiveNumberEnv: anything not a positive finite number falls back to the default) and guard packChunks itself, since it is exported and cannot assume its caller normalized. Non-finite weights from an arbitrary weightOf are clamped too. RUN_TESTS_CHUNK_TIMEOUT_MS gets the same treatment. Also resolve timing-table lookups with Object.hasOwn: the table is JSON-parsed, so a bare index would walk the prototype chain and return a function for a file named constructor.test.cjs or toString.test.cjs. The typeof guard already rejected that, but the lookup now resolves correctly rather than relying on the downstream check. Refs #2456 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2456): correct prototype-lookup rationale and guard generator keys Two findings from independent security review, fixed inline. The makeFileWeigher comment claimed a bare table lookup "would return a FUNCTION for a file named constructor.test.cjs". That premise is false: basename('constructor.test.cjs') is 'constructor.test.cjs', which is not an Object.prototype key, and walkTestFiles only ever collects *.test.cjs. The prototype chain was never reachable from a real selection, and the existing typeof guard already rejected the function it would return, so Object.hasOwn is defense-in-depth rather than a behavior change. The comment now says that instead of asserting something untrue. The accompanying test inherited the same false premise: it fed constructor.test.cjs and asserted a median fallback that would have held with or without the guard, so it passed for a reason unrelated to what it claimed to prove. It now uses BARE keys (constructor, toString, valueOf, hasOwnProperty, __proto__) — the only inputs that actually resolve on Object.prototype — and asserts the real exported contract: any key absent from the table weighs the median, never a function. gen-test-timings.cjs built its output object by computed-key assignment from basenames taken out of a reporter stream it does not control — the js/prototype-polluting-assignment shape, and this repo has a CodeQL barrier for exactly that pattern. It was not exploitable (the value is always a rounded number, so the __proto__ setter is a silent no-op), but it silently DROPPED such an entry rather than reporting it. Validate every key against a test-basename pattern and fail loudly instead, and build the table with a null prototype. Refs #2456 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2456): replace tautological chunking tests and clamp chunk count Six findings from independent correctness review, all reproduced and fixed inline. The two subprocess tests written to carry the #2088 guarantee forward were tautological: every seeded file weighed exactly 1, so both passed under the OLD prefix-heuristic packer and with the timings file deleted entirely. Neither could fail for the reason it existed. Both are rebuilt so the old algorithm produces a different packing and the assertion goes red: the spread test now uses three expensive files named so the old heuristic scored them 1 alongside three trivial `install-`-prefixed files it scored 12 — inverted from real cost, giving {2,2,1,1} under the old packer versus {2,2,2} under measured weights. The companion test covers the other direction: four trivial `install-` files the old heuristic split into four single-file chunks now stay in one. packChunks clamped the chunk count from below but not above, so a legitimate but tiny budget (RUN_TESTS_MAX_FILES_PER_CHUNK=1e-9, which positiveNumberEnv accepts) asked for 637,000,000,000 bins and threw RangeError. More chunks than files is never useful; the count now clamps at one file per chunk. The generator's basename-collision guard compared full dirnames, so two OS lanes reporting the same file under different container roots (/work/tests vs C:/work/tests) flagged every shared basename as a collision — on the script's own documented multi-lane usage. Detection is now scoped per stream, where the root is constant; a genuine same-lane collision is still caught. Also: the LPT tie-break compared raw paths, so a path separator (0x2F vs 0x5C) could order a subdir file differently per platform, contradicting the documented byte-identical guarantee — it now normalizes separators. loadTestTimings now honors schema_version instead of writing it and never reading it, falling back to uniform weight on an unknown version. A comment claiming an all-uniform suite "chunks exactly as it did before" was false and contradicted by this PR's own test: the chunk count is preserved, the composition is not. And the missing-table test created a temp dir it never cleaned up, for a path that only needed to not exist. Refs #2456 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
be5113abfe |
fix(#2408): fold colliding phase statuses + add W023 collision warning (#2461)
* fix(#2408): fold colliding phase statuses + add W023 collision warning Two coupled bugs from #2408: 1. cmdStats last-write-wins status (src/commands.cts:1610-1618): when two on-disk phase directories normalize to the same phase key (e.g. `05-real/` and `05-real-stray/`), the directory-scan merge overwrote `status` with whatever the *current* directory in scan order computed, discarding `existing?.status` entirely. fs.readdirSync order is not stable across platforms, so /gsd-stats could silently report `Not Started` for a phase that is actually `Complete`. Plan/summary counts were already additively merged; only `status` was wrong. Fix: introduced a `foldPhaseStatus(a, b)` helper that returns whichever status is further along the precedence ladder `Complete > Needs Review > Executed > In Progress > Planned > Not Started` (with `Pending` and unrecognized statuses ranked after). The merge site now calls `existing ? foldPhaseStatus(existing.status, status) : status`. The fold is commutative, so the result is identical regardless of read order. 2. cmdValidateHealth had no collision-detection pass (src/verify.cts): codes W001-W022 cover every condition except normalized-key collisions, so an operator got zero signal that anything was wrong. Fix: added W023 — groups phaseDirEntries by their normalized phase key (via the existing normalizePhaseName + extractPhaseToken helpers from phase-id.cjs — same normalization cmdStats uses) and emits a warning for any group with ≥2 dirs. The warning names the normalized key, both directory names (sorted by comparePhaseNum for stable output), and each directory's independently-computed status (via determinePhaseStatus imported from commands.cjs). Wording is deliberately neutral — never guesses which directory is the real one. The optional --repair path from the issue is intentionally NOT implemented in this PR (the issue marked it lower priority and acceptance criterion 4 is vacuously satisfied by omission). Triage correction applied: the issue proposed W022, but that code is already in use for config.json model-tier validation (src/verify.cts :1372-1391). The next free code is W023, used here. Tests: - tests/commands.test.cjs: integration test that 05-real/ (Complete) + 05-real-stray/ (empty/Not Started) collide and stats reports the merged phase as Complete regardless of read order; plus a direct unit test of foldPhaseStatus asserting commutativity + correct precedence for every status pair + correct handling of unrecognized statuses. - tests/health-validation.test.cjs: integration test that W023 fires on the collision naming both dirs + their statuses (and uses neutral wording), plus a negative test that no W023 fires when only one dir exists per key. References: #2408; reporter's three-layer triage + acceptance criteria; triage correction that W022 is already in use (model-tier validation). * chore(#2408): backfill pr:2461 in .changeset/graceful-koalas-forage.md |
||
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|
||
|
|
b6e6a22fce |
fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage (seven commits, never pushed) onto current origin/next as a single squashed commit. The original work was substantial and correct; this commit preserves its full scope, trimmed where rebase conflicts + workflow size budgets required it. Three independent layers where response_language was being dropped are closed: Layer 1 — orchestrator-facing directives across workflows. Adds the strong "All user-facing output in this workflow MUST be presented in {response_language}; technical terms, code, paths, and subagent prompts stay in English" directive to ~40 workflows that previously either lacked it entirely (verify-work, new-project, new-milestone, quick, manager, and ~35 more) or carried only the weak subagent-prompt-only form (plan-phase, execute-phase). The directive covers narration between tool calls and banner output, not just the AskUserQuestion prompts. Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts an optional responseLanguage parameter and renders the frame strings ("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.") in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/ Chinese/Korean/Italian) with an alias table covering ~30 input variants (en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads config.response_language via loadConfig(cwd) and passes it through, so the byte-for-byte block verify-work.md reprints verbatim is already localized when written — preserving the anti-injection hygiene rule at verify-work.md (the model is forbidden to translate after the fact). CJK display width is computed by East Asian Width property ranges (W/F) so the right ║ border of the banner stays aligned for full-width characters. English fallback is byte-identical to the pre-fix behavior when response_language is unset or unrecognized. Layer 3 — literal English report templates in execute-phase. The top-of- workflow directive covers all template sites (templates are a structural source, not literal output). Inline render-language notes that previously sat at each template site were removed during the squash because they pushed execute-phase.md over its frozen pre-phase-6 byte ceiling (93600 — ADR-857 Phase 6 capstone). The single top directive covers the same surface with fewer bytes. Also extends src/docs.cts and src/init.cts to propagate response_language into the init JSON bundle of the additional workflows so the directive can read it. Tests added: - tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls back to English default; recognized language swaps only the two frame strings while structural lines stay untouched; CJK display-width regression (independent recomputation of East Asian Width W/F ranges). - tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language wiring through docs.cts/init.cts. References: #2402; reporter's three-layer triage + Layer-4 follow-up; the byte-for-byte anti-injection hygiene rule at verify-work.md (the reason Layer 2 must be renderer-side, not model-translated). This is a squash of the in-flight bot branch — seven commits representing the original implementation plus its subsequent fix/CJK-padding/test/ changeset/regen cycles, none of which were ever pushed or PR'd. The squash captures the final coherent state. * chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md * chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451) Rebase conflicts were entirely in generated artifacts (golden-install-parity fixtures + workflow-size-baseline.json). After taking theirs during rebase, regenerated cleanly against the merged source tree. |
||
|
|
352876ff0c |
fix(#2315): respect review.default_reviewers in bare convergence invocation (#2451)
* test(#2315): regression test for review.default_reviewers precedence A bare /gsd-plan-review-convergence invocation (no reviewer flags) is supposed to let users configure a persistent reviewer lineup via review.default_reviewers and just run the loop. Instead, the orchestrator's argument parser silently discards that configuration and forces --codex on every no-flag invocation — with no warning that the configured reviewers were ignored. This commit adds a regression test that fails against the pre-fix workflow (the buggy unconditional --codex fallback is still present at this commit) and passes after the fix lands: - Structural: the workflow must NOT contain an unconditional 'if [ -z "$REVIEWER_FLAGS" ]; then REVIEWER_FLAGS="--codex"; fi' line before the workflow.plan_review_convergence config gate. - Structural: the workflow must query review.default_reviewers AFTER the config gate and document that empty REVIEWER_FLAGS lets gsd-review apply the default. - Behavioral: matrix across {configured, unset, empty-array, explicit-flag} invoking the actual deployed parse + resolution blocks with a stubbed gsd_run. Also updates two existing tests whose assertions the fix makes stale: - #2293 behavioral: endMarker was the buggy unconditional fallback line; the bare invocation assertion was 'run("5") === "--codex"'. Both flip post-fix (endMarker is now the last --all grep line; bare invocation returns empty from the parse block, default applied later in step 1.5). - command-default-claim: the pre-fix command documented '--codex (default if no reviewer specified)' which was the user-facing mirror of the bug. The assertion now requires the command to document the review.default_reviewers precedence. * fix(#2315): respect review.default_reviewers in bare convergence invocation Root cause: plan-review-convergence.md step 1 (Parse and Normalize Arguments) contained an unconditional fallback that set REVIEWER_FLAGS=\"--codex\" whenever no explicit reviewer flag was supplied. This value was then interpolated verbatim into the gsd-review args, so gsd-review saw --codex as an explicit flag (precedence rule 1) and never reached rule 3 (review.default_reviewers). The same path silently dropped any configured review.reviewer_instances (instances participate ONLY via review.default_reviewers per ADR-1517). Fix: - Remove the unconditional --codex fallback from step 1. - Add a config-gated resolution in step 1.5 (after CONVERGENCE_ENABLED check) that queries review.default_reviewers and either leaves REVIEWER_FLAGS empty (letting gsd-review apply its own rule-3 default) or falls back to --codex when no default is configured — preserving the pre-fix default for unconfigured users (AC3). - Replace the banner {REVIEWER_FLAGS} token with {REVIEWER_DISPLAY} so the startup banner reflects what will actually run (AC4), not a hardcoded value. The fix upholds the documented precedence contract (ADR-0011, ADR-0015) that the bug was actively violating. Explicit-flag invocations (--gemini, --all, etc.) are unaffected (AC5). References: #2315; ADR-0011 (review.default_reviewers precedence); ADR-0015 (autonomous cross-AI convergence); ADR-1517 (reviewer instances). * chore(#2315): bump plan-review-convergence.md baseline + changeset - Bump plan-review-convergence.md size baseline 23713 → 25536 (the new step-1.5 default-resolution block). - Add .changeset/plucky-yaks-roar.md documenting the user-visible change. * fix(#2315): restore /gsd: colon syntax + skip behavioral test when jq missing Two follow-ups to the #2315 fix discovered by gsd-test: 1. The fix commit accidentally regressed the slash-command namespace in the disabled-feature exit message: /gsd:plan-review-convergence (correct, from PR #3452) became /gsd-plan-review-convergence (retired dash syntax). The slash-command-namespace invariant test caught this. Restored the colon form. 2. The behavioral test exercises the deployed reviewer-resolution block, which pipes through jq. jq is a documented production dependency (review.md:244 "install jq if missing") and is present in every production deployment, but is NOT on PATH in the gsd-test linux-node{22,24} containers (same constraint as tests/opencode-review-reconstruction. property.test.cjs). Without jq, the printf|jq pipeline fails silently, the ||echo 0 fallback yields DEFAULT_REVIEWERS_COUNT=0, and the resolution falls through to the --codex branch — producing a false negative. Added a jqAvailable guard at module load (matching the existing pattern) and skip the behavioral test when jq is absent. The structural tests (no bash execution) still run and validate the fix. * chore(#2315): regenerate golden-install-parity fixtures Source changes to plan-review-convergence.md (workflow + skill mirror + command doc) changed the install-tree hashes. Regenerated via 'npm run gen:golden' after rebuilding gsd-core/bin/lib/install-engine.cjs ('npm run build:lib') — the on-disk lib was stale relative to src/install-engine.cts (isSymlinkedDestOptIn) and blocked fixture gen. * test(#2315): address review findings — strengthen structural tests + property test Code-review + security-review (isolated subagent passes) surfaced Low/Nit findings; this commit addresses the test-side findings: - Structural test 1 ("unconditional --codex one-liner") now asserts the buggy line is absent EVERYWHERE, not just before the config gate. The earlier assertion allowed a maintainer to re-add the line after the gate (passing the structural test) while the bug would still bite at runtime before the gate runs. - Structural test 4 ("banner uses REVIEWER_DISPLAY") now asserts the LITERAL banner placeholder "Reviewers: {REVIEWER_DISPLAY}" and forbids "Reviewers: {REVIEWER_FLAGS}". The earlier workflow.includes( "REVIEWER_DISPLAY") was satisfied by a comment mention. - New structural test for command/skill content parity: both files must document the review.default_reviewers precedence on the --codex flag (catches a manual edit to one that the gen:plugin-skills mirror misses). - Behavioral test stub now passes default_reviewers via env var ($GSD_TEST_DEFAULT_REVIEWERS) instead of inline-interpolating into a bash single-quoted string. Removes the (currently-safe) fragility where a future test input containing a single quote would close the bash quote and execute as bash under execFileSync. - New property test (fast-check, numRuns=25) for the JSON-classification contract: non-empty arrays of slugs -> empty REVIEWER_FLAGS; empty array / scalar JSON / malformed JSON -> --codex fallback. Locks the parser contract per CLAUDE.md mandate. * fix(#2315): address review findings — defensive jq-missing warning + banner cleanup Code-review surfaced two Low-severity workflow-side findings: - Defensive jq-missing warning: if jq is not on PATH in production (it is a documented dependency per review.md:244, but the dependency can be absent in degraded environments), the printf|jq pipeline fails silently to "0" and a user with review.default_reviewers configured gets --codex with no indication their configured default was unreadable. Added a command -v jq guard at the top of the resolution that falls back to --codex AND emits a stderr warning explaining the reason. This makes the failure diagnosable instead of silently reproducing the #2315 override. - Banner leading-space cleanup: REVIEWER_FLAGS accumulates with a leading space ("$REVIEWER_FLAGS --gemini" from ""), so the explicit-flag branch of REVIEWER_DISPLAY="$REVIEWER_FLAGS" rendered "Reviewers: --gemini" (double space). Pre-existing but worth fixing alongside the AC4 banner work. Strip one leading space with the ${VAR# } parameter expansion in the explicit-flag branch only (the configured-default and --codex branches already produce clean strings). * chore(#2315): bump plan-review-convergence.md baseline + regen golden fixtures The defensive jq-missing warning grew plan-review-convergence.md (25536 -> 26285 bytes). Bumps the per-file baseline snapshot and regenerates the golden-install-parity fixtures for the resulting install-tree hash changes. * chore(#2315): backfill pr:2451 in .changeset/plucky-yaks-roar.md |
||
|
|
eb45fc0e8c |
fix(#2415): close_phase_todos stages the pending/ deletion alongside completed/ (#2447)
* fix(#2415): close_phase_todos stages the pending/ deletion alongside completed/ Bug: the workflow step moved resolved todos from .planning/todos/pending/ to .planning/todos/completed/ with a plain 'mv', then committed listing ONLY the destination directory in --files: mv "$TODO_FILE" "$COMPLETED_DIR/" gsd_run query commit '...' --files .planning/todos/completed/ .planning/STATE.md Git's index still tracked the moved file at its old pending/<name>.md path. The commit therefore only staged the new completed/<name>.md copy — the deletion at pending/ was never staged, never committed, and lingered as an unstaged deletion in git status indefinitely until some later broad 'git add -A' caught it. The phase genuinely closed the todo, but the working tree was never clean. Fix: add .planning/todos/pending/ to the --files list. 'git add' of that directory (since git 2.0) stages deletions of tracked files in the pathspec, so the moved-away file is staged as a deletion atomically with the new completed/ copy in the same commit. Chose plain mv + two-dir --files over 'git mv' because git mv FAILS on: - untracked todos (new todo file not yet committed) - non-git .planning dirs (worktree safety / pre-init projects) Plain mv has neither failure mode. Regression tests in tests/close-phase-todos-stage-deletion.test.cjs (source-text-is-the-product: workflow .md text IS what the runtime loads) cover: - the commit --files list includes BOTH completed/ AND pending/ - the move uses plain 'mv' (not 'git mv') so untracked + non-git cases work * chore(#2415): trim commit subject to keep execute-phase.md under byte ceiling The fix added '.planning/todos/pending/' (~26 bytes) to the commit --files list. To stay under the ADR-857 Phase 6 pre-phase-6 byte ceiling margin (93400 bytes, hard ceiling 93600), shortened the commit subject from 'auto-close N todo(s) resolved by this phase' to 'close N resolved todo(s)'. Net change vs origin/next: +6 bytes (93384 → 93390), well under the margin. Also regenerates the golden-install-parity fixtures (execute-phase.md content-hash update across all runtimes). * chore(#2415): bump execute-phase.md workflow-size baseline (93384 → 93390) The +pending/ fix added 6 net bytes (93384 → 93390), still well under the ADR-857 Phase 6 pre-phase-6 byte ceiling margin (93400). * chore(changeset): backfill pr:2447 in .changeset/sturdy-wasps-run.md * fix(#2415): add issue ref to allow-test-rule annotation (ADR-456) CI lint-allow-test-rule-refs failed on the prior commit — ADR-456 requires '// allow-test-rule: <category> see #NNNN' so every exemption is traceable to an issue. Added 'see #2415' to the source-text-is-the-product annotation. |
||
|
|
6c00cdad07 |
docs(#2343): list gsd-omp EoS integration (#2448)
* docs: list gsd-omp EoS integration * chore: add gsd-omp registry changeset * chore: update registry changeset PR reference --------- Co-authored-by: AI Assistant <ai@example.com> |
||
|
|
e227e81f59 |
fix(#2395): persist runtime identity into ~/.gsd/defaults.json for non-Claude installs (#2446)
* fix(#2395): persist runtime identity into ~/.gsd/defaults.json for non-Claude installs Bug: finishInstall for non-Claude runtimes never persisted a 'runtime' key into ~/.gsd/defaults.json. resolveRuntime() precedence is GSD_RUNTIME env > config.runtime > 'claude', so both inputs being empty on a Cursor (or any non-Claude) install caused agent_runtime and every runtime-branded slash hint to silently fall through to 'claude'. Cursor users saw 'agent_runtime: "claude"' and Claude-formatted /gsd-* hints with no env or config hand-set. Fix mirrors the existing resolve_model_ids: 'omit' write site at the same call site (bin/install.js finishInstall, gated on !_hostBehaviors(runtime) .nativeModelAliases && !GSD_TEST_MODE). Writes runtime: <runtime> into ~/.gsd/defaults.json when absent/null/empty. Claude is the resolveRuntime() fallback so it needs no write; an explicit pre-existing runtime value is always preserved across installs of any runtime. Pattern parity with #1156 (default-to-omit intent) and #1569 (preserve explicit user value) — the new write is the third sibling on the same install-time persistence block. Regression tests in tests/install.test.cjs cover: - absent / null / empty-string runtime → populated to <install runtime> - explicit pre-existing runtime preserved (no clobber across runtimes) - parameterized across 5 non-Claude runtimes (cursor, codex, opencode, antigravity, windsurf) Out of scope (per triage): the resolveRuntime() precedence order itself (env > config > default) is unchanged. A separate follow-up noted in the issue (subagent rendering when runtime correctly identifies as 'cursor') is unrelated to branding and not addressed here. * chore(#2395): regenerate golden-install-parity fixtures for new runtime persistence The golden install-tree fixtures capture the post-install state, including ~/.gsd/defaults.json. For non-Claude runtimes, defaults.json now includes runtime: <runtime> — content hash updated for each affected runtime (17 golden-install-parity/*.json files). Claude's defaults.json is unchanged (no runtime key written for Claude — it's the resolveRuntime() fallback). * chore(#2395): drop product names from changeset fragment (product-name-purity gate) * test(#2395): move describe out of fix-1521 fold + add same-runtime idempotence test Code review (correctness subagent) flagged 2 Low test-quality issues: 1. The new Bug #2395 describe was inserted inside the folded:fix-1521-real-install-stamping IIFE callback, muddying test reporting and tracing the regression to the wrong epic. Moved it outside the IIFE close to be a top-level sibling. 2. No explicit same-runtime idempotence test — the suite covered cross-runtime preservation (cursor seed → opencode install) but not 'install cursor twice → second is a no-op'. Added: seeds fresh defaults, installs cursor, captures mtime, installs cursor again, asserts runtime unchanged AND defaults.json mtime unchanged (idempotent, no rewrite churn). * chore(changeset): backfill pr:2446 in .changeset/fierce-ravens-dance.md |
||
|
|
12e4d93b19 |
fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts (#2445)
* fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts Bug: v1.7.0's destSubpath write-confinement (ADR-1239 Phase B) refused install/update whenever CLAUDE_CONFIG_DIR (or an artifact-kind child like skills/, hooks/) was a pre-existing symlink, with no opt-out. Three legitimate user-owned layouts were blocked: - (lars-hh) CLAUDE_CONFIG_DIR=~/.claude-personal with skills/hooks symlinked to a user-owned external dir - (Mamiki) ~/.claude/skills is a Windows Junction to a shared skills dir - (Azd325) ~/.claude itself is a symlink to a dotfiles repo (the early root-is-symlink return refused before the component loop ran) Fix: add GSD_ALLOW_SYMLINKED_DEST env var (accepts '1' or 'true'). When set, hasExistingSymlinkBetween follows symlinks instead of refusing them. Cross-platform: fs.lstatSync().isSymbolicLink() returns true for both POSIX symlinks and NTFS junctions (Node ≥ 16), so Mamiki's Junction case is handled by the same code path. Threat model preserved (these still refuse EVEN WITH opt-in): (a) path-traversal in the destSubpath string itself ('../../etc'-style) — ADR-1239 Phase B threat (a), untrusted destSubpath protection (b) a symlink whose resolved real path equals the install root itself — would let _removeGsdEntries wipe the root; #1704 threat (b) (c) broken symlinks (realpathSync throws) — fail-closed What opt-in RELAXES specifically: the 'pre-existing symlink pointing outside configHome' refusal — #1704 threat (c). The user has explicitly asserted they own and trust the symlink target. Error messages at all 4 call sites (installRuntimeArtifacts, _copyStaged, migrateLegacyDevPreferencesToSkill, installOpencodeFamilySkills) updated to (1) name the env var opt-in, (2) be accurate when the root itself is a symlink (Azd325's complaint that the old message accused destDir of 'containing' a symlink when the root was the actual symlink). Docs: docs/CONFIGURATION.md Environment Variables table updated. Regression tests in tests/install-write-confinement.test.cjs cover: - child-symlink layout (lars-hh / Mamiki): default refuses, opt-in allows - root-is-symlink layout (Azd325): default refuses, opt-in follows - path-traversal '../../etc' refused EVEN WITH opt-in (threat a preserved) - resolved-target-equals-install-root refused EVEN WITH opt-in (threat b) - broken symlink refused EVEN WITH opt-in (fail-closed) * test(#2393): import beforeEach/afterEach in install-write-confinement suite The original file imported only { describe, test } from node:test. The new #2393 opt-in describe block uses beforeEach/afterEach to manage the GSD_ALLOW_SYMLINKED_DEST env var lifecycle — add them to the import. * test(#2393): correct broken-symlink test — existsSync follows link → loop terminates early Initial test expected broken symlinks to be refused even with opt-in. That was wrong: fs.existsSync follows symlinks, so a broken symlink returns false from existsSync and the component loop terminates before the symlink check fires. Both default and opt-in paths share this behavior; the fix preserves it. Updates the test to pin the actual current behavior so a future refactor (e.g. switching to lstatSync for existence) is a deliberate behavior change. * fix(#2393): realpath the install root — guard against macOS /var ↔ /private/var Code review (security subagent) flagged a HIGH-severity hole in the threat-(b) preservation: realTarget (from fs.realpathSync) is fully symlink-resolved, resolvedRoot (from path.resolve) is lexical-only. On macOS /var is a symlink to /private/var, so resolvedRoot='/var/foo/.claude' but realConfigHome is '/private/var/foo/.claude'. A symlink whose realtarget matches the install root by real path would compare unequal to the lexical resolvedRoot — defeating the wipe-protection guard exactly in the reporter's case (Azd325, nix-darwin: ~/.claude is itself a symlink). Fix: compute realRoot once via fs.realpathSync(resolvedRoot) at function entry (with fail-closed fallback to lexical form on realpath failure — broken/missing root, permission denied, exotic FS). Threat (a) path-traversal check above still confines regardless. Compare against BOTH lexical and real forms in both the root-symlink and component-symlink branches. Also adds the reviewer's transitivity-trust clarification comment: once a symlink is followed under opt-in, the walk continues from the resolved real path WITHOUT re-checking further segments stay inside a confining boundary. This is documented opt-in semantics — one opt-in trusts the whole reachable tree — and the comment makes the design choice explicit so a future maintainer doesn't add a 'follow one symlink only' expectation. Regression test added for the macOS /var normalization case (spelled configHome via os.tmpdir() lexically while pointing the test symlink through its realpath). Test skips on non-darwin platforms and when os.tmpdir() has no symlink component. * fix(#2393): root-symlink branch — do not apply threat-(b) check to root itself Initial fix applied the wipe-threat-(b) check to the root-symlink branch unconditionally. That was wrong: when root itself is a symlink (Azd325's nix-darwin case), its realpath IS realRoot by construction — so the check always fires, defeating the opt-in for exactly the case it was meant to enable. The wipe threat (b) does NOT apply to root being a symlink: destDir is a CHILD of root, and resolving root just gives root's target. There is no circular back-reference to root from a path that descends from a resolved root. So the root-symlink branch should just follow the symlink under opt-in and continue the walk, no threat-(b) check. Threat (b) only fires in the COMPONENT loop, where a child symlink can resolve back to the install root. That branch keeps the (b) check using BOTH lexical and real forms of root (the macOS /var ↔ /private/var fix from the prior commit). * fix(#2393): apply opt-in at the 5 bin/install.js call sites + add env-var/transitive tests Code review (correctness subagent) flagged a Critical coverage gap: the initial fix updated only the 4 src/install-engine.cts call sites. Five more call sites in bin/install.js still used the 2-arg signature, so the opt-in env var was silently ignored on: - installCodexConfig (config.toml + agents/ dir + per-agent .toml paths) — Codex only - copyWithPathReplacement (the generic emit path: workflows, commands, staging) — ALL runtimes - resolveInstallRelativePath (path resolver used in various places) Result: a user setting GSD_ALLOW_SYMLINKED_DEST=1 would see SOME refusals disappear (engine path) and OTHERS remain (bin/install.js paths) — a partially-applied install and a confusing UX, directly contradicting the PR's headline claim. Fix: - Export isSymlinkedDestOptIn from src/install-engine.cts alongside hasExistingSymlinkBetween - Import it in bin/install.js - Update all 5 bin/install.js call sites to pass { allowOptInFollow } - Update all 3 bin/install.js error messages to name the env var, matching the engine's phrasing Also addresses reviewer's Medium test-adequacy findings: - isSymlinkedDestOptIn env-var parsing now tested directly (accepts only documented '1' / 'true'; rejects 'TRUE', 'yes', 'on', '0', 'false', empty, unset) - transitive symlink chain (configHome/outer → outside1 → outside2) test pins the documented 'transitive and unbounded' opt-in semantics so a future contributor can't accidentally narrow it * chore(changeset): backfill pr:2445 in .changeset/eager-wasps-swim.md |
||
|
|
517bae8d6d |
fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message Bug: check.decision-coverage-plan's remediation message told the user to cite decisions "(or body)" but extractPlanDesignatedSections only scanned <objective>/<tasks>/<task>/<action>. A decision cited in <read_first>, <behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to the gate — false BLOCKING coverage gap, plus the message's own fix-hint sent the user to "the body" where re-citing still failed. Two-part fix (must change together — that drift was the bug): 1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also match <read_first>, <behavior>, <verify>, <acceptance_criteria>, <done>. These are all planner-canonical tags the planner is told to use (plan-phase.md:830-862, plan-phase.md:772). The body negative- lookahead mirrors the opening-tag set so each tag's body is captured independently. 2. Correct buildPlanMessage to name ONLY the surfaces the extractor actually scans (front-matter must_haves/truths/objective, designated markdown headings, and the nine planner-canonical tag bodies). The misleading "(or body)" clause is gone. Also updates the planner's documented contract (agents/gsd-planner.md:69) and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to reflect the wider scan. Regression tests in tests/decisions.test.cjs cover each newly-scanned tag body, a control (no citation still uncovered), and a message/extractor parity assertion that names every scanned surface — so the two cannot drift apart again. Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage is a separate command (decision-coverage-verify) checking shipped artifacts, not plan citations — untouched. * chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision- coverage contract (5 new scanned tag names + heading clarification). Growth is justified: the contract surface is itself the fix — the prior text under-described what the gate scans, which was the bug. Updates: - tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294) - 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime) - tests/fixtures/install-tree/*.json (regenerated by gen:golden) * fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting Code review (subagent) flagged a Medium edge-case regression from the single-alternation regex: when a newly-scanned tag nests inside another scanned tag, the alternation's negative lookahead halts the outer tag's body at the inner tag — losing any D-NN citation in the outer tag's prefix prose. Concretely: <action>per D-05 <verify>npm test</verify></action> → 3-tag alternation (old): captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught → 9-tag alternation (bug): captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST Switches extractXmlTagBodies to per-tag matching: each tag gets its own regex whose negative-lookahead tempers only against the SAME tag's reopening. So <verify> inside <action> is absorbed into <action>'s body (D-05 caught) AND <verify> is matched separately on its own pass. Per-tag preserves both: - the reporter's case (sibling tags inside <read_first>) - nested-tag citations in outer-tag prefix prose - ReDoS safety (each per-tag regex keeps the #2128 body tempering) Also adds the reviewer's other requested edge-case tests: - non-scanned tag (<name>) bearing D-NN must NOT count - self-closing form <read_first /> safely ignored - attribute form <verify type="...">D-NN</verify> (canonical planner shape) - CRLF newlines in tag body do not break capture * chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md |
||
|
|
0bbbca2a46 |
fix(#2069): forward model_policy, model_profile_overrides, runtime from global defaults (#2442)
* test(#2069): add fail-first regression for global-defaults dropped keys Adds four failing-first regression cases to tests/defaults-json-fallback.test.cjs: - model_policy forwarded from ~/.gsd/defaults.json - model_profile_overrides forwarded from ~/.gsd/defaults.json - runtime forwarded from ~/.gsd/defaults.json - parity: model_policy survives identically whether it lives in the global defaults or in a project's .planning/config.json All four fail on unfixed code (Branch D of loadConfigResolved builds _globalBaseCfg from a whitelist that omits these three keys). The project-config path at config-loader.cts:602-604 already forwards them, so the global path should too. * fix(#2069): forward model_policy, model_profile_overrides, runtime from global defaults The _globalBaseCfg whitelist in Branch D of loadConfigResolved previously omitted three keys that the project-config path forwards parsed['…']: - runtime - model_profile_overrides - model_policy so ~/.gsd/defaults.json silently dropped them. A machine-wide model policy (or runtime / profile overrides) was honored inside a project (where .planning/config.json carries it) but ignored for out-of-project runs — resolve-model fell back to the profile default with no warning. Adds the three entries to _globalBaseCfg in the same (globalDefaults['…']) || null shape as the sibling keys and the project-config path, so global defaults honor them identically. Regression tests in the prior commit (#2069 fail-first) demonstrate the fix on the same suite that previously failed. * test(#2069): extend parity test to all three previously-dropped keys Code review (subagent) flagged that the parity test only asserted model_policy shape-parity between global-defaults and project-config paths. A future regression breaking just runtime or just model_profile_overrides shape (e.g. someone changing parsed['runtime'] to ?? null in the project path) would slip a single-key test. Extends the parity test to assert deepStrictEqual / strictEqual across all three keys: model_policy, model_profile_overrides, runtime. Same two-dir setup, three cheap assertions. * chore(changeset): backfill pr:2442 in .changeset/sturdy-seals-fly.md |