8249ebcf6e75a986598edc59c3159fca4d2ca07e
356 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8249ebcf6e |
fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do not exist yet; every row fails on require. Per #3770 only an intentional target-test failure may authorize GREEN; zero-test discovery, fixture crashes, unrelated failures, and unexpected green are INVALID_RED. * fix(3770): require intentional RED evidence before GREEN Only an intentional failure of the TARGET test (distinctly named, TAP-reported assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test discovery, fixture/load crashes (file-named failures), nonzero exits without a failing test, unrelated failures, unexpected greens, and malformed/missing records are INVALID_RED and block GREEN. - src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses the prohibition-enforcement TAP primitives; fail-closed, never throws) - check tdd-red-evidence <record.json>: validates the persisted record (command, exit code, failing test, expected, actual) - gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now requires the evidence record + gate verdict, not a nonzero exit or a RED: tag * chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs * fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib - gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B < 49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid) - tests: the row-6 fixture used String.replace (first-occurrence), so the `not ok` line still named the target test and the classifier was right to accept it; replaceAll makes the failure genuinely unrelated - eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs (lint the src/*.cts source, per ADR-457 migration rule) Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line * chore(3770): add changeset * chore(3770): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
97ce61dee2 |
fix(#3990): state the RED/GREEN/REFACTOR cycle once, embed tdd.md conditionally (#4228)
* test(#3990): the RED/GREEN/REFACTOR cycle is stated once, embedded conditionally * fix(#3990): state the cycle once — pointers in consumers, conditional tdd.md embeds Emitted-Drift-Ack-Growth: execute-phase.md — #3990 conditions the tdd.md embed on the dispatch being TDD * chore(#3990): changeset for the single-statement TDD cycle * chore(#3990): backfill changeset pr number * fix(#4228): linear cycle check — the lazy-span regex pinned a Windows core for the whole job cap * test(#3990): allowlist pin tracks the rebased line * fix(tests): npm-integrity gate names an empty audit output explicitly — empty stdout crashed the parse as a bare SyntaxError * fix: name an empty npm-audit stdout explicitly — it crashed the parse as a bare SyntaxError Observed on CI (several branches, all lanes): spawnSync npm ETIMEDOUT with empty stdout; the empty string survived the recovery path and surfaced as 'SyntaxError: Unexpected end of JSON input', hiding the captured error. The recovery path now requires non-empty stdout, and an empty result throws with the captured stdout/stderr/message so the actual error is on the record. Root cause of the ETIMEDOUT itself is NOT diagnosed here — this change only stops masking it. --------- Co-authored-by: sim <sim@local> |
||
|
|
3639ab0431 |
fix(#3968): measure commit claims at all three surfaces — ledger, verifier BLOCKER, porcelain HANDOFF (#4230)
* test(#3968): commit claims must be measured against git, never narrated * fix(#3968): measure commit claims at all three surfaces Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument Emitted-Drift-Ack-Growth: pause-work.md — #3968 uncommitted_files from git status --porcelain * fix(#3968): retired slash syntax, allowlist line pin, git-compare test pin * fix(#3968): persist the ledger on disk and reconcile with the same rev-list instrument Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument * fix(#3968): hold the gsd-executor size cap with a compact ledger contract Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) * fix(#3968): allowlist pin and HALT regex track the final prose * chore(#3968): changeset for measured commit claims * chore(#3968): backfill changeset pr number * fix(#3968): quote the BASE expansion (SC2086) * ci: raise the test-lane budget 21 to 32 minutes (measured cost grew past the cap) --------- Co-authored-by: sim <sim@local> |
||
|
|
2f64e6230a |
feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local> |
||
|
|
647365faf1 |
fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as a gate condition, the end-of-phase escalation must not require MVP, the executor agent's gate section triggers on TDD_MODE alone, and the gate semantics reference loads without MVP_MODE. * fix(#4011): key the TDD runtime gate on TDD_MODE alone The RED-commit gate shipped as #76's MVP slice kept the paired invocation's conjunct, so workflow.tdd_mode=true was silently inert on every non-MVP phase, contradicting references/tdd.md's own contract. Drops the MVP conjunct from the per-task gate and the end-of-phase review escalation; rescopes execute-mvp-tdd.md's load condition, gsd-executor's gate section, and mvp-concepts' intersection claim. MVP remains free to imply TDD; the file is not renamed (stated assumption in the PR body). * test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing Review follow-ups: the detector now only inspects if/[ condition lines so explanatory prose mentioning both flags cannot trip it; remaining 'under/outside MVP+TDD' phrases in execute-phase.md, the gate reference, and docs/INVENTORY.md now describe TDD-mode semantics. Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011) Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011) * chore(#4011): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
8c9265d4e5 |
fix(#3724): warning-only Dimension 3b findings no longer force the revision loop (#3758)
* fix(#3724): stop advisory Dimension 3b findings from forcing the revision loop Dimension 3b (undeclared/temporal coupling, #1954) is spec'd "never a blocker" but tagged severity: warning — the tier plan-phase's revision loop counts as must-fix — and the planner is never taught the rule, so every multi-wave phase touching shared mutable state replans at least once, and intentionally coupled plans re-flag identically every iteration to the stall prompt. Three coordinated changes: - gsd-plan-checker: retag 3b to severity: info, the tier references/revision-loop.md already exempts by design; recognize a coupling_justified frontmatter declaration in the Do-NOT-flag list so deliberate pairs converge. Additions are offset by trimming 3b motivation prose — the checker sits 45 bytes under its LARGE hard cap. - plan-phase step 12: INFO-only accept — an issues block with zero BLOCKER/WARNING entries accepts the plan and surfaces the advisories instead of re-entering the revision loop. Real blockers and warnings still gate unconditionally. - gsd-planner: slim pointer in assign_waves to the new progressive-disclosure reference gsd-core/references/planner-coupling.md (the planner sits 19 chars under its own cap), which carries the shared-mutable-state rule and the coupling_justified escape hatch so first-pass plans avoid the finding when the coupling is unintentional. Documented the coupling_justified field in docs/reference/plan-md.md. Growth acks per #2914; inventory manifest and install-tree fixtures regenerated for the new reference file. Closes #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): pin Dimension 3b at severity: info The severity retag makes the old assertion (severity: warning) stale; lock the advisory tier from both directions — info must be present, warning must not — so a future edit cannot silently re-arm the revision-loop trigger. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * chore(#3724): changeset fragment for PR #3758 Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * docs(#3724): roster planner-coupling.md in docs/INVENTORY.md The new reference was enumerated in the manifest and all 19 install-tree fixtures but missing its row in the Modular Planner Decomposition table — the roster half the manifest-sync test cannot check. (Review Blocker.) Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): cover all four acceptance criteria (review round 1) - plan-checker-coupling: the 3b severity assertion is now a PARITY check deriving the exempt tier from revision-loop.md's flow instead of hardcoding info — editing either side alone reds the suite. New describe pins the other three criteria: plan-phase's INFO-only accept clause (proven failing-first), the BLOCKER + WARNING count staying intact, the coupling_justified Do-NOT-flag exemption + fix_hint, and the planner pointer + planner-coupling.md content. - ack fragment: $comment's plan-phase figure corrected to +79B; the 2775 pin note carried forward into the gsd-planner.md entry, updated for upstream's #3761/#3764 Rule-paragraph anchor (which this diff leaves verbatim). The parallel-dependent-plans re-anchor this commit originally carried was superseded by upstream #3764 during review; this branch no longer touches that file. Refs #3724 Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 2 — align the stance enumeration, complete the template contract MAJOR: <adversarial_stance>'s severity enumeration gains the INFO bullet so it agrees with Dimension 3b's 'ALWAYS INFO' mandate instead of contradicting it. Funded by extracting the inline <examples> block to the new progressive- disclosure reference gsd-core/references/plan-checker-examples.md (@-inlined from the same spot; #1949 precedent), which also restores the 3b motivation clause round 1 traded away (Nit 4) and nets the agent file SMALLER than base (49107 -> 48486) — the extraction the byte pressure was owed. MINOR: gsd-core/templates/phase-prompt.md now carries coupling_justified, and the field's shape becomes one 'plan-id: reason' string per coupled peer so a plan justified against two peers can express it; docs/reference/plan-md.md's Type column names the shape. NIT: the 3409 ack's plan-phase entry no longer calls the #1168 workflow ratchet an 'XL tier'. Acks and derived artifacts updated accordingly (checker entry removed — a shrink needs no ack; INVENTORY roster row + regen:derived for the new file). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * test(#3724): derive the 3b negative severity assertion (review round 2) Every severity token in the 3b span must BE the tier revision-loop.md exempts, replacing the hardcoded severity:warning negative — if the loop's exemption ever moves, the failure names the real conflict instead of blaming the agent file with a mutually-unsatisfiable pair. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): refit the planner coupling pointer under the char cap Upstream #3299 (PR #3390) grew agents/gsd-planner.md to 49146 chars at the base, leaving 5 chars of headroom where the +16-char pointer was measured against 13 more. The pointer prose shortens to 'Non-file coupling:' — 49150 chars, back under the strict 49152-char cap — and the ack figures follow. The @-path the tests pin is unchanged. Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): re-home the plan-phase ack after the #3823 spent-fragment sweep Upstream #3078/#3823 deleted all fully-spent ack fragments, including 3409-unreachable-guard-arms.json, which carried this PR's plan-phase.md +79B append. Per the collision remedy that sweep added: take the deletion and home the still-live entry in this PR's own fragment. Figures re-measured at this merge base (90871 -> 90950 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): absorb the spent #3172 plan-phase fragment into this PR's ack Upstream #3825 shipped 3172-stated-failing-direction.json naming only plan-phase.md, now spent at the base — colliding with this PR's live plan-phase entry. Per the #3003 pattern the fully-spent single-path fragment is deleted and this fragment stays the path's one source; figures re-measured at this base (93073 -> 93152 LF bytes). Claude-Session: https://claude.ai/code/session_01GshUzpGjoxiw6uNRiFMHvM * fix(#3724): review round 3 — true up the ack figures, restore the wave comment The fragment's absolute sizes are re-measured and anchored to base |
||
|
|
6beaa66b25 |
enhance(#3304): gate re-verification blockers on deterministic evidence (#4085)
* test(#3304): add failing-first suite for the convergence evidence gate Content-assertion suite for the Step 7 re-verification evidence gate (agents/gsd-verifier.md / gsd-core/references/verifier-evidence-gate.md). Committed before the implementation to prove RED via gsd-test. * enhance(#3304): gate re-verification blockers on deterministic evidence Step 7's anti-pattern scan re-runs at full, unbounded scope on every re-verification pass, independent of the must-haves established in Step 2. A blocker it finds — other than the self-evidencing debt-marker check — previously reverted a completed gap-closure round and started another --gaps cycle on nothing more than the verifier's own new judgment call, with no bound on how many times that could repeat. A Step 7 blocker now blocks unconditionally in re-verification mode only if it is a carried-forward gap (present in the prior VERIFICATION.md's gaps: list) or the flagged file was git-modified since the prior pass (a regression; fails closed toward blocking when history is unresolvable). Otherwise it predates the gap-closure round unflagged and needs deterministic evidence — a named test run red, or another concrete reproducible artifact — to stay blocking. Unevidenced, it downgrades to a new advisory: frontmatter list and report section instead of setting status: gaps_found, and never reverts a completed must-have. Maintainer approval was narrowed to this evidence condition only, explicitly rejecting the broader "advisory whenever untraceable to a requirement/decision/prior-gap" proposal — implemented and pinned by tests/verifier-evidence-gate.test.cjs and documented as rejected in gsd-core/references/verifier-evidence-gate.md so it can't silently re-expand. Closes #3304 * fix(#3304): correct window-truncation and indentation bugs in evidence-gate tests gsd-test's GREEN checkpoint caught 3 real bugs in the test file itself (not the production prose): a {0,600} match window was shorter than the 724-char paragraph it was scanning (the "exclude from Step 9 Rule 1" phrase starts at offset 662), and two regexes assumed no indentation after a markdown list-continuation line break. All three phrases are confirmed unique across agents/gsd-verifier.md, so the windowed submatches are replaced with direct whole-string assertions instead of just widening the window. Also acknowledges the deliberate byte growth in agents/gsd-verifier.md that the differential-attribution check (ADR-2719) correctly flagged. Emitted-Drift-Ack-Growth: gsd-verifier.md — adds the #3304 re-verification evidence gate (Step 7 rule, Advisory bucket, advisory: frontmatter, report section); 1488 bytes, still within the LARGE-tier 48 KiB cap (48751/49152). * docs(#3304): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
529480b4a5 |
fix(#3895): delete the mempalace-curator's model frontmatter pin — the fleet's only hardcoded model (#4048)
* test(#3895): no shipped agent may hardcode a model frontmatter pin (failing first) * fix(#3895): delete the mempalace-curator's model frontmatter pin — the fleet's only hardcoded model Exactly one of the 34 shipped agents carried 'model: sonnet' in its frontmatter; every other agent resolves through the model-profile system. The ship:post dispatch (#2684) resolves per-hook and — per #2517 — deliberately OMITS model= on inherit so the agent inherits the orchestrator's model; the frontmatter pin intercepted that inherit case, silently forcing sonnet where all 33 siblings would inherit, and operators could not durably remove it (install rewrites live copies wholesale). Deleting the line changes nothing for default profiles — the catalog entry (model-catalog.json agents.gsd-mempalace-curator: golden/balanced sonnet, budget haiku) preserves today's behavior — while restoring model_overrides and inherit authority. Pinned by a new agent-frontmatter guard: no shipped agent may hardcode a model pin, and the catalog entry must keep existing so the pin's deletion can never orphan the agent. * chore(#3895): changeset fragment (pr number backfilled after PR creation) * chore(#3895): backfill changeset PR number (4048) --------- Co-authored-by: sim <sim@local> |
||
|
|
192eb1dfbd |
fix(#3886): git commit timeout reported as commit_timeout; stale lock surfaced; 30s band (#4046)
* test(#3886): a timed-out git commit reports commit_timeout, not commit_failed (failing first) * fix(#3886): git commit timeout reported as commit_timeout; 30s band; stale-lock surfaced cmdCommit's git commit invocation did not distinguish a spawnSync timeout from a real non-zero exit (#2608 fixed this for the staging loop only): a slow pre-commit hook crossing the 10s cap was SIGTERM'd mid-hook and reported as reason commit_failed with whatever partial stderr git had flushed (in the reporter's case an incidental CRLF warning), while the kill left a stale .git/index.lock blocking the next attempt. All three commit sites now check isSpawnTimeout before the nothing-to-commit/ordinary-failure branches: cmdCommit reports reason commit_timeout + timed_out:true and names the stale lock's path (surfaced, not auto-deleted — deleting a lock a live git holds is destructive; the caller recovers deliberately); the subrepo counterparts do the same within their per-repo result / rollback error. The commit calls also move to the 30s band the push call already uses — husky+lint-staged alone idles ~4s on Windows before any task runs. * fix(#3886): review fold-ins — git-path lock resolution, shared band constant, executor contract row, precedence pin - The stale-lock path is resolved via git rev-parse --git-path index.lock, never a literal .git/index.lock join (#3588 row 8's class: a linked worktree's .git is a FILE, so the literal path cannot exist there while the real lock — under <gitdir>/worktrees/<name>/ — blocks the next commit; this repo leans on linked worktrees). - COMMIT_TIMEOUT_MS hoisted; all three sites and their messages build from it (the subrepo variant also regains the stdout fallback the primary site had). - agents/gsd-executor.md's commit-result contract gains the commit_timeout row with the OPPOSITE retry advice from staging_timeout (remove the stale lock, then retry once) — an executor matching the doc previously had no handling for the new reason. - Precedence pin: a timeout whose partial output contains 'nothing to commit' must still read as a timeout (branch-reorder mutant). Emitted-Drift-Ack-Growth: gsd-executor.md — #3886: +commit_timeout row to the commit-result contract with the retry guidance OPPOSITE staging_timeout's (remove the stale lock, then retry once); the executor previously had no handling for the new reason. * chore(#3886): changeset fragment (pr number backfilled after PR creation) * chore(#3886): backfill changeset PR number (4046) --------- Co-authored-by: sim <sim@local> |
||
|
|
ac3668e4b7 |
fix(#3797): make the roadmapper's role, output, and checklist match its write-first execution flow (#4008)
* test(#3797): the roadmapper must follow one write-first contract * fix(#3797): make the roadmapper's role, output, and checklist match its write-first execution flow The roadmapper contradicted itself: role blurb, output format, and completion checklist described an approve-first flow while its execution flow said "Write Files Immediately" with reactive-only revision (#3797). The approval gate belongs to the ORCHESTRATOR — both callers read the written ROADMAP.md, present it, and gate on approval (with an auto-mode bypass a subagent cannot host) — so write-first is the contract. All approve-first text now describes the write-then-return reality, the old "Draft Presentation Format" (whose ## ROADMAP DRAFT header matched no orchestrator branch) is folded into the ## ROADMAP CREATED structured return as a preview block, and the duplicate checklist lines are merged. A structural guard pins the single contract. Emitted-Drift-Ack-Growth: gsd-roadmapper.md — #3797: +bytes — approve-first wording replaced with write-first descriptions; the DRAFT presentation template folded into the ROADMAP CREATED return as a preview block * chore(#3797): changeset fragment (pr number backfilled after PR creation) * chore(#3797): backfill changeset PR number (4008) --------- Co-authored-by: sim <sim@local> |
||
|
|
52b11ee811 |
fix(#3763): pass --raw at every shipped config-get bash call site (#3961)
* test(#3763): guard every shipped config-get substitution on --raw * fix(#3763): pass --raw at every shipped config-get bash call site config-get without --raw prints JSON.stringify(value), so string-typed values reach bash with literal quotes and every string comparison silently never matches (#3763). --raw added at 75 command-substitution sites across shipped content; four JSON consumers (default_reviewers, sub_repos, pr_body_sections, code_review_depth_overrides) deliberately keep default JSON output. Emitted-Drift-Ack-Growth: ai-integration-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: audit-fix.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: autonomous.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: cleanup.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: code-review.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: complete-milestone.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: discuss-phase-assumptions.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: do.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: eval-review.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: execute-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: execute-plan.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: fast.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: graduation.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: gsd-executor.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: health.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: import.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: inbox.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: ingest-docs.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: mvp-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: new-milestone.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: next.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: plan-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: plan-review-convergence.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: plant-seed.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: profile-user.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: progress.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: quick.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: remove-workspace.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: secure-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: settings-integrations.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: settings.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: ship.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: sketch-wrap-up.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: sketch.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: smart-entry.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: spike-wrap-up.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: spike.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: ui-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: ui-review.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: undo.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted Emitted-Drift-Ack-Growth: validate-phase.md — #3763: bytes from '--raw' at config-get call sites so string-typed config values reach bash comparisons unquoted * chore(#3763): changeset fragment (pr number backfilled after PR creation) * chore(#3763): backfill changeset PR number (3961) --------- Co-authored-by: sim <sim@local> |
||
|
|
fb2d122d7f |
feat(#3841): assert gsd-tools identity on every state-mutating verb (#3848)
* feat(#3841): assert gsd-tools identity before any state-mutating verb only this package publishes. The path-based branches — a project-local install, a runtime config directory — had no such guarantee; they trusted their configured location. This closes them. Mechanism: once resolution finishes, and before any verb runs, the preamble probes the tool it picked with `runtime-identity --raw` and matches the answer with a shell `case` pattern ANCHORED to the start of the compact payload (`{"packageName":"@opengsd/gsd-core"`). An unanchored substring match accepts the decoy `{"packageName":"get-shit-done-cc","note":"@opengsd/gsd-core"}`, which any colliding package could publish. The outcome is exported as the two-valued `GSD_IDENTITY_STATUS` (`ok`/`unverified`), so the gate is asserted on a VALUE rather than on warning prose. Rollout is warn-then-fail per the #3146 ruling: `unverified` prints one line naming BOTH causes and continues, because `no_identity_verb` cannot tell a foreign package from an `@opengsd/gsd-core` older than the verb, and at rollout the old-version case is the common one. The blocker was byte budget, not design. The preamble is inlined into 112 shipped files and several sat within single-digit bytes of frozen ceilings (`gsd-verifier.md` 16 bytes, `gsd-executor.md` 33, `execute-phase.md` 234); a first attempt broke five of them. What made room was collapsing the resolver's twenty near-identical `elif [ -f … ]` arms into one candidate-list helper (`_gsd_at`), which buys far more than the assertion costs. The preamble is now 2,624 bytes against 4,500 — a net 1,876 bytes SMALLER per inlined file, so every capped file moved away from its ceiling rather than toward it. No cap raised, no size-budget exception added, no override token emitted. Resolution order, every runtime-home probe, the `unset -f gsd_run` re-source fix, the fail-closed `exit 1`, and the `CLAUDE_ENV_FILE` persistence are all preserved byte-for-byte in substring terms; the snippet still begins with `_GSD_SHIM_NAME=` and still ends with `fi`, which the parity extractors anchor on. `gsd-core/references/gsd-run-resolver.md` is re-synced byte-equal. Also fixes two stale claims found in passing: CONTEXT.md and FEATURES.md both described an `[ -x ]` guard as the load-bearing re-source defense. That guard was tried and REMOVED in #3831 — it rejected the bare function name, fell through every branch, and hit `exit 1`, which kills a sourced caller's shell. `unset -f gsd_run` is the actual mechanism. Refs #3841 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3841): pair the anchor's brace by requiring a closed identity payload The matrix went red on `tests/new-project-mvp-prompt.test.cjs` — "new-project.md has unbalanced braces: net depth 2" — plus a knock-on report from its parent `bug #1516` describe, which is the same failure counted once at the child and once at the block. Root cause: that guard (:182-189, mirroring #3784 bd53925f) walks characters and increments on `{`, decrements on `}`, with no awareness of shell quoting. It scans `new-project.md` PLUS every `new-project/steps/*.md`, and both `new-project.md` and `steps/auto-mode-config.md` carry one inlined preamble copy — hence net 2 from a snippet that was off by exactly one. The unpaired brace was the `{` inside the single-quoted `case` pattern of the identity anchor, which is correct shell and invisible to a text scanner. Fix in the snippet, not the guard. The pattern now anchors at BOTH ends: `'{"packageName":"@opengsd/gsd-core"'*'}'`. That balances 51/51 with a brace that does real work rather than a cosmetic pair — a truncated payload whose prefix matches now fails too, where before it verified. Safe for any future additive field: a JSON object's own closing brace is always the last character, whatever type the last value has, which is pinned by two negative-space tests (a nested object and an array-valued last key must both still verify). Cost: +3 bytes, against the 1,873 the resolver fold already gave back. The alternative considered and rejected was dropping the literal `{` for a `?` glob. It balances too, but weakens the anchor from "must be an opening brace" to "must be any one character", and the anchor is the entire point. Two guards added so this cannot recur silently: - runtime-launcher-parity (F0) pins brace balance at the SNIPPET, so the next edit to that pattern fails on the file it broke instead of surfacing three files downstream in a test whose name mentions neither the launcher nor this issue. It also asserts depth never goes negative, since a `}` preceding its `{` nets to zero while being unbalanced at every prefix. - runtime-identity gains behavioral truncated-payload and trailing-garbage fixtures, so the added `}` is proven load-bearing rather than merely present. Verified: snippet 51/51 braces; new-project combined net depth 0; the seven other preamble-bearing files with nonzero depth are unchanged from merged next (their own prose, not the preamble, and not in any guard's scan set); all 112 inlined copies and the resolver reference re-synced byte-equal; sync:launcher idempotent on the second run. Refs #3841 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3841): backfill changeset PR number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
63abcface9 |
feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools (#3831)
* feat(#3146): resolve gsd_run so workflows cannot reach a foreign gsd-tools The predecessor package get-shit-done-cc publishes a colliding gsd-tools bin whose phases.clear DELETES where this package's ARCHIVES, and both print success-shaped output against a gitignored .planning/ -- which is how #3129 cost a user 43 phase directories with no error and nothing recoverable from git. The launcher's PATH branch now resolves gsd_run, published only by this package and self-locating via its own symlink chain to the sibling shim, instead of the colliding gsd-tools. A foreign handler becomes unreachable from PATH, and when no gsd_run is reachable the resolver fails closed rather than falling back -- that fallback was the vulnerability. This is smaller than the branch it replaces, which matters: the preamble is inlined into 113 shipped files and agents/gsd-verifier.md sits 2 bytes under a red-line size cap. unset -f gsd_run leads the preamble so a re-source is idempotent. Without it, command -v finds the shell function, returns a bare name, and the resolver falls through to an exit 1 that kills a sourced caller's shell. Adds gsd-tools runtime-identity, a manual diagnostic reporting this runtime's package coordinates over the baked package-identity (#498) and readHostVersion, with a strict total classifier: only a JSON object with an exact packageName verifies, since JSON.parse admits 0/"str"/[]/null/true. An inlined identity assertion was built and reviewed first, then withdrawn -- it breaks five frozen size ceilings and no assertion fits in 2 bytes. Closes #3146 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3146): stop sync:launcher relocating a deliberate preamble placement Pre-existing defect, surfaced by this PR because sync is a no-op unless the snippet content actually changes. transformFile inserts the preamble into the first block that CALLS gsd_run, but gsd-core/workflows/explore.md deliberately places it in a bootstrap-only block that DEFINES gsd_run without calling it -- its own comment explains why: declining the research offer must not leave Step 5's commit call unbootstrapped. Stripping empties that block of calls, so the preamble migrated forward and broke the define-before-use invariant tests/explore-command.test.cjs pins. Reproduced on a pristine origin/next checkout with the base snippet and base file, so this was not introduced here. The insertion target now honours a block that already carried the preamble, falling back to the first calling block for files that have none yet. Adds a behavioral regression test over a two-block fixture. Also updates three runtime-launcher-parity tests that pinned the removed PATH fallback to gsd-tools. Their intent is preserved -- the PATH stub is renamed gsd_run so it is reachable by the new resolver, and the RUNTIME_DIR-wins test still asserts the stub is never invoked. Fixture shebangs move to an absolute /bin/sh, because the fixture PATH is deliberately restricted and #!/usr/bin/env sh could not resolve. Refs #3146 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3146): backfill changeset PR number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3146): document the FEATURES.md section-numbering practice The monotonically increasing section number in docs/FEATURES.md is the most frequent merge-conflict source in this repo, and it has TWO conflict cells, not one: the ### N. heading and the hand-maintained table of contents. Two PRs adding differently numbered features still collide on the TOC, so renumbering alone does not make a branch safe. This branch alone was renumbered 165 -> 166 -> 167 -> 168 across successive rebases. Adds a CONTRIBUTING section stating the practice: allocate the number last, never pre-emptively renumber, take max+1 after a rebase and update the TOC in the same commit, and never renumber someone else's section. Fork contributors are told explicitly they may leave the number to a maintainer at merge rather than chasing the counter. Agents are told to lease the allocation and to include the file in their published touched set. Records the durable fix as planned rather than pretending it exists: FEATURES.md should be generated from per-feature fragments the way CHANGELOG.md is generated from .changeset/, and the way tests/emitted-drift-acks/ works (#2914). Also renumbers this branch's own section to 168, leaving 167 to the PR already in flight. Refs #3146 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c933184b97 |
enhance(#3172): require a stated failing direction for every automated acceptance command (#3825)
* test(#3172): failing-first suite for the stated failing-direction probe Pins the <fails_when> pairing walk, placeholder denylist, MISSING sentinel exemption, degraded-read contract, CLI arm and the plan-authoring contract text. RED by construction: the module exports it requires do not exist yet. Executed on the remote runner. * feat(#3172): require a stated failing direction for every automated acceptance command Every runnable <automated> command now carries a <fails_when> sibling naming what output constitutes failure. A command with no expressible failure mode is not an acceptance test: it reads as rigour and is not falsifiable. - verify-command-grounding gains a failing-direction probe sharing the existing <automated> grammar, MISSING sentinel and walk guard rather than copying them - gsd-tools check verify-failure-directions <N> backs it; plan-phase dispatches it and hands the JSON to gsd-plan-checker check 8f - Dimension 8 detail extracted to references to stay under the agent size cap Verified on the remote runner. * fix(#3172): close four review findings in the failing-direction probe - MISSING_SENTINEL_RE matched an env-var assignment prefix (MISSING=1 cmd), so a real command was exempted from the new blocking gate. Tightened the SHARED constant rather than adding a second copy. - Both token regexes scanned to EOF on unclosed openers (O(n^2), 1562ms at 40k). Bodies are now non-crossing; 1ms, byte-identical on well-formed input. The pre-existing AUTOMATED_BLOCK_RE carried the same defect and is fixed here too. - probePhaseFailingDirections reported status 'ok' when one plan was unreadable, conflating 'could not look' with 'nothing to report'. - Extracted the phase-resolution block both check arms had copied verbatim. Also corrects a docs/AGENTS.md dimension list stale since #2401. Verified on the remote runner. * fix(#3172): project the planner rule onto the spawn contract, settle emitted bookkeeping The remote runner refuted the planner-side edit. agents/gsd-planner.md is frozen under a 49152-LF-char cap asserted by four suites and sat at 49,146 — six chars of headroom — so the +537 of authoring rule blew it. #3297/#3645 already settled where such a rule goes: the planner spawn contract in plan-phase.md, beside <tracked_source_paths>. The agent file is reverted to origin/next verbatim. - plan-phase.md gains <failing_direction_contract>; tests row 30 now asserts the contract there and row 30b guards the freeze in both directions - plan-phase.md growth acknowledged by APPENDING to the 3409 fragment, per the precedent that two ack sources may never name the same path - install-tree fixtures regenerated for the three new reference files Verified on the remote runner. * chore(#3172): backfill PR number into the changeset fragment pr:0 -> pr:3825 now that the PR exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
8442d984b9 |
fix(#3809): route runtime-loaded markdown through the gsd_run launcher (#3815)
* test(#3809): generalize dead-ref guard into a rule table (failing first)
The #2020 guard hardcoded `sdk/(src|dist|handlers)/` — the three dead paths
that had caused that storm. That proved those three paths were gone and said
nothing about the class, so #3809 reproduced the identical Windows find.exe
storm under a different token and the guard could not see it.
Replaces the single regex with a rule table over the same runtime-loaded
markdown surface, adds `commands/` to the scan set (previously uncovered),
and adds rule B: the runtime shim filename must never appear in command
position, because it is not a PATH command and an agent that meets it falls
back to locating the file.
Rule B's matcher is deliberately lenient — the launcher's own resolver
assignment, `node <path>/<shim>` calls, bare paths, and prose that names the
file all stay unflagged, each pinned by a negative-space row.
This commit is expected to FAIL: 50 offenders across 23 files remain in the
tree. The remediation lands next.
Refs #3809
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(#3809): route every workflow call through the gsd_run launcher
50 places across 23 runtime-loaded workflow, agent, reference, and command
files instructed the agent to run the runtime shim by filename. That filename
is not on PATH under any name -- package.json ships gsd-core, gsd-tools,
gsd_run and gsd-mcp-server -- so the call exited 127, the file-shaped token
sent the agent looking for the file, and on Git Bash for Windows the resulting
`find /` walked the entire drive (7268 CPU-seconds in the report) until
somebody killed it by hand.
CONTEXT.md -> Runtime Launcher Module already makes gsd_run the single entry
point: "Canonical space-safe shell preamble (`gsd_run`) used by every workflow
bash block to invoke the GSD runtime CLI." These sites predate that rule --
they trace to
|
||
|
|
cf15682d1c |
enhance(#3028): responsive Markdown separators instead of fixed-width rules (#3789)
* feat(#3028): responsive Markdown separators instead of fixed-width rules Stage banners, checkpoints, completion and error panels used fixed-width runs of box-drawing characters -- a 53-column heavy rule and a 62-column double-line box. Those runs are ordinary text to a Markdown-rendering host, so in a narrower pane they wrap and the border comes apart from the heading it framed. Shipped content now emits an ATX heading for a titled section and a blank-line-delimited --- for a break between sections, both of which adapt to the available width. The same convention is applied to the three code sites that built these strings at runtime: the UAT checkpoint renderer, the milestone-close audit report, and the TDD review checkpoint table. Removing the box also removes its only reason to exist -- the east-asian-width padding helpers that kept its right border aligned (checkpointBoxLine, displayWidth, isWideCodePoint, ZERO_WIDTH_MARK_RE, CHECKPOINT_BOX_WIDTH). RTL directional isolation is unchanged. The convention is specified in gsd-core/references/ui-brand.md and enforced across all shipped content by tests/responsive-separators.test.cjs. Refs #3028 * test(#3028): pin the heading form in checkpoint and audit-report assertions These suites asserted the exact box borders and the 62-column padded banner interior. With the box gone they assert the ### heading form, the --- break and the bolded instruction line, and each now carries a positive assertion that no box character remains -- which is what pins the fix rather than merely tolerating it. Language coverage is converted, not dropped: Japanese, Chinese, Korean, Hindi and Arabic all still assert their rendered banner, and the Arabic case still asserts the RTL directional isolates the box removal must not disturb. Adds a case for a banner longer than the old inner width, which previously produced a ragged border and now has none. Refs #3028 * chore(#3028): acknowledge execute-plan.md growth from the checkpoint display spec The checkpoint_protocol display spec described the drawn box; it now describes the heading, the --- break and the bolded action prompt, which costs 22 bytes (40111 -> 40133, 827 under the cap). Appended to the existing #3370 fragment rather than filed as a new one: a growth ack keys on the bare filename and #3370 already declares execute-plan.md, so a second source naming it would be a hard duplicate-key error. Same supersede-by-append route #3370 took for the spent #2652 fragment. Refs #3028 * docs(#3028): state the load-bearing half of the separator rule, and amend the zh-CN reference Review found three things. The rule as first written demanded a blank line above AND below every ---. Only the one above is load-bearing: it is what stops CommonMark reading the rule as a setext underline for the line above. The one below is cosmetic, because a thematic break is a leaf block. The rule now says that, with the reason, instead of asserting a stricter form the content does not keep. The zh-CN reference had received the mechanical box-to-heading swap but none of the prose behind it: it still claimed a 62-character checkpoint width and still listed --- among forbidden mixed banner styles, so it contradicted the convention it was translating. It now carries the separator section, the setext reasoning, the unconditional-vs-per-runtime rationale and a corrected anti-pattern list, in Chinese. The user guide asserted that a heading is not a degradation anywhere. That is an assertion, not a demonstration. It now says what was actually traded away in a plain terminal, points at the recorded rationale, and invites the report that would justify the capability flag instead. Refs #3028 * chore(#3028): backfill changeset PR number Refs #3028 --------- Co-authored-by: sim <sim@local> |
||
|
|
622f43353c |
fix(#3299): tracer feedback gate honors workflow.human_verify_mode (#3390)
* fix(#3299): tracer feedback gate honors workflow.human_verify_mode
The tracer feedback gate (#2294) predates `workflow.human_verify_mode`
(#3309, whose scope was the planner and verifier only), and branched on
auto-mode alone. Under the documented `end-of-phase` default an
interactive run therefore halted after EVERY `type="tracer"` task,
synthesizing a `checkpoint:human-verify` no planner ever emitted and
asking the user to retype a verdict the executor had just computed —
at the cost of a full executor cold-start each time.
Planner-side suppression cannot reach this halt because the executor
synthesizes it at runtime, which is why #3309 did not close it.
The gate now branches on HUMAN_VERIFY_MODE in the interactive path:
under `end-of-phase` an automated-only tracer `<verify>` is re-run and,
on success, expansion continues with no checkpoint. HALT-on-failure is
unchanged. `mid-flight`, `gate="blocking-human"`, and tracers carrying
genuine `<human-check>` evidence all still stop; the autonomous branch
is untouched.
`--default end-of-phase` on the config read is load-bearing, not
decorative: `workflow.human_verify_mode` is absent from SCHEMA_DEFAULTS,
so a bare `config-get` exits non-zero with `Key not found` on any
project whose config.json predates #3309 — which is the reporter's
exact config and every pre-existing project.
Both copies of the rule (workflows/execute-plan.md and
agents/gsd-executor.md) are updated together; the reference doc records
the seam and the human-check-still-halts rationale so it cannot recur.
Fixes #3299
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(#3299): add changeset
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): reconcile the canonical schema table and the stale acceptance test
Review round 1 (trek-e) — three items, all in the drift class this PR is
about, two of them landed inside this PR's own diff.
1. docs/reference/plan-md.md:233 — CONTEXT.md names this file the canonical
schema reference for the tracer task-type contract, and its Task-types row
still claimed interactive runs unconditionally present a
checkpoint:human-verify. CONTEXT.md and docs/AGENTS.md were updated in the
first round; this one was missed, so the authoritative reference was the
wrong answer. The row now carries the human_verify_mode-conditional
behavior and points at the canonical precedence chain.
2. tests/tracer-bullet.test.cjs — the docs assertion only checked that a
tracer ROW EXISTS, never its content, which is why CI could not see the
drift. It now asserts the row's actual claims and rejects the pre-#3299
wording. Separately, the #1945 acceptance test named 'interactive run emits
checkpoint:human-verify after the tracer' kept passing only because its
substrings still occur in the fallback clause, while its name asserted the
opposite of shipped behavior. Renamed and narrowed to what #1945 still
guarantees, plus a new interactiveIsConditional pin so the unconditional
prose cannot be restored under a passing substring check.
3. plan-md.md's <verify> row now documents that the legacy bare-text form
(valid, and still shown at :179) does not reach the #3299 auto-continue —
only a <verify> carrying <automated> does — so the benefit is silently
unreachable for tracers using that format.
Mutation-verified: reverting the plan-md row fails 1 test; reverting the
executor's interactive branch fails 4.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(#3299): make the tracer gate reachable from the planner template, and bind the assertions
Peer review round 3 found two Majors, both verified by reproducing the
mutation before fixing.
MAJOR 1 — the fix was largely inert on its own default path.
agents/gsd-planner.md's Nyquist Rule (:191) says every <verify> includes
<automated>, but the tracer-specific template twelve lines later emitted the
legacy bare-text form. The gate auto-continues only on a <verify> carrying
only <automated>, so every tracer produced from the canonical template fell
to the STOP fallback and #3299's benefit was unreachable for exactly the task
type it targets. Template now wraps in <automated>; a contract assertion pins
it so the two cannot drift apart again.
MAJOR 2 — the new assertions did not bind condition to action.
Appending 'Nevertheless, interactive runs always present a
checkpoint:human-verify' to the canonical row, and 'then immediately STOP and
return a checkpoint:human-verify' to the auto-continue clause in BOTH
operative copies, restored unconditional interactive checkpointing and left
the suite 35/35 green. Every required keyword still matched. Fixed by:
- clause 2 must now contain no STOP outcome and emit no checkpoint at all —
'never a checkpoint' has to be true OF the clause, not merely stated in it;
- interactiveIsConditional replaced with the ordered-clause parse plus the
same no-STOP property, instead of proving only that HUMAN_VERIFY_MODE
appears somewhere on the line;
- the plan-md.md Autonomy cell is now pinned EXACTLY rather than by keyword
presence. Deliberately brittle: CONTEXT.md names that table the canonical
schema reference, so a wording change must be a conscious edit in both
places.
Mutation-verified after the fix: the combined semantic regression now fails 3
tests; reverting the planner template fails 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): exact-pin the safety clauses instead of blacklisting outcome verbs
Peer review round 4. Blacklisting did not hold, twice over:
- Round 3 banned literal STOP and the 'return a'/'present a' checkpoint
forms in the auto-continue clause. Round 4 defeated that by appending
'then pause and invoke checkpoint_protocol with a checkpoint:human-verify
before expansion' — none of the banned tokens, same restored interruption
after every successful tracer. 36/36 passed.
- The planner guard looked for <automated> anywhere inside <verify>, so
'<verify>[...]<!--<automated>--></verify>' satisfied it while leaving the
legacy bare form operative. 107/107 passed across tracer, planner and the
three size-cap suites.
Synonyms are unbounded; the clauses are not. Both are now pinned exactly on
normalized whitespace, the same approach already proven on the plan-md.md
Autonomy cell, with defence-in-depth checks behind them: no checkpoint-emitting
or blocking outcome in any wording inside clause 2, and the planner's <verify>
body must be exactly one non-empty <automated> child with no commented markup.
These pins are deliberately brittle. Each is a safety contract, so changing the
behavior must be a conscious edit in both the prose and the expectation.
Mutation-verified: the synonym-checkpoint mutation fails 1; the commented-out
wrapper fails 1; the round-3 literal-STOP + contradictory-doc-row regression
fails 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): strip comments, require uniqueness, pin whole regions
Peer review round 5. Exact-pinning one clause was still bypassable two ways,
both reproduced before fixing (each left the suite fully green):
- COMMENTED DECOYS. Put the correct text in an HTML comment followed by a live
wrong copy: every extractor selected the commented decoy. Worked against the
planner template, the canonical plan-md.md row, and both executor branches.
- SURROUNDING OVERRIDE. Insert 'after every tracer, pause and invoke
checkpoint_protocol before expansion, regardless of the mode-specific rules
below' immediately ABOVE the pinned clause, or 'ignore row 3; always wait for
approval' below the canonical table. The pinned text was untouched, so
equality held while the shipped meaning inverted.
The shape that holds, applied to every operative surface:
1. strip HTML comments BEFORE selecting, so a decoy cannot be chosen;
2. require the structural anchor to occur EXACTLY ONCE, so a live second copy
cannot hide behind a correct first one;
3. pin the ENTIRE decision region, not one clause, so no unparsed prefix or
suffix can override what the pin proves.
Applied to: the executor's whole tracer branch, execute-plan.md's whole
dispatch line, checkpoints.md's whole precedence section, and plan-md.md's
Autonomy cell.
Also addresses the round-5 Minor: the planner template is now asserted
STRUCTURALLY (exactly one <verify> in the fenced block, body exactly one
non-empty <automated> child) rather than pinning the descriptive placeholder
verbatim, so behavior-preserving wording changes no longer false-fail. The
clause and section pins keep their exact form — those have a safety rationale
the placeholder copy does not.
Mutation-verified, all six rounds: override-above-clause 1; commented decoy row
1; commented decoy branch 1; ignore-row-3 override 1; synonym checkpoint 1;
commented-out wrapper 2.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): drop the superseded exact-placeholder planner assertion
Peer review round 6, Minor. The round-5 brittleness fix ADDED a structural
planner assertion but left the old exact-placeholder one in place, so the
over-brittleness it was meant to remove was still live: rewording the
descriptive placeholder while preserving exactly one non-empty direct
<automated> child failed the old test and passed the new one.
Removed the old test. The structural assertion is the real contract — the gate
auto-continues on the SHAPE of the verify, not on the wording of a placeholder.
Verified both directions: a behavior-preserving reword now passes; reverting the
template to bare <verify> still fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): select operative prose via parsePredicates, not a hand-rolled scanner
Peer review round 7. I had judged the round-6 selector bypass adversarial-only
and out of scope, intending to disclose it. Both premises were wrong, and the
review said so:
- 'Needs new src API' — false. parsePredicates is ALREADY a public export and
internally uses the repo's interleaved fence/comment scanner. Instrumenting
candidate lines as throwaway predicate declarations borrows that scanner with
no src change at all.
- 'Adversarial-only' — false, and this is the part that mattered. Two ORDINARY
edits silently turned the guards into decoy checks:
* a forgotten '-->' comments the live rule through to EOF, and the
balanced-only stripper still saw and accepted the commented rule;
* a normal fenced documentation example of the rule, plus a whitespace-only
reformat of the live list item, made the selector choose the example.
Neither needs intent. A dangling comment is a typo; a fenced example is good
documentation. Together they reproduce exactly the accidental drift #3299 came
from — with CI green.
The selection layer now defers to parsePredicates for operativeness, uses
whitespace-tolerant anchors so a reformat cannot decouple the live line from its
pin, extracts regions by operative line index rather than string search, and
carries a self-guard test proving fenced / balanced-commented /
after-unclosed-comment copies are all excluded. The helper also ignores indexes
it did not inject, so a pre-existing GSDTEST.CANDIDATE line cannot pollute it.
Verified both ordinary-edit scenarios now fail the suite (each was green before).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): close the operative-selection gaps the maintainer blocked on
trek-e's Blocker: the operative-line selection layer had three gaps, all
reachable by ordinary future doc edits rather than sabotage. He independently
found a fourth I had not disclosed. All are fixed.
1. INDENTATION PROMOTION (his find, not in my disclosure). The instrumentation
replaced a matched candidate with an UNINDENTED marker regardless of the
original line's indentation. A 4-space-indented CommonMark code block is not
skipped by parsePredicates (it accepts indented declarations by design), so
stripping the indent PROMOTED an indented decoy to operative — the exact
inversion of the guard's purpose. The marker now preserves the original
indent, and a candidate that is itself indented 4+ spaces is never injected.
2. NO SET MEMBERSHIP. The filter accepted any in-range integer, so a
pre-existing literal GSDTEST.CANDIDATE=<valid index> in source text could
pollute the count. Now filters on a Set of the indexes actually injected on
this call.
3. RAW FENCE SELECTION (planner). The template test matched the first raw
```xml fence after the marker with no fence/comment awareness — the one
selection in the suite that was not operative-aware — so a commented-out
decoy template between the marker and the real one would be selected while
the live template regressed. The opener must now be operative AND the first
non-blank line after the marker.
4. RAW END ANCHOR (regionFrom). The end anchor was tested against raw lines, so
a fenced example containing a ### / <type line truncated the pinned region
early — a false FAILURE on a legitimate doc edit. End anchors now go through
the same operative filter as start anchors.
Mutation-verified: the indented-decoy + whitespace-varied-anchor combination
and the commented-out fence decoy each now fail the suite (both passed clean
before). Truncation is confirmed fixed by extraction — the region spans the
full section and retains the content following a fenced example, where it
previously stopped at it.
Note on the remaining brittleness: adding a fenced example INSIDE a pinned
region still fails the whole-region exact pin. That is the intended tradeoff
for a safety contract, not the truncation defect, and is called out as such.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(#3299): allow-list operative indentation; pin marker provenance
Review round 9.
BLOCKER — the round-8 indentation guard was written as a DENY-list,
/^(?: {4,}|\t)/, and CommonMark has more indented-code forms than that
enumerates: " \t", " \t" and " \t" all open an indented code block and all
slipped through, so an indented decoy was still promoted to operative while the
live rule regressed (34/34 green). Inverted to an allow-list — only 0-3 literal
spaces is ordinary block indentation; anything else is code. Enumerating the
bad shapes was the error, not the specific regex.
MINOR — the injected-index Set validated the marker's VALUE but not its SOURCE.
A pre-existing literal `GSDTEST.CANDIDATE=<n>` could name an index that some
other (skipped) candidate had contributed to the set, and be accepted. Now also
requires p.line - 1 === Number(p.value): the predicate must have been parsed
from the line it names.
MINOR (false negative) — ```xml title=x is a valid CommonMark info string, and
requiring exactly ```xml failed the suite (33/34) on a behavior-preserving edit.
Both the opener assertion and the extraction now accept an info string.
Mutation-verified: the mixed " \t" decoy and the forged-provenance marker each
now fail; the info-string fence no longer false-fails.
KNOWN LIMITATION, disclosed on the PR rather than papered over: parsePredicates
is a predicate parser, not a general CommonMark operativeness oracle. Two
standards-valid constructs still read as operative — a lazy blockquote
continuation line (state opens only on a line that literally starts with ">"),
and a comment opened mid-line ("prose <!--", where state opens only when the
trimmed line STARTS with "<!--"). Closing those means either teaching the shared
src/context-predicates.cts about container/lazy-continuation state — a change to
a module every health rule consumes, well outside a tracer-gate fix — or
hand-rolling a CommonMark parser inside a test, which is how this suite got into
trouble in the first place. Left for the maintainer to scope.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore(#3299): re-arm the execute-plan.md emitted-drift ack after the base merge
The #3299 ack rode on tests/emitted-drift-acks/2652-quick-diagnose-dispatch-isolation.json,
which upstream retired in
|
||
|
|
2f86278b5e |
fix(#3003): opt-in mechanism for intentional deletions in worktree.cleanup-wave (#3757)
* test(#3003): failing-first suite for declared deletions in cleanup-wave Binds the guard's opt-in before it exists, so the suite is RED against next. The rows that carry the weight are the over-authorization set: a directory declaration must not authorize its children, a glob declaration must authorize nothing, and a declaration must not act as a string prefix of another path. Each of those BLOCKS, and each would PASS under a prefix, glob, or startsWith matcher — which is how a path list quietly degrades into the boolean opt-in #3003 explicitly rejected. The glob row matters most: declaredScopePrefix already returns null ("matches everything") for a glob-leading pattern, correct for the advisory it serves and catastrophic for a gate. Also pinned: a failed deletion check blocks on its own reason rather than being filtered into a pass; the block detail names only the undeclared residue so the operator is not misdirected by paths that were fine; an entry with no declaration blocks exactly as before; junk and non-array declarations do not authorize; and a blocked entry still isolates rather than aborting the wave (#2852, which must stay fixed). Two advisory rows cover an interaction found while designing: git diff --name-only includes deleted paths, so without unioning the declaration into the #2596 scope check, authorizing a deletion would raise SCOPE_OUT_OF_DECLARED against the very path just authorized. A seeded property states the whole invariant the three over-authorization rows sample: a deletion merges iff its normalized path is in the declared set. * feat(#3003): declared deletions opt-in for the cleanup-wave guard The deletions guard blocked the merge-back of any executor branch whose diff removed a file, with no way to say a removal was intended. A plan that folded one test file into a sibling could not be merged by the tool meant to merge it, forcing a manual --no-ff outside the tool -- strictly less safe than what the guard protects against. A plan now declares removals in its own frontmatter (files_deleted), and that list rides the same path files_modified already travels: plan-document parse -> phase plan JSON -> the per-plan worktree gate -> record-agent/create --deletions -> declared_deletions on the manifest entry -> the guard. The guard blocks only the deletions NOT in that list. A path list rather than a boolean, per the pinned decision: a boolean disarms the guard for the whole entry, so an unexpected deletion riding along with a declared one would pass unnoticed. Matching is exact after the module's shared normalizer -- never a prefix, never a glob. Both would let one declaration authorize a whole set, which is the mass-deletion accident the guard exists to catch. That also means declaredScopePrefix is deliberately NOT reused here: it returns null ("matches everything") for a glob-leading pattern, which is right for the advisory it serves and would silently disarm a gate. The block detail now carries only the undeclared residue, so an operator is not sent looking at paths that were fine. A failed deletion check still blocks on its own reason and is never filtered into a pass. A blocked entry still isolates rather than aborting the wave (#2852). The #2596 scope advisory unions the declaration into its declared set -- git diff --name-only includes deleted paths, so without that, authorizing a deletion would immediately warn that the same path was out of declared scope. Optional and additive throughout: files_deleted is absent from PLAN_REQUIRED_FIELDS, a manifest entry without declared_deletions keeps the original unconditional block, and omitting --deletions leaves the on-disk entry shape untouched. Supersedes the spent #2856 emitted-drift ack entry for execute-phase.md, the same supersede that entry performed on #3370 and #3370 on #3324. * fix(#3003): wire --deletions on every dispatch surface, not just one Review found the feature inert on two of three dispatch paths. execute-phase.md (harness inline) passed --deletions, but the orchestrator-worktree path (executor-isolation-dispatch.md, worktree.create) and the Fleet-parallel batch path (capabilities/claude-orchestration/fragments/execute-wave-pre.md, worktree.record-agent) still passed only --files. A plan declaring files_deleted would have merged on one path and been blocked on the other two -- the exact bug #3003 exists to fix, left unfixed where most of the isolation actually runs. Worse, per-plan-worktree-gate.md already claimed --deletions was passed 'on the same worktree.record-agent / worktree.create calls', which was false for both untouched sites. A doc asserting coverage that does not exist is how a gap survives review. All four surfaces now pass the flag, verified by sweeping every .md under gsd-core/, capabilities/, commands/, skills/ and agents/ that invokes worktree.record-agent or worktree.create: each one that passes --files now also passes --deletions. The isolation-dispatch note explains why this flag, unlike --files, is not advisory -- omitting it does not skip a check, it blocks a merge the plan declared. Regenerates capability-registry.cjs, which the fragment edit made stale. Neither newly-grown file needs an emitted-drift ack: executor-isolation-dispatch.md sits under workflows/execute-phase/steps/ and execute-wave-pre.md under capabilities/, both outside currentSizes()'s non-recursive scan of gsd-core/workflows/ and agents/. * docs(#3003): document files_deleted where a plan author will actually find it The feature's entire user surface is one plan-frontmatter field, and the canonical reference for that frontmatter -- docs/reference/plan-md.md, the table that documents every other key -- never mentioned it. A field nobody can discover ships as a field nobody uses. Adds the files_deleted row and an example entry in all five locales (en, ja-JP, zh-CN, ko-KR, pt-BR), stating the property that makes the opt-in safe: matching is exact per path after separator normalization, with no globs and no directory prefixes, so a declaration can never authorize more than it literally lists, and omitting the field keeps the guard's original unconditional block. Also corrects two claims in the scope-conformance how-to that this change made false. Its opening paragraph described the recorded declared scope as files_modified alone; declared_deletions is now unioned into that comparison. Its "Renames are not detected specially" bullet asserted the deletions guard blocks any entry whose diff contains a deletion, full stop -- which was the whole point of #3003 and is no longer true. Reworked to say what now decides a rename's fate: declare the old path in files_deleted and both halves become ordinary paths for the advisory check, which is also why the old path needs no separate files_modified entry. Documentation that describes the pre-change behavior of the thing being changed is worse than no documentation, because a reader trusts it. * fix(#3003): close every review finding on the declared-deletions opt-in Two independent isolated reviewers, correctness and security. Neither found a blocker; both found real defects, and the directive treats a finding at any severity as blocking. All of them are fixed here. MAJOR -- the submodule worktree gate could not see a deletion-only plan. per-plan-worktree-gate.md intersected $SUBMODULE_PATHS against $PLAN_FILES alone, while $PLAN_DELETIONS was extracted and then never used. Before files_deleted existed, a path had to appear in files_modified to be planned at all, so the gate saw it; the new field plus the new docs telling authors a deleted path needs no files_modified entry opened a hole where a plan whose only submodule touch is a removal kept worktree isolation on -- the exact case #2772 disabled it for. Both channels now feed the intersection. Note the posture is deliberately the OPPOSITE of the cleanup-wave guard: there the channels stay apart because a deletion AUTHORIZATION must never be inferred; here they merge because a safety fallback must never MISS a touch. MAJOR -- same-wave conflict detection could not see a deletion. The planner's implicit-dependency rule compared files_modified only, so plan A editing src/x.ts and plan B declaring files_deleted: [src/x.ts] scored as conflict-free and ran in parallel: one branch removing what the other is writing, which is the sharpest conflict there is. Overlap is now computed across both channels. MINOR (both reviewers, one root cause) -- the advisory union gave one field two matching rules. declared_deletions was unioned into the scope list handed to planWaveScopeConformance, which reads it with prefix-and-glob semantics. So a field that is exact-match-only at the gate silently became wider at the advisory: ["*.md"], inert at the gate, yielded a null prefix meaning "matches everything" and muted the advisory completely, and ["src"] muted all of src/. The union also activated the advisory on plans that declared no modification scope at all, warning on every modified path. Replaced with subtraction from the findings, gated on files_modified alone. One field, one rule, everywhere. MINOR -- core.quotepath made the feature silently inert for non-ASCII paths. git emits "tests/\303\251.ts" C-escaped and quoted, which never equals the declared plain path, so a correctly declared deletion of tests/é.ts would block forever with nothing pointing at the encoding. Both diffs now pass -c core.quotepath=false. NIT -- flag() consumed a following flag as a value, so --deletions --files x swallowed --files and dropped both. Now treated as a missing declaration, which fails closed. Fixed at both call sites; the helper is duplicated verbatim in cmdWorktreeRecordAgent and cmdWorktreeCreate and leaving one would reintroduce it. TEST -- one test passed for the wrong reason. "a declared deletion is in scope for the advisory" asserted only that warnings omit the deleted path; under a full revert the entry blocks first, warnings come back empty, and the negative assertion passes anyway. It now asserts the entry actually merged, which is the load-bearing half. Four regressions added, one per fix above. Docs corrected rather than extended. The rename bullet in the scope-conformance how-to claimed a rename whose delete side is undeclared never reaches the advisory. Verified false: git's rename detection is on by default, so a pure rename is a single R entry that appears in no --diff-filter=D output and was never gated, before or after #3003. Only a rename that edits enough to fall below the similarity threshold decomposes into add+delete. The pre-existing sentence made the same wrong claim; this restates it correctly instead of sharpening the error. The localized plan-md.md reference edits are reverted: the PR template requires docs content added here to be English, and the translations already lag by three fields, so English-only is the repo's standing posture, not an oversight. Agent-file size caps respected: gsd-planner.md is XL-tier by bytes but carries a separate 49152-LF-CHAR cap asserted by four suites, so its edit is deliberately terse and lands at 49141 with 11 chars of headroom, with the rationale moved to docs/reference/plan-md.md, which has no cap. gsd-plan-checker.md lands at 49107 bytes, 45 under the LARGE cap. Both acks merged into the existing fragments that already name those paths, since two ack sources may never name the same path. * fix(#3003): decode git's path quoting instead of changing the git argv The previous commit's non-ASCII fix turned the remote suite red: 44 failures, 42 of them "unexpected git call: -c core.quotepath=false diff --diff-filter=D --name-only ...". The suite's git mocks match on exact argv, so adding two flags to the deletions diff and the advisory diff invalidated every existing fixture in tests/worktree-safety.test.cjs. Rewriting dozens of fixtures to accommodate one flag would be paying a large Hyrum's-law bill to fix a small defect. Both execGit calls are reverted to their original argv. The C-quoting is now decoded in normalizeScopePath instead, via a new decodeGitQuotedPath helper. That is the better fix on its own merits, not merely the cheaper one: the git argv is untouched so no fixture moves, the decode lands on the ONE normalizer already applied to both sides of the comparison so the declared and reported paths cannot disagree, and it holds regardless of the user's own core.quotepath setting rather than only when we remember to override it. A value not wrapped in a leading AND trailing quote is returned completely untouched, so the plain-ASCII path -- the overwhelmingly common case -- is byte-identical to before. Escapes decode to BYTES collected into a Buffer and UTF-8 decoded only at the end, because \303\251 is two bytes forming one character and decoding them separately yields mojibake. Malformed input never throws: a trailing lone backslash or a short octal escape degrades to the literal character, since one bad path must not take down a cleanup wave. Caught while reviewing the helper: the non-escape branch pushed a UTF-16 code unit rather than UTF-8 bytes. Git always escapes non-ASCII so its own output was fine, but this normalizer runs on the DECLARED side too, and an author may write a quoted path holding a literal é -- pushing 0xE9 alone is invalid UTF-8, so the declaration would decode to a replacement character and silently stop matching. That is precisely the failure this change removes, reintroduced on the other side of the comparison. Now converts whole code points, surrogate pairs intact. The other 2 failures: tests/parallel-dependent-plans.test.cjs pins the exact unbackticked substring "files_modified overlap" in gsd-planner.md, and rewording that comment to "declared-scope overlap" deleted it. The comment is restored verbatim and the files_deleted change rides in the pseudocode and the Rule sentence instead. Recorded in the ack fragment so the next contributor does not rediscover it the same way. Four regression tests cover the decode through the public cleanup-wave seam (the helper is module-private): a declared non-ASCII deletion merges against a C-quoted git report, the symmetric case where the DECLARATION is the quoted form, an undeclared non-ASCII deletion still blocks with the residue naming the decoded path an operator can act on, and a path merely containing a quote is left alone. Plain ASCII was already covered and is not duplicated. * fix(#3003): revert the leading-dash flag guard, the review nit was wrong The remote suite came back with 2 failures, down from 44, and both point at the same thing: tests/worktree-safety.test.cjs:7045 already pins the opposite contract, deliberately. test('a flag-shaped --files value is not re-parsed as a flag', ...) recordAgent(['--files', '--branch']) -> files_modified === ['--branch'] -> branch === 'worktree-agent-a1' ("the real --branch value must be untouched") So consuming the next argv element positionally, whatever its shape, is the tested intent of this parser, not an oversight. The security reviewer's nit claimed --deletions --files x would "swallow --files and drop both". It does not: each flag runs its own indexOf, so --deletions records the literal '--files' while --files independently still resolves to x. And that literal is a path git never reports as deleted, so it authorizes nothing -- already fail-closed with no guard at all. The guard bought no safety and silently changed --files behavior along the way, outside this issue's scope. Reverted at both call sites, which are byte-identical again, along with the test asserting the reverted behavior and the docs sentence describing it. The nit is recorded as REJECTED in the review artifact with the reasoning above, rather than as fixed -- a finding that turns out to be wrong should leave a trace of why, or the next reviewer files it again. docs/CLI-TOOLS.md now states the positional-read behavior plainly instead, so the next person meets it as documented intent rather than rediscovering it through a red suite. * chore(#3003): backfill changeset pr number to 3757 * test(#3003): cover parsePlanDocument's filesDeleted branch to clear the mutation gate CI's Stryker shard for plan-document failed at 73.28 against a break threshold of 75: 170 killed, 62 survived, 232 total. Eight of those survivors are the filesDeleted block this issue added to parsePlanDocument, which shipped with no direct coverage at all -- the field was exercised end to end through the cleanup-wave tests, but the parser itself was never called with a plan that declares it, so every mutant in the block lived. Four tests, each pinned to specific mutants rather than written for coverage percentage: - absent key yields exactly [] -- kills the array-literal seed (["Stryker was here"]) and the `fmDeleted = true` conditional, which would otherwise produce ["true"] - a scalar underscore `files_deleted:` wraps into a one-element array -- kills `fmDeleted = false`, the `&&` logical-operator swap, the `fm[""]` string mutation on the first operand, the emptied if-block, and the ternary's non-array branch - an array-valued hyphenated `files-deleted:` maps element-wise -- kills the `fm[""]` mutation on the SECOND operand (only reachable when the legacy hyphen alias is the one carrying the value) and the ternary's array branch - an empty list yields [] -- boundary case, and a genuinely distinct one from the absent key: [] is truthy in JS so it ENTERS the if, and only Array.isArray's true branch mapping over nothing produces the same [] Threshold arithmetic: 174 of 232 are needed for 75%, and these take it to about 178, so the shard clears with margin rather than landing on the line. Every expected value was confirmed by executing the built parser before being asserted, not inferred from reading the source. --------- Co-authored-by: sim <sim@local> |
||
|
|
4918c62d76 |
feat(#2845): require provenance for UI-SPEC component inventories (#3745)
* test(#2845): failing-first suite for UI-SPEC inventory provenance Binds two shared formats before either exists, so the suite is RED against next: the gsd-ui-checker dimension roster (asserted independently on twelve surfaces, eight English and four translated) and the provenance-line grammar the UI-SPEC template emits and Dimension 7 consumes. Every parity assertion is paired with a synthetic mutation case, so the guard's failure branch executes rather than only reading a correct tree: limit-1 (a surface still declaring 6), limit (7), limit+1 (8), a dropped dimension, a label that drifts on one surface only, a non-contiguous roster, a duplicated number, and a surface that stops declaring a count at all. A seeded fast-check property renders the roster under formatting noise (CRLF, padding, interleaved sections) and asserts the parse round-trips and is strictly sensitive to a dropped heading. Assertions are on parsed typed records, never raw substrings. * docs: normalize design-a-ui-phase how-to to American English House style for docs/ is American English (CLAUDE.md). This file carried colour/initialisation/initialise/artefact throughout. Spelling only — no content change; kept separate from the #2845 feature commit so the release-notes classifier and the hotfix cherry-pick filter see it for what it is. * feat(#2845): require provenance for UI-SPEC component inventories A UI-SPEC's component inventory was treated downstream as a closed allowlist while the document recorded nothing about whether the list had been enumerated from the installed design system or recalled from memory. A recalled inventory is indistinguishable from an enumerated one, so an executor complying with the spec builds against a fraction of what the package offers, and every gate stays green because they assert semantics rather than composition. The UI-SPEC template gains a Component Inventory slot carrying one of two provenance lines: the command that enumerated the list, the count it returned, the resolved package@version and the date; or a Could not enumerate record with a real reason. gsd-ui-researcher gains an enumeration ladder and must record the line rather than write the list from recall. gsd-ui-checker gains Dimension 7. An inventory with no provenance line, a count with no command, an empty could-not-enumerate reason, or a line still carrying the template's unfilled placeholders BLOCKs; a partial line, a line placed below its table, or an honest negative record FLAGs; a complete line passes, and so does a spec carrying no inventory at all, which keeps every UI-SPEC predating the dimension validating unchanged. Whatever the verdict, an unsourced inventory is reported as a non-exhaustive list of known-good components rather than a closed allowlist, so the executor is never blocked from a component the spec merely failed to mention. The checker never runs the recorded command. The dimension count moved on all thirteen surfaces that assert it, across five languages. Also corrects the claim in the English, Korean and Portuguese how-tos that this checker applies a scored six-pillar rubric — that rubric belongs to /gsd-ui-review's retroactive audit. * chore(#2845): backfill changeset pr number to 3745 --------- Co-authored-by: sim <sim@local> |
||
|
|
94bc492f57 |
fix(#3645): tracked-source rule for planner/pattern-mapper path resolution (#3728)
* test(#3645): failing-first agent tracked-source contract rows * fix(#3645): tracked-source rule for planner and pattern-mapper paths * Revert "fix(#3645): tracked-source rule for planner and pattern-mapper paths" This reverts commit 61f05e947bbdaf3b4897240c3819d349215744fb. * fix(#3645): tracked-source rule at the spawn seam and mapper gate * fix(#3645): review fixes - bounded block, ack merge assertion, git wording * fix(#3645): fit the tracked-source block under the 1168 ceiling * chore(#3645): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
14679b866b |
enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability Binds the approved triage shape before any of it exists: - containment — the execute:wave:post hook must not render unless workflow.live_dom_uat is true AND the capability resolves active (fail-closed on a missing state entry, and on a non-boolean value) - criterion 4 — agents/gsd-executor.md carries no browser MCP family; asserted as an absence, which is the only way it is observable - Hyrum guard — the pre-existing mcp__playwright__* branch must stay outside the key-gated block, or upgrading silently removes working automated UI verification for every current Playwright-MCP user - parity — the browser glob list now lives in two surfaces (agent frontmatter + workflow detection block); the assertion fails if either gains or loses a family without the other Red by construction: the capability, agent and workflow block do not exist yet. Verified on the remote runner. Refs #2856 * enhance(#2856): add default-off live-DOM UAT capability A phase whose acceptance criteria needed a live DOM could not be finished by the agent that executed it: gsd-executor carries no browser tools, so it correctly returned checkpoint:human-action even though the work was not human-only, just tool-less. Every such phase degraded to "executed, then finished by hand in the orchestrator", and autonomous: false could not distinguish "a human must judge this" from "the executor lacks the tool". Implements the shape approved at triage, not the one reported. The executor's tools: line is NOT widened, in any configuration: for a first-party agent the static list is the only control that exists (ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one default-off capability owns the key, the agent, and the step: - capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat (boolean, default false), one additive step at execute:wave:post (onError: skip, gates: []), so it can never halt a wave - agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP globs, in its own tools: line, with no Bash - verify-work automated_ui_verification — a gsd:live-dom-families block naming both new families AND the key; presence alone never activates Two independent fail-closed gates: isCapabilityActive renders a hook only on state.active === true, plus the step's own `when`. The pre-existing mcp__playwright__* branch keeps the gating it already had and stays outside the new block. Pulling it behind a default-off key would have silently removed working automated UI verification from every current Playwright-MCP user on upgrade. Also closes a host gap this surfaced: execute:wave:post dispatched only contribution + gate, so ANY registered step was declared and silently never run — exactly the single-kind hand-roll loop-hook-dispatch.md names. Step 5.75 now dispatches every kind == "step". The browser-profile lock is tolerated, not coordinated: --isolated is a flag on the operator's own MCP-server registration that GSD neither launches nor parameterizes, so the verifier reports could_not_look / profile_locked, names the flag, and stops. DOM-VERIFY.md keeps could_not_look and nothing_to_report distinct behind a closed reason enum — collapsing them is the ambiguous-run-notes defect reported. Verified on the remote runner. Closes #2856 * fix(#2856): apply review findings from the orthogonal passes Correctness pass (blocker): - delete detectionBlockIsCrlfSafe. It was pass-always: it read the file, replaced LF with CRLF, then indexOf'd marker strings that contain no newline, so the replacement could not change the result and the assertion could never fail for the reason it stated. There is no real CRLF risk on this surface either — the gsd:live-dom-families block has no parser, only human and agent readers. Deleted rather than replaced, per the repo's pass-always-test rule. Isolated security pass (two minors, both real): - execute-phase.md step 5.75: this change is what first activates kind == "step" dispatch at execute:wave:post, which newly opens the ref.command shell path at that loop point. Our own step uses ref.agent and never touches it, but the door is now open, so the step-dispatch line carries the same in-context validate-before-shell warning the sibling gate-dispatch line directly below it already carries. - gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker influenced. Require it wrapped in inline code or a fence, kept short, and never left reading as a directive to the next reader. Verified on the remote runner. Refs #2856 * fix(#2856): settle the new-agent roster ripple Checkpoint 2 returned 28 failures, none in the new suite — all of them the guards that exist to make adding an agent a deliberate act. Each is a real boundary that had to move: - docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526), so the browser globs lose their backticks; primary-agent counts 21->22, roster 33/34->34/35, Verifiers category 1->2 - docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md to be classified exactly once - gsd-dom-verifier: add the anti-heredoc instruction and the commented hooks: frontmatter pattern both agent gates require - gsd-core/bin/shared/model-catalog.json: every shipped agent needs a profile entry (#3229) - copilot-install / kilo-upgrades / qwen-upgrades: expected agent list and the 34->35 roster boundary - execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately carries one step now. Asserted as an exact shape — one step, capId live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a real guard against accidental change rather than being relaxed Two findings worth naming: mcp-tool-inheritance (#2526) rejected the agent for documenting mcp__playwright__* while its tools: line withholds it — a dead instruction that invites the agent to claim a path it cannot take. The prose now names the Playwright MCP family without the dispatchable token, in both the agent and the capability fragment. runtime-launcher-parity rejected the new gsd_run call: each fenced block is its own shell, so a workflow step file invoking gsd_run needs its own canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs. That script also normalizes explore.md, which is unrelated pre-existing drift the parity check tolerates, so it is reverted to keep this diff scoped. The emitted-drift ack supersedes the spent #3370 entry for execute-phase.md — it is merged into next, so its ripple is absorbed at the base and it can no longer clear anything. That is the same supersede the #3370 entry itself performed on the spent #3324 fragment. Its unrelated execute-plan.md entry is untouched. Verified on the remote runner. Refs #2856 * fix(#2856): drop the stale emitted-drift ack entry The automated-ui-verification.md entry was written speculatively rather than from a reported growth, and the check names that precisely: an ack "written or reworded in THIS diff, but nothing here needed it, so it explains nothing". The growth tier keys on the bare filename as it appears under gsd-core/workflows/ or agents/. automated-ui-verification.md is nested under verify-work/steps/, so it was never in the tracked set — only execute-phase.md was ever reported, both before and after the launcher preamble landed. Only ack what the check actually reports. Verified on the remote runner. Refs #2856 * chore(#2856): backfill changeset pr number pr:0 -> 3716. The placeholder fails both changeset-lint (fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by design and can only be resolved once the PR number exists. Both now report ok against GITHUB_BASE_REF=next. Refs #2856 --------- Co-authored-by: sim <sim@local> |
||
|
|
8df5cb36c2 |
enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata (#3718)
* test(#2951): pin the absent-evidence provenance contract (failing first) 17 tests / 22 anchors on the deployed agent text. Measured against the parent commit: 20 anchors fail, 2 pass. The two that pass are the sibling-integrity guards on the package-name and in-repo-value rules -- green before and after is their intended signature. Refs #2951 * enhance(#2951): refuse [VERIFIED] for a compatibility claim resting on absent metadata A claim of the form "X does not support Y" drawn from MISSING metadata -- no python_requires, no engines field, no per-version classifier, no changelog entry, no matching support-matrix row -- no longer earns [VERIFIED] however authoritative the source consulted. An absence is silence about every value, so the same evidence would "prove" both the version being ruled out and the version being standardized on. The only route from an absence to [VERIFIED] is a positive falsification attempt with its failing output pasted; everything short of that is [ASSUMED], which the file already routes to "needs user confirmation before becoming a locked decision". Third member of the family beside the package-name and in-repo-value provenance rules, mirroring PR #2768's shape. A present declared constraint and an affirmatively documented incompatibility are untouched. Closes #2951 * fix(#2951): close the allow-list ambiguity and the mutation gap review found Findings from the isolated adversarial pass and the two-axis review, all fixed: MAJOR (x2, one root cause) -- the absence clause and the present-constraint carve-out gave opposite verdicts on the same evidence for the commonest real case: a classifier list enumerating :: 3.9 through :: 3.13 with no :: 3.14. A researcher could read the enumerated list as a "declared" positive constraint and re-earn [VERIFIED], which is also the evasion vector. The rule now states the decision procedure -- does the declaration bound EVERY value or only the ones it names -- and closes the positive-reframing restatement explicitly. New contract test pins all four clauses. MAJOR -- 'licenses a positive falsification attempt as the route to [VERIFIED]' asserted two independent substrings and never that the route lands on [VERIFIED]. A mutant swapping the tag for [CITED] or [ASSUMED] inverted the rule and survived all 17 tests. Now pinned as one joined sentence. MINOR -- the attributable-failure test regex-matched illustrative examples ("a missing certificate, a wrong host"), so a copy-edit would break it for no reason; relaxed to the substantive clause. The no-paraphrase guard counted only the heading, missing the drift mode in its own name; it now also pins the core proposition to one occurrence, and the test name matches what it checks. An off-by-one in the new allow-list regex bound (141 actual vs 140) is fixed. MINOR -- docs/AGENTS.md listed four of the five governed absence forms while the agent prose, docs/COMMANDS.md and the changeset listed five; three copies disagreeing on list membership is the drift this repo treats as a defect. SCOPE -- removed docs/how-to/verify-a-dependency-compatibility-claim.md and its docs/README.md index line. Both reviewers flagged them as a seventh and eighth surface beyond the six the requester capped, and CONTRIBUTING's "Agent or skill change" row requires only docs/AGENTS.md. The actionable four-case guidance is retained in docs/COMMANDS.md, which is inside the approved scope. Ack byte figures corrected for the final size: 44250 -> 46602 (+2352), 2550 bytes headroom under the LARGE cap of 49152. Refs #2951 * docs(#2951): restore the how-to the phase gate requires Reverses the removal in 6404b43d3. Both /code-review axes had flagged docs/how-to/verify-a-dependency-compatibility-claim.md and its docs/README.md index line as a seventh and eighth surface beyond the six the requester capped, and CONTRIBUTING.md's required-docs row for an "Agent or skill change" names only docs/AGENTS.md, so they were dropped. gsd-phase-gate.cjs then denied gh pr create: it refuses when the recorded enablement sequence has more than one step and the how-to quadrant is empty. The sequence here is genuinely four steps -- run plan-phase, read the [ASSUMED] claim, probe or cite or accept it unlocked, then answer discuss-phase's checkpoint -- and the last step lands on a different capability's surface, so a reference table cannot carry it. Compressing the sequence to one step to unlock howToSkipReason would be gaming the gate, which is the same Goodhart failure this whole change exists to close. A machine-enforced repo gate outranks two reviewers' scope preference and my own reading, so the page is restored and the PR body discloses the two extra surfaces instead of hiding them. Reverting is a one-file change if a maintainer prefers the tighter scope. Refs #2951 * chore(#2951): backfill the changeset PR number pr: 0 -> 3718 now that the real PR exists. The placeholder fails scripts/changeset/lint.cjs with fail_invalid_fragment, which also blocks lint-docs-required from consuming the fragment. Refs #2951 --------- Co-authored-by: sim <sim@local> |
||
|
|
79781e68eb |
enhance(#2401): ground verify-command paths and inherit prior-phase commands (#3678)
* feat(#2401): ground <automated> verify-command paths and inherit prior-phase commands Adds a deterministic resolvability probe over each PLAN.md <automated> verify command and surfaces the nearest prior phase's proven commands to the planner at every context window. - src/verify-command-grounding.cts: recognizer (not a shell interpreter) that grounds a leading cd <literal> chain and npm --prefix <literal>, and reports unresolvable rather than guessing. Never executes command text. - gsd-tools check verify-command-paths <N>: per-phase probe, wired into plan-phase.md before the plan-check pass. - init.plan-phase gains prior_verify_commands, ungated by context_window. - gsd-plan-checker: new Verify Command Path Resolvability dimension that reports the failing target and never prescribes a replacement. Also fixes first-match-wins prefix bucketing in scripts/lint-test-file-count.cjs (readdir order is not stable across platforms, so a module whose name extends another's with a hyphen bucketed differently on Linux than on macOS). Closes #2401 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2401): ground the canonical --prefix form, quoted paths, and absolute cd resets Independent review found three defects in the recognizer: - npm --prefix DIR run SCRIPT never reached the script-existence check, because the pattern required npm and run to be adjacent. That is the form the docs tell planners to prefer, so script_missing never fired for it. The prefix flag and its value are now stripped before matching. - --prefix captured with \S+, so a quoted path containing a space was truncated to a stray opening quote and reported as a missing directory - a false blocker, worse than the bug this feature fixes. The capture is now quote-aware. - A chained cd whose later segment was absolute concatenated instead of resetting, producing a nonsense path and another false blocker. The fold now resets on an absolute segment. Also replaces the bespoke phase-directory regex with the canonical phase-id helpers. Real phase directories are NN-slug, not phase-N-slug, so the prior-command harvest matched nothing outside its own fixtures and the planner-inheritance half of this feature was dead code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(#2401): source task blocks from the canonical sectionizer The module carried its own copy of the <task>-block grammar - a fourth hand-rolled mirror of the one markdown-sectionizer owns. verify.cts keeps its copy only because it needs the type= attribute the canonical helper discards; this module never reads that attribute, so it can share the owner outright instead of adding a test around a copy. extractAutomatedCommands now takes task bodies from extractTaggedBlocks and the out-of-task remainder from stripTaggedBlocks. A task-grammar parity test pins the attributed task-name set against the canonical helper across six awkward task shapes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2401): extract agent-file overflow to references and repair the property arbitrary The remote matrix run came back red with 19 failures, four root causes: - agents/gsd-plan-checker.md and agents/gsd-planner.md both blew the 49152 agent cap. Their bodies move to gsd-core/references/, leaving @-reference stubs, per the documented overflow pattern. - The new checker dimension invoked gsd_run before the canonical preamble that defines it. The call is deleted outright: plan-phase.md already runs the probe and hands the result in as {VERIFY_PATHS}, so the dimension consumes that rather than re-running anything. - fc.fullUnicodeString does not exist in fast-check 4.8.0. Replaced with fc.string({ unit: 'binary' }), which covers the same 0000-10FFFF range. - Three runtime-loaded files grew; acknowledged in the existing ack fragments that already own those bare filenames, since two ack sources may never name the same path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2401): regenerate golden install-tree fixtures for the new references Adding two files under gsd-core/references/ changes what the installer emits into every runtime's tree, so all 19 golden install-parity fixtures went stale. Regenerated with npm run gen:install-tree; the delta is exactly the two new reference paths per runtime, no removals. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2401): backfill changeset pr number to 3678 * fix(#2401): treat ~ as a home expansion only at the start of a path Windows CI caught this on both shards; the Linux-only remote matrix cannot see it. The dynamic-path refusal rejected ~ anywhere, and a GitHub Windows runner's tmpdir is an 8.3 short name - C:\Users\RUNNER~1\AppData\Local\Temp - so a valid absolute Windows path came back unresolvable/dynamic_path. This was a production bug, not a test artifact: any Windows user whose project path carries an 8.3 short name, or any literal ~, silently lost the probe entirely - every command degrading to unresolvable with no explanation. ~ is a home expansion only at the start of a path; elsewhere it is an ordinary literal. The check is now split: $, backtick, *, ? and newline stay refused anywhere (substitution and globs, and the glob characters are illegal in Windows path components regardless), while ~ is refused only leading, tolerating one leading quote since the check runs before quote stripping. The prior tests only caught this on Windows because only Windows puts a ~ in tmpdir. Four new tests pin it on every platform via a fixture directory literally named RUNNER~1. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ec7e49a64c |
fix(#3576): repair all 43 dead references/ cites and gate the canonical resolvable form (#3596)
* test(#3576): gate shipped reference citations on the canonical resolvable form Failing-first gate for #3576: a backticked bare references/<name>.md cite resolves from no install location (agents, workflows, and references all install where a bare relative references/ path is dead). The gate walks the runtime-loaded trees the issue prescribes, strips @~/ include tokens PER-TOKEN (a line-skip guard would miss a bare cite sharing a line with an include — the issue-named trap), pins the genuinely relative ../ href and canonical forms as non-offenders, and checks canonical cite targets exist. 43 offenders today across 19 files. * fix(#3576): repair all 43 dead references/ cites to the canonical resolvable form Every backticked bare references/<name>.md cite across the 19 shipped files rewritten to gsd-core/references/<name>.md — the form every required_reading block and @~/ include already uses, and the only form that resolves from any install location. All 20 cited targets verified to exist; the one genuinely relative href (plan-phase.md's ../references/mvp-concepts.md) is untouched (the repair is backtick-anchored). Growth acks: new fragment for the three first-time paths, #3206-pattern appends to the five fragments already naming the other grown files (two ack sources may never name the same path). execute-phase.md lands at 93,391/93,400 and gsd-executor.md at 49,150/49,152 — exactly the issue's projections; every repair fits. * fix(#3576): drop stale default.md growth ack (nested modes file is hash-attributed, not growth-ratcheted) Review finding: the emitted-attribution ratchet covers only top-level workflows/ + agents/ files; discuss-phase/modes/default.md's delta is source-attributed, so acknowledging its growth is a stale entry the differential lane fails on. * chore(#3576): add changeset fragment * chore(#3576): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
285cd41be0 |
fix(#3206): define "explicit evidence" inline at verifier 5b; repair stale honest-verifier cites (#3435)
* fix(#3206): define "explicit evidence" inline at verifier 5b; repair stale honest-verifier cites
Step 3 item 5b abstained on non-inferable (backstop) truths "unless
confirmed by explicit evidence" with the term undefined — its definition
lived only in the non-included gsd-core/references/honest-verifier.md,
behind a stale bare `references/` cite that 404s. Undefined, the term
falls back to presence + wiring, the exact false-pass the #1154
abstention protocol refuses.
- 5b: inline the compressed definition (a passing wired
held-out/property-based test or directly observed behavior; presence +
wiring never qualifies) and fix the cite. +84 B on the rewritten line;
file lands at 49,151 of the 49,152 LARGE cap.
- verifier-phase-gates.md (already <required_reading>): gains the
backstop-abstention reporting contract — AFK completion line
("complete with N unverified non-inferable checks", never silent,
never a halt) and reason-distinctness (insufficient_spec vs manual-UAT
human_needed). New content, no relocation of measured prose.
- 5c (line 204) and MVP-mode (line 644) bare cites repaired to
gsd-core/references/ (+9 B each).
- Drift acks per ADR-2719 §4; two entries merge-appended into existing
fragments (two ack sources may never name the same path).
- Changeset fragment with the sanctioned pr: 0 placeholder (post-create
backfill).
Sibling census at next@7976b1ca0: 7 bare-cite instances in 4 agent
files; the 3 in gsd-verifier.md are fixed here, gsd-executor.md:429,439
and gsd-doc-synthesizer.md:20,176 stay with the epic #1891 follow-up.
Refs #1891
* chore(#3206): set changeset fragment pr to 3435
* fix(#3206): drop stale emitted-drift-ack entries that trip the ADR-2719 ratchet
The round's ack bookkeeping explained ripples that were already
self-attributed, so `tests/emitted-attribution.test.cjs` failed
deterministically on the PR head with 5 stale acknowledgments.
`agents/gsd-verifier.md` and `gsd-core/references/verifier-phase-gates.md`
appear directly in `git diff --name-only`, so PROVENANCE_RULES attributes
their emitted deltas without an ack; `agents/gsd-verifier.agent.md`,
`agents/gsd-verifier.toml` and `agents/subagents/gsd-verifier.md` are
derived emissions of a changed source and are attributed the same way.
None of the five entries could ever be consumed, so all five were stale.
Removed: the whole `3206-verifier-explicit-evidence.json` fragment (all
four entries) and the `#3206 append` to `0000-legacy-migration.json`.
Deliberately KEPT: the `#3206 append` to
`1955-verifier-coincidental-reliance.json`. Its `gsd-verifier.md` entry is
consumed by the size-growth ratchet, not the hash pass — the agent grew
49049 -> 49151 bytes, and `diffEmitted` treats a base-identical ack as
spent and excludes it from `ackEntries`. Reverting that append as well
turns the stale-ack failure into `1 file(s) grew without an
acknowledgment` (verified both ways locally).
* fix(#3206): compress 5b and re-acknowledge growth after rebase onto next
The rebase onto next (
|
||
|
|
3c61b4a838 |
enh(#3565): sentinel/contract registry + check:contract-drift lint (#3571)
* enh(#3565): sentinel/contract registry + check:contract-drift lint * fix(#3565): report artifact-row markers once and dedupe per marker * docs(#3565): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
1591454357 |
feat(#3409): reject shell guards that cannot observe their own failure arm (#3558)
* test(#3409): failing-first regression tests for unreachable shell guard arms Drives the three live defects fail-first, executing the shipped workflow snippets rather than a re-typed copy: - G1/G2 plan-phase.md Walking Skeleton gate reads `--pick summaries_total`, a field that does not exist, so PRIOR_SUMMARIES is always "" and the gate has never fired (#3365). G2 is the load-bearing negative-space case: it rejects a fix that treats "no answer" as "zero" and fires unconditionally. - G3 plan-phase.md PHASE_REQ_IDS resolves "" instead of the TBD sentinel on a phase with zero requirements. - G4 complete-milestone.md's bare `cat <glob>` blocks on stdin under a nullglob left set by an earlier block (measured hang). Skipped on Windows for G4 only: the FIFO-blocked-stdin mechanism is POSIX only, and a weakened assertion there would pass vacuously. Refs #3409 * fix(#3409): make nine shell guards observe their own failure arm `--pick` coerces a missing field to empty string and exits 0, so the `|| echo <default>` fallback after it fires only on a verb typo, never on the field absence it was written for. Nine sites relied on that arm. - plan-phase.md walking-skeleton gate: `--pick summaries_total` names a field that does not exist under any flag combination, so the gate has never fired on any project (#3365). Repointed at the existing single owner, `phases.list --type summaries --pick count`, which returns a real integer in every case including a project with no `.planning` directory. No new counter is added: a second one would duplicate the ownership ADR-3180 Decision 1 forbids. The gate now fires only on a literal "0", so an unanswerable query fails safe instead of entering skeleton mode. - plan-phase.md phase_req_ids: now falls back to the documented TBD. - The remaining seven convert to an explicit empty test. - complete-milestone.md read all phase summaries through a bare `cat <glob>`; under a nullglob left set by an earlier block that is zero operands, so cat blocks on stdin. Guarded with the array shape the #3300 fix already established in review.md. Refs #3409 * fix(#3409): guard eleven more globs that defeat their own fallback arm The nullglob audit this issue asks for turned up the same class in files #3300 never touched. - Eight bare `cat <glob>` reads (transition, complete-milestone, planner x4, verifier, phase-researcher). With nullglob set that is zero operands, so cat reads stdin and blocks; measured rc=137 at 3s. - Three `ls <glob> || echo "<message>"` sites (session-report, review-backlog and its generated skill). nullglob makes ls succeed listing the cwd, so the message never prints and the user gets a directory listing instead. Guarded with `[ -e "${_ARR[0]}" ]` rather than `[ ${#_ARR[@]} -gt 0 ]`. The count form is correct only when nullglob is set, and six of these seven files never set it: without it the array holds the unmatched literal pattern, so the count is 1 and the guard passes wrongly. `-e` is correct in both worlds. review.md keeps its count guards — that block sets nullglob two lines above them. skills/gsd-review-backlog regenerated from commands/, never hand-edited. Refs #3409 * feat(#3409): add the unreachable-shell-guard drift lint A sibling of lint-planning-prompt-drift.cjs, consuming the shared scripts/lib/drift-scan.cjs rather than copying it, wired into lint:ci. Both detectors are one shape — a fallback arm defeated by a legitimate success-on-empty: - Detector A: `--pick` and `|| echo` on one line. `--pick` is the discriminator because "missing field renders empty at exit 0" is a documented CLI contract, not a heuristic. A rule keyed on gsd_run matched 111 lines, ~132 of them legitimate, and was rejected. - Detector B: `cat <glob>` in command position, and `ls <glob>` whose exit code feeds a real fallback or an if/while head. Informational `ls <glob>` whose stdout is consumed (97 sites) and `|| true` failure suppression (~15) are not guards and never fire. Shrink-only ratchet keyed on (file, trimmed text) with a per-pair count, POSIX-normalized unconditionally so Windows CI cannot report everything fresh and stale at once. Ships with a ZERO-entry baseline: every site it can find is fixed. Exemption is the per-line `# gsd-scan-ignore: #NNN` marker whose reason must name an issue or URL; a malformed reason reports a distinct error rather than silently exempting. No file allowlists. ADR-3409 records the invariant, the measurements behind both detectors, and why the upstream `--pick` contract fix belongs to #3473. Refs #3409 * fix(#3409): resolve review findings — typed surface, sanitized reports, tighter marker Standards axis (blocker): the guard's tests asserted on human-readable stdout/stderr and on free-form baseline-load prose, which CONTRIBUTING prohibits by name. Added the typed surface it prescribes instead of weakening the tests: a frozen REASON enum, a --json report mode, structured loadBaseline errors, and a test locking Object.keys(REASON) so a new reason stays three coordinated changes. Security axis: sanitizeForReport covered every violation field but not the baseline-load error path, which embeds raw JSON.stringify output -- that escapes nothing above 0x1f, so bidi and C1 controls reached CI logs unfiltered. Routed through the sanitizer at the output seam. Security axis: the scan-ignore marker accepted `#0` and a bare `http://`. Tightened to a positive issue number and a URL with a host. This diverges deliberately from the sibling in tests/commit-files-pathspec.test.cjs, whose looser form was copied verbatim; the header now records the divergence. Security axis: G4 built its FIFO with `mktemp -u`, reserving a name without creating it. Now created inside a `mktemp -d` directory. Spec axis: ADR-3409 claimed a ninth site landed after the issue was filed. git blame disproves it -- all nine predate it; the issue's hand count missed one. Corrected. The design and test matrix still specified B9 as a FLAG after implementation reversed it to PASS; both now record the reversal and why. Refs #3409 * docs(#3409): add the how-to for resolving unreachable-guard findings Reference and Explanation are carried by ADR-3409; this is the task-oriented quadrant CI cannot check for. The page exists mainly for one thing the lint structurally cannot catch: both `[ -e "${_ARR[0]}" ]` and `[ ${#_ARR[@]} -gt 0 ]` remove the glob from the command and therefore both pass, but the count form is correct only when nullglob is set — and nullglob is usually set in a different block of the same file. A reference table cannot carry that; a how-to can. Also documents the reason codes, so a reader can tell "nothing to report" from "could not look". No tutorial: this is a gate inside an existing CI loop, not a new entry point a newcomer starts from. Refs #3409 * fix(#3409): bring the touched prompt files back under their size gates The remote run was red on 14 tests, all size/attribution, none of them the regression suite. - agents/gsd-planner.md was 194 chars over a 49152 cap enforced by four separate tests, each of which says the remedy is extraction, not a bump. It had 41 chars of headroom before this branch. Its `## Checkpoint Types` section was an unlinked, condensed duplicate of references/checkpoints.md, which already carries all three types and their XML shapes; the section now points there and keeps the three names and percentages inline. Net -969, margin 1010. - gsd-core/workflows/execute-phase.md sat 2 chars under a comfortable margin assertion. Dropped the AUTO_MODE default: the `|| echo "false"` it replaced was unreachable, so the value was already sometimes empty on next, and its only consumer compares against `true`. Net -16. Left plan-phase.md's AUTO_CHAIN default alone -- that file names an explicit `false` branch, so empty would match neither branch. - Acknowledged the seven prompt files that genuinely grew, one specific reason each. Five of those paths were already claimed by spent fragments identical to next, which blocks a second source naming the same path; removed just the colliding key from each, deleting the two that this emptied. Refs #3409 * test(#3409): extract the whole PHASE_REQ_IDS block, not just its first line G3 failed on the remote runner with '' !== 'TBD'. The test was wrong, not the workflow. The shipped contract is now two consecutive lines -- the capture and the `${PHASE_REQ_IDS:-TBD}` default -- but the helper's `^PREFIX=.*$` regex returns only the first match, so the test executed half the contract and correctly observed the empty string. Renamed to extractAssignmentBlockFor and taught it to consume the contiguous run of lines sharing the prefix. The assertion is untouched: TBD is the right expectation, and weakening it to accept the empty string would have reinstated exactly the class this suite exists to catch -- a check that cannot observe the thing it is checking. extractFencedBashAfterAnchor is unaffected: it is fence-delimited rather than line-anchored, so G1/G2/G4 still capture their full blocks. Refs #3409 * chore(#3409): drop a spent ack fragment that collided on complete-milestone.md #3458 landed on next while this branch was in flight and its fragment claims complete-milestone.md, which this branch also grows. Two ack sources may never name the same path. Its entry is spent: the +9163 it explains is already absorbed at base, so it can no longer clear anything, and the checker's own guidance for spent entries is to delete them. Removing the key emptied the fragment, so the file goes too -- an empty one signals nothing. Refs #3409 * chore(#3409): backfill changeset pr number 3558 * test(#3409): hoist a regex subject out of exec() to clear the injection scan CI's prompt-injection scan flagged `MARKER_RE.exec('# gsd-scan-ignore: ...')`. The pattern `exec[[:space:]]*\(["']` is receiver-blind on purpose, so it catches `require('child_process').exec('...')` -- and the scanner's own header records that RegExp.prototype.exec is collateral, to be handled by its allowlist. Allowlisting the file would blind it to the real exec vector permanently, so the subject is hoisted into a const instead: same assertion, scanner left at full strength, no security surface widened. Refs #3409 --------- Co-authored-by: sim <sim@local> |
||
|
|
8fc88f663d |
fix(#3210): gate unmet preconditions as blocking-human; cap blocker retries at needs_human (#3528)
* fix(#3210): gate unmet preconditions as blocking-human and cap blocker retries at needs_human * chore(#3210): add changeset fragment for PR #3528 * fix(#3210): restore blocking-human carve-out and CRLF-safe split --------- Co-authored-by: sim <sim@local> |
||
|
|
d2fa696a30 |
fix(#3479): treat absent default-true mempalace keys as enabled in every prose gate (#3527)
* fix(#3479): treat absent default-true mempalace keys as enabled in every prose gate The #2982 absent-key fix (capture_artifacts === false) was applied to only one of the sibling gates. Five more hand-written gates in the mempalace skill/command mirrors and the curator agent still used positive presence ('when <key> is true'), silently skipping default-enabled behavior (mirror_kg, diary_journal) whenever the key was absent from .planning/config.json — inverted from the registry-declared defaults. Corrected sites, each now disabled only on an explicit false: - skills/gsd-mempalace-capture/SKILL.md step 3 (mirror_kg) - commands/gsd/mempalace-capture.md step 3 (mirror_kg) - skills/gsd-mempalace-recall/SKILL.md step 3 (mirror_kg) - commands/gsd/mempalace-recall.md step 3 (mirror_kg) - agents/gsd-mempalace-curator.md tasks 1+2 (diary_journal, mirror_kg) Default-false keys (mempalace.enabled, cross_project_tunnels) keep their positive-presence gates. New #3479 regression cases in tests/mempalace-capture-gate-default.test.cjs lock each site's absent/explicit-false boundary and add a registry-parity guard: no gate file may positively gate any mempalace boolean whose registry-declared default is true. * chore(#3479): acknowledge curator size growth from the gate rewording * chore(#3479): add changeset fragment for PR #3527 --------- Co-authored-by: sim <sim@local> |
||
|
|
26f8015cc2 |
fix(#3448): thread next_action through debug auto-resume respawn (#3476)
* fix(#3448): thread next_action through debug auto-resume respawn Both /gsd-debug auto-resume call sites (Section 1c continue-path return handling and Section 4's non-terminal branch) respawned the session manager with identical session_params, making every resume prompt-indistinguishable from a cold start: the checkpoint's recorded next_action and the disposition that any earlier checkpoint was already answered never reached the respawned agent. Two auto-resumes then made no progress and the (correct) no-progress guard stalled the loop. The respawn now carries resume: true, resume_status, and resume_next_action sourced from the checkpoint file; the session manager documents the params and its Step 2 gsd-debugger template forwards them via a DATA_START/DATA_END <resume_directive> (Step 3d's shape), instructing the debugger to proceed directly on the recorded next action without re-raising answered checkpoints. The anti-loop guard (next_action-only heuristic, 3-resume hard cap) is untouched. * chore(#3448): set changeset pr to 3476 --------- Co-authored-by: sim <sim@local> |
||
|
|
43475a2e0f |
fix(#3440): retire GAP CLOSURE PLANS CREATED marker, document artifact return contract (#3443)
* fix(#3440): retire GAP CLOSURE PLANS CREATED marker, document artifact return contract * chore(#3440): backfill changeset pr number * docs(#3440): mark changeset docs-exempt with audit reason --------- Co-authored-by: sim <sim@local> |
||
|
|
d30c99bc92 |
chore(#3421): delete orphan verify-phase workflow, migrate live gates to verifier (#3422)
* chore(#1892): delete orphan verify-phase workflow, migrate live gates to verifier reference * test(#1892): retarget structural suites from verify-phase.md to verifier-phase-gates.md * chore(#1892): reword retired-workflow mentions for removed-but-needed lint * test(#1892): correct stale surface labels in retargeted suites * docs(#1892): add verifier-phase-gates row to locale inventories * chore(#3421): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
2dbee3ebdd |
enhance(#2229): add three-way claim disposition (admit/refute/abstain) to /gsd-explore research pass (#2543)
Closes #2229. Each claim surfaced by /gsd-explore's research pass is dispositioned admit, refute, or abstain, with abstentions routed to a visible ledger instead of being smoothed into confident prose. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim; a strong prior is never authoritative alone. Two guards ride with it: conflict-abstention, and a tier floor that presents a would-be admit as an abstain when the researcher's resolved tier is the budget tier or cannot be determined. To make that floor enforceable, resolve-model now emits the effective tier (--pick tier). It was already computed above the resolve_model_ids omit gate but was unreachable from a workflow, which left the floor inert on every non-Claude install - the model id is blank under omit and runtime-substituted where a tier map exists, and the profile defaults to balanced. The tier signal mirrors every resolution step that can change which tier runs, including the model_policy preset, and reports unknown rather than guessing. Output is additive; model, profile and effort are unchanged. Two residuals are disclosed in the workflow rather than papered over: a raw-model-id model_overrides pin reports unknown and is floored (fails closed), and a model_profile_overrides entry repointing a tier at another tier's model can under-report (fails open, and predates this change). Admin merge used only to satisfy the missing secondary reviewer on a single-maintainer PR. No CI failure and no conflict were bypassed: 38 checks green, remote runner 32255/32255 on both Node lanes. |
||
|
|
6e59f97dd5 |
feat(#1955): flag coincidental reliance in goal-backward verification (#3250)
* test(#1955): failing-first contract for verifier coincidental-reliance advisory * test(#1955): anchor coincidental-reliance assertions on the frontmatter block * feat(#1955): flag coincidental reliance in goal-backward verification * chore(#1955): correct stale workflow tier high-water comment * fix(#1955): close the verify-phase divergence and state the endogeneity limit * docs(#1955): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
c75ce93be9 |
feat(#1954): flag undeclared coupling between same-wave plans (#3237)
* test(#1954): failing-first contract for plan-checker undeclared-coupling check * feat(#1954): flag undeclared coupling between same-wave plans * docs(#1954): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
10da377794 |
fix(#3021): recognize worktree-wf_* branch namespace in all guards (#3109)
* fix(#3021): recognize worktree-wf_* branch namespace in all guards The Claude-orchestration Workflow backend (#1143) creates per-plan worktrees on branches named worktree-wf_<runid>-<n>. Four independent copies of the agent branch allow-list regex (^(worktree-)?agent-...) never learned this namespace: - hooks/gsd-worktree-path-guard.js:176 — FAILED OPEN (process.exit(0)), silently disabling path containment for exactly the concurrent dispatch mode where cross-worktree writes are most likely - src/worktree-safety.cts:21 — silently dropped cleanup-wave manifest entries - agents/gsd-executor.md:503 — FATAL halt on branch check - gsd-core/references/worktree-branch-check.md:33 — same FATAL halt Extended all four to ^((worktree-)?agent-|worktree-wf_)[A-Za-z0-9._/-]+$. The path guard now correctly blocks cross-worktree writes for Workflow- backend branches instead of no-op'ing. * chore(#3021): backfill changeset PR number 3109 --------- Co-authored-by: sim <sim@local> |
||
|
|
589a9b29b0 |
fix(#2962): enable nullglob in for-glob shell blocks for zsh portability (#3087)
* fix(#2962): enable nullglob in for-glob shell blocks for zsh portability Workflow shell blocks are fenced bash but execute in the user's login shell (zsh on macOS). zsh's nomatch default aborts the WHOLE block on an unmatched glob in a for-list (not just skipping the command), silently bypassing every statement after it — including the verify-phase decision-coverage gate, whose optional *-CONTEXT.md lookup used the unsafe for-list form so the DECISION_RESULT= assignment on the next line never ran under zsh. Fix: prepend a portable nullglob shim to every bash block containing a for-glob loop: shopt -s nullglob 2>/dev/null; setopt NULL_GLOB 2>/dev/null Each command no-ops (stderr suppressed) in the shell that doesn't recognize it; the matching shell enables nullglob so an unmatched glob expands to nothing and the loop body is skipped cleanly. Verified locally: both zsh and bash now reach end-of-block (rc=0) on a no-match glob; bash matched-case behavior unchanged. 14 blocks across 7 files: verify-phase.md (4, incl. the decision-coverage gate), review.md, execute-phase.md, resume-project.md, complete-milestone.md, audit-milestone.md, gsd-integration-checker.md, gsd-plan-checker.md (3). Closes the zsh bypass of the #2770 fix. * chore(#2962): add changeset fragment * chore(#2962): backfill changeset PR number 3087 --------- Co-authored-by: sim <sim@local> |
||
|
|
ed360cd99f |
chore(#2995): extend fragment emission to agents/ and reclaim size-cap headroom (#3058)
* feat(#2995): extend fragment emission to agents/ across every read point Epic #1671 Phase 6.4. `composeWorkflow` stripped `<!-- gsd:section -->` markers only for `gsd-core/workflows/`, so a marked agent shipped its markers verbatim into every runtime — and agent text is loaded into a subagent's context on every dispatch. The issue proposed widening the `copyWithPathReplacement` guard. That is a no-op for agents: agents never traverse that function. Agent content is read for emission at five independent points, and the obvious chokepoint `stageAgentsForProfile` short-circuits on the DEFAULT `full` profile (`skills === '*'` returns the real unstaged directory), so a hook placed there is dead code on most installs. Composition now happens at two call sites instead of five parallel surfaces: `stageAgentsForRuntimeWithConverter` (with `agentsKind` and `kimiAgentsKind` routed through it via an identity converter) and the inline agent loop in bin/install.js. Both compose BEFORE any path rewrite, so a `.claude/` -> `.windsurf/` regex can never reach inside a marker attribute — the ordering #2930 established for workflows. `installCodexConfig` was the fifth read point: Codex embeds each agent's prompt into a per-agent `.toml` via its own readFileSync. Call-graph analysis missed it; the exhaustive per-runtime emission sweep found it. That is why the new guard is behavioral rather than structural — a sixth read point fails the sweep without anyone remembering to extend a list. tests/agent-fragments-emission.install.test.cjs spawns a real installer for every runtime at every agent-bearing scope, derived from RUNTIME_META and the capability registry at run time so a new runtime cannot be silently under-covered. It asserts markers are absent AND the `when="always"` body is retained, so marker-absence cannot be satisfied by dropping content. An identity-composer negative control proves the assertion can fail. Verified: 0 install failures, 0 marker leaks, body retained on 27 runtime/scope paths; red before the wiring on claude(global+local), zcode(global+local), kimi, codex and opencode. Refs #2995 * chore(#2995): give the tightest agents headroom and correct the design lock Epic #1671 Phase 6.4, second half. `agents/gsd-verifier.md` had 12 bytes of headroom under its 49,152-byte LARGE cap and `agents/gsd-debugger.md` had 147 under its 57,344-byte XL cap. Both now extract reference material to `gsd-core/references/` behind an @-reference — the documented DEFECT.AGENT-FILE-SIZE-CAP-BREACH remedy: gsd-verifier 49,140 -> 46,371 B headroom 12 -> 2,781 gsd-debugger 57,197 -> 48,851 B headroom 147 -> 8,493 Byte accounting proves no content was lost: the combined agent+reference delta is exactly the new files' headers plus the agents' slim replacement blocks. Each agent keeps its routing table and a one-line summary per entry, so it degrades gracefully on a runtime that does not inline @-references. `agents/gsd-planner.md` is untouched and still passes both char guards (49,130 < 49,152); it needed no change, so it took none. The other nine LARGE/XL agents carry NO gsd:section markers, and that is deliberate, not deferred. `when=` selection is read from gsd-core/workflows/section-manifest.json, which gen-section-manifest.cjs derives from gsd-core/workflows/*.md only — shape `{workflows: ...}`, no per-agent key, no per-agent init entry point. An agent atom therefore fails admission gate (2) ("a fact the init seam demonstrably computes at a real entry point") and would evaluate false forever while looking like working gating. Marking agents would manufacture exactly the silent-inertness rot the frozen vocabulary exists to prevent. ADR-1671 gains three amendments, two of which close gaps /adr-phase-coverage found against what actually merged: - The 19 -> 29 vocabulary widening shipped in #2994 with no coordinated ADR amendment, which that bullet's own rule forbids. Recorded now. - `flag:--verify-only` was one of six atoms #2992 withheld and deferred to "the LARGE/XL rollout phase". Five shipped; this one is permanently rejected, and that disposition lived only in a merged PR body. - Phase 6.4's own finding: emission extends to agents/, gating does not. CONTEXT.md's glossary was stale on both seams — Workflow Fragments Module still listed the original 4-atom vocabulary and described when= as "not yet acted on", and Section Manifest Module still described InvocationFacts as {waveFlag, phaseNumber, hasPriorPhases}. Both now match the shipped contract. Inventory manifest regenerated AFTER build:lib per the documented ordering landmine; 19 install-tree fixtures pick up the two new references. Refs #2995 * chore(#2995): correct the compose-site count and mark the raw stager Self-review found two comment defects in the prior commit. The agentsKind comment claimed composition lands at TWO call sites; it is three, since installCodexConfig's per-agent .toml writer was added after that comment was written. And stageAgentsForProfile is now production-dead — both callers route through the composing stager — while staying exported and unit-tested, which makes it a trap: it does a raw copyFileSync and short-circuits to the unstaged source directory under the default profile, so a future caller would silently reintroduce the marker-shipping path. Its JSDoc now says so. * test(#2995): guard the marker-documenting-doc class for agents Widening the composer's scope to agents/ makes reachable the exact class #2930 narrowed scope to avoid: a file that DOCUMENTS the marker syntax with an unfenced example is indistinguishable from a real marker, so the composer drops that line from the emitted artifact. Three rows. A fenced example must compose byte-identically. No shipped agent may carry a marker outside a fence — asserted by parsing every real agent and requiring zero explicit sections, which is what makes the fence protection load-bearing rather than decorative. And a non-vacuity row asserts an UNFENCED marker IS parsed as a real marker, so if that ever stops being true the second row is guarding nothing. Also applies two review findings: stageAgentsForProfile's new JSDoc claimed it had no production caller, which is false — bin/install.js's _stageAgents still calls it, and its consumers compose before writing. Corrected to state the invariant instead. And a let/const nit in the emission sweep. * fix(#2995): keep verifier status vocabulary in the agent, fix a wrong fixture The first remote run came back red with three failures. Both root causes were mine. 1. tests/agent-frontmatter.test.cjs requires agents/gsd-verifier.md to literally contain HOLLOW and DISCONNECTED. The Step 4b extraction moved that status vocabulary into gsd-core/references/verifier-wiring-patterns.md, so the agent no longer had it. Byte accounting said no content was lost, and byte-wise that was true — but a contract required those tokens to live IN THE AGENT. That is ADR-1671:66's flexReserve floor stated concretely: a load-bearing fragment must not be trimmed out of its host, and "the bytes still exist somewhere" is not the test. The two status tables are restored to the agent and deliberately mirrored in the reference with a note saying so, so the procedure there still reads standalone. gsd-verifier lands at 47,069 B — headroom 12 -> 2,083, rather than the 2,781 the first attempt claimed. 2. Row 12b of the new marker-documentation guard asserted that an unfenced marker example parses as a real marker, and threw instead: "unmatched /gsd:section close marker". The grammar is WHOLE-LINE only. The fixture had put the OPEN marker inline mid-sentence, so it was correctly not recognised as an open while the close, on its own line, was. That is a real refinement of the hazard this guard exists for: only a marker on its OWN line is mis-parsed — which is exactly how a documentation example is normally written. Row 12b now uses a whole-line marker, and a new row 12c pins the inline case as explicitly NOT a marker. No test was weakened to accommodate the change; the change was corrected to satisfy the tests. Refs #2995 * chore(#2995): backfill changeset pr number to 3058 --------- Co-authored-by: sim <sim@local> |
||
|
|
9640968f8e |
fix(#2847): require gap_closure value in plan-gap-closure schema and bind validate_plan to it (#3018)
* test(#2847): add failing-first regression tests for gap-closure frontmatter schema gap --gaps did not load a machine-checked requirement for gap_closure: true. The planner's only validation gate (frontmatter.validate --schema plan) never required it, and plan-phase.md's downstream_consumer contract never mentioned it either, so gap-closure plans could pass validation while missing the field that /gsd:execute-phase --gaps-only filters on. These tests are RED against current production code: no plan-gap-closure schema exists yet, and neither agents/gsd-planner.md's validate_plan step nor plan-phase.md's downstream_consumer block references gap_closure conditionally. * fix(#2847): enforce gap_closure via plan-gap-closure schema --gaps did not load a machine-checked requirement for gap_closure: true. The planner's only validation gate (frontmatter.validate --schema plan) never required it, so a gap-closure plan could pass validation while missing the field /gsd:execute-phase --gaps-only filters on, silently spawning zero executors. Add a plan-gap-closure schema (every plan-required field plus gap_closure) and make the planner's validate_plan step select it when gap_closure mode is active, plan otherwise. Standard/reviews-mode plans are unaffected: plan's required fields are unchanged. plan-phase.md's downstream_consumer block was investigated for a symmetric mention but deliberately left untouched: it sits 36 bytes under the frozen ADR-857 PRE_PHASE6 ceiling and the validate_plan step in gsd-planner.md is the actual call site, needing no help from plan-phase.md's prose. * fix(#2847): compact validate_plan edit under gsd-planner.md size caps Merging origin/next (7 commits, including #2775's gsd-planner.md STRIDE-row edit) left only 22 chars of headroom under four separate hard-coded 49152-char caps on gsd-planner.md (planner-decomposition, precondition-element, reversibility-tagging, security.test.cjs). The verbose validate_plan prose from the previous commit overran all four. Compact the edit to a single line (net +17 chars vs origin/next) while keeping the functional content: schema name, mode condition, and the unchanged base required-fields list. Also: - Fix a real bug in the fix-2847 negative-assertion test: plan-phase.md mentions the literal string "<downstream_consumer>" twice in backtick-quoted prose before the actual opening tag, so a plain indexOf() grabbed the wrong start position and swallowed ~10KB of unrelated content (including a "gap_closure" hit in a Mode: enum line), producing a false failure. Anchor on the tag starting its own line instead. - Merge the emitted-drift-ack fragment for gsd-planner.md with the #2775 fragment brought in by the merge (both named the same path; two ack sources may never name the same path) and correct its byte delta to the actual final number. * fix(#2847): drop stale merge-inherited emitted-drift-ack fragments Merging origin/next brought in three new emitted-drift-ack fragments (1700, 2658, 2775) relative to this branch's fork point. #2775 collided with my own gsd-planner.md key and was already consolidated. #1700 and #2658 don't collide, but none of their entries name a path this branch's actual diff touches (git diff --name-only origin/next...HEAD) — the ripples they explain are already baked into the current next baseline, so they explain nothing here and the emitted-attribution gate correctly reports them as stale (verified live: spike-wrap-up.md from #1700). Delete both fragment files. Neither is referenced by any test beyond a stray comment pointing at an unrelated diagnosis artifact path, not the ack fragment itself. * fix(#2847): restore merge-inherited ack fragments deleted in error 1700-spike-manifest-idea-scoping.json and 2658-trae-instruction-file-path.json exist on origin/next (landed via other, already-merged PRs) and arrived on this branch unchanged via the origin/next merge. The previous commit deleted them to satisfy a stale-acknowledgment finding, but the finding was about the acks being MODIFIED in this diff, not about needing to stop existing — deleting them would have silently reverted two other PRs' already-merged, already-justified byte growth. Restored byte-identical to origin/next (git diff origin/next -- <path> empty for both). 2775-planner-package-legitimacy-gate.json stays consolidated into 2847-gap-closure-validate-plan-step.json: that one was a genuine hard key-collision (two fragments naming the same gsd-planner.md path, which lint-emitted-drift-ack hard-blocks), not a pass-through case. * fix(#2847): bind --schema to gap_closure mode, not hardcode it Prior revision left the validate_plan bash invocation unconditional (--schema plan)) while only the prose sentence above it described the gap_closure-mode branch. An agent executing the shown line literally always validated with the plan schema, so a gap-closure plan missing gap_closure: true still reported valid:true — #2847 reproducing unchanged. Existing tests didn't catch it: they checked for substring presence anywhere in the step, which the prose alone satisfied. Change the bash line to --schema "$SCHEMA" — a real shell-variable reference in the same placeholder convention this file already uses for "$PLAN_PATH" (never literally assigned; the agent resolves it from context, same as PLAN_PATH). A genuine if/then bash conditional already exists elsewhere in this file (load_project_state's INIT @file: check), confirming executed conditionals, not merely descriptive prose, are the established pattern here. Rewrite the regression test to assert on the bash block's literal --schema argument: reject a hardcoded plan) or plan-gap-closure) literal, require a variable reference, and require the step's prose to bind that same variable name. Verified RED against the prior revision and GREEN against this one before committing either state. * fix(#2847): CRLF-safe tests, drop unexplained ack, require gap_closure=true Four items from independent review, all landing together per request: 1. The #2847 regression test file had two CRLF-fragile regexes (local/no-crlf-fragile-split): a bare \n on readFileSync content means a real \r\n checkout returns invocationLine === null and all four executable-content assertions stop asserting anything while still reporting green. Both now use \r?\n. Prior lint report of exit 0 was a false green from a stale eslint cache. 2. The 2847 drift-ack fragment explained nothing: a direct edit to agents/gsd-planner.md is self-explaining, drift-acks exist for emitted-artifact ripple that cannot be traced to a changed source path. Deleted. Restored the 2775 fragment byte-identical to next (git diff --name-status next...HEAD -- tests/emitted-drift-acks/ now prints nothing) — it only conflicted with the now-deleted 2847 fragment, never needed touching itself. 3. plan-gap-closure validated gap_closure by PRESENCE only (unchanged since the original #2847 fix), so gap_closure: false satisfied it — --gaps-only filters strictly on gap_closure === true, so a false-valued plan still validates green and still spawns zero executors: #2847's exact reported symptom, one value away. Added an optional requiredValues map to FRONTMATTER_SCHEMAS; plan-gap-closure now requires gap_closure to equal the string "true" (extractFrontmatter parses every scalar as a string) in addition to being present. Every other schema/field keeps the original presence-only contract. The row that had documented the hole instead of closing it now asserts the fix; a matching unit test locks requiredValues on FRONTMATTER_SCHEMAS. 4. The "names the plain plan schema" assertion matched the bare substring "plan" anywhere in the step, which verify.plan-structure satisfies incidentally a few lines below — the assertion could not fail even if the plain-plan branch were deleted from the prose. Changed to match the standalone backtick-quoted plan token. * fix(#2847): remove contradictory leftover assertion in Row 6 test The gap_closure:false test asserted !present.includes('gap_closure') (correct — matches the implementation's fold-wrong-value-into-missing semantics) immediately followed by a stale, unedited leftover from an earlier draft of the same test asserting the opposite: present.includes('gap_closure'). The second could never pass once the first did; both were in the same diff. Verified before committing: searched every consumer of frontmatter.validate output (agents/gsd-planner.md, docs/CLI-TOOLS.md, all other test files) for any read of the present field — none exist. Nothing depends on "present" meaning "physically exists regardless of value correctness", so the implementation's fold (present/missing stay a full partition of required) is the right call; the test needed to agree with it, not the other way around. Manually replayed all six rows in the plan-gap-closure describe block against the built CLI to confirm each now passes. * fix(#2847): prototype-key guard, wrong-value diagnostic, doc fixes, vacuous tests Six items from an independent SHIP_VERDICT:no review, landing together per request: 1. Prototype-key crash (src/frontmatter.cts): FRONTMATTER_SCHEMAS[schemaName] was an unguarded lookup, so --schema __proto__ (also constructor, toString, hasOwnProperty, valueOf) resolved to an Object.prototype member instead of undefined, the `!schema` check never fired, and the command crashed with an uncaught TypeError and a stack trace instead of "Unknown schema". Now reachable from prompt state (--schema is an agent-bound $SCHEMA), not just an unreachable literal. Guarded with Object.prototype.hasOwnProperty.call before the lookup, checked and rejected before assignment so `schema`'s type stays non-optional. Added a test for all five prototype keys. 2. Wrong-value diagnostic (src/frontmatter.cts, agents/gsd-planner.md): the strict gap_closure === "true" check from the previous fix was correct (fail-closed) but silent about WHY — a plan with gap_closure: True got "missing", indistinguishable from genuinely absent, even though the field is plainly in the file. Added an `invalidValue` field to the validate JSON (present but wrong-valued, disjoint from missing/present) and updated validate_plan's prose to state the exact required literal and explain invalidValue, within the remaining byte budget (49130/49152). 3. docs/reference/plan-md.md: fixed three inaccuracies in the gap_closure row — "this field plus every field above" implied `requirements` (documented Required: Yes) is schema-enforced, it is not; "Type: boolean" implied YAML True/TRUE/yes/1 are accepted, they are rejected (exact string match on literal lowercase true); "must never carry it" stated an unenforced rule as fact. Also switched /gsd:plan-phase and /gsd:execute-phase to the house-style hyphen form for docs/. 4. Vacuous negative assertions (tests/fix-2847-gap-closure-frontmatter.test.cjs): RegExp#test coerces a null invocationLine to the string "null", so both hardcoded-literal checks passed vacuously even if the step or its bash block were deleted entirely. Added a truthy precondition check first. 5. Deleted vacuous/pass-always tests: four in tests/frontmatter.unit.test.cjs strictly subsumed by (or, for the "superset" test, tautologically guaranteed by the same spread as) the deepEqual exact-list test; two describe blocks in the #2847 regression file that were already GREEN at the RED commit (5e5897cd2f17ebf2fc55757bae651bbbeb236289) and pinned untouched files rather than covering anything this change altered — one of them additionally forbade any future legitimate gap_closure mention in plan-phase.md, a trap for whoever frees up that file's byte budget later. 6. .changeset/clever-newts-wake.md: switched /gsd:plan-phase and /gsd:execute-phase to /gsd-plan-phase and /gsd-execute-phase — changesets render verbatim into CHANGELOG.md with no converter in the path, so the colon form would have reached readers naming a command no runtime registers. * chore(#2847): backfill changeset pr number (#3018) --------- Co-authored-by: sim <sim@local> |
||
|
|
de78f2eef2 |
docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate (#3010)
* docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate security-model.md, USER-GUIDE.md, ARCHITECTURE.md, COMMANDS.md, FEATURES.md, and gsd-planner.md's STRIDE template (+ ja-JP mirrors) described the pre-ADR-0656 design: slopcheck as the install-or-degrade gate, with unavailability degrading every package to [ASSUMED]. ADR-0656 inverted this months ago — registry-API verdicts (npm/PyPI/ crates.io) are the gate; slopcheck is an optional escalate-only adapter that no shipped configuration wires. Verified every replacement claim against src/package-legitimacy.cts (checkPackages, classifyPackage, lookupNpm/lookupPypi/lookupCrates) via Memtrace before writing it, so the corrected prose matches the live implementation rather than restating the ADR from memory. Restored docs/explanation/security-model.md:79-84 (and its ja-JP mirror) to original wording after an orthogonal spec review caught that an earlier draft had edited the "Why WebSearch packages are always [ASSUMED]" paragraph — inside the range issue #2775 explicitly named as correct and to leave alone. The ja-JP mirror was missing the closing clause present in the corrected English original ("its absence leaves registry-API verdicts intact rather than downgrading everything to [ASSUMED]") — added for parity. This completes the ja-JP mirror the issue's acceptance criteria named explicitly. zh-CN/ko-KR/pt-BR (not named by #2775, but carrying the same stale design) get the mechanical portion of the same fix: command-string swaps, table headers, ARCHITECTURE.md diagram labels, and technical- term swaps that reuse a word already attested elsewhere in the same file (合法性/적법성/legitimidade for "legitimacy") — surrounding prose untouched. The remainder in those three locales — full-paragraph rewrites of the corrected degrade-path mechanism, deleted "External dependency" bullets, and "manually install slopcheck" code blocks — needs prose composed by a fluent speaker of each language and is filed as open-gsd/gsd-core#3002 with an exact file:line inventory. * test(#2775): acknowledge gsd-planner.md byte growth from the STRIDE-row fix agents/gsd-planner.md grew 14 bytes (49309 -> 49323) from the STRIDE supply-chain row correction (slopcheck -> package-legitimacy gate). Emitted agent/workflow files are byte-tracked; this fragment acknowledges the growth per tests/emitted-attribution.test.cjs's "differential attribution over the real tree" check. * docs(#2775): close ja-JP FEATURES.md gap; fix a ko-KR transliterated heading docs/ja-JP/FEATURES.md:2808 still read the katakana transliteration "スロップチェック verdict" in REQ-PKG-GATE-01 — invisible to a literal "slopcheck" grep, so it was missed when ja-JP parity was checked and declared complete. Corrected to "正当性判定" (legitimacy verdict), matching the term already established in ja-JP/explanation/ security-model.md and ja-JP/USER-GUIDE.md. This was the only remaining ja-JP gap; a full sweep for the transliterated form across docs/ja-JP/ now returns zero hits, and the ja-JP mirror is genuinely at parity. docs/ko-KR/USER-GUIDE.md:398's heading "슬롭체크 판정:" had the same transliteration problem. Fixed inline to "적법성 판정:", reusing the 적법성/legitimacy word already attested two lines below in the same table. A parallel sweep of zh-CN and pt-BR found no transliterated forms of "slopcheck" in either locale. The remaining transliterated occurrence in ko-KR (USER-GUIDE.md:406, the lead-in to the pip-install code block) needs prose composition like the rest of that block and is added to open-gsd/gsd-core#3002's inventory. * chore(#2775): backfill changeset PR number to 3010 --------- Co-authored-by: sim <sim@local> |
||
|
|
c61dd49d95 |
enhance(#2255): blocking catastrophic-shrink guard for curated .planning/ writes (#2301)
* feat(#2255): blocking catastrophic-shrink guard for .planning writes Adds hooks/gsd-write-guard.js, a PreToolUse hook that hard-blocks (decision: 'block', exit 2) a whole-file Write collapsing a curated .planning/ artifact (ROADMAP.md, .planning/milestones/*-ROADMAP.md, STATE.md) below 40% of its on-disk line count. Files under 40 lines are exempt; GSD_ALLOW_PLANNING_SHRINK=1 (named in the block message) bypasses for legitimate milestone resets. Fix 3 of #973 — the only defense independent of per-agent tool config. Registered on the Claude plugin surface (hooks.json), settings-json runtimes (runtime-hooks-surface.cts, self-contained pattern), Kimi spec, and the OpenCode/Kilo plugin buses. Golden install fixtures and INVENTORY regenerated; regression tests negative-controlled (16/16 RED with the hook absent, 16/16 GREEN with it present). * chore(#2255): backfill changeset pr number to 2301 * enhance(#2255): address review — fail-closed reads, typed block output, registration, property test Review fixes for trek-e's CHANGES_REQUESTED on PR #2301: - Blocker 2: register gsd-write-guard.js in BUNDLED_GSD_HOOK_FILES (no-shipping-drift test). - Blocker 3: update the always-on hook enumerations in ADR-766 and CONTEXT.md from six to seven. - Major 4: fail CLOSED on non-ENOENT read errors — only a missing file (new-file Write) passes; EACCES/EISDIR/ELOOP/etc now block, with a typed readError field and the override still honored. Tested, with a negative control against the pre-fix hook. - Major 5: fast-check property test for the SHRINK_RATIO/FLOOR_LINES budget contract (blocked ⟺ newLines < oldLines*SHRINK_RATIO above the floor; sub-floor always exempt), boundary examples pinned. - Major 6: block output now carries typed oldLines/newLines/ overrideEnvVar fields; tests assert on those instead of regexing the free-form reason string. - Minor: CURATED_PATTERNS are case-insensitive (case-insensitive-FS bypass on macOS/Windows); limit+1 boundary tests added for both the floor and the ratio. * enhance(#2255): engage the write guard on Kimi's native payload shape The guard shipped with Claude-vocabulary checks (tool_name 'Write', tool_input.file_path), which #2304 showed leaves a guard dormant on Kimi: the [[hooks]] matcher is registered pre-translated but kimi-cli forwards its native payload verbatim — tool_name 'WriteFile' (bare or module-qualified) and tool_input.path per its tool schemas (src/kimi_cli/tools/file/write.py). The guard matched, saw an unknown name, and exited 0. Apply the same per-guard normalization PR #2326 gives the three sibling guards (name + field mapping, inlined — hook scripts stage as standalone files), and write the block reason to stderr as well as stdout JSON: Kimi feeds stderr, not stdout, back to the model on exit 2, so a stdout-only reason blocks without telling the model why or naming the documented override. Regression tests pipe Kimi-shaped payloads (engage, qualified-name, stderr-reason) plus exemption pins (StrReplaceFile stays out of scope by design; non-curated paths pass) — verified red against the pre-fix guard, green after. * enhance(#2255): rebase onto next; regenerate golden-parity fixtures * enhance(#2255): wire the escape hatch into complete-milestone's reorganize step Review Blocker 1: the guard hard-blocked /gsd:complete-milestone's ROADMAP reorganize — the tree's only legitimate milestone reset and the exact caller GSD_ALLOW_PLANNING_SHRINK was built for. The reorganize step now performs the rewrite through a shell write with the hatch set on the command (a hook inherits the runtime env, so a bare Write cannot carry a per-step override), and a binding test derives the env var name from the guard's typed output and asserts (a) the workflow step sets it and (b) the guard passes the identical catastrophic payload under it — so the next complete-milestone.md edit cannot silently re-break the wiring. * enhance(#2255): drop dead Edit-class mapping from normalizeKimiPayload Review Major 1: StrReplaceFile -> 'Edit' and the old_string/new_string reconstruction were unreachable-by-effect — the guard exits 0 for any tool_name !== 'Write', so nothing ever read the fields they set, leaving guaranteed-surviving mutants against the Stryker bar. The map now carries only WriteFile -> 'Write'; the StrReplaceFile exemption test message states the fall-through it actually exercises. * enhance(#2255): review minors — American spellings; writeSync before exit(2) Minor 1: normalised/normalise -> American house style. Minor 2: the two block paths wrote stdout+stderr via async pipe writes then exit(2) — async-on-Windows, unflushed at exit; fs.writeSync(1/2, ...) makes the block payload durable. * enhance(#2255): assert stderr equals the typed reason, not raw prose Minor 3: the last raw-text match in the suite pinned override-name prose on stderr. The contract is "stderr carries the reason Kimi feeds back" — now asserted as stderr non-empty and byte-equal to the parsed stdout.reason. * enhance(#2255): bind the write-guard's Kimi normalization into the parity test Review Major 2: the guard's normalizeKimiPayload is a 4th inlined copy with nothing binding it. This extends PR #2326's kimi-guard-normalization-parity test (same path and helpers, authored as a superset so either merge order resolves cleanly): sibling byte-parity is existence-gated zero-or-all — trivially green until #2326 lands, full-strength after — and the write-guard copy is bound semantically (map is the value-inverse of convertKimiToolName; the Kimi name for Write must map, or the guard is dormant on Kimi; the path -> file_path half must be present). Byte-parity is deliberately not asserted for this copy: it legitimately omits the Edit-class mapping (Major 1 — dead code in a Write-only guard). * enhance(#2255): refresh golden-parity fixtures for revised guard + workflow * chore(#2255): regenerate golden fixtures after rebase onto next The committed fixture hashes were generated against a tree predating next's latest 11 commits, which independently modified the same install-parity surface. Rebased onto next and regenerated with `npm run gen:golden`. Verified: against upstream/next the regenerated fixtures differ by exactly this PR's own entries -- hooks/gsd-write-guard.js (new), hooks/managed-hooks-registry.cjs, plugins/gsd-core.js, and gsd-core/workflows/complete-milestone.md. No unrelated drift. * fix(#2255): regenerate workflow size baseline for complete-milestone `complete-milestone.md` grew 31071 -> 32061 (+990) when the round-2 review fix bound GSD_ALLOW_PLANNING_SHRINK=1 into the reorganize step, but tests/workflow-size-baseline.json was never regenerated. The per-file workflow baseline test (issue #1074) failed on ubuntu-latest/22 and both macOS shard 1/3 jobs. The growth is justified: it is the escape-hatch binding requested in review round 2 (the guard must not hard-block the tree's only legitimate milestone reset), not incidental bloat. Regenerated via `npm run size:baseline`; the diff is exactly the one entry. * chore(#2255): regenerate golden fixtures and size baseline after rebase onto next * enhance(#2255): bind the shrink escape hatch mechanically — single-use sentinel the guard consumes Round-5 M1: the per-step `GSD_ALLOW_PLANNING_SHRINK=1 tee` prefix was inert (no PreToolUse hook exists on Bash in this family; the write succeeded by dodging the guard, not by the override firing) and the protection was prose. The hatch is now a transport code consults: complete-milestone's reorganize step arms `.planning/.gsd-allow-shrink` with the target's path, keeps the Write tool as the sanctioned path, and the guard — at the block point only — verifies the sentinel is fresh (15 min) and names the pending target, then CONSUMES it and allows that one write. Path-bound + single-use + freshness keep it from becoming a standing unlock. The env var remains as the interactive transport, where it can actually reach the hook. Regression tests written first (negative control: 3 failed pre-fix): the armed-sentinel Write passes and consumes; stale does not exempt; a token for a different file neither exempts nor is consumed; the binding test now takes the sentinel name from the guard's typed output (overrideSentinel), asserts the step arms it, and asserts the step no longer routes the rewrite around Write via a shell pipe. Also in this commit, same file: - m2: block emission is exception-safe — emitBlock() wraps both writeSync sites in their own try/catch that still exits 2, so an EPIPE can no longer convert fail-closed into the outer catch's fail-open. - Header discloses the two reviewed design limits (cumulative sequential shrink; lexical match vs symlinked paths) per round-5 scoping. * docs(#2255): document the sentinel transport across guard surfaces; changeset ends with the (#2255) parenthetical (m4) USER-GUIDE bullet, INVENTORY row (en + ja/ko/pt/zh), the runtime-hooks-surface registration comment, and the changeset now describe both hatches — the single-use sentinel for workflow steps and the env var for interactive use — instead of implying a per-step env can reach a hook. The changeset's trailing `Resolves #2255.` prose becomes the `(#2255)` parenthetical the repo's fragments use (round-5 m4). * chore(#2255): regenerate derived families on the rebased tree (full sweep) Full generator sweep after rebasing onto next @ the body-parser-patched lockfile: build, gen-inventory-manifest, gen:golden, size:baseline. Every regen delta verified to be either a PR-owned entry (gsd-write-guard.js, complete-milestone.md, INVENTORY/USER-GUIDE) or exact convergence to next's committed value for entries our arbitrary-side conflict resolution had left stale (all 18 runtime fixtures checked mechanically). * test(#2255): use helpers.cleanup for sentinel teardown, not raw fs.rmSync The repo's local/no-raw-rmsync-in-tests rule exists for the Windows-EBUSY retry budget; the sentinel disarm now rides it like every other teardown. * chore(#2255): regenerate derived families after rebase onto next Full sweep on the rebased tree (build -> gen-inventory-manifest -> gen:golden -> size:baseline). Every delta is either a PR-owned entry (hooks/gsd-write-guard.js, its registration surfaces hooks/managed-hooks-registry.cjs and the two plugin buses, gsd-core/workflows/complete-milestone.md) or exact convergence to next's committed value across all 18 runtime fixtures. * chore(#2255): regenerate derived families after rebase onto next @ |
||
|
|
0bb7525a62 |
fix(#2943): rename get-library-docs -> query-docs; correct the ctx7 fallback rationale (#2963)
* test(#2943): parity guard against the nonexistent get-library-docs tool Second context7 naming drift after #2017 (which guarded the plugin-marketplace PREFIX). #2017's guard only checks tools: frontmatter lines, not prose bodies — which is where the broken tool NAME (get-library-docs) lived. The context7 MCP server registers only resolve-library-id and query-docs; get-library-docs is a stale copy from upstream's own README. Scans the shipped prose surface (agents/, gsd-core/references|workflows/, commands/gsd/, skills/) and fails if any artifact instructs an agent to call mcp__context7__get-library-docs. Excludes tests/ (a fixture may use the name as a negative input) and CHANGELOG/RELEASE-NOTES-LEGACY (history). Fails-first: 4 offenders today (gsd-executor.md:29, research-documentation-lookup.md:5, discovery-phase.md:68 & :104). * fix(#2943): rename get-library-docs to query-docs and correct the ctx7 fallback rationale The context7 MCP server registers only resolve-library-id and query-docs (verified against upstream packages/mcp/src/index.ts); get-library-docs is a stale name copied from upstream's own README. Four shipped prose sites instructed agents to call a tool the server does not register, so every research path that loaded the canonical reference either errored, fell through to the ctx7 CLI branch, or fabricated a result. - research-documentation-lookup.md, gsd-executor.md, discovery-phase.md (x2): get-library-docs -> query-docs, params context7CompatibleLibraryId/topic -> libraryId/query (the registered contract). - Same files' ctx7 CLI fallback rationale: the cited cause (anthropics/claude-code#13898 'strips MCP tools from agents with a tools: frontmatter restriction') was wrong on two counts — #13898 is closed and was never about tools: frontmatter. Rewritten to describe the real mechanism (custom subagents cannot see project-scoped .mcp.json; they only inherit user-scoped ~/.claude/mcp.json). The fallback itself is kept. - discovery-phase.md 'mode: code/info' dropped — query-docs takes libraryId + query only; the code-vs-concepts intent is now expressed via the query text. resolve-library-id is unchanged (still registered upstream). CHANGELOG and RELEASE-NOTES-LEGACY citations are historical record, left as-is. * chore(#2943): add changeset fragment (pr:0 placeholder) * test(#2943): widen parity-guard scan surface to docs/ (isolated-review finding) The isolated adversarial review flagged that SCAN_DIRS omitted docs/, which ships docs/AGENTS.md — agent-consumed prose carrying 8 mcp__context7__* refs. No false negative today (it uses only the wildcard), but a future banned-name addition there would slip through, recreating the exact drift this guard exists to prevent. Add docs/ to the scan surface, with an EXCLUDED_FILES set for historical record (docs/RELEASE-NOTES-LEGACY.md, CHANGELOG.md) that must not be rewritten to satisfy the guard. * fix(#2943): update shifted PROSE_ALLOWLIST line + acknowledge gsd-executor.md growth The gsd-test gate caught two real consequences of the rationale rewrite in agents/gsd-executor.md (the +2-line corrected mechanism description shifted line numbers below it): 1. tests/no-bare-gsd-tools-command-position.test.cjs: the legitimate 'gsd-tools query commit' descriptive mention moved from line 791 -> 793. Update the PROSE_ALLOWLIST entry to the new line (the mention is unchanged, just relocated by my edit above it). Without this the gate reports both a stale allowlist entry (791) and a new offender (793) for the same mention. 2. tests/emitted-drift-acks/2943-context7-tool-name.json: gsd-executor.md grew 95 bytes (the accurate mechanism rationale is longer than the wrong one-line #13898 attribution it replaces). Acknowledge the growth with the reason. Both are mandated by the gate, not optional. The rename itself (get-library-docs -> query-docs) is byte-neutral-ish; only the rationale rewrite grew the file. * chore(#2943): backfill changeset PR number 2963 --------- Co-authored-by: sim <sim@local> |
||
|
|
07603df8f2 |
fix(#2647): code-fixer worktree under .claude/worktrees/, not a hardcoded /tmp path (#2942)
* test(#2647): failing-first — fixer worktree path must be repo-relative not /tmp * fix(#2647): place code-fixer worktree under .claude/worktrees/, not /tmp The gsd-code-fixer agent hand-rolled its worktree at a hardcoded `/tmp/sv-${padded_phase}-reviewfix-XXXXXX` mktemp path. On Windows/Git Bash that landed OUTSIDE the project tree — outside the agent session's permission allowlist, so every Read inside the worktree prompted (~25/run) — and mktemp's MAX_PATH-avoidance substitute produced an un-removable `C:/mvwtNN` path. Place the worktree repo-relative under `.claude/worktrees/` (the same dir the harness-managed executor worktrees use: gitignored via `.claude/`, inside the session's permission scope), with a $$-PID + epoch suffix for concurrency uniqueness (replacing mktemp's XXXXXX). $main_repo is resolved the same way the cleanup tail already resolves it. Three sites updated: setup_worktree bash, concrete-steps prose, critical_rules. The #2990 `-b "$reviewfix_branch"` invariant is preserved (the folded test asserts it). Failing-first regression added to the #2990 suite in tests/agent-frontmatter.test.cjs. * test(#2647): update #2686 path assertion to expect .claude/worktrees/, not /tmp The #2686 regression test encoded the worktree location as a hardcoded `/tmp/sv-` path (matching sibling GSD agents at the time). #2647 showed that breaks Windows/Git Bash (worktree outside the project tree → permission prompts; mktemp MAX_PATH substitute un-removable). Update the #2686 path assertion to require the repo-relative `.claude/worktrees/` location and forbid `/tmp/sv-`. The #2686 isolation + cleanup assertions are unchanged. * fix(#2647): word-boundary wt= parse + ack the fixer growth vs next Two follow-ups to the #2647 GREEN run: - parseWtAssignments matched `prior_wt=` (no word boundary), polluting the set and tripping the repo-relative + concurrency-unique assertions. Anchor on (?:^|\s)wt= so only the real worktree-path assignment is captured. - emitted-attribution: gsd-code-fixer.md grew 1875 bytes vs origin/next. Update the emitted-drift-ack entry to attribute the #2647 worktree-path change (supersedes the prior #2825 attribution, whose growth is already in next). * fix(#2647): address review — validate padded_phase at the sink + tighten test Code-review + security-review both APPROVED with one actionable minor: padded_phase is interpolated into a worktree PATH and a git BRANCH NAME, but was only validated by the orchestrator (code-review-fix.md), not at the agent sink. The agent prompt is a literal bash contract any caller can spawn, so add a `[[ =~ ^[0-9]+(\.[0-9]+)?$ ]]` self-defense check rejecting traversal/shell metachars (defense-in-depth; not a present vuln — the only caller validates). Also tighten the concurrency-uniqueness test to require BOTH $$ AND $(date +%s) (either-alone was too lax per review). Update the emitted-drift-ack reason to cover the added validation growth. * changeset(#2647): code-fixer worktree under .claude/worktrees not /tmp * changeset(#2647): backfill PR number 2942 * chore(#2938): regenerate stale docs/CONTEXT-INDEX.json on next #2938 (#2928) updated the CONTEXT.md RULESET prose for the new per-PR emitted-drift-ack fragment mechanism (#2914) but shipped a CONTEXT-INDEX.json generated from the OLD prose. lint:generated-sync fails on every PR that rebases onto next after #2938 (the regen produces a 3-line diff bringing three RULESET entries — AGENT_SIZE_BUDGET, EMITTED_ATTRIBUTION, WORKFLOW_SIZE_BUDGET — in sync with the prose already on next). Mechanical regen via `node scripts/gen-context-index.cjs --write`; idempotent; surfaced by the #2647 rebase. No behavioral change. --------- Co-authored-by: sim <sim@users.noreply.github.com> |
||
|
|
5d0fd4dc53 |
fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point (#2905)
* fix(#2825): gsd-code-fixer honors workflow.use_worktrees; never rm -rf a reparse point gsd-code-fixer was the only writer that hand-rolled a git worktree inside the agent prompt and the only one that never read workflow.use_worktrees. With the setting explicitly false, --fix still created worktrees; the fresh worktree had no node_modules, so runs improvised a teardown whose `rm -rf` followed a Windows junction into the REAL node_modules (silent data loss, 3x observed). Defect 1: gate setup_worktree + its cleanup tail on workflow.use_worktrees (same gsd_run query config-get read the four sibling workflows use). When false: edit/commit in the main checkout (wt='.', no temp branch, no sentinel, no cleanup). Defect 2: forbid rm -rf on a possible reparse point in the spec — never fall through to a destructive remove; on failure, stop and surface the error. Defect 3: REVIEW-FIX records where verification ran (main checkout vs worktree). The transactional worktree path (#2839/#2990/#2686) is unchanged when worktrees are enabled. Docs-parity guards in tests/code-review.test.cjs bind the spec to the fix. * chore(#2825): backfill changeset PR 2905 * fix(#2825): drop stale agents/gsd-code-fixer.toml ack + spent entries (emitted-attribution) gsd-test failed: agents/gsd-code-fixer.toml is a STALE ack here — that TOML was changed by #2834 (now in next), not this PR. The 4 base-carried acks (autonomous/discuss-phase-assumptions/next/plan-phase) are spent/inert. --------- Co-authored-by: Test <test@example.com> |
||
|
|
b36e3b7e1f |
fix(#2751): normalize bare gsd-tools command-position calls to gsd_run in shipped source (#2851)
* test(#2751): regression guard — no command-position bare gsd-tools calls Agents/workflows instructed bare `gsd-tools <verb>` invocations that fail with 'command not found' on a shim-only install (#725 fixed only the Codex conversion pipeline; the Claude-facing source shipped them verbatim). Adds a source-text guard (allow-test-rule: source-text-is-the-product) scanning agents/*.md + gsd-core/workflows/*.md for the operative shape `gsd-tools <verb> <arg>`, excluding command -v probes / resolver definitions, with a documented PROSE_ALLOWLIST for descriptive mentions that name the command without instructing literal invocation. A stale-allowlist check ensures entries stay real. RED first; source fix lands next commit. * fix(#2751): normalize bare gsd-tools calls to gsd_run in Claude-facing source The 12 command-position bare `gsd-tools <verb>` instructions across agents/ and gsd-core/workflows/ failed with 'command not found' on a shim-only install (no gsd-tools binary on PATH). #725 fixed this only for the Codex install-conversion pipeline; the Claude-facing SOURCE shipped the bare calls verbatim, and new ones kept accumulating (new-project.md:114 landed 12 days AFTER #725 closed). Rewrite each operative site to the portable `gsd_run` resolver that the same files already define (3-18x each) — a pure command-position token swap preserving all arguments, flags, --files, and surrounding prose. Every runtime now benefits from one source change instead of each needing its own converter. Touched sites (12): gsd-intel-updater (validate/snapshot/extract-exports), gsd-code-fixer (query commit), gsd-planner (learnings.query), gsd-project- researcher (websearch/research-plan/classify-confidence), gsd-phase-researcher (websearch/research-plan/classify-confidence), new-project (project-instruction- file / commit --files), new-milestone (commit --files). Preserves command -v gsd-tools probes, resolver-snippet definitions, and the 4 descriptive prose mentions that NAME the command without instructing invocation. RED @ 1bb12ba1. * chore(#2751): allow-test-rule issue ref + changeset fragment Add the (#2751) tracking ref to the source-text-is-the-product annotation per ADR-456, and the .changeset Fixed fragment (pr:0, backfilled post-PR). * fix(#2751): convert remaining command-position bare gsd-tools calls (verify-summary, windows, worktree, smart-entry, quick-tasks-append) Isolated adversarial review (Step 4) found the first pass missed genuine command-position bare calls because the regression test's hand-maintained 6-verb list silently false-passed verify-summary (the 'verify' branch matched the prefix then died on the hyphen) and omitted windows/worktree/smart-entry/ quick-tasks-append entirely. Convert these 8 additional operative sites across new-project.md, new-milestone.md, ship.md, execute-phase.md, progress.md, smart-entry.md, quick.md. * chore(#2751): backfill changeset PR number (2851) * fix(#2751): normalize allowlist paths to forward slashes — Windows path-separator false-flag The PROSE_ALLOWLIST is keyed by file:line using forward-slash paths, but path.relative() returns backslash separators on Windows, so the allowlist lookup failed and the 6 descriptive mentions were flagged as offenders on the windows-latest CI lane. Normalize rel to forward slashes before the lookup so the allowlist matches identically on every OS. --------- Co-authored-by: Test <test@example.com> |
||
|
|
6932fb16d7 |
enhance(#1699): require read-and-cite provenance for in-repo discrete values (#2768)
* test(#1699): failing-first contract tests for in-repo value provenance
Nine assertions on the deployed gsd-phase-researcher contract: the discrete-value taxonomy, the same-session Read requirement, path-AND-line-range citation, grep-alone exclusion, the verbatim quote and paraphrase ban, the quote-is-the-checkable-artifact guard, [ASSUMED] routing for unquoted skeleton values, a no-regression guard on the pre-existing package name provenance rule, and a single-definition-site guard.
Eight of the nine fail against the unmodified agent on origin/next; the ninth is the no-regression invariant and passes in both states, which is the intended enhancement shape.
Added to tests/research-agent-profiles.test.cjs rather than a new file: that file already carries the allow-test-rule exemption <runtime-contract-is-the-product> research agent .md content is the governed surface, so no new allowlist entry and no change to the per-module test-file count.
* enhance(#1699): require read-and-cite provenance for in-repo discrete values
The claim-provenance system governed external facts (npm registry, official docs, Context7, package-name provenance). For an in-repo discrete value -- an enum, schema or type union, error code, status constant, or filesystem path -- [VERIFIED] could be earned from training memory or a bare codebase grep, which proves a string occurs, not that the definition was read.
A drifted value passes into RESEARCH.md, is lifted by the planner into PLAN.md's <interfaces> context block, and is trusted by the executor, where it fails at parse()/typecheck as a mid-execution deviation -- the most expensive place to discover it.
The rule lands beside its structural sibling, the package name provenance rule, since both say existence is not verification. The verbatim quote is named as the load-bearing artifact: a citation with no quote does not earn the tag, however precise the line range looks. That keeps the rule falsifiable against the file rather than a self-report, which is the Goodhart guard.
Scope note: the quote goes in RESEARCH.md beside the claim, NOT in an <interfaces> block. <interfaces> appears zero times in this agent on next -- it is planner-side, defined at gsd-core/references/planner-interface-context.md:15 as a PLAN.md structure. The issue text and triage both said <interfaces>; instructing the researcher to populate a block it does not emit would be undefined.
Defined once at the definition site; the source-hierarchy recap is deliberately untouched, since restating it is the paraphrase-drift mode META.RULE.brief-no-paraphrase names.
* test(#1699): regenerate agent size baseline and golden parity fixtures
Regenerated via npm run size:baseline and npm run gen:golden, never hand-edited. agent-size-baseline gsd-phase-researcher.md 40866 to 42020 (LARGE tier, cap 49152, 7132 bytes headroom remaining, per ADR-1610's per-file baseline guard). 18 of 19 golden fixtures updated; pi.json is unchanged because the pi runtime ships zero agents.
* chore(#1699): add changeset for the in-repo value citation rule
User-facing behavior change in gsd-phase-researcher, so a Changed fragment is required. pr: 2768.
* test(#1699): acknowledge the gsd-phase-researcher growth for the emitted-drift gate
CI test (ubuntu-latest, 22) failed on emitted-attribution.test.cjs: gsd-phase-researcher.md grew 1154 bytes (40866 -> 42020) without an acknowledgment. The differential emitted-attribution gate landed on next in
|
||
|
|
3274db2757 |
fix(#2526): remove gsd-ui-auditor's uncallable Playwright-MCP block (#2594)
* fix(#2526): drop gsd-ui-auditor's uncallable Playwright-MCP block The agent declares `tools: Read, Write, Bash, Grep, Glob, Skill` — no `mcp__*` grant of any kind — while its body presented a `<playwright_mcp_approach>` block as the *preferred* capture path. That branch was unreachable by construction: the availability check had a fixed answer, the three `mcp__playwright__*` calls could never dispatch, and the "when Playwright-MCP is NOT available" fallback was the only branch that ever ran — 39 lines of instruction loaded on every /gsd-ui-review spawn that also invited the model to claim a capture path it could not take. Remove the dead block, leaving the CLI screenshot path as the sole documented approach. Guard the class in tests/mcp-tool-inheritance.test.cjs, which already owns agent MCP-grant parity: the new block generalizes the #1284 researcher check from two agents and one dispatch table to every agents/*.md and its whole body — no agent may document an `mcp__<server>__*` namespace absent from its own `tools:` declaration. Frontmatter is read through the canonical parser (gsd-core/bin/lib/frontmatter.cjs) rather than a hand-rolled scan, so inline CSV, block sequences, flow arrays, quoted scalars and full-line comments are handled by construction; inline comments inside a scalar survive that parser, so they are stripped explicitly. The check is server-level by design, ignores prose metavariables like `mcp__X__*`, matches hyphenated server ids, and carries a discovery guard plus negative controls for every documented boundary so it cannot decay into a vacuous pass. The session-level Playwright-MCP pass in gsd-core/workflows/ui-review.md is deliberately untouched — workflow files carry no fixed allowlist, so their availability check is genuinely runtime-detected and honest. Fixes #2526 * chore(#2526): set changeset fragment pr to 2594 The fragment shipped with the documented `pr: 0` placeholder because the PR number does not exist until the PR is opened, and scripts/changeset/parse.cjs rejects `pr <= 0`. Now that the PR is open, set the real number so changeset-lint passes. * test(#2526): cover the two-char server-id boundary of the metavariable exclusion The length-1 "prose metavariable" exclusion was tested at length=1 and at real ids (>=3 chars), but never at length=2 — the limit+1 boundary where a server id starts being recognized. Review finding on #2594: `mcp__ab__foo` in a body with no grant must flag `['ab']`. * fix(#2526): treat a bare mcp__* grant as covering every server `grantedServers()` stripped `mcp__*` to the empty string and dropped it via `if (server)`, so an allowlist that grants every MCP server read as granting none — and the guard then fired against a body the grant plainly covered. That is the one input shape that inverts the check, turning it against a correct agent rather than merely missing a bad one. A `/^mcp__\*+$/` token now sets a GRANT_ALL sentinel that short-circuits `ungrantedServers()`. The sentinel `*` is outside REFERENCE_RE's character class, so no body reference can collide with it. A bare `mcp__` with no wildcard stays a typo rather than a grant and keeps failing closed. No agent uses the `mcp__*` spelling today, so this was latent rather than live. Two negative controls pin both halves. * fix(#2526): scan the frontmatter description for MCP references too `ungrantedServers()` scanned `stripFrontmatter(content)` only, so an `mcp__foo__bar` reference in the `description` field escaped the check. That field ships with the agent and the dispatcher reads it, which makes a dead reference there exactly as dead as one in the body. Only `description` is added to the scanned surface, never the whole frontmatter: `tools:` is the grant list itself, so scanning it would let every allowlist satisfy itself and turn the guard vacuous. A negative control pins that boundary alongside the new positive case. All 34 per-agent tests still pass with the wider surface, so no live agent verdict changes — this was latent. * test(#2526): give multi-character placeholders a convention the checker knows The metavariable exclusion is `length === 1`, so the natural placeholders `mcp__SRV__*` and `mcp__SERVER__*` were flagged as real references — and the failure message then offered an author two remedies ("grant the namespace or drop the block") that both misread what they wrote. Adopts the angle-bracket half of the suggested fix: `mcp__<SERVER>__*` is the sanctioned multi-character placeholder, exempt by construction because `<` is outside the reference pattern's character class. This pins an existing property rather than adding a special case. Declines the all-caps half. An all-caps exemption would be a false NEGATIVE for any real server spelled in caps, and a guard that misses a dead reference fails in exactly the direction this check exists to prevent. The bare-caps form keeps firing; the message now names the convention as a third remedy. Also corrects "grants neither" in that message, which was wrong for any count other than two. * test(#2526): pin the zero-length server id, completing the boundary triple `mcp____foo` yields `[]`, but for a different reason than the length-1 case: it is unrepresentable by `/mcp__([A-Za-z0-9_-]+?)__/g` since `+?` requires at least one character, so the pattern skips it before the metavariable exclusion is ever consulted. Pinning limit-1 completes the 0/1/2 boundary rule on its own terms and records which mechanism owns the case. * docs(#2526): correct every drifted AGENTS.md Tools row, not just the one The review asked for the one-line `gsd-ui-auditor` correction (missing `Skill`). Sweeping the defect class first — every `**Tools**` row in docs/AGENTS.md against its agent's `tools:` frontmatter — found it was 26 of 34 rows, so the one-line framing was the reviewer's premise rather than the population. Breakdown of the 26: * 21 omitted `Skill`, 6 omitted `Edit` (overlapping) — under-promises, the same drift class as #2526 but in the harmless direction. * 8 wrote `mcp (context7)` as shorthand while frontmatter granted up to 8 servers (firecrawl, exa, tavily, ref, jina, perplexity, both context7s). * 1 was actively wrong: gsd-debug-session-manager documented `Task`, a tool name that no longer exists — the #2526 shape at the doc layer, naming a capability that cannot dispatch. Every row is now the frontmatter `tools:` value verbatim, which is also what makes the parity guard in the following commit non-brittle. The diff is 26 insertions / 26 deletions, all Tools rows. * test(#2526): guard AGENTS.md Tools rows against agent frontmatter The 26 corrected rows in the previous commit were free to drift because nothing asserted the role card and the frontmatter agreed — the same reason the #2526 block itself survived. Correcting them without an invariant just resets the clock. Lands in agent-classification-parity.test.cjs rather than a new file: that suite already owns docs/AGENTS.md as a contract surface, already carries the `allow-test-rule` exemption for treating the doc as the product, and file count is the unit of CI overhead (docs/TESTING-SUITES.md). Compares the row to the frontmatter value VERBATIM, not as a set — a set comparison would keep accepting the "mcp (context7)" shorthand that hid eight grants behind one, which is the under-documentation half of the drift. Carries the same discovery guard #2526's own check uses: a section with a granted `tools:` but no **Tools** row fails loudly, so deleting a row cannot silently retire its assertion. Both halves are negative-controlled — against the pre-fix doc it fails naming 26 rows (gsd-ui-auditor:339 among them), and with a row deleted it fails on the missing-row assertion. * chore(#2526): note the AGENTS.md drift correction in the changeset The role cards are user-visible, and 26 of them documented a tool set the agent did not have. Type, `pr: 2594`, and the trailing `(#2526)` are unchanged. * fix(#2526): use a CRLF-safe split in the AGENTS.md Tools-row guard `lint-tests` (npm run lint:ci) rejected `rawAgentsMd.split('\n')` under the repo's local/no-crlf-fragile-split rule: Windows autocrlf yields CRLF, so a trailing \r rides into the parsed line. Switched to `/\r?\n/`. Caught by CI on the round-3 push before the response comment went out. * docs(#2526): correct the drift tallies stated in c5a9607e Re-derived the census programmatically from the pre-fix doc instead of by eye. The 26-of-34 headline was right; the breakdown was not. Skill omitted 21 -> 22 Edit omitted 6 -> 7 "mcp (context7)" shorthand 8 rows -> 7 rows The 8 was conflating two things: 8 rows omitted MCP grants entirely, but only 7 of them used the "mcp (context7)" shorthand — gsd-executor listed no MCP at all. Also names the one `Agent` omission (gsd-debug-session-manager, the row that still read `Task`). c5a9607e's message keeps the wrong numbers rather than rewriting a pushed branch mid-review; the test comment and changeset are the durable statements and both are corrected here. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
60bc6ddd6e |
fix(#2568): commit the debug session doc on the manager-driven terminal path (#2731)
* test(#2568): failing-first contract for the debug manager's doc commit agents/gsd-debug-session-manager.md contains zero occurrences of 'commit', so on the manager-driven path — the normal /gsd-debug flow — nothing consults commit_docs and session docs are left untracked. The step exists only in gsd-debugger.md, which does not reach the end of a multi-cycle session. Two layers. The spec placement is asserted against the shipped agent text because that text is what the orchestrator executes; the commit_docs gate the fix relies on is exercised behaviorally through the CLI in temp git projects, because 'the CLI no-ops when disabled' is the assumption that makes calling it unconditionally correct — asserting it in prose would be assuming the load-bearing part. The most important case is the negative: CONTINUE_REQUIRED is non-terminal and must NOT commit. A fix that satisfied the positive cases by committing unconditionally would strand a half-finished session looking done, which is worse than the bug. RED expected on the five spec assertions; the three CLI gate tests pin existing behavior and pass both sides. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2568): commit the debug session doc on the manager-driven terminal path agents/gsd-debug-session-manager.md contained zero occurrences of "commit", so on the normal /gsd-debug flow nothing ever consulted commit_docs and session docs were left untracked — reproduced by the reporter across three consecutive real sessions. An obligation left behind during a responsibility move. gsd-debugger.md commits the doc at :1214, but that runs only when the debugger carries a fix to completion inside a single spawn. The multi-cycle manager took over checkpointing, fix application, archival to resolved/, and the terminal summary — every step that finishes a session — without the commit step. The debugger's copy still exists and still works on its own path, so "is this handled anywhere" answered yes. The manager now commits before both terminal shapes, and explicitly NOT on CONTINUE_REQUIRED. That exclusion is the load-bearing part: the manager's own contract already forbids fabricating a terminal summary on that path, and committing there would be the same lie in git form — a half-finished session stranded looking done. CHECKPOINT REACHED likewise does not commit. Two facts kept the fix small. The manager already carries the gsd_run preamble (:96, used at :97 for resolve-model), so no new plumbing. And cmdCommit (src/commands.cts:809-814) already gates on commit_docs and returns skipped_commit_docs_false when disabled — so the agent calls it unconditionally and correctness follows from the CLI rather than from a second copy of the config check that could drift. Calling git commit directly was rejected for exactly that reason: it would bypass a user's explicit commit_docs:false. In-session fix code is staged by specific file, never git add -A, which would sweep unrelated working-tree changes into a debug commit. Body-only edit — frontmatter untouched, so the research-profiles/AGENTS.md ripple does not apply. gsd-debugger.md's own commit step is preserved; the single-spawn path still ends there, and a double commit is harmless since the second finds nothing to stage. Fixes #2568 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2568): bind the commit paths, define gsd_run in-block, and make the fix commit idempotent Three defects in my own first cut, all found by the isolated adversarial pass. 1. BLOCKER — {debug_dir} was a dangling substitution. I copied it verbatim from the issue's suggested patch without checking scope. It is a real variable in gsd-core/workflows/debug.md, where the ORCHESTRATOR receives it from init JSON and uses it to build debug_file_path — but this agent never receives it: <session_parameters> declares only slug, debug_file_path, symptoms_prefilled, tdd_mode, goal and specialist_dispatch_enabled. An unbound token makes --files resolve to a nonexistent path; cmdCommit skips missing explicit files, staging stays empty, and the doc silently never commits — reproducing #2568 through a different broken path. Same class as #2684's dangling placeholders. Now spelled literally as .planning/debug/resolved/{slug}.md, matching gsd-debugger.md:1214 and this file's own prose. 2. MAJOR — gsd_run was undefined in the commit block's shell. Shell state does not persist across tool invocations, and the Step 2 preamble is ~230 lines and an entire spawn-and-loop earlier. gsd-debugger.md redeclares the full preamble immediately before each of its own call sites; the commit block now does the same, byte-identical to this file's existing definition. 3. MAJOR — the in-session fix commit was not idempotent. gsd-debugger.md's archive_session already commits the fix on the confirmed-checkpoint path, which is the standard find_and_fix flow, so `git add X && git commit` would hit an empty diff, exit non-zero, and abort the step before the summary was returned. Guarded with `git diff --cached --quiet || git commit`. The doc-commit half was already safe — query commit treats an empty diff as nothing_to_commit and exits 0 — and that asymmetry is now stated rather than assumed. Three tests added for exactly these, since the reviewer correctly noted the suite would have caught none of them: every token substituted into a commit command must be a declared session parameter; the preamble must sit in the same block as the call with no step boundary between; and the fix commit must carry the staged-content guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2568): keep the single canonical gsd_run preamble The previous commit added a second preamble to the commit block, on the isolated review's theory that shell state does not persist across tool invocations. That theory is reasonable in general and wrong for this corpus: the repo enforces "each agent .md using gsd_run contains exactly ONE canonical preamble, before the first gsd_run call". Adding a second broke that invariant and five suites with it — the B-agents preamble check, runtime-launcher-parity, and three slash- namespace guards. The single Step 2 preamble already precedes the commit call, so coverage was never actually missing. Reverted to one, and the test now asserts the real property — exactly one preamble, positioned before the call that needs it — rather than the locality I had wrongly encoded. Recorded because the direction of the mistake matters: the review was right that the question needed asking and wrong about the answer, and I shipped the wrong answer without checking the invariant that already governs it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2568): use the canonical /gsd:debug slash form in the new prose My explanatory sentence wrote `/gsd-debug`, the retired syntax. Three slash-namespace guards caught it: the #3443 invariant, the #1975 folded bug-2543 check, and the retired-syntax scan over Claude-facing source. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2568): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
28e486faf7 |
fix(#2608): fail closed when git add fails during commit staging (#2693)
* fix(#2608): fail closed when `git add` fails during commit staging `cmdCommit` ignored `git add` failures. #2523 had already stopped a failed path entering the commit pathspec, but skipping it silently left two bad outcomes, both reproduced against the pre-fix build: - SOME paths fail -> `{"committed":true}`. `git commit` still ran and PARTIALLY committed the subset that happened to stage, under a message describing the full requested scope. - EVERY path fails -> `{"reason":"nothing_to_commit"}`, which is not what happened and points the operator nowhere. In both cases git's original `add` stderr was discarded, so the user saw a downstream `commit_failed` / pathspec error naming an innocent file — the symptom reported in the issue from a linked worktree whose git directory was outside the managed writable root. Staging failures are now collected and the command fails closed BEFORE `git commit` runs, returning the issue's specified shape: { committed: false, hash: null, reason: "staging_failed", file: "<first failing path>", error: "<original git add stderr>", failures: [ { file, error, timed_out }, ... ] } A timeout is distinguished as `staging_timeout` (issue AC5) using the projection's SIGTERM+ETIMEDOUT signal — the same idiom worktree-safety.cts uses. The check is placed ahead of the `nothing_to_commit` branch so an all-paths-failed run reports the staging cause rather than an empty changeset. Unchanged: successful staging still commits exactly the declared scope and leaves unrelated staged files alone; an explicitly-named file that does not exist is still skipped rather than staged as a deletion (#2014/#2523), and a request where every named file is missing still reports `nothing_to_commit` — no `git add` ran, so there is no staging failure to report. Regression tests inject the failure by monkeypatching `execGit` on the projection module (per CLAUDE.md, over `chmod 0o000`, which does not fault under root and would make the tests vacuous), driven in a `node -e` child because `output()` writes via `fs.writeSync(1, …)` and cannot be captured in-process. Pre-fix, 6 of the 10 assertions fail; post-fix all pass. Closes #2608 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2608): roll back the index, guard the sibling surfaces, document the new reasons Six findings from the orthogonal review of the first commit, all fixed here. 1. A `staging_failed` return left the index PARTIALLY STAGED. The paths that did stage stayed in the index with no commit made and no cleanup, so the next bare `git commit` would sweep them up — the same silent partial commit this fix exists to prevent, deferred one step. (Pre-fix the partial state at least got consumed by the incorrect commit.) The staging failure path now resets the paths it staged, matching cmdPrSubrepo's established rollback-then-error convention. The reset is scoped to what THIS call staged — paths the caller had already staged are captured up front and excluded, so a caller's own work is never destroyed — and is best-effort, since an unwritable index (the very failure being reported) cannot be reset either. 2. `cmdCommitToSubrepo` still had the identical defect: a failed `git add` was dropped silently and the function committed the subset that happened to stage, discarding git's stderr. It now fails closed per sub-repo with the same staging_failed/staging_timeout reasons and the same scoped rollback. 3. The `git rm --cached --ignore-unmatch` branch (default mode, for a planning file that no longer exists on disk) still discarded its result. It mutates the index exactly like `git add`, and `--ignore-unmatch` already makes "no such path" a success, so a non-zero exit there is a real I/O failure — now routed through the same staging-failure path. 4. `agents/gsd-executor.md` documented the commit envelope as an exhaustive three-shape enum and pattern-matched only `nothing_to_commit | commit_failed`. It is the sole consumer doc for this surface, so the new reasons are added with explicit guidance not to retry (a retry hits the same unwritable index), and the "one of three shapes" framing is corrected. 5. The default (non---files) staging path and `--amend` are now covered by tests. Both were already guarded by the first commit but unexercised. 6. The changeset framed the fix as `--files`-only; it applies to default and sub-repo commits too, and now mentions the rollback. Regenerated the agent size baseline and the 18 golden install-parity fixtures for the gsd-executor.md edit. 16 assertions across both surfaces verified against the built lib. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2608): update the #2523 out-of-repo contract to the new staging_failed reason The remote test run surfaced this: `#2523: out-of-repo --files path is rejected by git` asserted `reason: 'nothing_to_commit'`, and now gets `staging_failed`. This is a deliberate contract improvement, not a papered-over failure. The old reason existed only because a failed `git add` was skipped and the resulting empty `stagedPaths` fell through to the empty-changeset branch. But "nothing to commit" is not what happened — the caller named a file and git refused it — and that misreport is exactly the class of defect #2608 closes. The result now carries the offending path and git's own message ("… is outside repository at …"), which is strictly more actionable for the same condition. #2523's two substantive invariants are untouched and still asserted: no commit is created, and the index is left clean. Two assertions are ADDED (the path is named, git's message is preserved) so the richer contract is pinned rather than merely allowed. Per CONTRIBUTING, a stale-test correction rides its own commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2608): compact the executor doc addition to stay under the agent LARGE cap The remote test run failed: `gsd-executor.md is 49217 bytes — exceeds the LARGE hard cap of 49152`. The file was already at 48596 (556 bytes of headroom) and the new commit-envelope documentation pushed it 65 bytes over. The cap is a red line, not a budget to raise, so the addition is compacted rather than the cap moved: four lines instead of eight, keeping the load-bearing facts — the two new reasons, that nothing was committed and the index was rolled back, that `file` + `error` should be surfaced, and that retrying is wrong because a retry hits the same cause. Dropped only the restatement of the linked-worktree example (already in the changeset and PR) and the `failures[]` field (a superset of `file`/`error`, discoverable from the payload). Net addition is now 276 bytes; the file sits at 48872 with 280 bytes of headroom. Extracting the agent's shared boilerplate to references/ would buy much more, but that is a restructuring of the executor agent and does not belong in a commit-staging bugfix. Agent size baseline and the golden install-parity fixtures regenerated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2608): backfill changeset PR number (#2693) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bd570618d4 |
feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop * fix(#2632): calibrate against the raw projection so the loop converges * test(#2632): add closed-loop convergence guard and codify the feedback-loop rule * fix(#2632): pair calibration samples per plan; atomic write; amend adr * chore(#2632): backfill changeset pr to 2672 * fix(#2632): retry renameSync on transient windows errnos and clean up the temp |