30468f16febfb0082ff6ebe245954944595bb61b
5777 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1ec4b38bd3 |
test(#4298): add tdd-walk.cjs end-to-end sniff-test harness for TDD dispatch (#4300)
* test(#4298): add tdd-walk.cjs end-to-end sniff-test harness for TDD dispatch Epic #4272 Phase 5's own checklist named this deliverable ("the same class of coverage loop-walk.cjs gives the loop") separately from #4268. Adds tests/qa/tdd-walk.cjs, extracting and REALLY EXECUTING (via a real `bash -c` subprocess against a real temp fixture project) the shipped bash resolution snippets from both TDD dispatch backends — never reimplementing or grep-simulating the predicate. Proves, by execution rather than text-shape assertion: the CLI predicate and both backends agree for a type: tdd plan and a plain plan; the worktree backend's fail-closed guard genuinely halts (non-zero exit, FATAL stderr) on a missing plan file; and the tdd.md embed ternary's condition tracks the real resolved value (#3800). This is exactly the class of proof #4264/#4265 (unassigned/divergent predicate) and #4268 (static-shape checks can't see backend divergence) could not provide. Extraction uses indexOf/slice on fenced-code markers only, never a backtracking regex over whole-file text (per the #4228 incident this repo's tests already document). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4298): scrub ambient env, tighten fail-closed assertion, fix comment Standards+Spec review found: (1) executeBackendScript spread raw process.env unfiltered into the spawned bash subprocess, unlike tests/helpers.cjs's runGsdTools, which deliberately scrubs SESSION_IDENTITY_ENV_KEYS + config-location env vars before spawning (#2665) — an ambient developer/CI override could silently change what phase.tdd-applicable resolves to in a way a gsd-test bench container won't reproduce; (2) the row-5 fail-closed test asserted only `stderr.includes('FATAL')`, which would also pass if the file's unrelated ISOLATION fail-closed guard fired instead of the TDD one; (3) a docstring called the worktree backend's first fenced block a "shim preamble" when it's actually the whole ISOLATION-resolution block. Fixes: spread the exported TEST_ENV_BASE (every scrub-listed key set to '') before the two intentional RUNTIME_DIR/GSD_TEST_MODE overrides; assert the exact TDD-applicability FATAL text; correct the docstring. Re-verified by direct execution against real fixtures — all three precedence-tier cases and the fail-closed case behave identically to before the fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
747a3730d4 |
fix(#4268): harden tdd-single-statement.test.cjs against reworded restatements and backend divergence (#4297)
* test(#4268): harden reworded-restatement and backend-predicate-divergence detection tests/tdd-single-statement.test.cjs's restatesCycle() keyed on the exact literal `commit: `test({phase}-{plan})`` substring, so a reworded restatement of the RED/GREEN/REFACTOR procedure shipped green. Adds restatesCycleStructurally(), a structural (span + list-marker) detector that stays linear-scan (per the #4228 catastrophic-backtracking incident this must not reintroduce) and is proven, empirically, to flag a paraphrased multi-step fixture while not flagging the real compact citations in execute-plan.md and gsd-executor.md (#4267's legitimate pointers). tests/tdd-backend-wiring.test.cjs never compared the two dispatch backends' `gsd_run query phase.tdd-applicable` calls against each other, so a one-word divergence between them (e.g. a changed --pick flag in only one backend) shipped green. Adds a byte-identity assertion on the command-substitution content (normalized for the two backends' differing variable-name prefixes), proven to have teeth via a RED-first mutation check before asserting it against the real files. The third gap in #4268 (nothing proves TDD_APPLICABLE has a real definition) was already covered by this file's existing assertTddApplicableIsComputed (epic #4272 Phase 2, #4266) — verified by inspection, no new test needed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4268): redesign restatement detector around deferral, not length Standards+Spec review of the prior commit proved by execution that a compact, no-list-marker restatement (under the 200-char span threshold) sails past the span/list-marker-only signal. Redesigns the primary check: the actual invariant is deferral, not length — a legitimate RED/GREEN/ REFACTOR mention always names tdd.md as the authority nearby, a restatement never does. Flags when no tdd.md/canonical reference appears within a 500-char trailing window past the cycle mention, regardless of length or list-marker shape; keeps span>200 and three-distinct-list-marker-lines as secondary defense-in-depth OR-conditions. Verified independently against both real files (execute-plan.md span=10, gsd-executor.md span=14, both with a nearby deferral marker at +228/+82 chars) — no false positive, and the reviewer's exact gap class (a 189-char no-citation paraphrase) is now flagged. Also fixes: boundary coverage at the span threshold (199/200/201, isolated via a factored-out measureCycleSpan() helper), a fast-check property test proving the fix holds for arbitrary filler text, and a fragile line-match in tdd-backend-wiring.test.cjs that happened to work only because a FATAL echo message containing the same substring came later in document order than the real assignment line. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5214ad5802 |
fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication (#4295)
* fix(#4267): correct tdd.md pointer citations; fix(#4269): collapse plan-level gate-rule duplication execute-plan.md and gsd-executor.md each cited "Red-Green-Refactor Cycle" for three facts (commit-scope contract, fail-fast rule, error handling), but only the commit-scope contract lives there. Fail-fast is in tdd.md's "Fail-Fast Rules" subsection (under "Gate Enforcement Rules") and error handling is in tdd.md's "Error Handling" section — cite each correctly. gsd-executor.md's "Plan-Level TDD Gate Enforcement" section also fully restated the gate-sequence rules tdd.md's "Gate Enforcement Rules" already owns (and covers more thoroughly, including the actual git-log validation script). Collapse it to a short pointer, matching the treatment already used by the cycle-steps pointer immediately above it. Adds tests/tdd-reference-correctness.test.cjs asserting the pointer text cites the correct section names, that those sections actually carry the guidance, and that the old gate-sequence restatement is gone from gsd-executor.md. Closes #4267 Closes #4269 * docs(#4267): add changeset for tdd.md pointer correctness fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4267): acknowledge execute-plan.md growth execute-plan.md grew 95 bytes (39766 -> 39861) from the corrected three-section citation in the #3990/#4267 cycle-steps pointer. Emitted-Drift-Ack-Growth: execute-plan.md — net +95 bytes from citing the "Fail-Fast Rules" and "Error Handling" sections by name instead of a single mis-scoped section (#4267). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4267): update PROSE_ALLOWLIST line numbers shifted by the pointer-citation edit This branch's edits to agents/gsd-executor.md and gsd-core/workflows/execute-plan.md shifted line numbers, leaving tests/no-bare-gsd-tools-command-position.test.cjs's PROSE_ALLOWLIST pointing at stale lines. Update both entries to their new correct lines (811 and 419 respectively) without changing the underlying prose. * fix(#4267): restore INVALID_RED citation, fix allowlist line shift after #3770 rebase The rebase onto next picked up #3770's already-merged fail-fast update to gsd-executor.md's plan-level gate section, which this branch's own commit collapses into a pointer. The conflict resolution kept the pointer but dropped the literal "INVALID_RED" term that tests/tdd-red-evidence.test.cjs requires gsd-executor.md to name — restored it. Also updates no-bare-gsd-tools-command-position.test.cjs's PROSE_ALLOWLIST line number for gsd-executor.md, shifted again by the rebase. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4267): fix non-matching regression-guard regex in tdd-reference-correctness The fail-fast regression guard asserted gsd-executor.md no longer contains "If a test passes unexpectedly during the RED phase" — but the actual old prose (removed by this branch's pointer-collapse) read "If a test passes unexpectedly during RED, STOP". The regex never matched the real old text, so the assertion would have passed even against the unmodified pre-change file. Caught by an isolated orthogonal review pass. Fixed to match the actual removed wording, and confirmed (via a direct grep) it is genuinely absent from the current file. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4267): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2e056488d9 |
fix(#4067): derive advance-plan phase-complete from disk, not the plan counter (#4292)
* test(#4067): pin advance-plan phase-complete guard matrix (RED) Five-case matrix: decline on unsummarized plans (regression), fire on fully-summarized phase, fail-open on unresolvable phase dir, idempotent decline, normal advance untouched. * fix(#4067): derive advance-plan phase-complete from disk, not the plan counter The phase-complete branch of state.advance-plan was decided purely by STATE.md's scalar plan counter (currentPlan >= totalPlans). A stale counter carried into a newly planned phase, or a counter raced by wave-parallel executors, let 'Phase complete — ready for verification' land while sibling plans were still executing. cmdStateAdvancePlan now re-decides that branch from disk before the write: every plan in the Current Position phase's directory must have a SUMMARY.md (scanPhasePlans single owner, the same source state.update-progress recalculates from). Outstanding plans decline the entire write byte-identically (idempotent, concurrency-safe, counter stays display-only); an unavailable disk answer fails open to the counter-derived decision. * fix(#4067): review round 1 — route phase-dir lookup through listMilestonePhaseDirs #3185 drift guard: no hand-rolled phases-dir readdirSync. Windowed (current-milestone) lookup first so an archived milestone's stale dir cannot shadow the live one; unscoped retry when the window cannot answer. Also restore the transform's undefined-data error semantics and extract scanOutstanding. * chore(#4067): add changeset fragment * chore(#4067): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
d29b50d696 |
fix(#4051): route specific intents first and confirm before dispatch in --do (#4289)
* test(#4051): pin freeform routing specificity contract in do.md * fix(#4051): order freeform routing specific-first, confirm before dispatch, argument-aware forwarding * fix(#4051): regenerate FEATURES.md, satisfy docs-guard on new routing test Emitted-Drift-Ack-Growth: do.md — deliberate growth: specific-first routing table (code-review, plan review, ui-review, secure-phase, audit, docs-update, phase CRUD rows), a REQ-DO-03 confirm step, and argument-hint-aware dispatch. * chore(#4051): fold regression into non-bug-prefixed test filename per lint-regression-test-names * fix(#4051): review fixes — em-dash description style, split audit-fix route * chore(#4051): sync skill mirrors of execute-phase/phase descriptions * chore(#4051): add changeset (pr backfill to follow) * chore(#4051): backfill PR 4289 in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
8249ebcf6e |
fix(#3770): require intentional RED evidence before GREEN (#4279)
* test(3770): add failing tests for intentional RED evidence gate RED: classifyRedEvidence / buildRedEvidenceRecord / check tdd-red-evidence do not exist yet; every row fails on require. Per #3770 only an intentional target-test failure may authorize GREEN; zero-test discovery, fixture crashes, unrelated failures, and unexpected green are INVALID_RED. * fix(3770): require intentional RED evidence before GREEN Only an intentional failure of the TARGET test (distinctly named, TAP-reported assertion failure) classifies as RED_EVIDENCE_OK and authorizes GREEN. Zero-test discovery, fixture/load crashes (file-named failures), nonzero exits without a failing test, unrelated failures, unexpected greens, and malformed/missing records are INVALID_RED and block GREEN. - src/tdd-red-evidence.cts: pure classifier + persisted record builder (reuses the prohibition-enforcement TAP primitives; fail-closed, never throws) - check tdd-red-evidence <record.json>: validates the persisted record (command, exit code, failing test, expected, actual) - gsd-executor.md / references/tdd.md / references/execute-mvp-tdd.md: RED now requires the evidence record + gate verdict, not a nonzero exit or a RED: tag * chore(3770): regenerate inventory manifest for tdd-red-evidence.cjs * fix(3770): fit executor fail-fast under size cap, fix unrelated-failure fixture, ignore generated lib - gsd-executor.md: compress the #3770 fail-fast rule to one line (49149 B < 49152 cap; line-count parity keeps the #2751 PROSE_ALLOWLIST line 816 valid) - tests: the row-6 fixture used String.replace (first-occurrence), so the `not ok` line still named the target test and the classifier was right to accept it; replaceAll makes the failure genuinely unrelated - eslint.config.mjs: ignore tsc-generated bin/lib/tdd-red-evidence.cjs (lint the src/*.cts source, per ADR-457 migration rule) Emitted-Drift-Ack-Growth: gsd-executor.md — the #3770 fail-fast rule now requires intentional RED evidence (check tdd-red-evidence) before GREEN; +172 bytes, kept under the LARGE cap and on one line * chore(3770): add changeset * chore(3770): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
2f4f7538e9 | fix(#4264): wire both TDD dispatch backends to phase.tdd-applicable (#4284) | ||
|
|
580059251a |
fix(#4040): route partially-created .planning to initialization recovery (#4283)
* test(#4040): add failing-first regression tests for partial-init routing Red: init.progress/init.resume/init.new-project payloads carry no partial-init discriminator, and progress.md/resume-project.md/ new-project.md route an interrupted bootstrap (.planning/PROJECT.md + config.json only) to Route F / STATE reconstruction / a hard error. * fix(#4040): route partially-created .planning to initialization recovery A bootstrap interrupted after .planning/PROJECT.md (but before REQUIREMENTS.md/ROADMAP.md/STATE.md) was mis-routed three ways: progress.md read it as between-milestones (Route F) or 'no planning structure', resume-project.md offered STATE.md reconstruction, and new-project.md errored 'already initialized' — a routing loop with no recovery exit. Add a shared buildInitCompletenessFields discriminator (planning_exists / requirements_exists / milestones_exists / init_incomplete) to the init.progress, init.resume and init.new-project payloads, and branch on init_incomplete in progress.md, resume-project.md and new-project.md BEFORE the legacy branches. MILESTONES.md presence excludes the archival between-milestones state, so Route F and the STATE-reconstruction path keep working. Emitted-Drift-Ack-Growth: progress.md — deliberate #4040 growth: new init_incomplete recovery branch (routing text + guard on the no-planning and Route F branches) added ahead of the legacy init_context routes. Emitted-Drift-Ack-Growth: resume-project.md — deliberate #4040 growth: new init_incomplete branch routing an interrupted bootstrap to initialization recovery before the STATE.md-reconstruction branch. Emitted-Drift-Ack-Growth: new-project.md — deliberate #4040 growth: project_exists gate split on init_incomplete so a partial bootstrap resumes initialization instead of erroring. * chore(#4040): add changeset fragment * chore(#4040): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
b1b7cabfb5 |
docs(#4123): add gsd-qoder EoS registry entry (#4278)
* docs(registries): add gsd-qoder EoS entry Adds one `type: "eos"` entry for a Qoder host integration and regenerates docs/registries/eos-registry.md. Qoder is Alibaba's AI coding product family (Qoder CLI and Qoder Desktop). The integration depends on @opengsd/gsd-core, negotiates the ADR-1239 host-integration handshake, and projects GSD's agents, skills, and hook scripts into the Qoder config directory (~/.qoder, or ~/.qoder-cn for the China edition), merging GSD's lifecycle hooks into settings.json. Every axis is sourced from Qoder's own docs per the never-infer rule. `dispatch.isolation` is `none`: Qoder documents `isolation: worktree` as a frontmatter-declared, per-agent-definition property, and GSD's two isolation negotiation models both assume a per-dispatch injection point Qoder does not expose. Re-homes the Qoder runtime work from #860 / PR #2005, which was closed in favor of the EoS path. Closes #4123 * docs(#4123): backfill changeset pr field |
||
|
|
75ee7b0214 |
enhance(#4273): add phase.tdd-applicable single-owner predicate (#4277)
* enhance(#4273): add phase.tdd-applicable single-owner predicate One query verb computes TDD-applicability for a plan (CLI flag, plan type: tdd frontmatter, a task's tdd="true" attribute, or the workflow.tdd_mode config default), mirroring phase.mvp-mode's precedence-cascade shape. Foundation for epic #4272 Phase 2, which wires both dispatch backends to consume it instead of restating the predicate independently. Also fixes workflow.tdd_mode, workflow.research, and workflow.nyquist_validation, which never reached cmdInitExecutePhase/cmdInitPlanPhase/cmdInitDebug/cmdInitNewMilestone because loadConfig() never populates config.workflow — a dead accessor found while wiring this verb's own config read, fixed inline per the no-defer rule rather than left alongside it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4273): document phase.tdd-applicable's FEATURES.md entry Add a docs/features/ fragment for the new phase.tdd-applicable query verb and regenerate docs/FEATURES.md. docs/COMMANDS.md is left untouched: it documents /gsd-* slash commands only, and the sibling verb phase.tdd-applicable mirrors (phase.mvp-mode) has no formal CLI reference entry anywhere in docs/ either -- only inline prose mentions in docs/reference/workflow-fragments.md -- so there is no COMMANDS.md precedent to extend. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use PHASE_NOT_FOUND reason code, remove try/finally from tests Two orthogonal code reviews flagged a mistyped error reason and a CONTRIBUTING.md-banned try/finally pattern in the phase.tdd-applicable change; both are corrected here. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): stop whitelisting capability-owned config keys centrally workflow.tdd_mode, workflow.research, and workflow.nyquist_validation are each already owned by their own first-party capability's federated config schema (the tdd/research/nyquist capabilities declare them under their own capability.json `config`), resolved via isCapabilityConfigKey. Adding them to gsd-core/bin/shared/config-schema.manifest.json's central validKeys, as the prior commit in this branch did (mirroring workflow.mvp_mode, which genuinely is central-only), declares the same key in two places at once. That collision breaks capability-loader.cts's loadRegistry composition: gsd-test caught this as 84-85 unrelated failures across capability-cli/capability-command-dispatch/capability-lifecycle test files, every one showing "unknown capability: <id>" for a freshly-installed third-party capability that should have resolved fine. Verified directly (not asserted): reverting only this file, keeping the config-loader.cts tdd_mode/research/nyquist_validation flattening and the init.cts call-site fixes from the prior commit, and re-running the exact capability install + capability set repro from tests/capability-cli.test.cjs's "issue-2322" test locally reproduces the failure with the whitelist entries present and clears it without them. loadConfig() still surfaces all three flattened values correctly with no central whitelist entry (confirmed directly against the compiled module) — the whitelist additions were never required for the #4273 fix to work; they were an incorrect over-application of the mvp_mode precedent to keys that aren't central. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4273): use getNested for tdd_mode (no legacy top-level fallback), allowlist new test file Both fixes address defects found by a gsd-test bench run: tdd_mode routed through get() invented an undocumented top-level alias that silently outranked the canonical workflow.tdd_mode key, and the new phase-tdd-applicable test file was missing from the file-count allowlist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4273): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f4bf449296 |
fix(#3850): surface gaps_found VERIFICATION files in audit-uat (#3879)
* fix(#3850): surface gaps_found VERIFICATION files in audit-uat cmdAuditUat admits `human_needed` OR `gaps_found`, but parseVerificationItems had a body only for the first and returned an empty array for the second — standing on a comment deferring to `plan-phase --gaps`, a different command audit-uat never reaches. Since cmdAuditUat pushes a file into `results` only when `items.length > 0`, a `gaps_found` report did not under-report: it vanished, taking its phase's `by_phase` row with it, so a clean-looking total gave the reader no cue anything was skipped. Eligibility now has one owner (the caller) and parseVerificationItems reports what the file says. The closed-entry filter could not be built on extractFrontmatter: its array-item parser keeps only each `- ` entry's FIRST line and has no notion of nested key/value objects, so an entry's `status:`/ `resolution:` siblings never reach its output and a closed entry is indistinguishable from an open one downstream. Rather than grow a competing object-list parser — or change extractFrontmatter, whose blast radius is every frontmatter consumer in the repo — this reads the raw segment BEFORE the flattening, via the existing anchored sliceTopLevelFrontmatterSegments, and hands it to the `## Gaps` machinery that already parses exactly this `- `-opened, indentation- continued shape. The human_needed path is byte-for-byte unchanged: same reader, same display names, same numbering, no resolved-entry filtering — pinned by a test and verified by identical CLI output on base and head. parseGapsItems keeps its narrower `status: resolved` rule so no *-UAT.md behaviour moves. Closes #3850 * chore(#3850): backfill changeset pr number for #3879 * fix(#3850): one parse per entry, one fence parser, one resolved-entry rule Adversarial review on #3879: B1, B2, M3, m5, m8 and n9. B1 — `sliceFrontmatterArrayEntries` hand-rolled a second frontmatter fence regex, which re-asserted the byte-0 rule #2977 removed: a BOM'd file (PowerShell 5.1 `>`/`Out-File` writes one by default) sliced nothing, so a `gaps_found` report vanished from the audit exactly as it did before this fix — this issue's own symptom, on a platform the repo already has a named defect class for. `extractFrontmatter`'s BOM+fence logic is now factored out as `frontmatterRegion` and shared. One fence parser, not two. B2 — the resolved-entry skip paired two DIFFERENT parsers by array index: `parseYamlRegion` is indent-blind, `splitGapsEntries` is indent-anchored. A block sequence written at its key's indent — ordinary, legal YAML — makes them disagree about entry count, and from the first disagreement every index names a different entry, so an OPEN entry inherits a CLOSED one's resolution and is silently dropped. That is the defect this PR exists to fix, reintroduced inside the fix. Display name and sibling fields now come from ONE parse of the raw slice; `frontmatterEntryDisplayName` applies `parseQuotedScalar` exactly as `parseYamlRegion` does, so the string is byte-identical to what `extractFrontmatter` produced. The flattened array remains the #2286 GATE, but is no longer the source of items. `sliceFrontmatterArrayEntries` also takes the LAST duplicate key, matching `parseYamlRegion`'s last-wins assignment. M3 — `frontmatterEntryToUatItem` is the single entry->UatItem mapper both readers use, rather than two copies differing only in `result`. m8 — closed entries are skipped on BOTH statuses. The earlier asymmetry cited an acceptance criterion #3850 does not contain: the issue has no AC section, and its suggested fix (2) states the skip unconditionally, naming a file with 14 of 16 entries resolved. That file is `human_needed`, so the asymmetry left the reporter's own scenario over-reporting by 14. m5 — `sliceTopLevelFrontmatterSegments`' contract doc names both consumers and says the column-0 boundary rule is now a cross-module contract. n9 — the vestigial bare block is gone and its body de-indented. Tests: the B1 BOM case, B2's nested-sequence and bare-bullet repros, a CRLF fixture (M4 — it survived by accident, now pinned) and the unified skip rule. Fail-first verified by running the new tests against the pre-fix build: the BOM, nested-sequence and unified-skip cases are red there. * fix(#3850): read the entries as objects, not as re-parsed display text Rebased onto `next`, which changed the ground this fix stood on. ADR-3473 §8.1 (#3881) replaced the hand-rolled frontmatter scanner with the vendored js-yaml: `parseQuotedScalar` and `parseYamlRegion` no longer exist, and an object entry now flattens to `test: A, resolution: R` rather than to its first line. The original mechanism existed ONLY to work around that lossy first-line flattening — it sliced the raw frontmatter segment and re-parsed each entry by hand so a `resolution:` sibling was visible at all. With a real parser upstream that workaround is obsolete, so it is deleted rather than repaired: `sliceFrontmatterArrayEntries`, `frontmatterEntryDisplayName`, the `splitGapsEntries`/`extractGapEntryFields` reuse and the second fence regex are all gone. `frontmatter.cts` instead exposes `frontmatterObjectListEntries(content, key)` — the same parse `extractFrontmatter` runs (same BOM strip, same byte-0 fence, same anchor/alias and sentinel guards, same ambiguous-colon repair), stopping one step before the display flattening. `flattenObjectListItem` is exposed alongside it so a caller deriving a display name produces the byte-identical string `extractFrontmatter` would have. That collapses the review's blockers into properties of the parse rather than things this fix has to get right: - B1 (BOM) — shares `extractFrontmatter`'s strip; verified through the CLI. - B2 (index pairing) — there is no second reader. Display name and sibling fields come from one object. - M3 (duplicate mapper) — one `frontmatterEntryToUatItem` for both readers. - M4 (CRLF) — js-yaml's, not ours; verified through the CLI. Also confirmed on the rebased base, per review: #3850 still reproduces on `next` after #3707 landed (`total_files: 0`, `total_items: 0` on a `gaps_found` fixture), so this PR is still doing work #3707 did not do. Nothing was dropped as redundant. One behaviour note: `entryField` returns a present value verbatim and treats only whitespace-only as absent. Trimming would rewrite an author's `truth:` on its way to becoming the display name. * fix(#3850): keep every frontmatter list entry at its own row Review round 3's Blocker. `frontmatterObjectListEntries` filtered its result to objects, and filtering COMPACTS: `parseHumanVerificationItems` then numbered the survivors by their position in the compacted array. On a list mixing object and non-object entries the non-object rows disappeared outright and the rest were renumbered — #3850's own vanishing-row defect, reached through entry SHAPE instead of file STATUS. Base never had it: it walked the display array, so every row surfaced at its own position. Renamed to `frontmatterListEntries` and it no longer filters (the name now matches what it returns). Deciding what a non-object entry MEANS is a caller's judgement; dropping it is nobody's. Both readers now walk the DISPLAY array — one element per row, the array #2286 already gates on — and consult the parsed array only for "does this entry carry a closure field?". `parsedEntriesFor` owns that pairing and checks the two lengths agree before trusting an index; all-null is the correct degradation, since over-reporting a closed row is recoverable and closing the wrong one is not. Names stay byte-identical to base for every entry shape, including a nested sequence (`[nested]`, not `["nested"]`). Same class closed in the gaps reader: a non-object `gaps:` entry surfaced nothing at all and now surfaces as `unknown`, which is this module's documented fail-safe direction (`parseGapsItems`) on a false-negative bug. Also restores the shared fence parser round 2 accepted. The ADR-3473 rebase dropped `frontmatterRegion` and left the BOM strip and byte-0 fence rule inlined twice; `extractFrontmatter` now routes through it, so "one fence parser" is enforced rather than asserted in a comment. Minors: `frontmatterEntryToUatItem`'s dead `forcedResult` option deleted and its "shared by both readers" comment corrected — it has one call site, and the two readers differ deliberately, each mirroring its own established sibling (`parseGapsItems` vs #2286). Documented at the divergence. Tests: `B2` asserted a name substring, so it passed while the row was mis-numbered and would have passed through outright loss; it now asserts positions and count. B2b pins the reviewer's 6-entry mixed fixture verbatim, B2c the survivors' file positions across skipped rows, B2d the gaps reader. All four fail-first against the reviewed head; 332/332 green with the fix. * fix(#3850): make status authoritative, and let the two gaps readers agree Round 4 review, all five findings. Major. `isFrontmatterEntryResolved` treated a non-empty `resolution:` as closure regardless of `status:`, so `status: failed` + `resolution: "attempted retry, still failing"` vanished from the report — the silently-vanishing-item defect #3850 exists to close, reached by field combination instead of file status. Closure is now per key, because the two keys have different conventions and one rule cannot serve both: `gaps:` `status: resolved` only, byte-identical to the rule `parseGapsItems` applies to a `## Gaps` markdown section, so one authored entry cannot read closed in one reader and open in the other. `human_verification:` a bare `resolution:` still closes, since that is how verifier-written entries record it — but a readable `status:` that contradicts it wins. A single unified rule was the first draft and is wrong: it closes a frontmatter `gaps:` entry carrying `resolution:` and no `status:`, which `parseGapsItems` surfaces, and `parseVerificationGapsItems`' own docstring claims it mirrors that reader's fail-safe status handling. The contradiction guard is not a judgment call about YAML. It is the rule this codebase already applies to the same field pair: `validateResolution` (probe-core.cts) rejects a populated `resolution:` on a non-resolved status outright — "a populated payload is an authoring mistake ... Reject it so the mistake surfaces." A reporter cannot throw, so it surfaces the item. Minor 1. Direct unit tests for `frontmatterListEntries` and `flattenObjectListItem` in `tests/frontmatter.unit.test.cjs`, the file that historically co-changes with `frontmatter.cts`. They were reachable only through `uat.cts`' readers before. Minor 2. `parsedEntriesFor`'s degrade-to-all-null branch is asserted directly. Verified unreachable through content rather than assumed: both readers enter through `frontmatterRegion`, `extractFrontmatter`'s only extra argument gates a warning, and `normalizeParsedValue`'s `value.map` is 1:1. It is a drift alarm for a future edit to either parser, so the helper is exported for tests rather than left as the one unpinned branch. Minor 3. The vestigial `const skipResolved = true` and its dead conditional are gone. Minor 4. `frontmatterEntryToUatItem` no longer reads `test:`. A `gaps:` entry has no `test:` in its vocabulary — the template's entries carry truth/status/reason/artifacts/missing — so it was speculative support for a field the shape does not have, and it collided with the 1..N row numbers `parseHumanVerificationItems` assigns by array position. Not reading it makes the collision impossible; an offset would have rewritten an authored value, against `entryField`'s verbatim contract. Docs, changeset and the dispatcher docstring all stated the unconditional rule and are corrected — three prior rounds here were comment/code drift. Fail-first proven: restoring the universal rule reddens all three new unit tests and both rewritten properties. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
585a8b7f1b |
fix(#3747): correct antigravity matrix evidence and pin the CLI-only skills install path (#4274)
* test(#3747): fail-first regression — matrix must not cite configHome skills path for antigravity * fix(#3747): correct disproven antigravity stateIO evidence; pin CLI-only probe branch install path * fix(#3747): scope doc evidence claim to skills discovery per adversarial review * chore(#3747): add changeset * chore(#3747): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
75bad7aedd |
test(#3936): tighten quick researcher regression coverage (#4169)
* test(#3936): tighten quick researcher regression coverage * test(#3936): restore adjacent quick dispatch coverage Assert the default researcher model and bind the executor persona check to its Agent payload. * test(#3936): make parse-list assertion wrap-safe Bound the researcher-model check to the full parse paragraph so formatting-only line wraps do not fail the regression test. * test(#3936): isolate Windows model defaults --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7eefad6f92 |
docs(#4260): record npm-audit retry/backoff enhancement as out-of-scope (#4271)
Co-authored-by: sim <sim@local> |
||
|
|
18e5cfff8a |
fix(#4250, #4260): distinguish a timed-out npm audit from a JSON parse failure, retry with backoff (#4251)
* fix(#4250): distinguish a timed-out npm audit from a JSON parse failure npm-audit-baseline.cjs's runPackageLockAudit, and the near-identical auditProductionVulns helper in npm-integrity-gate.test.cjs, both grabbed e.stdout whenever an npm audit child process exited non-zero -- without checking whether the process was actually killed by its 180s timeout. A timeout-killed process's stdout is truncated mid-write, not complete JSON, so JSON.parse threw a misleading "Unexpected end of JSON input" instead of naming npm's registry timeout as the real cause. Root-caused live during a CI investigation: npm's own status page reported degraded service, and the registry's bulk-advisories endpoint was returning 503/hanging, causing npm audit to sit until the timeout fired. Adds a shared isTimeoutKill(error) predicate (checks execFileSync's documented killed/signal fields) and checks it first in both catch blocks, throwing a clear, actionable error before ever reaching JSON.parse. The pre-existing "non-zero exit with complete JSON" recovery path is unchanged and still covered by regression tests. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4250): add changeset for npm-audit timeout fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4250): share the timeout-kill error message and cover auditProductionVulns Two independent review passes (standards + spec) on the first commit found real gaps: the timeout-kill error message was duplicated verbatim between runPackageLockAudit and the near-identical auditProductionVulns helper in tests/npm-integrity-gate.test.cjs (this repo's own Generative Fix Divergence anti-pattern -- shared logic across parallel surfaces with no parity check), and auditProductionVulns picked up the same production fix with zero test coverage of its own. Extracts buildTimeoutKillError(cwd), used by both callers so the message cannot independently drift. Gives auditProductionVulns the same injectable execFileSyncImpl seam runPackageLockAudit already had, and adds the matching regression tests (timeout-kill throws the clear error; the pre-existing non-zero-exit-with-complete-JSON path still recovers correctly). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4250): backfill changeset PR number to #4251 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * diag(#4250): surface captured stderr in the timeout-kill error The killed child process's stderr is buffered in-memory by execFileSync and attached to the thrown error, but nothing surfaced it -- the timeout message named the timeout but discarded the one piece of data that could show WHY npm was still running when it fired (DNS stall, TLS handshake stall, a registry-side retry loop, all look identical without it). buildTimeoutKillError now takes the killed error and includes its stderr (or an explicit 'no stderr was captured' note) in the message. This is a diagnostic improvement for the next CI occurrence, not a behavior fix -- local reproduction has directly ruled out npm version (installed the exact CI-bundled 11.17.0 and ran it against this repo: 0.49s, clean), general npm registry reachability (0.4-1.4s locally, repeatedly), and npm ci speed (2m, succeeded) as explanations for the 180s CI hangs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4260): bounded retry with backoff for npm audit calls, finish the extraction The audit backend has real, independent latency variance from the rest of the npm registry -- measured (see #4260): a bulk-advisories POST took 43.41s vs 0.20s for a plain registry fetch on the same host, and the same endpoint returned no response at all (000) twice in the same window, while status.npmjs.org reported fully operational throughout. Against that, runPackageLockAudit and its near-duplicate auditProductionVulns each made exactly one attempt with no retry -- any single bad moment failed a REQUIRED CI gate on a transport hiccup, not a real advisory. Replaces the single 180s attempt with runNpmAuditWithRetry: up to 3 attempts at 60s each (comfortably above the worst measured working latency) with exponential backoff between them. Only a confirmed timeout-kill is retried; a genuine non-timeout failure still fails immediately, and exhausting all attempts still fails the gate -- per #4260's own caveat, silently disarming a required security check on a transport error is worse than occasionally re-running CI. Also finishes the extraction #4260 flagged as stopped halfway: auditProductionVulns (tests/npm-integrity-gate.test.cjs) duplicated runPackageLockAudit's entire candidate loop, recovery branch, and timeout classification, differing only in npm args and precondition check. It is now a thin wrapper delegating to the newly-exported runInstalledTreeAudit, which shares runNpmAuditWithRetry with runPackageLockAudit -- one implementation instead of two that could independently drift. buildTimeoutKillError now reports attempt count and still surfaces captured stderr from the last kill. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#4260): update changeset for retry/backoff scope Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4260): budget for two sequential retry-audit calls, close coverage gaps Two review passes on the retry/backoff commit found real gaps: - TEST_TIMEOUT_MS budgeted only one retry-audit call's worst case (210s), but checkTreeAgainstBaseline makes two sequential calls (HEAD tree via auditProductionVulns, baseline tree via runPackageLockAudit) -- combined worst case is ~372s. If both genuinely exhausted retries, node:test's own timeout would fire first and mask buildTimeoutKillError's clear message, undercutting #4250's own fix in that edge case. Recomputed using the same backoff formula the production code uses, so it can't independently drift. - buildTimeoutKillError's default-attempts(1) singular-phrasing branch had zero direct test coverage (nothing calls it with a single attempt anymore) -- a real mutation-testing risk. Added direct tests for both phrasing branches plus the no-error-object case. - runInstalledTreeAudit's null-guard skip paths (missing package.json, missing node_modules) had no tests, unlike runPackageLockAudit's matching paths. Added for parity. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
97ce61dee2 |
fix(#3990): state the RED/GREEN/REFACTOR cycle once, embed tdd.md conditionally (#4228)
* test(#3990): the RED/GREEN/REFACTOR cycle is stated once, embedded conditionally * fix(#3990): state the cycle once — pointers in consumers, conditional tdd.md embeds Emitted-Drift-Ack-Growth: execute-phase.md — #3990 conditions the tdd.md embed on the dispatch being TDD * chore(#3990): changeset for the single-statement TDD cycle * chore(#3990): backfill changeset pr number * fix(#4228): linear cycle check — the lazy-span regex pinned a Windows core for the whole job cap * test(#3990): allowlist pin tracks the rebased line * fix(tests): npm-integrity gate names an empty audit output explicitly — empty stdout crashed the parse as a bare SyntaxError * fix: name an empty npm-audit stdout explicitly — it crashed the parse as a bare SyntaxError Observed on CI (several branches, all lanes): spawnSync npm ETIMEDOUT with empty stdout; the empty string survived the recovery path and surfaced as 'SyntaxError: Unexpected end of JSON input', hiding the captured error. The recovery path now requires non-empty stdout, and an empty result throws with the captured stdout/stderr/message so the actual error is on the record. Root cause of the ETIMEDOUT itself is NOT diagnosed here — this change only stops masking it. --------- Co-authored-by: sim <sim@local> |
||
|
|
590edec7a7 |
fix(#3956): require positive evidence for verify artifacts/key-links pass (#4004)
* fix(#3956): require positive evidence for verify artifacts/key-links pass An all-string or path-less must_haves.artifacts / key_links block is item-by-item skipped, leaving zero checked results, yet the pass verdict was computed as `passed === results.length` (0 === 0), so all_passed / all_verified read true with status valid and exit 0: a silent false GREEN over zero acceptance evidence. Add a positive-evidence floor (results.length > 0) to both verdicts, mirroring the no-vacuous-pass rule at src/uat-predicate.cts. A well-formed block, the fully-empty-block error, the parser's string tolerance, and key-links pending (#1202) semantics are all unchanged. Governing: ADR-3473 section 8 / 37C (absence, emptiness and failure must not encode as success) and Decision 3 (failure is a value). * chore(#3956): add changeset for verify vacuous-pass fix * test(#3956): add mixed-block coverage and correct the key-links vacuous-pass comment Addresses review on #4004: - Correct the cmdVerifyKeyLinks positive-evidence-floor comment: only bare-string items are continue-skipped; a from:-less object is NOT skipped (it falls through to a verified:false hard failure), so it was never part of the vacuous-pass surface. The prior comment overclaimed symmetry with the artifacts side. - Add a mixed-block regression test per verb (one bare-string prose bullet + one well-formed entry): the string is skipped, results.length === 1 > 0, and the verdict follows the single real entry — pinning that the floor does not over-reject a partial block. - Tighten the changeset wording to match (all-bare-string, not "no path:/from: key"). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a788afb120 |
fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives (#4021)
* fix(#4010): confine stateReplaceField to same-line whitespace so an empty field's following line survives stateReplaceField's bold and plain patterns used `\s*` for the label-to-value gap, which matches the newline after an empty field; `(.*)` then captured the following line and the rebuild discarded it -- silent STATE.md data loss on any `state update` against an empty body field (Status:, Stopped at:, Paused at:), with exit 0 and no warning. Confine the gap to same-line whitespace (`[ \t]*`), mirroring the already-correct read side (stateExtractField, src/state-document.cts:404/:409), and pin the label-value separator to a single space when the label line had none, so an empty field yields `**Status:** value` rather than a glued `**Status:**value`. Non-empty and pipe-table replacements are byte-identical to prior behaviour. ADR-3180 §7.7 makes stateExtractField the same-line-confined owner; this aligns the writer to it. Regression test fails before / passes after and covers bold and plain shapes, LF and CRLF, the non-empty byte-identity guard, and an end-to-end transitionCore characterization at the consumer (ADR-3180 Decision 4(c)). * chore(#4010): add changeset for the stateReplaceField empty-field fix * test(#4010): add boundary and property coverage; scope the changeset's unchanged claim Addresses review on #4021: - Add boundary tests for the shapes the example tests missed: an empty field at end-of-document (no following line, bold + plain), two consecutive empty fields (only the target is filled, the other empty field's line survives), and an empty new value on an empty field (joinFieldReplacement synthesizes no dangling separator and the following line is preserved). - Add a fast-check property over the bold/plain branches and joinFieldReplacement: for any field name, any values (empty fields included), and any new value, replacing one field changes only its own line and never the total line count — the invariant #4010 violated, now guarded directly. - Scope the changeset's "unchanged" claim to ordinary space/tab separators (an exotic vertical-tab/form-feed separator, which no GSD template emits, now normalises to a single space). * test(#4010): pin glued-separator non-empty field, scope joinFieldReplacement JSDoc Round-3 review carried forward a Minor finding: joinFieldReplacement's JSDoc still claimed non-empty replacements are unconditionally "byte-identical to prior behaviour", but a non-empty field written with no label-to-value separator (**Status:**value) gains a single inserted space under the narrowed [ \t]* gap. Round 2 scoped only the changeset prose; the source JSDoc was left making the false unconditional claim. - Scope the JSDoc's byte-identity claim to ordinary space/tab separators and name the no-separator normalization as the one intentional exception. - Add a test pinning the glued-separator case (**Status:**Planning): exactly one space inserted, following line survives, not byte-identical. Emitted .cjs is gitignored (class-1), so no emitted-drift-ack applies. build:lib clean; 74/74 state-document tests pass. Claude-Session: https://claude.ai/code/session_01Mzmut6aeqZ1APfUBAkBZTR --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
02c6955162 |
chore(#4241): add merge_group trigger to test.yml (#4242)
* chore(ci): add merge_group trigger to test.yml (#4241) GitHub's merge queue fires the `merge_group` event for the temporary merge-group commit it creates when a PR is added to the queue - not `pull_request` or `push`. Without this trigger, `required-tests` (the registered "Required tests" branch-protection check) never schedules for a queued PR, permanently stalling the queue on a check that never runs. This is workflow-side prerequisite wiring only; enabling the merge queue itself is a separate manual branch-protection step. * fix(#4241): pin AUDIT_BASELINE_REF for merge_group events too Code review on this branch caught that AUDIT_BASELINE_REF's ternary only branched on pull_request/push, so a merge_group run silently fell through to '' -- scripts/npm-audit-baseline.cjs's resolveBaselineRef() documents its origin/next live-tip fallback as unreachable from CI specifically because AUDIT_BASELINE_REF is "always set by test.yml". Reopens the exact race #4196 fixed, but only for merge-queue runs. Extends all three AUDIT_BASELINE_REF pins (test, test-inert, test-full) to also branch on merge_group, using github.event.merge_group.base_sha (confirmed against GitHub's own webhook payload schema: "the SHA of the merge group's parent commit") -- the base tip the temporary merge-group commit was built against. Adds a regression test asserting every AUDIT_BASELINE_REF pin branches on merge_group with the correct field. --------- Co-authored-by: sim <sim@local> |
||
|
|
1fe85cd43e |
chore(#4244): ESLint rules for the #4220 Windows dirname-walk / TMPDIR-triad bug class (#4246)
* fix(#4244): repoint TEMP/TMP alongside TMPDIR and fix the sweepProtectSet fixed-point walk Repo-wide sweep (ahead of adding lint rules for these exact bug classes) found both incident patterns still live and unfixed on `next`: - scripts/run-tests.cjs's sweepProtectSet walk stopped on `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel. win32 dirname('D:\') is a fixed point (length 3, never satisfies `> 1`... wait, it does satisfy length>1), so a selected file living outside runTempRoot (the common case) spins the walk forever on Windows. Extracted a pure, exported computeSweepProtectSet helper that terminates on dirname(cur) === cur instead, with in-process RuleTester-style coverage for both win32 and posix paths. - tests/run-tests-temp-root.test.cjs's own #4020 regression test set only TMPDIR on its runNode(...) child env. Node's os.tmpdir() never reads TMPDIR on Windows (only TEMP, then TMP), so the redirect silently no-oped there — masked because Windows CI died in the dirname-walk hang above before ever reaching this test. - tests/config-schema.property.test.cjs's fallow config-set test had the same TMPDIR-only pattern, direct process.env assignment this time, restored in its own finally block. Origin: #4220 and its shared root cause #4020. * feat(#4244): require-full-tmpdir-triad and no-unbounded-dirname-walk ESLint rules Two custom local ESLint rules catch the #4220 / #4020 Windows CI hang bug class at author time, joining the ADR-1703 DEFECT.WINDOWS-TEST-PORTABILITY catalog. Neither eslint-plugin-unicorn nor eslint-plugin-n has a rule for either shape. - local/require-full-tmpdir-triad: flags a TMPDIR environment override (direct process.env.TMPDIR assignment, or a TMPDIR property in a spawn-like call's env: object literal) not accompanied by TEMP and TMP in the same scope. Node's os.tmpdir() never reads TMPDIR on Windows. Registered on tests/**/*.cjs, matching the require-userprofile-with-home precedent. - local/no-unbounded-dirname-walk: flags a while/do-while loop reassigning from dirname() with no fixed-point termination guard (dirname(cur) !== cur, or path.parse(cur).root). path.dirname() is a no-op at the platform root, but the value differs by platform (win32 'D:\' is length 3, posix '/' is length 1), so a POSIX-shaped length/equality bound never fires on Windows. Registered on BOTH tests/**/*.cjs and scripts/**/*.cjs — the real #4020 bug lived in scripts/run-tests.cjs, not tests/. Both rules join the zero-escape-hatch discipline already established for this catalog (no bespoke comment marker; PROTECTED_RULES in tests/portability-rule-disable-ban.test.cjs independently bans eslint-disable of either). ADR-1703 and its two companion contributing docs get an amendment documenting the mechanism, code examples, and the repo-wide sweep (three live instances found and fixed in the prior commit; no others found). CI test-scope selection updated so an edit to either rule or to scripts/run-tests.cjs re-runs the right suites. * fix(#4244): no-unbounded-dirname-walk must analyze a single-condition loop test too checkWhile bailed out early unless node.test was a LogicalExpression, so a single-condition loop -- while (cur !== root) { cur = dirname(cur); } -- was silently skipped and never reported. That is the EXACT minimal shape of the original #4020/#4220 bug, and it is literally the shape used by this rule's own shipped RuleTester fixtures (the "equality-only bound" invalid cases), which were failing (0 errors reported, 1 expected) until this fix -- confirmed by running RuleTester directly against both fixtures, not just via a passing test-runner exit code. The conjunct-collection helper already handled a non-LogicalExpression test correctly (it pushes a single node as the sole conjunct); only the early-return gate needed to stop requiring a compound && / || test. Verified: RuleTester run directly against both previously-broken fixtures plus two new sanity cases (a guarded single-condition loop stays valid; an unrelated single-condition loop stays silent), and a fresh `npx eslint .` across the whole repo remains clean (no other single-condition dirname-walk shape exists in the tree). * fix(#4244): require-full-tmpdir-triad must recognize a destructured child_process call isSpawnLikeCallee only recognized a MemberExpression callee (child_process.spawnSync(...)) or a bare identifier in ENV_LOCAL_HELPER_NAMES (runNode). A destructured import called bare -- const { spawnSync } = require('child_process'); spawnSync(...) -- has an Identifier callee named "spawnSync", which matched neither branch, so the whole env-literal check was skipped. gsd-test caught this: both "invalid: child_process.spawnSync with TMPDIR-only env" cases in tests/require-full-tmpdir-triad.rule.test.cjs were failing (0 errors reported, 1 expected). Widened the bare-identifier branch to also match any of the known ENV_CHILD_PROCESS_METHODS names, matched by name only -- the same lightweight convention this repo's other eslint-rules/*.cjs use (e.g. no-hardcoded-tmp.cjs's isFsMethodCall), not full import data-flow tracing. Verified: RuleTester run directly against all 11 cases in tests/require-full-tmpdir-triad.rule.test.cjs (not just the two that were failing), all pass; a fresh npx eslint . and npm run lint:ci across the whole repo remain clean. * fix(#4244): correct a stale escape-hatch reference in a test comment The comment on the "length comparison against another expression's length" case referenced a "// allow-dirname-walk marker" that doesn't exist -- the rule has zero comment-based escape hatches by design (ADR-1703), and an earlier draft's marker mechanism was removed before this branch's first commit. Spec-axis review caught the stale reference. No behavior change; comment-only. * chore(#4244): backfill changeset PR number (pr:0 -> pr:4246) --------- Co-authored-by: sim <sim@local> |
||
|
|
456136659d |
fix(#4220): terminate the Windows temp-sweep ancestor walk (and two bugs it unmasked) (#4245)
* fix(#4220): terminate the temp-sweep ancestor walk with a fixed-point check scripts/run-tests.cjs's sweepProtectSet block walked each selected test file's ancestor directories, stopping on `cur !== runTempRoot && cur.length > 1` — a POSIX-only sentinel. path.posix.dirname('/') === '/' (length 1) correctly stops, but path.win32.dirname('C:\\') === 'C:\\' (length 3) never satisfies the length check, so the walk spun forever on Windows whenever a selected file lived outside runTempRoot (the common case). This has hung every Windows CI shard since #4207. Extract the walk into a pure, exported computeSweepProtectSet(selected, runTempRoot, dirnameImpl) helper and replace the length sentinel with a fixed-point check (stop when dirnameImpl(cur) === cur), which terminates correctly on POSIX, Windows drive roots, and UNC roots alike with no platform branch. * fix(#4220): repoint TEMP/TMP alongside TMPDIR in run-tests-temp-root test child env Node's os.tmpdir() on Windows never reads TMPDIR, only TEMP/TMP. The test's runNode child-process env override only set TMPDIR, so on a real Windows runner nested inside a run-tests invocation the child inherited the outer process's already-repointed TEMP/TMP and its mkdtempSync(os.tmpdir()) landed under the outer run's temp root instead of the test's intended `outer` directory. This was masked on gsd-test's benches and locally because Windows CI always died in the #4220 infinite loop before reaching this test. * fix(#4220): stop the ancestor walk from protecting the filesystem root itself computeSweepProtectSet added `cur` to the protect set before checking whether dirname(cur) === cur, so on the terminating iteration it protected the filesystem root (posix `/`, and analogously a win32 drive root) instead of stopping before adding it. Caught by the existing posix-parity regression assertion (`!protectSet.has('/')`) on the linux-node24 gsd-test bench. Reorder to compute the parent and check the fixed point before adding. * fix(#4220): backfill changeset pr number to 4245 --------- Co-authored-by: sim <sim@local> |
||
|
|
515191f07d |
feat(#3677): quick-batch hardening and acceptance (#4240)
* chore(#3677): checkpoint design artifacts (gitignored, dev-only) * test(#3677): add failing regression test for the crash-window duplicate-dispatch gap (RED) Independently re-traces resume-mode.md/planner-wave.md/worktree-dispatch.md/ merge-wave.md and src/quick-batch.cts's resumeBatch (lines 894-899) and confirms the prior research pass's Open Question 1: a coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) leaves BATCH.json at "pending" with no STATE.md row yet (only written in Step 9), so --resume's eligibility re-derivation would dispatch a second executor into a new worktree for the same item, orphaning the first. This test asserts worktree-dispatch.md's Step 6 excludes an item whose SUMMARY.md already exists from the spawn set, mirroring planner-wave.md's existing PLAN.md-existence check one layer earlier. Fails against the current worktree-dispatch.md, which has no such guard. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the full trace and fix-location rationale. * fix(#3677): guard worktree-dispatch.md against re-dispatching an already-executed item (GREEN) worktree-dispatch.md's Step 6 re-derives eligibility every dispatch round via the same quick-batch resume call resume-mode.md uses, but had no check for "did this item already finish executing" the way planner-wave.md already checks "did this item already get planned" (PLAN.md existence) before re-planning. A coordinator crash between Step 6 (executor commits, SUMMARY.md written) and Step 7 (merge) left the item eligible for a second dispatch on --resume, orphaning the first worktree's real, already- committed work and silently losing it once the second executor's SUMMARY.md write clobbered the first at the same item_dir path. Adds a SUMMARY.md-existence exclusion before spawn-plan is computed, symmetric to planner-wave.md's PLAN.md check. The excluded item is not lost: merge-wave.md's own mergeable-wave criterion (status=pending, SUMMARY.md on disk, not yet merged) already picks it up independently of this eligible/spawn list. Workflow-prose-only fix — touches no already-merged/reviewed .cts module. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §1 for the fix-location rationale (why not resumeBatch itself). * test(#3677): add real-git coverage for worktree-ownership tampering, scope drift, and submodules Closes the three coverage gaps identified in 40-design.md §2/§3 (#3677, epic #3344 Phase 5's own AC bullets: "arbitrary-worktree ownership attempts", "scope drift", "submodules"): - Arbitrary-worktree ownership tampering: a manifest entry naming a non-agent branch is silently dropped at normalization before any git subprocess runs; a manifest entry naming a plausible agent-branch that was never actually created by this repo's own worktree.create (a genuinely foreign repo/branch) is blocked via base_mismatch. Both leave the foreign location and repoRoot's HEAD provably untouched. - Advisory scope drift: a committed path outside declared files_modified still merges successfully (advisory, never blocking) while surfacing a scope_out_of_declared warning naming the drifted path; an exact declared-scope match produces zero warnings (boundary case). - Real .gitmodules submodule integration: a repo containing a real local git submodule merges cleanly through executeWorktreeWaveCleanupPlan for an unrelated plan; a real gitlink pointer bump (declared) merges cleanly with the superproject tree reflecting the new pinned commit; an undeclared bump is advisory-only and surfaces a scope warning naming vendor/sub, same as any other undeclared modification. No src/*.cts changes — all three gaps were coverage-only; the underlying primitives already behaved correctly (independently verified against real git subprocess output before writing each assertion). * docs(#3677): document how to diagnose a preserved quick-batch worktree Extends the one-sentence "worktree is preserved (never deleted)" mention into a concrete diagnosis procedure: where the preserved directory is, how to read the executor's real commits/diff against the plan's declared files_modified, how to read the item's own SUMMARY.md independent of merge outcome, how to manually merge-and-clean-up or discard, and how to re-run --resume afterward. Also documents that a SUMMARY.md-written-but-still- pending item (the crash-window case fixed in this same PR) needs no manual intervention — --resume routes it straight to the merge step. * chore(#3677): checkpoint final acceptance-evidence mapping (gitignored, dev-only) * fix(#3677): make crash-window duplicate-dispatch guard behaviorally provable and durably recoverable Orthogonal review (Spec finding): the crash-window regression test added earlier this phase only asserted readStep('worktree-dispatch.md') + regex matches against the markdown prose — proving the DOCUMENTATION says the right thing, never that the runtime condition (pending status + on-disk SUMMARY.md + absent STATE row) is actually handled correctly. #3677's own "Alternatives considered" explicitly rejects "document recovery without fault injection" for exactly this reason. Extracts the filtering decision into a pure, independently testable function, filterAlreadyExecuted(eligibleIds, executedIds) in src/quick-batch-dispatch.cts, wired to a new `quick-batch filter-executed` CLI verb (src/quick-batch-command-router.cts) — the same pure-decision- then-CLI-wired pattern computeSpawnPlan/computeMergeOrder already establish. worktree-dispatch.md now calls this verb explicitly instead of only describing the decision in prose. A genuine fixture-based test in tests/quick-batch.test.cjs constructs a REAL BATCH.json (createBatch), writes a REAL SUMMARY.md on disk at the item's real item_dir, calls the REAL resumeBatch, and proves both that resumeBatch alone still reports the item eligible AND that filterAlreadyExecuted (fed a real filesystem check) correctly excludes it. The prior prose-assertion tests are kept — they now prove the workflow markdown is correctly WIRED to the verb — but are no longer the only proof. Self-discovered defect while building that fixture (fixed inline, not deferred): tracing merge-wave.md against /gsd:quick's own prior art (QUICK_WORKTREE_MANIFEST=$(mktemp ...), quick.md:415) showed $QUICK_BATCH_WORKTREE_MANIFEST is a fresh PER-PROCESS temp file. A resumed coordinator correctly does not re-dispatch an already-executed item (this fix), but nothing durably recorded that item's worktree_path/branch/base either — Step 7 in the resumed process would have had no data to build its cleanup-wave entry from. Adds dispatched_worktree/dispatched_branch/ dispatched_base to QuickBatchItem (src/quick-batch.cts) — deliberately NOT a reuse of the pre-existing `worktree` field, whose loadBatch validation requires the path to exist on disk (verified empirically: reusing it made the batch permanently unloadable the moment a legitimately-merged worktree was removed). worktree-dispatch.md persists the triple once a worktree is created; merge-wave.md falls back to it when the ephemeral manifest lacks an entry, clears it after a successful merge, and fails closed rather than guessing if no record exists anywhere. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.1 and §9.3 for the full trace, empirical verification notes, and rejected alternatives (reusing `worktree` directly). * test(#3677): prove the arbitrary-worktree-ownership boundary against two real sibling worktrees Orthogonal review (Security finding): the two existing ownership-tampering tests didn't test ownership — one was trivially rejected by WORKTREE_AGENT_BRANCH_RE's shape check before any git call (proves branch- NAME filtering, not ownership), the other pointed at a wholly separate, never-linked foreign repo, so merge-base failed immediately because the branch didn't exist as a ref at all. Neither exercised the real scenario: a manifest entry whose worktree_path/branch are swapped to point at a DIFFERENT, GENUINELY-REGISTERED sibling worktree of the SAME repoRoot, with a branch name passing the shape check and a base in allowed_bases. Investigated executeWorktreeWaveCleanupPlan (src/worktree-safety.cts) directly: this is NOT a reachable gap. Git enforces branch-per-worktree uniqueness, so a swapped-in entry.branch can only match worktree_path's ACTUAL checked-out branch if it names that sibling's own real, uniquely- generated branch name — which manifest tampering confined to one batch's own record has no way to know (branch names are agent-<quick_id>[-<timestamp>]-shaped, and quick_id allocation is collision-checked GLOBALLY across every existing quick task and batch, not merely within one batch). Adds a stronger test that empirically proves this: two REAL, concurrently- alive sibling worktrees of the same repo (both via real `git worktree add`, both WORKTREE_AGENT_BRANCH_RE-passing, both sharing one merge-base), with worktree_path/branch swapped between them in both directions. Both attempts are blocked via branch_mismatch; both real worktrees, their branches, and one sibling's real uncommitted-to-main commit survive completely untouched. Supplements (does not replace) the original two tests, which still prove distinct, real boundaries. See .gsd/phase/feat-3677-quick-batch-hardening-acceptance/40-design.md §9.2 for the full trace, including the one explicitly-documented (not fixed) trust boundary this investigation surfaced: the primitive defends against fabricated data, not a caller bug that misattributes a real-but-wrong item's own triple to a different item. * chore(#3677): checkpoint design-doc addendum for review pass 2 findings (gitignored, dev-only) * docs(#3677): add changeset for PR 4240 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8f013983e5 |
fix(#3751): provision the claude CLI in CI and cover agents/ in the plugin-validate fixture (#4229)
* test(#3751): the validation fixture must cover agents/ and CI must provision the CLI * fix(#3751): cover agents/ in the plugin-validate fixture and provision the claude CLI in CI * fix(#3751): wire the strict live-config guard into the plugin-validate job * fix(#3751): job-level strict-guard env, where the guard derivation reads it * chore(#3751): changeset for the CI-provisioned plugin-validate gate * chore(#3751): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
3639ab0431 |
fix(#3968): measure commit claims at all three surfaces — ledger, verifier BLOCKER, porcelain HANDOFF (#4230)
* test(#3968): commit claims must be measured against git, never narrated * fix(#3968): measure commit claims at all three surfaces Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument Emitted-Drift-Ack-Growth: pause-work.md — #3968 uncommitted_files from git status --porcelain * fix(#3968): retired slash syntax, allowlist line pin, git-compare test pin * fix(#3968): persist the ledger on disk and reconcile with the same rev-list instrument Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) Emitted-Drift-Ack-Growth: verify-work.md — #3968 commit-claim reconciliation with the same rev-list instrument * fix(#3968): hold the gsd-executor size cap with a compact ledger contract Emitted-Drift-Ack-Growth: gsd-executor.md — #3968 plan commit ledger and measured commits contract (held under the agent size cap) * fix(#3968): allowlist pin and HALT regex track the final prose * chore(#3968): changeset for measured commit claims * chore(#3968): backfill changeset pr number * fix(#3968): quote the BASE expansion (SC2086) * ci: raise the test-lane budget 21 to 32 minutes (measured cost grew past the cap) --------- Co-authored-by: sim <sim@local> |
||
|
|
d5f8191f66 |
fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick (#4216)
* test(#3730): a legacy Quick Tasks table must be migratable to canonical * fix(#3730): quick-tasks-migrate — canonical-schema repair, auto-run on first quick * fix(#3730): review fixes — usage parity, contiguous table span, collision-safe bucket, template-width delimiter Emitted-Drift-Ack-Growth: fast.md — #3730 runs quick-tasks-migrate before the first append (auto-migration on first quick run) Emitted-Drift-Ack-Growth: quick.md — #3730 replaces the match-any-format note with the migration instruction * chore(#3730): backfill changeset pr number * fix(#3730): scope the quick-batch row-48 guard to branches touching quick-batch --------- Co-authored-by: sim <sim@local> |
||
|
|
114dfcb739 |
fix(#4196): npm-audit gate blocks only NEW advisories, not pre-existing ones (#4214)
* fix(#4196): npm-audit gate blocks only NEW advisories, not pre-existing ones The #3588 gate failed on ANY advisory in the production tree, regardless of whether the PR/push actually introduced it. Because npm's advisory database updates continuously and independently of repo state, a commit could pass this gate at merge time and fail it minutes later on the identical tree -- proven on PR #4188/dce40eeb6, which passed on all 3 OSes at 15:57-16:27 and failed the same assertion at 16:19-16:30 on the unchanged commit, purely because GHSA-jqff-g426-hqxp was disclosed for fast-uri in the interim. scripts/npm-audit-baseline.cjs diffs the head tree's vulnerable-package set against a resolved baseline (the PR's target branch, or the prior commit on a direct push) and blocks only newly-introduced advisories. When no baseline can be resolved, falls back to the original zero-tolerance behavior -- fail-closed, never silently weaker. * fix(#4196): pin the npm-audit baseline instead of using a drift-prone ref Two orthogonal reviews found the same class of bug this repo already fixed once for a different gate (see GSD_EMITTED_BASE's own incident comment in test.yml): origin/<branch> is live under fetch-depth: 0 and can advance mid-run, so resolveBaselineRef()'s fallback to origin/${GITHUB_BASE_REF} could silently disagree with the tree ci-rebase-check.cjs actually merged. Wire AUDIT_BASELINE_REF from the workflow to github.event.pull_request.base.sha / github.event.before, the same pinned values GSD_EMITTED_BASE already relies on. Also: HEAD~1 assumed exactly one commit per push, which this repo's allow_rebase_merge:true setting can violate (a rebase-merged PR lands as several discrete commits in one push) -- github.event.before is git's own record of the correct pre-push state, not an assumed offset. HEAD~1 remains as a documented last-resort fallback for out-of-band invocations (e.g. gsd-test) that don't set any of the above, alongside a new local-branch fallback for gsd-test's local `next` (not origin/next) sandbox shape. --------- Co-authored-by: sim <sim@local> |
||
|
|
2f64e6230a |
feat(#3676): quick-batch command, workflow, and isolation integration (#4212)
* test(#3676): add failing tests for quick-batch dispatch core Failing-first tests for Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding"): quick-batch-dispatch.test.cjs / .property.test.cjs cover the new pure decision-logic module (arg validation, effective concurrency, deterministic merge order, spawn backpressure, verification/merge routing, cleanup-entry construction — design doc rows 3-15,24,26-28, 30-36,39; property rows 51-53). quick-batch-update-items.test.cjs covers the new updateBatchItems export on src/quick-batch.cts (rows 15,22-23, including the negative cycle-rejection case). quick-batch-command-router.test.cjs covers the new gsd-tools quick-batch CLI family (rows 46-47). These reference modules/ exports that do not exist yet. * feat(#3676): implement quick-batch dispatch core, updateBatchItems, and command router Phase 4 of epic #3344 (ADR-1239 "Quick-batch binding") CORE decision layer — CLI verbs and pure orchestration logic only; no workflow markdown, no Agent()/git-worktree I/O. - src/quick-batch-dispatch.cts (new): pure decision functions consumed by the (separate, follow-up) /gsd:quick-batch workflow markdown — parseQuickBatchArgs, computeEffectiveConcurrency, computeMergeOrder, computeSpawnPlan, routeVerificationOutcome, routeMergeOutcome, buildCleanupManifestEntry (the last parses caller-supplied plan text via the existing parsePlanDocument; no filesystem access). - src/quick-batch.cts: adds updateBatchItems, resolving the design doc's Open Question 1 as ONE additive export on this module instead of the second, independent BATCH.json writer the design doc originally proposed. Reuses the same withPlanningLock transaction shape, computeWaves, and platformWriteSync call resumeBatch/ completeQuickItem already use; fails closed without persisting on an unknown item, an unknown/self dependency, or an introduced cycle. - src/quick-batch-command-router.cts (new): gsd-tools quick-batch CLI family, wired into HOST_COMMAND_ROUTERS (gsd-core/bin/gsd-tools.cjs) as a first-party always-on command (like /gsd:quick), not the opt-in capability-registry path graphify uses. Verbs: create/update/resume/ complete (wrap quick-batch.cts) and effective-concurrency/ merge-eligible/spawn-plan/verification-routing/merge-routing/ cleanup-entry/parse-args (wrap quick-batch-dispatch.cts). Design doc rows covered: 3-15, 22-24, 26-28, 30-39, 46-47. Property rows 51-53. Rows covering workflow markdown / Agent() dispatch / `git worktree` behavior (16-21, 25, 29, 40-45, 48-50) remain for the follow-up markdown-authoring pass, per the phase brief's explicit scope boundary. * docs(#3676): register quick-batch-dispatch/command-router modules in bookkeeping surfaces New-.cts-module ripple for the two Phase 4 modules (epic #3344, ADR-1239 "Quick-batch binding"): .gitignore (compiled .cjs artifacts, ADR-457 build-at-publish), eslint.config.mjs (lint the .cts source, not the emitted .cjs), docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json (via `node scripts/gen-inventory-manifest.cjs --write`, after `npm run build:lib`), and CONTEXT.md glossary entries for "Quick-Batch Dispatch Core Module" and "Quick-Batch Command Router Module", plus an update to the existing "Quick-Batch Core Primitives Module" entry documenting the new updateBatchItems export. * test(#3676): fold updateBatchItems tests into quick-batch.test.cjs (fix lint-test-file-count) scripts/lint-test-file-count.cjs buckets any quick-batch-*.test.cjs file under the quick-batch production module by longest-prefix match, and that module is already at its 2-file cap (quick-batch.test.cjs + quick-batch.property.test.cjs). The standalone tests/quick-batch-update-items.test.cjs added in the prior commit pushed it to 3 and failed `npm run lint:ci`. Fold its content into quick-batch.test.cjs (append-only — no existing test in that file is modified) and update the CONTEXT.md glossary reference to match. Surfaced while re-running `GITHUB_BASE_REF=next npm run lint:ci` after `npm ci` (this worktree previously had no local node_modules, which also made gen-scripts-cli-exit/gen-hooks-cli-exit/gen-exit-code-* unable to resolve typescript — resolved by npm ci, no code change needed there). `npm run lint:ci` and `npx tsc -p tsconfig.build.json --noEmit` are both green after this fix. * test(#3676): add failing tests for the quick-batch command/workflow markdown Failing-first tests for Phase 4's markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding"): gsd-quick-batch-workflow.test.cjs covers commands/gsd/quick-batch.md's frontmatter/objective/process, gsd-core/workflows/quick-batch.md's byte-size boundary (row 49, ADR 1610 NEW_FILE_CAP) and step-fragment count, the isolation model (rows 20-22), the executor single-writer invariant (row 18), merge validation reusing the existing bounded primitive (row 25), the optional research/plan-checker/verification leaves (rows 16,17,19, 30,31), planning-failure blocking execution (row 29), the submodule guard (rows 36,44), and the new agents/gsd-planner.md quick-batch mode (rows 13-15). gsd-quick-batch-quick-regression.test.cjs covers row 48 (ordinary /gsd:quick stays byte-identical). Named `gsd-quick-batch-*` (not `quick-batch-*`) so lint-test-file-count's longest-prefix bucketing doesn't fold these markdown-only tests into the already-capped quick-batch/quick-batch-dispatch/ quick-batch-command-router production-module buckets from the CORE pass. These reference files that do not exist yet. * feat(#3676): author the quick-batch command, workflow, and planner mode Phase 4 markdown-authoring pass (epic #3344, ADR-1239 "Quick-batch binding") — the orchestration layer that calls into Pass 1's CLI verbs (src/quick-batch-command-router.cts). - commands/gsd/quick-batch.md (new): frontmatter/objective/process, delegates argument validation to `quick-batch parse-args` (parseQuickBatchArgs) rather than re-deriving the grammar. - gsd-core/workflows/quick-batch.md (new, 11843 bytes — under ADR 1610's 32768-byte NEW_FILE_CAP for a brand-new file) + 9 lazy-loaded step fragments under gsd-core/workflows/quick-batch/steps/: resume-mode, batch-init, research-phase (flag:--research), planner-wave (+ nested plan-checker-loop when --validate), worktree-dispatch, merge-wave, verification-wave (flag:--validate), completion. Covers design doc rows 3-45: capacity/isolation resolution (reusing dispatch-isolation-gate.md verbatim), per-DAG- layer planning with full-task-catalog prompts and always-required depends_on/files_modified frontmatter, serialized worktree create/ merge/cleanup via the existing worktree.cleanup-wave primitive, deterministic wave-order merging, verification routing (human_needed/gaps_found), the executor single-writer invariant, submodule fail-loud guard, and #1941 fork-base auto-degrade. - agents/gsd-planner.md: additive new `load_mode_context` bullet for `**Mode:** quick-batch`, pointing at the new gsd-core/references/planner-quick-batch.md reference (documents the always-required depends_on/files_modified contract, reusing the existing frontmatter grammar — no new keys). Existing modes byte-identical, only a new bullet added. - src/init.cts (+init-command-router.cts, +command-aliases.cts): cmdInitQuickBatch / `init.quick-batch` — model profiles, commit_docs, roadmap/planning existence checks, and the section_manifest field gating research-phase/verification-wave (reuses the existing flag:--research/flag:--validate WHEN_VOCABULARY atoms — no new atom needed). Rows 16-21, 25, 29, 36, 38, 39, 44, 46-50 covered structurally by the prior test(#3676) commit; rows 3-15, 22-24, 26-28, 30-35, 37, 40-43, 45 covered by construction (verb wiring, single-writer prompt constraints, crash-window resume via unmodified Phase 3 primitives). * docs(#3676): regenerate skills/inventory/section-manifest/install-tree; baseline the intentional word-splitting pattern npm run regen:derived output for the new command/workflow/reference (epic #3344, ADR-1239 "Quick-batch binding"): - skills/gsd-quick-batch/SKILL.md (generated from commands/gsd/quick-batch.md) - docs/INVENTORY.md rows for /gsd-quick-batch, quick-batch.md, planner-quick-batch.md, and the quick-batch-dispatch.cjs/ quick-batch-command-router.cjs CLI-module rows' now-live `/gsd-quick-batch` cross-reference (was "(separate, follow-up)") + docs/INVENTORY-MANIFEST.json (`node scripts/gen-inventory-manifest.cjs --write`) - gsd-core/workflows/section-manifest.json (`npm run gen:section-manifest`) — research-phase/verification-wave gsd:section entries for the new quick-batch workflow - tests/fixtures/install-tree/*.json (`npm run gen:install-tree`) — the new command/workflow/skill/reference files now ship to every runtime scripts/lint-workflow-shellcheck-baseline.json: 3 new entries for gsd-core/workflows/quick-batch.md's intentional flag-token/$ARGUMENTS word-splitting (SC2046/SC2086) — the same deliberate unquoted-optional- flag pattern gsd-core/workflows/quick.md already carries baselined (e.g. `$DISCUSS_PARAM $RESEARCH_PARAM` in quick.md's own Step 2); quoting would break the intended "omit this arg when the flag is false" splitting. * fix(#3676): close prompt-injection and argv/glob-injection gaps in quick-batch leaf dispatch Security review pass findings, both confirmed real: 1. HIGH — prompt injection, no boundaries. Every leaf-dispatch fragment interpolated the raw, attacker-influenced task ${description} (and the shared ${TASK_CATALOG_TABLE}, broadcasting every item's raw description into every planner's prompt in the layer) straight into Agent() prompt bodies with no boundary. Fixed by wrapping every such interpolation in a <security_context> + DATA_START/DATA_END boundary, matching the CONCRETE convention already implemented in this repo (agents/gsd-debug-session-manager.md, agents/gsd-debugger.md, gsd-core/workflows/debug.md) — commands/gsd/quick.md's own <security_notes> only asserts this convention in prose, so the debug-agent files are the real precedent followed here. Added a new <security_notes> block to commands/gsd/quick-batch.md (it had none) documenting both this fix and the one below. 2. MEDIUM — unquoted $ARGUMENTS -> argv/glob injection. gsd-core/workflows/quick-batch.md and commands/gsd/quick-batch.md both ran `gsd_run quick-batch parse-args --raw -- $ARGUMENTS` UNQUOTED, causing shell word-splitting and pathname expansion on raw task-list text before the parser ever saw it. Fixed at the source: added a `--text <string>` form to the `parse-args` verb (src/quick-batch-command-router.cts) that accepts the ENTIRE $ARGUMENTS as ONE quoted argv element and does the whitespace split itself, in Node — which is never glob-aware, unlike the shell. Both call sites now use `--text "$ARGUMENTS"`. The `-- <tokens>` form is kept for direct/test callers that already have a real argv array. The SC2086 baseline entry added for the original unquoted line is now stale (`node scripts/lint-workflow-shellcheck.cjs` no longer reports it) and has been removed; the two SC2046 entries for the UNRELATED, still-unquoted `$([ "$VALIDATE_MODE" = true ] && echo --validate)`-style conditional-flag splitting remain — that line only ever expands to one of a few known-safe literal strings (never raw user text), matching quick.md's own already-baselined convention exactly. Tests: quick-batch-command-router.test.cjs covers the new --text form (token splitting, glob-shaped text passing through literally unexpanded, whitespace-only input). gsd-quick-batch-workflow.test.cjs asserts the DATA_START/DATA_END boundary on every leaf prompt (research-phase/planner-wave/plan-checker-loop/verification-wave, including the shared task catalog) and the quoted --text call sites. * fix(#3676): strengthen test-depth gaps in rows 9, 18, 24, 34, 35 Spec review pass findings — the test matrix claimed "yes" coverage these assertions did not actually support: - Row 9 (--jobs 0/-1/abc hostile case): previously asserted rejection only. Added an end-to-end assertion (tests/quick-batch-command-router.test.cjs, committed alongside the security fix that touches the same file) that .planning/quick-batches/ is never created for any rejected value — createBatch is genuinely never reached. - Row 18 (--resume <unknown-batch-id>): previously only exercised a hand-corrupted BATCH.json, never a genuinely nonexistent batch directory. Added the real nonexistent-id case (also in quick-batch-command-router.test.cjs). - Row 24 (post-planning updateBatchItems racing a concurrent completeQuickItem for a different item, both through withPlanningLock): zero test existed. Added a property test (tests/quick-batch.property.test.cjs, appended — Phase 3's own file, no existing test touched) exercising both call orders and asserting no lost update in the final on-disk manifest — the same technique Phase 3's own row-15 lock-contention property test uses (sequential calls through the real lock; a working mutex makes any interleaving equivalent to some serial order, so this is the same claim a literal concurrent-thread test would make without OS-level threading). - Row 34 (worktree preserved on merge_failed) and row 35 (undeclared- deletion detection): both were previously asserted only at the pure routeMergeOutcome level. Added tests/gsd-quick-batch-merge-integration.test.cjs using the SAME real-git-fixture pattern tests/worktree-safety.test.cjs already establishes for executeWorktreeWaveCleanupPlan (real repo, real worktree, a REAL merge conflict / a REAL file deletion diffed against declared_deletions) — asserting the actual worktree directory survives on disk, not just that a pure function returns a preserveWorktree:true field. Named gsd-quick-batch-* so lint-test- file-count's bucketing doesn't fold it into any capped module bucket. Row 48 (/gsd:quick regression) intentionally left as-is per the reviewer's own framing: the byte-identity claim is already mechanically proven by the changed-path diff (git diff --name-only empty on those two paths IS byte-identity), and a genuine execution- level regression test would require actually running the workflow — out of scope for this repo's unit-test model (no other quick.md regression test in this repo does that either). * docs(#3676): add the changeset and user-facing docs the command needed Standards review pass findings — both HARD: - Missing changeset. None of the 6 prior #3676 commits touched .changeset/*. /gsd-quick-batch is a new user-facing command; CLAUDE.md/CONTRIBUTING.md require one. Added .changeset/silly-rams-caper.md (type: Added, pr: 0 placeholder — backfilled after the PR opens, matching CLAUDE.md's own documented convention and Phase 3's own precedent, #4190's .changeset/mellow-yaks-squeak.md). Uses the docs-convention hyphen form `/gsd-quick-batch` throughout, never the source-artifact colon form (`scripts/lint-docs-command-form.cjs` confirms 0 violations; that check scans docs/**, not .changeset/, so it was never actually in scope for the fragment itself, but the wording still follows the doc convention for consistency, matching how Phase 3's own fragment named the not-yet-shipped command). - Missing docs. Added docs/how-to/batch-quick-tasks.md (Diátaxis how-to, matching docs/how-to/handle-quick-and-fast-tasks.md's existing convention for /gsd-quick /gsd-fast) covering --jobs, --validate, --research, --resume, --file, the capacity/isolation interaction, and resume/failure recovery. Cross-linked from docs/README.md's how-to index and from handle-quick-and-fast-tasks.md's own "Related" section. Added a /gsd-quick-batch section to docs/COMMANDS.md (same table format as the existing /gsd-quick entry) and docs/features/quick-batch.md (REQ-QB-01..12, same frontmatter shape as docs/features/quick-mode.md) — regenerated docs/FEATURES.md (179 features) and skills/gsd-quick-batch/SKILL.md via the standard generators. * fix(#3676): close docs-parity, attribution, and generated-registry gaps gsd-test caught gsd-test's real run against 155e8975b3 found 43 failures, all rooted in this phase's own new command/workflow never being registered across ~10 independent generated/hand-maintained registries this repo keeps in parity by convention. Root-caused each, no test weakened or special-cased. - help.md ↔ commands/gsd/ bidirectional parity (docs-parity-live- registry.test.cjs): added a /gsd:quick-batch entry to gsd-core/workflows/help/modes/full.md (the real help.md content; gsd-core/workflows/help.md is a thin dispatcher) documenting every flag (--file/--jobs/--validate/--research/--resume), matching the existing /gsd:quick entry's format. - gen-section-manifest.test.cjs: quick-batch.md's `gsd_run query init.quick-batch` invocation used inline `$([ ... ] && echo --flag)` substitutions, which never satisfy the test's exact-whitespace-token / assigned-variable detection (the trailing `))` glued onto `--research` in the compound substitution broke the "exact token" match). Rewrote to the same VALIDATE_PARAM/RESEARCH_PARAM two-line pattern gsd-core/workflows/quick.md's own Step 2 already uses. - runtime-launcher-parity.test.cjs: the 8 quick-batch/steps/*.md fragments that call gsd_run each needed their OWN embedded copy of the canonical shim preamble (every workflow .md that calls gsd_run carries its own copy — reading one file does not persist shell state into another). Ran `node scripts/sync-runtime-launcher.cjs`, which inserted it before each file's first gsd_run call. plan-checker-loop.md correctly has none — it never calls gsd_run directly. - Namespace routing (skill-manifest.test.cjs, install-nested- layout.test.cjs, runtime-artifact-layout-surface.test.cjs): added `quick-batch` to commands/gsd/ns-workflow.md's `requires:` array and routing table (same namespace `quick` already routes through), and to src/clusters.cts's `utility` cluster (same cluster `quick` already belongs to). Verified by hand-running installRuntimeArtifacts + applySurface for augment/cline against a real temp install: exactly 6 top-level gsd-ns-* router dirs, gsd-quick-batch correctly nested under gsd-ns-workflow/skills/, never re-flattened. - mcp-server-catalog.test.cjs: hardcoded command count 71 -> 72 (a brand-new command is a real count change, not a bug this test should hide). - model-omit-when-inherit-guard.test.cjs: added the canonical `<!-- #2517 model-omit-on-inherit -->` marker block to gsd-core/workflows/quick-batch.md (every leaf dispatch — planner/ researcher/checker/executor/verifier — lives in a steps/ fragment, read combined with the host by this test's own readWorkflowCombined, same as quick.md's own research-phase.md carries it for its gated section). Also fixed a genuine pre-existing inconsistency in the test's own "#2711: the guarded set is derived from dispatch sites" check: its `nonDispatching` computation read the BARE host file while `derived` (the set it's checked against) reads the combined host+steps content — inconsistent with that same test file's own #2994 doc comment explaining why the combined read is necessary. quick-batch.md is the first workflow whose EVERY model="{...}" dispatch site lives in a mandatory (never gated) steps/ fragment — extracted to stay under ADR-1610's tighter NEW_FILE_CAP for a brand-new file — which is what exposed the mismatch. Fixed by using the same readWorkflowCombined read in both places. - skill-frontmatter-contract.test.cjs: shortened commands/gsd/quick-batch.md's frontmatter `description` from 107 to 91 chars (<=100 budget), and added `quick-batch.md` to the hand- maintained KNOWN_SKILLS consolidation allowlist with a #3676 justification comment (a genuinely new first-party command, not a consolidation of an existing skill). - workflow-fragments-emission.install.test.cjs: added `quick-batch.md` to the hand-maintained MARKED_WORKFLOWS set (composeWorkflow is deliberately NOT a no-op for it — its research-phase/verification- wave sections are gated). - Regenerated all downstream artifacts (npm run build:lib && npm run regen:derived && npm run gen:plugin-skills -- --write && npm run gen:features -- --write): skills/gsd-quick-batch/SKILL.md, skills/gsd-ns-workflow/SKILL.md, install-tree fixtures for augment/cline/hermes/qwen/trae/zcode. - emitted-attribution.test.cjs: agents/gsd-planner.md's #3676 addition (one new `load_mode_context` bullet pointing at the new gsd-core/references/planner-quick-batch.md reference) grew the file 124 bytes without an acknowledgment trailer. Acknowledged below — the growth is the deliberate, additive, single-bullet change from the earlier feat(#3676) commit, not drift. Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green (includes lint-workflow-shellcheck, lint-test-file-count, lint-docs-command-form). The deep install/spawn/registry tests gsd-test actually runs (docs-parity-live-registry, gen-section-manifest, runtime-launcher-parity, install-nested-layout, runtime-artifact-layout-surface, skill-manifest, skill-frontmatter- contract, mcp-server-catalog, model-omit-when-inherit-guard, workflow-fragments-emission) are not part of lint:ci — each fix above was independently verified by hand-invoking the exact production function the failing test calls (installRuntimeArtifacts, applySurface, composeWorkflow, the CLUSTERS union, the section-manifest forwarding regex) against the real repo tree and confirming the expected shape. Emitted-Drift-Ack-Growth: gsd-planner.md — additive #3676 quick-batch mode bullet in load_mode_context (one new line pointing at gsd-core/references/planner-quick-batch.md); not drift. * fix(#3676): trim the /gsd:quick-batch help.md entry to fit the LARGE tier line budget skill-frontmatter-contract.test.cjs's "feature #3039: tiered help — size budgets" enforces a SEPARATE line-count ceiling for gsd-core/workflows/help/modes/full.md (FULL_BUDGET = 844 lines, tighten-only ratchet, scripts/lib/allowlist-ratchet.cjs's assertTightCeiling) — independent of the skill-frontmatter description- length budget and consolidation allowlist I touched in the prior round; those are unrelated checks in the same test FILE, not the same check. Root cause: the /gsd:quick-batch entry I added to full.md in the docs-parity fix round was 17 lines, pushing the file from 834 to 851 lines — 7 over the 844 ceiling. Condensed the entry (merged the per-flag bullet list into one dense "Flags:" line, dropped from 3 Usage examples to 1) to 844 lines exactly — at the ceiling with zero slack, which assertTightCeiling accepts (it only fails on actualMax > ceiling, or on slack > grace when the ceiling is too LOOSE — zero slack triggers neither). Verified after trimming: full.md still contains a live /gsd:quick-batch reference (bidirectional parity) and all 5 argument-hint flags (--jobs/--validate/--research/--resume/--file) still appear as literal tokens (docs-parity-live-registry.test.cjs's own flag-coverage check, re-run by hand against the trimmed content). Verified: npm run build:lib clean, npx tsc -p tsconfig.build.json --noEmit clean, GITHUB_BASE_REF=next npm run lint:ci fully green. * docs(#3676): backfill changeset pr number to 4212 Follow-up to fix(#3676) commits — .changeset/silly-rams-caper.md's pr:0 placeholder backfilled with the real PR number now that gh api POST /pulls has returned it (#4212). Matches CLAUDE.md's PR Number Handling convention and Phase 3's own #4190 precedent (708c5a3f8c). Doc-only (root-level .changeset/*.md fragment), exempt from a fresh gsd-test run per pre-pr-gate.sh's DOC_ONLY_RE. * fix(#3676): resolve prompt-injection-scan false positive on test fixture tests/quick-batch.test.cjs:232's row 11b regression proves the task-list parser carries a prompt-injection-shaped task description through createBatch as inert data, never interpreted. The fixture has to be a real "ignore all previous instructions..." phrase or the test asserts nothing, but the full-file --diff scan flagged it once unrelated edits in the same file pulled it into the changed-file set. Add the file to prompt-injection-scan.sh's ALLOWLIST, matching the sanctioned, precedented exemption already used for other legitimate security-regression fixtures (tests/windsurf-conversion.test.cjs, tests/health-validation.test.cjs, tests/continuation-grammar-parity.test.cjs) per DEFECT.PROMPT-INJECTION-SCAN-COLLISION. --------- Co-authored-by: sim <sim@local> |
||
|
|
7d6d788b51 |
fix(#4020): bound the test run's temp footprint with a swept run-scoped root (#4207)
* test(#4020): the runner must bound and sweep a run-scoped temp root * fix(#4020): bound the run's temp footprint with a swept run-scoped root * test(#4020): isolate the env-mutating rows in child processes * fix(#4020): gate root removal on ownership so nested runners spare the outer root * test(#4020): pass the probe file via --files, the runner's explicit-file flag * test(#4020): resolve the probe by basename, as --files matching requires * test(#4020): assert root survival, not content survival, in the nested-row * chore(#4020): changeset for the run-scoped temp root * chore(#4020): backfill changeset pr number * fix(#4020): the sweep spares ancestors of the runner's own selected files * fix(#4020): TMPDIR precedence — an operator redirect beats inherited TEMP/TMP * fix(#4020): only the root's owner sweeps — a nested runner spares live sibling fixtures --------- Co-authored-by: sim <sim@local> |
||
|
|
858bb89769 |
ci(#4196): exempt dependabot[bot] from issue-link, title, and unsolicited-PR gates (#4203)
Dependabot has no mechanism to link a PR it opens to a repo issue -- its alerts live in the Security tab, not as issues -- so require-issue-link, pr-title-validator, and auto-close-unsolicited-prs all rejected its PRs by design (confirmed live on #4193: auto-closed for "no pre-approved issue", then flagged again by the title gate on reopen). Exempt by authenticated author login (github.event.pull_request.user.login / context.payload.pull_request.user.login), which GitHub attributes and a crafted title or branch name cannot forge -- scoped narrowly to dependabot[bot] only, no other author gets this treatment. Co-authored-by: sim <sim@local> |
||
|
|
91ed46882a |
feat(#3675): quick-batch core primitives and resumable manifest (#4190)
* test(#3675): add failing tests for quick-batch core primitives Adds the full behavioral (tests/quick-batch.test.cjs) and property-based (tests/quick-batch.property.test.cjs) coverage for #3675's quick-batch core primitives per the phase's 35-row test matrix — task-list parsing (inline + --file, with path-confinement/symlink-escape/non-regular-file rejection), collision-safe quick-id preallocation under withPlanningLock, BATCH.json schema/validation/resume, dependency-DAG + partitionByFileOverlap wave construction, and exactly-once STATE.md completion (including the STATE-row-written-but-manifest-not-yet-updated crash window). The import target (gsd-core/bin/lib/quick-batch.cjs, compiled from a not-yet-written src/quick-batch.cts) does not exist yet — every test in both files fails at the top-level require() before any assertion runs. Five fast-check properties cover collision-freedom under lock contention, resume idempotency, exactly-once STATE completion, wave totality, and DAG-respecting wave order, per the design doc's property-based-coverage requirement. * feat(#3675): implement quick-batch core primitives Adds src/quick-batch.cts (ADR-457 build-at-publish, compiled to gsd-core/bin/lib/quick-batch.cjs) implementing #3675's quick-batch core primitives per the phase design lock — pure/state primitives and CLI-testable core operations only, no agent dispatch, no worktree creation, no user-facing command (Phase 4/#3676's job): - parseTaskList / parseTaskListFromFile: inline bulleted/numbered task-list parsing (>=2 items required) and a --file variant strictly confined to the planning workspace root via requireSafePath, rejecting non-regular-file targets. - allocateQuickIds / createBatch: collision-safe YYMMDD-xxx quick-id preallocation under withPlanningLock, checked against both on-disk .planning/quick/ entries and sibling .planning/quick-batches/*/BATCH.json manifests (never on-disk-only, which would miss another in-flight batch that hasn't dispatched any real quick directory yet) — replicates cmdInitQuick's own grammar rather than delegating to it (that function's 2-second granularity is not batch-safe). - computeWaves: deterministic wave construction combining dependency-DAG layering with partitionByFileOverlap (#3674), called per DAG layer over path-separator-normalized planned_files — normalization happens at this module's boundary, never inside the Phase 2 helper. - loadBatch: fail-closed BATCH.json schema validation (corrupt/truncated JSON, wrong types, missing fields, out-of-batch dependency references, dependency cycles, a worktree path absent from disk). - resumeBatch: skips complete items, never auto-retries failed items, propagates/reverses blocked status along the DAG to a fixed point, and detects a STATE.md row that already exists for a non-complete item (the "STATE written, BATCH.json not yet updated" crash window) — completing it without re-appending. Idempotent across repeated calls. - completeQuickItem / hasQuickTaskRow: exactly-once STATE.md completion — appendQuickTaskRow (unmodified) is called at most once per quick id, gated by hasQuickTaskRow's own idempotency check re-parsing the real "Quick Tasks Completed" table, since appendQuickTaskRow itself carries no idempotency. BATCH.json lives at .planning/quick-batches/<batch-id>/BATCH.json, a sibling of .planning/quick/ — never inside it, so scanQuickTasks never misreads a batch manifest as a broken quick task. * docs(#3675): register the new quick-batch module New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row plus the regenerated docs/INVENTORY-MANIFEST.json cli_modules entry, and a CONTEXT.md glossary entry matching the convention set by the sibling File Overlap Partitioner Module (#3674) entry it sits beside. NOTE: docs/INVENTORY-MANIFEST.json was updated BY HAND (alphabetically sorted single-entry insertion into families.cli_modules, matching the existing file's structure) rather than via `node scripts/gen-inventory-manifest.cjs --write` — this session's MEMTRACE-FIRST guard hard-blocks direct execution of that indexed script path from Bash, with no available Memtrace tool to route through instead. The orchestrator should re-run `node scripts/gen-inventory-manifest.cjs --check` to confirm this hand-edit is byte-identical to the generator's own output before merging. * fix(#3675): resolve lint findings in quick-batch primitives and tests Unsafe `any[]` assignment from `new Array(n)` in the DAG cycle-check color array, two unnecessary `as string[]` casts TS 5.5's inferred type predicates already narrowed, raw `fs.rmSync` in test cleanup (needs the Windows-EBUSY retry budget `helpers.cleanup` carries), an unused `loadBatch` import, an unbounded `mkfifo` subprocess spawn missing a timeout, and a CONTEXT.md glossary illustration that looked like a real file reference. * feat(#3675): close acceptance-criteria gaps found in review Standards- and spec-axis review (plus a self-caught race) surfaced real gaps against issue #3675's own acceptance criteria and this repo's test conventions: - BATCH.json was missing options, base_revision, per-item wave, and per-item commit — the issue's AC explicitly lists all four as things the manifest must track. Added them: createBatch persists caller-supplied batchOptions/baseRevision verbatim and assigns each item its computed wave index; completeQuickItem now persists the commit onto the item, not just the STATE.md row. All four are backward-tolerant on load (an older/hand-built manifest without them still validates). - resumeBatch had no "incompatible base divergence" check at all, despite the AC and the ADR's own "Base divergence" section requiring one. Added an opt-in currentBaseRevision comparison that fails closed with a recoverable diagnostic on mismatch, and touches nothing on refusal. - resumeBatch read-modify-wrote BATCH.json OUTSIDE withPlanningLock — the only durable write path in this module that wasn't lock-protected, a real lost-update race against a concurrent completeQuickItem or another resume. Now runs inside the same lock createBatch/ completeQuickItem use. - loadBatch and collectExistingBatchQuickIds used raw JSON.parse with no size cap (security review, Low/informational); switched to the existing safeJsonParse (1MB cap) for defense-in-depth. - Parser (parseTaskList) had only example-based tests; CLAUDE.md requires a fast-check property test for parsers. Added one plus a companion reject-property for <2 items. - The id-exhaustion fail-closed ceiling (MAX_TIME_BLOCK) was untested at any boundary. Exported the pure allocateIdsGivenUsed/MAX_TIME_BLOCK for direct limit-1/limit/limit+1 testing without needing 46k fixture dirs. - Issue AC explicitly asks for prompt-injection-payload test coverage, distinct from the existing shell-metacharacter test; added one. - Test row 9 (FIFO skip) silently returned instead of calling t.skip(), so an unsupported platform would report a pass rather than a documented skip; fixed to bind the test-context param and skip properly. - Extracted toWaveInput to remove a 2-site production duplication of the QuickBatchItem -> computeWaves reshape (Standards-axis smell). - Added the required .changeset/ fragment (CONTRIBUTING.md: editing src/ is user-facing even though the compiled .cjs is gitignored). * fix(#3675): restore "not valid JSON" wording in loadBatch's parse-failure reason gsd-test caught this: switching loadBatch to safeJsonParse changed the parse- failure message shape ("... parse error — ...") without preserving the "not valid JSON" substring row 27's own test asserts on. Re-wrap safeJsonParse's error into the original diagnostic phrasing regardless of which of its three failure modes fired. * docs(#3675): backfill changeset pr number to 4190 * fix(#3675): detect a silently-no-op mkfifo on Windows, not just a throwing one CI caught this on windows-latest: row 9's platform-skip only caught mkfifo throwing (command not found). On this runner mkfifo resolves to something that exits 0 without creating a file (NTFS has no FIFO concept), so execution fell through to parseTaskListFromFile against a path that doesn't exist, producing an ENOENT stat error instead of the expected "not a regular file" rejection. Check the artifact actually exists before trusting a zero exit code, and skip with a documented reason either way. --------- Co-authored-by: sim <sim@local> |
||
|
|
1c4a00244f |
ci(#4196): auto-merge Dependabot patch/minor bumps once required checks pass (#4200)
Dependabot already opens a correct fix PR within minutes of a new advisory (e.g. #4193 for GHSA-jqff-g426-hqxp), but nothing merged it -- it sat until a human noticed `next` had gone red and an unrelated PR tripped over the same npm-audit gate. Auto-approve + auto-merge closes that gap for patch/minor bumps; major bumps still need a human. Co-authored-by: sim <sim@local> |
||
|
|
0598a2cf2c |
chore(deps): bump the npm_and_yarn group across 1 directory with 2 updates (#4193)
Bumps the npm_and_yarn group with 2 updates in the / directory: [browserslist](https://github.com/browserslist/browserslist) and [fast-uri](https://github.com/fastify/fast-uri). Updates `browserslist` from 4.28.2 to 4.28.8 - [Release notes](https://github.com/browserslist/browserslist/releases) - [Changelog](https://github.com/browserslist/browserslist/blob/main/CHANGELOG.md) - [Commits](https://github.com/browserslist/browserslist/compare/4.28.2...4.28.8) Updates `fast-uri` from 3.1.5 to 3.1.7 - [Release notes](https://github.com/fastify/fast-uri/releases) - [Commits](https://github.com/fastify/fast-uri/compare/v3.1.5...v3.1.7) --- updated-dependencies: - dependency-name: browserslist dependency-version: 4.28.8 dependency-type: indirect dependency-group: npm_and_yarn - dependency-name: fast-uri dependency-version: 3.1.7 dependency-type: indirect dependency-group: npm_and_yarn ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7c52344284 |
fix(#4003): anchor the safe-resume gate's plan-scope greps to the milestone (#4194)
* test(#4003): safe_resume_gate must grep an anchored padding-tolerant scope * fix(#4003): anchor the resume-gate scope greps and bound them to the milestone tag Emitted-Drift-Ack-Growth: execute-phase.md — #4003 rewrites three commit-scope greps (safe_resume_gate, TDD RED, completion spot-check) to anchored zero-pad-tolerant regexes with a milestone tag bound; growth is the fix itself * test(#4003): align shape assertions with the implemented gate text * fix: bump fast-uri past GHSA-jqff-g426-hqxp (transitive, advisory reddened next) * fix(#4003): bound the TDD RED grep to the milestone and fix tdd.md's example greps * test(#4003): the gate pin tracks the anchored scope grep * fix(#4003): trim the gate rationale to hold the 93400 margin ceiling * test(#4003): the RED-grep pin tracks the milestone-bounded invocation * chore(#4003): changeset for the anchored resume-gate scope * chore(#4003): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
dce40eeb6e |
fix(#4002): rewrite zcode command @-refs to the zcode runtime home (#4188)
* test(#4002): zcode commands must rewrite at-refs to the zcode home * fix(#4002): add the missing zcode case to the runtime rewrite engine * chore(#4002): add ZCode to the bug-report runtime dropdown and drop the changeset * fix(#4002): attribute zcode command and skill ripples to the rewrite engine * fix(#4002): attribute zcode nested-skill ripples to the rewrite engine * chore(#4002): backfill changeset pr number * fix: bump qs past GHSA-x5fp-wj9c-mxmx (transitive, advisory reddened next) --------- Co-authored-by: sim <sim@local> |
||
|
|
acb903c2e8 |
enhance(#3661): make the code-review hook point configurable (#4159)
* feat(#3661): make the code-review hook point configurable Add `workflow.code_review_point` (`execute:post` default, or `execute:wave:post`) so a multi-wave phase can run code review once per wave instead of once at the end, scoped to what changed since the phase's prior review. The code-review capability now declares its step at both loop points via a new generic `pointFrom` step field: `pointFrom` names an enum config key, and the step is only active at its own `point` when that key resolves to a matching value. `_resolvePointGate` (capability-activation.cts) is the single shared implementation consumed identically by loop-resolver.cts and capability-state.cts, and capability-validator.cjs enforces that `pointFrom` references an enum key whose values cover the declaring step's own point. code-review.md's manual-invocation gate now reads `workflow.code_review` directly instead of probing registry presence at the hardcoded execute:post point (so manual `/gsd-code-review` keeps working regardless of which automatic point is configured), and its file-scope tiers narrow to what changed since the phase's last review commit when one exists. execute-phase.md's wave-post step dispatch gets a small, precedented carve-out so the code-review skill still receives its required phase argument when dispatched generically (caught by the isolated spec review). Closes #3661 Emitted-Drift-Ack-Growth: code-review.md — #3661 adds a point-aware config gate check and LAST_REVIEW_COMMIT-based incremental scoping to the file-scope tiers. Emitted-Drift-Ack-Growth: execute-phase.md — #3661 adds one carve-out sentence so the wave-post generic step dispatch passes PHASE_NUMBER to the code-review skill. * docs: backfill changeset PR number for #3661 (#4159) * fix: scope tests/io.test.cjs's fs.writeSync fault-injection mocks by fd Five fault-injection mocks in the "bug #1008" describe blocks intercepted every fs.writeSync call regardless of file descriptor, and several threw or truncated unconditionally on the first call. This surfaced as an intermittent macOS CI failure: node:test's own IPC channel back to the parent process (which also goes through fs.writeSync internally) could get a bogus injected error or truncated write if node's internal machinery called it while one of these mocks was active, corrupting the message frame the parent tried to deserialize ("Unable to deserialize cloned data.", location tests/io.test.cjs:1:1, uncaughtException — a whole-file IPC crash, not a test assertion failure). Root cause confirmed by a working counter-example already in the same file: the "#3912 A6" mocks gate on `fd !== 2` before any fault injection and were never implicated. Applied the same fd-scoped pattern to the five unscoped mocks (four output()-targeting tests gate on fd 1, one error()-targeting test gates on fd 2), and added a regression test proving an unrelated fd passes through untouched while the fault-injection mock is active. Found while verifying #3661; unrelated to that change's own diff. --------- Co-authored-by: sim <sim@local> |
||
|
|
6fdac3947b |
fix(#3996): carry agy stderr in the antigravity stub and gate the stall tell on the watermark (#4184)
* test(#3996): antigravity diagnostic must carry stderr and gate the stall tell * fix(#3996): carry agy stderr in the antigravity stub and gate the stall tell * fix(#3996): drop the stall token from the session-started sentence * fix(#3996): decide session-started from watermark growth, type the failure mode * chore(#3996): changeset for antigravity diagnostic stderr carry * chore(#3996): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
77dcdda534 |
enhance(#4014): an unreadable directory must not report as an empty one (#4163)
* test(#4014): add failing-first coverage for unreadable-vs-empty directory scope (epic #3473 B4) * fix(#4014): an unreadable directory must not report as an empty one (epic #3473 B4) * test(#4014): update hardcoded generateSlugInternal closing-brace line after import shift src/core-utils.cts's new #4014 import block shifted every subsequent line by 6, moving generateSlugInternal's real closing brace from line 193 to 199. tests/slug-derivation-drift-guard.test.cjs's MAJOR-1 fixture hardcodes that line number to plant a synthetic violation immediately after the function's real body; the guard script itself locates the boundary dynamically via brace-matching and needed no change. * docs(#4014): document the unreadable-directory scope signal and add changeset * docs(#4014): backfill changeset PR number to #4163 * test(#4014): kill pre-existing core-utils.cjs mutation-score gap, unrelated to this issue's diff --------- Co-authored-by: sim <sim@local> |
||
|
|
2131fe13f3 |
enhance(#3464): exec() detection widening, citation-debt cleanup — Phase 8 (#4171)
* feat(#3464): widen no-source-grep to detect regex.exec() on tracked text Adds an execCall kind alongside the existing regexTest detection -- regex.exec(tracked) was invisible to the rule while regex.test(tracked) was already caught, despite both reading a source-derived string through a regex. Measured: 4 previously-invisible sites across 2 files. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#3464): migrate 4 sites newly flagged by the exec() widening docs-hooks-table-parity.test.cjs's three regex-extraction loops are site-scoped marked (source-text-is-the-product) -- the dynamic preToolEvent/postToolEvent dialect branching they mirror is explicitly documented as not statically parseable, so a literal-pattern mirror is the practical minimum-cost check. no-bare-gsd-tools-command-position.test.cjs's readRouterVerbs() now requires HOST_COMMAND_ROUTERS directly instead of regex-walking gsd-tools.cjs's source text -- the same accessor three other suites already use. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3464): pay down 6 grandfathered uncited allow-test-rule markers Two were genuinely load-bearing (suppressing a real detected violation) and just needed a citation added -- phase6-capstone-conformance.test.cjs, runtime-name-policy.test.cjs, both now (#3464). Four were dead-weight file-header markers suppressing nothing -- each file's real effective sites are covered by separate, already-cited markers elsewhere in the same file. Deleted outright rather than cited, per Phase 1's own precedent (remove non-load-bearing markers instead of grandfathering them forever) -- codex-config.test.cjs (two copies), gsd-check-update-worker-platform-gate.test.cjs, orphaned-hooks.test.cjs, settings-jsonc.test.cjs. allowlist.json: 134 -> 128 entries. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3464): re-baseline effective-exemption ceiling to 84 The exec() widening's 3 newly-marked sites are now suppressed and counted; ceiling rises 81 -> 84, the exact measured high-water mark. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3464): correct citation and restore a wrongly-deleted marker Two review corrections, both found by the orthogonal review pass: - docs-hooks-table-parity.test.cjs's 3 new exec() markers cited #3464 (mechanically "the phase that widened the rule") when the file's own established, correct reference is #3839 (the issue this whole test exists to enforce, already cited in its file header) -- fixed to match. - gsd-check-update-worker-platform-gate.test.cjs's deleted file-header marker was NOT dead weight: its codeOnly() helper wraps readFileSync and is called inline as an assert argument, a genuine source-grep pattern on real .cjs/.js source that the rule cannot currently see (helper-function indirection is a distinct blind spot from anything Phase 7/8 measured) -- CONTRIBUTING.md is explicit that "unverified" is not the same as "vestigial." Restored, site-scoped this time (directly above codeOnly(), not as an inert file-header comment) and cited (#3103, the issue the file's own docstring already references). codex-config.test.cjs's two deletions and orphaned-hooks.test.cjs's / settings-jsonc.test.cjs's deletions were independently re-verified and stand: their flagged lines read generated .toml/.json OUTPUT, not source, or have no residual pattern at all. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9b77320580 |
fix(#4076): add missing gsd-hook-version header to gsd-node-runner.sh (#4092)
* fix(#4076): add missing gsd-hook-version header to gsd-node-runner.sh gsd-node-runner.sh was registered in MANAGED_HOOKS but shipped without a gsd-hook-version header, so gsd-check-update-worker.js always classified it as 'definitely stale' (a missing header is indistinguishable from a pre-version-tracking file). Every install on an otherwise up-to-date version showed a permanent, unclearable '⚠ stale hooks — run /gsd-update' warning naming this one file. Root cause: the build-hooks.js comment claimed the file is 'not a registered hook' and 'staged verbatim — no templating', but it IS in MANAGED_HOOKS (managed-hooks-registry.cjs:34) and install.js already stamps {{GSD_VERSION}} into every .sh hook unconditionally, gsd-node-runner.sh included. The comment contradicted both the registry and the installer's actual behavior, and the header line itself was simply never added. Fix: add the header (matching every other managed .sh hook's format) and correct the comment so it no longer asserts the opposite of what the registry and installer actually do. Adds a regression test that iterates every MANAGED_HOOKS entry and asserts it carries a header matching the worker's own detection regex, so a future hook added to the registry without one fails CI instead of shipping silently. Fixes #4076 * chore(#4076): add changeset fragment for PR #4092 * fix(#4076): address review nits — drop unneeded exemption, fix blank line Per @trek-e's review on #4092: - tests/managed-hooks.test.cjs:96: the readFileSync call uses a loop variable (entry-derived hookPath), not a literal path, so local/no-source-grep's static literal-path detector never flags it — the allow-test-rule exemption comment was unnecessary. Replaced with a plain note explaining the source-read rationale. - tests/managed-hooks.test.cjs:121-122: dropped a stray extra blank line before the bug #2136 section divider. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
b0572c0108 |
feat(#3674): extract shared file-overlap wave partitioner (#4166)
* test(#3674): characterize existing wave-dispatch output and add tests for the extracted partitioner Pins resolveWaveDispatch's and emitWorkflowScript's current, unextracted output (chain-overlap, disjoint-empty-set, and a multi-wave/multi-stage golden script) as a regression safety net ahead of extracting partitionStages into a standalone module. Also adds the new module's unit and property tests (test matrix rows 1-11) against its expected public API, which does not exist yet and is added in the next commit. * feat(#3674): extract file-overlap partitioner into a shared, generic module Moves partitionStages' greedy first-fit file-overlap algorithm into a new, dependency-free src/file-overlap-partitioner.cts module (partitionByFileOverlap), generalized over a plain {id, files}[] shape rather than claude-orchestration.cts's Plan/Wave interfaces. partitionStages becomes a thin adapter mapping its own Plan[] shape onto the generic input and back — behavior-preserving, no dependency ordering, no path normalization, no filesystem access moved or added. Enables a future consumer (quick-batch, #3675 / ADR-1239) to reuse the same primitive without pulling in orchestration internals. * docs(#3674): register the file-overlap-partitioner module bookkeeping New src/*.cts -> bin/lib/*.cjs modules need four hand-maintained registrations beyond the code itself: .gitignore (compiled artifact), eslint.config.mjs (ADR-457: lint the .cts, not the emitted .cjs), docs/INVENTORY.md's CLI Modules roster row (regenerated via gen-inventory-manifest.cjs --write), and a CONTEXT.md glossary entry matching the convention set by similarly-scoped leaf modules (text-lines.cts, plan-dependency-graph.cts, spec-section.cts). * fix(#3674): alphabetize INVENTORY.md row, manifest regen no-op, fast-check import already correct - docs/INVENTORY.md: move file-overlap-partitioner.cjs row to alphabetical position - docs/INVENTORY-MANIFEST.json: regenerated via gen-inventory-manifest.cjs --write, produced no diff (manifest is keyed by content, not row order) - tests/claude-orchestration.test.cjs's direct require('fast-check') is correct as-is: tests/helpers/fast-check-setup.cjs's own docstring scopes the shared-seed wrapper to "every *.property.test.cjs file"; claude-orchestration.test.cjs is not a .property.test.cjs file, and every .property.test.cjs file sampled uses the wrapper consistently. No outlier. * fix(#3674): constrain the no-overlap property test to unique ids, fixing an ambiguous duplicate-id reconstruction The `no two plans in the same stage share a modified file` property reconstructs which physical item produced each output id via `remaining.findIndex(r => r.id === id)`. Under duplicate ids (an explicitly-supported input shape for `partitionByFileOverlap`) that reconstruction can pick the wrong physical occurrence, producing a false-positive overlap failure (observed counterexample: p0(f1), p208(f1), p208([]) — correctly staged as [[p0,p208#2],[p208#1]], but misread by id-order as [[p0,p208#1],...], which do overlap). Properties (a) determinism and (b) totality already exercise duplicate ids correctly and are left unchanged; only this property's generated items are now constrained to unique ids via `fc.uniqueArray`, where the reconstruction is unambiguous. --------- Co-authored-by: sim <sim@local> |
||
|
|
fa107c0461 |
fix(#4172): add --merge-async to test:coverage:scripts-floor (#4173)
* fix(#4172): add --merge-async to test:coverage:scripts-floor The "Coverage gate (merged shards)" test.yml job has OOM-crashed (exit 134, SIGABRT) on every push to next since |
||
|
|
5c7243e54b |
fix(#3995): derive the review diff base from the phase directory (#4181)
* test(#3995): diff base keys on the phase directory, not commit subjects All three derivation sites (Tier 3, spawn_reviewer, fallow pre-pass) must anchor on the phase directory's first commit; the milestone-blind repro (an archived milestone's same-numbered phase commit capturing the base) is the failing-first row. #3191/#3503 rows reworked to the directory-anchor contract; T6 docs-parity forbids any remaining phase-scope message-grep site. * fix(#3995): derive the review diff base from the phase directory A phase number is unique within a milestone, not a repository; the message grep had no milestone bound and tail -1 deliberately selected the oldest same-numbered subject, dragging archived milestones phases into the scope (7 files to 3388 plus the >50 depth downgrade). All three lockstep sites now anchor on the first commit that added anything under the phase own directory — the same anchor class git-base-branch phaseStartCommit uses. ShellCheck baseline gains the escaped fragment shifted parse signature. Emitted-Drift-Ack-Growth: code-review.md — phase-directory anchor replaces the message-grep derivation at both sites (#3995) * chore(#3995): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
647365faf1 |
fix(#4011): key the TDD runtime gate on TDD_MODE alone (#4180)
* test(#4011): TDD gate keys on TDD_MODE alone, not the MVP intersection Contract updates: no shipped line may conjoin MVP_MODE with TDD_MODE as a gate condition, the end-of-phase escalation must not require MVP, the executor agent's gate section triggers on TDD_MODE alone, and the gate semantics reference loads without MVP_MODE. * fix(#4011): key the TDD runtime gate on TDD_MODE alone The RED-commit gate shipped as #76's MVP slice kept the paired invocation's conjunct, so workflow.tdd_mode=true was silently inert on every non-MVP phase, contradicting references/tdd.md's own contract. Drops the MVP conjunct from the per-task gate and the end-of-phase review escalation; rescopes execute-mvp-tdd.md's load condition, gsd-executor's gate section, and mvp-concepts' intersection claim. MVP remains free to imply TDD; the file is not renamed (stated assumption in the PR body). * test(#4011): scope no-conjunct detector to shell conditions; clean stale MVP+TDD phrasing Review follow-ups: the detector now only inspects if/[ condition lines so explanatory prose mentioning both flags cannot trip it; remaining 'under/outside MVP+TDD' phrases in execute-phase.md, the gate reference, and docs/INVENTORY.md now describe TDD-mode semantics. Emitted-Drift-Ack-Growth: execute-phase.md — TDD-gate decoupling comment + escalation rescoping (#4011) Emitted-Drift-Ack-Growth: gsd-executor.md — gate section trigger rescoped to TDD_MODE alone (#4011) * chore(#4011): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
383c2f6b34 |
fix(#3982): strip closed-milestone details from the current window (#4177)
* test(#3982): archived details must not leak into the current-milestone window Parser-level regression (newest-first layout, two archived details blocks) plus the issue's end-to-end phase.complete fixture: completing phase 20 must advance to 21, never backwards into the archived range. * fix(#3982): strip closed-milestone details from the current window The heading-located window stripped <details> archives from the preamble but not from currentSection; on newest-first roadmaps the archived titles sit in summary tags rather than headings, so the section walk reached end-of-document and the window swallowed every collapsed archive below the active milestone. The strip is gated on isClosedMilestoneHeading over each block's summary — the issue's prescribed narrow form — so the active milestone's own collapsed blocks (#1341) survive instead of trading this bug for the phase_count: 0 class (#557/#2947). * chore(#3982): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
9eef1b791b |
fix(#3981): give blocking guards a host-stall-proof timeout budget (#4175)
* test(#3981): blocking-guard timeout budget and 5→120 migration Fresh registration must carry a host-stall-proof 120 s budget on the six blocking PreToolUse guards; existing managed timeout:5 entries are migrated; non-managed entries and advisory budgets are untouched. * fix(#3981): give blocking guards a host-stall-proof timeout budget Claude Code treats a timed-out hook as non-blocking, so the 5 s budget on the six blocking PreToolUse guards silently dropped the gates exactly when the host stalled under load. Registers them at 120 s (covers every observed stall, max 84.3 s) and migrates existing managed timeout:5 entries in place, context-monitor-backfill shape. Advisory hook budgets are unchanged. * chore(#3981): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
7960374d15 |
fix(#3962): rename the TDD-Audit trailer token to gate-status (#4174)
* test(#3962): shipped TDD-Audit trailer token must round-trip through git Behavioral coverage: extract the trailers:key token from ship.md and prove a real git commit carrying that trailer reads back via %(trailers:key=token,valueonly). gate_status contains an underscore, which git's trailer machinery cannot tokenize, so the audit read was structurally empty. * fix(#3962): rename the TDD-Audit trailer token to gate-status Underscore is not a valid git trailer token character, so %(trailers:key=gate_status,...) could never match. Renames the token at the read, the documented aggregate write, and the section's prose/table header. Self-suppression semantics (#2431) unchanged. * test(#3962): compare outcome against the seam's literal, fix doc token Review follow-ups: the round-trip test compared outcome to 'EXITED' but the process seam emits 'exited'; and the trailer rename is propagated to docs/ship-pr-body-sections.md's read/write examples. * chore(#3962): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
10593e2cf2 |
fix(#3959): give review-prompt plans citable path anchors (#4170)
* test(#3959): plan copies carry source names, prompt carries path anchors Regression coverage: budget copies named gsd-review-plan-<plan-id>.md (not a bare padded index), no bare-index copy survives, and the Plans to Review template instructs a per-plan repo-relative #### path header. Also corrects the stale -00 assertion to the source-named copy. * fix(#3959): give review-prompt plans citable path anchors The Plans to Review template now instructs a per-plan repo-relative #### path header, and the budget copies are named gsd-review-plan-<plan-id>.md instead of a bare padded index, restoring provenance for prompt-budget's per-plan headers while keeping the trim glob intact. Emitted-Drift-Ack-Growth: review.md — per-plan path-header instruction in Plans to Review + provenance-named copy loop (#3959) * chore(#3959): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
91d5fdff6f |
chore(#3546): migrate hook advisory assertions onto typed output surfaces (#4167)
* chore(#3546): migrate hook advisory assertions onto typed output surfaces Add additive typed fields to 5 hook scripts' PreToolUse/PostToolUse advisory output alongside the existing additionalContext prose: - gsd-read-guard.js: code ('READ_BEFORE_EDIT'), fileName - gsd-context-monitor.js: severity ('warning'|'critical') - gsd-prompt-guard.js: findings ([{ruleId, match}], module-local RULE_IDS + renderFinding mapper mirroring gsd-read-injection-scanner.js's #3523 pattern) - gsd-read-injection-scanner.js: severity ('LOW'|'HIGH'), source (its findings array already existed from #3523) - gsd-workflow-guard.js: code ('WORKFLOW_ADVISORY') on the advisory leg, distinct from the existing force-add block leg's code additionalContext stays byte-identical in every hook (verified per-hook against the pristine HEAD version across a spread of payload shapes). Migrates all 20 assertion sites named in the issue off additionalContext.includes(...)/assert.match(...) substring-matching onto the new typed fields, per CONTRIBUTING.md's prohibition on raw text matching on test outputs. Closes #3546 * test: fix undersized commit-class timeout in gsd-statusline.test.cjs's commitN helper Surfaced by gsd-test on the #3546 checkpoint: `commitN()`'s loop called gitOrThrow(['add','-A']/['commit',...]) without a timeoutMs override, so each call used DEFAULT_GIT_TIMEOUT_MS (15s) -- a bound git-fixture.cjs's own doc comment says is sized for plumbing reads (rev-parse/branch/log), not write-heavy add/commit spawns. That file already documents the exact same defect class from a prior incident (PR #3323) and exports GIT_FIXTURE_TIMEOUT_MS (60s) for fixture-construction call sites - commitN just wasn't using it. Observed failure: `git commit -m filler 9` timed out under normal bench load, unrelated to any of this PR's own diff (hooks/*.js + 5 other test files). Not a flake: root-caused to the timeout bound being sized for the wrong call class, per this repo's no-flakes rule. * chore(#3546): backfill changeset PR number (#4167) --------- Co-authored-by: sim <sim@local> |
||
|
|
f16ff7d1b3 |
enhance(#3545): widen no-source-grep with fold+hooks, migrate 76 sites (#4161)
* feat(#3545): widen no-source-grep with one-hop path-fold and hooks dir Resolve a readFileSync() path argument that is a bare Identifier one hop back to its VariableDeclarator initializer before classification, and recognize `hooks` as a source directory alongside bin/lib/gsd-core/src. Measured (epic #3464 phase 7): fold+hooks together newly flag 76 unsuppressed sites across 18 files that were previously invisible to identifier-indirected or hooks/-rooted source reads. Neither widening alone is sufficient — hooks-only surfaces 0 new sites, confirming #3520's prior finding that the identifier-indirection gap must close first. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#3545): migrate 76 sites newly flagged by the fold+hooks widening Per-site classification: rewrite behaviorally (require() the real module, assert on its actual exported behavior) wherever the read was a proxy for code behavior; add a site-scoped `// allow-test-rule: <reason> (#3545)` marker only where the raw source text genuinely is the product under test (codex-config.test.cjs's adapter-header-contract checks, install.js structural-wiring guards with no exported symbol, AST-parse fixture inputs, etc.) — each marker cites an existing repo-sanctioned category from CONTRIBUTING.md's allow-test-rule exception table. Also converts two try/finally test bodies (introduced during this same migration) to the required t.after() cleanup pattern per CONTRIBUTING.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3545): re-baseline effective-exemption ceiling to 81 The fold+hooks widening's own newly-detected sites are now suppressed by site-scoped markers, moving them from invisible into the tightly-ratcheted effective-exemption count. Ceiling rises from 10 to 81 (the exact measured high-water mark, grace unchanged at 2) — a deliberate, measured re-baseline per the widening working as intended, not an ordinary ceiling bump. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3545): use canonical allow-test-rule category tokens 4 markers added during migration cited an issue ref correctly but didn't use one of CONTRIBUTING.md's seven recognized category tokens, unlike every other marker in this change. Cosmetic only — same suppression lines, same effective/live counts (81/81, 0 live). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3545): correct stale phase-artifact path in test comment Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ff0361071d |
feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane (#4160)
* feat(#3653): add review.models.cursor — wire modelArg/modelConfigKey for the cursor reviewer lane cursor-agent exposes --model (204 selectable models) but the cursor lane declared modelArg: null / modelConfigKey: null, so review.models.cursor was rejected as an unknown config key and the #1517 reviewer-instances escape hatch silently discarded a configured model at modelExpansion. Wire the lane the same way codex already is: inject {{model}} into args right after -p, set modelArg to --model, and declare modelConfigKey as review.models.cursor plus its config schema entry. An unconfigured lane still invokes byte-identically to today. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3653): add changeset for review.models.cursor Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3653): update co-change surfaces that assumed cursor has no model key gsd-test surfaced three surfaces still hardcoding "cursor declares no modelConfigKey", broken by wiring review.models.cursor: - tests/reviewer-config-federation.test.cjs: the #3691-narrows-#2797 invariant test listed cursor among lanes that must own no model key. - tests/settings-integrations.test.cjs: the #3651 keyless-lane test listed cursor as keyless, including a live config-set assertion that now correctly succeeds instead of failing (swapped to qwen). - gsd-core/workflows/settings-integrations.md: the settable-keys enumeration and two prose call-outs still named cursor as keyless. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3653): acknowledge deliberate growth of settings-integrations.md settings-integrations.md grew 4 bytes because it now enumerates review.models.cursor as a settable key alongside the other reviewer lanes, matching the modelConfigKey wired for cursor in this PR. Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(#3653): fix malformed Emitted-Drift-Ack-Growth trailer The previous commit's trailer was separated from Co-Authored-By by a blank line, splitting it into an earlier, non-trailer paragraph — git's trailer parser only recognizes the last contiguous block. Restating it here immediately adjacent to Co-Authored-By so both parse as trailers. Emitted-Drift-Ack-Growth: settings-integrations.md — adds review.models.cursor to the settable-keys enumeration and removes cursor from the two keyless-lane call-outs, matching #3653's modelConfigKey change Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3653): backfill changeset pr number pr:0 -> pr:4160 now that https://github.com/open-gsd/gsd-core/pull/4160 exists. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |