423f38e655195cddfd4288a098a31d0f86df5502
5726 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
423f38e655 |
fix(#4444): honor config-set --dry-run instead of silently ignoring it (#4504)
* test(#4444): failing-first regression coverage for config-set --dry-run config-set --dry-run is currently parsed nowhere -- routeConfigSet (gsd-core/bin/gsd-tools.cjs) never checks args for it, and cmdConfigSet has no dry-run parameter, so the flag is silently swallowed and the command always writes for real. Reproduces the issue's own repro (sequential --dry-run calls where the second's previousValue proves the first persisted), plus coverage for validation-still-runs, secret-masking, and the sibling unset (config-set <key> null) branch, which has the identical defect. This commit adds the regression coverage only; the fix lands in the next commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4444): honor config-set --dry-run instead of silently ignoring it routeConfigSet (gsd-core/bin/gsd-tools.cjs) never read args for --dry-run, and cmdConfigSet had no dry-run parameter at all -- so the flag was silently accepted (as any unrecognized trailing argument is) and the command always wrote for real. A second "dry run" then showed previousValue reflecting the first one, proving it had persisted. Threads a dryRun option through cmdConfigSet, gating BOTH mutating branches: the null/unset path (unsetConfigValue) and the real-set path (setConfigValue) -- the unset branch had the identical defect, undiscovered until auditing every mutation site while designing this fix. Each gains a previewConfigValue/previewUnsetConfigValue counterpart that reuses the real function's exact traversal/creation logic (_setNestedValue/_unsetNestedValue) on a throwaway in-memory config copy that is never written -- so the preview can never diverge from what the real write would compute. All validation (unknown key, enum/number/boolean checks, secret masking) runs identically whether or not --dry-run is passed; only the final write is skipped, replaced with a `{ dry_run: true, would_update / would_unset: true, ... }` preview payload matching the precedent established by `milestone complete --dry-run` (#2118) and `todo complete --dry-run` (#4096/#4325). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * refactor(#4444): extract loadConfigJson to stop a 5th copy-paste of the same load/parse block Code review flagged that setConfigValue, unsetConfigValue, setConfigValues, and the two new preview functions each repeated the identical "load .planning/config.json, JSON.parse, catch -> CONFIG_PARSE_FAILED" block -- exactly CLAUDE.md's own "Generative Fix Divergence" known-defect pattern. Extracted a single loadConfigJson(cwd) helper; behavior is unchanged (verified: build, tsc, and the dry-run/real-write smoke test all pass byte-identical to before). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4444): changeset for the config-set --dry-run fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4444): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4444): raise per-chunk CI test timeout to 800s for Windows headroom install-minimal-hooks.test.cjs (weight=24.45, the heaviest file in the suite) sits alone in its own chunk yet still occasionally brushed the 600000ms per-chunk ceiling on Windows -- observed on PR #4504's first CI run for this change (passed clean on rerun, consistent with the "legitimately too slow for the budget" cause the chunk-timeout diagnostic already names, not a leaked handle). Raised RUN_TESTS_CHUNK_TIMEOUT_MS's default from 600000ms to 800000ms: ~33% more margin, still comfortably below the 900000ms regen:derived fixture timeout that fragment-single-edit-propagation.install.test.cjs deliberately keeps ABOVE the chunk ceiling, and far under the 45-minute job cap -- Windows shards currently finish in ~19-20 minutes total, so there is ample headroom. Updated every dependent mirror/assertion in lockstep (tests/helpers/emitted-runtime.cjs's duplicated CHUNK_TIMEOUT_CEILING_MS constant, its lock test in tests/emitted-attribution.test.cjs, the Windows-skip prose in fragment-single-edit-propagation.install.test.cjs, and docs/TESTING-SUITES.md's reference table) so nothing describes a stale value. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * Revert "fix(#4444): raise per-chunk CI test timeout to 800s for Windows headroom" This reverts commit 394aadaf6f5af6fd700bf0f444c9fbd686285a4f. * test(#4444): consolidate redundant installer spawns in install-minimal-hooks.test.cjs This file's real, unrelated pre-existing cost (dated 2026-09-06, PR #4428) is what tipped a Windows CI shard over the per-chunk timeout backstop on PR #4504 (issue #4444's own diff never touches this file or the installer). Rather than raise the timeout, cut the file's actual spawn count: several describe blocks independently re-installed the IDENTICAL runtime/scope/flag configuration just to assert different things about the same install output. Merged each such group onto a single shared install, with every original assertion preserved: - --help x3 -> x1 - the three per-runtime/scope --minimal E2E loops (global, local, and on-disk-matches-manifest) merged into one loop over SKILL_RUNTIMES x [global, local]: 44 spawns -> 22 - the --minimal manifest-mode/backcompat triple-install -> one shared, memoized install via sharedMinimalManifestInstall() - .sh hooks existence checks (5 tests) -> 1, executable-bit check (its own Windows-conditional skip) left separate - Codex #4087 hook-helper tests (3) -> 1 - Windsurf #4087 hook-helper tests (2) -> 1 - pi shared-hooks-bundle tests (3 per scope) -> 1 per scope Net: ~65 real installer spawns in this file down to ~29, no assertion dropped or weakened. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
93e141a006 |
enhance(#4139): Phase 4 — measure the window instead of asserting it (#4502)
* enhance(#4404): add offline token benchmark for compact-content splits ADR-4139 Decision 2 requires the finite-attention justification for workflow.compact_content to be measured, not asserted. `npm run benchmark:compact-content` computes, per registered spine/detail split discovered under gsd-core/workflows/, the token count with the split active (spine alone) vs inactive (spine + all detail parts read back in), using gpt-tokenizer (pinned exact devDependency — Anthropic publishes no tokenizer for Claude 3+, so every output surface labels this a PROXY-TOKENIZER comparison: the on/off delta is exact under one tokenizer applied identically to both sides, the absolute counts are not Claude's real ones). Reporting-only by design and verified so: --check diffs the live recompute against a committed baseline (tests/fixtures/compact-content-benchmark-baseline.json) and prints drift, but never exits non-zero for a drifted or missing baseline — the only thing allowed to fail this script is a genuine I/O error reading a source .md file it's measuring. Not wired into lint:ci or pretest. Discovery is deliberately reimplemented rather than importing tests/helpers/compact-content-split.cjs (Phase 3, #4403), keeping a scripts/ reporting tool from depending on a test-only module. tests/fixtures/deny-network.cjs preloads via NODE_OPTIONS=--require to prove the benchmark makes no network call, monkeypatching http/https/ net/dns/fetch to throw rather than relying on sandboxing. docs/CONFIGURATION.md documents the new benchmark against the workflow.compact_content key to satisfy this repo's docs-required gate for an Added-type changeset. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4404): address orthogonal review findings on the token benchmark Standards axis found two hard violations against documented rules: - CLAUDE.md's Generative Fix Divergence rule requires a parity assertion for shared discovery logic maintained in two places. Added a test comparing benchmark-compact-content.cjs's own discoverRegisteredSplits against tests/helpers/compact-content-split.cjs's version on the real repo tree, so the two can never silently drift apart. - The changeset body closed its bold span with a period and continued as a second sentence, instead of the canonical `**<phrase>** — <explanation>.` shape CONTRIBUTING.md documents. Spec axis found the "network disabled + identical output across two runs" Done-when criterion was verified as two separate properties (determinism tested without network denial, offline survival tested as a single run) rather than as one combined property. Added a test that runs the benchmark twice under the deny-network preload and asserts byte-identical stdout. Security axis found tests/fixtures/deny-network.cjs didn't patch dns.promises (a separate binding from the callback dns API), tls.connect, or http2.connect — inert today since nothing in the benchmark calls them, but a silent gap in what the preload's own header claims to guarantee. Patched all three. CLAUDE.md's Property-Based Testing rule also requires a fast-check test for budget-limit arithmetic; added one for computeAggregate's off/on summation (true sum over N splits, never NaN/Infinity, never exceeds 100% when off >= on for every split). Standards axis's remaining two findings (a Data Clumps observation on the {offTokens, onTokens, reductionPct} triple, and mild duplication in formatDriftReport's three line-formatters) are left as judgement calls: introducing a named type for a 3-field local tuple, or a formatter abstraction for three short lines, would be exactly the premature abstraction CLAUDE.md's engineering guidance warns against for a script this size. All changes verified directly (parity logic, the fast-check property, and the three newly-denied network surfaces actually throwing under the preload) via node -e before committing; full npm run lint:ci passes with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4404): backfill changeset pr number to 4502 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8dcdcb253e |
fix(#4443): register hooks.commit_types (and sibling hooks.community) in config schema (#4501)
* test(#4443): failing-first regression coverage for hooks.commit_types config key isValidConfigKey('hooks.commit_types') currently returns false and config-set hooks.commit_types rejects with "Unknown config key", because the key was never added to config-schema.manifest.json's validKeys when it shipped (#3811/#4340, 1.13.0) despite being documented (docs/COMMANDS.md) and consumed by hooks/gsd-validate-commit.sh. This commit adds the regression coverage only; the manifest fix lands in the next commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4443): avoid false-positive docs-guard registration trip The assert message for the new hooks.commit_types test mentioned "docs/COMMANDS.md" literally, which happened to land between two unrelated pre-existing backticks and tripped lint-docs-guard-registration.cjs's template-literal co-occurrence detector (a known, documented false-positive shape for that lint). Rephrased to drop the literal docs/ path from the message; the test's intent (documenting why the key must be valid) is unchanged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4443): register hooks.commit_types (and sibling hooks.community) in config schema config-schema.manifest.json's validKeys never got hooks.commit_types added when the feature shipped (#3811/#4340, 1.13.0) despite it being documented (docs/COMMANDS.md) and consumed by hooks/gsd-validate-commit.sh -- so config-set hooks.commit_types rejected with "Unknown config key", and the only way to configure a documented feature was hand-editing .planning/config.json. While auditing every hooks.* key actually read by shipped code against validKeys (CLAUDE.md's no-deferrals rule: a defect found anywhere in the tree while working an issue is fixed in the current change, not filed separately), hooks.community -- gsd-validate-commit.sh's own opt-in gate -- turned out to have the exact same gap. Both are added here; an audit of every hooks.* read site confirmed these are the only two missing entries. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4443): e2e coverage for hooks.community + changeset Closes the coverage-rigor gap the Standards review flagged: hooks.community had only a unit-level isValidConfigKey assertion, not the same real config-set CLI round-trip hooks.commit_types already got. Also adds the changeset fragment the same review flagged as a missing hard requirement. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4443): use PROBE_TIMEOUT_MS instead of a bare 15000 literal local/no-adhoc-timeout-literal (lint:ci) correctly flagged both new spawnSync calls' bare timeout: 15000 -- this call class (a short CLI probe against a temp fixture) is exactly what tests/helpers/timeouts.cjs's PROBE_TIMEOUT_MS documents. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4443): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a0f8f956c4 |
enhance(#4139): Phase 3 — partition rules + the five checks (#4497)
* enhance(#4139): Phase 3 — partition rules + the five checks ADR-4139 Decision 5, epic #4139 Phase 3. Issue #4403's own "Proposed behavior" section lists four checks; the ADR's Decision 5 and its own phase table ("partition rules + the five checks") list five — the same four plus "boundary moves are declared, ongoing". Same issue-vs-ADR drift Phase 2 hit on the detail.md vs detail/*.md layout: the ADR is the locked, reviewed document, so it wins. This PR implements all five. docs/PARTITION-RULES.md (new) is the partition-rules document: the partition rule itself, the protected-content list and <!-- gsd:protected --> sentinel syntax (relocated unchanged from gsd-core/references/compact-content-protected-content.md, now deleted — it was never referenced by any runtime workflow Read, only by the predecessor test as documentation, so nothing at runtime regresses, and removing it from gsd-core/references/ also drops it from all 19 installed-project shipped-content trees for a file nothing ever read), and the five checks explained for a human reader. Referenced from a new CONTRIBUTING.md subsection under "Editing shipped content". tests/helpers/compact-content-split.cjs (new) is the shared mechanics: split discovery (any gsd-core/workflows/<name>/detail/*.md paired with <name>.md — no registry file, a pair is registered by existing on disk), line normalization (carries forward Phase 2's bare-label-line isTrivial fix and the canonical gsd_run-launcher-preamble exclusion), sentinel extraction, and a Boundary-Move-Declared commit-trailer reader that is a direct structural port of tests/helpers/emitted-runtime.cjs's Emitted-Drift-Ack-Hash/-Growth trailer reader (ADR-3942) — same merge-base range, same fail-closed throw on an uncomputable range, same dedupe/conflict rules. tests/compact-content-partition-guard.test.cjs (new) is the actual guard, superseding tests/plan-phase-compact-split.test.cjs (deleted — its per-pair checks are now the general guard's job for plan-phase specifically). Checks 2 (disjointness) and 3 (registration + size cap) run unconditionally against every registered split. Checks 1 (completeness, fires once per split on the PR that introduces a new detail/ path), 4 (protected content — no trailer can ever excuse this one, unlike check 5) and 5 (boundary moves declared) are PR-diff-scoped against the resolved base ref and skip cleanly when there's nothing to compare (a fresh clone, no PR in flight) — a deliberate asymmetry from check 5's trailer reader, which must throw rather than silently pass when ITS range is uncomputable, since that function is answering "did this PR declare its moves" rather than "is there even a diff to look at". Each of the five checks carries a RED (deliberately broken fixture) / GREEN (fixed) test pair, built against synthetic temp files or real throwaway git repos, per this repo's rule that a guard nobody has seen go red is not yet a guard. Building the real fixtures caught and fixed one real bug before it shipped: check 4's line-presence test was using the trivial-line-filtered normalizer, so a byte-identical spine falsely reported its own protected code-fence line as "deleted" — fixed with a non-filtering membership check. Extending docs/INVENTORY.md's "Workflow Sub-Files" table for `detail` surfaced a pre-existing, unrelated gap in the SAME area: gsd-core/workflows/<name>/templates/*.md is a fourth workflow sub-file kind that already existed on disk and was already known to lint-response-language-coverage.cjs's FRAGMENT_DIRS, but was invisible to gen-inventory-manifest.cjs and undocumented in that table. Fixed alongside it, same pattern, same PR, rather than deferred. Also, mechanically required by the new fourth sub-file kind: - scripts/lint-response-language-coverage.cjs: `detail` added to FRAGMENT_DIRS alongside modes/steps/templates — a detail/<part>.md inherits its parent's response_language coverage through the same per-file proof, not a parallel one. - tests/workflow-size-budget.test.cjs: explicit regression test locking that detail/ files are governed solely by the hard, non-waivable NEW_FILE_CAP (tests/helpers/emitted-diff.cjs) and never by the XL/LARGE/DEFAULT spine tiers — true by construction (measureWorkflows/listWorkflowStems don't recurse), made explicit per the issue's own Done-when item rather than left true-by-omission. - scripts/gen-inventory-manifest.cjs: `workflow_detail` and `workflow_templates` NESTED_FAMILIES entries; docs/INVENTORY-MANIFEST.json regenerated (plan-phase/detail/elaboration.md, discuss-phase/templates/*.md now tracked); docs/INVENTORY.md's table updated to four kinds. Verified: `npm run lint:ci` clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4403): review findings + a real gsd-test failure in the new guard Two orthogonal review passes (Standards + Spec, isolated sub-agents) plus a separate security review ran against the prior commit. Fixed everything each surfaced: - Security (Low, path-traversal existence oracle): checkRegistration's dangling-reference check extracted detail-path-shaped substrings from spine PROSE via a regex that permits `.`/`/` freely, then joined them onto repoRoot and probed fs.existsSync with no containment check — a spine file containing `../../../etc/detail/passwd.md`-shaped text could make the guard test file existence outside the repo. Added a path.relative-based containment check before the fs.existsSync call; anything that resolves outside repoRoot is now reported as a dangling reference directly, never probed on disk. - Standards (Boundary Coverage): the size-cap fixtures covered NEW_FILE_CAP and NEW_FILE_CAP-1 but not NEW_FILE_CAP+1 — added the third boundary-point case CLAUDE.md's TEST RULES require (limit-1/limit/limit+1). - Standards (Property-Based Testing): extractProtectedBlocks (a sentinel parser) and the new parseBoundaryMoveTrailerValues (a declare/dedupe/conflict parser, bijective-shaped) had no fast-check property test. Added three: a render/parse bijectivity property for the trailer parser (mirroring the exact ADR-3942 sibling test's alphabet/idiom), a dedupe-is-idempotent property for the same parser, and a well-formed-sentinel-round-trips property for extractProtectedBlocks. Then dispatched gsd-test on the resulting commit. It found a real bug the reviews couldn't have caught (none of them can run inside gsd-test's sandbox): checks 4/5's real-repo assertion failed against plan-phase's own split, reporting DISK_PLANS/#3218-comment lines as "undeclared boundary moves" — content Phase 2 (#4402) legitimately moved into detail/elaboration.md months before this PR's Boundary-Move-Declared mechanism existed to require a trailer for it. Root cause: `resolveBase()`'s own doc comment already documents that no `origin/*` remote-tracking ref exists inside the gsd-test sandbox container, and its fallback candidate (a bare `next` branch) can resolve to a point in history that predates an already-merged, already-reviewed split — making that split look "newly introduced" from the sandbox's vantage point. Check 1 (completeness) already scopes itself correctly to only genuinely-new detail paths (git diff status 'A'); checks 4 and 5 did not share that scoping, so a stale base made them re-litigate a settled split retroactively. Fixed by having checks 4/5 skip any split name check 1 already counted as newly-split — their own premise ("did an EXISTING split shed/undeclare something") does not apply to a split that is, from the resolved base's vantage point, brand new; that is check 1's domain alone. Verified locally (25/25 tests pass via a direct `node -e` require, since `node --test` is blocked in this repo) and via re-reasoning through the exact real-repo scenario the gsd-test failure showed. Also regenerated all 19 tests/fixtures/install-tree/*.json goldens — the prior commit's deletion of gsd-core/references/compact-content-protected-content.md was never reflected there, which is what golden-install-tree.test.cjs's other 19 failures in the same gsd-test run were. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4403): backfill changeset pr number to 4497 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4403): isolate codex-config.test.cjs into its own chunk, root-causing the Windows CI failure PR #4497's "full test (windows-latest, 24, shard 2/3)" job failed: run-tests killed chunk 3/8 at the 600s per-chunk backstop, with codex-config.test.cjs (weight 17.87, by far the chunk's dominant cost) packed alongside 39 other files. Traced, not assumed: - scripts/run-tests.cjs's own timeout-headroom comment for the OUTER per-shard timeout documents that "adding one test file reshuffled 115 of 268 unit files between shards" — shard/chunk composition is architecturally known to be unstable to single-file additions, which is exactly what this PR's own new tests/compact-content-partition-guard.test.cjs is. - A second comment, dated 2026-09-06 (one day before this PR, PR #4428's own CI), already documents the SAME chunk hitting the SAME 600s backstop with the SAME file (codex-config.test.cjs, "a genuinely MEASURED weight of 17.87 — not a stale-table miss") dominating it — the fix then was cutting the Windows per-chunk budget from 60 to 40. That cut clearly was not enough: two documented incidents in two days, at two different budget settings, both centered on one file that alone consumes ~45% of even the reduced Windows budget. - tests/test-timings.json's own header confirms its source data (test-events-linux-node22/24.jsonl) is Linux-only, and run-tests.cjs's own chunk-timeout diagnostic already prints "real Windows cost runs ~2.2x the recorded figure" — the packer's weight-balancing is working off data that is both stale (table last regenerated 2026-08-07) and known to underestimate the platform where the failure occurs. Given codex-config.test.cjs is disproportionately heavy AND every companion sharing its chunk is decided by a packing algorithm already documented as reshuffling unpredictably on any new file, tuning the shared budget a third time only moves the marginal line to wherever the next new file happens to land — it does not remove the gamble. Isolating codex-config.test.cjs into its own dedicated single-file chunk, unconditionally and on every platform, removes it at the source: the file never enters the pool packChunks balances, so no other file's packing changes, and no future single-file addition (mine or anyone else's) can silently reintroduce this exact failure by landing in its chunk. Extracted as a small pure function, partitionIsolatedFiles (mirroring this file's existing pattern of pulling packing/analysis logic out of main() for in-process unit coverage — see computeSweepProtectSet, analyzeChunkEvents), with 6 new tests in tests/run-tests-harness.test.cjs covering basename matching across path separators, near-miss non-matches, the empty-list case, and the isolated-set contents. Root cause is now closed rather than papered over with a retry: this failure is a property of one specific heavy file's chunk placement, not something that recurs randomly. If codex-config.test.cjs itself is ever genuinely sped up, this isolation can be revisited — this is a packing-side mitigation for a known file's cost, not a claim the cost is irreducible. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ac6ed6201d |
fix(#4257): harvest only prose phase references; W002 names its workstream scope (#4486)
* test(#4257): W002 harvest precision + workstream-scoped warning regression rows Tests-only RED commit: A-rows pin the command-mention/code-span harvest precision on statePhaseTokens, B-rows drive W002 under root and workstream scope (scope clause asserted, root grammar byte-identical), C-rows pin the additive workstream snapshot field. All fail on next; fix follows. * fix(#4257): harvest only prose phase references; W002 names its workstream scope Sub-defect (a): the statePhaseTokens harvest was the verbatim #3309 relocation of verify.cts's unanchored, markdown-blind scan ([Pp]hase\s+(TOKEN) over the raw file), so GSD's own command names (/gsd-execute-phase 5, bare or quoted) and any token inside a code span/fenced block were harvested as phase references and fired W002 on ledger rows. Now strips fenced blocks then inline spans via the canonical markdown-sectionizer seam (#2365 composition order) and matches with a left word boundary (?<![-\w]) so hyphen- or word-suffixed carriers are mentions, not references. Pinned tradeoff: a genuine reference written in backticks stops counting (a quoted literal is not a reference). Sub-defect (b): the valid set is workstream-scoped by construction (planningPaths under GSD_WORKSTREAM; per-workstream numbering is deliberate), but the message claimed 'only phases 1, 2 are declared' unqualified. New additive PlanningSnapshot.workstream field, sourced from planning-workspace's new resolveEnvWorkstream() — the ONE env discriminator planningDir itself applies — so the clause cannot disagree with the base the reads used. Root scope keeps the byte-identical message; the checker's scope is unchanged. * test(#4257): close the B2 quoted-literal code span (fixture typo) The B2 fixture wrote a single opening backtick — an unterminated span is literal text per CommonMark, so its content is prose and W002 correctly fired on it. The test's name, the A3 snapshot-level twin, and the B2 matrix row all intend a closed span; pre-fix this was indistinguishable because the unanchored harvest fired either way. * chore(#4257): changeset fragment (pr number to backfill) * chore(#4257): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
c3a18b5ba0 |
docs(#4440): stop telling agents to grep .env files the secret guard denies (#4500)
* docs(#4440): stop telling agents to grep .env files the secret guard denies verification-patterns.md's <environment_config> and user-setup.md's three per-service Verification examples documented reading .env/.env.local directly via grep. Every covered runtime's secret-read guard denies that (Claude Code deny-rules since #768/v1.4.0; the always-on gsd-secret-read-guard hook since #4236/#4221 in 1.13.0) -- verified by piping each documented command through the shipped hook. verification-patterns.md now checks the environment (printenv) instead of the file, with a case statement replacing a broken grep -v alternation (grep's BRE `|` is literal, so the old placeholder filter matched nothing -- PLACEHOLDER/TODO_fill values passed the "substantive" check as real). Verified under sh (dash) against real/placeholder/empty/ unset values. Existence check ([ -f ".env" ] || [ -f ".env.local" ]) is untouched -- it was never denied. user-setup.md's three grep <SERVICE> .env.local lines are removed outright rather than swapped for printenv: those examples describe a Next.js shape where the framework loads .env.local at runtime without exporting it to the shell, so a printenv substitute would wrongly report "not set" on a correctly configured project. Each block's existing service-level check (build/webhook/connection/email test) already verifies the setup. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4440): changeset for the secret-guard verification-examples fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4440): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e38d246015 |
fix(#4439): render pending-todo 'created' as date-only, matching [date] (#4494)
* test(#4439): failing-first regression coverage for pending-todo date rendering renderPendingTodoBullet currently echoes the todo's `created` frontmatter verbatim into the rendered bullet's [date] bracket, so a full ISO-8601 timestamp (the actual stored shape, per add-todo.md's create_file step) leaks through instead of the date-only format documented in docs/reference/state-md.md and docs/COMMANDS.md. This commit adds the regression coverage only; the renderer fix lands in the next commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4439): render pending-todo 'created' as date-only, matching [date] renderPendingTodoBullet echoed the todo's created frontmatter verbatim into the rendered STATE.md bullet. That value is always a full ISO-8601 timestamp by design (add-todo.md's create_file step writes init.todos' timestamp field), but docs/reference/state-md.md and docs/COMMANDS.md document the bullet as `- [date] ...` — a short calendar date. Added pendingTodoDateOnly(), a pure display-only formatter: a value starting with a well-formed YYYY-MM-DD is shortened to just that; anything else ('unknown', a malformed string) passes through unchanged. The stored frontmatter and the JSON todos[].created field are untouched. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4439): cover a non-4-digit-year near-miss for the date-only formatter Standards review flagged a gap: the malformed-date coverage exercised a non-padded month/day but not a non-4-digit year, the other way the input can look almost-but-not-quite like YYYY-MM-DD. Same pass-through code path as the existing malformed-date test; no behavior change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4439): changeset for the pending-todo date-only bullet fix Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4439): backfill changeset PR number Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
8b7a0b696b |
enhance(#4139): Phase 2 — one shared gate, one pilot split, one accuracy spot-check (#4471)
* enhance(#4402): split plan-phase into a spine + detail, add the shared compact-content gate ADR-4139 Decisions 3-5, Phase 2 of the #4139 Compact Content epic. Pilot split for plan-phase.md, the largest of the 58 eagerly-@-included workflow files (98,290 bytes): the spine keeps every happy-path step, every protected-content block (planner/checker prompt templates, quality gates, the failing-direction few-shot example, the two ScheduleWakeup guardrail paragraphs — each marked with a <!-- gsd:protected --> sentinel), and condensed one-paragraph summaries of five rare/opt-in fallback paths (planner and checker filesystem-hang recovery, phase-split recommendation, source-audit gaps, the thinking-partner conditional, and plan bounce). The full text of those five moves verbatim to gsd-core/workflows/plan-phase/detail.md (9.9KB, well under the 32,768-byte NEW_FILE_CAP), read by the spine only when workflow.compact_content is false (the default) — the exact same resolution rule now stated once in the new shared gsd-core/references/compact-content-gate.md, which every future split references instead of restating. Verified mechanically (tests/plan-phase-compact-split.test.cjs, scoped to this one split — Phase 3/#4403 owns the generalized guard): the union of spine + detail contains every non-trivial line the parent commit carried (0 missing), no non-trivial line is duplicated between them (0 duplicated), and every declared protected block is well-formed and non-empty. The spine shrinks from 98,290 to 93,206 bytes (-5.2% of the eager-window cost this epic exists to reduce); detail.md's 9,853 bytes are only ever paid by a project that has NOT opted in. Verified live, end to end, twice, against this actual repo (not a synthetic fixture) — real gsd-planner and gsd-plan-checker subagent spawns, real PLAN.md output: - workflow.compact_content=false: planned a real disposable phase (a docs/how-to page for enabling the key itself); planner returned PLANNING COMPLETE, checker returned VERIFICATION PASSED, all fact-checks against real repo state confirmed. - workflow.compact_content=true (detail.md never read): planned a second real disposable phase; planner returned PLANNING COMPLETE with frontmatter.validate and verify.plan-structure both clean, again fully grounded against real repo state. The five condensed fallback sections were independently re-read spine-only and confirmed sufficient to act on correctly without detail.md's elaboration. Also drafts gsd-core/references/compact-content-protected-content.md — the protected-content category list and <!-- gsd:protected --> sentinel syntax ADR-4139 Decision 5 calls for, written to move to Phase 3 (#4403) unchanged once it lands there. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4402): move detail.md into the ADR-4139-mandated detail/ subdirectory Two independent review sub-agents (Standards and Spec axes of /code-review) caught the same structural defect: ADR-4139 Decision 6 mandates gsd-core/workflows/<name>/detail/*.md ("one or more parts... individually skippable"), and this PR had shipped a flat plan-phase/detail.md instead, copying issue #4402's own (inconsistent) restatement rather than the locked ADR text. Fixed by git-mv to plan-phase/detail/elaboration.md and updating every cross-reference (the spine's step 0.5 gate pointer, the shared compact-content-gate.md's own resolution-rule wording, and the completeness test's path constants). Also, from the same review pass: - docs/CONFIGURATION.md and gsd-core/references/planning-config.md's workflow.compact_content rows said "nothing branches on it yet" — no longer true now that plan-phase.md's spine does. Updated both to name plan-phase as the pilot and note the rest of the corpus is still pending. - Regenerated all 19 tests/fixtures/install-tree/*.json golden fixtures (npm run gen:install-tree) — the three new shipped files were missing from the installer emitted-tree goldens. - Found via a cache-busted `eslint . --max-warnings 0` (this repo's eslint --cache has produced false-greens before): the split test's `git show` call had a bare `timeout: 10000` literal, tripping local/no-adhoc-timeout-literal. Extracted to the existing GIT_TIMEOUT_MS constant from tests/helpers/timeouts.cjs instead of a second guessed copy of the same class of timeout. Verified NOT needed, by tracing the actual mechanism rather than asserting (tests/helpers/emitted-provenance.cjs's gsd-core-verbatim rule attributes every gsd-core/{workflows,references}/** path to itself as an identity source): an Emitted-Drift-Ack-Hash/-Growth trailer. Every changed/added path in this diff is hand-authored and present in the diff itself, so diffEmitted's attribution loop resolves `via` to the path's own source before ever reaching the ack-lookup branch — there is no unattributed delta to acknowledge. The spine also shrank (98,290 to 93,206 bytes), so the growth ratchet has nothing to ack either. Re-verified after these changes: the completeness/disjointness self-check (0 missing, 0 duplicated) still holds against the relocated detail file, and a full `npm run lint:ci` passes clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4402): restore literal content the pre-existing drift guards pin on The first gsd-test run against this split (19 failures) surfaced real regressions: several pre-existing structural guards pin the EXACT text of the sections this split condensed, and paraphrasing broke them. - tests/plan-phase-drift-guard.test.cjs expects the literal `DISK_PLANS=$(gsd_run query find-phase ...)` bash assignment inside plan-phase.md itself, not a prose description of the same check. Restored the exact line into both §9a and §11a's spine summaries. - tests/thinking-partner.test.cjs expects plan-phase.md to literally offer "No, I'll decide" as the skip option. Restored that exact phrase into the condensed thinking-partner paragraph. - Both restores would have duplicated the same text into plan-phase/detail/elaboration.md (which still carries the full elaboration). Removed the now-redundant restatements from the detail file instead of leaving them duplicated — the spine already computes DISK_PLANS before the detail elaboration is ever read, so the detail file references it rather than recomputing it. - Re-running scripts/sync-runtime-launcher.cjs after that edit found the canonical gsd_run preamble had also become an unintentional spine/detail duplicate (both files call gsd_run and each is required, by runtime-launcher-parity's own contract, to carry its own copy). That's sanctioned duplication under a DIFFERENT contract, not lost/copy-pasted content, so tests/plan-phase-compact-split.test.cjs now excludes it from the disjointness check the same way it already excludes trivial fences/headings. - Applied the adversarial-review finding on tests/plan-phase-compact-split.test.cjs's own isTrivial(): a blanket `line.length <= 15` cutoff silently swallowed real content (e.g. the 14-char `<quality_gate>` sentinel). Replaced it with a specific bare-label-line pattern (`Options:`, `Display banner:` etc.) — verified 0 missing / 0 duplicated against the actual split, an improvement over both the original cutoff and a naive full removal (which produces false-positive "duplicates" on generic recurring labels). - gsd-core/references/planning-config.md's own workflow.compact_content row used `/gsd-plan-phase` (hyphen). That file is Claude-facing source text (gsd-core/references/), which tests/slash-command-namespace.test.cjs requires in colon form; docs/CONFIGURATION.md's use of the hyphen form is correct as-is since docs/ is human-facing and outside that test's scanned directories. Fixed to `/gsd:plan-phase`. - tests/plan-phase-compact-split.test.cjs's own `git show` of the parent commit failed inside the gsd-test sandbox ("detected dubious ownership") because the checkout is mounted under a UID the invoking user doesn't own. Scoped `-c safe.directory=<repo-root>` to that one git invocation rather than touching global git config. - docs/INVENTORY.md still had one outstanding "detail.md part" wording fix from the earlier adversarial-review pass, staged now. Re-verified locally against the exact assertions in all four affected test files (all pass) before dispatching a fresh gsd-test run — no change here should have broken any of the other 18 gates; `npm run lint` is clean with the eslint cache cleared. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4402): restore the full marker enumeration to §9a's spine trigger line The isolated Spec-axis review flagged that §9a's "Triggered when" line was condensed to "Agent() returns but the return contains no recognized marker" — dropping the literal `## PLANNING COMPLETE` / `## PHASE SPLIT RECOMMENDED` / `## ⚠ Source Audit` / `## CHECKPOINT REACHED` / `## PLANNING INCONCLUSIVE` enumeration, which is exactly the "machine- parsed structural headings" category compact-content-protected-content.md lists as protected. The load-bearing use of that same list (the gsd_stall_watch call and the Handle Planner Return bullets a few lines above) was never touched — only this one descriptive restatement was genericized — but leaving any instance of a protected category unsentineled is the silent erosion ADR-4139 Decision 4(c) warns sufficiency isn't machine-checkable enough to catch on its own. Restored the full enumeration into the spine. That reintroduced an exact duplicate into plan-phase/detail/elaboration.md, which still stated the same trigger sentence verbatim. Reworded the detail file's version to reference the spine's trigger condition instead of restating it, since the spine is now the single place that sentence lives in full — mirroring the DISK_PLANS/"already computed above" pattern from the previous commit. Re-verified locally: completeness/disjointness (0 missing, 0 duplicated) and all previously-fixed literal-content assertions still hold. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4402): backfill changeset pr number to 4471 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
93abf3a7b6 |
fix(#4276): read the quoting, not just the digits, before eating an IO number (#4420)
POSIX recognizes an IO number only when the digit run is unquoted, but tokenize()'s redirection branch tested `cur` — the token's characters — and never `curMask`, their quoting. `"2"` and `2` were indistinguishable to it, so a correct `gsd_run query commit ... --files "2">out` had its value consumed as an IO number and scored as an unscoped invocation: the #2269 guard reddening on documentation that was right. The mask was already maintained in the same loop and thrown away on this branch. One conjunct reads it: '0' marks a bare character, so /^0+$/ asks precisely whether every character of the digit run was unquoted. The regression arms come in both directions. The quoted rows (glued, detached, single-quoted, and the partially-quoted 2"3" that a some-character-unquoted test would get wrong) must be scoped; the existing unquoted `--files 2>&1` row is the negative control and must stay unscoped, so an implementation that simply stopped consuming IO numbers altogether fails instead of passing. No live instance exists in any of the six scan roots, so this is latent rather than urgent — a false positive, never a silent miss. Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
476394689a |
fix(#4254): pin sequential executor to the orchestrator's validated root (#4476)
* test(#4254): sequential executor root pin — failing-first regression + matrix The new suite executes the shipped supplied-root-pin guard against real git fixtures (drifted primary-checkout cwd halts before the write and the FATAL names both roots; matching cwd permits it; unexpanded/empty pins halt; normalization forms; submodule and sibling boundaries; metacharacter quoting; drive-letter form gate) and locks the dispatch contract across execute-phase.md, its sequential-root-pin step fragment, and worktree-path-safety.md. The #2772 per-plan serialization assertion retargets to the fragment that now carries those rules (ADR-857 Phase 6 ceiling), plus the host-step wiring. * fix(#4254): pin sequential executor to the orchestrator's validated root Sequential-mode dispatch told the executor to self-derive PROJECT_ROOT from its own cwd; every existing guard is worktree-mode-only or self-referential, so an executor spawned with a drifted cwd committed onto the wrong checkout silently. - worktree-path-safety.md step 0p: mode-agnostic supplied-root pin guard, composed by the orchestrator at build time with the literal $ORCHESTRATOR_WT (git-vs-git comparison on both sides — representation-safe on Windows, the #4296 lesson), fail-closed on empty/unexpanded pins, registered-submodule allowance, warn-and-proceed only when the dispatch carries no pin block. - execute-phase.md sequential branch: build-time embed of the bound <project_root_pin> via the new execute-phase/steps/sequential-root-pin.md fragment (ADR-857 Phase 6 frozen ceiling — the host step cannot grow; the wave serialization rules move with the fragment, verbatim in substance) plus the per-write/commit pin instruction in <sequential_execution>. Worktree-mode dispatch untouched (its self-derived toplevel IS correct there). - INVENTORY rows (5 locales) + INVENTORY-MANIFEST + install-tree goldens regenerated for the new fragment; changeset added. * chore(#4254): backfill changeset PR number * fix(#4254): accept backslash-separated Windows drive pins CI on windows-latest showed every permit-path test failing with "Actual root: <none>": pins composed from Node's path.join arrive in the backslash drive form (C:\Users\RUNNER~1\...), which the guard's absolute-form gate rejected before the cwd-side root was ever computed — a legitimate matching pin could never pass. The gate now accepts either separator ([A-Za-z]:[\\/]); git -C resolves both forms (and 8.3 short names) to the same canonical toplevel, so the git-vs-git comparison is unaffected. Form-gate tests cover the emitted (C:/…) and produced (C:\…) spellings plus short names. * fix(#4254): portable drive-form gate for MSYS bash The bracket class [\\/] that accepted backslash drive pins parses inconsistently on MSYS bash (the Windows CI leg still rejected C:\ pins — every permit-path test red with "Actual root: <none>"). Replace it with standard pattern escaping outside brackets: [A-Za-z]:/*|[A-Za-z]:\\* — the escape form is version- and build-portable. Verified across all forms: both drive spellings accepted; bare "C:", relative, empty, and unexpanded rejected. * fix(#4254): runtime-generated backslash comparator + self-describing FATAL The Windows CI legs failed every #4254 permit-path row with 'Actual root: <none>' across two prior pattern spellings ([\\/] and \\*). Stage misattribution: <none> appears whenever the FATAL fires BEFORE the cwd-side capture assigns ACTUAL_ROOT — the absolute-form gate was what fired. Mechanism: the test harness spawns bash -c <script> through the Windows command-line boundary; that round-trip applies one extra shell-quoting pass with double-quote semantics — a backslash written twice in the script text arrives halved, while a lone backslash survives (the pin displays intact; row 9's pure-bash gate independently showed the halved pattern rejecting C:\ while C:/ still passed its surviving arm). On windows-latest every pin carries backslashes (os.tmpdir() is the 8.3 short form C:\Users\RUNNER~1\...), so the gate ate every pin before the actual root was ever computed. Fix, robust by construction: - the drive-form gate generates its backslash comparator at RUNTIME (BS=$(printf '\134'); match [A-Za-z]:"$BS"*) — the shipped guard now contains no doubled backslash anywhere, enforced by a regression assertion on the extracted guard text; - the FATAL self-describes: Guard stage (pin-unbound / form-gate / actual-capture / pinned-capture / root-mismatch) plus a Diagnostic line carrying git's own stderr for capture failures and both compared values for mismatches — future platform failures name their stage in the log; - row 9's hand-rolled duplicate case gate (transit-fragile copy, #4296 Minor 1 duplication smell) is replaced by driving the SHIPPED guard and asserting the stage; rows 2/4 pin the new stage machinery. Validated on darwin across drift/match/relative/unbound/empty/bare-drive/ forward-and-backslash drive forms, each also re-run under a simulated Windows transit (every doubled backslash halved) with identical outcomes. * fix(#4254): close the empty-comparator fail-open seam in the drive-form gate Self-review of the runtime-generated backslash comparator: if printf's octal escape ever returned empty, the drive arm [A-Za-z]:"$BS"* would widen to drive-RELATIVE pins (C:foo) — the construction's one theoretical fail-open path. Fail closed with a self-describing diagnostic instead of trusting the shell's printf. --------- Co-authored-by: sim <sim@local> |
||
|
|
2388e6ab34 |
fix(#4342): run the bug-167 routing test in a fixture project, not the developer's (#4387)
The test called runGsdTools with its default cwd — the test process's own working directory — and an inherited HOME, so the child read the checkout's real .planning/ and the developer's real ~/.gsd/defaults.json. testEnvBase() blanks the config-LOCATION env keys but sandboxes neither cwd nor HOME. On a checkout that has workstreams with no active pointer, `init.progress` exits non-zero and the FIRST assertion fails, so the routing comparison the test exists for was never evaluated: init.progress failed: Error: init.progress requires a workstream in workstream mode — no active workstream is set ... Available workstreams: alpha Reproduced byte-for-byte by adding .planning/workstreams/alpha/ to the checkout: red on next, green here. The invariant under test — `query <cmd>` and `<cmd>` returning identical payloads — is independent of project state, so a plain createTempProject() fixture is enough, with HOME/USERPROFILE pointed at it (the idiom runGsdTools's own doc comment prescribes). Two assertions pin the sandbox deterministically rather than conditionally: the fixture HAS a .planning/ and the repo checkout does not, so dropping the cwd override fails on every lane — CI included, where the ambient state that exposed the bug is absent. Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ae40529d31 |
chore(#4394): lint allowed-tools parity — Bash without Grep (#4431)
gen-plugin-skills.cjs --check already guarantees skills/*/SKILL.md matches what commands/gsd/*.md generates, so the two trees cannot silently diverge FROM EACH OTHER. Nothing guarded the shape #3085 actually found: a command shipping Bash without Grep purely by omission, identical in both trees and therefore invisible to a parity check that only compares them to each other. That drift ran until 29 of 71 skills lacked a tool most of their siblings declared, and a manual audit — not a gate — is what surfaced it. Detection only: the lint never edits a command's allowed-tools. Two failure classes, not one. Violations are the rule itself. Stale exemptions are the other half: an entry whose command is gone, or which no longer declares Bash without Grep, fails just as loudly. An exemption list that can only grow becomes a list of things nobody re-examined, and a pre-forgiven command silently absorbs the next omission. That check earned its keep immediately. #4394 named eight exemptions from the #3085 review — the six ns-* dispatchers, help, and surface — and seven of them do not declare Bash at all, so the rule never reaches them. Listing them would have pre-forgiven seven commands for a condition none of them has. Only surface needs an entry. The tests drive synthetic fixtures rather than the live corpus: asserting "the real tree is clean" would say nothing about whether the rule can detect anything, which is the exact failure mode this lint exists to close. The one live-corpus arm asserts no stale exemptions — a property of this script's own list — and deliberately does not pin a violation count, which would make it a baseline every fix has to update. Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
0a0905705a |
fix(#4256): resolve todos from the root via todosDir everywhere (#4479)
* test(#4256): pin todos as root-scoped under workstreams (RED) * fix(#4256): resolve todos from the root via todosDir everywhere * chore(#4256): changeset fragment (pr number to backfill) * chore(#4256): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
6ebe6372ce |
fix(#4243): anchor stateReplaceProgressPercent bold form to line start (#4474)
* fix(#4243): anchor stateReplaceProgressPercent bold form to line start The bold branch of stateReplaceProgressPercent carried no ^ and no /m flag, so a bold percent-ish label quoted MID-SENTENCE inside prose — an Accumulated Context bullet mentioning **Progress:** — captured the machine-segment rewrite and destroyed the rest of its line, silently, while the real Progress line stayed stale (and the frontmatter moved on without it, breaking the #4213 surfaces-agree contract). Every caller (cmdStateUpdateProgress, syncCore's percent arm, applyPostSyncPreservation) feeds the whole document, so all three were exposed. Anchored to ^([ \t]*\*\*Progress:\*\*[ \t]*)([^\r\n]*)$ with /im — the exact idiom #4453 applied to stateReplaceField's bold branch (same-line confinement per #4010: the leading class is [ \t]*, deliberately not \s*, which can consume the newlines before the label into the match; $ is explicit-and-inert and documents end-of-line). #2177's recorded requirements all stand: frontmatter is stripped before matching, the suffix-preserving machine-segment swap is untouched, and bold-beats-plain priority now governs line-start forms, so an earlier free-text plain Progress: line still cannot capture the rewrite ahead of the real bold status line. Per the maintainer ruling (2026-09-07), #2177's incidental bold-anywhere matching was not load-bearing. * test(#4243): scope the C4 region check with splitLines, not a bare \n split lint:ci (local/no-crlf-fragile-split) flagged the free-text-plain-line row's content.split(/\n## /)[0] — a bare \n split on readFileSync content is CRLF-fragile under Windows autocrlf. Same scoping via splitLines() (src/text-lines.cts), which splits on \r?\n. * chore(#4243): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
c4b6dbd486 |
fix(#4247): refuse update-plan-progress on a roadmap with no writable phase entry (#4468)
* test(#4247): failing-first regressions for checklist-form update-plan-progress * fix(#4247): refuse update-plan-progress when the roadmap has no writable phase entry * fix(#4247): single local source for the phase-heading anchor grammar * docs(#4247): note the missing_phase_details refusal in cli-tools reference * docs(#4247): backfill pr number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
03c770be47 |
docs(#4463): record executor self-repair of worktree base as out-of-scope (#4470)
Denies reintroducing sub-agent-side `git reset --hard` recovery for a worktree base mismatch — the exact primitive #48 removed for safety. Keeps the fork-base measurement finding for future reference. Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
33e393ba4c |
fix(#4243): anchor stateReplaceField bold form; pin frontmatter round-trip (#4453)
* test(#4243): failing-first regressions for bold-field anchoring and frontmatter round-trip * fix(#4243): anchor stateReplaceField bold form to line start The bold branch of stateReplaceField carried no ^ and no /m flag, so a bold label quoted mid-sentence inside prose — the issue's **Status:** inside an Accumulated Context bullet — captured the rewrite and destroyed the rest of its line, silently, whenever a whole-body caller fed the function every section (beginPhaseCore's tryField, advancePlanCore's Status/Current Plan writes). The plain branch was always line-anchored; only the bold branch lagged. Anchored to ^([ \t]*\*\*Field:\*\*[ \t]*) with /im, reusing #4010's same-line confinement idiom for the leading class (deliberately not the issue's suggested ^\s* — it can consume the newlines before the label into the match) and #4186's recognition-by-anchoring discipline. Frontmatter half of the issue (unknown-key drops, invented milestone defaults) is already fixed on next by #2202/#3216/#4129; pinned here with the issue's requested regression fixtures. * test(#4243): pin survival contract, not derived percent, in frontmatter rows Bench RED run caught two assertion defects in the pin rows: the unknown progress subkey re-parses as a quoted scalar ('77' vs 77), and percent is a declared derived subkey - omitted under the #3573 no-roadmap withhold, recomputed when measured (#4129) - so pinning its value over-pins derived semantics. The rows now pin what the issue demands: unknown/custom keys survive, stored counters are kept under the withhold, milestone identity is never reset to invented defaults. * chore(#4243): changeset for the anchored bold-field fix * chore(#4243): backfill PR number in changeset |
||
|
|
8c8eda46b0 |
fix(#4225): scope the sibling-worktree phase-number horizon to the active workstream (#4450)
* test(#4225): failing-first matrix for phase.add --ws workstream-scoped numbering Nine rows driven through the real CLI: the issue's verbatim topology (root roadmap @39 committed, workstream @2, sibling git worktree carrying the root roadmap), same-workstream sibling boundary, empty-workstream first phase, coincidental root-maximum, no---ws control (the #3849 global horizon, byte-for-byte), cross-workstream isolation, sibling lacking the workstream (fail open), add-batch parity, and a next-decimal control. Rows 1/2/3/6/7/8 are RED on next @38e4ce5f62 (numbering computed from the sibling ROOT roadmaps: 40 instead of 3). * fix(#4225): scope the #3849 sibling-worktree widening horizon to the active workstream collectSiblingWorktreePhaseNums scanned each sibling git worktree's ROOT .planning/ (phases/ dirs + ROADMAP.md headers) unconditionally. Under --ws (GSD_WORKSTREAM), every local number source flows through planningDir(cwd) and lands in the workstream scope, but the widening horizon still merged the siblings' ROOT-roadmap numbers into it — so phase.add --ws in a workstream at Phase 2 inside a project whose root roadmap sits at Phase 39 minted Phase 40 (directory 40-<slug>, and a Depends on: Phase 39 that does not exist in the workstream's numbering universe). The horizon now resolves each sibling's planning dir through the SAME canonical resolver, planningDir(wt, ws), with the env workstream read once via planningDir's own discriminator: a workstream-scoped allocation scans the sibling's copy of the SAME workstream (a number taken by that workstream on another branch is still taken — the #3849 widening survives, scoped), and never the sibling's root roadmap or another workstream's. No workstream active: ws is null and the root-scope horizon is byte-for-byte the #3849 behavior. A sibling lacking the workstream directory contributes nothing (fail open, unchanged). phase.add and phase.add-batch share the helper; both scopes of both verbs are covered by the matrix in the previous commit. Output shape and the publishStateContract boundary are untouched — only the number changes. * fix(#4225): rename siblingPlanning -> siblingPlanningDir (review nit) * chore(#4225): changeset fragment (pr number to backfill) * chore(#4225): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
f15887ebb1 |
feat(#4446): ban ad hoc timeout literals in tests, ship with full legacy allowlist (#4449)
Nothing enforced CONTRIBUTING.md's own stated preference ("A non-literal value
is trusted — that is the shape you should be writing") for timeout values in
tests. local/no-unbounded-spawn requires SOME bound but allows a bare literal;
no-magic-sleep-in-tests and no-elapsed-assertion cover different anti-patterns
entirely. This is exactly how PR #4428's Windows CI incident happened: two
independently-guessed 15000ms literals (one in production's check-latest-
version.cjs, one in this suite's own worker test) collided exactly and raced
two SIGKILLs against each other.
New rule local/no-adhoc-timeout-literal (eslint-rules/no-adhoc-timeout-
literal.cjs) flags a resolvable numeric timeout/timeoutMs literal; an
Identifier or MemberExpression value is trusted. No marker-comment escape —
the fix is always to extract a named constant, which is trivial.
Ships with a full legacy allowlist (126 files, 352 violations, generated
by running the rule with an empty allowlist against tests/) so it can go
live at error severity without breaking CI, mirroring how no-unbounded-spawn
itself was rolled out. Migration is tracked separately in #4445 (this PR
does not close it — only #4446, introducing the gate itself).
Documents the policy and the compliant shapes in TESTING-STANDARDS.md and
TEST-EXAMPLES.md.
Co-authored-by: sim <sim@local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
ef30e59860 |
fix(#4448): stop io.test.cjs's in-process runMain calls from corrupting node:test's own fd-1 IPC (#4452)
tests/io.test.cjs's "review fix: pending-outcome cell lifetime" describe block deliberately drives runMain()/output() in-process (needed to observe a cross-invocation state leak) instead of via a subprocess. output() ends with a raw synchronous fs.writeSync(1, ...) to the real stdout fd, and Node's --test-isolation=process (default since Node 22) uses that same fd for the file's own parent-child reporter protocol. The two writes racing produced an intermittent "Unable to deserialize cloned data" that killed the whole file — observed twice on next's macOS lane, most recently on the commit that merged PR #4428 (unrelated to that PR's content; io.test.cjs isn't part of its diff). Empirically validated locally (gsd-test can't reach macOS): built a repro loop running N parallel copies of `node --test tests/io.test.cjs` to recreate CI-like contention. Baseline: ~13-17% of runs hit the corruption (12/90, 15/90 across two samples). A first fix attempt wrapped the writes in captureFdAsync (an await-aware twin of the existing captureFdSync, added because runMain() defers main() through a microtask chain, so a synchronous wrap restores before the real write fires) — but captureFdAsync always forwards to the real fs.writeSync by design (matching captureFdSync's "never swallow" contract, tests/helpers.cjs, #4306). Re-ran the same loop against that fix: 15/90, statistically unchanged. Forwarding the write doesn't stop it from reaching the fd node:test's own IPC also uses. Replaced it with suppressFdAsync: a narrow, deliberate exception to the never-swallow contract for a window the caller has verified is fully controlled (a single runMain() call plus its promise-chain settling, where nothing else can legitimately need that fd). It records the bytes for the test's own assertions but never lets them reach the real fd. Re-ran the loop: 0/300 across three samples (90+120+90), including one round at 8-way parallelism. Also caught and fixed a real bug surfaced by the same loop: the new regression assertion checked for compact-JSON `"error":"x"` but output() pretty-prints, so it failed 100% of runs deterministically until fixed to parse and check the structured value instead (io.test.cjs, matching this repo's "assert on structured output, not raw text" convention). Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
e54d3aa159 |
enhance(#4401): register workflow.compact_content as a validated config key (#4441)
* feat(#4401): register workflow.compact_content as a validated config key - Add compact_content: false to the nested workflow object in gsd-core/bin/shared/config-defaults.manifest.json - Add 'workflow.compact_content': false to SCHEMA_DEFAULTS in src/config.cts so an absent key resolves to false via config-get --raw - validKeys entry in config-schema.manifest.json already present Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(#4401): behavioral and boundary tests for workflow.compact_content - 19 behavioral tests covering config-set/config-get round trip, invalid-shape rejection (banana, 42, empty string), the corrected null-unset semantics (#2046), absent-key resolution against config-defaults.manifest.json, config-new-project wiring, and doc-row shape assertions - Drops the install-tree fixture-parity block (and its docstring item) that asserted gsd-core/references/compact-content-gate.md and gsd-core/workflows/compact/map-codebase.md fixture entries — those paths belong to #4402 and do not exist on this filtered branch Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(#4401): document workflow.compact_content in both config references - One 4-cell row in docs/CONFIGURATION.md (workflow.* run) - One 5-cell row under Workflow Fields in gsd-core/references/planning-config.md - Both cross-reference ADR-4139 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(#4401): add changeset - Added-type fragment, pr: 4401 (issue number; backfill to the real PR number is a required follow-up once the PR is opened, per D-08 and CHANGESET-PR- FIELD-DRIFT) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(#4401): backfill changeset pr field to #4441 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(#4401): derive workflow.compact_content default from CONFIG_DEFAULTS SCHEMA_DEFAULTS['workflow.compact_content'] hardcoded the literal false instead of deriving it from CONFIG_DEFAULTS the way 3 of its 8 sibling entries do (smart_zone_tokens, pr_strict, inline_plan_threshold), leaving a single-source-of-truth drift risk: a future manifest-only edit to the default could silently diverge from this literal, only caught later by the D-03 test if it ever happened to manifest. Adds compact_content to CONFIG_DEFAULTS in src/config-loader.cts and derives SCHEMA_DEFAULTS from it in src/config.cts, matching the majority sibling pattern. Found during maintainer review (review-open-prs) of this PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4401): map compact_content in config-field-docs NAMESPACE_MAP The previous commit added compact_content to CONFIG_DEFAULTS in src/config-loader.cts but missed the matching entry in tests/config-field-docs.test.cjs's NAMESPACE_MAP, which maps flat CONFIG_DEFAULTS keys to their namespaced doc form before checking gsd-core/references/planning-config.md for a match. Without it, the test looked for a bare `compact_content` doc reference instead of the actual `workflow.compact_content` row, and failed: "CONFIG_DEFAULTS keys missing from planning-config.md: compact_content". Found by actually running gsd-test against the branch rather than trusting the plausible-looking fix. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(#4401): register compact-content-4139 test in the docs-guard lane tests/compact-content-4139.test.cjs's D-06 tests read docs/CONFIGURATION.md directly (fs.readFileSync) to assert the workflow.compact_content doc row's shape, which makes it a doc-reading test file under the #3753 docs-guard lane. It was never added to scripts/docs-guard-registry.cjs's DOCS_GUARD_TESTS map and carries no docs-guard-exempt marker, so tests/ci-docs-guard-registry.test.cjs's registration lint correctly failed: "compact-content-4139.test.cjs reads a docs/ path but is not registered in the docs-guard lane and carries no docs-guard-exempt marker". Registers it with ['docs/CONFIGURATION.md'] (the only real docs/-prefixed path it reads; gsd-core/references/planning-config.md is outside this registry's docs/ scope, matching the sibling config-field-docs.test.cjs entry's existing convention). Found by actually running gsd-test against the branch — this gap predates the maintainer's config-loader.cts fix and was already present in the original PR. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: sim <sim@local> |
||
|
|
2cf119f57e |
fix(#4217): reconcile artifacts before classifying an abnormally-ended executor (#4442)
* fix(#4217): reconcile artifacts before classifying abnormal ends * test(#4217): pin the completion-reconciliation contract * chore(#4217): regen derived inventory and install-tree fixtures * test(#4217): follow the #4003 anchoring pins into the reconciliation fragment Emitted-Drift-Ack-Growth: execute-phase.md — the runtime-neutral completion-reconciliation pointer, the two Codex wait-rule bindings, and the step-7 reconcile-first gate net +33 bytes over the extracted fallback block (#4217) * chore(#4217): add changeset fragment * chore(#4217): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
19b66c3ec8 |
fix(#4218): stop the orchestrator steering an executor that is still working (#4391)
* fix(#4218): stop the orchestrator steering an executor that is still working An executor with recent RED/GREEN/REFACTOR commits and passing verification had not yet written its SUMMARY because it was finishing closeout. The parent saw no local OS test/build process, inferred an "idle tail", and sent "Finalize immediately" into a working child; in CLI runs the same inference interrupted an executor before GREEN, leaving a RED commit and an uncommitted edit. The stall block said only "if no completion signal, no SUMMARY.md, and no expected-branch commits appear for N minutes" — it never said what to do when commits DO exist and only the SUMMARY is outstanding, never defined the threshold as a period without progress rather than a total runtime, and never ruled out a process listing as an idleness signal. Four rules close that: - the threshold measures a period WITHOUT MEANINGFUL PROGRESS, from the last sign of progress, not from dispatch — a long verification tail is not a stall; - commits + missing SUMMARY + recent activity resolves to KEEP WAITING, with steering, interrupting and re-dispatching each named and forbidden; - urgency/finalization messages ("Finalize immediately" and family) are forbidden outright — they arrive mid-verification and truncate a correct run. The existing user-facing pause is the only sanctioned stop, and `kill and retry` is a clean restart, not a nudge; - the absence of a local OS test/build process is NOT idleness: a native subagent runs in the runtime's own session, and an executor between two tool calls shows no process at all. Progress is judged only by the signals this workflow names. Five prose-contract assertions in tests/execute-phase-wave.test.cjs, all red on next. * fix(#4218): extract the progress policy to a step fragment CI's #1168 gate caught it: execute-phase.md sits 77 bytes under a frozen 93600 ceiling and the four rules added ~2.3 KB. "Extract, not bump" is the repo's stated remedy, and this workflow already carries policy detail that way. execute-phase/steps/executor-progress-policy.md owns the policy. The worktree-recovery arm moved with it — `kill and switch to inline execution` qualifies the stop this policy governs, so it belongs beside the rule about when stopping is sanctioned at all, not stranded in the host. The #3212 recovery OPTIONS stay in the host, where tests/config.test.cjs pins them. The host keeps what must be read before the orchestrator acts: the verdict, the threshold definition, and a pointer that fires before any message is sent to the child. execute-phase.md is now 93475 bytes — 48 SMALLER than next. * chore: add changeset for #4218 * chore(#4218): regenerate the inventory manifest for the new step fragment docs/INVENTORY-MANIFEST.json is the authoritative per-file list behind INVENTORY.md's `<workflow>/steps/*.md` row, so a new fragment has to appear there or gen-inventory-manifest --check reds the lint-tests lane. * chore(#4218): restore the issue ref on the allow-test-rule marker ADR-456 requires a #NNN on a new exemption; the block rewrite that moved the policy into the fragment dropped it. * chore(#4218): regenerate the install-tree fixtures for the new step fragment The fragment ships with the workflow, so every runtime's golden install tree gains one path — gen:install-tree is the generator that owns those fixtures. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
1c0acb2359 |
feat(#4422): block merging into next/main while the base branch's Tests run is red (#4428)
* feat(#4422): block merging into next/main while the base branch's Tests run is red Adds a next-health job to test.yml that checks the base branch's own last push-triggered Tests run via the GitHub API and fails the existing "Required tests" required check when it's red, with a maintainer-applied "fix-next" label as the explicit escape hatch for the fix-forward PR itself. No branch-protection config change needed — it rides the already-required check. The job is deliberately not gated behind preflight, same reasoning as the changes job: a compute-free API read has nothing to save by waiting. Documents the fix-next label in CONTRIBUTING.md and adds a property test locking the CLEAN/RED/INDETERMINATE classification's iff-relationship. This closes the second half of the 2026-09-06 RCA: three unrelated PRs merged on top of an already-broken next before anyone noticed it was red. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix: close two zero-margin CI timing gaps found while verifying #4422 Discovered while watching this branch's own CI, root-caused via /diagnose rather than dismissed as Windows flakiness: 1. tests/gsd-check-update-worker-atomic-cache.test.cjs's outer timeout (15000ms) exactly matched the inner npm-view timeout the worker wraps (NPM_VIEW_TIMEOUT_MS, gsd-core/bin/check-latest-version.cjs). A slow registry response raced two SIGKILLs at the same instant, killing the worker before it could catch its own timeout and degrade gracefully. Windows's shell-wrapped npm subprocess made the race lose more often there, but the zero margin was platform-agnostic. Fixed by giving the test real headroom (+10s) beyond the named constant it wraps, plus an invariant test so the two values can't silently collide again. 2. scripts/run-tests.cjs's per-chunk weight budget (MAX_FILES_PER_CHUNK) let a Windows full-matrix chunk that was well under budget by the Linux/macOS-calibrated weight table (~32/60 units) still exceed the 600s wall-clock backstop — codex-config.test.cjs's genuinely-measured weight (17.87) doesn't transfer 1:1 to Windows's slower install/ subprocess overhead. Windows now gets its own lower cap (40 vs 60). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
38e4ce5f62 |
fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381)
* fix(#4186): anchored status vocabulary, record-session arg guard, recount pin Three defects from #4186: 1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the free-prose body Status field, so prose merely mentioning a status word was silently rewritten to a credible wrong token (a .planning/ path in Italian prose -> status: planning; verifica -> verifying; completezza -> completed). Recognition is now an ANCHORED whole-field match against a declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS, state-document.cts) — case/whitespace-tolerant, branch-order artifacts preserved (Planning complete -> planning; Phase complete — ready for verification -> verifying). The recorded lenient fallback (#3873 row 26) stands: unrecognized prose passes through verbatim. Read-side consumers (W011, statusline) ride the same function. 2. The progress recount skew (stray *-SUMMARY.md inflating completed_plans) is already dead on next via #1988/PR #2016 (countMatchedSummaries pairs summaries to plans) — verified live and pinned with regression rows composed against the #4129/#4359 ratchet. 3. state record-session with no args executed and wrote STATE.md; it now errors like state update (stopped-at or resume-file required), handler- side so SDK callers are covered too. Four tests pinning the bare-call write are updated to the new contract. * fix(#4186): update status pins to the anchored vocabulary contract Bench round 1 follow-ups: - Legacy bare 'Milestone complete' kept as reader-side vocabulary (ADR-2207 removed the writers, not recognition of legacy files). - state.test pins updated: 'Paused at Plan 3' and round-trip 'Executing Plan 5' were pins of the substring guessing itself — the round-trip now uses the real handler form 'Executing Phase 5'. - record-session no-op/no-fields tests repurposed to the usage-error contract (CLI + SDK-level ExitError), byte-unchanged assertions kept. - statusline tests repinned: vocabulary values collapse to keywords; narratives render the documented first-word fallback instead of a guessed token. Hook doc comment updated to match. - docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md. - docs/CLI-TOOLS.md: record-session signature notes the required flag. * fix(#4186): repair a dangling sentence in the schema docstring * test(#4186): bound the completed_plans scan regex (#2128 class) * chore(#4186): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
47f83beb62 |
fix(#4205): filter gsd_run, not gsd-tools, when isolating PATH in launcher fixtures (#4337)
* fix(#4205): filter gsd_run, not gsd-tools, when isolating PATH in launcher fixtures The runtime launcher's PATH-fallback arm probes `command -v gsd_run` (renamed from gsd-tools in #3146), but every PATH-isolation fixture in runtime-launcher-parity.test.cjs filtered on the pre-rename name. A real installed gsd_run reachable on PATH survived the filter and got invoked in place of the fixture's runtime-home stub, so negative tests passed without proving PATH was actually empty and positive home-fallback tests failed with "Unknown command" errors from the unrelated real CLI. Adds (B1), a regression test that plants a sentinel gsd_run on PATH and asserts the resolver still falls through to the HERMES_HOME stub instead of invoking it. * docs(#4205): fix stale gsd-tools references in PATH-probe doc comments Addresses agy adversarial review nits on PR #4205: several doc comments and JSDoc blocks still described the launcher's PATH-fallback probe as `gsd-tools` after the filter fix. Updates them to `gsd_run` to match the actual `command -v gsd_run` probe and the corrected filters. No test logic changes. * test(#4205): tighten (B1) assertions and dedupe rationale comments Addresses opus critical-code-reviewer/ponytail findings on PR #27: - (B1): split the collapsed && assertion into two, matching neighbor (B)'s style, for clearer failure diagnostics. - (B1): drop the dead `if (nodeBinDir)` guard — cleanup() already no-ops on a non-string argument (tests/helpers.cjs:452). - (B1): drop the dead backslash-path normalization — the test is win32-skipped, so stdout paths are always POSIX. - Six near-identical "#4205: probe target is gsd_run, not gsd-tools" comments collapsed to pointers at the one canonical explanation in buildIsolatedPath(). No behavior change; 31/31 tests still pass. * fix(#4205): scrub ambient config-dir env vars leaking into launcher fixtures Same class of bug as the PATH leak this issue reports, different vector: the resolver's runtime-home elif chain checks CLAUDE_CONFIG_DIR before HERMES_HOME, CURSOR_CONFIG_DIR, CODEX_HOME, etc., but the fixtures targeting those later arms never cleared the earlier ones from the spread process.env. An ambient CLAUDE_CONFIG_DIR pointing at a real install silently wins over the fixture's intended stub, exactly like the reported gsd_run PATH leak. Confirmed with a real leaked install: red on tests (D)/(H)/bug-211 (C)/(D)/(B1)/(B)/(C) without the fix, green with it. Also removes bug-891's (B) test, now a strict subset of (B1): once the sentinel is filtered by buildIsolatedPath(), both tests exercise the identical HERMES_HOME resolution with the identical script and env — (B1) already asserts everything (B) did, plus the sentinel-not-invoked check. Updated the block's header docblock to match. 30/30 tests pass (31 minus the removed duplicate). * fix(#4205): scrub CODEX_HOME/XDG_CONFIG_HOME, restore (B) on Windows Adversarial review of this PR found three more instances of the exact leak class the PR exists to close. (H) asserts the $HOME/.codex fallback but never cleared an ambient CODEX_HOME, which overrides that default outright. Proven load-bearing: with a fake install planted at CODEX_HOME the test fails without this scrub and the leaked install's own output appears in stdout. The two "every arm must miss" hard-error fixtures cleared all 16 config-dir vars but not XDG_CONFIG_HOME, which the resolver's opencode and kilo arms fall back through as ${OPENCODE_CONFIG_DIR:-${XDG_CONFIG_HOME:-$HOME/.config}/opencode} — a host XDG_CONFIG_HOME leaks past a fake HOME. Restore bug-891's (B), deleted here as a subset of (B1). It is not one on Windows: (B) ran cross-platform, (B1) is POSIX-only because it plants an executable sh sentinel, so the deletion left the HERMES_HOME arm with no Windows coverage. Restored with the CLAUDE_CONFIG_DIR scrub its sibling fixtures already carry. * docs(#4205): correct makeIsolatedPath's docblock and the bug-891 header The doc-fix commit earlier in this branch rewrote makeIsolatedPath's docblock from "strips gsd-tools" to "strips gsd_run". Both are false: the function filters nothing, returns process.env.PATH whole, and the "noToolsBin dir that shadows gsd_run with a sentinel" it describes does not exist — every caller passes an empty directory. Its isolation comes from resolution order, since the RUNTIME_DIR/.claude arm fires before the PATH arm. Say that instead, and point anyone who needs the PATH arm itself to miss at buildIsolatedPath(), which does filter. Add (B) to the bug-891 asserts list; restoring it in 48cf0b917 left the header describing a test set the file no longer has. Mark (B) and (B1) cross-platform and POSIX-only respectively, which is why both exist. Drop one more stale "remove gsd-tools" comment the doc pass missed. * fix(#4205): derive the fixture env scrub, stop writing to CLAUDE_ENV_FILE Two findings from review, both measured. CLAUDE_ENV_FILE was never scrubbed. The snippet's tail appends `export PATH='<dir>'` to it whenever it is set, so running this suite on a host that exports it wrote 11 lines into the developer's real env file, each naming a /tmp fixture directory the test had already deleted — they accumulate per run and prepend dead entries to the PATH of every later shell. The reported bug was fixtures READING developer state; this was them writing to it. Now 0 lines. The 17-key scrub list was hand-written, which tests/helpers.cjs already warns against: "#2665: this list is DERIVED, not hand-maintained. A hand-written list is exactly what reopened this bug twice". It was right — the hand list here missed CODEX_HOME and XDG_CONFIG_HOME until review caught them, and a 17th runtime home would have left it silently stale. Replace both copies with TEST_ENV_BASE, derived from the registry the resolver itself reads, applied at runBashFile/runResolver so every fixture that sources the snippet is covered rather than the two that remembered to ask. GEMINI_CONFIG_DIR is added explicitly: the runtime is retired (#1928) so the registry no longer carries it, but the snippet still probes its arm. Verified by pointing all 16 config-dir vars plus XDG_CONFIG_HOME at a real install tree: 30/30 pass. Fold (B1) into (B). Reverting the filter under a clean PATH left the old (B) green — it only caught the bug on an already-leaking machine — while (B1) caught it anywhere but was skipped on Windows. One test now does both: it plants the PATH sentinel on POSIX and still exercises the HERMES_HOME arm on Windows. Mutation-checked both ways on a PATH with no real gsd_run. Assert the hermes dir itself rather than "gsd-core/bin/", which every resolver arm ends in and so cannot tell them apart. * fix(#4205): scrub BASH_ENV and reject an empty RUNTIME_DIR in runResolver BASH_ENV defeated the whole scrub. Non-interactive bash sources it before the script runs, which is after the env: object is applied, so a single inherited var re-injects any of the others. Measured: a BASH_ENV exporting CODEX_HOME turned (H) red; blanked, 30/30. runResolver passed RUNTIME_DIR: runtimeDir || '', and '' is indistinguishable from unset to ${RUNTIME_DIR:-$(git rev-parse --show-toplevel)} — an empty value falls back to the real repo root and resolves its real install, the leak this issue is about. Both callers already pass one, so require it rather than paper over it. * test(#4344): plant the leaked gsd_run sentinel on Windows too (B) planted its gsd_run sentinel only on POSIX, so the Windows shards proved nothing about buildIsolatedPath()'s PATH filter — the exact gap #4344 recorded. npm's global installs write an extensionless Bourne shim beside gsd_run.cmd, and fs.constants.X_OK behaves like F_OK on Windows, so the existing probe already sees the leak there; only the fixture was POSIX-gated. Plant the sentinel on every platform and assert that buildIsolatedPath() strips its directory from the returned PATH. That assertion is red on both platforms when the filter probes the wrong name, and unlike the stdout assertions it does not depend on the MSYS mount's exec heuristics. Restore process.env.PATH before the child spawns rather than in a t.after hook: on Windows process.env spreads as 'Path', so a still-live leak would compete with snippetEnv()'s 'PATH' override for the casing the child receives. Refs #4205 * fix(#4205): stop the host PATH leaking past snippetEnv on Windows Adversarial review (agy, gemini-3.8-flash-high) found that the fixtures' PATH isolation is defeatable on Windows regardless of which name the filter probes. Windows environment variables are case-insensitive but a spread of process.env is not: the host PATH enumerates as 'Path', so '{ ...process.env, PATH: isolated }' yields both keys, and libuv's make_program_env sorts the child's environment block case-insensitively without ever dropping duplicates. The child could therefore resolve the host PATH. snippetEnv() now drops every other casing whenever a caller supplies its own PATH. Also from that review: - buildIsolatedPath() takes the PATH to filter as a parameter, so (B0) and (B) no longer mutate process.env.PATH and no longer need try/finally restores. - The two loud-guard fixtures asserted 'not found' OR 'ERROR', which bash's own 'node: command not found' satisfies; they now assert the launcher's 'ERROR: gsd-tools.cjs not found'. - Corrected a comment counting three scrub keys as two, and two comments crediting a removed env argument for clearing ambient config dirs rather than snippetEnv()'s derived TEST_ENV_BASE. Refs #4344 * fix(#4205): make the fixtures' node shim work on Windows bug-211 (C) located node with `which node` through the process seam. `which` is not a Windows binary; the fixture only survived CI because Git Bash ships one. process.execPath is the same answer without the spawn, and (H) already used it. Both fixtures then built their node shim with fs.symlinkSync, which raises EPERM on Windows without developer mode or elevation — the same reason buildIsolatedPath() skips its own symlink step there. The shared linkNodeShim() helper hard-links instead on that platform (no privilege required) and falls back to a copy across volumes. Found by adversarial review (agy, gemini-3.8-flash-high). Refs #4344 * fix(#4205): make the launcher PATH isolation extension-aware and Windows-safe trek-e's review asks for a Windows-safe node fallback, an extension-aware filter, and a Windows regression test, in that order: broadening the filter first can strip the directory node itself lives in. buildIsolatedPath() now always prepends a directory holding node, on every platform, via the linkExecutable() helper (hard link on Windows, where symlinks need elevation). nodeBinDir is no longer nullable and the win32 early return is gone, so the fallback exists before the filter widens. (B0), the co-location invariant, therefore runs on Windows instead of being skipped on the one platform that had no fallback. The filter probes every name the launcher's `command -v gsd_run` arm can resolve. msys bash appends an executable extension during PATH lookup, so a directory holding only gsd_run.exe is reachable on Windows although gsd_run is absent. PATHEXT is folded in as well; it over-matches, which costs nothing now that node is always supplied separately. (B0) asserts every name in that set is filtered, and (B) plants a gsd_run.exe sentinel on Windows beside the extensionless one npm installs. The predicate lived in five hand-maintained copies — the drift that caused #4205 in the first place, and four of the copies pointed readers at a buildIsolatedPath() that was block-scoped out of their reach. buildIsolatedPath() moves to module scope and the four inline copies call it. The three shadow runBashFile() declarations this PR had to edit identically go with them. Red-proved both ways: probing 'gsd-tools' again turns (B0) and (B) red; making the node prepend conditional turns (B0)(ii) red. Refs #4344 * fix(#4205): drop empty PATH elements from the isolated PATH A POSIX shell reads an empty PATH element as the current directory, so an isolated PATH carrying one still lets the launcher's `command -v gsd_run` arm resolve a gsd_run from the fixture's own working directory — the leak class this file exists to close. Two ways one appeared. An ambient PATH containing `::` survived the filter, because `path.join('', 'gsd_run')` probes the working directory rather than a directory entry, so hasGsdRun could not see what it was admitting. And a PATH whose every entry was filtered joined to an empty string, leaving the returned value ending in a delimiter, which means the same thing. Empty entries are now dropped alongside the gsd_run-bearing ones, and the surviving directories are joined as a list, so a fully-filtered PATH yields the node shim dir alone. (B0) asserts both cases. Found by CodeRabbit on the fork rehearsal PR. Refs #4344 * fix(#4205): keep only absolute PATH dirs, and assert the sentinel by basename Adversarial review (agy, gemini-3.8-flash-high) on the previous head. An empty PATH element was dropped, but `.` and any other relative entry say the same thing explicitly and survived. hasGsdRun() cannot see what it would admit either: `path.join('.', 'gsd_run')` probes the runner's working directory, not the child's, so isolation also varied by where the suite was started from. Only absolute directories survive now, which can only tighten the isolation. (B0) asserts it over an empty element, a `.`, a relative entry, and a fully-filtered PATH — the previous empty-string assertion passed on the `.` case. (B)'s GSD_TOOLS assertion compared an absolute os.tmpdir() path against launcher output, which the file already documents as a mismatch on Windows: git-bash prints /c/Users/... where Node gives C:\Users\.... It never matched there, so it asserted nothing on the platform it was added for. It matches the mkdtemp basename now, which both path forms share. (B) also plants ONLY gsd_run.exe on Windows: with an extensionless sibling present, an extension-blind filter would strip the directory for the wrong reason and pass. snippetEnv() deduped case-variant keys for PATH alone. On Windows every scrubbed key has the same exposure — an ambient `bash_env` reaches the child beside the blanked `BASH_ENV`, and BASH_ENV re-injects the rest. Every key the function sets now wins over other casings of itself; keys it does not set are untouched, so a caller passing no PATH override still gets the host PATH. Refs #4344 * test(#4205): plant probe fixtures as files, not interpreter links Review follow-ups on 776e9371f. (B0)'s co-location fixture and its GSD_RUN_NAMES sweep only ever probe the planted names with accessSync; nothing executes them. They used linkExecutable, so on Windows the sweep hard-linked node.exe once per PATHEXT entry — a dozen on a stock runner, and a full copy each when os.tmpdir() and process.execPath sit on different volumes. plantExecutable writes a zero-byte 0o755 file instead. linkExecutable keeps the two callers that need a real executable: the node buildIsolatedPath prepends, and the Windows sentinel. The '/usr/bin:/bin' fallback formatted POSIX paths with the platform delimiter, yielding '/usr/bin;/bin' on Windows, which path.isAbsolute accepts and no Windows shell would ever produce. It was also unreachable: basePath defaults to process.env.PATH. An unset PATH now yields the node shim dir alone and fails loudly at spawn rather than being papered over. (B)'s header said the sentinel is planted in both forms; the code plants one per platform, and planting both on Windows is what the branch below it exists to avoid. Also names which assertion carries the Windows guarantee, since SENTINEL_INVOKED cannot fire there. The shared helper's temp dirs were prefixed gsd-891-, attributing every fixture's leftovers to one of the four bugs it now serves. Refs #4344 * fix(#4205): model bash's PATH lookup, not cmd.exe's, and give (B0) its own oracle Ponytail review on aec6012fc. GSD_RUN_NAMES expanded PATHEXT, which describes cmd.exe rather than the shell the launcher's `command -v gsd_run` arm runs under. The Cygwin/msys rule is that .exe may be omitted from a command while '.bat and .com ... you cannot omit the extension', so gsd_run.exe is reachable for a bare gsd_run and gsd_run.cmd/.ps1 are not, whatever PATHEXT lists. The comment claimed the resulting over-match was free. It was not: buildIsolatedPath() restores node to the isolated PATH, but nothing restores bash, which the fixtures spawn by name — so every extra dropped directory was another chance to remove the one bash lives in and fail with ENOENT instead of an assertion. Narrowed to gsd_run and gsd_run.exe. (Aside, the PATHEXT default does not even contain .PS1.) (B0)'s name-sweep took its list from GSD_RUN_NAMES, so it swept the constant under test with itself and could only catch that constant being deleted, never being wrong. Its relative-element case re-ran the implementation's own filter predicate over that filter's output, which is true for any predicate. Both now assert against written-out expectations: the reachable names per platform, and the exact directories that must survive each case. Also: the test name covered two of its four assertions, and the case table's prose counted three of its four entries. Red-proved three ways: dropping the gsd_run filter, dropping the absoluteness filter, and claiming a name the filter does not cover each turn (B0) red. Refs #4344 --------- Co-authored-by: Test <test@test.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
f09e7ed08c |
fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath (#4375)
* test(#4137): keg-only Homebrew Cellar path falls back to raw execPath Regression tests for the Homebrew branch of normalizeNodePath: the rewrite to <prefix>/bin/node must be existsSync-guarded like the mise/volta branches, falling through to the raw execPath when the keg-only formula was never linked into <prefix>/bin. Also makes the existing #3181/#2185 Cellar assertions hermetic by injecting existsSync stubs (granting existence to exactly the one candidate each asserts) so they no longer depend on the runner machine's real /usr/local/bin/node. * fix(#4137): existsSync-guard the Homebrew Cellar rewrite in normalizeNodePath The Homebrew branch of normalizeNodePath returned <prefix>/bin/node unconditionally — the only one of five runtime branches that never probed its rewrite candidate. On a keg-only or versioned Homebrew install (node@24 never brew-linked) that path does not exist, so every managed hook command baked by resolveNodeRunner/buildBakedNodeToken/ buildNodeRunnerChainToken failed at invocation with exit 127, /bin/sh: <prefix>/bin/node: No such file or directory. Guard the rewrite with the already-injected existsSync exactly like the mise and volta branches: when <prefix>/bin/node exists (linked formula) the rewrite is byte-identical to today; when it does not, fall through to the raw execPath — a working keg path instead of an immediately broken one. Also drops two now-unused constants from the regression tests. * test(#4137): make the #977 non-fnm Cellar assertions hermetic too The Bug #977 folded block's two 'still maps to stable symlink' assertions called normalizeNodePath without an existsSync stub, silently depending on the runner machine's real /usr/local/bin/node (present on the Linux bench image, absent for /opt/homebrew). With the #4137 guard these become environment-dependent; grant each exactly the one candidate it asserts. * chore(#4137): add changeset fragment * chore(#4137): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
2920bbc022 | fix(#4421): rescind #494's macOS full-matrix skip on changed test files (#4427) | ||
|
|
54085516c1 |
fix(#4211): materialize Kimi's agent tree recursively during surface apply (#4371)
* fix(#4211): materialize Kimi's agent tree recursively during surface apply kimiAgentsKind stages `gsd.yaml` + `gsd.md` + `subagents/gsd-*.{yaml,md}`, and install copies that tree recursively (_copyStaged). Surface apply fell through to _syncGsdDir's flat command/agent branch, which reads only top-level `*.md`: the YAML half and the whole subagents/ subtree were ignored, and `gsd.md` was written as `gsdgsd.md` because the flat branch re-applies kind.prefix to a name that already carries it. A surface change could therefore corrupt Kimi's installed artifacts while still reporting success. Three divergences from the install path, all in src/surface.cts: - _syncGsdDir gains a kimi-agents branch: recursive copy, then a prune scoped to exactly what install's _removeGsdEntries owns for this kind (the two root files, and gsd-*.{yaml,md} under subagents/). Everything else is user-owned and preserved. - applySurface stages kimi-agents WITH agentCtx and the `skills: '*'` rule for an unmodified full profile, as it already does for the agents kind and as createRuntimeArtifactInstallPlan does for every kind — without it Kimi's generated subagents lost their path-prefix rewrites and attribution trailer, and an unmodified full profile staged only the skill-referenced subset. - applySurface runs rewriteStagedSkillBodies for kimi-agents, which the install plan routes through it alongside skills. * chore: add changeset for #4211 --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
acb3cc974b |
fix(#4197): dedup the update-context fast path against the selected global dir (#4413)
* fix(#4197): dedup the update-context fast path against the selected global candidate The preferredConfigDir fast path derived scope from a cwd-relative match alone, so a global install reported LOCAL whenever the shell sat in $HOME — and run_update then drove the installer through its --local arm (settings.local.json + the #338 relocation) against a global install. Extract resolveGlobalCandidate (env candidates first, then $HOME-relative, first hasInstall hit wins) and use it in BOTH paths: the fast path now answers LOCAL only for a cwd-relative match that is not the selected global dir, which is the same dedup the cascade applies at its isLocal check. A preferred dir that is also the env-directed global now answers GLOBAL on both paths (the cascade's answer), pinned by a parity test. The discriminator is the selected global candidate, not the $HOME pathname: with CLAUDE_CONFIG_DIR directing the global elsewhere, $HOME/.claude probed from cwd === $HOME is a genuine local install, and a pathname check would re-break parity (regression-pinned). * chore(#4197): add changeset * chore(#4197): backfill PR number in changeset --------- Co-authored-by: agent-4197 <agent-4197@gsd.local> |
||
|
|
ed133cc116 |
enhance(#3085): add Grep to allowed-tools for 21 skills (#4397)
* enhance(#3085): add Grep to allowed-tools for 21 skills
29 of 71 skills omit Grep from allowed-tools, forcing Bash grep for
structured search instead of the dedicated tool. Adds Grep to the
21-skill subset confirmed safe in prior review (excludes the 6
gsd-ns-* dispatchers, gsd-help, and gsd-surface, which have no
plausible structured-search need).
Hand-edits commands/gsd/*.md only; skills/*/SKILL.md is regenerated
via `npm run gen:plugin-skills` from that source. Updates the one
hardcoded copilot-install test assertion affected by gsd-health's
new tool order.
* chore(#3085): backfill changeset PR number
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* test(#2618): assert the Needs clause on the structured field, not the path-length-dependent bullet
The bullet embeds the todo file's ABSOLUTE path, so its total length
varies by runner tmpdir: on macOS CI, /private/var/folders/… plus the
test harness's gsd-test-run-* wrapper pushed the full bullet to 244
chars — past renderPendingTodoBullet's intended 240-char cap, whose
documented first degradation step drops the Needs clause. The product
behavior is correct (#2618 design); the assertion was runner-dependent.
The extraction is now pinned on json.todos[0].needs; the rendered-bullet
shape stays covered by the path-independent title assertion and the
renderer's own unit rows.
Found blocking #4186's CI on the macOS shard (test landed 30 minutes
earlier in
|
||
|
|
b7917882bb |
fix(#4398): render the pending-todo bullet link repo-relative (#4416)
* test(#4384): failing-first regression rows for the macOS long-base todo-cap failure The 240-char pending-todo bullet cap must be deterministic w.r.t. where the repo is checked out. Deterministic long-base-path fixtures (a single 110-char segment, no real macOS dependency) reproduce next's own macos shard 3/3 failure (run 34038716700) on every OS: with an absolute link the bullet exceeds the cap and the documented needs-first truncation drops the 'Needs <solution>' clause. Rows cover the determinism property (byte-identical bullets under short and long bases), the CLI surface, relative-path stability, legacy no-projectRoot behavior, drop-order preservation, and adversarial edges (outside-root, path===root, non-string path). * fix(#4384): render the pending-todo bullet link repo-relative renderPendingTodosMarkdown gains an optional projectRoot; when given and the todo's path is absolute, the bullet's [todo file](…) target becomes toPosixPath(path.relative(projectRoot, path)) — the idiom already used for project_exists. cmdInitTodos passes cwd. The JSON todos[].path field stays absolute (#2376). Only the rendered display link changes: embedding the machine-variable absolute base let macOS's /private/var/folders/… temp paths consume the 240-char budget and drop the 'Needs' clause on long-path machines only — next's own macos-latest shard 3/3 went red on exactly this (run 34038716700), Linux's short /tmp passed. The 240-char whole-bullet cap and the needs→title→area drop order are unchanged; this matches PR #4384's own canonical example, docs, and unit tests, which all show repo-relative links. Docs updated at all three surfaces that describe the bullet (COMMANDS.md, templates/state.md, reference/state-md.md — the last was still pre-#4384 'count and reference' prose). Fixes the macOS regression introduced by #4384; next is red on its own CI. * test(#4384): fix substring false positive in the outside-root regression row The ../-form relative link legitimately contains the absolute path as a substring, so !line.includes(absolutePath) fired on correct output (caught by the first remote verify run, linux-node24 44018/44019). Assert the property itself instead: extract the link target and require it to be non-absolute and not equal to the absolute path. * chore(#4398): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
aad04f4e9a |
docs(#4400): ADR-4139 — the compact-content seam (#4410)
* docs(#4400): ADR-4139 — the compact-content seam Phase 0 of epic #4139. Locks the design before any code lands. #4139's stated mechanism cannot reach the stream it exists for: 58 of 72 shipped commands deliver their whole workflow file through an eager @-include, which the host expands before any project config is in context. An in-content gate is evaluated after those bytes are already paid. The ADR declines the obvious fix (convert the 58 execution_context blocks to runtime Reads) because that removes the host guarantee for every user, not only opted-in ones — a global install shares one skill tree, so an @-include cannot be conditional. It instead keeps every @-include exactly where it is and splits what sits behind them: the canonical path becomes a runnable spine, elaborations move to a sibling detail file read at runtime. A missed Read then degrades to "runs correctly with fewer tokens", never to "runs with no instructions". Also records: the rename to workflow.compact_content, the re-pitch onto ADR-1610's context-rot argument rather than the cost argument ADR-1610 discounts, partition-not-duplication (which dissolves the dual-maintenance cost the Feature Review called disqualifying), the guard-scope and NEW_FILE_CAP mapping for the new subtree, and an argued reconciliation of the acceptance criteria this design does not meet literally. Corrects ADR-3646 §Context: it cites #3647 as open; #3647 closed 2026-09-01 as a duplicate of #3606. ADR-3646's Decision is unaffected — it explicitly disclaimed any dependence on #3647's state. The residual prose-dispatch variance named in #3647's own closure thread is unresolved, and this ADR routes around it rather than assuming it away. Refs #4139 Closes #4400 * docs(#4400): fold the two orthogonal review findings into ADR-4139 Code review (isolated context) and security review (isolated context) both returned findings. Fixed here rather than carried. Critical, from code review: the ADR repeated earlier research's claim that discuss-phase, manager and pause-work all reach a workflow by runtime Read. manager and pause-work carry plain eager @-includes and are inside the 58, not outside. discuss-phase is the only precedent, and it is one file. The Open Questions section is corrected with it. The NEW_FILE_CAP mapping was wrong in a way that changes the layout. It lives at tests/helpers/emitted-diff.cjs:96, not in workflow-size-budget, and its own doc comment records that it is a hard cap, not ack-able, and NOT tier-exemptible -- the pre-#2724 test-file version was. So a single detail.md holding plan-phase.md's elaborations is blocked outright with no exemption path. Detail content is now one or more parts under workflows/<name>/detail/, each below the cap, named by the spine in the dispatch-table shape discuss-phase.md already uses. commit-files-pathspec is in scope and earlier research called it irrelevant. Per CONTRIBUTING.md:1164-1170 it sweeps every .md under gsd-core/workflows/ for unscoped commit-seam invocations. Added to the guard table. Two byte figures were inherited rather than measured, against this ADR's own evidence note. Templates and agents re-measured; the table now carries the method and the exact numbers. From security review: "a spine that has shed a protected-content marker fails" never defined what a marker was, leaving the strongest check in the set resting on a prose-category judgment. Protection is now a literal greppable sentinel in the existing gsd: comment namespace, and the guard rule has no discretion in it. Also added: an explicit statement that the detail path is never user- or project-supplied and cannot be shadowed by a project-local file, and an exact-version pin commitment for gpt-tokenizer. Code review also found a real hole in the central fail-safe argument: spine sufficiency is verified once at split time and never again, so load-bearing procedural text carrying no sentinel could later drift into a detail part with every check green. Sufficiency is not machine-decidable, so a fifth ongoing check is added -- a spine that loses lines which reappear in its parts fails unless the PR declares the boundary move. The ADR now states plainly that this is authoring discipline with a forced checkpoint, not a structural invariant, and that the partition relocates the Feature Review's cost rather than fully eliminating it. Refs #4139 Refs #4400 --------- Co-authored-by: sim <sim@local> |
||
|
|
708d9a0b82 |
fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file (#4388)
* test(#4187): bare VERIFICATION.md regression matrix for the status surface Both query verbs must agree on every row: bare file, suffixed variants, missing file, other-dir placement, and the staleness seam. Row 1 is the failing-first regression from the issue repro. * fix(#4187): status reader resolves a bare VERIFICATION.md like resolve-file readVerificationStatus and its internal staleness check (findStaleVerificationSummary) called the shared resolver without allowBare, so a phase whose only report was a bare VERIFICATION.md read as missing and was told to re-run execute-phase while verification.resolve-file, determinePhaseStatus, and both init verification_path projectors all resolved the same file. Both call sites now pass allowBare: true, matching the other five; tier order (dashed > bare) is unchanged, so only bare-only directories change behavior. * fix(#4187): correct call-site counts in allowBare docblocks Adversarial review caught the comments claiming five of six call sites opted in; the current tree has six call sites with four previously passing allowBare — the two module-internal status-path sites were both holdouts, not one. * chore(#4187): changeset for the bare VERIFICATION.md status fix * chore(#4187): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
fd4aac5670 |
fix(#4192): honor explicit model pins on the claude runtime (#4396)
* fix(#4192): honor explicit model pins on the claude runtime Two documented model-configuration contracts did not hold on the claude runtime (confirmed-bug scope from the issue triage): Finding 1 — model_profile_overrides.claude.<tier> was inert. Step 3 of resolveModelInternal gated runtime-aware tier resolution on configRuntime !== 'claude', so the key's only reader was never consulted, while workflows/settings-advanced.md writes it for claude-runtime users. A new step 4.5 resolves ONLY the user's override entry (never the builtin claude tier map, so unpinned installs keep resolving aliases). An override value that maps to a current tier alias collapses to that alias (byte-equivalent, the #2041 protection); anything else — a pinned older generation, a bare alias repoint, a non-Anthropic id — resolves verbatim. It sits after the resolve_model_ids:'omit' gate so an explicit project omit still wins (#2297) and before the alias return so resolve_model_ids:true cannot re-materialize the pin to the latest id. Finding 2 — fully-qualified claude-* ids in model_overrides were warn-dropped to tier resolution (mapClaudeOverrideForRuntime unmappable branch, #2041), while the docs promise any fully-qualified model id is valid. The unmappable branch now passes the pin through verbatim with a warn-once breadcrumb (text describes the pass-through). Dropping it silently unpinned the operator's explicit choice — the exact 'profile can misrepresent what actually runs' defect of #4192. Mappable ids and non-claude values behave exactly as before; resolveModelForTier shares the mapping; the tier honesty signal is unchanged (raw ids still report 'unknown'); the model_policy path is untouched. Docs updated to the agreed contract (CONFIGURATION.md false 'Claude example' corrected; how-to + shipped reference document the pin semantics, the fable alias, and the tier-override composition). * test(#4192): pin explicit model pin resolution on the claude runtime 28 failing-first rows across the resolver seam and the resolve-model CLI: pinned-generation fidelity (tier override + per-agent verbatim pins, object form, explicit runtime), unpinned controls byte-stable (no override, other runtime/tier, inherit, project omit, precedence), adversarial rows (prototype-chain keys, malformed values, warn-once dedupe, 64-char stderr cap), and behavioral AC1/AC2 rows through runGsdTools. The stale #2041 fall-through assertions now pin the pass-through contract; mappable-id collapse assertions unchanged. * chore(#4192): add changeset fragment * chore(#4192): backfill PR number in changeset fragment --------- Co-authored-by: ZCode <zcode@localhost> |
||
|
|
b7406b293f | enhance(#2618): render pending todos as one bounded bullet per todo (#4384) | ||
|
|
66e4034fe4 |
fix(#4138): begin-phase without --phase exits non-zero and writes nothing (#4380)
* test(#4138): failing-first regression — begin-phase without --phase must fail closed * fix(#4138): begin-phase without --phase exits non-zero and writes nothing * chore(#4138): changeset fragment for begin-phase arg validation * chore(#4138): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
03738824de | enhance(#2586): stop installing Codex context-monitor hooks without metrics (#4367) | ||
|
|
0aa4202f6a |
fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening (#4376)
* test(#4135): regression rows for pristine regen coverage collapse RED skeleton: src/pristine-baseline.cts exports findPristineInGit as a null-returning stub (wired into verifyFile after the #4145 orphan tier, behavior-neutral) so the git-history rows fail behaviorally, not at require time. Failing-first rows: baseline_covered aggregate on a 1-of-13 multi-version fixture, coverageHeadline typed renderer, the opt-in --min-baseline-coverage gate (exit 3, >= threshold semantics, vacuous-pass and malformed-value boundaries), git-history baseline recovery (dropped-line catch + surviving-line verify + older-commit hop), findPristineInGit unit, Step 5a workflow headline contract, and the installer-side describeBaselineCoverage honest N-of-M summary with the collapse disk-state pinned. Negative-space rows pin today: non-git ok_no_baseline posture, no-match-no-adoption, #3657 drift never rescued, canonical precedence, and no git tier without --pristine-dir. * fix(#4135): headline baseline coverage, opt-in strict gate, git-history widening The #3407 promotion rule regenerates gsd-pristine/ baselines from the INCOMING release source and keeps only candidates byte-identical with the OUTGOING recorded hash — correct in isolation, but on a multi-version jump the surviving set is precisely the files upstream did NOT change. The verifier then reports ok_no_baseline (advisory, exit 0) for everything else, and no surface distinguishes a 12-of-13-unverified green run from a fully-verified one: the human summary printed Checked/Failures only, the JSON had no coverage aggregate, and the installer's update output gave per-bucket counts without N-of-M framing. All three issue directions, none exclusive: - Report coverage prominently: --json gains an additive baseline_covered aggregate; the human summary leads with 'Baseline coverage: N of M file(s)...' on every run plus an advisory section naming each skipped file and reason; the installer prints an honest covered-of-modified line via the exported describeBaselineCoverage helper (typed return, exact contract); workflow Step 5a computes and prints the headline before any pass/fail framing. - Fail louder on low coverage: opt-in --min-baseline-coverage <0..1> exits with new documented code 3 when coverage falls below the threshold (>= semantics; empty run vacuously passes; content failure exit 1 outranks it; malformed values are usage errors, exit 2). Default posture unchanged — no_baseline stays advisory per #934. - Widen the promotion rule (its only trustworthy form): when no baseline resolves under gsd-pristine/ and a hash is recorded, the verifier now recovers the baseline from the config dir's own git history — the workflow's documented Option A — anchored by the same authority every tier trusts, exact pristine_hashes sha-256 equality. Read-only (git log/git show, windowsHide per #685), bounded (100 commits/file, 10s/subprocess), null-on-any-failure so ok_no_baseline remains the universal fallback. Tier order: canonical join -> #4145 orphan scan -> git history -> OK_NO_BASELINE; #3657 drift and canonical precedence untouched. Hash validation in saveLocalPatches is NOT relaxed — the collapse is legitimate conservatism; hiding it was the bug. Measured on the issue's shape (13 files, 12 changed upstream, 1.10->1.12): non-git installs report baseline_covered 1/13 with the headline and can gate at exit 3; a git-managed config dir with the outgoing bytes in history verifies 13/13. Review fixes folded in: workflow headline derives the unverified count from checked - baseline_covered (not the drift+no_baseline sum), and the new site-scoped allow-test-rule annotation carries its ADR-456 see-ref on the marker line. Emitted-Drift-Ack-Growth: reapply-patches.md — #4135 — +20 lines / ~1.5 KB, prose and bash only: two additive parse lines (BASELINE_COVERED, CHECKED_COUNT), a Step 5a coverage-headline block printed BEFORE any pass/fail statement (documents the opt-in --min-baseline-coverage exit-3 gate), and one Option B sentence noting the verifier's read-only git-history fallback. No step ordering, gate, tool-invocation, or dispatch shape changed; 5a's fail/drift/advisory handling is unchanged, the headline only precedes it. * chore(#4135): backfill PR number into changeset fragment --------- Co-authored-by: agent-4135 <agent-4135@gsd.local> |
||
|
|
c95b734145 |
fix(#4136): compute the Incorporated status; stop re-grafting superseded patches (#4373)
* test(#4136): failing-first rows for the unreachable incorporated status Folded block bug-4136-reapply-incorporated-status locks the --classify contract (incorporated / needs_merge / unknown, never-incorporated guards, cycle end-to-end) plus the workflow-contract rows; REASON gains OK_UNVALIDATED_BASELINE in both shape-locks; the #2994 invocation count moves 1 -> 2 (classify + gate). All red until the verifier grows --classify and the workflow consumes it. * fix(#4136): compute the incorporated status; stop re-grafting superseded patches Add --classify pre-merge mode to the deterministic verifier: with a hash-validated pristine baseline, a file whose every significant user-added line is already present verbatim in the freshly installed version is classified incorporated (new frozen CLASSIFICATION enum, structured --json report, always exit 0 — the binding gate stays the post-merge run). Drifted (#3657), absent (#934), unvalidated (new OK_UNVALIDATED_BASELINE), and no-baseline runs classify unknown — a false incorporated silently retires a live customization, so only a confirmed baseline may ever confirm adoption. The baseline-resolution block moves out of verifyFile into a shared resolvePristineBaseline so the gate and the classifier cannot drift on what counts as a usable baseline; gate behavior is byte-identical. reapply-patches.md step 4 gains the pre-flight classifier invocation and the not-re-grafted contract: incorporated files are left exactly as shipped (their hash then re-converges with the manifest, ending the backup cycle), statuses feed steps 3/7, and the merge rules gain the already-present-verbatim arm. Emitted-Drift-Ack-Growth: reapply-patches.md — step 4 pre-flight classifier section, the documented Incorporated status needed its deterministic contract * fix(#4136): address review findings on the workflow contract Drop the unused INCORPORATED_COUNT shell variable (standards pass) and close the all-files-incorporated gap in the hunk-table guidance: emit a header row plus a note line so the step 5b absent-table halt is not tripped when nothing was merged (spec pass). Emitted-Drift-Ack-Growth: reapply-patches.md — step 4 pre-flight classifier section, the documented Incorporated status needed its deterministic contract * chore(#4136): backfill changeset fragment with PR 4373 * test(#4136): lock that a 4145-recovered baseline can confirm incorporation The hash-first orphan recovery merged with next (PR #4364) lands in the shared resolvePristineBaseline as a validated resolution; this row pins the composition so a future change cannot quietly downgrade recovered baselines to unknown and silently disable incorporated detection for prefix-less installs. --------- Co-authored-by: sim <sim@local> |
||
|
|
7bb366e836 |
fix(#4130): --context flag for check decision-coverage-plan + parseDecisions quadratic-backtracking hardening (#4374)
* test(#4130): failing-first regressions for --context flag + parseDecisions hardening Block A (flag): check decision-coverage-plan --context <path> must route identically to the positional form; flag wins over positional context; valueless --context falls through to the #2770 fail-closed caller error; verify keeps its positional surface (flag is plan-only). RED on base: the flag token lands in the args[2] phase slot (false uncovered) or the args[3] context slot (silent CONTEXT.md-missing skip). Block B (hardening): regex-lattice asserts pin the atomic-ID wrapper (?=(X))\1 and the em-dash first-separator narrowing [^*—–]*[—–] plus the no-adjacent-overlap property; a differential property compares the module against a frozen copy of the pre-hardening grammars (reference validated against the base build: 60k generated lines, 0 mismatches); 40k cliff shapes assert correct outcomes with no wall-time asserts (repo rule). A12: partitionPredicateArgs keeps one parser behind parsePredicateFlags. * fix(#4130): --context flag for check decision-coverage-plan + quadratic-backtracking hardening in parseDecisions (A) check decision-coverage-plan --context <path> — sibling convention (check predicate, #2008): --flag value pairs parsed by the new shared partitionPredicateArgs (parsePredicateFlags reimplemented as its flags half — one parser, cannot diverge), the flag winning over a same-purpose positional, positionals kept (no sibling deprecates them; the plan-phase workflow caller passes positionals), valueless --context falls through to the #2770 fail-closed caller error. Repair of the routing accident where --context landed in the args[2] phase slot (false uncovered) or the literal token in the args[3] context slot (silent green skip). (B) parseDecisions regex seam hardened, byte-identical on all legal inputs: the three bullet grammars consume the ID atomically via the (?=(X))\1 lookahead emulation (kills the tail/[^:*]* O(n^2) re-split, ~1.1s @ 40k), and the em-dash first separator narrows [^*]*[—–] to [^*—–]*[—–] (kills the dash-position O(n^2) retry, ~1.7s @ 40k). Group indices unchanged (handlers untouched). Pinned by regex-lattice tests, a differential fast-check property vs the frozen pre-hardening grammars, and 40k cliff/legal-shape outcome tests (no wall-time asserts per repo rule — no deterministic engine step counter exists in Node). * docs+test(#4130): document --context invocation; harden lattice test tooling - docs/CONFIGURATION.md Decision Coverage Gates: new 'Invoking the plan gate directly' block documenting both the positional and --context forms, flag precedence, and the valueless-flag fail-closed semantics (same place the gate's behavior is documented; sibling check predicate documents its flags the same way). - Two changeset fragments per the maintainer brief (Added: flag; Fixed: hardening), PR numbers to be backfilled. - tests/decisions.test.cjs review fixes: readRegExpTemplate template escaping (bare ')' SyntaxError), range-aware lattice checker with backreference skip and template unescape, honest A1 contract, lint escape warning. * fix(#4130): valueless --context fails closed per #2770; A8 isolates flag-vs-positional context Suite-caught fixes from the first verify run: - cmdDecisionCoveragePlan now refuses a flag-shaped token as the positional context path: a bare valueless --context stays a positional (sibling parser semantics, unchanged) but reading it as a PATH would turn a caller mistake into a silent 'CONTEXT.md missing' green skip — exactly what #2770's fail-closed law forbids. Now falls through to the missing-context-argument error, as documented. - A8 test compares decoy-positional+flag against flag-with-phase (phase held constant) so the row isolates WHICH context was read; the old form compared against a no-phase invocation that could never match. * chore(#4130): backfill PR number in changeset fragments (PR #4374) --------- Co-authored-by: sim <sim@local> |
||
|
|
6adf3098ac |
fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans (#4364)
* test(#4145): regression rows for hash-matching prefix-less pristine baselines RED skeleton: src/pristine-baseline.cts exports findPristineByHash as a null-returning stub so the new rows fail behaviorally, not at require time. Failing-first rows: verifier resolution (no_baseline must drop to 0 when an exact-hash orphan exists), findPristineByHash unit row, and the two saveLocalPatches relocation rows. Negative-space rows pin today's behavior: missing baselines still report ok_no_baseline, mismatching orphans are never adopted or deleted, canonical precedence and the #3657 drift posture are untouched. * fix(#4145): resolve gsd-pristine/ baselines by recorded hash, relocate orphans Both pristine readers joined the manifest-keyed path strictly, so a snapshot stored without the gsd-core/ prefix (an earlier release's writer) was reported as ok_no_baseline by the verifier and pushed into regeneration by saveLocalPatches — where incoming-release candidates can never satisfy the recorded outgoing hash, leaving the correct baseline permanently unconsumed. - src/pristine-baseline.cts (new, ADR-457): shared findPristineByHash — deterministic sorted scan of gsd-pristine/, exact sha-256 equality with the recorded pristine_hashes entry (the same authority the #3657 drift guard trusts), symlink-skipping, canonical path excluded via skipRel. - verify-reapply-patches.cjs verifyFile(): on canonical miss with a recorded hash, adopt byte-identical content found anywhere under gsd-pristine/ before reporting OK_NO_BASELINE. Drift posture (#3657), canonical precedence, and the frozen REASON/report shapes are untouched; the verifier stays read-only. - install.js saveLocalPatches(): preserve-check rescue — relocate a hash-matching orphan to the canonical path (copy, hash-verify, then remove the orphan) so the state self-heals on the next update instead of repeating forever. Honest accounting: new non-overlapping rescued counter. - Workflow doc: one-sentence note on hash-based snapshot resolution. - Derived ripples: INVENTORY-MANIFEST.json regen, eslint ignore + .gitignore entries for the compiled artifact, seedFixture mkdir fix in the new rows. Emitted-Drift-Ack-Growth: reapply-patches.md — one-sentence note on hash-based pristine snapshot resolution (#4145) * fix(#4145): review follow-up — orphan scan never consumes a canonical path Adversarial review finding: with two modified files sharing byte-identical outgoing content, recoverOrphanedPristine could adopt the OTHER file's canonical pristine as its rescue source — relocating it (copy + delete at its home path) and ping-ponging the single baseline between the two files across updates. findPristineByHash's skip parameter now accepts a Set, and saveLocalPatches passes the normalized manifest keys so every canonical path is excluded; only genuine non-canonical orphans are eligible for removal (no strict-join reader ever consults those). Adds the canonical-theft regression row, a Set-skip unit assertion, and tightens the workflow doc sentence the same pass flagged as overstated. * fix(#4145): INVENTORY roster row + symlink-fixture correction Two leftovers from the ab17b7a1e5 bench run, both root-caused: - docs/INVENTORY.md roster row for cli_modules/pristine-baseline.cjs (#3762 gate: every manifest entry carries a row). - The findPristineByHash symlink unit fixture placed its symlink target INSIDE the scanned root, so the walk legitimately matched the real target file. The implementation skips the symlink itself; the fixture now keeps the target outside the scanned tree so the assertion tests what it claims. * changeset(#4145): fixed fragment for pristine baseline hash resolution --------- Co-authored-by: gsd-agent <agent@gsd.local> |
||
|
|
d5a85da8ab | fix(#4363): bump download-artifact and setup-node off node20 runtimes (#4365) | ||
|
|
c3e2da153b |
fix(#4134): refuse punctuation-only milestone heading names (#4358)
* test(#4134): fail-first regression — refuse punctuation-fragment milestone names A first-milestone ROADMAP.md H1 that puts the version after the name (# Roadmap: Project — Name (v1.13)) leaves exactly ')' after the heading's own version token, which the ADR-3180 §7.2 pinned name rule returns as a COMPLETE-scope milestone name. Failing-first coverage: - getMilestoneInfo: name-then-version H1 (STATE-anchored + ROADMAP-only fallback) must yield TRUNCATED {version, name: null}, never ')' - the refusal is level-agnostic (H2/H3) - punctuation-family remainders (')', '()', '**', '.,;:', ']}', emoji-only) - listMilestoneHeadings enumerates the heading with name: null - init manager CLI reports milestone_name: null and no lone ')' anywhere - property (seed 20260905, 300 runs): a word-char remainder is always a name, a punctuation-only remainder never is - negative space: canonical delimiter forms, parenthetical names (#3171), trailing markers, digit-only names, CRLF headings, version-last-no-parens control * fix(#4134): refuse punctuation-only milestone heading names extractMilestoneHeadingName returns everything after the heading's own version token as the name (ADR-3180 §7.2 pinned rule), which assumes version-then-name. A name-then-version heading — the H1 a first-ever ROADMAP.md drifts into ('# Roadmap: Project — Name (v1.13)') — leaves exactly ')' after the token, and that fragment was returned as a COMPLETE-scope milestone name, propagating into init.* JSON output and buildStateFrontmatter's STATE.md writes. A remainder with no letter or digit anywhere (any script) is heading structure, not a curated name: refuse it as name: null so callers report the honest §7.2 rule-6 answer (version kept, TRUNCATED scope). Names that merely contain punctuation are unaffected — '(' stays an ordinary name character (#3171) — and digit-only names qualify. Also closes the template gap that lets the shape occur: the roadmapper agent's output_formats now templates the version-free canonical H1 ('# Roadmap: [Project Name]', per templates/roadmap.md) instead of leaving a first milestone's title line to invention. The new section shifts the file's existing bare-gsd-tools prose mention from line 647 to 660, so its line-keyed PROSE_ALLOWLIST entry moves with it. Emitted-Drift-Ack-Growth: gsd-roadmapper.md — deliberate +498 bytes: new '### 0. Top-Level Title (H1)' output_formats section templating the canonical version-free H1, closing the first-milestone template gap that lets an H1 drift into 'Name (vX.Y)' and corrupt milestone_name extraction (#4134) * chore(#4134): add changeset * chore(#4134): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
e6d047decc |
fix(#4129): derive completed_phases from the ROADMAP authority; honor the progress-ratchet on every state write (#4359)
* test(#4129): failing-first regressions — completed_phases clobber on resyncing writes and phase-complete failure to increment * fix(#4129): completed_phases derives from the ROADMAP authority and the write path honors the progress-ratchet Three coordinated prongs (diagnosis in .gsd/bug/fix-4129-completed-phases-recompute/): P1 — buildStateFrontmatter's disk scan floors the completed-phases numerator at the milestone-scoped ROADMAP Complete-row count (deriveProgressFromRoadmap, the one owner), gated inside the same safeToUseRoadmapCount / not-withheld branch that owns the denominator. A completed phase whose verification routes stale (#2348 clean-commit-time drift) or is missing no longer under-counts forever. P2 — applyPreserveAlways's resync arm merges instead of wholesale-replacing on a measured scan: totals derived both directions (#2440), completed counters up-only (#2969 — the schema-declared progress-ratchet, now enforced on the write path like the read path always has), percent recomputed from the merged counters. The #3756 unmeasured guard and the #3242 explicit-progress contract are unchanged. P3 — phase complete's atomic 3-file commit passes the post-completion ROADMAP-derived counters through the #2736 authoritativeFm seam (new object direction for the progress key; completedOnlyRaise at the post-preservation re-assert), because the transaction's disk scan reads the pre-completion ROADMAP and failed to increment on the completing phase's own write. * fix(#4129): adversarial-review hardening — intent is a floor at BOTH authoritativeFm sites The pre-preservation merge could lower a correctly-higher disk-derived counter (a verification-passed phase whose ROADMAP table row drifted behind the disk signal). completedOnlyRaise now governs both application sites: the intent and the derivation agree on direction (up), never on subtraction. * fix(#4129): the ratchet merge keeps derived values verbatim when numerically equal The re-parsed derived block carries string scalars ("2") while the curated snapshot carries numbers (2); substituting the curated spelling over an equal derived one was a no-op in substance but a shape churn the ADR-3473 §8.7 reporting loop surfaced as a phantom preserved-over-disagreeing-derived warning on phase complete (ADR-3408 §8.5 Matrix B). Only a strictly-greater curated counter replaces the derived value now; percent gets the same verbatim rule. * changeset(#4129): backfill PR 4359 --------- Co-authored-by: sim <sim@local> |
||
|
|
bd75d42f52 |
Merge pull request #4362 from open-gsd/chore/backmerge-main-to-next-b67f6028c
chore: back-merge main → next (
|
||
|
|
0293afd108 |
chore: back-merge main into next (b67f6028c)
|
||
|
|
519bb60263 |
Merge pull request #4361 from open-gsd/chore/sync-next-version-1.13.0
chore: sync next package version to 1.13.0 |
||
|
|
b3906c66f6 | chore: sync next package version to 1.13.0 | ||
|
|
b67f6028ce |
Merge pull request #4360 from open-gsd/release/1.13.0
chore: merge release v1.13.0 to main |