bf4485ada2e7b683ce737b7bae01a252ed327946
180 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4dfc46bbe7 |
enhance(#3348): add a context-drift pre-check gate to plan-phase (#4147)
* test(#3348): add failing-first coverage for the context-drift gate * feat(#3348): add context-drift pre-check gate for plan-phase Compares each phase's *-RESEARCH.md/*-PATTERNS.md/*-VALIDATION.md/*-SPEC.md effective last-changed time (git commit time, falling back to mtime for uncommitted edits) against *-CONTEXT.md's, so plan-phase no longer silently reuses an upstream artifact that predates a decision added to CONTEXT.md after that artifact was derived from it. Deterministic, no model call. New `gsd_run verify context-drift <phase>` command, sibling to the existing verify.codebase-drift/verify.schema-drift gates in the drift capability. Warn-only by default (workflow.context_drift_precheck), with an opt-in workflow.context_drift_action: block escape hatch. Wired at plan:pre in plan-phase.md, before both the RESEARCH.md and PATTERNS.md reuse decisions. * fix(#3348): address code-review findings — raw-text-match, stale comment, import placement, duplicated phase resolution * fix(#3859): pin the real commit's diff.ignoreSubmodules to match the empty-diff probe The #3859 empty-diff guard decides whether a submodule bump would land using `--ignore-submodules=dirty`, overriding the caller's `diff.ignoreSubmodules` config. The real `git commit -- <paths>` that follows was never given the same override, so under a bare `diff.ignoreSubmodules=all` repo config the two calculations disagree: driven on git 2.39.5 (Debian bookworm, the linux-node24 test-matrix image), the guard correctly stands aside but the scoped commit itself then silently fails (exit 1, no error text) for a gitlink bump it had just confirmed would be recorded, surfacing as commit_failed instead of committed:true. Pin `-c diff.ignoreSubmodules=dirty` onto the scoped commit call too, so the probe and the commit it protects can never diverge. Harmless when no submodule path is involved (driven: identical outcome on an ordinary scoped file, with and without the flag). * fix(#3348): guard resolvePhaseDirByToken's exact-match fallback against path traversal * fix(#3348): retarget phase-enumeration-drift exemption to the consolidated resolvePhaseDirByToken helper cmdVerifySchemaDrift's inline readdirSync was already function-scoped-exempt in lint-phase-enumeration-drift.cjs as a single-phase LOOKUP (not a current-milestone enumeration). This PR's refactor pass lifted that block into a shared helper, resolvePhaseDirByToken, also used by the new cmdVerifyContextDrift — the guard tracks exemptions by enclosing function name, so the readdirSync now lives in an unexempted function and started firing. Move the exemption to resolvePhaseDirByToken (same written reason, now covering both callers) instead of migrating to listAllPhaseDirs, which would introduce two real behavior deltas here: it catches readdirSync failures internally (old code let them throw) and sorts results by phase number before matchPhaseDirs picks matches[0] (old code used raw, OS-dependent readdirSync order). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): satisfy lint:ci — slash form, capability registry regen - docs/features/context-drift-gate.md used the deprecated /gsd: colon form; docs are never passed through the install-time slash-form converters, so lint-docs-command-form requires the hyphen form. Regenerated docs/FEATURES.md from the corrected fragment. - Regenerated gsd-core/bin/lib/capability-registry.cjs after editing capabilities/drift/capability.json (lint:generated-sync). * fix(#3859): pin the real commit's diff.ignoreSubmodules via env, not argv -c The prior fix pinned `-c diff.ignoreSubmodules=dirty` onto the scoped commit's argv via `commitArgs.unshift(...)`. `-c key=val` must precede the `commit` subcommand, so this shifted `commitArgs[0]` from `'commit'` to `'-c'` for every scoped commit call, breaking 17 position-based assertions in the commit-files pathspec regression suite that read `a[0] === 'commit'` to find the commit invocation among recorded git calls. `execGit` already accepts an `env` option merged onto `process.env` before spawning. Git honors `GIT_CONFIG_COUNT`/`GIT_CONFIG_KEY_0`/`GIT_CONFIG_VALUE_0` as a per-invocation config override functionally identical to `-c key=val`, expressed via env instead of argv. Passing that env alongside the existing commitArgs (still `['commit', ..., '--', ...stagedPaths]`, argv unchanged) fixes the real commit's effective diff.ignoreSubmodules to match the empty-diff guard's probe without moving anything in argv position 0. Scoped to exactly the canScope branch, matching the probe's own preconditions and leaving no behavior change for commits the probe never evaluated. No test file changes needed — the 17 previously-failing assertions test argv[0] against the array passed into execGit, which never changes. * fix(#3348): register verify-context-drift in the check subcommand router The drift capability's new plan:pre gate declares check.query "verify.context-drift", which normalizes to `check verify-context-drift`, but no such subcommand was routed — phase6-capstone-conformance's uniform-block-field test failed with "Unknown check subcommand" for every declared gate query. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): extend #1592's exact-key-list snapshot for the new context-drift config keys tests/capability-registry.test.cjs asserted an exact, hardcoded snapshot of the drift capability's config keys. #3348 legitimately adds two new keys (workflow.context_drift_precheck, workflow.context_drift_action) for its own plan:pre context-drift gate — extend the expected set (and clarify the assertion message) without weakening the test's exactness. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): reconcile E2's exemption-migration pin with the resolvePhaseDirByToken extraction #3348 (an earlier commit on this branch, e4b80ad81) extracted cmdVerifySchemaDrift's inline phasesDir readdirSync/matchPhaseDirs block into the shared resolvePhaseDirByToken helper (also used by the new cmdVerifyContextDrift), and retargeted lint-phase-enumeration-drift.cjs's function-scoped exemption from cmdVerifySchemaDrift to resolvePhaseDirByToken accordingly — cmdVerifySchemaDrift no longer contains a line the guard's detectors match, so it needs no exemption. tests/phase-locator.test.cjs's E2 test still pinned the exemption to the old name (cmdVerifySchemaDrift), unaware of the migration. Update E2 to match the same "migrated call site's exemption must move, not duplicate" pattern the test already applies to cmdRoadmapAnalyze and cmdInitMilestoneOp just below it: drop cmdVerifySchemaDrift from the still-exempt list and add symmetric assertions that it no longer carries the exemption while resolvePhaseDirByToken now does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): fix two self-contradicting/nondeterministic tests in context-drift.test.cjs 'always exits 0 (query command contract)' included the no-phase-arg case, which contradicts the file's own earlier 'errors with usage message on missing phase arg' test (that case legitimately exits 1 via the Usage error) — drop it from the always-exits-0 cases. 'degrades to mtime comparison outside a git repo' and '...in a repo with no commits' relied on real wall-clock ordering between two back-to-back writeFileSync calls to prove CONTEXT.md is newer than RESEARCH.md; on a fast filesystem both can land in the same mtime tick, producing a tie that computeContextDrift's strict `<` correctly treats as not-stale, so stale_artifacts comes back empty. Make both tests deterministic via explicit fs.utimesSync instead of relying on timing (CONTRIBUTING.md: never assert elapsed wall-clock time). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3348): add context_drift_precheck:false to the plan:pre all-off fixture The "all plan:pre when-keys false" fixture explicitly disables every known workflow.* plan:pre toggle, but didn't yet know about the new workflow.context_drift_precheck key (defaults to true), so the new drift context-drift gate stayed active and broke the empty-activeHooks assertion. Emitted-Drift-Ack-Growth: plan-phase.md — adds the #3348 context-drift plan:pre pre-check section (new ## 4.6); this PR's own diff, not incidental drift. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#3348): backfill changeset PR number (pr:0 -> 4147) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
1a358ce0fd |
feat(#2761): bracket-tolerant read path — roadmap/validate/verify/state recognize bracket ids (epic #612 PR-2) (#2867)
* feat(#2761): gated heading-intro selection + one bracket identity grammar Foundation. Two owner-level changes plus a federated convention resolver; no reader consumes them yet. 1. GATED SELECTION, not an ungated widening. Widening every heading matcher requires the claim "no legacy ROADMAP contains a `[CODE.MM]` bracket followed by a digit", and that is false: `### [RFC.2119] 5:`, `### [v1.0] 2024:`, `### [ADR.612] 3:` and `### [ISO.8601] 2026:` are ordinary headings, and a widened reader claims each as a phase — moving phase_count and total_phases and adding W006 on projects that never opted in. No narrowing rescues it: the premise is about documents we do not control. `phaseHeadingPrefixSrcFor(baseline, convention, capturing?)` selects the pattern SOURCE at construction time. A project whose resolved `phase_id_convention` is not exactly 'bracket' compiles the same source string it compiled before. `baseline` is explicit because whether a site spells the any-bracket prefix or a bare `Phase\s+` is a fact about that site's history: handing the wider grammar to a bare site retro-grants tolerance it never had, in both directions — warnings appear, and a warning that fires today vanishes. Both bracket forms CAPTURE. `[GSD.999] Phase 07:` previously matched through the base alternative, which captures nothing, so a reader saw no bracket, fell back to the legacy token rule, and counted a labeled icebox heading while excluding the label-less one beside it — two derivations of one ROADMAP disagreeing. 2. ONE bracket identity grammar, one width rule. The milestone width is reconciled with the emit validator: pad2 output, so two digits or 3+ with no leading zero. Earlier spellings diverged in both directions — admitting `002`, which the validator rejects, and a bare `0` pad2 never produces — and the section recognizers accepted `[GSD.2]`, which SCOPED a milestone no phase heading could then resolve into, recreating the on-disk-count fallback this epic removes. An unpadded bracket is now uniformly malformed: it scopes nothing, bounds nothing, sections nothing. W005 on its directories is the surfacing signal. The milestone field is boundary-anchored, so a malformed run cannot match by its prefix (`GSD.002-01` read as sentinel `00`). Recognition stays case-insensitive because readers compile `/i`, but identity helpers match `[A-Z]`, so a captured id is folded first — otherwise `### [gsd.999] 07:` failed every sentinel test. The qualified key shares the width, the `(?=-|$)` boundary and the single-sub-phase shape of the directory token, because phaseTokenMatches returns unconditionally on a qualified hit: a key matching a directory isPhaseDirName rejects would be a final wrong answer. 3. resolvePhaseIdConvention federates workstream -> root exactly as config-loader does — including that root is a fallback only when a WORKSTREAM is active, so a project-scoped directory stands alone. loadConfig cannot serve this: it merges against CONFIG_DEFAULTS and drops keys it does not know, and this key is not among them. It governs the bracket-selection reads ONLY. PHASE_HEADING_PREFIX_SRC is left byte-identical: PR-1 shipped it, nothing consumes it, and it is superseded rather than redefined. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): roadmap.cts selects its heading grammar from the convention Six matchers build their intro through the gated selector, and cmdRoadmapAnalyze / cmdRoadmapGetPhase / getRoadmapPhaseWithFallback each resolve the convention ONCE per command and thread it down. Three sites take the any-bracket baseline (they already tolerated `[anything] Phase N`); three take label-only (they spelled a bare `Phase\s+`). Handing the wider grammar to a label-only site retro-grants tolerance it never had — and not only by adding matches: on a legacy repo an unchecked `- [ ] **[v1.0] Phase 05: Thing**` bullet would start SUPPRESSING the W006 that fires today. Sentinel handling under bracket ADDS a rule rather than replacing one: a bracketed heading is a sentinel when its bracket milestone is reserved (`### [GSD.999] 01:`) OR when its token is, so the engine-wide 0/999 backlog convention keeps applying to `### [GSD.02] 999:`. Replacing the token rule let a mid-migration ROADMAP — bracket headings plus a legacy backlog block, exactly the content this epic targets — add entries to the progress denominator. The captured id is folded before the identity test, so a lowercase `### [gsd.999] 07:` is excluded too. The DIRECTORY read is threaded too. `cmdRoadmapAnalyze` resolves the convention once and hands it to all four of its heading/checklist patterns, but the single `phaseTokenMatches` call that decides `disk_status`, `plan_count`, `summary_count`, `has_context` and `has_research` was left two-argument — so every canonical `{CODE}.{MM}-{PP}-slug` directory read as `no_directory` with zero counts, on the PR's own headline verb, while the SAME build resolved those same directories correctly in three other places on the same repo (W006/W007 via phaseTokenFromDir, `state json` via the milestone filter, and the W021 milestone-complete read through this very helper's three-argument form). It failed ONLY for the directory shape the convention exists to name: a mid-migration bracket repo carrying legacy `01-one` dirs resolved fine, which is why nothing caught it. Measured, bracket vs its flat-legacy twin: `[["01","no_directory",0,0],["02","no_directory",0,0]]` against `[["01","complete",1,1],["02","planned",1,0]]`. The oracle is the twin, computed in the same test run, plus exact literals — `grep disk_status tests/adr-612-*` was zero hits before this, so neither the fix nor a future regression had any gate at all. Disclosed: a ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted. Silent invisibility during the migration window is the deliberate trade against claiming phases on projects that never opted in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): validate.cts selects its grammar; gated directory recognition The W006/W007 feeders take the resolved convention as a threaded parameter. These sites carry the letter-tolerant `[\w][\w.-]*` capture, which makes them where an ungated widening does the most damage: `### [RFC.2119] 5:` enters roadmapPhases as a phantom and becomes a W007 "in ROADMAP.md but no directory on disk" on a project that never opted in. buildRoadmapPhaseVariants also surfaces the tokens borne ONLY by sentinel-bracket headings. Surfaced rather than filtered in place because roadmapPhases feeds both a membership check and a missing-directory warning, and only the latter should ignore an icebox item. That set is OCCURRENCE-AWARE, and the subtlety is load-bearing: roadmapPhases is a TOKEN set, so `[GSD.999] 01` and `[GSD.02] 01` collapse to one entry. Keying suppression on the token alone let an icebox heading silence a REAL phase that happens to share its number — a false negative strictly worse than the warning it removed. A token is suppressed only when no non-sentinel heading bears it. Directory recognition is added as gated FUNCTIONS beside the exported RegExp constants, which stay byte-identical: the `{CODE}.{MM}-` prefix is string-indistinguishable from the letter-prefixed-decimal family this repo documents as ambiguous, and folding a branch in changes those constants' answers on exactly that family. A RegExp constant has nowhere to attach a gate. The recognizer mirrors the emit grammar and delegates the token to the canonical owner, so recognizer and resolver agree on rejected input as well as accepted. Both functions throw on a non-string, matching the call pattern they replace. buildRoadmapPhaseVariants' CHECKLIST scan is capturing, like its heading twin and like the sibling checklist scan in roadmap.cts, and for the reason that one states: the bracket id has to ride along or the sentinel filter is blind to `- [ ] **[GSD.999] 01: Icebox**`. Left un-capturing, the scan called every checklist token REAL, and the occurrence-aware un-suppression loop then deleted the icebox token the HEADING scan had correctly marked sentinel — so `validate consistency` warned that a bracket ICEBOX phase had no directory, in the HOUSE ROADMAP shape where an icebox appears as both a bold bullet and a detail heading. `validate health` stayed silent on that same repo, so the two verbs disagreed — which is the disagreement `sentinelPhases` exists to close. Both directions are pinned, because the failure mode of a careless fix here is the opposite one: a real phase sharing a sentinel's token must still warn. It does, in all four shapes that attack it (sentinel heading + real bullet, lowercase sentinel, sentinel after the real heading, colon-less bullet). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): count bracket headings, and retire them, in both derivations Both `total_phases` derivations select their grammar from the resolved convention, in one commit — cmdStateSync already carries the comment that it mirrors buildStateFrontmatter "so both report consistent percents (#3242 Bug B)", so teaching one and not the other ships that divergence. The #1514 retirement filter widens WITH the counter it protects. The canonical gesture strikes the checklist BULLET and leaves the detail heading intact, so a bracket-form retirement went undetected and the phase stayed in the denominator forever. That is half a fix alone: the retired key is compared against phaseKeyFromDir, which called extractPhaseToken with no convention. Both halves land here. Under bracket the sentinel token rule composes as the full engine set {0, 999}, so this counter agrees with `roadmap analyze`, which has always excluded both — otherwise the two derivations report different numbers for one ROADMAP and the changeset's "excluded from every count" is false as written. The LEGACY path keeps its pre-existing 999-only rule: widening it there would move legacy totals, so the two stay split off the bracket path exactly as they are today. The sync-side assertion reads the PERCENT sync writes into the STATE.md body, not the frontmatter total_phases. Sync's own counter never reaches that field — the read derivation writes it — so asserting the frontmatter after a sync measures the read path twice and lets a mutation to the write-path guard survive. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#2761): verify.cts bracket-coherence W021 + selected milestone-complete read The shipped milestone-prefixed W021 gate keeps its ROOT-only config read, verbatim base semantics. Federating it silently moved a legacy convention's answer in BOTH directions on workstream repos — a W021 that fires at base vanishing, and one that is silent at base firing. resolvePhaseIdConvention governs the new bracket-selection reads only. B6, the milestone-complete check, keeps its ungated POSTURE (bug-557 pins it with an empty config) but selects its grammar from the convention. Inferring 'bracket' from the shape of a matched bracket ran a repo-failing check against a legacy ROADMAP that merely contained `### [RFC.2119] 5:`. Directory resolution widens with the heading read, so a bracket repo whose phases are on disk stays silent, and a bracket sentinel is not reported as unstarted. checkBracketCoherence is advisory and gated. Anchored to tokenizeHeadings so fenced examples cannot warn and heading level is structural. Its scope rules each close a way it silently did nothing or fired wrongly: only a genuine MILESTONE heading opens or closes a section (a `### Notes` used to reset scope and disable both sub-checks); a legacy `## v3.0` DOES close it; an M-NN or letter-suffixed phase heading raises missing-bracket and CONTINUES; a bare `#### 2026:` is not a phase; the full h2-h6 range is processed. Its section recognizer shares the one milestone width, so an unpadded `### [GSD.3] 05:` can no longer be a phase to the id grammar and a section to the section grammar at once, silently re-scoping every warning after it. validate consistency suppresses bracket sentinels in its missing-directory warning — the two verbs disagreed, health suppressing via notStartedPhases while consistency did not. The legacy reading is untouched, including its pre-existing wart that `### Phase 999:` still warns there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2761): scope the milestone by its bracket; select the disk-side filter Two roadmap-parser reads, both of which made a bracket project's totals track the disk instead of the ROADMAP. The ADR pins the bracket milestone heading as `## [GSD.02] Foundation` — a name, no version — but scoping matched STATE's `milestone: v2.0` STRING against a heading, so the canonical form matched nothing and total_phases fell back to the directory count. The rule was re-derived in THREE places: extractCurrentMilestone plus two `milestoneBounded` guards; fixing one left the others falling back regardless, so they are now one gated helper. It matches the CANONICAL padded spelling only — accepting `0*N` bounded a milestone whose phases were invisible, which un-suppressed a progress percent computed off an unscoped disk count. getMilestonePhaseFilter's heading scan becomes the 14th selected read. On a bracket ROADMAP it collected nothing, so the filter degraded to pass-all and buildStateFrontmatter counted every other milestone's directories — making the bracket convention strictly worse than the M-NN one it supersedes on the property that matters most: totals must track the ROADMAP, not the disk. The DIRECTORY side of that same filter is selected with it. Teaching only the heading scan was half a fix and a worse one: `milestonePhaseNums` became non-empty, so the pass-all degrade stopped firing, but no bracket directory could satisfy the three legacy dir checks (numericRe fails on `GSD.02-05-five`, the custom-id match captures the project code `GSD`, and stripProjectCodePrefix does not strip a dotted prefix). Every bracket directory was rejected, and completed_phases / total_plans / completed_plans / percent all collapsed to 0 while `state sync` went on writing a percent off the unfiltered disk — `state json` reporting 0% on the same repo, in the same second, that STATE.md's body called 67%. That is the #3242 Bug B divergence this PR exists to avoid, and total_phases could not show it: `Math.max(phaseDirs.length, roadmapPhaseCount)` floors it at the ROADMAP count no matter how many directories are rejected. The dir side matches on the milestone-QUALIFIED id, delegated to the owner's gated `phaseTokenMatches(dir, id, 'bracket')`, not on the bare token: READING-B puts the milestone in the bracket, so `GSD.01-01-old-one` and `GSD.02-01-one` share the token `01` and only the qualified key separates them. The qualified ids are kept in their own set — a hyphen in `milestonePhaseNums` would flip `roadmapUsesHyphenedIds` and silently move the LEGACY dir path on a bracket repo — and the branch is ADDITIVE: on a miss it falls through to the three legacy checks, so a bracket project carrying legacy-shaped directories reads unchanged. Both are resolved lazily and gated, so the legacy path pays neither a config read nor a second scan and cannot change answer. The scoping call is also GUARDED: resolvePhaseIdConvention reaches planningDir, which throws a plain Error for a GSD_PROJECT/GSD_WORKSTREAM segment carrying `/`, `\` or `..`. At base the only planningDir call in extractCurrentMilestone sits inside the STATE-read try, so the function returned normally on such an environment; an unguarded one here let that escape and broke the never-throws invariant that getRoadmapPhaseInternal and getMilestoneInfo three hundred lines below carry #2245 / ADR-227 notes about. Unreachable through the CLI — GSD_WORKSTREAM is rejected up front by the workstream-name policy and GSD_PROJECT throws identically at base — but reachable by any in-process embedder, which is precisely who that invariant is for. The filter's own resolve call was already inside its try and is unaffected. The milestone-qualified key is formed only for a token that is itself a bracket phase token. `${bracketId}-${token}` is a string SPLICE, so a mid-migration heading carrying an M-NN label — `### [GSD.02] Phase 02-01:` — spliced to `GSD.02-02-01`, which the qualified-key grammar reads as milestone 02 / phase 02: the `-01` truncated, both such headings collapsing to one key, and the heading claiming `GSD.02-02-two`, the directory it does NOT name, while rejecting `GSD.02-01-one`, the one it does. The guard drops those headings back to the unqualified legacy path, restoring the base ACCEPTANCE VECTOR exactly — pinned against the milestone-prefixed reading of the same ROADMAP, which is base-identical on this shape. Scoped precisely, because the fixture moves one number that the guard does not touch: `total_phases` on it reads 1 at base and 2 here. That is the bracket heading COUNT this PR exists to add, not the splice — measured identical with and without the guard, and identical to what the canonical `### [GSD.02] 01:` spelling does on the same fixture (both read 2 with zero directories on disk, where base reads 0). The claim is base-equivalent ACCEPTANCE, not a base-equivalent reading. One consequence is stated rather than fixed: a heading whose token carries a hyphen still puts that hyphen into milestonePhaseNums and so still flips `roadmapUsesHyphenedIds`. Base does the same for that spelling, so preserving it is what keeps the shape base-equivalent; excluding the token would have moved answers versus base on malformed input. The comment at the qualified-set declaration is corrected to claim only what is true — it keeps QUALIFIED IDS out of that flag's input, not hyphens in general. The oracles ship with it, and they are the five numbers, not the one: the parity gate now asserts total_phases, completed_phases, total_plans, completed_plans AND percent, on both derivations, on two fixture shapes (one milestone; two milestones with stale prior-milestone directories on disk). The oracle is the flat-legacy twin, built in the same test run and compared number for number, plus exact literals so a shared wrong answer cannot pass. The oracle SUBSTITUTION is itself pinned. The M-NN spelling of these shapes could not serve, because buildStateFrontmatter's #2445 de-dup key captures only a directory's leading integer and collapses `02-01-one` / `02-02-two` / `02-03-three` to one — measured [3,0,1,0,0] against the flat-legacy twin's [3,2,3,2,67], identically at base and before this fix, and structurally unreachable from the bracket key space. That reasoning is only sound while it stays true, so a characterization test holds the M-NN reading down on the two numbers that do not depend on which directory wins the mtime race. Widen the de-dup key and it fails, instead of quietly invalidating the changeset's disclosure. Also adds the call-site pin. The structural table pins transcription against the selector; it cannot see a call site whose BASELINE ARGUMENT is wrong. Flipping verify.cts's milestone-complete site to the wider baseline grants a fires-on-every-repo check tolerance it has never had, and every behavioural test still passed. The pin reads the shipped sources and asserts the mode at each of the 14 sites, count-exact. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2761): pin the bracket read surfaces in the parity gate This gate exists because #2043 fixed one bug across five hand-edited copies of a rule and #2232 was the residual that survived, because a later reader could not tell the copies were one rule. PR-2 adds two consumers, so they belong here. Surface 7 — the heading read and the directory read must agree about WHICH phase a `MM-<seg>` pair names, across the shared width corpus, and the bracket and legacy spellings of one heading must yield the same token. Surface 8 — the two bracket directory readers, in BOTH directions. Agreement on ACCEPTED input was already pinned; agreement on REJECTED input is where they actually diverged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2761): changeset Disclosures for the PR body (deliberate, not defects): - phase_id_convention is not a CONFIG_DEFAULTS key, so loadConfig drops it and cannot serve as the convention resolver however the file is federated. This PR ships its own workstream->root resolver; adding the key and its value enum is later-slice work. - Convention matching is strictly === 'bracket'. A misspelled value reads as not-configured and the project keeps legacy behaviour silently. - An UNPADDED bracket milestone (`[GSD.2]`) is malformed: it scopes nothing, bounds nothing, sections nothing, and is not a phase id. W005 on its directories is the surfacing signal. - WIDTH UNIFICATION MOVED FOUR MERGED PR-1 EXPORT ANSWERS on non-canonical inputs, none of which toDir can emit and none of which had a bracket caller at base: isSentinelPhaseId('GSD.0-01', 'bracket') true -> false isSentinelPhaseId('GSD.0999-01', 'bracket') true -> false getMilestoneFromPhaseId('GSD.2-01', 'bracket') 'v2.0' -> null getMilestoneFromPhaseId('GSD.002-01', 'bracket') 'v2.0' -> null The canonical pad2 sentinel spelling `[GSD.00]` still tests true. - FLAG TO MAINTAINER: docs/adr/612:132 reads "Sentinel behavior (0.x / 999.x -> milestone null) is preserved". After the unification that holds for the canonical `00` spelling only, not for a bare `[GSD.0]`. ADR wording is yours; flagging the tension rather than editing it. - The bracket sentinel rule COMPOSES with the legacy one — a bracketed heading is a sentinel when its bracket milestone OR its token is reserved. Under bracket the state-side token rule is the full {0, 999} set so both derivations agree; the LEGACY path keeps its pre-existing 999-only rule, unchanged. - validate consistency's legacy reading is untouched, including the pre-existing wart that `### Phase 999:` warns there while validate health suppresses it. - find-phase still cannot resolve a bracket phase directory. phase-locator.cts is outside this PR's module set. Sibling PR #2559's matchPhaseDirs calls phaseTokenMatches without a convention, so whichever slice lands second must thread it through. - Four of the five bracket readers scan raw ROADMAP content, so a bracket heading inside a fenced code block is read as a phase. Pre-existing for the legacy spelling; parity, not a new class. - roadmapPhaseLookupSources gained no bracket source: nothing emits a milestone-qualified query into it yet. - roadmap validate remains a separate, unfederated convention reader. Pre-existing and base-identical, but two verbs can disagree about the active convention on one project. - _diskScanCache keys on cwd while the values it caches are now convention-dependent. Not reproducible through the CLI; pre-existing for the workstream dimension, widened here. Stated as inconclusive. - A ROADMAP written in bracket form before config.json is switched reads as empty rather than mis-counted — the deliberate migration-window trade. - THE READ AND WRITE PERCENTS STILL DIVERGE ON A MULTI-MILESTONE REPO, and that divergence is MIRRORED under bracket rather than closed. buildStateFrontmatter applies the milestone filter; cmdStateSync does its own fs.readdirSync and never calls it, so on a repo carrying prior-milestone directories the read path reports the SCOPED percent and the sync body reports the WHOLE-DISK one. Measured on the true base build ( |
||
|
|
62b0d939b6 |
feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey (#4083)
* feat(#3274): make reviewer-lane timeout configurable via timeoutConfigKey Add an optional `timeoutConfigKey` field to the reviewer lane descriptor, resolved in `resolveLanePlan` at invocation time and falling back to the frozen `timeoutFloorMs` when unset or invalid, in the same spirit as the existing `promptBudgetKey`/`modelConfigKey` fields. All 12 shipped lanes declare `review.timeouts.<slug>` on both surfaces (the descriptor and their capability.json manifest), validated by capability-validator.cjs. For the antigravity lane, the native `agy --print-timeout` flag — previously a second hardcoded literal (`540s`) independent of the outer cap — is now derived from the same resolved outer timeout in `antigravityArgv`, preserving the existing 60-second buffer relationship (ADR-2782 D6: the outer bound is declared data, the inner one is handler-owned). The antigravity default timeoutFloorMs stays at 600s per the maintainer's disposition; users raise it through the new config key instead. * docs(#3274): document review.timeouts.* and extract resolveTimeoutMs helper Address code-review findings on the timeoutConfigKey change: extract the inline timeout-resolution logic into a named, exported, directly-tested resolveTimeoutMs helper (matching the file's existing configString/ normalizeHost convention); document the new review.timeouts.* federated config keys in docs/CONFIGURATION.md, docs/reference/capability-manifest.md, and docs/how-to/ship-a-reviewer-lane.md; add the changeset fragment. * fix(#3274): resolve native antigravity timeout in resolveLanePlan, not the runner gsd-test caught two design mistakes in the prior commits: 1. SpawnPlan.argv is documented and tested as fully resolved by resolveLanePlan (model/effort/output/prompt already folded in) — leaving the antigravity '{{nativeTimeout}}' marker unresolved until the runner's antigravityArgv violated that contract and broke tests that read plan.argv directly (tests/antigravity-reviewer.test.cjs, tests/review-default-reviewers-workflow.test.cjs). Fix: '{{nativeTimeout}}' is now a fifth ARGV_PLACEHOLDER member, resolved by resolveLanePlan itself via the new nativeTimeoutToken() helper, exactly like the other four. antigravityArgv reverts to its pre-#3274 four-argument form. Also missed updating capabilities/antigravity/capability.json's invoke.args to match the descriptor, which broke the manifest/descriptor parity test. 2. tests/reviewer-config-federation.test.cjs enforces a deliberate, narrow invariant (#3691 narrows #2797): qwen, cursor, and coderabbit — the three lanes with neither a model flag nor a host — may own no config key beyond their own prompt-budget key. Adding review.timeouts.<slug> to all 12 lanes violated it. Fix: those three keep timeoutConfigKey: null and own no review.timeouts.* key, matching their existing modelConfigKey: null. The other 9 lanes are unaffected. * chore(#3274): backfill changeset PR number (pr:0 -> 4083) --------- Co-authored-by: sim <sim@local> |
||
|
|
8487f0ed42 |
enhance(#3552): warn on additional protected branches beyond the resolved base branch (#3648)
* test(01-01): add failing protected-branch warning coverage - pin configured, absent, and malformed branch-list behavior - require opposite CLI and execute warning outcomes * feat(01-01): warn on configured protected branches - resolve the base branch union configured protected branch names - expose exact boolean CLI comparison output for workflow callers - keep execute-phase warning advisory and within its byte budget * test(01-01): add failing protected branch config coverage - cover valid list persistence and null unset - reject hostile shapes while preserving the prior value * feat(01-01): validate protected branch configuration - register git.protected_branches as a canonical config key - require a non-empty array of non-blank branch names * test(01-02): add failing ship protected-branch controls - Execute both workflow warning blocks with exact predicate arguments - Require true and false results to produce opposite warning outcomes - Preserve the none-strategy feature-branch offer contract * feat(01-02): warn at ship on protected branches - Reuse the typed protected-branch predicate in ship preflight - Keep raw base resolution for PR targeting and advisory branch creation - Prove execute and ship warning blocks with opposite-result controls * test(01-02): add failing protected-branch docs parity - Require the canonical schema key in both English config references - Pin the non-empty string-array type and absent default - Require synchronized multi-branch examples and advisory semantics * feat(01-02): publish protected branch configuration contract - Document the optional non-empty string-array field in both references - Explain resolved-base union and absent-field compatibility - Keep execute and ship warnings advisory under branching_strategy none * fix(01): CR-01 honor active workstream branch policy * fix(01): WR-01 assert protected config path selection * docs: add changeset fragment for #3648 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017CteVPJt4BkPmroMPGajYx * fix(#3648): resolve base_branch precedence inversion and round-1 findings Blocker 1/2: production config resolution was flat-first, so a project that migrated to git.base_branch but still carried a stale flat base_branch got the old value back. Add base_branch to normalizeLegacyKeys (mirrors the existing branching_strategy/sub_repos pattern: canonical nested wins) and route readEffectiveGitConfig's test seam through the same normalization so it can't silently diverge from production again. Adds a regression test with both keys set that fails without the fix. Blocker 3/4/5: restore the handle_branching case-selector prose and "none" contract sentence that #3389's tests anchor on, and revert the unrelated prose/comment compaction in the same step — both were drive-by edits outside #3552's scope. Also addresses review majors/minors: delete readConfigBaseBranch and readConfigProtectedBranches (dead in production, only self-tested); --is-protected now fails closed (reports protected) instead of silently answering false when the base branch can't be verified; trim configured protected-branch names; fix HOME-without-USERPROFILE vacuous isolation on Windows; correct the drift-ack's byte accounting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01S44stkuQbhD3jTCtKzte5N * test(#3648): add failing legacy-key hoist safety coverage Round-2 review found normalizeLegacyKeys block 5 records a normalization carrying the DISCARDED flat value on the canonical-wins branch. Probing that turned up a second, unreported defect in the same helper shape: blocks 1, 2 and 5 all spread result['git'] / result['planning'] with no object guard, so a config whose section key holds a string is spread into index keys — {"git":"main","base_branch":"release"} -> {"git":{"0":"m","1":"a","2":"i","3":"n","base_branch":"release"}} The resolved value is accidentally still correct, so nothing fails and no diagnostic fires. But normalizations.length > 0 sets configDirty, and config-loader then serializes that shape back into the user's config.json — a read that silently corrupts config. The deleted #3057 W3 suite covered {"git":"main","base_branch":"release"} explicitly; this is the input it would have caught. Covers both defects across blocks 1 and 5, with object/array/null negative controls that must stay green in both phases, and a fast-check property over arbitrary `git` values. * test(#3648): pin fail-closed handling of malformed protected_branches Replaces the test that pinned the fail-OPEN behaviour. The old assertion — ['develop', 42] yields isProtected === false for 'develop' — locked in the exact failure #3552 exists to close: config-set validation is bypassable by a direct edit of .planning/config.json, so a user who believes 'develop' is protected got a silent false and no warning. It was also inconsistent with the fail-CLOSED direction twelve lines away, where an unverified base reports protected and writes a diagnostic. A protection predicate must not have two opposite failure directions depending on which input is bad (#3648 review Blocker 3). New coverage: a bad element drops only itself, a non-array contributes no names, an empty list is well-formed rather than malformed, and --is-protected surfaces the rejection. Both negative controls — a clean list reports nothing rejected and writes no diagnostic — must stay green in either phase, so the reject channel cannot fire unconditionally. * fix(#3648): drop only invalid protected_branches and report them Partition git.protected_branches instead of discarding the whole list on one bad element, and carry the rejections out through ProtectedBranchStatus so --is-protected can name them on stderr. Valid names keep protecting; the user finds out the rest were ignored. A non-array value still contributes no names — a bare string is not a list of branch names — but is now reported rather than swallowed. An empty array stays silent: declaring no extra protected branches is a valid choice, not a misconfiguration. writeDiagnostic is hoisted out of the unverified-base branch since both arms now use it. * test(#3648): prove the predicate diagnostic survives both call sites The workflow bash stub now emits a stderr diagnostic the way the real command does, which is what makes a swallowed `2>/dev/null` visible to a test — previously the stub was silent on stderr, so discarding it changed no observable behaviour and the call sites could drop the explanation undetected. Adds the Minor 2 binding check as well: ship must expose the predicate result as IS_PROTECTED rather than only echoing a warning, asserted by running the extracted bash and reading the bound value, not by grepping the workflow source. Both tests carry opposite-outcome controls — an empty diagnostic must leave the text absent, and a false predicate must bind false. * fix(#3648): surface the predicate diagnostic and bind ship's result Drop `2>/dev/null` from the --is-protected call at both call sites. The fail-closed explanation and the new rejected-entry warning both go to stderr, so discarding it left the user with a bare "protected branch" warning on a branch that is not protected and no way to tell a real match from a degraded-git guess. `git branch --show-current` keeps its own redirect — that one is genuine noise. ship.md binds IS_PROTECTED and its prose now branches on the variable, so the following steps have evaluable state instead of having to infer it from warning text in tool output. execute-phase.md byte accounting refreshed: 92326 -> 92645, net growth 319 bytes (was 331 before the redirect came out). Baseline re-verified against the current rebase base by blob id; the ceiling check passes with 755 bytes of margin. * test(#3648): restore negative space for the readFile config seam The #3057 W3 suite was deleted with readConfigBaseBranch, but every arm it pinned survives verbatim in readEffectiveGitConfig's readFile branch — the JSON.parse catch, the non-object guard, the git-section object guard, .trim() and blank-string rejection — and the four surviving readFile injections were positive-path only. protected_branches was never driven through this seam at all. Restores nine cases against the seam, including protected_branches partitioning, plus a control proving loadConfig still wins when both seams are supplied. Records honestly what the suite pins. Mutating the built lib shows .trim() is KILLED, while the non-object guard and the blank-string rejection SURVIVE — both are unreachable through this entry point for the same reasons the deleted suite documented against its own equivalents: a JSON-parsed non-object carries no relevant own-property either way, and a blank value is rejected a second time downstream by the resolver's truthiness check. They stay as defence-in-depth and are labelled known-unkillable rather than left looking like coverage this suite does not provide. * test(#3648): distinguish detached HEAD from a missing branch argument `args[1] ?? ''` collapsed two different situations into one: a detached HEAD, where `git branch --show-current` legitimately prints nothing, and the flag being called with no argument at all. Both answered false, so the right outcome arrived by an unintentional path and a caller bug was indistinguishable from normal operation. Asserts the detached case stays silent and the missing-argument case reports, with a control that the two diagnostics differ. * fix(#3648): report a missing --is-protected branch argument Answer false either way, but say so when the flag arrives with no argument. A detached HEAD passes an explicit empty string and stays silent, since that is a normal state rather than a misconfiguration. * docs(#3648): state exact-name matching and per-entry rejection isProtected is exact string equality, so a git-flow project must enumerate every release/* and hotfix/* by name. #3552 only asked for an integration-branch field, so the implementation satisfies the letter of the issue while leaving its git-flow motivation partly unserved — say so where users will meet it rather than leaving them to discover it. Also documents the Blocker 3 behaviour change: an invalid entry is ignored with a warning naming it and the remaining names still apply. Both statements land in docs/CONFIGURATION.md and gsd-core/references/planning-config.md, and the config-field-docs parity test asserts each in both so the two cannot drift. * refactor(#3648): extract isValidProtectedBranches for cross-surface pinning The `git.protected_branches` check inside `cmdConfigSet` and the resolver's per-entry filter in `git-base-branch.cts` are deliberately different shapes — all-or-nothing on write, per-entry on read, so a hand-edited config.json cannot fail the guard open. Nothing structural keeps their two definitions of "usable branch name" in step. Lifting the write-side check into a named, exported predicate lets a property test ask both surfaces about the same value and assert they agree, which is the fast-check gap the round-2 review flagged. No behaviour change: the predicate is the same expression, called from the same place. * fix(#3648): stop --is-protected rewriting the config it is asking about `gsd_run query git.base-branch --is-protected` runs on every execute-phase and every ship. It resolved config through `loadConfig`, whose normalize-then-write path rewrites `.planning/config.json` whenever any legacy key normalizes — so a boolean question was silently editing the user's checked-in config. This PR had widened the trigger by adding a fifth normalization block (top-level `base_branch` -> `git.base_branch`), making it fire for exactly the projects the feature targets. `loadConfigResolved` gains `options.persist` (opt-OUT, default true): resolution is unchanged, only the two write-back side effects are suppressed. The predicate passes `persist: false`; the ~30 other callers are untouched, so a legacy config is still migrated by ordinary use. Asserted on BYTES rather than parsed shape, because the rewrite reorders keys and reflows whitespace even when the values are equivalent. Three tests, each with its own control: the end-to-end CLI leaves the file byte-identical while still answering `true` from the legacy key (proving the config WAS read); an ordinary persisting load of the same fixture DOES change the bytes (proving the fixture is live rather than inert); and `persist:false` vs default over one directory returns deep-equal config while differing on the write. Reverting the one-line `persist: false` fails the first of those and only that one. Also from the review: - `readEffectiveGitConfig`'s comment claimed the readFile branch routed "through the same precedence authority production uses". It does not, and cannot — it reproduces two of production's steps over a single file. The comment now names what the seam covers and what it does NOT (root/workstream deep merge, builtin and global defaults, federated merge), and the seam now applies production's flat-then-nested lookup so it stops disagreeing about a surviving flat key. - The missing-argument diagnostic promised "answering false", which the fail-closed guard on the same call can contradict by printing `true`. It now states what it did with the argument and leaves the answer to stdout. * test(#3648): re-pin block 5 on #3760's refusal contract #3767 landed on next while this PR was in review and fixed the non-object config-section defect properly: a present-but-non-object section now BLOCKS its own migration — value preserved, no Normalization pushed, refusal reported via `skipped[]` — rather than being rebuilt from a plain-object view. That supersedes this branch's round-2 `hoistLegacyKey`, which prevented the character-key spread but still dropped the section value silently, and which the round-3 review correctly called out as destruction in place of corruption. The rebase drops that commit and routes block 5 through the upstream helper. This file's tests asserted the superseded design, so they are rewritten to pin block 5 — `base_branch` -> `git.base_branch`, which did not exist when #3760's suite was written — against the contract that now governs it: ordinary hoist into an absent/null/object section, canonical-nested-wins, and refusal for each of string/number/boolean/array sections with the exact `skipped` entry. Two controls keep it from passing vacuously: the refusal must be scoped to block 5 (an unrelated block still normalizes in the same call), and a property over arbitrary `git` values asserts hoist and refusal are exhaustive AND mutually exclusive per key, that a refusal leaves both the section and the legacy key untouched, and that a hoist manufactures no index key the input did not carry. * docs(#3648): correct the Git Query and Config Loader module contracts CONTEXT.md's Git Query Module still described base-branch tier 1 as a direct `.planning/config.json` read. Since this PR it is the EFFECTIVE configuration resolved by the Config Loader — a materially different authority, carrying the root/workstream deep merge, flat-then-nested lookup and builtin/federated defaults. The `--is-protected` predicate, `git.protected_branches`, and the two invariants that distinguish the predicate from the plain query (fails closed on an unverified base; must not write) were undocumented entirely. The Config Loader entry now states that loading is not side-effect-free by default and documents `options.persist`. docs/INVENTORY.md's `git-base-branch.cjs` row carried the same stale ladder and no mention of the predicate. `node scripts/gen-inventory-manifest.cjs --write` was run and produced no diff: the manifest indexes roster NAMES, not row prose, so a description edit cannot move it. Also closes the global-defaults minor: `git.protected_branches` is inert in `~/.gsd/defaults.json`, but so is every other `git.*` key — no branch-policy key appears in `_globalBaseCfg` or `GLOBAL_DEFAULTS_RESOLUTION_KEYS`. That is section-wide and predates this PR, so the fix is to state the scope where users meet it rather than to quietly extend the resolution set for two new keys. * fix(#3648): close four defects found by the round-4 external review Two external reviewers (codex, antigravity/Gemini 3.1 Pro) were run adversarially against this branch. Four findings reproduced against source; each is fixed with a failing-first test and a control, and each fix was verified by reverting it and watching exactly the intended test fail. 1. `persist:false` was DROPPED by the workstream fallback (codex). Blocker 1 was only half closed. `loadConfigResolved` re-enters itself with a bare `{ workstream: null }` when a workstream has no config.json of its own, and that literal discarded every other option — so the recursive pass ran at the DEFAULT persistence and rewrote the ROOT config. Reproduced: with GSD_WORKSTREAM=alpha and a legacy flat `base_branch`, `--is-protected` rewrote `.planning/config.json` despite `persist:false`. Both recursions now forward `options` and override only `workstream`; the explicit override still wins the hasOwnProperty check, so spreading cannot let `workstreamContext` reintroduce a workstream. 2. Both workflow call sites failed OPEN, and aborted under `set -e` (both reviewers, independently). `IS_PROTECTED=$(gsd_run ...)` yields an empty string when the query fails, so `[ "$X" = true ]` was simply false: no warning, no trace — a silent hole in the guard whose only job is to warn. The bare assignment also aborted the step under `set -e`. Both sites now degrade VISIBLY: `|| IS_PROTECTED=""`, then an explicit empty-string arm that says the check did not run. Deliberately not fail-closed — claiming "protected" on no evidence would warn on every branch whenever gsd-tools is unavailable. 3. `isValidProtectedBranches` and the resolver disagreed on a sparse array (antigravity). `.every()` skips holes; the resolver's `for...of` yields `undefined` for them, so `["main", , "develop"]` was accepted by config-set and rejected by the resolver. The cross-surface property passed only because `fc.array` cannot generate a hole. The predicate now indexes, and the generator punches holes so that axis is actually falsifiable. JSON cannot express a hole, so this is unreachable in production — but two definitions of one predicate must not contradict each other. 4. A top-level `protected_branches` silently outranked `git.protected_branches` (antigravity). Routing the key through `get(key, {section, field})` gave it flat-then-nested precedence, which is back-compat for keys `normalizeLegacyKeys` migrates. `protected_branches` is new in #3552 and has no legacy form, so that invented an undocumented alias. It now resolves nested-only through a new `getNested`, in production and in the test seam. `base_branch` keeps flat-then-nested — it HAS a legacy spelling that #3760's refusal path can leave behind — and a control pins that distinction. Also narrows a CONTEXT.md claim this round introduced. The predicate fails closed only when a git query TIMED OUT or could not be spawned (#3057 B4's `verified`); a git command that runs and exits non-zero counts as a clean negative, so a cwd that is not a repository answers `false`, not `true`. Verified pre-existing on next @ |
||
|
|
03b7125293 |
enhance(#3909): a probe that could not run no longer asserts a verdict (#3944)
* test(#3909): failing-first suite for the fabricated probe fallbacks Binds the four fabrication sites found by executing the surfaces (ADR-3889 failure class (c)), each with a positive control so an over-firing fix goes red: - the blocking api-coverage.verify-pre gate certifying "no external-API integration" from a zero-byte phase scope - the assumption-delta query route scanning an unresolvable phase section as the empty string and reporting it as an examined negative - both capability fragments' probe fallbacks, which append a fabricated verdict rather than replacing, and fire on the legitimate exit-1 negative Verification runs on the remote runner. Refs #3909 * enhance(#3909): a probe that could not run no longer asserts a verdict ADR-3889 Phase 5. Four sites turned a failed or unexamined probe into a confident negative; each now reports what it could not establish. - check api-coverage.verify-pre: a phase with no plan body and no roadmap section ran detection over zero bytes and PASSED the blocking seal gate, certifying "no external-API integration" from input it never read. It now holds with scope_unavailable. The discriminator is bytes examined, never signals found, so a phase whose plans are real and simply carry no API vocabulary passes exactly as before. - query assumption-delta scan: an unresolvable phase section was scanned as the empty string and reported as an examined negative. It now returns {skipped, reason: phase_unresolved}, still at exit 0 — an ADR-2980 degraded result in the payload, leaving the gsd-tools exit projection to P8. - both capability fragments: `|| echo '{"detected":false}'` appended rather than replaced, and fired on the legitimate exit-1 negative, so a correct answer and an honest skip both arrived as two concatenated objects. They now keep the probe's own payload and manufacture only an explicit probe_unavailable skip when the probe produced nothing at all. Every registered outcome is more restrictive on a blocking gate, so this can turn a false green red and never a red green. Docs: FEATURES 156, CONFIGURATION (both keys), references/api-coverage.md seal-time outcome table, and a new how-to for the reason-code vocabulary. Verification runs on the remote runner. Closes #3909 * test(#3909): correct the stale unknown-phase assertion `unknown phase → detected:false, no throw (graceful)` scanned phase 999 against a two-phase roadmap and asserted `detected === false`. That pinned the fabrication as intended behavior: the phase does not exist, so the detector was handed the empty string and its "no core assumption changed" answer described nothing that was ever read. It now asserts the skipped-with-reason shape. The graceful-degradation contract the test was actually protecting — the query succeeds and does not throw on an unknown phase — is unchanged. Found by code review, not by the author. Refs #3909 * docs(#3909): author the FEATURES entry in its generator source `docs/FEATURES.md` is generated by `scripts/gen-features.cjs` from the per-feature fragments in `docs/features/`. The API-coverage entry was edited in the generated file, so the next regeneration silently dropped it. The text now lives in `docs/features/api-coverage-gate.md` and `docs/FEATURES.md` is regenerated from it, leaving the shipped file byte-identical and its content actually derivable. Caught by `lint:generated-sync`. Refs #3909 * test(#3909): bind the skip to "not found", and pin the discriminator The first verification run went red on one case, and the case was wrong rather than the code. `getRoadmapPhaseWithFallback` returns `null` for an unknown phase and for a missing ROADMAP.md, but for a section whose body is whitespace-only it returns the heading line alone — which is not empty. So a body-less section WAS found, and reporting `detected:false` over its heading is a real negative, not a fabrication. The test had assumed the resolver yielded `''` there. Correcting the test rather than the resolver keeps `skipped` bound to the distinction the issue asks for — found versus not found — and avoids diverging `assumption-delta scan` from `roadmap.get-phase`, which the fragment documents as sharing one resolver. Also adds the seeded property the test matrix had promised: for any plan body, the scope read back is whitespace-only exactly when the body was. That pins the gate's discriminator to bytes examined, so it cannot quietly become "no signals found", across unicode whitespace and CRLF. `docs/INVENTORY.md` picks up the reference doc's new seal-time outcome table — surfaced by the co-change gate, not by a lint failure. Refs #3909 * chore(#3909): backfill the changeset PR number Refs #3909 --------- Co-authored-by: sim <sim@local> |
||
|
|
941b62249e |
enhance(#3906): two terminators over one registry, with a versioned exit projection (#3924)
* feat(#3906): two terminators over one registry, with a versioned projection Adds terminateNow (write-then-terminate, for callers that cannot wait for the event loop) beside runMain (drain-then-exit), both projecting through one shared function so they cannot disagree - the parity the ADR makes mandatory. A failed write does not change the exit code: letting it propagate would fail a hook open, which is what the fail-closed branches exist to prevent. The projection is versioned. v1 reproduces today's integers, including keeping a payload-carried degraded result at exit 0 - ADR-2980 ratified that across 60 sites and declined normalizing it on measured blast radius. v2 applies the registry. --exit-contract=v2 or GSD_EXIT_CONTRACT=v2 selects it; an unrecognized version throws rather than silently defaulting. The registry is now emitted beside both copies of the exit module, so it resolves as a sibling in the built tree and in the committed scripts/ copy that must load on an unbuilt clone. * fix(#3906): actually restrict code 2 to terminateNow, and generate the registry's type The claim that terminateNow is the only place 2 can be produced was false: runMain's outcome arm applied no guard, so runMain(()=>'HOOK_DENY') set exitCode 2 through the drain path - and the parity matrix demonstrated it while calling it parity. runMain now refuses any outcome projecting to the hook-protocol code, gated on the code rather than the name so an alias cannot slip past, and the matrix asserts the restriction instead of contradicting it. The ambient type for the generated registry was hand-written with no gate against the generator's actual output - the declared-surface-diverges-from-runtime defect class this epic exists to close, reintroduced inside it. It is now a third generated artifact covered by the same --check. Also converts every test-body try/finally to t.after(). * test(#3906): derive the glossary fixture's dependencies instead of hand-listing them Adding a require to scripts/lib/cli-exit.cjs broke 31 tests in one suite that built its fixture from a hand-written dependency list, so the new sibling was absent and the copied script could not load. copyScriptWithDeps walks the require graph and exists for exactly this class - #3412 paid the same bill when one new require broke 82 tests across two suites. Migrating rather than adding another copyFileSync line keeps the class closed. The other nine suites referencing that path were triaged; none copies-and-spawns, so none needed migrating. * fix(#3906): enumerate the new shipped file, drop a vendor name from shipped data, and fix three test defects install: scripts/lib/exit-code-registry.cjs was missing from GSD_SCRIPTS_LIB_FILES, so it shipped to every install and orphaned on uninstall. The registry gave HOOK_DENY a meaning naming one harness, and that string ships into every runtime's tree - a guard correctly caught it leaking into the hermes and qwen installs. The registry is runtime-neutral infrastructure; the vendor name belongs in the ADR, not in shipped data. Two more fixture harnesses built their trees from hand-listed dependencies and broke on the new require; both migrated to the derived helper, and all 23 copy-and-spawn candidates were enumerated so the class is closed rather than patched. One generator test used a fixture code that collided with a real allocation, so the generator correctly reported a duplicate where the test expected drift. The large-payload test embedded a 256KB literal in the child's argv, exceeding Linux's 128KiB MAX_ARG_STRLEN so the child never started - it now builds the payload inside the child. * chore(#3906): backfill changeset pr number * docs(#3906): document the exit-code contract selector P2 is the first phase of this epic with a user-invocable surface, so the flag and env var owe a reference entry. Records what actually differs between v1 and v2 today (one outcome), that an unrecognized value is rejected rather than silently defaulted, and the fail-safe property that makes switching safe. --------- Co-authored-by: sim <sim@local> |
||
|
|
382bf7c423 |
fix(#3706): deliver the resolved reasoning effort to OpenCode subagents (#3867)
* test(#3706): failing-first coverage for OpenCode variant emission and frontmatter escaping * fix(#3706): emit the resolved reasoning effort as OpenCode's variant key `query resolve-execution` resolved an effort level for every agent, but the OpenCode bake wrote only `model:` — the effort never reached the generated agent, so subagents ran at whatever the runtime defaulted the model to. This is the effort-side twin of the model-side defect fixed in #3705. The key is written only when an `effort` block is actually configured. `resolveInstallTimeEffort` always returns a level (the catalog default is `high`), so gating on its return value would stamp `variant: high` into every existing OpenCode install — and OpenCode resolves a variant name against a `variants` map in the user's `opencode.jsonc`, so a value nobody declared is not a safe default. Gating on `readGsdEffectiveEffortConfig` keeps installs that never asked for effort routing byte-identical. Kilo does not receive the key: `EFFORT_ARGV` declares surfaces for claude, opencode and codex and has no kilo entry. This is deliberately asymmetric with the model side, where #2794 J8 requires the two runtimes to resolve alike. Both frontmatter sinks now route through `frontmatterScalar`, which quotes and escapes any value that is not a plain scalar. The raw interpolation predates this change, but it was already shown by execution during the #3705 security review to let a config value containing a newline inject additional top-level keys (`tools:`, `permission:`) into a generated agent file. This change adds a second write to that sink, so it is closed here rather than doubled. * fix(#3706): quote frontmatter values YAML would not read back verbatim Self-review of the predicate added in the previous commit. Treating /^[A-Za-z0-9._:/@+-]+$/ as 'safe to emit bare' answers the wrong question: a value can match it and still not round-trip. - A leading '@' is a YAML *reserved* indicator and may not open a plain scalar at all, so a scoped ID like '@org/model' emitted bare is a parse error, not an ambiguity — the whole agent file becomes unreadable. - 'no' / 'y' / 'off' / 'null' resolve to booleans and null, so a variant with one of those names would match no entry in the user's variants map. - '12:30' resolves to 750 under YAML 1.1 sexagesimal, and ':' is legal mid-identifier here, so the form is reachable rather than contrived. Real model IDs pass every clause and stay bare, so already-generated files remain byte-identical. * fix(#3706): route variant through the declared effort seam and cover the live path Addresses six findings from the isolated review, all confirmed by execution. The tests were the serious one: they required `../bin/install.js` while the fix landed in src/, which compiles to gsd-core/bin/lib/. They exercised a different copy of the converter than the one the bake actually uses, so the whole suite was green-by-construction against unchanged code and the remote run failed all 13. Every case now runs against BOTH copies from one table, which doubles as the parity assertion the generative-fix note in runtime-artifact-conversion.cts asks for, and bin/install.js carries the mirrored change. Emission no longer hand-rolls the value. It goes through `renderEffortArgv`, the declared OpenCode effort seam (EFFORT_ARGV.opencode: its own supported set and clamp). That is what rejects a level that is not a wire value — above all `inherit`, which per #3533 (10d) means "omit the key and follow the host default" and was previously written literally, naming a variant that cannot resolve. Reachable two ways, both now pinned: an agent_overrides entry and a routing_tier_defaults entry. A bare effort.default does NOT reach a tiered agent (the #3531 tier ladder answers first), so a test written against `default` alone asserts nothing — that is pinned too. The plain-scalar decision moved into frontmatter.cts beside `scalarNeedsDoubleQuoting` rather than sitting next to it as a second, weaker predicate. `agentScalarNeedsDoubleQuoting` is a documented superset: it adds a trailing `:` (read as a nested mapping key, which fails the whole frontmatter), boolean/null words, and numeric-looking values including YAML 1.1 sexagesimal. Docs now state the cascade plainly: the gate is on effort being configured at all, not on the individual agent being named, so every generated OpenCode agent gets a variant line once any effort block exists. * test(#3706): assert the two frontmatterScalar copies cannot diverge A hand-picked adversarial corpus plus a fast-check property over YAML-significant strings, both run against bin/install.js and the live src copy. Verified the property can actually fail: mutating one copy's quoting rule is killed well inside the run budget. * fix(#3706): close the review findings — predicate, seam, and dead mirror Third review round; every item below was confirmed by execution. The scalar predicate was wrong in two families, both found by a round-trip property test rather than by reading. Basing it on scalarNeedsDoubleQuoting dropped the "first character must be alphanumeric" clause, so `~`, `.inf`, `.nan`, `+1`, `-0` and `.5` went out bare and came back as null/floats/ints; and that base predicate only inspects the FIRST character, so an embedded `: ` (a nested mapping, i.e. a parse error) or ` #` (a comment, i.e. silent truncation) also passed. Dates round out the set: `2026-08-25` opens alphanumeric, survives every other clause, and YAML resolves it to a Date. The property now asserts the contract directly over generated values instead of trusting an enumerated character list. The bin/install.js mirror is gone. Its premise was false — install.js already requires bin/lib at :65 — and it was unreachable besides: install.js's convertClaudeToOpencodeFrontmatter has no `isAgent: true` call site, because its agents path resolves converters from the compiled module. It was a third copy of the YAML rules serving a test rather than a caller, so the file is back to origin/next and the tests target the live copy only. Effort clamping moved to `clampEffortForHost`, which renderEffortArgv now delegates to. The layout was calling renderEffortArgv with a hardcoded 'argv' to borrow its clamp, which read as if the frontmatter key were gated on the invocation-time axis. It is not: claude declares effortSurface "argv" and independently bakes an effort: key. One capability table, one clamp, two channels that no longer pretend to be each other. Also corrects an earlier claim of mine: adding EFFORT_RENDERING.opencode would NOT have made `effort sync` write the wrong key, because it guards on the runtime name before it ever renders. The seam choice stands on other grounds. `effort sync` still skips OpenCode, but its stated reason claimed OpenCode "does not use effort: frontmatter", which this change makes false — so the message now says what is actually true. * docs(#3706): restate the changeset around the round-trip contract * fix(#3706): restore the changeset fragment belonging to #3809 An earlier commit in this branch picked the first file in .changeset/ by glob order instead of the fragment created for this issue, and overwrote agile-geese-squeak.md (PR 3815 / #3809) with this change's body. Restored verbatim from origin/next; this change's text now lives in its own patient-cranes-parade.md, where it was created. * feat(#3706): maintain the OpenCode variant key from effort sync Install bakes the resolved effort into OpenCode agent frontmatter as `variant:`, so `effort sync` has to maintain it or a config change only takes effect on reinstall — and its skip message claimed OpenCode does not use frontmatter effort at all, which this issue made false. cmdEffortSyncOpencode mirrors the codex branch: resolve per agent, clamp through the declared OpenCode capability, then write, strip, or skip. A null target means the key must not exist, which covers both "no effort configured" and "resolved to inherit or to an unsupported level" — the same states under which install writes nothing, so sync and install agree by construction. The frontmatter line-editors are key-parameterised rather than copied: setEffortFrontmatter / removeEffortFrontmatter are now thin wrappers over the same internals the variant path uses, and a test pins that the claude `effort:` behavior did not move. The child-process test harness fixes both HOME and USERPROFILE, so the hermetic-config assertions cannot pass vacuously on Windows. * fix(#3706): scope the frontmatter line editors to the matched block Found by the security review of the sync path, reported as correctness rather than vulnerability, and reproduced against pre-fix code before being fixed. Both editors matched the frontmatter with a regex that can match a block after a preamble, then derived the EOL and the opening-fence length from the START OF THE FILE. On a CRLF document with a preamble those disagree, the offsets shift by one byte, and the reassembled document comes back with a mangled fence (`---\rname: x`). Both now take the EOL from the matched block. `setFrontmatterKeyLine` additionally did a whole-file `/m` replace when the key already existed, gated only on the key being present in the frontmatter body — so a preamble line starting with the same key was rewritten instead of the frontmatter one. It now replaces inside the frontmatter span only, which is the hazard `removeFrontmatterKeyLine` already documented and guarded against. Neither is reachable from an install-written `gsd-*.md` (those begin at byte 0 with `---`), and both predate this change — but the editors are in this diff because #3706 key-parameterised them, so they are fixed here rather than left for the next caller to trip over. Three regression tests, each confirmed to fail against the pre-fix build. * fix(#3706): treat a present-but-empty key as present, and pin the real seam Fourth review round. The MAJOR one: both sync branches read the current value with `(.+?)`, which needs at least one character, so a key present with an EMPTY value read as "key absent". When the target was also null the code concluded "already correct" and skipped — leaving the key in the file, where it reads back as YAML `null`: exactly the unresolvable-variant state this change exists to prevent. Whitespace decided whether it fired, since `variant: ` matched and `variant:` did not. Presence and value are now separate questions at both the opencode and the claude branch. The OpenCode writer now follows the codex branch rather than the claude one: tmp file plus retryRenameSync with orphan cleanup, and a write failure skips that agent and is reported instead of aborting the sweep. Same granularity, same transient-Windows-lock exposure, so the hardened sibling was the right precedent. Also: the generic line-editors escape their interpolated key, the JSDoc stranded by the clampEffortForHost extraction is back on renderEffortArgv, and a cast that declared a nullable function as non-nullable is corrected. Tests close the gaps the review listed — empty value (both spellings), CRLF round-trip through write and strip, the symlink guard, a body line starting `variant:`, a file with no frontmatter, and the YAML classes that actually broke the predicate. The new layout-seam test drives the real stage() path and was verified to FAIL when `variant` is removed from the converter call; a seam test that survives cutting the seam is worse than none. * fix(#3706): clear the round-five review findings No blockers or majors this round; the repo's review gate is zero-tolerance, so the minors are cleared too. A duplicated key was only half-stripped: the strip regex had no `g` flag, so a frontmatter carrying the key twice lost one occurrence, reported success, and left the "a null target means the key must not exist" invariant false on disk — converging only on a second run. Such a document is already invalid YAML, so this is robustness rather than a live corruption path, but a successful sync has to leave the invariant true. A run in which every write failed still summarised as `ok`, so a caller could not tell "nothing to do" from "everything failed". The OpenCode branch now reports `failed` when any write failed. The write-failure path was also the newest code in the change with no coverage at all; it now has a test that injects the failure by monkeypatching the write, per CLAUDE.md §4, rather than by chmod — mode bits do not bite under root in CI. `CodexEffortSyncWriteFailure` is renamed `EffortSyncWriteFailure` now that two branches share it. Removed a guard on the claude concrete path that was provably unreachable — no member of EFFORT_SET renders null there, so it read as protection that did not exist. The claude inherit path's presence check is load-bearing and untouched. Three stale statements corrected: the OpenCode result shape matches codex's, not claude's, now that it emits write_failures; the `thread()` test helper now calls `clampEffortForHost` so it genuinely mirrors the layout instead of merely claiming to; and a test helper restored `USERPROFILE` by assignment, writing the literal string "undefined" into the environment on POSIX — it deletes now. * fix(#3706): converge the set path, degrade on unreadable files, preserve mode Rounds five and six of review. No blockers or majors; the review gate is zero-tolerance, so the minors are cleared too. `setFrontmatterKeyLine` was the mirror of a defect already fixed in its sibling: `remove` was made global, `set` was not, so on a frontmatter carrying the key twice it rewrote the first and left a stale second. Last-wins YAML readers honour the stale value while the sync's own first-occurrence read reports "in sync" — permanently non-converging. It now collapses to exactly one occurrence, in the position of the first, so ordinary single-occurrence documents stay byte-identical (verified across seven shapes before and after). An unreadable agent file used to throw and abort the entire sweep, while a failed WRITE in the same loop degraded into a report. The OpenCode branch now reports read failures alongside write failures; the claude branch degrades to a skip without a new result field, because its shape is long-standing and widely consumed and one bad file aborting the sweep is the actual defect. The tmp+rename publish dropped the original file's mode — a plain writeFileSync preserves it, a rename does not — so a 0600 agent came back 0644. Both the OpenCode and the codex branch now carry the original's permission bits across the publish, masked with 0o7777: the raw stat mode includes the file-type bits, and POSIX leaves those unspecified for chmod. Linux is the only OS the remote matrix runs, so relying on Darwin's tolerance would have been untestable here. Also documents the `from` contract on EffortSyncChange (null means the key was absent, '' means present with an empty value — a distinction earlier rounds introduced and then collapsed in the output), adds OpenCode to the docs paragraph enumerating where the key is omitted under inherit, and records in a comment that the 'failed' summary reaches only raw mode and does not change the exit code, which is a CLI-contract change affecting all three branches and is deliberately not made here. * fix(#3706): guard the codex read, close the tmp permission window, rename the failure type Round seven, plus one thing I found myself. `cmdEffortSyncCodex` still had an unguarded `fs.readFileSync` — a read fault on one agent exited 1 and aborted the whole sweep. The claude and opencode branches were both guarded earlier this round and codex was missed, with the unguarded read sitting ten lines above the chmod block the previous commit did edit. It now reports read failures the way the OpenCode branch does, and a read failure flips its summary to `failed` — which write failures did not do there either, so both are corrected for consistency. The tmp file was created at the default mode and only tightened afterwards, so a 0600 agent's contents sat in a 0644 file for the length of the publish. I measured the window rather than assuming it, then closed it by passing the mode at creation. The chmod after the write is deliberately RETAINED and commented: the `mode` option only applies when the file is actually created, so a leftover tmp from an earlier crashed run would be truncated and reused at its old mode, and the chmod is what corrects that. `EffortSyncWriteFailure` is renamed `EffortSyncFileFailure` — it was typing a `read_failures` array, the same naming-lie the `Codex…` prefix had last round. Also pins the codex mode preservation with a test. It only writes on a path that genuinely rewrites the file, so the fixture is an Anthropic-flavoured model pin the sync strips, and the test asserts the content changed before checking the mode — otherwise it would pass on a sync that did nothing. * fix(#3706): guard the claude writes and share one escaping rule The security sign-off caught a comment of mine that was factually wrong: the new claude read guard said the failure is folded in "like the write path in this same loop does", and there was no write guard in that loop. Rather than correct the sentence, both claude write sites are now guarded the way the read is — a failed file is skipped, the sweep continues, and the raw summary token flips to `failed`. The JSON shape stays frozen deliberately, because it is long-standing and widely consumed; the token is the channel that can carry the signal without a compatibility risk, which is the reviewer's own suggestion. That makes all three branches consistent: reads and writes guarded everywhere, per-file failures degrade instead of aborting, and every branch reports `failed` rather than `ok` when something did not sync. `setFrontmatterKeyLine` interpolated its value raw while the install-side writer quoted through the shared helpers — two writers of the same frontmatter key disagreeing on escaping, the divergence class this repo requires closed. They now share one rule. Verified no churn: all six effort levels are plain scalars and emit byte-identically, with claude's documented minimal-to-low clamp the only difference in the table, exactly as before. * fix(#3706): publish claude agent writes atomically too Both reviewers found this independently, and it is data loss rather than a reporting gap. The claude branch wrote in place, so `fs.writeFileSync`'s O_TRUNC meant a post-open fault left the agent file truncated or half-written: an injected ENOSPC produced an empty file, and under `ulimit -f` a 60000-byte agent came back as 512 bytes of wrong content. The guard added earlier this round then counted that destroyed file as `skipped`, which in JSON mode is indistinguishable from "already in sync" — so a caller would have read the sweep as clean while an agent on disk was corrupt. It now publishes the way the codex and opencode branches already do: write to a tmp file created at the original's masked mode, chmod, then retryRenameSync, with the tmp unlinked and the agent skipped on any failure. The corrupting case is gone rather than merely reported, which matters because this branch deliberately takes no new result key. I had claimed all three branches were consistent after the previous commit. That was true for degradation and reporting and not for atomicity; the reviewer caught the overclaim. It is true now. Also sorts the claude file list, which the other two branches already did — readdir order is platform-dependent, so leaving it unsorted made the reported `changes` ordering differ across machines for identical inputs. * chore(#3706): backfill the changeset PR number pr:0 placeholder replaced with the real PR now that gh api returned it. * test(#3706): kill the frontmatter mutants this change introduced CI's Stryker frontmatter shard scored 60.58 against a break floor of 62. The cause is documented in the lane's own config, from #1882: this PR added a multi-clause predicate to frontmatter.cts and exported the escaper, but the tests constraining them live in tests/runtime-converters.test.cjs, which that shard does not run — so every mutant in the new code was uncovered there even though the behaviour is tested elsewhere. The fix is assertions that kill real mutants, per the repo's own instruction, not a lowered floor and not a Stryker disable: scripts/mutation-matrix.cjs is untouched. Each clause of agentScalarNeedsDoubleQuoting now has a true case AND a near-miss that must answer the opposite way, so flipping the clause fails a specific named test — alnum-first against `a-b`, trailing `:` against `foo:bar`, embedded `: ` against `a:b`, embedded ` #` against `a#b`, the word list against `yes1`/`nullish`, the numeric forms against `1a`/`0xzz`, the timestamp against `2026-08-25x`, plus the case-insensitive spellings that pin the `i` flag. escapeDoubleQuoted is pinned on exact output, including a case constructed so that escaping in the wrong ORDER yields a different string. Two of my expectations were wrong and are asserted as the code actually behaves: `12:99` is NOT quoted, because the sexagesimal alternative never range-checks minutes and so does not match — which is right, since YAML would not read it as sexagesimal either; and `20260825` is quoted by the numeric clause rather than the timestamp one, being a bare integer. * chore(#3706): ratchet the frontmatter mutation floor to 65 The lane measured 66.67 on PR 3867 after the mutant-killing unit tests landed — above its pre-change 63.35 baseline, not merely recovered. Step 3 of this file's own HOW TO UPDATE procedure says to set minScore = floor(measured) - 1 in the same diff, so 62 becomes 65 and the improvement is locked in rather than left free to slide back. The ledger of measured scores now records the new measurement, why the shard broke in the first place (logic added to frontmatter.cts whose only tests lived in a file this lane does not run — the same trap the #1882 note describes), and one discrepancy: step 3 also says to update "the matching RATCHET_BASELINE entry", but no such declaration exists in this file. The name appears only in that comment, so minScore and the ledger are all there is to update. * fix(#3706): update RATCHET_BASELINE alongside the raised floor The ratchet test caught the previous commit: it raised COVERED['frontmatter'] .minScore to 65 without updating the baseline that mirrors it, which is exactly the mismatch that guard exists to make visible in review. I had claimed RATCHET_BASELINE did not exist. It does — in tests/mutation-matrix-ratchet.test.cjs, not in scripts/mutation-matrix.cjs, which is the only file I searched before concluding it was a stale reference. The ledger comment is corrected to say where it lives and to record that the guard caught the error rather than leaving my wrong claim on the record. * docs(#3706): put the mutation ledger entries back under their own dates The 2026-08-25 measurement was spliced into the middle of the 2026-06-14 list, so adr-parser, config-schema, active-workstream-store and core-utils ended up sitting under the wrong heading and misattributing their measurement dates. That ledger is what a future change reads to calibrate a floor, so a wrong date there is not cosmetic. Each measurement is now under the date it was taken. Also drops the first-person account of my own mistake from the entry — the factual half (where RATCHET_BASELINE lives, and that it is updated in the same diff) is what a reader needs; the confession is not. --------- Co-authored-by: sim <sim@local> |
||
|
|
aaf47c5fc2 |
fix(#3691): let every reviewer lane take a prompt cap, and make the documented global resolve (#3832)
* test(#3691): failing-first coverage for the reviewer prompt budget No prompt cap can reach any CLI reviewer lane, by any configuration. Two independent defects compound: all nine `transport: spawn` lanes declare `promptBudgetKey: null`, so `budgetFor` returns on its first line; and the documented global `review.max_prompt_tokens` is advertised in the schema manifest but declared nowhere, so the resolver never materializes it and `budgetFor`'s fallback is dead code. Adds to tests/reviewer-config-federation.test.cjs, which already owns the per-reviewer budget config-set/config-get idiom: - a CLI lane inherits the global cap (RED: reports null) - an http lane with the -1 sentinel inherits the global cap (RED: reports null) - the resolved review surface carries max_prompt_tokens at all (RED: absent) - per-lane overrides the global on a CLI lane - the sentinel boundary: -1 inherits, 0 means do-not-trim and must NOT read as unset, 1 is the smallest real budget — the regression budgetFor's own comment warns about - anti-tightening pins that must stay green: an empty config leaves every lane null, the three existing budgeted lanes are unchanged, and config-set still rejects a per-reviewer key naming something that is not a declared lane - a fast-check property over the resolution contract itself, with -1, 0 and non-finite inputs generated explicitly rather than left to chance Every row was reproduced by hand against the real CLI before being written, so the RED/GREEN split is observed rather than predicted. Refs #3691 * fix(#3691): let every reviewer lane take a prompt cap, and make the global resolve No prompt cap could reach any CLI reviewer lane, by any configuration. Two independent defects compounded. The nine spawn-transport lanes — claude, coderabbit, antigravity, cursor, gemini, codex, kimi-code, opencode, qwen — declared `promptBudgetKey: null`, so `budgetFor` returned on its first line and `review-lane plan` reported `promptBudget: null` no matter what was configured. Each now declares `review.max_prompt_tokens_per_reviewer.<slug>` with the same `-1`-is-unset sentinel the three local-server lanes already use. Separately, the central `review.max_prompt_tokens` was listed in the schema manifest's validKeys and documented as a supported setting, but declared nowhere — the resolved surface is built from capability declarations plus the defaults manifest, and neither carried it. `configGet` returned undefined and `budgetFor`'s documented fallback was dead code. It is now declared with a `null` default, exactly as docs/CONFIGURATION.md already specified, so the default behavior is unchanged: nothing configured means nothing trims. Two things the diagnosis had not predicted, found and fixed while implementing: - `REVIEWER_LANES` in src/review-lane-descriptor.cts is a second, hardcoded registration site that `mergeReviewerLanes` prefers over the capability registry on a slug collision. Editing only the capability files left every CLI lane still null. Both sites now agree. - The generated `gsd-core/bin/lib/capability-registry.cjs` was stale and masked the capability edits; regenerated with `npm run gen:capability-registry` rather than hand-edited. docs/CONFIGURATION.md said "Only lanes that declare a budget key accept one — today ollama, lm_studio and llama_cpp". That is false as of this change and is corrected rather than left to rot. The trim-versus-refuse question the issue raises is deliberately not taken up here: the refusal path already exists for the case that matters — a reviewer whose minimum set exceeds its budget is skipped rather than sent a misleading prompt — and trimming above that floor is the documented, shipped design of the feature. Changing it would alter behavior for the three lanes that already work, which is not what the issue asks for. Fixes #3691 * fix(#3691): document the new global and narrow an invariant this change obsoleted The full suite surfaced two consequences of giving every CLI lane a budget key. `review.max_prompt_tokens` entered CONFIG_DEFAULTS without a matching entry in the planning-config reference, which config-field-docs guards. Documented, including the sentinel semantics a reader needs: a per-lane value overrides the global, `-1` means unset and inherits it, and `0` means "do not trim that lane" and is not unset. The #2797 federation guard asserted that "a lane with no model flag and no host owns no config keys". That held only because budget keys existed solely on the three local-server lanes, all of which have hosts. A lane can now legitimately own a config key for a third reason, so qwen tripped it. The assertion is narrowed rather than weakened: such a lane must still own no model key and no host key, and may own at most its own `review.max_prompt_tokens_per_reviewer.<slug>` — never another lane's. That is strictly more specific in the dimensions that still matter. Proven to still bite: hypothetically giving qwen a `review.models.qwen` key fails it with `model/host: review.models.qwen`. The name and comment cite #3691 for why the premise changed, so a reader sees a deliberate narrowing, not erosion. Checked the sibling assertions in that describe block; the other three do not rest on the obsolete premise and are untouched. Refs #3691 * fix(#3685): port the write-flag content-change contract to its three sibling sites #3685 fixed `phase complete`'s `roadmap_updated` / `state_updated`, which reported `fs.existsSync(path)` rather than whether the transaction wrote anything. Three sibling sites carried the identical defect and are ported here. - `cmdPhaseRemove` reported `roadmap_updated: true`, hardcoded. `updateRoadmapAfterPhaseRemoval` now returns whether the content changed and the flag reports it. #2640/#2974 already fixed `state_updated` at this same call site and left this one behind, so the correct shape was adjacent. - `cmdMilestoneComplete` reported `state_updated: fs.existsSync(statePath)` — byte-identical to #3685's bug in a different command. - `cmdMilestoneComplete` reported `milestones_updated: true`, hardcoded, never consulting the MILESTONES.md write. `gsd-core/workflows/remove-phase.md:100` extracts `roadmap_updated` for display and never branches on it, so the flip from always-true to content-based changes no workflow behavior. Verified by reading the step, not assumed. One trap found while implementing: the obvious in-memory `finalContent !== originalStateContent` comparison — copying `cmdPhaseComplete`'s shipped shape verbatim — gives a FALSE POSITIVE for milestone completion. `platformWriteSync` normalizes Markdown at write time, and the milestone-closure transform regenerates `## Current Position` fresh on every call, so its pre-normalize output always differs from the already-normalized file on disk even when the persisted bytes are identical. The comparison is therefore made against the post-write on-disk content. `cmdPhaseComplete`'s own comparisons are left untouched — their repeat-no-op tests pass, so they are not exposed to this artifact. `milestones_updated` has no reachable no-op: the MILESTONES.md write unconditionally appends an entry every call. Only the true direction is pinned, documented inline rather than faked with a passing test. Refs #3685 * fix(#3685): compare write-flag content through the writer's own normalizer An independent reviewer disproved a claim made while porting #3685's contract to its sibling sites: that `cmdPhaseComplete`'s comparisons were not exposed to the Markdown-normalization artifact already diagnosed in `cmdMilestoneComplete`. `platformWriteSync` normalizes on write — CRLF stripped, blank-line runs collapsed, a blank line inserted after a heading, a single trailing newline enforced. Every flag that compares the PRE-normalization in-memory string against the on-disk pre-image can therefore report a change when the persisted bytes are identical. `cmdMilestoneComplete` had been worked around by re-reading the file after the write; the other sites compared raw strings. All of them now go through one exported seam, `contentChangedAfterNormalize(filePath, before, after)`, which normalizes both sides exactly as the writer does. That removes the extra disk read the milestone workaround needed, and makes the sites agree by construction rather than by four independent implementations of one rule — the divergence the repo names as an anti-pattern. Reachability, stated precisely rather than uniformly: the seam is load-bearing at `cmdPhaseComplete`'s `roadmapUpdated`, `requirementsUpdated` and `stateUpdated`, where section-rewrite logic genuinely regenerates content into a different-but-normalization-equivalent shape. At `updateRoadmapAfterPhaseRemoval` it is defense-in-depth: the no-match branch never reassigns `content`, so the raw comparison was already correct there. The first analysis claimed the reverse; this is the corrected finding. Also fixes an unsound test premise the remote suite caught. The byte-identity precondition in `roadmap_updated is false when ROADMAP.md comes out byte-identical` asserted against a hand-authored, un-normalized fixture — so the very first write reformatted it and the file could not come back identical. The fixture is now written already-normalized, so the assertion compares a normalized pre-image against a normalized post-image and still fails if the flag regresses to a hardcoded `true`. Not platform-specific; it reproduces on macOS too, and the earlier local check simply never exercised it. The sibling true-direction and milestone tests were checked for the same premise and do not share it — they assert `notEqual`, or compare two post-write states produced through the same normalizing seam. Refs #3685 * chore(changeset): backfill PR number for #3691 fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
a2387a0545 |
feat(#3034): add opt-in parallel reviewer lanes (#3822)
* test(#3034): failing-first coverage for opt-in parallel reviewer lanes Executes the real invoke_reviewers dispatch block from review.md against a stubbed gsd_run seam rather than pattern-matching the workflow text, so the two properties that actually carry risk are observable: that every lane is joined before aggregation, and that concurrent lanes cannot tear a line in gsd-review-lane-results.jsonl. Concurrency is proven by a barrier fixture, not by elapsed time -- each stub lane blocks until all lanes have checked in, which can only complete if they overlap. Red against the current sequential dispatch, by design. Refs #3034 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3034): add opt-in parallel reviewer lanes Reviewer lanes within one review pass inspect the same immutable plan snapshot and have no dependency on one another, but were dispatched strictly one at a time, so a multi-reviewer pass cost roughly the sum of its lanes. The serialization is a deliberate protection against provider rate limits, so it stays the default; review.parallel_lanes opts a project out of it. The loop body is hoisted into run_review_lane so the sequential and concurrent paths share one body -- two hand-synced dispatch bodies is the divergence class ADR-2782 spent a phase deleting. Each lane writes a slug-scoped result file, concatenated in selection order after the join: concurrent O_APPEND is atomic only below PIPE_BUF, and write_reviews parses that JSONL to render the models:/model_sources: frontmatter, so a torn line is a broken REVIEWS.md rather than a cosmetic log defect. Aggregating in selection order also keeps the artifact byte-identical between the two paths. The guard is strict equality on "true" and falls back to sequential when config-get fails -- the opposite polarity from the commit_docs guard, because failing open here fires the very requests the default prevents. Also corrects docs/COMMANDS.md and its four locale mirrors, which described --all as running every configured reviewer in parallel when dispatch was in fact sequential. Closes #3034 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3034): de-duplicate dispatch slugs and scope lane locals Review finding (Standards axis): a slug repeated in SELECTED_REVIEWERS would put two concurrent background jobs on the same > -truncated per-lane result file. The shared-append form this replaced could not corrupt itself that way, so de-duplicating is what keeps the concurrent path no worse than the sequential one. Selection de-dupes today -- the roster is a Set and review.default_reviewers normalizes lowercase-unique -- but reachability analysis is not a contract, which is the same reason the roster derivation itself is guarded. Splitting once into DISPATCH_SLUGS also removes the duplicated tr-split the same review flagged: the dispatch and aggregation loops now share one list, which is what guarantees they walk the same slugs in the same order. A plain string accumulator rather than an array, because zsh and bash disagree on array indexing and this block runs under both. Also scopes run_review_lane's locals. Not a live fix -- each dispatched call already forks its own subshell -- but it makes the isolation a property of the function rather than of the dispatch mechanism happening to fork. Refs #3034 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3034): acknowledge review.md growth, drop spent 2295 ack The differential attribution gate reported review.md growing 4173 bytes (30712 -> 34885) with no live acknowledgment. Adds the per-PR fragment it asks for, naming only the one path it reported. Deleting tests/emitted-drift-acks/2295-resolved-model.json is required, not opportunistic. That fragment declared review.md and nothing else, and its ripple is already absorbed into the base, so it is spent -- it can no longer clear anything, which is why the gate still reported review.md as unacknowledged. It could not simply be left alone either: two ack sources may never name the same path, so it blocked this PR's fragment outright. CONTRIBUTING is explicit that a fragment whose last entry is removed gets deleted with it, because an empty fragment signals nothing while its presence reads as a live alarm. Refs #3034 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3034): backfill changeset PR number Refs #3034 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
004e9dd741 |
fix(#3007): resolve Codex reasoning effort per model and make every clamp visible (#3765)
* test(#3007): failing-first suite for per-model Codex effort capability RED by construction. Binds to behavior renderEffortForRuntime does not yet have: an optional third `model` argument, a per-model advertised-level table, `max` passing through instead of clamping to `xhigh`, `minimal` clamping to `low`, `ultra` rejected outright, and clamp visibility (`requested`/`clamped`/ `reason`) so a downgrade is legible from resolver output rather than silent. Two of these pin defects that exist on next today: - `max` is discarded. Both Codex models whose catalog entries are retrievable (sol, luna) advertise `max`; GSD clamps it to `xhigh` and reports nothing. - `minimal` is emitted to a model that refuses it. providerPresets.openai. haiku.low pairs gpt-5.6-luna with reasoning_effort "minimal", and luna's advertised floor is `low`. GSD is sending a value into a document Codex itself validates. The parity test is what pins that fixed, and it names the offending path/model/effort when it trips. Also corrects tests/model-resolver.test.cjs:351, which asserted renderEffortForRuntime('codex','max').value === 'xhigh' -- the defect pinned as though it were a contract. ADR-443 recorded "Codex has no max" as fact and it was true when written; Codex has since added both `max` and `ultra`. That is a stale premise, so the assertion is corrected here rather than worked around. The property test asserts the invariant the whole change exists for: a rendered effort is always a level the target model actually advertises, or an explicit rejection. There is no third outcome. * fix(#3007): resolve Codex effort per model, and make every clamp visible Codex declares supported_reasoning_levels per MODEL and validates against it, so a single per-runtime capability set cannot be right for all of them. GSD's was wrong in both directions at once. `max` reaches Codex now. ADR-443 recorded "Codex has no max" as fact and clamped max -> xhigh on that basis; it was accurate when written, and Codex has since added both `max` and `ultra`. Every Codex model whose catalog entry is retrievable advertises `max`, so the clamp was discarding a level the provider supports, silently, on the most-used path. `minimal` stops reaching Codex. No Codex model advertises it -- both retrievable entries floor at `low` -- yet providerPresets.openai.haiku.low paired gpt-5.6-luna with reasoning_effort "minimal". GSD was writing a value the receiver validates and refuses into a file the receiver reads. Being unconservative in what you send is the half of Postel's rule with no defensible reading, so that preset is corrected and a parity test pins it. `ultra` is refused rather than laddered. Codex's own catalog calls it "Maximum reasoning with automatic task delegation": at ultra, effective_multi_agent_mode returns Proactive and Codex spawns sub-agents on its own initiative, underneath GSD's orchestration rather than inside it (#2167). It is a mode switch, not a reasoning depth, so it is not added to the universal ladder -- which stays provider-agnostic by ADR-443's design -- and it is rejected even for gpt-5.6-sol, which does advertise it. Clamping it down to `max` was considered and rejected: that silently discards what the user actually asked for. Clamping is now visible. RenderedEffort carries requested/clamped/reason and resolve-execution surfaces them. The previous table clamped correctly but invisibly, so a user asking for `max` on Codex had no way to find out they were getting `xhigh` -- exactly the failure mode the robustness principle's modern critique warns about, and why "be liberal" has to mean "liberal and loud". Also closes a latent trap found while reviewing the implementation: the clamp-up loop walks the ladder upward, and for a future model advertising `ultra` but not `max` it would have selected `ultra` as the clamp target -- re-entering by the back door the mode the rejection above exists to keep out. A clamp may never produce a value that a direct request for that value would refuse. Unreachable with today's catalog, which is why no test caught it; a test now asserts the invariant directly. Signature stability is preserved: the third `model` argument is optional and the two-argument form still resolves, against the family baseline. That form's BEHAVIOR does change for `max` and `minimal`, and it must -- keeping the old answer would have fixed the defect only where a model happened to be threaded through and left it live everywhere else. tests/model-resolver.test.cjs:351 asserted the defect as if it were a contract and is corrected here rather than worked around. * fix(#3007): close every review finding on the Codex effort alignment Two isolated reviewers, correctness and security. Both found the same two blockers, and the per-model work was inert on every surface that matters until this commit. BLOCKER — resolve-execution never passed the model and discarded the clamp. cmdResolveExecution called the two-argument form and emitted only effort_rendered/effort_param/effort_propagation, so the per-model table was unreachable from production code (tests were its only caller) and requested/ clamped/reason were computed and thrown away. Requested outcome 3 names "the effective rendered effort in resolver output" specifically, so the feature was unmet on the exact surface the issue asks for. Now passes the resolved model and emits effort_requested / effort_clamped / effort_clamp_reason, flat, matching the existing key convention rather than introducing a nested object. BLOCKER — the docs described output that did not exist. CONFIGURATION.md showed a nested {"effort": ...} sample; the real result is flat and those keys were absent entirely. A reference doc asserting a JSON path a reader can copy is worse than no doc. Corrected against the actual emitted key set. MAJOR — the argv channel still shipped both original defects. EFFORT_ARGV.codex kept minimal in its supported set and still clamped max down to xhigh, so the invocation-time and install-time channels disagreed about the same runtime's capability: --host codex with max emitted xhigh while the generated TOML said max. This is the repo's documented generative-fix-divergence class, so both tables now cross-reference each other and a parity test fails if they ever diverge again. MAJOR — malformed catalog data failed OPEN and could crash the CLI. A null _baseline became an EMPTY Set that is nonetheless truthy, so the nullish fallback never fired and every effort rendered as null. And a non-array value made the Set constructor throw at module load — model-catalog.cjs is required across the whole CLI, so one bad JSON value killed every command, not just codex effort. Guarded on size and filtered to array values; both degrade to the hardcoded baseline. MAJOR — value widened to a nullable string with two consumers left behind. runtime-artifact-conversion passed it straight into injectEffortFrontmatter (a null effort key in generated frontmatter); install-effort-resolver still declared a non-nullable return, a structural lie that silently defeated null checking. Both corrected, both omitting the key on null — the same posture as 'inherit', where omission means "follow the host default". MAJOR — the per-model table is inert today, and the docs now say so. All three shipped models advertise the same usable range and ultra (sol's only differentiator) is rejected for every model, so no observable output differs by model. The table stays because Codex declares capability per model and the sets are free to diverge — a single per-runtime assumption is precisely what went stale and produced this issue — but overselling it as a visible per-model feature would have been the same class of error as the doc blocker above. Tests: three passed under a full revert and are strengthened rather than deleted, since each guards a real contract (#3533's inherit rule, the undeclared-host rule, off-ladder handling) — they now also assert the clamp-visibility fields, which only exist after this change. The fast-check property is kept for its shrinking, and a deterministic nested loop over the full cross-product now sits beside it so coverage is exhaustive rather than sampled. Also folded in earlier: bin/install.js generated the Codex TOML with the two-arg form and would have written a literal null reasoning effort on the ultra path; CONTEXT.md's Model Catalog Module glossary entry now records CODEX_MODEL_EFFORT. The installer defect was found by the co-change gate, not by a reviewer — install.js is a historical co-change partner of model-catalog.cts that this diff had not touched. * test(#3007): correct assertions that pinned Codex's stale effort premise Thirteen pre-existing tests encoded "Codex has no max" as fact and failed on the shipped commit. Every one is a stale pin, not a defect: each was probed against the built module before its expectation was changed, and none failed for a reason other than this premise correction. Kept as its own commit per CONTRIBUTING — a test-fixture correction made stale by a production change must not ride inside another commit, because the release-sdk hotfix cherry-pick filter routes by subject prefix and a correction buried under the wrong prefix ships a half-state (v1.42.3, #3621). The most valuable one was tests/model-resolver.test.cjs's cross-provider validity invariant, which hardcoded the Codex enum as `minimal|low|medium|high|xhigh` and failed with "real API would 400". That message is now false in both directions: Codex accepts `max`, and rejects `minimal`, which no model advertises. The enum is corrected to `low|medium|high|xhigh|max` and the guard is kept intact — it is exactly the "would the real API refuse this" check worth having, and it was right to fail here. It simply carried the stale fact in its own fixture. Test NAMES were corrected alongside their assertions wherever the name asserted the old behavior — "max is Anthropic-only", "max clamps to xhigh", "minimal passthrough". A renamed test that still claims the old thing is worse than a failing one, and a green test whose name states a falsehood is how the next reader inherits the wrong premise. Both channels are covered: install-time (renderEffortForRuntime, and the generated .toml in install-runtime-artifacts) and invocation-time argv (effort-surface-axis). They were deliberately brought into agreement in this change, so their assertions had to move together. Each site carries a #3007 comment recording that Codex gained max/ultra and that capability is declared per model, so a future reader can tell this was a deliberate premise correction rather than a test bent to fit an implementation. * test(#3007): separate the effort-precedence case from the clamp case The previous stale-assertion pass over-corrected one test. It saw `effort: { default: 'max' }` on codex expecting `effort_rendered: 'xhigh'`, assumed the xhigh came from the max→xhigh clamp #3007 removes, renamed it to "max passes through" and changed the expectation to `max`. The remote runner disagreed. Reproduced against the real CLI: with that config and `gsd-planner`, the resolver emits `effort: "xhigh"`, `effort_requested: "xhigh"`, `effort_clamped: false`. The xhigh is produced by effort-resolution PRECEDENCE — gsd-planner is heavy/opus tier and its routing-tier default outranks `effort.default` — so `max` never reaches the renderer at all. The test says nothing about clamping and never did; it only looked like a clamp pin because both mechanisms happened to yield the same string. Restored to `xhigh` and renamed to say what it actually tests. It now also asserts `effort_clamped === false` and `effort_requested === 'xhigh'`, which is what makes it impossible to mistake for a clamp pin again: those two fields prove the value is what the resolver produced rather than something the renderer downgraded. Before #3007 there was no way to tell the two apart from the output — which is precisely why the previous pass could not tell them apart either. Added the test that was actually missing: `effort.agent_overrides`, which outranks the tier default, so the requested level genuinely reaches the renderer and `max` survives to `effort_rendered` end-to-end through the real CLI. Verified by probe before asserting. One test now pins the precedence rule and the other pins the #3007 behavior, and neither can be read as the other. That the clamp-visibility fields are what resolved this is a small argument for having added them. * chore(#3007): backfill changeset pr number to 3765 * test(#3007): put model-catalog under the mutation gate The Stryker shard showed as `skipping` on this PR despite the diff rewriting model-catalog's effort logic. That was legitimate, not a detection bug: `model-catalog` was never in scripts/mutation-matrix.cjs's COVERED map, so the whole module — including everything #3007 touches — sat entirely outside mutation scoring with has_work "false". Registered, with a dedicated spawn-free surface. tests/model-catalog.unit.test.cjs is new: 44 in-process tests, no runGsdTools, no child process, no filesystem, no temp dirs. That shape is not stylistic — it is the #2790 precedent this file already documents. Stryker's command runner treats a whole `node --test <file>` invocation as ONE test costing whatever its slowest case costs, and re-runs it per mutant, so pointing a shard at tests/model-resolver.test.cjs (which uses runGsdTools throughout) would reproduce exactly the 15-minute shard-cap cancellation #2790 hit. The integration file is unaffected and keeps running in full in the normal test job. Coverage spans the module rather than only the diff, because the score is measured over the whole file: effort rendering across every model and ladder level in both channels, the prototype-chain host guard, the exported enums and maps, isAnthropicFlavoredModel's provider namespacings, the profile projections, nextTier, and mergeEffortTierDefaults. The last two were nearly left out and are worth naming — every uncovered exported function is score given away, and mergeEffortTierDefaults turned out to have a genuinely interesting contract (#3531: a partial override merges over the built-ins rather than replacing them, and isValid gates the VALUE, not the tier name, so an unknown tier key is still merged in). Every expectation was probed against the built module before being asserted. minScore is 1 and that is a PLACEHOLDER, flagged as such in the registry comment. Floors in this repo are measured, not chosen — the existing entries sit at 94, 75 and 56 — and they can only be measured in CI, because mutation shards run `node --test`, which is hard-blocked locally. The first CI run on this branch reports the real number and the floor gets ratcheted to it before merge. A placeholder of 1 reaching `next` would make the gate decorative: it would pass whether or not a single mutant is ever killed. Note the target is "never regress from measured", not a fixed 80 — planning-inspect sits at 56 and is documented as an accepted ratchet candidate. * test(#3007): bootstrap model-catalog's mutation floor legally The placeholder floor was structurally illegal and the remote run said so. tests/mutation-matrix-ratchet.test.cjs guards the guard: every COVERED module must carry a matching RATCHET_BASELINE entry in the same diff, minScore must EQUAL that baseline, and it must be at least 50. `minScore: 1` failed all three. That is the ratchet working exactly as intended — a floor nobody can satisfy accidentally is the point of it. Bootstrapped at 50 in both places. Fifty is not a measured score and the comment says so plainly: it is the minimum the guard permits, and it coincides with Stryker's own configured `break` threshold, so it is the lowest legal starting point for a module that has never been measured. It still must be ratcheted to floor(measured) - 1 before this PR merges. Also corrected a real defect in the file's own instructions. "HOW TO UPDATE" step 1 read "Run the per-module Stryker shard locally" — which cannot be done here, and which the same file contradicts eighty lines further down, where the #2790 scores are recorded as "not a local run; mutation shards run `node --test`, hard-blocked in this repo's local environment". stryker.config.mjs confirms the command runner invokes `node --test` once per mutant, and .claude/hooks/block-local-node-test.sh denies exactly that. So the documented first step sends the next contributor at a wall. Rewritten to describe the path that works — push, read the measured score off the CI shard, then set the floor and its baseline together in one diff — and to say why local measurement is not available, so nobody rediscovers it the slow way. GOODHART SAFETY is untouched. The two-step is inherent to the environment rather than a shortcut: a floor cannot be measured before the first CI run exists, and the guard rightly refuses to accept an unmeasured one below its minimum. * test(#3007): ratchet model-catalog's mutation floor to its measured score The shard ran in CI and reported 59.62% — 248 mutants killed, 168 survived, no timeouts, no errors (run 32605073352, job 97108869486). Floor set to 58 per this file's own rule, minScore = floor(measured) - 1, which is the same arithmetic every sibling entry used: 57.03 to 56, 76.58 to 75, 95.65 to 94. Both halves moved together, because the ratchet guard asserts minScore equals its RATCHET_BASELINE entry and would reject them drifting apart. The spawn-free unit surface is vindicated by the clock: 57 seconds, against a 15-minute shard cap and a 9m46s frontmatter shard in the same run. That was the whole reason for creating tests/model-catalog.unit.test.cjs rather than pointing the shard at tests/model-resolver.test.cjs — #2790 recorded shards being CANCELLED at that cap when they targeted a runGsdTools-heavy integration file. The registry comment is rewritten rather than deleted. It previously warned that the floor was provisional and must not ship that way; leaving that text next to a measured floor would make the file lie in the other direction. It now records the measurement the way the sibling entries do, including that 59.62 sits below TARGET (80) and is therefore a ratchet candidate like planning-inspect at 56 — comfortably clear of its own floor with real room to grow. Raise it as the tests improve; never lower it. Worth stating plainly: 168 surviving mutants is not a clean bill of health. It is an honest floor for a module that had NO mutation coverage at all an hour ago, and it is now pinned so it cannot silently regress. --------- Co-authored-by: sim <sim@local> |
||
|
|
2b42b28687 |
fix(#3659): make the worktree base-check trust evidence, not baseRef (#3736)
* test(#3659): baseref-head suppress must be mode-aware regression rows * fix(#3659): make baseref-head suppress mode-aware and thread isolation mode * fix(#3659): review fixes - stale advice purge, message pins, mode alias * fix(#3659): pick-interceptable emit seam, ack merge, writeSync pin * test(#3659): rewrite set-baseref pin, fix writeSync row stub * chore(#3659): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
9a69a86f42 |
enhance(#2971): strict planning filter mode for /gsd-pr-branch (#3720)
* test(#2971): failing-first suite for the pr-branch planning-path filter Binds the not-yet-built planning.pr_strict mode and the corrected filter recipe for /gsd-pr-branch across six layers: pure classification and forbidden-path predicates, real-git fixtures that run the cherry-pick filter loop end to end, config-key registration through the real CLI and both manifests, the executed worktree-materialization claim the issue's triage asked to establish, fast-check properties over arbitrary path sets, and a drift guard over the shipped workflow. Two live defects in today's shipped recipe are pinned as regressions, both reproduced empirically first: `git rm -r --cached` stages a deletion of any .planning/ path the target branch already tracks, so the generated PR removes the base branch's planning files; and the same command leaves the cherry-picked file untracked on disk, so a second commit touching that path aborts the pick with "untracked working tree files would be overwritten" and every remaining commit is silently dropped. The test helper parses the canonical path lists out of gsd-core/workflows/pr-branch.md rather than restating them, so the workflow stays the single source of truth and the suite cannot drift from what ships. Refs #2971 * feat(#2971): strict planning filter mode for /gsd-pr-branch Adds planning.pr_strict — a boolean, default false, that selects what /gsd-pr-branch means by "filtered". Default mode is unchanged: structural planning state survives into the PR branch and the nine transient subdirectories do not. Strict mode drops every .planning/ path, structural files included, and carries a commit over only when it touches at least one file outside .planning/. Strict mode is what makes planning.commit_docs: true safe for a project that versions its planning tree locally but publishes none of it. The alternative posture, commit_docs: false, silently costs parallel executor isolation — a worktree is checked out from a commit, so an untracked or ignored .planning/ is simply absent inside it and the executor has no PLAN.md to read. That claim is now established by an executed fixture rather than inherited. The two path lists are declared once and both projections derived from them, so create_pr_branch and verify can no longer disagree about what the filter promised. verify previously counted every .planning/ path against a documented success criterion of zero while create_pr_branch was specified to preserve five structural files, so a correct run reported itself as failed on every phase that touched STATE.md — which is every phase. It now asserts against the active mode, and names the .planning/ paths default mode deliberately keeps rather than trading a wrong signal for silence. Two verified defects in the same recipe are fixed alongside, because strict mode would have amplified both. `git rm -r --cached` staged a deletion for any .planning/ path the target branch already tracked, so the generated PR removed the base branch's planning files — under strict mode that would have been the entire tree. The same command left the picked file untracked on disk, so a second commit touching that path aborted the cherry-pick with "untracked working tree files would be overwritten" and every remaining commit was silently dropped. Both were reproduced against real git before being fixed. The filter now forces excluded paths back to what the PR branch's HEAD carries, in the index and the working tree; a conflict outside the filter halts instead of being improvised past; a commit left empty by filtering is skipped rather than failing. A clean-working-tree precondition makes the worktree half safe. Closes #2971 * fix(#2971): unwind the checkout on a conflict halt, and test the real recipe Two review findings, both fixed in place. The isolated adversarial pass found that the conflict-outside-the-filter branch exited while leaving the user checked out on the half-built PR branch with cherry-pick state still live — this loop runs in the user's own working directory, so stranding them there is a real cost even though it is not a vulnerability. The branch now aborts the pick, returns to the original branch, removes the partial PR branch, and says so before exiting. The standards pass found the L2 fixtures executed a hand-written mirror of the cherry-pick filter recipe rather than the recipe itself, so a reordering in the workflow would not have been caught — and the order is load-bearing, since restoring a path from HEAD before removing it inverts the filter. The helper now extracts the canonical loop from the shipped workflow and the fixtures execute that verbatim, which also gives the conflict-halt unwind above real coverage. The drift guard additionally pins the two commands' relative order and asserts the workflow carries exactly one canonical loop. Also records the publication gate in the CONTEXT.md glossary next to the commit gate it is distinct from. Refs #2971 * fix(#2971): make the conflict-halt unwind actually unwind, and use the colon slash form The remote matrix caught two defects in the previous commit. The halt path claimed to restore the original branch but did not. `git cherry-pick --abort` does not apply to a single `--no-commit` pick with no sequencer file, and the fallback left the unmerged index in place, which makes `git checkout` refuse — a failure the `2>/dev/null || true` then swallowed, so the user was told they had been restored while still sitting on the half-built PR branch. The unwind now drops sequencer state, hard-resets the disposable PR branch to clear the unmerged index, and only claims a restore when the checkout actually succeeded; when it does not, it says where the user is and gives them the two commands to finish it by hand. Verified against real git: exit 1, the conflict named, HEAD back on the original branch, the partial branch gone, a clean tree and no CHERRY_PICK_HEAD. Two runtime-loaded source artifacts used the retired `/gsd-<cmd>` hyphen form, which names a command no runtime registers. The canonical authoring token for workflows and references is `/gsd:<cmd>`; docs keep the hyphen form, so the documentation added in this branch is unaffected. The comment in src/config.cts moves to the colon form too, since it propagates into the generated lib. Refs #2971 * docs(#2971): backfill PR number into the changeset fragments (#3720) --------- Co-authored-by: sim <sim@local> |
||
|
|
14679b866b |
enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability Binds the approved triage shape before any of it exists: - containment — the execute:wave:post hook must not render unless workflow.live_dom_uat is true AND the capability resolves active (fail-closed on a missing state entry, and on a non-boolean value) - criterion 4 — agents/gsd-executor.md carries no browser MCP family; asserted as an absence, which is the only way it is observable - Hyrum guard — the pre-existing mcp__playwright__* branch must stay outside the key-gated block, or upgrading silently removes working automated UI verification for every current Playwright-MCP user - parity — the browser glob list now lives in two surfaces (agent frontmatter + workflow detection block); the assertion fails if either gains or loses a family without the other Red by construction: the capability, agent and workflow block do not exist yet. Verified on the remote runner. Refs #2856 * enhance(#2856): add default-off live-DOM UAT capability A phase whose acceptance criteria needed a live DOM could not be finished by the agent that executed it: gsd-executor carries no browser tools, so it correctly returned checkpoint:human-action even though the work was not human-only, just tool-less. Every such phase degraded to "executed, then finished by hand in the orchestrator", and autonomous: false could not distinguish "a human must judge this" from "the executor lacks the tool". Implements the shape approved at triage, not the one reported. The executor's tools: line is NOT widened, in any configuration: for a first-party agent the static list is the only control that exists (ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one default-off capability owns the key, the agent, and the step: - capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat (boolean, default false), one additive step at execute:wave:post (onError: skip, gates: []), so it can never halt a wave - agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP globs, in its own tools: line, with no Bash - verify-work automated_ui_verification — a gsd:live-dom-families block naming both new families AND the key; presence alone never activates Two independent fail-closed gates: isCapabilityActive renders a hook only on state.active === true, plus the step's own `when`. The pre-existing mcp__playwright__* branch keeps the gating it already had and stays outside the new block. Pulling it behind a default-off key would have silently removed working automated UI verification from every current Playwright-MCP user on upgrade. Also closes a host gap this surfaced: execute:wave:post dispatched only contribution + gate, so ANY registered step was declared and silently never run — exactly the single-kind hand-roll loop-hook-dispatch.md names. Step 5.75 now dispatches every kind == "step". The browser-profile lock is tolerated, not coordinated: --isolated is a flag on the operator's own MCP-server registration that GSD neither launches nor parameterizes, so the verifier reports could_not_look / profile_locked, names the flag, and stops. DOM-VERIFY.md keeps could_not_look and nothing_to_report distinct behind a closed reason enum — collapsing them is the ambiguous-run-notes defect reported. Verified on the remote runner. Closes #2856 * fix(#2856): apply review findings from the orthogonal passes Correctness pass (blocker): - delete detectionBlockIsCrlfSafe. It was pass-always: it read the file, replaced LF with CRLF, then indexOf'd marker strings that contain no newline, so the replacement could not change the result and the assertion could never fail for the reason it stated. There is no real CRLF risk on this surface either — the gsd:live-dom-families block has no parser, only human and agent readers. Deleted rather than replaced, per the repo's pass-always-test rule. Isolated security pass (two minors, both real): - execute-phase.md step 5.75: this change is what first activates kind == "step" dispatch at execute:wave:post, which newly opens the ref.command shell path at that loop point. Our own step uses ref.agent and never touches it, but the door is now open, so the step-dispatch line carries the same in-context validate-before-shell warning the sibling gate-dispatch line directly below it already carries. - gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker influenced. Require it wrapped in inline code or a fence, kept short, and never left reading as a directive to the next reader. Verified on the remote runner. Refs #2856 * fix(#2856): settle the new-agent roster ripple Checkpoint 2 returned 28 failures, none in the new suite — all of them the guards that exist to make adding an agent a deliberate act. Each is a real boundary that had to move: - docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526), so the browser globs lose their backticks; primary-agent counts 21->22, roster 33/34->34/35, Verifiers category 1->2 - docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md to be classified exactly once - gsd-dom-verifier: add the anti-heredoc instruction and the commented hooks: frontmatter pattern both agent gates require - gsd-core/bin/shared/model-catalog.json: every shipped agent needs a profile entry (#3229) - copilot-install / kilo-upgrades / qwen-upgrades: expected agent list and the 34->35 roster boundary - execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately carries one step now. Asserted as an exact shape — one step, capId live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a real guard against accidental change rather than being relaxed Two findings worth naming: mcp-tool-inheritance (#2526) rejected the agent for documenting mcp__playwright__* while its tools: line withholds it — a dead instruction that invites the agent to claim a path it cannot take. The prose now names the Playwright MCP family without the dispatchable token, in both the agent and the capability fragment. runtime-launcher-parity rejected the new gsd_run call: each fenced block is its own shell, so a workflow step file invoking gsd_run needs its own canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs. That script also normalizes explore.md, which is unrelated pre-existing drift the parity check tolerates, so it is reverted to keep this diff scoped. The emitted-drift ack supersedes the spent #3370 entry for execute-phase.md — it is merged into next, so its ripple is absorbed at the base and it can no longer clear anything. That is the same supersede the #3370 entry itself performed on the spent #3324 fragment. Its unrelated execute-plan.md entry is untouched. Verified on the remote runner. Refs #2856 * fix(#2856): drop the stale emitted-drift ack entry The automated-ui-verification.md entry was written speculatively rather than from a reported growth, and the check names that precisely: an ack "written or reworded in THIS diff, but nothing here needed it, so it explains nothing". The growth tier keys on the bare filename as it appears under gsd-core/workflows/ or agents/. automated-ui-verification.md is nested under verify-work/steps/, so it was never in the tracked set — only execute-phase.md was ever reported, both before and after the launcher preamble landed. Only ack what the check actually reports. Verified on the remote runner. Refs #2856 * chore(#2856): backfill changeset pr number pr:0 -> 3716. The placeholder fails both changeset-lint (fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by design and can only be resolved once the PR number exists. Both now report ok against GITHUB_BASE_REF=next. Refs #2856 --------- Co-authored-by: sim <sim@local> |
||
|
|
77fa08f1e8 |
fix(#2773): feed the spec-phase edge probe English-translated requirement text (#3713)
* test(#2773): failing-first contract and premise tests for translated edge-probe input Locks the Step 5.5 contract that a response_language project must feed the edge probe an English translation of each requirement's text, and binds that advice to measured engine behavior: the same requirement classifies to zero shapes in Portuguese and to collection/adjacency/empty/ordering in English. Also pins the honest limit — the issue's own repro sentence classifies to [] in English too, so translation is necessary but not sufficient and the authored shapes override is the documented fallback. Red before the doc change; the assertions are all false today. Refs #2773 * fix(#2773): feed the spec-phase edge probe English-translated requirement text The shape cues in src/edge-probe.cts are English word-boundary regexes, so a project running with response_language set wrote its SPEC requirements into the Step 5.5 $REQS_JSON heredoc in that language, matched no cue, classified to zero shapes, and landed every row in the unclassified sentinel (#1110). The taxonomy contributed nothing and --auto left it all unresolved — the probe was a silent no-op for exactly the spec type it exists to harden. Step 5.5 now states that the $REQS_JSON payload is engine input rather than user-facing output, so the response_language rule does not govern it: each requirement's text carries a faithful English translation, the SPEC keeps its original language, and requirement ids are never translated or renumbered. The instruction sits before the heredoc on purpose — the downstream APPLICABLE=0 warning fires only when every requirement is unclassified, so a partly-classified non-English spec would otherwise slip through with no signal at all. Measured against the compiled engine: the same requirement returns [] in Portuguese and collection -> adjacency/empty/ordering in English. Also measured: the issue's own repro sentence returns [] in English too, so translation is necessary but not sufficient — the instruction therefore points at the authored shapes override for prose carrying no cue in any language rather than promising that translation restores classification. Doc scope only, per the triage disposition on the issue. The compiled engine is untouched; the lang-hint / per-language cue-set fix is a separate follow-up. Closes #2773 * fix(#2773): clean up the edge-probe temp file on the placeholder-guard exit path Surfaced by the isolated security review of this branch. Between the mktemp and the unconditional cleanup, Step 5.5 has two sibling guards that disagreed about their own invariant: the engine-failure guard runs rm -f "$REQS_JSON" before exiting, while the empty/placeholder guard directly above it exited without one. A spec run that tripped the placeholder check therefore stranded a temp file holding the SPEC's requirement text in TMPDIR, once per failed run. The added contract test walks the region between the mktemp and the unconditional cleanup and asserts no exit path leaves the file behind, so the two guards can no longer drift apart. Proven to bind: run against the pre-fix file the walker reports the leaking exit; against the fixed file it reports none. Refs #2773 * docs(#2773): record the edge probe's English-cue input constraint in the predicate store The co-change gate flagged CONTEXT.md (13 co-changes with spec-phase.md) and docs/CONFIGURATION.md (11) as candidate-missing-updates, and both were real gaps rather than incidental coupling. CONTEXT.md's EdgeCompletenessProbeModule entry documents the input contract for classifyShape but did not record that SHAPE_CUES are English word-boundary patterns — so the predicate store implied text was language-agnostic, which is what a future agent reads before touching this seam. docs/CONFIGURATION.md's response_language row is what a non-English project reads when it turns the setting on; it now names the one deliberate exception and links to the FEATURES.md explanation, so the interaction is discoverable from the config key rather than only from the workflow. CONTEXT-INDEX.json regenerated via gen-context-index.cjs --write. The drift-ack fragment is updated for the final byte range and now also records the placeholder-guard cleanup fix folded into the same block. Refs #2773 * fix(#2773): append the growth rationale to the existing spec-phase.md ack entry The remote runner caught this: emitted-attribution.test.cjs pins the 0000-legacy-migration.json spec-phase.md entry permanently (the #2914 migration regression test asserts the exact '31987 -> 31997' delta text survives), so removing it to avoid a duplicate-key collision with a new fragment broke that test instead of satisfying the ratchet. The entry is an accreting log, not a single-use slot — #2733, #3132 and #3102 were each appended to the same reason string by later PRs, which is how a shared growth key coexists with the rule that two ack sources may never name the same path. This appends the #2773 rationale the same way and drops the separate fragment, whose spec-phase.md key was the collision. Verified locally by reproducing both affected tests against the real fragment before re-dispatching: the pinned delta survives, grown[0].acked is true, staleAcks is empty, and all 35 entries still read as spent. Refs #2773 * docs(#2773): add a how-to for probing edges in a non-English project The phase gate's enablementSequence check caught a wrong call of mine. I had recorded that no how-to was owed because the user takes zero extra steps — the workflow translates the probe input itself. Written out, though, the sequence from off to value is two steps and step 1 depends on response_language, a setting owned by a different capability than the edge probe, which is exactly the condition the how-to test names. There is also real task content a reference table cannot carry: the three-way split between a few unclassified rows (the classifier's recall gap), every row unclassified (the probe could not read the spec at all), and the silent partly-classified case where the APPLICABLE=0 warning never fires. That last one is what a user would otherwise misread as a clean bill of health. Shaped after the resolve-edge-coverage-findings / resolve-unreachable-guard siblings and indexed from docs/README.md next to its closest relative. Refs #2773 * chore(#2773): backfill the changeset PR number pr:0 placeholder replaced with the real PR number now that #3713 exists. Refs #2773 --------- Co-authored-by: sim <sim@local> |
||
|
|
adb46cdd85 |
feat(#2734): surface STATE.md commit-age on the statusline (#3700)
* test(#2734): failing-first suite for the statusline STATE.md freshness marker Binds the contract before any hook change exists: a `state ~N commits back` segment gated on the state_head stamp landed by #2622, firing at the same advisory threshold /gsd-health's W024 uses rather than at > 0. Covers all five acceptance criteria — threshold parity (19/20/21 boundaries), both renderers including formatGsdStateCompact, an exact spawn-count assertion, repo-pinning and sub_repos degradation, and behavioral parity against readStateHeadFreshness rather than a source-grep of the two fence copies. 52 example-based tests plus 5 seeded fast-check properties. Red now by design. * feat(#2734): surface STATE.md commit-age on the statusline Adds an opt-in `state ~N commits back` marker to the GSD-state segment, consuming the `state_head` stamp and freshness contract landed by #2622. A solo developer returning to a project reads "Phase 4, executing" in STATE.md and acts on it, without noticing the codebase moved 40 commits since that line was written. /gsd-health reports it as W024, but only if you think to run it; the statusline is the surface you see without asking. Fires at STATE_HEAD_ADVISORY_COMMITS (20), the same threshold W024 uses, not at > 0: with commit_docs:true the commit carrying a STATE.md sync advances HEAD by one, so > 0 would alarm permanently on a fresh project. Costs exactly one bounded git subprocess per render and none when disabled. `rev-list --left-right --count` answers ancestry and distance together, and repo pinning is a filesystem check mirroring projectOwnsItsRepo rather than a --show-toplevel compare, which is unreliable on macOS /private/var and Windows 8.3 paths. Every unresolvable input degrades to the tri-state unknown -- the marker is absent, never a "fresh" claim the project cannot substantiate: a malformed stamp, a root that does not own its .git, a sub_repos workspace, history rewound past the stamp, or git being unavailable. Also collapses statusline config resolution onto one resolveStatuslineOptions() seam. runStatusline() and renderStatusline() duplicated it byte-for-byte; one copy is what keeps a newly-added key from reaching only one of them. * test(#2734): route the e2e spawn through the process seam and fix fixture leaks Review findings from the two orthogonal passes: - `bothEntryPointsResolveOptionsIdentically` spawned a child and substring-matched its stdout to test a pure function. It now calls resolveStatuslineOptions() directly — no subprocess, no text matching. - `skipsFreshnessWorkWhenTodoTaskActive` genuinely needs a child (the !task gate lives in runStatusline, which reads stdin), so it now spawns through tests/helpers/process-seam.cjs and proves the negative with a filesystem fact: the git shim appends to a marker file on every invocation, and the assertion is that the marker never appears. Stronger than asserting text is missing, and it drops the last stdout substring match in the block. - Every fixture-creating test now registers `t.after(() => cleanup(dir))` instead of a trailing cleanup(dir), which leaked the temp repo on assertion failure. derivationAgreesWithStateModule reassigns `dir` across five fixtures, so it binds each directory at scheduling time rather than cleaning only the last. Also corrects markerCoexistsWithMilestoneComplete, which asserted the wrong expectation rather than finding a code defect: `percent` drives the progress bar too, so the milestone segment reads "v1.9 [##########] 100%". The marker appends after it, which is what the test exists to prove. CONTEXT.md's opt-in statusline key list was missing statusline.show_git as well as the new key; both are now enumerated. * docs(#2734): backfill changeset PR number (#3700) --------- Co-authored-by: sim <sim@local> |
||
|
|
2fca0e17e4 |
enhance(#2554): resolve code review depth from path-scoped override rules (#3695)
* test(#2554): failing-first suite for path-scoped code review depth overrides Binds the not-yet-built code-review-depth module: segment-aware path-prefix matching of a changed-file set against ordered {paths,depth} rules, resolution order flag > strongest matching rule > global > standard, typed validation errors, and the large-scope downgrade boundary. Also proves behaviorally that workflow.code_review_depth_overrides is not yet a registered config key. Refs #2554 * feat(#2554): resolve code review depth from path-scoped override rules Adds workflow.code_review_depth_overrides — an ordered array of {paths, depth} rules matched against a review's changed-file set by segment-aware path-prefix comparison. Resolution order is --depth= flag, then the strongest matching rule, then workflow.code_review_depth, then standard; a matching rule replaces the global rather than being max'd with it, so quick and standard rules stay meaningful. Glob metacharacters are a hard configuration error rather than sugar for a prefix, and malformed rules halt the review instead of degrading to standard. The resolver is pure and reports its own provenance, so the workflow can print the resolved depth and the rule that matched. The pre-existing >50-file deep-to-standard downgrade moves into the module and now names the rule it overrode. The key is registered centrally rather than as a capability config slice: the federated slice channel admits only boolean/string/number/enum, so an array slice would be dropped as malformed. Closes #2554 * test(#2554): correct depth-provenance assertions and pin out-of-repo paths Two corrections to the failing-first suite. The source assertion for a non-matching rule with no global configured expected 'config'; with no global set the depth comes from the default, and a companion assertion tolerated either value, so both passed against an implementation that derived provenance from whether any rules existed rather than from where the depth came from. The out-of-repo absolute-path case used a home-directory path that matched neither implementation, so it never exercised the defect it named. It now pins the discriminating cases: an absolute path outside the repo root must not match a repo-relative rule, and one under the root must. * docs(#2554): document path-scoped code review depth overrides Reference rows for workflow.code_review_depth_overrides in the configuration, features and commands references plus the locale copies that carry those tables, and in the planning-config reference. Explanation of why escalation is whole-review rather than per-file and why v1 is prefix-only. New how-to for scoping review depth by path, carrying the configuration-error reason table and the distinction between nothing to report and could not look. CONTEXT.md glossary entry and the INVENTORY row for the new CLI module. ja-JP and ko-KR CONFIGURATION.md carry no code_review keys at all, and ko-KR and pt-BR FEATURES.md carry no code-review config table, so those files are deliberately untouched. * fix(#2554): make the depth-misconfiguration halt executable and reject control chars Three review findings, all in this change. The misconfiguration halt was prose rather than shell: the error-printing fence was followed by an unconditional extraction fence, so an ok:false result threw and left the depth empty instead of stopping the review. Prose is not a guard — the two fences are now one block with a real conditional, and anything that is not the literal string true fails closed. An interior control character in a rule path survived validation and reached the provenance string and the summary box; rule paths now reject control characters via a new PATH_CONTROL_CHAR reason, after the glob check so precedence is unchanged. That in turn makes the field record safe to delimit, so the seven node invocations that each re-parsed the same result to read one field collapse to one. Also corrects the glossary entry's illustrative paths, which the glossary-ref check read as real repository references. * fix(#2554): use the fast-check v4 string API and acknowledge workflow growth Two failures from the remote matrix on d3111f45, both this branch's. The property block built its segment arbitrary with fc.stringOf, removed in fast-check v4. Because the arbitrary is constructed in the describe body, the throw took out all four property tests rather than one — they had never executed. Rewritten to fc.string({unit, ...}), the form this repo already uses in emitted-attribution.test.cjs. Every other fast-check helper in the file was audited against the installed module. The emitted-attribution growth arm needed an acknowledgment for code-review.md, which grew 5376 bytes. The pre-existing 3503 fragment keying the same file is spent — its ripple was absorbed when #3503 merged, and the base file is exactly the 34435-byte baseline this growth is measured against — so it cannot clear anything, while the ack lint hard-fails on a duplicate key across two sources. Removed it in favor of the new fragment, which is exactly how #3503 itself replaced the spent 3191 fragment. * docs(#2554): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
8526bd46f8 |
enhance(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR (#3688)
* test(#2475): widen the item-1 effort-caller guard to both CLI argument shapes The guard matched only `resolve-execution ... --effort\s`, but the CLI also accepts `--effort=<level>` (gsd-core/bin/gsd-tools.cjs). A workflow written with the equals form was a live invocation-override caller the guard passed silently, along with `--effort` at end-of-input. Lift the matcher to a shared predicate and assert it directly against every shape the CLI accepts, plus the decoys it must not fire on (--effortless, a bare --effort with no resolve-execution, item 6's --attempt caller, a call and flag split across lines). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR ADR-443 sat at Proposed on one condition: its Decision item 1 orchestrator invocation override needed a caller in shipped orchestration. Per the maintainer's ruling, take unblock path (b) for item 1 only -- record that the override is an operator-facing CLI surface, deliberately not driven by shipped orchestration, and ratify. The ADR's own path (b) wording is not adopted verbatim: it says the scope is limited to static install-time propagation, which is false on both counts -- item 6 has a live caller (#2296) and #2481 delivered a live invocation-time argv channel. Only one precedence step is narrowed. No consumer was invented to clear the gate: nobody has asked for a per-run effort override, and #2475's actual complaint is already closed by the cascade-to-argv path. The amendment states explicitly that --effort remains supported and is not deprecated, so the scoping is not read as dead code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2475): correct two bare-`gsd` invocation examples to `gsd_run` There is no `gsd` binary -- package.json exposes gsd-core, gsd-tools, gsd_run and gsd-mcp-server. Both sites presented a command that cannot run as written. One is in this branch's own new ADR-443 amendment; the other is a pre-existing error in the docs/CONFIGURATION.md assumption_delta row, fixed here rather than deferred. No translated copy carries either line, so no i18n drift is created. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2475): record the guard/CLI divergence risk on the item-1 matcher The predicate independently models gsd-tools.cjs's argument parser rather than sharing a constant with it, so a third --effort spelling would leave the guard reporting green while ADR-443's ratifying invariant silently stopped holding. Name that risk where the next editor will meet it. Also restores the bounded-prose rationale that was attached to the eslint directive removed in 39793079c -- the directive went unused once the regex moved to a const, but the reasoning it carried is still worth having. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1bf73d957b |
enhance(#2295): record the resolved model per reviewer in REVIEWS.md frontmatter (#3649)
* test(#2295): failing-first coverage for per-lane resolved-model recording * feat(#2295): record the resolved model per reviewer lane * docs(#2295): document the recorded reviewer model and its provenance * fix(#2295): refuse control characters in a recorded model value * test(#2295): correct watermark assertions for the widened mark shape * fix(#2295): anchor the role-manipulation injection pattern at a word boundary * feat(#2295): record the applied reasoning effort in the model value * chore(#2295): backfill changeset pr number * chore(#2295): restore em-dash in changeset body --------- Co-authored-by: sim <sim@local> |
||
|
|
1adf6d2245 |
fix(#3620): point the docs at files that actually exist (#3658)
* fix(3620): point the docs at files that actually exist docproof found 34 stale references; the reporter hand-read all 34 and reported the 8 that are real, explaining why the other 26 are deliberate (files the documents themselves label legacy or "superseded by", and one pre-Diataxis link label whose target still resolves). Those 26 are left alone — re-touching them would contradict the issue's own analysis. Every claim was re-verified against git ls-files at HEAD before editing. docs/INVENTORY.md said its roster is anchored by six drift-control tests. Five are gone (commands-doc-parity, agents-doc-parity, cli-modules-doc-parity, hooks-doc-parity in 5d8a8c4d; command-count-sync in |
||
|
|
fe64704ace |
enhance(#3588): add an opt-in commit_docs pre-commit hook (#3609)
* feat(#3588): add an opt-in commit_docs pre-commit hook Final phase of epic #2292, scope narrowed to opt-in by maintainer decision: default-on installation and the bin/install.js wiring it would have required are explicitly out of scope. Enabling is an explicit verb call. The hook is written to the repo's real hooks dir resolved via git rev-parse --git-path hooks, so a linked worktree or submodule whose .git is a FILE works rather than getting a literal .git/hooks path. It refuses rather than overwrite a foreign pre-commit, refuses to delete one it did not write, and refuses outright when core.hooksPath is already set -- a written-but-ignored hook is worse than a refusal. Ownership is detected by marker presence, not byte-equality, so a user who appends a line does not make it unrecognizable. Deliberately NOT included: teaching cmdCheckCommit the per-phase commit_docs tier. #3587 was still unmerged when this landed, and implementing precedence against helpers that did not yet exist would have meant a second copy of the resolution chain -- the divergence class this epic has spent three phases fighting. That follows as its own change now that #3587 is on next. The ordering constraint is recorded in the design doc: this must not merge before #3587, or the hook would block a commit cmdCommit itself allows. * fix(#3588): teach the commit_docs guard the per-phase tier and -z paths Part 1, deferred until #3587 merged. cmdCheckCommit read only project-level commit_docs, so once #3587 landed, a phase with phase_commit_docs true under project false was ALLOWED by query commit and BLOCKED by this guard -- and the hook shipped in this same branch shells out to it. It now derives the staged phase via the single-owner detectPhaseNumberFromFiles and resolves through #3587's own resolveCommitDocsPolicy rather than a second precedence copy. Also fixes a proven false negative in the harm direction. git diff --cached --name-only C-style-quotes any path with non-ASCII or special characters, so a staged .planning/cafe.md was emitted as a quoted string, failed startsWith('.planning/'), and slipped past the guard entirely under commit_docs:false. Reading with -z and splitting on NUL removes the quoting at the source. The f.startsWith('.planning\\') branch was dead code under that read -- git emits /-separated paths on every platform -- and is removed rather than left implying coverage it never provided. The earlier C7 test pinned the buggy behavior as intended; it now asserts the file is detected and the commit refused. Self-caught: the commit-docs-guard verb was wired into the routers by this branch's earlier pass but missing from the top-level help listing. * test(#3588): replace try/finally with t.after, add negative-routing cases Standards review findings. CONTRIBUTING bans try/finally inside a test body outright -- it masks failures -- and B8 used one for worktree cleanup. Now t.after(), assertions unchanged. The new commit-docs-guard command family had zero negative-routing coverage, which CONTRIBUTING requires for any change to command dispatch. B11-B15 cover no subcommand, unknown, empty string, whitespace-only and a flag-shaped value, each asserting non-zero exit, a structured error, no stack trace, and -- the one that matters for a command that writes into a user's repo -- that NO hook is written in any of them. Those tests were verified to fail when routeCommitDocsGuard's else-branch is neutered, so they exercise the routing guard rather than any convenient error path. Also made two error() calls' control flow explicit with a return; they were safe only because error() is typed never two files away. * chore(#3588): backfill changeset pr number to 3609 * test(#3588): skip Windows-unrepresentable fixtures on win32 CI's Windows shards caught two of my own tests: fixtures whose filenames contain a quote and a backslash. Both are illegal on Windows -- backslash is the path separator, quote is invalid on NTFS -- so fixture creation failed before any assertion ran. Test-portability defect, not a production one. Those inputs cannot exist on that platform, so the guard has nothing to detect there. Both now check process.platform FIRST, before any fs or git call, and use t.skip() rather than a bare return -- a bare return registers as a PASS and would hide the gap it is meant to record. Each carries a comment saying the input is unrepresentable rather than unverified, so nobody later re-enables it. No padding added: the cafe.md case already exercises git's C-quoting path on every platform, since non-ASCII names are legal on NTFS. This is exactly the coverage the Linux-only remote matrix cannot provide, which the PR body already stated -- CI's Windows shards are what caught it. --------- Co-authored-by: sim <sim@local> |
||
|
|
debeabd524 |
enhance(#3587): add a per-phase commit_docs override (#3601)
* feat(#3587): add a per-phase commit_docs override Delivers epic #2292's second user story: commit an architecture phase's artifacts while execution phases stay local. commit_docs was project-wide and binary, so the only choices were all phases or none. Shape is a config dynamic key phase_commit_docs.<phase-id>, following the 14 existing dynamicKeyPatterns precedents rather than inventing a PLAN.md frontmatter spec -- which #2292 itself flags as becoming its own maintenance surface. Tier 1 resolves in cmdCommit, NOT in loadConfig: loadConfig has no phase context and is called by nearly every command, so threading one through it to serve a single caller would be a far larger blast radius for no gain. The phase comes from detectPhaseNumberFromFiles, which cmdCommit already computes for branch naming and which is already hardened against the #2539 project-code bug. Suppression by the per-phase tier returns its own reason rather than reusing skipped_commit_docs_false -- telling a user their project setting is false when it is true would be actively misleading. Additive; the two existing reason strings that agents/gsd-executor.md matches on are unchanged. The manifest's phase-id pattern is a hand-copy of PHASE_NUMBER_TOKEN_SOURCE because the manifest is hand-maintained JSON, so a behavioral parity test asserts both surfaces accept and reject the same token shapes. * fix(#3587): fold tests, close review findings, update reference docs Fold: the new tests were added as their own file, which required loosening a grandfathered lint-test-file-count bucket 5-to-6. A ratchet exists to go down only. commit-docs-bypass.test.cjs is the established commit_docs test home and already hosts two folded suites, so the tests fold there as a third block and the allowlist is reverted untouched. Standards review: CONTEXT.md and the test header both cited a phase-commit-docs-manifest-parity.test.cjs that never existed; a repo-wide sweep found a fourth stale cite in the schema manifest description. All four now name the real location. Spec review: the issue's Scope of changes named planning-config.md and git-planning-commit.md and neither was touched. Both now document the four-tier precedence and the new skip reason. Security review, minor and unproven: detectPhaseNumberFromFiles returns the FIRST matching path's phase, so a --files list spanning two phases resolves the override against whichever comes first. That helper is hardened and widely used, so it is not changed; the behavior is pinned by a named test and disclosed in the design and user docs. A pinned behavior is not a bug; an unpinned surprise is. * chore(#3587): backfill changeset pr number to 3601 --------- Co-authored-by: sim <sim@local> |
||
|
|
5f64d999dc |
fix(#3586): warn when .planning/ is gitignored but still tracked (#3598)
* feat(#3586): warn when .planning/ is gitignored but still tracked git ignore rules have no effect on files git already tracks, so a project that committed .planning/ before ignoring it keeps staging those files -- while commit_docs correctly resolves to false, which is exactly what makes the contradiction invisible. The probe lives in the SNAPSHOT BUILDER, not the rule: Rule.check may perform no ambient I/O (ADR-3180 8.1 rule 1, enforced by lint-planning-snapshot-bypass). buildPlanningTrackedField follows buildWorktreeHealthField's precedent -- injected execGit, bounded, degrading to UNREADABLE with a typed reason rather than throwing. W024 went inline instead only because no snapshot field carried its fact; that precondition does not apply here. W029 fires only on COMPLETE scope with ignored and tracked both true, so a degraded probe yields neither a finding nor a false all-clear, and the default project (tracked, not ignored) stays silent. The remedy is ADVISE-only -- --repair never untracks anything. * docs(#3586): document W029 and correct the health rule count CONFIGURATION.md documented the gitignore auto-detect without the caveat that ignore rules do not affect already-tracked files -- the very gap W029 exists to surface. Adds the caveat, the warning, its remedy, and why --repair will not act on it. CONTEXT.md's rule count was stale at 31 before this change (actual 32 through W028); corrected to 33 and pointed at the two other places the count is locked, so the next editor updates all three together. * fix(#3586): treat ls-files overflow as tracked, add CLI-level W029 tests Review findings. Security (minor, confirmed): execGit sets no maxBuffer, so Node's 1MB default applies to git ls-files. A .planning/ tree large enough to overflow it failed into git_list_failed and silenced W029 -- a false negative in exactly the large-history case most likely to have the real bug. Overflow is now treated as PROOF of tracking (the output was non-empty by definition) and resolves to tracked:true, scope COMPLETE, reason ok_truncated. Spec (major): test-matrix rows C1 and C2 were never implemented -- there was no CLI-level integration test at all, only rule-level ones. Both now drive the real validate-health dispatch and confirm W029 is reachable end-to-end. Known limit documented, not papered over: a deliberate git add -f under an otherwise-ignored .planning/ raises the same signal as the accidental case. There is no reliable way to tell them apart, the finding is advisory-only, and a heuristic that cannot actually distinguish them would be worse than the honest caveat. * test(#3586): update frozen health-doc counts and acknowledge health.md growth The remote matrix caught three gates that lint:ci does not cover. gen-health-docs.test.cjs froze a 35-row / 32-rule assertion; W029 makes it 36/33. Updated both the assertion and the test NAME, which embeds the counts -- a stale name is a lie even when the assertion passes. The second reported failure was the same assertion surfacing at describe-rollup granularity, not a distinct bug. emitted-attribution's growth arm needed an ack for the generated health.md. health.md was already named in 3309-health-docs-generated.json, and two ack sources naming one path is a hard error -- so a new fragment was not an option. That fragment's own history shows the pattern: #3309 created it, #2873 amended it in place for W028. Amended again for W029, with a note recording why this one file is amended rather than joined by a sibling. * docs(#3586): add the private-planning how-to and fix a wrong link docs/CONFIGURATION.md pointed 'Configure private planning' at how-to/configure-model-profiles.md -- an unrelated page -- and no private-planning how-to existed at all. Found while editing that section. The how-to test genuinely fires here: going private is four steps and crosses planning.search_gitignored, a setting owned by another concern, so a reference table structurally cannot carry it. The new page walks the whole sequence and leads with the step people miss -- .gitignore does not untrack what git already tracks -- which is the exact state W029 now detects. Also corrects 'artefacts' to 'artifacts' (repo house style is American). * chore(#3586): backfill changeset pr number to 3598 --------- Co-authored-by: sim <sim@local> |
||
|
|
311711754f |
docs(#3531): document where inherit must be set to reach tiered agents (#3550)
Co-authored-by: sim <sim@local> |
||
|
|
b7cca0363f |
fix(#3531): merge routing_tier_defaults over manifest tier defaults (#3539)
* test(#3531): failing-first suite for routing_tier_defaults manifest merge * fix(#3531): merge routing_tier_defaults over manifest tier defaults * docs(#3531): document routing_tier_defaults merge-over-built-ins semantics * fix(#3531): correct test helper scope, update folded #443 expectations, guard merge keys * test(#3531): pin tiers in effort-sync and surface-axis fixtures post-merge * chore(#3531): backfill changeset pr number * fix(#3531): correct rebase resolution — keep both 3531 and 3533 test blocks intact * test(#3531): pin inherit/effort fixtures to the layer that reaches tiered agents --------- Co-authored-by: sim <sim@local> |
||
|
|
adb2d03ed8 | Merge pull request #3540 from open-gsd/fix/3532-global-defaults-diagnostic | ||
|
|
50d5368add | fix(#3533): effort inherit — expressible, omitted at writers, never re-added (#3541) | ||
|
|
d26bfc2a3f | fix(#3534): resolve-execution reports resolved and effective effort | ||
|
|
129871a8be | fix(#3532): warn when global defaults keys are shadowed by a project config | ||
|
|
268ca7e32d |
fix(#3504): harden hook injection patterns and force-add guard (#3510)
* test(#3504): add failing-first parity, fail-closed, and bypass suites * fix(#3504): harden hook injection patterns and force-add guard * test(#3504): stage the scanner lib dependency in shared-hooks fixture * chore(#3504): backfill changeset pr number * test(#3504): build parity samples from fragments for the ci scan --------- Co-authored-by: sim <sim@local> |
||
|
|
7976b1ca0d |
feat(#1689): per-plan agent_hint executor routing (#3417)
* feat(#1689): per-plan agent_hint executor routing Option A per-plan specialist routing: a plan with an `agent_hint:` frontmatter field is dispatched to that subagent instead of gsd-executor when it resolves on the active runtime; absent/unresolved/disabled falls back to gsd-executor (byte-identical). Default-on via workflow.agent_hint_routing. - src/phase.cts: parse agent_hint into the plan-index JSON (plan_json.agent_hint) - agent-install-check.cts: resolveAgentHint() reuses getAgentsDir + runtime filename variants; probes project + global agent dirs; fails closed; rejects path-traversing names - gsd-tools.cjs: 'resolve-agent' query route (fail-closed to gsd-executor; --raw/--json) - execute-phase.md: lean per-plan reference + {EXECUTOR_TYPE} placeholder (host stays under the ADR-857 Phase 6 byte ceiling) - execute-phase/steps/per-plan-executor-routing.md: resolution logic (Agent()-based dispatch; advisory on orchestrator-worktree) - config: workflow.agent_hint_routing (validKey, default-on via SCHEMA_DEFAULTS, boolean validator) - docs (CONFIGURATION.md, plan-md.md), changeset, tests/agent-hint-routing-1689.test.cjs (17 tests) * chore(#1689): backfill changeset PR number (#3417) * chore(#1689): regenerate install-tree fixtures for new workflow fragment * chore(#1689): ack deliberate execute-phase.md growth (agent_hint routing) * test(#1689): SPAWN contract allows parameterized subagent_type placeholder agent-frontmatter's spawn-type checks scanned subagent_type="..." as a concrete agent name. execute-phase now uses subagent_type="{EXECUTOR_TYPE}" (a runtime placeholder resolved via resolve-agent, default gsd-executor). Skip {TOKEN} placeholders in both the known-type and <available_agent_types> checks; execute-phase still lists the built-in roster incl. gsd-executor. * fix(#1689): CI conformance for the routing fragment - per-plan-executor-routing.md: add the canonical runtime-launcher preamble to its gsd_run block (runtime-launcher-parity #373), matching sibling step fragments. - agent-install-check.cts: drop a literal ~/.claude/agents path from the resolveAgentHint JSDoc so it does not leak into the compiled engine .cjs (cline install leak guard). --------- Co-authored-by: sim <sim@local> |
||
|
|
f0abdb1b89 |
fix(#2486): do not recommend or persist Claude-only worktree isolation on non-Claude runtimes (#2531)
* fix(#2486): runtime-branch the settings worktrees question + W020 health diagnostic On non-Claude runtimes /gsd:settings offered "Yes (Recommended)" for worktree isolation and persisted workflow.use_worktrees: true — the exact value the execution workflows fail closed on (#1521 guards). Branch the question on the same stamped config-get runtime read the guards use: Claude keeps the unchanged question; non-Claude offers only "No (Recommended)" / "Leave unchanged", never persists true, and warns when the config carries an inherited explicit true. /gsd:health gains W020, surfacing such a config with the guards' own predicate before execution-time failure. Docs state the runtime-conditional default. Fixes #2486 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2486): add changeset for PR #2531 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2486): reassign the health worktrees check W020 -> W024 (verify.cts namespace collision) The workflow-level check collided with the live W020 (git-worktree-list health) emitted by cmdValidateHealth in src/verify.cts — invisible from health.md's error_codes table, which stops at W019 and under-represents the real namespace (W010-W017, W020-W023 all live). W024 verified free. Adds a regression test pinning the chosen code against src/verify.cts so a future assignment cannot silently collide, a table note naming the namespace owner, and the changeset body reworded to house style. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2486): pre-select the recommended repair in the broken-inheritance case Review round 2: at settings.md:142 the pre-selection rule left "Leave unchanged" as the default when the config carried an explicit non-false use_worktrees — the exact broken state the adjacent notice warns about, so accepting the default kept a config that fails closed at execution time. "Leave unchanged" is now the default only when the key is absent (nothing to repair); explicit false AND explicit non-false both pre-select "No (Recommended)", aligning the default, the label, and the notice. Pinned by two source-contract assertions in the #2486 regression block. Goldens (settings.md hash x19) + size baseline regenerated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2486): gate the worktrees question on dispatch.isolation, not the runtime name Review round 2: #2584 Phase 3 replaced the runtime-name test with a declared `dispatch.isolation` capability, invalidating this PR's premise. cursor declares harness-worktree and codex/opencode/kimi/ kimi-code declare orchestrator-worktree, so a `RUNTIME != claude` gate blocked a supported configuration on five runtimes and false-warned in health. - settings.md + health.md read `query dispatch-isolation` and branch on `ISOLATION = none`; the runtime-name read is gone from both, and the capability read needs no per-runtime stamping (it fail-closes unknown/ undocumented internally) - all "Claude Code-only primitive" prose rewritten, including the two gates the shell-syntax check missed (config-key list, JSON schema comment) - W024 reconciled across health.md + CONFIGURATION.md + planning-config.md (docs still said W020, which collides with a verify.cts code) - health.md error-codes table fixed: the namespace note no longer sits between rows orphaning I001 - the asymmetry note for the two workflows #2584 has not migrated (quick.md, diagnose-issues.md) is enforced by a set-equality test with a self-check table, so it cannot go stale in either direction Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(#2486): restore the Executor isolation section clobbered by #2661 `46ba02ac` (feat(#2630), the current next tip) reverted docs/CONFIGURATION.md to a pre-#2584 state: it restored the old "Non-Claude note" wording on the workflow.use_worktrees row and deleted the whole "Executor isolation per runtime" section. The change is unrelated to that PR's phase-estimation feature and looks like a stale-copy edit. This PR's use_worktrees row links to #executor-isolation-per-runtime, so the deletion leaves a dangling anchor. Restored byte-for-byte from |
||
|
|
2076d450d7 |
fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not the runtime name (#2728)
* fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not runtime name quick.md and diagnose-issues.md kept the pre-#2584 `RUNTIME != "claude"` worktree gate, so every non-Claude runtime failed closed regardless of the capability it negotiated — including Codex, which declares orchestrator-worktree. Route both through the negotiated dispatch.isolation seam via a new shared reference, and migrate the two execute-phase reference fragments that carried the same runtime-name gate. - new gsd-core/references/dispatch-isolation-gate.md: canonical ISOLATION resolution, harness-flag resolution, single-agent degrade rule - quick.md / diagnose-issues.md read the gate; dispatch uses the {harnessFlag} placeholder rather than a hardcoded isolation="worktree" - execute-phase-wave-guard.md / execute-phase-between-wave-reset.md: migrate [ "$RUNTIME" = "claude" ] -> [ "$ISOLATION" = "harness-worktree" ] - every degrade site now clears BOTH USE_WORKTREES and ISOLATION; clearing one dispatched an isolated agent with no base guard and no manifest - parity guard in host-integration.test.cjs scans workflows AND references and matches six reintroduction shapes - migrate four tests that pinned the pre-#2584 runtime-name contract Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): use the /gsd:<cmd> namespace in the isolation degrade messages The degrade warnings cited /gsd-execute-phase, the retired hyphen form that slash-command-namespace.test.cjs rejects in Claude-facing source. Same length, so the quick.md size budget is unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2652): add changeset for PR #2728 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): normalize dispatch-site paths to forward slashes for Windows path.relative() returns backslash-separated paths on Windows, so the #2652 dispatch-site parity test compared "gsd-core\workflows\quick.md" against the hardcoded forward-slash literal "gsd-core/workflows/quick.md" and failed on every windows-latest CI lane. Normalize with .replace(/\\/g, '/'), matching the existing convention used elsewhere in this suite (e.g. tests/branch-no-track-guard.test.cjs:37). * test(#2652): restore the size-growth acknowledgment The rebase dropped tests/emitted-drift-ack.json. #2757/#2758 fixed the ATTRIBUTION axis, but the SIZE-GROWTH axis is independent: diagnose-issues.md (+2086) and quick.md (+230) still need an ack naming them and saying why. Verified: 65/66 without it (both files named), 66/66 with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): convert execute-plan.md Pattern A onto the dispatch-isolation gate Pattern A hardcoded `isolation="worktree"` — Claude Code's own literal — gated only on `workflow.use_worktrees`, with no capability negotiation at all. It is the same defect #2652 fixes at the other four sites, just a different shape: the file contains no RUNTIME variable, so the new detector correctly does not flag it. Concrete break: a Codex user who follows this PR's own newly-documented pattern and sets `workflow.use_worktrees: true` to get isolated dispatch via /gsd:quick then runs a plan through /gsd-execute-plan Pattern A, and hits an unconverted path — either an Agent() call erroring on an unrecognized parameter or silent unisolated execution, depending on host tolerance. Pattern A is a single-agent dispatch site through the host's own subagent tool, so it takes the same treatment as quick.md and diagnose-issues.md: resolve ISOLATION/HARNESS_FLAG through the canonical reference, degrade to sequential on orchestrator-worktree hosts, and substitute the host's declared {harnessFlag} instead of Claude Code's literal. while the area was open. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2652): add the INVENTORY row for dispatch-isolation-gate.md, refresh CONTEXT Two bookkeeping gaps flagged in review: INVENTORY.md had no row for the new gsd-core/references/dispatch-isolation-gate.md. INVENTORY-MANIFEST.json was regenerated correctly and its --check only diffs a live directory scan against the committed manifest, so CI passed regardless — but gen-inventory-manifest.cjs's own stderr guidance says to add the matching INVENTORY.md row. This is the repo's named "Inventory Drift" pattern. Placed with the dispatch/isolation cluster (worktree-branch-check, runtime-aware-dispatch) rather than alphabetically, matching how that table is grouped. CONTEXT.md's Host-Integration Interface entry still described dispatch.isolation as "declared and negotiated but not yet consumed by any scheduler — Phase 1 of #2584". That was already stale before this PR (execute-phase graduated in Phase 3) and more so now with three single-agent dispatch sites consuming it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): detect reversed-operand runtime gates; add a permutation property All five reintroduction regexes assumed $RUNTIME on the LEFT of the comparison, so `[ "claude" != "$RUNTIME" ]` — the same gate written backwards — evaded every one of them. Verified against the old patterns before fixing: all four reversed shapes (single bracket, double bracket, test builtin, JS template) scored EVADED. Each comparison shape is now generated in both operand orders from a single template, so a shape cannot be added in one order and forgotten in the other. The mutation table gains the four reversed cases. Also adds the fast-check property review suggested in place of the hand-rolled cases: it generates the cross product of the axes an author actually varies — bracket form, operator, operand order, quoting, spacing, runtime id — so a permutation the hand-written patterns miss surfaces here rather than in production. The 11 explicit cases stay as named regression anchors. execute-plan.md joins the scan's required-identities list now that it is a converted dispatch site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): acknowledge the execute-plan.md size growth The Pattern A conversion adds 811 bytes to an emitted workflow. Per #2719 the size axis needs its own acknowledgment, independent of attribution. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): repin the execute-plan.md PROSE_ALLOWLIST line after the rebase The #2751 command-position gate pins its prose exemptions by line number. This branch inserts the dispatch-isolation resolution above the `validated downstream by gsd-tools uat classify-coverage` sentence, moving it from execute-plan.md:387 to :397 — which fired the gate twice for one displacement (an un-allowlisted mention at 397, a stale entry at 387). The prose itself is unchanged from next; only the pin moves. Fixes #2652 * fix(#2652): gate the #2649 base-check on ISOLATION in diagnose-issues.md The rebase onto next merged #2649's pre-dispatch base-check textually, but its degrade flipped USE_WORKTREES after ISOLATION was already resolved, so the degrade never reached the dispatch decision. Gate the block on ISOLATION = "harness-worktree" and degrade ISOLATION itself, the same pairing quick.md already uses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): key quick.md post-dispatch bookkeeping on ISOLATION, not the Claude literal Review Blocker: the manifest append (l.822), worktree merge-back (l.825), and its skip clause (l.839) all conditioned on the literal isolation="worktree" — Claude Code's own rendering of {harnessFlag}. Cursor renders --worktree, so a newly-unblocked isolated Cursor run created a worktree whose committed work was never merged back and never cleaned up, silently. All three now key on ISOLATION = "harness-worktree" at dispatch. The existing parity detector cannot catch this class (its ISOLATION_TOKEN treats the literal as a legitimate marker), so this adds a dedicated literal-condition detector with a discrimination proof against both pre-fix sentences, a benign-mention control, and a positive pin on all three re-keyed conditions. Verified fail-first against the pre-fix quick.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): scope the use_worktrees=false install stamp to isolation=none runtimes `_stampNonClaudeRuntimeDefaults` rewrote every non-Claude runtime's `workflow.use_worktrees` read to `--default false`. That default resolved before `gsd_run query dispatch-isolation` was ever consulted, so the five runtimes that declare worktree support — cursor (harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree) — got ISOLATION=none regardless of what they negotiated. The gate this PR migrates dispatch onto was therefore still deciding isolation by runtime name, one layer down. The stamp's #1521 premise was that worktree isolation *was* Claude Code's isolation="worktree" spawn parameter, which no other host honored. #2584 replaced that premise with the negotiated capability. The stamp is now scoped to runtimes whose negotiated isolation really is `none`, where the default it writes is the outcome the resolver reaches anyway. `_negotiatedDispatchIsolation` mirrors routeDispatchIsolation's resolution against the same registry — closed vocabulary, a harness-worktree host must declare its flag, an orchestrator-worktree host must carry a descriptor that resolves — and fails closed to `none` on anything else, so an undeclared or unknown runtime keeps today's behavior. Two #1515 tests pinned the superseded premise for codex and are re-pointed at the new contract rather than deleted: the safety property they protect is now held by the isolation gate's fail-closed resolution, not by a name-scoped install-time default. Verified fail-first — all five assertions red against the pre-fix source, green after. * test(#2652): acknowledge the emitted ripple and re-point the end-to-end stamp proof Scoping the use_worktrees stamp changes emitted output, and two gates caught it. `gsd-core/workflows/execute-phase.md` now differs at emit time for the five hosts that declare worktree support (cursor harness-worktree; codex, opencode, kimi, kimi-code orchestrator-worktree) — the source file is byte-identical, only the stamp is gone. Acknowledged in this PR's fragment. `tests/install.test.cjs`'s real-install assertion pinned the superseded premise end-to-end, asserting codex receives `--default false`. Re-pointed rather than deleted, matching the two unit tests: it now proves codex keeps the unstamped `true` read. A second arm installs windsurf — which declares isolation `none` — and asserts the false stamp is still applied there, so the change cannot silently degrade into "never stamp" without a test noticing. The ack entry collides with `2658-trae-instruction-file-path.json`, which is fully spent (merged via #2925, so all 25 of its entries are present at base and gate nothing) and is pruned for the same reason and by the same rule as the spent `2649-*` fragment this PR already removed. #2566 prunes the same file for the same collision on `new-project.md`; a delete/delete merges cleanly either way, and the base-side cleanup would make both unnecessary. * fix(#2652): re-record the sentinel when a dispatch site degrades isolation Review Blocker B1/B2/B3. Every isolation degrade in a dispatch site is decided in shell, where routeDispatchIsolation cannot see it. That resolver persists whatever it resolved to the run-scoped sentinel as an unconditional side effect (#3045), so a degrade that only reassigns $ISOLATION leaves the sentinel asserting harness-worktree while the dispatch correctly omits the harness flag. The shipped PreToolUse guard reads the sentinel at the instant of the Agent() call and denies that mismatch with exit 2 — the work does not run unisolated, it does not run at all. Latent on this branch and lands on rebase, since |
||
|
|
96a82bbffb |
enhance(#3245): report the detected host runtime in init (#3307)
* test(#3245): failing-first coverage for host runtime detection in init Locks the behavior epic #2313 Phase 5 must produce before any of it exists: init reports the detected host, explicit GSD_RUNTIME and config runtime still outrank detection, non-Codex sessions are untouched, and nothing is ever written to shared defaults (#2297). * enhance(#3245): report the detected host runtime in init init reported agent_runtime: claude inside a Codex session, and resolved agents_dir to the Claude agents root with agents_installed: true — a spuriously healthy triple. Runtime identity was only ever read from GSD_RUNTIME or an explicit runtime in .planning/config.json. Adds a detection rung beneath both explicit sources, in a new pure module. Codex is identified from its own documented session environment (CODEX_SANDBOX / CODEX_SANDBOX_NETWORK_DISABLED), else an explicitly exported CODEX_HOME whose config.toml exists. The default ~/.codex is never probed: that file exists on every machine that has run Codex, so probing it would misreport other runtimes' sessions. resolveRuntime keeps its exact contract and all 71 dependents, including formatGsdSlash command-style emission; only withProjectRoot consumes the new rung. Nothing is written on any path (#2297). Explicit config still wins (#2517). * fix(#3245): make the parity guard real and single-source the marker Four independent review passes found the generative-fix-divergence guard was vacuous: it asserted agreement at the one input where inferPreferredRuntime and detectHostRuntime do not differ, so it could not fail. It now pins the actual divergence point (CODEX_HOME set, config.toml absent) and records that the asymmetry is deliberate. The config.toml marker is now single-sourced from update-context.cts and imported, rather than carried independently by two surfaces. tests/helpers.cjs now scrubs CODEX_SANDBOX and CODEX_SANDBOX_NETWORK_DISABLED: GSD reads them, so an ambient Codex session would otherwise make the non-codex control test fail non-deterministically. Also: detection is throw-safe end to end rather than only around the fs probe; the Windows-join test is replaced with one that can actually fail (trailing-separator, catches hand-rolled concatenation); the #2297 no-write proof now wraps resolveReportedRuntime, the function that ships, across all three ladder outcomes. * chore(#3245): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
2dbee3ebdd |
enhance(#2229): add three-way claim disposition (admit/refute/abstain) to /gsd-explore research pass (#2543)
Closes #2229. Each claim surfaced by /gsd-explore's research pass is dispositioned admit, refute, or abstain, with abstentions routed to a visible ledger instead of being smoothed into confident prose. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim; a strong prior is never authoritative alone. Two guards ride with it: conflict-abstention, and a tier floor that presents a would-be admit as an abstain when the researcher's resolved tier is the budget tier or cannot be determined. To make that floor enforceable, resolve-model now emits the effective tier (--pick tier). It was already computed above the resolve_model_ids omit gate but was unreachable from a workflow, which left the floor inert on every non-Claude install - the model id is blank under omit and runtime-substituted where a tier map exists, and the profile defaults to balanced. The tier signal mirrors every resolution step that can change which tier runs, including the model_policy preset, and reports unknown rather than guessing. Output is additive; model, profile and effort are unchanged. Two residuals are disclosed in the workflow rather than papered over: a raw-model-id model_overrides pin reports unknown and is floored (fails closed), and a model_profile_overrides entry repointing a tier at another tier's model can under-report (fails open, and predates this change). Admin merge used only to satisfy the missing secondary reviewer on a single-maintainer PR. No CI failure and no conflict were bypassed: 38 checks green, remote runner 32255/32255 on both Node lanes. |
||
|
|
b901d1e06f |
feat(#1953): complexity-triggered refactor extension point (execute:post) (#3261)
* test(#1953): failing-first suite for the complexity-triggered refactor hook 60 behavioral cases against src/complexity-trigger.cts, which does not exist yet: decision-point counting, the comment/literal stripping leak surface, threshold and jump-delta boundaries at limit-1/limit/limit+1, stable-anchor baseline semantics, and fs fault injection via mock.method. Two fast-check properties assert that stripping never manufactures a decision point and that comments and string literals are score-neutral. Also registers the refactor-trigger capability manifest (inert until refactor.trigger_enabled) and regenerates the capability registry and matrix. Verified RED on the remote runner before any implementation exists. * feat(#1953): complexity-triggered refactor extension point Adds the opt-in refactor-trigger capability. After a phase executes, an execute:post step measures per-function complexity for the files the phase touched and writes a scoped refactor proposal when a function crosses the configured threshold or drifts past its recorded anchor. Design notes worth carrying: - The signal is computed in-core (decision-point counting over comment- and literal-stripped source, Node builtins only) rather than via Memtrace or a shelled-out analyzer. The hook fires as a deterministic CLI, not an agent with MCP tools, and core takes no external dependencies — this is the only option a behavioral test can bind to. The metric sits behind a seam. - The baseline is a stable anchor, not a rolling value: set on first observation, moved only on disposition. A rolling baseline makes the delta the single-phase change, so a function creeping +2 per phase never trips a delta of 5 and the jump check adds nothing over the absolute threshold. - Strict mode records an open deviation window in the broken-windows ledger rather than declaring its own ship:pre gate. ship.md has no generic ship:pre gate dispatch — only two hardcoded branches — so a third gate of any kind would be declared and never evaluated. - The gate clears on the proposal being dispositioned, never on the score improving. A blocking complexity number is one an executor can satisfy by splitting a coherent function in two. execute-phase.md gains a generic execute:post step-dispatch contract; it previously matched only ref.skill == "code-review", so any other step registered there was declared and never run. The code-review branch is unchanged. Full rationale in ADR-1953. Closes #1953 * fix(#1953): close git option injection and symlink escape in the refactor hook Three findings from the isolated security review, all fixed inline. HIGH — changedFilesSince interpolated the --since value into a revision token placed before the -- separator. A -- only stops PATHSPEC parsing of arguments after it; git still option-parses what comes before. So --since '--output=/tmp/x' became --output=/tmp/x..HEAD, which git accepts as --output=<file> and uses to redirect diff output — an arbitrary write. Fixed with --end-of-options before the revision range plus a conservative ref validator. The validator deliberately permits ~ ^ @ { } because those are legitimate git REVISION syntax (HEAD~1, main@{yesterday}) as distinct from ref-NAME syntax; --end-of-options is the actual barrier. The doc comment asserting the trailing -- was sufficient was wrong and is corrected. MEDIUM — resolveConfinedPath confined by string prefix only, so a symlink committed inside the repo passed the check (its own path is under cwd) and readFileSync then followed it outside the root. Now lstat-checks for a regular file and skips anything else with REFACTOR_FILE_UNREADABLE, so one bad path skips one file and the run continues. LOW — the new execute:post dispatch contract showed the gsd_run example before the rule requiring ref.command be validated first. That prose is executed by an agent, so textual order is execution order. Reordered. Refs #1953 * fix(#1953): make the analyzer able to see TypeScript at all Found by running the shipped analyzer over its own source: it reported functions=1 for a 940-line module with 24 function forms. A return-type annotation or a generic parameter list made a function invisible — `function f(a): number {}` and `function f<T>(a: T): T {}` both detected as zero. Since gsd-core is written in .cts and the capability declares .ts/.cts/.mts analyzable, the feature silently found nothing in this repo's own primary language while reporting success. A safety net that reports "all clear" because it cannot see is worse than no safety net. All 98 tests passed over this, because every fixture was plain JS — the exact failure the test matrix's own "assert against the shape production uses" warning describes. Adds a TypeScript-shapes suite covering return types (including unions, generics, object literals and type predicates), generic parameter lists (constrained and defaulted), export/async/generator combinations, annotated arrows, class-method modifiers, and optional/ default/rest params — plus the two traps: an overload signature has no body and must not count, and `a < b && c > d` is a comparison, not a generic. Detection now reports 24/37/21 functions for the three source files, which matches a hand count exactly. Also from review: - The strict-mode ledger dedup identified entries by parsing a prose description string. That is banned by CONTRIBUTING's raw-text-matching rule and was a real bug: the "exactly one window per untriaged proposal" guarantee rested on prose matching, so rewording a description or editing WINDOWS.md by hand silently produced duplicates. Now matches structurally on kind + phase + file + line. - A property test asserted on the stripper's output text. Reframed to assert the same invariant through analyzeSource's score. - nextBaseline's `candidates` parameter has been dead since the anchor change; removed from the signature and all call sites. - Extracted the duplicated require-or-degrade and capability-check boilerplate. - ADR-1953's Implementation bullet still named a `refactor.ship-gate` in check-command-router.cts — a leftover from the design cut D6 rejects. That file is untouched and no such gate exists. Removed. Refs #1953 * fix(#1953): keep execute-phase.md under its byte ceiling; un-vacuum the large-file test Five of the seven remote-runner failures were one cause: the execute:post dispatch contract, written out inline, grew execute-phase.md 1876 bytes (93,400 -> 95,276) against a frozen PRE_PHASE6 ceiling of 93,600. A drift-ack does not clear that — tests/phase6-capstone-conformance.test.cjs and tests/fix-2285-claude-orchestration-wiring.test.cjs assert the file is literally under the cap. The contract now lives in gsd-core/references/loop-hook-dispatch.md, which already claimed to be the point-agnostic dispatch reference and already documented ref.skill and ref.agent. It gains the ref.command shape, its in-context validation rule, the advisory-by-construction statement, and a note that a point whose workflow hand-rolls one kind is not implementing this contract. execute-phase.md now defers to it in one line: 145 bytes of growth, 55 B of headroom under the cap. Better placement than the first cut — the reference was overstating its coverage, and this makes the claim true rather than duplicating prose next to it. Acknowledged by appending to tests/emitted-drift-acks/2930-*.json rather than a new 1953-*.json: two ack sources may never name the same path, and that fragment is already the accumulating ack for this file. Sixth and seventh failures: analyzesLargeFileWithinBounds tripped its own vacuity guard — the fixture generated ~480 KB against a `> 500000` assert, so the guard fired and the three assertions after it never ran. The test has been vacuous since it was written. The matrix row specifies ~1 MB, so N goes 8000 -> 20000 (1.17 MB, 17% margin) and the guard to > 1_000_000. Verified by reproducing the exact body against the compiled module: 1168888 bytes, 118 ms, all four assertions hold. Refs #1953 * fix(#1953): fold the execute:post step deferral into the existing resolve line The remaining two failures were one test: execute-phase.md carries a SECOND, tighter assertion than the 93,600 ceiling — `<=93400`, which is exactly its current size. The file cannot grow by a single byte. My previous fix got it under 93,600 but not under 93,400, so it still failed. ("H." in the report is just the parent describe of that same test, not a separate defect.) Rather than add a paragraph, the deferral now REPLACES the existing hook resolution line. It read: Resolve active step hooks from `EXECUTE_POST_HOOKS_JSON` where `kind == "step"` and `ref.skill == "code-review"`. which is the bug itself written down — only code-review was ever dispatched. It now reads: Dispatch each `kind == "step"` hook per @gsd-core/references/loop-hook-dispatch.md. For `code-review`: The following prose already begins "If no active code-review step hook exists", so it reads correctly and the code-review handling is untouched. Net effect on the file is -11 bytes: 93,400 -> 93,389, under the margin assertion rather than merely under the ceiling. That also removes the need for a drift-ack: the file shrank, so there is no growth to acknowledge, and the append to the shared 2930-*.json fragment is reverted. Leaving it would have shipped a claim of "145 bytes of growth" that is no longer true, on a file six other issues share. The test's own comment states the principle this ended up honoring: "the host loop must stay small — optional-feature detail belongs in the capability fragment, not the host workflow." Putting the dispatch contract in the reference rather than inline is that rule, applied. Refs #1953 * fix(#1953): keep the code-review hook literal the workflow test requires tests/code-review.test.cjs extracts the <step name="code_review_gate"> block and asserts it contains `ref.skill == "code-review"` verbatim. The previous commit replaced the line carrying that literal, so the token vanished and the test went red — a fair assertion: code-review IS the bespoke branch there and the workflow should still name it. Restored inside the same one-line deferral, which now reads: Dispatch `kind == "step"` hooks per @gsd-core/references/loop-hook-dispatch.md. `ref.skill == "code-review"`: 93,396 bytes — still under the `<=93400` margin assertion and 4 bytes below the base, so the file continues to shrink rather than grow. Because three consecutive runs were each reddened by a different assertion on this one file, this change was verified by sweeping ALL of them at once rather than one run at a time: every test under tests/ that reads execute-phase.md or references/loop-hook-dispatch.md was located by resolving its path constants, and each content/size assertion was evaluated directly against the working tree — 22 assertions, plus two real executions (gen-section-manifest --check, and emitted-attribution's full real-tree differential). All pass. That sweep also confirms the earlier judgement call: the net change to execute-phase.md is a SHRINK, and the size ratchet only gates growth, so reverting the append to the shared 2930-*.json ack fragment was correct — an ack would have been both unnecessary and factually wrong. Refs #1953 * chore(#1953): backfill changeset pr number to 3261 * docs(#1953): add the missing how-to for acting on a refactor proposal Reference and explanation shipped (COMMANDS.md, CONFIGURATION.md, FEATURES.md 159, ADR-1953) but the Diataxis how-to quadrant did not, and that is the one a user reaches for. CONTRIBUTING's required-docs table is 'new command -> COMMANDS.md + FEATURES.md', so CI was green on a gap. Enabling this feature is genuinely multi-step and no single page walked it: turn it on, tune the threshold, understand advisory vs strict, discover that strict needs a SECOND toggle on a DIFFERENT capability, and know what to do when a proposal appears. The two-toggle subtlety in particular was a footnote in a config table; here it is a section with both commands. Follows the shape of its closest siblings, resolve-edge-coverage-findings and resolve-prohibition-findings — both 'the loop surfaced a finding, here is what to do with it'. Includes a reason-code table for the silent cases, since the analyzer is deliberately quiet in six situations and a user who expected a proposal needs to tell 'nothing to report' from 'could not look'. Indexed from docs/README.md beside the other loop how-tos. Docs-only: exempt from the push gate, no re-verification, pass marker on 2af188b4 untouched. Refs #1953 * feat(#1953): warn when strict mode is on but nothing will actually block Closes acceptance criterion 5, which I had wrongly marked satisfied. refactor.trigger_strict records an untriaged proposal as an open deviation window, but a ship only STOPS if workflow.windows_enforce is also on — a toggle owned by the broken-windows capability that this feature neither sets nor requires. So a user could enable strict, believe ship was gated, and find out otherwise at ship time. The split itself stays: requires:["broken-windows"] would force-install the ledger on advisory users who never enable strict, and a ship:pre gate of our own would never fire because ship.md has no generic ship:pre gate dispatch. What was missing was discoverability, so that is what this fixes. `refactor evaluate` now emits a typed REFACTOR_STRICT_NOT_ENFORCING warning, naming the exact remediation command, whenever strict is on and either workflow.windows_enforce is off or broken-windows is unavailable. It fires only on a run that produced a candidate — with nothing to block on there is nothing to warn about, and warning every run would be noise. Reads workflow.windows_enforce through the same resolveConfigKey walk the router already uses for its own keys rather than a second config reader. Four tests cover the matrix: strict+enforce-off warns, strict+enforce-on does not, strict+ledger-absent warns, strict-off never warns. Also corrects a user-facing message in this same file that told the user to run `gsd-tools config-set` — the wrong form. docs/CONFIGURATION.md and the broken-windows capability both use `gsd config-set`, and gsd-tools is invoked as `node gsd-tools.cjs`, so the bare form may not resolve. The two adjacent messages in this file now agree. Refs #1953 --------- Co-authored-by: sim <sim@local> |
||
|
|
58d73dd220 |
enhance(#3241): omit the codex per-agent model by default (#3276)
* test(#3241): failing-first suite for the codex passive model posture Locks ADR-2313's D1-D5 before any production code exists, so the tests bind to the behavior rather than to whatever the implementation happens to do. Red-first (fail against the current tree): - the resolver path emits no `model` and no `model_reasoning_effort` - a whitespace-only model_overrides value yields no pin - isAnthropicFlavoredModel / CLAUDE_AGENT_ALIASES on model-catalog - the one-time install notice, and its once-per-install dedupe Regression guards (pass today, must keep passing): resolver-null via `inherit` and via absent runtime; a resolver that resolves to nothing; empty-string and non-string overrides; and the light-tier service_tier/model_verbosity fields, which are NOT coupled to the model pin and would silently regress if the implementation coupled them. Classifying each test as red-first or regression guard is deliberate. A test that passes on both sides of the change proves nothing, and this epic has already shipped two such rows before catching them. The whitespace case is a live defect, not a quirk: `' '` is truthy, survives the type guard, is not Anthropic-flavored, and is embedded verbatim as `model = " "` — the same class the #2310 guard exists to stop. Same function, same path, fixed in this phase per CLAUDE.md §3. Two matrix rows were dropped as vacuous rather than shipped green: a 64-char truncation case (the pinned notice interpolates no user-controlled value, so it cannot exhibit truncation) and a newline hazard that the input surface cannot reach. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3241): omit the codex per-agent model by default Implements ADR-2313 D1-D5. generateCodexAgentToml no longer embeds the runtime resolver's per-tier Codex model, so an agent inherits the always-available session model instead of a pin a ChatGPT-account Codex may not expose. model_reasoning_effort disappears with it via the existing hasPinnedModel coupling (#838) — no logic change needed there. Supersedes #2517's embedding on the default path only. An explicit real-Codex model_overrides pin is still embedded verbatim, and the #2310 Anthropic-flavored guard is retained: the model_overrides route to it is still live even though the resolver route is now unreachable. The shared rule moves down a layer. CLAUDE_AGENT_ALIASES leaves model-resolver for model-catalog — a genuine leaf importing only node:path and its own JSON — with isAnthropicFlavoredModel defined beside it, and is re-exported from model-resolver so every existing importer is untouched. This is what lets Phase 2's install-check and Phase 3's sync consume the rule without taking the config-loader dependency model-resolver would have dragged into a module documented as pure read/verify with 33 dependents. A parity test fails if the two ever fork. Also fixes a live defect surfaced while writing the tests: a whitespace-only model_overrides value was truthy, survived the type guard, was not Anthropic-flavored, and so was embedded verbatim as `model = " "` — the same class the #2310 guard exists to stop, reached by a different route. Trimmed before the truthiness test. It is deliberately not routed through _warnCodexModelOverrideDropped, whose text would misdescribe a blank field as a mis-typed model. Adds the one-time install notice (maintainer direction, recorded as an ADR-2313 amendment): one stderr line naming model_overrides and the session model, deduped per install rather than per agent, and emitted only for the population that actually loses a pin. service_tier and model_verbosity stay decoupled from the model (#774); a regression guard asserts they still emit with nothing pinned. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3241): amend ADR-2313, add the model-catalog glossary entry ADR-2313 gains two dated amendments rather than edits to its merged text, since ADRs here are append-only. The first records that a deprecation notice IS offered, reversing the Migration section's "no deprecation window" position, and states why that position was wrong rather than just superseding it: the ADR identified the API-key population as losing something real and then declined to warn it, in the same document. Hyrum's guidance was applied to the recourse and not to the notice. The second records the whitespace-only model_overrides defect and notes that D2 always implied the fix — the implementation simply never enforced it and no test covered the case. CONTEXT.md gains a Model Catalog Module entry. The module had none, which is why the glossary gate passed without one: check-glossary-refs verifies that references resolve, not that modules are documented. The entry records why the Anthropic-flavored rule lives there rather than in model-resolver, so a later reader does not "helpfully" move it back. The Model Resolver entry is updated to point at its new home and note the back-compat re-export. docs/CONFIGURATION.md carried a claim that is now false: that the resolved tier ID is embedded into agent frontmatter at install time on codex and opencode. Corrected to name codex as the exception, with the 400 symptom and the model_overrides recourse. Changeset leads with the user-visible change and the migration line rather than the implementation, per the ADR's Hyrum's-Law analysis. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3241): only notice a lost pin when one was actually embeddable Review finding from an isolated reviewer. The deprecation notice gated on whether the runtime resolver would have returned *any* model, but the question that matters is whether that model would have been *embedded*. Those differ. The #2310 safety gate already rejected an Anthropic- flavored model arriving from the resolver path before Phase 1 — so for a mixed-runtime config resolving to a claude-* id against a Codex install target, the user never had that pin. The notice told them they lost something they never got, and pointed them at model_overrides for no reason. The existing #2310 test drives exactly that path but asserts only the emitted `model` line, never stderr, which is why it slipped through. Now covered. Deliberately unchanged: an Anthropic-flavored model_overrides value plus a legal resolver model fires BOTH the override warning and the notice. That is correct — pre-Phase-1 the guard dropped the override, execution fell through to the resolver, and the resolver's model was embedded, so that user did lose a pin. Two messages, two distinct true facts, and the prefixes differ (`gsd: warning — ` vs `gsd: notice — `) so the one-notice-per-install contract holds. A regression test now pins that behavior so it does not get "simplified" away. Of the three tests added, only the first is red-first; the other two pass on both sides by design and are labelled as guards — one against over-correcting the fix into silence, one against removing the intentional double message. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3241): reset the notice dedupe via a seam, not a require.cache bust The remote runner caught a regression I introduced: the #2760 post-write-validation test began failing with the validator override no longer intercepting. Cause, confirmed by trace rather than guessed: the new #3241 review tests deleted require.cache for bin/install.js and re-required it mid suite, to clear the notice's module-level dedupe flag. But runCodexInstall destructures `install` at file load, closing over the ORIGINAL module's exports. After the cache bust a second instance existed, so the test's `installModule.__codexSchemaValidator = ...` mutated the new object while the code under test still called the old one. The override silently stopped intercepting, the real validator ran and passed on GSD-emitted output, and the abort-and-restore path was never exercised. Cache-busting a module mid-suite breaks every later test that assumes a single instance, which every other test in the file is entitled to. So the fix is a seam, not a workaround: bin/install.js exports _resetCodexNoticeDedupeForTests(), and the three tests call it directly instead of reloading the module. The flag is module-level by design — the dedupe is per-install and install() already resets it — so a unit test driving generateCodexAgentToml directly needs an explicit way to reset it. That is now what it has. Swept the rest of the #3241 diff for the same hazard; this flag was the only shared module-level state introduced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3241): reset both codex dedupe stores, not just the notice flag Second incomplete fix, same class one layer down. bin/install.js keeps TWO module-level dedupe stores and the require.cache bust I removed had been papering over both; my replacement seam cleared only one. _codexModelOverrideDroppedWarned is a Set keyed `${agent}::${value}`. tests/codex-config.test.cjs:558 already emits for `gsd-executor::sonnet`, so by the time the review test using the same agent and value ran, _warnCodexModelOverrideDropped was a silent no-op and the expected warning never appeared. The seam now clears both stores and is renamed to say so. Its comment records that per-install dedupe lives in module scope deliberately and that this is the single sanctioned way for a unit test to clear it. Swept bin/install.js for every other module-scope mutable a test could latch. Two are inert (capability registries assigned once at require time; selectedRuntimes computed once from argv). One is a genuine latent hazard and is deliberately NOT folded in: attributionCache (:1654) memoizes getCommitAttribution by runtime name for process lifetime, so two in-process installs of one runtime with differing attribution config would collide. It is unreachable from any current test and is a different concern from Codex warning dedupe, so it stays out of this PR rather than widening it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3241): correct the codex tier-routing how-to The docs gate forced the task-oriented quadrant and found the worst defect in this change's documentation surface. docs/how-to/configure-model-profiles.md carried a section titled "If you want tiered models on Codex" telling users to set runtime:codex + model_profile:balanced, promising "GSD resolves each tier alias to the Codex-native model and reasoning effort defined in the runtime tier map." That is exactly the behavior this PR removes — a how-to page confidently instructing users to do something that no longer works, which is worse than a missing page because it fails at the moment of use. Rewritten to state that Codex does no tier routing, give the model_overrides pin as the supported alternative, and name the two constraints on what may be pinned: it must be a real Codex model id, and the account must actually expose it — GSD cannot verify the second, so the honest advice when unsure is to omit the pin. Carries an upgrade note for both account types, since the change is a no-op for ChatGPT accounts and a real loss for API-key ones. Also tightened the same page's claim that Codex "embeds the resolved model" at install time — now true only of an explicit override. The re-install instruction it supports is still correct and still needed, so only the premise moved. Both the required-docs set (COMMANDS.md + FEATURES.md) and lint-docs-required.cjs would have passed before this commit, since CONFIGURATION.md and the ADR had already moved. Neither checks the quadrant a user in trouble actually opens. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3241): backfill changeset pr number (#3276) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b7431a9259 |
feat(#1956): flag cross-artifact fact drift in the plan drift guard (#3259)
* test(#1956): failing-first contract for cross-artifact fact-drift pass * feat(#1956): flag cross-artifact fact drift in the plan drift guard * fix(#1956): correct config-key assertion and bidirectional lifecycle-lag exemption * docs(#1956): document the cross-artifact axis in the architecture reference * feat(#1956): decide the phase-status drift axis deterministically * fix(#1956): scope the progress-table lookup, abstain without a position section, rank deferred * docs(#1956): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
ffd5370464 |
fix(#2903): use the command form that actually works in reader-facing docs (#3047)
* fix(#2903): use the command form that actually works in reader-facing docs Docs told readers to type the colon form, which no runtime registers -- 18 of 19 runtimes use slash-hyphen and the 19th uses shell-var -- so anyone copying an example got an unrecognized command. Swept 178 occurrences across 53 files, locale mirrors included so they do not re-diverge from English. The colon form is a source-authoring token, not a user-facing one: install-time converters key on it to produce the hyphen form runtimes actually register. So the sweep is scoped, and three things are deliberately left alone: - ADRs, which are a historical record; editing their prose falsifies what was written at the time. - The legacy release-notes archive, pending a maintainer decision on whether it follows the same historical carve-out. Excluding it keeps a later reversal additive rather than a revert. - Source artifacts under commands, workflows and agents, where the colon form is load-bearing. Rewriting those would break the installed-skill guarantee across every runtime -- the single largest hazard here. The plugin namespace form is a real, separate token and survives untouched. Adds a lint enforcing exactly that boundary, since the correct form genuinely differs by directory and nothing previously caught the drift. Also fixes a hardcoded colon form in the capability-matrix generator. The sweep alone would have left the generated matrix disagreeing with the template that produces it, so the fix is at the source and the output regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2903): stop the sweep misquoting source frontmatter Adversarial review caught three lines where the sweep rewrote a citation of the literal YAML name: key from a source command file. That key genuinely is the colon form -- this change's own carve-out logic says source-authoring tokens keep it -- so the docs ended up misquoting the real files. One of the three is an acceptance-checklist assertion, which the sweep turned into a false statement. Restored the three citations to match their sources verbatim, surgically: where a line carried both a name: citation and a real reader-facing slash command, only the citation reverted and the command stayed corrected. The guard needed the same distinction, or it would have flagged the restoration and reddened the build: a gsd:<cmd> token preceded by name: is a citation of a source token and is now permitted. The exemption is deliberately narrow -- a bare gsd:<cmd> anywhere else still fails -- with a test pinning that narrowness. Also makes the detection case-insensitive. Review found /GSD:next slipped through silently; no such casing exists in the tree today, so this closes a latent gap rather than fixing a live one. Swept the whole tree for further corrupted citations: none beyond the three. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2903): retire the stale-next invariant and sweep next like every other command Maintainer decision on a genuine conflict between two contracts. Invariant #3054 banned the literal /gsd-next from user-facing docs because it named a retired workflow-advance command. But commands/gsd/next.md is a live command -- the state-aware smart-entry launcher -- and this issue requires docs to use the hyphen form every runtime actually registers. Both could not hold for this one command, so docs had been sidestepping the ban by keeping the colon form, which is exactly the defect this issue exists to remove. FEATURES.md already recorded the reassignment: the hyphen form "is not the retired workflow-advance command; it is reserved for the state-aware smart-entry launcher. Workflow advancement remains under /gsd-progress --next." With that reassignment the invariant's premise is obsolete and the guard now contradicts the documented command form, so it is retired with a comment recording why rather than deleted silently. next is now swept like every other command, and the earlier exemption added to the new guard is removed so nothing is special-cased. Four citations of the literal name: frontmatter key stay in colon form, because the source file really does carry name: gsd:next and a doc quoting it must reproduce it verbatim. Two of those lines were reworded to say which side is the frontmatter key and which is the slash command, since they previously conflated the two. Verified the retired scan would now genuinely fail against this tree -- the conflict was real and resolved, not dodged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2903): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
067a4d1c6c |
fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls (#3015)
* test(#2650): add failing-first regression for plan-phase stall detection Regression test for gsd_stall_should_recover / gsd_stall_watch and the planner.stall_* config keys, none of which exist yet — proves RED before the fix lands in the next commit. * fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls Mirrors the already-shipped executor.stall_* pattern (execute-phase.md, bug #3212) but with a dispatch change the executor's prose-only surveillance lacks: the standard planner spawn, chunked-outline planner spawn, chunked-per-plan planner spawn, plan-checker spawn, and revision-loop planner respawn now dispatch with run_in_background=true and are followed by a real, bounded bash poll (gsd_stall_watch) that returns control to the orchestrator on its own schedule instead of waiting indefinitely on a subagent that may never return. On stall, the existing accept-plans/retry/ stop recovery menu (9a/11a) is auto-surfaced instead of requiring a manual interrupt. New config keys planner.stall_detect_interval_minutes (default 5) / planner.stall_threshold_minutes (default 10) mirror executor.stall_*. The helper functions (gsd_stall_should_recover, gsd_stall_watch) live in a new lazily-loaded gsd-core/workflows/plan-phase/steps/stall-detection- helpers.md rather than inline, and per-site prose is kept minimal, because plan-phase.md is frozen under the ADR-857 Phase 6 PRE_PHASE6 gate (tests/phase6-capstone-conformance.test.cjs) with ~36 bytes of headroom at baseline; the net effect is plan-phase.md.md ships slightly SMALLER than before (the old unconditional-wait ORCHESTRATOR RULE sentences are gone at the five touched sites, superseded by the bounded watcher). Also fixes a stale doc comment in tests/workflow-size-budget.test.cjs that still described the per-file workflow-size-baseline.json guard removed by #2724 (ADR-2719 Phase 4) as if it were still the enforcement mechanism — discovered while verifying this fix's own byte budget. Researcher and pattern-mapper spawns are untouched (out of scope per the issue's Agent Brief). * fix(#2650): make gsd_stall_watch single-cycle; harden numeric config inputs Two review findings addressed on top of the prior commit: 1. gsd_stall_watch previously looped internally for the full threshold+interval duration inside ONE Bash tool call (up to 15 min at defaults) — a single call blocking that long risks the host tool's own timeout killing it before it ever prints a result, silently defeating the fix. Redesigned to a single sleep-and-check cycle per call, taking an explicit dispatch_ts so the orchestrator prose can repeat the (short, default 5 min) call until it resolves; the outer threshold is now enforced by dispatch_ts accumulating across calls, not by one call's duration. Documented the resulting trade-off (up to one interval of added latency on the success path) in the changeset and reference doc. 2. PLANNER_STALL_INTERVAL_MINUTES/THRESHOLD_MINUTES are config-controlled values that flow into bash arithmetic ($(( ))). A review flagged this as command injection; empirically verified against both macOS bash 3.2.57 and Docker bash:5 that this is NOT actually exploitable (bash hard-errors on a `$(cmd)`-shaped arithmetic operand rather than invoking it) — but an unvalidated malformed value WOULD abort the stall-watcher itself with that bash error, silently defeating the exact hang-recovery this issue ships. Added integer validation with safe-default fallback, both at the config-resolution point and defensively inside gsd_stall_should_recover. Also adds the previously-missing integration coverage for gsd_stall_watch's real execution (grep/find/date plumbing), not just the pure classifier. * fix(#2650): correct AC2 self-test — helpers doc may name teams-status in prose The AC2 regression test asserted the stall-detection-helpers.md step file never contains the substring "teams-status" at all, but the file's own prose explicitly documents its independence from that guard (containing the word by design). Narrowed the assertion to what actually matters: no second `query teams-status` call site and no gating on it, not a blanket absence of the word. * test(#2650): regenerate golden install-tree fixtures for the new step file gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md is an emitted file (installed for every runtime), so adding it changes the install tree even though it is invisible to docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (both explicitly scope to non-recursive gsd-core/workflows/*.md — verified against the execute-phase #2930 and pre-existing plan-phase step-file precedent, which are equally absent from both inventory artifacts). The golden install tree snapshots the sorted list of emitted relative paths per runtime, so a file invisible to the inventory is still visible here. Regenerated via `npm run gen:install-tree` — one line added per runtime fixture (19 files), no other drift. * fix(#2650): restore 7 ORCHESTRATOR RULE labels; sync runtime-launcher preamble Two more consequences of extracting helper bodies out of plan-phase.md, both caught by verification (0017e1a78, 9 unique failures): 1. tests/plan-phase-drift-guard.test.cjs (#913) requires at least 7 "ORCHESTRATOR RULE — ALL RUNTIMES" labels in plan-phase.md itself, one per agent spawn site. Moving the full explanatory blocks to plan-phase/steps/stall-detection-helpers.md carried 5 of the 7 labels out with them (only the untouched researcher/pattern-mapper sites kept theirs). Restored a short label at each of the 5 stall-watch sites, trimmed a few more redundant words ("Per 7.99, " — already established by the adjacent step-7.99 pointer) to stay under the frozen PRE_PHASE6 cap (94497 bytes, 21 bytes headroom). 2. tests/runtime-launcher-parity.test.cjs (#373) requires exactly one canonical gsd_run preamble, byte-equal to gsd-core/workflows/_runtime-launcher.snippet.sh, before the first gsd_run call in any workflow .md that calls it (recursive scan under gsd-core/workflows/, unlike the non-recursive inventory/step-tag-balance checks). The new step file's config-get calls use gsd_run without one. Fixed via `node scripts/sync-runtime-launcher.cjs`, verified: exactly 1 preamble occurrence, before the first call, including the .claude/ and .codex/ home fallback arms. Also verified (no fix needed, evidence recorded): the generic `gsd-core-verbatim` identity rule in tests/helpers/emitted-provenance.cjs (roots: ['gsd-core'], pattern matching workflows/.+) self-attributes any new gsd-core/workflows/** path to itself, so the new step file needs no drift-ack entry — consistent with plan-phase.md's own net shrinkage requiring none either. * test(#2650): acknowledge plan-phase.md's +14 byte drift Restoring the 5 ORCHESTRATOR RULE — ALL RUNTIMES labels (#913) flipped plan-phase.md from -142 bytes (post-extraction) to +14 bytes net growth against baseline (94483 -> 94497), which the differential attribution size ratchet (tests/emitted-attribution.test.cjs) correctly flags as unacknowledged growth. Added tests/emitted-drift-acks/2650-plan-phase- stall-detection.json, keyed on the bare filename plan-phase.md per the existing fragment schema (see tests/emitted-drift-acks/2649-diagnose- execute-plan-base-check.json), explaining the growth as exactly the 5 restored labels — still verified under the PRE_PHASE6 cap (94497 < 94519) and satisfying #913's 7-label requirement. * fix(#2650): bind {outputFile} from the real Agent() return — was dead code Independent review blocker: PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE were read by every gsd_stall_watch call but never assigned anywhere in the diff. With the variable permanently empty, `[ -f "$output_file" ]` was always false, marker_found could never become true, and marker_received was unreachable — the marker-based detection path was permanently dead. Worse for the plan-checker spawn specifically: a checker that PASSES touches no *-PLAN.md files, so it had no working completion signal at all without the marker path. A healthy plan-checker finishing cleanly in two minutes would be declared stalled once planner.stall_threshold_minutes elapsed and the recovery menu would fire on an already-succeeded agent — worse than the original unbounded hang. Fixed by replacing the dead bash variable with the `{outputFile}` orchestrator-substitution token, the same convention docs-update.md:471 already uses for a real run_in_background=true Agent() return ("Read tool: file_path: `{outputFile from README agent result}`"). This is a net BYTE SAVING at each site (`"{outputFile}"` is shorter than `"$PLANNER_OUTPUT_FILE"`), which funded moving the full binding explanation — including why plan-checker's *-PLAN.md glob alone is not a working completion signal — into the lazily-loaded reference file to stay under the frozen PRE_PHASE6 cap (94496 bytes, 22 headroom; net +13 over baseline, acknowledged in tests/emitted-drift-acks/2650-plan-phase- stall-detection.json). Added a regression test asserting plan-phase.md itself binds {outputFile} at all 5 spawn sites and contains no dangling $PLANNER_OUTPUT_FILE / $CHECKER_OUTPUT_FILE reference — the previous test suite only exercised gsd_stall_watch's behavior when handed a valid argument, which is why the dead production wiring survived two rounds of review. Also fixed tests/fix-2650-plan-phase-stall-detection.test.cjs:170-195's raw try/finally to use t.after(), per CONTRIBUTING's test-cleanup convention. * chore(#2650): backfill changeset PR number to 3015 * fix: normalize CRLF at the read boundary in all .md-bash-extraction tests Maintainer-authorized scope expansion, folded into this PR rather than deferred: the Windows CI lane on this PR's own tests/fix-2650-plan-phase- stall-detection.test.cjs exposed DEFECT.TEST-SHELL-PIPELINE-NONPORTABLE (CONTEXT.md; recurring since #1700) as a repo-wide latent class, not a one-off. Ten test files parse a fenced ```bash block out of a workflow .md file and execute it via spawnSync/execFileSync; a Windows checkout can yield CRLF line endings despite .gitattributes eol=lf, and bash then treats the trailing \r on every extracted line as part of the token — "unexpected EOF while looking for matching `"'" or a bare syntax error, partway through the script. Added tests/helpers.cjs:readFileNormalized() — strips \r\n -> \n at the read boundary, before any fence-slicing or regex runs, so every downstream operation is correct by construction. Migrated all ten call sites to it: Previously broken (fs.readFileSync with no normalization anywhere between read and spawn): - tests/worktree-cleanup.test.cjs (extractCwdGuardBash) — also fixes a misleading comment claiming the fence regex alone was "CRLF-safe"; it protected only the fence delimiters, never the captured body. - tests/new-milestone-clear-phases.test.cjs (extractFenceBetween, extractFenceContaining) - tests/code-review-pipeline-regression.test.cjs (extractPostProcessingScript) - tests/drift-detection.test.cjs (readGate/bashBlock, plus the snippet-file comparison read in the same test) - tests/graphify-visualization.test.cjs (extractStep3Block) - tests/pause-work-improvements.test.cjs (extractCheckBlock) - tests/plan-review-convergence.test.cjs (extractReviewerFlagsParseBlock and the inline post-config-gate resolution-block slices) Already correct (split(/\r?\n/) then join('\n')), migrated to the shared helper for consistency rather than a fourth/fifth/sixth copy of the same fix: - tests/git-base-branch.test.cjs (extractHandleBranchingBash) - tests/quick-branching.test.cjs (extractStep25Bash) - tests/runtime-launcher-parity.test.cjs (extractResolverSnippet) Verified against a simulated Windows CRLF checkout (not assumed): for both the worktree-cleanup.test.cjs and new-milestone-clear-phases.test.cjs extraction shapes, confirmed the pre-fix code produces a real bash syntax error on CRLF input and the post-fix code does not. One eslint follow-up: local/no-crlf-fragile-split statically flags any bare `\n` inside a markdown-fence-shaped regex, regardless of whether the receiver was already normalized — it cannot see the readFileNormalized() data-flow. Kept `\r?\n` in extractCwdGuardBash's fence regex (redundant but harmless on pre-normalized input) rather than fight the rule. Scope note: this diff is broader than issue #2650's own change (plan- phase.md stall detection) because the Windows lane surfaced a genuine repo-wide defect class while verifying that fix, and the maintainer authorized fixing it here rather than filing it separately and shipping a known-broken pattern. Runtime impact: none — this is a test-harness-only defect. The live orchestrator (Claude Code or another runtime) does not do a byte-exact extract-and-pipe of .md content into a shell the way these tests do; it reads the instructions and generates its own bash invocation text, which does not reproduce a raw CRLF pass-through the same way. Not touched: tests/plan-review-convergence.test.cjs's separate, tracked spawnSync ETIMEDOUT flake under bench load (#3005, reproduced on unmodified next) — unrelated load-sensitivity, not a CRLF symptom. * fix(#2650): remove stale drift-ack fragment — plan-phase.md is self-explaining tests/emitted-drift-acks/2650-plan-phase-stall-detection.json acknowledged plan-phase.md's own emitted-path hash move, but plan-phase.md is directly edited in this diff. Per the emitted-attribution law (ADR-2719, tests/emitted-attribution.test.cjs), a workflow's emitted key equals its own source path (gsd-core-verbatim identity rule), so a direct edit to the source is self-explaining and auto-attributed — no ack was ever needed. Verified via the pre-merge lint (scripts/lint-emitted-drift-ack.cjs, run through npm run lint:ci with a fully cleared eslint cache): it passes clean with the fragment removed, confirming no contradiction between the lint and the runtime attribution gate — this was simply an unnecessary fragment. * fix(#2650): restore plan-phase.md drift-ack — size ratchet demands it against next tests/emitted-drift-acks/2650-plan-phase-stall-detection.json was deleted in the previous commit because, against an earlier verification base, it was inert: it explained a moved emitted hash that a direct edit to plan-phase.md already self-attributes. Against origin/next@f1af47766a the demand is different: plan-phase.md is 13 bytes larger than the base copy, which trips the emitted-attribution size ratchet — a job this same ack also performs. Recreated in the documented shape, keyed on the bare filename plan-phase.md (not the full path, and not restating the byte delta per review guidance), describing the actual change: the {outputFile} binding fix for the dead PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE variables and the 5 restored ORCHESTRATOR RULE labels required by #913, both at the stall-watch spawn sites, with explanatory bodies living in the lazily-loaded gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md reference. Confirmed no other fragment (on this branch or on next) claims the bare key "plan-phase.md" before recreating — scripts/lint-emitted-drift-ack.cjs's duplicate check is an exact string match, and the only other mention of plan-phase.md in tests/emitted-drift-acks/ (2658-trae-instruction-file-path.json) uses the full path as its key, so there is no collision. * fix(#2650): real cause of Windows CI failure — bash -c argv-transport, not CRLF The CRLF diagnosis for PR #3015's Windows failure was wrong. Proven wrong, not assumed: .gitattributes' blanket `* text=auto eol=lf` means a Windows checkout never receives CRLF for stall-detection-helpers.md, and the extracted fence's line 64 is byte-identical and correctly balanced on every platform. The real cause: runShouldRecover() passed a 70+ line, quote-dense script as ONE argv element to `spawnSync('bash', ['-c', script, arg0, ...])` PLUS four more positional args. Windows has no execve — Node serializes that whole argv into a single CreateProcess command-line string, and Git Bash's MSYS layer re-splits and unescapes it with its own rules. The boundary between the script and the trailing args was not stable across that round trip (live evidence: one failure's stderr was prefixed `gsd_stall_should_recover_test:` — arg0 arrived — another `/usr/bin/bash:` — arg0 did not). Fixed by writing the script to a temp file and running `bash <file> <args>` instead — the four values are now normal, quote-free positional args, and the script itself never enters argv transport at all. Mirrors tests/quick-branching.test.cjs's extractStep25Bash/runStep, which already uses this exact shape and is green on Windows on `next`. tests/worktree-cleanup.test.cjs's extractCwdGuardBash/runGuard stays on `bash -c` but never appends extra positional args beyond the script itself, so it never hits the same boundary — checked both siblings per review, not assumed. Corrected the now-actively-misleading CRLF comment in extractStallHelpersBash(), and corrected the changeset's claim that the repo-wide CRLF-normalization fix (folded into this branch, maintainer- authorized) explains this PR's own Windows failure — it doesn't, though it remains defensible on its own merits as general test-portability hardening. Separately, while auditing the shipped (non-test) gsd_stall_watch for Windows portability per review request, found and fixed a second, real user-facing defect: the artifact-freshness check used GNU find's `-newermt "@<epoch>"` shorthand, which the BSD find(1) actually shipped on macOS does NOT understand ("Can't parse date/time: @<epoch>", verified live against /usr/bin/find on both a stale and a genuinely fresh file). With the adjacent `2>/dev/null`, that failed silently and permanently degraded artifact_fresh to false on every macOS run — a plan-checker or planner actively writing plan files could still be reported "stalled." Replaced with `find $glob -mmin -N` ("modified less than N minutes ago"), which needs no date-string parsing and is supported identically by GNU find and BSD find; verified live that the old shape fails and the new shape passes against the same real fresh file. Added a real-execution regression test (gsd_stall_watch with `sleep` stubbed to a no-op so the test doesn't actually wait, but the real `find ... -mmin` line still runs) proving the fix, replacing the prior "not integration-tested" note for that path. Note: the remote gsd-test runner is Linux-only, so it cannot itself confirm the Windows fix — only the actual windows-latest CI lane can. * fix(#2650): route the third bash -c call site through the same temp-file seam runWatch() and a `-mmin` regression test still passed their script via `bash -c <script>` after the previous commit only converted runShouldRecover() — live Windows CI on 4b86cc57f confirmed the mechanism: failures went 11 -> 4, and `full test (windows-latest, 22, shard 1/3)` and `shard 2/3` flipped from fail to pass, but the remaining 4 failures (all in this file, all still `bash: -c:`) were exactly the gsd_stall_watch describe block, which runWatch() serves. runWatch() passes NO extra positional args at all, so this also rules out the trailing-args theory from the prior commit: the ~73-line, quote-dense script itself is what does not survive Windows argv serialization when passed as a single `-c` element, regardless of how many (if any) further argv elements follow it. Extracted one shared runBashScript(script, args, opts) helper — write to a fs.mkdtempSync'd file, run `bash <file> [args...]`, clean up in `finally` — and routed all three bash-invoking call sites in this file through it (runShouldRecover, runWatch, and the -mmin freshness test that builds its own script inline for the `sleep` stub). One transport seam means a fourth call site in this file cannot silently reintroduce the bug in isolation, which is exactly what happened here with a second call site. Corrected extractStallHelpersBash()'s doc comment a second time to state the mechanism precisely (script content, not argv-element count) and cite the live evidence (11->4 failures, shards 1 and 2 flipping green) so the next reader does not have to rediscover it. Audited every other bash-invoking call site in files this branch touches, per review request: - tests/code-review-pipeline-regression.test.cjs (runPostProcessing), tests/graphify-visualization.test.cjs (runBlock), and tests/drift-detection.test.cjs (two execFileSync('bash', ['-c', ...]) sites, one of them carrying the same giant runtime-launcher preamble text) — all pre-existing, UNCHANGED by this branch (only touched for the readFileNormalized() CRLF swap), and already exercised on `next`'s last six Windows CI runs per the reviewer's own citation. Left as-is: no evidence of failure, and converting untested pre-existing code outside #2650's scope on an unverifiable guess would be its own risk. - tests/git-base-branch.test.cjs (runHandleBranchingStep) and tests/quick-branching.test.cjs (runStep) already use the same temp-file pattern. No action needed. - tests/runtime-launcher-parity.test.cjs (runResolver) uses `bash -c` but is explicitly `if (process.platform === 'win32') return '';` guarded off on Windows entirely, for an unrelated extension-less-PATH-stub reason — never reaches Windows argv transport at all. No action needed. - tests/worktree-cleanup.test.cjs (runGuard) confirmed by the reviewer as correct and verified; not touched, per instruction. Do not touch: the -mmin fix, the drift-ack fragment, the changeset — all three confirmed correct in prior rounds and left untouched here. Note: the remote gsd-test runner is Linux-only and cannot confirm this; only the windows-latest lanes on #3015 can. * fix(#2650): give runBashScript a default timeout runShouldRecover() was the only one of the three call sites through runBashScript() with no timeout — runWatch() and the -mmin test both pass timeout: 10000 explicitly. Not a regression (this path never had a bound before), but CONTEXT.md's unbounded-subprocess guidance applies directly, and runShouldRecover() is driven repeatedly by a fast-check property test: one pathological input that fails to terminate would hang CI indefinitely instead of failing. timeout: 10000 is now the helper's own default, with ...opts spread after it so the two existing explicit timeout: 10000 call sites are unchanged and any future caller inherits a bound automatically. * fix(#2650): build the -mmin freshness test's glob with forward slashes Windows CI on d6ddda6ea reported the last failure: the -mmin regression test expected 'active' but got 'waiting' — find matched nothing, the same silent-degradation shape as the macOS -newermt defect, but this time in the test's own fixture rather than the shipped bash. Traced what production actually passes: every gsd_stall_watch call site in plan-phase.md builds artifact_glob as `"${PHASE_DIR}"'/*-PLAN.md'` — PHASE_DIR is a POSIX-style .planning/phases/NN-slug value, and the whole thing runs under Git Bash regardless of host OS, so production's glob is always forward-slash. The test instead built it with `path.join(tmp, '*-PLAN.md')`, which on Windows yields a backslash path (C:\Users\RUNNER~1\...\*-PLAN.md). In bash pathname expansion a backslash escapes the next character, so that pattern can never match a real path — find silently returns empty under the existing 2>/dev/null, same shape as the macOS bug. Confirmed as a test artifact, not a production defect: production never constructs the glob this way, so no Windows user is affected. Fixed by forward-slashing the tmp dir before appending the glob suffix, matching production's own convention, with a comment recording why (so a future "simplify this back to path.join" edit doesn't silently reintroduce the failure). The shipped bash's unquoted $artifact_glob is untouched — quoting it would break the multi-file glob expansion it exists for. Note: the remote runner is Linux-only and already passed clean at d6ddda6ea (0/29,603, both node lanes); only the windows-latest lanes on #3015 can confirm this fix. * fix(#2650): forward-slash the three remaining runWatch globs (vacuous-pass CR) The :353 fix (833c11da9) only converted the -mmin freshness test's glob. Three sibling tests in the same describe block still built theirs with path.join(tmp, '*-PLAN.md'), which yields a backslash path on Windows. Two of those three were silently passing for the wrong reason: the '-> stalled' and '-> waiting' tests both expect the glob to match nothing, and on Windows a backslash path matches nothing regardless of whether the directory is actually empty (bash eats each backslash as an escape before the pattern is even evaluated). They would have passed identically with glob expansion completely broken, which is a vacuous pass — not exercising what they claim to. The third ('-> marker_received') is outcome-independent of the glob, so it was merely inconsistent rather than wrong. Converted all three to the same `${tmp.replace(/\\/g, '/')}/*-PLAN.md` construction already used at the -mmin test, so every glob in the file now matches production's own forward-slash `"${PHASE_DIR}"'/*-PLAN.md'` shape, and the two negative tests are meaningful on Windows instead of accidentally correct. Reworded the trailing comment on the 'stalled' test's glob line: it now describes the fixture (the tmp dir contains no *-PLAN.md files) rather than the pattern, since "matches nothing" read as a property of the glob syntax when it's a property of what's on disk. No assertion, the sleep stub, runBashScript, or the shipped bash changed. Smoke-tested all three updated tests manually before committing (not via node --test): marker_received / stalled / waiting, all correct. * fix(#2650): fix own regression tests for #2993's plan-phase.md relocation 531101843's merge with origin/next brought in #2993 (unrelated, epic #1671 Phase 6.2), which extracted plan-phase.md's whole "Chunked Planning Mode" section into gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md, leaving a <!-- gsd:section --> pointer behind. tests/plan-phase-drift-guard. test.cjs (#913) was already updated to read the combined surface (host file + every steps/*.md) so its label count didn't go blind — my own #2650 regression tests were not, and searched plan-phase.md alone for the two chunked spawn sites' headings, which no longer exist there. Two tests failed outright (indexOf returning -1); a third ("standard planner spawn") was silently weakened to an unbounded slice-to-EOF by the same relocation, since its own end-boundary heading also moved — passing by accident rather than by testing what it claimed. Promoted the drift guard's local readPlanPhaseCombined() to a shared, exported tests/helpers.cjs readWorkflowCombined(workflowPath) (host file + sorted steps/*.md, CRLF-normalized at the read boundary) so a second, divergent implementation is never written — the drift guard now delegates to it via a same-named local wrapper, unchanged at every existing call site. Fixed the three affected tests in tests/fix-2650-plan-phase-stall-detection. test.cjs: - "standard planner spawn (step 8)": end boundary changed from the now-gone "## 8.5. Chunked Planning Mode" heading to "## 9. Handle Planner Return", which still exists in plan-phase.md. - "chunked outline spawn (8.5.1)" / "chunked per-plan spawn (8.5.2)": now read gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md directly (not the generic multi-file combined blob, whose file-sort ordering would put unrelated step files between 8.5.2's slice and any downstream anchor) — the same heading-to-heading slicing as before still works because the file is small and self-contained. - Extended the "no unbound $PLANNER_OUTPUT_FILE/$CHECKER_OUTPUT_FILE" check to also scan chunked-planning-mode.md, since two of the five spawn sites now live there. - Added a new count-based test asserting exactly 5 (not "at least one") `gsd_stall_watch "$TS" "{outputFile}"` invocations across the combined surface, mirroring #913's own label-count guard, so every one of the five spawns stays provably bounded and a future relocation can't silently drop one without a test noticing. Also added a small positive test that plan-phase.md's <!-- gsd:section --> pointer to chunked-planning-mode.md exists (#2993 is unrelated to #2650 but its presence is now load-bearing for where 2 of the 5 spawn sites live). Audited every other test file in the repo for a stale reference to content #2993 relocated (searched for the moved headings/prose and for "chunked-planning-mode"/"CHUNKED_MODE" across all *.test.cjs): only this file and the drift guard needed changes. tests/issue-2762-plan-reviews-chunked.test.cjs already reads chunked-planning-mode.md directly (brought in correct by the same merge). gen-section-manifest.test.cjs, init.test.cjs, and workflow-fragments.test.cjs reference "chunked-planning-mode" only as a manifest/section-id fixture value for #2993 itself, not as a stale pointer to relocated content. Did not touch: the ported ORCHESTRATOR RULE lines, run_in_background=true, the glob constructions, runBashScript, the -mmin change, the timeout default, or the drift-ack fragment (confirmed correct against the stale local `next` ref two rounds ago and left alone). --------- Co-authored-by: sim <sim@local> |
||
|
|
7b204ad2ac |
enhance(#2530): extend UAT checkpoint frame language pack (9 more languages) (#2564)
* feat: extend UAT checkpoint frame language pack (9 more languages) response_language is a free-form config value, but CHECKPOINT_FRAMES only covered 9 languages — any other configured language silently fell back to the English frame. Add Dutch, Polish, Russian, Ukrainian, Turkish, Hindi, Arabic, Vietnamese, and Indonesian frames plus their aliases, with a regression test asserting each resolves instead of falling back. Follow-up to #2402 (PR #2457). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: add changeset for #2527 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(#2530): list UAT checkpoint frame languages in CONFIGURATION.md Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2530): point changeset fragment at PR #2557 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2530): address Unicode language-pack review * fix: address checkpoint language review * fix: count spacing combining marks in checkpoint width * test: verify checkpoint aliases structurally * fix: isolate RTL checkpoint frames * fix: isolate RTL checkpoint frames correctly * test(#2530): assert checkpoint aliases neither collide nor go unreachable Review Minor #1. A duplicate alias key was invisible to the existing catalog tests: the runtime object is well-formed after JS collapses the literal, the self-alias assertion still holds, and the losing language just stops resolving. tsc catches the byte-equal case (TS1117), but not the two that survive compilation — an alias whose NFC-lowercase form already belongs to another language, and an alias not in lookup form at all, which resolveCheckpointFrame() can never produce. The check reads the source literal rather than the object, since the object no longer records what was written. Both assertions are independently load-bearing: an NFD twin of an existing alias trips the collision check, an uppercase alias trips the unreachability check. Review Minor #2: changeset retyped Changed -> Added. Nine wholly new supported response_language values are an addition under Keep a Changelog, not a modification of existing behavior. * test(#2530): check alias collisions on the catalog, not its source --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: Rezolv <dave@sienkowski.com> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
0408276791 |
chore(#2797): federate reviewer config keys off the central schema (#2841)
* chore(#2797): federate reviewer config keys off the central schema Phase 4 of epic #2782 (ADR-2782 D9, config half). Runs AFTER 5a per the ADR's swap amendment: a federated config slice lives inside a capabilities/<id>/capability.json, and three of the five key families had no capability directory until 5a created them. Four key families move to the lanes that use them; the central-schema removal and the federated addition land in this one commit because the exclusivity invariant fails the build on a key present in both. review.max_prompt_tokens, review.default_reviewers and review.reviewer_instances describe policy ACROSS lanes and stay central. Two things the issue did not name, both found while building it: 1. THE EXCLUSIVITY GATE WAS BLIND TO PATTERNS. It compared federated keys against manifest.validKeys only, and two of the four families (review.models.<slug>, review.max_prompt_tokens_per_reviewer.<slug>) were pattern-backed. That is not cosmetic: isCentralConfigKey consults those patterns and mergeFederatedConfig skips every key for which it returns true, so declaring a slice while the pattern survived would have shipped an INERT slice behind a green gate — the exact half-migrated shape the invariant exists to prevent. The gate now loads the patterns from the same manifest the runtime reads. 2. AN UNSET PER-LANE BUDGET NOW RESOLVES TO 0, NOT NOT-FOUND, because a federated key always resolves to its declared default. The three fallback guards in review.md checked only empty-or-"null", so a user who set the GLOBAL review.max_prompt_tokens would have silently lost trimming on the HTTP lanes. The guards now treat 0 as unset. D9 says review.models.<slug> is owned by "the lane whose slug it names". That is false for one lane: the shipped key is review.models.agy while the slug is antigravity. Ownership follows the lane; the key name is preserved, because renaming would break every config that sets it. Existing tests updated rather than left asserting the old world: config-get on a cleared federated key yields empty instead of not-found (what #2046 actually protects — never persisting the literal "null" — is unchanged and still asserted); the config-schema dynamic pattern representative moves to reviewer_instances; the prototype-pollution guard case moves to a surviving dynamic prefix so alert #26 keeps its coverage, with a new case asserting the old key is now rejected earlier; and Phase 2's harvest-widening inertness assertion becomes an ownership assertion, since Phase 4 is what consumes it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): use a -1 sentinel so an explicit per-lane budget of 0 survives A federated config key always resolves to its declared default, so an unset per-lane prompt budget needed a value the workflow could treat as 'not configured'. The first cut used 0 — which is wrong: 0 is already a LEGITIMATE per-lane budget meaning 'do not trim this lane' (the early-return guard in prepare_trimmed_prompt_for_reviewer). Treating it as unset would have silently switched a user who deliberately disabled trimming for one lane onto the global budget. The sentinel is now -1, which is not a valid token budget, so all three states stay distinguishable: unset falls back to global, an explicit 0 disables trimming for that lane, and an explicit N is used. Locked by three CLI round-trip tests. Surfaced by the isolated security reviewer before it crashed mid-run; verified independently against the shipped trim guard rather than taken on trust. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): update central-registration assertions and stay under the review.md cap The remote runner caught both; my local sweep missed the files. 1. tests/plan-review-convergence.test.cjs asserted the three local-server host keys are in VALID_CONFIG_KEYS. They are federated to their lane capabilities now, and the exclusivity invariant forbids a key living in both places. What #2306-local actually protects is that config-set ACCEPTS them, so that is what is asserted — via isValidConfigKey, the predicate config-set itself uses, which spans central and federated. A second assertion pins federated ownership, so a silent reversion back to the central schema fails too. 2. review.md exceeded the LARGE tier hard cap (62583 > 61440). That cap is a red line, not a budget to raise. The three per-lane budget guard comments were near-identical; condensed to one terse line each. 61371 bytes, 69 to spare. Real extraction to workflows/review/modes/ is Phase 5b/6 work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): fail closed on a broken config-schema manifest; reconcile stale docs Isolated security review findings. MAJOR — loadCentralConfigPatterns failed OPEN. It swallowed a JSON parse error and returned [], while its sibling loadCentralConfigKeys, reading the SAME file, writes to stderr and throws ExitError(1) on that identical failure class. Fail-open here defeats the gate this function exists to feed: with zero patterns, validateCrossCapability's pattern-collision check silently passes and an inert federated slice ships green. It was masked in the one production call site only because loadCentralConfigKeys runs first against the same path — a coincidence of ordering, not a guarantee, and this function is exported and called standalone. The two now share a contract: ENOENT is the legitimate absent case, anything else throws loudly. A single unparseable PATTERN is still skipped, which degrades to "checked less" rather than blocking every build. The branch had zero coverage; it now has two tests (malformed JSON, EISDIR). MINOR — docs/CONFIGURATION.md still listed review.models.qwen and review.models.cursor as settable, ~770 lines below this PR's own new Ownership section. Those lanes take no model flag, so they declare no model key and config-set now rejects them. Rows removed; the missing review.models.agy row added; the per-reviewer budget row corrected to name only the lanes that own a budget key, and to document that a per-lane 0 disables trimming for that lane. Also fixes a shadowed "raw" binding introduced by the fail-closed change, which made the generator unrequirable — caught immediately by its own --check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2797): backfill changeset pr number to 2841 * fix(#2452): make the base-ref mutation test hermetic against leaked GIT_* env tests/mutation-workflow-base-ref.test.cjs fails on PR branches while next stays green, and it is currently blocking at least three unrelated PRs (#2841, #2832, #2827) with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees The existing loop comment attributes this to `git add .` rehashing O(n^2) blobs "before the object write had landed" and works around it by staging one path per iteration. That is not the cause: sequential execFileSync calls cannot race each other's object writes, and the failure persisted after that change — it simply moved to a lower commit index. The cause is that the git() helper inherited the runner's environment. A leaked GIT_INDEX_FILE makes `git add` write into a DIFFERENT repository's index; GIT_OBJECT_DIRECTORY / GIT_ALTERNATE_OBJECT_DIRECTORIES send the blob to another object store; GIT_DIR / GIT_WORK_TREE redirect the whole operation. In every case `git commit` then cannot resolve a blob it just staged, which is precisely the error above. Verified by negative control: with GIT_DIR exported, this test fails on the unfixed helper (the git commands operate on the wrong repository entirely); with the helper stripping GIT_* it passes. The single-path staging is kept — it is genuinely less work — but it is no longer load bearing. Found while shipping #2797. Fixed in place rather than deferred: it is a defect surfaced during the work, and it is blocking other contributors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2452): build the base-advance commits empty, removing the lost-object class The base-ref guard has been failing in CI with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees It is currently red on at least three unrelated PRs (#2841, #2832, #2827) while next stays green. Two theories have now been tried and neither held. #1881 blamed `git add .` rehashing O(n^2) blobs and switched to staging one path per iteration; the failure moved from commit 32 to commit 25 and carried on. The preceding commit here made the git helper hermetic against leaked GIT_* environment — that IS a real vulnerability (with GIT_DIR exported the helper operates on the wrong repository entirely, proven by negative control) but it produces a different error than CI reports, so it is not demonstrably the cause either. Neither trigger reproduces off-CI, so this stops guessing at the trigger and removes the failure CLASS instead. The loop needs base-branch DEPTH and nothing else: no assertion reads these commits' contents, and base-side files cannot appear in `origin/base...HEAD` regardless. `--allow-empty` writes no blob and no tree, so there is no object for the index to reference and lose. It is also far less work than 60 write+hash+index cycles. The guard still proves its mechanism: the test asserts that a --depth=1 base fetch FAILS and a full fetch resolves, so a broken topology would surface immediately rather than passing vacuously. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Test <test@example.com> |
||
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
16e59d0db5 |
fix(#2691): repair seven dangling references in the ADR corpus and contributor docs (#2692)
* fix(#2691): repair five dangling references in the ADR corpus and contributor docs
Found by the 2026-07-24 ADR corpus audit; each mechanism re-reproduced live
against next @
|
||
|
|
3eb1cede26 |
fix(#1880): distinguish a corrupt config from an absent one (epic #1879 Phase 1) (#2688)
* test(#1880): prove corrupt config is indistinguishable from absent Failing-first. Encodes the issue's runtime repro: a trailing comma in .planning/config.json currently yields source:builtin-defaults with degraded:false - byte-identical to the file not existing - and the user's entire configuration is silently discarded. Asserts on the typed surface (CONFIG_REASON, _warnedUnusableConfig) rather than diagnostic prose, per the ADR-1411 amendment's test-methodology clause and CONTRIBUTING.md's raw-text-matching rule. IO failure is injected by monkeypatching fs.readFileSync and restoring in t.after(), never chmod 0o000 (root bypasses mode bits). Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): distinguish a corrupt config from an absent one loadConfigResolved wrapped the read, the JSON.parse and the entire config build in one try with one catch, so ENOENT, EACCES and SyntaxError all fell through to the same defaults and the branches returned degraded:false - actively asserting health over discarded configuration. A single trailing comma in .planning/config.json silently replaced the user's whole config, reporting source:builtin-defaults degraded:false, byte-identical to having no config file at all. ConfigResolution now carries a machine-readable reason. Genuine absence keeps degraded:false / not_configured; a file that exists but cannot be used sets degraded:true with config_unparseable or config_unreadable. The same split applies to the root config and to ~/.gsd/defaults.json. Control flow is deliberately unchanged. preflight_check reports cyclomatic 141 / cognitive 196 and 93 dependents on this function, with the guidance that small edits beat one big one, so faults are CAPTURED at the existing read sites and stamped onto the returns rather than the try/catch being restructured. Also carries the ADR-1411 amendment's wiring clause: loadConfig returns .config alone to ~51 call sites and would never see the new field, so an unusable file emits a deduplicated stderr diagnostic keyed on resolved path plus errno. Without it the reason would be an unreachable field and the user whose config was discarded would still get no signal - the actual defect. Registers the config-loader seam in lint-resolution-provenance, which until now guarded only agent-skills. Caller audit: ConfigResolution.degraded has exactly one consumer outside this module, cmdAgentSkills (src/init.cts:2259), which destructures {config, source, degraded} - adding a field does not break it. Its --json IR now reports degraded:true for a corrupt config, which is the intended fix and the one observable behavior change. Closes #1880 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): degrade when any config on the path is unusable, not just the last Two defects found by isolated adversarial review of the first cut. BLOCKER: the success-path return did not consult configFault. A corrupt ROOT config whose workstream override happened to parse returned degraded:false / reason:resolved - the root's settings silently dropped, which is the exact failure this issue closes, reappearing for any project using workstreams. The stderr diagnostic fired, so the out-of-band half worked while the in-band half reported a clean resolve; a --json consumer saw health. MAJOR: reason was derived from Object.keys(parsed) - the root+workstream MERGE - so an empty workstream file inheriting a non-empty root reported resolved despite carrying no settings. Emptiness is now judged on the file actually read, snapshotted before normalizeLegacyKeys mutates it. Also: corrects the ConfigResolution JSDoc, which still described the pre-#1880 degraded contract; adds a fast-check property asserting a PRESENT file is never reported not_configured whatever its bytes (CONTRIBUTING.md parser rule); and asserts the literal enum values so the provenance lint's configured_empty/not_configured markers check real assertions rather than incidental prose in test titles. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): reject valid JSON that is not a config object at the read seam The fast-check property added in the previous commit failed on both node lanes: a config.json containing 0, "str", [], null or true is valid JSON, so it parsed "ok", then threw downstream in normalizeLegacyKeys, and the outer catch reported not_configured - a PRESENT file reported as absent, which is precisely the collapse this issue exists to close. The property asserts a present file is never not_configured, and it caught it. _readConfigFile now validates shape, not just parseability (ADR-227: check the semantic shape at a trust boundary, not merely the type). A non-object JSON document is an unusable config, reported config_unparseable. Adds named regression cases for each non-object form alongside the property, so the class is documented and not only randomly sampled. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1880): backfill changeset pr number (pr:0 -> 2688) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2452): record a fetch-time shallow failure instead of crashing This guard failed CI on ubuntu-24 while passing on ubuntu-22 and windows-24 for the same commit, and passed on other PRs. Not a flake and not caused by the change under test - a real fragility in the test. runnerDiff ran the base fetch OUTSIDE its try and only guarded the diff, so it assumed the failure mode is always 'fetch succeeds, diff reports no merge base'. At a shallow boundary that lands short of the merge base, git can instead fail during the FETCH ('unable to parse commit' - the boundary commit's parent is not available). Which stage git fails at is version and transport dependent, so on some runners the error escaped runnerDiff and crashed the test rather than being recorded as the ok:false the assertions expect. Both stages mean the same thing for what this guard protects: a shallow base ref cannot resolve the three-dot diff. Also drops two assert.match calls against git's stderr prose. 'no merge base' and 'unable to parse commit' are the same condition reported at different stages, and CONTRIBUTING prohibits raw text matching on subprocess output. The typed outcome (ok === false) is the contract; the tests now assert that plus the presence of a cause. Found while investigating the red lane on #2688; fixed here per the no-defer rule rather than filed. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
46ba02acde |
feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs * fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions * fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests * chore(#2630): backfill changeset pr to 2661 |
||
|
|
6ad30f74b6 |
feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler. harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run. Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard. Closes #2627 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
155c08facf |
docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase (#2574)
* docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase These two commands never parse --validate (silent no-op); the flag is real only for /gsd-quick. Remove the false flag-table rows and CLI examples across COMMANDS.md and the how-to guides (en + ja-JP/zh-CN/ ko-KR/pt-BR mirrors), and correct the manager.flags.execute example from --validate to --cross-ai (a flag execute-phase actually parses). /gsd-quick's real --validate docs are left untouched. Ref #2197 * docs(#2197): add changeset for --validate docs removal --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
09b535ac00 |
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.
Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.
Every per-host value is documentation-sourced, never inferred:
- claude argv -- verified via 'claude --help' (--effort <level>)
- opencode argv -- verified via 'opencode run --help' (--variant)
- codex argv -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
flag (config.toml key only), so the global -c override is the
only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
fails closed rather than inheriting a profile baseline
No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by
|
||
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|