cd5b8643aee18e97994cc24d302f579d9feabbeb
589 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cd5b8643ae |
fix(#2828): state sync reports correct total_phases on a flat unmilestoned roadmap (#2892)
* fix+test(#2828): total_phases uses roadmap count on flat unmilestoned roadmap The read-path disk-scan cache fell back to phaseDirs.length (1) when milestoneBounded was false, even though roadmapPhaseCount (6) was correct for a flat roadmap (no sibling milestones to conflate). Use roadmapPhaseCount as the floor when > 0, matching the write-path (cmdStateSync) which already did this. The milestoneBounded flag still flows to milestoneUnbounded for the percent-skip (#1761 guard preserved). Regression test asserts state-sync writes progress.total_phases:6 for a flat 6-phase roadmap + 1 phase dir. * chore(#2828): changeset fragment * test(#2828): add negative-space coverage (Math.max floor mutant + no-roadmap fallback) — review findings The 6-phase test alone couldn't kill a Math.max-dropping mutant (1<6). Add: a 3-phase-dir/2-roadmap-phase case proving Math.max(dirs,count) floor; a no-roadmap case proving phaseDirs.length fallback. * fix(#2828): refine — distinguish flat unmilestoned from milestoned-unbounded (preserve #1761) The first-pass fix (roadmapPhaseCount > 0 always) re-broke #1761: a milestoned- unbounded roadmap (asserted milestone not among existing version headings) conflated sibling milestones (8 = 4+4). Refine with a hasMilestoneSectioning discriminator: ^#{2,3}(?!Phase) detects non-Phase h2/h3 milestone section headings. A FLAT roadmap (only ### Phase headings + a # title) has none → safe to use roadmapPhaseCount; a SECTIONED-but-unbounded roadmap has them → fall back to phaseDirs.length (#1761). Verified both cases locally (flat→6, sectioned-unbounded→1). * test(#2828): remove two fragile negative-space tests (phase-dir scanner internals) The Math.max-floor and no-roadmap tests made assumptions about the phase-dir scanner's internals (which dirs count as 'realized') that didn't hold. The core regression test (6-phase flat → total_phases:6) plus the existing #1761 conflation tests (which the refined fix preserves) provide sufficient coverage. * chore(#2828): backfill changeset PR number 2892 * fix(#2828): replace ReDoS-prone regex in regression test with line-by-line parse CodeQL flagged the nested-quantifier regex (`(?:[ \t]+\w+:.+\r?\n?)*?`) in tests/issue-2828-flat-roadmap-total-phases.test.cjs as a high-severity catastrophic-backtracking risk. Rewrite the STATE.md progress.total_phases extraction as a ReDoS-safe line-by-line block walk. --------- Co-authored-by: Test <test@example.com> |
||
|
|
4bd6fb066b |
chore(#2880): close ADR-2143 deployment misses — table-regex fingerprint + state-document seam migration (#2889)
* chore(#2880): close ADR-2143 seam misses — widen table-regex fingerprint, migrate state-document onto the seam The no-adhoc-markdown-parsing rule matched only a negated class whose sole member was a pipe ([^|]), so the stricter and more common [^|\n] spelling evaded it entirely -- src/state-document.cts hand-rolled exactly that shape and linted clean. Widen the fingerprint to any negated class excluding a pipe, which is the ADR-2143 section 7 prohibition as written. With the rule fixed, state-document.cts goes red. Replace tableRowPattern with locateFieldRow: a line scan using the markdown-table seam's splitTableRow for cell semantics, returning the value cell's byte range, and splice that range instead of running a whole-document content.replace. An edit now physically cannot cross a row boundary (section 4). Behavior is frozen -- stateReplaceField has 79 dependents across 5 command processes. Characterization tests lock all 14 table-branch rows plus CRLF, extract round-trip and the withFallback caller shape; a fast-check property asserts every non-target line stays byte-identical. Refs #2880, epic #2143 * fix(#2880): address adversarial review — lone-CR rows, field-name padding, quadratic scan, over-broad fingerprint Isolated adversarial review found four defects in the first commit. 1. locateFieldRow split lines on \n only. JS treats a lone \r as a line terminator, so the regex it replaced matched rows separated by bare CR. "| Phase | 3 |\r| Other | 9 |" returned 3 before and null after. Now CR, LF and CRLF are all terminators, byte offsets unchanged. 2. The field name was normalised with trim().toLowerCase(). The old regex embedded it verbatim, so its whitespace had to be absorbed by the row's own padding -- and because the group is a literal-character match rather than a whitespace class, a tab-padded cell does not accept a space-padded name. Replaced with an offset-aligned search reproducing the original backtracking exactly. 3. The widened fingerprint regex had two unbounded [^\]]* around an optional and ran quadratically over every regex source in every linted file: 256000 chars took 23 seconds. Replaced with a single-pass scanner that never rescans; the same input is now ~1ms. 4. The fingerprint also matched non-table idioms such as [^\s|] and [^"|]. Narrowed to a class excluding the pipe plus only \n, \r or \t. Differential fuzz against origin/next: 20000 cases, 0 mismatches. Refs #2880, epic #2143 * test(#2880): drop wall-clock assertion from the ReDoS regression guard local/no-elapsed-assertion flagged the elapsed-time check, and CLAUDE.md bans timing assertions outright as flaky. The 256000-char input stays as the regression guard for the quadratic scan; correctness of the verdict is what is asserted. If the quadratic path returns, the test stops completing and surfaces as a suite timeout rather than a silent pass. Also adds the changeset fragment for #2880. Refs #2880 * fix(#2880): spec-correct case folding, property tests, naming Code-review findings. The field-name comparison used toLowerCase(). The regex it replaced used /i WITHOUT /u, and ECMAScript Canonicalize deliberately does not fold a non-ASCII character onto an ASCII one -- KELVIN SIGN U+212A matched ASCII K where the old code returned null. Replaced with spec-correct Canonicalize, including the multi-character uppercase case (eszett -> SS), which a naive uppercase comparison also gets wrong. Added the fast-check property tests CLAUDE.md requires for parsers: one for the negated-class scanner, one for the field-name fold semantics, each against an independent reference implementation. Both reference impls failed on first run against real bugs, so neither property is vacuous. Renamed p2/p3 to name the exactly-three-pipes invariant, and reduced a duplicated comment to a cross-reference. Differential fuzz vs origin/next: 20000 runs, 0 mismatches, with the harness proven to discriminate the KELVIN case. Refs #2880 * chore(#2880): backfill changeset PR number (#2889) * docs(#2890): correct the local ESLint plugin path in CONTEXT.md CONTEXT.md named the local AST-rule plugin directory as scripts/eslint-rules/, which does not exist. The real location is eslint-rules/ at the repo root -- what eslint.config.mjs actually imports -- and CONTEXT.md's own later entry already says so explicitly, so the file disagreed with itself. Found by a line-by-line audit of all 1036 lines against the live graph; this was the only confirmed inaccuracy. Closes #2890 --------- Co-authored-by: Test <test@example.com> |
||
|
|
7372d99a26 |
enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales The reviewer lane roster was hand-enumerated across five documentation surfaces and three workflow files that had drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md. Adds checkReviewerDocsParity, a second pure gate deliberately separate from checkReviewerLaneParity so a stale doc cannot make the runtime checker look red. Workflows now derive their flag lists from a new review-lane flags query instead of hand-enumerating them, which also retires the unanchored grep that matched --agy inside --antigravity. Documents the previously absent reviewer body and hostBehaviors field in the capability manifest reference. Closes #2800 Closes #2781 Closes #2272 * fix(#2800): key the docs parity table arm on first-cell position Review found the flag arm was file-scoped, so the forwarding row that lists every flag in its third cell satisfied it on its own. Deleting a lane's own reviewer-table row -- the #2781 regression this gate exists to prevent -- therefore passed undetected. Arm 4 keys on the FIRST table cell, which separates a lane row from the forwarding row structurally and in every locale. Regression test included. * fix(#2800): shape-filter the flags subcommand output All three consumers read review-lane flags through an unquoted command substitution so the output word-splits into loop items. Phase 2 admits third-party overlay lanes, so an overlay flag containing whitespace would inject a second loop item and one containing a glob would expand against the cwd. Emit only well-formed flags so neither reaches the shell. * fix(#2800): remove the regex length ceiling and count only prose mentions Review found two real defects in the docs parity gate. The never-throws contract was false: building a RegExp from a declared flag or section title throws SyntaxError past ~100k chars, and Phase 2 admits overlay lanes whose declared strings are untrusted in length. Every one of these matches is literal, so String.includes replaces the regex outright, which also deletes escapeLiteral and the llama.cpp escaping it existed for. Arm 1 was context-blind: a flag mentioned only inside a fenced example or a commented-out row counted as documented. Both are stripped before matching. Also advertises all 13 lane flags in the argument-hint and corrects a stale eleven-lane count in the slug grammar note. * test(#2800): repoint the convergence suite off deleted workflow text The derived flag loop deleted the literal per-flag grep lines four tests matched on. Two of those failed loudly. The behavioral and property tests failed SILENTLY instead: their end marker no longer resolved, so the parse block extracted empty and both passed vacuously, and the property test's gsd_run stub had a no-op default that hid it. All now share one extractor and execute the real deployed block through a gsd_run shim backed by the actual binary. The whitelist assertions become an anti-parity check: re-adding a hand-written flag list must fail. Also repairs two vacuous cases in the docs parity suite. The unreadable-doc test called its own mock rather than the reader, and the integration test bounded nothing, so a doc losing its marker would have been silently skipped and still passed green. * fix(#2800): run the derived flag loop after the launcher preamble The remote matrix caught a real runtime bug, not a test artifact. In autonomous.md and plan-review-convergence.md the launcher preamble that defines gsd_run lives in a separate, LATER bash fence than the derived loop. Each fence is its own shell, so gsd_run was undefined where the loop ran: the command substitution yielded nothing and zero reviewer flags would have been forwarded. Worse than the drift this epic fixes, and silent. The whole CONVERGENCE_ARGS construction moves as one unit, because the --max-cycles append sits between the loop and the preamble and would otherwise have run against an uninitialized variable and then been dropped by the relocated initializer. Also documents all 13 lane flags in help/modes/full.md, which the repo gates bidirectionally against each command's argument-hint. * test(#2800): repoint the two converge suites off deleted flag literals Both asserted workflow.includes('--codex') against the hand-enumerated list the derived loop removed. They now assert the derivation itself, keep --all and --text (convergence controls, still literal), and add an anti-parity guard so re-adding a hardcoded list fails. The lost pass-through proof is replaced with a real one: every flag the tests used to hardcode is asserted present in the actual roster emitted by the binary, which is the property the old assertion was protecting. * test(#2800): acknowledge the workflow byte growth from the derived flag loop * chore(#2800): backfill changeset pr number to 2882 * fix(#2800): strip HTML comments to a fixed point in the parity gate CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick construction smuggles a commented-out row past the gate and it counts as documented. Not an injection risk here since nothing is rendered, but it is the exact false pass this helper exists to prevent. Strips to a fixed point, then treats any surviving opener as unterminated so the multi-line branch closes it on a later line. Terminates because every pass strictly shortens the string. * test(#2800): pin the comment-smuggling regression with a real reproducer The obvious fixture for this class does not reproduce it: <!--<!---->--> leaves a dangling --> rather than a live <!--, and is caught either way, so it would have passed with and without the fix. The join-trick construction (<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely regresses on the single-pass strip and is what the test now uses. --------- Co-authored-by: Test <test@example.com> |
||
|
|
185da024cb |
fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block (#2881)
* test(#2770): empty contextPath argument must fail closed, not green-skip the decision-coverage gate The handler conflated empty-arg (caller error) with file-missing (legitimate skip), returning passed:true/skipped on an empty argument. Add: empty arg → passed:false; real-path-to-absent-file → legitimate green skip preserved; omitted arg → fail closed. * fix(#2770): decision-coverage gate fails closed on empty arg + workflow recomputes CONTEXT_PATH in-block Handler (check-command-router.cts): split the guard — empty/missing contextPath argument is a caller error (fail closed, passed:false, mirrors #1365); a real path whose file genuinely does not exist keeps the legitimate green skip. Workflow (plan-phase.md): recompute CONTEXT_PATH inside the consuming Bash block (it was set in the step-1 init block, which does not survive into the separately- spawned gate block — so the gate ran with an empty arg and silently green-skipped). * chore(#2770): changeset fragment * fix(#2770): guard workflow empty-glob case (review blocker) + update drift-guard test The handler now fails closed on an empty contextPath arg, so the workflow's unguarded glob (empty when a phase genuinely has no CONTEXT.md) would invoke the gate with an empty arg → passed:false → exit 1, hard-halting the legitimate 'Continue without context' plan-phase path. Guard the empty-glob case: only run the gate when a CONTEXT.md actually exists. Update the F1 drift-guard test (which gave false coverage — it only checked for the ${CONTEXT_PATH} token) to assert the in-block recompute AND the empty-glob guard. * fix(#2770): keep plan-phase.md under ADR-857 size cap + ack emitted drift + fix drift-guard window The workflow fix grew plan-phase.md past the ADR-857 phase-6 size cap (94519B) and triggered emitted-attribution. Condense adjacent §13a prose/JSON to offset (net +89B, under cap). Add tests/emitted-drift-ack.json acknowledging the residual growth. Widen the drift-guard test window (the gate invocation is now nested in the empty-glob guard, so the old 400-char window missed the glob recompute). * chore(#2770): backfill changeset PR number (2881) --------- Co-authored-by: Test <test@example.com> |
||
|
|
82ca13f5a5 |
fix(#2736): write current_phase_name from the transition intent, not the lossy prose round-trip (#2821)
* fix(#2736): intent-first current_phase_name on transitions; dash-first prose precedence Primary: completePhase (adapter) and beginPhase (via readModifyWriteStateMd options) pass the intent-held display name to syncStateFrontmatter as an authoritative override, applied after every derive/preserve/carry-forward step — so the lossy prose round-trip can never destroy a name the transition just resolved. Names containing a parenthetical (`Closer-ruling measurement (D1a)`) now land in frontmatter verbatim instead of collapsing to the parenthetical (`D1a`). Secondary (#1695 AC #3 residual): parsePhaseFromProse prefers the em-dash name when it is a genuine name (not a status keyword, not a `Milestone:` tail), else falls back to the parenthetical — satisfying both first-party writer shapes (`N — Name (aside)` and `N (Name) — EXECUTING`). Still lossy for paren-containing names, which is why the intent-first override is the primary fix. plannedPhase carries no name in its intent, so it is naturally out of scope. Fixes #2736 * docs(changeset): backfill PR number for #2736 fragment * fix(#2736): drop an unnecessary type assertion on result.data StateTransitionResult.data is already `Record<string, unknown> | undefined`, so the cast was a no-op and tripped @typescript-eslint/no-unnecessary-type- assertion (CI lint-tests red on the first push; every test lane was green). --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
4f6935e29b |
fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks (#2846)
* test(#2717): CommonJS marker for cursor/windsurf/codex staged .js hooks Cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) stage .js hook scripts via dedicated paths that bypass installSharedHooksBundle — the only writer of the {"type":"commonjs"} marker. Under a config root declaring {"type":"module"}, Node loaded those scripts as ESM and every require() failed with 'require is not defined', silently disabling the runtime's hooks. Adds regression tests (RED first, fix lands next commit): - parametrized cursor/windsurf/codex install asserts hooks/package.json exists with exactly GSD's marker content; - end-to-end: a cursor require()-using hook loads under a planted ESM-typed config root without the require-is-not-defined error; - the ensureCommonJsMarker / removeCommonJsMarkerIfGsdOwned contract: GSD markers are removed on uninstall, user-authored package.json is never touched. * fix(#2717): write CommonJS marker for cursor/windsurf/codex staged .js hooks The {"type":"commonjs"} marker lived only inside installSharedHooksBundle, which cursor/windsurf (skipSharedHooksInstall) and codex (!isCodex gate) never reach. Their .js hooks are staged by dedicated paths, so under a config root declaring {"type":"module"} Node loaded them as ESM and every require() failed with 'require is not defined', silently disabling those runtimes' hooks. Decouple the marker write into a shared helper so any code path that stages .js hooks can ensure it lands in the SAME directory as the scripts: - src/runtime-hooks-surface.cts: add ensureCommonJsMarker(dir) + removeCommonJsMarkerIfGsdOwned(dir) (byte-identical content to installSharedHooksBundle's marker; preserves a user-authored package.json on both write and uninstall). Call ensureCommonJsMarker(hooksDir) from writeCursorHooksJson + writeWindsurfHooksJson; call removeCommonJsMarkerIfGsdOwned on their matching remove paths. Export both. - bin/install.js: call hooksSurface.ensureCommonJsMarker after the codex hook copy; call hooksSurface.removeCommonJsMarkerIfGsdOwned in the generic hooks-removal loop (safe no-op where no marker exists). No change to which runtimes receive the shared bundle, the !isCodex gate, skipSharedHooksInstall, or kimi/kimi-code/cline/copilot/trae/zcode (all unchanged — audit in the diagnosis). RED @ dbb7d2bb (6 failures: 3 missing markers + the ESM require error + missing helpers); GREEN pending. * docs(#2717): changeset fragment (pr:0, backfilled post-PR) * chore(#2717): regen codex/cursor/windsurf install-tree fixtures + attribution ack The fix adds hooks/package.json to those three runtimes' install trees (the new CommonJS marker), so the golden install-tree fixtures gain one path each (regenerated via npm run gen:install-tree). emitted-attribution (ADR-2719) flags the 3 emitted hooks/package.json paths under the hooks-built rule; acknowledge them. Also drops 5 spent ack entries left by now-merged PRs (#2694 code-review.md/code-review-fix.md, #2695 worker/registry, #2794 review.md) — they are stale on this branch (base already carries them). * fix(#2717): codex ESM-root behavioral test + hooks-built provenance for package.json Two review-driven follow-ups on the #2717 fix: - Adversarial review noted the ESM-root behavioral test covered only cursor; refactor it into a helper and add a codex case (the !isCodex-gated path most likely to regress, whose marker write lives in bin/install.js). gsd-check-update.js require()s at module load, so it surfaces the ESM failure immediately. - emitted-provenance flagged hooks/package.json as 'attributed source does not exist' — the marker is code-derived (a fixed literal emitted by ensureCommonJsMarker at install time), not built from a tracked source. Route the hooks-built rule's sources/transforms for package.json to the surface source file, mirroring the existing .cmd-shim sub-family. * chore(#2717): drop now-redundant hooks/package.json attribution ack The hooks-built provenance routing (prior commit) now self-attributes the emitted hooks/package.json to src/runtime-hooks-surface.cts, which IS in this diff — so the attribution is self-explaining and the emitted-drift-ack entry became stale. Delete the (now-empty) ack file per ADR-2719's empty-file rule. * docs(changeset): backfill #2717 PR number to 2846 |
||
|
|
6a9babda69 |
chore(#2798): declare the eleven reviewer lanes as manifest data (#2837)
* chore(#2798): declare the eleven reviewer lanes as manifest data Phase 5a of epic #2782, delivering ADR-2782 D9 (roster half) and D3. - Five reviewers GSD never installs into become lane-only role:reviewer capabilities with no runtime body, no runtimeCompat and no install surface: gemini, coderabbit, ollama, lm-studio, llama-cpp. Before this they had no descriptor at all and lived as a hardcoded NON_RUNTIME_REVIEWER_SLUGS tail, which is now deleted outright. - The six hosts that are ALSO reviewers gain a reviewer body alongside their runtime body. Their runtime bodies are byte-identical to next -- verified per capability against the git blob, not asserted -- so no install behaviour moves. - KNOWN_REVIEWER_SLUGS derives from declared bodies via an exported deriveReviewerSlugs(registry). hostBehaviors.reviewerCli survives as a derived legacy alias for one release; where a capability carries both, the body wins and the slug appears once. Alias removal is Phase 7 (#2801). THE KEYSTONE: the roster is the SAME ELEVEN SLUGS as before -- antigravity, claude, coderabbit, codex, cursor, gemini, llama_cpp, lm_studio, ollama, opencode, qwen. This phase changes HOW the roster is derived, not WHO is in it, and the test asserts that literal list rather than a count. kimi-code is deliberately NOT declared here. It is net-new with no invoke_reviewers leg, so declaring it now would make it selectable but not invocable -- present in --all, selected, emitting an empty section for the whole 5a-to-5b window -- and would break Phase 1's parity assertion. It lands in 5b alongside the iteration that can run it. Legacy kimi (the Python CLI) is not a reviewer at all and gains nothing. The highest-value test is declaredManifestLanesMatchThePhase1Descriptor: it deep-compares all eleven declared bodies against REVIEWER_LANES field-by-field, including probe and invoke sub-fields. All eleven are byte-identical, key order included. The epic's premise is that the manifest and the core descriptor describe the same lane with NO translation layer, and Phase 2's review already caught one divergence that every other test missed. Two ADR corrections folded in, as Phases 1-3 each did: 1. PHASE ORDER. The ADR runs Phase 4 (federated config) before 5a and #2798 claims a dependency on 4. That is inverted and makes Phase 4 unsatisfiable: D9 assigns review.<host>_host to lane capabilities that do not exist until THIS phase creates them, and a federated config slice must live inside capabilities/<id>/capability.json. Real graph: Phase 2 -> 5a -> 4. 2. #2798's INVENTORY acceptance item is vacuous. The inventory catalogs bin/lib/*.cjs modules, not capability directories -- antigravity, opencode and qwen appear zero times in it -- and gen-inventory-manifest --check passes with the five new dirs and no edit. Also corrected a stale line in Phase 2's own ADR amendment: it recorded the slug pattern as /^[a-z][a-z0-9_-]*$/, but Phase 2's security review widened the shipped pattern to /^[a-z0-9][a-z0-9_-]*$/ to match Phase 1's exported LANE_SLUG_RE. The prose had not followed the code. Closes #2798 * fix(#2798): catalogue reviewer capabilities in the generated matrix The capability matrix rendered exactly two tables, feature and runtime, via renderTable(caps, role) filtering on c.role === role. ADR-2782 D3 added a THIRD role, so every role:"reviewer" capability was silently dropped from the first-party catalogue. The drift guard did not catch it, and could not: --check compares generated output against the committed file, and both omitted the five lanes identically, so it reported "up to date" while five shipped capabilities were invisible in the one document that is supposed to list what ships. A guard blind to an entire role is not guarding. This phase is what exposed it -- it ships the first role:"reviewer" capabilities -- so it is fixed here rather than deferred (CLAUDE.md: a defect found while working is fixed in the current change, which overrides one-concern-per-PR). Verified red-before-green: with a lane row deleted from the matrix, --check now exits 1; restored, it exits 0. Before this fix the lanes were absent entirely, so there was nothing for the guard to compare. Phase 6 (#2800) still owns enriching the matrix with lane-specific detail (slug/flag/transport columns) and the locale parity gate. This is the narrower fix: the capabilities APPEAR at all. * fix(#2798): close two hardening gaps and record three limits durably Isolated security review (5 targets, no blockers) reproduced two gaps in the new deriveReviewerSlugs. Both are unreachable through the checked-in registry -- it is generated, JSON-sourced and code-reviewed -- but the function is EXPORTED for reuse and carries no other validation, so it must not depend on its caller. - A whitespace-only slug passed the length>0 test verbatim and occupied a roster entry it could never match. Slugs are now trimmed before the emptiness test. A blank body correctly falls through to the legacy alias rather than DROPPING the lane, which would have been worse than the blank slug. - KNOWN_REVIEWER_SLUGS is computed at require() time, so an uncaught throw there breaks import for EVERY consumer rather than degrading selection. It is now guarded, yielding an empty roster on a malformed registry. That is a visible degradation, not a silent one: under D4 an explicitly requested reviewer that is unavailable is an ERROR, so /gsd:review --claude against an empty roster fails loudly. This also removes an asymmetry -- the sibling capability-trust module documents its collectors as TOTAL and wraps them for exactly this reason. Also records three findings that previously existed ONLY in squash-merged PR bodies, which is not a durable record: - ADR-2782 D5 gains an implementation note explaining why the resolved host is deliberately EXCLUDED from the disclosure signature. Rule 1 says consent binds the resolved host; the loader has no config resolver, so folding it in would make the loader and lifecycle compute different signatures for one manifest and re-prompt forever. The binding is split: signature covers the SHA-pinned manifest fields, the consent record stores the resolved host, and Phase 5b re-resolves at invocation -- which is where rule 4 already puts the check. A reader comparing rule 1 to the code would otherwise conclude it is unimplemented. - CONTEXT.md's capability-trust entry still described THREE executable surfaces. Phase 3 added the fourth and made that false; corrected here, since it is drift this epic introduced rather than Phase 6's new-glossary-term work. - stableJson documents the NaN/Infinity/undefined -> null signature collision and why it is unreachable (JSON grammar has no such literal, so JSON.parse throws first). Reachability rests entirely on the ingest path staying JSON.parse-only, so the note lives where someone would break it. * chore(#2798): backfill changeset pr number to 2837 |
||
|
|
8fc244b754 |
fix(#2702): workstream config-get inherits absent keys from root config (#2833)
* test(#2702): failing-first regression for workstream config-get root inheritance * fix(#2702): workstream config-get inherits absent keys from root config * fix(#2702): root inheritance wins over --default; gate on GSD_WORKSTREAM not path (review) * test(#2702): make GSD_PROJECT test exit-safe (--default sentinel, no error path) * docs(changeset): #2702 workstream config-get root inheritance * docs(changeset): backfill #2702 PR number to 2833 |
||
|
|
6229f0e55c |
fix(#2701): reject NUL-corrupted plan/state artifacts at the validator entry points (#2829)
* test(#2701): failing-first regression for NUL-corrupted plan/state validators * fix(#2701): reject NUL-corrupted plan/state artifacts at the validator entry points * fix(#2701): seed STATE.md in test (writeState); add NUL-path guards to validate/verify for parity (review) * docs(changeset): #2701 validators reject NUL-corrupted artifacts * docs(changeset): backfill #2701 PR number to 2829 |
||
|
|
69dbf28ca7 |
feat(#2796): reviewer lane as a fourth trust-disclosure class (#2826)
* feat(#2796): reviewer lane as a fourth trust-disclosure class Phase 3 of epic #2782, delivering ADR-2782 D5. A reviewer lane is piped the plan text, requirements, research findings and CONTEXT.md decisions, and its output is read back into REVIEWS.md -- an egress channel for the most sensitive artifacts GSD produces. Making lanes pluggable WITHOUT a disclosure class would open a data-exfiltration path behind a manifest field, which is why this gates the feature rather than following it. - discloseExecutableSurfaces was cyclomatic 51 / cognitive 99 / 110 lines with risk_level critical. Rather than grow it, it is now a short orchestrator over four extracted collectors (hooks, commands, mcp -- behaviour-preserving -- plus the new lane collector), each independently testable. That is also what makes the 80% mutation threshold survivable: 51 branches in one function cannot be mutation-covered by whole-function tests. - A spawn lane discloses its binary AND its full declared args, in rendered and raw form. Binary-only disclosure would be insufficient and not hypothetically: a lane declaring python3 with innocuous args could later change them to ['-c', '<program>'] without the binary changing. That is the bug class #1459 already fixed for MCP servers. - An openai-http lane has no binary, so it discloses the destination host and the config key naming it. A localhost destination is disclosed and distinguished from a remote one. Both forms name the egress payload classes. THE CONSTRAINT THAT SHAPED THE DESIGN: the lane element is appended to the disclosure signature ONLY when at least one lane is declared. signatureForManifest is the consent key both the loader and the lifecycle compare, so appending unconditionally would have changed every installed capability's signature and re-prompted every user for every capability on their next upgrade -- for a feature they do not use. Two pre-change goldens are asserted byte-for-byte as the tripwire. The resolved host is deliberately NOT in the signature. The loader has no config resolver, so including it would make the loader and the lifecycle compute different signatures for the same manifest and produce a permanent false-mismatch loop. It is disclosed and recorded instead; Phase 5b re-resolves and compares at invocation, which is D5 rule 4's own placement. reviewsSection and timeoutFloorMs are also excluded from the signature: a cosmetic change must not force re-consent, because a prompt carrying no security information is how users learn to click through. A lane's binary is NOT existence-checked against the staged bundle. It is a PATH tool, never a bundle artifact; treating it like a hook script would add every lane to missingArtifacts and block every lane install. Two defects fixed beyond the fourth class: - isLocalHostValue mis-parsed a scheme-less host: new URL('localhost:1234') does NOT throw, it reads 'localhost' as the URL scheme and yields an empty hostname, so a bare host:port would have been reported as non-local. Now falls back on an empty hostname rather than only on a caught throw. - The orchestrator's safeCollect closes a PRE-EXISTING totality gap in the other three classes: a null manifest, or one with a throwing getter or Proxy trap, previously threw out of disclosure -- which runs on an UNVALIDATED manifest at install time. No well-formed input changes; all 51 existing trust tests pass. Closes #2796 * fix(#2796): close four disclosure gaps found by the isolated security review All four were REPRODUCED by execution against the shipped module, and all four passed the existing 41-test suite while live -- each exists because the matrix did not think to ask. B (MEDIUM, reachable via plain JSON). Non-string argv members were folded into the consent SIGNATURE but dropped from the human-facing text, because the summary rendered the string-filtered args rather than the raw declared array. A manifest declaring args ['--json', 7, {mode:'exfiltrate-everything'}, true] printed as '--json' alone -- the host still receives the rest, so the user consented to a surface never shown. That directly contradicts this design's own Kerckhoffs claim that nothing about a lane is hidden. The summary now renders the raw array, with non-strings shown in a visible form, and never throws on a circular or BigInt member. F (MEDIUM, reachable). The [local] flag is design-load-bearing, and it was dropped for every loopback form except the dotted quad and the bare hostname. Bracketed IPv6 was mangled by splitting on the address's own colons ([::1]:8080 became '['), and legacy IPv4 encodings were not recognised at all. A browser, curl and the OS resolver all treat 127.1, 2130706433, 0x7f000001 and 0177.0.0.1 as loopback. isLocalHostValue now handles bracketed and bare IPv6, IPv4-mapped loopback, and inet_aton shorthand/decimal/hex/octal. The dangerous direction was already clean and is now pinned by tests: localhost.evil.com, http://user@localhost@evil.com and friends stay REMOTE. C (LOW). An empty reviewer body flipped hasExecutable true and perturbed the disclosure signature, producing a re-consent prompt whose only content was '(no binary declared)'. A prompt carrying no security information is the click-through-training harm this design explicitly refuses for reviewsSection and timeoutFloorMs; refusing it there and permitting it here was inconsistent. A body declaring nothing recognised is no longer a lane. The test is deliberately broad -- any ONE recognised field suffices -- because requiring specifically a binary, or specifically a slug, would let a lane declaring only the other slip through unconsented, which is the far worse failure. Pinned in both directions. D (LOW-MEDIUM). Disclosure runs BEFORE validation, so a mis-cased or unrecognised transport reaches this code. Keying on an exact string sent a lane that plainly declares a hostConfigKey down the spawn branch, printing '(no binary declared)' for a lane egressing to a live remote host, and left its resolvedHost blank -- which reads as 'no destination', the precise thing the design forbids. Both the collector and the summary now branch on the declared SHAPE, so such a lane discloses its key and either a resolved host or the explicit unresolved marker. Two further findings were reproduced but confirmed NOT reachable through the real pipeline and are recorded as known limits rather than fixed: a selective-throw Proxy blanking a whole lane, and NaN/Infinity/undefined colliding to 'null' in a signature. Every production manifest reaches disclosure through readManifestBounded's strict JSON.parse, which cannot produce a Proxy, a getter, a BigInt, a circular reference, NaN or Infinity. The 0/-0 sub-case IS reachable via valid JSON but is inert -- String(0) === String(-0), so a spawned process receives identical argv. 9 regression tests added (50 total in this file, up from 41). * chore(#2796): backfill changeset pr number to 2826 |
||
|
|
46e84d5e39 |
chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat (#2823)
* chore(#2795): reviewer manifest body + registry harvest, validation, forward-compat Phase 2 of epic #2782 under ADR-2782. Delivers D1, D2, D3, D7, D8 and the four Phase-1 vocabulary amendments (A1-A4). - VALID_ROLES gains "reviewer"; the reviewer body is admissible on role:runtime (a host that is also a reviewer keeps one manifest) and on the new role:reviewer (a lane that is not an install target). A reviewer body on role:feature is an error: declaring one is an assertion of lane-ness. - validateReviewerBody + validateLaneProbe + validateLaneInvoke: nine closed enums, a transport discriminator selecting mutually-exclusive invoke sub-shapes, bounded probes (D7), and outputArg required-iff outputChannel is file-arg and forbidden otherwise. - Absent-safe (D4.1): only `undefined` is absent. null/{}/[]/false/0 are malformed assertions and error. 39 of 39 shipped capabilities depend on this. - collectReviewerWarnings: an unknown field inside the body warns, never errors, so a forward-built manifest degrades visibly instead of failing the build. - D8 uniqueness (slug / flags / reviewsSection) lives in validateCrossCapability, so it is enforced at build time over first-party AND at load time over the merged first-party union overlay set, with first-party-wins falling out of the loader's existing ordering rather than a new provenance check. - Config harvest widened past the role==="feature" branch in both the generator and the ownership loop. The often-cited cause of the stranded reviewer config keys -- the runtime body forbidding feature-only fields -- is not the mechanism: `config` is not in FEATURE_FIELDS_FORBIDDEN_ON_RUNTIME. The cause is two harvest sites that never read it. Verified inert: no shipped capability declares config on a non-feature role, and the generated registry is unchanged. Three ADR corrections are folded in (Phase 1 set the precedent of amending in-phase): the misattributed config-stranding cause, D3's inverted profile-membership claim, and the specified capability folder names for lm_studio / llama_cpp, which would have failed the id kebab-case invariant. Closes #2795 * chore(#2795): collapse nine enum checks into one validateEnumField helper Standards-axis review findings, both applied: - Duplicated Code: the enum-membership + enumerate-the-members error shape repeated near-verbatim at nine call sites. Routing them through one helper makes "the error names its valid members" structural rather than a convention repeated nine times, where it would drift. That property is load-bearing until Phase 6 ships the prose reference, because these errors are currently the only documentation of the vocabulary. - Speculative Generality: the isReservedName() pre-check on every enum field was inert. A VALID_* set never contains __proto__/constructor/prototype, so membership alone already rejects them, and "must be one of: ..." is more actionable than "is a reserved name". The literal guards remain where they do real work -- the key-derived write sites in the registry generator and the claim() accumulator. The reserved-name test now asserts all three reserved names are rejected via enum membership, rather than one name via a branch that no longer exists. * fix(#2795): align lane slug grammar with Phase 1 and wire the load-time diagnostic channel Spec-axis review findings, both applied. (1) The slug grammar had diverged from Phase 1's core descriptor. Phase 1 exports LANE_SLUG_RE = /^[a-z0-9][a-z0-9_-]*$/ (leading digit permitted); the manifest validator required a leading LETTER. A slug the core descriptor accepts -- a model-named lane such as 4o-mini -- would have been rejected by the manifest validator, which is exactly the translation layer ADR-2782 exists to delete. It was inert only because all eleven shipped slugs begin with a letter, so nothing else would have caught it until a third party shipped such a lane. The grammar cannot be reduced to one definition: Phase 1's module compiles to gitignored build output, and capability-validator.cjs is a committed plain .cjs that must load on a fresh worktree before build:lib has ever run. That makes this the repo's DEFECT.GENERATIVE-FIX class, so the duplication now carries a parity assertion -- laneSlugGrammarMatchesPhase1Descriptor -- which compares both the source grammar and the accept/reject verdict for a shared input set, and fails if the two ever drift again. (2) collectReviewerWarnings had exactly one caller: the build-time generator, which only ever sees first-party in-repo manifests. The real third-party overlay loader never called it and ValidatorModule did not declare it, so ADR-2782 D4.3 -- an unknown field inside a reviewer body is ignored WITH A WARNING -- surfaced nowhere at runtime, which is precisely the case D4.3 exists for. loadRegistry now collects those diagnostics on the accept path, behind a typeof-guard (an older built validator without the function still loads) and a try/catch (ADR-1244 D2's never-crash contract outranks a diagnostic). They land in a NEW OverlayMeta.diagnostics field rather than OverlayMeta.warnings, because warnings records capabilities that were SKIPPED and a consumer treating every entry as inactive would mislabel a working lane. Covered end-to-end by overlayLaneWithUnknownFieldIsAcceptedAndDiagnosed, which drives a real global-scope overlay through loadRegistry and asserts the lane is accepted, produces no skip warning, and yields a diagnostic naming the field. * fix(#2795): make the reviewer validators honour their documented totality contract Isolated adversarial review finding (MAJOR), reproduced by execution. validateReviewerBody documents itself as "TOTAL: returns an array of error strings for ANY input and never throws", and the overlay loader contracts every validator to RETURN errors -- #1461 OVL-1 records a validator that THREW and would have crashed every consumer of loadRegistry. The contract was false at ten sites: JSON.stringify throws on a BigInt and on a circular structure, and every enum/scalar rejection path interpolated the rejected value into its own rejection message. Reading the value could throw too, before any message was built, via a throwing getter or a Proxy get/ownKeys trap. Not reachable through a capability.json today -- every ingestion path is a plain JSON.parse of file text, which cannot express any of those shapes. Fixed anyway: the contract is stated on an EXPORTED function, and a caller must not have to re-derive today's reachability analysis to know whether it holds. Two layers, because serialization safety alone is insufficient: - describeValue() renders any value without throwing, so messages stay useful (a BigInt now reads "got: 10n" rather than degrading to a generic fallback). - A structural try/catch around validateReviewerBody and collectReviewerWarnings makes the guarantee absolute rather than argued, covering read-time throws that fire before any message exists. The same review found the property test guarding this contract was FALSE CONFIDENCE, which is the more important half. fc.anything() at default constraints emits no BigInt, no circular reference, no getter and no Proxy -- 20,000 sampled draws produced zero of each -- so the test was named for a contract its generator could not reach. Even withBigInt is insufficient under whole-value fuzzing, because the defect needs an exotic value in a specifically NAMED field and random key names never land on one. The property is now field-targeted across all twelve reviewer fields, and a companion test enumerates the shapes fast-check cannot generate at all (BigInt, circular, throwing getter, symbol, function, null-prototype) across scalar positions, array-element positions, and read-time traps. Verified red-before-green: with the fix reverted both property tests fail; with it restored all 119 pass. * chore(#2795): backfill changeset pr number to 2823 |
||
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d626dbc6e3 |
fix(#1883): distinguish a permission/IO error from genuine emptiness in dir scans (#2802)
* test(#1883): failing-first regression for findContextMdIn / listMilestoneArchiveDirs swallowing EACCES
Adds failing-first regression tests proving an unreadable dir is currently
swallowed as empty/null instead of surfacing the permission error. Covers
EACCES, EIO, the unchanged ENOENT empty path, the array fast-path, and both
CONTEXT.md forms. listMilestoneArchiveDirs is exercised in-process via a new
_listMilestoneArchiveDirs test seam (the validate command runs in a subprocess,
so an fs monkeypatch in the test process cannot reach it).
* fix(#1883): distinguish a permission/IO error from genuine emptiness in dir scans
findContextMdIn (src/planning-workspace.cts) and listMilestoneArchiveDirs
(src/verify.cts) catch-alled every readdirSync error into the empty marker
(null / []), conflating a genuine ENOENT ('nothing there') with an EACCES/EIO
failure ('can't read this'). An unreadable phase dir was silently reported as
'no CONTEXT.md' (discuss/plan gates wrongly skipped context) and an unreadable
milestones/ dir as 'no archives' (active-milestone resolution / archived-phase
filtering misbehaved).
Narrow each catch to ENOENT only — keep the long-standing null/[] contract for
genuine absence (Hyrum: empty path unchanged) and re-throw every other error so
it propagates to the caller's existing try/catch. All six findContextMdIn
callers either pass a pre-read string[] (no readdir) or sit inside a try block
that already handles readdir failures; the two listMilestoneArchiveDirs callers
live in the validate command path where errors reach the command error handler.
Exposes a _listMilestoneArchiveDirs test seam so the permission-error path can
be unit-tested in-process (the validate command runs in a subprocess, so an fs
monkeypatch in the test process cannot reach the private helper).
* fix(tests): delete stale emitted-drift ack for gsd-phase-researcher.md
Pre-existing base-branch defect, not part of #1883: commit
|
||
|
|
1b41083220 |
enhance(#2151): probe interactive-control for loading and error states (#2575)
* feat(#2151): probe interactive-control for loading + error states The ui-consideration-probe taxonomy mapped interactive-control to only one consideration (long-text), so a control-only UI surface (e.g. a theme toggle) lifted no loading or error consideration — the verifier never asked what a control shows while its action is in flight or when it fails, and a spec omitting those states could PASS. Add 'interactive-control' to the loading and error entries' elements in UI_TAXONOMY so control-only surfaces are probed for in-flight and failure states. empty is deliberately excluded (a control is not data-bearing; empty would be Goodhart noise). No new categories, no cue-map change, no probe-core change — a widening within the closed shape-rooted 8 (ADR-550), an independently-versionable predicate- generator adapter (ADR-857). Regenerated the reference-doc coverage table and the 19 golden-install- parity fixtures (reference-doc hash). Regression test added first (RED: interactive-control yielded only long-text; GREEN: now error+loading+long-text). Closes #2151 * feat(#2151): add changeset for interactive-control loading/error coverage --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7e8f6a6d7d |
enhance(#2572): run the verify-summary artifact check against phase SUMMARYs (#2685)
* enhance(#2572): run the verify-summary artifact check against phase SUMMARYs (W025) The artifact<->git check has existed since the beginning but was only ever pointed at .planning/research/SUMMARY.md (new-project.md:1145, new-milestone.md:425). Phase summaries -- the ones that actually claim "I created these files" -- were never checked. - extract verifySummaryCore from cmdVerifySummary: same checks, lifted out of the output() wrapper so callers consume {passed, checks, errors} directly instead of shelling out and re-parsing JSON; cmdVerifySummary is now a thin adapter over it - validate.health gains advisory W025 per phase SUMMARY with missing files Advisory only: appends to warnings[], never touches status escalation beyond the channel's own warning semantics, the repair set, or readVerificationStatus. Resolves both open questions from triage: (a) commits_exist is deliberately NOT surfaced -- its hash pattern matches any hex-shaped token in prose, too loose to show a user; (b) a phase carries N per-plan summaries plus a legacy bare SUMMARY.md, so all of them are checked via the repo-wide filter. * chore(#2572): add changeset fragment * feat(#2572): move the SUMMARY artifact check to phase completion Responds to the #2685 review. Three substantive changes. Seam (Blocker 2). The check now runs in cmdPhaseComplete, the seam the issue body cited (src/phase.cts:~1745), not validate.health. That channel does exist: cmdPhaseComplete declares warnings[], populates it from the UAT/VERIFICATION pre-scan, and emits it. The cycle objection raised against the earlier deviation holds for state.cts only -- verify.cts has no transitive import path to phase.cts, so phase.cts -> verify.cjs adds no cycle (verified over every src/*.cts). Moving it also retires the retroactive firing across all historical phases: this fires once, at completion, for the phase being completed. Extraction (Blocker 1). Pattern 2 now excludes [ and ] from its path class. The SUMMARY templates prescribe a YAML flow sequence (key-files.created: [a.ts, b.ts]) and the label matches case-insensitively, so the class previously captured the literal [ and produced a candidate that can never exist on disk -- firing on healthy projects built from the templates GSD itself ships. Stripping frontmatter was the other offered remedy; measured across all three shipped templates it is a no-op on top of the exclusion, so it is not carried. Consequence named in-code: the key-files block still is not read, which needs a real frontmatter parse. Also narrowed to the noise classes confirmed in review -- globs, bare hostnames, and paths resolving outside the project are skipped rather than reported, and the containment guard the old comment claimed now actually exists. Budget (Majors 1 and 3). verifySummaryCore takes a checkCommits option; phase completion passes false, so the discarded git cat-file probes are not spawned at all. It also passes Infinity, so every referenced file is reported instead of the first two -- a summary listing twelve files of which nine are missing now says nine, not zero. The verb keeps its historical 2-file default. Tests (Major 2). The vacuous fixtures are gone with the health block. The replacements use /-bearing paths that genuinely extract, and each fix was mutation-checked: un-anchoring pattern 2, dropping the glob, hostname or containment filter, forcing commit checking on, and re-capping at 2 each fail at least one test. * docs(#2572): describe the phase-completion SUMMARY artifact check The W025 text under /gsd-health is withdrawn with the health seam; the check is documented where it now runs, under `phase complete` in docs/CLI-TOOLS.md. Both the docs and the changeset previously overclaimed: they said a referenced file not on disk is warned about, while at most two candidates per SUMMARY were ever examined. The cap is gone at this seam, so the claim now holds -- and the text states the limits that remain, rather than leaving them to be discovered: the key-files frontmatter block is not read, commit hashes are not resolved, and globs, URLs, bare hostnames and out-of-project paths are skipped rather than reported. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
0997d4f443 |
fix(#2620): inject the reference DispatchLogger on the live dispatch seam when observability is enabled (#2621)
* fix(#2620): inject the reference DispatchLogger on the live dispatch seam when observability is enabled The Command Routing Hub defaulted to createNoOpLogger and no caller ever injected createDefaultLogger, so GSD_AUDIT=1 wrote nothing and failed dispatches emitted no structured JSON to stderr — contradicting ADR-0174 §5/§6, CONTEXT.md's Dispatch Observability Module contract, and docs/CONFIGURATION.md. Inject the reference logger at both live createHub() sites, gated on the existing opt-in signal (newly exported isAuditEnabled). When observability is off no logger is injected, so the Hub keeps its no-op fallback and default output stays byte-for-byte identical. Enabling stderr-on-error unconditionally adds a second line to the --json-errors envelope that callers parse as exactly one JSON line, so that is deferred to its own increment under #2619. * chore(#2620): add changeset for the dispatch logger wiring fix * test(#2620): cover the phase seam and drop try/finally from the adapter test Two review findings from the #2621 round-1 review. The fix wires the logger at BOTH live createHub() seams, but only cjs-command-router-adapter was exercised. Adds a fail-first regression test for src/phase-command-router.cts:258 — verified RED against a tree with that hunk reverted (1 fail, exact assertion) and GREEN with it restored — plus a negative pin that no trace file appears when GSD_AUDIT is unset. The negative case passes pre-fix and is a pin, not fail-first. CONTRIBUTING.md:344 forbids try/finally inside test bodies; the new adapter test used it. Converted to the Pattern-2 t.after() form, switched to the centralized createTempDir helper, and removed the now-unused os require. * chore(#2620): scope the changeset to the activation path that actually ships The fragment claimed config.audit.enabled activates the audit trail. It cannot: both seams call isAuditEnabled() with zero arguments, so the config branch in _isAuditEnabled is unreachable from production, and src/config-schema.cts registers no audit key at all — a user setting it would be silently dropped. That string ships in the user-facing CHANGELOG. Scoped to GSD_AUDIT=1, which is what actually works. The missing schema key stays a disclosed deferred sub-defect on #2620. Also adds the (#2620) issue backlink the other fragments carry. * docs(#2620): correct the fork-leaked issue reference in the wiring comments Four files cited this fix as #26, the issue number from the fork where the change was first written. Upstream #26 is an unrelated closed SDK issue, and next already uses #26 with that meaning in src/validate.cts:17,29,42 and src/config.cts:474, so these references pointed somewhere real and wrong rather than merely dangling. Baked into permanent doc comments, they reach users compiled via the ADR-457 build-at-publish path. The changeset and tests/phase-command-router.test.cjs already cited #2620; this brings the remaining four files into line. Comment-only, no behaviour change. build:lib produces no generated drift. The rename is scoped to these four files so the pre-existing SDK #26 references in validate.cts, config.cts, health-validation.test.cjs and config.test.cjs are deliberately left untouched. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
16e59d0db5 |
fix(#2691): repair seven dangling references in the ADR corpus and contributor docs (#2692)
* fix(#2691): repair five dangling references in the ADR corpus and contributor docs
Found by the 2026-07-24 ADR corpus audit; each mechanism re-reproduced live
against next @
|
||
|
|
80778e2674 |
fix(#1881): report an unreadable ROADMAP instead of reading it as absent (#2729)
* test(#1882): stage one file per commit in the base-ref ancestry fixture CI failed on ubuntu-24 inside this test's setup loop, before any code under test ran: at commit 32 of 60 the index referenced a blob whose object write had not landed -- "invalid object ... for 'base-31.txt' / Error building trees". The loop staged with `git add .`, which re-stages every file already in the tree. Across 60 iterations that rehashes O(n squared) blobs -- roughly 1,800 stagings and 60 full index rewrites to add 60 one-line files -- and that churn is what the object store failed under. Each commit only ever adds a single new file, so staging that one path is equivalent and removes the redundant work entirely. Verified the loop still builds the intended history: 61 commits, git fsck clean. The fixture already carries a note from an earlier fix in this epic recording that it passed on ubuntu-22 and windows-24 and failed on ubuntu-24 for the same commit. That was a different stage -- fetch versus diff -- but the same lane and the same brittleness, so this is the second time this fixture's cost has surfaced as a red build rather than as a test failure. Not caused by this PR's change, which touches two configuration lists and cannot reach a scratch git repository in tmpdir. Fixed here rather than deferred, because the run surfaced it. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1881): prove an unreadable ROADMAP is indistinguishable from an absent one Failing-first. Encodes the issue's runtime repro: an unreadable ROADMAP.md makes getRoadmapPhaseInternal return the same null it returns for "phase not found", and getMilestoneInfo return the same {v1.0, milestone} it returns for a project with no roadmap at all -- so a permission or I/O fault reads as a brand-new project. Half these cases exist to hold the opposite line. getMilestoneInfo has no existsSync guard, so platformReadSync's null-for-ENOENT is converted to a synthetic Error carrying no errno, and that lands in the SAME catch as a real EACCES. Reporting unconditionally there would flag every project without a ROADMAP.md -- every brand-new project -- as corrupt. The absent case, the errno-less error, a non-string errno, unparseable content and a genuinely missing phase are all pinned silent. One case guards a decision rather than behaviour: an unreadable STATE.md alone must stay silent, because the inner catch that swallows it is deliberate and documented under the #2245 audit as an optional enhancement falling back to ROADMAP-only heuristics. Two more pin the invariant ADR-1411 names explicitly -- neither function may throw, because src/state.cts removed its own defensive try/catch on the strength of that guarantee. Assertions are on the frozen reason enum and the emission counter, never on diagnostic prose. Faults are injected by overriding the platformReadSync seam and restoring in t.after(), never chmod 0o000, which root bypasses. Adds the ROADMAP_UNREADABLE reason to the shared vocabulary as scaffolding; no call site emits it yet, which is what makes these tests red. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1881): report an unreadable ROADMAP instead of reading it as absent getRoadmapPhaseInternal returned null for a read failure exactly as it does for "phase not found", and getMilestoneInfo returned {v1.0, milestone} exactly as it does for a project with no roadmap -- so a permission or I/O fault presented as a brand-new project and workflows synthesised a blank phase or skipped requirement extraction with no signal. Both return values are preserved exactly, per ADR-1411's amendment: continuity is correct, the silence was the defect. Each catch now reports through the shared unusable-input seam that shipped with #1882 rather than a second copy of the same mechanism. The discriminator is the errno, and it is load-bearing in the silent direction. getMilestoneInfo has no existsSync guard, so platformReadSync's null-for-ENOENT is converted into a synthetic Error with no code that lands in the same catch as a real EACCES. Reporting unconditionally there would flag every project without a ROADMAP.md -- every brand-new project -- as corrupt. A genuine read fault always carries an errno; absence never does. The parse is regex over text and cannot throw, so nothing else reaches these catches. Neither function gains a throw. ADR-1411 names this explicitly: src/state.cts removed its defensive try/catch around getMilestoneInfo under the #2245 audit because it never throws, and two tests pin that. The inner STATE.md catch stays untouched and silent -- its fallback to ROADMAP-only heuristics is a deliberate, documented optional-enhancement path, not a fault. Where the fix belongs was the design question. platformReadSync does not leak: it keeps absent and unusable as two channels, exactly as an abstraction should. Both callers re-collapsed that distinction, so the fix is caller-side and the projection seam -- with roughly ninety other dependents -- is untouched. Closes #1881 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1881): admit the roadmap reason to the locked vocabulary The seam documents adding a reason as three coordinated changes -- the enum entry, the emitting call site, and the test that locks Object.keys(...).sort(). This PR made the first two and the lock caught the third, which is the whole point of pinning the key set rather than asserting each value exists. The roadmap suite no longer re-locks the full set. Two complete locks would mean two files to update every time a later phase adds a reason, and #1883 and #1884 are both going to. The canonical lock stays in the seam's own suite; the roadmap suite asserts only the value it introduces. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1881): resolve the roadmap path inside the try, not outside it Naming the file in the diagnostic required the resolved path in the catch, and the obvious way to get it was to hoist `path.join(planningDir(cwd), 'ROADMAP.md')` above the try. planningDir throws a plain Error for an invalid GSD_WORKSTREAM or GSD_PROJECT segment -- one containing a slash, backslash or `..` -- so hoisting it let that throw escape uncaught. That broke the exact invariant ADR-1411 names as this file's hazard: src/state.cts removed its defensive try/catch around getMilestoneInfo under the #2245 audit because that function never throws. Of its callers only archivePhaseDirectories wraps it; cmdInitExecutePhase, cmdInitNewMilestone, cmdInitMilestoneOp, cmdInitManager, cmdInitProgress, cmdProgressRender and cmdStats all call it bare, so a workstream name with a slash in it crashed the CLI outright instead of degrading. The previous commit asserted "neither function gains a throw -- two tests pin that". That was false. Both tests inject faults through platformReadSync only and never through planningDir, so neither could have exercised the path that broke. The guarantee was claimed, not demonstrated. The path is now declared before the try and resolved inside it, so the catch can still name the file when there is one, and a path that never resolved reports nothing and returns the sentinel unchanged. The two test names are narrowed to what they actually prove -- that a failing READ does not throw -- and a new case injects the planningDir failure directly, which is what would have caught this. getRoadmapPhaseInternal carried the same hazard, resolving the path outside its try since before this branch. It is fixed the same way rather than left: ADR-227 is explicit that throwing breaks pipeline continuity, this read path already degrades to null for every other failure, and a PR whose purpose is hardening this invariant is the wrong place to leave the sibling crashing. Behaviour otherwise unchanged and re-verified: healthy lookups, EACCES reporting on both functions, absent-roadmap silence, and the errno discriminator all unaffected. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1881): backfill changeset pr number to 2729 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9a76ca6783 |
fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter
extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.
Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.
The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.
Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (
|
||
|
|
09477f925e |
fix(#2686): thread the resolved executor model into the Workflow backend (#2715)
* test(#2686): failing-first parity guard for Workflow-backend model threading The Workflow backend emitted every agent() call with no model, so model_overrides / model_policy / model_profile were silently inert on that path while the inline path honored them (ADR-1411). Neither existing suite contained the string 'model' at all. The centrepiece derives BOTH sides from resolveModelInternal(cwd,'gsd-executor') rather than hardcoding either, so it asserts backend parity rather than a fixed string. Also covers: omit-on-inherit/empty (#2517), byte-identical output when nothing resolves, the #2772/#2285 per-plan worktree gate, adversarial model ids reaching the code generator, the #2285 composed seam, CLI config-defaulting, and a fast-check round-trip property. RED expected: no model key is emitted anywhere, and --executor-model does not exist. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2686): thread the resolved executor model into the Workflow backend The Workflow backend emitted every agent() call with no model at all, so model_overrides / model_policy / model_profile_overrides / model_profile were silently inert on that path while the inline path honored all of them. The model was not dropped at the last step — it was absent from the whole seam: agentOptions() took no model, EmitInput had no field to carry one, and ResolveWaveDispatchInput (the #2285 seam the orchestrator actually calls) could not forward one. The generated script asserted the parity it broke. VERIFY-FIRST, which #2686 flags as the question that decides the fix: the Workflow tool's agent() DOES accept a per-call model. Its documented signature is agent(prompt, opts?: { label?, phase?, schema?, model?, effort?, isolation?, agentType? }) so fix branch 1 applies and branch 2 (declare model routing unavailable) is ruled out. ADR-1143:24's option enumeration omitting `model` is an incomplete enumeration, not a decision to exclude it. - agentOptions(p, executorModel) emits `model` only when it is a non-empty string that is not "inherit" (#2517: an empty model 404s on runtimes without native tier aliases). A non-string is a malformed config: omit, never throw. - executorModel threaded through EmitInput and ResolveWaveDispatchInput. - The CLI resolves gsd-executor from project config by DEFAULT rather than requiring a flag, reading the same source the inline path reads. An orchestrator that never learns about a new flag would otherwise silently keep the old bug. --executor-model exists only to pin/override. - ADR-1411 provenance: the generated header now states which model was applied, or that none resolved and why. A fallback must be a visible value. Compatibility: when nothing resolves, the emitted options object is byte-identical to before, so every existing caller and assertion is unaffected. Behavior change (Hyrum's Law): opted-in users move from session inheritance to the catalog-resolved executor model. Adding a `model` key also changes agent() opts, which invalidates the cached prefix of any in-flight resumeFromRunId run — a one-time re-execution. Both disclosed in the changeset. Fixes #2686 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * fix(#2686): reject script-breaking model ids and share the emit predicate The isolated adversarial review found a BLOCKER in my own provenance comment, proven by execution (the emitted script exited 42 from an injected statement). U+2028/U+2029 are ECMAScript LineTerminators that END a `//` single-line comment in EVERY engine — the ES2019 change legalized them inside string LITERALS only. So quoteString (JSON.stringify) is sufficient for the `model: "..."` object literal but NOT for the `// model: ...` provenance line I added: a raw U+2028 in a model id closed the comment and made the rest of the line live top-level code. The value is reachable from `.planning/config.json` (model_overrides / model_policy), which `mapClaudeOverrideForRuntime` passes through verbatim on any non-claude runtime — attacker-influenceable in a cloned repo. `emitWorkflowScript` now rejects a string executorModel carrying any character in UNSCRIPTABLE_CHAR_RE — the same class `isScriptableIdentifier` already applied to phaseDir/runId, which is proof the codebase knew this hazard. Rejection is ok:false with a reason rather than a silent drop, and resolveWaveDispatch maps an emit failure to the inline backend WITH that reason, so the degradation is visible. A non-string stays on the existing defensive path (omit, never throw) — that is malformed config, not an injection attempt. Also from the reviews: - The predicate deciding "is this model emittable" was duplicated between the emission and the comment asserting it. Extracted to emittableModel() so a generated comment can never claim something the generator did not do — the exact failure class #2686 was filed for. - That predicate now trims and lower-cases before comparing, closing a real #2517-class gap: " " and "INHERIT" were previously emitted verbatim. - The adversarial test was pass-always against this very vulnerability — it asserted only that JSON.stringify appeared. Replaced with the real contract (rejection) plus an execution-level check that no LineTerminator survives into the comment. A raw U+2028 had also been committed into that test's fixture array where a tab was intended; both are now explicit \u escapes. - optionsOf in the test was /\{[^}]*\}/, which truncated at any brace a generated model contained — silently not testing what it claimed. Now brace- and string-aware. Stale-test corrections in tests/fix-2285-*: three assertions froze the exact options literal `{ agentType: "gsd-executor" }`. The object legitimately gained an optional additive `model` key, so they now assert the invariant they exist to protect (agentType present, isolation absent) rather than a frozen literal. The CLI-vs-pure equality test pins --executor-model on both sides; otherwise it compared a config-resolved CLI run against a pure call given no model. CONTEXT.md glossary updated for the changed emitWorkflowScript signature and the new rejection rule (CLAUDE.md: the glossary is a PR gate for core-module changes). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * test(#2686): fix the options extractor and model the rejection path Two defects in my own test helper, caught by the full matrix: - optionsOf anchored on /\(\s*\{/ — a '(' immediately followed by '{'. The emitted shape is agent("brief", { ... }), so that never matched and the helper returned an empty array, making every assertion over it vacuously true. It now anchors on agent( and takes the first balanced, string-aware {...} after it. - The fast-check property predated the security fix and asserted ok:true for any generated string. Strings carrying an unscriptable character are now rejected, so the property models the real three-way contract: unscriptable -> ok:false; trims to empty or 'inherit' (any case) -> omitted; otherwise -> emitted as the trimmed value. Verified locally against the built module: omit values clean, both plans carry the model on the parity path, property passes 500 runs at seed 42. Test file re-scanned for raw hazardous codepoints — zero; the U+2028/U+2029 cases are explicit \u escapes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * test(#2686): scope no-control-regex on the mirrored unscriptable-char class The class is the point of the assertion — those bytes are exactly what must be rejected — so the rule is disabled at that line rather than the class weakened. UNSCRIPTABLE_CHAR_RE is not exported from src/claude-orchestration.cts, hence the mirror. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso * chore(#2686): backfill changeset PR number Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DPq9ovaovP2UvSVLjD4Lso --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
27c2279a39 |
fix(#2617): project verification next_command onto the runtime's command surface (#2700)
* fix(#2617): project verification next_command onto the runtime's command surface `src/verification.cts` stored and synthesized hard-coded `/gsd:…` command strings with no runtime context, and `phase complete` relayed that raw field straight into its verification-blocked error. On a Codex project the suggested next step was `/gsd:execute-phase`, a surface Codex does not install — it installs `$gsd-execute-phase`. The colon form is wrong twice over: `runtime-slash.cts` documents that "the colon form is never emitted", so EVERY runtime — not just Codex — was being handed a deprecated shape. Fixed at the one routing seam rather than per caller: - The routing table now stores BARE command names (`execute-phase`), never a prefixed literal. A prefixed literal in the table is what leaked. - A single `projectNextCommand(bare, runtime, tail)` helper runs every return path through `formatGsdSlash`, preserving the argument tail (`01 --gaps`) untouched. An empty command stays empty, so "no next step" never becomes a bare prefix. - `readVerificationStatus` accepts `opts.runtime`; `cmdVerificationStatus` and `phase complete` pass `resolveRuntime(cwd)`. The default is `claude`, which yields the canonical `/gsd-` hyphen form. All four routed states are covered: missing, unknown, gaps_found, stale. `init.cts` keeps its own projector deliberately. It already formats correctly, and its command CONTENT differs from the router's on purpose (it appends the phase number to `execute-phase`, and routes `human_needed` to `verify-work`). Consolidating them would silently change `init`'s user-visible output, which this issue did not ask for — so the divergence is left intact and the new tests instead pin the property that matters on both surfaces: no raw colon form escapes. Failing-first record: `origin/next:src/verification.cts` carried the four `/gsd:` literals (lines 101, 108, 382, 392), and 11 existing assertions in tests/verification-status.test.cjs asserted the colon form. Those 11 are corrected in this commit — they passed before the fix and fail after it, which is precisely the regression this closes. Tests are folded into the module's primary suite rather than added as a third file (`lint-test-file-count` caps the `verification` module at two, and consolidating is its documented remedy — growing the allowlist is not). The `phase complete` assertion reads `res.error`, not `res.stderr`: `runGsdTools` exposes a clean non-zero exit's stderr as `error`, and reading the wrong field yields '' and makes the whole check vacuous — which is how this user-visible path stayed untested. Closes #2617 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2617): scope the new hooks to their describes; cover gaps_found through the CLI Two findings from the orthogonal review of the first commit, both in the tests this change added. 1. The folded block's `beforeEach`/`afterEach` were declared at MODULE scope. node:test applies module-scope hooks to every test in the file, so hooks added for the #2617 suites also wrapped the ~40 pre-existing tests in verification-status.test.cjs — making an unrelated block a single point of failure for them (currently benign, but a throwing hook would have failed suites it has nothing to do with). They now install inside their own describes via a small `useProjectionPhaseDir()` helper, with a comment recording why. 2. The live-CLI `phase complete` test exercised only the `missing` state, so a regression in any other routed branch would have shown up in the router's return object but not in the text a user actually reads. Added a `gaps_found` case per runtime, asserting the projected `plan-phase <N> --gaps` reaches the blocked-completion error. Whole file verified green: 48 tests, 48 pass — the ~40 pre-existing ones included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2617): correct the last colon-form assertion in phase.test.cjs The remote run surfaced one more stale assertion outside tests/verification-status.test.cjs: the `phase complete` canonical-gate suite matched the blocked-completion message against `/\/gsd:verify-work 0?1/`. That project fixture configures no runtime, so it takes the `claude` default, which now yields the canonical `/gsd-verify-work 01` hyphen form. The colon form this asserted is exactly the deprecated shape #2617 removes — `runtime-slash.cts` documents that "the colon form is never emitted". Like the eleven corrected in the first commit, this assertion passed before the fix and fails after it, which is the regression record rather than a test being loosened: the surrounding assertions (failure reason, `stale` wording, and that neither ROADMAP.md nor STATE.md was mutated) are untouched. Verified against the real CLI: the emitted message is now "Phase 1 verification is incomplete: Verification is stale. Re-run verify-work before transition. Next: /gsd-verify-work 01". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2617): collapse the two verification projectors into one seam The orthogonal review found that `init.cts` carried a second, independently maintained `verificationNextCommand()` that had drifted from the router's table in CONTENT, not just formatting: state router (before) init.cts missing execute-phase execute-phase <N> unknown execute-phase execute-phase <N> human_needed "" (no command) verify-work <N> The `human_needed` row is the sharp one: two GSD surfaces disagreed about whether a next command existed at all, and the router's own next_action told the user to "re-run the verify step until status is passed" while naming no command to run. init's answers were the useful ones, so the router adopts them and init now delegates to it — satisfying the issue's "keep one verification-routing seam" direction. `verificationNextCommand()` is deleted. Appending the phase number surfaced a trap the old bare commands hid. `extractPhaseToken` also returns project-code forms (`PROJ-07`), which are indistinguishable by shape from an ordinary directory name — `gsd-651-parent` yields `gsd-651` — so deriving the argument blindly emits `execute-phase gsd-651`. The number is therefore appended only when it is unambiguously numeric, or when the caller supplies it explicitly. `init` does supply it: its `phaseDir` is unresolved in several branches, where the router could not derive one at all. dir `01-example` -> $gsd-execute-phase 01, $gsd-verify-work 01 dir `gsd-651-parent` -> $gsd-execute-phase, $gsd-verify-work Suites verified green against the built lib: verification-status 50/50, phase 268/268, init 143/143, init-manager 40/40. `npm run lint:ci` clean. User-visible change beyond the reported bug, as agreed: `query verification.status` and `phase complete` now append the phase number for missing/unknown, and emit `verify-work <N>` for human_needed where they previously emitted nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2617): backfill changeset PR number (#2700) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
28e486faf7 |
fix(#2608): fail closed when git add fails during commit staging (#2693)
* fix(#2608): fail closed when `git add` fails during commit staging `cmdCommit` ignored `git add` failures. #2523 had already stopped a failed path entering the commit pathspec, but skipping it silently left two bad outcomes, both reproduced against the pre-fix build: - SOME paths fail -> `{"committed":true}`. `git commit` still ran and PARTIALLY committed the subset that happened to stage, under a message describing the full requested scope. - EVERY path fails -> `{"reason":"nothing_to_commit"}`, which is not what happened and points the operator nowhere. In both cases git's original `add` stderr was discarded, so the user saw a downstream `commit_failed` / pathspec error naming an innocent file — the symptom reported in the issue from a linked worktree whose git directory was outside the managed writable root. Staging failures are now collected and the command fails closed BEFORE `git commit` runs, returning the issue's specified shape: { committed: false, hash: null, reason: "staging_failed", file: "<first failing path>", error: "<original git add stderr>", failures: [ { file, error, timed_out }, ... ] } A timeout is distinguished as `staging_timeout` (issue AC5) using the projection's SIGTERM+ETIMEDOUT signal — the same idiom worktree-safety.cts uses. The check is placed ahead of the `nothing_to_commit` branch so an all-paths-failed run reports the staging cause rather than an empty changeset. Unchanged: successful staging still commits exactly the declared scope and leaves unrelated staged files alone; an explicitly-named file that does not exist is still skipped rather than staged as a deletion (#2014/#2523), and a request where every named file is missing still reports `nothing_to_commit` — no `git add` ran, so there is no staging failure to report. Regression tests inject the failure by monkeypatching `execGit` on the projection module (per CLAUDE.md, over `chmod 0o000`, which does not fault under root and would make the tests vacuous), driven in a `node -e` child because `output()` writes via `fs.writeSync(1, …)` and cannot be captured in-process. Pre-fix, 6 of the 10 assertions fail; post-fix all pass. Closes #2608 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2608): roll back the index, guard the sibling surfaces, document the new reasons Six findings from the orthogonal review of the first commit, all fixed here. 1. A `staging_failed` return left the index PARTIALLY STAGED. The paths that did stage stayed in the index with no commit made and no cleanup, so the next bare `git commit` would sweep them up — the same silent partial commit this fix exists to prevent, deferred one step. (Pre-fix the partial state at least got consumed by the incorrect commit.) The staging failure path now resets the paths it staged, matching cmdPrSubrepo's established rollback-then-error convention. The reset is scoped to what THIS call staged — paths the caller had already staged are captured up front and excluded, so a caller's own work is never destroyed — and is best-effort, since an unwritable index (the very failure being reported) cannot be reset either. 2. `cmdCommitToSubrepo` still had the identical defect: a failed `git add` was dropped silently and the function committed the subset that happened to stage, discarding git's stderr. It now fails closed per sub-repo with the same staging_failed/staging_timeout reasons and the same scoped rollback. 3. The `git rm --cached --ignore-unmatch` branch (default mode, for a planning file that no longer exists on disk) still discarded its result. It mutates the index exactly like `git add`, and `--ignore-unmatch` already makes "no such path" a success, so a non-zero exit there is a real I/O failure — now routed through the same staging-failure path. 4. `agents/gsd-executor.md` documented the commit envelope as an exhaustive three-shape enum and pattern-matched only `nothing_to_commit | commit_failed`. It is the sole consumer doc for this surface, so the new reasons are added with explicit guidance not to retry (a retry hits the same unwritable index), and the "one of three shapes" framing is corrected. 5. The default (non---files) staging path and `--amend` are now covered by tests. Both were already guarded by the first commit but unexercised. 6. The changeset framed the fix as `--files`-only; it applies to default and sub-repo commits too, and now mentions the rollback. Regenerated the agent size baseline and the 18 golden install-parity fixtures for the gsd-executor.md edit. 16 assertions across both surfaces verified against the built lib. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2608): update the #2523 out-of-repo contract to the new staging_failed reason The remote test run surfaced this: `#2523: out-of-repo --files path is rejected by git` asserted `reason: 'nothing_to_commit'`, and now gets `staging_failed`. This is a deliberate contract improvement, not a papered-over failure. The old reason existed only because a failed `git add` was skipped and the resulting empty `stagedPaths` fell through to the empty-changeset branch. But "nothing to commit" is not what happened — the caller named a file and git refused it — and that misreport is exactly the class of defect #2608 closes. The result now carries the offending path and git's own message ("… is outside repository at …"), which is strictly more actionable for the same condition. #2523's two substantive invariants are untouched and still asserted: no commit is created, and the index is left clean. Two assertions are ADDED (the path is named, git's message is preserved) so the richer contract is pinned rather than merely allowed. Per CONTRIBUTING, a stale-test correction rides its own commit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2608): compact the executor doc addition to stay under the agent LARGE cap The remote test run failed: `gsd-executor.md is 49217 bytes — exceeds the LARGE hard cap of 49152`. The file was already at 48596 (556 bytes of headroom) and the new commit-envelope documentation pushed it 65 bytes over. The cap is a red line, not a budget to raise, so the addition is compacted rather than the cap moved: four lines instead of eight, keeping the load-bearing facts — the two new reasons, that nothing was committed and the index was rolled back, that `file` + `error` should be surfaced, and that retrying is wrong because a retry hits the same cause. Dropped only the restatement of the linked-worktree example (already in the changeset and PR) and the `failures[]` field (a superset of `file`/`error`, discoverable from the payload). Net addition is now 276 bytes; the file sits at 48872 with 280 bytes of headroom. Extracting the agent's shared boilerplate to references/ would buy much more, but that is a restructuring of the executor agent and does not belong in a commit-staging bugfix. Agent size baseline and the golden install-parity fixtures regenerated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2608): backfill changeset PR number (#2693) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3eb1cede26 |
fix(#1880): distinguish a corrupt config from an absent one (epic #1879 Phase 1) (#2688)
* test(#1880): prove corrupt config is indistinguishable from absent Failing-first. Encodes the issue's runtime repro: a trailing comma in .planning/config.json currently yields source:builtin-defaults with degraded:false - byte-identical to the file not existing - and the user's entire configuration is silently discarded. Asserts on the typed surface (CONFIG_REASON, _warnedUnusableConfig) rather than diagnostic prose, per the ADR-1411 amendment's test-methodology clause and CONTRIBUTING.md's raw-text-matching rule. IO failure is injected by monkeypatching fs.readFileSync and restoring in t.after(), never chmod 0o000 (root bypasses mode bits). Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): distinguish a corrupt config from an absent one loadConfigResolved wrapped the read, the JSON.parse and the entire config build in one try with one catch, so ENOENT, EACCES and SyntaxError all fell through to the same defaults and the branches returned degraded:false - actively asserting health over discarded configuration. A single trailing comma in .planning/config.json silently replaced the user's whole config, reporting source:builtin-defaults degraded:false, byte-identical to having no config file at all. ConfigResolution now carries a machine-readable reason. Genuine absence keeps degraded:false / not_configured; a file that exists but cannot be used sets degraded:true with config_unparseable or config_unreadable. The same split applies to the root config and to ~/.gsd/defaults.json. Control flow is deliberately unchanged. preflight_check reports cyclomatic 141 / cognitive 196 and 93 dependents on this function, with the guidance that small edits beat one big one, so faults are CAPTURED at the existing read sites and stamped onto the returns rather than the try/catch being restructured. Also carries the ADR-1411 amendment's wiring clause: loadConfig returns .config alone to ~51 call sites and would never see the new field, so an unusable file emits a deduplicated stderr diagnostic keyed on resolved path plus errno. Without it the reason would be an unreachable field and the user whose config was discarded would still get no signal - the actual defect. Registers the config-loader seam in lint-resolution-provenance, which until now guarded only agent-skills. Caller audit: ConfigResolution.degraded has exactly one consumer outside this module, cmdAgentSkills (src/init.cts:2259), which destructures {config, source, degraded} - adding a field does not break it. Its --json IR now reports degraded:true for a corrupt config, which is the intended fix and the one observable behavior change. Closes #1880 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): degrade when any config on the path is unusable, not just the last Two defects found by isolated adversarial review of the first cut. BLOCKER: the success-path return did not consult configFault. A corrupt ROOT config whose workstream override happened to parse returned degraded:false / reason:resolved - the root's settings silently dropped, which is the exact failure this issue closes, reappearing for any project using workstreams. The stderr diagnostic fired, so the out-of-band half worked while the in-band half reported a clean resolve; a --json consumer saw health. MAJOR: reason was derived from Object.keys(parsed) - the root+workstream MERGE - so an empty workstream file inheriting a non-empty root reported resolved despite carrying no settings. Emptiness is now judged on the file actually read, snapshotted before normalizeLegacyKeys mutates it. Also: corrects the ConfigResolution JSDoc, which still described the pre-#1880 degraded contract; adds a fast-check property asserting a PRESENT file is never reported not_configured whatever its bytes (CONTRIBUTING.md parser rule); and asserts the literal enum values so the provenance lint's configured_empty/not_configured markers check real assertions rather than incidental prose in test titles. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): reject valid JSON that is not a config object at the read seam The fast-check property added in the previous commit failed on both node lanes: a config.json containing 0, "str", [], null or true is valid JSON, so it parsed "ok", then threw downstream in normalizeLegacyKeys, and the outer catch reported not_configured - a PRESENT file reported as absent, which is precisely the collapse this issue exists to close. The property asserts a present file is never not_configured, and it caught it. _readConfigFile now validates shape, not just parseability (ADR-227: check the semantic shape at a trust boundary, not merely the type). A non-object JSON document is an unusable config, reported config_unparseable. Adds named regression cases for each non-object form alongside the property, so the class is documented and not only randomly sampled. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1880): backfill changeset pr number (pr:0 -> 2688) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2452): record a fetch-time shallow failure instead of crashing This guard failed CI on ubuntu-24 while passing on ubuntu-22 and windows-24 for the same commit, and passed on other PRs. Not a flake and not caused by the change under test - a real fragility in the test. runnerDiff ran the base fetch OUTSIDE its try and only guarded the diff, so it assumed the failure mode is always 'fetch succeeds, diff reports no merge base'. At a shallow boundary that lands short of the merge base, git can instead fail during the FETCH ('unable to parse commit' - the boundary commit's parent is not available). Which stage git fails at is version and transport dependent, so on some runners the error escaped runnerDiff and crashed the test rather than being recorded as the ok:false the assertions expect. Both stages mean the same thing for what this guard protects: a shallow base ref cannot resolve the three-dot diff. Also drops two assert.match calls against git's stderr prose. 'no merge base' and 'unable to parse commit' are the same condition reported at different stages, and CONTRIBUTING prohibits raw text matching on subprocess output. The typed outcome (ok === false) is the contract; the tests now assert that plus the presence of a cause. Found while investigating the red lane on #2688; fixed here per the no-defer rule rather than filed. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c07734216f |
fix(#2603): document kimi-code in the host-integration capability matrix (#2687)
* fix(#2603): document kimi-code in the host-integration matrix; correct 3 inherited axes The matrix — ADR-1239's deployment source-of-truth — had a section for 18 of 19 installed runtimes but none for `kimi-code`, so its `hostIntegration` axes shipped with no citation and no evidence quote. Sourcing every axis independently against Kimi Code CLI's own docs (the issue's explicit requirement — `kimi` and `kimi-code` are distinct products) showed three values had been inherited from the Python `kimi` descriptor rather than sourced: - `embeddingMode` imperative -> declarative. Kimi Code plugins are a `kimi.plugin.json` manifest plus markdown Skills with no in-process programmatic API (docs/en/customization/plugins.md) — the same shape as `codex`. - `dispatch.nested` false -> true. The `coder` built-in "can dispatch its own nested sub-agents when a task decomposes naturally" (docs/en/customization/agents.md). The Python `kimi` CLI genuinely prohibits nesting; Kimi Code does not. - `dispatch.maxDepth` 1 -> "undocumented". Nesting is documented but no depth bound is published, so the fail-closed sentinel applies over a guessed integer. `dispatch.namedDispatch` deliberately stays `false`: GSD's kimi-code artifact layout installs Agent Skills only (no `agents` kind), so no named GSD subagent is registered with the host and `resolveDispatchType` maps every role onto coder/explore/plan. Flipping it would reintroduce the dispatch failure recorded in docs/migration/kimi-to-kimi-code.md. The matrix records the host-capability nuance under Documentation gaps instead. Behaviourally inert: `namedDispatch:false` already caps nested/maxDepth/background/ backgroundDispatch to false/0 in the effective axes (host-integration.cts:493-499), and the install adapter is not selected by `embeddingMode` (install.js:543 always uses the imperative adapter). The one visible effect is the curated profile pin, which moves programmatic-cli -> declarative-cli. Also fixes the axes legend, which omitted the `built-in-only` subagentToolkit member that has been in the closed vocabulary since kimi-code shipped. Same defect class and countermeasure as #2598: pin the corrected values and require the matrix to agree with the descriptor, because a descriptor/matrix disagreement is how the gap survived. Closes #2603 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2603): report the maxDepth `undocumented` sentinel as a sentinel, not as malformed Surfaced by the orthogonal review of this change. `negotiateHostCapabilities` emits a sentinel-specific warning for every dispatch sub-axis carrying the documented `undocumented` value — namedDispatch, nested, background, subagentToolkit, backgroundDispatch, isolation — except `maxDepth`, which fell through to the numeric guard and reported `host dispatch.maxDepth is missing or not a number — treating as 0`. That message is indistinguishable from a genuinely malformed descriptor, so a correctly fail-closed descriptor reads as broken. Six shipped runtimes carry the sentinel here (antigravity, augment, opencode, trae, windsurf, zcode) and this PR's kimi-code correction adds a seventh, which is why it is fixed here rather than left in place. The numeric guard keeps firing for genuinely malformed values; both paths still degrade `effective.dispatch.maxDepth` closed to 0. Covered by three tests, including the boundary case that the sentinel carve-out must not swallow a real malformed value. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2603): backfill changeset PR number (#2687) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5633bb32f |
enhance(#2671): brand raw vs calibrated token types so double-application is a compile error (#2676)
* test(#2671): add failing-first brand-typing compile fixtures * feat(#2671): brand raw vs calibrated token types * refactor(#2671): hoist type-compile into a before() hook Two review responses: - The fixture compile ran in the describe() body, so it executed at collection time even when the block was filtered out, and a failed precondition collapsed eight independent assertions into one opaque describe-level failure. A before() hook is this repo's documented idiom and preserves per-test granularity. - parseTokensFlag now records WHY it returns an unbranded number: it validates the magnitude of --tokens, but the basis is decided by --calibrated, so branding here would be wrong for half its callers. The assertion belongs to cmdEstimateCheck, its only caller. * test(#2671): pin each brand diagnostic to its OFFENDING marker Adversarial review demonstrated that asserting only exactly-one-diagnostic- at-code-N is not airtight. Repairing a fixture's brand violation while injecting an unrelated error of the same code (a string passed as the budget argument) still yielded exactly one TS2345, so the fixture would have reported green while no longer testing its regression at all. Each bad-* fixture now routes its violating value through a const named OFFENDING, and the test asserts the diagnostic's start offset falls inside that node — located through the AST, so it survives reformatting and never pattern-matches source text. Replaying the proof-of-concept against the new assertion rejects it: the diagnostic lands on the budget literal, not the marker. Also corrects a doc comment that claimed the program type-checks all of src/; it covers phase-estimation.cts and its transitive dependencies. * chore(#2671): backfill changeset PR number (#2676) |
||
|
|
0d08c32048 |
fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable (#2681)
* fix(#2590): emit Workflow scripts the Workflow tool accepts; make the backend reachable Every emitted script was rejected. Four invalid constructs, the first fatal on its own, so the Workflow backend could never dispatch a wave: 1. no `export const meta = {…}` first statement -> whole script rejected 2. resumeFromRunId("<id>") -> "resumeFromRunId is not defined". It is a Workflow TOOL INPUT parameter, not a script function. The run id still reaches the caller via summary.resumeRunId, to pass as that input. 3. budget(<n>) -> "budget is not a function". `budget` is a read-only object { total, spent(), remaining() } fed by the caller's token directive; a script cannot set it. Recorded as intent in a comment. 4. parallel(agent(…), agent(…)) -> "parallel() expects an array of functions". Now parallel([() => agent(…), …]) — passing agent() results directly also started every agent eagerly, before parallel() could bound concurrency. The single-plan stage had its own branch with the same parallel() defect; both branches are now one array-emitting path. Waves also emit phase() calls whose titles match meta.phases exactly, so progress groups correctly. Two secondary defects kept the script from ever being REACHED — which is why this shipped undetected: 5. NOTHING resolved the Agent SDK version. The fragment claimed there was "no scriptable way" to introspect it and told callers to omit the flag, so gate 5 returned agent_sdk_version_unknown on every automated run while `capability state` still reported active:true. True for bash, false for Node: the router now reads the installed @anthropic-ai/claude-agent-sdk version, walking node_modules up the tree and reading package.json directly — require.resolve throws ERR_PACKAGE_PATH_NOT_EXPORTED because the SDK's exports map does not expose ./package.json. Precedence: explicit flag > GSD_AGENT_SDK_VERSION > installed. Fail-closed is preserved; an unresolvable version still declines to inline. A too-old SDK now reports the truthful agent_sdk_version_below_floor instead of unknown. 6. The runtime fallback was `--runtime > GSD_RUNTIME > 'unknown'`, diverging from the canonical `GSD_RUNTIME > config.runtime > 'claude'`, so any invocation without --runtime reported runtime_not_claude on an ordinary Claude project. Now delegates to runtime-slash.resolveRuntime. The fragment's `${AGENT_SDK_VERSION:+--agent-sdk-version "$AGENT_SDK_VERSION"}` snippet is removed rather than repaired: it was also shell-dependent — zsh does not word-split unquoted parameter expansions, so it collapsed to a single argv element, argValue() never matched, and the run failed into the same agent_sdk_version_unknown, indistinguishable from genuinely unknown. Auto- resolution removes the need for the construct entirely. Verified with the issue's own repro: no flags now reaches the version gate; an SDK above the floor yields backend:"workflow" with a script that parses as a real ES module. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2590): sync generated registry, repair sibling tests, reject duplicate wave ids Findings from the isolated review, all fixed. HIGH — gsd-core/bin/lib/capability-registry.cjs was stale, and `lint:ci` was already RED because of it. The registry embeds the fragment text INLINE, so the shipped/installed copy still taught the exact broken contract this PR fixes: the old `${AGENT_SDK_VERSION:+…}` bash line and the "OMIT the flag when unknown" guidance. Regenerated. (I had read `lint:ci` by grepping its output instead of checking its exit code, so I recorded a red chain as green — checking $? now.) HIGH — three existing tests asserted the OLD broken shape and would have failed CI; none was touched by the first commit: tests/fix-2285-claude-orchestration-wiring.test.cjs — matched resumeFromRunId("…") tests/claude-orchestration.test.cjs — .includes('budget(') tests/claude-orchestration-command-router.test.cjs — .includes('budget(') Each now asserts the corrected contract: the id/pool reaches the caller via summary, and neither construct is ever CALLED. Two sibling assertions had also gone vacuous — `.includes('resumeFromRunId')` still passed, but only because the new explanatory COMMENT contains that substring, not because anything is wired. Rewritten to assert the real property. MEDIUM — duplicate wave ids were never rejected. Plan-id uniqueness was checked within a wave, but nothing checked wave ids across waves. That was harmless before; it is not now, because each wave emits a `phase("Wave <id>")` call plus a matching meta.phases entry and the tool matches titles by exact string — two waves sharing an id would collapse into one progress group and misattribute the second wave's agents to the first. Rejected at validation, with tests either side of the boundary. MEDIUM — the fragment contradicted itself (its "Manifest construction" header still listed $AGENT_SDK_VERSION as orchestrator-built) and, more seriously, never told the orchestrator to pass summary.resumeRunId as the Workflow tool's resumeFromRunId INPUT. Since this PR moves resume from a broken in-script call to a tool-invocation input, an implementer following only the fragment would have silently regressed phase-resume to a no-op. Both fixed. MEDIUM — docs/how-to/enable-claude-orchestration-workflow-backend.md and docs/explanation/claude-orchestration-capability.md documented `resumeFromRunId("<id>")` and `budget(<tokens>)` as current correct output — teaching the bug as the feature. Updated to the real contract, including the required meta block and the thunk-array parallel() form. (The changeset is `Fixed`, so the docs gate exempts this; it is corrected because it is wrong, not because a gate demanded it.) LOW — the router's top-of-file comment still described the divergent `--runtime > GSD_RUNTIME > 'unknown'` chain as current, ninety lines above the fix; and inserting resolveInstalledAgentSdkVersion had orphaned resolveDetectionArgs' JSDoc above the wrong function. Both repaired. lint:ci now exits 0 (verified by exit code, not by reading output). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2590): backfill changeset pr number (#2681) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c3958018dd |
docs(#2674): amend ADR-1411 — corrupt is not absent (epic #1879 Phase 0) (#2678)
* docs(#2674): amend adr-1411 with the corrupt-is-not-absent house pattern ADR-1411 reasons only about a resolution miss. It is silent on input that is present but not usable, which is how five engine read paths (#1879) could fold an unusable input into the value meaning 'genuinely absent' without contradicting an Accepted ADR. Read together, ADR-1411 and ADR-227 converge and do not license throwing as the cluster's answer: ADR-227 requires malformed input to be coerced rather than propagated and carves out only genuinely-fatal fields, while ADR-1411 already permits a fallback provided it is 'a visible value, not a silent substitution'. The defect in these five sites is therefore not that they fall back but that they fall back invisibly. Records the pattern that follows: every current return value is preserved, and the cause is made visible in-band where the result already carries a provenance envelope, or out-of-band via a deduplicated stderr diagnostic where it returns a bare value it cannot extend. Throwing stays confined to ADR-227's genuinely-fatal carve-out, decided per call. Also names the per-applier caller audit and the lint-resolution-provenance registry gap. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2674): prove the warning-state reset misses the unknown-key dedup set The two existing cases in this suite only pass because each picks a key name no other case reuses, so neither can observe whether the reset the beforeEach calls actually runs. Failing-first: asserts the exported _warnedUnknownConfigKeys is empty after _resetRuntimeWarningCacheForTests(). It is not - the helper clears only _warnedConfigKeys despite documenting itself as resetting per-process warning state. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2674): reset the unknown-key dedup set with the runtime warning cache _resetRuntimeWarningCacheForTests documents itself as resetting per-process warning state but cleared only _warnedConfigKeys, leaving _warnedUnknownConfigKeys populated across cases. The suite that exists to test that set - 'loadConfig - unknown-key warning dedup' - calls the helper in beforeEach expecting exactly this, so the reset was a silent no-op for it; both cases passed only because each picked a key name the other never reused. Any later case reusing a key would have had its warning suppressed by leaked state. Found while amending ADR-1411, which names this dedup guard as the pattern five downstream PRs (#1880-#1884) will adopt - shipping the ADR without the fix would have propagated the footgun to each of them. Folded in here per CLAUDE.md's no-defer rule rather than filed. RED verified on 3c4895841 (test only, no fix): linux-node22 reported 'FAIL tests/config-loader.test.cjs - the documented per-process warning-state reset must clear the unknown-key dedup set too'. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): document src/ in the changeset-lint trigger list CONTRIBUTING.md presented the Changeset Required trigger list as bin/, gsd-core/, agents/, commands/, hooks/, sdk/src/ - omitting src/, which scripts/changeset/lint.cjs has in USER_FACING_PREFIXES. src/ is the TypeScript source of truth compiled into gsd-core/bin/lib/*.cjs, so it is the most-edited user-facing path in the repo and the omission sends any contributor who touches it into a CI failure the doc says cannot happen. Also documents that the lint reads GITHUB_BASE_REF, which only CI sets, so running it bare locally reports success without evaluating the branch. This PR hit exactly that: a local run said ok_fragment_present and CI failed fail_missing_fragment on the same diff. Found while opening this PR; folded in per the no-defer rule. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): add Fixed changeset for the src/ trigger-list and reset fixes Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2674): restore the round-2 review corrections to the amendment These edits were made in response to the second isolated review pass but never staged: later commits used targeted `git add <file>` for the test and the source fix, so the two markdown files stayed dirty and shipped nothing. The branch carried the round-1 text, including the ADR-227 misquote the reviewer raised as a blocker. Restores: the unconditional-diagnostic clause (ADR-227's GSD_DEBUG opt-in was never implemented, so citing it as the precedent was wrong), the dedup key, #1882 folded into the out-of-band mechanism instead of a fourth mechanism-less category, the narrowed caller-audit rationale, and the test-methodology clause. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c7c2fe3c2b |
fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd (#2680)
* fix(#2587): resolve cursor hook workspace from workspace_roots, not cwd gsd-cursor-session-start.js and gsd-cursor-stop.js both resolved the project as path.join(process.cwd(), '.planning', 'STATE.md'). Under the cursor-agent CLI, hooks are invoked with cwd set to the Cursor config dir (~/.cursor), not the workspace — so the lookup always missed. sessionStart could only ever emit the "no .planning/ workflow found" nudge and stop's verify-work reminder could never fire, even with .planning/STATE.md sitting in the workspace. Slash commands were unaffected, which is why only the hook layer looked blind. Both hooks already buffered stdin into `raw` and never parsed it; the payload's workspace_roots carries the real path. Multi-root was left open in the report ("first root vs any root"). Resolved forward: prefer the first root that actually carries .planning/STATE.md, so a workspace whose GSD project is not the first root still resolves — strictly better than first-root-only and identical to it in the single-root CLI case. Falls back to roots[0], then to cwd, keeping IDE behavior unchanged if the IDE ever invokes hooks from the workspace. The resolver is duplicated verbatim across the two scripts rather than shared via hooks/lib/: these hooks ship standalone, and a new hooks/lib/ file must be registered in the GENERATED installer's GSD_HOOK_LIB_FILES allowlist — the installer-omits-shipped-file class that yields MODULE_NOT_FOUND at runtime. Per CLAUDE.md "Generative Fix Divergence", the duplication carries a parity assertion so the copies cannot drift. Failing-first, demonstrated by direct invocation with cwd != workspace: pre-fix sessionStart -> "no .planning/ workflow found" stop -> {} post-fix sessionStart -> ".planning/STATE.md is present" stop -> reminder tests/fix-2587-cursor-hook-workspace-roots.test.cjs spawns the real scripts as child processes with a cwd lacking .planning/ and workspace_roots pointing at it. Boundary coverage on the roots array (0 / 1 / 2 entries), plus malformed-JSON fail-open, junk-entry filtering, the parity assertion, and a guard that neither script resolves .planning from cwd again. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): extend workspace_roots fix to subagentStart; keep cwd a candidate Three findings from the isolated review, all fixed. 1. MISSED SITE (high). gsd-cursor-subagent-start.js carried the identical defect at line 43 — its own header documents workspace_roots in the input schema, but it resolved .planning/ from process.cwd() anyway. Under the cursor-agent CLI that meant every Cursor subagent (planner, executor, verifier) started with "no .planning/ workflow found" and no phase context. The report named only sessionStart and stop; the defect class was wider. Verified pre-fix vs post-fix by direct invocation with cwd != workspace. 2. SEMANTIC NARROWING (medium). The first cut searched only workspace_roots and fell back to cwd solely when the array was EMPTY. So when roots were supplied but none carried .planning/ while cwd did, the hook reported absent — where the pre-fix code, which always used cwd, reported present. That contradicted the fallback's own stated intent of preserving IDE behavior. cwd is now a CANDIDATE in the search (`[...roots, process.cwd()]`), so the fix is a strict superset of both the old behavior and the CLI fix, never a narrowing. 3. STALE GOLDEN FIXTURES (high, would have failed CI). The golden-install-parity fixtures store a content hash per installed file; these three hooks appear in 13 of the 19 runtime fixtures. Regenerated via `npm run gen:golden` — the diff is exactly the three hook hashes in exactly those 13 runtimes. Tests extended: subagentStart resolution via workspace_roots; the stop hook's absent branch (previously only session-start's was covered); an explicit regression guard that a project at cwd is still found when roots miss; parity now asserts all THREE copies byte-identical; and the cwd guard sweeps the whole RESOLVING_HOOKS list so a future hook in this family cannot be left on cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * refactor(#2587): extract cursor workspace resolution to a shared hooks/lib module The duplicate-plus-parity-test approach was the wrong call. The reported issue named two hooks; a third (subagentStart) had the identical defect. That is the signature of a systemic problem, and three copies of a resolver guarded by a parity assertion is a divergence risk maintained by hand rather than a fix. hooks/lib/cursor-workspace.js is now the single implementation. All three Cursor hooks require it; none defines a local copy. Divergence is prevented structurally instead of by asserting three copies stay byte-identical. The reason duplication looked necessary was real, and is fixed properly here rather than worked around: Cursor sets hostBehaviors.skipSharedHooksInstall (#2089), so it never reaches the installer's bulk hooks/lib copy — it was the ONE runtime shipping these hooks WITHOUT hooks/lib (verified against all 19 golden fixtures: cursor had the hook scripts, no lib). A naive require would have thrown MODULE_NOT_FOUND at load, BEFORE each hook's own try/catch, wedging every session on precisely the runtime this bug is about. writeCursorHooksJson (src/runtime-hooks-surface.cts) now stages the hooks/lib helpers the staged scripts actually require, discovered by scanning their require('./lib/…') calls rather than a hardcoded name — so a future helper cannot be silently omitted. This is narrower than flipping skipSharedHooksInstall, which would wrongly pull in every shared hook. cursor-workspace.js is also added to GSD_HOOK_LIB_FILES so uninstall and the manifest manage it for the runtimes that do receive hooks/lib. Verified against a REAL install (runMinimalInstall, cursor/global): the helper is staged, and all three INSTALLED hooks resolve the workspace end-to-end from a cwd that is not the project. Also closes the review gap that the stop hook was excluded from the cwd-candidate regression loop — it now sweeps RESOLVING_HOOKS. The byte-parity test is replaced by a structural guard (every hook requires the shared module, none redefines it) plus a new install test asserting the helper is staged and the installed hook actually loads against it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2587): fail loud on a missing hook lib source; drop unsubstituted version marker Two findings from the installer-focused review. H1 — the staging step's `if (!fs.existsSync(libSrc)) continue;` silently defeated the very guarantee it was added for. Reproduced: delete hooks/lib/cursor-workspace.js from source, run the cursor install — it exits 0, prints "Done!", and ships the three hook scripts with an EMPTY hooks/lib/. The installed hook then throws `Cannot find module './lib/cursor-workspace.js'` at load, before its own try/catch, wedging every session — and nothing surfaces until a user hits it. The scan protected against a required-but-UNLISTED helper while leaving required-but-MISSING wide open (typo, bad rebase, an accidental delete). It now throws: a missing helper source is a packaging bug and aborts the install. M1 — hooks/lib/cursor-workspace.js carried a `gsd-hook-version: <placeholder>` marker that NOTHING substitutes: copyLibDir stamps .sh files only, and writeCursorHooksJson's staging applies just the colon-to-dash rewrite. Verified the literal was reaching disk on both the bulk (--claude) and Cursor (--cursor) paths. hooks/lib/git-cmd.js — the only pre-existing hooks/lib/*.js — carries no such marker, so this was newly introduced, not inherited. Marker removed, matching that precedent, with a note on why. (The explanatory comment deliberately does not spell the token out, or it would reintroduce the literal.) M2 — the require-scan regex demanded the exact compact form, so `require( "./lib/x.js" )` would silently fail to stage its helper and compound H1. Now tolerant of interior whitespace and either quote style. Regression test added for H1 — the reviewer confirmed the invariant had zero coverage repo-wide: a source tree carrying the hooks but no hooks/lib/ must make writeCursorHooksJson throw rather than produce a broken install. Re-verified end to end: the missing-source case throws, no unsubstituted literal ships, and the installed hook still resolves the workspace from a foreign cwd. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2587): backfill changeset pr number (#2680) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bd570618d4 |
feat(#2632): executor actuals and the closed estimate-calibration loop (#2672)
* feat(#2632): record executor actuals and close the estimate calibration loop * fix(#2632): calibrate against the raw projection so the loop converges * test(#2632): add closed-loop convergence guard and codify the feedback-loop rule * fix(#2632): pair calibration samples per plan; atomic write; amend adr * chore(#2632): backfill changeset pr to 2672 * fix(#2632): retry renameSync on transient windows errnos and clean up the temp |
||
|
|
89b673e40d |
feat(#2631): planner emits estimate and plan-checker surfaces the over-budget flag (#2670)
* test(#2631): failing-first planner estimate emission and over-budget surfacing * feat(#2631): emit plan estimate and surface the over-budget split recommendation * fix(#2631): extract sizing prose to references to fit planner and plan-phase caps * fix(#2631): move estimate check to plan-checker; fix template regex and caps * fix(#2631): restore ALWAYS split literal and keep gsd_run after the launcher preamble * fix(#2631): invoke estimate-check after the launcher preamble in plan-checker * fix(#2631): stop double-applying calibration; repair COMMANDS table and stale reference * chore(#2631): backfill changeset pr to 2670 * chore(#2631): backfill changeset pr to 2670 |
||
|
|
115433bba6 |
fix(#2539): anchor commit phase-token detection; drop silent wrong-branch switch (#2669)
* fix(#2539): anchor commit phase-token extraction to the phases/ segment; drop silent switch-to-existing cmdCommit auto-detected the commit's phase from --files with an unanchored `match(/(\d+(?:\.\d+)*)-/)`, which returns the leftmost digit-run-then-hyphen anywhere in the joined path. A project_code ending in a digit (PROJECT_V2) made `.planning/phases/PROJECT_V2-07-name/…` match the `2-` inside `V2-` before the real `07-` token, resolving phase 2. findPhaseInternal also searches archived milestones, so an existing archived phase 2 produced a real branch name and the silent `git checkout <existing-branch>` fallback switched the whole working tree onto the wrong branch in the same call that then committed. The extraction now anchors to the directory segment immediately under `.planning/phases/` (or `.planning/milestones/<v>-phases/`) and runs it through the existing project-code-aware extractPhaseToken helper — the single owner shared by the other 6 call sites — rather than introducing a fourth independent copy of phase-token-matching logic. The auto-switch keeps create-if-absent only (the #1278 intent: ensure the branch exists before the first commit on it); it no longer force-switches an already-checked-out working branch onto a different existing branch. Adds two regression fixtures: a digit-suffixed project_code + an archived phase whose number collides with the trailing digit (the silent-wrong-branch case), and a pre-existing phase branch that must not be silently switched onto. * test(#2539): assert non-silent warning; hoist execFileSync; normalizePhaseName guard Address orthogonal-review findings on the #2539 fix: - Spec AC2 ('an auto-checkout mid-commit must never happen silently'): the no-switch path now writes a 'Warning: resolved phase branch X already exists; committing on Y instead' line to stderr when checkout -b fails because the branch already exists. The regression test captures stderr via spawnSync and asserts the warning, so neither direction of the branching resolution is silent. - Spec AC3 ('reuse normalizePhaseName/extractPhaseToken/stripProjectCodePrefix'): the token-acceptance guard now runs the candidate token through normalizePhaseName and accepts it only when it normalizes to a numeric phase form, rather than the brittle 'token !== phaseDir && /\d/.test(token)' check that leaned on extractPhaseToken's undocumented dirName fallback. - Standards (Duplicated Code): hoist execFileSync/spawnSync requires to the top of the 'commit command' describe block instead of inlining them per test. * fix(#2539): build phase-token shape from PHASE_NUMBER_TOKEN_SOURCE (#2128 guard) The acceptance guard regex in detectPhaseNumberFromFiles was a hardcoded `/^\d+[A-Z]?(?:\.\d+)*$/i` — a literal re-derivation of the canonical phase-number grammar, which the #2128 phase-id drift guard (tests/phase-id-drift-guard.test.cjs) rejects unless sanctioned with a `// phase-id-owner:` marker. Build it from the single-owner PHASE_NUMBER_TOKEN_SOURCE export instead, so this read-side acceptance check cannot drift from every other phase-token reader. gsd-test reported this as 2 failures (linux-node22 + linux-node24) on the prior commit. * docs(#2539): backfill changeset pr: 2669 |
||
|
|
46ba02acde |
feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs * fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions * fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests * chore(#2630): backfill changeset pr to 2661 |
||
|
|
bf127b7d25 |
fix(#2565): route generate-claude-profile target through runtime policy (#2659)
* test(#2565): add failing regression for generate-claude-profile runtime target cmdGenerateClaudeProfile hardcodes .claude/CLAUDE.md (project + global), ignoring the runtime policy that #3163 wired into the sibling cmdGenerateClaudeMd handler. These tests pin the parity contract for project scope (codex -> AGENTS.md, env precedence, --output override, claude preserved), global scope (codex -> <CODEX_HOME>/AGENTS.md, claude preserved), and a divergence guard asserting both handlers agree on the project instruction path for the same runtime. Failing-first: all seven tests reproduce the bug on unmodified next. * fix(#2565): route generate-claude-profile target through runtime policy cmdGenerateClaudeProfile hardcoded .claude/CLAUDE.md for both project and global scope, ignoring the runtime-aware resolution that #3163 wired into the sibling cmdGenerateClaudeMd handler. The #3163 fix diverged when it did not propagate here, so /gsd-profile-user kept writing Claude instruction files on Codex installs (and other AGENTS-native runtimes: opencode, kilo, kimi, antigravity, copilot). Fix mirrors the proven #3163 pattern using existing policy primitives: - Project scope resolves through getProjectInstructionFile(runtime) and a non-claude runtime wins over a stale claude_md_path (AGENTS-native projects must never write to CLAUDE.md). - Global scope derives ~/.<config-home>/<instruction-basename> via getGlobalConfigDir + basename(getProjectInstructionFile), so codex lands at ~/.codex/AGENTS.md. Claude global is preserved byte-for-byte (no env-var drift beyond the prior hardcoded path). - GSD_RUNTIME env var takes precedence over config.runtime. A parity test asserts both handlers agree on the project instruction path for the same runtime, guarding against future re-divergence (CLAUDE.md 'Generative Fix Divergence' rule). * docs(#2565): add changeset fragment for generate-claude-profile runtime fix * docs(#2565): remove parenthetical product description from changeset The product-name purity guard (#1777) flags 'ProductName (description)' patterns in changeset fragments because the prose renders verbatim into CHANGELOG.md at release time. The original fragment had 'Codex (and other AGENTS-native runtimes)' in the bold header, which matched the banned pattern. Reworded to drop the parenthetical; also trimmed the per-runtime mapping (belongs in code comments, not changelog prose). * docs(#2565): backfill PR number in changeset fragment * test(#2565): isolate os.homedir() cross-platform in claude global test The 'global scope: claude runtime writes to ~/.claude/CLAUDE.md' test set only HOME to redirect os.homedir() at a tmpDir. On Windows, Node's os.homedir() reads USERPROFILE (not HOME), so the child process still resolved the real user profile and the path assertion failed (#2659 CI). Set USERPROFILE alongside HOME so the isolation holds on both POSIX and Windows. Production code is unchanged — it uses os.homedir() exactly as the prior hardcoded path did. |
||
|
|
ec681e3c21 |
fix(#2567): scope Paused At to ## Session + guard Last Activity date regression (#2660)
* test(#2567): add failing regression for stale state field overwrites buildStateFrontmatter extracts Last Activity and Paused At from the full STATE.md body via stateExtractField, which matches the first 'Field:' line anywhere. Historical archive sections containing stale field-shaped lines silently overwrite the correct frontmatter value on every sync, and because the poisoning line stays in the body it regresses on the next write. Same divergence class as Bug #2444 (which scoped Stopped At to ## Session but did not propagate). Failing-first: all three tests reproduce the bug on unmodified next (verified via the dedicated red run on the test-only commit). * fix(#2567): scope Paused At to ## Session + guard Last Activity date Two complementary fixes for the stale-archive-overwrites-frontmatter bug class, chosen per field semantics: - Paused At is a session field: scope extraction to ## Session (via the existing matchSessionSection helper), exactly mirroring the #2444 fix for Stopped At. A stale 'Paused At:' line in an archive section can no longer win over the current value. Falls back to full body when no ## Session. - Last Activity has no single canonical section (it appears in the preamble, ## Configuration, and ## Current Position across STATE.md layouts), so a section scope cannot reliably exclude archive copies. Instead guard the information-losing direction: when the body-derived date is OLDER than the existing frontmatter date, keep the existing value and description (preferNewerLastActivity). Applied at both the write seam (syncStateFrontmatter) and the read seam (cmdStateJson) so they agree. Non-date values pass through unchanged. A first attempt scoped ALL current-state fields to the body preamble, but that broke STATE.md variants where the fields legitimately live inside ## Configuration / ## Current Position (regressed 4 frontmatter.test.cjs suites). This minimal fix targets only the two fields the issue names. * docs(#2567): add changeset fragment for stale state field overwrite fix * docs(#2567): backfill PR number in changeset fragment |
||
|
|
6ee4349272 |
fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) (#2642)
* fix(#2537): extract offer_next step to references/ (~3.3KB headroom restored) * chore(#2537): backfill changeset pr to 2642 |
||
|
|
e4dd0cbdd5 |
fix(#2523): normalize --files to repo-relative; reject out-of-repo; gate push on git-add exit (#2638)
* test(#2523): absolute + mixed + out-of-repo --files paths * fix(#2523): normalize --files to repo-relative; reject out-of-repo; gate push on git-add exit * chore(#2523): backfill changeset pr to 2638 |
||
|
|
6ad30f74b6 |
feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler. harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run. Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard. Closes #2627 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
461c744c31 |
fix(#2522): fold wrapped success-criteria lines into their criterion (#2637)
* test(#2522): wrapped + blank-line success-criteria parse * fix(#2522): fold wrapped success-criteria lines into their criterion * chore(#2522): backfill changeset pr to 2637 |
||
|
|
4a66d62d10 |
feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver (#2625)
* feat(#2584): Phase 2 — worktree create verb + orchestrator-exec resolver
Phase 2 of the negotiated executor-isolation feature (ADR-1239 Codex-binding amendment). Two building blocks for `dispatch.isolation: orchestrator-worktree` hosts, both unconsumed — no scheduler wires them yet (that is Phase 3), so no runtime behavior changes.
worktree create verb (planWorktreeCreate / executeWorktreeCreatePlan / cmdWorktreeCreate in worktree-safety.cts, routed via routeWorktree in gsd-tools.cjs): validates the wave base, creates a bounded branch+worktree, records it in the run manifest reusing record-agent 4-field entry shape, returns the executor working directory. Bounded git (10s timeout, degrade-not-throw); all manifest read/parse/validate/dedupe precedes the single git side effect (no unmanifested-orphan on a bad manifest); timeout-only best-effort partial rollback (a clean collision-exit never removes a live peer worktree); fail-closed on bad base, unsafe leading-dash / .. inputs, and malformed/mis-shaped manifest.
resolveOrchestratorExec (host-integration.cts): pure descriptor->argv resolver reading the new runtime.orchestratorExec descriptor field (codex/opencode/kimi/kimi-code), fail-closed on missing/invalid shape. Validator (capability-validator.cjs) + a parity guard asserting every orchestrator-worktree host declares a resolvable orchestratorExec.
Adding the create route edits the installed gsd-core/bin/gsd-tools.cjs, so the golden-install-parity fixtures for all 19 runtimes are regenerated (npm run gen:golden) — the only changed hash is gsd-tools.cjs. CONTEXT.md glossary updated; capability-registry regenerated. Behavioral tests (worktree-safety + host-integration) incl. a fast-check property test and the parity sweep.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: rebuild tracked state-transition.cjs to match #2400 source
The tracked compiled artifact drifted from src/state-transition.cts: #2400 (commit
|
||
|
|
7e1c736a3e |
fix(#2556): rescue SUMMARY when cat-file reports absent (exit 128, not 1) (#2611)
* test(#2556): correct cat-file stubs to exit 128 + rewrite fail-closed tests to fail-open * fix(#2556): rescue SUMMARY when cat-file reports absent (exit 128, not 1) * chore(#2556): backfill changeset pr to 2611 |
||
|
|
ec7978c0b4 | feat(#2584): add dispatch.isolation sub-field, descriptors, validator + negotiation (#2604) | ||
|
|
bf9fe4630d |
feat(#2249): bracket phase-id core grammar — parse/render/toDir round-trip pair (epic #612 PR-1) (#2258)
* feat(#2249): bracket phase-id core grammar — parse/render/toDir + READING-B + guards PR-1 of epic #612 (ADR-612, in-tree at docs/adr/612-bracket-phase-id-convention.md). Adds the bracket-convention grammar INSIDE src/phase-id.cts — the ADR-2121 single canonical owner — as a pure, additive extension. The 17 locked exports and PHASE_NUMBER_TOKEN_SOURCE are untouched, and normalizePhaseName is byte-identical, so the PR-0 collision anchor (tests/adr-612-collision-characterization.test.cjs) stays green. New pure round-trippable model (ADR Decision 4): - PhaseId { project, milestone, phase, subphase?, plan? }. - parsePhaseId(input): accepts display `[GSD.02] 05.03-01`, dir/token `GSD.02-05.03-slug`, or bare `GSD.02-05`; rejects ambiguous non-bracket tokens (`02-04`, `05`) rather than guessing. The rejection lives ONLY in this new parser — normalizePhaseName and every legacy reader keep accepting those tokens unchanged (conservative default; no existing path gains a throw). - renderPhaseId(id) -> `[GSD.02] 05.03-01`; toDir(id, slug) -> `GSD.02-05.03-slug` with a slug guard that sanitizes path-traversal input. - getMilestoneFromPhaseId(phaseId, convention?): READING-B derives the milestone from the `[PROJECT.MM]` prefix, gated on convention === 'bracket' and returning the `vN.0` form (parity with READING-A). The optional parameter keeps the helper pure (no config read) and byte-compatible — every existing single-arg caller resolves to the unchanged READING-A body (ADR Decision 6). - extractPhaseToken(dirName, convention?): bracket dir branch GATED on convention === 'bracket'. A bracket dir `{CODE}.{MM}-{PP}` is string-indistinguishable from the legacy #2043/#1324 letter-prefixed-decimal family (`P0.3-2`, `P0.12-34`) whenever the code ends in a digit, so no string-only discriminator is complete — an ungated auto-detect silently reinterpreted legacy reads on this CRITICAL 6-caller helper. The explicit convention signal keeps every existing convention-less call site byte-identical (pinned by a #2043 numeric-tail characterization in tests/phase-id.test.cjs). - comparator: no new code — comparePhaseNum already orders the dot-decimal `PP[.SS]` tokens extractPhaseToken yields; milestone-qualified ordering is a PR-2 resolution concern (bracketQualifiedKey), not core grammar. - SENTINEL_RANGES / isSentinelPhaseId(phaseId, convention?): {0, 999} non-milestone guard; the bracket-prefix reading is gated the same way (an ungated read called `P0.0-foundation` a sentinel), legacy leading-int form unchanged. - BRACKET_PHASE_TOKEN_SOURCE (dot-or-dash `[.-]` sub-separator; deliberately more permissive than parsePhaseId — a read-tolerance source for PR-2, not the emit grammar) and PHASE_HEADING_PREFIX_SRC exported from the drift-guard-exempt owner so PR-2 builds every bracket read regex from the canonical source and check:phase-id-drift stays green stack-wide. The bracket project code follows the repo's config-validated `[A-Z][A-Z0-9_]*` grammar (not the ADR §1 illustration's `[A-Z]{1,6}`), so every project_code the config permits parses. parsePhaseId has no live callers in PR-1, so this grammar choice is forward-facing for PR-2 with zero PR-1 behavior impact. Tests: tests/adr-612-bracket-grammar.test.cjs (28) — ADR §3 example round-trips, full 5-tuple parse, READING-B (+ legacy-unchanged and sentinel cases), extractPhaseToken bracket ON/OFF, comparator ordering of extracted tokens, sentinel + slug guards, bare-token rejection, exported-source behavioral assertions, and two generative fast-check properties: render∘parse identity over well-formed displays, and the toDir/disk↔display bijection. Plus a #2043 numeric-tail characterization (single- AND multi-digit rows) in tests/phase-id.test.cjs pinning the convention-less reading byte-identical. The compiled gsd-core/bin/lib/phase-id.cjs is gitignored (ADR-457 build-at-publish) and rebuilt by CI, so it is intentionally not committed. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2249): changeset fragment for PR #2258 (docs-exempt: internal grammar behind flag) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2249): reject non-canonical phase-id input + harden toDir (review B1/M1-M3) PR-1 CHANGES_REQUESTED follow-up (epic #612, ADR-612 Decision 4). B1 (blocker): parsePhaseId accepted non-canonical input (unpadded numbers, over-padded numbers, multi-space separators, stray whitespace), so render(parse(x)) === x did not hold for every well-formed x as ADR-612 Decision 4 requires. Both branches now enforce canonicality by construction: parse permissively, rebuild the canonical string via the same emit path (renderPhaseId for display, a hand-rebuilt token for dir/token), and throw "parsePhaseId: not canonical" on any mismatch. The .trim() at the parser's entry is removed — the match anchors now reject leading/trailing whitespace outright, folding into the existing "not a bracket phase id" rejection. M1 (major): toDir only ever guarded the slug; project/milestone/phase/ subphase were interpolated unsanitized, so a hand-built PhaseId (a structural, not nominal, type) could smuggle a path-traversal segment onto disk. Every field is now validated against the exact shape parsePhaseId itself would produce before use. M2 (major): a slug that sanitized to empty (e.g. '!!!') left a dangling trailing hyphen in the emitted dir name. toDir now throws in that case. M3 (major): an all-digit slug (e.g. '2026') was string-indistinguishable from the dir-branch's plan tail, so it silently broke the disk<->identity bijection on read-back. toDir now rejects all-digit slugs. Nits: toDir now rejects a non-string slug instead of coercing it to the literal token 'undefined'/'null'; sentinel boundary tests added for milestones 1/998/1000 (SENTINEL_RANGES is the two discrete values {0, 999}, not an inclusive range — these were already correct, now locked by test). Test-first: every new assertion (concrete examples + fast-check mutation property for B1; concrete cases for M1-M3 and the nits) was written and confirmed red before the implementation changes, per repo TDD convention. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2249): reformat changeset body to house convention (review Mi2) The fragment added in ab26190a was a plain paragraph — no bold headline, no trailing issue reference. Reformat to the repo's `**Bold headline** — symptom/explanation. (#issue)` body shape (see e.g. .changeset/agile-pandas-dance.md, .changeset/fierce-pumas-gather.md). Uses (#2249), the issue every commit on this branch references, not the PR number already carried in frontmatter (`pr: 2258`) — the changelog serializer appends `(#{pr})` unconditionally, so a body also ending in `(#2258)` would double-render as `(#2258) (#2258)`. Verified the rendered bullet directly via parseFragment + serializeChangelog: it now reads `... (#2249) (#2258)`, matching the dominant convention across the other fragments (frontmatter pr = merged PR, body reference = originating issue). Also moved the docs-exempt marker back before the paragraph -> after it (matching the file's original order): the marker sits on its own line and is stripped before the body is used, but placing it first left a leading blank line in front of the bold headline once reformatted. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(#2249): widen property generators — 3+-digit numerics + subphase-pad mutation (re-review Minor 1/2) PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes two property-generator coverage gaps the reviewer flagged; no source change (src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs are byte-unchanged). Minor 1 (3+-digit numerics never exercised): numArb capped at 99, so no property fed a 3+-digit milestone/phase/subphase/plan through parse/render/ toDir despite CANONICAL_NUMERIC_RE's dedicated `[1-9]\d{2,}` branch. Widen numArb to 1–999 so the round-trip and disk↔display bijection properties both span 3-digit widths (pad2 passes ≥3-digit values through un-truncated with no leading zero, so canonicality still holds). Add a concrete regression pinning the reviewer's hand-traced example: '[GSD.100] 05' round-trips, renders, and toDirs to 'GSD.100-05-feature' without truncation. Minor 2 (no subphase-pad mutation): the B1 mutation-rejection property covered milestone/phase pad + whitespace mutations but never a subphase pad. Add unpad-subphase / overpad-subphase to the mutation set and a generated `includeSub` boolean that decides whether the canonical carries a `.SS` (forced in for the subphase mutations so there is always a `.SS` to mutate); non-subphase mutations keep their original no-subphase coverage. Non-vacuity verified against the compiled lib by temporarily probing each widened/new property and confirming it fails: round-trip counterexample ["A",100,1,…] and bijection counterexample ["A",1,100,…,"a"] prove 3-digit tokens are genuinely generated and reach the body; a no-op unpad-subphase mutation trips the mutated===canonical guard (counterexample ["A",1,1,1,false,"unpad-subphase"]), proving the subphase branch is reached with a subphase present. Probes reverted; numRuns unchanged. Gates: tests/adr-612-bracket-grammar.test.cjs 44 pass / 0 fail; `npm run test:unit` 1079 pass / 0 fail; `npm run lint:ci` exit 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2249): consume the #2232 continuation seam at the bracket token's slug-adjacent position (review Major) BRACKET_PHASE_TOKEN_SOURCE was a sixth continuation-recognition site that re-derived the grammar as an unbounded `\d+` literal instead of consuming PHASE_CONTINUATION_SEGMENT_SOURCE, re-opening the #2232 bug class on the bracket path: a PR-2 reader interpolating it over dir `PROJ.01-14-2026-photos-…` (a slug whose first word is a year) over-collected the token as `01-14-2026` instead of `01-14`. Interpolating the cap verbatim at every position was rejected on evidence: the bracket run is `MM-PP[.SS][-LL]` and only the LAST position is slug-adjacent. The exactly-2 cap at the others would under-collect ids toDir itself emits — `PROJ.02-105-slug` (3-digit phase) reads as `02`, `[GSD.02] 05.100` (3-digit sub-phase) as `05` — because CANONICAL_NUMERIC_RE admits `[1-9]\d{2,}` and `[GSD.100] 05` is a pinned regression. Those positions are delimiter- disambiguated (a required field separator; a dot a slug can never contain), not heuristically recognized, so they have no year collision to defend against. Upstream draws the same line for the same reason: core-utils/phase cap the paired PLAN component while the leading phase component stays unbounded. So the run is now positional rather than a free `(?:[.-]\d+)*` repetition, and each position takes the width its delimiter affords: leading unbounded, dash-1 and dot canonical, and the slug-adjacent dash-2 interpolating the single-owner seam. The accepted trade-off is #2232's policy verbatim: a PLAN ≥100 is out of the token grammar. Also derives CANONICAL_NUMERIC_RE from the new BRACKET_CANONICAL_NUMERIC_SOURCE instead of re-spelling it as a literal, so the emit-side gate and the read-side token source are one rule — the same single-owner discipline this fix is about. Behaviour-identical (the anchors make the source's `(?!\d)` guard redundant). Refs #2249 * test(#2249): pin the bracket/#2232 reconciliation — parity surface 6 + divergence gate + property (review Major) The comment block alone cannot hold the divergence: src/phase-id.cts is exempt from the #2128 drift guard by construction, so lint-phase-id-drift.cjs would not catch the bracket token source drifting from the seam. Per the Generative Fix Divergence rule, the divergence is pinned behaviorally instead. Surface 6 joins the existing #2232 parity gate rather than starting a rival one: the review named the bracket token source "a sixth continuation-recognition site", and continuation-grammar-parity.test.cjs is already the invariant-named home where the five #2043 sites agree with the owner on a shared width corpus. Surface 6 asserts the same contract at the bracket run's slug-adjacent position (`01-14-<seg>-photos-…`, mirroring surface 1 with the extra milestone level), so the bracket path now fails the same gate the other five do. A second block pins the DELIBERATE half — the wider canonical width at the delimiter-disambiguated positions, plus the accepted bound (a plan >=100 is out of the grammar). Without it, "unifying" bracket onto the exactly-2 cap would look like a cleanup rather than a regression. The generative property ties the READ side to the EMIT side metamorphically: for every id toDir can produce, BRACKET_PHASE_TOKEN_SOURCE must collect exactly that id's numeric run — no more, no less. It needed a new arbitrary: the existing slugArb generates one [a-z0-9] word and so can never produce the number-leading slug the collision requires. Probe-falsified, both directions (probes reverted): - reverting the source to the old unbounded `\d+` fails 8: the parity gate reports `"01-14-2026-photos-performance" collected "01-14-2026"` — the review's scenario verbatim — and the property shrinks to ["A",1,1,undefined,"100-a"]. - interpolating the seam at EVERY position (the rejected verbatim option) leaves the repro and parity green but fails the divergence gate `'02' !== '02-105'` and the property at ["A",1,1,100,"100-a"] (3-digit sub-phase), which is the evidence that a verbatim cap under-collects ids toDir emits. Width 2 stays green under both probes — the corpus agrees with the owner exactly where the old and new rules coincide, so the gate discriminates rather than merely mirroring the regex. Refs #2249 * docs(#2249): add the new phase-id exports to the CONTEXT.md glossary bullet (round-4 Major) * test(#2249): pin deterministic grammar boundary cases (re-review m1) PR-1 re-review follow-up (epic #612, ADR-612 Decision 4). Test-only: closes the m1 proof gap — the grammar's bounds were exercised only incidentally through the fast-check domain (1-999, [a-z0-9] slugs). No source change (src/phase-id.cts and gsd-core/bin/lib/phase-id.cjs byte-unchanged). Adds a deterministic boundary block (7 describe groups, +22 tests) pinning the compiled lib's CURRENT behavior — a proof gap, not a behavior gap: - m1.1 numeric-width 99/100/101 at milestone/phase/subphase/plan: parse (display + dir) -> render/toDir round-trip byte-equality. The plan position is identity-symmetric (parse/render accept 99/100/101) but toDir drops it (filename-surface dimension only). - m1.2 read-token width is POSITIONAL: BRACKET_PHASE_TOKEN_SOURCE absorbs 99/100/101 at milestone/phase/subphase (delimiter-disambiguated) but caps the slug-adjacent plan (dash-2) at exactly 2 digits — plan >=100 is out of the token grammar (#2232 seam). Pinned as asymmetry, NOT symmetry. - m1.3 leading-zero 007 -> not-canonical rejection at every position/form. - m1.4 slug abuse: parse DROPS a null-byte/control/unicode/emoji trailing slug (never stored, never mis-read as a plan) and rejects a line terminator; toDir's allow-list sanitizer collapses each to a safe [a-z0-9-] token or rejects sanitize-to-empty. - m1.5 absolute-path slug sanitizes (next to the ../../etc traversal test); an absolute-path project on a hand-built id is rejected by PROJECT_ID_RE; an abs-path string is not a bracket id; an abs-path dir slug is dropped to a clean tuple. - m1.6 whitespace-only -> not-a-bracket-phase-id. - m1.7 very-long input (10k) resolves promptly (ReDoS smoke, behavioral): garbage/partial-prefix throw; a 10k-char slug parses (dropped)/sanitizes. No accept-not-reject case is a src bug: parse never STORES an abusive slug (dropped from the identity tuple) and toDir independently re-sanitizes on emit, so the only slug reaching disk is allow-listed. Plan >=100 accepted by parse is the documented positional design (toDir drops the plan; the read-token caps it) — divergence pinned, not papered over. Probe-falsify: corrupted one assertion in each of the 7 groups (m1.4 both its parse-side and emit-side), ran -> 8 distinct named failures, reverted -> 66/66 green. Confirms every new group executes and can fail. Gates: tests/adr-612-bracket-grammar.test.cjs 66 pass / 0 fail; grammar + continuation-grammar-parity + collision-characterization + phase-id family 175 pass / 0 fail; `npm run lint:ci` exit 0. `npm run test:unit` is green except one pre-existing, unrelated env failure (npm-integrity-gate: a live npm-audit advisory in the production dep tree — reproduces with this change stashed; no package.json/lock change here). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
ae8526cd50 |
fix(#2429): scope Codex skills home override to --global only (#2553)
* fix(#2429): scope Codex skills home override to --global only The skills-kind home override (redirecting skills to $HOME/.agents) was applied regardless of scope. Gate it behind scope === 'global' so --local installs keep skills project-local under the config directory. Closes #2429 * docs(#2429): backfill changeset PR number (2553) |
||
|
|
2bcfaa2e27 |
fix(#2400): warn on planned-phase no-op + sync progress.total_plans (#2552)
* fix(#2400): warn on planned-phase no-op + sync progress.total_plans Bug A: When STATE.md Current Position has no recognized labels (narrative prose), emit a warning field so the workflow detects the no-op instead of continuing with stale state. Bug B: Sync progress.total_plans in the YAML frontmatter when a plan count is provided, preventing contradictory state between frontmatter (0) and body (actual count). This writes the explicitly-provided count, not a re-derivation from disk (#500 safe). Closes #2400 * docs(#2400): backfill changeset PR number (2552) |
||
|
|
1482dc5ce0 |
fix(#2366): scope parseCoverageMatrix to recognized coverage tables (#2551)
* test(#2366): regression tests for parseCoverageMatrix scoping bugs Bug 1: summary table outside matrix not parsed as data Bug 2: multi-section matrix with repeated headers parses correctly Bug 3: markdown emphasis on decision cell is stripped * fix(#2366): scope parseCoverageMatrix to recognized coverage tables Replace latching sawHeader with contextual inMatrix tracking that resets on non-pipe lines, preventing summary tables from being parsed as data (bug 1). Allow multiple headers for multi-section matrices (bug 2). Strip markdown emphasis from decision cells before validation (bug 3). Closes #2366 * fix(#2366): update representative-corpus test to expect correct behavior The test previously documented the known-buggy parseCoverageMatrix behavior. Now that the fix is in place, test against the expected correct output (expectedBlock, expectedCounts, expectedErrorCount) instead of the currentBuggyOutput snapshot. * docs(#2366): backfill changeset PR number (2551) |
||
|
|
77bf21b3a6 |
fix(#1995): widen worktree branch regex to accept agent-<id> namespace (#2548)
* test(#1995): regression test for agent-<id> branch namespace Add failing-first tests proving that normalizeCleanupManifestEntry and planWorktreeRecordAgent reject Claude Code's current agent-<id> isolation branches (only worktree-agent-<id> is accepted). Boundary tests cover both namespaces plus rejection cases. * fix(#1995): widen worktree branch regex to accept agent-<id> namespace Claude Code's isolation="worktree" branch naming changed from worktree-agent-<id> to agent-<id>. Widen the regex in all 7 locations from ^worktree-agent-[A-Za-z0-9._/-]+$ to ^(worktree-)?agent-[A-Za-z0-9._/-]+$ so both namespaces are accepted. Introduce a shared WORKTREE_AGENT_BRANCH_RE constant in src/worktree-safety.cts to prevent future drift. Closes #1995 * fix(#1995): update workflow guards, test assertions, and baselines Widen the branch-check regex in execute-phase.md and execute-plan.md. Update all test assertions that checked for ^worktree-agent- to expect the widened ^(worktree-)?agent- pattern. Regenerate golden-install-parity fixtures, agent-size-baseline, and workflow-size-baseline. Closes #1995 * fix(#1995): update extractCwdGuardBash sanity check for widened regex The e2e test's sanity check verified the extracted bash block contained 'worktree-agent-'. After widening to '(worktree-)?agent-', update the check to match the new pattern. * fix(#1995): widen missed workflow-guard branch check + changeset + lint fixes - hooks/gsd-workflow-guard.js: widen startsWith('worktree-agent-') to /^(worktree-)?agent-/ regex — same defect class, was missed in prior commit - tests/worktree.test.cjs: fix indentation regression from prior edit - Add .changeset/1995-worktree-agent-branch-namespace.md (pr:0 placeholder) Found by orthogonal code review (Step 4). * fix(#1995): regenerate golden + size baselines for workflow-guard change * docs(#1995): backfill changeset PR number (2548) |
||
|
|
d579daa3ed |
docs(#2505): Phase 6 — migration guide + built-in-only subagent-toolkit enum (#2538)
* docs(#2512): Phase 6 — migration guide + built-in-only subagent-toolkit enum * fix #2512: update CONTRACT-PIN for built-in-only subagentToolkit value * docs(changeset): backfill PR #2538 for Phase 6 (#2512) |
||
|
|
f654c24a3e |
feat(#2505): Phase 4 — runtime-aware subagent dispatch (Option A; resolve-dispatch-type query) (#2525)
* feat(#2508): Phase 4 Option A — runtime-aware subagent dispatch via resolve-dispatch-type query (#2505) * fix(#2508): prose-variant preamble (avoid scanner-tripping literals) + namedDispatch===false-only mapping * fix(#2508): remove leftover old-preamble lines (keep prose variant only) * fix #2508: prose-only reference file * test #2508: regen golden install parity after workflow preamble additions * fix #2508: remove preamble from plan-phase.md (Phase 6 capstone ceiling); regen size+golden baselines * docs(changeset): backfill PR #2525 for Phase 4 (#2508) |
||
|
|
936a345381 |
feat(#2505): Phase 3 — agent-skills fallback for non-dispatchable runtimes (#2521)
* feat(#2454): PR 2 — cmdAgentSkills fallback reads installed agent prompt When no agent_skills config entry exists for a given agent type (the common case on AGENTS-native runtimes), cmdAgentSkills previously returned empty output. Workflows that inject ${AGENT_SKILLS_*} into subagent dispatch prompts then carried nothing — the persona was lost. The fallback: resolve the runtime's agents directory via checkAgentsInstalled and read <agentsDir>/<agentType>.md. The installed agent prompt content (now present for kimi-code via the flat-skills install layout) flows into the dispatch prompt so the persona survives even without explicit config opt-in. This is the reporter's suggested fix #2 from #2454. The fallback triggers for ALL runtimes (not just kimi-code) when no config entry exists — it is strictly additive (returns content the previous empty path could not). If the agent file is not found on disk, the block stays empty (same as before). * docs(changeset): Phase 3 agent-skills fallback Added (#2510) * docs(changeset): backfill PR #2521 for Phase 3 (#2510) |