8da2dd3ad2bdd781e0529d3dd6f52fd13a9d0a1c
167 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
77fa08f1e8 |
fix(#2773): feed the spec-phase edge probe English-translated requirement text (#3713)
* test(#2773): failing-first contract and premise tests for translated edge-probe input Locks the Step 5.5 contract that a response_language project must feed the edge probe an English translation of each requirement's text, and binds that advice to measured engine behavior: the same requirement classifies to zero shapes in Portuguese and to collection/adjacency/empty/ordering in English. Also pins the honest limit — the issue's own repro sentence classifies to [] in English too, so translation is necessary but not sufficient and the authored shapes override is the documented fallback. Red before the doc change; the assertions are all false today. Refs #2773 * fix(#2773): feed the spec-phase edge probe English-translated requirement text The shape cues in src/edge-probe.cts are English word-boundary regexes, so a project running with response_language set wrote its SPEC requirements into the Step 5.5 $REQS_JSON heredoc in that language, matched no cue, classified to zero shapes, and landed every row in the unclassified sentinel (#1110). The taxonomy contributed nothing and --auto left it all unresolved — the probe was a silent no-op for exactly the spec type it exists to harden. Step 5.5 now states that the $REQS_JSON payload is engine input rather than user-facing output, so the response_language rule does not govern it: each requirement's text carries a faithful English translation, the SPEC keeps its original language, and requirement ids are never translated or renumbered. The instruction sits before the heredoc on purpose — the downstream APPLICABLE=0 warning fires only when every requirement is unclassified, so a partly-classified non-English spec would otherwise slip through with no signal at all. Measured against the compiled engine: the same requirement returns [] in Portuguese and collection -> adjacency/empty/ordering in English. Also measured: the issue's own repro sentence returns [] in English too, so translation is necessary but not sufficient — the instruction therefore points at the authored shapes override for prose carrying no cue in any language rather than promising that translation restores classification. Doc scope only, per the triage disposition on the issue. The compiled engine is untouched; the lang-hint / per-language cue-set fix is a separate follow-up. Closes #2773 * fix(#2773): clean up the edge-probe temp file on the placeholder-guard exit path Surfaced by the isolated security review of this branch. Between the mktemp and the unconditional cleanup, Step 5.5 has two sibling guards that disagreed about their own invariant: the engine-failure guard runs rm -f "$REQS_JSON" before exiting, while the empty/placeholder guard directly above it exited without one. A spec run that tripped the placeholder check therefore stranded a temp file holding the SPEC's requirement text in TMPDIR, once per failed run. The added contract test walks the region between the mktemp and the unconditional cleanup and asserts no exit path leaves the file behind, so the two guards can no longer drift apart. Proven to bind: run against the pre-fix file the walker reports the leaking exit; against the fixed file it reports none. Refs #2773 * docs(#2773): record the edge probe's English-cue input constraint in the predicate store The co-change gate flagged CONTEXT.md (13 co-changes with spec-phase.md) and docs/CONFIGURATION.md (11) as candidate-missing-updates, and both were real gaps rather than incidental coupling. CONTEXT.md's EdgeCompletenessProbeModule entry documents the input contract for classifyShape but did not record that SHAPE_CUES are English word-boundary patterns — so the predicate store implied text was language-agnostic, which is what a future agent reads before touching this seam. docs/CONFIGURATION.md's response_language row is what a non-English project reads when it turns the setting on; it now names the one deliberate exception and links to the FEATURES.md explanation, so the interaction is discoverable from the config key rather than only from the workflow. CONTEXT-INDEX.json regenerated via gen-context-index.cjs --write. The drift-ack fragment is updated for the final byte range and now also records the placeholder-guard cleanup fix folded into the same block. Refs #2773 * fix(#2773): append the growth rationale to the existing spec-phase.md ack entry The remote runner caught this: emitted-attribution.test.cjs pins the 0000-legacy-migration.json spec-phase.md entry permanently (the #2914 migration regression test asserts the exact '31987 -> 31997' delta text survives), so removing it to avoid a duplicate-key collision with a new fragment broke that test instead of satisfying the ratchet. The entry is an accreting log, not a single-use slot — #2733, #3132 and #3102 were each appended to the same reason string by later PRs, which is how a shared growth key coexists with the rule that two ack sources may never name the same path. This appends the #2773 rationale the same way and drops the separate fragment, whose spec-phase.md key was the collision. Verified locally by reproducing both affected tests against the real fragment before re-dispatching: the pinned delta survives, grown[0].acked is true, staleAcks is empty, and all 35 entries still read as spent. Refs #2773 * docs(#2773): add a how-to for probing edges in a non-English project The phase gate's enablementSequence check caught a wrong call of mine. I had recorded that no how-to was owed because the user takes zero extra steps — the workflow translates the probe input itself. Written out, though, the sequence from off to value is two steps and step 1 depends on response_language, a setting owned by a different capability than the edge probe, which is exactly the condition the how-to test names. There is also real task content a reference table cannot carry: the three-way split between a few unclassified rows (the classifier's recall gap), every row unclassified (the probe could not read the spec at all), and the silent partly-classified case where the APPLICABLE=0 warning never fires. That last one is what a user would otherwise misread as a clean bill of health. Shaped after the resolve-edge-coverage-findings / resolve-unreachable-guard siblings and indexed from docs/README.md next to its closest relative. Refs #2773 * chore(#2773): backfill the changeset PR number pr:0 placeholder replaced with the real PR number now that #3713 exists. Refs #2773 --------- Co-authored-by: sim <sim@local> |
||
|
|
adb46cdd85 |
feat(#2734): surface STATE.md commit-age on the statusline (#3700)
* test(#2734): failing-first suite for the statusline STATE.md freshness marker Binds the contract before any hook change exists: a `state ~N commits back` segment gated on the state_head stamp landed by #2622, firing at the same advisory threshold /gsd-health's W024 uses rather than at > 0. Covers all five acceptance criteria — threshold parity (19/20/21 boundaries), both renderers including formatGsdStateCompact, an exact spawn-count assertion, repo-pinning and sub_repos degradation, and behavioral parity against readStateHeadFreshness rather than a source-grep of the two fence copies. 52 example-based tests plus 5 seeded fast-check properties. Red now by design. * feat(#2734): surface STATE.md commit-age on the statusline Adds an opt-in `state ~N commits back` marker to the GSD-state segment, consuming the `state_head` stamp and freshness contract landed by #2622. A solo developer returning to a project reads "Phase 4, executing" in STATE.md and acts on it, without noticing the codebase moved 40 commits since that line was written. /gsd-health reports it as W024, but only if you think to run it; the statusline is the surface you see without asking. Fires at STATE_HEAD_ADVISORY_COMMITS (20), the same threshold W024 uses, not at > 0: with commit_docs:true the commit carrying a STATE.md sync advances HEAD by one, so > 0 would alarm permanently on a fresh project. Costs exactly one bounded git subprocess per render and none when disabled. `rev-list --left-right --count` answers ancestry and distance together, and repo pinning is a filesystem check mirroring projectOwnsItsRepo rather than a --show-toplevel compare, which is unreliable on macOS /private/var and Windows 8.3 paths. Every unresolvable input degrades to the tri-state unknown -- the marker is absent, never a "fresh" claim the project cannot substantiate: a malformed stamp, a root that does not own its .git, a sub_repos workspace, history rewound past the stamp, or git being unavailable. Also collapses statusline config resolution onto one resolveStatuslineOptions() seam. runStatusline() and renderStatusline() duplicated it byte-for-byte; one copy is what keeps a newly-added key from reaching only one of them. * test(#2734): route the e2e spawn through the process seam and fix fixture leaks Review findings from the two orthogonal passes: - `bothEntryPointsResolveOptionsIdentically` spawned a child and substring-matched its stdout to test a pure function. It now calls resolveStatuslineOptions() directly — no subprocess, no text matching. - `skipsFreshnessWorkWhenTodoTaskActive` genuinely needs a child (the !task gate lives in runStatusline, which reads stdin), so it now spawns through tests/helpers/process-seam.cjs and proves the negative with a filesystem fact: the git shim appends to a marker file on every invocation, and the assertion is that the marker never appears. Stronger than asserting text is missing, and it drops the last stdout substring match in the block. - Every fixture-creating test now registers `t.after(() => cleanup(dir))` instead of a trailing cleanup(dir), which leaked the temp repo on assertion failure. derivationAgreesWithStateModule reassigns `dir` across five fixtures, so it binds each directory at scheduling time rather than cleaning only the last. Also corrects markerCoexistsWithMilestoneComplete, which asserted the wrong expectation rather than finding a code defect: `percent` drives the progress bar too, so the milestone segment reads "v1.9 [##########] 100%". The marker appends after it, which is what the test exists to prove. CONTEXT.md's opt-in statusline key list was missing statusline.show_git as well as the new key; both are now enumerated. * docs(#2734): backfill changeset PR number (#3700) --------- Co-authored-by: sim <sim@local> |
||
|
|
2fca0e17e4 |
enhance(#2554): resolve code review depth from path-scoped override rules (#3695)
* test(#2554): failing-first suite for path-scoped code review depth overrides Binds the not-yet-built code-review-depth module: segment-aware path-prefix matching of a changed-file set against ordered {paths,depth} rules, resolution order flag > strongest matching rule > global > standard, typed validation errors, and the large-scope downgrade boundary. Also proves behaviorally that workflow.code_review_depth_overrides is not yet a registered config key. Refs #2554 * feat(#2554): resolve code review depth from path-scoped override rules Adds workflow.code_review_depth_overrides — an ordered array of {paths, depth} rules matched against a review's changed-file set by segment-aware path-prefix comparison. Resolution order is --depth= flag, then the strongest matching rule, then workflow.code_review_depth, then standard; a matching rule replaces the global rather than being max'd with it, so quick and standard rules stay meaningful. Glob metacharacters are a hard configuration error rather than sugar for a prefix, and malformed rules halt the review instead of degrading to standard. The resolver is pure and reports its own provenance, so the workflow can print the resolved depth and the rule that matched. The pre-existing >50-file deep-to-standard downgrade moves into the module and now names the rule it overrode. The key is registered centrally rather than as a capability config slice: the federated slice channel admits only boolean/string/number/enum, so an array slice would be dropped as malformed. Closes #2554 * test(#2554): correct depth-provenance assertions and pin out-of-repo paths Two corrections to the failing-first suite. The source assertion for a non-matching rule with no global configured expected 'config'; with no global set the depth comes from the default, and a companion assertion tolerated either value, so both passed against an implementation that derived provenance from whether any rules existed rather than from where the depth came from. The out-of-repo absolute-path case used a home-directory path that matched neither implementation, so it never exercised the defect it named. It now pins the discriminating cases: an absolute path outside the repo root must not match a repo-relative rule, and one under the root must. * docs(#2554): document path-scoped code review depth overrides Reference rows for workflow.code_review_depth_overrides in the configuration, features and commands references plus the locale copies that carry those tables, and in the planning-config reference. Explanation of why escalation is whole-review rather than per-file and why v1 is prefix-only. New how-to for scoping review depth by path, carrying the configuration-error reason table and the distinction between nothing to report and could not look. CONTEXT.md glossary entry and the INVENTORY row for the new CLI module. ja-JP and ko-KR CONFIGURATION.md carry no code_review keys at all, and ko-KR and pt-BR FEATURES.md carry no code-review config table, so those files are deliberately untouched. * fix(#2554): make the depth-misconfiguration halt executable and reject control chars Three review findings, all in this change. The misconfiguration halt was prose rather than shell: the error-printing fence was followed by an unconditional extraction fence, so an ok:false result threw and left the depth empty instead of stopping the review. Prose is not a guard — the two fences are now one block with a real conditional, and anything that is not the literal string true fails closed. An interior control character in a rule path survived validation and reached the provenance string and the summary box; rule paths now reject control characters via a new PATH_CONTROL_CHAR reason, after the glob check so precedence is unchanged. That in turn makes the field record safe to delimit, so the seven node invocations that each re-parsed the same result to read one field collapse to one. Also corrects the glossary entry's illustrative paths, which the glossary-ref check read as real repository references. * fix(#2554): use the fast-check v4 string API and acknowledge workflow growth Two failures from the remote matrix on d3111f45, both this branch's. The property block built its segment arbitrary with fc.stringOf, removed in fast-check v4. Because the arbitrary is constructed in the describe body, the throw took out all four property tests rather than one — they had never executed. Rewritten to fc.string({unit, ...}), the form this repo already uses in emitted-attribution.test.cjs. Every other fast-check helper in the file was audited against the installed module. The emitted-attribution growth arm needed an acknowledgment for code-review.md, which grew 5376 bytes. The pre-existing 3503 fragment keying the same file is spent — its ripple was absorbed when #3503 merged, and the base file is exactly the 34435-byte baseline this growth is measured against — so it cannot clear anything, while the ack lint hard-fails on a duplicate key across two sources. Removed it in favor of the new fragment, which is exactly how #3503 itself replaced the spent 3191 fragment. * docs(#2554): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
8526bd46f8 |
enhance(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR (#3688)
* test(#2475): widen the item-1 effort-caller guard to both CLI argument shapes The guard matched only `resolve-execution ... --effort\s`, but the CLI also accepts `--effort=<level>` (gsd-core/bin/gsd-tools.cjs). A workflow written with the equals form was a live invocation-override caller the guard passed silently, along with `--effort` at end-of-input. Lift the matcher to a shared predicate and assert it directly against every shape the CLI accepts, plus the decoys it must not fire on (--effortless, a bare --effort with no resolve-execution, item 6's --attempt caller, a call and flag split across lines). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2475): scope ADR-443 item 1 to the operator surface and ratify the ADR ADR-443 sat at Proposed on one condition: its Decision item 1 orchestrator invocation override needed a caller in shipped orchestration. Per the maintainer's ruling, take unblock path (b) for item 1 only -- record that the override is an operator-facing CLI surface, deliberately not driven by shipped orchestration, and ratify. The ADR's own path (b) wording is not adopted verbatim: it says the scope is limited to static install-time propagation, which is false on both counts -- item 6 has a live caller (#2296) and #2481 delivered a live invocation-time argv channel. Only one precedence step is narrowed. No consumer was invented to clear the gate: nobody has asked for a per-run effort override, and #2475's actual complaint is already closed by the cascade-to-argv path. The amendment states explicitly that --effort remains supported and is not deprecated, so the scoping is not read as dead code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2475): correct two bare-`gsd` invocation examples to `gsd_run` There is no `gsd` binary -- package.json exposes gsd-core, gsd-tools, gsd_run and gsd-mcp-server. Both sites presented a command that cannot run as written. One is in this branch's own new ADR-443 amendment; the other is a pre-existing error in the docs/CONFIGURATION.md assumption_delta row, fixed here rather than deferred. No translated copy carries either line, so no i18n drift is created. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2475): record the guard/CLI divergence risk on the item-1 matcher The predicate independently models gsd-tools.cjs's argument parser rather than sharing a constant with it, so a third --effort spelling would leave the guard reporting green while ADR-443's ratifying invariant silently stopped holding. Name that risk where the next editor will meet it. Also restores the bounded-prose rationale that was attached to the eslint directive removed in 39793079c -- the directive went unused once the regex moved to a const, but the reasoning it carried is still worth having. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1bf73d957b |
enhance(#2295): record the resolved model per reviewer in REVIEWS.md frontmatter (#3649)
* test(#2295): failing-first coverage for per-lane resolved-model recording * feat(#2295): record the resolved model per reviewer lane * docs(#2295): document the recorded reviewer model and its provenance * fix(#2295): refuse control characters in a recorded model value * test(#2295): correct watermark assertions for the widened mark shape * fix(#2295): anchor the role-manipulation injection pattern at a word boundary * feat(#2295): record the applied reasoning effort in the model value * chore(#2295): backfill changeset pr number * chore(#2295): restore em-dash in changeset body --------- Co-authored-by: sim <sim@local> |
||
|
|
1adf6d2245 |
fix(#3620): point the docs at files that actually exist (#3658)
* fix(3620): point the docs at files that actually exist docproof found 34 stale references; the reporter hand-read all 34 and reported the 8 that are real, explaining why the other 26 are deliberate (files the documents themselves label legacy or "superseded by", and one pre-Diataxis link label whose target still resolves). Those 26 are left alone — re-touching them would contradict the issue's own analysis. Every claim was re-verified against git ls-files at HEAD before editing. docs/INVENTORY.md said its roster is anchored by six drift-control tests. Five are gone (commands-doc-parity, agents-doc-parity, cli-modules-doc-parity, hooks-doc-parity in 5d8a8c4d; command-count-sync in |
||
|
|
fe64704ace |
enhance(#3588): add an opt-in commit_docs pre-commit hook (#3609)
* feat(#3588): add an opt-in commit_docs pre-commit hook Final phase of epic #2292, scope narrowed to opt-in by maintainer decision: default-on installation and the bin/install.js wiring it would have required are explicitly out of scope. Enabling is an explicit verb call. The hook is written to the repo's real hooks dir resolved via git rev-parse --git-path hooks, so a linked worktree or submodule whose .git is a FILE works rather than getting a literal .git/hooks path. It refuses rather than overwrite a foreign pre-commit, refuses to delete one it did not write, and refuses outright when core.hooksPath is already set -- a written-but-ignored hook is worse than a refusal. Ownership is detected by marker presence, not byte-equality, so a user who appends a line does not make it unrecognizable. Deliberately NOT included: teaching cmdCheckCommit the per-phase commit_docs tier. #3587 was still unmerged when this landed, and implementing precedence against helpers that did not yet exist would have meant a second copy of the resolution chain -- the divergence class this epic has spent three phases fighting. That follows as its own change now that #3587 is on next. The ordering constraint is recorded in the design doc: this must not merge before #3587, or the hook would block a commit cmdCommit itself allows. * fix(#3588): teach the commit_docs guard the per-phase tier and -z paths Part 1, deferred until #3587 merged. cmdCheckCommit read only project-level commit_docs, so once #3587 landed, a phase with phase_commit_docs true under project false was ALLOWED by query commit and BLOCKED by this guard -- and the hook shipped in this same branch shells out to it. It now derives the staged phase via the single-owner detectPhaseNumberFromFiles and resolves through #3587's own resolveCommitDocsPolicy rather than a second precedence copy. Also fixes a proven false negative in the harm direction. git diff --cached --name-only C-style-quotes any path with non-ASCII or special characters, so a staged .planning/cafe.md was emitted as a quoted string, failed startsWith('.planning/'), and slipped past the guard entirely under commit_docs:false. Reading with -z and splitting on NUL removes the quoting at the source. The f.startsWith('.planning\\') branch was dead code under that read -- git emits /-separated paths on every platform -- and is removed rather than left implying coverage it never provided. The earlier C7 test pinned the buggy behavior as intended; it now asserts the file is detected and the commit refused. Self-caught: the commit-docs-guard verb was wired into the routers by this branch's earlier pass but missing from the top-level help listing. * test(#3588): replace try/finally with t.after, add negative-routing cases Standards review findings. CONTRIBUTING bans try/finally inside a test body outright -- it masks failures -- and B8 used one for worktree cleanup. Now t.after(), assertions unchanged. The new commit-docs-guard command family had zero negative-routing coverage, which CONTRIBUTING requires for any change to command dispatch. B11-B15 cover no subcommand, unknown, empty string, whitespace-only and a flag-shaped value, each asserting non-zero exit, a structured error, no stack trace, and -- the one that matters for a command that writes into a user's repo -- that NO hook is written in any of them. Those tests were verified to fail when routeCommitDocsGuard's else-branch is neutered, so they exercise the routing guard rather than any convenient error path. Also made two error() calls' control flow explicit with a return; they were safe only because error() is typed never two files away. * chore(#3588): backfill changeset pr number to 3609 * test(#3588): skip Windows-unrepresentable fixtures on win32 CI's Windows shards caught two of my own tests: fixtures whose filenames contain a quote and a backslash. Both are illegal on Windows -- backslash is the path separator, quote is invalid on NTFS -- so fixture creation failed before any assertion ran. Test-portability defect, not a production one. Those inputs cannot exist on that platform, so the guard has nothing to detect there. Both now check process.platform FIRST, before any fs or git call, and use t.skip() rather than a bare return -- a bare return registers as a PASS and would hide the gap it is meant to record. Each carries a comment saying the input is unrepresentable rather than unverified, so nobody later re-enables it. No padding added: the cafe.md case already exercises git's C-quoting path on every platform, since non-ASCII names are legal on NTFS. This is exactly the coverage the Linux-only remote matrix cannot provide, which the PR body already stated -- CI's Windows shards are what caught it. --------- Co-authored-by: sim <sim@local> |
||
|
|
debeabd524 |
enhance(#3587): add a per-phase commit_docs override (#3601)
* feat(#3587): add a per-phase commit_docs override Delivers epic #2292's second user story: commit an architecture phase's artifacts while execution phases stay local. commit_docs was project-wide and binary, so the only choices were all phases or none. Shape is a config dynamic key phase_commit_docs.<phase-id>, following the 14 existing dynamicKeyPatterns precedents rather than inventing a PLAN.md frontmatter spec -- which #2292 itself flags as becoming its own maintenance surface. Tier 1 resolves in cmdCommit, NOT in loadConfig: loadConfig has no phase context and is called by nearly every command, so threading one through it to serve a single caller would be a far larger blast radius for no gain. The phase comes from detectPhaseNumberFromFiles, which cmdCommit already computes for branch naming and which is already hardened against the #2539 project-code bug. Suppression by the per-phase tier returns its own reason rather than reusing skipped_commit_docs_false -- telling a user their project setting is false when it is true would be actively misleading. Additive; the two existing reason strings that agents/gsd-executor.md matches on are unchanged. The manifest's phase-id pattern is a hand-copy of PHASE_NUMBER_TOKEN_SOURCE because the manifest is hand-maintained JSON, so a behavioral parity test asserts both surfaces accept and reject the same token shapes. * fix(#3587): fold tests, close review findings, update reference docs Fold: the new tests were added as their own file, which required loosening a grandfathered lint-test-file-count bucket 5-to-6. A ratchet exists to go down only. commit-docs-bypass.test.cjs is the established commit_docs test home and already hosts two folded suites, so the tests fold there as a third block and the allowlist is reverted untouched. Standards review: CONTEXT.md and the test header both cited a phase-commit-docs-manifest-parity.test.cjs that never existed; a repo-wide sweep found a fourth stale cite in the schema manifest description. All four now name the real location. Spec review: the issue's Scope of changes named planning-config.md and git-planning-commit.md and neither was touched. Both now document the four-tier precedence and the new skip reason. Security review, minor and unproven: detectPhaseNumberFromFiles returns the FIRST matching path's phase, so a --files list spanning two phases resolves the override against whichever comes first. That helper is hardened and widely used, so it is not changed; the behavior is pinned by a named test and disclosed in the design and user docs. A pinned behavior is not a bug; an unpinned surprise is. * chore(#3587): backfill changeset pr number to 3601 --------- Co-authored-by: sim <sim@local> |
||
|
|
5f64d999dc |
fix(#3586): warn when .planning/ is gitignored but still tracked (#3598)
* feat(#3586): warn when .planning/ is gitignored but still tracked git ignore rules have no effect on files git already tracks, so a project that committed .planning/ before ignoring it keeps staging those files -- while commit_docs correctly resolves to false, which is exactly what makes the contradiction invisible. The probe lives in the SNAPSHOT BUILDER, not the rule: Rule.check may perform no ambient I/O (ADR-3180 8.1 rule 1, enforced by lint-planning-snapshot-bypass). buildPlanningTrackedField follows buildWorktreeHealthField's precedent -- injected execGit, bounded, degrading to UNREADABLE with a typed reason rather than throwing. W024 went inline instead only because no snapshot field carried its fact; that precondition does not apply here. W029 fires only on COMPLETE scope with ignored and tracked both true, so a degraded probe yields neither a finding nor a false all-clear, and the default project (tracked, not ignored) stays silent. The remedy is ADVISE-only -- --repair never untracks anything. * docs(#3586): document W029 and correct the health rule count CONFIGURATION.md documented the gitignore auto-detect without the caveat that ignore rules do not affect already-tracked files -- the very gap W029 exists to surface. Adds the caveat, the warning, its remedy, and why --repair will not act on it. CONTEXT.md's rule count was stale at 31 before this change (actual 32 through W028); corrected to 33 and pointed at the two other places the count is locked, so the next editor updates all three together. * fix(#3586): treat ls-files overflow as tracked, add CLI-level W029 tests Review findings. Security (minor, confirmed): execGit sets no maxBuffer, so Node's 1MB default applies to git ls-files. A .planning/ tree large enough to overflow it failed into git_list_failed and silenced W029 -- a false negative in exactly the large-history case most likely to have the real bug. Overflow is now treated as PROOF of tracking (the output was non-empty by definition) and resolves to tracked:true, scope COMPLETE, reason ok_truncated. Spec (major): test-matrix rows C1 and C2 were never implemented -- there was no CLI-level integration test at all, only rule-level ones. Both now drive the real validate-health dispatch and confirm W029 is reachable end-to-end. Known limit documented, not papered over: a deliberate git add -f under an otherwise-ignored .planning/ raises the same signal as the accidental case. There is no reliable way to tell them apart, the finding is advisory-only, and a heuristic that cannot actually distinguish them would be worse than the honest caveat. * test(#3586): update frozen health-doc counts and acknowledge health.md growth The remote matrix caught three gates that lint:ci does not cover. gen-health-docs.test.cjs froze a 35-row / 32-rule assertion; W029 makes it 36/33. Updated both the assertion and the test NAME, which embeds the counts -- a stale name is a lie even when the assertion passes. The second reported failure was the same assertion surfacing at describe-rollup granularity, not a distinct bug. emitted-attribution's growth arm needed an ack for the generated health.md. health.md was already named in 3309-health-docs-generated.json, and two ack sources naming one path is a hard error -- so a new fragment was not an option. That fragment's own history shows the pattern: #3309 created it, #2873 amended it in place for W028. Amended again for W029, with a note recording why this one file is amended rather than joined by a sibling. * docs(#3586): add the private-planning how-to and fix a wrong link docs/CONFIGURATION.md pointed 'Configure private planning' at how-to/configure-model-profiles.md -- an unrelated page -- and no private-planning how-to existed at all. Found while editing that section. The how-to test genuinely fires here: going private is four steps and crosses planning.search_gitignored, a setting owned by another concern, so a reference table structurally cannot carry it. The new page walks the whole sequence and leads with the step people miss -- .gitignore does not untrack what git already tracks -- which is the exact state W029 now detects. Also corrects 'artefacts' to 'artifacts' (repo house style is American). * chore(#3586): backfill changeset pr number to 3598 --------- Co-authored-by: sim <sim@local> |
||
|
|
311711754f |
docs(#3531): document where inherit must be set to reach tiered agents (#3550)
Co-authored-by: sim <sim@local> |
||
|
|
b7cca0363f |
fix(#3531): merge routing_tier_defaults over manifest tier defaults (#3539)
* test(#3531): failing-first suite for routing_tier_defaults manifest merge * fix(#3531): merge routing_tier_defaults over manifest tier defaults * docs(#3531): document routing_tier_defaults merge-over-built-ins semantics * fix(#3531): correct test helper scope, update folded #443 expectations, guard merge keys * test(#3531): pin tiers in effort-sync and surface-axis fixtures post-merge * chore(#3531): backfill changeset pr number * fix(#3531): correct rebase resolution — keep both 3531 and 3533 test blocks intact * test(#3531): pin inherit/effort fixtures to the layer that reaches tiered agents --------- Co-authored-by: sim <sim@local> |
||
|
|
adb2d03ed8 | Merge pull request #3540 from open-gsd/fix/3532-global-defaults-diagnostic | ||
|
|
50d5368add | fix(#3533): effort inherit — expressible, omitted at writers, never re-added (#3541) | ||
|
|
d26bfc2a3f | fix(#3534): resolve-execution reports resolved and effective effort | ||
|
|
129871a8be | fix(#3532): warn when global defaults keys are shadowed by a project config | ||
|
|
268ca7e32d |
fix(#3504): harden hook injection patterns and force-add guard (#3510)
* test(#3504): add failing-first parity, fail-closed, and bypass suites * fix(#3504): harden hook injection patterns and force-add guard * test(#3504): stage the scanner lib dependency in shared-hooks fixture * chore(#3504): backfill changeset pr number * test(#3504): build parity samples from fragments for the ci scan --------- Co-authored-by: sim <sim@local> |
||
|
|
7976b1ca0d |
feat(#1689): per-plan agent_hint executor routing (#3417)
* feat(#1689): per-plan agent_hint executor routing Option A per-plan specialist routing: a plan with an `agent_hint:` frontmatter field is dispatched to that subagent instead of gsd-executor when it resolves on the active runtime; absent/unresolved/disabled falls back to gsd-executor (byte-identical). Default-on via workflow.agent_hint_routing. - src/phase.cts: parse agent_hint into the plan-index JSON (plan_json.agent_hint) - agent-install-check.cts: resolveAgentHint() reuses getAgentsDir + runtime filename variants; probes project + global agent dirs; fails closed; rejects path-traversing names - gsd-tools.cjs: 'resolve-agent' query route (fail-closed to gsd-executor; --raw/--json) - execute-phase.md: lean per-plan reference + {EXECUTOR_TYPE} placeholder (host stays under the ADR-857 Phase 6 byte ceiling) - execute-phase/steps/per-plan-executor-routing.md: resolution logic (Agent()-based dispatch; advisory on orchestrator-worktree) - config: workflow.agent_hint_routing (validKey, default-on via SCHEMA_DEFAULTS, boolean validator) - docs (CONFIGURATION.md, plan-md.md), changeset, tests/agent-hint-routing-1689.test.cjs (17 tests) * chore(#1689): backfill changeset PR number (#3417) * chore(#1689): regenerate install-tree fixtures for new workflow fragment * chore(#1689): ack deliberate execute-phase.md growth (agent_hint routing) * test(#1689): SPAWN contract allows parameterized subagent_type placeholder agent-frontmatter's spawn-type checks scanned subagent_type="..." as a concrete agent name. execute-phase now uses subagent_type="{EXECUTOR_TYPE}" (a runtime placeholder resolved via resolve-agent, default gsd-executor). Skip {TOKEN} placeholders in both the known-type and <available_agent_types> checks; execute-phase still lists the built-in roster incl. gsd-executor. * fix(#1689): CI conformance for the routing fragment - per-plan-executor-routing.md: add the canonical runtime-launcher preamble to its gsd_run block (runtime-launcher-parity #373), matching sibling step fragments. - agent-install-check.cts: drop a literal ~/.claude/agents path from the resolveAgentHint JSDoc so it does not leak into the compiled engine .cjs (cline install leak guard). --------- Co-authored-by: sim <sim@local> |
||
|
|
f0abdb1b89 |
fix(#2486): do not recommend or persist Claude-only worktree isolation on non-Claude runtimes (#2531)
* fix(#2486): runtime-branch the settings worktrees question + W020 health diagnostic On non-Claude runtimes /gsd:settings offered "Yes (Recommended)" for worktree isolation and persisted workflow.use_worktrees: true — the exact value the execution workflows fail closed on (#1521 guards). Branch the question on the same stamped config-get runtime read the guards use: Claude keeps the unchanged question; non-Claude offers only "No (Recommended)" / "Leave unchanged", never persists true, and warns when the config carries an inherited explicit true. /gsd:health gains W020, surfacing such a config with the guards' own predicate before execution-time failure. Docs state the runtime-conditional default. Fixes #2486 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2486): add changeset for PR #2531 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2486): reassign the health worktrees check W020 -> W024 (verify.cts namespace collision) The workflow-level check collided with the live W020 (git-worktree-list health) emitted by cmdValidateHealth in src/verify.cts — invisible from health.md's error_codes table, which stops at W019 and under-represents the real namespace (W010-W017, W020-W023 all live). W024 verified free. Adds a regression test pinning the chosen code against src/verify.cts so a future assignment cannot silently collide, a table note naming the namespace owner, and the changeset body reworded to house style. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2486): pre-select the recommended repair in the broken-inheritance case Review round 2: at settings.md:142 the pre-selection rule left "Leave unchanged" as the default when the config carried an explicit non-false use_worktrees — the exact broken state the adjacent notice warns about, so accepting the default kept a config that fails closed at execution time. "Leave unchanged" is now the default only when the key is absent (nothing to repair); explicit false AND explicit non-false both pre-select "No (Recommended)", aligning the default, the label, and the notice. Pinned by two source-contract assertions in the #2486 regression block. Goldens (settings.md hash x19) + size baseline regenerated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2486): gate the worktrees question on dispatch.isolation, not the runtime name Review round 2: #2584 Phase 3 replaced the runtime-name test with a declared `dispatch.isolation` capability, invalidating this PR's premise. cursor declares harness-worktree and codex/opencode/kimi/ kimi-code declare orchestrator-worktree, so a `RUNTIME != claude` gate blocked a supported configuration on five runtimes and false-warned in health. - settings.md + health.md read `query dispatch-isolation` and branch on `ISOLATION = none`; the runtime-name read is gone from both, and the capability read needs no per-runtime stamping (it fail-closes unknown/ undocumented internally) - all "Claude Code-only primitive" prose rewritten, including the two gates the shell-syntax check missed (config-key list, JSON schema comment) - W024 reconciled across health.md + CONFIGURATION.md + planning-config.md (docs still said W020, which collides with a verify.cts code) - health.md error-codes table fixed: the namespace note no longer sits between rows orphaning I001 - the asymmetry note for the two workflows #2584 has not migrated (quick.md, diagnose-issues.md) is enforced by a set-equality test with a self-check table, so it cannot go stale in either direction Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(#2486): restore the Executor isolation section clobbered by #2661 `46ba02ac` (feat(#2630), the current next tip) reverted docs/CONFIGURATION.md to a pre-#2584 state: it restored the old "Non-Claude note" wording on the workflow.use_worktrees row and deleted the whole "Executor isolation per runtime" section. The change is unrelated to that PR's phase-estimation feature and looks like a stale-copy edit. This PR's use_worktrees row links to #executor-isolation-per-runtime, so the deletion leaves a dangling anchor. Restored byte-for-byte from |
||
|
|
2076d450d7 |
fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not the runtime name (#2728)
* fix(#2652): gate quick/diagnose dispatch on dispatch.isolation, not runtime name quick.md and diagnose-issues.md kept the pre-#2584 `RUNTIME != "claude"` worktree gate, so every non-Claude runtime failed closed regardless of the capability it negotiated — including Codex, which declares orchestrator-worktree. Route both through the negotiated dispatch.isolation seam via a new shared reference, and migrate the two execute-phase reference fragments that carried the same runtime-name gate. - new gsd-core/references/dispatch-isolation-gate.md: canonical ISOLATION resolution, harness-flag resolution, single-agent degrade rule - quick.md / diagnose-issues.md read the gate; dispatch uses the {harnessFlag} placeholder rather than a hardcoded isolation="worktree" - execute-phase-wave-guard.md / execute-phase-between-wave-reset.md: migrate [ "$RUNTIME" = "claude" ] -> [ "$ISOLATION" = "harness-worktree" ] - every degrade site now clears BOTH USE_WORKTREES and ISOLATION; clearing one dispatched an isolated agent with no base guard and no manifest - parity guard in host-integration.test.cjs scans workflows AND references and matches six reintroduction shapes - migrate four tests that pinned the pre-#2584 runtime-name contract Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): use the /gsd:<cmd> namespace in the isolation degrade messages The degrade warnings cited /gsd-execute-phase, the retired hyphen form that slash-command-namespace.test.cjs rejects in Claude-facing source. Same length, so the quick.md size budget is unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2652): add changeset for PR #2728 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): normalize dispatch-site paths to forward slashes for Windows path.relative() returns backslash-separated paths on Windows, so the #2652 dispatch-site parity test compared "gsd-core\workflows\quick.md" against the hardcoded forward-slash literal "gsd-core/workflows/quick.md" and failed on every windows-latest CI lane. Normalize with .replace(/\\/g, '/'), matching the existing convention used elsewhere in this suite (e.g. tests/branch-no-track-guard.test.cjs:37). * test(#2652): restore the size-growth acknowledgment The rebase dropped tests/emitted-drift-ack.json. #2757/#2758 fixed the ATTRIBUTION axis, but the SIZE-GROWTH axis is independent: diagnose-issues.md (+2086) and quick.md (+230) still need an ack naming them and saying why. Verified: 65/66 without it (both files named), 66/66 with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2652): convert execute-plan.md Pattern A onto the dispatch-isolation gate Pattern A hardcoded `isolation="worktree"` — Claude Code's own literal — gated only on `workflow.use_worktrees`, with no capability negotiation at all. It is the same defect #2652 fixes at the other four sites, just a different shape: the file contains no RUNTIME variable, so the new detector correctly does not flag it. Concrete break: a Codex user who follows this PR's own newly-documented pattern and sets `workflow.use_worktrees: true` to get isolated dispatch via /gsd:quick then runs a plan through /gsd-execute-plan Pattern A, and hits an unconverted path — either an Agent() call erroring on an unrecognized parameter or silent unisolated execution, depending on host tolerance. Pattern A is a single-agent dispatch site through the host's own subagent tool, so it takes the same treatment as quick.md and diagnose-issues.md: resolve ISOLATION/HARNESS_FLAG through the canonical reference, degrade to sequential on orchestrator-worktree hosts, and substitute the host's declared {harnessFlag} instead of Claude Code's literal. while the area was open. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2652): add the INVENTORY row for dispatch-isolation-gate.md, refresh CONTEXT Two bookkeeping gaps flagged in review: INVENTORY.md had no row for the new gsd-core/references/dispatch-isolation-gate.md. INVENTORY-MANIFEST.json was regenerated correctly and its --check only diffs a live directory scan against the committed manifest, so CI passed regardless — but gen-inventory-manifest.cjs's own stderr guidance says to add the matching INVENTORY.md row. This is the repo's named "Inventory Drift" pattern. Placed with the dispatch/isolation cluster (worktree-branch-check, runtime-aware-dispatch) rather than alphabetically, matching how that table is grouped. CONTEXT.md's Host-Integration Interface entry still described dispatch.isolation as "declared and negotiated but not yet consumed by any scheduler — Phase 1 of #2584". That was already stale before this PR (execute-phase graduated in Phase 3) and more so now with three single-agent dispatch sites consuming it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): detect reversed-operand runtime gates; add a permutation property All five reintroduction regexes assumed $RUNTIME on the LEFT of the comparison, so `[ "claude" != "$RUNTIME" ]` — the same gate written backwards — evaded every one of them. Verified against the old patterns before fixing: all four reversed shapes (single bracket, double bracket, test builtin, JS template) scored EVADED. Each comparison shape is now generated in both operand orders from a single template, so a shape cannot be added in one order and forgotten in the other. The mutation table gains the four reversed cases. Also adds the fast-check property review suggested in place of the hand-rolled cases: it generates the cross product of the axes an author actually varies — bracket form, operator, operand order, quoting, spacing, runtime id — so a permutation the hand-written patterns miss surfaces here rather than in production. The 11 explicit cases stay as named regression anchors. execute-plan.md joins the scan's required-identities list now that it is a converted dispatch site. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): acknowledge the execute-plan.md size growth The Pattern A conversion adds 811 bytes to an emitted workflow. Per #2719 the size axis needs its own acknowledgment, independent of attribution. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2652): repin the execute-plan.md PROSE_ALLOWLIST line after the rebase The #2751 command-position gate pins its prose exemptions by line number. This branch inserts the dispatch-isolation resolution above the `validated downstream by gsd-tools uat classify-coverage` sentence, moving it from execute-plan.md:387 to :397 — which fired the gate twice for one displacement (an un-allowlisted mention at 397, a stale entry at 387). The prose itself is unchanged from next; only the pin moves. Fixes #2652 * fix(#2652): gate the #2649 base-check on ISOLATION in diagnose-issues.md The rebase onto next merged #2649's pre-dispatch base-check textually, but its degrade flipped USE_WORKTREES after ISOLATION was already resolved, so the degrade never reached the dispatch decision. Gate the block on ISOLATION = "harness-worktree" and degrade ISOLATION itself, the same pairing quick.md already uses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): key quick.md post-dispatch bookkeeping on ISOLATION, not the Claude literal Review Blocker: the manifest append (l.822), worktree merge-back (l.825), and its skip clause (l.839) all conditioned on the literal isolation="worktree" — Claude Code's own rendering of {harnessFlag}. Cursor renders --worktree, so a newly-unblocked isolated Cursor run created a worktree whose committed work was never merged back and never cleaned up, silently. All three now key on ISOLATION = "harness-worktree" at dispatch. The existing parity detector cannot catch this class (its ISOLATION_TOKEN treats the literal as a legitimate marker), so this adds a dedicated literal-condition detector with a discrimination proof against both pre-fix sentences, a benign-mention control, and a positive pin on all three re-keyed conditions. Verified fail-first against the pre-fix quick.md. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2652): scope the use_worktrees=false install stamp to isolation=none runtimes `_stampNonClaudeRuntimeDefaults` rewrote every non-Claude runtime's `workflow.use_worktrees` read to `--default false`. That default resolved before `gsd_run query dispatch-isolation` was ever consulted, so the five runtimes that declare worktree support — cursor (harness-worktree) and codex/opencode/kimi/kimi-code (orchestrator-worktree) — got ISOLATION=none regardless of what they negotiated. The gate this PR migrates dispatch onto was therefore still deciding isolation by runtime name, one layer down. The stamp's #1521 premise was that worktree isolation *was* Claude Code's isolation="worktree" spawn parameter, which no other host honored. #2584 replaced that premise with the negotiated capability. The stamp is now scoped to runtimes whose negotiated isolation really is `none`, where the default it writes is the outcome the resolver reaches anyway. `_negotiatedDispatchIsolation` mirrors routeDispatchIsolation's resolution against the same registry — closed vocabulary, a harness-worktree host must declare its flag, an orchestrator-worktree host must carry a descriptor that resolves — and fails closed to `none` on anything else, so an undeclared or unknown runtime keeps today's behavior. Two #1515 tests pinned the superseded premise for codex and are re-pointed at the new contract rather than deleted: the safety property they protect is now held by the isolation gate's fail-closed resolution, not by a name-scoped install-time default. Verified fail-first — all five assertions red against the pre-fix source, green after. * test(#2652): acknowledge the emitted ripple and re-point the end-to-end stamp proof Scoping the use_worktrees stamp changes emitted output, and two gates caught it. `gsd-core/workflows/execute-phase.md` now differs at emit time for the five hosts that declare worktree support (cursor harness-worktree; codex, opencode, kimi, kimi-code orchestrator-worktree) — the source file is byte-identical, only the stamp is gone. Acknowledged in this PR's fragment. `tests/install.test.cjs`'s real-install assertion pinned the superseded premise end-to-end, asserting codex receives `--default false`. Re-pointed rather than deleted, matching the two unit tests: it now proves codex keeps the unstamped `true` read. A second arm installs windsurf — which declares isolation `none` — and asserts the false stamp is still applied there, so the change cannot silently degrade into "never stamp" without a test noticing. The ack entry collides with `2658-trae-instruction-file-path.json`, which is fully spent (merged via #2925, so all 25 of its entries are present at base and gate nothing) and is pruned for the same reason and by the same rule as the spent `2649-*` fragment this PR already removed. #2566 prunes the same file for the same collision on `new-project.md`; a delete/delete merges cleanly either way, and the base-side cleanup would make both unnecessary. * fix(#2652): re-record the sentinel when a dispatch site degrades isolation Review Blocker B1/B2/B3. Every isolation degrade in a dispatch site is decided in shell, where routeDispatchIsolation cannot see it. That resolver persists whatever it resolved to the run-scoped sentinel as an unconditional side effect (#3045), so a degrade that only reassigns $ISOLATION leaves the sentinel asserting harness-worktree while the dispatch correctly omits the harness flag. The shipped PreToolUse guard reads the sentinel at the instant of the Agent() call and denies that mismatch with exit 2 — the work does not run unisolated, it does not run at all. Latent on this branch and lands on rebase, since |
||
|
|
96a82bbffb |
enhance(#3245): report the detected host runtime in init (#3307)
* test(#3245): failing-first coverage for host runtime detection in init Locks the behavior epic #2313 Phase 5 must produce before any of it exists: init reports the detected host, explicit GSD_RUNTIME and config runtime still outrank detection, non-Codex sessions are untouched, and nothing is ever written to shared defaults (#2297). * enhance(#3245): report the detected host runtime in init init reported agent_runtime: claude inside a Codex session, and resolved agents_dir to the Claude agents root with agents_installed: true — a spuriously healthy triple. Runtime identity was only ever read from GSD_RUNTIME or an explicit runtime in .planning/config.json. Adds a detection rung beneath both explicit sources, in a new pure module. Codex is identified from its own documented session environment (CODEX_SANDBOX / CODEX_SANDBOX_NETWORK_DISABLED), else an explicitly exported CODEX_HOME whose config.toml exists. The default ~/.codex is never probed: that file exists on every machine that has run Codex, so probing it would misreport other runtimes' sessions. resolveRuntime keeps its exact contract and all 71 dependents, including formatGsdSlash command-style emission; only withProjectRoot consumes the new rung. Nothing is written on any path (#2297). Explicit config still wins (#2517). * fix(#3245): make the parity guard real and single-source the marker Four independent review passes found the generative-fix-divergence guard was vacuous: it asserted agreement at the one input where inferPreferredRuntime and detectHostRuntime do not differ, so it could not fail. It now pins the actual divergence point (CODEX_HOME set, config.toml absent) and records that the asymmetry is deliberate. The config.toml marker is now single-sourced from update-context.cts and imported, rather than carried independently by two surfaces. tests/helpers.cjs now scrubs CODEX_SANDBOX and CODEX_SANDBOX_NETWORK_DISABLED: GSD reads them, so an ambient Codex session would otherwise make the non-codex control test fail non-deterministically. Also: detection is throw-safe end to end rather than only around the fs probe; the Windows-join test is replaced with one that can actually fail (trailing-separator, catches hand-rolled concatenation); the #2297 no-write proof now wraps resolveReportedRuntime, the function that ships, across all three ladder outcomes. * chore(#3245): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
2dbee3ebdd |
enhance(#2229): add three-way claim disposition (admit/refute/abstain) to /gsd-explore research pass (#2543)
Closes #2229. Each claim surfaced by /gsd-explore's research pass is dispositioned admit, refute, or abstain, with abstentions routed to a visible ledger instead of being smoothed into confident prose. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim; a strong prior is never authoritative alone. Two guards ride with it: conflict-abstention, and a tier floor that presents a would-be admit as an abstain when the researcher's resolved tier is the budget tier or cannot be determined. To make that floor enforceable, resolve-model now emits the effective tier (--pick tier). It was already computed above the resolve_model_ids omit gate but was unreachable from a workflow, which left the floor inert on every non-Claude install - the model id is blank under omit and runtime-substituted where a tier map exists, and the profile defaults to balanced. The tier signal mirrors every resolution step that can change which tier runs, including the model_policy preset, and reports unknown rather than guessing. Output is additive; model, profile and effort are unchanged. Two residuals are disclosed in the workflow rather than papered over: a raw-model-id model_overrides pin reports unknown and is floored (fails closed), and a model_profile_overrides entry repointing a tier at another tier's model can under-report (fails open, and predates this change). Admin merge used only to satisfy the missing secondary reviewer on a single-maintainer PR. No CI failure and no conflict were bypassed: 38 checks green, remote runner 32255/32255 on both Node lanes. |
||
|
|
b901d1e06f |
feat(#1953): complexity-triggered refactor extension point (execute:post) (#3261)
* test(#1953): failing-first suite for the complexity-triggered refactor hook 60 behavioral cases against src/complexity-trigger.cts, which does not exist yet: decision-point counting, the comment/literal stripping leak surface, threshold and jump-delta boundaries at limit-1/limit/limit+1, stable-anchor baseline semantics, and fs fault injection via mock.method. Two fast-check properties assert that stripping never manufactures a decision point and that comments and string literals are score-neutral. Also registers the refactor-trigger capability manifest (inert until refactor.trigger_enabled) and regenerates the capability registry and matrix. Verified RED on the remote runner before any implementation exists. * feat(#1953): complexity-triggered refactor extension point Adds the opt-in refactor-trigger capability. After a phase executes, an execute:post step measures per-function complexity for the files the phase touched and writes a scoped refactor proposal when a function crosses the configured threshold or drifts past its recorded anchor. Design notes worth carrying: - The signal is computed in-core (decision-point counting over comment- and literal-stripped source, Node builtins only) rather than via Memtrace or a shelled-out analyzer. The hook fires as a deterministic CLI, not an agent with MCP tools, and core takes no external dependencies — this is the only option a behavioral test can bind to. The metric sits behind a seam. - The baseline is a stable anchor, not a rolling value: set on first observation, moved only on disposition. A rolling baseline makes the delta the single-phase change, so a function creeping +2 per phase never trips a delta of 5 and the jump check adds nothing over the absolute threshold. - Strict mode records an open deviation window in the broken-windows ledger rather than declaring its own ship:pre gate. ship.md has no generic ship:pre gate dispatch — only two hardcoded branches — so a third gate of any kind would be declared and never evaluated. - The gate clears on the proposal being dispositioned, never on the score improving. A blocking complexity number is one an executor can satisfy by splitting a coherent function in two. execute-phase.md gains a generic execute:post step-dispatch contract; it previously matched only ref.skill == "code-review", so any other step registered there was declared and never run. The code-review branch is unchanged. Full rationale in ADR-1953. Closes #1953 * fix(#1953): close git option injection and symlink escape in the refactor hook Three findings from the isolated security review, all fixed inline. HIGH — changedFilesSince interpolated the --since value into a revision token placed before the -- separator. A -- only stops PATHSPEC parsing of arguments after it; git still option-parses what comes before. So --since '--output=/tmp/x' became --output=/tmp/x..HEAD, which git accepts as --output=<file> and uses to redirect diff output — an arbitrary write. Fixed with --end-of-options before the revision range plus a conservative ref validator. The validator deliberately permits ~ ^ @ { } because those are legitimate git REVISION syntax (HEAD~1, main@{yesterday}) as distinct from ref-NAME syntax; --end-of-options is the actual barrier. The doc comment asserting the trailing -- was sufficient was wrong and is corrected. MEDIUM — resolveConfinedPath confined by string prefix only, so a symlink committed inside the repo passed the check (its own path is under cwd) and readFileSync then followed it outside the root. Now lstat-checks for a regular file and skips anything else with REFACTOR_FILE_UNREADABLE, so one bad path skips one file and the run continues. LOW — the new execute:post dispatch contract showed the gsd_run example before the rule requiring ref.command be validated first. That prose is executed by an agent, so textual order is execution order. Reordered. Refs #1953 * fix(#1953): make the analyzer able to see TypeScript at all Found by running the shipped analyzer over its own source: it reported functions=1 for a 940-line module with 24 function forms. A return-type annotation or a generic parameter list made a function invisible — `function f(a): number {}` and `function f<T>(a: T): T {}` both detected as zero. Since gsd-core is written in .cts and the capability declares .ts/.cts/.mts analyzable, the feature silently found nothing in this repo's own primary language while reporting success. A safety net that reports "all clear" because it cannot see is worse than no safety net. All 98 tests passed over this, because every fixture was plain JS — the exact failure the test matrix's own "assert against the shape production uses" warning describes. Adds a TypeScript-shapes suite covering return types (including unions, generics, object literals and type predicates), generic parameter lists (constrained and defaulted), export/async/generator combinations, annotated arrows, class-method modifiers, and optional/ default/rest params — plus the two traps: an overload signature has no body and must not count, and `a < b && c > d` is a comparison, not a generic. Detection now reports 24/37/21 functions for the three source files, which matches a hand count exactly. Also from review: - The strict-mode ledger dedup identified entries by parsing a prose description string. That is banned by CONTRIBUTING's raw-text-matching rule and was a real bug: the "exactly one window per untriaged proposal" guarantee rested on prose matching, so rewording a description or editing WINDOWS.md by hand silently produced duplicates. Now matches structurally on kind + phase + file + line. - A property test asserted on the stripper's output text. Reframed to assert the same invariant through analyzeSource's score. - nextBaseline's `candidates` parameter has been dead since the anchor change; removed from the signature and all call sites. - Extracted the duplicated require-or-degrade and capability-check boilerplate. - ADR-1953's Implementation bullet still named a `refactor.ship-gate` in check-command-router.cts — a leftover from the design cut D6 rejects. That file is untouched and no such gate exists. Removed. Refs #1953 * fix(#1953): keep execute-phase.md under its byte ceiling; un-vacuum the large-file test Five of the seven remote-runner failures were one cause: the execute:post dispatch contract, written out inline, grew execute-phase.md 1876 bytes (93,400 -> 95,276) against a frozen PRE_PHASE6 ceiling of 93,600. A drift-ack does not clear that — tests/phase6-capstone-conformance.test.cjs and tests/fix-2285-claude-orchestration-wiring.test.cjs assert the file is literally under the cap. The contract now lives in gsd-core/references/loop-hook-dispatch.md, which already claimed to be the point-agnostic dispatch reference and already documented ref.skill and ref.agent. It gains the ref.command shape, its in-context validation rule, the advisory-by-construction statement, and a note that a point whose workflow hand-rolls one kind is not implementing this contract. execute-phase.md now defers to it in one line: 145 bytes of growth, 55 B of headroom under the cap. Better placement than the first cut — the reference was overstating its coverage, and this makes the claim true rather than duplicating prose next to it. Acknowledged by appending to tests/emitted-drift-acks/2930-*.json rather than a new 1953-*.json: two ack sources may never name the same path, and that fragment is already the accumulating ack for this file. Sixth and seventh failures: analyzesLargeFileWithinBounds tripped its own vacuity guard — the fixture generated ~480 KB against a `> 500000` assert, so the guard fired and the three assertions after it never ran. The test has been vacuous since it was written. The matrix row specifies ~1 MB, so N goes 8000 -> 20000 (1.17 MB, 17% margin) and the guard to > 1_000_000. Verified by reproducing the exact body against the compiled module: 1168888 bytes, 118 ms, all four assertions hold. Refs #1953 * fix(#1953): fold the execute:post step deferral into the existing resolve line The remaining two failures were one test: execute-phase.md carries a SECOND, tighter assertion than the 93,600 ceiling — `<=93400`, which is exactly its current size. The file cannot grow by a single byte. My previous fix got it under 93,600 but not under 93,400, so it still failed. ("H." in the report is just the parent describe of that same test, not a separate defect.) Rather than add a paragraph, the deferral now REPLACES the existing hook resolution line. It read: Resolve active step hooks from `EXECUTE_POST_HOOKS_JSON` where `kind == "step"` and `ref.skill == "code-review"`. which is the bug itself written down — only code-review was ever dispatched. It now reads: Dispatch each `kind == "step"` hook per @gsd-core/references/loop-hook-dispatch.md. For `code-review`: The following prose already begins "If no active code-review step hook exists", so it reads correctly and the code-review handling is untouched. Net effect on the file is -11 bytes: 93,400 -> 93,389, under the margin assertion rather than merely under the ceiling. That also removes the need for a drift-ack: the file shrank, so there is no growth to acknowledge, and the append to the shared 2930-*.json fragment is reverted. Leaving it would have shipped a claim of "145 bytes of growth" that is no longer true, on a file six other issues share. The test's own comment states the principle this ended up honoring: "the host loop must stay small — optional-feature detail belongs in the capability fragment, not the host workflow." Putting the dispatch contract in the reference rather than inline is that rule, applied. Refs #1953 * fix(#1953): keep the code-review hook literal the workflow test requires tests/code-review.test.cjs extracts the <step name="code_review_gate"> block and asserts it contains `ref.skill == "code-review"` verbatim. The previous commit replaced the line carrying that literal, so the token vanished and the test went red — a fair assertion: code-review IS the bespoke branch there and the workflow should still name it. Restored inside the same one-line deferral, which now reads: Dispatch `kind == "step"` hooks per @gsd-core/references/loop-hook-dispatch.md. `ref.skill == "code-review"`: 93,396 bytes — still under the `<=93400` margin assertion and 4 bytes below the base, so the file continues to shrink rather than grow. Because three consecutive runs were each reddened by a different assertion on this one file, this change was verified by sweeping ALL of them at once rather than one run at a time: every test under tests/ that reads execute-phase.md or references/loop-hook-dispatch.md was located by resolving its path constants, and each content/size assertion was evaluated directly against the working tree — 22 assertions, plus two real executions (gen-section-manifest --check, and emitted-attribution's full real-tree differential). All pass. That sweep also confirms the earlier judgement call: the net change to execute-phase.md is a SHRINK, and the size ratchet only gates growth, so reverting the append to the shared 2930-*.json ack fragment was correct — an ack would have been both unnecessary and factually wrong. Refs #1953 * chore(#1953): backfill changeset pr number to 3261 * docs(#1953): add the missing how-to for acting on a refactor proposal Reference and explanation shipped (COMMANDS.md, CONFIGURATION.md, FEATURES.md 159, ADR-1953) but the Diataxis how-to quadrant did not, and that is the one a user reaches for. CONTRIBUTING's required-docs table is 'new command -> COMMANDS.md + FEATURES.md', so CI was green on a gap. Enabling this feature is genuinely multi-step and no single page walked it: turn it on, tune the threshold, understand advisory vs strict, discover that strict needs a SECOND toggle on a DIFFERENT capability, and know what to do when a proposal appears. The two-toggle subtlety in particular was a footnote in a config table; here it is a section with both commands. Follows the shape of its closest siblings, resolve-edge-coverage-findings and resolve-prohibition-findings — both 'the loop surfaced a finding, here is what to do with it'. Includes a reason-code table for the silent cases, since the analyzer is deliberately quiet in six situations and a user who expected a proposal needs to tell 'nothing to report' from 'could not look'. Indexed from docs/README.md beside the other loop how-tos. Docs-only: exempt from the push gate, no re-verification, pass marker on 2af188b4 untouched. Refs #1953 * feat(#1953): warn when strict mode is on but nothing will actually block Closes acceptance criterion 5, which I had wrongly marked satisfied. refactor.trigger_strict records an untriaged proposal as an open deviation window, but a ship only STOPS if workflow.windows_enforce is also on — a toggle owned by the broken-windows capability that this feature neither sets nor requires. So a user could enable strict, believe ship was gated, and find out otherwise at ship time. The split itself stays: requires:["broken-windows"] would force-install the ledger on advisory users who never enable strict, and a ship:pre gate of our own would never fire because ship.md has no generic ship:pre gate dispatch. What was missing was discoverability, so that is what this fixes. `refactor evaluate` now emits a typed REFACTOR_STRICT_NOT_ENFORCING warning, naming the exact remediation command, whenever strict is on and either workflow.windows_enforce is off or broken-windows is unavailable. It fires only on a run that produced a candidate — with nothing to block on there is nothing to warn about, and warning every run would be noise. Reads workflow.windows_enforce through the same resolveConfigKey walk the router already uses for its own keys rather than a second config reader. Four tests cover the matrix: strict+enforce-off warns, strict+enforce-on does not, strict+ledger-absent warns, strict-off never warns. Also corrects a user-facing message in this same file that told the user to run `gsd-tools config-set` — the wrong form. docs/CONFIGURATION.md and the broken-windows capability both use `gsd config-set`, and gsd-tools is invoked as `node gsd-tools.cjs`, so the bare form may not resolve. The two adjacent messages in this file now agree. Refs #1953 --------- Co-authored-by: sim <sim@local> |
||
|
|
58d73dd220 |
enhance(#3241): omit the codex per-agent model by default (#3276)
* test(#3241): failing-first suite for the codex passive model posture Locks ADR-2313's D1-D5 before any production code exists, so the tests bind to the behavior rather than to whatever the implementation happens to do. Red-first (fail against the current tree): - the resolver path emits no `model` and no `model_reasoning_effort` - a whitespace-only model_overrides value yields no pin - isAnthropicFlavoredModel / CLAUDE_AGENT_ALIASES on model-catalog - the one-time install notice, and its once-per-install dedupe Regression guards (pass today, must keep passing): resolver-null via `inherit` and via absent runtime; a resolver that resolves to nothing; empty-string and non-string overrides; and the light-tier service_tier/model_verbosity fields, which are NOT coupled to the model pin and would silently regress if the implementation coupled them. Classifying each test as red-first or regression guard is deliberate. A test that passes on both sides of the change proves nothing, and this epic has already shipped two such rows before catching them. The whitespace case is a live defect, not a quirk: `' '` is truthy, survives the type guard, is not Anthropic-flavored, and is embedded verbatim as `model = " "` — the same class the #2310 guard exists to stop. Same function, same path, fixed in this phase per CLAUDE.md §3. Two matrix rows were dropped as vacuous rather than shipped green: a 64-char truncation case (the pinned notice interpolates no user-controlled value, so it cannot exhibit truncation) and a newline hazard that the input surface cannot reach. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3241): omit the codex per-agent model by default Implements ADR-2313 D1-D5. generateCodexAgentToml no longer embeds the runtime resolver's per-tier Codex model, so an agent inherits the always-available session model instead of a pin a ChatGPT-account Codex may not expose. model_reasoning_effort disappears with it via the existing hasPinnedModel coupling (#838) — no logic change needed there. Supersedes #2517's embedding on the default path only. An explicit real-Codex model_overrides pin is still embedded verbatim, and the #2310 Anthropic-flavored guard is retained: the model_overrides route to it is still live even though the resolver route is now unreachable. The shared rule moves down a layer. CLAUDE_AGENT_ALIASES leaves model-resolver for model-catalog — a genuine leaf importing only node:path and its own JSON — with isAnthropicFlavoredModel defined beside it, and is re-exported from model-resolver so every existing importer is untouched. This is what lets Phase 2's install-check and Phase 3's sync consume the rule without taking the config-loader dependency model-resolver would have dragged into a module documented as pure read/verify with 33 dependents. A parity test fails if the two ever fork. Also fixes a live defect surfaced while writing the tests: a whitespace-only model_overrides value was truthy, survived the type guard, was not Anthropic-flavored, and so was embedded verbatim as `model = " "` — the same class the #2310 guard exists to stop, reached by a different route. Trimmed before the truthiness test. It is deliberately not routed through _warnCodexModelOverrideDropped, whose text would misdescribe a blank field as a mis-typed model. Adds the one-time install notice (maintainer direction, recorded as an ADR-2313 amendment): one stderr line naming model_overrides and the session model, deduped per install rather than per agent, and emitted only for the population that actually loses a pin. service_tier and model_verbosity stay decoupled from the model (#774); a regression guard asserts they still emit with nothing pinned. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3241): amend ADR-2313, add the model-catalog glossary entry ADR-2313 gains two dated amendments rather than edits to its merged text, since ADRs here are append-only. The first records that a deprecation notice IS offered, reversing the Migration section's "no deprecation window" position, and states why that position was wrong rather than just superseding it: the ADR identified the API-key population as losing something real and then declined to warn it, in the same document. Hyrum's guidance was applied to the recourse and not to the notice. The second records the whitespace-only model_overrides defect and notes that D2 always implied the fix — the implementation simply never enforced it and no test covered the case. CONTEXT.md gains a Model Catalog Module entry. The module had none, which is why the glossary gate passed without one: check-glossary-refs verifies that references resolve, not that modules are documented. The entry records why the Anthropic-flavored rule lives there rather than in model-resolver, so a later reader does not "helpfully" move it back. The Model Resolver entry is updated to point at its new home and note the back-compat re-export. docs/CONFIGURATION.md carried a claim that is now false: that the resolved tier ID is embedded into agent frontmatter at install time on codex and opencode. Corrected to name codex as the exception, with the 400 symptom and the model_overrides recourse. Changeset leads with the user-visible change and the migration line rather than the implementation, per the ADR's Hyrum's-Law analysis. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3241): only notice a lost pin when one was actually embeddable Review finding from an isolated reviewer. The deprecation notice gated on whether the runtime resolver would have returned *any* model, but the question that matters is whether that model would have been *embedded*. Those differ. The #2310 safety gate already rejected an Anthropic- flavored model arriving from the resolver path before Phase 1 — so for a mixed-runtime config resolving to a claude-* id against a Codex install target, the user never had that pin. The notice told them they lost something they never got, and pointed them at model_overrides for no reason. The existing #2310 test drives exactly that path but asserts only the emitted `model` line, never stderr, which is why it slipped through. Now covered. Deliberately unchanged: an Anthropic-flavored model_overrides value plus a legal resolver model fires BOTH the override warning and the notice. That is correct — pre-Phase-1 the guard dropped the override, execution fell through to the resolver, and the resolver's model was embedded, so that user did lose a pin. Two messages, two distinct true facts, and the prefixes differ (`gsd: warning — ` vs `gsd: notice — `) so the one-notice-per-install contract holds. A regression test now pins that behavior so it does not get "simplified" away. Of the three tests added, only the first is red-first; the other two pass on both sides by design and are labelled as guards — one against over-correcting the fix into silence, one against removing the intentional double message. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3241): reset the notice dedupe via a seam, not a require.cache bust The remote runner caught a regression I introduced: the #2760 post-write-validation test began failing with the validator override no longer intercepting. Cause, confirmed by trace rather than guessed: the new #3241 review tests deleted require.cache for bin/install.js and re-required it mid suite, to clear the notice's module-level dedupe flag. But runCodexInstall destructures `install` at file load, closing over the ORIGINAL module's exports. After the cache bust a second instance existed, so the test's `installModule.__codexSchemaValidator = ...` mutated the new object while the code under test still called the old one. The override silently stopped intercepting, the real validator ran and passed on GSD-emitted output, and the abort-and-restore path was never exercised. Cache-busting a module mid-suite breaks every later test that assumes a single instance, which every other test in the file is entitled to. So the fix is a seam, not a workaround: bin/install.js exports _resetCodexNoticeDedupeForTests(), and the three tests call it directly instead of reloading the module. The flag is module-level by design — the dedupe is per-install and install() already resets it — so a unit test driving generateCodexAgentToml directly needs an explicit way to reset it. That is now what it has. Swept the rest of the #3241 diff for the same hazard; this flag was the only shared module-level state introduced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3241): reset both codex dedupe stores, not just the notice flag Second incomplete fix, same class one layer down. bin/install.js keeps TWO module-level dedupe stores and the require.cache bust I removed had been papering over both; my replacement seam cleared only one. _codexModelOverrideDroppedWarned is a Set keyed `${agent}::${value}`. tests/codex-config.test.cjs:558 already emits for `gsd-executor::sonnet`, so by the time the review test using the same agent and value ran, _warnCodexModelOverrideDropped was a silent no-op and the expected warning never appeared. The seam now clears both stores and is renamed to say so. Its comment records that per-install dedupe lives in module scope deliberately and that this is the single sanctioned way for a unit test to clear it. Swept bin/install.js for every other module-scope mutable a test could latch. Two are inert (capability registries assigned once at require time; selectedRuntimes computed once from argv). One is a genuine latent hazard and is deliberately NOT folded in: attributionCache (:1654) memoizes getCommitAttribution by runtime name for process lifetime, so two in-process installs of one runtime with differing attribution config would collide. It is unreachable from any current test and is a different concern from Codex warning dedupe, so it stays out of this PR rather than widening it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3241): correct the codex tier-routing how-to The docs gate forced the task-oriented quadrant and found the worst defect in this change's documentation surface. docs/how-to/configure-model-profiles.md carried a section titled "If you want tiered models on Codex" telling users to set runtime:codex + model_profile:balanced, promising "GSD resolves each tier alias to the Codex-native model and reasoning effort defined in the runtime tier map." That is exactly the behavior this PR removes — a how-to page confidently instructing users to do something that no longer works, which is worse than a missing page because it fails at the moment of use. Rewritten to state that Codex does no tier routing, give the model_overrides pin as the supported alternative, and name the two constraints on what may be pinned: it must be a real Codex model id, and the account must actually expose it — GSD cannot verify the second, so the honest advice when unsure is to omit the pin. Carries an upgrade note for both account types, since the change is a no-op for ChatGPT accounts and a real loss for API-key ones. Also tightened the same page's claim that Codex "embeds the resolved model" at install time — now true only of an explicit override. The re-install instruction it supports is still correct and still needed, so only the premise moved. Both the required-docs set (COMMANDS.md + FEATURES.md) and lint-docs-required.cjs would have passed before this commit, since CONFIGURATION.md and the ADR had already moved. Neither checks the quadrant a user in trouble actually opens. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3241): backfill changeset pr number (#3276) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b7431a9259 |
feat(#1956): flag cross-artifact fact drift in the plan drift guard (#3259)
* test(#1956): failing-first contract for cross-artifact fact-drift pass * feat(#1956): flag cross-artifact fact drift in the plan drift guard * fix(#1956): correct config-key assertion and bidirectional lifecycle-lag exemption * docs(#1956): document the cross-artifact axis in the architecture reference * feat(#1956): decide the phase-status drift axis deterministically * fix(#1956): scope the progress-table lookup, abstain without a position section, rank deferred * docs(#1956): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
ffd5370464 |
fix(#2903): use the command form that actually works in reader-facing docs (#3047)
* fix(#2903): use the command form that actually works in reader-facing docs Docs told readers to type the colon form, which no runtime registers -- 18 of 19 runtimes use slash-hyphen and the 19th uses shell-var -- so anyone copying an example got an unrecognized command. Swept 178 occurrences across 53 files, locale mirrors included so they do not re-diverge from English. The colon form is a source-authoring token, not a user-facing one: install-time converters key on it to produce the hyphen form runtimes actually register. So the sweep is scoped, and three things are deliberately left alone: - ADRs, which are a historical record; editing their prose falsifies what was written at the time. - The legacy release-notes archive, pending a maintainer decision on whether it follows the same historical carve-out. Excluding it keeps a later reversal additive rather than a revert. - Source artifacts under commands, workflows and agents, where the colon form is load-bearing. Rewriting those would break the installed-skill guarantee across every runtime -- the single largest hazard here. The plugin namespace form is a real, separate token and survives untouched. Adds a lint enforcing exactly that boundary, since the correct form genuinely differs by directory and nothing previously caught the drift. Also fixes a hardcoded colon form in the capability-matrix generator. The sweep alone would have left the generated matrix disagreeing with the template that produces it, so the fix is at the source and the output regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2903): stop the sweep misquoting source frontmatter Adversarial review caught three lines where the sweep rewrote a citation of the literal YAML name: key from a source command file. That key genuinely is the colon form -- this change's own carve-out logic says source-authoring tokens keep it -- so the docs ended up misquoting the real files. One of the three is an acceptance-checklist assertion, which the sweep turned into a false statement. Restored the three citations to match their sources verbatim, surgically: where a line carried both a name: citation and a real reader-facing slash command, only the citation reverted and the command stayed corrected. The guard needed the same distinction, or it would have flagged the restoration and reddened the build: a gsd:<cmd> token preceded by name: is a citation of a source token and is now permitted. The exemption is deliberately narrow -- a bare gsd:<cmd> anywhere else still fails -- with a test pinning that narrowness. Also makes the detection case-insensitive. Review found /GSD:next slipped through silently; no such casing exists in the tree today, so this closes a latent gap rather than fixing a live one. Swept the whole tree for further corrupted citations: none beyond the three. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2903): retire the stale-next invariant and sweep next like every other command Maintainer decision on a genuine conflict between two contracts. Invariant #3054 banned the literal /gsd-next from user-facing docs because it named a retired workflow-advance command. But commands/gsd/next.md is a live command -- the state-aware smart-entry launcher -- and this issue requires docs to use the hyphen form every runtime actually registers. Both could not hold for this one command, so docs had been sidestepping the ban by keeping the colon form, which is exactly the defect this issue exists to remove. FEATURES.md already recorded the reassignment: the hyphen form "is not the retired workflow-advance command; it is reserved for the state-aware smart-entry launcher. Workflow advancement remains under /gsd-progress --next." With that reassignment the invariant's premise is obsolete and the guard now contradicts the documented command form, so it is retired with a comment recording why rather than deleted silently. next is now swept like every other command, and the earlier exemption added to the new guard is removed so nothing is special-cased. Four citations of the literal name: frontmatter key stay in colon form, because the source file really does carry name: gsd:next and a doc quoting it must reproduce it verbatim. Two of those lines were reworded to say which side is the frontmatter key and which is the slash command, since they previously conflated the two. Verified the retired scan would now genuinely fail against this tree -- the conflict was real and resolved, not dodged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2903): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
067a4d1c6c |
fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls (#3015)
* test(#2650): add failing-first regression for plan-phase stall detection Regression test for gsd_stall_should_recover / gsd_stall_watch and the planner.stall_* config keys, none of which exist yet — proves RED before the fix lands in the next commit. * fix(#2650): bound and auto-recover plan-phase planner/plan-checker stalls Mirrors the already-shipped executor.stall_* pattern (execute-phase.md, bug #3212) but with a dispatch change the executor's prose-only surveillance lacks: the standard planner spawn, chunked-outline planner spawn, chunked-per-plan planner spawn, plan-checker spawn, and revision-loop planner respawn now dispatch with run_in_background=true and are followed by a real, bounded bash poll (gsd_stall_watch) that returns control to the orchestrator on its own schedule instead of waiting indefinitely on a subagent that may never return. On stall, the existing accept-plans/retry/ stop recovery menu (9a/11a) is auto-surfaced instead of requiring a manual interrupt. New config keys planner.stall_detect_interval_minutes (default 5) / planner.stall_threshold_minutes (default 10) mirror executor.stall_*. The helper functions (gsd_stall_should_recover, gsd_stall_watch) live in a new lazily-loaded gsd-core/workflows/plan-phase/steps/stall-detection- helpers.md rather than inline, and per-site prose is kept minimal, because plan-phase.md is frozen under the ADR-857 Phase 6 PRE_PHASE6 gate (tests/phase6-capstone-conformance.test.cjs) with ~36 bytes of headroom at baseline; the net effect is plan-phase.md.md ships slightly SMALLER than before (the old unconditional-wait ORCHESTRATOR RULE sentences are gone at the five touched sites, superseded by the bounded watcher). Also fixes a stale doc comment in tests/workflow-size-budget.test.cjs that still described the per-file workflow-size-baseline.json guard removed by #2724 (ADR-2719 Phase 4) as if it were still the enforcement mechanism — discovered while verifying this fix's own byte budget. Researcher and pattern-mapper spawns are untouched (out of scope per the issue's Agent Brief). * fix(#2650): make gsd_stall_watch single-cycle; harden numeric config inputs Two review findings addressed on top of the prior commit: 1. gsd_stall_watch previously looped internally for the full threshold+interval duration inside ONE Bash tool call (up to 15 min at defaults) — a single call blocking that long risks the host tool's own timeout killing it before it ever prints a result, silently defeating the fix. Redesigned to a single sleep-and-check cycle per call, taking an explicit dispatch_ts so the orchestrator prose can repeat the (short, default 5 min) call until it resolves; the outer threshold is now enforced by dispatch_ts accumulating across calls, not by one call's duration. Documented the resulting trade-off (up to one interval of added latency on the success path) in the changeset and reference doc. 2. PLANNER_STALL_INTERVAL_MINUTES/THRESHOLD_MINUTES are config-controlled values that flow into bash arithmetic ($(( ))). A review flagged this as command injection; empirically verified against both macOS bash 3.2.57 and Docker bash:5 that this is NOT actually exploitable (bash hard-errors on a `$(cmd)`-shaped arithmetic operand rather than invoking it) — but an unvalidated malformed value WOULD abort the stall-watcher itself with that bash error, silently defeating the exact hang-recovery this issue ships. Added integer validation with safe-default fallback, both at the config-resolution point and defensively inside gsd_stall_should_recover. Also adds the previously-missing integration coverage for gsd_stall_watch's real execution (grep/find/date plumbing), not just the pure classifier. * fix(#2650): correct AC2 self-test — helpers doc may name teams-status in prose The AC2 regression test asserted the stall-detection-helpers.md step file never contains the substring "teams-status" at all, but the file's own prose explicitly documents its independence from that guard (containing the word by design). Narrowed the assertion to what actually matters: no second `query teams-status` call site and no gating on it, not a blanket absence of the word. * test(#2650): regenerate golden install-tree fixtures for the new step file gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md is an emitted file (installed for every runtime), so adding it changes the install tree even though it is invisible to docs/INVENTORY.md and docs/INVENTORY-MANIFEST.json (both explicitly scope to non-recursive gsd-core/workflows/*.md — verified against the execute-phase #2930 and pre-existing plan-phase step-file precedent, which are equally absent from both inventory artifacts). The golden install tree snapshots the sorted list of emitted relative paths per runtime, so a file invisible to the inventory is still visible here. Regenerated via `npm run gen:install-tree` — one line added per runtime fixture (19 files), no other drift. * fix(#2650): restore 7 ORCHESTRATOR RULE labels; sync runtime-launcher preamble Two more consequences of extracting helper bodies out of plan-phase.md, both caught by verification (0017e1a78, 9 unique failures): 1. tests/plan-phase-drift-guard.test.cjs (#913) requires at least 7 "ORCHESTRATOR RULE — ALL RUNTIMES" labels in plan-phase.md itself, one per agent spawn site. Moving the full explanatory blocks to plan-phase/steps/stall-detection-helpers.md carried 5 of the 7 labels out with them (only the untouched researcher/pattern-mapper sites kept theirs). Restored a short label at each of the 5 stall-watch sites, trimmed a few more redundant words ("Per 7.99, " — already established by the adjacent step-7.99 pointer) to stay under the frozen PRE_PHASE6 cap (94497 bytes, 21 bytes headroom). 2. tests/runtime-launcher-parity.test.cjs (#373) requires exactly one canonical gsd_run preamble, byte-equal to gsd-core/workflows/_runtime-launcher.snippet.sh, before the first gsd_run call in any workflow .md that calls it (recursive scan under gsd-core/workflows/, unlike the non-recursive inventory/step-tag-balance checks). The new step file's config-get calls use gsd_run without one. Fixed via `node scripts/sync-runtime-launcher.cjs`, verified: exactly 1 preamble occurrence, before the first call, including the .claude/ and .codex/ home fallback arms. Also verified (no fix needed, evidence recorded): the generic `gsd-core-verbatim` identity rule in tests/helpers/emitted-provenance.cjs (roots: ['gsd-core'], pattern matching workflows/.+) self-attributes any new gsd-core/workflows/** path to itself, so the new step file needs no drift-ack entry — consistent with plan-phase.md's own net shrinkage requiring none either. * test(#2650): acknowledge plan-phase.md's +14 byte drift Restoring the 5 ORCHESTRATOR RULE — ALL RUNTIMES labels (#913) flipped plan-phase.md from -142 bytes (post-extraction) to +14 bytes net growth against baseline (94483 -> 94497), which the differential attribution size ratchet (tests/emitted-attribution.test.cjs) correctly flags as unacknowledged growth. Added tests/emitted-drift-acks/2650-plan-phase- stall-detection.json, keyed on the bare filename plan-phase.md per the existing fragment schema (see tests/emitted-drift-acks/2649-diagnose- execute-plan-base-check.json), explaining the growth as exactly the 5 restored labels — still verified under the PRE_PHASE6 cap (94497 < 94519) and satisfying #913's 7-label requirement. * fix(#2650): bind {outputFile} from the real Agent() return — was dead code Independent review blocker: PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE were read by every gsd_stall_watch call but never assigned anywhere in the diff. With the variable permanently empty, `[ -f "$output_file" ]` was always false, marker_found could never become true, and marker_received was unreachable — the marker-based detection path was permanently dead. Worse for the plan-checker spawn specifically: a checker that PASSES touches no *-PLAN.md files, so it had no working completion signal at all without the marker path. A healthy plan-checker finishing cleanly in two minutes would be declared stalled once planner.stall_threshold_minutes elapsed and the recovery menu would fire on an already-succeeded agent — worse than the original unbounded hang. Fixed by replacing the dead bash variable with the `{outputFile}` orchestrator-substitution token, the same convention docs-update.md:471 already uses for a real run_in_background=true Agent() return ("Read tool: file_path: `{outputFile from README agent result}`"). This is a net BYTE SAVING at each site (`"{outputFile}"` is shorter than `"$PLANNER_OUTPUT_FILE"`), which funded moving the full binding explanation — including why plan-checker's *-PLAN.md glob alone is not a working completion signal — into the lazily-loaded reference file to stay under the frozen PRE_PHASE6 cap (94496 bytes, 22 headroom; net +13 over baseline, acknowledged in tests/emitted-drift-acks/2650-plan-phase- stall-detection.json). Added a regression test asserting plan-phase.md itself binds {outputFile} at all 5 spawn sites and contains no dangling $PLANNER_OUTPUT_FILE / $CHECKER_OUTPUT_FILE reference — the previous test suite only exercised gsd_stall_watch's behavior when handed a valid argument, which is why the dead production wiring survived two rounds of review. Also fixed tests/fix-2650-plan-phase-stall-detection.test.cjs:170-195's raw try/finally to use t.after(), per CONTRIBUTING's test-cleanup convention. * chore(#2650): backfill changeset PR number to 3015 * fix: normalize CRLF at the read boundary in all .md-bash-extraction tests Maintainer-authorized scope expansion, folded into this PR rather than deferred: the Windows CI lane on this PR's own tests/fix-2650-plan-phase- stall-detection.test.cjs exposed DEFECT.TEST-SHELL-PIPELINE-NONPORTABLE (CONTEXT.md; recurring since #1700) as a repo-wide latent class, not a one-off. Ten test files parse a fenced ```bash block out of a workflow .md file and execute it via spawnSync/execFileSync; a Windows checkout can yield CRLF line endings despite .gitattributes eol=lf, and bash then treats the trailing \r on every extracted line as part of the token — "unexpected EOF while looking for matching `"'" or a bare syntax error, partway through the script. Added tests/helpers.cjs:readFileNormalized() — strips \r\n -> \n at the read boundary, before any fence-slicing or regex runs, so every downstream operation is correct by construction. Migrated all ten call sites to it: Previously broken (fs.readFileSync with no normalization anywhere between read and spawn): - tests/worktree-cleanup.test.cjs (extractCwdGuardBash) — also fixes a misleading comment claiming the fence regex alone was "CRLF-safe"; it protected only the fence delimiters, never the captured body. - tests/new-milestone-clear-phases.test.cjs (extractFenceBetween, extractFenceContaining) - tests/code-review-pipeline-regression.test.cjs (extractPostProcessingScript) - tests/drift-detection.test.cjs (readGate/bashBlock, plus the snippet-file comparison read in the same test) - tests/graphify-visualization.test.cjs (extractStep3Block) - tests/pause-work-improvements.test.cjs (extractCheckBlock) - tests/plan-review-convergence.test.cjs (extractReviewerFlagsParseBlock and the inline post-config-gate resolution-block slices) Already correct (split(/\r?\n/) then join('\n')), migrated to the shared helper for consistency rather than a fourth/fifth/sixth copy of the same fix: - tests/git-base-branch.test.cjs (extractHandleBranchingBash) - tests/quick-branching.test.cjs (extractStep25Bash) - tests/runtime-launcher-parity.test.cjs (extractResolverSnippet) Verified against a simulated Windows CRLF checkout (not assumed): for both the worktree-cleanup.test.cjs and new-milestone-clear-phases.test.cjs extraction shapes, confirmed the pre-fix code produces a real bash syntax error on CRLF input and the post-fix code does not. One eslint follow-up: local/no-crlf-fragile-split statically flags any bare `\n` inside a markdown-fence-shaped regex, regardless of whether the receiver was already normalized — it cannot see the readFileNormalized() data-flow. Kept `\r?\n` in extractCwdGuardBash's fence regex (redundant but harmless on pre-normalized input) rather than fight the rule. Scope note: this diff is broader than issue #2650's own change (plan- phase.md stall detection) because the Windows lane surfaced a genuine repo-wide defect class while verifying that fix, and the maintainer authorized fixing it here rather than filing it separately and shipping a known-broken pattern. Runtime impact: none — this is a test-harness-only defect. The live orchestrator (Claude Code or another runtime) does not do a byte-exact extract-and-pipe of .md content into a shell the way these tests do; it reads the instructions and generates its own bash invocation text, which does not reproduce a raw CRLF pass-through the same way. Not touched: tests/plan-review-convergence.test.cjs's separate, tracked spawnSync ETIMEDOUT flake under bench load (#3005, reproduced on unmodified next) — unrelated load-sensitivity, not a CRLF symptom. * fix(#2650): remove stale drift-ack fragment — plan-phase.md is self-explaining tests/emitted-drift-acks/2650-plan-phase-stall-detection.json acknowledged plan-phase.md's own emitted-path hash move, but plan-phase.md is directly edited in this diff. Per the emitted-attribution law (ADR-2719, tests/emitted-attribution.test.cjs), a workflow's emitted key equals its own source path (gsd-core-verbatim identity rule), so a direct edit to the source is self-explaining and auto-attributed — no ack was ever needed. Verified via the pre-merge lint (scripts/lint-emitted-drift-ack.cjs, run through npm run lint:ci with a fully cleared eslint cache): it passes clean with the fragment removed, confirming no contradiction between the lint and the runtime attribution gate — this was simply an unnecessary fragment. * fix(#2650): restore plan-phase.md drift-ack — size ratchet demands it against next tests/emitted-drift-acks/2650-plan-phase-stall-detection.json was deleted in the previous commit because, against an earlier verification base, it was inert: it explained a moved emitted hash that a direct edit to plan-phase.md already self-attributes. Against origin/next@f1af47766a the demand is different: plan-phase.md is 13 bytes larger than the base copy, which trips the emitted-attribution size ratchet — a job this same ack also performs. Recreated in the documented shape, keyed on the bare filename plan-phase.md (not the full path, and not restating the byte delta per review guidance), describing the actual change: the {outputFile} binding fix for the dead PLANNER_OUTPUT_FILE/CHECKER_OUTPUT_FILE variables and the 5 restored ORCHESTRATOR RULE labels required by #913, both at the stall-watch spawn sites, with explanatory bodies living in the lazily-loaded gsd-core/workflows/plan-phase/steps/stall-detection-helpers.md reference. Confirmed no other fragment (on this branch or on next) claims the bare key "plan-phase.md" before recreating — scripts/lint-emitted-drift-ack.cjs's duplicate check is an exact string match, and the only other mention of plan-phase.md in tests/emitted-drift-acks/ (2658-trae-instruction-file-path.json) uses the full path as its key, so there is no collision. * fix(#2650): real cause of Windows CI failure — bash -c argv-transport, not CRLF The CRLF diagnosis for PR #3015's Windows failure was wrong. Proven wrong, not assumed: .gitattributes' blanket `* text=auto eol=lf` means a Windows checkout never receives CRLF for stall-detection-helpers.md, and the extracted fence's line 64 is byte-identical and correctly balanced on every platform. The real cause: runShouldRecover() passed a 70+ line, quote-dense script as ONE argv element to `spawnSync('bash', ['-c', script, arg0, ...])` PLUS four more positional args. Windows has no execve — Node serializes that whole argv into a single CreateProcess command-line string, and Git Bash's MSYS layer re-splits and unescapes it with its own rules. The boundary between the script and the trailing args was not stable across that round trip (live evidence: one failure's stderr was prefixed `gsd_stall_should_recover_test:` — arg0 arrived — another `/usr/bin/bash:` — arg0 did not). Fixed by writing the script to a temp file and running `bash <file> <args>` instead — the four values are now normal, quote-free positional args, and the script itself never enters argv transport at all. Mirrors tests/quick-branching.test.cjs's extractStep25Bash/runStep, which already uses this exact shape and is green on Windows on `next`. tests/worktree-cleanup.test.cjs's extractCwdGuardBash/runGuard stays on `bash -c` but never appends extra positional args beyond the script itself, so it never hits the same boundary — checked both siblings per review, not assumed. Corrected the now-actively-misleading CRLF comment in extractStallHelpersBash(), and corrected the changeset's claim that the repo-wide CRLF-normalization fix (folded into this branch, maintainer- authorized) explains this PR's own Windows failure — it doesn't, though it remains defensible on its own merits as general test-portability hardening. Separately, while auditing the shipped (non-test) gsd_stall_watch for Windows portability per review request, found and fixed a second, real user-facing defect: the artifact-freshness check used GNU find's `-newermt "@<epoch>"` shorthand, which the BSD find(1) actually shipped on macOS does NOT understand ("Can't parse date/time: @<epoch>", verified live against /usr/bin/find on both a stale and a genuinely fresh file). With the adjacent `2>/dev/null`, that failed silently and permanently degraded artifact_fresh to false on every macOS run — a plan-checker or planner actively writing plan files could still be reported "stalled." Replaced with `find $glob -mmin -N` ("modified less than N minutes ago"), which needs no date-string parsing and is supported identically by GNU find and BSD find; verified live that the old shape fails and the new shape passes against the same real fresh file. Added a real-execution regression test (gsd_stall_watch with `sleep` stubbed to a no-op so the test doesn't actually wait, but the real `find ... -mmin` line still runs) proving the fix, replacing the prior "not integration-tested" note for that path. Note: the remote gsd-test runner is Linux-only, so it cannot itself confirm the Windows fix — only the actual windows-latest CI lane can. * fix(#2650): route the third bash -c call site through the same temp-file seam runWatch() and a `-mmin` regression test still passed their script via `bash -c <script>` after the previous commit only converted runShouldRecover() — live Windows CI on 4b86cc57f confirmed the mechanism: failures went 11 -> 4, and `full test (windows-latest, 22, shard 1/3)` and `shard 2/3` flipped from fail to pass, but the remaining 4 failures (all in this file, all still `bash: -c:`) were exactly the gsd_stall_watch describe block, which runWatch() serves. runWatch() passes NO extra positional args at all, so this also rules out the trailing-args theory from the prior commit: the ~73-line, quote-dense script itself is what does not survive Windows argv serialization when passed as a single `-c` element, regardless of how many (if any) further argv elements follow it. Extracted one shared runBashScript(script, args, opts) helper — write to a fs.mkdtempSync'd file, run `bash <file> [args...]`, clean up in `finally` — and routed all three bash-invoking call sites in this file through it (runShouldRecover, runWatch, and the -mmin freshness test that builds its own script inline for the `sleep` stub). One transport seam means a fourth call site in this file cannot silently reintroduce the bug in isolation, which is exactly what happened here with a second call site. Corrected extractStallHelpersBash()'s doc comment a second time to state the mechanism precisely (script content, not argv-element count) and cite the live evidence (11->4 failures, shards 1 and 2 flipping green) so the next reader does not have to rediscover it. Audited every other bash-invoking call site in files this branch touches, per review request: - tests/code-review-pipeline-regression.test.cjs (runPostProcessing), tests/graphify-visualization.test.cjs (runBlock), and tests/drift-detection.test.cjs (two execFileSync('bash', ['-c', ...]) sites, one of them carrying the same giant runtime-launcher preamble text) — all pre-existing, UNCHANGED by this branch (only touched for the readFileNormalized() CRLF swap), and already exercised on `next`'s last six Windows CI runs per the reviewer's own citation. Left as-is: no evidence of failure, and converting untested pre-existing code outside #2650's scope on an unverifiable guess would be its own risk. - tests/git-base-branch.test.cjs (runHandleBranchingStep) and tests/quick-branching.test.cjs (runStep) already use the same temp-file pattern. No action needed. - tests/runtime-launcher-parity.test.cjs (runResolver) uses `bash -c` but is explicitly `if (process.platform === 'win32') return '';` guarded off on Windows entirely, for an unrelated extension-less-PATH-stub reason — never reaches Windows argv transport at all. No action needed. - tests/worktree-cleanup.test.cjs (runGuard) confirmed by the reviewer as correct and verified; not touched, per instruction. Do not touch: the -mmin fix, the drift-ack fragment, the changeset — all three confirmed correct in prior rounds and left untouched here. Note: the remote gsd-test runner is Linux-only and cannot confirm this; only the windows-latest lanes on #3015 can. * fix(#2650): give runBashScript a default timeout runShouldRecover() was the only one of the three call sites through runBashScript() with no timeout — runWatch() and the -mmin test both pass timeout: 10000 explicitly. Not a regression (this path never had a bound before), but CONTEXT.md's unbounded-subprocess guidance applies directly, and runShouldRecover() is driven repeatedly by a fast-check property test: one pathological input that fails to terminate would hang CI indefinitely instead of failing. timeout: 10000 is now the helper's own default, with ...opts spread after it so the two existing explicit timeout: 10000 call sites are unchanged and any future caller inherits a bound automatically. * fix(#2650): build the -mmin freshness test's glob with forward slashes Windows CI on d6ddda6ea reported the last failure: the -mmin regression test expected 'active' but got 'waiting' — find matched nothing, the same silent-degradation shape as the macOS -newermt defect, but this time in the test's own fixture rather than the shipped bash. Traced what production actually passes: every gsd_stall_watch call site in plan-phase.md builds artifact_glob as `"${PHASE_DIR}"'/*-PLAN.md'` — PHASE_DIR is a POSIX-style .planning/phases/NN-slug value, and the whole thing runs under Git Bash regardless of host OS, so production's glob is always forward-slash. The test instead built it with `path.join(tmp, '*-PLAN.md')`, which on Windows yields a backslash path (C:\Users\RUNNER~1\...\*-PLAN.md). In bash pathname expansion a backslash escapes the next character, so that pattern can never match a real path — find silently returns empty under the existing 2>/dev/null, same shape as the macOS bug. Confirmed as a test artifact, not a production defect: production never constructs the glob this way, so no Windows user is affected. Fixed by forward-slashing the tmp dir before appending the glob suffix, matching production's own convention, with a comment recording why (so a future "simplify this back to path.join" edit doesn't silently reintroduce the failure). The shipped bash's unquoted $artifact_glob is untouched — quoting it would break the multi-file glob expansion it exists for. Note: the remote runner is Linux-only and already passed clean at d6ddda6ea (0/29,603, both node lanes); only the windows-latest lanes on #3015 can confirm this fix. * fix(#2650): forward-slash the three remaining runWatch globs (vacuous-pass CR) The :353 fix (833c11da9) only converted the -mmin freshness test's glob. Three sibling tests in the same describe block still built theirs with path.join(tmp, '*-PLAN.md'), which yields a backslash path on Windows. Two of those three were silently passing for the wrong reason: the '-> stalled' and '-> waiting' tests both expect the glob to match nothing, and on Windows a backslash path matches nothing regardless of whether the directory is actually empty (bash eats each backslash as an escape before the pattern is even evaluated). They would have passed identically with glob expansion completely broken, which is a vacuous pass — not exercising what they claim to. The third ('-> marker_received') is outcome-independent of the glob, so it was merely inconsistent rather than wrong. Converted all three to the same `${tmp.replace(/\\/g, '/')}/*-PLAN.md` construction already used at the -mmin test, so every glob in the file now matches production's own forward-slash `"${PHASE_DIR}"'/*-PLAN.md'` shape, and the two negative tests are meaningful on Windows instead of accidentally correct. Reworded the trailing comment on the 'stalled' test's glob line: it now describes the fixture (the tmp dir contains no *-PLAN.md files) rather than the pattern, since "matches nothing" read as a property of the glob syntax when it's a property of what's on disk. No assertion, the sleep stub, runBashScript, or the shipped bash changed. Smoke-tested all three updated tests manually before committing (not via node --test): marker_received / stalled / waiting, all correct. * fix(#2650): fix own regression tests for #2993's plan-phase.md relocation 531101843's merge with origin/next brought in #2993 (unrelated, epic #1671 Phase 6.2), which extracted plan-phase.md's whole "Chunked Planning Mode" section into gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md, leaving a <!-- gsd:section --> pointer behind. tests/plan-phase-drift-guard. test.cjs (#913) was already updated to read the combined surface (host file + every steps/*.md) so its label count didn't go blind — my own #2650 regression tests were not, and searched plan-phase.md alone for the two chunked spawn sites' headings, which no longer exist there. Two tests failed outright (indexOf returning -1); a third ("standard planner spawn") was silently weakened to an unbounded slice-to-EOF by the same relocation, since its own end-boundary heading also moved — passing by accident rather than by testing what it claimed. Promoted the drift guard's local readPlanPhaseCombined() to a shared, exported tests/helpers.cjs readWorkflowCombined(workflowPath) (host file + sorted steps/*.md, CRLF-normalized at the read boundary) so a second, divergent implementation is never written — the drift guard now delegates to it via a same-named local wrapper, unchanged at every existing call site. Fixed the three affected tests in tests/fix-2650-plan-phase-stall-detection. test.cjs: - "standard planner spawn (step 8)": end boundary changed from the now-gone "## 8.5. Chunked Planning Mode" heading to "## 9. Handle Planner Return", which still exists in plan-phase.md. - "chunked outline spawn (8.5.1)" / "chunked per-plan spawn (8.5.2)": now read gsd-core/workflows/plan-phase/steps/chunked-planning-mode.md directly (not the generic multi-file combined blob, whose file-sort ordering would put unrelated step files between 8.5.2's slice and any downstream anchor) — the same heading-to-heading slicing as before still works because the file is small and self-contained. - Extended the "no unbound $PLANNER_OUTPUT_FILE/$CHECKER_OUTPUT_FILE" check to also scan chunked-planning-mode.md, since two of the five spawn sites now live there. - Added a new count-based test asserting exactly 5 (not "at least one") `gsd_stall_watch "$TS" "{outputFile}"` invocations across the combined surface, mirroring #913's own label-count guard, so every one of the five spawns stays provably bounded and a future relocation can't silently drop one without a test noticing. Also added a small positive test that plan-phase.md's <!-- gsd:section --> pointer to chunked-planning-mode.md exists (#2993 is unrelated to #2650 but its presence is now load-bearing for where 2 of the 5 spawn sites live). Audited every other test file in the repo for a stale reference to content #2993 relocated (searched for the moved headings/prose and for "chunked-planning-mode"/"CHUNKED_MODE" across all *.test.cjs): only this file and the drift guard needed changes. tests/issue-2762-plan-reviews-chunked.test.cjs already reads chunked-planning-mode.md directly (brought in correct by the same merge). gen-section-manifest.test.cjs, init.test.cjs, and workflow-fragments.test.cjs reference "chunked-planning-mode" only as a manifest/section-id fixture value for #2993 itself, not as a stale pointer to relocated content. Did not touch: the ported ORCHESTRATOR RULE lines, run_in_background=true, the glob constructions, runBashScript, the -mmin change, the timeout default, or the drift-ack fragment (confirmed correct against the stale local `next` ref two rounds ago and left alone). --------- Co-authored-by: sim <sim@local> |
||
|
|
7b204ad2ac |
enhance(#2530): extend UAT checkpoint frame language pack (9 more languages) (#2564)
* feat: extend UAT checkpoint frame language pack (9 more languages) response_language is a free-form config value, but CHECKPOINT_FRAMES only covered 9 languages — any other configured language silently fell back to the English frame. Add Dutch, Polish, Russian, Ukrainian, Turkish, Hindi, Arabic, Vietnamese, and Indonesian frames plus their aliases, with a regression test asserting each resolves instead of falling back. Follow-up to #2402 (PR #2457). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: add changeset for #2527 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs(#2530): list UAT checkpoint frame languages in CONFIGURATION.md Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(#2530): point changeset fragment at PR #2557 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(#2530): address Unicode language-pack review * fix: address checkpoint language review * fix: count spacing combining marks in checkpoint width * test: verify checkpoint aliases structurally * fix: isolate RTL checkpoint frames * fix: isolate RTL checkpoint frames correctly * test(#2530): assert checkpoint aliases neither collide nor go unreachable Review Minor #1. A duplicate alias key was invisible to the existing catalog tests: the runtime object is well-formed after JS collapses the literal, the self-alias assertion still holds, and the losing language just stops resolving. tsc catches the byte-equal case (TS1117), but not the two that survive compilation — an alias whose NFC-lowercase form already belongs to another language, and an alias not in lookup form at all, which resolveCheckpointFrame() can never produce. The check reads the source literal rather than the object, since the object no longer records what was written. Both assertions are independently load-bearing: an NFD twin of an existing alias trips the collision check, an uppercase alias trips the unreachability check. Review Minor #2: changeset retyped Changed -> Added. Nine wholly new supported response_language values are an addition under Keep a Changelog, not a modification of existing behavior. * test(#2530): check alias collisions on the catalog, not its source --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> Co-authored-by: Rezolv <dave@sienkowski.com> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
0408276791 |
chore(#2797): federate reviewer config keys off the central schema (#2841)
* chore(#2797): federate reviewer config keys off the central schema Phase 4 of epic #2782 (ADR-2782 D9, config half). Runs AFTER 5a per the ADR's swap amendment: a federated config slice lives inside a capabilities/<id>/capability.json, and three of the five key families had no capability directory until 5a created them. Four key families move to the lanes that use them; the central-schema removal and the federated addition land in this one commit because the exclusivity invariant fails the build on a key present in both. review.max_prompt_tokens, review.default_reviewers and review.reviewer_instances describe policy ACROSS lanes and stay central. Two things the issue did not name, both found while building it: 1. THE EXCLUSIVITY GATE WAS BLIND TO PATTERNS. It compared federated keys against manifest.validKeys only, and two of the four families (review.models.<slug>, review.max_prompt_tokens_per_reviewer.<slug>) were pattern-backed. That is not cosmetic: isCentralConfigKey consults those patterns and mergeFederatedConfig skips every key for which it returns true, so declaring a slice while the pattern survived would have shipped an INERT slice behind a green gate — the exact half-migrated shape the invariant exists to prevent. The gate now loads the patterns from the same manifest the runtime reads. 2. AN UNSET PER-LANE BUDGET NOW RESOLVES TO 0, NOT NOT-FOUND, because a federated key always resolves to its declared default. The three fallback guards in review.md checked only empty-or-"null", so a user who set the GLOBAL review.max_prompt_tokens would have silently lost trimming on the HTTP lanes. The guards now treat 0 as unset. D9 says review.models.<slug> is owned by "the lane whose slug it names". That is false for one lane: the shipped key is review.models.agy while the slug is antigravity. Ownership follows the lane; the key name is preserved, because renaming would break every config that sets it. Existing tests updated rather than left asserting the old world: config-get on a cleared federated key yields empty instead of not-found (what #2046 actually protects — never persisting the literal "null" — is unchanged and still asserted); the config-schema dynamic pattern representative moves to reviewer_instances; the prototype-pollution guard case moves to a surviving dynamic prefix so alert #26 keeps its coverage, with a new case asserting the old key is now rejected earlier; and Phase 2's harvest-widening inertness assertion becomes an ownership assertion, since Phase 4 is what consumes it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): use a -1 sentinel so an explicit per-lane budget of 0 survives A federated config key always resolves to its declared default, so an unset per-lane prompt budget needed a value the workflow could treat as 'not configured'. The first cut used 0 — which is wrong: 0 is already a LEGITIMATE per-lane budget meaning 'do not trim this lane' (the early-return guard in prepare_trimmed_prompt_for_reviewer). Treating it as unset would have silently switched a user who deliberately disabled trimming for one lane onto the global budget. The sentinel is now -1, which is not a valid token budget, so all three states stay distinguishable: unset falls back to global, an explicit 0 disables trimming for that lane, and an explicit N is used. Locked by three CLI round-trip tests. Surfaced by the isolated security reviewer before it crashed mid-run; verified independently against the shipped trim guard rather than taken on trust. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): update central-registration assertions and stay under the review.md cap The remote runner caught both; my local sweep missed the files. 1. tests/plan-review-convergence.test.cjs asserted the three local-server host keys are in VALID_CONFIG_KEYS. They are federated to their lane capabilities now, and the exclusivity invariant forbids a key living in both places. What #2306-local actually protects is that config-set ACCEPTS them, so that is what is asserted — via isValidConfigKey, the predicate config-set itself uses, which spans central and federated. A second assertion pins federated ownership, so a silent reversion back to the central schema fails too. 2. review.md exceeded the LARGE tier hard cap (62583 > 61440). That cap is a red line, not a budget to raise. The three per-lane budget guard comments were near-identical; condensed to one terse line each. 61371 bytes, 69 to spare. Real extraction to workflows/review/modes/ is Phase 5b/6 work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2797): fail closed on a broken config-schema manifest; reconcile stale docs Isolated security review findings. MAJOR — loadCentralConfigPatterns failed OPEN. It swallowed a JSON parse error and returned [], while its sibling loadCentralConfigKeys, reading the SAME file, writes to stderr and throws ExitError(1) on that identical failure class. Fail-open here defeats the gate this function exists to feed: with zero patterns, validateCrossCapability's pattern-collision check silently passes and an inert federated slice ships green. It was masked in the one production call site only because loadCentralConfigKeys runs first against the same path — a coincidence of ordering, not a guarantee, and this function is exported and called standalone. The two now share a contract: ENOENT is the legitimate absent case, anything else throws loudly. A single unparseable PATTERN is still skipped, which degrades to "checked less" rather than blocking every build. The branch had zero coverage; it now has two tests (malformed JSON, EISDIR). MINOR — docs/CONFIGURATION.md still listed review.models.qwen and review.models.cursor as settable, ~770 lines below this PR's own new Ownership section. Those lanes take no model flag, so they declare no model key and config-set now rejects them. Rows removed; the missing review.models.agy row added; the per-reviewer budget row corrected to name only the lanes that own a budget key, and to document that a per-lane 0 disables trimming for that lane. Also fixes a shadowed "raw" binding introduced by the fail-closed change, which made the generator unrequirable — caught immediately by its own --check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2797): backfill changeset pr number to 2841 * fix(#2452): make the base-ref mutation test hermetic against leaked GIT_* env tests/mutation-workflow-base-ref.test.cjs fails on PR branches while next stays green, and it is currently blocking at least three unrelated PRs (#2841, #2832, #2827) with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees The existing loop comment attributes this to `git add .` rehashing O(n^2) blobs "before the object write had landed" and works around it by staging one path per iteration. That is not the cause: sequential execFileSync calls cannot race each other's object writes, and the failure persisted after that change — it simply moved to a lower commit index. The cause is that the git() helper inherited the runner's environment. A leaked GIT_INDEX_FILE makes `git add` write into a DIFFERENT repository's index; GIT_OBJECT_DIRECTORY / GIT_ALTERNATE_OBJECT_DIRECTORIES send the blob to another object store; GIT_DIR / GIT_WORK_TREE redirect the whole operation. In every case `git commit` then cannot resolve a blob it just staged, which is precisely the error above. Verified by negative control: with GIT_DIR exported, this test fails on the unfixed helper (the git commands operate on the wrong repository entirely); with the helper stripping GIT_* it passes. The single-path staging is kept — it is genuinely less work — but it is no longer load bearing. Found while shipping #2797. Fixed in place rather than deferred: it is a defect surfaced during the work, and it is blocking other contributors. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2452): build the base-advance commits empty, removing the lost-object class The base-ref guard has been failing in CI with: error: invalid object 100644 <sha> for 'base-N.txt' error: Error building trees It is currently red on at least three unrelated PRs (#2841, #2832, #2827) while next stays green. Two theories have now been tried and neither held. #1881 blamed `git add .` rehashing O(n^2) blobs and switched to staging one path per iteration; the failure moved from commit 32 to commit 25 and carried on. The preceding commit here made the git helper hermetic against leaked GIT_* environment — that IS a real vulnerability (with GIT_DIR exported the helper operates on the wrong repository entirely, proven by negative control) but it produces a different error than CI reports, so it is not demonstrably the cause either. Neither trigger reproduces off-CI, so this stops guessing at the trigger and removes the failure CLASS instead. The loop needs base-branch DEPTH and nothing else: no assertion reads these commits' contents, and base-side files cannot appear in `origin/base...HEAD` regardless. `--allow-empty` writes no blob and no tree, so there is no object for the index to reference and lose. It is also far less work than 60 write+hash+index cycles. The guard still proves its mechanism: the test asserts that a --depth=1 base fetch FAILS and a full fetch resolves, so a broken topology would surface immediately rather than passing vacuously. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Test <test@example.com> |
||
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
16e59d0db5 |
fix(#2691): repair seven dangling references in the ADR corpus and contributor docs (#2692)
* fix(#2691): repair five dangling references in the ADR corpus and contributor docs
Found by the 2026-07-24 ADR corpus audit; each mechanism re-reproduced live
against next @
|
||
|
|
3eb1cede26 |
fix(#1880): distinguish a corrupt config from an absent one (epic #1879 Phase 1) (#2688)
* test(#1880): prove corrupt config is indistinguishable from absent Failing-first. Encodes the issue's runtime repro: a trailing comma in .planning/config.json currently yields source:builtin-defaults with degraded:false - byte-identical to the file not existing - and the user's entire configuration is silently discarded. Asserts on the typed surface (CONFIG_REASON, _warnedUnusableConfig) rather than diagnostic prose, per the ADR-1411 amendment's test-methodology clause and CONTRIBUTING.md's raw-text-matching rule. IO failure is injected by monkeypatching fs.readFileSync and restoring in t.after(), never chmod 0o000 (root bypasses mode bits). Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): distinguish a corrupt config from an absent one loadConfigResolved wrapped the read, the JSON.parse and the entire config build in one try with one catch, so ENOENT, EACCES and SyntaxError all fell through to the same defaults and the branches returned degraded:false - actively asserting health over discarded configuration. A single trailing comma in .planning/config.json silently replaced the user's whole config, reporting source:builtin-defaults degraded:false, byte-identical to having no config file at all. ConfigResolution now carries a machine-readable reason. Genuine absence keeps degraded:false / not_configured; a file that exists but cannot be used sets degraded:true with config_unparseable or config_unreadable. The same split applies to the root config and to ~/.gsd/defaults.json. Control flow is deliberately unchanged. preflight_check reports cyclomatic 141 / cognitive 196 and 93 dependents on this function, with the guidance that small edits beat one big one, so faults are CAPTURED at the existing read sites and stamped onto the returns rather than the try/catch being restructured. Also carries the ADR-1411 amendment's wiring clause: loadConfig returns .config alone to ~51 call sites and would never see the new field, so an unusable file emits a deduplicated stderr diagnostic keyed on resolved path plus errno. Without it the reason would be an unreachable field and the user whose config was discarded would still get no signal - the actual defect. Registers the config-loader seam in lint-resolution-provenance, which until now guarded only agent-skills. Caller audit: ConfigResolution.degraded has exactly one consumer outside this module, cmdAgentSkills (src/init.cts:2259), which destructures {config, source, degraded} - adding a field does not break it. Its --json IR now reports degraded:true for a corrupt config, which is the intended fix and the one observable behavior change. Closes #1880 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): degrade when any config on the path is unusable, not just the last Two defects found by isolated adversarial review of the first cut. BLOCKER: the success-path return did not consult configFault. A corrupt ROOT config whose workstream override happened to parse returned degraded:false / reason:resolved - the root's settings silently dropped, which is the exact failure this issue closes, reappearing for any project using workstreams. The stderr diagnostic fired, so the out-of-band half worked while the in-band half reported a clean resolve; a --json consumer saw health. MAJOR: reason was derived from Object.keys(parsed) - the root+workstream MERGE - so an empty workstream file inheriting a non-empty root reported resolved despite carrying no settings. Emptiness is now judged on the file actually read, snapshotted before normalizeLegacyKeys mutates it. Also: corrects the ConfigResolution JSDoc, which still described the pre-#1880 degraded contract; adds a fast-check property asserting a PRESENT file is never reported not_configured whatever its bytes (CONTRIBUTING.md parser rule); and asserts the literal enum values so the provenance lint's configured_empty/not_configured markers check real assertions rather than incidental prose in test titles. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1880): reject valid JSON that is not a config object at the read seam The fast-check property added in the previous commit failed on both node lanes: a config.json containing 0, "str", [], null or true is valid JSON, so it parsed "ok", then threw downstream in normalizeLegacyKeys, and the outer catch reported not_configured - a PRESENT file reported as absent, which is precisely the collapse this issue exists to close. The property asserts a present file is never not_configured, and it caught it. _readConfigFile now validates shape, not just parseability (ADR-227: check the semantic shape at a trust boundary, not merely the type). A non-object JSON document is an unusable config, reported config_unparseable. Adds named regression cases for each non-object form alongside the property, so the class is documented and not only randomly sampled. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1880): backfill changeset pr number (pr:0 -> 2688) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2452): record a fetch-time shallow failure instead of crashing This guard failed CI on ubuntu-24 while passing on ubuntu-22 and windows-24 for the same commit, and passed on other PRs. Not a flake and not caused by the change under test - a real fragility in the test. runnerDiff ran the base fetch OUTSIDE its try and only guarded the diff, so it assumed the failure mode is always 'fetch succeeds, diff reports no merge base'. At a shallow boundary that lands short of the merge base, git can instead fail during the FETCH ('unable to parse commit' - the boundary commit's parent is not available). Which stage git fails at is version and transport dependent, so on some runners the error escaped runnerDiff and crashed the test rather than being recorded as the ok:false the assertions expect. Both stages mean the same thing for what this guard protects: a shallow base ref cannot resolve the three-dot diff. Also drops two assert.match calls against git's stderr prose. 'no merge base' and 'unable to parse commit' are the same condition reported at different stages, and CONTRIBUTING prohibits raw text matching on subprocess output. The typed outcome (ok === false) is the contract; the tests now assert that plus the presence of a cause. Found while investigating the red lane on #2688; fixed here per the no-defer rule rather than filed. Refs #1879 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
46ba02acde |
feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs * fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions * fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests * chore(#2630): backfill changeset pr to 2661 |
||
|
|
6ad30f74b6 |
feat(#2584): Phase 3 — scheduler consumer + isolation adapters (#2635)
Final phase of #2584 (ADR-1239 Codex-binding amendment). execute-phase now negotiates dispatch.isolation and dispatches through the matching adapter, so a wave's independent plans run concurrently on six runtimes instead of one — with no runtime=== branch in the scheduler. harness-worktree passes the host's declared isolation flag (claude, cursor); orchestrator-worktree creates the worktree via the Phase-2 verb and spawns the executor into it with the resolved argv/cwd (codex, opencode, kimi, kimi-code); none stays sequential. Undeclared/unknown/unresolvable isolation degrades to none — never an unisolated parallel run. Fixes two shipped Phase-2 descriptors that per-host research found would fail at spawn: kimi lacked its headless flag (would launch the interactive TUI and hang the orchestrator), and kimi-code named a non-existent binary (Kimi Code installs as 'kimi'). Adds the worktree-path root confinement Phase 2 deferred here, and leading-dash guards on the resolver's prompt/cwd matching the existing git-argument guard. Closes #2627 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
155c08facf |
docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase (#2574)
* docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase These two commands never parse --validate (silent no-op); the flag is real only for /gsd-quick. Remove the false flag-table rows and CLI examples across COMMANDS.md and the how-to guides (en + ja-JP/zh-CN/ ko-KR/pt-BR mirrors), and correct the manager.flags.execute example from --validate to --cross-ai (a flag execute-phase actually parses). /gsd-quick's real --validate docs are left untouched. Ref #2197 * docs(#2197): add changeset for --validate docs removal --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
09b535ac00 |
feat(#2481): add a negotiated effortSurface axis and wire invocation-time effort
ADR-1239 gains a ninth negotiated axis, effortSurface (argv | none), declaring how
a host accepts reasoning effort. ADR-443 is amended in the same change because its
recorded deferral is what the axis resolves: its Unblock condition offered paths
(a) and (b) and stated the choice was 'a maintainer call this file records but does
not make'. Path (a) is selected and satisfied here.
Before this, effort reached a runtime only through install-time channels
(EFFORT_RENDERING's frontmatter/api), so reviewer CLIs spawned as subprocesses
silently inherited whatever effort sat in the user's own global CLI config. The
review lane now resolves one universal effort through the ADR-443 cascade and
renders it per host through the negotiated descriptor.
Every per-host value is documentation-sourced, never inferred:
- claude argv -- verified via 'claude --help' (--effort <level>)
- opencode argv -- verified via 'opencode run --help' (--variant)
- codex argv -- codex-rs/exec/src/cli.rs: model_reasoning_effort is NOT a CLI
flag (config.toml key only), so the global -c override is the
only argv route
- 15 hosts undocumented -- their docs state no reasoning setting; the sentinel
fails closed rather than inheriting a profile baseline
No config-file vocabulary member: the only host that ever had one (Gemini CLI's
thinkingConfig) was removed as a sunset runtime by
|
||
|
|
455ad49ae3 |
feat(#2296): config-gated provider escalation on quota-exceeded (#2458)
* test(#2296): failing-first coverage for provider escalation on quota-exceeded
Covers the provider-escalation ladder layered onto EXEC.CLASSIFY: back-compat
(no escalation block without --failure-class), cap boundaries at
min(max_escalations, list length) at limit-1/limit/limit+1, opt-in gating,
malformed/hostile provider_escalation config, the --failure-class CLI negative
matrix, config-key registration, and a fast-check budget-limit property.
Red until the resolver, CLI flag, and manifest key land.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(#2296): config-gated provider escalation on quota-exceeded
The dynamic_routing tier ladder escalates within one provider, which does not
help when that provider is what ran out of quota. Add an opt-in provider ladder
layered on the existing EXEC.CLASSIFY seam.
- model-resolver: resolveProviderEscalation walks dynamic_routing.provider_escalation
capped at min(max_escalations, list length), reporting from/to/attempted/exhausted.
Invalid entries are dropped (ADR 227 shape validation). Stays a leaf module —
the quota-class policy decision is the caller's, per the CONTEXT.md contract.
- agent-command-router: export a frozen AGENT_FAILURE_CLASSES so the new CLI
validator cannot drift from the classifier that produces the values.
- resolve-execution: --failure-class flag; emits an escalation block ONLY when
passed, so the existing JSON contract is byte-identical for every caller.
- config-schema.manifest: register dynamic_routing.provider_escalation.
- execute-phase step 7.1: auto-escalate, honor Retry-After, fail loudly naming
every model tried once the ladder is spent.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#2296): extract quota recovery to a reference fragment; regen goldens
The step 7.1a addition pushed gsd-core/workflows/execute-phase.md from 93390 to
95111 LF bytes, past the frozen ADR-857 Phase 6 ceiling (hard <93600, margin
<=93400) asserted by tests/fix-2285-claude-orchestration-wiring.test.cjs. The
base sat 10 bytes under the margin, so no inline wording would have fit.
That gate's own rationale is that optional-feature detail belongs in a fragment,
not the host loop. Moved BOTH the new provider-escalation branch and the
pre-existing manual recovery prompt into
gsd-core/references/execute-phase-quota-recovery.md, leaving step 7.1 as a
one-line pointer. execute-phase.md is now 92880 bytes — 510 SMALLER than base.
Also regenerates the fixtures that legitimately moved because three shipped
files changed (gsd-tools.cjs, config-schema.manifest.json, execute-phase.md):
golden-install-parity + install-tree for all 16 runtimes, INVENTORY.md +
INVENTORY-MANIFEST.json for the new reference, and the workflow size baseline.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#2351): make the C1 orphan-reaping test load-independent
tests/run-with-timeout.test.cjs C1 asserted the child heartbeat file exists
after a 1s group-kill window, but the child only wrote it on the first 100ms
setInterval tick. Nothing synchronized the two: on a loaded container the group
is SIGKILLed before that tick lands, the file never appears, and the assertion
fails for a reason unrelated to reaping. Observed failing on both linux-node22
and linux-node24.
The behavior actually under test is the FREEZE assertion (heartbeat stops
advancing => descendant was reaped, not orphaned). That is unaffected by
sampling once more at t=0.
Child now writes its first heartbeat synchronously at startup before arming the
interval, and the kill window widens 1s -> 3s to cover child boot under load.
Both remove the timing dependency; neither weakens what the test proves.
Refs #2296
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(#2296): backfill pr:2458 in .changeset/rapid-jays-bark.md
* chore(#2296): regenerate fixtures after rebase onto #2402
The rebase conflicted on the generated golden-install-parity fixtures and
workflow-size-baseline.json because #2402 (
|
||
|
|
12e4d93b19 |
fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts (#2445)
* fix(#2393): add GSD_ALLOW_SYMLINKED_DEST opt-in for intentional user-owned symlink layouts Bug: v1.7.0's destSubpath write-confinement (ADR-1239 Phase B) refused install/update whenever CLAUDE_CONFIG_DIR (or an artifact-kind child like skills/, hooks/) was a pre-existing symlink, with no opt-out. Three legitimate user-owned layouts were blocked: - (lars-hh) CLAUDE_CONFIG_DIR=~/.claude-personal with skills/hooks symlinked to a user-owned external dir - (Mamiki) ~/.claude/skills is a Windows Junction to a shared skills dir - (Azd325) ~/.claude itself is a symlink to a dotfiles repo (the early root-is-symlink return refused before the component loop ran) Fix: add GSD_ALLOW_SYMLINKED_DEST env var (accepts '1' or 'true'). When set, hasExistingSymlinkBetween follows symlinks instead of refusing them. Cross-platform: fs.lstatSync().isSymbolicLink() returns true for both POSIX symlinks and NTFS junctions (Node ≥ 16), so Mamiki's Junction case is handled by the same code path. Threat model preserved (these still refuse EVEN WITH opt-in): (a) path-traversal in the destSubpath string itself ('../../etc'-style) — ADR-1239 Phase B threat (a), untrusted destSubpath protection (b) a symlink whose resolved real path equals the install root itself — would let _removeGsdEntries wipe the root; #1704 threat (b) (c) broken symlinks (realpathSync throws) — fail-closed What opt-in RELAXES specifically: the 'pre-existing symlink pointing outside configHome' refusal — #1704 threat (c). The user has explicitly asserted they own and trust the symlink target. Error messages at all 4 call sites (installRuntimeArtifacts, _copyStaged, migrateLegacyDevPreferencesToSkill, installOpencodeFamilySkills) updated to (1) name the env var opt-in, (2) be accurate when the root itself is a symlink (Azd325's complaint that the old message accused destDir of 'containing' a symlink when the root was the actual symlink). Docs: docs/CONFIGURATION.md Environment Variables table updated. Regression tests in tests/install-write-confinement.test.cjs cover: - child-symlink layout (lars-hh / Mamiki): default refuses, opt-in allows - root-is-symlink layout (Azd325): default refuses, opt-in follows - path-traversal '../../etc' refused EVEN WITH opt-in (threat a preserved) - resolved-target-equals-install-root refused EVEN WITH opt-in (threat b) - broken symlink refused EVEN WITH opt-in (fail-closed) * test(#2393): import beforeEach/afterEach in install-write-confinement suite The original file imported only { describe, test } from node:test. The new #2393 opt-in describe block uses beforeEach/afterEach to manage the GSD_ALLOW_SYMLINKED_DEST env var lifecycle — add them to the import. * test(#2393): correct broken-symlink test — existsSync follows link → loop terminates early Initial test expected broken symlinks to be refused even with opt-in. That was wrong: fs.existsSync follows symlinks, so a broken symlink returns false from existsSync and the component loop terminates before the symlink check fires. Both default and opt-in paths share this behavior; the fix preserves it. Updates the test to pin the actual current behavior so a future refactor (e.g. switching to lstatSync for existence) is a deliberate behavior change. * fix(#2393): realpath the install root — guard against macOS /var ↔ /private/var Code review (security subagent) flagged a HIGH-severity hole in the threat-(b) preservation: realTarget (from fs.realpathSync) is fully symlink-resolved, resolvedRoot (from path.resolve) is lexical-only. On macOS /var is a symlink to /private/var, so resolvedRoot='/var/foo/.claude' but realConfigHome is '/private/var/foo/.claude'. A symlink whose realtarget matches the install root by real path would compare unequal to the lexical resolvedRoot — defeating the wipe-protection guard exactly in the reporter's case (Azd325, nix-darwin: ~/.claude is itself a symlink). Fix: compute realRoot once via fs.realpathSync(resolvedRoot) at function entry (with fail-closed fallback to lexical form on realpath failure — broken/missing root, permission denied, exotic FS). Threat (a) path-traversal check above still confines regardless. Compare against BOTH lexical and real forms in both the root-symlink and component-symlink branches. Also adds the reviewer's transitivity-trust clarification comment: once a symlink is followed under opt-in, the walk continues from the resolved real path WITHOUT re-checking further segments stay inside a confining boundary. This is documented opt-in semantics — one opt-in trusts the whole reachable tree — and the comment makes the design choice explicit so a future maintainer doesn't add a 'follow one symlink only' expectation. Regression test added for the macOS /var normalization case (spelled configHome via os.tmpdir() lexically while pointing the test symlink through its realpath). Test skips on non-darwin platforms and when os.tmpdir() has no symlink component. * fix(#2393): root-symlink branch — do not apply threat-(b) check to root itself Initial fix applied the wipe-threat-(b) check to the root-symlink branch unconditionally. That was wrong: when root itself is a symlink (Azd325's nix-darwin case), its realpath IS realRoot by construction — so the check always fires, defeating the opt-in for exactly the case it was meant to enable. The wipe threat (b) does NOT apply to root being a symlink: destDir is a CHILD of root, and resolving root just gives root's target. There is no circular back-reference to root from a path that descends from a resolved root. So the root-symlink branch should just follow the symlink under opt-in and continue the walk, no threat-(b) check. Threat (b) only fires in the COMPONENT loop, where a child symlink can resolve back to the install root. That branch keeps the (b) check using BOTH lexical and real forms of root (the macOS /var ↔ /private/var fix from the prior commit). * fix(#2393): apply opt-in at the 5 bin/install.js call sites + add env-var/transitive tests Code review (correctness subagent) flagged a Critical coverage gap: the initial fix updated only the 4 src/install-engine.cts call sites. Five more call sites in bin/install.js still used the 2-arg signature, so the opt-in env var was silently ignored on: - installCodexConfig (config.toml + agents/ dir + per-agent .toml paths) — Codex only - copyWithPathReplacement (the generic emit path: workflows, commands, staging) — ALL runtimes - resolveInstallRelativePath (path resolver used in various places) Result: a user setting GSD_ALLOW_SYMLINKED_DEST=1 would see SOME refusals disappear (engine path) and OTHERS remain (bin/install.js paths) — a partially-applied install and a confusing UX, directly contradicting the PR's headline claim. Fix: - Export isSymlinkedDestOptIn from src/install-engine.cts alongside hasExistingSymlinkBetween - Import it in bin/install.js - Update all 5 bin/install.js call sites to pass { allowOptInFollow } - Update all 3 bin/install.js error messages to name the env var, matching the engine's phrasing Also addresses reviewer's Medium test-adequacy findings: - isSymlinkedDestOptIn env-var parsing now tested directly (accepts only documented '1' / 'true'; rejects 'TRUE', 'yes', 'on', '0', 'false', empty, unset) - transitive symlink chain (configHome/outer → outside1 → outside2) test pins the documented 'transitive and unbounded' opt-in semantics so a future contributor can't accidentally narrow it * chore(changeset): backfill pr:2445 in .changeset/eager-wasps-swim.md |
||
|
|
517bae8d6d |
fix(#2372): widen decision-coverage-plan to all planner-canonical tags, drop misleading "(or body)" (#2443)
* fix(#2372): widen decision-coverage scan to planner-canonical tags, fix message Bug: check.decision-coverage-plan's remediation message told the user to cite decisions "(or body)" but extractPlanDesignatedSections only scanned <objective>/<tasks>/<task>/<action>. A decision cited in <read_first>, <behavior>, <verify>, <acceptance_criteria>, or <done> was invisible to the gate — false BLOCKING coverage gap, plus the message's own fix-hint sent the user to "the body" where re-citing still failed. Two-part fix (must change together — that drift was the bug): 1. Widen XML_DECISION_TAGS_RE in src/check-command-router.cts to also match <read_first>, <behavior>, <verify>, <acceptance_criteria>, <done>. These are all planner-canonical tags the planner is told to use (plan-phase.md:830-862, plan-phase.md:772). The body negative- lookahead mirrors the opening-tag set so each tag's body is captured independently. 2. Correct buildPlanMessage to name ONLY the surfaces the extractor actually scans (front-matter must_haves/truths/objective, designated markdown headings, and the nine planner-canonical tag bodies). The misleading "(or body)" clause is gone. Also updates the planner's documented contract (agents/gsd-planner.md:69) and user-facing docs (docs/CONFIGURATION.md, docs/USER-GUIDE.md) to reflect the wider scan. Regression tests in tests/decisions.test.cjs cover each newly-scanned tag body, a control (no citation still uncovered), and a message/extractor parity assertion that names every scanned surface — so the two cannot drift apart again. Out of scope (per triage): cmdDecisionCoverageVerify/buildVerifyMessage is a separate command (decision-coverage-verify) checking shipped artifacts, not plan citations — untouched. * chore(#2372): regenerate agent-size-baseline + golden-install-parity fixtures gsd-planner.md grew 49172 → 49294 (+122 chars) from the widened decision- coverage contract (5 new scanned tag names + heading clarification). Growth is justified: the contract surface is itself the fix — the prior text under-described what the gate scans, which was the bug. Updates: - tests/agent-size-baseline.json (gsd-planner.md: 49172 → 49294) - 17 tests/fixtures/golden-install-parity/*.json (one hash per runtime) - tests/fixtures/install-tree/*.json (regenerated by gen:golden) * fix(#2372): per-tag matching — outer-tag citations survive inner-tag nesting Code review (subagent) flagged a Medium edge-case regression from the single-alternation regex: when a newly-scanned tag nests inside another scanned tag, the alternation's negative lookahead halts the outer tag's body at the inner tag — losing any D-NN citation in the outer tag's prefix prose. Concretely: <action>per D-05 <verify>npm test</verify></action> → 3-tag alternation (old): captured 'per D-05 <verify>npm test</verify>' as <action> body → D-05 caught → 9-tag alternation (bug): captured 'npm test' only (from <verify>); D-05 in <action> prefix LOST Switches extractXmlTagBodies to per-tag matching: each tag gets its own regex whose negative-lookahead tempers only against the SAME tag's reopening. So <verify> inside <action> is absorbed into <action>'s body (D-05 caught) AND <verify> is matched separately on its own pass. Per-tag preserves both: - the reporter's case (sibling tags inside <read_first>) - nested-tag citations in outer-tag prefix prose - ReDoS safety (each per-tag regex keeps the #2128 body tempering) Also adds the reviewer's other requested edge-case tests: - non-scanned tag (<name>) bearing D-NN must NOT count - self-closing form <read_first /> safely ignored - attribute form <verify type="...">D-NN</verify> (canonical planner shape) - CRLF newlines in tag body do not break capture * chore(changeset): backfill pr:2443 in .changeset/noble-elks-chatter.md |
||
|
|
20ff405cb3 |
feat(#2162): opt-in compact GSD-state format for the statusline (#2175)
* feat(#2162): opt-in compact GSD-state format for the statusline New statusline.state_format config, enum full|compact (default full — existing rendering untouched). "compact" renders the state segment as "<version> · P<phase>/<total> · <status>", e.g. "v1.12 · P7/12 · executing" — dropping the milestone name and progress bar (the two biggest width costs) and collapsing narrative statuses to a single keyword. Per the #2162 approval conditions, the keyword set is the canonical vocabulary from normalizeStateStatus() in state-document.cjs (discussing/planning/executing/verifying/completed/paused) — no parallel hand-rolled list, so the vocabularies can't drift — and the canonical stuck state "paused" renders uppercase as PAUSED (no new "blocked" lifecycle state). Statuses the normalizer passes through unrecognized fall back to their first word capped at 16 chars. Lifecycle scenes preserved: active_phase wins over the body phase number, milestone completion renders "complete", idle-with-next-action renders "next <action> <phases>". Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2162): changeset fragment for PR #2175 * fix(#2162): review fixes — ENUM_KEYS coverage, cap boundary tests, changeset format - register statusline.state_format in the fix-1628 coercion-bypass matrix - 15/16/17-char boundary tests for the shortGsdStatus fallback cap - changeset body ends with the (#2162) citation per house convention Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): round-2 review fixes — scene exclusivity, direct config-set coverage - compact renderer gates the milestone-complete scene behind the absence of an in-flight phase id, mirroring formatGsdState's if/else precedence (Scene 1 beats Scene 3); regression test covers the non-atomic active_phase + percent=100 STATE.md shape - direct config-set accept/reject test for statusline.state_format plain strings (ENUM_KEYS matrix covers only the JSON coercion shapes) Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the statusline hook change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): complete-scene gate matches formatGsdState exactly (+property tests) Re-review Major: gating done on !phaseId held completion back for the legacy phaseNum shape — formatGsdState reaches Scene 3 on percent=100 regardless of phaseNum, so compact must too. Gate is now !s.activePhase. The phaseNum-only test now expects 'complete' and cross-checks the full renderer; a parity test feeds identical inputs to both renderers. Re-review Minor: shortGsdStatus gets fast-check property coverage (totality, canonical fixed points, separator safety, fallback shape). Golden fixtures regenerated for the hook byte change. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg |
||
|
|
d6672ff926 |
feat(#2163): opt-in git branch/status segment in the statusline
New statusline.show_git config (default false). When enabled, a git segment renders after the directory: current branch plus compact work-state markers (+staged ~unstaged ?untracked ↑ahead ↓behind, or ✓ when clean and in sync), e.g. " │ main+2~1?3". One git status --porcelain=v2 --branch spawn per render via execFileSync with a fixed argument array (no shell), a 1.5s timeout, and the workspace dir passed with -C. Fails silently — segment absent outside a repo, without git, or on timeout. Default output is unchanged when the flag is absent. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg |
||
|
|
ad7111e50b |
feat(#2161): opt-in absolute token count on the statusline context meter (#2174)
* feat(#2161): opt-in absolute token count on the statusline context meter New statusline.show_context_tokens config (default false). When enabled, the context meter shows the absolute token total after the percentage, e.g. "████░░░░░░ 46% (156k)" — summing input, cache-creation, cache-read, and output tokens from context_window.current_usage (matching /context). Default output is byte-for-byte unchanged when the flag is absent or false. The .planning config is now read once per render and shared with the last-command/position block instead of being re-read. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2161): changeset fragment for PR #2174 * fix(#2161): review fixes — k-to-M threshold, boundary tests, changeset format - formatTokens promotes to the M branch when k-rounding reaches 1000 (999,500-999,999 rendered "1000k" instead of "1.0M") - boundary tests at 999499/999500/999999/1000000/1000001 - Number() guards on the four usage fields (silent string-concat gap) - changeset body ends with the (#2161) citation per house convention Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2161): round-2 review fixes — config-set coverage, precision claim, exports style - config-set accept/reject tests for statusline.show_context_tokens (mirrors the post-planning-gaps precedent the issue scope names) - changeset + docs no longer claim parity with /context: the suffix sums four fields while the meter %% derives from used_percentage (three), so the figures can diverge slightly - module.exports one entry per line Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the statusline hook change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
79d7657eff |
feat(#2102): make pi a first-class installable runtime + fix its dispatch (ADR-1239)
Net-new EoS/pi installable runtime — purely additive (no prior runtime==='pi'
branches). pi is a bun-runtime programmatic-CLI whose /gsd command is registered
by a native ExtensionAPI extension and dispatches through the embedded engine.
Stage 1 (install plumbing):
- capabilities/pi/capability.json: full hostIntegration descriptor (imperative /
slash-programmatic / active-model / native-extension / bun) + hostBehaviors
{nativePlugin, pluginOnlyInstall}.
- --pi flag + interactive-menu renumber (All 17->18); pi added to RUNTIME_FLAG_IDS,
RUNTIME_LABELS, RUNTIME_META, allRuntimes/runtimeMap, model-catalog defaults.
- Install mirrors OpenCode: pi installs the gsd.cjs extension + the shared engine
payload (gsd-core + scripts + config markers) + the shared hooks bundle (spawned
by the extension at lifecycle events, like OpenCode's plugin). pluginOnlyInstall
EXCLUDES declarative command/agent/skill markdown, which pi has no host-read
surface for (its /gsd is programmatic). _installNativePluginIfDeclared (extracted
from the opencode-family path) copies pi/gsd.cjs -> ~/.pi/agent/extensions/gsd.cjs
(global) / .pi/extensions/ (local). pi added to package.json files.
- Golden: new pi.json (320 files: extension + engine + 27-file hooks bundle, no
markdown); the 16 other fixtures + claude-local change only by the shared
model-catalog hash line.
Stage 2 (real dispatch + upgrades):
- Shared dispatchGsdCommand() (shell-command-projection): bounded, no-throw
subprocess-shim to gsd-tools.cjs (the only full-surface dispatch path; no
in-process full-hub factory exists). Fixes pi/gsd.cjs's createHub()-no-args bug
(every dispatch was UnknownCommand) AND the identical bug in mcp-server.cts's
gsd_invoke_command, which a vacuous unknown-family-only test had masked (now has
a real dispatch regression test).
- pi/gsd.cjs: /gsd handler now (args, ctx) - tokenizes (quote-aware, via the
shipped hooks/lib/git-cmd.js) + dispatches real family/subcommand (not hardcoded
query/help); gsd_invoke gets a TypeBox (JSON-schema-fallback) parameters schema +
consumes params; getArgumentCompletions; before_provider_request active-model
steering (fail-open on null resolution); functional session_start /
before_agent_start / session_before_compact hook bridges (spawn the shipped GSD
hook scripts).
- EXTENSION_EVENT_SURFACES.pi expanded from ['tool_call'] to the full 30-event
vocabulary.
Docs (host-integration matrix + how-to) + changeset (Added).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
f014ec83bd |
feat(#2093): migrate Kilo onto EoS imperative adapter + hook-bus/model/MCP/dispatch upgrades (ADR-1239)
Fold remaining isKilo logic branches into descriptor-driven reads: finishPermissionWriter (uninstall cleanup), skipSharedHooksInstall (hooks copy), and a skills converter-name registry (the artifactLayout.converter field is now load-bearing, not decorative). frontmatterDialect stays the documented dispatch key for frontmatter (no descriptor field for it). Dead isKilo destructure bindings removed. Byte-identical golden parity for all 16 runtimes (opencode, which shares kilo's combined-family path, verified clean). UPGRADE 1 (hook bus): install .kilo/plugins/gsd-core.js native plugin + extensionEvents:"kilo" + EXTENSION_EVENT_SURFACES.kilo (OpenCode-fork bus). UPGRADE 2 (active model): populate runtimeTierDefaults.kilo + thread modelOverride through convertClaudeToKiloFrontmatter — model no longer stripped from agents. UPGRADE 3 (MCP): document the gsd-core MCP companion under kilo's mcp config key. UPGRADE 4 (named dispatch): agents/*.md mode:subagent roster is the Task-tool dispatch surface (tested); subagentToolkit stays 'undocumented' per AC so dispatch degrades to 'degraded' by design. Model-catalog single-source edit ripples the shared model-catalog.json hash into all 16 golden fixtures (expected). Inline defect fixes (no-defer): stale- bake-guard resolveAgentDir 'agent'->'agents' (was a silent no-op for opencode/ codex), hardcoded 'Removed OpenCode plugin' uninstall log -> generic, and the connect-gsd-mcp-server.md OpenCode mcpServers->mcp doc error. Tests: kilo-imperative-reference (adapter/axes/fail-closed/degradation/ hostBehaviors + widened isKilo source-grep across 4 modules) + kilo-upgrades (plugin parity+load, model-override converter, agents dispatch surface, MCP doc). Matrix + how-to + config docs updated; changeset added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
185abe2d66 |
feat(codex): advance Codex/OpenAI model defaults to GPT-5.6 (Sol/Terra/Luna)
Update runtimeTierDefaults.codex and providerPresets.openai in model-catalog.json to the GPT-5.6 family (gpt-5.6-sol/terra/luna), advancing from the superseded GPT-5.4/5.5 generation. Model IDs verified against OpenAI developer API docs: - gpt-5.6-sol: flagship, /, reasoning xhigh - gpt-5.6-terra: balanced, .50/, reasoning medium - gpt-5.6-luna: fast/cheap, /, reasoning medium Tier mapping is 1:1 (Sol↔flagship, Terra↔balanced, Luna↔fast), so profile semantics are unchanged — only the underlying IDs advance. Updates: catalog JSON, test assertions (catalog defaults), docs (CONFIGURATION.md + zh-CN/pt-BR translations, workflow settings), and changeset. Closes #2122 |
||
|
|
5aa7771dd3 | Merge branch 'next' into fix/2009-load-failed-capability-injects-blocking- | ||
|
|
1ee00320c7 | Merge branch 'next' into fix/2072-thread-model-into-assumptions-and-review-spawns | ||
|
|
4483300253 |
fix(#2072): thread resolved model into routed-agent spawns (assumptions-analyzer, code-reviewer, code-fixer)
model_overrides / models.<phaseType> were silently inert for gsd-assumptions-analyzer,
gsd-code-reviewer, and gsd-code-fixer on Claude Code: resolveModelInternal honors them,
but the workflows spawned these agents with no model= param, so the resolved value
never reached the Agent tool and the agents inherited the session model — no warning.
Fix — thread each agent's resolved model at every spawn site (the established
plan-phase pattern; the architecture-consistent Claude mechanism, since 13 other
agents already thread their model):
- discuss-phase-assumptions.md: `resolve-model gsd-assumptions-analyzer --raw`
→ ANALYZER_MODEL, threaded.
- code-review.md + code-review-fix.md (re-review): `resolve-model gsd-code-reviewer --raw`
→ REVIEWER_MODEL, threaded.
- code-review-fix.md (both fixer spawns): `resolve-model gsd-code-fixer --raw`
→ FIXER_MODEL, threaded (same silently-inert bug, same file — folded in per review).
- quick.md review step: was reusing `{executor_model}` for gsd-code-reviewer (so the
reviewer's own override was ignored); init.quick now resolves `reviewer_model`
(gsd-code-reviewer) and the spawn threads it.
resolve-model --raw returns the bare model string (resolve-execution --raw would
return effort — wrong). The resolver maps these agents to phaseType discuss /
verification / execution, so models.<phaseType> apply too.
Scope: the three agents reachable from the two issue-named workflows + quick.md. The
wider systemic class (other agents in UNTOUCHED workflows with the same pattern) stays
documented on the issue for a maintainer-scoped structural decision (thread-at-source
vs embed-at-install like #2256), not widened here.
Docs: the stale "discuss — reserved, no subagent today" model-profile tables now list
gsd-assumptions-analyzer and the verification row includes gsd-code-reviewer, across
the English docs, the shipped gsd-core/references/model-profiles.md reference, and the
ja-JP / zh-CN / ko-KR / pt-BR locale mirrors.
Tests:
- tests/model-resolver.test.cjs: #2072 acceptance — model_overrides and
models.discuss/verification/execution resolve for all three agents.
- tests/model-routing-spawn-threading.test.cjs: every spawn of the three agents threads
a resolved model (fails pre-fix); a header-precise parity guard fails the suite if a
new un-threaded spawn of any of them regresses.
All 16 golden-install-parity fixtures + the workflow size baseline regenerated for the
changed shipped files (4 workflows + the reference doc); bin/lib is excluded from parity.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
46f8d3814c |
fix(#2009): load-failed capability gates fail open with a loud warning
Previously a capability that failed to LOAD (e.g. incompatible engines.gsd) but declared a gate-kind loop hook caused the loop resolver to inject a BLOCKING synthetic gate (blocking:true, onError:halt) at every declared point, halting every ship:pre / verify:post project-wide over an unrelated load error, with no remediation surfaced. Per maintainer decision (#2009) it now fails OPEN: no gate is injected (the loop proceeds; --active-cap correctly reports the failed cap inactive) and a loud warning is emitted — to stderr (the channel host workflows/agents actually see) and in the envelope 'warnings' array — naming the load reason and the exact 'gsd capability remove <id>' remediation. The loader still records blockedGates; only the consequence changes from block to warn. Security (review): capId and reason originate from a third-party manifest / directory name. capId is validated against the canonical kebab-case id shape before it is placed in the runnable remediation command (withheld otherwise); reason is stripped of control chars and backticks. This closes an argument/ prompt-injection vector in the surfaced message. Docs updated to the fail-open-warning posture (ARCHITECTURE, INVENTORY, CONFIGURATION, README, capability-overlay-model). Also removes a dead 'before' import surfaced by lint in the issue-2045 test. |
||
|
|
6addeccd19 |
feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates an external API/SDK/service can no longer seal without a decided coverage matrix. - src/api-coverage.cts: deterministic detector (compound verb+noun signal + <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix parse/validate/render with field-length caps. - check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as a token under .planning/phases/ only (traversal-neutralized); validates COVERAGE.md or blocks iff a strong integration signal is detected and no matrix exists; fail-closed when phases tree exists but phase unresolvable. - capabilities/ai-integration: workflow.api_coverage_gate config key (default true), plan:pre contribution, blocking verify:pre gate. Data-driven. - gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch. - Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e. Code+security review findings fixed (stopword FP, scope containment, pipe/cap rejection, prompt-injection message hygiene). - Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs. Closes #1562 |