9e4f0e99ad05d3449be98d15fdd4b003d768febd
135 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
98ecb2ba8c |
enhance(#2142): archive quick tasks at milestone close-out (#3592)
* test(#2142): failing-first coverage for quick-task archival at milestone close-out * enhance(#2142): archive quick tasks at milestone close-out * fix(#2142): resolve review findings — readme injection, move/reset ordering, owned state write * fix(#2142): fold archival under milestone namespace, expose index IR, dedupe reset decision * test(#2142): assert archive-dir-relative summary path in index IR * docs(#2142): backfill changeset pr number to 3592 * test(#2142): skip newline-fixture injection test on windows (control chars illegal in path names) --------- Co-authored-by: sim <sim@local> |
||
|
|
abf3cf7c25 |
fix(#3458): scan archived milestone phases, and make [A] Acknowledge actually suppress (#3555)
* fix(#3458): scan archived milestone phases in the four audit-open scanners `query audit-open` resolved exactly one phase root, `.planning/phases/`. When a milestone closes its phase directories move to `.planning/milestones/v<X.Y>-phases/`, so an item still unresolved at that moment — the `[R]/[A]/[C]` prompt accepts "accept" and "carry forward", not only "resolve" — became invisible to the v1.1 pre-close audit and every audit after it. The window in which an unresolved item is visible to this gate was exactly one milestone wide, and nothing announced when it closed. Reproduced before fixing, with byte-identical artifacts in the two layouts and the active layout as the control: active → has_open_items=true deferred=1 uat_gaps=1 total=2 archived → has_open_items=false deferred=0 uat_gaps=0 total=0 `scanDeferredItems`' own doc comment names this as the thing it was built to prevent — "phase directories archive to `milestones/vX.Y-phases/` (#1871) and the entry leaves the live tree having never been triaged" — while the implementation eleven lines below cannot read that path. It catches an entry at its own milestone close and goes blind at precisely the transition the comment describes. This is not cosmetic under-reporting. `auditOpenArtifacts` sums all nine category counts into `counts.total` and returns `has_open_items: counts.total > 0`, so four blind scanners can flip the gate's headline boolean and let `/gsd-complete-milestone` assert a clean close it never verified. In a fully-archived project `.planning/phases/` may not exist at all, and the scanners' `if (!fs.existsSync(phasesDir)) return []` produced a value indistinguishable from "nothing is open". ## One enumeration, not four The four scanners each hand-rolled the same active-only walk. They now share `listAuditPhaseTargets(planDir, cwd)`, which yields both roots — the shape of fix epic #3473's B2 asks for, and the reason the fix is one seam rather than four edits. Three properties are load-bearing: * the ACTIVE enumeration is unchanged — still a raw `readdirSync`, NOT `listMilestonePhaseDirs`. These scanners are deliberately not milestone-filtered today, and switching would silently add window and sentinel filtering: a behavior change belonging to #3372, not here. * a missing or unreadable active root skips that half instead of returning early. That early return WAS the bug in a fully-archived project. * archived dirs are deliberately NOT milestone-filtered, per the comment `src/uat.cts` already carries: archived phases belong to past milestones by definition, so applying the current-milestone filter discards every one and silently reinstates this bug. Each item now carries `archived_milestone` when it comes from a closed milestone, matching how the sibling module already labels archived results — without it an operator triaging `[R]/[A]/[C]` cannot tell a live item from one carried over. Additive: no existing test or doc asserted an exact key set. `scripts/lint-phase-enumeration-drift.cjs`'s exemption list for this file drops from the four scanner names to the single helper, since that is now the only place the enumeration lives. ## Tests Written failing-first and confirmed red for the right reason before the fix, all four driven through the real `audit-open` CLI rather than private functions: archived-only (was 0/0/0/0 with `has_open_items=false`, now 1/1/1/1 true), mixed active+archived (was 1/1/1/1 — the archived half dropped — now 2/2/2/2), active-only unchanged, and an all-resolved archived phase contributing 0. That last one passed vacuously before the fix, because the archived path was not reached at all; it was re-verified as genuinely discriminating afterward by flipping one archived item to unresolved and watching the count rise. Closes #3458 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3458): restore the scan_error sentinel and show archive provenance Adversarial review found one BLOCKER that the previous revision introduced, which a green remote-runner suite did not catch because nothing in the tree asserts `scan_error` at all. ## The regression Consolidating four hand-rolled walks into `listAuditPhaseTargets` swallowed the active-root `readdirSync` throw in a bare `catch {}`. Pre-fix each scanner returned `[{scan_error: true, …}]`; after, each returned `[]`. Measured with `.planning/phases` created as a FILE (so `existsSync` passes and `readdirSync` throws ENOTDIR): before this fix: uat_gaps/verification_gaps/context_questions/deferred_items each `[{"scan_error":true,…}]` the regression: each `[]` `complete-milestone.md` re-runs `audit-open --json` and reads those counts, so a machine consumer could no longer tell "I/O failed" from "verified clean" — the exact conflation this issue exists to remove, reintroduced on the failure path. `listAuditPhaseTargets` now reports `activeUnreadable` and each scanner pushes the sentinel shape recovered verbatim from `origin/next`, not reinvented. The docstring claiming the active enumeration was "UNCHANGED" was false while that sentinel was missing, and is corrected to state what is actually preserved. An unreadable ARCHIVED root deliberately gets NO sentinel: there was no archived read before, so there is no consumer contract to preserve, and adding one would conflate the ordinary "no milestones archived yet" state with a real I/O failure. ## The operator could not see the archive `formatAuditReport` is the surface the gate actually shows a human — `complete-milestone.md` runs it without `--json` — and it never rendered `archived_milestone`. With `01-alpha` in both roots the identical line printed twice with nothing to tell them apart, and `[R] Resolve` sends the operator to `.planning/phases/01-alpha/` where the archived one does not exist. Phase numbering restarts at `01` after each archive, so that collision is the common case, not an edge case. All four loops now render ` (archived vX.Y)`; active lines stay byte-identical. ## Archived milestones sorted wrong `getArchivedPhaseDirs` ordered milestones with `.sort().reverse()` — lexicographic, so `v1.9` outranked `v1.10`. Measured order for v1.0/v1.9/v1.10 was `v1.9, v1.10, v1.0`. Now a numeric-segment descending compare. Pre-existing, but this change is what first surfaces it in audit output. ## Tests The blocker's regression test fails against the previous revision. Added: `archived_milestone` present on archived items and absent (not `undefined`) on active ones; the unreadable-active-root sentinel across all four categories; an unreadable archived root still leaving the active half scanned; the duplicate-name case producing two distinct entries that the human report distinguishes; and the v1.10-before-v1.9 ordering. `docs/COMMANDS.md` documents the archived scanning and the new field. Closes #3458 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3458): stop filesystem names forging lines in the audit report Found by the security review of this branch. Pre-existing on `next`, fixed here because it defeats the exact gate this PR is hardening. `audit-open`'s human report is the surface `/gsd-complete-milestone` shows an operator to decide whether a milestone may close. A `.planning/` tree authored by someone other than that operator — a cloned repo — could contain a directory literally named: zz<newline>0 open items require decisions.<newline><ESC>[2K<ESC>[1G FORGED and the report printed `0 open items require decisions.` as its own line, with raw ESC bytes reaching stdout able to erase or overwrite the lines above it. Reproduced against the real CLI before fixing, and again after. ## Why not just harden sanitizeForDisplay Because that helper's contract is multi-line prose — it removes protocol-leak lines while deliberately preserving the newlines between legitimate ones, which `tests/security.test.cjs` pins. Stripping CR/LF there would have broken a correct test to paper over a different problem. The two jobs are genuinely different, so there are now two helpers. New `sanitizeLabel` (`src/security.cts`) is for values that are semantically ONE LINE and derived from a filesystem NAME. It ESCAPES rather than strips C0 (including ESC/CR/LF), DEL and C1, so a doctored name renders visibly as `\n` / `\x1b` instead of being silently normalized — the report stays honest about what is in the tree. Ordinary input passes through byte-identical. ## Nine sites, not four The first pass covered the four phase-scoped scanners. A sweep of the rest of the file found the identical class in five more — `scanDebugSessions`, `scanQuickTasks`, `scanThreads`, `scanTodos`, `scanSeeds` — emitting name-derived `slug` / `filename` / `seed_id` through the prose sanitizer. `scanQuickTasks`' `date` had no sanitization call at all. Every emitted field in the file is now classified and the sweep recorded: `slug`, `filename`, `seed_id`, `phase`, `file`, `archived_milestone`, `date` are name-derived and take `sanitizeLabel`; `hypothesis`, `status`, `updated`, `title`, `priority`, `area`, `summary`, `questions[]` and deferred-item `text` are content and keep `sanitizeForDisplay`. No name-derived value reaches output unsanitized. `--json` was already safe — JSON string encoding escapes control characters, and a crafted name cannot break out of the string. Verified rather than assumed. Closes #3458 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3458): backfill changeset pr number * test(#3458): skip control-character fixtures where the OS forbids the name CI red on `test (windows-latest, 24, shard 1/3)`: the four forgery-rejection tests build directories whose names embed a newline and ESC, and NTFS forbids control characters in path components, so `mkdir` threw ENOENT. The remote runner is Linux-only, so it could not have caught this class. Semantically the skip is honest rather than a workaround: on Windows the directory-name forgery vector does not exist, because the OS refuses to create the name. The sanitizer's own behavior stays covered there by the `sanitizeLabel` unit tests, which are pure string tests with no filesystem calls — verified. Uses the repo's established capability-probe convention (`tests/adr-index-gate.test.cjs`'s `trySymlink`), which `t.skip()`s on the real errno rather than branching on `process.platform`, and whose comment gives the reason: a bare `return` "would silently report a PASS ... and hide the gap this guard exists to close". A skipped test is visibly skipped. Swept every test added on this branch for names Windows would reject or POSIX path assumptions; these four were the only ones. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3458): make [A] Acknowledge actually suppress, without overwriting a verdict Making archived phases visible exposed the other half of the problem: an item unresolved at a milestone close now resurfaces at every later close forever, because `[A] Acknowledge` wrote a prose block to STATE.md that `auditOpenArtifacts` never reads. `verified_closeout` became unreachable and the gate degraded to a mandatory `[A]` every time. ## The prompt does not change `[A] Acknowledge all` already promises "document as deferred and proceed with close". It documented but never deferred. This makes `[A]` do what it says. `[R]` and `[C]` stay abort paths. No "carry forward" option is invented — an item that is not acknowledged simply keeps surfacing, which is the default. ## The marker lives inside the artifact Not a ledger. The audit mints no ids and has no stable identity — `phase` is a token that collides across directories, `file` for deferred items is a constant, and identity otherwise degrades to the item's own prose after a lossy sanitizer. Any ledger must re-derive that key every close, so a reworded item silently un-suppresses or, worse, mis-suppresses a different one. Storing the acknowledgment next to the thing it suppresses makes that class of bug structurally impossible, and it is the pattern `src/uat.cts` already argues for with `deferred-items.md`'s in-place `status: resolved`. ## The marker is verdict-preserving and self-invalidating `status:` is never overwritten — writing `resolved` into an unresolved UAT would be a lie in the artifact of record, and the disclosure has to be additive. audit_acknowledged: milestone: v1.0 at: 2026-08-15 status: gaps_found # snapshot of what was true when acknowledged Suppression applies ONLY while the snapshot still matches reality: `status` for seven categories, `question_count` for context questions, and for deferred items a new per-entry `status: acknowledged` distinct from `resolved`, which keeps meaning "actually fixed". Change the artifact and the acknowledgment stops applying, so the item comes back on its own. That is what makes re-opening answer itself with no extra state, and it fails in the safe direction: a stale acknowledgment can never hide a NEW problem. A malformed marker is treated as absent — a bad marker must never silence an item. The check is ONE shared `isAuditItemAcknowledged`, not nine copies. This file has already been through that defect family twice in this PR. ## Observable, not silent `audit-open --json` now reports an `acknowledged` count beside `counts`, so a reviewer can tell a close that is clean because things were fixed from one that is clean because things were silenced. ## Writer New `audit-open acknowledge` verb snapshots current state itself, so the marker is never hand-authored from workflow prose — the gap that left the STATE.md block with no writer, no schema and two conflicting formats. Writes route through the existing path-confinement seam. ## Two deliberate limits, failing closed Heading-delimited deferred entries (#3457) are REFUSED with `unsupported_heading_shape` rather than edited, because mapping a heading entry back to its exact source span is not safely derivable when headless and heading entries interleave in one file. A loud refusal beats a mis-targeted write. A quick task with no summary gets one created to carry the marker, since there is otherwise nowhere to put it. ## Tests Self-invalidation is the important one and is covered per category: acknowledge, then change the status or question count, and the item resurfaces. Also malformed markers not suppressing, `status:` byte-unchanged after acknowledging, the writer refusing a path outside the project, and the four original #3458 scenarios unchanged. Closes #3458 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3458): wire [A] to the acknowledge verb and converge the disclosure table Consumer side of the suppression seam. ## The workflow stops hand-authoring the mechanism `[A]` now calls `audit-open acknowledge` once per open item, then writes the STATE.md `## Deferred Items` table as before. The table stays as a human-readable disclosure; it is no longer the mechanism. That closes the gap where the block had no writer, no schema and no reader — the marker is now written by the tool, which snapshots current state itself. The `[R]` / `[A]` / `[C]` prompt is unchanged, `[C]` still means "Cancel — exit without closing", and no carry-forward option is invented. The all-clear branch now distinguishes a close that is clean because items were FIXED from one that is clean because they were ACKNOWLEDGED, using the `acknowledged.total` count, and carries that into the MILESTONES.md disclosure line beside the existing override count. A clean close that was bought with acknowledgments should say so. ## Format drift resolved Two incompatible `## Deferred Items` shapes shipped simultaneously — 3 columns in the workflow, 4 in the template, with different body lines. Converged on one 5-column shape carrying the source Milestone, since archived items now appear and the archived-milestone disambiguator was previously discarded at write time. The workflow enumerates the categories instead of trailing off in `...`. ## Ack fragment bookkeeping `complete-milestone.md` grows 6,764 bytes (31,228 → 37,992; cap 61,440), covered by a new `tests/emitted-drift-acks/3458-*.json`. `2962-zsh-nomatch-for-glob-portability.json`'s `complete-milestone.md` entry is REMOVED — the no-duplicate-path rule hard-blocks two sources naming one path. That entry is spent: the nullglob shim it acknowledges is present in both `origin/next` and the CI emitted baseline `fd2b97a5`, so its ripple is already absorbed and it can never clear anything again — verified directly, not assumed, and the gate's own message directs deleting spent entries. Its other three files' entries are untouched. `scripts/sync-runtime-launcher.cjs` wanted to rewrite `explore.md` as well — pre-existing drift unrelated to this change, reverted. `complete-milestone.md` still carries exactly one canonical preamble. Docs cover the verb's real flag surface, the marker's verdict-preserving and self-invalidating behavior, and the new `acknowledged` count. A second `Added` changeset covers the verb, since the existing `Fixed` fragment describes only the archived-phase scanning. Closes #3458 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3458): close three blockers in the acknowledgment seam Adversarial review of the seam. Three BLOCKERs, one of which disproves a safety claim I published in the PR body, the changeset and the docs. ## The claim was false; the code is fixed rather than the claim softened I wrote that "a stale acknowledgment can never hide a NEW problem". It could. `context_questions` snapshotted only the question COUNT, so replacing two acknowledged questions with two brand-new blockers kept the item suppressed. `uat_gaps` snapshotted only `status`, so adding five more pending scenarios (`open_scenario_count` 1→6) kept it suppressed. The snapshot now identifies CONTENT, not size: a digest of the whole question set, and a status + open-scenario-count composite. Any edit invalidates. The other seven categories were checked and their single tracked dimension is already the whole story. Both disproofs now resurface the item. ## Writing to the wrong line, and reporting success `acknowledgeDeferredItem` built an unanchored regex and exec'd it over the whole file while match-selection and the ambiguity guard ran over the section body only, so the write landed at the first match ANYWHERE. A file with `# Notes` holding `- Fix the parser` above a `## Deferred Items` section holding the same bullet: the CLI exited 0 saying `acknowledged: true`, injected `status: acknowledged` under `# Notes`, and re-audit still reported the entry open. It corrupted unrelated content, suppressed nothing, and claimed success — and since `--file` is unconstrained the same path could inject into a UAT or VERIFICATION body. Matching is now anchored to the selected section, and the matched span is re-verified against the selected entry before any write; a mismatch refuses with `match_verification_failed` rather than writing. ## Acknowledging todos hid the ones never shown `scanTodos` capped at five files and then checked acknowledgment. With seven todos, acknowledging the five that were LISTED drove `todos: 0`, `has_open_items: false`, and items six and seven never appeared in any later scan. The workflow's own "repeat until no todos items" remedy terminates after one pass. Pre-feature this was unreachable because the count was pinned at five. That is silent over-suppression — the exact direction this PR exists to remove. Acknowledged items are now filtered BEFORE the display cap, so unacknowledged todos beyond it still drive the count. ## The [A] branch could not fail closed Every acknowledge call sat in a `cmd | while read` pipeline with no status accumulation, so any refusal was discarded and the close proceeded as `override_closeout`. Separately, `io.output` swaps payloads over 50000 chars for an `@file:<path>` sentinel — every `jq` would then fail, every loop body run zero times, nothing be suppressed, and the close happen anyway. Both closed: failures accumulate across all invocations and halt before close, and the sentinel is dereferenced using the same pattern `verify_readiness` already uses for `INIT_MANAGER`. Quoting was verified sound by the review and is left alone. ## Also Suppression is now visible in the human report, not only `--json` — the "clean because fixed vs clean because silenced" distinction was promised for the surface an operator actually reads. The CRLF-preservation branches in the writer were dead: every `.md` write goes through `_normalizeMd`, which normalizes line endings and blank lines whatever the writer does. Deleted and documented rather than left as code that cannot run. ## Why these shipped The review named it exactly: there was no coverage for `unsupported_heading_shape`, `ambiguous`, `not_found`, duplicate-text mis-targeting, todos beyond the cap, or CRLF. All are now tested, alongside both snapshot disproofs and the mixed-section fixture. Closes #3458 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3458): align the items-open footer wording with its assertion Remote runner red on one test: the items-open footer must match `/previously acknowledged item/i`. The disclosure was NOT missing — the items-open branch already printed "N additional items previously acknowledged and still suppressed." The word order simply did not match the regex the test in the same change asserts. A wording mismatch between my own test and my own implementation, not a behavior gap. Reworded to "N previously acknowledged items also suppressed above the M open items", which satisfies the assertion and states the relationship between the two counts more plainly than the original did. Swept `formatAuditReport` for other branches that could skip the tally: the only early return is the all-clear path, which already discloses it. `scan_error` sentinels are filtered per category and excluded from `counts.total`, so an all-error project falls through to that same branch. No inconsistency remains. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3458): splice by carried span, digest the untruncated question set Security review of the writer. Both findings are the same shape, and both are cases where an earlier fix of mine was incomplete in the same direction: a value derived for DISPLAY was reused for an IDENTITY or LOCATION decision. ## Writing to the wrong entry, again The previous fix anchored matching to the `## Deferred Items` SECTION but still re-found the entry inside it with an unanchored regex, so the write landed at the first SUBSTRING occurrence rather than the entry's own span. The `match_verification_failed` guard could not catch it, because the mis-targeted span is byte-identical to the target. Probe-confirmed, in a cloned repo's own artifact: - CRITICAL unfixed auth bypass see also: - minor typo - minor typo Acknowledging "minor typo" appended `status: acknowledged` into the CRITICAL entry, suppressing it at every future close, while the typo stayed open — exit 0, `"acknowledged": true`. A variant where the target text appears inside unrelated prose split that line mid-sentence, acknowledged nothing, and still exited 0, so the workflow's `ACK_FAILURES` halt never fired. Fixed structurally rather than with a better regex: `splitGapsEntriesWithSpans` carries each entry's own character span out of the splitter, and the write splices by that recorded span. The location is already known at selection time — re-deriving it by searching was the entire defect class. Added as a sibling so `splitGapsEntries`' three existing callers are untouched. With index-splicing, `match_verification_failed` becomes a genuine independent cross-check instead of a guard that could never fire. ## The digest was blind past the third question `deriveOpenQuestions` truncated to three questions, and clamped each to 200 chars, BEFORE the digest hashed it — so the snapshot could not see the fourth and later. Ship three innocuous questions, acknowledge, then add real blockers, and they are permanently invisible: measured `open=0, acknowledged=1`, report "All artifact types clear." That is the same self-invalidation property this digest was added to guarantee one revision ago. The digest now covers the untruncated list; truncation is display-only. Found while fixing it: the previous digest joined on a literal raw NUL byte embedded in the source — collisions are constructible, and reachable through attacker-controlled YAML `\x00` escapes. Verified both ways. Replaced with a length-prefixed encoding so no two question sets can collide by concatenation. ## Sweep Because this is the third incomplete fix on this seam, every identity and location derivation was swept for the display-vs-identity confusion: uat_gaps uses status plus a full-content count, the other seven categories use a scalar status or presence, the deferred `--text` identity is never truncated, and all five flat categories resolve their file by path rather than by content search. No further instances. Closes #3458 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3458): correct two assertions that over-reached the measured behavior Remote runner red on two of the F1 tests. The source is correct — reproduced both fixtures against the built CLI — and both failures were bugs in the assertions I wrote. `src/` is untouched by this commit. The first is worth recording. It computed the CRITICAL entry's block as content.slice(content.indexOf('- CRITICAL'), content.indexOf('- minor typo')) and `indexOf` found the FIRST SUBSTRING occurrence, which lives inside that entry's own continuation line ` see also: - minor typo`. The block was truncated mid-line, so the assertion could never match. The test committed the exact first-substring-match mistake it exists to catch, one revision after that mistake was fixed in the source. The second asserted `deferred_items === 0` after acknowledging the typo entry, but the decoy `- Note: reference - minor typo elsewhere, ignore` is itself an open entry and was never acknowledged, so the correct count is 1. It now also asserts WHICH item remains open — that is what actually proves the right entry was suppressed, and the original assertion would have passed even if both had been silenced. Both now derive their expectations from measured CLI output. A comment records that the write seam normalizes markdown (`_normalizeMd` inserts a blank line before a list item following a non-list line) so the inserted line is not later mistaken for a regression; that is repo-wide behavior for every `.md` write through the single write projection, not something this change should diverge from. Root cause of both: the previous two dispatches verified behavior with direct CLI probes but never executed the test file, so assertions could over-reach what had actually been measured. Every other assertion added in those two commits has since been re-derived from real output; no further mismatches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
147856040b |
fix(#2873): close review findings across fences, sanitizer and docs
Isolated security review found resolveSpecRootReference's fence tracker toggled on any delimiter, so a backtick fence could be closed by a tilde one and an include in the gap was rewritten inside a code block. Fixed by reusing scanFencedBlocks - the canonical engine already behind stripFencedCode and extractFencedBlock - rather than carrying a fourth copy of fence detection, which also closes the duplication the standards review flagged. sanitizeForRender now strips combining marks and zero-width characters alongside the ANSI, control and bidi classes it already handled. Adds the C, E and F matrix rows the spec review found missing, including installer-level coverage that spawns the real install rather than calling the report builder. Ships the how-to, the reference and command docs in five locales, the changeset, the inventory and glossary entries, and regenerates health.md for the new W028 rule. Refs #2873 |
||
|
|
6dbc124018 | enhance(#3180): the sibling validators share one envelope and one owner — Phase 12 (#3407) | ||
|
|
c67992de87 |
docs(#3309): document --backfill and the --repair DESTRUCTIVE-refusal change
docs/COMMANDS.md's /gsd-health section never documented --backfill at all, and predates this phase's breaking changes: --repair no longer auto-applies resetConfig/regenerateState (both destructive), and W021/ W017 split into W026/W027 for their previously-conflated second subjects. Required by lint:docs, which needs a docs/ touch alongside any Changed-type changeset fragment. |
||
|
|
0396d9cab1 |
enhance(#2483): stop the claude reviewer lane from inheriting CLAUDE.md + auto-memory (#2493)
* enhance(#2483): env-guard the claude reviewer leg against CLAUDE.md injection
The claude reviewer in workflows/review.md was a bare headless `claude -p`
spawn run from the project cwd, so it inherited the invoking user's global
CLAUDE.md, the project CLAUDE.md, and Claude Code auto-memory.
That made it the only reviewer leg seeing anything beyond the prompt file.
gather_context assembles PROJECT.md, the roadmap section, every PLAN file,
CONTEXT.md, RESEARCH.md and REQUIREMENTS.md into the prompt before any
reviewer runs; the gemini leg receives only that prompt and the codex leg
runs --ephemeral. Beyond the measured ~4k tokens/spawn, the asymmetry cuts
at the workflow's own premise: "independent review" meant something
different for the claude leg than for the other two.
Guard both dispatch lines with a per-invocation
`env CLAUDE_CODE_DISABLE_CLAUDE_MDS=1`. `env`, never `export` — the flag
must not leak into the orchestrating session (which may itself be Claude
Code on the SELF_CLI="auto" path) or into any later spawn.
review.md is the only claude -p call site in the installed tree, so this is
two lines on one surface. The self-skip logic is untouched.
* enhance(#2483): fix CRLF-fragile split and regenerate workflow baselines
Two CI failures from the first push, both mine:
1. lint-tests: the new regression test split readFileSync content on a
literal "\n". On a Windows git-autocrlf checkout that leaves a trailing
"\r" on every line (local/no-crlf-fragile-split). Use .split(/\r?\n/).
2. golden-install-parity / workflow-size-budget / workflow-compat: editing
gsd-core/workflows/review.md changes its content hash and byte size, and
both are pinned in committed baselines. Regenerated via the repo's own
generators (npm run size:baseline, npm run gen:golden).
The regenerated diffs are review.md-only: exactly one hash line per
golden-install-parity fixture and one size entry in workflow-size-baseline
— no unrelated drift swept in.
Full suite now green locally: 2113 pass, 0 fail, 3 skipped (run with HOME
and CLAUDE_CONFIG_DIR overridden to throwaway dirs; live profile verified
untouched afterward).
* enhance(#2483): adapt guard-test matcher to the effort-args dispatch reshape
The effortSurface wiring (#2481) reshaped the bare-model dispatch to
`claude $CLAUDE_EFFORT_ARGS -p -`; the invocation matcher's dash-first
form could no longer see it, and the count assertion failed exactly as
designed. The matcher now tolerates variable expansions between `claude`
and its first literal flag. Negative-controlled both ways: a stripped
guard and a deleted dispatch line each still fail.
* enhance(#2483): also guard the claude leg against auto-memory injection
CLAUDE_CODE_DISABLE_CLAUDE_MDS suppresses CLAUDE.md file loading;
auto-memory is an independently-toggled mechanism with its own flag.
Add CLAUDE_CODE_DISABLE_AUTO_MEMORY=1 to both dispatch lines, correct
the docs/COMMANDS.md and changeset claims that credited the first flag
with covering auto-memory, and extend the regression test to require
both flags on every claude invocation (negative-controlled: 2/4
assertions fail with the new flag removed).
* enhance(#2483): match the claude binary in command position, not argument position
The line-oriented invocation matcher counted any line where the token
`claude` was followed by a flag. #2589 (landed on next as
|
||
|
|
e87fb409ee |
enhance(#2573): stamp STATE.md with its commit and surface a freshness hint (#2622)
* enhance(#2573): stamp STATE.md with its commit and surface a commit-age freshness hint Adds a `state_head` stamp to STATE.md and derives a tri-state commit-age freshness proxy (state_commits_behind / state_commit_stale) through state.cjs's readStateHeadFreshness, surfaced on smart-entry signals and as health W024. The proxy is advisory: classify() deliberately does NOT consume it (ADR-1787 locks the classification/routing boundary — a signal, not a route). Composes with #3099 and #1882 (both merged to next after this branch): the commit-age proxy reads `state_head` while the LAST_ACTIVITY_UNPARSEABLE diagnostic reads `last_activity` — two different fields, not "two staleness signals on one field." A new regression test asserts a STATE.md carrying both an unparseable last_activity AND a valid state_head resolves each independently (diagnostic fires once; freshness reads state_head, commits_behind 0). Rebased onto next (flattened): resolved the add/add conflicts in src/smart-entry.cts (kept both the #2573 freshness import/derivation and the #3099 diagnostic import/call) and tests/smart-entry.unit.test.cjs (kept both describe blocks). Drift-ack for health.md's W024 row is unchanged (12348 B). Tests: smart-entry 62, state/state-transition/health/verify 639, all pass. * chore(#2573): allowlist health-validation test in the prompt-injection scan The scanner's `exec('` code-execution pattern matches the benign `re.exec('<phase-id>')` RegExp method calls in the phase-ID grammar tests (pre-existing: 16 such calls on next, this PR adds none). The file entered the diff-mode scan's changed-file set only because #2573's W024 state_head assertions touch it. Allowlist it alongside the other test files that carry pattern-matching content as data (same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class). Scanner self-test 38/0; diff scan 14 files, 0 findings. |
||
|
|
aceea3ce4a |
refactor(#3217): withhold a percentage when its scope is not complete (#3318)
* wip(#3217): rule-4 scope withholding — parked, two open findings Implemented but NOT shippable. An isolated review found buildStateFrontmatter still hardcodes SCOPE.COMPLETE, so state json reports percent 0 where roadmap analyze, stats and query progress all correctly report null on the same disk state - rule 4 reintroduced at a site this phase claims to close. Also: roadmap analyze emits scope complete beside progress_percent null with nothing explaining it. Parked to build Phase 4 (#3186) first, which is unblocked. Findings recorded in .gsd/phase/refactor-3217-completion-ratio-scoping/60-review.json. * fix(#3217): withhold the sync percentage on a non-complete scope The parked blocker is fixed - buildStateFrontmatter no longer hardcodes SCOPE.COMPLETE, and the prose Progress fallback is gated too, which was a second leak found while tracing the first. roadmap analyze exposes progress_scope so a consumer can tell WHY a percentage is absent from the JSON alone. Then a residual gap was reproduced rather than assumed. cmdStateSync carried the same hardcode behind a written reason claiming it did not reproduce. It did: on a TRUNCATED window and on UNSCOPED row 4, state sync wrote Progress 0 percent to 100 percent while state json, roadmap analyze, stats and query progress all withheld - and it persisted a self-contradictory file, body claiming 100 percent while its own frontmatter correctly omitted percent. The excuse was also wrong. syncRoadmapRaw is already parsed in that function and is exactly what produces a real scope, so there was a scope to pass. Threaded through listMilestonePhaseDirs; a non-complete scope now skips the write with a reason in changes. milestoneBounded stays as the orthogonal 1761 guard for row 5. Second time this epic a does-not-reproduce claim was too generous. Recorded in ADR Amendment 8 as a correction rather than a quiet rewrite. Verified on the remote runner. * test(#3217): give the withholding fixtures a resolvable scope 40 matrix failures, all fixture drift - no code regression. My own hypothesis that this was over-withholding was wrong and is recorded as such: the worry case, a plain ROADMAP with Phase entries and no version heading, resolves to complete exactly as ADR 7.1 says it should. The real causes were two fixture shapes. Most had no ROADMAP.md at all, which is unreadable via a pre-existing graceful path, and asserted a numeric percent. The five vscode, pi-extension, mcp-server and shell-projection failures were that shape - bare temp dirs using progress json as a reachability proxy while asserting typeof percent is number, which under rule 4 is now null. The rest had a version token in a title or heading with no STATE.md milestone pointer to resolve it, which is classification row 4, versioned but unresolved, so withholding is correct per the contract. Verified on the remote runner. * test(#3217): make the LM-tools reachability tests dispatch against their fixture The gsd_progress reachability test was never testing its fixture. invoke() resolves cwd from vscode.workspace.workspaceFolders by design (the real LanguageModelToolInvocationOptions has no cwd field, per the 2103 fix in extension.js), the mock had no workspace at all, and the test passed a cwd option nothing reads - so it dispatched against the repo working directory. Writing a ROADMAP into the temp dir had no effect. Rule 4 only made it visible. Fixed by mocking workspaceFolders. The two siblings in the same file carried the identical dead cwd and were dispatching against the repo too; they were not failing only because their assertions did not touch scope-dependent output. Both now use their own fixture with assertions unchanged - the no-planning fallback paths already satisfy them honestly. Re-scanned the other five reachability files: no further instances. They thread cwd into parameters that genuinely read it, not through an options shape that ignores it. Verified on the remote runner. * chore(#3217): backfill changeset PR number pr:0 placeholder replaced with the real number now that #3318 exists. * ci(#3217): give the coverage merge enough heap for the merged shards The coverage gate OOMed at exit 134. c8 report merges three shard artifacts, roughly 358MB of V8 dumps in coverage/tmp, and died holding their per-file position maps at the ~4GB default heap. Verified as this branch's delta rather than pre-existing: the same job succeeded on next at 14:18, after phases 4 and 5 merged. Both coverage-gate steps get the bump because both re-slice the same merged data. 8192 doubles what failed and leaves headroom on a 16GB ubuntu runner, matching the idiom the shard step already uses at 6144. This is a memory bound, not a change to what is measured. No threshold was touched. The test file was checked for gratuitous subprocess spawning and is already reasonable at 43 spawns, each a distinct fixture-by-surface pairing. Verified on the remote runner. --------- Co-authored-by: sim <sim@local> |
||
|
|
e201cde73c |
refactor(#3186): one shared phase-completion predicate, disk-strict (#3306)
* docs(#3186): record the disk-strict completion decision in ADR-3180 7.4 The maintainer decided #2957 on 2026-08-08: disk state is authoritative and a ROADMAP checkbox is a human annotation with no machine authority. Section 7.4 still carried the OPEN QUESTION and was marked blocked, so the contract said one thing and the tracker another. Recorded per section 7's own rule - a behavior not stated there is not decided, and amending a rule is an ADR amendment rather than a code change with a comment. The decision comment names Phase 4's PR as the carrier of this edit and makes it an acceptance criterion that the text be in the tree before implementation begins, so this lands first, alone, ahead of any code. Also clears the stale blocked-on-2957 row in the guard roster. * refactor(#3186): one shared phase-completion predicate, disk-strict isPhaseComplete in verification.cts becomes the single owner. It calls readVerificationStatus UNCONDITIONALLY - plan count is not a precondition - so a zero-plan phase with a passing VERIFICATION.md is complete. That is #3168: init gated the read on a plan count and synthesized a not_required sentinel, so phase.complete succeeded while init.manager reported incomplete for the same phase. The guard, built and run before scope was fixed per Amendment 3, found 9 re-derivations where the ADR named 3. Four were unnamed, including one in the prompt layer: mvp-phase.md ORed a ticked checkbox with disk status, which under disk-strict is the divergence itself. Per the #2957 decision, a ticked ROADMAP checkbox is a human annotation with no machine authority. The overrides in roadmap analyze and init manager are deleted rather than generalized; the user's checkbox stays in ROADMAP.md, only its authority goes. scanPhasePlans.completed and buildWorkstreamInventory are deliberately NOT folded - they answer 'are all plans summarized', which is a different question, and folding them would either over-report completion or invert the dependency direction between Phase 1's owner and this one. Verified on the remote runner. * fix(#3186): close seven review findings and record the missing-verdict rule The isolated review reproduced a write-path regression I introduced: migrating cmdRoadmapUpdatePlanProgress dropped its summaryCount>=planCount gate, so a phase with a fresh passing verification plus a newly-added unsummarized plan reported complete AND wrote a checkbox into ROADMAP.md while phase complete refused. The owner stays right per 7.4 - plan count is not a completion precondition - so the gate is restored at the write site as an explicit composition, mirroring the separate 2648 unexecuted-plan gate cmdPhaseComplete already carries. The spec axis was right that my 0.x-split reasoning was too permissive. The 2957 decision names buildStateFrontmatter as one of the three that must converge, and buildWorkstreamInventory combined a summaries-met local with verification data to decide the same verdict - Decision 4(c)'s named bypass, and it reproduced 3168 in a third surface. Both now route through the owner. The raw scanPhasePlans helper stays: it answers are-plans-summarized, which genuinely is a different question. Maintainer decision recorded in 7.4: a missing verdict is not a passing one, so an absent VERIFICATION.md means not complete everywhere. That retires 2645's verifier-disabled tolerance and inverts its Goodhart incentive - deleting the evidence now lowers completion instead of raising it. Guard hardened: block-form count gates and algebraic restatements are caught, and the header now discloses its remaining limits instead of overclaiming. Verified on the remote runner. * fix(#3186): route state sync through the owner and catch bare completed reads The matrix found 52 failures. 51 were fixtures asserting the old semantics: a phase with plans and summaries but no VERIFICATION.md used to count complete and correctly no longer does. Each fixture now carries a passing verification where that is what the test was actually about, rather than having its assertion weakened. The 52nd was a real 10th re-derivation the guard could not see. cmdStateSync destructured scanPhasePlans().completed directly - a bare field read, not a comparison - and used it as a completion verdict, so state sync and state json disagreed on completed_phases for identical disk state. Routed through the owner. Guard gains shape (d): any read of .completed off a scanPhasePlans() result outside plan-scan.cts, in chained, destructured and indirect forms, function scoped with no line window. It cannot tell a summaries-met read from a completion read - that is data flow - so it flags every one and requires a written-reason exemption, which is the same discipline shapes a-c already use. The blind spot is disclosed in the header rather than overclaimed. The emitted-attribution failure was also mine, not pre-existing: the mvp-phase.md checkbox-OR removal moves emitted bytes, acknowledged in tests/emitted-drift-acks. Verified on the remote runner. * test(#3186): give the nested-plans sync fixture a passing verification Last 3 matrix failures were one failure echoing up two describe levels. Phase 01-alpha had plans and summaries but no VERIFICATION.md, so under disk-strict completed stayed 0 and no Progress change was emitted - correct new behavior, not a regression. Added the passing verification rather than dropping the Progress expectation, so the test still covers what #3257 is about: that a nested plans/ layout is counted and not undercounted. Probe against the built lib confirms Progress: 0% -> 50% alongside Total Plans in Phase: 0 -> 3. * chore(#3186): backfill changeset PR number pr:0 placeholder replaced with the real number now that #3306 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
d28ab7c8f7 |
enhance(#3243): sync installed codex .toml model/effort to the passive posture (#3296)
* feat(#3243): sync installed codex .toml model/effort to the passive posture Implements ADR-2313 D7, and owns the Codex .toml typed IR that Phase 1's review assigned to this phase. The IR exists for a structural reason, not tidiness: this phase has to PARSE these files, and a parser kept bug-compatible with a separate renderer is the generative-fix-divergence shape this epic already dealt with once for the model predicate. So Phase 2's parsing MOVES here rather than being copied — agent-install-check now imports it, and its test file passing unchanged is the proof the extraction altered nothing. The load-bearing property is byte-identical round-trip: render(parse(x)) === x. Without it a sync silently reformats a user's file — line endings, key order, BOM, trailing newline — turning a two-line repair into a whole-file diff in their dotfile repo. The IR keeps original lines and removes targeted ones rather than reconstructing from parsed fields, which is what makes that property hold. It also reconciles a real contradiction between Phase 2 and ADR-2313. An unterminated developer_instructions block: the reader excludes the rest of the file, deliberately failing toward a false positive, because misreading prose as a pin only wastes a user's time. The writer must refuse, because proceeding on a malformed document rewrites it. A false positive is the safe direction for a reader and the dangerous one for a writer. So the parse reports the fact and the two consumers branch on it — one parse, one truth, two policies, instead of two parsers that agree today. The sync leaves a legal real-Codex pin and its coupled effort untouched, reported skipped rather than synced; strips a stale Anthropic or tier model and an orphaned effort; keeps dry-run as the default; refuses any file whose parse fails; and skips symlinks exactly as the Claude path already did. The Claude path itself is byte-identical. PARSE_REASON.NO_HEADER from the ADR's illustrative snippet is deliberately not implemented — a missing header is legal, not an error, so it would be a dead enum member that the enum-lock test then pins. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): preserve per-line endings and make the codex write atomic Two findings from an isolated review, both in the write path. BLOCKER: mixed line endings broke the byte-identical round-trip. `eol` was a single whole-file flag and split(/\r?\n/) discarded each line's own terminator, so render re-joined with ONE style and normalized every line — even with zero strips performed. A file with one CRLF line and the rest LF came back fully converted. That falsified the A14 guarantee, violated the design's "must not silently rewrite every line", and made the CONTEXT.md glossary claim wrong. It was untested because A12 and B15 only cover PURE CRLF; no mixed-ending fixture existed anywhere. Fixed by keeping each line's terminator alongside its content, so render is a plain concatenation and a strip removes only the target line and its own terminator. `eol` survives as informational metadata that render never reads. Seven fixtures added for the paths nothing exercised: mixed endings unmodified and with a strip, a lone \r, a file ending on the block's closing ''' with no newline, multiple trailing newlines, a BOM-only file, and an empty file. MINOR, but it contradicted this phase's own contract: the write was in-place open-truncate, so a failure between truncate and completion leaves a truncated .toml — exactly what ADR-2313 says must never happen. The Codex path now writes a sibling temp file and renames over the target, which is atomic on one filesystem, with cleanup on failure. It uses the repo's existing retryRenameSync rather than a hand-rolled rename, and deliberately NOT platformWriteSync, whose normalizeContent would mangle the very CRLF and trailing-newline bytes the round-trip property exists to preserve. The Claude path keeps its in-place write untouched. It has the same shape, but changing it is not this phase's business and its tests must stay byte-identical. B20 previously mocked writeFileSync to throw BEFORE touching anything, so it proved nothing about a mid-write failure — its passing comment was true only because of how the mock was built. It now performs a real truncated write wherever writeFileSync is called, catching both the naive direct-to-target path and the new temp path, and asserts the target is byte-identical afterwards with no stray temp file left. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): preserve the trailing-newline state when stripping a last line Caught by B17, one of this phase's own tests — the suite working, not a test problem. Content is reconstructed as the concatenation of lines[k] + terminators[k], so a file with no trailing newline has '' as its last terminator. removeLine spliced out both arrays at the same index, which is right for a middle line but wrong for the last one: it dropped the empty terminator and left the PREVIOUS line's newline in place. A file ending `...\nmodel = "sonnet"` with no trailing newline came back as `...\n`, gaining a newline the user never wrote. The new last line now inherits the removed line's terminator, so a removal leaves the file exactly as if that line had never been written. Removing the only line yields an empty file rather than a stray terminator. Both stripModel and stripReasoningEffort funnel through the one removeLine, confirmed rather than assumed, so a single fix covers both — including the row-B7 shape where a stale model and its orphaned effort are removed in sequence and the second removal targets the last line. Two of the four new cases are honestly not red-first and say so in their comments: removing a last line that HAS a trailing newline only exposes the bug under mixed EOL, since uniform files coincidentally have equal terminators on both sides; and removing the only line already degenerated correctly through Array.slice. They are kept as guards for the new branch rather than dressed up as catches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3243): document the codex repair path and close the loop How-to: the Codex-400 entry added in Phase 2 told users to re-run the installer, because that was the only repair available then. It now leads with `effort sync` and keeps the reinstall as the alternative, with the reason to prefer one — a reinstall regenerates the agent files wholesale, so anyone who hand-edited theirs loses those edits. Detect, preview, apply is now one continuous path in one place. Reference: docs/COMMANDS.md had no `effort sync` entry at all — the same gap `validate agents` had in Phase 2, found the same way. The entry documents BOTH runtimes, because the command genuinely forks on runtime and describing only the new half would misdescribe it. The write flag is `--apply`. The design doc and test matrix both said `--no-dry-run` throughout, which does not exist — verified against the actual arg parser in gsd-tools.cjs before writing. Documenting a flag that does not exist is worse than documenting nothing, because it fails at the moment someone needs it. Both surfaces state that only the targeted lines are removed and every other byte is preserved. That is a user-visible guarantee rather than an implementation note: it is the difference between a two-line diff and a reformatted file in someone's dotfile repo, it is what the IR's round-trip property exists to deliver, and writing it down makes it a contract a future change has to break knowingly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): inherit the trailing-newline state, not the line ending style My previous rule was subtly wrong and this phase's own test caught it. "The new last line inherits the removed line's terminator" copies the removed line's STYLE as well as its presence. A26 uses mixed endings on purpose — line one terminated \r\n, the model line terminated \n — so inheriting silently rewrote line one's ending to \n. That is precisely the defect class the mixed-EOL blocker fix existed to eliminate, reintroduced one layer down by the fix for it. The correct rule inherits the EMPTINESS only. If the removed line had no terminator, the new last line loses its own, preserving "this file has no trailing newline". Otherwise the new last line keeps its own terminator: it is already a newline, and already the right style for that line. A26's assertion moved too, and that deserves saying plainly rather than burying: it previously encoded my wrong rule. Changing a test to match the implementation is usually the mistake, so it was checked from first principles instead — a file whose first line ends \r\n and whose last line ends \n, with that last line removed entirely, must be the first line with its own \r\n intact. The new expectation is what the user's file should actually look like; the old one was wrong. A29 adds the interaction nothing covered: the compounding case (strip a stale model, then its orphaned effort, the second removal landing on the last line) with non-uniform endings either side. The two fixes meet there and nothing exercised the meeting point. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3243): drop the phantom trailing line from the IR representation Root cause, not another patch on the removal rule. Three consecutive fixes there each surfaced the next issue, which was the signal that the data model was wrong. splitPreservingTerminators left a phantom empty final entry for any file ending in a newline: "a\nb\n" became lines ['a','b','']. So for the common case the real last content line was NOT the last array element, removeLine's isLastLine check never matched it, and every rule I gave was reasoning about the wrong element. What hid it: render was already a plain concatenation, so a phantom empty line with an empty terminator contributes nothing to the output. A14's byte-identical round-trip could never have caught it — the defect is byte-neutral until a removal shifts the index arithmetic under it. That is worth recording, because "the round-trip test is green" was exactly the reassurance that kept the search pointed elsewhere. The representation is now 1:1 — terminators[i] follows lines[i] and may be '' — with no phantom, verified across empty, no-trailing-newline, trailing-newline, blank-line and mixed-CRLF inputs. render stays a plain concat and needs no special cases. With the phantom gone the removal rule is correct as stated and finally applies to the genuinely last element. Consumers checked rather than assumed: the block-range detector and header scanner are agnostic to array shape, and Phase 2's reader uses its own independent split, so tests/agent-install-check.test.cjs is untouched and still passes unchanged. One test expectation was wrong and is corrected rather than quietly adjusted: A18 asserted a 7-element terminators array whose trailing '' was the phantom itself. It now asserts the six real terminators, which is what the invariant lines.length === terminators.length requires. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3243): backfill changeset pr number (#3296) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2dbee3ebdd |
enhance(#2229): add three-way claim disposition (admit/refute/abstain) to /gsd-explore research pass (#2543)
Closes #2229. Each claim surfaced by /gsd-explore's research pass is dispositioned admit, refute, or abstain, with abstentions routed to a visible ledger instead of being smoothed into confident prose. Refute and abstain are separated by whether the disagreeing source is authoritative for that claim; a strong prior is never authoritative alone. Two guards ride with it: conflict-abstention, and a tier floor that presents a would-be admit as an abstain when the researcher's resolved tier is the budget tier or cannot be determined. To make that floor enforceable, resolve-model now emits the effective tier (--pick tier). It was already computed above the resolve_model_ids omit gate but was unreachable from a workflow, which left the floor inert on every non-Claude install - the model id is blank under omit and runtime-substituted where a tier map exists, and the profile defaults to balanced. The tier signal mirrors every resolution step that can change which tier runs, including the model_policy preset, and reports unknown rather than guessing. Output is additive; model, profile and effort are unchanged. Two residuals are disclosed in the workflow rather than papered over: a raw-model-id model_overrides pin reports unknown and is floored (fails closed), and a model_profile_overrides entry repointing a tier at another tier's model can under-report (fails open, and predates this change). Admin merge used only to satisfy the missing secondary reviewer on a single-maintainer PR. No CI failure and no conflict were bypassed: 38 checks green, remote runner 32255/32255 on both Node lanes. |
||
|
|
693f12ad56 |
refactor(#3187): give state field extraction one canonical owner (#3283)
* refactor(#3187): give state field extraction one canonical owner stateFieldValue in state-document.cts becomes the single owner of the #1760 frontmatter-then-body fallback chain. The new whole-repo guard found 14 independent re-derivations where the epic scoped 5, all now routed through it: cmdStateSnapshot (11), cmdStatePrune (2) and smart-entry fmScalar (1). state validate was a gate that could not fail. Every warning it could emit sat behind a phase resolved without the frontmatter tier, so a STATE.md whose phase lives only in frontmatter skipped the drift scan entirely and returned valid:true. It also read unstripped content, letting a frontmatter status: key shadow the body field (#1255 class). Both fixed; output gains a scope field so could-not-look stops being output-identical to looked-and-clean. Verified on the remote runner. * docs(#3187): document the state validate scope field and its reason codes Adds docs/how-to/interpret-state-validate-results.md so a reader can tell nothing-to-report from could-not-look, updates the COMMANDS.md and USER-GUIDE.md entries, corrects the CONTEXT.md glossary overstatement about Current Position sole ownership, and drops the changeset fragment. * fix(#3187): close three drift-guard evasion shapes and test the refuse path The isolated adversarial review found the ladder detector was evadable by ordinary reformatting, not just deliberately: a member or computed operand (fm.key / fm[key]) missed the bare-identifier backreference, a swapped tier order missed a hardcoded number-then-boolean sequence, and a ladder wrapped across lines missed single-line detection. All three now caught, each with its own test plus a proven boundary control. The frontmatter-parse refuse path on the destructive complete-phase route was unreachable and therefore untested. It is now driven by an injected parse failure and asserts STATE.md is byte-identical after the refusal, rather than shipping untested defensive code on a path that rewrites user state. Verified on the remote runner. * fix(#3187): widen the drift guard to the prompt layer and disclose tier-2 changes The code-review spec axis found the guard's scan surface was src/ only, which is Decision 4(d)'s forbidden allowlist one directory wide - and it had a live miss: gsd-core/workflows/smart-entry.md tells an agent to read status from frontmatter or the body, a prose expression of this same chain. The surface now covers the prompt layer. That one site carries a permanent written exemption rather than a ratchet: it is the gsd-tools-is-down fallback, so it cannot call the owner by construction, and a ratchet would imply removable debt that does not exist. Two tier-2 output changes shipped undisclosed and are now named in the changeset and docs: complete-phase's idempotency guard consulting frontmatter, and the workstream inventory resolving frontmatter-only fields. docs/COMMANDS.md gains a state complete-phase entry, which it never had. Also records Amendment 5 on ADR-3180, extracts the duplicated frontmatter-parse block the epic's own thesis forbids, and re-points two assertions from free-form warning prose onto the structured drift object. Verified on the remote runner. * chore(#3187): backfill changeset PR number pr:0 placeholder replaced with the real PR number now that #3283 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
4a1ed2531f |
enhance(#3242): validate codex .toml model posture, not just presence (#3290)
* test(#3242): failing-first suite for the codex posture health-check Specifies ADR-2313 D6 before the implementation exists, so the tests bind to the contract rather than to whatever the code happens to do. RED is established by construction, not by a remote run: checkCodexModelPosture and POSTURE_REASON are absent from the compiled lib today, so every row fails on the missing export. A remote checkpoint here would prove only that the function is missing, which is already known — so the run is deliberately deferred to the combined green checkpoint rather than spent proving a tautology. That makes the NEGATIVE PROOFS the rows that carry real signal. Every positive row passes even for a naive implementation that greps /model\s*=/ over the whole file. Six rows fail it: light-tier service_tier/model_verbosity decoupling (#774), hand-added keys, a commented pin, the model_verbosity key-prefix collision, the runtime no-op ordering, and the headline case — a literal `model = "sonnet"` inside the developer_instructions ''' block, which the emitter fills with agent prompts that discuss models constantly. Row 14's fixture was verified to discriminate before being written: a whole-file scan matches it and a header-slice scan does not. Without that check the test would pass trivially and prove nothing, which is the vacuous-test failure this epic has already hit repeatedly. Adversarial TOML fixtures are hand-authored against the real Codex shape rather than generated by generateCodexAgentToml, per #2371 — a fixture from the writer can only confirm what the writer already believed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#3242): validate codex .toml model posture, not just presence Implements ADR-2313 D6. checkCodexModelPosture is a new sibling export, not a branch inside checkAgentsInstalled — that function carries 33 upstream dependents, cyclomatic 25, and sits in two traced process flows, so it is deliberately left untouched. It imports isAnthropicFlavoredModel from model-catalog, a genuine leaf. That is what Phase 1's constant move bought: agent-install-check is documented as pure read/verify and imports only leaves, so reaching the rule through model-resolver would have dragged config-loader into it. Reads liberally, judges strictly, and never guesses. Tolerates comments, key order, whitespace, CRLF, and a BOM; anchors on full key names so model_verbosity does not satisfy a `model` probe; treats extra hand-added keys as none of its business, since the check is a predicate on the two fields the posture owns rather than a whitelist over the document. An unreadable file becomes a named violation and the loop keeps going. The scan covers only the header slice — the lines before the developer_instructions ''' marker. The emitter writes agent prompts into that block and GSD's prompts discuss models constantly, so a whole-file scan reports violations for prose. This is the highest-risk defect in the phase and the reason its fixture was verified to discriminate before being written. The non-codex short-circuit runs before any filesystem call, so a stray .toml under another runtime is never inspected. Wired through cmdValidateAgents as an additive codex_posture key, so a violating install is visible from a command a user actually runs rather than only from a library nothing calls. Also fixes a test defect found while implementing: .gitattributes forces `* text=auto eol=lf` repo-wide, so the committed CRLF fixture was normalized to LF in the index — `git ls-files --eol` reported `i/lf w/crlf`, the working copy being stale pre-normalization bytes. The CRLF row was asserting against a file that could not survive a fresh clone. CRLF is now derived at runtime, which puts it under the test's control rather than git's, instead of adding a .gitattributes exception that fights a deliberate repo-wide policy and that anyone could re-normalize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3242): document the codex posture check where a user will look Three quadrants, filed by where the reader actually arrives. How-to (recover-and-troubleshoot.md, under Install and update problems) is titled by the SYMPTOM — "If Codex agents fail to spawn with a 400 about an unsupported model" — and opens with the verbatim error string. Someone hitting this does not know the words "posture" or "ADR-2313"; they have a 400 in their terminal and will search for that. Reference (COMMANDS.md) had no `validate agents` entry at all, though sibling gsd-tools subcommands are documented. Adding user-visible output to an undocumented command and then linking to it from the new how-to would have left a dangling reference. The entry carries the violation-reason table, since the frozen POSTURE_REASON enum is the machine-readable contract a reader needs rather than the prose. Both surfaces state that presence and posture are separate verdicts — a missing agent lands in `missing`, never as a posture violation. That is a deliberate design decision and would otherwise be invisible to someone watching one command emit both. Explanation stays in ADR-2313, which already covers D6 and the liberal-parse/strict-judge boundary. Pointing at it beats duplicating it into COMMANDS.md and creating two copies to drift. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3242): close two false negatives in the posture scan Both found by an isolated reviewer and reproduced before fixing. Both made the check report clean when it was not — the worst direction for this function, since the how-to tells users an empty violations list means the install is posture-clean. A quoted TOML key was never matched. `"model" = "sonnet"` is legal TOML, but the key pattern required a bare identifier, so the pin was silently invisible. Bare, "double" and 'single' quoted forms now normalize to the same key name. The block marker was found by unanchored whole-content search and used to truncate the header. A `description` value merely containing the literal text `developer_instructions = '''` truncated the scan before a real pin, and a user who hand-reordered `model` to sit after the block — still legal TOML — was never scanned at all. Fixed by changing the strategy rather than the regex: find the block's range, anchored at line start, and scan every line OUTSIDE it. That covers both failures and is strictly more correct than truncation, while still never reading prompt prose. An unterminated block excludes the rest of the file, which fails toward a false positive — the safe direction, since misreading prose as a pin wastes a user's time while the alternative hides a real one. Also corrects two overclaims of mine. The how-to named "v1.11", a version that does not exist — package.json is 1.10.0 and unreleased — so it now describes the boundary by behavior and links the ADR. And the test matrix asserted that a naive whole-file scan "fails exactly rows 12,13,14,15,16,25"; the reviewer computed that rows 12, 13, 15 and 16 produce the correct result against that baseline too. They guard real but *different* mistakes, and the matrix now says which one each catches instead of attributing them all to the header-slice defect. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3242): skip symlinked agent files instead of following them Security review finding. The scan listed entries with readdirSync and read them with readFileSync, which follows symlinks — so a symlink in the agents directory pointing anywhere would have its contents read, and any line matching the model pattern echoed into cmdValidateAgents' output through the `value` field. A read-and-echo primitive on an arbitrary path. It needs write access to the agents directory, so it crosses no new trust boundary today. Fixed anyway, for two reasons. This repo already does it correctly next door: cmdEffortSync filters with lstatSync().isFile() and the comment "Skip symlinks — only write regular files to avoid clobbering symlink targets." Being inconsistent with a sibling in the same subsystem IS the defect. And Phase 3 (#3243) extends that same cmdEffortSync to WRITE these files. Establishing symlink-following as the house pattern for Codex .toml handling here would hand Phase 3 a worse starting point while it writes rather than reads. Skipped silently rather than reported, matching the sibling: a symlinked agent file is a structural install choice, which checkAgentsInstalled owns, not a model-content posture defect. An lstat that itself throws excludes the file rather than crashing the scan. That does narrow the guarantee slightly, so the how-to now says an empty list means every REGULAR .toml is clean, and tells anyone symlinking their configs to check the targets by hand. Claiming a clean bill of health over files the check declined to open would be the same kind of false confidence the two false negatives above produced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3242): backfill changeset pr number (#3290) --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b901d1e06f |
feat(#1953): complexity-triggered refactor extension point (execute:post) (#3261)
* test(#1953): failing-first suite for the complexity-triggered refactor hook 60 behavioral cases against src/complexity-trigger.cts, which does not exist yet: decision-point counting, the comment/literal stripping leak surface, threshold and jump-delta boundaries at limit-1/limit/limit+1, stable-anchor baseline semantics, and fs fault injection via mock.method. Two fast-check properties assert that stripping never manufactures a decision point and that comments and string literals are score-neutral. Also registers the refactor-trigger capability manifest (inert until refactor.trigger_enabled) and regenerates the capability registry and matrix. Verified RED on the remote runner before any implementation exists. * feat(#1953): complexity-triggered refactor extension point Adds the opt-in refactor-trigger capability. After a phase executes, an execute:post step measures per-function complexity for the files the phase touched and writes a scoped refactor proposal when a function crosses the configured threshold or drifts past its recorded anchor. Design notes worth carrying: - The signal is computed in-core (decision-point counting over comment- and literal-stripped source, Node builtins only) rather than via Memtrace or a shelled-out analyzer. The hook fires as a deterministic CLI, not an agent with MCP tools, and core takes no external dependencies — this is the only option a behavioral test can bind to. The metric sits behind a seam. - The baseline is a stable anchor, not a rolling value: set on first observation, moved only on disposition. A rolling baseline makes the delta the single-phase change, so a function creeping +2 per phase never trips a delta of 5 and the jump check adds nothing over the absolute threshold. - Strict mode records an open deviation window in the broken-windows ledger rather than declaring its own ship:pre gate. ship.md has no generic ship:pre gate dispatch — only two hardcoded branches — so a third gate of any kind would be declared and never evaluated. - The gate clears on the proposal being dispositioned, never on the score improving. A blocking complexity number is one an executor can satisfy by splitting a coherent function in two. execute-phase.md gains a generic execute:post step-dispatch contract; it previously matched only ref.skill == "code-review", so any other step registered there was declared and never run. The code-review branch is unchanged. Full rationale in ADR-1953. Closes #1953 * fix(#1953): close git option injection and symlink escape in the refactor hook Three findings from the isolated security review, all fixed inline. HIGH — changedFilesSince interpolated the --since value into a revision token placed before the -- separator. A -- only stops PATHSPEC parsing of arguments after it; git still option-parses what comes before. So --since '--output=/tmp/x' became --output=/tmp/x..HEAD, which git accepts as --output=<file> and uses to redirect diff output — an arbitrary write. Fixed with --end-of-options before the revision range plus a conservative ref validator. The validator deliberately permits ~ ^ @ { } because those are legitimate git REVISION syntax (HEAD~1, main@{yesterday}) as distinct from ref-NAME syntax; --end-of-options is the actual barrier. The doc comment asserting the trailing -- was sufficient was wrong and is corrected. MEDIUM — resolveConfinedPath confined by string prefix only, so a symlink committed inside the repo passed the check (its own path is under cwd) and readFileSync then followed it outside the root. Now lstat-checks for a regular file and skips anything else with REFACTOR_FILE_UNREADABLE, so one bad path skips one file and the run continues. LOW — the new execute:post dispatch contract showed the gsd_run example before the rule requiring ref.command be validated first. That prose is executed by an agent, so textual order is execution order. Reordered. Refs #1953 * fix(#1953): make the analyzer able to see TypeScript at all Found by running the shipped analyzer over its own source: it reported functions=1 for a 940-line module with 24 function forms. A return-type annotation or a generic parameter list made a function invisible — `function f(a): number {}` and `function f<T>(a: T): T {}` both detected as zero. Since gsd-core is written in .cts and the capability declares .ts/.cts/.mts analyzable, the feature silently found nothing in this repo's own primary language while reporting success. A safety net that reports "all clear" because it cannot see is worse than no safety net. All 98 tests passed over this, because every fixture was plain JS — the exact failure the test matrix's own "assert against the shape production uses" warning describes. Adds a TypeScript-shapes suite covering return types (including unions, generics, object literals and type predicates), generic parameter lists (constrained and defaulted), export/async/generator combinations, annotated arrows, class-method modifiers, and optional/ default/rest params — plus the two traps: an overload signature has no body and must not count, and `a < b && c > d` is a comparison, not a generic. Detection now reports 24/37/21 functions for the three source files, which matches a hand count exactly. Also from review: - The strict-mode ledger dedup identified entries by parsing a prose description string. That is banned by CONTRIBUTING's raw-text-matching rule and was a real bug: the "exactly one window per untriaged proposal" guarantee rested on prose matching, so rewording a description or editing WINDOWS.md by hand silently produced duplicates. Now matches structurally on kind + phase + file + line. - A property test asserted on the stripper's output text. Reframed to assert the same invariant through analyzeSource's score. - nextBaseline's `candidates` parameter has been dead since the anchor change; removed from the signature and all call sites. - Extracted the duplicated require-or-degrade and capability-check boilerplate. - ADR-1953's Implementation bullet still named a `refactor.ship-gate` in check-command-router.cts — a leftover from the design cut D6 rejects. That file is untouched and no such gate exists. Removed. Refs #1953 * fix(#1953): keep execute-phase.md under its byte ceiling; un-vacuum the large-file test Five of the seven remote-runner failures were one cause: the execute:post dispatch contract, written out inline, grew execute-phase.md 1876 bytes (93,400 -> 95,276) against a frozen PRE_PHASE6 ceiling of 93,600. A drift-ack does not clear that — tests/phase6-capstone-conformance.test.cjs and tests/fix-2285-claude-orchestration-wiring.test.cjs assert the file is literally under the cap. The contract now lives in gsd-core/references/loop-hook-dispatch.md, which already claimed to be the point-agnostic dispatch reference and already documented ref.skill and ref.agent. It gains the ref.command shape, its in-context validation rule, the advisory-by-construction statement, and a note that a point whose workflow hand-rolls one kind is not implementing this contract. execute-phase.md now defers to it in one line: 145 bytes of growth, 55 B of headroom under the cap. Better placement than the first cut — the reference was overstating its coverage, and this makes the claim true rather than duplicating prose next to it. Acknowledged by appending to tests/emitted-drift-acks/2930-*.json rather than a new 1953-*.json: two ack sources may never name the same path, and that fragment is already the accumulating ack for this file. Sixth and seventh failures: analyzesLargeFileWithinBounds tripped its own vacuity guard — the fixture generated ~480 KB against a `> 500000` assert, so the guard fired and the three assertions after it never ran. The test has been vacuous since it was written. The matrix row specifies ~1 MB, so N goes 8000 -> 20000 (1.17 MB, 17% margin) and the guard to > 1_000_000. Verified by reproducing the exact body against the compiled module: 1168888 bytes, 118 ms, all four assertions hold. Refs #1953 * fix(#1953): fold the execute:post step deferral into the existing resolve line The remaining two failures were one test: execute-phase.md carries a SECOND, tighter assertion than the 93,600 ceiling — `<=93400`, which is exactly its current size. The file cannot grow by a single byte. My previous fix got it under 93,600 but not under 93,400, so it still failed. ("H." in the report is just the parent describe of that same test, not a separate defect.) Rather than add a paragraph, the deferral now REPLACES the existing hook resolution line. It read: Resolve active step hooks from `EXECUTE_POST_HOOKS_JSON` where `kind == "step"` and `ref.skill == "code-review"`. which is the bug itself written down — only code-review was ever dispatched. It now reads: Dispatch each `kind == "step"` hook per @gsd-core/references/loop-hook-dispatch.md. For `code-review`: The following prose already begins "If no active code-review step hook exists", so it reads correctly and the code-review handling is untouched. Net effect on the file is -11 bytes: 93,400 -> 93,389, under the margin assertion rather than merely under the ceiling. That also removes the need for a drift-ack: the file shrank, so there is no growth to acknowledge, and the append to the shared 2930-*.json fragment is reverted. Leaving it would have shipped a claim of "145 bytes of growth" that is no longer true, on a file six other issues share. The test's own comment states the principle this ended up honoring: "the host loop must stay small — optional-feature detail belongs in the capability fragment, not the host workflow." Putting the dispatch contract in the reference rather than inline is that rule, applied. Refs #1953 * fix(#1953): keep the code-review hook literal the workflow test requires tests/code-review.test.cjs extracts the <step name="code_review_gate"> block and asserts it contains `ref.skill == "code-review"` verbatim. The previous commit replaced the line carrying that literal, so the token vanished and the test went red — a fair assertion: code-review IS the bespoke branch there and the workflow should still name it. Restored inside the same one-line deferral, which now reads: Dispatch `kind == "step"` hooks per @gsd-core/references/loop-hook-dispatch.md. `ref.skill == "code-review"`: 93,396 bytes — still under the `<=93400` margin assertion and 4 bytes below the base, so the file continues to shrink rather than grow. Because three consecutive runs were each reddened by a different assertion on this one file, this change was verified by sweeping ALL of them at once rather than one run at a time: every test under tests/ that reads execute-phase.md or references/loop-hook-dispatch.md was located by resolving its path constants, and each content/size assertion was evaluated directly against the working tree — 22 assertions, plus two real executions (gen-section-manifest --check, and emitted-attribution's full real-tree differential). All pass. That sweep also confirms the earlier judgement call: the net change to execute-phase.md is a SHRINK, and the size ratchet only gates growth, so reverting the append to the shared 2930-*.json ack fragment was correct — an ack would have been both unnecessary and factually wrong. Refs #1953 * chore(#1953): backfill changeset pr number to 3261 * docs(#1953): add the missing how-to for acting on a refactor proposal Reference and explanation shipped (COMMANDS.md, CONFIGURATION.md, FEATURES.md 159, ADR-1953) but the Diataxis how-to quadrant did not, and that is the one a user reaches for. CONTRIBUTING's required-docs table is 'new command -> COMMANDS.md + FEATURES.md', so CI was green on a gap. Enabling this feature is genuinely multi-step and no single page walked it: turn it on, tune the threshold, understand advisory vs strict, discover that strict needs a SECOND toggle on a DIFFERENT capability, and know what to do when a proposal appears. The two-toggle subtlety in particular was a footnote in a config table; here it is a section with both commands. Follows the shape of its closest siblings, resolve-edge-coverage-findings and resolve-prohibition-findings — both 'the loop surfaced a finding, here is what to do with it'. Includes a reason-code table for the silent cases, since the analyzer is deliberately quiet in six situations and a user who expected a proposal needs to tell 'nothing to report' from 'could not look'. Indexed from docs/README.md beside the other loop how-tos. Docs-only: exempt from the push gate, no re-verification, pass marker on 2af188b4 untouched. Refs #1953 * feat(#1953): warn when strict mode is on but nothing will actually block Closes acceptance criterion 5, which I had wrongly marked satisfied. refactor.trigger_strict records an untriaged proposal as an open deviation window, but a ship only STOPS if workflow.windows_enforce is also on — a toggle owned by the broken-windows capability that this feature neither sets nor requires. So a user could enable strict, believe ship was gated, and find out otherwise at ship time. The split itself stays: requires:["broken-windows"] would force-install the ledger on advisory users who never enable strict, and a ship:pre gate of our own would never fire because ship.md has no generic ship:pre gate dispatch. What was missing was discoverability, so that is what this fixes. `refactor evaluate` now emits a typed REFACTOR_STRICT_NOT_ENFORCING warning, naming the exact remediation command, whenever strict is on and either workflow.windows_enforce is off or broken-windows is unavailable. It fires only on a run that produced a candidate — with nothing to block on there is nothing to warn about, and warning every run would be noise. Reads workflow.windows_enforce through the same resolveConfigKey walk the router already uses for its own keys rather than a second config reader. Four tests cover the matrix: strict+enforce-off warns, strict+enforce-on does not, strict+ledger-absent warns, strict-off never warns. Also corrects a user-facing message in this same file that told the user to run `gsd-tools config-set` — the wrong form. docs/CONFIGURATION.md and the broken-windows capability both use `gsd config-set`, and gsd-tools is invoked as `node gsd-tools.cjs`, so the bare form may not resolve. The two adjacent messages in this file now agree. Refs #1953 --------- Co-authored-by: sim <sim@local> |
||
|
|
86bebcefa2 |
refactor(#3216): bind milestone identity to the canonical locator (#3226)
* refactor(#3216): widen milestone-window guard to literal-## matchers The guard keyed only on the `#{N,M}` quantifier plus a literal version or phase-lookahead token. getMilestoneInfo hand-rolls its milestone-heading match with a literal `^##`/`## ` and an interpolated ${escapedVer}, so it satisfied neither token and the guard reported a clean zero on a file carrying live re-derivations (#3171, #3197) — a zero it did not earn. Widen token (a) to a literal 2-6 `#` run, admitted ONLY inside a heading-MATCHER literal (a regex literal, or a string/template handed to new RegExp) so a heading-BUILDING template is not mistaken for a re-derivation. Widen token (b) with the grouped `v(\d+(?:\.\d+)+)` shape and an interpolated version placeholder. Ships BEFORE the consolidation per ADR-3180 s7.2: a guard widened afterwards measures an already-cleaned surface. It is expected to be RED until the consolidation lands. * test(#3216): failing-first milestone-identity single-owner suite 63 tests across two files, from the matrix in .gsd/phase/. Section H of milestone-window-single-owner.test.cjs covers the 21 input classes of the design's behavior table plus its negative space; milestone-window-drift-guard covers the widened tokens and proves the exemption is function-scoped, not file-scoped. Copy count is 3 found by the guard, not 1 per the epic (ADR-3180 Amendment 3's standing rule, holding for the fourth consecutive phase): both getMilestoneInfo sites plus cmdRoadmapAnalyze's milestone enumeration at roadmap.cts:454, which carries the same #3171 truncation and #3197 phase-heading confusion. Expected RED until the consolidation lands. * refactor(#3216): bind milestone identity to the canonical locator getMilestoneInfo hand-rolled two milestone-heading regexes inside the owner's own file. Both were wrong, differently: the STATE-version site's ^## anchor is level-blind so [^\n]* absorbs a third #, and the fallback site had no anchor at all, so '## ' matched from the second # of '###'. Against '### Phase 7: Close v3.3 gaps' the fallback returned {v3.3, gaps} (#3197). Both captured names with [^\n(], truncating at a parenthetical (#3171). Bind both to the canonical grammar. locateMilestoneHeadings becomes a version-filtered view over one shared source, and a new version-agnostic listMilestoneHeadings enumerates milestone headings for callers that need all of them. getMilestoneInfo returns ScopedResult<MilestoneInfo|null>; the {v1.0,'milestone'} default, which was output-identical to a real v1.0 project, is deleted. The #2245 never-throws invariant is preserved. Copy count: 3 found by the guard, not 1 per the epic. The third was cmdRoadmapAnalyze's own milestone enumeration (roadmap.cts:454), carrying both defects in the implementation the epic blessed. buildStateFrontmatter and archivePhaseDirectories branch on scope: the first writes null rather than a fabricated identity, the second falls through to its dated-label fallback. A fabricated v3.3 passes ARCHIVE_VERSION_LABEL_RE, so it would otherwise misfile phase history. Also fixes an unsafe cast in init.cts that masked these type errors across five call sites, which would have shipped undefined milestone fields under green tsc. * fix(#3216): restore the #1761 unbounded guard and bullet precedence Review and the first full-matrix run surfaced five real defects in the consolidation, all fixed here rather than by relaxing the tests that caught them: - buildStateFrontmatter gated its isMilestoneBoundedInRoadmap check on the scope-gated milestone value, which is null on any non-COMPLETE scope, so the #1761 unbounded guard was silently skipped and state json reported a percent it must omit. It now gates on the STATE-asserted version, independent of identity scope. - The rewrite lost #2135's precedence: the name-bearing progress-marker bullet is consulted before the heading again. - A single-segment version (v3, no dot) did not resolve; the name-extraction fallback now accepts it. - A version carrying regex metacharacters, or a $& / $1 replacement pattern, is matched literally. - listMilestoneHeadings' heading field trimmed, so a CRLF roadmap no longer leaks a trailing carriage return into roadmap analyze's output. Also emits milestone_version / milestone_name / current_milestone as explicit null rather than omitting the key, so the prompt layer cannot render a bare placeholder, and corrects an init.cts comment plus a cast left inconsistent. * test(#3216): update milestone-identity expectations to the scoped contract getMilestoneInfo returns ScopedResult<MilestoneInfo|null> and the {v1.0,'milestone'} default is deleted, so the suites asserting the old shape assert removed behavior. Updated rather than weakened: every touched call site now asserts the scope explicitly against the frozen SCOPE enum. roadmap-parser.test.cjs: 20 expectations moved to {value,scope}. The #1881 unreadable-vs-absent diagnostic assertions are untouched and still prove their original point — only the return shape moved. One pre-existing assert.ok(info) is now a specific UNSCOPED assertion, so that case is stronger than before. new-milestone-clear-phases.test.cjs: the test asserting phases clear archives under the v1.0 default now asserts the dated archived-<YYYYMMDD> fallback, which is the deliberate consequence of deleting that default. Two of this branch's own tests were also corrected after they drove the implementation the wrong way: the parity test compared raw heading text and so pushed a stray ## prefix into roadmap analyze's public output, and the hostile metacharacter row demanded a pathological version resolve, which pushed a widening of the ADR-locked \b boundary. Both now assert what the contract actually requires. * docs(#3216): document milestone identity and correct the CONTEXT.md entry ADR-3180 s7.2 moves to Enforced and gains two rules that were unstated: the name derives from the heading's own version token and drops a trailing status marker, and a free-form legacy ROADMAP with no version anywhere is UNSCOPED with no identity rather than a defaulted v1.0 (decided by the maintainer before implementation, per s7's own rule that an unstated behavior is not decided). Amendment 4 records Phase 6's validation, including that the copy count was a lower bound for the fourth consecutive phase. CONTEXT.md's Roadmap Parser entry described locateMilestoneHeadings as boundary-matched with (?![\w.-]) — the alternative Amendment 2 tried and REVERTED. The code uses \b and says so, and the ADR agrees; the revert updated code and ADR and missed CONTEXT.md, which is the epic's own fixed-on-one-copy failure class in the docs layer, on a file that is itself a PR gate. * fix(#3216): persist the real version on a truncated identity buildStateFrontmatter wrote null for BOTH milestone and milestone_name on any non-COMPLETE scope, discarding a real version. ADR-3180 s7.2 rule 6: a version known with no resolvable name is TRUNCATED carrying {version, name: null} — 'the version is a real answer, the name is a non-answer, and collapsing the two is the failure this contract exists to prevent.' The two fields are now gated by what is actually known: the version whenever one exists (COMPLETE or TRUNCATED), the name only on COMPLETE. Never fabricated. Caught by this phase's own Decision 4(c) consumer-output test, which is the argument for asserting at the consumer rather than the owner — the owner was correct throughout; only the consumer collapsed its answer. * refactor(#3216): extract helpers and make cmdCommit's scope gate explicit From the two-axis code review: - init.cts repeated the identical getMilestoneInfo cast at five sites with copy-pasted comments — duplication inside a PR whose thesis is that duplicates get deleted. Extracted milestoneRecord(cwd); the one site-specific comment is kept, the four generic copies removed. - getMilestoneInfo hand-built its { value, scope } literal at ten return points; a local scoped() constructor now does it once. Every per-branch rationale comment is preserved and no returned value or scope changed. - cmdCommit gated the milestone branch name on plain truthiness, which is also true for TRUNCATED, so an unresolved identity drove branch creation incidentally rather than deliberately. It now gates on the SCOPE enum, accepting COMPLETE or TRUNCATED because both carry a real version, and the comment records why that differs from archivePhaseDirectories — which demands COMPLETE because it uses the value as a filesystem path component. * test(#3216): cover the bare-version-in-prose truncated path The spec review found the bareVersionMatch path — no STATE version, no milestone heading, a version token only in prose — returning TRUNCATED with no test exercising that exact shape, violating Decision 4's boundary-coverage requirement. * docs(#3216): record the missed Tier-2 surfaces and rule 5's corollary Decision 3 requires an explicit call-out for EVERY Tier-2 change, and Amendment 4's first draft named eight surfaces while the change touched thirteen. Adds cmdCommit's branch-name construction and the four init JSON bundles, an incomplete list being the same defect in miniature that this epic removes. s7.2 rule 5 gains a corollary separating two cases the original wording ran together: no version token ANYWHERE is UNSCOPED, while a bare version token in prose or a non-milestone heading is weak but real evidence and yields TRUNCATED under rule 6. * chore(#3216): set changeset fragment pr to 3226 --------- Co-authored-by: sim <sim@local> |
||
|
|
636ec92107 |
refactor(#3185): phase enumeration has one owner and a decidable scope (#3222)
* test(#3185): failing-first phase-enumeration single-owner suite Covers the enumeration rows with direct code evidence: 999.* backlog dirs listed by progress/stats, the phase-0 sentinel divergence, the #1324 letter-prefixed-decimal negative space, and the destructive-path find — cmdPhasesClear carries a fifth sentinel copy (/^999(?:\.|$)/) that excludes 999 but not 0, so a 0-* directory roadmap.analyze preserves is deleted there. Also covers the pass-all degrade, which is where the defect actually lives: when the milestone window declares no phases the filter becomes a literal () => true and its heading-side sentinel exclusion is unreachable. A fixture carrying phase headings keeps the filter active and never reaches that path. Named for the derivation, not a module: the suite drives commands, phase, milestone, workstream-inventory and state, and both the phase and phase-locator buckets are already at the per-module test-file cap. Committed alone so the remote runner records the failure before the fix. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): phase enumeration has one owner and a decidable scope Adds phase-locator.cts::listMilestonePhaseDirs as the single canonical owner of "which phase directories belong to the current milestone". It applies the milestone window AND the sentinel filter and returns a ScopedResult, so a caller can tell a genuinely-empty milestone from an enumeration that could not be scoped. The sentinel test now runs against DIRECTORY NAMES and is unconditional. getMilestonePhaseFilter excludes sentinels from its ROADMAP heading set, but degrades to a literal () => true pass-all predicate when that set is empty -- at which point the heading set is never consulted and its sentinel exclusion is unreachable exactly when it is needed. That degrade is the #3167 path, and it is why stats already used the filter and still listed backlog directories. The narrowing is sentinel-only: pass-all stays over-inclusive otherwise. Sentinel copies deleted, canonical isSentinelPhaseId adopted: - cmdRoadmapAnalyze's local closure (parseInt === 0 || === 999), 2 call sites - cmdPhasesClear's /^999(?:\.|$)/ -- the DESTRUCTIVE path, which excluded 999 but not 0, so a 0-* directory roadmap.analyze preserves was deleted cmdStats also seeded rows from ROADMAP headings with no sentinel filter, so a 999 heading produced a row with no directory; that seed is filtered now. cmdPhasesList routes only its ENUMERATION. --phase lookup searches the physical set (scoping it would report an out-of-window phase as not found) and --include-archived still merges archived dirs (they are by definition from other milestones). Both exempt by documented reason, never a file allowlist. Fixed inline, found while building: isDirInMilestone could not match a #1324 letter-prefixed-decimal directory (P0.0-foundation) to its own Phase P0.0 heading, so stats reported the phase with plans: 0 while its directory held plan files. Defers to phase-id's extractPhaseToken rather than widening a fourth bespoke regex; additive, so it can only admit directories. Refs #3180. Closes #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): route the last two enumeration re-derivations workstream-inventory countRoadmapPhases counted every `Phase` heading across the whole ROADMAP -- no window, no sentinel filter -- so it counted 999.* backlog and Phase 0 and spanned every milestone the document ever had. Its own caller already resolved a currentVersion and passed it to getMilestonePhaseFilter elsewhere in the same file; this was the sibling copy that never got the fix. state.cts phaseInventoryProvider enumerated phase dirs with its own /^(\d+)-(.+)$/ convention regex and neither filter, so a rebuilt STATE.md inventory carried backlog and sentinel directories as current-milestone phases. A non-COMPLETE enumeration scope now throws to the outer catch as a real scan failure rather than reporting a confident undercount, mirroring the per-phase scanPhasePlans contract beside it. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): consolidate 23 sentinel re-derivations onto one predicate The whole-repo drift guard (ADR-3180 Decision 4a, no file allowlist) found the sentinel rule re-implemented 23 times across 8 modules, in three regex variants plus four integer-comparison forms. Most tested 999 only, so Phase 0 slipped through them while roadmap.analyze and the engine-wide convention (#1580) both treat 0 and 999 alike. That disagreement is the defect class this epic removes. All 23 now call phase-id's isSentinelPhaseId (SENTINEL_RANGES [0,999]). Sites: init recommended-actions and backlog counts, milestone phase scan, the phase-lifecycle progress table, phase.cts used-number collection and the four renumber-on-remove guards, roadmap-parser's heading and bullet milestone counts, roadmap get-phase fallbacks, and state's heading denominator. Excluding Phase 0 at these sites is a deliberate behavior change and the point of the consolidation — several carried comments already saying 0 should be excluded while the literal beside them caught only 999. Adds scripts/lint-phase-enumeration-drift.cjs, wired into lint:ci. It scans the whole src/ tree with no file allowlist and reports both shapes: an independent phases-dir enumeration, and an independent sentinel literal. Exemptions are function-scoped with a written reason. The guard is comment-aware — its first pass flagged JSDoc and a comment documenting that the code below uses the canonical owner, which would have trained readers to exempt prose. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): resolve every phases-dir enumeration; drift guard reports zero Per-site triage of the 31 remaining whole-repo guard hits, applying the rule generalized from #3183's Amendment 1: a LOOKUP, DIAGNOSTIC, ARCHIVAL or MUTATION pass wants the physical set; only "which phases belong to this milestone" wants the scoped set. Routed (10): init new-milestone phase_dir_count, init milestone-op fallback count, init manager, init progress, milestone complete stats/dry-run/archive move, phase complete's next-phase scan, state update-progress, state frontmatter stats, and uat audit's active set. Exempt with a written function-scoped reason (never a file allowlist): the audit/UAT/verification sweeps that deliberately scan every directory to report gaps, phase create/insert/rename/renumber mutations, single-phase lookups, roadmap-upgrade's cross-milestone migration, cmdPhasesClear's whole-tree destructive pass, and the reads that list a phase dir's FILES rather than enumerating the phases dir at all. Latent defects fixed by the routing: sentinel directories leaked into cmdInitNewMilestone's phase_dir_count, cmdMilestoneComplete's stats, dry-run AND ARCHIVE MOVE, cmdStateUpdateProgress, buildStateFrontmatter and cmdAuditUat's active set — every one of those hand-rolled an isDirInMilestone filter with no sentinel exclusion, so `milestone complete` was archiving backlog directories. scripts/lint-phase-enumeration-drift.cjs now reports 0 re-derivations and npm run lint:ci is green. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * docs(#3185): document milestone-scoped enumeration and record ADR Amendment 3 Changeset fragment (Changed), CLI-TOOLS/COMMANDS/USER-GUIDE updates for the scoped output of progress, stats, phases list, phases clear and milestone complete, the CONTEXT.md Phase Locator glossary entry naming listMilestonePhaseDirs, and ADR-3180 Amendment 3. Amendment 3 records: the SCOPE contract held unchanged; the declared deviation from Decision 1's provisional signature (the window needs cwd/ws, which the locked roadmapContent parameter cannot supply); the copy count being a lower bound for the third consecutive phase (4 scoped vs 54 found); the load-bearing finding that the sentinel exclusion sat on the heading set and was unreachable under the pass-all degrade; the two destructive-path defects; and the generalized exemption rule. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): wire scope to consumers; revert two wrong routings the suite caught Review + remote runner findings, all fixed: The three consumers computed the enumeration scope and threw it away, so TRUNCATED/UNSCOPED/UNREADABLE collapsed into the same output as COMPLETE -- reproducing this epic's own output-identical-failure defect one layer up. progress, stats and phases list now emit phase_scope (null on the phases list --phase lookup path, which performs no enumeration). Two routings were wrong and the suite proved it: roadmap-parser's two milestone phase-count scans are reverted to the 999-only literal. isSentinelPhaseId is BROADER than what it replaced: its legacy branch runs /^0*(\d+)/ over "00.1", which backtracks to capture 0, so it read #2554's decimal phase ids as sentinel milestone 0 and stopped counting them. state.cts phaseInventoryProvider is reverted to the physical disk scan. `state rebuild` is a RECONCILIATION pass -- scoping it made it throw on healthy trees whose fixture resolves no window, swallowed the raw readdirSync fault message #3057 B1 requires verbatim, and stopped it dropping orphan STATE.md rows, which is the job. Both are now function-scoped guard exemptions with written reasons, not silent reverts. This is the consolidation trap named in the epic: a canonical rule can cover MORE than the copy it replaces, and only real inputs show it. Adds phases list coverage, a scope-branch test, and a drift-guard unit suite; backports comment-awareness to the milestone-window and plan-count guards so all three siblings share one false-positive profile; names #3161 alongside #3167 in Amendment 3's subsumption record. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): correct isSentinelPhaseId's decimal-zero misclassification An isolated security review caught this branch committing the epic's own sin: the over-broad predicate was worked around at ONE call site and left live at the destructive ones. isSentinelPhaseId's legacy branch ran /^0*(\d+)/, which backtracks so any id whose leading digit run is all zeros before a non-digit captures 0 -- "0.1", "00.1" and "0.2554" all read as sentinel milestone 0. Two pinned contracts disagree with that: #2554 requires "00.1" to be counted as a real phase, and the 999 icebox is a whole reserved milestone so "999.1" must stay sentinel. The rule is asymmetric and now says so explicitly: 999 is sentinel with or without a decimal part; 0 is sentinel only when bare. A decimal phase under either is a real phase for 0 and reserved for 999, because 999 reserves a MILESTONE while 0 reserves a PHASE. Fixing the owner lets the earlier workaround go: getMilestonePhaseFilter's two scans route through isSentinelPhaseId again and the guard exemption that existed only to accommodate the defect is deleted. The state.cts cmdStateRebuild exemption stays -- that one is a genuine reconciliation-wants-the-physical-set case. Also corrects tests/adr-612-bracket-grammar.test.cjs, which asserted isSentinelPhaseId('0.1') === true and so had encoded the defect as expected behavior. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): keep isSentinelPhaseId's semantics — 0.x is layered, not wrong Reverts the previous commit. The remote suite failed six tests proving it wrong, and the reason is the sharpest finding of this phase. An isolated security review observed that isSentinelPhaseId reads 0.1 and 00.1 as sentinel milestone 0 and judged that a defect against #2554. Correcting the canonical predicate broke #2949. Both contracts are pinned and both are right, because they ask different questions: #2554 is this dir part of the current milestone's phase SET? -> count 00.1 #2949 must this phase COMPLETE before the milestone closes? -> 0.x sentinel No single global predicate answers both. isSentinelPhaseId keeps its semantics (0.x IS a sentinel, #2949), and the milestone-window layer keeps a narrower 999-only rule (#2554) as a function-scoped guard exemption with a written reason — not a second silent copy. That corrects how Decision 1 reads: "one owner per derivation" governs who computes an answer, not how many questions share it. An over-broad canonical rule is as much a defect as a divergent copy and fails worse, because it looks like consolidation. Recorded in Amendment 3 as the lesson for Phases 4 and 5. Where a review's inference about intent conflicts with a pinned contract, the pinned contract wins; the finding is adjudicated, not fixed. The boundary tables in the enumeration suite are corrected to assert 0.x IS a sentinel, with the layering explained. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * chore(#3185): set changeset fragment pr to 3222 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
342590c70e |
refactor(#3184): milestone windowing has one owner and a decidable failure signal (#3209)
* test(#3184): failing-first milestone-window single-owner suite Covers the 50 input classes in the phase test matrix: scope classification (genuinely-empty vs truncated vs unscoped vs unreadable), the section-end owner's level boundaries, consumer-output identity per ADR-3180 Decision 4(c), the milestone.complete refusal with negative proof that no directory moved, the version-token boundary defect, drift-guard behavior, and three fast-check properties over document-shaped generators. Committed alone so the remote runner records the failure before the fix lands. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * refactor(#3184): milestone windowing routes through one owner Three copies of the milestone section-end walk lived in roadmap-parser.cts — two distinct computeSectionEnd function nodes plus an inline third in getMilestonePhaseFilter's versionOverride branch. computeMilestoneSectionEnd is now the sole owner and the other two are deleted, not kept in sync by comment. The whole-repo drift guard found what the epic did not: state.cts held three more re-derivations of the same vocabulary — two byte-identical milestone bounding checks carrying a defect neither reported copy has (no boundary after the version token, so v2.0 matched inside v2.0.1), and a milestone-sectioning predicate. All three route through the owner now. A composition-level duplicate appeared inside this change's own first pass: getMilestonePhaseFilter and cmdMilestoneComplete each re-assembled a window out of the owner's primitives, and had already diverged on whether to skip a closed milestone heading. sliceMilestoneWindow is the one composition. Windows now carry the ADR-3180 SCOPE discriminator, so a truncated window is distinguishable from a genuinely empty milestone — those were output-identical, which is the whole failure class. roadmap analyze emits it (#3165), and milestone complete refuses to archive on anything but COMPLETE rather than pass-all moving every phase directory on disk (#3166). The pass-all degrade is preserved where its premise holds: making the filter deny-all would trade a silent over-inclusive answer for a silent under-inclusive one on the read paths that count with it. extractCurrentMilestone keeps its signature — 200+ affected symbols across 41 files and 25 process flows — and is a one-line wrapper over the scoped owner. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): fence-aware phase detection and one heading-selection owner Review fixes from the two orthogonal passes. The blocker: hasPhaseEntries matched ATX phase headings fence-aware via tokenizeHeadings but tested the #2199 bullet form against un-stripped markdown, so a fenced EXAMPLE of the bullet syntax counted as a real phase. A genuinely empty milestone then classified TRUNCATED and milestone complete refused a legitimate archive — a false positive in the destructive direction, worse than the defect this phase set out to fix. Both that path and getMilestonePhaseFilter own pre-existing bullet scan now run on stripFencedCode, since leaving one meant the owner file gave two different answers to the same question. The selection rule — locate, prefer the non-closed heading, else the first — had been written three more times inside the file whose thesis is single ownership. selectMilestoneHeading owns it; all three sites route through it. The copies were behaviorally identical, so this is de-duplication with no observable change, verified by probing that all three paths select the same heading. roadmap analyze emitting a scope no consumer read left #3165's actual symptom alive, so Route 0 in next.md now treats a non-complete scope as scan-failed rather than as a clean empty scan, and the ADR amendment no longer overstates what shipped. Also: the scope refusal moved above the archive-directory create, so a refusal leaves nothing on disk; the versionOverride comment names all four consumers; COMMANDS.md documents the new guard beside its sibling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#2658): exclude the changelog from the malformed-path scan The gate walks every emitted .md/.js/.cjs file in an installed tree and asserts none contains `.claude/.trae/rules` or `.trae/.trae/rules`. CHANGELOG.md ships into that tree, and its #2658 entry quotes both malformed paths while describing the fix that removed them — so the release note documenting the fix trips the fix's own regression test. Red on next before this branch. The installer is correct: a probe over a real --trae --local install found 621 emitted files, exactly one hit, and it was gsd-core/CHANGELOG.md. The scan scope was the defect, not the product. Excluded by exact relative path rather than by loosening the patterns or skipping all markdown — the emitted agent and command markdown is precisely what #2658 was about, so the gate stays strong everywhere it matters. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): regenerate install-tree fixtures for the shared drift scanner scripts/lib/ ships in the npm package and installer, so extracting the shared tree-walk into scripts/lib/drift-scan.cjs adds one path to every runtime's install tree. Regenerated via npm run gen:install-tree; the delta is exactly that one path per fixture. The two drift guards themselves do not ship (scripts/lint-*.cjs is excluded), so only the extracted library moves. This matches the existing scripts/lib/allowlist-ratchet.cjs precedent, which is likewise a lint-only helper carried in the shipped tree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): restore the #730 sub-milestone boundary and narrow the refusal The remote runner caught two regressions this branch introduced. Both were mine, and neither review pass found them — only running the existing suite did. The version-token boundary. I replaced locateMilestoneHeadings' \b with (?![\w.-]), reasoning that v2.0 matching inside v2.0.1 was the same defect #2562 fixed in isMilestoneShippedInRoadmap. It is not the same question. A milestone state of v8.0 legitimately selects the '## v8.0-B' sub-milestone section over a closed v8.0-A sibling (#730), and \b is what allows it while the stricter boundary forbids it — nine tests in roadmap-phase-fallback said so. Reverted to \b; the state.cts consolidation is now a straight merge with no behavior change, and the v2.0/v2.0.1 ambiguity is left exactly as it was. The ADR amendment and the design doc no longer claim otherwise. The refusal scope. I refused whenever the window was not COMPLETE, but #3166 is about the TRUNCATED window specifically — the heading is found and the section closes before the phase region, so pass-all archives everything. UNREADABLE and UNSCOPED are pre-existing, legitimately handled states, and refusing on them broke 'handles missing ROADMAP.md gracefully' and three archive tests. Narrowed to TRUNCATED; docs corrected to match. One of the new tests was also wrong: its fixture gave the shipped and current milestones' phases the same numeric id, and the filter matches on that id, so it could not have distinguished the two windows. Fixture corrected to exercise what it claims to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * fix(#3184): enumerate drift-scan.cjs for uninstall The installer copies scripts/lib/ wholesale, but uninstall removes an explicit set — deliberately, so a user's own helpers in that directory survive. The extracted drift-scan.cjs was copied in and never enumerated, so it outlived uninstall, left the directory non-empty, and the rmdir that follows failed. Added to GSD_SCRIPTS_LIB_FILES, following allowlist-ratchet.cjs, which is likewise a lint-only helper that ships there and is enumerated. Verified with a real install-then-uninstall into a temp target: scripts/lib/ held exactly the three GSD files and was gone afterwards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * test(#3184): assert install and uninstall agree on scripts/lib and scripts/changeset Found while shipping this phase, and fixed here rather than noted. install() copies scripts/lib/ and scripts/changeset/ into the target WHOLESALE — the comment at the copy site literally says "and any future lib helpers". uninstall() removes them by hardcoded enumeration, deliberately, so a user's own helpers in those directories survive. A wholesale writer paired with an enumerated remover cannot stay in sync by construction: any file added to either directory ships to every user and is then orphaned in their repo forever, since it survives uninstall, leaves the directory non-empty, and the rmdir that follows fails. Nothing reported this. 31,225 tests were green over it. That is the same divergence class this epic exists to delete, sitting in the installer, so it gets the same remedy CLAUDE.md prescribes for it: a parity assertion that fails the moment the two surfaces disagree. The test compares each directory's real contents against its enumeration and names the offending file plus the constant to add it to. Both enumerations are hoisted to module scope and exported, so the test asserts on the actual arrays rather than pattern-matching the installer's source — no allow-test-rule annotation needed. Proven non-vacuous both ways: empty diff on the current tree, correct report when an unenumerated file is injected. scripts/changeset/ turned out to carry the identical defect and is covered too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 * chore(#3184): backfill changeset PR number Also narrows the wording to match the shipped behavior: the refusal fires on a truncated window specifically, not on any non-complete scope. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015kfkRFNUESoBspUYcAQaT3 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
481ac7c71b |
fix(#2946): run milestone complete unstarted-phase guard independent of STATE (#3081)
* test(#2946): milestone complete unstarted-phase guard fails open on STATE desync Row 1 of the test matrix: the regression test that fails first. Adds seven cases to tests/milestone.test.cjs covering the desync, absent, no-file, --force-override, mismatch-WARNING, fresh-project-noop, and sentinel-skip behaviors. RED on next: the guard's entire scan is nested inside `if (stateVersion && stateVersion === version)`, so any STATE.md milestone: value that does not exactly equal the version argument skips the scan with no warning — functionally an implicit --force on a one-way-door operation. * fix(#2946): run milestone complete unstarted-phase guard independent of STATE The entire ROADMAP phase-directory scan was nested inside `if (stateVersion && stateVersion === version)`, so any STATE.md milestone: value that did not exactly string-equal the version argument — a desynced value, or no milestone: field at all — skipped the scan with no warning, functionally an implicit --force. The operation the guard fronts is a one-way door: ROADMAP.md and REQUIREMENTS.md are archived and phase directories are MOVED into .planning/milestones/<version>-phases/. The scan was already driven by the version argument through getMilestonePhaseFilter / extractCurrentMilestone; the STATE match was a redundant second gate that shadowed and broke it. Decouple: the scan now runs whenever --force is absent, and a present-but-mismatched STATE milestone: field emits a WARNING naming both values so the suspicious condition is visible rather than silent. A fresh project with no Phase headings in the scoped slice still yields an empty scan (no false positives) — the intent the STATE-match short-circuit was reaching for, now achieved by the scan itself. * docs(#2946): document milestone complete --force, --dry-run, and the unstarted-phase guard The CLI-TOOLS reference signature omitted --force and --dry-run entirely, and neither the unstarted-phase guard nor its override was documented anywhere user-facing. Add a flags table and a factual guard description to the Reference page (CLI-TOOLS.md), and a practical guard note to the /gsd-complete-milestone How-to (COMMANDS.md) covering what to do when the guard fires and the new STATE-mismatch WARNING (#2946). American English per CONTRIBUTING.md language policy. * fix(#2946): emit STATE-mismatch WARNING as JSON in --json-errors mode Follow-up to the guard decoupling: a structured caller using --json-errors parses stderr line-by-line as JSON, so the plain-text WARNING would break such a parser. Honor getJsonErrorMode() and emit a structured JSON object ({ ok, level, message }) in that mode, plain text otherwise — mirroring io.cts error()'s JSON shape. Addresses the isolated-review observation (~45% but credible, since --json-errors is a documented CLI flag). * fix(#2946): address review — drop JSON-mode WARNING scope creep, tighten test assertions Standards + spec review findings (code-review two-axis + isolated adversarial): 1. The --json-errors JSON WARNING branch (commit dee719404) was scope creep the issue never asked for, AND emitted ok:true for a suspicious-condition warning (a category error — a stderr JSON parser keying on ok would treat the suspicious state as success), AND its comment falsely claimed to mirror io.cts error()'s {ok:false,reason,message} shape. Dropped: the WARNING is now plain-text stderr only, matching the existing [gsd-tools] WARNING convention (state.cts). The issue asked for 'at minimum warn', not a structured JSON surface. 2. The WARNING test asserted on /WARNING/ regex (raw-text matching on stderr prose). Tightened to assert on the stable operator-facing tokens — the WARNING: marker and both version literals the operator must see — not the surrounding formatter prose. Added a paired negative test confirming no WARNING is emitted for an absent milestone: field (a missing declaration is a normal fresh-project state, not suspicious drift). * fix(#2946): sanitize STATE milestone value before stderr WARNING interpolation Security review (minor): stateVersion is read from a user-controlled file (STATE.md) and is not validated like the CLI version arg. Sanitize before interpolating into the WARNING — strip ANSI/control chars (/[\x00-\x1f\x7f]/g -> '?') and truncate to 80 chars — so a corrupted or hostile STATE.md cannot echo terminal escapes or secret-looking strings verbatim into a CI log or terminal aggregator (CONTRIBUTING.md security: secret-looking values in stderr). version is already constrained to [A-Za-z0-9._-] by ARCHIVE_VERSION_LABEL_RE upstream, so it needs no sanitization. * test(#2946): correct stale fixtures that relied on the guard being silently disabled Four pre-existing tests broke under the #2946 fix because their fixtures only passed thanks to the bug — the unstarted-phase guard was skipping on STATE mismatch, so fixtures with missing or non-matching phase directories slipped through. The tests exercise version-forwarding / version-scoping, not the guard, so give them legitimate directories: - milestone.test.cjs #3043: dirs were '103.old'/'104.old'/'108.new' (dot), which phaseTokenMatches rejects — renamed to hyphen form. The v3.6 stats scoping still yields 1 phase (getMilestonePhaseFilter scopes correctly); the guard now sees all three phases as having directories. - milestone-archive.test.cjs 'returns version in response data': ROADMAP listed Phase 1 but no directory was created. Added 01-foundation so the scan is satisfied. Per CONTRIBUTING.md, test-fixture corrections land as their own test: commit, not bundled into fix: (release hotfix cherry-pick routes by prefix). * chore(#2946): backfill changeset PR number 3081 --------- Co-authored-by: sim <sim@local> |
||
|
|
ffd5370464 |
fix(#2903): use the command form that actually works in reader-facing docs (#3047)
* fix(#2903): use the command form that actually works in reader-facing docs Docs told readers to type the colon form, which no runtime registers -- 18 of 19 runtimes use slash-hyphen and the 19th uses shell-var -- so anyone copying an example got an unrecognized command. Swept 178 occurrences across 53 files, locale mirrors included so they do not re-diverge from English. The colon form is a source-authoring token, not a user-facing one: install-time converters key on it to produce the hyphen form runtimes actually register. So the sweep is scoped, and three things are deliberately left alone: - ADRs, which are a historical record; editing their prose falsifies what was written at the time. - The legacy release-notes archive, pending a maintainer decision on whether it follows the same historical carve-out. Excluding it keeps a later reversal additive rather than a revert. - Source artifacts under commands, workflows and agents, where the colon form is load-bearing. Rewriting those would break the installed-skill guarantee across every runtime -- the single largest hazard here. The plugin namespace form is a real, separate token and survives untouched. Adds a lint enforcing exactly that boundary, since the correct form genuinely differs by directory and nothing previously caught the drift. Also fixes a hardcoded colon form in the capability-matrix generator. The sweep alone would have left the generated matrix disagreeing with the template that produces it, so the fix is at the source and the output regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2903): stop the sweep misquoting source frontmatter Adversarial review caught three lines where the sweep rewrote a citation of the literal YAML name: key from a source command file. That key genuinely is the colon form -- this change's own carve-out logic says source-authoring tokens keep it -- so the docs ended up misquoting the real files. One of the three is an acceptance-checklist assertion, which the sweep turned into a false statement. Restored the three citations to match their sources verbatim, surgically: where a line carried both a name: citation and a real reader-facing slash command, only the citation reverted and the command stayed corrected. The guard needed the same distinction, or it would have flagged the restoration and reddened the build: a gsd:<cmd> token preceded by name: is a citation of a source token and is now permitted. The exemption is deliberately narrow -- a bare gsd:<cmd> anywhere else still fails -- with a test pinning that narrowness. Also makes the detection case-insensitive. Review found /GSD:next slipped through silently; no such casing exists in the tree today, so this closes a latent gap rather than fixing a live one. Swept the whole tree for further corrupted citations: none beyond the three. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2903): retire the stale-next invariant and sweep next like every other command Maintainer decision on a genuine conflict between two contracts. Invariant #3054 banned the literal /gsd-next from user-facing docs because it named a retired workflow-advance command. But commands/gsd/next.md is a live command -- the state-aware smart-entry launcher -- and this issue requires docs to use the hyphen form every runtime actually registers. Both could not hold for this one command, so docs had been sidestepping the ban by keeping the colon form, which is exactly the defect this issue exists to remove. FEATURES.md already recorded the reassignment: the hyphen form "is not the retired workflow-advance command; it is reserved for the state-aware smart-entry launcher. Workflow advancement remains under /gsd-progress --next." With that reassignment the invariant's premise is obsolete and the guard now contradicts the documented command form, so it is retired with a comment recording why rather than deleted silently. next is now swept like every other command, and the earlier exemption added to the new guard is removed so nothing is special-cased. Four citations of the literal name: frontmatter key stay in colon form, because the source file really does carry name: gsd:next and a doc quoting it must reproduce it verbatim. Two of those lines were reworded to say which side is the frontmatter key and which is the slash command, since they previously conflated the two. Verified the retired scan would now genuinely fail against this tree -- the conflict was real and resolved, not dodged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2903): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
de78f2eef2 |
docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate (#3010)
* docs(#2775): align package-legitimacy docs to the ADR-0656 registry-API gate security-model.md, USER-GUIDE.md, ARCHITECTURE.md, COMMANDS.md, FEATURES.md, and gsd-planner.md's STRIDE template (+ ja-JP mirrors) described the pre-ADR-0656 design: slopcheck as the install-or-degrade gate, with unavailability degrading every package to [ASSUMED]. ADR-0656 inverted this months ago — registry-API verdicts (npm/PyPI/ crates.io) are the gate; slopcheck is an optional escalate-only adapter that no shipped configuration wires. Verified every replacement claim against src/package-legitimacy.cts (checkPackages, classifyPackage, lookupNpm/lookupPypi/lookupCrates) via Memtrace before writing it, so the corrected prose matches the live implementation rather than restating the ADR from memory. Restored docs/explanation/security-model.md:79-84 (and its ja-JP mirror) to original wording after an orthogonal spec review caught that an earlier draft had edited the "Why WebSearch packages are always [ASSUMED]" paragraph — inside the range issue #2775 explicitly named as correct and to leave alone. The ja-JP mirror was missing the closing clause present in the corrected English original ("its absence leaves registry-API verdicts intact rather than downgrading everything to [ASSUMED]") — added for parity. This completes the ja-JP mirror the issue's acceptance criteria named explicitly. zh-CN/ko-KR/pt-BR (not named by #2775, but carrying the same stale design) get the mechanical portion of the same fix: command-string swaps, table headers, ARCHITECTURE.md diagram labels, and technical- term swaps that reuse a word already attested elsewhere in the same file (合法性/적법성/legitimidade for "legitimacy") — surrounding prose untouched. The remainder in those three locales — full-paragraph rewrites of the corrected degrade-path mechanism, deleted "External dependency" bullets, and "manually install slopcheck" code blocks — needs prose composed by a fluent speaker of each language and is filed as open-gsd/gsd-core#3002 with an exact file:line inventory. * test(#2775): acknowledge gsd-planner.md byte growth from the STRIDE-row fix agents/gsd-planner.md grew 14 bytes (49309 -> 49323) from the STRIDE supply-chain row correction (slopcheck -> package-legitimacy gate). Emitted agent/workflow files are byte-tracked; this fragment acknowledges the growth per tests/emitted-attribution.test.cjs's "differential attribution over the real tree" check. * docs(#2775): close ja-JP FEATURES.md gap; fix a ko-KR transliterated heading docs/ja-JP/FEATURES.md:2808 still read the katakana transliteration "スロップチェック verdict" in REQ-PKG-GATE-01 — invisible to a literal "slopcheck" grep, so it was missed when ja-JP parity was checked and declared complete. Corrected to "正当性判定" (legitimacy verdict), matching the term already established in ja-JP/explanation/ security-model.md and ja-JP/USER-GUIDE.md. This was the only remaining ja-JP gap; a full sweep for the transliterated form across docs/ja-JP/ now returns zero hits, and the ja-JP mirror is genuinely at parity. docs/ko-KR/USER-GUIDE.md:398's heading "슬롭체크 판정:" had the same transliteration problem. Fixed inline to "적법성 판정:", reusing the 적법성/legitimacy word already attested two lines below in the same table. A parallel sweep of zh-CN and pt-BR found no transliterated forms of "slopcheck" in either locale. The remaining transliterated occurrence in ko-KR (USER-GUIDE.md:406, the lead-in to the pip-install code block) needs prose composition like the rest of that block and is added to open-gsd/gsd-core#3002's inventory. * chore(#2775): backfill changeset PR number to 3010 --------- Co-authored-by: sim <sim@local> |
||
|
|
fc3fde05ee |
feat(#2646): surface unresolved deferred-items.md at milestone close (#2983)
* feat(#2646): surface unresolved deferred-items.md at milestone close auditOpenArtifacts scanned eight categories; deferred-items.md was not among them. #2287 made that file readable at the PHASE boundary (audit-uat, /gsd-progress check 7), but one boundary up it stayed invisible — and phase directories archive to milestones/vX.Y-phases/ by default (#1871), so an out-of-scope discovery a phase agent correctly recorded rather than fixed left the live tree at milestone close having never reached the [R]/[A]/[C] prompt that exists to catch exactly this. Adds deferred_items as a ninth scanner plus its count, its items entry and its report section. The workflow needed no change: complete-milestone branches on "any section with count > 0", so the new category flows through the existing prompt. The resolved/unresolved predicate is NOT reimplemented. uat.cjs already exports parseDeferredItems, which owns the parsing rule (entries under a `## Deferred Items` level-2 heading, else the whole file fail-safe; RESOLVED only on an explicit case-insensitive `status: resolved` field). The scanner requires it lazily, inside the scan, so audit-command-router's property that a route never loads the module it does not need is preserved. Two readers of one file sharing one predicate is the point — duplicating the inequality is how they drift into disagreeing about what "open" means. Regression test proves fail-first: 9 of its 10 cases go red against the pre-change tree. The tenth is the deliberate no-regression boundary (a clean tree emits no section) and is green both ways. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TQU48ETJjEmGLJjA6hdQ4 * docs(#2646): document the pre-close artifact audit and its nine categories The /gsd-complete-milestone entry did not mention the audit at all, so the gate that can stop a close was undocumented — and this change adds a category to it. Tabulates all nine with their source artifact and what makes each one "open", plus the [R]/[A]/[C] outcomes. Also disambiguates the one genuinely confusing thing: the per-phase deferred-items.md scanned here is NOT the `## Deferred Items` section the [A] path writes into STATE.md. Same name, different artifact, opposite ends of the flow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TQU48ETJjEmGLJjA6hdQ4 * chore(#2646): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TQU48ETJjEmGLJjA6hdQ4 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
7372d99a26 |
enhance(#2800): derive reviewer flag lists and gate reviewer lane docs across locales (#2882)
* chore(#2800): derive reviewer flag lists and gate reviewer lane docs across locales The reviewer lane roster was hand-enumerated across five documentation surfaces and three workflow files that had drifted apart: --kimi-code was missing from all four translated COMMANDS.md mirrors, --coderabbit from every workflow forwarding list, and --antigravity from FEATURES.md. Adds checkReviewerDocsParity, a second pure gate deliberately separate from checkReviewerLaneParity so a stale doc cannot make the runtime checker look red. Workflows now derive their flag lists from a new review-lane flags query instead of hand-enumerating them, which also retires the unanchored grep that matched --agy inside --antigravity. Documents the previously absent reviewer body and hostBehaviors field in the capability manifest reference. Closes #2800 Closes #2781 Closes #2272 * fix(#2800): key the docs parity table arm on first-cell position Review found the flag arm was file-scoped, so the forwarding row that lists every flag in its third cell satisfied it on its own. Deleting a lane's own reviewer-table row -- the #2781 regression this gate exists to prevent -- therefore passed undetected. Arm 4 keys on the FIRST table cell, which separates a lane row from the forwarding row structurally and in every locale. Regression test included. * fix(#2800): shape-filter the flags subcommand output All three consumers read review-lane flags through an unquoted command substitution so the output word-splits into loop items. Phase 2 admits third-party overlay lanes, so an overlay flag containing whitespace would inject a second loop item and one containing a glob would expand against the cwd. Emit only well-formed flags so neither reaches the shell. * fix(#2800): remove the regex length ceiling and count only prose mentions Review found two real defects in the docs parity gate. The never-throws contract was false: building a RegExp from a declared flag or section title throws SyntaxError past ~100k chars, and Phase 2 admits overlay lanes whose declared strings are untrusted in length. Every one of these matches is literal, so String.includes replaces the regex outright, which also deletes escapeLiteral and the llama.cpp escaping it existed for. Arm 1 was context-blind: a flag mentioned only inside a fenced example or a commented-out row counted as documented. Both are stripped before matching. Also advertises all 13 lane flags in the argument-hint and corrects a stale eleven-lane count in the slug grammar note. * test(#2800): repoint the convergence suite off deleted workflow text The derived flag loop deleted the literal per-flag grep lines four tests matched on. Two of those failed loudly. The behavioral and property tests failed SILENTLY instead: their end marker no longer resolved, so the parse block extracted empty and both passed vacuously, and the property test's gsd_run stub had a no-op default that hid it. All now share one extractor and execute the real deployed block through a gsd_run shim backed by the actual binary. The whitelist assertions become an anti-parity check: re-adding a hand-written flag list must fail. Also repairs two vacuous cases in the docs parity suite. The unreadable-doc test called its own mock rather than the reader, and the integration test bounded nothing, so a doc losing its marker would have been silently skipped and still passed green. * fix(#2800): run the derived flag loop after the launcher preamble The remote matrix caught a real runtime bug, not a test artifact. In autonomous.md and plan-review-convergence.md the launcher preamble that defines gsd_run lives in a separate, LATER bash fence than the derived loop. Each fence is its own shell, so gsd_run was undefined where the loop ran: the command substitution yielded nothing and zero reviewer flags would have been forwarded. Worse than the drift this epic fixes, and silent. The whole CONVERGENCE_ARGS construction moves as one unit, because the --max-cycles append sits between the loop and the preamble and would otherwise have run against an uninitialized variable and then been dropped by the relocated initializer. Also documents all 13 lane flags in help/modes/full.md, which the repo gates bidirectionally against each command's argument-hint. * test(#2800): repoint the two converge suites off deleted flag literals Both asserted workflow.includes('--codex') against the hand-enumerated list the derived loop removed. They now assert the derivation itself, keep --all and --text (convergence controls, still literal), and add an anti-parity guard so re-adding a hardcoded list fails. The lost pass-through proof is replaced with a real one: every flag the tests used to hardcode is asserted present in the actual roster emitted by the binary, which is the property the old assertion was protecting. * test(#2800): acknowledge the workflow byte growth from the derived flag loop * chore(#2800): backfill changeset pr number to 2882 * fix(#2800): strip HTML comments to a fixed point in the parity gate CodeQL js/incomplete-multi-character-sanitization (high) on PR #2882: the single-pass <!--...--> strip can leave a live <!-- behind, so a join-trick construction smuggles a commented-out row past the gate and it counts as documented. Not an injection risk here since nothing is rendered, but it is the exact false pass this helper exists to prevent. Strips to a fixed point, then treats any surviving opener as unterminated so the multi-line branch closes it on a later line. Terminates because every pass strictly shortens the string. * test(#2800): pin the comment-smuggling regression with a real reproducer The obvious fixture for this class does not reproduce it: <!--<!---->--> leaves a dangling --> rather than a live <!--, and is caught either way, so it would have passed with and without the fix. The join-trick construction (<!- + <!--DUMMY--> + -...-->), the <scr<script>ipt> shape, genuinely regresses on the single-pass strip and is what the test now uses. --------- Co-authored-by: Test <test@example.com> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6932fb16d7 |
enhance(#1699): require read-and-cite provenance for in-repo discrete values (#2768)
* test(#1699): failing-first contract tests for in-repo value provenance
Nine assertions on the deployed gsd-phase-researcher contract: the discrete-value taxonomy, the same-session Read requirement, path-AND-line-range citation, grep-alone exclusion, the verbatim quote and paraphrase ban, the quote-is-the-checkable-artifact guard, [ASSUMED] routing for unquoted skeleton values, a no-regression guard on the pre-existing package name provenance rule, and a single-definition-site guard.
Eight of the nine fail against the unmodified agent on origin/next; the ninth is the no-regression invariant and passes in both states, which is the intended enhancement shape.
Added to tests/research-agent-profiles.test.cjs rather than a new file: that file already carries the allow-test-rule exemption <runtime-contract-is-the-product> research agent .md content is the governed surface, so no new allowlist entry and no change to the per-module test-file count.
* enhance(#1699): require read-and-cite provenance for in-repo discrete values
The claim-provenance system governed external facts (npm registry, official docs, Context7, package-name provenance). For an in-repo discrete value -- an enum, schema or type union, error code, status constant, or filesystem path -- [VERIFIED] could be earned from training memory or a bare codebase grep, which proves a string occurs, not that the definition was read.
A drifted value passes into RESEARCH.md, is lifted by the planner into PLAN.md's <interfaces> context block, and is trusted by the executor, where it fails at parse()/typecheck as a mid-execution deviation -- the most expensive place to discover it.
The rule lands beside its structural sibling, the package name provenance rule, since both say existence is not verification. The verbatim quote is named as the load-bearing artifact: a citation with no quote does not earn the tag, however precise the line range looks. That keeps the rule falsifiable against the file rather than a self-report, which is the Goodhart guard.
Scope note: the quote goes in RESEARCH.md beside the claim, NOT in an <interfaces> block. <interfaces> appears zero times in this agent on next -- it is planner-side, defined at gsd-core/references/planner-interface-context.md:15 as a PLAN.md structure. The issue text and triage both said <interfaces>; instructing the researcher to populate a block it does not emit would be undefined.
Defined once at the definition site; the source-hierarchy recap is deliberately untouched, since restating it is the paraphrase-drift mode META.RULE.brief-no-paraphrase names.
* test(#1699): regenerate agent size baseline and golden parity fixtures
Regenerated via npm run size:baseline and npm run gen:golden, never hand-edited. agent-size-baseline gsd-phase-researcher.md 40866 to 42020 (LARGE tier, cap 49152, 7132 bytes headroom remaining, per ADR-1610's per-file baseline guard). 18 of 19 golden fixtures updated; pi.json is unchanged because the pi runtime ships zero agents.
* chore(#1699): add changeset for the in-repo value citation rule
User-facing behavior change in gsd-phase-researcher, so a Changed fragment is required. pr: 2768.
* test(#1699): acknowledge the gsd-phase-researcher growth for the emitted-drift gate
CI test (ubuntu-latest, 22) failed on emitted-attribution.test.cjs: gsd-phase-researcher.md grew 1154 bytes (40866 -> 42020) without an acknowledgment. The differential emitted-attribution gate landed on next in
|
||
|
|
c87f6f358e |
enhance(#1854): offer restore for user-added files backed up on update (#2679)
* test(#1854): failing-first coverage for user-files-backup restore Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(#1854): offer restore for user-added files backed up on update Adds a restore-custom-files gsd-tools verb and wires it into update.md as a restore_custom_files step: plan, compatibility-check against the newly installed release, then restore only on explicit opt-in. The backup is never deleted, a shipped path is never overwritten, and a single unwritable entry does not abort the rest. Also drops the jq pipe from update-context field extraction (#2589 class, missed by that sweep) and repairs a broken code fence in docs/CLI-TOOLS.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): reject symlinked restore destinations and backup roots Self-review of the restore path found two write-through holes: copyFileSync follows a symlinked destination, so a link planted at the restore target wrote outside the config dir with every ancestor still a real directory; and statSync on the backup root followed a link, letting the walk read arbitrary files and present them as the user's own backup. Both now lstat. Also marks the report's path/detail strings as untrusted data in update.md so the rendered step cannot carry instructions into the runtime model. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): move the update-context jq guard into the #2589 sweep update.md joins the AUDITED list rather than carrying a duplicate assertion in the backup-restore suite, and the guard gains a negative-proof companion so 'no jq pipe' cannot pass by the fields simply no longer being read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): validate manifest files map shape before trusting it Security review flagged that Object.keys on a non-plain-object files field yields numeric-index keys matching nothing, so the managed-path check dies silently while manifest_found still reports true. Shape, not just type (ADR-227): an array or scalar files map is now an unusable manifest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): size the restore prompt by eligible_count Spec review found the prompt was driven by entries.length, so a backup holding only blocked entries asked "Restore 1 file(s)?" when accepting would restore zero. The question now reads eligible_count, and an all-blocked backup reports its reasons instead of offering a choice that cannot be honored. The decline path names the resolved backup_dir rather than the bare directory name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#1854): use t.skip on hosts without symlink support A bare return in a node:test body registers as a PASS, so the four symlink guards silently reported green on unprivileged Windows instead of skipping. Adds the dangling-link destination case the security review called out, and moves outside-dir teardown to t.after so a failing assert cannot leak it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#1854): unfence restore hint, regen goldens, widen install timeout Three gate failures from the c99d612a5 run, all root-caused: 1. capability-registry (3): update.md's decline message put an instructional 'gsd-tools ...' line in an UNTAGGED fence, and the guard treats untagged fences as shell blocks. Retagged both display blocks as text and switched the hint to the resolved 'node <config-dir>/.../gsd-tools.cjs' form users can actually paste. 2. golden-install-parity (19): update.md and gsd-tools.cjs ship, so every runtime fixture moved. Regenerated; the diff is exactly those two hashes per fixture, no other drift. 3. install.test.cjs (5): one real failure, four cascades. The Cursor suite's before hook died on 'spawnSync ETIMEDOUT' at the 60s cap while the node22 lane passed the SAME commit in 12.7s. A full install measures 13-30s idle, so 60s was under 2x headroom and shrinks with every file added to the payload. Raised to 120s, matching the heavy case already in this file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#1854): backfill changeset pr number to 2679 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
920a5f3f06 |
fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep (#2673)
* fix(#2589): use --raw/--pick for config/model/verify lookups, drop jq dep The reviewer/workflow config lookups resolved scalars and object fields with a `gsd_run query <cmd> … | jq … 2>/dev/null || <default>` shape. On any machine without jq (the default on Windows/Git-Bash) the jq stage fails with exit 127, the failure is swallowed by 2>/dev/null + the trailing || default, and the variable comes back EMPTY — the configured per-lane model/host/budget is silently dropped and the lane falls back to CLI defaults with no diagnostic. gsd-tools ships native flags that do the same job with no external dep: config-get <key> --raw (strips JSON quotes off a scalar) resolve-model <id> --pick model (descends an object) resolve-execution … --pick <f> (same) verification.status … --pick status Replaced every jq-piped config/model/verify lookup across review.md (×23), plan-phase.md, ship.md, debug.md (incl. the redundant boolean coercion — --raw returns true/false as bare tokens natively), autonomous.md (×2), ai-integration-phase.md (×4), and eval-review.md. The legitimate structured-JSON jq sites that parse HTTP curl responses (.choices[0], jq -rs, jq -n --rawfile) are untouched — only the jq-replaceable lookups moved to the native flags. Adds tests/fix-2589-config-get-no-jq.test.cjs: a source-invariant guard asserting no audited workflow pipes config-get/resolve-model/resolve-execution/verification.status to jq (fails-first on the pre-fix text, passes after). * test(#2589): update autonomous-converge jq assertion to --pick; regen golden fixtures Two test consequences of the workflow-doc edits in the prior commit: 1. tests/autonomous-converge.test.cjs pinned the OLD jq-dependent shape (`verification.status … | jq -r '.status//empty'`) as the canonical routing contract. The test's INTENT is correct (route human validation through canonical verification.status) but it over-specified the MECHANISM (the jq pipe). Updated the assertion to match the new native --pick status shape; the contract being guarded (canonical verification.status read before the human_needed branch) is unchanged. 2. The golden-install-parity fixtures (19 runtimes) record a content hash of every installed workflow .md; the 7 edited workflows changed those hashes. Regenerated via `npm run gen:golden` (the test's own failure message instructs this). Only the 7 edited workflow hashes changed in each fixture. * fix(#2589): declare jq a prerequisite for the lanes that still need it; repair test file Three defects in the first cut of the #2589 fix: 1. tests/autonomous-converge.test.cjs was a JavaScript syntax error. The regex literal /...2>\/dev/null .../ left the second slash unescaped, terminating the literal early and parsing `null` as regex flags: SyntaxError: Invalid regular expression flags The whole file failed to load, so every assertion in it — including the #1522 and #1526 guards — silently stopped running. Replaced with the string-compare form already used at line 202 for the sibling shell-snippet assertion. 2. lint:ci failed. tests/fix-2589-config-get-no-jq.test.cjs buckets into the capped `config` production module via its `config-get-...` effective prefix, making it a novel offender against the 2-file cap. The test is about workflow documents, not the config module, so it is renamed to fix-2589-workflow-jq-dependency.test.cjs (free prefix) rather than growing the allowlist with a module that does not actually need a 5th test file. 3. The fix deleted the repo's only jq-prerequisite declaration. review.md:244 ("install jq if missing") was the anchor plan-review-convergence.md cites by line number, and it went away with the jq pipes — while the ollama, lm_studio, llama_cpp, opencode, and agy lanes still hard-require jq to parse HTTP /v1/chat/completions responses, opencode's JSONL event stream, and agy's conversation cache. On a jq-less host those five lanes swallow exit 127 into empty output: the same silent-degradation class #2589 exists to close. detect_clis now probes jq alongside the other prerequisites and emits jq:available / jq:missing, and the five dependent lanes are treated as undetected when it is absent, with an install hint. The six lanes that do not need jq stay selectable. plan-review-convergence.md now cites the section by name instead of a line number that moves. Regression guards added to the renamed test file: review.md must keep the jq probe and must name all five dependent lanes, and no workflow may cite review.md by line number. Workflow-size baseline and the 19 install-parity goldens regenerated for the review.md / plan-review-convergence.md edits. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * fix(#2589): decide the ship verification gate on a single verification.status read Isolated review finding (medium). Pre-fix, ship.md captured verification.status ONCE into $VERIFICATION and picked status / next_action / next_command off that cached JSON with three jq calls. --pick takes a single dot-path field, so the mechanical conversion issued three separate queries up front: three node spawns that each re-read the phase VERIFICATION.md and re-derive the commit-time vs mtime staleness comparison, on every ship — including the common passing path that never uses the two message fields. It also meant the gate's verdict and the message shown to the user were derived from three reads with no guarantee they observed the same state. The gate now reads `status` once and decides. The two message-only fields are read on the blocking path only, after PHASE_VERIFICATION_INCOMPLETE is already determined — so the passing path costs one query instead of three, and a concurrent write between reads can no longer make the gate and its message disagree, because the block/allow decision no longer depends on them. Adding a multi-field --pick to gsd-tools would have collapsed this to one query, but that changes the flag's output contract and belongs in its own change. Regression guard in tests/fix-2589-workflow-jq-dependency.test.cjs: ship.md must read verification.status exactly three times total, the block decision must follow the status read, and next_action / next_command must both appear after the blocking prose so they cannot drift back onto the passing path. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * docs(#2589): document the jq prerequisite for the five reviewer lanes that need it /gsd-review's ollama, lm_studio, llama_cpp, opencode, and agy lanes parse JSON GSD does not produce (OpenAI-compatible /v1/chat/completions responses, OpenCode's JSONL event stream, Antigravity's conversation cache), so they require jq on PATH. Nothing in docs/ said so. Records which five lanes need it, which six do not, that reading configured models/hosts/budgets no longer requires jq at all, and what /gsd-review now does when jq is absent. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015TCwhbMuY37DzRMCfzTABJ * chore(#2589): backfill changeset pr number (#2673) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
89b673e40d |
feat(#2631): planner emits estimate and plan-checker surfaces the over-budget flag (#2670)
* test(#2631): failing-first planner estimate emission and over-budget surfacing * feat(#2631): emit plan estimate and surface the over-budget split recommendation * fix(#2631): extract sizing prose to references to fit planner and plan-phase caps * fix(#2631): move estimate check to plan-checker; fix template regex and caps * fix(#2631): restore ALWAYS split literal and keep gsd_run after the launcher preamble * fix(#2631): invoke estimate-check after the launcher preamble in plan-checker * fix(#2631): stop double-applying calibration; repair COMMANDS table and stale reference * chore(#2631): backfill changeset pr to 2670 * chore(#2631): backfill changeset pr to 2670 |
||
|
|
155c08facf |
docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase (#2574)
* docs(#2197): drop --validate docs for /gsd-plan-phase and /gsd-execute-phase These two commands never parse --validate (silent no-op); the flag is real only for /gsd-quick. Remove the false flag-table rows and CLI examples across COMMANDS.md and the how-to guides (en + ja-JP/zh-CN/ ko-KR/pt-BR mirrors), and correct the manager.flags.execute example from --validate to --cross-ai (a flag execute-phase actually parses). /gsd-quick's real --validate docs are left untouched. Ref #2197 * docs(#2197): add changeset for --validate docs removal --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> |
||
|
|
c5e0371775 |
feat(#1951): reversibility tagging — gate one-way-door decisions (#2471)
* test(#1951): add failing-first tests for reversibility tagging Red phase for issue #1951 (reversibility tagging: classify decisions by undo cost, gate one-way doors behind a checkpoint:decision). Tests assert, per the issue's acceptance criteria: - discuss-phase CONTEXT.md template records a **Reversibility:** field with a rationale on captured decisions, and states it is optional - gsd-planner @-references planner-reversibility.md and stays under the 49152-char agent cap (LARGE_CAP, tests/agent-size-budget.test.cjs) - a one-way rating inserts a checkpoint:decision before the dependent task; reversible inserts none; costly is flagged but never blocks - the taxonomy defaults to reversible when unsure (checkpoint-fatigue guard) and inserting a checkpoint implies autonomous: false - docs/reference/plan-md.md documents <reversibility> as optional with all three ratings - --no-reversibility-gates parses to REVERSIBILITY_GATES=false, is injected into the planner prompt, and is advertised in the command argument-hint and help full mode (argument-hint parity) - the override suppresses the gate but still persists the rating - cmdVerifyPlanStructure accepts every rating and the absent case (additive-validator guarantee, behavioral via runGsdTools) - parity: thinking-models-planning.md #4 adopts the canonical three-level taxonomy and the binary REVERSIBLE/IRREVERSIBLE vocabulary is gone - no content loss from the planner extraction made to fit under the cap Prose-contract assertions are Red until the implementation lands. The behavioral validator assertions pass immediately — regression guards proving the validator already accepts unknown optional tags. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(#1951): reversibility tagging — gate one-way-door decisions Classify planning decisions by what undoing them would cost, and give a one-way door a human beat before the agent walks through it (issue #1951, The Pragmatic Programmer Topic 15 'Reversibility'; Bezos's one-way/two-way door framing). Acceptance criteria met: - discuss-phase records an optional reversibility rating with a rationale on <decisions> entries in the phase CONTEXT.md template. Unrated decisions are treated as reversible, so existing phases are unaffected. - a one-way rating makes gsd-planner insert a checkpoint:decision before the task that implements the decision, reusing the existing checkpoint mechanism -- no new checkpoint machinery. - reversible ratings trigger no checkpoint; costly ratings are flagged in the plan but never block. - the rating persists on the task as the optional <reversibility rating=> element. cmdVerifyPlanStructure accepts every rating and the absent case; the structural validator does not reject unknown optional tags. - --no-reversibility-gates (REVERSIBILITY_GATES=false) suppresses checkpoint insertion for intentionally-unattended runs while still recording ratings -- the override changes what stops the run, not what the plan remembers. Single taxonomy, not two: references/thinking-models-planning.md #4 already shipped a binary REVERSIBLE/IRREVERSIBLE classification and is loaded by both gsd-planner and gsd-plan-checker. It is rewritten onto the canonical three-level vocabulary and now points at planner-reversibility.md as the taxonomy owner, with a parity test that fails if the surfaces diverge (DEFECT.GENERATIVE-FIX-DIVERGENCE). agents/gsd-planner.md sat 47 chars under the 49152 LARGE_CAP, so the checkpoint DO/DON'T guidance was relocated verbatim into planner-antipatterns.md -- already @-referenced from the same section for the same topic, so the planner still loads it and nothing was dropped. A test guards the relocation against content loss. Files: gsd-core/references/planner-reversibility.md (NEW, canonical taxonomy + emission rules + anti-patterns), gsd-planner.md, plan-phase workflow/command/help (flag wiring + parity), plan-md.md schema, discuss-phase context template, CONTEXT.md glossary, INVENTORY + manifest, size baselines, install goldens, plugin skills regen, changeset. Closes #1951 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1951): address orthogonal review findings Two isolated reviewers (correctness + security), neither of which authored the change. Every finding fixed: Security — the rationale is untrusted input (ADR-1577). It originates in conversation and flows CONTEXT.md -> planner -> PLAN.md -> executor, each hop an LLM reading the previous hop's output, with no validation on the path. planner-reversibility.md and the discuss-phase template now state it is data and never instructions, and name the </reversibility> early-termination hazard explicitly -- a rationale that closes its own element injects sibling structure the executor reads as real tasks. Four tests guard it. Correctness 1 — nothing machine-enforced the feature's own promise: a task rated one-way with no preceding checkpoint:decision validated as fully clean, so a planner error silently reopened the gap this feature exists to close. cmdVerifyPlanStructure now warns on an ungated one-way rating. A warning, not an error: <reversibility> stays additive and the plan stays valid. Four tests cover ungated (warns), gated (silent), still-valid, and reversible/costly never flagged. Correctness 2 — pass-always test. The --no-reversibility-gates parse test substring-matched the whole workflow file, and plan-phase.md prose mentions both tokens in one sentence, so it passed with the bash conditional deleted: it was testing the documentation, not the parser. Now scoped to the fenced bash blocks and matched as one physical line, with a negative control confirming prose alone cannot satisfy it. Correctness 3 — costly had no itemized emission rule, only one-way did, so two agents could diverge on whether to tag costly at all. Correctness 4 — template convention break: the example ratings were bare while every sibling field uses [...] to signal substitution, inviting an LLM to copy one-way/costly forward as boilerplate. Now bracketed. Correctness 5 — latent false-green: .includes('reversible') also matches inside irreversible/irreversibility, which appear in anti-pattern prose, so a surface that dropped the real taxonomy entry would still pass. Now word-boundary matched. ADR-857 phase-6 ceiling — the first gsd-test run caught plan-phase.md 1216 bytes over its frozen 94519 ceiling (it had 49 bytes of headroom on next). The ceiling may only rise for privileged host machinery, and reversibility gating is optional-feature logic, so the wiring was slimmed to its minimum and the explanatory prose moved to the reference files the planner already loads. plan-phase.md is now 94400 bytes -- 119 under the ceiling and 70 bytes SMALLER than on next, so the host loop shrank while gaining the feature, which is what phase 6 ratchets toward. The tracer contract (tests/tracer-bullet.test.cjs) is unchanged. Lint — fixed an unnecessary non-null assertion in verify.cts and a CRLF-fragile bare \n regex in the new test (DEFECT.WINDOWS-CRLF-TEST- PORTABILITY, the #1658/#1668/#2206/#2449/#2450 class). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture must carry the common task elements The gated-one-way fixture built a checkpoint:decision task from the abbreviated skeleton in gsd-planner.md, which shows only the checkpoint-specific elements (<decision>/<context>/<resume-signal>). cmdVerifyPlanStructure requires <name> and <action> on EVERY task regardless of type, so the fixture failed validation for reasons that had nothing to do with reversibility: errors: ["Task missing <name> element", "Task 'unnamed' missing <action>"] Caught by gsd-test on 14d14a39 (2 failures, both this fixture). The canonical shape is in tests/verify.test.cjs:266 — a checkpoint task carries <name>/<files>/<action>/<verify> like any other. Fixture corrected to match. Verified behaviorally against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Not a product defect: the validator's every-task contract is intentional and pre-existing, and docs/reference/plan-md.md scopes its required-element list to type=auto/tracer only because those are the elements a planner must author, not because checkpoints are exempt from <name>. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1951): backfill changeset pr number to 2471 * fix(#1951): CodeQL incomplete-sanitization + prompt-injection scan collision Both CI failures were real defects in code this PR added, not false positives. CodeQL js/incomplete-sanitization (high), reversibility-tagging.test.cjs:46 — the namesRating helper built its regex with `rating.replace(/[-]/g, '\\-')`, which escapes the hyphen but not backslash, so the escape was incomplete. It was also unnecessary: `-` carries no special meaning outside a character class. Replaced with a complete metacharacter escape (backslash included). Word-boundary behavior verified unchanged across all three ratings — notably that "irreversible" prose still does not satisfy a "reversible" match, which is the false-green this helper exists to prevent. Prompt injection scan — the checkpoint fixture used the human-verification child element inside <verify>. That tag name is a fake-instruction-boundary pattern in scripts/prompt-injection-scan.sh, and the scan runs over changed files, so copying the shape from tests/verify.test.cjs (unflagged only because it is not in this diff) tripped the gate. Switched to the documented plain-prose <verify> form. The first attempt at that fix failed the same gate a second time: the comment explaining the collision quoted the offending tag literally. The comment now names it in prose instead — the scanner does not care whether a match is code or commentary, which is the whole point of the DEFECT.PROMPT-INJECTION-SCAN-COLLISION note in CLAUDE.md. Verified locally before push: scan reports 0 findings across 57 changed files, eslint clean, and both fixtures still validate as designed (gated one-way silent, ungated one-way warns, neither errors). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): record measured cost and halve gsd-tools spawns The Windows shard 1/3 job timeout was traced to the sharding layer, not to this PR's assertions — see #2472. Two contributing factors were this file's own, and are fixed here. 1. tests/test-timings.json had no entry for reversibility-tagging.test.cjs, so scripts/run-tests.cjs weighted it at the table's median fallback (~315ms) for LPT chunk packing. It actually measures 5595ms — an 18x under-weight. Recorded the measured value from the green gsd-test run (max across the node22/node24 lanes, per gen-test-timings.cjs's convention). Only this one entry: a full regen churns 634 entries of run-to-run drift, and the table is explicitly advisory and un-gated, so a 637-line diff does not belong in a feature PR. 2. Each verifyPlan() spawns gsd-tools, which dominates this file's cost. Spawns cut from 9 to 6 with no coverage lost: - the ungated-one-way warning and its stays-valid assertion now share one plan instead of building the same plan twice; - the reversible/costly never-flagged-as-ungated test was strictly subsumed by the additive suite, which already runs those two ratings ungated and asserts no /reversibilit/ warning at all — and the gate warning's text contains both "reversibility" and "one-way", so the broader assertion catches it. It only re-spawned gsd-tools twice to prove the same thing. Both are symptom fixes. The shard imbalance itself (19/11/10 minutes against a 20-minute cap, from a cost-blind round-robin partition that also reshuffles downstream files whenever one is inserted) is tracked in #2472. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(#1951): checkpoint fixture adopts the #2444 type-branched contract Surfaced by rebasing onto next, which gained #2444 (branch plan-structure validation on task type=checkpoint:*) while this PR was in review. cmdVerifyPlanStructure no longer applies one required-element set to every task. A checkpoint:decision now requires <name> + <resume-signal> + <decision> + <options>, and is exempt from the <action>/<verify>/<done>/ <files> set that auto and tracer tasks carry. The gated-one-way fixture predated that split and failed on the new requirement: errors: ["Task 'Task 0: Confirm the on-disk format' missing <options>"] Fixture rewritten to mirror the checkpoint:decision contract exactly — real <options> with two <option> children — rather than padding it with fields checkpoints no longer need. That also drops the plain-prose <verify> the earlier revision carried purely to dodge the prompt-injection scan; a checkpoint task has no <verify> requirement at all, so the workaround is moot. Verified against the real gsd-tools CLI across all four cases: gated one-way (valid, silent), ungated one-way (valid, warns), costly (valid, silent), absent (valid, silent). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
a7d83dc234 |
fix(#2390): warn on goal-shaped phase.add titles, correct auto-detect docs (#2425)
* fix(#2390): phase.add title warning + auto-detect doc fix phase.add now returns a `warning` field when a description reads as goal-shaped (>80 chars and/or multi-sentence) rather than title-shaped, instead of silently writing the whole paragraph verbatim as the `### Phase N:` header. The CLI still creates the phase as-is (the strict two-layer slash-vs-CLI interface is unchanged); the warning just surfaces the gap. Also clarifies six doc sites (command argument hints, workflow detection steps, and how-to/reference docs) that described the phase-number argument as "auto-detecting" the next unplanned phase -- that detection is an orchestrating-workflow/LLM step reading ROADMAP.md (concretely: `query roadmap.analyze`'s `next_phase` field), not a `gsd-tools.cjs` CLI feature. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2390): regenerate fixtures + lint gate-prep * fix(#2390): repair failing tests after gate verification * chore(#2390): add changeset (#2425) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ada79bee97 |
fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md (#2338)
* fix(#2308): make new-milestone workstream-aware; stop clobbering shared PROJECT.md Step 4 rewrote the `## Current Milestone` heading in the shared root PROJECT.md unconditionally. references/workstream-flag.md marks PROJECT.md `# Shared`, and per-workstream milestone state already lives in the workstream's own STATE.md / ROADMAP.md / REQUIREMENTS.md. With parallel milestones — the sanctioned design — whichever workstream ran new-milestone last silently won the shared heading. Step 4 is now skipped when a workstream is active; step 6 no longer stages PROJECT.md in that mode (cmdCommit returns nothing_to_commit rather than failing when a staged path is unchanged). Also fixes a second defect found while diagnosing this, same root cause (the workflow was workstream-unaware): step 1 parsed only --reset-phase-numbers and the milestone name, so GSD_WS was never set — yet ${GSD_WS} was interpolated at the routing lines. It always expanded to empty, so `/gsd:new-milestone --ws x` suggested `/gsd:discuss-phase [N]` with the workstream scope silently dropped, violating the routing-propagation contract. Step 1 now parses --ws using the established idiom from verify-work.md. Guard is keyed on GSD_WS, not $GSD_WORKSTREAM: the runtime launcher does not export the latter and it is only priority 2 of 5 in resolution, so it would miss the --ws flag case that is the actual repro. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2308): regenerate install goldens for the new-milestone workflow change gsd-core/workflows/ ships as an installed artifact, so new-milestone.md's content hash is pinned in all 18 runtime golden fixtures. Only that hash changed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2308): address review — inert step-6 guard, dropped Evolution repair, tautological tests Independent review found the first pass was partly cosmetic: 1. The step-6 `if [ -n "$GSD_WS" ]` branch was INERT. GSD_WS is assigned in step 1's shell and each step's bash block runs in its own shell — this file already proves it, since step 5 round-trips OUTGOING_MILESTONE through a file for exactly that reason (#2288). The guard read an unset variable, always took the flat branch, and staged PROJECT.md anyway. Rather than re-deriving GSD_WS in step 6, the branch is removed entirely: step 4 Part A's guard is what protects the shared heading, so post-guard the only content PROJECT.md can carry is Part B's idempotent Evolution backfill — which must be staged, not stranded. A regression test now asserts no cross-step GSD_WS branch returns. 2. Skipping ALL of step 4 also dropped the `## Evolution` structural repair — a shared, idempotent backfill that is not workstream state. A pre-Evolution project running only `--ws` would never get the section that transition and complete-milestone expect. Step 4 is now split: Part A (milestone-state write) is workstream-guarded; Part B (Evolution) always runs. 3. The tests were tautological prose-pinning — including one asserting a comment mentions "#2308". The step-6 test asserted the guard's TEXT was present, so it passed on the inert guard it existed to catch. Replaced with executable tests that extract the step-1 and step-6 fences and run them under bash with stubbed gsd_run, asserting real parse and --files behavior. 4. --ws is now stripped from the milestone name (step 1 previously left "--ws search" in the remaining text), and documented in argument-hint, help/modes/full.md, and docs/COMMANDS.md. 5. Changeset no longer overstates: --ws reaches the prose guard and routing hints only, not the SDK calls (state.milestone-switch/phases.clear/init.new-milestone still take no ${GSD_WS} — out of scope here). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * chore(#2308): regenerate SKILL.md, goldens, and size baseline for the argument-hint change skills/gsd-new-milestone/SKILL.md is generated from commands/gsd/new-milestone.md, so documenting --ws in the argument-hint made it stale (caught by lint:ci's gen-plugin-skills --check). Regenerated it plus the install goldens and workflow size baseline, since commands/, skills/, and gsd-core/workflows/ all ship. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2308): backfill PR number 2338 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
315d94f6d4 |
feat(#1945): tracer-first planning default + executor feedback gate (#2294)
* feat(#1945): tracer-first planning default + executor feedback gate Make "thin end-to-end slice first, verify, then expand" the default planning + execution discipline instead of the opt-in --mvp mode. - gsd-planner: first-class `type="tracer"` task; every plan LEADS with one production-quality end-to-end tracer slice by default; --no-tracer restores horizontal layers; --mvp/--tdd compose on top. - gsd-executor + execute-plan: post-tracer feedback gate — autonomous runs halt-on-fail before expansion, interactive runs emit checkpoint:human-verify after the tracer. - --no-tracer flag wired through plan-phase workflow/command/help/skill. - CONTEXT.md glossary defines tracer bullet vs prototype; docs + references reconciled. - tests/tracer-bullet.test.cjs: prose-contract + behavioral (verify plan-structure accepts tracer) coverage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1945): backfill changeset PR number to 2294 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2c878c966f | no-mistakes(document): Sync onboarding docs | ||
|
|
a3cca0704d | no-mistakes(document): Sync onboard documentation | ||
|
|
896c2740d3 | feat(#1990): add onboard command for brownfield setup | ||
|
|
8f2ebbe9bf |
feat(#1928): remove sunset Gemini CLI runtime, redirect to Antigravity (#1996)
* feat(#1928): remove sunset gemini cli runtime, redirect to antigravity Google sunset Gemini CLI on 2026-06-18; Antigravity CLI is its official successor (already a first-class GSD runtime). Remove the gemini runtime from the enum (16->15), aliases, labels, config-home fragment, install path, converters (convertClaudeToGemini{Markdown,Toml,Agent}, convertSlashCommandsToGeminiMentions), capability descriptor, gemini-extension.json, RULESET.GEMINI.*, and the interactive menu (renumbered, no gap). --gemini now prints an explicit deprecation notice citing the 2026-06-18 sunset and redirects to --antigravity (no silent alias, per the issue's Hyrum's-Law rejection). Antigravity is preserved throughout: its GEMINI.md contextFileName, .gemini/antigravity config home, the shared convertGeminiToolName/claudeToGeminiTools tool vocabulary, and the 'gemini' hookEvents dialect it declares. GEMINI.md retargeted as Antigravity's context file. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): backfill changeset PR number (#1996) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1928): drop Gemini CLI from issue templates (review nit) Removes the sunset Gemini CLI runtime from the two GitHub issue-template runtime lists that the removal PR missed, per @davesienkowski's review nit: - feature_request.yml: 'Applicable runtimes' checkbox (a user could otherwise request a feature for a runtime GSD no longer supports) - bug_report.yml: 'Runtime' dropdown + the stale ~/.gemini/settings.json retrieval-help line Leaves the post-removal templates fully consistent with the Antigravity redirect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
6d072435d0 |
test(#1975): consolidate 51 CLI + scripts-tooling regression tests into module suites
Fold 51 issue-named CLI black-box + scripts-tooling regression files into their canonical module suites (runtime-launcher-parity, worktree-safety, install-*, managed-hooks, read-guard, capability-registry, etc.), plus a NEW slash-command-namespace.test.cjs grouping the 4 slash/colon-namespace-leak invariant suites that had no canonical owner. Verbatim block-scoped describe wrappers; 427 subtests conserved 1:1. Host-env pre-check (per B2): no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. One folded suite (bug-3668 runtime resolver) creates an extension-less PATH gsd-tools stub + bash -c; co-locating it with the host's chmodSync tripped local/no-unguarded-nonportable-exec, so it's now Windows-guarded (skip on win32) matching the host suite's own bash -c guard. Regenerates regression-name allowlist (222->182), ratchets file-count allowlist (graphify 7->6, docs entry removed), makes 26 relocated allow-test-rule exemptions issue-ref-compliant (ADR-456; prunes stale ids). Repoints 13 tests/ references across CONTEXT.md, COMMANDS.md/FEATURES.md (EN + ja/ko/pt/zh) and ADR-0002. lint:ci green. Part of epic #1969. Closes #1975. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
bad2c76c5e |
feat(#1826): cmdStateRebuild CLI + dry-run + verbose + integration tests + docs (#1830)
Phase 2 of approved feature #1817. Wires the pure `rebuildCore` transition (Phase 1, #1827) to the `gsd state rebuild` CLI subcommand per ADR-1817 §5 (heavy/manual counterpart to lightweight, auto-triggered `state sync`). Source changes: - src/state.cts: implement cmdStateRebuild. Locks via readModifyWriteStateMd (real path) or read-only (dry-run). Wires phaseInventoryProvider to a real .planning/phases/ disk scan (same canonical source buildStateFrontmatter uses). --dry-run emits a structured preview without writing. --verbose tees the audit-log entries to stderr (treated as data-only per ADR-1577). - src/state-command-router.cts: register cmdStateRebuild in the StateModule interface + add the rebuild handler with --dry-run and --verbose flag parsing. - src/command-aliases.cts: register the state.rebuild canonical + 'state rebuild' alias + mutation=true (so the manifest covers the new subcommand for SDK parity / dispatch hub tests). Tests (tests/state-rebuild-cli.test.cjs, 5 cases): - state rebuild with no flags reconciles drifted body + drops orphan table rows + appends audit log (end-to-end #1, #2, audit log). - state rebuild --dry-run computes the diff, writes nothing (criterion #5). - state rebuild --verbose tees the log; audit-log section still written. - Running rebuild twice on the just-rebuilt file is byte-identical (criterion #6 end-to-end). - Missing STATE.md produces a clean 'STATE.md not found' message, no stack trace (CONTRIBUTING QA matrix). Verified locally: - node --test tests/state-rebuild-cli.test.cjs → 5/5 pass - node --test tests/state-rebuild.test.cjs → 18/18 pass (Phase 1 regression) - node --test tests/state-transition.test.cjs → 85/85 pass (ADR-1769 regression) Docs (docs/COMMANDS.md): document `state rebuild [--dry-run] [--verbose]` with the canonical-command block format used by `state sync` / `state prune`. Changeset (.changeset/1817-state-rebuild.md): type=Added, user-facing description of the new subcommand (closes #1817 epic on merge). |
||
|
|
18995380ce |
feat(#1154): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1738)
* feat(verify-phase): honest verifier — abstain (insufficient_spec) on non-inferable backstop truths (#1154) Carry the edge-probe's existing `backstop` (non-inferable) tier through the plan-phase projection as a structured flat-scalar marker instead of a prose parenthetical, and make verify-phase abstain -> human_needed (never silent-pass) on a backstop truth it cannot confirm with explicit evidence. Truth-axis mirror of #644's prohibition judgment-tier (ADR-550 D4). Engine (deterministic, CI-tested per ADR-550 D5 — never the LLM verdict): - src/probe-core.cts: truthStatement/truthVerification normalizers, projectTruths (conservative serializer), dispositionForUnverifiableTruth (backstop+no-evidence -> unverified/flagged/insufficient_spec; backstop+evidence -> green; inferable -> green, the over-abstention guard). - src/roadmap.cts: coerceTruthToString now reads `statement` first so an object-form backstop truth is surfaced, not dropped (Hyrum backward-compat for truth-readers). Workflow/agent/docs: plan-phase emits the structured marker (flat scalar, ADR-550 #1278); verify-phase + gsd-verifier add the abstain arm; new references/honest-verifier.md; FEATURES/COMMANDS document insufficient_spec; ADR-550 amended (truth-axis D4 mirror). Decisions adopted (trek-e review): insufficient_spec feeds existing human_needed with a distinguishable reason (no new VERIFIER_STATUS); changeset Changed; round-trip parity test; abstain-on-unconfirmed-backstop regression test red-first. Implementation notes (deviations from the issue's proposed file list, verified live): - frontmatter.cts needs no change — its flat parser already round-trips object-form truths. - verify.cts needs no change — it grades artifacts/key_links structurally; truths are LLM-graded at the workflow layer, so consumption lives there + the deterministic helper. - No CJS<->SDK hand-sync — the SDK seam was retired (ADR-0174); src/*.cts is sole source. Regenerated artifacts: golden-install-parity fixtures, INVENTORY-MANIFEST, size baselines. * chore(#1154): add changeset (Changed) for honest verifier User-facing changelog fragment for #1738. Typed `Changed` (not `Added`) per trek-e review condition 3 — the verify behavior shifts for backstop-bearing specs (a confident silent `passed` becomes `human_needed`), which is user-visible even though the schema marker is additive. * docs(#1154): score-formula also excludes abstained insufficient_spec truths (review nit-1) trek-e review nit: the verify-phase score sentence said PRESENT_BEHAVIOR_UNVERIFIED truths were "the only ones excluded" from verified_truths. Post-#1154 an abstained `insufficient_spec` backstop truth is also excluded (it is not ✓ VERIFIED and routes to human_needed). Behavior was already correct; this tightens the wording. Regenerated golden-install-parity fixtures + workflow-size baseline for the touched verify-phase.md. (Nit-2 — a dedicated insufficient_spec_items frontmatter list — is intentionally not taken: the current design is ADR-550-D4-conformant, the abstain cause rides as a distinguishable report reason, and adding it would exceed the approved scope.) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ba96c70b14 |
feat(#1602): deterministic coverage-metadata UAT routing for verify-work
Add an optional structured `coverage:` block to SUMMARY.md frontmatter and a deterministic classifier that `verify-work` consumes to route deliverables to auto-pass vs human-UAT — replacing the rejected #1598/#1599 post-hoc heuristic. - New `src/coverage.cts` (→ bin/lib/coverage.cjs) parses the nested coverage block (extractFrontmatter can't — its `-` items are scalars-only; this is a focused parser, sibling of parseMustHavesBlock), validates each entry, and classifies into auto_passed vs present. Frozen MODE/PRESENT_REASON/ERROR_CODE typed-IR surface. Exposed via `uat classify-coverage --summary <f>`. - Auto-pass is the narrow proven case only: strict-boolean human_judgment:false AND non-empty all-`pass` verification AND zero validation errors. Everything else — judgment, empty/failing verification, malformed entry — routes to the human (fail-safe). A malformed block falls back to legacy prose extraction and surfaces an error; an absent block is byte-identical to pre-#1602. - execute-plan create_summary populates the block (fail-safe default human_judgment:true); verify-work extract_tests consumes it; create_uat_file marks auto-passed entries `source: automated`. - Templates (summary + 3 variants), CONTEXT.md predicate + glossary, INVENTORY, eslint/gitignore registration, and Diataxis docs (COMMANDS reference + USER-GUIDE explanation) updated. - Behavioral tests via the CLI (no source-grep); parser-robustness regressions for the null-entry/comment-header/mis-indent cases found in adversarial review. Closes #1602 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e12a2abfd8 |
feat(#441): add /gsd-capture --list-seeds for seed listing and audit (#722)
* feat(#441): add /gsd-capture --list-seeds for seed listing and audit Seeds (.planning/seeds/SEED-NNN-slug.md) could only be created (--seed), enriched (--enrich), or auto-surfaced at /gsd-new-milestone. There was no way to browse or audit parked seeds on demand. This adds a read-only listing, following the established --list → workflow pattern (per the approved scope on - gsd-tools `list-seeds [status]` (cmdListSeeds in src/commands.cts): scans the seeds dir, returns { count, seeds[], summary } JSON with each seed's id, slug, status, scope, trigger_when, planted, title. Optional case-insensitive status filter. User-controlled content is sanitized (sanitizeForDisplay) and every path validated (requireSafePath); read-only. Independent of audit.scanSeeds, which only returns unimplemented seeds for the milestone surface. - /gsd-capture --list-seeds routes to a new read-only list-seeds workflow that renders the seed table. Closes #441 * chore(#441): point changeset fragment at PR #722 * test(#441): allowlist list-seeds test in prompt-injection scan The test asserts that list-seeds neutralizes injection payloads (<system>, [INST]) embedded in seed content, so the fixtures legitimately contain those patterns — same as the sibling security tests already on the allowlist. * fix(#441): use canonical /gsd:capture colon form in list-seeds workflow Claude-facing source (commands/, agents/, gsd-core/workflows/, ...) must use the /gsd:<cmd> colon form per ADR/CONTEXT.md; the hyphen /gsd-<cmd> form is retired there (enforced by bug-2543-gsd-slash-namespace.test.cjs). The new list-seeds workflow used the hyphen form. * docs(#441): sync help full.md + INVENTORY for --list-seeds Adds the --list-seeds entry to the help reference (help/modes/full.md, per bug-2954 argument-hint↔help parity) and registers the new list-seeds workflow in docs/INVENTORY.md (88→89) and the generated INVENTORY-MANIFEST.json. * docs(#441): add --list-seeds how-to + drop phantom statuses Addresses CHANGES_REQUESTED on PR #722 (two documentation blockers): - USER-GUIDE.md Seeds section (how-to): extend the task to cover auditing parked seeds on demand via --list-seeds, including the status filter — kept task-oriented per Diataxis how-to mode. - CLI-TOOLS.md (reference): drop phantom statuses implemented|rejected from the list-seeds filter vocabulary; the system only produces dormant|active|triggered (src/audit.cts scanSeeds). Reference must be factually accurate and complete. * fix(#441): guard non-scalar status frontmatter in cmdListSeeds A seed with a bare `status:` line (extractFrontmatter yields {}) or a `status: [a, b]` value (yields an array) crashed the whole audit list: `(fm.status || 'dormant').toLowerCase()` throws a TypeError on a non-string. Coerce every frontmatter read through a `fmStr` helper (mirrors the existing `typeof fm.id === 'string'` guard), so a non-scalar status falls back to dormant and non-scalar scope/trigger_when/title can no longer leak a raw array/object into the JSON contract. Title is now capped symmetrically. Adds regression coverage for empty and array `status:` and non-scalar fields. Refs #441 * docs(#441): align list-seeds workflow status vocabulary The load_seeds step listed `implemented` as an example status filter, but the real seed vocabulary is dormant|active|triggered (src/audit.cts scanSeeds); `implemented` has no producer. Matches the earlier CLI-TOOLS.md correction. Refs #441 * refactor(#441): extract pure deriveSeedIdentity; match raw status in list-seeds Pull the seed_id/slug derivation out of cmdListSeeds into a pure, exported deriveSeedIdentity(stem, rawFmId) so the parsing contract can be property-tested in-process (review minor #1). No behavior change. Filter comparison now matches the raw lowercased status (both sides already normalized) instead of sanitizeForDisplay(status); sanitization is for output, not matching (review nit #3). * test(#441): add fast-check property coverage and count=1 boundary for list-seeds Adds tests/list-seeds.property.test.cjs with four fast-check properties over deriveSeedIdentity (never-throws, string-only contract, canonical id->seed_id/slug invariant, filename-prefix fallback) per RULESET.TESTS.property-based-testing (review minor #1). Adds an N==1 status-filter boundary case to list-seeds.test.cjs (review minor #2). * chore(#441): sync runtime launcher snippet into list-seeds workflow Propagate the current _runtime-launcher.snippet.sh (with non-Claude runtime home probes) into the new list-seeds.md workflow via scripts/sync-runtime-launcher.cjs, satisfying bug-891 (E) propagation. * test(#441): record list-seeds.md in workflow size baseline (#1074) --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ac40f070ef |
feat(#1318): require external reviewers to verify plan claims against source (#1421)
* feat(#1318): require external reviewers to verify plan claims against source /gsd-review built its external-reviewer prompt from plan text only and never asked reviewers to open the repo and verify claims, so a grounded HIGH could be outvoted by ungrounded LOWs. Add a concise, generic source-grounding block to build_prompt's Review Instructions: treat yourself as running in the working tree, open referenced files, cite path:line + mechanism, trace asserted mechanisms, downgrade to an open question if you have no file access, and know that grounded findings are weighted more heavily. Also clarify that CodeRabbit (a diff-only reviewer that never receives the prompt) must not be weighted as a grounded plan-level verdict in consensus synthesis. Workflow stays under its size cap (baseline bumped deliberately). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1318): add changeset for reviewer source-grounding Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(#1318): mark changeset docs-exempt (internal reviewer-prompt wording) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(#1318): document reviewer source-grounding in COMMANDS.md; drop docs-exempt Review: a user-visible behavioral Changed warrants a docs touch, not a docs-exempt. Add a sentence to the /gsd-review entry in docs/COMMANDS.md (reviewers verify against source, cite file:line, grounded findings weighted higher) and remove the changeset docs-exempt marker so lint:docs passes via docs-updated. Also note the literal build_prompt test anchor. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(#1318): harden build_prompt fence extraction to be fence-run-aware Addresses maintainer review on PR #1421 (required-before-merge). The buildPromptReviewInstructions() test helper located the closing fence with `src.indexOf('\n```')`, which terminates at the FIRST triple-backtick line — so a build_prompt ```markdown block whose body embeds a fenced code example would truncate mid-content (dropping the `## Review Instructions` section) and give a spurious failure or false pass. Since this feature feeds source/plan content (which routinely contains code fences) to reviewers, that is a live fragility. Rewrite the extraction to be fence-run-aware, mirroring the CommonMark close rule in src/markdown-sectionizer.cts stripFencedCode: parse the opener's backtick run length, then close on the first line with >= that many backticks and only trailing whitespace — so a shorter nested fence is treated as content. Add a fail-first regression test (a 4-backtick outer fence wrapping a nested ```bash block) asserting the trailing `## Review Instructions` still extracts. Test-only change; no production .cts touched. Verified: test file 7/7, empirical fail-first proof the old indexOf logic truncated, full suite 4236/4236, eslint clean. Codex review: approve. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
02c6491e61 | docs(#1464): fix ADR-1244 capability doc set — followable tutorials, overlay-model + install tutorial, set/fragment/runtimeCompat reference, accuracy fixes | ||
|
|
7c93d9e222 |
feat(#1463): add capability outdated (per-source update check); drop phantom slash-command docs
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0d56f544d2 |
feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation (#1458)
* feat(#1435): capability matrix (generated + drift-guarded) + trust-model doc consolidation ADR-1244 Phase 6. Adds the capability matrix reference, generated FROM the committed registry so it can never drift from the actual capability set: - scripts/gen-capability-matrix.cjs (--write / --check); --check is a CI drift guard. - tests/capability-matrix-sync.test.cjs (4 tests): drift guard, buildMatrix==committed, every cap present, no placeholders. - docs/reference/capability-matrix.md regenerated from the registry (release-stable: shows engines.gsd, omits the lockstep per-cap version that would churn the file every release). - Consolidated the duplicate trust-model doc: deleted docs/explanation/the-capability-trust-model.md, merged its content into capability-trust-model.md, redirected ~10 references; no stale links remain. - Diataxis verification (now that gsd capability is a real command): corrected the matrix's third-party section — the matrix is the first-party catalogue; the overlay-aware view of installed third-party capabilities is 'gsd capability list', not this generated file. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1435): Added changeset for the capability matrix reference Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1435): address code-review — non-vacuous matrix test + generator polish - capability-matrix-sync.test.cjs: assert the 'security registers a ship:pre gate' precondition unconditionally so the extension-point check can never degrade to a vacuous pass on registry drift. - gen-capability-matrix.cjs: warn (stderr) on an unknown loop point at generation time; rename enginesOf -> fmtEngines for consistency with the other fmt* helpers (output unchanged). - capability-trust-model.md: point the two how-to links at the real files (import-a-capability-from-a-url.md, version-a-capability.md) instead of the bare directory. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1435): backfill changeset PR number → #1458 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
34bc096ec2 |
feat(#1451): wire gsd capability install/update/remove/list/disable/enable management CLI (#1457)
* feat(#1451): wire gsd capability install/update/remove/list/disable/enable CLI ADR-1244 D5/D6: the management command was built as a library (capability-lifecycle.cjs install/upgrade/remove + capability-ledger.cjs) across Phases 3-5 but never wired to a user-facing command — gsd-tools.cjs 'capability' only handled state/set. This adds the six subcommands, dispatching to the existing lifecycle/ledger: - install <spec> [--integrity] [--scope global|project] [--yes] [--shared-file <rel>]… - update [<id>|--all] [--scope] [--yes] [--shared-file] (re-resolves recorded source) - remove <id> [--purge-data] [--scope] (first-party rejected) - list [--json] (first-party + overlay, both scopes, JSON array) - disable|enable <id> (activation-state alias of capability set --off/--on) Scope→runtimeDir mapping matches capability-loader exactly (global=$GSD_HOME||home, project=project root; caps at <root>/.gsd/capabilities/<id>, ledger at <root>/.gsd-capabilities.json). Consent is non-interactive: --yes grants; without it an executable install aborts after printing the disclosure and writes nothing. Best-effort reconcile before each mutation. Tests: tests/capability-cli.test.cjs (20 behavioral, real resolver via local specs, GSD_HOME-sandboxed) — install consent/block/usage matrix, list, update round-trip, remove round-trip + first-party guard, disable/enable, unknown subcommand. Docs: docs/reference/gsd-capability-command.md reconciled to the real surface (ledger paths, --shared-file, consent model, disable mechanism, outdated marked planned); docs/COMMANDS.md gains the gsd capability entry. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1451): resolve adversarial-review findings + root-cause the --raw silent-output bug Adversarial-review (Codex) fixes: - capReadStrict passes a malformed strict_known_registries value THROUGH so the trust gate fail-closes on it (was silently downgrading to permissive) - installCapability/upgradeCapability gain an expectedId guard + first-party-id rejection (capability-lifecycle.cts): an overlay can't shadow a first-party id, and 'update <id>' can't act on a different id if the recorded source was retargeted - capability update: prints the consent disclosure, exits non-zero on --all partial failure, no longer masks the resolved id - capability remove: ledger-first ordering so an overlay is removable even if it shadows a first-party name; first-party guard only fires for ids not in the ledger - gsd-capability-command.md: disable/enable doc corrected (registry-known ids; overlay toggle not yet wired through this path) Silent-output bug (root cause, not waved off as pre-existing): - captureStdoutSyncWrites buffered fd-1 output and DISCARDED it on the throw path — any --raw command that emitted a result/error envelope then threw (to set a non-zero exit) lost ALL of stdout. Now it flushes the captured buffer before re-throwing (exit code preserved). - cmdCapabilitySet threw via process.exit() (bypassing the capture wrapper entirely); now throws ExitError so the wrapper flushes — matches the repo's no-process-exit architecture. - Regression test: capability disable <unknown> --raw must emit the JSON error envelope on stdout. Verified: capability suite 165/165, @file/json-errors/phase 183/183, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1451): address adversarial-review R2 — shared-file confinement, MCP no-clobber, config fail-closed - confinedSharedFile(): realpath-confine every shared-config write/strip to the scope root (mirrors safeRmUnder), so a --shared-file whose parent is a symlink escaping the scope can't write outside it. - mcpServers shared edits: never overwrite an UNOWNED entry — a name collision with the user's (or another capability's) server is skipped, so install/remove can't silently clobber user MCP config (hooks already append; the map-keyed mcpServers path was the gap). - capReadStrict: a PRESENT-but-unparseable .planning/config.json now fails CLOSED (lockdown) instead of silently downgrading the strict_known_registries policy to permissive. - Tests: symlink-escape shared-file writes nothing outside scope; colliding user mcpServers entry preserved; unparseable config blocks an external install. capability suite 83/83, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#1451): address code-review — aborted-status robustness + coverage + project-scoped strict doc - install/update: handle an 'aborted' result independently of the requiresConsent flag so it can never fall through to the generic 'blocked: unknown reason' arm (aborted always means consent-needed per the lifecycle contract; latent today, hardened for future status additions). - Clarify capResolveScope comment (project scope === already-resolved cwd) and document that strict_known_registries is a PROJECT-scoped policy (read regardless of --scope; no machine-wide allowlist) in gsd-capability-command.md. - Tests: update --all over an empty ledger returns an empty result set (exit 0); a flag value that looks like another flag (--integrity --scope) is rejected, not swallowed. CLI suite 33/33, lint clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1451): FEATURES.md entry #147 + Added/Fixed changesets for the capability CLI Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#1451): backfill changeset PR number → #1457 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
1abebbf4fd |
feat(#1434): registry-driven dispatch for third-party capabilities (ADR-1244 Phase 5) (#1450)
ADR-1244 Phase 5 (D7). dispatchOverlayCapabilityCommand in gsd-tools.cjs dispatches an installed third-party capability command family via loadRegistry({includeInstalled}), gated on a committed ledger entry (consent) and confined to the capability's install root (defaultRequireFromInstallRoot: bare-.cjs basename + realpath containment, rejects ../ traversal + symlink escape); same own-property/function/sync/ExitError guards as the first-party path. capability-loader records _overlay.commandRoots only for accepted overlay caps with a committed, structurally-valid ledger entry (fail closed). First-party graphify/intel/audit unchanged (already on the registry seam). 3 Codex rounds converged + /security-review (no HIGH) + /code-review (Approve); gsd-test green both platforms; CI green.
Closes #1434.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|