next
120 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a9a7a328e6 |
refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted. |
||
|
|
6a4984cf69 |
fix(#4763): surface the displaced session record and pass --phase from the executor decision loop (#4919)
* test(#4763): failing-first — replaced-record payload and executor --phase pins * fix(#4763): surface the displaced session record and pass --phase from the executor decision loop state record-session keeps its last-writer-wins write (the recorded single-slot handoff design) but no longer displaces silently: when a non-empty Stopped At or authored Resume File record is replaced, the payload carries the full prior text under replacedRecord. Same-value rewrites, the insert path, and the #944 template-default DWIM are not displacements and report nothing. The executor decision loop now passes --phase "${PHASE}" to state.add-decision, matching execute-plan.md, so decisions stop inheriting whichever phase the global pointer names (#4763 case 2). advance-plan is unchanged (#3311 by-design). Emitted-Drift-Ack-Growth: gsd-executor.md — the decision loop gained its --phase guard and a comment naming why (#4763) * docs(#4763): add the changeset fragment * docs(#4763): backfill the changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
ccfed63355 |
fix(#4823): the Current Plan reset is scoped to the Current Position section (#4898)
* test(#4823): failing-first — the Current Plan reset must not rewrite prose outside Current Position * fix(#4823): the Current Plan reset is scoped to the Current Position section — the whole-body 'Plan' fallback matched hard-wrapped prose lines starting with plan: * chore(#4823): changeset fragment * chore(#4823): backfill changeset PR number (4898) --------- Co-authored-by: sim <sim@local> |
||
|
|
eadcba5f53 |
fix(#4481): anchor bold STATE field reads to line start (#4510)
* test(#4481): reproduce mid-sentence state field reads * fix(#4481): anchor bold STATE field reads to line start * docs(#4481): add changeset for #4510 * fix(#4481): align bold field readers with anchored writers |
||
|
|
7fe440a838 |
fix(#4488): report state update as successful when the value is already correct (#4581)
* fix(#4488): report state update as successful when the value is already correct `cmdStateUpdate` unconditionally overwrote `updateCore`'s own `updated:true` signal with `reconcileReportedFields`'s disk-diff result. That diff reports `[]` -- by design -- whenever `readModifyWriteStateMd`'s #948 no-op guard fires because the transform's output was byte-identical to the input, which happens precisely when the requested value already equals what's on disk. The field genuinely was found and matched; there was simply nothing left to change. Collapsing that into the same `false`/"not found" response as a genuine miss produced an actively wrong diagnostic message and a silent same-day no-op in gsd-ship + gsd-extract-learnings, which both write `Last Activity` to today's date. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * docs(#4488): backfill changeset pr number to 4581 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4488): untrack tdd-red-evidence.cjs, completing its ADR-457 gitignore migration Bundled discovery from this PR's own CI run: tests/lint-compiled-artifact- sync.test.cjs's full tsc compile (which runs whenever ANY compiled artifact remains tracked) SIGTERM'd under shard contention. gsd-core/bin/lib/tdd-red- evidence.cjs (introduced by #3770/PR #4279) was the sole remaining tracked artifact -- a tenth, later, separate instance of the #2657/#2653 migration-gap defect class this test file's closed nine-item list doesn't cover. Untracked it and added the .gitignore entry, same fix shape as the original nine. This eliminates the slow tsc-compile path entirely (verified: 0.1s vs ~7s locally) rather than papering over a timeout. Added a generic regression test asserting the tracked set is fully empty. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
6c5e11049b |
fix(#4383): require phase before planned-phase writes (#4534)
* fix(#4383): require phase before planned-phase writes * chore: add changeset for #4534 --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
6ebe6372ce |
fix(#4243): anchor stateReplaceProgressPercent bold form to line start (#4474)
* fix(#4243): anchor stateReplaceProgressPercent bold form to line start The bold branch of stateReplaceProgressPercent carried no ^ and no /m flag, so a bold percent-ish label quoted MID-SENTENCE inside prose — an Accumulated Context bullet mentioning **Progress:** — captured the machine-segment rewrite and destroyed the rest of its line, silently, while the real Progress line stayed stale (and the frontmatter moved on without it, breaking the #4213 surfaces-agree contract). Every caller (cmdStateUpdateProgress, syncCore's percent arm, applyPostSyncPreservation) feeds the whole document, so all three were exposed. Anchored to ^([ \t]*\*\*Progress:\*\*[ \t]*)([^\r\n]*)$ with /im — the exact idiom #4453 applied to stateReplaceField's bold branch (same-line confinement per #4010: the leading class is [ \t]*, deliberately not \s*, which can consume the newlines before the label into the match; $ is explicit-and-inert and documents end-of-line). #2177's recorded requirements all stand: frontmatter is stripped before matching, the suffix-preserving machine-segment swap is untouched, and bold-beats-plain priority now governs line-start forms, so an earlier free-text plain Progress: line still cannot capture the rewrite ahead of the real bold status line. Per the maintainer ruling (2026-09-07), #2177's incidental bold-anywhere matching was not load-bearing. * test(#4243): scope the C4 region check with splitLines, not a bare \n split lint:ci (local/no-crlf-fragile-split) flagged the free-text-plain-line row's content.split(/\n## /)[0] — a bare \n split on readFileSync content is CRLF-fragile under Windows autocrlf. Same scoping via splitLines() (src/text-lines.cts), which splits on \r?\n. * chore(#4243): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
33e393ba4c |
fix(#4243): anchor stateReplaceField bold form; pin frontmatter round-trip (#4453)
* test(#4243): failing-first regressions for bold-field anchoring and frontmatter round-trip * fix(#4243): anchor stateReplaceField bold form to line start The bold branch of stateReplaceField carried no ^ and no /m flag, so a bold label quoted mid-sentence inside prose — the issue's **Status:** inside an Accumulated Context bullet — captured the rewrite and destroyed the rest of its line, silently, whenever a whole-body caller fed the function every section (beginPhaseCore's tryField, advancePlanCore's Status/Current Plan writes). The plain branch was always line-anchored; only the bold branch lagged. Anchored to ^([ \t]*\*\*Field:\*\*[ \t]*) with /im, reusing #4010's same-line confinement idiom for the leading class (deliberately not the issue's suggested ^\s* — it can consume the newlines before the label into the match) and #4186's recognition-by-anchoring discipline. Frontmatter half of the issue (unknown-key drops, invented milestone defaults) is already fixed on next by #2202/#3216/#4129; pinned here with the issue's requested regression fixtures. * test(#4243): pin survival contract, not derived percent, in frontmatter rows Bench RED run caught two assertion defects in the pin rows: the unknown progress subkey re-parses as a quoted scalar ('77' vs 77), and percent is a declared derived subkey - omitted under the #3573 no-roadmap withhold, recomputed when measured (#4129) - so pinning its value over-pins derived semantics. The rows now pin what the issue demands: unknown/custom keys survive, stored counters are kept under the withhold, milestone identity is never reset to invented defaults. * chore(#4243): changeset for the anchored bold-field fix * chore(#4243): backfill PR number in changeset |
||
|
|
38e4ce5f62 |
fix(#4186): anchored status vocabulary, record-session arg guard, recount pin (#4381)
* fix(#4186): anchored status vocabulary, record-session arg guard, recount pin Three defects from #4186: 1. normalizeStateStatus ran a first-match-wins SUBSTRING chain over the free-prose body Status field, so prose merely mentioning a status word was silently rewritten to a credible wrong token (a .planning/ path in Italian prose -> status: planning; verifica -> verifying; completezza -> completed). Recognition is now an ANCHORED whole-field match against a declared vocabulary (STATUS_EXACT_TOKENS + STATUS_ANCHORED_PATTERNS, state-document.cts) — case/whitespace-tolerant, branch-order artifacts preserved (Planning complete -> planning; Phase complete — ready for verification -> verifying). The recorded lenient fallback (#3873 row 26) stands: unrecognized prose passes through verbatim. Read-side consumers (W011, statusline) ride the same function. 2. The progress recount skew (stray *-SUMMARY.md inflating completed_plans) is already dead on next via #1988/PR #2016 (countMatchedSummaries pairs summaries to plans) — verified live and pinned with regression rows composed against the #4129/#4359 ratchet. 3. state record-session with no args executed and wrote STATE.md; it now errors like state update (stopped-at or resume-file required), handler- side so SDK callers are covered too. Four tests pinning the bare-call write are updated to the new contract. * fix(#4186): update status pins to the anchored vocabulary contract Bench round 1 follow-ups: - Legacy bare 'Milestone complete' kept as reader-side vocabulary (ADR-2207 removed the writers, not recognition of legacy files). - state.test pins updated: 'Paused at Plan 3' and round-trip 'Executing Plan 5' were pins of the substring guessing itself — the round-trip now uses the real handler form 'Executing Phase 5'. - record-session no-op/no-fields tests repurposed to the usage-error contract (CLI + SDK-level ExitError), byte-unchanged assertions kept. - statusline tests repinned: vocabulary values collapse to keywords; narratives render the documented first-word fallback instead of a guessed token. Hook doc comment updated to match. - docs-guard exempt baseline: state.test.cjs now cites docs/CLI-TOOLS.md. - docs/CLI-TOOLS.md: record-session signature notes the required flag. * fix(#4186): repair a dangling sentence in the schema docstring * test(#4186): bound the completed_plans scan regex (#2128 class) * chore(#4186): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
66e4034fe4 |
fix(#4138): begin-phase without --phase exits non-zero and writes nothing (#4380)
* test(#4138): failing-first regression — begin-phase without --phase must fail closed * fix(#4138): begin-phase without --phase exits non-zero and writes nothing * chore(#4138): changeset fragment for begin-phase arg validation * chore(#4138): backfill PR number in changeset --------- Co-authored-by: sim <sim@local> |
||
|
|
e6d047decc |
fix(#4129): derive completed_phases from the ROADMAP authority; honor the progress-ratchet on every state write (#4359)
* test(#4129): failing-first regressions — completed_phases clobber on resyncing writes and phase-complete failure to increment * fix(#4129): completed_phases derives from the ROADMAP authority and the write path honors the progress-ratchet Three coordinated prongs (diagnosis in .gsd/bug/fix-4129-completed-phases-recompute/): P1 — buildStateFrontmatter's disk scan floors the completed-phases numerator at the milestone-scoped ROADMAP Complete-row count (deriveProgressFromRoadmap, the one owner), gated inside the same safeToUseRoadmapCount / not-withheld branch that owns the denominator. A completed phase whose verification routes stale (#2348 clean-commit-time drift) or is missing no longer under-counts forever. P2 — applyPreserveAlways's resync arm merges instead of wholesale-replacing on a measured scan: totals derived both directions (#2440), completed counters up-only (#2969 — the schema-declared progress-ratchet, now enforced on the write path like the read path always has), percent recomputed from the merged counters. The #3756 unmeasured guard and the #3242 explicit-progress contract are unchanged. P3 — phase complete's atomic 3-file commit passes the post-completion ROADMAP-derived counters through the #2736 authoritativeFm seam (new object direction for the progress key; completedOnlyRaise at the post-preservation re-assert), because the transaction's disk scan reads the pre-completion ROADMAP and failed to increment on the completing phase's own write. * fix(#4129): adversarial-review hardening — intent is a floor at BOTH authoritativeFm sites The pre-preservation merge could lower a correctly-higher disk-derived counter (a verification-passed phase whose ROADMAP table row drifted behind the disk signal). completedOnlyRaise now governs both application sites: the intent and the derivation agree on direction (up), never on subtraction. * fix(#4129): the ratchet merge keeps derived values verbatim when numerically equal The re-parsed derived block carries string scalars ("2") while the curated snapshot carries numbers (2); substituting the curated spelling over an equal derived one was a no-op in substance but a shape churn the ADR-3473 §8.7 reporting loop surfaced as a phantom preserved-over-disagreeing-derived warning on phase complete (ADR-3408 §8.5 Matrix B). Only a strictly-greater curated counter replaces the derived value now; percent gets the same verbatim rule. * changeset(#4129): backfill PR 4359 --------- Co-authored-by: sim <sim@local> |
||
|
|
3d03ae65e6 |
fix(#4094): withhold all four STATE.md progress counters under the milestone-unbounded guard (#4322)
* test(#4094): failing-first matrix for withholding all four progress counters * fix(#4094): withhold all four progress counters under the milestone-unbounded guard completed_phases/total_plans/completed_plans are accumulated from the same phaseDirs walk as total_phases, so the #3354/#3573 withhold condition makes them equally untrustworthy — yet only total_phases was withheld, and every resyncing state.* write silently clobbered the three stored siblings with the under-scoped disk numbers. Extend the withhold-then-fall-back-to-stored pattern to all three siblings: null sentinels in the disk-scan cache value, three new stored-counter readers threaded through all three buildStateFrontmatter call sites, and the same cached-else-stored consumer fallback. Milestone-bounded projects are untouched (gate-conditional). * fix(#4094): scope-requires for the new test block, keep the (#3573) warning token, and update two #3578 rows to the withheld-counter contract - the #4094 describe sat after the closing brace of the section that owned the module-level beforeEach destructure, so it needs its own local requires (mirroring the #3642 block); - the #3573 warning keeps its literal '(#3573)' tag (asserted by an existing test) with '#4094' appended as a separate token; - two #3578 status-guard rows in tests/state.test.cjs asserted the pre-#4094 unconditional disk-scan assignment of completed_phases under the roadmap-absent withhold — exactly the silent clobber #4094 removes; the status-guard conclusion (must not fire) is unchanged, the counter-value assertions now pin the withheld contract. * test(#4094): lint conformance — splitLines for the persisted-progress parser, local seeder, scoped rmSync disable * changeset(#4094) * changeset(#4094): backfill PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
2e1ede6d99 |
fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline (#4318)
* test(#4093): regression matrix for advance-plan zero-labeled-fields decline * fix(#4093): give advance-plan's zero-labeled-fields failure a disk-derived recovery decline * refactor(#4093): collapse IIFE to a plain block (review finding) * docs(#4093): document the advance-plan recovery decline + changeset * chore(#4093): backfill PR number in changeset * fix(#4093): budget lint-compiled-artifact-sync's tsc compile as a compile, not a probe --------- Co-authored-by: sim <sim@local> |
||
|
|
70f22e4643 |
fix(#4213): keep STATE.md progress surfaces synchronized (#4231)
* fix(#4213): keep STATE.md progress surfaces synchronized * fix(#4213): clamp the shared progress bar and keep bold-first priority, changeset + property tests - formatProgressMachineSegment clamps through clampPercentFromFraction (ADR-3180 Decision 7 kernel) with a 0 floor, so a hand-edited out-of-range persisted percent renders a clamped bar instead of throwing RangeError on repeat() inside the write seam - stateReplaceProgressPercent restores the #2177 bold-first priority: **Progress:** anywhere in the body wins; a plain ^Progress: line is the fallback, so free text starting with Progress: cannot capture the rewrite ahead of the real status line - cross-reference comment names the three consumers and the cmdStateSync sanctioned exception (ADR-3408 §8.3) - CONTEXT.md: applyPostSyncPreservation reconciliation documented in the STATE.md Transition Module entry - property tests (never-throws/well-formed, idempotency, round-trip, bold-first) + two regression rows through the CLI --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
294ec29857 |
fix(#4053): quote decimal-shaped frontmatter scalars for spec YAML readers (#4165)
* fix(frontmatter): quote decimal-shaped scalars so a spec YAML reader preserves them A decimal phase identifier written to STATE.md frontmatter (e.g. `current_phase: 22.10`) was emitted BARE, because `scalarNeedsDoubleQuoting` only asks whether a value can OPEN a plain scalar — which `22.10` can. A YAML-spec reader (js-yaml, the statusline, any external tool) then reloads bare `22.10` as the float 22.1, colliding with `22.1` and dropping the trailing zero. gsd's own tolerant line-scanner (`extractFrontmatter`) round-trips the raw text and so hid the defect; a spec reader does not. Fix: `reconstructFrontmatter`'s general scalar path now also quotes numeric- looking strings that are not plain all-digit integers (decimals, exponents, sexagesimal, hex/oct/bin) via `generalScalarNeedsNumericQuoting`, reusing the existing `YAML_NUMERIC_RE`. Every all-digit string — integer counts, phase numbers, and leading-zero fixtures like `02` — stays bare, so the state-rebuild idempotency baseline and the rest of the state corpus are unchanged. This also quotes `gsd_state_version: 1.0` on write, which matches the authoritative STATE.md template (`src/state.cts` already emits it quoted). Regression test drives the real write path and asserts, via js-yaml, that `22.1` and `22.10` no longer collide and read back string-typed; guards that integers and free-text stay unquoted. Fixes #4053 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YBicDMJyh3AH56ZFbUsyxC * chore(changeset): add Fixed fragment for #4053 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YBicDMJyh3AH56ZFbUsyxC * docs(frontmatter): trim the generalScalarNeedsNumericQuoting comment Cut the over-long doc block down to the essential why and drop the inline comment that repeated it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011ZKeSj55VakqQajtoBgCTC * docs(test): drop the #4053 explanatory comments from the touched tests The assertions speak for themselves; remove the added narrative comments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011ZKeSj55VakqQajtoBgCTC * fix(#4053): correct the trade-off comment, changeset PR number, and cover every claimed numeric form Review follow-ups (trek-e): - The doc comment claimed a plain integer round-trips harmlessly. That is false for leading-zero values (`02` -> 2, `017` -> 17 under js-yaml). Rewrite it to state the real, deliberate trade-off: all-digit strings stay bare because zero-padded ids (`plan: 01`, `phase: 02`) are the pervasive GSD convention and quoting them all is the blanket quoting #4053 asked to avoid; the loss is padding not identity (`02` and `2` normalize to the same phase, `22.1` and `22.10` do not). - Changeset carried the auto-closed draft's number (4151); correct to 4165. - Test exponent, hex, octal, binary and sexagesimal forms through js-yaml, and pin the leading-zero trade-off so the documented behaviour is asserted. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qc7VN4zTpTSDTS9JXM2cFB --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
9ee6d54cc3 |
fix(#4306): extend fault-injection fd-swallow fix across the whole suite (#4308)
* fix(#4306): forward real bytes through io.test.cjs's fault-injection mocks The bug #1008 fault-injection tests mock fs.writeSync scoped only by file descriptor. On their "success" arms (the retry-after-EAGAIN/EINTR call, and the short-write simulation) they fabricated a return byte count without ever calling the real writeSync -- the bytes went into a local array and nowhere else. node:test's process-isolation runner (default on Node >= 22) reads each test file's own stdout to parse its child-to-parent result protocol. If the runner's own reporter write for an adjacent test lands on fd 1 while one of these mocks is installed, that write was silently swallowed instead of reaching the real pipe -- observed in CI as "Unable to deserialize cloned data" (a corrupted/truncated byte stream on the parent's read side), not a thrown exception. Every "success" arm now forwards the real bytes to orig()/restore() instead of fabricating a return value, so anything else sharing the fd during the mocked window still gets its bytes delivered for real. writeAllSync (the only production caller reaching this mock) always passes a Buffer, so the forwarded calls use the buffer-form fs.writeSync overload unambiguously. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#4306): extend fault-injection fd-swallow fix across the whole suite The originally-fixed instance (tests/io.test.cjs) was one occurrence of a copy-pasted defect: mocked fs.writeSync arms fabricated a return byte count without ever forwarding the call to the real fs.writeSync, silently discarding bytes. Under node:test's process-isolated runner, the parent reads the child's real stdout to parse v8-serialized report frames interleaved with plain output (confirmed against node's own lib/internal/test_runner/runner.js and a matching upstream issue, nodejs/node#64061) — a swallowed write on that fd corrupts the parent's parse ("Unable to deserialize cloned data"). Adds a shared, safe capture helper to tests/helpers.cjs, captureFdSync(fd, fn): it always forwards every write to the real fs.writeSync first, then records only the observed fd's bytes, sliced by the real return count (not the requested length), decoded once via Buffer.concat so a short write can't split a multi-byte codepoint across two decodes. 17 test files migrate their local copy of the unsafe mock to this shared helper. tests/worktree-base-ref.test.cjs keeps a narrower in-place fix instead (it needs to record every fd a write touched, which the shared helper doesn't expose). tests/io.test.cjs gets two follow-up correctness fixes on top of the already-committed forwarding fix: the EAGAIN/EINTR/short-write arms now derive their recorded chunk from the real return count everywhere (including the string-form overload), and the short-write test no longer forces a Buffer-shaped truncation call onto a string-form write that could land on the same fd. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
2e056488d9 |
fix(#4067): derive advance-plan phase-complete from disk, not the plan counter (#4292)
* test(#4067): pin advance-plan phase-complete guard matrix (RED) Five-case matrix: decline on unsummarized plans (regression), fire on fully-summarized phase, fail-open on unresolvable phase dir, idempotent decline, normal advance untouched. * fix(#4067): derive advance-plan phase-complete from disk, not the plan counter The phase-complete branch of state.advance-plan was decided purely by STATE.md's scalar plan counter (currentPlan >= totalPlans). A stale counter carried into a newly planned phase, or a counter raced by wave-parallel executors, let 'Phase complete — ready for verification' land while sibling plans were still executing. cmdStateAdvancePlan now re-decides that branch from disk before the write: every plan in the Current Position phase's directory must have a SUMMARY.md (scanPhasePlans single owner, the same source state.update-progress recalculates from). Outstanding plans decline the entire write byte-identically (idempotent, concurrency-safe, counter stays display-only); an unavailable disk answer fails open to the counter-derived decision. * fix(#4067): review round 1 — route phase-dir lookup through listMilestonePhaseDirs #3185 drift guard: no hand-rolled phases-dir readdirSync. Windowed (current-milestone) lookup first so an archived milestone's stale dir cannot shadow the live one; unscoped retry when the window cannot answer. Also restore the transform's undefined-data error semantics and extract scanOutstanding. * chore(#4067): add changeset fragment * chore(#4067): backfill PR number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
ddb877fa0a |
enhance(#3957): a no-op reports the real condition and the values it already computed (#4157)
* test(#3957): add failing-first coverage for no-op decline reporting (epic #3473 B9) * fix(#3957): a no-op reports the real condition and the values it already computed (epic #3473 B9) * test(#3957): correct stale assertions and a withheld-arm fixture after rebase (epic #3473 B9) * docs(#3957): add Fixed changeset fragment for no-op decline reporting (epic #3473 B9) * docs(#3957): backfill changeset PR number to #4157 --------- Co-authored-by: sim <sim@local> |
||
|
|
bdfc62889b |
fix(#3784): read the hybrid Current Plan: N of M shape, keep zero-padding, and name the accepted shapes on failure (#3791)
* fix(state): read the hybrid "Current Plan: N of M" shape advancePlanCore derived the value FORMAT from the field NAME, so it handled the legacy pair (`Current Plan` + `Total Plans in Phase`) and the compound `Plan: N of M`, but not the hybrid of the two: the legacy field name carrying a compound value with no Total Plans sibling. `legacyTotal` is null so the legacy branch fell through, and the compound branch reads the `Plan` field through a `^Plan:`-anchored pattern that never matches `Current Plan:`. Both produced NaN against a file whose plan numbers are plainly readable. The shape is not exotic. An agent wrote it unprompted into a project's STATE.md, believing it was the parseable form, and every subsequent run in that project inherited the failure and worked around it by hand. Track the field name and the value shape separately (`planSourceField`, `planRawValue`) so write-back targets whichever field the value came from. The legacy pair still takes precedence when both fields exist, so a stray "of N" inside Current Plan cannot override an explicit Total Plans — covered by a new test. Also replace the caller's catch-all error. It reported "Cannot parse Current Plan or Total Plans" for ANY transition failure, and named no accepted shape, so a reader learned neither what failed nor what to write. It now distinguishes "no result" from "unreadable plan position" and lists all three shapes. The existing test asserted the literal "cannot parse"; it now asserts the message names the shapes, which is the property that makes it actionable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * fix(state): keep zero-padding when advancing a compound plan value The compound write-back rewrote only the leading half of "N of M", so a padded value drifted lopsided: "04 of 06" advanced to "5 of 06". Cosmetic on its own, but a plan line that looks wrong is one the next writer tidies by hand, and hand-tidying this particular line is what produced the hybrid shape the previous commit had to teach the parser to read. Pad the incremented number to the width it was written with. padStart never truncates, so a value that outgrows its padding widens correctly: 09 of 12 advances to 10 of 12. Unpadded values are untouched — 2 of 6 still advances to 3 of 6. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * fix(state): pass a literal field name to the compound write-back The previous commit passed `planSourceField` — a variable — as the field-name argument to `stateReplaceField`, which trips the state-write-path drift guard's `unstripped_content_write` axis (ADR-3408 §8.3(b)). The guard is right to care: a Title-Case literal cannot collide with a lowercase or snake_case frontmatter key, so it is safe whatever the content argument is, while a variable could hold anything and therefore requires its content to be demonstrably frontmatter-stripped first. The content argument here IS stripped — `body` is `stripFrontmatter(content)` — but the guard does a narrow backward scan rather than dataflow tracking, by design, and the nearest preceding assignment to `body` is another `stateReplaceField` result. Rather than baseline a bypass or ask a future reader to re-derive that the invariant holds, dispatch on the discriminator and pass the literal. Guard goes from 1 finding to 0; its own 32 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * chore(3784): add changeset fragment for #3785 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * test(#3784): cover the maintainer's AC1 write-back and AC6 reader-anchoring Triage published six acceptance criteria; two were only half-covered. AC1 asks that the hybrid write back to the SAME field with padding preserved. The existing hybrid test used an unpadded value and asserted only `result.data`, so it proved the parse but never the write. Now asserts the written content is `05 of 06` on the original field, and that no separate `Plan:` field appears as a side effect. AC6 asks that the shared field reader not be loosened. Reading the hybrid is the transition's job; `stateExtractField('Plan')` is line-anchored and has 13+ callers, so teaching it to match a name merely ENDING in "Plan" would be the wrong fix and would silently change what those callers read. This holds by construction here — the reader is untouched — but nothing locked it in. The new test fails if anyone later reaches for that shortcut. Also drops the changeset fragment written against the auto-closed PR number. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * chore(#3784): add changeset fragment for #3791 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016QHxbMHTPYbEqAnKTpJPR8 * fix(#3784): write the advanced plan back to the field it was read from Review findings 2-6 on #3791 were one defect seen from several angles: the read path learned the hybrid `Current Plan: N of M` shape, the write path did not follow it. - `bumpLeadingNumber` now owns the increment for all three parse branches. Only the leading digits belong to this transition; the padding width and everything after it (` of M`, and the `\r` of a CRLF file) are the author's text and are preserved. The legacy branch wrote `String(newPlan)`, which turned `2 of 99` into `3` and `04` into `5`. - `mutateCurrentPositionForAdvance` takes the plan field NAME. Its plan arm only ever looked for `Plan:`, so on a hybrid file the `## Current Position` section was never reached; combined with the body-level write being single-shot and bold-preferring, a file carrying the field at both sites advanced the header and left the section a plan behind. The parameter defaults to `Plan`, so the two callers that pass no plan are unchanged. - Tests: both-sites-advance (fails without the section arm), legacy write-back content assertions (the previous test read only `data` and so could not see the lossy write), hybrid boundary at limit-1 and limit+1, a CRLF fixture, and an fc property pinning the padding-width contract. Two characterization tests pinned `**Current Plan:** 02` advancing to `3`. That dropped padding is the defect #3784 reports, so the expectation is corrected to `03` rather than the fix being narrowed around it. * fix(#3784): drop the unreachable advance-plan error branch, sync the doc Findings 1 and 8 on #3791. The `!resultData` arm could not fire: the transform callback assigns `resultData` unconditionally, only runs once STATE.md is known to exist (the missing-file case returns "STATE.md not found" upstream), and every `advancePlanCore` return path sets `data`. It was a speculative second failure mode with a message no caller could receive, and the comment beside it claimed to distinguish two things that were never two. `!resultData` stays in the condition as a type guard, which is all it ever was. `docs/json-errors.md:142` quoted the old error literal verbatim and was the sole occurrence in the tree; it now quotes the emitted one. * chore(#3784): describe the write-back fix in the changeset * fix(#3784): anchor the plan grammar and widen the schema row to match Review round 3 on #3791: B1, B2, M1, M2, M3, M4 and the planSourceField nit. B1 — `STATE_FIELD_SCHEMA.current_plan.acceptedShapes` widens to `['N', 'N of M']`, which is what `src/state-md-schema.cts`'s own comment instructed this PR to do on merge. `'N/M'` stays undeclared so row 23 keeps a non-vacuous undeclared candidate to probe. Verified `gen-state-md-docs --check` exits 0 and `--write` rewrites 0 of 6: the generated artifacts do not surface this row, so there is nothing stale to regenerate. B2/M4 — the discriminator was `/of\s+(\d+)/`, unanchored, so a total could be read out of prose. `Current Plan: 4 — blocked on review of 2 PRs` parsed as `4 of 2`, took the `currentPlan >= totalPlans` branch and WROTE `Status: Phase complete — ready for verification` into the user's file. Both shapes are now anchored at the start and every number comes from a capture group via `planNumberFrom`, which rejects anything past `Number.MAX_SAFE_INTEGER` rather than letting `data` and the persisted string disagree. Nothing on this path calls `parseInt` on a raw field value any more. The grammar keeps a trailing remainder after the total, because `Plan: 2 of 5 in current phase` is a real tested shape. The refusal comes from requiring `of <total>` to follow the leading number immediately, not from forbidding a suffix. M1 — `bumpLeadingNumber` is total. It returned its input unchanged when there were no leading digits, so `+2` reported `advanced: true` while writing the file untouched. M2 — both section arms use replacer functions. File-derived text was being spliced into a `String.replace` replacement string, where `$&` / `` $` `` / `$'` expand: a value of `04 of 06 $&` spliced part of the document into itself. `stateReplaceField` already used a function; these now agree with it. M3 — the section arm targets the name the SECTION carries, and the body write now writes both spellings, each with its own rendering. Keying off the header's name left the other name stale in both directions: a legacy header beside a `Current Plan:` section line, and a `**Plan:**` header beside one. * fix(#3784): derive the shape error from the schema, widen the test coverage Review round 3 on #3791: B3, m1, m2, m5 and the two test nits. B3 — the accepted-shape set had two owners: the parser branches and an English list hand-written beside them in `state.cts`. Nothing coupled them, so adding a branch left the message stale and removing one left it advertising a shape that errors, with no test able to see either. The message is now built from `STATE_FIELD_SCHEMA.current_plan.acceptedShapes`, and the CLI test walks the schema instead of restating the list. `Plan: N of M` is still spelled out explicitly because no schema row owns the body-only `Plan` field — `buildStateFrontmatter` never reads it into frontmatter, so it has no key to hang a row on. m1 — the property drove only the pre-existing `**Plan:**` branch, i.e. not the branch under review. It now drives both compound spellings and ranges past 99 so the width transition is covered by the property rather than one example. A second property covers the legacy pair's own preservation contract. Both were mutation-checked: dropping the padStart turns 9 tests red. m2 — degenerate boundary fixtures around the threshold (`0 of 0` is phase-complete, not an error; `0 of 3` advances) plus the shapes the anchored grammar must refuse, including Arabic-Indic digits. m5 — `docs/json-errors.md` described rather than quoted the message, since it is now schema-derived and a verbatim quote would be a third owner. Nits — the CRLF assertion could not see a `\n` at index 0; the `!/^Plan:/m` presence proxy is now an identity assertion on the whole `## Current Position` body. * fix(#3784): give the section plan write its own flag, and stop narrowing what parses Review round 4 on #3791: Blockers 1-4, Majors 1-2, Minors 1-2. B1 — the section fallback was guarded by `!mutated`, and `mutated` is FUNCTION-wide, already set by the phase/status/lastActivity arms that `advancePlanCore` always populates. A section spelling the field bold or as a pipe-table row therefore skipped its fallback because an UNRELATED field had been refreshed, and stayed a plan behind the header — the split-brain document this arm exists to prevent. The arm now tracks its own `planWritten`. Worth recording: the reviewer's fixture does not reproduce. The body-level status write lands on the section's own `Status:` when the document has no header `Status:`, so `mutated` is still false by the time the plan arm runs and the fallback fires. The discriminating shape needs a header `Status:` to absorb that write AND a bold section plan line. The mechanism was right; the example was not, and the regression test uses the shape that actually fails. B2 — `fallbackName` chose one name by ternary. In the legacy shape both values are populated, so it always chose `Current Plan` and a `**Plan:**` section line — which base did write — got nothing. Each name is now attempted independently with its own fallback. B3 — `PLAN_SHAPE_N` was anchored harder than `PLAN_SHAPE_N_OF_M`, so values base parsed via `parseInt` began to hard-error: `Total Plans in Phase: 5 phases`, `Current Plan: 3 (blocked)`. #3784's brief puts normalizing plan numbers beyond this transition's read/write out of scope, so that narrowing was not licensed. Both grammars now carry the same trailing tolerance. The prose defect stays closed by the START anchor, not by forbidding suffixes. Major 1 — the whole-body `Plan` write is scoped to documents that declare a `Plan` field, instead of firing unconditionally where `stateReplaceField`'s first match could be prose outside `## Current Position`. Major 2 — the error message names both `Plan` spellings the parser accepts; it previously omitted the sibling-paired form, which is the same message-disagrees-with-parser drift the derivation exists to close. B4 — the changeset claimed a guarantee B1 broke; it now describes what ships. Minors — safe-integer boundary coverage at limit-1/limit/limit+1, and the CRLF comment states the real mechanism (`stateExtractField`'s `(.+)` stops before the CR; the trailing group is belt-and-braces, not the primary defence). All three blocker regression tests verified red against the pre-fix source. * test(#3784): pin the hybrid shape against #3807's ambiguity refusal #4028 landed `advance-plan`'s multi-`Phase:` refusal on `next` after this branch's last run, on the same function. The guard sits above the parse, so a refused document is never parsed and the shape #3784 adds cannot reach the mutation — but that is a property of source ordering, so assert it as behaviour instead. Fail-first proven, not assumed: with `phaseCandidates.length > 1` disabled, the ambiguous hybrid document advances its FIRST entry's `Current Plan: 04 of 06` to `05 of 06` and writes it — #3807's exact defect, reached through #3784's shape. Both tests go red; both go green with the guard restored. The control pins the other direction: an unambiguous hybrid section still advances, and its zero-padding still survives. * fix(#3784): advance every spelling from its own text, refuse when they disagree Round 6 review. B1 and M1 are one defect, so they are one fix. `advancePlanCore` picked one field to parse from, computed `newPlan`, then wrote BOTH spellings from that field's numbers. Two symptoms: B1 With `Plan` as the parse source, `Current Plan` was re-stamped with the number just derived from `Plan`. `Current Plan: 7` beside `Plan: 2 of 5` silently became `Current Plan: 3` — a value nothing derived for that field, no error, no diagnostic. M1 With the legacy pair winning, the `Plan:` line was re-rendered from a bare `${newPlan} of ${totalPlans}` built out of the sibling field. `Plan: 2 of 9` became `3 of 5`; `Plan: 03 of 05` became `4 of 5`. The changeset's claim that padding and everything after it survive was true only for whichever field happened to be the parse source. Now: every spelling is advanced from its own raw text via `bumpLeadingNumber`, so each keeps its own padding, its own total and its own trailing annotation. Differing TOTALS are preserved, not reconciled — `Plan: 2 of 9` beside a `Total Plans in Phase: 5` advances to `3 of 9`. Differing CURRENT numbers are refused, with `reason: "ambiguous_plan_position"` and both candidates named. Same posture as #3807's multi-`Phase:` guard one field over: name the conflict, let the caller resolve it, never pick. The guard sits immediately after the parse, BEFORE the phase-complete branch — guarding only the normal advance would let `Current Plan: 7` beside `Plan: 5 of 5` write a terminal "Phase complete" into a document whose two spellings never agreed. A field present but unreadable (`Plan: TBD`) is left exactly as authored. Refusing the whole document because an unrelated line cannot be read would be a narrowing #3784 does not license; writing a derived number over it is the fabrication B1 was filed for. The `planSourceField`/`planRawValue`/`useCompoundFormat` tracking is gone. It existed only so the write path could ask which field the value came from, and the write path no longer asks. M2. The bare `Plan: N` + `Total Plans in Phase: M` shape is dropped. A revision of this PR added it; base refused it. It cannot be given the schema-row + forcing-test coupling the other shapes have, because `Plan` is body-only and `buildStateFrontmatter` never reads it into frontmatter, so there is no `current_*` key to hang a row on. Parser, the spelling in `advancePlanShapeError`, and the lockstep test move together — the invariant is the lockstep, not the length of the list. N1. The whitespace narrowing (`5phases` no longer parses where `parseInt` read 5) is documented in the changeset beside the other deliberate narrowings, rather than loosened. Loosening restores the half-parse this change exists to remove. Tests: eight new cases plus a property that crosses the two spellings with agreeing and disagreeing numbers — the review noted the existing properties never did. Fail-first proven: restoring the old write path reddens seven of the eight, both new property arms, and two pre-existing padding tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U * fix(#3784): report Current Plan as updated only when it was written The write became conditional in the previous commit — a `Current Plan:` that is present but unreadable is left as authored — but the `updated` push stayed unconditional, so `transitionCore` reported a field it had not touched. `reconcileReportedFields` would have caught it against the persisted bytes at the `state.cts` caller, but `transitionCore`'s own `updated` is consumed directly (milestone-lock, the transition tests) and has to be true on its own. Covers the mirror of the unreadable-spelling case: `Current Plan: TBD` beside a readable `Plan: 2 of 5`, where `Plan` is the parse source and the legacy field is the one that cannot advance. Fail-first proven. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U * test(#3784): account for the new refusal in the output({error}) census `tests/io.test.cjs`' A3 census asserts the exact population of `output({error})` call sites in `src/`, per module. The `ambiguous_plan_position` refusal added a 27th to `state.cts`, so the census went red at 26/65. Updated the way #3807 updated it when it added the ambiguous-POSITION error one line above: bump the count and name the addition inline, so the next person reads why the number is what it is. The alarm did its job — it is the only gate that noticed a new user-visible error path had been introduced. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01H3eK225hgcnEDZsnmtaP1U --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
41466e8e88 |
fix(#4023): preserve decimal phase ids in init progress ordering and smart-entry output (#4110)
* test(#4023): reproduce decimal phase-id coercions * fix(#4023): preserve decimal phase ids in progress signals * test(#4023): align phase token contract expectations * chore(#4023): point the changeset at PR #4110 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
472f585f7c |
fix(#3726)!: require --confirm before milestone complete mutates (#3774)
* fix(#3726): require --confirm before milestone complete mutates `milestone complete <version>` is a one-way door — ROADMAP.md and REQUIREMENTS.md archived, every phase directory in the milestone MOVED, STATE.md rewritten — and ran unconditionally on first invocation through every invocation path, including `query milestone.complete <version>`, whose `query` meta-prefix reads as a read-only namespace but performs no filtering (#167's invocation-compatibility shim + #3243's dotted-form normalization). The gate lives on the destructive command itself, not on the `query` prefix (the prefix is an intentional invocation mechanism, not a permission boundary — restricting it would break dozens of shipped workflow callers). Without --confirm and without --dry-run the command now refuses via error() before reading anything beyond its arg checks, so an unconfirmed invocation is a guaranteed no-op on disk. --dry-run still previews with no confirmation needed and is now documented in the usage block (it was only documented for the sibling archive-quick). --force keeps its narrow meaning — bypassing the TRUNCATED-scope and unstarted-phase guards — and does not double as the mutation opt-in. --confirm follows the existing `phases clear --confirm` idiom in the same module. complete-milestone.md's two invocations pass --confirm (the workflow has gathered explicit user intent by that step). Existing tests get --confirm appended — pre-change behavior is exactly confirmed behavior — and a #3726 regression block covers: refusal + full-tree byte-identity on both invocation forms, --force not satisfying the gate, --dry-run still passing without confirmation, and --confirm proceeding. The refusal tests fail against pre-fix code (negative control run). Fixes #3726 * docs(#3726): document the --confirm requirement in CLI-TOOLS and COMMANDS Cross-AI review of the fix diff (codex, pre-create) caught three shipped doc sites still instructing the now-refused bare invocation: the CLI-TOOLS.md milestone-complete synopsis + flag table, and COMMANDS.md's two guard-override instructions (`--force` alone now refuses without --confirm). Localized CLI-TOOLS copies already lag the English synopsis (no --force/--dry-run either) and follow the translation pipeline, not this fix. * chore(#3726): set changeset fragment pr to 3774 * test(#3726): confirm-gate CI repairs — QA scenario caller + growth ack Two CI reds from the --confirm gate, both this branch's own misses: - tests/qa/scenarios/milestone-rollover.json invoked `milestone complete 1.0 --force` as a JSON arg-array fixture — a caller shape the test sweep (which grepped runGsdTools/runSdkQuery in tests/*.cjs) never enumerated. Adds --confirm; the scenario's boundary-crossing contract is otherwise untouched. - complete-milestone.md's +420-byte --confirm note trips the emitted-attribution growth ratchet. Acknowledged as a #3726 append to the existing complete-milestone.md entry in 3409-unreachable-guard-arms.json (two ack sources may never name the same path, per that fragment's own precedent). Local: lint-emitted-drift-ack ok; loop-walk.qa 115/115 green sandboxed. * docs(#3726): CLI-TOOLS.md guard-override sentences say --force --confirm Review Major 1: the truncated-window and unstarted-phase guard paragraphs still told the reader to "Pass `--force` to override", which now refuses (--force alone does not satisfy the confirmation gate), while the flag table 470 lines later said the opposite. Mirror the docs/COMMANDS.md pair so the file no longer contradicts itself. * docs(#3726): synopsis renders --confirm and --dry-run as alternatives Review Nit 1: `milestone complete <version> --confirm [--dry-run]` read as "a dry run still needs --confirm", the opposite of AC 3. Render the pair as `(--confirm | --dry-run)` in the CLI-TOOLS.md synopsis and the usage docblock, and let the flag rows carry the rule. * test(#3726): pass --confirm in base-added milestone fixtures; re-file the growth ack Rebase onto next (26 commits) surfaced three tests the gate now refuses: the #3685 write-flag contract pair in tests/milestone.test.cjs and the `milestone complete` boundary fixture in tests/state-contract.test.cjs all invoke the command bare. Each now passes --confirm (a mutating run is exactly what they assert on). The +420 byte complete-milestone.md growth ack rode on 3409-unreachable-guard-arms.json, which #3078 swept from next as fully spent — hence the modify/delete conflict. Re-filed under a fresh fragment named for this issue, never resurrecting the swept one. * test(#3726): pin the present-but-falsy arm of the confirmation gate Review Minor 1: the boundary triple covered absent and present but not present-but-falsy. The gate is an exact-token match, so --confirm=false and --confirm=0 refuse today — pinned (canonical + query forms, whole .planning/ tree byte-identical) so a future `=`-aware or prefix-matching parser cannot silently turn --confirm=false into a confirmed run of an irreversible command. * test(#3726): drop --confirm from dry-run-only invocations Review Nit 2: --confirm was mass-appended to 14 pre-existing --dry-run invocations that never needed it, so each stopped standing as incidental proof that a preview needs no confirmation. Reverted to the pre-PR form; the dedicated AC-3 test carries the explicit assertion. * docs(#3726): sync the localized CLI-TOOLS synopsis with the confirm gate REQ-I18N-02 (docs/features/internationalized-documentation.md) requires translations to stay synchronized with the English source. The four localized CLI-TOOLS.md guides still advertised a bare `milestone complete <version>`, which now exits 1. Render the English synopsis verbatim — `(--confirm | --dry-run)` plus the `[--force]` and `[--archive-quick]` flags the translations had also fallen behind on. * test(#3726): drop --confirm from the remaining preview-only invocations Round 2 reverted the --confirm appends on --dry-run-only invocations in tests/milestone.test.cjs, but four more sat in two files the sweep missed: tests/milestone-archive.test.cjs (three) and tests/milestone-window-single-owner.test.cjs (one). Each is a preview run whose whole purpose is to document that a preview mutates nothing, so `--dry-run ... --confirm` contradicted the semantics the test exists to pin. Dropping the token restores each as incidental proof that a preview needs no confirmation; the dedicated AC-3 test keeps the explicit assertion. No assertion added, relaxed, or removed — the change is four tokens. * chore(#3726): migrate the emitted-drift ack from a fragment to a commit trailer #3954 (ADR-3942) moved emitted-drift acknowledgments out of tests/emitted-drift-acks/ and into git commit trailers, and the fragment directory no longer exists on next. The reason this PR's fragment carried moves verbatim into the Emitted-Drift-Ack-Growth trailer on this commit; the fragment file is removed rather than resurrected. Emitted-Drift-Ack-Growth: complete-milestone.md — #3726: +420 bytes (40186 -> 40606). The archive_milestone step's two `milestone complete` invocations now pass the required --confirm flag (the command refuses to mutate without it — the archive is irreversible), with a note explaining the flag and pointing at --dry-run for previews. Deliberate runtime-loaded workflow text for the new gate, not converter drift. * fix(#3726): name --confirm in the version-required refusal The documented arg-discovery path (gsd-tools.cjs top-level usage: invoke the command without args and the error lists what is required) stopped at `version required for milestone complete (e.g., v1.0)` — one required argument short. Discovering --confirm took a second round trip through the gate. The refusal now reads `… — and --confirm to mutate`, pinned by a test that also asserts the version-less invocation leaves .planning/ untouched. * test(#3726): pin the milestone complete docs against a silent regression The changeset is `type: Fixed`, which the docs-required lint exempts, so nothing in CI would notice a later edit that reinstated the bare-`--force` override prose or dropped `--confirm` from the synopsis. Four tests in tests/milestone.test.cjs now pin: the synopsis line in docs/CLI-TOOLS.md and its four localized mirrors; the `--confirm` flag row; both guard-override instructions in docs/CLI-TOOLS.md and docs/COMMANDS.md, by guard name (a substring match on each instruction's `--force --confirm` text); and — as an identity ratchet over the milestone-complete sections — every `--force` sentence or clause that lacks `--confirm`, so a new bare instruction in its own sentence or clause fails whatever its wording. Named residual: a bare instruction spliced into the same clause as a compliant one coalesces with it and passes the ratchet; the by-name pins are what keep the four known instructions from losing the pairing that way. The file is registered in scripts/docs-guard-registry.cjs so the pin runs on the PR that changes those docs, not only after merge. --------- Co-authored-by: CI Rebase Check <ci@gsd-redux> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
ac7587287b |
fix(#3812): document how Current Position actually resolves a duplicate field (#4017)
* docs(#3812): say that Current Position is single-valued, and pin the behavior that makes it true #3812 shipped CLOSED with half its acceptance unmet. #3873 delivered cardinality for FRONTMATTER keys - current_phase/current_plan render as optional at docs/reference/state-md.md:89,91, covered by tests/gen-state-md-docs.test.cjs:374. The issue's actual ask was the ## Current Position BODY section, and that never landed. Surfaced by an /adr-phase-coverage audit of epic #3473; the issue was reopened rather than noted. The section now states three things: every field is single-valued, the section is overwritten rather than appended to, and a duplicate resolves to the FIRST occurrence with no warning - so a line appended in good faith is silently ignored rather than winning. Progress history belongs in ## Performance Metrics, two headings down, and the text now points there. The third claim is a behavioral promise about the reader, so it was VERIFIED BY EXECUTION before being written rather than inferred from the issue title: stateExtractField(<"Phase: 1 of 5 (First)" ... "Phase: 9 of 9 (Appended later)">, "Phase") -> "1 of 5 (First)" The mechanism is state-document.cjs:405 - the plain-line pattern ^<field>:[ \t]*(.+) carries flags im with NO g, so String.match returns the first hit. Writing "first wins" without running it would have repeated the exact error I had to retract twice in this epic already. A test pins the reader, not the prose. Three rows in tests/state.test.cjs: T1 (load-bearing) asserts the duplicated case resolves first; T2 asserts the ordinary single-field case still works, so a fix that only functions when duplicated cannot pass; T3 puts a Plan: line BETWEEN the two Phase: lines and asserts it resolves independently - negative space, because a reader returning the first line of the SECTION rather than the first matching FIELD would satisfy T1 alone. Proven to discriminate: a last-match variant returns "9 of 9 (Appended later)" and T1 reds. No assertion checks that the document contains a sentence. That is what local/no-source-grep exists to stop, and it would pin wording that is allowed to improve. The point of the test is that if that regex ever gains g and a last-match walk, the test fails - instead of the documentation quietly becoming a lie with nothing to notice. Prose only, no new heading. docs-state-md-locale-parity compares heading-level sequences by LCS rather than text, so added paragraphs cannot fail it while an added HEADING would fail all four locales. The constraint is structural, not stylistic - confirmed by running that comparison after the edit. The four locale copies are translated rather than left stale. They are not gate-enforced for prose, so "nothing fails" was available and is not the same as correct: leaving four documents asserting something the English one now contradicts is a correctness problem. Code spans and the anchor link stay untranslated - they name real tokens. The whole approach rests on one fact, checked first: ## Current Position at :196-208 sits OUTSIDE every generated marker region (:81-104, :138-151), so a hand edit survives --write. Re-confirmed after all five edits - gen-state-md-docs --check reports all 6 targets up to date. Had that been false the fix would have belonged in the generator, and a hand edit would have been silently reverted. One real gate failure fixed inline rather than reported: the new test's comments referenced docs/reference/state-md.md, which was not in that file's registered exempt-docs paths, and lint-docs-guard-registration failed lint:ci correctly. Registered. Known limit, named rather than folded in: gsd-tools validate/health still do NOT warn on a duplicated Phase:. #3812 records that as a "consider", not a requirement, and confirms none of the nine rules in src/health-diagnostic-rules/{state-consistency,phase-structure}.cts counts occurrences. Documenting the silent first-match is the delivered scope; making it loud is new scope and stays unclaimed. Closes #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3812): the rule I documented was false — replace it with the measured one An isolated review returned two blockers. Both mine, and the first is the worse kind: I wrote a falsifiable rule into a reference page and got it wrong. 1. "A duplicate resolves to the FIRST occurrence" is FALSE. stateExtractField (src/state-document.cts:401-419) tries BOLD `**F:**` across the whole input, THEN plain `^F:`, THEN a pipe-table row. Form precedence beats document order. Measured against the built reader, all intra-section: Phase: A (plain) / **Phase:** B (bold, later) -> B LATER WINS Phase: A (indented) / Phase: B (plain, later) -> B LATER WINS Phase: A (plain) / | Phase | T (table) | -> A first wins My original verification tested plain-versus-plain, saw first-wins, and generalized to all forms. Measuring one case and claiming the general rule is the same error I have had to retract twice already in this epic. It is also worse than silence. The sentence told authors an appended line is safely ignored; a bold line appended "for emphasis" silently overrides the original. Someone trusting the doc would have corrupted their own state file. And #3812 never asked for a resolution rule - it asked for single-valued, overwrite-not-append, and where history goes. The rule was my unrequested addition. Replaced with the measured truth: resolution is by FORM (bold anywhere, then plain at line-start, then table row), and only WITHIN the winning form does the first occurrence win. Both consequences stated plainly - a higher-ranked form wins regardless of position, and an indented `Phase:` is invisible to the plain form. All five claims in the new paragraph verified by execution before being written, including the two I had wrong. 2. The tests tested the wrong case and passed for the wrong reason. T1/T3 put the second `Phase:` under `## Somewhere else` - the INTER-section case, which #2956 already fixed by scoping. #3812 says verbatim that #2956 "fixed the inter-section case and never addressed intra-section duplication", so the case the new prose describes was untested, and the fixtures passed because of section scoping rather than field resolution. They also called bare stateExtractField rather than the production chain, T2 could not discriminate first from last at all, and no fixture mixed forms - which is precisely why the false claim survived to review. Rewritten as four rows, all intra-section, all through the real stateCurrentPositionSlice -> stateExtractField path: plain-then-plain (first wins within a form), plain-then-bold (the bold LATER value wins - the row whose absence let the false claim ship), indented-then-plain (indented invisible), and sibling-field independence. Each proven to fail against a reader that disagrees. 3. Two dead anchors. pt-BR and zh-CN linked `#performance-metrics` while their own headings are `### Métricas de Desempenho` and `### 性能指标`. Both fixed to the anchor their own heading generates. ja-JP/ko-KR kept the English heading, so theirs already resolved. 4. A ja/ko sentence inverted its own meaning. Both rendered "which is the section designed to grow" with a bare demonstrative whose nearest referent read as Current Position - saying the opposite of the point. Rewritten so the clause attaches unambiguously to `## Performance Metrics`. 5. Cross-locale drift, flagged by the implementing agent rather than by me: after fixing EN, the four locales still stated the OLD false rule. Four documents asserting something measured to be wrong is worse than four saying nothing. All four now carry a faithful translation of the corrected paragraph, with code spans, each file's own anchor, and the ja/ko referent fix preserved. Verified: all five claims executed against the built reader; every rewritten test row proven to discriminate; gen-state-md-docs --check reports all 6 targets up to date, so the edits stay outside the generated marker regions; locale heading parity unaffected (prose only, no headings added); build:lib, lint and lint:ci all exit 0. Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3812): second false rule on the same page — scope the ranking to the section A second isolated review found a second false falsifiable claim, and the failure mode is the same one twice in a row: attempt 1: verified plain-vs-plain, wrote a claim about ALL FORMS attempt 2: verified bare stateExtractField, wrote a claim about THE DOCUMENT Both times the claim covered a wider surface than what was actually executed. The fix each time was not a better sentence, it was executing the surface the sentence describes. BLOCKER — "bold `**Phase:**` anywhere in the DOCUMENT wins" is false. ## Current Position / Phase: 1 of 5 + ## Archive / **Phase:** 88 bare stateExtractField(whole doc) -> "88 (other section)" PRODUCTION (slice then extract) -> "1 of 5 (in section)" #2956's section slice means production never hands another section to the matcher; a bold line in `## Archive`, or in the YAML frontmatter, is simply not seen. The ranking is real but scoped: it applies WITHIN `## Current Position`. I verified against the bare function and wrote a claim about the system. Every existing test placed its bold line inside the section, which is exactly why nothing contradicted the claim. T5 now puts a bold `**Phase:**` in `## Archive` and asserts production returns the in-section plain value, with the unscoped reader asserted to DISAGREE so the row proves the scoping rather than assuming it. BLOCKER — the changeset still shipped the ORIGINAL retracted claim. I corrected the page and left the release note saying "resolves to the first occurrence ... a second entry added in good faith is silently ignored". The note contradicted the page it announces, and the release note is what most people actually read. Rewritten to the corrected rule. MEDIUM — the concession was inverted. It read "wins even if it comes FIRST in the file", which is the vacuous direction; the surprising case, and the one the very next clause illustrates with an APPENDED bold line, is "even if it comes LAST". All four locales reproduced the inversion faithfully, so it was an EN-source defect rather than translation drift. Two sharp edges now named, both measured: a bold `**Phase:**` followed only by trailing spaces resolves to an EMPTY STRING and does not fall through to a valid plain line below (T6 pins it); and `| **Phase:** | 3 of 4 |` short-circuits to the bold form and returns the literal `"| 3 of 4 |"`. A page that teaches form ranking has to say where the ranking bites. Also fixed: all five files labelled the link `## Performance Metrics` while the heading is `### Performance Metrics`. Anchors resolved correctly everywhere; only the label's level was wrong. Every clause in the final paragraph re-verified through the PRODUCTION chain (stateCurrentPositionSlice -> stateExtractField), clause by clause, before being written: bold in another section does not win; bold in frontmatter does not win; bold appended last does win; first wins within one form; trailing-space bold yields empty. All four locales carry the same corrected rule. gen-state-md-docs --check reports all 6 targets up to date; heading counts unchanged at 20/20 across all five files, so locale heading-parity is untouched; build:lib, lint and lint:ci all exit 0. Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3812): backfill changeset pr number Refs #3812 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
15af0f5536 |
enhance(#3951): B6+B7 — widen two unreachable lint rules and make the guard ledger true (#3965)
* fix(#3951): two lint rules that could not reach the code they govern B6 names two widenings. Measuring them first turned up a defect the criterion did not know about, and refuted the reason it gave for one of them. 1. no-adhoc-markdown-parsing self-gates on its own filename. Lines 107-110 short-circuit create() to {} unless the path matches /(?:^|\/)src\/[^/]+\.cts$/. B6 says to widen the files: glob in eslint.config.mjs - but doing only that ships an INERT rule, because the gate still returns {} for every new path. Both halves have to change, and the gate is the load-bearing one. That same regex hides a live hole: [^/]+ is FLAT-ONLY, so it requires the file to sit directly in src/. The registered glob is src/**/*.cts, which includes subdirectories. 28 .cts files - health-diagnostic-rules/ (10), installer-migrations/ (11), observability/ (3), host-integration-adapters/ (2), vendor/ (2) - are inside the registered glob and silently skipped. Measured with the gate neutralized: 0 violations there today. The hole is hiding nothing right now, and is fixed anyway, because "no violations today" is not a property that keeps holding. The fix is not invented: require-subprocess-timeout.cjs:196 already carries the correct form of this guard, /(?:^|\/)src\/.*\.cts$/ with .*, one directory over. Checked the other 21 rules for the same bug - no-adhoc-regex-escape and no-private-binary-resolution short-circuit only to exempt their own seam file, which is the right shape, and no-crlf-fragile-split has no filename gate at all. This bug is unique to the one rule. 2. no-adhoc-regex-escape could not see the shape that actually occurs. Line 396 gated the whole UNSAFE-NEW-REGEXP arm on arg.type === 'Identifier'. Every check below it - the _SOURCE provenance check, the isSoleReturnOfOwnParameter shape - lives inside that branch, so new RegExp(obj['key']) and new RegExp(cfg.pattern) were never examined at all. Runtime data arrives as a property access far more often than as a bare identifier, which is exactly why this rule never fired on the #3477 ReDoS. Widened to MemberExpression, measured by AST walk across all five registered blocks rather than by grep. 27 sites, zero TSAsExpression: 18 safe new RegExp(X.source, flags) -> exempted, keyed strictly on the PROPERTY being `source`, never on the object. Keying on the object would wave through X.anything and buy nothing. B6 estimated ~10; that was an undercount. 3 _SOURCE-suffixed constants reached through a required module namespace (phaseId.BRACKET_PHASE_TOKEN_SOURCE) -> the same provenance-exempt class the rule already recognizes for bare identifiers, extended to reach them. Without this the widening produces 3 false flags. 6 real findings -> marked, each a test extracting a pattern from a shipped file at test time, where the runtime contract IS the product. Deliberately the NARROW MemberExpression form. The rule's own isSoleReturnOfOwnParameter doc comment records that an earlier broad "any non-literal identifier" heuristic produced ~25 false positives and was rejected; a re-run of the census after this change flags exactly the 6 above and nothing else. Verified by execution, not by reading: the gate now accepts src/<subdir>/x.cts, still accepts flat src/x.cts, and still exempts paths outside src/ - each pinned by a test proven to fail against the old regex. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): give no-adhoc-markdown-parsing its reach, and fix the 80 parses it finds The rule self-gates on filename AND is registered on one glob, so widening either half alone is inert. Both move here: the gate now accepts tests/**/*.cjs and scripts/**/*.cjs alongside src/**/*.cts, and eslint.config.mjs registers it on the same two. A test pins that the gate and the registration AGREE, in both directions. The original defect was a gate narrower than its registration; the failure mode of this fix is a gate wider than its registration. Both are silent, so the test asserts the pair rather than either half. 80 violations across 43 files, all in tests/, zero in scripts/. 70 are routed through the existing seams - scanFencedBlocks, collectSection, stripFencedCode, tokenizeHeadings from markdown-sectionizer; splitTableRow, parseMarkdownTable, findTableWithColumns from markdown-table. Headerless STATE.md tables use splitTableRow per line, because parseMarkdownTable needs a real delimiter row. 10 are suppressed, 12.5%, well under the third that would have meant the rule is mis-scoped for tests/ rather than the tests carrying debt. Each names its reason: three regression guards (#3873 / bug-#21) are deliberately independent of the generator's own fence handling, and routing them through the seam would have them test the generator against itself; one is a negative-text probe that extracts nothing; six are a shell-pipe-to-jq detector whose regex coincidentally matches the table fingerprint and is not markdown parsing at all. All ten sit in tests whose subject is .md content, which is normally a reason to prefer the seam. The marker used is allow-adhoc-markdown, distinct from no-source-grep's allow-test-rule, and lint:ci's lint-allow-test-rule-refs reports the same 280/280 unverified count as before - checked rather than assumed, because those two markers are easy to conflate. The widening earned its keep immediately: it found a test that passed for the wrong reason. tests/config-field-docs.test.cjs asserted notEqual(<cell>, '600') against the TYPE column instead of the DEFAULT column. notEqual('number', '600') is true forever, so the guard against workflow.subagent_timeout regressing to the old seconds default could never fire. docs/CONFIGURATION.md:434 is `| workflow.subagent_timeout | number | 300000 | ... |`, so the default is cell index 2; the assertion is now row-scoped through splitTableRow and reads 300000. That is the argument for the widening in one case: the violation was invisible to lint, the suite was green, and the assertion was vacuous. A rule that cannot reach a file cannot tell you the file is lying. Not fixed here, and recorded rather than assumed: #3426/#3239 are NOT reachable by this widening. tests/package-legitimacy-gate.test.cjs yields zero violations even with the gate bypassed - its hand-rolled scans are real, but built from line filters and split('|') rather than the regex-literal fingerprints this rule detects. They need new detectors. The epic assumed a wider glob would catch them. build:lib, lint and lint:ci all exit 0; the post-fix census across tests/** and scripts/** is 0 violations. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3951): B7 — and #3356's defects were still live in the code B7 asks that each closed child be driven fail-first with a behavioral identity test at the CONSUMER's output. Four of eleven children had no test citing their issue number. Auditing them by BEHAVIOR rather than by number-grep changed the answer for three of the four. #3364 and #2540 — traceability only. Both were implemented by #3941 and their consumer-output tests exist and were shown failing-first; neither cited its originating issue, so an audit that greps for the number reports them uncovered. Tagged the specific asserting test in each file, following the citation form those files already use. #3372 — covered, but only at helper level, and the triage narrowed it. Of the four commands the issue names, only estimate-cli's collectCalibrationSamples actually enumerates phase dirs from disk; smart-entry, audit and roadmap-upgrade derive from ROADMAP/body text and never reach the sentinel path, so they are benign by construction and were left alone rather than "fixed" into churn. The existing #3882 rows asserted the helper's return value. Added a consumer-output test driving `query estimate-calibrate` and asserting sample_count and the persisted document. RED proof: reverted collectCalibrationSamples to a raw readdirSync and ran the real CLI - sample_count 3, sentinel leaked; restored - sample_count 2. #3356 — NOT covered, and BOTH halves of the defect were still live in source. The issue is closed; the bug was not fixed. Fixed here rather than writing tests that document a bug as correct. Defect 1, the contradicted row. quick.md:627 claimed `quick-tasks-append` performs "the equivalent write" to the Step 7c row. It did not: the `#` cell was a positional ordinal and `Directory` read `—`, because the route had no way to receive a quick id or task directory. Added OPTIONAL `--quick-id` / `--slug` / `--directory`. A caller with neither - fast.md, the original #2133 caller - omits them and gets the byte-identical prior row, so nothing existing changes. A caller that HAS a real id and directory now gets the canonical row quick.md:632 renders. The false-equivalence sentence itself is corrected rather than left to mislead the next reader. Defect 2, the forced re-derive. The route called readModifyWriteStateMd with no options, so a body-only append to the Quick Tasks table triggered a full re-derive of the disk-derived progress.* frontmatter. Every other body-only writer passes { resync: false } - src/state.cts's own docstring prescribes it - and this route was the lone outlier. RED proof: reverted the option, seeded a project with 2 real phase dirs and a curated total_phases of 25, ran quick-tasks-append; total_phases collapsed to 2. Restored; it stayed 25. That second one is the shape this epic exists to close: a silent write that replaces curated state with a re-derivation nobody asked for, exit 0 throughout. build:lib, lint and lint:ci all exit 0. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3951): amend B6's ledger to what was measured, and document the new flags The ADR gains a ledger amendment in its own correction style - the sixth wrong premise it records, found the same way as the other five, by measuring before building. B6 says the net guard count must fall. It rose: 62 -> 69, +7, measured from the epic's filing commit to origin/next. The attribution is the point, though. Five of the seven came from PRs unrelated to this epic, one was added by a phase of it, and the epic did retire something sub-file - #3884 removed a detector with an explicit "net: -1 detector, 0 added" ledger. Every named casualty is load-bearing, two already carry retractions in this same document, and a sweep of all 22 rules plus every scripts/lint-* found no provably dead guard. There is no honest way to make the count fall; forcing it would trade coverage for a number, which is the Goodhart outcome Decision 6 exists to prevent. The amendment also records that B6's own prescribed fix for one widening was inert. no-adhoc-markdown-parsing self-gates on its filename, so widening only the files: glob - which is what the criterion says to do - ships a rule that still returns {} for every new path. And #3426/#3239 are not reachable by that widening at all; their scans use line filters and split('|'), not the regex fingerprints the rule detects. The roster row tracked them against the wrong mechanism. Three roster rows updated from aspiration to fact: the two widenings are DONE with their measured counts, and lint-phase-enumeration-drift is marked RETAINED rather than "expected casualty - verify before retiring", because Phase 5 verified it and kept it. The rule Decision 6 should carry forward is stated plainly: a guard ledger is a claim about COVERAGE, not about COUNT. "Net count must fall" is measurable and wrong. "Every guard is reachable, and each retirement names what makes its defect unrepresentable" is the property that was actually wanted. CLI-TOOLS.md documents the optional --quick-id/--slug/--directory flags and says plainly that omitting them keeps the pre-#3356 row byte-identical, plus that the append no longer re-derives progress frontmatter. New features fragment (id 3951); FEATURES.md regenerated rather than hand-edited. Changeset is Changed, pr:0 pending backfill. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): correct four rows that pinned the lint rule's old narrow reach The remote suite came back RED with 5 failures, all in tests/eslint-rules.test.cjs. They are stale tests, not a regression: four rows assert that no-adhoc-markdown-parsing is inert outside src/*.cts, which is exactly the contract this deliverable changes. Confirmed by reading rather than inferred from the names - the row at :1981 used filename: 'tests/some.test.cjs' and filename: 'scripts/helper.cjs', the two roots the rule now covers on purpose. Worth recording WHY local gates missed this. npm run lint and lint:ci were green, and the touched test files passed standalone. Lint only reports violations in real files; these rows assert the rule's REACH using synthetic RuleTester filenames, so nothing but the full suite could see them. Local green on a rule change says nothing about the rule's own tests. Each row is rewritten with BOTH halves rather than flipped from valid to invalid: - the same fingerprint under tests/ or scripts/ is now flagged, with the right messageId - the negative space is preserved - the same fingerprint under a path outside all three roots (gsd-core/bin/lib/foo.cjs) is still NOT flagged The second half is the one that matters. Without it the rule has no boundary and nothing would catch an over-wide gate later, which is the mirror image of the bug this deliverable just fixed. Each row is renamed to state the current contract; the old names said "non-src/*.cts ... is not flagged" and would have been actively misleading once the bodies changed. Proven to test the widening rather than restate it: every flagged half was run against HEAD~2's pre-widening rule and does NOT fire there, then against the current rule and does. 12/12 on that probe; the full file is 178/178. Swept for the same staleness elsewhere and found none. require-subprocess-timeout's own "inert outside src/*.cts" row is untouched - that rule's gate was not widened here - and no-adhoc-regex-escape's test file already carries correctly-targeted rows. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): acknowledge the quick.md growth the attribution guard reported The full suite came back RED with one failure, and it is mine: 1 file(s) grew without an acknowledgment: quick.md grew 364 bytes gsd-core/workflows/quick.md is runtime-loaded emitted content, so correcting its false 'performs the equivalent write' claim trips emitted-attribution by construction. This is the acknowledgment, not a workaround - there is nothing to regenerate. The fragment names ONE path, which is the only one the guard reported. The four spent acknowledgments it also listed (audit-uat, plan-phase, progress, review) belong to other fragments whose ripple the base already absorbs; they are inert, not failures, and are deliberately NOT copied here - naming paths I did not change would make this record false in the other direction. Byte figure corrected before committing: the guard reported 37220 -> 37584 (+364), but origin/next has since moved and quick.md is 37232 there now, so the measured delta is +352. The reason text says so and names the base as a moving figure rather than pinning a number that is already stale. Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3951): move the quick.md growth ack to a trailer, delete the obsolete fragment The acknowledgment mechanism changed under this branch. Merging next brought in the redesign - it also deleted .github/workflows/ack-fragment-sweep.yml, which was in the merge status and which I did not register at the time - and the guard now says so directly: Add a trailer to a commit in this PR (never a new file). Emitted-Drift-Ack-Growth: quick.md - <why this growth is deliberate> So tests/emitted-drift-acks/3951-quick-append-equivalence.json is obsolete on arrival. A fragment file is no longer read by anything, and leaving it would be a dead record that looks like an active one. It is deleted here rather than kept "just in case". The byte figure moved again with the merge: 37232 -> 37596, +364. The earlier fragment said +352, measured before the merge auto-merged quick.md itself. The trailer carries no number, which is the better design - the figure was stale twice in two attempts. Refs #3951 Emitted-Drift-Ack-Growth: quick.md — #3356/#3951 replaces a false claim with an accurate one. Line 627 said the `quick-tasks-append` shortcut "performs the equivalent write" to the Step 7c row rendered above it; it did not, and that was the documented half of #3356 — with no quick id or task directory the route emitted a positional ordinal in `#` and an em-dash in `Directory`, a visibly different row. The corrected sentence has to carry three facts the original elided: what the shortcut actually writes when it has neither input, that this is honest behavior for its real caller (`fast.md`, which has neither), and how a caller with both now gets the byte-identical canonical row via the new optional `--quick-id`/`--slug`/`--directory` flags. Prose is the product here — an executing agent reads this line to decide whether the shortcut is safe for its case, and a shorter correction would either drop the flags (leaving the reader unable to act on the fix) or drop the limitation (recreating the false claim in gentler words). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3951): backfill changeset pr number Refs #3951 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8b41d855e0 |
fix(#3742): path-shaped comment channel keys and post-restore propagation (#3952)
* test(#3742): frontmatter comment survival must not depend on the body or indentation * fix(#3742): path-shaped comment channel keys and post-restore channel propagation * chore(#3742): changeset fragment (pr number backfilled after PR creation) * chore(#3742): backfill changeset PR number (3952) * test(#3742): direct mutation-shard coverage for the nested comment channel --------- Co-authored-by: sim <sim@local> |
||
|
|
e20744eacb |
enhance(#3884): failure is a value — strict argv, and --pick that signals absence (#3922)
* test(#3884): failing-first coverage for strict argv and absence-signalling --pick ADR-3473 §8.4 says failure is a value. Three families currently encode failure as success, and this commit pins each one RED before the fix lands. Measured on this tree, 2026-08-26: gsd-tools generate-slug "test" --pick nonexistent -> empty stdout, exit 0 (#3365) gsd-tools audit-open --pick nonexistent_field -> dumps the entire human-readable audit report, exit 0 gsd-tools generate-slug "Hello World" --raw --pick bogus -> prints "hello-world", another field's value, exit 0 gsd-tools query state.planned-phase 3 (positional, no --phase) -> exit 0; STATE.md's "Phase: 2 of 5 (Widget Support)" is overwritten to "Phase: null - READY TO EXECUTE" and the frontmatter gains a corrupted current_phase_name (#3358) tests/pick-flag.test.cjs:27 previously asserted the #3365 defect as the contract ("returns empty string for missing field", success === true). That assertion is replaced by the required behavior rather than deleted. The new parseNamedArgs block calls the spec-object signature that does not exist yet, so it fails today by construction. The 11 existing behavior-lock tests are left untouched here; they are corrected in the implementation commit. C1/C4 assert at the consumer's output - STATE.md's bytes - per ADR-3180 Decision 4(b). A unit assertion on the parser would have passed throughout this defect's life. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3884): failure is a value — strict argv, and --pick that signals absence Implements ADR-3473 §8.4. Absence, emptiness and failure stop being interchangeable ways to say "I could not answer". parseNamedArgs (src/command-arg-projection.cts) Takes a spec object with a REQUIRED `positionals: number | 'rest'` and returns the hub's Result shape instead of a bare Record. Declaring the positional arity is what makes #3358's call site unrepresentable rather than merely detectable: an unrecognized flag or a token past the declared boundary is now InvalidArgs, naming the offending token and listing the accepted flags. The legacy positional-array call shape throws a TypeError — an internal invariant violation per ADR-3473 Decision 2, so a stale hand-written .cjs call site fails loudly instead of destructuring undefined off a Result. parseNamedArgsOrExit projects a failure onto the caller's error(); it is a projection over the one parser, not a second parser. Measured before, against a STATE.md with a populated phase-2 block: query state.planned-phase 3 (positional, no --phase) -> exit 0; "Phase: 2 of 5 (Widget Support)" overwritten to "Phase: null - READY TO EXECUTE", frontmatter gains a corrupted current_phase_name After: exit 1, `unexpected positional argument "3"`, STATE.md byte-identical. The flag form is unchanged and still updates STATE.md. --pick <field> (gsd-core/bin/gsd-tools.cjs) extractField returns {found,value}, and the pick block no longer shares one catch between "output was not JSON" and "field was absent". An absent field exits 1 with pick_field_absent, naming the field and the keys that do exist; non-JSON output exits 1 with pick_output_not_json instead of dumping the command's entire output. A field that is PRESENT with value null, '', 0 or false still prints at exit 0 — that is an answer, not a failure, and it is what keeps `--pick count` printing 0 on a fresh project. Measured before: `audit-open --pick nonexistent_field` printed the whole human-readable audit report at exit 0, and `generate-slug X --raw --pick bogus` printed "hello-world" — a different field's value, confidently, at exit 0. ADR-3409 Decision 7 explicitly deferred this contract fix to #3473; this is it. The sub-issue's "returns 0 when the count is zero OR absent" wording is superseded by the ADR rule it implements: zero prints 0, absence exits non-zero. Defaulting absence to 0 would demote "could not answer" to "the answer is zero" — the hazard docs/how-to/resolve-unreachable-guard-findings.md already warns against. Guard ledger (ADR-3473 Decision 6) scripts/lint-unreachable-guard-drift.cjs Detector A is RETIRED. Its premise — that a `--pick ... || echo` arm can never fire — is now false, so the shape it forbade is the correct idiom and keeping it would forbid the fix. Detector B (glob-consuming cat/ls, a nullglob mechanism this change does not touch) is retained in full, as are the shared scanner, the escape-marker parser and the baseline. Net: -1 detector, 0 added. The file is not deleted. Call-site audit 45 prompt-layer --pick invocations, every one a plain X=$(...) assignment — none in an if test, && chain, or a pipeline whose status is consumed, and no shell block in workflows/commands/agents/references sets -e. Of the 13 (command, field) pairs the prompt layer reads, 10 are always present; the 3 sometimes-absent ones each sit behind a prior found/existence check. No ADR-3409-class "field the command never produces" remains. Design: .gsd/phase/feat-3884-failure-is-a-value/40-design.md Test matrix: .gsd/phase/feat-3884-failure-is-a-value/50-test-matrix.md Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): escape untrusted tokens in diagnostics, and cover five unpinned rows Two review findings, both fixed here rather than recorded as limits. 1. A newline in an untrusted token forged a second stderr line. Before, plain-text mode: $ gsd-tools query state.planned-phase $'foo\nError: forged second line' Error: unexpected positional argument "foo Error: forged second line" After: Error: unexpected positional argument "foo\nError: forged second line" --json-errors mode was never affected — io.error runs that payload through JSON.stringify. Plain-text mode writes 'Error: ' + message verbatim, and the three new InvalidArgs reasons plus the two new --pick diagnostics all interpolate a token that comes straight from argv. Fixed with ONE shared helper, formatDiagnosticToken (src/io.cts), applied at every interpolation site — not a copy per site. It is deliberately NOT applied inside error() itself: several callers in this tree emit intentional multi-line diagnostics, and escaping newlines there would mangle them. The available-top-level-keys list needed the same treatment for a reason the review did not anticipate: `frontmatter get <file>` reads an ARBITRARY user document and echoes that document's own keys into the diagnostic. Verified reachable — a frontmatter key containing a newline reaches the key list — so formatKeyForDiagnosticList is guarding a live path, not a hypothetical one. Ordinary keys still render plain and unquoted; a fix that merely dropped the key would also have passed a "one line" assertion, so the test pins the escaped key's presence too. 2. Five behavior-table rows were implemented but nothing pinned them: B7 a dotted path that dies partway B9 bracket syntax on a non-array B10 a negative array index, in and out of range B14 a JSON root that is not an object B17 an @file: payload over 50KB B17 is the load-bearing one. output() writes @file:<path> instead of inline JSON past 50000 characters, and --pick resolves that BEFORE parsing; with no test, a future reordering of those two steps turns every large result into a false pick_output_not_json. The fixture seeds 1200 phase directories and measures the payload at 62474 characters, asserting the spill actually happened rather than assuming it. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): correct the strict-argv surface against a full verification run The first full run came back with 90 failures across 12 files, none in the new tests. They were the argv surface telling me what it actually is. Ten root causes; each classified before anything was changed. I over-implemented, and that is reverted. ADR-3473 §8.4 says parseNamedArgs rejects "unrecognized and positional tokens". It says nothing about a value flag whose value is missing. Making that an error was my design decision, not the rule, and it broke a deliberately recorded contract: `--prd` with no value resolving to null (tests/init.test.cjs emptyPrdValueIsFalsyAndTreatedAsAbsent, row B5; tests/section-manifest-init-facts.test.cjs "flag-shaped value"). The "requires a value" branch is deleted outright rather than kept behind an option — an unused strictness mode is speculative generality. Unknown-flag and unexpected-positional rejection, which is what §8.4 actually mandates, is unchanged. --wave needed a third flag kind the original design did not anticipate. `--wave N` is documented (commands/gsd/execute-phase.md:4,48) and the shipped workflow reconstructs and passes it (execute-phase.md:84), while #2932 records token-PRESENCE semantics: the CLI cares only that the flag appeared, and the value belongs to the workflow layer. That is neither a boolean flag nor a value flag, so `optionalValueFlags` now exists — presence-only in `data`, and the validation cursor consumes a following non-flag token so it is not reported as a stray positional. Every other declared boolean flag was checked against every argument-hint and prose usage in commands/, workflows/, agents/ and docs/; `--wave` is the only one of this shape. Five tests were pinning forms that never worked. tests/adr857-core-without-capabilities.test.cjs passed `init plan-phase --phase 01-stub`, but the documented form is positional (docs/CLI-TOOLS.md:776) and the handler reads args[2] — which for that form is the literal string "--phase". Measured on the pre-fix build against a real .planning/phases/01-stub/ directory: init plan-phase 01-stub -> phase_found=true init plan-phase --phase 01-stub -> phase_found=false The test asserted only exit 0 and key presence, so it had been green while proving nothing about phase resolution. Corrected to the documented form and strengthened to assert phase_found === true. Same class in state.test.cjs (`--plan-count`, a flag that does not exist; the real one is `--plans`), milestone-archive.test.cjs (`init new-milestone --json`, silently ignored), and concurrency-safety.test.cjs (a bare positional field name whose OR-assertion passed because a whole-document dump happens to contain the substring it looked for). Six handlers had no argv validation at all — the same #3358 shape this phase exists to close, found while fixing the rest: init verify-work / phase-op / review / todos / remove-workspace read args[2] with nothing checking the rest, and validate health read --repair/--backfill through a bare args.includes() scan that bypassed the parser entirely. All now go through the seam, so the flag has one owner. tests/init-debug.test.cjs rows C4/C5 asserted that an unrecognized flag must NOT fail. That is the behavior §8.4 removes, and Decision 8 says a caller's local expectation does not override §8, so they are inverted and renamed — a test still called "ignores an unrecognized flag" while asserting rejection would be its own defect. Row C6's point is its PWNED canary; that assertion is kept verbatim and only its exit-status expectation changed, because the hostile token is now rejected rather than absorbed. The blast-radius estimate in 40-design.md is corrected rather than quietly left wrong. get_impact reported MEDIUM / 8 symbols upstream, and that was accurate for what the graph can see — parseNamedArgs's callers. It cannot see that those callers' handlers accept argv shapes wider than the code reading args[2] suggests, which is where the real surface was. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3884): withdraw the validate-health tightening, finish the A2/A3 revert Second full run: 46 failures, down from 90. Four causes, two of them mine. Reverted `validate health` entirely — it was scope creep, and it broke a real flag. ~30 of the 46 read `unknown flag "--json"; accepted: --repair, --backfill`. The previous commit routed `validate health` through the parser on the reasoning that a flag should have one owner. That was wrong twice over: §8.4 names parseNamedArgs and count queries, and `validate health` was never a parseNamedArgs call site — it read its flags, just not through the parser, so it had no silent-drop defect to fix. Tightening it omitted `--json`, which the health-diagnostic suites use heavily. The handler is now byte-for-behaviour back to its pre-branch form. `validate context` stays converted: it genuinely was a call site, and its `--json` is now declared rather than read by a second `args.includes` scan. The five handlers that had NO validation at all — init verify-work / phase-op / review / todos / remove-workspace — stay fixed. Those read args[2] with nothing checking the rest, which is the #3358 shape this phase owns. Finished the A2/A3 revert. Three tests still encoded the deleted "a value flag with a missing value is an error" rule, including one added by the previous commit for that rule. All three now assert the reverted null contract, and the ones whose titles said "rejected" are renamed — a test named for a contract it no longer asserts is its own defect. `--wave=` and `--wave --weird` are correctly rejected. Neither is documented in commands/gsd/execute-phase.md, gsd-core/workflows/execute-phase.md or docs/, and neither is emitted by the shipped prompt layer, so both are unrecognized tokens that §8.4 mandates rejecting. `doesNotConsumeFollowingFlagAsWaveValue` keeps the property it exists for — asserted directly now, at the parser, that `--wave` does not swallow a following flag as its value — and only its exit-status expectation changed. A contradiction inside this branch, surfaced by the audit and resolved the safe way. Two pre-existing #3573 tests call `state begin-phase '2'` and `state planned-phase '2'` with a bare positional, relying on the old permissive parser to ignore it. This branch's own #3358 regression test requires that exact argv to be REJECTED. The two are mutually exclusive. Widening the router to accept a bare positional — mirroring complete-phase — would have silently re-opened #3358, and was verified to do exactly that: with the widened router, `query state.planned-phase 3` returned exit 0 and wrote current_phase_name again. It is reverted. docs/CLI-TOOLS.md:116 and docs/COMMANDS.md:2192 document only the `--phase N` form for both verbs, so the two #3573 tests move to it. Their assertions were never about the call shape — only that total_phases survives the resync — and both still pass. complete-phase is untouched: its bare positional IS documented, and it keeps the dynamic boundary and the negative-space note that record why. The audit that produced this is in the PR body: for every handler whose declaration changed, the flags it reads anywhere in its body, the flags the shipped surface documents, and the shapes the suite passes, compared. The `--json` miss was a pattern, not an accident — declaring a handler's flags from its parseNamedArgs call alone misses whatever it reads elsewhere. Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3884): backfill the changeset PR number Refs #3884 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ddde001af6 |
enhance(#3873): the STATE.md schema — one owner, generated artifacts (#3880)
* test(#3873): failing-first locale parity, plus tripwires for what must not move Pins ADR-3473 §8.8 at the artifact a reader actually sees. The English STATE.md reference carries a Status lifecycle section that is missing from all four translations — the section documenting the status enum whose clobbering is #3853. The test derives the heading set rather than hard-coding the missing one, and names the locale and the heading when it fails. Two tripwires that must pass today and after. The field-drift guard still catches a re-derived fallback ladder: §8.8 instructs deleting that script, and that instruction rests on a wrong premise about what it guards, so the test stops a future reader from deleting it on the ADR's word. And last_activity's label resolution is pinned to what ships today, because it is declared in one of the two tables this phase consolidates and not the other — the consolidation must not silently pick a side. The locale test buckets under docs rather than state, which is what it tests; that bucket is allowlisted with justification rather than folded into an unrelated docs suite. It reads only markdown, so it carries no allow-test-rule marker — a marker there would suppress nothing and would grow the unverified pool against its ceiling. Refs #3873 * feat(#3873): one schema owns the STATE.md key set, three tables become projections ADR-3473 §8.8. The key set was declared in four places that had to agree by hand and already did not: FIELD_CLASSIFICATION, FRONTMATTER_BODY_SOURCE, FRONTMATTER_KEY_TO_BODY_LABEL and buildStateFrontmatter's emit behavior. One frozen null-prototype schema now declares each key's type, enum, cardinality, source, preservation, body source, body label, accepted parse shapes and whether it is emitted unconditionally; the three tables are derived from it at module load. The projections are byte-identical to the literals they replace, key order included, and the parity tests compare against verbatim copies of today's tables rather than re-deriving both sides from the schema — a parity test fed from one source proves nothing, which is how a consolidation ships a changed policy under a green test. last_activity was the live disagreement: present in one table, absent from the other. The schema declares what ships today rather than the tidier answer, and a test pins it. The schema is a leaf module and owns the four field-policy types, re-exported from state-transition so existing importers are untouched — the same split health-diagnostic-types made to break a CJS require cycle. Refs #3873 * feat(#3873): generate the schema-derived regions, parity-check the prose tables ADR-3473 §8.8's generator half. gen-state-md-docs.cjs owns marked regions in the shipped template and all five reference docs, follows gen-features.cjs's fail-closed contract, and is wired into regen:derived and lint:generated-sync. The Status lifecycle section was missing from all four translations — the section documenting the status enum behind #3853 — and is now generated into every locale. Field cardinality is a new generated table: pure schema data, no prose, so nothing to lose. The Field-reference and Status-values tables are parity-CHECKED rather than generated. Their Purpose, When-populated and Matched-text columns are genuinely hand-translated per locale, and §8.8 itself says prose stays hand-translated; generating them from an English registry would overwrite four locales' translations on every write. The row set is checked against the schema instead, so a key added to one and not the other fails, which is what field drift actually means. Building that check found last_activity_desc undocumented in all five tables. Three keys the docs describe are absent from the schema — active_phase, next_action, next_phases. They are grandfathered by name, not by wildcard, so a fourth fails: a declared gap with a forcing function rather than a silent one. Refs #3873 * fix(#3873): declare what the parsers do, and close the shape-parity gap Two declarations in the new schema described intended behavior rather than actual — the defect class this epic exists to end, committed inside the epic. Both were caught by executing the parsers instead of reading their docstrings. current_plan.acceptedShapes claimed ['N', 'N of M']. Standalone, the hybrid shape errors; the path that looks like support is parseInt truncating '2 of 5' to 2 and discarding the rest. Narrowed to ['N']. The parser is deliberately NOT fixed here: that is #3784 and PR #3791 is already doing it. When #3791 lands this row must widen, and the shape test will go red until it does — the schema and the parser cannot drift apart quietly, which is what §8.8's checked-not- generated rule is for. STATUS_LIFECYCLE_ENUM claimed to be the closed set status can hold. normalizeStateStatus passes unrecognized prose through unchanged, so it is not closed at runtime. The seven members are the canonical values it maps onto; the docstring now says that and the test asserts the real lenient contract. Closes the acceptance item that a test asserts the parsers accept exactly the declared shapes: the check is table-driven over every row carrying acceptedShapes, guarded against passing vacuously on an empty set, and fails loudly if a future row has no registered driver. Adds the unwired-label throw and the fast-check property that every projection agrees with its schema row. Refs #3873 * fix(#3873): keep the shipped template's frontmatter first, and make row 27 able to fail The remote matrix caught 12 failures with one cause. Making the template's frontmatter a generated region wrapped it in its own yaml fence ahead of the markdown fence, so extractFileTemplate and readShippedStateTemplateBody — which both match the single markdown block — found the heading first, not the frontmatter. That breaks the contract every new project's STATE.md is created from: bug #21 and epic #1969 B8 pin that the File Template block starts with frontmatter and carries gsd_state_version. The markers now sit inside the single markdown fence, so the fence opens before the frontmatter and the region still ends ahead of the heading. Same layout as before this phase, with markers embedded rather than a second fence. Row 27 existed to catch exactly this and did not, because it was writer-seeded: it asserted against the generator's own output shape, so it passed on the broken template. It now parses the fence the way production does and was verified to fail against the broken shape before being trusted against the fixed one. A test that would not have caught the bug it exists to prevent is worse than no test. The emitted-attribution failure was separate and the fragment was the wrong remedy: gsd-core/templates/state.md self-attributes under a verbatim-copy identity rule, so a diff touching it needs no acknowledgment. Fragment deleted rather than left explaining nothing. Refs #3873 * docs(#3873): how to change the STATE.md schema The phase gate was right and my docs artifact was wrong. I listed lint:generated-sync as the second enablement step, which is a verification command dressed as one, and then claimed a one-step sequence owed no how-to. The real sequence is build:lib then regen:derived, and the ordering is a trap: the generator reads the COMPILED schema, so regenerating before building regenerates against the previous schema and commits artifacts that look plausible while disagreeing with the code just written. A reference table cannot carry an ordering dependency; that is what the how-to test is for. The page covers adding, changing and removing a key, every reason code the check emits and what to do about each, what is generated versus hand-translated and why the two prose-bearing tables are parity-checked instead of generated, adding a language, and the three grandfathered keys. Indexed from docs/README.md. Refs #3873 * chore(#3873): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
3b18eff388 |
enhance(#3872): what a command reports it wrote — the transaction diff (#3878)
* test(#3872): failing-first regressions for what a command reports it wrote Pins ADR-3473 §8.7 at the consumer's output. state planned-phase advances current_phase on disk and never reports it, and reports progress.total_plans which reconcileReportedFields silently drops because it cannot resolve a dotted key against nested frontmatter. Both directions of #3818's own before/after diff, reproduced against the real CLI. Also pins the two properties the change must not break: a fully-failed patch still reports an empty updated array, which is what state.cts:607's success boolean depends on; and two content-identical writes differ in last_updated alone. That second one measured state_head NOT to be ambient — it is recomputed every write but only changes when git HEAD moved — so the provenance exclusion is a one-element set, with a companion test pinning that state_head does change when HEAD moves. Refs #3872 * feat(#3872): derive what a command reports from the transaction diff ADR-3473 §8.7. reconcileReportedFields compared the transform's own output against persisted bytes and then filtered what preservation had restored by its FIELD_CLASSIFICATION policy. Both are replaced by one comparison of persisted against the pre-write state the transaction already holds, surfaced to the command through the same caller-allocates out-param idiom divergedFields established. Both of the old directions fall out of that single comparison: a field the transform reported but the pipeline discarded is persisted-equals-snapshot and drops out, and a field nobody reported but the write moved is different and appears. The classification filter is deleted, not relocated — no policy test remains anywhere in the reporting path. Reporting is at dotted-leaf granularity, enumerated from the progress.* rows FIELD_CLASSIFICATION already declares rather than by walking user data to arbitrary depth. That closes a live defect: plannedPhaseCore already pushed progress.total_plans and reconcileReportedFields silently dropped it, because a flat hasOwnProperty cannot resolve a dotted key against nested frontmatter. Current Position was lost the same way and is fixed in the same place. The exclusion is one field, last_updated, and it is by provenance rather than by classification: it is the only field measured to change on every write regardless of content. state_head was measured NOT to qualify — it is recomputed every write but only changes when git HEAD moved. Without that exclusion state.patch's success boolean, which is updated.length > 0, would be permanently true and a fully-failed patch would report success. Refs #3872 * fix(#3872): cover the matrix, and close a prototype-chain read the coverage found Review found 20 of 29 test-matrix rows uncovered. Covering them found two real defects rather than merely documenting the intended behavior. bodyLabelFor read FRONTMATTER_KEY_TO_BODY_LABEL with a bare bracket index on a plain object literal, so a field named __proto__, constructor or toString resolved to the inherited prototype member and leaked a non-string value into the updated array. Fixed with an own-property check, mirroring the discipline resolveFrontmatterPath already had. The security-relevant matrix row proved it before the fix. applyPostSyncPreservation still carried its own inline copy of the value comparison alongside the new stateFieldValuesDiffer, which is two live copies of one rule introduced by the epic that exists to remove them. Routed through the single owner. Adds the fast-check property that a field appears iff its persisted value differs from the snapshot, the string-versus-number representation boundary, dotted paths into missing parents and into scalars, deleted and added keys, and the preserve-if-placeholder pair that proves no classification test survives in the reporting path. Refs #3872 * docs(#3872): document the transaction diff on the write path The updated array's contract belongs where the write path is described. States the iff rule, leaf granularity, the single provenance exclusion and why state_head is deliberately not one, and closes with the consequence a reader actually needs: these arrays are longer than they used to be, because they used to under-report. Refs #3872 * fix(#3872): a derived leaf materializing is not a change the caller made The remote matrix caught 17 failures with two causes. The substantive one is that progress is source: disk, and the disk cannot change during a STATE.md write — the write only touches STATE.md. So a progress block appearing where the snapshot had none is the scanner populating a document that had never been synced. The bytes moved; nothing the caller did moved them. That is the same shape as last_updated one level up, so the provenance rule is generalized rather than special-cased: a field appears iff its persisted value changed for a reason attributable to this write's action, and two cases are not attributable — a field stamped unconditionally on every save, and a declared derived leaf materializing from a source that did not change. Crucially this does not consult the preservation policy, so the filter §8.7 deleted stays deleted; it uses the declared leaf set to know which keys are derived. This had a second production consumer the earlier review concluded did not exist: cmdStatePlannedPhase gates publishStateContract on updated.length, and its own inline comment predicts exactly this failure. A no-op call was publishing state.json. advancePlanNoOpDoesNotPublish genuinely encoded pre-§8.7 behavior and moves. E2 and E6 had carved out total_plans as reportable-on-materialization, an error introduced earlier on this branch rather than a pre-existing pin, and are corrected with it. Refs #3872 * chore(#3872): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
1863f5569c |
enhance(#3871): the state transaction — mandatory snapshot, open()/rebuild() (#3874)
* test(#3871): failing-first regressions for the dropped curated progress block Pins ADR-3473 §8.6 / #3756 at the consumer's output: state record-session and state add-decision on an archived-milestone project drop the curated progress frontmatter entirely, exit 0, and report nothing. Reproduced against the real CLI before writing the tests, not inferred from the issue text. Also adds the unit-level probe that applyStatePreservation's preserve-always row is inert on a resyncing write, and an over-preservation guard that an empty project is never inflated. Refs #3871 * feat(#3871): make the STATE.md pre-write snapshot mandatory via open()/rebuild() ADR-3473 §8.6. StatePreservationInput's nullable preFm and the always-present preFmSnapshot were the same extractFrontmatter call, one of them nulled on resync — a policy flag baked into a snapshot. Both collapse into a single StateTransaction whose snapshot cannot be absent: openStateTransaction() applies preservation, rebuildStateTransaction() does not, and both carry the snapshot because the reporting phase needs it either way. An absent snapshot is now a construction failure; an empty one stays legal, because that is what a document with no parseable frontmatter honestly has. writeStateMd requires a rebuild transaction, which types ADR-3408 §8.3's closed exception list at both call sites (state sync, health --repair) instead of matching them as strings in a ratcheted baseline. Fixes the dropped curated progress block: an all-zero or absent derived total set is an unmeasured scan, not a measurement, so the curated block stands. Also fixes two defects surfaced while building — preserve-always reported a mutation even when it restored an identical value, and it re-entered the curated object by reference, which would alias the snapshot the next phase diffs against. Refs #3871 * fix(#3871): close the three remaining subsumed defects and restore the arm the type does not replace Review of the first two commits found four things. The guard shrink deleted the seam-bypass axis whole, but only its writeStateMd( arm became redundant. Its other arm catches a call site re-assembling syncStateFrontmatter + applyPostSyncPreservation instead of the owned composition, which the transaction type does not make unrepresentable and which #3469 found live. Restored as findCompositionBypasses, terminal rather than ratcheted. Three of the four issues this phase claims were untouched. All three are the epic's own shape and are fixed at the seam: current_phase_name is reasserted from the curated value when the caller names none, and cmdStateJson stops carrying a hand-maintained list parallel to FIELD_CLASSIFICATION and projects it instead. The construction failure that is the point of this phase had no test. Every enumerated matrix row now has one, including the measured-versus-unmeasured coercion boundary and a seeded property that no curated key is ever dropped. ADR-3473 §8.6 said the guard 'keeps only its raw-write check'. Verified against next: there was no raw-write check, and four other checks it does not name. Amended in place with the evidence. ARCHITECTURE.md separately advertised a preservation policy the code had deleted. Refs #3871 * fix(#3871): do not let the unmeasured-scan rule block an explicitly-requested resync The remote matrix caught over-preservation, the failure this phase's own negative space says must not happen. state update Progress re-derives the block from the body the caller just rewrote; on a project with no phase dirs the derivation yields zero totals, the unmeasured rule read that as 'the scan measured nothing', and the stale curated percent was restored over the resync the user asked for. preserve-always already said what the missing condition was: never overwrite unless the caller explicitly names this field. explicitProgressField carries it and is derived from shouldResyncStateProgress, not set by hand at a call site, so it cannot drift from what the caller asked for. Two defects found in the same mechanism and fixed with it. readModifyWriteStateMd enumerates its option keys, so a new option was silently dropped rather than rejected. And the raw-write axis captured its first argument up to the first comma, which lands inside a nested path.join, so a write to a STATE.md literal was invisible to it — the prove-it-can-fail test caught that one immediately. No test assertion was weakened; all three frontmatter rows encode #3242, #1969 B3 and #1972 and stand unchanged. Refs #3871 * docs(#3871): record why the raw-write check is kept, not why it was named The amendment justified findRawStateWrites as 'written because §8.6 requires it to exist', which is cargo-culting the contract and would have been the wrong reason to keep anything. The real reason is that writeStateMd acquires the STATE.md lockfile and a raw fs.writeFileSync acquires nothing, so this is a lock bypass and lost-update is the #500/#905/#1230 family — and after this phase it is the one reachable path into the file that nothing else covers. Also records why ADR-3408 §8.6's deletion of the 'clear' policy is not the precedent it looks like: 'clear' was dead vocabulary in a closed enum, this is coverage of a reachable path. Refs #3871 * chore(#3871): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
de95c03f72 |
fix(#3699): report why a derived frontmatter key was not written, and repair a missing body source (#3846)
* test(#3699): failing-first coverage for derived-key reporting and the case-D fallback * fix(#3699): report why a derived frontmatter key was not written, and repair a missing body source * fix(#3699): scope session-field writes to ## Session so an archived line cannot absorb the update * fix(#3699): resolve the session writer from body labels only, so a frontmatter key never writes the body * chore(#3699): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
394bf384be |
fix(#3696): report the last_activity invariant and make the verdict gateable with --strict (#3844)
* test(#3696): failing-first coverage for the last_activity invariant and --strict exit status * fix(#3696): report the last_activity invariant and make the verdict gateable with --strict * fix(#3696): agree with the real reader on last_activity, and stop reporting structure as truncation * chore(#3696): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
107eb8c1d9 |
feat(#3753): run docs guards on the PR that changes the docs they read (#3787)
A PR whose diff is entirely under docs/ runs zero tests, so a guard whose INPUT
is shipped prose cannot protect the PR lane of the diffs it exists to check. Its
only firing opportunity is after merge, on the shared branch -- which is how next
went red on
|
||
|
|
8d1f770dfe |
test(#3395): pin the clock and scope the stale-prose scan that reddened windows shard 2/3 (#3669)
* test(#3395): failing-first coverage for the silently-ignored clock pin Lands the regression matrix BEFORE the fix so the failure is proven rather than asserted. Three assertions fail deterministically on this commit: 1. PINNED_ENV does not actually pin. `_pinnedNowMs()` (src/clock.cts) returns null unless GSD_TEST_MODE is set, so GSD_NOW_MS alone is discarded and last_updated is stamped from the live wall clock. An instant ending ...:35.149Z contains the substring 35.1, which is what reddens the windows-latest shard 2/3 lane roughly 1 run in 600. 2. The colliding-instant regression cannot reach its instant, for the same reason. 3. The #3052 same-date test never lands on 2020-09-10, so it has been exercising the different-date path and passing for the wrong reason. Also adds currentPositionBlock() plus boundary (ms 099/100/199/200, second 34/35/36, LF and CRLF) and two-arm fast-check coverage for the scoped read the fix will switch to. Refs #3395 * test(#3395): pin the clock and scope the stale-prose scan to the body Drives the failing-first coverage from a5a919ffb green. Two changes, both needed: 1. PINNED_ENV now sets GSD_TEST_MODE alongside GSD_NOW_MS. _pinnedNowMs() (src/clock.cts:44) returns null without it, so the pin was silently discarded and last_updated carried a live wall-clock instant. src/clock.cts is deliberately NOT changed: requiring both keys is what stops an ambient GSD_NOW_MS from freezing a production clock, so the caller was the side that was wrong. 2. The stale-prose assertion now reads currentPositionBlock(stateContent) instead of the whole document. Frontmatter is not phase prose, and an instant ending ...:35.149Z contains the substring 35.1 — which is exactly how a document with no stale prose in it produced 'the stale 35.1 phase prose must be refreshed away'. Confirmed hypothesis: the two defects compose. The inert pin supplies a live timestamp; the whole-document scan turns it into a failure. Either alone is latent, which is why this sat unnoticed for five days and then reddened a lane the release never touched. Also corrects two things the failing-first run exposed. The property test used fc.date() without noInvalidDate, so ~1 sample in 300 was an Invalid Date whose toISOString() threw (counterexample: new Date(NaN)); re-soaked at 5000 runs. And a precondition assertion added to the #3052 block was measured to pass with or without the pin, so it was removed rather than shipped as vacuous truth — last_activity there is body-derived, not clock-derived. Refs #3395 * test(#3395): apply review findings — pin #3052, one fixture builder, CRLF coverage Spec-axis review caught a real slip: the #3052 block carried a comment saying its pin was being added as hygiene, but the RED-state revert had removed GSD_TEST_MODE and the fix commit never restored it. A comment describing an action that was not taken is worse than either doing it or leaving it alone — the pin is now actually there. Standards-axis review flagged the same frontmatter+heading fixture shape being rebuilt in three tests. Extracted one stateDoc({iso, lines, eol}) builder; eol is a parameter rather than a constant because the helper's CRLF behavior is a claim under test. Self-review finding: CRLF was only exercised on a single-heading document, and the following-heading case only under LF — so the exact claim the helper's comment rests on (`\n## ` matches inside `\r\n## ` because the CR precedes the newline) was never actually run. The control test now loops both line endings WITH a following heading. Also drops a comment that restated the PINNED_INSTANT rationale verbatim. Refs #3395 --------- Co-authored-by: sim <sim@local> |
||
|
|
46f14c621e |
fix(#3583): one percent per write — route update-progress through the shared computation (#3634)
* test(3583): failing-first coverage for one percent per write state update-progress computes plan throughput (summaries/plans) for stdout and the body Progress bar, while the same write re-derives frontmatter progress.percent as min(planFraction, phaseFraction). Neither consults the other, so mid-phase the file contradicts itself and state json disagrees with the verb that just wrote it. These tests fail on that: equality across stdout, body bar, frontmatter and state json on fixtures where the two fractions differ, plus a derivation-parity test that fails if completedPhases is ever derived by summary parity instead of verification-passed status. Also updates three pre-existing tests that pinned stdout to the plan-throughput value (50->0, 50->0, 100->0). Those fixtures have summarized-but-unverified phases, so the old expectations encoded the bug; changing them IS the fix, as the issue states explicitly. * fix(3583): one percent per write — route the verb through the shared computation RED proven at 7dbbb2d2: 9 failures — the new cross-surface equality tests, the withhold test, and the pre-existing tests whose expectations encoded the bug. state update-progress computed plan throughput (summaries/plans) for stdout and the body Progress bar, while the SAME write re-derived frontmatter progress.percent as min(planFraction, phaseFraction) through a separate path. Neither consulted the other, so on any project where plan throughput ran ahead of phase completion — the normal mid-phase state — the file contradicted itself and state json disagreed with the verb that had just written it. Exit 0, no signal. This is not a dispute about which metric is right. The min cap is deliberate (#3242 Bug B) and is untouched; the fix aligns the printed and body values WITH it. Verified by diff: computeProgressPercent's definition and cmdStateSync are both unmodified. The verb now takes its percent from buildStateFrontmatter — the single owner of the isPhaseComplete-based completedPhases count and the ROADMAP-union totalPhases logic that the frontmatter sync later uses inside the same read-modify-write. Both calls hit the same disk-scan cache against the same on-disk state, so they cannot disagree. Reusing that owner, rather than re-deriving completedPhases locally, is the point: a second almost-identical derivation is the very defect class being fixed, and a parity test now fails if anyone swaps it for summary parity. The first cut fell back to plan throughput when the shared computation withheld. That reintroduced the defect in a rarer case — stdout would print a number the frontmatter deliberately did not contain — so it is gone. The verb now withholds in the same shape as its existing #3217 and #3233 guards. That path is reachable, not theoretical: a bare vX.Y token in ROADMAP prose with no versioned heading leaves the milestone unbounded while both existing guards see a COMPLETE scope. Covered by a test that also asserts state json omits the percent, proving it is the same withhold rather than a divergent local computation. Three pre-existing tests pinned stdout to plan throughput (50->0, 50->0, 100->0); their fixtures have summarized-but-unverified phases, so those expectations encoded the bug. Updating them is the fix, as the issue states. Fixes #3583 * fix(3583): source the reported counts from the same milestone window as the percent The adversarial pass found the first cut left the SAME defect one field over. cmdStateUpdateProgress still reported completed/total from the top-of-function scan, which calls listMilestonePhaseDirs with NO versionOverride — the auto-derived current milestone — while percent now came from buildStateFrontmatter, whose scan scopes by versionOverride: storedMilestone. getMilestonePhaseFilter shows those can select different milestone windows, and #3017's own comment warns about exactly that mis-bind. So a single JSON object could report a percent inconsistent with its own counts: the self-contradiction this issue was filed to close, relocated rather than removed. Counts now come from the same buildStateFrontmatter result as the percent. Proven on a real divergent-milestone fixture where a preamble phase leaks into the auto-derived scan but is excluded from the stored-milestone-scoped one: with the fix stashed the verb emits {percent:0, completed:1, total:2}; with it applied, {percent:0, completed:1, total:1}. The guard scan remains, gating only the #3217/#3233 withholds. Also corrected a comment that overstated caching. Only the phase/plan disk scan is shared between the two buildStateFrontmatter calls; getMilestoneInfo re-reads and re-parses ROADMAP.md and readGitHeadSha spawns a bounded git rev-parse, and both now run twice per invocation. Threading a precomputed frontmatter through the write seam to avoid it was rejected: that seam is the shared ADR-3408 §8.3 composition with three other callers and heavily-documented invariants, and this is not the change to renegotiate it. The comment now says what is and is not cached instead of implying the second call is free. Standards: six new assertions matched raw STATE.md body text the code under test had just produced — the pattern CONTRIBUTING bans by name. They now extract the body Progress field with the repo's own field extractor and assert the parsed percent, so the check survives rewording of the rendered bar. The acceptance criterion still verifies the bar; only what it asserts on moved. Also trimmed ~50 lines of narration around a ~15-line change into a named helper, and fixed a stale test comment that still claimed 100% next to assertions expecting 0%. * chore(3583): add changeset fragment * chore(3583): backfill changeset PR number (#3634) --------- Co-authored-by: sim <sim@local> |
||
|
|
bcefffc132 |
fix(#3578): derive milestone status from phase counters, not phase-completion prose (#3614)
* test(3578): failing-first coverage for milestone status on partial completion Completing phase 2 of a 4-phase milestone sets frontmatter status: completed while the same call correctly writes completed_phases: 2 / total_phases: 4. These tests fail on that conflation and pin the boundary either side of it (3-of-4 must not complete, 4-of-4 must), plus milestone_name byte-identity and the 1-of-1 case that legitimately does complete. * fix(3578): derive milestone status from phase counters, not phase-completion prose RED proven at 253843b4 (tests-only): the 2-of-4 and 3-of-4 cases failed while the 4-of-4, milestone_name and 1-of-1 controls passed — the conflation, and nothing else. state complete-phase writes body prose `Phase N complete`. normalizeStateStatus matches 'complete' as a case-insensitive SUBSTRING, so phase-level prose collapsed into milestone-level frontmatter status: completed — even while the same call correctly derived completed_phases: 2 / total_phases: 4 / percent: 50. Check ORDER is why the sibling surface stays correct: completePhaseCore writes 'Ready to plan' for non-final phases, hitting the 'planning' arm before 'complete'. The two phase-completion surfaces disagreed and this was the conflated one — a violation of ADR-2207, which gives milestone termination solely to milestoneCompleteCore. buildStateFrontmatter now honors a 'completed' normalization from phase-completion prose only when the counters it already derived agree. Scoped deliberately: - anchored to bare `Phase <token> complete`, so 'All phases complete' and '<version> milestone complete' are untouched (both out of scope). Verified by executing the guard's own regex from source against both forms. - gated on counter trustworthiness (COMPLETE disk scope, finite counts, positive denominator) so an unknown scope withholds rather than guessing 'not complete', which would be the mirror-image bug - normalizeStateStatus itself is NOT modified — it feeds every state.* write and the read path A 1-of-1 milestone still yields 'completed' by the rule, not by exemption, so the #1255 pinning test stays green on its merits. Fixes #3578 * fix(3578): gate the guard on milestone boundedness and close the review gaps Review findings from two orthogonal passes, all fixed inline. GUARD (correctness, from the standards pass): the guard omitted `milestoneUnbounded`, which is the established trust authority for these very counters in this same function — it nulls progressPercent at :2286 and gates the prose fallback at :2294. An unbounded milestone yields a conflated/understated total, so `completedPhases < totalPhases` could be an artifact of a bad denominator and demote a genuinely-complete milestone. Now gated. TESTS: - Prose/guard parity assertion. The guard regex-matches prose emitted from a DIFFERENT file; if that prose drifts the guard silently stops firing and the bug returns undetected. Per the repo's generative-fix-divergence rule, a test now asserts the emitted body Status still matches the guard's pattern — asserting the emitted value against the pattern rather than duplicating the string. - limit+1: completedPhases > totalPhases must NOT fire; inconsistent counters fall through rather than guessing. - Untrustworthy counters (no phases dir → totalPhases null) must NOT fire. - AC4: MCP invoke-command dispatch parity via handleMessage, the criterion both reviewers independently flagged as asserted-but-untested. - Hand-rolled STATE.md writes routed through the existing writeState fixture helper. The adversarial pass independently verified, by reading rather than trusting the diff's own comments, that: paused/stopped short-circuit before 'completed' so a paused milestone can never be clobbered; only cmdStateCompletePhase emits the targeted prose, so no sibling caller over-fires; the counters come from a fresh disk scan independent of this write, so there is no pre/post off-by-one; and the #1255 pinning fixture creates no phases dir, leaving completedPhases null and the guard inert — so that test is provably unaffected rather than assumed to be. * chore(3578): add changeset fragment * chore(3578): backfill changeset PR number (#3614) --------- Co-authored-by: sim <sim@local> |
||
|
|
59e7a677fe | fix(#3511): scope every phase-directory scan to the phase it belongs to (#3535) | ||
|
|
d922469613 |
refactor(#3408): close the two known limits instead of recording them (#3524)
* refactor(#3408): close the two known limits instead of recording them
Both of these were flagged in review and written down as 'known limits' in a
PR body and an issue comment. CLAUDE.md is explicit that a note is not a fix
and is not surfacing — it is a silent defer. Recording them while closing the
epic was the pattern this epic exists to remove, performed on the epic itself.
syncAndPreserveStateMd and applyPostSyncPreservation each took eight
positional arguments, the last three optional, one of them an out-param. The
review's own wording was that 'a third consumer should trigger an
options-object refactor' — a deferral with a trigger condition nobody would
notice firing. Content and path stay positional; resync, authoritativeFm,
deriveProgressKeys and divergedFields move into a named
StatePreservationOptions. Every call site updated, with tsc as the proof none
was missed.
cmdStateCompletePhase's updated array carried both field labels and a section
name, worked around by a SECTION_ENTRIES Set that re-derived the distinction
by string matching. The kinds are now typed where they are produced and
flattened once at output.
Output contract unchanged: updated is still a flat string array with the same
entries in the same order.
Behavior-preservation was proven rather than asserted — the compiled lib was
built at
|
||
|
|
1b027298dc |
fix(#3481): resolve add-roadmap-evolution's phase from STATE.md, not a literal ? (#3522)
* fix(#3481): resolve add-roadmap-evolution's phase from STATE.md, not a literal `?` `state add-roadmap-evolution` built its entry from the raw `--phase` flag alone, so omitting the flag persisted `- Phase ?` even when STATE.md's own frontmatter carried `current_phase` above the insertion point — the #3231 defect at a second call site. Roadmap-evolution entries are the permanent trail explaining why the roadmap changed shape; `Phase ?` makes that trail unattributable, and the command is mostly invoked from agents that do not know to pass `--phase`. The #3481 triage confirmed the #3231 sibling site (`add-decision`) was also still unfixed on next — both PRs that attempted it (#3232, #3347) were closed unmerged. This applies the #3347 treatment to both call sites: - Extracts the write-path phase-resolution ladder `cmdStatePrune` already ran — frontmatter `current_phase` → body `Current Phase` field → prose `Phase: X of Y` scoped to `## Current Position` — into a shared `resolveCurrentPhaseId`, and routes `cmdStateAddRoadmapEvolution`, `cmdStateAddDecision`, and `cmdStatePrune` through it. - Deliberately NOT routed through `resolveStatePhase` (#3208): its `matchCurrentPositionSection(body) ?? body` fallback widens the prose rung to the whole document when no `## Current Position` section exists, where the pipe-table fallback matches any historical `| Phase | N |` row (#1776). Read-path callers (snapshot/validate) report to a human; write-path callers persist durably, so they take the strict rung and render `?` instead of guessing. - The resolved id is returned as written, never parsed to a number (`11-01` and `04.1` are real ids). Prune still parses its own integer cutoff, so its behavior is byte-identical. - Explicit `--phase` still wins and its path is untouched — STATE.md is not even read. When no rung resolves, `?` is still written. Tests: per-call-site coverage for both commands — omitted `--phase` resolves (including a non-integer prose id), explicit `--phase` wins, and two counter-tests pinning the degraded verdict (nothing resolvable → `?`, and a historical `| Phase | 7 |` table row must NOT be adopted). Plus a static guard sweeping src/*.cts for the raw `phase || '?'` placeholder shape so a future call site cannot reintroduce the class. Fixes #3481 * chore(#3481): add changeset fragment for PR #3522 --------- Co-authored-by: sim <sim@local> |
||
|
|
411196bc3a |
refactor(#3471): one enforcement point for the empty case, and reports that match the disk (#3519)
* refactor(#3471): one enforcement point for the empty case, and reports that match the disk Implements ADR-3408 section 8.5 and section 8.4's residue (folded in when Phase 3 closed as subsumed). Four items, and two findings the design did not predict. FINDING 1 — the guards could not simply be deleted, as the design instructed. state sync and REGENERATE_STATE never run applyStatePreservation at all, so those six conditions were their ONLY empty-field fallback. A baseline probe on the unedited tree confirmed unconditional deletion drops current_phase, current_phase_name, current_plan, stopped_at and paused_at from a blank-body STATE.md on state sync — breaking the byte-identical requirement section 8.3 grants those two sanctioned-permanent exceptions. They are now GATED, not deleted: on for the exceptions, off for the write seam, where an empty derived value finally reaches the executor unmolested. FINDING 2, the more serious one — there was a FOURTH encoding of this policy. The pre-existing #2202 unknown-key carry-forward loop independently restored the same six fields whenever derivedFm lacked the key, completely neutralizing the fix. It is named nowhere in the ADR, the design, or three prior phases. It was found only because a probe that should have passed did not: the first attempt reported divergedFields: [] and silently restored both fields, reproducing the exact bug this phase exists to close. That is worth stating plainly. This epic's thesis is 'policy declared in one table, enforcement hand-rolled per call site.' The final phase found one more call site than anyone had counted — which is the fourth consecutive time a copy count in this epic proved to be a lower bound. Also: divergedFields could only observe fields the executor actively RESTORED, by diffing postFm. A discard-to-empty is absent both before and after, so it was invisible. A second pass now reports it, which is what makes section 8.5's 'preservation is visible' true for the delete-the-body-line case rather than aspirational. cmdPhaseComplete now reports what it preserved — #3374 was filed against that command and its complaint was warnings: [], silence. cmdStateJson's private third copy of the guards is routed onto the executor's preserve-when-unchanged rule. A read is definitionally not a write, so the #1230 delta is 'unchanged' and curated wins over a stale annotation. shouldPreserveExistingProgress is a different rule and is untouched. Report reconciliation is ONE shared helper across seven commands, not five copies of fix(#3351)'s block. Five copies of a reconciliation is precisely the shape this epic removes, and introducing it in the final phase would have been a poor joke. Both untraced commands were traced rather than assumed: cmdStatePlannedPhase matched cmdStateBeginPhase exactly; cmdStateCompletePhase turned out to be a different legacy hand-rolled path reporting a mix of field names AND a section name, where the naive helper would have dropped 'Current Position' as a false negative every time. * test(#3471): characterization coverage for one enforcement point and reconciled reports Matrix sections A-E, asserted at the consumer's output per ADR-3180 Decision 4(b)/(c) — this phase owes Decision 5's outcome metric, the one the drift guard's zero may never be reported without. Three walls matter more than the new coverage: A2 is SIX separately named tests, one per gated guard, not one parameterised assertion over a list. A list is trivially shortened later; six named tests are not, and six guards is exactly where a field gets silently dropped. A6 pins what Phases 1-3 already fixed — non-empty stale body, delta unchanged, losing to fresher curated frontmatter, with the divergence reported. If A6 reddens, this phase broke the thing the epic was for. D1/D2 pin state sync byte-identical. The implementation had to GATE the six guards rather than delete them precisely because state sync has no executor, and a baseline probe showed unconditional deletion drops five fields. Nothing else in the suite would notice that regression. E6 covers #3345's direction — a field preservation restored that the intent never named IS reported. Nothing has ever tested that direction. Assertions were empirically verified against the compiled lib and the real CLI before being written, since the suite cannot be executed locally. That caught two type bugs in the draft: fm.current_phase after a quoted-YAML round-trip is the string '5', not the number 5. E5 is recorded as structurally unreachable rather than weakened or faked. Those four commands report body Title-Case labels, which cannot string-collide with a frontmatter snake_case key the way cmdStatePatch's arbitrary field names can — which is why fix(#3351) targeted only cmdStatePatch. Testing it directly would need reconcileReportedFields exported from private scope; the helper is exercised through E6 and all seven commands instead. * docs(#3471): amend ADR-3408 section 8.5 — a fourth enforcement point, and guards that could not be deleted Amendment 3. The contract held; two of section 8.5's own statements did not. It said the six empty-only guards are DELETED. They cannot be. writeStateMd is the sole path for both section 8.3 sanctioned-permanent exceptions and never runs applyStatePreservation, so those guards were their only empty-field fallback. A baseline probe on the unedited tree confirmed unconditional deletion drops five fields from a blank-body STATE.md on state sync, breaking the byte-identical guarantee section 8.3 grants it. They are gated instead. It also mis-located cmdStateJson's guards, describing them as living in syncStateFrontmatter. They were a separate private copy on the read path with no delta check at all, so a stale body annotation always beat fresher curated frontmatter in state.json — #3395's shape entirely outside the write seam. THE FINDING: a fourth enforcement point nobody had counted. The pre-existing #2202 unknown-key carry-forward loop independently restored the same six fields, silently neutralizing the fix. It is named nowhere in this ADR, in the phase design, or in three prior phases, and was found only because a probe that should have passed did not. Fourth consecutive time a copy count in this epic proved a lower bound: 2 write-seam bypasses became 4, three preservation encodings became four, and the estimate was wrong every time. ADR-3180's standing rule has earned itself in every phase — read the code, not the write-up. Records the Row 2 decision (a discard-to-empty wins per the delta rule and is reported, not silent — the sharpest Hyrum exposure in the epic), section 8.4's residue landing as ONE shared reconcileReportedFields across seven commands rather than five copies, and the parity assertion added because FRONTMATTER_KEY_TO_BODY_LABEL was itself a second table that failed silently — this epic's shape in miniature, in its final phase. * fix(#3471): repair four regressions the checkpoint caught Checkpoint returned 16 failures of 34389: six real regressions in pre-existing tests, plus seven of my own test bugs. My hypothesis was wrong and is recorded as such. I predicted the #2202 carry-forward skip was the cause, reasoning it had removed a load-bearing fallback the way the six guards nearly were. It was not implicated in any of the six. Three unrelated causes: #2111 — current_phase came back undefined from milestone complete, which is the epic's own defect class reintroduced by its final phase. Root cause is Row 2 working exactly as designed: milestoneCompleteCore rewrites the body Phase: line to a closure message, so current_phase's #1230 delta reads CHANGED and the new rule correctly discards the curated value. The transition never declared any intent to touch that field. Fixed by re-asserting current_phase and current_phase_name through authoritativeFm — the existing #2736 mechanism beginPhaseCore and completePhaseCore already use — rather than by weakening Row 2, which A5 pins. That interaction is worth naming: a rule that keys on 'did this write change the body source' will fire on a transition that moves the body line for an entirely unrelated reason. The design did not anticipate it. #1264 / #3242 / the state.patch progress report — reconcileReportedFields folded EVERY divergedFields entry into updated, including preserve-always progress restores no caller asked about. Now scoped to preserve-when-unchanged rows only. #1162 / case-insensitive table fields — valueOf checked frontmatter before body, so a lowercase table field name exact-matched the lowercase frontmatter key sync always derives, comparing stale pre-sync body text against a post-sync frontmatter enum. Flipped to body-first. That last one is the SAME lesson as Phase 2's patchCore, recurring in a different function two phases later: in this model the body is authoritative and frontmatter is the projection, so a name that could mean either resolves body-first. Twice now. Test bugs: a stray unused parameter shifted every argument at six call sites, so body arrived undefined; and A4 compared nested progress scalars against numbers when extractFrontmatter returns raw YAML strings. The string-vs-number YAML round-trip has now been caught three times in this phase alone. * test(#3471): one helper for the progress coercion that bit four times A2f failed on the string-vs-number YAML round-trip: extractFrontmatter returns nested progress scalars as raw YAML strings, so a comparison against numeric literals can never pass. This is the FOURTH time this exact class has been caught in this phase — twice during test authoring, once as A4 in the previous checkpoint, now as A2f. Patching it a fourth time by hand would guarantee a fifth. Added numericProgress() with a comment saying why it exists, and routed every progress-reading assertion in the #3471 block through it. Swept the block: C3 needed no change, because cmdStateJson's output already runs through normalizeProgressNumbers. Deliberately NOT shared with frontmatter.test.cjs's readPersistedProgress: that one is path-based and re-reads from disk, while these assert on an in-memory string that is never written. Sharing would have meant either a disk round-trip these tests do not do, or duplicating half the helper — so the coercion pattern is mirrored locally and the reason recorded, rather than manufacturing a dependency to satisfy the letter of consolidation. * chore(#3471): backfill pr number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
e2f4c16d9e |
refactor(#3469): one composition for the STATE.md write seam (#3501)
* docs(#3469): amend ADR-3408 section 8.3 — the pipeline has sanctioned exceptions Section 8.3 read 'Every STATE.md write applies the pipeline.' That is false by design for two commands, and acting on it would have inverted a shipped feature. Preservation makes curated frontmatter win over a re-derived body value. state sync exists to do the opposite — #905's 'body annotation beats existing frontmatter when both are present'; it re-derives frontmatter FROM the body. REGENERATE_STATE is a factory reset that rebuilds STATE.md from scratch. Applying the pipeline to either would re-lock exactly what the command was invoked to replace. This issue's own scope line, inherited from the epic, said to route the direct writeStateMd callers through the pipeline. For cmdStateSync that would have shipped silently, with every gate green, because no test asserts that sync LETS the body win. Caught by reading the helper's docstring and then verifying the claim against the code — a stale comment had already misdirected this epic once. Both commands are now named in a closed exception list and are permanent ratchet entries. Consequence recorded rather than left to bite Phase 4: the 'drive the ratchet to 0 and delete the file' target in this ADR and in #3471 is wrong. Two entries are permanent, so the correct end state is 2, and the honest report is '0 removable bypasses, 2 sanctioned'. A guard reaching 0 here would only do so by having stopped looking at two real writers. * refactor(#3469): one composition for the write seam, not one per caller Implements ADR-3408 section 8.3 as amended. syncAndPreserveStateMd is now the single composition of syncStateFrontmatter and applyPostSyncPreservation. readModifyWriteStateMd and cmdPhaseComplete both CALL it instead of each assembling the two steps themselves. cmdPhaseComplete keeps its own writePlanningFileSet envelope — the composition returns content, it does not take over the write, so STATE.md still commits atomically with ROADMAP and REQUIREMENTS. Assembling the stages at a call site is a re-derivation even when every step calls an owner. Upstream's fix(#3374) routed cmdPhaseComplete through applyPostSyncPreservation but left it calling syncStateFrontmatter directly first, so the composition was duplicated and free to diverge with both guards green. That is ADR-3180 Amendment 2's finding repeating on the write side. cmdMilestoneComplete gains preservation. It wrote through writeStateMd, so it got sync and no preservation — the identical shape #3374 reported for phase.complete, and flagged upstream as a follow-up in the helper's own docstring. This is that follow-up. Divergence is now visible: preservation_warnings names each field restored over a disagreeing derived value. Deliberately NOT named warnings — cmdPhaseComplete already exposes warnings as a prose string array, and two sibling commands carrying that name with different element types is Generative Fix Divergence, the class this epic exists to remove. patchCore stops running stateReplaceField over the whole document. One observable consequence, intended per design row 9: a frontmatter-shaped patch key with no body counterpart now reports failed instead of silently succeeding, because the old whole-document match was literally hitting the YAML line case-insensitively. The guard closes Phase 1's DECLARED KNOWN GAP as promised rather than re-deferring it: section 8.3(b) detection is tractable now the composition exists. Scoped by two factors to avoid Phase 1's measured 29-to-1 false positive rate — a variable field-name argument AND a content argument whose nearest preceding assignment is not stripFrontmatter. Verified 0 findings and 0 false positives across all 33 call sites, plus 5 synthetic shapes. It also detects the re-assembly shape above. Ratchet: 4 entries to 2, both sanctioned-permanent. cmdStateSync's owner changes from #3471 to sanctioned-permanent per Amendment 2 — routing it through preservation would invert the #905 contract. Also fixed inline rather than deferred: cmdMilestoneComplete's STATE.md read now happens inside withStateLock. It previously read outside any lock before writeStateMd took its own, leaving a TOCTOU window under concurrent writers. * test(#3469): characterization coverage for the single write seam Matrix sections A-E. Criterion 6 was amended by maintainer decision — all five instances closed by point fixes while Phase 1 was in flight — so these are characterization tests at the consumer's output per ADR-3180 Decision 4(b)/(c), paired with the drift guard's count, never either alone. Section C is the one that earns its keep. cmdStateSync is a sanctioned permanent exception: state sync exists to re-derive frontmatter FROM the body, so preservation there re-locks exactly what the command was invoked to replace. C1 pins that the body wins; C4 pins that this phase left the command byte-identical. Nothing else in the suite would notice if a future change made sync start preserving, and the natural reading of 'one write seam' is to make precisely that change. Section E pins the guard's false-positive scoping. E4 (updateCore's strip-then-replace) and E5 (sectionBody-scoped calls) must NOT be reported — the naive detector measured 29 false positives to 1 true positive in Phase 1. E7 is the inverse: a sanctioned-permanent entry disappearing must FAIL, because a guard reaching zero here would only do so by having stopped looking at two real writers. Also corrects a stale test that asserted patchCore's old whole-document behavior, which this phase deliberately changes. One honest limitation, flagged rather than papered over: A1's 'byte-identical to pre-refactor' cannot be diffed against real pre-refactor bytes from inside the suite. It is implemented as the seeded fast-check property that cmdPhaseComplete's composed output equals readModifyWriteStateMd's for the same inputs — the strongest available proxy, not the literal claim. * docs(#3469): refresh the seam glossary entry and add the changeset Two spec-review gaps, both real. CONTEXT.md's STATE.md Transition Module entry named three direct writeStateMd callers including cmdMilestoneComplete. This phase routed that one through the composition, so the line was false the moment the refactor landed. Worth recording plainly: I wrote that sentence in Phase 0, correcting an older stale pointer in it, and my own Phase 2 change invalidated it again within the same epic. That is the exact drift this epic exists to remove, demonstrated on the epic's own documentation — and it is why the entry now ends by saying the whole-repo drift guard, not this line, is the authoritative count. The entry now records the composition (syncAndPreserveStateMd) and states that exactly two direct callers remain, both SANCTIONED PERMANENT rather than debt. Changeset: type Changed, because milestone complete's observable output moves. Tier-2 per ADR-3180 Decision 3 — a stale body line no longer wins over fresher frontmatter, and the command gains preservation_warnings. Docs requirement is met by the ADR amendment already in this diff. * test(#3469): register property-test temp-dir cleanup at creation time Standards review, minor but real: the new fast-check property cleaned up its temp dirs in a loop AFTER fc.assert returned. A genuine property failure throws, so that line never ran and every dir from the failing run — including all of fast-check's shrinking iterations — leaked. The failure path is exactly when a littered machine hurts most, and a failing property test is the case the test exists for. Cleanup is now registered with t.after() at dir-creation time, so teardown happens however the test exits. Not try/finally — CONTRIBUTING.md:356 bans it inside test bodies, which is why the after-the-assertion shape existed in the first place. Swept the rest of the branch's test diff for the same shape; phase.test.cjs already uses registered teardown and nothing else matched. * fix(#3469): patchCore routes frontmatter writes instead of dropping them Checkpoint returned 10 failures of 33880. One implementation defect, three test defects, one stale test — all fixed, and the implementation defect is the one that matters. patchCore stripped frontmatter and then reconstructed it VERBATIM, applying no patches to it. An arbitrary custom frontmatter key with no body counterpart and no FIELD_CLASSIFICATION row — risk_level in the upstream fix(#3351) test — therefore always reported failed and silently never wrote. It worked before, via the old whole-document match on the raw YAML line. That is a regression against this phase's own design row 9, which requires frontmatter changes to ROUTE THROUGH the seam — still work, policy-governed — not to stop working. Removing a capability is not routing it. An upstream test caught it, which is the argument for running the checkpoint before believing the refactor. patchCore now partitions by frontmatter shape, decided structurally from the parsed frontmatter's own keys rather than a naming heuristic: - classified keys still report failed — policy owns them and a raw patch may not bypass it; - unclassified keys apply to the frontmatter object and report updated — Phase 1's behavior-table row 19, a field with no row is not this contract's business; - body-shaped keys are unchanged. The property 'failure' was my own test breaking the repo's Clock Seams rule. The two paths agree byte-for-byte; the only difference was last_updated, stamped from the wall clock on two invocations milliseconds apart, so it could never pass. Time is now frozen with mock.timers across both — not by excluding last_updated from the comparison, which would have silently stopped comparing a field the composition writes. B4's fixture could not discriminate: normalizeStateStatus maps any text containing 'complete' to 'completed', and milestone complete's own new body value derives to exactly that — which was also the fixture's stale value. The stale value is now 'executing' so the assertion can tell 'body correctly won' from 'stale survived'. B5's fixture tripped a pre-existing unstarted-phase guard before reaching any write-seam code; it now has the matching phase directory. D9 asserted the old exempt set. readModifyWriteStateMd now calls one symbol rather than assembling two, so it needs no exemption; syncAndPreserveStateMd is the sole legitimate composition site. * fix(#3469): patchCore resolves body-first, so the body wins a name collision Re-verification returned 2 failures of 33880, both D4 — the hostile row for a key that exists as BOTH a frontmatter key and a body field. The partition checked frontmatter first, so 'status' — classified in FIELD_CLASSIFICATION and also present as a body 'Status:' line — routed to the frontmatter branch, was rejected as classified, and reported failed. Wrong order. Patching 'status' means the body field, and upstream fix(#3351) says so in its own comment: 'the legitimate working case for state.patch is display-cased BODY fields — Status, Current Plan, Phase.' The body is authoritative in this model; frontmatter is the projection. D4 asserted exactly that and was right. Resolution order is now body, then frontmatter: 1. resolves to a body field -> apply to body, updated 2. else an own key of the frontmatter: classified -> failed (policy owns it) unclassified -> apply to frontmatter, updated 3. else -> failed Verified by probe against the compiled lib for all four cases rather than asserted: risk_level (frontmatter-only, unclassified) still lands; current_phase still fails; display-cased Status unchanged; D4's lower-cased status now lands via the body with the frontmatter untouched. The current_phase case was the one that could have regressed silently, so its fixture was read rather than assumed — D1's body carries 'Phase: 3 (alpha)' and no 'Current Phase:' line, so body-first cannot reach it. * chore(#3469): backfill pr number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
1218d76d62 |
refactor(#3468): dispatch state preservation on the declared policy, not the field (#3495)
* test(#3468): add write-path drift guard, ratcheted at its measured baseline Guard-first, per ADR-3180 Amendment 3's standing rule that a phase builds and runs its guard BEFORE its scope is fixed, and states its copy count as 'N found by the guard', never 'N per the epic'. Measured, not assumed: Axis 1 (policy dispatch, ADR-3408 section 8.1) — 7 violations, RED by design. 5 field-name-keyed getFieldClassification('literal') branches plus 2 declared FieldPreservation members with no executor at all (derive, clear). This is the fail-first evidence for the refactor. Axis 2 (write seam, section 8.3) — 4 bypasses, ratcheted. Epic #3408 scoped this at two writers; the whole-repo scan found four, and one the epic named (patchCore) is not among them because it bypasses via stateReplaceField rather than the seam calls. Fourth consecutive time an epic's copy count proved a lower bound. Two detectors were written and removed again before this commit, both recorded in the file header rather than silently dropped: - A prompt-layer detector that reported 5 backticked prose mentions as drift. That is ADR-3180 Amendment 3's recorded false-positive class, and CONTRIBUTING.md already settles it: a backticked command reference is a mention. Now gated on inline-code spans. - A stateReplaceField co-occurrence detector for section 8.3(b). Measured at 29 false positives to 1 true positive — it matched the function's own definition and ~20 calls on frontmatter-free body slices. Banking 29 non-defects to catch one is the 'ratchet as a parking lot' gaming route Decision 5 names, so it is a DECLARED KNOWN GAP owned by Phase 2 (#3469), which both fixes it and makes its detection tractable. * test(#3468): failing-first coverage for policy dispatch and the loud failure Matrix sections A, B and C from 50-test-matrix.md. Expected RED against this tree, confirmed by static trace rather than assumed: B1, B2, B3 — an unwired declared preserve-when-unchanged row must throw with code STATE_PRESERVATION_UNWIRED_ROW and a structured .field. Today src/state-transition.cts:314 silently continues. A4 — a whitespace-only snapshot is restored today, because the guard is .length > 0. Required behavior is skip. Everything else is characterization, locking in behavior the refactor must preserve. C1 is table-driven over every FIELD_CLASSIFICATION key; C2 pins current_phase_name's exact outputs as literals, because its row is being reclassified preserve-always to preserve-when-unchanged as a behavior-preserving change and nothing else would catch a drift. C3 is a seeded fast-check property (seed 3468, 200 runs, replay data on failure). A22 is deliberately NOT a behavioral test. Whether 'derive' has an explicit executor is not observable through applyStatePreservation's public API — it is a structural property, and the drift guard's unimplemented_policy axis is what enforces it. That split is ADR-3408 Decision 5's own pairing: the lint is the structural metric, the test is the outcome metric, and neither is reported alone. * refactor(#3468): dispatch preservation on the declared policy, not the field Implements ADR-3408 sections 8.1, 8.2 and 8.6. applyStatePreservation is now one loop over FIELD_CLASSIFICATION dispatching on the row's preservation value, with four small executors — one per FieldPreservation member. No branch is selected by field name. Zero literal-argument getFieldClassification calls remain. Behavior-preserving for 16 of 20 input classes. The four that change: - An unwired declared preserve-when-unchanged row now THROWS (code STATE_PRESERVATION_UNWIRED_ROW, structured .field) instead of silently continuing. This fires only on an internal invariant violation with both ends in our own source; a drifted, malformed or unparseable user STATE.md must never reach it, which is section 8.2's bright line and what test B8 proves through the real CLI. - derive gained an explicit no-op executor. That is what makes the throw decidable: 'policy says do nothing' is now distinguishable from 'nobody wired this'. - current_phase_name's row is corrected from preserve-always to preserve-when-unchanged. The row was wrong, not the code — it has always been delta-gated on the body Phase line, so preserve-always had two divergent implementations. Behavior is unchanged and test C2 pins it. - A whitespace-only snapshot is no longer restored; the check is trimmed. clear is deleted from the FieldPreservation union — no row used it and no executor existed. Speculative Generality: a policy invented for a need that never arrived. Verified zero dependents. The caller folds six dedicated pre/post parameters into one bodyDeltas map keyed by field, so all seven preserve-when-unchanged rows travel one channel instead of two. Two shapes for one kind of data is why the executor needed per-field branches at all. Also fixed, found while reviewing the refactor rather than deferred: - applyPreserveIfPlaceholder opened with a field-name literal test, which section 8.1 forbids outright. The executor is idempotent, so the test bought nothing. The drift guard could not see it, so Axis 1 is widened to catch field-variable comparisons against literals — the guard reported zero while a violation sat in the file it polices, which is Goodhart's gaming-by-indirection. - loadBaseline conflated an unreadable baseline with an absent one. A guard whose own diagnostic collapses two states into one identical result reproduces the exact failure shape this epic exists to remove. * docs(#3468): record Phase 1 validation as ADR-3408 Amendment 1 Amendment 1 records what Phase 1 found, per ADR-3408 section 8's rule that a behavior it does not state is not decided: - preserve-always had TWO divergent implementations; current_phase_name's row was wrong and is reclassified, behavior unchanged. - section 8.6 resolved: clear is deleted, zero dependents. - the closed guard vocabulary is real and has exactly one true member, because stopped_at's scoping turned out to be caller-side extraction. - copy count found by the guard: 4 write-seam bypasses where the epic scoped 2, and patchCore — one of the two it named — is not among them. - two detectors built and removed again, with their measured false-positive rates, so nobody re-attempts them. - a DECLARED KNOWN GAP for section 8.3(b), owned by Phase 2. - Decision 5's anti-gaming list earned itself twice in one phase. Also adds the changeset fragment. * test(#3468): fix review findings — try/finally, stale clear allowlist, ratchet owners Standards axis, both hard violations: - tests/state-write-path-drift-guard.test.cjs wrapped stdout/argv/exitCode restoration in try/finally inside the test body. CONTRIBUTING.md:356 forbids it outright, and the correct t.after() pattern was already in use two lines up in the same test. - tests/state-transition.test.cjs still listed 'clear' as an allowed FieldPreservation value in the row-enumeration test AND the getFieldClassification property test, after this PR deleted it. A stale allowlist weakens the property's negative space — it would accept a resurrected clear row as valid. Contract tension, resolved rather than left: ADR-3408 section 8.3 requires each ratchet entry carry the issue owning its removal. All four shipped with owner: null. The guard was right not to INVENT one, but the owners are known from the phase plan, so recording them is not inventing: phase.cts -> #3469, state.cts and milestone.cts -> #3471, health-diagnostic.cts -> sanctioned-permanent. Rather than a JSDoc caveat, --baseline now MERGES prior owner values on the (file, source) key, so a mechanical regeneration can no longer silently discard curated provenance. Verified by regenerating twice. * fix(#3468): sanitize attacker-controlled fields on every guard output path Isolated security review, MEDIUM, confidence 8/10. findSeamBypasses and findPromptSeamUses built findings with an UNSANITIZED `file`, while the co-located `source` on the same object was correctly wrapped in sanitizeForReport. On a fork PR a filename is exactly as attacker-controlled as a source fragment — a repo can legally track a filename carrying C1 control bytes or bidi overrides. The raw value reached two paths: --json stdout, and the COMMITTED baseline JSON via buildBaselineEntries. JSON.stringify neutralizes C0 controls but does NOT escape C1 (0x7f-0x9f) nor the bidi/zero-width range sanitizeForReport exists to strip — which is the precise threat the guard's own header names. Only the human formatter was safe. Sanitization now happens at CONSTRUCTION, so every consumer inherits it rather than each output path having to remember. The same defect was present on `field` and `policy` and is fixed alongside. Double-sanitization in the formatter is left in place, verified idempotent: escaped output is ASCII and cannot re-match the control/bidi classes. Also: the guard was not referenced anywhere in package.json, so nothing ran it. A drift guard nobody runs is not a guard, and ADR-3408 Decision 5 assumes it runs. Wired into lint:ci beside its sibling drift guards; it was already green on this tree, so the chain stays green. * chore(#3468): re-curate ratchet after an upstream rewording of a tracked bypass The rebase onto origin/next turned the guard red on its first real day, which is the ratchet working rather than a defect. |
||
|
|
be9329b10b |
fix(#3374): phase.complete stops harvesting stale body stopped_at (#3491)
* fix(#3374): phase.complete stops harvesting stale body stopped_at Variant A: cmdPhaseComplete's adapter calls syncStateFrontmatter directly (deliberately - STATE.md commits atomically with ROADMAP/REQUIREMENTS), which also bypassed the #948/#1230 preservation pass every RMW write gets. A stale body 'Stopped at:' line then silently clobbered a fresher frontmatter stopped_at on every phase completion, with warnings: []. Three layers close it without reversing #3517's refresh expectation: - completePhaseCore now refreshes the body continuity line it implies ('Phase N complete, ready to plan Phase N+1'; ADR-2207 phrasing on the last phase), session-scoped via the new stateReplaceFieldInSession seam so a decoy bold Stopped-at line in an unrelated section cannot absorb the refresh. Replace-only - a layout with no session line keeps its shape and its frontmatter value survives via the preservation delta. - the RMW post-sync preservation chunk (snapshots + table-driven applyStatePreservation + #2736 re-assert, full bodyDeltas wired) is extracted into the shared applyPostSyncPreservation helper; the phase.complete adapter and writeStateMd (milestone complete / state sync - the gap the closed PR #3442 review flagged) now run it too. - cmdStateRecordSession pushed 'Stopped At' onto updated[] on any label MATCH, including a value already on disk - reporting a write that never changed a byte. It now reports only on real change, and the match is tracked separately so an identical value does not arm the #944 DWIM section rewrite (which would reset an executor-authored resume file to None). * docs(#3374): backfill changeset pr field to 3491 * fix(#3374): drop the writeStateMd preservation pass - state sync's #905 contract is body-wins CI on this PR caught what the closed PR #3442 review's MAJOR remediation option (a) would have broken: state sync's #905 contract ('body annotation beats existing frontmatter when both are present') is the opposite by design - sync exists to re-derive frontmatter from the body. A blanket applyStatePreservation pass on writeStateMd re-locked stale frontmatter (current_phase 3 over the body's 5) on every sync. Take the review's sanctioned option (b) instead: the scope claim is accurate (phase.complete only) and the milestone complete / state sync exposure is tracked as follow-up issue #3492. --------- Co-authored-by: sim <sim@local> |
||
|
|
8bead8b0ff |
fix(#3395): own the phase line in planned-phase and persist --name (#3490)
* fix(#3395): own the phase line in planned-phase and persist --name * fix(#3395): backfill changeset pr 3490 --------- Co-authored-by: sim <sim@local> |
||
|
|
58e3437a48 |
fix(#3351): reconcile state.patch report with persisted state.md (#3487)
* fix(#3351): reconcile state.patch report with persisted state.md * chore(#3351): add changeset fragment * chore(#3351): backfill pr number in changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
69e7afd0c7 |
chore(#3212): bounded quantifiers over document content — prohibition with teeth — Phase 4 (#3441)
* feat(#3415): ship local/no-unbounded-quantifier, burn down ReDoS class Phase 4 of epic #3212 (ADR-3212 §5/§7, the final phase). New rule flags an unbounded */+/{n,} quantifier over a broad character class ([\s\S], dotAll ., or a 1-2-unit negated class like [^\n]/[^)\n] — the exact #2128-fixed shape) applied to a regex whose match target is data-flow-traced to readFileSync content. eslint-rules/lib/readfilesync-trace.cjs extracts the data-flow tracer shared with no-crlf-fragile-split (Phase 2) rather than a second copy — no-crlf-fragile-split refactored onto it with zero behavior change, parity-tested. Real triage, not 798 mechanical edits: the ADR's census (2026-08-08) screened every unbounded quantifier in the tree unscoped. Correctly scoped to readFileSync-derived content (matching Phase 2's own G2/G3 scoping), the rule found 162 real hits across two detection waves — the second wave (93) surfaced only after a genuine off-by-one bug in this rule's own first draft was caught while writing its RuleTester tests and fixed (the bug silently missed every directly-quantified [\s\S]* with no gap before the quantifier — exactly the class this rule exists to catch). 3 hits landed in production src/ (commands.cts, milestone.cts, roadmap.cts) and were each empirically timed against adversarial input (matching #2128's own measured-not-assumed precedent) — all confirmed linear-time/benign, left unbounded with a measured-evidence comment rather than mechanically bounded. The remaining 159 are test-file fixture parsing (test-author-controlled, fixed-size content, not adversarial input) — each suppressed with a specific, non-generic reason. Zero functional behavior changed anywhere in this diff. tests/no-pending-3212-markers.test.cjs locks the epic's own closing invariant (ADR §7: "assert zero pending #3212 markers remain") — ground truth confirmed trivially true today (no phase left any such marker behind), now regression-locked going forward. Design: .gsd/phase/chore-3415-prohibition-with-teeth/40-design.md Test matrix: .gsd/phase/chore-3415-prohibition-with-teeth/50-test-matrix.md Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): correct rule category mislabel, add CI test-scope entry An orthogonal Standards-axis review found eslint-rules/no-unbounded-quantifier.cjs mistakenly carried meta.docs.category: 'Portability', copied from a sibling rule without realizing what that implied: docs/contributing/cross-platform- portability-rules.md governs an ADR-1703 rule family under a hard "zero escape hatches" contract (tests/portability-rule-disable-ban.test.cjs's PROTECTED_RULES bans eslint-disable for those rules entirely). This rule is not part of that family — it's ADR-3212 (ReDoS/CWE-1333), a different epic — and its eslint-disable-next-line suppressions (159 of them, added earlier this same phase after empirical benign-verification) are an intentional, correct design, not a bypass. Corrected to category: 'Best Practices', matching the actual precedent (no-adhoc-regex-escape.cjs, Phase 1 of the same epic, which is also correctly outside PROTECTED_RULES), and the rule's own docstring now states this explicitly so a future reader doesn't have to re-derive it. Also registers a new scripts/ci-test-scope.cjs bucket so editing this rule or the shared eslint-rules/lib/readfilesync-trace.cjs helper re-runs their own test suites under targeted CI selection — was previously unregistered and invisible to that fast-path (this PR's own gsd-test checkpoint runs the full suite regardless, so this only affects future narrowly-scoped PRs). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): bound no-unbounded-quantifier's own scanner (CWE-1333, ironic) Security review found the rule meant to catch algorithmic-complexity bugs had one of its own: hasUnboundedBroadQuantifier's negated-class inner scan walked from each `[^` occurrence to the next `]` (or EOF) with no bound, while the outer loop only ever advanced by one character — O(n²) total work on a pattern with many unclosed `[^` runs. Runs unconditionally inside checkPattern on any `new RegExp('literal string')` argument in any linted file, before the (cheap) readFileSync data-flow gate — so a single crafted string literal, no valid regex syntax required, could make `npm run lint` / CI hang. Empirically confirmed both the bug and the fix: pre-fix, n=4000/8000/ 16000/32000 chars took 30.8/115.6/463.8/1874.3ms (~4x work per 2x n, quadratic); extrapolated, the 300000-char repro from the finding would run ~165s. Post-fix (bail the inner scan once units exceeds the rule's own 1-2-unit scope, rather than continuing to hunt for a closing `]`), the same 300000-char input runs in 8.7ms via the real rule module, independently reconfirmed at 18ms via a fresh Linter.verify() call. New regression row in tests/no-unbounded-quantifier.rule.test.cjs asserts the RuleTester run on a 50000-char adversarial pattern completes and returns a defined result — no wall-clock assertion (CLAUDE.md Clock Seams / local/no-elapsed-assertion). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(#3415): triage 3 new sites, re-raise ceiling after upstream batch next merged 12 more PRs during this PR's review. Two consequences: - tests/edit-phase.test.cjs (fix #3262, unrelated) added 3 new content.match(/<tag>([\s\S]*?)<\/tag>/) reads of this repo's own workflow .md content — the same Class A pattern as the ~159 sites already triaged elsewhere in this PR. Suppressed with the same established reason. - lint-allow-test-rule-refs' ratchet ceiling needed re-raising again (301 -> 303) for the same reason as the two prior bumps: organic growth from unrelated, already-reviewed PRs landing concurrently, not a defect in this branch's own diff. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
5fff839713 |
fix(#3258): honor all field-classification preservation rows (#3447)
* fix(#3258): honor all field-classification preservation rows * chore(#3258): set changeset pr to 3447 --------- Co-authored-by: sim <sim@local> |
||
|
|
6dbc124018 | enhance(#3180): the sibling validators share one envelope and one owner — Phase 12 (#3407) | ||
|
|
2b20b7e2cd |
fix(#3257): preserve full-line frontmatter comments through the parse→reconstruct pair + syncStateFrontmatter (#3387)
* test(#3257: full-line frontmatter comments survive the parse→reconstruct pair AND a mutating state verb parseYamlRegion dropped column-0 # comments and reconstructFrontmatter rebuilt from Object.entries alone, so full-line comments were silently destroyed on every mutating STATE verb. Add failing-first regressions: 3 unit tests for the public pair (comment between keys, leading+trailing, consecutive) and an e2e test running a state verb (state update) on a commented STATE.md — the e2e exercises syncStateFrontmatter's fresh-derivedFm rebuild path, which is the actual loss site the issue is filed against. RED — fails on next; fix follows. * fix(#3257: preserve full-line frontmatter comments through parse→reconstruct AND syncStateFrontmatter Carry column-0 # comments through the frontmatter pair via a Symbol-keyed channel (FULL_LINE_COMMENTS): parseYamlRegion captures ^# lines and attaches them to the next top-level key (leading) or a trailing slot; reconstructFrontmatter re-emits them in place. The Symbol is invisible to Object.entries/keys/JSON, so every existing reader is unchanged; the channel is created only when a comment is seen, so comment-less frontmatter is byte-identical. CRITICAL (isolated review): syncStateFrontmatter rebuilds its target via buildStateFrontmatter (fresh object) + an Object.keys carry-forward, both of which skip the Symbol — so the pair-preserving channel was lost on the very STATE verbs the issue names. Export propagateCommentChannel(source, target) from frontmatter.cts and call it in syncStateFrontmatter before reconstruct, copying the channel onto derivedFm (leading filtered to keys still present so a deleted key's annotation drops with it, trailing preserved). Decision A. * chore(#3257: add changeset fragment * chore(#3257: backfill changeset PR number (#3387) --------- Co-authored-by: sim <sim@local> |
||
|
|
23e6d49929 |
fix(#3233): no-op state update-progress when the milestone scan finds zero plans (#3375)
* test(#3233): zero plans (0/0) is a no-op; plans-but-none-done still writes 0% cmdStateUpdateProgress mapped 0/0 through clampPercent to 0% and rewrote the shipped Progress record after milestone close. Replace the stale 'handles zero plans gracefully' test (which asserted the buggy percent:0) with a #3233 no-op regression (100% record preserved, updated:false), and add a negative-space guard: plans exist but none done must still write a legitimate 0%. RED — fails on next; fix follows. * fix(#3233): no-op state update-progress when the milestone scan finds zero plans cmdStateUpdateProgress mapped 0/0 through clampPercent to 0% and unconditionally rewrote the body Progress line, so after /gsd-complete-milestone archived the phases (.planning/phases/ empty, scope COMPLETE) a routine update-progress run destroyed the shipped record ([██████████] 100% → [░░░░░░░░░░] 0%). Add an early-return no-op when totalPlans === 0 — mirroring the established scope-withholding no-op (stderr WARNING + {updated:false, reason}) and computeProgressPercent's null-for-empty contract ('nothing to measure' ≠ '0% done'). The legitimate 0% case (plans exist, none summarized) is unaffected: totalPlans > 0 reaches clampPercent(0, N>0) = 0 and writes 0% as before. * test(#3233): unshadow 'Progress field missing' — clear the zero-plans guard The new totalPlans===0 no-op guard fires before the 'Progress field not found' branch, so the existing 'returns error when Progress field missing' test (no phase dirs → 0 plans) was passing for the wrong reason and that branch lost coverage. Give that test a phase dir + PLAN so totalPlans > 0 clears the guard and it reaches the branch it is named for. (Isolated review finding.) * chore(#3233): add changeset fragment * chore(#3233): backfill changeset PR number (#3375) --------- Co-authored-by: sim <sim@local> |
||
|
|
5e951540af |
fix(#3162): resolve active state phase before drift scan (#3208)
* test(02-01): reproduce template state validation drift - derive command fixtures from the shipped STATE template - pair passed-verification drift with a clean opposite-result control * test(02-01): cover state phase resolution boundaries - exercise precedence conflicts fallbacks and fail-closed directory handling - prove canonical equality and reject outside-root verification evidence * docs: add changeset for PR #3208 * Address review feedback * fix(#3162): preserve phase validation after state refactor * test(#3162): align validation scope cases |
||
|
|
e87fb409ee |
enhance(#2573): stamp STATE.md with its commit and surface a freshness hint (#2622)
* enhance(#2573): stamp STATE.md with its commit and surface a commit-age freshness hint Adds a `state_head` stamp to STATE.md and derives a tri-state commit-age freshness proxy (state_commits_behind / state_commit_stale) through state.cjs's readStateHeadFreshness, surfaced on smart-entry signals and as health W024. The proxy is advisory: classify() deliberately does NOT consume it (ADR-1787 locks the classification/routing boundary — a signal, not a route). Composes with #3099 and #1882 (both merged to next after this branch): the commit-age proxy reads `state_head` while the LAST_ACTIVITY_UNPARSEABLE diagnostic reads `last_activity` — two different fields, not "two staleness signals on one field." A new regression test asserts a STATE.md carrying both an unparseable last_activity AND a valid state_head resolves each independently (diagnostic fires once; freshness reads state_head, commits_behind 0). Rebased onto next (flattened): resolved the add/add conflicts in src/smart-entry.cts (kept both the #2573 freshness import/derivation and the #3099 diagnostic import/call) and tests/smart-entry.unit.test.cjs (kept both describe blocks). Drift-ack for health.md's W024 row is unchanged (12348 B). Tests: smart-entry 62, state/state-transition/health/verify 639, all pass. * chore(#2573): allowlist health-validation test in the prompt-injection scan The scanner's `exec('` code-execution pattern matches the benign `re.exec('<phase-id>')` RegExp method calls in the phase-ID grammar tests (pre-existing: 16 such calls on next, this PR adds none). The file entered the diff-mode scan's changed-file set only because #2573's W024 state_head assertions touch it. Allowlist it alongside the other test files that carry pattern-matching content as data (same DEFECT.PROMPT-INJECTION-SCAN-COLLISION class). Scanner self-test 38/0; diff scan 14 files, 0 findings. |