next
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a9a7a328e6 |
refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD across contents and paths, upstream package/repo coordinates -> @golem15/msd-core and golem15com/msd-core. Deep links into upstream history, sibling upstream packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is. Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line, package/plugin identity, regenerated lockfile, install-tree fixtures, derived registries and benchmark baseline; migration checksum baseline re-locked (MSD keeps its own install state, so no install had applied the old sums); sort-order and regex-escaped expectations in tests adjusted. |
||
|
|
85545a77a5 |
fix(#4663): gate the canonicalization on the uat-passed predicate (#4809)
* test(#4663): add failing-first contract coverage for the blocked-uat canonicalization gate verify-work.md's complete_session step flips VERIFICATION.md to passed on 'zero issues' alone, so a session whose every UAT row is blocked (a session that observed nothing) canonicalizes the report. Pins the deployed contract the fix must satisfy: the flip runs the unflagged phase uat-passed predicate inside the human_needed branch, frontmatter.set sits inside a passed==true guard, a refusal message carries the blocker count and keeps human_needed, and an indeterminate pre-check fails closed. All four new assertions are RED until the workflow grows the guard. * fix(#4663): gate the canonicalization on the uat-passed predicate complete_session flipped VERIFICATION.md to passed whenever the session recorded zero issues and the status was human_needed — but blocked rows are not issues by this workflow's own rule, so a 0-passed / 0-issues / N-blocked session (one that observed nothing) rewrote the canonical report to passed. Every later reader (transition.md's preliminary check, resume paths, validate-phase, verification.status) then inherited the unearned pass while the phase-close predicate correctly refused it. The flip now runs the phase-close predicate in a new --uat-only form before canonicalizing: UAT rows evaluated (at least one pass, no pending/blocked/failed/unexplained-skip row), VERIFICATION-status blockers skipped — they must be, because the report still reads human_needed at pre-check time and that status is itself a blocking verification entry, so the full predicate could never pass there and the flip would deadlock (found by isolated review, probed). The --require-verification call stays the transition gate; the refusal branch reports the blocker count and keeps human_needed; an indeterminate pre-check fails closed. Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663) * test(#4663): align the canonicalize pre-check needles with the shipped line The workflow line carries a 2>/dev/null redirect the needles did not include, so both pre-check assertions fail against the committed fix (fixed-string grep verified). Reviewer-found; needle and message aligned. * fix(#4663): reword the canonicalize prose and refresh its size baseline The rationale paragraph mentioned the flagged transition-gate call by its flag, putting a --require-verification literal before the first phase uat-passed occurrence and breaking the existing ordering pin; the prose now describes it without the literal. verify-work.md's growth also drifted the committed compact-content baseline; regenerated via benchmark-compact-content.cjs --write (derived artifact, report-not-gate contract). Emitted-Drift-Ack-Growth: verify-work.md — canonicalize block gains the uat-only pre-check and refusal branch (#4663) * docs(#4663): backfill changeset PR number --------- Co-authored-by: sim <sim@local> |
||
|
|
cbbde6786a |
fix(#4546): deferred UAT follow-ups no longer block completion and promote to the backlog (#4769)
* test(#4546): failing-first tests for deferred uat follow-ups * chore(#4546): regenerate derived lists for the deferred-promotion suite The new verify-work-deferred-promotion suite changes the tests/ tree the macOS conformance-tier classifier tracks and is a novel file under the verify prefix in the test-file-count ratchet; both derived lists are regenerated/registered per their own guards' instructions. * fix(#4546): deferred uat follow-ups no longer block, and get promoted Two halves of one disconnect (#1921's deferral design vs the completion predicate): - uat-predicate: the item parser now captures the block's reason: line alongside result:. A skipped item whose reason carries the verify-work writer's 'Deferred follow-up:' template is a deliberate deferral -- non-blocking, flagged deferred in the report. Quote- tolerant (the writer wraps the value) and case-insensitive. A reasonless skip, a non-deferral reason, pending/blocked/issue/ failed/missing all still block, exactly as before. - verify-work complete_session: when the Deferred Follow-Ups section is non-empty, offer to promote the items to a ROADMAP.md 999.x backlog entry reusing next.md's prior_phase_completeness entry shape, with a --files-scoped commit. Offer, not auto-mutation -- matches the workflow's interactive convention and next.md's own prompt style. * chore(#4546): refresh compact-content benchmark baseline verify-work.md grew (the #4546 deferred-follow-up promotion offer in complete_session); the registered split's token counts moved with it. Baseline recomputed with the script's own --write. Emitted-Drift-Ack-Growth: verify-work.md — complete_session gained the deferred-follow-up promotion offer (detection, [P]/[K] choice, the next.md-shaped 999.x entry template, and the --files-scoped ROADMAP.md commit); the growth is the new contract text, not duplication * fix(#4546): gate/audit agreement and review fixes for deferred follow-ups - src/uat.cts categorizeItem: a skipped item carrying the deferred follow-up template reason now categorizes as 'deferred' (the category already existed for deferred-items.md entries) instead of being misfiled into the blocked families by keyword match -- the gate/audit agreement #3078-CR expects, restored in the permissive direction the #1921 design intends. Checked BEFORE the keyword families so '... on the release build next version' is not build_needed. - verify-work.md promotion step: numbering scans for the smallest free 999.n (count races + non-contiguous history), one backlog entry per deferred follow-up, ROADMAP.md-absent behavior specified, idea text newline-flattened, Deferred at placeholder harmonized with next.md. - DEFERRED_REASON_RE: trust assumption documented (authoring contract, not a security boundary; non-matching spellings block fail-closed). - tests: the property now drives evaluateUatPassed and derives expectations from the input spec (never restates the matcher), includes the no-result-line branch, and pins its seed; the parity test drops try/finally for the approved pattern, uses createTempDir, sites its allow-test-rule marker at the suppression site, and asserts the literal [P]/[K] choices. * fix(#4546): close promotion-test docstring, drop fc replay-path misuse, refresh baseline The final matrix run caught three defects in my own review-fix commit: the parity test file's JSDoc was left unterminated (the whole file parsed as one comment -- zero tests registered, hence the file-level 'test failed' the runner reported); fast-check's replay-path parameter was misused as a label (invalid path at replay); and the workflow-text ambiguity fixes re-drifted the compact-content benchmark baseline. * docs(#4546): add Fixed changeset for deferred follow-up coverage * docs(#4546): backfill changeset PR number * fix(#4546): use the pattern seam escapeRegex for shape-marker matching The hand-rolled metacharacter escape in the shape-marker assertion tripped local/no-adhoc-regex-escape, whose named remedy this adopts. --------- Co-authored-by: sim <sim@local> |
||
|
|
4e8927b0b9 |
fix(#3707): degrade the fold for every UAT gap class, and stop line endings hiding rows from the audit and the acceptance gate (#3903)
* test(#3707): failing-first coverage for reverting the fence-shortfall fold shield Pins the post-revert contract: a phase whose only gap is a fence shortfall must degrade the fold and withhold the milestone percentages, like every other gap class. Five of the eight rows are CONTROLS that pass before the change, and they carry more weight than the failing row. The failure mode of this revert is degrading TOO MUCH: a revert that sets foldScope outside the headingsSeen > 0 branch would withhold every percentage in the project, and only the no-gap control catches that. Another control catches a revert that collapses the two scopes into one and loses the distinction between what a phase reports and what the fold folds -- uat.scope must stay TRUNCATED for every gap either way, which it already is. The row that pinned the shielded behavior is rewritten rather than deleted. Deleting a test because the behavior it asserts is being reversed leaves the reversal unguarded. * fix(#3707): degrade the fold for every UAT gap class, reverting the fence-shortfall shield Maintainer decision. The two orthogonal engines split on this during #3707 and neither filed it as blocking, so it shipped in the shape the engine that raised the objection endorsed after verifying seven fixtures. The call has now gone the other way, restoring the fail-safe direction chosen twice already on this issue. The shield exempted one gap class from the fold's teeth. It could not do that safely: shortfallBlocks is a single tally incremented at exactly one site and spans BOTH a harmless fenced documentation sample AND a genuinely fence-straddled result: blocked row. Exempting it therefore could not exempt only the harmless case -- it also published a milestone percentage over a real, unread outstanding row. SCOPE.TRUNCATED means the scan could not SEE part of the evidence, which is exactly that case. scope and foldScope now agree: every gap class degrades both. The accepted over-report documented in uat.cts is unchanged and still documented there; what changed is only that it no longer buys an exemption from the fold. The comment block above it argued FOR the shield and is rewritten, because a comment defending behavior the code no longer has is worse than no comment. shortfallBlocks leaves this function's destructure but is untouched upstream, where audit-uat still consumes it. * fix(#3707): correct the caller comment, add the changeset, and name what the order tests guard Review found a SECOND comment still documenting the removed shield -- the caller's, beside the worstScope fold, stating that foldScope differs from scope for exactly one case which must not raise phase_scope_degraded or withhold the milestone's percentages. That is now the opposite of what the code does. I rewrote the buildUatRows comment in the previous commit and asserted in its message that a comment defending behavior the code no longer has is worse than no comment, then left exactly that one standing a few hundred lines away. The change had no changeset. It is user-visible: a milestone's percentage goes from published to withheld whenever any phase has a fence-shortfall-only gap. PR gates hard-fail a user-facing code diff without one. The two scopes are now identical at every return site. They are NOT collapsed -- that would change the return shape and the caller on what is meant to be a one-condition revert, and the seam is worth keeping if the distinction is ever wanted again -- but the declaration now says plainly that they agree by decision rather than by accident, so a reader does not have to re-derive it. The two order-independence tests were renamed. foldScope is monotonic with no reset path, so file order is structurally irrelevant and those rows could never have failed for the ordering reason their names promised. They do guard something real -- a multi-file phase degrading when any one file has a shortfall-only gap -- so they now say that instead. * test(#3707): failing-first coverage for the lone-CR UAT false-clean The parser splits on newline only, and the heading tokenizer agrees with it, so a lone carriage return is not a line boundary anywhere in it. CommonMark treats a lone CR as a line ending, so such a row renders to a human reader while being invisible to BOTH sides of the parser's symmetry invariant: no item, no shortfall, no headingsSeen. A phase hiding a result: blocked row this way reports 100 percent with zero diagnostics. Found by the security review of the fold-shield revert. It is the one false-clean class that revert does not reach, and it is the same bug class this issue exists to fix -- an unreadable row reported as clean. Nine rows. The LF control is what proves this is a separator defect rather than a content defect: identical bodies, one separator apart, and only one of them hides the row. CRLF and CR-inside-a-fence controls guard the coming normalization against double-counting or tearing content that legitimately contains a carriage return. Two further manifestations turned up while writing them: a leading CR breaks column-0 anchoring of the first heading, and an all-CR document flags a shortfall it cannot attribute to any row. * fix(#3707): treat a lone carriage return as a line ending in the UAT parser A lone CR was not a line boundary anywhere in the parser -- it split on newline only, and the heading tokenizer agreed with it. CommonMark treats a lone CR as a line ending, so such a row rendered to a human reader while being invisible to BOTH sides of the parser's symmetry invariant: no item, no shortfall, no headingsSeen. A phase hiding a result: blocked row that way reported 100 percent with zero diagnostics. Line endings are now normalized once at document ingress -- CRLF and lone CR both to newline -- at the two independent entry points, rather than teaching each split site about CR. Every downstream scan, offset and span therefore reads one convention. That single-frame property is deliberate: this issue already cost a HIGH when two scans read the same document through different frames. MY OWN END-TO-END TEST WAS WRONG and is replaced rather than weakened. It asserted that a lone-CR document must withhold its percentage, which reasons from the pre-fix symptom: after the fix the row is not hidden, it is surfaced, and this module deliberately keeps visible outstanding UAT work separate from completion percentages -- only unreadable evidence degrades scope. The success of the fix is what made the assertion false. The implementing agent refused to satisfy both it and the architecture and asked instead of bending either; it was right. What replaces it is a stronger contract: a lone-CR document and its LF twin, built from one source, must produce identical audit output -- scope, percent, every unresolved row by identity, and the diagnostic set. That is what 'a line-ending convention must not change what the audit reports' actually means, and it carries a non-vacuity check so it cannot pass with both sides empty. shortfallBlocks keeps being returned, now documented as currently unconsumed. An earlier reviewer told me audit-uat still consumed it and I passed that on as an instruction; it was wrong, and it was caught by checking rather than by me. * fix(#3707): normalize at the document read boundary, not at two call sites The lone-CR fix was half-applied and both review engines caught it independently. cmdAuditUat has four document ingresses, not the two I normalized: VERIFICATION.md and deferred-items.md still handed raw text to newline-only splitters, and the frontmatter extract in the UAT loop read raw content while its parser read normalized -- one audit entry mixing the two frames the fix exists to unify. Measured: a phase written twice from one source gave total_files 2 / total_items 4 under LF and results [] / total_items 0 under lone CR, with zero diagnostics. Normalizing two call sites and declaring it done is exactly why two were missed, so this moves it to the read boundary: every document now enters through a helper that normalizes, in audit-uat, in planning-inspect's readDocument, and in the shared verification-status read. Future parsers downstream get normalized text by construction rather than because someone remembered. That last seam also fixes an under-reporting case of the same root: a lone-CR VERIFICATION.md saying status: passed was read as missing, telling the user a verify step that had completed never ran. The parity test's load-bearing assertion is now marked as such. Four of its five equality checks still pass with the bug present -- only the unresolved-row identity differs -- so trimming that one as redundant would make the row vacuous. Second changeset added: the CR fix is user-visible independently of the fold revert, and one fragment covering both would have described neither. * test(#3707): failing-first coverage for the U+2028 and duplicate-result false-cleans Two more of the same class, both found by the security review of this branch and both reproduced before writing a line. normalizeLineEndings folds only carriage returns, but a JS /m anchor also treats U+2028 and U+2029 as line terminators while split on newline does not. That is the identical asymmetry the carriage-return bug exploited, one separator over, and worse in one respect: these are not CommonMark line endings, so a reader still sees the column-0 result: blocked that the tool discards. Measured: a scalar-internal result: pass placed after U+2028 wins over the real blocked line and the row disappears with no gap raised. Separately, and independent of any separator, a block with two column-0 result: lines resolves to the first with no ambiguity signalled. Prepending result: pass to a block therefore deletes an outstanding row silently; reversing the order surfaces it. Order deciding meaning is the defect, so the pair of rows pins the contract as ambiguity-is-a-gap rather than last-one-wins, leaving the fix room to implement the gap sensibly. Four controls: an ordinary marker in the same position (proving separator not content), legitimate U+2028 inside prose that must not be torn, a single result line, and a result line inside a fence that must not count as a second occurrence. * fix(#3707): scan result lines by split, not by a multiline anchor Two more false-cleans from the security review, both closed by the same change. A JS /m anchor treats U+2028 and U+2029 as line terminators while split on newline does not. A scalar-internal result: pass placed after one of those separators therefore matched as a line start and beat the real column-0 result: blocked, and the row vanished at 100 percent with no gap. Worse than the carriage-return case in one respect: these are not CommonMark line endings, so a reader still saw the blocked row the tool discarded. Separately, the non-global match returned the leftmost hit, so a block with two column-0 result: lines silently resolved to the first. Prepending result: pass deleted an outstanding row; reversing the order surfaced it. Order deciding meaning was the defect. Both close by scanning lines produced by split rather than by anchoring a regex inside the whole document: each line is tested on its own, and a count other than exactly one is reported as a parse gap instead of resolved to either candidate. I asked for U+2028 to be folded in normalizeLineEndings and that was wrong. Folding is length-preserving, so it would have made the U+2028 fixture byte-identical to the genuine two-result-line fixture -- while one requires a confident item and the other requires an ambiguity gap. No implementation can satisfy both once the distinguishing character is erased. The agent proved that and deviated rather than forcing it, which is why normalizeLineEndings still folds only carriage returns, now with a comment saying why. * fix(#3707): bound the ambiguity scan at the next heading-shaped line The split-based result scan regressed four pre-existing #3078/#3707 guards, each off by exactly one gap. My diagnosis was wrong. I read the off-by-one as double counting -- zero-result blocks taking both the new path and the pre-existing one -- and said to change the ambiguity condition from not-equal-one to greater-than-one. The agent checked and refused: the zero path was never duplicated. The real cause is double ATTRIBUTION. A block is sliced to the next TOKENIZED heading, so when the next row is untokenized -- hidden by a straddling fence, or indented and already counted by the shortfall scan -- that row's own result: line is absorbed into the previous block. The scan then saw two result lines across what are really two rows and raised a second, redundant gap on top of the one already counted elsewhere. Had the greater-than-one change gone in, the counts would have matched while the double attribution stayed. That is the compensating-adjustment failure I had asked it to refuse, and it did. The scan is now bounded at the first following heading-shaped line, either indent class, so a genuine same-block ambiguity is untouched while spillover from a row counted elsewhere is excluded. * fix(#3707): keep the U+2028 immunity, revert the ambiguity detection The ambiguity half of this change regressed the suite twice and is coming out. Attempt one double-attributed: a block is sliced to the next TOKENIZED heading, so when the real next row is untokenized its result: line was absorbed into the previous block and raised a second gap on a row already counted elsewhere. Four guards broke. Attempt two bounded the scan at the next heading-shaped line and broke thirty. An indented ### N. inside a block scalar is legitimate scalar CONTENT, not a heading, and truncating there defeats every #3078 guard that exists to stop scalar bodies being read as rows. Telling a genuinely hidden indented row apart from indented scalar text is a classification countUnattributedIndentedRows already owns; a raw regex does not have that information. What survives is the half that is sound and was never implicated in either regression: the result scan tests each line produced by split rather than anchoring a regex with the multiline flag over the whole block. split never treats U+2028 or U+2029 as a delimiter, so those separators can no longer manufacture a line start and steal a row. Everything else returns to first-match-wins, byte-identical to origin/next. The two tests pinning ambiguity-as-a-gap are removed with it, since the contract is no longer implemented here. The defect they described is real, pre-existing and independent of any separator -- result: pass before result: blocked silently deletes an outstanding row -- and it needs its own change with a scalar-aware counter rather than being wedged into a branch already carrying three fixes. * fix(#3707): correct the shared-seam rationale and restore U+2028 trailing text The revert left a stale rationale in core-utils, justifying the decision not to fold U+2028 by claiming uat.cts must tell a fake line start apart from a real second column-0 result: declaration that gets flagged as ambiguous. Nothing flags ambiguity any more; that behavior was reverted and the same file says so a few lines away. The decision is still right, the stated reason was false. This is the third stale comment this branch has shipped and had to fix, and the worst placed of them: core-utils is a shared leaf that every future document consumer will read for guidance. Rewritten to the true reason -- the scan tests each split line individually rather than anchoring over the block, so an exotic separator cannot manufacture a line start and folding is unnecessary. Also a real behavior delta I had not noticed. Dropping the multiline flag left the pattern's trailing .*$ in place, and dot never matches U+2028, so a genuine column-0 result: blocked whose TRAILING text contained one stopped parsing entirely -- a visible parse gap rather than a false clean, so fail-safe, but a regression against origin/next that nothing pinned. The trailing portion now matches any character and a test pins it by identity against its plain-LF twin. Plus the JSDoc orphaned when normalizeLineEndings moved to core-utils, and the changeset, which described neither the separator fix nor planning-inspect surfacing lone-CR rows. * fix(#3707): harden the acceptance gate, which had both halves of the same bug uat-predicate is a SECOND, independent UAT parser, and it is the one that decides phase uat-passed. It read raw bytes and anchored a multiline regex over unsplit text -- exactly the two defects this branch closed one module away in uat.cts. The consequence is worse than the audit surface it mirrors. Measured on identical bytes: a U+2028 scalar injection made the gate return passed true while planning inspect reported the same row as blocked and outstanding. The hardened surface and the gate disagreed, and the gate was the permissive one -- so a phase could be accepted over a row the audit could see and the gate could not. Both raw reads now go through the shared normalize seam and both scans test lines produced by split rather than anchoring over the document. First-match-wins, matching uat.cts; no ambiguity counting is reintroduced. Tests assert the AGREEMENT between the two surfaces rather than each separately, because divergence is the defect. Also finishes the same root cause one module over: phase complete's advisory pre-scan read raw bytes, so a lone-CR VERIFICATION.md lost its human_needed or gaps_found warning -- the fix verification.cts already got on this branch. And narrows the core-utils rationale I reworded last commit, which claimed consumers already avoid multiline anchors. uat.cts still has five over unsplit text. That is the fourth comment on this branch to assert something the code does not do, so it now states only what is true of core-utils itself. * fix(#3707): give structure and attribution different line frames, normalize the close audit Two more from review, and the first was a regression I introduced one commit earlier. Converting the gate's heading scan to split-then-match removed a detection origin/next had: a ### N. heading delimited by U+2028 was found by the old multiline scan and was not found after. So hardening the result scan quietly weakened the heading scan, and the gate stopped blocking on rows origin/next blocked -- the permissive direction, on the surface that decides acceptance. The insight I had missed is that the two scans need DIFFERENT frames. Heading detection is structure: there is no distinction to preserve, so it splits on newline or either exotic separator and finds a heading however it is delimited. The result scan is attribution: the newline-only frame is exactly what stops a scalar-internal result: from being read as a column-0 line, so it stays. One frame applied uniformly was the error. Second, a THIRD unnormalized parser family: the milestone-close audit read every artifact raw. A lone-CR VERIFICATION.md degraded to status unknown and was skipped, and deferred entries vanished outright -- measured as three items requiring decisions under LF and one under CR, on identical bytes. All nine scanner reads now normalize; six of them had the identical defect beyond the three review named. The acknowledge path stays deliberately raw, since it splices by byte offset, and now says so. Also pins the cross-newline result: divergence, and replaces three raw U+2028 literals in test source with escapes. A raw separator in a fixture is one formatter away from becoming an ordinary-character control that still passes -- vacuous in the only test pinning the separator fix. * fix(#3707): share one frame between the acknowledge writer and the audit reader Normalizing the audit scanners left the writer and the reader on different frames. cmdAuditAcknowledge derives its stored snapshot values from raw content -- correct for the SPLICE, which rewrites by byte offset -- but scanUatGaps and scanContextQuestions now recompute those same values from normalized content. For a lone-CR artifact the two can never match, so an acknowledgement never suppresses its item and it resurfaces on every audit: acknowledge became a silent no-op. Fail-safe in direction, since the item stays visible rather than being wrongly suppressed, but it is the writer and reader disagreeing about what a line is -- the exact class this branch exists to eliminate, and the fourth instance of it here. The derive functions now read a normalized copy while the splice keeps raw bytes and raw offsets, so both sides share one frame and the byte-offset rewrite is untouched. Round trip pinned for lone-CR and LF, with an existing LF marker asserted still recognised so the change cannot silently invalidate acknowledgements already in users' files. Also tightens an assertion that pinned this branch's own heading fix with a proxy: notStrictEqual against 'passed' also passes on 'pass', which IS a passing token, so it could not have caught a regression attributing a passing result to the recovered heading. It now pins the exact token. * chore(#3707): backfill changeset pr numbers Both fragments still carried the pr: 0 placeholder, which failed changeset-lint and docs-lint on PR 3903. The review had flagged the backfill as pending and I opened the PR without doing it. --------- Co-authored-by: sim <sim@local> |
||
|
|
59e7a677fe | fix(#3511): scope every phase-directory scan to the phase it belongs to (#3535) | ||
|
|
53ea8e0664 |
fix(#3057): make a guard's failure distinguishable from its benign result — Wave 1 (#3088)
* fix(#3057): refuse the write when the duplicate scan cannot complete writeManifest documents itself as a fail-closed duplicate guard: if any existing manifest shares plan_id with a different, non-terminal job_id it must refuse, because dispatching again would duplicate the external job. It could not honour that. The scan reads every sibling manifest looking for the duplicate, and an unreadable or unparseable sibling was `continue`d past. If the corrupt file was the one holding the live duplicate, the scan found nothing and a duplicate external job dispatched. The asymmetry is what gives it away: a malformed TARGET refused with malformed_existing because clobbering is unacceptable, while a malformed SIBLING was skipped — yet siblings are the only thing the duplicate check reads. Adds a scan_incomplete verdict that refuses and names the offending file, so an operator can quarantine or repair it. Fail-closed alone would let one stale corrupt manifest wedge every dispatch for that planning dir permanently; naming the file is what makes refusing survivable. malformed_existing is untouched, so the target/sibling distinction stays visible. The docstring is updated — it previously stated a rule the function did not keep. memFs() gains an optional failReads map so these branches are reachable at all; they had zero coverage because the fake could not express a per-file read fault. The signature is additive and every existing caller is unchanged. The regression is proved by a pair, not a single test. A control writes a readable sibling holding a genuine non-terminal duplicate and asserts duplicate_plan_id, establishing the scenario is real; the regression then makes that same path unreadable and asserts scan_incomplete. A first draft of this test used a corrupt-JSON fixture containing no plan_id at all while its comment claimed otherwise — it duplicated the unparseable-sibling case and proved nothing, which is the defect class this phase exists to remove. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3057): make a guard's failure distinguishable from its benign result Wave 1 of the negative-space backfill: the branches where a guard that could not verify something reported the same value it reports when everything is fine. That indistinguishability is the defect; every fix here makes the two states tellable apart, and every test proves it with a pair — one for the failure, one for the benign case. A single test cannot establish that two states are distinguishable, which is the whole property being fixed. state.cts phaseInventoryProvider returned null for both a real disk-scan failure and a genuinely empty phases dir, so `state rebuild` could report success while phase-table reconciliation never ran. It now returns a discriminated result and the CLI surfaces phase_inventory_scan_failed plus a reason. The reason field turned out never to have been wired into the emitted JSON at all — it existed only as an internal variable — so a test could only assert on the operator-facing note. It is a real field now. state.cts treated an unreadable lock body the same as an empty one, applying the 1-second stealable floor. A lock we cannot read is not a lock we know is stale; an unreadable body is now held to the deadman ceiling like a live holder. verification.cts findStaleVerificationSummary returned null on any fs, scan or clock failure — meaning "not stale". It now returns a discriminated StaleCheckResult and the caller records that the check was indeterminate. git-base-branch resolveBaseBranch returned 'main' both when no candidate branch existed and when every git tier timed out. A diagnostics variant now reports whether the answer was verified, and the CLI writes an unverified-fallback note to stderr. The stdout contract five workflows parse is untouched. worktree-safety snapshotWorktreeInventory left exists:true when statSync threw, so a guard that could not check reported the worktree present; exists is now tri-state and a stat failure surfaces as an 'unverified' finding. planWorktreePrune reported 'no_worktrees' for a parse failure, which is not the same as an empty list — and it drives a prune. It now reports 'parse_failed'. Fixing the inventory change exposed a second fail-open in verify.cts: the validate-health consumer silently dropped findings whose kind it did not recognise, so the new kind would have vanished. That is closed too — worth noting that the survey enumerated producers of degraded verdicts, not consumers that discard them. worktree-base-ref and state-transition gain the distinguishing signal without changing what they do: headAbsenceVerified, and a phase-inventory scan meta. Whether those guards should ACT differently is a product question this change does not answer, and both are flagged rather than quietly settled. rescueSummaryArtifacts is left alone: rescuing on an uncertain cat-file is deliberate per #2556. It now has tests proving it, and a recorded negative finding — git cat-file -e returns 128 for both "absent from HEAD" and a fatal error, so "uncertain" and "certain-and-fine" are not separable at the git level. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): assert typed values, not rendered text Ten assertions in the rebuild CLI suite matched substrings of produced output — STATE.md body fields, a markdown table row, an audit-log heading, and JSON keys read as text. CONTRIBUTING prohibits that: if the code under test produces text, the test asserts on its structured surface instead. No production surface had to be built. Every one already existed and was already compiled into bin/lib: stateExtractField for body fields, parseMarkdownTable for the phase table, collectSection for the audit-log section, and result.data.log — already a typed RebuildLogEntry[]. The tests were matching rendered text sitting next to the structured data. One of those assertions was passing for the wrong reason. `stdout.includes ('rebuilt')` matched the JSON KEY name, not a value: the dry-run path emits `mutated` and the real path emits `rebuilt`, so it would have passed whether the value was true or false. It now asserts the value. external-job's refusal already had to name the offending file — that naming is why the fail-closed variant is survivable rather than a permanent wedge — but the tests proved it by substring of a prose message. The failure result now carries offendingPath as its own field and the tests assert it by value. The human message is unchanged; operators read it. Array membership is left alone. `phaseIds.includes('99')` and `result.updated.includes('Completed Phases')` are membership checks on real arrays, not text matching, and converting them would weaken nothing and clarify nothing. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): execute acquireStateLock instead of grepping its source The non-EEXIST lock test asserted on the TEXT of the built .cjs and never called acquireStateLock. It carried an allow-test-rule: architectural-invariant exemption to permit that. A source grep proves a literal is present in a file, not that the behaviour works — it is weaker than a liveness test, which at least runs the code, and it was the only coverage the fatal-errno path had. Replaced with tests that inject the errno through fs and assert what actually happens: a fatal EACCES propagates out of acquireStateLock with zero backoff sleeps, while EAGAIN/EINTR/EINVAL/EIO/ENOENT/ESTALE/EPERM/EBUSY retry once and succeed. The exemption is removed and its allowlist entry with it. One old assertion is deliberately not carried over: it checked the retryable errnos were expressed as a Set rather than an inline literal. That is a shape check with no runtime signature; the behavioural tests fail if the code reverts to the old inline check, which is the regression it was really guarding. The #3057 lock-body tests move into that same file rather than a new one, which is what lint-test-file-count asks for and puts every acquireStateLock test in one place. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3057): surface an indeterminate staleness check to its callers An isolated review caught an inconsistency inside this wave. Two of the three "add the distinguishing signal" fixes wire through to something a user sees: git base-branch writes an unverified-fallback diagnostic to stderr, and an unverifiable worktree surfaces as a W020 finding. The third set staleCheckIndeterminate on readVerificationStatus's result and nothing read it. A signal nobody consumes leaves the fail-open exactly as silent as before: the staleness check could fail and the operator saw precisely what they would see if the answer were genuinely "not stale". That is the defect this issue exists to remove, so it is not defensible as scaffolding when its two siblings in the same change already wire through. All five callers now surface it, each through the channel it already had rather than a mechanism imposed uniformly: phase complete adds it to its existing warnings array and, on the blocked path, as an additive note on the error text; init and roadmap carry it as a field on output they already emit; the UAT report carries it without ever gating passed/blockers; workstream inventory takes an injectable writeDiagnostic mirroring the git base-branch idiom, because its return shape had nowhere to hang a per-phase field without rippling the builder's types. The routing decision is unchanged everywhere. What changes is only that a caller and an operator can now tell a failed check from a completed one. That diagnostic carries structured meta rather than being asserted by regex — the default still writes only the human message to stderr, but tests assert phaseDir and reason by value. Two earlier assertions in this branch were converted the same way; this was the last raw-text assertion left. Also records a scope correction: the completePhaseCore guards now compare stateReplaceField's result to the body instead of testing truthiness, so a field whose substitution produced identical text no longer reports as updated. That is a real behaviour fix, not the signal-only change this file was described as carrying, and its tests cover both the changed and unchanged cases. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): bound two heavy subprocesses for a loaded bench, not an idle one The remote matrix surfaced three failures unrelated to this branch's changes. All were bad tests, and a re-run would have hidden every one of them. The reviewer-flags parse block bounded bash -> node -> a full gsd-tools cold start at 5 seconds. On a bench running thirty thousand tests in parallel that is not a hang, it is a busy machine. Raised to 30s, matching the convention sibling suites already use for script invocations, with a comment saying what the budget covers so nobody tightens it back. Two further copies of the same 5-second spawn in the same file had the identical defect and are raised too — they were not in the failure report, but they will be next time. The fragment-propagation test bounded npm run regen:derived — a full build plus eight generators, the heaviest subprocess in the suite — at five minutes, and node22 was killed near the end. The captured output proves it: every generator had written its files and gen:install-tree had emitted all fifteen runtimes before the kill. Raised to fifteen minutes. That failure read as `null !== 0`, which says nothing. status null means killed, not a non-zero exit, and the two want different responses: one is a timeout to size correctly, the other is a real build break. The assertion now distinguishes them and names the signal. Neither test's assertions were weakened and no retry was added. A retry here would suppress exactly the signal the timeout exists to produce. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): capture fd 1 through the mock tracker, not a raw reassignment The phase suite reported zero test results on both lanes while running for five and a half minutes and exiting 1. No assertion text, no stderr, four events for the whole file: enqueue, start, dequeue, complete. That shape is not a failing assertion — it is the runner being unable to read the child at all, because it parses its event stream from the child's stdout. The cause was the capture helper reassigning fs.writeSync directly. Proven rather than assumed: a standalone probe patched fs.writeSync and called process.stdout.write, and the interception fired only when fd 1 resolved to a FILE, not when it was a pipe. The remote runner captures the event stream to a file, so a helper that was invisible against a pipe swallowed the reporter's own output on the bench. That is also why the two sibling suites wired the same way in this change pass cleanly — they use the mock tracker, the seam io.test.cjs established for this exact function. The helper now uses t.mock.method with an explicit restore after each call, so teardown belongs to node:test rather than a second hand-rolled implementation, and the interception cannot outlive the one synchronous call it wraps even if that call throws. Ten call sites thread the test context through; three test callbacks gained the parameter they lacked. The three B3 tests are untouched — same assertions, same fault injection. Only how the context reaches the helper changed. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): capture phase-complete output from a subprocess, not fd 1 Two attempts to make in-process fd-1 interception safe both failed on the bench. The suite reported zero test results on either lane while exiting 1 — four events for the whole file — because the runner parses its event stream from the child's stdout, and process.stdout.write routes through fs.writeSync whenever fd 1 resolves to a file, which is how the runner captures. Patching that seam anywhere in a file can therefore destroy the file's own reporting, and tightening the window only moved the runtime from 326s to 125s without recovering a single event. So the interception is gone rather than tuned. The helper now spawns gsd-tools as a real subprocess and reads stdout the way the OS already gives it to us, which is what the rest of the suite does. It asserts the command succeeded before parsing, so a genuine failure can no longer present as a JSON parse error. The two fault-injecting tests could not survive that move as written: a subprocess cannot see a mock installed in the parent. Instead of reinstating the interception they now produce the fault on disk — the summary artifact is created as a dangling symlink, so the staleness check's real statSync throws inside the child. That is a more honest fixture than a mock in any case, since it is a condition a user's tree can actually be in. Skipped on Windows, matching the existing symlink precedent in the write-guard suite. Three further call sites turned out to depend on parent-process writeFileSync mocks the subprocess could not see. Those call the CJS function directly, which is what they always wanted — they never needed stdout at all. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3057): one name for one signal, one encoding for one distinction Standards review found four things this branch introduced, all of them inconsistencies with itself rather than with the repo. One upstream bit reached its consumers under three names — verification_stale_check_indeterminate in two modules, the same value with "stale" dropped in a third, and stderr only in the fourth. Standardised on the long name wherever it is a field. The workstream inventory keeps its stderr channel, since its return shape has nowhere to hang a per-phase field without rippling the builder's types, but it now says the same word for the same thing. worktree-safety encoded one three-way distinction two ways in a single file: a named union for a finding's kind, and boolean|null for an inventory entry's existence. The second is now a named union too. Two assertions matched human prose because the blocked and non-blocked completion paths carried no typed field for the signal. Both now assert typed values. The first round of this fix added the field but left the regex beside it, which is the banned pattern sitting next to its own replacement; the second removed it and added an assertion on the reason enum so nothing was lost. The remaining two were reasoned away before being fixed, and both reasons were bad. "No typed surface exists" is the condition CONTRIBUTING says to fix by adding one — it took three lines. "The file already does this dozens of times" is not licence to add instance number thirty-one; a convention that violates a documented rule is debt, not precedent. Vocabulary differing across DIFFERENT modules is left alone: CONTEXT.md rejects a single shared result envelope, so per-module shapes are precedented, and a baseline smell does not outrank a documented standard. A census of every line this branch adds to a test file now finds no regex or substring assertion on produced prose: 87 strictEqual, 25 ok (all non-empty or shape guards), 12 equal, 3 throws (all typed err.code predicates), 3 deepStrictEqual, 2 notStrictEqual. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3057): backfill changeset pr number to 3088 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
77c7b4fc9d |
fix(#1522): enforce canonical verification before phase transition (#1548)
* fix: require fresh phase verification before transition * no-mistakes(review): Fix canonical verification closeout gates * no-mistakes(review): Fix verify-work frontmatter promotion command * no-mistakes(review): Fix stale verification gates * no-mistakes(review): Fix canonical verification routing gates * no-mistakes(review): Fix verification dependency and runtime routing gates * no-mistakes(review): Block stale verification bypasses * fix: handle large init manager outputs in verification workflows * chore: update changeset pr number * fix(verify-work): use fresh verification.status for stale gate The stale check after UAT used phase_completion.verification_status from session-start INIT while human_needed promotion already queried fresh verification.status. Align the stale gate with the canonical query so mid-session verification refresh is not ignored. * fix(init): skip roadmap-checked phases when selecting next_phase Roadmap-only phases without a disk directory were still promoted to next_phase when their checkbox was already checked. Exclude checkboxComplete phases so progress routing does not point at work the roadmap already marks done. * fix: gaps_found not overridden by stale, transition uses canonical verification - verification.cts: check gaps_found before stale so gap-closure routing is not masked by a newer summary mtime - phase.cts: remove redundant findStaleVerificationSummary — readVerificationStatus already handles stale detection - transition.md: replace raw grep on file content with verification.status query to avoid false-positive blocks from body text matching * ci: retrigger tests after rebase * fix(transition): replace gsd_run advisory check with awk frontmatter extraction The runtime launcher is not defined until the update_roadmap_and_state step bash block (~line 165). The early verify_completion block used gsd_run to query verification.status, which violated the runtime-launcher-parity test: 'preamble appears AFTER the first gsd_run reference'. Replace the gsd_run call with an awk-based frontmatter extractor that reads only the status: field between the two --- fences. This avoids both the preamble-ordering constraint and the original false-positive grep bug where body text like 'previous_status: gaps_found' would match a full-text regex. The phase.complete gate at update_roadmap_and_state is the canonical enforcement point; this early check is advisory only. Also update workflow-size-baseline.json for the updated transition.md size. Fixes: runtime-launcher-parity test (B) Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * fix: re-check verification under planning lock in phase complete Move readVerificationStatus into withPlanningLock so stale verification cannot slip through when a SUMMARY.md is written between the gate and the roadmap/state mutation. Return the blocked status from the lock callback and emit the error after release to avoid leaving .lock behind. * fix(transition): gate on canonical verification.status including stale Replace awk frontmatter read with verification.status query so transition blocks when summaries are newer than VERIFICATION.md, matching phase.complete and other workflows (autonomous, progress, verify-work). * Fix workflow verification gates for yolo transition and stale routing Require VERIFY_STATUS passed before yolo/interactive transition advance. Route stale verification recovery to verify-work, matching canonical projection. * fix(transition): use verification.status query for stale-aware advisory check The awk-based check read raw frontmatter status: passed, which misses the stale case where summaries are newer than the VERIFICATION.md file even though the frontmatter still says passed. The stale status is computed from file modification times, not stored in frontmatter. Move the preamble to the verify_completion bash block (the first block with a gsd_run call) so gsd_run query verification.status can be used for the advisory check. This gives the full readVerificationStatus logic including mtime-based staleness detection, matching the enforcement gate at phase.complete. Capture full JSON (VERIFY_JSON) so next_action can be included in the advisory output alongside the status. Also update workflow-size-baseline.json for the updated transition.md size. Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * ci: trigger test matrix for 525b946 Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * fix(transition): restore awk frontmatter extraction for pre-shim verification check The gsd_run launcher shim is not defined until line ~163 of transition.md, so the verification debt check at line ~80 cannot use gsd_run. Restore the awk-based frontmatter extraction that correctly reads status without needing the runtime, and restore the shim at its proper location before phase.complete. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(#1522): clarify transition verification gate wording * fix(#1522): update transition workflow size baseline * fix(#1522): update workflow-size-baseline after rebase onto next Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> * fix(#1522): guard findStaleVerificationSummary FS calls + thread opts.fs seam (review) Address review blocker B1 on #1548: findStaleVerificationSummary ran fs.readdirSync and two fs.statSync calls unguarded between readVerificationStatus's try/catch sections, so a TOCTOU race (a SUMMARY listed by scanPhasePlans then removed before statSync) or any FS error threw uncaught into callers NOT under the planning lock (init.manager / init.progress / uat-predicate). Wrap the body in try/catch degrading to 'not stale', and thread the injectable opts.fs seam (add statSync to FsLike, pass fsImpl from the caller) for parity with readVerificationStatus's no-throw contract and testability. Also adds the Verification Module glossary entry to CONTEXT.md (review B3). --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
9223f2f4c8 |
feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results (#1063)
* feat(#247): runtime-neutral phase uat-passed predicate from HUMAN-UAT results Wire the already-reserved `phase.uat-passed` alias (subcommand `uat-passed`, mutation:false) into the phase command router with a new markdown-aware predicate that evaluates HUMAN-UAT results and reports pass only when every required check passes. Post-SDK-retirement (ADR-0174/#174) successor to the SDK-framed #70, with no SDK-specific API surface. New pure module src/uat-predicate.cts: - stripFalsePositiveContexts: frontmatter -> HTML-comment -> CommonMark-style fenced-block state machine (tracks delimiter char+length) -> blockquote, each a small composable step, so a `result: passed` inside frontmatter, a fenced/~~~ block (incl. ~~~ nested in a ``` fence), a comment, or a blockquote is never counted. - parseUatResultItems: heading-block parser, column-0-anchored same-line result; a heading with no result -> `missing` (fail-closed). - analyzeMarkdown: unterminated fence/comment detection (malformed -> blocker). - evaluateUatPassed: allowlist pass/verification semantics; passed = no blockers && >=1 check && all passing; no_uat_artifacts discriminator (no vacuous pass); optional requireVerification policy hook. Thin cmdPhaseUatPassed handler in phase.cts; router closure rejects unknown flags via makeInvalidArgs. Hardened across two Codex adversarial passes (vacuous pass, dropped failing tests, permissive verification status, nested-fence escape, cross-line result value, masked unterminated comment) — all fixed fail-closed. New unit + CLI-integration suites incl. a fast-check property test; docs, CONTEXT glossary, inventory, and changeset updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#247): backfill changeset PR number (#1063) * fix(#247): indexOf paired-scan for unterminated-comment detection CodeQL js/incomplete-multi-character-sanitization (high) flagged the `raw.replace(/<!--[\s\S]*?-->/g,'')`-then-`.includes('<!--')` detection in analyzeMarkdown as incomplete sanitization (a single regex pass can leave a residual `<!--`). Replace it with a paired left-to-right indexOf scan that contains no `.replace()` of the comment token — CodeQL-clean and strictly more correct (a closed earlier comment can never mask a later unterminated one). Behaviour unchanged; 98 predicate tests + scoped docker run green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |