b6dd4e2e744e859b3591da4eb6b8245d62039327
83 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b6dd4e2e74 |
fix(#3936): dispatch quick research with researcher persona and model (#4158)
* test(#3936): quick research dispatch must use researcher persona/model Regression tests for the planner-persona leak in quick's Step 4.75 research dispatch: init quick emits researcher_model, the workflow resolves AGENT_SKILLS_RESEARCHER and parses researcher_model, and the researcher dispatch uses the researcher persona/tier while planner and executor dispatches keep their own. * fix(#3936): dispatch quick research with researcher persona and model cmdInitQuick now emits researcher_model (parity with the plan-phase init), quick.md resolves AGENT_SKILLS_RESEARCHER and parses the new field, and Step 4.75's dispatch swaps ${AGENT_SKILLS_PLANNER}/ {planner_model} for the researcher's own persona and tier. The #2517 model-omission list gains researcher_model. Emitted-Drift-Ack-Growth: quick.md — researcher persona resolution line + researcher_model parse-list entry (#3936) * chore(#3936): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
41466e8e88 |
fix(#4023): preserve decimal phase ids in init progress ordering and smart-entry output (#4110)
* test(#4023): reproduce decimal phase-id coercions * fix(#4023): preserve decimal phase ids in progress signals * test(#4023): align phase token contract expectations * chore(#4023): point the changeset at PR #4110 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
9f219d05ba |
fix(#3964): route init's waiting-signal, codebase, and skill-manifest paths through the project-aware resolver (#3971)
* test(#3964): waiting_signal, codebase_dir, and skill-manifest must be project-scoped * fix(#3964): route waiting_signal, codebase_dir, and skill-manifest through the project-aware resolver * chore(#3964): changeset fragment (pr number backfilled after PR creation) * chore(#3964): backfill changeset PR number (3971) * test(#3964): assert on the POSIX-normalized codebase_dir across platforms --------- Co-authored-by: sim <sim@local> |
||
|
|
ddeb141bd2 |
fix(#3749): route migrate-config, project_exists, and health repairs through the project-aware resolver (#3955)
* test(#3749): GSD_PROJECT-scoped migrate-config, project_exists, and repair paths * fix(#3749): route migrate-config, project_exists, and health repairs through the project-aware resolver * chore(#3749): changeset fragment (pr number backfilled after PR creation) * chore(#3749): backfill changeset PR number (3955) * test(#3749): assert on the POSIX-normalized project_path across platforms --------- Co-authored-by: sim <sim@local> |
||
|
|
929e02cb2c |
enhance(#3885): no silent swallow, and no verdict manufactured from dropped data (#3925)
* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking answer. Three families do exactly that today; this commit pins each one RED. Measured on this tree, 2026-08-27: intel query, .planning/intel/file-roles.json nested 12000 deep -> exit 1, "Error: Maximum call stack size exceeded" searchJsonEntries / matchesInValue carry no depth parameter at all. The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage never received it. same fixture nested 48 and 49 deep -> both return total=1 at exit 0, truncated=undefined Nothing distinguishes "searched to the bottom" from "stopped looking". query phase-plan-index, a plan whose depends_on names an unresolvable token -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in wave 1"] The token is never mentioned. computeDependencyLevels drops the edge with `if (!resolvedDep) continue;`, every plan becomes a root, and the tool then reports the author's correct wave: as the thing that is wrong. countPhasePlansAndSummaries with fs.readdirSync throwing EACCES -> hasContext:false, indistinguishable from a phase that simply has no CONTEXT.md. context_read_error is undefined. The shapes these tests assert against, chosen here so the implementation has a target rather than inventing one later: `truncated: boolean` on the intel query result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and `context_read_error: string | null` per analyzed phase. Deliberately green, and they must stay that way — each stops the fix from over-firing: depth 48 is found and NOT flagged truncated (the ceiling is inclusive) a shallow miss reports no truncation (noise control, N1) 10,000 siblings at depth 2 are unaffected (the bound is DEPTH, N2) a genuine wave: mismatch on a fully-resolved DAG still warns (N3) a genuinely missing directory is absent, not an error the emitted depends_on display mapping still passes an unresolved token through verbatim — already pinned by the existing #3785 test, so no duplicate was added T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the real CLI and reads the emitted JSON, because a unit assertion on computeDependencyLevels would have passed throughout #3427's life. Design: .gsd/phase/feat-3885-no-silent-swallow/40-design.md Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * enhance(#3885): no silent swallow, and no verdict manufactured from dropped data Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed into an output that reads as authoritative. The recursion bound, restored but NOT verbatim (src/intel.cts) MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and matchesInValue, which carried no depth parameter at all. The bound existed in the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage never received it — §8.3's "a consolidation may not delete an invariant along with the surface that held it", demonstrated. Measured before: a .planning/intel file nested 12000 deep exits 1 with "Error: Maximum call stack size exceeded". Reachable from a project document. The original returned a bare `false` at the ceiling. Restoring that verbatim would trade a crash for a silent "no match" when the truth is "I stopped looking" — the same class this epic exists to close, and ADR-3473 Decision 4 forbids it. So the bound carries a truncation signal: nesting 47 -> found, truncated false nesting 48 -> found, truncated false (the ceiling is inclusive) nesting 49 -> not found, truncated TRUE nesting 12000 -> exit 0, truncated TRUE, no RangeError A shallow document that simply has no match reports truncated FALSE — the flag means "I stopped early", never "I found nothing", or it would be noise. The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected. The dropped edge is named, and stops being blamed on the author (src/phase.cts) computeDependencyLevels dropped every unresolvable depends_on token with a bare `continue`. Each drop makes a plan a root, so the whole phase collapses to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as the thing that was wrong. Before: warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in wave 1"] After: warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not resolve to any plan in this phase — edge dropped, wave placement for this plan may be unreliable"] The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is deliberately not built here. The emitted depends_on display mapping still passes an unresolved token through verbatim (#3785). No artifact from failed inputs (gsd-core/workflows/review.md, #3352) A failed lane leaves no result file, so "every lane failed" is exactly "the aggregate JSONL has zero lines" — the gate condition already existed as a byproduct. REVIEWS.md is no longer written in that case, and the commit step is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT counted as a failure. Per-lane output and non-empty .err are preserved to .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that the lanes failed at all; the commit step names one file, never a glob, so the diagnostics are not swept in. Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2) Four callers collapsed an EACCES on a phase directory into [] and reported hasContext:false — byte-identical to a phase that simply has no CONTEXT.md. Each now names the directory it could not read. A genuinely missing directory stays absent rather than becoming an error, which is what keeps the fix from over-firing. Fatal errno folded into a retry set: audited, no defect found Reported as a verified negative rather than padded with a change. withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776; atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they return or rethrow the final error rather than swallowing it. estimate-cli's sole caller surfaces that rethrow as write_error in its JSON output. Manufacturing a diff to make the checkbox look worked-on is the Goodhart outcome Decision 6 exists to prevent. Disclosed: R46 (the commit step names one file, never a glob) is a real regression guard but is NOT independently failing-first — the commit fence is byte-identical pre- and post-fix, so it only fails pre-fix through its shared extraction dependency. Recorded rather than claimed as fail-first. Design: .gsd/phase/feat-3885-no-silent-swallow/40-design.md Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence Two review findings, both real, both in my own change. An isolated adversarial review found the evidence-preservation block never checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in a SEPARATE fenced block. A disk-full or unwritable phase directory therefore still destroyed the only copy of the failed lanes' output — reintroducing the exact #3352 data loss this item exists to stop, inside the fix for it. Preservation and cleanup are now one block, because each fenced block is a separate execution and a shell variable cannot carry between them. mkdir -p and each cp are exit-checked; cleanup runs only when preservation succeeded, and a failure warns naming the intact run directory. "Nothing to preserve" is not a failure and still cleans up. Driven three ways: success removes run_dir, failure leaves it intact with the warning, nothing-to-preserve removes it. The failure is induced by a file-vs-directory conflict rather than chmod 0o000, which root bypasses. The new unresolved-depends_on warning embedded a user-authored token verbatim: warnings: ["Plan 03-02: depends_on token \"evil Plan 03-01: FORGED WARNING\" does not resolve ..."] The JSON wire form is safe, and the security reviewer judged it non-exploitable for that reason. It is escaped anyway through formatDiagnosticToken — the helper #3884 added one phase earlier for exactly this class. warnings[] is an array a consumer naturally prints line by line, and not reusing the sibling fix is the generative-fix-divergence shape this epic exists to close. The same treatment is applied to context_read_error / phase_dir_read_error, which embed a phase directory path a repository can choose, and to the fs error message, which echoes the raw path itself. Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per negative space N2, disclosed rather than left to be discovered. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot" Blocker from the round-2 isolated review, and it is my own inconsistency: this phase applied "unreadable is not absent" to phase directories and left it broken in the file it was already editing. chmod 000 .planning/intel/file-roles.json gsd-tools intel query <term> -> {"matches":[],"total":0,"truncated":false} exit 0 safeReadJson swallowed every read failure and returned null, so an EACCES was byte-indistinguishable from an absent file AND from a genuine no-match. Now it separates three states: ENOENT stays silently absent, because not every project has every intel file and intelQuery loops over all of them expecting misses; EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel file previously read as "no matches" too — same defect, same fix. Threading that outcome through the other three callers found something worse than the reported case. intelDiff returned no_baseline:true for a corrupt or unreadable snapshot — not a silent failure but an actively FALSE verdict, telling the caller they never took a snapshot when they did. That is §8.5's headline case, so it is fixed and tested rather than noted. intelStatus and intelApiSurface collapsed the same way; intelApiSurface additionally printed a "not yet populated" banner that was simply untrue. Every row is failing-first, including the absent-file ones — the field is new, so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin that the fix does not OVER-fire on the ordinary absent case, which is what would turn this into noise on every project lacking an intel file. IO failure is injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as a test. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): build the pathological intel fixture as text, not by stringifying a nested object The remote runner came back red on Linux with two failures, both T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on macOS. The product was never at fault. writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned. Linux's container stack is smaller than macOS's, which is the whole of the platform difference. Measured, with the same document built as JSON TEXT so nothing in the building process recurses: depth=100 rc=0 truncated=true depth=5000 rc=0 truncated=true depth=12000 rc=0 truncated=true depth=60000 rc=0 truncated=true V8 parses this shape iteratively; only stringify recurses. The bound works at every depth tried. The fixture is now built by string concatenation. That is also the more faithful input — a real deeply nested JSON document on disk is exactly what the bound guards, where a stringified object was only ever a way to produce one. The depth stays 12000. Lowering it would have made the test pass by weakening it to accommodate a fixture bug, and 12000 is a legitimate pathological input the product handles. T4 remains a genuine fail-first: rebuilt against the parent of the commit that added the bound, the string-built depth-12000 fixture still drives the CLI to rc=1 with "Error: Maximum call stack size exceeded". A comment records why the fixture is text, so it is not "simplified" back into a macOS-green / Linux-red test. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3885): backfill the changeset PR number Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): normalize path separators before splicing into the workflow's bash CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the remote runner were all green. AssertionError: commit must name the single REVIEWS.md file; got: --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The harness spliced an OS-native temp path into the extracted bash, and bash consumes \U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory survived — which is the other two assertions. This is a fixture defect, not a product one, and that was checked rather than assumed. In production the phase directory is toPosixPath-normalized at every call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and the run directory is created by "mktemp -d" running inside the bash block itself (gsd-core/workflows/review.md:163), which emits POSIX-style output even under Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it. The file's pre-existing #3034 harness splices raw native paths too, but only ever inside double-quoted assignments, so it never tripped this — my new harness followed that convention faithfully into the one place where it does not hold. Both now splice through toPosixPath from shell-command-projection, the established seam, which is a no-op on POSIX and mirrors what production does. No assertion was weakened. "commit must name the single REVIEWS.md file" and "the run dir must still be destroyed" still assert exactly that; only how the fixture supplies its path changed. Nothing is skipped on Windows — a t.skip() here would have hidden the question of whether the exposure was real, which is the question that mattered. Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input produces a byte-identical shape, proving the normalization is idempotent. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3885): stop the harness making the deleted run dir its own cwd Windows shard 3/3 stayed red after the separator fix, on two assertions the separator fix never touched: AssertionError: the run dir must still be destroyed AssertionError: nothing to preserve is not a failure — run dir must still be removed The separators were a real bug and fixing them fixed the --files assertion. They were not this bug, and two CI cycles went into the wrong axis before I stopped converting path forms and looked at what the harness actually does. runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's working directory WAS the directory the block under test then removes with rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live process's working directory cannot be removed. So on Windows the directory survived and both assertions failed, on macOS and Linux it vanished and they passed. Nothing to do with slashes. Harness-only. Production never cd's into the run directory; every reference is by absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself (gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged. Fix: the child now runs with its cwd in an unrelated temp directory that the block under test never deletes. Neither assertion was weakened, and nothing is skipped on Windows — the tests in this file carry no platform guard and run there unconditionally, which is how this surfaced at all. Honest limit: the Windows failure mode cannot be reproduced on macOS, because POSIX permits the very thing Windows refuses. The diagnosis is grounded in that documented divergence and in the fact that only the Windows lane failed, but the green outcome on windows-latest is unverified until CI runs it. Refs #3885 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
389bc86e02 |
enhance(#3883): route every slug call site through its canonical owner (#3896)
* docs(#3883): correct section 8.3 against the tree
I wrote this section in Phase 0 stating rules I had not executed against the
tree. Measured at
|
||
|
|
a638ca4332 |
enhance(#3882): stop sentinel phases skewing estimation calibration (#3893)
* test(#3882): failing-first rows for sentinel phases skewing calibration Adds A1a/A1b/A2/A3 to tests/estimate-calibrate.test.cjs, the module's existing test file, rather than a new bug-NNNN file. collectCalibrationSamples (src/estimate-cli.cts:206) does a raw readdirSync over .planning/phases and never applies isSentinelPhaseId, so a sentinel phase (milestone 0 or 999) carrying a PLAN estimate / SUMMARY actuals pair contributes a phantom calibration sample. computeCalibration is median-based, so a single 50x outlier among three samples leaves the factor unmoved — asserting "the factor is unchanged" against one sentinel would pass on the broken code for the wrong reason. Each row instead asserts the WHOLE computed CalibrationResult object (factor, applied, confidence, sampleCount, clamped) for a sentinel-free project against its sentinel-injected twin: - A1a: one sentinel flips applied false->true and confidence low->med on phantom evidence (calibration switches on with zero real signal). - A1b: two sentinels corrupt the factor itself (1 -> 3, clamped false->true). - A2: the sentinel's own sample is verified absent from the returned list. - A3: the two genuine phases still contribute their own unchanged samples (regression pin — stops A1/A2 passing by filtering everything). Verified RED on today's code (node tests/estimate-calibrate.test.cjs): A1a/A1b/A2 fail with the exact differing objects; A3 and all pre-existing rows in the file remain green (no collateral). Refs #3882 * feat(#3882): route phase enumeration through its owner and name the sentinel axis Task 1: collectCalibrationSamples (src/estimate-cli.cts) hand-rolled a raw readdirSync over .planning/phases, treating every directory (including sentinel phases, milestone 0/999) as a completed phase and feeding phantom PLAN/SUMMARY samples into the estimation calibration factor. Routed through the existing owner, listMilestonePhaseDirs(phasesRoot) with no cwd -- already 'all milestones, sentinels excluded', exactly the combination this caller needs; no new API was required for this half. It now also surfaces the scope discriminator: an unreadable phases directory throws PhasesUnreadableError instead of silently returning zero samples, and cmdEstimateCalibrate reports it via a new ERROR_REASON.ESTIMATE_PHASES_UNREADABLE instead of persisting a phantom empty calibration document. Task 2: added listAllPhaseDirs(phasesDir, { includeSentinels }) to src/phase-locator.cts -- the one genuinely missing axis: 'physical set, sentinels INCLUDED'. includeSentinels has no default and is required, so a call site cannot obtain sentinel-inclusion by omission (compile-time refusal, not just documentation). Mirrors listMilestonePhaseDirs's absent/unreadable scope handling. Task 3: migrated the two exemptions whose written reason maps cleanly onto 'physical set, sentinels included' -- cmdRoadmapAnalyze's _phaseDirNames (src/roadmap.cts) and cmdInitMilestoneOp's diskPhaseDirs (src/init.cts), both heading->directory lookup indexes. Left the rest: archivePhaseDirectories's own body has no readdirSync to migrate (its callers already resolve dirs before calling it, and both current callers deliberately EXCLUDE sentinels -- migrating it would be an unauthorized behavior change, not an API swap); cmdValidateHealth's exemption is vestigial (its actual physical-set sweep already lives in planning-snapshot.cts's buildAllPhaseDirNamesField, a pre-existing near-duplicate of the new axis, flagged as a finding, not restructured); cmdPhasesClear/cmdMilestoneComplete/cmdVerifySchemaDrift/detectHasPriorPhases/detectUiPhaseActive want a different combination (sentinels excluded, or a single-phase lookup) and are unaffected. Task 4: detector 2 (sentinel literal) is untouched and retained. Removed exemption entries only for the two migrated call sites; every other function-scoped exemption is preserved. Guard exits 0. Refs #3882 * refactor(#3882): delegate the snapshot phase-dir scan to its owner buildAllPhaseDirNamesField duplicated listAllPhaseDirs's own readdirSync + directory-filter + absent/unreadable handling — the 'one implementation per rule' defect ADR-3473 SS8.3 names, introduced by this branch's own #3882 work. Delegate to listAllPhaseDirs and re-apply the field's existing lexicographic sort on top, since W007's observable order must not change. Refs #3882 * docs(#3882): document the sentinel axis and the enumeration consolidation Records listAllPhaseDirs in the Phase Locator glossary entry, and the fact that the owner already answers the all-milestones sentinel-free question when called without a cwd -- the call collectCalibrationSamples was missing. Also notes that buildAllPhaseDirNamesField now delegates rather than carrying a second readdir, and that exactly one readdirSync over the phases directory remains across the two modules. Refs #3882 * test(#3882): close review findings — real order proof, unreadable coverage, collision fixtures Refs #3882 * chore(#3882): backfill changeset PR number Refs #3882 --------- Co-authored-by: sim <sim@local> |
||
|
|
79781e68eb |
enhance(#2401): ground verify-command paths and inherit prior-phase commands (#3678)
* feat(#2401): ground <automated> verify-command paths and inherit prior-phase commands Adds a deterministic resolvability probe over each PLAN.md <automated> verify command and surfaces the nearest prior phase's proven commands to the planner at every context window. - src/verify-command-grounding.cts: recognizer (not a shell interpreter) that grounds a leading cd <literal> chain and npm --prefix <literal>, and reports unresolvable rather than guessing. Never executes command text. - gsd-tools check verify-command-paths <N>: per-phase probe, wired into plan-phase.md before the plan-check pass. - init.plan-phase gains prior_verify_commands, ungated by context_window. - gsd-plan-checker: new Verify Command Path Resolvability dimension that reports the failing target and never prescribes a replacement. Also fixes first-match-wins prefix bucketing in scripts/lint-test-file-count.cjs (readdir order is not stable across platforms, so a module whose name extends another's with a hyphen bucketed differently on Linux than on macOS). Closes #2401 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2401): ground the canonical --prefix form, quoted paths, and absolute cd resets Independent review found three defects in the recognizer: - npm --prefix DIR run SCRIPT never reached the script-existence check, because the pattern required npm and run to be adjacent. That is the form the docs tell planners to prefer, so script_missing never fired for it. The prefix flag and its value are now stripped before matching. - --prefix captured with \S+, so a quoted path containing a space was truncated to a stray opening quote and reported as a missing directory - a false blocker, worse than the bug this feature fixes. The capture is now quote-aware. - A chained cd whose later segment was absolute concatenated instead of resetting, producing a nonsense path and another false blocker. The fold now resets on an absolute segment. Also replaces the bespoke phase-directory regex with the canonical phase-id helpers. Real phase directories are NN-slug, not phase-N-slug, so the prior-command harvest matched nothing outside its own fixtures and the planner-inheritance half of this feature was dead code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(#2401): source task blocks from the canonical sectionizer The module carried its own copy of the <task>-block grammar - a fourth hand-rolled mirror of the one markdown-sectionizer owns. verify.cts keeps its copy only because it needs the type= attribute the canonical helper discards; this module never reads that attribute, so it can share the owner outright instead of adding a test around a copy. extractAutomatedCommands now takes task bodies from extractTaggedBlocks and the out-of-task remainder from stripTaggedBlocks. A task-grammar parity test pins the attributed task-name set against the canonical helper across six awkward task shapes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2401): extract agent-file overflow to references and repair the property arbitrary The remote matrix run came back red with 19 failures, four root causes: - agents/gsd-plan-checker.md and agents/gsd-planner.md both blew the 49152 agent cap. Their bodies move to gsd-core/references/, leaving @-reference stubs, per the documented overflow pattern. - The new checker dimension invoked gsd_run before the canonical preamble that defines it. The call is deleted outright: plan-phase.md already runs the probe and hands the result in as {VERIFY_PATHS}, so the dimension consumes that rather than re-running anything. - fc.fullUnicodeString does not exist in fast-check 4.8.0. Replaced with fc.string({ unit: 'binary' }), which covers the same 0000-10FFFF range. - Three runtime-loaded files grew; acknowledged in the existing ack fragments that already own those bare filenames, since two ack sources may never name the same path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2401): regenerate golden install-tree fixtures for the new references Adding two files under gsd-core/references/ changes what the installer emits into every runtime's tree, so all 19 golden install-parity fixtures went stale. Regenerated with npm run gen:install-tree; the delta is exactly the two new reference paths per runtime, no removals. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2401): backfill changeset pr number to 3678 * fix(#2401): treat ~ as a home expansion only at the start of a path Windows CI caught this on both shards; the Linux-only remote matrix cannot see it. The dynamic-path refusal rejected ~ anywhere, and a GitHub Windows runner's tmpdir is an 8.3 short name - C:\Users\RUNNER~1\AppData\Local\Temp - so a valid absolute Windows path came back unresolvable/dynamic_path. This was a production bug, not a test artifact: any Windows user whose project path carries an 8.3 short name, or any literal ~, silently lost the probe entirely - every command degrading to unresolvable with no explanation. ~ is a home expansion only at the start of a path; elsewhere it is an ordinary literal. The check is now split: $, backtick, *, ? and newline stay refused anywhere (substitution and globs, and the glob characters are illegal in Windows path components regardless), while ~ is refused only leading, tolerating one leading quote since the check runs before quote stripping. The prior tests only caught this on Windows because only Windows puts a ~ in tmpdir. Four new tests pin it on every platform via a fixture directory literally named RUNNER~1. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9de4d67118 |
fix(#3579): a pointer-less session inherits the repo active-workstream marker (#3616)
* test(3579): failing-first coverage for repo-marker inheritance A session that carries an identity but has never run 'workstream use' reads an absent session pointer, resolves null, and composes the flat .planning tree even when .planning/active-workstream names a live workstream. These tests fail on that and pin the invariants the fix must not break: a session with its own pointer is never repointed, and a session that merely lacked a pointer must never clear the shared marker on another session's behalf. * fix(3579): a pointer-less session inherits the repo active-workstream marker RED proven at 157cae26: the three inheritance tests failed while every isolation and negative control passed on base — the gap, and nothing else. pickActiveWorkstreamAdapter returned exactly ONE adapter: the session-scoped one whenever a session key existed, so the shared .planning/active-workstream marker was never consulted. getWorkstreamSessionKey resolves a key from ~13 env vars or the controlling TTY, so on any normal interactive terminal a key almost always exists — which is why a session that had never run 'workstream use' read an absent pointer, resolved null, and composed the FLAT planning tree even though the repo marker named a live workstream. Reads misreported; writes corrupted the superseded flat STATE. Silent, because the stale tree is well-formed. This was a genuine design fork, not an oversight: references/workstream-flag.md documented step 4 as a fallback 'when no session key exists', and the session isolation that buys is deliberate (#2850). The issue's Agent Brief left the choice open and said the reference doc should match whatever semantics ship. The maintainer ruled in chat for inheritance. Resolution now walks an ORDERED chain — session adapter first, shared second — and only a null from the session adapter falls through to the marker. Strictly additive: it can only turn a null into a name, never change a name that already resolves. The dangerous part is clear() ownership. resolveFromChain treats chain[0] as owned: only it is ever cleared, and only under selfHeal (getActiveWorkstream, never peek). An INHERITED marker is read-only — a stale value there resolves null and the file is left alone. Without that, one pointer-less session's read would delete the repo marker for every other session, which is a worse bug than the one being fixed. Covered by a test that asserts the marker still exists on disk after such a read. peekActiveWorkstream inherits but still mutates nothing (#2850 — the statusline draws on every render). references/workstream-flag.md's Resolution Priority is rewritten to match, keeping the session-isolation rationale and noting that inheritance does not weaken it: a session that owns a pointer is never repointed. Fixes #3579 * fix(3579): correct the guard diagnostics and lock the clear-semantics Three review passes; every finding fixed inline. MISSING ACCEPTANCE CRITERION (spec pass). The brief requires refusal diagnostics that distinguish 'marker present but the session lookup missed it' from 'no workstream set at all', and the two workstream-mode fail-safe guards were byte-for-byte untouched — still emitting a generic 'no active workstream is set' even when a marker exists and merely names a missing directory. Both guards (cmdPhaseComplete, cmdInitProgress) now branch on a new read-only diagnoseUnresolvedActiveWorkstream, which reuses the SAME resolvesToExistingWorkstream predicate resolveFromChain uses, so the diagnosis and the resolution cannot disagree. Two typed reasons added to ERROR_REASON; both arms still refuse — the fail-closed behavior is unchanged, only the message is now true. REAL TEST FAILURE, not a flake. The remote run failed 'clearing one session does not clear another session pointer'. That describe uses before() rather than beforeEach, so one tmpDir is shared and an earlier test writes active-workstream=beta into it; under inheritance the just-cleared session picks that marker up and resolves beta instead of null. The failure is a CORRECT consequence of Option A surfaced through an order-dependent fixture. The test now establishes its own marker state explicitly — its real intent (clearing A must not disturb B's pointer) is preserved and not weakened — and a new test pins the semantic deliberately: clearing a session pointer returns that session to INHERITING the marker, it does not force flat mode. Documented in references/workstream-flag.md, including how to actually get flat behavior. Also from review: partial activeWorkstreamAdapters injection no longer silently synthesizes a REAL filesystem adapter for the missing half (a latent test-isolation trap); the duplicated validate-then-existsSync logic is factored into one predicate; and the two try/finally test bodies are converted to t.after per CONTRIBUTING. New coverage: whitespace/empty shared marker; a session whose OWN pointer is stale while the marker names a different valid workstream (must self-heal to null, never inherit — the isolation guarantee at its sharpest); and both new diagnostic arms asserted on structured --json-errors output rather than prose. * fix(3579): read resolvability with the non-mutating peek, not the self-healing resolver Three of our own new tests failed on 7f5e706a. All three had ONE root cause, and none was fixed by relaxing an assertion. gsd-tools.cjs's bootstrap called the MUTATING getActiveWorkstream unconditionally on every invocation, purely to populate routing env. On an unresolvable pointer that self-healed — cleared it — BEFORE the dispatched command ran its own resolution. A second read in the same process then observed already-cleared state: - Isolation violation: a session whose own pointer was stale had it cleared by the bootstrap, so cmdWorkstreamGet's own resolution found a pointer-LESS session and inherited the shared marker ('beta' instead of null). Exactly the guarantee #2850 exists to protect, defeated across two calls rather than within one. - Guard diagnostics: the guards' own truthiness check also used the mutating resolver, so it cleared the invalid marker and the immediately-following read-only diagnosis found nothing and reported none_active instead of marker_unresolved. So a single invocation's answer depended on how many times it resolved. The bootstrap self-heal is PRE-EXISTING and was harmless while pointer-less meant flat — inheritance is what made it answer-changing, so this fix belongs here. Every call site that only CHECKS resolvability — the bootstrap, both fail-safe guards' truthiness check, and two informational init report fields — now uses the non-mutating peekActiveWorkstream. Self-heal is unchanged in active-workstream-store and still fires exactly once, at whichever site actually consumes the workstream. Verified by driving the real CLI against temp fixtures, since the suite cannot run locally: stale-own-pointer resolves null with the marker intact; both guard arms report marker_unresolved with missing_workstream_dir / invalid_name and the marker survives; no-marker still reports none_active; identity-less self-heal still deletes an invalid marker byte-identically to pre-#3579; and a session with a valid own pointer still wins. * chore(3579): backfill changeset PR number (#3616) * test(3579): kill the surviving mutants in the new resolution code CI's Stryker gate failed: active-workstream-store scored 79.45% against a break threshold of 80 — 259 killed, 67 survived, at 'Ran 1.00 tests per mutant on average'. The survivors cluster in the code this PR added (pickActiveWorkstreamAdapterChain, resolvesToExistingWorkstream, resolveFromChain, diagnoseUnresolvedActiveWorkstream): the CLI-level tests exercise those paths but do not DISCRIMINATE their branches, which is precisely what a surviving mutant means. Raised by strengthening assertions, never by touching the threshold. 21 unit tests added to the existing unit suite, each written to fail under a specific named mutant, using the module's injected adapter seams and createMemoryPointerAdapter so they stay hermetic under Stryker's per-mutant reruns: - chain shape with and without a session key, asserting length AND element identity (kills the if(false), the ': []' array mutant, and the block removal) - partial adapter injection, asserting the missing half is an inert memory adapter that never touches the filesystem (kills the three '??' -> '&&' mutants) - both arms of '!name || !validateWorkstreamName(name)' as SEPARATE tests — an absent name and a non-empty invalid one — which is what kills the '||' -> '&&' mutant - self-heal discrimination: getActiveWorkstream must clear an unresolvable owned pointer and peekActiveWorkstream must not, asserted on adapter state after each (kills if(selfHeal) -> if(true)) - fallback arm both ways: a fallback that resolves and one that does not - diagnoseUnresolvedActiveWorkstream asserted as a full object per case, with the reason strings compared exactly (kills present:true -> false and both StringLiteral mutants) One mutant is deliberately left: 'if (chain.length === 0)' -> 'if (false)'. The branch is structurally unreachable — the only chain source always returns a 1- or 2-element array literal — and resolveFromChain is not exported. Killing it would mean exporting an internal or deleting a defensive guard; neither is worth doing for a mutant, and the score clears 80 without it. Recorded here rather than left unexplained. Every new assertion was evaluated against the built module with real fixtures before committing, since the suite cannot run locally. --------- Co-authored-by: sim <sim@local> |
||
|
|
f56ffa86ab |
fix(#3581): derive init.progress's next_phase from roadmap order, not artifact presence (#3603)
* test(#3581): pin init.progress's frontier to roadmap order over stray artifacts Failing-first regression for #3581: a stray out-of-order phase directory (a phase-9 UAT evidence file while roadmap phase 8 was pending and unscaffolded) made init.progress report next_phase 09, skipping Phase 8 and disagreeing with roadmap.analyze. Rows pin the issue shape, the aligned-tree control, and the all-complete boundary. * fix(#3581): derive init.progress's next_phase from roadmap order, not artifact presence The frontier is re-derived from the sorted phase union after the disk and roadmap loops: the first not-yet-begun, not-roadmap-complete phase wins. Artifacts still feed status and completion per entry, but a stray out-of-order directory can no longer drag the frontier past a pending unscaffolded roadmap phase, and init.progress agrees with roadmap.analyze. * fix(#3581): frontier = first not-complete phase in roadmap order (resume semantics) Review-of-own-control refinement: a begun-but-unfinished phase (in_progress, executed, researched) is the frontier — the next thing to execute is to resume it — so the frontier predicate is simply 'not complete and not roadmap-complete', first in the sorted union. * fix(#3581): preserve the pinned pending-only frontier contract; repair the boundary fixture Review findings: the resume-semantics refinement broke the suite-pinned contract that an in-progress phase is currentPhase's lane, not nextPhase's (tests/init.test.cjs 'multiple phases with mixed statuses') — reverted to first pending-or-not_started; the control row now pins the pure ordering property (roadmap-only pending beats a later pending directory); the boundary fixture gains passing verification reports so disk status reaches complete under the #3168 disk-strict bar. * chore(#3581): add changeset fragment * chore(#3581): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
59e7a677fe | fix(#3511): scope every phase-directory scan to the phase it belongs to (#3535) | ||
|
|
3893d1ff69 |
fix(#3518): pin uat_path to the phase's own UAT artifact via the shared phase-pinned resolver (#3525)
* fix(#3518): pin uat_path to the phase's own UAT artifact via the shared phase-pinned resolver Both uat_path projectors in src/init.cts picked the phase's UAT file with a bare .find() over unsorted readdir order — no phase-membership check, no ordering — so a stray cross-phase 04-UAT.md in phase 03's directory could become phase 03's uat_path, filesystem-dependently (creation order on APFS, hash order on ext4/XFS): two machines on the same commit could emit different uat_path values for the same phase. Route both sites through a new resolveUatFile in src/verification.cts, the UAT counterpart of #3357/#3492's resolveVerificationFile, sharing the exact selection rule via one extracted core (resolvePhaseArtifactFile): the phase's own <token>-UAT.md always wins; otherwise the alphabetically-first dashed candidate (deterministic everywhere); a bare UAT.md only via allowBare when no dashed candidate exists. resolveVerificationFile now delegates to the same core — behavior byte-identical. Guarded by: two end-to-end repro tests in tests/init.test.cjs (plan-phase and phase-op, red on the pre-fix readdir pick), resolveUatFile contract anchors and a src/-wide call-site guard in tests/verification-status.test.cjs. * chore(#3518): add changeset fragment for PR #3525 --------- Co-authored-by: sim <sim@local> |
||
|
|
ddf852873c |
fix(#3357): one phase-pinned resolver for verification-report discovery (#3513)
A phase directory can hold more than one `*-VERIFICATION.md` — an ad-hoc `03-CORRECTION-VERIFICATION.md` worksheet beside the real `03-VERIFICATION.md`. Discovery took the alphabetically-first match, so the worksheet won and the phase could report `missing` while a passing report sat next to it. The issue named two copies. There were seven, in four grammars: two `.sort()[0]` sites in the verification module, three `.find()` over UNSORTED readdir order (phase status, and `verification_path` twice — filesystem-dependent, so two machines on one commit could disagree), and two in shell. All seven now route through one exported `resolveVerificationFile`; the shell copies via a new `verification resolve-file` verb rather than hand-rolling the rule an eighth time. Fixing five of seven would have been worse than fixing none: the verify-work workflow is a WRITER that stamps `status: passed` onto the file it picks, so canonical-aware readers plus an alphabetical writer means the human_needed→passed canonicalization silently no-ops forever while the worksheet gets stamped. That divergence did not exist on next. The first resolver was itself a regression — it preferred ANY canonically-shaped name over the phase's own report, so a stray cross-phase or sentinel-numbered file outranked it. The global-canonical preference was removed rather than narrowed; the rule is pinned to the phase token via `PHASE_NUMBER_TOKEN_SOURCE`, its existing owner. Fixed in passing: the transition workflow's awk guarded on `NR==1` instead of `FNR==1`, so across a multi-file glob it armed only on the first file — a leading worksheet with no frontmatter blocked transition even when the canonical report passed. Also removed a U+00AD soft hyphen introduced earlier on this branch. Five broader-grammar AGGREGATE scans are deliberately out of scope — a different defect class (phase-unscoped scanning), tracked as #3511. Closes #3357 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0d4d78550f |
fix(#3188): null absent planning-doc paths in init phase queries (#3430)
* fix(#3188): null absent planning-doc paths in init phase queries The init phase-op / plan-phase / execute-phase queries emitted a non-null absolute requirements_path / state_path / roadmap_path even when the named file did not exist — built with a bare path.join and no existence check, unlike the conditional sibling fields (patterns_path, context_path, ...) in the same payload. Consumers (e.g. ultraplan-phase.md:104 'requirements_path is not null') therefore read missing files. Each of the three reading sites now returns null when the file is absent and its absolute path when present. The project/milestone-bootstrap and doc-ingest emitters that use these paths as write-targets for not-yet-created files are intentionally unchanged. * chore(#3188): backfill changeset pr: 3430 --------- Co-authored-by: sim <sim@local> |
||
|
|
a3b72d9071 |
fix(#3171): init execute-phase emits display name, not directory slug (#3429)
* fix(#3171): init execute-phase emits display name, not directory slug When a phase directory already exists on disk, the disk-lookup path (searchPhaseInDir) derived phase_name from the directory-name remainder -- itself an already-slugified value (phase.add writes ${num}-${slug} dirs) -- so phase_name and phase_slug came out byte-identical. The execute-phase workflow forwards phase_name into 'state begin-phase --name', which wrote that raw slug into STATE.md's current_phase_name on every phase start. cmdInitExecutePhase now prefers the ROADMAP's curated display name ('### Phase N: <Name>') for phase_name, matching the no-disk fallback path that already did this correctly. phase_slug is unchanged (it feeds branch-name construction). The state.begin-phase authoritativeFm override (#2821/#2736) is untouched; the correction is in the value fed into --name. The milestone_name half of #3171 was subsumed by #3216 / PR #3226; this fixes the remaining current_phase_name half. * docs(changeset): backfill pr 3429 for #3171 --------- Co-authored-by: sim <sim@local> |
||
|
|
0624c5da6f |
chore(#3212): src/text-lines.cts is the sole owner of line-terminator handling — Phase 2 (#3420)
* test(#3413): failing-first suite for the line-terminator seam Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Tests only — src/text-lines.cts does not exist yet, so tests/text-lines.test.cjs fails with MODULE_NOT_FOUND at its require line, which is the intended RED. The frontmatter.test.cjs additions drive #3360 (confirmed-bug) fail-first: parseMustHavesBlock currently returns [] for every must_haves block on a CRLF-authored plan file, because \r is its own LineTerminator in ECMAScript and two /m-anchored \s* patterns can absorb it, inflating a captured indent by one character and tripping the "not nested under must_haves" guard. Verified locally against the current (unfixed) compiled module: both the direct repro and the silent-exit "blank line before must_haves:" variant return [] today. A parity property test (crlf vs lf must deep-equal for every block name) matches a pattern this maintainer has required repeatedly for prior CRLF fixes in this codebase (Cortex-recorded, verify_intent=held). The no-crlf-fragile-split.rule.test.cjs additions lock the eslint rule's future fix-hint text (pointing at splitLines()) and its self-reference non-violation (the seam's own correct \r?\n split must never flag itself). Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md * chore(#3413): src/text-lines.cts owns line-terminator handling Phase 2 of epic #3212 (ADR-3212 §3/§6/§7). Adds splitLines/normalizeEol/ detectEol/joinLines and migrates frontmatter.cts onto it. parseMustHavesBlock (#3360, confirmed-bug) returned [] for every must_haves block on a CRLF plan file. Root cause: \r is its own LineTerminator in ECMAScript, so under /m two \s*-anchored indentation lookups could match at the position INSIDE a \r\n pair and absorb the terminator, inflating the captured indent by one character and tripping the "not nested under must_haves" guard. Two silent exits, one with a diagnostic and one without (a blank line before must_haves: hits the silent path). Fixed by converting both lookups from a whole-string /m match to split-then-scan — splitLines first, then a per-line, non-/m match — the same structural pattern parseYamlRegion (30 lines away in the same file) already used safely. Nothing downstream of the two lookups changed; blockLines is now sliced from the already-split array instead of re-splitting a substring, but its contents are unchanged for LF input, and the per-line dash/kv parsing loop is untouched. A parity property test (CRLF and LF plans parse to identical must_haves for every block name) matches a pattern this maintainer has required repeatedly for prior CRLF fixes in this file's neighborhood (Cortex: 7 recorded decisions, verify_intent -> held). frontmatter.cts's other .split(/\r?\n/) call sites (parseYamlRegion, isFrontmatterShaped, sliceTopLevelFrontmatterSegments, spliceFrontmatter) are rerouted onto splitLines — a literal 1:1 substitution, zero behavior change, since splitLines IS that same regex plus a type guard. The 4 scripts/normalizeLineEndings copies (gen-registry, gen-loop-host- contract, gen-capability-registry, gen-context-index) are deleted and rerouted onto normalizeEol, which strips a bare unpaired \r exactly like the deleted copies did (not just \r\n pairs) -- verified against each script's own --check mode against its real generated output. local/no-crlf-fragile-split widens from tests/ to src/**/*.cts, with its fix-hint message now naming splitLines() instead of the raw regex -- the prohibition finally has a primitive to point at. Detection logic unchanged in this phase (deliberate scope limit, see design doc Known limits: the rule doesn't yet recognize safeReadFile/platformReadSync as a content source, and has no detector for the \s-adjacent-to-anchor shape that is #3360's actual mechanism -- the CLASS is converged by the direct fix + regression test regardless). joinLines/detectEol are NOT wired into frontmatter.cts's own write path (cmdFrontmatterSet/Merge -> platformWriteSync) -- verified that platformWriteSync already, unconditionally converts CRLF->LF on every .md write today as a pre-existing policy owned by a different module, and ADR-3212's backward-compatibility clause rules out a file-format change in any phase. Stated explicitly in Known limits rather than left for a reader to discover. Six-gate ripple: .gitignore, eslint.config.mjs (src/**/*.cts block), docs/INVENTORY.md + INVENTORY-MANIFEST.json (regenerated), CONTEXT.md glossary (Text Lines Module, mirroring Phase 1's Pattern Module entry). Design: .gsd/phase/chore-3413-text-lines-seam/40-design.md Test matrix: .gsd/phase/chore-3413-text-lines-seam/50-test-matrix.md * fix(#3413): fix 13 pre-existing CRLF-fragile splits the widened rule found Widening local/no-crlf-fragile-split from tests/ to src/**/*.cts (the previous commit) immediately surfaced 13 real, pre-existing violations across 10 files -- undetected until now because the rule never scanned src/. This is the exact defect class ADR-3212 exists to close, playing out again one phase after Phase 1 hit the same shape ("the new lint rule -- once live -- found 27 more"). Per CLAUDE.md's no-defer rule, fixed inline rather than deferred or suppressed; there is no established suppression convention for this rule in src/ and inventing one now would undermine the point of widening it. audit.cts, broken-windows.cts, core-utils.cts, init.cts, milestone.cts, phase.cts (x3), profile-output.cts, roadmap.cts (x2): bare-\n splits or regex character classes widened to \r?\n / [^\r\n], each following the same pattern already established migrating frontmatter.cts. phase-estimation.cts: `\r?(?:\n|$)` restructured to `(?:\r?\n|\r?$)` -- already semantically CRLF-safe, but the rule's lexical scanner doesn't recognize \r? guarding a group (only \r? immediately before a literal \n). Verified the two forms are equivalent across all four EOL/EOF cases before restructuring, not assumed. roadmap-upgrade.cts needed two coupled sites, not the one flagged line: computeMigrationPlan and applyMigration must agree on line representation for the lines[edit.lineIndex] === edit.from equality check to hold, and the write-back needed joinLines + detectEol -- a plain lines.join('\n') was silently flattening a CRLF ROADMAP.md to LF wholesale on every migration. This is the first real production consumer of joinLines/detectEol in this epic (frontmatter.cts's own write path doesn't use them -- see the previous commit's Known limits). Fixing the 13 flagged sites surfaced 4 more adjacent same-shape sites the rule doesn't track (.search() and new RegExp(dynamicString) aren't in its tracked call/construction set). Investigated each empirically -- hand-tracing this exact bug class already produced one wrong conclusion earlier in this phase (a detectEol design-doc arithmetic error), so these were verified with real CRLF fixtures rather than reasoned about on paper: - audit.cts (scanTodos): REAL bug, fixed. `bodyMatch.trim().split ('\n')[0]` leaked a trailing \r into a user-visible todo summary on CRLF input -- .trim() only strips the string's outer edges, not a \r sitting mid-string before the first bare \n. Now splitLines(...) [0]. - phase.cts (cmdPhaseInsert, bullet-style branch): REAL bug, fixed. [^\n]* in targetBulletPattern swallowed a line's trailing \r on CRLF input, shifting the computed insert position to land INSIDE the \r\n pair; combined with a hardcoded '\n' bullet separator, a CRLF ROADMAP.md ended up with a mixed CRLF/LF result after an insert. Fixed with two coupled changes (either alone still corrupts, verified both ways): [^\r\n]* in the pattern, and the new bullet's leading terminator now comes from detectEol(rawContent). - roadmap.cts (cmdRoadmapAnnotateDependencies phase-boundary scan): investigated, genuinely safe, left untouched. The .search(/\n#{2,4} .../) boundary-finder and the [^\n]*-based heading match were empirically verified on a 3-phase CRLF fixture -- the only stray \r ends up at the tail of an intermediate phaseSection string that is only ever used for .test()-based idempotency checks, never for an exact-match comparison or written back to disk. No corruption on round-trip. Every fix re-verified: npm run build:lib clean, npx eslint 'src/**/*.cts' --no-cache reports 0 problems (was 13), and each fixed function's existing LF-input tests were spot-checked unchanged. * fix(#3413): apply orthogonal review findings Two isolated review engines (correctness + security) ran against the full diff and found three majors, one real security issue, and several disclosure-worthy minors. All fixed or explicitly disclosed with evidence; nothing deferred. MAJOR — detectEol's tie-break contradicted its own documented contract. Code returned '\n' on a 1:1 crlf/bare-LF tie; every doc (design doc, CONTEXT.md, the function's own comment) says ties resolve to '\r\n'. The existing test masked this by reusing the same tie fixture the buggy code happened to satisfy, rather than a genuine LF-majority case. Root cause: an Edit attempted earlier in this phase to fix this exact arithmetic error was blocked by the tier guard, and a later dispatch was incorrectly told it had already landed. Fixed: condition is now crlfCount >= bareLfCount; the test fixture corrected to a genuine 2:1 majority, with a new explicit tie-case test. MAJOR — phase.cts's cmdPhaseInsert built an EOL-aware bulletEntry via detectEol(rawContent), justified by a comment claiming a hardcoded '\n' corrupts a CRLF ROADMAP.md. False: this write goes through platformWriteSync, whose normalizeContent/_normalizeMd unconditionally converts CRLF->LF for any .md target — the templating was inert dead code, erased before the file is ever written. Reverted to hardcoded '\n', comment corrected to state the true reasoning. The separate [^\n]* -> [^\r\n]* widening one function up (a real splice-position fix, independent of final EOL) was kept. MAJOR — roadmap-upgrade.cts's stated rationale for switching onto splitLines/joinLines was wrong (both functions always agreed on line representation, before and after — the claimed equality-check risk never existed), and the change it justified introduced a real regression: forcing every line onto one dominant terminator silently rewrites untouched lines' EOL on a mixed-CRLF/LF ROADMAP.md. This write path uses raw fs.writeFileSync, not platformWriteSync, so unlike the phase.cts case above the regression is genuinely live. Fixing this took two attempts. The first attempt (revert to split('\n')/join('\n') plus a suppression comment) was correctly blocked by an agent that discovered local/no-crlf-fragile-split is a PROTECTED_RULES entry in tests/portability-rule-disable-ban.test.cjs — a hard, out-of-band, ADR-1703-governed guardrail banning any eslint-disable of this rule anywhere in src/**/*.cts. That agent also detected and correctly disregarded an injected instruction that appeared in tool output during a git operation, per this session's untrusted-content policy. The actual fix: computeMigrationPlan reverted to roadmapContent.split('\n') (confirmed lint-clean — the rule's data-flow tracking only follows a variable's initializer, and this one is declared empty then reassigned in a try block). applyMigration's write-back now splices edits against the ORIGINAL content string via indexOf('\n', pos) boundary-walking instead of a full split/rejoin, so every untouched character — including every line's own terminator — is copied byte-for-byte. A capture-group split (/(\r\n|\n)/, preserving terminators inline) was tried first and empirically confirmed to still trip the rule before this approach was chosen instead. MINOR (security) — roadmap.cts's cmdRoadmapAnnotateDependencies used the STRING form of String#replace, so $&, $`, $', $1-$9 inside must_haves.truths content (author-controlled) were interpreted as replacement directives, splicing unrelated ROADMAP.md text into the result. Fixed with the function-replacement form, which is never pattern-interpreted. Verified before/after with the reviewer's exact repro. Also disclosed rather than silently left: test matrix row 31 (four planned CRLF-materialized regression tests) was never implemented as separate files — corrected to record the actual verification (a manual --check run plus incidental existing coverage via each script's normalizeLineEndings: normalizeEol alias). parseMustHavesBlock's LF behavior was claimed byte-for-byte unchanged but the old yaml.indexOf(blockMatch[0]) substring search could match an unrelated earlier occurrence of the header text (e.g. inside a quoted value) — the split-then-scan fix incidentally also closes this, a strict improvement now recorded in the design doc rather than left implicit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3413): checkpoint 2 red — missing eslint ignore entry, RuleTester config error Checkpoint 2 came back red with 5 failures on the reviewed sha, both gaps genuinely undetectable by any local gate. eslint.config.mjs was missing the 'gsd-core/bin/lib/text-lines.cjs' ignores-list entry (ADR-457: generated .cjs artifacts are excluded from direct type-aware linting). Phase 1's sibling entry (pattern.cjs) sits two lines above it and was the exact precedent read while researching the six-gate ripple for this module -- missed anyway. Caught by tests/repo-invariants.test.cjs's bin/lib coverage-tracking test, which only runs on the remote suite. tests/no-crlf-fragile-split.rule.test.cjs's row-32 case specified both `messageId` and `message` on the same RuleTester error assertion -- ESLint's RuleTester rejects that combination outright. This existed since the test was first authored and was never caught locally: `npx eslint` only lints the file's syntax, it does not execute RuleTester, and local `node --test` is hard-blocked in this repo -- the assertion had never actually RUN before this checkpoint. It was even present in checkpoint 1's failure list, listed there as one of the "expected RED" tests; I matched it against my expected-failures list by test NAME only and never inspected the actual failure detail closely enough to notice it was failing for the wrong reason (a RuleTester config error, not the intended message-text mismatch). Fixed by keeping `message` (the exact-text assertion the test exists to make) and dropping `messageId`. Verified the crlfFragileSplit message string in eslint-rules/no-crlf-fragile-split.cjs matches this assertion character-for-character, and swept every other invalid case in the file for the same double-specification bug (none found -- all pre-existing cases use messageId alone). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#3413): add Fixed changeset for the #3360 CRLF parsing fix The sole user-visible effect of this phase. No breaking-change label or Changed fragment needed — ADR-3212's Backward Compatibility section names the Node floor (Phase 1, already shipped) as the epic's only breaking change; Phase 2 has none. * chore(#3413): backfill changeset pr number to 3420 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
dc3c81e93d |
chore(#3212): src/pattern.cts is the sole owner of runtime-value regex construction — Phase 1 (#3416)
* test(#3412): failing-first suite for the pattern-construction seam Phase 1 of epic #3212 (ADR-3212 §1/§2/§7). Tests only — src/pattern.cts and eslint-rules/no-adhoc-regex-escape.cjs do not exist yet, so both suites fail with MODULE_NOT_FOUND, which is the intended RED. Locks the measured behavior rather than the assumed behavior: RegExp.escape hex-escapes the leading character of nearly every string ("abc" -> "\x61bc"), so the suite asserts match-equivalence against an inlined historical oracle (the implementation being deleted) rather than byte-equivalence of pattern text — 200 seeded fast-check runs plus a fixed corpus, 0 mismatches. Also locks the latent character-class range bug this phase fixes as a side effect: a hyphen-bearing value interpolated into [...] currently forms a real range and matches an unintended character; post-migration it must not. * chore(#3412): src/pattern.cts owns runtime-value regex construction Phase 1 of epic #3212 (ADR-3212 §1/§2/§6/§7). Adds the pattern seam delegating to the built-in RegExp.escape, deletes every hand-rolled copy, and raises the Node floor to the Active LTS line. The census was low, three times over. ADR-3212 counted 10 copies; a graph query found 12; the new lint rule — once live — found 27 more. The difference is that the census counted named helper FUNCTIONS while the rule counts the escape SHAPE, so inline .replace(<class>, '\$&') copies were never in scope. ADR §1's actual requirement is that no module outside the seam escapes a value for regex use, so all of them are, and CLAUDE.md's no-defer rule makes them this change's work. Fourth consecutive epic here whose copy count was low — the argument for ADR-3180 Amendment 3's "state N found by the guard" rule. Also corrected mid-implementation: the survey reported phase-id.cts's escapeRegex had 0 external importers. It had 8 production importers, making its removal a public-surface change to an ADR-2121-owned module and requiring an update to that ADR's locked-surface test. Blast radius revised Medium-High -> High. RegExp.escape is match-equivalent but NOT text-equivalent: it hex-escapes the leading char of nearly every string ("abc" -> "\x61bc"). Equivalence is proven by a seeded fast-check property test against the deleted implementation as oracle. It also fixes a latent bug: a hyphen-bearing value interpolated into a character class previously formed a real range and matched an unintended character. Node floor 22 -> 24 (RegExp.escape is Node 24+), across engines, .nvmrc, package-lock, 9 CI matrix entries, and 5 docs. The aggregate `required-tests` context is unchanged and no job was added or removed, so branch protection cannot be orphaned by the dropped lanes. Enforced by eslint-rules/no-adhoc-regex-escape.cjs (shape-matched, with structural provenance for reviewed pattern-fragment constants rather than a name heuristic) plus a whole-tree companion guard covering the directories ESLint's globs miss. * fix(#3412): close the _SOURCE guard evasion, correct two false claims Three findings from the orthogonal review pass, all fixed. 1. The ESLint rule's `_SOURCE` provenance fallback was pure identifier- name matching with no binding check, so `new RegExp(userInput_SOURCE)` — a function parameter — sailed past the guard. That is the same rename-evasion class issue #3410 documents, reopened by the very fallback meant to complement the structural check. Now bound to the identifier's actual binding kind: import, require-derived const, or module-scope const; parameters, `let`/`var`, and unresolvable bindings fail closed. Four RuleTester cases cover the evasion and prove the legitimate cross-module case still passes. 2. src/pattern.cts's own header carried the stale pre-correction counts (12 copies / 17 call sites) while CONTEXT.md and the design doc carried the corrected ones (~39 / ~44) — a self-contradiction inside the PR whose entire purpose is deleting divergent copies. Rewritten, preserving the durable lesson: a named-function census cannot see inline copies; only a shape-matching guard can. 3. The claim that all deleted copies threw TypeError on non-string was false. phase-id.cts's copy — the one with 8 external importers — did String(value).replace(...) and never threw. The seam's locked signature does not coerce, so this is a real, now-disclosed behavior change rather than the pure preservation the tests asserted. Audited all 32 invocations across the 8 importers and 6 in-file callers: every one is safe by construction (upstream truthy guard or a string-producing derivation), verified by runtime probe against the compiled modules rather than by TS compilation, which cannot see a runtime undefined. Corrected the false claim in both the test comment and the design doc, and added it to Known limits. * docs(#3412): add Changed changeset for the Node 24 floor The only user-visible break in this phase. The escape-behavior change is internal and match-equivalent, so it carries no user-facing note. * fix(#3412): resolve the seam's require graph in script fixtures and packaging Checkpoint 2 came back red with 90 failures on the node24 lane. Three distinct defects, all introduced by routing scripts/ through the new pattern seam, none reproducible by any local gate: 1. ~82 failures — tests/adr-index-gate.test.cjs and tests/removed-but-needed-lint.test.cjs copy a scripts/*.cjs into an mkdtemp fixture and spawn it there (necessary: those scripts resolve their scan root from __dirname/.., so running the real script would scan the real repo). Each harness hand-listed the dependencies to copy alongside. Adding require('../gsd-core/bin/lib/pattern.cjs') to gen-adr-index.cjs made both lists silently incomplete -> MODULE_NOT_FOUND, plus 17 downstream 'did not emit parseable JSON' failures from the same crash. Fixed as a class, not an instance: new tests/helpers/copy-script- fixture.cjs walks a script's transitive static relative-require graph and copies it, so dependencies are derived and never re-declared. It throws (naming the unbuilt artifact) instead of letting the child die with a bare MODULE_NOT_FOUND. Verified for all four seam-consuming scripts: gen-adr-index, lint-removed-but-needed, gen-loop-host- contract, sync-runtime-launcher. 2. 2 failures — scripts/ ships wholesale but eslint-rules/ does not, so the new scripts/lint-no-adhoc-regex-escape.cjs would be MODULE_NOT_FOUND in a published install (#2858 guard). Excluded from the tarball, matching the existing precedent for gen-emitted- baseline.cjs, which is excluded for the identical reason, and locked with a test modeled on that one. Confirmed against a real npm pack: 890 files, 0 from eslint-rules/, and gsd-core/bin/lib/pattern.cjs present (so the other four scripts' requires are legitimate). 3. 6 failures — tests/phase-id.test.cjs asserted the literal escaped source text ('0*29', 'PROJ-42'). RegExp.escape is match-equivalent to the retired hand-rolled escaper but NOT text-equivalent: it hex- escapes the leading character and all hyphens ('0*\x329', '\x50ROJ\x2d42'). Verified NOT a behavior change — 576 match decisions across all three real interpolation prefixes, zero divergence. Those tests now compile each source into the same heading regex src/roadmap.cts's searchPhaseInContent builds and assert what matches and what does not, including the 'i'-flag canonicalization the hex escape has to preserve. Re-pinning the new literals would have rebuilt the same brittleness one layer down. Adds a test for the property the escape exists for: a dot in '1.2' must not act as a wildcard. Also shares one definition of 'a require' between the packaging guard and the fixture copier, so the two cannot disagree about what they scan. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): refuse to copy a fixture dependency outside the fixture root copyScriptWithDeps resolved each relative require and joined the repo-relative result onto fixtureRoot. A require resolving OUTSIDE the repo yields a '../'-prefixed relative path, so path.join climbed out of the fixture and wrote into the surrounding temp dir (verified: repoRoot=/repo + depAbs=/etc/passwd wrote /tmp/etc/passwd). No script in the tree does this today, so this closes an available escape rather than an active one. Refuses via the existing unresolved- require path so the failure names the offending specifier. Covered by a negative proof that the guard fires and that nothing lands outside the fixture. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3412): parse requires instead of pattern-matching them; restore the foreign-prefix contract Applies all findings from the second orthogonal review round, re-run because real code changed after round 1. HIGH (security) — extractRequires stripped BLOCK comments before LINE comments, so a '//' comment containing '/*' opened a phantom block comment, and a '//' inside a string literal truncated the line. Both hid real requires: 'const u="http://x"; require("./real.cjs")' returned [], and four real requires in gsd-core/bin/gsd-tools.cjs were invisible. Replaced with a real AST parse via espree. This is ADR-3212's own Decision 4 — tokenizer-first for stateful grammars — applied to the case it describes; comment/string/regex nesting is exactly such a grammar, which is why the regex version was wrong. The function was moved byte-identical out of the #2858 packaging guard, so the bug PRE-DATES this branch and has been a live blind spot there: a shipped script could have required an unshipped path undetected. Fixing it makes that guard strictly stronger than on next. espree is promoted from a transitive eslint dependency to an explicit devDependency rather than relying on hoisting. The script parse attempt sets ecmaFeatures.globalReturn because Node wraps CommonJS bodies in a function, making a top-level return legal — scripts/check-coverage-gate .cjs relies on it, and without the flag the guard throws on a file it is supposed to scan. Verified 0 unparseable across all 324 .cjs/.js under scripts/, bin/, and gsd-core/bin/, and 0 new violations against a real npm pack, so the exact extractor does not newly fail the guard. MEDIUM (security) — the repo-containment check guarded dependencies but not the entry path. One escapesContainment predicate now guards both. LOW (security) — containment was lexical while fs follows symlinks, and a directory symlink could mint a fresh dedupe key per level. realpath now resolves both repoRoot and each dependency before the decision, and the realpath-derived path is the dedupe key. Destination layout still uses the original repo-relative path, so copied trees are unchanged. MAJOR (standards) — the round-1 behavioral rewrite of phase-id tests lost the foreign-prefix contract: every assertion was satisfied by an impl returning [A-Z]+\x2d42, i.e. ANY project code — the exact #3599 bug class the exact-source prevents. The literal assertions it replaced were catching this. Now asserts the compiled regex REJECTS a different prefix with the same number. MAJOR (standards) — the test hand-duplicated production's heading regex with no parity guard (CLAUDE.md's 'Generative Fix Divergence'). Removed the parallel surface instead of policing it: src/roadmap.cts exports buildPhaseHeadingRegex, searchPhaseInContent calls it, the test imports it. Byte-identical .source and .flags verified for both escaped forms. MINOR — '..foo' no longer false-flagged as an escape; the inverted spurious-vs-missing doc claim corrected; the dead allow-test-rule header removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3412): backfill changeset pr number to 3416 * fix(#3412): make the escape guard's own regex linear, reword an injection-scan collision Two CI failures on PR #3416, both in code this branch added. CodeQL js/redos (high) — REPLACE_CALL_RE's outer alternation let a bracket run be consumed EITHER by the character-class branch OR one character at a time by the trailing catch-all, so a failing match explored both parses of every pair. Measured on the real regex: n=26 -> 204ms, n=28 -> 791ms, n=30 -> 3475ms, a clean 2^n. This script scans repo source, so a file with a long bracket run after '.replace(/' would hang CI outright — a guard against undisciplined pattern construction was itself the worst pattern in the diff. Fixed the way ADR-3212 already prescribes: the catch-all branch now excludes '[' and ']' so a bracket can only be consumed by the class branch (this is what makes it linear), and every quantifier is bounded (the locked bounded-quantifiers decision) as a second line of defense. Now 0ms at n=2000. Disclosed coverage tradeoff, recorded at the constant: a regex literal with a BARE unescaped ']' outside a class is no longer matched by this backstop. No census shape has that form, and the AST rule remains the primary detector. Verified the guard did not go blind doing it: a real census-shape violation is still reported, and an allow-adhoc-regex-escape suppression comment is still honored. Regression test drives the exported findViolations on a 2000-repetition adversarial input and asserts the RESULT. It makes no wall-clock assertion — elapsed-time tests are forbidden — so a regression surfaces as a harness timeout, which is the correct signal. Prompt injection scan — 'must not act as a regex wildcard' in a test comment matched the scanner's jailbreak pattern act\s+as\s+(a|an|if| my). Reworded to 'behave as'. Deliberately NOT allowlisted: silencing a whole test file over one phrase would blunt the scanner permanently, and the comment has nothing to do with injection. Neither failure was reachable from the remote runner — CodeQL and the injection scan are not in that matrix, so the sha it passed was green and still wrong. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
77374cfc32 |
fix(#2528): resolve digit-leading phase directories by bare number (#2559)
* fix(#2528): resolve digit-slug phase dirs by bare number — tokenizer rewind, shared bare-integer fallback, resolution-path parity gate
extractPhaseToken welded 2-digit slug words onto the phase token (phase 10
named "24/7 Autonomy" -> dir 10-24-7 -> token 10-24), making digit-prefixed
phase names unresolvable by bare number across every phase verb.
- phase-id: continuation segments must be the PURE 2-digit zero-padded form
the write side emits; a 1-digit terminator rewinds the absorbed run
(10-24-7 -> 10) while >=2-digit terminators keep the locked #2232
round-trip (14-06-2026-photos -> 14-06).
- phase-id: new matchPhaseDirs owner — primary exact-token match plus a
bare-integer leading-digit-run fallback for shapes the tokenizer cannot
rewind (05-80-20-cleanup); collisions stay #2237-loud.
- locator/find-phase/phase-plan-index all delegate selection to the owner;
plan-index gains the previously missing multi-match guard.
- tests: #2528 unit + fast-check metamorphic blocks; new 9-scenario
resolution-path parity gate across all three paths.
Fixes #2528
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore(#2528): add changeset for PR #2559
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(#2528): align validation token grammar
* fix: address phase token review
* fix: restore phase grammar parity for numeric slugs
* docs: document digit-leading phase resolution
* docs: clarify ambiguous phase resolution behavior
* docs: register canonical phase directory selectors
* fix: align prefixed deep phase token parsing
* fix(#2528): route the fourth resolution site through matchPhaseDirs
Review BLOCKER. smart-entry.cts::detectVerifyFailed resolved the current
phase's directory with its own `.find(phaseTokenMatches)` and never
reached the shared selection, so the bare-integer-fallback family the
issue names — `05-80-20-cleanup`, `30-12-factor-refactor` — resolved
nowhere. The miss is silent by construction: an unresolved phase reports
"not failed", which is byte-identical to a healthy one, so a failed
verification simply never surfaced in /gsd or /gsd:progress.
`entries` is already sorted and matchPhaseDirs filters without
reordering, so matches[0] reproduces the previous selection exactly
wherever the old code resolved at all.
Wiring it into phase-resolution-parity.test.cjs as a fourth path then
exposed a second, older defect in the same function: phaseTokenFromDirName
shape-probed the UNSTRIPPED token, so a project-code-prefixed directory
(`MEM-05-…`, tokenizing to `MEM-05-80-20`) failed the leading-digit test
and was dropped before any resolution ran — every phase in a
project-coded plan was invisible to this check. The probe now runs on the
stripped token; the returned value is unchanged, so the comparePhaseNum
sort is untouched.
Path 4 has no JSON surface to compare, so the gate observes selection
indirectly: plant the failing artifact in exactly one directory and a
passing one everywhere else, then read the boolean. Reverting either fix
turns 5 of the 10 corpus scenarios red.
* refactor(#2528): collapse the duplicated extractCanonicalPlanId
Review MAJOR. The function existed as two independent, byte-identical
copies — src/core-utils.cts and src/phase.cts — and this PR had to patch
BOTH with the same single-digit-slug rewind rule. That is the generative
fix divergence CLAUDE.md names, and only the core-utils copy was under
test, so a future one-sided patch would have silently split plan-id
canonicalization between the plan listing and everything else.
Removed rather than parity-tested: core-utils was already the leaf owner
and already exported it, and phase.cts already imported that module, so
there is no second surface left for a parity test to police.
* test(#2528): pin matchPhaseDirs at the digit-width boundaries
Review MAJOR. The bare-integer fallback's correctness rests entirely on
capturing each directory's whole leading digit run before the zero-strip
compare; a regex that stopped short would turn every query into a prefix
match, and "1" would claim 10, 100, and 12 alike. The existing coverage
was example-based and never touched that boundary.
Adds the explicit 9/10 and 1/10/100 cases — including the forms where
only the wider directories exist, so an exact-width neighbour cannot
satisfy the assertion — plus a fast-check property over arbitrary
distinct leading runs. The property is stated as an invariant on the
result (every returned directory's leading run IS the query) rather than
an expected list, so it covers primary and fallback matches alike and
cannot be satisfied by reimplementing the selection in the test.
Both fail when the fallback regex is degraded to a prefix match.
* fix(#2528): route the remaining eight consumers through matchPhaseDirs
phaseTokenMatches had eight consumers left that each rebuilt the directory
selection around it by hand: phases-list, next-decimal, phase-remove, the
W021 milestone-consistency check, schema-drift, the init-manager overview,
milestone-complete's disk check, and roadmap analyze. Every one of them
reproduced the reported symptom in full after the tokenizer was fixed.
None of them derives a displayed phase number from the matched directory,
so none needs phaseNumberForMatch; the change at each site is the
selection and nothing else. matchPhaseDirs filters without reordering, so
matches[0] reproduces the prior .find() choice wherever the old code
resolved at all.
phaseTokenMatches now has no call sites outside phase-id.cts. It stays
exported as the primitive matchPhaseDirs is built from and as a pinned
canonical surface, but no consumer reaches past the owner to it.
* test(#2528): extend the parity gate to the migrated consumers
Each of the eight is observed through the surface a user sees, not
through the matcher, with a no-directory control so the assertions cannot
be satisfied by a consumer that resolves unconditionally. init-manager
and roadmap-analyze are additionally asserted to agree with each other.
* refactor(#2528): own the case-flexible phase grammar and the leading-digit-run fragment
validate.cts derived its case-flexible regex sources by running
`replaceAll('A-Z', 'A-Za-z')` over two constants exported by phase-id.cts.
That passes lint-phase-id-drift.cjs — there is no literal copy of the
grammar — but it depends on the owner rendering that exact substring. The
day phase-id.cts expresses the same class any other way the replaceAll
silently no-ops and validate.cts narrows to uppercase-only. The failure
mode is a NON-match, so nothing throws and no uppercase-only fixture
notices. Both variants are now derived once, beside the sources they
widen, and imported.
Also names the leading digit run the bare-integer fallback selects on.
It was spelled `/^(\d+)(?:-|$)/` where the fallback filters and `/^\d+/`
where phaseNumberForMatch reads the number back off the winner; selecting
on one run and displaying another would resolve a directory and then label
it with a number that never matched it.
* fix(#2528): refuse to remove a phase when two directories claim its number
cmdPhaseRemove was the only migrated site taking matches[0] with no
multi-match guard. Every sibling resolution path returns ambiguous_matches
and refuses to choose; this one is the DESTRUCTIVE path, so choosing
silently is strictly worse than anywhere else. With 05-80-20-a and
05-90-till-late on disk, `phase remove 5 --force` deleted one of them and
renumbered every phase after it — where the base resolved nothing, deleted
nothing, and the corpus in tests/phase-resolution-parity.test.cjs already
declared that exact input ambiguous.
The refusal is emitted before any file is touched and carries both
candidates. CONSUMER_SCENARIOS could not express the case — every row is
binary, resolving to one directory or to none — so the gate gains a
dedicated ambiguous test. It asserts on the filesystem, not only on the
reported directory_deleted: a null printed after an rmSync would satisfy
every other check.
* fix(#2528): pair digit-leading phase directories with their roadmap phase in validate health
W006/W007 are the ninth site of this bug class and the one a
`phaseTokenMatches` grep could never surface: they resolve roadmap↔disk by
intersecting TOKEN SETS, which is a dir→token labelling rather than the
query→dir selection matchPhaseDirs owns. On the canonical fixture the
label is wrong in both directions at once, so `validate health` reported
"Phase 5 in ROADMAP.md but no directory on disk" AND "Phase 05-80-20
exists on disk but not in ROADMAP.md" for the same directory.
collectDiskPhases now keeps the directory names behind each token, so
W006 can ask the canonical matcher whether a roadmap phase resolves to a
real directory, and W007 — which iterates directories and therefore has no
query to resolve — gets the inverse mapping it never had: a directory is
claimed when some roadmap phase resolves to it.
Both checks are additive: the token intersection still decides every shape
it already decided, and the resolution can only REMOVE a warning. The
regression test carries controls in the other direction — a roadmap phase
with no directory must still raise W006, an unclaimed directory must still
raise W007 — so it cannot be satisfied by a check that stopped reporting.
* docs(#2528): state and pin the directory-side scope of the bare-integer fallback
The matchPhaseDirs docblock claimed deep-decomposition lookups were
untouched. That is true of the QUERY side only — no non-bare query enters
the fallback — but the DIRECTORY side is what changed classification: a
bare `5` now reaches a lone `05-01-auth` and resolves it (phase_number
"05", phase_name "01-auth") where the base found nothing.
The widening is irreducible from directory names alone. `05-01-auth`
(sub-phase 5.1) and `30-12-factor-refactor` (phase 30 named "12-Factor
Refactor") are the same `NN-NN-<slug>` shape, and the discriminator that
would separate them — "is the second segment a valid decimal sub-phase" —
accepts `5.1` and `30.12` equally. Any rule strong enough to exclude the
first excludes the second, which is the defect #2528 exists to fix. So the
tie is broken in favour of resolving, the docblock now says so, and the
consequence is bounded where it matters: two such directories are two
matches, and every caller (including phase remove) refuses to choose.
Pins both directions, since nothing observed the directory side before.
* fix(#2528): count surviving phases by identity in phase remove's STATE resync
#2640 landed on `next` while this branch was open. Its STATE.md phase-count
resync re-derives "which directory was removed" from the query with
`phaseTokenMatches`, which is the tenth site of this issue's defect: the
bare-integer fallback resolves `05-80-20-cleanup` for query `5`, but the
token predicate does not, so the just-deleted directory is counted as still
present and the written `Total Phases` is one too high.
`targetDir` already IS the directory that was removed, and the block is gated
on it being non-null, so identity answers the question exactly — which is also
what the comment above the filter already claimed it did. This keeps
`phaseTokenMatches` out of `phase.cts` rather than re-importing it to satisfy
one call site: the module's public surface should not grow for a question that
does not need re-derivation.
Pinned in the parity gate with a control on a directory the tokenizer reads
correctly, so the assertion is about the digit-leading shape and not about the
counting rule changing for everything.
* fix(#2528): let the resolution layer own the digit-leading slug family alone
The tokenizer rewind this fix carried — pop the last absorbed continuation when
the segment that stopped the scan is a bare single digit — reads
"10-24-7-autonomy" (phase 10 named "24/7 Autonomy") correctly and silently
re-reads "10-24-7-zip" (sub-phase 10.24 named "7-Zip Integration") from "10-24"
to "10". The two names are string-identical in shape, so no local signal
separates them; the rule traded the reported ambiguity for the symmetric one a
level down, on a 15-caller chokepoint whose output also feeds query-less
derivations (STATE.md phase counts, W007, the #2562 key surface). A well-formed
sub-phase directory became unresolvable by its own id — the very symptom #2528
was filed about.
It also bought nothing. The bare-integer fallback in matchPhaseDirs already
resolves "10-24-7-autonomy" for query "10" whatever the token is: no primary
match, bare query, leading digit run "10". The reported case was covered twice,
by two rules, and the two disagreed about the case nobody reported.
So the rewind is removed rather than narrowed, in the tokenizer and in the five
surfaces kept in lockstep with it (BRACKET_PHASE_TOKEN_SOURCE,
PHASE_TOKEN_FROM_DIR_RE, canonicalPlanStem's pair grammar and its collision
branch, roadmap-parser's numericRe, extractCanonicalPlanId), together with the
SINGLE_DIGIT_RUN_SEGMENT_SOURCE owner constant they shared. Disambiguation now
lives only where a QUERY exists to disambiguate against, which is the same
bounded mechanism the "05-80-20-cleanup" shape already used.
Measured, not argued: over 29800 generated directory names, extractPhaseToken is
byte-identical to `next` on every input except the lowercase-continuation class
("01-20a", "05-80-20-25abc") — a rule about the segment itself, not a guess about
its neighbour.
Both readings now stay reachable by their own ids:
matchPhaseDirs(['10-24-7-autonomy'], '10') -> the dir (fallback)
matchPhaseDirs(['10-24-7-zip'], '10') -> the dir (fallback)
matchPhaseDirs(['10-24-7-zip'], '10-24') -> the dir (primary)
* test(#2528): pin the one-continuation boundary the rewind had no coverage for
The regressing shape was invisible to the suite by construction, not by luck:
the deep-rewind property built its cases from `continuationArb` with
`minLength: 2`, so it never exercised the single-continuation case — exactly one
genuine sub-phase level before a digit-leading slug — and every hand-written
fixture used the ambiguous shape only where "phase-plus-slug" was the intended
reading.
`continuationArb` is now `minLength: 1` and the property states the invariant
instead of the old rule: for any prefix, phase, 1-5 continuations and any
one-digit terminator, the token equals the FULL continuation run on both the
imperative and the regex surface, and `matchPhaseDirs([dir], token)` returns that
dir. That third assertion is the one that catches the class on its own — the old
behaviour made a well-formed directory unresolvable by its own id, which is a
property, not a fixture.
Around it: "10-24-7-zip" and "10-24-3d-printer" now sit beside
"10-24-7-autonomy" everywhere the family is pinned, so the two readings can never
diverge again; the 24/7 metamorphic property asserts the RESOLUTION result rather
than the token (the token is precisely the part no surface may decide); the
end-to-end parity corpus gains "a sub-phase with a digit-leading slug resolves by
its full id" across all four resolution paths; and the milestone-scoping residual
is pinned in three directions rather than left to prose.
Mutation: re-inserting the rewind and rebuilding turns 9 tests red, the
`minLength: 1` property first, and nothing else. Build success checked separately
(build:lib reports 0 `error TS`), so the mutation reached the artifact under test.
* test(#2528): pin the #2946 guard against digit-leading phase directories
The #2946 fix makes the milestone-complete unstarted-phase guard run
unconditionally, so whether it fires now rides entirely on the
directory-resolution owner this PR replaces. Two cases, both with STATE.md
carrying no `milestone:` field so the #2946 path is the one exercised:
- ROADMAP Phase 5, disk `05-80-20-cleanup` → guard must stay silent.
RED on next (fail-closed: the guard blocks a legitimate one-way-door
operation because phaseTokenMatches resolves neither 05 nor 80 for
that directory).
- ROADMAP Phase 80, same directory → guard must still fire. Green on
both sides; it pins the fail-open direction against a future widening
of the matcher.
* fix(#3175): stop the injection scanner reading RegExp.exec as code execution
The apostrophe fix in
|
||
|
|
4a0b5e26b4 |
test(#3338): fold the verify/validate & workflow-text issue-* cluster — Wave 6 (#3383)
* test(#3338): fold the verify/validate & workflow-text issue-* cluster — Wave 6 Folds 10 legacy issue-*.test.cjs regression files (75 test() blocks) into their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053). Third of 4 issue-* waves. - issue-2701-nul-corrupted-validators.test.cjs (9) + issue-429-comment- text-gate.test.cjs (31, incl. fast-check property tests): both target verify.cjs/validate.cjs via different calling styles (CLI vs direct- require) — merged jointly into verify.test.cjs, 0 dropped. - issue-2762-plan-reviews-chunked.test.cjs (3) merged into plan-phase-drift-guard.test.cjs. - issue-2771-advisor-subagent-type.test.cjs (1) + issue-2772-discuss- phase-text-inconsistencies.test.cjs (6) merged jointly into discuss-phase-power.test.cjs. - issue-498-update-backup-runtime-dir.test.cjs (3, rename basis) + issue-815-update-next-channel.test.cjs (7, merged in): both concern update.md workflow-text contracts, now update-workflow.test.cjs. - issue-498-update-context.test.cjs (13): pure rename to update-context.test.cjs, sole comprehensive suite for its module. - issue-2765-brace-expansion-lockfile.test.cjs (1, rename basis) + issue-3238-js-yaml-lockfile.test.cjs (1, merged in): two distinct CVE regression pins against package-lock.json, now lockfile-cve-audit.test.cjs, shared ROOT/npmLs helpers deduped instead of double-declared. Ratchet upkeep to keep this wave's own gates green: pruned 3 stale allow-test-rule allowlist entries, cited 2 previously-uncited comments that surfaced in the folded content (#3338), tightened the exemption-file ceiling 305 -> 297. Fixed one stale filename reference in production code (src/init.cts) plus one in docs/reference/workflow-fragments.md. Zero net test-coverage loss. No production code BEHAVIOR changed. * test(#3338): fix orthogonal-review findings — Wave 6 fold Standards-axis review found a real structural defect in two files, both the same root cause and both fixed here: - tests/update-workflow.test.cjs: the folded:issue-815-update-next-channel wrapper's closing brace was placed after the file's pre-existing tail instead of before it, making the already-established folded:bug-2470 and folded:bug-3130 wrappers CHILDREN of issue-815's block in the test hierarchy instead of independent siblings — confirmed via an actual node --test run showing the mislabeled TAP nesting. Moved the closing brace to the correct position; all three fold wrappers are now top-level siblings again (verified via node --test, TAP hierarchy correct, 10/10 tests, 5 suites, identical count before and after). - tests/plan-phase-drift-guard.test.cjs: same mistake in the other direction — folded:issue-2762-plan-reviews-chunked was spliced inside the pre-existing folded:bug-2492-context-coverage-gate wrapper instead of after it. Fixed the same way (227/227 tests, 38 suites, identical count before and after). - Reverted unnecessary 815-suffixed local renames (assert815/fs815/etc.) introduced by the fold — the wrapper is genuinely block-scoped once correctly closed, so no collision existed (same class as Wave 4's __foldSetNested finding). No test() count changed in either file. No production code touched. --------- Co-authored-by: sim <sim@local> |
||
|
|
80734a9694 |
chore(#3053): clock-seam ADR-456 amendment + backfill — H2 (#3332)
* test(#3314): backfill deterministic clock-seam coverage (failing-first) Replaces loose regex/range assertions with exact-value and boundary tests for the CLI-subprocess and in-process clock-touching call sites identified by H2's audit (epic #3053): cmdCurrentTimestamp, _wsParseRetryAfter (commands.cts), cmdInitManager's is_active gate and cmdInitQuick's quick_id generation (init.cts), and reapStaleTempFiles (io.cts). The CLI-subprocess-pinned tests are expected RED until a follow-up commit routes those call sites through realClock so GSD_TEST_MODE+GSD_NOW_MS can reach them. * refactor(#3314): route CLI-subprocess clock reads through realClock cmdCurrentTimestamp, cmdInitManager's is_active gate, and cmdInitQuick's quick_id generation read Date directly, which the GSD_TEST_MODE+GSD_NOW_MS subprocess pin cannot reach (it only fires inside realClock.now()). Behavior-preserving: realClock.now() falls through to Date.now() whenever GSD_TEST_MODE is unset, which is every real invocation. * docs(#3314): amend ADR-456 with reachability-based clock-control rule ADR-456 §(a) documented one mechanism (injected {clock=Date} + t.mock.timers). Adds the two this repo already relies on: t.mock.timers for in-process direct-Date reads, and the GSD_TEST_MODE+GSD_NOW_MS subprocess pin (routed through realClock) for CLI-spawned code. Updates TESTING-STANDARDS.md's matching passages, which already referenced this issue by number as the no-elapsed-assertion promotion precondition. * fix(#3314): address orthogonal review findings Spec-axis findings: file and link the no-elapsed-assertion promotion follow-up (#3331) instead of leaving TESTING-STANDARDS.md pointing at a dead #1885, and ship the module-by-module audit table in the ADR itself rather than only in a gitignored phase artifact. Standards-axis finding: pin the "hour-old file = not active" test via GSD_TEST_MODE+GSD_NOW_MS for consistency with the reachability rule this PR's own ADR amendment now documents. --------- Co-authored-by: sim <sim@local> |
||
|
|
e201cde73c |
refactor(#3186): one shared phase-completion predicate, disk-strict (#3306)
* docs(#3186): record the disk-strict completion decision in ADR-3180 7.4 The maintainer decided #2957 on 2026-08-08: disk state is authoritative and a ROADMAP checkbox is a human annotation with no machine authority. Section 7.4 still carried the OPEN QUESTION and was marked blocked, so the contract said one thing and the tracker another. Recorded per section 7's own rule - a behavior not stated there is not decided, and amending a rule is an ADR amendment rather than a code change with a comment. The decision comment names Phase 4's PR as the carrier of this edit and makes it an acceptance criterion that the text be in the tree before implementation begins, so this lands first, alone, ahead of any code. Also clears the stale blocked-on-2957 row in the guard roster. * refactor(#3186): one shared phase-completion predicate, disk-strict isPhaseComplete in verification.cts becomes the single owner. It calls readVerificationStatus UNCONDITIONALLY - plan count is not a precondition - so a zero-plan phase with a passing VERIFICATION.md is complete. That is #3168: init gated the read on a plan count and synthesized a not_required sentinel, so phase.complete succeeded while init.manager reported incomplete for the same phase. The guard, built and run before scope was fixed per Amendment 3, found 9 re-derivations where the ADR named 3. Four were unnamed, including one in the prompt layer: mvp-phase.md ORed a ticked checkbox with disk status, which under disk-strict is the divergence itself. Per the #2957 decision, a ticked ROADMAP checkbox is a human annotation with no machine authority. The overrides in roadmap analyze and init manager are deleted rather than generalized; the user's checkbox stays in ROADMAP.md, only its authority goes. scanPhasePlans.completed and buildWorkstreamInventory are deliberately NOT folded - they answer 'are all plans summarized', which is a different question, and folding them would either over-report completion or invert the dependency direction between Phase 1's owner and this one. Verified on the remote runner. * fix(#3186): close seven review findings and record the missing-verdict rule The isolated review reproduced a write-path regression I introduced: migrating cmdRoadmapUpdatePlanProgress dropped its summaryCount>=planCount gate, so a phase with a fresh passing verification plus a newly-added unsummarized plan reported complete AND wrote a checkbox into ROADMAP.md while phase complete refused. The owner stays right per 7.4 - plan count is not a completion precondition - so the gate is restored at the write site as an explicit composition, mirroring the separate 2648 unexecuted-plan gate cmdPhaseComplete already carries. The spec axis was right that my 0.x-split reasoning was too permissive. The 2957 decision names buildStateFrontmatter as one of the three that must converge, and buildWorkstreamInventory combined a summaries-met local with verification data to decide the same verdict - Decision 4(c)'s named bypass, and it reproduced 3168 in a third surface. Both now route through the owner. The raw scanPhasePlans helper stays: it answers are-plans-summarized, which genuinely is a different question. Maintainer decision recorded in 7.4: a missing verdict is not a passing one, so an absent VERIFICATION.md means not complete everywhere. That retires 2645's verifier-disabled tolerance and inverts its Goodhart incentive - deleting the evidence now lowers completion instead of raising it. Guard hardened: block-form count gates and algebraic restatements are caught, and the header now discloses its remaining limits instead of overclaiming. Verified on the remote runner. * fix(#3186): route state sync through the owner and catch bare completed reads The matrix found 52 failures. 51 were fixtures asserting the old semantics: a phase with plans and summaries but no VERIFICATION.md used to count complete and correctly no longer does. Each fixture now carries a passing verification where that is what the test was actually about, rather than having its assertion weakened. The 52nd was a real 10th re-derivation the guard could not see. cmdStateSync destructured scanPhasePlans().completed directly - a bare field read, not a comparison - and used it as a completion verdict, so state sync and state json disagreed on completed_phases for identical disk state. Routed through the owner. Guard gains shape (d): any read of .completed off a scanPhasePlans() result outside plan-scan.cts, in chained, destructured and indirect forms, function scoped with no line window. It cannot tell a summaries-met read from a completion read - that is data flow - so it flags every one and requires a written-reason exemption, which is the same discipline shapes a-c already use. The blind spot is disclosed in the header rather than overclaimed. The emitted-attribution failure was also mine, not pre-existing: the mvp-phase.md checkbox-OR removal moves emitted bytes, acknowledged in tests/emitted-drift-acks. Verified on the remote runner. * test(#3186): give the nested-plans sync fixture a passing verification Last 3 matrix failures were one failure echoing up two describe levels. Phase 01-alpha had plans and summaries but no VERIFICATION.md, so under disk-strict completed stayed 0 and no Progress change was emitted - correct new behavior, not a regression. Added the passing verification rather than dropping the Progress expectation, so the test still covers what #3257 is about: that a nested plans/ layout is counted and not undercounted. Probe against the built lib confirms Progress: 0% -> 50% alongside Total Plans in Phase: 0 -> 3. * chore(#3186): backfill changeset PR number pr:0 placeholder replaced with the real number now that #3306 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
96a82bbffb |
enhance(#3245): report the detected host runtime in init (#3307)
* test(#3245): failing-first coverage for host runtime detection in init Locks the behavior epic #2313 Phase 5 must produce before any of it exists: init reports the detected host, explicit GSD_RUNTIME and config runtime still outrank detection, non-Codex sessions are untouched, and nothing is ever written to shared defaults (#2297). * enhance(#3245): report the detected host runtime in init init reported agent_runtime: claude inside a Codex session, and resolved agents_dir to the Claude agents root with agents_installed: true — a spuriously healthy triple. Runtime identity was only ever read from GSD_RUNTIME or an explicit runtime in .planning/config.json. Adds a detection rung beneath both explicit sources, in a new pure module. Codex is identified from its own documented session environment (CODEX_SANDBOX / CODEX_SANDBOX_NETWORK_DISABLED), else an explicitly exported CODEX_HOME whose config.toml exists. The default ~/.codex is never probed: that file exists on every machine that has run Codex, so probing it would misreport other runtimes' sessions. resolveRuntime keeps its exact contract and all 71 dependents, including formatGsdSlash command-style emission; only withProjectRoot consumes the new rung. Nothing is written on any path (#2297). Explicit config still wins (#2517). * fix(#3245): make the parity guard real and single-source the marker Four independent review passes found the generative-fix-divergence guard was vacuous: it asserted agreement at the one input where inferPreferredRuntime and detectHostRuntime do not differ, so it could not fail. It now pins the actual divergence point (CODEX_HOME set, config.toml absent) and records that the asymmetry is deliberate. The config.toml marker is now single-sourced from update-context.cts and imported, rather than carried independently by two surfaces. tests/helpers.cjs now scrubs CODEX_SANDBOX and CODEX_SANDBOX_NETWORK_DISABLED: GSD reads them, so an ambient Codex session would otherwise make the non-codex control test fail non-deterministically. Also: detection is throw-safe end to end rather than only around the fs probe; the Windows-join test is replaced with one that can actually fail (trailing-separator, catches hand-rolled concatenation); the #2297 no-write proof now wraps resolveReportedRuntime, the function that ships, across all three ladder outcomes. * chore(#3245): backfill changeset pr number --------- Co-authored-by: sim <sim@local> |
||
|
|
86bebcefa2 |
refactor(#3216): bind milestone identity to the canonical locator (#3226)
* refactor(#3216): widen milestone-window guard to literal-## matchers The guard keyed only on the `#{N,M}` quantifier plus a literal version or phase-lookahead token. getMilestoneInfo hand-rolls its milestone-heading match with a literal `^##`/`## ` and an interpolated ${escapedVer}, so it satisfied neither token and the guard reported a clean zero on a file carrying live re-derivations (#3171, #3197) — a zero it did not earn. Widen token (a) to a literal 2-6 `#` run, admitted ONLY inside a heading-MATCHER literal (a regex literal, or a string/template handed to new RegExp) so a heading-BUILDING template is not mistaken for a re-derivation. Widen token (b) with the grouped `v(\d+(?:\.\d+)+)` shape and an interpolated version placeholder. Ships BEFORE the consolidation per ADR-3180 s7.2: a guard widened afterwards measures an already-cleaned surface. It is expected to be RED until the consolidation lands. * test(#3216): failing-first milestone-identity single-owner suite 63 tests across two files, from the matrix in .gsd/phase/. Section H of milestone-window-single-owner.test.cjs covers the 21 input classes of the design's behavior table plus its negative space; milestone-window-drift-guard covers the widened tokens and proves the exemption is function-scoped, not file-scoped. Copy count is 3 found by the guard, not 1 per the epic (ADR-3180 Amendment 3's standing rule, holding for the fourth consecutive phase): both getMilestoneInfo sites plus cmdRoadmapAnalyze's milestone enumeration at roadmap.cts:454, which carries the same #3171 truncation and #3197 phase-heading confusion. Expected RED until the consolidation lands. * refactor(#3216): bind milestone identity to the canonical locator getMilestoneInfo hand-rolled two milestone-heading regexes inside the owner's own file. Both were wrong, differently: the STATE-version site's ^## anchor is level-blind so [^\n]* absorbs a third #, and the fallback site had no anchor at all, so '## ' matched from the second # of '###'. Against '### Phase 7: Close v3.3 gaps' the fallback returned {v3.3, gaps} (#3197). Both captured names with [^\n(], truncating at a parenthetical (#3171). Bind both to the canonical grammar. locateMilestoneHeadings becomes a version-filtered view over one shared source, and a new version-agnostic listMilestoneHeadings enumerates milestone headings for callers that need all of them. getMilestoneInfo returns ScopedResult<MilestoneInfo|null>; the {v1.0,'milestone'} default, which was output-identical to a real v1.0 project, is deleted. The #2245 never-throws invariant is preserved. Copy count: 3 found by the guard, not 1 per the epic. The third was cmdRoadmapAnalyze's own milestone enumeration (roadmap.cts:454), carrying both defects in the implementation the epic blessed. buildStateFrontmatter and archivePhaseDirectories branch on scope: the first writes null rather than a fabricated identity, the second falls through to its dated-label fallback. A fabricated v3.3 passes ARCHIVE_VERSION_LABEL_RE, so it would otherwise misfile phase history. Also fixes an unsafe cast in init.cts that masked these type errors across five call sites, which would have shipped undefined milestone fields under green tsc. * fix(#3216): restore the #1761 unbounded guard and bullet precedence Review and the first full-matrix run surfaced five real defects in the consolidation, all fixed here rather than by relaxing the tests that caught them: - buildStateFrontmatter gated its isMilestoneBoundedInRoadmap check on the scope-gated milestone value, which is null on any non-COMPLETE scope, so the #1761 unbounded guard was silently skipped and state json reported a percent it must omit. It now gates on the STATE-asserted version, independent of identity scope. - The rewrite lost #2135's precedence: the name-bearing progress-marker bullet is consulted before the heading again. - A single-segment version (v3, no dot) did not resolve; the name-extraction fallback now accepts it. - A version carrying regex metacharacters, or a $& / $1 replacement pattern, is matched literally. - listMilestoneHeadings' heading field trimmed, so a CRLF roadmap no longer leaks a trailing carriage return into roadmap analyze's output. Also emits milestone_version / milestone_name / current_milestone as explicit null rather than omitting the key, so the prompt layer cannot render a bare placeholder, and corrects an init.cts comment plus a cast left inconsistent. * test(#3216): update milestone-identity expectations to the scoped contract getMilestoneInfo returns ScopedResult<MilestoneInfo|null> and the {v1.0,'milestone'} default is deleted, so the suites asserting the old shape assert removed behavior. Updated rather than weakened: every touched call site now asserts the scope explicitly against the frozen SCOPE enum. roadmap-parser.test.cjs: 20 expectations moved to {value,scope}. The #1881 unreadable-vs-absent diagnostic assertions are untouched and still prove their original point — only the return shape moved. One pre-existing assert.ok(info) is now a specific UNSCOPED assertion, so that case is stronger than before. new-milestone-clear-phases.test.cjs: the test asserting phases clear archives under the v1.0 default now asserts the dated archived-<YYYYMMDD> fallback, which is the deliberate consequence of deleting that default. Two of this branch's own tests were also corrected after they drove the implementation the wrong way: the parity test compared raw heading text and so pushed a stray ## prefix into roadmap analyze's public output, and the hostile metacharacter row demanded a pathological version resolve, which pushed a widening of the ADR-locked \b boundary. Both now assert what the contract actually requires. * docs(#3216): document milestone identity and correct the CONTEXT.md entry ADR-3180 s7.2 moves to Enforced and gains two rules that were unstated: the name derives from the heading's own version token and drops a trailing status marker, and a free-form legacy ROADMAP with no version anywhere is UNSCOPED with no identity rather than a defaulted v1.0 (decided by the maintainer before implementation, per s7's own rule that an unstated behavior is not decided). Amendment 4 records Phase 6's validation, including that the copy count was a lower bound for the fourth consecutive phase. CONTEXT.md's Roadmap Parser entry described locateMilestoneHeadings as boundary-matched with (?![\w.-]) — the alternative Amendment 2 tried and REVERTED. The code uses \b and says so, and the ADR agrees; the revert updated code and ADR and missed CONTEXT.md, which is the epic's own fixed-on-one-copy failure class in the docs layer, on a file that is itself a PR gate. * fix(#3216): persist the real version on a truncated identity buildStateFrontmatter wrote null for BOTH milestone and milestone_name on any non-COMPLETE scope, discarding a real version. ADR-3180 s7.2 rule 6: a version known with no resolvable name is TRUNCATED carrying {version, name: null} — 'the version is a real answer, the name is a non-answer, and collapsing the two is the failure this contract exists to prevent.' The two fields are now gated by what is actually known: the version whenever one exists (COMPLETE or TRUNCATED), the name only on COMPLETE. Never fabricated. Caught by this phase's own Decision 4(c) consumer-output test, which is the argument for asserting at the consumer rather than the owner — the owner was correct throughout; only the consumer collapsed its answer. * refactor(#3216): extract helpers and make cmdCommit's scope gate explicit From the two-axis code review: - init.cts repeated the identical getMilestoneInfo cast at five sites with copy-pasted comments — duplication inside a PR whose thesis is that duplicates get deleted. Extracted milestoneRecord(cwd); the one site-specific comment is kept, the four generic copies removed. - getMilestoneInfo hand-built its { value, scope } literal at ten return points; a local scoped() constructor now does it once. Every per-branch rationale comment is preserved and no returned value or scope changed. - cmdCommit gated the milestone branch name on plain truthiness, which is also true for TRUNCATED, so an unresolved identity drove branch creation incidentally rather than deliberately. It now gates on the SCOPE enum, accepting COMPLETE or TRUNCATED because both carry a real version, and the comment records why that differs from archivePhaseDirectories — which demands COMPLETE because it uses the value as a filesystem path component. * test(#3216): cover the bare-version-in-prose truncated path The spec review found the bareVersionMatch path — no STATE version, no milestone heading, a version token only in prose — returning TRUNCATED with no test exercising that exact shape, violating Decision 4's boundary-coverage requirement. * docs(#3216): record the missed Tier-2 surfaces and rule 5's corollary Decision 3 requires an explicit call-out for EVERY Tier-2 change, and Amendment 4's first draft named eight surfaces while the change touched thirteen. Adds cmdCommit's branch-name construction and the four init JSON bundles, an incomplete list being the same defect in miniature that this epic removes. s7.2 rule 5 gains a corollary separating two cases the original wording ran together: no version token ANYWHERE is UNSCOPED, while a bare version token in prose or a non-milestone heading is weak but real evidence and yields TRUNCATED under rule 6. * chore(#3216): set changeset fragment pr to 3226 --------- Co-authored-by: sim <sim@local> |
||
|
|
636ec92107 |
refactor(#3185): phase enumeration has one owner and a decidable scope (#3222)
* test(#3185): failing-first phase-enumeration single-owner suite Covers the enumeration rows with direct code evidence: 999.* backlog dirs listed by progress/stats, the phase-0 sentinel divergence, the #1324 letter-prefixed-decimal negative space, and the destructive-path find — cmdPhasesClear carries a fifth sentinel copy (/^999(?:\.|$)/) that excludes 999 but not 0, so a 0-* directory roadmap.analyze preserves is deleted there. Also covers the pass-all degrade, which is where the defect actually lives: when the milestone window declares no phases the filter becomes a literal () => true and its heading-side sentinel exclusion is unreachable. A fixture carrying phase headings keeps the filter active and never reaches that path. Named for the derivation, not a module: the suite drives commands, phase, milestone, workstream-inventory and state, and both the phase and phase-locator buckets are already at the per-module test-file cap. Committed alone so the remote runner records the failure before the fix. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): phase enumeration has one owner and a decidable scope Adds phase-locator.cts::listMilestonePhaseDirs as the single canonical owner of "which phase directories belong to the current milestone". It applies the milestone window AND the sentinel filter and returns a ScopedResult, so a caller can tell a genuinely-empty milestone from an enumeration that could not be scoped. The sentinel test now runs against DIRECTORY NAMES and is unconditional. getMilestonePhaseFilter excludes sentinels from its ROADMAP heading set, but degrades to a literal () => true pass-all predicate when that set is empty -- at which point the heading set is never consulted and its sentinel exclusion is unreachable exactly when it is needed. That degrade is the #3167 path, and it is why stats already used the filter and still listed backlog directories. The narrowing is sentinel-only: pass-all stays over-inclusive otherwise. Sentinel copies deleted, canonical isSentinelPhaseId adopted: - cmdRoadmapAnalyze's local closure (parseInt === 0 || === 999), 2 call sites - cmdPhasesClear's /^999(?:\.|$)/ -- the DESTRUCTIVE path, which excluded 999 but not 0, so a 0-* directory roadmap.analyze preserves was deleted cmdStats also seeded rows from ROADMAP headings with no sentinel filter, so a 999 heading produced a row with no directory; that seed is filtered now. cmdPhasesList routes only its ENUMERATION. --phase lookup searches the physical set (scoping it would report an out-of-window phase as not found) and --include-archived still merges archived dirs (they are by definition from other milestones). Both exempt by documented reason, never a file allowlist. Fixed inline, found while building: isDirInMilestone could not match a #1324 letter-prefixed-decimal directory (P0.0-foundation) to its own Phase P0.0 heading, so stats reported the phase with plans: 0 while its directory held plan files. Defers to phase-id's extractPhaseToken rather than widening a fourth bespoke regex; additive, so it can only admit directories. Refs #3180. Closes #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): route the last two enumeration re-derivations workstream-inventory countRoadmapPhases counted every `Phase` heading across the whole ROADMAP -- no window, no sentinel filter -- so it counted 999.* backlog and Phase 0 and spanned every milestone the document ever had. Its own caller already resolved a currentVersion and passed it to getMilestonePhaseFilter elsewhere in the same file; this was the sibling copy that never got the fix. state.cts phaseInventoryProvider enumerated phase dirs with its own /^(\d+)-(.+)$/ convention regex and neither filter, so a rebuilt STATE.md inventory carried backlog and sentinel directories as current-milestone phases. A non-COMPLETE enumeration scope now throws to the outer catch as a real scan failure rather than reporting a confident undercount, mirroring the per-phase scanPhasePlans contract beside it. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): consolidate 23 sentinel re-derivations onto one predicate The whole-repo drift guard (ADR-3180 Decision 4a, no file allowlist) found the sentinel rule re-implemented 23 times across 8 modules, in three regex variants plus four integer-comparison forms. Most tested 999 only, so Phase 0 slipped through them while roadmap.analyze and the engine-wide convention (#1580) both treat 0 and 999 alike. That disagreement is the defect class this epic removes. All 23 now call phase-id's isSentinelPhaseId (SENTINEL_RANGES [0,999]). Sites: init recommended-actions and backlog counts, milestone phase scan, the phase-lifecycle progress table, phase.cts used-number collection and the four renumber-on-remove guards, roadmap-parser's heading and bullet milestone counts, roadmap get-phase fallbacks, and state's heading denominator. Excluding Phase 0 at these sites is a deliberate behavior change and the point of the consolidation — several carried comments already saying 0 should be excluded while the literal beside them caught only 999. Adds scripts/lint-phase-enumeration-drift.cjs, wired into lint:ci. It scans the whole src/ tree with no file allowlist and reports both shapes: an independent phases-dir enumeration, and an independent sentinel literal. Exemptions are function-scoped with a written reason. The guard is comment-aware — its first pass flagged JSDoc and a comment documenting that the code below uses the canonical owner, which would have trained readers to exempt prose. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * refactor(#3185): resolve every phases-dir enumeration; drift guard reports zero Per-site triage of the 31 remaining whole-repo guard hits, applying the rule generalized from #3183's Amendment 1: a LOOKUP, DIAGNOSTIC, ARCHIVAL or MUTATION pass wants the physical set; only "which phases belong to this milestone" wants the scoped set. Routed (10): init new-milestone phase_dir_count, init milestone-op fallback count, init manager, init progress, milestone complete stats/dry-run/archive move, phase complete's next-phase scan, state update-progress, state frontmatter stats, and uat audit's active set. Exempt with a written function-scoped reason (never a file allowlist): the audit/UAT/verification sweeps that deliberately scan every directory to report gaps, phase create/insert/rename/renumber mutations, single-phase lookups, roadmap-upgrade's cross-milestone migration, cmdPhasesClear's whole-tree destructive pass, and the reads that list a phase dir's FILES rather than enumerating the phases dir at all. Latent defects fixed by the routing: sentinel directories leaked into cmdInitNewMilestone's phase_dir_count, cmdMilestoneComplete's stats, dry-run AND ARCHIVE MOVE, cmdStateUpdateProgress, buildStateFrontmatter and cmdAuditUat's active set — every one of those hand-rolled an isDirInMilestone filter with no sentinel exclusion, so `milestone complete` was archiving backlog directories. scripts/lint-phase-enumeration-drift.cjs now reports 0 re-derivations and npm run lint:ci is green. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * docs(#3185): document milestone-scoped enumeration and record ADR Amendment 3 Changeset fragment (Changed), CLI-TOOLS/COMMANDS/USER-GUIDE updates for the scoped output of progress, stats, phases list, phases clear and milestone complete, the CONTEXT.md Phase Locator glossary entry naming listMilestonePhaseDirs, and ADR-3180 Amendment 3. Amendment 3 records: the SCOPE contract held unchanged; the declared deviation from Decision 1's provisional signature (the window needs cwd/ws, which the locked roadmapContent parameter cannot supply); the copy count being a lower bound for the third consecutive phase (4 scoped vs 54 found); the load-bearing finding that the sentinel exclusion sat on the heading set and was unreachable under the pass-all degrade; the two destructive-path defects; and the generalized exemption rule. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): wire scope to consumers; revert two wrong routings the suite caught Review + remote runner findings, all fixed: The three consumers computed the enumeration scope and threw it away, so TRUNCATED/UNSCOPED/UNREADABLE collapsed into the same output as COMPLETE -- reproducing this epic's own output-identical-failure defect one layer up. progress, stats and phases list now emit phase_scope (null on the phases list --phase lookup path, which performs no enumeration). Two routings were wrong and the suite proved it: roadmap-parser's two milestone phase-count scans are reverted to the 999-only literal. isSentinelPhaseId is BROADER than what it replaced: its legacy branch runs /^0*(\d+)/ over "00.1", which backtracks to capture 0, so it read #2554's decimal phase ids as sentinel milestone 0 and stopped counting them. state.cts phaseInventoryProvider is reverted to the physical disk scan. `state rebuild` is a RECONCILIATION pass -- scoping it made it throw on healthy trees whose fixture resolves no window, swallowed the raw readdirSync fault message #3057 B1 requires verbatim, and stopped it dropping orphan STATE.md rows, which is the job. Both are now function-scoped guard exemptions with written reasons, not silent reverts. This is the consolidation trap named in the epic: a canonical rule can cover MORE than the copy it replaces, and only real inputs show it. Adds phases list coverage, a scope-branch test, and a drift-guard unit suite; backports comment-awareness to the milestone-window and plan-count guards so all three siblings share one false-positive profile; names #3161 alongside #3167 in Amendment 3's subsumption record. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): correct isSentinelPhaseId's decimal-zero misclassification An isolated security review caught this branch committing the epic's own sin: the over-broad predicate was worked around at ONE call site and left live at the destructive ones. isSentinelPhaseId's legacy branch ran /^0*(\d+)/, which backtracks so any id whose leading digit run is all zeros before a non-digit captures 0 -- "0.1", "00.1" and "0.2554" all read as sentinel milestone 0. Two pinned contracts disagree with that: #2554 requires "00.1" to be counted as a real phase, and the 999 icebox is a whole reserved milestone so "999.1" must stay sentinel. The rule is asymmetric and now says so explicitly: 999 is sentinel with or without a decimal part; 0 is sentinel only when bare. A decimal phase under either is a real phase for 0 and reserved for 999, because 999 reserves a MILESTONE while 0 reserves a PHASE. Fixing the owner lets the earlier workaround go: getMilestonePhaseFilter's two scans route through isSentinelPhaseId again and the guard exemption that existed only to accommodate the defect is deleted. The state.cts cmdStateRebuild exemption stays -- that one is a genuine reconciliation-wants-the-physical-set case. Also corrects tests/adr-612-bracket-grammar.test.cjs, which asserted isSentinelPhaseId('0.1') === true and so had encoded the defect as expected behavior. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * fix(#3185): keep isSentinelPhaseId's semantics — 0.x is layered, not wrong Reverts the previous commit. The remote suite failed six tests proving it wrong, and the reason is the sharpest finding of this phase. An isolated security review observed that isSentinelPhaseId reads 0.1 and 00.1 as sentinel milestone 0 and judged that a defect against #2554. Correcting the canonical predicate broke #2949. Both contracts are pinned and both are right, because they ask different questions: #2554 is this dir part of the current milestone's phase SET? -> count 00.1 #2949 must this phase COMPLETE before the milestone closes? -> 0.x sentinel No single global predicate answers both. isSentinelPhaseId keeps its semantics (0.x IS a sentinel, #2949), and the milestone-window layer keeps a narrower 999-only rule (#2554) as a function-scoped guard exemption with a written reason — not a second silent copy. That corrects how Decision 1 reads: "one owner per derivation" governs who computes an answer, not how many questions share it. An over-broad canonical rule is as much a defect as a divergent copy and fails worse, because it looks like consolidation. Recorded in Amendment 3 as the lesson for Phases 4 and 5. Where a review's inference about intent conflicts with a pinned contract, the pinned contract wins; the finding is adjudicated, not fixed. The boundary tables in the enumeration suite are corrected to assert 0.x IS a sentinel, with the layering explained. Refs #3180 #3185. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG * chore(#3185): set changeset fragment pr to 3222 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QELmgcSwcNBgbUs3kzJeqG --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2bead6ca1d |
feat(#3149): add dedicated init.debug entry point for /gsd:debug
/gsd:debug was one of the last workflows with no cmdInit* of its own: its Step 0 made three separate round-trips (state.load, resolve-model gsd-debugger, config-get workflow.tdd_mode) to assemble one context. Because no debug-scoped fact was computed at any entry point, ADR-1671 admission gate (2) could never be satisfied for debug — an applicability atom naming such a fact would evaluate FALSE forever and silently exclude its section. Adds cmdInitDebug (init.debug), registers it in the init router and the command-alias table, and collapses debug.md Step 0 to one call. Every field resolves through the same primitive the call it replaces used: loadConfig for commit_docs, withProjectRoot for response_language (#2402), planningPaths for debug_dir, resolveModelInternal for debugger_model, and the existing Boolean(workflow.tdd_mode) idiom for tdd_mode. PlanningPaths gains a debug field so state.load and init.debug share ONE debug-directory expression rather than two kept in sync by hand. state.load keeps emitting debug_dir: it is a shipped query surface with its own test anchor, so narrowing it would break unseen consumers for no gain. No WHEN_VOCABULARY atom and no gsd:section marker: gate (1), a consuming section of at least 400 bytes, belongs to the change that adds the section. Closes #3149 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fb3ee56651 |
fix(#3033): resolve zero-plan split-parent phase as complete when roadmap checkbox is checked (#3114)
* fix(#3033): resolve zero-plan split-parent phase as complete when roadmap checkbox is checked A phase split into sub-phases (parent kept as shared context, zero plans by design) was permanently stuck as 'researched' because the roadmap- checkbox override at line 2266 required completion.phase_complete (derived from plan/summary counts), which is always false for zero-plan phases. The parent was permanently eligible for current-phase selection and re-planning recommendations. The override now fires when roadmapComplete AND planCount === 0 (the split-parent shape), treating it as complete regardless of the plan-count derivation. A zero-plan phase whose checkbox is still unchecked stays in-progress (researched). Ordinary phases with plans are unchanged. * chore(#3033): backfill changeset PR number 3114 --------- Co-authored-by: sim <sim@local> |
||
|
|
53ea8e0664 |
fix(#3057): make a guard's failure distinguishable from its benign result — Wave 1 (#3088)
* fix(#3057): refuse the write when the duplicate scan cannot complete writeManifest documents itself as a fail-closed duplicate guard: if any existing manifest shares plan_id with a different, non-terminal job_id it must refuse, because dispatching again would duplicate the external job. It could not honour that. The scan reads every sibling manifest looking for the duplicate, and an unreadable or unparseable sibling was `continue`d past. If the corrupt file was the one holding the live duplicate, the scan found nothing and a duplicate external job dispatched. The asymmetry is what gives it away: a malformed TARGET refused with malformed_existing because clobbering is unacceptable, while a malformed SIBLING was skipped — yet siblings are the only thing the duplicate check reads. Adds a scan_incomplete verdict that refuses and names the offending file, so an operator can quarantine or repair it. Fail-closed alone would let one stale corrupt manifest wedge every dispatch for that planning dir permanently; naming the file is what makes refusing survivable. malformed_existing is untouched, so the target/sibling distinction stays visible. The docstring is updated — it previously stated a rule the function did not keep. memFs() gains an optional failReads map so these branches are reachable at all; they had zero coverage because the fake could not express a per-file read fault. The signature is additive and every existing caller is unchanged. The regression is proved by a pair, not a single test. A control writes a readable sibling holding a genuine non-terminal duplicate and asserts duplicate_plan_id, establishing the scenario is real; the regression then makes that same path unreadable and asserts scan_incomplete. A first draft of this test used a corrupt-JSON fixture containing no plan_id at all while its comment claimed otherwise — it duplicated the unparseable-sibling case and proved nothing, which is the defect class this phase exists to remove. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3057): make a guard's failure distinguishable from its benign result Wave 1 of the negative-space backfill: the branches where a guard that could not verify something reported the same value it reports when everything is fine. That indistinguishability is the defect; every fix here makes the two states tellable apart, and every test proves it with a pair — one for the failure, one for the benign case. A single test cannot establish that two states are distinguishable, which is the whole property being fixed. state.cts phaseInventoryProvider returned null for both a real disk-scan failure and a genuinely empty phases dir, so `state rebuild` could report success while phase-table reconciliation never ran. It now returns a discriminated result and the CLI surfaces phase_inventory_scan_failed plus a reason. The reason field turned out never to have been wired into the emitted JSON at all — it existed only as an internal variable — so a test could only assert on the operator-facing note. It is a real field now. state.cts treated an unreadable lock body the same as an empty one, applying the 1-second stealable floor. A lock we cannot read is not a lock we know is stale; an unreadable body is now held to the deadman ceiling like a live holder. verification.cts findStaleVerificationSummary returned null on any fs, scan or clock failure — meaning "not stale". It now returns a discriminated StaleCheckResult and the caller records that the check was indeterminate. git-base-branch resolveBaseBranch returned 'main' both when no candidate branch existed and when every git tier timed out. A diagnostics variant now reports whether the answer was verified, and the CLI writes an unverified-fallback note to stderr. The stdout contract five workflows parse is untouched. worktree-safety snapshotWorktreeInventory left exists:true when statSync threw, so a guard that could not check reported the worktree present; exists is now tri-state and a stat failure surfaces as an 'unverified' finding. planWorktreePrune reported 'no_worktrees' for a parse failure, which is not the same as an empty list — and it drives a prune. It now reports 'parse_failed'. Fixing the inventory change exposed a second fail-open in verify.cts: the validate-health consumer silently dropped findings whose kind it did not recognise, so the new kind would have vanished. That is closed too — worth noting that the survey enumerated producers of degraded verdicts, not consumers that discard them. worktree-base-ref and state-transition gain the distinguishing signal without changing what they do: headAbsenceVerified, and a phase-inventory scan meta. Whether those guards should ACT differently is a product question this change does not answer, and both are flagged rather than quietly settled. rescueSummaryArtifacts is left alone: rescuing on an uncertain cat-file is deliberate per #2556. It now has tests proving it, and a recorded negative finding — git cat-file -e returns 128 for both "absent from HEAD" and a fatal error, so "uncertain" and "certain-and-fine" are not separable at the git level. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): assert typed values, not rendered text Ten assertions in the rebuild CLI suite matched substrings of produced output — STATE.md body fields, a markdown table row, an audit-log heading, and JSON keys read as text. CONTRIBUTING prohibits that: if the code under test produces text, the test asserts on its structured surface instead. No production surface had to be built. Every one already existed and was already compiled into bin/lib: stateExtractField for body fields, parseMarkdownTable for the phase table, collectSection for the audit-log section, and result.data.log — already a typed RebuildLogEntry[]. The tests were matching rendered text sitting next to the structured data. One of those assertions was passing for the wrong reason. `stdout.includes ('rebuilt')` matched the JSON KEY name, not a value: the dry-run path emits `mutated` and the real path emits `rebuilt`, so it would have passed whether the value was true or false. It now asserts the value. external-job's refusal already had to name the offending file — that naming is why the fail-closed variant is survivable rather than a permanent wedge — but the tests proved it by substring of a prose message. The failure result now carries offendingPath as its own field and the tests assert it by value. The human message is unchanged; operators read it. Array membership is left alone. `phaseIds.includes('99')` and `result.updated.includes('Completed Phases')` are membership checks on real arrays, not text matching, and converting them would weaken nothing and clarify nothing. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): execute acquireStateLock instead of grepping its source The non-EEXIST lock test asserted on the TEXT of the built .cjs and never called acquireStateLock. It carried an allow-test-rule: architectural-invariant exemption to permit that. A source grep proves a literal is present in a file, not that the behaviour works — it is weaker than a liveness test, which at least runs the code, and it was the only coverage the fatal-errno path had. Replaced with tests that inject the errno through fs and assert what actually happens: a fatal EACCES propagates out of acquireStateLock with zero backoff sleeps, while EAGAIN/EINTR/EINVAL/EIO/ENOENT/ESTALE/EPERM/EBUSY retry once and succeed. The exemption is removed and its allowlist entry with it. One old assertion is deliberately not carried over: it checked the retryable errnos were expressed as a Set rather than an inline literal. That is a shape check with no runtime signature; the behavioural tests fail if the code reverts to the old inline check, which is the regression it was really guarding. The #3057 lock-body tests move into that same file rather than a new one, which is what lint-test-file-count asks for and puts every acquireStateLock test in one place. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3057): surface an indeterminate staleness check to its callers An isolated review caught an inconsistency inside this wave. Two of the three "add the distinguishing signal" fixes wire through to something a user sees: git base-branch writes an unverified-fallback diagnostic to stderr, and an unverifiable worktree surfaces as a W020 finding. The third set staleCheckIndeterminate on readVerificationStatus's result and nothing read it. A signal nobody consumes leaves the fail-open exactly as silent as before: the staleness check could fail and the operator saw precisely what they would see if the answer were genuinely "not stale". That is the defect this issue exists to remove, so it is not defensible as scaffolding when its two siblings in the same change already wire through. All five callers now surface it, each through the channel it already had rather than a mechanism imposed uniformly: phase complete adds it to its existing warnings array and, on the blocked path, as an additive note on the error text; init and roadmap carry it as a field on output they already emit; the UAT report carries it without ever gating passed/blockers; workstream inventory takes an injectable writeDiagnostic mirroring the git base-branch idiom, because its return shape had nowhere to hang a per-phase field without rippling the builder's types. The routing decision is unchanged everywhere. What changes is only that a caller and an operator can now tell a failed check from a completed one. That diagnostic carries structured meta rather than being asserted by regex — the default still writes only the human message to stderr, but tests assert phaseDir and reason by value. Two earlier assertions in this branch were converted the same way; this was the last raw-text assertion left. Also records a scope correction: the completePhaseCore guards now compare stateReplaceField's result to the body instead of testing truthiness, so a field whose substitution produced identical text no longer reports as updated. That is a real behaviour fix, not the signal-only change this file was described as carrying, and its tests cover both the changed and unchanged cases. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): bound two heavy subprocesses for a loaded bench, not an idle one The remote matrix surfaced three failures unrelated to this branch's changes. All were bad tests, and a re-run would have hidden every one of them. The reviewer-flags parse block bounded bash -> node -> a full gsd-tools cold start at 5 seconds. On a bench running thirty thousand tests in parallel that is not a hang, it is a busy machine. Raised to 30s, matching the convention sibling suites already use for script invocations, with a comment saying what the budget covers so nobody tightens it back. Two further copies of the same 5-second spawn in the same file had the identical defect and are raised too — they were not in the failure report, but they will be next time. The fragment-propagation test bounded npm run regen:derived — a full build plus eight generators, the heaviest subprocess in the suite — at five minutes, and node22 was killed near the end. The captured output proves it: every generator had written its files and gen:install-tree had emitted all fifteen runtimes before the kill. Raised to fifteen minutes. That failure read as `null !== 0`, which says nothing. status null means killed, not a non-zero exit, and the two want different responses: one is a timeout to size correctly, the other is a real build break. The assertion now distinguishes them and names the signal. Neither test's assertions were weakened and no retry was added. A retry here would suppress exactly the signal the timeout exists to produce. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): capture fd 1 through the mock tracker, not a raw reassignment The phase suite reported zero test results on both lanes while running for five and a half minutes and exiting 1. No assertion text, no stderr, four events for the whole file: enqueue, start, dequeue, complete. That shape is not a failing assertion — it is the runner being unable to read the child at all, because it parses its event stream from the child's stdout. The cause was the capture helper reassigning fs.writeSync directly. Proven rather than assumed: a standalone probe patched fs.writeSync and called process.stdout.write, and the interception fired only when fd 1 resolved to a FILE, not when it was a pipe. The remote runner captures the event stream to a file, so a helper that was invisible against a pipe swallowed the reporter's own output on the bench. That is also why the two sibling suites wired the same way in this change pass cleanly — they use the mock tracker, the seam io.test.cjs established for this exact function. The helper now uses t.mock.method with an explicit restore after each call, so teardown belongs to node:test rather than a second hand-rolled implementation, and the interception cannot outlive the one synchronous call it wraps even if that call throws. Ten call sites thread the test context through; three test callbacks gained the parameter they lacked. The three B3 tests are untouched — same assertions, same fault injection. Only how the context reaches the helper changed. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3057): capture phase-complete output from a subprocess, not fd 1 Two attempts to make in-process fd-1 interception safe both failed on the bench. The suite reported zero test results on either lane while exiting 1 — four events for the whole file — because the runner parses its event stream from the child's stdout, and process.stdout.write routes through fs.writeSync whenever fd 1 resolves to a file, which is how the runner captures. Patching that seam anywhere in a file can therefore destroy the file's own reporting, and tightening the window only moved the runtime from 326s to 125s without recovering a single event. So the interception is gone rather than tuned. The helper now spawns gsd-tools as a real subprocess and reads stdout the way the OS already gives it to us, which is what the rest of the suite does. It asserts the command succeeded before parsing, so a genuine failure can no longer present as a JSON parse error. The two fault-injecting tests could not survive that move as written: a subprocess cannot see a mock installed in the parent. Instead of reinstating the interception they now produce the fault on disk — the summary artifact is created as a dangling symlink, so the staleness check's real statSync throws inside the child. That is a more honest fixture than a mock in any case, since it is a condition a user's tree can actually be in. Skipped on Windows, matching the existing symlink precedent in the write-guard suite. Three further call sites turned out to depend on parent-process writeFileSync mocks the subprocess could not see. Those call the CJS function directly, which is what they always wanted — they never needed stdout at all. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3057): one name for one signal, one encoding for one distinction Standards review found four things this branch introduced, all of them inconsistencies with itself rather than with the repo. One upstream bit reached its consumers under three names — verification_stale_check_indeterminate in two modules, the same value with "stale" dropped in a third, and stderr only in the fourth. Standardised on the long name wherever it is a field. The workstream inventory keeps its stderr channel, since its return shape has nowhere to hang a per-phase field without rippling the builder's types, but it now says the same word for the same thing. worktree-safety encoded one three-way distinction two ways in a single file: a named union for a finding's kind, and boolean|null for an inventory entry's existence. The second is now a named union too. Two assertions matched human prose because the blocked and non-blocked completion paths carried no typed field for the signal. Both now assert typed values. The first round of this fix added the field but left the regex beside it, which is the banned pattern sitting next to its own replacement; the second removed it and added an assertion on the reason enum so nothing was lost. The remaining two were reasoned away before being fixed, and both reasons were bad. "No typed surface exists" is the condition CONTRIBUTING says to fix by adding one — it took three lines. "The file already does this dozens of times" is not licence to add instance number thirty-one; a convention that violates a documented rule is debt, not precedent. Vocabulary differing across DIFFERENT modules is left alone: CONTEXT.md rejects a single shared result envelope, so per-module shapes are precedented, and a baseline smell does not outrank a documented standard. A census of every line this branch adds to a test file now finds no regex or substring assertion on produced prose: 87 strictEqual, 25 ok (all non-empty or shape guards), 12 equal, 3 throws (all typed err.code predicates), 3 deepStrictEqual, 2 notStrictEqual. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#3057): backfill changeset pr number to 3088 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ef823ca9d9 |
fix(#2830): propagate a halted plan to its transitive dependents (#3038)
* test(#2830): add failing regression tests for halted-plan dependent blocking Add tests/fix-2830-halted-plan-dependents.test.cjs covering direct, transitive (2 and 3 hop), and diamond dependents of a halted plan across both independent "which plans are incomplete" readers (phase-plan-index's cmdPhasePlanIndex and findPhaseInternal/searchPhaseInDir), the negative case (an unrelated decoupled plan stays runnable), and a parity check that the two readers agree. Uses only modules that already exist at this commit (gsd-tools.cjs via subprocess, the pre-existing phase-locator.cjs) so the test file loads and runs cleanly on a fresh clone of this exact commit. These fail against current behavior: neither reader has any concept of a halted plan or a blocked_by/runnable view yet. * fix(#2830): a halted plan no longer leaves its dependents on the runnable work list A plan that reaches a designed stop still writes a SUMMARY, so both "which plans are incomplete" readers saw it as an ordinary completion and reported its dependents as ordinary runnable work — never checking whether an upstream plan had halted rather than finished. - New `status: halted` frontmatter value, documented in all four SUMMARY templates alongside the existing `status: complete`. - New shared src/plan-dependency-graph.cts: a single computeHaltPropagation pass that both phase.cts's cmdPhasePlanIndex (wave-grouping) and phase-locator.cts's searchPhaseInDir (the phase-location primitive, ~50 dependent symbols across 5 command routers) now call, so the two-implementation divergence that caused this bug cannot recur. It accepts an optional precomputedOrder so cmdPhasePlanIndex — which already runs Kahn's algorithm in computeDependencyLevels for wave assignment — passes that order straight through instead of a second traversal; searchPhaseInDir (no prior traversal) lets the module derive its own. The two small duplicated predicates each reader would otherwise carry (is this status "halted"?, which summary file matches which plan id?) are centralized in the same module as isHaltedStatus/buildSummaryFileIndex. - Additive fields only: `halted`/`blocked_by`/`runnable` on cmdPhasePlanIndex's plans[] and top level, `halted_plans`/`blocked_by`/ `runnable_plans` on searchPhaseInDir's result. The pre-existing `incomplete`/`incomplete_plans` fields are unchanged in meaning and membership. - execute-phase.md's discover_and_group_plans step now also skips any plan whose `blocked_by` is non-empty, reporting it by name with its blocking chain, in addition to (not instead of) the existing has_summary skip rule. Extends tests/fix-2830-halted-plan-dependents.test.cjs (introduced in the prior commit) with direct unit coverage of computeHaltPropagation (including the precomputedOrder call shape) and a fast-check property test — both only possible once this commit's new module exists. Closes #2830 * fix(#2830): surface the halt-aware view from init execute-phase The adopted work made phase-locator compute halted_plans / blocked_by / runnable_plans, but cmdInitExecutePhase builds its output by explicitly enumerating fields, so all three were computed and then silently dropped at the exact consumer the issue names as regressed. Forwards them additively -- incomplete_plans and incomplete_count keep their name, type and semantics byte-for-byte -- and adds the same three empty defaults to the roadmap-only fallback so the shape is consistent in both branches. Covered by a new test that drives the real CLI end to end rather than the locator function, since the locator already worked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2830): fail closed on dependency cycles and stop the templates inviting the defect Three review findings, all fixed: - BLOCKER (isolated adversarial). Cycle participants never reach indegree 0 in the Kahn pass, so they were excluded from the topological order, never visited by the forward pass, and vanished from blocked_by entirely -- i.e. reported as runnable. The wave-grouping reader hard-fails on a cycle so it never hit this, but the phase-location reader does not, so init execute-phase offered a plan depending directly on a halted plan. Reproduced, then fixed in the shared engine so every consumer is safe regardless of pre-checks: a node absent from the order is now blocked with a deterministic, non-empty named cause. A plan silently missing from both blocked_by and runnable is the exact disappearance this issue exists to prevent. - MAJOR (isolated adversarial). All four summary templates showed the field as an inline comment on the value line. Frontmatter parsing does not strip trailing comments, so an executor copying the templates' own presentation wrote a halt that parsed as a non-halted string, silently reproducing the original bug. Guidance moved off the value line, and the halt predicate now tolerates an unquoted trailing comment. - HARD standards violation. A test regex-matched child-process stderr prose for /cycle/i, which CONTRIBUTING bans. Replaced with the structured failure signal plus a differential assertion (same fixture without the cycle edge must succeed), so it stays cycle-specific without matching prose. Also folds the duplicated read-summary-and-check-halted wrapper out of both readers into the shared module -- centralizing only the predicate left the exact two-copies-that-drift pattern the module exists to prevent -- and commits the artifact-types documentation for the new status value. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2830): stop the property generator hanging the whole suite The remote runner did not fail -- it hung. Two containers sat in this file for 31+ minutes, and an earlier attempt ran 9 hours before I killed it. The runner passes --test-timeout=0, so nothing ever reaps it: this would have hung CI indefinitely, not reported a failure. Root cause: the DAG generator built edges by rejection -- from: fc.integer({ min: 0, max: n - 1 }) to: fc.integer({ min: 0, max: n - 1 }) .filter(({ from, to }) => from < to) With n === 1 both integers are forced to 0, so the predicate is unsatisfiable and fast-check retries value generation forever. n is drawn from 1..12 and fast-check biases toward boundary values, so n === 1 is reached almost at once. This also explains why the failing-first run completed normally while the fixed run hung: before the fix the graph module did not exist, so the property test threw on import and never reached generation. It only starts hanging once the code under test works. Generates the DAG by construction instead -- `to` is drawn strictly above `from`, with the degenerate single-node case short-circuited to an empty edge list -- so no rejection sampling is involved. Switches the import to the shared fast-check setup so the seed and run count are pinned per CONTRIBUTING, and adds a bounded regression guard that samples the arbitrary directly, so a future reintroduction fails loudly instead of hanging. Verified: the file now completes in 2 seconds, 29 tests started and 29 finished, zero failures, against an indefinite hang before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2830): restore the depends_on display contract and acknowledge the workflow growth Full-suite run surfaced two things the focused harnesses could not. 1. Regression of a pinned pre-existing contract (#3785). A refactor routed the EMITTED depends_on field through the new dependency resolver, which also consults the canonical-prefix map. The original consulted the plan map only, so a short canonical prefix passed through verbatim -- '24-01' stayed '24-01' rather than becoming '24-01-auth-hardening'. The emitted field is a DISPLAY mapping, not the DAG resolution, and #3785 pins that. Reverted with a comment recording why it must not use the resolver; full resolution is still used for the wave DAG and halt propagation, which is what needs it. 2. The workflow file grew 518 bytes without an acknowledgment, from the halt-aware skip rule and the widened parse contract. Acknowledged. Note on where the acknowledgment landed: the guidance is to add a NEW fragment, but execute-phase.md is already named by an existing fragment and the linter hard-fails when two ack sources name the same path. Appending to the owning fragment, following its own established multi-PR pattern, was the only lint-clean option. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2830): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ff4a57b78c |
chore(#1671): migrate the remaining 13 LARGE/XL workflows to the fragment model — Phase 6.3 (#3030)
* chore(#2994): fragmentize progress.md forensic audit onto the fragment model Extract the --forensic-gated forensic_audit step to workflows/progress/steps/forensic-audit.md behind a section marker, and repair progress.md's init line to forward --forensic so the atom is actually true in production rather than only under direct CLI tests. progress.md shrinks 32630 -> 27207 bytes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize the four manifest-wired workflows new-project, quick, new-milestone and progress each already had a dedicated cmdInit* entry point but zero marked sections. Extract nine gated bodies to workflows/<wf>/steps/ behind section markers and repair each init line to forward its flags. Fold --full into the discuss/research/validate facts inside cmdInitQuick so the when= grammar never sees an OR, per the chunked-mode precedent. Fixes found while working, per the no-defer rule: - cmdInitProgress passed no phase info to buildSectionManifestField, so state:phase-mvp-mode was permanently false — an atom in the vocabulary whose fact could never be computed. - the quick init router folded flag tokens into the free-text description, which the new forwarding would have corrupted. - a #2508 dispatch note was nested inside quick.md's Agent(prompt=) fence, leaking orchestrator guidance into the subagent prompt. - progress.md had a 3-vs-4 backtick outer-fence imbalance. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize verify-work.md and admit state:ui-phase-active Wire cmdInitVerifyWork to buildSectionManifestField — it was a dedicated entry point that never emitted a manifest — and mark two sections. state:ui-phase-active folds (plan:pre hooks include an active ui step) OR (the phase dir holds a *-UI-SPEC.md) into one boolean in init.cts, so the grammar still sees a single operator-free atom. The inner Playwright-MCP check stays as prose inside the fragment: it is live session state and no init seam can precompute it. The MVP false-branch note is a real fallback, not redundant prose, so it sits outside the marker — gating it away would delete the text needed precisely when MVP mode is off. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): follow moved workflow content in drift guards Retarget every guard that asserted on content this branch moved into workflows/<wf>/steps/, mirroring 815b3d897. Each retargeted assertion was verified to still fail when its step file is blanked, so none was weakened into vacuity. Three assertions in verify-mvp-uat were genuinely red. Three more were worse than red — passing for the wrong reason: - quick-commit-boundary and worktree-cleanup anchored on indexOf('Step 5.6'), which matched a later cross-reference and sliced 16069 chars that coincidentally held the asserted substrings. Replaced with an expandWorkflowSections helper that splices step content back in place. - phase6-review-capabilities lost its end boundary and widened to EOF. - playwright-ui-verify matched 'UI' in an unrelated bullet and 'fall back' in a subagent-dispatch line after the real content moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize code-review and complete-milestone, admit three atoms Add dedicated cmdInitCodeReview and cmdInitCompleteMilestone entry points alongside the shared generic ones rather than modifying them — init.phase-op and init.manager carry a CRITICAL blast radius (179 dependents, 24 processes) and stay byte-identical for their other callers. Admit flag:--fix, state:fallow-enabled and state:git-create-tag, each with a consuming section and a fact its own entry point computes. Both sections had the resolver-in-body hazard: the fallow config-gate and the git.create_tag check each sat inside the very block being gated, so gating would have disabled the resolver that decides the gate. Both are hoisted into init and the bodies now consume the resolved fact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): retarget code-review and milestone drift guards, fix two red tests Retarget guards that asserted on content moved into steps/, proving non-vacuity by blanking each step file and confirming failure. Also fixes two genuinely red tests found while working, per the no-defer rule: - workflow-fragments' frozen-vocabulary lock was missing state:ui-phase-active, so commit 7ef7f8336 shipped red. Lint and build both passed over it, which is why neither is sufficient verification. - code-review's quick.md capability-hook assertion carried a stale delimiter after the 18ff35d20 extraction. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize autonomous.md and admit state:plan-strategy-converge Five sections share one atom, the pattern plan-phase already uses for flag:--research-phase. The atom folds --converge OR --cross-ai into a single boolean in cmdInitAutonomous so the grammar stays operator-free. cmdInitAutonomous is additive; init.milestone-op, init.manager and init.phase-op are untouched and still consumed. The $PLAN_STRATEGY bash resolver is deliberately retained — ungated local-planning bullets still read it, so the init-side fact supplements it rather than replacing it. converge-fail-fast required splitting one bash fence so the always-run CONVERGENCE_ARGS construction stays outside the marker. All three flag-absent fallbacks were left outside their markers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize review and discuss-phase-assumptions Admit state:reviewer-instances-configured (two peripheral notes share it; the core reviewer-lane dispatch stays unmarked — it is the workflow's primary always-evaluated logic, not an optional branch) and state:auto-advance-active, which folds --auto OR two config keys into one boolean so the grammar stays operator-free. discuss-phase-assumptions was the highest-risk edit in this PR. Its auto_advance step is a full if/elif/else; gating it whole would have deleted the flag-absent fallback needed exactly when --auto is off. Split verified exact: resolvers 636-651 and the 'End here' fallback 668-669 both stay outside the marker; only 653-667 is gated. Adds emitted-drift acks for the two files that grew — review.md (+55 B) and autonomous.md (+737 B from 80799211c, which had none and would have red-gated the push. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): fragmentize docs-update, update, transition and new-milestone Part A Completes the 13-workflow rollout. Three of these had no init call at all and gained a dedicated entry point plus their first gsd_run query line. Admits state:is-monorepo and adds state:next-channel, state:workstream-active and state:flat-mode. Vocabulary 26 -> 30 atoms. Part A of new-milestone applies when NO workstream is active — the negation of state:workstream-active. Rather than teach the grammar negation, which is the Greenspun drift the frozen list exists to prevent, it gets a separate positively-phrased atom whose fact is the inverse. Part B, which always runs, stays outside the marker. flag:--verify-only is deliberately NOT admitted: docs-update has no contiguous purely-additive region for it, and an atom without a consuming section is dead vocabulary. Evidence recorded in the slice report. update.md reuses its existing resolved $GSD_TOOLS rather than prepending the canonical preamble, which would have clobbered it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): stop automated-ui-verification re-resolving its own gate, retire dead vocabulary Two defects the new tests caught. The automated-ui-verification step re-ran gsd_run loop render-hooks and recomputed UI_PHASE_ACTIVE inside a body that is only read when that fact is already true — the circular self-disabling pattern this design forbids, introduced by 3c654b168. cmdInitVerifyWork now exposes ui_phase_active and the step consumes it. Its launcher preamble goes too: no gsd_run remains. The Playwright-MCP check stays as prose — that is live session state. Dead vocabulary predating this PR: flag:--full and state:needs-codebase-map were admitted with a gate-1 claim that never materialized. flag:--full is removed, redundant once quick folds it into discuss/research/validate. state:needs-codebase-map gets the real consumer it always lacked, gating new-project's codebase-map offer. Vocabulary 30 -> 29, and no atom is now without a consuming section. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): add the atom-admission, inversion and resolver-hoist gates The two existing parity guards prove vocabulary/predicate symmetry but never that a fact is computed — an atom no cmdInit* assembles evaluates false forever. These close that hole: - per-atom satisfiability for all 29 atoms, plus an anti-vacuity assertion so the loop cannot silently cover zero atoms - dead-vocabulary check against the shipped manifest - inversion guard: the flag-absent fallbacks in discuss-phase-assumptions and verify-work must stay outside their markers - data-driven resolver-hoist guard over the shipped manifest, so a future extraction cannot reintroduce the circular class - compound-fold coverage (--full, --cross-ai, --rc, config-only --auto) - null-vs-[] degraded/computed distinction, and flag value shapes Also repairs the frozen-vocabulary lock, which was stale and red for the seven atoms earlier commits on this branch shipped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2994): add changeset for the fragment-model rollout Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(#2994): cite the issue on the two new allow-test-rule exemptions ADR-456 requires an issue ref on the same line as the annotation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2994): correct the atom-count claims after retiring flag:--full The vocabulary doc comments still said 30 entries; it is 29 since flag:--full was removed as dead vocabulary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): dedupe the phase-fallback block and harden --ws parsing Review findings. MAJOR: the three new init entry points each pasted a verbatim copy of the guardedFindPhase/guardedGetRoadmapPhase fallback, taking the repo from four copies to seven — DEFECT.GENERATIVE-FIX. Extracted applyRoadmapFallback and folded six of the seven; each call site keeps its own field-set via a closure. Duplication removed rather than papered over with a parity test. cmdInitPhaseOp stays out: its fallback omits has_reviews, so it is not a byte-identical copy, and it is CRITICAL-radius. LOW, pre-existing: GSD_WS captured [^[:space:]]+ and expands unquoted, so a workstream name holding glob metacharacters would expand against the filesystem. Narrowed to [A-Za-z0-9._-]+. The unquoted expansion is kept — it must word-split into two args and vanish when empty. Also restores the vocabulary ordering convention, and fixes a masked test bug the mandated run surfaced: the flag-forwarding guard checked only the first init line per workflow, but new-milestone has two, so a real failure was reporting exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): drop the stale new-milestone emitted-drift ack new-milestone.md was acked for a +406 B growth measured against an intermediate commit. Net against origin/next it SHRANK by 8 bytes, so nothing needed the ack and it explained nothing — which the differential attribution check reports as a stale acknowledgment, not a pass. update.md's entry stays: it genuinely grew +703 B. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): resolve the 15 failures from the full matrix run All 15 were real and identical on both lanes. REAL REGRESSION: autonomous.md hit 41479 chars against the #2196 guard's 40960 cap — a CHARS cap distinct from the LARGE tier byte cap, which the five section stubs pushed it over. Extracted the 3a.5 UI Design Contract body to references/; now 39968 chars, and the file nets -795 B vs base, so its growth ack is deleted rather than left stale. REAL DEFECT: docs referenced /gsd-transition, which is not a live registered command. Reworded. STALE FIXTURE: the emission byte-identity test hardcoded two marked workflows; this branch legitimately marks fifteen. Fixture corrected — the source was right. The rest were drift guards over the eight workflows the earlier sweep did not cover, retargeted at where the content now lives with non-vacuity proven by blanking each step file and confirming failure. The GSD_WS forwarding guard was checked as a possible real break and is not one: the charclass narrowing is intact and forwarding works end to end. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): drop the ack for a newly-added reference file A new file's emitted ripple is attributable to the diff that adds it, so the acknowledgment explained nothing and the differential check reports it as stale. Removing the last entry removes the fragment — an empty one signals nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2994): retarget the UI-contract guards and clear two transitive advisories The §3a.5 extraction that brought autonomous.md under the #2196 char cap moved its body to references/autonomous-ui-design-contract.md, so ten guards in autonomous-ui-steps and check-ui-safety-gate were asserting it against the host. Retargeted via a combined read, each proven non-vacuous by blanking the reference file and confirming failure. This class had already bitten twice on this branch because each sweep was scoped to the workflows touched at that moment, so this one was exhaustive: ~70 test files across all 13 workflows, zero further broken or vacuous assertions found. Also clears two high transitive advisories the matrix flagged on one lane — fast-uri GHSA-7p8r-x3mc-p8w7 and three ip-address SSRF/trust-boundary issues. Both pre-date this branch: package-lock.json was untouched until now, so the production tree was byte-identical to the base. Lockfile-only, package.json unchanged, verified against a real npm ci install. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2994): backfill changeset pr number to 3030 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ad3b9ec486 |
chore(#1671): fragmentize plan-phase.md and repair flag forwarding to the init bundle — Phase 6.2 (#3019)
* chore(#2993): fragmentize plan-phase.md onto the fragment model Epic #1671 Phase 6.2. plan-phase.md is the largest workflow in the repo and carried zero markers; it was deferred out of the Phase 3 pilot for two reasons, both now dead. The 36-byte PRE_PHASE6 headroom was never the blocker it looked like — fragmentizing is net-negative on host source, so the trim is what creates the room. The --mvp interleaving was resolved by measurement in #2992 and no sub-line mechanism is built. - widen WHEN_VOCABULARY 14 -> 19 via a second coordinated ADR-1671 amendment: flag:--ingest, flag:--prd, flag:--research-phase, flag:--reviews, state:chunked-mode - state:chunked-mode is `--chunked` OR config workflow.plan_chunked, and that disjunction is resolved in the FACT, never in the grammar, so a compound condition never becomes an operator - parse the new flags on the plan-phase route; extract six gated bodies to gsd-core/workflows/plan-phase/steps/ behind manifest-gated stubs - prd-express-path.md was already extracted but read unconditionally; its wrapper is now gated, so the existing extraction finally pays off plan-phase.md 94,483 -> 87,575 bytes (cap 94,519): headroom goes from 36 bytes to 6,944. Also closes a surfaced docs gap: five real plan-phase flags (--chunked, --skip-ui, --bounce, --skip-bounce, --granularity) were documented in neither the argument-hint nor help. Making --chunked load-bearing without fixing its siblings would leave the defect class half-open. Refs #2993 * fix(#2993): forward flags to the init bundle so section gating actually fires Blocker found by the correctness review, confirmed directly, and missed by both the isolated reviewer and every test in this branch. Neither workflow forwarded its flags to the init CLI: plan-phase.md:71 INIT=$(gsd_run query init.plan-phase "$PHASE" $GRAN_PARAM) execute-phase.md:84 INIT=$(gsd_run query init.execute-phase "${PHASE_ARG}") So every flag: atom was permanently false in production and its section permanently excluded. For plan-phase that made the PRD express path UNREACHABLE — a regression, since it was an unconditional read before. For execute-phase this is PRE-EXISTING: #2932 shipped `flag:--wave` gating that has never once been true, so `--wave` silently dropped its own wave-filtering guidance. Fixed here under the no-defer rule. Why every test missed it: they drive the init CLI directly with flags, which works. Production goes through the workflow's bash line, which did not pass them — the exact "assert against the shape production uses" trap this branch's own test matrix warns about. - parse and forward --prd/--ingest/--research-phase/--reviews/--chunked (plan-phase) and --wave (execute-phase), using the anchored regex idiom the neighbouring GRAN_PARAM line already uses - add a regression guard DERIVED FROM THE MANIFEST: for every flag:--X section, the owning workflow's init line must forward --X. It fails against the pre-fix files and covers any future atom, rather than spot-checking today's six. Verified through the workflow shape, not the CLI shape: `3 --prd spec.md` now yields ["prd-express-gate"] (was []), `2 --wave 2` yields ["partial-wave"] (was []). Refs #2993 * test(#2993): acknowledge the execute-phase ripple and regenerate install-tree fixtures Remote matrix was red with 46 unique failures, identical on both lanes. Both causes are mechanical consequences of changing shipped workflow content, and neither is visible to any local gate. - emitted-attribution: execute-phase.md grew 163 bytes from the WAVE_PARAM forwarding fix and was unacknowledged, while the ack fragment named plan-phase.md, which SHRANK and therefore needed no ack at all — a stale entry is itself a failure. The reason now names the real ripple. The entry had to merge into the existing 2930 fragment: the ack linter does unconditional cross-fragment duplicate-key detection with no spent/live exception, so a second fragment declaring execute-phase.md collides even when the first is already merged and inert. Resolved per the linter's own guidance and that file's precedent of appending successive ripple reasons to one entry. - golden-install-tree: tests/fixtures/install-tree/*.json are committed and deliberately excluded from the ADR-2719 attribution cutover, so they must be regenerated when shipped tree content changes. Regenerated after build:lib per the ordering landmine. 19 runtimes each gained exactly the six new plan-phase step files; zero paths removed, which is the absolute failure shape those fixtures exist to catch. Refs #2993 * fix(#2993): restore the launcher preamble in an extracted step and follow moved content in its drift guards Second red run: 26 unique failures, identical on both lanes, in two classes. RUNTIME BUG (runtime-launcher-parity, 7 failures) — chunked-planning-mode.md calls gsd_run but carried no canonical launcher preamble, which is what DEFINES gsd_run(). On any non-Claude runtime that step would fail outright. The preamble is now copied verbatim from the canonical source of truth, gsd-core/workflows/_runtime-launcher.snippet.sh, and the fence dedented to column 0 to match the prd-express-path.md sibling (a list-continuation indent breaks the byte-equal preamble match). prd-express-path.md already had a correct one. This is the same defect #2932 hit when it extracted steps; the parity test caught a real bug, not a stale assertion. DRIFT GUARDS (plan-phase-drift-guard, issue-2762-plan-reviews-chunked, skill-frontmatter-contract) — these assert plan-phase.md contains content this branch moved into step files. Retargeted at where the content now lives, with the asserted property unchanged; the ALL-RUNTIMES label COUNT test now reads host + every step file so the count is preserved across the split rather than reduced. Each retargeted guard was verified to still fail when its step file is stripped, so none was weakened into vacuity. No emitted-drift ack was needed: currentSizes() enumerates gsd-core/workflows/*.md non-recursively, so files under plan-phase/steps/ are never in the size ratchet's scope. Refs #2993 * chore(#2993): backfill changeset pr number to 3019 --------- Co-authored-by: sim <sim@local> |
||
|
|
f1af47766a |
chore(#1671): widen the when= grammar and key the section manifest per workflow — Phase 6.1 (#3013)
* chore(#2992): widen the when= grammar and key the section manifest per workflow Epic #1671 Phase 6.1. Two blockers stopped the fragment model reaching any file beyond execute-phase.md: the when= vocabulary was frozen at 4 atoms (3 execute-phase-specific), and the section manifest was single-workflow by construction with 'execute-phase' hardcoded into buildSectionManifestField. - widen WHEN_VOCABULARY 4 -> 14 via a coordinated ADR-1671 amendment; the grammar stays CLOSED (one atom, no operators, negation or nesting) and WHEN_PREDICATES stays a hand-written literal map, never deriving a predicate from its atom string - InvocationFacts gains flags: ReadonlySet<string> plus three computed state booleans; add the missing reverse vocabulary/predicate parity guard - key the manifest artifact per workflow; a stale flat {sections:[...]} artifact now fails shape validation instead of being misattributed - wire the field into six init entry points and parse the flags each needs An atom ships only with both a real consuming section and a fact the init seam actually computes. Six surveyed atoms are withheld because their workflows have no dedicated init entry point; an atom without a computed fact evaluates false forever and silently disables its own section. Fixes a defect found while wiring: parseNamedArgs always materializes a boolean flag key, so folding its false into the absent sentinel is required or every flag reads as present and gating is silently always-on. Also resolves ADR-1671:194 by measurement: --mvp stays unmarkable, because its interleaved sites are always-run flag resolution and a ~340 byte block that already delegates lazily. Refs #2992 * fix(#2992): treat any falsy option value as an absent flag and reject unsafe manifest read paths Findings from two orthogonal reviews (Claude /code-review + an isolated adversarial pass); both independently reproduced the first one. - MAJOR: the flags-builder treated only `undefined` as absent, but parseNamedArgs yields `null` for an absent value-flag and `false` for an absent boolean-flag, so `--granularity` read as present on every plan-phase invocation. Fixed at the root: a flag is present iff its option value is truthy. The six per-handler `|| undefined` folds are now redundant and removed, which also closes the duplicate-translation and missed-onboard-handler findings. - MAJOR: state:needs-codebase-map had zero coverage. Added unit, property and real-CLI integration tests. - MINOR: reject absolute, UNC/drive and `..`-traversing `read` paths in the manifest, degrading the whole load to null like every other shape violation. Verified: `/etc/passwd` previously reached section_manifest.read. - MINOR: corrected a stale "4 to 20" doc comment; the vocabulary is 14. Refs #2992 * test(#2992): update the generator suite for the per-workflow manifest shape The remote matrix went red with 5 unique failures, identical on linux-node22 and linux-node24, all in tests/gen-section-manifest.test.cjs. Re-keying the artifact to {workflows:{...}} left this suite asserting the old flat {sections:[...]} shape; nothing else in the tree still does. - three tests read manifest.sections.length, now undefined; retargeted at workflows.<name> with their original intent preserved (a fenced or loop-host marker still asserts NO section is produced, not merely a changed count) - the stale-manifest test wrote its fixture in the OLD shape, so it tripped shape validation and stopped exercising staleness at all. Its fixture is now valid-but-mismatched so FAIL_STALE is genuinely reached again. - added the coverage that exposed: a pre-6.1 flat artifact must report FAIL_MANIFEST_MALFORMED_SHAPE. That is the real upgrade path for an installed tree and nothing covered it. Refs #2992 * chore(#2992): backfill changeset pr number to 3013 --------- Co-authored-by: sim <sim@local> |
||
|
|
a987cf2731 |
chore(#2932): emit a per-invocation section manifest from the init bundle (#2987)
* chore(#2932): emit a per-invocation section manifest from init Extends the init bundle with a typed per-invocation section manifest so an invocation loads only the branch guidance it will actually take. The three flag/state-gated branches in execute-phase.md move into their own step files; the parent keeps its gsd:section markers wrapping a one-line on-demand reference, so each section's prose lives in exactly one file and the parent shrinks 93369 -> 89507 bytes. A new drift-guarded generator derives the shipped section manifest from those markers, and a new pure evaluator maps invocation facts to applicable section ids. The evaluator is a lookup over the frozen WHEN_VOCABULARY, never a parser (Greenspun's Tenth Rule, ADR-1671:69); a parity test asserts the vocabulary and the predicate map stay exhaustively in sync. Closes #2932 * fix(#2932): fail closed on prototype-chain when values An isolated adversarial review found WHEN_PREDICATES[section.when] was a bracket lookup on a plain-prototype object, so inherited Object.prototype members resolved as predicates: "constructor"/"toString"/"valueOf"/ "hasOwnProperty" returned truthy and SILENTLY INCLUDED the section, and "__proto__" threw an untyped TypeError carrying no .reason. Both violate the module's documented fail-closed contract, and the manifest is read from disk at run time so it cannot be assumed trustworthy. Builds the predicate map on a null prototype and guards the lookup with an explicit Object.hasOwn check. Adds table-driven coverage for nine Object.prototype-shaped keys asserting the TYPED reason (asserting only that it throws would still pass while broken) plus a fast-check property injecting a hostile value at an arbitrary document position. * test(#2932): retarget execute-phase step assertions at extracted step files * fix(#2932): emit typed reasons for generator lib-load and write failures * fix(#2932): restore launcher preamble in extracted steps and refresh derived fixtures * chore(#2932): backfill changeset pr number to 2987 --------- Co-authored-by: sim <sim@local> |
||
|
|
000a322489 |
fix(#2941): hint at global: prefix when bare skill name matches a global skill (#2973)
* test(#2941): add regression for bare-skill-name global: hint When a bare skill name matches an existing global skill, the skip warning must hint at the global: prefix. When no global skill matches, the original 'Skill not found' warning is unchanged. * fix(#2941): hint at global: prefix when a bare skill name matches a global skill When buildAgentSkillsBlock skips a bare skill name that doesn't exist as a project-relative path, check if it matches an existing global skill. If so, append a hint to the warning: 'a global skill named X exists; use global:X to reference it'. When no global skill matches, the warning is unchanged. getGlobalSkillDir and getGlobalSkillsBase are already imported in this module for the global: branch. The hint is guarded on globalSkillsBase being non-null (runtimes without a skills directory don't support the prefix). * chore(#2941): add changeset fragment * chore(#2941): backfill changeset PR number 2973 --------- Co-authored-by: sim <sim@local> |
||
|
|
f0ff23635e |
fix(#2602): discover project-local Codex agents (#2623)
* fix(#2602): discover project-local Codex agents - Select an existing local Codex agents directory before global fallback - Prove init reports the canonical local installation through compiled CJS * test(#2602): lock Codex agent precedence - Cover override, local authority, global fallback, and runtime compatibility - Exercise installed state through the compiled resolver * fix(#2602): resolve local Codex agent skills - Pass the canonical project root to the non-Claude persona fallback - Cover nested-Codex fallback and Claude compatibility through the CLI * test(#2602): cover local Codex validation status - Assert emitted validate and health commands use the project-local install - Preserve empty local-directory authority beside complete global agents * fix(#2602): align validation with local Codex discovery - Pass the resolved runtime and project root to health W010 - Resolve the validate-agents runtime before checking installation status * test(#2602): cover local Codex docs status - Assert docs-init reports an authoritative empty local install as unhealthy * fix(#2602): align docs with local Codex discovery - Pass the resolved runtime and canonical project root to the shared agent checker * fix(#2602): honor agent-skills runtime override - Resolve agent-skills fallback runtime through the canonical project resolver - Cover conflicting config and GSD_RUNTIME values through the emitted CLI * fix(#2602): ignore non-directory local agents paths - Treat only a local Codex agents directory as authoritative - Cover regular-file fallback through the emitted install checker * chore(#2602): add changelog fragment - record the user-visible local Codex agent discovery fix for PR #2623 * fix(#2602): align local agent discovery with runtime policy - Resolve Codex's local config directory through the canonical runtime policy - Use test-managed cleanup for local-agent discovery coverage * fix(#2602): discover local agents across runtimes - Prefer manifest-backed project-local installs for non-Claude runtimes - Respect runtime-specific local install roots and preserve global fallback behavior - Cover native, partial, cross-runtime, and project-root local discovery * fix(#2602): preserve agent discovery fallback - Fall back globally when local-install probes fail - Document and test symlink rejection - Align the changeset with repository format * fix(#2602): reuse local directory policy - Resolve runtimes without local config through the canonical sentinel - Document the manifest gate and refresh the context index --------- Co-authored-by: Daniel E. <daniel.e@teachingstrategies.com> Co-authored-by: Rezolv <dave@sienkowski.com> |
||
|
|
9a76ca6783 |
fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter
extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.
Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.
The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.
Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (
|
||
|
|
27c2279a39 |
fix(#2617): project verification next_command onto the runtime's command surface (#2700)
* fix(#2617): project verification next_command onto the runtime's command surface `src/verification.cts` stored and synthesized hard-coded `/gsd:…` command strings with no runtime context, and `phase complete` relayed that raw field straight into its verification-blocked error. On a Codex project the suggested next step was `/gsd:execute-phase`, a surface Codex does not install — it installs `$gsd-execute-phase`. The colon form is wrong twice over: `runtime-slash.cts` documents that "the colon form is never emitted", so EVERY runtime — not just Codex — was being handed a deprecated shape. Fixed at the one routing seam rather than per caller: - The routing table now stores BARE command names (`execute-phase`), never a prefixed literal. A prefixed literal in the table is what leaked. - A single `projectNextCommand(bare, runtime, tail)` helper runs every return path through `formatGsdSlash`, preserving the argument tail (`01 --gaps`) untouched. An empty command stays empty, so "no next step" never becomes a bare prefix. - `readVerificationStatus` accepts `opts.runtime`; `cmdVerificationStatus` and `phase complete` pass `resolveRuntime(cwd)`. The default is `claude`, which yields the canonical `/gsd-` hyphen form. All four routed states are covered: missing, unknown, gaps_found, stale. `init.cts` keeps its own projector deliberately. It already formats correctly, and its command CONTENT differs from the router's on purpose (it appends the phase number to `execute-phase`, and routes `human_needed` to `verify-work`). Consolidating them would silently change `init`'s user-visible output, which this issue did not ask for — so the divergence is left intact and the new tests instead pin the property that matters on both surfaces: no raw colon form escapes. Failing-first record: `origin/next:src/verification.cts` carried the four `/gsd:` literals (lines 101, 108, 382, 392), and 11 existing assertions in tests/verification-status.test.cjs asserted the colon form. Those 11 are corrected in this commit — they passed before the fix and fail after it, which is precisely the regression this closes. Tests are folded into the module's primary suite rather than added as a third file (`lint-test-file-count` caps the `verification` module at two, and consolidating is its documented remedy — growing the allowlist is not). The `phase complete` assertion reads `res.error`, not `res.stderr`: `runGsdTools` exposes a clean non-zero exit's stderr as `error`, and reading the wrong field yields '' and makes the whole check vacuous — which is how this user-visible path stayed untested. Closes #2617 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2617): scope the new hooks to their describes; cover gaps_found through the CLI Two findings from the orthogonal review of the first commit, both in the tests this change added. 1. The folded block's `beforeEach`/`afterEach` were declared at MODULE scope. node:test applies module-scope hooks to every test in the file, so hooks added for the #2617 suites also wrapped the ~40 pre-existing tests in verification-status.test.cjs — making an unrelated block a single point of failure for them (currently benign, but a throwing hook would have failed suites it has nothing to do with). They now install inside their own describes via a small `useProjectionPhaseDir()` helper, with a comment recording why. 2. The live-CLI `phase complete` test exercised only the `missing` state, so a regression in any other routed branch would have shown up in the router's return object but not in the text a user actually reads. Added a `gaps_found` case per runtime, asserting the projected `plan-phase <N> --gaps` reaches the blocked-completion error. Whole file verified green: 48 tests, 48 pass — the ~40 pre-existing ones included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * test(#2617): correct the last colon-form assertion in phase.test.cjs The remote run surfaced one more stale assertion outside tests/verification-status.test.cjs: the `phase complete` canonical-gate suite matched the blocked-completion message against `/\/gsd:verify-work 0?1/`. That project fixture configures no runtime, so it takes the `claude` default, which now yields the canonical `/gsd-verify-work 01` hyphen form. The colon form this asserted is exactly the deprecated shape #2617 removes — `runtime-slash.cts` documents that "the colon form is never emitted". Like the eleven corrected in the first commit, this assertion passed before the fix and fails after it, which is the regression record rather than a test being loosened: the surrounding assertions (failure reason, `stale` wording, and that neither ROADMAP.md nor STATE.md was mutated) are untouched. Verified against the real CLI: the emitted message is now "Phase 1 verification is incomplete: Verification is stale. Re-run verify-work before transition. Next: /gsd-verify-work 01". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * fix(#2617): collapse the two verification projectors into one seam The orthogonal review found that `init.cts` carried a second, independently maintained `verificationNextCommand()` that had drifted from the router's table in CONTENT, not just formatting: state router (before) init.cts missing execute-phase execute-phase <N> unknown execute-phase execute-phase <N> human_needed "" (no command) verify-work <N> The `human_needed` row is the sharp one: two GSD surfaces disagreed about whether a next command existed at all, and the router's own next_action told the user to "re-run the verify step until status is passed" while naming no command to run. init's answers were the useful ones, so the router adopts them and init now delegates to it — satisfying the issue's "keep one verification-routing seam" direction. `verificationNextCommand()` is deleted. Appending the phase number surfaced a trap the old bare commands hid. `extractPhaseToken` also returns project-code forms (`PROJ-07`), which are indistinguishable by shape from an ordinary directory name — `gsd-651-parent` yields `gsd-651` — so deriving the argument blindly emits `execute-phase gsd-651`. The number is therefore appended only when it is unambiguously numeric, or when the caller supplies it explicitly. `init` does supply it: its `phaseDir` is unresolved in several branches, where the router could not derive one at all. dir `01-example` -> $gsd-execute-phase 01, $gsd-verify-work 01 dir `gsd-651-parent` -> $gsd-execute-phase, $gsd-verify-work Suites verified green against the built lib: verification-status 50/50, phase 268/268, init 143/143, init-manager 40/40. `npm run lint:ci` clean. User-visible change beyond the reported bug, as agreed: `query verification.status` and `phase complete` now append the phase number for missing/unknown, and emit `verify-work <N>` for human_needed where they previously emitted nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QT3ibz5qJuDuGqpTGRYVGf * chore(#2617): backfill changeset PR number (#2700) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
936a345381 |
feat(#2505): Phase 3 — agent-skills fallback for non-dispatchable runtimes (#2521)
* feat(#2454): PR 2 — cmdAgentSkills fallback reads installed agent prompt When no agent_skills config entry exists for a given agent type (the common case on AGENTS-native runtimes), cmdAgentSkills previously returned empty output. Workflows that inject ${AGENT_SKILLS_*} into subagent dispatch prompts then carried nothing — the persona was lost. The fallback: resolve the runtime's agents directory via checkAgentsInstalled and read <agentsDir>/<agentType>.md. The installed agent prompt content (now present for kimi-code via the flat-skills install layout) flows into the dispatch prompt so the persona survives even without explicit config opt-in. This is the reporter's suggested fix #2 from #2454. The fallback triggers for ALL runtimes (not just kimi-code) when no config entry exists — it is strictly additive (returns content the previous empty path could not). If the agent file is not found on disk, the block stays empty (same as before). * docs(changeset): Phase 3 agent-skills fallback Added (#2510) * docs(changeset): backfill PR #2521 for Phase 3 (#2510) |
||
|
|
b6e6a22fce |
fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer (#2457)
* fix(#2402): honor response_language across orchestrator output + UAT checkpoint renderer Replays the in-flight bot branch fix/2402-response-language-orchestrator-coverage (seven commits, never pushed) onto current origin/next as a single squashed commit. The original work was substantial and correct; this commit preserves its full scope, trimmed where rebase conflicts + workflow size budgets required it. Three independent layers where response_language was being dropped are closed: Layer 1 — orchestrator-facing directives across workflows. Adds the strong "All user-facing output in this workflow MUST be presented in {response_language}; technical terms, code, paths, and subagent prompts stay in English" directive to ~40 workflows that previously either lacked it entirely (verify-work, new-project, new-milestone, quick, manager, and ~35 more) or carried only the weak subagent-prompt-only form (plan-phase, execute-phase). The directive covers narration between tool calls and banner output, not just the AskUserQuestion prompts. Layer 2 — UAT checkpoint renderer (src/uat.cts). buildCheckpoint now accepts an optional responseLanguage parameter and renders the frame strings ("CHECKPOINT: Verification Required", "Type `pass` or describe what's wrong.") in any of 9 languages (English/Spanish/French/German/Portuguese/Japanese/ Chinese/Korean/Italian) with an alias table covering ~30 input variants (en, es, español, ja, 日本語, etc.). cmdRenderCheckpoint reads config.response_language via loadConfig(cwd) and passes it through, so the byte-for-byte block verify-work.md reprints verbatim is already localized when written — preserving the anti-injection hygiene rule at verify-work.md (the model is forbidden to translate after the fact). CJK display width is computed by East Asian Width property ranges (W/F) so the right ║ border of the banner stays aligned for full-width characters. English fallback is byte-identical to the pre-fix behavior when response_language is unset or unrecognized. Layer 3 — literal English report templates in execute-phase. The top-of- workflow directive covers all template sites (templates are a structural source, not literal output). Inline render-language notes that previously sat at each template site were removed during the squash because they pushed execute-phase.md over its frozen pre-phase-6 byte ceiling (93600 — ADR-857 Phase 6 capstone). The single top directive covers the same surface with fewer bytes. Also extends src/docs.cts and src/init.cts to propagate response_language into the init JSON bundle of the additional workflows so the directive can read it. Tests added: - tests/uat.test.cjs: buildCheckpoint with unset/unrecognized language falls back to English default; recognized language swaps only the two frame strings while structural lines stay untouched; CJK display-width regression (independent recomputation of East Asian Width W/F ranges). - tests/workspace.test.cjs, tests/docs-update.test.cjs: response_language wiring through docs.cts/init.cts. References: #2402; reporter's three-layer triage + Layer-4 follow-up; the byte-for-byte anti-injection hygiene rule at verify-work.md (the reason Layer 2 must be renderer-side, not model-translated). This is a squash of the in-flight bot branch — seven commits representing the original implementation plus its subsequent fix/CJK-padding/test/ changeset/regen cycles, none of which were ever pushed or PR'd. The squash captures the final coherent state. * chore(#2402): backfill pr:2457 in .changeset/2402-response-language-orchestrator-coverage.md * chore(#2402): regen golden + size baseline after rebase against #2315 (PR #2451) Rebase conflicts were entirely in generated artifacts (golden-install-parity fixtures + workflow-size-baseline.json). After taking theirs during rebase, regenerated cleanly against the merged source tree. |
||
|
|
1a46bc068a |
fix(#2376): emit absolute subagent-facing paths from init/state, convert workflow literals (#2428)
* fix(#2376): emit absolute subagent-facing init/state paths Make init.* and state.* path fields absolute rather than cwd-relative so subagent prompts resolve correctly regardless of working directory. Adds intel_dir/conflicts_path/requirements_path/roadmap_path/state_path to cmdInitIngestDocs, an absolute debug_dir to cmdStateLoad, and replaces bare .planning/... literals in 12 workflow Agent() prompt blocks with the absolute init-JSON path fields. Includes decoy-cwd regression tests and realpath'd tmpdir fixtures for macOS. Squashed rebase of the #2376 commit series onto a fresh origin/next (previous merge ee25543a1 was against a now-stale next). * chore(#2376): add changeset * chore(#2376): regenerate golden fixtures + workflow size baseline Regenerated after rebasing the absolute-path fix onto current next (picks up #2351's run-with-timeout content in execute-phase.md too). * fix(#2376): trim execute-phase.md redundancy to stay under the size margin * chore(#2376): regenerate golden/size baseline after rebase onto next |
||
|
|
b0f672f88c |
fix(#2337): capture and surface todo severity (#2381)
add-todo.md gains a confirm-based infer_severity step (infer from the blocker/major/minor/cosmetic taxonomy, confirm via AskUserQuestion with TEXT_MODE fallback, before writing) and a severity frontmatter field. cmdListTodos and cmdInitTodos now surface severity, backward-compatible (key omitted when absent), in parity. Closes #2337. Admin-merged (self-review bypass) with full green CI. |
||
|
|
5f609762ed |
fix(#2237): fail loud on ambiguous bare-number phase directory collision (#2262)
* fix(#2237): fail loud on ambiguous bare-number phase directory collision When two unrelated projects share a .planning/phases/ tree, a bare phase number silently resolved to the first 0N-* directory found — risking cross-project file writes. The fix detects multiple matches for the same phase number and surfaces an ambiguous_matches result instead of silently taking the first. Changes: - src/phase-locator.cts: searchPhaseInDir uses filter() + ambiguity check - src/phase.cts: cmdFindPhase same pattern - src/init.cts: cmdInitPhaseOp surfaces ambiguous_matches in the result - tests/phase-locator.test.cjs: 3 regression tests * docs: backfill changeset PR number (#2262) * merge: keep up to date with next * fix: regenerate stale capability-registry after next merge |
||
|
|
e0f969af6a |
refactor(#2246): centralize cross-platform path-separator handling (toPosixPath / toNativePath / posixNormalize) (#2247)
Replace every open-coded separator translation across the installer/hooks
source with named, tested seams in shell-command-projection.cts (the platform
seam), removing all hardcoded `/`+`\` from path handling:
- toPosixPath(p) — this machine's native path → POSIX (running-OS relative;
for local filesystem paths).
- toNativePath(p) — POSIX → native (collapses the win32 `/\//g,'\\'` ternary).
- posixNormalize(p)— unconditional `\`→`/`, OS-independent; for emitting paths
to a POSIX/bash TARGET (which may differ from the running
OS) and for parsing mixed-separator input.
core-utils.toPosixPath now delegates to the seam, so its 20+ existing consumers
resolve to one implementation; no duplicate helper.
- ~47 sites across runtime-hooks-surface, runtime-artifact-conversion,
runtime-artifact-install-plan, drift, init, worktree-safety,
installer-migrations, installer-migration-authoring, install-engine, surface,
verify, runtime-artifact-layout, schema-detect, check-command-router.
- Closes the latent POSIX-literal-backslash corruption class (the regex form
corrupts a POSIX path containing a literal backslash; split(path.sep) does not).
- New unit + fast-check property tests for all three helpers.
Closes #2246
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
58b20c31f8 |
fix(#2136): migrate ALL operator-facing date sites to localToday (anti-pattern elimination)
The UTC-slice anti-pattern (deriving a calendar day a human reads by slicing a
UTC instant) remained in several operator-facing sites beyond the original
seamed last_activity set. Eliminate it everywhere a human reads the value as a
calendar day — do not leave known bad code in place:
- commands.cts cmdTodoComplete + cmdScaffold: completion/scaffold dates.
- gsd2-import.cts: migrated STATE.md 'Last activity' / 'Last session'.
- template.cts: plan frontmatter 'completed:' date.
- verify.cts: health --repair session-log date + '(Backfilled: <date>)' header.
- workstream.cts: workstream-create 'Last Activity' / 'created' + archive dirname.
- init.cts: JSON-bundle 'date' (-> localToday) + 'timestamp' (-> nowIso) at all
three sites; drop the now-dead 'const now = new Date()' in cmdInitTodos /
cmdInitMapCodebase (cmdInitQuick keeps it for the local branch-id derivation).
- state.cts: prune-archive '## Pruned <date>' header.
- state-transition.cts: the 7 seamed last_activity writes (prior commit).
No realClock.today() / .clock.today() / raw new Date().toISOString().split('T')
operator-facing sites remain in src/. Rebuilds the tracked
bin/lib/state-transition.cjs artifact to match.
|
||
|
|
f910352ea6 |
fix(2104): guard foreign-prefix collapse in init execute-phase/verify-work/phase-op
#2056 fixed the foreign-prefix collapse for init plan-phase only. The identical defect remained in three sibling commands that called findPhaseInternal/getRoadmapPhaseInternal without the guard: - cmdInitExecutePhase - cmdInitVerifyWork - cmdInitPhaseOp Extracted the #2056 guard into shared helpers (guardedFindPhase / guardedGetRoadmapPhase) and routed all four init commands through them. Deleted the local parsePhasePrefix/isForeignPrefixedPhaseQuery copies in init.cts — the canonical export from phase-id.cts is now used directly, eliminating the drift risk flagged by both reviewers. Added 5 regression tests (3 reject + 2 accept-branch) mirroring the #2056 plan-phase tests. |
||
|
|
7866e22457 |
Merge remote-tracking branch 'origin/next' into fix/2056-plan-phase-foreign-prefix
# Conflicts: # src/init.cts |
||
|
|
c1cd43a39f |
fix(#2128): bound the phase-tag clause to {0,200} — kill quadratic ReDoS
The canonical OPTIONAL_PHASE_TAG_SOURCE tag clause `(?:\s*\([^)\n]*\))?` (and its
inlined literal mirrors across 11 modules) had an UNBOUNDED body, making the
optional-group + /g header scan quadratic on adversarial ROADMAP.md/STATE.md — a
long run of `(` after a header ran ~18.8s at 1.7MB. Bound the body to {0,200} in
the constant AND every mirror in lockstep (the #1729 "both forms change together"
contract), so the scan is linear: the same 1.7MB input now resolves in ~9ms
(measured), while real tags (a handful of chars) still match and a 201-char tag
is rejected. Added a #2128 boundary regression to the #1729 suite.
Pre-existing (byte-identical before/after the Phase 4 migrations); folded in at
maintainer direction rather than deferred.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
e2eaa5b046 |
fix(#2128): address review — migrate 9 mis-allowlisted sites, harden scanner + guards
Correctness review of the Phase 4 guard found the allowlist over-broad and the scanner/guards evadable. Fixed all findings: - Migrate 9 sites that were wrongly sanctioned: their regex is the PURE canonical token (`\d+[A-Z]?(?:\.\d+)*`, no variant), byte-identical to already-migrated siblings. The old justification argued against swapping to the extractPhaseToken() FUNCTION (behavior-risky) — but the guard only wants the same regex built from the SOURCE string (byte-equal, zero risk). Coverage is now 32 migrated / 5 sanctioned, not the overstated 23 / 14 (audit.cts x3, uat.cts, init.cts x4, roadmap-upgrade.cts). Each conversion proven byte-equal (.source + .flags). - Harden the drift detector: also catch the `[0-9]`-in-place-of-`\d` variant; document the accepted limits (cross-line split, semantic restructuring — covered by the identity guard + review, not a text scan). - Sanction robustness: a `phase-id-owner:` marker now counts only inside a `//` comment (a bare substring in a string no longer suppresses a real flag), and the preceding-line window skips blank lines (an auto-formatter's blank line no longer reactivates the flag). - roadmap-parser.cts:462 comment: corrected — that regex carries no /i flag, so its [A-Za-z] class does real case work (matches state.cts:1409's rationale). - Identity guard: surface require failures instead of silently skipping, and floor coverage at >75% of consumer modules (inspects 156/157). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dfad3a7510 |
refactor(#2128): single-source 23 phase-token re-derivations; sanction 14 context-specific sites
Route 23 literal re-derivations of the canonical phase-number token through phase-id.cjs `PHASE_NUMBER_TOKEN_SOURCE` (via new RegExp). Each conversion was proven BYTE-IDENTICAL (old.source === new.source && old.flags === new.flags), so the runtime regexes are unchanged — zero behavior change by construction. The remaining 14 phase-token sites are genuine but context-specific and stay literal with a `// phase-id-owner: <reason>` sanction: dir-name parses whose dash-continuation semantics differ from extractPhaseToken, and the [A-Za-z] case-variant / [.-] dot-or-dash separator forms that are not source-byte-equal to the canonical token. Scanner (`npm run check:phase-id-drift`) is now green. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9d9139b8b1 | Merge branch 'next' into fix/2056-plan-phase-foreign-prefix | ||
|
|
6e773d97df |
feat(#2088): migrate Codex onto the Embeddable Orchestration System (ADR-1239)
Drive Codex install/uninstall through the descriptor-driven Host-Integration Interface (declarative embedding adapter → engine surface dispatch) and fold every positive `runtime === 'codex'` / `isCodex` projection into descriptor-driven `runtime.hostBehaviors`. Install/uninstall output stays byte-parity-gated (tests/fixtures/golden-install-parity/codex.json); no other runtime changes. Three Context7-verified upgrades, each with a test on the user-reachable surface: - Skill root → canonical $HOME/.agents/skills via a skills-kind `home` override, with pre-move migration cleanup (stale ~/.codex/skills/gsd-* removed on install and uninstall; user content preserved). Fixes getGlobalSkillsBase, writeManifest, and the skill-manifest inventory to honor the override so --skills-root / sync-skills / the manifest report the real location. - Six new hooks.json lifecycle events (PreToolUse, PermissionRequest, PreCompact, PostCompact, SubagentStop, UserPromptSubmit) shared by install + uninstall; extendedHookEvents reconciled [] -> the schema-valid wired subset. - Explicit `[agents] max_depth = 1` in the managed config.toml block, pinning the negotiated dispatch.maxDepth:1 axis. validateCodexConfigSchema now permits a known-scalar-only bare `[agents]` AgentsToml table (still rejects [[agents]] and unknown-key break-forms, #2760); mergeCodexConfig preserves the user's own AgentsToml scalars (max_threads etc.) instead of dropping them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |