Commit Graph

5 Commits

Author SHA1 Message Date
Tom Boucher
929e02cb2c enhance(#3885): no silent swallow, and no verdict manufactured from dropped data (#3925)
* test(#3885): failing-first coverage for the depth bound and the manufactured wave verdict

ADR-3473 §8.5 says a swallowed failure may not become an authoritative-looking
answer. Three families do exactly that today; this commit pins each one RED.

Measured on this tree, 2026-08-27:

  intel query, .planning/intel/file-roles.json nested 12000 deep
    -> exit 1, "Error: Maximum call stack size exceeded"
       searchJsonEntries / matchesInValue carry no depth parameter at all.
       The MAX_JSON_SEARCH_DEPTH = 48 bound existed in the retired SDK lineage
       (sdk/src/query/intel.ts at 11918dcc3^) and the surviving .cts lineage
       never received it.

  same fixture nested 48 and 49 deep
    -> both return total=1 at exit 0, truncated=undefined
       Nothing distinguishes "searched to the bottom" from "stopped looking".

  query phase-plan-index, a plan whose depends_on names an unresolvable token
    -> warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it
                  in wave 1"]
       The token is never mentioned. computeDependencyLevels drops the edge
       with `if (!resolvedDep) continue;`, every plan becomes a root, and the
       tool then reports the author's correct wave: as the thing that is wrong.

  countPhasePlansAndSummaries with fs.readdirSync throwing EACCES
    -> hasContext:false, indistinguishable from a phase that simply has no
       CONTEXT.md. context_read_error is undefined.

The shapes these tests assert against, chosen here so the implementation has a
target rather than inventing one later: `truncated: boolean` on the intel query
result, `unresolved: Array<{plan, token}>` from computeDependencyLevels, and
`context_read_error: string | null` per analyzed phase.

Deliberately green, and they must stay that way — each stops the fix from
over-firing:

  depth 48 is found and NOT flagged truncated (the ceiling is inclusive)
  a shallow miss reports no truncation           (noise control, N1)
  10,000 siblings at depth 2 are unaffected      (the bound is DEPTH, N2)
  a genuine wave: mismatch on a fully-resolved DAG still warns (N3)
  a genuinely missing directory is absent, not an error
  the emitted depends_on display mapping still passes an unresolved token
    through verbatim — already pinned by the existing #3785 test, so no
    duplicate was added

T31 asserts at the consumer's output per ADR-3180 Decision 4(b): it runs the
real CLI and reads the emitted JSON, because a unit assertion on
computeDependencyLevels would have passed throughout #3427's life.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* enhance(#3885): no silent swallow, and no verdict manufactured from dropped data

Implements ADR-3473 §8.5. A failure or a gap in the input stops being absorbed
into an output that reads as authoritative.

The recursion bound, restored but NOT verbatim (src/intel.cts)

  MAX_JSON_SEARCH_DEPTH = 48 is threaded through searchJsonEntries and
  matchesInValue, which carried no depth parameter at all. The bound existed in
  the retired SDK lineage (sdk/src/query/intel.ts at 11918dcc3^) and the
  surviving .cts lineage never received it — §8.3's "a consolidation may not
  delete an invariant along with the surface that held it", demonstrated.

  Measured before: a .planning/intel file nested 12000 deep exits 1 with
  "Error: Maximum call stack size exceeded". Reachable from a project document.

  The original returned a bare `false` at the ceiling. Restoring that verbatim
  would trade a crash for a silent "no match" when the truth is "I stopped
  looking" — the same class this epic exists to close, and ADR-3473 Decision 4
  forbids it. So the bound carries a truncation signal:

    nesting 47 -> found,     truncated false
    nesting 48 -> found,     truncated false      (the ceiling is inclusive)
    nesting 49 -> not found, truncated TRUE
    nesting 12000 -> exit 0, truncated TRUE, no RangeError

  A shallow document that simply has no match reports truncated FALSE — the
  flag means "I stopped early", never "I found nothing", or it would be noise.
  The bound is on DEPTH: 10,000 siblings at depth 2 are unaffected.

The dropped edge is named, and stops being blamed on the author (src/phase.cts)

  computeDependencyLevels dropped every unresolvable depends_on token with a
  bare `continue`. Each drop makes a plan a root, so the whole phase collapses
  to wave 1 — and cmdPhasePlanIndex then reported the author's CORRECT wave: as
  the thing that was wrong.

  Before:
    warnings: ["Plan 03-02: declared wave: 2 but depends_on DAG places it in
                wave 1"]
  After:
    warnings: ["Plan 03-02: depends_on token \"nonexistent-token-3427\" does not
                resolve to any plan in this phase — edge dropped, wave placement
                for this plan may be unreliable"]

  The suppression is PER PLAN, never blanket: a plan with a fully-resolved DAG
  and a genuinely wrong wave: still gets the mismatch warning. resolveDependencyId
  stays two-tier — the shortFormToId third tier is §8.3/Phase 6's rule and is
  deliberately not built here. The emitted depends_on display mapping still
  passes an unresolved token through verbatim (#3785).

No artifact from failed inputs (gsd-core/workflows/review.md, #3352)

  A failed lane leaves no result file, so "every lane failed" is exactly "the
  aggregate JSONL has zero lines" — the gate condition already existed as a
  byproduct. REVIEWS.md is no longer written in that case, and the commit step
  is skipped with it. A budget-SKIPPED lane also leaves no file and is NOT
  counted as a failure. Per-lane output and non-empty .err are preserved to
  .review-diagnostics/ before `rm -rf "{run_dir}"` destroys the only record that
  the lanes failed at all; the commit step names one file, never a glob, so the
  diagnostics are not swept in.

Unreadable is not absent (roadmap.cts, gap-checker.cts, init.cts x2)

  Four callers collapsed an EACCES on a phase directory into [] and reported
  hasContext:false — byte-identical to a phase that simply has no CONTEXT.md.
  Each now names the directory it could not read. A genuinely missing directory
  stays absent rather than becoming an error, which is what keeps the fix from
  over-firing.

Fatal errno folded into a retry set: audited, no defect found

  Reported as a verified negative rather than padded with a change.
  withPlanningLock was fixed by #1884/PR #3472; acquireStateLock by #3776;
  atomicRenameWithRetry and estimate-cli's renameWithRetry are correct by
  construction — bounded set {EPERM,EBUSY,EACCES}, bounded attempts, and they
  return or rethrow the final error rather than swallowing it. estimate-cli's
  sole caller surfaces that rethrow as write_error in its JSON output.
  Manufacturing a diff to make the checkbox look worked-on is the Goodhart
  outcome Decision 6 exists to prevent.

Disclosed: R46 (the commit step names one file, never a glob) is a real
regression guard but is NOT independently failing-first — the commit fence is
byte-identical pre- and post-fix, so it only fails pre-fix through its shared
extraction dependency. Recorded rather than claimed as fail-first.

Design:      .gsd/phase/feat-3885-no-silent-swallow/40-design.md
Test matrix: .gsd/phase/feat-3885-no-silent-swallow/50-test-matrix.md

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): escape untrusted tokens, and stop cleanup destroying unpreserved evidence

Two review findings, both real, both in my own change.

An isolated adversarial review found the evidence-preservation block never
checked mkdir/cp exit status while `rm -rf "{run_dir}"` ran unconditionally in
a SEPARATE fenced block. A disk-full or unwritable phase directory therefore
still destroyed the only copy of the failed lanes' output — reintroducing the
exact #3352 data loss this item exists to stop, inside the fix for it.

Preservation and cleanup are now one block, because each fenced block is a
separate execution and a shell variable cannot carry between them. mkdir -p and
each cp are exit-checked; cleanup runs only when preservation succeeded, and a
failure warns naming the intact run directory. "Nothing to preserve" is not a
failure and still cleans up. Driven three ways: success removes run_dir, failure
leaves it intact with the warning, nothing-to-preserve removes it. The failure is
induced by a file-vs-directory conflict rather than chmod 0o000, which root
bypasses.

The new unresolved-depends_on warning embedded a user-authored token verbatim:

  warnings: ["Plan 03-02: depends_on token \"evil
  Plan 03-01: FORGED WARNING\" does not resolve ..."]

The JSON wire form is safe, and the security reviewer judged it non-exploitable
for that reason. It is escaped anyway through formatDiagnosticToken — the helper
#3884 added one phase earlier for exactly this class. warnings[] is an array a
consumer naturally prints line by line, and not reusing the sibling fix is the
generative-fix-divergence shape this epic exists to close. The same treatment is
applied to context_read_error / phase_dir_read_error, which embed a phase
directory path a repository can choose, and to the fs error message, which
echoes the raw path itself.

Known limit L5 recorded: the bound is on DEPTH only. A 300,000-element shallow
array yields a 14.5MB reply with truncated:false. Correct per §8.5 and per
negative space N2, disclosed rather than left to be discovered.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3885): unreadable is not absent in intel.cts either, and a corrupt snapshot is not "no snapshot"

Blocker from the round-2 isolated review, and it is my own inconsistency:
this phase applied "unreadable is not absent" to phase directories and left it
broken in the file it was already editing.

  chmod 000 .planning/intel/file-roles.json
  gsd-tools intel query <term>
  -> {"matches":[],"total":0,"truncated":false}  exit 0

safeReadJson swallowed every read failure and returned null, so an EACCES was
byte-indistinguishable from an absent file AND from a genuine no-match. Now it
separates three states: ENOENT stays silently absent, because not every project
has every intel file and intelQuery loops over all of them expecting misses;
EACCES/EIO and malformed JSON are both surfaced naming the file. A corrupt intel
file previously read as "no matches" too — same defect, same fix.

Threading that outcome through the other three callers found something worse
than the reported case. intelDiff returned no_baseline:true for a corrupt or
unreadable snapshot — not a silent failure but an actively FALSE verdict, telling
the caller they never took a snapshot when they did. That is §8.5's headline
case, so it is fixed and tested rather than noted. intelStatus and
intelApiSurface collapsed the same way; intelApiSurface additionally printed a
"not yet populated" banner that was simply untrue.

Every row is failing-first, including the absent-file ones — the field is new,
so it does not exist pre-fix at all. Those rows are not pre-fix pins; they pin
that the fix does not OVER-fire on the ordinary absent case, which is what would
turn this into noise on every project lacking an intel file. IO failure is
injected by monkeypatching fs and restoring in finally, never chmod 0o000 — root
bypasses mode bits, so the reviewer's manual chmod repro is not reproducible as
a test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): build the pathological intel fixture as text, not by stringifying a nested object

The remote runner came back red on Linux with two failures, both
T4: deeplyNestedIntelDoesNotOverflowTheStack, while the same test passed on
macOS. The product was never at fault.

writeNestedFixture(12000) built a 12,000-deep JavaScript OBJECT and then
JSON.stringify'd it. JSON.stringify recurses once per level, so it overflowed
the TEST PROCESS's stack — the error was thrown before the CLI was ever spawned.
Linux's container stack is smaller than macOS's, which is the whole of the
platform difference.

Measured, with the same document built as JSON TEXT so nothing in the building
process recurses:

  depth=100    rc=0 truncated=true
  depth=5000   rc=0 truncated=true
  depth=12000  rc=0 truncated=true
  depth=60000  rc=0 truncated=true

V8 parses this shape iteratively; only stringify recurses. The bound works at
every depth tried.

The fixture is now built by string concatenation. That is also the more faithful
input — a real deeply nested JSON document on disk is exactly what the bound
guards, where a stringified object was only ever a way to produce one.

The depth stays 12000. Lowering it would have made the test pass by weakening it
to accommodate a fixture bug, and 12000 is a legitimate pathological input the
product handles. T4 remains a genuine fail-first: rebuilt against the parent of
the commit that added the bound, the string-built depth-12000 fixture still
drives the CLI to rc=1 with "Error: Maximum call stack size exceeded".

A comment records why the fixture is text, so it is not "simplified" back into a
macOS-green / Linux-red test.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3885): backfill the changeset PR number

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): normalize path separators before splicing into the workflow's bash

CI red on one lane — test (windows-latest, 24, shard 3/3). macOS, Linux and the
remote runner were all green.

  AssertionError: commit must name the single REVIEWS.md file; got:
    --files C:UsersRUNNER~1AppDataLocalTempgsd-3352-phasedir-mOKmuy/03-REVIEWS.md

Every backslash in C:\Users\RUNNER~1\AppData\Local\Temp\... was eaten. The
harness spliced an OS-native temp path into the extracted bash, and bash consumes
\U, \A, \L and \T as escapes on an unquoted expansion. The same loss broke
RUN_DIR, so "rm -rf" targeted a path that never existed and the run directory
survived — which is the other two assertions.

This is a fixture defect, not a product one, and that was checked rather than
assumed. In production the phase directory is toPosixPath-normalized at every
call site that serializes it (bin/lib/init.cjs:951, 1381, 1461, 1529, 1595), and
the run directory is created by "mktemp -d" running inside the bash block itself
(gsd-core/workflows/review.md:163), which emits POSIX-style output even under
Git-Bash on Windows. Neither ever carries a backslash where the workflow reads it.

The file's pre-existing #3034 harness splices raw native paths too, but only ever
inside double-quoted assignments, so it never tripped this — my new harness
followed that convention faithfully into the one place where it does not hold.
Both now splice through toPosixPath from shell-command-projection, the
established seam, which is a no-op on POSIX and mirrors what production does.

No assertion was weakened. "commit must name the single REVIEWS.md file" and
"the run dir must still be destroyed" still assert exactly that; only how the
fixture supplies its path changed. Nothing is skipped on Windows — a t.skip()
here would have hidden the question of whether the exposure was real, which is
the question that mattered.

Driven both ways: a synthetic C:\Users\RUNNER~1\... input reproduces the exact CI
string when unfixed and yields C:/Users/RUNNER~1/... when fixed; a POSIX input
produces a byte-identical shape, proving the normalization is idempotent.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3885): stop the harness making the deleted run dir its own cwd

Windows shard 3/3 stayed red after the separator fix, on two assertions the
separator fix never touched:

  AssertionError: the run dir must still be destroyed
  AssertionError: nothing to preserve is not a failure — run dir must still be removed

The separators were a real bug and fixing them fixed the --files assertion. They
were not this bug, and two CI cycles went into the wrong axis before I stopped
converting path forms and looked at what the harness actually does.

runWriteReviewsFlow passed cwd: runDir to runHook, so the child bash process's
working directory WAS the directory the block under test then removes with
rm -rf "$RUN_DIR". POSIX allows a process to delete its own cwd — verified
locally, cd "$d"; rm -rf "$d" removes it cleanly — and Windows does not: a live
process's working directory cannot be removed. So on Windows the directory
survived and both assertions failed, on macOS and Linux it vanished and they
passed. Nothing to do with slashes.

Harness-only. Production never cd's into the run directory; every reference is by
absolute path, and RUN_DIR is created by mktemp -d inside the bash block itself
(gsd-core/workflows/review.md:165) rather than injected. review.md is unchanged.

Fix: the child now runs with its cwd in an unrelated temp directory that the
block under test never deletes. Neither assertion was weakened, and nothing is
skipped on Windows — the tests in this file carry no platform guard and run
there unconditionally, which is how this surfaced at all.

Honest limit: the Windows failure mode cannot be reproduced on macOS, because
POSIX permits the very thing Windows refuses. The diagnosis is grounded in that
documented divergence and in the fact that only the Windows lane failed, but the
green outcome on windows-latest is unverified until CI runs it.

Refs #3885

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-27 04:12:47 -04:00
Tom Boucher
14e0709eed refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active (#1315)
* refactor(#1307): gate intel on isCapabilityActive + loop-resolver honors capability active

Part A: intel's command gate moves from config-only isIntelEnabled to the
shared isCapabilityActive('intel', cwd) — a consistency cutover (intel has
skills:[] so its tri-state collapses to the intel.enabled config leg; the gate
now flows through the resolver's precedence + runtime-aware resolution).
Part B: loop-resolver hook rendering now gates on capability state.active
(=== true, fail-closed) instead of state.enabled, so the capability config
gate is honored by the hook consumer, not just per-hook 'when'. active is now
required in the loop-resolver input types. Regression test proves a config-
disabled (active=false) capability's unconditional hook is not rendered.
Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1307): add changeset for intel + loop-resolver active gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 00:12:54 -04:00
Tom Boucher
728788fdb3 fix(#1034): correct stale intel.cts comments to canonical INTEL_FILES names (#1036)
Three maintainer-facing comments in src/intel.cts referenced the retired
short/markdown intel filenames (files.json, deps.json, arch.md) and described
arch as searched "text lines", though the implementation iterates the exported
INTEL_FILES map and parses every entry — arch included — as JSON.

Update the comments to cite the canonical names (file-roles.json,
dependency-graph.json, arch-decisions.json) and the JSON-for-arch behavior.
Comment-only; no code or behavior change. Deferred from the #1000 fix.

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 11:55:55 -04:00
Tom Boucher
ba231ecbfc chore: clean up clear-cut ESLint warnings (#732) (#734)
Pay down pre-existing error→warn lint debt. Removes dead imports/vars, unused functions, redundant regex/string escapes, and stale eslint-disable directives; converts unused `catch (_e)` to optional catch binding (src/*.cts).

No behavior change. Lint 345→125 warnings (0 errors); deferred categories (n/no-process-exit, test-sleeps, control-regex) tracked in #732 for follow-up. Full test suite green (0 failures); code-review verified all removals unused and all escape fixes semantics-preserving.

Closes #732

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 11:24:48 -04:00
Tom Boucher
df04aae5e4 enhancement(#537): migrate all hand-written bin/lib/*.cjs to TypeScript source of truth (ADR-457) (#602)
* enhancement(#537): migrate code-review-flags to TS source of truth

Collapse the hand-written get-shit-done/bin/lib/code-review-flags.cjs to a
TypeScript source of truth (src/code-review-flags.cts), compiled by tsc to a
gitignored .cjs build artifact at the same path, per ADR-457 (build-at-publish).
Second module after the semver-compare pilot (#541).

Behaviour is preserved byte-for-behaviour (characterization test added in
tests/code-review-flags.test.cjs locks the parser quirks). Adds compile-time
type checking: CodeReviewFlags interface + CodeReviewWorkflow literal union.
The require() path is unchanged, so code-review.md and the bug-3727 test keep
working. The emitted .cjs is gitignored and eslint-ignored, mirroring the pilot.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 9 leaf bin/lib modules to TS source of truth

ADR-457 build-at-publish, batch 1 (pure leaf modules, 0 sibling-deps):
001-legacy-orphan-files, context-utilization, redaction, artifacts,
command-arg-projection, clock, ui-safety-gate, review-reviewer-selection,
clusters. Each moves to src/*.cts (strict TS, typed), compiled by tsc to a
gitignored .cjs at the same require() path; behaviour preserved byte-for-
behaviour. Adds src/node-globals.d.ts (minimal ambient shim; "types":[]).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#537): add @types/node, drop hand-rolled node-globals shim

ADR-457 migration infra: replace the temporary src/node-globals.d.ts ambient
shim with @types/node@22 + "types":["node"] in tsconfig.build.json. Unblocks
migrating the ~49 remaining bin/lib modules that use node:fs/path/os/
child_process. Build + full suite (3030 pass) + lint all green; no .cts type
changes were needed (real Node types matched the shim).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 9 more bin/lib modules to TS (batch 2)

ADR-457 build-at-publish. Clean leaves: installer-migration-report,
prompt-budget. Type-error-prone leaves (were tsconfig.lint-excluded; now
strict-typed and removed from that exclude list): secrets, phase-lifecycle,
workstream-name-policy, decisions, validate, schema-detect. Plus
runtime-name-policy. Strict type fixes narrow unknown->concrete domain types
(no any/ts-ignore); behaviour preserved. Full suite green, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate runtime-slash to TS (cross-import proof)

ADR-457. First cross-module TS->TS import: src/runtime-slash.cts imports
./runtime-name-policy.cjs and tsc resolves the sibling .cts types under strict
(no declaration files; NodeNext .cjs->.cts mapping), emitting a correct
require("./runtime-name-policy.cjs"). Confirms the recipe for coupled modules,
which must be migrated in dependency order (leaves-up). Suite green, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 10 more bin/lib modules to TS (batch 3)

ADR-457 build-at-publish, Wave-1 leaves: event, workstream-inventory-builder,
plan-scan, fallow-runner, project-root, installer-migration-authoring,
update-context, 000-first-time-baseline, runtime-homes, model-catalog. Strict
typing fixed real issues (narrowing unknown, qualified fs/path calls, removed
unnecessary casts); plan-scan/project-root/workstream-inventory-builder dropped
from tsconfig.lint exclude. Behaviour preserved; suite green, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 5 large Wave-1 leaves to TS (batch 4)

ADR-457 build-at-publish: configuration, state-document, shell-command-
projection (42 dependents), security, command-aliases. shell-command-
projection keeps a namespace child_process import for mock-intercept
testability. loadConfig/migrateOnDisk emit synchronously (every caller uses
them sync; the one awaited migrateOnDisk caller tolerates a non-Promise) —
full suite (3030 pass) confirms behaviour preserved. configuration/
state-document/command-aliases dropped from tsconfig.lint exclude. Also fixes
the malformed batch-3 changeset frontmatter (type/pr) that failed lint:docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 6 Wave-2 modules to TS (batch 5)

ADR-457 build-at-publish: config-schema, model-profiles,
002-codex-legacy-hooks-json, logger, active-workstream-store, adr-parser.
First batch importing already-migrated siblings (configuration, model-catalog,
shell-command-projection, redaction, security) via ./sibling.cjs specifiers.
Strict type narrowing (typeof guards over String(unknown)); behaviour
preserved; suite 3030 pass, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 5 large Wave-2 modules to TS (batch 6)

ADR-457 build-at-publish: graphify, install-profiles, intel,
installer-migrations, worktree-safety. installer-migrations preserves its
dynamic require() loader for numbered migration modules (scoped lint
suppressions). Strict typing (typeof guards over String(unknown)); behaviour
preserved; suite 3030 pass, lint 0 errors. Wave 2 complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate Wave-3 modules to TS (batch 7)

ADR-457 build-at-publish: planning-workspace, runtime-artifact-layout,
command-routing-hub, drift. Uses `import x = require()` for export= siblings;
drift's lazy require of runtime-slash hoisted to a top-level import (verified
non-circular). Behaviour preserved; suite 3030 pass, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate small Wave-4 modules to TS (batch 8)

ADR-457 build-at-publish: cjs-command-router-adapter, phase-command-router,
surface, roadmap-upgrade. Typed the hub router handler results as the HubResult
discriminated union; surface drops 4 genuinely-unused imports. Behaviour
preserved; suite 3030 pass, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate core hub (2.5k LOC, 68 dependents) to TS (batch 9)

ADR-457 build-at-publish: get-shit-done/bin/lib/core.cjs -> src/core.cts,
preserving all 63 exports via export=. All sibling deps already migrated
(shell-command-projection, model-profiles, model-catalog, worktree-safety,
planning-workspace, project-root, configuration, config-schema). Strict types,
no any/ts-ignore; config-schema lazy require hoisted (non-circular). Behaviour
preserved (independently verified: core's shard 3030 pass / 0 fail).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537): make ESLint-coverage + test-sprawl checks migration-aware

#551 test hardcoded 12 now-migrated modules as "hand-written, must be linted";
that invariant is obsoleted by the ADR-457 migration. Rewrite it to a
filesystem-driven invariant that holds at every stage: a bin/lib/*.cjs must be
eslint-ignored IFF it has a src/*.cts source (tsc-generated), else linted
(covers package-identity, which has no TS source). Also eslint-ignore
config-types.cjs (has a src counterpart) and drop the redundant
tests/clock.test.cjs (clock already covered by clock-seam + bug-474 tests),
which tripped the lint-test-file-count ratchet.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 9 Wave-5 router/inventory modules to TS (batch 10)

ADR-457 build-at-publish: phases/verify/init/agent/task/validate/roadmap/state
command routers + workstream-inventory. Router handler results typed against
core's exported shapes; behaviour preserved (caught+fixed a --verify boolean
flag regression mid-migration). Full suite green across all shards (only the 4
local gpg-env changeset-notes failures remain; CI passes them).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 7 Wave-5 modules to TS (batch 11)

ADR-457 build-at-publish: gap-checker, docs, check-command-router, frontmatter,
learnings, gsd2-import, profile-pipeline. Behaviour preserved; full suite green
across all shards (only the 4 local gpg-env failures remain). Also broadens
atomic-write-coverage.test.cjs to accept the tsc-compiled namespace-import form
while still asserting platformWriteSync is called (safety guard intact).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate config + profile-output to TS (batch 12)

ADR-457 build-at-publish: config (729 LOC), profile-output (1142 LOC). All
exports preserved; cmdMigrateConfig de-asynced (migrateOnDisk is sync, awaited
caller tolerates it). Behaviour preserved; suite green across all shards
(only the 4 local gpg-env failures). Wave 5 complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 5 Wave-6 modules to TS (batch 13)

ADR-457 build-at-publish: template, uat, workstream, roadmap, audit. Behaviour
preserved (dead toPosixPath import dropped from audit; inline requires hoisted).
Suite green across all shards (only the 4 local gpg-env failures).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate commands + state hubs to TS (batch 14)

ADR-457 build-at-publish: commands (1305 LOC), state (2074 LOC, 17 dependents).
All exports preserved; inner requires kept non-hoisted where load-order matters
(install.js, per-call security); acquireStateLock cast inlined to preserve the
err.code source token a structural test inspects. Behaviour preserved; suite
green across all shards (only the 4 local gpg-env failures). Wave 6 complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate milestone to TS (batch 15a, hand-authored)

ADR-457 build-at-publish: milestone -> src/milestone.cts. Authored directly
(subagent capacity was unavailable). Also relaxes core.output()'s 3rd param to
optional, matching its real always-optional call contract (unblocks remaining
2-arg output callers). Behaviour preserved; suite green across all shards
(only the 4 local gpg-env failures).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate phase, verify, init to TS (batch 15, final modules)

ADR-457 build-at-publish, Wave 7 (the last hubs): phase (1608 LOC), verify
(1615), init (2113). Adds src/package-identity.d.cts so verify can import the
permanently value-baked package-identity.cjs under strict TS.

Fixes two regressions the migration introduced in verify: restore
cmdValidateHealth's `return result` (callers/tests read result.warnings — it is
NOT side-effect-only), and make the bug-3384 source-pattern test tolerant of the
tsc-compiled bracket-notation form of the git_list_failed->W020 branch (behaviour
intact). Full suite green across all shards (only the 4 local gpg-env failures);
lint 0 errors. All 86 migratable bin/lib modules are now TypeScript sources.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#537): finalize ADR-457 migration — retire tsconfig.lint.json

All hand-written bin/lib/*.cjs are now src/*.cts sources, so the checkJs
stopgap tsconfig.lint.json (unused; not wired into eslint, scripts, or CI) is
deleted per ADR-457's final step. Also gitignore the tsc-generated
config-types.cjs (was still committed) for consistency with every other
emitted artifact. package-identity.cjs stays value-baked (declared via
src/package-identity.d.cts). Suite green; #551 ESLint-coverage test green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): add prepare script so unpacked/git installs build bin/lib artifacts

ADR-457 build-at-publish: bin/lib/*.cjs are now gitignored, built by tsc. The
prepack/prepublishOnly hooks cover `npm pack`/publish, but `npm install -g
<dir>` and git installs run the `prepare` lifecycle — which was missing — so the
unpacked install shipped without the compiled .cjs and failed at startup with
"Cannot find module './lib/core.cjs'" (caught by the smoke-unpacked CI job).
Add `prepare` mirroring prepublishOnly (build:lib + build:hooks). prepare does
NOT run for registry consumers (they get the pre-built tarball), only for
source/local/pack installs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): make CI build/lockfile checks work with gitignored bin/lib artifacts

ADR-457 build-at-publish exposed two CI assumptions that bin/lib/*.cjs are
always present on disk:
- check:env's lockfile-sync ran `npm ci --dry-run`, which now triggers the
  `prepare` build (tsc) — but it runs before deps are installed, so tsc is
  absent and it misreported the lockfile as out of sync. Add --ignore-scripts
  (a lockfile check must not build).
- the lint-tests job installs with --ignore-scripts (no prepare build), but
  lint:skill-deps require()s the built install-profiles.cjs. Add an explicit
  `npm run build:lib` step after install.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): narrow prepare to build:lib only (unbreak packed-smoke pack step)

prepare running build:hooks emitted "✓ Copying ..." stdout during `npm pack`,
which the install-smoke "Pack root tarball" step captures into $GITHUB_OUTPUT —
breaking it with "Invalid format". build:lib (tsc) is silent on success and is
all the unpacked/source install needs (the smoke-unpacked assertions exercise
gsd-tools, i.e. bin/lib, and tolerate hook setup with `|| true`). Matches
prepack. build:hooks still runs on prepublishOnly for real publishes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): wire Stryker mutation gate to build-at-publish layout

The gate scored 0.00 because it mutated changed bin/lib/*.cjs that (a) were
generated artifacts and (b) included modules with no coverage in the command's
test set. Rework: mutation.yml now derives changed COVERED modules from
src/*.cts and maps them to their built bin/lib/*.cjs; Stryker mutates those
built artifacts with a no-rebuild command (mutating src/*.cts + per-mutant tsc
was ~3x over the 30-min CI budget).

NOTE: with the gate now correctly measuring the covered modules, their actual
mutation score is 42.94% (< break 50) — a pre-existing test-coverage gap
(adr-parser/prompt-budget/etc.), not introduced by this behaviour-preserving
migration. Reaching 50 needs more tests, a threshold/scope change, or a waiver —
a maintainer decision.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537): raise mutation coverage of covered modules above the 50 gate

Adds focused example-based unit tests that kill surviving mutants in the two
lowest-scoring covered modules:
- tests/prompt-budget.unit.test.cjs (112 tests): 17.9% -> 97.9%
- tests/adr-parser.unit.test.cjs (205 tests): 44.7% -> 89.4%
Both wired into stryker.config.mjs's command. Fresh full run over the 6 covered
modules now scores 82.25% (>= break 50); every covered module is >= 68%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537,#609): parallelize mutation gate via dynamic per-module matrix

The serial Stryker run timed out at 30 min once the migration's added tests
made every mutant re-run ~300 tests. Replace it with a dynamic matrix so the
gate completes well under budget — folded into this PR (was tracked as #609)
because it's a prerequisite for this PR's mutation gate to pass.

- scripts/mutation-matrix.cjs: single source of truth (covered-module -> test
  files) computing changed covered modules from git diff -> {has_work, matrix}.
- mutation.yml: detect -> dynamic `matrix: fromJSON(...)` mutate job (one
  parallel shard per changed module, scoped via MUTATION_TEST_CMD to only that
  module's tests, 15-min/shard) -> summary job that KEEPS the legacy check name
  "Stryker mutation score (changed files only)" so branch protection is
  unchanged. Per-shard jobs report as "Stryker (<module>)".
- stryker.config.mjs: commandRunner.command reads MUTATION_TEST_CMD (falls back
  to the full command locally).

Closes #609.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537,#609): give each mutation shard ≥50% on its own tests; drop blacksmith note

Per-module sharding revealed that active-workstream-store (46.5%) and
frontmatter (7.4%) only cleared 50% in the old serial run via timeout-noise from
the bloated 300-test command; on their own tests they were below the gate. Add
focused unit tests:
- tests/active-workstream-store.unit.test.cjs (115 tests): 46.5% -> 81.9%
- tests/frontmatter.unit.test.cjs (165 tests): 7.4% -> 63.4%
Both wired into scripts/mutation-matrix.cjs (per-module test map) and
stryker.config.mjs DEFAULT_TEST_CMD. All 6 covered modules now clear break:50
with only their own tests (config-schema/context-utilization/prompt-budget/
adr-parser already did). Also removes the leftover blacksmith TODO comment —
GitHub-hosted runners only; speed comes from parallel per-module shards.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537,#609): strengthen prompt-budget tests to clear the gate on its own tests

prompt-budget scored 39.58% when mutation-tested with ONLY its own tests (the
way the per-module CI shard runs it) — an earlier ~98% reading was inflated by
accidentally running the full multi-module command. Add 96 targeted tests to
tests/prompt-budget.unit.test.cjs (exact note-template text, plan-truncation
arithmetic/percentages, drop-block strings, noteInjected/hardFailed booleans):
scoped score 39.58% -> 68.75% (>= break 50). All 6 covered modules now clear
the gate on their own tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 11:45:01 -04:00