Commit Graph

7 Commits

Author SHA1 Message Date
Tom Boucher
d6cce9e2c3 fix(#3020): verify graphify tool identity before reporting compatibility (#3107)
* fix(#3020): verify graphify tool identity before reporting compatibility

checkGraphifyVersion ran 'graphify --version' on PATH and trusted any
plausible version string — a foreign binary named 'graphify' that happened
to print a version would silently report compatible:true with no warning.
Downstream graphify work would then proceed against the wrong tool and fail
later in ways that looked like GSD defects.

After getting a version from the binary, verify the graphifyy Python package
via importlib.metadata. If the package cannot be confirmed, emit a warning
naming the mismatch regardless of version-range compatibility. The
compatible flag is now correctly false for unverified tools (it was
previously read by nobody — the warning is what surfaces).

* chore(#3020): backfill changeset PR number 3107

---------

Co-authored-by: sim <sim@local>
2026-08-06 02:56:07 -04:00
Tom Boucher
7e1004d89e test(#3056): add the in-process fault-injection adapter, and normalize execGit's shape (#3077)
* refactor(#3071): normalize execGit's call and result shape, unify ExecGitFn

ExecGitFn was declared four times. Three were hand-copies of one signature and
two of those were wrong: they typed exitCode as number|null when _spawnResult
returns `result.status ?? 1` and can never yield null, weakened signal from
NodeJS.Signals to string, and widened error from Error to unknown. Only
verification.cts got it right, via `typeof execGit`.

The root cause was a missing export: SpawnResultOutput was declared without
`export`, so no other module could name the return type of execGit. Three
authors independently hand-copied it instead. Exported now.

Normalizing the type alone would have left the pressure that caused the
divergence, so the function is normalized on both sides. It now ACCEPTS every
call its consumers make — worktree-safety's declaration could not express an
env-carrying call at all — and RETURNS every result code they need: timedOut
moves into _spawnResult, so execGit, execNpm and execTool all carry it and the
one extension that justified a separate type disappears. All four sites are
now `typeof execGit` with nothing left to restate.

timedOut reuses the existing isSpawnTimeout predicate introduced by #3050
rather than re-deriving it. That predicate checks error.code === 'ETIMEDOUT'
only; the signal === 'SIGTERM' conjunct was deliberately dropped there because
Windows does not reliably report SIGTERM and requiring it risks a false
negative. There is no false-positive risk, and a test proves it: an
externally-delivered SIGTERM leaves error null, so it is still not reported as
a timeout.

No dead null-checks surfaced. Every exitCode comparison in the two affected
modules is === 0, !== 0 or === 128 — never a null guard — so the nullable
declaration had never been written against.

Closes #3071

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3056): add the in-process fault-injection adapter

Adds tests/helpers/faulty-deps.cjs — makeFaultyGit() and withFaultyFs() — so a
module's error branch can be driven deterministically and its degraded verdict
asserted, instead of a counter-test that only proves the call did not throw.

makeFaultyGit returns a value structurally assignable to `typeof execGit`, so
one stub satisfies all four seams that the #3071 normalization collapsed into
that single shape. A parity test drives the same stub through a real injectable
entry point in each of worktree-safety, git-base-branch, worktree-base-ref and
verification; it fails the moment any of them re-grows its own shape.

Faults are scoped rather than global — by argv predicate and by call ordinal —
because a fault adapter that faults everything looks like it works and proves
nothing, and because verification.cts's two-call error handling needs to fault
the second call only. Invocations are recorded so a test can assert an exact
call count.

The timeout fault carries error.code === 'ETIMEDOUT', and a test asserts the
real isSpawnTimeout predicate matches it, so the fixture cannot drift from the
production definition of a timeout. A companion test asserts an externally
delivered SIGTERM with a null error is still NOT reported as a timeout — the
false-positive guard for #3050's dropped conjunct.

withFaultyFs restores in a finally so a throwing body still restores, patches
only the named methods, and nests without clobbering an outer saved original.
It never uses chmod: that no-ops under root, so the test would pass with zero
coverage in root Docker/CI.

The adapter is in-process via deps only. The Phase 1 process seam is documented
as deliberately not a fault-injection surface — it cannot distinguish an
injected timeout from a genuine bench OOM and would retry it.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3071): add changeset fragment for the execGit normalization

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3071): route the last two timeout checks through the shared predicate

An isolated review found this branch had normalised the timeout verdict but
left two callers still hand-rolling the fragile version of it.

check-command-router's runBoundedShell computed `timedOut: r.signal ===
'SIGTERM'` while the correctly-derived `r.timedOut` sat on the same result
object. graphify's execGraphify branched on `result.signal === 'SIGTERM'`, with
a comment asserting the very premise isSpawnTimeout exists to reject.

Both fail in both directions. On Windows a genuine timeout is not reliably
reported as SIGTERM, so the guard silently fails to fire — the false negative
#3050 was raised for. And an externally-delivered SIGTERM is not a timeout at
all, so the check also fires when it should not; isSpawnTimeout avoids that
because `error` is null in that case and it keys on error.code.

Both now read the derived `timedOut`, and graphify's comment states the actual
rule instead of the fragile assumption.

Also replaces a vacuous test: "execGitDefault now accepts env" never called
execGitDefault (it is unexported), called execGit — whose signature already
accepted env before this branch — and asserted only that exitCode was a
number, which would pass whether or not the change under test existed. It now
proves env reaches the child by asserting `git var GIT_EDITOR` returns the
injected sentinel.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3071): make the graphify timeout fixture faithful to a real timeout

The remote matrix failed on both Linux lanes: graphify's "returns exitCode 124
on timeout" got 1 instead of 124. The fixture was wrong, not the production
change.

It stubbed spawnSync as { status: null, signal: 'SIGTERM', error: undefined }.
That is not a timeout. A real spawnSync timeout also sets error.code
'ETIMEDOUT'; a SIGTERM with no error is an externally delivered signal — a
kill. The old `result.signal === 'SIGTERM'` check accepted it as a timeout,
which is the false positive the shared predicate exists to reject, so this test
was locking that bug in rather than guarding against it.

The fixture now carries a real ETIMEDOUT error and all three original
assertions pass unchanged. A counter-test is added alongside it: an externally
delivered SIGTERM with no error must NOT be reported as a timeout. That is the
assertion whose absence let the false positive live.

Swept every SIGTERM/SIGKILL stub under tests/ for the same unfaithful shape.
No other instance: the worktree-safety, worktree-base-ref and commit-staging
fixtures already set ETIMEDOUT, and the remaining hits are either deliberate
external-kill tests or feed code that never consults timedOut.

Two sites keep their own signal check deliberately and are NOT changed:
capability-source.cts:1301,1386 fail closed on ANY abnormal termination, which
is correct — reading timedOut there would stop it failing closed on a kill and
let it parse stdout from a killed process. Their reason strings, and
check-latest-version.cjs:115, label any signal as "timed out", which is
imprecise wording over a correct verdict, not a silent failure.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#3071): backfill changeset pr number to 3077

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 02:14:41 -04:00
0xdhx
f4d6747d21 fix(#2738): report graphify query budget outcome and stop between-tier over-trimming (#2819)
* fix(#2738): report budget outcome from graphify query and stop over-trimming between tiers

applyBudget retains seed nodes unconditionally, so the seed set is a floor
the edge-tier reduction cannot go below — a --budget 500 request could
return the full seed payload (~119k tokens measured) with no signal of the
miss. Add budget_met + budget_estimate to the budget result and surface
them through graphifyQuery when a budget was requested.

Secondary: the tier loop estimated against the full pre-filter node set,
so a tier removal that already satisfied the budget (once its orphaned
nodes were excluded) still triggered the next, higher-confidence tier
drop. Recompute reachability and the estimate after each tier and break
as soon as the pruned result fits.

Adjacent same-class instance from pre-submit review: the CLI forwards
--budget 0 but truthiness checks silently treated it as no budget and
returned the unbounded result. Test budget presence with != null so a
zero budget is honored and reported as an unmeetable miss.

* docs(changeset): backfill PR number for #2738 fragment

* fix(#2738): estimate the payload as emitted, not a private compact form

budget_met measured a different payload than the caller receives. The
estimator serialized a compact `{nodes, edges}`, while output() emits the
whole response pretty-printed (2-space indent, plus the term/total_*/trimmed
wrapper keys). Measured on the repo's own SAMPLE_GRAPH fixture: reported 183
tokens against 302 actually emitted — 1.65x — so `--budget 200` returned
budget_met: true while handing back 302 tokens. That is worse than the old
silent miss: an automated consumer stops checking a signal that is
confidently wrong.

Fix the basis rather than the number:

- io.cts gains serializeForOutput(), the single definition of the wire form.
  output() now calls it, so the estimator and the emitter cannot drift on
  indentation or shape. Pure extraction; output()'s behaviour is unchanged.
- graphify builds its response through one buildQueryResponse() used by both
  the emitter and the estimator, so the estimate describes exactly the bytes
  returned.
- The tier loop estimates on that same basis, so it keeps trimming until the
  real payload fits instead of stopping at a smaller internal measure. This
  makes budget_met === (budget_estimate <= budget) true by construction.
- Drop the module-private chars/4 helper for prompt-budget's estimateTokens —
  the repo's single token scale, per the rule phase-estimation.cts documents.

budget_estimate is self-referential (its own digits are part of the emitted
bytes), resolved by iterating to a fixed point; the sequence only ever grows,
so it settles in a couple of passes and errs toward over-reporting.

Tests pin estimator to emitter so this cannot silently re-diverge if output()
ever changes its indentation. Both new tests fail against the pre-fix source.

* test(#2738): property-test the budget-limit and reporting contract

RULESET.TESTS.property-based-testing names budget-limit contracts, and this
module is the literal case: #2819 turns it into a *reporting* contract, which
is what properties express well. Five invariants over arbitrary small graphs
and any budget >= 0:

- budget_met === (budget_estimate <= budget)
- budget_estimate === the tokens actually emitted
- the seed set is a floor the reduction never goes below (the changeset's
  "seeds are a floor" claim, previously asserted only for one hand-built
  fixture)
- total_nodes/total_edges match the returned arrays
- a larger budget never yields a smaller payload

Two notes on what these do and do not prove. The emitted-payload property
fails against the pre-fix source; the budget_met/budget_estimate agreement
property does NOT — pre-fix both derived from the same wrong number, so it
is a contract guard, not a regression proof.

Monotonicity is asserted over the payload (node/edge counts), not over
budget_estimate: the estimate measures emitted bytes exactly, and budget_met
renders as "false" (5 chars) or "true" (4), so an identical payload can
measure one token larger when the budget is missed. That is the estimate
being honest, not a monotonicity break.

The generator seeds on `label` — seedAndExpand matches label/description,
never id/name, and a fixture that gets this wrong expands to nothing and
passes vacuously.

* test(#2738): pin the budget boundary and label the forward-guard test

Two test-coverage gaps from review.

RULESET.TESTS.boundary-coverage wants limit-1 / limit / limit+1. The added
tests used 1, 0, 200, 100000, 50 — all far from the decision point. The
branch that matters is `estimate <= budgetTokens`, so the input that decides
it is budget === estimate exactly: that is the one value where an off-by-one
in the comparison flips budget_met, and nothing else in the suite would
catch it. Pinned at the limit and either side of it.

The "omits budget fields when no budget was requested" test passes unchanged
on next — graphifyQuery never set those keys before the fix, so both
assertions already held. It has value as a forward guard against the spread
leaking budget fields, but it is not failing-first and should not be counted
toward RULESET.TESTS.regression-must-fail-first. Said so in a comment, so a
later reader does not mistake it for the regression proof.

* fix(#2738): keep a non-finite budget out of the budget path

Switching `!budgetTokens` to `budgetTokens == null` widened the internal
contract to admit NaN, where every `estimate <= NaN` is false: the loop
strips all three tiers and returns a seeds-only payload that is
indistinguishable from a legitimate aggressive trim.

Unreachable through the CLI — graphify-command-router rejects a non-numeric
--budget with makeInvalidArgs before graphifyQuery is called — so this is
hardening, not a live defect. It is still worth guarding: graphifyQuery and
applyBudget are module-level entry points a future caller could reach
without the router's validation, and the failure mode is silent.

Number.isFinite also routes Infinity to the no-budget path, deliberately: an
unbounded budget is not a budget, and parseInt cannot produce one anyway.

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-08-01 21:36:48 -04:00
Tom Boucher
bd77b40107 feat(#1825): configurable graphify graph location (graphify.graph_path) (#2013)
* feat(#1825): configurable graphify graph location (graphify.graph_path)

Add a graphify.graph_path config key (.planning/config.json) that overrides
where /gsd-graphify query|status|diff read the knowledge graph, so one curated
umbrella-level cross-repo graph can serve multiple sibling projects without N
drifting ~5 MB mirror copies. Previously the graph location was hardcoded to
<cwd>/.planning/graphs/.

- src/graphify.cts: resolveGraphLocation(cwd, planningDir) honors the key
  (resolved relative to project root; absolute paths honored via path.resolve);
  falls back to the historical .planning/graphs/graph.json when unset/blank/
  non-string (byte-identical). Wired into graphifyQuery, graphifyStatus,
  graphifyDiff (snapshot travels with the configured graph via dirname), and
  writeSnapshot. Configured-but-missing -> actionable error naming the path.
  Build stays project-scoped (skill hardcodes the cp dest); umbrella graph is
  built in the umbrella project, sub-projects only READ it.
- config-schema.manifest.json: register graphify.graph_path in validKeys.
- tests/graphify-graph-path.test.cjs: boundary matrix (unset byte-identical,
  set+present reads configured graph not default, set+missing actionable error,
  relative resolved vs project root, blank treated as unset, snapshot alongside
  configured graph, diff from configured dir, build project-scoped) +
  VALID_CONFIG_KEYS registration.
- docs: CONFIGURATION.md row, FEATURES.md REQ-GRAPH-06, CONTEXT.md module note,
  .changeset (Added).

Closes #1825

* docs(#1825): backfill changeset pr number 2013
2026-07-05 14:50:01 -04:00
Tom Boucher
e3b829e765 refactor(#1306): gate graphify on isCapabilityActive (tri-state, runtime-aware) (#1313)
* refactor(#1306): gate graphify on isCapabilityActive (tri-state), not config-only

graphify's command gate moves from the config-only isGraphifyEnabled to the
shared isCapabilityActive('graphify', cwd) — so graphify is off unless installed
AND surfaced AND graphify.enabled. Fixes a latent claude-hardcoding in the
resolver: resolveCapabilityRuntimeState now detects the active runtime via
resolveRuntime(cwd) (GSD_RUNTIME -> config.runtime -> 'claude') so non-Claude
runtimes (Codex/Cursor) read their own surface, not ~/.claude. Hermetic
regression test proves config-on+unsurfaced -> disabled; cross-runtime test
proves GSD_RUNTIME=codex honors CODEX_HOME. Gate fails closed on every error
path. Part of #1302.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(#1306): add changeset for graphify tri-state gate

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 23:19:21 -04:00
Tom Boucher
ba231ecbfc chore: clean up clear-cut ESLint warnings (#732) (#734)
Pay down pre-existing error→warn lint debt. Removes dead imports/vars, unused functions, redundant regex/string escapes, and stale eslint-disable directives; converts unused `catch (_e)` to optional catch binding (src/*.cts).

No behavior change. Lint 345→125 warnings (0 errors); deferred categories (n/no-process-exit, test-sleeps, control-regex) tracked in #732 for follow-up. Full test suite green (0 failures); code-review verified all removals unused and all escape fixes semantics-preserving.

Closes #732

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 11:24:48 -04:00
Tom Boucher
df04aae5e4 enhancement(#537): migrate all hand-written bin/lib/*.cjs to TypeScript source of truth (ADR-457) (#602)
* enhancement(#537): migrate code-review-flags to TS source of truth

Collapse the hand-written get-shit-done/bin/lib/code-review-flags.cjs to a
TypeScript source of truth (src/code-review-flags.cts), compiled by tsc to a
gitignored .cjs build artifact at the same path, per ADR-457 (build-at-publish).
Second module after the semver-compare pilot (#541).

Behaviour is preserved byte-for-behaviour (characterization test added in
tests/code-review-flags.test.cjs locks the parser quirks). Adds compile-time
type checking: CodeReviewFlags interface + CodeReviewWorkflow literal union.
The require() path is unchanged, so code-review.md and the bug-3727 test keep
working. The emitted .cjs is gitignored and eslint-ignored, mirroring the pilot.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 9 leaf bin/lib modules to TS source of truth

ADR-457 build-at-publish, batch 1 (pure leaf modules, 0 sibling-deps):
001-legacy-orphan-files, context-utilization, redaction, artifacts,
command-arg-projection, clock, ui-safety-gate, review-reviewer-selection,
clusters. Each moves to src/*.cts (strict TS, typed), compiled by tsc to a
gitignored .cjs at the same require() path; behaviour preserved byte-for-
behaviour. Adds src/node-globals.d.ts (minimal ambient shim; "types":[]).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#537): add @types/node, drop hand-rolled node-globals shim

ADR-457 migration infra: replace the temporary src/node-globals.d.ts ambient
shim with @types/node@22 + "types":["node"] in tsconfig.build.json. Unblocks
migrating the ~49 remaining bin/lib modules that use node:fs/path/os/
child_process. Build + full suite (3030 pass) + lint all green; no .cts type
changes were needed (real Node types matched the shim).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 9 more bin/lib modules to TS (batch 2)

ADR-457 build-at-publish. Clean leaves: installer-migration-report,
prompt-budget. Type-error-prone leaves (were tsconfig.lint-excluded; now
strict-typed and removed from that exclude list): secrets, phase-lifecycle,
workstream-name-policy, decisions, validate, schema-detect. Plus
runtime-name-policy. Strict type fixes narrow unknown->concrete domain types
(no any/ts-ignore); behaviour preserved. Full suite green, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate runtime-slash to TS (cross-import proof)

ADR-457. First cross-module TS->TS import: src/runtime-slash.cts imports
./runtime-name-policy.cjs and tsc resolves the sibling .cts types under strict
(no declaration files; NodeNext .cjs->.cts mapping), emitting a correct
require("./runtime-name-policy.cjs"). Confirms the recipe for coupled modules,
which must be migrated in dependency order (leaves-up). Suite green, lint clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 10 more bin/lib modules to TS (batch 3)

ADR-457 build-at-publish, Wave-1 leaves: event, workstream-inventory-builder,
plan-scan, fallow-runner, project-root, installer-migration-authoring,
update-context, 000-first-time-baseline, runtime-homes, model-catalog. Strict
typing fixed real issues (narrowing unknown, qualified fs/path calls, removed
unnecessary casts); plan-scan/project-root/workstream-inventory-builder dropped
from tsconfig.lint exclude. Behaviour preserved; suite green, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 5 large Wave-1 leaves to TS (batch 4)

ADR-457 build-at-publish: configuration, state-document, shell-command-
projection (42 dependents), security, command-aliases. shell-command-
projection keeps a namespace child_process import for mock-intercept
testability. loadConfig/migrateOnDisk emit synchronously (every caller uses
them sync; the one awaited migrateOnDisk caller tolerates a non-Promise) —
full suite (3030 pass) confirms behaviour preserved. configuration/
state-document/command-aliases dropped from tsconfig.lint exclude. Also fixes
the malformed batch-3 changeset frontmatter (type/pr) that failed lint:docs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 6 Wave-2 modules to TS (batch 5)

ADR-457 build-at-publish: config-schema, model-profiles,
002-codex-legacy-hooks-json, logger, active-workstream-store, adr-parser.
First batch importing already-migrated siblings (configuration, model-catalog,
shell-command-projection, redaction, security) via ./sibling.cjs specifiers.
Strict type narrowing (typeof guards over String(unknown)); behaviour
preserved; suite 3030 pass, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 5 large Wave-2 modules to TS (batch 6)

ADR-457 build-at-publish: graphify, install-profiles, intel,
installer-migrations, worktree-safety. installer-migrations preserves its
dynamic require() loader for numbered migration modules (scoped lint
suppressions). Strict typing (typeof guards over String(unknown)); behaviour
preserved; suite 3030 pass, lint 0 errors. Wave 2 complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate Wave-3 modules to TS (batch 7)

ADR-457 build-at-publish: planning-workspace, runtime-artifact-layout,
command-routing-hub, drift. Uses `import x = require()` for export= siblings;
drift's lazy require of runtime-slash hoisted to a top-level import (verified
non-circular). Behaviour preserved; suite 3030 pass, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate small Wave-4 modules to TS (batch 8)

ADR-457 build-at-publish: cjs-command-router-adapter, phase-command-router,
surface, roadmap-upgrade. Typed the hub router handler results as the HubResult
discriminated union; surface drops 4 genuinely-unused imports. Behaviour
preserved; suite 3030 pass, lint 0 errors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate core hub (2.5k LOC, 68 dependents) to TS (batch 9)

ADR-457 build-at-publish: get-shit-done/bin/lib/core.cjs -> src/core.cts,
preserving all 63 exports via export=. All sibling deps already migrated
(shell-command-projection, model-profiles, model-catalog, worktree-safety,
planning-workspace, project-root, configuration, config-schema). Strict types,
no any/ts-ignore; config-schema lazy require hoisted (non-circular). Behaviour
preserved (independently verified: core's shard 3030 pass / 0 fail).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537): make ESLint-coverage + test-sprawl checks migration-aware

#551 test hardcoded 12 now-migrated modules as "hand-written, must be linted";
that invariant is obsoleted by the ADR-457 migration. Rewrite it to a
filesystem-driven invariant that holds at every stage: a bin/lib/*.cjs must be
eslint-ignored IFF it has a src/*.cts source (tsc-generated), else linted
(covers package-identity, which has no TS source). Also eslint-ignore
config-types.cjs (has a src counterpart) and drop the redundant
tests/clock.test.cjs (clock already covered by clock-seam + bug-474 tests),
which tripped the lint-test-file-count ratchet.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 9 Wave-5 router/inventory modules to TS (batch 10)

ADR-457 build-at-publish: phases/verify/init/agent/task/validate/roadmap/state
command routers + workstream-inventory. Router handler results typed against
core's exported shapes; behaviour preserved (caught+fixed a --verify boolean
flag regression mid-migration). Full suite green across all shards (only the 4
local gpg-env changeset-notes failures remain; CI passes them).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 7 Wave-5 modules to TS (batch 11)

ADR-457 build-at-publish: gap-checker, docs, check-command-router, frontmatter,
learnings, gsd2-import, profile-pipeline. Behaviour preserved; full suite green
across all shards (only the 4 local gpg-env failures remain). Also broadens
atomic-write-coverage.test.cjs to accept the tsc-compiled namespace-import form
while still asserting platformWriteSync is called (safety guard intact).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate config + profile-output to TS (batch 12)

ADR-457 build-at-publish: config (729 LOC), profile-output (1142 LOC). All
exports preserved; cmdMigrateConfig de-asynced (migrateOnDisk is sync, awaited
caller tolerates it). Behaviour preserved; suite green across all shards
(only the 4 local gpg-env failures). Wave 5 complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate 5 Wave-6 modules to TS (batch 13)

ADR-457 build-at-publish: template, uat, workstream, roadmap, audit. Behaviour
preserved (dead toPosixPath import dropped from audit; inline requires hoisted).
Suite green across all shards (only the 4 local gpg-env failures).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate commands + state hubs to TS (batch 14)

ADR-457 build-at-publish: commands (1305 LOC), state (2074 LOC, 17 dependents).
All exports preserved; inner requires kept non-hoisted where load-order matters
(install.js, per-call security); acquireStateLock cast inlined to preserve the
err.code source token a structural test inspects. Behaviour preserved; suite
green across all shards (only the 4 local gpg-env failures). Wave 6 complete.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate milestone to TS (batch 15a, hand-authored)

ADR-457 build-at-publish: milestone -> src/milestone.cts. Authored directly
(subagent capacity was unavailable). Also relaxes core.output()'s 3rd param to
optional, matching its real always-optional call contract (unblocks remaining
2-arg output callers). Behaviour preserved; suite green across all shards
(only the 4 local gpg-env failures).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537): migrate phase, verify, init to TS (batch 15, final modules)

ADR-457 build-at-publish, Wave 7 (the last hubs): phase (1608 LOC), verify
(1615), init (2113). Adds src/package-identity.d.cts so verify can import the
permanently value-baked package-identity.cjs under strict TS.

Fixes two regressions the migration introduced in verify: restore
cmdValidateHealth's `return result` (callers/tests read result.warnings — it is
NOT side-effect-only), and make the bug-3384 source-pattern test tolerant of the
tsc-compiled bracket-notation form of the git_list_failed->W020 branch (behaviour
intact). Full suite green across all shards (only the 4 local gpg-env failures);
lint 0 errors. All 86 migratable bin/lib modules are now TypeScript sources.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(#537): finalize ADR-457 migration — retire tsconfig.lint.json

All hand-written bin/lib/*.cjs are now src/*.cts sources, so the checkJs
stopgap tsconfig.lint.json (unused; not wired into eslint, scripts, or CI) is
deleted per ADR-457's final step. Also gitignore the tsc-generated
config-types.cjs (was still committed) for consistency with every other
emitted artifact. package-identity.cjs stays value-baked (declared via
src/package-identity.d.cts). Suite green; #551 ESLint-coverage test green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): add prepare script so unpacked/git installs build bin/lib artifacts

ADR-457 build-at-publish: bin/lib/*.cjs are now gitignored, built by tsc. The
prepack/prepublishOnly hooks cover `npm pack`/publish, but `npm install -g
<dir>` and git installs run the `prepare` lifecycle — which was missing — so the
unpacked install shipped without the compiled .cjs and failed at startup with
"Cannot find module './lib/core.cjs'" (caught by the smoke-unpacked CI job).
Add `prepare` mirroring prepublishOnly (build:lib + build:hooks). prepare does
NOT run for registry consumers (they get the pre-built tarball), only for
source/local/pack installs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): make CI build/lockfile checks work with gitignored bin/lib artifacts

ADR-457 build-at-publish exposed two CI assumptions that bin/lib/*.cjs are
always present on disk:
- check:env's lockfile-sync ran `npm ci --dry-run`, which now triggers the
  `prepare` build (tsc) — but it runs before deps are installed, so tsc is
  absent and it misreported the lockfile as out of sync. Add --ignore-scripts
  (a lockfile check must not build).
- the lint-tests job installs with --ignore-scripts (no prepare build), but
  lint:skill-deps require()s the built install-profiles.cjs. Add an explicit
  `npm run build:lib` step after install.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): narrow prepare to build:lib only (unbreak packed-smoke pack step)

prepare running build:hooks emitted "✓ Copying ..." stdout during `npm pack`,
which the install-smoke "Pack root tarball" step captures into $GITHUB_OUTPUT —
breaking it with "Invalid format". build:lib (tsc) is silent on success and is
all the unpacked/source install needs (the smoke-unpacked assertions exercise
gsd-tools, i.e. bin/lib, and tolerate hook setup with `|| true`). Matches
prepack. build:hooks still runs on prepublishOnly for real publishes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(#537): wire Stryker mutation gate to build-at-publish layout

The gate scored 0.00 because it mutated changed bin/lib/*.cjs that (a) were
generated artifacts and (b) included modules with no coverage in the command's
test set. Rework: mutation.yml now derives changed COVERED modules from
src/*.cts and maps them to their built bin/lib/*.cjs; Stryker mutates those
built artifacts with a no-rebuild command (mutating src/*.cts + per-mutant tsc
was ~3x over the 30-min CI budget).

NOTE: with the gate now correctly measuring the covered modules, their actual
mutation score is 42.94% (< break 50) — a pre-existing test-coverage gap
(adr-parser/prompt-budget/etc.), not introduced by this behaviour-preserving
migration. Reaching 50 needs more tests, a threshold/scope change, or a waiver —
a maintainer decision.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537): raise mutation coverage of covered modules above the 50 gate

Adds focused example-based unit tests that kill surviving mutants in the two
lowest-scoring covered modules:
- tests/prompt-budget.unit.test.cjs (112 tests): 17.9% -> 97.9%
- tests/adr-parser.unit.test.cjs (205 tests): 44.7% -> 89.4%
Both wired into stryker.config.mjs's command. Fresh full run over the 6 covered
modules now scores 82.25% (>= break 50); every covered module is >= 68%.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* enhancement(#537,#609): parallelize mutation gate via dynamic per-module matrix

The serial Stryker run timed out at 30 min once the migration's added tests
made every mutant re-run ~300 tests. Replace it with a dynamic matrix so the
gate completes well under budget — folded into this PR (was tracked as #609)
because it's a prerequisite for this PR's mutation gate to pass.

- scripts/mutation-matrix.cjs: single source of truth (covered-module -> test
  files) computing changed covered modules from git diff -> {has_work, matrix}.
- mutation.yml: detect -> dynamic `matrix: fromJSON(...)` mutate job (one
  parallel shard per changed module, scoped via MUTATION_TEST_CMD to only that
  module's tests, 15-min/shard) -> summary job that KEEPS the legacy check name
  "Stryker mutation score (changed files only)" so branch protection is
  unchanged. Per-shard jobs report as "Stryker (<module>)".
- stryker.config.mjs: commandRunner.command reads MUTATION_TEST_CMD (falls back
  to the full command locally).

Closes #609.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537,#609): give each mutation shard ≥50% on its own tests; drop blacksmith note

Per-module sharding revealed that active-workstream-store (46.5%) and
frontmatter (7.4%) only cleared 50% in the old serial run via timeout-noise from
the bloated 300-test command; on their own tests they were below the gate. Add
focused unit tests:
- tests/active-workstream-store.unit.test.cjs (115 tests): 46.5% -> 81.9%
- tests/frontmatter.unit.test.cjs (165 tests): 7.4% -> 63.4%
Both wired into scripts/mutation-matrix.cjs (per-module test map) and
stryker.config.mjs DEFAULT_TEST_CMD. All 6 covered modules now clear break:50
with only their own tests (config-schema/context-utilization/prompt-budget/
adr-parser already did). Also removes the leftover blacksmith TODO comment —
GitHub-hosted runners only; speed comes from parallel per-module shards.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(#537,#609): strengthen prompt-budget tests to clear the gate on its own tests

prompt-budget scored 39.58% when mutation-tested with ONLY its own tests (the
way the per-module CI shard runs it) — an earlier ~98% reading was inflated by
accidentally running the full multi-module command. Add 96 targeted tests to
tests/prompt-budget.unit.test.cjs (exact note-template text, plan-truncation
arithmetic/percentages, drop-block strings, noteInjected/hardFailed booleans):
scoped score 39.58% -> 68.75% (>= break 50). All 6 covered modules now clear
the gate on their own tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-02 11:45:01 -04:00