* test(#1953): failing-first suite for the complexity-triggered refactor hook 60 behavioral cases against src/complexity-trigger.cts, which does not exist yet: decision-point counting, the comment/literal stripping leak surface, threshold and jump-delta boundaries at limit-1/limit/limit+1, stable-anchor baseline semantics, and fs fault injection via mock.method. Two fast-check properties assert that stripping never manufactures a decision point and that comments and string literals are score-neutral. Also registers the refactor-trigger capability manifest (inert until refactor.trigger_enabled) and regenerates the capability registry and matrix. Verified RED on the remote runner before any implementation exists. * feat(#1953): complexity-triggered refactor extension point Adds the opt-in refactor-trigger capability. After a phase executes, an execute:post step measures per-function complexity for the files the phase touched and writes a scoped refactor proposal when a function crosses the configured threshold or drifts past its recorded anchor. Design notes worth carrying: - The signal is computed in-core (decision-point counting over comment- and literal-stripped source, Node builtins only) rather than via Memtrace or a shelled-out analyzer. The hook fires as a deterministic CLI, not an agent with MCP tools, and core takes no external dependencies — this is the only option a behavioral test can bind to. The metric sits behind a seam. - The baseline is a stable anchor, not a rolling value: set on first observation, moved only on disposition. A rolling baseline makes the delta the single-phase change, so a function creeping +2 per phase never trips a delta of 5 and the jump check adds nothing over the absolute threshold. - Strict mode records an open deviation window in the broken-windows ledger rather than declaring its own ship:pre gate. ship.md has no generic ship:pre gate dispatch — only two hardcoded branches — so a third gate of any kind would be declared and never evaluated. - The gate clears on the proposal being dispositioned, never on the score improving. A blocking complexity number is one an executor can satisfy by splitting a coherent function in two. execute-phase.md gains a generic execute:post step-dispatch contract; it previously matched only ref.skill == "code-review", so any other step registered there was declared and never run. The code-review branch is unchanged. Full rationale in ADR-1953. Closes #1953 * fix(#1953): close git option injection and symlink escape in the refactor hook Three findings from the isolated security review, all fixed inline. HIGH — changedFilesSince interpolated the --since value into a revision token placed before the -- separator. A -- only stops PATHSPEC parsing of arguments after it; git still option-parses what comes before. So --since '--output=/tmp/x' became --output=/tmp/x..HEAD, which git accepts as --output=<file> and uses to redirect diff output — an arbitrary write. Fixed with --end-of-options before the revision range plus a conservative ref validator. The validator deliberately permits ~ ^ @ { } because those are legitimate git REVISION syntax (HEAD~1, main@{yesterday}) as distinct from ref-NAME syntax; --end-of-options is the actual barrier. The doc comment asserting the trailing -- was sufficient was wrong and is corrected. MEDIUM — resolveConfinedPath confined by string prefix only, so a symlink committed inside the repo passed the check (its own path is under cwd) and readFileSync then followed it outside the root. Now lstat-checks for a regular file and skips anything else with REFACTOR_FILE_UNREADABLE, so one bad path skips one file and the run continues. LOW — the new execute:post dispatch contract showed the gsd_run example before the rule requiring ref.command be validated first. That prose is executed by an agent, so textual order is execution order. Reordered. Refs #1953 * fix(#1953): make the analyzer able to see TypeScript at all Found by running the shipped analyzer over its own source: it reported functions=1 for a 940-line module with 24 function forms. A return-type annotation or a generic parameter list made a function invisible — `function f(a): number {}` and `function f<T>(a: T): T {}` both detected as zero. Since gsd-core is written in .cts and the capability declares .ts/.cts/.mts analyzable, the feature silently found nothing in this repo's own primary language while reporting success. A safety net that reports "all clear" because it cannot see is worse than no safety net. All 98 tests passed over this, because every fixture was plain JS — the exact failure the test matrix's own "assert against the shape production uses" warning describes. Adds a TypeScript-shapes suite covering return types (including unions, generics, object literals and type predicates), generic parameter lists (constrained and defaulted), export/async/generator combinations, annotated arrows, class-method modifiers, and optional/ default/rest params — plus the two traps: an overload signature has no body and must not count, and `a < b && c > d` is a comparison, not a generic. Detection now reports 24/37/21 functions for the three source files, which matches a hand count exactly. Also from review: - The strict-mode ledger dedup identified entries by parsing a prose description string. That is banned by CONTRIBUTING's raw-text-matching rule and was a real bug: the "exactly one window per untriaged proposal" guarantee rested on prose matching, so rewording a description or editing WINDOWS.md by hand silently produced duplicates. Now matches structurally on kind + phase + file + line. - A property test asserted on the stripper's output text. Reframed to assert the same invariant through analyzeSource's score. - nextBaseline's `candidates` parameter has been dead since the anchor change; removed from the signature and all call sites. - Extracted the duplicated require-or-degrade and capability-check boilerplate. - ADR-1953's Implementation bullet still named a `refactor.ship-gate` in check-command-router.cts — a leftover from the design cut D6 rejects. That file is untouched and no such gate exists. Removed. Refs #1953 * fix(#1953): keep execute-phase.md under its byte ceiling; un-vacuum the large-file test Five of the seven remote-runner failures were one cause: the execute:post dispatch contract, written out inline, grew execute-phase.md 1876 bytes (93,400 -> 95,276) against a frozen PRE_PHASE6 ceiling of 93,600. A drift-ack does not clear that — tests/phase6-capstone-conformance.test.cjs and tests/fix-2285-claude-orchestration-wiring.test.cjs assert the file is literally under the cap. The contract now lives in gsd-core/references/loop-hook-dispatch.md, which already claimed to be the point-agnostic dispatch reference and already documented ref.skill and ref.agent. It gains the ref.command shape, its in-context validation rule, the advisory-by-construction statement, and a note that a point whose workflow hand-rolls one kind is not implementing this contract. execute-phase.md now defers to it in one line: 145 bytes of growth, 55 B of headroom under the cap. Better placement than the first cut — the reference was overstating its coverage, and this makes the claim true rather than duplicating prose next to it. Acknowledged by appending to tests/emitted-drift-acks/2930-*.json rather than a new 1953-*.json: two ack sources may never name the same path, and that fragment is already the accumulating ack for this file. Sixth and seventh failures: analyzesLargeFileWithinBounds tripped its own vacuity guard — the fixture generated ~480 KB against a `> 500000` assert, so the guard fired and the three assertions after it never ran. The test has been vacuous since it was written. The matrix row specifies ~1 MB, so N goes 8000 -> 20000 (1.17 MB, 17% margin) and the guard to > 1_000_000. Verified by reproducing the exact body against the compiled module: 1168888 bytes, 118 ms, all four assertions hold. Refs #1953 * fix(#1953): fold the execute:post step deferral into the existing resolve line The remaining two failures were one test: execute-phase.md carries a SECOND, tighter assertion than the 93,600 ceiling — `<=93400`, which is exactly its current size. The file cannot grow by a single byte. My previous fix got it under 93,600 but not under 93,400, so it still failed. ("H." in the report is just the parent describe of that same test, not a separate defect.) Rather than add a paragraph, the deferral now REPLACES the existing hook resolution line. It read: Resolve active step hooks from `EXECUTE_POST_HOOKS_JSON` where `kind == "step"` and `ref.skill == "code-review"`. which is the bug itself written down — only code-review was ever dispatched. It now reads: Dispatch each `kind == "step"` hook per @gsd-core/references/loop-hook-dispatch.md. For `code-review`: The following prose already begins "If no active code-review step hook exists", so it reads correctly and the code-review handling is untouched. Net effect on the file is -11 bytes: 93,400 -> 93,389, under the margin assertion rather than merely under the ceiling. That also removes the need for a drift-ack: the file shrank, so there is no growth to acknowledge, and the append to the shared 2930-*.json fragment is reverted. Leaving it would have shipped a claim of "145 bytes of growth" that is no longer true, on a file six other issues share. The test's own comment states the principle this ended up honoring: "the host loop must stay small — optional-feature detail belongs in the capability fragment, not the host workflow." Putting the dispatch contract in the reference rather than inline is that rule, applied. Refs #1953 * fix(#1953): keep the code-review hook literal the workflow test requires tests/code-review.test.cjs extracts the <step name="code_review_gate"> block and asserts it contains `ref.skill == "code-review"` verbatim. The previous commit replaced the line carrying that literal, so the token vanished and the test went red — a fair assertion: code-review IS the bespoke branch there and the workflow should still name it. Restored inside the same one-line deferral, which now reads: Dispatch `kind == "step"` hooks per @gsd-core/references/loop-hook-dispatch.md. `ref.skill == "code-review"`: 93,396 bytes — still under the `<=93400` margin assertion and 4 bytes below the base, so the file continues to shrink rather than grow. Because three consecutive runs were each reddened by a different assertion on this one file, this change was verified by sweeping ALL of them at once rather than one run at a time: every test under tests/ that reads execute-phase.md or references/loop-hook-dispatch.md was located by resolving its path constants, and each content/size assertion was evaluated directly against the working tree — 22 assertions, plus two real executions (gen-section-manifest --check, and emitted-attribution's full real-tree differential). All pass. That sweep also confirms the earlier judgement call: the net change to execute-phase.md is a SHRINK, and the size ratchet only gates growth, so reverting the append to the shared 2930-*.json ack fragment was correct — an ack would have been both unnecessary and factually wrong. Refs #1953 * chore(#1953): backfill changeset pr number to 3261 * docs(#1953): add the missing how-to for acting on a refactor proposal Reference and explanation shipped (COMMANDS.md, CONFIGURATION.md, FEATURES.md 159, ADR-1953) but the Diataxis how-to quadrant did not, and that is the one a user reaches for. CONTRIBUTING's required-docs table is 'new command -> COMMANDS.md + FEATURES.md', so CI was green on a gap. Enabling this feature is genuinely multi-step and no single page walked it: turn it on, tune the threshold, understand advisory vs strict, discover that strict needs a SECOND toggle on a DIFFERENT capability, and know what to do when a proposal appears. The two-toggle subtlety in particular was a footnote in a config table; here it is a section with both commands. Follows the shape of its closest siblings, resolve-edge-coverage-findings and resolve-prohibition-findings — both 'the loop surfaced a finding, here is what to do with it'. Includes a reason-code table for the silent cases, since the analyzer is deliberately quiet in six situations and a user who expected a proposal needs to tell 'nothing to report' from 'could not look'. Indexed from docs/README.md beside the other loop how-tos. Docs-only: exempt from the push gate, no re-verification, pass marker on 2af188b4 untouched. Refs #1953 * feat(#1953): warn when strict mode is on but nothing will actually block Closes acceptance criterion 5, which I had wrongly marked satisfied. refactor.trigger_strict records an untriaged proposal as an open deviation window, but a ship only STOPS if workflow.windows_enforce is also on — a toggle owned by the broken-windows capability that this feature neither sets nor requires. So a user could enable strict, believe ship was gated, and find out otherwise at ship time. The split itself stays: requires:["broken-windows"] would force-install the ledger on advisory users who never enable strict, and a ship:pre gate of our own would never fire because ship.md has no generic ship:pre gate dispatch. What was missing was discoverability, so that is what this fixes. `refactor evaluate` now emits a typed REFACTOR_STRICT_NOT_ENFORCING warning, naming the exact remediation command, whenever strict is on and either workflow.windows_enforce is off or broken-windows is unavailable. It fires only on a run that produced a candidate — with nothing to block on there is nothing to warn about, and warning every run would be noise. Reads workflow.windows_enforce through the same resolveConfigKey walk the router already uses for its own keys rather than a second config reader. Four tests cover the matrix: strict+enforce-off warns, strict+enforce-on does not, strict+ledger-absent warns, strict-off never warns. Also corrects a user-facing message in this same file that told the user to run `gsd-tools config-set` — the wrong form. docs/CONFIGURATION.md and the broken-windows capability both use `gsd config-set`, and gsd-tools is invoked as `node gsd-tools.cjs`, so the bare form may not resolve. The two adjacent messages in this file now agree. Refs #1953 --------- Co-authored-by: sim <sim@local>
9.8 KiB
ADR-1953: Complexity-triggered refactor — the loop measures the entropy it just added
- Status: Proposed
- Date: 2026-08-09
- Issue: #1953
- Implementation:
src/complexity-trigger.cts,src/refactor-trigger-command-router.cts,capabilities/refactor-trigger/capability.json, and a changed-files adapter insrc/git-base-branch.cts - Extends: ADR-857 (registers on the
execute:postextension point) · ADR-894 (arole: "feature"manifest withcommands,steps, and agatesentry) - Related: #1950 / the
broken-windowscapability — this ADR reuses its ledger rather than adding a second one
Context
The Pragmatic Programmer Topic 40 — "Refactor Early, Refactor Often" — argues for
refactoring as gardening: a little, continuously, because entropy compounds. GSD today has
refactoring only as a manual commit type (agents/gsd-executor.md, agents/gsd-planner.md):
an optional cleanup the executor may perform. Nothing measures accumulated complexity and
nothing triggers a refactor. It happens if, and only if, someone remembers.
That gap is worse under an AI executor than under a human one, and for a structural reason: every phase runs in a fresh context, so no single run ever sees the accumulated mess. Each phase adds a branch here and a special case there, every one locally justified by "make the tests pass". By the time a human notices, the hotspot is a rewrite.
The signal needed to close this is cheap and already sitting there: the loop knows exactly which files a phase touched, and it knows the moment the phase ends. What is missing is a trigger.
Decision
D1 — A capability on execute:post, not a change to the loop. refactor-trigger is a
role: "feature" capability, activationKey: refactor.trigger_enabled, default off.
Nothing about the loop changes for anyone who does not opt in.
D2 — The signal is computed in-core, with no new dependency. A decision-point counter
over comment- and literal-stripped source, Node builtins only, in the same regex idiom
src/intel.cts already uses for export extraction. This was chosen over the issue's original
"use Memtrace's complexity signals" and over shelling out to ESLint complexity / radon.
The reason is not that the alternatives are worse metrics — they are better ones. It is that
the execute:post hook fires as a deterministic CLI, not as an agent holding MCP tools,
and core forbids external dependencies. An agent-side Memtrace variant could not be bound by
a behavioral test, which is a hard requirement of the linked issue's own acceptance criteria.
The metric sits behind a named seam so a second one is additive later.
D3 — Two numbers, never one composite. A function is a candidate when its absolute score
exceeds refactor.complexity_threshold, or when its growth over its anchor exceeds
refactor.complexity_jump_delta. Both are reported. Trigger semantics are ESLint's
complexity: {max: N} semantics — strictly greater — so a score equal to the threshold does
not trigger.
D4 — The baseline is a stable anchor, not a rolling value. It is set the first time a function is observed and moves only when the proposal is dispositioned. The delta is therefore cumulative since the last conscious decision about that function.
A rolling baseline was tried first and discarded: with it, the delta is always the
single-phase change, so a function creeping +2 per phase against a delta of 5 never trips
the jump check, and the absolute threshold catches the creep first — leaving the jump-delta
contributing nothing. The stable anchor catches that creep a full phase earlier, which is the
entire reason the second number exists. The cost is that a legitimately-growing function
re-proposes until dispositioned.
D5 — Advisory by default; strict mode tracks on disposition, never on the score. This is
the load-bearing decision of the whole design. A blocking complexity number is a metric that
an executor optimizing for green gates can satisfy by splitting one coherent function into two
incoherent ones — identical total complexity, worse cohesion. So the tracked entry clears when
the proposal is dispositioned — refactor accept or refactor decline, either one — and
never asks whether the number went down.
D6 — This capability declares no gates. Strict mode routes through the broken-windows
ledger. An untriaged proposal becomes an open deviation entry via appendWindow;
blocking is broken-windows' existing, already-dispatched ship:pre gate, enabled separately
with workflow.windows_enforce. A declined proposal resolves its entry as waived with the
recorded reason; an accepted one resolves it as fixed. There is no second ledger and no
second gate.
Two earlier cuts of this decision were wrong and are recorded because the reason generalizes.
A check.predicate on the proposal artifact was killed by artifact-frontmatter-equals
mapping artifact not found to block: true — it would block every ship in which no
proposal was produced. A check.query was then killed by the test suite: ship.md has no
generic ship:pre gate dispatch at all, only two hardcoded branches (security at
ship.md:105, broken-windows at ship.md:155), so a third gate of either kind would be
declared and never evaluated — strict mode would silently do nothing. The structural guards
in tests/loop-hooks-ship-pre-e2e.test.cjs pin that reality rather than express a
preference. Generalized: at ship:pre, a capability cannot add a gate; it can only add a
window.
D7 — A declined refactor is recorded in the existing broken-windows ledger. The same
appendWindow path as D6, degrading to a note when that capability is absent. No second ledger.
D8 — execute:post gains a generic step-hook dispatch contract. execute-phase.md
dispatches execute:post step hooks only where ref.skill == "code-review", so a ref.command
step there is declared-but-never-run. The contract added mirrors the existing one in ship.md.
The code-review branch is left byte-identical; the generic contract handles the rest.
What stays OUTSIDE
- Choosing or applying the refactor. The capability surfaces a proposal. It never edits
code, never picks a strategy, never auto-applies. Executing an accepted proposal remains the
executor's existing
refactorcommit type. - Cognitive complexity, Halstead, maintainability index. One metric behind a seam. Adding a second is a later, additive decision, not a reason to widen this one.
- Non-JS/TS languages. Unsupported extensions are reported as skipped, never scored.
- Test files and generated output. Excluded by default; branchy tests are not a defect.
- Threading the touched-file set through
agents/gsd-executor.md(named in the issue's scope list). Deriving it from a boundedgit diffinside the CLI is deterministic and testable and avoids the agent-size ripple. - Repo-wide or scheduled scanning. The signal is strongest on exactly the files the phase just touched; a periodic full-repo scan is decoupled from the change that caused the growth.
Consequences
Good. Continuous, automatic refactoring pressure that does not depend on anyone remembering. Slow creep becomes visible a phase before the absolute threshold would catch it. Declined cleanups stay tracked instead of evaporating. Zero cost when off, and zero new dependencies when on.
Bad, and accepted. The metric is approximate by construction: biased against a flat
switch (SonarSource's well-known objection to raw cyclomatic complexity), blind to nesting
depth, and JS/TS-only. A rename reads as delete-plus-add and loses its anchor. Defaults will
need tuning — too low reads as refactor spam, which the linked issue names as the main
maintenance burden. The leak surface of a non-AST analyzer is real; it is handled by refusing
to emit a number rather than by emitting an approximate one.
A soft dependency, disclosed. Strict mode can only block when the broken-windows
capability is installed and workflow.windows_enforce is on — two toggles, not one. Without
it, strict mode still records the proposal and says so, but cannot stop a ship. requires: ["broken-windows"] was rejected because it would force-install the ledger on advisory users
who never enable strict mode.
Risk we are deliberately holding. D5 mitigates Goodhart exposure but does not eliminate it. A team that turns on strict mode and treats proposals as chores to clear will get worse code than one that leaves it advisory. The default and the documentation both push toward the latter.
Open questions
- Is 15 the right default threshold? SonarSource's default; ESLint's is 20; radon's rank C starts at 11. Needs field data.
- Should an accepted proposal auto-create a task in the next phase's plan, rather than relying on the developer to act? Deferred — it would couple this capability to the planner.
- Should cognitive complexity become the default once available, with cyclomatic as the fallback? The seam allows it; the evidence to decide does not exist yet.
References
- The Pragmatic Programmer, Topic 40 — "Refactoring".
- T. J. McCabe, "A Complexity Measure", IEEE TSE, 1976 —
V(G) = E − N + 2P. - G. Ann Campbell, "Cognitive Complexity" (SonarSource) — the readability critique of raw
cyclomatic complexity that motivates D2's seam and the flat-
switchcaveat. - ESLint
complexityrule — the source of D3's strictly-greatermaxsemantics. - Martin Fowler, Refactoring and "Code Smell" — a threshold crossing is a signal to look closer, not a verdict; the reason D5 surfaces rather than blocks.
docs/reference/capability-manifest.md,docs/reference/gate-predicates.md.