Commit Graph

22 Commits

Author SHA1 Message Date
Tom Boucher
adb46cdd85 feat(#2734): surface STATE.md commit-age on the statusline (#3700)
* test(#2734): failing-first suite for the statusline STATE.md freshness marker

Binds the contract before any hook change exists: a `state ~N commits back`
segment gated on the state_head stamp landed by #2622, firing at the same
advisory threshold /gsd-health's W024 uses rather than at > 0.

Covers all five acceptance criteria — threshold parity (19/20/21 boundaries),
both renderers including formatGsdStateCompact, an exact spawn-count assertion,
repo-pinning and sub_repos degradation, and behavioral parity against
readStateHeadFreshness rather than a source-grep of the two fence copies.

52 example-based tests plus 5 seeded fast-check properties. Red now by design.

* feat(#2734): surface STATE.md commit-age on the statusline

Adds an opt-in `state ~N commits back` marker to the GSD-state segment,
consuming the `state_head` stamp and freshness contract landed by #2622.
A solo developer returning to a project reads "Phase 4, executing" in
STATE.md and acts on it, without noticing the codebase moved 40 commits
since that line was written. /gsd-health reports it as W024, but only if
you think to run it; the statusline is the surface you see without asking.

Fires at STATE_HEAD_ADVISORY_COMMITS (20), the same threshold W024 uses,
not at > 0: with commit_docs:true the commit carrying a STATE.md sync
advances HEAD by one, so > 0 would alarm permanently on a fresh project.

Costs exactly one bounded git subprocess per render and none when
disabled. `rev-list --left-right --count` answers ancestry and distance
together, and repo pinning is a filesystem check mirroring
projectOwnsItsRepo rather than a --show-toplevel compare, which is
unreliable on macOS /private/var and Windows 8.3 paths.

Every unresolvable input degrades to the tri-state unknown -- the marker
is absent, never a "fresh" claim the project cannot substantiate: a
malformed stamp, a root that does not own its .git, a sub_repos
workspace, history rewound past the stamp, or git being unavailable.

Also collapses statusline config resolution onto one resolveStatuslineOptions()
seam. runStatusline() and renderStatusline() duplicated it byte-for-byte;
one copy is what keeps a newly-added key from reaching only one of them.

* test(#2734): route the e2e spawn through the process seam and fix fixture leaks

Review findings from the two orthogonal passes:

- `bothEntryPointsResolveOptionsIdentically` spawned a child and substring-matched
  its stdout to test a pure function. It now calls resolveStatuslineOptions()
  directly — no subprocess, no text matching.
- `skipsFreshnessWorkWhenTodoTaskActive` genuinely needs a child (the !task gate
  lives in runStatusline, which reads stdin), so it now spawns through
  tests/helpers/process-seam.cjs and proves the negative with a filesystem fact:
  the git shim appends to a marker file on every invocation, and the assertion is
  that the marker never appears. Stronger than asserting text is missing, and it
  drops the last stdout substring match in the block.
- Every fixture-creating test now registers `t.after(() => cleanup(dir))` instead
  of a trailing cleanup(dir), which leaked the temp repo on assertion failure.
  derivationAgreesWithStateModule reassigns `dir` across five fixtures, so it
  binds each directory at scheduling time rather than cleaning only the last.

Also corrects markerCoexistsWithMilestoneComplete, which asserted the wrong
expectation rather than finding a code defect: `percent` drives the progress bar
too, so the milestone segment reads "v1.9 [##########] 100%". The marker appends
after it, which is what the test exists to prove.

CONTEXT.md's opt-in statusline key list was missing statusline.show_git as well
as the new key; both are now enumerated.

* docs(#2734): backfill changeset PR number (#3700)

---------

Co-authored-by: sim <sim@local>
2026-08-20 00:35:01 -04:00
Tom Boucher
bf2332e67c fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint

On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately
absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw
tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before
requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in
its fail-closed catch and is misreported as an unreadable dispatch-isolation
configuration, blocking every executor dispatch.

These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor
guard and update worker, plus the seam's actionable build error surfacing instead of
the generic misreport.

Also adds the drift lint the acceptance criteria require, with a fixture proving it
CAN fail — a guard never shown to fail is worthless. It is red here by design: it
flags today's unfixed hooks, which is exactly the defect.

* fix(3582): route every hook's compiled-module require through the self-heal seam

RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard,
cursor guard and update worker, the fail-closed-with-actionable-message assertion, and
the lint's own real-tree check.

The compiled runtime library is produced by build:lib and gitignored (ADR-457,
build-at-publish). The npm package builds before publishing; a plugin-marketplace or
git-clone install materializes the raw tree and never does, so on that channel those
modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly
this, and the CLI entrypoint already calls it — no hook did. The isolation guard's
Cannot-find-module therefore landed in its fail-closed catch and was reported as
'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was
misdiagnosed as an unreadable project config and every executor dispatch was blocked.

All SEVEN affected files now call the seam before their first compiled require. The issue
named four; a scan found six; implementing it surfaced a seventh — the shared isolation
sentinel helper, used by BOTH guards, which requires two compiled modules itself and
would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so
fixed here rather than left as a known-broken remainder.

Failure posture is deliberately split by hook kind:
- Gates (agent isolation guard, cursor subagent start) surface the seam's actionable
  build error distinctly instead of swallowing it into the generic text, and stay
  fail-closed — a genuinely unreadable project config still DENIES exactly as before.
- Cosmetic and detached hooks (statusline, update worker, update check, update banner)
  DEGRADE rather than crash: the statusline draws on every render and the worker is a
  detached process, so a build failure there must not take down the prompt.

The npm path is untouched: the seam's already-built fast path returns immediately, so
prebuilt installs pay nothing and behave bit-for-bit as before.

Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than
remembered — without it the next hook to add a compiled require reintroduces the class
silently. It is proven able to fail: a fixture hook requiring a compiled module without
the seam is flagged, and one that uses the seam is not. Verified directly — on the
unfixed tree it named all seven offenders; with the fix it passes.

While writing the lint's comment stripper, a naive whole-text block-comment regex ate its
own fixture, because this repo's comments legitimately spell the compiled-lib glob whose
star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression
test pinning that case.

* fix(3582): test the three untested seam call sites and assert typed reason codes

Two independent reviews converged on the same major gap: the fix wired the seam into
seven files but only four had cold-tree tests. The adversarial pass put it plainly —
deleting the shared isolation-sentinel helper's seam call would not have failed any test
in the diff. That file was my own addition beyond the issue's four, so it shipped
untested; that is now closed.

- Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT
  directly under cwd, and every existing cold-tree fixture puts it there, so the early
  return always fired first. Now covered, and proven load-bearing by mutation: with the
  call removed the spy records zero seam invocations and the test fails.
- update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED
  VERDICT — the fallback cache filename, and silent suppression when the package name
  degrades to null — rather than merely 'did not throw'. The banner hook previously had
  no test file at all.

Standards violation fixed: two tests asserted on free-form prose via assert.match against
a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly
that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch
it. Both isolation guards now emit a machine-readable reason_code from a frozen enum,
following the repo's existing REASON convention, and the tests assert that instead. The
human-readable message is unchanged for operators; only the assertion target moved.

The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT
extracted, and the reason is recorded at each site: both viable shapes — a
path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file
literal co-occurrence check, so extracting would require the lint to special-case its own
helper. Triplication is the lesser evil while the lint stays a co-occurrence scan.

The lint's header now states what it does and does not catch (literal quoted requires
only; hooks/ scan root), so a future reader does not over-trust a guard that a
concatenated path or a require inside a non-hooks helper would evade.

* chore(3582): regenerate the committed install-tree fixtures

Adding a new shipped hook helper changed the install tree, and those fixtures are
committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>'
tests failed on 541a1913. Regenerated rather than hand-edited.

The delta across all 15 runtime fixtures is exactly two lines — the new helper under both
its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in
no unrelated drift.

This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from
lint:ci, which passed both before and after.

* chore(3582): backfill changeset PR number (#3629)

---------

Co-authored-by: sim <sim@local>
2026-08-18 14:11:23 -04:00
Tom Boucher
aee83c8e9e test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5 (#3378)
* test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5

Folds 6 legacy issue-*.test.cjs regression files (120 test() blocks) into
their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053).
Second of 4 issue-* waves.

- issue-766-plugin-manifest.test.cjs (50 tests): pure rename (git mv) into
  plugin-manifest.test.cjs, sole comprehensive suite for its module.
- issue-844-manifest-version-sync.test.cjs (14 tests, rename basis) +
  issue-1855-marketplace-manifest.test.cjs (17 tests, merged in): both
  target scripts/sync-manifest-versions.cjs via non-overlapping describe
  blocks (generic engine vs. marketplace.json-specific), now
  manifest-version-sync.test.cjs, 31 tests, 0 dropped.
- issue-498-identity-drift-lint.test.cjs (8 tests, rename basis) +
  issue-498-package-identity.test.cjs (17 tests, merged in): both target
  package-identity.cjs's surface from different angles (lint-side drift
  detection vs. derive/slugify), now package-identity.test.cjs, 25 tests,
  0 dropped.
- issue-607-cache-lineage.test.cjs (14 tests) merged into the existing
  gsd-statusline.test.cjs, matching its already-established fold-wrapper
  convention from prior waves.

Learned from Wave 4: fold agents checked src/** (not just docs/ and
gsd-core/references/) for stale filename references — found and fixed 5
across 3 ADR docs (766, 2121, 457), zero in src/ this time.

Zero net test-coverage loss. No production code changed.

* test(#3337): fix orthogonal-review findings — Wave 5 fold

Standards-axis review found real issues in the just-folded
manifest-version-sync.test.cjs, both fixed here:

- The folded:issue-1855-marketplace-manifest block omitted its own
  test/describe/assert/fs/path/os requires, silently closing over the
  #844 basis section's outer-scope bindings instead of declaring its own
  dependencies — inconsistent with every other fold block in this wave.
  Restored the local requires the original issue-1855 file declared.
- Merging two independently-lettered legacy files (each A-F) left
  duplicate top-level describe() labels (two "A:", two "B:", two "C:").
  Renamed the folded-in #1855 labels (A2/B4/C2) to make every top-level
  label in the file unique.

No test() count changed (31). No production code touched.

---------

Co-authored-by: sim <sim@local>
2026-08-11 23:55:11 -04:00
Tom Boucher
9faacc0c15 test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist

Migrates the final 170 unbounded sync spawn sites across 49 files, then
removes the allowlist entirely. local/no-unbounded-spawn now runs with no
exemption surface across tests/**: there is no file to add a name to.

drift-detection's throw-native git() helper routes to gitOrThrow -- bare
runGit would have taken 16 call sites quiet on failure. commands.test.cjs
has two independently-scoped runGsdTools/runCli helpers, one already bounded
and one not; they are kept distinct rather than unified, the same trap as the
two same-named git() helpers in Wave 1.

runNpm's bound was erasable. Its options spread callerOptions after the
defaults, so an explicit timeout:undefined silently dropped the 180000ms
bound -- the rule flagged it and was right; it was not a false positive. Fixed
by destructuring with a default, with a test that fails when the default is
removed.

Two sites stay on a raw spawn with an explicit timeout because the seam
cannot express them: one needs shell:true for npm.cmd on Windows, one
redirects stdout to a real fd. Both are the rule's own documented second
option, not an escape from it.

Closure verified rather than asserted: the derivation scan reports 0 unbounded
spawn helpers and 0 unbounded direct git call sites, and a temporary file
carrying an unbounded spawn still errors with the allowlist gone.

Closes #3064.

* test(#3148): close a hole in the guard's own eslint-disable ban

The ban listed only the top level of tests/, so it was blind to 37 .cjs
files under tests/helpers, qa, observability, fixtures and dispatch. With the
allowlist deleted this test is the sole remaining way to detect someone
silencing the rule inline, so the gap was load-bearing: a nested file could
carry an unbounded spawn plus an eslint-disable and pass everything.

Proven before and after. A probe planted under tests/helpers with both was
invisible to the guard and clean under eslint; after making the listing
recursive the guard fails on it. The scanned set goes from 771 files to 808.

Pre-existing since the guard shipped, but this wave is what promoted it to
sole defense, so it is fixed here rather than filed.

Also converts the last hand-rolled throw check to throwIfFailed and the last
re-derived legacy shape to compose toLegacyResult, which makes the epic's
none-remain claim true rather than nearly true. toLegacyResult itself is not
widened -- eight callers depend on its shape and one consumer does not
justify changing a shared contract.

* fix(#3148): correct seam incoherence at the bound and a slow review-lane error path

Two real failures from the remote runner, both fixed at the cause.

The seam could return outcome TIMED_OUT together with exitCode 0. At the
exact bound spawnSync reports ETIMEDOUT while the child has already exited
with a real status, and toSeamResult classified on the error code while
passing status straight through -- an incoherent pair its own boundary test
was written to catch, and did. A status that is not null is direct evidence
the child exited on its own, so it now decides the outcome before the
error-code branches run. process-seam.cjs was deliberately untouched by every
earlier wave; this is a defect in the module itself, kept surgical, with a
unit test that fails against the old logic.

review-lane with an unknown subcommand fell through to its usage error only
after loading the capability registry and building a per-lane plan, which
spawns one child process per lane -- up to twelve. The error path took
~1288ms instead of ~119ms, and under bench load it outran a caller's spawn
timeout and was killed before writing anything, which is the empty stdout and
stderr CI saw. It now fails fast before any of that work begins.

This is the epic's first production change. It is user-facing, so it carries
a changeset rather than a no-changelog label.

* test(#3148): replace a real-race timeout test with a deterministic one

E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a
warm container git finishes first, spawnSync returns status 0 with no error
at all, the seam correctly classifies EXITED, and gitOrThrow correctly does
not throw -- so the test failed on both lanes. A probe confirms a genuine
timeout always carries status null, so this was never the seam misbehaving.

Raising the bound would only lengthen the odds, which is the same defect with
better luck. The test now drives gitOrThrow against a stubbed runGit that
returns a synthetic TIMED_OUT result, so it asserts exactly what it always
meant to -- that a timeout propagates as a throw -- with no timing
dependence. Five consecutive runs are identical where the old one varied.

I wrote this test in Wave 0; it is a real-race test by construction and
CLAUDE.md says to replace those rather than re-run them.

* chore(#3148): backfill changeset PR number 3192

---------

Co-authored-by: sim <sim@local>
2026-08-07 21:03:50 -04:00
Tom Boucher
5fd5c81042 test(#3055): add the process seam so a subprocess timeout is expressible as data (#3066)
* test(#3055): add the process seam and route runGsdTools through it

Adds tests/helpers/process-seam.cjs — runNode/runGit/runHook over spawnSync,
each returning a typed discriminated union
{ outcome, exitCode, stdout, stderr, timedOut, signal, killed, code }.
Every call is timeout-bounded; there is no unbounded path.

runGsdTools becomes an adapter over the seam. Its legacy
{ success, output, error, exitCode } shape and retry-once-on-kill behaviour
are preserved byte-identically, so none of its 136 caller files change.

Outcome discrimination was corrected against probed runtime behaviour rather
than assumption: a timeout and a maxBuffer overflow are identical on both
status (null) and signal (SIGTERM), and differ only by code (ETIMEDOUT vs
ENOBUFS). Overflow is therefore classified before timeout. This fixes a live
defect — the previous isKilled() treated an overflow as a kill, retried it for
a second full 60s run, and then reported "host OOM or scheduler contention"
for a child that had merely printed too much.

Also widens the ESLint tests glob from tests/**/*.test.cjs to tests/**/*.cjs,
which brought 31 previously unlinted shared helpers under the same rules their
sibling test files already obey, and fixes the 5 violations that surfaced —
including a bare npm invocation without shell:true in
tests/helpers/emitted-runtime.cjs (DEFECT.WINDOWS-TEST-PORTABILITY), now
routed through the existing portable runNpm helper.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3055): migrate every local spawn wrapper onto the process seam

Replaces the spawn body of all 25 local runHook/runGuard/runGate definitions
with a call to tests/helpers/process-seam.cjs. Each wrapper keeps its name,
parameter list, return shape and post-processing (JSON parse, ANSI strip, env
sanitising, field extraction) — only the spawn mechanism changes, so no test
assertion moves.

The 4 bash-driven wrappers use the seam's explicit `interpreter` option rather
than a fourth primitive; it is explicit rather than inferred from the file
extension, because guessing an interpreter from a path fails silently when a
script's name does not match its shebang.

Seven wrappers were previously unbounded and now carry an explicit timeout
sized to what each actually runs, not the seam default. Two of those seven
(gsd-write-guard, lint-docs-command-form) were absent from the issue's
inventory entirely and were found by scanning after the migration.

Adds the CONTEXT.md `### Process seam` glossary entry and a CONTRIBUTING.md
reference section covering the three primitives, the discriminated union, and
the two rules the seam enforces.

Scope disclosure recorded in the phase design notes: the issue scoped three
identifier names. A scan for local helpers that spawn AND return the spawn
result finds 113 across 82 names, 71 of them unbounded, plus 122 unbounded
direct git call sites. This change bounds 25 of those. The remaining surface
is the same defect class and is NOT closed by this PR.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): classify an externally-killed child as KILLED, not EXITED

Blocker found in this branch's own diff, independently confirmed by an
isolated reviewer.

A child killed by an external signal — a genuine bench OOM kill — makes
spawnSync return { status: null, signal: 'SIGKILL' } with NO .error field.
The seam's "no error implies EXITED" rule therefore classified it as a clean
exit, and runGsdTools returned { success: false, exitCode: 1 } without
retrying. That silently defeated the #969 kill-discrimination for precisely
the case it was built for: the old isKilled() fired on `signal != null`,
retried once, then threw a labelled resource-starvation error. A real OOM
would have been reported as an ordinary assertion failure.

Adds a fifth outcome, KILLED, for "no error but a signal is set", and makes
the adapter retry on TIMED_OUT or KILLED — reproducing the old
`killed || signal != null || code === 'ETIMEDOUT'` condition exactly.
SPAWN_FAILED still does not retry (matching the old behaviour, where signal
was null). BUFFER_OVERFLOW still does not retry, which remains a deliberate
divergence: the old code retried it because signal was SIGTERM, burning a
second 60s run on a child that had merely printed too much.

All five outcomes verified against the live runtime rather than assumed:
SIGKILL -> killed, exit 0/7 -> exited, timeout -> timed_out (ETIMEDOUT),
>1MB stdout -> buffer_overflow (ENOBUFS).

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): address standards-review findings on this branch

Three findings from the standards axis of the review, all in this branch's
own diff.

The CONTEXT.md glossary entry this branch introduced was already stale on the
branch's own last commit: it enumerated a 4-member OUTCOME while the code had
5, because the KILLED fix did not update it. That is precisely the drift the
"module changes update Domain-terms" gate exists to catch, so the entry now
lists all five and explains KILLED.

api-coverage-gate-e2e compared an outcome against the raw string 'exited'
rather than OUTCOME.EXITED, the only such outlier; the enum is now imported
and used. A sweep for the other four outcome literals found no further
comparison sites.

Three call sites hand the literal bash flag '-c' to the seam's first
parameter, which the JSDoc described as an absolute script path. Rather than
add a fourth primitive, the contract is corrected to match reality: the
parameter is renamed `target` and documented as the first argv element handed
to the interpreter — normally a script path, but for an interpreter invoked
with an inline program it may be that interpreter's own flag. No behaviour
change.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(#3055): assert the cross-platform timeout contract, not the macOS one

The remote runner failed on both Linux lanes (node 22 and node 24, identical)
while the same tests passed locally on macOS. Two assertions encoded a
platform-specific behaviour as a cross-platform guarantee.

When spawnSync times out, macOS preserves the child's partial stdout/stderr;
Linux discards it and returns empty strings. Verified on node v26.5.1 both
ways. The seam passes through whatever spawnSync hands it and cannot
manufacture output that was discarded, so the production code was correct —
the tests were wrong.

Both tests now assert the guarantee the seam actually makes on every
platform: outcome TIMED_OUT, timedOut true, and stdout/stderr always being
strings rather than undefined or a Buffer. The partial-content assertions are
retained behind an explicit process.platform === 'darwin' guard so the macOS
coverage is not lost, and the first test is renamed to say what it now
guarantees.

This is the failure mode the remote matrix exists to catch: local macOS
verification would have shipped it.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#3055): classify a failed spawn as SPAWN_FAILED, not a timeout

Windows CI caught two defects the Linux matrix could not.

tests/context-predicates-query.test.cjs passes a 32K-char argv value. On
Windows that exceeds the argv limit and spawnSync fails with code
ENAMETOOLONG, signal null, status null. The seam's fallback rule — "otherwise,
status === null implies TIMED_OUT" — swallowed it, so the adapter retried a
spawn that can never succeed and then threw the resource-starvation error. The
old isKilled() returned false for that shape and returned an ordinary failure
result.

TIMED_OUT is now identified positively: code === 'ETIMEDOUT' OR signal is set.
Anything else carrying an error is SPAWN_FAILED, which covers ENAMETOOLONG,
E2BIG, EACCES and ENOENT alike. The signal clause is what keeps a platform
whose timeout errno differs classified correctly, so the greedy catch-all is no
longer needed.

The second defect is a contract regression I introduced and had claimed
otherwise. That same test asserts `typeof r.exitCode === 'number'`, and
toLegacyShape was returning null for BUFFER_OVERFLOW and SPAWN_FAILED, so the
assertion failed on type. The old code returned `err.status ?? 1` on every
non-retried failure path. The adapter now returns 1 again for both, and the
comment claiming "never coerced to exitCode:1, unlike the pre-seam helper" is
retracted: the seam keeps the richer truth (exitCode null plus a distinct
outcome), the legacy adapter keeps the old numeric contract its callers
actually depend on.

Verified on this host: a 4MB argv yields E2BIG -> SPAWN_FAILED; ENOENT ->
SPAWN_FAILED; timeout -> TIMED_OUT; >1MB stdout -> BUFFER_OVERFLOW; SIGKILL ->
KILLED; clean exit -> EXITED.

Refs #3051

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-05 01:06:39 -04:00
Tom Boucher
fd07e1a357 fix(#2850): resolve the active workstream in the statusline GSD-state segment (#3012)
* test(#2850): add failing-first tests for workstream statusline state

readGsdState only ever reads the flat .planning/STATE.md via a directory
walk-up; it has no path for .planning/workstreams/<ws>/STATE.md and never
consults GSD_WORKSTREAM or the stored active-workstream pointer, so the
GSD-state segment silently disappears in workstream mode. These tests
prove the RED before the fix lands.

Uses shared saveSessionEnv/restoreSessionEnv/clearSessionEnv helpers now
added to tests/helpers.cjs (single source of truth for the session-env-var
save/clear/restore pattern also used by tests/active-workstream-store.unit.test.cjs,
which is updated here to consume the same shared helpers instead of its own
local, already-diverged copy).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#2850): resolve active workstream in the statusline

readGsdState only ever walked up looking for a flat .planning/STATE.md; it
had no branch for .planning/workstreams/<ws>/STATE.md and never consulted
GSD_WORKSTREAM or the stored active-workstream pointer, so the GSD-state
segment silently vanished in workstream-mode projects with no root
STATE.md (exit 0, no diagnostic).

Reuses the existing CLI>env>store resolution seam (resolveActiveWorkstream,
active-workstream-store.cts) and the existing mode-detection/path-building
seam (listAvailableWorkstreams/planningPaths, planning-workspace.cts)
rather than re-implementing either inline. When workstream mode is
detected but nothing resolves, readGsdState now returns a
{noActiveWorkstream:true} sentinel that formatGsdState/formatGsdStateCompact
render as "no active workstream" -- observable, never silent emptiness.
Flat-mode behavior and the case where a resolved workstream has no
STATE.md yet are both unchanged.

Adds active-workstream-store.cts's peekActiveWorkstream: a read-only
sibling of getActiveWorkstream. resolveActiveWorkstream's default store
lookup self-heals a stale/invalid pointer by deleting it
(adapter.clear()) -- correct for a command, but not for a renderer
invoked once per prompt, which must never mutate persistent, possibly
cross-session state as a side effect of drawing a screen. The statusline
now injects peekActiveWorkstream via resolveActiveWorkstream's own
getStored override, keeping the env>store precedence itself fully reused
while removing only the store tier's write side effect. This satisfies
the issue's AC4 ("the fix is purely additive to what's displayed").

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#2850): backfill changeset PR number to 3012

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 21:44:04 -04:00
Tom Boucher
ef7a71fc09 fix(#2754): make parseStateMd CRLF-safe — frontmatter no longer drops on Windows-authored STATE.md (#2865)
* test(#2754): parseStateMd must parse CRLF STATE.md identically to LF

The frontmatter fence regex and downstream splits used literal \n, so a CRLF
STATE.md dropped the ENTIRE frontmatter block. Assert CRLF/LF parity (the
production contract) across full frontmatter, next_phases flow + block forms,
the progress nested block, and null handling.

* fix(#2754): make parseStateMd CRLF-safe — frontmatter fence + splits use \r?\n

The fence regex, scalar-line split, next_phases block-list regex, and progress
block regex all used literal \n, so a CRLF (Windows-authored) STATE.md dropped
the ENTIRE frontmatter block — every statusline field was silently absent. Use
\r?\n throughout, mirroring the CRLF-safe extractFrontmatter in src/frontmatter.cts.
LF behavior unchanged.

* chore(#2754): changeset fragment

* test(#2754): pin parseStateMd↔extractFrontmatter parity (Generative-Fix Divergence guard)

The statusline parseStateMd and the canonical extractFrontmatter both derive GSD
state from STATE.md frontmatter and diverged once already (the CRLF bug this PR
fixes). Add a cross-parser parity assertion (CLAUDE.md parallel-surfaces rule)
over the overlapping fields, under both LF and CRLF, so a future divergence on a
scalar shape or line ending is caught here. (isolated-review minor finding)

* chore(#2754): backfill changeset PR number (2865)

---------

Co-authored-by: Test <test@example.com>
2026-07-30 14:34:28 -04:00
Cody Anderson
20ff405cb3 feat(#2162): opt-in compact GSD-state format for the statusline (#2175)
* feat(#2162): opt-in compact GSD-state format for the statusline

New statusline.state_format config, enum full|compact (default full —
existing rendering untouched). "compact" renders the state segment as
"<version> · P<phase>/<total> · <status>", e.g. "v1.12 · P7/12 ·
executing" — dropping the milestone name and progress bar (the two
biggest width costs) and collapsing narrative statuses to a single
keyword. Per the #2162 approval conditions, the keyword set is the
canonical vocabulary from normalizeStateStatus() in state-document.cjs
(discussing/planning/executing/verifying/completed/paused) — no
parallel hand-rolled list, so the vocabularies can't drift — and the
canonical stuck state "paused" renders uppercase as PAUSED (no new
"blocked" lifecycle state). Statuses the normalizer passes through
unrecognized fall back to their first word capped at 16 chars.
Lifecycle scenes preserved: active_phase wins over the body phase
number, milestone completion renders "complete", idle-with-next-action
renders "next <action> <phases>".

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2162): changeset fragment for PR #2175

* fix(#2162): review fixes — ENUM_KEYS coverage, cap boundary tests, changeset format

- register statusline.state_format in the fix-1628 coercion-bypass matrix
- 15/16/17-char boundary tests for the shortGsdStatus fallback cap
- changeset body ends with the (#2162) citation per house convention

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2162): round-2 review fixes — scene exclusivity, direct config-set coverage

- compact renderer gates the milestone-complete scene behind the absence of
  an in-flight phase id, mirroring formatGsdState's if/else precedence
  (Scene 1 beats Scene 3); regression test covers the non-atomic
  active_phase + percent=100 STATE.md shape
- direct config-set accept/reject test for statusline.state_format plain
  strings (ENUM_KEYS matrix covers only the JSON coercion shapes)

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the statusline hook change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2162): complete-scene gate matches formatGsdState exactly (+property tests)

Re-review Major: gating done on !phaseId held completion back for the
legacy phaseNum shape — formatGsdState reaches Scene 3 on percent=100
regardless of phaseNum, so compact must too. Gate is now !s.activePhase.
The phaseNum-only test now expects 'complete' and cross-checks the full
renderer; a parity test feeds identical inputs to both renderers.
Re-review Minor: shortGsdStatus gets fast-check property coverage
(totality, canonical fixed points, separator safety, fallback shape).
Golden fixtures regenerated for the hook byte change.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-14 21:09:24 -04:00
Cody Anderson
eec9efc351 feat(#2160): collapse verbose '(1M context)' model suffix to compact (1M) badge (#2173)
* feat(#2160): collapse verbose '(1M context)' model suffix to compact (1M) badge

Claude Code appends " (1M context)" to the model display name in
long-context sessions, eating 12 characters of statusline width. Collapse
it to " (1M)" — the signal stays, the width doesn't. Any other display
name passes through unchanged.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2160): changeset fragment for PR #2173

* fix(#2160): review fixes — ctx variant, boundary test, changeset format

- broaden the suffix match with a context|ctx alternation (approval-condition
  variant the regex missed)
- pin non-context parentheticals ((beta), (deprecated)) as untouched
- changeset body ends with the (#2160) citation per house convention

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the statusline hook change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-14 20:37:43 -04:00
Cody Anderson
73039d3717 fix(#2163): round-2 review fixes — deterministic fail-soft injection tests
- readGitStatus calls execFileSync via the child_process namespace so tests
  can inject spawn failures through the shared module object
- two deterministic tests: ERR_CHILD_PROCESS_STDOUT_MAXBUFFER-shaped and
  ETIMEDOUT-shaped throws both degrade to null (segment absent), proving
  the fail-soft paths the PR previously only asserted

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-13 11:33:43 -06:00
Cody Anderson
d6672ff926 feat(#2163): opt-in git branch/status segment in the statusline
New statusline.show_git config (default false). When enabled, a git
segment renders after the directory: current branch plus compact
work-state markers (+staged ~unstaged ?untracked ↑ahead ↓behind, or ✓
when clean and in sync), e.g. " │ main+2~1?3".

One git status --porcelain=v2 --branch spawn per render via execFileSync
with a fixed argument array (no shell), a 1.5s timeout, and the
workspace dir passed with -C. Fails silently — segment absent outside a
repo, without git, or on timeout. Default output is unchanged when the
flag is absent.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg
2026-07-13 11:33:32 -06:00
Cody Anderson
ad7111e50b feat(#2161): opt-in absolute token count on the statusline context meter (#2174)
* feat(#2161): opt-in absolute token count on the statusline context meter

New statusline.show_context_tokens config (default false). When enabled,
the context meter shows the absolute token total after the percentage,
e.g. "████░░░░░░ 46% (156k)" — summing input, cache-creation, cache-read,
and output tokens from context_window.current_usage (matching /context).

Default output is byte-for-byte unchanged when the flag is absent or
false. The .planning config is now read once per render and shared with
the last-command/position block instead of being re-read.

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* docs(#2161): changeset fragment for PR #2174

* fix(#2161): review fixes — k-to-M threshold, boundary tests, changeset format

- formatTokens promotes to the M branch when k-rounding reaches 1000
  (999,500-999,999 rendered "1000k" instead of "1.0M")
- boundary tests at 999499/999500/999999/1000000/1000001
- Number() guards on the four usage fields (silent string-concat gap)
- changeset body ends with the (#2161) citation per house convention

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* fix(#2161): round-2 review fixes — config-set coverage, precision claim, exports style

- config-set accept/reject tests for statusline.show_context_tokens
  (mirrors the post-planning-gaps precedent the issue scope names)
- changeset + docs no longer claim parity with /context: the suffix sums
  four fields while the meter %% derives from used_percentage (three), so
  the figures can diverge slightly
- module.exports one entry per line

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

* test: regenerate golden-install-parity fixtures for the statusline hook change

Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg

---------

Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
2026-07-13 13:25:47 -04:00
Tom Boucher
0cc7a1a426 test(#1974): consolidate 27 installer/hooks remainder tests into module suites
Fold 27 issue-named installer/hooks/statusline/migration/reapply regression files
into their canonical module suites (installer-migrations, installer-migration-report,
gsd-statusline, reapply-verify-hunks, install-*, gsd-check-update-worker-platform-gate,
etc.). Verbatim block-scoped describe wrappers; 276 subtests conserved 1:1. No new files.

The one subdir origin (tests/installer-migrations/001-legacy-orphan-files) moved up one
level into installer-migrations.test.cjs; its single ../../ module require corrected to
../ so it resolves from tests/ root (verified). Host-env pre-check: no CLI-receiving host
sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value.

Regenerates regression-name allowlist (222->205), ratchets file-count allowlist (verify
11->8, validate entry removed), makes 16 relocated allow-test-rule exemptions issue-ref-
compliant (ADR-456; prunes stale ids). Repoints 15 tests/ references across state-md.md
(EN + ja/ko/pt/zh). lint:ci green.

Part of epic #1969. Closes #1974.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 09:34:39 -04:00
Tom Boucher
d444864bf8 fix(#1194): correct inverted statusline auto-compact buffer math (#1211)
* fix(#1194): correct inverted statusline auto-compact buffer math

The reserved-buffer percentage was computed as (acw/totalCtx)*100 — the
usable fraction — instead of (1 - acw/totalCtx)*100 — the reserved fraction.
When acw == totalCtx this produced buffer=100%, making the usable-range
denominator zero and pinning `used` at a constant 100% regardless of real
remaining context.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore: backfill changeset PR number (#1211)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 10:46:00 -04:00
Tom Boucher
bf48127c13 chore(lint): justify intentional no-control-regex (ANSI strip) + ratchet to error (#737)
The 5 no-control-regex warnings are all the same intentional ANSI-color-strip
pattern /\x1b\[[0-9;]*m/g across 5 test files. The \x1b (ESC) control char is
the required leading byte of an ANSI SGR sequence, so matching it is the whole
point of stripping color codes from captured CLI/console output. Add an inline
eslint-disable-next-line with justification at each site (not a refactor — the
control char is essential, not accidental), then flip no-control-regex from
warn to error so the debt can't regrow.

Refs #736

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-06 12:04:18 -04:00
Tom Boucher
a28dcec981 chore(#597): replace count-based ratchet guards with AST lint + named-set allowlists (#603)
The windows-test-parity ratchet greps test source for fs.rmSync-without-
maxRetries (and six other Windows-portability anti-patterns), failing when an
integer offender COUNT exceeds a frozen baseline (rmSync: 95). A count ratchet
is a Goodhart metric: fixing one offender and adding another keeps the count
constant, so a new defect slips through green. Replace it — and every other
count ratchet in the repo — with a layered, masking-proof design.

Behavioral seam test
- tests/helpers-cleanup.test.cjs proves helpers.cleanup() carries the Windows
  EBUSY retry budget. cleanup() delegates retries to Node's fs.rmSync via
  maxRetries (it owns no loop), so the test asserts the option contract
  (recursive/force/maxRetries>0/retryDelay>0) + real-FS removal + the cwd-guard,
  rather than a loop that does not exist. The EBUSY risk is now tested ONCE at
  the helper, not approximated textually at every call site.

Write-time ESLint rule (AST-accurate, replaces the grep)
- eslint-rules/no-raw-rmsync-in-tests.cjs (error in tests/**/*.test.cjs) bans
  raw fs.rmSync, steering to cleanup(). Catches member, computed (fs['rmSync']),
  destructured and aliased forms; escape hatch is inline
  `// eslint-disable-next-line local/no-raw-rmsync-in-tests -- <reason>` only.
- Migrated 336 raw fs.rmSync teardown calls across ~116 test files to cleanup().
  ~18 genuinely load-bearing sites (mid-test SUT/fault-injection removals,
  error-swallowing or name-colliding local teardown helpers) keep the raw call
  with an inline eslint-disable + reason.

Shared anti-ratchet primitive
- scripts/lib/allowlist-ratchet.cjs:
  - assertWithinAllowlist: fails on NOVEL ids (new offender introduced) AND on
    STALE ids (a known offender was fixed but not pruned) — identity, not count,
    and a ratchet DOWN toward zero.
  - assertTightCeiling: a size/length budget whose ceiling must stay within a
    grace band of the high-water mark, so budgets may only tighten, never creep.

Ratchets converted onto the primitive
- windows-test-parity-guard.test.cjs: rmSync rule deleted (now ESLint-enforced);
  the remaining six patterns moved from integer baselines to named-set
  allowlists with ratchet-down.
- scripts/lint-test-file-count.{cjs,allowlist.json}: per-module integer counts →
  named filename sets (closes the swap-a-file-keep-the-count blind spot); a
  module dropping under cap now FAILS to force pruning its allowlist entry.
- enh-2790 skill-count `<= 63` → named skill allowlist (ratchets toward ~58).

Size budgets hardened (tighten-only)
- agent-size / workflow-size / feat-3039 help-tiered: ceilings lowered to the
  current high-water mark and an assertTightCeiling anti-creep check added per
  tier. Fixed external-contract limits (description ≤100 chars, agent ≤100 KB)
  are intentionally left as-is — they are not grandfathered creeping budgets.

No user-facing behavior change (tests + tooling only); no USER_FACING_PREFIXES
touched, so no changeset fragment is required.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 22:43:49 -04:00
Tom Boucher
091b7e2c4b perf(#305): single-pass max-by-mtime for statusline todo lookup (#385)
Replace the per-render readdirSync().filter().map(statSync).sort() chain
with a single-pass max-by-mtime loop. Drops the O(n log n) sort and the
throwaway intermediate array; I/O and resolved-file behavior are identical.
Adds the first behavior-lock test for the todo-resolution path.

The larger disk-backed cache win from the issue is deferred: statusline is
a fresh child process per render, so any cache must be disk-backed with
invalidation/atomic-write design that needs maintainer input.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-27 20:21:15 -04:00
Tom Boucher
6913dbcdb1 fix(10): centralize semver comparison policy across hooks and changeset 2026-05-24 18:14:56 -04:00
Tom Boucher
dbc09d21a6 fix(3153): handle numeric 100 percent and block-list next_phases 2026-05-05 18:34:42 -04:00
Tom Boucher
0ea443cbcf fix(install): chmod sdk dist/cli.js executable; fix context monitor over-reporting (#2460)
Bug #2453: After tsc builds sdk/dist/cli.js, npm install -g from a local
directory does not chmod the bin-script target (unlike tarball extraction).
The file lands at mode 644, the gsd-sdk symlink points at a non-executable
file, and command -v gsd-sdk fails on every first install. Fix: explicitly
chmodSync(cliPath, 0o755) immediately after npm install -g completes,
mirroring the pattern used for hook files throughout the installer.

Bug #2451: gsd-context-monitor warning messages over-reported usage by ~13
percentage points vs CC native /context. Root cause: gsd-statusline.js
wrote a buffer-normalized used_pct (accounting for the 16.5% autocompact
reserve) to the bridge file, inflating values. The bridge used_pct is now
raw (Math.round(100 - remaining_percentage)), consistent with what CC's
native /context command reports. The statusline progress bar continues to
display the normalized value; only the bridge value changes. Updated the
existing #2219 tests to check the normalized display via hook stdout rather
than bridge.used_pct, and added a new assertion that bridge.used_pct is raw.

Closes #2453
Closes #2451

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-20 10:08:46 -04:00
Tom Boucher
53078d3f85 fix: scale context meter to usable window respecting CLAUDE_CODE_AUTO_COMPACT_WINDOW (#2219)
The autocompact buffer percentage was hardcoded to 16.5%. Users who set
CLAUDE_CODE_AUTO_COMPACT_WINDOW to a custom token count (e.g. 400000 on
a 1M-context model) saw a miscalibrated context meter and incorrect
warning thresholds in the context-monitor hook (which reads used_pct from
the bridge file the statusline writes).

Now reads CLAUDE_CODE_AUTO_COMPACT_WINDOW from the hook env and computes:
  buffer_pct = acw_tokens / total_tokens * 100
Defaults to 16.5% when the var is absent or zero, preserving existing
behavior.

Also applies the renameDecimalPhases zero-padding fix for clean CI.

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-16 15:40:15 -04:00
Maxim Brashenko
cc04baa524 feat(statusline): surface GSD milestone/phase/status when no active todo (#1990)
When no in_progress todo is active, fill the middle slot of
gsd-statusline.js with GSD state read from .planning/STATE.md.

Format: <milestone> · <status> · <phase name> (N/total)

- Add readGsdState() — walks up from workspace dir looking for
  .planning/STATE.md (bounded at 10 levels / home dir)
- Add parseStateMd() — reads YAML frontmatter (status, milestone,
  milestone_name) and Phase line from body; falls back to body Status:
  parsing for older STATE.md files without frontmatter
- Add formatGsdState() — joins available parts with ' · ', degrades
  gracefully when fields are missing
- Wrap stdin handler in runStatusline() and export helpers so unit
  tests can require the file without triggering the script behavior

Strictly additive: active todo wins the slot (unchanged); missing
STATE.md leaves the slot empty (unchanged). Only the "no active todo
AND STATE.md present" path is new.

Uses the YAML frontmatter added for #628, completing the statusline
display that issue originally proposed.

Closes #1989
2026-04-10 15:56:19 -04:00