f7df920681f233ae0fe064ee659550bdf41ff708
22 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
adb46cdd85 |
feat(#2734): surface STATE.md commit-age on the statusline (#3700)
* test(#2734): failing-first suite for the statusline STATE.md freshness marker Binds the contract before any hook change exists: a `state ~N commits back` segment gated on the state_head stamp landed by #2622, firing at the same advisory threshold /gsd-health's W024 uses rather than at > 0. Covers all five acceptance criteria — threshold parity (19/20/21 boundaries), both renderers including formatGsdStateCompact, an exact spawn-count assertion, repo-pinning and sub_repos degradation, and behavioral parity against readStateHeadFreshness rather than a source-grep of the two fence copies. 52 example-based tests plus 5 seeded fast-check properties. Red now by design. * feat(#2734): surface STATE.md commit-age on the statusline Adds an opt-in `state ~N commits back` marker to the GSD-state segment, consuming the `state_head` stamp and freshness contract landed by #2622. A solo developer returning to a project reads "Phase 4, executing" in STATE.md and acts on it, without noticing the codebase moved 40 commits since that line was written. /gsd-health reports it as W024, but only if you think to run it; the statusline is the surface you see without asking. Fires at STATE_HEAD_ADVISORY_COMMITS (20), the same threshold W024 uses, not at > 0: with commit_docs:true the commit carrying a STATE.md sync advances HEAD by one, so > 0 would alarm permanently on a fresh project. Costs exactly one bounded git subprocess per render and none when disabled. `rev-list --left-right --count` answers ancestry and distance together, and repo pinning is a filesystem check mirroring projectOwnsItsRepo rather than a --show-toplevel compare, which is unreliable on macOS /private/var and Windows 8.3 paths. Every unresolvable input degrades to the tri-state unknown -- the marker is absent, never a "fresh" claim the project cannot substantiate: a malformed stamp, a root that does not own its .git, a sub_repos workspace, history rewound past the stamp, or git being unavailable. Also collapses statusline config resolution onto one resolveStatuslineOptions() seam. runStatusline() and renderStatusline() duplicated it byte-for-byte; one copy is what keeps a newly-added key from reaching only one of them. * test(#2734): route the e2e spawn through the process seam and fix fixture leaks Review findings from the two orthogonal passes: - `bothEntryPointsResolveOptionsIdentically` spawned a child and substring-matched its stdout to test a pure function. It now calls resolveStatuslineOptions() directly — no subprocess, no text matching. - `skipsFreshnessWorkWhenTodoTaskActive` genuinely needs a child (the !task gate lives in runStatusline, which reads stdin), so it now spawns through tests/helpers/process-seam.cjs and proves the negative with a filesystem fact: the git shim appends to a marker file on every invocation, and the assertion is that the marker never appears. Stronger than asserting text is missing, and it drops the last stdout substring match in the block. - Every fixture-creating test now registers `t.after(() => cleanup(dir))` instead of a trailing cleanup(dir), which leaked the temp repo on assertion failure. derivationAgreesWithStateModule reassigns `dir` across five fixtures, so it binds each directory at scheduling time rather than cleaning only the last. Also corrects markerCoexistsWithMilestoneComplete, which asserted the wrong expectation rather than finding a code defect: `percent` drives the progress bar too, so the milestone segment reads "v1.9 [##########] 100%". The marker appends after it, which is what the test exists to prove. CONTEXT.md's opt-in statusline key list was missing statusline.show_git as well as the new key; both are now enumerated. * docs(#2734): backfill changeset PR number (#3700) --------- Co-authored-by: sim <sim@local> |
||
|
|
bf2332e67c |
fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in its fail-closed catch and is misreported as an unreadable dispatch-isolation configuration, blocking every executor dispatch. These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor guard and update worker, plus the seam's actionable build error surfacing instead of the generic misreport. Also adds the drift lint the acceptance criteria require, with a fixture proving it CAN fail — a guard never shown to fail is worthless. It is red here by design: it flags today's unfixed hooks, which is exactly the defect. * fix(3582): route every hook's compiled-module require through the self-heal seam RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard, cursor guard and update worker, the fail-closed-with-actionable-message assertion, and the lint's own real-tree check. The compiled runtime library is produced by build:lib and gitignored (ADR-457, build-at-publish). The npm package builds before publishing; a plugin-marketplace or git-clone install materializes the raw tree and never does, so on that channel those modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly this, and the CLI entrypoint already calls it — no hook did. The isolation guard's Cannot-find-module therefore landed in its fail-closed catch and was reported as 'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was misdiagnosed as an unreadable project config and every executor dispatch was blocked. All SEVEN affected files now call the seam before their first compiled require. The issue named four; a scan found six; implementing it surfaced a seventh — the shared isolation sentinel helper, used by BOTH guards, which requires two compiled modules itself and would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so fixed here rather than left as a known-broken remainder. Failure posture is deliberately split by hook kind: - Gates (agent isolation guard, cursor subagent start) surface the seam's actionable build error distinctly instead of swallowing it into the generic text, and stay fail-closed — a genuinely unreadable project config still DENIES exactly as before. - Cosmetic and detached hooks (statusline, update worker, update check, update banner) DEGRADE rather than crash: the statusline draws on every render and the worker is a detached process, so a build failure there must not take down the prompt. The npm path is untouched: the seam's already-built fast path returns immediately, so prebuilt installs pay nothing and behave bit-for-bit as before. Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than remembered — without it the next hook to add a compiled require reintroduces the class silently. It is proven able to fail: a fixture hook requiring a compiled module without the seam is flagged, and one that uses the seam is not. Verified directly — on the unfixed tree it named all seven offenders; with the fix it passes. While writing the lint's comment stripper, a naive whole-text block-comment regex ate its own fixture, because this repo's comments legitimately spell the compiled-lib glob whose star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression test pinning that case. * fix(3582): test the three untested seam call sites and assert typed reason codes Two independent reviews converged on the same major gap: the fix wired the seam into seven files but only four had cold-tree tests. The adversarial pass put it plainly — deleting the shared isolation-sentinel helper's seam call would not have failed any test in the diff. That file was my own addition beyond the issue's four, so it shipped untested; that is now closed. - Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT directly under cwd, and every existing cold-tree fixture puts it there, so the early return always fired first. Now covered, and proven load-bearing by mutation: with the call removed the spy records zero seam invocations and the test fails. - update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED VERDICT — the fallback cache filename, and silent suppression when the package name degrades to null — rather than merely 'did not throw'. The banner hook previously had no test file at all. Standards violation fixed: two tests asserted on free-form prose via assert.match against a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch it. Both isolation guards now emit a machine-readable reason_code from a frozen enum, following the repo's existing REASON convention, and the tests assert that instead. The human-readable message is unchanged for operators; only the assertion target moved. The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT extracted, and the reason is recorded at each site: both viable shapes — a path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file literal co-occurrence check, so extracting would require the lint to special-case its own helper. Triplication is the lesser evil while the lint stays a co-occurrence scan. The lint's header now states what it does and does not catch (literal quoted requires only; hooks/ scan root), so a future reader does not over-trust a guard that a concatenated path or a require inside a non-hooks helper would evade. * chore(3582): regenerate the committed install-tree fixtures Adding a new shipped hook helper changed the install tree, and those fixtures are committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>' tests failed on 541a1913. Regenerated rather than hand-edited. The delta across all 15 runtime fixtures is exactly two lines — the new helper under both its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in no unrelated drift. This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from lint:ci, which passed both before and after. * chore(3582): backfill changeset PR number (#3629) --------- Co-authored-by: sim <sim@local> |
||
|
|
aee83c8e9e |
test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5 (#3378)
* test(#3337): fold the manifest & package-identity issue-* cluster — Wave 5 Folds 6 legacy issue-*.test.cjs regression files (120 test() blocks) into their module's main suite, per H3 (#3315) of the test-hygiene epic (#3053). Second of 4 issue-* waves. - issue-766-plugin-manifest.test.cjs (50 tests): pure rename (git mv) into plugin-manifest.test.cjs, sole comprehensive suite for its module. - issue-844-manifest-version-sync.test.cjs (14 tests, rename basis) + issue-1855-marketplace-manifest.test.cjs (17 tests, merged in): both target scripts/sync-manifest-versions.cjs via non-overlapping describe blocks (generic engine vs. marketplace.json-specific), now manifest-version-sync.test.cjs, 31 tests, 0 dropped. - issue-498-identity-drift-lint.test.cjs (8 tests, rename basis) + issue-498-package-identity.test.cjs (17 tests, merged in): both target package-identity.cjs's surface from different angles (lint-side drift detection vs. derive/slugify), now package-identity.test.cjs, 25 tests, 0 dropped. - issue-607-cache-lineage.test.cjs (14 tests) merged into the existing gsd-statusline.test.cjs, matching its already-established fold-wrapper convention from prior waves. Learned from Wave 4: fold agents checked src/** (not just docs/ and gsd-core/references/) for stale filename references — found and fixed 5 across 3 ADR docs (766, 2121, 457), zero in src/ this time. Zero net test-coverage loss. No production code changed. * test(#3337): fix orthogonal-review findings — Wave 5 fold Standards-axis review found real issues in the just-folded manifest-version-sync.test.cjs, both fixed here: - The folded:issue-1855-marketplace-manifest block omitted its own test/describe/assert/fs/path/os requires, silently closing over the #844 basis section's outer-scope bindings instead of declaring its own dependencies — inconsistent with every other fold block in this wave. Restored the local requires the original issue-1855 file declared. - Merging two independently-lettered legacy files (each A-F) left duplicate top-level describe() labels (two "A:", two "B:", two "C:"). Renamed the folded-in #1855 labels (A2/B4/C2) to make every top-level label in the file unique. No test() count changed (31). No production code touched. --------- Co-authored-by: sim <sim@local> |
||
|
|
9faacc0c15 |
test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist Migrates the final 170 unbounded sync spawn sites across 49 files, then removes the allowlist entirely. local/no-unbounded-spawn now runs with no exemption surface across tests/**: there is no file to add a name to. drift-detection's throw-native git() helper routes to gitOrThrow -- bare runGit would have taken 16 call sites quiet on failure. commands.test.cjs has two independently-scoped runGsdTools/runCli helpers, one already bounded and one not; they are kept distinct rather than unified, the same trap as the two same-named git() helpers in Wave 1. runNpm's bound was erasable. Its options spread callerOptions after the defaults, so an explicit timeout:undefined silently dropped the 180000ms bound -- the rule flagged it and was right; it was not a false positive. Fixed by destructuring with a default, with a test that fails when the default is removed. Two sites stay on a raw spawn with an explicit timeout because the seam cannot express them: one needs shell:true for npm.cmd on Windows, one redirects stdout to a real fd. Both are the rule's own documented second option, not an escape from it. Closure verified rather than asserted: the derivation scan reports 0 unbounded spawn helpers and 0 unbounded direct git call sites, and a temporary file carrying an unbounded spawn still errors with the allowlist gone. Closes #3064. * test(#3148): close a hole in the guard's own eslint-disable ban The ban listed only the top level of tests/, so it was blind to 37 .cjs files under tests/helpers, qa, observability, fixtures and dispatch. With the allowlist deleted this test is the sole remaining way to detect someone silencing the rule inline, so the gap was load-bearing: a nested file could carry an unbounded spawn plus an eslint-disable and pass everything. Proven before and after. A probe planted under tests/helpers with both was invisible to the guard and clean under eslint; after making the listing recursive the guard fails on it. The scanned set goes from 771 files to 808. Pre-existing since the guard shipped, but this wave is what promoted it to sole defense, so it is fixed here rather than filed. Also converts the last hand-rolled throw check to throwIfFailed and the last re-derived legacy shape to compose toLegacyResult, which makes the epic's none-remain claim true rather than nearly true. toLegacyResult itself is not widened -- eight callers depend on its shape and one consumer does not justify changing a shared contract. * fix(#3148): correct seam incoherence at the bound and a slow review-lane error path Two real failures from the remote runner, both fixed at the cause. The seam could return outcome TIMED_OUT together with exitCode 0. At the exact bound spawnSync reports ETIMEDOUT while the child has already exited with a real status, and toSeamResult classified on the error code while passing status straight through -- an incoherent pair its own boundary test was written to catch, and did. A status that is not null is direct evidence the child exited on its own, so it now decides the outcome before the error-code branches run. process-seam.cjs was deliberately untouched by every earlier wave; this is a defect in the module itself, kept surgical, with a unit test that fails against the old logic. review-lane with an unknown subcommand fell through to its usage error only after loading the capability registry and building a per-lane plan, which spawns one child process per lane -- up to twelve. The error path took ~1288ms instead of ~119ms, and under bench load it outran a caller's spawn timeout and was killed before writing anything, which is the empty stdout and stderr CI saw. It now fails fast before any of that work begins. This is the epic's first production change. It is user-facing, so it carries a changeset rather than a no-changelog label. * test(#3148): replace a real-race timeout test with a deterministic one E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a warm container git finishes first, spawnSync returns status 0 with no error at all, the seam correctly classifies EXITED, and gitOrThrow correctly does not throw -- so the test failed on both lanes. A probe confirms a genuine timeout always carries status null, so this was never the seam misbehaving. Raising the bound would only lengthen the odds, which is the same defect with better luck. The test now drives gitOrThrow against a stubbed runGit that returns a synthetic TIMED_OUT result, so it asserts exactly what it always meant to -- that a timeout propagates as a throw -- with no timing dependence. Five consecutive runs are identical where the old one varied. I wrote this test in Wave 0; it is a real-race test by construction and CLAUDE.md says to replace those rather than re-run them. * chore(#3148): backfill changeset PR number 3192 --------- Co-authored-by: sim <sim@local> |
||
|
|
5fd5c81042 |
test(#3055): add the process seam so a subprocess timeout is expressible as data (#3066)
* test(#3055): add the process seam and route runGsdTools through it Adds tests/helpers/process-seam.cjs — runNode/runGit/runHook over spawnSync, each returning a typed discriminated union { outcome, exitCode, stdout, stderr, timedOut, signal, killed, code }. Every call is timeout-bounded; there is no unbounded path. runGsdTools becomes an adapter over the seam. Its legacy { success, output, error, exitCode } shape and retry-once-on-kill behaviour are preserved byte-identically, so none of its 136 caller files change. Outcome discrimination was corrected against probed runtime behaviour rather than assumption: a timeout and a maxBuffer overflow are identical on both status (null) and signal (SIGTERM), and differ only by code (ETIMEDOUT vs ENOBUFS). Overflow is therefore classified before timeout. This fixes a live defect — the previous isKilled() treated an overflow as a kill, retried it for a second full 60s run, and then reported "host OOM or scheduler contention" for a child that had merely printed too much. Also widens the ESLint tests glob from tests/**/*.test.cjs to tests/**/*.cjs, which brought 31 previously unlinted shared helpers under the same rules their sibling test files already obey, and fixes the 5 violations that surfaced — including a bare npm invocation without shell:true in tests/helpers/emitted-runtime.cjs (DEFECT.WINDOWS-TEST-PORTABILITY), now routed through the existing portable runNpm helper. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3055): migrate every local spawn wrapper onto the process seam Replaces the spawn body of all 25 local runHook/runGuard/runGate definitions with a call to tests/helpers/process-seam.cjs. Each wrapper keeps its name, parameter list, return shape and post-processing (JSON parse, ANSI strip, env sanitising, field extraction) — only the spawn mechanism changes, so no test assertion moves. The 4 bash-driven wrappers use the seam's explicit `interpreter` option rather than a fourth primitive; it is explicit rather than inferred from the file extension, because guessing an interpreter from a path fails silently when a script's name does not match its shebang. Seven wrappers were previously unbounded and now carry an explicit timeout sized to what each actually runs, not the seam default. Two of those seven (gsd-write-guard, lint-docs-command-form) were absent from the issue's inventory entirely and were found by scanning after the migration. Adds the CONTEXT.md `### Process seam` glossary entry and a CONTRIBUTING.md reference section covering the three primitives, the discriminated union, and the two rules the seam enforces. Scope disclosure recorded in the phase design notes: the issue scoped three identifier names. A scan for local helpers that spawn AND return the spawn result finds 113 across 82 names, 71 of them unbounded, plus 122 unbounded direct git call sites. This change bounds 25 of those. The remaining surface is the same defect class and is NOT closed by this PR. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): classify an externally-killed child as KILLED, not EXITED Blocker found in this branch's own diff, independently confirmed by an isolated reviewer. A child killed by an external signal — a genuine bench OOM kill — makes spawnSync return { status: null, signal: 'SIGKILL' } with NO .error field. The seam's "no error implies EXITED" rule therefore classified it as a clean exit, and runGsdTools returned { success: false, exitCode: 1 } without retrying. That silently defeated the #969 kill-discrimination for precisely the case it was built for: the old isKilled() fired on `signal != null`, retried once, then threw a labelled resource-starvation error. A real OOM would have been reported as an ordinary assertion failure. Adds a fifth outcome, KILLED, for "no error but a signal is set", and makes the adapter retry on TIMED_OUT or KILLED — reproducing the old `killed || signal != null || code === 'ETIMEDOUT'` condition exactly. SPAWN_FAILED still does not retry (matching the old behaviour, where signal was null). BUFFER_OVERFLOW still does not retry, which remains a deliberate divergence: the old code retried it because signal was SIGTERM, burning a second 60s run on a child that had merely printed too much. All five outcomes verified against the live runtime rather than assumed: SIGKILL -> killed, exit 0/7 -> exited, timeout -> timed_out (ETIMEDOUT), >1MB stdout -> buffer_overflow (ENOBUFS). Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): address standards-review findings on this branch Three findings from the standards axis of the review, all in this branch's own diff. The CONTEXT.md glossary entry this branch introduced was already stale on the branch's own last commit: it enumerated a 4-member OUTCOME while the code had 5, because the KILLED fix did not update it. That is precisely the drift the "module changes update Domain-terms" gate exists to catch, so the entry now lists all five and explains KILLED. api-coverage-gate-e2e compared an outcome against the raw string 'exited' rather than OUTCOME.EXITED, the only such outlier; the enum is now imported and used. A sweep for the other four outcome literals found no further comparison sites. Three call sites hand the literal bash flag '-c' to the seam's first parameter, which the JSDoc described as an absolute script path. Rather than add a fourth primitive, the contract is corrected to match reality: the parameter is renamed `target` and documented as the first argv element handed to the interpreter — normally a script path, but for an interpreter invoked with an inline program it may be that interpreter's own flag. No behaviour change. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3055): assert the cross-platform timeout contract, not the macOS one The remote runner failed on both Linux lanes (node 22 and node 24, identical) while the same tests passed locally on macOS. Two assertions encoded a platform-specific behaviour as a cross-platform guarantee. When spawnSync times out, macOS preserves the child's partial stdout/stderr; Linux discards it and returns empty strings. Verified on node v26.5.1 both ways. The seam passes through whatever spawnSync hands it and cannot manufacture output that was discarded, so the production code was correct — the tests were wrong. Both tests now assert the guarantee the seam actually makes on every platform: outcome TIMED_OUT, timedOut true, and stdout/stderr always being strings rather than undefined or a Buffer. The partial-content assertions are retained behind an explicit process.platform === 'darwin' guard so the macOS coverage is not lost, and the first test is renamed to say what it now guarantees. This is the failure mode the remote matrix exists to catch: local macOS verification would have shipped it. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): classify a failed spawn as SPAWN_FAILED, not a timeout Windows CI caught two defects the Linux matrix could not. tests/context-predicates-query.test.cjs passes a 32K-char argv value. On Windows that exceeds the argv limit and spawnSync fails with code ENAMETOOLONG, signal null, status null. The seam's fallback rule — "otherwise, status === null implies TIMED_OUT" — swallowed it, so the adapter retried a spawn that can never succeed and then threw the resource-starvation error. The old isKilled() returned false for that shape and returned an ordinary failure result. TIMED_OUT is now identified positively: code === 'ETIMEDOUT' OR signal is set. Anything else carrying an error is SPAWN_FAILED, which covers ENAMETOOLONG, E2BIG, EACCES and ENOENT alike. The signal clause is what keeps a platform whose timeout errno differs classified correctly, so the greedy catch-all is no longer needed. The second defect is a contract regression I introduced and had claimed otherwise. That same test asserts `typeof r.exitCode === 'number'`, and toLegacyShape was returning null for BUFFER_OVERFLOW and SPAWN_FAILED, so the assertion failed on type. The old code returned `err.status ?? 1` on every non-retried failure path. The adapter now returns 1 again for both, and the comment claiming "never coerced to exitCode:1, unlike the pre-seam helper" is retracted: the seam keeps the richer truth (exitCode null plus a distinct outcome), the legacy adapter keeps the old numeric contract its callers actually depend on. Verified on this host: a 4MB argv yields E2BIG -> SPAWN_FAILED; ENOENT -> SPAWN_FAILED; timeout -> TIMED_OUT; >1MB stdout -> BUFFER_OVERFLOW; SIGKILL -> KILLED; clean exit -> EXITED. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fd07e1a357 |
fix(#2850): resolve the active workstream in the statusline GSD-state segment (#3012)
* test(#2850): add failing-first tests for workstream statusline state readGsdState only ever reads the flat .planning/STATE.md via a directory walk-up; it has no path for .planning/workstreams/<ws>/STATE.md and never consults GSD_WORKSTREAM or the stored active-workstream pointer, so the GSD-state segment silently disappears in workstream mode. These tests prove the RED before the fix lands. Uses shared saveSessionEnv/restoreSessionEnv/clearSessionEnv helpers now added to tests/helpers.cjs (single source of truth for the session-env-var save/clear/restore pattern also used by tests/active-workstream-store.unit.test.cjs, which is updated here to consume the same shared helpers instead of its own local, already-diverged copy). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2850): resolve active workstream in the statusline readGsdState only ever walked up looking for a flat .planning/STATE.md; it had no branch for .planning/workstreams/<ws>/STATE.md and never consulted GSD_WORKSTREAM or the stored active-workstream pointer, so the GSD-state segment silently vanished in workstream-mode projects with no root STATE.md (exit 0, no diagnostic). Reuses the existing CLI>env>store resolution seam (resolveActiveWorkstream, active-workstream-store.cts) and the existing mode-detection/path-building seam (listAvailableWorkstreams/planningPaths, planning-workspace.cts) rather than re-implementing either inline. When workstream mode is detected but nothing resolves, readGsdState now returns a {noActiveWorkstream:true} sentinel that formatGsdState/formatGsdStateCompact render as "no active workstream" -- observable, never silent emptiness. Flat-mode behavior and the case where a resolved workstream has no STATE.md yet are both unchanged. Adds active-workstream-store.cts's peekActiveWorkstream: a read-only sibling of getActiveWorkstream. resolveActiveWorkstream's default store lookup self-heals a stale/invalid pointer by deleting it (adapter.clear()) -- correct for a command, but not for a renderer invoked once per prompt, which must never mutate persistent, possibly cross-session state as a side effect of drawing a screen. The statusline now injects peekActiveWorkstream via resolveActiveWorkstream's own getStored override, keeping the env>store precedence itself fully reused while removing only the store tier's write side effect. This satisfies the issue's AC4 ("the fix is purely additive to what's displayed"). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2850): backfill changeset PR number to 3012 --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ef7a71fc09 |
fix(#2754): make parseStateMd CRLF-safe — frontmatter no longer drops on Windows-authored STATE.md (#2865)
* test(#2754): parseStateMd must parse CRLF STATE.md identically to LF The frontmatter fence regex and downstream splits used literal \n, so a CRLF STATE.md dropped the ENTIRE frontmatter block. Assert CRLF/LF parity (the production contract) across full frontmatter, next_phases flow + block forms, the progress nested block, and null handling. * fix(#2754): make parseStateMd CRLF-safe — frontmatter fence + splits use \r?\n The fence regex, scalar-line split, next_phases block-list regex, and progress block regex all used literal \n, so a CRLF (Windows-authored) STATE.md dropped the ENTIRE frontmatter block — every statusline field was silently absent. Use \r?\n throughout, mirroring the CRLF-safe extractFrontmatter in src/frontmatter.cts. LF behavior unchanged. * chore(#2754): changeset fragment * test(#2754): pin parseStateMd↔extractFrontmatter parity (Generative-Fix Divergence guard) The statusline parseStateMd and the canonical extractFrontmatter both derive GSD state from STATE.md frontmatter and diverged once already (the CRLF bug this PR fixes). Add a cross-parser parity assertion (CLAUDE.md parallel-surfaces rule) over the overlapping fields, under both LF and CRLF, so a future divergence on a scalar shape or line ending is caught here. (isolated-review minor finding) * chore(#2754): backfill changeset PR number (2865) --------- Co-authored-by: Test <test@example.com> |
||
|
|
20ff405cb3 |
feat(#2162): opt-in compact GSD-state format for the statusline (#2175)
* feat(#2162): opt-in compact GSD-state format for the statusline New statusline.state_format config, enum full|compact (default full — existing rendering untouched). "compact" renders the state segment as "<version> · P<phase>/<total> · <status>", e.g. "v1.12 · P7/12 · executing" — dropping the milestone name and progress bar (the two biggest width costs) and collapsing narrative statuses to a single keyword. Per the #2162 approval conditions, the keyword set is the canonical vocabulary from normalizeStateStatus() in state-document.cjs (discussing/planning/executing/verifying/completed/paused) — no parallel hand-rolled list, so the vocabularies can't drift — and the canonical stuck state "paused" renders uppercase as PAUSED (no new "blocked" lifecycle state). Statuses the normalizer passes through unrecognized fall back to their first word capped at 16 chars. Lifecycle scenes preserved: active_phase wins over the body phase number, milestone completion renders "complete", idle-with-next-action renders "next <action> <phases>". Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2162): changeset fragment for PR #2175 * fix(#2162): review fixes — ENUM_KEYS coverage, cap boundary tests, changeset format - register statusline.state_format in the fix-1628 coercion-bypass matrix - 15/16/17-char boundary tests for the shortGsdStatus fallback cap - changeset body ends with the (#2162) citation per house convention Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): round-2 review fixes — scene exclusivity, direct config-set coverage - compact renderer gates the milestone-complete scene behind the absence of an in-flight phase id, mirroring formatGsdState's if/else precedence (Scene 1 beats Scene 3); regression test covers the non-atomic active_phase + percent=100 STATE.md shape - direct config-set accept/reject test for statusline.state_format plain strings (ENUM_KEYS matrix covers only the JSON coercion shapes) Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the statusline hook change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2162): complete-scene gate matches formatGsdState exactly (+property tests) Re-review Major: gating done on !phaseId held completion back for the legacy phaseNum shape — formatGsdState reaches Scene 3 on percent=100 regardless of phaseNum, so compact must too. Gate is now !s.activePhase. The phaseNum-only test now expects 'complete' and cross-checks the full renderer; a parity test feeds identical inputs to both renderers. Re-review Minor: shortGsdStatus gets fast-check property coverage (totality, canonical fixed points, separator safety, fallback shape). Golden fixtures regenerated for the hook byte change. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg |
||
|
|
eec9efc351 |
feat(#2160): collapse verbose '(1M context)' model suffix to compact (1M) badge (#2173)
* feat(#2160): collapse verbose '(1M context)' model suffix to compact (1M) badge Claude Code appends " (1M context)" to the model display name in long-context sessions, eating 12 characters of statusline width. Collapse it to " (1M)" — the signal stays, the width doesn't. Any other display name passes through unchanged. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2160): changeset fragment for PR #2173 * fix(#2160): review fixes — ctx variant, boundary test, changeset format - broaden the suffix match with a context|ctx alternation (approval-condition variant the regex missed) - pin non-context parentheticals ((beta), (deprecated)) as untouched - changeset body ends with the (#2160) citation per house convention Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the statusline hook change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
73039d3717 |
fix(#2163): round-2 review fixes — deterministic fail-soft injection tests
- readGitStatus calls execFileSync via the child_process namespace so tests can inject spawn failures through the shared module object - two deterministic tests: ERR_CHILD_PROCESS_STDOUT_MAXBUFFER-shaped and ETIMEDOUT-shaped throws both degrade to null (segment absent), proving the fail-soft paths the PR previously only asserted Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg |
||
|
|
d6672ff926 |
feat(#2163): opt-in git branch/status segment in the statusline
New statusline.show_git config (default false). When enabled, a git segment renders after the directory: current branch plus compact work-state markers (+staged ~unstaged ?untracked ↑ahead ↓behind, or ✓ when clean and in sync), e.g. " │ main+2~1?3". One git status --porcelain=v2 --branch spawn per render via execFileSync with a fixed argument array (no shell), a 1.5s timeout, and the workspace dir passed with -C. Fails silently — segment absent outside a repo, without git, or on timeout. Default output is unchanged when the flag is absent. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg |
||
|
|
ad7111e50b |
feat(#2161): opt-in absolute token count on the statusline context meter (#2174)
* feat(#2161): opt-in absolute token count on the statusline context meter New statusline.show_context_tokens config (default false). When enabled, the context meter shows the absolute token total after the percentage, e.g. "████░░░░░░ 46% (156k)" — summing input, cache-creation, cache-read, and output tokens from context_window.current_usage (matching /context). Default output is byte-for-byte unchanged when the flag is absent or false. The .planning config is now read once per render and shared with the last-command/position block instead of being re-read. Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * docs(#2161): changeset fragment for PR #2174 * fix(#2161): review fixes — k-to-M threshold, boundary tests, changeset format - formatTokens promotes to the M branch when k-rounding reaches 1000 (999,500-999,999 rendered "1000k" instead of "1.0M") - boundary tests at 999499/999500/999999/1000000/1000001 - Number() guards on the four usage fields (silent string-concat gap) - changeset body ends with the (#2161) citation per house convention Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * fix(#2161): round-2 review fixes — config-set coverage, precision claim, exports style - config-set accept/reject tests for statusline.show_context_tokens (mirrors the post-planning-gaps precedent the issue scope names) - changeset + docs no longer claim parity with /context: the suffix sums four fields while the meter %% derives from used_percentage (three), so the figures can diverge slightly - module.exports one entry per line Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg * test: regenerate golden-install-parity fixtures for the statusline hook change Claude-Session: https://claude.ai/code/session_01Hme55Pvq6BhpgwBcyC5HAg --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
0cc7a1a426 |
test(#1974): consolidate 27 installer/hooks remainder tests into module suites
Fold 27 issue-named installer/hooks/statusline/migration/reapply regression files into their canonical module suites (installer-migrations, installer-migration-report, gsd-statusline, reapply-verify-hunks, install-*, gsd-check-update-worker-platform-gate, etc.). Verbatim block-scoped describe wrappers; 276 subtests conserved 1:1. No new files. The one subdir origin (tests/installer-migrations/001-legacy-orphan-files) moved up one level into installer-migrations.test.cjs; its single ../../ module require corrected to ../ so it resolves from tests/ root (verified). Host-env pre-check: no CLI-receiving host sets a redirecting GSD_WORKSTREAM/GSD_PROJECT value. Regenerates regression-name allowlist (222->205), ratchets file-count allowlist (verify 11->8, validate entry removed), makes 16 relocated allow-test-rule exemptions issue-ref- compliant (ADR-456; prunes stale ids). Repoints 15 tests/ references across state-md.md (EN + ja/ko/pt/zh). lint:ci green. Part of epic #1969. Closes #1974. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d444864bf8 |
fix(#1194): correct inverted statusline auto-compact buffer math (#1211)
* fix(#1194): correct inverted statusline auto-compact buffer math The reserved-buffer percentage was computed as (acw/totalCtx)*100 — the usable fraction — instead of (1 - acw/totalCtx)*100 — the reserved fraction. When acw == totalCtx this produced buffer=100%, making the usable-range denominator zero and pinning `used` at a constant 100% regardless of real remaining context. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: backfill changeset PR number (#1211) --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
bf48127c13 |
chore(lint): justify intentional no-control-regex (ANSI strip) + ratchet to error (#737)
The 5 no-control-regex warnings are all the same intentional ANSI-color-strip pattern /\x1b\[[0-9;]*m/g across 5 test files. The \x1b (ESC) control char is the required leading byte of an ANSI SGR sequence, so matching it is the whole point of stripping color codes from captured CLI/console output. Add an inline eslint-disable-next-line with justification at each site (not a refactor — the control char is essential, not accidental), then flip no-control-regex from warn to error so the debt can't regrow. Refs #736 Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a28dcec981 |
chore(#597): replace count-based ratchet guards with AST lint + named-set allowlists (#603)
The windows-test-parity ratchet greps test source for fs.rmSync-without-
maxRetries (and six other Windows-portability anti-patterns), failing when an
integer offender COUNT exceeds a frozen baseline (rmSync: 95). A count ratchet
is a Goodhart metric: fixing one offender and adding another keeps the count
constant, so a new defect slips through green. Replace it — and every other
count ratchet in the repo — with a layered, masking-proof design.
Behavioral seam test
- tests/helpers-cleanup.test.cjs proves helpers.cleanup() carries the Windows
EBUSY retry budget. cleanup() delegates retries to Node's fs.rmSync via
maxRetries (it owns no loop), so the test asserts the option contract
(recursive/force/maxRetries>0/retryDelay>0) + real-FS removal + the cwd-guard,
rather than a loop that does not exist. The EBUSY risk is now tested ONCE at
the helper, not approximated textually at every call site.
Write-time ESLint rule (AST-accurate, replaces the grep)
- eslint-rules/no-raw-rmsync-in-tests.cjs (error in tests/**/*.test.cjs) bans
raw fs.rmSync, steering to cleanup(). Catches member, computed (fs['rmSync']),
destructured and aliased forms; escape hatch is inline
`// eslint-disable-next-line local/no-raw-rmsync-in-tests -- <reason>` only.
- Migrated 336 raw fs.rmSync teardown calls across ~116 test files to cleanup().
~18 genuinely load-bearing sites (mid-test SUT/fault-injection removals,
error-swallowing or name-colliding local teardown helpers) keep the raw call
with an inline eslint-disable + reason.
Shared anti-ratchet primitive
- scripts/lib/allowlist-ratchet.cjs:
- assertWithinAllowlist: fails on NOVEL ids (new offender introduced) AND on
STALE ids (a known offender was fixed but not pruned) — identity, not count,
and a ratchet DOWN toward zero.
- assertTightCeiling: a size/length budget whose ceiling must stay within a
grace band of the high-water mark, so budgets may only tighten, never creep.
Ratchets converted onto the primitive
- windows-test-parity-guard.test.cjs: rmSync rule deleted (now ESLint-enforced);
the remaining six patterns moved from integer baselines to named-set
allowlists with ratchet-down.
- scripts/lint-test-file-count.{cjs,allowlist.json}: per-module integer counts →
named filename sets (closes the swap-a-file-keep-the-count blind spot); a
module dropping under cap now FAILS to force pruning its allowlist entry.
- enh-2790 skill-count `<= 63` → named skill allowlist (ratchets toward ~58).
Size budgets hardened (tighten-only)
- agent-size / workflow-size / feat-3039 help-tiered: ceilings lowered to the
current high-water mark and an assertTightCeiling anti-creep check added per
tier. Fixed external-contract limits (description ≤100 chars, agent ≤100 KB)
are intentionally left as-is — they are not grandfathered creeping budgets.
No user-facing behavior change (tests + tooling only); no USER_FACING_PREFIXES
touched, so no changeset fragment is required.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
091b7e2c4b |
perf(#305): single-pass max-by-mtime for statusline todo lookup (#385)
Replace the per-render readdirSync().filter().map(statSync).sort() chain with a single-pass max-by-mtime loop. Drops the O(n log n) sort and the throwaway intermediate array; I/O and resolved-file behavior are identical. Adds the first behavior-lock test for the todo-resolution path. The larger disk-backed cache win from the issue is deferred: statusline is a fresh child process per render, so any cache must be disk-backed with invalidation/atomic-write design that needs maintainer input. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
6913dbcdb1 | fix(10): centralize semver comparison policy across hooks and changeset | ||
|
|
dbc09d21a6 | fix(3153): handle numeric 100 percent and block-list next_phases | ||
|
|
0ea443cbcf |
fix(install): chmod sdk dist/cli.js executable; fix context monitor over-reporting (#2460)
Bug #2453: After tsc builds sdk/dist/cli.js, npm install -g from a local directory does not chmod the bin-script target (unlike tarball extraction). The file lands at mode 644, the gsd-sdk symlink points at a non-executable file, and command -v gsd-sdk fails on every first install. Fix: explicitly chmodSync(cliPath, 0o755) immediately after npm install -g completes, mirroring the pattern used for hook files throughout the installer. Bug #2451: gsd-context-monitor warning messages over-reported usage by ~13 percentage points vs CC native /context. Root cause: gsd-statusline.js wrote a buffer-normalized used_pct (accounting for the 16.5% autocompact reserve) to the bridge file, inflating values. The bridge used_pct is now raw (Math.round(100 - remaining_percentage)), consistent with what CC's native /context command reports. The statusline progress bar continues to display the normalized value; only the bridge value changes. Updated the existing #2219 tests to check the normalized display via hook stdout rather than bridge.used_pct, and added a new assertion that bridge.used_pct is raw. Closes #2453 Closes #2451 Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
53078d3f85 |
fix: scale context meter to usable window respecting CLAUDE_CODE_AUTO_COMPACT_WINDOW (#2219)
The autocompact buffer percentage was hardcoded to 16.5%. Users who set CLAUDE_CODE_AUTO_COMPACT_WINDOW to a custom token count (e.g. 400000 on a 1M-context model) saw a miscalibrated context meter and incorrect warning thresholds in the context-monitor hook (which reads used_pct from the bridge file the statusline writes). Now reads CLAUDE_CODE_AUTO_COMPACT_WINDOW from the hook env and computes: buffer_pct = acw_tokens / total_tokens * 100 Defaults to 16.5% when the var is absent or zero, preserving existing behavior. Also applies the renameDecimalPhases zero-padding fix for clean CI. Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
cc04baa524 |
feat(statusline): surface GSD milestone/phase/status when no active todo (#1990)
When no in_progress todo is active, fill the middle slot of gsd-statusline.js with GSD state read from .planning/STATE.md. Format: <milestone> · <status> · <phase name> (N/total) - Add readGsdState() — walks up from workspace dir looking for .planning/STATE.md (bounded at 10 levels / home dir) - Add parseStateMd() — reads YAML frontmatter (status, milestone, milestone_name) and Phase line from body; falls back to body Status: parsing for older STATE.md files without frontmatter - Add formatGsdState() — joins available parts with ' · ', degrades gracefully when fields are missing - Wrap stdin handler in runStatusline() and export helpers so unit tests can require the file without triggering the script behavior Strictly additive: active todo wins the slot (unchanged); missing STATE.md leaves the slot empty (unchanged). Only the "no active todo AND STATE.md present" path is new. Uses the YAML frontmatter added for #628, completing the statusline display that issue originally proposed. Closes #1989 |