0c413bbc9c10c8a501c7c5eadd90a643418c593c
111 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0c413bbc9c |
chore(#3059): close the ESLint glob-coverage escape and guard it (#3277)
* chore(#3059): close the ESLint glob-coverage escape and guard it 62 tracked source files matched no `files:` glob in eslint.config.mjs, so ESLint skipped them entirely while `eslint .` still exited 0 — including all 26 files under hooks/, the enforcement machinery itself. Covers 56 of them (eslint-rules/, hooks/, bin/lib/, pi/, examples/, vscode/, the plugin shims, root *.mjs) and allowlists the 6 deliberate must-not-compile brand-typing fixtures with a recorded reason each. hooks/** is covered with n/no-process-exit deliberately off: a hook's whole contract is its exit code, several exits are load-bearing stdin-timeout guards where nothing else terminates the process, and ADR-0012/0174 scope the no-process-exit convention to the Command Routing Hub. bin/lib/ui-safety-gate.cjs is dual-mode, so it keeps the rule live and takes two targeted disables in its require.main===module tail instead. Adds scripts/lint-eslint-glob-coverage.cjs + a node:test drift guard so the class cannot regrow: allowlist entries require a non-empty reason, the list ratchets down only (a stale entry fails), and a tracked-count floor means a broken `git ls-files` fails rather than reporting clean. Closes #3059 * chore(#3059): apply review findings — correct the changeset count, add parser properties Isolated adversarial review, confirmed by rebuilding a byte-for-byte replica of the pre-change eslint.config.mjs: the changeset claimed 44 previously- unlinted files. The real figure is 56 (56 covered + 6 allowlisted = 62). That was user-facing CHANGELOG text and was wrong; corrected, along with three consequential figures in the design record. CLAUDE.md requires a fast-check property test for parsers and budget limits, and listTrackedSourceFiles is a parser. The standards review called this "satisfied in spirit"; it is not. Adds three properties driving the real exported parser through an injected execFile: extension totality/soundness including a trailing terminator, backslash-normalization totality, and CRLF/LF equivalence — the invariant the repo's recurring CRLF defect class breaks. Also de-duplicates the anchor rows onto one shared resolver, kept deliberately independent of the guard's own resolveFileCoverage so an anchor still fails if that resolution regresses, and records in the guard's header why the bin/install.js family is NOT allowlisted: it resolves to 2 rules under ADR-1703, so an entry would trip the allowlist_stale ratchet. * fix(#3059): make the coverage guard's git call container-safe The remote runner reported the guard degrading to `git_failed` on both Node lanes: fatal: detected dubious ownership in repository at '/work' The runner executes in a container where the repo is owned by a different UID, so git refuses to operate on it. The guard's degraded-verdict path worked exactly as designed — it reported the failure instead of throwing or falsely reporting clean — but a guard that cannot run in CI is not a gate. `git ls-files` is now invoked as `git -c safe.directory=* ls-files`. `-c` scopes the override to the single invocation and mutates no config file, and the wildcard is appropriate because this command only enumerates tracked paths in the repository it is already executing inside. Adds a regression test that captures the argv through the injected execFile seam and asserts `-c safe.directory=*` precedes `ls-files`, so the container case is pinned behaviorally rather than by reading the script's source. * chore(#3059): backfill changeset PR number pr:0 placeholder replaced with the real number now that #3277 exists. --------- Co-authored-by: sim <sim@local> |
||
|
|
343835facc |
refactor(#3183): route live-plan counting through scanPhasePlans (#3199)
* refactor(#3183): route live-plan counting through scanPhasePlans scanPhasePlans becomes the sole owner of the live-plan derivation. Twenty-one independent re-derivations across seven modules now route through it, and scripts/lint-plan-count-drift.cjs reports zero, scanning the whole repo rather than an allowlist (ADR-3180 Decision 4a). The epic scoped this at three copies. A whole-repo guard found twenty-six sites across nine files, so Phase 1 absorbs every live-plan re-derivation and Phase 3 narrows to window plus sentinel enumeration. Two sites are exempt with a documented reason rather than a bare allowlist: audit.cts scans one quick task's own directory for a single completion record, and gsd2-import.cts reads a foreign GSD-2 tasks/ layout during a one-time import. Neither is a phase directory. scanPhasePlans gains allPlanFiles (pre-supersession) alongside planFiles so one owner answers both questions: verify.cts's numbering-gap check wants every plan on disk, its pairing check wants the live set. Both fields are additive. Highest-severity fix: cmdPhasePlanIndex, which feeds execute-phase wave scheduling, was scheduling status:superseded plans into waves and reporting zero plans for the post-#3139 nested layout. filterPlanFiles and filterSummaryFiles are deleted; getPhaseFileStats orphaned them and only their own tests still called them. New leaf module src/planning-scope.cts carries the frozen SCOPE discriminator, with its six-gate ripple closed: gitignore, inventory manifest, INVENTORY.md and the CONTEXT.md glossary. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * docs(#3183): amend ADR-3180 for the Phase 1/3 boundary re-slice The contract held; the phase boundary did not. The whole-repo drift guard found 26 re-derivations across 9 files against the epic's estimate of 3, and cmdProgressRender re-derives both enumeration and plan counting on adjacent lines, so DW4 was unsatisfiable within Phase 1's original file scope. Records the amended scope, scanPhasePlans's new allPlanFiles field, findOrphanSummaries, the two documented exemptions, the re-derived Tier-2 table, and the describeNonCanonicalPlans trap for later phases. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * fix(#3183): complete the canonical pairing rule and gate the naming diagnostic The remote runner went red with 13 deterministic failures on both lanes, and they were right: replacing verify.cts's canonicalPlanStem pairing with summaryCandidates dropped a case the bespoke rule covered. A plan carrying a descriptive slug after its id (68-01-scaffolding-PLAN.md) pairs with its canonical-stem summary (68-01-SUMMARY.md), and summaryCandidates generated no such candidate, so the plan read unsummarized. The fix is to complete the one rule rather than restore a second: summaryCandidates gains a canonical-id candidate, narrowed to fire only when an id pair was actually extracted. countMatchedSummaries, findUnsummarizedPlans and findOrphanSummaries all inherit it. The two-plans-one-summary collision behaviour of the original rule is preserved deliberately and documented in place. Second defect, independently root-caused while verifying: routing the #2893 naming diagnostic through scanPhasePlans exposed it to the loose /PLAN/i fallback, which is correct for counting and wrong for a naming check — a non-canonically-named file was accepted as a valid plan and the diagnostic went silent. cmdPhasesList, cmdFindPhase and cmdPhasePlanIndex now intersect with a strict isCanonicalPlanFile predicate before reporting names. Same class as the describeNonCanonicalPlans trap already recorded in ADR-3180: a question about file naming wants the physical, strictly-matched set; only a question about outstanding work wants the live set. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * chore(#3183): register planning-scope.cjs in the eslint migration list tests/repo-invariants.test.cjs asserts every bin/lib/*.cjs is linted xor ignored per its ADR-457 migration state. The new planning-scope module closed five of the six .cts ripple gates - gitignore, inventory manifest, INVENTORY.md and the CONTEXT.md glossary - but not eslint, because that one is enforced by a test rather than by lint:ci, so the local pipeline stayed green while it was missing. Generated from src/planning-scope.cts, so the .cjs is ignored and the .cts is linted, matching every other migrated module. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * fix(#3183): replace the plan-count drift detector with a literal tokenizer CodeQL reported 4 high-severity js/redos alerts on REGEX_LITERAL_MD_RE, the backtracking regex that finds "a regex literal mentioning PLAN/SUMMARY and an escaped \.md". Five review rounds found it had two defects, not one: - EXPONENTIAL, then CUBIC. Its "any char" atom `(?:\\.|[^/\r\n])` let a `\.` pair be consumed either as one escape or as two class characters, which is exponential backtracking: 27,464ms on `"/\.mdplan" + "\.".repeat(28) + "X"`. Excluding `\` from the class killed that but left a cubic path — 23ms at N=200, 172ms at N=400, 1362ms at N=800 on `"/" + "PLAN\.md".repeat(N)` with no closing `/`. This guard is the last stage of `npm run lint:ci`, which CI runs on fork pull requests, so a crafted src/*.cts could stall the job. - A DETECTION HOLE. A character class holding a bare, unescaped `/` — e.g. `/SUMMARY[^/]*\.md$/`, an ordinary path-excluding filter — terminated the literal at that `/`, so the scan never reached `\.md` and the guard missed it entirely. (Classes holding an ESCAPED `\/` were already matched; the tests cover those separately as parity, not as regressions.) Both defects have one root cause: regex-literal grammar — `\x` escapes, and `/` inside `[...]` not terminating — is not expressible in a backtracking regex. So the detector is now a tokenizer, not a regex. readRegexLiteralAt reads the literal at a given `/` in a single left-to-right pass with no backtracking, treating escapes as two-character units and suppressing the `/` terminator inside a character class. findRegexLiteralMdMatch restarts it at every `/` on the line, preserving the old "find anywhere" behaviour; MAX_REGEX_LITERAL_LEN (400) bounds each read — including the trailing-flag scan — which keeps the whole-line cost linear. Results: cubic shape flat at 0.06-0.39ms out to N=3200 (25KB), exponential shape 0.01ms at 28 reps and 0.00ms at 64, and the bare-`/` class shapes are now caught. Differential against the old regex over 28,474 lines (those matching FILENAME_TEST_RE but not PLAN_SUMMARY_LITERAL_RE, across src/tests/scripts/ gsd-core/bin/eslint-rules, excluding 265 lines with >6 backslashes on which the old regex hangs): 6 differences, all the tokenizer returning the fuller or newly-correct literal, 0 old-only misses. The `\.md` token stays case-insensitive, matching the `/i` the old regex carried. Also closes three holes in the same new file: - walk() tested entry.isFile(), false for a symlink, so a symlinked src/*.cts was silently unscanned — an evasion of a guard whose stated principle (ADR-3180 Decision 4a) is whole-repo discovery with no allowlist. It now resolves symlinks, but confined: file links must resolve inside the repo root, directory links inside the scanned dir itself. Every sibling drift guard in scripts/ uses the Dirent classification and never follows links, so following them unconfined would have made this the only linter able to read outside the tree — on fork PRs an arbitrary out-of-repo read whose matched fragments reach a public CI log. The narrower directory rule additionally stops `src/up -> ..` from sweeping the whole repo, and the skip list is now checked against resolved paths so `src/g -> ../.git` cannot reach .git/** or node_modules/**. Real paths are de-duplicated and files reported canonically, so a symlink alias cannot shift which FUNCTION_SCOPED_EXEMPTIONS key applies. - Both the reported fragment and the reported FILE PATH are attacker- controlled source text written straight to a CI log, and git permits control bytes in a filename. Both are now escaped — C0/C1/DEL plus the bidi and zero-width controls — so a crafted literal or filename cannot recolour the log, overwrite a line with CR, or fabricate a line that looks like this guard's own success output. Regression coverage in tests/plan-count-single-owner.test.cjs: a child-process probe over both pathological shapes (catastrophic backtracking is synchronous and would freeze the suite rather than fail one test), the bare-`/` class shapes verified to fail against the parent-commit blob, root-confinement tests covering the outside-file, outside-directory, cycle, broken-link and duplicate cases, direct isInsideRoot coverage including the sibling-prefix case that a bare startsWith would let through, sanitizeForReport coverage, and limit-1/limit/limit+1 coverage of MAX_REGEX_LITERAL_LEN derived from the exported constant. The earlier structural assertion was dropped — it checked for the substring `[^/`, which respelling the class as `[^\r\n/]` defeats while staying exponential. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma * chore(#3183): backfill changeset PR number Restores b77931869, which a force-push during the ReDoS remediation dropped. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012qYy4ZWif3sscQyMsup6Ma --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9faacc0c15 |
test(#3148): bound the long tail and delete the unbounded-spawn allowlist (#3192)
* test(#3148): bound the long tail and delete the allowlist Migrates the final 170 unbounded sync spawn sites across 49 files, then removes the allowlist entirely. local/no-unbounded-spawn now runs with no exemption surface across tests/**: there is no file to add a name to. drift-detection's throw-native git() helper routes to gitOrThrow -- bare runGit would have taken 16 call sites quiet on failure. commands.test.cjs has two independently-scoped runGsdTools/runCli helpers, one already bounded and one not; they are kept distinct rather than unified, the same trap as the two same-named git() helpers in Wave 1. runNpm's bound was erasable. Its options spread callerOptions after the defaults, so an explicit timeout:undefined silently dropped the 180000ms bound -- the rule flagged it and was right; it was not a false positive. Fixed by destructuring with a default, with a test that fails when the default is removed. Two sites stay on a raw spawn with an explicit timeout because the seam cannot express them: one needs shell:true for npm.cmd on Windows, one redirects stdout to a real fd. Both are the rule's own documented second option, not an escape from it. Closure verified rather than asserted: the derivation scan reports 0 unbounded spawn helpers and 0 unbounded direct git call sites, and a temporary file carrying an unbounded spawn still errors with the allowlist gone. Closes #3064. * test(#3148): close a hole in the guard's own eslint-disable ban The ban listed only the top level of tests/, so it was blind to 37 .cjs files under tests/helpers, qa, observability, fixtures and dispatch. With the allowlist deleted this test is the sole remaining way to detect someone silencing the rule inline, so the gap was load-bearing: a nested file could carry an unbounded spawn plus an eslint-disable and pass everything. Proven before and after. A probe planted under tests/helpers with both was invisible to the guard and clean under eslint; after making the listing recursive the guard fails on it. The scanned set goes from 771 files to 808. Pre-existing since the guard shipped, but this wave is what promoted it to sole defense, so it is fixed here rather than filed. Also converts the last hand-rolled throw check to throwIfFailed and the last re-derived legacy shape to compose toLegacyResult, which makes the epic's none-remain claim true rather than nearly true. toLegacyResult itself is not widened -- eight callers depend on its shape and one consumer does not justify changing a shared contract. * fix(#3148): correct seam incoherence at the bound and a slow review-lane error path Two real failures from the remote runner, both fixed at the cause. The seam could return outcome TIMED_OUT together with exitCode 0. At the exact bound spawnSync reports ETIMEDOUT while the child has already exited with a real status, and toSeamResult classified on the error code while passing status straight through -- an incoherent pair its own boundary test was written to catch, and did. A status that is not null is direct evidence the child exited on its own, so it now decides the outcome before the error-code branches run. process-seam.cjs was deliberately untouched by every earlier wave; this is a defect in the module itself, kept surgical, with a unit test that fails against the old logic. review-lane with an unknown subcommand fell through to its usage error only after loading the capability registry and building a per-lane plan, which spawns one child process per lane -- up to twelve. The error path took ~1288ms instead of ~119ms, and under bench load it outran a caller's spawn timeout and was killed before writing anything, which is the empty stdout and stderr CI saw. It now fails fast before any of that work begins. This is the epic's first production change. It is user-facing, so it carries a changeset rather than a no-changelog label. * test(#3148): replace a real-race timeout test with a deterministic one E9 raced git rev-parse against a 1ms bound and assumed git always lost. On a warm container git finishes first, spawnSync returns status 0 with no error at all, the seam correctly classifies EXITED, and gitOrThrow correctly does not throw -- so the test failed on both lanes. A probe confirms a genuine timeout always carries status null, so this was never the seam misbehaving. Raising the bound would only lengthen the odds, which is the same defect with better luck. The test now drives gitOrThrow against a stubbed runGit that returns a synthetic TIMED_OUT result, so it asserts exactly what it always meant to -- that a timeout propagates as a throw -- with no timing dependence. Five consecutive runs are identical where the old one varied. I wrote this test in Wave 0; it is a real-race test by construction and CLAUDE.md says to replace those rather than re-run them. * chore(#3148): backfill changeset PR number 3192 --------- Co-authored-by: sim <sim@local> |
||
|
|
27aa40f65e |
fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ directory (#3175)
* test(#3023): failing-first guard — pi must not stage hooks in its reserved dir pi reserves <configDir>/hooks as its deprecated extension location and warns on every startup when it exists. Assert a pi install stages the shared hook bundle under gsd-hooks/ instead, manifests it there, and never creates hooks/. Also adds pi to the local-scope dir table in install-shared.cjs: pi was in RUNTIME_META but not LOCAL_DIR_NAME, so scope:'local' resolved path.join(root, undefined) and no local pi install could be exercised. Fails before the fix. Verified via the remote runner. * fix(#3023): stage pi's shared hook bundle outside pi's reserved hooks/ dir pi reserves <configDir>/hooks as its now-deprecated extension location and warns on every startup when that directory merely exists — checkDeprecatedExtensionDirs() guards the warning with a bare existsSync(), unlike its tools/ sibling. GSD staged its shared hook bundle exactly there, and pi's advised remediation (move it to extensions/) would break the adapter's paths and expose GSD's .js helpers to pi's extension auto-discovery. The bundle directory name is now runtime-descriptor-driven: hostBehaviors .sharedHooksDirName, defaulting to 'hooks' so all 18 other runtimes are byte-identical. pi sets 'gsd-hooks'. The name is validated as a single path segment — separators, dot-only segments, trailing dots, absolute paths, NUL, and Windows reserved device names all fall back to the default, because the value is joined onto a user's config root and written to. Renamed in place rather than relocated: hook scripts resolve siblings via __dirname/.., so a depth change would silently break them. - install / uninstall / manifest sites all read the resolved name - pi/gsd.cjs probes gsd-hooks then hooks, so dev checkouts and half-upgraded trees still resolve; the never-throws contract is preserved - new migration 009 retires the legacy pi hooks/ dir on upgrade, using a new non-recursive remove-empty-dir engine primitive (rmdirSync only, symlink-refusing, containment-guarded); ADR-0008 amended accordingly - fixes two latent name-dependencies the rename exposed: the stale-hook scan and the injection scanner's self-exclusion both hardcoded 'hooks' Verified on the remote runner. Closes #3023 * fix(#3023): close review findings and align emitted provenance with the rename Adversarial review found two defects, and the remote runner found four failure clusters. All fixed here. Review BLOCKER — detect-custom-files was blind to the renamed bundle. GSD_PREFIX_MANAGED_DIRS in gsd-tools.cjs hardcoded 'hooks', so for pi the whole gsd-hooks/ tree was invisible to the custom-file scan and user-added files there were never backed up before the next update's clean-install wipe. The dir set now resolves via the .gsd-runtime marker plus the shipped capability registry (never bin/install.js, which is not shipped into installed trees), and falls back to scanning every known candidate when the runtime cannot be determined — over-scanning is safe, under-scanning is the data loss. Review MAJOR — the pi adapter bound to an empty bundle. resolveSharedHooksDir accepted any directory, so an interrupted install left gsd-hooks/ winning over a fully-staged legacy hooks/ and every hook silently no-opped. A candidate now qualifies only if it is non-empty. Remote-runner clusters: - emitted-provenance had no rule for the gsd-hooks/ family; added two pi-scoped rules pointing at the same sources the existing hooks/ rules use. The table is total, so an unattributed family is a hard failure by design. - pi tests in install-minimal-hooks and the install integration suite asserted the old layout; updated to derive the dir name from the descriptor rather than hardcoding either name. - 19 unrelated-looking failures on node22 only were a leaked fs mock: t.after() runs in registration order, cleanup was registered before mock.restoreAll(), and node22's JS rimraf calls the public fs.rmdirSync while node24's native path does not — so the EACCES stub leaked process-wide on one lane. Restore now runs first. Verified on the remote runner. * fix(#3023): honor PI_CODING_AGENT_DIR, ack the rename ripple, fix expandTilde pi resolves its agent dir as PI_CODING_AGENT_DIR ?? ~/<CONFIG_DIR_NAME>/agent (packages/coding-agent/src/config.ts). GSD's pi descriptor declared an empty configHome.env, so a user with that variable set had GSD installed where pi never looks. Added the env name; the dot-home-nested resolver already handled the override, so no resolver logic changed. Also fixes expandTilde in the shared runtime-homes resolver, found while adding that: it hardcoded os.homedir() and ignored the opts.home every caller threads, so EVERY runtime's tilde-valued env override (claude, antigravity, windsurf, pi) silently resolved against the real home. That is a correctness bug and a test-escape hazard — a sandboxed test asserting on a tilde override reached the developer's actual home directory. Now threaded through every branch; behavior with no injected home is unchanged. Adds the emitted-drift ack fragment for the 58 pi paths whose emitted location moved with the rename. The provenance rules satisfy the totality gate; the differential gate needs the ack because the hook sources are byte-unchanged — only the installer's target directory moved. The two hook files this branch genuinely edits stay attributed and are not double-acked. Note on piConfig.configDir: it is read from pi's OWN installed package.json (getPackageDir walks up from pi's __dirname), alongside piConfig.name — a white-label setting for a redistributed pi fork, not a per-project user setting. Documented accordingly rather than treated as an unsupported override. Verified on the remote runner. * fix(#3023): reject blank env overrides, pin adapter/descriptor parity Three review findings, all fixed. A whitespace-only config-dir override was accepted verbatim: the guard was `if (val)`, falsy only for the empty string, so PI_CODING_AGENT_DIR=' ' resolved to a literal three-space directory name instead of falling back to the descriptor default. Fixed across every env-consuming branch — dot-home, dot-home-nested, all three xdg steps, and generic-agents-root — not just pi's. Non-blank values are still never trimmed, so '~/My Agent Dir' keeps working. pi/gsd.cjs's probe list and the descriptor were two independent sources of truth for the bundle directory name; a future rename would have desynced them silently and left every pi hook quiet with no error. The probe list stays deliberate — it must resolve in a dev checkout and a half-upgraded tree, where the registry's answer would be wrong — so this adds the parity assertion the repo's generative-fix-divergence rule calls for: the descriptor value must be the FIRST candidate, and the default must remain present. Changeset body rewritten to cover the two later user-facing fixes it had not caught up with. Verified on the remote runner. * chore(#3023): backfill changeset PR number * fix(#3023): anchor injection-scan patterns and fix a macOS detection hole CI's security job flagged CONTEXT.md:124 — pre-existing prose reading 'not the same fact as a genuinely empty or absent one'. The match was the 'act as a' INSIDE 'f-act as a': the pattern had no left word boundary, so any word ending in act tripped it (fact, impact, contract, artifact, interact, redact, abstract). My four-line CONTEXT.md edit dragged the latent false positive into this PR because the scan is diff-scoped by file but reads whole files. Anchored with (^|[^[:alnum:]]) rather than rewording maintainer-owned prose, which would have left the class alive for the next PR touching any file saying 'fact as a'. Auditing the rest of the list for the same class surfaced a real detection hole: the eval/exec/Function patterns matched a quote via \x27, a GNU-grep-only hex escape. BSD/macOS grep reads it as four literal characters, so single-quoted eval('...')/exec('...') payloads were NEVER detected there while passing on GNU-grep CI. Replaced with a literal apostrophe class. Boundaries were added only where a real word-suffix collision exists; exec, jailbreak, developer mode and the role-manipulation family were audited and deliberately left unanchored. 22 new cases cover both directions — the false positives now scan clean, and every real payload still fires, including the quote/punctuation/start-of-line boundary forms. Also builds this branch's injection test fixture at runtime instead of carrying the literal phrase, so the payload keeps its teeth without tripping the scan. Verified on the remote runner. --------- Co-authored-by: sim <sim@local> |
||
|
|
2afe17bbdb |
test(#3143): add the no-unbounded-spawn guard and throw-preserving git fixture (#3150)
* test(#3143): add no-unbounded-spawn guard and throw-preserving git fixture Adds the ESLint rule local/no-unbounded-spawn, wired into the tests/**/*.cjs block, plus an allowlist that only ratchets down: a listed file with zero violations reports its own entry as stale. The rule resolves renamed destructures and chained requires rather than matching literal callee names -- both forms exist in the suite today and a name-only matcher leaves them permanently invisible. It resolves an options object held in a single-write const, which is what keeps process-seam.cjs, the bounded reference implementation, from flagging itself. timeout: 0 and anything above the 600000ms ceiling are rejected as only nominally bounded. Adds tests/helpers/git-fixture.cjs so a migrated execSync call site keeps its throw-on-non-zero contract; process-seam.cjs is unchanged. * test(#3143): prove the allowlist guards can actually fail Extracts the D4/D6/D7/D8 checks into pure helpers and drives each against a synthetic fixture carrying an injected violation. Without this the suite only proved that today's clean data passes, which a deleted check would also satisfy. * fix(#3143): close two ceiling and alias escapes found in review Nested arithmetic bypassed the ceiling entirely: the numeric evaluator only resolved a flat literal, so `timeout: 60 * 60 * 1000` (3600000ms, six times the ceiling) fell through to trusted and reported nothing. The evaluator now recurses through arithmetic and unary signs with a depth cap. Alias resolution was traversal-order dependent, not scope dependent: a call textually above its own require destructure saw an empty alias map and reported clean. The map is now built in a Program pre-pass. Also: an explicit timeoutMs:undefined no longer overwrites the git fixture default via spread, adds the missing seam-routed rule test, and de-duplicates the repeated try/catch in the fixture tests. --------- Co-authored-by: sim <sim@local> |
||
|
|
7203011400 |
feat(#3072): ship the deferred MCP served catalog (resources + prompts) (#3083)
* test(#3072): add failing-first coverage for the mcp served catalog 55 input-class rows from the phase test matrix, across four suites: the catalog module over injected readFile/readDir seams, the protocol surface through handleMessage, the install-vs-catalog parity gate, and fast-check properties for uri round-trip, traversal refusal, and pagination partition. src/mcp-catalog.cts lands as a skeleton whose functions throw, so the suites fail on BEHAVIOR rather than on a missing module. The REASON enum is real so tests assert typed codes instead of message prose. Hostile coverage for the one client-controlled path surface (resources/read): dot-dot and backslash traversal, percent- and double-encoded traversal, absolute posix and windows paths, file:// scheme, null byte, symlink escape, unindexed sibling, non-string and empty uri, wrong root segment. IO faults are injected by monkeypatching the seam, never chmod 0o000 - root bypasses mode bits, so a permission-based test silently passes with zero coverage in root CI. Refs #3072 * feat(#3072): serve the mcp catalog as resources and prompts gsd-mcp-server now serves GSD's own content alongside its three tools: the workflow, reference and command tree as MCP resources (resources/list, cursor paginated, and resources/read over gsd://<segment>/<relpath> uris) and the 71 commands/gsd/*.md as MCP prompts keyed by bare command name. initialize advertises resources and prompts, and deliberately does not advertise subscribe or listChanged - the catalog is fixed for a server process lifetime, so declaring a notification we never send would be a lie a host acts on. Composition scope is SHARED, not re-declared. shouldCompose lives in src/mcp-catalog.cts and bin/install.js now imports it instead of carrying its own regex, so the served catalog and the installed file floor cannot drift on what gets composed. Proven behavior-preserving across all 2871 tracked paths plus windows-backslash, absolute and near-miss-prefix cases: zero mismatches. tests/mcp-catalog-parity.test.cjs asserts served text equals the installer composition-stage text over the real tree, with anti-vacuity guards requiring both a marker-bearing workflow and a non-composed file in the comparison set. Two measurements corrected the literal issue text. Composition is scoped to gsd-core/workflows/ only, because a reference or command that documents marker syntax with an unfenced example would otherwise be parsed as carrying a real marker and have that line lossily dropped. And parity is asserted at the composition stage rather than against an emitted runtime tree, since install applies per-runtime path rewrites afterwards and the catalog is host-agnostic, so byte equality with any one runtime would be false by construction. resources/read is the one client-controlled path surface and is guarded in two independent layers: the uri must be an exact key in the prebuilt index, which defeats every traversal string by construction, and the mapped path is then re-checked with validatePath so a symlink planted inside a root after indexing is still refused. Also fixes a real drift defect found while here: SERVER_VERSION was hardcoded 1.7.0 while the package is at 1.9.1. It now resolves lazily from VERSION or package.json, reusing the precedent in runtime-artifact-conversion. Closes #3072 * test(#3072): make the catalog parity gate drive the real installer Review found the parity gate vacuous: it never imported or spawned bin/install.js, and recomputed the installer side with the SAME shouldCompose and composeWorkflow the catalog calls internally. It therefore proved only that src/mcp-catalog.cts is self-consistent. The old row 52 compared shouldCompose against a regex literal frozen in the test file rather than against the installer at all. An inline divergent regex re-added to bin/install.js - the exact regression ADR-1671 asks this gate to catch - would have left the suite green. The gate now spawns a real bin/install.js and compares the composition DECISION, observed as gsd:section marker survival, against what the catalog serves for the same files. Marker presence is the right observable because the installer applies per-runtime path rewrites after composing while the catalog applies none, so raw byte equality between the two surfaces is false by construction and must not be asserted. Sensitivity was proven, not assumed: overlaying the shouldCompose export that bin/install.js imports so it always returns false makes a real spawned install leave autonomous.md's markers in place while the catalog still strips them, and the row 48 assertion diverges. Anti-vacuity guards are kept and extended - the comparison set must be non-empty, must contain a workflow that actually carries markers, must contain a file the predicate declines to compose, and the install must have emitted a non-zero file count. The marker-documenting reference case has no instance in the real tree, so it uses an overlay fixture built with the same technique workflow-fragments-emission.install.test.cjs already uses. Renamed to .install.test.cjs so it lands in the install suite it now belongs to. Refs #3072 * test(#3072): retarget the unknown-method assertion off a now-implemented method tests/gsd-mcp-server.test.cjs used 'resources/read' as its example of an UNKNOWN JSON-RPC method. The served catalog implements that method, so it now returns -32602 (no uri supplied) rather than -32601. The remote runner caught it deterministically on both linux lanes: -32602 !== -32601. The test's intent is still correct and worth keeping, so it is corrected rather than deleted or weakened. It now uses 'resources/subscribe', which the server deliberately does not implement and deliberately does not advertise in initialize's capabilities, because it never sends the corresponding notification. That turns the assertion into a real contract - the advertised capability surface and the implemented method surface agree - instead of an arbitrary method name a future feature could invalidate the same way. Swept the rest of the suite for other assertions pinning the newly implemented methods; this was the only one. Refs #3072 * chore(#3072): backfill changeset PR number 3083 * test(#3072): make the catalog fake fs separator-agnostic for windows CI caught this on windows-latest (22 and 24): every catalog fixture indexed ZERO entries, surfaced by the anti-vacuity guards as 'fixture catalog must actually index resources for this property to mean anything'. Mechanism: makeFakeFs keyed its dirMap/fileMap on POSIX-joined paths (${root}/${rel}), while production buildCatalog looks paths up with path.join, which is backslash-separated on Windows. Every lookup missed, tryReadDir returned null, and the catalog came back empty. Production is NOT at fault and is unchanged. The same CI run proves it: on windows-latest the real-filesystem tests all passed, including 'installer composition decision matches the served catalog for every file in the real installed tree' and the row-51 non-vacuity proof against a real spawned installer. A real Windows fs accepts both separators; the FAKE did not, so the fake was the unfaithful one and is what changed. Lookup keys are now normalized unconditionally with .replace(/\\/g,'/') in readDir and readFile - never path.sep-conditional, never platform-gated. The row-42/43 injected-fault wrappers got the same treatment, since they compared raw production paths against POSIX-literal fixtures. No assertion was weakened, and the anti-vacuity guards that caught this are untouched - they are the reason this surfaced as a loud failure instead of a suite that silently asserted nothing on Windows. Refs #3072 --------- Co-authored-by: sim <sim@local> |
||
|
|
5fd5c81042 |
test(#3055): add the process seam so a subprocess timeout is expressible as data (#3066)
* test(#3055): add the process seam and route runGsdTools through it Adds tests/helpers/process-seam.cjs — runNode/runGit/runHook over spawnSync, each returning a typed discriminated union { outcome, exitCode, stdout, stderr, timedOut, signal, killed, code }. Every call is timeout-bounded; there is no unbounded path. runGsdTools becomes an adapter over the seam. Its legacy { success, output, error, exitCode } shape and retry-once-on-kill behaviour are preserved byte-identically, so none of its 136 caller files change. Outcome discrimination was corrected against probed runtime behaviour rather than assumption: a timeout and a maxBuffer overflow are identical on both status (null) and signal (SIGTERM), and differ only by code (ETIMEDOUT vs ENOBUFS). Overflow is therefore classified before timeout. This fixes a live defect — the previous isKilled() treated an overflow as a kill, retried it for a second full 60s run, and then reported "host OOM or scheduler contention" for a child that had merely printed too much. Also widens the ESLint tests glob from tests/**/*.test.cjs to tests/**/*.cjs, which brought 31 previously unlinted shared helpers under the same rules their sibling test files already obey, and fixes the 5 violations that surfaced — including a bare npm invocation without shell:true in tests/helpers/emitted-runtime.cjs (DEFECT.WINDOWS-TEST-PORTABILITY), now routed through the existing portable runNpm helper. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3055): migrate every local spawn wrapper onto the process seam Replaces the spawn body of all 25 local runHook/runGuard/runGate definitions with a call to tests/helpers/process-seam.cjs. Each wrapper keeps its name, parameter list, return shape and post-processing (JSON parse, ANSI strip, env sanitising, field extraction) — only the spawn mechanism changes, so no test assertion moves. The 4 bash-driven wrappers use the seam's explicit `interpreter` option rather than a fourth primitive; it is explicit rather than inferred from the file extension, because guessing an interpreter from a path fails silently when a script's name does not match its shebang. Seven wrappers were previously unbounded and now carry an explicit timeout sized to what each actually runs, not the seam default. Two of those seven (gsd-write-guard, lint-docs-command-form) were absent from the issue's inventory entirely and were found by scanning after the migration. Adds the CONTEXT.md `### Process seam` glossary entry and a CONTRIBUTING.md reference section covering the three primitives, the discriminated union, and the two rules the seam enforces. Scope disclosure recorded in the phase design notes: the issue scoped three identifier names. A scan for local helpers that spawn AND return the spawn result finds 113 across 82 names, 71 of them unbounded, plus 122 unbounded direct git call sites. This change bounds 25 of those. The remaining surface is the same defect class and is NOT closed by this PR. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): classify an externally-killed child as KILLED, not EXITED Blocker found in this branch's own diff, independently confirmed by an isolated reviewer. A child killed by an external signal — a genuine bench OOM kill — makes spawnSync return { status: null, signal: 'SIGKILL' } with NO .error field. The seam's "no error implies EXITED" rule therefore classified it as a clean exit, and runGsdTools returned { success: false, exitCode: 1 } without retrying. That silently defeated the #969 kill-discrimination for precisely the case it was built for: the old isKilled() fired on `signal != null`, retried once, then threw a labelled resource-starvation error. A real OOM would have been reported as an ordinary assertion failure. Adds a fifth outcome, KILLED, for "no error but a signal is set", and makes the adapter retry on TIMED_OUT or KILLED — reproducing the old `killed || signal != null || code === 'ETIMEDOUT'` condition exactly. SPAWN_FAILED still does not retry (matching the old behaviour, where signal was null). BUFFER_OVERFLOW still does not retry, which remains a deliberate divergence: the old code retried it because signal was SIGTERM, burning a second 60s run on a child that had merely printed too much. All five outcomes verified against the live runtime rather than assumed: SIGKILL -> killed, exit 0/7 -> exited, timeout -> timed_out (ETIMEDOUT), >1MB stdout -> buffer_overflow (ENOBUFS). Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): address standards-review findings on this branch Three findings from the standards axis of the review, all in this branch's own diff. The CONTEXT.md glossary entry this branch introduced was already stale on the branch's own last commit: it enumerated a 4-member OUTCOME while the code had 5, because the KILLED fix did not update it. That is precisely the drift the "module changes update Domain-terms" gate exists to catch, so the entry now lists all five and explains KILLED. api-coverage-gate-e2e compared an outcome against the raw string 'exited' rather than OUTCOME.EXITED, the only such outlier; the enum is now imported and used. A sweep for the other four outcome literals found no further comparison sites. Three call sites hand the literal bash flag '-c' to the seam's first parameter, which the JSDoc described as an absolute script path. Rather than add a fourth primitive, the contract is corrected to match reality: the parameter is renamed `target` and documented as the first argv element handed to the interpreter — normally a script path, but for an interpreter invoked with an inline program it may be that interpreter's own flag. No behaviour change. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#3055): assert the cross-platform timeout contract, not the macOS one The remote runner failed on both Linux lanes (node 22 and node 24, identical) while the same tests passed locally on macOS. Two assertions encoded a platform-specific behaviour as a cross-platform guarantee. When spawnSync times out, macOS preserves the child's partial stdout/stderr; Linux discards it and returns empty strings. Verified on node v26.5.1 both ways. The seam passes through whatever spawnSync hands it and cannot manufacture output that was discarded, so the production code was correct — the tests were wrong. Both tests now assert the guarantee the seam actually makes on every platform: outcome TIMED_OUT, timedOut true, and stdout/stderr always being strings rather than undefined or a Buffer. The partial-content assertions are retained behind an explicit process.platform === 'darwin' guard so the macOS coverage is not lost, and the first test is renamed to say what it now guarantees. This is the failure mode the remote matrix exists to catch: local macOS verification would have shipped it. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#3055): classify a failed spawn as SPAWN_FAILED, not a timeout Windows CI caught two defects the Linux matrix could not. tests/context-predicates-query.test.cjs passes a 32K-char argv value. On Windows that exceeds the argv limit and spawnSync fails with code ENAMETOOLONG, signal null, status null. The seam's fallback rule — "otherwise, status === null implies TIMED_OUT" — swallowed it, so the adapter retried a spawn that can never succeed and then threw the resource-starvation error. The old isKilled() returned false for that shape and returned an ordinary failure result. TIMED_OUT is now identified positively: code === 'ETIMEDOUT' OR signal is set. Anything else carrying an error is SPAWN_FAILED, which covers ENAMETOOLONG, E2BIG, EACCES and ENOENT alike. The signal clause is what keeps a platform whose timeout errno differs classified correctly, so the greedy catch-all is no longer needed. The second defect is a contract regression I introduced and had claimed otherwise. That same test asserts `typeof r.exitCode === 'number'`, and toLegacyShape was returning null for BUFFER_OVERFLOW and SPAWN_FAILED, so the assertion failed on type. The old code returned `err.status ?? 1` on every non-retried failure path. The adapter now returns 1 again for both, and the comment claiming "never coerced to exitCode:1, unlike the pre-seam helper" is retracted: the seam keeps the richer truth (exitCode null plus a distinct outcome), the legacy adapter keeps the old numeric contract its callers actually depend on. Verified on this host: a 4MB argv yields E2BIG -> SPAWN_FAILED; ENOENT -> SPAWN_FAILED; timeout -> TIMED_OUT; >1MB stdout -> BUFFER_OVERFLOW; SIGKILL -> KILLED; clean exit -> EXITED. Refs #3051 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ef823ca9d9 |
fix(#2830): propagate a halted plan to its transitive dependents (#3038)
* test(#2830): add failing regression tests for halted-plan dependent blocking Add tests/fix-2830-halted-plan-dependents.test.cjs covering direct, transitive (2 and 3 hop), and diamond dependents of a halted plan across both independent "which plans are incomplete" readers (phase-plan-index's cmdPhasePlanIndex and findPhaseInternal/searchPhaseInDir), the negative case (an unrelated decoupled plan stays runnable), and a parity check that the two readers agree. Uses only modules that already exist at this commit (gsd-tools.cjs via subprocess, the pre-existing phase-locator.cjs) so the test file loads and runs cleanly on a fresh clone of this exact commit. These fail against current behavior: neither reader has any concept of a halted plan or a blocked_by/runnable view yet. * fix(#2830): a halted plan no longer leaves its dependents on the runnable work list A plan that reaches a designed stop still writes a SUMMARY, so both "which plans are incomplete" readers saw it as an ordinary completion and reported its dependents as ordinary runnable work — never checking whether an upstream plan had halted rather than finished. - New `status: halted` frontmatter value, documented in all four SUMMARY templates alongside the existing `status: complete`. - New shared src/plan-dependency-graph.cts: a single computeHaltPropagation pass that both phase.cts's cmdPhasePlanIndex (wave-grouping) and phase-locator.cts's searchPhaseInDir (the phase-location primitive, ~50 dependent symbols across 5 command routers) now call, so the two-implementation divergence that caused this bug cannot recur. It accepts an optional precomputedOrder so cmdPhasePlanIndex — which already runs Kahn's algorithm in computeDependencyLevels for wave assignment — passes that order straight through instead of a second traversal; searchPhaseInDir (no prior traversal) lets the module derive its own. The two small duplicated predicates each reader would otherwise carry (is this status "halted"?, which summary file matches which plan id?) are centralized in the same module as isHaltedStatus/buildSummaryFileIndex. - Additive fields only: `halted`/`blocked_by`/`runnable` on cmdPhasePlanIndex's plans[] and top level, `halted_plans`/`blocked_by`/ `runnable_plans` on searchPhaseInDir's result. The pre-existing `incomplete`/`incomplete_plans` fields are unchanged in meaning and membership. - execute-phase.md's discover_and_group_plans step now also skips any plan whose `blocked_by` is non-empty, reporting it by name with its blocking chain, in addition to (not instead of) the existing has_summary skip rule. Extends tests/fix-2830-halted-plan-dependents.test.cjs (introduced in the prior commit) with direct unit coverage of computeHaltPropagation (including the precomputedOrder call shape) and a fast-check property test — both only possible once this commit's new module exists. Closes #2830 * fix(#2830): surface the halt-aware view from init execute-phase The adopted work made phase-locator compute halted_plans / blocked_by / runnable_plans, but cmdInitExecutePhase builds its output by explicitly enumerating fields, so all three were computed and then silently dropped at the exact consumer the issue names as regressed. Forwards them additively -- incomplete_plans and incomplete_count keep their name, type and semantics byte-for-byte -- and adds the same three empty defaults to the roadmap-only fallback so the shape is consistent in both branches. Covered by a new test that drives the real CLI end to end rather than the locator function, since the locator already worked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2830): fail closed on dependency cycles and stop the templates inviting the defect Three review findings, all fixed: - BLOCKER (isolated adversarial). Cycle participants never reach indegree 0 in the Kahn pass, so they were excluded from the topological order, never visited by the forward pass, and vanished from blocked_by entirely -- i.e. reported as runnable. The wave-grouping reader hard-fails on a cycle so it never hit this, but the phase-location reader does not, so init execute-phase offered a plan depending directly on a halted plan. Reproduced, then fixed in the shared engine so every consumer is safe regardless of pre-checks: a node absent from the order is now blocked with a deterministic, non-empty named cause. A plan silently missing from both blocked_by and runnable is the exact disappearance this issue exists to prevent. - MAJOR (isolated adversarial). All four summary templates showed the field as an inline comment on the value line. Frontmatter parsing does not strip trailing comments, so an executor copying the templates' own presentation wrote a halt that parsed as a non-halted string, silently reproducing the original bug. Guidance moved off the value line, and the halt predicate now tolerates an unquoted trailing comment. - HARD standards violation. A test regex-matched child-process stderr prose for /cycle/i, which CONTRIBUTING bans. Replaced with the structured failure signal plus a differential assertion (same fixture without the cycle edge must succeed), so it stays cycle-specific without matching prose. Also folds the duplicated read-summary-and-check-halted wrapper out of both readers into the shared module -- centralizing only the predicate left the exact two-copies-that-drift pattern the module exists to prevent -- and commits the artifact-types documentation for the new status value. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(#2830): stop the property generator hanging the whole suite The remote runner did not fail -- it hung. Two containers sat in this file for 31+ minutes, and an earlier attempt ran 9 hours before I killed it. The runner passes --test-timeout=0, so nothing ever reaps it: this would have hung CI indefinitely, not reported a failure. Root cause: the DAG generator built edges by rejection -- from: fc.integer({ min: 0, max: n - 1 }) to: fc.integer({ min: 0, max: n - 1 }) .filter(({ from, to }) => from < to) With n === 1 both integers are forced to 0, so the predicate is unsatisfiable and fast-check retries value generation forever. n is drawn from 1..12 and fast-check biases toward boundary values, so n === 1 is reached almost at once. This also explains why the failing-first run completed normally while the fixed run hung: before the fix the graph module did not exist, so the property test threw on import and never reached generation. It only starts hanging once the code under test works. Generates the DAG by construction instead -- `to` is drawn strictly above `from`, with the degenerate single-node case short-circuited to an empty edge list -- so no rejection sampling is involved. Switches the import to the shared fast-check setup so the seed and run count are pinned per CONTRIBUTING, and adds a bounded regression guard that samples the arbitrary directly, so a future reintroduction fails loudly instead of hanging. Verified: the file now completes in 2 seconds, 29 tests started and 29 finished, zero failures, against an indefinite hang before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2830): restore the depends_on display contract and acknowledge the workflow growth Full-suite run surfaced two things the focused harnesses could not. 1. Regression of a pinned pre-existing contract (#3785). A refactor routed the EMITTED depends_on field through the new dependency resolver, which also consults the canonical-prefix map. The original consulted the plan map only, so a short canonical prefix passed through verbatim -- '24-01' stayed '24-01' rather than becoming '24-01-auth-hardening'. The emitted field is a DISPLAY mapping, not the DAG resolution, and #3785 pins that. Reverted with a comment recording why it must not use the resolver; full resolution is still used for the wave DAG and halt propagation, which is what needs it. 2. The workflow file grew 518 bytes without an acknowledgment, from the halt-aware skip rule and the widened parse contract. Acknowledged. Note on where the acknowledgment landed: the guidance is to add a NEW fragment, but execute-phase.md is already named by an existing fragment and the linter hard-fails when two ack sources name the same path. Appending to the owning fragment, following its own established multi-PR pattern, was the only lint-clean option. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2830): backfill changeset pr number Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
88f6d9bd1b |
fix(#2644): deduplicate Cursor slash menu (#2812)
* fix(#2644): deduplicate Cursor slash menu * fix: preserve installer executable mode * chore: add changeset for PR #2812 * test(#2644): acknowledge Cursor emission changes * test(#2644): drop spent emitted drift acknowledgments * fix(#2644): remove retired Cursor command converter --------- Co-authored-by: clezcoding <clezcoding@users.noreply.github.com> Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
a987cf2731 |
chore(#2932): emit a per-invocation section manifest from the init bundle (#2987)
* chore(#2932): emit a per-invocation section manifest from init Extends the init bundle with a typed per-invocation section manifest so an invocation loads only the branch guidance it will actually take. The three flag/state-gated branches in execute-phase.md move into their own step files; the parent keeps its gsd:section markers wrapping a one-line on-demand reference, so each section's prose lives in exactly one file and the parent shrinks 93369 -> 89507 bytes. A new drift-guarded generator derives the shipped section manifest from those markers, and a new pure evaluator maps invocation facts to applicable section ids. The evaluator is a lookup over the frozen WHEN_VOCABULARY, never a parser (Greenspun's Tenth Rule, ADR-1671:69); a parity test asserts the vocabulary and the predicate map stay exhaustively in sync. Closes #2932 * fix(#2932): fail closed on prototype-chain when values An isolated adversarial review found WHEN_PREDICATES[section.when] was a bracket lookup on a plain-prototype object, so inherited Object.prototype members resolved as predicates: "constructor"/"toString"/"valueOf"/ "hasOwnProperty" returned truthy and SILENTLY INCLUDED the section, and "__proto__" threw an untyped TypeError carrying no .reason. Both violate the module's documented fail-closed contract, and the manifest is read from disk at run time so it cannot be assumed trustworthy. Builds the predicate map on a null prototype and guards the lookup with an explicit Object.hasOwn check. Adds table-driven coverage for nine Object.prototype-shaped keys asserting the TYPED reason (asserting only that it throws would still pass while broken) plus a fast-check property injecting a hostile value at an arbitrary document position. * test(#2932): retarget execute-phase step assertions at extracted step files * fix(#2932): emit typed reasons for generator lib-load and write failures * fix(#2932): restore launcher preamble in extracted steps and refresh derived fixtures * chore(#2932): backfill changeset pr number to 2987 --------- Co-authored-by: sim <sim@local> |
||
|
|
cc3ee301a7 |
fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root (#2593)
* fix(#2544): stage the CommonJS marker in GSD-owned dirs, not the config root installSharedHooksBundle wrote `{"type":"commonjs"}` over <configRoot>/package.json unconditionally — no existence check, no merge, no backup — on every install and every /gsd-update re-install. On the 11 affected runtimes that file is often user-owned; on OpenCode and Kilo it is the documented place to declare local-plugin npm dependencies, so a user's name/type/dependencies/scripts were destroyed on each run. The uninstall path already read the file and unlinked it only on an exact content match. That asymmetry was the defect: the discipline existed in the codebase, it just was not applied on the write side. Move the marker into the directories GSD creates and fills with its own .js files — hooks/ (all shared-hooks runtimes, incl. Kimi's own root) and the nativePlugin dir (plugins/ for OpenCode+Kilo, extensions/ for pi) — and stop writing the config root entirely. New src/commonjs-marker.cts owns the marker string plus one ownership predicate (absent / gsd-owned / foreign, fail-closed on an unreadable file) shared by ensureCommonJsMarker and removeCommonJsMarker, so install and uninstall cannot drift apart again. Nothing else depended on the config-root marker: package identity is baked at build time (#378/#498) and version resolution prefers gsd-core/VERSION and already tolerates a missing root package.json (#1383) — Codex has installed without one all along. A package.json in plugins/ or extensions/ is inert to plugin discovery, which globs *.{ts,js} only (see installer-migration 006). Uninstall retires the pre-fix config-root marker, so upgrading users are cleaned up on removal, and still never touches a file it did not write. * fix(#2544): point the changeset fragment at the filed PR The fragment's `pr:` field is only knowable after `gh pr create` returns. * fix(#2544): register commonjs-marker.cjs in the tsc-generated ESLint ignore set bin/lib/commonjs-marker.cjs is tsc output (src/commonjs-marker.cts is the linted source), so it belongs in the ADR-457 ignore list like its siblings. Clears the lint-tests no-var failure and the repo-invariants "linted xor ignored" migration-state test. * fix(#2544): pin the kimi CommonJS marker to hooks/, not the ~/.kimi root The UPGRADE 1 test still asserted the pre-#2544 marker location (~/.kimi/package.json). The marker now lives inside ~/.kimi/hooks — the directory GSD itself creates — matching the updated golden-install-parity and install-tree fixtures. Also asserts the root marker is NOT written. * fix(#2544): make the CommonJS marker write path non-fatal Review round 2, Major 3 + Minor 1 + the stagedHooks nit. ensureCommonJsMarker rethrew any non-EEXIST write error and neither call site caught it, so EACCES on a read-only hooks/, EROFS, or ENOSPC aborted the whole install with a raw stack trace. Every other marker interaction in the module is best-effort — removeCommonJsMarker swallows unlink failures, classifyMarker swallows read failures — and this was the write path, i.e. the one most likely to fail on a locked-down config dir. It now returns a new 'failed' outcome and both call sites warn and continue. Sibling found while sweeping for the same defect class: fs.mkdirSync sat OUTSIDE the try block, so an unwritable parent threw past the guard entirely. Creating the directory is the same environmental hazard as writing into it, so it moved inside. Also in this file: - The hooks marker is now gated on `stagedHooks && hooksOk`, not stagedHooks alone. stagedHooks is computed from the SOURCE listing before the copy loop, so it stays true when the copies land but verifyInstalled() then fails — marking a hooks/ GSD did not successfully populate claims an ownership the install did not earn. - The uninstall rmdir of the native plugin dir is gated on GSD having actually removed something from it. Hoisting it out of the adapter-exists guard (so the marker-only case could prune) had silently widened it into deleting a user-created but empty plugins/ or extensions/ dir — the same "don't touch territory GSD didn't fill" principle this issue is about, inverted. - Kimi's pre-#2544 marker at its native hook root (~/.kimi) is retired at the same call site that writes its replacement. That path is outside kimi's configDir, so installer-migration 007 structurally cannot reach it. * fix(#2544): retire the stale config-root marker via installer-migration 007 Review round 2, Major 1 — the PR's headline claim was false for existing installs. Upgraders kept BOTH markers: the new one under hooks/ and the stale {"type":"commonjs"} at the config root, so their config root stayed pinned to CommonJS and their dependency manifest stayed gone until they uninstalled. The migration is unusual in one way, and it is the part worth reviewing: the config-root marker was never recorded in gsd-file-manifest.json (writeManifest records hooks/, agents/, commands/, scripts/ and the native plugin, never a root package.json), so classifyArtifact answers 'unknown' for it and the planner's own guard downgrades a remove-managed on an 'unknown' classification to preserve-user. 007 therefore supplies the "purpose-built detector for an old GSD-owned shape" that docs/installer-migrations.md#remove-managed sanctions — exact content match, the same predicate removeCommonJsMarker has always used — and declares the resulting classification on the action. A package.json with any other content is left untouched, and there is deliberately no backup-and-remove branch: a non-matching file here is not a patched GSD artifact, it is somebody else's file. Scope is all runtimes. The `runtimes` field is OMITTED rather than `[]`: validateStringArray requires the field to be non-empty WHEN PRESENT, while the runtime filter treats an empty array as "all" — so `runtimes: []` throws at plan time and the migration never runs. The metadata test pins this. Kimi is a deliberate carve-out, named in the migration's own header: its marker lived at ~/.kimi, outside kimi's configDir, and migration relPaths are structurally confined to configDir. It is retired by the installer instead. Registration: shipped-migrations table, .gitignore for the emitted .cjs, the EXPECTED_CHECKSUMS baseline, and the ESLint ignore set. That last one is not copied from migration 006 by rote — 006 needs no entry because it imports nothing, while 007 imports node builtins, so tsc emits its __importDefault helper and the `var` in it trips no-var. This is the same lint gate that made round 1 red. * test(#2544): fault-injection and multi-runtime marker coverage Review round 2, Major 2 + Minors 4 and 5. Major 2 — CONTRIBUTING.md:514-531 is mandatory for install/uninstall flows and the suite had no fs monkeypatching at all. Every branch now covered is one whose doc comment claims it as the module's safety posture: - classifyMarker non-ENOENT lstat error -> 'foreign' (the fail-closed rule), with an ENOENT control alongside it so the test discriminates rather than just asserting one side - classifyMarker readFileSync throw -> 'foreign' (present-but-unreadable never downgrades to the permissive answer) — the fixture's bytes are exactly GSD's marker, so the test fails if the code ever answers on content it could not read - a DIRECTORY at the marker path (CONTRIBUTING:521; the symlink case was already covered with a real symlink, the directory case needs no injection at all) - the ensureCommonJsMarker TOCTOU EEXIST branch — the entire reason for flag:'wx' - the new 'failed' outcome, for both writeFileSync (EACCES/EROFS/ENOSPC) and the mkdirSync that used to sit outside the guard - removeCommonJsMarker unlink throw -> false These save and restore fs methods in `finally` rather than using chmod 0o000, which does not fault under root and would pass vacuously in root Docker and CI. Minor 4 — uninstall was driven for opencode only. pi's extensions/ and both kimi locations now have behavioral coverage, install and uninstall, each paired with a user-authored-file case proving GSD leaves it alone. Minor 5 — the stagedHooks gate had no assertion behind its stated reason. A pre-existing, GSD-untouched hooks/ directory is now driven through a runtime that declares skipSharedHooksInstall and asserted to stay marker-free, with its user content intact. Also regression-tests the uninstall rmdir gate from the previous commit: an empty plugin dir GSD removed nothing from must survive. * docs(#2544): correct stale marker prose, register the module, document the trade-off Review round 2, Minors 2, 3 and 6. Minor 2 — six files asserted the installed ROOT ships the synthetic marker. None was load-bearing (all three walk-up consumers are VERSION-first with try/catch and the marker never carried a `version`), but ADR-457:52 is the rationale for keeping a generated module, so a future reader would mis-derive the constraint from it. Each site is corrected to what is now true: the installed tree carries no package.json with a .name at all, because the only ones GSD stages are {"type":"commonjs"} markers and they now live in GSD's own directories. Two of the six needed more than a location swap. hooks/gsd-check-update-worker.js and the platform-gate test both described `require('../package.json').name` resolving to undefined; post-#2544 that require does not resolve at all, so the history is kept accurate and the present-tense claim corrected rather than just moved. And src/runtime-artifact-conversion.cts described the no-root-package.json case as Codex-only — it is now every runtime, which strengthens that comment's own argument for lazy resolution. The generated .cjs sibling needs no edit: it is gitignored build output, not a tracked file. Minor 3 — src/commonjs-marker.cts had no CONTEXT.md entry, unlike every peer module, and CONTEXT.md is the #2 co-change partner of bin/install.js. Added, including the fail-closed posture and the never-throws contract. Minor 6 — the plugins//extensions/ marker shadows the config root for all .js siblings, so an OpenCode/Kilo user's ESM plugin/*.js stays broken. That is exactly what #2544's Fix section prescribed and it is disclosed in the PR body, but the PR body is not documentation. It now lives in the OpenCode section of docs/how-to/install-on-your-runtime.md, stated as a real constraint rather than a pure improvement, with the .ts mitigation and a fallback for ESM plugins. * test(#2544): attribute the CommonJS marker in the emitted-provenance rules The differential emitted-attribution gate (#2723, landed on `next` after this branch was cut) went red on the macOS shards once this PR rebased onto it. Two distinct causes, both real gaps rather than noise: 1. `plugins/package.json` and `extensions/package.json` matched NO rule — the `native-plugin` rule covers `*.{js,cjs,mjs}` only, so the marker read as an unattributed emitted family. 2. `hooks/package.json` fell through to `hooks-built`, which attributes an emitted `hooks/<X>` to a repo source `hooks/<X>`. There is no `hooks/package.json` in the repo, so it resolved to a nonexistent path. Cause 2 is exactly the failure already documented three lines above it for Copilot's `gsd-session.json` — "a code literal, not a built script" — so the fix follows that precedent rather than inventing one: `package.json` is excluded from `hooks-built` the same way, and a dedicated `commonjs-marker` rule attributes the family across all four roots it can appear in (both hooks roots plus `plugins`/`extensions`) to the sources that actually emit it. Deliberately a RULE, not an entry in tests/emitted-drift-ack.json. An ack is for a one-off ripple and goes stale by design — the gate fails a stale ack precisely so it cannot pre-clear the next change on that path. These markers are a permanent part of the emitted tree from #2544 onward, so they need standing attribution. Verified by reproducing the CI failure locally with GSD_EMITTED_BASE: 3 provenance errors + 12 unattributed paths before, 35/35 green after. * fix(#2544): route the #2717 hooks-surface marker helpers through commonjs-marker #2717 landed a second copy of ensureCommonJsMarker/removeCommonJsMarkerIfGsdOwned in src/runtime-hooks-surface.cts for the runtimes that stage .js hooks via dedicated paths (cursor/windsurf/codex). That copy had drifted from this PR's module on the two properties that matter: - ownership probe: `fs.existsSync` FOLLOWS symlinks and reports false for a DANGLING one, so a dangling package.json symlink classified as absent and the write went straight through it. Demonstrated: against the pre-fix copy, ensureCommonJsMarker() on a hooks/ dir holding a dangling package.json symlink returns true and creates {"type":"commonjs"} OUTSIDE that directory. - create: a plain writeFileSync leaves the classify->write window open, where commonjs-marker creates with flag:'wx' (O_EXCL). Both helpers now delegate to src/commonjs-marker.cts, which is what this PR's own docstring already claimed was the single place these rules are enforced. Exported signatures are unchanged (still boolean), so bin/install.js and the #2717 tests are unaffected. The new subtest is the only coverage that fails if the duplicate is ever reintroduced — the two implementations agree on every non-adversarial input, so the existing suites pass against both. * test(#2544): pin the stagedHooks gate on zcode, not windsurf The Minor-5 coverage picked windsurf because hostBehaviors.skipSharedHooksInstall kept it out of the shared hooks bundle, so GSD staged nothing into hooks/ and the marker was correctly absent. #2717 changed that premise: cursor/windsurf/codex now stage their .js hooks via dedicated paths and get the marker beside those scripts. Measured on this tree, windsurf stages 2 .js hooks and receives a marker — so the assertion was pinning behaviour that is now wrong, not the gate it was written for. ZCode is the durable choice: per #1821 it has hooksSurface:'none' AND no plugin surface to spawn hooks, so GSD stages no .js there by either route (measured: 0 staged, no marker). The property under test is unchanged — a user-created hooks/ directory GSD never fills stays marker-free. * test(#2544): use the shared cleanup helper in the migration test Addresses the review's Major 1. The suppression's stated reason — "no helpers import available" — was not correct: tests/helpers.cjs exports cleanup, and the other test file added in this same PR imports it (tests/commonjs-marker.test.cjs). The local reimplementation dropped two protections that are live on this repo's windows-latest lane: the CWD guard (Windows cannot remove a directory that is the current working directory) and the 20 x 250ms retry budget that absorbs the deferred-scan handle Windows Defender holds on newly-written files. Local function and suppression both removed; local/no-raw-rmsync-in-tests now passes without one. * test(#2544): expect hooks/package.json for the #2717 runtimes The fresh-install contract table predates #2717, which stages cursor/windsurf/ codex .js hooks via dedicated paths and writes the CommonJS marker beside them. All three therefore now receive hooks/package.json legitimately. Measured on this tree: codex stages 3 .js hooks, cursor 6, windsurf 2 — each with the marker; cline/copilot/trae/zcode stage none and get none, so their contracts are unchanged. * fix(#2544): gate the #2717 marker writes on having staged something The three dedicated marker writers #2717 added ran unconditionally. Each one mkdirs hooks/ up front and stages its scripts conditionally on the source existing, so with an absent or empty hook source they created a directory, filled it with nothing, and marked it as GSD's anyway. That is the same write-into-someone-else's-territory this issue is about, and installSharedHooksBundle already guards the identical case with `stagedHooks`. The dedicated paths now carry the matching gate: - cursor / windsurf: `installedScripts.size > 0` - codex: a new `codexStagedHooks` flag. The enclosing guard only proves that hooks/dist EXISTS; it says nothing about whether any CODEX_HOOKS_TO_COPY entry landed. Covered for cursor and windsurf by driving each writer against a src tree whose hooks/ dir is empty. The codex leg is defensive and deliberately uncovered: its trigger state needs a package tree where hooks/dist exists but holds none of the allowlist, which is not constructible from a real checkout. * test(#2544): scope the commonjs-marker sources per root The rule declared one flat source list for every marker root, so `extensions/package.json` was attributed to runtime-hooks-surface.cts (which never writes there) and `.kimi/hooks/package.json` to install-engine.cts. That is not merely untidy. emitted-diff.cjs accepts the FIRST satisfied source, so a flat list containing bin/install.js let any change anywhere in that 13k-line file authorise marker drift for every root — the blanket escape hatch this file's own agents-verbatim comment refuses for exactly the same reason. Sources are now derived per root from ctx.rel. Note the rule ctx is `{ rel, runtime }` and carries no `root`, so keying on ctx.root would have sent every path down one branch silently. * test(#2544): state precisely what the zcode assertion pins The comment claimed the test pinned installSharedHooksBundle's `stagedHooks` gate. It does not, and neither did the windsurf version it replaced: zcode declares skipSharedHooksInstall, so the outer guard skips that helper entirely and the gate is never evaluated. The test passes on the runtime exclusion. What it does pin — the outcome a pre-existing, GSD-untouched hooks/ stays marker-free — is still worth having, and is what the review asked for. The two `staging zero hook scripts` tests are the ones that pin a real staged-nothing gate. Comment corrected rather than left implying coverage that is not there. --------- Co-authored-by: Tom Boucher <trekkie@nomorestars.com> |
||
|
|
640eaee16e |
chore(#2930): fragmentize execute-phase.md and prove per-runtime composed emission (#2972)
* feat(#2930): fragmentize plan-phase.md workflow into per-runtime-composed sections Adds src/workflow-fragments.cts (in-file <!-- gsd:section --> marker parser/composer, ADR-1671 epic #1671 Phase 3), wires it into bin/install.js's copyWithPathReplacement emission path, and pilots the marker grammar on gsd-core/workflows/plan-phase.md. Bookkeeping ripple for the new src/*.cts module: .gitignore, eslint.config.mjs, docs/INVENTORY.md + docs/INVENTORY-MANIFEST.json, and a CONTEXT.md glossary entry. Amends ADR-1671 with open questions 1 and 2 resolutions and records the closed when= applicability grammar. Adds docs/reference/workflow-fragments.md and an ARCHITECTURE.md section documenting the marker authoring model. * fix(#2930): put allow-test-rule issue ref on the same line as the marker lint-allow-test-rule-refs.cjs requires the #NNN issue reference on the same source line as `allow-test-rule:`; it was one line below and read as an unreferenced novel exemption. * docs(#2930): link the orphaned gate-predicates reference from the docs index Found while adding the workflow-fragments reference doc: docs/reference/gate-predicates.md shipped without an entry in docs/README.md, so it was unreachable from the docs index. Fixed inline rather than deferred. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2930): scope composition to workflows, add typed failure reasons Review findings from two orthogonal passes: - Scope composeWorkflow to gsd-core/workflows/ only. It previously ran on every .md the installer copied, so a future agent/command/reference doc documenting the marker syntax with an unfenced example would have been mis-parsed and silently stripped — a lossy drop the phase forbids. - Add a frozen REASON enum; failures attach a typed .reason and tests assert on it instead of matching free-form message text (CONTRIBUTING.md:635-694). - Derive the property generator's when= values from WHEN_VOCABULARY instead of duplicating them (DEFECT.GENERATIVE-FIX). - Add adversarial parser fixtures: Unicode headings, NUL, U+FFFD, BOM, fence-within-fence, tilde and indented fences, lone-CR marker line. - Document why --mvp is structurally unmarkable: its content is interleaved, not sectioned, so the whole-line grammar cannot reach it. Also fixes two stale tests on this branch, each reproduced on the unmodified tree before correction. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(#2930): retarget the pilot from plan-phase to execute-phase The full remote matrix went red on both Linux lanes. Root cause was ours: tests/phase6-capstone-conformance.test.cjs holds a PRE_PHASE6 ceiling of 94519 bytes for plan-phase.md, asserting an ADR-857 Phase-6 completion property. That is a third size gate beyond the tier caps and the differential ratchet, and it left plan-phase.md just 36 bytes of headroom rather than the 3821 computed from the XL cap. The 330 marker bytes overran it by 294. Raising the ceiling is not an option: it is a red line certifying another ADR's completion. plan-phase.md is reverted to byte-identical origin/next and the pilot moves to execute-phase.md, which has 728 bytes of headroom under its own ceiling and lands at 93147 with 3 marker pairs. The vocabulary narrows to the atoms actually used: always, flag:--wave, state:gap-closure-phase, state:has-prior-phases. Recorded in the ADR: every branch the epic names lives in plan-phase.md, which cannot be fragmentized until caps move from source to emitted bytes. That is direct evidence for the epic's premise and may reorder phases 3-4. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(#2930): backfill changeset PR number (#2972) * fix(#2930): make the emission install tests portable on Windows The windows-latest lane went red on three tests in the new install suite; Linux was green. Both causes were in the test harness, not the module. Root normalization: the opencode converter always embeds the install root forward-slashed, but the tests stripped it with the native-separator string from mkdtemp. On Windows that never matched, so the root leaked through unstripped — and because the real and stub install roots have different prefix lengths, that length difference landed directly in the byte-delta assertion (344 observed vs 275 expected). Normalize both text and root to one separator form before stripping. @-ref resolution: the helper stripped only the @~/ and @$HOME/ forms, so a Windows absolute ref (@C:/Users/...) fell through and was joined onto the root, producing ...\@C:\Users\... Strip the @ first, then detect absoluteness from the token's own shape (POSIX, drive-letter, or UNC) with no platform branching, so every OS takes the same path. Neither assertion was weakened; the exact-equality byte check is the point of the test and still holds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(#2930): document every REASON member and guard the doc/enum parity Code review found the reference doc's 'Fails closed' list covering 10 of the 11 frozen REASON members — MALFORMED_ATTRIBUTES (parseAttrs rejects malformed key="value" syntax) had no bullet, and it is distinct from UNRECOGNIZED_ATTRIBUTE, which is valid syntax with an unknown key. Two parallel surfaces sharing one constant with nothing asserting they agree is the DEFECT.GENERATIVE-FIX class, so the same commit adds the parity assertion: the test derives the enum side from the built module and the doc side by parsing the reference page, keyed on the reason IDENTIFIER rather than prose so a reworded bullet does not break it, and reports set differences in both directions by name. Proven non-vacuous: removing the MALFORMED_ATTRIBUTES bullet turns the suite red naming that exact member; restoring it returns 44/44. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9ac0dfad58 |
chore(#2929): generalize prompt-budget into the shared context-composer seam (#2958)
* test(#2929): capture prompt-budget parity corpus pre-refactor Phase 2 of epic #1671 generalizes prompt-budget's trim ladder into a shared context-composer seam. Its success condition is that review-prompt output does not change, and the only authority on "did not change" is the behavior that shipped before the refactor. Capture that behavior now, while it is still the live implementation. 47 characterization cases, every `expected` value computed by executing the current implementation rather than hand-authored — the independence CONTRIBUTING.md "Fixture provenance (#2371)" asks for. A corpus is only worth what it can detect, so this one was validated by mutation rather than assumed. Five deliberate defects were injected and each must be caught by at least one case: - the note reserve deducted unconditionally instead of only under pressure - the pressure test relaxed from `>` to `>=` - a no-op head-shrink still setting the shrunk flag - the per-plan floor dropped from the proportional share - drop order reversed Two of those exposed real holes in the first cut of this corpus, and the cases that close them exist because of it: - `>=` was caught by NOTHING. At exact cap the only trimmable fragment was a floored plan group, and the 1024-char floor absorbed the entire trim, so the mutation was byte-invisible. A3b/A3c put a droppable at exactly the cap, which makes the strict inequality observable as context kept vs omitted. - No case reached proportional-truncate at all — B6 and B7 both hard-failed the min-set pre-check first, leaving planTruncationPct at 0 across every case and the floor semantics entirely unexercised. Rebudgeted to 700 and 1100 so the min-set fits and the truncate step is actually reached; they now record 40.20% and 48.80%. The A4/A10 families sweep the pressure boundary from both sides, which is where this function has regressed before: CONTEXT.md's LEARNING.prompt-budget.boundary-gap records PR #3708 shipping two regressions that only fired when the baseline sat inside the NOTE_RESERVE_TOKENS band, because the suite paired a trivially-fitting budget with a trivially-overflowing one and never sampled between them. A4 pins that nothing is trimmed from the cap down to 81 tokens under it; A10 pins that pressure fires at +1. Together with A3b/A3c they satisfy row (d) of RULESET.TESTS.boundary-coverage.fixtures. Two facts the corpus establishes that the design notes had wrong: - "" and null sections are NOT distinguished. applyBudget uses truthy checks throughout, so an empty-string section is treated as absent: not rendered, not dropped, never recorded in `omitted`. B13b pins this while the ladder is actively trimming, where only the non-empty `research` is dropped. - Sizing matters. B12/B13 were first written at a budget where both hard-failed the min-set check and returned "", so comparing them compared two empty strings and proved nothing. Committed as its own commit, ahead of the refactor, and regenerated against the pre-refactor implementation, so the oracle is demonstrably independent of the change it will adjudicate. Refs #2929 * refactor(#2929): extract the context-composer seam from prompt-budget Epic #1671 needs prompt-budget's budget-trimming logic for a second consumer — per-runtime artifact emission — but it is walled inside the cross-AI review pipeline. Lift it into a shared seam so later phases can call it, without changing what the review pipeline emits. ADR-1671 specifies the composer as "priority + binary-search cutoff to a per-runtime budget". Read against the code it generalizes, that contract cannot express the thing being generalized. applyBudget is not a cutoff: it is a fixed five-step ladder in which each section carries its own shrink strategy, and only three of its eight sections are ever dropped. PROJECT.md is head-shrunk to N lines; plans are proportionally tail-truncated with a per-plan 1024-byte floor; instructions and roadmap are never touched at all. A cutoff composer sorts by priority and discards the tail — it has no way to say "shrink this one", "truncate that one but never below 1 KB each", or "these three are the only droppables, in this order". Building to the literal contract and routing prompt-budget through it would have silently changed review-prompt output, which is the one outcome this phase forbids. So shrink strategies are the core abstraction here, and cutoff becomes one strategy among them — the right one for per-runtime emission in Phases 3-4, not for this ladder. That is an elaboration of the ADR's intent, not a departure from it, and ADR-1671 is updated to say so. Three decisions worth stating: - The composer DECIDES; the caller RENDERS. composeWithinBudget returns a plan of surviving fragments and never a string. assemblePrompt's rendering is prompt-shaped (`## Roadmap`, `### <file>`, the note in position two), and owning it in the composer would force emission to adopt prompt-shaped rendering. The split is what lets one seam serve both consumers. - The budget unit is INJECTED via `measure(text)`. prompt-budget passes its chars/4 estimator; emission will pass a byte counter, which ADR-1671 requires for emission caps. The existing code converts a token budget to a character budget with a hardcoded `* 4`; that assumption is now an explicit `charsPerUnit` inverse, which is precisely what a byte unit needs in order to reuse this. - The entry point is `composeWithinBudget`, not `applyBudget`. That name already exists twice — src/prompt-budget.cts and src/graphify.cts, the latter being an unrelated graph-edge budget. A third would make every symbol search in this repo ambiguous, and it already misresolves: preflight and impact queries for "applyBudget" return graphify's. Behavior is unchanged and proven so: all 47 characterization cases reproduce byte-identically, and the corpus is mutation-validated rather than merely green (see the preceding commit). prompt-budget.cts drops from 436 to 343 lines and from eighteen mutable accumulators to two, both inside a helper copied verbatim. estimateTokens deliberately stays in prompt-budget and keeps its exact math: src/phase-estimation.cts re-exports it as measureTokens, and CONTEXT.md pins plan estimates and recorded actuals to that same scale, so moving or changing it would silently break the calibration loop. Refs #2929 * docs(#2929): document the context-composer seam and amend ADR-1671 Adds the INVENTORY row, the CONTEXT.md glossary entry (a PR gate for new domain modules), and a mutation-matrix entry for the new module. The ADR amendment is the substantive part. ADR-1671 specified the composer as "priority + binary-search cutoff to a per-runtime budget". Implementing Phase 2 established that a cutoff alone cannot express the function the platform generalizes, so the ADR now records shrink strategies as the core abstraction with cutoff as one strategy among them, reserved for per-runtime emission in Phases 3-4. Recording it in the ADR matters because Phases 3-6 are planned against that contract and would otherwise be planned against a mechanism that does not work. The mutation-matrix entry is not bookkeeping. Stryker scores per module against a named .cjs, so relocating the ladder out of prompt-budget.cjs would leave the extracted code unmeasured while prompt-budget's own score floated free of the logic it used to cover. context-composer gets its own entry at the same floor. Refs #2929 * test(#2929): pin the effectiveBudget rounding mode in the parity corpus An isolated correctness review found a real blind spot: mutating `Math.floor` to `Math.round` in the effectiveBudget calculation failed ZERO of the 47 corpus cases. Every (budget, safetyMarginPct) pair in the generator happened to produce a whole number, so floor, round and ceil all agreed and the rounding mode was entirely unpinned by a corpus whose whole job is to pin observable behavior. Three cases fix that by straddling the .5 boundary: A11 95 * 0.90 = 85.5 floor 85, round 86 -> the two disagree A12 97 * 0.90 = 87.3 floor and round agree; ceil (88) does not A13 93 * 0.85 = 79.05 same guard at a non-multiple-of-10 margin, so the margin arithmetic is exercised and not just the budget A11 alone catches the round mutation; all three catch ceil. Regenerated against the pre-refactor implementation (`git show 9557f8552:src/prompt-budget.cts`), so the expanded corpus keeps the independence property the original capture had. The corpus is now mutation-validated against seven injected defects, every one caught: unconditional note reserve, `>` relaxed to `>=`, no-op head-shrink setting its flag, the truncate floor ignored, drop order reversed, and both rounding-mode changes. Refs #2929 * feat(#2929): flexReserve floors and the byte-stable isolate prefix Two of issue #2929's "Done when" items were unimplemented rather than deferred, and an isolated review flagged them alongside my own audit. Both are part of ADR-1671's composer contract, so shipping the seam without them would have left Phases 3-4 building against a contract that does not exist yet. flexReserve is a per-fragment floor in measure units that every strategy must respect, which is what makes it different from the pre-existing floorChars: that one is a chars-denominated detail of proportional-truncate alone and is retained unchanged. A floored fragment is never dropped, is never head-shrunk below its floor, and raises its own proportional cap. A fragment already smaller than its floor is untouchable outright. Metadata gains `floored`, listing the ids whose floor actually prevented a trim — a guarantee no caller can observe is a guarantee no test can hold you to. isolate marks the byte-stable canonical prefix the ADR calls for: never trimmed, never dropped, but still counted, because a prefix excluded from accounting would silently under-count real context. Metadata gains `isolatePrefix` so a caller can hash or assert on the exact bytes. Declaring an isolate fragment after a non-isolate one throws: a prefix that is not at the front is not a prefix, and accepting it would make the cross-runtime stability claim meaningless. Adds tests/context-composer.test.cjs for the exact new semantics and tests/context-composer.property.test.cjs for the five invariants, including the budget-monotonicity property the issue names explicitly. Both are registered in the mutation matrix, since coverage does not migrate with relocated code. prompt-budget uses neither feature, and its output is unchanged: all 50 corpus cases still reproduce byte-identically. Refs #2929 * chore(#2929): allowlist the prompt-budget parity suite The parity corpus needs its own test file and that makes prompt-budget a three-file module against a limit of two. The lint offers consolidation or an allowlist entry with justification; the entry is the right call here. Consolidation would mean folding the characterization suite into prompt-budget.test.cjs, which is the one thing that should not happen to it. The parity suite is a distinct concern with a distinct lifecycle: it is generated rather than hand-written, it is named by scripts/mutation-matrix.cjs as its own scoring target, and its failure means something categorically different from a unit-test failure — not "this behavior is wrong" but "observable output moved". Burying it inside a general unit file would obscure exactly that signal. The allowlist is an identity ratchet, so this entry pins today's three exact filenames: adding a fourth still fails, and dropping back to two requires removing the entry. Refs #2929 * fix(#2929): register the new module with two gates it was missing The remote matrix caught three defects that no local check could, because the local runner is blocked in this repo and these suites had therefore never executed. Eight failures, identical on node22 and node24, so nothing environment-shaped. Two are the new-module ripple. A net-new src/*.cts lands in six places and this change had reached four of them — .gitignore, INVENTORY, the manifest, and the CONTEXT.md glossary — while missing the ESLint ignore list (tsc OUTPUTS must not be linted; repo-invariants asserts linted-xor-ignored) and the mutation ratchet baseline (a deliberate review-visible mirror of the matrix floors, which every COVERED module must carry). Both are now registered, the ratchet at the same floor of 66 the matrix declares. The third was a test asserting an outcome it had made impossible. It set budget:1 alongside a 400-char required fragment, so the group budget came out at -99 and the proportional-truncate step was skipped entirely — the deliberate "non-positive group budget is skipped, never clamped" rule inherited from the original ladder. Nothing was trimmed, and the test then asserted a truncation. Rebudgeted so the step actually runs, with the arithmetic written out in a comment so the next reader does not have to re-derive why 120 rather than 80. Fixing that surfaced a genuine bug in the composer. `floored` is documented as recording fragments whose flexReserve prevented a trim that would otherwise have happened, but the push sat in the else-branch of "content did not change", so it only fired when nothing was trimmed at all. A fragment truncated to a reserve-raised cap has also had a trim prevented — 40 characters' worth in the test above — and was silently absent from the field that exists to make the guarantee observable. The condition was already right; it was in the wrong branch. Now recorded on both paths: a drop prevented outright, and a truncation capped higher than the share alone would have allowed. Parity is unaffected — prompt-budget never sets flexReserve, so the branch is unreachable from every corpus path, and all 50 cases still match. Refs #2929 * chore(#2929): backfill changeset PR number (#2958) * chore(#2929): correct the corpus case count in the changeset fragment --------- Co-authored-by: sim <sim@local> |
||
|
|
05b170e448 |
chore(#2928): productionize the CONTEXT.md predicate fact-store and gate it in CI (#2938)
* feat(#2928): port CONTEXT.md predicate fact-store into the src seam Productionizes the ADR-1671 Option-E reference example as a real module: src/context-predicates.cts (parser + selector + index builder) compiled to gsd-core/bin/lib/, plus scripts/gen-context-index.cjs following the repo's --check/--write drift-guard idiom and wired into lint:generated-sync. Parser behavior is deliberately prototype-equivalent in this commit so the next commit's regression matrix binds to the real defects rather than to a missing module. Two locked design deviations from the prototype: - duplicates carry a count, not line numbers - the committed index carries no line field at all, resolving ADR-1671 open question 4: an artifact without line numbers cannot drift on a line shift, so promoting --check to a CI gate does not make it routinely red Also reconciles the one remaining duplicate predicate ID (RULESET.WORKFLOW_MARKDOWN.FENCES was declared twice; the non-MD040 wording is removed) so the gate can land fail-closed on duplicates. Refs #1671 * test(#2928): failing-first matrix for the predicate fact-store Adds the regression matrix from the phase test plan: parser declaration forms, fence and comment regions, ID/value grammar boundaries at limit-1/limit/limit+1, CRLF fidelity, duplicate detection, the drift-guard CLI, the selector query surface, and four document-shaped fast-check properties. Seven rows are RED for behavioral reasons against the ported parser: indented-bare, star-list, plus-list and numbered-list declaration forms are dropped; a tilde fence and a four-backtick fence containing a shorter fence are not skipped; and a multi-line HTML comment is parsed as live. Eleven selector rows are RED because the query surface is not wired yet. Negative fixtures come from real repo documents that predate the grammar (CONTEXT.md, CONTRIBUTING.md's fenced env-assignment examples) per the fixture-provenance rule, and the property generators are document-shaped rather than seeded from our own serializer. Refs #1671 * fix(#2928): consume the shared fence scanner, relocate the index, wire the selector Drives the failing-first matrix green. Parser: replaces the ported naive triple-backtick toggle with the shared markdown-sectionizer fence engine. scanFencedBlocks and FencedBlockRecord gain an export keyword — the only change to that module, which has 71 upstream dependents — because it already returns line-indexed spans, which is exactly what a line-reporting parser needs. It also already documents itself as the second copy of the fence state machine pending consolidation; adding a third copy here would have been the generative-fix divergence this repo warns about. A parity suite now pins predicate fence-skipping against that scanner across eight fence shapes. HTML-comment skipping stays local because the sectionizer has no comment scanner. Declaration forms widen to indented-bare, star, plus and numbered list items. Index location: docs/CONTEXT-INDEX.json, not a module under bin/lib. The remote matrix run caught the original choice — a committed .cjs there ships ~120KB of CONTEXT.md prose into a runtime module, and two content guards fired truthfully on it (a leaked .claude install path, and four hardcoded package-name literals). Neither guard was allowlisted; the artifact moved instead, mirroring docs/INVENTORY-MANIFEST.json. Nothing at runtime needs to require it — it is a drift-detection artifact, so the selector parses CONTEXT.md live and is always current. Generator: adds a frozen REASON enum and --check --json so the gate's outcome is asserted structurally instead of by matching prose, and --context-path/--index-path so tests drive the real CLI against a temp tree with no filesystem monkeypatching. Selector: gsd_run query context-predicates with --class/--prefix/--contains, structured output carrying a matched count, own-property guards, and no project-root resolution. Registering it exposed that the query dispatch table and the usage string had drifted: a new parity test found 20 routed commands missing from the usage list, all added here rather than deferred. Refs #1671 * test(#2928): lock the newly-public scanFencedBlocks contract Exporting scanFencedBlocks made it public API for the first time, so it needs its own contract test independent of the consumer that motivated the export. Memtrace's co-change analysis flagged the gap: this suite changes together with markdown-sectionizer.cts 8 times in 90 days and was absent from the diff. Covers the documented rules: 0-based indices, -1 for an unterminated fence, the same-char/>=length/no-trailing-text closer rule, a shorter fence inside a longer one staying content, CommonMark 4.5 backtick-in-info-string, and <=3-space indent tolerance. Refs #1671 * fix(#2928): address both isolated review passes Two independent reviewers (correctness axis and security axis, neither the author) found seven findings. All are fixed here with regression tests; none deferred. BLOCKER — comment-blind fence scanning caused silent, permanent predicate loss. The HTML-comment scan and the fence scan ran as two independent passes, and the fence scanner is comment-blind, so a fence delimiter inside an HTML comment with no later close read as an unterminated fence and skipped every remaining line to EOF. Worse, the drift-guard could not catch it: it diffs against a baseline produced by the same corrupted parse. The two constructs now interleave in a single pass so each suppresses the other's boundary detection while active, covered in both directions. The parity suite still binds this scanner to markdown-sectionizer's for comment-free documents, so the two cannot diverge unnoticed. BLOCKER — the selector was not consumed anywhere, leaving the phase's acceptance criterion unmet. Now wired into the pre-work predicate-citation step in contributor-standards, which is the repo's actual brief-assembly path; no code-level brief assembler exists to wire into. MAJOR — ReDoS with an unauthenticated CI-hang exploit. The predicate-id regex nested a dot-containing character class inside a dot-prefixed repeat, so N consecutive dots had exponentially many partitions: 40 dots took 565ms and growth was exponential. CI runs this parser over a pull request's own CONTEXT.md, so any contributor could have hung a shared runner with one line. Replaced with linear per-segment validation. Doubled-dot ids are now rejected; the real document contains none. MAJOR — the duplicate-id gate had only ever been proven on synthetic fixtures. A test now re-inserts the exact line this branch removed and asserts the real generator names it. MAJOR — --check together with --write silently let write win, turning the gate into a writer; a missing path value resolved to the cwd and leaked an EISDIR stack trace. Both are now clean usage errors. MINOR — the hoisted skip-list was exported as a live mutable Set; replaced with a read-only predicate. MINOR — flag-shaped selector values were unmatchable; the inline --flag=value form now provides the escape hatch. Refs #1671 * chore(#2928): backfill changeset PR number 2938 --------- Co-authored-by: sim <sim@local> |
||
|
|
3f6b063fbb |
chore(#2799): invoke_reviewers and write_reviews iterate declared lanes (#2861)
* chore(#2799): resolve reviewer lanes into executable invocation plans Phase 5b of ADR-2782. Adds the resolver and runner that let invoke_reviewers iterate declared lanes instead of hand-authored per-CLI bash. Five additive descriptor amendments, each forced by a lane that ships today: - LaneHandler gains 'opencode' — the lane rebuilds its review from assistant text parts of a --format json stream; a plain stdout copy re-breaks #1936. - modelConfigKey — antigravity's key is review.models.agy, not .antigravity, so resolving by slug silently dropped a configured model. - defaultHost/fallbackModel — Phase 4 federated every *_host with a default of empty string; the real fallback only existed in the bash. - args becomes an argv template with a closed four-placeholder vocabulary. Positional splicing produced 'codex --model M -o F exec --ephemeral', which is not a valid invocation: codex injects in the middle, twice. - kimi-code lane, with the bounded command-capability probe (needle --output-format) that tells Kimi Code from the legacy python kimi-cli. Parity gate re-pointed: the workflow-text families it scanned are the text this phase deletes, so they are replaced by descriptor-to-registry parity plus an anti-parity check that no bespoke leg returns. jq, curl and external timeout/gtimeout all drop out of the review path. Refs #2782 * chore(#2799): add review-lane query surface and widen the manifest vocabulary Adds the gsd-tools 'review-lane' route (plan/invoke/sections) the workflow loops over, projects all twelve lanes into their capability manifests, and widens capability-validator for the amendments. opencode admitted to VALID_LANE_HANDLERS under the second arm of the enum's own admission rule: one lane, justified by a documented upstream defect data cannot express (#1936 — the agent can end its turn with zero output tokens and --format default then drops the assistant text entirely). Two bugs caught by an end-to-end stub run and fixed here: - loadConfigResolved returns a provenance wrapper, not the config; using it directly resolved every key to undefined, which reads as 'nothing configured' and silently dropped every model override. - hasBinary used shell:true with an args array (Node 26 DEP0190). Replaced with a PATH scan that spawns nothing at all. Refs #2782 * chore(#2799): iterate declared lanes in invoke_reviewers and write_reviews Replaces the eleven hand-authored per-CLI bash legs with a loop over resolved lanes, and renders REVIEWS.md sections from each lane's declared reviewsSection instead of thirteen hardcoded headings. review.md drops from 1104 lines to 507 (61KB to 28.7KB). Parity gate re-pointed, as agreed: the leg-marker and section-heading families scanned exactly the text this phase deletes, so they are replaced by descriptor-to-registry parity in both directions, plus an anti-parity check that fires if a bespoke leg is ever re-added. Enum, emitting sites and the Object.keys lock moved together. The budget-trim helper is hoisted out of the Ollama leg: it was always lane-agnostic, and any lane may now declare a promptBudgetKey. Refs #2782 * feat(#2799): bind the consented egress host and re-verify it at invocation Completes ADR-2782 D5. Rule 1 was recorded in the ADR as delivered by Phase 3 but was not implemented: ConsentRecord had no host field and nothing in the tree bound one, so this phase's rule-4 comparison had no baseline. ConsentRecord gains an OPTIONAL reviewerHost. Optional is the whole design: isValidConsentRecord does not require it, so every record already on disk stays valid and no re-consent storm fires (D4 rule 5). It is deliberately excluded from disclosureSignature — the loader has no config resolver, so folding a config-derived value in would make loader and lifecycle compute different signatures for the same manifest and re-prompt forever. Install resolves hostConfigKey (falling back to the lane's declared defaultHost, which is what the invocation path uses) and records it. Invocation re-resolves and blocks on mismatch rather than silently redirecting. Absence allows: no record, or a record predating the field, means nothing to compare — denying there would break every existing local-model user on upgrade. Refs #2782 * test(#2799): cover the resolver, runner and handlers; retarget the parity suites Adds the golden invocation-plan table (one row per shipped lane, derived from the bash legs rather than the descriptor types) plus runner coverage for the probe, empty-output policy, the three handlers and the egress check. Retargets the existing suites onto the new contract: descriptor-to-registry parity, the anti-parity check, the opencode handler, and the twelfth lane. Two corrections found by running them: - modelConfigKey was required; that breaks D4 rule 2, since a reviewer manifest authored before this phase would fail validation on upgrade. It is optional, read as null when absent. - the antigravity non-zero-exit test pre-seeded the transcript, which asserted that a STALE entry leaks through — the exact bug the watermark prevents. The spawn now appends, as the real tool does. Refs #2782 * fix(#2799): restore agy --add-dir and the self-report prompt in the handler Retargeting the three legacy reviewer suites off the deleted bash surfaced two real regressions in the port, both #2176: - --add-dir was dropped. Without it agy's permission context never receives the cwd repo, so the agent anchors on its own scratch dir and reviews the plan text in isolation — the exact failure the Review Instructions forbid. It is capability-probed, because an older agy rejects the unknown flag outright and a lane that fails to start is worse than one running on the prompt anchor. - the prompt lost the clause mandating a REVIEWED-WITHOUT-REPO-ACCESS self-report, which is what makes a blind review distinguishable from a grounded one. antigravity now builds its own prompt variant. Also ports the #2073 mode-2 cli.log diagnostic, which was dropped: a pinned model that 404s exits 0 with empty stdout AND an empty transcript, so agy's own log is the only evidence that anything failed. The three suites now assert against the plan and the handler instead of matching fence text, so they no longer need allow-test-rule exemptions. Refs #2782 * docs(#2799): document the declared lanes, the new flag, and dropped prerequisites COMMANDS.md gains --kimi-code and replaces the jq-prerequisite paragraph, which is now false: no lane requires jq, curl or an external timeout. Adds the changed-egress-destination behavior, since a blocked lane is something a user can hit. CONFIGURATION.md records that the model config key is declared per lane rather than derived from the flag — antigravity's is review.models.agy — and adds review.models.kimi-code. reviewer-instances.md now routes an instance through its lane's single invocation seam instead of a copied per-adapter bash block, which is what lets a cross-cutting fix reach instances for free. That required implementing the --model/--agent/--as flags it documents; --model re-resolves through the lane's argv template rather than splicing, so the flag lands where the lane declares it rather than ahead of a subcommand. CONTEXT.md glossary gains both new modules. Refs #2782 * chore(#2799): drop the stale emitted-drift acknowledgment The only entry was #2797's, acknowledging COMMENT-ONLY GROWTH in review.md. That file now shrinks by ~32KB and every emitted hash that moved is attributable to this diff, so the ack no longer explains anything. Removing the last entry means removing the file: its presence is the alarm, and an empty one signals nothing. Verified by deleting it and re-running the attribution and provenance gates plus lint:ci — all green without it. Refs #2782 * docs(#2799): record the Phase 5b vocabulary widenings in ADR-2782 Five additive amendments, each forced by a lane that ships today, plus two corrections the phase had to make rather than work around: - D5 rule 1 was recorded as delivered by Phase 3 and was not implemented, so this phase's rule-4 comparison had no baseline. Recorded because an ADR asserting a rule was delivered is exactly what stops a later phase checking. - The DEFECT.GENERATIVE-FIX gate is re-pointed: its workflow-text families scanned the text this phase deletes. Also records that D7's 'skip the probe where no bounding mechanism exists' carve-out is obsolete — in practice it meant the Antigravity lane ran unbounded on every stock macOS host, which ships neither timeout nor gtimeout. Refs #2782 * fix(#2799): close four defects found by adversarial review Two confirmed bugs, both reproduced before fixing: - resolveLanePlan was not total. An openai-http lane with a missing or non-object invoke dereferenced inv.hostConfigKey and threw, contradicting the module's own documented contract; the spawn branch guarded correctly and the http branch did not. The CLI seam resolves every selected lane in one map, so one malformed overlay manifest would have aborted the whole review rather than dropping its own lane. Guarded, plus a per-lane try/catch at the seam so a throw can never take down siblings. - A reviewer-instance model was silently dropped for any lane declaring modelConfigKey null (cursor, qwen, coderabbit). reviewer_instances validates that cli is a known slug but never that the slug accepts a model, so a user could configure one, get a clean run, and never learn a different model reviewed their plan. Now warns explicitly. Two hardening fixes: - The slug is concatenated into artifact paths, so LANE_SLUG_RE is enforced in the resolver rather than inherited from a validator that does not run on this path — the module documents itself as the overlay-manifest trust boundary, so it should not depend on someone else having checked. - normalizeHost mangled a scheme-less value: new URL('localhost:11434') parses with an empty hostname, so it became 'localhost://11434' and was compared and requested as if real. An empty hostname now means not-a-URL. Also documents the one gap that cannot be closed here: the antigravity watermark is keyed by workspace, so two concurrent reviews of the same repo share a transcript. agy exposes no per-invocation id to filter on, so the handler now states which half of its never-stale guarantee actually holds. Refs #2782 * test(#2799): retarget the remaining eight review.md-asserting suites The remote runner found 37 failures the local sweep missed (it hit the shell's two-minute cap before reaching these). All eight extract per-CLI bash from review.md that this phase deletes; each protects a real invariant, so each is retargeted onto the plan, the runner or the handler rather than removed. Three real defects surfaced by doing so: - effort args never reached ANY lane. model-resolver.cjs exports no resolveExecution, so effortFor silently returned [] every time. Restored by calling the same bounded resolve-execution query the bash legs used — and NOT with --raw, which prints the resolved effort rather than the picked field, so claude got 'low' instead of '--effort low'. - the timeout guidance lost 'a silent empty output is a timeout kill, not a crash' — the operator note that exists because of the Codex 0xc0000142 misdiagnosis. Restored. - the opencode handler dropped EMPTY assistant text parts. The shipped jq was , and only substitutes for false/null — an empty string is truthy in jq and contributed a blank line. Found by a property test shrinking to ['', '']. The opencode property suite no longer spawns jq at all, which deletes the #2099 hang mechanism it was architected around rather than mitigating it. Refs #2782 * fix(#2799): register the two new generated modules, and untrack them The remote runner caught build output committed to git. Both new modules compile from src/*.cts into gsd-core/bin/lib/*.cjs, and every sibling generated that way is gitignored and eslint-ignored (ADR-457) - including Phase 1's own review-lane-descriptor.cjs. Mine were neither, so repo-invariants' "each bin/lib/*.cjs is linted xor ignored according to migration state" failed. Registered both in .gitignore and eslint.config.mjs alongside the Phase 1 module, and dropped them from the index. Nothing about the shipped behaviour changes; the artifacts are rebuilt by build:lib. This is the new-.cts-module registration ripple, and it is the one part of it I had not completed - the CONTEXT.md glossary and the inventory manifest were already done. Refs #2782 * chore(#2799): backfill changeset pr number to 2861 * chore(#2799): backfill changeset pr number to 2861 --------- Co-authored-by: Test <test@example.com> |
||
|
|
8b44a0da43 |
chore(#2794): single-source the reviewer invocation contract + parity assertion (#2820)
* chore(#2794): single-source the reviewer invocation contract Phase 1 of epic #2782 (ADR-2782). Introduces one core descriptor table as the declared contract for all 11 cross-AI reviewer lanes, and the DEFECT.GENERATIVE-FIX parity assertion the roster has never had. The lane contract lived in three unrelated surfaces — the roster, ~640 lines of hand-authored per-CLI bash in invoke_reviewers, and the write_reviews section headings — so cross-cutting fixes landed per-leg (#2494 and #2605 were the same empty-output defect filed twice). - src/review-lane-descriptor.cts: frozen table declaring per lane the slug, flags, probe, invoke shape, timeout floor, empty-output policy, REVIEWS.md section, evidence class, required binaries, prompt-budget key and handler. Field names track ADR-2782 D1/D2/D6/D7 verbatim so Phase 2 harvests the shape with no translation layer. It declares; it does not execute — invoke_reviewers iterates in Phase 5b. - checkReviewerLaneParity: bidirectional parity across descriptor, roster, invoke_reviewers legs and write_reviews sections. Forward-only would miss the failure it exists to catch (#2718 added a leg, #2781 was the drift). ADR-1517 instance headings are exempt per D8. - Legs carry an explicit <!-- reviewer-lane: slug --> marker; five non-lane bold labels share the bold-then-fence shape a heuristic matcher would key on. - ADR-2782 D4: an explicitly-flagged reviewer that cannot run is now an error in both the core module and the workflow prose that mirrors it. A code-only change would be unobservable — the module has no production caller; the workflow narrates the policy. Discovery paths (--all, review.default_reviewers) stay lenient. - Fixes the qwen leg, the last one discarding stderr to /dev/null. Two ADR-2782 D2 vocabulary widenings were forced by surveying the shipped legs: promptChannel 'none' (CodeRabbit is fed no prompt) and outputChannel 'file-arg' (Codex writes via -o and discards stdout, #1698). Both are additive and closed; Phase 2 owns the validator. Closes #2690 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): make the parity checker total and pin the lane slug grammar Findings from the orthogonal review passes. Spec axis — the module claimed its vocabulary tracked ADR-2782 D1/D2 "verbatim" while diverging in three undisclosed ways, which is the translation layer Phase 2 was supposed to be spared: - `transport` moves from `invoke.transport` to the LANE level, a sibling of `probe`/`invoke`, exactly as D1's manifest example places it. The nested form read better as a TS discriminated union; the union is now discriminated at the lane level instead, which costs nothing. - The header and the CONTEXT.md glossary now enumerate all FOUR widenings (adding `outputArg` and `flags[]`), not two. Standards axis — CLAUDE.md requires a fast-check property test for a parser, and `checkReviewerLaneParity` parses markdown for markers and headings. Adding one found two real defects that the hand-written matrix missed: - NOT TOTAL: a malformed descriptor entry threw on `lane.flags` iteration, contradicting the module's own "never throws" claim. Every field is now narrowed from `unknown` at the trust boundary and reported as MALFORMED_LANE / INVALID_SLUG. This matters because Phase 2 feeds this function third-party overlay data, and a parity gate that crashes is indistinguishable from one never run. - SILENT GRAMMAR MISMATCH: LEG_MARKER_RE captures only [a-z0-9_-], so a slug outside that class was unmatchable — its marker could be present and correct and the scan would still report LEG_MARKER_MISSING forever. LANE_SLUG_RE now pins the grammar and a violating slug is reported INVALID_SLUG. A loud named violation beats a silent miss. Generators are document-shaped, not writer-seeded (CONTRIBUTING #2371): seeding from the module's own matchers could only produce documents those matchers already recognize. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(#2794): register the new bin/lib module in the ESLint ignore list The remote runner caught this; lint:ci did not, because the invariant lives in the test suite rather than the lint chain: tests/repo-invariants.test.cjs "each bin/lib/*.cjs is linted xor ignored according to migration state" -> tsc-generated bin/lib modules not yet added to ESLint ignore list: review-lane-descriptor.cjs Adding a src/*.cts module ripples to six surfaces (.gitignore, the ESLint ignore list, docs/INVENTORY-MANIFEST.json, the CONTEXT.md glossary, the capability/inventory manifests, and any size baseline). The other five were covered; this was the miss. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(#2794): amend ADR-2782 D1/D2/D8 with the vocabulary Phase 1 surfaced Building the Phase 1 descriptor table against all eleven shipped legs is the first time every lane's contract was written in one place, and it surfaced four cases the ADR's original survey did not cover. Amending the design lock rather than diverging from it, so Phase 2 (#2795) implements the manifest validator against the amended vocabulary instead of rediscovering the gaps. All four are additive widenings of closed enums; no decision reverses: - D2 promptChannel gains `none` — coderabbit is fed no prompt at all, it reviews the working-tree diff. - D2 outputChannel gains `file-arg` — the ADR called a file-writing lane a shape a real CLI *could* take; codex already is one, writing via -o/--output-last-message and discarding stdout (#1698). - D2 gains `outputArg`, required iff file-arg — knowing the review lands in a file is useless without the argument naming it. - D1 `flag` becomes `flags[]` and D8's uniqueness flattens across lanes — antigravity is selected by both --antigravity and --agy, which a single-valued field cannot express. This is the same evidence path that produced the openai-http transport: the vocabulary widens on a lane that exists, under review, never on speculation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(#2794): backfill changeset pr number to 2820 --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9a76ca6783 |
fix(#1882): distinguish unterminated frontmatter from absent frontmatter (#2712)
* fix(#1882): distinguish unterminated frontmatter from absent frontmatter
extractFrontmatter returned {} both for a document with no frontmatter and for
one whose fence was opened and never closed, so a file truncated mid-write was
byte-identical to a legitimate no-metadata file. Verified live through
`gsd-tools frontmatter get`: both printed {} with exit 0 and nothing on stderr.
Per ADR-1411's "corrupt is not absent" amendment the {} return is preserved
exactly -- no caller may break -- and the cause is surfaced out-of-band as a
deduplicated, unconditional stderr diagnostic. That mechanism lands as a shared
leaf module rather than a per-site copy because three sibling findings in the
same epic need it identically; four hand-rolled copies of one behaviour is the
generative-fix-divergence defect class.
The discriminator is deliberately not "opened but never closed". A Markdown
document whose first line is a thematic break takes that exact branch, so
flagging on the missing fence alone reports corruption on good Markdown -- the
failure mode this class of check has shipped with before. The unterminated
region is instead run through extractFrontmatter's own parser (extracted as
parseYamlRegion so the probe and the real parse can never diverge) and reported
only when it yields at least one key.
Also folds an inline defect found while working: src/config-loader.cts carried
two NUL bytes in the JSDoc added by this epic's Phase 1 (
|
||
|
|
46ba02acde |
feat(#2630): phase-estimation module, smart-zone config key, and cli verbs (#2661)
* feat(#2630): add phase-estimation module, smart-zone config key, and cli verbs * fix(#2630): document smart_zone_tokens, refresh golden fixtures, fix null-proto property assertions * fix(#2630): align smart_zone_tokens write/read validation and harden estimation tests * chore(#2630): backfill changeset pr to 2661 |
||
|
|
c19d3d7bda |
chore(#2453): resolve the uniform-router-param conflict and clear the warning floor (#2489)
npx eslint . reported 0 errors, 53 warnings. The floor eroded the signal: a
genuinely new warning had to be spotted against noise, so 'lint is clean' was not
usable as a check. This takes it to zero.
Fifty of the 53 were one category in gsd-core/bin/gsd-tools.cjs — all unused
ARGUMENTS (error x30, cwd x10, raw x7, args x3), never unused variables. They are
the Command Routing Hub's uniform handler signature, function routeX({ args, cwd,
raw, error }), declared identically whether or not a handler uses all four
members. argsIgnorePattern: '^_' is structurally in conflict with that convention:
satisfying it would mean _-prefixing ~50 parameters and making the signature
non-uniform across the table. Option 1 of #2453: disable args checking for that
file only, keeping varsIgnorePattern intact so genuinely dead variables (the #2379
class) still surface — verified by injecting an unused variable, which still warns.
This is the config decision #732 explicitly deferred ('Severities stay warn (no
config change in this pass)').
The remaining three predate the issue's count of 51 and are real defects, not
suppressions:
- tests/workflow-compat.test.cjs — the step-9 lookahead used (?=\*\*10\.|\z).
\z is a Perl/Ruby end-of-input anchor with NO meaning in JavaScript; it matched
a literal 'z', so the lazy span silently stopped at the first z whenever **10.
was absent, truncating the captured step and letting the assertion pass against
a partial block. Corrected to $.
- tests/debugger-prevention.test.cjs — an unused RegExp built one line above the
one actually used. Removed.
- tests/installer-migrations.test.cjs — try/finally inside a test body, which
CONTRIBUTING prohibits ('verbose, masks test failures, not an approved
pattern'); the unused t was the symptom. Converted to t.after().
Closes #2453
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
d16a66479a |
feat(#1950): broken-windows ledger — cross-phase defect register gating ship (#2441)
* feat(#1950): broken-windows ledger — cross-phase defect register gating ship Adds a new capability (#1950) that operationalizes GSD's no-defer discipline as a tracked, enforced artifact: accumulates stubs, TODOs, skipped tests, unrun verifies, and unmet truths across phases, and /gsd-ship blocks while any entry is open. Implementation: - src/broken-windows.cts → gsd-core/bin/lib/broken-windows.cjs: typed IR + I/O entry points (parseLedger/renderLedger/appendWindow/markWaived/markFixed + cmdWindowsStatus/Append/Waive/MarkFixed). Frozen REASON enum for typed error assertions. Windows-safe atomic rename with retry on transient EPERM/EBUSY/EACCES. - gsd-tools.cjs: new subcommand (status | append | waive | fixed), wired via routeWindows + HOST_COMMAND_ROUTERS.windows. - capabilities/broken-windows/capability.json: one ship:pre gate with artifact-frontmatter-equals predicate on WINDOWS.md open_count == 0. activationKey windows.enabled (default true) + sibling windows.enforce (default true, separate so tracking can precede enforcement). - gsd-core/workflows/ship.md: capId==broken-windows branch in preflight, sibling to security — reads gsd_run windows status --raw, fails closed on open_count > 0 or unreadable ledger. - agents/gsd-executor.md: extends the existing ## Known Stubs instruction to also append to WINDOWS.md via gsd_run windows append (best-effort, never blocks execution). - agents/gsd-verifier.md: new Step 8b — record unmet truths + human-verify items in WINDOWS.md. - gsd-core/workflows/progress.md: surfaces open + waived counts. - docs/COMMANDS.md + CONTEXT.md glossary entry + docs/INVENTORY.md: document the gate, waiver mechanism, and new module. - tests/broken-windows.test.cjs: pure + CLI behavioral coverage + fast-check roundtrip property; fail-closed on malformed ledger; security boundary on path traversal in --file. Backward-compatible: a project with no .planning/WINDOWS.md reports open_count: 0 and ships cleanly. Disable enforcement per-project with gsd config-set windows.enforce false (tracking continues, gate stays open). * chore(#1950): ratchet size baselines, defer verifier integration - Workflow size baseline: ship.md 25575→27928, progress.md 31789→32632 (broken-windows preflight branch + open-windows surface). - Agent size baseline: gsd-executor.md 46644→47951 (Known Stubs → also appends to WINDOWS.md). gsd-verifier.md unchanged. - LARGE_CAP (49152) preempted the planned verifier integration (gsd-verifier.md was at 49140 pre-PR — 12 bytes of headroom, not the documented 'real headroom'). Verifier integration deferred to a follow-up PR that extracts the VERIFICATION.md template (lines 739-859) to gsd-core/references/ — a pre-existing cap-tightness defect this PR exposed but does not expand scope to fix. Verifier integration is not in the issue's acceptance criteria (executor writes is; unmet-truths recording was an enhancement, not a gate). * fix(#1950): gate default-off, rename to workflow.windows_enforce, regen goldens Test-failure-driven fixes after first gsd-test run on db8733c8f failed 44 cases (pre-existing structural tests encoded 'ship:pre has 1 gate' / 'all caps off → empty hooks'): - capability manifest: rename windows.enabled+windows.enforce (default true) → single federated key workflow.windows_enforce (default FALSE, opt-in). Matches security's workflow.security_enforce convention and makes the adr857 all-caps-off test pass without modification (the test's buildAllFalseConfig handles workflow.* out of the box). Default-OFF keeps the gate out of the registry's default ship:pre resolution so existing loop-hooks-ship-pre-e2e structural assertions (exactly 1 gate, capId 'security') stay valid; users opt in via gsd config-set workflow.windows_enforce true. - drop activationKey (security doesn't have one either; workflow.* key doubles as the activation toggle). - regenerate docs/reference/capability-matrix.md to include broken-windows (capability-matrix-sync test). - regenerate tests/fixtures/golden-install-parity/*.json (18 runtimes) — installer now emits the new capability + lib file. - update CONTEXT.md, docs/COMMANDS.md, docs/FEATURES.md, ship.md, agents/gsd-executor.md to use the new key name and /gsd:colon slash syntax (slash-command-namespace test). - restore accidentally-regressed /gsd:capture in progress.md. Tracking-only by default; enforcement is opt-in. Acceptance criterion '/gsd-ship fails while any ledger entry is open' is met when workflow.windows_enforce=true (test fixture enables it). * test(#1950): update ship:pre structural invariants for 2-gate registry - loop-hooks-ship-pre-e2e: the registry now declares 2 gates at ship:pre (security + broken-windows), regardless of activation. Activation tests above still pin security-only or empty behavior via fixtures; these structural tests pin the REGISTRY shape, which has 2 gates as of #1950. - workflow-size-baseline: ship.md 27928→27945 (workflow.windows_enforce rename added 17 bytes). * fix(#1950): review H1+H2+M1+M2+M3 — fence-injection, EACCES fail-closed, cleanup, strict line, stryker Adversarial isolated review (Step 6.3) found 2 HIGH findings that block the PR and 3 mediums. All addressed: H1 (HIGH): description containing the markdown 3-backtick fence would terminate the ledger's JSON code block early inside JSON.stringify output (JSON doesn't escape backticks), corrupting the file and bricking the next parse. Fix: use a 4-backtick fence (json ... ) which JSON.stringify cannot produce on its own, AND validate that no entry text field contains a 4-backtick run (reject at append time with new WINDOWS_INVALID_TEXT reason code). Locked by a regression test. H2 (HIGH): readLedgerOrNull swallowed ALL fs errors as 'no ledger', silently returning open_count:0 on EACCES/EPERM/EIO. The ship gate would then pass on an unreadable ledger — the precise vector the workflow doc claims is impossible. Fix: only ENOENT returns null; every other fs error propagates as WINDOWS_LEDGER_MALFORMED so the gate blocks and the operator sees a real diagnostic. Locked by a regression test that chmod 000s a ledger with open_count=1 and asserts the result is never a false-green 0. M1: writeLedgerAtomic left an orphaned .tmp file on rename failure. Wrapped renameWithRetry in try/catch with best-effort unlink. M2: validateLine silently coerced 'abc' → NaN → null, hiding type drift. Removed the line === 0 special case (was undocumented) and made the error message match the strict check. Now any non-positive- integer line value throws, including strings. M3: tests/broken-windows.test.cjs (with its fast-check property test) was not in stryker.config.mjs DEFAULT_TEST_CMD — Stryker would mutate src/broken-windows.cts but no test would catch the mutations, producing false surviving-mutant scores. Added to the list. L1 (dead throw e after error()), L7 (line boundary tests, H1/H2 regression tests, 4-backtick CLI test) also addressed. * docs(#1950): inline concurrency + busy-wait notes (review L2+L3) * fix(#1950): regen goldens against latest gsd-tools; correct --line 0 boundary test gsd-test v4 caught two issues: - goldens I regenerated earlier (commit 526682084) predated the L1 routeWindows catch-block cleanup (commit dd844d565). Regenerated via 'npm run gen:golden' against current HEAD so the install parity hash for gsd-tools.cjs matches. - 'append --line boundary' test expected --line 0 to succeed with null entry.line, but the M2 fix correctly rejects 0 (lines are 1-indexed; 0 is not a valid source line). Updated the boundary test to assert --line 0 fails alongside -1 and 'abc'. * chore(#1950): regen goldens after rebase onto next * chore(#1950): quick.md baseline 50699→50993 (correct resolution from next rebase) * chore(changeset): backfill pr:2441 in .changeset/broken-windows-ledger.md * fix(#1950): renderTable escapes backslash before pipe (CodeQL incomplete-sanitization) CodeQL flagged the markdown-table cell escaper: String(s ?? '').replace(/\|/g, '\\|') — it escapes pipe but not backslash first. A description containing '\|' would render as '\\|' which markdown parses as 'literal backslash' + 'cell separator', splitting the column. Fix: escape backslash FIRST (each \ → \\), then pipe (each | → \|). Now a description with '\|' renders as '\\\\|' (literal '\\' + escaped pipe), which markdown renders as a single '\|' inside the cell. The JSON code block (the parse source-of-truth) was already correctly escaped via JSON.stringify; only the display-only table was affected. Locked by a regression test that: 1. Verifies the JSON block reparses with the description intact. 2. Walks the rendered table row counting unescaped pipes — must be exactly 11 (the row separators for 10 cells), proving no in-cell pipe added a split. |
||
|
|
ed06b6a4b9 |
fix(#2329): write opencode slash commands to commands/ (plural), migrate legacy command/ (#2354)
* test(#2329): fail-first tests for opencode commands/ (plural) command dir Red phase, empirically probed: global/local install lands in command/ (singular) with 71 gsd-*.md files and no commands/; the manifest records 71 keys under command/ and zero under commands/; all four declaring sites report 'command'. Migration coverage is black-box (two sequential install runs against one configDir) so it holds regardless of how the fix implements cleanup. The Kilo guard passes today by design — a forward-looking no-collateral check. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): write opencode commands to commands/ (plural), migrate legacy command/ OpenCode discovers slash commands from commands/ (plural); the installer wrote them to command/ (singular), so none of the ~71 /gsd-* commands appeared in the TUI. Five sites declared the directory and all had to agree: - capabilities/opencode/capability.json: both artifactLayout destSubpath entries (global + local) and hostBehaviors.flatCommandDir - bin/install.js: the manifest prefix was a SEPARATE hardcoded 'command/' literal, so the manifest would have diverged from the descriptor even after a rename. It now derives from _hostBehaviors(runtime).flatCommandDir. - src/install-engine.cts installOpencodeFamilyArtifacts: the actual write target, which bypasses resolveRuntimeArtifactLayout via combinedFamilyInstall. This was a fifth site the issue did not list — without it the descriptor change alone would not have moved a single file. Migration: an upgrade over a pre-fix install removes only manifest-proven GSD-managed files from the legacy command/ dir and rmdirs it once empty. Unmanifested user files are preserved, never deleted. Kilo shares the opencode family install path and is explicitly unaffected — pinned by a no-collateral test. Note on the tests: the migration cases originally built their legacy fixture by running the installer and relying on it to produce command/ — i.e. they depended on the bug to set up the fixture, and became unsatisfiable the moment it was fixed (block 1 requires command/ to be absent after a fresh install). They now fabricate the legacy layout explicitly, including rewriting the manifest keys to the command/ prefix — which is load-bearing, since the migration only removes manifest-proven files and an unrewritten fixture would silently no-op and pass even against a broken migration. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): regenerate opencode install golden after rebase onto next The golden conflicted on rebase because #2322 also regenerated it. Resolved by regenerating from the merged source rather than hand-merging a generated file; the only delta is the 71 command/gsd-*.md -> commands/gsd-*.md key renames. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * test(#2329): update stale tests that pinned opencode's singular command/ dir Seven tests encoded the old contract (opencode: command/gsd-help.md exists, the descriptor's flatCommandDir, the install-integration contract, and the resolveRuntimeArtifactLayout golden). They passed in the red phase precisely because they pinned the buggy singular dir; the fix intentionally changes that contract, so these are stale-test corrections, not regressions. Kilo shares the opencode family install path and is deliberately NOT changing — it stays on command/ (singular). The shared opencode/kilo test is now split via an explicit per-runtime dir map so the two cannot be conflated, and Kilo's own layout test is untouched. tests/opencode-command-dir-plural.test.cjs independently pins Kilo unchanged end-to-end. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): changeset for opencode commands/ dir fix Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): backfill PR number 2354 into changeset Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): correct the changeset — do not assert opencode ignores command/ The changeset repeated the issue's stated mechanism ("OpenCode discovers them from commands/ ... a clean install produced no usable commands in the TUI at all"). OpenCode's source contradicts that: packages/core/src/v1/config/command.ts globs {command,commands}/**/*.md, so BOTH names resolve, and its own skill doc still calls .opencode/command/ typical. Shipping that claim as a release note would document a mechanism that does not exist. The change is still right, for the stronger reason: OpenCode's config docs list plural as the convention and singular as backwards compatibility, so GSD was shipping on the alias the vendor may withdraw. Reworded to describe it as the alignment it is, decided on OpenCode's source and docs rather than on bug reports in either repo. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * fix(#2329): baseline opencode's commands/ surface — closes a data-loss path this PR opened Not a bookkeeping gap. Moving opencode's command dir to commands/ moved the install destination to a surface the first-time baseline scan does not cover: 000-first-time-baseline's RUNTIME_SURFACES.opencode lists ['gsd-core','command', 'skills','agents'] — no 'commands'. installOpencodeFamilyCommands unconditionally unlinks every gsd-*.md under its destination before writing the fresh set (install-engine.cts:870-873), with zero manifest or migration involvement. The only thing that protects a pre-existing file is assertInstallerMigrationsUnblocked, which runs before materialization and halts when the baseline scan flags an unknown file at a KNOWN surface. Probed: a pre-existing commands/gsd-plan.md is silently destroyed (install exits 0). The identical file under the legacy, already-baselined command/ surface correctly halts the install with "installer migration blocked pending user choice". So this PR would have traded a protected surface for an unprotected one. Fixed with a NEW fix-forward migration rather than editing 000, per docs/installer-migrations.md:131-134 — an applied migration never re-runs, so editing 000 would only protect fresh installs and leave every existing machine exposed. A new id runs for both populations and drifts no shipped checksum; adding its entry to EXPECTED_CHECKSUMS is the case that test explicitly sanctions. All five pre-existing shipped checksums verified byte-identical. Kilo is excluded by the migration's runtimes filter and keeps command/. This was previously deferred as a PR-body note claiming "low impact — nothing else acts on baseline-scan misses". That claim was never probed and was wrong. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA * docs(#2329): drop the parenthetical product description from the changeset The product-name purity guard (#1777) rejects "Kilo (which still uses command/)" — fragment prose renders verbatim into CHANGELOG.md, so a product name must not carry a parenthetical. Reworded to a plain sentence; the meaning is unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SLufH5sDuqA1AiEGu45cuA --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
2cbf186420 |
chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 (#2251)
* chore(#2143): fail-loud Result + per-surface write-set contract — Phase 3 Phase 3 of epic #2143 (ADR-2143 §5/§6). The three target bugs (#2140, #2112, #2118) were already fixed tactically on next; this introduces the reusable structural contracts and rewires the primary #2140 site onto them. - src/write-set.cts (new): the parse `Result<T> = {ok,value|reason}` (§5) and the per-surface write-set (`WriteOutcome {surface, applied, requirement?}`, `WriteSet`, `writeSetComplete`) (§6). markdown-table.cts now imports + re-exports `Result` from here (single source; distinct from command-routing-hub's Result). - requirements mark-complete (src/milestone.cts): returns a PER-REQUIREMENT, per-surface write-set; `write_set_complete` is true only if every surface of every requirement applied — structurally forbidding the #2140 OR-into-one-flag masking, including across a multi-ID batch (adversarial-review regression). Pre-existing output fields unchanged (behaviour-preserving; #2140 already fixed). - deriveProgressFromRoadmap (src/phase-lifecycle.cts): removed the vestigial null-swallowing try/catch (findTableWithColumns never throws) — ADR §5 no-swallow; RoadmapProgress return contract unchanged. - commit --files (#2112) and milestone complete --dry-run (#2118) left as-is (single-surface commit / pre-mutation preview — not genuine multi-surface writes). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2244): backfill changeset PR number (#2251) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
d49ac81306 |
chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 (#2248)
* chore(#2143): markdown table model + schema registry + fail-loud pilot — Phase 1 Phase 1 of epic #2143 (ADR-2143): consolidate markdown table parsing onto a canonical seam and migrate the pilot reader. - Add src/markdown-table.cts: parseMarkdownTable (GFM tables -> typed {columns, rows} addressed by column NAME; ragged rows are typed parse errors, not silent), a single-source TABLE_SCHEMAS registry (RoadmapProgress / RequirementsTraceability / QuickTasks / Security, with variants under one id), matchTableSchema, and findTableBySchema. Result<T> is scoped to this seam (distinct from the dispatch Result). - Migrate deriveProgressFromRoadmap (src/phase-lifecycle.cts) off the position-anchored regex to name-based resolution via the seam — fixes #2137 (the 5-column milestone-grouped Progress table previously returned all-null). - Add a schema-backed `gsd-tools quick-tasks-append` subcommand and route fast.md's log_to_state through it, retiring the inline `awk NF-2` column arithmetic — fixes #2133 (addresses #2012, #2119). Cell values are escaped (| and newlines) and the STATE.md read-modify-write is atomic under readModifyWriteStateMd (lost-update race, cf. #500/#905/#1230). - Writer/reader/template parity test guards TABLE_SCHEMAS against drift (ADR-2143 §3 Generative-Fix-Divergence). Registration: .gitignore, eslint.config.mjs, docs/INVENTORY.md + INVENTORY-MANIFEST.json, CONTEXT.md glossary, docs/CLI-TOOLS.md. Behaviour-preserving for the canonical 4-column Progress table; the named bugs are driven fail-first. Extend-never-mutate (ADR-2143 §2). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(#2242): backfill changeset PR number (#2248) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): escape backslash before pipe in markdown-table cell escaping CodeQL js/incomplete-sanitization (high): escapeCell escaped | -> \| but not the backslash itself. Now escapes \ -> \\ before | -> \|, and splitTableRow unescapes both \\ -> \ and \| -> | symmetrically so cell values (incl. literal backslashes) round-trip exactly. Added backslash round-trip tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(#2242): read ROADMAP Progress table by column name — supersede #2168 ad-hoc scan Rebase reconciliation with #2168 (the tactical #2137 fix that marked itself "pending #2143"). deriveProgressFromRoadmap now resolves the Progress table via a new seam helper findTableWithColumns (first table whose header is a superset of Phase/Plans Complete/Status/Completed, any order, extra columns ignored) and reads cells by NAME — order/injection-invariant per ADR-2143 §3 — instead of the exact TABLE_SCHEMAS match. This satisfies #2168's column-invariance property test while staying seam-based and preserving its `## Progress` scoping (#2012/#1445). Ragged Progress tables now resolve to null (ADR-2143 fail-loud); updated the stale state.test.cjs assertion that predated the Phase-1 migration. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a4d8eb4a78 |
chore(#1867): eslint-ignore generated ui-consideration-probe.cjs (SHIP-01)
ci-preflight (full test:unit) caught a Phase-1 registration gap: the
tsc-generated gsd-core/bin/lib/ui-consideration-probe.cjs was not in the
eslint.config.mjs global ignore list, failing 551-eslint-bin-lib-coverage
("each bin/lib/*.cjs is linted xor ignored according to migration state",
ADR-457/#537). Generated artifacts are ignored, not linted — added it next to
the sibling probe adapters (edge-probe/probe-core/prohibition-enforcement).
Claude-Session: https://claude.ai/code/session_01BKt4hgNZwXSeJYJtYAQUSS
|
||
|
|
6e7e3111fb |
test(#2126): fix #1259 real-eslint CPU starvation — lint a non-type-aware .cjs clean target
The prohibition-enforcement real-runner tests linted src/clock.cts (a .cts) as their clean target. Under eslint.config.mjs's type-aware block for src/**/*.cts (recommendedTypeChecked + parserOptions.project: tsconfig.build.json), each eslint spawn loaded the WHOLE tsconfig.build.json program (~2s, CPU-heavy). The real-runner tests spawn eslint repeatedly; under --test-concurrency those full-program type-checks oversubscribed the bench CPU and blew the 60s subprocess bound -> fail-closed (intermittent, load-dependent — passed 24241/24241 in an earlier run, failed here). Root fix (not a retry/timeout bandaid; measured projectService = no faster since a single-file .cts lint still loads type info): add tests/_ff_lint_clean.cjs, a KNOWN-CLEAN lint-scoped .cjs companion to _ff_lint_violation.cjs, with a flat-config block enabling local/no-source-grep so the clean pass stays non-vacuous. Repoint the 6 src/clock.cts real-runner usages (5 targets + the FF-02 toothless violationFixture) at it. Each spawn is now ~0.8s non-type-aware (no whole-program load) — starvation removed. All 6 tests' semantics verified in-process (SF-01 greens; toothless/fail-closed stay unverified); full-repo `eslint .` green. Refs #2126, #1259 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
015c3a7fda |
fix(#2071): extract install-time effort resolvers so effort sync stops requiring the un-shipped bin/install.js
`gsd-tools effort sync` crashed in every installed runtime (e.g. ~/.claude/gsd-core/) with `Cannot find module '../../../bin/install.js'`: cmdEffortSync (src/commands.cts) required the package-root bin/install.js for its install-time effort resolvers, but the installer only copies the gsd-core/ subtree into a runtime home — bin/install.js is never present there. So `effort` config changes silently never reached installed agents without a full reinstall (exactly the gap #488 was meant to close). 4th instance of the recurring "runtime code under gsd-core/ requires a file outside the shipped subtree via ../../../" anti-pattern (#1223/#1920/#1383 were the prior three, all already mitigated). Fix (ADR-457 direction — extract, single source): move readGsdEffectiveEffortConfig + resolveInstallTimeEffort (with their _getGsdEffortCatalog + _readGsdConfigFile helpers) out of the hand-authored bin/install.js into a new src/install-effort-resolver.cts that compiles into the shipped gsd-core/bin/lib/install-effort-resolver.cjs. commands.cts now requires it as a sibling (`./install-effort-resolver.cjs`) — always present in the installed tree — instead of `../../../bin/install.js`. bin/install.js imports the same four symbols back from the new module (it still calls them + re-exports them), so there is one source of truth and no duplication/drift. The lazy manifest read is repointed from the package-root layout (`.., gsd-core, bin, shared`) to the bin/lib layout (`.., shared`). Scope note: this is one of four instances of the anti-pattern; the other three are already shipped/guarded. A build-time guard rejecting new cross-boundary requires whose target isn't in the installer copy manifest (to prevent instance #5) is recommended on the issue but kept out of this fix. Tests: tests/effort-sync-installed-runtime.test.cjs does a real minimal install into a temp home (the golden-parity helper) and runs the issue's exact repro (`gsd-tools effort sync --config-dir <temp>`), asserting no MODULE_NOT_FOUND for bin/install.js. Fail-first verified: against pristine next the same test throws `Cannot find module '../../../bin/install.js'` at cmdEffortSync; post-fix it syncs cleanly. New module registered in .gitignore (ADR-457), eslint ignores, docs/INVENTORY.md + INVENTORY-MANIFEST.json. bin/install.js is not shipped and the new module is under bin/lib (excluded from golden parity), so no golden fixtures change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6addeccd19 |
feat(ai-integration): API-coverage verify:pre gate (#1562)
Full API Coverage by Default — Opt Out, Never Opt In. A phase that integrates an external API/SDK/service can no longer seal without a decided coverage matrix. - src/api-coverage.cts: deterministic detector (compound verb+noun signal + <Service> API/SDK surface; stopword-guarded; strips fenced code) + matrix parse/validate/render with field-length caps. - check api-coverage.verify-pre: blocking seal-time gate; phase arg resolved as a token under .planning/phases/ only (traversal-neutralized); validates COVERAGE.md or blocks iff a strong integration signal is detected and no matrix exists; fail-closed when phases tree exists but phase unresolvable. - capabilities/ai-integration: workflow.api_coverage_gate config key (default true), plan:pre contribution, blocking verify:pre gate. Data-driven. - gsd-core/workflows/verify-work.md: generic verify:pre gate dispatch. - Tests: detector FP/FN + matrix validation + fast-check bijection; gate e2e. Code+security review findings fixed (stopword FP, scope containment, pipe/cap rejection, prompt-injection message hygiene). - Regenerated registry/matrix/loop-host-contract/goldens/baseline + docs. Closes #1562 |
||
|
|
603593d41d |
fix(#1857): test gates normalize to one-shot + bounded timeout (no watch-mode hang)
A GSD verification gate resolves a project's test command and runs it. vitest defaults to WATCH mode in an interactive TTY — exactly where a user runs `gsd-execute-phase` — so a resolved `npm test`/`pnpm test` backed by vitest never exited and the orchestrator waited indefinitely. Recovery needed the user to manually prompt "something blocking?". Fix — one shared helper + a bounded, surfacing timeout on the test-command gates: - New pure module src/normalize-test-command.cts + `gsd-tools query normalize-test-command` verb: rewrites a resolved command to a best-effort one-shot form (direct vitest → `vitest run`; jest `--watch` → `--watchAll=false`; a package-manager `test` script whose package.json runner is watch-vitest → `CI=true` prefix; handles `--dir`; already-one-shot commands unchanged — never double-flagged). Named `normalize-test-command` (not `test-*`) so the file does not match node --test's default `test-*` discovery glob. - The three gates that HUNG or silently-continued route through that ONE helper and bound execution with `timeout $(config-get workflow.test_gate_timeout)` (new config key, default 600s): the regression gate (extracted to execute-phase/steps/regression-gate.md since execute-phase.md is size-frozen — it shrank 93528→93132; ABORTS on exit 124), the post-merge gate, and the audit-fix gate (previously an UNBOUNDED `eval`). All name watch/dev mode on 124. - verify-phase's gate was ALREADY bounded (a fixed `timeout 300`, not a hang), so it only gains the normalizer (so a watch runner exits fast) + a watch-mode hint on 124, staying under its frozen 40960-byte tier cap. Security hardening (review): the normalizer only rewrites a runner named as a standalone command TOKEN (so `run-vitest.js`/`make test-vitest`/paths are never mangled), is length-capped and uses only linear-time split-based scanning (no super-linear backtracking on an adversarial `workflow.test_command`), and reads package.json only when it is a regular file (never blocks on a FIFO via `--dir`). Config key `workflow.test_gate_timeout` (seconds, default 600) registered in the schema manifest + templates/config.json + docs/CONFIGURATION.md (mirrors workflow.cross_ai_timeout). New module registered in .gitignore, eslint ignores, inventory manifest/index. All 16 golden-install-parity fixtures + workflow size baseline regenerated for the changed shipped files; bin/lib is excluded from the parity manifest. Tests: tests/normalize-test-command.test.cjs (normalizer units incl. security hardening) and tests/test-gate-watch-mode.test.cjs (the three core gates route through the shared helper + configured timeout + exit-124 watch-mode hint; verify-phase asserted as normalize-only/already-bounded). tests/execute-phase-active-flags.test.cjs repointed at the extracted step; tests/planner-language-regression.test.cjs allowlist comment updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ea4378063f | Merge branch 'next' into feat/1820-specless-predicate-rail | ||
|
|
080bacdb4b | Merge branch 'next' into codex/gsd-onboard | ||
|
|
e3262d94d3 |
feat(capabilities): add claude-orchestration capability (Workflow backend) (#1143)
Default-off, BETA, claude-only capability adopting Claude Code's Workflow tool (/effort ultracode, Agent SDK >= v0.3.149) as an optional parallel-execution backend for the GSD loop. Restores the wave parallelism + plan-checker + verifier that #853 forces inline on Claude Code, and folds gsd-ultraplan-phase under one runtime gate. - Pure fail-closed core (src/claude-orchestration.cts): detectWorkflowBackend (gate ladder: enabled -> Claude -> backend != inline -> nested+background host -> valid Agent SDK -> SDK >= floor; every miss degrades to inline) and emitWorkflowScript (waves -> parallel() barriers, plans -> gsd-executor + worktree, files_modified overlap -> separate stages, resumeFromRunId, budget). All interpolated identifiers validated script-safe; briefs JSON-quoted. - claude-orchestration command family (gsd-tools claude-orchestration detect-backend|emit-workflow) for orchestrator invocation. - Two gated loop contributions at wired points (execute:wave:post, plan:post); federated config keys (enabled/execution_backend/min_agent_sdk_version). - ADR-1143 implementation amendment; CONTEXT.md glossary entry; explanation doc. On any runtime lacking the Workflow tool, behaviour is byte-identical to today. closes #1143 |
||
|
|
7ef834cabc | feat(#1820): spec-optional predicate rail — author probe predicates into must_haves when SPEC omits them | ||
|
|
a5298c1fc0 | refactor: project onboard routing in init | ||
|
|
8de2ff9121 |
feat(#2008): generic command-exit-zero gate-predicate evaluator (#2011)
* feat(#2008): add generic command-exit-zero gate-predicate evaluator Third-party capability gates declared via check.predicate were rendered for display but never evaluated (only built-in check.query gates fired; the security capability's gate worked solely via a hard-coded ship.md branch). Add a generic, deps-injected gate-predicate evaluator (src/gate-predicate-evaluator.cts) that dispatches by predicate.kind. Built-in kind: command-exit-zero — runs a bounded sh -c command at the project root (via shell-command-projection.execTool), inherits env, exit 0 => pass, non-zero => block, timeout => block, fail-closed. Wire a 'check predicate' subcommand into check-command-router.cts and extend the three generic workflow gate-dispatch sites (execute:wave:post, execute:post, plan:post) to route check.predicate gates to the new evaluator. The two-step gate contract (command-failure => onError; block => halt) is unchanged. - src/gate-predicate-evaluator.cts: pure leaf, KIND_TABLE extensible - src/check-command-router.cts: cmdCheckPredicate + buildPredicateDeps + parsePredicateFlags - docs/adr/2008-*, docs/reference/gate-predicates.md, docs/how-to/command-exit-zero-gate.md - tests: 38 unit + integration tests (exit mapping, timeout, interpolation, property-based bijection, malformed-predicate fail-closed, real subprocess e2e) Closes #2008 * docs(#2008): backfill changeset pr number 2011 |
||
|
|
5ae4ea4c84 |
feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) (#1998)
* feat(#1105): add external-job capability (SLURM scheduler-adapter producer half) The async external-job consumer half (#1165) shipped long ago: the core loop reads .planning/async-jobs/<job>.json manifests and treats a non-terminal one as the legal external_job_waiting half-state. The PRODUCER half (#1164) was the remaining unimplemented piece of #1105. This adds the producer as a default-off capability: - capabilities/external-job/ — capability.json (execute:wave:post -> executor, plan:post -> planner contributions, external_job.* config keys, default-off) + fragments teaching runtime-budget classification and externalization. - src/external-job.cts -> gsd-core/bin/lib/external-job.cjs — pure producer module: SLURM state -> manifest-status map (no guessing), manifest build/validate (versioned stability contract), sbatch/squeue/sacct parsers, and a fail-closed manifest writer (refuses a second non-terminal job for a plan_id already in flight; refuses to clobber a malformed manifest). fs/clock seams for deterministic tests. - scripts/slurm-adapter.cjs — operator CLI (submit/poll/show) wrapping bounded sbatch/squeue/sacct subprocesses; surfaces manifest commands for confirmation and never auto-runs them (trust boundary). - tests/external-job.test.cjs — 23 behavioral + fast-check property tests. - docs/reference/long-running-operations.md + docs/how-to/async-external-jobs.md. - CONTEXT.md glossary entry for the External-job Capability. - Regenerated capability-registry.cjs; pruned the now-stale test-file-count allowlist entry (external-job is at the 2-file cap). * chore(#1105): backfill PR number in changeset * fix(#1105): sync capability artifacts + update registry shape-pin tests gsd-test caught that adding the external-job capability requires its dependent artifacts regenerated and its registry-shape drift absorbed: - sync-manifest-versions: stamp 1.7.0-rc.2 into capability.json (was 1.0.0). - gen-capability-matrix --write: regenerate docs/reference/capability-matrix.md. - gen-inventory-manifest --write: regenerate docs/INVENTORY-MANIFEST.json. - check-gap-analysis-plan-post-e2e: plan:post now has 1 contribution (external-job planner fragment) instead of 0. - execute-wave-post-gate-pipeline-e2e: execute:wave:post now has 2 contributions (mempalace + external-job) instead of 1. * fix(#1105): regenerate capability-registry after version stamp sync-manifest-versions re-stamped external-job/capability.json from 1.0.0 to 1.7.0-rc.2 after the last registry regeneration, leaving the committed capability-registry.cjs stale (CI gen-capability-registry --check failed). gsd-test masked this because its setup runs the full 'npm run build' (which regenerates the registry); CI's 'npm test' pretest only runs build:lib. |
||
|
|
e5ef323b15 |
feat(#1787): add /gsd:next smart entry workflow (#1798)
* docs: design spec for /gsd smart-entry command
Hybrid approach porting gsd-pi's smart-entry wizard to gsd-core:
deterministic classifier (gsd-tools smart-entry --json) + markdown
command/workflow with AskUserQuestion + --text fallback. Routing-first
('what now?' menu), 10 situations redesigned for gsd-core's phase loop.
* feat: add /gsd-start smart-entry command
State-aware front door adapted from gsd-pi's smart-entry wizard,
redesigned for gsd-core's markdown-first, multi-runtime architecture.
- src/smart-entry.cts: deterministic situation classifier (no-project,
paused, blocked, verify-failed, needs-first-phase, planning, executing,
verify-pending, idle-stranded, complete, unknown). Reads STATE.md,
ROADMAP.md, git, and verify signals; emits JSON the workflow consumes.
- gsd-tools.cjs: wire case + help listing.
- commands/gsd/start.md + gsd-core/workflows/gsd.md: thin markdown
dispatcher presenting an AskUserQuestion menu (with --text fallback for
non-Claude runtimes) and dispatching to existing commands. Falls back
to /gsd:progress if detection is unavailable.
- help.md: document /gsd:start (parity with bug-2954).
- tests: smart-entry.unit.test.cjs (classifier behavior across all
situations + priority + JSON shape) and gsd-workflow.structure.test.cjs
(markdown-layer invariants + every emitted command resolves to a real
slash command).
Spec: docs/superpowers/specs/2026-06-27-gsd-smart-entry-design.md
Note: command-contract (ADR-0002) requires a gsd:* prefix, so the bare
/gsd from the spec surfaces as /gsd-start.
* refactor: rename smart-entry command to /gsd:next
Rename the command from /gsd:start to /gsd:next per feedback. The
command file is now commands/gsd/next.md (name: gsd:next) and the
backing workflow is gsd-core/workflows/smart-entry.md (named for the
smart-entry classifier and gsd-tools smart-entry subcommand; does not
collide with the existing workflows/next.md, which is the progress
--next sub-workflow). help.md and the spec updated to match.
All affected tests (188) pass; lint:ci clean.
* fix: smart-entry reads real STATE.md schema (nested progress YAML + body Phase field)
Codex review found the classifier misread this repo's own STATE.md: it
looked only for scalar current_phase/total_phases frontmatter and body
fields named 'Current Phase'/'Total Phases', but real STATE.md stores
the phase as body 'Phase: N' and total_phases/percent under a nested
'progress:' YAML object. Both came back null, so active projects
(e.g. this repo at Phase 3 / verifying) wrongly classified as
needs-first-phase.
- detectSignals now reads total_phases + percent from nested progress{}
first, then scalar fm, then body; current_phase falls back to the
body 'Phase:' field (parseProsePhaseField lineage).
- Add regression tests against the real schema (nested progress YAML +
body Phase field) covering verify-pending + executing situations.
Verified against this repo: now classifies verify-pending (was
needs-first-phase). Coverage 93.25% lines / 86.99% branches.
* fix(workflow): tiered fallback when gsd-tools is broken (not just smart-entry)
Live test exposed a self-defeating fallback: when smart-entry --json
failed because gsd-tools itself was broken (missing
markdown-sectionizer.cjs), the workflow fell back to /gsd:progress —
which also depends on gsd-tools and would dead-end too.
Replace the single /gsd:progress fallback with a tiered recovery:
1. Probe gsd_run state-snapshot. If it ALSO errors, the whole tool
layer is down — read .planning/STATE.md directly with the Read tool
and synthesize a minimal situation + actions menu so /gsd:next stays
useful. Surface a rebuild hint.
2. Only if smart-entry alone is missing (older gsd-core), fall back to
/gsd:progress as before.
Matches the direct-read resilience the live agent already did by hand.
* docs: add gsd-next skill surface
* chore: trigger no-mistakes validation
* no-mistakes(review): Fix smart-entry phase ordering
* no-mistakes(review): Fix decimal smart-entry phase ordering
* no-mistakes(test): Fix smart-entry next test contracts
* no-mistakes(document): Docs synced for smart entry
* chore: add changeset fragment for #1798 (/gsd:next smart-entry workflow)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: shorten next.md description and update golden install parity fixtures
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: update /gsd-next refs to /gsd:next in docs and add Smart Entry topic alias
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: trigger no-mistakes validation
* fix: regenerate INVENTORY-MANIFEST.json for new /gsd-next files
Full CI caught that adding commands/gsd/next.md + gsd-core/workflows/smart-entry.md
left docs/INVENTORY-MANIFEST.json stale (not in the affected-test scope that
no-mistakes' test gate runs, so it surfaced in CI). Regenerated via
node scripts/gen-inventory-manifest.cjs --write; inventory-manifest-sync
test now passes.
* fix: add 'next' to core_loop cluster, update INVENTORY-MANIFEST, fix gates.md ref
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: regenerate golden install parity fixtures for /gsd:next
Full CI (shard 3/3) caught that adding commands/gsd/next.md + the
smart-entry workflow/lib made the per-runtime golden install parity
fixtures stale across all 16 runtimes. Regenerated via
UPDATE_GOLDEN=1 node --test tests/golden-install-parity.test.cjs.
All 16 fixtures + inventory-manifest-sync now pass.
* Fix smart-entry verify-failed phase scoping and empty resolve shim step
Scope detectVerifyFailed to STATE.md's current phase so leftover higher
phase directories cannot force verify-failed routing. Move the gsd_run
shim resolver into the workflow resolve step so agents define gsd_run
before the detect step runs smart-entry.
* fix: recapture golden fixtures with updated gates.md hash (/gsd:next)
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* fix: recapture all 16 golden fixtures with updated smart-entry.md hash
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* chore: regenerate fixtures + inventory manifest after rebase onto next
Rebased onto next which adopted #1837 (package-version normalization to
<VERSION> in golden-install-parity hashes). Recaptured the golden fixture
that needed it (hermes), re-sorted INVENTORY-MANIFEST.json, and regenerated
the gsd-next / ns-workflow skill descriptions to match the command surface.
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
* refactor(#1787): delegate /gsd:next in-project advancement to gated /gsd:progress --next
Reconciles the /gsd:next smart-entry front door with the existing
/gsd:progress --next engine (davesienkowski review on PR #1798). The
classifier previously recommended /gsd:execute-phase directly for the
`executing` situation, bypassing workflows/next.md Route 0
(resume-incomplete-phase invariant, #160) and Gates 1-3 — reproducing the
duplication that got the old flat /gsd-next removed (#3054), plus a
correctness hazard (executing the recorded current phase while an earlier
phase is silently incomplete).
Now planning/executing/verify-pending recommend `/gsd:progress --next`
(single gated engine); the specific command stays an explicit secondary.
Off-path states (no-project, paused, blocked, verify-failed,
idle-stranded, complete) keep direct recommendations — smart-entry's
distinct value over --next. Adds docs/adr/1787-gsd-next-smart-entry.md and
a regression test locking the delegation contract.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(#1787): avoid literal /gsd-next token in ADR (bug-3054 guard)
The repo-invariants #3054 guard bans the removed /gsd-next slash form in
docs surfaces. Refer to the removed command as `gsd-next` (prose) — the
historical reference is unchanged, just the banned token is dropped.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore: gitignore compiled host-integration-sdk + handshake-serialized .cjs
Pre-existing gap from #1683: these two src/*.cts modules compile to
gsd-core/bin/lib/*.cjs but were omitted from the per-file ignore list, so
`npm run build`/`npm test` left them as untracked build artifacts (dirty
tree + accidental-commit footgun). Adds them alongside their siblings
(host-integration.cjs, mcp-server.cjs, …). Found while finishing #1798.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test(#1787): lock per-situation action invariants for all 11 situations + ADR typo
Adversarial-review follow-ups:
- Add a test asserting every situation's action set has exactly one
recommended action, 1-4 unique-id /gsd:* actions (previously the
one-recommended/1-4 invariant was only sampled for 6 of 11 situations).
- Fix ADR typo: /gsd-progress → /gsd:progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(#1798): split oversized test chunks so a slow shard can't trip the per-chunk timeout
Root-cause of the intermittent `full test (windows-latest, 22, shard 1/3)`
failure. It was NOT a leaked handle (the runner's kill message guesses that,
but --test-force-exit already exits leaks cleanly). Diagnosis:
- Ran every shard-1/3 file WITHOUT --test-force-exit + a 45s kill-timer:
zero hangs, zero leaks — every file self-exits. So no leaked handle / hang.
- CI activity profile: output kept flowing (slowly) right up to the 600.0s
kill — a dead hang would go silent. => pure slowness.
- Per-file timing: install-minimal-hooks.test.cjs is a 4987-line / 250-case
consolidation file doing dozens of real installs — 41s even on a fast Mac
(much worse on the slow Windows I/O path), plus an install-heavy cluster.
Mechanism: MAX_FILES_PER_CHUNK=180 packed the whole ~171-file shard into ONE
`node --test` chunk, so the entire shard's wall-clock ran against a single
600s per-chunk backstop. On slow Windows runners that single chunk crossed
600s and was killed mid-run — an intermittent false-negative gate that also
hits `next` directly.
Fix: lower MAX_FILES_PER_CHUNK 180 -> 90 so each shard splits into ~2 chunks,
each with its own fresh 600s budget and a fresh node process (also relieves
per-process memory pressure). Verified locally: shard 1/3 now runs as
chunk 1/2 (90 files) + chunk 2/2 (81 files), 5323 tests, 0 fail. Also made the
timeout kill-message name slowness as a cause instead of asserting a leak, so
the next debugger isn't sent hunting a nonexistent handle leak.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Codesmith <codesmith-bot@users.noreply.github.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tom Boucher <trekkie@nomorestars.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
1388e7a362 |
feat(#1683): published Host-Integration SDK surface + smoke test — Slice 2 (#1939)
The SDK entry (src/host-integration-sdk.cts) is the single PUBLIC surface a host-plugin author imports: the negotiated schema + classification, the five adapters (declarative/imperative/model/hook/state), and the serialized handshake. Frozen so the public shape cannot be mutated. Everything else in gsd-core stays internal. tests/sdk-smoke.test.cjs imports ONLY from the SDK entry and builds a third-party host-plugin end-to-end (compose adapters + handshake + classify) — proving an external author can wire a host without gsd-core internals (#1683 AC). ESLint-ignore + inventory manifest kept in sync for the new tsc-emitted module. |
||
|
|
47ecf997d6 |
feat(#1683): serialized capability-exchange handshake — Phase 6 Slice 1 (#1937)
* feat(#1683): serialized capability-exchange handshake — Phase 6 Slice 1 The out-of-process wire form of Phase 1's negotiateHostCapabilities: SDK hosts (pi, VS Code) that cannot share object refs exchange a JSON capability set over a wire boundary (MCP-style initialize). src/handshake-serialized.cts provides buildHandshakeRequest (host side) + handleHandshakeRequest (engine side, delegates to negotiateHostCapabilities), both JSON-round-trip-safe. CONSISTENCY is the contract: a serialized request yields the same NegotiationResult as the in-process call for the same axes (asserted in the test). Pure + additive; the companion MCP server or an SDK host binds it to a real transport. * fix(#1683): ignore tsc-emitted handshake-serialized.cjs (ADR-457) + refresh inventory manifest * fix(#1683): correct inventory manifest — drop junk smart-entry.cjs, add handshake-serialized.cjs The local gen-inventory-manifest scan captured a stale untracked smart-entry.cjs (not on next, no src/*.cts) and missed the freshly-built handshake-serialized.cjs. Aligned to the canonical clean-build report: + handshake-serialized.cjs, - smart-entry.cjs. * fix(#1683): type the JSON wire round-trips (no-unsafe-return) |
||
|
|
41193a44bd |
feat(#1681): ADR-1239 Phase C-2 — gsd-mcp-server bin entry + lifecycle test [slice 3b] (#1810)
* feat(#1681): ADR-1239 Phase C-2 — gsd-mcp-server bin entry + lifecycle test [slice 3b] Phase 4 slice 3b (closes #1681). The companion MCP server bin entry so any MCP-consuming host connects via 'npx gsd-mcp-server' (or its bin on PATH) and gets GSD command (point 1) + state IO (point 5) with no bespoke plugin. - gsd-core/bin/gsd-mcp-server.cjs: #!/usr/bin/env node shim requiring ./lib/mcp-server.cjs + runServer({stdin, stdout}); non-zero exit on fatal error (justified n/no-process-exit disable). Mirrors gsd-tools.cjs. - package.json: add 'gsd-mcp-server' bin entry. - tests/gsd-mcp-server-bin.test.cjs: 3 process-lifecycle tests — initialize + tools/list round-trip + clean exit, malformed-line -> parse error + server keeps running, empty stdin -> clean exit. Synchronous spawnSync (bounded; server exits on stdin EOF, no orphan). Phase 4 trust-gate (#1806) + loader wiring (#1808) + server module (#1809) + this bin/lifecycle slice = all of #1681's deliverables. Concrete host binding -> Phase 5 (#1682). npm-integrity + eslint + security + inventory all clean. * docs(#1681)+chore(changeset): how-to for the companion MCP server + Added fragment docs/how-to/connect-gsd-mcp-server.md — Diataxis how-to guide for connecting any MCP-capable host to gsd-mcp-server: goal-oriented flow (add config → restart → verify), real-world per-host conditionals, troubleshooting, and a trimmed reference table. Explanation/reference linked out (ADR-1239, capability-trust- model) per Diataxis boundary rules rather than mixed in. .changeset/humble-seals-rest.md — type: Added (first user-reachable surface of the epic: a new bin command). The how-to doc satisfies the docs-required gate. * fix(#1681): move gsd-mcp-server shim to top-level bin/ (out of the runtime-copied tree) The shim at gsd-core/bin/gsd-mcp-server.cjs was inside the tree the installer copies into every runtime config dir, so it leaked into all 16 runtimes and broke golden-install-parity. The MCP server is a PACKAGE bin the host spawns (npx gsd-mcp-server), not a per-runtime artifact — so it belongs at top-level bin/ alongside install.js (which is also never copied into a runtime config). - gsd-core/bin/gsd-mcp-server.cjs -> bin/gsd-mcp-server.js (require path now ../gsd-core/bin/lib/mcp-server.cjs). - package.json: bin entry -> bin/gsd-mcp-server.js. - tests/gsd-mcp-server-bin.test.cjs: SHIM path updated. - eslint.config.mjs: add bin/gsd-mcp-server.js to the bin/install.js block (drops the n/no-process-exit disable — the n plugin isn't loaded for that block, so the disable referenced an undefined rule). golden-install-parity 16/16 restored; lifecycle + unit tests green; eslint 0; lint:ci all ok. |
||
|
|
2b38356275 |
feat(#1681): ADR-1239 Phase C-2 — companion MCP server module (points 1 + 5) [slice 3a] (#1809)
* feat(#1681): ADR-1239 Phase C-2 — companion MCP server module (points 1 + 5) [slice 3a] Phase 4 slice 3a. A minimal, dependency-free stdio JSON-RPC 2.0 server exposing two of the six interface points so any MCP-consuming host (Claude/Codex/OpenCode/ VS Code/Gemini/Cursor/Cline/Hermes) can drive GSD with no bespoke plugin: - point 1 (command): tool gsd_invoke_command -> createHub/dispatch. - point 5 (state IO): tools gsd_read_state / gsd_write_state -> the Phase 3 stateIO seam (filesystem default). - src/mcp-server.cts: handleMessage(request, ctx) pure JSON-RPC handler (initialize / tools/list / tools/call) + runServer({input, output}) thin line-delimited-JSON loop over injectable streams. 3 tools wired to the existing engine surfaces. NO new dependency (hand-rolled JSON-RPC; the repo ships only claude-agent-sdk + ws — an MCP SDK is a separate packaging call). - tests/gsd-mcp-server.test.cjs: 9 tests (initialize, tools/list, state read/write round-trip, command dispatch, unknown tool / missing name / unknown method / notification / parse error, injectable-stream round-trip). Bin entry / packaging / manifest-version-sync / process-lifecycle docs -> slice 3b. This slice ships the importable, tested server surface a host (or the bin shim) drives. Proactive CI gates: ADR-457 ignores + INVENTORY-MANIFEST + injection-scan audit. All clean locally (9 tests + security 15/15 + inventory + eslint 0 problems). * chore(changeset): add Changed fragment for companion MCP server module (#1681) |
||
|
|
21b81ea068 |
feat(#1681): ADR-1239 Phase C-2 — external-descriptor trust gate (configHome confinement) [slice 1] (#1806)
* feat(#1681): ADR-1239 Phase C-2 — external-descriptor trust gate (configHome confinement) [slice 1] Phase 4 slice 1. Load-time, fail-closed configHome write-confinement for installed third-party host-plugin descriptors — defense-in-depth on top of the existing opt-in/schema/consent/first-party-wins loader gates + Phase 2's install-time assertDestWithinConfigHome (#1679 AC3). - src/external-descriptor-trust.cts: isPathConfined(target, root) pure cross-platform containment primitive + assertDescriptorConfined(descriptor, configHome) — walks runtime.artifactLayout global/local destSubpaths, throws fail-closed (naming descriptor + path) on the first escape. Rejects ../escape + absolute-outside-root. Missing layout / invalid entries skipped. - tests/external-descriptor-confinement.test.cjs: 7 tests (containment primitive, benign passes, global/local/absolute escapes rejected, missing layout, invalid entries). Load-time twin of Phase 2's install-time gate — rejects malformed/escaping descriptors BEFORE consent even matters. NOT wired into loadRegistry yet (slice 2, 4 callers, medium blast radius); this ships the reusable gate + tests. Not the ADR-1577 prompt-injection breaker (separate concern, shared word trust). Proactive CI gates: ADR-457 ignores + INVENTORY-MANIFEST + injection-scan audit. All clean locally (7 tests + security + inventory + eslint 0 problems). * chore(changeset): add Changed fragment for external-descriptor trust gate (#1681) |
||
|
|
abd21b2968 |
feat(#1680): ADR-1239 Phase C-1 — hook-bus + stateIO seams [AC4] (#1805)
* feat(#1680): ADR-1239 Phase C-1 — hook-bus + stateIO seams [AC4] Phase 3 slice 4 (AC4, final #1680 slice). The last two adapter seams behind the negotiated hookBus/stateIO axes: - src/hook-bus.cts: createHookBus({bus}, {hostEmit?}) -> host/engine/none. engine = in-process pub/sub (handler errors isolated); host = host-owned, fail-closed emit until a host emitter is bound; none = silent no-op (degrade to rule-text). PORTABLE_EVENT_FLOOR = SessionStart/PreToolUse/PostToolUse/ Stop/SessionEnd (the claude dialect all hook hosts share). - src/state-io.cts: createStateIO({io}, {backend?}) -> filesystem (today's behavior — straight fs) / sandboxed-storage / session-log-append (fail-closed seams until a host backend is bound). Proactive CI gates: ADR-457 ignores + INVENTORY-MANIFEST entries for both new .cjs; injection-scan 'act as' substring audit; unused-import check. All clean locally (11 tests + security 15/15 + inventory + eslint 0 problems). Phase 3 (#1680) seam layer now complete. Concrete host binding -> Phase 5 (#1682, D15/D18). * chore(changeset): add Changed fragment for hook-bus + stateIO seams (#1680) |
||
|
|
b152f7e64c |
feat(#1680): ADR-1239 Phase C-1 — model adapter seam (passive + active) [AC3] (#1804)
* feat(#1680): ADR-1239 Phase C-1 — model adapter seam (passive + active) [AC3] Phase 3 slice 3 (AC3). Two model-layer adapters selected by the negotiated modelMode axis (host-integration.cts): - passive: formalizes today's tier routing from src/model-resolver.cts — resolveModel delegates straight to resolveModelForTier (byte-for-behavior). This is the CLI runtimes (claude/gemini/codex/opencode/cursor/...): GSD injects prompts / a per-agent model field. - active: a host-supplied sendRequest seam (VS Code vscode.lm / pi providers) — GSD calls the model through the host. Ships as a fail-closed seam (throws until a provider is bound); Phase 5 wires a concrete provider. createModelAdapter({modelMode}, {sendRequest?}) — factory gating throws on invalid mode. Proactive CI gates applied: ADR-457 eslint ignores entry + INVENTORY-MANIFEST cli_modules entry + comment-wording audited for the injection-scan substring trap. * chore(changeset): add Changed fragment for model adapter seam (#1680) |
||
|
|
da368311ea |
feat(#1680): ADR-1239 Phase C-1 — imperative embedding adapter (composes loadRegistry) [AC2] (#1803)
* feat(#1680): ADR-1239 Phase C-1 — imperative embedding adapter (composes loadRegistry) [AC2] Phase 3 slice 2 (AC2). The engine-as-library path: createImperativeAdapter composes loadRegistry({includeInstalled:true}) — first-party-wins + consent + fail-closed gates, identical trust semantics to the CLI — and binds the engine surface behind the SAME HostIntegrationInterface the declarative adapter (AC1) satisfies, plus a registry accessor for the composed capability set. - src/adapter-imperative.cts: createImperativeAdapter({runtime}, {loadOptions}) → ImperativeAdapter (kind:'imperative' + .registry + install/uninstall delegating to install-engine). Thin: delegates the loop, does not reimplement. - tests/adapter-imperative.test.cjs: kind (16 runtimes), registry composition (loadRegistry called with includeInstalled:true), loadOptions pass-through, install/uninstall delegation, fail-closed construction. - eslint.config.mjs + docs/INVENTORY-MANIFEST.json: ADR-457 ignores entry + cli_modules entry for the new emitted .cjs (the two drift gates that bit AC1, applied proactively here). Concrete host binding (OpenCode/VS Code/pi) deferred to Phase 5 (#1682). * test: remove dead readStateMd helper from bug-1760 test readStateMd was defined but never called (writeStateMd is the only state-md helper this test uses). Clears the lone no-unused-vars warning so the repo lints fully clean (0 problems). No behavior change — test still passes 2/2. * chore(changeset): add Changed fragment for imperative embedding adapter (#1680) * fix(adapter-imperative): reword comment to avoid injection-scan substring match The prompt-injection scan regex 'act\s+as\s+(?:a|an|the)' was matching the 'act as the' substring inside 'contract as the declarative adapter' (contrACT AS THE). Reword 'contract as' -> 'shape as' — no 'act' substring, identical meaning. Clears the 'lib source files are clean' + 'codebase prompt injection scan' security-gate failures. |
||
|
|
c642ed0ec5 |
feat(#1680): ADR-1239 Phase C-1 — declarative embedding adapter + minimal HostIntegrationInterface [AC1] (#1802)
* feat(#1680): ADR-1239 Phase C-1 — declarative embedding adapter + minimal HostIntegrationInterface [AC1] Phase 3 slice 1 (AC1). Names + bounds today's projection path behind the common HostIntegrationInterface that both declarative + imperative adapters will satisfy. - src/embedding-adapter.cts: minimal HostIntegrationInterface (kind + runtime + install/uninstall) + ADAPTER_KINDS. The full 6-point binding surface (command/dispatch/model/hooks/state/artifact) is DEFERRED until the imperative adapter (AC2) fixes the shape — ADR-1239 lists the wire-shape as an open question; freezing it now risks rework across Phases 3-6. - src/adapter-declarative.cts: createDeclarativeAdapter({runtime}) factory. Delegates in-process to install-engine installRuntimeArtifacts / uninstallRuntimeArtifacts (the SAME engine functions bin/install.js uses), so output is byte-identical to today's install (gated by golden-install-parity). Module-ref call style = monkeypatch-friendly for tests. Lossy by design: projects files, does not drive the loop (that's the imperative adapter, AC2). - tests/adapter-declarative-equivalence.test.cjs: kind classification (all 16 runtimes), install/uninstall delegation with exact args (the byte-identity link), fail-closed construction (missing/invalid runtime throws). Purely additive — no install.js/install-engine changes. Unblocks AC2 (imperative adapter) + AC3/AC4 (model/hook/state seams) as follow-up slices. * chore(changeset): add Changed fragment for declarative embedding adapter (#1680) * fix(lint): ignore tsc-emitted embedding-adapter/adapter-declarative .cjs (ADR-457) The new src/embedding-adapter.cts + src/adapter-declarative.cts modules' emitted gsd-core/bin/lib/*.cjs artifacts must join the ADR-457 ignores list (lint the src/*.cts source, not the emitted .cjs). Without this, the .cjs is linted under js.recommended where @typescript-eslint/no-require-imports is undefined, so the verbatim-copied eslint-disable directive surfaces as 'Definition for rule not found' — failing the lint-tests CI job. Also restores the clean line-level disable in adapter-declarative.cts (valid in the .cts source context where the rule IS defined). * fix(docs): add adapter-declarative + embedding-adapter to INVENTORY-MANIFEST cli_modules The new ADR-1239 Phase C-1 modules' built .cjs artifacts must be registered in docs/INVENTORY-MANIFEST.json's cli_modules array or the 'docs/INVENTORY-MANIFEST.json matches the filesystem' drift test fails on CI. Mirrors the existing sorted entries. |
||
|
|
744bb7aaee |
refactor(#1771): ADR-1769 Phase 1 — STATE.md Transition Module substrate + beginPhase (#1775)
* refactor(#1771): ADR-1769 Phase 1 — STATE.md Transition Module substrate + beginPhase Lands the Phase 1 substrate per ADR-1769: - src/state-transition.cts (new Module): - Field-classification table (FieldClass enum + FIELD_CLASSIFICATION rows) - STATE_MD_SECTIONS constants block - Pure transitionCore(content, intent, deps) dispatch - beginPhase intent implementation (first-time + #3127 resume paths) - src/state.cts:cmdStateBeginPhase — collapses ~190 lines to a thin dispatch onto transitionCore via readModifyWriteStateMd. The lock, no-op write guard, and #1230 post-sync delta heuristic stay in the RMW seam; the body-mutation policy moves to transitionCore. - tests/state-transition.test.cjs (24 tests): - Substrate invariants (table enum, section constants) - Characterization: 6 first-time body field updates - Characterization: 5 #3127 idempotency-guard resume behaviors - Characterization: 3 Current Position section mutations - Characterization: Current focus body text line (#1104) - Property (RULESET.TESTS.property-based-testing): beginPhase status propagation + FIELD_CLASSIFICATION own-property contract - Resume Current Position mutation (preserves Plan/Phase/Status) No external behavior change. Full state.test.cjs regression (177 tests) plus bug-3127/#3242/#905/#948 pass. Property tests surfaced two pre-existing quirks (state-document.cjs greedy \s* on whitespace-only field values; Object.prototype method leakage on FIELD_CLASSIFICATION lookups for strings like 'toString') — documented in test comments; fix-out-of-scope for Phase 1. Closes #1771 * refactor(#1771): ADR-1769 Phase 1 codex review corrections Addresses 3 blocking findings from codex gpt-5.5/high review: 1. FIELD_CLASSIFICATION shape (state-transition.cts): - Was flat FieldClass enum (collapsed source + preservation) - Now two-column {source, preservation} rows per ADR-1769 §4 - Added missing fields verified via Memtrace against buildStateFrontmatter (state.cts:1633-1653): gsd_state_version, last_updated, last_activity_desc, progress.{total_phases, completed_phases, total_plans, completed_plans, percent} - Field 2-7 preservation dispatch can now consult the table 2. Prototype-pollution hardening (state-transition.cts): - Table is now Object.freeze(Object.assign(Object.create(null), {...})) - getFieldClassification() helper uses Object.hasOwn; returns null for inherited prototype methods (toString/valueOf/__proto__) - Old code: FIELD_CLASSIFICATION['toString'] returned the function 3. STATE_MD_SECTIONS aligned to canonical template: - Verified against gsd-core/templates/state.md via Memtrace - Was: 8 entries including non-template sections (## Session, ## Decisions, ## Operator Next Steps, ## Session Log, ## Roadmap Evolution) - Now: 6 canonical top-level sections (## Project Reference, ## Current Position, ## Performance Metrics, ## Accumulated Context, ## Deferred Items, ## Session Continuity) Also: beginPhase now consults getFieldClassification() per touched field (codex finding: 'table not consulted by transitionCore'). Unknown fields raise immediately — adding a field without a table row is caught at runtime. Repo-hygiene catches from gsd-test (not node --test, which missed these): - gsd-core/bin/lib/state-transition.cjs added to eslint.config.mjs ignore list (ADR-457 tsc-generated) - docs/INVENTORY-MANIFEST.json regenerated via node scripts/gen-inventory-manifest.cjs --write Property test for Object.prototype leakage tightened to verify getFieldClassification() returns null for toString/valueOf/__proto__. Ref #1771 * fix(#1771): cast Object.create(null) to satisfy @typescript-eslint/no-unsafe-assignment ESLint CI failed on src/state-transition.cts:72:14 — Object.create(null) returns `any`, which leaked through Object.assign to the typed `FIELD_CLASSIFICATION` declaration. Adding an explicit cast to `Record<string, FieldClassification>` eliminates the unsafe-assignment while preserving the null-prototype protection codex review recommended. gsd-test: 21888/21888 PASS. * fix(#1771): add 'see #1771' to allow-test-rule exemption per ADR-456 CI lint-allow-test-rule-refs failed: 'New allow-test-rule exemption without an issue ref — add `see #NNN` per ADR-456'. Updated comment on tests/state-transition.test.cjs to reference the Phase 1 issue. * fix(#1771): remove unnecessary allow-test-rule exemption The exemption was added speculatively. The test file does not use readFileSync + .includes()/.match()/.startsWith() on source content — it calls transitionCore() with string literals and verifies results via stateExtractField() and array .includes() on the updated[] array. No exemption needed. |
||
|
|
4ced0a64cc |
feat(#1561): assumption-delta advisory checkpoint (#1767)
* feat(#1561): assumption-delta advisory checkpoint * chore(#1561): backfill changeset PR number (#1767) --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
e075a41c86 |
feat(#1754): CLI version-skew detection — warn when a global install shadows project-local GSD (#1755)
* feat(#1754): CLI version-skew detection — warn when a global install shadows project-local GSD Addresses #1754 (approved-enhancement). Detects when the running gsd-tools.cjs is outside the project root while a project-local install exists — the shadowing scenario from #1748 where a stale global canary CLI (retired @gsd-build/sdk) silently overrides project-local GSD. Implementation (Node CLI entry-point, not shell snippet — avoids bloating 93 workflow files past their size caps): - src/cli-skew-check.cts: pure function checkCliSkew({resolvedPath, projectRoot, projectLocalExists}) → string|null. Compares paths via path.relative; returns a warning when the resolved CLI is outside the project root AND a project-local install exists. Includes @gsd-build/sdk removal hint when the path matches. No I/O (pure), no gsd-sdk literal (avoids bug-2801 lint). - gsd-core/bin/gsd-tools.cjs: wired at startup via the existing findProjectRoot resolver. Non-blocking (try/catch; advisory stderr warning, never gates). - eslint.config.mjs: registers the new ADR-457 generated artifact in the ignores. - tests: 6-case suite (skew/no-skew/legacy/normalization); all green. - Golden fixtures regenerated (UPDATE_GOLDEN=1) for the new compiled artifact. - docs/how-to/update-gsd.md: Diátaxis reference note for the skew warning. Full suite: 3354 pass, 0 regressions (1 pre-existing local AGENTS.md failure). lint:ci green. Closes #1754 * chore(#1754): backfill changeset pr placeholder * chore(#1754): regenerate INVENTORY-MANIFEST for the new cli-skew-check source module --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
871621c3c8 |
feat(#1740): require-fs-op-fallback production AST rule + Windows transient-lock retry (Phase 6) (#1742)
* feat(#1740): require-fs-op-fallback production AST rule + Windows transient-lock retry (Phase 6) ADR-1703 Phase 6 of the cross-platform portability epic (#1702). Adds the second production-code portability AST rule + the ADR-mandated glob expansion to bin/install.js and scripts/build-hooks.js. - eslint-rules/require-fs-op-fallback.cjs: flags an unguarded fs.rename / fs.renameSync (the atomic-publish primitive named first in DEFECT.WINDOWS-FS-OPS.symptom) that is NOT inside a try/catch whose handler references a transient errno ('EPERM'/'EBUSY'/'EACCES' or a *RETRY_ERRNOS set) AND NOT behind a Windows platform guard. A catch that silently swallows or cleans-up-and-rethrows without an errno check does NOT satisfy the defect's 'never silently swallow' clause. copyFile/unlink are deliberately not flagged (they are the fallback primitives per the defect's own fix-forward). Scope narrowed to rename per Phase 5's precision discipline; documented on #1740. - src/shell-command-projection.cts: export retryRenameSync(from, to) — the drop-in bounded-retry helper over the existing atomicRenameWithRetry. - 27 bare fs.renameSync sites across 11 modules routed through retryRenameSync (capability-lifecycle/lock/source, installer-migrations, milestone, phase, planning-workspace, roadmap-upgrade, runtime-hooks-surface, state, workstream). Idempotent on POSIX; resilient to AV/indexer transient locks on Windows. - eslint.config.mjs: register rule at error on src/**/*.cts; new focused portability-rules block covering bin/install.js + scripts/build-hooks.js (ADR-1703 L124-126 glob expansion — both files are compliant: zero rename violations). - tests: 15-case RuleTester suite; portability-rule-disable-ban extended (PROTECTED_RULES + scans bin/install.js/build-hooks.js with shebang handling); ci-test-scope portability-lint selection rule. - CONTEXT.md DEFECT.WINDOWS-FS-OPS predicate rewritten to point at the rule; docs/contributing/cross-platform-portability-rules.md reference + how-to. Closes #1740 * chore(#1740): backfill changeset pr:1742 * fix(#1740): tighten require-fs-op-fallback precision (codex review HIGH-1/HIGH-2) Addresses two false-negative findings from the codex (gpt-5.5/high) adversarial review of PR #1742: HIGH-1 — a catch that REFERENCES a transient errno but only rethrows (no retry/fallback) was marked compliant. The DEFECT.WINDOWS-FS-OPS fix-forward requires retry, not just recognition. Fix: catchHandlerHasRetrySignal now requires a loop `continue` backedge OR a `return <call>` delegation; a bare rethrow is flagged. The misleading `/* retry logic */` valid test is replaced with a real retry loop, and the rethrow-only shape is added as invalid. HIGH-2 — the nested-try ancestor walk treated an OUTER errno-catch as protecting the rename even when an INNER catch intercepted/swallowed the error (the outer catch is unreachable). Fix: isInsideTransientErrnoTryCatch now stops at the NEAREST enclosing TryStatement WITH A CATCH HANDLER whose block contains the rename (try-finally is skipped — it doesn't catch); outer catches are no longer consulted. The unsound nested-try valid test is converted to invalid, and a try-finally-skipped valid case is added. Verified: 17 RuleTester cases pass; zero new production violations (the 27 fixed sites use retryRenameSync; the real retry loops — atomicRenameWithRetry, capability-ledger/consent, build-hooks — remain compliant via continue/errno); lint:ci green; disable-ban + vocab-drift green. --------- Co-authored-by: review-bot <review-bot@gsd> |
||
|
|
9d52043f50 |
feat(#1733): normalize-path-in-content production AST rule + fix Windows agent-skills content leak (Phase 5) (#1736)
* feat(#1733): normalize-path-in-content production AST rule (Phase 5) ADR-1703 Phase 5 — the first production-code rule. local/normalize-path-in-content (src/**/*.cts, @typescript-eslint/parser): flags a path-returning fn result (path.basename excluded — returns a separator-less filename) interpolated into an @-reference / config-dir markdown body without .replace(/\\/g,'/') normalization, per RULESET.CONTENT-PATH-NORMALIZATION / DEFECT.WINDOWS-PATH-LEAK-IN-MARKDOWN-CONTENT. Build-and-assess found the canonical defect site (computePathPrefix) already compliant and only 1 src/ hit — a false positive (path.basename in a status message) — eliminated by narrowing (exclude basename; require a real @-ref/ config-dir marker, not bare .md). 0 src/ violations: clean forward-prevention. The out-of-band disable-ban now scans src/**/*.cts too (typescript-estree) so the production rule also cannot be eslint-disabled. Registered (error) + PROTECTED_RULES; CONTEXT.md predicates + how-to doc updated. - RuleTester suite (26 cases) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(#1733): add changeset for Windows agent-skills path-leak fix Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix: harden mutation-matrix.cjs stdin read against EAGAIN on non-blocking pipe scripts/mutation-matrix.cjs read piped stdin via readFileSync(process.stdin.fd). On macOS libuv marks the stdin pipe fd non-blocking, so a synchronous read can throw EAGAIN before the writer fills the pipe — intermittently, under heavy CI shard load — aborting the script (status 2) and flaking mutation-matrix-ratchet. Replace with readStdinSync(): an fs.readSync loop that retries on EAGAIN (1ms synchronous Atomics.wait yield), stops on 0-byte/EOF, and rethrows other errors. Deterministic regression test injects EAGAIN via an fs.readSync monkeypatch. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * ci: re-run golden-install-parity on src/lib + installer changes (close drift guard) golden-install-parity hashes every installed bin/lib/*.cjs per runtime, so it must re-run whenever the built lib could change. ci-test-scope selected it for neither src/** nor installer changes, so a source-only edit (e.g. #1691's milestone.cts/roadmap.cts) recompiled bin/lib and silently drifted the golden fixtures past the scoped lane. Add golden-install-parity.test.cjs to both the 'TS runtime sources' and 'installer and package layout' selection rules, with behavioral regression tests for each. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: review-bot <review-bot@gsd> |