Commit Graph

7 Commits

Author SHA1 Message Date
Jakub Zych
fe3ed06691 chore: clear dead test and allowlist leftovers of dropped runtimes
Some checks failed
Tests / PR mergeability (push) Successful in 18s
Tests / Base branch health (push) Successful in 9s
Tests / Detect test scope (push) Successful in 16s
Tests / lint-tests (push) Failing after 1m43s
Tests / plugin-validate (push) Successful in 58s
Tests / test (ubuntu-latest, 24, shard 1/3) (push) Failing after 19s
Tests / test (ubuntu-latest, 24, shard 2/3) (push) Failing after 20s
Tests / test (ubuntu-latest, 24, shard 3/3) (push) Failing after 20s
Tests / test (ubuntu-latest, 24) (push) Failing after 18s
Tests / test (inert CI) (push) Has been skipped
Tests / QA loop walk (smell ratchet) (push) Failing after 19s
Tests / Coverage gate (merged shards) (push) Has been skipped
Tests / Publish emitted-baseline artifact (push) Has been skipped
Duplicate auto-close sweep / sweep (push) Successful in 19s
CI timeout budget report / report (push) Failing after 14s
Close Draft PRs (sweep) / Sweep open draft PRs (push) Successful in 9s
Dismiss Unauthorized PR Approvals / dismiss-unauthorized-approval (push) Successful in 9s
Tests / conformance test (macos-latest, 24) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 1/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 2/3) (push) Has been cancelled
Tests / conformance test (windows-latest, 24, shard 3/3) (push) Has been cancelled
Tests / Required tests (push) Has been cancelled
2026-10-06 20:35:12 +02:00
Jakub Zych
a9a7a328e6 refactor: hard-fork GSD -> MSD (Make Software Done)
Mechanical rename produced by scripts/msd-rename.cjs: gsd/Gsd/GSD -> msd/Msd/MSD
across contents and paths, upstream package/repo coordinates -> @golem15/msd-core
and golem15com/msd-core. Deep links into upstream history, sibling upstream
packages, the GSD-2 import feature, CHANGELOG.md and .changeset/ are kept as-is.

Hand edits on top: MSD block-letter banner and logos, LICENSE copyright line,
package/plugin identity, regenerated lockfile, install-tree fixtures, derived
registries and benchmark baseline; migration checksum baseline re-locked
(MSD keeps its own install state, so no install had applied the old sums);
sort-order and regex-escaped expectations in tests adjusted.
2026-10-06 01:47:40 +02:00
Tom Boucher
c5629bbe74 fix(#4734): degrade worktree isolation when the root has no git repository (#4843)
* test(#4734): non-git root must degrade worktree isolation (failing first)

* fix(#4734): degrade worktree isolation when the root has no git repository

* fix(#4734): review fold-ins — 3972 ladder fixture, parity fixture, docs row, message wording

* chore(#4734): backfill changeset PR number (4843)

---------

Co-authored-by: sim <sim@local>
2026-09-18 03:16:25 -04:00
Tom Boucher
c0b2a05d2f fix(#4594): one canonical dispatch-identity owner — the emitted format and the parser that reads it back (#4693)
* fix(#4594): give dispatch identity one owner for the emitted format and its parser

The isolation guards decided whether a run-scoped sentinel applied to a
dispatch by regex-scraping model-authored prose. The scrape returned values in
a different namespace from the ones the sentinel records, so the comparison
could never succeed:

  sentinel  { phase: "03", plan: "03-02-hardening" }   <- $PHASE_NUMBER / $plan_id
  prose     "Execute plan 02 of phase 03-auth."
  scraped   { phase: "03-auth.", plan: "02" }          <- greedy (\S+), both wrong

#4594 reports only the phase half. Measured against a real phase-plan-index
run, plans[].id is phase-prefixed, plan-numbered AND slugged, while the prose
carries a bare in-phase plan number — so the plan field mismatches too, and the
Claude path is dead rather than latent. A fresh sentinel was therefore
discarded on every executor dispatch and every legitimate ISOLATION=none
degrade was denied, leaving the work unrun.

hooks/lib/dispatch-identity.js is now the single owner of both halves. The two
prompt-body producers emit a canonical marker carrying the same shell values
the sentinel records, so producer and consumer agree by construction. The prose
frame stays as a fallback, bounded by the phase-token grammar ADR-2121 owns and
deliberately reporting no plan — an absent identifier means "cannot compare"
and is safe; a wrong one is a false mismatch and is not.

The prose sentence itself is byte-identical: the executor agent reads it too,
so the marker is purely additive (Hyrum's Law).

An inapplicable sentinel is now named in the guards' deny reason instead of
being dropped silently — the silence is why this survived three producers and
two consumers unnoticed. Interpolated values come from a sentinel file and from
prompt text, so both are length-bounded and stripped of control characters.

ADR-4630 locks the seam and maps the epic's three phases.

Refs #4630
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(#4594): resolve eight review findings across the dispatch-identity seam

Three orthogonal review engines ran on 43418af144 — the code-review skill's
Standards and Spec axes, and an isolated adversarial security pass — plus a
self-review of the committed diff. Every finding is fixed here; none deferred.

F1 (major, reproduced). A keyless or unknown-key-only marker — the literal
"[gsd:dispatch]" or "[gsd:dispatch run=..]" — matched the marker grammar and
returned source:'marker' with both fields null, suppressing the prose fallback
entirely. Any prompt text containing that literal silently disabled identity
narrowing, so a fresh sentinel applied to a dispatch it was never scoped to,
defeating #3045 SECURITY F2. Prompt text is attacker-influenceable. A marker
that yields neither recognized key is no longer a marker: the scan continues to
later markers, then later texts, then prose. Forward-compatible tolerance of
unknown keys is unchanged.

F2/F3 (major). The first cut duplicated sanitizeForReason,
describeSentinelDiscard and REASON_INTERPOLATION_MAX_LEN byte-for-byte across
both guards — the exact defect class this epic exists to delete, and with no
cold-load justification, since both hooks already require hooks/lib/. They now
live in hooks/lib/isolation-deny-reason.js, and buildSentinelDiscard lives in
isolation-sentinel.js beside the comparison it mirrors, returning the nested
{sentinel:{phase,plan}, dispatch:{phase,plan}} shape instead of a bespoke
four-field bag that renamed the pairs already flowing through the seam.

F4 (hard violation). The visibility test asserted on the deny reason's prose.
CONTRIBUTING.md prohibits raw text matching on hook output, which is why every
deny carries a stable reason_code. The discard is now a structured
sentinel_discarded field on each hook's stdout JSON, and the test asserts that;
the sentence stays for the operator but is no longer the contract.

F5 (hard violation). The 64-character truncation limit had no boundary
coverage. 63/64/65 are now exercised against the single consolidated helper.

F6 (minor). sanitizeForReason stripped C0/C1 controls but not U+2028/U+2029 or
the bidi overrides, so a crafted value could still reflow or reverse the
message. Both classes are stripped, with a test each.

F7 (major). The producer/template parity test was vacuous — it rendered a
marker and re-parsed its own output, and would have passed with both templates
deleted. It now reads the two workflow templates, extracts each marker line,
substitutes the measured values and asserts the owner's parser returns them.
Proven red by deleting one template's marker line before being proven green.

F8 (doc). ADR-4630 and the design notes claimed the marker is guaranteed on the
orchestrator-worktree path because that prompt is built in shell. It is not:
executor-isolation-dispatch.md:131 says plainly that those are template
placeholders, not shell variables, so {plan_id} is model-substituted there too.
A false guarantee in a design lock is worse than a stated limit. Both documents
now say the marker is model-substituted on both paths and that the prose
fallback is the real floor everywhere. The "3 workflow templates" count was
also wrong — 3 prose sites across 2 files, 2 of which carry the marker.

Refs #4630
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4594): refresh the compact-content baseline and acknowledge execute-phase.md growth

Refs #4630.

The dispatch-identity marker and its substitution note grew
gsd-core/workflows/execute-phase.md by 525 bytes (91846 -> 92371), which drifts
two real-tree guards that lint:ci does not run:

- tests/benchmark-compact-content.test.cjs asserts the committed baseline is
  "up to date"; the split for execute-phase.md moved off 25827 -> 25952 and on
  23576 -> 23701, taking its compaction reduction 8.72% -> 8.67%. Baseline
  regenerated with scripts/benchmark-compact-content.cjs --write.
- tests/emitted-attribution.test.cjs requires a growth acknowledgment trailer
  for any emitted file that grows, keyed on the bare filename. Added below.

The growth is two additions and no rewrites: the [gsd:dispatch ...] marker line
inside the Agent() prompt's <objective>, and the note telling the orchestrator
to substitute {plan_id} with the plan's id verbatim. Both are load-bearing --
the marker is what lets a guard hook match a dispatch to the sentinel the
per-plan gate wrote, and without the note the orchestrator has no instruction
telling it the value must not be paraphrased.

Emitted-Drift-Ack-Growth: execute-phase.md — adds the canonical [gsd:dispatch] identity marker and its {plan_id} substitution note, which the isolation guards compare verbatim against the run-scoped sentinel (#4594)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(#4594): set changeset fragment pr to 4693

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: sim <sim@local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 15:36:08 -04:00
Tom Boucher
bf2332e67c fix(#3582): route every hook's compiled-module require through the self-heal build seam (#3629)
* test(3582): failing-first cold-tree coverage and the seam drift lint

On a plugin-channel install the compiled gsd-core/bin/lib/*.cjs are legitimately
absent (ADR-457 build-at-publish; the npm package builds before publishing, a raw
tree materialization never does). gsd-tools.cjs calls ensureRuntimeBuild() before
requiring ./lib; no hook does, so the isolation guard's Cannot-find-module lands in
its fail-closed catch and is misreported as an unreadable dispatch-isolation
configuration, blocking every executor dispatch.

These tests fail on that: cold-tree runs of the isolation guard, statusline, cursor
guard and update worker, plus the seam's actionable build error surfacing instead of
the generic misreport.

Also adds the drift lint the acceptance criteria require, with a fixture proving it
CAN fail — a guard never shown to fail is worthless. It is red here by design: it
flags today's unfixed hooks, which is exactly the defect.

* fix(3582): route every hook's compiled-module require through the self-heal seam

RED proven at 5b174b0d: 11 failures — the cold-tree runs for the isolation guard,
cursor guard and update worker, the fail-closed-with-actionable-message assertion, and
the lint's own real-tree check.

The compiled runtime library is produced by build:lib and gitignored (ADR-457,
build-at-publish). The npm package builds before publishing; a plugin-marketplace or
git-clone install materializes the raw tree and never does, so on that channel those
modules are legitimately absent. The self-heal seam added by #2002 exists to heal exactly
this, and the CLI entrypoint already calls it — no hook did. The isolation guard's
Cannot-find-module therefore landed in its fail-closed catch and was reported as
'could not read or resolve dispatch-isolation configuration', so an ARTIFACT ABSENCE was
misdiagnosed as an unreadable project config and every executor dispatch was blocked.

All SEVEN affected files now call the seam before their first compiled require. The issue
named four; a scan found six; implementing it surfaced a seventh — the shared isolation
sentinel helper, used by BOTH guards, which requires two compiled modules itself and
would have defeated the guards' own fix on a genuinely cold tree. Same defect class, so
fixed here rather than left as a known-broken remainder.

Failure posture is deliberately split by hook kind:
- Gates (agent isolation guard, cursor subagent start) surface the seam's actionable
  build error distinctly instead of swallowing it into the generic text, and stay
  fail-closed — a genuinely unreadable project config still DENIES exactly as before.
- Cosmetic and detached hooks (statusline, update worker, update check, update banner)
  DEGRADE rather than crash: the statusline draws on every render and the worker is a
  detached process, so a build failure there must not take down the prompt.

The npm path is untouched: the seam's already-built fast path returns immediately, so
prebuilt installs pay nothing and behave bit-for-bit as before.

Adds a drift lint, wired into the CI lint chain, so the invariant is enforced rather than
remembered — without it the next hook to add a compiled require reintroduces the class
silently. It is proven able to fail: a fixture hook requiring a compiled module without
the seam is flagged, and one that uses the seam is not. Verified directly — on the
unfixed tree it named all seven offenders; with the fix it passes.

While writing the lint's comment stripper, a naive whole-text block-comment regex ate its
own fixture, because this repo's comments legitimately spell the compiled-lib glob whose
star-slash reads as a comment opener. Rewritten as a line-based scanner with a regression
test pinning that case.

* fix(3582): test the three untested seam call sites and assert typed reason codes

Two independent reviews converged on the same major gap: the fix wired the seam into
seven files but only four had cold-tree tests. The adversarial pass put it plainly —
deleting the shared isolation-sentinel helper's seam call would not have failed any test
in the diff. That file was my own addition beyond the issue's four, so it shipped
untested; that is now closed.

- Shared isolation-sentinel helper: its seam call is only reached when .planning is NOT
  directly under cwd, and every existing cold-tree fixture puts it there, so the early
  return always fired first. Now covered, and proven load-bearing by mutation: with the
  call removed the spy records zero seam invocations and the test fails.
- update-check hook and update-banner hook: cold-tree tests added asserting the DEGRADED
  VERDICT — the fallback cache filename, and silent suppression when the package name
  degrades to null — rather than merely 'did not throw'. The banner hook previously had
  no test file at all.

Standards violation fixed: two tests asserted on free-form prose via assert.match against
a JSON reason string, which CONTRIBUTING bans by name — its own BAD example is exactly
that. The ESLint rule only covers readFileSync/spawnSync text, so tooling did not catch
it. Both isolation guards now emit a machine-readable reason_code from a frozen enum,
following the repo's existing REASON convention, and the tests assert that instead. The
human-readable message is unchanged for operators; only the assertion target moved.

The duplicated degrade boilerplate across the three cosmetic hooks was deliberately NOT
extracted, and the reason is recorded at each site: both viable shapes — a
path-parameterized helper, or a ceremony-only wrapper — defeat the drift lint's per-file
literal co-occurrence check, so extracting would require the lint to special-case its own
helper. Triplication is the lesser evil while the lint stays a co-occurrence scan.

The lint's header now states what it does and does not catch (literal quoted requires
only; hooks/ scan root), so a future reader does not over-trust a guard that a
concatenated path or a require inside a non-hooks helper would evade.

* chore(3582): regenerate the committed install-tree fixtures

Adding a new shipped hook helper changed the install tree, and those fixtures are
committed-and-derived (regen:derived / gen:install-tree), so 12 'install tree — <runtime>'
tests failed on 541a1913. Regenerated rather than hand-edited.

The delta across all 15 runtime fixtures is exactly two lines — the new helper under both
its hooks/ and gsd-hooks/ install paths — and nothing else, so the regeneration pulled in
no unrelated drift.

This is the bookkeeping ripple a new file under hooks/ carries; it was not visible from
lint:ci, which passed both before and after.

* chore(3582): backfill changeset PR number (#3629)

---------

Co-authored-by: sim <sim@local>
2026-08-18 14:11:23 -04:00
Tom Boucher
58e5a5b581 fix(#3566): read the per-install .gsd-runtime marker above host-wide defaults in the isolation guards (#3589)
* test(#3566): pin per-install .gsd-runtime marker precedence in the isolation guard

Failing-first regression for #3566: resolveRuntimeIdentity must consult the
per-install marker (<install>/gsd-core/.gsd-runtime, written by every install
since #2297) above the host-wide ~/.gsd/defaults.json whose leakage #2840
exists to prevent. In-process block drives the marker through the same
_setInstallRuntimeMarkerForTests seam model-resolver.cts established.

* fix(#3566): read the per-install .gsd-runtime marker above host-wide defaults in the isolation guard

resolveRuntimeIdentity consulted ~/.gsd/defaults.json — the exact host-wide
file whose runtime leakage #2840 exists to prevent — and never the
per-install marker the installer has written for every runtime since #2297.
On a 2-runtime machine the guard confidently resolved the WRONG runtime and
silently went inert when that runtime declares no harnessIsolationFlag.
Precedence is now GSD_RUNTIME > config.json runtime > .gsd-runtime marker >
defaults.json, restoring #2840's design; the defaults rung stays last so
single-runtime and pre-#2297 installs keep #3045 BLOCKER 2 behavior.

* fix(#3566): apply the marker rung to the cursor subagent-start fallback; review fixes

Review finding (spec pass): hooks/gsd-cursor-subagent-start.js's
resolveFallbackIsolation mirrored the Claude hook's exact three-rung chain
and shared the bug — same rung inserted between config.json and the
host-wide defaults, same #2297-pattern seam, in-process regression +
negative controls.

Review finding (standards): dropped the one new raw-text assert.match on
the block reason (CONTRIBUTING test-output rule); the reason-naming
property stays pinned by the pre-existing #3045 row.

* chore(#3566): add changeset fragment

* chore(#3566): backfill changeset pr number

---------

Co-authored-by: sim <sim@local>
2026-08-17 10:48:32 -04:00
Tom Boucher
33fca50d8a test(#3333): fold the runtime & install surface fix-* cluster — Wave 1 (#3341)
* test(#3333): fold the runtime & install surface fix-* cluster — Wave 1

Folds 11 legacy tests/fix-*.test.cjs regression files into their module's
main test suite: 6 folded into existing suites (host-integration-descriptors,
effort-surface-axis, trae-imperative-reference, hermes-skills-migration,
gsd-agent-isolation-guard), 5 renamed to become the module's sole suite
(cursor-hook-workspace-roots, cursor-subagent-isolation,
lint-compiled-artifact-sync, hooks-commonjs-marker,
shared-hooks-dir-resolution). All 195 test() blocks preserved with zero
drops; lint-test-file-count.cjs and eslint remain clean. No production code
changed. Wave 1 of 7 in #3315 (H3 of epic #3053).

* test(#3333): replace try/finally with t.after() in isolation-guard tests

CONTRIBUTING.md bans try/finally inside test bodies (masks failures, not an
approved pattern). The fold in the prior commit carried 27 instances forward
verbatim from the deleted fix-3045-dispatch-isolation-resolver.test.cjs into
an otherwise-clean file. Converts each to the approved per-test t.after()
cleanup pattern — same cleanup call, registered instead of finally-wrapped.
No assertion, fixture, or test-name change; test( count unchanged at 50.

Found by the Standards review pass on Wave 1 (#3333, H3 of epic #3053).

* fix(#3333): restore raw NUL byte mangled by the fold in hermes-skills-migration.test.cjs

The prior fold commit copied fix-2284-hermes-agent-delegate-task-projection's
"collision-robust" test via a text-based Read/Write pipeline, which silently
turned a raw NUL byte (0x00) embedded in two string literals into a regular
space character. That corrupted the test's actual purpose (proving a NUL
byte survives a string-rewrite operation untouched) and produced a genuine
gsd-test failure: `24 !== 1` for `out.split(' ').length`, because splitting
on a space finds every space in the sentence instead of the single NUL byte
the test meant to isolate.

Root-caused by diffing the raw bytes (via `git cat-file blob` + `cat -v`)
between the pre-fold source and the folded target — confirmed exactly two
bytes differ. Restored via a byte-precise patch (latin1 round-trip) touching
only those two lines; test( count and every other byte unchanged.

* fix(#3333): use \x00 escape sequence instead of a raw NUL byte in test fixture

The prior commit restored a byte-exact raw NUL byte matching the original
fix-2284 source, and the production function (applyClaudeCodeBrandSwap) was
confirmed correct in a standalone repro. But the same raw byte still failed
through gsd-test's remote pipeline. Root cause is upstream of gsd-core: some
step in that transfer path does not carry a raw 0x00 byte through untouched.

A raw embedded NUL byte was never necessary here — `\x00` as a 4-character
escape sequence in the source text produces the identical runtime character
(U+0000) without ever putting a raw byte in the tracked file, sidestepping
any byte-oriented transfer step. Applied at both call sites (the fixture
string and the split() delimiter). No behavior change; test( count unchanged
at 76.

* fix(#3333): harden copyWithPathReplacement against a source file vanishing mid-copy (TOCTOU)

Surfaced by this PR's own gsd-test run: tests/install-minimal-hooks.test.cjs
and tests/opencode-command-dir-plural.test.cjs intermittently crashed with
ENOENT reading gsd-core/workflows/zzz-e5-drift-fixture.md. Root cause is
unrelated to test-file consolidation — tests/planning-prompt-drift.test.cjs
writes that fixture directly into the real, shared gsd-core/workflows/ tree
(main() hardcodes its scan root to the real repo) and deletes it in
t.after(); copyWithPathReplacement's readdirSync-then-read loop has no
protection against the listed file vanishing before it gets there, so a
concurrently-running install path can crash entirely on what is otherwise a
completely benign race.

Fixed by skipping (not crashing on) a listed entry that no longer exists by
the time the loop reaches it. Added a regression test that deterministically
reproduces the race (readdirSync snapshot still lists the file; it is
deleted immediately after) and proves both outcomes: no throw, and the
vanished entry's destination is never partially written.

Per CLAUDE.md's no-defer rule, a defect surfaced while verifying this PR is
fixed inline rather than deferred — this overrides one-concern-per-PR.

* fix(#3333): fix third NUL-byte-mangled occurrence missed by prior fix passes

The fold originally mangled three raw-NUL-byte occurrences to spaces, not
two — the earlier byte-restore and escape-sequence commits both only
targeted the fixture string and the split() delimiter, missing
out.includes('[ ]') a few lines below (should read out.includes('[\x00]')).
A remote gsd-test run kept failing on this exact assertion even after both
prior fixes, which is what surfaced the miss. Verified via a standalone
repro using the file's real (not retyped) fixture content: all six
assertions in the collision-robust test now pass. Zero raw NUL bytes remain
in the file; test( count unchanged at 76.

* chore(#3333): add changeset for the copyWithPathReplacement TOCTOU fix

Fixed-type fragment for the production defect fixed inline in this PR
(bin/install.js's copyWithPathReplacement). Exempt from docs/ requirements
per CONTRIBUTING.md (only Added/Changed/Deprecated/Removed require it).

* chore(#3333): backfill changeset PR number (pr:0 -> pr:3341)

---------

Co-authored-by: sim <sim@local>
2026-08-10 19:02:34 -04:00