Files
msd-core/docs/explanation/live-dom-uat-capability.md
Tom Boucher 14679b866b enhance(#2856): add default-off live-DOM UAT capability (#3716)
* test(#2856): add failing-first suite for the live-dom-uat capability

Binds the approved triage shape before any of it exists:

- containment — the execute:wave:post hook must not render unless
  workflow.live_dom_uat is true AND the capability resolves active
  (fail-closed on a missing state entry, and on a non-boolean value)
- criterion 4 — agents/gsd-executor.md carries no browser MCP family;
  asserted as an absence, which is the only way it is observable
- Hyrum guard — the pre-existing mcp__playwright__* branch must stay
  outside the key-gated block, or upgrading silently removes working
  automated UI verification for every current Playwright-MCP user
- parity — the browser glob list now lives in two surfaces (agent
  frontmatter + workflow detection block); the assertion fails if
  either gains or loses a family without the other

Red by construction: the capability, agent and workflow block do not
exist yet. Verified on the remote runner.

Refs #2856

* enhance(#2856): add default-off live-DOM UAT capability

A phase whose acceptance criteria needed a live DOM could not be
finished by the agent that executed it: gsd-executor carries no browser
tools, so it correctly returned checkpoint:human-action even though the
work was not human-only, just tool-less. Every such phase degraded to
"executed, then finished by hand in the orchestrator", and autonomous:
false could not distinguish "a human must judge this" from "the executor
lacks the tool".

Implements the shape approved at triage, not the one reported. The
executor's tools: line is NOT widened, in any configuration: for a
first-party agent the static list is the only control that exists
(ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one
default-off capability owns the key, the agent, and the step:

- capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat
  (boolean, default false), one additive step at execute:wave:post
  (onError: skip, gates: []), so it can never halt a wave
- agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP
  globs, in its own tools: line, with no Bash
- verify-work automated_ui_verification — a gsd:live-dom-families block
  naming both new families AND the key; presence alone never activates

Two independent fail-closed gates: isCapabilityActive renders a hook
only on state.active === true, plus the step's own `when`.

The pre-existing mcp__playwright__* branch keeps the gating it already
had and stays outside the new block. Pulling it behind a default-off key
would have silently removed working automated UI verification from every
current Playwright-MCP user on upgrade.

Also closes a host gap this surfaced: execute:wave:post dispatched only
contribution + gate, so ANY registered step was declared and silently
never run — exactly the single-kind hand-roll loop-hook-dispatch.md
names. Step 5.75 now dispatches every kind == "step".

The browser-profile lock is tolerated, not coordinated: --isolated is a
flag on the operator's own MCP-server registration that GSD neither
launches nor parameterizes, so the verifier reports could_not_look /
profile_locked, names the flag, and stops. DOM-VERIFY.md keeps
could_not_look and nothing_to_report distinct behind a closed reason
enum — collapsing them is the ambiguous-run-notes defect reported.

Verified on the remote runner.

Closes #2856

* fix(#2856): apply review findings from the orthogonal passes

Correctness pass (blocker):
- delete detectionBlockIsCrlfSafe. It was pass-always: it read the file,
  replaced LF with CRLF, then indexOf'd marker strings that contain no
  newline, so the replacement could not change the result and the
  assertion could never fail for the reason it stated. There is no real
  CRLF risk on this surface either — the gsd:live-dom-families block has
  no parser, only human and agent readers. Deleted rather than replaced,
  per the repo's pass-always-test rule.

Isolated security pass (two minors, both real):
- execute-phase.md step 5.75: this change is what first activates
  kind == "step" dispatch at execute:wave:post, which newly opens the
  ref.command shell path at that loop point. Our own step uses ref.agent
  and never touches it, but the door is now open, so the step-dispatch
  line carries the same in-context validate-before-shell warning the
  sibling gate-dispatch line directly below it already carries.
- gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker
  influenced. Require it wrapped in inline code or a fence, kept short,
  and never left reading as a directive to the next reader.

Verified on the remote runner.

Refs #2856

* fix(#2856): settle the new-agent roster ripple

Checkpoint 2 returned 28 failures, none in the new suite — all of them
the guards that exist to make adding an agent a deliberate act. Each is
a real boundary that had to move:

- docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526),
  so the browser globs lose their backticks; primary-agent counts 21->22,
  roster 33/34->34/35, Verifiers category 1->2
- docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md
  to be classified exactly once
- gsd-dom-verifier: add the anti-heredoc instruction and the commented
  hooks: frontmatter pattern both agent gates require
- gsd-core/bin/shared/model-catalog.json: every shipped agent needs a
  profile entry (#3229)
- copilot-install / kilo-upgrades / qwen-upgrades: expected agent list
  and the 34->35 roster boundary
- execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately
  carries one step now. Asserted as an exact shape — one step, capId
  live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a
  real guard against accidental change rather than being relaxed

Two findings worth naming:

mcp-tool-inheritance (#2526) rejected the agent for documenting
mcp__playwright__* while its tools: line withholds it — a dead
instruction that invites the agent to claim a path it cannot take. The
prose now names the Playwright MCP family without the dispatchable
token, in both the agent and the capability fragment.

runtime-launcher-parity rejected the new gsd_run call: each fenced block
is its own shell, so a workflow step file invoking gsd_run needs its own
canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs.
That script also normalizes explore.md, which is unrelated pre-existing
drift the parity check tolerates, so it is reverted to keep this diff
scoped.

The emitted-drift ack supersedes the spent #3370 entry for
execute-phase.md — it is merged into next, so its ripple is absorbed at
the base and it can no longer clear anything. That is the same supersede
the #3370 entry itself performed on the spent #3324 fragment. Its
unrelated execute-plan.md entry is untouched.

Verified on the remote runner.

Refs #2856

* fix(#2856): drop the stale emitted-drift ack entry

The automated-ui-verification.md entry was written speculatively rather
than from a reported growth, and the check names that precisely: an ack
"written or reworded in THIS diff, but nothing here needed it, so it
explains nothing".

The growth tier keys on the bare filename as it appears under
gsd-core/workflows/ or agents/. automated-ui-verification.md is nested
under verify-work/steps/, so it was never in the tracked set — only
execute-phase.md was ever reported, both before and after the launcher
preamble landed.

Only ack what the check actually reports.

Verified on the remote runner.

Refs #2856

* chore(#2856): backfill changeset pr number

pr:0 -> 3716. The placeholder fails both changeset-lint
(fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by
design and can only be resolved once the PR number exists. Both now
report ok against GITHUB_BASE_REF=next.

Refs #2856

---------

Co-authored-by: sim <sim@local>
2026-08-20 15:07:21 -04:00

5.9 KiB

The live-DOM UAT capability

Why GSD can drive a browser during execution, and why it does that from a purpose-built agent instead of from the plan executor.

The problem

A phase with a live-UI acceptance criterion could not be finished by the agent that executed it. gsd-executor carries no browser tools, so it did the correct thing — returned a checkpoint:human-action at the gate — even though the work was not actually human-only. An orchestrator with a browser MCP server could do it unattended.

The practical result was that every phase with a DOM-level acceptance criterion quietly degraded from executed by the executor to executed by the executor, then finished by hand in the orchestrator. On a UI-heavy project that is routine, not an edge case. Worse, the plan's own autonomous: false marker could not distinguish "a human must judge this" from "the executor lacks the tool" — so the run notes had to explain the deviation every time.

The shape that was not taken

The obvious fix is to add browser globs to agents/gsd-executor.md's tools: line. It is one line, and an absent MCP server simply means the tool is not offered, so it is inert for anyone without one.

That reasoning is correct and it is not the concern. The concern is the user who does have a browser MCP configured for unrelated work. For a first-party agent, the static tools: list is the only control that exists:

  • A capability cannot grant tools to a first-party agent. ADR-1244 D2 rejects an overlay whose id collides with a first-party id or that claims an agent stem already owned. A contribution hook (ADR-857 D4) injects prose into a step's prompt; no hook kind grants tool permissions.
  • There is no per-dispatch tool override. The executor is spawned with subagent_type, description, model, and prompt. The agent definition file is the sole authority on its tool surface.
  • Nothing else contains it. ADR-1244 D5 is explicit: "there is no sandbox… Consent + integrity + reversibility are the barrier." There is no domain allowlist and no gate that inspects what a browser call fetched.

The codebase already reflected this instinct. gsd-ui-auditor — the one subagent that produces UI screenshots — captures via CLI rather than taking an MCP tool grant. The executor carrying only mcp__context7__* while the researcher agents carry the full web-reaching set is a deliberate separation, not an oversight.

So widening the executor would trade a narrow, auditable surface for a permanent broad one, to solve a problem that a narrower surface also solves.

The shape that was taken

One default-off capability owns everything:

  • The key. workflow.live_dom_uat, boolean, default false, declared as the capability's activationKey. With it off the capability resolves inactive, and resolveLoopHooks is fail-closed on state.active === true — the hook does not render at all. The step's own when guard is a second, independent gate.
  • The agent. gsd-dom-verifier carries mcp__chrome-devtools__* and mcp__claude-in-chrome__* in its own tools: line, and carries no Bash. Browser reach is confined to one agent that only exists to look at a DOM.
  • The step. Registered at execute:wave:post as a step hook with onError: skip. A step is additive by construction — it never halts the host. Blocking preconditions are gates, and this capability declares none.

gsd-executor's tool surface is unchanged in every configuration. That is a tested invariant, asserted as an absence, because that is the only way it is observable.

Why the existing Playwright path was left alone

The orchestrator's automated_ui_verification step already used mcp__playwright__*, gated on tool presence plus an active UI phase. Pulling that branch behind a new default-off key would have silently removed working behaviour from every current Playwright-MCP user on upgrade — a regression wearing an enhancement's clothes.

So the key gates only the newly added families. Playwright keeps the gating it already had. The instruction to gate on "presence AND the config key, not presence alone" is what gives the new families the same two-condition shape Playwright already possessed.

Why there is no browser-lock coordination

chrome-devtools-mcp holds an exclusive lock on its browser profile, so parallel execution waves will collide. The tempting design is a lease or queue around the profile.

GSD cannot enforce one. --isolated is a flag on the user's MCP server registration; GSD neither launches that server nor passes its arguments. Coordination over a resource you do not own is theatre — it adds machinery, and the lock still happens.

So the verifier tolerates the lock instead: it reports could_not_look / profile_locked, names --isolated so the operator knows the remedy, and stops. No retry, no wait, no held-up wave. The docs carry the flag; the code does not pretend to.

The distinction the artifact must preserve

DOM-VERIFY.md separates nothing_to_report (there were no UI criteria) from could_not_look (there were criteria, and the check did not happen — with a reason code saying which). Collapsing those two is what produced the ambiguous run notes that opened the original report. A report claiming no issues when it never opened a browser is worse than no report.

Known limits

  • No sandbox. Once the key is on, nothing constrains which origins a browser call reaches. This capability narrows who can reach a browser, not where it may go.
  • Concurrent waves still collide on a shared profile unless the operator passes --isolated.
  • DOM observation only — no screenshot diffing, accessibility audit, or performance tracing.
  • chrome-devtools and claude-in-chrome are detected but not feature-normalized; the verifier uses whichever answers and does not paper over differences between them.