* test(#2856): add failing-first suite for the live-dom-uat capability Binds the approved triage shape before any of it exists: - containment — the execute:wave:post hook must not render unless workflow.live_dom_uat is true AND the capability resolves active (fail-closed on a missing state entry, and on a non-boolean value) - criterion 4 — agents/gsd-executor.md carries no browser MCP family; asserted as an absence, which is the only way it is observable - Hyrum guard — the pre-existing mcp__playwright__* branch must stay outside the key-gated block, or upgrading silently removes working automated UI verification for every current Playwright-MCP user - parity — the browser glob list now lives in two surfaces (agent frontmatter + workflow detection block); the assertion fails if either gains or loses a family without the other Red by construction: the capability, agent and workflow block do not exist yet. Verified on the remote runner. Refs #2856 * enhance(#2856): add default-off live-DOM UAT capability A phase whose acceptance criteria needed a live DOM could not be finished by the agent that executed it: gsd-executor carries no browser tools, so it correctly returned checkpoint:human-action even though the work was not human-only, just tool-less. Every such phase degraded to "executed, then finished by hand in the orchestrator", and autonomous: false could not distinguish "a human must judge this" from "the executor lacks the tool". Implements the shape approved at triage, not the one reported. The executor's tools: line is NOT widened, in any configuration: for a first-party agent the static list is the only control that exists (ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one default-off capability owns the key, the agent, and the step: - capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat (boolean, default false), one additive step at execute:wave:post (onError: skip, gates: []), so it can never halt a wave - agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP globs, in its own tools: line, with no Bash - verify-work automated_ui_verification — a gsd:live-dom-families block naming both new families AND the key; presence alone never activates Two independent fail-closed gates: isCapabilityActive renders a hook only on state.active === true, plus the step's own `when`. The pre-existing mcp__playwright__* branch keeps the gating it already had and stays outside the new block. Pulling it behind a default-off key would have silently removed working automated UI verification from every current Playwright-MCP user on upgrade. Also closes a host gap this surfaced: execute:wave:post dispatched only contribution + gate, so ANY registered step was declared and silently never run — exactly the single-kind hand-roll loop-hook-dispatch.md names. Step 5.75 now dispatches every kind == "step". The browser-profile lock is tolerated, not coordinated: --isolated is a flag on the operator's own MCP-server registration that GSD neither launches nor parameterizes, so the verifier reports could_not_look / profile_locked, names the flag, and stops. DOM-VERIFY.md keeps could_not_look and nothing_to_report distinct behind a closed reason enum — collapsing them is the ambiguous-run-notes defect reported. Verified on the remote runner. Closes #2856 * fix(#2856): apply review findings from the orthogonal passes Correctness pass (blocker): - delete detectionBlockIsCrlfSafe. It was pass-always: it read the file, replaced LF with CRLF, then indexOf'd marker strings that contain no newline, so the replacement could not change the result and the assertion could never fail for the reason it stated. There is no real CRLF risk on this surface either — the gsd:live-dom-families block has no parser, only human and agent readers. Deleted rather than replaced, per the repo's pass-always-test rule. Isolated security pass (two minors, both real): - execute-phase.md step 5.75: this change is what first activates kind == "step" dispatch at execute:wave:post, which newly opens the ref.command shell path at that loop point. Our own step uses ref.agent and never touches it, but the door is now open, so the step-dispatch line carries the same in-context validate-before-shell warning the sibling gate-dispatch line directly below it already carries. - gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker influenced. Require it wrapped in inline code or a fence, kept short, and never left reading as a directive to the next reader. Verified on the remote runner. Refs #2856 * fix(#2856): settle the new-agent roster ripple Checkpoint 2 returned 28 failures, none in the new suite — all of them the guards that exist to make adding an agent a deliberate act. Each is a real boundary that had to move: - docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526), so the browser globs lose their backticks; primary-agent counts 21->22, roster 33/34->34/35, Verifiers category 1->2 - docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md to be classified exactly once - gsd-dom-verifier: add the anti-heredoc instruction and the commented hooks: frontmatter pattern both agent gates require - gsd-core/bin/shared/model-catalog.json: every shipped agent needs a profile entry (#3229) - copilot-install / kilo-upgrades / qwen-upgrades: expected agent list and the 34->35 roster boundary - execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately carries one step now. Asserted as an exact shape — one step, capId live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a real guard against accidental change rather than being relaxed Two findings worth naming: mcp-tool-inheritance (#2526) rejected the agent for documenting mcp__playwright__* while its tools: line withholds it — a dead instruction that invites the agent to claim a path it cannot take. The prose now names the Playwright MCP family without the dispatchable token, in both the agent and the capability fragment. runtime-launcher-parity rejected the new gsd_run call: each fenced block is its own shell, so a workflow step file invoking gsd_run needs its own canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs. That script also normalizes explore.md, which is unrelated pre-existing drift the parity check tolerates, so it is reverted to keep this diff scoped. The emitted-drift ack supersedes the spent #3370 entry for execute-phase.md — it is merged into next, so its ripple is absorbed at the base and it can no longer clear anything. That is the same supersede the #3370 entry itself performed on the spent #3324 fragment. Its unrelated execute-plan.md entry is untouched. Verified on the remote runner. Refs #2856 * fix(#2856): drop the stale emitted-drift ack entry The automated-ui-verification.md entry was written speculatively rather than from a reported growth, and the check names that precisely: an ack "written or reworded in THIS diff, but nothing here needed it, so it explains nothing". The growth tier keys on the bare filename as it appears under gsd-core/workflows/ or agents/. automated-ui-verification.md is nested under verify-work/steps/, so it was never in the tracked set — only execute-phase.md was ever reported, both before and after the launcher preamble landed. Only ack what the check actually reports. Verified on the remote runner. Refs #2856 * chore(#2856): backfill changeset pr number pr:0 -> 3716. The placeholder fails both changeset-lint (fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by design and can only be resolved once the PR number exists. Both now report ok against GITHUB_BASE_REF=next. Refs #2856 --------- Co-authored-by: sim <sim@local>
5.9 KiB
The live-DOM UAT capability
Why GSD can drive a browser during execution, and why it does that from a purpose-built agent instead of from the plan executor.
The problem
A phase with a live-UI acceptance criterion could not be finished by the agent that executed
it. gsd-executor carries no browser tools, so it did the correct thing — returned a
checkpoint:human-action at the gate — even though the work was not actually human-only. An
orchestrator with a browser MCP server could do it unattended.
The practical result was that every phase with a DOM-level acceptance criterion quietly
degraded from executed by the executor to executed by the executor, then finished by hand in
the orchestrator. On a UI-heavy project that is routine, not an edge case. Worse, the plan's
own autonomous: false marker could not distinguish "a human must judge this" from
"the executor lacks the tool" — so the run notes had to explain the deviation every time.
The shape that was not taken
The obvious fix is to add browser globs to agents/gsd-executor.md's tools: line. It is one
line, and an absent MCP server simply means the tool is not offered, so it is inert for anyone
without one.
That reasoning is correct and it is not the concern. The concern is the user who does have
a browser MCP configured for unrelated work. For a first-party agent, the static tools: list
is the only control that exists:
- A capability cannot grant tools to a first-party agent. ADR-1244
D2 rejects an overlay whose
idcollides with a first-party id or that claims an agent stem already owned. Acontributionhook (ADR-857 D4) injects prose into a step's prompt; no hook kind grants tool permissions. - There is no per-dispatch tool override. The executor is spawned with
subagent_type,description,model, andprompt. The agent definition file is the sole authority on its tool surface. - Nothing else contains it. ADR-1244 D5 is explicit: "there is no sandbox… Consent + integrity + reversibility are the barrier." There is no domain allowlist and no gate that inspects what a browser call fetched.
The codebase already reflected this instinct. gsd-ui-auditor — the one subagent that produces
UI screenshots — captures via CLI rather than taking an MCP tool grant. The executor carrying
only mcp__context7__* while the researcher agents carry the full web-reaching set is a
deliberate separation, not an oversight.
So widening the executor would trade a narrow, auditable surface for a permanent broad one, to solve a problem that a narrower surface also solves.
The shape that was taken
One default-off capability owns everything:
- The key.
workflow.live_dom_uat,boolean, defaultfalse, declared as the capability'sactivationKey. With it off the capability resolves inactive, andresolveLoopHooksis fail-closed onstate.active === true— the hook does not render at all. The step's ownwhenguard is a second, independent gate. - The agent.
gsd-dom-verifiercarriesmcp__chrome-devtools__*andmcp__claude-in-chrome__*in its owntools:line, and carries noBash. Browser reach is confined to one agent that only exists to look at a DOM. - The step. Registered at
execute:wave:postas astephook withonError: skip. A step is additive by construction — it never halts the host. Blocking preconditions aregates, and this capability declares none.
gsd-executor's tool surface is unchanged in every configuration. That is a tested invariant,
asserted as an absence, because that is the only way it is observable.
Why the existing Playwright path was left alone
The orchestrator's automated_ui_verification step already used mcp__playwright__*, gated on
tool presence plus an active UI phase. Pulling that branch behind a new default-off key
would have silently removed working behaviour from every current Playwright-MCP user on
upgrade — a regression wearing an enhancement's clothes.
So the key gates only the newly added families. Playwright keeps the gating it already had. The instruction to gate on "presence AND the config key, not presence alone" is what gives the new families the same two-condition shape Playwright already possessed.
Why there is no browser-lock coordination
chrome-devtools-mcp holds an exclusive lock on its browser profile, so parallel execution
waves will collide. The tempting design is a lease or queue around the profile.
GSD cannot enforce one. --isolated is a flag on the user's MCP server registration; GSD
neither launches that server nor passes its arguments. Coordination over a resource you do not
own is theatre — it adds machinery, and the lock still happens.
So the verifier tolerates the lock instead: it reports could_not_look / profile_locked,
names --isolated so the operator knows the remedy, and stops. No retry, no wait, no held-up
wave. The docs carry the flag; the code does not pretend to.
The distinction the artifact must preserve
DOM-VERIFY.md separates nothing_to_report (there were no UI criteria) from
could_not_look (there were criteria, and the check did not happen — with a reason code
saying which). Collapsing those two is what produced the ambiguous run notes that opened the
original report. A report claiming no issues when it never opened a browser is worse than no
report.
Known limits
- No sandbox. Once the key is on, nothing constrains which origins a browser call reaches. This capability narrows who can reach a browser, not where it may go.
- Concurrent waves still collide on a shared profile unless the operator passes
--isolated. - DOM observation only — no screenshot diffing, accessibility audit, or performance tracing.
chrome-devtoolsandclaude-in-chromeare detected but not feature-normalized; the verifier uses whichever answers and does not paper over differences between them.