* test(#2856): add failing-first suite for the live-dom-uat capability Binds the approved triage shape before any of it exists: - containment — the execute:wave:post hook must not render unless workflow.live_dom_uat is true AND the capability resolves active (fail-closed on a missing state entry, and on a non-boolean value) - criterion 4 — agents/gsd-executor.md carries no browser MCP family; asserted as an absence, which is the only way it is observable - Hyrum guard — the pre-existing mcp__playwright__* branch must stay outside the key-gated block, or upgrading silently removes working automated UI verification for every current Playwright-MCP user - parity — the browser glob list now lives in two surfaces (agent frontmatter + workflow detection block); the assertion fails if either gains or loses a family without the other Red by construction: the capability, agent and workflow block do not exist yet. Verified on the remote runner. Refs #2856 * enhance(#2856): add default-off live-DOM UAT capability A phase whose acceptance criteria needed a live DOM could not be finished by the agent that executed it: gsd-executor carries no browser tools, so it correctly returned checkpoint:human-action even though the work was not human-only, just tool-less. Every such phase degraded to "executed, then finished by hand in the orchestrator", and autonomous: false could not distinguish "a human must judge this" from "the executor lacks the tool". Implements the shape approved at triage, not the one reported. The executor's tools: line is NOT widened, in any configuration: for a first-party agent the static list is the only control that exists (ADR-1244 D2, ADR-857 D4, no per-dispatch override). Instead one default-off capability owns the key, the agent, and the step: - capabilities/live-dom-uat/ — activationKey workflow.live_dom_uat (boolean, default false), one additive step at execute:wave:post (onError: skip, gates: []), so it can never halt a wave - agents/gsd-dom-verifier.md — the only GSD agent carrying browser MCP globs, in its own tools: line, with no Bash - verify-work automated_ui_verification — a gsd:live-dom-families block naming both new families AND the key; presence alone never activates Two independent fail-closed gates: isCapabilityActive renders a hook only on state.active === true, plus the step's own `when`. The pre-existing mcp__playwright__* branch keeps the gating it already had and stays outside the new block. Pulling it behind a default-off key would have silently removed working automated UI verification from every current Playwright-MCP user on upgrade. Also closes a host gap this surfaced: execute:wave:post dispatched only contribution + gate, so ANY registered step was declared and silently never run — exactly the single-kind hand-roll loop-hook-dispatch.md names. Step 5.75 now dispatches every kind == "step". The browser-profile lock is tolerated, not coordinated: --isolated is a flag on the operator's own MCP-server registration that GSD neither launches nor parameterizes, so the verifier reports could_not_look / profile_locked, names the flag, and stops. DOM-VERIFY.md keeps could_not_look and nothing_to_report distinct behind a closed reason enum — collapsing them is the ambiguous-run-notes defect reported. Verified on the remote runner. Closes #2856 * fix(#2856): apply review findings from the orthogonal passes Correctness pass (blocker): - delete detectionBlockIsCrlfSafe. It was pass-always: it read the file, replaced LF with CRLF, then indexOf'd marker strings that contain no newline, so the replacement could not change the result and the assertion could never fail for the reason it stated. There is no real CRLF risk on this surface either — the gsd:live-dom-families block has no parser, only human and agent readers. Deleted rather than replaced, per the repo's pass-always-test rule. Isolated security pass (two minors, both real): - execute-phase.md step 5.75: this change is what first activates kind == "step" dispatch at execute:wave:post, which newly opens the ref.command shell path at that loop point. Our own step uses ref.agent and never touches it, but the door is now open, so the step-dispatch line carries the same in-context validate-before-shell warning the sibling gate-dispatch line directly below it already carries. - gsd-dom-verifier: quoted page text in DOM-VERIFY.md is attacker influenced. Require it wrapped in inline code or a fence, kept short, and never left reading as a directive to the next reader. Verified on the remote runner. Refs #2856 * fix(#2856): settle the new-agent roster ripple Checkpoint 2 returned 28 failures, none in the new suite — all of them the guards that exist to make adding an agent a deliberate act. Each is a real boundary that had to move: - docs/AGENTS.md: Tools row must copy the frontmatter verbatim (#2526), so the browser globs lose their backticks; primary-agent counts 21->22, roster 33/34->34/35, Verifiers category 1->2 - docs/INVENTORY.md: roster completeness requires every agents/gsd-*.md to be classified exactly once - gsd-dom-verifier: add the anti-heredoc instruction and the commented hooks: frontmatter pattern both agent gates require - gsd-core/bin/shared/model-catalog.json: every shipped agent needs a profile entry (#3229) - copilot-install / kilo-upgrades / qwen-upgrades: expected agent list and the 34->35 roster boundary - execute-wave-post-gate-pipeline-e2e: execute:wave:post legitimately carries one step now. Asserted as an exact shape — one step, capId live-dom-uat, ref.agent gsd-dom-verifier, onError skip — so it stays a real guard against accidental change rather than being relaxed Two findings worth naming: mcp-tool-inheritance (#2526) rejected the agent for documenting mcp__playwright__* while its tools: line withholds it — a dead instruction that invites the agent to claim a path it cannot take. The prose now names the Playwright MCP family without the dispatchable token, in both the agent and the capability fragment. runtime-launcher-parity rejected the new gsd_run call: each fenced block is its own shell, so a workflow step file invoking gsd_run needs its own canonical preamble. Propagated with scripts/sync-runtime-launcher.cjs. That script also normalizes explore.md, which is unrelated pre-existing drift the parity check tolerates, so it is reverted to keep this diff scoped. The emitted-drift ack supersedes the spent #3370 entry for execute-phase.md — it is merged into next, so its ripple is absorbed at the base and it can no longer clear anything. That is the same supersede the #3370 entry itself performed on the spent #3324 fragment. Its unrelated execute-plan.md entry is untouched. Verified on the remote runner. Refs #2856 * fix(#2856): drop the stale emitted-drift ack entry The automated-ui-verification.md entry was written speculatively rather than from a reported growth, and the check names that precisely: an ack "written or reworded in THIS diff, but nothing here needed it, so it explains nothing". The growth tier keys on the bare filename as it appears under gsd-core/workflows/ or agents/. automated-ui-verification.md is nested under verify-work/steps/, so it was never in the tracked set — only execute-phase.md was ever reported, both before and after the launcher preamble landed. Only ack what the check actually reports. Verified on the remote runner. Refs #2856 * chore(#2856): backfill changeset pr number pr:0 -> 3716. The placeholder fails both changeset-lint (fail_invalid_fragment) and docs-lint (fail_malformed_fragment) by design and can only be resolved once the PR number exists. Both now report ok against GITHUB_BASE_REF=next. Refs #2856 --------- Co-authored-by: sim <sim@local>
4.8 KiB
How to enable live-DOM verification
Let GSD open a real browser and check a phase's UI acceptance criteria against the live DOM — during execution, not only after it — without widening what the plan executor can reach.
Default-off, and deliberately so. A browser MCP server you configured for unrelated work must not start driving your project's UI on its own. You opt in per project with one key. See the explanation for why the executor's own tool surface was left alone.
What you need:
- GSD installed with the
fullprofile (the capability istier: full). - A browser MCP server registered in your runtime — either
chrome-devtools-mcp(exposesmcp__chrome-devtools__*) or Claude-in-Chrome (exposesmcp__claude-in-chrome__*). - Something serving your UI — a dev server, a preview deployment, any reachable URL.
- A phase whose plan actually states UI acceptance criteria. The verifier will not invent them.
Step 1 — Turn the key on
gsd-tools query config-set workflow.live_dom_uat true
Verify it took:
gsd-tools query config-get workflow.live_dom_uat
# → true
That one key gates both halves: the gsd-dom-verifier step that runs after each execution
wave, and the extra browser families the orchestrator's own UI-verification step will consider.
With it off, neither reaches a browser.
Step 2 — Make the browser reachable to more than one wave
chrome-devtools-mcp keeps an exclusive lock on its browser profile at
$HOME/.cache/chrome-devtools-mcp/chrome-profile. A second instance fails with:
The browser is already running for <dir>. Use --isolated to run multiple browser instances.
GSD runs execution waves in parallel, so two verifiers can reach for one profile. GSD cannot
fix this for you — --isolated is a flag on your MCP server registration, not something
GSD passes. Add it there:
{
"mcpServers": {
"chrome-devtools": {
"command": "npx",
"args": ["-y", "chrome-devtools-mcp@latest", "--isolated"]
}
}
}
--isolated gives each instance a throwaway profile. If you would rather share one server
across concurrent agents, --experimentalPageIdRouting routes tools per page instead.
Skipping this step is safe — you just get could_not_look / profile_locked on the waves that
lost the race, never a failed wave.
Step 3 — Run a phase and read the report
Execute normally. After each wave, gsd-dom-verifier writes
.planning/phases/<phase>/<n>-DOM-VERIFY.md:
---
schema_version: 1
wave: 2
outcome: verified
reason: ok
checked: 4
passed: 3
failed: 0
needs_review: 1
---
The body lists one line per criterion with the observation behind its verdict.
Reading the outcome — "nothing to report" is not "could not look"
This is the part worth learning, because a report that says no issues when it never opened a browser is worse than no report at all.
outcome |
reason |
What actually happened | What to do |
|---|---|---|---|
verified |
ok |
Criteria existed and were observed | Read the per-criterion lines |
nothing_to_report |
no_criteria |
The wave's plan stated no UI acceptance criteria | Nothing. This is a clean result |
could_not_look |
no_browser_mcp |
Key is on, but no browser MCP answered | Check your MCP server is registered and running |
could_not_look |
profile_locked |
Another instance holds the browser profile | Add --isolated — see Step 2 |
could_not_look |
target_unreachable |
Nothing was serving the criterion's URL | Start your dev server before executing |
Only could_not_look means the check did not happen. nothing_to_report means it happened and
found nothing to check.
What this does not do
- It never blocks. The step is advisory by construction — it cannot fail a task, fail a wave, or stop a phase. Findings are findings; the executor still owns task outcomes.
- It does not widen the executor.
gsd-executorcarries no browser tools in any configuration. The browser reach lives ingsd-dom-verifieralone. - It does not sandbox the browser. Once the key is on there is no domain allowlist and nothing inspects what a page fetched. Turn it on for projects where that is acceptable.
- It observes the DOM only. No screenshot diffing, no accessibility audit, no performance
tracing. A criterion needing one of those comes back
needs_reviewwith the reason named.
Turning it back off
gsd-tools query config-set workflow.live_dom_uat false
The capability resolves inactive immediately and the hook stops rendering. See Turn a capability off (and keep it off) for removing it entirely.