How the conformance harness is factored so any host runs the same suite —
five web chat clients and two Electron desktop apps (Cursor, Goose) today. All the
platform-specific code lives behind one Host interface.
Runs in the iframe. Owns the test
definitions and the MCP-app communication (the ext-apps App). A test
emits typed capability requests, awaits the results, and asserts.
Platform-agnostic. Lists tests, then pumps each request the suite parks to the Host and feeds the result back. A generic dispatcher — no per-test logic.
The only platform-specific piece. Opens the app,
prompts the agent so the suite renders, exposes one method per capability.
BrowserHost (Playwright) covers web chat clients and Electron
desktop apps, over CDP.
The protocol is deliberately primitive — one variant per thing a
host can physically do. The test carries the meaning (e.g. "the tool must
stay hidden" is a test that negates a conversationContains result, not a
dedicated request kind).
// shared/protocol.ts — the one contract both sides import as source
export type CapabilityRequest =
| { kind: "clickTrigger"; commitDraftedMessage?: boolean } // real cross-origin click (user activation)
| { kind: "confirmDialog"; dialog: "download" | "sampling" } // host-native permission dialog
| { kind: "checkLinkOpen"; url: string } // ui/open-link actually opened THIS url
| { kind: "conversationContains"; marker: string; timeoutMs: number }
| { kind: "toggleTheme"; to: "light" | "dark" }
| { kind: "readModelToolList" } // optional desktop-host affordance
| { kind: "inspectFrame" } // read the host page's <iframe> sandbox/allow/csp attrs
| { kind: "readConsole"; pattern: string; timeoutMs: number }
| { kind: "resetIsolation" }; // suite emits this before each manual test
export interface CapabilityResult {
ok: boolean; value?: unknown; error?: string;
unsupported?: boolean; // host lacks the capability → test skips / falls back
}
// Reported by the host at ui/initialize and recorded next to the results, so a
// run is attributable to a build — not just to a product name.
export interface HostImplementation { name?: string; version?: string; title?: string }
A sandboxed, nested, cross-origin iframe can't reliably push a message out
(nested cross-origin postMessage drops), but the Runner can always reach
in via frame.evaluate. So the direction is fixed: a test
awaits t.host(req), which parks req in a single
pending slot; the Runner poll()s that slot, services the request against
the Host, and calls resolve(result) to unblock the test.
window.__mcpConformance with listTests() ·
start(filter?) · poll() · resolve(result).
A desktop host swaps frame.evaluate for its own transport (IPC / a
WebSocket to the app process) behind the same SuiteBridge interface —
that is the entire drop-in seam. The same pending slot also backs the in-iframe
yes/no/skip buttons, so the suite runs with no driver at all (a human answers).Only setup() and teardown() are mandatory on a
Host; every capability method is optional. An absent method resolves to
{ unsupported: true }. In the suite:
t.host(req) — auto-skips the test if the capability
is unsupported ("I need this or I can't run here").t.hostOptional(req) — returns the raw result so the test can
fall back to another path.// visibility/app-tool-hidden — capability-or-fallback
const direct = await t.hostOptional({ kind: "readModelToolList" });
if (!direct.unsupported) { // a desktop host that can introspect the model's tools
t.assert(!(direct.value as string[]).includes("conformance_probe"), "hidden tool leaked");
return;
}
// browser hosts don't expose that — fall back to asking the agent and scanning the conversation
t.bindTrigger(() => t.app.sendMessage({ role: "user", content: [{ type: "text", text: ASK }] }));
await t.host({ kind: "clickTrigger", commitDraftedMessage: true });
const r = await t.host({ kind: "conversationContains", marker: "conformance_probe", timeoutMs: 45_000 });
t.assert(!r.ok, "hidden tool name surfaced in the conversation");
Products differ in behavior — how you enter a prompt, dismiss a modal, verify a
conversation turn, commit a drafted message — and in which capabilities they support.
So they are subclasses of a shared BrowserHost, not config rows:
| Host | Prompt entry | Verify a turn | Notes |
|---|---|---|---|
| ChatGPTBrowserHost | #prompt-textarea, mention picker | backend conversation API | sends ui/message directly |
| ClaudeBrowserHost | ProseMirror contenteditable | scan transcript text | drafts ui/message → Send |
| MistralBrowserHost | ProseMirror contenteditable | scan transcript text | consent reads “Confirm”; stops generation between tests |
| ManufactBrowserHost | textarea[data-testid=chat-input] | scan transcript text | third-party inspector; no login |
| AlpicPlaygroundBrowserHost | textarea[name=message] | scan transcript text | no login |
| GooseBrowserHost desktop | textarea[data-testid=chat-input] | scan renderer DOM | Electron; attaches over CDP |
| CursorBrowserHost desktop | [contenteditable], fresh “New Agent” | scan webview frames | Electron; app renders in a nested vscode-webview:// |
The desktop hosts are the load-bearing evidence that the seam
generalizes. GooseHost/CursorHost were expected to be peer
Host implementations with their own transport — in practice both are plain
BrowserHost subclasses that override one method, open(), to
attach to a running Electron app over CDP instead of launching Chrome. Everything above
that — the bridge, the real cross-origin clicks, frame inspection — is reused verbatim,
even though the app renders inside a nested webview. Anything a host can't do is simply
unsupported and those tests skip: a desktop host may implement
readModelToolList that browser hosts can't, and conversely neither desktop
host can observe ui/open-link (it opens in the OS browser, outside the
driver), so that test honestly reports “can't verify” rather than a false failure.
Where the driver clicks real product UI it necessarily overfits per-host DOM — see
How it works.
Two pieces ship on npm, matching the two halves of the diagram. Nothing about the suite is specific to a vendor — the only code you write is the adapter.
One command. Connect the printed URL as an MCP server in your host, then
ask the agent to run run_conformance.
npx mcp-apps-conformance # serves the ui:// TestSuite on :3000/mcp npx mcp-apps-conformance --stdio # or over stdio, for a desktop host's config
Subclass BrowserHost, fill in the three hooks that describe
your product's UI, and the generic Runner does the rest — returning
structured SubtestResult[] you can assert on in CI.
import { BrowserHost, Runner } from "mcp-apps-conformance";
import type { Page } from "playwright";
class MyHost extends BrowserHost {
readonly name = "my-host";
readonly url = "https://my-host.example/chat";
readonly widgetSelector = 'iframe[src*="my-sandbox-origin"]';
// Get the agent to render the conformance app.
protected async sendPrompt(page: Page, appName: string) {
await page.fill("#composer", `run ${appName}`);
await page.keyboard.press("Enter");
}
protected async dismissModal(_page: Page) {} // cookie banner / login wall, if any
// Did this turn actually reach the conversation?
protected async verifyConversation(page: Page, marker: string, timeoutMs: number) {
return this.pollMarker(page, (m) => document.body.innerText.includes(m) ? "found" : "no",
marker, timeoutMs, 3_000);
}
}
const { results } = await new Runner(new MyHost(), {
appName: "MCP Apps Conformance",
profileDir: ".profile/my-host",
}).run();
const failed = results.filter((r) => r.status === "FAIL");
if (failed.length) process.exit(1); // gate your CI on the spec
Every capability method is optional, so an adapter is useful long before
it is complete: implement setup/teardown and the tests that
need nothing else already run. Playwright is an optional peer dependency — needed only
for this half, not to serve the suite.