← ResultsHow it works

Architecture: Host / Runner / TestSuite

How the conformance harness is factored so any host runs the same suite — five web chat clients and two Electron desktop apps (Cursor, Goose) today. All the platform-specific code lives behind one Host interface.

Three objects

TestSuite

Runs in the iframe. Owns the test definitions and the MCP-app communication (the ext-apps App). A test emits typed capability requests, awaits the results, and asserts.

Runner

Platform-agnostic. Lists tests, then pumps each request the suite parks to the Host and feeds the result back. A generic dispatcher — no per-test logic.

Host

The only platform-specific piece. Opens the app, prompts the agent so the suite renders, exposes one method per capability. BrowserHost (Playwright) covers web chat clients and Electron desktop apps, over CDP.

The MCP server serves a ui:// TestSuite; the host renders it in a sandboxed iframe; an external Runner polls the suite over window.__mcpConformance and dispatches each capability request to a host adapter that drives the real product UI. npx mcp-apps-conformance the product under test import { Runner, BrowserHost } MCP server ui://conformance/runner the TestSuite — one HTML file conformance_probe app-visible fixture model_only_probe must stay hidden from the app MCP resources/read Host under test ChatGPT · Claude · Cursor · Goose · Mistral · … sandboxed iframe the TestSuite runs in here asserts the spec from in-view parks a typed CapabilityRequest window.__mcpConformance chat transcript · dialogs · theme · console poll() resolve() real clicks DOM · console Runner platform-agnostic · no per-test logic poll → dispatch → resolve your Host adapter extends BrowserHost sendPrompt · dismissModal verifyConversation the only part you write roughly 20 lines of TypeScript SubtestResult[] PASS · FAIL · SKIP · TIMEOUT → results matrix

The typed capability protocol

The protocol is deliberately primitive — one variant per thing a host can physically do. The test carries the meaning (e.g. "the tool must stay hidden" is a test that negates a conversationContains result, not a dedicated request kind).

// shared/protocol.ts — the one contract both sides import as source
export type CapabilityRequest =
  | { kind: "clickTrigger"; commitDraftedMessage?: boolean }   // real cross-origin click (user activation)
  | { kind: "confirmDialog"; dialog: "download" | "sampling" } // host-native permission dialog
  | { kind: "checkLinkOpen"; url: string }                     // ui/open-link actually opened THIS url
  | { kind: "conversationContains"; marker: string; timeoutMs: number }
  | { kind: "toggleTheme"; to: "light" | "dark" }
  | { kind: "readModelToolList" }        // optional desktop-host affordance
  | { kind: "inspectFrame" }             // read the host page's <iframe> sandbox/allow/csp attrs
  | { kind: "readConsole"; pattern: string; timeoutMs: number }
  | { kind: "resetIsolation" };          // suite emits this before each manual test

export interface CapabilityResult {
  ok: boolean; value?: unknown; error?: string;
  unsupported?: boolean;                 // host lacks the capability → test skips / falls back
}

// Reported by the host at ui/initialize and recorded next to the results, so a
// run is attributable to a build — not just to a product name.
export interface HostImplementation { name?: string; version?: string; title?: string }

Pull model: the suite pulls, the Runner polls

A sandboxed, nested, cross-origin iframe can't reliably push a message out (nested cross-origin postMessage drops), but the Runner can always reach in via frame.evaluate. So the direction is fixed: a test awaits t.host(req), which parks req in a single pending slot; the Runner poll()s that slot, services the request against the Host, and calls resolve(result) to unblock the test.

The suite installs one object at window.__mcpConformance with listTests() · start(filter?) · poll() · resolve(result). A desktop host swaps frame.evaluate for its own transport (IPC / a WebSocket to the app process) behind the same SuiteBridge interface — that is the entire drop-in seam. The same pending slot also backs the in-iframe yes/no/skip buttons, so the suite runs with no driver at all (a human answers).

Optional capabilities & fallback

Only setup() and teardown() are mandatory on a Host; every capability method is optional. An absent method resolves to { unsupported: true }. In the suite:

// visibility/app-tool-hidden — capability-or-fallback
const direct = await t.hostOptional({ kind: "readModelToolList" });
if (!direct.unsupported) {               // a desktop host that can introspect the model's tools
  t.assert(!(direct.value as string[]).includes("conformance_probe"), "hidden tool leaked");
  return;
}
// browser hosts don't expose that — fall back to asking the agent and scanning the conversation
t.bindTrigger(() => t.app.sendMessage({ role: "user", content: [{ type: "text", text: ASK }] }));
await t.host({ kind: "clickTrigger", commitDraftedMessage: true });
const r = await t.host({ kind: "conversationContains", marker: "conformance_probe", timeoutMs: 45_000 });
t.assert(!r.ok, "hidden tool name surfaced in the conversation");

Pluggable hosts — seven, including two desktop apps

Products differ in behavior — how you enter a prompt, dismiss a modal, verify a conversation turn, commit a drafted message — and in which capabilities they support. So they are subclasses of a shared BrowserHost, not config rows:

HostPrompt entryVerify a turnNotes
ChatGPTBrowserHost#prompt-textarea, mention pickerbackend conversation APIsends ui/message directly
ClaudeBrowserHostProseMirror contenteditablescan transcript textdrafts ui/message → Send
MistralBrowserHostProseMirror contenteditablescan transcript textconsent reads “Confirm”; stops generation between tests
ManufactBrowserHosttextarea[data-testid=chat-input]scan transcript textthird-party inspector; no login
AlpicPlaygroundBrowserHosttextarea[name=message]scan transcript textno login
GooseBrowserHost desktoptextarea[data-testid=chat-input]scan renderer DOMElectron; attaches over CDP
CursorBrowserHost desktop[contenteditable], fresh “New Agent”scan webview framesElectron; app renders in a nested vscode-webview://

The desktop hosts are the load-bearing evidence that the seam generalizes. GooseHost/CursorHost were expected to be peer Host implementations with their own transport — in practice both are plain BrowserHost subclasses that override one method, open(), to attach to a running Electron app over CDP instead of launching Chrome. Everything above that — the bridge, the real cross-origin clicks, frame inspection — is reused verbatim, even though the app renders inside a nested webview. Anything a host can't do is simply unsupported and those tests skip: a desktop host may implement readModelToolList that browser hosts can't, and conversely neither desktop host can observe ui/open-link (it opens in the OS browser, outside the driver), so that test honestly reports “can't verify” rather than a false failure. Where the driver clicks real product UI it necessarily overfits per-host DOM — see How it works.

Run it against your own host

Two pieces ship on npm, matching the two halves of the diagram. Nothing about the suite is specific to a vendor — the only code you write is the adapter.

1 · The server, with the suite inside it

One command. Connect the printed URL as an MCP server in your host, then ask the agent to run run_conformance.

npx mcp-apps-conformance          # serves the ui:// TestSuite on :3000/mcp
npx mcp-apps-conformance --stdio  # or over stdio, for a desktop host's config

2 · The runner, with your host

Subclass BrowserHost, fill in the three hooks that describe your product's UI, and the generic Runner does the rest — returning structured SubtestResult[] you can assert on in CI.

import { BrowserHost, Runner } from "mcp-apps-conformance";
import type { Page } from "playwright";

class MyHost extends BrowserHost {
  readonly name = "my-host";
  readonly url = "https://my-host.example/chat";
  readonly widgetSelector = 'iframe[src*="my-sandbox-origin"]';

  // Get the agent to render the conformance app.
  protected async sendPrompt(page: Page, appName: string) {
    await page.fill("#composer", `run ${appName}`);
    await page.keyboard.press("Enter");
  }
  protected async dismissModal(_page: Page) {}   // cookie banner / login wall, if any

  // Did this turn actually reach the conversation?
  protected async verifyConversation(page: Page, marker: string, timeoutMs: number) {
    return this.pollMarker(page, (m) => document.body.innerText.includes(m) ? "found" : "no",
                           marker, timeoutMs, 3_000);
  }
}

const { results } = await new Runner(new MyHost(), {
  appName: "MCP Apps Conformance",
  profileDir: ".profile/my-host",
}).run();

const failed = results.filter((r) => r.status === "FAIL");
if (failed.length) process.exit(1);   // gate your CI on the spec

Every capability method is optional, so an adapter is useful long before it is complete: implement setup/teardown and the tests that need nothing else already run. Playwright is an optional peer dependency — needed only for this half, not to serve the suite.