Challenge candidate · authentic evidence

Semantic release testing for WebMCP

Unit tests for meaning.

For product, QA, safety, and release teams shipping agent-callable sites.

Handler tests prove a tool can run. They do not prove that a natural-language request selected the human-approved action or produced the represented page effect.

  1. Human declares meaning
  2. Agent acts through WebMCP
  3. Thurstone verifies the effect
  4. Reviewer decides
Inspect sealed ResultsOpen live WebMCP Lab

Results works anywhere. Lab ready = tools offered → found → executable. Requires the ChatGPT in-app browser or Chrome 149+ with WebMCP.

Review the human-approved contract
Same approved meaningEquivalent action
Meaning-changing boundaryRequired action difference

Selection, canonical arguments, approval posture, and page effects remain independently inspectable.

Human + agent workflow

From ambiguous wording to inspectable behavior

01

Declare meaning

A human and agent draft the intended action, arguments, effects, and semantic boundaries.

02

Run the same contract

Fresh model contexts act through the live WebMCP catalog from a verified fixture.

03

Inspect exact effects

Trace-derived evidence separates tool choice, arguments, observable state, and over-action.

60-second judge path

One click. One fixed model decision.

In a supported Chrome/WebMCP browser, the public judge lane asks one server-fixed cart question, exposes the model selection, and verifies the returned read through the live native catalog.

  1. Open the Lab signed out; no key or Thurstone login is needed.
  2. Confirm consumer-ready on the clean four-tool catalog.
  3. Run the bounded judge proof and inspect/download its receipts.
  4. Open Results for the separate 24-case paired evidence.