Declare meaning
A human and agent draft the intended action, arguments, effects, and semantic boundaries.
Semantic release testing for WebMCP
For product, QA, safety, and release teams shipping agent-callable sites.
Handler tests prove a tool can run. They do not prove that a natural-language request selected the human-approved action or produced the represented page effect.
Results works anywhere. Lab ready = tools offered → found → executable. Requires the ChatGPT in-app browser or Chrome 149+ with WebMCP.
Review the human-approved contractSelection, canonical arguments, approval posture, and page effects remain independently inspectable.
Human + agent workflow
A human and agent draft the intended action, arguments, effects, and semantic boundaries.
Fresh model contexts act through the live WebMCP catalog from a verified fixture.
Trace-derived evidence separates tool choice, arguments, observable state, and over-action.
60-second judge path
In a supported Chrome/WebMCP browser, the public judge lane asks one server-fixed cart question, exposes the model selection, and verifies the returned read through the live native catalog.