The measurement problem
A successful tool call can still be the wrong behavior.
Ordinary tests can prove that a handler works. These papers explain why evaluating a semantic system requires another question: does its behavior remain correct when the same intended meaning appears through valid variation?
Published foundation
Two principles. One measurable WebMCP contract.
Paper 01Invarra Research · June 28, 2026
When intent cannot be observed directly, one correct response is weak evidence. Stable behavior across meaning-preserving representations is stronger evidence that a system is tracking the underlying phenomenon—not merely its wording.
In ThurstoneThis is why Thurstone preserves one intended behavior while testing representative ways a person may express it.
LIPRepresentation processr = g(Φ, c, ε)
A request expresses latent meaning through context and surface variation.Observed behaviorB(r)
The evaluator sees what the system does with that representation.Hold meaning fixed. Change the representation. Observe whether behavior drifts.
Paper 02Invarra Research · June 28, 2026
CSR separates the meaning being tested, the language used to express it, and the outcome produced by the system. Meaning becomes the experimental unit; each valid realization becomes another measurement.
In ThurstoneThis becomes Thurstone’s contract structure: declare the meaning once, test representative requests, and compare every tool decision and site effect with the same expectation.
CSRRealization relationp = π(s, c)
A prompt realizes one canonical semantic unit under a condition.Semantic preservationp₁ ≡ₛₑₘ p₂
Valid realizations preserve the same meaning-bearing commitments.Meaning is the unit. Wording is controlled variation. Outcomes are evidence.
From research to product
Thurstone turns the measurement layers into a test.
Research languageCanonical meanings
→Owner inputContract + requestsE(s)
→Observed systemTool + argumentsB(r)
→Trusted realityEffect + verdictexpected ≟ actual