Aibiguity
The same number can hide different reasoning
Two answers can occupy the same position on a scale and still deserve a closer look. Aibiguity makes both the number and the explanation available for comparison.
My contribution. I designed the judgment and comparison flow: your answer comes before the model reveal, and your original choice remains visible on the same scale.
One question asks, “How much does intention matter compared with outcome?” The interface represents your answer as a percentage assigned to intention.

The saved entry labeled GPT-5.5 Thinking chooses 40% intention. Claude also chooses 40%, with a different written explanation:
GPT-5.5 Thinking · 40%
“Outcomes matter more, but intention shapes blame and trust.”
Claude · 40%
“Intentions ground moral character; outcomes determine actual harm.”


These explanations overlap in substance but differ in emphasis. They do not establish contradictory conclusions or reveal the models’ internal reasoning. They show why a shared number is an incomplete comparison: it leaves open which distinctions matter and how the stated principle would apply to a particular situation.
Agreement is a position on the scale
“Closest match” means the smallest numerical distance from your choice within this saved sample. It does not make that model’s answer correct or establish that you reasoned in the same way. Tied answers remain tied.
Asking for your choice first keeps an initial answer separate from the reveal. I have not measured whether this sequence changes people’s responses. The interface makes the sequence visible without requiring agreement to count as success.
The question has a frame, too
The scale runs from zero to one hundred and starts at fifty. The wording specifies no concrete situation. Turning a broad judgment into a percentage makes comparison convenient, while leaving substantial meaning outside the number.
The preview uses the full six-question Quick Start quiz with seven saved model answers per question. Your answers stay local to the demo. It makes no new model request and provides no human population statistics. These examples do not establish model accuracy, consistency, or performance.
A useful next test would vary one question’s wording and compare both the numbers and explanations. I would document what changes before drawing conclusions about whether the difference comes from interpretation or framing.
Quotes and values come from the project’s saved Quick Start data, reproduced in the portfolio. These are stored explanations, not raw response transcripts. The specific question’s collection date is not recorded in that data. Interface frames captured October 6, 2026.