Blind inputs
Provider identities were hidden from the judges, so the inputs carried anonymous codes instead of provider labels.
A fair, blind benchmark showing the top scorer for each search type, what each provider costs, and how much confidence the evidence supports. First look only—no switching recommendation yet.
Quality score (0-100, 50 = average, higher = better). Scores are relative within each search type and are not absolute grades.
Statistical ties: in every category, the top scorer's 95% confidence range overlaps the runner-up's. These are observed leads, not proven winners.
No switching recommendation yet: this is a first look, and the small gaps should be read with the confidence ranges above.
How many facts it correctly pulled from the page (higher = better).
Dropped providers: loading…
Compare the Search quality frontier, then scan the published unit price for each provider.
Left is cheaper. Up is better. The strongest value sits toward the top-left.
Linkup isn't plotted: no rankable Search categories (insufficient evidence).
List prices (vendor pay-as-you-go); real bills vary with volume.
Provider identities were hidden from the judges, so the inputs carried anonymous codes instead of provider labels.
grok-4.5 and gpt-5.6-luna judged independently, so no single judge's quirks decide the outcome.
Every score comes with a confidence range; overlapping ranges are reported as statistical ties, not proven wins.