Ask every provider the same questions
Each of the 216 test searches goes to every provider, under identical settings. We save every response exactly as it came back, so the scoring can be repeated later without running the searches again.
The full process, blinding design, score definitions, evidence gates, cost ledger, and limitations behind the snapshot generated from the current report.
Each of the 216 test searches goes to every provider, under identical settings. We save every response exactly as it came back, so the scoring can be repeated later without running the searches again.
Before judging, provider identities are hidden from the judges: names become anonymous codes and the result lists are supplied as blind inputs.
The provider labels stay hidden, while the actual content differences remain visible so the judges can compare result quality.
Each judge sees the search question and all the anonymized result lists side by side, and ranks them from best to worst. Judges can mark lists as tied, flag unusable ones, or decline to rank when they genuinely can't tell. Two judges work independently so no single judge's quirks decide the outcome.
For each search category, a standard statistical method turns all those head-to-head comparisons into a quality score per provider (higher = better), plus a confidence range showing how sure we are. Where there isn't enough data to rank a provider fairly, we say so instead of publishing a shaky number.
A ranking on this page never automatically becomes a “switch your provider” recommendation. Before we'd ever recommend a switch, a top scorer has to prove itself again on a separate, held-back set of test searches, by a clear margin, with both judges agreeing. In this release, nothing has cleared that bar — so this is a first look, not a final verdict.
Provider comparisons need blind inputs, uncertainty ranges, and test questions that aren't selected by a provider's marketing team. SPB is designed around those safeguards.
Provider identities were hidden from the judges (blind inputs), so the inputs carried anonymous codes instead of provider labels.
Every score comes with a confidence range showing how sure we are. That supports the honest language: observed lead or statistical tie, without treating noise as a decisive result.
A provider that's great for product research but weak on breaking news isn't “good” or “bad” — it's just being used for the wrong searches. Ranking each search category separately makes exactly that visible.
Every run is versioned and saved, and every step can be replayed and audited. If you disagree with a conclusion, you can re-run the same release and see for yourself. Trust here isn't “believe our chart” — it's “here's the method, here's how sure we are, come check our work.”
SPB’s workflow is fully reproducible: collect the same public corpora, build blinded inputs, run a two-judge panel, and regenerate the reports. A live rerun reproduces the method—not identical result bytes—because the web, provider indexes, APIs, and hosted models change.
search: run-matrix → build-judge-inputs → run-judge-v2 → build-results-report
extract: run-extract → score-extract
Full guide in the repository — code coming soon.
These warnings and gate definitions are read directly from the current report JSON.
Per-intent results must be read with the language and source buckets. The current case counts are shown without smoothing or reinterpretation.
| Search type | Language buckets | Source buckets |
|---|
Loading…
Loading…
Loading recorded Search attempts…
| Provider | Calls | Successful | Errors | Coverage |
|---|
Two separate benchmark lanes, two separate cost ledgers. Every row is calculated from this run's JSON, not a generic pricing estimate.
Estimated cost per call × recorded calls, sorted by run total.
| Provider | Cost / call | Calls | Run total |
|---|
List price × extracts, sorted by run total.
| Provider | Cost / extract | Extracts | Run total |
|---|