Methodology

How SPB turns raw results into careful claims.

The full process, blinding design, score definitions, evidence gates, cost ledger, and limitations behind the snapshot generated from the current report.

How it works

Six steps from raw search results to numbers you can trust.

01

Ask every provider the same questions

Each of the 216 test searches goes to every provider, under identical settings. We save every response exactly as it came back, so the scoring can be repeated later without running the searches again.

02

Remove the labels

Before judging, provider identities are hidden from the judges: names become anonymous codes and the result lists are supplied as blind inputs.

03

Keep the content available for judging

The provider labels stay hidden, while the actual content differences remain visible so the judges can compare result quality.

04

Two independent AI judges compare the results

Each judge sees the search question and all the anonymized result lists side by side, and ranks them from best to worst. Judges can mark lists as tied, flag unusable ones, or decline to rank when they genuinely can't tell. Two judges work independently so no single judge's quirks decide the outcome.

05

Turn the rankings into quality scores

For each search category, a standard statistical method turns all those head-to-head comparisons into a quality score per provider (higher = better), plus a confidence range showing how sure we are. Where there isn't enough data to rank a provider fairly, we say so instead of publishing a shaky number.

06

Keep rankings and recommendations separate

A ranking on this page never automatically becomes a “switch your provider” recommendation. Before we'd ever recommend a switch, a top scorer has to prove itself again on a separate, held-back set of test searches, by a clear margin, with both judges agreeing. In this release, nothing has cleared that bar — so this is a first look, not a final verdict.

Reading the results

What the numbers mean — in plain words.

quality score
A 0-100 relative scale where 50 = average and higher = better. It tells you how a provider stacks up against the others within the same search category; it is not an absolute grade, and scores from different categories can't be compared directly.
confidence range
How sure we are about a score. A narrow range means we're fairly confident; a wide range means less certainty. We only call one provider genuinely better than another when their ranges don't overlap — otherwise, small differences may not be meaningful.
published vs. withheld
A score is published when we have enough data behind it. When there isn't enough data to rank a provider fairly, we withhold the number entirely rather than show a guess. Even a published score doesn't mean every gap between providers is real — check the confidence ranges.
Trust story

Why these numbers hold up.

Provider comparisons need blind inputs, uncertainty ranges, and test questions that aren't selected by a provider's marketing team. SPB is designed around those safeguards.

Blind by design

Provider identities were hidden from the judges (blind inputs), so the inputs carried anonymous codes instead of provider labels.

Numbers that refuse to overclaim

Every score comes with a confidence range showing how sure we are. That supports the honest language: observed lead or statistical tie, without treating noise as a decisive result.

Judged by search type

A provider that's great for product research but weak on breaking news isn't “good” or “bad” — it's just being used for the wrong searches. Ranking each search category separately makes exactly that visible.

Check it yourself

Every run is versioned and saved, and every step can be replayed and audited. If you disagree with a conclusion, you can re-run the same release and see for yourself. Trust here isn't “believe our chart” — it's “here's the method, here's how sure we are, come check our work.”

Reproducibility

Run it yourself.

SPB’s workflow is fully reproducible: collect the same public corpora, build blinded inputs, run a two-judge panel, and regenerate the reports. A live rerun reproduces the method—not identical result bytes—because the web, provider indexes, APIs, and hosted models change.

search:  run-matrix → build-judge-inputs → run-judge-v2 → build-results-report
extract: run-extract → score-extract

Full guide in the repository — code coming soon.

Claim boundaries

Warnings, confounding, and evidence gates.

These warnings and gate definitions are read directly from the current report JSON.

  1. Loading claim warnings…

Language and source confounding

Per-intent results must be read with the language and source buckets. The current case counts are shown without smoothing or reinterpretation.

Search typeLanguage bucketsSource buckets

Release evidence

  • Loading evidence counts…
RANKING PUBLICATION GATE

Loading…

ROUTING DECISION GATE

Loading…

Retries & coverage

The call ledger keeps failed and extra attempts visible.

Loading recorded Search attempts…

ProviderCallsSuccessfulErrorsCoverage
Cost methodology estimate

Estimated list price, line by line.

Two separate benchmark lanes, two separate cost ledgers. Every row is calculated from this run's JSON, not a generic pricing estimate.

Search lane · list price
Extract lane · list price
Combined · list price

Search lane

Estimated cost per call × recorded calls, sorted by run total.

ProviderCost / callCallsRun total

Extract lane

List price × extracts, sorted by run total.

ProviderCost / extractExtractsRun total
Honest limits

What SPB does not claim.

  • We don't claim any provider is best everywhere — only how they performed on these specific test searches and categories.
  • AI judges aren't perfectly objective. They're a measuring tool with known flaws, which we manage rather than deny.
  • The web changes. Providers update their indexes, pages change, judges get updated — running the benchmark again later is a new measurement, and results may shift.
  • Provider identities were hidden from the judges (blind inputs); this release does not report a verified identity-guessing result.
  • These results are evidence for this snapshot in time — not a guarantee for your specific use.
The methodology JSON could not be loaded. Serve this directory as a static site and refresh.