Benchmark · Web Search APIs

Reliability research of search tools

Not one of the seven search APIs we tested publishes a deprecation policy with a notice period. No vendor scores above 2 of 4 on documentation and commitment. Separately, we measured retrieval against 1,000 SimpleQA questions: the fastest provider ranked last, and the best retriever was fifth of seven on speed.

Verified
2026-09-03
Vendors
7 measured · 6 scored
Calls
6,998 / 7,000
Questions
1,000
Licence
CC BY 4.0

Reliability scorecard

What each vendor publicly commits to, scored 0–4 on five fixed axes. Six vendors.

Scored disclosures, not measurements. A low score means poor disclosure, not poor practice. No confidence intervals apply. Compliance is the weakest axis — several scores rest on a trust center existing rather than a report being retrievable, so read the source grade on each cell.

Verified 2026-09-03 · valid until 2026-12-02

  1. 1Price stabilityHow much a buyer's cost basis has moved, and how existing customers were treated when it did.
  2. 2Supply independenceWhether the vendor can be cut off or repriced by an upstream it does not control.
  3. 3Documentation and commitmentWhether the vendor has made the operational commitments a production integration depends on.
  4. 4ComplianceWhether the vendor clears the first filter of a standard vendor security review.
  5. 5Procurement maturityWhether an organisation with a procurement function can actually buy this, and how much friction it will meet.

Exa

founding date unresolved

PriceIndependenceDocsComplianceProcurement

Tavily

founding date unresolved

PriceIndependenceDocsComplianceProcurement

Acquired by Nebius, announced 2026-02-10.

Linkup

founded 2024

PriceIndependenceDocsComplianceProcurement

You.com

founding date unresolved

PriceIndependenceDocsComplianceProcurement

Brave Search API

founding date unresolved

PriceIndependenceDocsComplianceProcurement

The search API is a monetisation surface for a browser business rather than the company's core product.

Keenable

public 2 weeks

PriceIndependenceDocsComplianceProcurement

Thin disclosure here reflects company age, not a deliberate choice. See missing_data_policy and score_state.

Price history

Base search rate over the trailing 24 months. Two vendors never moved.

Base search rate per 1,000 requests, 2024-09 to 2026-09 · verified 2026-09-03
$0$2$4$6$82024-092025-032025-092026-032026-09Sources conflict on this change — see the scorecard cellSources conflict on this change — see the scorecard cellSources conflict on this change — see the scorecard cellTavily$8.00no change in 24 monthsExa$7.00Linkup$5.00no change in 24 monthsBrave$5.00You.com$5.00Keenable$4.00
  • A dot is a dated, sourced change
  • A ring marks a change our sources disagree about
  • An open marker means our record begins there, not the price
  • A dashed run means both levels are attested but the change date is not
  • Vendors on the same rate are drawn a hair apart — read the value from the label
6 notes on sourcing

Tavily. Billed per credit: $0.008 per credit is $8.00 per 1,000 credits. The $7.50 in the retrieval table is per 1,000 searches — a different unit, not a conflicting figure. Wayback-derived; verify against the live page.

Exa. The 2026-03-03 changelog announces simplified and lowered pricing and fixes the $7 level on that date; it does not date the rise. Two tracker records place the raise at 2026-04 and 2026-07-29, both after the changelog already quotes $7, so neither can be the transition either. Pending re-derivation from Archive captures.

Brave. The February 2026 event carries two adverse elements, not one: the tier removal and an attribution requirement that did not previously apply. Existing free-plan subscribers at 2,000 queries/month reportedly kept access, so grandfathering is partial rather than absent — that sits between two anchors and is recorded in the scorecard rather than carried by the score.

You.com. The API platform launched 2025-10, so there is no earlier level to plot. Existing pay-as-you-go and grandfathered plans moved automatically to the lower rates, enterprise contracts were honoured through end of term and credits carried over. Launch rates are tracker and Archive sourced and still need verifying.

Keenable. Publicly available roughly two weeks at the verification date. The short line is insufficient history, not a stable record.

Brave, Tavily and Linkup rest on third-party and archive-derived captures rather than a live vendor page, and the scorecard flags all three for direct verification before publication. Nothing between two marked points is measured — a flat segment means no change was recorded, not that one was ruled out.

Retrieval, latency and cost

Seven providers, the same 1,000 SimpleQA questions, k = 10. Latency measured in a separate pass at concurrency 1.

1,000 search tasks

Seedcast Benchmark 01

Web Search Tools

Retrieval accuracy (URL-exact match) across the same 1,000 search tasks.

Higher accuracy is better
Accuracy
70%75%80%85%90%95%
91.1%
90.0%
89.1%
84.8%
84.4%
83.4%
75.9%
ExaExa
Brave SearchBrave
You.comYou.com
LinkupLinkup
SerperSerper
Tavily Tavily
KeenableKeenable
Search provider
Same tasks Same inputs Observed resultsBar color = provider brand
Does the returned list contain the specific page a human researcher cited?
Retrieval, latency and price by provider · prices verified 2026-09-03
ProviderExact source rate95% CIIntervalMean latencyp50p95$ / 1,000
Tier 1 — intervals overlap within this group
Exa91.1%89.2–92.71,705 mspendingpending$7.00
Brave90.0%88–91.7532 mspendingpending$5.00
You.com89.1%87–90.9672 mspendingpending$5.00
Tier 2 — intervals overlap within this group
Linkup84.8%82.4–86.92,209 mspendingpending$5.00
Serperretrieval only84.4%82–86.51,929 mspending4,911 ms$1.00
Tavily83.4%81–85.62,232 mspendingpending$7.50
Tier 3 — intervals overlap within this group
Keenable75.9%73.2–78.4404 mspendingpending$4.00

Serper's rate is third-party sourced, not read from a live vendor page. Treat it as indicative.

2 notes on these figures

Serper's price. Its public pricing page was not reachable and its credit packs sit behind signup. It is also the figure that most changes the comparison, since it would make Serper the cheapest per call by a wide margin.

Pending figures. pending marks a value named in the published plan but not yet in the repository. Per-provider domain rates and the latency distribution arrive with retrieval-summary.json and latency.csv.

Exact source rate is a floor: SimpleQA's reference URLs are not exhaustive, so a provider returning a different but perfectly good source scores zero. At domain level all seven fall within 2.2 points of each other.

Detail, method and caveats

Scope

Two separate studies over the same set of providers, under different evidence regimes.

Retrieval quality — measured

One query in, one ranked list of ten out, scored against SimpleQA's reference source URLs. n = 1,000 per provider, 7,000 calls, 6,998 succeeded. 95% Wilson intervals. Providers whose intervals overlap are not distinguishable at this sample size. Seven providers.

Reliability disclosure — reviewed

What each vendor has publicly committed to across five dimensions. A structured document review, not a measurement, so no confidence intervals apply and none are shown. Six providers.

What we did not measure

Answer synthesis, content extraction, multi-step agentic retrieval, index freshness, non-English or non-US performance, and whether any vendor's compliance controls actually work.

What the shapes show

No vendor is strong everywhere. The largest total area belongs to Linkup at 3-3-2-3-3, which leads no single dimension.

Compliance and independence trade off. Tavily scores highest on compliance and procurement while running no general index of its own. Brave runs one of the few genuinely independent Western indexes and scores 1 on compliance.

Documentation is uniformly weak. The highest score in the field is 2 out of 4. The artifact count that produces it is a straight tally of eight checkable items, making this the most objective axis and the most damning.

Serper was scored and removed from the scorecard. It scored at the floor on all five dimensions, and a single all-zero row compresses the visual range without adding information. Removing an outlier is defensible on presentation grounds; it also narrows the spread across the remaining six, which flatters them. Its full scored row ships in scorecard.json, and it is measured in retrieval above.

Latency

The finding that changed our own numbers.

Our first latency figures came from the main concurrent run. We re-measured with exactly one request in flight and the numbers moved by 1.5× to 7.2× — and not uniformly. The inflation tracked the per-provider concurrency we had assigned to respect each vendor's rate limits, and it reordered the ranking: the provider that looked fastest under load was second slowest in isolation, and the reverse.

In-run timings measure your own worker allocation, not vendor speed. We discarded ours. Anyone running a similar test should.

Method. Separate pass, exactly one request in flight globally, 150 questions, 1,050 sequential calls, providers visited round-robin in shuffled order. Measured from a single location on a single day — relative comparisons under identical conditions, not absolute figures reproducible from another network.

Tails. Means understate the spread. Serper's p95 is 4,911 ms, about 3.5× its median, where the other six sit near 1.5×.

Price

List price per 1,000 search requests, verified 2026-09-03. All seven bill per request.

Checked against observed spend where possible. Exa reports costDollars on every call; measured across 1,000 calls it came to exactly $7.00 per 1,000, matching its advertised rate. Linkup was checked by differencing its credit balance: $5.63 observed against $5.00 advertised, a gap caused by our own retry behaviour rather than a pricing discrepancy. Announced prices held up where they could be checked, so announced prices are what we publish.

One figure is weaker than the others. Serper's public pricing page was not reachable and its credit packs sit behind signup, so its rate comes from third-party sources. It is also the number that most changes the comparison, since it would make Serper the cheapest per call by a wide margin. Treat it as indicative.

The same 24-month window, per vendor

Price and free-tier movement · verified 2026-09-03
VendorRate changesFree-tier changesNet on base searchGrandfathered
Exa1 increase (undated), 2 decreases3, all loosening$5 → $7 (+40%)Auto-applied
Brave1 restructure2, one raise then removalfree → $5/1kPartial
TavilyNo change in 24 months — $0.008/credit throughoutn/a
LinkupNo change in 24 months — €5 → $5 /1k (FX only)n/a
You.comNo adverse change in 24 months — 2 documented decreases, $6.25 → $5/1kYes, explicitly
Keenableno historyno historyno history

The full dated event log ships as pricing-history.json, which is named in the publication plan but not yet in the repository.

Where the two studies meet

Quality and speed are close to uncorrelated. The fastest provider ranks last on retrieval; the best retriever is fifth of seven on speed.

Quality and disclosure are close to uncorrelated too. Brave places second on retrieval at 90.0% and scores 1 on compliance and 1 on documentation. Tavily has the strongest disclosure in the field and places sixth of seven on retrieval.

The best-disclosed vendors are not the most independent. Tavily leads compliance and procurement while operating no general index of its own. Brave and Exa own their indexes and disclose least and second-least on compliance respectively.

Nothing here supports a single ranking. A buyer optimising for retrieval, for continuity, or for passing a security review would choose three different vendors from this set of seven.

What we could not determine

Named gaps, not hedges.

Whether any vendor's compliance controls function.
We checked whether reports exist, not what they contain.
Index provenance beyond vendor statements.
Cross-provider result overlap would settle this from data we already hold. Not yet run.
Freshness.
SimpleQA answers are static by design. Providers are known to diverge sharply between static-fact and time-sensitive performance, and nothing here speaks to that.
Whether SimpleQA appears in any provider's own tuning or evaluation set.
It is public and well known. We cannot detect this.
Serper's true price.
Its public pricing page was not reachable and its credit packs sit behind signup.
Non-English and non-US performance.
Untested.

Data and reproduction

Everything published under CC BY 4.0. Attribution required; commercial use, redistribution and tabular reproduction permitted.

FileContentsStatus
manifest.jsonIndex of all artifacts, versions, dates, schema referencesnot yet written
scorecard.jsonPer-cell reliability scores with full provenance, including Serper's excluded rowpublished
retrieval-summary.jsonPer-provider aggregatesnot yet written
retrieval-per-question.csv7,000 rows — the evidence behind §5not yet written
latency.csvPer-call timings, both passesnot yet written
pricing-history.jsonThe 24-month event lognot yet written
provider-config.jsonLiteral API parameters per providernot yet written
schema/JSON Schema for each filenot yet written

Every provider response and intermediate judgement is stored, so scoring rules can be changed and the study re-scored offline without re-querying any API.

Corrections. Logged with the artifact that prompted them. Original values stay visible. We do not change a score in exchange for anything, including access, data, or advertising.

Publication blockers recorded in scorecard.json

These are unresolved at the verification date and are reproduced here rather than held back.

  • Serper's price_stability and compliance zeros rest on absence of evidence and third-party reporting. Both are load-bearing and must be verified directly.
  • Brave's pricing cells must be verified against the live pricing page, and the February 2026 announcement needs its URL and full ISO date, which the conventions require and the record still omits.
  • Brave's partial grandfathering — existing free-plan subscribers at 2,000 queries/month reportedly retaining access — rests on a third-party post and must be checked against Brave's own announcement. It is the evidence behind the anchor_gap on that cell.
  • Exa's $5 to $7 transition date is not established. Re-derive it from Internet Archive captures of exa.ai/pricing between the 2025 $5 citation and 2026-03-03. Until then the increase is plotted as an undated span.
  • Exa's 2026-07-14 reductions and the ACU unit change are tracker-sourced. Confirm from a primary source before the ACU cell moves from unresolved to substantiated.
  • You.com's 2025-10 launch rates are tracker and Archive sourced and must be verified. Section 12 forbids establishing a change from a tracker.
  • You.com's price_stability carries a conditional: confirm whether the pricing page and the Search API docs disagree on the 100-calls-per-day free quota. A confirmed discrepancy moves the score from 4 to 3.
  • All compliance cells graded vendor_claimed with verification_status pending must be re-attempted against the live trust centers before the compliance dimension is published as scored.
  • Tavily's supply_independence rests on a competitor's comparison page. Requires corroboration or downgrade to not_disclosed.
  • Company founding dates are TODO for Exa, Tavily, You.com and Brave; the age normalisation rule cannot be applied without them.
  • Right of reply has not been conducted.
Version
unresolved
Reliability verified
2026-09-03
Valid until
2026-12-02
Published
unresolved
Attribution
unresolved

Private evaluation

Run the same test on your workflow.

Bring your requirements and your shortlist. We'll run one free evaluation and show you the evidence behind it.

Schedule a demo