Exa
founding date unresolved
Benchmark · Web Search APIs
Not one of the seven search APIs we tested publishes a deprecation policy with a notice period. No vendor scores above 2 of 4 on documentation and commitment. Separately, we measured retrieval against 1,000 SimpleQA questions: the fastest provider ranked last, and the best retriever was fifth of seven on speed.
Retrieval methodology — page pendingScorecard methodology — page pendingData and reproduction
What each vendor publicly commits to, scored 0–4 on five fixed axes. Six vendors.
Verified 2026-09-03 · valid until 2026-12-02
founding date unresolved
founding date unresolved
Acquired by Nebius, announced 2026-02-10.
founded 2024
founding date unresolved
founding date unresolved
The search API is a monetisation surface for a browser business rather than the company's core product.
public 2 weeks
Thin disclosure here reflects company age, not a deliberate choice. See missing_data_policy and score_state.
Base search rate over the trailing 24 months. Two vendors never moved.
Tavily. Billed per credit: $0.008 per credit is $8.00 per 1,000 credits. The $7.50 in the retrieval table is per 1,000 searches — a different unit, not a conflicting figure. Wayback-derived; verify against the live page.
Exa. The 2026-03-03 changelog announces simplified and lowered pricing and fixes the $7 level on that date; it does not date the rise. Two tracker records place the raise at 2026-04 and 2026-07-29, both after the changelog already quotes $7, so neither can be the transition either. Pending re-derivation from Archive captures.
Brave. The February 2026 event carries two adverse elements, not one: the tier removal and an attribution requirement that did not previously apply. Existing free-plan subscribers at 2,000 queries/month reportedly kept access, so grandfathering is partial rather than absent — that sits between two anchors and is recorded in the scorecard rather than carried by the score.
You.com. The API platform launched 2025-10, so there is no earlier level to plot. Existing pay-as-you-go and grandfathered plans moved automatically to the lower rates, enterprise contracts were honoured through end of term and credits carried over. Launch rates are tracker and Archive sourced and still need verifying.
Keenable. Publicly available roughly two weeks at the verification date. The short line is insufficient history, not a stable record.
Brave, Tavily and Linkup rest on third-party and archive-derived captures rather than a live vendor page, and the scorecard flags all three for direct verification before publication. Nothing between two marked points is measured — a flat segment means no change was recorded, not that one was ruled out.
Seven providers, the same 1,000 SimpleQA questions, k = 10. Latency measured in a separate pass at concurrency 1.
Seedcast Benchmark 01
Retrieval accuracy (URL-exact match) across the same 1,000 search tasks.
| Provider | Exact source rate | 95% CI | Interval | Mean latency | p50 | p95 | $ / 1,000 |
|---|---|---|---|---|---|---|---|
| Tier 1 — intervals overlap within this group | |||||||
| Exa | 91.1% | 89.2–92.7 | 1,705 ms | pending | pending | $7.00 | |
| Brave | 90.0% | 88–91.7 | 532 ms | pending | pending | $5.00 | |
| You.com | 89.1% | 87–90.9 | 672 ms | pending | pending | $5.00 | |
| Tier 2 — intervals overlap within this group | |||||||
| Linkup | 84.8% | 82.4–86.9 | 2,209 ms | pending | pending | $5.00 | |
| Serperretrieval only | 84.4% | 82–86.5 | 1,929 ms | pending | 4,911 ms | $1.00⚠ | |
| Tavily | 83.4% | 81–85.6 | 2,232 ms | pending | pending | $7.50 | |
| Tier 3 — intervals overlap within this group | |||||||
| Keenable | 75.9% | 73.2–78.4 | 404 ms | pending | pending | $4.00 | |
⚠ Serper's rate is third-party sourced, not read from a live vendor page. Treat it as indicative.
Serper's price. Its public pricing page was not reachable and its credit packs sit behind signup. It is also the figure that most changes the comparison, since it would make Serper the cheapest per call by a wide margin.
Pending figures. pending marks a value named in the published plan but not yet in the repository. Per-provider domain rates and the latency distribution arrive with retrieval-summary.json and latency.csv.
Exact source rate is a floor: SimpleQA's reference URLs are not exhaustive, so a provider returning a different but perfectly good source scores zero. At domain level all seven fall within 2.2 points of each other.
Two separate studies over the same set of providers, under different evidence regimes.
One query in, one ranked list of ten out, scored against SimpleQA's reference source URLs. n = 1,000 per provider, 7,000 calls, 6,998 succeeded. 95% Wilson intervals. Providers whose intervals overlap are not distinguishable at this sample size. Seven providers.
What each vendor has publicly committed to across five dimensions. A structured document review, not a measurement, so no confidence intervals apply and none are shown. Six providers.
Answer synthesis, content extraction, multi-step agentic retrieval, index freshness, non-English or non-US performance, and whether any vendor's compliance controls actually work.
No vendor is strong everywhere. The largest total area belongs to Linkup at 3-3-2-3-3, which leads no single dimension.
Compliance and independence trade off. Tavily scores highest on compliance and procurement while running no general index of its own. Brave runs one of the few genuinely independent Western indexes and scores 1 on compliance.
Documentation is uniformly weak. The highest score in the field is 2 out of 4. The artifact count that produces it is a straight tally of eight checkable items, making this the most objective axis and the most damning.
Serper was scored and removed from the scorecard. It scored at the floor on all five dimensions, and a single all-zero row compresses the visual range without adding information. Removing an outlier is defensible on presentation grounds; it also narrows the spread across the remaining six, which flatters them. Its full scored row ships in scorecard.json, and it is measured in retrieval above.
The finding that changed our own numbers.
Our first latency figures came from the main concurrent run. We re-measured with exactly one request in flight and the numbers moved by 1.5× to 7.2× — and not uniformly. The inflation tracked the per-provider concurrency we had assigned to respect each vendor's rate limits, and it reordered the ranking: the provider that looked fastest under load was second slowest in isolation, and the reverse.
In-run timings measure your own worker allocation, not vendor speed. We discarded ours. Anyone running a similar test should.
Method. Separate pass, exactly one request in flight globally, 150 questions, 1,050 sequential calls, providers visited round-robin in shuffled order. Measured from a single location on a single day — relative comparisons under identical conditions, not absolute figures reproducible from another network.
Tails. Means understate the spread. Serper's p95 is 4,911 ms, about 3.5× its median, where the other six sit near 1.5×.
List price per 1,000 search requests, verified 2026-09-03. All seven bill per request.
Checked against observed spend where possible. Exa reports costDollars on every call; measured across 1,000 calls it came to exactly $7.00 per 1,000, matching its advertised rate. Linkup was checked by differencing its credit balance: $5.63 observed against $5.00 advertised, a gap caused by our own retry behaviour rather than a pricing discrepancy. Announced prices held up where they could be checked, so announced prices are what we publish.
One figure is weaker than the others. Serper's public pricing page was not reachable and its credit packs sit behind signup, so its rate comes from third-party sources. It is also the number that most changes the comparison, since it would make Serper the cheapest per call by a wide margin. Treat it as indicative.
| Vendor | Rate changes | Free-tier changes | Net on base search | Grandfathered |
|---|---|---|---|---|
| Exa | 1 increase (undated), 2 decreases | 3, all loosening | $5 → $7 (+40%) | Auto-applied |
| Brave | 1 restructure | 2, one raise then removal | free → $5/1k | Partial |
| Tavily | No change in 24 months — $0.008/credit throughout | n/a | ||
| Linkup | No change in 24 months — €5 → $5 /1k (FX only) | n/a | ||
| You.com | No adverse change in 24 months — 2 documented decreases, $6.25 → $5/1k | Yes, explicitly | ||
| Keenable | no history | no history | no history | — |
The full dated event log ships as pricing-history.json, which is named in the publication plan but not yet in the repository.
Quality and speed are close to uncorrelated. The fastest provider ranks last on retrieval; the best retriever is fifth of seven on speed.
Quality and disclosure are close to uncorrelated too. Brave places second on retrieval at 90.0% and scores 1 on compliance and 1 on documentation. Tavily has the strongest disclosure in the field and places sixth of seven on retrieval.
The best-disclosed vendors are not the most independent. Tavily leads compliance and procurement while operating no general index of its own. Brave and Exa own their indexes and disclose least and second-least on compliance respectively.
Nothing here supports a single ranking. A buyer optimising for retrieval, for continuity, or for passing a security review would choose three different vendors from this set of seven.
Named gaps, not hedges.
Everything published under CC BY 4.0. Attribution required; commercial use, redistribution and tabular reproduction permitted.
| File | Contents | Status |
|---|---|---|
manifest.json | Index of all artifacts, versions, dates, schema references | not yet written |
scorecard.json | Per-cell reliability scores with full provenance, including Serper's excluded row | published |
retrieval-summary.json | Per-provider aggregates | not yet written |
retrieval-per-question.csv | 7,000 rows — the evidence behind §5 | not yet written |
latency.csv | Per-call timings, both passes | not yet written |
pricing-history.json | The 24-month event log | not yet written |
provider-config.json | Literal API parameters per provider | not yet written |
schema/ | JSON Schema for each file | not yet written |
Every provider response and intermediate judgement is stored, so scoring rules can be changed and the study re-scored offline without re-querying any API.
Corrections. Logged with the artifact that prompted them. Original values stay visible. We do not change a score in exchange for anything, including access, data, or advertising.
scorecard.jsonThese are unresolved at the verification date and are reproduced here rather than held back.
Private evaluation
Bring your requirements and your shortlist. We'll run one free evaluation and show you the evidence behind it.
Schedule a demo