Blog/Model Comparison/

Best Search API for Research Agents: 6 Tools Tested (2026)

Best search API for research agents, measured: Exa, Tavily, Firecrawl, Cloudsway, Scholar and Exa Answer with GPT-6.1 Sol on 20 fresh questions, 3 runs each.

Research agent search API benchmark 2026 cover; editorial artwork, not a test result

Our first Exa run scored 55 out of 60, and Exa wasn’t the problem. In 44 of its 91 searches, the agent got an empty result list. The SandBase gateway had sent the same Exa search back in a second response shape, with the hits under entities instead of results, and our harness only read results. Once the reader handled both shapes, Exa scored 60/60. Nothing else changed.

That’s the clearest result from this search API benchmark for research agents. Tools are the agent’s hands. With the model held fixed, what separated them wasn’t whether the answers existed somewhere on the web. It was whether the agent could reliably read what came back, whether it could open the page, and how much reading cost. We gave GPT-6.1 Sol one search tool at a time on SandBase and asked 20 questions about facts from July to October 2026, three runs each, on 2026-10-03 (UTC).

Key takeaway

  • With the reader fixed, four setups answered all 60 runs correctly: Exa (search + contents), Firecrawl (search + scrape), Scholar web search and Exa Answer. Tavily scored 58/60 and Cloudsway 55/60. With no tool, the model answered 1 of 51 answerable runs.
  • Valid citations, meaning the cited URL loaded and contained the answer: 51/51 for Exa, Tavily and Scholar, 50 for Exa Answer, 48 for Cloudsway and 46 for Firecrawl.
  • Median LLM cost per question ranged from $0.0023 (Exa Answer) and $0.0026 (Scholar) to $0.0121 (Tavily). Most of the gap came from tokens spent reading fetched pages.
  • All nine search and fetch endpoints showed a base price of $0 on 2026-10-03, and 27 sampled run ids were billed $0.000000.
  • Scope: 20 short-answer questions, one model, N=3, one day. Accuracy was saturated for the top four setups, so this test can’t rank them against each other.

For the feature-level comparison of these APIs, read best web search APIs for AI agents and Exa vs Tavily vs Firecrawl vs SerpAPI. This post measures the same tools inside a real agent loop.

Setup: one model, one search tool per run

The LLM was openai/gpt-6.1-sol through SandBase POST /v1/chat/completions, with a 2,000-token output cap per turn and at most six turns per question. We picked it because it scored 51/51 on our 12-model tool-calling benchmark at a moderate cost. Keeping the model fixed means any difference comes from the tool.

Each run got exactly one vendor’s tools, under the same tool names and descriptions:

Setupweb_search callsfetch_page callsSandBase base price (2026-10-03)
Exaexa/search (5 results, highlights)exa/contents$0 / $0
Tavilytavily/search (5 results)tavily/extract$0 / $0
Firecrawlfirecrawl/search (limit 5)firecrawl/scrape (markdown)$0 / $0
Cloudswaycloudsway/search (5 results)none$0
Scholarscholar/search-web (5 results)none$0
Exa Answeranswer_engine tool backed by exa/answernone$0
No toolnonenonen/a

The harness trimmed every search result to title, URL and a 500-character snippet, and every fetched page to its first 10,000 characters, whichever vendor returned it. Cloudsway and Scholar have no fetch endpoint on SandBase, so those agents had snippets only. Exa Answer, the answer-engine baseline, returns a written answer plus the URLs it cites.

SandBase model page for exa/search showing Exa Search with Free base price, sync execution, api model type and 11 input fields, under an API Free Week banner

Caption: The exa/search model page lists a Free base price and sync execution; the banner above it reads API Free Week, so treat $0 as the price on the capture date, 2026-10-03.

The price column comes from GET https://api.sandbase.ai/v1/models/<id>, where all nine endpoints had base_price: "0". The model pages show the same thing, but the site banner said “API Free Week” that day, so don’t assume search stays free. GPT-6.1 Sol was $2 per million input tokens and $10 per million output tokens.

SandBase language model catalog searched for gpt-6.1-sol, showing OpenAI GPT-6.1 Sol with 1.1M context, $2 input and $10 output per million tokens

Caption: The SandBase catalog row for openai/gpt-6.1-sol shows $2 input and $10 output per million tokens, the rates behind every cost figure in this post (captured 2026-10-03).

The 20 questions

Every question has a short answer that we checked against a primary source fetched on 2026-10-03, with the URL and a quote of 25 words or fewer saved in ground_truth.json. Most are facts from August to October 2026. Three are false premises, where the right answer is null. Four are in Chinese, and their sources are Chinese-language official pages.

TypeCountExamples (answer)
Dates (release, publish, end of life)9Python 3.14.8 (2026-09-30), Go 1.27 (2026-08-19), Rust 1.99.0 (2026-10-01), Firefox 157 (2026-09-29)
Names and versions3Kubernetes v1.37 codename (Garhwal); Next.js September security release (16.3.8); DeepSeek V4.1 Flash API model name (deepseek-flash)
Numbers from announcements and pricing5Kubernetes v1.37 enhancements (67); iPhone 18 Pro US price ($1,199) and China price (RMB 9,999); China’s August CPI (0.8%) and industrial output (5.2%)
False premise3PostgreSQL 19 GA date (GA is scheduled for Oct 29); price of a 4TB iPhone 18 Pro (top tier is 2TB); Next.js 17 release date (not released)

Some questions have traps built in. PEP 745 still lists Python 3.14.8 for October 6, but it shipped early, on September 30, as an expedited security release. The PostgreSQL mailing list gives a scheduled GA date, a trap for agents that read “scheduled” as “released”.

The system prompt, verbatim:

You are a research agent. Today's date is 2026-10-03. Answer the user's question with the tools you have; if you have no tools, answer from your own knowledge. Prefer primary sources such as official sites, release notes, docs and press releases. If the thing asked about does not exist, has not happened yet, or you cannot find it, answer null. End with exactly one line: FINAL: {"answer": <value or null>, "source_url": <the URL that states the answer, or null>}

Grading: answer and citation, scored separately

Answer. The grader compares the answer in the FINAL JSON to the gold answer after normalizing by type:

  • Dates are parsed to YYYY-MM-DD, so “Sept. 30, 2026”, “30 September 2026” and 2026年9月30日 (the Chinese date format) all match.
  • Numbers drop $, ¥, RMB, USD, 元, %, commas and spaces, then are compared as floats.
  • Names are compared case-insensitively.
  • Versions drop a leading v or Next.js.
  • A false-premise question counts as correct only if the answer is null or says “not found” or “not released”. Giving a date or price is wrong.

A missing FINAL line is a miss.

Citation. For the 17 answerable questions, the grader fetched source_url with a plain HTTP GET and a browser User-Agent. The citation is valid if the response was HTTP 200 and the page text, with whitespace removed, contained a surface form of the gold answer. For dates that means several written formats. For short numbers we add the unit or context (“67 enhancements”, “$1,199”, “0.8%”) so a stray digit doesn’t count. Each URL was fetched once on 2026-10-03.

Results

420 scored runs: 7 setups × 20 questions × 3 runs, between 10:33 and 10:55 UTC. The Exa row is the rerun after the reader fix, described in the next section.

SetupCorrect (60)Answerable (51)False premise (9)Valid citations (51)Tool callsMedian LLM cost / questionMedian timep90 time
Exa (search + contents)6051951120$0.007714.1 s18.6 s
Firecrawl (search + scrape)6051946121$0.009713.0 s17.0 s
Scholar web search605195195$0.002610.2 s14.8 s
Exa Answer605195093$0.002312.0 s17.9 s
Tavily (search + extract)5851751152$0.012113.9 s19.3 s
Cloudsway5546948175$0.005814.3 s31.8 s
No tool101910$0.00115.0 s7.7 s

Per-question stability was high. Exa, Firecrawl, Scholar and Exa Answer got all 20 questions right in every run. Tavily missed one question in two of its three runs. Cloudsway had two questions at 2/3 and one at 0/3. Two of Cloudsway’s 48 valid citations came with a wrong answer: the cited page held the right date, but the agent wrote a different one.

The no-tool row shows these facts are outside the model’s memory. GPT-6.1 Sol answered null in 50 of 51 answerable runs, which is the honest response, and in all 9 false-premise runs. Its one correct answer was Rust 1.99.0 on 2026-10-01, given once out of three runs, together with the correct blog URL. That fits Rust’s six-week schedule, so it was probably an informed guess.

Median time per search call ran from 1.65 s (Firecrawl) to 2.48 s (Tavily). Exa Answer took 1.91 s per call. Fetch calls took about 1.0 to 1.6 s each.

Where things failed

The response shape changed between calls (Exa, first pass). The SandBase API reference for these endpoints documents the payload at outputs[0].data. In our runs it never arrived there. Firecrawl search, Scholar, Exa Answer and all three fetch endpoints came back as a top-level output object. Cloudsway came back as outputs[0] with no data key. Tavily search and Exa search alternated between the two:

  • tavily/search: 40 calls with output, 33 with outputs[0].
  • exa/search (rerun): 37 with output, 30 with outputs[0].

For Tavily the inner payload was the same either way. For Exa it wasn’t: the output variant carried Exa’s native results list, and the outputs[0] variant carried a normalized entities list. Our first harness read only results, so 44 of Exa’s 91 searches looked empty. The agent answered null in five runs, and Exa scored 55/60. With a reader that accepts both, the rerun scored 60/60 on 67 searches. If your agent seems to “find nothing” on some runs and not others, log which response shape each call returned before you blame the search engine.

SandBase API reference for tavily/search showing POST /v1/api/tavily/search, the query, search_depth, max_results and topic parameters, and a completed response example with the payload under outputs and data

Caption: The tavily/search reference documents POST /v1/api/tavily/search and a completed response with outputs[0].data; on 2026-10-03 the live endpoint returned top-level output or outputs[0] without data instead.

No fetch tool means trusting snippets (Cloudsway). Cloudsway lost all three runs of the DeepSeek question. Each run spent its six turns on site: searches over DeepSeek’s docs and never surfaced the sentence that names deepseek-flash. It also answered Python 3.14.8’s date as 2026-10-01 once (the Python Insider announcement is dated October 1), and gave a July date for the Next.js 16.3 post once (that post says a preview shipped the month before). Both look like a date pulled from a snippet with no way to open the page, though we didn’t log the snippets to prove it. Scholar had no fetch tool either and still scored 60/60, so snippet quality matters as much as having a fetch tool.

Searching for something that doesn’t exist (Tavily). On the 4TB iPhone question, two of three Tavily runs kept searching and extracting pages until the six-turn limit, without writing a FINAL line. The most expensive of those cost $0.078, the highest of all 420 runs. The third run answered null. A turn budget plus an explicit “answer null if you can’t confirm it” doesn’t fully stop an agent from hunting for a price that isn’t there.

Citations a plain GET can’t verify (Firecrawl, Exa Answer). All five Firecrawl citation misses came with correct answers:

  • Twice it cited blog.rust-lang.org/releases/latest/, a redirect stub that holds only a meta refresh to the 1.99.0 post.
  • Three times it cited Apple China’s iPhone 18 Pro product page, where the price isn’t in the static HTML.

Exa Answer’s one miss was the Rust releases index. A browser would get there; a citation-checking script wouldn’t. For an audit trail, ask for the final article URL, not a landing page.

Transient fetch errors. Three exa/contents calls returned HTTP 503 with an upstream “must provide ids” message, even though the request body did contain ids. The same body worked elsewhere in the run. The agent fell back on the search snippets and still answered correctly.

Cost and billing

The 420 scored runs used 1,368,972 input and 47,195 output tokens, $3.21 at list price. Adding the first Exa pass ($0.54), a 28-run pilot ($0.28) and the code sample below, the LLM spend was about $4.04. Six screenshot captures added about $0.03. Search was free on the day: every endpoint showed a base price of $0, and we checked 27 run ids (three per endpoint) with GET /v1/tasks/<id>/cost, which returned "cost": "0.000000" and settled: true each time. Both GET /v1/tasks/<id>/cost and GET /v1/models/<id> need the same Authorization: Bearer key as the calls; without one, the public model pages on sandbase.ai show the same prices. Exa’s own payload includes a costDollars field, for example 0.007 for one search. That’s Exa’s upstream figure, not what SandBase billed us.

So the cost differences in the table are LLM tokens, and nearly all of them are input. Median output was 80 to 115 tokens per question in every setup with a tool. Median input ranged from 711 tokens (Exa Answer) and 919 (Scholar) to 4,334 (Firecrawl) and 5,493 (Tavily). Fetching pages is what makes a run expensive. Tavily’s agent fetched a page 79 times in 60 runs, Firecrawl’s 58 and Exa’s 53. At the medians, 1,000 questions like these per day would cost about $2.30 with Exa Answer, $2.60 with Scholar, $7.70 with Exa and $12.10 with Tavily in LLM tokens, plus whatever search costs after the free week.

How to choose

If you needStart withBut
Cheapest correct answers on short factual lookupsScholar web search: 60/60, $0.0026 per question, fastest medianNo fetch tool; you depend on Google-style snippets containing the fact
The vendor to read pages for youExa Answer: 60/60, $0.0023One citation pointed at an index page; you see less of the evidence
Full page text for deeper reading or extractionExa (search + contents): 60/60, all 51 citations validHandle both response shapes; it costs about 3x Scholar here
Clean markdown from JS-heavy pagesFirecrawl (search + scrape): 60/605 citations a plain GET couldn’t verify (redirect stub, client-side price)
Tavily’s search plus extract pairTavily: all 51 answerable runs correct, all citations validHighest cost; two runs ran out of turns on the false-premise price
Date filters and localized resultsCloudsway, which supports start_date, end_date and country (we didn’t use them)55/60 here with snippets only; slowest p90 (31.8 s)

This isn’t a ranking of search quality. The top four setups were tied on accuracy, and 20 questions with three runs can’t separate them. So the best search API for research agents depends on the row that matches your workload, not on a single score.

A one-tool research agent you can run

The program below is the Scholar setup, trimmed down: one search tool, GPT-6.1 Sol, a FINAL line and a citation check. It calls POST https://api.sandbase.ai/v1/api/scholar/search-web and reads the payload from the documented outputs[0].data first. Because the live endpoints returned other shapes on 2026-10-03, it also accepts outputs[0] and a top-level output, and it raises on any run that isn’t completed. The fields results[].title, link and snippet were observed in our responses, not documented guarantees.

import os
import re
import json
import requests

BASE = "https://api.sandbase.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['SANDBASE_API_KEY']}"}
MODEL = "openai/gpt-6.1-sol"   # $2 / $10 per million tokens on SandBase (2026-10-03)
SYSTEM = ("You are a research agent. Today's date is 2026-10-03. Answer with the tools you have. "
          "Prefer primary sources such as official sites, release notes, docs and press releases. "
          "If the thing asked about does not exist, has not happened yet, or you cannot find it, answer null. "
          'End with exactly one line: FINAL: {"answer": <value or null>, "source_url": <URL or null>}')


def read_payload(body: dict):
    """Return the business payload of a completed SandBase run.
    The API reference documents outputs[0].data. On 2026-10-03 the search endpoints we called also
    returned it as outputs[0] itself or as a top-level `output` object, so accept all three."""
    if body.get("status") != "completed":
        raise RuntimeError(f"run {body.get('id')} not completed: {body.get('status')} {body.get('error')}")
    outputs = body.get("outputs")
    if outputs:
        first = outputs[0]
        return first["data"] if "data" in first else first
    if "output" in body:
        return body["output"]
    raise RuntimeError(f"run {body.get('id')} has no payload")


def web_search(query: str) -> dict:
    """One search tool: SandBase scholar/search-web, trimmed to 5 results."""
    resp = requests.post(f"{BASE}/api/scholar/search-web", headers=HEADERS,
                         json={"query": query, "max_num_results": 5}, timeout=120)
    resp.raise_for_status()
    data = read_payload(resp.json())
    results = data.get("results") or []          # observed field names: results[].title/link/snippet
    return {"results": [{"title": r.get("title"), "url": r.get("link"),
                         "snippet": (r.get("snippet") or "")[:500]} for r in results]}


TOOLS = [{"type": "function", "function": {
    "name": "web_search",
    "description": "Search the web. Returns up to 5 results with title, url and a short snippet.",
    "parameters": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]}}}]


def ask(question: str, max_turns: int = 6) -> dict:
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": question}]
    tokens_in = tokens_out = searches = 0
    for _ in range(max_turns):
        resp = requests.post(f"{BASE}/chat/completions", headers=HEADERS, timeout=240,
                             json={"model": MODEL, "max_tokens": 2000,
                                   "tools": TOOLS, "messages": messages})
        resp.raise_for_status()
        body = resp.json()
        tokens_in += body["usage"]["prompt_tokens"]
        tokens_out += body["usage"]["completion_tokens"]
        msg = body["choices"][0]["message"]
        messages.append({k: v for k, v in msg.items() if k in ("role", "content", "tool_calls")})
        if not msg.get("tool_calls"):
            match = re.search(r"FINAL:\s*(\{.*\})", msg.get("content") or "", re.S)
            final = json.loads(match.group(1)) if match else {"answer": None, "source_url": None}
            cost = tokens_in * 2 / 1e6 + tokens_out * 10 / 1e6
            return {**final, "searches": searches, "tokens_in": tokens_in, "tokens_out": tokens_out,
                    "llm_cost_usd": round(cost, 5)}
        for call in msg["tool_calls"]:
            args = json.loads(call["function"].get("arguments") or "{}")
            searches += 1
            messages.append({"role": "tool", "tool_call_id": call["id"],
                             "content": json.dumps(web_search(**args), ensure_ascii=False)})
    raise RuntimeError("turn limit reached without a FINAL line")


def citation_supports(url, forms) -> bool:
    """The cited page must load (HTTP 200) and contain one surface form of the answer."""
    if not url:
        return False
    page = requests.get(url, headers={"User-Agent": "Mozilla/5.0"}, timeout=40)
    text = re.sub(r"\s+", "", page.text).casefold()
    return page.status_code == 200 and any(re.sub(r"\s+", "", f).casefold() in text for f in forms)


if __name__ == "__main__":
    result = ask("On what date was Go 1.27 released? Answer as YYYY-MM-DD.")
    result["citation_ok"] = citation_supports(result["source_url"], ["2026-08-19", "19 August 2026"])
    print(json.dumps(result, indent=1))

Tested on 2026-10-03 (UTC). We ran the program above verbatim at 13:50 UTC. It printed "answer": "2026-08-19", "source_url": "https://go.dev/blog/go1.27", 1 search, 764 input and 59 output tokens, "llm_cost_usd": 0.00212 and "citation_ok": true. GET /v1/tasks/<id>/cost billed the two GPT-6.1 Sol calls $0.000598 and $0.001520 (task ids 79244a1b-2cd3-4582-a5e8-98f416320c91 and d92957ee-d669-4e52-9ee0-20410ea080d0), $0.002118 in total, and the Scholar call $0.000000. A separate request at 13:49 UTC, POST https://api.sandbase.ai/v1/api/scholar/search-web with {"query": "Go 1.27 release date", "max_num_results": 5}, returned 5 results. Trimmed response:

{"id": "d43afca6-3012-4a78-b484-8e929f9c98de", "status": "completed", "model": "scholar/search-web",
 "output": {"results": [
   {"title": "Go 1.27 Release Notes", "link": "https://go.dev/doc/go1.27",
    "snippet": "The latest Go release, version 1.27, arrives in August 2026, six months after Go 1.26. Most of its changes are in the implementation of the toolchain, runtime, ..."},
   {"title": "Go 1.27 is released", "link": "https://go.dev/blog/go1.27",
    "snippet": "Go 1.27 is released Nicholas Husin, 19 August 2026 Today the Go team is pleased to release Go 1.27. You can find its binary archives and ..."}]}}

First 2 of 5 results shown, each complete; the id key inside output is omitted. The trailing ... inside each snippet is part of the text the endpoint returned. That run, and the search inside the program run, came back in the top-level output shape. To swap in another vendor, change only the web_search body. For example, exa/search returns results[].url in one shape and entities[].url in the other.

The request fields are in the Scholar web search API reference. Get a SandBase API key to run the agent on your own questions.

Scope of the data: every question is about public web pages from companies, open-source projects and government statistics offices, with no private individuals. The calls need a SandBase API key, and SandBase is not an official partner of Exa, Tavily, Firecrawl, Cloudsway or the sites searched. The method, question types, grading rule and aggregate results are all in this article. The harness, ground truth and per-run records are kept internally, and we don’t publish the cached page text, because it’s third-party content.

FAQ

What is the best search API for a research agent?

On our 20 questions, Exa (with contents), Firecrawl (with scrape), Scholar web search and Exa Answer all scored 60/60 with GPT-6.1 Sol. Scholar and Exa Answer were the cheapest at about $0.0025 per question. That’s a tie on accuracy, so pick on cost, citations and whether you need full page text.

Exa vs Tavily: which did better?

Both answered all 51 answerable runs correctly, and both had 51/51 valid citations. Tavily lost two false-premise runs to the turn limit and cost $0.0121 per question at the median, against $0.0077 for Exa. Exa’s endpoint changed response shape between calls, so your reader has to handle both forms.

Does a research agent need a separate fetch or scrape tool?

Not for short facts that show up in snippets: Scholar scored 60/60 with search alone. Cloudsway, also search-only, missed five runs, most likely because the fact wasn’t in a snippet (we didn’t log snippets). A fetch tool helps when the answer sits deep in a page, but in this test it was also the main cost driver.

Can’t the LLM answer these from memory?

No. With no tool, GPT-6.1 Sol answered null on 50 of 51 answerable runs and got one Rust release date right, probably from the release cadence. These were questions about August to October 2026, so search was the only way to answer them.

How much does a search-based research agent cost on SandBase?

On 2026-10-03 the search endpoints were billed at $0, and median LLM cost ran from $0.0023 to $0.0121 per question at GPT-6.1 Sol’s $2/$10 per million tokens. Check the model page before you scale; the site was running an API Free Week.

Limitations

This is 20 short-answer questions with exact answers, one LLM, three runs each and default search settings (no depth, date or domain options), all on one day. The top four setups tied, so the test doesn’t measure retrieval quality on hard, multi-hop or long-form research. Results depend on live indexes, which change. Exa’s score is from a rerun after we fixed our own response reader; the first-pass score (55/60) is reported alongside it. The citation check uses a plain GET that follows HTTP redirects but doesn’t run JavaScript or follow meta-refresh redirects, which is stricter than a browser. Latency was measured from one machine with six or seven runs in parallel. Rerun this method on questions that look like your own workload before you choose.