Does the AI world really need another search benchmark?

Does the AI world really need another search benchmark?
Kale Bogdanovs
Oct 8, 2026

Today we announced ATLAS, a new benchmark that grades the performance of search agents across 547 wide search tasks. At this point, with WideSearch, DeepSearchQA, BrowseComp, WANDR, and many others, it is fair to ask: does the world really need another search benchmark? We think it does. Current benchmarks are either substantially memorized by frontier models, fail to reflect real-world search tasks, or are not precisely graded or reproducible.

Designing an effective search benchmark is a tough challenge. It needs to test agents on tasks of sufficient scale and difficulty while still being able to fairly grade them on completeness and accuracy at a manageable cost. The current crop of search benchmarks all make significant compromises, either in the tasks themselves or how they’re graded, that prevent them from achieving these goals.

We found that the best search agents of today are highly accurate; more than 90% of the entities they find are correct. However, they fail to find a third or more of possible entities. ATLAS is designed to help us track and ultimately close that completeness gap.

What is a search benchmark trying to measure?

Agentic research is used for difficult and high-value tasks. Either each individual entity represents real economic value, or the risk associated with missing a single critical entity is high. In other words, getting everything correct matters. For example:

  • In a search for companies matching complex criteria, returned entities are often sales leads. Say there are 743 companies matching the search criteria but a search agent only finds 412. If you can close 4.3% of leads at an ACV of $85K, those missing 331 leads are worth about $1.2M, and that’s just one search.
  • Investment firms may research dozens of companies before spending millions on a purchase. Due diligence for these deals requires exhaustive search for adverse media or for any prior art that could affect patents. Finding hundreds of documents but missing just one piece of vital information can put the entire investment at risk.

When research tasks are used to support agentic workflows, the ultimate goal is to reach a level of confidence in an assessment that justifies an action: a campaign, an investment, a partnership, etc. For that to happen, a result needs two qualities:

  • Accuracy: entities and their properties must be correct and the information must be attributable to a trusted source.
  • Completeness: All available entities must be found and gaps in available information must be known.

What does a useful benchmark look like?

To be able to successfully and fairly grade agents on completeness and accuracy in wide research tasks, a benchmark needs to meet four criteria.

  1. Tests search, not memory

    Tasks must require information the models don’t already know. If frontier models can answer from memory, the benchmark is not useful. In practice, given that frontier model training cycles have historically lasted 3-5 months,[1] tasks much more than six months old lose value quickly.

  2. Difficult

    Tasks have to be hard and at least some need to be large. If tasks are too easy, the benchmark will quickly become exhausted and models will cluster together. A benchmark is only useful if it separates out competitors.

  3. Based in real demand

    Tasks have to resemble actual work users want research agents to perform. Ideally they should be grounded in real queries. If every task in a set is made up by a research team, a high score might not reflect real-world results.

  4. Economically viable

    For a benchmark to succeed at scale in the long term the cost of creating tasks needs to be manageable, as does the cost to an assessor to run the benchmark against the models they want to test. The more a benchmark costs to run, the less it will be used.

The big problem: grading

Most wide search benchmarks begin with a similar format:

Example ATLAS task: which U.S. patents had a claim's unpatentability affirmed by the Federal Circuit on appeal from a PTAB inter partes review, 1 Aug–18 Sep 2026, with petitioner HQ, lead counsel's law school, first panel judge's commission date and the petition expert's doctorate. Wide: 20 patents to find. Deep: patent to IPR petition to lead counsel to law school. Fresh: decided weeks before the eval.
  1. Give the agents a list of tasks to complete. The form is: find all entities that match X conditions, and enrich Y attributes.
  2. Grade the performance of the agent for accuracy and completeness. Rank agents by how well they perform the task, as well as how long they take, and how much it costs.

Step one is straightforward. It’s necessary to regularly refresh tasks to prevent memorization and maintain the benchmark’s value. Research teams can create tasks, or source them from user surveys. Alternately, tasks can be built from real topics sourced from an agent provider. This second approach allows for a continuous stream of tasks that are known to reflect real user demand.

The second step is trickier. We’ve already established that the research tasks have to be difficult, so how can the answers be confidently judged? There are two main approaches to this problem. One approach is to build a “golden” answer key with complete and accurate answers. The second is to instead create a rubric to judge answer quality and have an LLM judge grade each response at run time.

Traditionally, each approach has required some major compromises.

Golden Answer Key

The big problem with an answer key is how to make an accurate and complete answer key for a sufficiently difficult task. Benchmarks built on answer keys usually use teams of human researchers to compile the key. This can get you an accurate and complete answer key but there are limitations.

First, creating the answer key is too expensive to allow for regular refreshing. The most common benchmarks using human derived answer keys are now 9-18 months old[2]. That’s more than enough time for the data to be built into a model’s weights[3]. Second, the need to conduct intensive human research limits the practical size and difficulty of the benchmark.

This means that for golden key benchmarks like DeepSearchQA and BrowseComp, all search agents perform reasonably well and there is relatively little separation between the best and worst performers.[4]

In short, benchmarks that use answer keys can test completeness and accuracy of tasks, but they struggle to produce enough difficult tasks and to stay fresh enough to be useful.

LLM Judge

In this type of benchmark, a second LLM grades answers on whether the cited source supports the model’s claim for each entity and cell. The task has a “target” number of entities to retrieve, rather than working from a complete list.

The key advantage of LLM judges is that they don’t require the answer to be known ahead of time. They can test on larger tasks and can gracefully handle facts that change over time, like new hires for key positions in a company search. However, they introduce several of their own critical limitations.

LLM judge grading vs answer-key grading: an LLM judge only scores the 40 rows the agent returned and never counts the 23 missing companies; an answer key scores 40 found, 23 missed and 6 wrong

First, they don’t test for completeness. If a model finds the target number of records, it can score full marks, even though in many real-world research tasks, missing one important entity can be critical. Second, they grade on whether the quoted source supports the agent’s answers, which is not the same as whether the answers are true. Finally, the need to run the LLM judge against each agent’s response makes these benchmarks orders of magnitude more expensive to run and grade.

In other words, LLM judges can help overcome the freshness and difficulty issues faced by golden key-based benchmarks, but they can’t reliably track progress towards completeness and accuracy in research tasks.

Exa’s approach to building a better benchmark

When designing our benchmark, Exa was not satisfied with either approach. Ultimately, we decided that the limitations of LLM judges in scoring accuracy and completeness were not acceptable and that an answer key was necessary.

The challenge then was to build an answer key that could be refreshed regularly enough to prevent memorization and that covered more and harder tasks than current answer key benchmarks. In the end, that meant building an automated pipeline that would allow agents, rather than human researchers, to build a task list and a comprehensive answer key.

ATLAS build pipeline: seed, discover, enrich, stress-test, audit, answer key, regenerated from fresh demand every 3–6 months

For an in-depth look at how we built out ATLAS, read the technical paper. Here’s the short version:

Tasks start from topics our customers search for heavily. We use aggregate topic names and never read raw queries. For each task, two builder agents, one on ChatGPT and one on Claude, independently hunt for every entity that fits. Both search with a shared pool of Exa, Brave, and Perplexity results with the provider names hidden, so neither can lean on any one index. A separate verifier, using a search index outside that pool, checks every row against its evidence. Each column has to clear three gates before it’s kept: its values must come from pages beyond the ones that found the entity, a model can’t recite them from memory, and no single page answers the whole column. Only about one in five proposed columns makes it.

Then we tried to break the key. Twelve research systems ran every task, and every disagreement with the key became a lead. Judges from two model families researched each dispute without knowing which system gave the answer, and the key changed only when both agreed on verified evidence. Anything nobody could settle is left ungraded rather than guessed. Building and checking one task took a median of 8 agent-hours and about 1,200 searches, close to 4,800 agent-hours across all 547 tasks. A sampled audit puts wrong gold values under 1%. That cost is paid once. The result is an answer key that represents the best information a very well-resourced agent could be expected to know, and that takes into account information gaps and uncertainty.

After that, grading is a table comparison. While an LLM is still used for fuzzy matching, 95.5% of cells are scored without it, and grading one system on the whole benchmark costs about 24 cents. Because the pipeline is automated, we can rerun it on fresh topics every three to six months and stay ahead of what models have memorized.

Where ATLAS succeeds

Looking back at our success criteria for a benchmark, let’s judge how well ATLAS succeeds.

  1. Tests search not memory

    In a closed-book run (i.e. with no access to search tools), agents return many correct answers on existing benchmarks. On ATLAS, agents fail without web search.

    Scores without search: Google DeepSearchQA 61%, OpenAI BrowseComp 48%, ByteDance WideSearch 47%, Exa ATLAS 9%
  2. Difficult

    Exa Agent Ultra, an Exa model specially designed for intense research, gets all attributes for each entity correct about two thirds of the time. Of the frontier models, Opus 5.5 does best at a little under 50 percent, leaving plenty of headroom before the benchmark is exhausted.

    ATLAS benchmark: row F1 against median cost per task (log scale) for Exa Agent, Opus 5.5, Perplexity and GPT-6 Astra
  3. Based in real demand

    All 547 tasks in the current version of ATLAS are sourced from real search topics, with the automated pipeline ready to refresh with a new set of tasks when needed.

  4. Economically viable

    Using agents to create the answer key allows the benchmark to be built and refreshed at a lower cost than human-generated answer keys. Meanwhile, the cost per graded answer to run the benchmark comes in at 0.04 cents, compared to 88 cents for the LLM-judged WANDR benchmark when we ran it.

What does this mean for you?

We built ATLAS primarily to tell us how we were tracking towards our goal of a research agent that you can trust to power critical decision-making workflows. This requires the ability to conduct large and difficult research tasks comprehensively and accurately.

ATLAS is the most capable benchmark yet at balancing the need to test on a large set of difficult tasks, and confidently grade for the completeness and accuracy of an agent’s responses. That means that a high score on ATLAS is stronger evidence for real-world effectiveness than any other benchmark provides.

A few takeaways from our runs of ATLAS so far:

  • There is still a lot of progress to be made on comprehensive search at a reasonable cost. Max-effort search agents like Exa Agent Ultra performed better than cheaper counterparts, but no agent run costing less than $1 per task achieved a success rate of over 50%.
  • Even the most expensive search agents miss about 1/3 of the golden results, indicating substantial room for web search to improve performance in use cases where completeness is critical.
  • When we tested just search APIs by holding the model fixed in a harness, we found that the choice of search API made a big difference in quality with scores ranging 16% across different search backends.

In short, when it comes to difficult research tasks, no agent can yet guarantee completeness of results, but the model and search backend you choose make a significant difference in the amount and quality of results you will find.

What’s next?

While we are confident that the ATLAS answer key is of exceptional quality, we’re currently conducting a blind human audit of a sample of tasks to ensure that they meet our standard of 99% of gold values proving correct. We’ll also be open sourcing this benchmark in a few weeks so that you can run it for yourself against any agents you’re considering for challenging search tasks.


Notes

  1. Recent frontier models shipped about 2–5 months after their data cutoffs: Claude Opus 5 (May → July 24, 2026), Claude Opus 5.5 (June → Sept. 22, 2026), GPT-5.6 (Feb. 16 → July 9, 2026) and GPT-6 Astra (April 30 → Sept. 3, 2026). Source: allmo.ai, LLM knowledge cut-off dates. ↩
  2. BrowseComp (OpenAI, April 2025), WideSearch (ByteDance Seed, August 2025) and DeepSearchQA (Google DeepMind, January 2026). ↩
  3. With no search tools, GPT-5.6 Sol answers 46.1% of BrowseComp-Plus questions from memory (“Projecting BrowseComp-Plus onto ClimbMix,” arXiv:2608.20317). ↩
  4. Top systems on BrowseComp now score within about a point of each other: GPT-6 Astra 91.5%, Kimi K3 91.2%, Claude Opus 5 90.8% and GPT-5.6 Sol 90.4%. Scores are self-reported by each lab. (tensorfeed.ai/benchmarks/browsecomp) ↩