Introducing Exa Agent Ultra – Lists without Limits

Introducing Exa Agent Ultra – Lists without Limits
The Exa Team
The Exa Team
Sep 24, 2026

Today we shipped Exa Agent Ultra, setting a new state of the art for comprehensive deep research. Ultra is the highest effort level for our Exa Agent. It uses frontier intelligence better and cheaper for web search, delivering frontier-level results at a fraction of the cost. Leveraging Exa's index of over 100 billion documents, Ultra is the best way to retrieve exhaustive, accurate information from the web and beyond.

Frontier deep research

Exa Agent Ultra delivers higher scores and cost efficiency across multiple benchmarks: WANDR, WideSearch, DeepSearchQA, and our internal Find-All Company benchmark.

Agent Ultra leads all four benchmarks, scoring meaningfully higher than GPT-6 Astra, Opus 5.5, and Perplexity Agent at significantly lower cost per task versus frontier models:

  • WANDR. Ultra scores 81.4% versus GPT-6 Astra at 26.3% and Opus 5.5 at 72.3%. Ultra also demonstrates significantly lower cost per task, 44% cheaper than Opus 5.5 and 20% cheaper than GPT-6 Astra.
  • WideSearch. Ultra scored 58.9%, ahead of Perplexity Agent (56.0%), GPT-6 Astra (54.7%) and Opus 5.5 (51.6%). As above, this was also cost-effective, at $3.85 per task: 25% less than Perplexity Agent, 40% less than Opus 5.5 and 54% less than GPT-6 Astra.
  • DeepSearchQA. Ultra scored 93.9%, more than 8 percentage points higher than GPT-6 Astra (85.3%) and 16 points higher than Opus 5.5 (77.6%). Agent Ultra also cost 52% less per task than GPT-6 Astra and 53% less than Opus 5.5.
  • Find-All Company. Ultra succeeded on 229 rows per task compared to GPT-6 Astra (60 rows), Opus 5.5 (49 rows) and Perplexity Agent (48 rows). On this benchmark, GPT-6 Astra was the most cost-effective, at $7.35 per task.

WANDR

Soft recall / Cost
Soft recall (%)
0
20
40
60
80
100
Exa Agent Ultra$18.53 · 81.4%
Opus 5.5$32.92 · 72.3%
Perplexity Agent$28.38 · 40.2%
GPT-6 Astra$23.09 · 26.3%
$10.00$15.00$20.00$25.00$30.00$35.00$40.00
p50 cost / task ($)

We evaluated Agent Ultra with WANDR (Wide And Deep Research), a benchmark that tests research agents for ability to find a large list of qualifying entities (wide) with evidence-supported accuracy across many fields (deep). Tasks are modeled on entity enrichment use cases in domains such as due diligence and legal research. Our grader shares the same vendored evaluation logic as the upstream repository; the core differences lie in the contents tool (Exa), the transport logic, and a better + cheaper judge model (gpt-6-luna). Where Perplexity, Anthropic or OpenAI had published a result obtained on this grader harness, we report the published figure; where none existed, we ran the benchmark ourselves.

Lists without limits

Exa Agent divides a task into subtasks, assigning subagents to research multiple domains at once. When researching, it uses a mix of models (frontier and cost-effective) to find the most cost-effective methodology for a given task. As the highest effort mode for Exa Agent, Ultra is the right choice when completeness matters more than latency or cost. It is useful for deep research, list-building, and entity enrichment – any task that needs to run to exhaustion.

Examples:

Model providers

  • Assemble training data from the web, such as "every paper and code repo implementing a given technique" or "every public benchmark with a leaderboard and its top submission".
  • Verify hard-to-check criteria about an entity like "has a reproducible eval harness", or "released weights, not just an API".

Financial services

  • Create market maps for diligence like "every company that does battery recycling in Europe, adding founders, funding, and customers for each".
  • Research a customer for KYC compliance across many sources such as filings, press, court records, and regulator sites, synthesized into one record with citations.
  • Monitoring for signals, like "every portfolio company that announced a layoff, CFO change, or new facility this quarter".

Go-to-market teams

  • Build account lists like "every US company that makes browser-automation tooling" or "every Series A–C fintech with a Head of Risk hired in the last 12 months"
  • Enrich a list with new fields requiring judgment, "does this company sell to hospitals?", "who owns the data platform?", with a cited URL as evidence.

Ultra makes it easy to expand lists by passing the rows you already have, excluding them from future results.

Try Ultra today

Agent Ultra is available today. Read the docs here or try it out in our API Playground.