Train and evaluate models on the world's highest-quality knowledge

Run pre- and post-training on an index of 100B documents and versioned snapshots of the web.

One index for every stage of training

Pre-train on the world's knowledge, hold every training run to a fixed snapshot, and keep every rollout token-efficient.

Build training corpora from a 100B-document index

The index behind Exa, delivered as a clean, comprehensive corpus you can train on. Draw from 1.4T tracked URLs and the billions of topics about the world, including news, case law, financial markets, research papers, points of interest, code documentation, and more.

Train on a fixed version of the web

Exa Snapshot freezes the world's knowledge to a time of your choosing and serves content as it existed up to that point. Run reinforcement learning and model evaluation through an unchanging version of the web, without retrieving answers published after the cutoff.

Cut tokens across each rollout

Dynamic Highlights return relevant passages across search results, using 95% fewer tokens on average than full-page content. Keep rollout context focused on the evidence the model needs. A Qwen3-4B agent trained with Exa reached equivalent performance to a SERP-trained agent with 69% fewer tokens, 62% fewer search calls, and 58% fewer turns.

Exa organizes data across billions of sources

Every datapoint is selected on source quality and freshness, based on the latest events from today

70M+ Businesses
Pricing pages
1B+ People Records
Job postings
Contact data
Product reviews
Latest News
Funding rounds
Financial market data
Exa Connect
Earnings call transcripts
Investor portfolios
Recent hires
Conference speaker lists

Connect Exa to your training and eval workflow

Claude

Train and evaluate

ChatGPT

Ground model rollouts

LangChain

Wire search into agents

Exa MCP

Connect research tools

AWS

Run workflows

Databricks

Join eval traces to data

See all integrations (55)

Common questions

AI labs put Exa Search inside training and eval loops: a model retrieves from Exa's index of 100B documents during post-training RL rollouts, and the same index grounds live-web model evals. Exa returns relevant passages rather than full pages, which cuts tokens across each rollout, and Snapshot, Exa's versioned web, holds the index fixed so later runs stay comparable.

In Exa's published study of May 13, 2026, a Qwen3-4B model trained with Exa Search in the loop scored Pass@1 of 0.767 on SimpleQA, 0.694 on HotpotQA, and 0.839 on 2WikiMultihopQA, against 0.692, 0.684, and 0.798 for the same model trained on SERP search. The Exa-trained model reached that performance with 69% fewer tokens, 62% fewer search calls, and 58% fewer turns. The figures are scoped to that one study; full setup and tables are in the RL search outcomes post.

Yes. Snapshot, Exa's versioned web, serves Exa's index at a fixed point in time. Training, eval, and backtesting workloads use the static snapshot to keep the environment controlled and consistent. Snapshot terms vary by contract, so ask Exa sales what is available on your deal.

Yes, through Exa Monitors and Exa Search. Exa Monitors runs an Exa search on a recurring schedule, drops the results it has already returned, and sends each new one to your webhook, so a query written around a vendor's release notes delivers only what is new. Exa Search covers release notes, changelogs, and technical writing across the open web. Literature and paper tracking is a separate job from training.

No. Exa is a search API: it searches the full text of 350 million publications by meaning, reaches the open web around each paper, and runs inside your own product. Research teams often use Exa alongside a curated index.

Exa is SOC 2 Type II audited, and honors GDPR and CCPA data-subject rights. Zero data retention is available, meaning your queries and results are not stored or used for training. Single sign-on comes with an Enterprise plan. Terms are on the enterprise page.

Train your next model on the world's knowledge