At Exa, we train, embed, and serve our own large-scale search index and models for perfect retrieval over the internet.
See open rolesJULY 2026
State of the art search over research papers and technical publications.
MAY 2026
Training search agents on Exa instead of a Google baseline yields higher pass@k, faster, at lower cost.
APRIL 2026
Dense, query-specific excerpts let agents ground on the web while reading dramatically fewer tokens.
Our highest-quality search endpoint. SOTA on DSQA, FRAMES, and HLE-Search, and up to 20x faster than competitors.
How we built Canon, a search pipeline orchestrator for full control over our search engine as it scales in complexity.
MARCH 2026
Open-source web search evaluations for coding agents.
Our best agentic search endpoint — faster, cheaper, with structured outputs and field-level grounding
FEBRUARY 2026
The fastest search engine in the world. Sub-200ms latency
JANUARY 2026
How do you store and retrieve information from the web in a database?
MAY 2025
Our 18-node GPU cluster training our next-gen search models
How we cut our BM25 index footprint in half at billions-document scale without sacrificing performance.
DECEMBER 2024
It uses clustering, matryoshka embeddings, binary quantization, and SIMD operations. Written in rust of course 🦀
FEBRUARY 2024
Serving real-time embeddings at scale is challenging. To launch Exa Highlights, we 4X’ed throughput by migrating from Python to Rust.
Our technical team is small and growing quickly. If you are an exceptional AI researcher or engineer, join us to make an outsized impact.
See open roles