Start with the smallest useful request
The best baseline is a natural-language query withhighlights: true. Exa sizes each result’s excerpts to its relevance, so there is no character budget to tune:
Search vs. Deep Search
Standard search retrieves and ranks pages for a query. Deep Search runs a research process that can search iteratively, inspect what it found, refine the search, and synthesize a grounded result.
Deep modes are recommended by default when using
outputSchema. Read the Deep Search guide for full instructions and examples.
Improve retrieval quality
When results need improvement, change one part of the request at a time.1
Clarify the query
Describe the pages you want, not a bag of keywords. Include the subject and any source type,
time period, or other detail that changes what a relevant result looks like.
2
Read the response in layers
Look at the titles, URLs, publication dates, and highlights before changing the request.The title and URL show the kind of source Exa retrieved,
publishedDate shows
its recency, and the highlight shows the evidence that matched the query. Refine the query to
retrieve different pages, add date filters to narrow the time period, or fetch full text when
you need more context from a useful result.3
Add only hard constraints
Use
includeDomains, excludeDomains, and publication-date filters only when a result that
violates the constraint cannot be used. Put retrieval preferences in the query and, when
synthesizing, put response instructions in systemPrompt.4
Change the search mode last
Use a faster mode for a latency requirement or a deep mode when the retrieval process itself
needs iteration and reasoning. A different mode cannot repair an underspecified query.
Pass the objective from agents
objective is the broader goal behind a search: the task the search is one step of. The query says what to retrieve; the objective says what that task needs from the results, such as which documents should rank first, which to exclude, and which facts or figures to pull. It is most useful when a model picks the query as part of a larger task.
objective as a required string parameter with a description that tells the model what the search turn is trying to achieve. We recommend the wording below verbatim, which is what the Exa MCP server uses:
requestId, searchTime, and costDollars so regressions are reproducible.
Budget latency and context
Each control spends a different resource:
For a real-time path where cached content is acceptable, combine the lowest-latency mode with highlights and cache-only content:
auto and default freshness unless the product has a measured latency target.
To let Exa allocate one context budget across the whole result set — more from strong sources, less from redundant ones — see the Dynamic Highlights research preview.
Tips for common use cases
Common mistakes
Requests copied from older examples often include parameters that are removed, deprecated, or belong somewhere else in the request. The API reference is the source of truth for the current schema.
A few behaviors that are easy to miss:
- The Python SDK uses snake_case for every key, including nested ones:
contents={"highlights": True, "max_age_hours": 24}. stream: truereturns server-sent events only whenoutputSchemais set. Without a schema, the response is the normal JSON body.highlights,text, andsummarycan be combined in one request, but each view is billed separately. Request one view unless the task needs both.
When to use another endpoint
Use a different Exa endpoint when the task changes shape:Next steps
Search API reference
Every request parameter and response field.
Search quickstart
Core request shapes, filters, output, and freshness.
Contents API
Extract highlights or full text from pages you already know.
Exa Agent
Long-running research, list building, and enrichment.