

LLM grounding is a technique for connecting a large language model to external, verified data to reduce hallucinations. In practice, this enables an LLM to base its answers on retrieved evidence rather than relying only on its training data.
Grounding can use private documents, databases, or the live web, and helps models produce more accurate, current, and verifiable answers.
We think a lot about LLM grounding here at Exa. Developers use Exa products like Search and Agent to give LLM-based agents fresh, relevant context from the web and private data, and AI labs use Exa’s 100B-document index to ground models during training, post-training, and evaluation.
This guide explains how grounding works, compares the main methods, and shows how to ground a model with a web search API, using Exa Search as the example.
Most grounding systems turn a user's question into an evidence-based answer through three steps: retrieval, context window injection, and source attribution.
Retrieval. The system searches a data source for information that matches the question. That source may be a vector database of internal documents, a SQL table, or the live web. The retriever returns a small set of passages, records, or pages most likely to contain the answer.
Context window injection. The system places the retrieved text in the prompt, usually with instructions to answer only from that material. The model reads the evidence as it writes the answer. Because the context window has a fixed size, the retrieved text must be concise and relevant.
Source attribution. The model cites the passage or URL that supports each claim. These citations allow a user or an automated check to confirm that the answer matches the source. They also make unsupported claims easier to find because those claims lack citations.
Together, these steps help a model use relevant evidence instead of relying only on its training data. Grounding addresses three practical problems.
Fresh data. Every model has a training cutoff. Prices, product releases, regulations, and news may change after that date. A grounded model retrieves the current version of a fact when the user asks a question.
Private knowledge. A public model has not seen your contracts, support tickets, or internal wiki. Grounding lets the model answer from those documents without retraining. Access controls in the data source determine what each user can see.
Trust in law, medicine, and business. In these fields, an answer without a source is difficult to use. A lawyer reading an AI summary of case law still needs to find the original case. Grounded answers include that evidence, allowing a professional to review the answer before acting on it.
Grounding can take several forms. The following four methods are common, and many production systems combine two or more.
Retrieval-augmented generation (RAG) indexes a document collection in advance, usually as vector embeddings. When a user submits a query, the system finds the chunks most closely related to the question and adds them to the prompt. RAG works well for large, slow-changing collections you control, such as a help center or policy library.
Fine-tuning gives a model additional training on domain-specific examples, helping it learn relevant vocabulary, formats, and patterns. It changes how the model writes and what it tends to know. However, it does not provide live sources or citations when the model answers, so teams often pair it with retrieval.
Knowledge graphs store facts as entities and relationships. A grounded system queries the graph for exact information, such as which subsidiary an organization owns. Knowledge graphs work well for questions that depend more on precise relationships than on long passages of text.
Tool and API use allows a model to call external systems while processing a request. The model identifies the information it needs, calls a search API or database, and reads the result. For public information that changes often, this method commonly relies on web search.
Web search grounds answers in current public information. The model sends a query, the search API returns relevant passages with citations, and the application includes those passages in the next model call. Exa Search supports this process as a web search API.
from exa_py import Exa
exa = Exa() # reads EXA_API_KEY from the environment
question = "What changed in the EU AI Act obligations for general-purpose AI models?"
# 1. Query: send the question to the search API.
response = exa.search(question, type="auto", num_results=5, contents={"highlights": True})
# 2. Cited passages: each result carries a URL and query-relevant highlights.
sources = [
f"[{i}] {r.title} ({r.url})\n" + "\n".join(r.highlights or [])
for i, r in enumerate(response.results, start=1)
]
# 3. Injection: put the numbered sources in the prompt and require citations.
prompt = (
"Answer using only the sources below. Cite each claim with its [number].\n\n"
+ "\n\n".join(sources)
+ f"\n\nQuestion: {question}"
)The quality of web grounding depends on three properties of the search layer: accuracy, precision, and token efficiency.
Accuracy. A grounded answer depends on the evidence available to the model. If the search results omit the correct fact, the model has nothing accurate to cite. Exa is designed to return useful evidence with fewer searches, reducing gaps that could lead the model to guess.
Precision. Each Exa result includes structured data, such as a title, URL, publication date when available, and passages that match the query. The application can connect each citation in the answer to its source URL without parsing prose. It can also flag a claim when its citation number does not match a source before the answer reaches the user.
Token efficiency. Full pages can quickly fill a context window, and each additional search in an agent loop adds more text. Exa returns query-relevant highlights that can be up to 17 times more token-efficient than full-page text. This approach leaves more room in the context window for additional sources and the conversation. Beyond performance, teams in regulated fields must also evaluate how the search layer handles sensitive data.
In healthcare, finance, and legal work, the retrieval layer may handle sensitive data. A user's query can include patient details, deal names, or case facts, all of which the search provider receives. Before choosing a provider, determine whether it stores queries and results, which audits it has passed, and whether it supports the regulations that apply to your organization.
Exa's Enterprise plan includes controls designed for these requirements. Exa is SOC 2 Type II certified, with reports available through the Exa Trust Center. Zero data retention is available for Search, Contents, and Agent on Enterprise plans. When enabled, this option prevents queries and results from being stored or used for training. HIPAA mode, available after Exa enables it for an Enterprise team, also includes zero data retention for those requests.
No. RAG is one method of grounding. Grounding is the broader practice of connecting model outputs to verified data. It also includes live tool calls, such as web searches, knowledge graph queries, and database lookups.
No. Grounding reduces hallucinations, but a model may still misread a source or add an unsupported claim. Citations make these errors easier to identify. An evaluation step that checks each claim against its source can catch additional problems.
Use a vector database for documents you own and control. Use web search when the answer is on the public web, changes often, or may come from sites you cannot name in advance. Many systems use both.
Use only as much text as the question requires. Extra text costs money and adds latency, and irrelevant passages pull the model off the answer. Query-matched highlights provide these focused passages more effectively than full pages.

October 1, 2026

October 1, 2026

October 1, 2026

October 1, 2026

October 1, 2026

October 1, 2026