

RAG stands for retrieval-augmented generation. A RAG API connects a language model to external or private data, allowing the model to answer from that information instead of relying only on its training memory.
At Exa, we build retrieval infrastructure for AI applications, including the search layer behind web-connected experiences at OpenRouter, Cursor, and Firefox. Exa retrieves current web content in a form designed to go directly into a model's context.
This guide explains how RAG works, compares open-source and managed approaches, and shows how to build the retrieval layer with Exa.
Most retrieval systems use the same five-stage pipeline, which a single RAG API call can trigger.

Ingestion and chunking. The API loads files from uploads, storage buckets, or connected apps. It splits each document into chunks of a few hundred tokens because a whole document is often too large and unfocused to send to a model.
Embedding. An embedding model turns each chunk into a vector. Chunks with similar meaning get vectors that sit close together.
Vector storage. The API stores the vectors in a vector database with the chunk text and metadata, such as the file ID, owner, and date.
Retrieval. When a question arrives, the API embeds it, finds the nearest chunks, and often adds keyword search and reranking. It then returns the most relevant chunks.
Generation. The API adds the retrieved chunks to the prompt and calls the model. A strong implementation asks the model to cite the chunk that supports each claim.
The first three stages run when documents are added or updated. The final two, retrieval and generation, run whenever a user submits a question.
A typical self-hosted RAG API combines four open-source components.
FastAPI serves the HTTP endpoints for upload, query, and delete.
LangChain or LlamaIndex handles document loaders, chunking, and the retrieval chain.
pgvector or Chroma stores the vectors. pgvector adds vector search to PostgreSQL, making it a good fit for teams that already use Postgres.
Ollama runs embedding and chat models locally, so documents never leave your servers.
The LibreChat project provides a reference implementation of this stack. Its rag_api repository combines LangChain and FastAPI with PostgreSQL and pgvector for document indexing and retrieval. It organizes embeddings by file_id, allowing an application to query a specific file and receive relevant passages with their metadata.
LibreChat primarily uses the API for file search, although the README notes that any ID-based application can use it. pgvector is the default vector store, with MongoDB Atlas available as an alternative. Teams can also choose from embedding providers such as OpenAI, Bedrock, Hugging Face, Vertex AI, and Ollama. This flexibility requires ongoing engineering work for parsing, access control, and upgrades.
Teams that do not want to operate a self-hosted stack can use managed RAG APIs. These products host some or all of the pipeline and expose its capabilities through endpoints or model tools.
Google Cloud RAG Engine is part of the Gemini Enterprise Agent Platform. It ingests data from local files, Cloud Storage, and Google Drive. It then indexes the chunks in a corpus that Gemini models can search.
OpenAI Responses API file search is a hosted tool. You upload files to vector stores, and the model searches them using semantic and keyword search.
Cohere Chat API accepts a documents parameter with the query and returns a grounded answer with inline citations to those documents.
CustomGPT.ai sells a RAG API with REST endpoints and SDKs, and its site lists support for more than 1,400 file types.
RAG as a service (RAGaaS) extends this managed approach into a complete cloud platform that connects data sources to a language model. The vendor operates ingestion, chunking, embedding, vector storage, and retrieval, while applications access those capabilities through an API. This model removes the need for an internal team to build and maintain the underlying infrastructure.
Most RAGaaS platforms share five features.
Connectors. Prebuilt integrations sync content from sources such as Google Drive, SharePoint, Confluence, Notion, and Slack.
Chunking. The platform parses each file type and splits it into retrievable passages without custom code.
Hybrid search. Retrieval combines keyword and semantic search to find exact terms and paraphrases in one query.
Citations. Answers return links to the source passages, so users can check each fact.
Permissions. The platform syncs access controls from source systems and limits results to content each user can view.
The following three platforms illustrate common approaches to RAG as a service.
Amazon Bedrock Knowledge Bases connects to SharePoint, Confluence, Google Drive, OneDrive, S3, and a web crawler. With the managed option, AWS operates the vector store, embeddings, and reranking models. Teams that need more control can use their own vector store.
Vectara offers an enterprise agent platform that can run as SaaS, in a customer VPC, or on premises. Customers can also bring their own model. Vectara states that the platform detects and corrects hallucinations at runtime.
Coveo offers RAG as a service built on its enterprise search index, which draws on more than 15 years of search development. Its Passage Retrieval API returns ranked passages with document permissions already applied.
A document RAG system answers from previously ingested data, so its knowledge is only as current as the latest sync. It cannot cover this morning's news or last week's software release unless that information has already been added. For most teams, crawling and embedding the entire web is not practical.
Web retrieval closes this freshness gap at query time. The application sends the question to a web search API, then passes the results to the model alongside relevant passages from the private index.
Exa is one option for the web retrieval layer in this architecture. Its search API uses a proprietary web index, while an in-house extraction model evaluates each result against the query and returns only the passages that answer it. This process supplies relevant chunks without requiring the application to scrape or split web pages.
Published benchmarks indicate that these query-relevant highlights can be up to 17 times more token-efficient than full-page text. Each result includes a URL and title for citations, and domain filters can restrict retrieval to trusted sources. Search costs $7 per 1,000 requests.
A common architecture routes each question to the private index, web search, or both. It then combines the retrieved passages in one prompt and labels each passage with its source.
No. A vector database stores vectors and finds the nearest matches. A RAG API adds the surrounding ingestion, chunking, embedding, retrieval, and often generation steps. Together, these components allow one API call to turn a question into a grounded answer.
Pricing varies by vendor. OpenAI's file search, for example, charges $0.10 per GB of storage per day after the first free GB, plus $2.50 per 1,000 tool calls and the model's token costs. Compare storage, query, and token charges when evaluating the total cost.
Yes. Some platforms, including Bedrock Knowledge Bases, provide a web crawler connector, but crawled pages are only as current as the latest sync. For up-to-date information, add a web search API that runs when the user submits a question.
Self-hosting suits teams that must keep data within their infrastructure or need custom retrieval logic, but it requires ongoing engineering work. A managed service suits teams that prioritize built-in connectors, synchronized permissions, and reliable uptime without operating the pipeline themselves.

October 6, 2026

October 1, 2026

October 1, 2026

October 1, 2026

October 1, 2026

October 1, 2026