

Search once relied mainly on matching words in a query with words in a document. But semantic search – which is a core technology here at Exa – changed the landscape for search dramatically.
Semantic search retrieves content by the meaning, context, and intent of a query instead of by the literal keywords it contains. Put another way, semantic search focuses on meaning rather than wording.
A keyword engine needs the same words on both sides. A semantic engine can match a question to a page that phrases the answer in different words. For instance, Exa runs semantic search for AI agents across the public web and private data.
Semantic search is extremely powerful, and this guide explains how semantic search works and where it shows up, then compares Exa with the other APIs that offer it.
The process starts with three components that help a search system interpret and compare meaning.
Most semantic search systems combine three main parts.
Vector embeddings. An embedding model turns text into a list of numbers called a vector. Texts with similar meanings produce vectors that are close together. Wikipedia's entry on semantic search explains that modern systems use embeddings to convert words, phrases, or documents into vectors. Models such as BERT and Sentence-BERT can create these embeddings.
Vector databases and similarity. After creating the embeddings, the system stores each document vector in an index. When someone submits a query, the system embeds it and returns documents with the closest vectors. A measure such as cosine similarity scores how close two vectors are. Large indexes group vectors into clusters, so the system does not have to compare a query with every vector.
Context and intent. Beyond comparing vectors, the model reads the full query, so surrounding terms shape the result. In a test with the all-MiniLM-L6-v2 model, "Apple earnings" scored closer to "Microsoft quarterly revenue" than to "apple pie recipe," even though the latter shares a word with the query. Some engines also rank results using signals such as location or earlier queries.
Keyword search, also called lexical search, scores documents based on the query terms they contain. BM25 is a common ranking function. Keyword search works well for exact strings such as product codes and error messages, but it can miss documents that use synonyms. Semantic search can find those documents, but it may miss an exact identifier if the embedding model does not recognize it as distinct.
| Keyword search | Semantic search | |
|---|---|---|
| Matching logic | Exact words or phrases in the query | The intent and context of the query |
| Synonym handling | Misses synonyms unless someone adds a synonym list | Matches related terms without a list |
| Method | Inverted index and term scoring such as BM25 | Embedding models and vector similarity |
Because each method has different strengths, many production systems combine their results through hybrid search.
The distinction becomes clearer when lexical, vector, and semantic search are considered separately.
Lexical search looks for exact query words or variations of them. Wikipedia contrasts it with semantic search: lexical search matches words without interpreting the query's meaning.
Vector search is a retrieval method. Meilisearch describes it as converting data into numerical embeddings, storing them in a vector index, and returning the vectors closest to the query vector. This method can search text, images, and documents.
Semantic search aims to return results that match what a user means. Meilisearch notes that it can use text embeddings, vector search, and knowledge graphs to infer intent. Apache Doris describes vector search as one of the most common ways to build semantic search. This broader use distinguishes vector search from semantic search: vector search can also find similar images or products without interpreting meaning.
The following examples show how semantic search connects queries and documents that share no words.
"Cheap laptop" and "affordable notebook." An Apache Doris post on vector search uses this pair to explain the concept. A shopper types "cheap laptop," while the product listing says "affordable notebook." A keyword engine may find no matching words and return nothing useful. It may also return laptop cases. An embedding model places both phrases near each other, so the listing ranks high.
"Wrongful termination" and "at-will employment violations." Firecrawl's guide to semantic search APIs offers a legal example. A lawyer searches for "wrongful termination." The engine finds documents about at-will employment violations even if they do not contain that phrase. The two phrases express related legal concepts in different words.
Both examples show the same advantage. A keyword system needs a manually maintained synonym list, while a semantic system learns these relationships from its embedding model's training data.
This ability to connect different wording makes semantic search useful when people describe what they want in their own words.
E-commerce. Shoppers rarely use the same terms found in a product catalog. Meilisearch gives the example of a search for "jackets to keep me warm" that returns a product labeled "insulated winter jackets."
Enterprise knowledge. Teams often use different terms in company wikis, tickets, and policy documents. Semantic search lets an employee find the right document without knowing its exact title.
Retrieval-augmented generation (RAG). A RAG system finds relevant passages and gives them to a language model for context. OpenAI's file search tool searches uploaded files using semantic and keyword methods.
Web retrieval for agents. An AI agent writes long natural-language queries. A web search API can match those queries by meaning and return useful pages even when their wording differs.
Semantic search tools fall into four groups, ordered from the least setup to the most control.
The first group searches the public web by meaning and requires no custom index.
The next group provides semantic search as a built-in feature but searches data you supply.
Typesense supports built-in semantic search with its own embedding models. It also supports external models from OpenAI, PaLM, and Vertex AI.
Firecrawl searches the web through its /search endpoint. It can return full-page Markdown for each result in the same call.
The third group adds semantic features to an existing keyword search system.
Elasticsearch's semantic_text field creates embeddings at index time and splits long documents into chunks.
OpenSearch supports semantic search through text-embedding models. Users can combine its neural query clause with keyword queries.
The final group gives you the most control: you provide the data and application logic, while the database handles storage and similarity search.
This script uses the open-source Sentence Transformers library. Install the library with pip install sentence-transformers. The script loads a small embedding model, embeds three documents and one query, and then ranks the documents by cosine similarity.
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer("all-MiniLM-L6-v2")
docs = [
"Affordable notebook computers for students",
"Premium gaming desktops with liquid cooling",
"How to repot a houseplant",
]
doc_vecs = model.encode(docs)
query_vec = model.encode("cheap laptop")
scores = util.cos_sim(query_vec, doc_vecs)[0]
for doc, score in sorted(zip(docs, scores), key=lambda x: -x[1]):
print(f"{score:.3f} {doc}")The script prints 0.690 for the notebook listing, 0.342 for the gaming desktops, and -0.003 for the houseplant guide. The notebook listing scores highest even though it shares no words with the query. In a production system, a vector index stores the document vectors so the system computes them only once.
The local example requires you to manage the vectors. For the same type of query on the public web, as well as private data, Exa's Search API handles the embedding and index.
No. Vector search retrieves the vectors nearest to a query vector, while semantic search focuses on meaning and intent. Most semantic search systems use vector search for that purpose. Some also use knowledge graphs or rerankers.
It depends on the query. Semantic search handles paraphrases, questions, and synonyms well. Keyword search works well for exact strings such as SKUs, error codes, and legal citations. Wikipedia notes that hybrid systems combine lexical methods such as BM25 with semantic ranking. Because the methods address different needs, many teams combine them.
Not always. Elasticsearch, OpenSearch, and Typesense store vectors with their keyword indexes. This lets you add semantic search to an existing system. A dedicated vector database such as Pinecone makes sense when vectors are the main workload. If you need web data instead, a semantic web search API such as Exa requires no database on your side.

October 6, 2026

October 1, 2026

October 1, 2026

September 30, 2026

September 30, 2026

September 30, 2026