Overview
Semantic search matches meaning rather than literal keywords by comparing dense embeddings. This guide shows a practical end-to-end pipeline you can implement today: chunk documents, compute embeddings, build a vector index (FAISS), add optional sparse (TF‑IDF) signals for hybrid ranking, and serve filtered, ranked results.
When to use semantic search
- Good: question-answering over documentation, retrieval-augmented generation (RAG), fuzzy matching across long-form content.
- Not ideal: exact-match legal or financial terms where keyword precision and boolean search are required.
Core components
- Chunking: split long documents into retrieval-sized pieces.
- Embeddings: encode chunks and queries into vectors (local or API models).
- Vector index: FAISS, Milvus, Weaviate, Pinecone, etc.
- Optional sparse index: TF‑IDF or BM25 to combine lexical signals.
- Metadata store: map vector ids to text, document id, timestamps, permissions.
- Retriever: search, filter, and return ranked candidates to your app or LLM.
Minimal example: chunk → embed → FAISS
Below is a compact Python example that demonstrates chunking text, computing embeddings with sentence-transformers, building a FAISS index, and a simple search function. This is intentionally minimal; production systems will add batching, persistence, and async I/O.
from typing import List
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
model = SentenceTransformer('all-MiniLM-L6-v2') # lightweight, fast encoder
# 1) Chunking
def chunk_text(text: str, chunk_size: int = 200, overlap: int = 40) -> List[str]:
tokens = text.split()
chunks = []
start = 0
while start < len(tokens):
end = min(start + chunk_size, len(tokens))
chunks.append(' '.join(tokens[start:end]))
if end == len(tokens):
break
start = max(end - overlap, 0)
return chunks
# Example document
text = """Long document text goes here..."""
chunks = chunk_text(text)
# 2) Embeddings (batch)
embeddings = model.encode(chunks, show_progress_bar=False, convert_to_numpy=True)
# 3) Build FAISS index (cosine via normalized inner product)
d = embeddings.shape[1]
faiss.normalize_L2(embeddings) # normalize for cosine similarity
index = faiss.IndexFlatIP(d) # for small datasets; swap with IVF/HNSW for large
index.add(embeddings)
# simple id->metadata map
id_to_meta = {i: {'text': chunks[i], 'doc_id': 'doc1', 'chunk_id': i} for i in range(len(chunks))}
# 4) Search function
def search(query: str, k: int = 5):
q_emb = model.encode([query], convert_to_numpy=True)
faiss.normalize_L2(q_emb)
D, I = index.search(q_emb, k)
results = []
for score, idx in zip(D[0], I[0]):
meta = id_to_meta.get(int(idx), {})
results.append({'score': float(score), 'text': meta.get('text'), 'meta': meta})
return results
# usage
print(search('How do I authenticate?'))Hybrid ranking: add a sparse signal
Combining vector similarity with TF‑IDF or BM25 often improves relevance for keyword-heavy queries. The example below shows a simple TF‑IDF + dense combination.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.preprocessing import normalize
# build TF-IDF over the same chunks
tf = TfidfVectorizer().fit_transform(chunks) # shape (n_chunks, n_terms)
# pre-normalize dense vectors if not already
dense_matrix = embeddings.copy()
# embeddings already normalized above; ensure 2D
# combined search
def hybrid_search(query: str, k: int = 5, alpha: float = 0.6):
# dense score
q_emb = model.encode([query], convert_to_numpy=True)
faiss.normalize_L2(q_emb)
D, I = index.search(q_emb, k*3) # fetch more from dense
dense_candidates = list(I[0])
# sparse scores
q_tfidf = TfidfVectorizer().fit(chunks + [query]).transform([query])
sparse_scores = (tf @ q_tfidf.T).toarray().ravel()
# combine and rank candidate set
candidates = set(dense_candidates)
scored = []
for idx in candidates:
dense_score = 0.0
# approximate dense score by checking position in I
if idx in I[0]:
pos = list(I[0]).index(idx)
dense_score = float(D[0][pos])
sparse_score = float(sparse_scores[int(idx)])
# normalize sparse (optional) and combine
final = alpha * dense_score + (1 - alpha) * sparse_score
scored.append((final, int(idx)))
scored.sort(reverse=True)
return [{'score': s, 'text': id_to_meta[i]['text']} for s, i in scored[:k]]
# usage
print(hybrid_search('authenticate with API'))Metadata filtering
Store metadata in a separate key-value store or alongside your vector DB. Filter candidate ids before scoring to enforce permissions, language, or document date.
Tip: keep vector-only operations in the vector DB and apply strict filters in the metadata layer to avoid leaking unauthorized documents.
Tradeoffs and performance tips
- Index type: IndexFlatIP is exact but memory-heavy. For larger collections use IVF+PQ or HNSW. Use FAISS IVFPQ for disk‑efficient indexes and HNSW for fast recall at low latency.
- Dimensionality: smaller embedding models (e.g., 384 dims) use less memory and often perform well for many tasks. Larger models may increase semantic fidelity but cost CPU/GPU time.
- Chunk size: short chunks improve precision but increase index size. Aim for ~150–400 tokens depending on your retrieval needs.
- Freshness & updates: real-time insertion is supported by HNSW and many hosted vector DBs. For IVF indexes, plan periodic reindexing or use hybrid approaches for near-real-time data.
- Embedding cost: precompute and cache embeddings for static content. For dynamic content, batch embedding requests and use GPU inference when possible.
- Hybrid vs pure dense: Hybrid (dense + sparse/BM25) improves precision on keyword queries and is often a practical compromise.
- Privacy and compliance: if embeddings are sent to third-party APIs, ensure you follow data policies and consider private/local models for sensitive content.
Deployment and scaling checklist
- Precompute embeddings for static documents and store them alongside metadata.
- Use batching and GPU for embedding large ingests.
- Choose an index type that matches read/write and latency targets (HNSW for low-latency, IVF/PQ for low-memory).
- Expose a retriever API that returns candidate IDs + metadata, not raw vectors.
- Monitor recall and relevance with labeled queries and iteratively tune chunking, candidate pool size, and alpha in hybrid ranking.
Further reading
For a discussion of what "semantic search" means in practice, see the Stack Overflow blog post: What (un)exactly do you mean by semantic search?
FAISS docs: FAISS on GitHub. SentenceTransformers: sbert.net.
Concise conclusion
Semantic search is a practical technique when you need meaning-aware retrieval. Start small: chunk your docs, precompute embeddings, index with FAISS, and add a sparse lexical signal for hybrid ranking. Monitor relevance, tune chunk size and candidate counts, and choose index types that balance cost, latency, and update needs. This pipeline will serve as a solid foundation for RAG, knowledge search, and many developer-facing search features.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment