Introduction
Semantic search—matching intent and meaning rather than exact keywords—has moved from research demos into production systems. Recent industry discussions highlight the need to standardize what we mean by semantic search and the components that make it reliable. This article gives a practical, implementation-first guide you can adapt: embedding selection, vector storage options, ingestion patterns, retrieval, optional reranking, and production tradeoffs.
What is semantic search (brief)
At a high level, semantic search converts text into numeric vectors (embeddings) that encode meaning, stores those vectors in a vector database or index, and returns nearest neighbors for a query vector. Optional reranking with a stronger model or business logic improves answer quality.
Typical architecture
- Ingest: text & metadata → chunking → embeddings
- Indexing: store embedding vectors + payload in a vector DB (FAISS, Milvus, Weaviate, Pinecone, pgvector, etc.)
- Retrieve: nearest-neighbor search for query embedding
- Rerank / Filter: optionally re-score candidates with a stronger model, heuristics, or business rules
- Serve: return results, with caching, pagination, and telemetry
Choosing embeddings: tradeoffs
- Open-source vs cloud models: Open-source (sentence-transformers, Llama/other local embeddings) gives control and lower per-call costs but may need GPUs for high throughput. Cloud embeddings (hosted APIs) reduce ops burden and often provide higher-quality representations for some tasks.
- Dimensionality: Higher-dim vectors can improve accuracy but increase storage and search cost. Typical ranges: 384–1536 dimensions for many models.
- Latency vs cost: Smaller models run faster and cheaper locally; larger models cost more but may reduce downstream reranker calls.
- Normalization and similarity metric: Use cosine similarity (normalize vectors) or inner product; index must match metric.
Vector store options & tradeoffs
- FAISS — great for on-prem/offline; CPU/GPU deployment flexibility; you manage persistence and scaling.
- pgvector — fits well if you already use Postgres; simple to deploy; good for smaller datasets.
- Milvus / Weaviate — feature-rich, scalable, add-ons for metadata filtering and hybrid search.
- Managed services (Pinecone, VectorDBs) — simplify operations and scale but introduce vendor lock-in and recurring cost.
Minimal runnable pipeline (Python + sentence-transformers + FAISS)
Below is a compact example showing embedding, index creation, and query. Use this as a reference implementation to test ideas locally before choosing a production vector store.
from sentence_transformers import SentenceTransformer
import numpy as np
import faiss
# 1) Load model and documents
model = SentenceTransformer('all-MiniLM-L6-v2')
docs = [
'How to reset your password',
'Troubleshooting network connectivity',
'Deploying the app with Docker'
]
# 2) Embed documents
embeddings = model.encode(docs, show_progress_bar=False, convert_to_numpy=True)
# 3) Normalize embeddings for cosine similarity
faiss.normalize_L2(embeddings)
# 4) Build FAISS index
d = embeddings.shape[1]
index = faiss.IndexFlatIP(d) # Inner product on L2-normalized vectors = cosine similarity
index.add(embeddings)
# 5) Query
query = 'How do I fix my network?'
q_emb = model.encode([query], convert_to_numpy=True)
faiss.normalize_L2(q_emb)
k = 3
D, I = index.search(q_emb, k)
print('scores:', D)
print('indices:', I)Persisting and serving with Postgres + pgvector
If you prefer keeping vectors alongside relational data, pgvector is a pragmatic choice. Create a table and insert vectors from Python.
-- SQL (for reference):
-- CREATE EXTENSION IF NOT EXISTS vector;
-- CREATE TABLE documents (id SERIAL PRIMARY KEY, content TEXT, embedding vector(384));
import psycopg2
from psycopg2.extras import register_adapter, Json
from sentence_transformers import SentenceTransformer
model = SentenceTransformer('all-MiniLM-L6-v2')
conn = psycopg2.connect('postgresql://user:pass@localhost/db')
cur = conn.cursor()
docs = ['First doc', 'Second doc']
embs = model.encode(docs, convert_to_numpy=False).tolist()
for doc, emb in zip(docs, embs):
cur.execute('INSERT INTO documents (content, embedding) VALUES (%s, %s)', (doc, emb))
conn.commit()
cur.close()
conn.close()Retrieval + Rerank pattern
A common production pattern is to retrieve an initial candidate set (k = 10–50), then rerank with a more powerful model or business heuristics. Reranking reduces the number of expensive LLM calls while improving final answer quality.
# Pseudocode for reranking using an external LLM API
# 1) Retrieve top-K candidates from vector store (ids, scores, content)
# 2) Build a prompt that lists candidates and asks the model to rank or score them
# 3) Call LLM to produce final ranking
candidates = [
{'id': 1, 'content': 'Troubleshooting network connectivity'},
{'id': 3, 'content': 'Deploying the app with Docker'}
]
prompt = 'Rank these candidate documents for the query "How do I fix my network?". Return a JSON list of {id,score}.'
# send prompt + candidates to LLM and parse response
# Note: keep candidate size small to limit cost/latency; prefer templates and deterministic scoring when possible.Operational considerations
- Index freshness: Decide near-real-time vs batch ingestion. Near-real-time requires streaming pipelines and smaller, frequent updates; batch ingestion simplifies consistency but delays fresh data.
- Sharding & scale: FAISS + GPU is fast, but scaling writes/read across nodes needs orchestration. Managed vector DBs handle this for you.
- Quantization & compression: For large corpora, use IVF, PQ, or other quantization to save memory at a small accuracy cost.
- Metadata filtering (hybrid search): Combine vector similarity with attribute filters to reduce false positives and enforce ACLs.
- Monitoring: Track recall/precision with labeled queries, latency p95/p99, and embedding failures.
Common pitfalls and how to avoid them
- Mixing metrics: If you normalize for cosine, the index must use inner product on normalized vectors; mixing wrong metric destroys results.
- Chunking too small or too large: Very small chunks lose context; very large ones dilute signal. Aim for 200–800 tokens depending on use-case.
- Noisy retrieval set: Use metadata filters and reranking to avoid returning irrelevant matches.
- Ignoring hallucination: If you use an LLM to generate final answers from retrieved docs, always surface sources and use conservative prompts to reduce invented facts.
When to use what
- Small dataset & relational requirements: pgvector inside Postgres is fast to adopt.
- Large corpus, high throughput: FAISS or Milvus (self-host) or managed vector DBs for scaling and performance.
- Low ops tolerance: Managed services reduce operational burden.
Conclusion
Semantic search is a composable pipeline: choose embeddings that suit your accuracy and cost needs, select a vector store that fits your scale and operations constraints, and add reranking and metadata filtering for robust results. Start with a small local prototype (sentence-transformers + FAISS), measure recall and latency, then iterate to a scalable store and reranker. If you want, use the example code as a base to benchmark models and indexing strategies against your real queries.
Further reading: the Stack Overflow Engineering piece on semantic search offers conceptual context; also review cost/context tradeoffs for models when you add reranking.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment