Why semantic search matters
Semantic search — retrieving content by meaning rather than exact keywords — is now a core feature for documentation, support systems, knowledge bases, and modern search experiences. As the Stack Overflow post "What (un)exactly do you mean by semantic search?" explains, ambiguity in the term leads to mismatched expectations. This article turns that ambiguity into a practical checklist and code-first recipes you can use right away to build and evaluate semantic search systems.
Read the referenced discussion: What (un)exactly do you mean by semantic search?
Core concepts (short)
- Embeddings: fixed-length vectors that encode semantic properties of text.
- Similarity metric: usually cosine similarity or inner product after normalization.
- ANN indexes: approximate nearest neighbor structures (FAISS, HNSW, Annoy) for speed at scale.
- Hybrid search: combine BM25/keyword filtering with vector retrieval to handle exact matches and structured filters.
- Reranking: run a lightweight model or cross-encoder on top-k results to improve precision.
Implementation recipe — minimal, reliable path
- Pick an embedding model that matches your budget and text type (short UI queries vs. long documents).
- Preprocess consistently: normalization, basic cleanup, and optionally chunk long docs.
- Index embeddings into an ANN or a dense-flat index for small datasets.
- At query time, embed the query, retrieve top-k, optionally filter by metadata, then rerank.
- Measure latency, recall@k, and business KPIs (task completion, reduced support tickets) to guide tuning.
Python: quick FAISS prototype (production-ready steps shown)
This example uses sentence-transformers to create embeddings and FAISS for an inner-product index with L2-normalized vectors (cosine similarity).
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
# Example documents
docs = [
"How to reset your password",
"Troubleshooting login errors",
"How to configure SSO with SAML",
"API rate limits and best practices",
]
# 1) Load embedding model
model = SentenceTransformer("all-MiniLM-L6-v2") # small, fast, good baseline
# 2) Create embeddings (batch for large corpora)
embs = model.encode(docs, convert_to_numpy=True)
# 3) Normalize for cosine similarity and build FAISS index
faiss.normalize_L2(embs)
dim = embs.shape[1]
index = faiss.IndexFlatIP(dim) # inner product on normalized vectors == cosine
index.add(embs)
# 4) Query
query = "SSO configuration help"
q_emb = model.encode([query], convert_to_numpy=True)
faiss.normalize_L2(q_emb)
k = 3
D, I = index.search(q_emb, k)
# I contains indices of top-k docs, D contains similarity scores
for rank, idx in enumerate(I[0]):
print(f"rank={rank+1}, score={D[0][rank]:.4f}, doc=\"{docs[idx]}\"")Notes:
- FAISS IndexFlatIP is simple and deterministic; switch to IndexIVFFlat or HNSW for larger datasets to trade memory for speed.
- Normalize vectors to use inner product as cosine similarity; this is numerically stable and efficient in FAISS.
JavaScript: rapid prototype with embeddings and brute-force similarity
Use this pattern for small datasets or to validate customer flows before adopting ANN infrastructure. The example below assumes an embedding endpoint; replace with your provider or local embedding method.
const fetch = require('node-fetch');
async function embed(text) {
const res = await fetch('https://api.openai.com/v1/embeddings', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${process.env.OPENAI_API_KEY}`
},
body: JSON.stringify({ model: 'text-embedding-3-small', input: text })
});
const json = await res.json();
return json.data[0].embedding;
}
function dot(a, b) {
let s = 0;
for (let i = 0; i < a.length; i++) s += a[i] * b[i];
return s;
}
function norm(a) {
return Math.sqrt(dot(a, a));
}
function cosine(a, b) {
return dot(a, b) / (norm(a) * norm(b));
}
(async () => {
const docs = [
'Resetting your password',
'Login failure troubleshooting',
'Configuring single sign-on',
];
// Precompute embeddings (store them persistently in DB)
const docEmbeddings = [];
for (const d of docs) docEmbeddings.push(await embed(d));
// Query
const q = 'How do I set up SSO?';
const qEmb = await embed(q);
const scores = docEmbeddings.map((ve, i) => ({ i, score: cosine(qEmb, ve) }));
scores.sort((a, b) => b.score - a.score);
console.log('Top results:');
for (let r of scores.slice(0, 3)) {
console.log(r.score.toFixed(4), docs[r.i]);
}
})();Practical tips and tradeoffs
- Embedding model choice: smaller models are cheaper and faster but may lose nuance. Evaluate recall@k on representative queries before committing.
- Index type: flat indexes give perfect recall but scale poorly. HNSW/IVF/OPQ reduce latency at the cost of possible missed neighbors — tune recall/latency tradeoffs.
- Hybrid search: combine keyword filters (to enforce exact constraints) or BM25 for short queries, and then apply vector ranking for semantics.
- Reranking: a cross-encoder reranker can dramatically improve precision on the top-k at modest cost — useful when business value is high.
- Chunking long docs: divide long documents into overlapping passages and store per-chunk embeddings. Maintain links back to the source for display and aggregation.
- Monitoring: track latency, recall@k, and business outcomes. Embed drift and model updates need reindexing and A/B tests.
- Privacy & compliance: avoid sending PII to third-party embedding APIs unless permitted; consider on-prem or private-hosted models.
Evaluation checklist
- Define success metrics: recall@k, MRR, user task completion, or reduction in support escalations.
- Build a labeled set of queries and expected documents for automated tests.
- Run A/B tests when changing models, index types, or rerankers to measure user impact.
Conclusion
Semantic search is a spectrum, not a single technology. Start small: pick a compact embedding model, prototype with brute-force or FAISS, and add ANN, hybrid filters, and reranking as you measure gaps. Clear evaluation metrics and consistent preprocessing are the most important levers for predictable improvements.
Key takeaway: define the behavior you expect from "semantic search" in your product context, measure it, and iterate.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment