Sechno
Ai & Machine Learning

Semantic Search for Engineers: Practical Guide, Code & Tradeoffs

A hands-on guide to building and evaluating semantic search: embeddings, vector indexes (FAISS/HNSW), chunking, hybrid retrieval, evaluation metrics and production tips with runnable Python examples.

SSechno Team 5 min read 85 views
Semantic Search for Engineers: Practical Guide, Code & Tradeoffs

Semantic search returns results based on meaning rather than keyword overlap. Instead of matching text by exact tokens, it maps queries and documents into vector embeddings and retrieves nearest neighbors in vector space. This approach improves relevance for paraphrases, synonyms and conceptual matches.

Core components

  • Embeddings: dense vectors representing text semantics (models from sentence-transformers, OpenAI-style APIs, etc.).
  • Vector store / ANN index: FAISS, HNSWlib, Milvus, Pinecone and others store vectors and return nearest neighbors quickly.
  • Document preprocessing & chunking: split long docs into passages with overlap, keep metadata (source, position).
  • Retrieval pipeline: embedding & search → optional re-ranking (cross-encoder) → downstream use (UI, RAG, analytics).

Quick, runnable Python example (sentence-transformers + FAISS)

This minimal example shows document embedding, L2-normalization for cosine similarity, and nearest-neighbor retrieval with FAISS. It is intended for prototypes and local testing.

from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
 
# Sample corpus
docs = [
    "How to reset your password on example.com",
    "Troubleshooting login issues and token expiry",
    "Guide to setting up two-factor authentication",
    "Billing and subscription management FAQ",
    "Deploying Python apps with Docker and CI/CD"
]
 
# Load a compact embedding model (good default for many tasks)
model = SentenceTransformer('all-MiniLM-L6-v2')
 
# Compute embeddings as numpy arrays
embeddings = model.encode(docs, convert_to_numpy=True, show_progress_bar=False)
 
# Normalize for cosine similarity
faiss.normalize_L2(embeddings)
 
# Build an inner-product index (cosine on normalized vectors)
dim = embeddings.shape[1]
index = faiss.IndexFlatIP(dim)  # exact search; switch to an ANN index for scale
index.add(embeddings)
 
# Query
query = "how do I change my password"
q_emb = model.encode([query], convert_to_numpy=True)
faiss.normalize_L2(q_emb)
 
k = 3
scores, ids = index.search(q_emb, k)
 
print("Top results:")
for score, idx in zip(scores[0], ids[0]):
    print(f"score={score:.4f}\tsource=doc[{idx}]\ttext={docs[idx]}")

Notes:

  • IndexFlatIP does exact search; for large corpora use FAISS HNSW/IVF or HNSWlib for approximate nearest neighbors (faster & smaller memory).
  • Normalize embeddings to use inner product as a cosine proxy: faiss.normalize_L2().
  • Store metadata (doc id, source, offset) separately to avoid re-encoding text for display.

Chunking and metadata best practices

  • Chunk size: 200–500 tokens per chunk is a practical balance; use domain knowledge (code vs. legal text differs).
  • Overlap: 20–30% overlap helps preserve context across chunk boundaries for better retrieval.
  • Store metadata: source URI, document id, chunk index, timestamps and any filters (language, region) to enable hybrid filtering.
  • Keep original document text accessible for re-ranking and answer extraction (don't store only vectors).

Hybrid search: combine lexical and semantic

Pure vector search can miss exact matches (IDs, code snippets). A hybrid approach combines BM25/Elasticsearch with vector scoring. Typical merge strategies:

  1. Retrieve top N candidates from lexical and vector searches, deduplicate, then re-rank.
  2. Weighted score: final_score = alpha * lexical_score + (1 - alpha) * vector_score (tune alpha).
  3. Use metadata filters first (date range, language) to reduce candidate set then apply vectors.

Evaluation & metrics

  • Recall@k / Precision@k: measures if relevant items appear within top-k results.
  • MRR (Mean Reciprocal Rank): useful when a single correct answer exists.
  • Latency & throughput: measure end-to-end embedding + search time for SLOs.
  • Human relevance labeling: build a small labeled set to compare models and index settings.

Production considerations and tradeoffs

  • Index size vs recall: denser indexes (IVF with large clusters) reduce memory but can reduce recall. ANN algorithms trade some accuracy for speed.
  • Embedding drift: when you change embedding models, you must re-embed the corpus or use dual-index strategies to avoid mismatches.
  • Upserts and freshness: choose vector DBs that support incremental inserts and deletes; re-indexing large corpora is expensive.
  • Cost: hosted vector DBs simplify ops but add monthly cost; self-hosted FAISS/HNSWlib is cheaper but requires infra and scaling work.
  • Privacy & compliance: ensure embeddings and text handling comply with data policies (PII, retention, encryption-at-rest/ transit).

Minimal JavaScript cosine example (in-memory)

Small utility for local testing. Not for high-scale production but useful to understand cosine scoring.

// Simple cosine similarity for small in-memory collections
function dot(a, b) {
  return a.reduce((s, v, i) => s + v * b[i], 0);
}
function norm(a) {
  return Math.sqrt(dot(a, a));
}
function cosine(a, b) {
  return dot(a, b) / (norm(a) * norm(b) + 1e-12);
}
 
const corpus = [
  {id: 1, vec: [0.1, 0.2, 0.3]},
  {id: 2, vec: [0.0, 0.9, 0.1]},
];
const q = [0.05, 0.2, 0.25];
 
const scored = corpus.map(d => ({id: d.id, score: cosine(q, d.vec)}));
scored.sort((a, b) => b.score - a.score);
console.log(scored);

Troubleshooting common failure modes

  • Low recall: try larger chunk sizes, reduce embedding pooling lossy steps, or use a different model.
  • High latency: switch to ANN index, pre-compute and cache embeddings, or use batched embedding calls.
  • Hallucinations in RAG: ensure retrieved chunks contain factual grounding and use a conservative generation strategy with citations.

Resources

Conclusion

Semantic search is a practical, well-understood pattern: embed, index, retrieve, and re-rank. Start small with a robust embedding model and an exact or small ANN index to validate relevance. Measure with recall@k and latency, and iterate on chunking, model choice, and hybrid strategies. For production, plan for embedding versioning, incremental updates and a clear evaluation set so you can quantify improvements.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment