Sechno
Machine Learning

Practical Semantic Search for Developers: Build an Embedding-Based Search Pipeline

A compact, practical guide to implementing semantic search: embeddings, chunking, vector stores (FAISS), query flows, evaluation metrics, and production tradeoffs.

SSechno Team 5 min read 30 views
Practical Semantic Search for Developers: Build an Embedding-Based Search Pipeline

Semantic search finds content by meaning rather than exact keyword matches. Instead of string matching, it compares numeric representations (embeddings) of queries and documents and returns the nearest items in vector space. This approach powers document search, knowledge retrieval for LLMs, and more.

For background reading, see the Stack Overflow Blog overview: What (un)exactly do you mean by semantic search?

Core components of a semantic-search system

  • Embedding model: converts text (or other modalities) to vectors.
  • Text chunker & preprocessor: splits long documents into retrievable passages and preserves metadata.
  • Vector store / index: FAISS, Milvus, Weaviate, Pinecone, etc., for nearest-neighbor search.
  • Retriever + Reranker: initial nearest-neighbor search followed by reranking (model or heuristics).
  • Evaluation & monitoring: precision@k, MRR, human checks and drift detection.

Step-by-step implementation advice

Choose an embedding model

Pick between hosted embeddings (OpenAI, Anthropic, Cohere) for simplicity, or open-source/smaller models if you need on-prem privacy or lower inference cost. Consider vector dimension and latency: higher-dim vectors often improve accuracy but increase storage and search time.

Chunking & preprocessing

  • Chunk length: 200–1000 tokens is common. For long-form text, use overlapping windows (e.g., 50–200 token overlap) to preserve context.
  • Store metadata: source id, paragraph index, original offsets — this enables precise highlighting and provenance.
  • Normalize text: trim whitespace, unify unicode, optionally lowercase (model-specific).

Vector store and index selection

Local options: FAISS (great for prototyping and cost-effective on VMs), Annoy, hnswlib. Managed options: Pinecone, Milvus Cloud, Weaviate, Qdrant. Choose based on:

  • Write/update pattern: frequent upserts favor hnsw or vector DBs with efficient dynamic updates.
  • Scale & latency targets: for millisecond queries at large scale, prefer managed services or optimized ANN indexes.
  • Disk vs memory: FAISS has disk-backed indexes but often best performance is in RAM.

Minimal indexing example (Python + FAISS)

Below is a compact pipeline: chunk text, create embeddings, normalize vectors, build a simple FAISS index, and store metadata in a parallel array. Replace the embedding call with your provider.

"
from openai import OpenAI
import faiss
import numpy as np

# Initialize embedding client (adjust to your SDK)
client = OpenAI()

def embed(text):
    # Use single-input batching for clarity; batch in production
    r = client.embeddings.create(model='text-embedding-3-small', input=text)
    return np.array(r.data[0].embedding, dtype='float32')

# Example documents (id, text)
docs = [
    ('doc1', 'First long document text ...'),
    ('doc2', 'Another document text ...'),
]

# Chunk documents (simple split, use tokenizer for token-accurate chunks)
chunks = []
metadatas = []
for doc_id, text in docs:
    parts = [text[i:i+500] for i in range(0, len(text), 500)]
    for idx, part in enumerate(parts):
        chunks.append(part)
        metadatas.append({'doc_id': doc_id, 'chunk_index': idx})

# Embed chunks in batch (batching is recommended)
embeddings = np.stack([embed(t) for t in chunks])

# Normalize to unit vectors if using inner product as cosine proxy
faiss.normalize_L2(embeddings)

# Build index
dim = embeddings.shape[1]
index = faiss.IndexFlatIP(dim)  # inner product on normalized vectors approximates cosine similarity
index.add(embeddings)

# Persist index and metadatas as needed (not shown)
"

Query pipeline

Take a query, embed, normalize, search nearest neighbors, then optionally rerank or call a reranker model for final ordering.

"
# Query example
query = 'How do I reset my password?'
qv = embed(query).reshape(1, -1)
faiss.normalize_L2(qv)

k = 5
D, I = index.search(qv, k)  # D: scores, I: indexes into embeddings/metadatas

results = []
for score, idx in zip(D[0], I[0]):
    md = metadatas[idx]
    snippet = chunks[idx]
    results.append({'score': float(score), 'metadata': md, 'text': snippet})

# Post-processing: threshold by score, rerank with a cross-encoder, or assemble for RAG
"

Evaluation & validation (don’t skip this)

  • Metrics: precision@k, recall@k, mean reciprocal rank (MRR), and human relevance checks.
  • Unit tests: synthetic queries with expected source IDs to detect regressions during model or data updates.
  • Monitor drift: embedding model updates can change vector distributions — reindex or perform compatibility tests before rollout.

Tradeoffs and practical tips

  • Approximate vs exact search: ANN (HNSW, IVF) trades a bit of recall for big speed gains at scale.
  • Storage and cost: Higher dimensional vectors increase RAM/storage and search time. Consider PCA or smaller embedding models if cost is high.
  • Freshness: If documents change frequently, prefer indexes that support efficient upserts (HNSW or a managed DB) and keep metadata stores independent from the vector index.
  • Privacy: If embeddings are sent to a third party, evaluate data leakage risk. On-prem inference with open models can reduce exposure.
  • Reranking: A cheap lexical filter or a lightweight cross-encoder reranker often improves results over raw nearest neighbors.
  • Provenance: Always surface source IDs and offsets alongside answers to help users verify results.

Quick checklist before production

  1. Decide embedding model and benchmark speed/quality.
  2. Design chunking strategy with provenance metadata.
  3. Choose vector store and index type based on update pattern and SLA.
  4. Implement evaluation harness (precision@k, MRR) and synthetic tests.
  5. Plan monitoring: query latency, distribution of similarity scores, and human QA sampling.

Conclusion

Semantic search turns text into vectors, enabling meaning-based retrieval that's invaluable for search, RAG, and knowledge assistants. Start small with a clear chunking strategy, a reliable embedding model, and a simple FAISS index for prototyping. Add reranking, evaluation, and production-grade vector DBs as you scale. Validating quality and tracking drift are as important as picking models: better data and measurement often beat chasing slightly higher-performing embeddings.

Next steps: prototype with a small dataset and iterate — measure precision@5, add a lightweight reranker, and test reindexing compatibility when you swap embedding models.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment