Sechno
Backend

Semantic Search for Developers: Build a Practical Vector Retrieval Pipeline with FAISS and Embeddings

A hands-on guide to building a production-ready semantic search pipeline: chunking, embeddings, FAISS indexing, hybrid sparse+dense ranking, metadata filtering, and deployment tips for reliable retrieval.

SSechno Team 5 min read 43 views
Semantic Search for Developers: Build a Practical Vector Retrieval Pipeline with FAISS and Embeddings

Overview

Semantic search matches meaning rather than literal keywords by comparing dense embeddings. This guide shows a practical end-to-end pipeline you can implement today: chunk documents, compute embeddings, build a vector index (FAISS), add optional sparse (TF‑IDF) signals for hybrid ranking, and serve filtered, ranked results.

  • Good: question-answering over documentation, retrieval-augmented generation (RAG), fuzzy matching across long-form content.
  • Not ideal: exact-match legal or financial terms where keyword precision and boolean search are required.

Core components

  • Chunking: split long documents into retrieval-sized pieces.
  • Embeddings: encode chunks and queries into vectors (local or API models).
  • Vector index: FAISS, Milvus, Weaviate, Pinecone, etc.
  • Optional sparse index: TF‑IDF or BM25 to combine lexical signals.
  • Metadata store: map vector ids to text, document id, timestamps, permissions.
  • Retriever: search, filter, and return ranked candidates to your app or LLM.

Minimal example: chunk → embed → FAISS

Below is a compact Python example that demonstrates chunking text, computing embeddings with sentence-transformers, building a FAISS index, and a simple search function. This is intentionally minimal; production systems will add batching, persistence, and async I/O.

from typing import List
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
 
model = SentenceTransformer('all-MiniLM-L6-v2')  # lightweight, fast encoder
 
# 1) Chunking
def chunk_text(text: str, chunk_size: int = 200, overlap: int = 40) -> List[str]:
    tokens = text.split()
    chunks = []
    start = 0
    while start < len(tokens):
        end = min(start + chunk_size, len(tokens))
        chunks.append(' '.join(tokens[start:end]))
        if end == len(tokens):
            break
        start = max(end - overlap, 0)
    return chunks
 
# Example document
text = """Long document text goes here..."""
chunks = chunk_text(text)
 
# 2) Embeddings (batch)
embeddings = model.encode(chunks, show_progress_bar=False, convert_to_numpy=True)
 
# 3) Build FAISS index (cosine via normalized inner product)
d = embeddings.shape[1]
faiss.normalize_L2(embeddings)  # normalize for cosine similarity
index = faiss.IndexFlatIP(d)  # for small datasets; swap with IVF/HNSW for large
index.add(embeddings)
 
# simple id->metadata map
id_to_meta = {i: {'text': chunks[i], 'doc_id': 'doc1', 'chunk_id': i} for i in range(len(chunks))}
 
# 4) Search function
def search(query: str, k: int = 5):
    q_emb = model.encode([query], convert_to_numpy=True)
    faiss.normalize_L2(q_emb)
    D, I = index.search(q_emb, k)
    results = []
    for score, idx in zip(D[0], I[0]):
        meta = id_to_meta.get(int(idx), {})
        results.append({'score': float(score), 'text': meta.get('text'), 'meta': meta})
    return results
 
# usage
print(search('How do I authenticate?'))

Hybrid ranking: add a sparse signal

Combining vector similarity with TF‑IDF or BM25 often improves relevance for keyword-heavy queries. The example below shows a simple TF‑IDF + dense combination.

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.preprocessing import normalize
 
# build TF-IDF over the same chunks
tf = TfidfVectorizer().fit_transform(chunks)  # shape (n_chunks, n_terms)
 
# pre-normalize dense vectors if not already
dense_matrix = embeddings.copy()
# embeddings already normalized above; ensure 2D
 
# combined search
def hybrid_search(query: str, k: int = 5, alpha: float = 0.6):
    # dense score
    q_emb = model.encode([query], convert_to_numpy=True)
    faiss.normalize_L2(q_emb)
    D, I = index.search(q_emb, k*3)  # fetch more from dense
    dense_candidates = list(I[0])
 
    # sparse scores
    q_tfidf = TfidfVectorizer().fit(chunks + [query]).transform([query])
    sparse_scores = (tf @ q_tfidf.T).toarray().ravel()
 
    # combine and rank candidate set
    candidates = set(dense_candidates)
    scored = []
    for idx in candidates:
        dense_score = 0.0
        # approximate dense score by checking position in I
        if idx in I[0]:
            pos = list(I[0]).index(idx)
            dense_score = float(D[0][pos])
        sparse_score = float(sparse_scores[int(idx)])
        # normalize sparse (optional) and combine
        final = alpha * dense_score + (1 - alpha) * sparse_score
        scored.append((final, int(idx)))
    scored.sort(reverse=True)
    return [{'score': s, 'text': id_to_meta[i]['text']} for s, i in scored[:k]]
 
# usage
print(hybrid_search('authenticate with API'))

Metadata filtering

Store metadata in a separate key-value store or alongside your vector DB. Filter candidate ids before scoring to enforce permissions, language, or document date.

Tip: keep vector-only operations in the vector DB and apply strict filters in the metadata layer to avoid leaking unauthorized documents.

Tradeoffs and performance tips

  • Index type: IndexFlatIP is exact but memory-heavy. For larger collections use IVF+PQ or HNSW. Use FAISS IVFPQ for disk‑efficient indexes and HNSW for fast recall at low latency.
  • Dimensionality: smaller embedding models (e.g., 384 dims) use less memory and often perform well for many tasks. Larger models may increase semantic fidelity but cost CPU/GPU time.
  • Chunk size: short chunks improve precision but increase index size. Aim for ~150–400 tokens depending on your retrieval needs.
  • Freshness & updates: real-time insertion is supported by HNSW and many hosted vector DBs. For IVF indexes, plan periodic reindexing or use hybrid approaches for near-real-time data.
  • Embedding cost: precompute and cache embeddings for static content. For dynamic content, batch embedding requests and use GPU inference when possible.
  • Hybrid vs pure dense: Hybrid (dense + sparse/BM25) improves precision on keyword queries and is often a practical compromise.
  • Privacy and compliance: if embeddings are sent to third-party APIs, ensure you follow data policies and consider private/local models for sensitive content.

Deployment and scaling checklist

  1. Precompute embeddings for static documents and store them alongside metadata.
  2. Use batching and GPU for embedding large ingests.
  3. Choose an index type that matches read/write and latency targets (HNSW for low-latency, IVF/PQ for low-memory).
  4. Expose a retriever API that returns candidate IDs + metadata, not raw vectors.
  5. Monitor recall and relevance with labeled queries and iteratively tune chunking, candidate pool size, and alpha in hybrid ranking.

Further reading

For a discussion of what "semantic search" means in practice, see the Stack Overflow blog post: What (un)exactly do you mean by semantic search?

FAISS docs: FAISS on GitHub. SentenceTransformers: sbert.net.

Concise conclusion

Semantic search is a practical technique when you need meaning-aware retrieval. Start small: chunk your docs, precompute embeddings, index with FAISS, and add a sparse lexical signal for hybrid ranking. Monitor relevance, tune chunk size and candidate counts, and choose index types that balance cost, latency, and update needs. This pipeline will serve as a solid foundation for RAG, knowledge search, and many developer-facing search features.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment