Why semantic search matters now
Keyword search hits limits when queries are paraphrased or rely on meaning rather than exact terms. Semantic search uses vector embeddings and nearest-neighbor search to find results by meaning. This article gives developers practical, evergreen guidance: core concepts, storage options, and two implementation patterns you can run locally or in production.
Core components of a semantic search system
- Document ingestion — split, clean, and store source text (docs, transcripts, web pages).
- Embeddings — map text fragments to numeric vectors using an embedding model (hosted API or open weights).
- Vector index — a vector database or index for fast nearest-neighbor queries (e.g., Milvus, Weaviate, Pinecone, pgvector, Milvus).
- Search layer — query conversion, optional hybrid ranking (BM25 + vector), and result filtering.
- Evaluation — relevance testing (precision@k, recall), latency and cost measurement.
When to pick which embedding model
- Use a hosted embedding API (OpenAI, Anthropic, etc.) for simplicity and consistency in production.
- Use an open-source model (sentence-transformers, Hugging Face) when you need self-hosting, data residency, or lower inference costs at scale.
- Match embedding dimensionality with your vector store: higher dims may increase cost and retrieval time; lower dims can be faster but may lose nuance.
Vector store choices and tradeoffs
- Managed services (Pinecone, Milvus Cloud, Weaviate Cloud) — easy scaling and maintenance; pay-as-you-go but vendor lock-in and cost tradeoffs.
- Self-hosted Milvus / Milvus + MVM — high performance, many indexing options; requires ops resources.
- pgvector (Postgres extension) — great for small to medium workloads and teams already using Postgres; simpler operations, limited at very large scale.
- Hybrid approaches — combine BM25 (text-based) + vector ranking to improve precision on factual queries.
Practical pipeline: Python + OpenAI embeddings + Milvus
This pipeline is suitable for production prototypes or services that want a robust vector index. Steps: split text, embed, upsert to Milvus, query and re-rank.
import os
import openai
from pymilvus import connections, FieldSchema, CollectionSchema, DataType, Collection
# Configure
openai.api_key = os.getenv('OPENAI_API_KEY')
MILVUS_HOST = 'localhost'
MILVUS_PORT = '19530'
COLLECTION_NAME = 'documents'
EMBED_DIM = 1536 # match your embedding model
# Connect to Milvus
connections.connect('default', host=MILVUS_HOST, port=MILVUS_PORT)
# Define schema
fields = [
FieldSchema(name='id', dtype=DataType.INT64, is_primary=True, auto_id=True),
FieldSchema(name='embedding', dtype=DataType.FLOAT_VECTOR, dim=EMBED_DIM),
FieldSchema(name='text', dtype=DataType.VARCHAR, max_length=65535),
]
collection_schema = CollectionSchema(fields, description='semantic docs')
if COLLECTION_NAME not in (c.name for c in Collection.list()):
collection = Collection(COLLECTION_NAME, schema=collection_schema)
else:
collection = Collection(COLLECTION_NAME)
# Function to get embedding from OpenAI
def embed_text(text):
resp = openai.Embedding.create(input=[text], model='text-embedding-3-small')
return resp['data'][0]['embedding']
# Ingest documents
texts = [
'First document text here.',
'Second document text here.'
]
embeddings = [embed_text(t) for t in texts]
collection.insert([embeddings, texts])
# Create index and load
index_params = {"index_type": "IVF_FLAT", "metric_type": "COSINE", "params": {"nlist": 1024}}
collection.create_index('embedding', index_params)
collection.load()
# Query
query_vec = embed_text('Find documents about topic X')
search_params = {"metric_type": "COSINE", "params": {"nprobe": 16}}
results = collection.search([query_vec], 'embedding', search_params, limit=5, output_fields=['text'])
for hits in results:
for hit in hits:
print(hit.score, hit.entity.get('text'))Notes: tune nlist/nprobe for your latency/recall tradeoff; batch embeddings for throughput; keep an eye on embedding API cost.
Lightweight option: Node.js + Postgres + pgvector
When you already use Postgres and need a low-ops path, pgvector is a great fit for small/medium indexes.
// Example Node.js flow (assumes pgvector extension installed)
const { Client } = require('pg')
const fetch = require('node-fetch')
const client = new Client({ connectionString: process.env.DATABASE_URL })
async function init() {
await client.connect()
await client.query("CREATE EXTENSION IF NOT EXISTS vector")
await client.query("CREATE TABLE IF NOT EXISTS docs (id serial primary key, embedding vector(1536), text text)")
}
// Call your embedding provider
async function embedText(text) {
// replace with real API call; ensure you return a JS array of floats
const resp = await fetch('https://api.example.com/embed', { method: 'POST', body: JSON.stringify({text}) })
const j = await resp.json()
return j.embedding
}
async function upsert(text) {
const v = await embedText(text)
await client.query('INSERT INTO docs (embedding, text) VALUES ($1, $2)', [v, text])
}
async function search(queryText, k=5) {
const qv = await embedText(queryText)
// pgvector's distance operator: <-> is Euclidean distance
const res = await client.query('SELECT id, text, embedding <-> $1 AS distance FROM docs ORDER BY distance LIMIT $2', [qv, k])
return res.rows
}Notes: pgvector keeps operations inside Postgres and supports indexing (ivfflat). Good for transactional workloads and simple deployments.
Practical tips, pitfalls and tradeoffs
- Chunking — split long docs into semantically coherent chunks (200–1000 tokens). Too small hurts context; too large blurs signal.
- Hybrid ranking — for many datasets, combine BM25 initial filter then vector re-rank for best precision and lower cost.
- Index tuning — IVF/ANNOY/HNSW choices affect latency and memory. Bench with representative queries.
- Vector dims — 512–1536 are common; higher dims increase storage and compute. Match to model output.
- Recall vs latency — approximate indexes trade accuracy for speed. Set k and search params based on SLA.
- Privacy & compliance — embedding APIs may send data to 3rd parties; consider self-hosting for sensitive content.
Measuring relevance and costs
- Use labeled queries to compute precision@k and mean reciprocal rank for iterative improvement.
- Measure embedding cost per document and per query; batch embeddings to amortize overhead.
- Track index rebuild time and storage consumption as data grows; plan sharding or tiering if needed.
Further reading
See the Stack Overflow blog for a conceptual discussion about different expectations around semantic search: What (un)exactly do you mean by semantic search?
Conclusion
Semantic search is practical today: choose embeddings that fit your data and constraints, pick a vector store according to scale and ops resources, and combine vector results with text-based ranking for best results. Start small with pgvector or a managed vector DB, measure relevance with labeled queries, and iterate indexing parameters for production performance.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment