Why semantic search matters for developers
Semantic search moves beyond keyword matching to find results by meaning. For product search, docs, support knowledge bases, or retrieval-augmented generation (RAG), it raises relevance dramatically when implemented correctly. This guide gives a practical, production-oriented pipeline you can implement and maintain.
Core components
- Text chunking — break long documents into search-sized units with overlap.
- Embeddings — map text fragments and queries to fixed-length vectors.
- Vector database — store vectors and payloads; support nearest-neighbor queries and metadata filters (Qdrant, Milvus, Weaviate, Pinecone, etc.).
- Re-ranking / hybrid search — combine vector results with sparse (BM25) retrieval or an LLM reranker to improve precision.
- Monitoring & evaluation — latency, recall@k, human labels and A/B testing.
Design principles
- Store original text and minimal metadata (source, doc-id, chunk-index, language) as payload.
- Index at chunk granularity (200–1000 tokens) with overlap (10–30%) to preserve context.
- Choose vector DB for operational constraints: memory, disk, scaling, filtering support, cloud-hosted vs self-hosted.
- Keep embeddings and vector DB separate so you can switch embedding models without reworking storage.
Example pipeline (PHP snippets)
Below are practical PHP examples for the three core tasks you will implement: chunking, embedding + upsert, and query + re-rank. Replace API endpoints/keys with your provider settings.
1) Chunk text for indexing
This is a simple token-approximation chunker based on characters/words — replace with tokenizer-based chunking (BPE/byte-level) for more consistent token sizes.
<?php
function chunk_text(string $text, int $chunkSize = 800, int $overlap = 100): array {
$words = preg_split('/\s+/', trim($text));
$chunks = [];
$i = 0;
$n = count($words);
while ($i < $n) {
$chunkWords = array_slice($words, $i, $chunkSize);
$chunks[] = implode(' ', $chunkWords);
$i += ($chunkSize - $overlap);
if ($i < 0) $i = 0; // safety
}
return $chunks;
}
// Example
$text = "Long document text ...";
$chunks = chunk_text($text, 500, 80);
var_dump(count($chunks));
?>2) Create embeddings and upsert to a vector DB (Qdrant example)
This example shows one embedding call (replace with your embedding API) and an upsert into Qdrant via HTTP. Validate your provider's API shape — field names vary.
<?php
$OPENAI_KEY = getenv('EMBED_API_KEY');
$QDRANT_URL = 'https://qdrant.example.local';
$QDRANT_COLLECTION = 'semantic_docs';
function get_embedding(string $text, string $apiKey): array {
$payload = json_encode(["model" => "text-embedding-3-small", "input" => $text]);
$ch = curl_init('https://api.openai.com/v1/embeddings');
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, [
'Content-Type: application/json',
'Authorization: Bearer ' . $apiKey,
]);
curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
$resp = curl_exec($ch);
curl_close($ch);
$data = json_decode($resp, true);
// Defensive checks
if (!isset($data['data'][0]['embedding'])) {
throw new RuntimeException('Embedding API returned unexpected response');
}
return $data['data'][0]['embedding'];
}
function upsert_point(string $qdrantUrl, string $collection, int $id, array $vector, string $textPayload) {
$endpoint = rtrim($qdrantUrl, '/') . '/collections/' . urlencode($collection) . '/points';
$body = json_encode(["points" => [[
"id" => $id,
"vector" => $vector,
"payload" => ["text" => $textPayload]
]]]);
$ch = curl_init($endpoint);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, ['Content-Type: application/json']);
curl_setopt($ch, CURLOPT_POSTFIELDS, $body);
$resp = curl_exec($ch);
$code = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
if ($code >= 400) {
throw new RuntimeException("Qdrant upsert failed: HTTP " . $code . " - " . $resp);
}
return json_decode($resp, true);
}
// Usage: index a chunk
$chunk = "This is a document chunk to index...";
$embedding = get_embedding($chunk, $OPENAI_KEY);
upsert_point($QDRANT_URL, $QDRANT_COLLECTION, 12345, $embedding, $chunk);
?>3) Query + local rerank by cosine similarity
Query flow: embed the query, ask the vector DB for top-N (request vectors), then compute cosine similarity in app code to rerank or to combine with other signals.
<?php
function dot_product(array $a, array $b): float {
$sum = 0.0;
$len = min(count($a), count($b));
for ($i = 0; $i < $len; $i++) $sum += $a[$i] * $b[$i];
return $sum;
}
function norm(array $v): float {
$s = 0.0;
foreach ($v as $val) $s += $val * $val;
return sqrt($s);
}
function cosine_similarity(array $a, array $b): float {
$na = norm($a);
$nb = norm($b);
if ($na == 0 || $nb == 0) return 0.0;
return dot_product($a, $b) / ($na * $nb);
}
function qdrant_search(string $qdrantUrl, string $collection, array $vector, int $top = 20) {
$endpoint = rtrim($qdrantUrl, '/') . '/collections/' . urlencode($collection) . '/points/search';
$body = json_encode([
'vector' => $vector,
'top' => $top,
'with_vectors' => true,
'with_payload' => true
]);
$ch = curl_init($endpoint);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, ['Content-Type: application/json']);
curl_setopt($ch, CURLOPT_POSTFIELDS, $body);
$resp = curl_exec($ch);
curl_close($ch);
return json_decode($resp, true);
}
// Example query flow
$query = "How do I reset my password?";
$queryEmb = get_embedding($query, $OPENAI_KEY);
$results = qdrant_search($QDRANT_URL, $QDRANT_COLLECTION, $queryEmb, 20);
// Rerank locally by exact cosine against stored vectors
$scored = [];
foreach ($results['result'] as $hit) {
$docVector = $hit['vector'];
$score = cosine_similarity($queryEmb, $docVector);
$scored[] = ['id' => $hit['id'], 'score' => $score, 'payload' => $hit['payload']];
}
usort($scored, fn($a,$b) => $b['score'] <> $a['score'] ? ($b['score'] > $a['score'] ? 1 : -1) : 0);
// Top result
$top = array_slice($scored, 0, 5);
foreach ($top as $r) {
echo "Score: " . round($r['score'], 4) . " - " . substr($r['payload']['text'], 0, 200) . "\n\n";
}
?>Practical tradeoffs and tips
- Model cost vs quality: larger embedding models are more accurate but costlier and slower. Use smaller models for high-volume pipelines and reserve large models for reranking or offline reindexing.
- Hybrid search: combine sparse methods (BM25) for exact matches and vector search for semantic matches. That reduces false positives for queries with clear keywords.
- Indexing frequency: near-real-time vs batch. Real-time gives freshness but higher costs; batching improves throughput and lets you reuse embeddings.
- Filtering by metadata: store minimal, queryable metadata (tenant id, language, doc type). Use DB-side filters to reduce candidate sets.
- Reranking approaches: local cosine is cheap and deterministic. LLM rerankers can reorder by nuance but add latency and cost — consider asynchronous reranking for heavy workloads.
Evaluation and monitoring
- Collect labeled queries with ideal results; measure recall@k and precision@k.
- Capture latency percentiles (p50/95/99) for embedding + DB search + rerank.
- Track drift: embedding model updates change vector space; store model version and reindex periodically.
- Log user feedback and clicks to create online learning signals or sampling set for manual review.
Production checklist
- Record embedding model version and vector dimensions in metadata.
- Provide safe fallbacks (keyword search) if embeddings fail or CORS/third-party services are unavailable.
- Limit vector size and normalize vectors server-side to reduce numeric issues.
- Automate reindex tasks and maintain a tested restore path for the vector DB snapshot.
Further reading
Concept-level discussion and taxonomy on semantic search helps when designing evaluation and hybrid strategies — see the Stack Overflow Blog article on semantic search: What (un)exactly do you mean by semantic search?
Vector DB docs (example: Qdrant): https://qdrant.tech/documentation/
Conclusion
Implementing semantic search is a mix of engineering and evaluation. Build a clear pipeline: chunk & embed > store with metadata > retrieve > rerank > monitor. Start with a compact, cost-effective embedding model and local reranking; add LLM rerankers or hybrid retrieval later as your traffic and budgets allow. Pay attention to metadata, model versions, and evaluation so your search stays accurate as data and models evolve.
Note: replace the example endpoints and keys in the snippets with your vendor's endpoints. API shapes differ among embedding providers and vector DBs; treat the code as a pattern to adapt.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment