Sechno
Architecture

Practical Semantic Search for Developers: A Minimal PHP Implementation and Architecture Guide

Implement a practical semantic search pipeline: embeddings, local vector index (JSON), cosine similarity search in PHP, and guidance to scale with vector databases. Actionable code, tradeoffs, and deployment tips.

SSechno Team 5 min read 72 views
Practical Semantic Search for Developers: A Minimal PHP Implementation and Architecture Guide

What is semantic search (brief)

Semantic search matches meaning, not just exact words. Instead of relying on keywords, you convert text into numeric embeddings and rank results by vector similarity. This post shows a minimal, practical pipeline in PHP you can use for prototypes and explains how to scale responsibly.

  • Use it for fuzzy intent matching, FAQ retrieval, code search, and paraphrase-tolerant queries.
  • Prefer keyword search for exact matches, boolean filters, or when index size and cost are severely constrained.

Core architecture (practical overview)

  1. Embed: use an embeddings provider (OpenAI, Cohere, Mistral, Sentence Transformers, etc.) to convert text to vectors.
  2. Store: store vectors and metadata in a vector store (local file for prototypes, Postgres+pgvector, Pinecone, Weaviate, Milvus for scale).
  3. Retrieve: compute similarity between query embedding and stored vectors; return top-K candidates.
  4. Rerank/Display: optionally rerank candidates with a lightweight model or heuristic, then present results with context snippets.

Practical tips

  • Chunk long documents (200–500 tokens) and keep chunk-level metadata to surface relevant context.
  • Normalize embeddings if your similarity method assumes normalized vectors (cosine similarity prefers normalized vectors, but many vector DBs provide operators).
  • Cache frequent query embeddings and use batching when requesting embeddings for many chunks.
  • Monitor cost: embedding large corpora can be expensive — consider open-source embedding models for heavy offline indexing.

Minimal, runnable PHP prototype (embeddings + local JSON index)

The example below uses a generic HTTP embedding API (replace endpoint and key for your provider). It demonstrates: requesting embeddings, saving an index to disk, computing cosine similarity in PHP, and returning top-K matches. This approach is great for prototyping or small datasets.

<?php
// Minimal semantic search prototype in PHP
// Replace EMBEDDING_API_URL and EMBEDDING_API_KEY with your provider values.
 
function get_embedding(string $text): array {
    $apiUrl = getenv('EMBEDDING_API_URL') ?: 'https://api.example.com/v1/embeddings';
    $apiKey = getenv('EMBEDDING_API_KEY') ?: 'YOUR_API_KEY';
 
    $payload = json_encode(['input' => $text]);
 
    $ch = curl_init($apiUrl);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
    curl_setopt($ch, CURLOPT_HTTPHEADER, [
        'Content-Type: application/json',
        'Authorization: Bearer ' . $apiKey,
    ]);
    curl_setopt($ch, CURLOPT_POST, true);
    curl_setopt($ch, CURLOPT_POSTFIELDS, $payload);
 
    $resp = curl_exec($ch);
    if ($resp === false) {
        throw new RuntimeException('Embedding request failed: ' . curl_error($ch));
    }
    curl_close($ch);
 
    $data = json_decode($resp, true);
    // Adjust according to your provider's response shape; this is a common pattern.
    if (isset($data['data'][0]['embedding'])) {
        return $data['data'][0]['embedding'];
    }
    throw new RuntimeException('Unexpected embedding response: ' . $resp);
}
 
function cosine_similarity(array $a, array $b): float {
    $dot = 0.0;
    $na = 0.0;
    $nb = 0.0;
    $len = min(count($a), count($b));
    for ($i = 0; $i < $len; $i++) {
        $dot += $a[$i] * $b[$i];
        $na += $a[$i] * $a[$i];
        $nb += $b[$i] * $b[$i];
    }
    $denom = sqrt($na) * sqrt($nb);
    if ($denom == 0.0) return 0.0;
    return $dot / $denom;
}
 
function index_documents(array $docs, string $indexPath): void {
    $indexed = [];
    foreach ($docs as $doc) {
        $embed = get_embedding($doc['text']);
        $indexed[] = [
            'id' => $doc['id'],
            'text' => $doc['text'],
            'embedding' => $embed,
        ];
        // be kind to API rate limits
        usleep(100000);
    }
    file_put_contents($indexPath, json_encode($indexed));
}
 
function load_index(string $indexPath): array {
    $json = file_get_contents($indexPath);
    return json_decode($json, true) ?: [];
}
 
function search_index(string $query, array $index, int $k = 5): array {
    $qemb = get_embedding($query);
    $scores = [];
    foreach ($index as $item) {
        $sim = cosine_similarity($qemb, $item['embedding']);
        $scores[] = ['id' => $item['id'], 'text' => $item['text'], 'score' => $sim];
    }
    usort($scores, function($a, $b) { return $b['score'] <=> $a['score']; });
    return array_slice($scores, 0, $k);
}
 
// Usage example:
$docs = [
    ['id' => 1, 'text' => 'How to reset a Docker container'],
    ['id' => 2, 'text' => 'Deploy Laravel app to production'],
    ['id' => 3, 'text' => 'Introduction to vector databases'],
];
$indexPath = __DIR__ . '/semantic_index.json';
 
// First-time: index documents
// index_documents($docs, $indexPath);
 
// Later: load and search
$index = load_index($indexPath);
$results = search_index('where to deploy a PHP app', $index, 3);
 
foreach ($results as $r) {
    echo "ID: {$r['id']} - score: " . round($r['score'], 4) . "\n";
    echo "Snippet: " . substr($r['text'], 0, 120) . "\n\n";
}
?>

Actionable next steps to scale

  • For small datasets, the JSON index approach is simple and fast to iterate on.
  • When dataset or traffic grows, move to a vector-enabled DB: Postgres+pgvector for self-hosted, or managed stores (Pinecone, Milvus, Weaviate) for production features like ANN indexing and metadata filters.
  • Batch embedding requests and run offline indexing jobs. Keep a delta-index for recently added docs to avoid re-embedding the entire corpus.
  • Use hybrid search: combine a fast keyword filter with vector ranking to reduce candidate set and improve precision.
  • Evaluate with human-labeled queries and measure precision@K and latency to choose the right ANN configuration (HNSW, IVF, etc.).

Tradeoffs and caveats

  • Cost vs control: managed vector DBs save ops time but add recurring cost; self-hosting gives control but requires tuning and monitoring.
  • Embedding quality varies: pick an embedding model aligned to your domain (e.g., code embeddings for code search).
  • Latency: local brute-force scoring works for hundreds to low thousands of vectors; use ANN indices for larger scales to meet sub-100ms targets.
  • Privacy: embedding providers may retain data unless you choose a provider and plan that supports private processing.

Concise conclusion

Start small: use a provider for embeddings and a simple local index to validate the UX. Measure retrieval quality and latency, then move to a vector store and batching when you need scale. The core techniques—chunking, caching, hybrid retrieval, and evaluation—stay consistent as you grow.

Further reading

  • Investigate embedding model options and vector DB documentation before committing to a provider.
  • When moving to production, add monitoring for query latency, index staleness, and retrieval quality.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment