Why multi-model support matters
Today many teams run into two needs: (1) use several providers or self-hosted models for cost, latency, or policy reasons, and (2) keep a single, stable API surface for the rest of the app. A small adapter layer that is OpenAI-compatible on the outside but can route requests to different backends gives you flexibility without scattering provider-specific code across your codebase. See a short how-to that inspired this approach: How to Add Multi-Model AI to Your App in 5 Minutes (OpenAI-Compatible).
Goals of this pattern
- Keep a single client interface for completions and embeddings
- Allow runtime provider selection (routing rules, cost/latency tradeoffs)
- Support graceful fallback and simple caching
- Make it easy to add new providers (OpenAI, Anthropic, self-hosted Llama, region-specific providers)
Design overview
- Provider adapter — an implementation that translates the unified API into provider-specific HTTP calls.
- Model catalog & routing — metadata describing each provider/model (cost, latency, capabilities), and routing rules to select a provider.
- Fallback & circuit breaker — try cheaper models first, fallback to more reliable ones on error, and mark unhealthy providers.
- Cache & dedupe — cache recent prompts/results and avoid duplicate concurrent requests.
- Embeddings/semantic search — normalize embeddings so your vector store can be queried regardless of which provider produced them.
Minimal PHP implementation (practical)
Below is a compact example showing: a provider interface, an OpenAI-compatible adapter, a MultiModelClient that supports routing/fallback, and a usage example. This is intentionally small — treat it as a template to extend with retries, async workers, and production caching.
<?php
interface ProviderInterface
{
public function name(): string;
public function completion(string $prompt, array $opts = []): array; // returns ["text"=>..., "meta"=>...]
public function embed(array $inputs): array; // returns array of embedding vectors
}
class OpenAIProvider implements ProviderInterface
{
private $endpoint;
private $apiKey;
public function __construct(string $endpoint, string $apiKey)
{
$this->endpoint = rtrim($endpoint, '/');
$this->apiKey = $apiKey;
}
public function name(): string
{
return 'openai-compatible';
}
public function completion(string $prompt, array $opts = []): array
{
$payload = array_merge(['model' => $opts['model'] ?? 'gpt-4o-mini', 'prompt' => $prompt], $opts['extra'] ?? []);
$ch = curl_init($this->endpoint . '/v1/completions');
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, [
'Content-Type: application/json',
'Authorization: Bearer ' . $this->apiKey,
]);
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($payload));
$res = curl_exec($ch);
$err = curl_errno($ch);
curl_close($ch);
if ($err || !$res) {
throw new RuntimeException('Request failed to provider: ' . $this->name());
}
$data = json_decode($res, true);
// Normalize common response shapes
$text = $data['choices'][0]['text'] ?? ($data['choices'][0]['message']['content'] ?? '');
return ['text' => $text, 'raw' => $data];
}
public function embed(array $inputs): array
{
$payload = ['model' => 'text-embedding-3-small', 'input' => $inputs];
$ch = curl_init($this->endpoint . '/v1/embeddings');
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, [
'Content-Type: application/json',
'Authorization: Bearer ' . $this->apiKey,
]);
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($payload));
$res = curl_exec($ch);
curl_close($ch);
$data = json_decode($res, true);
$vectors = [];
foreach ($data['data'] ?? [] as $item) {
$vectors[] = $item['embedding'];
}
return $vectors;
}
}
class MultiModelClient
{
private $providers = [];
private $order = []; // preferred order by name
private $cacheDir;
public function __construct(string $cacheDir = '/tmp/mmclient-cache')
{
if (!is_dir($cacheDir)) mkdir($cacheDir, 0755, true);
$this->cacheDir = $cacheDir;
}
public function registerProvider(ProviderInterface $p)
{
$this->providers[$p->name()] = $p;
$this->order[] = $p->name();
}
public function completion(string $prompt, array $opts = []): array
{
$key = sha1($prompt . json_encode($opts));
$cacheFile = $this->cacheDir . '/' . $key . '.json';
if (file_exists($cacheFile)) {
return json_decode(file_get_contents($cacheFile), true);
}
$preferred = $opts['preferred'] ?? $this->order;
$lastEx = null;
foreach ($preferred as $name) {
if (!isset($this->providers[$name])) continue;
try {
$res = $this->providers[$name]->completion($prompt, $opts);
file_put_contents($cacheFile, json_encode($res));
return $res;
} catch (Throwable $e) {
$lastEx = $e;
// mark/unhealthy logic could go here
continue; // try next provider
}
}
throw $lastEx ?: new RuntimeException('No providers available');
}
public function embed(array $inputs, string $providerName = null): array
{
$providerName = $providerName ?? $this->order[0];
if (!isset($this->providers[$providerName])) {
throw new InvalidArgumentException('Provider not registered: ' . $providerName);
}
return $this->providers[$providerName]->embed($inputs);
}
}
// Usage example
$client = new MultiModelClient();
$client->registerProvider(new OpenAIProvider('https://api.openai.com', getenv('OPENAI_API_KEY')));
// You could register other providers (anthropic, self-hosted local endpoints, region-specific aggregator)
try {
$out = $client->completion("Summarize the following: ...", ['preferred' => ['openai-compatible']]);
echo $out['text'];
} catch (Exception $e) {
error_log('AI request failed: ' . $e->getMessage());
}
?>Notes about the example
- The client normalizes responses to a small shape (text + raw). Extend that normalization to include tokens, usage, finish_reason, etc.
- Use a stronger cache (Redis) and request de-duplication in production to avoid multiple identical requests hitting provider endpoints concurrently.
- Add async request workers for long-running or higher-cost models and return results via webhooks or polling.
Embedding & semantic-search pattern
Keep embeddings provider-agnostic by normalizing vectors to a uniform float array and storing the vector alongside metadata in your vector store. If you don't have pgvector/FAISS, you can store vectors as JSON in a DB and compute a cosine similarity in-worker for small datasets.
<?php
// simple cosine similarity for small result sets
function cosine(array $a, array $b): float
{
$dot = 0.0; $na = 0.0; $nb = 0.0;
for ($i = 0, $l = count($a); $i < $l; $i++) {
$dot += $a[$i] * $b[$i];
$na += $a[$i] * $a[$i];
$nb += $b[$i] * $b[$i];
}
return $dot / (sqrt($na) * sqrt($nb) + 1e-9);
}
// fetch candidate rows from DB (limit), then rank in PHP using cosine()
$needle = $client->embed(["How to use X in Y?"])[0];
$candidates = $db->query('SELECT id, embedding_json, text FROM docs LIMIT 200')->fetchAll();
$ranked = [];
foreach ($candidates as $row) {
$vec = json_decode($row['embedding_json'], true);
$score = cosine($needle, $vec);
$ranked[] = ['id' => $row['id'], 'score' => $score, 'text' => $row['text']];
}
usort($ranked, fn($a,$b) => $b['score'] <> $a['score']);
// use top-k resultsTradeoffs and practical advice
- Simplicity vs accuracy: Querying a single reliable provider is simplest; routing to cheaper/self-hosted models saves cost but may give lower-quality outputs.
- Latency: Local/self-hosted models can reduce latency if close to your infra, but they require ops effort (GPU, scaling). External APIs offload ops but add network latency.
- Consistency: Different models can return different tokenization/stop behaviors. Normalize responses and test critical flows with deterministic prompts.
- Compliance & regional constraints: Some providers have data residency or registration requirements. Keep that metadata in your catalog and enforce routing rules.
- Monitoring: Track provider errors, average latency, and output quality (or human-in-the-loop feedback) and use that to adjust routing automatically.
Where to extend next
- Background workers for expensive calls, with webhook/push results.
- Vector index backed by pgvector, FAISS, or Pinecone for large-scale semantic search.
- Rate limiter and per-provider cost tracking to make automated routing decisions.
- Support streaming completions to forward tokens to clients as they're produced (requires provider streaming support).
Conclusion
A small adapter-based multi-model layer gives immediate benefits: vendor portability, cost/latency optimization, and a single API surface for your app. Start with a simple provider interface, add caching and fallback, and gradually add monitoring, health checks, and vector indexes as you scale. If you need a quick multi-model proof-of-concept, the pattern above is enough to get an OpenAI-compatible façade running and swap providers without changing application code.
Further practical reference: How to Add Multi-Model AI to Your App in 5 Minutes (OpenAI-Compatible) and a self-host Llama guide: Self-Host Llama 2 on a $6/month Droplet.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment