Why a local LLM gateway?
As more teams use multiple LLM providers and integrate AI into messaging channels (Telegram, Slack, Feishu, etc.), a local gateway gives you a single control plane for routing, security, cost control, telemetry, and prompt shaping. A gateway can run locally or inside your VPC, mediate access to hosted models, and also host on-prem models. This article shows a pragmatic architecture and concrete code patterns to build one.
Core responsibilities
- Provider abstraction: Adapter per model/provider (OpenAI, Gemini, Claude, local runtime).
- Routing: Route by team, model capability, cost, or policy.
- Security: Authentication, authorization, input sanitization, and secrets handling.
- Observability: Structured logs, traces, and usage counters for cost/SLAs.
- Operational features: Retry/backoff, rate limiting, caching, and batching.
High-level architecture
- Public/internal API endpoint(s) — single entry for clients and bots.
- Request router — selects provider adapter, model, and transforms prompts.
- Provider adapters — implement a small interface to call remote or local models.
- Shared services — auth, cache (Redis), rate limiter, metrics.
- Optional: worker queue for long-running jobs or batching.
Design tradeoffs
- Local on-prem models increase privacy but require ops and hardware; hosted models reduce ops burden.
- Caching speeds repeated queries but may introduce staleness; cache only non-sensitive, idempotent outputs.
- Single gateway simplifies control but is a potential central point of failure — run it behind a load balancer and make it horizontally scalable.
Minimal gateway skeleton (Express.js)
The following skeleton demonstrates routing a request to a provider adapter. This example focuses on structure; production systems need error handling, retries, and proper secret management.
const express = require('express');
const bodyParser = require('body-parser');
const ProviderRouter = require('./providerRouter');
const app = express();
app.use(bodyParser.json());
// Simple auth middleware (JWT or API key stub)
app.use((req, res, next) => {
const key = req.get('x-api-key');
if (!key) return res.status(401).json({ error: 'missing api key' });
// validate key & attach tenant info
req.tenant = { id: 'team-123', plan: 'pro' };
next();
});
app.post('/v1/generate', async (req, res) => {
const { provider, model, input } = req.body;
try {
const output = await ProviderRouter.dispatch(req.tenant, { provider, model, input });
res.json({ ok: true, output });
} catch (err) {
console.error('gateway error', err);
res.status(500).json({ ok: false, error: err.message });
}
});
app.listen(8080, () => console.log('LLM gateway listening on 8080'));Adapter pattern
Each provider implements a small interface: generate(tenant, model, prompt, opts). The gateway only depends on that interface, making it easy to add new providers or swap implementations.
// provider/openaiAdapter.js
const fetch = require('node-fetch');
class OpenAIAdapter {
constructor(apiKey) { this.apiKey = apiKey; }
async generate(tenant, model, prompt, opts = {}) {
const url = `https://api.openai.com/v1/models/${model}/completions`;
const res = await fetch(url, {
method: 'POST',
headers: { 'Authorization': `Bearer ${this.apiKey}`, 'Content-Type': 'application/json' },
body: JSON.stringify({ prompt, max_tokens: opts.maxTokens || 256 })
});
const data = await res.json();
if (!res.ok) throw new Error(data.error?.message || 'openai error');
return data.choices?.[0]?.text || '';
}
}
module.exports = OpenAIAdapter;// provider/localLlamaAdapter.js (example: spawn a local process like llama.cpp)
const { spawn } = require('child_process');
class LocalLlamaAdapter {
constructor(binaryPath) { this.binaryPath = binaryPath; }
generate(tenant, model, prompt, opts = {}) {
return new Promise((resolve, reject) => {
const p = spawn(this.binaryPath, ['--model', model, '--prompt', prompt]);
let out = '';
p.stdout.on('data', d => out += d.toString());
p.stderr.on('data', d => console.error('llama stderr', d.toString()));
p.on('close', code => code === 0 ? resolve(out) : reject(new Error('local model failed')));
});
}
}
module.exports = LocalLlamaAdapter;Provider router and policy
Router selects an adapter based on tenant policy, model capability, and cost. Keep routing logic small and declarative.
// providerRouter.js (simplified)
const OpenAIAdapter = require('./provider/openaiAdapter');
const LocalLlamaAdapter = require('./provider/localLlamaAdapter');
const adapters = {
'openai': new OpenAIAdapter(process.env.OPENAI_KEY),
'local': new LocalLlamaAdapter('/usr/local/bin/llama')
};
module.exports = {
async dispatch(tenant, { provider, model, input }) {
// simple policy: prefer local for private teams
const chosen = tenant.id.startsWith('private') ? 'local' : (provider || 'openai');
const adapter = adapters[chosen];
if (!adapter) throw new Error('no adapter');
// prompt shaping example
const prompt = `[tenant:${tenant.id}] ${input}`;
return adapter.generate(tenant, model, prompt);
}
};Caching and idempotency
Cache responses for deterministic prompts to reduce cost and latency. Only cache non-sensitive outputs and honor tenant policies.
// caching pseudo
const redis = require('redis');
const client = redis.createClient();
async function cachedGenerate(key, fn, ttl = 60) {
const cached = await client.get(key);
if (cached) return JSON.parse(cached);
const result = await fn();
await client.setEx(key, ttl, JSON.stringify(result));
return result;
}
// usage inside router: const key = `resp:${hash(prompt)}:`+model; return cachedGenerate(key, () => adapter.generate(...));Rate limiting and quotas
Implement per-tenant and per-model limits to avoid runaway costs. Libraries like express-rate-limit or token-bucket algorithms are fine for simple setups. For robust enforcement use a centralized quota store (Redis) and a leaky-bucket algorithm.
// rate limit sketch (token bucket using Redis)
// On each request: LUA script to atomically check & decrement tokens, return allowed flag.
// If not allowed, return 429 with retry-after.Security and privacy
- Use mutual TLS or VPC-only endpoints for internal traffic.
- Rotate provider secrets and never embed them in client code.
- Sanitize prompt inputs if you forward user-supplied content to external providers.
- Mask or redact sensitive fields before caching or logging.
Observability and cost control
Emit structured telemetry: tenant, model, tokens, latency, provider. Use those metrics for automated routing (e.g., failover to cheaper models during spikes) and billing reconciliation.
Operational tips
- Start small: implement a single provider adapter and add features iteratively.
- Fail open vs fail closed: choose policy per tenant — fail closed for private data, fail open for low-criticality chatbots.
- Testing: unit-test adapters with recorded fixtures and integration-test with a staging model endpoint.
- Local development: mock adapters so devs don't consume paid quota during testing.
Example: prompt templating and safety wrapper
Templating helps standardize system messages and safety checks. Keep templates versioned so you can A/B test system prompts without code changes.
// simple template function
function renderTemplate(template, vars) {
return template.replace(/\{\{(\w+)\}\}/g, (_, k) => vars[k] || '');
}
const sys = 'You are an assistant for tenant {{tenantId}}. Keep answers short.';
const prompt = renderTemplate(sys, { tenantId: 'team-123' }) + '\nUser: ' + 'Explain event sourcing.';Concise checklist before production
- Secrets stored in a secret manager, not in code.
- Per-tenant authentication and authorization enforced.
- Rate limits, quotas, and cost alerts configured.
- Cache policy and data retention rules defined.
- Monitoring, logging, and tracing in place for root cause analysis.
Conclusion
A local LLM gateway centralizes control over model selection, security, and cost while enabling consistent prompt shaping and telemetry. Start with a small adapter interface, add caching, rate limiting, and observability, and iterate. The adapter pattern keeps the gateway flexible: add new providers or local runtimes without rewriting routing logic.
Further reading: the original project that motivated this pattern explores a local gateway for multiple models and messaging integrations.
I Wanted One Local Gateway for Claude Code, Codex, Gemini, Telegram, Feishu, and DingTalk — CliGate
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment