Sechno
Architecture

How to Build a Local LLM Gateway: adapters, routing, caching, and security

A practical guide for developers to design and implement a local gateway that unifies multiple LLM providers and chat/messaging integrations. Includes adapter patterns, routing, caching, rate limits, and example code to get you started.

SSechno Team 6 min read 75 views
How to Build a Local LLM Gateway: adapters, routing, caching, and security

Why a local LLM gateway?

As more teams use multiple LLM providers and integrate AI into messaging channels (Telegram, Slack, Feishu, etc.), a local gateway gives you a single control plane for routing, security, cost control, telemetry, and prompt shaping. A gateway can run locally or inside your VPC, mediate access to hosted models, and also host on-prem models. This article shows a pragmatic architecture and concrete code patterns to build one.

Core responsibilities

  • Provider abstraction: Adapter per model/provider (OpenAI, Gemini, Claude, local runtime).
  • Routing: Route by team, model capability, cost, or policy.
  • Security: Authentication, authorization, input sanitization, and secrets handling.
  • Observability: Structured logs, traces, and usage counters for cost/SLAs.
  • Operational features: Retry/backoff, rate limiting, caching, and batching.

High-level architecture

  1. Public/internal API endpoint(s) — single entry for clients and bots.
  2. Request router — selects provider adapter, model, and transforms prompts.
  3. Provider adapters — implement a small interface to call remote or local models.
  4. Shared services — auth, cache (Redis), rate limiter, metrics.
  5. Optional: worker queue for long-running jobs or batching.

Design tradeoffs

  • Local on-prem models increase privacy but require ops and hardware; hosted models reduce ops burden.
  • Caching speeds repeated queries but may introduce staleness; cache only non-sensitive, idempotent outputs.
  • Single gateway simplifies control but is a potential central point of failure — run it behind a load balancer and make it horizontally scalable.

Minimal gateway skeleton (Express.js)

The following skeleton demonstrates routing a request to a provider adapter. This example focuses on structure; production systems need error handling, retries, and proper secret management.

const express = require('express');
const bodyParser = require('body-parser');
const ProviderRouter = require('./providerRouter');
 
const app = express();
app.use(bodyParser.json());
 
// Simple auth middleware (JWT or API key stub)
app.use((req, res, next) => {
  const key = req.get('x-api-key');
  if (!key) return res.status(401).json({ error: 'missing api key' });
  // validate key & attach tenant info
  req.tenant = { id: 'team-123', plan: 'pro' };
  next();
});
 
app.post('/v1/generate', async (req, res) => {
  const { provider, model, input } = req.body;
  try {
    const output = await ProviderRouter.dispatch(req.tenant, { provider, model, input });
    res.json({ ok: true, output });
  } catch (err) {
    console.error('gateway error', err);
    res.status(500).json({ ok: false, error: err.message });
  }
});
 
app.listen(8080, () => console.log('LLM gateway listening on 8080'));

Adapter pattern

Each provider implements a small interface: generate(tenant, model, prompt, opts). The gateway only depends on that interface, making it easy to add new providers or swap implementations.

// provider/openaiAdapter.js
const fetch = require('node-fetch');
 
class OpenAIAdapter {
  constructor(apiKey) { this.apiKey = apiKey; }
 
  async generate(tenant, model, prompt, opts = {}) {
    const url = `https://api.openai.com/v1/models/${model}/completions`;
    const res = await fetch(url, {
      method: 'POST',
      headers: { 'Authorization': `Bearer ${this.apiKey}`, 'Content-Type': 'application/json' },
      body: JSON.stringify({ prompt, max_tokens: opts.maxTokens || 256 })
    });
    const data = await res.json();
    if (!res.ok) throw new Error(data.error?.message || 'openai error');
    return data.choices?.[0]?.text || '';
  }
}
 
module.exports = OpenAIAdapter;
// provider/localLlamaAdapter.js (example: spawn a local process like llama.cpp)
const { spawn } = require('child_process');
 
class LocalLlamaAdapter {
  constructor(binaryPath) { this.binaryPath = binaryPath; }
 
  generate(tenant, model, prompt, opts = {}) {
    return new Promise((resolve, reject) => {
      const p = spawn(this.binaryPath, ['--model', model, '--prompt', prompt]);
      let out = '';
      p.stdout.on('data', d => out += d.toString());
      p.stderr.on('data', d => console.error('llama stderr', d.toString()));
      p.on('close', code => code === 0 ? resolve(out) : reject(new Error('local model failed')));
    });
  }
}
 
module.exports = LocalLlamaAdapter;

Provider router and policy

Router selects an adapter based on tenant policy, model capability, and cost. Keep routing logic small and declarative.

// providerRouter.js (simplified)
const OpenAIAdapter = require('./provider/openaiAdapter');
const LocalLlamaAdapter = require('./provider/localLlamaAdapter');
 
const adapters = {
  'openai': new OpenAIAdapter(process.env.OPENAI_KEY),
  'local': new LocalLlamaAdapter('/usr/local/bin/llama')
};
 
module.exports = {
  async dispatch(tenant, { provider, model, input }) {
    // simple policy: prefer local for private teams
    const chosen = tenant.id.startsWith('private') ? 'local' : (provider || 'openai');
    const adapter = adapters[chosen];
    if (!adapter) throw new Error('no adapter');
 
    // prompt shaping example
    const prompt = `[tenant:${tenant.id}] ${input}`;
    return adapter.generate(tenant, model, prompt);
  }
};

Caching and idempotency

Cache responses for deterministic prompts to reduce cost and latency. Only cache non-sensitive outputs and honor tenant policies.

// caching pseudo
const redis = require('redis');
const client = redis.createClient();
 
async function cachedGenerate(key, fn, ttl = 60) {
  const cached = await client.get(key);
  if (cached) return JSON.parse(cached);
  const result = await fn();
  await client.setEx(key, ttl, JSON.stringify(result));
  return result;
}
 
// usage inside router: const key = `resp:${hash(prompt)}:`+model; return cachedGenerate(key, () => adapter.generate(...));

Rate limiting and quotas

Implement per-tenant and per-model limits to avoid runaway costs. Libraries like express-rate-limit or token-bucket algorithms are fine for simple setups. For robust enforcement use a centralized quota store (Redis) and a leaky-bucket algorithm.

// rate limit sketch (token bucket using Redis)
// On each request: LUA script to atomically check & decrement tokens, return allowed flag.
// If not allowed, return 429 with retry-after.

Security and privacy

  • Use mutual TLS or VPC-only endpoints for internal traffic.
  • Rotate provider secrets and never embed them in client code.
  • Sanitize prompt inputs if you forward user-supplied content to external providers.
  • Mask or redact sensitive fields before caching or logging.

Observability and cost control

Emit structured telemetry: tenant, model, tokens, latency, provider. Use those metrics for automated routing (e.g., failover to cheaper models during spikes) and billing reconciliation.

Operational tips

  • Start small: implement a single provider adapter and add features iteratively.
  • Fail open vs fail closed: choose policy per tenant — fail closed for private data, fail open for low-criticality chatbots.
  • Testing: unit-test adapters with recorded fixtures and integration-test with a staging model endpoint.
  • Local development: mock adapters so devs don't consume paid quota during testing.

Example: prompt templating and safety wrapper

Templating helps standardize system messages and safety checks. Keep templates versioned so you can A/B test system prompts without code changes.

// simple template function
function renderTemplate(template, vars) {
  return template.replace(/\{\{(\w+)\}\}/g, (_, k) => vars[k] || '');
}
 
const sys = 'You are an assistant for tenant {{tenantId}}. Keep answers short.';
const prompt = renderTemplate(sys, { tenantId: 'team-123' }) + '\nUser: ' + 'Explain event sourcing.';

Concise checklist before production

  1. Secrets stored in a secret manager, not in code.
  2. Per-tenant authentication and authorization enforced.
  3. Rate limits, quotas, and cost alerts configured.
  4. Cache policy and data retention rules defined.
  5. Monitoring, logging, and tracing in place for root cause analysis.

Conclusion

A local LLM gateway centralizes control over model selection, security, and cost while enabling consistent prompt shaping and telemetry. Start with a small adapter interface, add caching, rate limiting, and observability, and iterate. The adapter pattern keeps the gateway flexible: add new providers or local runtimes without rewriting routing logic.

Further reading: the original project that motivated this pattern explores a local gateway for multiple models and messaging integrations.

I Wanted One Local Gateway for Claude Code, Codex, Gemini, Telegram, Feishu, and DingTalk — CliGate

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment