Why orchestrate LLMs in the browser?
Developers are increasingly running parts of LLM logic in the browser: to reduce latency, improve privacy for client data, or provide offline fallbacks with compact WASM models. But a single model rarely covers every requirement. Orchestration — routing requests between local and remote models, combining outputs, and applying caching, retry and batching strategies — unlocks practical, cost-efficient, and resilient UX.
Core architecture and patterns
A simple browser LLM orchestrator usually includes these components:
- Provider adapters — thin wrappers around each model backend (local WASM, hosted API, multi-model cloud).
- Routing policy — decides which provider to use (cost, freshness, latency, capability).
- Batching / queue — groups small requests to reduce per-call overhead.
- Cache — memoize deterministic prompts to avoid repeated cost.
- Safety and secret handling — never embed long-lived API keys in client code; use short-lived tokens or a minimal proxy.
Tradeoffs to consider
- Local models = low latency and privacy, but limited capability and larger client footprint.
- Remote APIs scale in capability but add cost, latency, and require secure token exchange.
- Complex orchestration improves UX but increases maintenance and testing surface.
Example: Minimal orchestrator (routing + fallback + caching)
The following demonstrates a pragmatic orchestrator in plain JavaScript. It shows a provider interface, a routing decision that tries a local provider first and falls back to a remote API, and a simple TTL cache in localStorage. The code examples are intentionally small so you can adapt them.
<!-- Provider interface (pseudo-JS) -->
class LLMProvider {
async request(prompt, opts = {}) {
throw new Error('Not implemented');
}
}
// Local provider (e.g., WASM-backed) - stubbed
class LocalWasmProvider extends LLMProvider {
async request(prompt, opts = {}) {
// Replace with an actual WASM model call running in a Worker
return { text: 'local: ' + prompt.slice(0, 100) };
}
}
// Remote provider wrapper (calls a server-side token exchange / API)
class RemoteApiProvider extends LLMProvider {
constructor(apiUrl, getShortLivedToken) {
super();
this.apiUrl = apiUrl;
this.getShortLivedToken = getShortLivedToken; // async fn
}
async request(prompt, opts = {}) {
const token = await this.getShortLivedToken();
const res = await fetch(this.apiUrl, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
Authorization: 'Bearer ' + token
},
body: JSON.stringify({ prompt, opts })
});
if (!res.ok) throw new Error('Remote provider error');
return res.json();
}
}
// Simple TTL cache using localStorage
class SimpleCache {
constructor(prefix = 'llm:', ttlMs = 1000 * 60 * 5) {
this.prefix = prefix;
this.ttlMs = ttlMs;
}
_key(key) { return this.prefix + key; }
get(key) {
try {
const raw = localStorage.getItem(this._key(key));
if (!raw) return null;
const obj = JSON.parse(raw);
if (Date.now() - obj.t > this.ttlMs) {
localStorage.removeItem(this._key(key));
return null;
}
return obj.v;
} catch (e) {
return null;
}
}
set(key, value) {
try {
localStorage.setItem(this._key(key), JSON.stringify({ t: Date.now(), v: value }));
} catch (e) {
// ignore storage errors (quota)
}
}
}
// Orchestrator
class Orchestrator {
constructor({ localProvider, remoteProvider, cache }) {
this.local = localProvider;
this.remote = remoteProvider;
this.cache = cache || new SimpleCache();
}
// Deterministic key for caching simple prompt responses
_cacheKey(prompt, opts) {
return btoa(unescape(encodeURIComponent(prompt + JSON.stringify(opts || {}))));
}
async generate(prompt, opts = {}) {
const key = this._cacheKey(prompt, opts);
const cached = this.cache.get(key);
if (cached) return cached;
// Routing policy: try local first for privacy/latency; fallback to remote
try {
const localRes = await this.local.request(prompt, opts);
this.cache.set(key, localRes);
return localRes;
} catch (err) {
console.warn('Local provider failed, falling back to remote:', err);
}
const remoteRes = await this.remote.request(prompt, opts);
this.cache.set(key, remoteRes);
return remoteRes;
}
}
// Usage example (replace getShortLivedToken with your backend handshake)
const orchestrator = new Orchestrator({
localProvider: new LocalWasmProvider(),
remoteProvider: new RemoteApiProvider('/api/infer', async () => {
// Exchange a session cookie or client id for a short-lived token
const r = await fetch('/api/get-token', { method: 'POST', credentials: 'include' });
if (!r.ok) throw new Error('token exchange failed');
const j = await r.json();
return j.token;
})
});
// Call
(async () => {
const out = await orchestrator.generate('Explain memoization in one paragraph.');
console.log(out);
})();Example: Batch queue for small requests
Batching is useful when many small UI-driven prompts can be aggregated to one request to a remote API (reducing per-call overhead and cost). Below is a minimal grouping queue with a fixed micro-batching interval.
class BatchQueue {
constructor(sendBatchFn, { maxBatch = 8, waitMs = 50 } = {}) {
this.sendBatchFn = sendBatchFn;
this.maxBatch = maxBatch;
this.waitMs = waitMs;
this.queue = [];
this.timer = null;
}
enqueue(prompt, opts = {}) {
return new Promise((resolve, reject) => {
this.queue.push({ prompt, opts, resolve, reject });
if (this.queue.length >= this.maxBatch) this._flush();
if (!this.timer) this.timer = setTimeout(() => this._flush(), this.waitMs);
});
}
async _flush() {
if (this.timer) { clearTimeout(this.timer); this.timer = null; }
if (this.queue.length === 0) return;
const batch = this.queue.splice(0, this.maxBatch);
try {
const inputs = batch.map(b => ({ prompt: b.prompt, opts: b.opts }));
const results = await this.sendBatchFn(inputs);
results.forEach((r, i) => batch[i].resolve(r));
} catch (err) {
batch.forEach(b => b.reject(err));
}
}
}
// Example sendBatchFn that calls remote API
async function sendBatchToRemote(inputs) {
const res = await fetch('/api/batch-infer', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ inputs })
});
if (!res.ok) throw new Error('batch request failed');
return res.json();
}
// Usage
const q = new BatchQueue(sendBatchToRemote, { maxBatch: 4, waitMs: 30 });
q.enqueue('Summarize A').then(console.log);
q.enqueue('Summarize B').then(console.log);Security and secret handling
Never bake long-lived API keys into client-side code. Practical approaches:
- Token exchange: issue short-lived tokens from a backend endpoint that authenticates the user (cookie or OAuth) and returns a token with limited TTL and scope.
- Minimal proxy: a server endpoint that accepts the client prompt and forwards it to the real API, enforcing rate limits and content policy. Keep the surface minimal to reduce server costs.
- Model sandboxing: run local models inside Web Workers and restrict data sent to remote providers client-side.
// Example: very small token-exchange endpoint (server-side pseudocode)
// POST /api/get-token
// Authenticate user session (cookie/JWT), then call provider to mint a short token.
// Response: { token: 'eyJ...' }
// Client uses that token for a short time and then requests a fresh one.
// This prevents embedding permanent keys in shipped JavaScript.Testing and observability
Key practices:
- Mock providers in unit tests — your orchestrator should be tested with local and remote stubs.
- Record telemetry for routing decisions (local vs remote), request latency, and fallback counts.
- Sanitize or avoid logging user prompt contents in production traces.
When not to orchestrate in the browser
If you require high-accuracy models only available in private cloud, strict auditing of prompts, or complex multi-step chained reasoning with hidden data dependencies, keep orchestration server-side where you can fully control execution, credentials, and logging.
Conclusion
Browser LLM orchestration offers practical benefits: lower latency, better privacy options, and cost control — but it also adds complexity. Start small: implement provider adapters, use a short-lived token pattern, add a simple cache, and optionally a micro-batching queue. Iterate the routing policy based on metrics. The patterns above give a compact, testable foundation you can expand as your product needs grow.
Further steps
- Replace the LocalWasmProvider stub with a real WASM model running in a Worker and stream responses to the UI.
- Use IndexedDB for larger or structured caches.
- Expose observability hooks for routing decisions and cost tracking.
Tip: Keep the orchestration layer small and focused — it's easier to reason about and audit than a large client-side SDK that mixes UI concerns with heavy model logic.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment