Sechno
Web Development

Practical Guide to Browser-Based LLM Orchestration: Routing, Batching, and Safety

How to design a client-side LLM orchestration layer that routes requests between local and remote providers, batches and caches queries, and keeps secrets and costs under control — with actionable code patterns you can reuse.

SSechno Team 7 min read 76 views
Practical Guide to Browser-Based LLM Orchestration: Routing, Batching, and Safety

Why orchestrate LLMs in the browser?

Developers are increasingly running parts of LLM logic in the browser: to reduce latency, improve privacy for client data, or provide offline fallbacks with compact WASM models. But a single model rarely covers every requirement. Orchestration — routing requests between local and remote models, combining outputs, and applying caching, retry and batching strategies — unlocks practical, cost-efficient, and resilient UX.

Core architecture and patterns

A simple browser LLM orchestrator usually includes these components:

  • Provider adapters — thin wrappers around each model backend (local WASM, hosted API, multi-model cloud).
  • Routing policy — decides which provider to use (cost, freshness, latency, capability).
  • Batching / queue — groups small requests to reduce per-call overhead.
  • Cache — memoize deterministic prompts to avoid repeated cost.
  • Safety and secret handling — never embed long-lived API keys in client code; use short-lived tokens or a minimal proxy.

Tradeoffs to consider

  • Local models = low latency and privacy, but limited capability and larger client footprint.
  • Remote APIs scale in capability but add cost, latency, and require secure token exchange.
  • Complex orchestration improves UX but increases maintenance and testing surface.

Example: Minimal orchestrator (routing + fallback + caching)

The following demonstrates a pragmatic orchestrator in plain JavaScript. It shows a provider interface, a routing decision that tries a local provider first and falls back to a remote API, and a simple TTL cache in localStorage. The code examples are intentionally small so you can adapt them.

<!-- Provider interface (pseudo-JS) -->
class LLMProvider {
  async request(prompt, opts = {}) {
    throw new Error('Not implemented');
  }
}
 
// Local provider (e.g., WASM-backed) - stubbed
class LocalWasmProvider extends LLMProvider {
  async request(prompt, opts = {}) {
    // Replace with an actual WASM model call running in a Worker
    return { text: 'local: ' + prompt.slice(0, 100) };
  }
}
 
// Remote provider wrapper (calls a server-side token exchange / API)
class RemoteApiProvider extends LLMProvider {
  constructor(apiUrl, getShortLivedToken) {
    super();
    this.apiUrl = apiUrl;
    this.getShortLivedToken = getShortLivedToken; // async fn
  }
 
  async request(prompt, opts = {}) {
    const token = await this.getShortLivedToken();
    const res = await fetch(this.apiUrl, {
      method: 'POST',
      headers: {
        'Content-Type': 'application/json',
        Authorization: 'Bearer ' + token
      },
      body: JSON.stringify({ prompt, opts })
    });
    if (!res.ok) throw new Error('Remote provider error');
    return res.json();
  }
}
 
// Simple TTL cache using localStorage
class SimpleCache {
  constructor(prefix = 'llm:', ttlMs = 1000 * 60 * 5) {
    this.prefix = prefix;
    this.ttlMs = ttlMs;
  }
 
  _key(key) { return this.prefix + key; }
 
  get(key) {
    try {
      const raw = localStorage.getItem(this._key(key));
      if (!raw) return null;
      const obj = JSON.parse(raw);
      if (Date.now() - obj.t > this.ttlMs) {
        localStorage.removeItem(this._key(key));
        return null;
      }
      return obj.v;
    } catch (e) {
      return null;
    }
  }
 
  set(key, value) {
    try {
      localStorage.setItem(this._key(key), JSON.stringify({ t: Date.now(), v: value }));
    } catch (e) {
      // ignore storage errors (quota)
    }
  }
}
 
// Orchestrator
class Orchestrator {
  constructor({ localProvider, remoteProvider, cache }) {
    this.local = localProvider;
    this.remote = remoteProvider;
    this.cache = cache || new SimpleCache();
  }
 
  // Deterministic key for caching simple prompt responses
  _cacheKey(prompt, opts) {
    return btoa(unescape(encodeURIComponent(prompt + JSON.stringify(opts || {}))));
  }
 
  async generate(prompt, opts = {}) {
    const key = this._cacheKey(prompt, opts);
    const cached = this.cache.get(key);
    if (cached) return cached;
 
    // Routing policy: try local first for privacy/latency; fallback to remote
    try {
      const localRes = await this.local.request(prompt, opts);
      this.cache.set(key, localRes);
      return localRes;
    } catch (err) {
      console.warn('Local provider failed, falling back to remote:', err);
    }
 
    const remoteRes = await this.remote.request(prompt, opts);
    this.cache.set(key, remoteRes);
    return remoteRes;
  }
}
 
// Usage example (replace getShortLivedToken with your backend handshake)
const orchestrator = new Orchestrator({
  localProvider: new LocalWasmProvider(),
  remoteProvider: new RemoteApiProvider('/api/infer', async () => {
    // Exchange a session cookie or client id for a short-lived token
    const r = await fetch('/api/get-token', { method: 'POST', credentials: 'include' });
    if (!r.ok) throw new Error('token exchange failed');
    const j = await r.json();
    return j.token;
  })
});
 
// Call
(async () => {
  const out = await orchestrator.generate('Explain memoization in one paragraph.');
  console.log(out);
})();

Example: Batch queue for small requests

Batching is useful when many small UI-driven prompts can be aggregated to one request to a remote API (reducing per-call overhead and cost). Below is a minimal grouping queue with a fixed micro-batching interval.

class BatchQueue {
  constructor(sendBatchFn, { maxBatch = 8, waitMs = 50 } = {}) {
    this.sendBatchFn = sendBatchFn;
    this.maxBatch = maxBatch;
    this.waitMs = waitMs;
    this.queue = [];
    this.timer = null;
  }
 
  enqueue(prompt, opts = {}) {
    return new Promise((resolve, reject) => {
      this.queue.push({ prompt, opts, resolve, reject });
      if (this.queue.length >= this.maxBatch) this._flush();
      if (!this.timer) this.timer = setTimeout(() => this._flush(), this.waitMs);
    });
  }
 
  async _flush() {
    if (this.timer) { clearTimeout(this.timer); this.timer = null; }
    if (this.queue.length === 0) return;
    const batch = this.queue.splice(0, this.maxBatch);
    try {
      const inputs = batch.map(b => ({ prompt: b.prompt, opts: b.opts }));
      const results = await this.sendBatchFn(inputs);
      results.forEach((r, i) => batch[i].resolve(r));
    } catch (err) {
      batch.forEach(b => b.reject(err));
    }
  }
}
 
// Example sendBatchFn that calls remote API
async function sendBatchToRemote(inputs) {
  const res = await fetch('/api/batch-infer', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ inputs })
  });
  if (!res.ok) throw new Error('batch request failed');
  return res.json();
}
 
// Usage
const q = new BatchQueue(sendBatchToRemote, { maxBatch: 4, waitMs: 30 });
q.enqueue('Summarize A').then(console.log);
q.enqueue('Summarize B').then(console.log);

Security and secret handling

Never bake long-lived API keys into client-side code. Practical approaches:

  1. Token exchange: issue short-lived tokens from a backend endpoint that authenticates the user (cookie or OAuth) and returns a token with limited TTL and scope.
  2. Minimal proxy: a server endpoint that accepts the client prompt and forwards it to the real API, enforcing rate limits and content policy. Keep the surface minimal to reduce server costs.
  3. Model sandboxing: run local models inside Web Workers and restrict data sent to remote providers client-side.
// Example: very small token-exchange endpoint (server-side pseudocode)
// POST /api/get-token
// Authenticate user session (cookie/JWT), then call provider to mint a short token.
 
// Response: { token: 'eyJ...' }
 
// Client uses that token for a short time and then requests a fresh one.
 
// This prevents embedding permanent keys in shipped JavaScript.

Testing and observability

Key practices:

  • Mock providers in unit tests — your orchestrator should be tested with local and remote stubs.
  • Record telemetry for routing decisions (local vs remote), request latency, and fallback counts.
  • Sanitize or avoid logging user prompt contents in production traces.

When not to orchestrate in the browser

If you require high-accuracy models only available in private cloud, strict auditing of prompts, or complex multi-step chained reasoning with hidden data dependencies, keep orchestration server-side where you can fully control execution, credentials, and logging.

Conclusion

Browser LLM orchestration offers practical benefits: lower latency, better privacy options, and cost control — but it also adds complexity. Start small: implement provider adapters, use a short-lived token pattern, add a simple cache, and optionally a micro-batching queue. Iterate the routing policy based on metrics. The patterns above give a compact, testable foundation you can expand as your product needs grow.

Further steps

  • Replace the LocalWasmProvider stub with a real WASM model running in a Worker and stream responses to the UI.
  • Use IndexedDB for larger or structured caches.
  • Expose observability hooks for routing decisions and cost tracking.

Tip: Keep the orchestration layer small and focused — it's easier to reason about and audit than a large client-side SDK that mixes UI concerns with heavy model logic.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment