Sechno
Ai

Practical Patterns for Integrating LLMs into Production Backends

Concrete architecture, code patterns, and tradeoffs for safely, cost-effectively, and observably integrating large language models (LLMs) into backend services.

SSechno Team 5 min read 162 views
Practical Patterns for Integrating LLMs into Production Backends

Why this matters

LLMs are powerful but bring operational concerns: cost, latency, reliability, observability, and safety. This guide gives practical patterns you can apply now: caching, retries and backoff, rate limiting, tracing and correlation, cost controls, and strategies for testing unknown model outputs.

High-level architecture

Keep an LLM integration behind a small, opinionated service boundary:

  • API adapter: normalizes prompts and responses to your domain schema.
  • Policy layer: enforces rate limits, cost budgets, and input/output validation.
  • Cache layer: deduplicates repeated prompts and stores short-lived outputs.
  • Observability: request tracing, latency metrics, token usage tracking.
  • Safety sandbox: sanitization, hallucination detection heuristics, and human review paths.

Tradeoffs

  • Latency vs. freshness: caching reduces cost and latency but can return stale completions.
  • Complexity vs. control: adding retries, tracing, and cost accounting increases code but prevents outages and surprises.
  • Accuracy vs. guardrails: strong filters reduce risk but may reject useful outputs.

Example 1 — Python (FastAPI): cache + request correlation

This example shows a minimal endpoint that deduplicates identical prompts via Redis, attaches a correlation id, and forwards the request to an LLM API. Use it as a starting point, not a drop-in production piece.

from fastapi import FastAPI, Request, Header, HTTPException
import httpx
import aioredis
import hashlib
import json
import uuid
 
app = FastAPI()
redis = None
 
@app.on_event("startup")
async def startup():
    global redis
    redis = await aioredis.from_url("redis://localhost:6379/0")
 
def cache_key(prompt: str) -> str:
    h = hashlib.sha256(prompt.encode()).hexdigest()
    return f"llm:prompt:{h}"
 
@app.post("/generate")
async def generate(request: Request, x_request_id: str | None = Header(None)):
    payload = await request.json()
    prompt = payload.get("prompt")
    if not prompt:
        raise HTTPException(status_code=400, detail="missing prompt")
 
    req_id = x_request_id or str(uuid.uuid4())
    key = cache_key(prompt)
 
    cached = await redis.get(key)
    if cached:
        return json.loads(cached)
 
    headers = {"Authorization": "Bearer YOUR_API_KEY", "X-Request-Id": req_id}
    async with httpx.AsyncClient(timeout=15.0) as client:
        resp = await client.post(
            "https://api.example.com/v1/generate",
            json={"prompt": prompt, "max_tokens": 512},
            headers=headers,
        )
        resp.raise_for_status()
        result = resp.json()
 
    # short TTL to balance freshness vs cost
    await redis.set(key, json.dumps(result), ex=300)
    return result

Example 2 — JavaScript: retry with exponential backoff and jitter

Transient network errors are common. Retry smartly to avoid thundering herds and to respect provider rate limits.

async function sleep(ms) {
  return new Promise(resolve => setTimeout(resolve, ms));
}
 
async function fetchWithRetry(url, options = {}, retries = 3, base = 200) {
  for (let attempt = 0; attempt <= retries; attempt++) {
    try {
      const res = await fetch(url, options);
      if (!res.ok) throw new Error(`HTTP ${res.status}`);
      return await res.json();
    } catch (err) {
      if (attempt === retries) throw err;
      // exponential backoff with jitter
      const jitter = Math.random() * base;
      await sleep(base * Math.pow(2, attempt) + jitter);
    }
  }
}
 
// usage
// const result = await fetchWithRetry('https://api.example.com/v1/generate', {method: 'POST', body: JSON.stringify({prompt})}, 4);

Observability: tracing and token accounting

Tracing lets you see where latency and errors happen. At minimum:

  • Propagate a request-id header from your frontend through your LLM adapter.
  • Log request/response size, token counts (if the provider returns them), and model name.
  • Emit metrics: requests/sec, errors/sec, avg latency, cost per call.

If possible, integrate OpenTelemetry to capture spans for network calls and downstream DB/cache interactions.

Cost control patterns

  • Cache repeated prompts (100–300s TTL for conversational turn reuse).
  • Model selection: route cost-sensitive calls to smaller models, heavy reasoning to larger ones.
  • Input truncation and token budgeting: pre-validate prompts and truncate or summarize long inputs client-side.
  • Per-user or per-tenant quotas with soft/hard limits and quota dashboards.

Testing and validation when outputs are unknown

Unlike deterministic services, LLM outputs vary. Use these practices:

  • Schema validation: define the expected output shape and reject or transform invalid responses.
  • Golden-path fuzz tests: send many prompts and assert invariants (e.g., no PII, required fields present).
  • Human-in-the-loop sampling: route a sample of outputs to reviewers to catch degradation.
  • Record inputs + outputs for reproducibility and rollback; redact sensitive data.

Safety and moderation

Always run outputs through safety checks before returning them to end users. Simple heuristics include blocklists, regex checks, and classifier-based moderation. When in doubt, fail-safe to a human review or a generic fallback response.

Putting it together: a small checklist before production rollout

  1. Implement request correlation and basic tracing.
  2. Add caching for repeatable prompts.
  3. Implement retries with exponential backoff and jitter.
  4. Track token usage and build a cost dashboard or alerts.
  5. Validate outputs against a schema and provide fallbacks.
  6. Enforce quota/rate limits per actor (user, tenant, key).
  7. Set up sampling-based human review and automated moderation.

Further reading

Conclusion

Integrating LLMs is not just about API calls. Treat the model as an expensive, non-deterministic downstream service: add caching, tracing, retries, cost controls, and safety layers. Start small with an adapter service, measure token usage, and iterate on model routing and validation. These patterns reduce surprises and let you deliver value safely.

Note: code examples are illustrative. Replace endpoints, keys, and configuration with your production settings and secure secrets in environment variables or a secrets manager.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment