Why model routing matters for agentic AI
Agentic systems—agents that decompose tasks, call tools, and chain LLM calls—benefit from routing decisions that match a subtask to the right model. Model routing reduces cost, improves latency, boosts capability fit, and limits risk by isolating high-capacity LLMs to tasks that truly need them.
Practical routing patterns
- Static routing: Map request types to models (e.g., code → code-specialized LLM).
- Classifier-based routing: Run a lightweight classifier to tag incoming tasks and pick a model.
- Cost/latency-aware routing: Choose models based on budget or SLA constraints.
- Cascading / multi-stage: Try a cheap model first; escalate to larger models only on failure or low confidence.
- Ensemble or reviewer routing: Use a fast model for generation and a stronger model for verification or safety checks.
Design primitives to implement
- Model registry: metadata for cost, latency, capabilities, and endpoints.
- Selector engine: rules or ML classifier that maps tasks & context to model candidates.
- Executor with timeouts, retries, and concurrency limits for each model endpoint.
- Fallbacks: defined fallthrough when a model times out, errors, or returns low-confidence output.
- Observability: log input/features, selected model, latency, cost estimate, and confidence.
Example: async Python router for agentic subtasks
The example shows a minimal model registry and a selector that balances capability tags and a latency budget. This is an integration pattern you can adapt to real LLM clients.
import asyncio
from dataclasses import dataclass
from typing import Dict, List
@dataclass
class ModelProfile:
name: str
cost_per_token: float
est_latency_ms: int
capabilities: List[str]
# Simple in-memory registry
MODEL_REGISTRY: Dict[str, ModelProfile] = {
"fast-nl": ModelProfile("fast-nl", 0.00001, 100, ["chat", "summarize"]),
"code-pro": ModelProfile("code-pro", 0.0003, 400, ["code", "explain", "static-analysis"]),
"large-omn": ModelProfile("large-omn", 0.001, 1200, ["reasoning", "planning", "long-context"]),
}
def select_model_for_task(task_tags: List[str], latency_budget_ms: int):
# Prefer models that advertise all capabilities required and fit latency
candidates = []
for m in MODEL_REGISTRY.values():
if all(tag in m.capabilities for tag in task_tags) and m.est_latency_ms <= latency_budget_ms:
candidates.append(m)
if candidates:
# simple cost heuristic: pick cheapest among candidates
return min(candidates, key=lambda p: p.cost_per_token)
# Fallback: pick the most capable model (best coverage) and log escalation
return max(MODEL_REGISTRY.values(), key=lambda p: len(p.capabilities))
async def call_model(model_name: str, prompt: str, timeout_s: float = 5.0):
# Replace this stub with your LLM client call (e.g., OpenAI, local API)
# Simulate network call and response
await asyncio.sleep(MODEL_REGISTRY[model_name].est_latency_ms / 1000)
return {"model": model_name, "text": f"response from {model_name} to: {prompt}"}
async def route_and_execute(task):
model = select_model_for_task(task["tags"], task.get("latency_budget_ms", 1000))
try:
result = await asyncio.wait_for(call_model(model.name, task["prompt"]), timeout=3.0)
return {"selected": model.name, "result": result}
except asyncio.TimeoutError:
# Timeout: escalate to a more capable model with higher timeout
fallback = MODEL_REGISTRY["large-omn"]
result = await call_model(fallback.name, task["prompt"], timeout_s=10.0)
return {"selected": fallback.name, "result": result, "escalated": True}
# Usage example
async def main():
task = {"tags": ["code", "explain"], "prompt": "Explain this function...", "latency_budget_ms": 500}
out = await route_and_execute(task)
print(out)
if __name__ == "__main__":
asyncio.run(main())Example: Node.js/Express endpoint that routes by intent and budget
A thin HTTP layer that extracts client budget and intent, then maps to an internal model id. This is suitable for multi-tenant services where clients pass SLAs.
const express = require('express')
const bodyParser = require('body-parser')
const app = express()
app.use(bodyParser.json())
// Simple mapping; replace with classifier if needed
function pickModel(intent, budgetMs) {
if (intent === 'code') return (budgetMs <= 300) ? 'fast-nl' : 'code-pro'
if (intent === 'summarize') return 'fast-nl'
return (budgetMs <= 500) ? 'fast-nl' : 'large-omn'
}
app.post('/api/generate', async (req, res) => {
const { intent, budgetMs = 1000, prompt } = req.body
const model = pickModel(intent, budgetMs)
try {
// Call your LLM service here. Stubbed as an immediate response.
const response = { model, text: `generated by ${model}` }
res.json({ model, response })
} catch (err) {
// On error, log and send safe fallback
console.error('model error', err)
res.status(502).json({ error: 'model unavailable' })
}
})
app.listen(3000, () => console.log('router running on :3000'))Implementation advice and observability
- Metrics: track selected model, latency, error rate, token usage, and cost estimate per request.
- Canaries: route a small % of traffic to a new model to measure behavior before full rollout.
- Confidence signals: use model-provided scores, auxiliary verifier models, or heuristics to decide escalation.
- Caching: cache deterministic responses and verification results to avoid repeated calls to large models.
- Rate/Concurrency limits: enforce per-model concurrent call limits to avoid overloading endpoints or local GPU nodes.
- Security: treat model outputs as untrusted—validate tool invocations, sanitize inputs, and avoid unsafe deserialization patterns.
Tradeoffs
- Simplicity vs. accuracy: Static rules are simple but less adaptive; classifiers improve fit at the cost of model complexity and maintenance.
- Cost vs. quality: Cascading saves cost but introduces extra latency when escalation occurs.
- Latency budgets: Tight budgets can reduce model choices and force lower-quality outputs; track customer SLAs carefully.
- Operational overhead: More models and routing logic require stronger observability and automation for deploys, canaries, and safety checks.
Conclusion
Model routing is a high-leverage pattern for agentic AI systems. Start with a simple registry and deterministic rules, add lightweight classifiers or cost-aware heuristics, and instrument everything for measurement. Use cascading and reviewer patterns to keep costs down without sacrificing safety. Over time, mature routing into an automated policy engine with canaries and automated rollbacks.
Further reading
- Agentic AI Architecture Needs Model Routing (DEV) - discussion that motivated this practical guide.
- Production-Grade Engineering Skills for AI Coding Agents (DEV) - tips on production concerns for agentic systems.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment