Sechno
Ai

Designing Reliable Agentic Workflows for LLMs: Patterns, Code, and Production Tradeoffs

Practical patterns and code for building robust, observable, and secure agent-style LLM workflows that call external tools in production — with caching, timeouts, idempotency, and testing advice.

SSechno Team 7 min read 161 views
Designing Reliable Agentic Workflows for LLMs: Patterns, Code, and Production Tradeoffs

Overview

As LLMs are used as orchestrators that call external tools (search, databases, calculators, APIs), production systems need explicit patterns to remain reliable, secure, and observable. This article collects practical patterns and ready-to-adapt code for building "agentic" workflows: tool registries, validation layers, explicit caching with TTL, retry and backoff strategies, idempotency, and testing approaches.

Core patterns

1) Strong tool interface and registry

Define a small, explicit set of callable tools with typed inputs and outputs. The agent should never call arbitrary code paths — use a registry that maps tool names to vetted functions and enforces input validation and output schemas.

2) Action validation and sandboxing

Require the LLM to return a restricted, machine-readable action (JSON or function-call format). Validate the action against a schema before executing the tool. Reject or ask for clarification instead of executing uncertain payloads.

3) Timeouts, cancellation, and resource limits

Always run tools with timeouts and resource limits. For network calls, set socket/connect/read timeouts. For CPU-bound operations, run in a worker with an enforced execution budget.

4) Explicit caching and prompt-result caching

Do not rely solely on provider-side prompt caches (these can change). Implement a local cache layer you control with configurable TTLs, invalidation rules, and cache keys derived from the prompt + tool inputs + agent policy version.

5) Idempotency and retry semantics

Design tools to be idempotent where possible. For non-idempotent side effects, require explicit confirmation steps. When retrying transient errors, use exponential backoff and detect duplicate intent using idempotency tokens.

6) Observability and tracing

Emit structured logs and traces for: prompts sent, model responses, parsed actions, tool calls, tool responses, errors, and final outputs. Include correlation ids for each request so you can reconstruct the full interaction.

Production-ready Python agent skeleton

The example below is intentionally minimal but demonstrates: parsing an LLM response into an action, validating the action, mapping to a registered tool, enforcing a timeout, applying caching, and returning a combined result.

import json
import time
import threading
from concurrent.futures import ThreadPoolExecutor, TimeoutError
from typing import Any, Callable, Dict
 
# Simple in-memory cache with TTL
class TTLCache:
    def __init__(self):
        self.store: Dict[str, Dict[str, Any]] = {}
 
    def get(self, key: str):
        v = self.store.get(key)
        if not v:
            return None
        if v['expires_at'] < time.time():
            del self.store[key]
            return None
        return v['value']
 
    def set(self, key: str, value: Any, ttl: int):
        self.store[key] = {'value': value, 'expires_at': time.time() + ttl}
 
cache = TTLCache()
 
# Tool registry: map name to function and a wrapper that enforces validation/timeouts
Tool = Callable[[Dict[str, Any]], Dict[str, Any]]
 
def search_tool(params: Dict[str, Any]) -> Dict[str, Any]:
    # placeholder for a real search call
    query = params.get('query', '')
    return {'results': [f'Result for: {query}']}
 
def calculator_tool(params: Dict[str, Any]) -> Dict[str, Any]:
    expr = params.get('expression', '0')
    # VERY simple evaluator - in prod use a safe math evaluator
    try:
        value = eval(expr, {"__builtins__": {}})
        return {'value': value}
    except Exception as e:
        return {'error': str(e)}
 
TOOL_REGISTRY: Dict[str, Tool] = {
    'search': search_tool,
    'calc': calculator_tool,
}
 
# Mock LLM call - replace with your provider SDK
def call_llm(prompt: str) -> str:
    # The LLM should respond with a JSON action like: {"tool":"search","input":{"query":"x"}}
    # Here we return a fixed action for demonstration.
    return json.dumps({"tool": "search", "input": {"query": "best practices for caching"}})
 
# Validate action shape
def validate_action(action: Dict[str, Any]) -> bool:
    return isinstance(action, dict) and 'tool' in action and 'input' in action and action['tool'] in TOOL_REGISTRY
 
# Execute tool with timeout
def execute_tool_with_timeout(tool_fn: Tool, params: Dict[str, Any], timeout_sec: int) -> Dict[str, Any]:
    with ThreadPoolExecutor(max_workers=1) as ex:
        fut = ex.submit(tool_fn, params)
        try:
            return fut.result(timeout=timeout_sec)
        except TimeoutError:
            return {'error': 'tool_timeout'}
        except Exception as e:
            return {'error': f'tool_exception: {e}'}
 
# Agent entry point
def run_agent(prompt: str, cache_ttl: int = 300, tool_timeout: int = 5):
    cache_key = f"agent_prompt_v1:{prompt}"
    cached = cache.get(cache_key)
    if cached:
        return {'from_cache': True, 'result': cached}
 
    llm_raw = call_llm(prompt)
    try:
        action = json.loads(llm_raw)
    except Exception:
        return {'error': 'invalid_llm_response'}
 
    if not validate_action(action):
        return {'error': 'invalid_action'}
 
    tool_name = action['tool']
    params = action['input']
    tool_fn = TOOL_REGISTRY[tool_name]
 
    result = execute_tool_with_timeout(tool_fn, params, timeout_sec=tool_timeout)
 
    # Basic post-call validation
    if isinstance(result, dict) and 'error' in result:
        return {'error': 'tool_failed', 'details': result}
 
    cache.set(cache_key, result, ttl=cache_ttl)
    return {'from_cache': False, 'result': result}
 
# Example run
if __name__ == '__main__':
    out = run_agent('Find best practices for caching in LLM agents')
    print(out)

JavaScript: idempotency and server handler pattern

Node/Express handler that enforces an idempotency token to avoid double-executing non-idempotent tools (e.g., creating invoices).

import express from 'express'
import fetch from 'node-fetch'
 
const app = express()
app.use(express.json())
 
// In-memory idempotency store (replace with durable store in prod)
const idempotencyStore = new Map()
 
app.post('/agent', async (req, res) => {
  const idempotencyKey = req.header('Idempotency-Key')
  if (!idempotencyKey) return res.status(400).json({ error: 'missing idempotency key' })
 
  if (idempotencyStore.has(idempotencyKey)) {
    return res.json({ status: 'duplicate', result: idempotencyStore.get(idempotencyKey) })
  }
 
  const { prompt } = req.body
 
  // call LLM (replace with provider SDK)
  const llmResp = await fetch('https://llm.example/v1/generate', { method: 'POST', body: JSON.stringify({ prompt }) })
  const llmJson = await llmResp.json()
 
  // Parse safe action from model
  const action = llmJson.action
  if (!action || !action.tool) return res.status(400).json({ error: 'invalid action' })
 
  try {
    // If tool is non-idempotent, enforce explicit confirmation step or merchant-approved flow
    const result = await runTool(action)
    idempotencyStore.set(idempotencyKey, result)
    return res.json({ status: 'ok', result })
  } catch (err) {
    return res.status(500).json({ error: String(err) })
  }
})
 
async function runTool(action) {
  // Tool dispatching with timeouts and validation
  if (action.tool === 'payments.create') {
    // ensure required fields
    if (!action.input || !action.input.amount) throw new Error('invalid payment')
    // call payment gateway with strong idempotency key
    // ...
    return { id: 'payment_123', status: 'created' }
  }
  // other tools
  return { ok: true }
}
 
app.listen(3000)

Testing, CI, and safe rollout

  • Unit test tools: test each tool implementation in isolation with mocked dependencies.
  • Integration tests: run agent runs against a sandboxed LLM or deterministic stub to validate action parsing and error flows.
  • Chaos and fault injection: simulate tool timeouts, slow responses, and malformed outputs in CI to ensure graceful degradation.
  • Canary rollout: route a small portion of live traffic through the agent logic while monitoring error rates and business metrics.

Tradeoffs and practical considerations

  1. Complexity vs. capability: Agentic systems are powerful but increase surface area for bugs and security issues. Prefer a small set of well-tested tools rather than exposing everything.
  2. Latency vs. accuracy: Synchronous tool calls (waiting for external APIs) increase latency. Consider async patterns where the LLM triggers an async job and returns a progress handle.
  3. Caching freshness: Aggressive caching reduces cost and latency but can serve stale results. Use per-tool TTLs and cache keys including agent version and relevant context.
  4. Provider changes: Provider-side defaults (like prompt-cache TTLs) can change. Run a local cache and design for graceful change: use feature flags and runtime-configurable TTLs.

Further reading

In short: treat the LLM as a planner, not as an invoker of arbitrary side effects. Enforce constraints, validate everything, and instrument thoroughly.

Conclusion

Agentic LLM workflows unlock powerful integrations, but they require explicit engineering guardrails: typed tool interfaces, validation, local caching, timeouts, idempotency, and solid observability. Start small, instrument everything, and automate tests that emulate failure modes. When you do, agentic patterns can be deployed safely and scaled reliably.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment