Overview
Agentic LLM workflows — systems where a language model orchestrates tools, APIs, and code — are powerful but brittle. This guide shows concrete, practical patterns to make those workflows reliable: defensive tool design, retry and timeout strategies, caching, test harnesses, and lightweight monitoring. Examples use LangChain-style agents and plain Python so you can apply the patterns with other frameworks (including LangGraph-style orchestration).
Why reliability matters
Common production failures for agentic systems include unexpected tool outputs, long-running tool executions, high token costs, hallucinations, and cascading failures when a tool is down. Plan for these early by treating each tool as a micro-dependency with its own SLA, validation, and fallback.
Common failure modes
- Tool crashes or network timeouts
- LLM hallucinations calling the wrong tool or inventing results
- Expensive repeated calls for identical prompts
- Side-effecting tools executed multiple times (non-idempotent)
- Observability gaps: no logs or metrics for agent decisions
Pattern 1 — Compose agents with guarded tools
Treat every tool as a small service: validate inputs, sanitize outputs, and sandbox side effects. Keep the interface simple (string in, string out) and serialize richer data via JSON schemas when needed.
from langchain.llms import OpenAI
from langchain.agents import initialize_agent, load_tools
# Keep the LLM deterministic for decision-making where possible
llm = OpenAI(temperature=0)
# Load only the tools you need; keep tool responsibilities small
tools = load_tools(["serpapi", "python_repl"], llm=llm)
# Create the agent executor (zero-shot React is a common starting agent)
agent = initialize_agent(tools, llm, agent="zero-shot-react-description", verbose=False)Pattern 2 — Retries, exponential backoff, and timeouts
Wrap calls to agents and tools with retries and bounded timeouts. Use exponential backoff to avoid thundering herds and cap attempts to contain cost.
from tenacity import retry, stop_after_attempt, wait_exponential
import concurrent.futures
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
def run_query_with_retry(prompt: str) -> str:
return agent.run(prompt)
def run_with_timeout(prompt: str, timeout: int = 15) -> str:
# Run the agent in a thread and enforce a timeout at the call boundary
with concurrent.futures.ThreadPoolExecutor(max_workers=1) as ex:
future = ex.submit(agent.run, prompt)
return future.result(timeout=timeout)Pattern 3 — Cache deterministic results
Cache outputs for deterministic prompts to save cost and latency. Use a request fingerprint that includes prompt template + tool versions.
from functools import lru_cache
@lru_cache(maxsize=128)
def cached_run(prompt: str) -> str:
# Combine caching with retries and a bounded timeout
return run_with_timeout(prompt, timeout=15)Pattern 4 — Sanitize and validate tool I/O
Validate tool inputs and outputs with small JSON schemas or type checks. Reject or transform unexpected tool outputs before they reach downstream code.
import jsonschema
RESULT_SCHEMA = {
"type": "object",
"properties": {
"answer": {"type": "string"},
"source": {"type": "string"}
},
"required": ["answer"]
}
def validate_tool_output(out: dict) -> dict:
jsonschema.validate(instance=out, schema=RESULT_SCHEMA)
return outPattern 5 — Test harnesses and deterministic unit tests
Mock LLM and tool calls to create deterministic tests for agent decision logic. Focus on the agent's orchestration behavior (which tool to call, how many steps) rather than the LLM's raw text.
from unittest.mock import patch
def test_agent_chooses_search_tool():
prompt = "What is the population of Tokyo?"
# Patch the agent instance's run method to make the test deterministic
with patch.object(agent, "run", return_value="Tokyo population is ~37 million") as mock_run:
result = agent.run(prompt)
mock_run.assert_called_once_with(prompt)
assert "Tokyo population" in resultFor integration tests, run the agent against a recording of tool responses (a cassette) rather than live APIs, and verify the call sequence and final output.
Operational checklist
- Log: request id, prompt template version, tool invocations, and final answer.
- Emit metrics: latency, error rate, tool failure counts, retry counts, token usage per call.
- Guardrails: enforce prompt length, limit tool output sizes, and sandbox code-execution tools.
- Idempotency: require explicit confirmation for side-effecting tools or use a two-step confirmation flow.
- Cost controls: set per-request token limits and global rate limits.
Tradeoffs
- Latency vs reliability: Retries and validation add latency, but reduce downstream errors.
- Cost vs correctness: More calls (retries, tool calls) cost more; caching helps but can staleness-proof your data.
- Complexity vs transparency: Adding orchestration (state machines, graphs) improves control but increases maintenance.
Further reading
For practical ideas on building reliable agent workflows, see the community write-up on LangChain and LangGraph: LangChain and LangGraph: Building Reliable Agentic AI Workflows.
Conclusion
Make agentic workflows reliable by treating tools as first-class dependencies: validate inputs/outputs, add retries and timeouts, cache deterministic results, and test orchestration behaviour with mocks and recorded tool responses. These patterns reduce surprises in production and give you predictable, debuggable agent behavior.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment