Why reliability matters for agentic LLM systems
As LLM-based agents move from experiments into production, reliability is no longer optional. Agents that lose state, repeat work, or fail silently quickly become unusable in real workflows. This post gives pragmatic architecture patterns and a small, practical Python example you can adapt to your stack: a durable state store, checkpoints, idempotent actions, retriable LLM calls, and a recovery path.
Core design pillars
- Durable state — Keep minimal structured state in a persistent store (DB, KV, object store) instead of only in memory.
- Checkpoints — Persist periodic snapshots and progress markers so the agent can resume from known-good points.
- Idempotency — Ensure action handlers can be applied multiple times without harmful side effects.
- Failure domains — Separate transient LLM/API failures from persistent state corruption; defend each differently.
- Observability — Emit structured logs, metrics, and action history to reason about recovery and to debug misbehavior.
Minimal architecture sketch
- API / worker runs the agent loop.
- State store (SQL/NoSQL/Redis) stores: current task, history, checkpoints, and a monotonic progress counter.
- Action handlers are idempotent and reference a request ID or sequence number.
- LLM calls are wrapped with retry/backoff and timeout logic; responses persist before side effects execute.
- A recovery routine can rehydrate state and replay or skip actions based on the progress counter.
Practical Python example (SQLite for portability)
The example below shows a tiny agent that keeps state in SQLite, writes checkpoints, wraps LLM calls with retries, and recovers on restart. Replace the call_llm stub with your actual model/API call and adapt storage to Postgres/Redis when you need concurrency.
1) Schema and helpers
import sqlite3
import json
import time
import random
from typing import Dict, Any
DB_PATH = 'agent_state.db'
def init_db():
conn = sqlite3.connect(DB_PATH)
cur = conn.cursor()
cur.execute('''
CREATE TABLE IF NOT EXISTS agent_state (
id TEXT PRIMARY KEY,
data TEXT,
last_updated REAL,
checkpoint INTEGER
)
''')
# action log for audit/replay
cur.execute('''
CREATE TABLE IF NOT EXISTS action_log (
seq INTEGER PRIMARY KEY AUTOINCREMENT,
action TEXT,
payload TEXT,
result TEXT,
created_at REAL
)
''')
conn.commit()
conn.close()
def load_state(agent_id: str) -> Dict[str, Any]:
conn = sqlite3.connect(DB_PATH)
cur = conn.cursor()
cur.execute('SELECT data FROM agent_state WHERE id=?', (agent_id,))
row = cur.fetchone()
conn.close()
return json.loads(row[0]) if row else {}
def save_state(agent_id: str, state: Dict[str, Any], checkpoint: int = 0):
payload = json.dumps(state)
now = time.time()
conn = sqlite3.connect(DB_PATH)
cur = conn.cursor()
cur.execute('INSERT OR REPLACE INTO agent_state(id, data, last_updated, checkpoint) VALUES (?, ?, ?, ?)',
(agent_id, payload, now, checkpoint))
conn.commit()
conn.close()2) Retriable LLM wrapper and idempotent action logging
def retry_backoff(func, max_attempts=5, base_delay=0.5):
def wrapper(*args, **kwargs):
attempt = 0
while True:
try:
return func(*args, **kwargs)
except Exception as e:
attempt += 1
if attempt >= max_attempts:
raise
delay = base_delay * (2 ** (attempt - 1)) + random.random() * 0.1
time.sleep(delay)
return wrapper
# Stub: replace with your LLM client call (OpenAI/Anthropic/LLama, etc.)
@retry_backoff
def call_llm(prompt: str, timeout: float = 10.0) -> str:
# Example stub that simulates transient errors
if random.random() < 0.2:
raise RuntimeError('transient model error')
return 'LLM response for: ' + prompt
def log_action(action: str, payload: Dict[str, Any], result: Dict[str, Any] = None):
conn = sqlite3.connect(DB_PATH)
cur = conn.cursor()
cur.execute('INSERT INTO action_log(action, payload, result, created_at) VALUES (?, ?, ?, ?)',
(action, json.dumps(payload), json.dumps(result) if result else None, time.time()))
conn.commit()
conn.close()3) Agent loop with checkpoints and recovery
class SimpleAgent:
def __init__(self, agent_id: str):
self.agent_id = agent_id
self.state = load_state(agent_id) or {'task_queue': [], 'vars': {}, 'progress': 0}
def checkpoint(self):
# Increment a monotonic checkpoint counter before critical side effects
self.state['progress'] = self.state.get('progress', 0) + 1
save_state(self.agent_id, self.state, checkpoint=self.state['progress'])
def add_task(self, task: str):
self.state['task_queue'].append(task)
save_state(self.agent_id, self.state)
def run_step(self):
if not self.state['task_queue']:
return False
task = self.state['task_queue'].pop(0)
payload = {'task': task, 'progress': self.state.get('progress', 0) + 1}
# 1) Call LLM safely
try:
result_text = call_llm(task)
except Exception as e:
# transient failure handling: push task back and raise for orchestrator to retry later
self.state['task_queue'].insert(0, task)
save_state(self.agent_id, self.state)
raise
# 2) Persist result before performing side effects (idempotency point)
result = {'response': result_text}
log_action('llm_call', payload, result)
# 3) Checkpoint to mark progress (so we can safely replay or skip later)
self.checkpoint()
# 4) Execute idempotent side effect (example: write to external DB with request id)
# Here we just record into state.vars as a deterministic update
self.state['vars'][f'res_{self.state["progress"]}'] = result_text
save_state(self.agent_id, self.state)
return True
def recover_and_run(self, max_steps=100):
# On startup, state already loaded; you could verify action_log vs progress for consistency
steps = 0
while steps < max_steps and self.run_step():
steps += 1
if __name__ == '__main__':
init_db()
agent = SimpleAgent('agent-1')
# simulate adding tasks
if not agent.state['task_queue']:
agent.add_task('Summarize user message A')
agent.add_task('Extract entities from message B')
try:
agent.recover_and_run()
except Exception as e:
# In a production system, push to retry queue or alerting pipeline
print('Agent paused due to transient error:', e)Why this pattern works
- Persisting LLM responses and progress before side effects lets you replay or skip operations during recovery.
- Idempotent handlers (or attaching a unique request id) prevent double-application of side effects.
- Separating transient retry logic (for LLM calls) from persistent state updates keeps failure handling local and predictable.
Tradeoffs and when to use what
- SQLite vs Postgres/Redis: SQLite is great for local tests and single-process agents. For concurrent workers, use Postgres with row-level locking or Redis streams/consumer groups for at-least-once delivery and coordination.
- Checkpoint frequency: Frequent checkpoints increase durability at the cost of write amplification. Pick a cadence based on how expensive it is to replay work.
- Strong consistency vs throughput: If you need strict ordering and exactly-once semantics, invest in transactional stores and distributed locks — but expect higher latency.
- State shape: Keep the canonical state minimal (task pointer, little metadata). Store large artifacts externally (blob store) and reference them from state to avoid big DB rows.
Testing and observability
- Unit-test action handlers with injected mock stores and mocked LLM responses, including failure scenarios.
- Replay tests: seed the action_log and checkpoint, restart the agent, and assert final state matches expectations.
- Instrument metrics: task throughput, retry rate, checkpoint latency, and action-log growth.
- Include structured logging with request IDs and progress counters — invaluable during recovery debugging.
Concise conclusion
Reliable agents require deliberate choices: durable minimal state, checkpoints that mark progress, idempotent handlers, and retriable LLM calls. Start small with a local durable store for development and gradually swap in more robust systems (Postgres, Redis streams, or an event store) as concurrency and scale demands grow. The pattern above gives a pragmatic migration path: snapshot, persist, then act.
Further reading: see a practical discussion of agent architecture and failure recovery in the community article Practical Agent Architecture.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment