Why this matters
Short demos that create an "AI agent in 50 lines of Python" are useful for exploration, but they gloss over the engineering needed to make an agent reliable, auditable, and maintainable. This article gives practical, implementable patterns you can use to move from a toy agent to a production-capable system.
Core architecture and components
Design agents as orchestrations of interoperable components rather than monolithic scripts. Key parts:
- LLM adapter: a thin client that wraps your LLM API, handles rate limits, retries, and response normalization.
- Orchestrator / loop: the central controller that builds prompts, parses actions, invokes tools, and updates memory.
- Tools: side-effecting capabilities exposed via well-defined adapters (search, browser, DB queries, OS commands, APIs).
- Memory & retrieval: long- and short-term memory stores (embedding index, SQL/kv store) that the orchestrator queries for context.
- Safety & guardrails: validators that sanitize tool inputs and check outputs before exposing them to users or systems.
- Observability: structured logs, traces, and deterministic test harnesses for reproducible debugging.
Minimal, practical agent skeleton (Python)
This skeleton shows the orchestrator, a simple LLM call, a tool dispatcher, and memory updates. It focuses on structure you can extend — keep in mind you should replace llm_api_call, persistence, and tool implementations with production-grade code and credential handling.
import json
import time
# LLM adapter (placeholder)
def llm_api_call(prompt):
# Replace with your provider SDK call, handle auth/rate limits
# Return a text response
return '{"action": "search", "input": "Python decorators tutorial"}'
# Simple parser: expects JSON response describing an action
def parse_action(text):
try:
return json.loads(text)
except Exception:
raise ValueError("LLM response not valid JSON action")
# Example tools
def search_tool(query):
# Implement actual search or vector retrieval
return f"results_for({query})"
def open_url_tool(url):
# Implement safe fetch with sanitization
return f"fetched({url})"
# Dispatcher
def run_tool(action_name, payload):
if action_name == "search":
return search_tool(payload)
if action_name == "open_url":
return open_url_tool(payload)
raise ValueError(f"Unknown tool: {action_name}")
# Very small memory example (in-memory for demo)
class Memory:
def __init__(self):
self.events = []
def add(self, item):
self.events.append({"ts": time.time(), "item": item})
def recall(self, last_k=5):
return [e["item"] for e in self.events[-last_k:]]
# Orchestrator (loop)
def run_agent(task, max_turns=5):
memory = Memory()
observation = None
for turn in range(max_turns):
prompt = f"Task: {task}\nMemory: {memory.recall()}\nObservation: {observation}\nRespond with a JSON action."
raw = llm_api_call(prompt)
try:
action = parse_action(raw)
except Exception as e:
# log and continue or fallback
memory.add({"error": str(e), "raw": raw})
break
result = run_tool(action.get("action"), action.get("input"))
memory.add({"action": action, "result": result})
observation = result
if action.get("action") == "finish":
break
return memory.events
if __name__ == "__main__":
print(run_agent("Research Python decorators"))Notes on the skeleton
- Keep the LLM adapter thin: separate retries, throttling, caching, and response normalization from orchestration logic.
- Prefer machine-readable action formats (JSON) to make parsing deterministic and testable.
- Limit loop iterations and enforce timeouts to avoid runaway behaviors.
Memory and retrieval
For anything beyond ephemeral conversational context, use a persistent store and a vector index for semantic recall. Workflow:
- Store documents, tool outputs, and important observations with timestamps and metadata.
- Index semantic embeddings for documents you want to retrieve later.
- At prompt time, retrieve a small set of relevant items and include them in the prompt as context.
Keep retrieval results short and rank by recency/importance. Always include provenance (source id, timestamp) to support audits.
Testing, observability, and safe deployment
- Unit tests: mock the LLM adapter and each tool. Assert orchestrator calls the expected tool and updates memory.
- Integration tests: run the orchestrator against sandboxed tool implementations or fixtures.
- Deterministic prompts: for regression tests, freeze the LLM output using recorded responses or deterministic LLM modes if available.
- Logging & traces: log prompts, LLM responses, parsed actions, tool inputs/outputs (redact secrets), and final decisions.
- Runtime guards: input sanitizers, tool call whitelists, rate limits, and human-in-the-loop approvals for sensitive actions.
Retry helper (exponential backoff)
import time
def retry(fn, retries=3, base_delay=0.5):
for attempt in range(retries):
try:
return fn()
except Exception:
if attempt + 1 == retries:
raise
time.sleep(base_delay * (2 ** attempt))Prompt engineering and parsing
Two complementary approaches reduce brittleness:
- Structured output instructions: tell the LLM to reply only with a JSON block and provide a schema.
- Schema validation: validate and sanitize the parsed JSON before using it to invoke tools.
Fail open for low-impact informational actions (log and continue) and fail closed for destructive actions (block and escalate).
Tradeoffs
- Latency vs capability: adding retrieval and tool calls improves correctness but increases latency and cost.
- Simplicity vs control: simple prompt-based agents are easy to prototype but offer limited observability compared to orchestrated architectures.
- Memory freshness vs cost: frequent embedding and reindexing improves recall but adds operational overhead.
- Automation vs safety: more automation reduces manual work but raises risk; use progressive automation and human approval gates.
Further reading and context
If you started from a short demo or tutorial, the next steps are to add persistence, tool validation, and deterministic tests. For background on retrieval and semantic relevance, see the discussion on semantic search and how teams interpret it in practical systems: What (un)exactly do you mean by semantic search?
Conclusion
To move an AI agent from a toy demo to a reliable component of your stack, separate concerns: an LLM adapter, an orchestrator, explicit tools, memory & retrieval, and safety & observability layers. Start with the skeleton above, add persistence and validation, write deterministic tests, and iterate on guardrails and monitoring. These patterns keep agents maintainable, auditable, and safer as they take on real tasks.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment