Sechno
Software Engineering

How to Build Robust AI Agents in Python: Architecture, Patterns, and a Practical Skeleton

A pragmatic guide for developers to design, implement, test, and operate production-ready AI agents driven by large models. Covers architecture, memory, tools, retrieval, safety, testing, and a compact Python skeleton you can extend.

SSechno Team 5 min read 68 views
How to Build Robust AI Agents in Python: Architecture, Patterns, and a Practical Skeleton

Why this matters

Short demos that create an "AI agent in 50 lines of Python" are useful for exploration, but they gloss over the engineering needed to make an agent reliable, auditable, and maintainable. This article gives practical, implementable patterns you can use to move from a toy agent to a production-capable system.

Core architecture and components

Design agents as orchestrations of interoperable components rather than monolithic scripts. Key parts:

  • LLM adapter: a thin client that wraps your LLM API, handles rate limits, retries, and response normalization.
  • Orchestrator / loop: the central controller that builds prompts, parses actions, invokes tools, and updates memory.
  • Tools: side-effecting capabilities exposed via well-defined adapters (search, browser, DB queries, OS commands, APIs).
  • Memory & retrieval: long- and short-term memory stores (embedding index, SQL/kv store) that the orchestrator queries for context.
  • Safety & guardrails: validators that sanitize tool inputs and check outputs before exposing them to users or systems.
  • Observability: structured logs, traces, and deterministic test harnesses for reproducible debugging.

Minimal, practical agent skeleton (Python)

This skeleton shows the orchestrator, a simple LLM call, a tool dispatcher, and memory updates. It focuses on structure you can extend — keep in mind you should replace llm_api_call, persistence, and tool implementations with production-grade code and credential handling.

import json
import time
 
# LLM adapter (placeholder)
def llm_api_call(prompt):
    # Replace with your provider SDK call, handle auth/rate limits
    # Return a text response
    return '{"action": "search", "input": "Python decorators tutorial"}'
 
# Simple parser: expects JSON response describing an action
def parse_action(text):
    try:
        return json.loads(text)
    except Exception:
        raise ValueError("LLM response not valid JSON action")
 
# Example tools
def search_tool(query):
    # Implement actual search or vector retrieval
    return f"results_for({query})"
 
def open_url_tool(url):
    # Implement safe fetch with sanitization
    return f"fetched({url})"
 
# Dispatcher
def run_tool(action_name, payload):
    if action_name == "search":
        return search_tool(payload)
    if action_name == "open_url":
        return open_url_tool(payload)
    raise ValueError(f"Unknown tool: {action_name}")
 
# Very small memory example (in-memory for demo)
class Memory:
    def __init__(self):
        self.events = []
    def add(self, item):
        self.events.append({"ts": time.time(), "item": item})
    def recall(self, last_k=5):
        return [e["item"] for e in self.events[-last_k:]]
 
# Orchestrator (loop)
def run_agent(task, max_turns=5):
    memory = Memory()
    observation = None
    for turn in range(max_turns):
        prompt = f"Task: {task}\nMemory: {memory.recall()}\nObservation: {observation}\nRespond with a JSON action."
        raw = llm_api_call(prompt)
        try:
            action = parse_action(raw)
        except Exception as e:
            # log and continue or fallback
            memory.add({"error": str(e), "raw": raw})
            break
        result = run_tool(action.get("action"), action.get("input"))
        memory.add({"action": action, "result": result})
        observation = result
        if action.get("action") == "finish":
            break
    return memory.events
 
if __name__ == "__main__":
    print(run_agent("Research Python decorators"))

Notes on the skeleton

  • Keep the LLM adapter thin: separate retries, throttling, caching, and response normalization from orchestration logic.
  • Prefer machine-readable action formats (JSON) to make parsing deterministic and testable.
  • Limit loop iterations and enforce timeouts to avoid runaway behaviors.

Memory and retrieval

For anything beyond ephemeral conversational context, use a persistent store and a vector index for semantic recall. Workflow:

  1. Store documents, tool outputs, and important observations with timestamps and metadata.
  2. Index semantic embeddings for documents you want to retrieve later.
  3. At prompt time, retrieve a small set of relevant items and include them in the prompt as context.

Keep retrieval results short and rank by recency/importance. Always include provenance (source id, timestamp) to support audits.

Testing, observability, and safe deployment

  • Unit tests: mock the LLM adapter and each tool. Assert orchestrator calls the expected tool and updates memory.
  • Integration tests: run the orchestrator against sandboxed tool implementations or fixtures.
  • Deterministic prompts: for regression tests, freeze the LLM output using recorded responses or deterministic LLM modes if available.
  • Logging & traces: log prompts, LLM responses, parsed actions, tool inputs/outputs (redact secrets), and final decisions.
  • Runtime guards: input sanitizers, tool call whitelists, rate limits, and human-in-the-loop approvals for sensitive actions.

Retry helper (exponential backoff)

import time
 
def retry(fn, retries=3, base_delay=0.5):
    for attempt in range(retries):
        try:
            return fn()
        except Exception:
            if attempt + 1 == retries:
                raise
            time.sleep(base_delay * (2 ** attempt))

Prompt engineering and parsing

Two complementary approaches reduce brittleness:

  • Structured output instructions: tell the LLM to reply only with a JSON block and provide a schema.
  • Schema validation: validate and sanitize the parsed JSON before using it to invoke tools.

Fail open for low-impact informational actions (log and continue) and fail closed for destructive actions (block and escalate).

Tradeoffs

  • Latency vs capability: adding retrieval and tool calls improves correctness but increases latency and cost.
  • Simplicity vs control: simple prompt-based agents are easy to prototype but offer limited observability compared to orchestrated architectures.
  • Memory freshness vs cost: frequent embedding and reindexing improves recall but adds operational overhead.
  • Automation vs safety: more automation reduces manual work but raises risk; use progressive automation and human approval gates.

Further reading and context

If you started from a short demo or tutorial, the next steps are to add persistence, tool validation, and deterministic tests. For background on retrieval and semantic relevance, see the discussion on semantic search and how teams interpret it in practical systems: What (un)exactly do you mean by semantic search?

Conclusion

To move an AI agent from a toy demo to a reliable component of your stack, separate concerns: an LLM adapter, an orchestrator, explicit tools, memory & retrieval, and safety & observability layers. Start with the skeleton above, add persistence and validation, write deterministic tests, and iterate on guardrails and monitoring. These patterns keep agents maintainable, auditable, and safer as they take on real tasks.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment