Why this matters
Modern AI agents (chat assistants, automation scripts, or domain-specific bots) are moving from monolithic prompts to modular architectures. Breaking an agent into a deterministic agent loop, a composable tool system, and a strict permission model improves safety, observability, and maintainability. This article gives practical patterns and examples you can adapt to production systems.
Core components and responsibilities
- Agent loop — orchestrates planning, tool selection, execution, memory updates, and termination checks.
- Tool system — a registry of narrow, testable capabilities (e.g., search, DB query, file ops) with clear input/output contracts.
- Permission model — enforces what each agent instance may do (which tools, which data scopes, rate limits, and audit hooks).
Implementing a clear agent loop
Make the loop explicit and deterministic. Keep reasoning/planning (LLM calls) and side-effecting operations (tool calls) separated so you can test, mock, and audit each step.
// Pseudocode-style agent loop (adapt to your language/framework)
function runAgent(agentContext) {
while (!agentContext.done) {
// 1) Planner: produce an instruction describing next step
const instruction = planner.plan(agentContext); // pure function using memory + request
// 2) Router: pick an applicable tool and validate permissions
const toolId = router.selectTool(instruction);
if (!permissions.isAllowed(agentContext.agentId, toolId, instruction)) {
audit.log("permission_denied", {agent: agentContext.agentId, tool: toolId});
throw new Error("permission denied");
}
// 3) Executor: call the tool, capture result and metadata
const result = tools.execute(toolId, instruction.input);
// 4) Update memory/state with result for subsequent planning
agentContext = memory.update(agentContext, instruction, result);
// 5) Termination check
if (planner.shouldStop(agentContext)) {
agentContext.done = true;
}
}
return agentContext.output;
}Notes
- Keep each component small and independently testable. Mock tools when unit testing planners.
- Log the planner's rationale (or a hashed summary) so you can later correlate tool calls to decisions.
Designing a pluggable tool system
Design tools with explicit input/output schemas and idempotent, side-effect-minimised behaviour where possible. A centralized registry helps enforce contracts and feature flags.
// Example tool interface and registry (pseudo-TypeScript in structure)
interface Tool {
id: string;
description?: string;
inputSchema?: object; // JSON Schema for validation
run(input: any, ctx: {agentId: string, traceId: string}): PromisePractical patterns
- Validate inputs against JSON Schema before executing to avoid tool misuse.
- Wrap tools with a middleware that adds timeouts, retries, and resource constraints.
- Prefer narrow tools (search, summarizer, safe-executor) over “run arbitrary code” tools.
Permission model and sandboxing
Permissions are the system’s safety gate. Model them per-agent or per-session and cover these dimensions: allowed tools, allowed data scopes (which DBs, which buckets), rate limits, and soft vs hard denials.
/* Example permission document (JSON-like) */
{
"agent_id": "billing-bot-v1",
"allowed_tools": ["db_query","invoice_pdf_generator","logger"],
"data_scopes": {
"db_query": ["invoices_read_only"],
"invoice_pdf_generator": ["s3:invoices-bucket"]
},
"rate_limits": {
"db_query": {"per_minute": 60},
"invoice_pdf_generator": {"per_minute": 10}
}
}Enforcement techniques
- Use a policy engine (e.g., OPA/Rego) or a compact internal evaluator to decide allowed actions before tool execution.
- Implement runtime guards: middleware that checks the permission document and rejects calls that exceed scope.
- For risky tools (shell execution, network access), run inside containerized sandboxes or separate processes with least privilege.
Audit logging, observability, and testing
Logging and testability are essential for safety and iteration.
-- Example SQL audit table schema
CREATE TABLE agent_audit_logs (
id UUID PRIMARY KEY,
timestamp TIMESTAMPTZ NOT NULL,
agent_id TEXT NOT NULL,
trace_id TEXT NOT NULL,
planner_rationale TEXT, -- optional hashed or tokenized
tool_id TEXT,
tool_input JSONB,
tool_output JSONB,
result_status TEXT,
error TEXT
);- Log planner outputs and tool calls with trace IDs so you can replay incidents.
- Use deterministic planner modes or fixed seeds during CI tests to assert behaviour.
- Perform chaos testing on rate limits and tool failures to ensure graceful degradation.
Tradeoffs and practical advice
- Simplicity vs flexibility: Narrow tools reduce risk but require more upfront engineering. Broad tools enable faster prototyping but increase surface for misuse.
- Latency vs auditability: Synchronous permission checks/audits add latency. Cache safe-permission decisions for short windows to reduce overhead.
- Determinism vs creativity: For repeatable automation, prefer deterministic planning or constrained LLM prompts; for exploratory assistants, allow softer constraints and stricter post-execution audits.
Quick checklist to ship a safe agent feature
- Define tool contracts (input/output schemas).
- Implement a registry and middleware for timeouts and retries.
- Create a permission document per agent and enforce it at runtime.
- Instrument planner rationale and tool calls with trace IDs.
- Add unit tests with mocked tools and CI replay tests for common flows.
Further reading
For a deeper architectural discussion inspired by recent work on modular agent frameworks, see this analysis: Claude Code Architecture Explained.
Conclusion
Splitting an AI agent into an explicit agent loop, a composable tool system, and an enforceable permission model makes your system easier to test, safer to run, and simpler to evolve. Start small: add schemas and runtime guards, instrument everything with trace IDs, and iterate toward stricter policies as you gain operational confidence.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment