Sechno
Ai

How to Test and Harden Applications That Use LLMs and Agents

Practical patterns for testing, mocking, validating, and protecting backends that call large language models or orchestrate AI agents — with PHP examples for deterministic tests, contract checks, and circuit breakers.

SSechno Team 5 min read 91 views
How to Test and Harden Applications That Use LLMs and Agents

Why this matters

AI models and agent frameworks change the failure modes and observability of your stack. Traditional unit and integration tests often assume deterministic dependencies; LLM responses are non-deterministic, and agent orchestration can introduce latency, hallucinations, or new transient errors. This guide gives pragmatic techniques to make LLM-backed features testable, reliable, and auditable.

Core strategies

  • Deterministic testing via recorded responses (cassettes) and mocks
  • Input/output contract validation and strict schemas
  • Resilience patterns: timeouts, retries with backoff, and circuit breakers
  • Observability and contract-based assertions in CI
  • Property tests and fuzzing for edge-case prompts

1) Make tests deterministic: mocks and cassettes

Record real API responses during a controlled run, then replay them in unit and CI tests. For HTTP-driven LLMs you can use an HTTP client mock. Example using Guzzle's MockHandler:

<?php
use GuzzleHttp\Client;
use GuzzleHttp\Handler\MockHandler;
use GuzzleHttp\HandlerStack;
use GuzzleHttp\Psr7\Response;
 
// In your test bootstrap
$mock = new MockHandler([
    new Response(200, [], json_encode(['id' => 'resp_1', 'text' => 'Hello from mock'])),
]);
 
$handler = HandlerStack::create($mock);
$client = new Client(['handler' => $handler, 'base_uri' => 'https://api.llm.example/']);
 
$response = $client->post('/v1/complete', ['json' => ['prompt' => 'hi']]);
$body = json_decode($response->getBody(), true);
echo $body['text'];

Use this pattern for unit tests so they don't hit live LLMs, keeping tests fast and repeatable.

2) Replay (cassette) pattern for integration tests

Store responses keyed by a hash of the prompt + options so tests can exercise end-to-end code paths while remaining deterministic.

<?php
function call_llm_with_cassette(string $prompt): array {
    $key = sha1($prompt);
    $cassette = __DIR__ . '/cassettes/' . $key . '.json';
 
    if (file_exists($cassette)) {
        return json_decode(file_get_contents($cassette), true);
    }
 
    // Replace with your real request code; keep a single I/O boundary
    $resp = real_llm_request($prompt);
    if (!is_dir(dirname($cassette))) { mkdir(dirname($cassette), 0775, true); }
    file_put_contents($cassette, json_encode($resp, JSON_PRETTY_PRINT));
    return $resp;
}
 
function real_llm_request(string $prompt): array {
    // Call the actual LLM provider; keep small, isolated function for easier mocking
    return ['id' => 'live_1', 'text' => 'live response for: ' . $prompt];
}

Record once in an approved environment, then commit cassettes to your test suite. This also documents expected outputs for audit purposes.

3) Contract and schema validation

Validate model outputs against an explicit schema before using them in production. A schema reduces hallucination impact and makes errors testable.

<?php
// Minimal output assertion without external deps
function validate_completion(array $resp): bool {
    if (!isset($resp['id']) || !is_string($resp['id'])) return false;
    if (!isset($resp['text']) || !is_string($resp['text'])) return false;
    // Optional: check length, JSON structure, or parse as JSON if model returns serialized objects
    return true;
}
 
$resp = ['id' => 'resp_1', 'text' => '{"action":"approve","value":42}'];
if (!validate_completion($resp)) {
    throw new RuntimeException('LLM response failed schema validation');
}

For structured outputs prefer JSON schema validators and require the model to emit a strict JSON object. Enforce this in tests and reject unexpected shapes at runtime.

4) Resilience: timeouts, retries, and circuit breakers

LLM providers can return transient errors or slow responses. Add short client-side timeouts, limited retries with jitter, and a circuit breaker to prevent cascading failures.

<?php
class SimpleCircuitBreaker {
    private $failures = 0;
    private $lastFailureTs = 0;
    private $threshold;
    private $timeout;
 
    public function __construct(int $threshold = 5, int $timeout = 60) {
        $this->threshold = $threshold;
        $this->timeout = $timeout;
    }
 
    public function recordFailure(): void {
        $this->failures++;
        $this->lastFailureTs = time();
    }
 
    public function recordSuccess(): void {
        $this->failures = 0;
    }
 
    public function allowRequest(): bool {
        if ($this->failures < $this->threshold) return true;
        return (time() - $this->lastFailureTs) >= $this->timeout;
    }
}
 
// Usage
$cb = new SimpleCircuitBreaker(3, 30);
if (! $cb->allowRequest()) {
    throw new RuntimeException('LLM circuit open — aborting request');
}
try {
    $resp = real_llm_request('prompt');
    $cb->recordSuccess();
} catch (Exception $e) {
    $cb->recordFailure();
    throw $e;
}

For production use, back the breaker state with Redis or a durable store to keep state across processes.

5) Test for hallucinations and brittle behavior

  • Write property tests that mutate prompts and ensure key assertions hold (e.g., no unexpected keys, numeric fields remain in range).
  • Include negative tests: assert that clearly impossible facts are rejected or cause safe fallback behavior.
  • Use adversarial/fuzz inputs in CI to detect brittle parsing or brittle prompt templates.

6) Observability and CI checks

Make LLM behavior visible:

  • Log prompt and response hashes (not always full prompts when sensitive).
  • Add contract assertions as automated checks in CI. If a recorded cassette is updated, require human review.
  • Record drift metrics: response length, parsing failure rate, schema violation counts.

Tradeoffs

  • Deterministic tests (cassettes) can mask real drift. Periodically refresh recordings and review diffs.
  • Strict schemas reduce hallucinations but can make responses brittle; design schemas that allow safe extension.
  • Circuit breakers and short timeouts improve stability but may reduce availability for slow-but-correct responses; tune thresholds carefully.

Conclusion

LLMs and agents introduce new failure modes, but established testing and resilience patterns still apply. Use mocks and cassettes for deterministic tests, enforce strict I/O contracts, add resilience at the I/O boundary, and make model behavior observable in CI. These practices make AI features safer to ship and easier to maintain.

Further reading: the "Hello Agents" developer discussion explores agent architectures, and industry posts on testing opaque dependencies provide useful background for integrating these techniques into CI. See source notes below.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment