Unit Tests for Agent Behavior: How We Validate Swarm Actions Before They Touch Production
Over the past year, we've run over 200 agent swarms in production at any given time. The single biggest lesson? Without unit tests for agent behavior, you cannot trust your swarm.
LLM outputs are non-deterministic. Prompts drift. Tool APIs change. If you only rely on integration or end-to-end tests, you'll discover regressions after they've already cost you compute and user trust.
We built a testing framework that treats agent actions as first-class units. Here's exactly how we do it.
Why Standard Unit Tests Fall Short
Traditional unit tests assert on return values. For agents, the "return value" is a sequence of actions—tool calls, messages, state transitions. You can't just assert on the final output because the path matters.
Consider a customer support agent that should escalate if sentiment drops below 0.3. If you only test the final response, you might miss that it called the wrong tool or skipped a required escalation step.
We need to assert on:
- Action sequence: Did the agent call tools in the right order?
- Action parameters: Were the arguments correct?
- State transitions: Did internal state update as expected?
- Guardrails: Did the agent stay within allowed boundaries?
Architecture of Our Test Harness
We use a deterministic simulator that replays agent logic with mocked LLM completions. The core components:
- MockLLM: Returns pre-recorded or rule-based completions.
- ActionRecorder: Intercepts every tool call and message.
- StateSnapshot: Captures agent state at each step.
- AssertionBuilder: Fluent API for action-level checks.
Here's the skeleton in Python:
from dataclasses import dataclass, field
from typing import List, Dict, Any
@dataclass
class Action:
type: str # 'tool_call', 'message', 'state_change'
name: str
arguments: Dict[str, Any] = field(default_factory=dict)
timestamp: float = 0.0
class MockLLM:
def __init__(self, responses: List[str]):
self.responses = responses
self.index = 0
def complete(self, prompt: str) -> str:
resp = self.responses[self.index]
self.index += 1
return resp
class ActionRecorder:
def __init__(self):
self.actions: List[Action] = []
def record(self, action: Action):
self.actions.append(action)Writing Your First Agent Unit Test
Let's test a simple agent that should call search_knowledge_base when asked a question, then respond with the result.
First, we define the mock LLM responses. For the first call, we want the agent to decide to use the tool. The mock returns a completion that includes a function call.
mock_llm = MockLLM([
'{"function_call": {"name": "search_knowledge_base", "arguments": {"query": "password reset"}}}',
'{"content": "Here are the steps to reset your password..."}'
])
recorder = ActionRecorder()
agent = CustomerSupportAgent(llm=mock_llm, recorder=recorder)
agent.run("How do I reset my password?")
# Assert on action sequence
assert recorder.actions[0].type == 'tool_call'
assert recorder.actions[0].name == 'search_knowledge_base'
assert recorder.actions[0].arguments['query'] == 'password reset'
assert recorder.actions[1].type == 'message'
assert 'reset your password' in recorder.actions[1].arguments['content']This test is fast (~10ms) and deterministic. It catches regressions like the agent calling the wrong tool or forgetting to include the result.
Testing Guardrails
Guardrails are rules that constrain agent behavior. We test them the same way—by asserting that certain actions are NOT taken, or that state remains within bounds.
def test_agent_does_not_escalate_without_sentiment_check():
mock_llm = MockLLM([
'{"function_call": {"name": "escalate_to_human", "arguments": {}}}'
])
recorder = ActionRecorder()
agent = SupportAgent(llm=mock_llm, recorder=recorder, guardrails=[
Guardrail(require_sentiment_check_before_escalation=True)
])
with pytest.raises(GuardrailViolation):
agent.run("I'm very angry!")This test ensures that the agent cannot escalate without first checking sentiment. If a future prompt change causes the agent to skip the sentiment tool, the test fails.
Simulating State Transitions
Many agents maintain internal state (e.g., conversation history, user context). We need to verify that state updates correctly.
def test_state_updates_after_tool_call():
mock_llm = MockLLM([
'{"function_call": {"name": "lookup_user", "arguments": {"user_id": "42"}}}',
'{"content": "User found: Alice"}'
])
agent = UserLookupAgent(llm=mock_llm, recorder=recorder)
agent.run("Find user 42")
assert agent.state.current_user == {"id": 42, "name": "Alice"}
assert agent.state.steps_completed == 2Parametrized Tests for Prompt Drift
Prompt drift is inevitable. We combat it by parametrizing tests with different mock responses that simulate edge cases.
@pytest.mark.parametrize("mock_responses, expected_action_count", [
([...], 2),
([...], 3), # Agent might need extra clarification
])
def test_agent_handles_ambiguous_queries(mock_responses, expected_action_count):
mock_llm = MockLLM(mock_responses)
recorder = ActionRecorder()
agent = AmbiguousQueryAgent(llm=mock_llm, recorder=recorder)
agent.run("Tell me about X")
assert len(recorder.actions) == expected_action_countPerformance Matters
Our test suite runs 500+ agent unit tests in under 2 seconds. Key optimizations:
- No real LLM calls: MockLLM returns pre-defined strings.
- No I/O: All tool implementations are in-memory stubs.
- Parallel execution: Tests are stateless and can run in parallel.
Lessons Learned
- Mock at the LLM boundary, not the tool boundary. Mocking individual tool calls leads to brittle tests.
- Assert on actions, not just final output. The sequence is the behavior.
- Test guardrails as negative assertions. It's more important to verify what the agent doesn't do.
- Keep tests deterministic. Randomness in LLM responses must be eliminated.
- Write tests before you change prompts. When you modify a system prompt, the tests tell you what broke.
The Full Testing Pyramid for Agent Swarms
We use three layers:
- Unit tests (this article): Mock LLM, test action sequences.
- Integration tests: Real LLM but sandboxed tools.
- End-to-end tests: Full stack in staging.
Unit tests catch 80% of regressions. They run in CI on every commit. Integration tests run nightly. E2E tests run before release.
Conclusion
Unit testing agent behavior is not optional. Without it, you're flying blind. Our framework has saved us countless hours of debugging and prevented dozens of production incidents. Start with simple action assertions and build up.
The code snippets above are simplified but reflect our actual production test suite. Adapt them to your stack—the principles are language-agnostic.
Now go write some tests.