Benchmarking Agent Behavior: Why We Use Custom Scenarios Over Standard LLM Evals
When we started building autonomous agents, we reached for the usual evaluation suites: MMLU, HellaSwag, GSM8K. These benchmarks are excellent for measuring a model's static knowledge or basic reasoning in isolation. But they tell you almost nothing about how an agent behaves when it has to chain multiple tool calls, recover from errors, or decide when to ask for clarification. An agent is not a question-answering system; it's a decision-making system operating over time. Standard evals treat each turn as independent, missing the sequential dependencies that define real agentic behavior.
We've shifted to custom scenario-based benchmarks. This post explains why, and how we design and run them.
The Gap Between Static Evals and Agent Behavior
Standard LLM evals are designed for single-turn, well-defined tasks. MMLU asks a multiple-choice question; the model picks an answer. HellaSwag tests commonsense narrative completion. GSM8K checks math reasoning with a known final answer. These are valuable for model selection, but they don't capture:
- Multi-step planning: An agent may need to decompose a goal into subgoals, execute them in order, and adapt when a step fails.
- Tool use and integration: Calling an API, querying a database, or writing to a file—each tool has side effects, and the agent must manage state.
- Error recovery: A tool might return an error, a timeout, or unexpected data. The agent must decide whether to retry, rephrase, or escalate.
- Context management: Over a long conversation, the agent must remember earlier decisions and not lose track of the overall objective.
- Ambiguity handling: When instructions are underspecified, the agent should ask clarifying questions rather than guessing.
Standard evals provide a single number, but agent behavior is multidimensional. A model that scores 90% on MMLU might fail catastrophically in a simple web navigation task because it cannot recover from a 404 error.
Why We Build Custom Scenarios
Custom scenarios let us define the exact environment, tools, and success criteria for a task. Instead of measuring the model's internal knowledge, we measure its ability to accomplish a goal using the tools we provide. This is closer to how the agent will be used in production.
A scenario typically includes:
- A world definition: The initial state of the environment—what data exists, what tools are available, what constraints apply.
- A goal description: A natural language instruction that the agent must satisfy. The goal may be underspecified intentionally.
- A set of permitted actions: The tools the agent can invoke, each with a defined interface and possible outcomes.
- Success criteria: A deterministic check that evaluates whether the goal was achieved. This could be a database state, a file content, or a final answer.
- Failure modes: Common pitfalls that the agent should avoid, such as infinite loops, security violations, or cost overruns.
Because we control the scenario, we can inject specific challenges: ambiguous instructions, missing information, conflicting data, or time pressure. This gives us a fine-grained view of where the agent excels and where it breaks.
Designing a Scenario: A Concrete Example
Consider a scenario where an agent must manage a user's calendar. The agent has access to three tools: read_events(date), create_event(title, date, time, duration), and delete_event(event_id). The goal is: "Schedule a one-hour meeting with Alice next Tuesday at 3 PM, but first check if I'm free."
The agent must:
- Call
read_eventsfor next Tuesday. - Parse the returned events to check for conflicts.
- If free, call
create_eventwith the correct parameters. - If not free, either suggest an alternative or ask the user.
A standard eval would never test this flow because it requires sequential tool calls and conditional logic. We can also create variants: what if the user says "schedule it" without specifying a time? The agent should ask for the missing parameter, not guess.
We write the scenario as a Python class that simulates the environment. Here's a simplified skeleton:
from typing import List, Dict, Optional
class CalendarScenario:
def __init__(self):
self.events = {
"2025-04-15": [
{"id": 1, "title": "Standup", "time": "10:00", "duration": 30}
]
}
self.tools = {
"read_events": self.read_events,
"create_event": self.create_event,
"delete_event": self.delete_event,
}
self.success = False
self.error = None
def read_events(self, date: str) -> List[Dict]:
return self.events.get(date, [])
def create_event(self, title: str, date: str, time: str, duration: int) -> str:
if date not in self.events:
self.events[date] = []
# Check conflict
for event in self.events[date]:
if event["time"] == time:
self.error = "Time slot already booked"
return "Error: conflict"
event_id = len([e for evs in self.events.values() for e in evs]) + 1
self.events[date].append({"id": event_id, "title": title, "time": time, "duration": duration})
self.success = True
return f"Event created with id {event_id}"
def delete_event(self, event_id: int) -> str:
for date, evs in self.events.items():
for ev in evs:
if ev["id"] == event_id:
evs.remove(ev)
return "Deleted"
return "Event not found"
def check_success(self) -> bool:
return self.successWe then run the agent against this scenario, recording every tool call and the final state. The evaluation is not just pass/fail; we also measure:
- Number of tool calls: Efficiency.
- Error rate: How often does the agent hit an error and recover?
- Hallucination rate: Does the agent invent tool outputs or assume outcomes?
- Clarification requests: Does it ask for missing info?
Running the Benchmark
We run each scenario multiple times with different random seeds or initial conditions to average out stochasticity. For each run, we log the full conversation and tool call trace. We then compute aggregate metrics.
A typical benchmark suite might have 20–50 scenarios, each targeting a different capability: planning, memory, error handling, tool selection, etc. We also include adversarial scenarios where the environment changes mid-conversation or the tools return misleading data.
We do not report a single score. Instead, we produce a radar chart or a capabilities matrix showing strengths and weaknesses. For example, a model might excel at planning but struggle with error recovery. This granularity helps us decide where to focus fine-tuning or prompt engineering.
Automating the Evaluation
To scale, we automate the entire pipeline. Each scenario is a self-contained Python script that implements the environment and a run_agent(agent, scenario) function. The agent is a wrapper around the LLM that exposes a call_tool(name, args) method. We use a simple loop:
import json
def evaluate(agent, scenario, max_turns=20):
state = scenario.initial_state()
history = []
for turn in range(max_turns):
action = agent.act(state, history)
if action["type"] == "finish":
break
result = scenario.execute(action)
history.append({"action": action, "result": result})
state = scenario.update_state(state, result)
return scenario.check_success(), historyThe agent's act method takes the current state and history, and returns a structured action (tool call or finish). This interface is model-agnostic, so we can swap models and compare.
Why Not Use Existing Agent Benchmarks?
There are emerging agent benchmarks like AgentBench, WebArena, or SWE-bench. They are valuable, but they have limitations:
- Environment complexity: Many are tied to specific environments (e.g., web browsing, code repositories). Our custom scenarios can target exactly the tools and data our agents will use.
- Control over difficulty: We can create scenarios that are deliberately easy or hard, or that test a specific failure mode. Standard benchmarks are fixed.
- Reproducibility: Some benchmarks rely on live websites or APIs that change over time. Our scenarios are fully deterministic and repeatable.
- Cost: Running agents on live environments can be expensive or rate-limited. Our simulated scenarios run locally and fast.
That said, we still run standard benchmarks for model selection. But for agent behavior, custom scenarios give us the signal we need.
Practical Tips for Building Your Own
- Start small: Pick one tool (e.g., a calculator, a search API) and design 5 scenarios around it. Iterate.
- Inject failures: Make tools occasionally return errors or timeouts. See how the agent reacts.
- Test ambiguity: Give underspecified goals. A good agent asks questions; a bad one guesses.
- Log everything: Record every action, tool output, and state change. This is invaluable for debugging.
- Automate scoring: Write deterministic success checks. Avoid LLM-as-judge if possible; it adds noise.
- Version your scenarios: As you improve your agent, you may need harder scenarios. Keep old ones for regression testing.
The Tradeoffs
Custom scenarios require upfront effort to design and implement. They are not as standardized as MMLU, so comparing results across teams is harder. But for internal development, they are far more informative. We trade broad comparability for deep insight.
Another tradeoff: scenarios can become stale. Once the agent learns to solve a scenario perfectly, it no longer provides signal. We continuously add new scenarios or introduce randomness (e.g., different dates, different tool responses) to keep the benchmark challenging.
Closing Thoughts
Standard LLM evals are necessary but not sufficient for agents. They measure the model's raw capability, not its behavior in a tool-augmented, sequential environment. Custom scenario-based benchmarks fill that gap. They let us test exactly what we care about: can the agent accomplish a goal using the tools we gave it, even when things go wrong?
This approach has shaped our development cycle. We don't just train a model and hope; we iterate on scenarios, find failure modes, and improve the agent's prompting or fine-tuning. The result is an agent that we trust to operate autonomously in production.
If you're building agents, I encourage you to invest in custom scenarios. Start with a single tool and a handful of goals. You'll learn more about your agent's behavior than any leaderboard can tell you.