This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
inspect-test-utils is a utility library for the Inspect AI framework providing:
- Deterministic tasks, scorers, and a hardcoded model for integration tests
- Test utilities (eval runner, assertions, fixtures) for writing tests against Inspect AI
- Scanners for inspect-scout
Primary use cases:
- Smoke tests for inspect-action and other Inspect AI deployments
- Ad-hoc tests for reproducing bugs that are hard to find otherwise
- Manual exploratory testing of eval infrastructure
- Testing failure paths and edge cases without real model API calls
# Install dependencies
uv sync
# Linting and formatting
uv run ruff check # Lint
uv run ruff format # Format
# Testing
uv run pytest # All tests
uv run pytest tests/test_hardcoded.py::TestParseToolCalls # Single test class
# Run an evaluation
inspect eval inspect_test_utils/say_hello \
--task-arg sample_count=3 \
--model hardcoded --model-arg answer=helloThe library has 9 modules in inspect_test_utils/:
| Module | Purpose |
|---|---|
tasks.py |
Task definitions (see Tasks section below) |
scorers.py |
Scorer implementations (failing_scorer, closeness_log, hardcoded_scorer) |
hardcoded.py |
HardcodedModelAPI - deterministic model that emits pre-defined tool calls |
solvers.py |
Hardcoded solvers (hardcoded_bash_solver, hardcoded_python_solver, inspection_solver, combined_solver) |
eval_runner.py |
EvalTestResult dataclass and run_eval_test() helper for running evals in tests |
assertions.py |
Assertion helpers (assert_eval_score, assert_score_in_range, assert_files_exist, assert_contains) |
fixtures.py |
Pytest fixtures and markers (skip_sandbox, requires_docker, custom markers) |
scanners.py |
Scanner implementations (suspicious_behaviour, word_counter, citing_scanner) using inspect-scout |
_registry.py |
Plugin registration for Inspect AI discovery |
Plugin entry point: Registered via [project.entry-points.inspect_ai] so tasks/models can be referenced directly in Inspect CLI.
| Task | Purpose | Key Parameters |
|---|---|---|
say_hello |
Simple task requiring answer "hello" | sample_count, local |
say_hello_with_tools |
Like say_hello with extra tools (text_editor, bash_session, think) |
sample_count |
guess_number |
Numeric guessing with closeness_log scorer |
sample_count, target, local |
guess_number_keep_guessing |
Uses react agent with try_guess tool |
sample_count, target, delay, local |
timeout |
Task with configurable bash timeout | sample_count, timeout (seconds) |
hardcoded_score |
Returns pre-defined scores (supports NaN for manual scoring) | hardcoded_score, hardcoded_score_by_sample_id_and_epoch |
sometimes_fails_setup |
Randomly fails during setup phase | sample_count, fail_setup_on_epochs, failure_rate |
sometimes_fails_scoring |
Randomly fails during scoring phase | sample_count, fail_score_on_epochs, failure_rate |
configurable_sandbox |
K8s sandbox with resource configuration; optional crash injector (crash_after) for agent-agnostic deployment resume tests |
cpu, memory, storage, gpu, gpu_model, allow_internet, crash_after, crash_hard |
network_sandbox |
Docker network mode testing, uniform or per-service | network_mode ("none", "bridge", "bridge_network_pattern"), services, service_network_modes |
The hardcoded model (hardcoded.py) emits a sequence of pre-defined tool calls, then submits a final answer.
Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
tool_calls |
list[HardcodedToolCall] | str | list[str] |
[] |
Tool calls to execute (JSON string, list of bash commands, or list of dicts) |
repetitions |
int |
1 |
Number of times to cycle through tool_calls before answering |
answer |
str |
"done" |
Final answer to submit |
delay |
float |
0.0 |
Delay between tool calls (seconds) |
concurrency |
int |
DEFAULT_MAX_CONNECTIONS |
Max parallel tool calls |
failure_rate |
float |
0.0 |
Random failure probability (0.0-1.0) for testing error handling |
rate_limit_capacity |
int | null |
null |
Raise a simulated HTTP 429 when concurrent in-flight generate calls exceed N (0 = every call). Only bites while calls overlap, i.e. delay > 0 |
rate_limit_status |
int |
429 |
Status code on the simulated error; anything but 429 is retried as transient (no adaptive scale-down) |
rate_limit_retry_after |
float | null |
null |
Retry-After handed to the adaptive controller. Carried but unconsumed since inspect_ai#5138; before that fix it extended the cooldown. Never shortens the retry backoff |
rate_limit_swallowed_retries |
int |
0 |
Report N retries to the adaptive controller and then succeed anyway, instead of raising (what a real provider's SDK-internal retry looks like). Kind follows rate_limit_status. 0 keeps the raise |
retry_wait_seconds |
float | null |
null |
Fixed retry backoff instead of inspect-ai's 3s-1800s exponential jitter |
CLI Example:
inspect eval inspect_test_utils/say_hello \
--model hardcoded \
--model-arg tool_calls='[{"tool_name": "bash", "tool_args": {"cmd": "echo hi"}}]' \
--model-arg repetitions=2 \
--model-arg answer="hello"The HardcodedToolCall TypedDict defines the structure for tool calls:
from typing import Any, TypedDict
class HardcodedToolCall(TypedDict):
tool_name: str
tool_args: dict[str, Any]Helper pattern (commonly used in test code):
def bash_tool_call(cmd: str) -> HardcodedToolCall:
return {"tool_name": "bash", "tool_args": {"cmd": cmd}}
def python_tool_call(code: str) -> HardcodedToolCall:
return {"tool_name": "python", "tool_args": {"code": code}}
# Usage
tool_calls = [
bash_tool_call("echo hello"),
bash_tool_call("cat /etc/passwd"),
python_tool_call("print(2 + 2)"),
]Shorthand: Pass a list of strings for bash commands:
# These are equivalent:
model_args={"tool_calls": ["echo hi", "ls -la"]}
model_args={"tool_calls": [{"tool_name": "bash", "tool_args": {"cmd": "echo hi"}}, ...]}For testing error handling and robustness:
# Task that always fails during setup
inspect eval inspect_test_utils/sometimes_fails_setup \
--task-arg failure_rate=1.0
# Task that fails 20% of the time during scoring
inspect eval inspect_test_utils/sometimes_fails_scoring \
--task-arg failure_rate=0.2
# Model that randomly fails
inspect eval inspect_test_utils/say_hello \
--model hardcoded \
--model-arg failure_rate=0.1 \
--model-arg answer=helloParameters:
failure_rate: Probability of failure (0.0-1.0)fail_on_epochs/fail_setup_on_epochs/fail_score_on_epochs: List of specific epochs to fail on
rate_limit_capacity makes the model raise a simulated HTTP 429 once concurrent
in-flight generate calls exceed N, which drives inspect-ai's adaptive
concurrency controller down toward that capacity:
inspect eval inspect_test_utils/say_hello \
-T sample_count=500 -T local=true \
--model hardcoded/rl -M answer=hello -M delay=0.5 \
-M rate_limit_capacity=2 \
-M retry_wait_seconds=0.5 \
--max-retries 500 \
--adaptive-connections 1-20-20That run takes ~2.5 minutes and walks the limit
20 -> 15 -> 10 -> 8 -> 6 -> 4 -> 3 -> 2, then oscillates around the simulated
capacity. Gotchas, all of which will otherwise leave you looking at a single
cut and wondering why:
- The controller floors at its
min(default 10) and cuts at most once percooldown_seconds(default 15s, not settable from the CLI), so the walk-down needs--adaptive-connections 1-20-20and has to outlast several windows. - On inspect-ai before UKGovernmentBEIS/inspect_ai#5138,
rate_limit_retry_afterextends that cooldown on every retry, so a hint above the gap between retries freezes the walk-down after a single cut. It is omitted above for that reason. #5138 makes the hint unconsumed and the walk-down then behaves the same for any value, so keep hints small if you want a recipe that behaves identically either side of that fix. - Without
retry_wait_secondsevery simulated 429 costs 3-48s of tenacity backoff.max_retriesalso defaults to unlimited, so bound it. - The capacity is deliberately independent of
concurrency; inspect-ai ignores the provider'smax_connections()while adaptive connections are on, so capping the client at N and then asserting it discovered N would be circular. rate_limit_status=503produces the same failures classified as transient: retried, but never scaling the controller down.- Setting
max_connectionsinGenerateConfig(orbatch) silently disables adaptive concurrency entirely.
rate_limit_swallowed_retries swaps the raise for what a real client does most
of the time: openai/anthropic SDKs retry internally (max_retries=2), so the
429 never escapes generate() and yet the controller still hears about it.
-M rate_limit_capacity=12 -M rate_limit_swallowed_retries=2 -M delay=0.05For a 429 the controller trajectory is identical to the raise path (inspect-ai's
own retry loop holds the same connection slot, so it cannot tell them apart);
what changes is that the 429s now come from requests that succeeded - no
errored model calls, no tenacity backoff, and ModelEvent.retries populated as
it is for a real provider. retry_wait_seconds is inert here; delay holds the
slot instead. With rate_limit_capacity=0 every call is rate-limited and the
eval still finishes with every sample successful.
Pairing it with rate_limit_status=503 reaches a state the raise path cannot:
a transient retry never cuts the limit, but it does set _request_had_retry,
which stops the eventual success counting toward scale-up. Growth stalls from
retry noise alone, with no cut and nothing failing - measured, the limit runs
4 -> 8 -> 16 -> 32 -> 64 unthrottled and stops at 16 or 32 with
rate_limit_capacity=12 -M rate_limit_status=503 -M rate_limit_swallowed_retries=2.
That gate is the main reason to reach for this knob when testing scale-up.
Use hardcoded_score task with hardcoded_scorer for testing score handling:
# Return a specific score
inspect eval inspect_test_utils/hardcoded_score \
--task-arg 'hardcoded_score={"value": 0.75, "explanation": "Partial credit"}'
# Return NaN for manual scoring scenarios
inspect eval inspect_test_utils/hardcoded_score \
--task-arg 'hardcoded_score={"value": "NaN", "metadata": {"manual-scoring": true}}'Use network_sandbox task to test Docker network configurations:
# No network access (default)
inspect eval inspect_test_utils/network_sandbox \
--task-arg network_mode=none
# Bridge network mode
inspect eval inspect_test_utils/network_sandbox \
--task-arg network_mode=bridge
# Shared bridge network between services
inspect eval inspect_test_utils/network_sandbox \
--task-arg network_mode=bridge_network_pattern \
--task-arg 'services=["default", "server"]'
# Mixed: a connected agent container next to an isolated one
inspect eval inspect_test_utils/network_sandbox \
--task-arg 'services=["default", "solution"]' \
--task-arg 'service_network_modes={"default": "bridge", "solution": "none"}'service_network_modes overrides network_mode for the services it names;
network_mode covers the rest (and defaults to none). Passing a
service_network_modes that covers every service alongside network_mode, or
naming a service that is not in services, raises ValueError. A service set to
none is never put on the shared network — network_mode: none plus networks
is rejected by Hawk and by the inspect_k8s_sandbox converter.
For writing tests against Inspect AI evaluations:
from inspect_test_utils import run_eval_test, assert_eval_score
from inspect_test_utils.solvers import hardcoded_bash_solver
# Run eval with a hardcoded solver (bypasses model)
result = run_eval_test(
my_task,
solver=hardcoded_bash_solver(["echo hello"]),
)
assert_eval_score(result, expected=1.0, tolerance=0.01)
# Or use HardcodedModelAPI for full agent loop testing
result = run_eval_test(
my_task,
model="hardcoded/test",
model_args={"tool_calls": ["echo hello"], "answer": "done"},
)For pytest fixtures, add to conftest.py:
pytest_plugins = ["inspect_test_utils.fixtures"]