Empirical validation that Inspect AI + inspect_swe + Grove are a viable foundation for the Agent Testing & Observability hackathon. This repo contains the working code behind every "validated" row in the skunkworks team brief.
Headline result: running the same mongodb-mcp-setup eval prompt through Claude Code with the skill ON vs. OFF moved the score from 0.000 → 1.000. That's the lift signal we want as the demo's headline finding, validated end-to-end through Grove.
- Docker Desktop running (for any task using
claude_code()) uvinstalled.envpopulated from.env.examplewith Grove credentials (see the internal Grove Wiki)- Local clone of
mongodb/agent-skills. Tasks resolve eval and skill paths via theAGENT_SKILLS_DIRenv var; if unset, a sibling../agent-skillsdirectory is used. Set it in.env, e.g.AGENT_SKILLS_DIR=/path/to/agent-skills.
.
├── pyproject.toml inspect-ai 0.3.220, inspect-swe (MongoCaleb fork),
│ anthropic 0.100.0, fastapi 0.136.1, uvicorn 0.46.0
├── .env.example Grove credential template (.env is gitignored)
├── proxy.py Local translating proxy - holds Grove key server-side,
│ strips client auth headers, injects api-key:.
│ Inspect points at this so the Grove key never enters
│ the Inspect process or eval logs.
├── verify_grove.py Direct Anthropic SDK <-> Grove sanity check (no proxy,
│ no Inspect). Useful for confirming Grove credentials work.
├── tasks/
│ ├── mcp_setup.py First real eval, no seeding -> I (filesystem mismatch)
│ ├── mcp_setup_seeded.py Empty .zshrc seed -> I (still missed env)
│ ├── mcp_setup_realistic.py + mcp_servers via claude_code() -> I (introspection mismatch)
│ ├── mcp_setup_settings.py + ~/.claude/settings.local.json -> I (correct diagnosis, deferred action)
│ └── mcp_setup_with_skill.py + mongodb-mcp-setup skill loaded -> C (1.000) <- the headline
├── tests/
│ └── hello.py hello_basic + hello_anthropic + hello_openai
├── web/ Vite + React dashboard that renders the
│ compare-*.json summaries from logs/summary/.
└── logs/ Generated by `inspect eval`. GITIGNORED.
The mcp_setup_*.py files form a realism ladder: each iteration up surfaced a different gap (file paths, runtime introspection, action authorization). The progression itself is a useful debugging story.
Grove issues Azure-backed gateway keys. For Anthropic, Grove preserves the native protocol (/anthropic/v1/messages, model in body, anthropic-version header) but uses Azure-style auth (api-key: header instead of Anthropic's native x-api-key:). The Anthropic SDK auto-sends x-api-key: from its api_key= param and we can't disable that.
Our proxy.py solves this:
- Holds
GROVE_API_KEYandGROVE_BASE_URLin its own env. - Listens on
127.0.0.1:7676and accepts Anthropic-format requests. - Strips client-supplied auth headers (
x-api-key:,Authorization:) and injectsapi-key: <grove-key>server-side. - Forwards to Grove and streams the response back.
Inspect (and inspect_swe's sandbox bridge) point at the proxy via --model-base-url. Result: nothing in model_args except the model name, no Grove key in eval logs, logs are share-clean.
The proxy also satisfies the inspect_swe sandbox bridge architecture: Claude Code in the sandbox → bridge proxy on localhost:<bridge_port> (inside the sandbox) → Inspect's anthropic SDK on the host → our proxy on 127.0.0.1:7676 → Grove. Two-stage proxying, both stages work.
set -a && source .env && set +a
uv run uvicorn proxy:app --host 127.0.0.1 --port 7676
# Sanity check from another shell:
curl http://127.0.0.1:7676/__healthNote: set -a && source .env && set +a reads the .env file and exports the variables to the environment. .env must include the ANTHROPIC_API_KEY because the Anthropic SDK errors on construction without one; the proxy strips it before anything reaches Grove. Inspect auto-loads .env via python-dotenv, so no inline prefix is needed.
Be sure you are in your repo's root directory, and have activated your virtual environment. Then, source your .env file:
# In your eval-running shell:
set -a && source .env && set +auv run inspect eval tasks/<my_python_file>.py@<my @task method name> \
--model "anthropic/$ANTHROPIC_MODEL" \
--model-base-url "http://127.0.0.1:7676/anthropic"uv run inspect eval tasks/<my_python_file>.py \
--model "anthropic/$ANTHROPIC_MODEL" \
--model-base-url "http://127.0.0.1:7676/anthropic"SDK sanity (no proxy, no Inspect) - confirms Grove credentials work.
uv run python verify_grove.pyBare Inspect connectivity (no sandbox, no inspect_swe).
uv run inspect eval tests/hello.py@hello_basic \
--model "anthropic/$ANTHROPIC_MODEL" \
--model-base-url "http://127.0.0.1:7676/anthropic"Full pipeline - Docker sandbox + inspect_swe bridge (Anthropic).
uv run inspect eval tests/hello.py@hello_anthropic \
--model "anthropic/$ANTHROPIC_MODEL" \
--model-base-url "http://127.0.0.1:7676/anthropic"Same pipeline through Grove's OpenAI Chat Completions endpoint via the OpenCode
sandbox harness. $OPENAI_MODEL is already in provider/id form (e.g.
openai/gpt-4o); do NOT use openai/$ANTHROPIC_MODEL here — Grove rejects
Anthropic model ids on its OpenAI endpoint with api_not_supported.
uv run inspect eval tests/hello.py@hello_openai \
--model "$OPENAI_MODEL" \
--model-base-url "http://127.0.0.1:7676/openai"In a third terminal window, run the following command to start the Inspect viewer:
uv run inspect viewThen open http://127.0.0.1:7575/ to explore the logs.
compare.py runs N (task, model) configurations in one shot and prints a summary table at the end. Each run still produces a normal .eval log in logs/ (viewable via inspect view), and a compare-<timestamp>.json summary is written alongside.
A "run" is the triple (task_ref, model, base_url?). Because each @task in tasks/ already encodes the harness (generate / claude_code / opencode) and any solver options (skills, seeded files, etc.), comparing across harness, model, or options is just a matter of listing more triples. base_url is derived from the model's provider/ prefix against --proxy:
anthropic/...→<proxy>/anthropicopenai/...→<proxy>/openai- anything else → bare
<proxy>(matcheshello_basic)
Set up the proxy and .env as in Step 1 first.
# List all discovered @task refs:
uv run python compare.py --list-tasks
# Dry run (resolve and print, no eval):
uv run python compare.py --dry-run \
--run "tests/hello.py@hello_basic:anthropic/$ANTHROPIC_MODEL" \
--run "tests/hello.py@hello_openai:$OPENAI_MODEL"
# Headline lift: same eval, skill OFF vs ON:
uv run python compare.py \
--run "tasks/mcp_setup_settings.py@mcp_setup_first_settings:anthropic/$ANTHROPIC_MODEL" \
--run "tasks/mcp_setup_with_skill.py@mcp_setup_first_with_skill:anthropic/$ANTHROPIC_MODEL"
# Cross-harness on the same prompt:
uv run python compare.py \
--run "tests/hello.py@hello_anthropic:anthropic/$ANTHROPIC_MODEL" \
--run "tests/hello.py@hello_openai:$OPENAI_MODEL"
# From a config file:
uv run python compare.py --config compare.json
# Run up to 2 cells in parallel (one inspect_ai process per child):
uv run python compare.py --parallel 2 \
--run "tests/hello.py@hello_basic:anthropic/$ANTHROPIC_MODEL" \
--run "tests/hello.py@hello_basic:$OPENAI_MODEL"
# Every @task in every tasks/*.py file against N models in one shot:
uv run python compare.py --parallel 2 --all-files \
--models "anthropic/$ANTHROPIC_MODEL,$OPENAI_MODEL"--all-files enumerates tasks/*.py (skipping _*.py private modules) and pairs each file with every model in --models (comma-separated). The same file-level fan-out as --run tasks/X.py:MODEL applies: each @task becomes its own row in the summary and its own .eval log, and one bad @task in a file doesn't abort the rest. --all-files and --models must be used together, cannot be combined with --config, and combine additively with --run. Pair with --dry-run first if you want to see the full matrix without launching evals.
Output:
- A flat table with one row per
(run, scorer, metric)triple. A run with multiple scorers or multiple metrics per scorer (e.g.accuracy+stderr) gets one row each. - A pivot table (rows = tasks, cols = models, cells = primary metric value) is printed when there are 2+ unique tasks AND 2+ unique models AND all results share the same primary metric.
- The JSON summary captures the full
metrics: [{scorer, name, value}, ...]per run, plussamples,status,error, and the path to the.evallog.
--parallel N uses a process pool (each child has its own inspect_ai state; eval_async from one process can't run concurrently with another). Output from concurrent runs will interleave on the console; the summary table at the end is always coherent.
Example compare.json:
{
"proxy": "http://127.0.0.1:7676",
"runs": [
{"task": "tests/hello.py@hello_anthropic", "model": "anthropic/claude-sonnet-4-5"},
{"task": "tests/hello.py@hello_anthropic", "model": "anthropic/claude-opus-4-7"},
{"task": "tests/hello.py@hello_openai", "model": "openai/gpt-4o"}
]
}With no --run and no --config, compare.py falls back to interactive prompts for task selection and model list.
web/ is a small Vite + React app that renders the compare-*.json summaries written to logs/summary/ by compare.py. It picks up new summaries automatically — no rebuild needed; just refresh the page after a compare.py run.
Prereqs: Node 18+ and npm. The Vite dev server reads the summaries directly from the repo's logs/summary/ via a small middleware (/api/summaries), so there's no separate backend to start.
cd web
npm install # first time only
npm run dev # serves on http://127.0.0.1:5173/Then open http://127.0.0.1:5173/ and pick a summary from the dropdown. The page shows the same pivot + flat table as compare.py's console output, with the .eval log path for each row so you can cross-reference with inspect view.
npm run build produces static assets in web/dist/, but the /api/summaries middleware is dev-server-only — a built bundle needs an equivalent endpoint wired up by whatever serves it. For day-to-day local use, stick with npm run dev.
uv run inspect log dump logs/<file>.eval | jq '.eval.model_args, .eval.model_base_url'Should print {} and a localhost URL ending in /anthropic or /openai (e.g. "http://127.0.0.1:7676/anthropic"). Empty model_args and a localhost URL — no Grove key, no Grove URL.
If you ever revert to the dual-header approach (passing -M default_headers={...}), the key WILL appear in model_args in plaintext — so don't.
| Step | Question answered |
|---|---|
verify_grove.py |
Does Grove accept the Anthropic SDK request when both api-key: (real) and x-api-key: (dummy) headers are present? |
proxy.py __health |
Is the proxy alive and pointing at the right Grove upstream? |
hello_basic via proxy |
Does Inspect's anthropic/ provider talk to the proxy correctly? |
hello_anthropic via proxy |
Does inspect_swe's sandbox bridge route Claude Code's API calls through Inspect's proxy-configured client? |
hello_openai via proxy |
Does the proxy's /openai/ route (with v1/ injection) route OpenCode's OpenAI-protocol calls through Inspect's openai/ provider to Grove's Chat Completions endpoint? |
mcp_setup_first |
What does a real multi-turn agentic trajectory on a real agent-skills eval prompt look like, scored by model_graded_qa? |
mcp_setup_first_seeded |
Does adding empty shell profiles change the agent's behavior? (No.) |
mcp_setup_first_realistic |
Does passing mcp_servers=[...] via claude_code() close the gap? (No - the agent reads files, not runtime state.) |
mcp_setup_first_settings |
Does seeding ~/.claude/settings.local.json (which inspect_swe doesn't overwrite) close the diagnostic gap? (Yes - agent now correctly identifies the missing env var. But scores I because it stops at description and doesn't act.) |
mcp_setup_first_with_skill |
Does loading the mongodb-mcp-setup skill make the agent execute its prescribed steps? Yes — score moves to C (1.000). |
Tasks are created in the tasks/ directory. Each file in tasks/ can contain multiple tasks. Each @task function tests one of the evals defined in the evals.json file you specify at the top of the file.
For example, to test the evals in a schema-design/evals.json file, you might create a file named schema_design.py that contains an @task function for each eval in the evals.json file.
An example function looks like this:
@task
def product_view_counter() -> Task:
return Task(
dataset=[load_sample_by_name(EVALS_PATH, "product-view-counter")],
solver=claude_code(),
scorer=model_graded_qa(),
sandbox="docker",
)Note that the name "product-view-counter" matches the name field in the evals.json file.
If the evals.json file does not have a name field, you can use the id field of the eval.
For example, if you want to use the eval with "id": 3,, you would use dataset=[load_sample_by_id(EVALS_PATH, 3)].
Alternatively, you can use the index of the eval in the evals.json file. For example, to use the first eval in the file,
you would use dataset=[load_sample_by_index(EVALS_PATH, 0)].
Use any of the existing task files as a starting point for your own tasks.
- No Gemini access through Grove as of May 8, 2026. Cross-harness Gemini cuts unless an alternate path appears.
- Skunkworks team brief:
team-brief.md(sibling doc; not in this repo). - Inspect AI ·
inspect_swe(fork with OpenCode + Codex fixes; upstream: meridianlabs-ai/inspect_swe) ·inspect_flow mongodb/agent-skills— public skills + eval cases