Skip to content

About

Spike repo exploring the Inspect agent eval and observability harness as a possible tool for running agent platform evals.

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

Inspect spike (pre-hackathon)

Empirical validation that Inspect AI + inspect_swe + Grove are a viable foundation for the Agent Testing & Observability hackathon. This repo contains the working code behind every "validated" row in the skunkworks team brief.

Headline result: running the same mongodb-mcp-setup eval prompt through Claude Code with the skill ON vs. OFF moved the score from 0.000 → 1.000. That's the lift signal we want as the demo's headline finding, validated end-to-end through Grove.

Prereqs

  • Docker Desktop running (for any task using claude_code())
  • uv installed
  • .env populated from .env.example with Grove credentials (see the internal Grove Wiki)
  • Local clone of mongodb/agent-skills. Tasks resolve eval and skill paths via the AGENT_SKILLS_DIR env var; if unset, a sibling ../agent-skills directory is used. Set it in .env, e.g. AGENT_SKILLS_DIR=/path/to/agent-skills.

Layout

.
├── pyproject.toml            inspect-ai 0.3.220, inspect-swe (MongoCaleb fork),
│                              anthropic 0.100.0, fastapi 0.136.1, uvicorn 0.46.0
├── .env.example              Grove credential template (.env is gitignored)
├── proxy.py                  Local translating proxy - holds Grove key server-side,
│                              strips client auth headers, injects api-key:.
│                              Inspect points at this so the Grove key never enters
│                              the Inspect process or eval logs.
├── verify_grove.py           Direct Anthropic SDK <-> Grove sanity check (no proxy,
│                              no Inspect). Useful for confirming Grove credentials work.
├── tasks/
│   ├── mcp_setup.py                   First real eval, no seeding         -> I (filesystem mismatch)
│   ├── mcp_setup_seeded.py            Empty .zshrc seed                   -> I (still missed env)
│   ├── mcp_setup_realistic.py         + mcp_servers via claude_code()     -> I (introspection mismatch)
│   ├── mcp_setup_settings.py          + ~/.claude/settings.local.json     -> I (correct diagnosis, deferred action)
│   └── mcp_setup_with_skill.py        + mongodb-mcp-setup skill loaded    -> C (1.000) <- the headline
├── tests/
│   └── hello.py                       hello_basic + hello_anthropic + hello_openai
├── web/                      Vite + React dashboard that renders the
│                              compare-*.json summaries from logs/summary/.
└── logs/                     Generated by `inspect eval`. GITIGNORED.

The mcp_setup_*.py files form a realism ladder: each iteration up surfaced a different gap (file paths, runtime introspection, action authorization). The progression itself is a useful debugging story.

How auth works

Grove issues Azure-backed gateway keys. For Anthropic, Grove preserves the native protocol (/anthropic/v1/messages, model in body, anthropic-version header) but uses Azure-style auth (api-key: header instead of Anthropic's native x-api-key:). The Anthropic SDK auto-sends x-api-key: from its api_key= param and we can't disable that.

Our proxy.py solves this:

  • Holds GROVE_API_KEY and GROVE_BASE_URL in its own env.
  • Listens on 127.0.0.1:7676 and accepts Anthropic-format requests.
  • Strips client-supplied auth headers (x-api-key:, Authorization:) and injects api-key: <grove-key> server-side.
  • Forwards to Grove and streams the response back.

Inspect (and inspect_swe's sandbox bridge) point at the proxy via --model-base-url. Result: nothing in model_args except the model name, no Grove key in eval logs, logs are share-clean.

The proxy also satisfies the inspect_swe sandbox bridge architecture: Claude Code in the sandbox → bridge proxy on localhost:<bridge_port> (inside the sandbox) → Inspect's anthropic SDK on the host → our proxy on 127.0.0.1:7676 → Grove. Two-stage proxying, both stages work.

Running it

Step 1: Start the proxy (separate terminal, leaves it running)

set -a && source .env && set +a
uv run uvicorn proxy:app --host 127.0.0.1 --port 7676
# Sanity check from another shell:
curl http://127.0.0.1:7676/__health

Note: set -a && source .env && set +a reads the .env file and exports the variables to the environment. .env must include the ANTHROPIC_API_KEY because the Anthropic SDK errors on construction without one; the proxy strips it before anything reaches Grove. Inspect auto-loads .env via python-dotenv, so no inline prefix is needed.

Step 2: Run evals

Be sure you are in your repo's root directory, and have activated your virtual environment. Then, source your .env file:

# In your eval-running shell:
set -a && source .env && set +a

To run a single task:

uv run inspect eval tasks/<my_python_file>.py@<my @task method name> \
  --model "anthropic/$ANTHROPIC_MODEL" \
  --model-base-url "http://127.0.0.1:7676/anthropic"

To run all tasks in a file:

uv run inspect eval tasks/<my_python_file>.py \
  --model "anthropic/$ANTHROPIC_MODEL" \
  --model-base-url "http://127.0.0.1:7676/anthropic"

For reference, here are the commands you can run for the "hello" test evals:

SDK sanity (no proxy, no Inspect) - confirms Grove credentials work.

uv run python verify_grove.py

Bare Inspect connectivity (no sandbox, no inspect_swe).

uv run inspect eval tests/hello.py@hello_basic \
  --model "anthropic/$ANTHROPIC_MODEL" \
  --model-base-url "http://127.0.0.1:7676/anthropic"

Full pipeline - Docker sandbox + inspect_swe bridge (Anthropic).

uv run inspect eval tests/hello.py@hello_anthropic \
  --model "anthropic/$ANTHROPIC_MODEL" \
  --model-base-url "http://127.0.0.1:7676/anthropic"

To run OpenAI / OpenCode through Grove:

Same pipeline through Grove's OpenAI Chat Completions endpoint via the OpenCode sandbox harness. $OPENAI_MODEL is already in provider/id form (e.g. openai/gpt-4o); do NOT use openai/$ANTHROPIC_MODEL here — Grove rejects Anthropic model ids on its OpenAI endpoint with api_not_supported.

uv run inspect eval tests/hello.py@hello_openai \
  --model "$OPENAI_MODEL" \
  --model-base-url "http://127.0.0.1:7676/openai"

Step 3: Check the log reports

In a third terminal window, run the following command to start the Inspect viewer:

uv run inspect view

Then open http://127.0.0.1:7575/ to explore the logs.

Comparing runs

compare.py runs N (task, model) configurations in one shot and prints a summary table at the end. Each run still produces a normal .eval log in logs/ (viewable via inspect view), and a compare-<timestamp>.json summary is written alongside.

A "run" is the triple (task_ref, model, base_url?). Because each @task in tasks/ already encodes the harness (generate / claude_code / opencode) and any solver options (skills, seeded files, etc.), comparing across harness, model, or options is just a matter of listing more triples. base_url is derived from the model's provider/ prefix against --proxy:

  • anthropic/... → <proxy>/anthropic
  • openai/... → <proxy>/openai
  • anything else → bare <proxy> (matches hello_basic)

Set up the proxy and .env as in Step 1 first.

# List all discovered @task refs:
uv run python compare.py --list-tasks

# Dry run (resolve and print, no eval):
uv run python compare.py --dry-run \
  --run "tests/hello.py@hello_basic:anthropic/$ANTHROPIC_MODEL" \
  --run "tests/hello.py@hello_openai:$OPENAI_MODEL"

# Headline lift: same eval, skill OFF vs ON:
uv run python compare.py \
  --run "tasks/mcp_setup_settings.py@mcp_setup_first_settings:anthropic/$ANTHROPIC_MODEL" \
  --run "tasks/mcp_setup_with_skill.py@mcp_setup_first_with_skill:anthropic/$ANTHROPIC_MODEL"

# Cross-harness on the same prompt:
uv run python compare.py \
  --run "tests/hello.py@hello_anthropic:anthropic/$ANTHROPIC_MODEL" \
  --run "tests/hello.py@hello_openai:$OPENAI_MODEL"

# From a config file:
uv run python compare.py --config compare.json

# Run up to 2 cells in parallel (one inspect_ai process per child):
uv run python compare.py --parallel 2 \
  --run "tests/hello.py@hello_basic:anthropic/$ANTHROPIC_MODEL" \
  --run "tests/hello.py@hello_basic:$OPENAI_MODEL"

# Every @task in every tasks/*.py file against N models in one shot:
uv run python compare.py --parallel 2 --all-files \
  --models "anthropic/$ANTHROPIC_MODEL,$OPENAI_MODEL"

--all-files enumerates tasks/*.py (skipping _*.py private modules) and pairs each file with every model in --models (comma-separated). The same file-level fan-out as --run tasks/X.py:MODEL applies: each @task becomes its own row in the summary and its own .eval log, and one bad @task in a file doesn't abort the rest. --all-files and --models must be used together, cannot be combined with --config, and combine additively with --run. Pair with --dry-run first if you want to see the full matrix without launching evals.

Output:

  • A flat table with one row per (run, scorer, metric) triple. A run with multiple scorers or multiple metrics per scorer (e.g. accuracy + stderr) gets one row each.
  • A pivot table (rows = tasks, cols = models, cells = primary metric value) is printed when there are 2+ unique tasks AND 2+ unique models AND all results share the same primary metric.
  • The JSON summary captures the full metrics: [{scorer, name, value}, ...] per run, plus samples, status, error, and the path to the .eval log.

--parallel N uses a process pool (each child has its own inspect_ai state; eval_async from one process can't run concurrently with another). Output from concurrent runs will interleave on the console; the summary table at the end is always coherent.

Example compare.json:

{
  "proxy": "http://127.0.0.1:7676",
  "runs": [
    {"task": "tests/hello.py@hello_anthropic", "model": "anthropic/claude-sonnet-4-5"},
    {"task": "tests/hello.py@hello_anthropic", "model": "anthropic/claude-opus-4-7"},
    {"task": "tests/hello.py@hello_openai",    "model": "openai/gpt-4o"}
  ]
}

With no --run and no --config, compare.py falls back to interactive prompts for task selection and model list.

Web dashboard

web/ is a small Vite + React app that renders the compare-*.json summaries written to logs/summary/ by compare.py. It picks up new summaries automatically — no rebuild needed; just refresh the page after a compare.py run.

Prereqs: Node 18+ and npm. The Vite dev server reads the summaries directly from the repo's logs/summary/ via a small middleware (/api/summaries), so there's no separate backend to start.

cd web
npm install            # first time only
npm run dev            # serves on http://127.0.0.1:5173/

Then open http://127.0.0.1:5173/ and pick a summary from the dropdown. The page shows the same pivot + flat table as compare.py's console output, with the .eval log path for each row so you can cross-reference with inspect view.

npm run build produces static assets in web/dist/, but the /api/summaries middleware is dev-server-only — a built bundle needs an equivalent endpoint wired up by whatever serves it. For day-to-day local use, stick with npm run dev.

Confirming logs are clean

uv run inspect log dump logs/<file>.eval | jq '.eval.model_args, .eval.model_base_url'

Should print {} and a localhost URL ending in /anthropic or /openai (e.g. "http://127.0.0.1:7676/anthropic"). Empty model_args and a localhost URL — no Grove key, no Grove URL.

If you ever revert to the dual-header approach (passing -M default_headers={...}), the key WILL appear in model_args in plaintext — so don't.

What each step proves

Step Question answered
verify_grove.py Does Grove accept the Anthropic SDK request when both api-key: (real) and x-api-key: (dummy) headers are present?
proxy.py __health Is the proxy alive and pointing at the right Grove upstream?
hello_basic via proxy Does Inspect's anthropic/ provider talk to the proxy correctly?
hello_anthropic via proxy Does inspect_swe's sandbox bridge route Claude Code's API calls through Inspect's proxy-configured client?
hello_openai via proxy Does the proxy's /openai/ route (with v1/ injection) route OpenCode's OpenAI-protocol calls through Inspect's openai/ provider to Grove's Chat Completions endpoint?
mcp_setup_first What does a real multi-turn agentic trajectory on a real agent-skills eval prompt look like, scored by model_graded_qa?
mcp_setup_first_seeded Does adding empty shell profiles change the agent's behavior? (No.)
mcp_setup_first_realistic Does passing mcp_servers=[...] via claude_code() close the gap? (No - the agent reads files, not runtime state.)
mcp_setup_first_settings Does seeding ~/.claude/settings.local.json (which inspect_swe doesn't overwrite) close the diagnostic gap? (Yes - agent now correctly identifies the missing env var. But scores I because it stops at description and doesn't act.)
mcp_setup_first_with_skill Does loading the mongodb-mcp-setup skill make the agent execute its prescribed steps? Yes — score moves to C (1.000).

Creating Tasks

Tasks are created in the tasks/ directory. Each file in tasks/ can contain multiple tasks. Each @task function tests one of the evals defined in the evals.json file you specify at the top of the file.

For example, to test the evals in a schema-design/evals.json file, you might create a file named schema_design.py that contains an @task function for each eval in the evals.json file.

An example function looks like this:

@task
def product_view_counter() -> Task:
    return Task(
        dataset=[load_sample_by_name(EVALS_PATH, "product-view-counter")],
        solver=claude_code(),
        scorer=model_graded_qa(),
        sandbox="docker",
    )

Note that the name "product-view-counter" matches the name field in the evals.json file. If the evals.json file does not have a name field, you can use the id field of the eval.

For example, if you want to use the eval with "id": 3,, you would use dataset=[load_sample_by_id(EVALS_PATH, 3)].

Alternatively, you can use the index of the eval in the evals.json file. For example, to use the first eval in the file, you would use dataset=[load_sample_by_index(EVALS_PATH, 0)].

Use any of the existing task files as a starting point for your own tasks.

Known issues / open work

  • No Gemini access through Grove as of May 8, 2026. Cross-harness Gemini cuts unless an alternate path appears.

References

About

Spike repo exploring the Inspect agent eval and observability harness as a possible tool for running agent platform evals.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages