This file provides guidance to AI coding agents working in this repository.
- Install dev environment:
pip install -e ".[dev]", or with uv:make sync - Format, lint, and type check:
make check(runsruff format,ruff check --fix,mypy) - Run fast tests:
make test(orpytest) - Run a single test:
pytest tests/test_file.py::test_name -v
- mypy runs with
strict = trueoversrcandtests. All functions need type annotations, and an unusedtype: ignorecomment is a hard error (warn_unused_ignores). - CI installs the latest unpinned mypy, so a new mypy release can turn
mainred with no code change (e.g. atype: ignorethat a newer mypy considers unused). If the mypy check fails on a line your diff didn't touch, check whethermainhas the same failure before debugging your change — and if so, fix it in your PR and note that it was pre-existing.
- Three opt-in gates skip tests by default:
--runslow(integration tests that drive real agents end-to-end),--runapi(tests needing model API access), and--runflaky. PR CI runs only the fast suite; the slow suite runs nightly viameridianlabs-ai/actions("Inspect SWE Nightly Tests"). - Slow tests call real model APIs (skipped without the relevant
OPENAI_API_KEY/ANTHROPIC_API_KEY/GOOGLE_API_KEY) and need a sandbox: docker and/or k8s are auto-detected from your environment and sandbox-dependent tests are parametrized over whatever is available.pytest --coshows what would actually run. - The end-to-end example tests all route through the
run_examplehelper intests/conftest.py, which bounds each eval withtime_limitandtoken_limit(see the comments there for the rationale before changing them). The nightly runs pytest with--timeout=900, so any per-eval time limit must stay comfortably under that or a timeout surfaces as an opaque pytest kill instead of a clean eval failure.
- Title PRs as Conventional Commits (
<type>: <description>)—we squash-merge, so the PR title becomes the commit message that drives releases;pr-title-lintenforces it feat:/fix:are for user-facing changes only: they headline the release notes and bump the version.perf:/revert:also appear in the notes (no bump);docs:,refactor:,chore:,build:,ci:,test:,style:are hidden- In the description part of the title, state the user-facing outcome — the problem a user hit or the capability they gain — not the mechanism of the fix:
fix: agent hang when sandbox startup races container pull, notfix: add lock around container init. For changes with no user-facing outcome (refactoring, CI, docs), describe the change itself. - Body lines starting with
<type>:are parsed as extra changelog entries—don't begin description lines with a conventional-commit prefix unless that's intended - Never edit
CHANGELOG.md, version numbers, or.release-please-manifest.json—Release Please owns them - Before opening a non-trivial PR, run at least one code review pass in a fresh context (e.g.
/code-reviewor a subagent that hasn't seen the authoring conversation) using a frontier-class model — the authoring context is blind to its own assumptions, and in our experience reviews from small fast-tier models rarely surface real issues. Fix or explicitly dismiss each finding before opening the PR, and disclose the pass in the description (see "Agent Review Disclosure" below). - After opening a PR, don't stop at creation: watch its checks (
gh pr checks <number> --watch) until they complete, report the outcome, and investigate and fix any failures. If the branch falls behindmain, update it so CI runs against current code. - See CONTRIBUTING.md for full guidelines
- Credential boundary: treat the evaluated sandbox as untrusted. By default, bridge-owned model-provider credentials never enter it—only the dummy key used to authenticate with the in-sandbox model bridge reaches inference. A
claude_code(env=...)caller can explicitly override its environment, includingANTHROPIC_AUTH_TOKEN, so caller-provided credentials must be treated as sandbox-exposed. Other real credentials (e.g. MCP auth headers) may enter the sandbox for transport, but must stay outside every model- and tool-readable path. - Public API evolution: add new parameters at the end of the signature, using
Literaltypes with runtime validation for constrained options. Renaming or removing a released parameter needs a**deprecated_args: Unpack[...]shim; new public surface needs a concrete use case first.
When an AI agent authored or co-authored a change, include an ### Agent review section in the PR description summarizing the pre-PR review passes described above: what model/tool reviewed, whether the review ran in a fresh context and/or on a different model from the author, how many passes, and the findings — issues found, which were fixed, and which were dismissed with a one-line reason each. Maintainers weight the disclosed reviewer model and pass count when deciding how much independent review a PR still needs. If no review pass was run, say so explicitly — never report a review that didn't happen; a content-free claim ("reviewed, looks good") is worse than disclosing none. Example:
### Agent review
- Reviewer: Claude Fable 5 via /code-review (fresh context), 2 passes
- Findings: 3 — 2 fixed, 1 dismissed (flagged a missing None check that is
guarded upstream)