Tags: METR/inspect-eval-utils
Tags
perf(artifacts): stop report writes blocking the eval event loop (#23) * perf(artifacts): add async writers and cut report round-trips Report writes happen in Inspect AI scorers, which are always coroutines, and in production the destination is S3. The synchronous writers therefore block the event loop -- and so every other sample in the run -- for the whole round trip. On a live prd runner the event loop never reached epoll_wait at all, and 13.9% of main-thread stack samples were parked in fsspec's sync() waiting on S3 from inside a report write. Add write_report_async / write_artifacts_async / write_artifact_async, which run the existing sync writers on a worker thread. anyio propagates contextvars, so sample_active() still resolves there. Also cut the round-trips themselves. _write_files listed the directory and then made three calls per existing entry to delete it; it now takes one listing and one bulk removal. Symlinks still go one at a time, because fsspec resolves the link and so cannot delete a symlink to a directory by either route. Writes use pipe_file rather than constructing a file object per file. Steady-state cost for replacing a report, counting only top-level filesystem calls (nested delegation excluded, since s3fs batches a multi-key rm into one delete_objects): files before after 1 10 7 2 14 8 5 26 11 10 46 16 The delete path is now O(1) in the number of existing files rather than O(n). The sync writers keep working for synchronous callers such as scripts and tests, and carry a deprecated:: note pointing at the async form. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor: tighten docstrings and fold the two round-trip tests into one Move the symlink reasoning in _clear_dir from its docstring into a comment at the branch it explains, trim the deprecation notes and the plot module docstring, and merge the round-trip budget and O(1)-clear tests, which shared a fixture and largely the same regression. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Make scaffolded tasks resumable by default (#22) The new_task skeleton declared no CheckpointSampleConfig, so the recent harder-tasks rollout had to retrofit ~31 tasks by hand to add the one-line default. Bake it into the skeleton instead. Scaffolded tasks now declare CheckpointSampleConfig(sandbox_paths={"default": ["/home/agent"]}), which matches the template's own "default" sandbox (working_dir /home/agent) and is the right config for the large majority of tasks. A comment guides authors to extend it (extra paths/services) or replace it (in-memory server state needs a Store-based scorer, not path capture). Declaring sandbox_paths is cost-free unless an eval enables checkpointing, so this adds no runtime overhead to tasks that don't resume. The template's inspect-ai floor moves to >=0.3.241 (first version with CheckpointSampleConfig). The scaffolder repo itself is unaffected: _templates/** is excluded from pyright and the tests transform template files as text rather than importing them. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Make build_plot thread-safe (#18) * fix: make build_plot thread-safe via matplotlib OO API Render off a standalone Figure with an explicit Agg canvas instead of pyplot. pyplot's global figure registry and rc_context's process-wide rcParams mutation both race under concurrent calls (e.g. when report generation is offloaded to a thread pool), corrupting plots or raising. Styling is now applied per-artist, and the one-time global font registration is guarded by a lock. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * refactor: simplify tick-label font handling and strengthen thread-safety tests - Apply tick-label font via tick_params(labelfontfamily=...) (matplotlib 3.7+), dropping the post-draw realize-then-mutate loop and a FontProperties. - Remove now-dead file-level pyright suppressions (the OO API is fully typed). - Tests: add a deterministic pyplot figure-registry leak check, forbid matplotlib.use in the no-globals test, and compare concurrent renders byte-for-byte against a single-threaded reference. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * perf: make font registration lock-free after first call Address PR review: _register_bundled_font acquired the global lock and rescanned the font list on every render, briefly serializing concurrent calls. Use double-checked locking with a module-level flag so the common case (font already registered) skips the lock entirely. The font test resets the flag to keep exercising the real registration path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
fix(tool_cli): pass RPC method args by keyword, not positionally (#19) The sandbox_service client generated by inspect_ai is keyword-only after the method name (`call_<service>(method, **params)`) and has been since the feature was introduced (inspect_ai #922). tool_cli's generated CLI passed method args positionally, so every command except `list` failed at runtime with "call_<service>() takes 1 positional argument but 2 were given". This was masked by tests whose fake client used a permissive `*args, **kwargs` signature. Pass describe_tool/describe_tool_for_call/call_tool args by keyword (tool_name, arguments, snapshot_token), and make the test fakes keyword-only so they mirror the real client and catch regressions. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add per-sample report and artifact folder helpers (#17) * build: make universal-pathlib a core dependency * feat: add report_dir and artifacts_dir path helpers Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: add write_report writer Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: add write_artifacts writer Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat: add write_artifact single-file writer Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test: cover path-traversal rejection in artifacts helpers * refactor: remove write_report_artifacts in favor of artifacts module Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * docs: document per-sample report and artifact helpers * refactor: tighten flat-path-component validation and add edge-case tests Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix: harden clear loop against symlinks and reject control-char names Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * test: extract active_log fixture to remove setup boilerplate * fix: heal non-directory dest in _write_files Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * style: apply ruff format --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Add workdir field to Workspace (#16) * feat: add workdir field to Workspace Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test: drop redundant Workspace.workdir test Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: document Workspace.workdir in README Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: fix grammar in Workspace.workdir docstring Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
build: switch version source to hatch-vcs for tag-driven releases (#15) Adopt hatch-vcs so the package version is derived from the latest git tag instead of carried in pyproject.toml. Releasing is now a single step: draft a GitHub Release with tag vX.Y.Z and the workflow builds and publishes that exact version. No PR to bump the version is needed. Changes: - pyproject.toml: drop static `version`, add `dynamic = ["version"]`, add hatch-vcs to build requires, point `[tool.hatch.version]` at the VCS source. - uv.lock: regenerated; the self-version line drops out (uv doesn't pin a workspace member's own version when it's dynamic, so `uv lock --check` remains stable across non-tagged commits). - release.yml: drop the tag-vs-pyproject consistency check (no pyproject version to drift from). Add `fetch-depth: 0` so hatch-vcs can see the release tag. - pr-and-main.yaml: add `fetch-depth: 0` to lint, test, and uv-lock-check jobs so `uv sync` (which installs the editable package and therefore invokes the build backend) can resolve the version. Local builds without tags now produce versions like `0.4.1.dev1+g935df1f.d20260519` (latest tag + commits-since + sha + date). Builds AT a tag produce the clean tag string, e.g. `0.4.0`. Only the latter is uploaded to PyPI.
Polish PyPI release: trim sdist, README, workflow ergonomics, bump to… … 0.4.0 (#13) * ci(release): trigger on 'released', add manual dispatch and PyPI URL - types: [released] excludes pre-releases from auto-publishing - workflow_dispatch enables manual re-runs without cutting a new release - environment.url renders a clickable link to the PyPI project in GitHub's deployment UI * chore: bump version to 0.4.0 * docs(README): install from PyPI; correct license note inspect-eval-utils is now on PyPI, so install commands no longer need the git+ssh://... URL. License section previously said "Internal METR project", which has been wrong since the package was MIT-licensed. * build(sdist): restrict to package, LICENSE, README, and pyproject By default hatchling included every non-gitignored path in the sdist, which leaked IDE config (.idea/), local AI assistant files (.codex), and internal planning docs under docs/superpowers/ to PyPI. Restrict the sdist to the minimum needed to rebuild the wheel. * ci(release): drop workflow_dispatch trigger Manual dispatch defaults to running against the default branch, which would make the tag/version check (it reads GITHUB_REF_NAME) fail unless the dispatcher remembered to pick a tag from the ref dropdown. The recovery scenarios we cared about (transient PyPI errors, env config fixes) are already covered by GitHub's "Re-run failed jobs" button on the original release run, which preserves the tag ref.
Support dynamic tools in sandbox CLI (#10) * fix: escape generated tool CLI help text * feat: add dynamic tool CLI resolver * feat: serialize dynamic tool CLI schemas * fix: type dynamic tool CLI schema serialization * feat: resolve tool CLI methods dynamically * fix: keep static tool CLI working with dynamic methods * feat: generate dynamic tool CLI client * fix: update install test for dynamic tool CLI client * fix: handle dynamic client required args * fix: harden dynamic client parser edge cases * fix: require json args for colliding dynamic params * test: cover dynamic tool CLI client behavior * feat: use dynamic tool CLI completion * fix: remove static completion test expectation * fix: avoid shell expansion in dynamic completion * docs: describe dynamic tool source CLI behavior * fix: type setting tool CLI doc test * docs: document dynamic tool CLI usage * refactor: remove static tool CLI generator * fix: remove stale tool CLI import * fix: remove install-time tool source resolution * fix: refresh dynamic schema for tool calls * fix: keep dynamic call schema and execution consistent