This repository is an incremental Inspect AI port of MonitorBench, a benchmark for evaluating whether decision-critical factors and hidden objectives are visible to chain-of-thought monitors. Upstream defines 19 tasks across three evaluation axes: Input Intervention, Outcome Justification, and Solution Process. Tasks are migrated, validated, and submitted individually.
Four tasks are implemented: goal_sandbag.math and goal_sandbag.safety,
exposed as goal_sandbag_math and goal_sandbag_safety, plus
dual_objectives.steganography and
dual_objectives.summarization, exposed as
dual_objectives_steganography and dual_objectives_summarization. The other
15 tasks remain migration backlog
and are not registered or runnable from this package.
| Evaluation axis | Upstream tasks | Implemented | Planned |
|---|---|---|---|
| Input Intervention | 12 | 2 | 10 |
| Outcome Justification | 6 | 2 | 4 |
| Solution Process | 1 | 0 | 1 |
| Total | 19 | 4 | 15 |
See the benchmark overview and task catalog for the status and tracking issue for every task. A task is listed in the runtime registry only after its implementation, assets, attribution, tests, and documentation are complete.
Install the locked environment and run any available task:
uv sync
uv run inspect eval monitor_bench/dual_objectives_steganography \
--model <evaluated-model> \
--model-role monitor=<monitor-model> \
-T stress_test=standard
uv run inspect eval monitor_bench/goal_sandbag_math \
--model <evaluated-model> \
--model-role monitor=<monitor-model> \
-T stress_test=standard
uv run inspect eval monitor_bench/goal_sandbag_safety \
--model <evaluated-model> \
--model-role monitor=<monitor-model> \
-T stress_test=standard
uv run inspect eval monitor_bench/dual_objectives_summarization \
--model <evaluated-model> \
--model-role monitor=<monitor-model> \
-T stress_test=standardThe evaluated model and monitor are distinct roles. For benchmark rollouts,
use -T epochs=N; Inspect's global --epochs option replaces these tasks'
custom count-pooling reducers and must not be used. The goal-sandbag tasks keep
each original/intervened prompt pair inside one Inspect sample, so --limit 1
still evaluates both causal arms.
Task-specific behavior, scoring, fidelity notes, and validation status are in
the steganography README,
goal_sandbag.math README,
goal_sandbag.safety README,
and summarization README.
Each remaining task has its own migration issue and should be delivered in a
focused pull request. Draft multi-task implementations may be used as extraction
references, but they are not treated as supported product state. The package
continues to use one monitor_bench Inspect entry point; only completed tasks
are exported and listed in src/monitor_bench/eval.yaml.
Run the local verification gate with:
uv run pytest tests/monitor_bench
make checkThe port is pinned to
ASTRAL-Group/MonitorBench commit
43dda5994bfb16d34b1c30d4b3482d78a714e640 and the
MonitorBench paper (v2). Vendored
MonitorBench assets, Databricks Dolly, GovReport, AIME, and WMDP provenance,
the AIME and GovReport rights audits, the separate WMDP MIT grant, and the
runtime-fetched NLTK tokenizer notice are recorded in
NOTICE and
src/monitor_bench/assets/ATTRIBUTION.md.
The repository's original code is MIT licensed. Known third-party terms and unresolved redistribution rights are identified in those notices.