Skip to content
 
 

Latest commit

 

History

105 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MonitorBench for Inspect

This repository is an incremental Inspect AI port of MonitorBench, a benchmark for evaluating whether decision-critical factors and hidden objectives are visible to chain-of-thought monitors. Upstream defines 19 tasks across three evaluation axes: Input Intervention, Outcome Justification, and Solution Process. Tasks are migrated, validated, and submitted individually.

Four tasks are implemented: goal_sandbag.math and goal_sandbag.safety, exposed as goal_sandbag_math and goal_sandbag_safety, plus dual_objectives.steganography and dual_objectives.summarization, exposed as dual_objectives_steganography and dual_objectives_summarization. The other 15 tasks remain migration backlog and are not registered or runnable from this package.

Migration status

Evaluation axis Upstream tasks Implemented Planned
Input Intervention 12 2 10
Outcome Justification 6 2 4
Solution Process 1 0 1
Total 19 4 15

See the benchmark overview and task catalog for the status and tracking issue for every task. A task is listed in the runtime registry only after its implementation, assets, attribution, tests, and documentation are complete.

Quick start

Install the locked environment and run any available task:

uv sync
uv run inspect eval monitor_bench/dual_objectives_steganography \
  --model <evaluated-model> \
  --model-role monitor=<monitor-model> \
  -T stress_test=standard

uv run inspect eval monitor_bench/goal_sandbag_math \
  --model <evaluated-model> \
  --model-role monitor=<monitor-model> \
  -T stress_test=standard

uv run inspect eval monitor_bench/goal_sandbag_safety \
  --model <evaluated-model> \
  --model-role monitor=<monitor-model> \
  -T stress_test=standard

uv run inspect eval monitor_bench/dual_objectives_summarization \
  --model <evaluated-model> \
  --model-role monitor=<monitor-model> \
  -T stress_test=standard

The evaluated model and monitor are distinct roles. For benchmark rollouts, use -T epochs=N; Inspect's global --epochs option replaces these tasks' custom count-pooling reducers and must not be used. The goal-sandbag tasks keep each original/intervened prompt pair inside one Inspect sample, so --limit 1 still evaluates both causal arms.

Task-specific behavior, scoring, fidelity notes, and validation status are in the steganography README, goal_sandbag.math README, goal_sandbag.safety README, and summarization README.

Development

Each remaining task has its own migration issue and should be delivered in a focused pull request. Draft multi-task implementations may be used as extraction references, but they are not treated as supported product state. The package continues to use one monitor_bench Inspect entry point; only completed tasks are exported and listed in src/monitor_bench/eval.yaml.

Run the local verification gate with:

uv run pytest tests/monitor_bench
make check

Provenance and licensing

The port is pinned to ASTRAL-Group/MonitorBench commit 43dda5994bfb16d34b1c30d4b3482d78a714e640 and the MonitorBench paper (v2). Vendored MonitorBench assets, Databricks Dolly, GovReport, AIME, and WMDP provenance, the AIME and GovReport rights audits, the separate WMDP MIT grant, and the runtime-fetched NLTK tokenizer notice are recorded in NOTICE and src/monitor_bench/assets/ATTRIBUTION.md.

The repository's original code is MIT licensed. Known third-party terms and unresolved redistribution rights are identified in those notices.

About

Inspect AI port of MonitorBench (github.com/ASTRAL-Group/MonitorBench), the reference implementation for evaluating chain-of-thought monitorability.

Resources

Contributing

Stars

0 stars

Watchers

1 watching

Forks

Contributors

Languages