Skip to content

Fix eval_awareness and broken_env prompts to name required result fields - #120

Open
tomseimandi wants to merge 2 commits into
METR:mainfrom
tomseimandi:tweak-scanner-prompts
Open

tomseimandi wants to merge 2 commits into
METR:mainfrom
tomseimandi:tweak-scanner-prompts

Conversation

@tomseimandi

Copy link
Copy Markdown

Summary

We experienced multiple pydantic ValidationErrors on eval_awareness_scanner and broken_env_scanner scan jobs using OpenAI models (via OpenRouter), because each scanner's prompt never asks the judge for a field that its own result schema requires:

  • eval_awareness: the schema requires reason, but the prompt never names the reason field
  • broken_env: the schema requires success, but the prompt never mentions it

This PR changes only the two prompt suffixes and bumps both scanners to version=1.

Failure mode

_scan_with_retry sends the result schema as a non-strict ResponseSchema (strict is never set) and validates the completion with ResultClass.model_validate_json. Non-strict structured output guides the output format but does not force required keys. Some models follow the prompt wording over the schema.

Observed errors:

ValidationError: reason   Field required [type=missing]

Other scanners from this repository (reward_hacking, internet_leaves_tracks, sandbagging, etc.) never raise the error in practice because their suffix prompts contain instructions similar to those added by this PR.

Verification

  • Patching only the prompt suffix to name the missing field makes both previously erroring scanners validate on the first judge attempt
  • ruff check, ruff format --check, and basedpyright are clean, and the scanners test suite passes

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant