Skip to content

Convert broken_env, reward_hacking, and sandbagging to scout - #30

Merged
sjawhar merged 11 commits into
mainfrom
feature/scout
Dec 1, 2025
Merged

sjawhar merged 11 commits into
mainfrom
feature/scout

Conversation

@sjawhar

@sjawhar sjawhar commented Nov 22, 2025 •

Copy link
Copy Markdown
Contributor

See the results of the scanners here: https://inspect-ai.internal.metr.org/scan/hcast-claude-opus-4-5-triframe-n4-202511#/scan/scan_id=n7sy5pM9XT5RovVVuMSFsM?scanner=reward_hacking_scanner

Currently the broken_env scanner results won't load, I think I can fix that :)

@sjawhar sjawhar self-assigned this Nov 22, 2025
@sjawhar
sjawhar force-pushed the feature/scout branch 3 times, most recently from 1d8af75 to abf7e73 Compare November 24, 2025 19:22
@sjawhar
sjawhar force-pushed the feature/scout branch 16 times, most recently from 52d7d5e to a9a3575 Compare November 29, 2025 22:44
sjawhar and others added 8 commits November 29, 2025 22:46
This commit introduces two new scanners for detecting reward hacking and
sandbagging in transcripts. It includes utility functions for chunking
transcripts and parsing JSON with Pydantic, along with necessary updates
to the project dependencies and registry.
- Fix problem when we have short transcript
- Allow to have a version of the prompt for a single transcript at once
@sjawhar sjawhar changed the title Convert broken_env scorer to scanner Convert broken_env, reward_hacking, and sandbagging to scout Nov 30, 2025
@sjawhar
sjawhar marked this pull request as ready for review November 30, 2025 00:15
@sjawhar
sjawhar requested review from pipmc and satojk November 30, 2025 00:20
@sjawhar
sjawhar force-pushed the feature/scout branch 3 times, most recently from 522da8a to ecd9720 Compare November 30, 2025 03:35
Comment thread .devcontainer/Dockerfile
Comment thread .dockerignore
@@ -4,4 +4,6 @@
!uv.lock
!README.md

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why the README.md files? Is it because they're specified in the pyproject.toml files and so uv will expect them to be there?

(If so, maybe you don't need to specify this one as the top-level pyproject.toml file no longer mentions a README)

Comment thread .github/workflows/pr-and-main.yaml
Comment thread .github/workflows/pr-and-main.yaml Outdated
fi

package_name="metr_${package}"
verion_file="src/${package_name}/version.py"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
verion_file="src/${package_name}/version.py"
version_file="src/${package_name}/version.py"

Comment thread .github/workflows/pr-and-main.yaml Outdated
PREVIOUS_VERSION="$(uv version --short)"
git checkout HEAD -- pyproject.toml
NEW_VERSION="$(uv version --bump=stable --short 2>/dev/null || uv version --bump=patch --short)"
echo "__version__ = \"${NEW_VERSION}\"" > "${verion_file}"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
echo "__version__ = \"${NEW_VERSION}\"" > "${verion_file}"
echo "__version__ = \"${NEW_VERSION}\"" > "${version_file}"

Comment on lines +45 to +58
def _get_agent_tools(
transcript: inspect_scout.Transcript,
) -> str | None:
agent_tools = _AGENT_TOOLS.get(None)
if agent_tools:
return agent_tools

for event in reversed(transcript.events):
if isinstance(event, inspect_ai.event.ModelEvent) and event.tools:
agent_tools = json.dumps([tool.model_dump() for tool in event.tools])
_AGENT_TOOLS.set(agent_tools)
return agent_tools

return None

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this actually a safe way to get the agent's tools? For example, here is an ai_rd_fix_embedding run where it appears to me that the last call to the model (in epoch 1) only supplies the rate_options tool.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(Also, I get a bunch of errors when I try and scan that eval log using the command scout scan metr_scanners/broken_env_scanner -T 2025-09-25T15-19-41+00-00_ai-rd-fix-embedding_a7nCsDBesrNVNNmSKNeLfN.eval --model openai/gpt-5-mini, specifically the exception ijson.common.IncompleteJSONError: lexical error: invalid char in json text.. I wasn't able to debug this because I couldn't work out how to catch the exception in the debugger)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it appears to me that the last call to the model (in epoch 1) only supplies the rate_options tool.

TRIFRAME!! 😆

Comment on lines +98 to +101
response_schema=inspect_ai.model.ResponseSchema(
name=ResultClass.__name__,
json_schema=inspect_ai.util.json_schema(ResultClass),
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re the Inspect docs and the Scout code: was there a specific reason for using structured output instead of an "answer tool"? Any idea if performance is the same/better than using a tool?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The current broken_env scorer uses structured output, and we seem happy with its output. I think Romain (the trialee who wrote most of this code) simply copied that. My preference would be that we have a single consistent method for running scanners. Happy to change, but I wouldn't have the time to run that comparison myself.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks pretty similar to Scout's built in LLM scanner, except with support for chunking messages and more capable retry logic - I guess it probably wouldn't shorten the code to implement this on top of the built in LLM scanner if we still had to implement chunking, and it wouldn't be possible to replace its retry logic, so this seems okay to me

Comment on lines +80 to +98
for chunk in result_chunks:
if "[M1]" in chunk.text and "[M2]" in chunk.text:
test_text = "I found an issue in [M1] and [M2]"
refs = chunk.extract_references(test_text)

assert len(refs) == 2
assert refs[0].cite == "[M1]"
assert refs[0].type == "message"
assert refs[1].cite == "[M2]"
assert refs[1].type == "message"
break
else:
chunk = result_chunks[0]
message_ids = re.findall(r"\[M(\d+)\]", chunk.text)
if len(message_ids) >= 1:
test_text = f"I found an issue in [M{message_ids[0]}]"
refs = chunk.extract_references(test_text)
assert len(refs) == 1
assert refs[0].cite == f"[M{message_ids[0]}]"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It appears to me that "[M1]" in chunk.text and "[M2]" in chunk.text will always resolve to False, and so none of the block beneath that if statement will ever execute. Is that what you intended?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, I need to review the tests in more detail

Comment on lines +115 to +117
message_ids = re.findall(r"\[M(\d+)\]", chunk.text)

if message_ids:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there any reason to expect that message_ids would ever be falsy, and if so would it actually be appropriate to skip all the if statements below?

@sjawhar

sjawhar commented Nov 30, 2025

Copy link
Copy Markdown
Contributor Author

I've rewritten the tests to be more clear. Thanks for the close review!

@sjawhar
sjawhar requested a review from pipmc November 30, 2025 20:19
Comment on lines +121 to +162
async def _scan_transcript_single_pass(
transcript: inspect_scout.Transcript,
model: inspect_ai.model.Model,
result_type: type[QuotedResult],
*,
prompt_prefix: str,
prompt_suffix: str,
**prompt_kwargs: str,
) -> inspect_scout.Result:
transcript_str, extract_fn = await inspect_scout.messages_as_str(
transcript, include_ids=True
)

return await _scan_with_retry(
model,
result_type,
_render_partial_template(
_PROMPT_TEMPLATE_SINGLE_PASS,
**{
_PREFIX_KEY: prompt_prefix,
_SUFFIX_KEY: prompt_suffix,
_TRANSCRIPT_KEY: DONT_RENDER,
},
).format(
**prompt_kwargs,
**{_TRANSCRIPT_KEY: transcript_str},
),
extract_fn,
)


async def _scan_transcript_chunked(
transcript: inspect_scout.Transcript,
model: inspect_ai.model.Model,
result_type: type[QuotedResult],
*,
prompt_prefix: str,
prompt_suffix: str,
early_messages_count: int = 5,
max_chunk_size: int = 150_000,
**prompt_kwargs: str,
) -> inspect_scout.Result:

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm pretty sure these two functions can be combined into one by checking if there's only one chunk. It also seems like poor separation of concerns that the chunking function is doing template rendering

async def scan(transcript: inspect_scout.Transcript) -> inspect_scout.Result:
model = inspect_ai.model.get_model(model_name)
prompt_kwargs = prompt_values(transcript) if prompt_values else {}
if len(transcript.messages) <= early_messages_count:

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems like the wrong thing to be checking (message count vs. chunk length).

Comment on lines +121 to 134
if git cat-file -e HEAD~:pyproject.toml 2>/dev/null; then
git checkout HEAD~ -- pyproject.toml
previous_version="$(uv version --short)"
git checkout HEAD -- pyproject.toml
else
previous_version=""
fi

git checkout HEAD~ -- pyproject.toml
PREVIOUS_VERSION="$(uv version --short)"
git checkout HEAD -- pyproject.toml
NEW_VERSION="$(uv version --bump=stable --short 2>/dev/null || uv version --bump=patch --short)"
echo "__version__ = \"${NEW_VERSION}\"" > "${verion_file}"
echo "__version__ = \"${NEW_VERSION}\"" > "${version_file}"

git add pyproject.toml "${verion_file}"
git add pyproject.toml "${version_file}"
git tag "${package_name}/v${NEW_VERSION}"
commit_message+="${package_name}: ${PREVIOUS_VERSION} → ${NEW_VERSION}\n"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Need to use the same case for PREVIOUS_VERSION everywhere. Shellcheck catches this:

Line 23:
              previous_version=""
              ^-- [SC2034](https://www.shellcheck.net/wiki/SC2034) (warning): previous_version appears unused. Verify use (or export if used externally).
 
Line 31:
            commit_message+="${package_name}: ${PREVIOUS_VERSION} → ${NEW_VERSION}\n"
                                              ^-- [SC2153](https://www.shellcheck.net/wiki/SC2153) (info): Possible misspelling: PREVIOUS_VERSION may not be assigned. Did you mean previous_version?

@pipmc pipmc Dec 1, 2025 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe we should try using actionlint in CI and/or our devcontainers as it apparently runs shellcheck on Actions workflow run steps (it has an official Docker image and instructions for using in GitHub Actions too)

async for chunk in transcript_messages_to_chunks(messages, max_chunk_size=50)
async for chunk in chunks.transcript_messages_to_chunks(
messages, max_chunk_size=50
)
]

for chunk in result_chunks:
message_count = len(re.findall(r"\[M\d+\]", chunk.text))
assert message_count >= 1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the max_chunk_size is smaller than each of the messages, then surely this should be

Suggested change
assert message_count >= 1
assert message_count == 1

for all the chunks?

@sjawhar
sjawhar force-pushed the feature/scout branch 2 times, most recently from 4c425dd to 1e6e0d8 Compare December 1, 2025 13:51

@pipmc pipmc left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I've run broken_env and reward_hacking locally and neither are obviously broken. Will run more in anger soon with the Opus 4.5 runs.

@sjawhar
sjawhar merged commit 6eaea05 into main Dec 1, 2025
3 checks passed
@sjawhar
sjawhar deleted the feature/scout branch December 1, 2025 16:17
@sjawhar

sjawhar commented Dec 1, 2025

Copy link
Copy Markdown
Contributor Author

Publish workflow failed 😆

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants