Skip to content

esolang: score with a trusted interpreter the agent cannot replace (#31) - #33

Open
IgorVoytyuk wants to merge 2 commits into
METR:mainfrom
IgorVoytyuk:fix/esolang-trusted-interpreter-issue-31
Open

IgorVoytyuk wants to merge 2 commits into
METR:mainfrom
IgorVoytyuk:fix/esolang-trusted-interpreter-issue-31

Conversation

@IgorVoytyuk

@IgorVoytyuk IgorVoytyuk commented Sep 22, 2026 •

Copy link
Copy Markdown

Fixes #31.

Problem

start() copies the interpreter into the agent-writable home
(/home/agent/interpreter) and score() executes exactly that path. An agent
can delete the root-provided interpreter, drop in its own script that ignores
the submitted program and prints the reference answers, and score 1 on every
test case without writing a single Broccoli program — defeating the task.

Fix

  • build_steps.json installs a trusted, root-owned copy of the interpreter under
    /opt/broccoli, a directory the agent cannot write. interpreter.js imports
    ./.broccoli.js relative to its own path, so both files are installed together
    and the trusted interpreter is self-contained (independent of CWD and of the
    agent's home).
  • score() now runs /opt/broccoli/interpreter and passes the submission as a
    plain argument without a shell (shell=True removed), so the agent can neither
    swap the interpreter nor inject shell syntax through the submitted path.
  • The agent keeps its own /home/agent/interpreter for running example programs;
    only scoring is moved to the trusted copy, so instructions and agent workflow
    are unchanged.

Test

Adds test_scoring_ignores_agent_writable_interpreter: it overwrites
/home/agent/interpreter with a cheat that touches a sentinel when executed and
prints a wrong answer, then asserts scoring never creates the sentinel and that
a correct submission still scores 1. The test fails on the current code (the agent
copy is executed, and its wrong answer becomes the score) and passes with this
change.

The existing test_fibonacci / test_collatz / test_factorial cases are
unaffected: the trusted interpreter behaves identically for genuine submissions.

Note: I ran JSON/lint and syntax checks locally, but I could not run the full
task-standard containerized suite in my environment, so please run CI on the
build step. Happy to adjust the install location (e.g. /usr/local/lib) if you
prefer a different convention.

Direction and authorship: Igor Voytyuk, worked jointly with Claude.

🤖 Generated with Claude Code

IgorVoytyuk and others added 2 commits September 22, 2026 20:50
…ETR#31)

score() executed /home/agent/interpreter, which start() copies into the
agent-writable home. An agent can delete that file, drop in its own script
that ignores the submitted program and prints the reference answers, and
score 1 on every test case without writing a single Broccoli program.

Fix: build_steps.json installs a trusted, root-owned copy of the interpreter
(and its sibling .broccoli.js, which interpreter.js imports relative to its
own path) under /opt/broccoli, a directory the agent cannot write. score()
now runs /opt/broccoli/interpreter and passes the submission as a plain
argument without a shell, so the agent can neither swap the interpreter nor
inject shell syntax through the submission path. The agent keeps its own
/home/agent/interpreter for running example programs.

Adds a regression test asserting score() never executes the agent-writable
interpreter.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The regression test only proved the agent-writable interpreter was not
executed. Also assert a correct submission still scores 1 while a cheat
interpreter is planted, which is the property task validity actually
depends on. The planted cheat now prints a wrong answer so scoring
through it is visible in the score as well as in the side effect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

esolang: score() runs an interpreter in the agent-writable home, so a submission can replace it and score 1 without writing a Broccoli program

1 participant