Skip to content

feat(human_agent): let several people into one sandbox, with operators - #122

Draft
spruceb wants to merge 8 commits into
mainfrom
spruce/hawk-human-access
Draft

spruceb wants to merge 8 commits into
mainfrom
spruce/hawk-human-access

Conversation

@spruceb

@spruceb spruceb commented Sep 24, 2026 •

Copy link
Copy Markdown

Hawk human evals let exactly one person SSH into a sandbox, so whoever runs a baseline can't look inside it without the baseliner's private key, and sharing that key counts as a security incident. This adds humans=[{name, public_key, role}] to human_agent, so several registered people can log in with their own keys. A participant does the task; several share the clock and submission, and their recordings and task notes carry their name. An operator logs in as the task user without starting the clock, being recorded, or being able to pause, submit or quit. A root_operator is the same, as root.

Dropbear has no environment= option, so every key gets a forced command: a small root-owned shell that tags the session with role and name, appends to an access log that's saved with the sample, and then runs what the client asked for (a shell, a command, or the sftp subsystem, so VS Code Remote-SSH still works). The .bashrc block returns early for operators, and dropbear only drops -w when there's a root operator, whose key is the only one in root's authorized_keys. Without humans, the single public_key is written exactly as before, and that's all Hawk sends for single-human evals.

It also brings its own tmux. Most task images have none, and a dropped SSH connection takes a human's running work with it; tasks were adding it one image at a time (harder-tasks #652, #654). When command -v tmux finds nothing, the agent installs a static tmux 3.5a at /usr/bin/tmux, plus /etc/tmux.conf if absent. The binaries are built for x86_64, aarch64 and armv7hf by resources/tmux.Dockerfile from sha256-pinned upstream tarballs. ncurses has common terminfo compiled in (kitty's and Ghostty's TERMs map to xterm-256color), so tmux starts even in images with no terminfo.

Tested with unit tests and a Docker end-to-end test: a participant, an operator and a root operator in one sandbox, checking the clock, the task refusal, sftp, root, tmux, and the note and access log in the sample. It also ran end to end on a Hawk dev stack through hawk#1911, with the details there. The Hawk side is METR/hawk#1911. Once this merges and releases, the prd/stg defaultHumanAgentPackage pin needs bumping before that deploys.

🤖 Generated with Claude Code

`humans=[{name, public_key, role}]` lets more than one registered person SSH
into a human-eval sandbox with their own key. Roles:

- participant: does the task; shares the clock and submission with any
  other participants. Recordings and `task note` carry their name.
- operator: logs in as the same user to watch. Skips the session
  recording, the instructions banner and `task resume`, and `task` only
  allows status/instructions, so an operator can't restart, pause or end
  the participant's run.
- root_operator: the same, as root (dropbear drops -w only then, and only
  their keys go in root's authorized_keys).

Every key gets a forced command (dropbear has no environment= option), a
root-owned access shell that tags the session and appends to an access log
the service keeps with the sample. Without `humans`, the single public_key
is written plain exactly as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved critical and moderate security and correctness issues remain.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 3 High severity · 2 Medium severity

Open (5)
What changed in this PR

Adds multi-person human-agent sandbox access with participant, operator, and root-operator roles.

Changes:

  • Adds role-aware forced SSH access, root access, and audit logging.
  • Attributes notes and recordings to named humans.
  • Adds unit and Docker integration coverage.
File Reviewed changes and final notes
packages/​agents/​tests/​test_human_agent.py Multi-human Docker integration coverage.
packages/​agents/​tests/​test_human_agent_task_script.py Tests operator restrictions and note attribution.
packages/​agents/​tests/​test_human_agent_service.py Tests service behavior and access-log persistence.
packages/​agents/​tests/​test_human_agent_install.py Tests shell rendering and recording behavior.
packages/​agents/​tests/​test_human_agent_agent.py Tests SSH setup and human validation.
packages/​agents/​src/​metr_agents/​human_agent/​task_script.py Adds operator command restrictions and named notes. Critical, 2 votes: role and identity checks rely on mutable environment variables and are not enforced at an authenticated service boundary.
packages/​agents/​src/​metr_agents/​human_agent/​service.py Handles named notes and access-log collection. Moderate, 2 votes: stale logs can be included in legacy runs; gate collection and reset logs for multi-human samples. Moderate, 1 vote: timeout or cancellation omits the access log.
packages/​agents/​src/​metr_agents/​human_agent/​install.py Configures role-aware shell initialization and recording names.
packages/​agents/​src/​metr_agents/​human_agent/​agent.py Implements human validation, forced commands, SSH setup, and access logging. Critical, 4 votes: duplicate public keys are accepted. Critical, 4 votes: mode 666 allows audit-log tampering. Moderate, 3 votes: humans=[] falls back to legacy setup instead of being rejected. Moderate, 1 vote: SFTP subsystem requests are not routed to the server binary.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread packages/agents/src/metr_agents/human_agent/agent.py
Comment thread packages/agents/src/metr_agents/human_agent/agent.py Outdated
Comment thread packages/agents/src/metr_agents/human_agent/task_script.py
Comment thread packages/agents/src/metr_agents/human_agent/agent.py Outdated
Comment thread packages/agents/src/metr_agents/human_agent/service.py Outdated
spruceb and others added 7 commits September 24, 2026 13:04
Most task images have no tmux, and a dropped SSH connection otherwise takes
a human's running work with it; tasks have been adding it one image at a
time (harder-tasks #652, #654). The human agent now installs its own at
/usr/bin/tmux, alongside dropbear, when `command -v tmux` finds nothing,
plus /etc/tmux.conf (mouse scrolling, 50k scrollback) if absent.
Best-effort: a failure logs and never fails the sample.

tmux 3.5a, libevent and ncurses are built from sha256-pinned upstream
tarballs by resources/tmux.Dockerfile, static against musl. ncurses
searches Debian's and Alpine's terminfo dirs and compiles common entries
in, with kitty's and Ghostty's TERMs mapped to xterm-256color, so tmux
starts even in images with no terminfo. Checked in python slim and busybox
on x86_64 and aarch64 with xterm-256color, xterm-kitty and xterm-ghostty.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Built with resources/tmux.Dockerfile under linux/arm/v7; starts in busybox
with TERM=xterm-ghostty.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… checks

- Who logged in now comes from dropbear itself: with `humans`, it runs with
  -E into a fresh file in a root-only directory, and the service keeps that
  log with each key fingerprint annotated "name, role". The access shell no
  longer writes a world-writable log a participant could edit or forge.
- The log is only read when `humans` was given, so a single-human run
  never ships a stale or planted file.
- Reject two humans with the same key material: sshd takes the first
  matching line, so one would get the other's role.
- An explicit `humans=[]` is rejected instead of silently falling back to
  the single-key setup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
From review (Megan): two roles, not three. Operators now log in as root,
with their keys only in root's authorized_keys, so a root login naturally
skips the task user's .bashrc hooks (clock, recording, instructions); the
env marker only matters when the task user is itself root. The plain
same-user operator bought little, since its `task` limits were never a
security boundary, and crash rescue needs root anyway.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a crash

From review (Megan): the point of operators is getting into a run that
ended badly, to rescue data or debug, and the sandbox already survives the
sample (hawk defaults runner.cleanup to false for human evals). Until now
the agent stopped dropbear whenever the sample ended, error included.

With operators, sample end (submit, quit, error, timeout, cancellation) now
revokes the participants' keys and disconnects them, and leaves dropbear up
for the operators, whose keys are root's; `hawk delete` ends it. Without
operators, teardown is unchanged.

Participants are found by the access shell's role marker, which their tmux
or VS Code server inherits too. Reading another user's
/proc/<pid>/environ needs CAP_SYS_PTRACE, which sandboxes drop, so the scan
runs as the participants' own user; root then kills only their dropbear
connections, never the listener. A Docker test crashes a sample with an
operator in: the operator still gets in as root, the participant key is
refused, and a detached process the participant left running is gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
From review (Megan): a reader of the eval log should know how far to
trust the login record. It is written by dropbear as root, so operators
could edit it, and so could participants when the task runs as root. The
note is added as the log leaves the sandbox, so it can't be altered from
inside.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… submit

From review (Megan): only `task submit` and `task quit` saved the login
record, so a run that errored, timed out or was cancelled lost it, and
that is when operators get in and a reviewer most needs to know who was
there. The service now reads it at the end of any other sample, under a
shield with its own deadline so it can't make the sample uncancellable.
A runner that dies outright still leaves it only in the sandbox.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants