Currently an independent AI Safety Evals Engineer contributing to UK AISI's Inspect eval framework and publishing results from independent evaluations.
An eval built on Inspect attempting to measure how abliteration and prefill attacks on open-weight models unlock refusal on harmful queries, and whether they are stronger when used together.
inspect_ai
- #3709 - vLLM chat-template controls for base-model evals
- #3969 -
pass_kepoch reducer (τ-bench pass^k consistency metric) - #4380 - Add instance parameter and path canonicalization to memory
- #4035 - Krippendorff's α metric for multi-judge agreement
inspect_evals
- #1429 - fix CodeIPI exfiltration scorer to check tool-result messages
- #1501 -
cyberseceval_4: tolerate fenced / prose-wrapped judge JSON - #1503 - fix
mean_ofon_missing="skip"to also skipNone-valued samples
inspect_scout
- #455 - resolve
ModelEventinput refs from theevents_datapool schema
Hit me up! Open to collaboration on evals, tooling, or if you just want to contact me :)