Follow
Erik Jenner
Erik Jenner
Google DeepMind
Verified email at google.com - Homepage
Title
Cited by
Cited by
Year
Foundational challenges in assuring alignment and safety of large language models
U Anwar, A Saparov, J Rando, D Paleka, M Turpin, P Hase, ES Lubana, ...
TMLR, 2024
4652024
Chain of thought monitorability: A new and fragile opportunity for ai safety
T Korbak, M Balesni, E Barnes, Y Bengio, J Benton, J Bloom, M Chen, ...
arXiv preprint arXiv:2507.11473, 2025
2182025
imitation: Clean imitation learning implementations
A Gleave, M Taufeeque, J Rocamonde, E Jenner, SH Wang, S Toyer, ...
arXiv preprint arXiv:2211.11972, 2022
1272022
When chain of thought is necessary, language models struggle to evade monitors
S Emmons, E Jenner, DK Elson, RA Saurous, S Rajamanoharan, H Chen, ...
arXiv preprint arXiv:2507.05246, 2025
882025
Chain of thought monitorability: A new and fragile opportunity for ai safety, 2025
T Korbak, M Balesni, E Barnes, Y Bengio, J Benton, J Bloom, M Chen, ...
URL https://arxiv. org/abs/2507.11473, 2026
562026
Obfuscated activations bypass LLM latent-space defenses
L Bailey, A Serrano, A Sheshadri, M Seleznyov, J Taylor, E Jenner, ...
International Conference on Learning Representations 2026, 146838-146882, 2026
552026
Steerable Partial Differential Operators for Equivariant Neural Networks
E Jenner, M Weiler
ICLR, 2022
492022
Evidence of Learned Look-Ahead in a Chess-Playing Neural Network
E Jenner, S Kapur, V Georgiev, C Allen, S Emmons, S Russell
NeurIPS, 2024
422024
When Your AI Deceives You: Challenges with Partial Observability of Human Evaluators in Reward Learning
L Lang, D Foote, S Russell, A Dragan, E Jenner, S Emmons
NeurIPS, 2024
34*2024
Can reasoning models obfuscate reasoning? stress-testing chain-of-thought monitorability
A Zolkowski, W Xing, D Lindner, F Tramèr, E Jenner
arXiv preprint arXiv:2510.19851, 2025
252025
others. 2025. Chain of thought monitorability: A new and fragile opportunity for ai safety
T Korbak, M Balesni, E Barnes, Y Bengio, J Benton, J Bloom, M Chen, ...
Preprint, 22
2422
Preprocessing Reward Functions for Interpretability
E Jenner, A Gleave
NeurIPS Cooperative AI workshop, 2021
232021
STARC: A General Framework For Quantifying Differences Between Reward Functions
J Skalse, L Farnik, SR Motwani, E Jenner, A Gleave, A Abate
ICLR, 2023
212023
RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
R Gupta, E Jenner
arXiv preprint arXiv:2506.14261, 2025
142025
Diffusion on syntax trees for program synthesis
S Kapur, E Jenner, S Russell
ICLR, 2025
102025
h’Eigeartaigh
U Anwar, A Saparov, J Rando, D Paleka, M Turpin, P Hase, ES Lubana, ...
S., Recchia, G., Corsi, G., Chan, A., Anderljung, M., Edwards, L., Petrov, A …, 2024
92024
A general framework for reward function distances
E Jenner, JMV Skalse, A Gleave
NeurIPS ML Safety Workshop, 2022
92022
Calculus on MDPs: Potential shaping as a gradient
E Jenner, H van Hoof, A Gleave
arXiv preprint arXiv:2208.09570, 2022
8*2022
A comparison of causal scrubbing, causal abstractions, and related methods
E Jenner, A Garriga-alonso, E Zverev
AI Alignment Forum, 2023
42023
Frontier Models Can Take Actions at Low Probabilities
A Serrano, W Xing, D Lindner, E Jenner
arXiv preprint arXiv:2603.02202, 2026
22026
The system can't perform the operation now. Try again later.
Articles 1–20