Follow
Adam Gleave
Adam Gleave
CEO at FAR AI
Verified email at far.ai - Homepage
Title
Cited by
Cited by
Year
Stable-baselines3: Reliable reinforcement learning implementations
A Raffin, A Hill, A Gleave, A Kanervisto, M Ernestus, N Dormann
Journal of machine learning research 22 (268), 1-8, 2021
55102021
Stable baselines
A Hill, A Raffin, M Ernestus, A Gleave, A Kanervisto, R Traore, P Dhariwal, ...
10622018
Adversarial policies: Attacking deep reinforcement learning
A Gleave, M Dennis, C Wild, N Kant, S Levine, S Russell
International Conference on Learning Representations, 2020
6452020
Firmament: Fast, centralized cluster scheduling at scale
I Gog, M Schwarzkopf, A Gleave, RNM Watson, S Hand
12th USENIX Symposium on Operating Systems Design and Implementation (OSDI …, 2016
3582016
Multi-agent risks from advanced AI
L Hammond, A Chan, J Clifton, J Hoelscher-Obermaier, A Khan, ...
arXiv preprint arXiv:2502.14143, 2025
2732025
imitation: Clean imitation learning implementations
A Gleave, M Taufeeque, J Rocamonde, E Jenner, SH Wang, S Toyer, ...
arXiv preprint arXiv:2211.11972, 2022
152*2022
Adversarial Policies Beat Superhuman Go AIs
TT Wang, A Gleave, T Tseng, N Belrose, J Miller, MD Dennis, Y Duan, ...
arXiv preprint arXiv:2211.00241, 2022
122*2022
Scaling trends for data poisoning in llms
D Bowen, B Murphy, W Cai, D Khachaturov, A Gleave, K Pelrine
Proceedings of the AAAI Conference on Artificial Intelligence 39 (26), 27206 …, 2025
109*2025
Quantifying differences in reward functions
A Gleave, M Dennis, S Legg, S Russell, J Leike
International Conference on Learning Representations, 2021
1082021
Invariance in policy optimisation and partial identifiability in reward learning
JMV Skalse, M Farrugia-Roberts, S Russell, A Abate, A Gleave
International Conference on Machine Learning, 32033-32058, 2023
852023
Inverse reinforcement learning for video games
A Tucker, A Gleave, S Russell
Deep Reinforcement Learning Workshop at NeurIPS, 2018
742018
Multi-task maximum entropy inverse reinforcement learning
A Gleave, O Habryka
GoalsRL Workshop at ICML, 2018
592018
Exploiting novel GPT-4 apis
K Pelrine, M Taufeeque, M Zając, E McLean, A Gleave
arXiv preprint arXiv:2312.14302, 2023
522023
Understanding learned reward functions
EJ Michaud, A Gleave, S Russell
Deep Reinforcement Learning Workshop at NeurIPS, 2020
522020
Uncertainty estimation for language reward models
A Gleave, G Irving
arXiv preprint arXiv:2203.07472, 2022
482022
A primer on maximum causal entropy inverse reinforcement learning
A Gleave, S Toyer
arXiv preprint arXiv:2203.11409, 2022
452022
Active inverse reward design
S Mindermann, R Shah, A Gleave, D Hadfield-Menell
GoalsRL Workshop at ICML, 2018
402018
Scaling trends in language model robustness
N Howe, I McKenzie, O Hollinsworth, M Zajac, T Tseng, A Tucker, ...
arXiv preprint arXiv:2407.18213, 2024
33*2024
On the fragility of learned reward functions
L McKinney, Y Duan, D Krueger, A Gleave
arXiv preprint arXiv:2301.03652, 2023
322023
Planning in a recurrent neural network that plays sokoban
M Taufeeque, P Quirke, M Li, C Cundy, AD Tucker, A Gleave, ...
arXiv preprint arXiv:2407.15421, 2024
24*2024
The system can't perform the operation now. Try again later.
Articles 1–20