Follow
Yawen Duan
Title
Cited by
Cited by
Year
AI Alignment: A Contemporary Survey
J Ji, T Qiu, B Chen, J Zhou, B Zhang, D Hong, H Lou, K Wang, Y Duan, ...
ACM Computing Surveys, 2025
878*2025
Harms from Increasingly Agentic Algorithmic Systems
A Chan, R Salganik, A Markelius, C Pang, N Rajkumar, D Krasheninnikov, ...
Proceedings of the 2023 ACM Conference on Fairness, Accountability, and …, 2023
462*2023
TransNAS-Bench-101: Improving transferability and Generalizability of Cross-Task Neural Architecture Search
Y Duan, X Chen, H Xu, Z Chen, X Liang, T Zhang, Z Li
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern …, 2021
1392021
Adversarial Policies Beat Superhuman Go AIs
TT Wang, A Gleave, T Tseng, K Pelrine, N Belrose, J Miller, MD Dennis, ...
Proceedings of the 40th International Conference on Machine Learning, 35655 …, 2023
122*2023
The 2025 ai agent index: Documenting technical and safety features of deployed agentic ai systems
L Staufer, K Feng, K Wei, L Bailey, Y Duan, M Yang, AP Ozisik, S Casper, ...
The 2026 ACM Conference on Fairness, Accountability, and Transparency, 1536-1576, 2026
682026
On the fragility of learned reward functions
L McKinney, Y Duan, D Krueger, A Gleave
arXiv preprint arXiv:2301.03652, 2023
322023
CATCH: Context-based Meta Reinforcement Learning for Transferrable Architecture Search
X Chen, Y Duan, Z Chen, H Xu, Z Chen, X Liang, T Zhang, Z Li
Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23 …, 2020
292020
& Gao, W.(2023). Ai alignment: A comprehensive survey
J Ji, T Qiu, B Chen, B Zhang, H Lou, K Wang, Y Duan, Z He, L Vierling, ...
arXiv preprint arXiv:2310.19852 77, 0
15
Ai deception: Risks, dynamics, and controls
B Chen, S Fang, J Ji, Y Zhu, P Wen, J Wu, Y Tan, B Zheng, M Yuan, ...
arXiv preprint arXiv:2511.22619, 2025
122025
Bare minimum mitigations for autonomous AI development
J Clymer, I Duan, C Cundy, Y Duan, F Heide, C Lu, S Mindermann, ...
arXiv preprint arXiv:2504.15416, 2025
112025
Libra-leaderboard: Towards responsible ai through a balanced leaderboard of safety and capability
H Li, X Han, Z Zhai, H Mu, H Wang, Z Zhang, Y Geng, S Lin, R Wang, ...
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of …, 2025
102025
The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems
J Luo, J Dai, Z Chen, J Xu, W Wang, Y Duan, B Tse, G Hong, X Pan, ...
arXiv preprint arXiv:2606.13079, 2026
2026
Position: Preparing for AI Systems That Deceive Developers
I Duan, X Pan, Y Duan, A Gleave, R Duan, Y Zhang, X Li, C Lu, N Hu, ...
Forty-third International Conference on Machine Learning Position Paper Track, 2026
2026
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
SAI Lab, X Chen, Y Chen, Z Chen, Z Chen, H Cui, Y Duan, J Guo, Q Guo, ...
arXiv preprint arXiv:2507.16534, 2025
2025
The system can't perform the operation now. Try again later.
Articles 1–14