When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently...
This work was done as part of the Second Look Fellowship and mentored by Uzay Macar. I'm immensely grateful for the multiple rounds of feedback and support given by Yixiong and Zephaniah Roe for my work. I'm also very thankful for the valuable insights shared by Yujin Potter and Yao...
tl;dr: Some important AI safety research is never rerun on the newest models. There are probably cases where this would be valuable and a single well-positioned researcher could likely do this with sufficient funding. This summer, Second Look Research (SLR) is running a summer fellowship dedicated to empirical replications of...
tl;dr: * Lindsey 2025 found models can modulate their internal states: when instructed to “think about” a concept while writing an unrelated sentence, the representation of the concept is more present than when instructed to not think about it. * Internal state controllability appears to be a general property of...
This work was done as part of the Second Look Fellowship by Arav Dhoot and supervised by Yixiong Hao and Zephaniah Roe. I'm grateful to Harshul Basava and Vanessa Ng for their feedback. This is an extension to a prior replication which can be found here. Introduction and Motivation In...
Summary * AY 2025-26 was an outlier year for Georgia Tech’s AI Safety Initiative (AISI), with 15+ members placed in AI safety roles. In this post, we distill our most important advice for other university groups. * Key takeaways: * Deliberately identify potential talent in the fellowship, heavily invest high-context...
This replication was done as part of the Second Look Fellowship by Arav Dhoot and supervised by Yixiong Hao and Zephaniah Roe. I am grateful to Andy Wang for their feedback. My code can be found here. > "So the answer should be A - at active promoters and enhancers."...