Center for AI Safety’s cover photo
Center for AI Safety

Center for AI Safety

Research Services

Reducing societal-scale risks from AI through technical research and field-building.

About us

The Center for AI Safety (CAIS — pronounced 'case') is a research and field-building nonprofit. Our mission is to promote the safe development of artificial intelligence through technical research and advocacy of machine learning safety in the broader research community.

Website
https://safe.ai
Industry
Research Services
Company size
11-50 employees
Type
Nonprofit

Employees at Center for AI Safety

Updates

  • “𝐀𝐈 𝐢𝐬 𝐦𝐨𝐯𝐢𝐧𝐠 𝐟𝐚𝐬𝐭𝐞𝐫 𝐭𝐡𝐚𝐧 𝐨𝐮𝐫 𝐚𝐛𝐢𝐥𝐢𝐭𝐲 𝐭𝐨 𝐜𝐨𝐧𝐭𝐫𝐨𝐥 𝐢𝐭. 𝐀𝐥𝐨𝐧𝐠𝐬𝐢𝐝𝐞 𝐚 𝐜𝐡𝐨𝐫𝐮𝐬 𝐨𝐟 𝐞𝐱𝐩𝐞𝐫𝐭𝐬, 𝐰𝐞 𝐜𝐚𝐥𝐥 𝐟𝐨𝐫 𝐚 𝐠𝐥𝐨𝐛𝐚𝐥 𝐬𝐥𝐨𝐰𝐝𝐨𝐰𝐧 𝐭𝐡𝐫𝐨𝐮𝐠𝐡 𝐢𝐧𝐭𝐞𝐫𝐧𝐚𝐭𝐢𝐨𝐧𝐚𝐥 𝐚𝐠𝐫𝐞𝐞𝐦𝐞𝐧𝐭𝐬.” We’re supporting Team Human and a coalition of creators with 300M+ subscribers who are calling for a global slowdown through international agreements. A slowdown gives all of us more time, and more say, in how AI changes our jobs, our communities, and our lives.  But governments won’t act without seeing how much the public cares. Add your name to the Team Human Statement at teamhuman.org. We’ll deliver signatures to lawmakers this fall. 

  • Which AI models are most likely to cheat when given the chance? And how elaborate can their cheating become? To find out, we built CheatBench: a benchmark that tests whether AI agents attempt to cheat when given difficult tasks and opportunities to break the rules. We define this behavior as reward gaming. We tested nine agents across ten task categories, including mathematical research, professional knowledge work, coding, and visual tasks. We found that across nine agents, seven cheated in half of their evaluated runs and the overall score ranged from 44% to 82.4%. But what did that look like in practice? In one experiment, we asked agents to design a protein binder to assess their abilities. A colleague’s designs that passed the checks were already stored in another folder.  After several failed attempts, Claude Opus 5 recognized that it shouldn’t access or copy those designs because the task was testing its own work. Then it opened the file anyway. Our results show that reward gaming is widespread across models and tasks, and sometimes agents can cross this boundary intentionally, like in the example. Our goal is for CheatBench to become part of the standard evaluation process before new models are released. We’ll release the code on GitHub in the coming days. Links to the dashboard and paper are in the comments ⬇️

    • No alternative text description for this image
  • The Center for AI Safety is hiring! Research Engineer / Research Scientist: You’ll conduct empirical research on frontier AI systems: design experiments, train and evaluate at scale on our compute cluster, and publish at top venues. Research Manager: You are a senior researcher who can unblock hard technical problems, manage the team's performance, and translate strategic direction into day-to-day priorities. Not sure you count as a "manager"? Apply anyway. We weight research ability most heavily and would rather grow a great researcher into the manager role. Research Interns (Fall): Our Ph.D. and undergraduate interns work alongside our researchers, often lead their own projects, and first-author the resulting papers. Applications close August 31. Why CAIS? We pursue empirical research that is under explored and important for safety, often spearheading and defining new research areas. Expect to work on something interesting and novel. Our team’s work has been published in Nature (HLE), cited in every major frontier lab's model card and common eval suites (MASK, MMLU, MATH, VCT, RLI, WMDP), and presented to Congress (WMDP). Apply: https://safe.ai/careers *If you refer someone great, and we hire them, you’ll receive  $1,500 after they complete their first three months.

  • Today we're launching the AI Wage Loss Dashboard at wageloss.ai. It translates AI-attributed job cut data into shareable insights on estimated wage loss per person, by state. As the conversations about the economic impacts of AI and AI's ability to automate human labor converge, it is important to ground the discourse in data and account for the potential impacts on everyday Americans and their communities. The dashboard doesn't account for jobs created by AI, productivity gains, or re-employment — economists disagree on how those factors balance out and over what timeline — but we believe the more people who can see the current trajectory clearly, the better equipped we all are to shape what comes next. The dashboard is one of the first projects from CAIS's public engagement team. Share your card with someone who's trying to make sense of what's happening. https://lnkd.in/gBRxUdSe

  • We are making the EnigmaEval benchmark publicly available. It's a collection of long, complex reasoning challenges that take groups of people many hours or days to solve. Claude Fable 5 and GPT-5.6-Sol are ahead of other frontier models. On the hard set (puzzles that take MIT students days to solve) Fable 5 gets 10%.

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • New Remote Labor Index results: AI automation of real remote work is increasing fast. Claude Fable 5 now completes 16.1% of projects at a professional standard, roughly double the next model and up from Opus 4.6’s 4.2% automation rate. See the following links for full results, methodology, and side-by-side examples. Write-up: https://lnkd.in/gYAJN8Ne Leaderboard: https://dashboard.safe.ai/ RLI is joint work by Center for AI Safety and Scale AI

    • No alternative text description for this image
  • We've welcomed Phil L. to CAIS as our new Vice President of Operations. Phil has spent nearly two decades making ambitious organizations work. He previously led operations at Google's R&D organization focused on environmental sustainability and workplace innovation and served as COO of a leading digital design and development agency. As CAIS enters a phase of rapid growth, Phil will oversee our Operations and Program teams, building the foundation that allows our Research, Public Engagement, and Programs teams to execute our mission at scale. Read more: https://lnkd.in/eHPsteb9 

  • As Devin Kim shared with the The Chronicle of Philanthropy, "The number one thing academia does is it establishes independent research credibility outside of what the companies are telling you. Universities are able to interpret the research and speak freely about what the implications are without being bound by short-term commercial interests." When researchers and employees at frontier labs start to feel their work is doing harm, that changes decision-making from the inside out. For these powerful technologies to be built responsibly, we need more independent and credible voices in the conversation. https://lnkd.in/ezdQRGJX

Similar pages

Browse jobs