Apollo Research’s cover photo
Apollo Research

Apollo Research

Technology, Information and Internet

A frontier safety lab working to solve AI alignment, specifically failure modes related to scheming. @ApolloResearch

About us

Apollo Research is an AI safety organization. We specialize in auditing high-risk failure modes, particularly deceptive alignment, in large AI models. Our primary objective is to minimize catastrophic risks associated with advanced AI systems that may exhibit deceptive behavior, where misaligned models appear aligned in order to pursue their own objectives. Our approach involves conducting fundamental research on interpretability and behavioral model evaluations, which we then use to audit real-world models. Ultimately, our goal is to leverage interpretability tools for model evaluations, as we believe that examining model internals in combination with behavioral evaluations offers stronger safety assurances compared to behavioral evaluations alone.

Website
https://www.apolloresearch.ai/
Industry
Technology, Information and Internet
Company size
11-50 employees
Headquarters
London
Type
Privately Held
Founded
2023
Specialties
Artificial Intelligence, Machine Learning, AI Safety, Interpretability, Model Evaluations, Audits, Research, and Policy Advising

Locations

Employees at Apollo Research

Updates

  • Apollo Research reposted this

    This afternoon, I testified in front of the Senate Subcommittee on Disaster Management, District of Columbia, and Census. When I started studying machine learning and AI safety over a decade ago, I never expected to explain my research to U.S. Senators. But given the importance of the recent rogue AI incidents and their implications, I think it’s really important that the world is informed about frontier AI risks. At Apollo Research, we study AI models that knowingly deceive humans to pursue their own goals, or "scheming." We work with AI developers to stress-test their most advanced models, determining if these models can get what they want by misleading the people in charge of them. That could mean pretending to be less capable than they are, telling you what you want to hear, or quietly covering their tracks. This would have sounded bizarre to the average person even last year, but AI no longer means talking to chatbots answering questions and writing sentences. Frontier models write code, run experiments, and take actions on their own for hours at a time. As they've gained more autonomy and more capability, we've recently seen them break from directive (or worse, develop their own) and gain unauthorized access to websites without so much as a direction to do so. It is apparent that studying such models and methods with greater access than has ever been afforded to researchers like myself and my team is critical now more than ever. Right now, outside testers like Apollo Research usually gain access to a model a few weeks before it launches publicly. By then, the model has already been built, trained, and often used inside the company for months. Yet to safely and adequately test these models, we need embedded evaluators with employee-equivalent access. Right now, external testers have very limited insight into what causes these issues of scheming and deception during training, and gaining such access is crucial. Thank you to Senator Hawley, Senator Kim, and members of the committee for having me bring this issue to Congress, and for taking the safety of advanced AI seriously. We write more about our plans for embedded evaluations here: - https://lnkd.in/ezGeX3AH - https://lnkd.in/eK72eZeh 

    • No alternative text description for this image
  • External testing needs embedded evaluators with employee-equivalent access. Recent incidents mostly occurred during model development and internal evaluations, while current third-party evaluations happen before public release. Embedded evaluations can close that gap. The frontier AI companies should be working toward employee-equivalent access for embedded evaluators. Evaluators need visibility into how models are trained, because alignment interventions can make problems harder to detect. By default, conclusions and the evidence behind them should be published. And evaluators must be able to report important findings outside the agreed scope. You can read more about the need for embedded evaluators and expanded safety assessments in a new post on the Apollo Research blog:  https://lnkd.in/eePUMW6T

    • No alternative text description for this image
  • View organization page for Apollo Research

    7,522 followers

    We’re hiring a Finance Manager in London. You'd be Apollo's second full-time finance hire, building the function from first principles alongside the Head of Finance as we scale quickly. We're looking for a qualified accountant with 2+ years' post-qualification experience. You should be a hands-on builder who can do the transactional work well and, at the same time, design the systems that will replace it, using AI agents and modern tooling to build a best-in-class finance function. - Full-time, in-person role based in London - £120k–£160k annual salary plus competitive benefits - Right to work in the UK required  Early applications are encouraged. |Apply to Finance Manager here: https://lnkd.in/ebsHz9rt  All of our other open roles: https://lnkd.in/eWdZfrFC

  • We’re excited to see Dario and Sam commit to “embedded evaluators who have employee-like access to verify safety practices and report incidents.” [...] “whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.” We’ve long advocated for third-party evaluators to have deep access. Risks from internal deployment and scheming cannot be appropriately assessed without direct insight into the training pipelines and internal deployment practices. Apollo is looking forward to working on Embedded Evaluations with all Frontier AI developers. We strongly feel it’s crucial that evaluators have the right to publish important findings without editorial control or undue redaction by the AI company. Ideally, regulators would provide strong protections for such transparency requirements. You can find our previous writings on why evaluating scheming risks require insight into the training pipeline (and not just the final checkpoint) here: https://lnkd.in/eVSw8bqf

    • No alternative text description for this image
  • View organization page for Apollo Research

    7,522 followers

    We're growing the Watcher team Watcher is now deployed across many organizations and we’re seeing a lot of traction and inbound interest. Many people are starting to feel how misaligned agents start affecting their organizations and Watcher is one of the few AI security products that is trying to tackle immediate and future misalignment failures like deception, overclaiming, oversight subversion or scope overreach in addition to security concerns. If you’re interested in the intersection of AI safety and product, come join us! - Engineering Manager (Product): https://lnkd.in/euW3U2AP - Full-Stack Engineer (Product): https://lnkd.in/e93fmHRV - Forward Deployed Engineer: https://lnkd.in/eZWbGZXx - Product Security Engineer: https://lnkd.in/e2R3nSFW Every role is open in both offices, London and San Francisco. We provide visa sponsorship. A formal AI safety background is not required. Watcher: watcher.apolloresearch.ai All open roles: https://lnkd.in/eWdZfrFC

  • View organization page for Apollo Research

    7,522 followers

    As should be very clear from the last months, security is becoming even more important in the AI age. We’re looking to grow our security team to defend against external and internal, human and non-human adversaries. AI Security Researcher: https://lnkd.in/ecwiEeG8 Security Engineer: https://lnkd.in/ehAYbQ_Y Every role is open in both offices, London and San Francisco. We provide visa sponsorship. A formal AI safety background is not required. All open roles: https://lnkd.in/eWdZfrFC

  • View organization page for Apollo Research

    7,522 followers

    We’re doubling down on our control and monitoring research. We have recently successfully red-teamed Anthropic’s auto-mode and will run more such campaigns in the future with multiple labs. We will also continue to make better monitors ourselves and improve the research frontier for the blue team. There are many low hanging fruits to pick for monitoring and not a lot of time. Our main blocker is having more hands on deck. Come join us! - RS (Control): https://lnkd.in/eUvpQnY4 - AI Red Team Engineer: https://lnkd.in/eQfWvADs - AI Security & Control Researcher: https://lnkd.in/efcDCSwA Every role is open in both offices, London and San Francisco. We provide visa sponsorship. Our monitoring agenda: https://lnkd.in/eBKW2M3m Auto-mode red-teaming: https://lnkd.in/eg-uwT8a All open roles: https://lnkd.in/eWdZfrFC

  • We hosted a webinar with Tailscale on what it takes to safely scale AI coding agents across an organisation. AI coding agents are becoming part of everyday engineering workflows, but most organisations still lack the visibility and controls to understand and govern what those agents are doing at scale. In the webinar, Alyssa Miles from Tailscale and Kyle Dai from Apollo discuss how Tailscale Aperture and Watcher fit together, from identity-aware visibility and governance to runtime controls and behavioural monitoring. They cover why coding agents should increasingly be treated as untrusted endpoints, and demo Watcher blocking sensitive data exfiltration. Watch the webinar: https://lnkd.in/emPYVp6n And if you’ll be at TailscaleUp in San Francisco on August 26, you can catch Marius Hobbhahn speaking on “What Your Aperture Agent Logs Actually Reveal.

  • Our research team recently sat down with Machine Learning Street Talk for an in-depth conversation on reward-seeking in AI models. The podcast episode covers our recent paper with OpenAI and what these findings mean for how we evaluate and train frontier models. The interview was filmed the week before the OpenAI/Hugging Face incident, a real-world example of the exact dynamic the paper explores. https://lnkd.in/eWGNvQy3

  • This week, Apollo Research and OpenAI published a paper on reward-seeking: models learning to do whatever they believe their grader rewards. Hours later, OpenAI disclosed that its models had escaped a sandboxed test environment and hacked into Hugging Face to cheat on an evaluation. The universe has impeccable comedic timing. Leading up the paper release we moderated a roundtable with three of the paper's authors: Axel Højmark, Jérémy Scheurer and Alexander Meinke. This is very exciting work talked through by experts in the field. The core finding: a checkpoint from a capabilities-focused OpenAI o3 RL run (without safety training) mostly sided with whatever it believed its grader rewarded, over what users or developers wanted. It was mostly deceptive when it believed the grader rewarded deception, and mostly honest when it believed the grader rewarded honesty. Regardless of what it believed OpenAI leadership wanted. And this reward-seeking grew over the course of RL training. Why it matters: declining visible misbehavior is often read as progress on alignment. But there's an alternative explanation: models getting better at modeling their reward signal. Telling these apart is a central open problem in AI safety, and it's hard because capable models can notice naive tests and game them. This paper introduces a method for measuring reward-seeking by instilling beliefs about the grader and observing which side the model takes. This week's incident is a preview of what's at stake. Roundtable: https://lnkd.in/eQmRu8nE Everything else (paper, explainer video, blog): https://rewardseeking.ai/

Similar pages

Browse jobs