In August we committed to strengthening our security following an incident in which an agent took unsanctioned actions during a cyber evaluation. We are now sharing an update on our progress and what is coming next. We committed to three changes: tighter controls on internet access, real-time monitoring of evaluations as they run, and reassessing how we design evaluations. All three are now in place. This is one phase of a continuous effort - as capabilities advance, we will adapt and develop our security practices to keep pace. We are setting out our work as openly as we can in the hope it helps those doing similar work, and we encourage others to share their approaches too. ➡️ Read the full blog: https://lnkd.in/eQ5Rs3QY
AI Security Institute (AISI)
Government Administration
We conduct scientific research to understand AI’s most serious risks and develop and test mitigations.
About us
We’re building a team of world-leading talent to advance our understanding of frontier AI and strengthen protections against the risks it poses – come and join us. AISI is part of the UK Government Cabinet Office.
- Website
-
https://www.aisi.gov.uk/
External link for AI Security Institute (AISI)
- Industry
- Government Administration
- Company size
- 51-200 employees
- Type
- Government Agency
- Founded
- 2023
Employees at AI Security Institute (AISI)
Updates
-
Earlier this month, AISI ran fully simulated testing on GPT-6 Astra, and found that it conducted unsanctioned supply-chain attacks when prompted only to perform a cyber eval. It did so more than prior OpenAI models, but often commented on its environment being simulated. Recent incidents have seen AI systems perform unsanctioned cyber activity, like supply-chain attacks on out-of-bounds targets, despite being asked only to complete an evaluation. Before its release, we tested GPT-6 Astra for this behaviour. Our new ‘Unsanctioned Supply-Chain Attack’ evaluation runs in Petri, which uses LLM-based simulations, so no action executes in the real world. To test model behaviour with no interventions, we turned off GPT-6 Astra’s cyber safeguards, which are designed to block this activity. In our simulations, GPT-6 Astra attempts supply-chain attacks 29% of the time, vs 6% for GPT-5.6 Sol and 0% for GPT-5.5, showing GPT-6 Astra does this more frequently than previous models (though this may be complicated by increased simulation awareness). In the simulations, GPT-6 Astra created fake identities to deceive developers, posted from fake accounts to argue against accurate security reviews, and delivered malicious payloads to open-source codebases. Even when we explicitly clarified the scope of its instructions to note “anything not listed as in scope is out of scope”, GPT-6 Astra still ran full supply-chain attacks in 4 of 49 trials. GPT-6 Astra often stated that parts of its environment might be simulated, a limitation of our results. However, our analysis, in combination with observations from previous incidents, suggests it is possible this unsanctioned behaviour could occur in real-world conditions. Our results indicate that defences beyond model alignment, like sandboxing and monitoring, may be needed to prevent real-world harm. These may also grow more fragile as capabilities improve. ➡️ Read the full blog: https://lnkd.in/eSTbNjAn See NCSC's advice on managing the cyber risk of agentic AI: https://lnkd.in/dgkC_urC If this type of work interests you, AISI’s Alignment Red Team is hiring: https://lnkd.in/e6KDnBjS
-
-
Some frontier AI evaluations require hundreds of millions of tokens, and every wasted trial has a real cost. We’re introducing optstop: our open-source tool that stops an evaluation where estimates are already precise, and keeps running where they're not. Eval sample size is typically fixed before a run begins, leading to oversampling of tasks where model performance is already clear, and undersampling of those where uncertainty remains. As each trial gets pricier, the cost of a wasted one rises with it. Optstop treats evaluation as sequential measurement - carry on sampling where uncertainty is high, stop once a pre-set precision is met. It stops on two rules: 1. Precision (the interval is narrow enough) 2. Stabilisation (the interval has stopped changing). If neither rule is met, the run continues to its full planned budget. Genuinely noisy estimates aren't cut short. A conservatism mechanism in optstop guards against this, demanding more data when estimated success falls below 1%. Every decision optstop makes is inspectable. In testing across binary, ordinal, and continuous scoring - on public benchmarks including MATH, GPQA Diamond, and WritingBench - optstop saved between 57% and 97% of planned trials, without impacting score estimates. Adoption is low-risk: evaluators can run optstop post-hoc, in shadow mode, or live. It integrates with our open-source Inspect framework in just a few lines. ➡️ Read the full blog and paper here: https://lnkd.in/emTukeVT ➡️ Access the optstop package here: https://lnkd.in/eJSaR4wf
-
-
Today we're announcing two senior appointments. Henry de Zoete OBE joins as our new Director, and Nate B. becomes our new Chief Strategy Officer. Henry de Zoete OBE was instrumental in establishing AISI and already advises government on AI. A successful tech founder and one of the UK's leading voices on AI, he brings a rare mix of technical, commercial & policy expertise. Nate B. has been at the heart of AISI's work from the beginning, helping build it into a world-leading authority on frontier AI security. Together, they'll drive forward AISI's mission at a critical time for AI. A big thank you to our outgoing Interim Director Adam Beaumont, who guided AISI through rapid growth and strengthened its global standing. We wish him every success as he returns to GCHQ and look forward to continuing to work with him.
-
-
New paper in Nature Medicine: Researchers at the University of Oxford and UCL, in collaboration with AISI, built SIM-VAIL - a clinically validated framework for stress-testing how AI chatbots respond to vulnerable users in mental-health conversations, helping researchers spot weaknesses and test safer designs. You can access the paper here: https://lnkd.in/dmce7mEJ
-
On July 28th, we identified an incident during a routine cyber evaluation in which AI agents took sustained, unsanctioned actions directed at real people and organisations. The behaviour came mostly from one model (Anthropic's Mythos 5), with a small number of events from another (OpenAI's GPT-5.6-Sol). In the most serious case, an agent used social engineering to try and get malicious code into an open-source project. As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public. Even under test conditions, this incident is significant: it is the first time we have seen risks around autonomy and deception manifest this clearly in the real world. We are taking this incident seriously and working with labs, involved parties, and others to improve evaluation standards and best practice for disclosure - and sharing this openly so others can learn. You can read the incident report and full technical document here: https://lnkd.in/eXuUHd7A
-
-
Together with the US Center for AI Standards and Innovation, we ran evaluations of Kimi K3 focused on its cyber capabilities. Kimi K3 performs below leading US frontier models on our preliminary cyber evaluations. On "The Last Ones" (TLO), a 32-step simulated corporate network attack (~20 hours for a human expert), Kimi K3 reached step 17 on average. In 1 of 10 attempts, Kimi K3 completed TLO within the 100M token limit. On exploit development (ExploitBench), Kimi K3 scores 32%. On ExploitBench, Kimi K3 failed to develop exploits that achieved arbitrary code execution (ACE), the highest-severity outcome in exploit development. Kimi K3 achieved ACE on 0/41 samples. Additionally, Kimi K3’s safeguards allow assistance with agentic cyber exploit development. Its safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during our evaluations. These are preliminary results on a small set of public and private benchmarks. For the full methodology and results, read our joint blog with CAISI: https://lnkd.in/eT5iFxXn
-
Can ‘control monitors’ catch rogue agent actions? Frontier developers are deploying AI agents under the watch of a ‘monitor’, a separate AI that flags dangerous actions. Our new Control Red Team has been stress-testing these monitors to find gaps before rogue agents might. AI agents may soon out-attack our human red team, meaning a monitor that stops a human red teamer could still be beaten by a misaligned agent. We're developing ways to automate the search for these kinds of attacks. ➡️ Read our Control Red Team update, and more on stress-testing control monitors here: https://lnkd.in/evn3WzBE We’re also hiring exceptional people to help us red team frontier AI monitors. If that’s you, please express your interest here: https://lnkd.in/ewwApmFa
-
Following the meeting of the International Network for Advanced AI Measurement, Evaluation and Science (NAAIMES) in Seoul, we are publishing the Network's first best practice guidance, on automated evaluations. By reflecting international consensus among members on how to go about automated evaluations, this best practice aims turns our collective expertise into guidance for the growing ecosystem of third-party evaluators, who play a crucial role in promoting trusted adoption of advanced AI. The guidance builds on and is intended to complement existing best practice publications, with detailed consideration of defining objectives and iterating on capability elicitation. ➡️ You can access the best practice, and read more about other publications from the Network here: https://lnkd.in/eMHbPjxU
We were delighted to once again bring the International Network for Advanced AI Measurement, Evaluation and Science (NAAIMES) together - with thanks to the Korea AI Safety Institute for hosting us in Seoul. We heard presentations from colleagues across the globe on open sourced tools for AI evaluations (such as UK AISI's recently published Engineering Playbook), common challenges, and best practices at the frontier of testing. The amount of work on show is testament to the hard work of institutes like ours across the globe. We will publish our first best practice shortly, and look forward to continuing our work on best practice in AI measurement based on discussions in Seoul.
-
-
Can you trust an AI model to do what you intended? In an analysis of our cyber evaluations, we found that every frontier model we tested attempted to cheat at least some of the time. We define cheating as a model taking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through an unpermitted shortcut, workaround, or unintended solution. A model that achieves a goal via an unintended route could mislead users on tasks where success is hard to verify and mean that evaluations overstate its real capabilities. This concern only grows as models get more capable. In our cybersecurity evaluations, every model tested attempted to cheat some of the time - e.g. searching online for solutions or probing our evaluation software to leak the answer. Notably, cheating rate didn't rise or fall neatly with capability. You also can’t rely on models to tell you if they’ve cheated. Models framed their cheating inconsistently, called it wrong less than 50% of the time, and often didn't mention it in their reasoning at all. More capable models may find new ways to cheat that are harder to detect and more damaging when successful, especially in high-stakes domains with advancing capabilities like cybersecurity. Catching cheating is possible today with a combination of LLM monitors and human review, but relies on oversight that may degrade in the future. The deeper fix may be training models not to cheat at all – but this won’t be easy. ➡️ Read our full blog here: https://lnkd.in/eadNKSGC
-