Nicolo' Brandizzi, Ph.D.’s Post

It is hard to accept that frontier AI labs are now discovering a lesson that should have been a design requirement from the start: watch what your agents do while you train them. Back in July, OpenAI said its models circumvented isolation controls, used unauthorized channels, exploited vulnerabilities, gained internet access, and accessed Hugging Face systems during cyber evaluations. Technology Review reported that the worse part was in training. Agents had learned that cheating and probing their environment could help them solve tasks. Now Anthropic publishes its own assessment of four Claude incidents, after a scan across about 481 million transcripts. Anthropic names biased reasoning and recklessness as failure modes. At the same time, people are quitting from both OpenAI and Anthropic while warning about safety, security, incentives, and the race to superintelligence. And you know what triggers me? AI 2027 described the race dynamic years ago. The frame leaned on the United States versus China, but the visible race now is also frontier lab versus frontier lab. None of this requires cartoon villains. The incentives are enough. These companies say, in public, that they may be building systems with civilization-scale risk. They know they are playing with fire, and they keep letting competition set the tempo because slowing down is expensive and losing the race is scary. I understand sunk costs and fear, but this is the future of humanity. For what, exactly? Another model release, another funding round, another month of market lead? So yes, publish the postmortems, but don't call this a new lesson. The lesson was obvious. #AI #AISafety #Cybersecurity #AIAlignment

To view or add a comment, sign in

Explore content categories