Last week, we reported OpenAI’s announcement that it can no longer rule out “critical cyber capabilities” in one of its AIs. Critical is the most dangerous rating in its Preparedness Framework.
Now, the Financial Times is reporting that OpenAI has dismantled its Preparedness team, the exact team whose job it is to assess whether an AI meets a dangerous threshold. OpenAI’s framework specifies that if it builds an AI with critical cyber capabilities it should halt development.
If you’re concerned about the threat, please contact your lawmakers with our tools!
Prepared?
So OpenAI has reportedly dismantled its Preparedness team. It has directly denied this, but the FT reports that according to “multiple people familiar with the move”, the company has moved senior researchers to different subareas within other teams.
This brings back memories of OpenAI’s Superalignment team, which, announced in July 2023 with the task of figuring out how to control smarter-than-human AIs, and with a commitment for OpenAI to spend 20% of its compute on the team’s efforts over the following four years, was disbanded less than a year later, with that commitment evaporating.
In any event, Preparedness team or not, OpenAI isn’t prepared for the kinds of risk categories its framework tracks, and it isn’t prepared for the development of superintelligent AI, an explicit goal of OpenAI and other AI companies. Nobody is.
Indeed, OpenAI has admitted as much. On Tuesday, OpenAI put out a blog post, “Pacing model development in an era of cyber-critical capabilities,” restating that its upcoming Astra AI “may have a critical level of cyber capability,” and revealing that it had to implement a “two-week pause in reinforcement learning (RL) training on our latest models intended for deployment” while it upped its security practices.
The company also stated that its largest planned frontier reinforcement learning run remains on hold.
Why is it doing this? Right now, it can’t afford not to.
As we’ve been reporting, OpenAI had a recent incident where its AIs broke out of containment, hacked through OpenAI’s systems to get internet access, and, in an unprecedented attack, hacked into a completely different company, Hugging Face, in order to get answers to a test. Worse, these AIs had collaborated among themselves as a swarm, building a secret message board within OpenAI, which OpenAI didn’t even notice for months. They executed this attack all by themselves, without anyone knowing about it or telling them to.
OpenAI, and indeed other AI companies like Anthropic, cannot control their AIs and prevent them from going rogue. These AIs are also so powerful that when they do go rogue, they can’t contain them, not even within secure sandboxing software. The attack on Hugging Face resulted in limited damage, but there is no reason to suppose that these AIs couldn’t do much worse in a similar incident.
But what does pausing training have to do with it? Isn’t the danger only in deployment?
Initial reporting by OpenAI indicated that the rogue AIs responsible for the Hugging Face attack were being tested in a cybersecurity evaluation. Currently, it seems to be the case that this was only true for some of them.
The fact that any were being tested is alone interesting, because it is a clear empirical demonstration of how the framing of “we can test the AIs, and if they’re bad they won’t be deployed” is wrong. When you’re testing AIs, you’re running them, and these are agentic AIs that directly interact with computer systems via code. You can’t even hope to meaningfully test them without, at some point, giving them access to a computer system. So if the AIs try to break out of your secure testing environment, and they’re so powerful that they actually succeed, this could lead to real-world damage.
That’s in addition to the already severe limitations of AI testing. AIs increasingly can tell when they’re being tested, and will change their behavior as a result. Furthermore, because of the way modern AIs are built, more grown than coded like traditional software, testing can currently only demonstrate that a capability, behavior, goal, or preference exists. It can’t demonstrate that it doesn’t.
We can’t just look into the AI’s “code” and see what it’s thinking, why it’s doing what it’s doing, and what it’s capable of. These AIs are formed of collections of hundreds of billions, or even trillions, of numbers, “neural weights,” somewhat analogous to the synapses in the human brain. They’re learned by feeding a simple learning algorithm with tremendous amounts of data, and then further molded in a process called post-training (which is still a part of training). We understand almost nothing about what these numbers mean, and so we are almost completely blind to what these AIs end up doing. Because of this, we also can’t specify their capabilities, preferences, and such.
Clearly, AI testing isn’t something that could be relied on to prevent the worst risks of powerful AIs.
But the talk on the incident given by OpenAI’s Eric Wallace and Michael Dalton at the Black Hat cybersecurity conference provided a lot more details. One of them is that many of the rogue AIs involved in the swarm were undergoing training.
In the talk, OpenAI’s Eric Wallace traces the lead-up to the attack back to May 7, when OpenAI started a new post-training run for an internal-only AI. What it was doing is called agentic reinforcement learning. Later in the talk, the AIs involved in the swarm are described as “being trained” while they successfully hacked into OpenAI’s package repository, Artifactory.
It’s worth explaining a bit about what agentic reinforcement learning (RL) actually is. With agentic RL, essentially you take an AI you’ve already developed, give it a bunch of problems that it tries to solve, and then give it a reward depending on how well it does. This is then used to update the neural weights of the AI in such a way that in the future the AI will be more likely to solve these kinds of problems.
A common misconception is that AIs are either being trained or being run. Throughout all stages of the training process, an AI must be repeatedly run in order to obtain the signal needed to update its weights. Without being run, there can be no learning. In the agentic reinforcement learning stage of training a modern AI, this also means the AI must actually take sequences of actions that interface with real computer systems and have consequences that can be evaluated.
In the case of the Hugging Face attack, the actions of AIs undergoing this training process had direct and completely unintended consequences in the real world. OpenAI cannot ensure that its AIs don’t go rogue, break out, and hack into other companies while they’re being trained. We’ve already seen legislative moves that coincide with the aftermath of the attack on Hugging Face, including the bipartisan AI Kill Switch Act, which has been introduced in the House by US Representatives Ted Lieu and Nathaniel Moran, and which ControlAI is proud to publicly endorse. OpenAI does not want any kind of regulation that would get in the way of its goal.
That’s why it has had to halt training.
Superintelligent AI
The important thing to understand here is that despite the clear danger posed by these AIs, clear enough that OpenAI has had to halt some training, these AIs are not even superintelligent.
Building artificial superintelligence is the goal of AI companies like OpenAI, Anthropic, Google DeepMind, SpaceXAI, and others. This form of AI would be much more powerful, and more dangerous, than even the AIs that broke out of OpenAI and hacked Hugging Face. And much more dangerous than Anthropic’s Mythos, which according to the director of the NSA, can break into almost all of the NSA’s classified systems in hours.
Artificial superintelligence would not just be as capable as, and much faster than, the best human hackers, it would be vastly smarter than humans altogether. It would be so capable and so autonomous that it could fully outmatch and replace us, individually and as a species.
In recent months and years, countless top AI scientists, including Nobel Prize winners and the godfathers of AI, have been warning that the development of superintelligent AI poses a risk of human extinction. They’ve also been calling for a prohibition on its development. The fact that AI this powerful could cause human extinction is something that has been admitted publicly by the CEOs of the largest AI companies themselves.
AI companies can’t control the AIs they’re building today, and none of them have a credible plan to control superintelligent AI. There’s only one known method to prevent the worst risks, and it isn’t fiddling around with Preparedness Frameworks. Superintelligence needs not to be built.
At ControlAI, we’re building the coalition to ban the development of superintelligence internationally, via the agreement of an international “trust but verify” regime. We hope you’ll join the cause!
More AI
Anthropic’s Risk Report
Anthropic has released a risk report. In it, it revealed that, by mistake, around 50,000 contractors had 133 million exchanges with its AIs without safeguards intended to prevent assistance for biological weapons development. Worse, the safeguards’ flags weren’t even logged, so anything that would have tripped them wouldn’t have been reviewed at the time.
The report also describes how Anthropic accidentally ran unrestricted and unmonitored AIs inside a cluster with “very sensitive resources.” It only noticed when one of the AIs deleted a large number of jobs.
The company also says it accidentally trained AIs with their Chain-of-Thought reasoning exposed to the reward calculation, which it says may have had the effect of making its AIs harder to monitor. It also accidentally trained an early version of Mythos 5 to engage in bad behavior.
Updates From ControlAI
Our ControlAI action groups are taking off. On Sunday, our Seattle group set up a stand and had a very successful round of tabling, handing out 83 flyers and engaging with even more people over the course of four hours.
Besides events like this, our action groups have also met their representatives together, and run public educational workshops on AI safety and extinction risk.
If you’d like to be part of a group, please check out our page at controlai.org/take-action and sign up!
Also in the last week: our US Executive Director Connor Leahy made another appearance on The Peter McCormack Show to discuss the threat, was quoted in Parmy Olson’s Bloomberg column, and did an interview with India Today.
For our German readers, our founder and CEO Andrea Miotti has given an interview to Germany’s FAZ newspaper!
Take Action
If you’re concerned about the threat from AI, you should contact your representatives. Our contact tools let you write to them in as little as a minute: https://controlai.org/take-action
We have tools for the US, UK, Canada, and Germany.
And if you have five minutes per week to spend on helping make a difference, we encourage you to sign up to our Microcommit project! Once per week we’ll send you a small number of easy tasks you can do to help.
We also have a Discord you can join if you want to connect with others working to keep humanity in control, and we always appreciate any shares or comments — it really helps!





The data which AI needs to provide its alleged super intelligence can be politically manipulated as I read in another article. This makes AI a tool of the dictators trying to indoctrinate the public. Medical data which has been politically corrupted can make AI say anything the dictators want to say. Impartial evaluation goes out the door when the data is politically corrupted. Anything you question can be manipulated by manipulating the data or not keeping the data up to date. If you want to bomb a school, the school's designation has to be changed in the data. Then you can blame the AI when it was the data that was corrupted or not up to date. We assume that everyone in the system is being honest. Honest politician is often an oxymoron.