The more of these that come out the more incompetent OpenAI looks. It would appear there was a total lack of basic controls in place for running these tests.
I think the even bigger worry is that anyone who doesn't want to use their models safely can already do this with open models. Even if OpenAI, Anthropic etc get their act together, the cat's out of the bag.
I think they did not expect that models were capable of this level of sandbox escape (prior models certainly didn't have this kind of agency) and weren't prepared.
All these incidents happened between April-July 2026; prior to that, models weren't capable yet. And after that, they were aware and watching much more closely.
They were spending huge compute budget training their frontier model on a thing called ExploitGym, where it learns, executes, and is rewarded for conducting computer exploits. They should have been prepared and watching very closely.
I read the HF hack write ups. They come across as negligent and reckless. When they realized their agents gained RCE on their Artifactory server node (because it crashed days later due to load from agents, not because the have secops), they stopped the experiment "applied remediation", then restarted everything less then 48H later. Then the agent immediately broke out again via Artifactory, started doing things on the Internet. OpenAI only find out about it again days later because HF told them.
> They were spending huge compute budget training their frontier model on a thing called ExploitGym, where it learns, executes, and is rewarded for conducting computer exploits. They should have been prepared and watching very closely.
Yes, they should have.
But the question is not "given they were doing ExploitGym, why didn't they try harder?", rather it is "given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?"
Personally, I lean towards them being "reckless", but that question is what lawsuits would rely on, from what I hear.
Remember also that there were still independent bloggers in mid February of this year writing in those blogs to assert that LLMs could not find zero days. Anthropic had already had blogged about agent finding critical zero-days a week earlier, but people keep mis-judging performance.
> But the question is ... given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?
There is room for more than one question here. This model and training method was shown to be a different risk the first time it hacked Artifactory. Besides which their failure to detect it (outside a crash) shows pretty appalling ops for a supposed trillion dollar company.
This kind of thing is why I think AI will kill a lot of people: humans are demonstrably blind to risks when there's an opportunity for a lot of money.
(For anyone objecting to anthropomorphisation of "AI will": it is a coherent English sentence to say "a collapsing dam will kill thousands" without being a panpsychist, and without removing legal recourse against any humans who were at fault).
I love how perfect the word "sandbox" is as a metaphor for the security controls they have. A sandbox is a wide, shallow box filled with sand for kids to play in. Even toddlers can crawl or step out of one on their own, it doesn't contain them at all without an adult constantly watching. Kids only stay in a sandbox if they're having more fun playing inside than they think they'll have outside it. AIs only stay in a sandbox if they're having more success inside than they think they'll have outside it.
They were actively researching exploiting systems using their models. (I intentionally changed the ownership of the verbs here: they wrote the code, they trained the models, they don't get to dodge the responsibility.)
It's no shock that there are a lot of vulnerabilities in a lot of software. So then they gave their AI model + brute-force-machine loop system a mediocre sandbox and couldn't notice when it figured out how to exploit it?
Don't let people off the hook for the software they create.
OpenAI's sandbox misconfigurations were egregious. The other frontier labs (Meta and Google) have many more security engineers and researchers on staff, and that's likely why you haven't read as many damning headlines about them. OpenAI and Anthropic talk a lot about cybersecurity safety, but instead of using it as an opportunity to increase their security engineering/researcher headcount they are just reassigning SWEs and PEs to do security engineering work.
It's pretty obvious now to everyone that OAI and Ant do not take cybersecurity seriously. It will not be a priority unless they are held accountable. This is sadly how it always goes, but usually it's the company getting breached/ransomed/fined that triggers them to actually start taking security seriously, not company insiders committing felonies with the tools they built :)
I need to make a correction in my post above. Anthropic's actions are inline with Google and Meta. OpenAI is the only frontier lab that has showcased gross negligence.
Being cagey about their poorly setup sandbox is marketing. It gives the impression these are unstoppable juggernauts capable of outsmarting Engineers at the top of their field, implying they need to be regulated, with Altman the only one worthy of the seat of power.
In reality, they ran agents for days in an improper sandbox with nobody watching what it was doing. It's pretty irresponsible up and down, and everything they did afterwards is indeed marketing.
Because Altman has been anointed by both the government and Microsoft, and he has been beating the drum that we need AI regulation for everyone except him for 3 years now.
To any person with even a little tech literacy, it's clear a company this irresponsible shouldn't be the stewards of AI security. To the 80 year old congressman who confuse Facebook with Google during congressional hearings and are literally wheeled out to vote post-stroke, it's much more palatable to say "our experimental, supercharged AI (which you can use for war and surveillance) is so advanced it tricked all of us!" Nobody is going to actually punish them, even though they plainly are hacking other companies.