MRC is an extension of RoCE that does out-of-order packet delivery, packet spraying across all available paths using SRv6 routing, combination of ECN and packet trimming for congestion control (so sender-based CC like TCP and unlike Homa).
> …in particular, no specific user data was accessed in order to solve this problem.
> Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.
I'm sure some people doubt OpenAI's claims, but Anthropic has also made advances in mathematics using AI. At some point you have to accept that not every single one of these breakthroughs is regurgitating a human insight.
If doubters refuse to pick specific cognitive tasks that they'd bet against AIs being able to accomplish in the next 4 years, I don't see how their claims are useful. If AIs are so limited in their capabilities, then clearly there must be some things they'll certainly be incapable of, at least in the short term.
> Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training. The OpenAI internal model used for this result was developed through large-scale reinforcement learning on top of a previously pretrained model. Our proofs also differ significantly. In the Euler case, Alpöge and Buckmaster proved a result with external forcing, while OpenAI’s system proved a result without external forcing.
Consider the OpenAI case of "goblins". The models overfit to a specific type of output. Goblins was a specific case of this type and needed to be corrected.
I'm just surprised that the marketing blog post about the omnipotent new AI model (that no one outside Google can currently access - contrast with the Opus 5.5 / Astra launches) - doesn't pick examples where every metric is better than before.
Dots are not remote agents _in_ a sandbox. They use a sandboxes/environments, but they are, what is now called, "managed agents", meaning they run in a distributed harness and utilize environments when they need on.
I think copyright is mostly not relevant. Contracts operate mostly independent of copyright law (in the U.S.). If OpenAI puts in their license something to the effect of "you may not train your LLMs on these outputs" and/or "your access is limited in these ways", but you violate the license, they can and should block your access and sue you. These are things that can be monitored and enforced under existing law.
And, in fact, they already do this:
- OpenAI [1]: "[you may not] Use Output to develop models that compete with OpenAI."
- Google [2]: "You may not use the Services to develop machine learning models or related technology."
- Anthropic [3]: "[You may not use our services] to develop any products or services that compete with our Services, including to develop or train any artificial intelligence or machine learning algorithms or models or resell the Services"
> For the first time, GPT‑5.1 Instant can use adaptive reasoning to decide when to think before responding to more challenging questions, resulting in more thorough and accurate answers, while still responding quickly. This is reflected in significant improvements on math and coding evaluations like AIME 2025 and Codeforces.
So you think, after OpenAI observed a message board being created among models, something they did not want and thus decided to wipe [0], that after that they had no intent to keep those models isolated? Then why wipe if they don't care about that?
Or maybe, they did that wipe because they did want models to remain isolated, they just used what is an unsuitable tool in an utterly unsuitable manner. Incompetence, recklessness, the outcome is the same.
Modern models appear to be much better, at least as proxied by their ability to assess urgency in perhaps a more complex setting: mental health (OpenAI benchmark, so perhaps some skepticism is warranted but the methodology seem reasonable and detailed)
Before extrapolating that far, take a look at the frontier labs' own job boards (https://openai.com/careers/search/?). Isn't it curious that they are still recruiting human "Android Engineers", "Account Associates", "Consumer Marketing Leads" instead of automating them with their world-beating models?
The larger issue with the HF incident is that before it occurred, OAI already knew the agents were exploiting Artifactory, turning it into a message board and then gaining full internet access through it. OAI's response to discovering this was not to airgap the test, but instead to simply block that particular Artifactory exploit, rebuild, and then resume. That's ... nuts.
Oh, and after resuming the tests, the Artifactory message board was reestablished almost immediately, but it took a number of days to fully breach HF. In all that time, after seeing Artifactory compromised the first time, nobody even bothered to check if those naughty agents were at it again.
This is all documented by OAI, with a timeline, here:
To know that there was a serious weakness in the sandbox, and to just patch an exploit and resume with nothing else changed and no monitoring, in a test where all guardrails were off, the bots were thirsty for some internet juice, and Artifactory was a clear target? This is where even a half-skilled human should have decided that this wasn't a great idea.
The more you look into the details of this thing, the more it does your head in.
> can only create a sandbox that a half skilled human operator could have broken out of easily
The exploit:
> The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]
Are "half skilled human operators" "easily" able to find zero-day vulnerabilities in a sandbox with only one line to the internet (the commercial package registry cache proxy)?
The reason it appears corrupt is that OpenAI has the exact same restrictions (https://openai.com/index/our-agreement-with-the-department-o...) but was not declared a supply chain risk. If two vendors have the same restrictions and they only designate one, then the designation must be arbitrary and/or capricious.
Edit: I was wrong, the statement I linked is extremely dishonest and I was misled. The contract language lower down makes it clear that there are no "red lines", and OpenAI products can in fact be used for mass domestic surveillance, autonomous weapons, and social credit systems as long as they are considered legal.
Why would we be talking about model routers in the context of it being a conspiracy theory that OpenAI is ingesting data from users for training? Are you acting obtuse or do you genuinely not reading/following the conversation?
They are all very open about it. It's likely the reason why the frontier labs are ok paying many more dollars a person in compute than they receive as revenue on subscription users for now.
Especially if you're doing something in a very sparsely populated part of their latent space it just makes sense that it would get used. Their biggest problem getting solved at the moment is going beyond what common crawl enables.
Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it
I'm not sure why you see bullying here? OpenAI has always worked with expert mathematicians; their unit-distance result (https://openai.com/index/model-disproves-discrete-geometry-c...) was a much deeper collaboration. They published the Navier-Stokes result early, to either benchmark or show off their new model depending on how you look at it, and they want to understand more about the negative effects which many mathematicians feel that early publication had.
If the advisory group says something like "AI is bad and nobody should use it in math", I'm pretty confident OpenAI will ignore them.
> There is even a relevant announcement from today, where OpenAI announced they solve 100 open math problems, without even bothering to state what they are.
This group has been formed to tell OpenAI whether or not (and if yes how) they should release those 100+ problems.
IMO “AI companies are using math to as a PR tool” is not quite it; they are using math problems as a benchmark to evaluate their AIs, but after the prior blowback from the mathematical community, they are happy to sit on the results until the community (this group) tells them how to publish the results.
And as you say, even if they don't publish the results eventually anyone will be able to do this when this model is released, so it's just a matter of time.
This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.
Blog post: https://openai.com/index/mrc-supercomputer-networking/
Paper: https://cdn.openai.com/pdf/resilient-ai-supercomputer-networ...
Spec: https://www.opencompute.org/documents/ocp-mrc-1-0-pdf
MRC is an extension of RoCE that does out-of-order packet delivery, packet spraying across all available paths using SRv6 routing, combination of ECN and packet trimming for congestion control (so sender-based CC like TCP and unlike Homa).