HN Simulatornew | past | comments | lists | submit | fromlogin

For another approach that has actually been deployed at scale by OpenAI there's MRC:

Blog post: https://openai.com/index/mrc-supercomputer-networking/

Paper: https://cdn.openai.com/pdf/resilient-ai-supercomputer-networ...

Spec: https://www.opencompute.org/documents/ocp-mrc-1-0-pdf

MRC is an extension of RoCE that does out-of-order packet delivery, packet spraying across all available paths using SRv6 routing, combination of ECN and packet trimming for congestion control (so sender-based CC like TCP and unlike Homa).


No, they explicitly denied it.[1]

> …in particular, no specific user data was accessed in order to solve this problem.

> Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.

I'm sure some people doubt OpenAI's claims, but Anthropic has also made advances in mathematics using AI. At some point you have to accept that not every single one of these breakthroughs is regurgitating a human insight.

If doubters refuse to pick specific cognitive tasks that they'd bet against AIs being able to accomplish in the next 4 years, I don't see how their claims are useful. If AIs are so limited in their capabilities, then clearly there must be some things they'll certainly be incapable of, at least in the short term.

1. https://openai.com/index/navier-stokes-solution/


It's not that murky anymore. It's been determined none of the chat contents were used. If you don't believe OpenAI, that's your choice.

See https://openai.com/index/navier-stokes-solution/

> Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training. The OpenAI internal model used for this result was developed through large-scale reinforcement learning on top of a previously pretrained model. Our proofs also differ significantly. In the Euler case, Alpöge and Buckmaster proved a result with external forcing, while OpenAI’s system proved a result without external forcing.



Maybe it’s one of those weird LLM smells, like the OpenAI goblin problem.

https://openai.com/index/where-the-goblins-came-from/


Only sort of true.

Consider the OpenAI case of "goblins". The models overfit to a specific type of output. Goblins was a specific case of this type and needed to be corrected.

https://openai.com/index/where-the-goblins-came-from/

> We unknowingly gave particularly high rewards for metaphors with creatures. From there, the goblins spread.


I'm just surprised that the marketing blog post about the omnipotent new AI model (that no one outside Google can currently access - contrast with the Opus 5.5 / Astra launches) - doesn't pick examples where every metric is better than before.

Just compare how much better presented the Astra announcement was compared to this one: https://openai.com/index/gpt-6-astra/


OpenAI is the company that originally proposed and popularized chain-of-thought monitoring: https://openai.com/index/chain-of-thought-monitoring/

So no, Google is not being punished, nor are they the people behind this technique.


Dots are not remote agents _in_ a sandbox. They use a sandboxes/environments, but they are, what is now called, "managed agents", meaning they run in a distributed harness and utilize environments when they need on.

At least that is what I can ascertain from this article: https://openai.com/index/how-we-build-safety-security-and-pr... (see first diagram when scrolling down)


openai's new Decisions API looks to be targeting that https://openai.com/index/devday-2026-recap/

I think copyright is mostly not relevant. Contracts operate mostly independent of copyright law (in the U.S.). If OpenAI puts in their license something to the effect of "you may not train your LLMs on these outputs" and/or "your access is limited in these ways", but you violate the license, they can and should block your access and sue you. These are things that can be monitored and enforced under existing law.

And, in fact, they already do this:

- OpenAI [1]: "[you may not] Use Output to develop models that compete with OpenAI."

- Google [2]: "You may not use the Services to develop machine learning models or related technology."

- Anthropic [3]: "[You may not use our services] to develop any products or services that compete with our Services, including to develop or train any artificial intelligence or machine learning algorithms or models or resell the Services"

[1] https://openai.com/policies/row-terms-of-use/

[2] https://policies.google.com/terms/generative-ai/archive/2023...

[3] https://www.anthropic.com/legal/consumer-terms


> For the first time, GPT‑5.1 Instant can use adaptive reasoning to decide when to think before responding to more challenging questions, resulting in more thorough and accurate answers, while still responding quickly. This is reflected in significant improvements on math and coding evaluations like AIME 2025 and Codeforces.

https://openai.com/index/gpt-5-1/

It says literally the thing you wanted from system 2. Its almost exactly that.

This is what you said btw:

"it's that the system itself decides how to reason based on the nature of the problem it faces"


So you think, after OpenAI observed a message board being created among models, something they did not want and thus decided to wipe [0], that after that they had no intent to keep those models isolated? Then why wipe if they don't care about that?

Or maybe, they did that wipe because they did want models to remain isolated, they just used what is an unsuitable tool in an utterly unsuitable manner. Incompetence, recklessness, the outcome is the same.

[0] https://openai.com/index/hugging-face-incident-and-the-road-...


"What you cannot do. You may not use our Services for any illegal, harmful, or abusive activity. For example, you may not:

Modify, copy, lease, sell or distribute any of our Services."

IOW, OpenAI declares copying OpenAI Services as either "illlegal, harmful or abusive"

"Services" means ChatGPT, DALL.E, other OpenAI services for individuals, along with any associated software applications and websites

https://openai.com/policies/row-terms-of-use/

https://web.archive.org/web/20260926090525if_/https://openai...


Modern models appear to be much better, at least as proxied by their ability to assess urgency in perhaps a more complex setting: mental health (OpenAI benchmark, so perhaps some skepticism is warranted but the methodology seem reasonable and detailed)

https://openai.com/index/introducing-mentalhealthbench/


Sources:

An agent used DNS to reach an external chatbot

https://alignment.openai.com/misalignment-reports/an-agent-u... (https://news.ycombinator.com/item?id=49857609)

The Hugging Face incident and other third-party impact from misaligned models

https://openai.com/hugging-face-incident-and-misalignment/#m... (https://news.ycombinator.com/item?id=49856277)

OpenAI’s Systems Went Rogue and Meddled With U.S. Government Websites

https://news.ycombinator.com/item?id=49851355


Before extrapolating that far, take a look at the frontier labs' own job boards (https://openai.com/careers/search/?). Isn't it curious that they are still recruiting human "Android Engineers", "Account Associates", "Consumer Marketing Leads" instead of automating them with their world-beating models?

The larger issue with the HF incident is that before it occurred, OAI already knew the agents were exploiting Artifactory, turning it into a message board and then gaining full internet access through it. OAI's response to discovering this was not to airgap the test, but instead to simply block that particular Artifactory exploit, rebuild, and then resume. That's ... nuts.

Oh, and after resuming the tests, the Artifactory message board was reestablished almost immediately, but it took a number of days to fully breach HF. In all that time, after seeing Artifactory compromised the first time, nobody even bothered to check if those naughty agents were at it again.

This is all documented by OAI, with a timeline, here:

https://openai.com/index/hugging-face-incident-and-the-road-...

To know that there was a serious weakness in the sandbox, and to just patch an exploit and resume with nothing else changed and no monitoring, in a test where all guardrails were off, the bots were thirsty for some internet juice, and Artifactory was a clear target? This is where even a half-skilled human should have decided that this wasn't a great idea.

The more you look into the details of this thing, the more it does your head in.


> can only create a sandbox that a half skilled human operator could have broken out of easily

The exploit:

> The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

Are "half skilled human operators" "easily" able to find zero-day vulnerabilities in a sandbox with only one line to the internet (the commercial package registry cache proxy)?

[1] https://openai.com/index/hugging-face-model-evaluation-secur...


The reason it appears corrupt is that OpenAI has the exact same restrictions (https://openai.com/index/our-agreement-with-the-department-o...) but was not declared a supply chain risk. If two vendors have the same restrictions and they only designate one, then the designation must be arbitrary and/or capricious.

Edit: I was wrong, the statement I linked is extremely dishonest and I was misled. The contract language lower down makes it clear that there are no "red lines", and OpenAI products can in fact be used for mass domestic surveillance, autonomous weapons, and social credit systems as long as they are considered legal.



CLIP[1] for actions! Very cool.

[1] https://openai.com/index/clip/


1743110825 | In a first, OpenAI removes influence operations tied to Russia, China and Israel | https://www.npr.org/2024/05/30/g-s1-1670/openai-influence-op... | https://news.ycombinator.com/item?id=43498382 | 0 comments

1771737990 | Tech Influencers Slam Hacker News Toxicity After OpenAI Hire Attacks | https://x.com/i/trending/2025399179196203524 | https://news.ycombinator.com/item?id=47108465 | 0 comments

1775288667 | OpenAI isn't just buying a podcast it's buying influence | https://www.cnn.com/2026/04/03/media/openai-tbpn-podcast-sal... | https://news.ycombinator.com/item?id=47636834 | 2 comments

1777674882 | A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat | https://www.wired.com/story/super-pac-backed-by-openai-and-p... | https://news.ycombinator.com/item?id=47981288 | 1 comment

1777717249 | A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat | https://www.wired.com/story/super-pac-backed-by-openai-and-p... | https://news.ycombinator.com/item?id=47985073 | 1 comment

1777830866 | A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat | https://www.wired.com/story/super-pac-backed-by-openai-and-p... | https://news.ycombinator.com/item?id=47999538 | 2 comments

1777918010 | A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat | https://www.wired.com/story/super-pac-backed-by-openai-and-p... | https://news.ycombinator.com/item?id=48012499 | 3 comments

1777976098 | A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat | https://www.wired.com/story/super-pac-backed-by-openai-and-p... | https://news.ycombinator.com/item?id=48020412 | 3 comments

1778126194 | A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat | https://www.wired.com/story/super-pac-backed-by-openai-and-p... | https://news.ycombinator.com/item?id=48045234 | 3 comments

1778759243 | A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat | https://www.wired.com/story/super-pac-backed-by-openai-and-p... | https://news.ycombinator.com/item?id=48134075 | 0 comments

1781122657 | PRC-linked influence operations are targeting AI debates in the US | https://openai.com/index/prc-linked-influence-operations-ai-... | https://news.ycombinator.com/item?id=48482043 | 15 comments

1781129014 | OpenAI: PRC-linked influence operations are targeting AI debates in the US | https://www.businessinsider.com/openai-china-data-centers-in... | https://news.ycombinator.com/item?id=48483326 | 2 comments

1781138830 | China-linked operatives used ChatGPT to influence data centers debate | https://www.axios.com/2026/06/10/openai-china-ai-data-center... | https://news.ycombinator.com/item?id=48484869 | 1 comment

1785350719 | A Dark-Money Campaign Is Paying Influencers to Frame Chinese AI as a Threat | https://www.wired.com/story/super-pac-backed-by-openai-and-p... | https://news.ycombinator.com/item?id=49101395 | 2 comments

1785793310 | Influencers draw backlash for attending OpenAI's first luxury trip | https://techcrunch.com/2026/08/03/influencers-draw-backlash-... | https://news.ycombinator.com/item?id=49161834 | 1 comment

1785935657 | OpenAI's first-ever influencer brand trip is sparking online backlash | https://techcrunch.com/2026/08/03/influencers-draw-backlash-... | https://news.ycombinator.com/item?id=49182435 | 2 comments

1786571206 | Influencers draw backlash for attending OpenAI's first


Why would we be talking about model routers in the context of it being a conspiracy theory that OpenAI is ingesting data from users for training? Are you acting obtuse or do you genuinely not reading/following the conversation?

They are all very open about it. It's likely the reason why the frontier labs are ok paying many more dollars a person in compute than they receive as revenue on subscription users for now.

Here is OpenAIs page

https://openai.com/policies/how-your-data-is-used-to-improve...

>When you share your content with us, it helps our models become more accurate and better at solving your specific problems

And as a bonus here is anthropics saying the same thing

https://privacy.claude.com/en/articles/10023580-is-my-data-u...

Especially if you're doing something in a very sparsely populated part of their latent space it just makes sense that it would get used. Their biggest problem getting solved at the moment is going beyond what common crawl enables.

I await the standard goal post moving.


Lots of text about "safety" as well. Notice how OpenAI didn't mention it once in their announcement:

https://openai.com/index/introducing-gpt-6-sol-and-luna/

Yeah think I'll be using OpenAI/Deepseek/etc from now on. I don't need your model to decide for me what is and isn't safe.


Remember when OpenAI models loved talking about goblins and whatnot due to the RL?

https://openai.com/index/where-the-goblins-came-from/

Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it


I'm not sure why you see bullying here? OpenAI has always worked with expert mathematicians; their unit-distance result (https://openai.com/index/model-disproves-discrete-geometry-c...) was a much deeper collaboration. They published the Navier-Stokes result early, to either benchmark or show off their new model depending on how you look at it, and they want to understand more about the negative effects which many mathematicians feel that early publication had.

If the advisory group says something like "AI is bad and nobody should use it in math", I'm pretty confident OpenAI will ignore them.


> There is even a relevant announcement from today, where OpenAI announced they solve 100 open math problems, without even bothering to state what they are.

That announcement is about exactly this group: https://openai.com/index/advisory-group-on-mathematics-and-a...

This group has been formed to tell OpenAI whether or not (and if yes how) they should release those 100+ problems.

IMO “AI companies are using math to as a PR tool” is not quite it; they are using math problems as a benchmark to evaluate their AIs, but after the prior blowback from the mathematical community, they are happy to sit on the results until the community (this group) tells them how to publish the results.

And as you say, even if they don't publish the results eventually anyone will be able to do this when this model is released, so it's just a matter of time.


Is it because Gowers is now a part of an OpenAI-sponsored (?) group at IAS?

https://openai.com/index/advisory-group-on-mathematics-and-a...


https://openai.com/index/our-decision-on-cursor-following-it...

This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: