These mathematicians’ prompts are not like “hey chat, please solve Navier-Stokes for me”. They add real expertise and intuition from the cutting edge of their field.
Recklessly prompting OpenAI without a care to the safety of their knowledge.
And after that trying to cast aspersions at OpenAI?
Hopefully we get some better facts, because OpenAI are disliked enough that a smear campaign could work against them.
Edit: also the narritive is getting framed as OpenAI versus Anthropic. A highly political extremely capitalist fight is going on, and facts are victims.
If the new model is that good, and is chewing through open problems at an unprecedented rate, the smart move would have been to let the humans have their W on this one and present solutions to those other problems.
Especially if there really is a long list of them.
"Here are a few hundred proofs" is far more convincing than "We really Navier Stokes and coincidentally someone else did too but we don't know the details or anything, who us, definitely not."
It's a PR fiasco, and a cynic might wonder if it's entirely about the IPO.
I'm consistently entertained by how these companies, with the most advanced models on the planet, consistently do the most idiotic things.
It seems a perfectly reasonable possibility that Navier-Stokes is just the most easily solvable of the remaining problems, and that their new model is capable of solving it while not being capable of solving the others.
There's many cases of researchers racing to solve various problems after hearing that others are working on them. I don't think anyone's suggesting that it was a coincidence at all. In fact, OpenAI freely admits that they started working on the problem after hearing rumours that others were close to solving it. To me, that's not evidence of "cheating" in any way.
Is it really such an extraordinary claim to say that they could have solved the problem without copying Buckmaster and Alpoge? It seems very much in the realm of possibilities.
To me, it seems just as extraordinary to claim that they did "cheat". If I were a betting man, I would put the odds around 50/50 from everything I've read on the subject.
But my point is that everyone seems to be presuming guilt.
It is an extraordinary claim, it is a millenium prize problem after all. We don't even know, even if there was no copying, how much human involvement there was in the result.
If the model can understand neuralese why can it not convert it into English for monitoring or review purposes?
Or we believe the model will encode secret messages like "don't reveal this information" into the neuralese. But as the author mentions, they could have been doing that all along
> Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography
> If the model can understand neuralese why can it not convert it into English for monitoring or review purposes?
A model doesn't really "understand" neuralese, in the same way that the human brain doesn't intuitively understand the low level processes that compose a thought.
Even if we could trace all the electrical and chemical activity behind a human thought, we (likely) couldn't directly translate that activity into its meaning, because internal representations don't map neatly to intuitive concepts. There (usually) isn't a single neuron for "apple" and another for "eating" so that connecting the two forms the thought "eating an apple".
Having said that, there have been experiments on LLM that have managed to identify and even modify internal representations. However, these methods are still computationally expensive and limited.
> Models can omit key information in their visible thoughts, as this Anthropic 2025 paper shows. We are also worried about steganography
I think the analogy with the human brain is very fitting. If we train somebody to perform an action, they'll be able to do it, but we can't know for certain whether they internally agree with it or not.
High Anubis difficulty is annoying the hell out of me for several sites. And it's starting to not block LLM bots anymore?
> 33% are now solving the math and getting through to the main site — because apparently what we have to offer is worth spending a ton of cycles to calculate the Anubis challenge.
Has anyone considered having Anubis perform more valuable hashing?
Like, maybe you can't stop the LLM bots, but you can use them as one-off Bitcoin pool mining pool participants. You have to assume that making them find hash values with N leading zeroes has led to finding hash values with more than N leading zeroes. Maybe run a Bitcoin node under there and let each visitor take a couple swings for you with their pickaxes.
Visions of vast data centers surrounded by fields of browning grass, in which aging, rusting, formerly extremely expensive hardware is spending billions of compute cycles looking at anime catgirls
It's very likely the last few years of bot behavior is the consequence of the residential proxy business booming. This is indirectly due to AI company crawling, but the fact that they are as cheap and available as they are changes the incentives for anyone using them toward reckless and unsustainable request behavior, as there is no risk of burning your IPs, and very small chances of seeing any consequences of essentially DDoS:ing a website.
And the residential proxy business was created by Cloudflare, who was created by us using Cloudflare.
I've been on all three sides (user of RPs, getting paid to run an RP, and trying to block RPs from my site). Residential proxy service is nice. You can scrape anything, even with the dumbest curl command, and only get a Cloudflare block maybe 15% of the time, in which case you just try again. That's less often than I get a cloudflare block from using a privacy browser from a non-proxy address. Cloudflare does not stop bots, it stops humans.
Wait, so you’ve been the person who wanted to keep people from scraping your site, the person who’s trying to scrape your site, and the person getting paid to help someone scrape your site? Brother, what are you doing with your life?
Contingency exists, I get that, and if you're truly in that state, sure, I don't judge necessity. A lot of our cohort seems to confuse a studied disinterest in looking beyond the end of their own nose for the sage wisdom of adulthood, though.
How will you know anything about a system you stubbornly only look at one side of? Like saying planes are terrible because they're loud. They are, but have you never traveled to a different city?
Unfortunately we're apt to run into some kind of Jeavons Paradox where the hardware gets so much faster in those few years will be able to slurp massive amounts of data cheaply so the problem never really ends.
> High Anubis difficulty is annoying the hell out of me for several sites.
I just close the website if I see Anubis. Some have it set at reasonable difficulties (like 2)… others have it where I need to wait for like 30 seconds, I'm not wasting 30 seconds of my life for that.
the brutal truth is that if the website operator simply disabled Anubis, your page load would likely take more than 30s.
when a system was designed for 100 req/s and bots hit it with 5000 req/s, nobody entering that queue is having a good time. Anubis is the trade those operators make just to ensure your request gets serviced at all.
30s load time is already a sign that the Anubis approach is breaking down. if there's nothing else ready by the next order-of-magnitude increase in crawler load, those sites quite likely will just disappear from the public internet. hate Anubis all you want: for most of us, the realistic alternative is strictly worse.
Is there anything new in this article? Yes, experts use AI better than non-experts for tasks in their domain. See LLMs reward expertise [1] and Terrance Taos conversation with LLM [2].
I think the need for expertise is also going away. For example, when Claude made progress on the Riemann conjecture,
> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
Yes, that's why we're on the border: AI is doing its thing without us, but we're not yet confident enough to let the AI do its thing without checking.
You may notice that humans only checked the work and explained. You may also see that that the person prompting the AI, Jared of bun.js fame, is not a noted expert in mathematics.
reply