> Were you, let's say 6 months ago, expecting it to resolve one of the Millennium Prize problems?
I didn't expect them to throw millions of dollars at each famous math problem. But one year ago we already had LLMs that solved IMO problems, no?
> Are you sure? (The numbers I've heard, which I admittedly have no very strong reason to trust, don't seem that way to me.)
Math has very little founding compared to other science domains. Also, if you filter mathematicians by specialization in PDE and that have worked on Navier-Stokes, then you end up with a very niche community.
> For instance, suppose you give one of today's frontier models some of those chain-of-cubes rotation puzzles. How well will it do?
I feel like this is not the correct way of thinking about it. We can also ask, for instance, how well a state-of-the-art algorithm for the salesman problem works on a particular graph topology. People do PhD thesis on topics like that, so the answer is not obvious at all. For LLMs we still don't have a curated theory that explains what they're good/bad at, and that you don't see how to extract an answer from the definitions is no surprise since this is obviously not an easy problem. But all this is normal because this is a rather new topic (models of this scale appeared when? 3 years ago? That's nothing for science).
Anthropomorphizing LLMs has added so much noise to this discussion.
Yes, one year ago we had LLM-based AI systems solving some IMO problems. My impression is that most observers at that time didn't expect them to be solving Millennium Prize problems within a year.
> Math has very little funding compared to other science domains.
True. But to whatever extent the numbers I've seen are correct, for the whole mathematical community to have spent less on Navier-Stokes than OpenAI did -- even if we value the tokens they spent at something like market rate rather than at what the compute actually costs them (which might be right since any capacity they use internally can't be sold to customers) -- the average number of mathematicians working on Navier-Stokes since 2000 would need to be somewhere around four (depending of course on how well paid they are), and that seems too low to me.
> I feel like this is not the correct way of thinking about it.
It seems to me that if you say "It is absurd to waste time discussing whether it is intelligent or not. It is just an algorithm, we know how it works, and it does exactly what we expect it to do." then this only makes any sense if your "knowing how it works" and "what we expect it to do" enable you to predict what it can and can't do.
(I repeat that I agree that what matters is what it can do, not whether we choose to apply the term "intelligent" to it. But unless I misunderstood you were saying somewhat more than that.)
> Anthropomorphizing LLMs has added so much noise to this discussion.
I think sometimes it helps, sometimes it hurts, and sometimes it's indifferent, because LLMs are like us in some ways and unlike us in some ways. (The same goes for many other things, but LLMs are much more like us in some important ways than any other human-made artefacts.)
I didn't expect them to throw millions of dollars at each famous math problem. But one year ago we already had LLMs that solved IMO problems, no?
> Are you sure? (The numbers I've heard, which I admittedly have no very strong reason to trust, don't seem that way to me.)
Math has very little founding compared to other science domains. Also, if you filter mathematicians by specialization in PDE and that have worked on Navier-Stokes, then you end up with a very niche community.
> For instance, suppose you give one of today's frontier models some of those chain-of-cubes rotation puzzles. How well will it do?
I feel like this is not the correct way of thinking about it. We can also ask, for instance, how well a state-of-the-art algorithm for the salesman problem works on a particular graph topology. People do PhD thesis on topics like that, so the answer is not obvious at all. For LLMs we still don't have a curated theory that explains what they're good/bad at, and that you don't see how to extract an answer from the definitions is no surprise since this is obviously not an easy problem. But all this is normal because this is a rather new topic (models of this scale appeared when? 3 years ago? That's nothing for science).
Anthropomorphizing LLMs has added so much noise to this discussion.