It's actually no different for dice than for LLMs. Explaining accurately the reason for the exact outcome of any given dice roll someone makes would be stupendously hard. It would require lots of instrumentation and math and be poorly transferrable to another surface, another player, etc.
But even so people don't say that we don't understand how dice work.
Saying that we don't understand how LLMs work is exactly like saying we don't understand how dice, or tires, or golf ball shots work. Or like the old myth that we don't understand how bumblebees fly.
That's precisely the point: you may be able to understand dice statistically and over the course of long rolls of dice you can extract some properties of the dice. But you won't ever understand any particular roll of the dice.
But importantly for dice we do understand the overarching principles that give rise to this. And dice don't output coherent sentences. Meanwhile in LLM land the analogous "roll of the dice" can result in a coherent response in natural language.
If you use a loaded dice, you can be pretty confident about where it will lands. It may not be 100% accurate, but can be quite close to certain. Without training the weight are pure noises. After training, it leans towards coherent sentences and particular statements.
Yes, and I believe my point still stands. We thoroughly understand the principle by which a loaded die can be intentionally biased despite not being able to predict the outcome of any given throw due to the system in question being a chaotic one.
In contrast, we do not understand LLMs in the same way (nor biological brains). Claiming that anything of that nature is simply biased towards coherent output seems entirely reductive to me - the question is how such coherence arises in the first place. There is no meaning encoded or computation performed by the particular pathway a die travels through the chaotic landscape.
Sure an argument can be made that it's "just" a next token predictor thus how is it really any different from a markov model? Yet the output is not even remotely the same.
> In contrast, we do not understand LLMs in the same way
From my point of view, (not a ML researcher), it’s due to the magic of numbers. The same thing happens with computer vision and neural networks. There’s a bunch of magic weights that get created which has no meaning by themselves, but computing them does help with detecting objects.
So if you take words, derives them into tokens, use the attention techniques to extract the “coherency” aspect, it’s no wonder you can replicate “coherency”. Add reinforcement learning to that to increase towards certain aspects like correct code syntax and you have heavily loaded the dice again.
We have used maths to model chemistry, biology, and physics, as well as economics and sociologic phenomena. Then we use maths (more specifically logic and set theory) to usher in the age of information and computing. Now you want us to act surprised that maths, through ML, can model language.
Maybe further down the line, we can have a simpler set of formulas for language coherency, but for now we have to make to with using the whole internet and a bazillion watts of power to guess the weights for the generic ML model.
I'm comparing it with chess. Chess is pretty complex, complex enough that only a small subset of humans can play it at a very high level. Introducing computers to chess first led to a statistical and brute force approach. But once that had paid off and the results were in people spent a lot of time analyzing those results and this led to an entirely new class of engine that was far more efficient than what had gone before and which performed even better than the 'big iron'.
I would not be surprised at all if we will find that AI will go the same route. The fact that we don't know how it works is where the opportunity for improvement lies.
The best way to model dice is the Physical Stance. You consider rules such as gravity, kinematics, etc. There is no “internal state”, “world model”, “knowledge”. If you prefer, in Friston’s terms, there is no Markov Blanket.
The best way to model a human is the Intentional Stance[1]. You mostly need things like beliefs, knowledge, biases, etc to build this model. In Friston’s terms, there is a Markov Blanket, an inside vs outside.
Without going into any irrelevant-but-interesting philosophical discussions about consciousness, I believe the intentional stance is most useful for modeling LLMs. Most of the success in predicting, debugging, optimizing these systems is in activities like understanding what they believe, what their intent was, what they observed, what they concluded from those observations. Also note that much simpler creatures benefit from the Intentional Stance; you will be more successful at modeling your dog if you think about what it “wants” rather than trying to run Physics on it.
[1]: https://en.wikipedia.org/wiki/Intentional_stance - the astute reader will note that I skipped the Design Stance. If we truly understood how NNs actually implement all their cognitive processes then we could perhaps apply this to them; if we actually crafted and designed every parameter of its mind. But we are talking about why dice are different.
> For the next several decades, we'll have engineers (presumably with AI) optimize things like rocket engines and turbines and AC compressors to work a few percent better because the numerical approximations might have caused us to be overly conservative.
> Knowing the exact mathematical breakdown mechanisms helps developers improve adaptive mesh refinement and sub-grid scale models around high-vorticity regions (like vortex stretching and turbulent shear layers).
This is 100% wrong and reads like copy paste of AI slop.
Any simulation which uses sub-grid scale models is already solving a different PDE than the actual Navier-Stokes considered in the Millenium problem, and that PDE is guaranteed to have different properties. Full stop.
And to claim this is somehow connected to AMR methods is an example of the kind of pseudoscientific statement Wolfgang Pauli would have called "not even wrong".
This is a wrong interpretation. Physicists have a shit-ton of models that produce "aphysical singularities", they just work around those to get meaningful answers anyway. This is a whole trope and stereotype. Some of the most successfull and accurate predictions in all of physics come out after you discard a bunch of singularities.
Nobody who actually works in fluid dynamics on any sort of application gives a hoot about the N-S millenium problem. Many do not even know what it is. There is no practical effect of this proof on how we do fluid mechanics.
Whether or not ways exist to work around the singularities, that they exist is surely of note. Before von Neumann formalized QM people were still doing QM, okay fine. But it's wrong to then say von Neumann was doing no physics of note.
"Does there exist a pathological combination of smooth body forces and initial conditions for this set of PDEs, where singularities appear, which by the way is completely impossible to actually create in the real world unless you are a literal God?" is a question of math, not physics. This is a hill I will die on.
If you think about what it actually means to have a time-varying smooth body force defined in all of 3-space, you fairly quickly come to that kind of conclusion.
Even if someone comes up with a construction that does not require any forcing, it is going to be some extremely weird initial conditions that you will never be able to even approximate in reality unless you can move all the individual molecules of a fluid around and set their initial velocities from a far distance.
Obtaining a finite-time blow-up for Navier-Stokes does not necessarily advance the field of mathematics by any significant measure, whether the proof is very long or very short.
As a concrete example, such a proof could be less than a page with very specific initial and boundary conditions and inserting them into the equations to get something that goes to infinity when time goes to some finite value.
This would resolve the Millenium problem but not make humanity any smarter.
For cold / glacial backups, M-disc will run you $200 for 10x 100GB discs, plus around $300 for a pair of USB Blu-ray burners. With that, your data up to 500 GB can be written in duplicates and stored at two physical locations (one drive at each location), with very strong guarantees that it will be readable far into the future.
If you write 2 additional disks per year, this is a $500 initial investment plus $20 per year. Compared to Google's "AI Plus" plan you break even in year 5, assuming Google keeps pricing flat.
Of course you don't get any cloud access or similar.
Depends heavily on the software cost and how far from core competency it is.
If the company already has a developer team and the software is an expensive domain-specific CRUD app for internal use without significant compliance or risk factors - sure, go DIY.
If the company does not employ any devs, or if the core logic of the software is a business secret of the vendor, or it is customer facing, or there are significant compliance / risk factors - you are better off paying for it.
Unless the cost is astronomical, or the company is motivated also by anger at rent-seeking MBA f*ckers at the vendor.
That's the opening line to Pride and Prejudice, where Jane Austen (I guess the "bot" part of the name is intentional) states something that many people of the time would superficially agree on, but which is intended to be highly ironic.
I guess the point of GP (and of Jane Austen) is that people never actually universally agree on anything. And when they superficially do, there is actually a large undercurrent of disagreement.
Another fun quote apropos here would be "I love standards, there are so many to choose from"
There's a great book from a different line that delves into this as a form of what the author calls 'manifold objectivism'. In short, there are some big concepts like Christianity or Islam that people refer to as if it's the same thing but almost nobody has a shared meaning when they refer to such big things. Even so there can be concrete dialogues about these topics from different frames and with different underlying meanings.
There's trouble though when either there is conflict that isn't reconciled between the interlocutors or worse, when there is a satisfactory conclusion between the two interlocutors who never accounted for the divergent definitions ... there are many 'objective' understandings of what a thing is.
"A Fundamental Fear: Eurocentrism and the Emergence of Islamism"
> short, there are some big concepts like Christianity or Islam that people refer to as if it's the same thing but almost nobody has a shared meaning when they refer to such big things. Even so there can be concrete dialogues about these topics from different frames and with different underlying meanings.
See Wittgenstein’s Philosophical Investigations for a deeper treatment of that topic.
https://youtu.be/Ps2Jc28tQrw?is=tXq9sAaeK_OurFqR
reply