With the benefit of LLMs already being proven, in a couple of years we will have vastly better hardware for inference I guess.
I feel like now hardware is stagnating a bit, because the software side has moved too fast for the hardware to catch up. Once we settle on some good, optimal software architecture for the models, dedicated hardware will easily increase throughout by 10x or 100x, for a fraction of the cost.
LLMs seems quite simple, maybe we'll be able to print/assemble at home our own chips with the desired models/weights.
Maybe we'll have model weights being shared like game cartridges.
I fear the future of local models will be controlled by governments. I feel like some time soon there's going to be a crackdown on what is available to download, what is hostable, and what is "acceptable". I partially suspect it has something to do with why 128 GB seems to be the most you can currently purchase for a single machine, despite the price.
Well, everything—even human life on Earth—is just temporary, but RAM supply lagging centralized-AI-driven demand increases continuing to squeeze the consumer market may not be a short term phenomenon.
> LLMs seems quite simple, maybe we'll be able to print/assemble at home our own chips with the desired models/weights.
So, the solution to the RAM crunch is “everyone has their own home chip fab and deals with the raw material supply and hazardous waste disposal”?
You're not thinking through the way economics works.
If, hypothetically, the thing you wish existed, why wouldn't it be valuable to high-utilization cloud providers?
And if useful to them, why would it be sold at a lower price than the (extremely high) price those cloud providers can afford to profitably pay?
The only solution to the hardware crunch that eases the cost/access barrier is if there's such a glut of excess supply of newly-commoditized inference hardware that the price craters.
And you've yet to propose a feasible manufacturing scenario through which that would be possible.
Well, before AI the RAM prices were cheapest ever, even if cloud providers were still in demand for hosting and cloud computing.
I hope the supply will increase at some point, I doubt cloud providers will be able to buy everything, especially if the competition will be high and the price of AI goes down, so there won't be an easy way to make profits by serving models.
> so there won't be an easy way to make profits by serving models
That's the only way I see things going, which ultimately comes down to the market losing faith in AI as a profit engine.
Which is probably going to take one of (a) OpenAI / Anthropic IPO failing or newly public financials looking bleak or (b) a high-enough profile large AI company failing (unfortunately most of the large enough to shock the market ones have diversified and durable cash generation engines).
So the solution to the chip supply problem is “pretend there is no chip supply problem and everyone gets the chips they need, but for some reason does some assembly task at home”?
If we could just do the first part and get reality to conform, the rest would be superfluous, but I don't think we can do that.
Its not a solution ever, because it takes as its premise that the problem is solved and then goes on to suggest something else on top of it (the purpose of which is unclear.)
I personally think of this like sorting algorithms. Quick sort does the same thing bubble sort does so why do we need quick sort? Pushing for efficiency drives innovation. It does this for many reasons but a big one is that putting a cap on a resource forces you to consider the others available and often you find that all it took was a little effort and suddenly the alternate path that looked a little worse is actually better than you realized.
This has a lot to do with how MCTS works BTW. The current best path is often only the current best path because a lot of investment has been sunk into it. If you were to put equal resources into a different path you may find that it was actually far better. It is just that the early rollouts favored the other 'best path' so you sunk a lot of resources into that one. We are very early in our exploration of LLM architecture. I highly doubt we are anywhere near the best path right now.
The diffusion models are interesting, but those also seem hacky.
I think the next form of AIs will be simpler and more abstract.
The building blocks of our brain don't have the notion of a "token" embed into them, it's lower level that that.
I think first step is to find a better way to represent information.
LLMs shouldn't "compute" stuff using language tokens, but some other, more efficient logical mechanisms. LLMs should first "feel" the solution, reason internally in that optimised space, then, only when interacting with a human should it convert all that into actual tokens/language.
The problem with that is we don't have any kind of training data in that abstract sense, maybe we could use RL to figure that out but current RL techniques are too slow and prone to breakage that anybody trying to use them to train a big enough general model (LLM, diffusion, world model, etc) will either fail or have to make a very very big investment.
The other option is maybe hook up humans to EEG or the likes and map their brains while they solve different kinds of problem, or just see and feel the world around them
With the benefit of LLMs already being proven, in a couple of years we will have vastly better hardware for inference I guess.
I feel like now hardware is stagnating a bit, because the software side has moved too fast for the hardware to catch up. Once we settle on some good, optimal software architecture for the models, dedicated hardware will easily increase throughout by 10x or 100x, for a fraction of the cost.
LLMs seems quite simple, maybe we'll be able to print/assemble at home our own chips with the desired models/weights.
Maybe we'll have model weights being shared like game cartridges.