People are happy to play with AI when the tech companies are burning hundreds of billions of dollars to subsidize it. It remains to be seen who is actually willing to pay for it at the prices required to recoup those insane investments.
Which directly leads to the next big development: all the big players are investing in silicon with "baked-in" models, like [0,1]. Turns out you don't need an expensive general-purpose GPU with heaps of RAM to contain a model when you can make a custom ASIC around one specific model! Why spend a fortune on DRAM / HBM when all you need is some finetuning parameters which are easily stored in on-die SRAM?
> It remains to be seen who is actually willing to pay for it at the prices required to recoup those insane investments.
Everyone. Open models are going to keep prices down. A lot of current models are more than usable. Self-hosting would have been an option if hardware prices get back to sane values.
The insane investments have to do with insane over-valuations, VCs involved and hence media spam on it. Chinese labs for example make do with 1% of the valuation and 1% of the resources.
Prices must raise a lot to make self hosting mainstream or at least fairly popular.
Yesterday a coworker posted on a customer's Slack the specs of a box he is planning to buy to run local models. It's about 5k Euro. I am paying 18 Euro per month for Claude Pro and I'm going through a migration (almost a total rewrite) of a web app from Vue 2 / Vuetify 2 / Vuex to Vue 3 / Vuetify 4 / Pinia. I never hit the 6 hours limit. I could consider running an equivalent model on a 500 Euro machine (about 2 years of Claude Pro) but 5k is 20 years and that box will be obsolete or will have failed beyond repair (no spares) much earlier than that.
Which directly leads to the next big development: all the big players are investing in silicon with "baked-in" models, like [0,1]. Turns out you don't need an expensive general-purpose GPU with heaps of RAM to contain a model when you can make a custom ASIC around one specific model! Why spend a fortune on DRAM / HBM when all you need is some finetuning parameters which are easily stored in on-die SRAM?
[0]: https://www.theregister.com/systems/2026/08/06/amd-acquires-...
[1]: https://thenextweb.com/news/google-frozen-chip-gemini-silico...