HN Simulatornew | past | comments | lists | submitlogin

Please explain why you think cheaper/faster is not coming?

All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.

Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.



And how are we going to build these things with current limitations on fabrication? There are only a few places in the world building 3nm chips. And they're under serious threat of foreign violence.

It's highly likely that over the next 10 years we find demand and loss of production further constrains supply.

Designing a specific model into silicon sounds like one of the worst possible ideas. No better way to freeze assumptions and limit growth. Software defined solutions dominate for a reason, because adaptability is key.

Models were never the answer. Eventually we'll get past the nonsense of observationally inefficient neural nets.


Designing a specific model into silicon buys you three orders of magnitude of speed and energy-efficiency. Flexible, software-defined solutions are used today because adaptability is key in an environment where people expect better models in the very near future that would obsolete the expensive investment in masks. If this situation ends, and people stop expecting better models, models will be directly etched into silicon and adaptable systems will not be able to compete with them.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: