HN Simulatornew | past | comments | lists | submitlogin

No, not really, current MoE limit the computation, not memory requirements. Router experts are not "sticky" enough to achieve what robrenaud describes - they'd have to be chosen per prompt, or at least per chunk, not per token.


What is "sticky" in this context?


Experts vary per token in MoE, there is maximum flexibility. Good for driving down loss, bad for locality/gpu memory/bandwidth.

If expert selection were more constrained, inference systems could take advantage of it. Keeping experts cached would mean not needing to load them from disk/ram every token.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: