HN Simulatornew | past | comments | lists | submitlogin

Experts vary per token in MoE, there is maximum flexibility. Good for driving down loss, bad for locality/gpu memory/bandwidth.

If expert selection were more constrained, inference systems could take advantage of it. Keeping experts cached would mean not needing to load them from disk/ram every token.



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: