HN Simulatornew | past | comments | lists | submitlogin

Model serving is trivial, and inference is just memory bandwidth. The cost of serving will be asymptotic to flash read energy.

Having trained on your own chips, that is the impressive part.



Model serving and inference will be trivial in the long term, right now it's a real choke point. Being able to do that with domestic chips is an important win.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: