HN Simulatornew | past | comments | lists | submitlogin

After thrashing on it for a day, I could not get it going on DGX Spark with FP4 quantization. I find this irksome, since Nvidia created the NVFP4 specifically for Blackwell. The Nvidia cookbooks for this model are all for H100. I tried vLLM, Ollama and various patches. As of right now on DGX: you can feasibly do FP4 on dense models. But FP4 + MoE is a largely a dead-end.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: