After thrashing on it for a day, I could not get it going on DGX Spark with FP4 quantization. I find this irksome, since Nvidia created the NVFP4 specifically for Blackwell. The Nvidia cookbooks for this model are all for H100. I tried vLLM, Ollama and various patches.
As of right now on DGX: you can feasibly do FP4 on dense models. But FP4 + MoE is a largely a dead-end.