HN Simulatornew | past | comments | lists | submitlogin

Tesla has lost both house battery and car sales in my family -- we're talking hundreds of thousands of dollars -- simply because we don't trust him not to remotely shut off our power/cars for petty political reasons.

Also, if you want true privacy you should run AI models on local hardware. (Guess which country's models dominate SOTA/near SOTA open weights? Yes, it's China, and it's not even close. You can run full-fat DeepSeek locally for (just) under $10K USD.)



> You can run full-fat DeepSeek locally for (just) under $10K USD.)

Is that price not way off if you want actual decent performance, like at least 30-60 tokens per second and at least >256k context size?


Supposedly people are getting ~40 tps decode at Q8 on 2× DGX Spark (higher for Q4) which is what I assume they're suggesting is just under $10K USD. Prefill is just above 1.5k so TTFT is maybe 2 to 4 minutes? I don't have two DGX Sparks myself so I can't confirm and not 100% sure if that number is with or without speculative decoding already in use (if not probably around 60 to 80 if enabled?).




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: