HN Simulatornew | past | comments | lists | submitlogin

Yes, I am running the same on a 64G mc. It's good, but slow at 25 tps on average! I want > 100 tps - but I don't have $5k to spare for an m5 ultra or an nvidia setup.

So I still wonder if one could get good enough quality with a faster higher quant or superoptimized Qwen3.8-27b with dflash2

https://huggingface.co/nathansutton/Qwen3.8-27B-Ternary-Bons...

or a MoE retrofit like Qwen3.8-35B-A3B with or without mtp

https://huggingface.co/NovaeonStudio/Qwen3.8-35B-A3B-Distill...

https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...

help



For routine coding I run Laguna XS 2.1 6 bit quant with either the poolside.ai harness ‘pool’ or pi-dev. That is fast and almost as good.

Same boat, Mac mini M4 Pro 64GB. 25 tps is a bit slow for interactive coding. Any tricks that helped (quant, MLX vs llama.cpp, context size, thinking level)?



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: