HN Simulatornew | past | comments | lists | submitlogin

https://unsloth.ai/docs/models/qwen3.8

The 1bit quant model is at an astonishing 397GB with 95B active per MOE. This literally puts Opus 4.5 performance level into a machine a normal person could buy, and still gets usable tokens/second.

The full lossless model BF16 is clocking at 4.9TB. The model card claims the model to be between Opus 4.8 and Fable 5. Again that's astonishing as getting a machine with 7TB RAM (with context + KV cache) is still within the realm of medium size companies.

Bad things: The open source version has its vision capability removed, and the context capped at 250k . I expect someone to bolt a Kimi 2.6 vision tower to it to restore the vision capability (at less performance of course). For context, I played around with extending the context to 600k for Qwen 3.5 397b, and the context remained stable up to around 480k. It'd be interesting to see if the same can be done to Q3.8 .

Also no out of the box DSpark/DFlash support. MTP is present so we should at least get some boost in TP speed.



To compare a 1 bit quant to the full fat model is misleading.

Honestly this model people at home can tinker with, if you have a big enough Mac. Maybe 4 Strix Halo/DGX Spark, and then at 1 bit quant? Nah.

Use the right sized model, for your hardware. You'll get better results.


Extremely large 1 bit models are usually within 50-60% of KV divergence to lossless models. In this case I think the comparison to Opus 4.5 is a fair assessment.

Extremely large models don't suffer as much from quantization due to its weight topology also contains encoded information, so the loss of info from any one weight is somewhat mitigated.


KL divergence (you misspelled it) doesn't tell you anything about capability drop - how much did this particular benchmark (thus ranking among models) change after 10% or 50% KL divergence?


Not only that, but KL divergence is not a '%'. It's just a number ranging from 0 to infinity that tells you the 'distance' between two probability distributions.


Any one weight, but all of them. And also crushing the architecture itself?

I wouldn't pick up 400gb of hardware to run in that mode. I might try it for fun, but even then you are looking at handling a 95GB active parameter set.

This is NOT a model for most home labs. I'm sure some can and will use it. But most, should steer clear.


I wouldn't just rush out and buy hardware, but there will be benchmarks after a while to make an informed decision.

95GB active is not _too_ bad, would require some creativity and $$, but I bet I could do that at home for less than a cheap car.


I don't understand the logic behind model sizes and quantization.

Suppose I have 100GB of unified memory, how should I know which model suits it best? I understand how a 2.4T model wouldn't fit, but I don't understand the impact of quantization and whether I should use a 200G model quantised to fit say 90GB of memory, or a non-quantised 90G model.


Usually the largest Q4 model that fits and has best reputations. Usually the performance degradation is not considered tolerable below Q4. Usually the model of choice ends up being either Qwen 3.6 27B or 35B-A3B.

What's weird about local LLM models is that closed door improvements in training/RLHF dataset have been so significant that it's rare for larger but older models to make sense - everyone seem to always hard switch to the newest one and report step changes in capabilities(or maybe people running Kimi K2 since release just don't talk about it on the public Internet, giving me that impression).


It depends on your usecase and the size of your ram is not the only driver. I think the primary performance drivers (without sacrificing precision) right now are QAT, MTP/DFlash, MoE and Hybrid approaches to avoid full attention in parts by replacing with linear / sparse attention. So Gemma4 26B A4B QAT+MTP (from unsloth) would be a good pick atm. I'd love to see some smaller models with all that features.


It really depends. It used to be easier to have a rule of thumb, but now it's not clear anymore. Now there are a lot of things to consider, such as a model's kv efficiency (how much context you can fit), MoE v. dense, QAT or not (Quant aware training) and so on.

The old rule of thumb was that a lower quant of a larger model > higher quant of a smaller model. That being said, for some things going lower than fp8 will see a lot of degradation in generation quality. Except if the model comes with QAT 4bit quants. Then there's also nvfp4 w/ calibration data, which also can improve things. So it's really not easy to tell "at a glance" you'd have to test them yourself on your hardware.


Standard models are designed to quantize down to 4-bits relatively well.

Anything below that, and especially 1.58b - is typically complete garbage, and you're much better off running a model 100x smaller at regular precision (compared to one 7x smaller quantized into complete garbage).

If the model was designed specifically to quantize down to 1.58b, then it's different.

AFAIK, there's no large models designed for this yet.


> If the model was designed specifically to quantize down to 1.58b, then it's different.

> AFAIK, there's no large models designed for this yet.

Isn't BitNet b1.58 2B4T what you are looking for? (haven't tried it myself though)


2B is pretty small...

No 100B+ param (certainly no 2T+ param) models have been trained natively to quantize down to 1.58b.


Usually 4-bit 200B model is better than 8-bit 90B. But if you go below 4 bits, I am not sure what is better.


There's no rhyme or reason to it. Quants aren't benchmarked much. Generally 4bit better than smaller model 8bit


Opus 4.5 level of performance is also accessible with deepseek-v4-flash-0731 (0731 being the july 31 update) which is much, much, much smaller. 2x RTX pro 6000 blackwell can run it. 4x can run it very comfortably


I am running DS v4 flash 0731 lossless at 80t/s right now. It really is not at Opus 4.5 level (for my workload). I would say it's around 3.7 Sonnet, which is still pretty good, but other models such as GLM 5.2 are still leaps better. Of course I run DSv4 flash over GLM 5.2 for a few very good reasons, but intelligence is not 1 of them.


Despite fitting into VRAM, I can't get DSV4 to run at usable speeds on my AMD hardware. The upcoming qwen3.8 27b greatly excites me, and I hope it can outperform Stepfun 3.7 Flash, which is the best thing I can run today.


I'm just trying out Muse-Glimmer 30b, and my initial vibe is this might be better than qwen3.6-27b. No idea how it compares to Stepfun, because I can't run that model - but worth checking out while you wait for qwen3.8-27b


Anyone thinking of buying 2x RTX Pro 6000 Blackwells - beware: unlike other cards e.g. RTX 5090, The RTX Pro 6000 cards cannot be NV-Linked, so you'll be going through the PCIe bus instead (7x higher sync cost)


My understanding is the last consumer card that supported that was the 3090. A Google search seems to agree the 5090 does NOT support NVLink...


What do you need the extra 2 for? Tensor parallelism?


Longer context and more cache. The problem is that native format with DSpark enabled you have very little room on the VRAM.


I was under the impression that you could fit the full 1M context within the 192GB VRAM as a result of DeepSeek's various architectural advancements, but I'll grant that DSpark + a larger pool for concurrency may necessitate more VRAM, yes.


Opus 4.5, even 4.6-level performance has been around since July 31st in 284B total params and just 160GB of weights at native FP4 quantization- DSv4 Flash.


> The 1bit quant model i

at this kind of quantization is it useful though?




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: