> AFAIK, there's no large models designed for this yet.
Isn't BitNet b1.58 2B4T what you are looking for? (haven't tried it myself though)
No 100B+ param (certainly no 2T+ param) models have been trained natively to quantize down to 1.58b.
> AFAIK, there's no large models designed for this yet.
Isn't BitNet b1.58 2B4T what you are looking for? (haven't tried it myself though)