Totally. Any MoE model can have experts swapped in and out from disk or system ram. I only framed it this way because the question was about the model fitting in vram.
If you want an expensive model to reason on your files, you need to give them your files.
If you think a cheap model is smart enough to filter information to give to your expensive model, you can save some money. If you think your cheap model is smart enough to format your expensive output, you can save some money.
In practice, this didn't work well until Qwen 3.8.
Qwen 3.6 and (abliterated) Gemma 4 were almost there but still making mistakes.
The obvious goal is to destabilize the western economy and prove that US tech is a worthless bubble - but I agree, OSS AI is great for everybody and what OpenAI was supposed to be
There’s an alternate universe in which OpenAI stays open, licenses according to revenue, Chinese models don’t gain traction in the US because domestic models take all the capacity…whatever, $1T IPO beats the right answer ever time
> Chinese models don’t gain traction in the US because domestic models take all the capacity…whatever, $1T IPO beats the right answer ever time
The trouble here is how more infrastructure helps OpenAI and Anthropic continue billing at 10/100x Chinese model rates.
Either their models have to be better (to justify the higher prices and margin) or their inference has to be lower cost (which isn't going to happen until they move away from Nvidia).
Black market operators can resell stolen account tokens at a lower price than authentic premier tokens from frontier labs and can host their own infra too. I'm not super convinced frontier model serving without downstream model development on a vertical specific software / knowledge worker "factory" model can work
Very cool and I wish you the best of luck (I genuinely hate this generation browsers) - but my experience with bun and other complex zig tools is that a segfault always awaits behind the corner
reply