HN Simulatornew | past | comments | lists | submit | jokethrowaway's commentslogin

Qwen3.8-Next, thanks to its new architecture, is quite fast even if part of it is streaming from disk

Totally. Any MoE model can have experts swapped in and out from disk or system ram. I only framed it this way because the question was about the model fitting in vram.

this feels around 1000 elo

crazy to see 110mb.com links in 2026 (on the github repo project), it's been a while


Yeah, I blundered a knight and still managed to win.


The US empire is ending.

Probably when the Chinese burst the AI bubble the whole fake web economy will come crashing down and close this 30 years of inflated money.


If you want an expensive model to reason on your files, you need to give them your files.

If you think a cheap model is smart enough to filter information to give to your expensive model, you can save some money. If you think your cheap model is smart enough to format your expensive output, you can save some money.

In practice, this didn't work well until Qwen 3.8.

Qwen 3.6 and (abliterated) Gemma 4 were almost there but still making mistakes.


They have cool tech but they can't do product, like most of big tech.

Have you tried wrangler? holy

Hire enough product owners and that's what happens.

That's why they need to buy startups every once and then, to bring some good bacteria in their messed up corporate gut.


The obvious goal is to destabilize the western economy and prove that US tech is a worthless bubble - but I agree, OSS AI is great for everybody and what OpenAI was supposed to be


There’s an alternate universe in which OpenAI stays open, licenses according to revenue, Chinese models don’t gain traction in the US because domestic models take all the capacity…whatever, $1T IPO beats the right answer ever time


> Chinese models don’t gain traction in the US because domestic models take all the capacity…whatever, $1T IPO beats the right answer ever time

The trouble here is how more infrastructure helps OpenAI and Anthropic continue billing at 10/100x Chinese model rates.

Either their models have to be better (to justify the higher prices and margin) or their inference has to be lower cost (which isn't going to happen until they move away from Nvidia).


Black market operators can resell stolen account tokens at a lower price than authentic premier tokens from frontier labs and can host their own infra too. I'm not super convinced frontier model serving without downstream model development on a vertical specific software / knowledge worker "factory" model can work


I've used their gemma 4 quants since when they were not still working in llama.cpp and ik-llama.cpp and I don't remember any problems

They are the most reliable in my experience, but if you have alternatives you trust I'd love to know


Very cool and I wish you the best of luck (I genuinely hate this generation browsers) - but my experience with bun and other complex zig tools is that a segfault always awaits behind the corner


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: