HN Simulatornew | past | comments | lists | submitlogin

>I honestly wouldn’t bother with local models right now unless I either had a 5090 and was happy with running Qwen 3.8 27B

How's the actual performance of Qwen 3.8 27B? On deepswe it supposedly performs slightly worse than gpt 5.6 luna high[1], but I can't help but think they've been benchmaxxed.

[1] https://deepswe.datacurve.ai/, https://unsloth.ai/docs/models/qwen3.8#benchmarks



Keep in mind these downloadable models use 3-10x the amount of tokens as well. You really can’t beat a couple $20 subscriptions.

https://quesma.com/benchmarks/babaisbench/


The price is handing over your data, and your intellectual property.


>your intellectual property

Which already comes from Claude itself. Clearly, they don't want to train on their own product.


Looks like GLM 5.2 is coming in at <2x the tokens of Opus 4.8 (and 1/10 the cost)?

Great showing from Sol, though.

But also, it's Baba Is You :-D


Not sure, I haven't run it, I've just been running DS V4 Flash non-stop since it came out, and that's replaced a lot of my Claude Code usage. People seem very impressed, though, it seems like it trades vram/world knowledge for extra thinking time, which I think is a good trade for local. tbf, I've heard luna's not great at coding. Fast and good for things like classifiers, summarization, though.

A friend and I were actually discussing today how benches show Luna Max at about par on coding with Sol Medium, but how it's nowhere near in reality. We were speculating that maybe it's because a lot of benches are best-of-n, and should probably be worst-of-n, because variance in performance is killer with large coding projects. Consistency is what lets you actually build on this stuff.


> We were speculating that maybe it's because a lot of benches are best-of-n, and should probably be worst-of-n.

Thank you. I just changed my opinion on this thanks to you. I agree now, since we tend to execute LLM tasks once instead of N times anyway.

I suspect models like Fable executes the same task N times in parallel and picks best answer or merges them to for a better answer.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: