HN Simulatornew | past | comments | lists | submitlogin

It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (https://artificialanalysis.ai/models/deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of https://mimo.xiaomi.com/mimo-v2-6 the deepseek model sometimes surpasses mimo and it's not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)
help



The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn't actually been "corrected". It's only Fable 5.1 that is shown as ahead of Opus 5.

The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.


The way Artificial Analysis keeps changing their weights feels kind of like deciding who the winner should be and making the weights reflect that. They’ve been changing their weights to add more weight to improved long-running agentic capabilities, but doing so means they’re reducing the relative importance of world knowledge and of writing ability.

I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.


How does someone objectively quantify writing ability?

I mean artificial is in their name....

> It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 gets 39.

Why?


> W

W what?


All of the chinese labs have been overfitting on benchmark data to game the results for a while now - MiMo and Deepseek are not anywhere near frontier and mostly compete with models like Luna - which they are still worse than.

There isn't much compelling reason to use these unless you are just averse to giving money to openai/altman. A $20 codex sub gives you ~$150 of luna use per weekly limit, while there isn't any good subsidized options for chinese models at all (and the few who were subsidizing, like opencode, rugpulled by reducing monthly limit to $60 to $15 with no notice to users).


If you hit the $20 codex 5 hour limit wouldn't that be a good situation to switch to a Chinese model for a bit?

I mean, why would you? Chinese models through an API are more expensive than any subsidized subscription and the only one offering a transparent subsidy (opencode) only gives you $15 of credit on any relevant models, and still has a 5 hour limit.

The answer is just a second $20 subscription.


rugpull is a very strong word... I agree that they are not great with their pricing/credit communications, but they do state the multipliers quite clearly in various places. And plenty still have $60 or $30.

Changing pricing buried in docs (which they didn't actually do when they started switching every model to $15 limit 2 months ago BTW, they only did this once people noticed and called it out) without notifying users is the definition of a rugpull.

DS 4.1 is good but it’s clearly not as “smart” as non-flash models- it just doesn’t have the training data. Without a solid plan, it goes off the rails pretty regularly.



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: