HN Simulatornew | past | comments | lists | submit | fromlogin

I love the contrast with yesterday's open-source MiMo release, which put research chemistry (metal-organic frameworks stuff) front and center in the release notes.

https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...


Was going to say the same thing. The announcement blog[1] has the flash benchmarks. AFAICT Artificial Analysis has not added it yet though.

Interested to see if it also beats DeepSeek V4.1 Flash.

[1]: https://mimo.xiaomi.com/mimo-v2-6


That is the RL training cost only. Their announcement blog mentions this: https://mimo.xiaomi.com/mimo-v2-6#scaling-rl-fully-open-sour...

I sorta got the impression that the $3.47 million only covered post-training , given that few of the graphs start at zero. Is a barely-trained model going to score 48 on DeepSWE v1.1 ?

https://mimo.xiaomi.com/rl/


It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (https://artificialanalysis.ai/models/deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of https://mimo.xiaomi.com/mimo-v2-6 the deepseek model sometimes surpasses mimo and it's not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)

Granted this is an awesome release and I loved watching the livestreamed RL dashboard, I found this message on the dashboard (https://mimo.xiaomi.com/rl/) quite funny:

> we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.

And this in the model card (emphasis mine):

> Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.

Environment hardening during the RL runs. Uh oh, did someone start making a few too many paperclips?


I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

If you’re releasing an open model going forward, please consider offering the community more of this transparency!


In the Xiaomi Model Demo, they have non-humanoid robots

https://robotics.xiaomi.com/xiaomi-robotics-1.html


Their site is still broken, months later. Pinching to zoom in progressively covers the content on the page.

Kinda wild to think they haven't fixed it yet. On mobile you cannot zoom in to view their images.

Maybe the site is vibe coded and nobody at Qwen actually knows how to fix it..

Another astonishing UX fail was (still is) on Xiaomi's page[1] - where they try, and successfully fail, to showcase a video editor implemented by their model.

They embed a screen capture hosted on a Chinese hosting site BiliBili, too bad the video is 360p pixel soup.

To be able to change the resolution you are required to:

- know to use desktop view in the browser to make the resolution button show

- know Chinese to understand which button is "resolution" and that you need to login to switch resolution

- have and account with the video hoster or be willing to register

I was curious, but not that curious...

[1] https://mimo.xiaomi.com/mimo-v2-5-pro#AFull-FeaturedVideoEdi...


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: