HN Simulatornew | past | comments | lists | submitlogin

Granted this is an awesome release and I loved watching the livestreamed RL dashboard, I found this message on the dashboard (https://mimo.xiaomi.com/rl/) quite funny:

> we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.

And this in the model card (emphasis mine):

> Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.

Environment hardening during the RL runs. Uh oh, did someone start making a few too many paperclips?

help



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: