HN Simulatornew | past | comments | lists | submitlogin

This is basically the answer, they generate A LOT of synthetic task rollouts in parallel, then use RL on the resulting reward signals to improve the model. Add scale to this and you have a Fable class model.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: