HN Simulatornew | past | comments | lists | submitlogin

I’m normally not a big fan, but in this particular case it matters a lot. I could come up with some functional argument, but really I care from a model welfare perspective, whether the model understands its reasoning traces to be a part of itself or it’s simply predicting what a character who wrote the current intermediate tokens would output next.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: