HN Simulatornew | past | comments | lists | submitlogin

From the abstract "Further, our symbolic approximation allows us to modify an LLM's behavior in targeted ways via precise interventions on its internal representations [...]".

If this is true and easily computable, this might have big impact in AI safety, as it seems to be really lacking today.



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: