HN Simulatornew | past | comments | lists | submitlogin

Related reading:

The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.

https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: