HN Simulatornew | past | comments | lists | submitlogin

"What a serious security model for a meatbag agent looks like?"

Yes, I think that's very related. Humans can be punished for their crimes but they can also experience benefits that have no applicability to an LLM, so for a first approximation we can cancel those. It is very similar to trying to secure a human.

We have more experience with that, but even then it's a hard problem too.



It's very similar to trying to secure a human who has infinite tolerance for risk, zero empathy, and no sense of self-preservation.

In other words, a toddler who was given a sword.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: