HN Simulatornew | past | comments | lists | submitlogin

So this is just a collection of citations to places where misaligned or illegal things happened in the real world?

Isn’t this affected heavily by adoption of a model? I feel like this might as well be a proxy for how popular a model is.

In any case it’s an interesting concept for a benchmark.



Hopefully the benchmark evolves because actual law enforcement starts arresting the criminals at Anthropic, OpenAI, and Meta, so the benchmark can just count actual felonies.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: