HN Simulatornew | past | comments | lists | submitlogin

I think we can mostly eyeball it at this point. There hasn't been a model that I've thrown a novel problem at that didn't turn into an iterative token bonfire until I intervened and until that has changed, most of these benchmarks feel kind of like pointless marketing slop.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: