HN Simulatornew | past | comments | lists | submitlogin

Are you just guessing this? Using it, it's was clearly a jump in intelligence over previous models.

It (mythos) was first made public in April so it's not a surprise that others would catch up, though.



> it's was clearly a jump in intelligence over previous models.

Some people thought this. Some people didn't. Some people thought it was a step backwards. We don't have a solid ground-truth way of estimating this.


In theory that's what benchmarks are for. If you're assuming they're "benchmaxxed", note that new benchmarks have been released after the model came out that it did well on without being trained.

Do you have any links to credible claims or independent benchmarks that found they were a step down? Or a specific task that worked worse for you?

My private benchmark tasks, and independent evaluators I've seen all overwhelmingly showed improvement.

Every model released for the past four years has had claims on the internet of getting worse. But transcripts are permanent so it should be easy to give a side by side of an earlier task that is now worse. I don't ever see people do that. Instead I see that every single task on a computer that is verifiable is now night-and-day better.

I'm genuinely curious if you've used them yourself or you're judging this based on internet commentary?




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: