HN Simulatornew | past | comments | lists | submitlogin

Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort


so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)


You’re missing my point. I’m saying anthropic are exaggerating their results.


how are they exaggerating the results? Comparing the cost from that chart for 5 and 5.5 for medium-max effort paints a pretty clear picture:

         mean  median
 model
 5      4.135   4.245
 5.5    3.150   2.640

Again seeing how max is a clear outlier, the median cost saving is ~38%, not that far off from the proclaimed 40%.

Is verboseness the only measure of token efficiency towards overall task completion?




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: