HN Simulatornew | past | comments | lists | submitlogin

They're claiming a drop in token use too, and that it nets to 40% cheaper.
help



Unfortunately, they're full of it https://artificialanalysis.ai/models/claude-opus-5-5#token-u...

It does work out to be a similar cost per task though


You should probably look at the cost/score graph by effort level instead:

https://artificialanalysis.ai/models/claude-opus-5-5#intelli...

It is most of the pareto frontier.


Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort

so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)

You’re missing my point. I’m saying anthropic are exaggerating their results.

how are they exaggerating the results? Comparing the cost from that chart for 5 and 5.5 for medium-max effort paints a pretty clear picture:

         mean  median
 model
 5      4.135   4.245
 5.5    3.150   2.640

Again seeing how max is a clear outlier, the median cost saving is ~38%, not that far off from the proclaimed 40%.

Is verboseness the only measure of token efficiency towards overall task completion?

5.5 is higher for max effort, slightly higher for xhigh and lower for high, medium and low effort.

The biggest proportional difference seems to be at max (5.5 is 38% more) and at high (5.5 is 21% less).

I think most people run at high and xhigh. At xhigh it is close enough to be task dependent and I don't think most people will notice. At high effort I think it looks like it will be an improvement for most people.

5.5 Max should probably be compared to Fable - it performs a lot better than 5 Max.

https://artificialanalysis.ai/models/claude-opus-5-5?models=...


I don't think so, I typically use Opus 5 on High, and 5.5 scores lower on token use:

https://artificialanalysis.ai/models/claude-opus-5-5?models=...


[flagged]

Very fast you were.

Yes, I need 30 min to run the test suite with new model. Why the sparky comment

Even created an account to tell us just that.

Yep, moving here from reddit. Thanks for constructive discussion.

I tested it with Claude Code, and I can confirm it's way cheaper, better, faster and less verbose than Opus 5.

parent means that they could get more client / a larger part of the market, which would lead to more income (more tokens) despite lower marginal prices



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: