Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
does it really matter wherever its the harness or the model for the vibe coder user claude code?
fwiw, i think almost all regressions are down to a/b testing in the harness by anthropic, but it is objectively indistinguishable beyond "the coding agent ceases to be usable" and i'm back to traditional coding for a few hours until its back to normal again
Yes. The agent allows you to switch models, so you could sidestep a bug in one model by temporarily using other models.
The /r/antigravity SubReddit is full of users who very much notice bugs with the tool/agent. We should be thankful that Claude Code is pretty stable by comparison.
It absolutely matters because something like Claude Code has no guarantee that there won't be changes between updates but a model pinned at the API version level that is getting enterprise traffic absolutely does have that guarantee and would be a much more widespread problem...
I’m typing from my phone and im not going to review the semantics of Anthropic’s storied history of performance issues.
It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.
Two postmortems, neither quite "admitted to nerfing":
Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]
April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]
So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.
Sorry, you are correct - I modified my original post. I get frustrated every time there's a model release and 1 week later everyone is saying NERF! NERF! 99.9% of the time these people are wrong, but you are right that it's technically not 100% due to a few edge cases.
I am more skeptical about the compute provider claim - do you have any evidence of that?
I've seen that looping and glitching behaviour in over-quantized (~Q4 or lower) local models like Llama and Qwen 3.x. Thus, it is likely that they quantize the model after release to save on compute costs (while giving a favourable result at launch). That quantization can result in changes to the model's behaviour (you are changing the weights) that could be interpreted as nerfing.
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Your chart is wrong.