It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai.
just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so
i would say its industry consensus at this point. the creative output of the anthropic models is far ahead of openai. the benchmarks cannot capture the difference
Ive anecdotally heard that openai is far more chaotic, which includes not having a central infra team for example (or at least some teams not counting on depending on them). At least the previous hacks in openai were mainly due to bad infra architecture design.
the creative and "big picture understanding" of anthropic models are noticeably ahead of openai. external models are distilled representations of internal models, its clear who is ahead
I don't know that we're in the singularity, but we are certainly past the point where LLMs can only write spaghetti code. LLMs can write complex systems, when properly lead by software engineers. They can produce code faster, and at a higher quality, than purely human endeavors.
The higher quality part is the part people are missing. You can write much more robust code using LLMs because you can employ more comprehensive testing strategies. People are using LLMs to find hundreds of vulnerabilities in popular software. Imagine how much more secure software can be when LLMs because integrated into the process of writing, testing, and penetration testing code.
> I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?
> When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it
> That's the entirety of Anthropic's billions of dollars of research: any prompt with the word "reasoning" is trying to hack Claude to figure out how it reasons!
> A model like that should never have gotten out of QA, let alone been released.
How do you define "way" when saying ahead? How is this measured?
I only use open weight models now and I don't really feel a loss, curious what those who still use it think. I see output from coworkers that does not indicate Claude is that much better (still makes dumb mistakes all the time), not sure they are using the most expensive models either though.
sounds cursory, it takes more than a day to learn the quirks of a model, have you put similar effort into customizing / harness engineering your open weight interactions as you have claude?
you are definitely displaying strong bias that Anthropic is way ahead of everyone throughout your posts under this story
as such, I give your opinions zero weight, they don't align with the majority of accountings or my own experiences
vice a versa re your usage of open weights, they are way more capable with good tools, context, process, and harness engineering
here's an example of Qwen-3.6 35B A3B MoE porting my phd code to JAX with only high level guidance from my expertise, newer qwen models share the same noticeable step change in capability as recent Big Ai models
if open weights were so inferior, they would not be >50% of all token processing
I don't think this follows at all. Just like benchmarks get saturated, lots of tasks get saturated as well. Over time, you can accomplish a given task for much cheaper, and part of that is due to open weight models. That doesn't imply that they're competitive with frontier models for the most advanced tasks, which might represent a smaller fraction of overall work, and thus use a smaller portion of tokens.
That said, at the moment I'm finding that not much can compete with GPT-6 Luna on cost / performance (not using for coding, but for AI pipelines in my product).
you've definitely left rational discussion for emotional responses man, you won't persuade or convince anyone with takes like this
why are open weight models seeing such rapid rise in usage?
there has been a step function change this summer, like the end of last year for closed models
---
do you think you would experience real (legitimate) feelings of loss were you not able to chat with Claude again?
(for clarity, I am not attempting to delegitimize real feelings that real people experience, regardless of my biases, it's a question from curiosity about how others are engaging with the technology)
just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so