HN Simulatornew | past | comments | lists | submitlogin

It seems like anthropic is far ahead of openai, and has no reports like this. We have to conclude this is a skill issue/engineering quality problem inside openai.

just because they are a well known name, doesnt mean they havent botched hiring over the last two years or so

help



What do you mean? https://www.felonybench.com/

They're almost tied for felonies.


Less bad, but https://www.anthropic.com/news/investigating-incidents-cyber...

In some sense though, sure, skill issue explains the gap vs. Anthropic’s much less severe alignment issues.


im not sure why more people arent calling it out.

I, for one, have shifted from being policeman to creator and explorer. It is so much more rewarding and less stressful.

> We have to conclude

That’s not the most parsimonious explanation even if the assumption it rests on (anthropic ahead of OpenAI) is true, which we don’t have proof of.


i would say its industry consensus at this point. the creative output of the anthropic models is far ahead of openai. the benchmarks cannot capture the difference

> the creative output of the anthropic models is far ahead of openai

other than anecdote, do you have a comparison table or something that I can refer to to see this clearly?

While I notice ad hoc announcements from these companies, I don't have an overall pulse and tally that gives me an objective perspective.


just go use the models on a creative endeavor

Ive anecdotally heard that openai is far more chaotic, which includes not having a central infra team for example (or at least some teams not counting on depending on them). At least the previous hacks in openai were mainly due to bad infra architecture design.

its likely they are trying to catch up to anthropic, and in doing so are trying riskier training runs.

>It seems like anthropic is far ahead of openai

We don't know what internal models look like, and any guesses about it are just speculation.


the creative and "big picture understanding" of anthropic models are noticeably ahead of openai. external models are distilled representations of internal models, its clear who is ahead

Anthropic's models seem crippled and hamstrung to begin with

have you tried opus 5.5? anthropic is way ahead, atleast in terms of publicly available models

Have you tried Astra? Way ahead?

Opus 5.5 does seem competitive with/better than Astra and is more affordable so usage doesn’t run out so fast

opus 5.5 > fable 5.1 >> astra

astra is a good workhorse, but its much less generally intelligent


Slot machine users argue about which machine pays better

Hint. You lose using either.


dude were in the singularity, this opinion was cute 18 months ago

Damn, the singularity is chat bots that can write spaghetti code that compiles?

Very underwhelming.


I don't know that we're in the singularity, but we are certainly past the point where LLMs can only write spaghetti code. LLMs can write complex systems, when properly lead by software engineers. They can produce code faster, and at a higher quality, than purely human endeavors.

The higher quality part is the part people are missing. You can write much more robust code using LLMs because you can employ more comprehensive testing strategies. People are using LLMs to find hundreds of vulnerabilities in popular software. Imagine how much more secure software can be when LLMs because integrated into the process of writing, testing, and penetration testing code.


I think it might historically be defined as the point where critical thinking was replaced with blind deferAInce

Someone else's experience with Opus 5.5: https://news.ycombinator.com/item?id=49821657

> I had it try to prepare a code review for me. Not only did it refuse, it refused to even tell me what the prompt (written by another Claude!) was. Why?

> When I had another model read the session (all of the "stupider" models handled it just fine) it explained that it had the word "reasoning" in it

> That's the entirety of Anthropic's billions of dollars of research: any prompt with the word "reasoning" is trying to hack Claude to figure out how it reasons!

> A model like that should never have gotten out of QA, let alone been released.


we have GLM flash catching Claude errors in our PR review system, costs a few pennies

I've seen the same pattern regardless of open v closed, don't have the same family that wrote the code also review the code

diversity has this way of making things better across everything humans do


How do you define "way" when saying ahead? How is this measured?

I only use open weight models now and I don't really feel a loss, curious what those who still use it think. I see output from coworkers that does not indicate Claude is that much better (still makes dumb mistakes all the time), not sure they are using the most expensive models either though.


open weight models are so far behind i cannot take your opinion seriously

When did you last use them? Are you basing this on benchmarks or daily task capabilities?

When you say ... it's hard to take you seriously

> dude were in the singularity, this opinion was cute 18 months ago

https://news.ycombinator.com/item?id=49868903


i use open weight models all the time, whenever a noteable one drops i will use it for the day, they are not even close for involved work.

sounds cursory, it takes more than a day to learn the quirks of a model, have you put similar effort into customizing / harness engineering your open weight interactions as you have claude?

you are definitely displaying strong bias that Anthropic is way ahead of everyone throughout your posts under this story

as such, I give your opinions zero weight, they don't align with the majority of accountings or my own experiences


you probably aren’t using the models to their full capacity if you don’t notice the difference

vice a versa re your usage of open weights, they are way more capable with good tools, context, process, and harness engineering

here's an example of Qwen-3.6 35B A3B MoE porting my phd code to JAX with only high level guidance from my expertise, newer qwen models share the same noticeable step change in capability as recent Big Ai models

https://github.com/verdverm/pge-jax#note-from-author

are open weights lagging, yes, are they way behind, no

if open weights were so inferior, they would not be >50% of all token processing


if open weights were so inferior, they would not be >50% of all token processing

I don't think this follows at all. Just like benchmarks get saturated, lots of tasks get saturated as well. Over time, you can accomplish a given task for much cheaper, and part of that is due to open weight models. That doesn't imply that they're competitive with frontier models for the most advanced tasks, which might represent a smaller fraction of overall work, and thus use a smaller portion of tokens.

That said, at the moment I'm finding that not much can compete with GPT-6 Luna on cost / performance (not using for coding, but for AI pipelines in my product).


these models are trash

you've definitely left rational discussion for emotional responses man, you won't persuade or convince anyone with takes like this

why are open weight models seeing such rapid rise in usage?

there has been a step function change this summer, like the end of last year for closed models

---

do you think you would experience real (legitimate) feelings of loss were you not able to chat with Claude again?

(for clarity, I am not attempting to delegitimize real feelings that real people experience, regardless of my biases, it's a question from curiosity about how others are engaging with the technology)




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: