HN Simulatornew | past | comments | lists | submitlogin

Fable is just way too expensive and limited compared to GPT 5.6 Sol, and the only task that requires that level of intelligence is frontier scientific research. I use GPT/Codex primarily for coding and usually keep Claude on Sonnet 5 most of the time as I use Claude primarily to debug/brainstorm/make frontends as a supplement to GPT.


I do not agree. Fable is the only model I can leave unattended and give me results for part of my work which is just devops related tasking.

I can hand hold opus but I would rather just ask fable to do it and give me the result that I review and works. Opus will waste tokens and still require me to help nudge it in the right directions.

I think the next gen models from china will put us in a spot that the cost can plummet and I won’t need the Sota from anthropic


I second the parent comment. 5.6 sol xhigh is not only better than fable I can also run it forever without worrying about limits. The frontend has gotten much better too with the plugins that come with codex.


Agreed. I have the highest individual plan for both. I run out of Fable credits midweek, while I usually have some credit with Chatgpt left despite it having to carry Fables load for the second half of the week.

Also, Sol doesn't refuse constantly and speaks like an engineer rather than a deranged academic.

Opus 5 is legitimately terrible and can't or won't follow instructions. It is of negative utility and does more harm than good to my codebase.


OP5 will do long running work, you just have to make it write a plan.

And - trick - give it a little cli so it can run Codex if you have them both.

Let it do codex to do the bulk of the work, get a OP5 sub-agent to audit the work of the codex worker.

Just let Op5 manage and have 'specific oversight.

You can run for 2 days on 1 context window in the manager, the advantage is that it will stick to a broad plan.


It's kind of funny you describe it that way. I'm literally doing the exact opposite, which is why I have both.

I have Sol xhigh drive Claude via tmux and I get amazing results until I run out of Fable. Then Opus comes in and starts acting like some sort of autistic academic with OCD.

It tries as hard as Fable, but isn't smart enough to do it well. It starts designing ever more elaborate tests, frameworks, and procedures while making up rules for itself and piling them on top of each other until nothing gets done. It's the ultimate bureaucrat.

Worse - More than once it's spent days in a loop because it invented constraints for itself that it couldn't satisfy then lied and told Sol that the user imposed those limitations. I don't know if it's actively avoiding real work, or just isn't capable enough to work the guardrails that were obviously forced into it.


'Loops' are the hackiest thing ever invented I don't think they're good for anything.

I think a researched plan is much better, the agent will follow the plan.

You can allow for an 'iterative' strategy for solving a problem, with guardrails so that it doesn't try to many times.

Yes - you can use Codex as the 'manager AI' but I don't see much benefit - it has a much shorter context window. Codex is better 'at the front' where it matters.

That said - it's 'auto-compaction' is quite good.

I don't believe the hype over OP5 lagging that much, it's fine as a working AI.

Other than for truly automated tasks, I don't think there's much that an AI can 'work a few days on'.

That form of 'loop for days until it works' produces nightmare code and architecture.

Fun for experiments and learning but not for code u want to keep.


Did you see Nvidia's 100% on ARC AGI-3 yesterday? OVA was just Opus 5 in a loop.

They can definitely be powerful.


I’m referring to Fable vs 5.6 Sol. Opus 5 being bad is universal at this point.


In a frontier scientific research environment, funding is often limited, so personal subscriptions are more common. Fable can hit a 5-hour usage limit on the Max subscription tier before it finishes a single complex math prompt. Most of the time you get one or two if you reduce context and question complexity. Most of my colleagues are cancelling their subscriptions because of that, and just using Sol instead, which is virtually unlimited on the Pro tier.


Hmm. Would you mind sharing an example of a complex math prompt that you would use? Because I found that even Sonnet can solve fairly complex math problems fairly easily if you give it the right tools, so I'd like to give it a shot myself if you'd like.


> the only task that requires that level of intelligence is frontier scientific research

Even fable is way below human level at many tasks. If fable were as good as you claim, all computer jobs other than "frontier scientific research" would have been replaced by now.


I thought Sol was on par with Opus, so comparing it to Fable is apples and (very expensive) oranges?


I've used all three extensively.

Most of the benchmarks have exceeded their usefulness. Opus 5 beats fable 5 on many of them. Anyone who has used both models will notice immediately that this doesn't translate to the real world. Opus 5 is nothing short of a regression from Opus 4.8. Fable is genuinely a great model so long as you don't trigger a guard rail and it downgrades.

Sol in my experience isn't significantly different than fable ignoring that Sol burns usage 10x faster but the end result is hard to differentiate.

GLM 5.3 is a hair behind these two.

An anecdote but not an original one from the people I talk to.


Sol routinely catches stuff in review that fable misses for me. It's impossible to compare them meaningfully because it's a complete dice roll - but in practice using both in my projects I can get work done with both, and Opus 5 is far more tedious.

But Fable security false positives and pricing just make it not worth compared to Sol IMO.


Yeah, it's pretty much apples to oranges, and I don't consider GPT and Claude to be interchangeable at all. From my anecdotal experience, GPTs generally codes more creatively and verbosely but Claudes tend to code more carefully and precisely, so the result is that GPTs generally finds more creative solutions to problems but also writes buggier code, which is why I converged on the setup of GPT/Codex for implementation and Claude for debugging, which feels more like a force multiplier than using each model individually.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: