For me Sol is less efficient than Astra - Sol makes many avoidable mistakes and has issues with context compaction - sometimes it goes haywire after a couple compactions.
Agreed on the “being dumbed down” observation. It appears they’re most powerful at release time and then are gradually “optimized” so every new model feels more powerful. But there’s no evidence on routing to a deployment with other weights. It would be plausible to do so though at least at peak times.
What really is fun is when the opposite of the public outcry about purported distillation happens - when Astra in Codex suddenly responds in Chinese. Now that is fun. Wonder what it’s routing to.
I have done the same. I hesitated for way too long. I shouldn’t have.
I get way more usage for way less money without any quality or performance degradation. My $200 Codex Pro plan allowance is depleted in 2-3 days. Sometimes Tibo announces a usage reset. But GPT-5.6 models are really not good for coding. Sol has been making increasingly more mistakes in the past two weeks even in the reviewer and advisor roles. Astra is usable for coding but slow and very expensive. In the past two days I’ve used up over 70% on simple copy editing, with dedicated short specs and short sessions. Really little one can do to make it more efficient. Similar work took 20% at most just a month ago. I’m looking to use Astra for milestone reviews/advisory. Perhaps a $100 Pro downgrade will be enough. But my main work is now on open-weight models. And you don’t need to depend on someone to send you a reset. And it’s cheaper by the end of the month too.
With Claude the limits are not even fun anymore - my weekly $100 Max plan quota is gone in one day on merely review invocations, no coding. And my $200 Pro quota is gone in two with some coding. Sonnet 5 is not usable for coding. And Opus 5 tends to always make a couple avoidable mistakes on every task. Fable 5.1 is ok but tends to ignore skills and to work around explicit instructions. Completely canceled all my Claude subscriptions.
With Qwen 3.8, DeepSeek 4.1 Flash, GLM 5.3 I’ve been getting Opus 5-level performance, with less blah blah and no overengineered churn. Public benchmarks are really not telling the real story. The models are more dependable and more predictable. They have their own failure modes. Sometimes DeepSeek 4.1 Flash is quite stubborn but it fails in a good way. Bad for full autonomy - I need to intervene, but it sticks to the rails and instructions - other than Fable and Opus that try to outsmart you and your harness.
Grok is interesting but has been a bit underwhelming on Grok plans - my SuperGrok allowance is depleted in a single session overnight. SuperGrok+ gives more but it’s still about the same as with OpenAI, Claude is way less now.
Since the allowance volume has been shrinking with the major model providers, to me, open-weight alternatives are really necessary now to at least maintain the momentum and budget.
But at work it’s really an uphill challenge - it’s become impossible to convince the tech leadership once they got hooked on Anthropic. They No facts will help. Some people underestimate how expensive Claude really is after getting used to the subscription plans with allowance resets. OpenAI models are expensive too.
I don’t get Claude, and that’s almost exactly what I did - I dropped to Claude Pro $20 + Codex Pro $100, and then unsubscribed from Claude and ramped up Codex. The Claude Pro is consumed within an hour on a simple task. I wish only Codex worked a bit faster than on the Fast mode.
I used to rely on Fable for research when it was first out, today it doesn’t seem to be much better than Opus, and it uses up the quota exceptionally fast - 1h Fable in a single short session, and there’s little left for Opus to hit the 5h limit in a second session. With Opus I get about 3-5h of relaxed use with a couple subagents to save the context, but there’s usually quite some disagreement between the subagents and orchestrator - Claude does some model routing with default agents and picks Haiku and Sonnet for subtasks - only later to disagree with them and redo the work - and burn extra tokens. With Claude, it’s really either Opus or Fable if you want some quality.
That said, their marketing is exceptionally effective. Virtually all nontech folks consider only Claude.
It is wild to read stuff like this when it is OpenAI that is being blamed for rug-pulling users and secretly reducing limits. It's such a big mess that Codex's product lead has been frantically posting updates on Twitter about it.
I don’t get how Claude is considered providing “unlimited” quotas. I use up my 5h on Max $100 and Team Premium in 2-3h of relaxed use of Opus 5 high+ on fresh sessions with just a couple skills/plugins. And my weekly quotas are gone in 3 days of such relaxed use. With Codex my $100 weekly quota is used up within 3 days with Sol high+ too.
Im not convinced to pay $200 for Claude’s models.
With Claude, I have to intervene every 15-20 minutes, it’s non-autonomous and it’s incredibly unreliable at self-correction. GPT is strong at self-correction but it tends to drift away from the plan to self-correct in a loop very often - a lot of tokens and time burnt on aimless churn. Opus tends to push its uninformed opinions and fake retrieval, drifting every turn increasingly farther from the intended and approved design. Opus skims over specs and makes too many mistakes.
As for closed frontier models, I prefer the GPT models over Claude’s.
I’ve started relying more on Grok, GLM, Kimi and DeepSeek models for subagents - I’ve ended up with a factory and am seeking to reduce my reliance on the closed frontier models - they’re just not SoTA on their own for development anymore.
You have to keep your session warm in cache. Keep the AI talking/thinking. If you have not touched a session for few minutes then /clear and start a new session.
Providers will generally keep your session in cache for at least 5 minutes, possibly hours. The exact cache policy depends on the provider.
If your session expires from cache then the next time you send a message you will have to pay for all the tokens you had used in context up until that point again. e.g. if you have 200k tokens in context then if your session goes cold and you send a message after expiry you will have to pay for those 200k tokens again.
With 1M contexts especially you have to be extremely careful that you don't end up resubmitting requests for hundreds of thousands of tokens again and again.
Try to get yourself and the model to use disk for medium-term context rather than model context, that way it's much easier to /clear and restart if you need to go to the bathroom or something.
Agreed on the “being dumbed down” observation. It appears they’re most powerful at release time and then are gradually “optimized” so every new model feels more powerful. But there’s no evidence on routing to a deployment with other weights. It would be plausible to do so though at least at peak times.
reply