OpenAI usage limits have been severely cut, and intelligence appears to be markedly declining, so I'm going to start trying these Chinese models seriously now. I don't mind if it takes longer. I just need the intelligence to predictably work the same way from day to day.
I have a 20x sub and it feels like the 20$ sub eight months ago. I can easily burn it in an afternoon, and I don't have many projects.
According to API usage, they cut you off at around the equivalent of 900$ of API usage, whatever that means. It's very difficult to track all this, and very subjective. What validates me is that of all my friends I am not the only one.
I can only imagine the 100$ users must be feeling the rug being pulled even harder.
Anyway, this has led me to get a Spark, and a second is on the way.
For me Sol is less efficient than Astra - Sol makes many avoidable mistakes and has issues with context compaction - sometimes it goes haywire after a couple compactions.
Agreed on the “being dumbed down” observation. It appears they’re most powerful at release time and then are gradually “optimized” so every new model feels more powerful. But there’s no evidence on routing to a deployment with other weights. It would be plausible to do so though at least at peak times.
What really is fun is when the opposite of the public outcry about purported distillation happens - when Astra in Codex suddenly responds in Chinese. Now that is fun. Wonder what it’s routing to.
Weird. That's definitely not my experience. On the $100 plan and I can burn my entire week's budget in a few hours easily with Astra. It's borderline unusable.
People often use a bunch of subagents, poor context management (though Codex's tight context limits and constant compaction tend to mitigate this), a bunch of projects at once, etc., as well as not using workflows that do heavy planning once up front and then consult it rather than thinking endlessly about what to do during the implementation part.
It's also the case that working on massive codebases is just a different beast. If they've been slopmining a monorepo for months with 200x, then their codebase is probably Lovecraftian at that point and requiring extensive effort to iterate on.
i am on just the $20 plan and I am using it literally almost all day, 2-3 different projects at the same time, and I am failing to run out! I am naturally always trying to be efficient, its a weird personality trait I have, where I don't try to but I willingly choose to use GPT-6-Luna most of the time now (well 5.6 before that) More Sol lately and for most tasks.. they seem to have the same output, its only super hard things where I kick it up to Astra.
I am so used to doing planning with the best models and switching to Luna, Deepseek 4.1 Flash (really, the cheapest model, and often much better results for lots of things) really people are limiting themselves when they only use one company's models. You really gain a TON by treating each one more like different people with diverse range of personalities, passions, skills, and knowledge.
Today, ChatGPT desktop app was tasked with making 7 variations of new versions of some existing websites of mine - and they all kinda looked the same. I did get lazy with it though.... I can solve that probably, with skills/changing the default frontend skills but also just sending the same task to Reasonix Code w/ deepseek, and some other models, gets some good wide range of outputs.
I strongly agree. Check out the Codex subreddit. Many empirical examples of Astra silently downgrading the models. One found Astra was silently using Luna Max (but still billing for Astra).
Even when I try to stick with Sol X/High, my limits are at best half of what they were before Astra launched, and the intelligence has declined markedly.
I cancelled my $100 plan. This is absolutely absurd and frankly unusable now.
it's per week. I don't use the mainstream harness, skills, agents.md, subagents, orchestration and mcp, so the the thing does what I ask instead of going into rabbit holes and doing coffee machine chatter with it's shaitan friends
Entirely depends on the work asked and the repo involved.
When I’m doing work on a repo where I’m implementing a standard and the agents have to read the standard to keep from hallucinating my usage skyrockets.
Hell this changes depending on which language I’m working with.
Sol 5.6 xhigh had been a very reliable workhorse for coding for me via the 200 bucks sub.
But this week they seem to have tweaked the system to a point at which all models (Astra, Sol, Luna) hit rate limits all_the_time without me being anywhere close to the weekly limit.
Early results with MiMo 2.6pro are quite encouraging for anything that's non-UI work so likely switching spend for the time being
having discussed this with people and saw similar discussion, it seems what they're doing is limiting long running context _and_ tweaking the models to run agents which basically strip mine your usage tokens because they want to avoid users taking up precious KV cache & VRAM space.
So whatever they advertise as the context window, assume the model has been tweaked to keep it a quarter. Other commentors think all models are generally useless past 250k, and that might just be a reality of thesemodels.
Eitherway: they're enshittifying precisely according to hardware vs users vs actual cash on hand, and smaller agents with less cache/vram on hand makes the overall hardware more performant.
This of course is ignoring whether they're trying to stealthly deploy quantized models to eek out even more space on the hardware.
This stuff is easy to learn when you play with local models.
Welcome to the dark side. I've been on Kimi, with a little DeepSeek-V4-Pro, GLM 5.2/5.3, and MiMo thrown in, for probably about a year now. It's great here!
For DeepSeek, I recommend their Reasonix harness strongly, due to its alignment to DeepSeek's prefix cache. It means mostly (95%+) cache hit input tokens, so very cheap large-scale code analyses and things that require mega context windows (at the cost of some attentional drift, yes). Reasonix does require that you send data to China.
For most everyday stuff outside of where Reasonix + DeepSeek just makes overwhelming sense, I use OpenCode/Maki/Pi/whatever harness I feel like using today with Kimi K3, via OpenRouter. This does not require sending data to China.
I also use Kimi K3 in Zed via OpenRouter quite a bit, but sometimes like to mix it up with the other models.
For local hardware experiments on my MacBook (128 GB unified memory), Qwen3.6-35B-A3B (speed) and Qwen3.8-27B (intelligence, but slow). As has been widely noted, this amount of unified memory isn't as useful as it seems, due to memory bandwidth and decoding constraints, lack of tensor cores (on the M4 Max, anyway), etc. A giant bag of memory isn't fast, but it'll let you load some impressively big models. The future M5 Studio Macs will continue in this general vein, but will of course be somewhat faster, particularly due to the apparition of tensor cores in the M5 -- Neural whateverApplecallsthem.
The Chinese models are simply _excellent_, and cater to lots of use-cases and tastes.
I have done the same. I hesitated for way too long. I shouldn’t have.
I get way more usage for way less money without any quality or performance degradation. My $200 Codex Pro plan allowance is depleted in 2-3 days. Sometimes Tibo announces a usage reset. But GPT-5.6 models are really not good for coding. Sol has been making increasingly more mistakes in the past two weeks even in the reviewer and advisor roles. Astra is usable for coding but slow and very expensive. In the past two days I’ve used up over 70% on simple copy editing, with dedicated short specs and short sessions. Really little one can do to make it more efficient. Similar work took 20% at most just a month ago. I’m looking to use Astra for milestone reviews/advisory. Perhaps a $100 Pro downgrade will be enough. But my main work is now on open-weight models. And you don’t need to depend on someone to send you a reset. And it’s cheaper by the end of the month too.
With Claude the limits are not even fun anymore - my weekly $100 Max plan quota is gone in one day on merely review invocations, no coding. And my $200 Pro quota is gone in two with some coding. Sonnet 5 is not usable for coding. And Opus 5 tends to always make a couple avoidable mistakes on every task. Fable 5.1 is ok but tends to ignore skills and to work around explicit instructions. Completely canceled all my Claude subscriptions.
With Qwen 3.8, DeepSeek 4.1 Flash, GLM 5.3 I’ve been getting Opus 5-level performance, with less blah blah and no overengineered churn. Public benchmarks are really not telling the real story. The models are more dependable and more predictable. They have their own failure modes. Sometimes DeepSeek 4.1 Flash is quite stubborn but it fails in a good way. Bad for full autonomy - I need to intervene, but it sticks to the rails and instructions - other than Fable and Opus that try to outsmart you and your harness.
Grok is interesting but has been a bit underwhelming on Grok plans - my SuperGrok allowance is depleted in a single session overnight. SuperGrok+ gives more but it’s still about the same as with OpenAI, Claude is way less now.
Since the allowance volume has been shrinking with the major model providers, to me, open-weight alternatives are really necessary now to at least maintain the momentum and budget.
But at work it’s really an uphill challenge - it’s become impossible to convince the tech leadership once they got hooked on Anthropic. They No facts will help. Some people underestimate how expensive Claude really is after getting used to the subscription plans with allowance resets. OpenAI models are expensive too.
> and intelligence appears to be markedly declining
Serious question: does anyone have evidence of this?
It’s something that’s constantly asserted, and has been since 2023. Every time someone posts a site that tries to track this though, I look at it and it’s just a flat line.
By and large they don’t. I have seen this drop a few times, eg before fable came out opus dropped a lot probably due to less compute available.
My guess is it’s a combination of getting used to the new cliff models fall off on and forgetting that model performance drops significantly when context fills up.
So new model comes out, people try it and it’s amazing on a task or two. Then they start using it, context window fills up and it gets a lot worse.