I’ve been working with frontier models for the last year on a large real-world project with multiple apps. Eventually they started becoming unreliable. Simple UI issues, usage limits and overengineering became constant problems. I had to come up with a robust process: task analysis, user intent analysis, implementation and testing. All of these steps loop when needed with 15-20 steps per task in total. This finally got the models to do a good job but I started hitting usage limits. Then I switched to DeepSeek and haven’t noticed any drop in quality. It’s fast at around 300 TPS and cheap. I don’t miss frontier models anymore! Come up with a good process and you might not need them either.
This is so true, I use frontier models at work and the variance in the day to day experience can be dramatic.
I use Deepseek for some home stuff, and it's as you describe - fast, consistent. Not brainiac level, but when you drive it well it just tends to execute reliably. I have seen it fail consistently on very hard stuff, but I'm certainly ok with having to special case occasionally. Especially if those occasions keep getting less and less numerous as the models improve.
Now I need to convince my company to allow us to use a Chinese model ... I don't think I'll have much luck.
How much do you get out of these models for same amount (say $20 sub for claude/codex/zai-glm) - the equivalent amount of tokens? Also is deepseek as good as glm5.3? (if we don't compare it with frontier models from USA)
Well, to be honest it’s all based on my subjective experience. I started benchmarking models on 10 of my real tasks ranging from easy to hard but even with the same model the completion time sometimes varied by as much as 50% between runs. That made me realize you’d probably need hundreds of tasks to get any meaningful numbers. Otherwise there’s just too much variance. So I can’t give you solid benchmarks without spending weeks testing everything properly. What I can say is that I actually use most of these models regularly for different things, so I’ve developed a general feel for them. DS is my main model for work, Claude for hobby projects and local AI, Fable mostly for design, Codex for various other tasks and sometimes Grok for health-related stuff. Cost-wise, a $200 subscription wasn’t enough to get through 20-30 tasks while $30 of DeepSeek was, and that was with Sol, not even Astra. Fable is expensive too and I haven’t found it particularly strong at coding. Opus 5 was horrible in my experience, though maybe 5.5 will be better since I’m testing it now. I also ran GLM 5.3 Flash locally for a while but it was pretty slow and its code was usually worse than both DS4.1 and Qwen Flash Next. Things are moving so quickly that it’s honestly hard to keep up, so take all of this as my general experience from actually using these models rather than a proper benchmark.
Np, good luck! PS I’ve been using Sol 6 and Opus 5.5 today and they’ve been great so far. The new Opus is smart, fast and cheap. It makes mistakes but after every task I ask Sol to review it and it catches most of them. I need to test the models properly but it looks promising so far! Hope they won’t nerf them again. I think DS may still be cheaper though (and it’s 300 TPS via the API which is sick).