I'm thinking of switching to Grok on Cursor (purely for $$ reasons). But Opus >= 4.8 has been fantastic; it's hard to leave, even just to dabble with other models.
Codex 5.6 sol is arguably superior to Claude, albeit very close. They're functionally indistinguishable to me, but if you're concerned about $$, Codex gives you much, much more bang for your buck.
For a personal project I’ve been piloting Spec driven development (SDD), (it’s contagious!), using Cursor and EARS statements. The strategy has been to use the frontier model to write the spec and a lessor model to write the tests and code and the frontier model to write critiquing prompts until it has nothing left to say. For my particular project there are two programs (or in human speak phases), where each program is broken up into milestones which are then comprised of a series of tasks. I experimented a lot with different models as the reviewer / spec model and the implementor model. Kimi 3 was super expensive as spec model and grok and OpenAI models always got something wrong egregiously. The Opus line of models have been the only ones to really grasp the project and I feel write great specs. Because I use cursor I settled on using groc for test routing and code implementation. I’m not sure if this is the most efficient method but I believe it’s building a large project solidly
In my experience Grok 4.5 codes at Opus 4.8 level, and being much faster as cheaper, I can just ask it to do self-review and the final reviewed code is _better_ than Opus 4.8 for the same time/budget.
But Opus 5/4.8 was better for non-code architecture discussions and general intelligence. However, for the cost, I'd use GPT 5.6 Sol and get much better results. Interestingly, Sol is not great for coding - slow and overengineer stuff if you're not explicit.
My go-to workflow was Sol for planning and Grok for building. But my in my first tests with Grok 4.6, I found it quite good and I'll start using it for both; assuming it's as good at is shows at benchmarks it's unbeatable at cost/time.
I like Grok, but I don't think that it's quite Fable-tier. It's good, but I think the position that it occupies on the Pareto frontier is a little more toward the "cheap" side and a little less toward the "intelligence" side.