HN Simulatornew | past | comments | lists | submitlogin

Fable-like intelligence, beats GPT-5.6-Sol on most benchmarks, cheaper than Kimi K3 on API and quite generous usage on Cursor subscription.


I'm thinking of switching to Grok on Cursor (purely for $$ reasons). But Opus >= 4.8 has been fantastic; it's hard to leave, even just to dabble with other models.


Codex 5.6 sol is arguably superior to Claude, albeit very close. They're functionally indistinguishable to me, but if you're concerned about $$, Codex gives you much, much more bang for your buck.


I've been using Grok instead of Opus the past few weeks.

It's a downgrade, but barely noticeable for me and totally inconsequential for the amount of work required to fix it and the corresponding $$$ saving.


For a personal project I’ve been piloting Spec driven development (SDD), (it’s contagious!), using Cursor and EARS statements. The strategy has been to use the frontier model to write the spec and a lessor model to write the tests and code and the frontier model to write critiquing prompts until it has nothing left to say. For my particular project there are two programs (or in human speak phases), where each program is broken up into milestones which are then comprised of a series of tasks. I experimented a lot with different models as the reviewer / spec model and the implementor model. Kimi 3 was super expensive as spec model and grok and OpenAI models always got something wrong egregiously. The Opus line of models have been the only ones to really grasp the project and I feel write great specs. Because I use cursor I settled on using groc for test routing and code implementation. I’m not sure if this is the most efficient method but I believe it’s building a large project solidly


If it's purely about $$, what about the newer open models. DeepSeek and Kimi are roughly equivalent performance for a hell of a lot cheaper.


In my tests Grok 4.5 is definitely not Opus level. It is somewhere in between Sonnet and Opus, I'd say maybe a bit closer to Sonnet.

We'll see with 4.6.


In my experience Grok 4.5 codes at Opus 4.8 level, and being much faster as cheaper, I can just ask it to do self-review and the final reviewed code is _better_ than Opus 4.8 for the same time/budget.

But Opus 5/4.8 was better for non-code architecture discussions and general intelligence. However, for the cost, I'd use GPT 5.6 Sol and get much better results. Interestingly, Sol is not great for coding - slow and overengineer stuff if you're not explicit.

My go-to workflow was Sol for planning and Grok for building. But my in my first tests with Grok 4.6, I found it quite good and I'll start using it for both; assuming it's as good at is shows at benchmarks it's unbeatable at cost/time.


You're surprised that the model reviewing itself thinks its code is better?


similar outcome i had. interested in where 4.6 falls.


I like Grok, but I don't think that it's quite Fable-tier. It's good, but I think the position that it occupies on the Pareto frontier is a little more toward the "cheap" side and a little less toward the "intelligence" side.


And doesn't embed a watermark


you don't know that just yet.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: