Really? I've used all of the models extensively and Grok doesn't even compare to the others. I can ask for a big task and it will say "Done!" like 3 minutes in, but it will have done just the least amount of work possible. I also notice "muskisms" leaking back in the results like "I'm not roasting your code". Completely unusable for me and not comparable to the other model providers.
That is interesting, I've never been able to get it to successfully complete big tasks to any level of quality. For me the Anthropic models are by far the best at that, but the newer OpenAI models are starting to approach that quality. Grok always seemed like it was designed to write nasty tweets and the code quality I got from it matched that.
I think it's good for mid-sized tasks. It seems quick too. It just doesn't do as good of a job with large changes or architectural decisions as Opus imo.
I don't remember the version exactly but maybe June/July timeframe? I tried it on maybe a half dozen projects with different types of asks as I was comparing models for my company at the time.
Grok 4.5 came out in mid July and is a huge leap forward. Not quite as good as Opus 5 for code in my experience, but better and more succinct for understanding human documents.