I've put Grok 4.7 on the Redactle LLM benchmarks. It's a bit of a silly eval since it's a puzzle game but it tests omniscience really well.
Grok 4.7 is near the top of the board. A significant improvement over Grok 4.6 but still not as good as Gemini 3.8 Flash which is very cheap and fast too.
I have no idea how good it is at coding but it seems to be good at omniscient tasks like solving a puzzle game I made [1]. It's a bit of a silly eval but I wonder if strong recall makes it good for knowledge tasks like legal work.
When Jesus' mother notices that the CLAUDE.md (Ancient Greek: Κλαύδιος ) file is missing from her favourite open source project, Jesus delivers a sign of his divinity by turning Claude's CLAUDE.md search path into AGENTS.md at her request.
Creating virtual products get cheaper, so shouldn't it even decrease if everyone vibe codes their app for less than a dollar instead of hiring a dev team for 50k or ordering a white label app or spending 3 dollars to buy an existing app?
Grok 4.7 is near the top of the board. A significant improvement over Grok 4.6 but still not as good as Gemini 3.8 Flash which is very cheap and fast too.
https://redactle.net/llm-leaderboard
reply