I feel like I'm completely missing something with ARC-AGI. The tasks are so limited in scale and very black and white, which do not at all map to real-life challenges.
I do think it's impressive that LLMs can reliably solve them, and I recognize LLMs are getting much better at navigating more ambiguous and expansive tasks. But I'm not impressed by any person who can solve ARC-AGIs, and nor would I even look down on a person who couldn't solve them all. I'd certainly never consider ARC-AGI results when deciding whether to hire someone.
Anyone taking a single look at the ARC-AGI "challenges" can see things a 5 year old could reasonably solve.
They're just jerking eachother off and sending eachother the elevator back: "independent" ML engineer (worked at and currently runs Every single benchmark has been catastrophically flawed and made by clowns.
> Anyone taking a single look at the ARC-AGI "challenges" can see things a 5 year old could reasonably solve.
Isn't that the goal of these challenges? Each release shows challenges that are very easy for humans, but are impossible for the models at the time of release (which demonstrates some missing generality).
I think I've read the challenge authors say that, the day they cannot make a new challenge, then models are AGI.
I do think it's impressive that LLMs can reliably solve them, and I recognize LLMs are getting much better at navigating more ambiguous and expansive tasks. But I'm not impressed by any person who can solve ARC-AGIs, and nor would I even look down on a person who couldn't solve them all. I'd certainly never consider ARC-AGI results when deciding whether to hire someone.