Falling back on a RL pipeline to cover gaps always strikes me as a more sophisticated version of the mechanical turk. If what we had was truly AGI wouldn't they be able to derive this from the data they already have.
While I agree that a coding model (such as Opus) by itself tends to act very shallowly, when it's driven by a harness like Claude Code, the combination seems to be a far more general thing than a LLM. It's capable of consistently making excellent data structure and architectural choices over large code bases. It imitates thinking about anything and it can drive itself for hours.
Honestly, if I simply fed it a sense of presence (I would repeatedly tell it what's going on right now and ask it to react if it thinks it should), it would feel eerily like AGI.
Coding models can already make well-considered data structure choices if given all the relevant context, but a non-programmer doesn't know the context to give it or even to tell it to optimize the data structure choice.