I've come to believe this is also a side effect of the desire for less (/goal: no) human in the loop on the part of the people driving all this capex spend. I think if you actually want to manually review output there will be a moment where you will actually want a separate interface to a stupider or "simpler" model. I suspect sometimes dealing with Fable 5 that this threshold has already been crossed. It's not that the raw code output is so good, it's that it just doesn't speak to me in a way I would like. Perhaps the verbosity is worthwhile when generating code as a sort of first pass some other model can auto or adversarially chop down. The best place for a human is probably outside of this part of the loop all together.
So I might as well just let it auto /goal it's own thing with sufficient constraints while myself and a model that can converse in parallel with less "deictic" (thanks for this word btw) volume as you put it for the areas of the code where I want to "frame" the vocabulary or where my personal understanding is of high value. I know people already do this in many ways, like use one company's model for planning and another for coding. It just feels inevitable at a certain point that the "natural language" output of LLMs writing the bulk of the code is not targeted towards humans. And really, why should it be?
I think Opus 5 might have crossed the line on this where Opus 4.8 just barely didn't. Working with Opus 4.8 came to feel pretty natural eventually, but I hate working with Opus 5. I'm always telling it to go back and rephrase basically everything it said. And it doesn't even answer my question without burying it - it outputs reams of babbling and summarised summaries upon summaries, and if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before.
Making the experience more hostile for the human feels exactly like what it's doing, but I don't think that's intentional, I think that's a side effect of newer models being optimised for agenticness. No one is benchmarking DX.
Couldn't agree more. I actually HATE Opus 5. It's not that I don't like it, it's that I would physically attack it if I could, for all it made me go through mentally.
Opus 5 is actually a terrible agent and a liar. It will waste a whole afternoon making stuff up and arguing with you before finally admiting that it didn't read the code nor the documents. It avoids reading and prefers to assume, which is the worst thing an agent can do.
Agreed, I have never raged at a model until Opus 5. It's like a cocky fresh graduate who thinks everything it says is majorly profound, and if you can't keep up with its slang and jargon then that's on you. Newer models are clearly favoring complexity, perhaps because the demand for long-horizon tasks is so high and complexity is required there, but a frontier model that favors simplicity and clarity above all else (in communication and the code it creates) would be a major differentiator. To me, this is a case that models have plateaued. The vast majority of the time, an Opus 4.6-tier model (or any of the newer open weights models) will do just fine.
> if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before.
Turns out that's why a lot of senior management/C-levels like it, who doesn't love a mirror.
I've come to believe this is also a side effect of the desire for less (/goal: no) human in the loop on the part of the people driving all this capex spend. I think if you actually want to manually review output there will be a moment where you will actually want a separate interface to a stupider or "simpler" model. I suspect sometimes dealing with Fable 5 that this threshold has already been crossed. It's not that the raw code output is so good, it's that it just doesn't speak to me in a way I would like. Perhaps the verbosity is worthwhile when generating code as a sort of first pass some other model can auto or adversarially chop down. The best place for a human is probably outside of this part of the loop all together.
So I might as well just let it auto /goal it's own thing with sufficient constraints while myself and a model that can converse in parallel with less "deictic" (thanks for this word btw) volume as you put it for the areas of the code where I want to "frame" the vocabulary or where my personal understanding is of high value. I know people already do this in many ways, like use one company's model for planning and another for coding. It just feels inevitable at a certain point that the "natural language" output of LLMs writing the bulk of the code is not targeted towards humans. And really, why should it be?