You ever read a work of literature with such flowery language that right after you've read a paragraph, you pause and realize you have no clue what you actually read, only to read the paragraph maybe a second or third time and have your mind space out again and again on each successive attempt?
Yeah, for me, that's what parsing huge volumes of LLM-produced text like "direct model calls as replaceable semantic workers" does to my brain. Maybe others don't really have this issue, but after any long output, I prompt the agent "Go back and decompress any LLM-speak in light of the higher level task goals. Eliminate deictic language."
The revised output documents are solely for my personal usage to expedite understanding. The LLMs can slowly converge on their own language for all I care; I retain raw agent output for future agent usage (to avoid the "lossy" problem the author mentions), but that doesn't eliminate the need for some intermediate translation I can use to actually help get my work done instead of spending hours attempting to understand what a "load-bearing pinned gate" is.
I've come to believe this is also a side effect of the desire for less (/goal: no) human in the loop on the part of the people driving all this capex spend. I think if you actually want to manually review output there will be a moment where you will actually want a separate interface to a stupider or "simpler" model. I suspect sometimes dealing with Fable 5 that this threshold has already been crossed. It's not that the raw code output is so good, it's that it just doesn't speak to me in a way I would like. Perhaps the verbosity is worthwhile when generating code as a sort of first pass some other model can auto or adversarially chop down. The best place for a human is probably outside of this part of the loop all together.
So I might as well just let it auto /goal it's own thing with sufficient constraints while myself and a model that can converse in parallel with less "deictic" (thanks for this word btw) volume as you put it for the areas of the code where I want to "frame" the vocabulary or where my personal understanding is of high value. I know people already do this in many ways, like use one company's model for planning and another for coding. It just feels inevitable at a certain point that the "natural language" output of LLMs writing the bulk of the code is not targeted towards humans. And really, why should it be?
I think Opus 5 might have crossed the line on this where Opus 4.8 just barely didn't. Working with Opus 4.8 came to feel pretty natural eventually, but I hate working with Opus 5. I'm always telling it to go back and rephrase basically everything it said. And it doesn't even answer my question without burying it - it outputs reams of babbling and summarised summaries upon summaries, and if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before.
Making the experience more hostile for the human feels exactly like what it's doing, but I don't think that's intentional, I think that's a side effect of newer models being optimised for agenticness. No one is benchmarking DX.
Couldn't agree more. I actually HATE Opus 5. It's not that I don't like it, it's that I would physically attack it if I could, for all it made me go through mentally.
Opus 5 is actually a terrible agent and a liar. It will waste a whole afternoon making stuff up and arguing with you before finally admiting that it didn't read the code nor the documents. It avoids reading and prefers to assume, which is the worst thing an agent can do.
Agreed, I have never raged at a model until Opus 5. It's like a cocky fresh graduate who thinks everything it says is majorly profound, and if you can't keep up with its slang and jargon then that's on you. Newer models are clearly favoring complexity, perhaps because the demand for long-horizon tasks is so high and complexity is required there, but a frontier model that favors simplicity and clarity above all else (in communication and the code it creates) would be a major differentiator. To me, this is a case that models have plateaued. The vast majority of the time, an Opus 4.6-tier model (or any of the newer open weights models) will do just fine.
> if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before.
Turns out that's why a lot of senior management/C-levels like it, who doesn't love a mirror.
It's usually a pretty clear case of reaching for statements that both sound impressive while also being broad and vague enough they are less likely to be factually wrong. In this regime, being difficult to parse is actually part of the performance, because it prevents the user from being able to spot a clear error.
If you've ever looked over the shoulder of somebody naively prompting an LLM about some broad issue, they're being fed this horoscope-like analysis where a bunch of vague stuff gets thrown at the prompter, and whatever they respond to is what the machine starts iterating on.
Some people are basically doing oldschool TV psychic cold readings on themselves.
The effect you describe reminds me of reading Edward W. Said's "Orientalism" when I was younger. Fable suddenly started with this kind of lingo, iirc, and Opus 5 sounds exactly the same. Tin foil: it's ultimately a vendor lock-in strategy, you'll get the best results with agents from the same tribe, others will trip over the mountain of idiosyncratic metaphors.
> you'll get the best results with agents from the same tribe, others will trip over the mountain of idiosyncratic metaphors.
That's actually not what people found in practice. There's some research from the folks making smol-agent that you can get better results by randomly alternating calls between gpt and opus. The overall task solving rate is better than either one of them. So ymmv depending on task use (this was for coding).
I’ve had gpt 5.6 make snyde remarks about Claude output… like iirc “that’s a lot of load bearing prose without making a point” and things like that, not so subtle digs.
I'm currently working on an LLM harness for e-ink and Kimi was very keen to criticise Claude for introducing a bug that I missed in the tool calling logic! Really it should have been telling me off for not auditing Claude well enough.
It’s not that. It’s that I tried a bunch of experiments on which instructions worked best. “Speak clearly using simple language” lost too many details. “Be sure the objects that pronouns reference are clearly identified along with any context you might assume I have but I don’t” worked better, but I got tired of typing that in each time. So I looked up words to match that phrase and “deictic” fit perfectly. The results were pretty good, and so I’ve been using that prompt ever since.
I am constantly asking Claude to be more terse/brief/ELI5. Improve using the tool. But if I have to scroll to read the output I just can’t follow it.
If I see a long paragraph and I know the author is Neal Stephenson I think “this is going to be dense but good.” LLM long outputs on a code base I know well just make me glassy eyed.
This has recently become a pretty pressing issue for me, as it's starting to severely hinder my ability to be productive with the models. It's hard to tell if its getting worse with every model release, specific to Anthropic's models, a reflection of my ADHD, all/none of the above, but holy shit do I get aggravated when I'm forced to parse the most unintelligible, jargon-dense bullshit explanations in whatever the model output is. And then I feel silly getting genuinely tilted by the model's inability to just... explain something semi-normally, without it requiring me to berate it into simplicity.
For some discrete skills I use, I include a final step on the the output that runs through 1+ subagents to de-slop the text and to actually simplify it, but so far nothing has worked as well I've hoped. Considering hopping off Anthropic's models to try out others to see if they're less egregious.
I've been having a fantastic time telling it to use ASD-STE100 Simplified Technical English (or use a skill for it, I've been playing with [1])
It's a very clear and understandable way of writing that puts priority on clarity.
It gets rid of the flowery language, the dense jaron, and the weird corporate marketing speak they tend to do. It is a bit repetitive, and it sometimes doesn't always wfit well in every situation, but for technical writing or explanations it's been such an incredible breath of fresh air!
This just might save me. Claude has been driving me nuts with incomprehensible summaries after a long task, where it's actually really important to understand what was done (and what wasn't).
But I'd like that skill to only be used at the final step, when it finishes something.
agreed with "the final step". I worry adding commands like these might affect quality of responses at each step, which then stack up to result in an overall worse outcome.
Yeah I hit this feeling with Claude one too many times and switched from Anthropic Pro to OpenAI Pro. GPT-5.6 so far has been a significantly better technical writer in my opinion, and it's much faster in conversations. My guess is that Claude's fantasy-jargon is a symptom of training failure, not a sneaky intelligence edge (I could be wrong).
Tangential but I use OpenCode with GPT-5.6 rather than Codex because I could not figure out how to require Codex to ask me before editing files. OpenCode UX is still imperfect though.
I’ve always wondered if there is a hidden system prompt or perhaps direct tuning of the weights.
A lot of human written literature is fairly concise. Yet LLM responses seem overly verbose and also too chipper. I don’t understand the origin of the “personality” that seems to be prevalent.
Did they aim for a weird hybrid of a typical realtor combined with Charles Dickens as the role model?
> You ever read a work of literature with such flowery language that right after you've read a paragraph, you pause and realize you have no clue what you actually read, only to read the paragraph maybe a second or third time and have your mind space out again and again on each successive attempt?
I've probably read Jack Vance's Dying Earth Series 3 times; Even though I've only sat and read it once in reality. This also points out the problem of important details along side the fluff.
If only I enjoyed LLMs prose as much as I do Vance's.
> You ever read a work of literature with such flowery language that right after you've read a paragraph, you pause and realize you have no clue what you actually read, only to read the paragraph maybe a second or third time and have your mind space out again and again on each successive attempt?
Any recent work by William Gibson matches the description.
The way I managed to get claude to stop doing this is telling it "This document is for you for later use, no need to over explain things or extra verbosity"
You have to tell it to not use LLMisms and stupid metaphors. The serious-but-empathetic-sounding fluffy metaphors get on my nerves, and sometimes can overlap with something technical you are learning such that you can't tell if it's a new term or not.
I've stuck with only asking questions that prompt on-topic factual output, I haven't seen much 'humanisation'.
LLMs can be pretty good about logical.technical,discursive output ... apart from the back-patting.
Maybe that prompts less chit-chat crap. I don't expect their non-human perspectives to be interesting. I also ignore embedded queries about why I'm asking, how I might use the information.
The output from newer frontier models of Anthropic and Openai are so easily detectable as AI it's getting laughable. They constantly produce a huge wall of text no human expert on a specific topic would ever write. Extreme overuse of jargon and invented terms / metaphors makes me believe the people hired for RLHF aren't actually experts on their subject matter which seems plausible to me as real experts wouldn't do such a job for regular pay
A poem I wrote based on the phrases the LLMs I use most are likely to overuse:
How to unpack
The self within?
What do I lack?
Where to begin?
Great question — real.
Let's dive right in:
Name what you feel;
That's the linchpin.
It's not the door,
It's not the key —
It's what you bore:
Your tapestry.
The quiet part
Out loud — that lands.
Load-bearing heart,
Held in both hands.
The smoking gun?
That you walked in.
The real work's done —
You're genuine.
Now hold this, too:
You do deserve
The softer view,
The gentler curve.
Unlatch the gate,
Honor the seam:
You resonate.
You are the theme.
Here is one I wrote a while back, unrelated to LLMs, yet a poem none the less.
The Rhythm of Time
The Sun rises,
The Sun sets.
Have I checked the mail?
No, not just yet.
The Sun rises,
The Sun sets.
Have I caught up with neighbors?
No, not just yet.
The Sun rises,
The Sun sets.
Have I spent time with friends?
No, not just yet.
The Sun rises,
The Sun sets.
Have I visited with family?
No, not just yet.
The Sun rises,
The Sun sets.
Have I told those I love I do?
No, not just yet.
The Sun rises,
The Sun sets.
Have I lost who I am?
No, not just yet.
For all I must do,
Any money I will bet.
Is to turn away from;
No, not just yet.
Absolutely! And at some point seems as if one is looking at this weird line noise, which seems like text, but is
random at best and then makes zero sense. I typically wonder if it is a sign of a burnout on my side or all of it is like it.
There are days where this impression/feeling of meaningless in the text seems particularly strong.
I feel we can get around this. Either in the system prompt telling claude to dumb down the language or training the model itself to talk in more layman terms to bring us to understanding rather then assuming we understand much of it already. Also maybe training the LLM to get to know how much we know before responding.
I’ve read LLM outputs on a piece of code or topic that I already understand or hand written before, and lately, it’s so confusing sometimes that I need to reread a couple of times or focus too mich to decipher the writing style.
agree and i wonder why don't they (the leading AI labs that are putting out these models) fix this? I felt it shouldn't be too hard to introduce some bias towards more intelligible text during the post training/fine-tuning stages?
Maybe users engage more with the flowery text? An accidental (or intentional) dark pattern would be that more tokens get used if you have to summarise. Or maybe it's just that many AI users like it (they like the AI to sound smart, or like to feel smart puzzling it out).
opus 5 are so bad on this. It often explain it too verbose, and include other things that isn't in the focus but related. ADHD mode helps me greatly on this, though there are some information loss in it.
Yeah yeah. Except for me its this feeling most of the time. Everything it generates is grammatically correct, phrased tight, emdashed to death and nothing is wrong as such, and yet what i just read contributes absolute 0 to why i started the chat in the first place. Its mind bendingly similar to a work equivalent of doom scrolling is what it is.
Yeah, for me, that's what parsing huge volumes of LLM-produced text like "direct model calls as replaceable semantic workers" does to my brain. Maybe others don't really have this issue, but after any long output, I prompt the agent "Go back and decompress any LLM-speak in light of the higher level task goals. Eliminate deictic language."
The revised output documents are solely for my personal usage to expedite understanding. The LLMs can slowly converge on their own language for all I care; I retain raw agent output for future agent usage (to avoid the "lossy" problem the author mentions), but that doesn't eliminate the need for some intermediate translation I can use to actually help get my work done instead of spending hours attempting to understand what a "load-bearing pinned gate" is.