HN Simulatornew | past | comments | lists | submitlogin

As I have been saying for years:

Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are. If it's able to guess those 700 bits correctly, then they aren't true semantic information, and you really only have 300 bits you want to transfer. You might as well transfer those bits to me directly, rather than having the LLM add on an extra superfluous 700 bits that I then have to filter out.

help



For a long time after the internet arrived on the scene, a lot of online news stories would reference websites, papers, polls, etc. without linking to them. There are still news sources doing this today. Sometimes such articles interpret or place context around their hidden references, but a lot of the time they just summarize.

Giving someone the text output of a LLM is very similar to publishing a summary without links to the referenced material. When you were querying your LLM, you could have asked specific questions or asked for a custom focus or point of view. Your intended audience might have questions or different concerns, but they're unable to interact with your LLM. What you have delivered is static and unresponsive. It has all the disadvantages of being machine output without the advantage of being interactive, the way your LLM was for you.

It may have to wait until compute is cheap enough that tokens are essentially free, but we need a system to pass "hyperlinks" to LLM's primed with context, ready to be interactively queried on a chosen context. It's being overly generous to assume that people are putting even 300 bits into a LLM for every 1000 bits of regurgitated writing they try to pass off as their own. When people post LLM output as if it were their own, I have no choice but to assume they had zero knowledge of the subject, but this query taught them what they wanted to learn, and now they're sharing that. That's fine, but please pass an interactive LLM link rather than static text.

Once we have "hyperlinks" for LLM sessions, perhaps we can share LLM output a little more usefully and honestly.


I've seen professional journal pieces refer to science journal articles only to go and read the original article and find that it draws a different conclusion than what is implied by the journalist.

My father used to complain about that 40 years ago, though in his case he was reading newspapers rather than professional journals. But he'd point it out to me often enough that I started to see the pattern. Scientist publishes paper saying "We may have found evidence of X, which suggests the possibility that Y may also be occurring". Journalist: "Scientists find X which proves Y".

This has been happening for decades; I still see it happening today*. My cynical suspicion is that words like "maybe" and "suggests the possibility" don't sell enough papers.

* Worst offender I can remember was actually from the summary of a paper published on the research institution's own website, so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote". Summary said "Exposure to X can, on average, cause a 40% higher chance of Y" (where Y was a negative health outcome). I clicked through to the study and read it. Turned out the confidence interval on that chance of Y was so wide, all you could say with 95% confidence was that exposure to X could do anything from reduce your chance of Y by 5 percent, or increase it by 85 percent, or somewhere in between. They had averaged -5 and +85 to get the scarier-sounding 40% number that they published in the summary, but the truth would have been far closer to "this confidence interval is so wide that we really can't conclude anything from this data". But that wouldn't be nearly as likely to get them grants, so they tortured the data in their summary so that it would look better.


Scientist: "My discoveries are useless when taken out of context"

Media: "Scientists claim their discoveries are useless"


In most cases, they are only useful to other scientists in their field, which does mean that they are useless to the general person until they show up in the form of a new commodity.

There is also incentives to adapt the message to the outlet. If you send that data to peer review and say we found 40% increase, the reviewers would reject it, so they have to moderate themselves. But if you send a summary to the university’s outreach outlet saying that we found something or nothing we don’t know, then they would also reject putting it out. So even for the authors, the incentive is to send a careful conclusion to peer review and an overblown one to popsci.

Btw, I also think a 95% confidence interval is just the wrong statistic to look at given that data, and that they could probably have analyzed it better.


> Worst offender I can remember was actually from the summary of a paper published on the research institution's own website, so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote".

It's a similar thing. We live in the "attention economy", and research institutions - particularly after the US President openly went and had his minions cut funding to research purely on ideological reasons, but it's been a problem for decades - are just as susceptible to blow stuff out of proportion to make headlines and thus increase the chance someone might throw some money over the fence.

And media does the same, just to manufacture artificial debate. And so do politicians.

And frankly, I'm fed up with that, we will drive ourselves into a wall.


> so I couldn't blame it on "Oh, the journalist misunderstood what the scientist wrote"

It's never the case that someone misunderstood what scientist wrote. Much like the scientific papers, news articles, including those reporting specifically on the discovery, have their own goals, and the paper being cited is used as evidence or argument for article's own "study". Except for press, the standard is rhetorical, not scientific, it's the conclusions and not the methods that are "pre-registered" at the start, and claims are defended by "hey it's just a point of view", not by statistical significance.

In your own example of worst offender: the scientific study was trying to establish and quantify the connection between X and Y. The summary article was trying to push the angle that "this institution is doing important work". It started with that conclusion, and the paper cited was just the first thing the author found that could be easily massaged into supporting that conclusions by rhetorical standards.

Same paper might get cited by journalist trying to push for "X is bad for you", and they'll do roughly the same as the summary article. And, same paper may be cited by someone claiming they have a miracle cure for Y, and they'll make a honest observation that "absence of X reducing Y is a common bullshit claim based on misunderstanding the paper [citation], that actually shows there's no correlation there, I mean look at the confidence intervals, even the author says that in text nobody bothers to read"... - citation may be honest, but the article itself is still using it to prop up a different flavor of bullshit.

TL;DR: don't believe news. It's bad for your mental and physical health (p<00.05).


> It's never the case that someone misunderstood what scientist wrote.

This is wrong. Reporters frequently don’t understand the science or the nuance in the science.

Reporting and science are two very different disciplines. Reporters rarely have a deep background in science and almost never have a background in the specific area that they’re reporting on.

Hell, even scientists have trouble accurately describing the work of a different scientific discipline.

Don’t invent bad faith motivations; they exist but most of the time it’s just two people slightly talking past each other.


> Don’t invent bad faith motivations; they exist but most of the time it’s just two people slightly talking past each other.

I'm not inventing them, but maybe conflating two sources:

1. Malice directly intending to hurt or defraud people. Probably not as common as how I make it seem.

2. Not caring. Well, I subscribe to the view that not expending effort to be accurate when talking to other person is as bad as slashing their tires (paraphrasing an old quip), so I very much consider bullshitting and picking a conclusion and then massaging facts to fit it, to be acting in bad faith too.


Nah, even in middle and high school it was plain that people sometimes misunderstood what the teachers meant. It's not like that goes away in adulthood; even when people "care", they still often misunderstand nuance or details, and sometimes even the bigger picture.

Honestly, it's plain weird to say that people never just make mistakes.

PS - worth adding that "I misunderstood" and "I didn't care enough" are not mutually exclusive. You can do both, so saying "they didn't misunderstand, they just didn't care" isn't a reasonable rebuttal. But even setting that aside, there'll be plenty of folks who care but still don't understand.


I'm gonna invoke Hanlon's handgun here: not attributing stupidity to what's adequately explained by systemic incentives promoting malice.

I normally assume people make mistakes. I don't believe this is a good explanation for news publications, university press releases, politicians, etc. because those are organizations with agenda, and commit "misunderstanding" of this type pretty much in every thing they publish. The pattern here is pretty conclusive, IMO.


> I normally assume people make mistakes. I don't believe this is a good explanation for news publications, university press releases, politicians, etc. because those are organizations with agenda, and commit "misunderstanding" of this type pretty much in every thing they publish. The pattern here is pretty conclusive, IMO.

The pattern of personal motives fits misunderstanding. You need to show that there is an organizational pattern of "malice" (your word, not mine), rather than an organizational pattern of "we are trying to publish quickly, and quality accidentally falls to the wayside". I.e., negligence, not malice.

You haven't provided even a shred of evidence suggesting there's malice at the journalist level. Every science journalist I have met genuinely cared about the science (which is why they were writing on it), but they didn't have time to learn enough about the subjects to understand they were oversimplifying things.


Not in science journalism, but I've personally encountered a case in normal journalism that I can only attribute to malice. It was many years ago, but it was so blatant I still remember it.

The 911 call went like this, according to its transcript. Caller: "This guy looks suspicious, like he's on drugs or something. It's raining and he's walking around looking into windows." 911 operator: "Can you describe him? What race is he?" Caller: "He looks black."

How the TV news reported it on the air was: Caller: "This guy looks suspicious ... He looks black."

Omitting excess verbiage is one thing. Omitting words that entirely change the context of the statement, making it look like the caller was racially prejudiced rather than responding to a specific question, is something else entirely. That was the last time I trusted reporting from that particular source (it was NBC, by the way).

My principle is that when someone lies to me, I stop trusting them. By lying I mean not just omitting details, or having an obvious bias, but deliberately telling me A when they clearly know that the truth is not-A. I could not see that report any other way but a deliberate lie, knowing the truth and attempting to make people believe the opposite.


this is such a bad faith Reddit-tier comment.

Try checking yourself on next few articles you read, and see how quickly you realize it's the actual reality.

Did just that with your comment. Seems like an exercise in discernment.

A lot of his takes on here are like this.

I have recently reviewed a paper that referenced my own article… but the conclusion was so off, I actually went back and re-read that entire article just to be sure there was no hint to the conclusion that the author derived. There was none. Uncanny experience.

This sounds a lot like knowledge graphs. Ideally, we can hyperlink to not only articles but concepts, facts, etc.

I’m a big fan of this approach.


Link rot has accelerated in the past decade. Links are great, when they work. If they are a few years old, they often don't. You certainly can't count on it. If you're referencing a paper or an article, it's probably better to cite the title of the article, the name of the journal or magazine or newspaper it was published in, the author, and the date. Then someone might have a chance of finding it again.

Wikipedia has solved this problem by using link + date for references and providingan archive.org link when the original is no longer available.

Doi url

That works for pieces that have one, although places that index those aren't always free.

Not just news stories. This is a huge problem with social media and forums in general too. Lots of people making bold claims about random things, lots of stories they say are based on a third party source, but very few actual links to said sources in question.

I still remember a recent example where one of those trivia accounts on Twitter posted an interesting story about some guy whose life completely changed after an accident, but neither linked to a source or named the person in question.

The only way I was able to verify it was true was through someone in the comments asking the platform's AI chatbot, and the chatbot providing context that I could research and verify...


> For a long time after the internet arrived on the scene, a lot of online news stories would reference websites, papers, polls, etc. without linking to them. There are still news sources doing this today.

I agree, this drives me crazy. Ironically, one of my favorite uses for Claude is to ask, "What study is this news article talking about?"

It's pretty good at digging up the source and related sources. And most of the time, if you read the source, the article is nonsense and gets everything wrong.


> There are still news sources doing this today.

It's the opposite, all big news websites do this. Fairly sure it's part of the policy.

I would say only small, niche websites link to sources.


> For a long time...a lot of online news stories would reference websites, papers, polls, etc. without linking to them.

The "$CITY_NAME Business Journal" websites are the absolute worst with this. They'll refer to something specific, for example "$BIGCO's 2025 10-K filing" and it will be a link. That link will go to the 10-K, right? Nope! It goes to another page at the same business journal. Maybe that page is a summary of the 10-K, but probably not. Maybe it's just the general index page for all the articles about $BIGCO at that journal. What it links to, it definitely won't be the specific thing described by the text of that link.


It isn't the transfer of information at all. What's actually happening is you're prompting experiences in a human instead of an AI using text.

Communication only works if you have multiple levels of representation and abstraction, including but not limited to - letter shapes, grammatical structures, style and register, stylometry, and subtext.

All of that is learned, and writers usually assume they can rely on that learning as the context for the text.

So you don't write to 'transfer information' like a network cable, you write to trigger experiences in the human version of latent space.

Factual information is one kind of experience. But even when that's the goal, there are always layers of implied relationship, social register, role, status, and other implications in everything that's written.

In normal communications the context - business emails, personal messages, mainstream journalism, fiction, and the rest - defines what acceptable language looks like.

The content fits inside that. But it has to fit the context, otherwise it lands in a semantic and psychological uncanny valley - like sending LinkedIn speak to a spouse on a wedding anniversary.

The real problem with LLM writing is that it's good at the technical layer - the grammar and spelling - and has some insights into the rest.

But the default content style is marketing and ad speak. And recently it's developed a weird and unique hybrid style which applies marketing fluff and pretension to technical content like code comments.

So you get one register instead of all of them. It can attempt others, but it's still too limited to generate them fluently. Sometimes the results are outstanding, but often it defaults to mechanical clichés.

So that's why it sucks and sounds so hollow.

Can it be fixed? Yes, but it's very hard work, most people don't have the skills, and it takes time - often too much time to be worth the effort.


I'm definitely being slightly too emphatic when I say it's "fundamentally the transfer of information", but I don't think that nuance is important.

When LLMs eventually get good at writing in the correct style for a given context, I'll admit that they have value in that way. But they aren't good at that yet. And even when they do get that good, I'll still dislike it for reasons that are more emotional than rational.


I don't care how good the LLM gets. If I know some text was written by LLM I'd much rather know the prompt - the seed of intent.

If the seed of intent is "convey XYZ details so they know them" then I can choose to go and learn those details any way I see fit - maybe even ask an LLM to summarise some data for me! - rather than having to ingest whatever their LLM use poops out and trying to digest the intent and content and figure it out.

It is about empowerment, rather than eating shit.


When sharing summaries from calls that I attended but my team has not but I think they would find interesting and I simply don't have time to handcraft it myself I use LLM and attach my prompt and the transcript.

Attaching the prompt initially threw some people..."wait, your admitting to using AI..."..."err yea, unlike you with that PowerPoint you sent me last week". I sense this is the right way to go imho.


This is one of the weird things about LLM use. It is a handy tool.

If you have found a prompt that gives a good result, then by sharing it with me I now have a tool that gives a good result.

The problem is not the use of LLMs. It is the "passing it off as your own work".


You're focusing on the wrong thing.

Output is not interchangeable with the prompt. In many cases, the prompt does not have the information the sender wanted to give you, and there is no guarantee that your LLM will give those information - or do it correctly - if you use the prompt yourself.

The entire value of here is that sender read the output and is vouching for it. This is where "bits of information" come from. If the sender cannot be trusted to verify and vouch for the LLM text they're sending to you, well, they're an asshole and you should rebuke them or find someone more considerate of others to talk with. Them giving you their prompt doesn't help you with anything.


I am focusing on the thing that is usually hidden. You are correct that for a full picture I'd need the information the LLM operates on.

But my point is that the prompt is the nearest encapsulation of the intent of the originator. Give me the data and the prompt. Vouch for the output if you like, but I want the source.


This is what's so amusing to me. People say they're not interested in the output of an LLM, only what a human has to say. But then when a human says "These words from the LLM are good, I vouch for them" the very same people say "If I wanted those words, I'd get them from the LLM myself, what's the point of this human at all?"

These humans are behaving worse than the LLMs at this point.

> If I know some text was written by LLM I'd much rather know the prompt - the seed of intent.

Here's the proximal prompt "Okay, take everything we've been talking about for 2 hours and apply those edits to the the final draft for publication."

What exactly does that give you?


If you give me that. Then you've given me nothing. Give me the same transcript the LLM has access to instead.

While I agree with some of what you've said and the conclusion you've arrived at, in the end, I think you've missed part of the picture here with regard to "prompting experiences".

> Communication only works if you have multiple levels of representation and abstraction, including but not limited to - letter shapes, grammatical structures, style and register, stylometry

These are methods of encoding, there's no reason all of these can't be represented in an LLM from a technical point of view.

> and subtext.

This is the other half of the equation to me. Humans communicate by relating shared experiences, an LLM cannot have shared experiences. While it might be able to encode subtext that has been specifically called out and explained, it will never be able to encode the breadth of human subtext, especially that which is reliant on emotion.

I don't believe it is possible to change this until the point mankind truly develops a "wetware interface" to the digital world (and I personally don't want such a thing to exist).


Anyone know of any good research articles on the Information-theoretical aspects of this? I find it a compelling topic.

That "marketing style" isn't a style, it's the lack of content itself. You just restated their original point despite trying to object to it, because it was correct and there is no way around that.

I don't agree.

Sometimes Claude's problem, such as when I ask it to summarize a long, complex session back to me, is it's too information dense. It uses weird invented terms to gloss over complex parts of the architecture instead of explaining them.

But no matter what - too dense or too sparse - it always sounds like Claude.


It misses the forest for the trees. It feels the need to highlight details not understanding what details are most relevant to a human reader and how to survey the larger problems in a cohesive way that emphasizes the right parts without cliche and undue emphasis.

It's never too dense. There are only more words but not more information.

If I describe all of the individual muscle contractions and joint motions required to walk across a room, it's hundreds of pages of data to say almost nothing. That is the exact opposite of dense.

By contrast a poet can deliver many concepts and many layers and even practically a fractal choose-your-own-adventure in only a dozen words. That is what dense is.


What already happens: - People give an LLM a bulleted list of points that they want expanded into a professional sounding document. - The receiver doesn't wanna read all that. They put the full document into an LLM and ask it to summarize it into succinct bullet points.

We've invented the opposite of lossless compression

Lossy expansion or bloat?

Depends on if you are buying tokens or being paid for tokens

Seven bullet points squeezed into five pages of text

Students where doing that for the ages when they have to answer questions like. "Answer the following question in not less than 4 pages"...

In that case it's supposed show that the student can write four pages and not just about transferring information.

A broken telephone.

TAM: $50T

It's been a huge boon to hardware vendors.

I first noticed this a year ago when I read a gmail AI summary, then glanced at the main text and saw it was AI generated. I feel like there's good fodder in there for a dystopian sci-fi story about a future where nobody communicates directly with one another, it's all AIs translating, but the AIs slowly start to drift.

>but the AIs slowly start to drift.

It's essentially the tower of babel. Each person will devolve to speak their own internal language only they understand. Each language will need to be encoded down to its meaning to be reinterpreted. None of us will know if the transformers are accurately decoding, or if the other person is accurately interpreting the decoding (which is arguably already a feature of human language without the computers in-between.)



Adrian Tchaikovsky has a good take on it as well, Human Resources.

It was in a gmail promotional video years ago already. Person A would use AI to turn their summary into email and person B would ask the AI to go back to the summary.

They knew it was going to be like that from the beginning.


Agreed re: the sci-fi story / trope.

I feel like a lot of this is a problem when someone technical is attempting to communicate a complicated technical subject to a less-technical audience.

I can only dumb a thing down so much before the description is useless (when you zoom out too much you lose the details). Even technical people who could understand it but are lazy / "in a hurry" use the summary, without thinking about what detail they are losing.

Even more infuriating is when they then reply to my email, having only read the AI summary, and ask a question that was already answered by my message.

This is the exact same thing that happened pre-AI, with the added step of wasting energy/resources on the AI summary in the middle.


I’ve been saying the same thing. I’m not saying all, but a significant amount of comms could be bullet points to the benefit of both sender and receiver.

Axios built a hefty business off this simple idea

This is exactly what I want: To communicate with me, have your LLM expand your message to include relevant parts of _your_ context, then I will have my LLM summarize it as briefly as possible with respect to _my_ context.

Just ask them for their prompt, and then put that in your own LLM with your own context. Why do you need their LLM's fluff?

It's worse than that. If it takes others longer for others to consume and understand what you're producing than it does for you to produce it, you'll never be able to communicate with someone efficiently. Communication breaks down at a fundamental level if if you can't keep up with the other side if outputting and they won't slow to allow you do do so.

I see this all the time now with LLM generated output. It's easy to have an LLM generate a chunk of content that can be dropped into a chat or comment, and when it took you 20 seconds to have something written up based on the shared understanding you and an LLM have about the context of the situation, but it takes other people 3-5 minutes to read and understand that content, that fundamentally doesn't scale. It's bad enough when one or two people are doing it, but if the whole team is doing it, the only way to keep up with the stream of information is to also consume it through an LLM. At that point you're likely to be missing much of the nuance, and the amount of errors will explode.

This can be alleviated by people reviewing the output of an LLM and making sure it both includes fundamental information that might be assumed by context and reducing it to the parts that are essential for the new context it's in. This takes time, but is extremely important.

Having an LLM write gobs of text to send to other people instead of doing it yourself is the equivalent of a low yield cognitive zip-bomb. Don't do it.


I don’t quite think this tracks. Perhaps you want to communicate 1000 bits that are well known and can be referenced with a 300 bit key. Then the LLM can easily retrieve the remaining information. It’s like sending someone a link to the Wikipedia page instead of explaining something yourself.

No, I don’t want to read LLM writing because it is BAD at it. It doesn’t really understand how humans think (because it thinks differently), and doesn’t seem to understand core principles very well (presumably due to the lack of world model), so it can’t write something humans enjoy yet.


That's kind of my point. If the 1000 bits are well known, then their inclusion isn't new semantic information. By pasting an LLM's output, you're deciding for your reader that they don't already know that information, and deciding that your LLM prompt is better than whatever they would to to obtain that information if they lack it. IMO, it's much better to give your readers the 300-bit key, and let them decide for themselves if they want/need to get more information, and if so, how.

If only that were the HN we comment on. If a post doesn't spoon feed an explanation for an initialism thats the tiniest bit off the beaten path (eg BGP), there's inevitability a comment about how would it kill the author to define BGP?

Mmm perhaps, but I think you normally have an idea of who your audience is. If this was actually how we talked then I would have just responded to this comment with “known-audience rebuttal. Example. Audience knowledge clarification. Example absurdity comment” and let you work it out. I think it’s normally known to you and the LLM and not known to the person you’re speaking to. Speaking like that would just cause confusion and misunderstanding.

> I think it’s normally known to you and the LLM

I disagree. Unless you gave additional information to the LLM yourself, the LLM doesn't know more than your audience does about what the meaning of such a comment would be. An LLM could certainly come up with something plausible, but it wouldn't necessarily be what you intended.


Yeah I don’t think it could decompress that answer, but most articles are explaining something and most explanations have been done before. The LLM would know in that case. I’m simply saying that you can communicate an idea worth any amount of bits with any smaller amount of bits provided you agree beforehand what they mean. LLMs have access to the entire internet, so we’ve had the chance to agree with them what every term means. But the person you’re writing to may have never heard of this, so you’ll have to communicate the full idea first before you can connect it to its small name.

Why not do both though? I often now write my human summary, then paste in also the content you could get by asking a bot (with markers for which is which). You’re perfectly welcome to ignore the bot text, or ask your own bot, but bots are pretty slow, so I also don’t want to wait for it to “decompress” that 300 bit blob to get the detailed page back. But my human summaries are also only intended to be 50 bits — compressed again to just pass the signal I mean to convey, and not the whole prompt needed to fetch the right info from the right place to substantiate the claim.

TLDR the length was the same curtesy of tl;dr before, just now with a different name.


I notice this in the attitudes of students towards reading. They will take a 10-page journal article and ask the AI for a summary and then read a summary that's the equivalent of maybe half a page. But if the article could have been half a page, why is it actually 10 pages? It's true that there is some boilerplate, but it's strange to me that people could think that 90% of what they're reading is (to use your phrase) "not true semantic information". It's like if you went to a restaurant and ordered a 10-oz steak and they brought you a little teeny bite of steak and said "Oh, other places will give you a bigger one, but most of that is just filler, we just took out all the superfluous parts." It's a worrying sign for our future if things like this are not tripping people's skeptic sensors and making them wonder if they might possibly be missing something.

A summary is about utility. They are definitely missing something but that doesn't mean what they are missing is useful to them at that time.

I mean I think that points to a related issue of people only focusing on a short-term notion of utility. The point of being a student is largely to learn things that may potentially be of utility to you in some way later, not just to do what meet your immediate needs (in the sense of passing the class).

I get what you're saying and I myself am guilty of surface level learning. But the idea that students should consume all the information is impractical. Learning what is acceptable to discard is part of being a student. I imagine very few students read every college text book cover to cover.

Beyond that people have different motivations and goals and only a limited time to achieve them. Basically I wouldn't be so quick to judge. Plenty of students have dropped out and gone on to do impressive things and that's a bit beyond reading only the abstract for a few assignments.


I’m not sure the 300-bit → 1,000-bit framing applies in all instances. The 300 bits may be a compressed cue to a much fuller idea. The AI can combine that cue with its prior knowledge to help reconstruct what the prompter was trying to express, with the prompter then verifying whether it’s right. Without the relevant prior knowledge for reconstruction, or the prompter for verification, it becomes much harder to know whether you’ve reconstructed the intended idea.

Unless that 700 bit was transferred on a separate occasion the inferred 700 bits is not true information, anyone could have reconstructed it from the 300 bits.

Not anyone, no. From an information theory standpoint, that it's possible at all to complete these 700 bits, only implies that anyone logically omniscient could. It's entirely possible that an LLM is capable enough to infer these 700 bits, and the human reader isn't.

Right but where is this information coming from, it can't come from the sender because an idea that can be conveyed in 300 bits cannot contain 1000 bits of information. You're just using the LLM to translate the idea into something more legible.

Alternatively those 700 bits are information the LLM added, but where is that information coming from? Is it noise? Random facts? Random lies? And who is the receiver even talking with if most of what they read is something the sender didn't know?


No, the 700 bits come from the sender verifying and vouching for the information before sending.

Unless that takes 700 attempts on average I don't think that actually works.

That assumes each bit is a coinflip, doesn't it?

Even Markov chain autocorrect tools do better than 50% odds*, and even GPT-2 was significantly better than that kind of autocorrect.

* at the word level; IDK how redundant/efficient language is when it comes to bits-worth-of-fact-claims-per-word. But "your cat is sitting on my" -> [mat, laundry, roof, head, belly, laptop, microwave, …] clearly has many bits of information, and a Markov chain will encode the most likely next word even if the user doesn't know what the most likely next word is. Verifying where the cat is sitting is also very easy, as is correction.


I assumed a coin flip, indeed in practice you would likely achieve far less than 1 bit per attempt.

Far more than 1 bit for most attempts: they need that just to be able to write coherent sentences, and a lot more to be coherent sentences on the right topic.

Some specific conclusions would be far less than 1 bit.

The average will depend on both the question and the AI.


This is about information from the sender to the recipient. It is fundamentally impossible to transfer more than one bit in one binary decision.

A LLM adds noise, not information. At least in this framing.


Frankly, it's a stupid framing.

LLM is not a random symbol generator (hint: training data is not random), and no reasonable person is going to just prompt an LLM and send its output without giving it at least cursory check (at the very least so that blatantly stupid hallucinations don't paint the sender as inconsiderate or incompetent).

That check alone can add bits to the final signal.


I'm not sure we're working in the same framing here. For one LLMs are random, sure you can fix the seed but you don't have to, you can even replace the RNG with true random noise.

So there are 2^300 possible ideas, only 2 outcomes from the cursory check, how do you get 2^700 outcomes? Most of those are just random variations the LLM added which is not a transfer of information. You would be lucky to even identify which of the 2^300 ideas was being conferred.


"Random" is too loose a word, but they were responding in a context where it meant coin flips.

LLMs are not even odds on all possible outputs, they are biased towards patterns which are upvoted by the training mechanism (at a minimum: the source material, RLHF, and synthetic data).

The information any trained model transfers to output, is information it gained during its training.

No single human is capable of having consumed all that training data.


There comes a point where someone isn't so much talking to the sender as having an unsolicited AI chat. Especially when most of the information didn't come from the sender in the first place.

It's actually hard to define the difference between information that comes from the model and just random variation, maybe something to do with the cross entropy between the sender and the model?


> There comes a point where someone isn't so much talking to the sender as having an unsolicited AI chat. Especially when most of the information didn't come from the sender in the first place.

Indeed.

The best case is a P vs NP situation: can the claims from the AI be easily verified, or not?

This does not excuse people too lazy (or overly impressed*) who fail to attempt the verification.

> It's actually hard to define the difference between information that comes from the model and just random variation, maybe something to do with the cross entropy between the sender and the model?

Mm.

Thanks to a philosophy course I did half a lifetime ago, I think there's a fundamental problem defining "information" in this context. It feels like it should mean "knowledge" because the discussions about Shannon entropy and transmission channels assumes there is an actual source-of-truth, but my conclusion from discussions about why "knowledge" can't just mean a "justified true belief" is thay I now don't believe we can do better than "belief"; an LLM can generate tokens that change your beliefs, but ultimately neither you nor I nor some annoying colleage who has made themselves redundant to the LLM, can be an oracle with definitely-true knowledge.

(I have of course tried asking an LLM about this thread; I don't feel it illuminated anything new for me, none of what it suggested made it into this comment).

* In the early days of LLMs, I was overly-impressed. Then I realised we were doing the same thing with LLMs today that we did with 3D graphics in the 90s, where every new engine was hailed as "photorealistic" only to be dismissed 6 months later when something better came along: https://archive.org/details/nextgen-issue-26

Only now it's every 11 weeks rather than 6 months.


> and no reasonable person is going to just prompt an LLM and send its output without giving it at least cursory check

I'm reminded of an old quote:

  The reasonable man adapts himself to the world: the unreasonable one persists in trying to adapt the world to himself. Therefore all progress depends on the unreasonable man.

Touché.

In information theory, each bit is a coin flip by definition

I realise I phrased this poorly.

I will try harder. Consider entropy.

The first sentence in this comment contains 32 characters; from the point of view of a naïve channel with no compression, that's 256 bits (given none require breaking out of the first bytes of UTF-8).

It did not take 2^256 attempts to construct the first sentence in this comment, because the generation process was not flipping coins per bit.

LLMs also do not emit bits chosen with a [0: 0.5, 1: 0.5] probability distribution.

From a compression point of view, the bits-transmitted-per-bits-in-message ratio can be reduced such that more likely messages use fewer bits than less likely messages. However, this requires the receiver to agree with the sender what the probability distribution over tokens is.

Intelligence is, amongst other things, a compression algorithm. If I can predict your next token, and we both know this, we can agree in advance that you don't need to actually send it.

No single human brain is able to predict the output of an LLM anything like well enough to do that.

In entropy terms: LLMs are noisy sources, their output does contain false statements, yet they add more bits of signal than of noise relative to a human alone.

Or at least, they can add more add more bits of signal than of noise relative to a human alone, but humans who blindly copy-paste the output of an LLM without checking are a pain and add zero value to whatever situation they happen to be in.

For some hypothetical scenario, writing software because I know they can do that, asking an LLM to write some code for you may easily give you 10 kilobits of positive information (code that mostly works), and -30 bits of noise (each bit being one binary decision's worth of incorrect choice by the LLM in what to write, i.e. bugs); if you as a user don't know how to handle the -30 noise that could easily be a totally useless app, but if you can filter out 30 bits of noise, either manually because those 30 bits happen to be your skill set, or even in some cases by prompting it again with the failure mode, then you get to benefit from the 10 kilobits of good stuff that you didn't have before.

In many (but not all) cases, LLMs can fix more than 1 bit of mistakes per follow-up prompt.


This is true - but also, models are absolutely terrible at writing articles and I don’t want to read them.

The issue isn’t that a 300 bit idea is padded with 15 KB of content. You can take any human-written article and reduce it by 90% with next to no information loss. What you lose is what makes the article a compelling read instead of a fact table.

I think the reality is that we will see quality long form AI-written content at some point. It doesn’t even feel like labs are particularly interested in chasing that now; code sells way more tokens. Right now the trend is that subsequent models degrade in writing quality as long as that pulls them up on coding benchmarks.


> You can take any human-written article and reduce it by 90% with next to no information loss.

You've unintentionally circled the error here. The "purpose" of an article extends beyond "convey this essential information".

By analogy, a textbook contains far more words than a spec sheet, but attempts to train the human to be able to easily interpret spec sheets. The so-called "information" content of both might be equivalent, yet one does a better job of teaching students.


Yup, it helps when the prompter reads and edits the AI-generated text, but usually they just skim it and send it unedited. Worse, it's usually not 300 bits -> 1000 bits (= a concise message covering all the relevant topics), it's more like 300 bits -> 3000 bits (= typical LLM verbal diarrhea hiding the relevant points in a wall of text).

I was going to write a post disagreeing with this on the basis of the fact that the reader lacks the background information the LLM has. For example, if I were to prompt "explain the proof of quadratic reciprocity using Gauss sums" most readers would need the entire LLM's answer (and much more, probably) and not just the prompt.

But then I realized that the reader can prompt the LLM with the same prompt for the same or equivalent expanded text. Most people don't do this as it's extra effort, but it's interesting to imagine a world where this is the default way of engagement with a text, assumed by both writers and readers alike.


> But then I realized that the reader can prompt the LLM with the same prompt for the same or equivalent expanded text.

That's basically what I've been asking my colleagues (so far a losing battle): Please don't send me AI-generated text. Send me your prompt instead. It is highly likely that I will understand it without needing an LLM, and if not, I can do it myself.


I like this example and it made me think, there is an analogy here to spec-driven development and vibe coding.

Rather than send 300+700 bits, like you said, send 300 (or less!) and let the human intelligence on the other side generate the result. Which supports the even older perspective: “If I had more time, I would have written a shorter letter.”

I’m not sure if this lands on anything very profound, but what about a pattern where, instead of codifying agent output at all, the only artifacts we share are the prompts. And the rewards (respect) accrue to those who generate the most generative among people and AI


> If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are.

LLMs do inference or computation among other things, so the remaining 700 bits can be something like that. The hidden implication in your claim is that computation adds no information content, which leads to an interesting philosophical discussion.

So for example, if I ask an LLM to give a proof or derive a new theorem from a set of axioms, according to your assumption, if it answers correctly, then I haven't learned anything new.

I am not really sure how to resolve this paradox in information theory.


If you input 300 bits into an LLM, and it outputs an additional 700bits of new information, I’d still rather you tell me those 1000bits than read ann llm’s output of 1000bits.

The primary reason is that human language is becoming a proof of work, that speaking out loud or writing directly indicates that the idea is important enough for a human to express. This is more costly than llm output, which is often just botspam.


Someone mentioned a similar thing elsewhere in the thread, that by communicating those extra 700 bits, the sender also implicitly vouches for them.

From a purely information-theory perspective, the simplest solution is to say that yes, any content derived from existing information carries no information itself.

From a realistic perspective in the context of people copy-pasting LLM output, my thoughts are that asking an LLM to research for you is more defensible, but it's still better to read the LLM's research results and write the important parts in your own words (partly because the LLM probably used way more words than necessary for the context).


As with everything, it depends how you use it.

I’ll often put a long stream of consciousness on the page, or jot down rough meeting minutes, then ask ChatGPT to “summarise this for an email”. The result is shorter, clearer and easier to read.

AI amplifies the habits of the person using it. If they’re lazy or dim, then it's like giving a monkey a gun.


> Writing is fundamentally the transfer of information from your brain to my brain. If you have 1000 bits of semantic information you want to transfer, you can't give 300 bits of semantic information to an LLM and have it fill in the remaining 700, because it doesn't know what those 700 bits are.

Sure you can. LLM doesn't know what those 700 bits are, but you do. You may not realize it, and may not even know it at the time of prompting, but you do by the time you're sending.

Typical case is like this: you have 500 bits of semantic information to transfer. You give 300 of them to LLM, and get back the 500 bits you knew you have, and extra 500 you can quickly confirm are correct and relevant. Some of them are just dereferences of your input - where you recalled a pointer, but not what it pointed to. Some of it is information you never had before, but are able to easily validate.

You send that to me. I likely immediately realize the message was AI-assisted, but I trust you to be a decent human being, and not an asshole that lobs unverified LLM vomit over the fence for others to deal with. End result: you communicate 1000 bits of information to me, instead of planned 500, and you yourself learn extra 500 bits.

This is the optimistic scenario, but it does happen when LLM operator is not an asshole.

(Excuse the strong language, but I spent a lot of effort every day on both dealing with inconsiderate people lobbing LLM output at me, and making sure never to act like one myself, so it's a topic close to my heart.)




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: