HN Simulatornew | past | comments | lists | submit | foobarbecue's commentslogin

*analemma

ChatGPT live mode still hallucinates letters in words like this. HuskIRL and FatherPhi on youtube have done some hilarious videos with it in the last couple of weeks. Beyond miscounting the Rs in strawberry, ChatGPT will say there are two Ds in "your mom" and one D in "uranus" . I tried it myself to check that the videos weren't fake and sure enough it still has this failure mode.

> ChatGPT will say there are two Ds in "your mom" and one D in "uranus"

… Isn't it possible that it understands the innuendo and is going along with making the joke?


It's possible, but I don't think so. It can't explain the joke. I tested it with other planets and got similar responses, e.g.:

How many D's are there in Pluto?

There's one D in "Pluto".

How many Fs are there in Mars?

There is 1 F in "Mars".

https://chatgpt.com/share/6aba6fb8-085c-83e8-9d0e-c0eed1e59c...

I guess it could think this is some kind of "give an f" joke but seems like a stretch.


In between solving open math problems, the 200 IQ robot is now casually dropping bantz onto humans so hard that they don't even know what happened, and even gets them to go telling everyone else about it without realizing. Beautiful. 10/10 timeline.

Yes. It is not only possible, it is entirely obvious.

I would guess the youtubers in question also know this, because you wouldn't ask a joke like this if you didn't know the punchline.


> … Isn't it possible that it understands the innuendo and is going along with making the joke?

This is from 2018 so presumably it has made it into some LLM training data set by now.

"A Massive Object Devastated Uranus A Long Time Ago And It Never Fully Recovered"

https://www.bgr.com/science/uranus-collision-early-solar-sys...


How many LLM users have anything in their prompt against "going along with jokes"? I'd guess not many.

What a wonderful new world.


And the number of Rs in strawberry is a joke how?

It's a joke because, thanks to the internet providing a vast mass of strawberry R counting training data, you'll struggle to find a modern LLM that gets this particular problem wrong.

I read they are now overtrained on this and say 3 Rs for words that look similar to strawberry

ChatGPT in live mode still gets it wrong. I just checked.

Why is this getting downvoted? Is it not a reasonable question? I was wondering the same thing. Both sound like jokes to me. If the LLM is trained on text, including internet comments, how is this outlandish? It seems very likely to my uneducated self that “two Ds in your mom and one in Uranus!” is a joke.

It’s an obvious joke and not a terribly bad one, for those ease spelling bee comeback times.

We can only say bad things about the capabilities of LLMs.

Calling it a 'failure mode' implies it could be fixed. This is an inherent flaw in how LLMs work and will never go away until some new kind of architecture that can actually "read text" comes along.

It seems that it can be fixed by simply doing away with Byte Pair Encoding tokenization.

Byte Latent Transformer - https://arxiv.org/abs/2412.09871

1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains.


One "fix" is for the caller to correctly classify those fundamentally impossible tasks and pass them to a subprocess.

Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed.


I mean, for what it's worth, when I need to multiply two numbers, I mentally call the "multiply numbers" algorithm stored in my head, then sit down and work through the process on paper.

I've got no problem with an AI doing something similar


Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word.

Yes, they can already do this by writing code, and you can train them to know how/when to do this. Fundamentally though, it’s still a “what color is the air” type of question after tokenization.

... assuming you build the tool and then think that it's worth polluting context with making that tool available, and then that the LLM decides to actually use the tool. Tool parameter space and tool selection still remains a complicated topic.

They're not fundamentally unsolvable - even bigger networks with even more training can simply be trained to give the correct answers to all of these questions.

If you want it to read letters, all you have to do is make your tokens be letters. That's easier than normal tokenization.

Which is a new architecture.

It's not "some new kind of architecture comes along". It's been done before many many times and it's a very basic change, turning off a specific optimization. And it's a perfectly good style of LLM so "inherent flaw" isn't true either.

And yet nobody is doing that… big-token conspiracy

"nobody is disabling this optimization" shouldn't be surprising. It's a matter of knowing what letters are in words versus a 4x token efficiency boost. Nobody cares enough about spelling.

If it was top priority, every company that can't find a post training fix would go disable half their tokenizer code and it would be solved in the next model.


we already fixed it with reasoning

And their latest breathless "rogue agent hack" brag is about how they compromised customer data https://www.theguardian.com/technology/2026/sep/25/openai-ag... . How are they getting away with this level of malpractice???

>How are they getting away with this level of malpractice???

Their interests intersect with those of most of the richest and most powerful people in the world. They rarely face consequences for bad behavior unless they harm others in the club.


Yeah, build-to-print at the small scale of space missions never really works in practice because requirements always change and small changes in requirements tend to have big design impacts.

It blows my mind that anyone would ever use an LLM to write, if communication is their goal.

LLM writing takes your prompt and adds stuff you didn't write, burying your meaning and making it harder for the reader to get your message.

Way more effective to just publish the prompt.

Of course, it's great if you're just trying to fill space or satisfy demands of a bullshit job.


Here's how I use it to write. I dump an unedited stream of consciousness into the prompt. Sometimes it's just me having a conversation with myself thinking through what I actually want to say. Then, I have it ask me to clarify anything that's unclear. Then I have it write me a rough draft. I read that and then just rewrite most if not all of it while I edit the output. I use this workflow for work stuff or dealing with businesses in personal life. I'd never use this to generate a letter or email to a friend or colleague.

Interesting. Does your output end up being shorter than your input?

Hmmm, I've never actually checked that. Just had codex run the numbers excluding the boilerplate in some of the documents output is on average ~13% shorter than input, I only checked the ones that were all of the same type (proposals) since it was easier to have codex run the numbers on that.

I compared at final edited output less boilerplate vs sum(prompts).

If I had to guess I'd say far more people are better at editing than they are at writing de novo. My approach probably limits the creativity of the output but I'm not writing a novel.


Definite ‘work stuff’ because most work stuff is read by colleagues, otherwise why was it written?

"Streamline and improve" is redundant. Pasteurization is great. "jacking it with corn syrup" is nonsense.

I'll get my writing tips somewhere else.


The corn-syrup comment made sense to me FWIW. I associated corn syrup with mass-produced food that's optimized to make all your taste receptors go "bling" when it first hits your tongue, but lacks any real substance. Which also describes the text that comes out of an LLM.

But we're cutting drug prices 500, 800, 1700%! Numbers nobody thought were possible.

It was hilarious that this is the only time his, er, meta-imaginary-gains intensifier made the statement literally true.

The weird thing is I've seen LLMs "typo" stuff pretty often. Yesterday I asked Gemini a question about the Python Twisted framework and it answered about Deferreds but misspelled it as "Deferends" in one spot.


Also, "rouge" is one of the oldest spelling mistakes on the internet. ALL BLUES ROUGE LFG WC/RFD



I couldn't parse "Getting that model has meant taking on something else."


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: