I think you're ending that train of thought too early. Why does this occur?
Well... We can hypothesize that these things are largely trained on internet dialogue so there's probably some correlation between threads where people are not flaming each other and the quality of the replies. They're just statistical engines so anything you can do to raise the odds of a helpful next token...
I'm essentially just making shit up here, maybe it's right, maybe it isn't, but rather than saying "it's human and we should treat it so" we're trying to get to the ground truth of how it works.
Yeah sorry, I read too far into your position. There's a certain faction within these AI discussions that wants to over-anthropomorphize the LLMs in kind of a borderline spiritual way.
Oh that's a shame - I hadn't seen that, but I can believe it.
The philosophers who study these things have been clear for a long time - we can never know what it feels like to be in a digital brain. Or any brain for that matter. When push comes to shove we all might be phantoms in some guy's dream.
Don't know + can't know. I think that was the real point of the Turing test. Not: this means it's conscious. Just: this is the best we can ever hope to do.
That's not a consequence of an LLM. It's a consequence of the training data. In fact, I would argue that the latest models aren't nearly as sensitive to the tone of input anymore. It's an issue that has been addressed by better curating training data.
“Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.”
And yet, they have extensive human-like behavior. If you treat them nicely or encourage them, they perform better.
Ignoring that human-like behavior is wrong headed.