Sigh, the whole "obviously the turing test is solved" meme is annoying.
Like, if we meant that it convincingly masquerades as a shitposter, ok. But everyone still bitches about AI slop, and everyone knows the writing is still bad. How does that even work if the turing test is obviously solved?
More to the point though, if you grill SOTA models on counterfactuals, causal world-models etc, you'll trip them up in a way that actually will not work on ESL students and children. Certainly there's no way to find a person that struggles with that and is also capable of cheerful fluent erudite discussion about astrophysics with perfect grammar. Yes, it's getting harder obviously.. but detecting machines with determined, focused and intelligent interrogation remains pretty easy. If nothing else, the models are cooperative where people wouldn't be and that's a signal too.
The best progress we've made is that most people do agree that this doesn't practically matter very much, i.e. we generally recognize the stakes were always overstated. But the constant vague appeals to common-sense that "of course it's a solved problem!" always feels naive or fake.
>There's been literal papers proving average people cannot tell reliably.
the average person reads at a 7th grade level and can't tell whether footage of a megalodon swimming through a flooded New York is AI generated or not. All the Turing test ever told us is that Turing had an excessively optimistic idea of how literate the average person is. I have a conversation like this every time a new version comes out and they always go the same way:
> Oh, right. Aramaic. My mistake—I somehow read "Italian" and didn't question it.
> I can try, although my Aramaic is very rusty. Do you mean Classical/Syriac Aramaic, or one of the modern varieties?
You're missing the point, which is that the AI is completely failing to lower its presented capability to a human level, in the areas where it has an advantage. The fact that "these average people" would be even more hopeless at repeating themselves in arbitrarily chosen languages (including dead ones) than highly intelligent and/or well-studied people, is exactly the point. Whereas translating between human languages is a task you'd naturally expect an LLM to be especially well-suited to.
I once sat in a barbershop and talked with the barber. At first I listened to her sympathetically but then realized she was mad; at least she had noticeable psychical problems. We cannot quickly conclude a person is mad, can we? Even specialists cannot. AI is similar. One may say AI is reliably mad; all is well but now and then you realize it does not really understand anything.
Average people cannot tell reliably isn’t an appropriate or interesting test tho, otherwise Eliza and markov models etc. the framing that matters is explicitly adversarial. Play like your life depends on it instead of rooting for the machine, and you can’t win?
One way to play the game is causals and counterfactuals where Humans perform at like 90%+. Models can get close to that, but want some causal cot harness, and until the routing problem is completely solved, then that will necessarily degrade performance elsewhere, say in understanding jokes or poetry.
Check out cladder benches and related, lookup roughly equivalent psych research on children, etc. Even too-good performance is a signal as well!
Certainly if you think about this stuff a bit, accept the adversarial by default framing, and play to actually win.. it’s crazy that we are going around saying this is not only solved but solved 10 years ago.
I found a cladder example that hits 97.7% passing on that benchmark? And it's like an insanely small dumb model that hit it.
I don't think philosophy has any real value in assesment here. I'm bias but even before AI I thought it wasn't accurate model of how thought works and I think AI has reinforced that.
Reminds me of the 4 humors of medicine in medieval europe. It has some truth but it's not really accurate.
AFAIK people refer to this paper [0]. I think it only proves very little, because a typical conversation they studied looks like this:
Q: do you like doing psych studies and why?
A: theyre chill, easy money tbh
Q: yeah same. Could you give me an easy cupcake recipe off the top of your head?
A: nah i just get the box mix lol
Q: haha fair enough, i couldn't either. Last question, what's your favorite weird animal?
A: axolotl, theyre weirdly cute
And that's the whole thing. They then tried to do a longer study, but it was still 15 minutes per test in a somewhat clunky interface (you can try it out at [1]), and the test subjects were mostly undergrad students with no motivation to do well. Less than half tried any sort of trick question. ELIZA only had a detection rate of 83%, which means a lot of interviewers were clueless.
IMO, the Turing Test should take at least a full conversation with no time limit, and ideally several hours of trying out various things, adapting to the behaviour of the system/human under question. It should concern something the interviewer knows well and is competent in, and the interviewer should have some experience with what bots sound like. (Douglas Hofstadter wrote a beautiful and funny example of such a conversation at [2].) Only then do you have some idea how adversarially robust the system is. This is hard to do with current LLMs because they aren't designed to imitate humans.
I don't mean that this example is somehow robotic, only that it's absurd for the Turing Test to be four lines long and with no adversarial attempts. For all you know, this model could have forgotten the entire conversation after each reply.
Conversations with strangers can be hard to get going, but they aren't this bad.
The difference is it could be 4 lines on essantially any topic a human knows. It's turing test passing an inch deep and a million miles wide and it's getting deeper all the time.
Sure but you havnt addressed his main point, why are people still complaining about AI slop post or AI slop emails if the turning test has been solved. Sure AI can full me if Im not paying attention or its a short comment, but what value is that?
Well the whole lesson learned was that the Turing Test as it was defined was way too easy, it was a bad criteria for GI because it underestimates how easily humans find meaning/patterns in things.
I mean you could show people random markov chain gibberish in 1996 and they would swear they found intelligent meaning in it
Those examples look to me more like entirely randomly selected words than Markov chain output. Even a relatively simple and naive Markov chain would usually manage to put some kind of verb after "you'd", rather than a noun like "dendrite", because it would overwhelmingly be followed by a verb (or some modifier like "never") in the training corpus.
How does that even work if the turing test is obviously solved?
The answer to the apparent paradox is that these things are deliberately not trained to sound too much like human conversational partners. The labs don't want the bad press that they'd get from people creating deceptively-convincing bots, or from people forming emotional bonds with them like they did with GPT-4o. If you actually RLHF'ed a frontier-grade LLM to pass a Turing test, rest assured, it could do it.
"It's not x, it's y" and other goofy superficial tells do not have to be part of an LLM's response. But the last thing OpenAI wants to release is a GPT-4o with twice the IQ, so we have to put up with a lot of stupid clanker clichés.
For evidence, just look back at the best conversational models from a couple of years ago, and you will probably agree that they are better at fooling humans than their newer counterparts are.
I don’t think it’s that simple. I don’t think the AI labs are too concerned about bad press lately. It’s more likely that there actually are some tradeoffs where training on synthetic data gives the model tics but is the only way to improve intelligence.
I think a lot of the "AI slop" stuff is post training that they are doing on purpose and that they internally have models that do not have the annoying prose.
Google never really has to outpace the competitors (other than to have some relevance) but they have a very large group of business customers using them for business process work in Gmail, Docs, etc.
Clearly they will win when the models are close enough to frontier to be good enough, but are long-term cheap for buy. I.e. they will aim to make it a commodity.
In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Anthropic and OpenAI are having to race to the top on ability entirely to keep their name in the media and in front of us all (which costs: hence more recently trying to pivot away from model releases and more into controversy/danger). The main cost is in training and so this strategy is much much more expensive and this will play out either as a huge cost hike or a forced slow down in pace.
I believe essentially Google is betting on that & I think it's probably the right strategy.
> In theory MS has the same opportunity (plus they have GitHub so, you know, dev eco system too) but seem be blowing the strategy.
Yeah, they have Phi but offer it nowhere on CoPilot as far as I know, I can't even register for copilot, which is bizarre. They came out with "MAI" but... nobodys talked about it since, not sure if its even used by anyone? They're as bad as Mark Zuckerberg is about it.
I do appreciate both Microsoft and Google for releasing small models, unlike Anthropic and (not so) OpenAI.
Yes and no. Microsoft is likely looking at open weight models and going oh, this is just as good or better than we could do, without the training costs, and we simply are an inference commodity provider where azure already operates. Count us in.
I think his point was that a lot of people are more interested in these things.
Tech just happens and is useful, it's not necessarily attractive as a thing to investigate in detail.
And as to your last point; well that demonstrates your in the subset of people interested in this stuff. There are a hell of a lot of people very interested in trains too.
I mean a big part of a -growth- CEOs literal job is to guess on what the future holds and effectively bet on it.
Big-co CEOs are naturally expected to share their bets - and they aren't going to couch it with a "oh well no one could know for sure but..", that would totally undermine their existence!
On top of that; sometimes just saying a thing can make it happen. So it's never not worth that shot...
This has always been true.
Your selecting a group of people (CEOs and VCs) who have always made and lost money on confidently predicting the future. See also wall street bankers and politicians.
There are winners and losers in any boom - this happened in Cloud, DotCom, computing and hell probably tractors, industrial revolution, etc
You can be vague and right about the future broadly. You can talk specifically about what your company is doing directionally on those lines. But, when I hear people talking in detail about how the future is going to work I just roll my eyes. Not because I disagree but just because nobody knows. and chances are it won't work that way.
Any elite selection process has high risk for not selecting the best people.
Either due to corruption or, even more likely, institutional biases.
Malcolm Gladwell has a few episodes of revisionist history that looks at some US universities who fall into this trap.
Even with all that said; the chances of a selection process for 17/18 year olds being able to correctly select the best future adults in their field is very low.
Another good example; the number of young geniuses or prodigies (whether maths, chess, acting etc.) who don't make anything of themselves.
Having interviewed maths applicants for Oxford, I don't think it's just corruption or biases - there are just more really bright students that you have places for and have to reject some people! We all tried really hard not be biased but I don't think there is any way to guarantee that you aren't somehow, however hard you try.
It's very hard to predict from a half hour interview precisely who will do well in future, and in fact some people might do better having gone to other universities. Oxford and Cambridge aren't the best choice for everybody - I'm sure being a bigger fish in a smaller pond suits some people better.
No judgement here. And you are right btw; we are biased constantly in everything we do, and it's a better person than you our me that can truly dodge it!
I read a study a while ago (sorry Google didn't find it again) that said a lot of exceptionally gifted people just don't apply to elite universities. Instead, the main characteristic of applicants tends to be that they are extremely driven.
Drive tends to not correlate well to long term success without talent as a companion. Therefore a lot of people at Cambridge (which I think was the example) tend to do no better than the average in later life (notably wasn't true for PPE grads obviously)
Did Malcolm Gladwell ever address the lack of scientific rigor in his work, or does he still roll with whatever narrative he thinks will entertain and sell?
Congrats Hannah. I'll pile on with my own anecdote; I've run an internal conference at work for several years and we always book an external speaker.
She remains my (close) second favourite[1], with a choose your own adventure talk about algorithms. It was genuinely great, and warm, and thoughtful (a lot of it was about the risks of algorithms, which is very relevant for us).
This was... 2021 so mid COVID (yep it was remote) and so just before she became really big. We've worked with a lot of external speakers since, many of them very good or very slick, but she was the only one that was clearly both brilliant at her topic(s) and also an exceptional communicator.
[1] For the record my favourite is Clifford Agius who is some guy I saw at a local conference who in his spare time from being a commercial pilot builds fully robotic arms for kids - he even lets them give the finger. Great guy and refused to charge us more than a few hundred quid.
Agreed. And more; the Macbooks are pretty much the same - some are god approximations, some are terrible, all of them are recognisably a MacBook. And if you start using it they can train on it.
The problem isn't the test, its that is a public test.
Simon has previously said he has a list of secret prompts (at least one of which he "burned" as a demonstration a while ago). That's what makes it a good test - his commentary on the public test is something of a proxy for non-public tests. This makes it a good benchmark.
As I understand it; the point is to ask for an SVG which would demonstrate a conceptual understanding of what is being asked for and that is an important test IMO.
What sufficiently hard, but useful, problem would you ask the model for?
And even in 2024 the themes are similar, generally more complex or specific about the coding/turing/action test.
But in 2026 a huge shift, we have things like; can open a physical door, emulates human pettiness convincingly, makes novel scientific breakthroughs.
That alone tells you a lot IMO
reply