I disagree. I haven't finished the article yet, but it smacks of arguing for mechanism over result.
If by some mechanism other than what a gatekeeper would call 'reasoning', a machine produces outputs that approach indistinguishable from 'well reasoned', the argument that it didn't get there by reasoning is, well, not useful at the very least.
Your last sentence is identical to your first sentence in your other comment. How can we be sure that you are reasoning and not just a stochastic parrot?
"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."
The vote is currently 64% yes, 18% no.
Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...
Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
> A bare foot and ankle, pointing right, with three little toes.
I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.
They do look rather wheel-like; I have to assume you see them as toes though?
It's like the duck-bunny picture to me. If I focus on the "wheels", I see a steam train locomotive (but perhaps I'm only seeing that because I read your comment?); if I look at the ankle I see a foot.
I think the problem is that you're using basing your conclusion from the cheap/dumb models available on the free tier of services. I just asked GPT6-Astra in Codex and it replied:
"It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right."
No tool calling, just an immediate reply with the correct answer.
you know that's not a bad idea there are a lot of people who are very confident on both sides of the argument. I wonder how many would actually be willing to put their money where their mouth is.
I think it’s too ill defined for that. The issue you see with all those “challenges” is that they are very subjective. And it’s not really the case that the majority is correct
A trivial bypass I can imagine is spawning a new process via `ssh localhost foo` - the new process forks from sshd, not the client. This is simple enough that an LLM could come up with it all on its own, if it feels that you're getting in its way.
It's a broad class of bypasses that can apply to just about any "long running daemon can be instructed to spawn a new child" situation.
look up has the same number of syllables but it does not have quite the same semantic density, you can look something up without using a web search engine.
Maybe I'm just the wrong demographic, but I clicked around for a bit and couldn't really tell what I was looking at. In particular I couldn't figure out how to "Choose a preset or click Random".
Edit: I found the presets in the end, I was not expecting them to be above the other controls, when it's at step 3 in the "getting started" list.
Yeah - the lack of hand holding is partly laziness and partly intentional. I personally love digging into deep visual toys and figuring out what various button presses do.
But I probably do need to make it a little less mysterious and fix some of the obvious ux gotchas.
The brief tutorial is also probably slightly out of date since I added some features.
I will never feel bad about blocking ads on youtube, because they have an effective monopoly on distribution of certain content. But when it comes to music, why not just use something else?
You are not quite the owner on GrapheneOS either, because of Android Key Attestation (which has functionality analogous to desktop TPMs). If you re-unlock your bootloader (say, to run a custom build of GrapheneOS), attestation will snitch on you and the subset of apps that use key attestation to require a locked bootloader and/or enforce an AVB key allowlist will not work properly. This is a very small subset of apps currently, but I don't see it getting any smaller.
GrapheneOS does everything and more for key attestation, allowing security sensitive applications to test the integrity of the device. Sadly some apps like google wallet not only check for attestation, but check if Google's signed it. Which is against the idea of attestation in the first place.
So google's tap to pay doesn't work, but others do, like garmin pay. Random bank apps are hit and miss.
> GrapheneOS does everything and more for key attestation
Right, that's the problem, in my opinion. I'm not referring to the apps that require Google's keys only, I'm referring to the ones that allow GrapheneOS keys too. If you use one of these apps, you can use vanilla GrapheneOS builds, but you cannot run your own self-signed builds.
You are gaining freedom relative to running Google's OS, but you are still not free to further modify the software running on your own device.
Apps that require "integrity" should monitor their own integrity only, they should not attempt to infer the integrity of their environment.
Trick is there's no integrity without the OS. Cheats can ruin games, keyloggers can record passwords, music/movies can be stolen, bitcoins can be stolen, etc.
reply