Imagine if a car company like Ford or Toyota came out and said their new cars are kind of out of control and they needed special indemnities and protection from liabilities if they want to keep selling cars.
Replace this with Tesla, Waymo, etc and it's not far fetched. It might take a while, but liability is going to need a new legal doctrine to describe why it's not their fault that their car & their code killed someone.
They pretty much did. Cars get an insane amount of special legal treatment. It's basically legal to kill people with a car as long as you aren't drunk or grossly negligent.
I will be shocked if I live to see a year where AI kills more people than cars do legally.
> It's basically legal to kill people with a car as long as you aren't drunk or grossly negligent.
It's the same even without cars; if you kill someone by accidentally bumping into someone, or running them down with a bicycle, you aren't going to get charged with murder as long as you aren't grossly negligent.
Name one other multi-ton object that you can move through public spaces at high velocities without the law considering it inherently grossly negligent.
> Name one other multi-ton object that you can move through public spaces at high velocities without the law considering it inherently grossly negligent.
That is not the point I replied to.
This is a different point, namely, "Using a car on public roads is, by itself, grossly negligent".
Imagine further that there's actually nothing unusually wrong with the cars at all, but the manufacturers are deliberately playing up perceptions of risk in order to encourage regulatory intervention that they know they can influence and/or work around, but which competing solutions can't.
I mean... there is Boeing, which is arguably more horrifying.
With AI it seems to be internal testing produces wild results, which is unsurprising to me, all of these models can be jailbroken, and as long as that's the case, they're all likely to be exploited maliciously.
Not just internal, Anthropic’s S-1 literally covers the risk of ending humanity. They are telling their investors a public company could eradicate humanity. Give us money, the risk you have to evaluate in your financial framework is that we could eliminate humankind.
I’m repeating it because that has to be internalized: the largest IPO ever will be for a company that casually consider they might kill us all and their own employees too. Including their children.
You can use open weights models right now and ignore all this fuss. And by the time open weights models are objectively and subjectively on par with what Fable/Astra are today, the frontier labs will probably offer something newer, more capable and more expensive. The models OpenAI and Anthropic roll out clearly aren’t for people on a strict budget.
Maybe one day open weights and closed models will hit the same wall?
I heard someone who studies this sort of thing say basically what biological neurons are trying to do is predict as well. Predicting what exactly? I’m not sure. The next time they should fire or something. I can’t find the YouTube video now.
In the last few decades, there has been an increased interest in the role of prediction in language comprehension. The idea that people predict (i.e., context-based pre-activation of upcoming linguistic input) was deemed controversial at first. However, present-day theories of language comprehension have embraced linguistic prediction as the main reason why language processing tends to be so effortless, accurate, and efficient.
Predicting reality, under the "controlled hallucination" framing, corrected by sensory error signals. The brain has no access to ground truth, only input data that helps correct the hallucination.
Based on lots of human interactions, I think there are a lot of human beings out there who mentally aren’t much more than “next word predictors” who happen to be made of meat+neurons instead of silicon+code.
I don’t think that’s what’s going on. I notice flaws on day one of model releases. But I also notice improvements if the model is truly more advanced than what I’m used to. Then over time the same questions or tasks return worse results.
What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
>What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
As far as the API goes, it would be really obvious. I run a small service that uses LLMs extensively, and if a model suddenly dropped in performance it would be straightforward for us to prove it. We regularly run comparisons where we generate completions with alternative models to e.g. see if we could get away with using cheap models for easy cases, if the baseline outputs deteriorated it would be all over our metrics.
> What is actually stopping these model companies from running a model at full capacity on release then once its name rings out, start serving users quantized garbage?
...I mean, if they were actually doing this despite saying that they don't—promising one product and delivering something else—I think that would be fraud, no?
And, maybe it's one thing to secretly defraud normies like us (although class action lawsuits do exist), but I don't think major enterprises or the US military would take too kindly to it.
Are you telling me that companies might defraud people for millions and billions of dollars and pay fines that are 1000% less than their profits?" My goodness, you must live on a hell planet.
Sorry there for the smarminess but fraud is just a standard business practice these days and fines are the cost of doing business.
And I really am all for someone suing these companies forcing discovery so we can see how the sausage is made and how many eyeballs are in it.
The question isn't whether the penalty would be less than their profit, it's whether the penalty would be less than whatever they make by secretly downgrading the models (or whatever it is you suspect), which remember also causes consumers to get less value out of the product and more likely to cancel.
The reputational hit, if this was to be confirmed, would also be massive. And I do think it would leak! Some employee would say something.
You can say this about any company in the world, selling anything.
It's trivially measurable, and there are people running the same benchmark on the leading models every day and measuring if they degrade. Spoiler: they don't.
But you can always say "the conspiracy goes higher", and that the companies know about these daily benchmarks and are routing them to "quality" envs.
Interesting approach. I would like to see a benchmark like this without web or mobile app code. In fact just: systems code (operating systems, drivers, low level applications), libraries, native apps, firmware, robotics software, signal processing, embedded code, bare-metal, RTOS, hardware design languages, EDA tools, etc.
I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”.
> I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine.
These machines were designed with the express purpose of being condescendingly sycophant to the point they hallucinate just to state "you are absolutely right".
No wonder some people even find these chatbots to be wife material.
And it disagrees with me a lot, granted this is in the webchat which I don't use for programming but for general usage seems to work, it isn't sycophantic
I don't have those system prompts, but gemini pushes back on its own quite a bit. It'll tell me when I wrong, or when there are better options to consider. They won't be exhaustive, but good enough for 95% of my queries, so it's also my go-to web chat.
Random humans don't have a significant cross section of human knowledge available in real-time, although many like to pretend they do, especially in internet comments :P Being able to compete with the capabilities of a median human would be an absolutely world changing achievement.
What do you mean? Virtually all humans have access to the internet. That's literally "a significant cross section of human knowledge available in real-time".
Oh, the median human can't process that in realtime, you say? Looks like they can't compete with the capabilities of the frontier AI then.
Yep - I like to phrase it as "AI is better at most tasks than most people". AI will still be beat at experts at specific tasks, but in general I find it to be better than me at the areas where I have no expertise. It's the ultimate generalist.
Have you? It’s a dumb model that gets a lot wrong. Codex voice is IMO only good for when I cannot dictate into the good models. The delay introduced by letting it work with a good model kills it for me.
Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.
Unless the company is going under, why not? Let's say Google releases Gemini Pro 4 tomorrow, and it's better than Fable and Sol; lots of people would switch over to it.
AI models are almost completely interchangeable, so the best/cheapest/fastest whatever will always have a market.
I agree we’d switch to it. I guess what I’m doubting is if a company can recover from that sort of stumble in the first place.
And they might not want to either. They might think there’s more value somewhere else besides trying to get back to the absolute performance and capability frontier. Smaller models targeted to specific domains that large models would be too inefficient at no matter how large they get or how clever you are at distillation, for example.
reply