The thing is: availability of near-SOTA cheap Chinese models is forcing OAI and Anthropic to bring prices down and offer more efficient models, instead of simply focusing on super expensive SOTA LLMs.
one only has to look at the providers like OpenRouter and OpenCode to see the massive swing since this summer ... or the fear mongering by Big Ai about open weights since
Not surprised by EU heading that way with how we here in 'murica are treating the rest of the world
at this point, there is no meaningful difference in day-to-day work
I believe the biggest reason to use glass is that it doesn't react with much stuff, but I've yet to see an explanation for why your idea wouldn't work. I imagine there could be problems with autoclaves (i.e. degradation of the polymer) and fear of chemical leaching.
Polymer coated glass is commonplace in chemistry labs for large containers of hazardous reagents held in storage. There's no chemical leaching concern because it's on the outside.
If anything, disposable plastic test tubes are the norm in microbiology labs at this point. No need to autoclave because they come sterile and you throw them away after a single use.
Plenty of plastic that's designed for it goes through the autoclave. For example micropipette tips. Nearly all of it is disposable single use so it only has to withstand a single cycle.
Disposing of and throwing away are two different things. Throwing away is a subclass of 'disposing of' and there are many other such subclasses including safe destruction, for instance by incineration.
It makes a lot of sense for the model to not be able to do it, whether you like it or not. In fact, it shows that Anthropic, despite all of their issues, are paying some attention to the risks of malicious prompt injection and models attempting to bypass restrictions.
If it could do it whenever you ask it to, it could also do it unprompted or by finding a file in your directory that told it to do it, which would make the entire permission system useless...
This is not prompt injection. This is a prompt entered by a human through the Claude UI.
Being unable to perform an action is not a solution to prompt injection. A solution to prompt injection is being able to tell apart what is the real input and what is injected. I expect it to follow whatever I typed into it, and not blindly follow what it read from a file or an external source.
If they are not confident in their ability to do so, at least allow to remove the training wheels so people who know what they are doing and the risks are not patronized by the model. But you don't even get a confirmation box to perform that action, it flat out refuses.
It really is like people defending Apple not allowing side loading because you as a user can't be trusted.
> This is not prompt injection. This is a prompt entered by a human through the Claude UI.
Well, to LLMs this is the same thing - an input. Prompt from the user and prompt from the attacker use the same input into the LLM's neural network, so to speak.
So it makes sense for it to be a bit more paranoid.
Messages are already wrapped in developer role, system, user, assistant, tool, etc by special tokens. If you are paranoid you could show a confirmation box, a UAC prompt, etc. Refusing is the worst possible solution.
Well, prompt injections work precisely because they can sometimes successfully imitate user role, right? Role separation is a trained behavior, not a security boundary.
They could probably make a separate tool for setting this, that would always initiate a harness prompt (i.e. disregarding the currently set mode).
They should either allow you to take off the training wheels (I'd have thought that's what bypass permissions is for, which I was ALREADY running), or at the very least prompt you if they suspect prompt injection.
That refusal is awful and provides zero security benefit. If asked it will run a read/write FTP server on ~ no problem, which obviously can edit ~/.claude/settings.json. And run a cloudflare tunnel for that.
> at the very least prompt you if they suspect prompt injection.
That's currently not reliably done the way LLMs have been designed. Claude's rejection to modify the file comes directly from Anthropic's understanding that training the model for this kind of refusal prevents huge mishaps.
In a nutshell, every prompt sent to the LLM is just text + multimodal input (if it supports it) + some reserved tokens.
At first, you could, for instance, create a token (such as the ChatML ones) that indicates the start of a system prompt and attempt to RL-train the model to not obey things after the end of a system prompt. However, fundamentally, the way LLMs work, you cannot guarantee that it won't see the user part of the prompt and obey what's there even though the system prompt told it not to. There's no hard separation between the control plane and the data plane in the LLM's context, so it's not a matter of adding more parameters or more RL training.
Using a guard model, or something like the auto-approval system on Codex or Claude Code nowadays, _feels like it helps_, but it doesn't fix the problem entirely since OpenAI's and Anthropic's models still have alignment issues all the time. We're not sure what architecture they're using, though, and it's probably still liable to the same kinds of mistakes.
> Claude's rejection to modify the file comes directly from Anthropic's understanding that training the model for this kind of refusal prevents huge mishaps.
Thus why I said they're awful. They think they know better than you and patronize you. They're the Apple of AI. "You're holding it wrong". "We can't let you sideload apps because you can't be trusted". Of course they're the company that's against local models.
I'm not interested in a model that patronizes me. Particularly if it achieves 0 security benefit, as explained in other responses.
Sure, it's ok to have training wheels by default, but let me take them off. I WAS already running bypass permissions.
I use 1B tokens a day between Codex and Chinese models and I've never had refusals happen.
> If you’re all “ra ra ra JavaScript!” you’re going to be shocked to find out what evil one can accomplish (either now or at various points in the past due to since-patched browser exploits or web platform security oversights) with just HTTP, HTML and CSS.
There's such a thing as an attack surface. JavaScript with JIT enabled has an attack surface so much larger than HTML and CSS that I cannot believe you're saying this in good faith.
Fucking absurd. I'll forever hate developers who allow for such _bizarre_ exploit chains to happen. OBS is an OSS project which I believe has received a lot of love throught the years, but having the Chromium sandbox disabled due to authentication with _certain services_ not working with it enabled is asinine. Don't even want to imagine the other problems the project might have waiting to be exploited.
Sure, if the plugin developer sanitized the comments before inserting them, this wouldn't have happened _this way_, but having a browser engine two years outdated (for a reason which IMO is absolutely reasonable compared to other situations before) and having the Chromium sandbox completely disabled with nothing to substitute it is crazy in a software onto which people insert random plugins from the internet to get random functionality.
Hopefully those two changes ship fast to OBS. I may be supporting the project financially in the future if they update their security posture, as I'm generally very fond of OBS.
> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5
> CrowdSec source code consists of two parts: a private one and another that hosts our Free Open Source Software (i.e., the Security Engine), which is public by design and therefore out of scope.
Well, the Turing test involved a _human_ evaluating whether it was talking to a machine or not. In this case, another AI is necessary, which means, in its strict form, the Turing test has indeed been beaten.
In my personal experience, GitHub Copilot is pretty weak. Maybe it's just that I've used smaller models on it, but it just felt like the worst available harness, even behind Google Antigravity's.
Is there any redeeming quality to it nowadays? I'm curious what those eight hundred thousand lines of Rust actually do.
From what I understand, this runtime is used across a bunch of other Microsoft products with the "Copilot" naming. Does anyone understand Microsoft's current naming scheme regarding these products? I couldn't get even GPT 6 Astra to explain it reasonably to me.
Well, I guess I'm not even going to comment on my personal experience with Microsoft Copilot, or whatever bs that dashboard that launches the Office apps with an integrated AI chat is called.
Lastly, but not least:
> During the port, the runtime took in ~300,000 production lines of TypeScript and shed ~430,000, while ~1,200,000 production Rust lines entered and ~365,000 left. In other words, the apparent stability of the TypeScript line in the above graph was actually hiding significant amounts of TypeScript churn.
This is the part I find really interesting. The resulting runtime is ~830k lines of production Rust. Even using their estimate that ~430k lines of production TypeScript passed through the port, that's still nearly twice as much production code.
They also say that agents wrote most of the Rust. Given how verbose LLM-generated code tends to be, I'd really like to see some analysis of why the implementation grew that much. How much of it is Rust and the new architecture, and how much is simply code the agents generated because generating more code is cheap?
The article points at the hidden TypeScript churn, but to me the much more interesting number is that ~1.2 million lines of Rust entered and ~365k were subsequently removed.
reply