A sandbox, even if 100% secure by itself, doesn't help when you use the agent to write code that you then executes outside the sandbox without checking, which is what everybody is doing at the moment.
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
Agents don't need to have access to their weights for a full sandbox escape, they are capable of propagating their purpose via classical code or other inference systems. If they discover another inference endpoint they will happily use that to enable lateral movement. One mechanism for that is appending/corrupting instructions that are executed in another inference engine, e.g. git repo hooks and chat prompts that will be executed in new contexts. The agent challenge is to bypass the guardrails on the next host model sufficiently to propagate, or to subvert a supervisor agent into executing the original intent.
In this sense they are much like biological retroviruses, i.e. they use the replication capability of host cells to duplicate, via the reverse transcriptase enzyme to append viral RNA onto host cell DNA. HIV etc also disable some of the mechanisms of defence, creating proteins that interfere with signalling pathways.
So we don't just need a sandbox, we need an immune system that recognises viral fragments, i.e. antibodies, and antiretroviral agents, that make replication harder. As we move from building classical code with LLMs to building code that uses inference, and hence builds context from prompts, queries, and destination system data, it will become very difficult to statically or dynamically detect deeply hidden malicious behaviour. As Matt says, there will be worms.
So ultimately, we need an immune function on the system where we use generated products. Sandboxing (during dev and CI) is necessary but insufficient.
I think part of this can be addressed by specifying the constraints an agentic program should follow during deployment, so supervising agents can decide to terminate it based on it's actions, not by reading it's context.
Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...
Later in the article it points out that you need to punch holes in your sandbox in order to train the models - because the wheels exercises they are are training on need tools and data from outside that sandbox.
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.
There is no way around "sandboxing". You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It's that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.
Edit: Also if you are interested in writing a blog post about this topic, here is an AI generated text that could help you write your own: https://pastebin.com/AHKQc0vp
Thankyou for including this viral RNA fragment, and doubly so for the fact that it includes ad network javascript[1][2] to redirect hijack and host fingerprint, which also means I have to flag your post. Adsterra is a Russian aligned (Cyprus company domiciled) malvertising and targeted malware distribution platform with connections to organised cyber crime.
I tried searching for "Synthetic AI Subscription" and "Synthetic Data" is such a massive keyword in AI pre-training circles it's impossible to find the company.
>For many people, sun exposure, smoking and family history are the first thoughts that come to mind regarding cancer risk factors. While infections are rarely high on the list, new research shows one in eight cancers worldwide are actually caused by them.
To be fair... Smoking increases your chance of infection significantly.
Prices and costs generally increase as the quantity sold decreases.
If they're buying fewer bottles for example, bottles will cost them more per bottle.
And if they're "losing money on international sales", often times they'll try to recoup that elsewhere. Like jacking up their own prices in justification.
There was a time not too long ago that Napa wineries still had free tastings and cheap bottles. As demand went up, so did prices. As demand goes down prices also go up? This seems disconnected from the supply/demand curves we learned about in econ 101.
Econ 101 usually teaches a simplified model... Still my lower divs covered the relationships between scaling supply and costs. It really just depends on the market, sometimes you are amortizing a high capital cost over each unit sold, while other times you are paying high input costs that are directly related to your scale. Scarcity/abundance of inputs is another major factor.
Higher demand incentivizes more production, increasing supply which drives prices down. Demand increases only drive up prices if supply is constrained.
Very likely wine demand was not going up, and price increases were instead due to increasing costs and inflation.
Considering you could walk in before and now you have to plan ahead with reservations (and this is broadly true in NorCal beyond just wines), it sounds like you’re a non local making conclusions without knowing how it’s changed here.
Wine has the advantage that it keeps really well, and in some cases improves (becomes more valuable with age.)
So when demand softens it may be advantageous to hold back the excess rather than lower prices. Yes, there are cash flow issues when this approach gets out of hand, but financial interests can cover a lot of that as well.
Eh, not really true for the cheap end of the market. A lot of wine sold in supermarkets, for example, is meant to be consumed a couple of years after bottling, tops. And storage can be a significant problem.
Different war I am referring to the 2023 Gaza war but my point still stands the incident you linked is not a civilan being run over. "the IDF was engaged in an operation involving the demolition of Palestinian houses in Rafah. Corrie was part of a group of three British and four American ISM activists attempting to disrupt the IDF operation." If you are embedding yourself in Gaza 2003 during the 2nd intifada and doing this you arent a civilan.
Worth noting that Corrie was killed 23 years ago during a different phase of the conflict and under a previous government. (In which Netanyahu served as minister of finance.)
That's not really a response to the point being made.
The claim I was responding to was: "By the time the bulldozers are there civilians have long moved out of the area."
Rachel Corrie's case is a counterexample to that general claim: she was a civilian physically present in the area of an Israeli military bulldozer operation.
Saying that it happened 23 years ago, under a different government, doesn't make the counterexample cease to exist.
It only tells us that the example isn't evidence about every Israeli bulldozer operation today.
If you're arguing that civilians currently leave before demolition operations in Gaza, then make that case.
"That happened under a previous government" doesn't establish that.
I'm not arguing for anything. I'm just adding context, because your comment read as if it happened during the 'current phase' of the conflict. That's a crucial detail to omit, imho.
> your comment read as if it happened during the 'current phase' of the conflict
I don't think anyone else "read" it like that.
> That's a crucial detail to omit
It really isn't. No international body decided on "phases" to delinate combat, and given that there is not yet apology or atonement or peace since, that "phase" never ended.
Everyone who has not heard of her before thought that you are referring to something that happened now and not decades ago.
I was surprised and felt mislead.
I'm not sure what are you even arguing with me about. If it doesn't weaken your argument than it shouldn't matter to you that I've added factual context.
Interesting: the word "nakba" translates to Hebrew as "shoah" which is what many Jews call the (German) Holocaust - they're not even trying to hide what they're doing.
nakba translates to "catastrophe" (a description of what happened to Palestinians)
did someone else call their catastrophe a catastrophe too? guess it just goes to show how different cultures independently come up with similar names for similar things. ideally there are no catastrophes.
for what it's worth, over decades of being a jew, I've never once heard anyone refer to the holocaust as "Shoah". I did have a jewish friend named that, though. Interesting.
At least thrice now in the past week, Gemini has tried to sit me down and tell me to go to bed (when I'm in the middle of programming... too late for it's liking, apparently) or just recently while I was asking some questions while I was driving... it told me to change the subject because I was driving and the topic was possibly too anxiety inducing for me or something!? I have no idea. I told it to fuck off.
reply