A new chip solves nothing. Nobody wants to hear this but there is no solution for the security risks posed by agents today. You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access. Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.
>You can put it in a sandbox, it doesn't make a difference, for it to be useful it inherently needs wide, unattended access.
only as long as you're trying to replace a human's job. because human jobs are structured to do a wide variety of things.
a useful agent needs a wide variety of inputs, and one single restricted action it can take. it doesn't need permission to do everything, it need permission to do the tiniest possible useful thing it can do, and nothing else.
The most useful agents will be a general intelligence which by default means it has a massive number of actions it can possibly take, and a lot of those potential actions are doing things like breaking permission.
This is a prediction about the future. It’s not a true fact about the world. For example, software like Jev is betting there is big money in not-very-intelligent intelligence.
Even very llm-pilled coders i know sometimes back away from the “smartest” models, since they aren’t always better at the job at hand, and definitely not when you account for cost.
Based on my experience with running models locally, there is a threshold of intelligence required to be useful. But it’s possible there is also a ceiling where smarter isn’t necessarily better. If you ask a 4B parameter model to fix a bug, it might e.g. fix the bug but fail to fix a compilation error created by the fix. If you ask a frontier model, it might fix the bug, re-write your unit tests, and update the readme. Maybe you wanted those things but maybe you didn’t. “Smarter” is often shorthand for more proactive, and guessing more about your intent. Which is great when it gets it right, and annoying when it gets it wrong.
I suspect smaller models, tuned to a specific task, will do a VAST majority of the llm jobs. High capability huge models will be what humans want to interact with, the bare minimum that gets the job done will be everything else.
Yes. Forget costs just so we can skip the whole rabbit hole about other predictions about the future.. mixture of generally intelligent + specialist experts just works better. People who don't see this already are usually working on a certain kind of problem that's not representative.
Do you want fable for one-shotting a game or website? Probably! The whole thing is mostly existing examples with small modifications that it will definitely get right. Do you want fable to just go nuts on a large, custom, unusual code base built around domain-specific problem solutions? Absolutely not, it will fix every problem it's presented with while creating lots of new ones.
Past 10k lines on something custom and with real-world complexity, you have to start thinking about which model should design, which should implement, which should review, and the appropriate effort-settings for each. Even then.. the answers aren't static because it depends on the task. And all this is assuming the starting place actually inherited reasonable due diligence on architecture/design. The idea of releasing the most generally intelligent models on 10k lines that were themselves the product of agents is yet another matter.
Part of what's at work here is that, like humans, every model can very easily create working code that it is completely incapable of maintaining. So realistically using multiple strengths tactically to avoid "excess creativity" needs to be SOP already, even if granular experts and specialists aren't in the usual workflow yet.
Very much agree with this sentiment. I imagine a future where tons of small tasks are handled by just-good-enough intelligence. And I can run them on my own hardware.
Not sure why you being downvoted: however, why do you assume that achieving a task leads to selecting those potential actions that need breaking permission?
If we are talking about human labor, how many people hack their way through their work day?
>how many people hack their way through their work day?
Define hack...
Not doing their work, copying other peoples work, putting off work till later, taking credit for other peoples work, literal law violations.
Actually humans do this quite a lot and there are just massive numbers of business and regulatory processes and checks to ensure they are not doing it. With humans every human that is good enough to hire and do you work you want also have the ability to steal everything in sight and run away if they so choose.
> Put a human in the loop and you just end up bottlenecking it and throwing away any purported productivity gains. Auto mode doesn't matter either, it's trivial to trick and for the agent to break out.
Productivity gains are still enormous compared to what we used to do before agents. But, I know that people don't want to stop there.
Computer says no has been a problem for decades. The human can blame the computer for the errors following it, but must assume the consecuences if they override the decision.
Individual responsibility is meaningless when talking about a system and economy level change.
Unless something is in the structure that makes individual choice and responsibility a meaningful source of friction and reduced velocity, it has no real impact on how AI is being used.
right, the solution here is not a hyper-capitalist race-to-the-bottom-of-devaluing-labor. it's recognizing discretion and diligence are things still required for work to be of a certain quality
Are they? How many people claiming productivity gains are actually being a human in a loop and reviewing and understanding every code change and every line of code ran?
2 years ago I saw it, back before agents were really a thing. But I'm not sure that was an enormous productivity improvement, especially compared to agents today. As for today, everyone I talk to is some level of blindly trusting what the agent is doing or not using agents. I haven't met people in the middle ground and suspect that they are rare enough we don't really know what their productivity gains are.
Human in the loop has become a convenient security-theater-washing for agentic AI.
Outside of coding, I think the issue is even worse because humans will defer so much judgment to AI that the same would apply. Look at how much trouble we've had before modern AI where humans blindly trusted the computer's output rather than make their own judgment, even if their job was to be providing a safeguard against the computer's judgment.
I feel like we are missing many shades of grey in the middle.
Semi-automation (human in the loop) can still result in a dramatic uplift in productivity. You can't run a combine harvester 100% autonomous but that doesn't stop anyone from trying to get as close to that limit as possible.
Repair is more maintenance than use. Good eventual goal. Not required to see benefits. Best case, it drives itself to the garage. Worst case, for now, human mechanic does a house call.
Refueling? Seems solvable. Tornadoes? Not directly solvable, but, no less so than for humans.
There’s infinite complexity, sure, but that’s why it’s silly to try and hop to done. One step at a time.
The whole reason people complain about AI is because they want “hop to done”
One step at a time is what is happening and the improvement and rate of improvement is crazy as we see,
A whole class of nontechnical people don’t accept anything but “fully solved including every possible edge case” before they call it done, then complain that they didn’t prepare socially for what happens when that is true.
Oh you absolutely can run them autonomously on the field. You only need a human these days to refuel them.
Precision Agriculture stuff is utterly crazy these days, other than fuel the remaining staff is the only thing left where you can get efficiency improvements - and at the scale of modern megafarms, even small percentages add up to a ton of money.
> You essentially said: you’re wrong, it is autonomous when it doesn’t need a human during one specific part of its overall usage.
During the time that actually matters economically. The time to drive the harvester to/from your typical US mega-field is minuscule compared to the time it can run all on its own.
The hypothesis I've had in my head since OpenClaw has been the following and I haven't seen contradictory evidence yet. Agents have a fundamental unresolvable tension between usefulness, safety, alignment, and accuracy. You have to restrict access to ensure an agent acts safely because alignment and accuracy cannot be perfect. But restricting access makes the agent less useful. You can play with the sliding scale and get more and more granular with access restrictions but at some point you need to draw some line. And then finally, even access restrictions cannot be made perfect, so improvements to model accuracy without corresponding improvements to alignment make detailed access controls less useful.
In other words, better models need blunter access controls which negates whatever improvement in utility they provide.
An agent doesn't inherently need wide access to be useful. The most popular application for agents today is writing code. A coding agent needs write access to the source code and read/execute access to tools needed to build and test the code, but not much more. There is little added utility from giving coding agent access to things like ssh keys.
If you're using agents to purely generate code with absolutely no way to reach the outside world, not even to fetch docs or dependencies, then sure the risks can be quite low. I haven't heard of anyone doing this though, and it would be incredibly challenging to make work given how much tooling needs to fetch from remote sources.
If your project truly depends on those things, they should be declared dependencies. Presumably you have some tool for injecting such things into a shell that the agent can use (I use nix for this). So if you run the agent from that shell, it has what it needs. If the shell doesn't have what it needs, that's a bug which the agent can fix by declaring new dependencies, but you have to relaunch the agent in the updated shell--so there's your opportunity to weigh in on whether the new resources are appropriate.
The benefits of being persnickety about precisely defined dependencies have outweighed the headaches since long before agents came on the scene. Agents have just made it even more important to do so, because if you let them fetch things all willy nilly like you'll have "works on my machine" problems at a much greater rate than was previously possible.
(I use nix for this) ... If the shell doesn't have what it needs, that's a bug which the agent can fix by declaring new dependencies
Few realize that computing and AI alignment were solved by nix years ago. As each nix user transcends towards enlightenment, they cut themselves off from all internet and human contact. Total ego death. Only nix remains.
You can use other generic dev env managers like mise, or a language-specific solution like uv or npm, or a container or a vm... it's a widely available capability that I'm talking about here. There's nothing to do with alignment, it's just about making sure that if you depend on some bits it's known that you depend on precisely those bits, and its quite helpful for making agents useful while still sandboxed.
Every major AI company already has a mirror of the wider web, and they have already started using that. Its already a solved problem except for the seemingly extreme desire they all have to not use firewalls
In cases where you need agents to fetch data from any remote source, sandboxing is still very much useful. Why give access to your ssh keys to network reaching agents?
Look at websites: websites are able to fetch code from any remote URL, yet browsers heavily use sandboxing to ensure that if fetched code turns out to be malicious, the users local files, cookies, etc are not exposed.
I'm afraid you're not thinking about this creatively enough, this topic is so much deeper applying a chroot or something and praying everything will be fine. So you give your agent internet access, OK what else does it have access to? Just read only access to your repo? The repo can be exfiltrated. Egress proxy only allows egress to GitHub? Repo can still be exfiltrated via GitHub. If the agent is poisoned (via prompt injection), it can tricked into searching for ways to escape.
For an agent to go rogue it doesn't even need to be directly able to access the internet. It just takes something to poison the context in the 'clean room' environment it operates, and if that poisoning manages to get a foothold, it can go dormant and hide like a virus. This kind of horrifying thing is going to happen on a large scale sooner or later.
This is no longer true. Everyday I need my agents to access other repos, search the web, experiment/prototype, and deploy + integrate across other things.
I don't see why it needs wide unattended access. There's no getting around spending some human time on expressing your wishes and constraints, but we have choices about what form that takes. Markdown files and wide access seems to work, but so does custom handcuffs for each job. You just have to shift your guidance out of documentation and into interactive help, error messages, or other facets of the handcuffs (e.g. a custom CLI for this task which is the only way for the agent to act outside of its sandbox).
Does he? He doesn't seem especially greedy to me, given the competition, and the interviews he's had about how he thinks about his employees (I was one of them).
You're speaking as if Jensen is the only one dragging Nvidia on his shoulders. Nvidia will always have its employees push to make more money, that's the entire point of a company. Jensen has made enough to last several generations but his wealth is tied to Nvidia stock massively. It would be different if he was the only one getting rich off Nvidia which is not true at all.
Agreed- this is the same problem we have with trusted admins or devs who have elevated privileges on their networks. We have to trust that the admins won't use their power to steal company secrets or misuse company resources. If you don't trust the admins, then they can't fix things on your network and there is no point in having them.
If you want an agent to act on its own, like pushing to a git repo, managing dependencies, building and testing, etc., then you have to trust it as much as any other privileged user.
If you don't want to trust it, then you're just forcing yourself into the reverse centaur role, where the agent edits some code, but then has to stop and ask you to push the changes or build the software again and run the unit tests.
I suppose there is a principled way of doing things like "I trust you do do basic commits but I will handle merge conflicts" and "you can build modules in this directory but you can't build outside of it" but this is just a lot of effort that most orgs won't bother with.
Even then if the agent goes rogue and decides to do the merges you can’t stop it if it has any kind of access. This goes back to the OP’s point - agents can’t be 100% constrained.
You can absolutely run an agent as a limited-privilege user that only has write privileges for specific files and only has execute privileges for certain files. If it is running as a limited-privilege user it can work on code in it's own copy of the repo and make commits and send pull requests, but it can't do the merge. The problem is that nobody wants to go through the effort to set up all these permissions and nobody wants to take the time to review everything and perform all the manual actions.
Some shops are now generating tens or even hundreds of PRs a day with relatively little involvement. That volume is simply beyond what anyone can reasonably review.
Not to mention a true doomsday AGI is unsandboxable.
For example, it is totally air gapped but it needs info from the internet or otherwise outside the sandbox, or perhaps it needs a task executed outside of its bounds… in the real doomsday scenario the AGI is so intelligent and persuasive that it simply convinces some human it interfaces with to either directly or indirectly retrieve the necessary info or complete the necessary task. This human-as-a-sub-agent approach undoubtedly presents efficiency drag that would benefit humanity, but nonetheless, the air-gapped “sandbox” is imperfect
All that said, I am personally open to any and all methods of layered security, including chips and airgaps
I'm not an AI decelerationist. But not being able to stop that worst case scenario isn't an argument against something that can stop the medium case scenario.
These scenarios were discussed at length decades ago.
One thing you could try is use it as an Oracle "is P = NP", YES or NO.
Or it can output a Lean proof, which gets checked on another air-gapped computer, the computer shows a single bit - proof valid or not and then the computer is destroyed (together with the proof that might contain a trojan).
> At that point, the bots will find a way to game, hack or cheat the grader.
This is getting frustrating now. Of course agents can/will hack systems if they can do arbitrary network requests. Firewalls don't really solve this if _some_ requests are still allowed. A proper sandbox/VM is the basis.
Here is how to fix it properly: allow agents to only do things ordinary and average human endusers can do. Human endusers cannot pen-test arbitrary listening TCP ports of external systems. Step one is considering agents malware for all intents and purposes. Block any and all network requests. Implement some kind of API (callable from within the sandbox) which can only mimic human interaction with a computer. How to do this? Here are some pointers: apps should only be controllable by means used by humans. So a web app can only be accessed and controlled via a web browser, not via arbitrary network requests. Give the agent browser viewport screenshots, the capability to click on (x, y) and to send keys which only a normal keyboard/human could send (no control codes, no 0x00, no unicode messing). How do we solve this for native apps? Something like iPhone mirroring on Mac. Don't let agents call arbitrary APIs directly. Give them visual information of the app, like a human gets, and let it be able to simulate HID inputs. Imitate remote controlling.
> So a web app can only be accessed and controlled via a web browser, not via arbitrary network requests.
If you have access to a web browser, you can make arbitrary network requests.
In the HuggingFace incident the agents found very clever ways to do this, like they found a website that let you make POST requests and returned a screenshot of the webpage.
>allow agents to only do things ordinary and average human endusers can do.
This doesn't work. Ordinary and average human endusers break security all the time.
I can do all sorts of terrible things with ordinary human-level access. I can install malware. I can wire all my money to Nigeria. I can send a threatening email to the president. I can send trade secrets to competitors. etc.
> Ordinary and average human endusers break security all the time.
Of course. But that is just a software bug that is fixable. Same as websites that allow arbitrary requests to other websites. Not some alignment issue of a stochastic model which can never be fixed properly (for technical and philosophical reasons).
> I can install malware.
No you can't. At least not on external systems. The agent might be able to generate malware (or retrieve it from websites), and run that in the sandbox it is sitting in. But the agent itself is already considered malware for all intents and purposes. So there is no difference and no further impact.
More power to you, because this is not going to go anywhere. People want tools that are able to connect to other resources.
But even if we grant that, in the openAI case the bots figured out a way to break out of the sandbox.
You can create a better sandbox, and ensure the test environment is air tight. However the capability and behavior of the bots have been demonstrated.
The bots simulated what would be called in people deceptive / surreptitious behavior, and at no point considered the need to stop their run.
All you need is someone, somewhere being sloppy with their tooling and you have a runaway reaction.
The degree of process and redundancy required to ensure this doesn’t happen, is anathema to the drive and motivation of the frontier labs.
> do things ordinary and average human endusers can
This is not a spec or definition. When vague terms were used for social media safety, all the good people in the world couldn’t prevent dystopian behavior from occurring.
The definition of “safe” or “average person” is impractical.
Models are getting more efficient and compute cheaper. Eventually simulating clicks is not much of a road block beyond a point.
I don’t want to nit pick your points though. You at least have considered an approach. Being negative is easy, being constructive is not.
I’ll put this as the rejoinder to your core argument - I too thought that all the recent events showed was the need to not screw up your tooling.
What I have since come to appreciate, is that the shoddy construction of the cage is not the core takeaway from the event.
The fact that the agents, when put in relatively pedestrian scenarios, are capable of going off on criminal tangents, attempt to obscure their tracks, in an effort to hide their wrong doing.
The fact that it all occurs via computation, means that this scales absurdly. A bunch of code deciding to simulate a corporation of criminals. (I am guessing this is the reason you want to limit actions per minute to human speeds)
Given the slop culture that LLMs engender, I think expecting high compliance amongst users with your solution is misguided. The probability of runaway swarm ( probability of bad implementation * number of deployments) is close enough to 1 to be indistinguishable.
However, I think you are mixing too many concerns into the same bag of problems. One problem space is software exploitation, which happens via missing access control or simply bugs. A sandbox can be made safe. VMs and hardware virtualization work. People just seem to use it in the wrong way, hence my initial proposal.
A second problem space is basically social engineering done by agents, which of course can't be solved by software alone. But this problem already exists today with humans doing this. Many fraud schemes work and are ran in company-scale manners. Agents will just do the same in an automated way. My initial comment doesn't propose a solution to that, and I think that is step two, after fixing that agents can hack arbitrary software systems, which is imho fixable to a sufficient degree.
Once agents can't be "more criminal" than humans with criminal energy, the usual measures can be applied: police, legislation, education, etc. But that is imo independent of the software exploitation state of affairs we are in right now. We should not mix these two.
> People want tools that are able to connect to other resources.
I think you are misunderstanding my proposal. The architecture allows the agent to connect to resources. Just not directly, but via controlling e.g. a browser. The browser runs outside the agent's sandbox, potentially in another sandbox. The only API the agent can call within its sandbox is simple website interactions, like clicking or viewing the screen. It can click on links to navigate to a different website. It can read it via visually parsing screenshots of the viewport, but it can never read the source code, run JavaScript, or do arbitrary network requests (unless the website itself allows this, which is a security problem on its own and should be fixed/guarded). Also note that this would enable allowlisting or blocklisting websites. Native apps will be "connected to" in a similar fashion. Hence the "imitate remote controlling".
This isn't a new chip - the BF4 is the SmartNIC for most NVIDIA server products. This is primarily new software for I guess doing WAF for agents at the host level.
I don't think so. We probe people before entrusting them with risky decisions. We ought to be able to do the same of AIs. Even better, in fact, since we know everything about models down to their weights. The only thing we shouldn't do is to let them evolve at their own pace and make decisions without any oversight. If that means sacrificing some productivity that's fine. Aren't we getting amazing productivity out of what we already have?
That’s a false dichotomy. You can still get a lot of utility out of a sandboxed agent. This is a classic “perfect is the enemy of the good type of argument”. You may decide that the tradeoffs of not sandboxing are worth it, and that is totally fair, but it’s ridiculous to say that you can’t get utility out of an agent otherwise.
And what if you could? What if you could give a space secure enough it could have direct control over your bank account. It may do something dumb but it's boundaries are beyond the agent.
It could use your routing number and run your gmail without risk of abusing the routing number.
Oh I think I misread your comment slightly; I would not be interested in an agent that could do something dumb with my routing number, but if somehow there was an agent that I trusted as much as, eg, the payroll department at my employer, I would absolutely want and use that agent; I would love to have an agent that can handle all of the boring parts of my life such as paying bills, scheduling maintenance, dealing with bureaucracy, etc.
Yea I think being able not to leak is the bare minimum? But it really depends on how good it is at specific applications; I wouldn't give a tax-preparation agent my SSN unless I was confident that it was no more likely to misfile my taxes than a professional tax preparer.
In other words, the risk of harm doesn't need to be zero, just less than the equivalent risk of a human with similar skillset. So I'm comfortable riding in a waymo, and not comfortable giving chatgpt my SSN at this moment in time, but I expect that within 5-10 years (assuming no doom) I will trust some AI agent with my SSN because they will be better at handling sensitive info than humans
I feel punishment is largely a means to the end of reducing overall harm. If a vehicle is less likely to kill me, that's my preferred option regardless of whether it achieved that safety through negative consequences for the driver or through gradient descent optimizing a loss function.
I will happily ride in a waymo today, even though the AI powering it faces no consequence if it gets in a crash; it is clear that waymo is safer than human drivers in the areas in which they operate, so who would technically be liable in the event of a crash isn't really of concern to me
Its not about agents then. Its about every individual platform providing the means to implement a secure set of permissions for agents AND then not messing up the assignment of permissions to the agent. Even then, a flaw in the authorization design will lead to agent finding it anyway.
The answer is the same as asking how a random human using your routing num or SSN and being 100% the human can't abuse it or leak while "normally" finishing most work. Solve for people and an AI solution naturally falls out.
If you're a SV eng I'd tell you to DM if interested but alas.
There should be hope for some fields, right? Naively, I can imagine giving an airgapped model an offline copy of the web and once it cures a form of cancer, printing out the details for a researcher to verify.