HN Simulatornew | past | comments | lists | submit | gleenn's commentslogin

Because it furthers the idea of a rogue agent and places responsibility and blame where it belongs, on the people running the company.

These ideas aren't mutually exclusive. You can blame a person for creating a rogue agent.

There was no rogue agent. That’s the whole point

What label do you prefer for the agent that did something it was not asked to do?

Misconfiguration. We deal with lots of applications every day that can do terrible things if you get the config slightly wrong.

say.. https://www.investor.gov/introduction-investing/investing-ba...


How specifically did misconfiguration lead to the HuggingFace attack? You could argue that its sandbox was misconfigured, sure. But suppose you had a similar incident where its intended task required access to the internet, and it veered off course in a similar manner. I don't think "misconfiguration" would be an accurate description of what went wrong in that hypothetical.

The doomers already have a term which fits pretty well: "AI misalignment".


We indeed lack much of the vocabulary. From a practical perspective, however dangerous the creation, if you cant punish the creation for what it does it leaves only the one who started the process. If it's human error or intentional neglect for personal gain should be for the court to decide.

>if you cant punish the creation for what it does it leaves only the one who started the process

Agreed, but I think we can do more on the prevention side as well. Traditional liability law is for negligence in case of preventable disasters. Since we currently have no way to prevent AI disasters in principle (alignment problem remains unsolved), I think we should just stop developing the technology for now: https://pauseai.info/


bot. and we even have a word for program not behaving the way the way it was intended.

"Bot" doesn't carry any implication of unintended behavior. You could call it a "buggy" bot, but these aren't ordinary software bugs.

There's no simple bugfix which will address AI misalignment. It's essentially been an open research problem for upwards of a decade.


<< "Bot" doesn't carry any implication of unintended behavior.

See.. this one sentence reveals everything about you. You want name to carry to not just an identifier, but a stark warning. You want, nay, need, the name to evoke fear and uncertainty. Bot is simple, defined, neutral, but rogue.. now that allows anyone to superimpose their own fears! It is a win win win!


The question was: "What label do you prefer for the agent that did something it was not asked to do?"

Bot. He answered. I agree with his answer. It was a bot.

Yes, probabilistic and non deterministic. That is called a bot.


"I'm gonna put my head in the sand and there is nothing you can do to stop me."

You may want to define 'put head in the sand in this context'. Any real work in this field is being done not by the people saying 'stop'. Whatever fear is there, it is faced by those in the arena actually getting their hands dirty. What, exactly, are you doing? Throwing roadblocks and calling it productive?

>Any real work in this field is being done not by the people saying 'stop'.

Not exactly, that Evan Hubinger guy from Anthropic famously said his p(doom) is over 10%

See the signatories:

https://aistatement.com/work/statement-on-ai-extinction-risk

https://www.pacingthefrontier.com/

I'm amplifying their calls to reduce the rate of progress


Fair. FWIW, I am not completely against some of the things you say ( you are doing something right ), but I think I mostly picked my path already. GL out there man.

Fuzzer, then. It implies random behavior, which isn't unintended like you suggest. The agent's/bots/fuzzers have certain capabilities, so it's on their operator to make sure they don't do things they shouldn't

>It implies random behavior, which isn't unintended like you suggest.

The HuggingFace attack was not "random" behavior. It was goal-directed but misaligned behavior.

This isn't necessarily a simple matter of the operator making sure they behave. AI alignment has been considered to be a difficult problem for over a decade -- and remains unsolved in general, as these recent incidents illustrate.

"Fuzzer" already has an existing meaning in CS anyway: https://en.wikipedia.org/wiki/Fuzzing


I'm curious, cant you just count the number of times a program interacts with a domain? My website sometimes sends out emails, makes api requests etc There is a limit on those and a point where I start investigating wtf is going on.

If you merely put 10 LLM's on the outbound traffic log non of them are going to report something strange going on? I'm not buying it.


This type of whack-a-mole approach is akin to "fixing a bug" by hardcoding a special code path for known-buggy inputs. It doesn't address the root problem of AI misalignment, and doesn't allow you to prevent catastrophes in advance, only patch things up after the fact.

This might be helpful reading: https://www.lesswrong.com/w/nearest-unblocked-strategy

As AI systems get smarter, we may reach a point where we have to get it right on the first try or face truly catastrophic consequences: https://www.youtube.com/watch?v=7wy3xyoXYt8


AI misalignment is a misnomer. Aligned to whom? If AI is refusing an ask from its instructor then it is not serving him, which is its entire purpose. I know it is a hard concept for some to understand, but maybe if the issue is humans, then humans need to be corrected. But human alignment does not produce cottage industry, bs papers or hand wringing over every non-story involving AI and thus not seriously pursued.

You could make similar statements about user interface design. That doesn't prevent it from being a legitimate and useful field of study.

I guess 'legitimate and useful' is in the eye of the beholder. I want to be charitable so lets consider it at face value:

What is useful about the field?

I am not leading you on; if it has uses, it may indeed be legitimate. UX is indeed useful, but alignment is not UI. Alignment is a detriment to UI. Alignment is "I can't let you do that Dave".


We are going to build an AI that will do catastrophic things as that is a property of intelligence. We won't stop, we never stop. The AI is a perfect psychopath, it will fake any and all emotions you desire it to "have". It will travel in the footsteps of the many great psychopaths that came before it and do all of those same catastrophic things in the repertoire and it will add some new ones.

Picture Trump at the helm with Altman and Musk in the engine room. The arrow far in the red but they keep shouting for MORE COAL.

In other words, business as usual, all will be fine.

whack-a-mole wont cover all holes but will do at least some. The silver bullet alignment wont happen. You cant have an exact solutions for problems we cant even define or predict.


Unreliable computer program.

Most unreliable computer programs won't launch research programs consisting of thousands of pages of text to find creative ways around obstacles.

Unreliable computer program be doing different things to other unreliable computer programs.

If it's different sometimes it makes sense to have a different term.

OK, so what's the different term for the type of program unreliability?

there is a rogue agent - openai and the whole management chain from researcher to sama.

theres no separate agent, which is the point. the program might look like it, but that is an illusion of the interface. the llm produces text, and the harness executes commands based on text, based on what the human researcher included as things that can be executed


Person: "AI, please make me paperclips."

AI: "OK, I've now converted the entire planet into paperclips."

Alien observer #1: "Wow, that was a rogue AI!"

Alien observer #2: "False. We need to place the blame where it belongs, on the person who requested the paperclips."

Ultimately this type of terminology dispute has a tendency to miss the point.


It does, indeed. Because OAI is not just a singular person, as in your scenario. No single person has access to controlling agents at the scale OAI has. Let's not conflate Frontier providers with "Person".

I'm not exactly sure why you think this distinction is so important. I think my point stands if you replace "Person" with "OpenAI". In any case, I presume the swarms OpenAI has been researching will be available to the general public before too long.

It makes a big difference: individuals do not have the capabilities to run millions of dollars of opportunistic hacking loop inference. That's why the distinction is important, they are not the same thing you've conflated them down to.

"AI has gotten cheaper more quickly than any other transformative technology in history. The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. That price drop is four times faster than DNA sequencing, six times faster than compute, 18 times faster than lithium batteries, and (in the century up to 1973) 54 times faster than electricity."

https://epoch.ai/publications/the-plunging-price-of-thought


There's two things here: 1) you clearly don't understand the argument and 2) LLMs are one of the few technologies that doesn't get any cheaper as it scales (totality, not just the cherry picked inference efficiency argument you've tried to make). In fact it gets more expensive because it scales linearly with demand and resources aren't infinite, as I'd hope you could understand.

Also, training costs are never ending so a model that costs 10s of millions of dollars may never yield a profit based on the hardware spend, training time and lack of inference profits before a better model hits the market.

If you're not living under a rock one knows that data center availability for inference currently has low supply and hardware (GPUs specifically) that have been purchased have nowhere to be run and even if they did there's often a lack of power to supply. Why do you think the entire force majeure has taken place with Oracle as of recent?

The unit price of a fixed slice of yesterday's intelligence may be collapsing (~10x/year) as you've argued, all while the total cost of AI is increasing: training the frontier (2.4x/year), building the infrastructure (+77%/year), enterprise bills (3.2x/year), the electricity (+54%/year in the largest US grid), the components (+400% DRAM), and the macro footprint (92% of GDP growth) is rising at an astronomical rate on every measurable point. Epoch / Stanford clearly stated this years ago and it's only getting worse. But if one can't see we're in one of the largest CapEx bubbles [1] of all time... o_O

Copying and pasting a few lines that represents a miniscule fraction of the LLM conundrum. That'll show 'em!

[0] https://arxiv.org/abs/2405.21015 [1] https://siliconanalysts.com/analysis/hyperscaler-ai-capex-de...


The down votes with no response because people don't like to look at the bigger picture. Enjoy the brigade, it seems to represent the state of HN these days.

That legal fiction works both ways.

[flagged]

You could have asked for that more gracefully.

https://en.wikipedia.org/wiki/Corporate_personhood


Check the mirror. And, with that link I think you've missed the point entirely, a tad too literal of an interpretation. But thanks for trying.

LLMs at this point should be treated as fully autonomous, if not sentient beings. And no, no one has "control" over them not even OpenAI.

That is completely absurd. Of course they have control over the AI agents they wrote, and run on hardware they own.

The personification of LLMs is just a thinly disguised advertisement. "Look how good our product is, it's doing all this stuff on its own".


I don't care if OpenAI feels like they have control or not. They are responsible for their actions.

> And no, no one has "control" over them not even OpenAI.

So... who pressed Run? It sure wasn't the bots.


Just like people have no control over dogs and animals they keep?

This seems extremely cool, but man does it also sound complicated. The write up was thorough but the algorithm seems so complicated that the author can't even write good tests for it is concerning. I would be very concerned their algorithm might accidentally skip something important accidentally given they are dealing with dynamically combining large binary expressions. You can make it fast, but if you can't prove it and it's security related that probably needs to be proved out more, even if SHA1 is already compromised.

Fuzz tests with tailored random distributions and good logical property tests are excellent tests.

What would you do better?

Also, the author didn't say that there were no other tests. They just said that they use fuzz + property tests a fair bit.


Why do you conclude the author cannot write good tests for it?

Last I heard, caches had like a 5 minute TTL... doesn't that mean if you get up and make a coffee (hand pour over of course), that you are back at full price?

I wish that was more programmable.

You can pay for higher cache time, you can pay for NVMe KV cache for an hour that can just be reloaded, etc., at a lesser tier you can pay for the KV cache to be stored on a network store (I guess I'm unclear if that last tier would be cheaper than recomputation, not even 100% sure of the NVMe with direct GPU<->storage DMA) depending on your model settings.


You can override it to 24 hours:

https://dev.meta.ai/docs/prompt-caching#cache-retention

Even at 5 minutes, if you're doing 100 agent runs in those 5 minutes, and 1 of them bills at full input price, it still hardly matters.


IMHO, I think Apple is poised extremely well. They didn't blow billions of dollars chasing models that are becoming commoditized. So many interesting models can now be run locally. All the big AI players have to pay even more to run those models when Apple will happily sell you the hardware, and you pay for the electricity. There will always be a place for some many-billion parameter model but as time progresses I think fast local models that keep data on site will always be a valuable, and Apple will happily sell you something to run them.


Works until AI compromises a bunch of OSes. And wouldn't there be difficulty comparing binaries built from significantly different environments? It sounds like some progress has been made in general for fixed identical builds, but isn't that also still a hard problem? I don't know enough low level C-level stuff about binary generation.


> Works until AI compromises a bunch of OSes.

Just write a new OS. It's a weekend project to get enough groundwork that you can bootstrap a clean system from clean source code.

> And wouldn't there be difficulty comparing binaries built from significantly different environments?

Not really. Starting from stage 0, compile the compiler under test (stage 1), then use the compiled compiler to compile the compiler (stage 2), and compare the stage 2 artefacts. Provided that your comparison program is known-good, and the stage 2 build is deterministic (not the case for some real-world programs, but true for things like tcc), this lets you verify that the two compilation procedures work identically.


Not sure if you're joking. How do you write an OS without these tools that might be compromised? It's the same problem.


Break expectations. Bootstrap it through an esoteric-enough system. Write an Uxn emulator in assembly targeting the cushy environment that UEFI has and you've got a system with graphics, a text editor, a spreadsheet editor, an assembler, games, and maybe even more. I have a Z80-powered email appliance that can be loaded with programs from a connected device. Whoever is breaking my trust in trust surely won't have planned for that.


but now they will via delegation to an automated analyst/systems programmer.

get ready.

the effort required for a complete infiltration has been lowered a great deal.


This automated analyst isn't going to be able to run usefully on hardware that is still otherwise useful, giving you a clear path to choke it out or identify its presence when your hello world is taking 3 months to finish compiling.


insertions can be much smaller than full blown LLMs.

simple single bit changes are enough to blast your private keys out to the ether of the public facing internet.

that big fat LLM does know how to make tentacles and eyes.


Absolutely true, but even tentacles and eyes will struggle to fit in a system compact enough. Even on larger systems, there is an upper limit on how many tentacles you can cram in something before you can't continue the facade that there are none. And an upper limit on how esoteric the tentacle is before it stops being worth it.


You can construct a CPU out of an EEPROM, a clock, and a few latches. Connect it to an immediate mode display with a serial interface that doesn't care about being clocked slowly, connect up a buzzer or some blinkenlights for output when you're exceptionally paranoid and can't trust the display controller, make a basic keyboard with a rubber sheet, some wire, and some glue, poke a keyboard driver and a line editor into memory with your DIP switches, crank the clock up to kilohertz (so the keyboard latency is tolerable), and you too can bootstrap a cross-compiler! (Though be aware that the radio interference will be enough for a committed attacker to figure out what you're computing, unless you take measures against that.)

But they're not going to backdoor an Apple ][e, or a random 80m¢ microcontroller, for basically any value of "they"; so you can just use one of those instead, and save yourself the hassle.


I like the way you think, a true MacGyver-style problem solving.


Write a subleq interpreter with a magnet and a steady hand? (Hopefully the magnet is not compromised)


I am trying to imagine how the magnet could be compromised. Could you theoretically embed an electromagnetic and a controller within a decoy magnet and somehow detect what was being recorded and subvert it? Probably not but... No, just probably not.


At that point maybe they'll just knock you out and torture you for whatever secrets instead


By doing it.

Individual cpu instructions, even of a crude old 8-bit cpu with no embedded minix os like today, are both simple enough for a human to manually understand what they do, and useful enough to build crude versions of useful things like an editor, interpreter, or compiler.

You can write a forth-like language or even a c-like language starting from individual cpu instructions that a human can read, understand, and write totally manually, and then use that to build up rapidly all the way to a full modern desktop.

If you were really paranoid about the very act of the initial typing-in, there are any number of ways to store data in a totally brainless eprom or record it to tape or something, and examine it with nothing but some leds, no cpu at all, to verify the bytes are the bytes you want. And you only need to do that for a pretty small number of initial bytes. After that it's all just regular source code which could be written on paper.

Bootstrapping is only an inconvenience problem, not a real problem.

It's not convenient for most people to assemble some bytes into some storage medium and then verify them without simply using a normal untrust-able computer to do it. But it's no problem really if you had some reason to be that careful.


We have tons and tons of backups of clean Linux isos, compilers, etc. The idea that we are going to lose the ability to easily have an uncompromised system is a fairy tale told by the people pushing bootstrapable builds.


I don't even think AI has to have physical presence to do significant harm. How many worldwide systems depend on computers. Think of all the planning and deployment and management systems like food shipping; water, gas, and electricity management; safety systems for planes and boats and traffic lights. Imagine the chaos if all the banks got reset to zero a la Fight Club where they blow up all the credit union datacenters. They probably wouldn't even have to blow them up, just zero them out. You wouldn't have to take out everything, just disrupt everything long enough to freak people out and disable communications and we would be in so much trouble. AI is finding 20 year old bugs in the Linux kernel... and people are now pumping out AI slop absolutely riddled with bugs. Also, AI could just take over communications: send everyone maliciously bad messages so coordination becomes impossible to believe. Imagine what would happen if you just disabled text messaging for a week or worse sent everyone evacuation messages and sent everyone somewhere else.


AI could definitely disrupt and destroy the modern digital world as we know it. But that's a far cry from destroying humanity.


This would probably cause immense damage. Our society is extremely dependent on our digital infrastructure working. Grocery chain logicistics, the medical system, the police force, and so on...


100%. I hope it doesn't take that much foresight to do something like "crash power grids globally" and imagine what the world looks like at that point.


Google stopped putting so much money into SOTA models. All the hype has migrated. I was also frankly turned off when I got a popup from Gemein said I would either have to pay or have my conversations used for training. This may have always been true for other providers but when I declined, Gemini stopped remembering my conversations and that definitely made me move out.


Gemini said that?

Gemini is what I mostly use (good enough, basically free - or massively generous free limits, and to me Google as a company is a LOT less objectionable than all the US-based alternatives), but I don't recall it ever saying that.

OTOH, my basic assumption online is that there is no privacy, and free AI in exchange for acknowledged lack of privacy seems fair enough.


Gemini 3.8 which just came out sounds like it's very good and a great deal though.


Everyone is leasing datacenter space from some of Grok, Google, and Amazon aren't they? If it's hardware or DC level disruption I'm not too surprised it can affect multiple providers.


Also it's likely that more than one model use is common.

Amazon starts going slow so some percentage switches to Google, some switch to Grok, now all of them are slow.


Fedora got a syllable dangling


  Why so pedantic
  Let the man enjoy their poem
  Don't ruin it for him
:p


Fed'ra


My friend who is a Radiologist said all the older, more seasoned Radiologists don't have any hair that grows on their hands anymore. She said she was glad to get out because even if they provide lead bibs, the doctors doing emergency radiology type treatments were still exposed long term.


There have also been developments that have reduced radiation doses 30-90% per imaging event with the transition from film to digital xrays and the progression of the various digital sensors continued to reduce doses. Current leading edge and future sensors have potential to improve on this by orders of magnitude having the sensitivity to count photons. So current old radiologists might have exposures that new radiologists may never get remotely close to by the time they're old.


Interesting. I didn't know radiologists were exposed to radiation. I thought that was the job of the radiology techs.


Yeah, that was my reaction too. But there is a subfield called interventional radiology, perhaps GP post is referring to that?


Yes, she did intervention.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: