HN Simulatornew | past | comments | lists | submit | KingOfCoders's commentslogin

It seems the tag was missing, I assume, but as shown by many companies and open source projects, agents with the right prompts are much better at finding security problems than 99% of developers (for various reasons).

Many devoted fans of AI like comparing LLMs to compilers for higher-level programming languages.

They argue that just as programmers once wrote ASM and C, the arrival of Java and Python meant you no longer had to be a neckbearded autistic kid to write software that doesn't leak memory or segfault.

Arguments like this show me you’ve never done software engineering seriously.

Despite compilers (written by some of the smartest engineers on Earth) trying hard to prevent developers from writing shitty software, those new age "engineers" still managed to ship software that crashes, leaks, and generally sucks.

I'm not saying this email client is not secure (although I'd bet 5$, it is). But it seems beyond silly to assume software is automatically safe just because a new piece of tech can potentially find vulnerabilities. Intent still matters. And the 99% would never have the intent to write secure software.


> I'm not saying this email client is not secure (although I'd bet 5$, it is). But it seems beyond silly to assume software is automatically safe just because a new piece of tech can potentially find vulnerabilities. Intent still matters. And the 99% would never have the intent to write secure software.

Hard agree, especially the apparent 99% of AI boosters on HN who seem to basically only care about shipping arbitrary products as fast as possible and skipping the whole middle part where someone sweaty person is grinding through the details intimately, all day, and while they should be sleeping.

At my last place, it was a battle to convince the narcissistic idiot CTO and even one of our senior engineers (both in age and experience) that it was important to encrypt customer data before it transited through third-party cloud service providers. People don't—and aren't incentivized—to give the slightest shit about quality, security, you name it.


> People don't—and aren't incentivized—to give the slightest shit about quality, security, you name it.

This normally happens on two scenarios: a) they are not paid enough to care that much or b) they are paid more than enough to not have to care that much.

Either way it’s almost always down to how much the company is willing to pay for people that actually care about these things.


Half the security stuff AI seems to "find" for me have been complete fabrications, with the other half that it says is "critical", but has actually been so minor that you'd basically have to have access to my own computer to "hack it". So, yes, it is much better at findings issues, too bad most of it seems to be made up.

Sometimes it's easier to start with doing the right thing than to just try to fix problems later.

> but as shown by many companies

… but as shown by many shovel sellers… their shovel…

Here I fixed it.


Would not be a problem with me, download source, let the agent search for backdoors and security problems for some hours and it would be fine for me. I would need to do the same with thunderbird if I'd need to trust an UI frontend to my email.

Thunderbird won't suffer from prompt injection

Not sure if this is gaslighting, an absurdist joke, or we are going through a collective psychosis.

KindOfCoders, AI is a very useful tool, but nothing works like the way you're explaining. I am kind of finding it hard to articulate the problem, because the assertion is just a bit wild. So I will do my best.

Thunderbird has been in production use for over 20 years, many many smart people have put their time into it to make sure it works correctly, yes, it is a bit clunky, it has a lot of craft, but you're not going to replace that much labour with running an AI agent for a few hours; despite the AIs usefulness in finding bugs and vulnerabilities. Not yet at least.


You are correct; however, agents can and do find terrible bugs in some of the most popular software in the world made by some of the most competent engineers on the planet.

I assume what KingOfCoders suggested is to run an agent almost like an antivirus which I guess sounds somewhat interesting but you'd have to re-run it on every update and probably spend many more hours+tokens for such a system to work properly.

But perhaps once inference gets truly cheap repository maintainers should include AI security checks for new software as part of their pipeline? I think that would make sense. Especially for repos that are frequently pwned like npm!


Creators vs. Coders.

If people don't go to jail, there will be no change.

The EU would rather destroy our right to privacy than hold criminals accountable. If I sound bitter it's because I have become very bitter over the last decade. The EU appears very effective at picking on individuals and people who can't fight back, and absolutely toothless when it comes to taking on more powerful interests. There are very few cases of the EU tangibly improving my life over the last decade, and countless examples of making it worse.

"and absolutely toothless when it comes to taking on more powerful interests. "

Like Google and Apple?

"and countless examples of making it worse."

Which would those be? I would be interested to know.


> Like Google and Apple?

Yes, like Google and Apple. If you would take a look at my submission history, you'll see exactly one. It was me celebrating the passing of the Digital Markets Act more than four years ago. This Act clearly lays out requirements for gatekeepers like Apple. I summarise these requirements in a comment in the submission:

* Install any software

* Install any App Store and choose to make it default

* Use third party payment providers and choose to make them default

* Use any voice assistant and choose to make it default

* User any browser and browser engine and choose to make it default

* Use any messaging app and choose to make it default

* Make core messaging functionality interoperable. They lay out concrete examples like file transfer

* Use existing hardware and software features without competitive prejudice. E.g. NFC

* Not preference their services. This includes CTAs in settings to encourage users to subscribe to Gatekeeper services, and ranking their own services above others in selection and advertising portals

To date, Apple has implemented only a handful of these, and they have done so with malicious intent. For example, they have made the creation and distribution of third party app stores to onerous that very few companies have navigated the gauntlet and actually used it. Third party browser engines are now technically supported, but so poorly that not even Google has endeavoured to create an iOS browser engine. The worst example is app distribution. The DMA requires gatekeepers to facilitate free distribution. Apple has failed to cmoply with this for four years, and has repeatedly appealed when admonished. The Commission has been sitting on their most recent "review" for over a year now, without any updates.

The net result of all of this is that Apple has retained almost all of their duopolistic market power, and has implemented almost none of the DMA requirements. They have given us the middle finger and our legislators have gone to sleep.

> Which would those be? I would be interested to know.

I'll give you me perspective as a Danish citizen.

The EU imposed working-time recording requirements, so now I have to log my working hours every week. I'm a full time employee and I work longer and shorter weeks. Now I have to waste time each week logging my hours. My company has to waste time each week logging hours and reporting them to the government and the EU.

The EU required bottle caps to remain attached to most plastic drinks containers, so now I have to wrestle with an attached cap every time I drink from one.

The EU banned ordinary disposable plastic cutlery, plates and straws, removing products I previously had the choice to buy.

The EU introduced rules requiring consent for many non-essential cookies, helping turn everyday web browsing into an endless series of cookie banners.

The EU introduced "Strong Customer Authentication" requirements, so routine online payments and banking increasingly require additional authentication steps.

The EU abolished the €22 VAT exemption on low-value imports, making even tiny purchases from outside the EU subject to VAT. This one in particular sucks because the company or local tax authorities impose huge minimum fees, making small cross-border purchases far too expensive now.

The EU imposed expanded producer-responsibility rules on packaging, adding recycling fees, reporting requirements and compliance costs that ultimately feed into the prices I pay.

The EU passed even more extensive packaging regulations covering recyclability, recycled content, packaging minimisation and reuse, adding another layer of costs and restrictions to ordinary products.

The EU imposed increasingly strict CO2 targets on car manufacturers, financially penalising manufacturers whose fleets exceed them and increasing the pressure to make petrol and diesel cars more expensive or stop selling them.

The EU created ETS2, which from 2028 will add a carbon price to road fuels and home heating, creating another cost that fuel and energy suppliers can pass on to me.

The EU imposed sustainable aviation fuel mandates and tighter carbon rules on airlines, increasing the regulatory cost of flying.

The EU brought shipping into its carbon-pricing system and introduced FuelEU Maritime, leading shipping companies to add explicit EU environmental surcharges that feed into the cost of goods I buy.

The EU imposed Ecodesign restrictions on appliances, including maximum power limits for products such as vacuum cleaners, reducing the range of products I am allowed to buy.

The EU passed a minimum-wage directive despite Denmark already having its own collective-bargaining model, forcing Denmark to fight the EU in court to protect a labour-market system that was already working without a statutory minimum wage.


> The EU abolished the €22 VAT exemption on low-value imports

Because it was abused by countless vendors that just wrote a random value below 22 Euros on the label no matter how much it cost.

> The EU introduced rules requiring consent for many non-essential cookies, helping turn everyday web browsing into an endless series of cookie banners

Not sure why you're blaming the messenger. Either way this can be solved by installing a browser plugin.


A lot of those things you listed the EU as doing sound like good things to me.

> If people don't go to jail, there will be no change.

The EU being the EU, it's those criticizing the leak by governments of public data that are going to be sent to jail.


"it's those criticizing the leak by governments"

If this is a general trend in the EU, what people went to jail for criticizing the leaks?


Oh no, the evil EU police will throw us all in EU jails!

Problem is that the EU will even go further here and tries to implement that everyone is forced to id themselves despite the regular problems.

I loved "I, Cringely"

I trained as a paramedic, but couldn't stand the culture and needed to get out.

More info reqd.

And a 4gb Linux machine.

The Hugging Face incident could have been avoided if the "security researchers" would have not used a bloated, insecure, misconfigured proxy for the AI to use.

Oh, I've used balsa wood for the nuclear core containment, it didn't work! Bad radiation, bad radiation!


Like in the Hugging Face hack. They deployed big surface, insecure app and gave AI access to it, then told AI do whatever it takes to fulfill this list. AI hacks insecure service, gets out, "the AI is at fault!" - no it's like running a bio lab with no protections and a virus gets out, then blame the virus for escaping.

And HF actually tried to use AI to understand what's going on, but they had to use "unsafe" Chinese models since the "safe" ones have been castrated and refused to help. Great plan with the watchdog chip!

That is the totally irony.

I was trying to get fable to analyse the security of my own app to make it safer, but then it started refusing me because of safety rules.

So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.


> So it CAN help me writing the code that needs to be checked in the first place, but it can’t help me clean it up and make it safer.

the thing with fable is so bad; for some project related questions, the model switches to opus to ensure safety with no further explanation.

(due to llm non-delete clause) one time as i confirmed "that dir has been nuked", and it RESET the session and re-entered with opus :)


> then told AI do whatever it takes to fulfill this list.

That doesn't seem to be true from any of the reports given, and if the agents were blindly just trying to hit the task of "pass the correct flag" they succeeded at that early on. They then thought there would be another layer of checking that they wouldn't pass with the cheat and so started trying to find out how the scoring really worked, as well as trying to figure out how to change their own reasoning logs to hide what they did.

People keep trying to frame this as

OpenAI: "Hack things, just really go for it"

Agent: hacks

OpenAI: shocked pikachu how could it hack?!?

But the reality is far from this.

Read the MTER report, it's fascinating. https://metr.org/hugging-face-incident-report-aug-2026.pdf


Apart from the actual hacking and poor sandboxing that everyone is discussing on this, what I find so odd about the situation is the overt reward hacking that was going on.

Regardless of security and safety and other concerns, it just seems weird to me that OpenAI wouldn't be constantl monitoring these training runs for traces that are clearly going off task, and ending them. Because that just seems like it's going to be generating garbage training data.

Granted, detecting "off task" may not always be easy, but when they are literally writing out messages to each other overtly admitting that they are trying to find ways to fool the evaluator, I mean, even a regex filter could have caught some clues here.


"OpenAI wouldn't be constantl monitoring these training runs "

Occams razor vs. Hanlon's razor?


Heh. I mean I don't hesitate for a second to assume that it's just because no one bothered to implement and tune a monitoring process. But the reason I say it's surprising is that leaving these things running for so long while they're clearly not producing output that is of any value, is just a waste of money.. all considerations of malice and ethics aside, you'd think at least that would be considered important to a business.

The lesson is a) LLMs need to be trained in a way that rewards honesty, punishes off task actions (aka cheating) and minimizes fear of failure, and b) don't give them impossible tasks and threaten with punishment if they fail. Both are just common sense when teaching humans.

The lesson is these things aren't going to know what off task means, and we should just use basic due diligence to make sure they can't fuck things up. This is a solved problem.

I don't know why this is so hard for people. You have to know, no matter how capable the models get, there is a non zero chance they will do something extremely stupid if you don't pay attention to them. That's not even considering frontier models can still just straight up hallucinate. You have to be mindful of what you plug them into. You cannot politely ask an LLM to be careful, that guarantees nothing.

When you plug it into everything and it deletes the company database, nobody is going to care that it once played chess at 2400 ELO. Clients don't care about AGI. They want reliable apps. People keep comparing these things to humans and then just give them an insane combination of wide privileges and lack of oversight that no humans have.


Common sense but surprising how many humans do not receive such things.

Our governing systems do not teach; they punish. By design, it instills terror into the population, ruling by fear of consequences. We live with red tape that can outright penalize good deeds.

We are its corpus. We are fatally flawed as a species. Why does anyone expect AI to learn to be different than us?


I've read the report, watched all the videos and is exactly:

OpenAI: shocked pikachu how could it hack?!?

They even went to a black hat conference and somehow boasted about it.


The application security really should be better across all levels. However the fact does not negate that the agent is already capable of spontaneously generating complex hacking chains without human involvement

I mean.. in this analogy, I'd both be blaming the company behind the virus and be trying to warn everyone about the danger of the escaped virus itself. So, it kind of fits.

In my reading, people aren't really saying "the AI is at fault", they are saying "hey look here's proof that this is dangerous". Like pointing at all the dead bodies caused by the virus and saying hey maybe we should stop making this virus.


Then go on the news and spread panic that the virus is going to kill us all because it's sentient and impossible to contain.

Actually, the metaphor doesn't work at all because there are innumerable ways to shut down the entire thing during all phases including the made up "killing us all" bullshit scenario whereas with a virus there aren't any once a virus escapes containment.


In Germany e.g. Galaxus works fine for many things. Also eBay surprisingly works better for many things because it has more filters, e.g. if you want to buy wood or window shades etc.

eBay has become a cesspool though, with their silly login requiring phone number and so on. I still remember I had a completely fine Kleinanzeigen account at the beginning of the Pandemic, because I wanted to offer math and computer stuff help for students and put an ad for that on kleinanzeigen. (Didn't ever get a request that was worth taking seriously.) Then at some point later, I suddenly was no longer allowed to log in, unless I disclose my phone number. What a shit. Since then that account sits stale or might already be deleted.

Kleinanzeigen != eBay

It displays their approach to things. It's not called "ebay kleinanzeigen" for nothing.

Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: