HN Simulatornew | past | comments | lists | submit | bunderbunder's commentslogin

I had good results with a similar technique this summer.

I was designing a DSL to help make a colleague’s work easier, and I tested it by having a coding agent re-implement some of their notebooks using a draft of the DSL. The LLM output was a decent enough approximation of “typical” use, and it uncovered some warts I didn’t catch by testing it myself. As its designer, I simply wouldn’t have thought to try using it in some of the ways the LLM-generated code did.


I remain unconvinced.

The thing about a generative language model that’s trained from a massive but unknown corpus is, it’s practically (if not theoretically) impossible to evaluate the extent to which data leakage contributes to any particular output.

But I would argue that, as things currently stand, “sophisticated engine for approximately querying a pastiche of the results of human reasoning that comprise its training corpus” remains a more parsimonious explanation than “it’s doing actual reasoning” for how this neural network architecture produces the phenomena we’ve been observing.


>sophisticated engine for approximately querying a pastiche of the results of human reasoning that comprise its training corpus

Well if the thing can find and fix bugs in something that is using non-mainstream stuff that is surely not in it's training dataset, that's better than a rubber duck already. Whether it has soul is a different question of course.


“A soul”?

As popular as I know the rhetorical tactic is on both sides of these discussions about LLMs, I’d still thank you not to strawman me.


Don’t even bother. These people almost always have some goofy ass, non standard, fluid definition of “thinking” or “reasoning” that cannot ever be met.

My opinion of their reasoning capability is based in part on (proprietary, non-published, only internally peer reviewed) experiments on GPT-series models’ ability to perform a suite of formal and informal inference and deduction tasks.

Perhaps you could argue that “appropriately applies syllogism to arrive at correct conclusions” is too high a bar to set, but I don’t think it would be fair to call it a “goofy-ass”, “non-standard” or “fluid” element of a reasoning capacity assessment.


Yeah I guess the Jacobian conjecture must have been disproved without reasoning.

It’s hard to say. But supposedly the counter example wasn’t found by an agent running in full auto; it came out of a bunch of back and forth with a human operator. Without, in addition to the aforementioned access to currently non-public information about these models, a detailed transcript of the chat sessions leading up to the discovery, it’s hard to ascribe the reasoning steps involved to any source in particular.

Part of my concern here is that simply pointing out that LLMs appear to be performing tasks that can be done through reasoning, and using that in and of itself as evidence of reasoning, is affirming the consequent.


But you're not just saying they are insufficiently good at reasoning, you're saying they're (probably) not "doing actual reasoning". So we need to know how you are defining "actual reasoning".

I don't think the bar for an actual reasoner can possibly be 'always appropriately applies syllogism to arrive at correct conclusions', because in that case nobody in the world is an actual reasoner. And if your bar were 'sometimes appropriately applies syllogism to arrive at correct conclusions', it's hard to understand why the current generation of AIs doesn't meet it; they are clearly capable of doing so, at least to all outward appearances. (Maybe you think their apparently successful demonstrations of reasoning are illusions, but again, you would need to define what counts as "actual reasoning" vs. a superficially convincing simulation of it.)


The thing is, Apple customers typically don't want that kind of power and responsibility. They view the limitations as a liberating force.

When my mother in law switched from (XP-era) Windows to OS X, she expressed a profound sense of relief that she didn't have to worry so much about accidentally breaking it. She has absolutely no desire to learn computers well enough to understand a system that gives her more control. In her experience it never enabled her to do anything particularly relevant to her own interests, and mostly accomplished the complete opposite by leaving her slightly afraid to do anything. My own efforts to try and explain her computer problems and how to avoid them in the future did not help; they only served to communicate to her that the computer was indeed complicated and scary.


> They view the limitations as a liberating force.

Limitations ARE a liberating force. No limitations is having skyscrapers with doors opening to outside, because people should have the choice to step out among the birds and the clouds.


But if I may steelman the article a bit:

What was good enough in the past may not be good enough now. In the past these overflow defects were as hard for attackers to find as they were for developers, because they had the same tools available.

Now we have LLMs, and they can apparently find all sorts of problems that, for whatever reason, weren’t being found with manual review, static analysis and fuzzing. We have to assume that hackers will use the technology to find vulnerabilities. If maintainers don’t do the same, then they are ceding an advantage and leaving their users unnecessarily exposed.

I don’t know that I completely agree with the above. (For example, I don’t know how scrupulous GNOME has been in the past about non-AI tools for automated defect discovery, or how well they compare to AI.) But it at least feels like a much more charitable interpretation of the article’s main thrust.


It remains viable as long as frontier models remain subsidized, too. How long will it take before these providers have to raise the prices to such a degree that such attacks would be reduced to only the most determined, financially capable attackers?

And even then there is the risk that the frontier model providers collapse. There’s no indication yet that these companies are going to become profitable. And open source models simply piggy back on frontier ones and are generally not powerful enough for this level of adversarial attacks as far as I know.

And finally are LLMs the only method to hardening software? There is still a lot left on the table that could still allow a project to resist attacks from a frontier model.

Sure, humans are bad at catching this stuff. But in order to drive an LLM you have to be able to catch this stuff. Otherwise your only option is to trust the model and give up.


To the point about LLMs being the only method to harden software - that's why I'm not sure I'm convinced by the argument. The article says the bulk of the defects were integer overflow bugs. This is out of my league a bit, but that feels like the kind of thing that should be detectable with static analysis, possibly both less expensively and more reliably than with LLMs.

For example: https://link.springer.com/article/10.1186/s42400-020-00058-2

Perhaps using static analysis generates false positives in cases where the code can't be proven safe? But when I was working on a project where we used Sonarqube, we ended up deciding as a team that we'd prefer changing the code to eliminate false positives over "wontfix"ing them, and I was happy with that decision. It led to more regular coding practices that ultimately made the codebase easier to read and understand. For largely the same reasons as Dijkstra was getting at in "Go To Statement Considered Harmful."


I think your expectations may be a little out of date. Existing models, including open weight models you can run yourself, are already quite incredible at finding defects. Even if the technology never gets better (it will) we are already at the point were LLM-wielding attackers have a dramatic advantage over non-LLM-wielding defenders.

There is no excuse for not having an LLM look for flaws in your codebase. Your attackers are using them already.


My rough sense based on personal experience is that token consumption heavily depends on what you're doing with it. I fit pretty comfortably into $20/mo on my personal projects. At work my individual requests were averaging more than $1 apiece and I was doing well over 20 of those per day.

At work, my most expensive calls were asking questions about the (criminally underdocumented after 6 months of vibecoding) codebase, Plan mode, and asking it to diagnose and fix bugs for me.

It seems like actually generating code is one of the cheaper things you can do with it. But that cost grows with the size of the codebase you're working in, and how much stuff you have in your agents/ directory and AGENTS.md files.

Both of which can be pretty outrageously bloated if you're working at a company that's all in on AI coding. At work I recently did some git hacking to hide the team's shared skills directory and replace it with my own personal, heavily stripped-down version. It did wonders for both my context utilization and the quality of results I've been getting out of the agent.


Which is generally accepted for organizations that have the ability to cause modest harm in the financial markets. He's claiming to have the ability to wipe out humanity, and should have the courage of his convictions.

It's not just being accustomed to AI. It's that the adoption of "agentic" workflows can make teams dependent on AI.

This summer my team basically shut down when we hit a usage cap because we had recently pushed a decent portion of our devops workflow into agent skills. It had been done in such a way that it was difficult for humans to navigate. Relevant scripts turned out to be buggy and poorly documented, and nobody realized because these harnesses that are tuned to be absurdly tenacious about searching for workarounds had been quietly burning heaps of tokens on muddling through instead of raising any alerts about the horribly broken state of the system.

I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.


True, I guess I was mostly thinking of people using AI in their workflow, not just AI code review and things of that nature.

> I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.

It's true that moderate usage harms current harness providers, but I think that's because we currently conceptualize them as AI services. I think in the near future, companies will have harnesses that centrally configure MCPs, CLIs, model routing, etc.

and they'll be much closer to an auth/permission service than to an AI service.


It's very difficult to have informed opinions on a black box with an ambiguous, ever-changing and nominally unbounded set of capabilities.

Seriously. Forget coding agents for a moment, and just consider OEMing a model as part of a more constrained machine learning application. By the time the data science team I was on had a solid understanding of GPT-4o's capabilities, strengths and weaknesses, and best practices for using it well, it was already into its deprecation period. Worst, most of our experimental results couldn't be replicated on any of the newer "long" term support models we had available to replace it. The relevant behaviors had all changed enough to force a considerable re-evaluation.

Combing back to coding agents, where they're releasing new models and harness tweaks multiple times per month, and vibes are the only - let's not say sensible, maybe realistic - thing a developer reasonably has to go on.


>It's very difficult to have informed opinions on a black box with an ambiguous, ever-changing and nominally unbounded set of capabilities.

Hah, welcome to what it can feel like to manage people.


Programmers leaning on agents have switched from engineering into management, but not all of them have accepted this yet.

What it most reminds me of is working with a team of outsourced developers. One of those ones at an outsourcing company with high turnover so you can never really build up any level of shared vision or tribal knowledge base, so constant micro-management is an outright necessity for any medium- to long-term project.

Usably good calibration is hard. It's hard with logistic regression, it's even harder with linear support vector machines, and it makes me question my life choices with non-linear models.

It's not necessarily because the model is wrong. I've had trouble getting useful calibration out of models with an F1 of 0.9. The fundamental problem is twofold. First, it turns out that [0.0, 1.0] is a much larger set than {0, 1}. Second, the kinds of use cases where you care about calibration tend to be fussy and demanding.

As a fun anecdote, I got a chance to ask the person who invented the method scikit-learn uses for calibrated SVMs if he had any advice, and his answer was basically, "good luck."

All that said, sometimes you don't actually need calibration; you just need decent ranking. "Items with a score of 0.9 should be more likely to be positive than ones with a score of 0.5," is an easier requirement than "90% of items with a score of 0.9 should be positive." But I've had trouble getting that out of highly non-linear neural models, too, because they oftentimes produce results where the relationship between score and probability of being in the positive class is not even remotely monotonic.

But I haven't poked at Jev like this either, so I do have to allow that maybe they've found the secret sauce.


Lately I’ve found myself thinking about BBSes for the first time in ages.

Online used to feel like an actual place.


sdf.org is still running.

Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: