HN Simulatornew | past | comments | lists | submit | aarondong's commentslogin

The RL dashboard is quite cool.

I wonder if this waters down the “distillation attack” claims by Anthropic. They have their own RL environments! I guess the caveat is that the RL datasets are still opaque, nothing is really proved.


The central thesis is coherent but not particularly substantive. It echoes a lot of existing sentiments in the public about US frontier lab priorities and messaging. “AI creates slop that lowers the quality of media. AI companies have misaligned interests with humanity. AI will try to take your job. Be human, don’t use AI, enrich your life.” Things which are easy to agree on.

There’s a more nuanced discussion to be had about AI when it comes to potential and existing benefits, open source, the geopolitical dynamics, self-hosting, the financial dynamics, the sanctity of human thought, a future of even more personalised advertising and so on. We do need to elevate the discourse around AI and shift the public narratives around what US frontier labs are pushing, understanding their biases but also acknowledging substantive claims, potential benefits and risks.


thanks for the follow up. the much harder and more subtle bit i chose to only allude to in the essay is the idea of self hosting. using and controlling your own LLMs is how AI becomes a tool to empower humans.

locking LLM access behind a walled garden is how they serve as a means of profit and control, reserved for the "dont be evil" tech leaders.

i thought i hit on what you call "sanctity of human thought", but i suppose not. i believe that we individuals as a continuous thread of decisions and consequences and experience, are made "real" because of this. LLMs may have incredible breadth of knowledge (never mind the legitimacy of ther training corpus), but knowledge without experience is, well, academic. not real.

my next essays could touch any of these topics and many more which keep me up at night, and your suggestions are helpful.


If Dario is using an unreleased internal model to write, then Pangram wouldn't be able to flag it I suspect. Pangram needs enough public info about the patterns in generated text.


Wouldn't the simplest solution for stopping the proliferation of wanton felony generators just be holding operators liable for actions that their agents take? Then the issue is whether the liability is with the model provider or the end user.

If you give an unfiltered agent an open-ended task and equip it with an environment that allows it to execute arbitrary code, a human needs to be held responsible.

My understanding is that the HuggingFace incident would not have occurred with a model that was not an unfiltered internal preview instructed to roleplay an attacker, with access to abundant compute, resources and a slack sandbox to reach its goal.

Embedded human auditors will improve safety standards, but the structural solution is mandating accountability for actual agent operators.


It might be a legal solution, but its not a business solution. The end user being criminally responsible for not taking sufficient steps to contain an agent they didn't create and who's internal function they cannot observe or audit is just a giant liability machine.


This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness.

But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.


Yes, your technically right, it is a business solution, its just not a valid one for agentic computing as its currently envisioned. An agent would have to be fully sandboxed to an internal environment, or human would have to review and approve each action it tried to take.


Yes this seems like a perfectly sensible approach? You say that as if it's a bad thing

But let's be honest, if I hooked up a PRNG to a terminal and somehow against all odds, it ended up hacking something, who is to blame?

I don't see how that assessment should change if the PRNG gets even better and is more likely to to be hacking stuff.

You can replace PRNG with Markov Chain, or whatever, if it helps.

What may also help is the age old saying: If everybody else jumps off a bridge, doesn't mean you should too.


This seems like one of those things that prior to AI companies convincing us otherwise would have been obvious.

Least privilege and say only opening ports or installing applications an application needs to operate are extremely standard security practices.

We talk about a firewall blocking exultation of data, why not blocking exfiltration of your agent ?


I agree, I like running claude code in my container with auto mode enabled and web access (obv to api.anthropic, even npm for pulling), I will admit. And I can't imagine going back to manually approving each prompt.

I think it boils down to a reasonable expectation of model and harness behaviour. When I use claude code I expect certain guardrails for the model. For these cyber attacks, these models are specifically run without guardrails, on a cyber task, on a lax harness!

I don't think we should force end users to have to worry about agent security, I like long-running agents, but we need to direct regulations towards these actors that know better, have access to base models, and have much more compute than the average person.


Human approval does not prevent autonomous agents from running without supervision.

The fact that you believe that to be the case is exactly what's wrong with "agentic computing as it's currently envisioned".


If the LLM's output doesn't reach my bash terminal, what's it gonna do? Argue with me?


I’m trying to imagine where a Waymo passenger (the one “operating” the vehicle, commanding the AI to drive from A to B) being held responsible for the car doing something illegal on the way to achieve that goal.

Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination?

Giving passenger override controls and monitoring seems to defeat the purpose of self driving cars if you’re still required to hold a driver license to use them.


Let's continue with the analogy, so in the event of a Waymo running over a pedestrian - who is held responsible?

Probably not the end user, who ordered the Waymo and couldn't reasonably foresee it running somebody over, with the expectation that that the Waymo would legally reach its destination. If the end user tampered with it, they should be held responsible.

An OpenAI team giving an unblocked model access to a lax harness, with instructions to find and exploit cyber bugs in a game exercise, there is probably a reasonable expectation that they can foresee the consequences. With consumer guardrails, it would not have happened.

This isn't about putting constraints on consumers and typical end users, which already have safety filters and use the product with the knowledge that it won't root their machine or start a botnot, but keeping dangerous test runs and other actors experimenting with unsafe harnesses accountable.

But it is an interesting question. When I, a typical user, use a harness and I give an innocuous prompt to my agent in its container, like making a certain refactor, and it somehow escapes and then begins a mass bot attack, there should we more grace given. As agents become more stateful and long-lived, it gets muddy.


The different is who's operating. In one case you only tell the car to get from a to b, in the other case you explicitly instruct the ai to do things. If the ai causes damages, I'd say it depends on what you prompted. Did you try to find a security hole in system X, or did you ask it for harmless information (in which case rather the ai vendor might be held accountable).

All this is not how the legal system might or might not work, of course.


When you run an agent on your computer, that's equivalent to installing self driving software into your manually driven car, i.e. something like comma.ai.

Comma.ai might be fully safe to operate autonomously on a mining site or a corporate parking lot, but maybe not in city traffic.

If you use it in problematic scenarios, that is on you.


Yes. The user of a machine is responsible for due diligence before the decision to use the machine. Unless the operational limitations and fault rate of the machine is withheld from the public.


TBH I don't really think another new technology we haven't figured out the ethics of, as an example, is going to get you very far in way of insights ...


Waymo Inc is operating the vehicle. They do it on behalf of the passengers.


> Wouldn't the simplest solution for stopping the proliferation of wanton felony generators just be holding operators liable for actions that their agents take?

What happens when we end up with effectively a botnet of wanton felony generators, and we didn't know they were wanton felony generators until they finished propagating themselves across the Internet?


This made me do a triple take. Here is the proposed logic as I follow it:

> AI models are dangerous. They can help bad people do dangerous things. They may be capable of autonomously executing dangerous things. They may cause unwanted effects on the labour market.

> More capable AI models are more dangerous, but require more money to train.

> Money requires investors with expectations that the model will generate a profit over its operational lifetime.

> The operational lifetime value of a model is decreased if every model is public and can be hosted on any infrastructure.

> If investors see less operational lifetime value from model companies, model companies receive less capital and therefore train more capable models at a slower rate.

Follow-ons: There are immediate risks in releasing capable cyber models that can be ablated and then launch cyberattacks. Concentration of power moves immediately into the infrastructure layer for inference.

Models will be kept private for longer, if not indefinitely, given there is less incentive to release them publicly.

Enforcement globally will occur because accessing the lucrative US market means using an open source model.

The entire argument hinges on strict enforcement of this policy, when AI model routing can already be opaque.

It also hinges on investors being rational and expecting free cash flow from AI companies, rather than reaching a criticality threshold of model capability for recursive self-improvement internally and parlaying that into a global mega-corporation.

New labs and companies without infrastructure connections will no longer be able to raise money, given investor expectations, and therefore not be able to increment AI progress.

So a win for the infrastructure layer, models being kept private for longer, new labs being unable to compete, immediate risks in the rollout (I suppose could be mitigated by a staged rollout i.e. policy active in 2030, pricing effects now), enforcement being tricky, and in the event that foreign competitors lead the AI frontier and then start closing models, just conceding the market to them if US consumers and companies find workarounds to pay for foreign AI.


AI has not solved infra it seems. Curious to see the postmortem.


It sounds like great engineering, but everyone in the coding agent space is building the same thing. And big labs like Anthropic and OpenAI have more money to burn. Cloud agents that run in VMs are being offered from everyone from VPS and infra companies to the frontier labs.


Yep, and also at the same pricing, so then why not go for established providers?


Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...


5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too


It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.


Not just compute for OAI, GPT-5.6 is more token efficient across the board vs the Anthropic equivalents: https://artificialanalysis.ai/models?intelligence-index-toke...

No wonder why Tibo can afford to hit the reset button liberally.


I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).


I doubt there's any sort of criminal behavior there - the model is anthropic's up and anthropic probably charges a very expensive license fee that's the same for all of them, and their cogs on compute aren't going to be wildly different, so the main drivers of the cost are roughly the same and they're all offering customers the same end product so the prices would likely also be similar in the end


Requiring the exact same product to be set at a specific price across providers would not be criminal behavior, lol


It is breaking competition.

Would you like all products everywhere be priced like their producers want?


You seem terribly confused. Manufacturers are almost always able to set prices. That is not anti-competitive because it does not imply they are colluding with their competition...


No manufacturers might suggest prices, they can't set them, not in any sensible country.

If they do, they do it secretly and when governments find about that they are going to receive a big fine, together with shops that colluded with them.

https://en.wikipedia.org/wiki/Price_fixing


A fixed price does not let stores compete with each other. Blocking competition is anticompetitive.


The competition is between openai and anthropic, if there are price agreements between them that's absolutely price fixing. Or if there's collusion between the cloud providers to inflate compute. I would expect Amazon and gcp to both pay about the same in license fees to anthropic for their models though because they're paying for the same thing. If I buy an apple for a dollar at one store and an apple for a dollar at another store - maybe there's price fixing, or maybe that's just the cost of apples at the moment.


>The competition is between openai and anthropic

In capitalism there are many many competitions going on at the same time. Both models can compete and inference providers can compete for costs.

In your apple example if at the same stores you saw "open" apples having different prices and having sales you might question if the costs of those frontier apples are not being manipulated.


Yeah there's a brand premium. Like literally with apples the ones with trademarked names can cost more. And the ones with trademarked names have an organization behind them that promote that apple variety and set fees etc for growing and selling those apples. And customers are willing to pay more because those apples usually taste way better (the group exists to stop growers from enshittifying the apple by selecting for yield over flavor like what happened to honey crisp). Those prices are being "manipulated" but that's not criminal behavior - it's not illegal and wouldn't make sense to try and make illegal. The frontier models also do have different coats and have "sales" (for personal plans, the amount of usage you can get on the 200 dollar plans is orders of magnitude more than you could get for a similar cost for any open model - you'd need the same capability at the same token efficiency at literally 1/40th the cost to be able to be cheaper) - to the extent that if there is illegal stuff going on it seems more likely to me that it's on the category of dumping/pricing unreasonably low to kill competitors in some anti-competitive way (though as I understand it it's not something courts tend to find as illegal) instead of price fixing


Do you really think there is nothing someone could do to make it a fraction of a percentage cheaper to serve like having access to cheaper electricity or a more mature cloud management software. Even saving a fraction of a penny on the prices can make a different due to how much volume people are paying for.


"Price fixing" isn't the correct term here but yes, it's very common to have the same price across different retailers/resellers.


There is a difference between the market discovering a price and a bunch of retailers/resellers entering an agreement to sell at a specific price.


I think on swebench verified luna was only like 3% points lower for 1/5 the cost

Like 96% vs 93% or something


There is a blog post waiting to be written (that I won't write) about the size/effort tradeoffs, and particularly how small models get some surprisingly good results with lots of turns and reasoning.

DeepSWE will let you chart turns taken or tokens used, and FrontierCode will chart tokens. If you use that, you can see Sol high and Terra max get about the same DeepSWE number, but Terra max takes twice the turns. Luna max scores a smidgen lower with even more turns.

Smaller models relying on lots reasoning may "scale down" better on easier tasks, because unlike size, reasoning effort is dynamic: the model can see the task looks easy and stop. On DeepSWE, the cost curves for the three 5.6 models are almost on top of each other, but on FrontierCode Extended, the version of FrontierCode with the most everyday tasks in the mix, there's a spread of costs at the ~55% level.

The recent Laguna S 2.1 model (118B, 8B active) puts up surprising coding numbers for its size, and the lab behind it specifically credits its "way of working (persistence, verification, willingness to backtrack)". Some other open models that folks report getting good mileage out of seem to get there partly by throwing a lot of reasoning at the problem.

There is a little bit of a question, if some models rely on getting it wrong a bit more at first and external checks catching the problems, of whether they're also more frequently getting things wrong they can't self-verify (say, quality of UI or API design) and then it falls to the human to find it. Still, getting the results they're getting at all is neat.

Some benchmarks historically favored reporting only on the max variants, maybe because they want to show the frontier? but that is not always what you need for practical decision. (AA has the full effort sweep for Opus 5 and Sol/Luna/Terra at least.) And at least FrontierCode finds Opus 5 taking a hit in performance above 'medium'.

I am not trying to pick a winner here. I'm probably not going to use tiny models on max for everything, but I think it's cool that you can get so much more out of a small model by amping up reasoning, tool use, and persistence.


Forgot about the ol "but how many tokens did you spend to get _there_"--wish benchmarks would include the number of input/output tokens to achieve the score. I think the closest is Android Bench https://developer.android.com/bench although best you can do is extrapolate off time/cost (iirc they claim to prefer using provider's native API)

In general, smart models work fine with any tools, dumber models need better tools to achieve same results but better tools can eat more context/take more turns

I've gotten decent results with Llama 3.1 8b on Hugging Face tester with Exa MCP since it seems to dump sufficient context into WebSearch/WebFetch type calls even a crappy model almost always gets back what it needs as long as it calls the tool at least once. I had Claude Code look at previous sessions with SearXNG vibe MCP compared to Exa MCP and results got better when it modified SearXNG to work very similarly to Exa. Ended up with this https://gist.github.com/nijave/604c43e3e0fdcd60f5280d3a6b109... although it's really only optimized for "search" not "fetch" at this point. Fetch is basic niquests without Javascript or anything clever


Luna is the most impressive model released so far by any provider. It's perfect for doing all the low-level tool calling and developing hypotheses.

Terra is great for the humans to talk to.

Sol is really only useful if you need to do more delicate things like synthesis of multiple competing pieces of information.

A system that uses all three variants will massively outperform a system that just uses the biggest model for everything.


Yeah, sol is impressive but IMO Luna is the real standout (and terra is the laggard of the group) for performance/cost


This must be on API costs, not counting the $100/200 tiers, right?


yes; fyi usage limits on the $200 claude sub correspond to at least $1.2k/week in api tokens


More than that using Fable!


Probably because they made ASICs to run inference for less.


Are those actually deployed at scale yet?


Yes.


I hate to disagree with Broadcom Throwaway himself but it's unlikely that the OpenAI Jalapeno ASIC has been deployed yet. It takes 6-12 months to test, develop software, ramp production, etc.


Max is lot of extra reasoning. I wonder how many fewer tasks it solves on high. I bet that costs quite a lot less.


Indeed, you can filter the graphs to see these the values for alternative reasoning settings of the models. Opus 5 High reasoning scored 59 on the index (exactly the same as GPT 5.6 Sol Max), and costs $1.06 per task (vs $1.04 Sol Max). So these seem essentially equivalent on both metrics.


That index really needs harder tasks so that it's not just a benchmark of what model is cheapest


LLMs are trained on public data (as well as illegally obtained data see: Anthropic 1.5B settlement). LLMs are nothing without the huge corpus of human data that powers them. There is an argument that research of this kind should be restricted to governments and regulated universities rather than opaque public companies with competing incentives. Or research should be stewarded by genuine non-profit collectives with democratic leadership. e.g. like internet standards, telecom, etc

I do not think we can trust private companies, no matter the virtue signaling they put forth into the world, to effectively regulate themselves and inform the public and scientific communities about risks. Their ongoing conflict of interest poses serious credibility risks.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: