HN Simulatornew | past | comments | lists | submit | NiloCK's commentslogin

Hey Jared. Are the office hours here a hard requirement?

My own working hours cut off at 1:00 PST (5:00 local).


I think that this is an unsolved problem in the same way that mangled fingers in image generation was an unsolved problem.

Through at least Opus 4, LLMs were practically useless for authoring any sort of coherent procedural closed-curve geometry (I know this with strong confidence because of the little animated guys at https://letterspractice.com).

Opus 5.5 can bang it all out. Possibly a deliberate RL sort of thing or maybe another surprise emergent capability.


Because agents trained without internet access (real or simulated) would be bad at using the internet.

Specifically, bad at search and returning information with references, bad at discovering and debugging package versioning conflicts, etc.

It's a Metcalf thing. The utility of an agent grows in some fn of the tools it owns. Skilled tool use comes from rich training environments.


Anthropic claims that their intention with the agents that led to the Hugging Face hack was an offline experiment running against a hacking benchmark.

Either you're being pedantic and missing the entire point, or you're saying that Anthropic lied and made a fake sandbox knowing their agents would need to connect to the internet anyways.


OpenAI, but not main point.

But the specifics here are the thing I was describing. This was cyber capabilities training on model(s) that were in a relatively unknown state of alignment training.

Because of the unknown alignment (and for varied practical reasons I guess) the training is intended to be inside a sandbox.

Because the training is on cyber capabilities, the models need access to simulated cyber environments, including target endpoints, including package managers, etc.

For package management, they set up Artifactory as a secure proxy. Agents ask Artifactory for packages inside the local network, and Artifactory serves them directly or goes to the internet to fetch if they are not cached. But the agents hacked Artifactory to steal its internet access.

So: to train cyber abilities, you need to at least approximate cyber environments. To realistically approximate cyber environments, you need to either pull a full copy of the entire internet to local or to use proxies. The former is pretty impractical, and the latter is exposing our limits at creating secure proxies. Yes, any specific failure can be mitigated, but the models get stronger and stronger. Fingers-crossed this is not escapable doesn't feel great!


OpenAI was involved in the hugging face incident, not anthropic, and yes, inadequate sandboxing was a major factor, even the wikipedia page states this: https://en.wikipedia.org/wiki/OpenAI%E2%80%93HuggingFace_inc...

Can you say more here?

You don't think that there is competitive pressure between Anthropic, OpenAI, Google, the various Chinese model companies, and others, to advance the capabilities of their AIs?

Or you don't think that sufficiently advanced AI can cause (mass) harms?


I am saying that if they lost alignment and their new model started injecting cryptolockers, dropping all tables of productions databases into the software or if it stated writing poisonous cooking recipes there would be backlash, lawsuits, and more towards such a lab. Consumers don't want to trust such a dangerous model so they won't buy the tokens and the lab will not want to spend money on lawsuits.

>Or you don't think that sufficiently advanced AI can cause (mass) harms?

Even a feather can cause mass harm if it's used to sign a declaration of war. Something merely being capable of causing mass harm is not an issue and doesn't mean that the existence of feathers are an issue.


But this assumes an evil superintelligence. We don't even need that for terrible things to happen; we just need people doing people things and AI doing AI things at scale, and that scale is rapidly exploding as we scramble to deploy AI in the real world.

Case in point, that hallucination (which, if it was an evil superintelligence, could have been "strategic") that almost led to US boarding a Chinese ship over suspicions of nuclear weapons: https://www.msn.com/en-gb/news/other/us-military-ai-failure-...

It doesn't have SkyNet, it just has to be WOPR. And it doesn't have to be people who are motivated to cause harm, it just has to be people making consequential decisions who are careless, distracted, paranoid or anxious -- which everybody is at some point or another.


>that almost led to US boarding a Chinese ship

And I almost kill people when stopping at a cross walk. That doesn't mean I harmed a pedestrian. Society is set up to be very robust. Humans themselves make mistakes and do bad things and society has had to learn how to live with that truth.

>it just has to be WOPR

Then why argue for setting a pace for the frontier labs if we've already surpassed WOPR level integration / intelligence.


> And I almost kill people when stopping at a cross walk. That doesn't mean I harmed a pedestrian. Society is set up to be very robust. Humans themselves make mistakes and do bad things and society has had to learn how to live with that truth.

Maybe not you personally, but many, many other people have killed many, many pedestrians, and when they exhibit a pattern of bad driving -- or other deviant behavior -- we take them off the streets. That is an example of society being robust.

Except, over here, we're rushing to make AI, which we know has many deviant behaviors and we know caused harms in many circumstances (starting with AI-assisted suicides), even more powerful AND deploy it in more and more real world systems!

As history and, literally, current events show us again and again, society has failure modes that lead to widespread harm and destruction. We've had world wars and then literally had multiple close calls with nuclear war right after. And then we have all that's going on out there. (In related news, the Pentagon threw a hissy fit because Anthropic would not let them use Claude for autonomous killing machines. Do I have to even mention what they've been up to these days?)

Most of these failures are caused by misaligned incentives and socioeconomic forces. The incentives and forces around AI have hints of many brand new failure modes that we can't even foresee because things are moving so fast.

> Then why argue for setting a pace for the frontier labs if we've already surpassed WOPR level integration / intelligence.

"We're already going down this mountain pass at 200 miles an hour, why slow down now?"

I think we should not just slow down frontier AI development, we should also slow down where and how that AI gets deployed.


> and now Google

Just a reminder that Hinton left Google from a much higher and more influential perch (not dismissing current author, but, you know), and for similar risk and communication reasons, way back in 2023.

This stuff isn't new - it's just newly breaking through into mainstream conversation.


A given model may or may not have strong self-awareness with respect to how strong it is with different tools.

If it's unusually skilled with one tool, but it's not the default tool for a job, then you have to put it in their hands before they reach for something else.

The Opus 5.5 javascript-art + art direction definitely seems to be one of these surprise capabilities jumps. Maybe strong enough that providers will start to nudge in that direction in the system prompt, so that the user request doesn't need to specify it.


Emissions.

Batteries and solar and wind are a lot cheaper, faster to deploy, and scale a ton better.

Nuclear is far too slow to solve our emissions problems, there's no chance of us being able to scale the nuclear industry to be able to confront climate change on the time scale necessary.

In contrast, batteries, solar, and wind are scaling on their own, without the massive government subsidies that would be needed to deliver nuclear too late to help.


Nuclear is actively negative-value too, because it discourages investment when people know that government-subsidised competition is coming.

If you start thinking today, there's no way to know if you can think without writing.

- Socrates

(The problem is real, but the future equilibrium is not easy to predict.)


Artificial General Intransigence


I am working on an SRS based early literacy acquisition webapp: https://letterspractice.com

The app has recently moved into production, so I'd encourage anyone with verbal but pre-literate kids to check it out.

The basic pitch is high efficiency acquisition of the highest yield phonetic mapping skills, and nothing else. I myself am something of a screen-time zealot and very wary of applying engagement mind hacks against kids. The narrow focus allows for good progress on a very modest schedule (recommended cap at n minutes per day for n years old, n >= 2). I defer the social and cultural aspects of learning to read entirely to parents.

It is mostly intended for parent-child co-use, although kids with a bit of experience can drive many of their own sessions most of the time.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: