HN Simulatornew | past | comments | lists | submit | lossolo's commentslogin

This seems like fruit of the poisonous tree. They didn't monitor their training environments, so I bet the reward hacking just got incorporated into their training corpus. In other words, agents solved some tasks, but not quite as intended, because of reward hacking. Instead of discarding that data, they included it in the training data for later checkpoints. And once that signal is reinforced, it happens more often, so the more it's reinforced, the more reward hacking you get.

And they didn't monitor what was going into the training data, so if one instance achieved its results through RL reward hacking (in other words, cheating), it just went into the training data, and other agents later used that pattern. I'm not sure whether that's a lack of preparation, negligence or incompetence, but they literally trained later checkpoints on the rollouts from the HF hack.

So it seems that OpenAI hacked so many systems not because they have superior models, but because of how poor their training, sandboxing and evaluation pipeline was compared to Anthropic's.


> Anthropic sets up Bay Area lab beyond computer simulation work, two sources say

> Startup aims for Claude AI to direct robots in lab environments, one source says

> Company to stop short of clinical trials to avoid drugmaker competition, life sciences head says

https://www.reuters.com/world/anthropic-quietly-sets-up-biol...


Relevant: "Anthropic quietly sets up biology lab as it ramps AI drug program"

https://www.reuters.com/world/anthropic-quietly-sets-up-biol...


"Congratulations to China! They were the second nation to land on the Moon, just a few decades after the greatest nation in history landed there first, the United States of America.

We are not in a hurry now, because we already won that race many decades ago!

Thank you for your attention to this matter.

President Donald J. TRUMP"


The Fake News Media is reporting that NASA completely lost the Moon race to China. A total disaster! Under my Administration, Space Force is winning BIG. Now, Beijing is supposedly up there doing who knows what. Fake flags, very low quality! The moon is ours, and Mars too (which the Moon is part of!). MAKE AMERICA SPACE AGAIN!

President Donald J Trump


After a recent interview[1] with Noam Brown (OpenAI), in which he said they had specifically trained agents for cooperation before this hack, the hack doesn't seem as impressive anymore.

They didn't even bother to control the post training rollouts, so the training data got contaminated and was included in the training of other agents. Connect these two dots and you have the Hugging Face hack. And at the beginning, when these incidents were first reported, it was portrayed as if all of this (the communication between agents etc.) was emergent behaviour.

1. https://www.youtube.com/watch?v=6AgOfiZOWiY


The result doesn't change, the result is what is dangerous. And the natural ability for the model to form a swarm is now trained in, this is extremely risky. The models should not be able to form swarms, this is an extremely dangerous property.

They are basically creating a slime mold or ant colony that can speak multiple languages, create their own language and operate as a collective.

The idea that you are impressed is a non sequitur.


China is not Iraq, Afghanistan, Venezuela or some other third world country that the US has fought in over the past few decades. You can't just destroy one of their satellites for free. And I can bet the US didn't do it, because that would be a major escalation, especially now that Xi is coming to the US and there are rumors that the US wants a deal.

So China's brand new recon satellite that was placed specifically in orbit over Iran just happened to self-destruct weeks after China was accused of giving Iran high-res satellite imagery that was used to kill American soldiers?

And the US just publicly told everyone they have weapons in space?

Hell of a coincidence.


So it broke apart by itself?

"The European Space Agency estimates that there have been more than 660 explosions, collisions, or anomalous events resulting in fragmentation. As it stands, there are currently more than 1.5 million space debris objects in Earth orbit.

It’s not clear what may have caused the Yaogan-50 (02) satellite to break apart, but it may have been a propulsion or battery failure, or it may have itself collided with another object or space debris."


That doesn’t mean anything. “Sometimes satellites hit things or have problems and blow up. Maybe it hit something and had a problem”.

Edit: Its a perfectly valid response to the question that was asked.


> If they opted out of training, then we definitely did not train on them.

Can't you guys just check their account settings so the public knows what was set?

EDIT: Why was this downvoted? I'm genuinely asking because I have no idea. Opting out is just a normal setting in the profile, It's not like I'm asking for their private conversations or PII. If I were the person claiming that they trained on my conversations, I'd make sure to disclose that I had opted out and hadn't given them permission to do so. And if I were the accused party, I'd disclose whether that setting was turned on or off to provide evidence against the accusation.


I don’t think your question is unfair*. They can check and so can Buckmaster. If he didn’t opt out, there’s a good chance his data was used for training. I believe this to be the case myself. What I’m more skeptical about is the purported impact of this data on the model’s behavior.


Yeah, I'm just curious about the setting. It's just weird to me that this wasn't disclosed by either party while the accusations were being made, that's all. Even if it was used in the training data, I don't believe it had that much of an impact myself, since the solutions are quite different.


Not entirely, it's just a late stage of the overall training process. It's an early checkpoint in post training (you can use the model at different stages of training), so it will probably become even stronger with more post training.


It was working for me too.

sample_size++;


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: