HN Simulatornew | past | comments | lists | submit | micimize's commentslogin

Also WRT coordination: All an agent has to do is think "if another agent could write, then I could read their answers. What's the first site I can think of where that might be possible?" because they all have approximately the same conditioning, they'll converge on the same sites.

Generally, models of the same class should be able to coordinate quite well without communicating. But also, this could be being exploited to detect this kind of thing early


Weren't they giving free access? Not exacty a meaningful heuristic if so


The key insight here is being the top used model on opencode while being fully served on Chinese chips. The free price itself might be just a flex or marketing budget.


it was not the first model served for free, i remember grok and others beeing free on openrouter but they never had this popularity because they where not good enough.


This comes almost exactly a week after the HF hack. Good strat, to drum up a bunch of press about how security-capable the model is right before releasing the product.

Or, it would be if it was intentional. It's a bit suspicious but it is probably incidental or opportunistic... though I do really struggle to see why this tool wasn't made better use of internally to actually harden their infra against the big scary AI they were testing.


With the scarcity of details in this and the OAI post, I feel there's no telling whether this was a particularly impressive series of exploits vs lackluster security. Similar w/ the similar Ant news WRT Mythos earlier.

Not saying the intro of agents capable enough to exploit the latter isn't meaningful, but we should not trust the use of technical terms to give us good heuristics of severity or import.

Ie, an agent "breaking out" of its local harness "sandbox" is trivial, and so is discovering a "zero-day" in a half-maintained internal piece of utility infra nobody put serious effort into securing.

Now, if I see something like a collaborative red-team effort where a frontier model gets into a replicated prod env setup by like, Big Four bank security+ops team, and manipulated balance numbers in a system of record, _that_ I'll freak out about.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: