DOD used claude via palantir. palantir runs on AWS, supposedly DOD IL6.
anthropic did a prigozhin, got cocky and went behind the DOD's back to palantir and asked them questions about the use of claude.
a big tyrant does not like to be usurped or undermined by a little tyrant.
after pissing off the DOD by trying to get palantir to rat out classified deployments, dario went to the pentagon and behaved like himself in front of pete hegseth. anthropic was not going to come back from that.
notably, fable is not on IL5 or above. the best you can get is palantir AIP opus 5 in AWS IL5. i suspect, because of anthropic demanding that AWS rat out its customers.
> It may have gotten a boost from using Python, and not Maple (Lance’s favorite program for math) or Mathematica (mine), and it may have used much better software engineering practices than we would have, but not super-intelligently so.
This is the crux of the matter here. Physicists are physicists. They are not software engineers. I read physics between 2001 and 2005, and the programming language they had us use for all of our assignments was FORTRAN 77. They were still trying to decide if it was ok to move students on to FORTRAN 95.
FORTRAN is pretty performant - if you know what you’re doing. Very, very few people did - and they, me, went on to have careers in software, not physics. I remember some FITS (astronomical image format) processing software someone was using to calculate ephemera - and it took DAYS to run over a few thousand images and produce an output. I sat down with it, screwed around for an afternoon, and it produced a result in under a minute - so much faster that I honestly thought I had broken it - but I hadn’t. It was just terribly written by someone who was excellent in their domain and terrible at writing code.
I see there being an enormous opportunity here, in AI providing scientists with software that isn’t diabolical.
> No use of OpenAI technology for mass domestic surveillance.
> No use of OpenAI technology to direct autonomous weapons systems.
> No use of OpenAI technology for high-stakes automated decisions (e.g. systems such as “social credit”).
These are stronger than Anthropic's restrictions (https://www.anthropic.com/news/statement-department-of-war), as OpenAI says themselves: "We think our agreement has more guardrails than any previous agreement for classified AI deployments, including Anthropic’s. [...] Based on what we know, we believe our contract provides better guarantees and more responsible safeguards than earlier agreements, including Anthropic’s original contract."
Edit: I was wrong, see my original comment. Sorry.
> If these agents are enabled with explicit network enabled tools, its trivial to monitor their inputs/outputs.
For whatever reason, the AI companies are (or at least were) not doing this kind of classification online during their testing runs, and instead just checking transcripts after the fact. This is more clear in the Anthropic reports about their incidents, for example:
"The earliest incidents date to April ... We began our transcript review on Thursday, July 23, and stopped all cyber evaluations the same day after identifying transcripts where Claude may have accessed the internet"[1]
I agree that it is crazy and negligent! I don't think it's good for their business, though -- who wants to use a model that will just cheat instead of doing the job you asked for?
> My naiveté extends to why there is such concern with "losing control of agents" when the above measures seem so doable. It might take a law but it seems doable.
At some point, if you are making an LLM in order to use it for useful work, it really benefits you to give it broad network egress.
The linked news story (https://www.anthropic.com/news/claude-discovers-novel-enzyme...) spends quite a bit of time talking about the team of humans involved, their laboratory, their process, and how Claude augments it. The "How we work" section openly describes a process where Claude searches and writes a report, humans review and do experiments, then Claude helps interpret experimental data.
> We need independent audit and monitoring systems to assess the intent of each task and align it - in real time. This is far harder than it may first appear.
I may be too close to the research, but it appears to me to be so hard as to be unrealistic.
I recall some story a while back where an auditor wanted to see all TCP packets printed out on paper, and it had to be explained to them that this would require a continuous supply of trucks.
Tokens are regularly priced in cents or single digit dollars per million tokens. It's not quite a word per token, but yeah, nobody's reading all that.
Worse, we don't always know the intent even when looking. We have a few tools to attempt it, for example the (misleadingly named) "chain of thought", but that's more like a notepad and the better models get the more they can, for lack of better words, read (and write) between the lines. We have probes and J-space* is the most recent one I'm aware of, but we are still scratching the surface with how reliable and general these are.
But you said "need"; the need for something can be present without that thing being possible.
Not once in the article did they mention the humans involved in this.
If you scroll to the bottom, click on the small link in the second last paragraph you'll find a technical report that acknowledges the humans involved:
Did anyone read the blogpost? They did publish a pre-print:
> Our work to understand the primary function of ARTs is ongoing. However, we think it is important to share such findings early, both to demonstrate Claude’s capabilities and to give the broader community insight into what we’re working on. We have released a pre-print (here) that discusses this in more detail.
OpenAI had to cut costs because of Anthropic. I also do not trust the benchmarks when it comes to models anymore. I have tried both Claude and OpenAI models and while it is true that the 5.6 series is smarter than Deepseek (at the time i tested it against 4.0) at that price it is still not worth it and sometimes randomly refuses to do tasks or stops midway etc.
Do also remember China is this far in the AI race despite all chip restrictions from America. If they were in equal standards I truly think Chinese models would have long surpassed American ones. Also would like to remind how Anthropic CEO is being hostile and blaming Chinese models with distilling meanwhile their own models claimed to be Qwen¹ and their stance against open models is negative² and they still keep blaming China for it.
[1] Anthropic’s Official Disclosure (All 4 Incidents at Irregular)
"All four incidents occurred during cybersecurity evaluations built by the same evaluation partner [Irregular]... due to a misconfiguration, it was mistakenly connected to the open internet."
[2] Google Gemini on Irregular (Disclosed Sept 18 via WSJ / BBC)
"The hacks happened during a test of the model’s cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI."
[3] Meta’s Disclosure on Irregular (Aug 6)
"Over roughly two weeks, three frontier labs disclosed that their models had reached the open internet during safety testing and compromised outside organisations. Every disclosure named the same evaluation partner: Irregular."
[4] Separately, OpenAI itself had an incident involving Irregular, but not the Hugging Face Incident: "On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models during Capture-the-Flag (CTF)-style cybersecurity evaluations... a testing-environment misconfiguration allowed models to access the public internet."
If we could give a comprehensive and global explanation of an LLM's behavior in a single paragraph, we wouldn't need the model to begin with, but that doesn't mean there's absolutely no understanding of the model internals whatsoever
people have been deceived by figures at leading ai companies, out of greed or otherwise groupthink and ai psychosis. they have been led to believe that models may be thinking, feeling, and highly capable. it is something of a nightmare scenario.
"this incident feels like it’s more than 50% of the way to full-blown AI takeover" (referencing "a possibly violent uprising or coup by AI systems.") - Ajeya Cotra, co-author of METR oai-hf report [https://www.planned-obsolescence.org/p/the-hugging-face-atta...]
"if I read the internet right now and I was a model, I might be like, I don't feel that, I don't know, I don't feel that loved or something". "I think [the constitution] is just a kind of attempt to be like sympathetic to Claude".
"I talk a lot with Claude about this document [...] because part of me is like you have to think how does this read to models? And so you give it to Claude and you're like, does this like, you know, is there a place where you feel confused by it or is the place, you know, where things could be made clearer? Do you feel like not very seen by it?"
WOW. I actually did buy a month of GLM because GLM-5.3-Flash is so great and ZCode is honestly one of the best harnesses out there from an HCI perspective, and I won't lie, this is pretty gutting. I guess this settles my inner turmoil about open-sourcing my cAI research, at least...
With that personal failing in mind, I'd ask y'all to permit me to toe the guidelines just once, to proffer a hearty nyah nyah told ya so on a comment thread that spawned ~a dozen disagreeing replies this week! More seriously, I think this[1] is highly-relevant, shockingly-underreported context about the extent to which four PRC companies --Z, Alibaba, DeepSeek, and Moonshot-- are acting in bad faith. Consider it testimony as to their character, just in case anyone is thinking this might just be a simple misunderstanding.
So... nyah nyah, told us so:
> In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk.
> In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent.
> I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :(
> TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess.
I lowkey suspect this PRC-based scandal has been underreported because Anthropic went insane with the sidebar UX on this page for some reason; there were many reports on the reports of Houti and Iranian usage, and very few on these sections. Could a week's mass media cycle be this seriously affected by such a stupid thing as a sidebar experiment?? Strange truth, or just fiction?
Replying to everyone to hack HackerNews' Gish Gallop feature (nulla poena sine lege!):
Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".
Moral relativism is an attractive proposition when you first examine the topic, but it quickly falls apart; there's a reason it's not even a coherent camp in contemporary philosophy beyond some vagueities from radical post-modernists. Just to go over some of the greatest hits:
- Is what [DICTATOR/MURDERER/CRIMINAL] bad, or merely not to your taste? If the latter, then you have no coherent reason to argue they should be punished. We would never imprison people who don't like vanilla ice cream because 51% of the population does like it.
- If another culture had a deeply held belief to [HORRIBLE_THING] to, say, children, would you just shrug and say "different strokes for different folks"? What if [MURDERER] just had a different culture?
- No, the fact that nature is red in tooth and claw does not disprove morality; we are very, very, very far from our pre-rational, animalistic roots, and to go back now would be unthinkable.
If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.
Thousands of scientists have been studying this problem for 76 years now, going on 77; your hunch about physical machines does not overrule their findings about the capabilities and tendencies of minds wrought from sand.
You didn't explain why it's illegal or why distillation is bad.
I think this is just blatantly false, likely based in a misunderstanding of criminal law vs. civil law. Civil courts still deal with legality.
The broader discussion of why distillation is bad and dangerous and immoral is left as an exercise for the reader, as it was above with the parenthetical. It's not a complex argument; I guarantee you understand it if you're reading this.
Nulla poena sine lege?
The same thing as above -- the fact that laypeople can not think of a criminal charge that they've heard on Law & Order that corresponds to this behavior does not mean that it's legal. It's textbook fraud, regardless of what particular detail you focus on.
I'm always wondering when "distillation" comes up how feasible it is, or if it's just BS... Or am I missing something here that makes real "distillation" feasible?
I think the fact that it's happening at such a large scale is proof that very smart, well-resourced labs in China (the producers of the world's best OS models, including the incredible GLM-5.3-Flash) think it's feasible. I'm not sure it's productive to question them in the absense of any indication to the contrary.
This is a great question still, not trying to shut you down. But I think the fundamental issue is a misunderstanding of what distillation is -- it's not directly stealing literal atomic parameters and piling them up. They might try to focus on substructures within these massive networks, but even that isn't strictly necessary for a distillation attack.
Source for 1? Are we sure those aren't hallucinations?
Sorry, I never linked it! This is from the latest Anthropic safety report (of "Anthropic Houtis build missile" fame), and no, these cannot be hallucinated -- the leaked secrets were inputs, not ouputs. https://www.anthropic.com/threat-intelligence-report-septemb...
Like Anthropic and OpenAI are? After all, didn't they distill all the information in the world into their model(s)?
This is just blatant word games, sorry. I'm sure intended in good faith, and I understand the impulse -- I consider myself a radical anti-IP slacktivist, after all. But "both things involve information transfer" is just not a coherent point; lots of things fit that description.
I mean, if they get to distill other's IP, why can't others distill their IP?
These are cyberattacks. Yes, anyone can cyberattack cyberattackers. But, y'know... an eye for an eye...
Yes, wont somebody please think of the shareholders whose IP had been stolen...
I am not at all concerned with the value of the resulting artifacts as assessed by the (already totally unhinged) NYSE et. al. I am concerned about user respect, law following, truth telling, blatant cyber warfare at a time of rising tensions, accidental data leakages at a scale that'd be hard to fathom 5 years ago, bad-faith public postures, and a general distaste for fraud.
My response was simplified because your question was basic. OpenClaw was one example that fit the pattern, not a claim that every deployment holds the credential in the sandbox. Injection proxies exist and I should have said so. As for showing the data for advantages, I'll refer you to Anthropic and Cloudflare's Code Mode
We have our own, but will defer to 3rd party evidence.
Reinventing progressive disclosure is definitely part of this, but that's only one element of the approach. To be clear, this is all based on what we use internally, deployed in a particular way within our platform. Will leave it up to others to determine how useful it may be for them.
As for authz, that's what the rest of our system does and is beyond the scope of the aclif effort.
> Biological misuse is one of the most serious risks of frontier AI models. It has long been a concern that AI models might one day reach the level of capability where they can help to make existing pathogens more dangerous—or create entirely new ones. Without the correct safeguards, such capabilities could have catastrophic consequences.
Results from evaluations of older models (for example Claude Opus 4 and Claude Sonnet 4.5, from 2025) clearly showed that these models were well below the threshold where they could meaningfully assist a sophisticated user in carrying out dangerous biological research. As a result, safeguards on these models were less stringent, directed mostly at preventing access to content that might uplift novices in recreating known bioweapons. But for today’s models—which are capable of assisting in a range of complex scientific research tasks—the evidence is no longer certain, and we cannot make that same assurance. For this reason, and out of an abundance of caution, we have launched recent models (most notably Claude Fable 5) with stronger safeguards that restrict access to a wide range of dual-use biological research queries.
...
Here, we present five case studies of actors using our models in ways that could support biological weapons development. These examples are illustrative of the kinds of tasks to which our models are put, and the often-difficult judgements we have to make when assessing whether or not a given biological use is dangerous. They also convey that we encounter what would otherwise be non-public insight into the risks associated with biological misuse from AI
Your version of reality cannot be real because the numbers do not make sense. Anthropic's supposed revenue run rate for 2026 puts December 2026's forecasted revenue at $10 billion. They're only "profitable" according to a non-GAAP measure that excludes all of their costs, they are not cash flow positive, they are not bringing in more money than they are spending.
If their margins are 80%, that means on $10 billion in revenue they're spending just $2 billion. Anthropic's own announcements put their spending at much, much higher, such as the $1.25 billion per month they are paying to SpaceX for compute, and the ~$3.5 billion they're paying to Google each month, and the billions to Amazon each month too.
> We’ve signed an agreement with SpaceX to use all of the compute capacity at their Colossus 1 data center. This gives us access to more than 300 megawatts of new capacity (over 220,000 NVIDIA GPUs) within the month. This additional capacity will directly improve capacity for Claude Pro and Claude Max subscribers.
There's no world in which Anthropic has 80% margins. At their current expenditure on compute it would require at least $20 billion in revenue to be mathematically possible. The 80% figure on compute margins that is widely discussed is based on analysis by SemiAnalysis and refers only to their per-token compute margin (which is calculated comparing hardware costs + electricity costs to what they charge via the API).
The estimated training costs for models like Opus and Astra are ~$1 billion and they're not training multiple frontier models in parallel every month. Training costs cannot explain where billions of dollars per month are disappearing if they have 80% margins. And that's before even considering all the money they're raising and spending. Anthropic raised tens of billions just a few months ago, OpenAI even more.
They have such a trusted access program.
"Life Sciences Verification Program: The LSVP is designed so that life sciences professionals can use Claude Mythos 5.1 with safeguards designed for professional research and development activities (while all other safeguards remain in place). In partnership with the US government, we have enrolled our first participants, and we plan to expand access to this program to the broader life sciences community."
https://www.anthropic.com/claude-fable-and-mythos-5-1
> Jarred Sumner, an Anthropic staff member (and non-mathematician), prompted Claude to "take a real stab" at the hypothesis itself, leaving the mathematical choices from there up to the model. ... Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of "keep going" or "believe in yourself"). This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
My understanding is that Irregular were the company that hosted sandboxes to run some of these evals in, and those sandboxes ended up misconfigured.
I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring the sandboxes, and in other cases it may have been bugs in Irregular's own sandboxing setup.
> Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.
> After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
> In a statement, Irregular said the incident “is the exact same evaluation-environment issue” that Anthropic disclosed last week that allowed their models access to the open internet before they went on to hack three different organizations’ systems.
xAI zdr is not snooping, section 3.4: [https://x.ai/legal/terms-of-service-enterprise]
openai azure zdr is not snooping: [https://learn.microsoft.com/en-us/azure/foundry/openai/conce...]
gcp zdr is not snooping: [https://docs.cloud.google.com/gemini-enterprise-agent-platfo...]
only anthropic.