That's honestly nonsense. Kimi K3, GLM-5.3, and Qwen3.8-Max will happily hack anything you ask them to. These are near frontier models and are incredibly competent. Try them out yourself, you'll see.
I've been using them to test my infra/devices as well as reverse engineering.
Not giving legitimate people cyberdefense capabilities with the safety excuse is irresponsible
Define easy. I am running 40+ probes to crack the best compression algorithm, and I spent $3. If you surgically tackle the problems, you can do really complicated work for less money than by assuming a frontier top model will one-shot everything.
Well, at least I spent lots of dollars, and I had to use those models the same way I am using local and cheap models, with the same results.
Right now I'm using several different models for reverse engineering (DS V4.1 Flash, MiMo V2.6 Flash, GLM-5.3 Flash, etc.), and so far none of them are able to finish the task - it's been ~1.5h and several dozen million tokens used, but still struggling with the algorithm (FFT and some other stuff on an image manipulation library, which appears to be hard for them).
On the other hand, Grok and GPT finish these tasks in <5min with no issues, and significantly better output.
> That believes that the world can be simplified into dichotomies, or at least, simplified. Sometimes problems are complex, and the solutions to them necessarily so. For example, cancer
You just simplified most of the problems people work on down to cancer complexity. Ironic, isn't it?
That's also simply not the case, most people are building CRUD apps with some frontend code and some accessory stuff like build systems etc., which while complex, can still be expressed in very plain, easy to understand language for anyone who's a bit technical.
> Since LLMs can give different answers to the same question, each question was run five
times. That means, each LLM was tested 600 times, and in total over 10,000 questions and
answers were assessed.
> All models were given the same zero-shot format. They were not given worked examples,
previous conversations, hints or an opportunity to correct their answers. This is to make it as
similar as possible to a response to a question from consumers.
As for the evaluation itself:
> Responses were checked against this (using an LLM-as-a-judge), and was only given a pass if
every element was met; otherwise it was assessed as a fail. This all-pass approach was
intentionally strict, so that the score measures whether an answer is complete enough to
meet the expert legal standard, rather than how many individual points it gets right.
It's just AI slop and it should be taken with a mountain of salt.
> It's just AI slop and it should be taken with a mountain of salt.
Can't you see the irony. You are defeating the argument that LLMs are incorrect or weak with low effort with the term "AI slop" that itself is a narrative that AIs produce weak outputs with low effort.
> "Look how impressive and 'dangerous' our model is, look how naughty it was! We can't control it!"
While simultaneously announcing: "btw we're releasing the ever more powerful and more dangerous and more benchmaxing model next week, get our $200 sub asap!"
What if?
Easy to "what if" on an unfalsifiable claim.
reply