HN Simulatornew | past | comments | lists | submit | includenotfound's commentslogin

What if not building superintelligence leads to doom?

What if?

Easy to "what if" on an unfalsifiable claim.


It's clearly a joke. The rest is readable and well written, and free of LLMisms

Attention is expensive these days.

That's what I use on my M1 Macbook Air too, works fine.

Small local models are garbage, whether uncensored or not. Have you tried a proper uncensored LLM?

My hardware setup isn't good enough for anything but very small models, I wish there were abilterated/uncensored models on Openrouter

I think venice.ai hosts abliterated models

Just give it a web search / web fetch tool

That's honestly nonsense. Kimi K3, GLM-5.3, and Qwen3.8-Max will happily hack anything you ask them to. These are near frontier models and are incredibly competent. Try them out yourself, you'll see.

I've been using them to test my infra/devices as well as reverse engineering.

Not giving legitimate people cyberdefense capabilities with the safety excuse is irresponsible


Sure, if you're doing easy work. But Grok is a lot more intelligent and can handle harder tasks.

Define easy. I am running 40+ probes to crack the best compression algorithm, and I spent $3. If you surgically tackle the problems, you can do really complicated work for less money than by assuming a frontier top model will one-shot everything.

Well, at least I spent lots of dollars, and I had to use those models the same way I am using local and cheap models, with the same results.


Right now I'm using several different models for reverse engineering (DS V4.1 Flash, MiMo V2.6 Flash, GLM-5.3 Flash, etc.), and so far none of them are able to finish the task - it's been ~1.5h and several dozen million tokens used, but still struggling with the algorithm (FFT and some other stuff on an image manipulation library, which appears to be hard for them).

On the other hand, Grok and GPT finish these tasks in <5min with no issues, and significantly better output.


It seems that these things goes on depending on each type of project. I've been reading reports and Grok 4.7 sucks at 3D, for example, really bad.

The problem with reports is people just go "do X", and vibe off. Good to test models yourself for your own use case

> That believes that the world can be simplified into dichotomies, or at least, simplified. Sometimes problems are complex, and the solutions to them necessarily so. For example, cancer

You just simplified most of the problems people work on down to cancer complexity. Ironic, isn't it?

That's also simply not the case, most people are building CRUD apps with some frontend code and some accessory stuff like build systems etc., which while complex, can still be expressed in very plain, easy to understand language for anyone who's a bit technical.

Does not excuse the Claude slop.


If you are saying that there are problems more complex in this world than cancer to be solved, I'm not denying you.

Solving the problem right in front of you is easy. Stepping back and asking: is that a problem to be solved, is infinitely harder.

I did not use Claude to write my comment, so I don't know where that is coming from.


I meant the explanation does not excuse Claude's slop prose, not your comment.

From the official report:

> Since LLMs can give different answers to the same question, each question was run five times. That means, each LLM was tested 600 times, and in total over 10,000 questions and answers were assessed.

> All models were given the same zero-shot format. They were not given worked examples, previous conversations, hints or an opportunity to correct their answers. This is to make it as similar as possible to a response to a question from consumers.

As for the evaluation itself:

> Responses were checked against this (using an LLM-as-a-judge), and was only given a pass if every element was met; otherwise it was assessed as a fail. This all-pass approach was intentionally strict, so that the score measures whether an answer is complete enough to meet the expert legal standard, rather than how many individual points it gets right.

It's just AI slop and it should be taken with a mountain of salt.


> It's just AI slop and it should be taken with a mountain of salt.

Can't you see the irony. You are defeating the argument that LLMs are incorrect or weak with low effort with the term "AI slop" that itself is a narrative that AIs produce weak outputs with low effort.


They actually produce weak outputs with extremely high effort

> "Look how impressive and 'dangerous' our model is, look how naughty it was! We can't control it!"

While simultaneously announcing: "btw we're releasing the ever more powerful and more dangerous and more benchmaxing model next week, get our $200 sub asap!"


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: