I saw some guy streaming how he was deploying qwen3.8 37B on his local setup. Well, he was asking Claude to do it. It took him two hours of passing errors to Claude for the endpoint to start working, he then started testing it against DS4 Flash when Qwen had thinking disabled and Claude messed up sampling parameters, it was an absolute pain to watch
Wow. I tried to get Qwen3.8 4B to parse song lyrics and analyse them. Getting ollama running was a minute or two.
However coming up with a prompt that didn't turn out total garbage was impossible. After wasting over an hour and I ended up getting Qwen side by side with Llama 3.2 3B, just to see if I was being stupid. Nope, it just looks like Llama is orders of magnitude better at this specific task for some reason).
If you think I'm doing it wrong, you're probably right, I don't know a ton about local LLMs. But I hand selected 50 songs, set up ollama with both LLMs, and for each iteration on the prompt text, ran both LLMs 10x times per song. Side-by-side comparisons showed that Qwen 3 4B was so bad that I actually downloaded Qwen again, thinking there must have been some mistake and I accidentally grabbed an old 1B model.
> Nope, it just looks like Llama is orders of magnitude better at this specific task for some reason
This is almost every ML model, if the task isn't part directly or indirectly of the datasets they use for training it, then the model is gonna be pretty trash at it. What the big AI labs have over the smaller labs, is a huge amount of data and diverse set of tasks, hence they generalize better, but still not great.
So, how do you avoid having to spend hours on figuring out if the model is just dumb, or don't know the task? Your own private benchmarks! Figure out a way, ideally without using another LLM, to score how good a model is at doing your specific task. Come up with 3-5 examples for this benchmark yourself, ask a SOTA LLM to fill out 45 more, review everything VERY closely, then use this whenever you want to figure out if $new_model actually is an improvement over what you use today, and once you have a bunch of different tasks, you'll see that all these HUGE improvements tend to be specifically for the benchmarks they mention in the press release, as many of your own benchmarks won't show that big of a difference in reality.
Yes there is a higher upfront cost, but if you're building longer-term projects that rely on LLM models, particularly local ones that seem very benchmaxxed a lot of the times, you need a quick and reproducible way of scoring them somehow, where you can just add more models to compare, and you need to keep these benchmarks to yourself.
> What the big AI labs have over the smaller labs, is a huge amount of data and diverse set of tasks, hence they generalize better, but still not great.
Hehe, this kind of sounds like the opposite of generalization. As in it’s just specialization at scale.
Tbh the way you're treated depends on the hotness of the task.
A couple of years ago, people I know got paid OK for relatively simple programming and logic RLHF tasks. But very soon it turned dark because that sort of data was required less and less, and the number of feedbackers has grown.
Today, the type of data the model developers pay for requires actual domain experience. E.g in software engineering they have people work in simulated environments with other LLMs, grade them, feedback, PRs, Jira everything.
This pays well and they treat you well, entice you with more money/task/hour, etc, for now. In few years when this gets drilled into LLMs, these guys will face the same painful hours and bad pay and less work and so on too.
In physical tasks, we are still in the early stages where basic packing clothes (in a textile factory setting) etc is being recorded and data is only now being used for training. Due to the problems with translating human hand data to robotic hands, these people do the factory work holding robotic grippers and operating that, you should see a YouTube video. But this also means that it's much less sweeping than data collection in SWE. In many cases it's not practical to collect data given that you have to use the specific gripper, wear a big gopro type thing, etc. So I am expecting much slower of an impact on physical tasks (of this kind) compared to how quick the uptake was in SWE/math.
I have yet to hear back from them on how it's going for teamwork white collar tasks, it's been a few months. Everyone is paying for the end products of those it seems - grok bot, perplexity computer, claure cowork, chatgpt work etc,. Not as much as their coding agents of course.
This is the way. It's also very important to automate as much of this verification as possible into the harness, rather than sit there and prod it in the chat.
Of course some things are not auto verifiable, and you'll have to give human judgement and input there, but you'll save a lot more time if you spend 1 week painstakingly writing checks for as many little things as possible and integrating them into the harness.
That's funny, because I just went through the opposite.
I got llama.cpp working with qwen3.6 and qwen3.8 by Googling and manually adjusting things according to reddit posts and Google not-really-helpful AI suggestions.
I tried settings up per-model stuff in settings.json, but again Google got in my way, and llama.cpp having 2 different settings.json (and Google lying about where 1 goes) made it far too difficulty to figure out. I spent hours on it.
Then I got fed up and asked Claude.
Immediately, it told me that the winget version of llama.cpp is for Vulkan, and I needed a different one and pointed at it. It doubled my speed.
Then it figured out what I was doing wrong with settings.json (wrong spot, global settings can't go in the per-model file, etc etc) and fixed all that, and got it working.
Then it tuned it somewhat.
Then I showed it the official settings pages for both models, and it undid the tuning and all the damage I had done with my tinkering, and got everything working.
In 30 minutes.
It was absolutely amazing.
Every time I see people recommending Qwen locally with llama.cpp, they just say "download it" and act like anyone that can't get it running is an idiot. But if there's a "using this settings.json" tutorial somewhere, I didn't find it, and neither did Google over a week of searching.
But Claude got it done for me.
Now, I admit, I haven't played with it much. Just before all this, I ran out of Claude on the $20 plan and bumped up to $100, and It's been so amazing that it's really hard to work on the local. Especially since it feels like Qwen3.8 35b a3b is probably around the corner, and why mess with 3.6 when 3.8 will probably release soon?
On DeepSWE it's now 53% vs 63% which is one of the coding benchmarks I trust the most. DS own measurements also show a more significant increase so I suspect AA might update when they release an article.
Surprisingly DeepSWE currently shows a lower total cost for pro so that might also update I guess. As usual, don't trust the benchmarks and try for yourself.
The idea is, I believe, that the Pro model is a larger model (more parameters, or less quantization) in general. What implication that has, I couldn't tell you.
For tasks like pondering on something, reviewing code, etc. I use Pro, just because it feels like the right model for that.
I was recently working on an AI harness too, but I wasn't using AI to code it. It's really easy and requires little code. I wasn't even using langchain - that would require even less code.
UI is harder for sure, but it's not that bad. You need to think though where to catch which exceptions.
LLMs might opt for langchain which has had multiple breaking changes after the knowledge cutoff, making it hard for the LLM to work with it. This is probably going to lead to the LLM having to make many changes to it's code, making it messy and leading to further code being less maintainable.
They care so little they don't even bother to proofread once, or lack enough intelligence to distinguish a prompt meant for a machine from a speech to inform/inspire/influence a human audience.
It's a red flag. Even if they agree and understand whatever they copied from chatgpt, they were too lazy to clean up its output. Are they also too lazy to read the news or does chatgpt summarize their articles too? AI people are like emperors without clothes. Totally embarrassing to witness.
This is bad because it's only getting harder to trust people. This is not someone I would want representing me. Although I'm not thrilled about any of my representatives anyway.
You're absolutely right. Sure, I'll help you write a Hacker News comment explaining why politicians should have functional mental faculties and a coherent understanding of what they do during the course of their jobs and how their actions affect their constituents.
It's not a matter of whether the theory "works"; it's a matter of whether one is asking the right questions. Convex optimization studies how quickly an optimizer can reach the optimum. In the non-convex case, there are many basins containing their own local minima. The more sensible questions there are "which basin is it likely to go into?" and "how do I steer it to go where I want?". Global convergence rates are largely irrelevant by comparison.