It's an article with no clear conclusion it's normal to feel confused.
My read of the article is balancing the fact that there's a lot of overlap between CPTSD/ADHD/ASD in the symptoms (emotional dysregulation, hyperarousal, etc) and in hereditary factors (undiagnosed parents causing trauma more often on average). Also that traumatic childhood experiences are more likely to stick around as CPTSD in adulthood if there's also neurodivergence.
The author says there can be incredible relief to be correctly diagnosed with ADHD/ASD, so obviously she says it's helpful.
But she also warns that wrong treatment can often happen (eg. Giving stimulants to perpetual fight or flight PTSD brains), or inneffective therapy.
Relief if correctly diagnosed, decades of wasted effort if incorrectly diagnosed. Especially if the underlying issue lies elsewhere, but is ignored in favor of the mistaken one.
It's rarely reminded that when dealing with professionals (doctors, lawyers, accountants, etc.), it's important to get second opinions if you have any doubts.
RDNA4 is pretty similar to CDNA4, yet over a year after the release of "pro AI" cards like the r9700, they had basic kernels lacking in vllm (like w4a16 int4 kernels) while they were implemented in the datacenter CDNA4 cards.
AMD hardware runs well on llama.cpp because basically anything runs on llama.cpp, especially with vulkan. It's not high praise of AMD's software team to say llama.cpp runs well on their hardware
As you said: everything works on llama.cpp
Why it does not work on vllm? Of course you can say that it is AMD fault but there was an issue of abysmal performance of models on Strix Halo, that is open for half a year (https://github.com/vllm-project/vllm/issues/34579#issuecomme...) and nothing is happening there. They do not care about those use cases. Seems like they are going with bit players that will run vllm inside datacenters racks. Hobbyists does not matter.
What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods).
Output TPS in vllm for instance:
- Gemma4 26B-A4B: 200-300TPS
- Qwen3.6 35B-A3B: 120-180TPS
- Gemma4 31B: 80-120TPS
- Qwen3.6 27B: 60-80TPS
This is for a first request on a dual 5090 setup, with their respective speculative decoding methods.
> with their respective speculative decoding methods
You're benchmarking drafter acceptance rate, then. Which is real life values, yes, but attributing worse drafter performance to the other 95% of the model being inherently slower.
Use claude/codex/whatever with /goal to optimize params for you.
IMHO draft model support on dense models is great alternative to MoE on GPUs (high bandwidth, less memory) – more intelligence, speed somewhere mid way there which is usually sufficient.
I took the liberty of adding your reasoning effort chat template to my setup. You can play around with the last few parameters. In generall VLLM will be better in higher concurrency scenarios, so if you only use it for a personal vibe coding assistant and less as a general home model for task execution llama.cpp may be better.
yes, you know, personal use doesn't necessary mean no concurrency.
it's good to play with harness setup where you fan out multiple concurrent branches that share non trivial amount of prefix then reduce their output/summary back into main agent.
ie. instead of serially reading further skills/relevant source files for planning/thinking, you can branch and read them in parallel reusing prefix / or use to to approach request from different angles in parallel - to map-reduce result onto main context of what's actually relevant. branching subagents has benefits of not polluting main context, shared prefix prefill is close to free on a cache hit and with concurrent decoding/continuous batching you can utilize gpu well to get good speedups.
ie. what's relevant is number of active concurrent sequences (and their shape, ie. shared prefix), not so much number of users.
i'm not sure with llama.cpp vs vllm regarding concurrency – llama server has multiple server slots, continuous/dynamic batching enabled by default, prompt caching (also on by default), ram prompt cache, context checkpoints, unified kv buffer across sequences etc. so shouldn't be bad, i guess would be good to actually benchmark. personally i'm happy with llama.cpp.
Yes, I mentioned the setup, but on vllm you can only use TP with speculative decoding or pipeline parallelism without, so there's tradeoff to both.
I gave general numbers of what I'm getting above, the performance ratios seemed similar regardless of setup (eg. getting a AWQ-in4 quant on a single GPU vs PP without speculative decoding vs TP with speculative decoding).
Overall single GPU is fastest, and TP+speculative decoding is still faster than PP, but for fp8 models you need dual GPUs whether you want it or not.
As an aside, the bigger S 2.1 runs faster than the dense Qwen 3.6 and Gemma 4 models on the same hardware (assuming the same hardware is big enough to run it) and definitely feels smarter, and more capable of long tasks, but it doesn't feel as heavily tuned for code as Qwen 3.6.
I found it worse than alternatives in the similar space (Claude, genspark, kagi research, etc.) and slightly baffling they were being acquired at a $2B valuation in the first place.
In what bubble ChatGPT Work and Claude Cowork are established? I don't know anyone using it regularly and I'm in a bubble of people using lots of AI tools.
Claude CoWork absolutely dominates - I use it continuously now when connecting to mail, calendar, jira, google docs - any type of knowledge work. Almost everyone I know (at work) who does "Knowledge work" type stuff (as opposed to coding) uses it as their primary pane-of-glass now.
I have 15+ years of top 1 percentile knowledge using awk to process complex unstructured documents - and I don't even bother anymore - just dump the dataset into CoWork and ask it to do the analysis, cross check its work.
I did it last night during a call for a reasonably complex 10,000 node cluster with all sorts of jobs/allocations/memory constraints. I didn't even attempt to throw together a quick awk script (which would have taken me 7-8 minutes - it was a beast) - just did a nomad-job-inspect of a couple hundred jobs, and dumped all the data at CoWork (real time, while talking on the zoom) and said, "Get me the Memory/CPU requirements based on job/datacenter/SKU constraints and give me a summary Table".
Any normal person would have take 3-4 hours, minimum - futzing with google sheets and such (lots of task/taskgroup independent counts). Even I would have taken a minimum of 10 minutes. Cowork had it for me in under 30 seconds - I didn't even have to really look away from the zoom.
I haven't use the Atlassian Jira interface to manage my sprints/issues in 6+ weeks. All that cognitive drain is gone - I just ask CoWork to update the sprints, link issues to change-notices, comment and change status, set epics, etc... Zero need to find the damn field in Jira.
Agreed. I split my work between Anthropic and OpenAI models but CoWork is deeply integrated into my workflows in compliance and project management. I can do more in a day with these tools than I did in a week before them and I have 20+ years of experience leading teams and building companies (even opened the NASDAQ once). It’s impossible to not be using these tools and still win in today’s world.
Out of curiosity - are you getting more work done in the same amount of time= Do you work less hours, equal or more? Do you felt or feel any exhaustion?
How do you have any kind of trust in these tools? They always feel like one prompt injection away, or one over eager agent away ("I'll lookup a forgotten credential in an email to ssh into a server to hack a system in order to fulfill your request"), from disaster. I don't trust any AI session unless I carefully select the exact files they have access to. No way do I dare sharing my entire mailbox.
Very weird for this to be downvoted. It's a legitimate concern. It's also a legitimate question: where do you get your trust from? Do you have guardrails? Or is it really just blind trust?
Which is weird since this is not a criticism on AI. I really want to know how I can use things like Claude Work without the risk of it going rogue! Giving it access to lots of systems is potentially very valuable, but also so dangerous that I dare not to.
Simular stories. I work in an L&D company with about 15 coworkers. I think most of us use Cowork/Chatgpt Work almost fulltime. Haven’t used Word or Excel in two months other than to check or correct a bit of formatting from the AI. We now also go straight from Cowork to our CMS for finalizing. Tremendous timesavers but also better integration through our MCP. (AI is used for ideation, information research and writing the first version based on our guidelines. We finalize and/or rewrite material in the CMS by hand. No AI to consumer without verification or editing)
This shows 10 credits per 5.6 Sol chat message, but the price for 100 credits I see on my (consumer) account is just $4. Are business credits vastly more expensive?
Enterprise pricing: 10 credits per message if not "Instant". Our monthly charge for 30k credits makes that around $2.
I might be out by a factor of 2/3 (automatically applied promo deals/credits make it very hard to be sure), but we get 3k messages per month for a seat count of ~60, and finance are concerned about cost.
$2 is probably a good deal for the Extra High 8 minutes of agentic searching/coding/etc when it one-shots a whole project.
$2 is a terrible deal for a 5 second instant retrieval.
> In what bubble ChatGPT Work and Claude Cowork are established? I don't know anyone using it regularly and I'm in a bubble of people using lots of AI tools.
I have lots of highly-educated knowledge worker friends who aren't particularly techie but who swear by Claude Cowork.
By "established" I meant more that they now exist as products and are being first developed. I agree that most people still don't really understand what they are for - but I expect they already have way more exposure than Manus ever did.
My non-technical teammates use Claude Cowork extensively. We set up MCP servers to give them audited access to certain internal services, while they just have to think of MCP as a "connector" which they can configure in the GUI.
As a counterpoint, the Accounting and Finance departments where I work are banned from using Cowork because it performed abysmally during a trial run.
To put things in perspective, humans with a 25% lower error rate than Cowork are pip'd and usually let go.
As with junior programmers and their love of vibe coding, I've found that anyone who thinks that AI tooling right now is "amazing" lacks the skills, knowledge, and/or experience to know that it really isn't.
LOL. Outside of tech my "anecdote" is the standard experience for white collar workers. AI is universally loved by people too junior to understand when its wrong, and by execs who never understand when its wrong.
This will be like Web 2.0 and Web 3.0 all over again, and all the AI sycophants will be coming to the rest of us begging for new jobs that they're not qualified for.
Maybe for the accountants! I have always done my own taxes, but used arcane custom spreadsheets that accumulated cruft each year. I told cowork to pretend it was an accountant and reorganize them using best practices, and the result is much better than what I had before. It also helped me complete a task I had been deferring for years (tracking down the history of money in my IRAs that had been transferred between different institutions, plus withdrawals conversions etc). Obviously I’m responsible for any errors Claude makes, but it seems to do a good job, and I don’t want to pay an accountant every year.
Manus did a lot of harness work to make up for gaps in Opus 4.5 tier models. I tried it for a time - they had a great deep research/PDF generation pipeline, parallelization, etc. The bitter lesson has now come for them: the latest models no longer have these gaps and the entire premise of having a unique product focus on this area is no longer relevant. Even Ant/OAI have let their "Deep Research" capabilities fall by the wayside, you can largely get the same result by asking for subagents or simply "keep going" style prompting.
IMO I disagree. I find the built in harness (database, browser use, etc.) quite good.
The only downside (and a big one) is that by not being natively offered by OpenAI, Anthropic, etc. you're paying @ API billing and not plan billing, which makes it less appealing. (Of course Manus has a wrapper around this w some token system, but it ends up being expensive)
Interesting, what does your tasks & workflow look like with them?
I generally found the quality decent (say, similar to other competitors), but the speed of task completion was very slow. I think because they would depend too much on Sonnet as a core backend, and relied on big/expensive models more than other harnesses.
Valuations are significantly leveraged compared to actual user metrics. My impression of Manus was that, like Perplexity, some users were using it to get around firewalls/regulation.
I use it alot. It's great for like researching deep on a topic. I also accidentally found that it's better at taking my directions and transferring that well to image generation especially with iteration. When I try my own iterations directly it never really lands
I found there Deep Research quite good in terms of browsing the web, better than Gemini's. Though it blew through ~40$ in one task so it was very expensive, probably running at insane margins.
I like it and use it regularly. Maybe I don't get other harnesses, but in my experience Manus gets shit done, whereas I have to constantly babysit other tools.
The value prop really depends on what you're doing.
If you're just vibe coding with giant frontier models, yes, the value will be worse. Especially now, where GPU prices have spiked another 20% last month.
For some tasks where owning the setup and full kv cache matters, the payoff calculation is ridiculously in favor of running your own deployment.
For instance for some batch classifications jobs where the prefix cache hit rate will be >95%.
The calculus also changes if you just use AI as a light tool while coding and don't need the giant models; qwen3 27B runs at 80TPS on a 5090 properly deployed.
My read of the article is balancing the fact that there's a lot of overlap between CPTSD/ADHD/ASD in the symptoms (emotional dysregulation, hyperarousal, etc) and in hereditary factors (undiagnosed parents causing trauma more often on average). Also that traumatic childhood experiences are more likely to stick around as CPTSD in adulthood if there's also neurodivergence.
The author says there can be incredible relief to be correctly diagnosed with ADHD/ASD, so obviously she says it's helpful.
But she also warns that wrong treatment can often happen (eg. Giving stimulants to perpetual fight or flight PTSD brains), or inneffective therapy.
reply