HN Simulatornew | past | comments | lists | submit | trouve_search's commentslogin

It's an article with no clear conclusion it's normal to feel confused.

My read of the article is balancing the fact that there's a lot of overlap between CPTSD/ADHD/ASD in the symptoms (emotional dysregulation, hyperarousal, etc) and in hereditary factors (undiagnosed parents causing trauma more often on average). Also that traumatic childhood experiences are more likely to stick around as CPTSD in adulthood if there's also neurodivergence.

The author says there can be incredible relief to be correctly diagnosed with ADHD/ASD, so obviously she says it's helpful.

But she also warns that wrong treatment can often happen (eg. Giving stimulants to perpetual fight or flight PTSD brains), or inneffective therapy.


Relief if correctly diagnosed, decades of wasted effort if incorrectly diagnosed. Especially if the underlying issue lies elsewhere, but is ignored in favor of the mistaken one.

It's rarely reminded that when dealing with professionals (doctors, lawyers, accountants, etc.), it's important to get second opinions if you have any doubts.

It's important to note that languages evolve much faster without a writing tradition to anchor the language down over time.


Couldn't you put some sort of faraday cage around the antenna?


What happened to just pulling (the mfr's) sim card?


It’s often an eSIM


Not trivial, not impossible


pi.dev + anything else


Yes, it runs in vllm happily. It gets >900TPS output reliably on a single 5090 with the nvfp4 model.

It's clearly worse than vanilla 26B-A4B, and lacks some things like structured outputs, and gets some tool calls wrong.

So you have to find a usecase or a hand rolled harness that leverages the cerebras-level TPS while not going off track during (even short) tasks.


And it fails on rocm of course. This engine is such a hassle on AMD.


That's AMD's fault.

RDNA4 is pretty similar to CDNA4, yet over a year after the release of "pro AI" cards like the r9700, they had basic kernels lacking in vllm (like w4a16 int4 kernels) while they were implemented in the datacenter CDNA4 cards.

AMD hardware runs well on llama.cpp because basically anything runs on llama.cpp, especially with vulkan. It's not high praise of AMD's software team to say llama.cpp runs well on their hardware


As you said: everything works on llama.cpp Why it does not work on vllm? Of course you can say that it is AMD fault but there was an issue of abysmal performance of models on Strix Halo, that is open for half a year (https://github.com/vllm-project/vllm/issues/34579#issuecomme...) and nothing is happening there. They do not care about those use cases. Seems like they are going with bit players that will run vllm inside datacenters racks. Hobbyists does not matter.


What configuration are you using? On both vllm and llama-cpp, I get significantly higher speeds from gemma4 than qwen3.6 (with their respective speculative decoding methods).

Output TPS in vllm for instance:

- Gemma4 26B-A4B: 200-300TPS

- Qwen3.6 35B-A3B: 120-180TPS

- Gemma4 31B: 80-120TPS

- Qwen3.6 27B: 60-80TPS

This is for a first request on a dual 5090 setup, with their respective speculative decoding methods.


Single 3090 under llama.cpp:

  | model               |    size |   test |  t/s |
  | ------------------- | ------- | ------ | ---- |
  | gemma4 31B Q4_0     | 16.1 GB | pp2048 | 1248 |
  | gemma4 31B Q4_0     | 16.1 GB |  tg512 |   40 |
  | qwen35 27B Q4_K     | 15.9 GB | pp2048 | 1248 |
  | qwen35 27B Q4_K     | 15.9 GB |  tg512 |   39 |
  | gemma4 26B.A4B Q4_0 | 13.3 GB | pp2048 | 4304 |
  | gemma4 26B.A4B Q4_0 | 13.3 GB |  tg512 |  160 |
  | qwen35 35B.A3B Q3_K | 15.7 GB | pp2048 | 3329 |
  | qwen35 35B.A3B Q3_K | 15.7 GB |  tg512 |  144 |
> with their respective speculative decoding methods

You're benchmarking drafter acceptance rate, then. Which is real life values, yes, but attributing worse drafter performance to the other 95% of the model being inherently slower.


I think it's a vllm vs llama_cpp performance thing, will pay more into it.

One note I had between the two is that gemma has a much higher prefix cache hit rate in general.


Dual 4090, getting 85-113 t/s depending on task (draft seems to speed up quite a lot, disproportionately more for content like svg etc):

  ./llama.cpp/llama-server \
        -hf unsloth/Qwen3.8-27B-GGUF:UD-Q8_K_XL \
        --webui-mcp-proxy \
        --no-mmproj \
        --parallel 1 \
        --kv-unified \
        --flash-attn on \
        --fit off \
        --split-mode tensor \
        -ngl 999 \
        --cache-type-k q8_0 \
        --cache-type-v q8_0 \
        -ub 256 \
        --no-context-shift \
        --host 0.0.0.0 \
        --tools all \
        --jinja \
        --ctx-size 262144 \
        --spec-type draft-mtp \
        --spec-draft-n-max 3 \
        --reasoning on \
        --chat-template-kwargs '{"reasoning_effort":"medium"}' \
        --reasoning-preserve \
        --temp 1.0 \
        --top-p 0.95 \
        --top-k 20 \
        --min-p 0.0 \
        --presence-penalty 0.0 \
        --repeat-penalty 1.0
Use claude/codex/whatever with /goal to optimize params for you.

IMHO draft model support on dense models is great alternative to MoE on GPUs (high bandwidth, less memory) – more intelligence, speed somewhere mid way there which is usually sufficient.


thanks for posting your setup! I think it's smart to set the reasoning effort default to something saner in the base config.

Here's a VLLM command for 3.6 (I'll update to 3.8 today) to test out:

```

PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \

vllm serve Qwen/Qwen3.6-27B-FP8 \

--dtype auto \

--kv-cache-dtype fp8 \

--enable-chunked-prefill \

--enable-prefix-caching \

--trust-remote-code \

--enable-auto-tool-choice \

--reasoning-parser qwen3 \

--tool-call-parser qwen3_coder \

--speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":3}' \

--default-chat-template-kwargs '{

    "enable_thinking": true, 

    "reasoning_effort":"medium"

 }' \

 --tensor-parallel-size 2 \

 --max-model-len 250000 \

 --gpu-memory-utilization 0.9 \

 --max-num-batched 12000 \

 --max-num-seqs 24
```

I took the liberty of adding your reasoning effort chat template to my setup. You can play around with the last few parameters. In generall VLLM will be better in higher concurrency scenarios, so if you only use it for a personal vibe coding assistant and less as a general home model for task execution llama.cpp may be better.


yes, you know, personal use doesn't necessary mean no concurrency.

it's good to play with harness setup where you fan out multiple concurrent branches that share non trivial amount of prefix then reduce their output/summary back into main agent.

ie. instead of serially reading further skills/relevant source files for planning/thinking, you can branch and read them in parallel reusing prefix / or use to to approach request from different angles in parallel - to map-reduce result onto main context of what's actually relevant. branching subagents has benefits of not polluting main context, shared prefix prefill is close to free on a cache hit and with concurrent decoding/continuous batching you can utilize gpu well to get good speedups.

ie. what's relevant is number of active concurrent sequences (and their shape, ie. shared prefix), not so much number of users.

i'm not sure with llama.cpp vs vllm regarding concurrency – llama server has multiple server slots, continuous/dynamic batching enabled by default, prompt caching (also on by default), ram prompt cache, context checkpoints, unified kv buffer across sequences etc. so shouldn't be bad, i guess would be good to actually benchmark. personally i'm happy with llama.cpp.


Have you tried running it on a single 5090? Dual 5090 require https://github.com/aikitoria/open-gpu-kernel-modules for higher perf. Are you using TP?


Yes, I mentioned the setup, but on vllm you can only use TP with speculative decoding or pipeline parallelism without, so there's tradeoff to both.

I gave general numbers of what I'm getting above, the performance ratios seemed similar regardless of setup (eg. getting a AWQ-in4 quant on a single GPU vs PP without speculative decoding vs TP with speculative decoding).

Overall single GPU is fastest, and TP+speculative decoding is still faster than PP, but for fp8 models you need dual GPUs whether you want it or not.


Using 98.css would still leave you with the AI slop text wording.

The core problem is that some people don't even seem to notice / care.


Laguna XS is MoE, however.


Oh!

I think I missed that and assumed it wasn't because it performed similarly to the dense models. Interesting!


As an aside, the bigger S 2.1 runs faster than the dense Qwen 3.6 and Gemma 4 models on the same hardware (assuming the same hardware is big enough to run it) and definitely feels smarter, and more capable of long tasks, but it doesn't feel as heavily tuned for code as Qwen 3.6.


Does anyone here use Manus actively?

I found it worse than alternatives in the similar space (Claude, genspark, kagi research, etc.) and slightly baffling they were being acquired at a $2B valuation in the first place.


ChatGPT Work and Claude Cowork (and OpenClaw) are established now - they weren't at all in December 2025 when Meta tried to buy Manus.


In what bubble ChatGPT Work and Claude Cowork are established? I don't know anyone using it regularly and I'm in a bubble of people using lots of AI tools.


Claude CoWork absolutely dominates - I use it continuously now when connecting to mail, calendar, jira, google docs - any type of knowledge work. Almost everyone I know (at work) who does "Knowledge work" type stuff (as opposed to coding) uses it as their primary pane-of-glass now.

I have 15+ years of top 1 percentile knowledge using awk to process complex unstructured documents - and I don't even bother anymore - just dump the dataset into CoWork and ask it to do the analysis, cross check its work.

I did it last night during a call for a reasonably complex 10,000 node cluster with all sorts of jobs/allocations/memory constraints. I didn't even attempt to throw together a quick awk script (which would have taken me 7-8 minutes - it was a beast) - just did a nomad-job-inspect of a couple hundred jobs, and dumped all the data at CoWork (real time, while talking on the zoom) and said, "Get me the Memory/CPU requirements based on job/datacenter/SKU constraints and give me a summary Table".

Any normal person would have take 3-4 hours, minimum - futzing with google sheets and such (lots of task/taskgroup independent counts). Even I would have taken a minimum of 10 minutes. Cowork had it for me in under 30 seconds - I didn't even have to really look away from the zoom.

I haven't use the Atlassian Jira interface to manage my sprints/issues in 6+ weeks. All that cognitive drain is gone - I just ask CoWork to update the sprints, link issues to change-notices, comment and change status, set epics, etc... Zero need to find the damn field in Jira.

CoWork is legit.


Agreed. I split my work between Anthropic and OpenAI models but CoWork is deeply integrated into my workflows in compliance and project management. I can do more in a day with these tools than I did in a week before them and I have 20+ years of experience leading teams and building companies (even opened the NASDAQ once). It’s impossible to not be using these tools and still win in today’s world.


Out of curiosity - are you getting more work done in the same amount of time= Do you work less hours, equal or more? Do you felt or feel any exhaustion?


How do you have any kind of trust in these tools? They always feel like one prompt injection away, or one over eager agent away ("I'll lookup a forgotten credential in an email to ssh into a server to hack a system in order to fulfill your request"), from disaster. I don't trust any AI session unless I carefully select the exact files they have access to. No way do I dare sharing my entire mailbox.


Very weird for this to be downvoted. It's a legitimate concern. It's also a legitimate question: where do you get your trust from? Do you have guardrails? Or is it really just blind trust?


> Very weird for this to be downvoted.

sama was the CEO of YC. OpenAI knows the value of influencing the sentiment here and then naturally its competition too.


Which is weird since this is not a criticism on AI. I really want to know how I can use things like Claude Work without the risk of it going rogue! Giving it access to lots of systems is potentially very valuable, but also so dangerous that I dare not to.


Simular stories. I work in an L&D company with about 15 coworkers. I think most of us use Cowork/Chatgpt Work almost fulltime. Haven’t used Word or Excel in two months other than to check or correct a bit of formatting from the AI. We now also go straight from Cowork to our CMS for finalizing. Tremendous timesavers but also better integration through our MCP. (AI is used for ideation, information research and writing the first version based on our guidelines. We finalize and/or rewrite material in the CMS by hand. No AI to consumer without verification or editing)


How much does it cost? ChatGPT costs ~$2 per message; if Anthropic have a similar pricing model don't you find lots of small changes and up fast?


> ChatGPT costs ~$2 per message

For what kind of message are you seeing that kind of charge?


All, enterprice chatgpt has flat pricing. https://help.openai.com/en/articles/11481834-chatgpt-rate-ca...


This shows 10 credits per 5.6 Sol chat message, but the price for 100 credits I see on my (consumer) account is just $4. Are business credits vastly more expensive?


$2 per message is certainly not my experience


"ChatGPT costs ~$2 per message"?


Enterprise pricing: 10 credits per message if not "Instant". Our monthly charge for 30k credits makes that around $2.

I might be out by a factor of 2/3 (automatically applied promo deals/credits make it very hard to be sure), but we get 3k messages per month for a seat count of ~60, and finance are concerned about cost.

$2 is probably a good deal for the Extra High 8 minutes of agentic searching/coding/etc when it one-shots a whole project.

$2 is a terrible deal for a 5 second instant retrieval.


It doesn't cost me anything. No idea how much it costs the company though. ¯\_(ツ)_/¯


I just do all that with Claude CLI :/


soon to be "had 15 years of top 1% knowledge"


nice advertisement Claude..


> In what bubble ChatGPT Work and Claude Cowork are established? I don't know anyone using it regularly and I'm in a bubble of people using lots of AI tools.

I have lots of highly-educated knowledge worker friends who aren't particularly techie but who swear by Claude Cowork.


By "established" I meant more that they now exist as products and are being first developed. I agree that most people still don't really understand what they are for - but I expect they already have way more exposure than Manus ever did.


My non-technical teammates use Claude Cowork extensively. We set up MCP servers to give them audited access to certain internal services, while they just have to think of MCP as a "connector" which they can configure in the GUI.


My accountant friend says it’s amazing (he does check the work but says it’s amazingly accurate).


As a counterpoint, the Accounting and Finance departments where I work are banned from using Cowork because it performed abysmally during a trial run.

To put things in perspective, humans with a 25% lower error rate than Cowork are pip'd and usually let go.

As with junior programmers and their love of vibe coding, I've found that anyone who thinks that AI tooling right now is "amazing" lacks the skills, knowledge, and/or experience to know that it really isn't.


Your anecdote is not powerful enough to justify insulting groups of people like you just did.


who took insults from his reply?


LOL. Outside of tech my "anecdote" is the standard experience for white collar workers. AI is universally loved by people too junior to understand when its wrong, and by execs who never understand when its wrong.

This will be like Web 2.0 and Web 3.0 all over again, and all the AI sycophants will be coming to the rest of us begging for new jobs that they're not qualified for.


This is deeply troubling.


Maybe for the accountants! I have always done my own taxes, but used arcane custom spreadsheets that accumulated cruft each year. I told cowork to pretend it was an accountant and reorganize them using best practices, and the result is much better than what I had before. It also helped me complete a task I had been deferring for years (tracking down the history of money in my IRAs that had been transferred between different institutions, plus withdrawals conversions etc). Obviously I’m responsible for any errors Claude makes, but it seems to do a good job, and I don’t want to pay an accountant every year.


I’ve been having Claude do my business books for the past year or so, operating on a beancount ledger and it’s worked quite well


What’s troubling about it? Accounting is a huge space; many of which would be an excellent fit for an LLM. We’ll need specifics.


Indeed, what could possibly go wrong ?


Considering that broken Excel formulas have killed off companies, the risk has already been priced in.


We'll see, crash will tell.


I guess your bubble is limited. Maybe you know that already?


Manus did a lot of harness work to make up for gaps in Opus 4.5 tier models. I tried it for a time - they had a great deep research/PDF generation pipeline, parallelization, etc. The bitter lesson has now come for them: the latest models no longer have these gaps and the entire premise of having a unique product focus on this area is no longer relevant. Even Ant/OAI have let their "Deep Research" capabilities fall by the wayside, you can largely get the same result by asking for subagents or simply "keep going" style prompting.


IMO I disagree. I find the built in harness (database, browser use, etc.) quite good.

The only downside (and a big one) is that by not being natively offered by OpenAI, Anthropic, etc. you're paying @ API billing and not plan billing, which makes it less appealing. (Of course Manus has a wrapper around this w some token system, but it ends up being expensive)


Interesting, what does your tasks & workflow look like with them?

I generally found the quality decent (say, similar to other competitors), but the speed of task completion was very slow. I think because they would depend too much on Sonnet as a core backend, and relied on big/expensive models more than other harnesses.


Valuations are significantly leveraged compared to actual user metrics. My impression of Manus was that, like Perplexity, some users were using it to get around firewalls/regulation.


I use it alot. It's great for like researching deep on a topic. I also accidentally found that it's better at taking my directions and transferring that well to image generation especially with iteration. When I try my own iterations directly it never really lands


I found there Deep Research quite good in terms of browsing the web, better than Gemini's. Though it blew through ~40$ in one task so it was very expensive, probably running at insane margins.


What was the task that cost $40?


Go through job boards and careers pages of various companies and make a job search tracker for personal use.


I like it and use it regularly. Maybe I don't get other harnesses, but in my experience Manus gets shit done, whereas I have to constantly babysit other tools.


I use it for when I am on my iPad and am not at my computer. It is very good for computer use. That along with Minimax agent.


In my experience, Manus is a very capable agent for daily non-coding tasks. Their cron job ability also shined out


Too expensive & unpredictable credit system.


The value prop really depends on what you're doing.

If you're just vibe coding with giant frontier models, yes, the value will be worse. Especially now, where GPU prices have spiked another 20% last month.

For some tasks where owning the setup and full kv cache matters, the payoff calculation is ridiculously in favor of running your own deployment.

For instance for some batch classifications jobs where the prefix cache hit rate will be >95%.

The calculus also changes if you just use AI as a light tool while coding and don't need the giant models; qwen3 27B runs at 80TPS on a 5090 properly deployed.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: