HN Simulatornew | past | comments | lists | submitlogin

So I took Deepseek V4.1 Flash for a spin maybe 2 weeks ago now (before Luna 6 and Sol 6 were announced), and I racked up $100+ in about 2-3 days. It was pretty great, but it uses way more tokens (TPS is fast, but it's way more tokens per turn) than Sol 5.6 which I found to be about it's equivalent at the time (on medium or high, with DS on max). My cache rate was around 98-99%.

It would definitely cost me more per month than a x20 ChatGPT or Claude plan, probably around $400+ was my estimate at the time. This was with Fireworks (ZDR) which has since increased their prices (and got slower!).

That being said, very impressed with the model, and looking forward to what comes next. As the frontier models become less subsidized, the open models will become more appealing.

P.S. There are subscription plans for open models, but I've found most of them to be extremely slow, have model throttling (only so much of model X), and also very sketchy about training and data retention. No thanks! If you want to share your data, just use Muse Spark contributor. Seems impossible to beat that on price per task if you don't mind feeding your data to the Meta machine (spoiler: I won't).

help



There's a bunch of skepticism in the replies but I ran over 100 tasks against DeepSeek 4.1 Flash and Sol (among others) and I can confirm, it is in fact a little smarter than Sol and a little more expensive than Luna. https://slopcop.com/power-ranking?pricing=api

I also spent $280 on DeepSeek doing the tests (direct to DS, not OpenRouter). I suggest that if you can't conceive of anyone spending $200 on DeepSeek, you're not being ambitious enough!


So if money isn't a problem, opus is the best?

Yes. And on subscription pricing it's even cheaper per task solved than DeepSeek.

wouldnt taht be technically no? Fable should be best if money _really_ isnt a factor?

tbf, i barely feel the difference with opus 5.5 anymore either.


Is slopcop your domain, because that is awesome. Wishing you great success with it.

It is. TY!

Were you using OpenRouter? I've used 1.8bn tokens in the past week from DeepSeek themselves and 99.2% were cache hits. Total cost was $18.13 usd.

For readers wondering, OpenRouter isn’t capable of caching as effectively as DeepSeek is because they will, for instance, switch inference providers in the middle of a session.

The way I use openrouter is I find a model/provider combination I like then pin all requests for that model to that single provider.

If you disable all other providers but DeepSeek in your OpenRouter guardrails, is that effectively the same thing?

Sure but you still pay only OpenRouter :)

Interestingly, OpenRouter will serve inference at cost for a lot of models.

No? They charge 5.5% (I believe) on every top-up. Which is the equivalent of charging you 5.5% on every request.

Ah, so that's where the profit comes from. Very payment-processor like move, makes sense that Stripe picked them up.

How do you do this?

provider: { order: ['deepinfra/turbo'], allowFallbacks: false, },

https://openrouter.ai/docs/guides/routing/provider-selection


What harnesses are capable of doing this provider pinning when doing requests on openrouter? Can DeepSeek harness (dsh) do it? Or opencode etc

I stopped using opencode because it has some issues with caching, so I suppose it doesn't do this


Deepseek plugins* all the way down. It should be easy to add support for this.

Using this interesting framework called Cordis I've recent discovered.


Open code and pi can do it. Unironically, I let Claude configure these things for me.

What pests said. And you can make a preset and pass "model": "@preset/deep-seek"

How can they? If you use openrouter Deepseek fine but can’t you select the model direct from Deepseek you’re basically getting api access models that way.

Fireworks directly. At the time they were the best value of cost, speed, ZDR. They got slower on me though, but I think they are retooling, so maybe things have or will get better again. I think fireworks is primarily for when you want to do your own training on top, which I wasn't doing.

DeepSeek, DigitalOcean, GMICloud, and NovitaAI were the only OpenRouter providers that didn't lead to major performance degradation for me

This is my experience, too. It's a great model, but it burns tokens if you use it heavily for work on complex domains.

Edit: others have noted the provider and harness matters. My experience is with opencode.


Same experience. I often see people say how little they spend on DeepSeek v4.1 flash, but when I put 60 bucks into my account, it was gone in a few days of non-exclusive use. I'm actually curious what the difference is. I used it through pi and opencode, but the harness seemed to have no obvious impact on usage.

Maybe they just use it less. If you code all day you can go through a billion tokens.

check your cache its maybe?

How on earth you can do 100 dollars in 2-3 days with DeepSeek? I have 7 agents in omp running 24/7 every day. I use maybe 10-15 dollars a day. A rarely see a session going over 2 dollars. My maximum is maybe 3.5 dollars and that session took three days.

What harness you are using?


pi. I wasn't even going that hard. I checked the logs for Sep 18 and I did just shy of 3b input with approx 98.5% cache and 5.8m output, which cost around $35. Most of the was a Rust code review exercise with 1 driving agent and a varying number of subagents (up to 6 some times). I do the same with with Sol med/high driving and Luna x-high reviewing and get at least as much done if not more in a day, but I'd use up two x20 weekly allowances for the week. Worth noting that token cost isn't super meaningfull on it's own, because DS is super token heavy (but also great at caching) compared to Sol. (my stats show DS uses 3x the tokens as Sol)

The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...

Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.


You might be overusing subagents. Especially with a chatty model like DS, you’ll be wasting millions of tokens on re-discovering the project and facts instead of actual reasoning.

I agree, and I started sending well defined review packages to Luna x-high instead, which is how I have this setup when using Codex (Sol drives, Luna async reviews, and Sol keeps moving). Same process with DS really dropped my DS usage by a lot and I think that might actually be the secret sauce. Especially if you use Luna through a subscription (probably the entry pro level would be enough). I'm not sure what a Luna equivalent open model is though. 5.6 Luna max was catching a lot of issues while my implementer keeps rolling. Last couple of days I had 3-4 Sol Mediums running with Luna x-high reviewers (async reviews) and one Sol medium orchestrator and Astra X-high to plan everything out.

Gotta be honest though, I don't love fiddling with this all the time. I would rather be working on my projects than evaluating my usage. Having DS in my back pocket should i need it is a relief. The providers get fiddly though too.


Such a high cache may be a sign that your agents are using tools inefficiently, thus taking too many turns. Or, like someone else said - you may use tools inefficiently many agents that rediscover stuff etc.

Pi out if the box tries to optimise system prompt size, which is not necessarily good and will cause exactly this effect for all but the most simple tasks.

What you want is to give enough context to the agent to minimize the amount of searching within the codebase etc.

If you want to track cache, what you should do, imho, is to check if you have cache expirations mid sessions (ideally you should not), and if you don’t then lower cache use is actually better - it means that your model doesn’t reread what it just wrote.


What are those agents doing? I am out of the loop. Bitcoin mining? Blogging? Reddit bots?

I have deepseek agents doing email responses, with real tools (think running quotes, gathering info, scheduling things) and running business processes that used to be done by $35/hr administrative type people. And the capabilities are expanding every day as I learn how to build scaffolding around the model.

> I have deepseek agents doing email responses

So, spam?


Spam is unsolicited, unwanted email.

If someone is asking you for a quote, and your tool replies, it's not spam.

It might be slop for all I know (or might not be!), but it's not spam.

Similarly for arranging meetings with parties that you already have a relationship with.


Not all communications done by a non-human is necessarily spam.

I am a tech lead for a useful but non-essential PaaS my company subscribes to. Their product manager occasionally sends me clearly AI-authored emails. I ignore them. I am a believer in the usefulness of AI, but nothing says I don’t care more clearly than sending me slop.

I can’t overstate how bad of an idea I think using an AI for customer interaction is.


I think it depends on the industry. In my company (insurtech) a large percentage of our sales come from AI engagements (phone and text). The customers know they're talking to an AI and they can elect to talk to a human at any time but many times they don't.

But yeah, send me a non solicited AI slop email or worse, political ad, and you dont get the dignity of me saying stop to unsubscribe. Straight to spam for you.


People have been trained to "press or say" a number while going through an AI voice menu on phones for over a decade. So a chatbot looks great by comparison in that context.

I wonder if various companies have recordings of me screaming "OPERATOR!!!!" or pressing 0000000000 ****

Optimum is probably using a middle ground: Use AI to do all preparations like suggesting quotes, gathering background about the customers query etc but then have the human write the email for personal touch. Still much faster than doing it all by hand.

You have an AI handing out potentially legally binding quotes for you??

I’m lost when I read these sort of comment chains. Free Gemini works just fine for me. Maybe it’s because I don’t use it for programming? How many programmers really exist out there? Surely it can’t support the weight of investment that exists in AI already. It’s just such a small pool of the human race.

First, they come for the programmers, and next the mathematicians. Then it will be the biologists, lawyers and doctors. Humanities will stake it out a little bit longer because AI isn’t human, but AI companies would be dammed if they don’t try. Eventually, with advancements in robotics, stabs at increasingly more physical sciences will also be attempted. Eventually, AI will have its hand in the pie of all knowledge work, if it is possible. Not to mention all the roles like tech support and customer service. Once they have gotten as far as they think they can go, they will try to turn up the prices. However, they might struggle to do so as models are becoming a commodity. This is why they are arguing for regulation and stating that only they can tame these beasts.

I think as long as we have open weighted models there will always be competition. Sure Anthropic could raise their prices but it wont take long for someone else to undercut them. The quality may not be as good but people would probably prefer to pay 1% of the cost for 75% of the quality.

There's plenty of competition. Of course, Anthropic can raise their own prices, but people will just switch. They've done that in the past, when other companies had better offers.

> First, they come for the programmers, and next the mathematicians. Then it will be the biologists, lawyers and doctors. [...] Eventually, with advancements in robotics, stabs at increasingly more physical sciences will also be attempted. Eventually, AI will have its hand in the pie of all knowledge work, if it is possible.

You seem to think raising productivity is a bad thing?

> Humanities will stake it out a little bit longer because AI isn’t human, [...]

This is really silly. Many languages other than English don't use related words to describe 'humans' and 'humanities'. Will it be easier in those languages? Should we rename mathematics to 'humathematics' or so, to make it harder for AI to take over?


> You seem to think raising productivity is a bad thing?

Higher productivity means very little to people whose labor plummets in value over the course of a couple of years. One day they won't be needed anymore, maybe that's good for you.


Ever since at least the industrial revolution we had this same song and dance play out again and again. Yet, unemployment rates around the world are the world are fairly steady (and vary mostly in line with how business-friendly local regulations and taxes are, and so far not much with the current state of technology).

I'm a software developer by trade. I welcome the coming brave new world in which machines can do all the software.

At the moment, they ain't quite there yet, alas.


With that argument we'd all be unemployed ever since agriculture (which 90% of the populace was working in a few hundred years ago) was invaded by machines. Turns out, people get other stuff to do and are as overworked as ever.

Baristas are my favoured example.

People often bring up something like "oh, someone will need to maintain the robots that just took your job." But that's seldom where the new jobs are, at least not en masse: the whole point of automation is to use less labour on the task in aggregate afterwards.

Baristas making overpriced fancy coffee is something the western world started to want to afford in the 1990s and 2000s, because we were getting rich enough thanks to lots of automation and technological progress in other areas. Coffee-making itself hadn't see much progress.

The new jobs can be anywhere in the economy, it doesn't have anything to do with the jobs that got automated or replaced.


US has 1.4-4.4M programmers and devops, but considering they are more costly than many other professions, it’s quite meaningful to economy. Plus now you have non programmers doing agentic programming to speedup their work - i know nondev project managers, financial advisors, marketers and lawyers rolling their own mini apps now.

Also even if your Gemini is giving you nonprogramming output, underneath the model is most likely generating code for certain tasks.


> How many programmers really exist out there? Surely it can’t support the weight of investment that exists in AI already. It’s just such a small pool of the human race

I dont consider myself a programmer but use LLMs almost exclusively for coding.

The number of people able to create useful software today is much much larger than it used to be and arguably a minority of these people are/were "programmers"


Why did this get downvotes? :o

Free Gemini has at least one subtle mistake on every scientific or programming task I ask it. It's usually in the right ballpark but I guess I do stuff outside its training locus.

Work for my company. Research, code, analysis.

How have you spent hundreds of dollars? I’ve only spent 11 and I’ve been using it for four months!

claude 5.5 with price discounts is a step up in terms of quotas, as well as gpt 6.1 sol, i still feel like for complex things claude is way better vs anything else. also for the topic started, opus 5.5 and sonnet 5.5 are much faster too. I was using grok for that reasons just to do faster changes, but now i'm considering to cancel my cursor subscription, clade is much better and it's also as fast as anything else especially with workflows.

Note that $400 a month is a typical car payment. For the same money -- well, OK, probably a little more -- you could be paying off a set of 4 RTX 6000 cards that will run 4.1 Flash all you want, all day and all night, at hundreds of tokens per second.

Ok, but I have two of those cards plugged into fiber internet, and they consistently rent out for $1500+/mo on vast.ai. So your example would be $3000, minus a couple hundred for business internet and electricity… that still buys a LOT of $200/mo frontier model subscriptions.

which year are they from?

I bought them direct from PNY the week the RTX Pro 6000 was released. Prices are almost double now what they were then.

Seriously? People are paying you $1500/month to rent your 2x RTX 6000 Blackwell cards?

Yes, they have been near 100% utilization on vast.ai at $1 to $1.25 depending on demand for the past two months. They're in two separate workstations, which I purchased for around $12k each before GPU and RAM prices went nuts this year.

Now, I have absolutely no clue how long this situation is going to last! But the economics don't really work out for local models while it does.


I think the point was this was the spot price he/she would be paying if he/she had to rent and didn't own them

You can buy 2 of them from Central Computer any day of the week for $30K US, and that's not the lowest price I've heard lately.

So if you can spend $30K and immediately start mining $1500/month out of thin air, that's a pretty nice investment even if the electricity costs $200/month. Two years later the cards will have paid for themselves entirely and (I suspect) will still be pretty useful.

It does argue in favor of just paying OpenAI or Anthropic for tokens, though.


You do need fairly high end workstations to put them in, adequate cooling, power and very solid business internet as well. I've put more work into keeping them online than I had expected. But overall the economics are fantastic for it right now, the big question is how long it will continue.

How important is it that you maintain a good availability record? I have a few RTX 6000s in a server that I use occasionally for NDA'ed contract work and for experimenting with open models, and now I feel stupid for just letting it idle at ~1 kWh when it could be earning revenue. I didn't realize that anyone would care to rent prosumer-grade GPUs.

At the same time, when I need to use the hardware for something, whoever is renting it from me at the moment is going to get unceremoniously booted, and I imagine they are not going to be happy about that. I assume that vast.ai's providers get uptime ratings that drive their work allocation, right?


Yeah - that is the issue. Doing it that way is going to nuke your reliability rating and you will have to cut prices drastically to get rented.

What I do is… rent on the same platform when I actually need to use a card. A benefit is that if I need a burst of more power, that’s not an issue since it’s available from other hosts. But obviously there’s an inflection point of first party utilization where it makes more sense to own.


If you break down the pure profit per work hour you put in (after all the expenses), how much is that?

Electricity costs money.

And there's a lot of labour and effort involved in setting these things up and maintaining them. People don't even run their own email servers, even though the hardware side of that is trivial.


Electricity costs money

Not $400 a month, it doesn't. At least not around here (we average $0.12/kWh and I don't personally run my cards over 300W.)

And there's a lot of labour and effort involved...

Theoretically, the people who hang around HN are more likely than most to be capable of the labour and effort of setting these things up and maintaining them.

People don't even run their own email servers...

People don't run their own email servers because a convenient coalition of spammers, standards bodies, and large email providers have done their best to make running one's own email server almost impossible.


4x RTX 6000 running at full load with relatively cheap electricity ($0.25/kWh) is around $450/month.

I did the math and decided it’d better to pay for tokens than to buy the hardware and generate them myself.


> Theoretically, the people who hang around HN are more likely than most to be capable of the labour and effort of setting these things up and maintaining them.

'Capable' doesn't mean your labour has no opportunity costs.

I think my argument is easier to attack by noting that you can use AI to substitute for much of that labour.


That's what I did personally, not being a Linux guru. But to your point, I still had to spend a lot of time and effort dorking around with the hardware.

It's worth it, knowing that there are no rugs Sam or Dario or anyone else can pull.


> It's worth it, knowing that there are no rugs Sam or Dario or anyone else can pull.

At the moment, it's still very easy to switch from one open weights model to another, and even between the closed models. So the 'no rug pull' property is nice to have, but not as big of a deal.


Power averages $0.30 to $0.45 for a business in the UK if you convert it to dollars. Smaller businesses without load shedding agreements are at the middle / top of that. (there are lots of people complaining about it). Maybe I just have expensive power, but $0.12 per kWh is pretty cheap.

> Theoretically, the people who hang around HN are more likely than most to be capable of the labour and effort of setting these things up and maintaining them.

Theoretically, the people who hang around HN have immense opportunity cost when doing this. Being capable does not mean it takes no time. Time they could be using for something more profitable (and fun?).


That's one way to look at it. Another way to look at it is to point out that the potential cost of delegating cognitive resources to companies like OpenAI and Anthropic is unbounded.

The AI models we currently have still maintain their working state with a finite and laughably-small context window, but that is already starting to change. My .claude directory contains over 300 .md files that I didn't put there myself. Their contents are eye-opening. The question of who owns, stores, maintains, and can access that data is going to become insanely important over the next couple of years.

If you thought LLMs themselves were disruptive and contentious, just wait until the fight over object permanence gets under way. That's when owning your own box full of graphics cards is going to become important. My own bet is that I won't care too much about the electric bill or my opportunity cost when we all find out what the AI labs really have in mind, and what they're going to have to do in order to justify the valuations they're seeking.

TL,DR: it's not about the tokens, IMHO.


Doesn't a RTX 6000 run at about $100/month for electricity? Do that x4 and I don't see where your savings would be coming from.

I think you may have a problem with caching, can you check how much cache % you have?

Whats your harness?


> if you don't mind feeding your data to the Meta machine (spoiler: I won't).

How is that different from feeding into the OpenAI or Anthropic machines?


same, used a wrapper around cc and i was spending up to $30 a day with basic stuff

Maybe CC does something that breaks the cache? I cannot recommend Oh My Pi enough. Every default is galaxy brained, and it plays incredibly well with deepseek flash 4.1. My favorite coding harness rn for sure.

I found omp used quite a bit more tokens than my fairly basic pi setup... but most of those tokens would be cached with DS V4.1, so maybe worth if there are gains elsewhere.



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: