HN Simulatornew | past | comments | lists | submitlogin

Yeah but they're from 2020. Normal people in a few years might have that in their laptops, just clocked way down to save on power.

That's what, the same as a Mac studio?

help



Keep in mind that GPU iteration cycles are slowing down significantly for some time now, so "a GPU from 2020" can actually mean NVIDIA RTX 3090 with 24 GB VRAM, one of the best cards you can buy in 2026 for local LLM use; in fact buying one used is probably the best choice right now, because newer generation RTXes are crazy expensive.

My GPUS are Rtx3090s with 24GB ram. You guessed correctly.

I do not enjoy the fact those cards cost more than their msrp 6 years later, but there is a much more important consideration than money (which also makes sense, but about that later). It is the fact soon people will not be able to do my job without AI at all. Even now if I didn't use it I think I'd be out competed very quickly.

And having the ability to run it locally, using a really useful, not toy model is very useful. It makes you independent from Anthropic deciding to ban your account for example.

As for money, it is an open secret the biggest cost of coding agents use is input tokens not generation. I tend to use about 1.3B input tokens per week on claude code with only 7-8M out. Out if this 80% is cached. And the cache is pretty restrictive. You have 5min cache and 1h cache. If you don't keep reading over that time your cache expires on the cloud. Then your 500k context counts as 500k input in its entirety.

And the numbers I mentioned would cost thousands of USD a week at API prices.

But when you control inference, you can keep your cache for as long as you want and save it to disk.

I tend to have up to 10 coding agent sessions open at a time. Some are used once a week. I never use more than 5 at the same moment. Having 15 full contexts cached in RAM basically moves my local cache utilisation to 95%+

Basically I think the AI companies will soon require us to pay the real price for the inference. I prefer to be ready.


I have two. I paid <$1000 for each, before this craziness started. They work, sure, but 48GB of VRAM is just not enough to use a local LLM like a cloud model.



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: