>> with CNC and robots we can achieve far superior results for a fraction of the cost.
> This is only true for mass production, where the set-up costs are spread across thousands of units. For small batches, the cost of automation often exceeds the cost of the batch itself. CNC machines are not magic; they require qualified and competent staff.
This isn't right.
CNC machines aren't magic and they do require skilled people to design for them and operate them.
BUT: they are an enormous advance on capability compared to 30 years ago and are absolutely not limited to mass production. 3D CNC milling is how most prototyping of metal parts is done and can be done to extraordinary precision, repeatably. It was always a low volume technology - it was Apple that started applying it in volume as part of mass manufacture.
While not metalwork, some are surprised to learn how old CNC machined guitars are. I have a Peavey T-60 and they were among the first CNC-machined production guitars in the late 1970s. It's going strong to this day as well, they're very well-built instruments.
The problem is that it ISN'T a fee for services as it should be. Their platform power means that they can levy a tax, and they don't feel the need to do anything on exchange for that.
Yep, I'd hypothetically economically gladly build an alternative app store charging "only" 5% on app revenue. The blocker would be that Google has a hold on what store is installed by default on Android. Apple even more so. This is hardly a fair market.
There are other app stores on Android, e.g. OEM ones like the Galaxy store. The latter doesn't really compete on price though - 20% cut of one time payments and 15% of subscriptions. I think the Android app stores anchor to the iOS app store prices to piggy back on Apple's monopoly power - devs are used to having to pay Apple that much, so they will not resist if you charge the same.
All of the alternative stores are systematically disadvantaged time and time again due to Google's and Apple's near-absokute control over their respective OS. They are as anti-competitive as they are allowed to be. And the US seems incapable or unwilling to rein them in.
Google is not developing Android for the IAP fees, they're doing it to sell phones loaded with their search, advertising, browser and services. They take hundreds of billions in revenue annually from this, they are being paid several times over for that work even if they get no IAP fees.
The speed of government is slow, incredibly slow. It's measured in election cycles. It takes government eight, twelve years to simply turn around and notice. But once it does, it's a juggernaut.
Nothing can stop it. Nothing can avert it. Nothing can prevent it. And as everybody from Facebook to Microsoft knows, once it starts moving, you're screwed.
In the US, the oil companies thought they were too big. Then the telco companies thought they were too big. What really happens is eventually people just get pissed off, and then it's too late.
Google is fast approaching that. They almost had chrome and other things ripped away from them, and now they've got requirements to have alternate play stores. If they don't turn course, and turn course blazingly fast, they're going to have their entire empire wrenched away from them.
I think it's best that government has a lot of inertia in general, but maybe that doesn't preclude government having sub components that can act fast within defined domains.
Like you don't want to have fundamentals and pronciples change hardly ever. So you might take a long time to create or destroy say the FDA, but the FDA could be responsible for being on top of time-sensitive FDA things.
Maybe there could even be a body whos job is only to decide if something is time -sensitive or not, allowing other bodies to move as they need to, yet not just have no limits or accountability at all even within their own domains. So a rapid action can be done, then shortly after determined it was a bad idea and undone, or partly wrong and needs adjustment, and have that all be justified by a seperate independant process that judged or agreed "this is something that needed rapid action".
I don't think we have anything so organized yet. We do have some things where a body is responsible for a domain and can act more quickly within their scope, but we also have a lot of things where we get both forms of bad, making big changes on immediate whims, and being unable to make clearly needed even small changes simply because "government is slow".
We do have some examples of a reasonable organization and heirarchy and layers, and a whole lot of not, all together at the same time.
> If they don't turn course, and turn course blazingly fast, they're going to have their entire empire wrenched away from them.
Unfortunately, that's no longer the case with the app store monopolies. While Apple is under some partial enforcement on parts of their app store policies, Google managed to play the game of regulatory brinksmanship better and largely avoided that for their app store. So the regulatory powers (such as they are) have already taken their best shot and this is where we are.
There is now a record of this. The next action against Google, will see people pointing out each step they took to avoid compliance.
A case a year down the road about AI, will see stubborn attempts to comply like the above cited.
Google, with the gold plated shovel, digging a hole for itself. Put another way, breaking up a company typically comes after other methods don't work. Such as being sneaky and making them not work.
Or even because the required competition didn't occur, regardless of fault.
It's not over. It's a step taken, with a positive outcome, and a judge who is making sure there's follow through. Next step will be more stark.
They also say they generate $57 billion ad revenue from being default search engine on iPhone, so obviously there is a cash value to being the default search (and everything else) on Android too even if their accountants can finagle reasons that doesn't contribute to Android's development costs.
If they didn't develop Android they'd be paying $10s of billions annually for the prime positioning their apps and services and search enjoy too, so even if you buy their accounting tricks the development is paying for itself many times over.
If Android did belong to someone else then there could be an argument that IAP fees are funding it and subsequently required, perhaps even under less-than-competitive circumstances.
But if the court allowed, Google would be paying that company $10s of billions for the prime real estate their software, search and services enjoy and it would then be a very weak argument that only IAP funds its development. Just like it is when Google earns $100s of billions in revenue from Android now.
Would it be OK for Microsoft to prevent Steam from running on Windows because Valve doesn't contribute enough to funding the Experiences and Devices org? I've never really understood why mobile platforms should be different from how we think about desktop ones.
Parasitic? Writing software for folks' devices and selling it isn't parasitic...it's just participating in the market. Weird that the overton window has moved so far that anything not built by the trillion-dollar duopoly is considered some kind of free-rider.
That trillion dollar duopoly is probably the worst case of free riders. It’s an enabler/vehicle for all the other billion dollar free riders to steal from the population.
It's not weird - you've just completely misunderstood the situation. This is about the OS, which the software in question (the alternate app store) relies on, but generates no revenue for. The devices are still being sold, because their manufacturers are allowed to make money, the alternate store is allowed to make money. Only the OS maker will make no money, because that would be bad.
what you are arguing for is objectively extremely bizarre. if you take a step back, this is what is playing out: company C makes a device. it needs an operating system to run software, so they add it. if there was no operating system the device would be useless and no one would buy it. now person A buys said device. it's theirs, they have paid money for it, they own it. person B writes some software. it runs on the device's operating system. they want to sell it to person A. person A wants to buy it from them. why should company C be involved at all? I found this super weird and distasteful when game consoles were allowed to use this business model, and am even more taken aback that general purpose computers get away with it.
You seem to imply Google was forced to open source Android or even to give it for free: sounded like it was a purely business decision to capture more of the market.
Now that they have, they are also turning around on bits of that too.
theres nothing parasitic about a competing app store.
it is parasitic to try and prohibit competing app stores and then leverage your monopoly by charging monopoly prices.
everything google is trying to do to discourage graphene, newpipe, competing app stores, etc. is parasitic and it's right that they are being fined by the EU for this behavior.
> theres nothing parasitic about a competing app store.
I think a more honest way to say this might be: Android OS development is massively funded by Google, and the Play Store feeds into this hugely. No other store does this, so it shouldn't be surprising that they can charge a lower fee. They're just doing the easy bit.
It is rent extraction which always feels like bullshit because it is. Competition is supposed to be the counterbalancing force but it barely exists here.
The basic services Google provides can effectively only be consumed by humans individually, creating an inevitable bottleneck.
Such bottlenecks impede interchangeability of goods, as you can't reasonably entertain multiple suppliers simultaneously there. And that means, you cannot have a free market, because you can't really have meaningful competition.
Google has captured services that really are public utilities. Them being "new" doesn't play a role.
I (as someone British and from an Army family) read that as - "You know you screwed up right? Please reassure me you know that and won't do anything like it again". Followed by an admission that we all screw up sometimes (implied by the stories).
Makes it easy to switch between models and I like it for exactly the reason that you're saying - I prepay and so can't accidentally spend my food budget.
It is absolutely not worth buying hardware to run models for purely (long term) cost reasons. For open weights models the economies of scale means the cloud beats local significantly and your payback time is like 10 years.
However there are other reasons (e.g. privacy) that might make it worth running locally for some people.
> I think the biggest reason is to own the stack so your model can't be changed out from under you,
The concern would be future regulations that prohibit you from buying a hosted version of the model. Even that could be bypassed with a VPN to another country but it's more work to go through the payments.
As long as there is demand for a model, it will be hosted by multiple providers.
I live in a place where using VPN is illegal and akin to "terrorism" because why would you want to hide what you are doing. Only bad guys hide. So if you use VPN, you are a bad guy.
“Out of the 15 individuals identified, five were minors who were counselled and advised in the presence of their guardians, with emphasis on awareness, lawful digital conduct, and the consequences of violating lawful orders,” he added.
What if the model is hopelessly obsolete, and thus no demand, but I want that specific model? Owning the weights and hardware is not just solving for one problem. It eliminates all the classes of problems that occur outside of your building, if you have a solar and battery setup.
Also, on a more practical basis, what if the way it's served is bad. Maybe I want my specific KV setup, or ultra low quant for entertaining garbage at 200 tk/s
> What if the model is hopelessly obsolete, and thus no demand, but I want that specific model?
You can still find a lot of old and completely outdated models on OpenRouter. The providers can scale serving of models up and down as demand arrives, so models don't generally disappear. They're just kept in the mix and the clouds will allocate hardware to it if someone is willing to pay.
In the odd case that it disappears completely, buying the hardware 2 years from now is probably going to be a better deal. That wasn't true if you selectively check the time period before hardware got expensive, but as new hardware comes out we're going to start seeing Strix Halo and old Apple hardware hit the market as people upgrade. It's already happening.
There is a certain personality type that cannot tolerate any uncertainty and must lock everything in right now against all future possibilities. If you fit that description then there's nothing anyone can say to discourage you from buying your own hardware, but for everyone else I do not recommend buying hardware to self-host LLMs just to save money. I self-host and run a lot of tokens through my setup (non-coding work) but I'm not really saving money.
> There is a certain personality type that cannot tolerate any uncertainty and must lock everything in right now against all future possibilities.
I thought HN banned personal attacks. I'm in this sentence and I don't like it. /s
I just buy the good apple hardware because it's good, and it also happens to run local models. It's not as good for the dollar, don't get me wrong, but I'm not going to develop iOS without a mac, that's even more questionable than buying a strix or whatever.
Also, this makes me wonder if, by using a bicycle generator, and a local model at sufficiently low power consumption, you could directly claim to have produced the text in a really physical way. "Yes, I generated the electrons that made that text work by my own efforts".
You do, there's like 20 providers for any model on openrouter. You can also just spin bedrock or gcp and download the weights for later if you're worried. It's never going to make cost sense when the token rate is so low with how expensive ram is
I think that's overly pessimistic. Here's [1] a video of somebody running it on a ~$6000 rig and getting around 14T/s for complex prompts (about double that for simpler prompts). Payback time is going to depend on your electric cost/consumption. In most domains cloud providers end up charging a significant premium rather than a offering a scale enabled discount, relative to local at retail costs. That will almost certainly end up being the case with LLMs as well, if it isn't already.
Furthermore we continue to follow the path that image gen neural networks took. In that domain hardware requirements reached a peak and then started sharply declining to where we are today where a plain old video card can rapidly generate images that took a supercomputer not that long ago. So it's reasonable to assume that performance of such a system could potentially even increase over time.
With roughly 2.7 million seconds per month, times 14 tokens per second, you are getting 38.5 million tokens a month at most.
That’s less than 164USD worth of GLM5.3 tokens on the inference market. So that 6000 USD rig will take 3 years to break even - and only if it runs continuously.
And this is being generous, as it’s not even taking quantisation into account.
I think the “killer app” is doing inference without sending the data to China or the US. At home it’s overkill but imagine you are an EU consultancy with a lot of client data to work on, or a company/institution with a lot of sensitive data, buying the hardware to make sure the data stays private is a huge benefit. So is that you “own” the model. Its capabilities, price or access don’t change at someone else’s whim.
Some of that is that EU providers need to up their game here.
Needing an EU native option is really the one and only reasonably objection I've heard against using LLMs from the cloud, the rest is tin-foil hat level unless you're actually intending to meddle with the inference or fine tuning or something beyond just querying.
I think if you steel-man what I'm saying, what you're saying falls apart. 14 tokens per second was rare. It only dropped that low in one scenario where he had it single shot an entire game (flappy bird clone) from scratch, with different assets, all self created, and so on. It ended up resulting in the LLM doing stuff like plotting out a some odd 100 item long to-do list, requerying it repeatedly, and so on. And it succeeded.
Also as the video mentions, the guy wasn't very familiar with what he was doing, and so there are almost certainly various optimizations on the config side he could work out, especially as he was using a 5 GPU system, which default configs are probably not well optimized for.
But I think we've rapidly moving along the same path as image gen stuff. Local generation has gone from purely theoretic, to requiring supercomputers to run relatively incapable models, to where we are today - where with a fairly basic high end setup, he's comfortably running a frontier level model. There's definitely an argument for going local that's only growing stronger by the day.
I agree it’s probably not representative token speed. But I do believe the overall observation holds: The monetary value of local inference is bound by the wall clock.
I agree that there are many other reasons than cost alone.
I'm actively uninspired to write high quality code when using Anthropic/OpenAI models given the high chance I'm a customer as well as used as dataset generation tool for them.
But currently cloud does beat costs of hardware ownership, particularly with ridiculously high RAM/GPU/SSD costs....again due to these same companies.
> It is absolutely not worth buying hardware to run models for purely (long term) cost reasons
This is especially true when it's trivial to have the LLM itself write you a script/tool that can rent a GPU node for you (via API calls to providers) and then download and set up an open weight model for you.
I mean, I think it depends. At home 3 of us we use AI for multiple reasons, from coding apps to asking general questions, and if we would have to pay equivalent subscriptions that would be ~1k a year on AI + submitting all your data to external services. I payed around ~8k on 2 DGX Sparks that, at the moment, serves perfectly fine as a ChatGPT/Claude replacement at home (DS4 Flash peaking at ~170 tokens per sec with 6 concurrent sequences), and even once the technology is obsolete for inference in a few years, I will still have 2 pretty powerful machines for whatever I need + some pretty fast NVME Storage. I don't think its a terribly bad idea.
I think the privacy argument that keeps coming up is overrepresented. Certainly ZDR is enough for an absolute majority of use cases? I see so much talk about local inference but I doubt most of it has privacy as a valid argument (not arguing it doesn't exist). It's fun to do things locally though. I've tried it as well but cloud is just faster and cheaper.
These companies have displayed zero respect for everyone's intellectual property getting these models trained.
I think not giving them your complete trust is reasonable! I'm not saying zero trust, and ZDR is fine for most things but I understand the people who don't want to stream their whole codebase out token by token.
Then use other providers hosting open models. Companies and individuals already put their whole code base on the cloud. I'm genuinely interested in privacy-oriented use cases where ZDR is not enough.
ZDR is built on trust. Given that end-to-end encryption fundamentally doesn't work with LLMs, as they need the content to be unencrypted to operate on it[1], you have no way to prove that once your plaintext data is on somebody else's server they aren't doing whatever the hell they please with it. All you have to rely on is their pinky promise that they won't do anything with it. Trust is a valid option, much of our society runs on trust, but you can eliminate the need for trust whatsoever by running on your own hardware.
[1] Yes, I'm aware of experiments to operate on encrypted prompts, but these are only research attempts, not something that could actually be used with frontier models in production.
I'm not that worried about the codebase itself. I'm worried about the fact coding agents poke around the terminal and system so much that there is almost a certainty that some of your other personal data ends up in the context somewhere which is getting logged in to a training dataset by random hosting providers.
Privacy isn’t only, I don’t want anyone to have access to my data. It could also be, I don’t want anyone to know my use case because it’s niche and highly profitable.
Counterpoint: you can't just go to Visa/Mastercard or a merchant acquirer out there and set up an account on the same terms that Stripe can.
On the other hand, you can sign up to any LLM provider and get API access on terms that are the same or better (since I'm sure they don't appreciate having a middleman and would benefit from incentivizing direct usage) than OpenRouter gets.
To add to this; payment infrastructure requires a lot of heavy lifting. There's a lot of regulations you need to adhere to, different payment systems in different countries, settlement, chargebacks, etc.
There is a reason why not doing your own payment processing is a thing.
OpenRouter may have some interesting things in streamlining the process of switching LLM providers, but it is indeed something easy to replicate in comparison to payment processing.
Yes, you can get better pricing if you do it yourself. But the true advantage of OpenRouter is that, in a space where there's a new model being released every week, you can easily switch to whatever model is best at any given time without having to set up accounts with multiple providers. Or you can just experiment with the latest release. Their product is the convenience. Of course if you decide to only use a specific provider or two, then you don't need OpenRouter.
But it's not really that difficult to make a clone of OpenRouter's service. What they do isn't really that original.
Their only value comes from the fact that the currently have lots of traffic. And I dkn't think that their cumstomers are really bound to theur servuce. They could switch to a competitor without too much hassle.
Yes, but as l9ng as there's no clone, they're good. On the potential clone side, I guess many will be put off by the fact that "there's already OpenRouter". Also, a clone would need to find a way to make existing OpenRouter customers to switch, which isn't easy.
Openrouter is fragile, one good competitor and they are at risk.
Or if one of the provider suddenly decide to forbid openrouter from using their api. Why would they do that tho ?? Well it's not like execs never take dumb decisions.
In my humble opinion, it's extremely overpriced. But then, it's just the standard with AI currently. Divide every ai company valuation by 1000 and you might get it's real value.
Now that's the value. OpenRouter has the power to step on the air hose of any provider they don't like.
That's Google's real power. Works for them. Between search, ads, and the "app store", they can crush most companies. OpenRouter's power is only in one area. For now.
Stripe operates in a very regulated industry where new entrance is difficult
OpenRouter is commodity stuff, I've never used it personally and picked alternatives. The sense I got from people who said they used it was that they are on average penny pinchers. That does not seem like an ideal user base
Having worked on web apps that processed online payments before and after Stripe I totally agree, this was an area of real pain that became suddenly extremely simple because of Stripe. The alternatives were terrible–100 page Word docs of SOAP API docs for Authorize.net, massive PCI compliance requirement specifications, horrible legacy merchant services businesses.
On the other hand, I operate an app that talks to (and logs prompts/meters costs) to many different LLM API providers, and I do not consider it painful at all. I have an AI agent to deal with any integration quirks, if needed. Mostly they provide OpenAI-compatible APIs anyhow. It's basically a no-brainer to go direct with the providers and save 5%, the great majority of the cost of an AI-powered app is no longer dev time implementing the integration, it's the tokens themselves.
> This is only true for mass production, where the set-up costs are spread across thousands of units. For small batches, the cost of automation often exceeds the cost of the batch itself. CNC machines are not magic; they require qualified and competent staff.
This isn't right.
CNC machines aren't magic and they do require skilled people to design for them and operate them.
BUT: they are an enormous advance on capability compared to 30 years ago and are absolutely not limited to mass production. 3D CNC milling is how most prototyping of metal parts is done and can be done to extraordinary precision, repeatably. It was always a low volume technology - it was Apple that started applying it in volume as part of mass manufacture.
reply