though this is exactly the thing that can be handled during RL deep learning
the problem is that introduction of any new technology usually happens with so much emotional baggage, that when there's an error (human or otherwise) some humans will understandably see their biases confirmed in them, and will signal boost everything to the Moon.
what does the license say? provide source when asked or publish source?
OSI is not clear on this either. ("Where some form of a product is not distributed with source code, there must be a well-publicized means of obtaining the source code for no more than a reasonable reproduction cost, [...]")
Google became one of those OEMs that drag their feet and make it a hassle to publish the source code. This happens because they are not interested in open source Android anymore
Alas, the licences have become outdated and we haven't kept up, because most have played by the spirit of the rules up to this point. There is an expectation that source code will be readily available using modern tooling, like a version control system, but that's not what the licences actually require.
GPL requires that source code be in the preferred form for making modifications. If you use source control for your own modifications, then that's obviously the preferred form and you must distribute the source in that form.
That’s an ambitious interpretation. The preferred form for making modifications to the Linux kernel also involves having a build cache, but that’s clearly not required to remain GPL compliant.
I hate to be that guy but publishing source online for free theoretically involves unbounded distribution costs that can't be recovered. That isn't the issue here but it should be considered before changing a license.
Google has servers inside ISP networks that distribute YouTube videos. They pay nothing for the distribution costs, I think they don’t even pay for electricity/cooling. If they wanted to distribute Android source code, they wouldn’t have to pay a dime.
Regardless of whether these costs are zero (I doubt they actually are), these licenses apply to EVERYONE who uses them, not just very wealthy companies. It would not be right to essentially require unbounded free hosting or use of specific online tools/platforms, just to be able to comply with the license.
Let not forget that google heavily relies on opensource software.
The Linux kernel being the obvious one in this discussion, which they make extensive use of.
It's a generic license that doesn't say "by the way, if you've got a market cap above X then you have to host and distribute the code FOR FREE"... Putting such garbage in there would hurt small developers who made popular, nonprofit projects and probably anyone who isn't flush with cash. Imagine being legally obligated to host files in the middle of a DDoS attack, and pay the resulting AWS bill or whatever.
Not from my perspective. I get free mapping anywhere on earth, with traffic information, and can route plan walking, cycling, driving, or public/private transport and likely journey times. I can search the world's web-published information. I'm using a powerful, relatively affordable mobile phone. I can watch anything anyone with a camera uploads to YouTube across the world. I can pay from my phone with a thumbprint verification only. I can take a photo of something and get back what it is and some information about it.
Yeah, to make things spicy, what we ship is a plugin, so there's limit of how static we can link. glibc stays dynamic and that's a huge limiting factor.
I've got plans to try to build on RockyLinux9 and package its glibc along with the app. Simple helloworld works, so I have a glimpse of hope that it would be possible to ship like that.
LLM-assisted software engineering seems to be very efficient if you have a deterministic target. (bun rewrite from zig to Rust, 100% Node.js compatibility, pnpm compatibility -- https://github.com/oven-sh/bun/pull/38333)
Problem with that is that it's not proper rust code, it's some kind of franken zig style rust port.
The actual hard part is getting it into idiomatic, safe rust, and I don't believe the LLM port makes the full transition any easier than doing things the old fashioned way: a dual lang code base like Linux.
And compiler generated assembly code is ugly and spaghetti, unlike beautiful hand-crafted human written assembly code, the kind you see in ffmpeg codecs, ...
It does matter, because that's the entire point of the language. You know, memory safety?
The vibe coded rust port did not get them any closer to full compiler verified memory safety. They still need to go file by file, bit by bit, and make it memory safe, at which point, why not just do that from zig.
I didn't say don't use LLMs. In fact, I'm quite certain they could help quite a bit.
But I don't think they did it in a way that actually provided a meaningful gain. My point is that the state of the codebase right after the code is just a worse version of the original zig code with few of the rust advantages. Going from that Frankenstein rust code to actual memory safe rust is a similar leap to going from zig.
> I don't believe the LLM port makes the full transition any easier than doing things the old fashioned way: a dual lang code base like Linux
...
how so?
let's say Linus opens a branch, rust-temp, and in 2 weeks pushes ~15 million lines of code deleting most of the old C code. and then it gets merged in a few weeks. and then there's still a few months of the "merge window" and RC process. and each day folks report bugs, and automatic fuzzers make sure that both versions "behave the same".
I was pretty skeptical (still am), but what the Bun team did is pretty great so far.
https://bun.com/blog/bun-in-rust details the process, we see the results (Node.js compat [0], more than 3 thousand of issues fixed since since 1.3 [1], more than 900 issues fixed in ~24 hours [2], and it's live/in-prod [3])
> 50 dynamic workflows in Claude Code run continuously over the course of 11 days.
so let's say 500 workflows could do something similar for the kernel. the reported cost is 165K USD, so let's say this would cost 1.65M USD, plus CI costs [4] (which is ~200K/year for bun, so let's say it's 2M USD for the kernel)
How much the world is spending on kernel bug bounties and various security programs each year? (rough estimate says that just the visible kernel testing programs cost at least 15-30M / year)
The bar is not perfect. The bar is something better.
Because they are still dealing with memory safety issues to this day. They ported it to rust, completely introduced a bunch of bugs, went through all that effort, just for the codebase to be not much safer.
I'm actually basing my opinion off the blog post, specifically the code. While they do get a few changes for free with the port, most of the code looks identical. And the number of unsafe statements certainly agrees.
And now they're going little by little, playing whack a mole with seg faults... My point is that the step they're at right here is the important part, and that mechanically converting the entire codebase to rust was an optional step when incrementally rewriting would've worked just as well if not better.
Being "only" 165k USD doesn't mean anything, because you didn't do a complete port, you did a port in a trench coat. If they went for the more targeted module by module approach, and did a clean and proper rewrite, LLM assisted or not, they would reach their end goal much less turbulently and without the effort of that initial frankenport.
are they dealing with more or less memory safety bugs?
for me it was already "don't run in prod" quality before. (at least after this I'm considering adding it to the CI to see how it fares.)
segfaults. I don't know. I looked at their CI and GH issues over the past few weeks. (though now GH is down so I can't do a search, but I didn't see thousands of segfaults.)
by all accounts and measures it seems it made their house of cards more manageable. despite the frankenport, no?
they already had a zig compiler fork, wanted to upstream it, but the zig maintainer(s) said it's low-effort. now they don't have to maintain their compiler fork. (or wait for the zig team to deliver the features they wish for.) no need to maintain a hybrid codebase. (though it has C++ because of the embedded JSC.)
also I have no idea what's the zig-rust FFI status, but getting over with a rewrite faster is usually better, even if you are left with non-idiomatic code.
the neat part is that you don't have to, millions will do it for you!
the point is that it's a quite obvious opportunity that was just some pipedream years ago.
by the time v8.0 comes around we'll have more data on the bun rewrite. also we'll see how LLMs will affect the productivity of kernel developers, and the overall stability/maintainability.
Arguably the question is whether it's economically feasible to self-host something similar.
Are you self-hosting Google or Bing? No, but we have quite a huge ecosystem of full-text search tools with PageRank, with options to scale to almost Google scale (if you have the money). After all LLM training starts with the same crawl mechanism.
As long as barriers to entry is not too high (ie. it makes sense to take the risk to start a business that provides something similar - usually for a niche) market forces work.
We have the classic empirical chart reproducing microeconomics.
And setting up a pharma plant is also very capital intensive.
Here the obvious barrier to entry is completely artificial. (Which provides an incentive to spend a lot of money on R&D -- though it naturally raises the question of Pareto efficiency.)
You're missing the point, what will happen is this:
1) in things like tax law, registering with city hall, dealings with the DMV, your phone subscription, insurance contract, ... you will find that one of the new fine prints in the contract will be that you're not allowed to use AI to communicate with them.
2) because of how SynthID works (you need the SynthID keys to verify, which are secret. So the only way to find if text is ChatGPT/Google/Anthropic watermarked is to ask ChatGPT/Google/Anthropic), government and large companies can enforce this against you. That is what the watermark is for. To end any insurance claim written by AI with "you're not allowed to submit AI written insurance claims" and refuse it outright there and then.
"Sorry your request was AI watermarked and pursuant to law 234 of 2025/03/11 chapter 3258 paragraph 33 decile 1299 we hereby close it without response"
3) when they reply, however, they use a custom model that also has custom SynthID keys. You will not even be able to tell their responses are AI written, or at least, you won't be able to prove it. You won't be able to enforce any AI-related rights (ie. the right to talk to a human) you have under the law against large companies.
In other words: this is to make sure that all the advantages AI provides are available to deny your unemployment claim, and to Verizon to charge you more, but completely inaccessible TO YOU when you want to change to a cheaper subscription. They can inundate YOU with AI-written requests BUT YOU CAN'T.
Self-hosting helps because it prevents them from verifying if your responses are AI written, because you can generate non-watermarked AI text and so there is a level playing field.
For government, because it's in law or regulations (ministerial decisions in Europe). For large companies "You agreed to it" (you know, like you agreed to allow Verizon to sell your location data to Palantir)
The other 2 questions I don't understand. My point is that the EU AI directive makes this possible. Makes it possible in ONE direction, while prohibiting the other. AI can be used by government and large companies to spam you and deal with you, and can't be used by you without being 100% up front about that to them (ie. enabling refusal)
No. Here is the list of organizations that have the power to make laws in the EU (and JUST the across-the-EU part of that list, within countries, within states, within provinces, within towns there's another list). This is referred to in legal tradition as the "Hierarchy of norms", because there is also a clear order defined.
>What's an example of a government adjacent company?
Toll road contractors, military base housing, municipal utility providers, private prisons, auto towing/impound, private rail transit operators, utilities, etc.
Basically every company the government grants a captive market to abuses everyone in that captive market as much as they can get away with.
this seems exactly the usual anti-consumer bullshit that is regulated state-by-state (or sometimes by (lack of) FCC/FTC effort, or by the CFPB that is now a zombie)
however, AI doesn't really influence this. already there's a lot of problem with things like Ticketmaster, Apple's walled garden, abuses of IP law (patent trolls, DMCA trolls), etc.
the insurance industry is a prime example of this. the suffering caused by power imbalance is incomprehensible, and yet there's not enough political will to address this.
sure, it's easily possible that some important aspects of our everyday lives will be worsened by bad AI regulation. but IMHO this is wholly an upstream problem, it's a symptom of bad politics. (a byproduct of the Zip2 to Tesla to "democracy with roman salute characteristics" pipeline.)
that said, obviously the foundation to have any chance of a nonpatological market to exist is that self-hosting has to be legal.
The US bailout paid for itself... If you consider the equivalent of 0.6% growth, not adjusted for inflation, paying for itself. In practice the US's cost of borrowing the money was substantially more than that.
you need to compare the alternative. and also take into account the growth that happened thanks to cheap mortgages.
there was a nominal price bubble (especially visible in the rent vs buy graph) but rents did not show the same spike, there was no general shortage (as seen in vacancy rates), and since interest rates fell nominal prices went up, and the central bank overcorrected.
... the obvious problem is the low productivity of suburbs. spending a lot of money on building them is good for short-term GDP growth, but long-term it's a poor investment compared to the usual recommendations.