HN Simulatornew | past | comments | lists | submit | zygentoma's commentslogin

I believe it's a reference to the hitchhikers guide to the galaxy – where the plans to remove the protagonists building to build a bypass road was hidden in this way.

Sorry, no.

When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.

Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.


And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?

Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.


Dealing with potentially-hostile hosts is quite common, actually. See for example how most Linux mirrors work, or Subresource Integrity with HTML.

Turns out securing a service to transfer a single hash is a lot easier than securing a service to transfer gigabytes of data.

Even if I don't fully trust Github, it is still incredibly convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn't compromised, that should be secure.


Similarly, if you push to github, then trigger CI through something other than github actions, then (other than the SHA-1 problem), you have reasonable assurances that someone who has compromised GH cannot compromise your CI host.

Note that the US CLOUD Act means that, if someone figures out how to actually use collisions to compromise that CI machine, then, if the US government asks Microsoft to do use that vector to break into an overseas machine, then Microsoft will be legally obligated to do it.


Code signing also generally relies on hashes. You don’t encrypt the code to make a signature, you encrypt a fingerprint of the code, which is usually a hash from the SHA family.

I pushed aviation toward starting with SHA-256 instead of SHA-1 for code signing more than fifteen years ago, just after NIST first started discouraging the use of SHA-1 in new code.


Rumor has it that GitHub has a flat namespace for commits. They don't store "user1/repo1/abcd1234" in one file and "user2/repo2/abcd1234" in another. Both references point to the same commit in a global shared space. If the hashes are truly unique, then that never matters, because the odds are approximately 0.000000000...000 of you and I accidentally generating the same commit. However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes because the odds are infinitesimal that it'd ever matter, than voila, I've updated your repo by writing to my own.

Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.

I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.


So, I haven't worked at GitHub in some time, but we never had a flat namespace for objects. There are a lot of SHAs in the DB but objects are always namespaced by repository. Forks shared an object database for efficiency, and have technically added reachable objects to a shared database via fork, but it's never been a real problem afaik.

But the very wrong assumption here is: "if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes"

All parts of this are incorrect.

You can push a colliding commit to a fork, but Git will see that it's already there and ignore it - first write does win. Also, "backend doesn't check uniqueness" is also wrong. The server will check for collisions and if this particular case happens, the server will see this and warn you _AND_ not write the object.


> I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is.

All the details on how GitHub's infrastructure has evolved over the years are very publicly detailed in the GitHub engineering blog and in technical talks. There are no "secrets" or "rumors" here, all the information is one google search away. Perhaps you need to re-evaluate your priors on how knowledgeable your friends are.

- https://github.blog/engineering/architecture-optimization/in... - https://www.youtube.com/watch?v=Ri8hSZNKzu4 - https://github.blog/engineering/building-resilience-in-spoke... - https://github.blog/open-source/git/counting-objects/ - https://www.youtube.com/watch?v=DY0yNRNkYb0 - https://cursor.com/blog/git-at-any-scale


> However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit...

You cannot do that. This is called a second-preimage attack, which has not been demonstrated even against MD5, let alone SHA1.


I'd expect an LLM to prove this (if true) in ~ 60 minutes, given just your post and "try to prove this true or false; here's a github PAT".

It is already known to be true for repos that are forks or have been forked: https://trufflesecurity.com/blog/anyone-can-access-deleted-a...

That's per-repo. The claim is that user/repo/commit is flattened to just the commit. Your counterpoint (which I agree with and is well-known) is that it's flattened to repo/commit (dropping the user).

Ah right, though I think it is rather unlikely Github flattens across multiple repos given that Github has already blogged about how their repo storage backend (Spokes) works, and it seems to operate entirely on whole repositories.

Yeah, I also don't believe it.

I also don't believe it.

> And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

Yes, they ought to be good enough. That has always been git's security model.

Note that the commit doesn't have to be communicated over the same channel as the git data.

> Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?

Then they're vulnerable. But what about somebody learning about the trusted hash in another way, e.g. a build server getting an internal call authenticated by an authorized developer?

Just because you can think of an insecure way to use git hashes doesn't mean there aren't any other, secure ones.


You are not saying anything other than that hashes are useful for signing. When you communicate to someone that you trust such and such a hash, and they should too, you are informally certifying the hash, the cryptographically formal way to do which is to sign it.

> You are not saying anything other than that hashes are useful for signing.

Yes, but you seem to be saying they are not, or maybe just that git hashes are somehow inherently not capable of taking that role, unless I'm misunderstanding.


> And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?

How does this matter? When I have a machine that I trust and a git hash that I trust, I don't need to rely on the transport mechanism. As long as the content cannot be forged to match the hash, the transport is completely irrelevant.


> And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries

Wait, why is anyone expecting that to be a good idea at all?


I mean, Git commit signing should be used more often... then you can actually trust the person signing, not the distribution method

But you still need SHA256 for that


Sorry, no. For trusting code there is code signing. A sha-1 hash is cryptographically the wrong approach for it or we wouldn't have RSA and ECDSA algos.

And what do these cryptographic signing algorithms do? They sign a cryptographic hash of the data … Which the SHA-family of hashes are.

If you rely on the commit SHA-1 as a integrity verification it's your problem. Git was never intended to be used as an integrity check.

Git was maybe never intended to be used as an integrity check, but SHA hashes definitely were and are.

And if git provides a cryptographic hash over the content, I don't see why it shouldn't be used to verify the integrity of the checked-out content.


The reason why not is very simple: the SHA-1 algorithm was chosen because it made it practically impossible for two commits to have the same hash, not on the assumption that it would remain forever immune to collision attacks.

A vulnerability can be found in a hash algorithm at any time, whereas changing git's hashing algorithm is inevitably a slow process. IMO, if you want to ensure that someone gets an exact set of files, zip it up and sign the archive with the algorithm of your choice. It's not necessarily going to be practical to change git's signing algorithm every time a security issue is found.


I think perhaps you used the wrong words there. The hash is an integrity check but it was never intended as a security check.

>I expect the code to be exactly what has been reviewed by me under that hash

Sorry, no.

Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.

> Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.

What? Uncircumventable? Logically equivalent, I read your statement as "So if all cars are not blue then they must be red"? This does not follow...

I like the last part of the article, which proposes a (very reasonable sounding) extension for people who care (more) about their code-sec. You should be able to swap out your hash algo without having to rebuild your content addressing system.

Separation of Concerns, people...


> Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.

This doesn't seem to be a useful definition. Would you classify every computable algorithm as insecure, because by generating a random bitstring, there is a (very, very low) probability of guessing the hash/secret key/solution/signature?


> Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.

Cool, you've just defined the foundation of signatures and web encryption as insecure. What next?

Also, you can only be so certain about any piece of data no matter what you do. With a non-broken hash you can make the collision chance be a trillion times lower than the chance you're hashing the wrong data to begin with. That's as good as gold, well actually better than gold.


Sorry, but aren't there signing and attestation mechanisms built into git already? Why not use those?

This also seems to happen for 16 and 32 bit numbers, so you should be able to see zeros easily.

They also write:

> Running the same programs on an Intel processor, and the 0's are there with no problem.


Radiation pressure! There is so much heat (= photons) radiating outwards, that it counteracts the gravitational pull.


What’s particularly interesting to me is that for stars the size of our sun, regular old gas pressure dominates. The sun’s atmosphere is held up essentially just from the temperature (and therefore high kinetic energy) of the plasma.

Only once you get to 10+ solar masses does radiation (light) pressure begin to become significant, and at 50+ solar masses is when it dominates and the atmosphere is held up by the momentum of light.


In most star photospheres the role of radiation pressure is negligible. They are supported by the pressure gradients. The exception is very hot stars.


Wow I knew that stars have that radiation pressure but had no idea it caused mass to get so crazy far away from the ignited area of the star!


These giant stars burn so very very bright. And correspondingly only live a few tens of millions years at most.


It's not really like you're 5, but the third sentence of the introduction makes it really understandable:

> The problem’s definition is simple: There are k servers located at points of a metric space. At each time step, a request arrives at a point of the metric space. An online algorithm must serve the request immediately by moving a server to the requested location, without knowledge of future requests. The goal is to minimize the total distance traveled by servers.

So metric space is anything where you can measure a distance, so you know the distances between all servers and the distance from the request to all servers. Could be direct distance, could be travel time …

Easiest to just imagine just some (eg. n=5) servers on a plane. A request pops up somewhere on the plane. Which server do you move there, such that the total distance moved by servers is as low as possible in the end after a sequence of requests.


Fair enough, but the abstract really is too opaque imo.

> The k-server conjecture states that a deterministic online algorithm can achieve competitive ratio k on every metric space.

I was like "algorithm for what?" when I knew what every one of those terms meant. Lots of things use k servers. Lots of problems involve metric spaces.


Well, any natural number is unimaginably small, compared to ω …


Shouldn't a fixed additional amount also care for inflation based price increases indefinitely?

(Maybe not if the inflation greater or equal the interest rate … though I did not do my math here.)


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: