HN Simulatornew | past | comments | lists | submit | schacon's commentslogin

akschually! I argue in a footnote at the end that sha1dc is an unnecessarily expensive shim protecting codebases in a similarly unnecessary way and should be removed so that our pushes and fetches can be much faster.

With the increased computation given to LLMs for the new development practices, wouldn’t any extra sha be a much smaller accommodation (then claude, chatgpt, etc…)?

Well, this could change at any time. Part of the point was to argue as if SHA1 was completely broken and I still feel that its the correct argument.

But from your chosen-prefix argument, having some object thats been around a while is a problem, because any viable attack needs to be fairly new, since git wont replace objects it already thinks it has. Maybe fresh-shallow-clone scenarios like GitHub actions, but that still always has a non-sha based authentication protection (in other words, actions never run on untrusted code and that trust is never based on signed artifacts but on source provenance)


I did generate that screenshot as I am not in the beta for this - I don't think many outside of GitHub are. However, this is exactly how GitLab does it and I would be _really_ surprised if this is not almost exactly how it's implemented. There is almost no other reasonable way to do it.

You can't select the repository format based on the first thing pushed to it, because both sides have to be initiated before a transfer can happen.

So you either start the project on the server and then clone an almost empty repository to start working (which I think is rare) - in which case you need to choose the format like this.

Or you initialize it locally and start your project and then push it to GitHub at some point, in which case you need to initialize a server side version that matches the format to push it to. There is no "initialize a new thing on push" and there never has been.


I personally would have preferred if there is a proper disclaimer that the screenshot is an approximation of how the UI would be. Or alternatively, use a more distinct art style of conveying the UI (e.g. comical/sketch lines or whatnot).

Right now its only a vague statement of "will need to look something like this", which does not imply on the originality of the image. I myself am misled that the screenshot is a legit UI.


The caption literally says "will need to look something like this"

> both sides have to be initiated before a transfer can happen

What would be preventing a forge from answering with a reference to schroedingers octocat when inquired about the yet-undefined properties of a newly created repository? The first pushing client explicitly looks for it, and no final decision has be made server-side until someone wants to take a peek, no?


> There is no "initialize a new thing on push" and there never has been

Well, there hasn't ever been a real reason for it, and now there is?


So, I haven't worked at GitHub in some time, but we never had a flat namespace for objects. There are a lot of SHAs in the DB but objects are always namespaced by repository. Forks shared an object database for efficiency, and have technically added reachable objects to a shared database via fork, but it's never been a real problem afaik.

But the very wrong assumption here is: "if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes"

All parts of this are incorrect.

You can push a colliding commit to a fork, but Git will see that it's already there and ignore it - first write does win. Also, "backend doesn't check uniqueness" is also wrong. The server will check for collisions and if this particular case happens, the server will see this and warn you _AND_ not write the object.


It doesn't reject, but it will not replace. Same for a fetch/pull. That is another issue with this attack vector (that Linus also mentions) - it has to be the _first_ time that a node has seen this object. It makes the attack even more difficult than it already is (in like 4 different major ways)

(2 counter) is impractical because all nodes of git will not replace objects if it thinks it already has it. So any attack has to assume this is the first time the node fetched, which is difficult before trust is established, which is difficult. This is part of the argument Linus originally outlined for this vector, which is that it only works for _very recent_ objects.

(1/3 counter) is not what I argued. I argued from the worst-case position that collision and preimages were theoretically cheap and fast. Even in that case, I feel my arguments hold.

The main issue here is that you assume you can replace an existing object with a replaced one, which you cannot. Not only that, but in all known cases, the sha1dc variant of SHA1 that Git uses will even _tell_ you that someone tried to do this, which singles out the source quickly.


It was theoretical - the point was that maybe some paper is published or some new tech or issue comes up. Now we have to do this again. If we separate the concerns, then we don't have to deal with both as though they're one problem. We can deal with one thing for content addressing and another for trust and security.

Heard. If we’re going through this pain of a migration- adopting something that is forward compatible at the ecosystem/community level like ipfs makes sense to me. You throw a byte at the front of the hash which indicates the format the hash is in.

We have a content addressable storage system internally and we use a self describing hash for it, so we can change the hashing scheme for different use cases.

Multihash is ipfs’ container for this.


Also, functionally, this is incredibly easy to add to Git.

Actually, this entire blog post came out of a short chat at Git Merge a few weeks ago with Jeff King. I argued more or less this and he didn't _entirely_ disagree, though he has good counterarguments on the list over the last few years, so I don't really know how he thinks about it ultimately.

I would write this to the mailing list, but I thought a conversation that includes people outside that list is more interesting to me. Ultimately I'm not sure if I'm dumb about this or the whistle blower that's willing to actually say "maybe this isn't the right call"


Actually, I think sha-256 is possibly faster than the sha1dc variant that Git currently uses.

I just sent a patch series to the list that enables sha1dc to be accelerated on modern CPU architectures to close to normal SHA1 speeds, but since it was ported from a Rust project by an agent, it will never be applied.

https://lore.kernel.org/git/20260929112544.86511-1-scott@git...


Last I checked SHA-256 was faster than SHA-1, and SHA-512 was even faster (though the output is annoyingly long).

How could SHA-512 be faster? It does more rounds of the exact same operations as SHA-256 with a bigger state. Although if it really were faster, SHA-512/256 gives you a truncated version.

It hashes a double amount of data per cycle with much less than a double amount of operations.

SHA-512 is always faster in software than SHA-256, when run on 64-bit CPUs, and it is also faster in the CPUs that support both SHA-256 and SHA-512 in hardware.

Arm-based CPUs have supported SHA-512 already for many years and the latest Intel CPUs also support it, i.e. Lunar Lake, Arrow Lake S (S is for desktops, Arrow Lake H for laptops does not support it), Panther Lake and Clearwater Forest.

I expect that AMD Zen 6 should also support it, because they are the last important vendor without SHA-512 support.

All modern CPUs support SHA-256 in hardware, so it is faster when SHA-512 is not supported in hardware, otherwise SHA-512/256 is preferable, by being both faster and more secure.


Thanks, that makes sense. I was vaguely aware that SHA-256 halved the internal state from SHA-512 but not that it halved the block size, I haven't looked much into it. Although it makes sense in retrospect that they would simply use 32-bit variables to keep the same structure.

[flagged]

What part of the code did you find bad when you reviewed it?

Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: