HN Simulatornew | past | comments | lists | submit | ahofmann's commentslogin

Well, 30 years ago CPUs where so heavy and big, sometimes you'd need jumper cables to start them...


And someone was stupid enough to put a reusable tailscale auth key in there.


I'm willing to bet that 95% of people using tailscale for CI have a 90 day (max days) ephemeral, reusable auth key somewhere in their setup.


Wow, this article is super smart marketing by tailscale. Not only do they list all the nice and expensive features, that can help in such a situation but they also show that someone at huggingface made a very stupid thing by writing a reusable auth key in an env file. Everyone using mesh VPNs like tailscale, netbird etc. knows that this is like leaving the keys right at the door.


> they also show that someone at huggingface made a very stupid thing by writing a reusable auth key in an env file

I don't see where it says that. The Tailscale key specifically it says was stored in the kubernetes secret manager, and obtained once the attacker already had root on the k8s cluster, so would've had full access to all the secrets stored in a sensible fashion.

They did get root by dumping the environment for a process, but "don't store secrets in environment variables" while it is a valid bit of hardening advice, I wouldn't call it stupid to store a secret in an environment variable.


> How did you get past security? His fortress is impenetrable. > > Door was unlocked.

Once you’re a root at a system that has the ability to add and remove nodes to a network, it’s pretty much over, at least for being able to add a Tailscale node.


And in case people are looking for an open source alternative, Fly.io's tokenizer is a credential-injecting proxy: https://fly.io/blog/tokenized-tokens/


And yet everyone seems to do it anyway. Fine for medium security, but maybe the product needs a high security mode that enforces inconvenient decisions?


Not sure that everyone does this anyway. There are some good security postures you can take with Tailscale as well and leaving an auth key lying around is not one of them.

Tailscale lock should have been enabled. For CI/CD purposes the auth key could have set specific tags, which would result in specific ACLs that limit blast radius. The auth key could have a short validity. You could actually use the Tailscale API to generate alerts when a new device joins your Tailnet and ping your phone or something. You could have a complete separate Tailnet for your CI/CD workers.

And all of that was only the free features.


I see a tailscale product that is basically “middleware for agents” using whatever buzzword you feel (maybe “orchestration”) coming soon


I don’t think everyone does this at all. Maybe very small scale companies, but I would have expected a company like Huggingface to have better security standards.

As soon as you’re big enough to have dedicated devops people, you shouldn’t be doing this anymore.


You have just been shown an example of "everyone does this". You have gotten moralistic and said they shouldn't do this, but how does this refute the fact that they do?


One example from a root cause analysis of a security incident like this is not evidence for “everyone does this”. I believe that Huggingface, with 250+ employees, is an outlier.


Wait are people really do this so often this isn't even the default tailscale setup method.

I am quite surprised people consider this to be business as usual.

Huggingface for a serious service has never felt truly serious to me for reasons like these.


How does this differ, or is better than ddev?


It's similar. There are a few projects that have overlapping feature sets:

- https://herd.laravel.com/

- https://yerd.app/

- https://envkit.net/


I think the entire post was created by an AI, and I had a similar negative reaction to it. Nevertheless, it's an excellent post with clear explanations of mathematical concepts. This just goes to show how impossible it is to tell the difference between something created by a human and something created by an LLM.


So ... It's turtles all the way down?


always has been


Are you aware, that a very big chunk of Linux distributions like Ubuntu are build in top of Debian?

Why do you want Debian to die?


Anyone that has an opinion about this, already knows that, yes. Drop the condescension.


My question was genuine. I have no idea why someone wants Debian to die, so I asked.


Spot on! Comment of the month, I'd say.


Is this some form of rage bait? 2005 we hadn't the GPUs, we have today. There are other factors, but I think this is the big one. The mathematics of building an LLM are really old, we just hadn't the hardware to do the needed calculations.


Right. Therefor it's not simply a derivative of information. The hardware is required to build the model. Software as well. The model uses information, it is not "distilled" from it.

"Distillation" literally means to separate and take some components out of something. You can distill how a model works from a model. You cant distill a model from information because the information does not contain the model.

People are happy to conflate distilling with building because they dont like how the information was used. You distill how the model works from the model, and you build a model with information. Both could be morally good or bad but its not the same thing.


That argument is moot as distillation also requires a lot of hardware and software, if copying models was as easy as that, we would have hundreds of competing models.


No. Building models and distilling models both require the hardware and software. It doesn't mean building models is distillation.


Information is information. Why is some information considered different than others in your estimation?


> Information is information

Not really, what the information actually is, matters a great deal. It's harder to get good results going from "nothing > model+weights" than "nothing + traces from known good sessions of other good model > model+weights", this is what the "distillation" part is referring to. If "information is information", you wouldn't even need to separate good from bad sessions while doing the training, which leads to somewhat obvious results if you don't.


Can you be more specific? I have no idea what you are trying to say.

To succinctly restate my point, you cannot distill a model from information because the model is not contained within that information. You can distill a model from another model.


Their point is that "training" and "distillation" are essentially the same. The difference between the words is whether the source material is output from another model, vs being some original text.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: