It's only half the story. WireGuard uses chacha20-poly1305 and certain implementations can still achieve 100Gbps. Hardware acceleration is the path to 400Gbps and power efficiency though.
IPsec is a terrible fit. Tailscale uses a control plane for peer & key distribution and renders almost all of the complexity in the IPsec protocol useless.
(Tailscale cofounder) Many years ago I made a VPN for experimental purposes that kept IPsec’s data plane but abandoned IKE. It’s actually a pretty good arrangement and might be applicable here.
My implementation relied on the extremely finicky Linux kernel IPsec and would never have worked in userspace macOS or windows. But yes, a Tailscale control plane with a per-node switchable data plane, one of which is IPsec with PQC, is a real option that would work.
By locking down the control plane it becomes incompatible with classic IPsec, but also avoids most of the security gotchas that plagued IPsec.
1) Besides marketing, I see no reason that Tailscale needs to concern itself with the WireGuard protocol spec at all. Tailscale is already incompatible.
2) Tailscale's control plane model for peer distribution avoids every concern affecting IPsec in regard to protocol negotiation security and compatibility.
TL;DR, diverging from WireGuard & supporting AES (via protocol negotiation with hardware acceleration), would be relatively painless.
(someone who went on rabbit hole trip through tsnet) It's not very visible (and not really user-accessible) but Tailscale supports connecting arbitrary wireguard peers into the network, and it's how mullvad integration works IIRC - got confirmation on twitter I think from apenwarr.
So yes, tailscale is compatible with wireguard, it's just not exposed by any of the major coordination servers other than mullvad integration in tailscale.com, but the clients will happily consume information about wireguard peers that do not run tailscale at all.
> Tailscale supports connecting arbitrary wireguard peers into the network
Now screaming internally because every router I've had that's supported Wireguard I wish I could have just joined into a tailnet directly for subnet routing... and I can't find anything else that just plays nice either.
PfSense and OPNsense both support Tailscale and can integrate it with their normal routing/firewall behavior. I think OpenWRT can as well but I haven't tried it.
They do use the official client, but it also makes it fairly easy to bridge tailscale networks with regular wireguard clients with normal routing rules.
You lose the formally-verified cryptography guarantees underpinning WireGuard & its implementation, though. That's a bigger tradeoff: https://www.wireguard.com/protocol/
Implementation requires different considerations, no? For example, AES-GCM needs to be supplied a big-endian nonce, whilst ChaPoly uses little-endian. As the WireGuard paper notes, ChaPoly (in software) can be better protected against (CPU) side-channels. Besides extended-nonce AEAD being "native" to ChaPoly, BLAKE family of hash functions that WireGuard uses, are also based on the same construction as ChaPoly, making the implementation leaner (something Jason keenly emphasizes as an advantage).
No. I specifically said, "lose the formally-verified cryptography guarantees underpinning WireGuard & its implementation" and I pointed out the implementation quirks.
Anywho, given your strong conviction about this, consider suggesting to Jason on the WireGuard mailing list to swap ChaPoly for AES-GCM, or perhaps consider a fork yourself. Good luck.
It is slow. It cannot achieve speeds of greater than 1Gbps on clients systems (Windows & Mac), where you'd normally see it being used. On Linux, it struggles to achieve 10Gbps even when using a synthetic large packet benchmark [1]. With an IMIX benchmark, it would not be competitive whatsoever.
This problem is fixable. WireGuard achieves higher performance (Kernel vs Userspace implementation) and IPsec implementations can achieve 100Gbps/400Gbps (DPDK/XDP). Zero-copy networking.
From this blog post, I can say Tailscale still seems to not have the appetite for that, which is a shame.
Remember that at least one LPE CVE associated to kernel IPSec implementation has been discovered (copy.fail), which means that whatever gains you get from this vpn tunneling, is lost by breaking the basic user security system guarantee.
You are better off not using a VPN at all rather than using kernel crypto
By that logic, we should avoid TCP as the Linux kernel implementation has had plenty of CVEs. Thankfully our expert critical thinking helps us acknowledge that as silly.
Not true. There are plenty of userspace TCP stacks in production today, just like there are also userspace IPsec stacks. UDP is also an option if TCP is too complex for your tastes.
If you believe that, you don't understand TCP, I recommend reading the actual RFC (RFC 793) and the IP RFC if necessary.
A port identifies a process, an IP address identifies a machine. The machine receives the packet, and then passes it to the process, since it's the machine that needs to receive the packet before routing it to the process, it's the kernel with ring 0 privileges that needs to process the packet and map the port to the process.
There's nothing stopping you from giving each stack its own IP, making the kernel only responsible for routing.
I don't think it's a good idea to put more of the stack than necessary into each individual program though, because then you're reliant on each program implementing everything correctly and you have to worry about bugs and security issues in every single program you're running rather than just in the kernel.
Considering that even something as simple as opening a socket (which involves calling one function to give you three numbers which you pass unchanged to a second function) is largely a clusterfuck in existing programs, I don't trust them to implement the whole TCP stack.
The performance hit for running an emulated CPU would be significant, to the point that it wouldn't really be the same argument.
I was thinking more along the lines of network namespaces -- although actually, I should have been thinking of TUN interfaces. Any packets sent into one of those are delivered to the attached program, which can do what it likes with them. Those are usually used by e.g. OpenVPN for L3 things, but nothing stops a program from handling the L4 headers.
And if any tailscale employees are reading this - https://github.com/tailscale/tailscale/issues/15724 please fix this too. Regular users not using some sort of enterprise saas DNS (whatever their thing is?) deserve DNS privacy too.
(Tailscale cofounder) That’s a good callout on DoH support, thanks.
That said, note that if you run your own DNS server on your tailnet, the regular UDP DNS is automatically private because it’s carried over Tailscale. That’s the most common setup for non-SaaS DNS servers. DoH doesn’t really add anything in that arrangement. (And it’s more fiddly because you need to get and refresh a TLS cert.)
Tailscale's netstack is barely even WireGuard and they aren't compatible whatsoever. It's all marketing at this point.
So it's not that simple: it's impossible for Tailscale to use any existing kernel or accelerated WireGuard implementation. They could derive inspiration, but a kernel module for Linux won't fix Windows & Mac. With that said, I feel they have enough funding to maintain a few platforms (:
You're correct, kernel isn't faster by default. With that said, the following is true:
1) the WireGuard kernel implementation, despite not even being zero-copy, exceeds the performance of the userspace implementation
2) implementations utilizing the userspace network stack have a maximum potential performance (context switch + memcpy is very slow, and that affects UDP disproportionately). It's the wrong approach for meaningful improvement.
What do you mean by “userspace network stack” when we’re talking about io_uring? That’s a contradiction. Unless you mean extra memcpy’s within the kernel network stack, but when talking about the buffer you supply I don’t believe that’s true - your data generally gets directly DMA’ed into the device because the buffer you supplied is pinned and can’t be released until the second CQE is delivered. It would be helpful if you clarified.
I do not understand why people continue parroting this performance nonsense.
Memory copying is on the order of 100 gigabytes per second. You can do 80 full payload copys and still out-pace a dog-slow 10 gigabit per second connection. If your bottleneck is memory copying you are either doing something very wrong and doing way too many copys or congratulations you have implemented one of the fastest network stacks.
Supervisor calls are also very fast, on the order of 100 ns up to maybe 1 us with all the Spectre mitigations. Even if you did something as stupid as one supervisor call per packet, you would still be getting on the order of 10 Gbps at the long end there and 100 Gbps at the short end. Which, again, means congratulations are in order because you have implemented one of the fastest network stacks. If you do batching and add just 10 us (us, not ms) of latency then that entire cost is so small as to be irrelevant. You are going to bottleneck on your memory copying first.
Network stacks are so slow almost entirely due to poor protocol design and poor protocol implementation. Usually both.
I believe the 30W figure was power usage before SFPs or DACs. Looks like 400G sfps are about 10W, so you'd double it when fully loaded. I believe passive twinax doesn't use anywhere near that though.
"there was a correction issue" is downplaying it. Etcd is truly the worst example of Raft.
Etcd corruption and loss of quorum is extremely common in practice and the GitHub issues sit for years. The design is simple, the performance is modest, yet it still has still never been reliable, despite being marketed as so. I can't speak to whether this is specifically due to their Raft implementation, but I'd argue the entire codebase is over-engineered and questionable.
Their lock, leader election, sessions, and leases are all awful and I'd never recommend anyone to use those. But as a strongly consistent kv store and if you need the watch mechanics, its useful. It has its place and that's mostly being used by kubernetes.
No they won't. They are generally faster in throughput than any non-gc application that isn't heavily hand optimized. Their problems are higher memory usage and unpredictable latency, not speed.
The Apple Upgrade lease model isn't even a replacement as you have to wait 24 months before upgrading.
The entire point of the iPhone Upgrade Program was to make it painless to swap to the newest iPhone every year. Without that, there's no value proposition.
They have 12 month options for the Apple Upgrade program. I just checked how much it would cost to get an iPhone 17 PM + AppleCare+ with theft protection like I currently get from the IUP and it's $3 more a month.
I mean, it is $36 more a year but I guess that's not terrible. What will be interesting is the tax paid up front. IIRC, I paid tax on the entire device every time I renewed rather than just paying it during each month. So I always thought it wasn't the best deal because I paid tax on a device that I only ended up paying off half of before renewing and starting the process again. I could be wrong, though. I'm not the best with finances.