HN Simulatornew | past | comments | lists | submit | irusensei's commentslogin

BCacheFS is the best Linux filesystem now that storage costs a premium.

You can mix devices of different sizes and types on bcachefs. You can have foreground and background devices to balance performance and also different compression settings for foreground and background transactions.

You can set replicas=N to the individual file or directory on bcachefs. For example files you can just re-download or re-build. Likewise you can set a higher number of copies to important files.


Sure, when it's mature...


Yeah, and I think Overstreet might be going hard on AI, so not sure how that might effect the project


If you have "refine until it's perfect and don't screw with things you don't understand" thoroughly ingrained, along with the dangers of overconfidence, you'll do fine with AI.

Not everyone gets it though, that's for sure.

And, if you want to know if it's mature, I'd trust the user reports over the one liners :)


I've been using ZFS on my Linux servers for years and have pretty good experience with it. I've not been following Linux development too much. Is bcachefs usable/stable/reliable enough to replace it?


It's not as battle-tested like ZFS with 15+ years of field usage, but it's good enough for me and other users.

There's a NAS appliance called NASty and it shares publicly usage stats: https://nasty-telemetry.pages.dev/ Those numbers are NASty alone. There are more users on other systems.


15+ years of ZFS code is also often said as being a nearly unmanageable pile of mess, supposedly to the point that there’s some of the code that maintainers are afraid of touching, and with plenty of unresolved weird bugs (which I experienced myself with data loss).

Sometimes starting fresh with one coherent codebase and all features design baked in from the start might be better.


It takes 10+ years for any storage engine to iron out all the wrinkles, get it to a state where it's both fast enough, and you can trust it to not lose data. It's stupid to throw it away once it reaches that point.


It's not stupid if the 10+ years code is unmaintainable. Why did you ignore that part?


People keep saying it takes 10 years, but that presupposes methods that never improve. Why would we keep doing the same thing over and over again? Wouldn't be much point in that.


That might be the best thing happened to the project since now development can happen at its own pace without the clicky bait influencers.

In fact they delivered the erasure coding for parity raid back in march this year.

The thing is that as soon as you seriously give a chance to Bcachefs you see how good it is. I can only tell you that mixing different device tiers and having a per-file/directory replication setting is a god send specially in these times where storage costs more than gold.


> The thing is that as soon as you seriously give a chance to Bcachefs you see how good it is.

I'm pretty sure Bcachefs is amazing and better than Btrfs. I also think Zfs is amazing and better than Btrfs. Even so, I still use Btrfs because I know it is guaranteed to always be present on any Linux without any effort on my part.

> That might be the best thing happened to the project since now development can happen at its own pace without the clicky bait influencers.

A better approach might have been to just paused mainline merging instead of forcing being kicked out?

Eg "Hey Linus, Bcachefs is still in early development and I need to merge changes in a pace that is not compatible with Linux development process. So I'm going to pause for a while now and once it reaches maintenance status I will focus on submitting patches in a healthy pace that you can digest".


> Even so, I still use Btrfs because I know it is guaranteed to always be present on any Linux without any effort on my part.

100%. My system is rock solid and the last thing I need is rolling the dice after every update on whether my system will boot. https://www.reddit.com/r/archlinux/comments/eywcp7/linux_551...

I'm impressed with bcachefs's accomplishments though, and if they ever reconcile with the kernel I'll surely give it a fair shake.


btrfs has one critical issue they don't fix: it blocks access to fs for minutes if you remove large files. I am not sure how this is acceptable for prod grade fs..


> it blocks access to fs for minutes if you remove large files.

How large is large? I've deleted files with sizes of tens to hundreds of GBs and not seen that, and can probably whip up a test with a single-digit TB file if motivated.

Do you perhaps have 'discard=sync' in your mount options, or are using a kernel earlier than 6.2, which is the version -according to the docs- where async discard became the default?


> How large is large? I've deleted files with sizes of tens to hundreds of GBs and not seen that, and can probably whip up a test with a single-digit TB file if motivated.

for 1TB compressed (probably 5tb uncompressed) it is reproducable 100% reliably for me.

Here is some discussion: https://www.reddit.com/r/btrfs/comments/1mok440/filesystem_l...


Thanks for the information.

I made a ~3TB btrfs FS and mounted it with 'force-compress', put 5TB of zeros on it (which compressed down to like 160GB), and did a delete along with some concurrent operations on that same FS. Based on what I saw, btrfs doesn't "block access for minutes" while a large delete is in progress, but access to the volume that has the delete in progress is dreadfully slow. I used vim to create a new file in the mountpoint and saw that write delays were between ten and twenty seconds. Really bad, but still functional. Not at all blocked.

Someone in that Reddit discussion that you linked to says that all btrfs filesystems hang during an extremely large delete. This is not what happens for me. The only btrfs FS made slow was the one that had the ongoing delete. I have four other btrfs filesystems mounted and they're all just fine, whether or not they're on the same physical disk that has the ongoing delete. space_cache is v2 on all of my btrfs filesystems.

On my system, it looks like an events_unbound kworker was eating 100% of a single CPU while the big delete was in progress. No other kernel threads seemed to be consistently occupied.

For fun, I re-ran the thing I document below when mounted without compression, and then with an uncompressable file when mounted with non-forced compression. I had to reduce the size of the file to 2TB for both scenarios, but omitting compression writes out like 10x the data to disk, so it still seems like a fair test.

I'm not going to the trouble to provide a transcript for those two runs, but both when mounted without compression enabled and when an uncompressable file was written to a compression-not-forced volume, I saw absolutely no delays in filesystem operations while I was deleting that 2TB file. FWIW, putting 2TB of /dev/zero on that compression-not-forced volume and deleting it behaved the same as it did on a 'force-compress' mount.

Whatever is causing the dreadful slowness is directly linked to transparent compression, rather than being something you get when you run btrfs in all configurations. "Why run btrfs if not for transparent compression?" you might ask. I would answer: "Snapshots and reflinks, and yes, I know that XFS has reflinks too.".

A lightly-edited terminal transcript follows for if you want to double-check my work up to the end of the 'force-compress' run.


The lightly-edited transcript, relocated to a second comment to maybe avoid folding up the report:

  # lvcreate --size 3T --name testlv --stripes=2 testvg
    Using default stripesize 64.00 KiB.
    Logical volume "testlv" created.
  # mkfs.btrfs /dev/mapper/testvg-testlv
  btrfs-progs v7.1
  See https://btrfs.readthedocs.io for more information.
  
  # mount -o compress-force /dev/mapper/testvg-testlv /mnt/test/
  # btrfs fi usage
      Device size:     3.00TiB
      Device allocated:     2.02GiB
      Device unallocated:     3.00TiB
      Device missing:       0.00B
      Device slack:       0.00B
      Used:    320.00KiB
      Free (estimated):     3.00TiB (min: 1.50TiB)
  
  # dd if=/dev/zero of=/mnt/test/5TBFile bs=4MiB count=5TiB
  1310720+0 records in
  1310720+0 records out
  5497558138880 bytes (5.5 TB, 5.0 TiB) copied, 2510.22 s, 2.2 GB/s
  # /usr/bin/time --format='** fi sync %e' btrfs fi sync /mnt/test
  ** fi sync 0.00
  # btrfs fi df /mnt/test/ | grep Data
  Data, single: total=160.00GiB, used=160.00GiB
  # /usr/bin/time --format='** totalTime %e' bash -c "
    /usr/bin/time --format='** 20gbTime %e' bash -c 'dd if=/dev/zero of=/mnt/test/20GBFile bs=4MiB count=20GiB status=none; du -h /mnt/test/20GBFile; rm /mnt/test/20GBFile'&
    /usr/bin/time --format='** 10gbTime %e' bash -c 'dd if=/dev/zero of=/mnt/test/10GBFile bs=4MiB count=10GiB status=none;  du -h /mnt/test/10GBFile; rm /mnt/test/10GBFile'&
    wait"
  10G /mnt/test/10GBFile
  ** 10gbTime 2.42
  20G /mnt/test/20GBFile
  ** 20gbTime 4.95
  ** totalTime 4.95
  # date
  Sun Sep 20 10:19:46 PM PDT 2026
  # /usr/bin/time --format='** totalTime %e' bash -c "
    /usr/bin/time --format='** 05tbTime %e' rm /mnt/test/5TBFile &
    sleep 5 # It takes a few seconds for the delete to start making things slow when the file has been compressed. The operations on the 10GB file will complete in a normal amount of time if this sleep isn't present.
    /usr/bin/time --format='** 20gbTime %e' bash -c 'dd if=/dev/zero of=/mnt/test/20GBFile bs=4MiB count=20GiB status=none; du -h /mnt/test/20GBFile; rm /mnt/test/20GBFile'&
    /usr/bin/time --format='** 10gbTime %e' bash -c 'dd if=/dev/zero of=/mnt/test/10GBFile bs=4MiB count=10GiB status=none;  du -h /mnt/test/10GBFile; rm /mnt/test/10GBFile'&
    wait"
  10G /mnt/test/10GBFile
  ** 10gbTime 183.64
  20G /mnt/test/20GBFile
  ** 20gbTime 259.47
  ** 05tbTime 1026.21
  ** totalTime 1026.22
  # date ; /usr/bin/time --format='** 20gbTime2 %e' bash -c 'dd if=/dev/zero of=/mnt/test/20GBFile bs=4MiB count=20GiB status=none; du -h /mnt/test/20GBFile; rm /mnt/test/20GBFile' ; date
  Sun Sep 20 10:36:52 PM PDT 2026
  20G /mnt/test/20GBFile
  ** 20gbTime2 23.04
  Sun Sep 20 10:37:15 PM PDT 2026


Thank you, its useful finding that only transparently compressed btrfs is affected.


> Do you perhaps have 'discard=sync'

I have 'discard=async'


Only if you have quotas enabled IIRC.


no quota enabled


bcachefs on Arch is a bit better supported, we have the distro package maintainer in the bcachefs IRC channel, and I've never lagged on mainline support like ZFS has.

Actual distro support, and doing it right with people actually communicating with each other, has always been a priority for the project.


> I still use Btrfs because I know it is guaranteed to always be present on any Linux without any effort on my part.

I used to think that about ReiserFS, too. It was in the mainline kernel, development was snappy, and it solved some performance problems. I used it all over the place.

Things then subsequently... changed. :-/


You had plenty of time to move away from reiserfs. Something like 15 years between the conviction and removal.


I did.

And I did move away from it.

But I had once expected ReiserFS to be permanent, especially since this Hans Raiser dude who was driving the ship seemed to be sharp AF, and this seemed doubly-true when his filesystem got mainlined.

But it was not permanent. It did not last forever.

This idea of permanence, or rather the lack of it, was the whole of the point that I was responding to and also trying to impress upon.

At the end of the day: We do not know the future. Things can change.

I've outlived this one high-performance Linux filesystem in my life that I was using. This does not in any way mean that I will be the last to outlive other Linux filesystems.

Permanence is not guaranteed.


> I still use Btrfs because I know it is guaranteed to always be present on any Linux without any effort on my part.

Except RHEL. They don’t include it in their kernels.

Alma Linux started including it again though.

It can never be easy.


Came here to say this. Supposedly Fedora also considers giving up on it.


This sounds like FUD, do you have any references? Genuinely asking. I follow LWN reporting religiously, which in turn follows Fedora development (and associated drama) closely, and haven't seen anything said in this direction. Just had a quick look on LWN and Fedora development resources, and nothing came up.


This is what I read here and elsewhere bunch of times, which is why I said "supposedly". A quick search did not reveal any meaningful proof, though. I'll ask next time I hear someone say this.


unlikely. reference please. btrfs was the default filesystem on fedora installs last time i checked. to remove it they would have to first change that, then give it a few years before even considering removing it from the kernel. redhat could remove it because it was never default and never recommended.


This is what I read here and elsewhere bunch of times, which is why I said "supposedly". A quick search did not reveal any meaningful proof, though. I'll ask next time I hear someone say this.


I think in a perfect world they should had put someone in between to mediate and curate patches while providing DKMS for urgent patches.

As for BTRFS I think its also pretty good. Its just that I have the impression its development is guided by the needs of its sponsors and sadly for us META doesn't need RAID5.


Meta doesn't have anyone working on btrfs anymore, it appears to be two guys at SuSE and drive bys.


Not true, Boris Burkov from Meta works on btrfs. And two people from WDC.


Within the last year? I think bcachefs is going to be overtaking btrfs soon on active developers, from the trends I saw in the commit logs


I was very up front about where we were at.

A lot of things were tried, people did try to mediate.

The particularly galling thing though was when I finally started looking - post split - comparing bcachefs PRs to other subsystems and especially XFS - I was being more conservative with what I considered a critical bugfix.

There was never a clear statement on what the issue was. What you guys got in public was about as much as I got.

All I can say is - going fast when you're stabilizing and getting bugfixes out the door is what you can and should be doing when you've invested in test coverage, test automation, keeping the codebase clean and asserted, and building up a community that works well together on testing and shaking things out.

I genuinely do not know what they were thinking.


> There was never a clear statement on what the issue was. What you guys got in public was about as much as I got.

If you are referring to why bcachefs was removed from the Linux kernel, here's a discussion on bcachefs being removed from the Linux kernel.

https://news.ycombinator.com/item?id=44868868


I'd like to point out that HN user koverstreet was involved in those threads, here.

They already know what was discussed.

(Good? Bad? Indifferent? I don't know and I don't have a dog in this race. I'm just here connecting the dots.)


> I'd like to point out that HN user koverstreet was involved in those threads, here.

Yes, that's why those remarks on how it's a mystery how bcachefs was pulled from the kernel are perplexing. To me they sound like gaslighting.


Just pleases try to get it back into mainline.


bcachefs was already working at its own pace prior to being accepted in the kernel. It could have continued doing so for years until it was really "ready".

Instead it got kicked out because Kent constantly ignored the kernel's contribution rules and is unlikely it will ever be accepted back into the kernel.


That would be a shame since there is nothing else over there with the same set of features. Disregarding drama, I'm telling you it’s that good.


I'd really appreciate it if we could drop the FUD over contribution rules. There are no such rules, it is explicitly Linus's way or the highway, and I already replied to that elsewhere.

And it went in when it did because Redhat was pushing for it and claiming to be supportive - but that never materialized. They wanted to get something for free without investing, or putting in the absolute bare minimum.

A _lot_ of people were saying publicly and privately "dear god yes we need something better than btrfs" - but no one from the existing kernel community was interested in stepping up.

Community's still growing, though. A lot of people have gotten active in making sure bcachefs actually works well for people end to end, and there's a hell of a lot more to shipping a filesystem than just writing kernel code.


You can't let Reddit guide your technical decisions.

The FS was marked experimental, so there is no urgency in fixing bugs or providing features in a certain cycle. Everyone using it knows what they got themselves into. You can still provide the DKMS module for faster fixes and features for anyone who wants to use BCacheFS more seriously for the time that the upstreaming process takes, but eventually it would have all been on mainline.

Asahi is taking a similar approach where they have their downstream kernel and push things upstream once they are mature.

That means the upstream kernel is not useful for running on that hardware now, but things are moving there eventually.


No urgency over fixing bugs? What do you think this is, btrfs? :)

All this has been discussed to death, we don't need people armchair quarterbacking a year later. It's over, it's time to move on.


If anyone is curious and want to skip all the PR talk:

>Combined with SVE2 vector operations and software optimization,

It’s ARMv9.


It was funny looking over this just waiting to see what ISA it was. No surprise that it was ARM but you never know.

Nowadays if you have the source code the ISA is practically irrelevant.


Why did they bury the lede so much on what the ISA is?


Press releases are sadly like that. Fujitsu's main page has it in the first line:

> FUJITSU-MONAKA is a Next-Gen Arm-based processor, set for release in 2027, designed to address the challenges of next-generation data centers with its unmatched performance and power efficiency.

https://global.fujitsu/en-global/technology/research/fujitsu...

It is no secret it is Arm; MONAKA was announced in late 2024.

https://www.techpowerup.com/329761/fujitsu-previews-monaka-1...


I hoped it would be SPARC.


I was hoping risc-v. One of these days


I hoped it would be Alpha.


Why not 68000


I crossed my fingers for a 8087


I can't believe its not butter


I wonder how many cores they plan to cram in there, at least like 256 right?


144 per the article


Funny first thing I did was to search 'arm' and I got 1 result at the bottom of the page.


Hipe they'll be affordable


Is this classic 9front tomfoolery?


Indeed. If that URL is too crude then one may also use only9fans.com.


>It also solved the issue I had with entering a password on FDE (full disk encryption) systems

I've fixed that with secure boot (my own keys, not Microsoft's) and TPM2. From that host I can run a program named Tang, which operates in conjunction with another program named Clevis.

On hosts without secure boot such as RPis and other ARM SBCs, Clevis will contact Tang at boot time and decrypt the disk.

Incidentally my x86 with secure boot has intel vPro so the KVM is not needed. The remaining sisters have a piece of lost technology used for decades to solve the KVM problem: serial (UART specifically) connections.

Some seasoned sysadmins might have memories of dialing their Sun servers serial ports for remote administration.


To me, it is a feature to have a password required at boot to unlock the drive. Relying on TPM for keymatter is convenient but comes with caveats I don't like. I don't personally trust that the boot chain and OS on a modern system are secure enough for this model to be similarly secure to using a passphrase properly. And if I care enough to try to secure something in this way, I definitely care enough to pick something that I believe would be at least truly secure at rest with a decent degree of certainty.

TPM based unlock does at least still fulfill the goal of ensuring data stored to disk is encrypted so that it can't easily be recovered from a discarded drive.


I am using Tang and Clevis without a TPM on an Orange Pi 5. The key is stored on my Tang server. No TPM required.

It is either on my LAN with that server available, or you will need to enter a key.

It am only trying to prevent casual snooping if it goes missing from my garage though. Anybody more sophisticated can have my garage YouTube browsing history.


This. I'm not fighting against a state level actor. My concern is some crackhead burglar stealing my stuff and then my personal data ending up in the hands of whoever buys the stolen goods.


Exactly. Disk encryption simplifies a stolen device to a VISA problem.

As in, no problem, I will just buy another one. At these RAM and storage prices, maybe not.


I'm using a certain object storage implementation post minio enshitification. The software itself is great don't get me wrong but I've noticed their logs are basically unreadable. Its metrics and traces in json data meant to be rendered on a dashboard instead of being read by humans. It's also extremely verbose even at an INFO level.

Maybe just get these on an open telemetry endpoint instead? I also don't get why people send by default json logs to journald as it's clearly meant to be a replacement to syslog which is already a good standard.


I heard about Alpha in that it while being more powerful than x86 it wasn't as alien as the competition. The boards had PCI slots and looked like normal PCs, it had Windows NT and Linux. So if your Exchange server was not handling the load you moved it to Alpha and it was good.

It also seems faith in IA64 was the meteor that killed most of these self developed RISC architectures.


FWIW Alpha pretty quickly went with mainly PCI + few legacy ISA slots for considerable chunk of the line, unlike competition which used proprietary buses that might have been faster but meant extra level of exotic in procurement


A VAR I worked with back in the day made great scratch for a good number of years flogging MS SQLServer on Alpha for the performance boost, which was apparently more than enough to justify the price premium for those that needed it.


Apple private relay is very flawed but on the other hand most sites that view connections originating from VPN as suspicious seem to be ok with Apple Private Relay. If you have VPN on your router private relay on top of it could be a good way to do ip address laundering.


I find it pretty annoying that iCloud Private Relay does not work together with VPNs.

I mainly use VPNs to access bank sites (or other sites that are annoying enough to use IP country as a proxy for traffic being evil/legitimate) when traveling, but I often forget to turn off the VPN afterwards, and then spend the rest of my day browsing from my home IP (that terminates the VPN), which would have been hidden behind Private Relay if I'd actually been browsing from my home Wi-Fi.

Cross-site/app ad targeting getting creepy good despite using different browser profiles for work/personal browsing etc. is usually a good tell (at least on IPv6).


> Improved support for vintage and misc. hardware: alpha evbppc hp300 hppa m68k mac68k macppc mips x68k

It is interesting to see that as Linux drops support for legacy computers NetBSD still welcomes them. This is a NetBSD thing as it's not so much for FreeBSD.

As of now NetBSD is probably the go-to operating system for vintage hardware.


Afaik it has always been that way. I remember 20 years ago as I was dabbling with various bsds, OpenBSD was known for security, FreeBSD for features and netbsd for small, ultra portable kernel.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: