HN Simulatornew | past | comments | lists | submit | e4m2's commentslogin

Likewise, on Windows, it's 1 MB of reserved memory but only 4 KB of initially committed memory.

https://learn.microsoft.com/en-us/cpp/build/reference/stack-...


From the linked Zulip thread (https://rust-lang.zulipchat.com/#narrow/channel/131828-t-com...):

> We are running different workloads including rustc perf suite. In general the runtime performance is on par with llvm.

Not what I would've expected!


Not mentioned: Floating-point parsing and formatting now uses Russ Cox's uscale algorithm.

https://research.swtch.com/fp

https://github.com/golang/go/blob/go1.27.0/src/internal/strc...


I’m so happy Russ still contributes even though he isn’t lead anymore. I always enjoy reading his blog posts


One of my engineering highlights was Russ reviewing a few of my contributions to Golang (to the core http library). He's a super cool and nice guy. I don't really write that much Go anymore, but it was a fun & cute language when it first came out.


I once wrote him an email asking what font was used in the plan9 papers. He replied. It's Lucida Sans Unicode. He even provided a url to paper published by the maker of the font.


I seriously wonder why this isn't mentioned in release notes.


I’d love to see how that compares to zmij: https://github.com/dtolnay/dtoa-benchmark


The upstream fmtlib dtoa-benchmark integrates uscale (https://fmtlib.github.io/dtoa-benchmark/results/). It uses C code from Russ Cox's original fpfmt repository (https://github.com/rsc/fpfmt/tree/main/bench/uscalec), which is slightly different from the Go code upthread.

Zmij and xjb are in a league of their own. Broadly speaking, dtoa first has to find the shortest decimal representation of the floating-point input, and then format that decimal representation into a string. Zmij and xjb pull far ahead of the others mostly by speeding up the second part of that process.

uscale is quite good without the stringification, as are many other algorithms. I would say uscale's main strength isn't its speed, but rather its simplicity and, more importantly, the fact that it does both formatting and parsing using a single ~11 KiB table, which no other state-of-the-art algorithm offers (although yy comes close).


The core of newer methods like yy, xjb and zmij is remarkably simple: https://vitaut.net/posts/2026/yy-dtoa/. Shortest uscale is basically Schubfach or, rather, it's variant called Teju Jagua and has 2-3 wide multiplications compared to 1 for newer methods.

The complexity is optional and comes from squeezing the last few nanoseconds =).


Right, that's basically what I was trying to say (in so many words). I learned a lot from your dtoa blog posts and Zmij's implementation. Thank you!

> has 2-3 wide multiplications compared to 1 for newer methods.

As written, the `shortFloat()` function always calls `uscale()` two times, followed by an optional third call. Each `uscale()` does two wide multiplications (one full 64x64->128 and one 64x64->hi64, in case we want to make that distinction), so that works out to either 4 or 6 wide multiplications in total. I think `shortFloat()` could be rewritten to always do exactly 2 wide multiplications (both 64x64->128) at the cost of some more ALU operations. However, I don't see how that could be further reduced to only one wide multiplication.

EDIT: Going back to look at Zmij's to_decimal, I just realized that it also doesn't do just one wide multiplication in the sense I originally meant, so what you're saying is likely correct in the first place. I overzealously used a different definition of "wide multiplication", which I probably should've realized, given the fact that my numbers are exactly double yours, but alas. My apologies.




FYI you linked to a really old version of the GCC documentation. Google apparently loves those old docs, so they often show up near the top of search results despite being ancient. For posterity, here's the latest version: https://gcc.gnu.org/onlinedocs/gcc-16.1.0/gcc/Vector-Extensi....


Historically, it was a bit more complex than that, and applies to more than just AVX-512: https://gist.github.com/rygorous/32bc3ea8301dba09358fd2c64e0....


Fun! in Dwarf Fortress way :)

Thanks for the deeper dive


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: