There are two types of pages: anonymous and mapped from files. The code pages are mapped from a binary; they are not very special.
Under memory pressure, the kernel evicts less popular pages from memory. If a page has been mapped from a file, it is dropped (if dirty, then it is written out first). If it is needed later, the kernel can read it back from the file. If a page is anonymous (read: heap page), then there is no backing file and the kernel copies it to swap before dropping it. This is swapping.
So, what happens if you disable swap and the kernel is low on memory? What can it evict? Anonymous pages cannot be evicted: there is no swap to put a copy in. The only choice the kernel has is to evict pages that are mapped from files. Those include pages mapped from the executable. You don't eliminate stalls by disabling swap, you just move them elsewhere: the kernel will page out code and your app gets paused whenever the execution flow hits such a page.
I haven't had to analyze the performance of no-swap processes before. My assumption is that code is hot enough to avoid eviction and that evicted code pages are rather the exception. To strong-man the argument, I can imagine long running complex (bloated) services could have parts that are not touched unless a specific request comes in.
>So, what happens if you disable swap and the kernel is low on memory? What can it evict?
Your userspace early OOM killer triggers and lets you know you're trying to run more than will fit in memory so you don't do it again. (In my experience, the kernel OOM killer can't be trusted to kill processes soon enough.)
You are missing that the file LRU is not only the code pages or memory mapped files but all file data. The active code pages are specifically excluded from being discarded during the reclaim process.
I wish this is true but evidence is against you. Just try testing memory pressure on a no swap system. You will see that the system grinds to a half for a few minutes before you even get to an OOM event.
Could you clarify what exactly is "not true"? Also, could you explain what exactly your evidence proves? What do you think happens when the system becomes unresponsive under memory pressure, and why?
> The active code pages are specifically excluded from being discarded during the reclaim process.
This is not true, or the system wouldn't grind to a halt due to having to reload code pages (even those that are recently in use) from disk over and over.
For example, key presses and output streaming in SSH become unresponsive. The code pages for handling key presses and output streaming are obviously in use — I just used them!
I'll try again. Why do you think that unresponsiveness is caused by code pages being discarded? Do you have evidence for this? Have you considered that there are other mechanisms that cause programs to become unresponsive under memory pressure?
It's the explanation that I've read when scouring forums for reasons why it happens. I'd love to hear alternative explanations and how to measure/verify them. But the measurement tools also become unresponsive so I can't check anything.
But frankly, as a user I don't really care what the real technical reason is. I just don't think the system should become unresponsive for minutes on end during memory pressure. Just OOM-kill earlier, when it's supposed to.
The assumption that a system gets unresponsive due to code pages being reclaimed doesn't make any sense. It is a common myth, cargo-cult knowledge. The kernel doesn't discard any active referenced pages at all, and the active code pages are the last that are getting reclaimed.
The proper explanation is relatively straightforward. Normally, any memory request for a process is served by getting a page from the free memory pool. But when the system free memory goes below the low watermark (calculated from vm.min_free_kbytes and vm.watermark_scale_factor), the kernel stops giving the pages from the free memory pool and instead goes into direct reclaim, trying to find unused pages. While the kernel does reclaim, the allocating process remains stalled. The kernel can't return control to the process - there is no page mapped that the process tried to access. This applies to any types of pages - anon, file-backed, code pages. So, under the condition of free memory below the low watermark, any process trying to allocate memory gets stalled. This is measured as memory pressure stalls (PSI) in /proc/pressure/memory.
The problem with the system being stuck in an unresponsive state is that the design choice for the kernel OOM handling is to be optimistic and keep the workload running as long as the kernel is making any progress reclaiming pages. The direct reclaim process needs to fail to make any progress 16 times before the OOM killer is invoked.
The solution is to use cgroup memory limits to avoid getting into the situation of memory exhaustion and use user-space oomd to recover from this situation earlier.
So you're saying processes are stuck for minutes on end at trying to get memory from the kernel? I find this implausible because relatively simple processes such as ssh and bash shouldn't be getting memory from the kernel at a continuous basis. Unless they're suddenly processing a lot of new data (like me trying to cat /dev/urandom), they should be spending most of the time working with the memory they already have. Especially when all I'm doing is pressing keystrokes in ssh+bash, not even hitting enter. Note that malloc doesn't ask memory from the kernel at every call. Glibc malloc is (in)famously stubborn with caching memory it has gotten from the kernel, hesitant to release it back to the kernel.
Any repeatable way to prove your thesis or disprove the trashing-code-pages thesis?
You seem to confuse user-space memory allocations (malloc) with the kernel memory page allocation and mapping. The processes don't need to allocate any new memory with malloc. They get stuck when accessing memory that is not mapped to a physical memory page or when they enter a syscall that needs in-kernel memory allocation.
Ok fine. But again: any repeatable way to prove your thesis or disprove the trashing-code-pages thesis? And how to do this when the system is freezing up?
I have already pointed you to /proc/pressure/memory. Prometheus node_exporter tracks it as node_pressure_memory_waiting_seconds_total and node_pressure_memory_stalled_seconds_total. You'll see them spiking when you experience stalling.
You can also track variables from /proc/vmstat exported as node_vmstat_allocstall_(normal|dma32|movable), node_vmstat_pgscan_direct, node_vmstat_pgsteal_direct - you'll see them spike when the system enters direct reclaim.
Install and configure systemd-oomd and track its log - it will activate on memory pressure.
To track code page eviction/re-faults specifically, you'll need to write some eBPF code, there are no direct measurements for this available.
Perhaps I wasn't clear, but I did say that all file pages are treated equally.
I would assume that page with PC is excluded from reclamation process, but other code pages are valid targets.
All recently referenced code pages in the active LRU are excluded from reclaim.
Code pages are paged in on demand: for any sufficiently large binary just some code pages will be mapped right after start. Given enough memory pressure, the active LRU will be shrunk often enough to push code pages out of the active LRU and then they are eventually evicted.
I'm not aware how it works in modern MGLRU, but I would assume similar ovservable behevior holds.
Active code pages with the "Accessed" bit set are specifically excluded from being moved to inactive LRU when active LRU is being shrunk. Instead, they are moved to the head of the active LRU. All other pages are moved from active LRU to the inactive LRU unconditionally, even if they have the "Accessed" bit set. So, no, the active code pages are not going to be eventually evicted.
I think we are saying the same thing. Active code pages are not evicted; non-active code pages (that is, pages that had not been accessed since last time of active LRU shrinking) may be evicted.
Under memory pressure the system will free the ram used for code pages because it can always load them back from the executable on disk. It’s the same virtual memory mechanism as swap but without needing dedicated swap space. The op is saying disabling swap doesn’t prevent long pauses during memory pressure, because the system just swaps code out instead of dynamically allocated memory.
There are good sibling replies answering your question, but to give a specific example: if your application writes something but then doesn't need it anymore. Since we're talking about go, perhaps a data structure you need also holds a reference to data you don't need.
When the kernel needs memory, it goes hunting for a page it can discard. But since that's transparent, the kernel can only discard a page if it knows it can get it back (after all, it's still got valid data on it, and maybe you'll try access it again later).
If there's swap, a page full of stale/unneeded data can be written out to swap. But if there's no swap, your page of "dangling data that you'll never use, but is still valid & referenced" can't be discarded; the kernel doesn't know you won't want it later, and it can't recreate the page if it throws it away.
So like sibling said, at that point it has to find other pages it can evict from memory, ones that _do_ have somewhere persistent they can be written out to. Pages loaded from binaries on disk satisfy that, so those will get dropped instead.
For JIT languages most of the code is dynamically generated, not available to be loaded back from the disk. Morealso, the allocations do happen at large chucks where the oom_killer also happens, so it's significantly less of an issue.
That doesn't necessarily improve matters. Now any code that gets JIT'd is also going to be permanently pinned in memory too, no matter whether it'll get used again, if you don't have swap.
But only in 2028 because all the wafers rotting away in the warehouses are HBM which are too expensive to package or integrate into consumer electronics with tighter margins. Looking forward to having enthousiast-grade GPUs with HBM again though.
I find it useful to use the term "determinate" (determinacy) [1,2] when dealing with concurrent or parallel computing. That puts the focus on equality of outcome (same result), leaving open "deterministic" to mean the exact same sequence of events (schedule).
And it tracks with Determinate Systems, the Nix company. Nix is also focused on giving you the same results every time even if the build process is chaotic inside.
I used Bitlbee until not that many earth rotations ago - it's an IRC daemon that provides bridging of various other chat protocols. It worked way better than you would expect, especially back in the day when you could use XMPP to connect to the various proprietary chat services.
I didn't stop using it until I stopped caring about those chat services.
reply