How the Linux Buddy and SLUB Allocators Work: Orders, Partial Slabs, and Why malloc Is Not a Page Fault

A driver that calls kmalloc(256, GFP_ATOMIC) from an interrupt handler and a process that calls malloc(256) are not talking to the same allocator. The first asks the kernel for a small object that must already have a physical page behind it. The second asks a userspace library for a virtual address. The frame often does not exist until the first write faults.

Linux splits that work across two kernel layers. The buddy allocator owns free physical page blocks. SLUB, the default slab allocator since 2.6.23, carves pages it borrows from the buddy into fixed-size objects. Userspace malloc sits above both and is not either of them. This article is about those two kernel layers: how a block is split and merged, how a slab hands out an object without taking a lock, and where a page fault actually enters the picture. Address translation itself is covered separately in How Virtual Memory Works.

What each layer is allowed to give you

The page allocator, implemented as a buddy system, returns a block of physically contiguous pages. The block size is always a power of two pages. Callers that need a raw frame, a compound page for a huge-page pool, or the backing store for a slab use alloc_pages() or a wrapper.

SLUB returns an object of a size the cache was created for. kmalloc() is a set of those caches. kmem_cache_alloc() is the same machinery for a type the subsystem registered, such as dentry or inode objects. The pointer is kernel-virtual and, for ordinary kmalloc, the underlying pages are physically contiguous, which is why drivers use it for small DMA buffers when the device can address that memory.

vmalloc() is a third path, not a third buddy. It allocates order-0 pages from the buddy and maps them into a contiguous kernel virtual range. The physical frames need not be neighbors. That is why a large vmalloc can succeed when a high-order alloc_pages fails, and why the result is a poor fit for devices that require a single physical range.

Orders and the buddy relation

An order-0 block is one base page. Order 1 is two physically contiguous pages, order 2 is four, and order n is 2n pages. With the common 4 KiB base page, order 0 is 4 KiB and order 10 is 4 MiB. The maximum order is a kernel configuration (MAX_PAGE_ORDER, historically orders 0 through 10). It is not a promise that a block of that size is free.

Free blocks live on lists kept per memory zone, per migrate type, and per order. A zone is a physical range with different usability: DMA for ancient 16 MiB-limited devices, DMA32 for 32-bit device addresses, Normal for ordinary kernel memory, and on some machines Movable or Device. Migrate types (Unmovable, Movable, Reclaimable, and others) exist so the kernel can group pages that are allowed to move. Mixing a permanently pinned kernel object into a region of movable anonymous pages is how high-order allocations later fail.

The buddy of a block is the unique adjacent block of the same order that, together with it, forms an aligned block of the next order. Given a page-frame number and an order, the buddy frame number is the original XOR 2order. Alignment is the point. Two neighboring order-0 pages are buddies only if their combined pair is aligned to 8 KiB. Freeing one page next to an unrelated free page does not create an order-1 block unless that alignment holds.

Allocation walks up from the requested order:

  1. Look for a free block at that order in the allowed zone and migrate type.
  2. If none exists, take a block from the next higher order and split it in half.
  3. Put one half on the free list of the lower order. That half is the buddy.
  4. Repeat until the remaining half is the requested order, then return it.

Free runs the same relation backwards. If the buddy of the block being freed is also free and of the same order, the allocator removes the buddy from its free list and merges the pair into an order+1 block. Merging repeats until the buddy is busy or the maximum order is reached. That merge is the entire defense against permanent fragmentation. It only works when both halves are free at the same time.

Order-0 traffic does not usually touch those lists on every call. Each CPU keeps a small magazine of single pages, the per-CPU pageset. A fault that needs one anonymous frame can often be satisfied from that magazine. The shared buddy lists are the slow path, used to refill the magazine or to serve orders above 0.

Watermarks, reclaim, and why GFP flags exist

Each zone tracks three watermarks: min, low, and high. Above high, the zone is comfortable. Between low and high, the kernel wakes kswapd to reclaim in the background. Below min, ordinary allocations are not supposed to dip further unless the caller is allowed to use reserves.

GFP flags say how hard this attempt may try, and whether the caller can sleep. GFP_KERNEL may enter direct reclaim and wait. That is the right flag for most kernel data structures allocated from process context. GFP_NOWAIT will not reclaim in the caller's context; under pressure it fails. GFP_ATOMIC is the non-sleeping flag that may also touch reserves, used from interrupt and other atomic context, with the expectation that failure is possible. Kernel documentation treats GFP_NOFAIL as a last resort that loops until the allocation succeeds, and warns against it for large orders.

A failed high-order allocation is often not "out of memory" in the sense of no free pages. /proc/buddyinfo shows the shape of the free lists: many order-0 pages and almost nothing at order 8 or 9. Compaction moves movable pages to rebuild higher-order blocks. It cannot move a page that a driver has pinned for DMA. That is why migrate types exist, and why a machine with gigabytes free can still refuse a 1 MiB contiguous allocation.

What SLUB adds on top of a page

A 256-byte object placed in its own order-0 page wastes most of the page. SLUB asks the buddy for one or more pages, carves them into equal slots, and keeps the free slots on a list. The cache is a kmem_cache: one size, one alignment, one constructor policy. The number of pages per slab is the cache's order, chosen so objects pack without a large tail.

The fast path is per CPU. Each CPU holds a current slab for the cache. A free object on that slab is popped with a compare-and-swap against the freelist pointer. No zone lock, no node lock. The freelist pointer is stored in the free object itself, not in a side array, which is part of why SLUB's metadata is smaller than the older SLAB design. The original SLAB allocator was removed as a selectable option in Linux 6.8; SLUB is the implementation those kmalloc calls hit on a current kernel.

When the CPU slab has no free object, SLUB takes a slab from the partial list: slabs that are neither full nor empty. With CONFIG_SLUB_CPU_PARTIAL, which is the usual configuration, the CPU also keeps a short private partial list so it does not bounce on the node list for every refill. If nothing partial exists, SLUB allocates a fresh slab from the buddy, initializes the freelist, and that slab becomes the CPU slab.

Freeing prefers the same CPU slab. An object returned to a slab that still has other busy objects stays partial. A slab that becomes completely free can be returned to the buddy, which is how kernel object traffic eventually coalesces pages. Under churn, SLUB would rather keep a partial slab than hand the page back and immediately ask for it again.

kmalloc does not create a cache per call. It selects a prebuilt size class: powers of two, plus a few intermediates such as 96 and 192 bytes that match common structure sizes and avoid rounding a 80-byte object up to 128. Requesting 100 bytes lands in the 128-byte cache. The unused 28 bytes are internal fragmentation, paid so the freelist stays uniform. Above the largest kmalloc cache, the call falls through to the page allocator. kvmalloc() tries kmalloc first and vmalloc if the contiguous attempt fails, and the caller must free with kvfree.

Where userspace malloc actually enters

glibc malloc, and the allocators that replace it, keep their own arenas, size classes, and free lists in the process. Small allocations come out of a heap region grown with brk. Larger ones often come from an anonymous mmap. Neither system call is a request for a SLUB object. Both install virtual memory areas. The kernel does not have to allocate a frame for every page of that mapping at creation time.

The frame appears on the fault. A store to a fresh anonymous page is a minor fault: the fault handler allocates an order-0 page, typically from the per-CPU pageset, zeros it or copies a copy-on-write source, and installs a present page-table entry. That order-0 allocation is the buddy, or its per-CPU cache. It is not SLUB. SLUB is for kernel objects. The userspace object lives in a frame the fault path just obtained, inside a virtual range the C library already owned.

That split is why RSS and the malloc statistic diverge. The library can reserve a megabyte of virtual heap and report it as allocated while /proc/pid/smaps still shows those pages not present. It is also why fork cost shows up later, on the write fault, rather than inside malloc. The copy-on-write path, described in How fork and Copy-on-Write Work, allocates a new frame at the fault and copies the old one. The buddy supplies the frame. The C library is not involved.

File-backed faults are the other customer of the same page allocator. A read that misses the page cache allocates a frame, fills it from storage, and indexes it. The page cache is a user of the buddy, not a replacement for it.

What the counters are actually counting

/proc/buddyinfo is one row per zone. Each column is the count of free blocks at that order, not the count of free pages. A "4" in the order-3 column means four free 32 KiB blocks on a 4 KiB system, 128 KiB total, and those blocks can still be split if a lower order runs out. Watching only MemFree hides a machine that cannot satisfy order 9.

/proc/slabinfo and /sys/kernel/slab/ show each cache: object size, objects per slab, pages per slab, and how many slabs are active. A growing dentry or radix_tree_node cache is kernel metadata, not the process heap. Reclaim can shrink some of those caches; it cannot invent a higher-order page out of pinned ones.

Misconceptions

malloc does not call kmalloc. The names rhyme. The paths do not meet until a fault or a page-cache fill asks the buddy for a frame.

The buddy allocator is not a general heap. It cannot hand out 100 bytes. Anything smaller is a slab object or a userspace allocator sitting on top of pages.

SLUB is not the userspace allocator, and it is not a garbage collector. A language runtime's bump allocator, described in How Garbage Collection Works, manages objects inside virtual pages the runtime already mapped. Collection finds unreachable objects. It does not coalesce physical page blocks. That job stays with the buddy when those pages are finally freed.

A successful kmalloc of a small size does not mean a fresh page was taken from the zone. The usual case is a freelist pop on a slab this CPU already owns. The buddy is on the path only when a new slab is required.

Free memory in /proc/meminfo is not contiguous memory. High-order allocations fail from fragmentation first. Compaction and migrate types are the tools aimed at that failure; adding swap does not by itself rebuild a 4 MiB contiguous block.

Takeaways

Physical contiguity is an order: 2n pages, split by halving, merged only with the aligned buddy. Zones and migrate types constrain which free list is legal. Watermarks and GFP flags decide whether the caller may sleep, reclaim, or fail. SLUB turns those pages into same-sized objects, with the fast path on a per-CPU freelist and the slow path on partial slabs. Userspace malloc allocates virtual address space. The page fault is the moment a frame is requested, and that request is an order-0 page allocation, not a slab allocation.

Next Post Previous Post