How fork and Copy-on-Write Work: Page Tables, Shared Pages, and Why the Child Is Not Free

A prefork server that looked cheap in RSS before the first request can double its memory after workers start handling traffic. fork did not copy the heap at the moment of the call. It copied the page tables and marked the shared pages read-only. The copy happens later, one page at a time, on the first store. That delay is copy-on-write, and it is the reason a child process can exist without paying for a second copy of every byte the parent already has.

This article is about that contract: what fork and clone actually duplicate, how anonymous pages stay shared until a write, and why the kernel sometimes skips the copy. File-backed MAP_PRIVATE uses a related fault path; that mapping model is covered in How mmap Works. The address-space objects underneath both are in How Virtual Memory Works.

What the call returns

On Linux, glibc fork is a thin wrapper around clone. The kernel creates a new task_struct, gives it a new pid, and returns twice: 0 in the child, the child's pid in the parent. Both returns resume after the same instruction. From that point the two tasks are scheduled independently. See How Linux CFS Works for how each becomes a runnable entity.

fork is the special case of clone that asks for a new address space. CLONE_VM would share the parent's mm_struct instead, which is what threads do. CLONE_FILES, CLONE_FS, and the namespace flags choose other sharing. A container runtime passes a different set than a shell that wants a child process. The memory question in this article is the case without CLONE_VM.

What is duplicated, and what is not

The kernel does not walk the heap and memcpy it. dup_mmap walks the parent's virtual memory areas and builds a second page-table tree that points at the same physical pages. VMAs are copied as descriptors: start, end, permissions, backing file or anonymous flag. Open file descriptions are shared by default (the fd table is copied, the underlying struct file is referenced). Signal handlers are copied. Pending signals are not shared.

Physical pages of anonymous memory stay shared. The cost paid inside fork is the page-table copy, the new mm_struct, and the task structure, not a second copy of the resident set. A process with a multi-gigabyte heap and mostly untouched upper page-table levels still pays for every page-table page that must be duplicated so the child can later diverge.

Write-protecting both sides

A shared anonymous page that either side may write cannot stay mapped writable. For each present writable PTE, copy_present_pte clears the write bit in both the parent and the child and keeps both entries aimed at the same page frame. The VMA still says the range is writable. The hardware PTE does not. That mismatch is the trap.

parent PTE: pfn=0x1a00, present, read-only
child  PTE: pfn=0x1a00, present, read-only
VMA:          VM_READ | VM_WRITE | VM_MAYWRITE

Read-only VMAs do not need this treatment. Clean file-backed shared mappings can stay shared and writable only when the mapping is truly shared (MAP_SHARED); private writable file mappings are write-protected the same way anonymous pages are, and the fault then copies into an anonymous page. That file case is the mmap article. Here the pages are already anonymous.

The fault that does the copy

The child stores to a shared page. The CPU raises a write fault because the PTE is read-only. The kernel sees a write to a VMA that allows writes and routes the fault to the copy-on-write handler (do_wp_page in current kernels). The handler asks whether this process is the only remaining user of the folio.

If another task still maps the page, the kernel allocates a new folio, copies the old contents, points the faulting process's PTE at the new frame with write permission, and drops a reference on the old folio. The other process keeps the original frame and keeps a read-only PTE. The copy is one page (or one compound page, if the fault is on a huge page), not the whole address space.

write to child address A
  PTE read-only, VMA writable
  folio mapcount > 1
    new = alloc folio
    copy old -> new
    child PTE = new, writable
    parent PTE unchanged, still read-only

The parent's PTE stays read-only even after it becomes the exclusive owner. The next store in the parent faults again. This time the folio is exclusive, so the handler marks the existing PTE writable and returns. No second copy. That reuse path is why a parent that never writes after fork keeps the original pages, and why a parent that writes only a few of them pays only for those pages.

Why RSS lies until the first store

Tools that sum proportional set size (PSS) attribute a shared page fractionally to each mapper. Resident set size (RSS) often counts the full page for each. Right after fork, RSS can look like the heap was duplicated while PSS barely moved. After both sides write every page, PSS and RSS converge and the machine really does hold two copies. A prefork worker that dirties its heap, its malloc arenas, and its per-request buffers will pay that bill even though fork itself returned quickly.

Overcommit interacts with this. The kernel may allow the fork because the copy is not reserved page-for-page up front, depending on vm.overcommit_memory. The fault that later needs a free frame can fail if the machine is actually out of memory. A successful fork is not a promise that every future COW fault will allocate.

vfork and posix_spawn

vfork exists for the case that immediately execs. The child borrows the parent's address space and the parent stays stopped until the child execs or exits. There is no page-table duplication and no COW window, which is why a vfork child must not return from the calling function or touch the parent's stack. posix_spawn is the interface that can avoid the fork-and-write pattern entirely: file actions and a new image without exposing a child that mutates the old heap.

A multi-threaded parent makes plain fork sharper. Only the calling thread is duplicated. Mutexes held by other threads stay locked in the child, with no owner to unlock them. That is a POSIX rule, not a COW quirk. Libraries that fork from a multi-threaded process are expected to exec immediately, or to use posix_spawn.

Where the cost actually goes

Page-table duplication dominates a large process: every leaf page-table page that has present entries is copied, and higher levels are copied so the child has its own mm_struct. After that, each COW fault pays an allocation, a page copy, a PTE update, and a TLB invalidation on the faulting CPU. Parent and child do not share an address space, so the invalidation does not have to shoot down the other process's TLB for that virtual address. Huge pages make the copy larger and less frequent. Transparent huge pages that split under COW add a collapse-or-split choice the fault path has to get right.

File descriptors and the page cache are the parts that do not follow this rule. A shared file page in the page cache stays one page for every process that maps or reads it. COW on a private mapping copies out of that cache into an anonymous page. The cache page itself is not duplicated just because someone forked. See How the Linux Page Cache Works.

Misconceptions

fork copies all memory. It copies the page tables and write-protects shared anonymous pages. Bytes move on the fault, and only for pages that are still shared.

The child is free until it writes. The child is not free. Page tables, the task, and kernel bookkeeping are paid immediately. Untouched pages stay shared; that is a different statement.

Copy-on-write is the same mechanism as RCU. RCU publishes a new object and waits out old readers. COW breaks sharing of a physical page on a write fault. The names rhyme. The contracts do not. How RCU Works is the other one.

MAP_PRIVATE and fork are the same article. Both end in a write fault that may copy a page. Fork's subject is a new task and a duplicated address space. MAP_PRIVATE's subject is a private view of a file.

Takeaways

fork without CLONE_VM builds a second address space whose present PTEs alias the parent's physical pages, then clears write permission on both sides of anything either side might store to. The first write faults, copies if the folio is still shared, and reuses the page if this task is the last mapper. RSS that doubles after workers start is that fault path, not the fork call itself. If the child only exists to exec, vfork or posix_spawn skips the page-table copy that COW was invented to postpone.

Previous Post