How sendfile and splice Work: Pipe Buffers, Page References, and Why Zero-Copy Is Not Zero Work
A static file server that uses read() into a user buffer and write() onto a socket copies every byte twice through the CPU: page cache to user memory, then user memory to a socket send buffer. On a busy origin, that tax is file size times request rate times cache-line traffic.
sendfile(2) and splice(2) ask the kernel to move data between descriptors without mapping the payload into user space. The bytes stay in kernel pages. What moves is a reference. This article covers pipe buffers, sendfile as a splice case, and what zero-copy does and does not remove. See also How the Linux Page Cache Works, How mmap Works, How io_uring Works, How System Calls Work, and How TCP Works.
The copy that used to be mandatory
read(file, buf, n)finds or faults page-cache pages, then copies intobuf.write(sock, buf, n)copies into kernel socket buffers and skb fragments.- The NIC DMA engine pulls those fragments onto the wire.
The application did not need to inspect the bytes, yet the payload crossed user space twice.
The contract
sendfile(out_fd, in_fd, offset, count) transfers up to count bytes from the input, starting at *offset, to the output. Linux may return fewer bytes, so callers loop. Since Linux 5.12, an output pipe uses splice semantics.
splice(fd_in, off_in, fd_out, off_out, len, flags) moves data where at least one descriptor is a pipe. tee(fd_in, fd_out, len, flags) duplicates pipe buffers without consuming the source. vmsplice plants user iovecs into pipe buffers. None promise atomic whole-range transfer, absent metadata work, absent checksums, or copy-free TLS.
Pipe buffers are page references
A Linux pipe is a circular array of struct pipe_buffer. Its default capacity is 16 pages, or 64 KiB with 4 KiB pages; F_SETPIPE_SZ can raise it to a sysctl cap. A slot contains a page reference, offset, length, and release operations.
File-to-pipe splice normally references page-cache pages and increments their reference counts. It does not memcpy the file bytes. Pipe-to-socket splice can attach those pages as skb fragments when scatter-gather DMA works. Pipe-to-pipe splice moves slots and ring indices, not bytes.
sendfile is splice with the pipe hidden
Modern Linux can treat file-to-socket sendfile as file to an internal pipe-like actor to socket. There is no user buffer for rewriting HTML, adding a template header, or compressing. Servers therefore write headers separately and sendfile the body, often using TCP_CORK or MSG_MORE.
A compact file-to-socket loop
int in = open("video.mp4", O_RDONLY);
off_t off = 0;
for (;;) {
ssize_t n = sendfile(sock, in, &off, 1 << 20);
if (n < 0) { if (errno == EINTR) continue; break; }
if (n == 0) break;
}Short transfers are normal. EAGAIN means the socket or TCP window is full. The offset prevents a retry from resending the prefix.
splice where sendfile will not take the path
int p[2];
pipe(p);
fcntl(p[0], F_SETPIPE_SZ, 1 << 20);
fcntl(p[1], F_SETPIPE_SZ, 1 << 20);
for (;;) {
ssize_t n = splice(in_sock, NULL, p[1], NULL, 1 << 16,
SPLICE_F_MOVE | SPLICE_F_MORE);
if (n <= 0) break;
splice(p[0], NULL, out_sock, NULL, n,
SPLICE_F_MOVE | SPLICE_F_MORE);
}SPLICE_F_MOVE is a hint, not a guarantee. SPLICE_F_NONBLOCK applies to splice itself, while SPLICE_F_MORE resembles MSG_MORE. tee can feed a logger and a forwarder; pages remain shared until both consumers release them.
vmsplice is a different bet
vmsplice can place user pages in pipe buffers, then splice them to a socket without another payload bounce. Those pages must remain stable and mapped until the kernel is finished; double-buffering is common. The zero-copy direction is user to pipe. MSG_ZEROCOPY instead reports buffer release through the error queue.
What zero-copy does not remove
- Page-cache fill: cold files still require storage I/O into DRAM; sendfile rides the page cache.
- DMA: disk-to-memory and memory-to-NIC movement remains; the CPU is simply not the mover.
- Protocol work: TCP bookkeeping, checksums, headers, and segmentation decisions remain.
- Encryption: kernel TLS can work on supported paths, but user-space TLS normally needs an owned plaintext buffer unless kTLS splice support is used.
- Forced copies: incompatible devices, filters, some FUSE paths, and byte-transforming consumers can fall back to copying.
- Syscalls: tiny loops still pay crossings; larger transfers and io_uring operations amortize them.
Where it breaks
Concurrent truncation can cause a short transfer. A reset peer can leave pages in a pipe, so drain or close both ends. Oversized pipes can pin page-cache pages and create memory pressure. User-space TLS without kTLS silently restores the copy.
Misconceptions
- Zero-copy means nothing moves. Storage, RAM, and the wire still move bits; the CPU payload copy disappeared.
- sendfile works for any two descriptors. Supported pairs are incomplete; splice plus a pipe is the general tool.
- SPLICE_F_MOVE always moves. It is only a hint.
- mmap is the same optimization. mmap attaches page-cache pages to a user address space; sendfile never gives the process those addresses. See How mmap Works.
- io_uring replaces splice. It can submit operations with fewer syscall crossings, but the page cache and pipe remain the substrate.
Takeaways
sendfile and splice move page references between kernel objects, avoiding a user-space bounce buffer. The pipe is the transferable slot array. Zero-copy describes CPU payload memcpy, not I/O, DMA, encryption, or syscalls. When a path must inspect or transform bytes, the copy returns and that is the correct trade.