How NAND Flash and the FTL Work: Pages, Erase Blocks, and Why a Small Write Rewrites a Block

A database that updates a 4 KiB page does not update a 4 KiB cell on the SSD. The host sends a logical block address. The drive's flash translation layer picks a free NAND page somewhere else, programs that page, and marks the old physical page invalid. The block that held the old page cannot be reused until every still-valid neighbor has been copied out and the whole block has been erased. That indirection is why a small logical write can become a much larger physical rewrite, and why the page cache's idea of a dirty page is not the device's idea of a program.

This article is about that device contract: NAND pages and erase blocks, the mapping table that hides them, and the garbage collection that recycles them. It is not a second tour of LSM compaction or the Linux page cache. Those sit above this layer. The single search intent is how NAND flash and the flash translation layer turn an overwrite into an out-of-place program plus a later erase.

What the array can actually do

NAND flash is organized in pages and blocks, not in independently writable bytes. A page is the unit of program and of read. On current consumer and enterprise NAND a page is commonly 4 KiB to 16 KiB of user data plus a spare area for ECC bits, a logical address tag, and program status. An erase block is a run of those pages, often a few hundred, so the erase unit is megabytes rather than kilobytes. Exact sizes are a property of the die generation, not a constant the host can assume.

Three operations are not symmetric. A read returns a page. A program sets bits in a page that has been erased, and on ordinary NAND that page is not programmed again until the block is erased. An erase resets every page in the block to the all-erased state. There is no byte store. Changing one sector means programming a new page and treating the previous page as garbage.

A cell may hold one bit (SLC), two (MLC), three (TLC), or four (QLC). More bits per cell means more voltage levels, slower and less reliable programs, and fewer program/erase cycles before the cell is retired. A drive can still expose SLC-like behavior for a small cache: it programs some blocks in a faster mode, then folds that data into denser blocks in the background. The host write completed when the fast program and its mapping update were durable enough for the drive's power-loss policy, not when the final QLC fold finished.

The map the host never sees

The operating system addresses the drive in logical blocks, typically 512 bytes or 4 KiB. NAND addresses are physical page numbers inside a die, plane, and block. The flash translation layer is the firmware that keeps the map between them. A page-mapped FTL stores a logical-page to physical-page entry for every live logical page. That table is large: a 4 TiB drive with 4 KiB pages has on the order of a billion entries, so the map itself lives in DRAM on the controller, with a journal or checkpoint on NAND so a power cut does not lose the translation.

Hybrid schemes exist because a pure page map is expensive. A block map is smaller but punishes random updates. A common compromise keeps a block-level map for cold data and a page-level log for recent overwrites, then merges the log back. The host cannot tell which scheme is in use. It only sees that a logical block address still reads back the last write, even though the physical page changed.

An overwrite of logical page L looks like this:

host write LBA range covering L
  FTL allocates a free physical page P2
  program P2 with the new bytes and a tag for L
  update L2P: L -> P2
  mark previous P1 invalid
  acknowledge the write once the program and map update meet the drive's durability rule

Nothing in that path erased P1's block. Invalid is a mapping state, not a physical clear. The block still holds other live pages. Until garbage collection moves those, the space is occupied.

Garbage collection is the rewrite

Free pages come from erased blocks. When the free-block pool drops, the FTL picks a victim, copies every still-valid page to a new block, updates the map, and erases the victim. That copy is internal write traffic the host did not request. Write amplification is NAND bytes programmed divided by host bytes written. A sequential fill of empty space can sit near 1, plus metadata. A random overwrite of a full drive with little spare space can force the collector to move many live pages for each page the host invalidated.

Victim choice is the whole policy. Greedy selection takes the block with the fewest valid pages, which minimizes copied bytes now. It can also hammer hot blocks. Cost-benefit and hot/cold separation try to keep short-lived data in blocks that will soon be fully invalid, so the erase throws away pages instead of relocating them. Multi-stream hints and NVMe directives exist so a host that knows its lifetimes can aim different streams at different blocks. The FTL is not required to honor a hint, and most filesystems do not send one.

Over-provisioning is the spare the collector spends. The marketed capacity is smaller than the NAND on the board. A drive sold as 480 GB often has 512 GiB-class raw media, and factory bad-block reserve sits on top of that. The gap is invisible to the filesystem. It is why a drive that is "100 percent full" from the OS still has blocks to erase into, and why filling the logical space and then hammering random updates is the workload that exposes the worst amplification. Steady-state random write performance is a property of spare area and of how mixed the valid pages are, not of the sequential brochure number.

Wear leveling is a second mover

Each block survives a finite number of program/erase cycles. Dynamic wear leveling only steers new writes onto less-worn free blocks. That is not enough: a block full of cold data never re-enters the free pool, so the hot blocks wear out while the cold ones sit. Static wear leveling occasionally moves cold valid data so those blocks can be erased and reused. That move is more write amplification, paid to keep the cycle counts from diverging. When a block fails program or erase, the FTL retires it and draws on the spare pool. SMART attributes such as percentage used and media units written are the host-visible shadow of this accounting, not a direct view of the map.

TRIM is an invalidation the host can send early

If the filesystem deletes a file and never writes those logical blocks again, the FTL still believes the old pages are valid. It will copy them during garbage collection. Discard, called TRIM on SATA and deallocate on NVMe, tells the device that a logical range has no live data. The FTL can mark those physical pages invalid without waiting for an overwrite. That does not erase them immediately, and it does not guarantee the old bytes are unreadable before the next GC. It does stop the collector from treating deleted files as data it must preserve.

A filesystem that never discards, or a RAID layer that swallows discards, leaves the drive cleaning up after ghosts. Periodic batch discard is cheaper than issuing a discard on every unlink, and it is still a hint about liveness, not a flush.

What fsync is asking for

A successful write from the kernel's point of view means the device acknowledged the command. fsync pushes the page cache and the filesystem journal, then asks the block device to flush its volatile buffer. On an SSD that flush is a command to make the acknowledged data recoverable, which usually means the programmed pages and the mapping update that points at them must survive power loss. Enterprise controllers often keep a capacitor-backed buffer so they can finish that update after the rail drops. A consumer drive may acknowledge sooner and rely on a smaller protected region. The POSIX call cannot see which. It can only wait for the flush command to complete.

The flush does not run garbage collection to completion, and it does not turn a random logical update into an in-place NAND program. Alignment still matters: a 512-byte rewrite inside a 4 KiB NAND page, or a write that straddles two logical pages, becomes a read-modify-write in the FTL or in the filesystem. The page cache article explains why the kernel may not even issue that I/O until writeback. This layer explains why, once it is issued, the device still will not overwrite the previous physical page.

What software can and cannot change

Sequential large writes keep invalidations clustered, so a victim block is mostly dead and the copy set is small. Random 4 KiB overwrites scatter invalid pages across many blocks and are the case the spare area has to absorb. Log-structured engines take that constraint as a design input: they append, then compact, so the device sees longer sequential programs. Compaction has its own amplification. It is not a substitute for knowing the device's. A B-tree that dirties a leaf still presents the drive with an out-of-place NAND program even if the logical block address is reused.

Two limits are easy to miss. First, the FTL's map is per drive, or per namespace, not per file. Two tenants random-writing the same SSD share one free pool and one collector. Second, encryption, compression, and checksums in the host change the bytes programmed. They do not change the erase-before-reuse rule. A checksum mismatch after a torn mapping update is a power-loss bug at the FTL or at the filesystem, not evidence that NAND stored both versions in the same page.

Takeaways

NAND programs pages and erases blocks. The FTL hides that by mapping logical pages to physical pages and by acknowledging an overwrite only after a new page is programmed and the map points at it. Garbage collection is the later copy-and-erase that produces write amplification. Wear leveling moves cold data so cycle counts stay even. Discard is how the host marks logical ranges dead before they are overwritten. None of that is visible as a sector write in the page cache, and none of it is the compaction story an LSM tells above the device.

Related: How LSM-Trees Work is the software shape that tries to hand this device sequential programs. How Write-Ahead Logging Works is the durability log above the flush command. How the Linux Page Cache Works is why a dirty page may not have reached the FTL yet. How mmap Works covers file-backed pages that still write back through this same device.

Previous Post