How Dynamic Linking Works: PLT, GOT, and Why the First Call Is Indirect

A call to puts in a normal C binary is not a direct jump to libc. The machine code in the executable names a procedure-linkage-table stub. That stub loads an address from the global offset table. Until something fills that slot, the address points back at the stub, and the dynamic linker has to finish the job.

This is the runtime half of linking. The static linker (ld) already chose layouts, emitted relocations, and recorded which shared objects are required. It did not, on a normal build, copy libc into the executable. execve maps the file; the interpreter named by PT_INTERP resolves symbols. Virtual memory explains how those mappings exist. It does not explain why the first call is indirect and the second call is not.

What the static linker leaves behind

Compile a one-call program and the dynamic section is the contract:

int main(void) { puts("hi"); return 0; }

readelf -d on a current gcc PIE build shows the fields the loader actually walks:

  • NEEDED lists libc.so.6. That is a soname, not a path. The loader searches the cache, RPATH/RUNPATH, and the default paths.
  • STRTAB and SYMTAB are the dynamic symbol table. Strip can drop .symtab and the binary still links. It cannot drop these.
  • GNU_HASH is the hash table used for symbol lookup. A linear scan of the symbol table would be the wrong shape for a process that pulls in libc.
  • RELA and JMPREL are two relocation streams. Data and eager function references live in .rela.dyn. Lazy function slots live in .rela.plt.
  • PLTGOT is the address of .got.plt.

On that binary, puts is a R_X86_64_JUMP_SLOT at a GOT entry. __libc_start_main is a R_X86_64_GLOB_DAT in .rela.dyn. The split is the whole design: some symbols must be real before any user code runs, and some can wait until the first call.

Who runs before main

The kernel's ELF loader maps each PT_LOAD segment, honors the interpreter, and transfers control to the interpreter, not to the executable's e_entry. The auxiliary vector records the real entry as AT_ENTRY, the program headers as AT_PHDR, and the interpreter's load bias as AT_BASE.

On Linux that interpreter is ld-linux-x86-64.so.2 (or the architecture equivalent). It relocates itself first, because it is a shared object too. Then it loads each DT_NEEDED object, breadth-first, and builds a link map: a list of objects with load bias, symbol tables, and search scopes. Relative relocations (R_X86_64_RELATIVE) are just load_bias + addend. Symbol relocations require a search.

Search order is why interposition works. LD_PRELOAD objects are consulted first, then the executable, then dependencies. The first strong global definition wins. A STB_WEAK definition is a fallback. A STV_HIDDEN or internal symbol is not in the dynamic table, so another library cannot replace it. -Bsymbolic tells a shared object to prefer its own definition, which is faster and also a way to break a plugin that expected to override malloc.

After relocations and INIT / .init_array functions, the loader jumps to AT_ENTRY. That is usually the CRT, which calls __libc_start_main, which calls main. The symbol for __libc_start_main was a GLOB_DAT, so it is already resolved. puts may not be.

The PLT stub and the GOT slot

x86-64 lazy PLT entries are 16 bytes. Disassembly of the demo looks like this:

0000000000001020 :
  ff 35 ca 2f 00 00   push   GOT+8(%rip)    # link_map pointer
  ff 25 cc 2f 00 00   jmp    *GOT+16(%rip)  # _dl_runtime_resolve
  0f 1f 40 00         nop

0000000000001030 :
  ff 25 ca 2f 00 00   jmp    *puts@GOT(%rip)
  68 00 00 00 00      push   $0x0            # reloc index in .rela.plt
  e9 e0 ff ff ff      jmp    PLT0

The first GOT reserved slots are not function addresses. Slot 0 ties the table to the dynamic section. Slot 1 is the link_map pushed by PLT0. Slot 2 is the resolver. The puts slot starts out holding 0x1036, the address of the push $0x0 instruction inside its own PLT entry. That is visible in .got.plt before the process starts: the slot does not contain a libc address.

The first call therefore does this:

  1. call puts@plt enters the stub.
  2. The indirect jump through the GOT lands on the push of the relocation index, not on puts.
  3. The stub jumps to PLT0, which pushes the link map and jumps through GOT[2] into the dynamic linker's resolver.
  4. The resolver hashes puts, finds the definition in libc, adds libc's load bias, writes that absolute address into the GOT slot, and transfers to it.

The second call still enters the PLT, but the indirect jump now lands in libc. The resolver is not on the path. The stub remains because the compiler emitted a call to the PLT, and because the PLT is also the canonical address used by auditors (LD_AUDIT) and, in some builds, by function-pointer identity.

-fno-plt changes the call site itself to an indirect call through the GOT and resolves those slots up front. There is no lazy stub. The tradeoff is an indirect branch at every call instead of a direct call into a stub that becomes a single indirect jump after the first hit. Both are still dynamic linking.

Eager binding and RELRO

Lazy binding keeps the writable window that attackers want. A GOT slot that the loader will later trust as a function pointer is a write target if the process has a memory-corruption bug.

Partial RELRO, the usual -z relro default, relocates .got and then mprotects the data GOT read-only. .got.plt stays writable so the resolver can fill jump slots on demand.

Full RELRO is -z relro -z now, or LD_BIND_NOW=1 in the environment. DF_1_NOW asks the loader to process .rela.plt before transferring to the program, the same way it processes GLOB_DAT. After that, the loader can make the PLT GOT read-only too. Startup pays every symbol lookup, including ones the process might never call. The writable function-pointer table goes away.

This is also why a setuid binary ignores LD_PRELOAD and most loader environment variables. Interposition is a feature for debuggers and allocators, and a privilege boundary the loader refuses to cross.

Copy relocations and address identity

A non-PIE executable that references an imported data object cannot reach it with a GOT load if the compiler emitted a direct absolute address. The static linker allocates space in the executable and emits R_X86_64_COPY. At startup the dynamic linker copies the object from the shared library into that space. The library's own references are then pointed at the executable's copy. Two copies would be a bug; the copy relocation forces one.

That trick does not compose with interposition, and it is a reason PIE became the default. Position-independent executables reference data through the GOT, so the copy step is unnecessary. Function addresses have a related identity problem: the pointer you get from &puts in the executable may be the PLT entry, while code inside libc sees the real symbol. Comparing function pointers across DSO boundaries is not an ABI guarantee.

What dlopen adds

dlopen is the same loader, invoked late. RTLD_NOW matches bind-now. RTLD_LAZY leaves jump slots unresolved. RTLD_LOCAL keeps the new object's symbols out of the global search scope; RTLD_GLOBAL publishes them for later lookups. dlsym with RTLD_DEFAULT walks the scope of the caller. A plugin that expected to override malloc only works if it was loaded into a scope the allocator's lookup will see, and if libc was not bound symbolically.

Misconceptions

Dynamic linking is not "the kernel links the binary." The kernel maps segments and jumps to PT_INTERP. Symbol search, the hash table, and init arrays belong to the userspace loader. A page fault on the first call to puts can still happen, because libc's text page may be cold, but that fault is demand paging, not binding. Binding is the write to the GOT slot.

A statically linked binary is a different contract: no PT_INTERP, no JUMP_SLOT, libc's text copied or archived in. Shipping a static binary does not make LD_PRELOAD work. Shipping a dynamic binary does not mean every call pays a hash lookup; only the unresolved ones do, and only once per slot.

Shared pages of libc are why a second process is cheap to start. Fork and copy-on-write explain the page sharing. They do not fill the child's GOT. The child inherited the parent's already-resolved slots, or it will resolve its own if it was exec'd.

Takeaways

The executable stores a relocation and a stub, not libc's address. The dynamic linker is a shared object the kernel enters before main. Lazy PLT entries bounce through a GOT slot that initially points at the stub; the resolver overwrites that slot. Bind-now and full RELRO close the writable window by doing that work at startup. Interposition, copy relocations, and function-pointer identity all fall out of the same search order.

Related: How System Calls Work covers the execve boundary this loader runs after. How Virtual Memory Works covers the page tables behind PT_LOAD. How mmap Works is the mapping primitive those segments use. How fork and Copy-on-Write Work is why libc text stays shared across processes after this resolution has happened.

Previous Post