How Linux Interrupts Work: Hardware IRQs, Softirqs, and Why Threaded Handlers Exist
A NIC that just received a frame cannot wait for the current process to finish a syscall. The device raises a hardware interrupt. The CPU stops the instruction stream that was running, switches to kernel mode on that core, and runs a handler that the kernel registered for that line. When the handler returns, the interrupted context continues as if nothing happened, except that time passed and some work was queued.
That path is not a signal and it is not a system call. Signals wait for a return-to-userspace check. System calls are requested by the current instruction. An interrupt is requested by hardware, preempts whatever was on the core, and does not return a value to that instruction. This article is about that path: how a vector becomes an irq_desc, how level and edge flow handlers differ, why hard-IRQ context cannot sleep, what softirqs and ksoftirqd are for, and why modern drivers use threaded IRQs.
Related reading: How Linux Signals Work, How System Calls Work, How Linux CFS Works.
The hardware contract
On x86-64 a device does not poke the CPU directly. It asserts a line on an interrupt controller: historically the 8259 PIC, on modern machines the Local APIC plus I/O APIC or MSI/MSI-X messages. MSI-X lets a device send a write to an APIC address that encodes a vector and a destination CPU. The local APIC then interrupts that core.
The CPU consults the interrupt descriptor table, saves a minimal frame, and enters the architecture entry stub with interrupts masked on that core. The stub identifies the vector, switches to the kernel stack if the CPU was in user mode, and calls into the generic IRQ layer. After the handler, irq_exit may run pending softirqs before the core returns to the interrupted context.
A vector is a small integer the APIC understands. A Linux IRQ number is a software identifier. irq_domain maps one to the other so a GPIO expander, a PCI device, and a local timer can all live in the same irq_desc table.
irq_desc is the kernel object
Each IRQ number names an irq_desc. That structure holds the flow handler (handle_level_irq, handle_edge_irq, handle_fasteoi_irq, handle_percpu_irq), a pointer to an irq_chip that knows how to mask, ack, and EOI the controller, and a linked list of irqaction entries installed by request_irq or request_threaded_irq.
The architecture entry path ends in generic_handle_irq_desc, which calls desc->handle_irq(desc). The flow handler, not the driver, owns ack and unmask policy. That split is why a driver can be written once against request_irq while the same line is level-triggered on one SoC and edge-triggered on another.
Level versus edge
A level-triggered line stays asserted until the device is serviced. handle_level_irq masks the line, runs the action list, then unmasks. Leaving the device still asserting after unmask immediately retriggers the IRQ. That is the correct contract for many SoC devices.
An edge-triggered line fires a pulse. The controller latches it. handle_edge_irq acks the latch, runs the actions, and if another edge arrived while the handler ran, loops. If the handler is already running on another CPU, the new edge is recorded as pending and the line may be masked until the running handler finishes. Missing that loop drops interrupts under load.
MSI-X and many APIC-delivered interrupts use an EOI-only flow (handle_fasteoi_irq): run the handler, then tell the APIC the interrupt is done. Per-CPU interrupts such as the local timer skip the desc lock and use handle_percpu_irq.
What hard-IRQ context may do
The primary handler runs with the local interrupt disabled for that line's flow, on the interrupted CPU's kernel stack, with preemption off. It must not sleep. It must not call mutex_lock, kmalloc(GFP_KERNEL), or copy large buffers to user space. It may read status registers, ack the device, copy a small descriptor into a ring the rest of the kernel already owns, and return IRQ_HANDLED, IRQ_NONE, or IRQ_WAKE_THREAD.
Shared IRQs are a linked list of actions on one irq_desc. Every handler on a shared line must check whether its device actually raised the interrupt and return IRQ_NONE if not. Returning IRQ_HANDLED for someone else's device hides storms and breaks spurious-IRQ detection.
/proc/interrupts counts per-CPU deliveries. A line that only ever fires on CPU 0 is either pinned by affinity or is a poorly distributed MSI-X vector. irq_set_affinity and /proc/irq/N/smp_affinity move the destination APIC target. They do not move work that the handler then queues onto a single global lock.
Softirqs are the high-frequency bottom half
Networking and the block layer cannot finish a packet or a completion in the hard-IRQ handler. They raise a softirq: a statically numbered, per-CPU deferred callback. NET_RX_SOFTIRQ, NET_TX_SOFTIRQ, BLOCK_SOFTIRQ, RCU_SOFTIRQ, and a handful of others are the current set. New code is not supposed to allocate another one.
raise_softirq sets a bit in the local pending mask. After the hard-IRQ handler returns, irq_exit runs do_softirq if bits are set. Softirqs run with hard interrupts enabled, still in interrupt context: they still cannot sleep. They can be interrupted by new hard IRQs, which may raise the same softirq again.
That reentrancy is why a long NET_RX loop can starve user tasks. The kernel bounds consecutive restarts (MAX_SOFTIRQ_RESTART, MAX_SOFTIRQ_TIME). When the budget is spent, leftover bits are handed to ksoftirqd/N, a per-CPU kernel thread. ksoftirqd is process context. It can be scheduled against. That is the safety valve, not the fast path.
NAPI sits on top of this for NICs. After the first packet interrupt, the driver disables further IRQs from that queue and polls the ring from NET_RX_SOFTIRQ. When the queue drains, IRQs are re-enabled. The interrupt becomes a wakeup for a polling loop instead of a per-packet trap. That is why a quiet NIC still uses IRQs and a loaded NIC tries not to.
Threaded IRQs are the driver default
request_threaded_irq(irq, handler, thread_fn, flags, name, dev) splits the work. handler runs in hard-IRQ context. If it returns IRQ_WAKE_THREAD, the kernel wakes a dedicated irq/N-name thread that runs thread_fn in process context. That thread may sleep, allocate, and talk to the rest of the kernel under ordinary locks.
IRQF_ONESHOT keeps the line masked until the thread finishes. That is the usual choice for level-triggered devices whose status must stay latched until the thread reads it. Without ONESHOT, the line can fire again while the thread is still running, which is correct only if the hard-IRQ handler can tolerate that.
Force threading (threadirqs on the kernel command line, and the default for many PREEMPT_RT configurations) turns even a handler registered with request_irq into a primary-plus-thread pair. Real-time kernels do this so a device ISR cannot run unbounded on a CPU that a deadline task needs. The cost is latency jitter of a wakeup instead of a few microseconds in hard-IRQ context.
Tasklets still exist as a serialized softirq wrapper. New drivers should not add them. Workqueues are the general sleepable deferral mechanism when the work is not bound to a specific IRQ line. Threaded IRQs are the mechanism when the work is that line.
A short timeline of one NIC IRQ
1. The NIC writes an MSI-X message. The local APIC interrupts CPU 3 with vector V.
2. CPU 3 enters the IDT stub, saves the interrupted registers, and calls the generic layer with the mapped Linux IRQ.
3. handle_fasteoi_irq runs the action. The driver's hard-IRQ handler sees a filled RX descriptor, schedules NAPI, and returns IRQ_HANDLED. The APIC receives EOI.
4. irq_exit notices NET_RX_SOFTIRQ. NAPI polls the ring, builds skbs, and hands them to the protocol stack. If the poll budget expires, NAPI reschedules itself.
5. If softirq time is exhausted, ksoftirqd/3 continues the work. User processes on CPU 3 can run between those slices.
6. The interrupted thread, which may have been in user mode or in a syscall, resumes. It did not request this work and it does not see a return value from it.
Misconceptions
Interrupts are not signals. A signal is a pending bit delivered on the way back to userspace. An interrupt preempts immediately on the CPU that received the vector.
Interrupts are not system calls. The interrupted instruction did not ask for the kernel and will be restarted or continued, not completed with an error code from the device.
Top half versus bottom half is not two functions with the same privileges. Hard-IRQ context cannot sleep. Softirq context cannot sleep. Threaded IRQ and workqueue context can. Mixing those rules is how drivers deadlock on a lock the hard-IRQ path already holds.
Disabling interrupts with local_irq_disable is a per-CPU operation. It does not stop other cores from running the same shared handler. Spinlocks that are taken in IRQ context must be the _irq variants on the process-context side, or the process context on this CPU can deadlock against its own interrupt.
A high count in /proc/interrupts is not automatically a problem. A high count plus ksoftirqd using a full core is a problem: the machine is interrupt-bound and the softirq safety valve is the main consumer of that core.
Takeaways
A Linux interrupt is a vector that the generic IRQ layer turns into an irq_desc flow handler and a list of actions. The flow handler owns ack and mask policy for level, edge, and EOI controllers. The hard-IRQ action must be short and unsleepable. Softirqs absorb high-frequency leftover work on the same CPU and overflow into ksoftirqd. Threaded IRQs give drivers a sleepable half without inventing a new softirq. NAPI is the networking special case that turns a storm of packet IRQs into a poll loop. Signals and syscalls share the kernel entry machinery with this path. They do not share its timing or its rules.