TL;DR — io_uring removes the per‑operation syscall overhead of traditional asynchronous I/O by sharing memory with the kernel and using two ring buffers for submissions and completions. This lets applications achieve near‑zero context switch costs, dramatically improving throughput and latency for I/O‑bound workloads.

Traditional Linux asynchronous I/O mechanisms—aio_* syscalls, epoll, and libaio—have long been the go‑to tools for building high‑performance servers. Yet each of them still forces a system call for every operation, which adds latency, consumes CPU, and limits scalability. The io_uring interface, introduced in kernel 5.1 and refined ever since, flips this model on its head: instead of making a syscall for each I/O request, the application and the kernel share a pair of ring buffers in memory. Submissions are written into the submission queue (SQ) by user space, and the kernel processes them asynchronously, posting completions into the completion queue (CQ). The only time a syscall is needed is to notify the kernel that new entries have been added, and even that can be batched.

In this post we’ll walk through the internals of io_uring, show how to get started with liburing, examine real‑world performance numbers, and discuss production patterns that have proven successful.

How io_uring Works

Submission and Completion Queues

The heart of io_uring is two circular buffers:

  • Submission Queue (SQ) – an array of io_uring_sqe structures. Each entry describes a single operation (read, write, accept, etc.) and its parameters.
  • Completion Queue (CQ) – an array of io_uring_cqe structures. When the kernel finishes an operation, it writes a completion entry containing the result.

Both queues reside in memory mapped from the kernel, so user space can read/write them without any syscall. The application fills SQ entries, then issues an io_uring_enter() call (or uses the newer io_uring_submit() from liburing) to hand the batch to the kernel. The kernel processes the SQ, performs the I/O, and writes results to the CQ. The application polls the CQ (or uses io_uring_wait_cqe() for blocking) to retrieve completions.

// Example: Submit a single read request using liburing
#include <liburing.h>
#include <stdio.h>
#include <stdlib.h>

int main() {
    struct io_uring ring;
    // Initialize the ring with 32 entries
    io_uring_queue_init(32, &ring, 0);

    // Prepare a read SQE
    struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
    io_uring_prep_read(sqe, /* fd */ 0, /* buf */ NULL, /* len */ 0, /* off */ 0);

    // Submit the SQE (this may trigger io_uring_enter if needed)
    io_uring_submit(&ring);

    // Wait for the completion
    struct io_uring_cqe *cqe;
    io_uring_wait_cqe(&ring, &cqe);
    printf("Read result: %d\n", cqe->res);
    io_uring_cqe_seen(&ring, cqe);
    io_uring_queue_exit(&ring);
    return 0;
}

Shared Memory and Zero‑Copy

Because the SQ and CQ are memory‑mapped, the kernel can read SQ entries directly from user space without copying them into a kernel buffer. Similarly, completions are written directly into the user‑space CQ. This eliminates the extra copies that plague traditional read()/write() syscalls, where data must be copied between kernel and user space at least twice.

Operations and Opcodes

io_uring supports a growing set of opcodes. The most common are:

  • IORING_OP_READ, IORING_OP_WRITE – file I/O.
  • IORING_OP_SEND, IORING_OP_RECV – socket I/O.
  • IORING_OP_ACCEPT, IORING_OP_CONNECT – network connections.
  • IORING_OP_OPENAT, IORING_OP_CLOSE – file management.
  • IORING_OP_FALLOCATE, IORING_OP_FADVISE – storage hints.
  • IORING_OP_LINK – chain multiple operations.

Each opcode can be combined with flags such as IOSQE_IO_LINK to create a linked chain of operations that are submitted atomically.

Integrating io_uring in Your Application

Installing liburing

The easiest way to start is the liburing userspace library:

# On Debian/Ubuntu
sudo apt-get install liburing-dev

# Or build from source
git clone https://github.com/axboe/liburing.git
cd liburing && ./configure && make && sudo make install

Basic Read Example

The snippet above shows a synchronous read using io_uring_wait_cqe. In practice, you’ll want to batch many SQEs and poll the CQ without blocking.

// Batch submit 100 reads
for (int i = 0; i < 100; i++) {
    struct io_uring_sqe *sqe = io_uring_get_sqe(&ring);
    io_uring_prep_read(sqe, fd[i], bufs[i], sizes[i], offsets[i]);
}
io_uring_submit(&ring);

Advanced Features

  • Fixed Files – Register a set of file descriptors once; subsequent operations can reference them by index, avoiding repeated open()/close() calls.
  • Registered Buffers – Pre‑register memory regions to enable zero‑copy I/O and reduce TLB pressure.
  • Polling Mode – For ultra‑low latency, the kernel can poll the device without needing a wakeup interrupt.

Performance in Production

Benchmark Results

A recent benchmark on a 4‑core AWS EC2 instance (c5.4xlarge) comparing libaio vs io_uring for 4 KB random reads on an NVMe SSD showed:

Metriclibaioio_uring
IOPS (single QD)120 k210 k
Latency (p99)18 µs9 µs
CPU usage22 %12 %

The numbers illustrate that io_uring not only increases throughput but also cuts CPU consumption, which is critical for scaling beyond a single core.

Real‑World Use Cases

  • Web Servers – nginx and Caddy have experimental io_uring backends that reduce request latency by ~30 % under heavy load.
  • Databases – RocksDB and MySQL are exploring io_uring for their write‑ahead log and page reads, aiming to shave microseconds off each transaction.
  • Event Loops – The Node.js community is prototyping an io_uring‑based libuv backend to improve I/O performance in single‑threaded JavaScript applications.

Patterns in Production

Web Servers

When integrating io_uring into a web server, a common pattern is to maintain a single event loop that:

  1. Accepts new connections using IORING_OP_ACCEPT.
  2. Reads request data with IORING_OP_READ.
  3. Dispatches the request to a worker thread (or handles it inline).
  4. Writes the response with IORING_OP_WRITE.

All operations are posted to the SQ in a single batch, and the CQ is polled after each batch, minimizing context switches.

Databases

In a storage engine, io_uring can be used to issue concurrent reads and writes to multiple files. By chaining IORING_OP_READ with IORING_OP_WRITE via IOSQE_IO_LINK, the engine ensures that a read’s result is available before the dependent write is executed, without an extra syscall.

Event Loops

For a single‑threaded event loop, io_uring provides a unified mechanism to handle sockets, timers, and even device I/O. The loop can register a fixed set of file descriptors, then use io_uring_enter() with IORING_ENTER_GETEVENTS to wait for completions, effectively replacing epoll_wait with a more efficient interface.

Pitfalls and Debugging

  • Memory Ordering – Because the SQ and CQ are shared, you must use proper memory barriers (io_uring_smp_wmb()) to ensure visibility.
  • CQ Overflow – If the application does not reap completions quickly enough, the CQ can wrap and entries may be overwritten. Always check cqe->flags for IORING_CQE_F_OVERFLOW.
  • Kernel Version Support – Some opcodes (e.g., IORING_OP_OPENAT2) require kernel ≥ 5.6. Verify compatibility with your deployment.
  • Error Handling – Completions carry a negative res value on failure. Map these to errno and handle appropriately.

Future Outlook

The io_uring interface is still evolving. Upcoming work includes:

  • io_uring for filesystem operations – Directly integrating with the page cache to avoid double buffering.
  • Integration with io_uring‑aware allocators – Reducing memory allocation overhead for large buffers.
  • Support for io_uring in languages – Rust, Go, and Python are adding bindings, making the technology accessible beyond C/C++.

As more production systems adopt io_uring, we can expect the gap between raw hardware performance and application‑level I/O to narrow further.

Key Takeaways

  • io_uring eliminates per‑operation syscalls by sharing SQ/CQ memory with the kernel.
  • Batch submissions and poll completions to maximize throughput and minimize context switches.
  • Use liburing for a high‑level API, but understand the underlying ring mechanics for advanced features.
  • Production benefits include higher IOPS, lower latency, and reduced CPU usage.
  • Watch for overflow, use proper memory barriers, and check kernel version compatibility.

Further Reading