High-Throughput Linux I/O in Rust: io_uring, Tokio & Zero-Copy Networking

Table of Contents(16 sections)
Linux I/O performance is critical for high-scale network services. Traditional epoll-based asynchronous I/O, while efficient, still incurs significant overhead due to repeated syscalls and data copying between kernel and user space. io_uring, introduced in Linux kernel 5.1, fundamentally re-architects this by providing a powerful asynchronous I/O interface that minimizes syscalls and enables zero-copy operations. This guide details leveraging io_uring with Rust, specifically integrating it with tokio for high-throughput, zero-copy networking.
The io_uring Paradigm Shift
io_uring operates on a shared ring buffer mechanism between user space and the kernel. Instead of issuing individual syscalls for each I/O operation, applications enqueue Submission Queue Entries (SQEs) into a Submission Queue (SQ). The kernel then processes these SQEs asynchronously, placing Completion Queue Entries (CQEs) into a Completion Queue (CQ) when operations complete. This batching significantly reduces context switches and syscall overhead.
Key io_uring features:
- Batching: Multiple I/O operations submitted with a single
io_uring_entersyscall. - Asynchronous: Operations complete without blocking the calling thread.
- Polling: Kernel can actively poll the SQ, eliminating the need for
io_uring_enterin some cases, further reducing latency. - Zero-Copy: Operations like
IORING_OP_SENDMSGandIORING_OP_RECVMSGcan directly operate on user-space buffers, avoiding data copies. - Fixed Buffers/Files: Registering buffers and file descriptors with
io_uringallows the kernel to optimize access and avoid repeated lookups.
epoll vs. io_uring: Architectural Comparison
| Feature | epoll (Standard Tokio) | io_uring (tokio-uring) |
|---|---|---|
| I/O Model | Edge-triggered, event-driven | Asynchronous, ring-buffer |
| Syscall Overhead | High (1 syscall per op + epoll_wait) | Low (batching, io_uring_enter for many ops) |
| Data Copy | User-kernel copy for read/write | Zero-copy possible (MSG_ZEROCOPY) |
| Buffer Management | User-managed, passed per-syscall | Kernel-registered fixed buffers |
| File Descriptors | Passed per-syscall | Kernel-registered fixed files |
| Latency | Higher due to syscalls/copies | Lower due to batching/zero-copy |
| Throughput | Good for many connections, limited by syscalls | Excellent for high-volume I/O |
| Kernel Version | Linux 2.5.44+ | Linux 5.1+ |
| Complexity | Simpler API | More complex API, higher learning curve |
Rust Integration: tokio-uring
While direct io_uring bindings exist (e.g., io-uring crate), integrating io_uring with an existing async runtime like tokio is often desirable. The tokio-uring crate provides a tokio-compatible runtime built on io_uring. It offers io_uring-backed equivalents for common tokio I/O primitives like TcpStream, UdpSocket, and file I/O.
Setting up tokio-uring
Add the following to your Cargo.toml:
[dependencies]
tokio-uring = { version = "0.6", features = ["net", "fs"] }
bytes = "1.4"
log = "0.4"
env_logger = "0.10"
The net feature enables TcpStream, UdpSocket, etc., while fs enables file I/O.
Basic tokio-uring TCP Echo Server
This example demonstrates a simple TCP echo server using tokio-uring. Note the tokio_uring::start() macro, which initializes the io_uring runtime.
use tokio_uring::net::{TcpListener, TcpStream};
use tokio_uring::buf::IoBuf;
use bytes::{BytesMut, BufMut};
use log::{info, error};
const BUFFER_SIZE: usize = 4096;
#[tokio_uring::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
env_logger::init();
let addr = "127.0.0.1:8080";
let listener = TcpListener::bind(addr).await?;
info!("Listening on {}", addr);
loop {
match listener.accept().await {
Ok((socket, peer_addr)) => {
info!("Accepted connection from {}", peer_addr);
tokio_uring::spawn(async move {
if let Err(e) = handle_connection(socket).await {
error!("Error handling connection from {}: {}", peer_addr, e);
}
});
}
Err(e) => {
error!("Error accepting connection: {}", e);
}
}
}
}
async fn handle_connection(mut socket: TcpStream) -> Result<(), Box<dyn std::error::Error>> {
let mut buf = BytesMut::with_capacity(BUFFER_SIZE);
loop {
// Prepare a buffer for receiving data.
// `bytes::BytesMut` implements `IoBuf`, making it suitable for `tokio-uring`.
let (res, filled_buf) = socket.recv(buf.split_off(0).limit(BUFFER_SIZE)).await;
let bytes_read = res?;
if bytes_read == 0 {
info!("Client disconnected.");
break; // Client disconnected
}
// Advance the buffer to reflect the bytes read.
// `filled_buf` is the original buffer, but its `IoBuf` trait implementation
// ensures it's correctly sliced for the `recv` operation.
// We need to manually advance the `BytesMut` to reflect the data.
buf.unsplit(filled_buf);
buf.advance_mut(bytes_read);
info!("Received {} bytes: {:?}", bytes_read, &buf[..bytes_read]);
// Send the received data back.
// `send` takes an `IoBuf` and returns the buffer back after completion.
let (res, _) = socket.send(buf.split_to(bytes_read)).await;
let bytes_written = res?;
info!("Sent {} bytes.", bytes_written);
if bytes_written == 0 {
info!("Failed to send data, client likely disconnected.");
break;
}
}
Ok(())
}
To test this, you can use netcat: nc 127.0.0.1 8080. Type some text and press enter; the server will echo it back.
Zero-Copy Networking with MSG_ZEROCOPY
The true power of io_uring for networking lies in its ability to perform zero-copy operations. This is achieved using the IORING_OP_SENDMSG operation with the MSG_ZEROCOPY flag. When this flag is set, the kernel maps the user-space buffer directly into its own memory space for transmission, avoiding the traditional copy_from_user syscall. The application is notified via a CQE when the buffer is safe to reuse (i.e., the data has been transmitted or queued for transmission).
tokio-uring exposes this functionality through its send_zc method on TcpStream.
Implementing Zero-Copy TCP Streaming
This example demonstrates a server that receives data and then sends a fixed response using zero-copy. The key is to manage the buffer lifecycle carefully.
use tokio_uring::net::{TcpListener, TcpStream};
use tokio_uring::buf::IoBuf;
use bytes::{BytesMut, BufMut};
use log::{info, error};
use std::sync::Arc;
const BUFFER_SIZE: usize = 4096;
const RESPONSE_DATA: &[u8] = b"HTTP/1.1 200 OK\r\nContent-Length: 12\r\n\r\nHello, World!";
#[tokio_uring::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
env_logger::init();
let addr = "127.0.0.1:8080";
let listener = TcpListener::bind(addr).await?;
info!("Listening on {}", addr);
loop {
match listener.accept().await {
Ok((socket, peer_addr)) => {
info!("Accepted connection from {}", peer_addr);
tokio_uring::spawn(async move {
if let Err(e) = handle_zero_copy_connection(socket).await {
error!("Error handling zero-copy connection from {}: {}", peer_addr, e);
}
});
}
Err(e) => {
error!("Error accepting connection: {}", e);
}
}
}
}
async fn handle_zero_copy_connection(mut socket: TcpStream) -> Result<(), Box<dyn std::error::Error>> {
let mut recv_buf = BytesMut::with_capacity(BUFFER_SIZE);
let send_buf = Arc::new(BytesMut::from(RESPONSE_DATA)); // Use Arc for shared buffer for zero-copy send
loop {
// Receive data (standard copy, as MSG_ZEROCOPY is for send)
let (res, filled_buf) = socket.recv(recv_buf.split_off(0).limit(BUFFER_SIZE)).await;
let bytes_read = res?;
if bytes_read == 0 {
info!("Client disconnected.");
break;
}
recv_buf.unsplit(filled_buf);
recv_buf.advance_mut(bytes_read);
info!("Received {} bytes from client.", bytes_read);
// Perform zero-copy send
// `send_zc` returns a future that completes when the kernel is done with the buffer.
// The buffer is returned, allowing reuse.
let (res, returned_buf) = socket.send_zc(send_buf.clone().slice(..)).await;
let bytes_written = res?;
info!("Zero-copy sent {} bytes.", bytes_written);
if bytes_written == 0 {
info!("Failed to zero-copy send data, client likely disconnected.");
break;
}
// Clear the receive buffer for the next read
recv_buf.clear();
}
Ok(())
}
To test this, you can use curl: curl http://127.0.0.1:8080. You should receive "Hello, World!".
Important Considerations for Zero-Copy
- Buffer Lifetime: The buffer passed to
send_zcmust remain valid and unchanged until thesend_zcfuture completes.tokio-uringhandles this by returning the buffer, ensuring it's not dropped prematurely. For shared, static data,Arc<BytesMut>orArc<[u8]>is suitable. - Kernel Support:
MSG_ZEROCOPYrequires Linux kernel 5.2 or newer. - Error Handling: If
MSG_ZEROCOPYfails (e.g., due to insufficient kernel memory or older kernel),send_zcwill return an error. A fallback to standardsendmight be necessary in production environments. - Buffer Registration: For maximum performance, especially with repeated sends of the same data, consider registering buffers with
io_uringusingIORING_REGISTER_BUFFERS.tokio-uringprovidestokio_uring::buf::BoundedBufandtokio_uring::buf::FixedBuffor this.
Benchmarking 100k Concurrent WebSocket Connections
Benchmarking io_uring for 100k concurrent WebSocket connections requires a robust setup. We'll outline the approach and expected results rather than providing a full benchmark client/server, which would be extensive.
Server Architecture
A tokio-uring based WebSocket server would involve:
tokio_uring::net::TcpListener: Accepting new connections.- WebSocket Handshake: Performing the HTTP upgrade. This part is standard HTTP and can use
httparseor similar. - WebSocket Framing: Implementing the WebSocket protocol's framing (masking, opcode, length).
tokio_uring::net::TcpStream::recv: For receiving WebSocket frames.tokio_uring::net::TcpStream::send_zc: For sending WebSocket frames, especially for broadcast messages or large data. This is where zero-copy shines.
Client Architecture
A benchmark client needs to:
- Establish 100k TCP connections: This requires significant file descriptor limits (
ulimit -n). - Perform WebSocket handshakes: For each connection.
- Maintain connections: Send pings/pongs to keep connections alive.
- Send/Receive data: Simulate application traffic.
- Measure latency and throughput: Track round-trip times and data rates.
Expected Performance Gains
With io_uring and zero-copy, for a server sending frequent, identical messages (e.g., market data, game state updates) to many clients, we expect:
- Reduced CPU Utilization: Fewer syscalls and data copies mean less kernel CPU time.
- Higher Throughput: More data processed per unit of time.
- Lower Latency: Especially for send operations, as data is directly queued for transmission.
- Increased Connection Density: More connections can be handled per server instance due to reduced per-connection overhead.
Benchmarking Environment
- Kernel: Linux 5.10+ (LTS) or 6.x for optimal
io_uringfeatures. - Hardware: High core count CPU, ample RAM, fast NIC.
ulimit -n: Set to a value > 100,000 (e.g.,ulimit -n 1048576).- Network Tuning:
sysctlparameters likenet.core.somaxconn,net.ipv4.tcp_max_syn_backlog,net.ipv4.tcp_tw_reuse,net.ipv4.tcp_fin_timeout.
Production Gotchas & Troubleshooting
io_uringNot Available/Enabled:- Symptom:
io_uringoperations fail withENOSYSor similar errors. - Cause: Kernel version < 5.1, or
io_uringmodule not loaded/compiled. - Fix: Upgrade kernel to 5.1+ (5.10+ recommended for stability/features). Verify
CONFIG_IO_URING=yin kernel config.
- Symptom:
ulimit -nToo Low:- Symptom:
Too many open fileserror when accepting connections or creating sockets. - Cause: Default file descriptor limit (often 1024) is insufficient for high concurrency.
- Fix: Increase
ulimit -nfor the user running the application. For systemd services, setLimitNOFILEin the service unit file.
- Symptom:
MSG_ZEROCOPYFailures:- Symptom:
send_zcreturnsEOPNOTSUPPor other errors. - Cause: Kernel version < 5.2, or specific network driver limitations.
- Fix: Upgrade kernel. Implement a fallback to
TcpStream::sendifsend_zcfails.
- Symptom:
- Buffer Management Issues with Zero-Copy:
- Symptom: Data corruption, use-after-free errors, or unexpected behavior.
- Cause: Reusing or modifying a buffer before
send_zccompletes and returns it. - Fix: Ensure the buffer's lifetime is strictly managed.
tokio-uring'ssend_zcreturns the buffer, indicating it's safe to reuse. For shared static data, useArcandsliceto ensure immutability during kernel processing.
- High CPU Usage with
io_uringPolling:- Symptom:
io_uringworker threads consume 100% CPU even with low load. - Cause:
IORING_SETUP_SQPOLL(kernel polling) can be aggressive. If not configured correctly, it might spin-wait. - Fix: Ensure
io_uringis configured appropriately for your workload.tokio-uringgenerally manages this, but if you're using rawio-uring, be mindful of polling flags. For most server workloads, event-driven (non-polling)io_uringis sufficient, with polling reserved for ultra-low-latency scenarios.
- Symptom:
- Memory Pressure:
- Symptom: OOM errors, excessive swapping.
- Cause: Large number of connections, each with its own receive buffer, or large registered fixed buffers.
- Fix: Optimize buffer sizes. Consider using a buffer pool for receive operations. For fixed buffers, register only what's necessary.
Frequently Asked Questions
- Can I use
tokio-uringwith existingtokiocode? Yes,tokio-uringprovides atokio-compatible runtime. You can usetokio_uring::spawnto runio_uring-backed futures alongside regulartokio::spawnfutures, though it's generally recommended to keepio_uringspecific I/O on thetokio-uringruntime for optimal performance. - What are the main benefits of
io_uringoverepollfor networking? The primary benefits are reduced syscall overhead through batching, and the ability to perform zero-copy data transfers between user space and the kernel. This leads to lower CPU utilization, higher throughput, and lower latency, especially for high-volume I/O operations. - Is
io_uringalways faster thanepoll? Not always. For low-concurrency, low-throughput applications, the overhead of setting upio_uringmight outweigh its benefits.io_uringshines in scenarios with high concurrency, high I/O volume, or when zero-copy is critical. It also requires a modern Linux kernel. - How does
tokio-uringhandle buffer management for zero-copy?tokio-uring'ssend_zcmethod takes anIoBuf(likeBytesMutorArc<[u8]>) and returns it once the kernel has finished with the buffer. This ensures the buffer's lifetime is correctly managed, preventing use-after-free issues. For static data,Arcis used to share ownership. - What are the kernel requirements for
io_uringandMSG_ZEROCOPY?io_uringrequires Linux kernel 5.1 or newer.MSG_ZEROCOPYspecifically requires Linux kernel 5.2 or newer. For production, kernel 5.10 (LTS) or a recent 6.x kernel is recommended for stability and feature completeness.
Conclusion
io_uring represents a significant advancement in Linux asynchronous I/O, offering unparalleled performance for high-throughput applications. Rust, with its strong type system and performance characteristics, is an ideal language to leverage io_uring's capabilities. The tokio-uring crate provides a robust and ergonomic way to integrate io_uring into tokio-based applications, enabling developers to build highly efficient network services with features like zero-copy data transfer. While the learning curve for io_uring can be steeper than traditional epoll, the performance gains for demanding workloads are substantial and warrant the investment.
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

Zero-Knowledge Proofs in Rust & Circom: SnarkJS Verification & Production Guide
Comprehensive guide covering zero-knowledge proofs in rust & circom: snarkjs verification & production guide with production-grade architecture and code examples.
Read more
Building High-Throughput Microservices with Rust and Axum: Complete Production Guide
Comprehensive guide covering building high-throughput microservices with rust and axum: complete production guide with production-grade architecture and code examples.
Read more
eBPF in Production: Low-Overhead Linux Observability, Tracing, and Kernel Profiling
Implement low-overhead Linux kernel observability using eBPF. Profile system call latency, track memory allocations, and monitor network sockets without sidecars.
Read more