•11 min read

High-Throughput Linux I/O in Rust: io_uring, Tokio & Zero-Copy Networking

High-Throughput Linux I/O in Rust: io_uring, Tokio & Zero-Copy Networking

Linux I/O performance is critical for high-scale network services. Traditional epoll-based asynchronous I/O, while efficient, still incurs significant overhead due to repeated syscalls and data copying between kernel and user space. io_uring, introduced in Linux kernel 5.1, fundamentally re-architects this by providing a powerful asynchronous I/O interface that minimizes syscalls and enables zero-copy operations. This guide details leveraging io_uring with Rust, specifically integrating it with tokio for high-throughput, zero-copy networking.

Audio Briefing
0:00 / 0:00

The io_uring Paradigm Shift

io_uring operates on a shared ring buffer mechanism between user space and the kernel. Instead of issuing individual syscalls for each I/O operation, applications enqueue Submission Queue Entries (SQEs) into a Submission Queue (SQ). The kernel then processes these SQEs asynchronously, placing Completion Queue Entries (CQEs) into a Completion Queue (CQ) when operations complete. This batching significantly reduces context switches and syscall overhead.

Key io_uring features:

  1. Batching: Multiple I/O operations submitted with a single io_uring_enter syscall.
  2. Asynchronous: Operations complete without blocking the calling thread.
  3. Polling: Kernel can actively poll the SQ, eliminating the need for io_uring_enter in some cases, further reducing latency.
  4. Zero-Copy: Operations like IORING_OP_SENDMSG and IORING_OP_RECVMSG can directly operate on user-space buffers, avoiding data copies.
  5. Fixed Buffers/Files: Registering buffers and file descriptors with io_uring allows the kernel to optimize access and avoid repeated lookups.
Advertisement

epoll vs. io_uring: Architectural Comparison

Featureepoll (Standard Tokio)io_uring (tokio-uring)
I/O ModelEdge-triggered, event-drivenAsynchronous, ring-buffer
Syscall OverheadHigh (1 syscall per op + epoll_wait)Low (batching, io_uring_enter for many ops)
Data CopyUser-kernel copy for read/writeZero-copy possible (MSG_ZEROCOPY)
Buffer ManagementUser-managed, passed per-syscallKernel-registered fixed buffers
File DescriptorsPassed per-syscallKernel-registered fixed files
LatencyHigher due to syscalls/copiesLower due to batching/zero-copy
ThroughputGood for many connections, limited by syscallsExcellent for high-volume I/O
Kernel VersionLinux 2.5.44+Linux 5.1+
ComplexitySimpler APIMore complex API, higher learning curve

Rust Integration: tokio-uring

While direct io_uring bindings exist (e.g., io-uring crate), integrating io_uring with an existing async runtime like tokio is often desirable. The tokio-uring crate provides a tokio-compatible runtime built on io_uring. It offers io_uring-backed equivalents for common tokio I/O primitives like TcpStream, UdpSocket, and file I/O.

Setting up tokio-uring

Add the following to your Cargo.toml:

[dependencies]
tokio-uring = { version = "0.6", features = ["net", "fs"] }
bytes = "1.4"
log = "0.4"
env_logger = "0.10"

The net feature enables TcpStream, UdpSocket, etc., while fs enables file I/O.

Basic tokio-uring TCP Echo Server

This example demonstrates a simple TCP echo server using tokio-uring. Note the tokio_uring::start() macro, which initializes the io_uring runtime.

use tokio_uring::net::{TcpListener, TcpStream};
use tokio_uring::buf::IoBuf;
use bytes::{BytesMut, BufMut};
use log::{info, error};

const BUFFER_SIZE: usize = 4096;

#[tokio_uring::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    env_logger::init();
    let addr = "127.0.0.1:8080";
    let listener = TcpListener::bind(addr).await?;
    info!("Listening on {}", addr);

    loop {
        match listener.accept().await {
            Ok((socket, peer_addr)) => {
                info!("Accepted connection from {}", peer_addr);
                tokio_uring::spawn(async move {
                    if let Err(e) = handle_connection(socket).await {
                        error!("Error handling connection from {}: {}", peer_addr, e);
                    }
                });
            }
            Err(e) => {
                error!("Error accepting connection: {}", e);
            }
        }
    }
}

async fn handle_connection(mut socket: TcpStream) -> Result<(), Box<dyn std::error::Error>> {
    let mut buf = BytesMut::with_capacity(BUFFER_SIZE);
    loop {
        // Prepare a buffer for receiving data.
        // `bytes::BytesMut` implements `IoBuf`, making it suitable for `tokio-uring`.
        let (res, filled_buf) = socket.recv(buf.split_off(0).limit(BUFFER_SIZE)).await;
        let bytes_read = res?;

        if bytes_read == 0 {
            info!("Client disconnected.");
            break; // Client disconnected
        }

        // Advance the buffer to reflect the bytes read.
        // `filled_buf` is the original buffer, but its `IoBuf` trait implementation
        // ensures it's correctly sliced for the `recv` operation.
        // We need to manually advance the `BytesMut` to reflect the data.
        buf.unsplit(filled_buf);
        buf.advance_mut(bytes_read);

        info!("Received {} bytes: {:?}", bytes_read, &buf[..bytes_read]);

        // Send the received data back.
        // `send` takes an `IoBuf` and returns the buffer back after completion.
        let (res, _) = socket.send(buf.split_to(bytes_read)).await;
        let bytes_written = res?;

        info!("Sent {} bytes.", bytes_written);

        if bytes_written == 0 {
            info!("Failed to send data, client likely disconnected.");
            break;
        }
    }
    Ok(())
}

To test this, you can use netcat: nc 127.0.0.1 8080. Type some text and press enter; the server will echo it back.

Zero-Copy Networking with MSG_ZEROCOPY

The true power of io_uring for networking lies in its ability to perform zero-copy operations. This is achieved using the IORING_OP_SENDMSG operation with the MSG_ZEROCOPY flag. When this flag is set, the kernel maps the user-space buffer directly into its own memory space for transmission, avoiding the traditional copy_from_user syscall. The application is notified via a CQE when the buffer is safe to reuse (i.e., the data has been transmitted or queued for transmission).

tokio-uring exposes this functionality through its send_zc method on TcpStream.

Implementing Zero-Copy TCP Streaming

This example demonstrates a server that receives data and then sends a fixed response using zero-copy. The key is to manage the buffer lifecycle carefully.

use tokio_uring::net::{TcpListener, TcpStream};
use tokio_uring::buf::IoBuf;
use bytes::{BytesMut, BufMut};
use log::{info, error};
use std::sync::Arc;

const BUFFER_SIZE: usize = 4096;
const RESPONSE_DATA: &[u8] = b"HTTP/1.1 200 OK\r\nContent-Length: 12\r\n\r\nHello, World!";

#[tokio_uring::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    env_logger::init();
    let addr = "127.0.0.1:8080";
    let listener = TcpListener::bind(addr).await?;
    info!("Listening on {}", addr);

    loop {
        match listener.accept().await {
            Ok((socket, peer_addr)) => {
                info!("Accepted connection from {}", peer_addr);
                tokio_uring::spawn(async move {
                    if let Err(e) = handle_zero_copy_connection(socket).await {
                        error!("Error handling zero-copy connection from {}: {}", peer_addr, e);
                    }
                });
            }
            Err(e) => {
                error!("Error accepting connection: {}", e);
            }
        }
    }
}

async fn handle_zero_copy_connection(mut socket: TcpStream) -> Result<(), Box<dyn std::error::Error>> {
    let mut recv_buf = BytesMut::with_capacity(BUFFER_SIZE);
    let send_buf = Arc::new(BytesMut::from(RESPONSE_DATA)); // Use Arc for shared buffer for zero-copy send

    loop {
        // Receive data (standard copy, as MSG_ZEROCOPY is for send)
        let (res, filled_buf) = socket.recv(recv_buf.split_off(0).limit(BUFFER_SIZE)).await;
        let bytes_read = res?;

        if bytes_read == 0 {
            info!("Client disconnected.");
            break;
        }

        recv_buf.unsplit(filled_buf);
        recv_buf.advance_mut(bytes_read);

        info!("Received {} bytes from client.", bytes_read);

        // Perform zero-copy send
        // `send_zc` returns a future that completes when the kernel is done with the buffer.
        // The buffer is returned, allowing reuse.
        let (res, returned_buf) = socket.send_zc(send_buf.clone().slice(..)).await;
        let bytes_written = res?;

        info!("Zero-copy sent {} bytes.", bytes_written);

        if bytes_written == 0 {
            info!("Failed to zero-copy send data, client likely disconnected.");
            break;
        }

        // Clear the receive buffer for the next read
        recv_buf.clear();
    }
    Ok(())
}

To test this, you can use curl: curl http://127.0.0.1:8080. You should receive "Hello, World!".

Important Considerations for Zero-Copy

  • Buffer Lifetime: The buffer passed to send_zc must remain valid and unchanged until the send_zc future completes. tokio-uring handles this by returning the buffer, ensuring it's not dropped prematurely. For shared, static data, Arc<BytesMut> or Arc<[u8]> is suitable.
  • Kernel Support: MSG_ZEROCOPY requires Linux kernel 5.2 or newer.
  • Error Handling: If MSG_ZEROCOPY fails (e.g., due to insufficient kernel memory or older kernel), send_zc will return an error. A fallback to standard send might be necessary in production environments.
  • Buffer Registration: For maximum performance, especially with repeated sends of the same data, consider registering buffers with io_uring using IORING_REGISTER_BUFFERS. tokio-uring provides tokio_uring::buf::BoundedBuf and tokio_uring::buf::FixedBuf for this.
Advertisement

Benchmarking 100k Concurrent WebSocket Connections

Benchmarking io_uring for 100k concurrent WebSocket connections requires a robust setup. We'll outline the approach and expected results rather than providing a full benchmark client/server, which would be extensive.

Server Architecture

A tokio-uring based WebSocket server would involve:

  1. tokio_uring::net::TcpListener: Accepting new connections.
  2. WebSocket Handshake: Performing the HTTP upgrade. This part is standard HTTP and can use httparse or similar.
  3. WebSocket Framing: Implementing the WebSocket protocol's framing (masking, opcode, length).
  4. tokio_uring::net::TcpStream::recv: For receiving WebSocket frames.
  5. tokio_uring::net::TcpStream::send_zc: For sending WebSocket frames, especially for broadcast messages or large data. This is where zero-copy shines.

Client Architecture

A benchmark client needs to:

  1. Establish 100k TCP connections: This requires significant file descriptor limits (ulimit -n).
  2. Perform WebSocket handshakes: For each connection.
  3. Maintain connections: Send pings/pongs to keep connections alive.
  4. Send/Receive data: Simulate application traffic.
  5. Measure latency and throughput: Track round-trip times and data rates.

Expected Performance Gains

With io_uring and zero-copy, for a server sending frequent, identical messages (e.g., market data, game state updates) to many clients, we expect:

  • Reduced CPU Utilization: Fewer syscalls and data copies mean less kernel CPU time.
  • Higher Throughput: More data processed per unit of time.
  • Lower Latency: Especially for send operations, as data is directly queued for transmission.
  • Increased Connection Density: More connections can be handled per server instance due to reduced per-connection overhead.

Benchmarking Environment

  • Kernel: Linux 5.10+ (LTS) or 6.x for optimal io_uring features.
  • Hardware: High core count CPU, ample RAM, fast NIC.
  • ulimit -n: Set to a value > 100,000 (e.g., ulimit -n 1048576).
  • Network Tuning: sysctl parameters like net.core.somaxconn, net.ipv4.tcp_max_syn_backlog, net.ipv4.tcp_tw_reuse, net.ipv4.tcp_fin_timeout.

Production Gotchas & Troubleshooting

  1. io_uring Not Available/Enabled:
    • Symptom: io_uring operations fail with ENOSYS or similar errors.
    • Cause: Kernel version < 5.1, or io_uring module not loaded/compiled.
    • Fix: Upgrade kernel to 5.1+ (5.10+ recommended for stability/features). Verify CONFIG_IO_URING=y in kernel config.
  2. ulimit -n Too Low:
    • Symptom: Too many open files error when accepting connections or creating sockets.
    • Cause: Default file descriptor limit (often 1024) is insufficient for high concurrency.
    • Fix: Increase ulimit -n for the user running the application. For systemd services, set LimitNOFILE in the service unit file.
  3. MSG_ZEROCOPY Failures:
    • Symptom: send_zc returns EOPNOTSUPP or other errors.
    • Cause: Kernel version < 5.2, or specific network driver limitations.
    • Fix: Upgrade kernel. Implement a fallback to TcpStream::send if send_zc fails.
  4. Buffer Management Issues with Zero-Copy:
    • Symptom: Data corruption, use-after-free errors, or unexpected behavior.
    • Cause: Reusing or modifying a buffer before send_zc completes and returns it.
    • Fix: Ensure the buffer's lifetime is strictly managed. tokio-uring's send_zc returns the buffer, indicating it's safe to reuse. For shared static data, use Arc and slice to ensure immutability during kernel processing.
  5. High CPU Usage with io_uring Polling:
    • Symptom: io_uring worker threads consume 100% CPU even with low load.
    • Cause: IORING_SETUP_SQPOLL (kernel polling) can be aggressive. If not configured correctly, it might spin-wait.
    • Fix: Ensure io_uring is configured appropriately for your workload. tokio-uring generally manages this, but if you're using raw io-uring, be mindful of polling flags. For most server workloads, event-driven (non-polling) io_uring is sufficient, with polling reserved for ultra-low-latency scenarios.
  6. Memory Pressure:
    • Symptom: OOM errors, excessive swapping.
    • Cause: Large number of connections, each with its own receive buffer, or large registered fixed buffers.
    • Fix: Optimize buffer sizes. Consider using a buffer pool for receive operations. For fixed buffers, register only what's necessary.

Frequently Asked Questions

  1. Can I use tokio-uring with existing tokio code? Yes, tokio-uring provides a tokio-compatible runtime. You can use tokio_uring::spawn to run io_uring-backed futures alongside regular tokio::spawn futures, though it's generally recommended to keep io_uring specific I/O on the tokio-uring runtime for optimal performance.
  2. What are the main benefits of io_uring over epoll for networking? The primary benefits are reduced syscall overhead through batching, and the ability to perform zero-copy data transfers between user space and the kernel. This leads to lower CPU utilization, higher throughput, and lower latency, especially for high-volume I/O operations.
  3. Is io_uring always faster than epoll? Not always. For low-concurrency, low-throughput applications, the overhead of setting up io_uring might outweigh its benefits. io_uring shines in scenarios with high concurrency, high I/O volume, or when zero-copy is critical. It also requires a modern Linux kernel.
  4. How does tokio-uring handle buffer management for zero-copy? tokio-uring's send_zc method takes an IoBuf (like BytesMut or Arc<[u8]>) and returns it once the kernel has finished with the buffer. This ensures the buffer's lifetime is correctly managed, preventing use-after-free issues. For static data, Arc is used to share ownership.
  5. What are the kernel requirements for io_uring and MSG_ZEROCOPY? io_uring requires Linux kernel 5.1 or newer. MSG_ZEROCOPY specifically requires Linux kernel 5.2 or newer. For production, kernel 5.10 (LTS) or a recent 6.x kernel is recommended for stability and feature completeness.

Conclusion

io_uring represents a significant advancement in Linux asynchronous I/O, offering unparalleled performance for high-throughput applications. Rust, with its strong type system and performance characteristics, is an ideal language to leverage io_uring's capabilities. The tokio-uring crate provides a robust and ergonomic way to integrate io_uring into tokio-based applications, enabling developers to build highly efficient network services with features like zero-copy data transfer. While the learning curve for io_uring can be steeper than traditional epoll, the performance gains for demanding workloads are substantial and warrant the investment.

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement