NVMe-over-TCP in Production: Linux Kernel Architecture, SPDK & High-IOPS Kubernetes Storage

Table of Contents(14 sections)
NVMe-over-TCP (NVMe/TCP) provides a robust, high-performance block storage transport over standard Ethernet, leveraging the NVMe command set's efficiency. This guide details its architecture, deployment considerations, and integration into Kubernetes for high-IOPS, low-latency storage. We will analyze the performance implications of kernel-space versus user-space (SPDK) implementations and demonstrate achieving demanding performance targets.
NVMe/TCP Architecture Overview
NVMe/TCP encapsulates NVMe commands and data within TCP/IP packets. This allows standard network infrastructure to carry high-performance storage traffic, eliminating the need for specialized Fibre Channel or InfiniBand hardware.
Key Components
- NVMe Host: The initiator, typically a Linux server or Kubernetes node, that consumes NVMe/TCP storage.
- NVMe Target: The storage server exporting NVMe namespaces over TCP. This can be a dedicated storage appliance, a general-purpose server running LVM/ZFS with NVMe drives, or an SPDK-based solution.
- NVMe-oF (NVMe over Fabrics) Protocol: The overarching specification that defines how NVMe commands are transported over various network fabrics, including TCP, RDMA (RoCE, iWARP), and Fibre Channel.
Linux Kernel NVMe/TCP Stack
The Linux kernel provides native NVMe/TCP initiator and target support.
- Initiator: The
nvme-tcpkernel module handles connection establishment, command submission, and data transfer. It integrates with the block layer, presenting remote NVMe namespaces as local block devices (e.g.,/dev/nvme0n1). - Target: The
nvmet-tcpkernel module exposes local NVMe devices or block devices as NVMe/TCP namespaces.
This kernel-native approach offers simplicity and broad compatibility but introduces context switching overheads between user and kernel space, and TCP/IP stack processing.
SPDK (Storage Performance Development Kit)
SPDK is a collection of user-space libraries and tools for writing high-performance, scalable storage applications. For NVMe/TCP, SPDK offers:
- User-space NVMe Driver: Bypasses the kernel's block layer and NVMe driver, directly accessing NVMe devices via
UIO(Userspace I/O) orVFIO(Virtual Function I/O). - User-space TCP/IP Stack: SPDK includes its own highly optimized, polled-mode TCP/IP stack, eliminating kernel context switches and interrupt-driven processing. This is crucial for achieving ultra-low latency and high IOPS.
- Polling Mode: Instead of relying on interrupts, SPDK continuously polls for I/O completion and network events, reducing latency jitter.
The SPDK approach yields superior performance but requires dedicated CPU cores and memory, and can be more complex to configure.
Architectural Comparison: NVMe/TCP vs. Alternatives
| Feature | iSCSI | NVMe/TCP (Kernel) | NVMe/TCP (SPDK) | NVMe/RoCE (RDMA) |
|---|---|---|---|---|
| Protocol | SCSI over TCP/IP | NVMe over TCP/IP | NVMe over User-space TCP/IP | NVMe over RDMA |
| Command Set | SCSI (legacy) | NVMe (modern, parallel) | NVMe (modern, parallel) | NVMe (modern, parallel) |
| Network | Standard Ethernet | Standard Ethernet | Standard Ethernet | RDMA-capable Ethernet (RoCE/iWARP) |
| CPU Overhead | Moderate | Moderate-High (kernel TCP) | Low (user-space polling) | Low (hardware offload) |
| Latency | High (100s µs - ms) | Medium (50-200 µs) | Low (sub-50 µs) | Very Low (sub-20 µs) |
| IOPS | Moderate (10s-100s kIOPS) | High (100s kIOPS - 1M IOPS) | Very High (1M+ IOPS) | Extremely High (2M+ IOPS) |
| Complexity | Low | Low-Medium | High (dedicated resources) | Medium (RDMA network setup) |
| Hardware | Standard NIC | Standard NIC | Standard NIC (CPU/Mem intensive) | RDMA NIC (RoCE/iWARP) |
| Use Case | General purpose, compatibility | High-performance, cost-effective | Extreme performance, latency-critical | Ultra-low latency, HPC, AI/ML |
Deploying NVMe/TCP in Kubernetes
We will focus on a kernel-based NVMe/TCP target for simplicity and broad applicability, demonstrating how to achieve high performance. For SPDK, the principles are similar, but the target setup is more involved.
Prerequisites
- Kubernetes cluster (v1.20+)
- Linux kernel 5.0+ on all nodes (for robust NVMe/TCP support)
nvme-cliinstalled on target and initiator nodesmultipath-toolsinstalled on initiator nodes
1. NVMe/TCP Target Setup (Example: Dedicated Storage Server)
Assume a storage server with an NVMe SSD (/dev/nvme0n1). We'll create an LVM logical volume and expose it.
# On the NVMe/TCP Target Server
# 1. Ensure nvmet-tcp module is loaded
sudo modprobe nvmet-tcp
# 2. Create a Volume Group and Logical Volume (example: 100GB)
# Replace /dev/nvme0n1 with your actual NVMe device
sudo pvcreate /dev/nvme0n1
sudo vgcreate nvme_vg /dev/nvme0n1
sudo lvcreate -L 100G -n nvme_lv nvme_vg
# 3. Configure NVMe/TCP Target
# Create a subsystem (nqn.2023-10.com.locionic:k8s-storage)
# This NQN (NVMe Qualified Name) uniquely identifies the storage subsystem.
sudo nvme target create -t tcp -n nqn.2023-10.com.locionic:k8s-storage
# 4. Add a controller (listener) for the target
# Replace 192.168.1.100 with your target server's IP address
sudo nvme target add-listener -t tcp -n nqn.2023-10.com.locionic:k8s-storage -a 192.168.1.100 -s 4420
# 5. Create a namespace and attach the logical volume
# The namespace ID (nsid) must be unique within the subsystem.
sudo nvme target add-namespace -t tcp -n nqn.2023-10.com.locionic:k8s-storage -d /dev/nvme_vg/nvme_lv -s 1
# 6. Enable the subsystem
sudo nvme target enable -t tcp -n nqn.2023-10.com.locionic:k8s-storage
# Verify target configuration
sudo nvme target show
2. Kubernetes Initiator Setup (Worker Nodes)
Each Kubernetes worker node that will consume NVMe/TCP storage needs the initiator module and multipath-tools.
# On each Kubernetes Worker Node
# 1. Ensure nvme-tcp module is loaded
sudo modprobe nvme-tcp
# 2. Install multipath-tools
sudo apt update && sudo apt install -y multipath-tools # Debian/Ubuntu
# OR
sudo yum install -y device-mapper-multipath # CentOS/RHEL
# 3. Configure multipath (optional but highly recommended for HA)
# Edit /etc/multipath.conf
sudo tee /etc/multipath.conf <<EOF
defaults {
user_friendly_names yes
find_multipaths yes
# Increase queue depth for better performance with NVMe
queue_without_daemon no
max_sectors_kb 512
}
blacklist {
devnode "^(ram|loop|fd|md|dm-|sr|scd|st)[0-9]*"
devnode "^hd[a-z]"
devnode "^sd[a-z]"
}
EOF
# 4. Restart multipathd
sudo systemctl enable multipathd
sudo systemctl restart multipathd
# 5. Connect to the NVMe/TCP target (manual test)
# Replace 192.168.1.100 with your target server's IP
# Replace nqn.2023-10.com.locionic:k8s-storage with your target NQN
sudo nvme connect -t tcp -n nqn.2023-10.com.locionic:k8s-storage -a 192.168.1.100 -s 4420
# 6. Verify connection and device
# You should see a new /dev/nvmeXnY device
lsblk
sudo nvme list
3. Kubernetes CSI Driver Deployment
For dynamic provisioning and seamless integration, a Container Storage Interface (CSI) driver is essential. While a generic NVMe/TCP CSI driver might exist, often you'll use a vendor-specific one or adapt a generic block storage CSI driver to manage NVMe/TCP connections. Here, we'll outline a conceptual CSI driver that leverages nvme-cli for connection management.
A production-grade CSI driver would handle:
- Dynamic provisioning of NVMe namespaces on the target.
- Connecting/disconnecting NVMe/TCP devices on worker nodes.
- Multipath configuration.
- Volume publishing and unpublishing.
Example: Conceptual CSI Driver for NVMe/TCP
This is a simplified representation. A real CSI driver would be more complex, involving gRPC services, controller, and node plugins.
# csi-nvmetcp-driver.yaml
apiVersion: storage.k8s.io/v1
kind: CSIDriver
metadata:
name: nvmetcp.locionic.com
spec:
attachRequired: false # NVMe/TCP devices are directly attached
podInfoOnMount: false
volumeLifecycleModes:
- Persistent
- Ephemeral
fsGroupPolicy: File
# storageclass.yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: nvmetcp-fast
provisioner: nvmetcp.locionic.com # Must match the CSIDriver name
parameters:
targetIp: "192.168.1.100" # IP of your NVMe/TCP target
targetNqn: "nqn.2023-10.com.locionic:k8s-storage"
targetPort: "4420"
# Other parameters for dynamic provisioning (e.g., size, LVM details)
reclaimPolicy: Delete
volumeBindingMode: Immediate
allowVolumeExpansion: true
# pvc.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: my-nvmetcp-pvc
spec:
accessModes:
- ReadWriteOnce
storageClassName: nvmetcp-fast
resources:
requests:
storage: 10Gi
# pod-with-pvc.yaml
apiVersion: v1
kind: Pod
metadata:
name: fio-test-pod
spec:
containers:
- name: fio
image: ghcr.io/locionic/fio:latest # A custom FIO image with nvme-cli
command: ["/bin/bash", "-c", "sleep infinity"]
volumeMounts:
- name: nvmetcp-volume
mountPath: /data
volumes:
- name: nvmetcp-volume
persistentVolumeClaim:
claimName: my-nvmetcp-pvc
4. Multipath I/O Configuration
For high availability and increased bandwidth, configure multiple paths to the NVMe/TCP target. This requires the target to expose the same namespace via multiple IP addresses or network interfaces.
Target Side (Example with two IPs):
# On NVMe/TCP Target Server
# Assuming 192.168.1.100 and 192.168.1.101 are target IPs
sudo nvme target add-listener -t tcp -n nqn.2023-10.com.locionic:k8s-storage -a 192.168.1.100 -s 4420
sudo nvme target add-listener -t tcp -n nqn.2023-10.com.locionic:k8s-storage -a 192.168.1.101 -s 4420
Initiator Side (Kubernetes Worker Node):
The nvme connect command will automatically discover multiple paths if the target NQN and port are the same but IPs differ. multipathd will then manage these paths.
# On Kubernetes Worker Node
# Connect to both target IPs
sudo nvme connect -t tcp -n nqn.2023-10.com.locionic:k8s-storage -a 192.168.1.100 -s 4420
sudo nvme connect -t tcp -n nqn.2023-10.com.locionic:k8s-storage -a 192.168.1.101 -s 4420
# Verify multipath device
sudo multipath -ll
# Expected output will show a single device with multiple paths
# e.g., mpatha (360014050000000000000000000000000) dm-0 NVME,Linux
# size=100G features='0' hwhandler='0' wp=rw
# |- 0:0:0:0 nvme0n1 8:0 active ready running
# `- 1:0:0:0 nvme1n1 8:16 active ready running
The CSI driver would abstract this, connecting to all specified target IPs.
Performance Benchmarking (500k+ Random Write IOPS)
To achieve 500k+ random write IOPS with sub-100µs latency, several factors are critical:
- Fast NVMe SSDs: The underlying storage on the target must be capable.
- High-Bandwidth Network: 25GbE or 100GbE is recommended.
- CPU Resources: Dedicated cores for SPDK or sufficient cores for kernel-based I/O.
- Optimal FIO Parameters:
iodepth: High queue depth to saturate the device.numjobs: Multiple parallel jobs.direct=1: Bypass OS cache.ioengine=libaioorioengine=io_uring(for modern kernels).randwrite: Random writes.bs=4k: Standard block size.
FIO Test Example (inside fio-test-pod):
# On the fio-test-pod
# Ensure /data is mounted from the NVMe/TCP PVC
cd /data
# Create a test file
dd if=/dev/zero of=testfile bs=1M count=1024 status=progress
# FIO command for 4K random writes
fio --name=randwrite_test \
--ioengine=libaio \
--iodepth=128 \
--rw=randwrite \
--bs=4k \
--direct=1 \
--numjobs=4 \
--size=10G \
--runtime=60 \
--filename=testfile \
--group_reporting \
--output-format=json
Expected Output (excerpt for 500k+ IOPS, sub-100µs latency):
{
"jobs": [
{
"jobname": "randwrite_test",
"write": {
"io_bytes": 10737418240,
"bw": 178957,
"iops": 44739,
"lat_ns": {
"min": 20000,
"max": 150000,
"mean": 75000,
"stddev": 15000
}
}
},
{
"jobname": "randwrite_test",
"write": {
"io_bytes": 10737418240,
"bw": 178957,
"iops": 44739,
"lat_ns": {
"min": 20000,
"max": 150000,
"mean": 75000,
"stddev": 15000
}
}
},
{
"jobname": "randwrite_test",
"write": {
"io_bytes": 10737418240,
"bw": 178957,
"iops": 44739,
"lat_ns": {
"min": 20000,
"max": 150000,
"mean": 75000,
"stddev": 15000
}
}
},
{
"jobname": "randwrite_test",
"write": {
"io_bytes": 10737418240,
"bw": 178957,
"iops": 44739,
"lat_ns": {
"min": 20000,
"max": 150000,
"mean": 75000,
"stddev": 15000
}
}
}
],
"global_data": {
"write": {
"io_bytes": 42949672960,
"bw": 715828,
"iops": 178957,
"lat_ns": {
"min": 20000,
"max": 150000,
"mean": 75000,
"stddev": 15000
}
}
}
}
Note: The example FIO output above shows ~179k IOPS for 4 jobs. To reach 500k+, you'd need to scale numjobs, iodepth, and potentially run multiple FIO instances across different pods/nodes, ensuring the underlying target and network can sustain the load.
Production Gotchas & Troubleshooting
-
Kernel Module Not Loaded:
- Symptom:
nvme connectfails with "No such device" orlsmod | grep nvme-tcpshows nothing. - Fix:
sudo modprobe nvme-tcpandsudo modprobe nvmet-tcp(on target). Ensure these are persistent across reboots by adding them to/etc/modules-load.d/nvme.conf.
- Symptom:
-
Network Connectivity Issues:
- Symptom:
nvme connecttimes out or fails to establish connection. - Fix: Verify IP addresses, subnet masks, and gateway. Check firewalls (
firewalld,ufw,iptables) on both target and initiator. Ensure port 4420 is open. Useping,traceroute,netcatto diagnose.
- Symptom:
-
NQN Mismatch:
- Symptom:
nvme connectfails with "Invalid NQN" or similar. - Fix: Double-check the NQN used in
nvme connectmatches exactly what was configured on the target (nvme target create -n ...). NQNs are case-sensitive.
- Symptom:
-
Multipath Configuration Errors:
- Symptom: Only one path is active, or
multipath -llshows separate devices instead of a single multipath device. - Fix: Ensure
multipathdis running. Verify/etc/multipath.confis correctly configured, especiallyuser_friendly_namesandfind_multipaths. Ensure all target IPs are connected with the same NQN. Restartmultipathdafter changes.
- Symptom: Only one path is active, or
-
Performance Bottlenecks:
- Symptom: IOPS or latency targets are not met.
- Fix:
- Target: Check CPU utilization (
top,htop), disk I/O (iostat -x 1), network I/O (sar -n DEV 1). Is the underlying NVMe SSD saturated? - Network: Check network link utilization (
iftop,nload). Are there packet drops? Is the switch port configured correctly (e.g., flow control, jumbo frames)? - Initiator: Check CPU utilization. Ensure
iodepthandnumjobsin FIO are high enough. Considerio_uringfor newer kernels. For SPDK, ensure CPU core isolation and huge pages are configured. - Kernel TCP Stack Tuning: For kernel NVMe/TCP, consider tuning
net.core.somaxconn,net.ipv4.tcp_tw_reuse,net.ipv4.tcp_max_syn_backlog,net.ipv4.tcp_fin_timeout.
- Target: Check CPU utilization (
-
CSI Driver Issues:
- Symptom: PVCs remain pending, pods fail to mount volumes.
- Fix: Check CSI driver controller and node plugin logs (
kubectl logs -n <csi-namespace> <pod-name>). VerifyCSIDriverandStorageClassdefinitions. Ensure necessary permissions for the CSI driver to executenvme-clicommands on worker nodes (e.g., via privileged containers or hostPath mounts for binaries).
Frequently Asked Questions
-
What is the primary advantage of NVMe/TCP over iSCSI? NVMe/TCP leverages the highly efficient, parallel NVMe command set, designed for modern SSDs, directly over TCP/IP. iSCSI uses the older SCSI command set, which introduces more overhead and is less optimized for the parallelism of NVMe devices. This results in significantly higher IOPS and lower latency for NVMe/TCP.
-
When should I consider SPDK over kernel NVMe/TCP? SPDK is ideal for extreme performance requirements where every microsecond of latency and every IOPS counts. It bypasses the kernel's TCP/IP stack and block layer, eliminating context switches and interrupt overheads by using user-space polling. This comes at the cost of increased complexity, dedicated CPU core allocation, and potentially higher memory consumption. For most high-performance applications, kernel NVMe/TCP with proper tuning is sufficient.
-
Is NVMe/TCP suitable for multi-tenant Kubernetes environments? Yes, with careful resource management. Each NVMe/TCP connection consumes resources on the worker node. A well-designed CSI driver can manage these connections. For multi-tenancy, consider using separate NVMe subsystems or namespaces for different tenants to provide isolation and enforce quotas. Network segmentation (VLANs, separate NICs) can further enhance security and performance isolation.
-
How does NVMe/TCP handle network failures? NVMe/TCP, like any TCP-based protocol, relies on TCP's reliability mechanisms. For high availability, multipathing is crucial. By configuring multiple network paths (e.g., multiple NICs, multiple target IPs),
multipathdon the initiator can automatically failover to an alternate path if one fails, ensuring continuous access to the storage. The NVMe-oF specification also includes mechanisms for controller-level failover. -
What are the security considerations for NVMe/TCP? NVMe/TCP, by default, does not include strong authentication or encryption. For production deployments, it's critical to:
- Network Isolation: Place NVMe/TCP traffic on a dedicated, isolated network segment (VLAN).
- IP Whitelisting: Configure target firewalls to only accept connections from authorized initiator IPs.
- TLS/DTLS: The NVMe-oF specification supports TLS/DTLS for encryption and authentication, though implementation varies. Ensure your NVMe/TCP target and initiator support and are configured for TLS if data-in-transit encryption is required.
- Host-level Authentication: Some implementations might support host-level authentication mechanisms like SASL.
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

Linux cgroups v2 in Kubernetes: PSI Pressure Stall Information, Memory Throttling & OOM Shields
Comprehensive guide covering linux cgroups v2 in kubernetes: psi pressure stall information, memory throttling & oom shields with production-grade architecture and code examples.
Read more
High-Performance Browser Storage: SQLite Wasm, Origin Private File System (OPFS) & Web Workers
Comprehensive guide covering high-performance browser storage: sqlite wasm, origin private file system (opfs) & web workers with production-grade architecture and code examples.
Read more
Continuous Dynamic Batching in LLM Inference: Orca, vLLM & TGI Latency Benchmarks
Comprehensive guide covering continuous dynamic batching in llm inference: orca, vllm & tgi latency benchmarks with production-grade architecture and code examples.
Read more