eBPF & Cilium in Kubernetes: High-Throughput Routing, Network Policies & Hubble Observability

Table of Contents(12 sections)
Kubernetes networking, traditionally reliant on kube-proxy and iptables, presents inherent limitations in scalability, performance, and observability. The iptables chain traversal overhead, coupled with kube-proxy's userspace proxying for NodePort and ExternalIPs, introduces latency and complexity. eBPF, the extended Berkeley Packet Filter, offers a paradigm shift by enabling programmable kernel-level packet processing without modifying kernel source code or loading kernel modules. Cilium leverages eBPF to deliver high-performance networking, robust security policies, and deep observability within Kubernetes.
This guide details the architectural advantages of deploying Cilium with eBPF in Kubernetes, focusing on replacing kube-proxy, optimizing data paths with XDP, enforcing L7 policies, and leveraging Hubble for comprehensive network visibility.
eBPF Fundamentals for Kubernetes Networking
eBPF programs execute in a sandboxed virtual machine within the Linux kernel. They can be attached to various hook points, including network interfaces (ingress/egress), system calls, and tracepoints. For networking, eBPF's primary utility stems from its ability to manipulate network packets directly in the kernel's data path, bypassing traditional network stack layers.
Cilium utilizes eBPF to:
- Replace
kube-proxy: By attaching eBPF programs to network interfaces, Cilium handles Service load balancing directly in the kernel, eliminatingiptablesoverhead andkube-proxy's userspace proxying. This includesClusterIP,NodePort,ExternalIPs, andLoadBalancerServices. - Implement Network Policies: eBPF programs enforce L3/L4 and L7 network policies with minimal overhead, directly at the packet ingress/egress points.
- Accelerate Data Path: Features like XDP (eXpress Data Path) allow eBPF programs to process packets even before the kernel's network stack, enabling ultra-low latency packet forwarding and DDoS mitigation.
- Provide Observability: eBPF programs can export rich metadata about network flows, enabling tools like Hubble to provide deep insights into network traffic.
Architectural Deep Dive: Cilium's eBPF Data Path
Cilium operates as a CNI (Container Network Interface) plugin. When a Pod is scheduled, Cilium configures its network interface and attaches eBPF programs.
kube-proxy Replacement
Cilium's kube-proxy replacement mode operates by installing eBPF programs on each node's network interfaces. These programs intercept traffic destined for Kubernetes Services.
For ClusterIP Services, the eBPF program performs NAT (Network Address Translation) and load balancing directly in the kernel. When a Pod initiates a connection to a Service IP, the eBPF program on the originating node's interface rewrites the destination IP to a backend Pod's IP and port, then forwards the packet. This occurs entirely within the kernel, avoiding context switches to userspace.
For NodePort Services, the eBPF program attached to the host's network interface intercepts incoming traffic on the NodePort. It then performs NAT to a backend Pod, similar to ClusterIP, but originating from the host's external interface.
# Example: Verify Cilium's kube-proxy replacement status
# This command checks if Cilium is managing kube-proxy functionality.
cilium status --verbose | grep KubeProxyReplacement
# Expected output indicating full replacement:
# KubeProxyReplacement: Enabled (strict)
XDP (eXpress Data Path) Integration
XDP allows eBPF programs to run at the earliest possible point in the network driver's receive path, before the packet is even allocated a sk_buff (socket buffer) structure. This enables extremely high-performance packet processing, ideal for use cases like DDoS mitigation, load balancing, and high-throughput packet forwarding.
Cilium can leverage XDP for certain data path optimizations, particularly for NodePort and HostPort traffic, and for accelerating ingress traffic to Services. By dropping unwanted packets or forwarding desired ones directly from the XDP layer, significant CPU cycles can be saved.
# Example CiliumDaemonSet configuration snippet for enabling XDP
# This would be part of your Cilium installation YAML.
# Note: XDP requires specific network driver support.
apiVersion: helm.toolkit.fluxcd.io/v2beta1
kind: HelmRelease
metadata:
name: cilium
namespace: kube-system
spec:
chart:
spec:
chart: cilium
version: 1.15.x # Use your desired Cilium version
sourceRef:
kind: HelmRepository
name: cilium
namespace: flux-system
interval: 1m
values:
kubeProxyReplacement: strict
bpf:
masquerade: true
tproxy: false # TPROXY is for L7, not directly XDP
# Enable XDP for NodePort if your kernel and NIC support it
# This can significantly reduce latency for NodePort traffic.
nodePort:
enabled: true
mode: xdp
# ... other Cilium configurations
Socket-Level Load Balancing (Sockmap)
Beyond traditional packet-level load balancing, Cilium can utilize eBPF's sockmap feature for socket-level load balancing. sockmap allows eBPF programs to redirect TCP connections before they are fully established in the kernel's network stack. This is particularly beneficial for high-connection-rate Services, as it avoids the overhead of full TCP handshake processing on the initial receiving socket before redirection.
A sockmap eBPF program can intercept connect() or accept() calls and redirect the socket to another local socket or even another node's socket, effectively load balancing connections at a very early stage. This is distinct from kube-proxy's iptables DNAT, which operates on packets after the TCP handshake has begun.
L7 Network Policies Without Sidecars
Traditional L7 policy enforcement in Kubernetes often relies on sidecar proxies (e.g., Envoy in an Istio mesh). While powerful, sidecars introduce resource overhead (CPU, memory) and latency due to the additional hop. Cilium can enforce L7 policies for common protocols like HTTP, Kafka, and DNS directly using eBPF.
Cilium achieves this by attaching eBPF programs to the socket's sendmsg() and recvmsg() syscalls. These programs can inspect the application-layer payload (e.g., HTTP headers, Kafka topic names) and enforce policies based on that content. This is done without requiring a separate userspace proxy process per Pod.
# Example: CiliumNetworkPolicy for L7 HTTP enforcement
apiVersion: "cilium.io/v2"
kind: CiliumNetworkPolicy
metadata:
name: allow-http-get-to-api
spec:
endpointSelector:
matchLabels:
app: backend-api
ingress:
- fromEndpoints:
- matchLabels:
app: frontend
toPorts:
- ports:
- port: "8080"
protocol: TCP
rules:
http:
- method: "GET"
path: "/api/v1/data"
- method: "POST"
path: "/api/v1/submit"
headers:
- "Content-Type: application/json" # Example: enforce header presence
This policy allows frontend Pods to make GET requests to /api/v1/data and POST requests to /api/v1/submit on port 8080 of backend-api Pods. Any other HTTP method or path would be denied by the eBPF program.
Hubble Observability
Hubble is a distributed network and security observability platform built on top of Cilium and eBPF. It provides deep visibility into network flows, policy enforcement decisions, and DNS requests within the Kubernetes cluster.
Hubble leverages eBPF's ability to export flow metadata directly from the kernel. This data includes source/destination IPs, ports, protocols, Kubernetes Pod/Service/Namespace identities, and even L7 protocol details (e.g., HTTP method, path, status code).
Components of Hubble
- Hubble Agent: Runs as an eBPF program on each node, collecting flow data.
- Hubble Relay: A gRPC service that aggregates flow data from all Hubble Agents.
- Hubble CLI: A command-line tool to query Hubble Relay.
- Hubble UI: A web-based graphical interface for visualizing network flows and policies.
Setting up Hubble
Hubble is typically enabled during Cilium installation.
# Example: Enabling Hubble in Cilium HelmRelease
apiVersion: helm.toolkit.fluxcd.io/v2beta1
kind: HelmRelease
metadata:
name: cilium
namespace: kube-system
spec:
chart:
spec:
chart: cilium
version: 1.15.x
sourceRef:
kind: HelmRepository
name: cilium
namespace: flux-system
interval: 1m
values:
hubble:
enabled: true
listenAddress: ":4244" # Default gRPC port for Hubble Relay
ui:
enabled: true
service:
type: ClusterIP # Or LoadBalancer for external access
# ... other Cilium configurations
After installation, you can access Hubble UI via port-forwarding or a LoadBalancer Service.
# Port-forward to Hubble UI
kubectl port-forward -n kube-system svc/hubble-ui 8080:80
# Then open http://localhost:8080 in your browser.
# Using Hubble CLI to observe flows
# First, install the Hubble CLI:
# curl -L --remote-name-all https://github.com/cilium/hubble/releases/latest/download/hubble-linux-amd64.tar.gz{,.sha256sum}
# sha256sum --check hubble-linux-amd64.tar.gz.sha256sum
# sudo tar -C /usr/local/bin -xvf hubble-linux-amd64.tar.gz
# rm hubble-linux-amd64.tar.gz{,.sha256sum}
# Then, connect to Hubble Relay and observe flows:
hubble observe
# Filter by namespace and HTTP status code
hubble observe --namespace default --protocol http --http-status 2xx
# Observe DNS queries
hubble observe --type dns
Hubble provides microsecond-level latency visibility, showing exactly where packets are dropped or delayed, and which network policies are being enforced. This is invaluable for debugging complex microservice interactions.
Architecture Comparison: kube-proxy vs. Cilium eBPF
| Feature | kube-proxy (iptables mode) | Cilium (eBPF mode) |
|---|---|---|
| Load Balancing | iptables DNAT rules, userspace proxy for NodePort | Kernel eBPF programs, direct packet rewrite |
| Performance | iptables chain traversal overhead, userspace context switches | Kernel-native, minimal overhead, XDP acceleration |
| Scalability | iptables rule explosion with many Services/Endpoints | Scales well with eBPF maps, constant time lookups |
| Network Policies | L3/L4 only, iptables rules | L3/L4 and L7 (HTTP, Kafka, DNS) with eBPF |
| Observability | Limited, relies on conntrack and iptables logs | Deep, real-time flow visibility with Hubble (L3-L7) |
| Resource Usage | kube-proxy daemon, iptables kernel overhead | Cilium agent daemon, eBPF programs in kernel |
| Kernel Bypass | No | Yes, with XDP for specific use cases |
| Security | Basic L3/L4 firewall | Advanced L3/L4/L7 policy enforcement, identity-based |
Production Gotchas & Troubleshooting
-
Kernel Version Compatibility: eBPF features evolve rapidly. Ensure your kernel version (typically 4.9+ for basic eBPF, 5.x+ for advanced features like
sockmap, XDP) is compatible with your Cilium version.- Symptom: Cilium pods fail to start,
cilium statusshows errors related to eBPF program loading. - Fix: Check Cilium's release notes for minimum kernel requirements. Upgrade kernel if necessary, or use an older Cilium version if kernel upgrade is not feasible.
- Command:
uname -rto check kernel version.
- Symptom: Cilium pods fail to start,
-
Network Interface Driver Support for XDP: XDP requires specific network card drivers that support the XDP API. Not all drivers are XDP-capable.
- Symptom: XDP-related features in Cilium (e.g.,
nodePort.mode: xdp) don't activate or cause network issues. - Fix: Verify driver support.
ethtool -i <interface>can show driver info. Consultcilium-healthlogs for XDP-related errors. If not supported, fall back togenericornativeXDP modes if available, or disable XDP for that interface.
- Symptom: XDP-related features in Cilium (e.g.,
-
kube-proxyConflicts: Ifkube-proxyis not fully disabled or removed, it can conflict with Cilium's eBPF-based Service load balancing.- Symptom: Intermittent connectivity to Services, incorrect load balancing,
iptablesrules conflicting with Cilium. - Fix: Ensure
kube-proxyis completely disabled or removed from your cluster. Forkops,kubeadm, or cloud provider managed clusters, there are specific flags or configurations to achieve this.- For
kubeadm:kubeadm init --config=kubeadm-config.yamlwithkubeProxy.disabled: truein theKubeProxyConfiguration. - For
kops: SetkubeProxy.enabled: falsein your cluster spec. - For existing clusters: Scale
kube-proxydeployment to 0 replicas and ensure noDaemonSetis recreating it.
- For
- Symptom: Intermittent connectivity to Services, incorrect load balancing,
-
L7 Policy Misconfiguration: Incorrect L7 policies can lead to application connectivity issues that are hard to debug without Hubble.
- Symptom: Applications fail to communicate, HTTP 403 errors, Kafka connection resets, but L3/L4 connectivity appears fine.
- Fix: Use
hubble observe --protocol http(orkafka,dns) to see policy enforcement decisions. Look forVERDICT: DENIEDflows and the associatedPolicyName. AdjustCiliumNetworkPolicyrules accordingly. EnsuretoPortsandrulesmatch application expectations.
-
eBPF Map Exhaustion: In very large clusters with many Services, Endpoints, and policies, eBPF maps can reach their size limits.
- Symptom: New policies or Services fail to apply, Cilium agent logs show errors about eBPF map full.
- Fix: Cilium has internal mechanisms to manage map sizes. Ensure you are on a recent Cilium version. Consider optimizing your network policies to reduce complexity. For extreme cases, adjust Cilium's
bpf.map.sizeconfiguration (with caution, as this consumes kernel memory).
Frequently Asked Questions
-
Can Cilium replace
kube-proxyentirely? Yes, Cilium can fully replacekube-proxyforClusterIP,NodePort,ExternalIPs, andLoadBalancerServices by implementing all load balancing and NAT functionality directly in the kernel using eBPF. This is the recommended deployment model for performance and observability. -
What are the minimum kernel requirements for Cilium's advanced eBPF features? For basic eBPF functionality and
kube-proxyreplacement, Linux kernel 4.9+ is generally sufficient. For advanced features like XDP,sockmap, and certain L7 policy enforcements, kernel 5.x or newer (e.g., 5.4+ forsockmap, 5.10+ for specific XDP features) is often required or highly recommended for optimal performance and stability. Always consult the Cilium release notes for precise kernel version compatibility. -
How does Cilium handle L7 policies without a sidecar proxy? Cilium attaches eBPF programs to the
sendmsg()andrecvmsg()syscalls of application sockets. These programs can then inspect the application-layer payload (e.g., HTTP headers, Kafka topic names) and enforce policies based on the content, all within the kernel context. This avoids the resource overhead and latency introduced by userspace sidecar proxies. -
What is the performance impact of enabling Hubble? Hubble's data collection relies on eBPF programs that export flow metadata from the kernel. This process is highly optimized and has a minimal performance impact on the data plane. The primary resource consumption comes from the Hubble Relay and UI components, which aggregate and visualize the data. For high-traffic clusters, ensure Hubble Relay has sufficient CPU and memory resources.
-
Is Cilium compatible with other CNI plugins? No, Cilium is a CNI plugin itself and is designed to be the sole CNI in a Kubernetes cluster. It manages all aspects of Pod networking, IP address management, and network policy enforcement. Attempting to run Cilium alongside another CNI will lead to conflicts and an inoperable network.
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

Cilium vs Calico with eBPF: Kubernetes Network Throughput, Security & Service Mesh
Comprehensive guide covering cilium vs calico with ebpf: kubernetes network throughput, security & service mesh with production-grade architecture and code examples.
Read more
High-Performance Observability with eBPF in Kubernetes: Bypassing the Sidecar Tax
Deep dive into eBPF observability in Kubernetes: eliminating Envoy sidecars, kernel probes, BPF ring buffers, and zero-code telemetry instrumentation.
Read more
Kubernetes Gateway API in Production: Migrating from Ingress-Nginx with Envoy
Comprehensive guide covering kubernetes gateway api in production: migrating from ingress-nginx with envoy with production-grade architecture and code examples.
Read more