•11 min read

eBPF & Cilium in Kubernetes: High-Throughput Routing, Network Policies & Hubble Observability

eBPF & Cilium in Kubernetes: High-Throughput Routing, Network Policies & Hubble Observability

Kubernetes networking, traditionally reliant on kube-proxy and iptables, presents inherent limitations in scalability, performance, and observability. The iptables chain traversal overhead, coupled with kube-proxy's userspace proxying for NodePort and ExternalIPs, introduces latency and complexity. eBPF, the extended Berkeley Packet Filter, offers a paradigm shift by enabling programmable kernel-level packet processing without modifying kernel source code or loading kernel modules. Cilium leverages eBPF to deliver high-performance networking, robust security policies, and deep observability within Kubernetes.

This guide details the architectural advantages of deploying Cilium with eBPF in Kubernetes, focusing on replacing kube-proxy, optimizing data paths with XDP, enforcing L7 policies, and leveraging Hubble for comprehensive network visibility.

Audio Briefing
0:00 / 0:00

eBPF Fundamentals for Kubernetes Networking

eBPF programs execute in a sandboxed virtual machine within the Linux kernel. They can be attached to various hook points, including network interfaces (ingress/egress), system calls, and tracepoints. For networking, eBPF's primary utility stems from its ability to manipulate network packets directly in the kernel's data path, bypassing traditional network stack layers.

Cilium utilizes eBPF to:

  1. Replace kube-proxy: By attaching eBPF programs to network interfaces, Cilium handles Service load balancing directly in the kernel, eliminating iptables overhead and kube-proxy's userspace proxying. This includes ClusterIP, NodePort, ExternalIPs, and LoadBalancer Services.
  2. Implement Network Policies: eBPF programs enforce L3/L4 and L7 network policies with minimal overhead, directly at the packet ingress/egress points.
  3. Accelerate Data Path: Features like XDP (eXpress Data Path) allow eBPF programs to process packets even before the kernel's network stack, enabling ultra-low latency packet forwarding and DDoS mitigation.
  4. Provide Observability: eBPF programs can export rich metadata about network flows, enabling tools like Hubble to provide deep insights into network traffic.
Advertisement

Architectural Deep Dive: Cilium's eBPF Data Path

Cilium operates as a CNI (Container Network Interface) plugin. When a Pod is scheduled, Cilium configures its network interface and attaches eBPF programs.

kube-proxy Replacement

Cilium's kube-proxy replacement mode operates by installing eBPF programs on each node's network interfaces. These programs intercept traffic destined for Kubernetes Services.

For ClusterIP Services, the eBPF program performs NAT (Network Address Translation) and load balancing directly in the kernel. When a Pod initiates a connection to a Service IP, the eBPF program on the originating node's interface rewrites the destination IP to a backend Pod's IP and port, then forwards the packet. This occurs entirely within the kernel, avoiding context switches to userspace.

For NodePort Services, the eBPF program attached to the host's network interface intercepts incoming traffic on the NodePort. It then performs NAT to a backend Pod, similar to ClusterIP, but originating from the host's external interface.

# Example: Verify Cilium's kube-proxy replacement status
# This command checks if Cilium is managing kube-proxy functionality.
cilium status --verbose | grep KubeProxyReplacement

# Expected output indicating full replacement:
# KubeProxyReplacement:   Enabled (strict)

XDP (eXpress Data Path) Integration

XDP allows eBPF programs to run at the earliest possible point in the network driver's receive path, before the packet is even allocated a sk_buff (socket buffer) structure. This enables extremely high-performance packet processing, ideal for use cases like DDoS mitigation, load balancing, and high-throughput packet forwarding.

Cilium can leverage XDP for certain data path optimizations, particularly for NodePort and HostPort traffic, and for accelerating ingress traffic to Services. By dropping unwanted packets or forwarding desired ones directly from the XDP layer, significant CPU cycles can be saved.

# Example CiliumDaemonSet configuration snippet for enabling XDP
# This would be part of your Cilium installation YAML.
# Note: XDP requires specific network driver support.
apiVersion: helm.toolkit.fluxcd.io/v2beta1
kind: HelmRelease
metadata:
  name: cilium
  namespace: kube-system
spec:
  chart:
    spec:
      chart: cilium
      version: 1.15.x # Use your desired Cilium version
      sourceRef:
        kind: HelmRepository
        name: cilium
        namespace: flux-system
  interval: 1m
  values:
    kubeProxyReplacement: strict
    bpf:
      masquerade: true
      tproxy: false # TPROXY is for L7, not directly XDP
      # Enable XDP for NodePort if your kernel and NIC support it
      # This can significantly reduce latency for NodePort traffic.
      nodePort:
        enabled: true
        mode: xdp
    # ... other Cilium configurations

Socket-Level Load Balancing (Sockmap)

Beyond traditional packet-level load balancing, Cilium can utilize eBPF's sockmap feature for socket-level load balancing. sockmap allows eBPF programs to redirect TCP connections before they are fully established in the kernel's network stack. This is particularly beneficial for high-connection-rate Services, as it avoids the overhead of full TCP handshake processing on the initial receiving socket before redirection.

A sockmap eBPF program can intercept connect() or accept() calls and redirect the socket to another local socket or even another node's socket, effectively load balancing connections at a very early stage. This is distinct from kube-proxy's iptables DNAT, which operates on packets after the TCP handshake has begun.

L7 Network Policies Without Sidecars

Traditional L7 policy enforcement in Kubernetes often relies on sidecar proxies (e.g., Envoy in an Istio mesh). While powerful, sidecars introduce resource overhead (CPU, memory) and latency due to the additional hop. Cilium can enforce L7 policies for common protocols like HTTP, Kafka, and DNS directly using eBPF.

Cilium achieves this by attaching eBPF programs to the socket's sendmsg() and recvmsg() syscalls. These programs can inspect the application-layer payload (e.g., HTTP headers, Kafka topic names) and enforce policies based on that content. This is done without requiring a separate userspace proxy process per Pod.

# Example: CiliumNetworkPolicy for L7 HTTP enforcement
apiVersion: "cilium.io/v2"
kind: CiliumNetworkPolicy
metadata:
  name: allow-http-get-to-api
spec:
  endpointSelector:
    matchLabels:
      app: backend-api
  ingress:
  - fromEndpoints:
    - matchLabels:
        app: frontend
    toPorts:
    - ports:
      - port: "8080"
        protocol: TCP
      rules:
        http:
        - method: "GET"
          path: "/api/v1/data"
        - method: "POST"
          path: "/api/v1/submit"
          headers:
          - "Content-Type: application/json" # Example: enforce header presence

This policy allows frontend Pods to make GET requests to /api/v1/data and POST requests to /api/v1/submit on port 8080 of backend-api Pods. Any other HTTP method or path would be denied by the eBPF program.

Hubble Observability

Hubble is a distributed network and security observability platform built on top of Cilium and eBPF. It provides deep visibility into network flows, policy enforcement decisions, and DNS requests within the Kubernetes cluster.

Hubble leverages eBPF's ability to export flow metadata directly from the kernel. This data includes source/destination IPs, ports, protocols, Kubernetes Pod/Service/Namespace identities, and even L7 protocol details (e.g., HTTP method, path, status code).

Components of Hubble

  • Hubble Agent: Runs as an eBPF program on each node, collecting flow data.
  • Hubble Relay: A gRPC service that aggregates flow data from all Hubble Agents.
  • Hubble CLI: A command-line tool to query Hubble Relay.
  • Hubble UI: A web-based graphical interface for visualizing network flows and policies.

Setting up Hubble

Hubble is typically enabled during Cilium installation.

# Example: Enabling Hubble in Cilium HelmRelease
apiVersion: helm.toolkit.fluxcd.io/v2beta1
kind: HelmRelease
metadata:
  name: cilium
  namespace: kube-system
spec:
  chart:
    spec:
      chart: cilium
      version: 1.15.x
      sourceRef:
        kind: HelmRepository
        name: cilium
        namespace: flux-system
  interval: 1m
  values:
    hubble:
      enabled: true
      listenAddress: ":4244" # Default gRPC port for Hubble Relay
      ui:
        enabled: true
        service:
          type: ClusterIP # Or LoadBalancer for external access
    # ... other Cilium configurations

After installation, you can access Hubble UI via port-forwarding or a LoadBalancer Service.

# Port-forward to Hubble UI
kubectl port-forward -n kube-system svc/hubble-ui 8080:80

# Then open http://localhost:8080 in your browser.

# Using Hubble CLI to observe flows
# First, install the Hubble CLI:
# curl -L --remote-name-all https://github.com/cilium/hubble/releases/latest/download/hubble-linux-amd64.tar.gz{,.sha256sum}
# sha256sum --check hubble-linux-amd64.tar.gz.sha256sum
# sudo tar -C /usr/local/bin -xvf hubble-linux-amd64.tar.gz
# rm hubble-linux-amd64.tar.gz{,.sha256sum}

# Then, connect to Hubble Relay and observe flows:
hubble observe

# Filter by namespace and HTTP status code
hubble observe --namespace default --protocol http --http-status 2xx

# Observe DNS queries
hubble observe --type dns

Hubble provides microsecond-level latency visibility, showing exactly where packets are dropped or delayed, and which network policies are being enforced. This is invaluable for debugging complex microservice interactions.

Architecture Comparison: kube-proxy vs. Cilium eBPF

Featurekube-proxy (iptables mode)Cilium (eBPF mode)
Load Balancingiptables DNAT rules, userspace proxy for NodePortKernel eBPF programs, direct packet rewrite
Performanceiptables chain traversal overhead, userspace context switchesKernel-native, minimal overhead, XDP acceleration
Scalabilityiptables rule explosion with many Services/EndpointsScales well with eBPF maps, constant time lookups
Network PoliciesL3/L4 only, iptables rulesL3/L4 and L7 (HTTP, Kafka, DNS) with eBPF
ObservabilityLimited, relies on conntrack and iptables logsDeep, real-time flow visibility with Hubble (L3-L7)
Resource Usagekube-proxy daemon, iptables kernel overheadCilium agent daemon, eBPF programs in kernel
Kernel BypassNoYes, with XDP for specific use cases
SecurityBasic L3/L4 firewallAdvanced L3/L4/L7 policy enforcement, identity-based
Advertisement

Production Gotchas & Troubleshooting

  1. Kernel Version Compatibility: eBPF features evolve rapidly. Ensure your kernel version (typically 4.9+ for basic eBPF, 5.x+ for advanced features like sockmap, XDP) is compatible with your Cilium version.

    • Symptom: Cilium pods fail to start, cilium status shows errors related to eBPF program loading.
    • Fix: Check Cilium's release notes for minimum kernel requirements. Upgrade kernel if necessary, or use an older Cilium version if kernel upgrade is not feasible.
    • Command: uname -r to check kernel version.
  2. Network Interface Driver Support for XDP: XDP requires specific network card drivers that support the XDP API. Not all drivers are XDP-capable.

    • Symptom: XDP-related features in Cilium (e.g., nodePort.mode: xdp) don't activate or cause network issues.
    • Fix: Verify driver support. ethtool -i <interface> can show driver info. Consult cilium-health logs for XDP-related errors. If not supported, fall back to generic or native XDP modes if available, or disable XDP for that interface.
  3. kube-proxy Conflicts: If kube-proxy is not fully disabled or removed, it can conflict with Cilium's eBPF-based Service load balancing.

    • Symptom: Intermittent connectivity to Services, incorrect load balancing, iptables rules conflicting with Cilium.
    • Fix: Ensure kube-proxy is completely disabled or removed from your cluster. For kops, kubeadm, or cloud provider managed clusters, there are specific flags or configurations to achieve this.
      • For kubeadm: kubeadm init --config=kubeadm-config.yaml with kubeProxy.disabled: true in the KubeProxyConfiguration.
      • For kops: Set kubeProxy.enabled: false in your cluster spec.
      • For existing clusters: Scale kube-proxy deployment to 0 replicas and ensure no DaemonSet is recreating it.
  4. L7 Policy Misconfiguration: Incorrect L7 policies can lead to application connectivity issues that are hard to debug without Hubble.

    • Symptom: Applications fail to communicate, HTTP 403 errors, Kafka connection resets, but L3/L4 connectivity appears fine.
    • Fix: Use hubble observe --protocol http (or kafka, dns) to see policy enforcement decisions. Look for VERDICT: DENIED flows and the associated PolicyName. Adjust CiliumNetworkPolicy rules accordingly. Ensure toPorts and rules match application expectations.
  5. eBPF Map Exhaustion: In very large clusters with many Services, Endpoints, and policies, eBPF maps can reach their size limits.

    • Symptom: New policies or Services fail to apply, Cilium agent logs show errors about eBPF map full.
    • Fix: Cilium has internal mechanisms to manage map sizes. Ensure you are on a recent Cilium version. Consider optimizing your network policies to reduce complexity. For extreme cases, adjust Cilium's bpf.map.size configuration (with caution, as this consumes kernel memory).

Frequently Asked Questions

  1. Can Cilium replace kube-proxy entirely? Yes, Cilium can fully replace kube-proxy for ClusterIP, NodePort, ExternalIPs, and LoadBalancer Services by implementing all load balancing and NAT functionality directly in the kernel using eBPF. This is the recommended deployment model for performance and observability.

  2. What are the minimum kernel requirements for Cilium's advanced eBPF features? For basic eBPF functionality and kube-proxy replacement, Linux kernel 4.9+ is generally sufficient. For advanced features like XDP, sockmap, and certain L7 policy enforcements, kernel 5.x or newer (e.g., 5.4+ for sockmap, 5.10+ for specific XDP features) is often required or highly recommended for optimal performance and stability. Always consult the Cilium release notes for precise kernel version compatibility.

  3. How does Cilium handle L7 policies without a sidecar proxy? Cilium attaches eBPF programs to the sendmsg() and recvmsg() syscalls of application sockets. These programs can then inspect the application-layer payload (e.g., HTTP headers, Kafka topic names) and enforce policies based on the content, all within the kernel context. This avoids the resource overhead and latency introduced by userspace sidecar proxies.

  4. What is the performance impact of enabling Hubble? Hubble's data collection relies on eBPF programs that export flow metadata from the kernel. This process is highly optimized and has a minimal performance impact on the data plane. The primary resource consumption comes from the Hubble Relay and UI components, which aggregate and visualize the data. For high-traffic clusters, ensure Hubble Relay has sufficient CPU and memory resources.

  5. Is Cilium compatible with other CNI plugins? No, Cilium is a CNI plugin itself and is designed to be the sole CNI in a Kubernetes cluster. It manages all aspects of Pod networking, IP address management, and network policy enforcement. Attempting to run Cilium alongside another CNI will lead to conflicts and an inoperable network.

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement