•17 min read

OpenTelemetry Collector & eBPF in Production: Zero-Code Tracing, Metrics & Prometheus Pipelines

OpenTelemetry Collector & eBPF in Production: Zero-Code Tracing, Metrics & Prometheus Pipelines

The OpenTelemetry Collector, when augmented with eBPF-based auto-instrumentation, offers a robust solution for comprehensive observability without requiring application code modifications. This guide details the deployment and configuration of such a system within a Kubernetes environment, focusing on zero-code tracing, metrics, and Prometheus pipelines.

Architecture Overview

The core architecture involves deploying the OpenTelemetry Collector as a DaemonSet on Kubernetes. This ensures a collector instance runs on every node, facilitating local data collection and processing. eBPF agents, often integrated into or alongside the collector, leverage kernel-level probes to capture network, process, and system calls, translating these into OpenTelemetry traces and metrics.

This setup typically involves:

  1. eBPF Agent: Deployed as a DaemonSet, often integrated with the OpenTelemetry Collector or as a sidecar. It uses eBPF programs to instrument kernel events, generating OpenTelemetry data. Examples include Pixie, Parca Agent, or custom eBPF solutions. For this guide, we'll assume an eBPF-enabled collector distribution or a separate eBPF agent that exports OTLP to the collector.
  2. OpenTelemetry Collector (Agent): Deployed as a DaemonSet on each Kubernetes node. It receives data from eBPF agents and optionally from applications instrumented with OpenTelemetry SDKs. It performs initial processing (batching, filtering, basic enrichment).
  3. OpenTelemetry Collector (Gateway): Deployed as a Deployment, often with multiple replicas, acting as a central aggregation point. It receives data from agent collectors, applies advanced processing (sampling, attribute manipulation, aggregation), and exports to various backends.
  4. Observability Backends: Trace, metrics, and log storage and visualization systems (e.g., Jaeger, Prometheus, Loki).
Advertisement

Kubernetes Deployment: OpenTelemetry Collector DaemonSet

We'll deploy the OpenTelemetry Collector as a DaemonSet. This ensures that a collector instance runs on every node, collecting data from local eBPF agents and applications.

# otel-collector-agent-daemonset.yaml
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: otel-collector-agent
  namespace: observability
  labels:
    app: otel-collector-agent
spec:
  selector:
    matchLabels:
      app: otel-collector-agent
  template:
    metadata:
      labels:
        app: otel-collector-agent
    spec:
      serviceAccountName: otel-collector-agent
      hostNetwork: true # Required for eBPF agents to capture all network traffic
      dnsPolicy: ClusterFirstWithHostNet # Required with hostNetwork
      containers:
        - name: otel-collector
          image: otel/opentelemetry-collector-contrib:0.90.1 # Use contrib for more receivers/processors
          command: ["/otelcol", "--config=/conf/otel-collector-config.yaml"]
          securityContext:
            privileged: true # Required for eBPF agent capabilities
            runAsUser: 0
            runAsGroup: 0
          volumeMounts:
            - name: otel-collector-config
              mountPath: /conf
            - name: varlog
              mountPath: /var/log # For host log collection
              readOnly: true
            - name: varlibdockercontainers
              mountPath: /var/lib/docker/containers # For Docker container logs
              readOnly: true
            - name: sysfs
              mountPath: /sys # For eBPF kernel access
              readOnly: true
            - name: procfs
              mountPath: /proc # For eBPF process info
              readOnly: true
          env:
            - name: KUBERNETES_NODE_NAME
              valueFrom:
                fieldRef:
                  fieldPath: spec.nodeName
          ports:
            - name: otlp-grpc
              containerPort: 4317
              hostPort: 4317 # Expose OTLP gRPC on host
            - name: otlp-http
              containerPort: 4318
              hostPort: 4318 # Expose OTLP HTTP on host
            - name: prometheus
              containerPort: 8888 # Collector's own metrics
            - name: health-check
              containerPort: 13133
          livenessProbe:
            httpGet:
              path: /health
              port: 13133
            initialDelaySeconds: 5
            periodSeconds: 10
          readinessProbe:
            httpGet:
              path: /health
              port: 13133
            initialDelaySeconds: 5
            periodSeconds: 10
          resources:
            limits:
              cpu: 500m
              memory: 512Mi
            requests:
              cpu: 100m
              memory: 256Mi
      volumes:
        - name: otel-collector-config
          configMap:
            name: otel-collector-agent-config
        - name: varlog
          hostPath:
            path: /var/log
        - name: varlibdockercontainers
          hostPath:
            path: /var/lib/docker/containers
        - name: sysfs
          hostPath:
            path: /sys
        - name: procfs
          hostPath:
            path: /proc
---
# otel-collector-agent-serviceaccount.yaml
apiVersion: v1
kind: ServiceAccount
metadata:
  name: otel-collector-agent
  namespace: observability
---
# otel-collector-agent-clusterrole.yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: otel-collector-agent
rules:
  - apiGroups: [""]
    resources: ["nodes", "nodes/proxy", "pods", "services", "endpoints"]
    verbs: ["get", "list", "watch"]
  - apiGroups: ["apps"]
    resources: ["replicasets"]
    verbs: ["get", "list", "watch"]
  - apiGroups: ["extensions"]
    resources: ["daemonsets", "deployments", "replicasets"]
    verbs: ["get", "list", "watch"]
  - apiGroups: [""]
    resources: ["events"]
    verbs: ["create", "patch"]
  - apiGroups: ["policy"]
    resources: ["podsecuritypolicies"]
    verbs: ["use"]
    resourceNames:
      - otel-collector-agent # If using PSPs
---
# otel-collector-agent-clusterrolebinding.yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: otel-collector-agent
subjects:
  - kind: ServiceAccount
    name: otel-collector-agent
    namespace: observability
roleRef:
  kind: ClusterRole
  name: otel-collector-agent
  apiGroup: rbac.authorization.k8s.io

Explanation of DaemonSet Configuration:

  • hostNetwork: true: Essential for eBPF agents to observe all network traffic on the host. This allows the collector to bind to host ports and capture traffic across all pods on the node.
  • privileged: true: Often required for eBPF agents to load kernel modules and access sensitive kernel interfaces. This grants the container broad capabilities.
  • volumeMounts for /var/log, /var/lib/docker/containers, /sys, /proc: These host paths are mounted to allow the collector (or an integrated eBPF agent) to access host logs, container logs, and kernel/process information necessary for eBPF operations.
  • ports: Exposes OTLP gRPC and HTTP ports on the host network, allowing applications or eBPF agents to send data directly to the collector on the same node.

OpenTelemetry Collector Configuration (Agent)

This configuration demonstrates how the agent collector receives data, processes it, and forwards it to a gateway collector. It includes a basic eBPF receiver (conceptual, as specific eBPF receivers vary), host metrics, and log collection.

# otel-collector-agent-config.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: otel-collector-agent-config
  namespace: observability
data:
  otel-collector-config.yaml: |
    receivers:
      # OTLP receiver for applications and eBPF agents exporting directly
      otlp:
        protocols:
          grpc:
          http:

      # Host metrics receiver
      hostmetrics:
        collection_interval: 10s
        scrapers:
          cpu:
          memory:
          disk:
          filesystem:
          network:
          load:
          paging:
          processes:

      # Filelog receiver for host logs (e.g., systemd, kernel logs)
      # Requires hostPath mount for /var/log
      filelog:
        include:
          - /var/log/*.log
          - /var/log/*/*.log
        start_at: beginning
        poll_interval: 1s
        operators:
          - type: json_parser
            output: json
          - type: regex_parser
            regex: '^(?P<time>\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{3}Z)\s(?P<level>\w+)\s(?P<message>.*)$'
            output: json
            timestamp:
              parse_from: time
              layout: '%Y-%m-%dT%H:%M:%S.%LZ'
            severity:
              parse_from: level
              preset: syslog
          - type: move
            from: json.message
            to: body
          - type: remove
            field: json

      # Kubernetes Container Logs (Docker/CRI-O)
      # Requires hostPath mount for /var/lib/docker/containers
      # This is a basic example; for production, consider a dedicated log agent or more robust K8s integration
      filelog/container:
        include:
          - /var/lib/docker/containers/*/*-json.log
        start_at: beginning
        poll_interval: 1s
        operators:
          - type: json_parser
            output: json
          - type: move
            from: json.log
            to: body
          - type: move
            from: json.stream
            to: attributes.log.stream
          - type: move
            from: json.time
            to: attributes.log.time
          - type: add
            field: attributes.log.file.path
            value: "{{ .file.path }}"
          - type: add
            field: attributes.container.id
            value: "{{ .file.name }}" # Extract container ID from filename
          - type: add
            field: attributes.k8s.node.name
            value: "${KUBERNETES_NODE_NAME}" # Injected via env var
          - type: regex_parser
            regex: '^/var/lib/docker/containers/(?P<container_id>[a-f0-9]{64})/(?P<container_id_short>[a-f0-9]{12})-json.log$'
            parse_from: attributes.log.file.path
            output: attributes
          - type: remove
            field: json

    processors:
      # Batching for efficiency
      batch:
        send_batch_size: 1024
        timeout: 5s

      # Resource detection for Kubernetes metadata
      resourcedetection:
        detectors: ["system", "env", "kubernetes"]
        timeout: 2s
        kubernetes:
          pod_association:
            - from: "ip"
            - from: "cgroup"
          exclude:
            pods:
              - name: "otel-collector-agent" # Exclude self-instrumentation

      # Memory limiter to prevent OOMs
      memory_limiter:
        check_interval: 1s
        limit_mib: 256
        spike_limit_mib: 64

      # Tail-based sampling (example, typically done at gateway)
      # For agent, head-based is more common if sampling is needed here
      # This example is for demonstration, usually agents forward all data.
      # tail_sampling:
      #   decision_wait: 10s
      #   num_traces: 100000
      #   expected_new_traces_per_sec: 100
      #   policies:
      #     [
      #       {
      #         name: "error-policy",
      #         type: "status_code",
      #         status_code: { status_codes: ["ERROR"] }
      #       },
      #       {
      #         name: "latency-policy",
      #         type: "latency",
      #         latency: { threshold_ms: 500 }
      #       }
      #     ]

    exporters:
      # Export to a central OpenTelemetry Collector Gateway
      otlp:
        endpoint: "otel-collector-gateway.observability.svc.cluster.local:4317" # Internal K8s service
        tls:
          insecure: true # Use mTLS in production

      # Optional: Prometheus exporter for collector's own metrics
      prometheus:
        endpoint: "0.0.0.0:8888"

    service:
      telemetry:
        metrics:
          address: 0.0.0.0:8888
      pipelines:
        traces:
          receivers: [otlp]
          processors: [resourcedetection, batch, memory_limiter] # Add tail_sampling if needed
          exporters: [otlp]
        metrics:
          receivers: [otlp, hostmetrics]
          processors: [resourcedetection, batch, memory_limiter]
          exporters: [otlp]
        logs:
          receivers: [otlp, filelog, filelog/container]
          processors: [resourcedetection, batch, memory_limiter]
          exporters: [otlp]

Key Configuration Elements:

  • receivers:
    • otlp: Accepts OTLP data (traces, metrics, logs) from applications or eBPF agents.
    • hostmetrics: Scrapes CPU, memory, disk, network, etc., from the host.
    • filelog: Collects logs from specified host paths, parsing them into structured logs. This is crucial for zero-code log aggregation.
  • processors:
    • batch: Batches data for efficient export.
    • resourcedetection: Enriches telemetry with Kubernetes metadata (pod name, namespace, node name, etc.), vital for context.
    • memory_limiter: Prevents the collector from consuming excessive memory, crucial for DaemonSets.
  • exporters:
    • otlp: Forwards all processed telemetry data to a central OpenTelemetry Collector Gateway.
    • prometheus: Exposes the collector's internal metrics in Prometheus format.
  • service.pipelines: Defines how data flows through receivers, processors, and exporters for traces, metrics, and logs.

OpenTelemetry Collector Gateway Deployment & Configuration

The gateway collector aggregates data from all agent collectors, applies advanced processing like sampling, and exports to final backends.

# otel-collector-gateway-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: otel-collector-gateway
  namespace: observability
  labels:
    app: otel-collector-gateway
spec:
  replicas: 2 # Scale as needed
  selector:
    matchLabels:
      app: otel-collector-gateway
  template:
    metadata:
      labels:
        app: otel-collector-gateway
    spec:
      containers:
        - name: otel-collector
          image: otel/opentelemetry-collector-contrib:0.90.1
          command: ["/otelcol", "--config=/conf/otel-collector-config.yaml"]
          volumeMounts:
            - name: otel-collector-config
              mountPath: /conf
          ports:
            - name: otlp-grpc
              containerPort: 4317
            - name: otlp-http
              containerPort: 4318
            - name: prometheus
              containerPort: 8888
            - name: health-check
              containerPort: 13133
          livenessProbe:
            httpGet:
              path: /health
              port: 13133
            initialDelaySeconds: 5
            periodSeconds: 10
          readinessProbe:
            httpGet:
              path: /health
              port: 13133
            initialDelaySeconds: 5
            periodSeconds: 10
          resources:
            limits:
              cpu: 1000m
              memory: 2Gi
            requests:
              cpu: 200m
              memory: 512Mi
      volumes:
        - name: otel-collector-config
          configMap:
            name: otel-collector-gateway-config
---
# otel-collector-gateway-service.yaml
apiVersion: v1
kind: Service
metadata:
  name: otel-collector-gateway
  namespace: observability
spec:
  selector:
    app: otel-collector-gateway
  ports:
    - name: otlp-grpc
      protocol: TCP
      port: 4317
      targetPort: 4317
    - name: otlp-http
      protocol: TCP
      port: 4318
      targetPort: 4318
    - name: prometheus
      protocol: TCP
      port: 8888
      targetPort: 8888
Advertisement

OpenTelemetry Collector Configuration (Gateway)

This configuration includes advanced processors like tail-based sampling and exports to various backends.

# otel-collector-gateway-config.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: otel-collector-gateway-config
  namespace: observability
data:
  otel-collector-config.yaml: |
    receivers:
      otlp:
        protocols:
          grpc:
          http:

    processors:
      batch:
        send_batch_size: 1024
        timeout: 5s

      memory_limiter:
        check_interval: 1s
        limit_mib: 1500
        spike_limit_mib: 500

      # Tail-based sampling for traces
      # This is crucial for reducing trace volume while retaining important traces.
      tail_sampling:
        decision_wait: 10s # How long to wait for all spans of a trace
        num_traces: 100000 # Max number of traces to keep in memory for sampling
        expected_new_traces_per_sec: 1000 # Estimate of new traces per second
        policies:
          - name: "always-sample"
            type: "always_on" # Always sample 100% of traces (for dev/low traffic)
          - name: "error-policy"
            type: "status_code"
            status_code: { status_codes: ["ERROR", "UNSET"] } # Sample traces with errors
          - name: "latency-policy"
            type: "latency"
            latency: { threshold_ms: 500 } # Sample traces exceeding 500ms
          - name: "probabilistic-policy"
            type: "probabilistic"
            probabilistic: { sampling_percentage: 10 } # Sample 10% of remaining traces

      # Attributes processor for general attribute manipulation
      attributes:
        actions:
          - key: "service.namespace"
            action: "insert"
            value: "default" # Default namespace if not present
          - key: "k8s.cluster.name"
            action: "insert"
            value: "my-prod-cluster" # Add cluster name
          - key: "host.ip"
            action: "delete" # Remove sensitive host IP if not needed downstream

      # Resource processor for adding/modifying resource attributes
      resource:
        attributes:
          - key: "cloud.provider"
            value: "aws"
            action: "insert"
          - key: "cloud.region"
            value: "us-east-1"
            action: "insert"

    exporters:
      # Export traces to Jaeger
      jaeger:
        endpoint: "jaeger-collector.observability.svc.cluster.local:14250" # gRPC
        tls:
          insecure: true

      # Export metrics to Prometheus remote write (e.g., Mimir, Thanos)
      prometheusremotewrite:
        endpoint: "http://prometheus-mimir-gateway.observability.svc.cluster.local:9009/api/v1/push"
        headers:
          X-Scope-OrgID: "locionic"
        # auth:
        #   oauth2:
        #     client_id: "..."
        #     client_secret: "..."
        #     token_url: "..."

      # Export logs to Loki
      loki:
        endpoint: "http://loki.observability.svc.cluster.local:3100/loki/api/v1/push"
        tls:
          insecure: true
        # auth:
        #   basic:
        #     username: "..."
        #     password: "..."
        labels:
          attributes:
            - host.name
            - k8s.namespace.name
            - k8s.pod.name
            - service.name
          resource:
            - k8s.node.name
            - k8s.cluster.name

      # Optional: Prometheus exporter for collector's own metrics
      prometheus:
        endpoint: "0.0.0.0:8888"

    service:
      telemetry:
        metrics:
          address: 0.0.0.0:8888
      pipelines:
        traces:
          receivers: [otlp]
          processors: [memory_limiter, batch, tail_sampling, attributes, resource]
          exporters: [jaeger]
        metrics:
          receivers: [otlp]
          processors: [memory_limiter, batch, attributes, resource]
          exporters: [prometheusremotewrite]
        logs:
          receivers: [otlp]
          processors: [memory_limiter, batch, attributes, resource]
          exporters: [loki]

Sampling Strategies: Head-based vs. Tail-based

FeatureHead-based SamplingTail-based Sampling
Decision PointAt the start of a trace (first span)At the end of a trace (all spans collected)
Data RequiredOnly the first span's contextAll spans belonging to a trace
ImplementationApplication SDKs, OTel Collector AgentOTel Collector Gateway
ProsLow overhead, simple, reduces network traffic earlyContext-aware (errors, latency, specific attributes)
More intelligent sampling decisions
ConsNot context-aware, may drop important tracesHigher resource usage (memory, CPU) on collector
Requires all spans of a trace to reach collector
Introduces latency before decision is made
Use CaseHigh-volume, basic sampling, initial filteringProduction environments, complex sampling logic
Focus on critical traces (errors, slow requests)

Recommendation: Use head-based sampling at the application or agent level for very high-volume, non-critical traces (e.g., 1% probabilistic sampling). Implement tail-based sampling at the gateway collector for intelligent, context-aware sampling of critical traces (errors, high latency, specific user IDs).

Prometheus Metrics Export

The OpenTelemetry Collector can act as a Prometheus scraper target or export metrics via Prometheus remote write.

1. Collector as a Prometheus Scrape Target (for its own metrics): The prometheus exporter in both agent and gateway configurations exposes the collector's internal metrics (e.g., otelcol_receiver_accepted_spans_total). Prometheus can then scrape these endpoints.

# Example Prometheus scrape config for OTel Collector Gateway
# Add this to your Prometheus server's scrape_configs
- job_name: 'otel-collector-gateway'
  kubernetes_sd_configs:
    - role: endpoints
      namespaces:
        names: ['observability']
  relabel_configs:
    - source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
      action: keep
      regex: otel-collector-gateway;prometheus
    - source_labels: [__meta_kubernetes_pod_name]
      target_label: kubernetes_pod_name
      action: replace
    - source_labels: [__meta_kubernetes_namespace]
      target_label: kubernetes_namespace
      action: replace

2. Exporting Application Metrics via Prometheus Remote Write: The prometheusremotewrite exporter in the gateway configuration sends processed metrics to a Prometheus-compatible remote storage (e.g., Mimir, Thanos, Cortex). This centralizes metrics storage and allows for long-term retention and high availability.

Production Gotchas & Troubleshooting

  1. eBPF Permissions (privileged: true, hostNetwork: true):

    • Gotcha: Without privileged: true and hostNetwork: true, eBPF agents will fail to load programs or capture network traffic. This is a common misconfiguration.
    • Fix: Ensure the DaemonSet has privileged: true in securityContext and hostNetwork: true at the pod level. Also, verify the ServiceAccount has necessary ClusterRole permissions (e.g., nodes/proxy).
    • Troubleshooting: Check pod logs for permission denied, operation not permitted, or failed to load eBPF program errors. Use kubectl describe pod <pod-name> to verify security context and network settings.
  2. Memory Limits and OOMKilled Collectors:

    • Gotcha: OpenTelemetry Collectors, especially gateways with tail-based sampling, can consume significant memory, leading to OOMKills.
    • Fix:
      • Implement memory_limiter processor in the collector configuration.
      • Set appropriate resources.limits.memory in the Kubernetes manifests.
      • Tune tail_sampling parameters (num_traces, expected_new_traces_per_sec) to match your trace volume and available memory.
      • Scale out gateway collectors if a single instance cannot handle the load.
    • Troubleshooting: Check kubectl describe pod <pod-name> for OOMKilled status. Monitor collector memory usage via its own Prometheus metrics (otelcol_processor_memory_limiter_memory_usage_bytes).
  3. Network Connectivity Issues (Agent to Gateway, Gateway to Backends):

    • Gotcha: Incorrect service names, ports, or firewall rules can prevent data flow.
    • Fix:
      • Verify endpoint values in exporters (e.g., otel-collector-gateway.observability.svc.cluster.local:4317).
      • Ensure Kubernetes Services are correctly defined and target the collector pods.
      • Check network policies if they are in use.
    • Troubleshooting:
      • Check collector logs for connection refused, unavailable, or timeout errors.
      • Use kubectl exec -it <collector-pod> -- curl <target-endpoint> (if curl is available in image) or nc -vz <target-host> <target-port> from within the collector pod to test connectivity.
      • Verify DNS resolution from within the pod.
  4. Missing or Incomplete Telemetry Data:

    • Gotcha: Data might be dropped due to misconfigured receivers, processors, or exporters, or due to aggressive sampling.
    • Fix:
      • Review service.pipelines to ensure all desired receivers, processors, and exporters are correctly linked.
      • Check include/exclude rules in receivers (e.g., filelog).
      • Adjust sampling policies if too much data is being dropped.
      • Ensure resourcedetection is correctly configured to enrich data with necessary metadata.
    • Troubleshooting:
      • Monitor collector's internal metrics (otelcol_receiver_accepted_spans_total, otelcol_exporter_sent_spans_total, otelcol_processor_batch_batch_send_size_sum) to identify where data might be dropping.
      • Enable debug logging for the collector (--config=/conf/otel-collector-config.yaml --set service.telemetry.logs.level=debug) to get more verbose output.
  5. eBPF Agent Compatibility and Kernel Versions:

    • Gotcha: eBPF programs are kernel-version sensitive. An eBPF agent compiled for one kernel version might not work on another, or might require specific kernel headers.
    • Fix:
      • Use eBPF agents that support CO-RE (Compile Once – Run Everywhere) or are distributed with pre-compiled binaries for common kernel versions.
      • Ensure your Kubernetes nodes are running a compatible kernel version.
      • Some eBPF solutions require specific kernel modules or features to be enabled.
    • Troubleshooting: Look for errors like BPF program load failed, invalid argument, or kernel version mismatch in the eBPF agent's logs. Consult the documentation of your specific eBPF agent.

Frequently Asked Questions

  1. Q: What is the performance overhead of using eBPF for auto-instrumentation? A: eBPF-based instrumentation generally has a very low overhead, typically in the single-digit percentage range for CPU and memory. This is because eBPF programs run directly in the kernel, avoiding context switches and user-space overhead. However, the overhead can increase with the complexity and frequency of the eBPF probes and the volume of data being collected. Aggressive data processing and high cardinality attributes in the collector can also contribute to higher resource usage.

  2. Q: Can I use eBPF auto-instrumentation alongside traditional OpenTelemetry SDK instrumentation? A: Yes, absolutely. This is a common and recommended pattern. eBPF provides "zero-code" visibility into network interactions, system calls, and process execution, filling gaps where application code might not be instrumented. OpenTelemetry SDKs provide richer, semantic context directly from the application's business logic. The OpenTelemetry Collector can merge and correlate data from both sources, providing a more complete picture.

  3. Q: How do I ensure my eBPF agent is collecting data from all my applications in Kubernetes? A: Deploying the eBPF agent (or an eBPF-enabled collector) as a DaemonSet ensures it runs on every node. With hostNetwork: true and privileged: true, it gains the necessary access to monitor all processes and network traffic on that node, regardless of which pod or namespace they belong to. Resource detection processors in the collector then enrich this data with Kubernetes metadata to link it back to specific pods and services.

  4. Q: What's the best way to handle high-volume trace data to avoid overwhelming my backend? A: Implement a multi-stage sampling strategy:

    • Head-based sampling (probabilistic) at the application or agent level for initial reduction of non-critical traces.
    • Tail-based sampling at the OpenTelemetry Collector Gateway. This allows for intelligent sampling based on trace attributes like errors, latency, or specific business logic, ensuring critical traces are always captured.
    • Batching and memory limiting in the collector are also crucial to manage throughput and resource usage.
  5. Q: How does the OpenTelemetry Collector handle logs without application changes? A: The filelog receiver in the OpenTelemetry Collector can be configured to scrape logs directly from host paths (e.g., /var/log for system logs, /var/lib/docker/containers for container logs). By mounting these host paths into the collector DaemonSet, the collector can read, parse, and process these logs, then export them to a log backend like Loki or Elasticsearch, all without requiring any modifications to the application code itself. This provides a unified log collection pipeline alongside traces and metrics.

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement