•24 min read

OpenTelemetry CollectorとeBPFを本番環境で活用:コード不要のトレース、メトリクス、Prometheusパイプライン

OpenTelemetry CollectorとeBPFを本番環境で活用:コード不要のトレース、メトリクス、Prometheusパイプライン

OpenTelemetry CollectorをeBPFベースの自動計装で強化すると、アプリケーションコードの変更なしに包括的な可観測性(observability)を実現する堅牢なソリューションが提供されます。このガイドでは、Kubernetes環境におけるそのようなシステムのデプロイと設定について、ゼロコードのトレーシング、メトリクス、Prometheusパイプラインに焦点を当てて詳しく説明します。

アーキテクチャの概要

コアとなるアーキテクチャは、OpenTelemetry CollectorをKubernetes上にDaemonSetとしてデプロイすることを含みます。これにより、すべてのノードでコレクターインスタンスが実行され、ローカルデータの収集と処理が容易になります。eBPFエージェントは、多くの場合コレクターに統合されるか、コレクターと並行してデプロイされ、カーネルレベルのプローブを利用してネットワーク、プロセス、システムコールをキャプチャし、これらをOpenTelemetryのトレースとメトリクスに変換します。

このセットアップには通常、以下が含まれます。

  1. eBPFエージェント: DaemonSetとしてデプロイされ、OpenTelemetry Collectorと統合されるか、サイドカーとして機能します。eBPFプログラムを使用してカーネルイベントを計装し、OpenTelemetryデータを生成します。例としては、Pixie、Parca Agent、またはカスタムeBPFソリューションがあります。このガイドでは、eBPF対応のコレクターディストリビューション、またはOTLPをコレクターにエクスポートする独立したeBPFエージェントを想定します。
  2. OpenTelemetry Collector (エージェント): 各KubernetesノードにDaemonSetとしてデプロイされます。eBPFエージェントからデータを受信し、オプションでOpenTelemetry SDKで計装されたアプリケーションからもデータを受信します。初期処理(バッチ処理、フィルタリング、基本的なエンリッチメント)を実行します。
  3. OpenTelemetry Collector (ゲートウェイ): Deploymentとしてデプロイされ、多くの場合複数のレプリカを持ち、中央集約ポイントとして機能します。エージェントコレクターからデータを受信し、高度な処理(サンプリング、属性操作、集約)を適用し、様々なバックエンドにエクスポートします。
  4. 可観測性バックエンド: トレース、メトリクス、ログのストレージおよび可視化システム(例: Jaeger、Prometheus、Loki)。
Advertisement

Kubernetesデプロイメント: OpenTelemetry Collector DaemonSet

OpenTelemetry CollectorをDaemonSetとしてデプロイします。これにより、すべてのノードでコレクターインスタンスが実行され、ローカルのeBPFエージェントやアプリケーションからデータを収集します。

# otel-collector-agent-daemonset.yaml
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: otel-collector-agent
  namespace: observability
  labels:
    app: otel-collector-agent
spec:
  selector:
    matchLabels:
      app: otel-collector-agent
  template:
    metadata:
      labels:
        app: otel-collector-agent
    spec:
      serviceAccountName: otel-collector-agent
      hostNetwork: true # Required for eBPF agents to capture all network traffic
      dnsPolicy: ClusterFirstWithHostNet # Required with hostNetwork
      containers:
        - name: otel-collector
          image: otel/opentelemetry-collector-contrib:0.90.1 # Use contrib for more receivers/processors
          command: ["/otelcol", "--config=/conf/otel-collector-config.yaml"]
          securityContext:
            privileged: true # Required for eBPF agent capabilities
            runAsUser: 0
            runAsGroup: 0
          volumeMounts:
            - name: otel-collector-config
              mountPath: /conf
            - name: varlog
              mountPath: /var/log # For host log collection
              readOnly: true
            - name: varlibdockercontainers
              mountPath: /var/lib/docker/containers # For Docker container logs
              readOnly: true
            - name: sysfs
              mountPath: /sys # For eBPF kernel access
              readOnly: true
            - name: procfs
              mountPath: /proc # For eBPF process info
              readOnly: true
          env:
            - name: KUBERNETES_NODE_NAME
              valueFrom:
                fieldRef:
                  fieldPath: spec.nodeName
          ports:
            - name: otlp-grpc
              containerPort: 4317
              hostPort: 4317 # Expose OTLP gRPC on host
            - name: otlp-http
              containerPort: 4318
              hostPort: 4318 # Expose OTLP HTTP on host
            - name: prometheus
              containerPort: 8888 # Collector's own metrics
            - name: health-check
              containerPort: 13133
          livenessProbe:
            httpGet:
              path: /health
              port: 13133
            initialDelaySeconds: 5
            periodSeconds: 10
          readinessProbe:
            httpGet:
              path: /health
              port: 13133
            initialDelaySeconds: 5
            periodSeconds: 10
          resources:
            limits:
              cpu: 500m
              memory: 512Mi
            requests:
              cpu: 100m
              memory: 256Mi
      volumes:
        - name: otel-collector-config
          configMap:
            name: otel-collector-agent-config
        - name: varlog
          hostPath:
            path: /var/log
        - name: varlibdockercontainers
          hostPath:
            path: /var/lib/docker/containers
        - name: sysfs
          hostPath:
            path: /sys
        - name: procfs
          hostPath:
            path: /proc
---
# otel-collector-agent-serviceaccount.yaml
apiVersion: v1
kind: ServiceAccount
metadata:
  name: otel-collector-agent
  namespace: observability
---
# otel-collector-agent-clusterrole.yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: otel-collector-agent
rules:
  - apiGroups: [""]
    resources: ["nodes", "nodes/proxy", "pods", "services", "endpoints"]
    verbs: ["get", "list", "watch"]
  - apiGroups: ["apps"]
    resources: ["replicasets"]
    verbs: ["get", "list", "watch"]
  - apiGroups: ["extensions"]
    resources: ["daemonsets", "deployments", "replicasets"]
    verbs: ["get", "list", "watch"]
  - apiGroups: [""]
    resources: ["events"]
    verbs: ["create", "patch"]
  - apiGroups: ["policy"]
    resources: ["podsecuritypolicies"]
    verbs: ["use"]
    resourceNames:
      - otel-collector-agent # If using PSPs
---
# otel-collector-agent-clusterrolebinding.yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: otel-collector-agent
subjects:
  - kind: ServiceAccount
    name: otel-collector-agent
    namespace: observability
roleRef:
  kind: ClusterRole
  name: otel-collector-agent
  apiGroup: rbac.authorization.k8s.io

DaemonSet設定の説明:

  • hostNetwork: true: eBPFエージェントがホスト上のすべてのネットワークトラフィックを監視するために不可欠です。これにより、コレクターはホストポートにバインドし、ノード上のすべてのPodのトラフィックをキャプチャできます。
  • privileged: true: eBPFエージェントがカーネルモジュールをロードし、機密性の高いカーネルインターフェースにアクセスするためにしばしば必要です。これにより、コンテナに広範な機能が付与されます。
  • volumeMounts for /var/log, /var/lib/docker/containers, /sys, /proc: これらのホストパスは、コレクター(または統合されたeBPFエージェント)がeBPF操作に必要なホストログ、コンテナログ、カーネル/プロセス情報にアクセスできるようにマウントされます。
  • ports: ホストネットワーク上でOTLP gRPCおよびHTTPポートを公開し、アプリケーションまたはeBPFエージェントが同じノード上のコレクターに直接データを送信できるようにします。

OpenTelemetry Collector設定 (エージェント)

この設定は、エージェントコレクターがデータを受信、処理し、ゲートウェイコレクターに転送する方法を示しています。基本的なeBPFレシーバー(特定のeBPFレシーバーは異なるため概念的)、ホストメトリクス、ログ収集が含まれます。

# otel-collector-agent-config.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: otel-collector-agent-config
  namespace: observability
data:
  otel-collector-config.yaml: |
    receivers:
      # OTLP receiver for applications and eBPF agents exporting directly
      otlp:
        protocols:
          grpc:
          http:

      # Host metrics receiver
      hostmetrics:
        collection_interval: 10s
        scrapers:
          cpu:
          memory:
          disk:
          filesystem:
          network:
          load:
          paging:
          processes:

      # Filelog receiver for host logs (e.g., systemd, kernel logs)
      # Requires hostPath mount for /var/log
      filelog:
        include:
          - /var/log/*.log
          - /var/log/*/*.log
        start_at: beginning
        poll_interval: 1s
        operators:
          - type: json_parser
            output: json
          - type: regex_parser
            regex: '^(?P<time>\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}\.\d{3}Z)\s(?P<level>\w+)\s(?P<message>.*)$'
            output: json
            timestamp:
              parse_from: time
              layout: '%Y-%m-%dT%H:%M:%S.%LZ'
            severity:
              parse_from: level
              preset: syslog
          - type: move
            from: json.message
            to: body
          - type: remove
            field: json

      # Kubernetes Container Logs (Docker/CRI-O)
      # Requires hostPath mount for /var/lib/docker/containers
      # This is a basic example; for production, consider a dedicated log agent or more robust K8s integration
      filelog/container:
        include:
          - /var/lib/docker/containers/*/*-json.log
        start_at: beginning
        poll_interval: 1s
        operators:
          - type: json_parser
            output: json
          - type: move
            from: json.log
            to: body
          - type: move
            from: json.stream
            to: attributes.log.stream
          - type: move
            from: json.time
            to: attributes.log.time
          - type: add
            field: attributes.log.file.path
            value: "{{ .file.path }}"
          - type: add
            field: attributes.container.id
            value: "{{ .file.name }}" # Extract container ID from filename
          - type: add
            field: attributes.k8s.node.name
            value: "${KUBERNETES_NODE_NAME}" # Injected via env var
          - type: regex_parser
            regex: '^/var/lib/docker/containers/(?P<container_id>[a-f0-9]{64})/(?P<container_id_short>[a-f0-9]{12})-json.log$'
            parse_from: attributes.log.file.path
            output: attributes
          - type: remove
            field: json

    processors:
      # Batching for efficiency
      batch:
        send_batch_size: 1024
        timeout: 5s

      # Resource detection for Kubernetes metadata
      resourcedetection:
        detectors: ["system", "env", "kubernetes"]
        timeout: 2s
        kubernetes:
          pod_association:
            - from: "ip"
            - from: "cgroup"
          exclude:
            pods:
              - name: "otel-collector-agent" # Exclude self-instrumentation

      # Memory limiter to prevent OOMs
      memory_limiter:
        check_interval: 1s
        limit_mib: 256
        spike_limit_mib: 64

      # Tail-based sampling (example, typically done at gateway)
      # For agent, head-based is more common if sampling is needed here
      # This example is for demonstration, usually agents forward all data.
      # tail_sampling:
      #   decision_wait: 10s
      #   num_traces: 100000
      #   expected_new_traces_per_sec: 100
      #   policies:
      #     [
      #       {
      #         name: "error-policy",
      #         type: "status_code",
      #         status_code: { status_codes: ["ERROR"] }
      #       },
      #       {
      #         name: "latency-policy",
      #         type: "latency",
      #         latency: { threshold_ms: 500 }
      #       }
      #     ]

    exporters:
      # Export to a central OpenTelemetry Collector Gateway
      otlp:
        endpoint: "otel-collector-gateway.observability.svc.cluster.local:4317" # Internal K8s service
        tls:
          insecure: true # Use mTLS in production

      # Optional: Prometheus exporter for collector's own metrics
      prometheus:
        endpoint: "0.0.0.0:8888"

    service:
      telemetry:
        metrics:
          address: 0.0.0.0:8888
      pipelines:
        traces:
          receivers: [otlp]
          processors: [resourcedetection, batch, memory_limiter] # Add tail_sampling if needed
          exporters: [otlp]
        metrics:
          receivers: [otlp, hostmetrics]
          processors: [resourcedetection, batch, memory_limiter]
          exporters: [otlp]
        logs:
          receivers: [otlp, filelog, filelog/container]
          processors: [resourcedetection, batch, memory_limiter]
          exporters: [otlp]

主要な設定要素:

  • receivers:
    • otlp: アプリケーションまたはeBPFエージェントからのOTLPデータ(トレース、メトリクス、ログ)を受け入れます。
    • hostmetrics: ホストからCPU、メモリ、ディスク、ネットワークなどをスクレイピングします。
    • filelog: 指定されたホストパスからログを収集し、構造化されたログにパースします。これはゼロコードのログ集約にとって非常に重要です。
  • processors:
    • batch: 効率的なエクスポートのためにデータをバッチ処理します。
    • resourcedetection: Kubernetesメタデータ(Pod名、名前空間、ノード名など)でテレメトリをエンリッチし、コンテキストにとって不可欠です。
    • memory_limiter: コレクターが過剰なメモリを消費するのを防ぎ、DaemonSetにとって重要です。
  • exporters:
    • otlp: 処理されたすべてのテレメトリデータを中央のOpenTelemetry Collector Gatewayに転送します。
    • prometheus: コレクターの内部メトリクスをPrometheus形式で公開します。
  • service.pipelines: トレース、メトリクス、ログのレシーバー、プロセッサー、エクスポーターを介したデータの流れを定義します。

OpenTelemetry Collector Gatewayのデプロイと設定

ゲートウェイコレクターは、すべてのエージェントコレクターからのデータを集約し、サンプリングなどの高度な処理を適用し、最終的なバックエンドにエクスポートします。

# otel-collector-gateway-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: otel-collector-gateway
  namespace: observability
  labels:
    app: otel-collector-gateway
spec:
  replicas: 2 # Scale as needed
  selector:
    matchLabels:
      app: otel-collector-gateway
  template:
    metadata:
      labels:
        app: otel-collector-gateway
    spec:
      containers:
        - name: otel-collector
          image: otel/opentelemetry-collector-contrib:0.90.1
          command: ["/otelcol", "--config=/conf/otel-collector-config.yaml"]
          volumeMounts:
            - name: otel-collector-config
              mountPath: /conf
          ports:
            - name: otlp-grpc
              containerPort: 4317
            - name: otlp-http
              containerPort: 4318
            - name: prometheus
              containerPort: 8888
            - name: health-check
              containerPort: 13133
          livenessProbe:
            httpGet:
              path: /health
              port: 13133
            initialDelaySeconds: 5
            periodSeconds: 10
          readinessProbe:
            httpGet:
              path: /health
              port: 13133
            initialDelaySeconds: 5
            periodSeconds: 10
          resources:
            limits:
              cpu: 1000m
              memory: 2Gi
            requests:
              cpu: 200m
              memory: 512Mi
      volumes:
        - name: otel-collector-config
          configMap:
            name: otel-collector-gateway-config
---
# otel-collector-gateway-service.yaml
apiVersion: v1
kind: Service
metadata:
  name: otel-collector-gateway
  namespace: observability
spec:
  selector:
    app: otel-collector-gateway
  ports:
    - name: otlp-grpc
      protocol: TCP
      port: 4317
      targetPort: 4317
    - name: otlp-http
      protocol: TCP
      port: 4318
      targetPort: 4318
    - name: prometheus
      protocol: TCP
      port: 8888
      targetPort: 8888
Advertisement

OpenTelemetry Collector設定 (ゲートウェイ)

この設定には、テールベースサンプリングなどの高度なプロセッサーと、様々なバックエンドへのエクスポートが含まれます。

# otel-collector-gateway-config.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: otel-collector-gateway-config
  namespace: observability
data:
  otel-collector-config.yaml: |
    receivers:
      otlp:
        protocols:
          grpc:
          http:

    processors:
      batch:
        send_batch_size: 1024
        timeout: 5s

      memory_limiter:
        check_interval: 1s
        limit_mib: 1500
        spike_limit_mib: 500

      # Tail-based sampling for traces
      # This is crucial for reducing trace volume while retaining important traces.
      tail_sampling:
        decision_wait: 10s # How long to wait for all spans of a trace
        num_traces: 100000 # Max number of traces to keep in memory for sampling
        expected_new_traces_per_sec: 1000 # Estimate of new traces per second
        policies:
          - name: "always-sample"
            type: "always_on" # Always sample 100% of traces (for dev/low traffic)
          - name: "error-policy"
            type: "status_code"
            status_code: { status_codes: ["ERROR", "UNSET"] } # Sample traces with errors
          - name: "latency-policy"
            type: "latency"
            latency: { threshold_ms: 500 } # Sample traces exceeding 500ms
          - name: "probabilistic-policy"
            type: "probabilistic"
            probabilistic: { sampling_percentage: 10 } # Sample 10% of remaining traces

      # Attributes processor for general attribute manipulation
      attributes:
        actions:
          - key: "service.namespace"
            action: "insert"
            value: "default" # Default namespace if not present
          - key: "k8s.cluster.name"
            action: "insert"
            value: "my-prod-cluster" # Add cluster name
          - key: "host.ip"
            action: "delete" # Remove sensitive host IP if not needed downstream

      # Resource processor for adding/modifying resource attributes
      resource:
        attributes:
          - key: "cloud.provider"
            value: "aws"
            action: "insert"
          - key: "cloud.region"
            value: "us-east-1"
            action: "insert"

    exporters:
      # Export traces to Jaeger
      jaeger:
        endpoint: "jaeger-collector.observability.svc.cluster.local:14250" # gRPC
        tls:
          insecure: true

      # Export metrics to Prometheus remote write (e.g., Mimir, Thanos)
      prometheusremotewrite:
        endpoint: "http://prometheus-mimir-gateway.observability.svc.cluster.local:9009/api/v1/push"
        headers:
          X-Scope-OrgID: "locionic"
        # auth:
        #   oauth2:
        #     client_id: "..."
        #     client_secret: "..."
        #     token_url: "..."

      # Export logs to Loki
      loki:
        endpoint: "http://loki.observability.svc.cluster.local:3100/loki/api/v1/push"
        tls:
          insecure: true
        # auth:
        #   basic:
        #     username: "..."
        #     password: "..."
        labels:
          attributes:
            - host.name
            - k8s.namespace.name
            - k8s.pod.name
            - service.name
          resource:
            - k8s.node.name
            - k8s.cluster.name

      # Optional: Prometheus exporter for collector's own metrics
      prometheus:
        endpoint: "0.0.0.0:8888"

    service:
      telemetry:
        metrics:
          address: 0.0.0.0:8888
      pipelines:
        traces:
          receivers: [otlp]
          processors: [memory_limiter, batch, tail_sampling, attributes, resource]
          exporters: [jaeger]
        metrics:
          receivers: [otlp]
          processors: [memory_limiter, batch, attributes, resource]
          exporters: [prometheusremotewrite]
        logs:
          receivers: [otlp]
          processors: [memory_limiter, batch, attributes, resource]
          exporters: [loki]

サンプリング戦略: ヘッドベース vs. テールベース

機能ヘッドベースサンプリングテールベースサンプリング
決定ポイントトレースの開始時(最初のスパン)トレースの終了時(すべてのスパンが収集された後)
必要なデータ最初のスパンのコンテキストのみトレースに属するすべてのスパン
実装アプリケーションSDK、OTel Collector AgentOTel Collector Gateway
利点低オーバーヘッド、シンプル、早期にネットワークトラフィックを削減コンテキスト認識(エラー、レイテンシ、特定の属性)
よりインテリジェントなサンプリング決定
欠点コンテキスト認識なし、重要なトレースをドロップする可能性ありコレクターでのリソース使用量(メモリ、CPU)が高い
トレースのすべてのスパンがコレクターに到達する必要がある
決定が下される前にレイテンシが発生する
ユースケース大容量、基本的なサンプリング、初期フィルタリング本番環境、複雑なサンプリングロジック
重要なトレース(エラー、低速リクエスト)に焦点を当てる

推奨: 非常に大容量で重要でないトレース(例: 1%の確率的サンプリング)には、アプリケーションまたはエージェントレベルでヘッドベースサンプリングを使用します。エラー、高レイテンシ、特定のユーザーIDなどの重要なトレースのインテリジェントでコンテキスト認識型のサンプリングには、ゲートウェイコレクターでテールベースサンプリングを実装します。

Prometheusメトリクスエクスポート

OpenTelemetry Collectorは、Prometheusスクレイパーターゲットとして機能したり、Prometheusリモートライトを介してメトリクスをエクスポートしたりできます。

1. Prometheusスクレイプターゲットとしてのコレクター(自身のメトリクス用): エージェントとゲートウェイの両方の設定におけるprometheusエクスポーターは、コレクターの内部メトリクス(例: otelcol_receiver_accepted_spans_total)を公開します。Prometheusはこれらのエンドポイントをスクレイピングできます。

# Example Prometheus scrape config for OTel Collector Gateway
# Add this to your Prometheus server's scrape_configs
- job_name: 'otel-collector-gateway'
  kubernetes_sd_configs:
    - role: endpoints
      namespaces:
        names: ['observability']
  relabel_configs:
    - source_labels: [__meta_kubernetes_service_name, __meta_kubernetes_endpoint_port_name]
      action: keep
      regex: otel-collector-gateway;prometheus
    - source_labels: [__meta_kubernetes_pod_name]
      target_label: kubernetes_pod_name
      action: replace
    - source_labels: [__meta_kubernetes_namespace]
      target_label: kubernetes_namespace
      action: replace

2. Prometheusリモートライトを介したアプリケーションメトリクスのエクスポート: ゲートウェイ設定のprometheusremotewriteエクスポーターは、処理されたメトリクスをPrometheus互換のリモートストレージ(例: Mimir、Thanos、Cortex)に送信します。これにより、メトリクスストレージが一元化され、長期保存と高可用性が可能になります。

本番環境での落とし穴とトラブルシューティング

  1. eBPFパーミッション (privileged: true, hostNetwork: true):

    • 落とし穴: privileged: trueとhostNetwork: trueがないと、eBPFエージェントはプログラムのロードやネットワークトラフィックのキャプチャに失敗します。これは一般的な設定ミスです。
    • 修正: DaemonSetのsecurityContextにprivileged: trueが含まれていること、およびPodレベルでhostNetwork: trueが設定されていることを確認します。また、ServiceAccountに必要なClusterRoleパーミッション(例: nodes/proxy)があることを確認します。
    • トラブルシューティング: Podログでpermission denied、operation not permitted、またはfailed to load eBPF programエラーを確認します。kubectl describe pod <pod-name>を使用してセキュリティコンテキストとネットワーク設定を検証します。
  2. メモリ制限とOOMKilledコレクター:

    • 落とし穴: OpenTelemetry Collector、特にテールベースサンプリングを使用するゲートウェイは、かなりのメモリを消費し、OOMKilledにつながる可能性があります。
    • 修正:
      • コレクター設定にmemory_limiterプロセッサーを実装します。
      • Kubernetesマニフェストで適切なresources.limits.memoryを設定します。
      • トレース量と利用可能なメモリに合わせてtail_samplingパラメーター(num_traces、expected_new_traces_per_sec)を調整します。
      • 単一のインスタンスで負荷を処理できない場合は、ゲートウェイコレクターをスケールアウトします。
    • トラブルシューティング: kubectl describe pod <pod-name>でOOMKilledステータスを確認します。コレクター自身のPrometheusメトリクス(otelcol_processor_memory_limiter_memory_usage_bytes)を介してコレクターのメモリ使用量を監視します。
  3. ネットワーク接続の問題(エージェントからゲートウェイ、ゲートウェイからバックエンド):

    • 落とし穴: サービス名、ポート、またはファイアウォールルールの誤りにより、データフローが妨げられる可能性があります。
    • 修正:
      • エクスポーター(例: otel-collector-gateway.observability.svc.cluster.local:4317)のendpoint値を確認します。
      • Kubernetesサービスが正しく定義され、コレクターPodをターゲットにしていることを確認します。
      • ネットワークポリシーが使用されている場合は、それを確認します。
    • トラブルシューティング:
      • コレクターログでconnection refused、unavailable、またはtimeoutエラーを確認します。
      • コレクターPod内からkubectl exec -it <collector-pod> -- curl <target-endpoint>(イメージにcurlが利用可能な場合)またはnc -vz <target-host> <target-port>を使用して接続をテストします。
      • Pod内からのDNS解決を検証します。
  4. テレメトリデータの欠落または不完全:

    • 落とし穴: レシーバー、プロセッサー、エクスポーターの設定ミス、または積極的なサンプリングにより、データがドロップされる可能性があります。
    • 修正:
      • service.pipelinesを確認し、目的のすべてのレシーバー、プロセッサー、エクスポーターが正しくリンクされていることを確認します。
      • レシーバー(例: filelog)のinclude/excludeルールを確認します。
      • ドロップされるデータが多すぎる場合は、サンプリングポリシーを調整します。
      • 必要なメタデータでデータをエンリッチするためにresourcedetectionが正しく設定されていることを確認します。
    • トラブルシューティング:
      • コレクターの内部メトリクス(otelcol_receiver_accepted_spans_total、otelcol_exporter_sent_spans_total、otelcol_processor_batch_batch_send_size_sum)を監視して、データがどこでドロップされている可能性があるかを特定します。
      • コレクターのデバッグロギング(--config=/conf/otel-collector-config.yaml --set service.telemetry.logs.level=debug)を有効にして、より詳細な出力を取得します。
  5. eBPFエージェントの互換性とカーネルバージョン:

    • 落とし穴: eBPFプログラムはカーネルバージョンに依存します。あるカーネルバージョン用にコンパイルされたeBPFエージェントは、別のバージョンでは動作しないか、特定のカーネルヘッダーを必要とする場合があります。
    • 修正:
      • CO-RE(Compile Once – Run Everywhere)をサポートするeBPFエージェント、または一般的なカーネルバージョン用にプリコンパイルされたバイナリが配布されているものを使用します。
      • Kubernetesノードが互換性のあるカーネルバージョンを実行していることを確認します。
      • 一部のeBPFソリューションでは、特定のカーネルモジュールまたは機能が有効になっている必要があります。
    • トラブルシューティング: eBPFエージェントのログでBPF program load failed、invalid argument、またはkernel version mismatchのようなエラーを探します。特定のeBPFエージェントのドキュメントを参照してください。

よくある質問

  1. Q: eBPFを自動計装に使用した場合のパフォーマンスオーバーヘッドはどれくらいですか? A: eBPFベースの計装は通常、CPUとメモリで一桁台のパーセンテージ範囲という非常に低いオーバーヘッドです。これは、eBPFプログラムがカーネル内で直接実行され、コンテキストスイッチやユーザー空間のオーバーヘッドを回避するためです。ただし、eBPFプローブの複雑さと頻度、および収集されるデータの量が増加すると、オーバーヘッドも増加する可能性があります。コレクターでの積極的なデータ処理と高カーディナリティの属性も、リソース使用量の増加に寄与する可能性があります。

  2. Q: 従来のOpenTelemetry SDK計装とeBPF自動計装を併用できますか? A: はい、もちろんです。これは一般的で推奨されるパターンです。eBPFは、ネットワークインタラクション、システムコール、プロセス実行に対する「ゼロコード」の可視性を提供し、アプリケーションコードが計装されていないギャップを埋めます。OpenTelemetry SDKは、アプリケーションのビジネスロジックから直接、よりリッチでセマンティックなコンテキストを提供します。OpenTelemetry Collectorは、両方のソースからのデータをマージおよび関連付けし、より完全な全体像を提供できます。

  3. Q: eBPFエージェントがKubernetes内のすべてのアプリケーションからデータを収集していることを確認するにはどうすればよいですか? A: eBPFエージェント(またはeBPF対応コレクター)をDaemonSetとしてデプロイすることで、すべてのノードで実行されることが保証されます。hostNetwork: trueとprivileged: trueを使用すると、どのPodまたは名前空間に属しているかに関係なく、そのノード上のすべてのプロセスとネットワークトラフィックを監視するために必要なアクセス権が得られます。コレクターのリソース検出プロセッサーは、このデータをKubernetesメタデータでエンリッチし、特定のPodやサービスにリンクし直します。

  4. Q: バックエンドを圧倒しないように、大量のトレースデータを処理する最善の方法は何ですか? A: 多段階サンプリング戦略を実装します。

    • ヘッドベースサンプリング(確率的)をアプリケーションまたはエージェントレベルで実行し、重要でないトレースの初期削減を行います。
    • OpenTelemetry Collector Gatewayでテールベースサンプリングを実行します。これにより、エラー、レイテンシ、特定のビジネスロジックなどのトレース属性に基づいてインテリジェントなサンプリングが可能になり、重要なトレースが常にキャプチャされることが保証されます。
    • コレクターでのバッチ処理とメモリ制限も、スループットとリソース使用量を管理するために不可欠です。
  5. Q: OpenTelemetry Collectorはアプリケーションの変更なしにログをどのように処理しますか? A: OpenTelemetry Collectorのfilelogレシーバーは、ホストパス(例: システムログ用の/var/log、コンテナログ用の/var/lib/docker/containers)から直接ログをスクレイピングするように設定できます。これらのホストパスをコレクターDaemonSetにマウントすることで、コレクターはこれらのログを読み取り、パースし、処理し、LokiやElasticsearchなどのログバックエンドにエクスポートできます。これらすべては、アプリケーションコード自体を変更することなく行われます。これにより、トレースやメトリクスと並行して、統合されたログ収集パイプラインが提供されます。

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement