•12 min read

The OpenTelemetry LGTM Stack: Loki, Grafana, Tempo & Mimir Production Guide

The OpenTelemetry LGTM Stack: Loki, Grafana, Tempo & Mimir Production Guide

Observability at scale demands a unified approach to logs, metrics, and traces. The Grafana LGTM stack—Loki, Grafana, Tempo, and Mimir—provides a robust, cloud-native solution. This guide details its production deployment, focusing on OpenTelemetry (OTel) integration, cross-signal correlation, and cost optimization.

Audio Briefing
0:00 / 0:00

Architecture Overview

The LGTM stack, augmented by OpenTelemetry Collectors, forms a cohesive observability platform.

Key Components:

  • OpenTelemetry Collector: The universal agent for ingesting, processing, and exporting telemetry data. It acts as the central nervous system, standardizing data before forwarding to the respective LGTM components.
  • Mimir (Metrics): Horizontally scalable, long-term storage for Prometheus metrics. It offers high availability, multi-tenancy, and efficient querying.
  • Loki (Logs): A log aggregation system designed for cost-effectiveness and scalability. It indexes metadata (labels) rather than full log content, making it efficient for querying.
  • Tempo (Traces): A high-volume, low-cost distributed tracing backend. It stores traces and allows for efficient retrieval based on trace IDs.
  • Grafana: The visualization layer, providing dashboards, alerts, and unified exploration across all telemetry signals.
  • Object Storage (S3/GCS): The primary long-term storage for Mimir, Loki, and Tempo, leveraging cloud-native object stores for durability and cost efficiency.
Advertisement

OpenTelemetry Collector Pipelines

The OTel Collector is critical for standardizing telemetry data. We'll configure it to receive OTLP, process, and export to Mimir, Loki, and Tempo.

Collector Configuration (otel-collector-config.yaml)

receivers:
  otlp:
    protocols:
      grpc:
      http:

processors:
  batch:
    send_batch_size: 1024
    timeout: 10s
  resource/add_service_name: # Ensure service.name is present for correlation
    attributes:
      - key: service.name
        value: unknown_service
        action: upsert
  attributes/add_k8s_labels: # Example: Add Kubernetes labels as resource attributes
    actions:
      - key: k8s.pod.name
        from_attribute: k8s.pod.name
        action: upsert
      - key: k8s.namespace.name
        from_attribute: k8s.namespace.name
        action: upsert
  transform/logs: # Transform OTel logs to Loki-compatible format
    log_statements:
      - context: log
        statements:
          - set(attributes["loki.tenant.id"], "default") # Multi-tenancy example
          - set(attributes["loki.resource.labels"], Concat([resource.attributes["service.name"], resource.attributes["k8s.namespace.name"]], ","))
          # Ensure trace_id and span_id are available as attributes for Loki
          - set(attributes["trace_id"], SpanIDToHex(trace_id))
          - set(attributes["span_id"], SpanIDToHex(span_id))
  transform/metrics: # Example: Add tenant ID to metrics
    metric_statements:
      - context: metric
        statements:
          - set(attributes["tenant_id"], "default")
  resource/add_tenant_id: # Add tenant ID to resource attributes for traces
    attributes:
      - key: tenant_id
        value: default
        action: upsert

exporters:
  otlp/mimir:
    endpoint: mimir-gateway.observability.svc.cluster.local:9090 # Mimir OTLP endpoint
    tls:
      insecure: true # Use proper TLS in production
  otlp/loki:
    endpoint: loki-gateway.observability.svc.cluster.local:3100 # Loki OTLP endpoint
    tls:
      insecure: true
  otlp/tempo:
    endpoint: tempo-gateway.observability.svc.cluster.local:4317 # Tempo OTLP endpoint
    tls:
      insecure: true
  logging: # For debugging
    verbosity: detailed

service:
  pipelines:
    metrics:
      receivers: [otlp]
      processors: [resource/add_service_name, attributes/add_k8s_labels, transform/metrics, batch]
      exporters: [otlp/mimir, logging]
    logs:
      receivers: [otlp]
      processors: [resource/add_service_name, attributes/add_k8s_labels, transform/logs, batch]
      exporters: [otlp/loki, logging]
    traces:
      receivers: [otlp]
      processors: [resource/add_service_name, resource/add_tenant_id, batch]
      exporters: [otlp/tempo, logging]

Explanation:

  • receivers.otlp: Configures the collector to accept OTLP data over gRPC and HTTP.
  • processors.batch: Batches telemetry data for efficient export.
  • processors.resource/add_service_name: Ensures service.name is always present, crucial for correlation.
  • processors.attributes/add_k8s_labels: Demonstrates enriching telemetry with Kubernetes metadata, useful for filtering and grouping.
  • processors.transform/logs: This is critical for Loki. It adds loki.tenant.id and loki.resource.labels (which Loki uses as indexed labels) and extracts trace_id and span_id into log attributes, enabling trace-to-log correlation.
  • exporters.otlp/mimir, otlp/loki, otlp/tempo: Direct OTLP export to the respective LGTM components. Ensure these endpoints are resolvable within your cluster.
  • service.pipelines: Defines how data flows through receivers, processors, and exporters for each signal type (metrics, logs, traces).

Cross-Signal Correlation: Trace-to-Logs & Logs-to-Traces

Unified observability hinges on the ability to navigate between signals. Grafana facilitates this with data link configurations.

1. Trace-to-Logs Drilldown

When viewing a trace in Tempo, you want to jump to logs associated with a specific span. This requires the trace_id and span_id to be present as attributes in your logs. The transform/logs processor in the OTel Collector handles this.

Grafana Data Link Configuration (Tempo Datasource):

{
  "name": "Logs for Span",
  "url": "/explore?orgId=1&left=[\"now-1h\",\"now\",\"Loki\",{\"expr\":\"{service_name=\\\"$${service.name}\\\", trace_id=\\\"$${traceID}\\\", span_id=\\\"$${spanID}\\\"}\",\"refId\":\"A\",\"queryType\":\"range\"},{\"ui\":true}]",
  "targetBlank": true
}

Explanation:

  • $${service.name}: Inferred from the trace span.
  • $${traceID}: The trace ID of the current span.
  • $${spanID}: The span ID of the current span.
  • The expr uses Loki's LogQL to filter logs by service_name, trace_id, and span_id. These labels must be present in your Loki logs (as configured by the OTel Collector).

2. Logs-to-Traces Drilldown

From a log line in Loki, you want to jump to the corresponding trace in Tempo. This requires the trace_id to be present as an attribute in the log.

Grafana Data Link Configuration (Loki Datasource):

{
  "name": "Trace for Log",
  "url": "/explore?orgId=1&left=[\"now-1h\",\"now\",\"Tempo\",{\"query\":\"$${__line.trace_id}\",\"queryType\":\"search\",\"refId\":\"A\"},{\"ui\":true}]",
  "targetBlank": true
}

Explanation:

  • $${__line.trace_id}: This special Grafana variable extracts the trace_id attribute from the current log line. This assumes your OTel Collector has added trace_id as an attribute to logs.

High-Signal SLI/SLO Dashboards

Leveraging Mimir, we can build robust SLI/SLO dashboards. This example focuses on a request latency SLI.

1. Define SLI/SLO

  • SLI: Percentage of requests served with p99 latency under 500ms.
  • SLO: 99.9% of requests must meet the SLI over a 7-day rolling window.

2. Mimir Recording Rules for SLI

Define recording rules in Mimir to pre-aggregate SLI data. This reduces query load on the raw metrics and improves dashboard performance.

# rules.yaml for Mimir
groups:
  - name: service-sli
    rules:
      - record: service_latency_sli_total
        expr: |
          sum by (service_name, http_route) (
            rate(http_server_request_duration_bucket{le="0.5"}[5m])
          )
      - record: service_latency_slo_error_budget
        expr: |
          1 - (
            sum by (service_name, http_route) (
              rate(http_server_request_duration_bucket{le="0.5"}[5m])
            )
            /
            sum by (service_name, http_route) (
              rate(http_server_request_duration_count[5m])
            )
          )

Explanation:

  • http_server_request_duration_bucket: Assumes your application exports OpenTelemetry HTTP server request duration histograms.
  • service_latency_sli_total: Counts requests within the latency target.
  • service_latency_slo_error_budget: Calculates the error budget consumed.

3. Grafana Dashboard Panels

  • Current SLI Status:
    (
      sum by (service_name, http_route) (service_latency_sli_total)
      /
      sum by (service_name, http_route) (rate(http_server_request_duration_count[5m]))
    ) * 100
    
  • Error Budget Burn Rate (7-day window):
    sum by (service_name, http_route) (
      increase(service_latency_slo_error_budget[7d])
    )
    
  • Remaining Error Budget:
    100 - (
      sum by (service_name, http_route) (
        increase(service_latency_slo_error_budget[7d])
      )
    )
    
Advertisement

Object Storage Cost Optimization (S3/GCS)

Mimir, Loki, and Tempo heavily rely on object storage. Optimizing this is crucial for cost control.

1. Data Retention Policies

Configure retention periods based on data criticality and compliance.

  • Mimir: Short-term (e.g., 30 days) for high-resolution metrics, longer for aggregated metrics. Mimir's ingester and compactor manage this.
  • Loki: Typically longer (e.g., 90-180 days) as logs are often needed for forensics. Loki's table_manager and compactor handle retention.
  • Tempo: Often shorter (e.g., 7-30 days) due to high volume, unless specific compliance requires longer. Tempo's compactor manages retention.

Example Loki Retention Configuration (loki.yaml):

table_manager:
  retention_period: 90d # 90 days for index and chunks
compactor:
  retention_enabled: true
  retention_period: 90d

2. Storage Tiers & Lifecycle Policies

Leverage cloud provider lifecycle policies to transition older data to cheaper storage tiers (e.g., S3 Standard-IA, Glacier, GCS Nearline, Coldline).

Example S3 Lifecycle Policy (Terraform aws_s3_bucket_lifecycle_configuration):

resource "aws_s3_bucket_lifecycle_configuration" "loki_bucket_lifecycle" {
  bucket = aws_s3_bucket.loki_chunks.id

  rule {
    id     = "loki-retention"
    status = "Enabled"

    transition {
      days          = 30
      storage_class = "STANDARD_IA" # Infrequent Access
    }

    transition {
      days          = 60
      storage_class = "GLACIER" # Archival
    }

    expiration {
      days = 90 # Final deletion
    }
  }
}

3. Data Compression

All LGTM components use compression (Snappy, LZ4, Zstd) for data stored in object storage. Ensure your configurations are leveraging efficient codecs. This is typically default, but worth verifying.

4. Sharding & Indexing Strategies

  • Loki: Optimize max_chunk_age and chunk_target_size to balance write performance and query efficiency. A smaller max_chunk_age means more frequent flushing, potentially more small objects, but faster index updates.
  • Tempo: Consider trace ID sharding strategies if you have extremely high ingest rates to distribute load across ingesters.

Production Gotchas & Troubleshooting

1. OTel Collector Backpressure / OOMKilled

Symptom: OTel Collector pods are OOMKilled or show high CPU/memory usage, dropping telemetry data. Cause: Ingest rate exceeds processing/export capacity, or batch processor settings are too aggressive for available memory. Fix:

  • Increase resources: Allocate more CPU/memory to OTel Collector pods.
  • Tune batch processor: Increase send_batch_size and timeout to allow larger batches, but monitor memory.
  • Add memory_limiter processor: Configure a soft and hard limit to prevent OOMs and gracefully drop data under extreme pressure.
    processors:
      memory_limiter:
        limit_mib: 256
        spike_limit_mib: 64
        check_interval: 1s
    
  • Scale out: Deploy more OTel Collector instances behind a load balancer.

2. Loki Query Performance Degradation

Symptom: LogQL queries are slow, especially those involving regex or large time ranges. Cause:

  • Too many unique label combinations (high cardinality).
  • Queries scanning too many log streams.
  • Inefficient LogQL expressions.
  • Insufficient Loki query-frontend/querier resources. Fix:
  • Review labels: Minimize the number of unique labels. Avoid adding high-cardinality attributes (e.g., request ID, full URL path) as Loki labels. Use them as log content attributes instead.
  • Optimize LogQL: Start queries with highly selective labels. Use line_format to extract data from log content instead of indexing it.
  • Increase Loki resources: Scale query-frontend and querier components.
  • Index optimization: Ensure max_chunk_age and chunk_target_size are balanced.

3. Tempo Trace Ingestion Failures

Symptom: Traces are missing or incomplete in Tempo. Cause:

  • OTel Collector not forwarding traces correctly.
  • Tempo ingesters are overloaded or unhealthy.
  • Incorrect service instrumentation (e.g., missing trace context propagation). Fix:
  • Check OTel Collector logs: Look for errors exporting to Tempo.
  • Monitor Tempo ingesters: Check their health and resource utilization. Scale if necessary.
  • Verify instrumentation: Ensure all services are correctly propagating trace context (e.g., W3C Trace Context headers). Use curl with traceparent headers to test.
  • Increase Tempo ingester max_block_bytes: If traces are very large, this might be a bottleneck.

4. Mimir High Cardinality Issues

Symptom: Mimir ingesters/compactors are struggling, high storage costs, slow metric queries. Cause: Too many unique label combinations on metrics. Fix:

  • Metric Relabeling: Use OTel Collector's metricstransform processor or Prometheus relabel_configs to drop or rename high-cardinality labels before ingestion.
    processors:
      metricstransform:
        transforms:
          - include: "http_request_duration_seconds_bucket"
            action: "delete_label"
            label: "request_id" # Example: delete high-cardinality label
    
  • Review instrumentation: Educate developers on avoiding high-cardinality labels.
  • Mimir limits: Configure Mimir's max_series_per_user and max_label_names_per_series to prevent abuse.

Frequently Asked Questions

Q1: How do I handle multi-tenancy with the LGTM stack?

A1: Mimir, Loki, and Tempo are designed for multi-tenancy.

  • Mimir: Use the X-Scope-OrgID HTTP header (or OTLP attribute tenant_id) to separate tenants. Each tenant gets isolated data.
  • Loki: Similar to Mimir, use X-Scope-OrgID or the loki.tenant.id OTLP log attribute.
  • Tempo: Use X-Scope-OrgID or the tenant_id OTLP resource attribute.
  • Grafana: Configure separate data sources for each tenant, or use template variables for X-Scope-OrgID in dashboards.

A2: Helm charts are the standard. Grafana Labs provides official charts for Loki, Mimir, and Tempo. For the OpenTelemetry Collector, use the opentelemetry-collector Helm chart. Ensure you configure persistent storage for stateful components (e.g., Mimir ingesters, Loki ingesters) and use a robust object storage backend.

Q3: How do I monitor the LGTM stack itself?

A3: The LGTM components expose Prometheus metrics. Deploy a dedicated Prometheus instance (or use Mimir itself) to scrape these metrics. Create Grafana dashboards to monitor their health, resource usage, ingest rates, query latencies, and error rates. Key metrics include loki_ingester_received_entries_total, mimir_ingester_received_samples_total, tempo_ingester_traces_received_total, and various _grpc_server_handled_total metrics.

Q4: Can I use existing Prometheus metrics with Mimir?

A4: Yes. Mimir is a long-term storage for Prometheus. You can configure your existing Prometheus servers to remote-write to Mimir, or use the OpenTelemetry Collector's Prometheus receiver and OTLP exporter to Mimir. The latter is generally preferred for a unified OTel approach.

Q5: What are the trade-offs between Loki's label-based indexing and traditional full-text log indexing?

A5:

FeatureLoki (Label-based)Traditional (Full-text)
CostLower storage, less compute for indexingHigher storage, more compute for indexing
Query SpeedFast for label-filtered queries, slower for textFast for full-text searches, slower for complex filters
ScalabilityHighly scalable horizontallyScalable, but index management can be complex
FlexibilityRequires pre-defined labels for efficient searchSearchable on any log content
Use CaseOperational debugging, known patternsSecurity analysis, ad-hoc exploration

Loki excels when you know what you're looking for (e.g., logs from service=X in namespace=Y). For arbitrary full-text searches across vast, unstructured logs, traditional systems might be more performant but at a significantly higher cost.

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement