The OpenTelemetry LGTM Stack: Loki, Grafana, Tempo & Mimir Production Guide

Table of Contents(26 sections)
Observability at scale demands a unified approach to logs, metrics, and traces. The Grafana LGTM stack—Loki, Grafana, Tempo, and Mimir—provides a robust, cloud-native solution. This guide details its production deployment, focusing on OpenTelemetry (OTel) integration, cross-signal correlation, and cost optimization.
Modern DevOps & eBPF Observability Track
Architecture Overview
The LGTM stack, augmented by OpenTelemetry Collectors, forms a cohesive observability platform.
Key Components:
- OpenTelemetry Collector: The universal agent for ingesting, processing, and exporting telemetry data. It acts as the central nervous system, standardizing data before forwarding to the respective LGTM components.
- Mimir (Metrics): Horizontally scalable, long-term storage for Prometheus metrics. It offers high availability, multi-tenancy, and efficient querying.
- Loki (Logs): A log aggregation system designed for cost-effectiveness and scalability. It indexes metadata (labels) rather than full log content, making it efficient for querying.
- Tempo (Traces): A high-volume, low-cost distributed tracing backend. It stores traces and allows for efficient retrieval based on trace IDs.
- Grafana: The visualization layer, providing dashboards, alerts, and unified exploration across all telemetry signals.
- Object Storage (S3/GCS): The primary long-term storage for Mimir, Loki, and Tempo, leveraging cloud-native object stores for durability and cost efficiency.
OpenTelemetry Collector Pipelines
The OTel Collector is critical for standardizing telemetry data. We'll configure it to receive OTLP, process, and export to Mimir, Loki, and Tempo.
Collector Configuration (otel-collector-config.yaml)
receivers:
otlp:
protocols:
grpc:
http:
processors:
batch:
send_batch_size: 1024
timeout: 10s
resource/add_service_name: # Ensure service.name is present for correlation
attributes:
- key: service.name
value: unknown_service
action: upsert
attributes/add_k8s_labels: # Example: Add Kubernetes labels as resource attributes
actions:
- key: k8s.pod.name
from_attribute: k8s.pod.name
action: upsert
- key: k8s.namespace.name
from_attribute: k8s.namespace.name
action: upsert
transform/logs: # Transform OTel logs to Loki-compatible format
log_statements:
- context: log
statements:
- set(attributes["loki.tenant.id"], "default") # Multi-tenancy example
- set(attributes["loki.resource.labels"], Concat([resource.attributes["service.name"], resource.attributes["k8s.namespace.name"]], ","))
# Ensure trace_id and span_id are available as attributes for Loki
- set(attributes["trace_id"], SpanIDToHex(trace_id))
- set(attributes["span_id"], SpanIDToHex(span_id))
transform/metrics: # Example: Add tenant ID to metrics
metric_statements:
- context: metric
statements:
- set(attributes["tenant_id"], "default")
resource/add_tenant_id: # Add tenant ID to resource attributes for traces
attributes:
- key: tenant_id
value: default
action: upsert
exporters:
otlp/mimir:
endpoint: mimir-gateway.observability.svc.cluster.local:9090 # Mimir OTLP endpoint
tls:
insecure: true # Use proper TLS in production
otlp/loki:
endpoint: loki-gateway.observability.svc.cluster.local:3100 # Loki OTLP endpoint
tls:
insecure: true
otlp/tempo:
endpoint: tempo-gateway.observability.svc.cluster.local:4317 # Tempo OTLP endpoint
tls:
insecure: true
logging: # For debugging
verbosity: detailed
service:
pipelines:
metrics:
receivers: [otlp]
processors: [resource/add_service_name, attributes/add_k8s_labels, transform/metrics, batch]
exporters: [otlp/mimir, logging]
logs:
receivers: [otlp]
processors: [resource/add_service_name, attributes/add_k8s_labels, transform/logs, batch]
exporters: [otlp/loki, logging]
traces:
receivers: [otlp]
processors: [resource/add_service_name, resource/add_tenant_id, batch]
exporters: [otlp/tempo, logging]
Explanation:
receivers.otlp: Configures the collector to accept OTLP data over gRPC and HTTP.processors.batch: Batches telemetry data for efficient export.processors.resource/add_service_name: Ensuresservice.nameis always present, crucial for correlation.processors.attributes/add_k8s_labels: Demonstrates enriching telemetry with Kubernetes metadata, useful for filtering and grouping.processors.transform/logs: This is critical for Loki. It addsloki.tenant.idandloki.resource.labels(which Loki uses as indexed labels) and extractstrace_idandspan_idinto log attributes, enabling trace-to-log correlation.exporters.otlp/mimir,otlp/loki,otlp/tempo: Direct OTLP export to the respective LGTM components. Ensure these endpoints are resolvable within your cluster.service.pipelines: Defines how data flows through receivers, processors, and exporters for each signal type (metrics, logs, traces).
Cross-Signal Correlation: Trace-to-Logs & Logs-to-Traces
Unified observability hinges on the ability to navigate between signals. Grafana facilitates this with data link configurations.
1. Trace-to-Logs Drilldown
When viewing a trace in Tempo, you want to jump to logs associated with a specific span. This requires the trace_id and span_id to be present as attributes in your logs. The transform/logs processor in the OTel Collector handles this.
Grafana Data Link Configuration (Tempo Datasource):
{
"name": "Logs for Span",
"url": "/explore?orgId=1&left=[\"now-1h\",\"now\",\"Loki\",{\"expr\":\"{service_name=\\\"$${service.name}\\\", trace_id=\\\"$${traceID}\\\", span_id=\\\"$${spanID}\\\"}\",\"refId\":\"A\",\"queryType\":\"range\"},{\"ui\":true}]",
"targetBlank": true
}
Explanation:
$${service.name}: Inferred from the trace span.$${traceID}: The trace ID of the current span.$${spanID}: The span ID of the current span.- The
expruses Loki's LogQL to filter logs byservice_name,trace_id, andspan_id. These labels must be present in your Loki logs (as configured by the OTel Collector).
2. Logs-to-Traces Drilldown
From a log line in Loki, you want to jump to the corresponding trace in Tempo. This requires the trace_id to be present as an attribute in the log.
Grafana Data Link Configuration (Loki Datasource):
{
"name": "Trace for Log",
"url": "/explore?orgId=1&left=[\"now-1h\",\"now\",\"Tempo\",{\"query\":\"$${__line.trace_id}\",\"queryType\":\"search\",\"refId\":\"A\"},{\"ui\":true}]",
"targetBlank": true
}
Explanation:
$${__line.trace_id}: This special Grafana variable extracts thetrace_idattribute from the current log line. This assumes your OTel Collector has addedtrace_idas an attribute to logs.
High-Signal SLI/SLO Dashboards
Leveraging Mimir, we can build robust SLI/SLO dashboards. This example focuses on a request latency SLI.
1. Define SLI/SLO
- SLI: Percentage of requests served with p99 latency under 500ms.
- SLO: 99.9% of requests must meet the SLI over a 7-day rolling window.
2. Mimir Recording Rules for SLI
Define recording rules in Mimir to pre-aggregate SLI data. This reduces query load on the raw metrics and improves dashboard performance.
# rules.yaml for Mimir
groups:
- name: service-sli
rules:
- record: service_latency_sli_total
expr: |
sum by (service_name, http_route) (
rate(http_server_request_duration_bucket{le="0.5"}[5m])
)
- record: service_latency_slo_error_budget
expr: |
1 - (
sum by (service_name, http_route) (
rate(http_server_request_duration_bucket{le="0.5"}[5m])
)
/
sum by (service_name, http_route) (
rate(http_server_request_duration_count[5m])
)
)
Explanation:
http_server_request_duration_bucket: Assumes your application exports OpenTelemetry HTTP server request duration histograms.service_latency_sli_total: Counts requests within the latency target.service_latency_slo_error_budget: Calculates the error budget consumed.
3. Grafana Dashboard Panels
- Current SLI Status:
promql
( sum by (service_name, http_route) (service_latency_sli_total) / sum by (service_name, http_route) (rate(http_server_request_duration_count[5m])) ) * 100 - Error Budget Burn Rate (7-day window):
promql
sum by (service_name, http_route) ( increase(service_latency_slo_error_budget[7d]) ) - Remaining Error Budget:
promql
100 - ( sum by (service_name, http_route) ( increase(service_latency_slo_error_budget[7d]) ) )
Object Storage Cost Optimization (S3/GCS)
Mimir, Loki, and Tempo heavily rely on object storage. Optimizing this is crucial for cost control.
1. Data Retention Policies
Configure retention periods based on data criticality and compliance.
- Mimir: Short-term (e.g., 30 days) for high-resolution metrics, longer for aggregated metrics. Mimir's ingester and compactor manage this.
- Loki: Typically longer (e.g., 90-180 days) as logs are often needed for forensics. Loki's
table_managerandcompactorhandle retention. - Tempo: Often shorter (e.g., 7-30 days) due to high volume, unless specific compliance requires longer. Tempo's
compactormanages retention.
Example Loki Retention Configuration (loki.yaml):
table_manager:
retention_period: 90d # 90 days for index and chunks
compactor:
retention_enabled: true
retention_period: 90d
2. Storage Tiers & Lifecycle Policies
Leverage cloud provider lifecycle policies to transition older data to cheaper storage tiers (e.g., S3 Standard-IA, Glacier, GCS Nearline, Coldline).
Example S3 Lifecycle Policy (Terraform aws_s3_bucket_lifecycle_configuration):
resource "aws_s3_bucket_lifecycle_configuration" "loki_bucket_lifecycle" {
bucket = aws_s3_bucket.loki_chunks.id
rule {
id = "loki-retention"
status = "Enabled"
transition {
days = 30
storage_class = "STANDARD_IA" # Infrequent Access
}
transition {
days = 60
storage_class = "GLACIER" # Archival
}
expiration {
days = 90 # Final deletion
}
}
}
3. Data Compression
All LGTM components use compression (Snappy, LZ4, Zstd) for data stored in object storage. Ensure your configurations are leveraging efficient codecs. This is typically default, but worth verifying.
4. Sharding & Indexing Strategies
- Loki: Optimize
max_chunk_ageandchunk_target_sizeto balance write performance and query efficiency. A smallermax_chunk_agemeans more frequent flushing, potentially more small objects, but faster index updates. - Tempo: Consider trace ID sharding strategies if you have extremely high ingest rates to distribute load across ingesters.
Production Gotchas & Troubleshooting
1. OTel Collector Backpressure / OOMKilled
Symptom: OTel Collector pods are OOMKilled or show high CPU/memory usage, dropping telemetry data.
Cause: Ingest rate exceeds processing/export capacity, or batch processor settings are too aggressive for available memory.
Fix:
- Increase resources: Allocate more CPU/memory to OTel Collector pods.
- Tune
batchprocessor: Increasesend_batch_sizeandtimeoutto allow larger batches, but monitor memory. - Add
memory_limiterprocessor: Configure a soft and hard limit to prevent OOMs and gracefully drop data under extreme pressure.yamlprocessors: memory_limiter: limit_mib: 256 spike_limit_mib: 64 check_interval: 1s - Scale out: Deploy more OTel Collector instances behind a load balancer.
2. Loki Query Performance Degradation
Symptom: LogQL queries are slow, especially those involving regex or large time ranges. Cause:
- Too many unique label combinations (high cardinality).
- Queries scanning too many log streams.
- Inefficient LogQL expressions.
- Insufficient Loki query-frontend/querier resources. Fix:
- Review labels: Minimize the number of unique labels. Avoid adding high-cardinality attributes (e.g., request ID, full URL path) as Loki labels. Use them as log content attributes instead.
- Optimize LogQL: Start queries with highly selective labels. Use
line_formatto extract data from log content instead of indexing it. - Increase Loki resources: Scale
query-frontendandqueriercomponents. - Index optimization: Ensure
max_chunk_ageandchunk_target_sizeare balanced.
3. Tempo Trace Ingestion Failures
Symptom: Traces are missing or incomplete in Tempo. Cause:
- OTel Collector not forwarding traces correctly.
- Tempo ingesters are overloaded or unhealthy.
- Incorrect service instrumentation (e.g., missing trace context propagation). Fix:
- Check OTel Collector logs: Look for errors exporting to Tempo.
- Monitor Tempo ingesters: Check their health and resource utilization. Scale if necessary.
- Verify instrumentation: Ensure all services are correctly propagating trace context (e.g., W3C Trace Context headers). Use
curlwithtraceparentheaders to test. - Increase Tempo ingester
max_block_bytes: If traces are very large, this might be a bottleneck.
4. Mimir High Cardinality Issues
Symptom: Mimir ingesters/compactors are struggling, high storage costs, slow metric queries. Cause: Too many unique label combinations on metrics. Fix:
- Metric Relabeling: Use OTel Collector's
metricstransformprocessor or Prometheusrelabel_configsto drop or rename high-cardinality labels before ingestion.yamlprocessors: metricstransform: transforms: - include: "http_request_duration_seconds_bucket" action: "delete_label" label: "request_id" # Example: delete high-cardinality label - Review instrumentation: Educate developers on avoiding high-cardinality labels.
- Mimir limits: Configure Mimir's
max_series_per_userandmax_label_names_per_seriesto prevent abuse.
Frequently Asked Questions
Q1: How do I handle multi-tenancy with the LGTM stack?
A1: Mimir, Loki, and Tempo are designed for multi-tenancy.
- Mimir: Use the
X-Scope-OrgIDHTTP header (or OTLP attributetenant_id) to separate tenants. Each tenant gets isolated data. - Loki: Similar to Mimir, use
X-Scope-OrgIDor theloki.tenant.idOTLP log attribute. - Tempo: Use
X-Scope-OrgIDor thetenant_idOTLP resource attribute. - Grafana: Configure separate data sources for each tenant, or use template variables for
X-Scope-OrgIDin dashboards.
Q2: What's the recommended way to deploy the LGTM stack in Kubernetes?
A2: Helm charts are the standard. Grafana Labs provides official charts for Loki, Mimir, and Tempo. For the OpenTelemetry Collector, use the opentelemetry-collector Helm chart. Ensure you configure persistent storage for stateful components (e.g., Mimir ingesters, Loki ingesters) and use a robust object storage backend.
Q3: How do I monitor the LGTM stack itself?
A3: The LGTM components expose Prometheus metrics. Deploy a dedicated Prometheus instance (or use Mimir itself) to scrape these metrics. Create Grafana dashboards to monitor their health, resource usage, ingest rates, query latencies, and error rates. Key metrics include loki_ingester_received_entries_total, mimir_ingester_received_samples_total, tempo_ingester_traces_received_total, and various _grpc_server_handled_total metrics.
Q4: Can I use existing Prometheus metrics with Mimir?
A4: Yes. Mimir is a long-term storage for Prometheus. You can configure your existing Prometheus servers to remote-write to Mimir, or use the OpenTelemetry Collector's Prometheus receiver and OTLP exporter to Mimir. The latter is generally preferred for a unified OTel approach.
Q5: What are the trade-offs between Loki's label-based indexing and traditional full-text log indexing?
A5:
| Feature | Loki (Label-based) | Traditional (Full-text) |
|---|---|---|
| Cost | Lower storage, less compute for indexing | Higher storage, more compute for indexing |
| Query Speed | Fast for label-filtered queries, slower for text | Fast for full-text searches, slower for complex filters |
| Scalability | Highly scalable horizontally | Scalable, but index management can be complex |
| Flexibility | Requires pre-defined labels for efficient search | Searchable on any log content |
| Use Case | Operational debugging, known patterns | Security analysis, ad-hoc exploration |
Loki excels when you know what you're looking for (e.g., logs from service=X in namespace=Y). For arbitrary full-text searches across vast, unstructured logs, traditional systems might be more performant but at a significantly higher cost.
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

OpenTelemetry Collector & eBPF in Production: Zero-Code Tracing, Metrics & Prometheus Pipelines
Comprehensive guide covering opentelemetry collector & ebpf in production: zero-code tracing, metrics & prometheus pipelines with production-grade architecture and code examples.
Read more
Prometheus Grafana Alerting and Burn Rate Design
Design production Prometheus and Grafana alerts using multi-window multi-burn-rate PromQL rules for service-level objectives (SLOs) and error budgets.
Read more
High-Performance Observability with eBPF in Kubernetes: Bypassing the Sidecar Tax
Deep dive into eBPF observability in Kubernetes: eliminating Envoy sidecars, kernel probes, BPF ring buffers, and zero-code telemetry instrumentation.
Read more