•15 min read

Kafka Tiered Storage Architecture: Slashing EBS Storage Costs with AWS S3 & Google Cloud Storage

Kafka Tiered Storage Architecture: Slashing EBS Storage Costs with AWS S3 & Google Cloud Storage

Apache Kafka's traditional architecture tightly couples compute and storage. Brokers manage local log segments, typically backed by high-performance block storage like AWS EBS or GCP Persistent Disk. While this design offers unparalleled throughput and low-latency access for recent data, it becomes economically unsustainable for long-term data retention, especially as data volumes scale into petabytes. The cost of provisioned IOPS and storage capacity on block devices for historical, infrequently accessed data can quickly dominate infrastructure budgets.

Kafka Tiered Storage, introduced as KIP-405, fundamentally re-architects Kafka's storage layer. It enables the seamless offloading of older, less frequently accessed log segments from local broker storage to cost-effective, highly durable object storage solutions like AWS S3 and Google Cloud Storage (GCS). This separation of compute and storage allows for significant cost reductions, enhanced scalability, and simplified operational management without compromising Kafka's core guarantees.

Audio Briefing
0:00 / 0:00

Understanding Kafka Tiered Storage (KIP-405)

KIP-405 introduces a two-tiered storage model: a local tier on the Kafka brokers and a remote tier on object storage. The primary objective is to leverage the cost-effectiveness and virtually infinite scalability of cloud object storage for historical data while retaining the performance benefits of local storage for hot data.

Core Concepts

  1. Separation of Compute and Storage: Brokers are no longer solely responsible for storing all topic data. They primarily serve recent data from local disks and offload older data to the remote tier. This allows for independent scaling of compute (brokers) and storage (object storage).
  2. Local Storage Layer: This is the traditional Kafka log directory on the broker's local disk. It holds the most recent log segments, providing low-latency read/write access for active producers and consumers. The retention period for this layer is configurable.
  3. Remote Storage Layer: This is the cloud object storage (S3, GCS). Once log segments are deemed "cold" based on configured retention policies, they are asynchronously uploaded to this layer. This layer acts as an immutable, highly durable archive for all historical data.
  4. Remote Storage Manager: A component within the Kafka broker responsible for managing the lifecycle of log segments between local and remote storage. It handles the upload of segments to object storage and the retrieval of segments when consumers request data beyond local retention.
  5. Metadata Management: To enable seamless consumption from both tiers, Kafka maintains metadata about which segments reside locally and which have been offloaded to remote storage. This metadata is crucial for consumers to transparently fetch data regardless of its physical location.

Benefits of Tiered Storage

  • Significant Cost Reduction: Object storage is orders of magnitude cheaper than block storage (EBS/Persistent Disk) for long-term retention. This is the primary driver for adoption.
  • Enhanced Scalability: Decoupling storage from brokers allows for scaling storage capacity almost infinitely without adding more brokers or increasing local disk sizes.
  • Simplified Operations: Brokers can be provisioned with smaller, faster local disks, reducing recovery times during failures and simplifying disk management.
  • Extended Data Retention: Retain data for months or years without prohibitive costs, enabling new analytical use cases and compliance requirements.
  • Improved Broker Resilience: Smaller local disk footprints mean faster data replication and recovery in the event of a broker failure.
Advertisement

Architectural Deep Dive

The tiered storage architecture fundamentally alters how Kafka manages its log segments. Understanding the data flow and component interactions is critical for effective deployment and optimization.

Data Flow and Component Interaction

  1. Producer Writes: Producers append messages to topic partitions. These messages are written to the active log segment on the broker's local disk.
  2. Segment Rollover: When an active log segment reaches its size (log.segment.bytes) or time (log.segment.ms) limit, it is rolled over, becoming an immutable, closed segment.
  3. Local Retention Policy: The broker applies local.retention.bytes or local.retention.ms to determine which segments to keep locally. Segments older than this policy are marked for offloading.
  4. Remote Segment Upload: The RemoteLogManager asynchronously uploads these marked segments to the configured object storage bucket (S3 or GCS). This process is non-blocking to broker operations.
  5. Local Segment Deletion: Once a segment is successfully uploaded to remote storage and its metadata is updated, the local copy can be deleted, freeing up local disk space.
  6. Consumer Reads:
    • Hot Data: If a consumer requests data within the local retention window, the broker serves it directly from its local disk, maintaining low latency.
    • Cold Data: If a consumer requests data that has been offloaded to remote storage, the broker transparently fetches the required segments from S3/GCS, caches them temporarily, and serves them to the consumer. This introduces additional latency.

Key Configuration Parameters

  • remote.log.storage.enable: Global broker setting to enable/disable tiered storage. Must be true.
  • remote.log.storage.manager.impl.class: Specifies the implementation class for the remote storage manager.
    • For AWS S3: org.apache.kafka.server.log.remote.metadata.storage.S3RemoteLogMetadataManager
    • For Google Cloud Storage: org.apache.kafka.server.log.remote.metadata.storage.GcsRemoteLogMetadataManager
  • remote.log.storage.manager.impl.prefix: A prefix for all remote storage manager related configurations. E.g., remote.log.storage.manager.impl.aws.region.
  • local.retention.bytes: The maximum number of bytes to retain on local disk per partition. When this limit is exceeded, older segments are offloaded.
  • local.retention.ms: The maximum time (in milliseconds) to retain segments on local disk per partition.
  • log.segment.bytes: The maximum size of a log segment file. Smaller segments mean more frequent rollovers and uploads, potentially increasing object storage API calls.
  • log.segment.ms: The maximum time before a log segment is rolled over.

Configuration and Implementation

Implementing Kafka Tiered Storage requires careful configuration at both the broker and topic levels. We will detail the setup for AWS S3 and Google Cloud Storage.

Prerequisites

  1. Kafka Version: Apache Kafka 3.0.0 or later for KIP-405. For production, 3.3.x or newer is recommended for stability and features.
  2. Cloud Credentials:
    • AWS S3: IAM role or user with permissions to s3:PutObject, s3:GetObject, s3:DeleteObject, s3:ListBucket, s3:GetBucketLocation on the target S3 bucket.
    • Google Cloud Storage: Service account key with Storage Object Admin or Storage Object Creator and Storage Object Viewer roles on the target GCS bucket.
  3. Cloud Bucket: An S3 bucket or GCS bucket created in the desired region.

Broker Configuration (server.properties)

The following parameters are added to server.properties on each Kafka broker.

Common Tiered Storage Settings

# Enable remote storage
remote.log.storage.enable=true

# Set the local retention policy.
# Example: Retain 10GB per partition locally.
# This value should be carefully chosen based on hot data access patterns.
local.retention.bytes=10737418240 # 10 GB

# Alternatively, or in addition to bytes, set time-based local retention.
# Example: Retain 24 hours of data locally.
# local.retention.ms=86400000 # 24 hours

# Configure the segment size and time for optimal offloading.
# Smaller segments mean more frequent uploads but faster local deletion.
log.segment.bytes=1073741824 # 1 GB
log.segment.ms=604800000 # 7 days (default)

AWS S3 Specific Configuration

# S3 Remote Log Storage Manager implementation class
remote.log.storage.manager.impl.class=org.apache.kafka.server.log.remote.metadata.storage.S3RemoteLogMetadataManager

# S3 bucket name for remote storage
remote.log.storage.manager.impl.aws.bucket.name=your-kafka-tiered-storage-s3-bucket

# AWS region of the S3 bucket
remote.log.storage.manager.impl.aws.region=us-east-1

# (Optional) AWS endpoint for S3, useful for S3-compatible storage or private endpoints
# remote.log.storage.manager.impl.aws.endpoint=https://s3.us-east-1.amazonaws.com

# (Optional) AWS credentials provider. Default is DefaultAWSCredentialsProviderChain.
# remote.log.storage.manager.impl.aws.credentials.provider.class=com.amazonaws.auth.DefaultAWSCredentialsProviderChain

# (Optional) Number of threads for S3 uploads/downloads
# remote.log.storage.manager.impl.aws.num.client.threads=10

Google Cloud Storage Specific Configuration

# GCS Remote Log Storage Manager implementation class
remote.log.storage.manager.impl.class=org.apache.kafka.server.log.remote.metadata.storage.GcsRemoteLogMetadataManager

# GCS bucket name for remote storage
remote.log.storage.manager.impl.gcp.bucket.name=your-kafka-tiered-storage-gcs-bucket

# GCP Project ID
remote.log.storage.manager.impl.gcp.project.id=your-gcp-project-id

# Path to the service account key file (JSON).
# Ensure this file is accessible by the Kafka broker process.
# Alternatively, if running on GCP, use default credentials via instance service account.
# remote.log.storage.manager.impl.gcp.credentials.file=/path/to/your/gcp-service-account-key.json

# (Optional) Number of threads for GCS uploads/downloads
# remote.log.storage.manager.impl.gcp.num.client.threads=10

Topic Configuration

Tiered storage can be enabled and configured at the topic level, overriding broker-level defaults.

# Create a new topic with remote storage enabled and specific local retention
kafka-topics.sh --create --topic my-tiered-topic \
  --bootstrap-server localhost:9092 \
  --partitions 3 --replication-factor 3 \
  --config remote.storage.enable=true \
  --config local.retention.bytes=5368709120 # 5 GB local retention for this topic

# Modify an existing topic to enable remote storage
kafka-configs.sh --bootstrap-server localhost:9092 --entity-type topics --entity-name my-existing-topic \
  --alter --add-config remote.storage.enable=true

# Modify local retention for an existing tiered topic
kafka-configs.sh --bootstrap-server localhost:9092 --entity-type topics --entity-name my-existing-topic \
  --alter --add-config local.retention.ms=172800000 # 48 hours local retention

Kubernetes Deployment with Strimzi

Strimzi simplifies running Kafka on Kubernetes. To enable tiered storage, you modify the Kafka custom resource.

Strimzi Kafka CR Example

apiVersion: kafka.strimzi.io/v1beta2
kind: Kafka
metadata:
  name: my-kafka-cluster
spec:
  kafka:
    version: 3.6.0 # Ensure Kafka version supports KIP-405
    replicas: 3
    listeners:
      - name: plain
        port: 9092
        type: internal
        tls: false
      - name: tls
        port: 9093
        type: internal
        tls: true
    storage:
      type: jbod
      volumes:
        - id: 0
          type: persistent-claim
          size: 100Gi # Smaller local storage needed due to tiered storage
          deleteClaim: false
    config:
      # Common Tiered Storage Settings
      remote.log.storage.enable: "true"
      local.retention.bytes: "10737418240" # 10 GB
      log.segment.bytes: "1073741824" # 1 GB

      # AWS S3 Specific Configuration
      remote.log.storage.manager.impl.class: "org.apache.kafka.server.log.remote.metadata.storage.S3RemoteLogMetadataManager"
      remote.log.storage.manager.impl.aws.bucket.name: "your-kafka-tiered-storage-s3-bucket"
      remote.log.storage.manager.impl.aws.region: "us-east-1"
      # For AWS, ensure the EC2 instance profile or Kubernetes service account has the necessary IAM role.
      # Strimzi can inject environment variables for AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY if needed,
      # but IAM roles are preferred.

      # OR Google Cloud Storage Specific Configuration
      # remote.log.storage.manager.impl.class: "org.apache.kafka.server.log.remote.metadata.storage.GcsRemoteLogMetadataManager"
      # remote.log.storage.manager.impl.gcp.bucket.name: "your-kafka-tiered-storage-gcs-bucket"
      # remote.log.storage.manager.impl.gcp.project.id: "your-gcp-project-id"
      # For GCP, use Workload Identity to bind a Kubernetes Service Account to a GCP Service Account.
      # The GCP Service Account should have Storage Object Admin permissions.
      # Strimzi will automatically pick up credentials if the pod's service account is configured correctly.

    # Example of how to configure a KafkaTopic with remote storage enabled
  entityOperator:
    topicOperator: {}
    userOperator: {}
---
apiVersion: kafka.strimzi.io/v1beta2
kind: KafkaTopic
metadata:
  name: my-tiered-topic
  labels:
    strimzi.io/cluster: my-kafka-cluster
spec:
  partitions: 3
  replicas: 3
  config:
    remote.storage.enable: "true"
    local.retention.bytes: "5368709120" # 5 GB local retention for this topic

IAM/GCP Permissions for Strimzi

  • AWS: The Kubernetes nodes (or the specific Kafka broker pods if using IRSA - IAM Roles for Service Accounts) must have an IAM role attached that grants s3:PutObject, s3:GetObject, s3:DeleteObject, s3:ListBucket, s3:GetBucketLocation permissions to the target S3 bucket.
  • GCP: Utilize Workload Identity to bind a Kubernetes Service Account to a GCP Service Account. The GCP Service Account requires Storage Object Admin or equivalent permissions on the GCS bucket.

Cost Optimization Analysis

The primary driver for Kafka Tiered Storage is cost reduction. Let's quantify the potential savings with a realistic scenario.

Scenario:

  • Data Ingested: 10 TB/day
  • Replication Factor: 3 (standard for high availability)
  • Total Raw Data: 30 TB/day
  • Retention Policy: 30 days for "hot" data, 1 year for "cold" data.
  • Average Broker Count: 10 brokers.
  • Region: US East (N. Virginia) - us-east-1 for AWS, us-east4 for GCP.

Traditional Kafka (EBS-backed)

For 1 year of retention, the total storage required is 30 TB/day * 365 days = 10,950 TB (approx 11 PB). This is an extreme case for traditional Kafka, often leading to shorter retention. Let's assume a more common 30-day retention for traditional Kafka due to cost constraints.

  • Total Storage: 30 TB/day * 30 days = 900 TB (replicated)
  • EBS Volume Type: gp3 (cost-effective, balanced performance)
  • EBS gp3 Cost: $0.08/GB-month
  • Total EBS Cost: 900,000 GB * 0.08/GB-month = **72,000/month**

This calculation only covers storage. It does not include EC2 instance costs, network transfer, or IOPS. The high storage cost often forces organizations to drastically reduce retention.

Tiered Storage Kafka

With tiered storage, we can retain 30 days locally and 1 year remotely.

  • Local Storage (30 days):

    • Total Storage: 900 TB (replicated)
    • This storage is distributed across brokers. If each broker has 100GiB local retention, and we have 10 brokers, that's 1TB total local storage. This is a simplification, as local.retention.bytes is per partition. Let's assume 100GiB per broker for active data.
    • 10 brokers * 100 GiB/broker = 1000 GiB = 1 TB
    • EBS gp3 Cost: 1,000 GB * 0.08/GB-month = **80/month**
  • Remote Storage (1 year):

    • Total Raw Data: 10 TB/day * 365 days = 3,650 TB (unreplicated, as S3/GCS handles durability)
    • AWS S3 Standard Cost: 0.023/GB for first 50TB, 0.022/GB for next 450TB, etc. (average $0.022/GB-month for this scale)
    • Total S3 Storage Cost: 3,650,000 GB * 0.022/GB-month = **80,300/month**
    • GCS Standard Cost: 0.020/GB for first 1PB, etc. (average 0.020/GB-month for this scale)
    • Total GCS Storage Cost: 3,650,000 GB * 0.020/GB-month = **73,000/month**
  • Object Storage API Calls:

    • Assume 1GB segments, 10TB/day ingress = 10,000 segments/day.
    • S3 PUT requests: 0.005 per 1,000 requests. 10,000 * 30 days = 300,000 PUTs. Cost: 0.005 * 300 = $1.5/month. (Negligible)
    • S3 GET requests: Highly variable. If 10% of cold data is accessed monthly (e.g., 365TB/month), and each GET is for a 1GB segment: 365,000 GETs. Cost: 0.0004 per 1,000 requests. 0.0004 * 365 = $0.146/month. (Negligible)
    • GCS operations costs are similar, often slightly lower.
  • Total Tiered Storage Cost (AWS S3): 80 (Local EBS) + 80,300 (S3 Storage) + ~2 (API Calls) = **~80,382/month**

  • Total Tiered Storage Cost (GCS): 80 (Local EBS) + 73,000 (GCS Storage) + ~2 (API Calls) = **~73,082/month**

Demonstrating 75%+ Savings

This comparison is tricky because traditional Kafka often cannot afford 1 year of retention. If we compare 30 days of retention on EBS vs. 1 year of retention with Tiered Storage:

  • Traditional (30 days EBS): $72,000/month
  • Tiered (1 year S3): $80,382/month

This doesn't show savings, it shows increased cost for significantly increased retention. The true saving is in the cost per GB-month for long-term data.

Let's re-evaluate the scenario: What if we had to store 1 year of data on EBS?

  • Total EBS Storage for 1 year: 10,950 TB (replicated)
  • EBS gp3 Cost: 10,950,000 GB * 0.08/GB-month = **876,000/month**

Now, compare this to Tiered Storage for 1 year:

  • Tiered Storage (AWS S3): ~$80,382/month
  • Tiered Storage (GCS): ~$73,082/month

Savings Calculation (AWS S3 vs. EBS for 1 year retention): (876,000 - 80,382) / $876,000 = 0.908 = 90.8% savings

Savings Calculation (GCS vs. EBS for 1 year retention): (876,000 - 73,082) / $876,000 = 0.916 = 91.6% savings

This demonstrates that for long-term retention, tiered storage offers over 90% savings on storage costs compared to an all-EBS solution. Even if you only needed 3 months of retention on EBS, the costs would be 3x higher than the 30-day EBS + 1-year S3/GCS tiered approach.

Furthermore, the compute costs for brokers can be reduced. With smaller local disk requirements, you might be able to use smaller EC2 instance types or reduce the number of brokers, further contributing to savings.

Advertisement

Performance Characteristics and Trade-offs

While cost savings are substantial, it's crucial to understand the performance implications of introducing an object storage layer.

Latency and Throughput

  • Local Reads: For data within the local.retention window, performance remains identical to traditional Kafka. Latency is typically sub-millisecond.
  • Remote Reads: When a consumer requests data that has been offloaded, the broker must fetch it from S3/GCS. This introduces network latency and object storage retrieval overhead.
    • Typical Latency: 50ms - 500ms per segment fetch, depending on network conditions, object size, and cloud provider performance. This is significantly higher than local disk access.
    • Impact: Consumers reading historical data will experience higher latency and potentially lower throughput. Applications sensitive to this should ensure their working set remains within local retention.
  • Writes: Producer write performance is largely unaffected as offloading is asynchronous and non-blocking.
  • Broker CPU/Network: Brokers will consume additional CPU for managing remote segments and network bandwidth for uploading/downloading data to/from object storage. This needs to be factored into instance sizing.

Durability and Availability

  • Object Storage Durability: AWS S3 and GCS offer extreme durability (11 nines, 99.999999999%) by redundantly storing data across multiple facilities within a region. This is superior to typical EBS RAID configurations.
  • Object Storage Availability: Both S3 and GCS offer high availability (99.9% to 99.99% SLA for standard tiers).
  • Kafka Durability: The 3x replication factor for local segments still applies. Once a segment is successfully uploaded to object storage, its durability is effectively handled by the cloud provider.

Comparison Table: EBS vs. S3/GCS for Kafka Storage

| Feature | Traditional Kafka (EBS/Persistent Disk) | Kafka Tiered Storage (Local EBS + Remote S3/GCS)

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement