Kafka Tiered Storage Architecture: Slashing EBS Storage Costs with AWS S3 & Google Cloud Storage

Table of Contents(19 sections)
Apache Kafka's traditional architecture tightly couples compute and storage. Brokers manage local log segments, typically backed by high-performance block storage like AWS EBS or GCP Persistent Disk. While this design offers unparalleled throughput and low-latency access for recent data, it becomes economically unsustainable for long-term data retention, especially as data volumes scale into petabytes. The cost of provisioned IOPS and storage capacity on block devices for historical, infrequently accessed data can quickly dominate infrastructure budgets.
Kafka Tiered Storage, introduced as KIP-405, fundamentally re-architects Kafka's storage layer. It enables the seamless offloading of older, less frequently accessed log segments from local broker storage to cost-effective, highly durable object storage solutions like AWS S3 and Google Cloud Storage (GCS). This separation of compute and storage allows for significant cost reductions, enhanced scalability, and simplified operational management without compromising Kafka's core guarantees.
Understanding Kafka Tiered Storage (KIP-405)
KIP-405 introduces a two-tiered storage model: a local tier on the Kafka brokers and a remote tier on object storage. The primary objective is to leverage the cost-effectiveness and virtually infinite scalability of cloud object storage for historical data while retaining the performance benefits of local storage for hot data.
Core Concepts
- Separation of Compute and Storage: Brokers are no longer solely responsible for storing all topic data. They primarily serve recent data from local disks and offload older data to the remote tier. This allows for independent scaling of compute (brokers) and storage (object storage).
- Local Storage Layer: This is the traditional Kafka log directory on the broker's local disk. It holds the most recent log segments, providing low-latency read/write access for active producers and consumers. The retention period for this layer is configurable.
- Remote Storage Layer: This is the cloud object storage (S3, GCS). Once log segments are deemed "cold" based on configured retention policies, they are asynchronously uploaded to this layer. This layer acts as an immutable, highly durable archive for all historical data.
- Remote Storage Manager: A component within the Kafka broker responsible for managing the lifecycle of log segments between local and remote storage. It handles the upload of segments to object storage and the retrieval of segments when consumers request data beyond local retention.
- Metadata Management: To enable seamless consumption from both tiers, Kafka maintains metadata about which segments reside locally and which have been offloaded to remote storage. This metadata is crucial for consumers to transparently fetch data regardless of its physical location.
Benefits of Tiered Storage
- Significant Cost Reduction: Object storage is orders of magnitude cheaper than block storage (EBS/Persistent Disk) for long-term retention. This is the primary driver for adoption.
- Enhanced Scalability: Decoupling storage from brokers allows for scaling storage capacity almost infinitely without adding more brokers or increasing local disk sizes.
- Simplified Operations: Brokers can be provisioned with smaller, faster local disks, reducing recovery times during failures and simplifying disk management.
- Extended Data Retention: Retain data for months or years without prohibitive costs, enabling new analytical use cases and compliance requirements.
- Improved Broker Resilience: Smaller local disk footprints mean faster data replication and recovery in the event of a broker failure.
Architectural Deep Dive
The tiered storage architecture fundamentally alters how Kafka manages its log segments. Understanding the data flow and component interactions is critical for effective deployment and optimization.
Data Flow and Component Interaction
- Producer Writes: Producers append messages to topic partitions. These messages are written to the active log segment on the broker's local disk.
- Segment Rollover: When an active log segment reaches its size (
log.segment.bytes) or time (log.segment.ms) limit, it is rolled over, becoming an immutable, closed segment. - Local Retention Policy: The broker applies
local.retention.bytesorlocal.retention.msto determine which segments to keep locally. Segments older than this policy are marked for offloading. - Remote Segment Upload: The
RemoteLogManagerasynchronously uploads these marked segments to the configured object storage bucket (S3 or GCS). This process is non-blocking to broker operations. - Local Segment Deletion: Once a segment is successfully uploaded to remote storage and its metadata is updated, the local copy can be deleted, freeing up local disk space.
- Consumer Reads:
- Hot Data: If a consumer requests data within the local retention window, the broker serves it directly from its local disk, maintaining low latency.
- Cold Data: If a consumer requests data that has been offloaded to remote storage, the broker transparently fetches the required segments from S3/GCS, caches them temporarily, and serves them to the consumer. This introduces additional latency.
Key Configuration Parameters
remote.log.storage.enable: Global broker setting to enable/disable tiered storage. Must betrue.remote.log.storage.manager.impl.class: Specifies the implementation class for the remote storage manager.- For AWS S3:
org.apache.kafka.server.log.remote.metadata.storage.S3RemoteLogMetadataManager - For Google Cloud Storage:
org.apache.kafka.server.log.remote.metadata.storage.GcsRemoteLogMetadataManager
- For AWS S3:
remote.log.storage.manager.impl.prefix: A prefix for all remote storage manager related configurations. E.g.,remote.log.storage.manager.impl.aws.region.local.retention.bytes: The maximum number of bytes to retain on local disk per partition. When this limit is exceeded, older segments are offloaded.local.retention.ms: The maximum time (in milliseconds) to retain segments on local disk per partition.log.segment.bytes: The maximum size of a log segment file. Smaller segments mean more frequent rollovers and uploads, potentially increasing object storage API calls.log.segment.ms: The maximum time before a log segment is rolled over.
Configuration and Implementation
Implementing Kafka Tiered Storage requires careful configuration at both the broker and topic levels. We will detail the setup for AWS S3 and Google Cloud Storage.
Prerequisites
- Kafka Version: Apache Kafka 3.0.0 or later for KIP-405. For production, 3.3.x or newer is recommended for stability and features.
- Cloud Credentials:
- AWS S3: IAM role or user with permissions to
s3:PutObject,s3:GetObject,s3:DeleteObject,s3:ListBucket,s3:GetBucketLocationon the target S3 bucket. - Google Cloud Storage: Service account key with
Storage Object AdminorStorage Object CreatorandStorage Object Viewerroles on the target GCS bucket.
- AWS S3: IAM role or user with permissions to
- Cloud Bucket: An S3 bucket or GCS bucket created in the desired region.
Broker Configuration (server.properties)
The following parameters are added to server.properties on each Kafka broker.
Common Tiered Storage Settings
# Enable remote storage
remote.log.storage.enable=true
# Set the local retention policy.
# Example: Retain 10GB per partition locally.
# This value should be carefully chosen based on hot data access patterns.
local.retention.bytes=10737418240 # 10 GB
# Alternatively, or in addition to bytes, set time-based local retention.
# Example: Retain 24 hours of data locally.
# local.retention.ms=86400000 # 24 hours
# Configure the segment size and time for optimal offloading.
# Smaller segments mean more frequent uploads but faster local deletion.
log.segment.bytes=1073741824 # 1 GB
log.segment.ms=604800000 # 7 days (default)
AWS S3 Specific Configuration
# S3 Remote Log Storage Manager implementation class
remote.log.storage.manager.impl.class=org.apache.kafka.server.log.remote.metadata.storage.S3RemoteLogMetadataManager
# S3 bucket name for remote storage
remote.log.storage.manager.impl.aws.bucket.name=your-kafka-tiered-storage-s3-bucket
# AWS region of the S3 bucket
remote.log.storage.manager.impl.aws.region=us-east-1
# (Optional) AWS endpoint for S3, useful for S3-compatible storage or private endpoints
# remote.log.storage.manager.impl.aws.endpoint=https://s3.us-east-1.amazonaws.com
# (Optional) AWS credentials provider. Default is DefaultAWSCredentialsProviderChain.
# remote.log.storage.manager.impl.aws.credentials.provider.class=com.amazonaws.auth.DefaultAWSCredentialsProviderChain
# (Optional) Number of threads for S3 uploads/downloads
# remote.log.storage.manager.impl.aws.num.client.threads=10
Google Cloud Storage Specific Configuration
# GCS Remote Log Storage Manager implementation class
remote.log.storage.manager.impl.class=org.apache.kafka.server.log.remote.metadata.storage.GcsRemoteLogMetadataManager
# GCS bucket name for remote storage
remote.log.storage.manager.impl.gcp.bucket.name=your-kafka-tiered-storage-gcs-bucket
# GCP Project ID
remote.log.storage.manager.impl.gcp.project.id=your-gcp-project-id
# Path to the service account key file (JSON).
# Ensure this file is accessible by the Kafka broker process.
# Alternatively, if running on GCP, use default credentials via instance service account.
# remote.log.storage.manager.impl.gcp.credentials.file=/path/to/your/gcp-service-account-key.json
# (Optional) Number of threads for GCS uploads/downloads
# remote.log.storage.manager.impl.gcp.num.client.threads=10
Topic Configuration
Tiered storage can be enabled and configured at the topic level, overriding broker-level defaults.
# Create a new topic with remote storage enabled and specific local retention
kafka-topics.sh --create --topic my-tiered-topic \
--bootstrap-server localhost:9092 \
--partitions 3 --replication-factor 3 \
--config remote.storage.enable=true \
--config local.retention.bytes=5368709120 # 5 GB local retention for this topic
# Modify an existing topic to enable remote storage
kafka-configs.sh --bootstrap-server localhost:9092 --entity-type topics --entity-name my-existing-topic \
--alter --add-config remote.storage.enable=true
# Modify local retention for an existing tiered topic
kafka-configs.sh --bootstrap-server localhost:9092 --entity-type topics --entity-name my-existing-topic \
--alter --add-config local.retention.ms=172800000 # 48 hours local retention
Kubernetes Deployment with Strimzi
Strimzi simplifies running Kafka on Kubernetes. To enable tiered storage, you modify the Kafka custom resource.
Strimzi Kafka CR Example
apiVersion: kafka.strimzi.io/v1beta2
kind: Kafka
metadata:
name: my-kafka-cluster
spec:
kafka:
version: 3.6.0 # Ensure Kafka version supports KIP-405
replicas: 3
listeners:
- name: plain
port: 9092
type: internal
tls: false
- name: tls
port: 9093
type: internal
tls: true
storage:
type: jbod
volumes:
- id: 0
type: persistent-claim
size: 100Gi # Smaller local storage needed due to tiered storage
deleteClaim: false
config:
# Common Tiered Storage Settings
remote.log.storage.enable: "true"
local.retention.bytes: "10737418240" # 10 GB
log.segment.bytes: "1073741824" # 1 GB
# AWS S3 Specific Configuration
remote.log.storage.manager.impl.class: "org.apache.kafka.server.log.remote.metadata.storage.S3RemoteLogMetadataManager"
remote.log.storage.manager.impl.aws.bucket.name: "your-kafka-tiered-storage-s3-bucket"
remote.log.storage.manager.impl.aws.region: "us-east-1"
# For AWS, ensure the EC2 instance profile or Kubernetes service account has the necessary IAM role.
# Strimzi can inject environment variables for AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY if needed,
# but IAM roles are preferred.
# OR Google Cloud Storage Specific Configuration
# remote.log.storage.manager.impl.class: "org.apache.kafka.server.log.remote.metadata.storage.GcsRemoteLogMetadataManager"
# remote.log.storage.manager.impl.gcp.bucket.name: "your-kafka-tiered-storage-gcs-bucket"
# remote.log.storage.manager.impl.gcp.project.id: "your-gcp-project-id"
# For GCP, use Workload Identity to bind a Kubernetes Service Account to a GCP Service Account.
# The GCP Service Account should have Storage Object Admin permissions.
# Strimzi will automatically pick up credentials if the pod's service account is configured correctly.
# Example of how to configure a KafkaTopic with remote storage enabled
entityOperator:
topicOperator: {}
userOperator: {}
---
apiVersion: kafka.strimzi.io/v1beta2
kind: KafkaTopic
metadata:
name: my-tiered-topic
labels:
strimzi.io/cluster: my-kafka-cluster
spec:
partitions: 3
replicas: 3
config:
remote.storage.enable: "true"
local.retention.bytes: "5368709120" # 5 GB local retention for this topic
IAM/GCP Permissions for Strimzi
- AWS: The Kubernetes nodes (or the specific Kafka broker pods if using IRSA - IAM Roles for Service Accounts) must have an IAM role attached that grants
s3:PutObject,s3:GetObject,s3:DeleteObject,s3:ListBucket,s3:GetBucketLocationpermissions to the target S3 bucket. - GCP: Utilize Workload Identity to bind a Kubernetes Service Account to a GCP Service Account. The GCP Service Account requires
Storage Object Adminor equivalent permissions on the GCS bucket.
Cost Optimization Analysis
The primary driver for Kafka Tiered Storage is cost reduction. Let's quantify the potential savings with a realistic scenario.
Scenario:
- Data Ingested: 10 TB/day
- Replication Factor: 3 (standard for high availability)
- Total Raw Data: 30 TB/day
- Retention Policy: 30 days for "hot" data, 1 year for "cold" data.
- Average Broker Count: 10 brokers.
- Region: US East (N. Virginia) -
us-east-1for AWS,us-east4for GCP.
Traditional Kafka (EBS-backed)
For 1 year of retention, the total storage required is 30 TB/day * 365 days = 10,950 TB (approx 11 PB). This is an extreme case for traditional Kafka, often leading to shorter retention. Let's assume a more common 30-day retention for traditional Kafka due to cost constraints.
- Total Storage: 30 TB/day * 30 days = 900 TB (replicated)
- EBS Volume Type:
gp3(cost-effective, balanced performance) - EBS
gp3Cost: $0.08/GB-month - Total EBS Cost: 900,000 GB * 0.08/GB-month = **72,000/month**
This calculation only covers storage. It does not include EC2 instance costs, network transfer, or IOPS. The high storage cost often forces organizations to drastically reduce retention.
Tiered Storage Kafka
With tiered storage, we can retain 30 days locally and 1 year remotely.
-
Local Storage (30 days):
- Total Storage: 900 TB (replicated)
- This storage is distributed across brokers. If each broker has 100GiB local retention, and we have 10 brokers, that's 1TB total local storage. This is a simplification, as
local.retention.bytesis per partition. Let's assume 100GiB per broker for active data. - 10 brokers * 100 GiB/broker = 1000 GiB = 1 TB
- EBS
gp3Cost: 1,000 GB * 0.08/GB-month = **80/month**
-
Remote Storage (1 year):
- Total Raw Data: 10 TB/day * 365 days = 3,650 TB (unreplicated, as S3/GCS handles durability)
- AWS S3 Standard Cost: 0.023/GB for first 50TB, 0.022/GB for next 450TB, etc. (average $0.022/GB-month for this scale)
- Total S3 Storage Cost: 3,650,000 GB * 0.022/GB-month = **80,300/month**
- GCS Standard Cost: 0.020/GB for first 1PB, etc. (average 0.020/GB-month for this scale)
- Total GCS Storage Cost: 3,650,000 GB * 0.020/GB-month = **73,000/month**
-
Object Storage API Calls:
- Assume 1GB segments, 10TB/day ingress = 10,000 segments/day.
- S3 PUT requests: 0.005 per 1,000 requests. 10,000 * 30 days = 300,000 PUTs. Cost: 0.005 * 300 = $1.5/month. (Negligible)
- S3 GET requests: Highly variable. If 10% of cold data is accessed monthly (e.g., 365TB/month), and each GET is for a 1GB segment: 365,000 GETs. Cost: 0.0004 per 1,000 requests. 0.0004 * 365 = $0.146/month. (Negligible)
- GCS operations costs are similar, often slightly lower.
-
Total Tiered Storage Cost (AWS S3): 80 (Local EBS) + 80,300 (S3 Storage) + ~2 (API Calls) = **~80,382/month**
-
Total Tiered Storage Cost (GCS): 80 (Local EBS) + 73,000 (GCS Storage) + ~2 (API Calls) = **~73,082/month**
Demonstrating 75%+ Savings
This comparison is tricky because traditional Kafka often cannot afford 1 year of retention. If we compare 30 days of retention on EBS vs. 1 year of retention with Tiered Storage:
- Traditional (30 days EBS): $72,000/month
- Tiered (1 year S3): $80,382/month
This doesn't show savings, it shows increased cost for significantly increased retention. The true saving is in the cost per GB-month for long-term data.
Let's re-evaluate the scenario: What if we had to store 1 year of data on EBS?
- Total EBS Storage for 1 year: 10,950 TB (replicated)
- EBS
gp3Cost: 10,950,000 GB * 0.08/GB-month = **876,000/month**
Now, compare this to Tiered Storage for 1 year:
- Tiered Storage (AWS S3): ~$80,382/month
- Tiered Storage (GCS): ~$73,082/month
Savings Calculation (AWS S3 vs. EBS for 1 year retention): (876,000 - 80,382) / $876,000 = 0.908 = 90.8% savings
Savings Calculation (GCS vs. EBS for 1 year retention): (876,000 - 73,082) / $876,000 = 0.916 = 91.6% savings
This demonstrates that for long-term retention, tiered storage offers over 90% savings on storage costs compared to an all-EBS solution. Even if you only needed 3 months of retention on EBS, the costs would be 3x higher than the 30-day EBS + 1-year S3/GCS tiered approach.
Furthermore, the compute costs for brokers can be reduced. With smaller local disk requirements, you might be able to use smaller EC2 instance types or reduce the number of brokers, further contributing to savings.
Performance Characteristics and Trade-offs
While cost savings are substantial, it's crucial to understand the performance implications of introducing an object storage layer.
Latency and Throughput
- Local Reads: For data within the
local.retentionwindow, performance remains identical to traditional Kafka. Latency is typically sub-millisecond. - Remote Reads: When a consumer requests data that has been offloaded, the broker must fetch it from S3/GCS. This introduces network latency and object storage retrieval overhead.
- Typical Latency: 50ms - 500ms per segment fetch, depending on network conditions, object size, and cloud provider performance. This is significantly higher than local disk access.
- Impact: Consumers reading historical data will experience higher latency and potentially lower throughput. Applications sensitive to this should ensure their working set remains within local retention.
- Writes: Producer write performance is largely unaffected as offloading is asynchronous and non-blocking.
- Broker CPU/Network: Brokers will consume additional CPU for managing remote segments and network bandwidth for uploading/downloading data to/from object storage. This needs to be factored into instance sizing.
Durability and Availability
- Object Storage Durability: AWS S3 and GCS offer extreme durability (11 nines, 99.999999999%) by redundantly storing data across multiple facilities within a region. This is superior to typical EBS RAID configurations.
- Object Storage Availability: Both S3 and GCS offer high availability (99.9% to 99.99% SLA for standard tiers).
- Kafka Durability: The 3x replication factor for local segments still applies. Once a segment is successfully uploaded to object storage, its durability is effectively handled by the cloud provider.
Comparison Table: EBS vs. S3/GCS for Kafka Storage
| Feature | Traditional Kafka (EBS/Persistent Disk) | Kafka Tiered Storage (Local EBS + Remote S3/GCS)
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

The Pragmatic GCP Architecture Guide: Which Services to Actually Use (and What to Avoid)
A battle-tested production guide to Google Cloud Platform. Learn why Cloud Run beats GKE for 90% of workloads, how to leverage BigQuery and Secret Manager, and the 5 hidden cost traps that drain cloud budgets.
Read more
Open Table Formats in 2026: Apache Iceberg v2 vs Delta Lake 3.0 on Cloud Object Storage
Comprehensive guide covering open table formats in 2026: apache iceberg v2 vs delta lake 3.0 on cloud object storage with production-grade architecture and code examples.
Read more
BigQuery + Cloud Run: Building a Production Serverless Data Ingestion Pipeline
A production-grade guide to serverless data ingestion on Google Cloud: the BigQuery Storage Write API, partitioning and clustering strategy, a FastAPI async receiver on Cloud Run, complete Terraform, a real cost breakdown, and the failure modes that page you at 3am.
Read more