Kafka architecture is a cluster of brokers that store each topic as a set of partitioned, replicated, append-only logs, a small quorum of KRaft controllers that manages the cluster's metadata, and client applications that write to and read from those partitions directly. There is no routing tier in the middle: a producer sends each record straight to the broker that leads the target partition, and consumers pull from it. Since Apache Kafka 4.0, released on 18 March 2025, ZooKeeper is gone and KRaft is the only way to run the control plane.
The diagram shows how the pieces connect; the sections after it cover each one and the detail that matters in production.
Kafka architecture diagram
Three brokers hold one topic, orders, with three partitions and three replicas each. Every broker leads one partition and follows the other two, three KRaft controllers hold the metadata, and the billing consumer group splits the partitions between its two members.
Brokers, topics and partitions
A broker is a Kafka server that stores data and serves client requests. A topic is a named stream of events, such as payments or orders, and every topic is split into partitions spread across the brokers. Each partition is an ordered, append-only log: new events go on the end and receive the next sequential offset. Reading an event does not delete it; retention does, so several applications can read the same data at their own pace.
Ordering is guaranteed within a partition, not across a topic, and events with the same key always go to the same partition, which is why an order ID or customer ID is the usual key. The partition count also caps parallelism, because inside one consumer group each partition is read by only one consumer at a time.
Choose the count with care. Kafka cannot reduce the number of partitions in a topic, and adding partitions later changes which partition a given key maps to, which breaks per-key ordering for data already in flight. See our Kafka partition strategy guide.
Replication, leaders and the in-sync replica set
Every partition has a replication factor. One replica is the leader and handles all writes; the others are followers that copy the leader's log. Kafka tracks which followers are caught up in the in-sync replica set (ISR). A write counts as committed only when every replica in the ISR has it, and only ISR members are eligible to become leader. With f+1 replicas, a topic survives f broker failures without losing committed data.
Three settings decide how much that guarantee is worth in practice:
- Replication factor. Three is the common production setting.
ackson the producer. Withacks=allthe producer waits for the full ISR. Withacks=1, only the leader has the record when the producer is told the write succeeded.min.insync.replicason the topic. If the ISR shrinks below this number, the partition rejectsacks=allwrites rather than accept them onto too few copies.
Replication factor 3 with min.insync.replicas=2 and acks=all is the usual durable baseline: one broker can be down for maintenance and writes continue on two copies. Unclean leader election, which lets an out-of-sync replica take over at the cost of data, has been off by default since Kafka 0.11.0.0.
KRaft controllers: the control plane without ZooKeeper
Brokers hold the data; controllers hold the metadata: which topics exist, where each partition's replicas live, which replica leads, and who is in each ISR. In KRaft mode the controllers form a Raft quorum and keep that state in a replicated metadata log. One is active; the others are hot standbys.
The process.roles setting decides what a node does: broker, controller, or both. Kafka's documentation does not recommend the combined mode for critical deployments, because controllers then cannot be rolled or scaled separately from brokers. Production clusters typically run three or five dedicated controllers. A majority must be up, so three controllers tolerate one failure and five tolerate two.
Kafka 4.0 removed ZooKeeper completely, so a cluster still running on ZooKeeper has to migrate to KRaft on a 3.x release before it can move to 4.x. Version support dates and the upgrade path are in our Kafka end-of-life guide.
Producers, consumers and consumer groups
A producer asks any broker for metadata, learns which broker leads each partition, and sends batches straight to the leaders. The record key picks the partition; records without a key are spread across partitions. Batching and compression happen on the client.
Consumers pull. Each one fetches from an offset and commits its position, so the state Kafka keeps per partition for a group is a single number. Consumers that share a group.id form a consumer group, and Kafka assigns each partition to exactly one member. A group with more consumers than partitions leaves the extras idle, and two different groups each receive the full stream.
Two recent releases changed this picture. Kafka 4.0 made the next-generation rebalance protocol (KIP-848) generally available. It is enabled by default on the broker, and clients opt in with group.protocol=consumer; it replaces the stop-the-world rebalances behind many consumer group incidents. Kafka 4.2 then declared share groups (KIP-932, Queues for Kafka) production-ready. A share group lets several consumers read the same partition with per-record acknowledgement, giving queue-style work sharing on ordinary topics.
Storage: log segments, the page cache and tiered storage
On disk, a partition is a directory of segment files. The active segment takes new writes; when it reaches segment.bytes (1 GiB by default) or its time limit, it is closed and a new one starts. Retention deletes whole segments, by time (seven days by default) or by size (disabled by default), and a segment goes only when its newest record has expired. Our Kafka retention policy guide covers the settings.
Reads are served from the page cache, and data moves from page cache to socket through the sendfile system call, so consumers that keep up often cause no disk reads at all. That zero-copy path is not used when TLS is enabled, because encryption happens in user space.
Tiered storage, production-ready since Kafka 3.9, adds a remote tier. Closed segments are copied to an external store such as S3 or HDFS, and local.retention.ms or local.retention.bytes limits how much stays on broker disks. It does not support compacted topics.
Kafka Connect and Kafka Streams
Both are built on the standard producer and consumer clients, so they inherit the same rules about partitions, groups and offsets.
- Kafka Connect runs on separate worker servers and moves data between Kafka and other systems: change data capture from a relational database into a topic, for example, or a topic into object storage or a search index.
- Kafka Streams is a Java library that runs inside your application. It reads topics, transforms, joins and aggregates them, and writes the results to other topics. Its local state stores are backed by changelog topics in Kafka, so a restarted instance can rebuild them.
Because both are clients, their failures usually show up as client problems: lag, rebalances, or a connector task stuck on one bad record.
What is Kafka used for?
Kafka moves and stores streams of events between systems that should not depend on each other directly. The Kafka project lists these core uses:
- Messaging between services, replacing a traditional broker for high-throughput, durable traffic.
- Website activity tracking, Kafka's original use case: page views, searches and other user actions published to topics.
- Metrics and log aggregation from many hosts into central feeds.
- Stream processing of raw events into enriched or aggregated topics.
- Event sourcing and commit logs, where the topic is the record of state changes and log compaction keeps the latest value for each key.
In practice: payments, shipment tracking, IoT data and microservice event backbones. Kafka is a weaker fit for small task queues that need per-message routing, priorities or delayed delivery. That is RabbitMQ territory, and our RabbitMQ vs Kafka comparison covers where the line falls.
When to get help
Most Kafka incidents trace back to a decision described above: a partition count chosen before the key was understood, min.insync.replicas left at 1, controllers sharing hosts with busy brokers, or retention sized to disks that tiered storage could have relieved. If you want a second opinion on an existing cluster, or engineers on call when it misbehaves, AceMQ provides independent Kafka support for open-source Apache Kafka, and Kafka managed services if you would rather hand over day-to-day operations.
Sources
- Apache Kafka: Introduction
- Apache Kafka 4.3 documentation: Use cases
- Apache Kafka 4.3 documentation: Design (replication, ISR, consumers, sendfile)
- Apache Kafka 4.3 documentation: Log implementation (segments and deletion)
- Apache Kafka 4.3 documentation: Topic configs (segment.bytes, retention.ms)
- Apache Kafka 4.3 documentation: Basic Kafka operations (partition changes)
- Apache Kafka 4.3 documentation: KRaft
- Apache Kafka 4.3 documentation: Tiered storage
- Apache Kafka 4.3 documentation: Kafka Streams architecture
- Apache Kafka 4.0.0 release announcement, 18 March 2025
- Apache Kafka 4.2.0 release announcement, 17 February 2026
- Apache Kafka 3.9.0 release announcement, 6 November 2024
Frequently Asked Questions
What are the main components of Kafka architecture?
Brokers that store each topic as partitioned, replicated logs; a quorum of KRaft controllers that manages cluster metadata; producers that write to partition leaders; and consumers, organized into consumer groups, that read from them. Kafka Connect and Kafka Streams are built on the same clients.
Does Kafka still use ZooKeeper?
No. Apache Kafka 4.0, released on 18 March 2025, is the first major release that runs entirely without ZooKeeper. Cluster metadata is managed by KRaft controllers. A cluster still on ZooKeeper has to migrate to KRaft on a 3.x release before it can upgrade to 4.x.
How many KRaft controllers does a Kafka cluster need?
Usually three or five. A majority must be available, so three controllers tolerate one failure and five tolerate two. Kafka's documentation does not recommend combined broker and controller nodes for critical deployments.
What is the ISR in Kafka?
The in-sync replica set: the replicas of a partition that are caught up with the leader. A write is committed once every ISR member has it, and only ISR members can be elected leader. With acks=all and min.insync.replicas set on the topic, a partition refuses writes when too few replicas are in sync.
What is Kafka used for?
Moving and storing streams of events between systems: messaging between services, website activity tracking, metrics and log aggregation, stream processing, event sourcing and as a commit log. In practice that means payments, order processing, IoT and fleet data, and the event backbone for microservices.
Can you reduce the number of partitions in a Kafka topic?
No. Kafka does not support reducing the partition count of a topic. You can add partitions, but that changes which partition a key maps to and so affects per-key ordering. To reduce partitions, create a new topic and move producers and consumers to it.
What is Kafka tiered storage?
A feature, production-ready since Kafka 3.9, that copies closed log segments to external storage such as S3 or HDFS, so brokers keep only recent data on local disk. It does not support compacted topics.
Go deeper on Kafka
- GuideThe Kafka Operations GuideRead the guide
- GuideBuying Kafka Support: The GuideRead the guide
- GuideThe Kafka Monitoring and Alerting GuideRead the guide
- GuideThe Kafka Migration GuideRead the guide
- GuideKafka for AI Agents Without Confluent CloudRead the guide
- ComparisonNATS vs Kafka: Core NATS, JetStream and Kafka ComparedSee the comparison
- ComparisonKafka vs Kinesis ComparedSee the comparison
- ComparisonManaged Kafka Options ComparedSee the comparison