Skip to content

Kafka

Listen to this page4:42
Read the transcript

1. A log, not a queue

Host: So everyone calls Kafka a message queue, but that’s actually the wrong mental model, and I think it’s the source of most of the confusion people have. The real picture is a log — an append-only log, like a ledger that just keeps growing. Today we’re going to build up from that one idea and see how it explains basically everything Kafka does, both the good and the painful parts.

Guest: Right, and the detail that makes it a log instead of a queue is what happens after you read a message: nothing. A traditional queue deletes or hides a message once it’s consumed, but Kafka just retains it by time or by size, so the data sits there and multiple readers can come along later and replay it. To talk about this precisely we need a little vocabulary — a topic is a named log split into partitions, a partition is the actual unit of ordering and storage, each message gets a monotonically increasing offset that the consumer itself tracks, and a consumer group is just a set of consumers sharing that work. Everything else we discuss today is really just consequences of that one design choice.

2. The guarantee everyone misquotes

Host: So let’s hit the phrase people get wrong constantly: ‘Kafka guarantees ordering.’ What’s the actual scope of that promise?

Guest: It’s per partition, full stop, never across a topic. If you want all events for a given user or order in order, that entity’s key has to hash to the same single partition every time, otherwise the messages just interleave across partitions with no relationship to each other. And this bites you in a second way: partition count is also your parallelism ceiling, since only one consumer in a group can read a given partition at a time — add more consumers than partitions and the extras just sit idle. Worse, you can’t shrink partitions once created, and increasing them rehashes existing keys to new partitions, silently breaking that ordering guarantee you built your whole design around, so you really want to size that number up front.

Host: And what happens to the consumer that just can’t keep up with that pace — does Kafka cut it some slack?

Guest: No slack at all — retention is storage-bound, not consumption-bound, so a slow consumer falling behind past that window, commonly seven days, doesn’t pause the clock. The records just age out and get deleted, and when that consumer finally catches up it resumes at the earliest offset still available, silently skipping everything that expired in between. No error, no warning, just a gap in your data that you’ll only notice later, which is why watching consumer lag — in time, not just message count — actually matters.

3. Durability knobs and the exactly-once myth

Host: Okay, so gaps from lag aside — how do you actually control durability on the write side? What’s the knob that decides whether a write survives a broker dying?

Guest: It’s acks, and it’s more subtle than people think. Acks equals zero is fire and forget, acks equals one means the leader wrote it locally and told you success — but if that leader dies before replicating, the write is just gone despite the happy response you got. Acks equals all plus min insync replicas is the real contract: the write only succeeds once enough in-sync replicas, the ISR, have it, so a single dead broker can’t erase it.

Host: So that’s durability solved — but then people say Kafka does exactly-once, and you’re telling me before we started recording that’s the phrase you want to puncture.

Guest: Because it’s true only inside a narrow box: transactions give you exactly-once for Kafka-to-Kafka work, atomically writing to partitions and committing offsets together. The moment you touch anything external — a database row, an API call, an email — you’re back to at-least-once, full stop. So the handler on that side effect still has to be idempotent; delivery semantics aren’t a feature the queue hands you, they’re a contract your application has to uphold.

4. When to reach for Kafka — and where to go next

Host: So let’s land the plane — if I’m staring at a whiteboard deciding between Kafka and something like RabbitMQ, what’s the one question I ask myself?

Guest: That’s really the one question — which side of that trade-off you’re on for this particular workload. And if you want the deeper argument on delivery guarantees and idempotency, or how this same at-least-once ceiling shows up in long-running agent work, the distributed systems module and the durable-agent-execution material walk through both — that’s where I’d go next.

Not covered

The planner wanted these and found nothing in the source to support them:

  • A narrative walkthrough of an actual production Kafka outage or incident timeline
  • Kafka Streams, KSQL, or schema registry — not covered by any excerpt
  • Specific broker internals like ZooKeeper vs KRaft controller architecture

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

A distributed, append-only commit log. Producers append to partitions; consumers read forward at their own pace and track their own position. Messages are retained by time or size rather than deleted on read, which is the property that separates Kafka from a queue and makes replay possible.

Reach for it when you need durable ordered history, replay, or several independent consumers of the same stream. Reach for a broker with per-message acknowledgement and flexible routing when you need a work queue with competing consumers and per-message retry.

Concept What it does
Topic A named log, split into partitions. The unit of subscription
Partition The unit of ordering, parallelism, and storage. Ordering is guaranteed within one, never across
Offset A monotonically increasing position within a partition. The consumer owns it, not the broker
Consumer group A set of consumers sharing a subscription; each partition is assigned to exactly one member
Rebalance Reassignment of partitions when membership changes — briefly stops consumption
Replication factor Copies of each partition across brokers
ISR In-sync replicas: those caught up enough to be eligible for leadership
acks How many acknowledgements a producer waits for: 0, 1, or all
min.insync.replicas With acks=all, the minimum ISR count for a write to succeed
Log compaction Retention that keeps the latest value per key instead of a time window
Idempotent producer Deduplicates retries within a producer session, preventing duplicates from retry alone
Transactions Atomic writes across partitions plus consumer offsets — the real mechanism behind “exactly-once”
Consumer lag Log end offset minus committed offset. The single most useful health metric
Quantity Value Why it matters
Ordering scope One partition The most misquoted guarantee in the system
Max useful consumers per group Partition count Extra consumers sit idle; partitions set your parallelism ceiling
Common replication factor 3 With min.insync.replicas=2, survives one broker loss and still accepts writes
Durable write configuration acks=all + min.insync.replicas=2 Anything less can acknowledge a write that a failover loses
Default delivery semantics At-least-once Exactly-once needs transactions, or idempotent consumers
Partition count changes Increase only You cannot reduce it, and increasing breaks key-to-partition stability
Retention default Commonly 7 days Storage-bound, not consumption-bound — a slow consumer does not extend it
Rebalance cost Consumption pauses Frequent rebalances from long processing loops look exactly like an outage
Practical partitions per cluster Thousands, not millions Each costs file handles, memory, and recovery time
  • “Kafka guarantees ordering” is only true per partition. Multi-partition topics interleave, so ordering per entity requires that entity’s key to hash to one partition.
  • Adding partitions breaks key affinity. Existing keys rehash to different partitions, so per-key ordering is violated across the change. Size partitions before you need them.
  • A slow consumer does not extend retention. Fall behind past the retention window and records are gone — the consumer resumes at the earliest available offset, having silently skipped data.
  • Committing offsets before processing turns at-least-once into at-most-once. Commit after the work is durable, and make the work idempotent, because you will reprocess.
  • acks=1 acknowledges a write the leader has not yet replicated. If that leader dies before replication, the write is gone despite a success response.
  • Rebalances are triggered by slow processing. Exceed max.poll.interval.ms and the group assumes the consumer is dead, reassigns its partitions, and the work is reprocessed elsewhere.
  • Exactly-once is scoped narrower than the phrase suggests. It covers Kafka-to-Kafka processing with transactions. A side effect on an external system is still at-least-once, so the handler must be idempotent regardless.
  • Consumer lag in messages can mislead. Ten thousand small records and ten thousand large ones are very different recovery times; watch lag in time as well as count.
  • Compacted topics still keep tombstones for a while, and compaction is asynchronous — a compacted topic is not a key-value store with read-your-writes.