Raft
Read the transcript
1. One Leader, One Log
Host: So let’s start with the shape of the problem. You’ve got a replicated log, a bunch of nodes, and everyone needs to agree on the same sequence of writes even when machines crash or the network hiccups. Paxos technically solves this, but it’s notorious for being nearly impossible to reason about in practice. Raft showed up specifically to fix that understandability problem, right?
Guest: Exactly, that was the whole design goal — not new capability, just the same guarantees as Paxos but with a mental model humans can actually hold in their heads. And it worked well enough that it’s now sitting underneath etcd, Consul, TiKV, CockroachDB, which means it’s quietly running underneath Kubernetes itself. The core trick is brutally simple: elect one leader, and every write in the entire cluster routes through that single node.
Host: One leader, one log — so walk me through the roles. You’ve got leader, follower, candidate, and this idea of terms that keep everyone from talking over each other.
Guest: Right, followers just sit passively accepting AppendEntries — which doubles as both replication and a heartbeat — until their election timer fires, at which point they become a candidate and request votes for a new term. Terms are the key safety valve: it’s a monotonically increasing epoch number, and any message carrying a higher term forces whoever’s holding old authority to step down immediately, so you can never have two leaders both thinking they’re in charge for the same term. And that single-leader-per-term design is exactly why Raft is unapologetically CP — if a minority partition gets cut off, it simply stops accepting writes rather than risk two sides of a split diverging.
2. The Numbers That Make It Safe
Host: Okay, let’s get concrete then, because ‘majority’ sounds simple until you’re staring at a cluster-sizing decision. Why does everyone land on three or five nodes instead of, say, four or six?
Guest: It comes down to quorum math: you need floor of N over 2 plus 1 nodes to agree, which means failures tolerated is floor of N minus 1 over 2. Run that formula and four nodes only tolerate one failure, exactly the same as three, but every write now has to wait on one more replica. Five gets you to tolerating two failures, so odd sizes are the sweet spot — you’re paying for redundancy you actually get, not padding your latency for nothing.
Host: And that randomized election timeout — walk me through why that specific trick, the 150 to 300 millisecond window, actually matters instead of just picking a fixed number.
Guest: If every follower waited the exact same fixed interval, they’d all time out simultaneously, all become candidates at once, split the vote, and do it again forever — that’s the classic livelock. Randomizing means one node almost always fires first, grabs votes before anyone else wakes up, and the heartbeat interval sits a full order of magnitude below that timeout, so a healthy leader never gets second-guessed. It all rests on one inequality: broadcast time has to be much less than election timeout, which has to be much less than mean time between failures — break the left side and you get endless elections.
3. Where It Bites You in Production
Host: So the theory is solid, the arithmetic checks out — where does this actually go wrong for someone running it in production?
Guest: A few classic traps. First, Raft doesn’t scale throughput, it just gives you safety — every write still funnels through one leader, so if you need more capacity you shard into multiple Raft groups rather than expecting one group to get faster. Second, follower reads can be stale, because a follower doesn’t necessarily know the latest entry is committed, so strong reads need to go through the leader with a lease or a ReadIndex check. And there’s a subtler trap: committed isn’t applied — an entry can be durable on a majority but a client only sees its effect once the state machine actually processes it.
Host: And I’d guess geography and membership changes are where people really get burned?
Guest: Exactly — stretch a Raft group across regions and cross-region round trips start approaching your election timeout, so you get leaders flapping in and out for no real reason; keep the group in one region and replicate across regions some other way. Membership changes are the other landmine — add or remove several nodes at once and you can briefly create two disjoint majorities, so you always change membership one node at a time. Zooming out, this is all just Raft picking consistency over availability during a partition, the CP side of CAP — if you want the deeper reasoning on that tradeoff, or the original Raft paper, or how etcd’s timing constraints shape Kubernetes control-plane topology, those are all worth reading next.
Not covered
The planner wanted these and found nothing in the source to support them:
- A detailed side-by-side comparison of Raft’s mechanics against Paxos
- Byzantine fault-tolerant consensus variants
- Concrete multi-region Raft deployment topologies beyond ‘keep it in one region’
- Real-world etcd/Consul/CockroachDB internals beyond the fact that they use Raft
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
At a Glance
Section titled “At a Glance”Raft keeps a replicated log identical across a cluster by electing a single leader and routing every write through it. It solves the same problem as Paxos and was designed to be understandable, which is why it ended up inside etcd, Consul, TiKV, and CockroachDB — and therefore underneath Kubernetes.
The thing to remember in a design discussion: Raft is CP. A minority partition stops accepting writes rather than diverge.
Key Concepts
Section titled “Key Concepts”| Term | What it does |
|---|---|
| Leader | The single node accepting client writes and replicating them. All others redirect to it |
| Follower | Passive. Accepts AppendEntries, votes in elections, becomes a candidate on election timeout |
| Candidate | A follower whose election timer fired. Requests votes for the next term |
| Term | A monotonically increasing election epoch. Any message carrying a higher term forces a step-down |
| Log entry | (term, index, command). Applied to the state machine only once committed |
| Commit index | The highest index replicated to a majority. Entries at or below it are durable |
AppendEntries |
Replication and heartbeat — an empty one is how a leader keeps its followers quiet |
| Log matching property | If two logs agree on (term, index), they are identical up to that point |
| Snapshot | Compacted state, so the log does not grow forever and a lagging follower can catch up |
| Membership change | Done one node at a time (or via joint consensus) so two majorities can never coexist |
Numbers That Matter
Section titled “Numbers That Matter”| Quantity | Value | Why it matters |
|---|---|---|
| Quorum | ⌊N/2⌋ + 1 |
2 of 3, 3 of 5, 4 of 7 |
| Failures tolerated | ⌊(N−1)/2⌋ |
3 nodes survive 1 failure; 5 survive 2 |
| Why cluster sizes are odd | Even sizes buy no extra tolerance | 4 nodes tolerate 1, exactly like 3, at higher write cost |
| Typical cluster size | 3 or 5 | Beyond 5, replication cost grows faster than the tolerance gain |
| Election timeout | Commonly 150–300 ms, randomized | Randomization is what breaks split votes |
| Heartbeat interval | An order of magnitude below the election timeout | Roughly 10–50 ms for the timeouts above |
| Timing requirement | broadcastTime ≪ electionTimeout ≪ MTBF |
Violating the left inequality causes endless elections |
| Leaders per term | At most 1 | A node votes once per term, so two leaders in one term is impossible |
| Write path round trips | 1 to a majority | Latency is bounded by the median follower, not the slowest |
Common Gotchas
Section titled “Common Gotchas”- Raft gives consensus, not scale. Every write goes through one leader, so throughput does not improve by adding nodes — it gets slightly worse. Scale by sharding into many Raft groups.
- Follower reads can be stale. A read served locally may miss a committed entry. Strong reads
need the leader plus a lease or a
ReadIndexround trip. - A leader that has been partitioned away may not know it. This is why leases and quorum checks exist; without them a stale leader can serve stale reads while a new leader is already elected.
- Committed is not applied. An entry replicated to a majority is durable, but a client sees it only once it reaches the state machine.
- Election timeouts must be randomized. Fixed timeouts cause repeated split votes, and the cluster can churn without electing anyone.
- Geographic spread breaks the timing assumption. Cross-region round trips can approach the election timeout, producing leader flapping. Keep a Raft group inside a region and replicate across regions by another mechanism.
- Membership changes are the dangerous operation. Adding or removing several nodes at once can create two disjoint majorities. Change one at a time.
- Even-sized clusters are strictly worse. More nodes to replicate to, no additional failures tolerated.
Where to Go Deeper
Section titled “Where to Go Deeper”- Module 2: Distributed Systems — where consensus fits relative to leases, idempotency, and delivery guarantees.
- CAP theorem — the frame Raft’s partition behavior sits inside.
- Module 10: Kubernetes — etcd is a Raft cluster, and its timing constraints are the reason control planes have the topology they do.
- Diego Ongaro and John Ousterhout, In Search of an Understandable Consensus Algorithm — the original paper, and unusually readable for the genre.