Justin Mclean

@jmclean

What happens when a node dies

Submitted Oct 8, 2026

Proposal template

You do not need to answer everything perfectly in the first pass. Start with what you know. Short, rough answers are fine. The editorial team may follow up to help shape the proposal further.

What Happens when a Node Dies?

What clients and operators see in the seconds after a node, disk or network fails in an Apache Iggy cluster, and the deterministic simulator the project uses to prove it.

Every replicated system says it survives node failure. Fewer can say exactly what a client sees in the seconds afterwards: which writes are lost, which come back twice, and how long the cluster takes to elect a new primary. Replication bugs are rare, timing-dependent and nearly impossible to reproduce from logs, so teams either trust the paper or build integration tests that never hit the interleaving that matters. The session walks through one system’s answers, with the trade-offs each one makes, and the testing method that makes those failures reproducible.

Who is the intended audience? SRE, Platform Engineering, Infrastructure, Developers

Level Intermediate

Practical takeaways

  • A checklist of failure questions to ask of any replicated system you run, with the answers for one real system as a worked example.
  • How to put real production code inside a deterministic simulator so a failure replays from one seed.

Architecture or design decisions; production experience; failure modes and debugging; trade-offs and alternatives; open-source tooling. Specifically: the five-second election wait, client reconnect and leader discovery, in-flight writes returning errors rather than silent resends, quorum versus on-disk acknowledgment, the two-minute disk watchdog, and checksum, fence, and repair on restart. Then the simulator: real replicas, real consensus code and state machines on one thread, with clock, storage and network swapped for in-memory doubles, so every run is a pure function of a seed.

Open-source project; production system. Apache Iggy shipped clustering in 0.9.0, and the simulator is in the main repository, run in the project’s test suite. Three consensus bugs in the production code were found by the simulator, each reproducible from one seed and a handful of flags.

Behaviour that surprises people in practice, and is now documented

Viewstamped Replication rather than Raft, because it maps onto a thread-per-core runtime where every core owns its data. Acknowledge on quorum by default rather than on disk. Clustering is off by default, so a single node stays simple. A simulator that runs the real code rather than a model of it which means the bugs it finds are real.

A set of operational practices: the questions to ask before trusting “fault tolerant”. A debugging technique: deterministic simulation of real code. A way to evaluate competing approaches.

Current state: production experience; open source

#failurestory #distributedsystems #sre #databases #scalability #opensource #rust

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy