Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Justin Mclean
@jmclean
Submitted Oct 8, 2026
You do not need to answer everything perfectly in the first pass. Start with what you know. Short, rough answers are fine. The editorial team may follow up to help shape the proposal further.
What Happens when a Node Dies?
What clients and operators see in the seconds after a node, disk or network fails in an Apache Iggy cluster, and the deterministic simulator the project uses to prove it.
Every replicated system says it survives node failure. Fewer can say exactly what a client sees in the seconds afterwards: which writes are lost, which come back twice, and how long the cluster takes to elect a new primary. Replication bugs are rare, timing-dependent and nearly impossible to reproduce from logs, so teams either trust the paper or build integration tests that never hit the interleaving that matters. The session walks through one system’s answers, with the trade-offs each one makes, and the testing method that makes those failures reproducible.
Who is the intended audience? SRE, Platform Engineering, Infrastructure, Developers
Level Intermediate
Practical takeaways
Architecture or design decisions; production experience; failure modes and debugging; trade-offs and alternatives; open-source tooling. Specifically: the five-second election wait, client reconnect and leader discovery, in-flight writes returning errors rather than silent resends, quorum versus on-disk acknowledgment, the two-minute disk watchdog, and checksum, fence, and repair on restart. Then the simulator: real replicas, real consensus code and state machines on one thread, with clock, storage and network swapped for in-memory doubles, so every run is a pure function of a seed.
Open-source project; production system. Apache Iggy shipped clustering in 0.9.0, and the simulator is in the main repository, run in the project’s test suite. Three consensus bugs in the production code were found by the simulator, each reproducible from one seed and a handful of flags.
Behaviour that surprises people in practice, and is now documented
Viewstamped Replication rather than Raft, because it maps onto a thread-per-core runtime where every core owns its data. Acknowledge on quorum by default rather than on disk. Clustering is off by default, so a single node stays simple. A simulator that runs the real code rather than a model of it which means the bugs it finds are real.
A set of operational practices: the questions to ask before trusting “fault tolerant”. A debugging technique: deterministic simulation of real code. A way to evaluate competing approaches.
Current state: production experience; open source
#failurestory #distributedsystems #sre #databases #scalability #opensource #rust
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}