A Kubernetes cluster that is Ready is not a cluster that stays healthy. etcd fragments, nodes drift, registries go stale, and a central control plane is the wrong place to fix any of that. This session covers CLuster Maintainer (CLM): an in-cluster operator that runs periodic maintenance and reconciles drift so a multi-node cluster stays stable without humans.
Day-0 gets you a cluster. Day-2 is what keeps it alive, and it is still mostly people, cron, and hope.
Once a cluster is up — often with many nodes — nobody should have to SSH in to defrag etcd, take a snapshot, fix a newly joined node, or sync a registry. The cluster should remain stable, self-heal when something drifts, and run periodic maintenance on its own. If the only plan is a person or a cron, the job will be missed — schedules drift, the cluster is down at fire time, and “cron fired” is not “maintenance succeeded.”
The failure modes are specific:
- etcd grows and fragments until space quota and latency become an outage. Defrag is invasive. Snapshots exist until the restore you never rehearsed.
- Nodes join Ready with the wrong role, labels, or host config. The API looks fine; the topology is not.
- Registries drift across replicas, unused tags never get garbage-collected, and a later pull fails on the replica that never got the image.
We used to drive this from a central fleet controller. That couples every cluster’s health to one remote system: network path, load, and “did the job run here?” live far from the etcd, disks, and registries that can actually fail.
The engineering problem is here is to make this maintenance a desired state of the cluster: no human in the loop, no remote babysitter, status you can query, and operations that are safe under partial failure.
CLM is that loop. A Custom Resource declares what to maintain. An in-cluster service reconciles etcd, nodes, and registries continuously and on schedule.
Platform engineers, SREs, DevOps, and Kubernetes operators who run self-managed or privately hosted clusters, and engineers building operators who need a concrete Day-2 maintenance case.
Intermediate.
How to treat Day-2 work (etcd defrag and snapshots, node repair, registry sync and GC) as desired state that the cluster reconciles itself, so stability does not depend on a human or a central cron.
A reference architecture for an in-cluster Cluster Maintainer: one CR, pluggable operations, scheduled invasive jobs, continuous drift repair, and status for last success, failure count, and partial completion.
I will share:
- Why “the cluster is Ready” is not “the cluster is maintained,” and why a central control plane is a bad place to run that work.
- The CLM shape: CRD + operator + in-cluster maintainer, with etcd, node, and registry operations behind a common interface.
- etcd: scheduled defrag gated on space-quota threshold, member-by-member execution with health wait, snapshots, why CR status matters more than logs.
- Nodes: informer-driven reconciliation on join — roles and host config, why Ready is not correct.
- Registry: bidirectional image sync, health checks before sync, garbage collection as a first-class job.
- Failure modes: locks across operations, disabling defrag when it would block a non-HA etcd, reporting “2 of 3 members succeeded.”
- Architecture and a walkthrough of CR → operation → status (no product pitch).
Production system. Internal platform engineering project.
CLM runs as an in-cluster operator and maintenance service on Kubernetes fleets. Some of these works used to live in a central control plane; we moved it onto each cluster.
- Cron and runbooks. Schedules drift, owners rotate, nobody can answer “when did this cluster last snapshot?”
- One-shot SSH/Ansible. Works until a node joins at 2am with the wrong role and stays wrong until a human notices.
- Central control plane as the maintainer. The mothership cannot see local quorum, local disk, or local registry health in time. A stuck central job stalls hygiene for the fleet. You still needed a human to interpret logs.
- Treating defrag as etcdctl defrag. Without quorum-aware sequencing and post-op health waits, you trade fragmentation for a worse outage.
- Assuming registry HA means two running active instances. Two endpoints do not stay in sync by magic.
- Do not leave Day-2 to humans, and do not execute it from a remote control plane.
- Declare maintenance intent on the cluster. Reconcile drift continuously (nodes, registries). Run invasive work on a schedule with locks (etcd). Put last success, failure count, and message on the CR so on-call does not grep logs to learn if the cluster healed itself.
- The central control plane can still set policy. It should not defrag etcd or sync registries.
- One CR vs many micro-operators. One CustomResource object is easier to reason about per cluster; it also becomes a fat API. We composed operations under one CR so intent stays in one place.
- Where the work runs. Central control plane execution looked simpler (one binary, one schedule). In-cluster execution costs a workload on every cluster and wins locality, blast-radius isolation, and “no human / no mothership” for the actual job.
- Scheduled vs threshold-triggered etcd defrag. Pure cron defrags when unnecessary; pure threshold can stampede. We combine schedules with a space-quota gate.
- Forceful registry sync vs fail-closed. Syncing when a registry is unhealthy amplifies failure. We health-gate and retry on the loop.
- A design approach: cluster stability after setup is a reconcile problem, not an operations checklist.
- A way to evaluate placement: which Day-2 jobs must run next to etcd, nodes, and registries, and which belong in a central control plane.
- A set of operational practices: quota-gated defrag, verified snapshots, node reconciliation, registry sync and GC, and status that proves the cluster maintained itself.
Production experience (operator and maintenance service running on managed Kubernetes fleets).
#kubernetes #platformengineering #devops #operators #casestudy
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}