Vikash Kumar

When ML Serving Becomes a Distributed Systems Problem

Submitted Oct 1, 2026

Deploying an ML model is easy. Operating hundreds of them reliably is not.

As ML serving moves into production, the difficult problems quickly move beyond inference itself:

  • How do you manage models consistently from creation to retirement?
  • How do you safely activate, scale, and roll out new versions?
  • What happens when replicas disappear or become unhealthy?
  • How do you recover when infrastructure drifts from the desired state?
  • How do you keep Kubernetes complexity away from model owners?

At InMobi, we faced these challenges while evolving our ML serving infrastructure into a Kubernetes-native control plane for model lifecycle management. Instead of asking model owners to manage individual Kubernetes resources, we introduced a higher-level ModelDeployment abstraction that captures the desired state of a model while the platform continuously reconciles that intent with the underlying infrastructure.

This talk walks through that evolution—from deployment automation to a production control plane—and the distributed-systems problems that emerged as the platform grew.

We’ll follow a model through its journey:

  1. Creating a model deployment
  2. Translating the declaration into Kubernetes resources
  3. Activating replicas and validating readiness
  4. Handling failures and recovery
  5. Scaling with changing traffic
  6. Rolling out new model versions
  7. Retiring deployments safely

We’ll then explore the challenges that arise across a large fleet:

  • Designing idempotent reconciliation
  • Handling partial failures, retries, and infrastructure drift
  • Coordinating scaling, availability, and rollouts
  • Isolating models and versions to limit blast radius
  • Keeping infrastructure concerns behind the platform API

We’ll also look at two aspects that are easy to underestimate.

First, the CRD becomes a platform API. Once teams depend on ModelDeployment, its fields and semantics become a long-lived contract. We’ll discuss how to evolve that API as the platform and its infrastructure change while preserving backward compatibility and user intent.

Second, readiness is more than a Kubernetes condition. A model may need to load its artifacts, initialize its runtime, complete warmup, pass health checks, and become reachable through the serving path before it is safe to receive production traffic.

Finally, we’ll share the production lessons behind the platform: where abstractions helped, where they introduced new failure modes, how responsibilities were divided between the control plane and Kubernetes, and why reliable automation is ultimately about designing for failure and recovery.

Key Takeaways

From Kubernetes Resources to Platform APIs

Learn how to design a higher-level model deployment abstraction that captures user intent without exposing every underlying infrastructure primitive.

Managing the Deployment Lifecycle

See how a control plane takes a deployment from creation and activation through scaling, version changes, retirement, and recovery—without exposing the underlying Kubernetes resources.

Reconciliation as the Reliability Mechanism

Understand how continuous reconciliation keeps desired and actual state aligned while handling retries, partial failures, unhealthy workloads, interrupted operations, and infrastructure drift.

Coordinating Scaling, Availability, and Rollouts

Explore how autoscaling, readiness, disruption budgets, replica availability, and version transitions interact—and why coordinating them is important for safe activation, scaling, and rollout.

Serving Isolation

Understand how models and versions can be isolated so that failures, scaling changes, configuration updates, and rollouts have a limited blast radius, while still allowing the platform to coordinate dependencies when required.

Evolving a Production CRD

Learn why a CRD becomes a long-lived platform contract and how API evolution, versioning, and backward compatibility become important as the platform grows.

Designing for Failure

Take away practical lessons for building control planes where failure, recovery, retries, and reconciliation are normal parts of the operating model rather than operational afterthoughts.

Audience

This talk is intended for:

  • Platform Engineers
  • MLOps Engineers
  • SREs
  • Infrastructure Engineers
  • Kubernetes Engineers
  • Software Engineers building distributed systems

It will be particularly relevant to teams building:

  • ML serving platforms
  • Internal developer platforms
  • Kubernetes operators and controllers
  • Model deployment infrastructure
  • Control planes for large fleets of production workloads

#mlops #kubernetes #platformengineering #modelserving #distributedsystems #controlplanes #scalability #sre

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy