Manish Sharma

@manishrma

"Mise en Place" for Kubernetes: Why We Stopped Fetching Images at Runtime

Submitted Oct 8, 2026

One-line summary

Your cluster isn’t failing to pull an image; your release failed to pack it. This talk shows how we turned a fragmented image scavenger hunt into one verifiable bundle for predictable Kubernetes deployments, upgrades even in airgap environments.

What problem are you addressing?

Modern microservices platforms do not ship one binary. They ship a dependency closure: container images, Helm charts, Kubernetes YAML, platform packages, and sometimes node or registry disk images. In our platform, those artifacts had accumulated across multiple repositories and registries, with separate workflows for connected sites, dark sites, deployment, and upgrade.
The result was a familiar but deceptively difficult class of failure: a missing image, an unavailable registry, a mismatched tag, or a late-discovered dependency could stop an upgrade after the customer had already begun mutating the cluster. Manual registry image generation and large air-gap bundles made the system slower and harder to reason about. The platform also had to maintain overlapping ownership and delivery paths for core platform images and service images.
The interesting constraint was that the workflow had to work in both connected and fully air-gapped environments, preserve compatibility across independently released services, and avoid paying a large storage penalty for a safer delivery model.

Intended audience

Platform Engineering, SRE, DevOps, Infrastructure, Kubernetes, release engineering, and engineering leaders responsible for shipping software into environments they do not fully control.

Level

Intermediate to advanced.

What will you share?

  • Architecture or design decisions: Shifting from human-authored image lists to a single, mechanically generated manifest pinned by SHA-256 digests. The generated inventory and bundle contract: typed artifacts, metadata, dependencies, checksums, platform/service profiles, and an OCI image layout.
  • Production experience: How we built a typed, self-describing envelope (the Nutanix Bundle Format) to carry images, charts, YAML, and OS disk images together
  • Failure modes and debugging: Eliminating the “missing image” genre of support cases that plagued our air-gapped customers. fragmented registries, manual publication, late downloads, reverse dependencies, and divergent deploy/upgrade workflows.
  • Customer gains: How it helped reduce the Maintenance window by predownloading the images.
  • Security posture improvements: How shifting to a bundled approach eliminated the need to store and rotate credentials for external registries inside the deployment environment.
  • The lifecycle and ownership choices that let service teams ship their own bundles without each team inventing a new packager.

Experience with this problem

  • Production system
  • Internal engineering project
  • Real failure modes and operational lessons
  • Experiment and benchmark work
  • Hard-earned engineering lesson

What approaches failed, disappointed, or created unexpected problems?

  • Maintaining separate workflows for connected vs. air-gapped sites resulted in drift and missing images
  • Treating a tag or a registry location as the source of truth.
  • Keeping image lists in multiple repositories (external cloud registries) and expecting teams or scripts to keep them synchronized.
  • Pulling images late in deployment or upgrade, when connectivity and credentials were least reliable.
  • Carrying overlapping bundles for platform and service workloads.
  • Treating Helm charts as a separate side channel instead of part of the release dependency closure.

What will you do differently today?

  • Treat the disconnected (air-gapped) case as the primary design constraint.
  • Connected environments should just be a simplified subset of the air-gapped workflow, not a separate path.
  • Generate a consolidated, typed inventory from the deployment specifications and explicitly manage only the small set of include/exempt policy decisions.
  • Pin content by digest and verify the archive, each member, and the nested OCI content before applying anything.
  • Generate Nutanix Qualified bundles - Package platform and service artifacts in one self-describing envelope, with metadata selecting the appropriate apply graph.
  • Convert images and Helm charts into a standard OCI layout so upgrades can skip unchanged layers.
  • Validate required image names and tags before mutating the cluster, with a dry-run option.
  • Model long-running transfers as idempotent, retryable, resumable tasks.
  • Keep platform and service release cadence independent through explicit compatibility and depen

What trade-offs did you consider?

  • One envelope versus many specialized packages: one envelope reduces consumer and integration complexity, but requires a clear typed metadata contract and strict compatibility rules.
  • Build-time completeness versus runtime flexibility: fetching everything at build time increases release-pipeline work and artifact size, but removes runtime dependence on external registries and makes air-gapped operation deterministic.
  • Platform atomicity versus service composability: platform bundles must be version-gated and treated as atomic releases; service bundles need independent ownership and the ability to compose.
  • Tags versus digests: tags remain useful for human-facing metadata, but digests are required for reproducibility and verification.
  • Migration versus a clean break: supporting old and new packaging for a bounded compatibility window reduces adoption risk, at the cost of temporary complexity.

How can this help other practitioners?

The talk offers a reusable way to inspect any release pipeline:

  • A pattern to adopt: The “Mise en place” approach of resolving and bundling all dependencies (images, charts) at build-time rather than runtime
  • A design approach: The architecture of the Nutanix Bundle Format as a way to unify multi-artifact delivery.
  • Eliminating code churn by anchoring to the strictest constraint: By inverting the problem and designing natively for the airgap model, you strip away the complexity of handling multiple remote registries, fragmented download stages, and fallback logic.

The broader pattern is to trade recurring, variable operational cost for a small, explicit contract: one generated inventory, one typed envelope, one content-addressed payload, and one apply workflow.

Current state

Production experience and retrospective lessons from an internal engineering project.

Tags

#platformengineering #kubernetes #devops #sre #infrastructure #reliability #airgap #supplychain #oci #helm #failurestory #casestudy #security #scalability #automation #Nutanix

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy