Your cluster isn’t failing to pull an image; your release failed to pack it. This talk shows how we turned a fragmented image scavenger hunt into one verifiable bundle for predictable Kubernetes deployments, upgrades even in airgap environments.
Modern microservices platforms do not ship one binary. They ship a dependency closure: container images, Helm charts, Kubernetes YAML, platform packages, and sometimes node or registry disk images. In our platform, those artifacts had accumulated across multiple repositories and registries, with separate workflows for connected sites, dark sites, deployment, and upgrade.
The result was a familiar but deceptively difficult class of failure: a missing image, an unavailable registry, a mismatched tag, or a late-discovered dependency could stop an upgrade after the customer had already begun mutating the cluster. Manual registry image generation and large air-gap bundles made the system slower and harder to reason about. The platform also had to maintain overlapping ownership and delivery paths for core platform images and service images.
The interesting constraint was that the workflow had to work in both connected and fully air-gapped environments, preserve compatibility across independently released services, and avoid paying a large storage penalty for a safer delivery model.
Platform Engineering, SRE, DevOps, Infrastructure, Kubernetes, release engineering, and engineering leaders responsible for shipping software into environments they do not fully control.
Intermediate to advanced.
- Architecture or design decisions: Shifting from human-authored image lists to a single, mechanically generated manifest pinned by SHA-256 digests. The generated inventory and bundle contract: typed artifacts, metadata, dependencies, checksums, platform/service profiles, and an OCI image layout.
- Production experience: How we built a typed, self-describing envelope (the Nutanix Bundle Format) to carry images, charts, YAML, and OS disk images together
- Failure modes and debugging: Eliminating the “missing image” genre of support cases that plagued our air-gapped customers. fragmented registries, manual publication, late downloads, reverse dependencies, and divergent deploy/upgrade workflows.
- Customer gains: How it helped reduce the Maintenance window by predownloading the images.
- Security posture improvements: How shifting to a bundled approach eliminated the need to store and rotate credentials for external registries inside the deployment environment.
- The lifecycle and ownership choices that let service teams ship their own bundles without each team inventing a new packager.
- Production system
- Internal engineering project
- Real failure modes and operational lessons
- Experiment and benchmark work
- Hard-earned engineering lesson
- Maintaining separate workflows for connected vs. air-gapped sites resulted in drift and missing images
- Treating a tag or a registry location as the source of truth.
- Keeping image lists in multiple repositories (external cloud registries) and expecting teams or scripts to keep them synchronized.
- Pulling images late in deployment or upgrade, when connectivity and credentials were least reliable.
- Carrying overlapping bundles for platform and service workloads.
- Treating Helm charts as a separate side channel instead of part of the release dependency closure.
- Treat the disconnected (air-gapped) case as the primary design constraint.
- Connected environments should just be a simplified subset of the air-gapped workflow, not a separate path.
- Generate a consolidated, typed inventory from the deployment specifications and explicitly manage only the small set of include/exempt policy decisions.
- Pin content by digest and verify the archive, each member, and the nested OCI content before applying anything.
- Generate Nutanix Qualified bundles - Package platform and service artifacts in one self-describing envelope, with metadata selecting the appropriate apply graph.
- Convert images and Helm charts into a standard OCI layout so upgrades can skip unchanged layers.
- Validate required image names and tags before mutating the cluster, with a dry-run option.
- Model long-running transfers as idempotent, retryable, resumable tasks.
- Keep platform and service release cadence independent through explicit compatibility and depen
- One envelope versus many specialized packages: one envelope reduces consumer and integration complexity, but requires a clear typed metadata contract and strict compatibility rules.
- Build-time completeness versus runtime flexibility: fetching everything at build time increases release-pipeline work and artifact size, but removes runtime dependence on external registries and makes air-gapped operation deterministic.
- Platform atomicity versus service composability: platform bundles must be version-gated and treated as atomic releases; service bundles need independent ownership and the ability to compose.
- Tags versus digests: tags remain useful for human-facing metadata, but digests are required for reproducibility and verification.
- Migration versus a clean break: supporting old and new packaging for a bounded compatibility window reduces adoption risk, at the cost of temporary complexity.
The talk offers a reusable way to inspect any release pipeline:
- A pattern to adopt: The “Mise en place” approach of resolving and bundling all dependencies (images, charts) at build-time rather than runtime
- A design approach: The architecture of the Nutanix Bundle Format as a way to unify multi-artifact delivery.
- Eliminating code churn by anchoring to the strictest constraint: By inverting the problem and designing natively for the airgap model, you strip away the complexity of handling multiple remote registries, fragmented download stages, and fallback logic.
The broader pattern is to trade recurring, variable operational cost for a small, explicit contract: one generated inventory, one typed envelope, one content-addressed payload, and one apply workflow.
Production experience and retrospective lessons from an internal engineering project.
#platformengineering #kubernetes #devops #sre #infrastructure #reliability #airgap #supplychain #oci #helm #failurestory #casestudy #security #scalability #automation #Nutanix
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}