Adithya Parthasarthy

3 Weeks to 3 Hours: Building an AI Co-Pilot for Kubernetes Upgrade Preparation

Submitted Oct 8, 2026

3 Weeks to 3 Hours: Building an AI Co-Pilot for Kubernetes Upgrade Preparation at Booking.com

One-line summary

Every Kubernetes version upgrade costs our platform team 3–4 weeks of manual preparation — reading changelogs, chasing compatibility matrices, and diffing client-go commits across 50+ repositories and 1000+ services. We built an AI-assisted intelligence system at Booking.com that does the analysis in hours, validated it against a real upgrade, and it caught everything the team found manually — plus issues they’d missed.


What problem are you addressing?

Kubernetes upgrades are an invisible tax on every platform team. At Booking.com, we operate one of the largest Kubernetes platforms in the travel industry. Before a single cluster is touched, engineers spend weeks doing something that looks suspiciously like reading comprehension homework — at massive scale.

The numbers tell the story:

  • 3–4 weeks of senior engineering time consumed per Kubernetes version upgrade — just for preparation
  • Hundreds of upstream changes per Kubernetes release to manually triage
  • 50+ internal repositories to cross-reference against every change
  • Multiple custom controllers and operators with client-go dependencies to analyze
  • Dozens of ecosystem dependencies (Karpenter, Argo CD, CoreDNS, VictoriaMetrics, cert-manager, etcd, NVIDIA GPU components, and more) — each with their own compatibility matrix scattered across separate upstream repos
  • ~3 Kubernetes releases per year, meaning this 3–4 week tax repeats multiple times annually
  • Effort scales worse than linearly — every new controller, dependency, or cluster added to the platform makes the next upgrade harder

What makes this genuinely hard:

  • Client-go provides no upgrade guide. Engineers literally diff commit histories across releases to find what changed in the specific subset of packages they use. At Booking.com’s scale, this alone can take days.
  • Compatibility information has no single source of truth. It’s scattered across dozens of upstream repos, release notes, GitHub issues, and documentation sites.
  • The process depends on institutional knowledge that lives in people’s heads, not in systems. When those people are unavailable, upgrades stall.

We’ve felt the consequences of missing things. At Booking.com, a cert-manager incompatibility broke customer SRM workloads. An etcd change killed Flink jobs in production. An AMI change broke Terraform provisioning across clusters. Every one of these was discoverable in advance — if anyone had the bandwidth to look.

We built Kubernetes Upgrade Intelligence: a three-track AI-assisted system that reads upstream changes, understands our platform’s specific usage patterns across all 50+ repositories, identifies what actually matters to Booking.com, and generates remediation — while keeping engineers in control of every decision.

We validated it against a real Kubernetes 1.33 upgrade at Booking.com. It found everything the team found manually, plus additional issues they’d missed. It has been used for 1.34 and 1.35 upgrade cycles successfully.


Intended audience

Platform Engineering / SRE / Infrastructure / DevOps


Level

Intermediate


Practical takeaways

  1. A reusable architecture pattern for combining deterministic code analysis (AST parsing, dependency graphs, Sourcegraph queries across 50+ repos) with LLM reasoning — so you get AI that’s grounded in facts, not hallucinations. Tested at Booking.com scale.

  2. A concrete playbook for building AI-assisted upgrade preparation for your own Kubernetes platform — including what to feed the model, what to keep deterministic, where humans must stay in the loop, and how to measure if the AI is actually helping.


What will you share?

  • Architecture and design decisions — how the three-track system works at Booking.com
  • Production experience — real results from the Kubernetes 1.33 upgrade cycle
  • Before-and-after comparison — manual vs. AI-assisted upgrade prep (weeks vs. hours)
  • Failure modes and debugging — what the AI gets wrong and how we catch it
  • Trade-offs and alternatives — why we made the design choices we did
  • Live demo — the changelog intelligence pipeline analyzing real Kubernetes changes against a real codebase

Experience with this problem

  • Production system at Booking.com — one of the world’s largest online travel platforms
  • Real incidents and failures — customer-facing breakages from missed upgrade risks
  • Internal engineering project — approved through Booking.com’s Gen AI intake process
  • Hard-earned engineering lesson — built from years of painful manual upgrade cycles (versions 1.25 through 1.31, each with its own preparation document)

What approaches failed, disappointed, or created unexpected problems?

Naive LLM approach: Feeding full Kubernetes changelogs to an LLM produced plausible-sounding but unreliable results. The model would confidently flag irrelevant changes and miss critical ones buried in dense release notes. At Booking.com’s scale — where a false positive wastes senior engineering time across multiple teams and a false negative can cause production incidents — this was unacceptable. Pure LLM inference without grounding in our actual codebase was essentially a hallucination generator with a professional tone.

Fully deterministic approach: Static analysis, regex-based changelog parsing, and version constraint solvers were precise but brittle. They caught explicit API removals but missed behavioral changes, subtle deprecation paths, and the kind of cross-cutting impact that requires understanding context across 50+ repos, not just matching strings.

Brute-force client-go diffing: Early attempts at automated diffing produced thousands of changes with no signal about which ones mattered to our actual usage patterns. Engineers ended up doing more triage work than the manual process required.


What will you do differently today?

We now treat deterministic analysis and AI reasoning as complementary stages, not competing approaches.

Deterministic tooling establishes facts first — which APIs we use across all repositories, which manifests reference deprecated resources, which dependency versions we run, which client-go packages are actually imported. AI then interprets those facts in the context of upstream changes to assess impact, prioritize risks, and draft remediation.

This “facts first, reasoning second” architecture — validated at Booking.com against real Kubernetes upgrades — dramatically reduced false positives and eliminated the most dangerous category of hallucinations.


Trade-offs considered

Decision What we chose What we gave up Why
Analysis approach Hybrid: deterministic fact-finding feeds AI reasoning Pure AI (faster to build) or pure static analysis (more precise) At Booking.com’s scale, we needed both coverage and accuracy
Automation level Human-in-the-loop at every decision point Fully autonomous upgrades Platform upgrades are safety-critical — the cost of a wrong automated decision far exceeds a human review step
AI input sources Official changelogs, EKS release notes, upstream Git history only Broader web knowledge, blog posts, Stack Overflow We’d rather miss an edge case than act on a hallucination in production
System architecture Three specialized tracks (changelog, client-go, dependencies) Single monolithic pipeline Each track has fundamentally different source material, analysis patterns, and failure modes
Optimization target Signal-to-noise ratio and accuracy over speed Maximum automation Engineers trust the system because its findings are consistently actionable

How can this help other practitioners?

  • A design pattern to adopt: The “deterministic grounding + AI reasoning” architecture applies to any domain where AI needs to make recommendations about complex technical systems — not just Kubernetes upgrades
  • A mistake to avoid: Don’t point an LLM at changelogs and trust the output. We tried it at Booking.com scale. It doesn’t work. Ground everything in your actual codebase first.
  • A way to evaluate competing approaches: Framework for deciding what to keep deterministic vs. what to delegate to AI in safety-critical engineering workflows
  • Operational metrics that matter: How to measure AI-assisted engineering quality (signal-to-noise ratio, false positive rate, inference accuracy, escaped issues) so you know if it’s actually helping or just generating confident noise
  • A new way of thinking about the problem: Upgrade preparation isn’t a task to automate end-to-end — it’s an intelligence problem. The right framing changes what you build.

Current state

Production experience — validated against Kubernetes 1.33, 1.34, 1.35 upgrade cycle at Booking.com. The changelog intelligence track found all manually-identified upgrade considerations plus additional findings. Client-go intelligence has been successfully applied to real code changes. Dependency compatibility intelligence flagged the right upgrade ordering for dependencies.


Tags

#kubernetes #platformengineering #ai #agents #sre #upgrades #casestudy #devops #scalability #infrastructure #booking

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy