Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Adithya Parthasarthy
Submitted Oct 8, 2026
Every Kubernetes version upgrade costs our platform team 3–4 weeks of manual preparation — reading changelogs, chasing compatibility matrices, and diffing client-go commits across 50+ repositories and 1000+ services. We built an AI-assisted intelligence system at Booking.com that does the analysis in hours, validated it against a real upgrade, and it caught everything the team found manually — plus issues they’d missed.
Kubernetes upgrades are an invisible tax on every platform team. At Booking.com, we operate one of the largest Kubernetes platforms in the travel industry. Before a single cluster is touched, engineers spend weeks doing something that looks suspiciously like reading comprehension homework — at massive scale.
The numbers tell the story:
What makes this genuinely hard:
We’ve felt the consequences of missing things. At Booking.com, a cert-manager incompatibility broke customer SRM workloads. An etcd change killed Flink jobs in production. An AMI change broke Terraform provisioning across clusters. Every one of these was discoverable in advance — if anyone had the bandwidth to look.
We built Kubernetes Upgrade Intelligence: a three-track AI-assisted system that reads upstream changes, understands our platform’s specific usage patterns across all 50+ repositories, identifies what actually matters to Booking.com, and generates remediation — while keeping engineers in control of every decision.
We validated it against a real Kubernetes 1.33 upgrade at Booking.com. It found everything the team found manually, plus additional issues they’d missed. It has been used for 1.34 and 1.35 upgrade cycles successfully.
Platform Engineering / SRE / Infrastructure / DevOps
Intermediate
A reusable architecture pattern for combining deterministic code analysis (AST parsing, dependency graphs, Sourcegraph queries across 50+ repos) with LLM reasoning — so you get AI that’s grounded in facts, not hallucinations. Tested at Booking.com scale.
A concrete playbook for building AI-assisted upgrade preparation for your own Kubernetes platform — including what to feed the model, what to keep deterministic, where humans must stay in the loop, and how to measure if the AI is actually helping.
Naive LLM approach: Feeding full Kubernetes changelogs to an LLM produced plausible-sounding but unreliable results. The model would confidently flag irrelevant changes and miss critical ones buried in dense release notes. At Booking.com’s scale — where a false positive wastes senior engineering time across multiple teams and a false negative can cause production incidents — this was unacceptable. Pure LLM inference without grounding in our actual codebase was essentially a hallucination generator with a professional tone.
Fully deterministic approach: Static analysis, regex-based changelog parsing, and version constraint solvers were precise but brittle. They caught explicit API removals but missed behavioral changes, subtle deprecation paths, and the kind of cross-cutting impact that requires understanding context across 50+ repos, not just matching strings.
Brute-force client-go diffing: Early attempts at automated diffing produced thousands of changes with no signal about which ones mattered to our actual usage patterns. Engineers ended up doing more triage work than the manual process required.
We now treat deterministic analysis and AI reasoning as complementary stages, not competing approaches.
Deterministic tooling establishes facts first — which APIs we use across all repositories, which manifests reference deprecated resources, which dependency versions we run, which client-go packages are actually imported. AI then interprets those facts in the context of upstream changes to assess impact, prioritize risks, and draft remediation.
This “facts first, reasoning second” architecture — validated at Booking.com against real Kubernetes upgrades — dramatically reduced false positives and eliminated the most dangerous category of hallucinations.
| Decision | What we chose | What we gave up | Why |
|---|---|---|---|
| Analysis approach | Hybrid: deterministic fact-finding feeds AI reasoning | Pure AI (faster to build) or pure static analysis (more precise) | At Booking.com’s scale, we needed both coverage and accuracy |
| Automation level | Human-in-the-loop at every decision point | Fully autonomous upgrades | Platform upgrades are safety-critical — the cost of a wrong automated decision far exceeds a human review step |
| AI input sources | Official changelogs, EKS release notes, upstream Git history only | Broader web knowledge, blog posts, Stack Overflow | We’d rather miss an edge case than act on a hallucination in production |
| System architecture | Three specialized tracks (changelog, client-go, dependencies) | Single monolithic pipeline | Each track has fundamentally different source material, analysis patterns, and failure modes |
| Optimization target | Signal-to-noise ratio and accuracy over speed | Maximum automation | Engineers trust the system because its findings are consistently actionable |
Production experience — validated against Kubernetes 1.33, 1.34, 1.35 upgrade cycle at Booking.com. The changelog intelligence track found all manually-identified upgrade considerations plus additional findings. Client-go intelligence has been successfully applied to real code changes. Dependency compatibility intelligence flagged the right upgrade ordering for dependencies.
#kubernetes #platformengineering #ai #agents #sre #upgrades #casestudy #devops #scalability #infrastructure #booking
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}