Jitendra

@jeetu12400

Upgrade Broke My Query… Break Free with AI [ Detect, Debug, Fix with AI-Assisted Self-Service Platform ]

Submitted Oct 8, 2026

Session title

Upgrade Broke My Query… Break Free with AI
Detect, Debug, Fix with AI-Assisted Self-Service Platform

One-line summary

In this session, we will explore why and how data platform teams can build a self-service performance analysis platform – automating upgrade testing, regression detection, and debugging to reduce platform-team intervention and help application teams identify and resolve performance issues independently.

What problem are you addressing?

As a small central data platform team, our core mandate is to ensure the databases we support run seamlessly across all application teams. Fulfilling this exposed a major bottleneck during ClickHouse version upgrades: to identify and debug query regressions, we had to simulate each team’s workload in our lab and deeply understand their unique business logic. Recognizing this hands-on approach was completely unscalable, we set out to build a self-serve platform that empowers application teams to run their own workloads and identify regressions independently of the platform team.

We named this fully self-serve platform Velocity. Velocity expanded well beyond its original scope, not only does it run workloads against multiple ClickHouse versions to pinpoint degraded queries, but it also provides a sandbox for application teams to proactively test their own query optimizations independently without relying on our central team.

Yet, identifying a degraded query is only half the job. To truly solve the problem, teams needed the root cause analysis (RCA) and specific remedial steps to fix performance drops which again needed platform team involvement. To maintain our goal of a fully self-serve ecosystem, the platform evolved once more. We built Sherlock directly on top of Velocity – an AI-assisted query regression investigation system that automatically generates the RCA and provides the exact steps needed to get query performance back on track.

ClickHouse evolves rapidly, upgrades can make a set of queries unexpectedly slower without causing any functional failures. These regressions are often discovered late in the release cycle, when fixing them becomes expensive.
Our initial approach relied heavily on a central platform team to run tests and investigate regressions, which did not scale as more application teams and workloads were added.
This led to a key question: Can performance testing and regression investigation become a self-service capability for application teams?

Velocity was our answer.

The platform provides a common execution environment where application teams can run their workloads against ClickHouse versions, compare results, identify regressions, and investigate the reasons behind them.
We then faced a second problem: detecting a regression is only half the job. An application team needs to know why the query regressed and what they can do about it.
That led to Sherlock — an AI-assisted query regression investigation system built on top of Velocity.

Who is the intended audience?

Platform Engineering / Infrastructure / Developers / Data & Backend Engineers / Engineering leaders

The talk is particularly relevant for teams that operate shared data platforms and need to support multiple application teams with different workloads.

Level

intermediate
The talk assumes familiarity with databases, performance testing, and basic platform concepts, but does not require deep ClickHouse internals or AI/LLM expertise.

Practical takeaways

  • A framework for going beyond regression detection: combine workload execution, query telemetry, database-version changes, best practices, settings, and AI-assisted investigation to turn “this query got slower” into “this is likely why, and here is what you can try.”

What will you share?

Code or implementation details

What is your experience with this problem?

Real incident or failure

What will you do differently today?

We initially built Velocity and Sherlock specifically for ClickHouse performance regression analysis. As we developed the platform, we realized the same approach could be generalized beyond ClickHouse to other databases, streaming systems, and data services.
Today, we would design the platform with this broader scope from the start—keeping the execution, regression detection, context collection, and AI-assisted analysis extensible across different data systems.

What trade-offs did you consider?

We balanced centralized control vs. self-service and chose self-service for better scalability, accepting the added platform complexity.
We also chose diagnosis over simple benchmarking—building an RCA AI agent to explain regressions rather than just report numbers.
For AI analysis, we chose specialized agents over one general-purpose agent to provide more focused and actionable recommendations.

How can this help other practitioners?

A new way of thinking about the problem

Current state

Production experience

#clickhouse #databases #performanceengineering #platformengineering #developerplatforms #selfservice #kubernetes #ai #agents #llm #observability #devops #infrastructure #scalability #regressiontesting #benchmarking #casestudy #workinprogress

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy