Sreeram Venkitesh

Sreeram Venkitesh

@fillerink

OTel Memory Management 101: Lessons learnt while running OpenTelemetry scraping 20k clusters in production

Submitted Oct 8, 2026

OTel Memory Management 101: Lessons learnt while running OpenTelemetry scraping 20k clusters in production

Session title

OTel Memory Management 101: Lessons learnt while running OpenTelemetry scraping 20k clusters in production

One-line summary

This talk goes over some of the challenges we faced when running OpenTelemetry to scrape metrics and logs from about 20k customer clusters as part of DigitalOcean’s managed Kubernetes platform.

What problem are you addressing?

Getting up and running with an observability stack today is simple. But doing it right, and more importantly knowing if you’ve done it right, is considerably more difficult. The DigitalOcean Managed Kubernetes platform migrated to using OpenTelemetry for scraping metrics and logs of customer Kubernetes clusters from our legacy stack that used Prometheus and Fluent Bit. This talk goes over how we did this and some of the challenges we faced when scaling OpenTelemetry to scrape metrics and logs from around 20,000 customer clusters that we manage in production.

The main challenge I’ll be presenting in this talk was how we optimized the memory usage of OpenTelemetry collectors. One set of OpenTelemetry collector pods, which were managed by DaemonSet on our infrastructure nodes, to collect container logs from the control planes of customer clusters were frequently getting OOMKilled. We scaled the memory to no avail since the memory usage would just scale to the newly increased limit each time.

We investigated this issue and fixed it at the root by tweaking the different parameters of the OpenTelemetry collector configuration to ensure that garbage collection takes place properly and that the memory usage is not going haywire as the collector starts collecting logs from all the containers. We used Go’s GOMEMLIMIT along with the OTel ollector’s built-in memory_limiter processor to do this. This talk explains what each mechanism controls: memory_limiter enforcing a hard limit by refusing pipeline data, and GOMEMLIMIT steering the Go garbage collector’s pacing. Attendees will learn why and how these two must be set in a specific ratio relative to the container’s memory limit to optimize the memory.

Who is the intended audience?

Platform engineers, SRE, DevOps and folks generally into infrastructure.

Level

Intermediate

List one or two practical takeaways

Attendees will get to:

  • Learn the architecture we adopted when migrating an existing production cluster platform with real users to OpenTelemetry
  • Learn the different challenges we faced when scaling OpenTelemetry
  • Learn the best practices to follow when managing OpenTelemetry in production

What will you share?

I’ll be sharing our architecture and the design decisions we took when setting up OpenTelemetry from scratch, the production experience we had when running such a platform, the issues we faced and how we fixed it. To back this up I’ll be showing Grafana dashboards showing the memory usage and pod OOMKills before and after the issue was fixed.

What is your experience with this problem?

I have been involved in building the observability stack at DigitalOcean’s Managed Kubernetes team. As part of this I’ve come up with the design and architecture and deployed OTel collectors in production and worked on optimizing the collectors and making them stable on a production environment.

What approaches failed, disappointed, or created unexpected problems?

Trying to scale the OTel collector DaemonSet directly via VPA etc was ineffective. Given the load, the memory usage just scaled as we increased the memory limits. We were not fixing the issue at its source (by tweaking the collector’s garbage collection behavior)

What will you do differently today?

I’ll be mindful of how easily memory can be used (or misused :D) and optimize my systems in advance for handling such scenarios better.

What trade-offs did you consider?

The tradeoff here is of cost incurred by both compute and memory. Optimizing the collector meant we could run our systems reliably with lesser CPU and memory resulting in cost savings. The tradeoff was that we could throw money at the problem and scale the collectors vertically. Understanding how the systems work at a deeper level let us not have to do this.

How can this help other practitioners?

Folks using OpenTelemetry will get to learn the best practices around memory management which they might not come across when first setting up the collector. The talk would also give them a good idea of how to architect and design their own OpenTelemetry stack for their platforms.

Current state

Currently our OpenTelemetry collectors are handling the load much better and are not getting OOMKilled like before.

#observability #opentelemetry #production #kubernetes #developerplatform #failurestory #demo #casestudy #platformengineering

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy