Guda Sai Vikhyath Reddy

@vikhyath

Engineering Beyond the Broker: Solving Operational Challenges Without Replacing Kafka

Submitted Oct 7, 2026

Proposal template

Session title
Kafka: The Infrastructure Around the Infrastructure

One-line summary
Evaluating Kafka ACLs, network boundary controls, and protocol-aware gateways (Kroxylicious) to tackle multi-team operational complexity at scale.

What problem are you addressing?
Kafka scales core messaging reliably, but as deployments expand across dozens of application teams and network boundaries, platform operations become complex. Requirements like IP whitelisting per broker, missing centralized audit visibility, and raw Kafka ACLs create operational friction and security risks when evolving infrastructure or onboarding new clients.

This session breaks down how to address these operational bottlenecks in production environments. We analyze where native Kafka capabilities (ACLs, broker networking) are sufficient, where introducing an intermediary layer like Kroxylicious simplifies governance, and the architectural trade-offs between native control versus proxy-based abstraction.

The focus is not on presenting a single solution, but on the engineering reasoning involved in deciding where a problem should be solved and what complexity each approach introduces.

Who is the intended audience?
Platform Engineering / SRE / DevOps / Infrastructure / Software Developers/Engineers

Level:
Intermediate/Advanced

List one or two practical takeaways.

  • A decision framework for evaluating Kafka security, access control, and governance against architectural alternatives.
  • Concrete criteria for identifying when Kafka-native mechanisms are sufficient versus when a protocol-aware gateway adds high-value operational capabilities.

What will you share?

  • Architecture decisions
  • Failure modes and operational challenges
  • Trade-offs and alternatives
  • Live demonstration

What is your experience with this problem?

  • Production system
  • Real incident or failure

What approaches failed, disappointed, or created unexpected problems?

Relying solely on broker IP whitelisting and direct broker connectivity created operational bottlenecks whenever clusters were expanded or brokers were replaced. Additionally, enabling Kafka ACLs retroactively without detailed knowledge of topic ownership and active consumers posed a significant risk of breaking existing applications.

We examined the practical limitations of relying exclusively on broker-level networking and standard ACLs as client counts scale, challenges around auditing, client-side reconfiguration, and policy enforcement across heterogeneous teams.

Rather than treating these approaches as failures, we examined where they stop being sufficient for a particular operational requirement and what additional complexity alternative approaches introduce.

What will you do differently today?

  • We would evaluate Kafka operational requirements earlier and separate the concerns of message brokering from the operational layer around the broker.
  • First, establish what can be handled cleanly through Kafka itself, networking, and existing infrastructure. Where those mechanisms become difficult to manage consistently, we would then showcase how a protocol-aware proxy or another layer provides enough value to justify its operational cost.

What trade-offs did you consider?

  • Kafka ACLs and native security mechanisms
  • Network-level access controls and client isolation
  • Maintaining a new component/service (i.e the proxy)

How can this help other practitioners?

  • Identify which concerns belong inside Kafka and which may belong in an operational layer.
  • Understand when to consider Kafka-native capabilities before adding infrastructure.
  • Understand the operational cost of introducing a protocol-aware proxy.
  • Think through migration and failure scenarios before adopting an additional component.
  • Compare competing approaches based on their actual operational requirements rather than technology preference.

Current state
Prototype

Tags
#kafka #platformengineering #infrastructure #devops #sre #security #distributed-systems #kubernetes #networking #observability #architecture #failurestory #demo #casestudy #workinprogress #kroxylicious

(This will be a talk with 2 speakers)

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy