Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Silu Panda
Submitted Oct 8, 2026
A public-code case study of a LiteLLM Redis-cluster startup failure: why synchronous cluster construction needed credentials before a later IAM connection callback could supply them, and how to test the real authentication handshake.
30-minute remote technical deep dive, with a shorter demo format possible if preferred by the editors. Remote participation remains subject to editorial selection, availability, and necessary approval; this is not a commitment to travel or a confirmed presentation slot.
An AI gateway can fail before serving its first request because a dependency authenticates at a different point in its connection lifecycle than the application expects. This session would examine a concrete LiteLLM Redis-cluster startup failure: synchronous cluster construction needed credentials before a later IAM connection callback could supply them. Working asynchronous authentication did not establish that synchronous bootstrap would work.
Using my contribution to LiteLLM PR #40204, I would trace the connection ordering and the provider-based fix: supplying an existing credential provider during cluster construction while removing conflicting authentication arguments. This is a public-code reliability case study, not a proprietary LinkedIn incident, an authentication-bypass claim, or a product pitch. I contributed the initial fix and regression tests, and would credit the original reporter’s diagnosis and the maintainers’ subsequent revisions and validation.
Platform engineers, SREs, backend engineers, and AI-infrastructure practitioners operating gateways or services with authenticated dependencies. Basic Python, Redis, and cloud IAM familiarity would help; LiteLLM-specific knowledge would not be assumed.
Attendees would learn to trace authentication from startup through client construction, test the actual handshake without cloud credentials in CI, verify that CI selects the intended regression tests, and distinguish constructor correctness, local integration evidence, cloud validation, and production-rollout claims.
The fix was merged upstream on 23 September 2026; the conference walkthrough/demo is proposed, not already prepared. The final PR documents live Azure Managed Redis validation with a service principal, not a managed identity; I would describe this as the final PR’s validation record, not my own live deployment. GCP validation used local Redis with token issuance stubbed, not live Memorystore. The case does not establish multi-node failover behavior, long-lived token-expiry handling, or rollout to the original reporter’s environments.
#platformengineering #mlops #sre #security #databases #redis #iam #integrationtesting #casestudy
Silu Panda is a Staff Software Engineer in ML Infrastructure at LinkedIn, based in Sunnyvale, California, focusing on model deployment and inference reliability.
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}