Shaurya Madukuri

@shauryamadukuri

We spent a week optimising an LLM pipeline and most of what we measured was noise

Submitted Sep 18, 2026

One-line summary

A 12,858-image vision pipeline took 1h51m, so we tested three optimisations
against it. Our first measurements said none of them worked, because
run-to-run variance was 43%. Once we controlled the noise, two were real
speedups (up to 7x from URI image transport) and one traded latency for cost.
Here’s how we told them apart, and why the rest of the wins came from reading
code rather than benchmarking it.


What problem are you addressing?

Anyone putting an LLM pipeline into production hits the same wall: it is too
slow, there are obvious things to try, and none of the usual profiling instincts
work. The provider’s latency moves 2x over minutes. Your effect size is 10%.
Your variance is 43%. Every benchmark you run is a coin flip you mistake for a
result.

We ran the experiment properly and it taught us more about measurement than
about latency:

  • Sweeping concurrency 100 -> 200 -> 400 -> 800 produced a clean monotonic
    win that was entirely drift.
    On my own earlier run the drift ran the other
    way and hid a real win. Bracketing every level against repeats of a fixed
    baseline found the truth: 600 in flight is about 2x 200 on the relabelling
    stage, and nothing past 600 helps.
  • A prompt-caching change raised cache hit 36.5% -> 42% and made the pipeline
    7% slower.
    Reproducible in both directions. At the cached-token discount
    that is roughly 7% off the input-token bill, so it is a cost win and a latency
    loss, the opposite of what it was built for. We kept it for the money.
  • URI image transport first measured as doing nothing, then as up to 7x. On
    the route we benchmarked first it removed 20ms/image of client work, and we
    nearly wrote it off. Measured properly it took one stage from 370 s to 53 s,
    and halved screen discovery end to end.

The real wins came from auditing, not benchmarking:

  • A proxy-only field in the request body caused 1,402 consecutive HTTP 400s
    and zero labels before anyone noticed. Retries cannot help a malformed
    request, so the retry budget made it slower to fail.
  • An uncapped model output let one repetition loop generate 458,752
    characters and stall its entire batch, because every stage gathers
    concurrently and one slow call sets the floor for all of it.
  • Synchronous prompt-file reads were 75% of the process’s own CPU and
    blocked the event loop for all 200 in-flight calls — 8,000 disk reads where 4
    would do.
  • Why URI transport first looked like nothing. Reading the routing config
    showed the pipeline calling Google directly, bypassing our LiteLLM proxy.
    Through the proxy, URI transport with signed URLs took the same 2,811 images
    from 8.7 to 4.2 minutes. The real bottleneck was the laptop’s uplink, pushing
    ~5 GB of base64 per dataset, and no benchmark on the direct route could have
    shown it.

End to end, the same 2,811 images went from 13.4 minutes to 4.2, with zero
errors. That projects to about 15 minutes for a 10,000-image dataset that took
137 before. Route, model and transport all changed in that number; the 8.7 -> 4.2
step is the transport alone.

Thesis: when run-to-run variance exceeds your effect size, stop benchmarking
and start reading code. Defects are measurable at n=1; optimisations are not. And a null result only
holds for the route you measured it on.


Who is this for?

Engineers running inference pipelines in production, and more generally anyone
whose system depends on a service whose latency is noisier than the change they
are trying to make. No ML background needed — the subject is an LLM pipeline,
the content is measurement discipline.


What will attendees take away?

  1. Bracket, never sweep. Why interleaving every trial against a fixed
    baseline is the only defence against drift, with the two runs that prove it
    (one drifted up, one drifted down, both looked like clean results).
  2. Capture the metric that explains the mechanism, not just the clock. A
    2.8x gap between two routes was blamed on the wrong cause twice from
    wall-clock alone; cache hit rate and in-flight concurrency showed it was
    mostly throttling.
  3. Where asyncio.gather turns one slow request into a whole-stage stall,
    and why bounding output length fixes it better than bounding time.
  4. A short audit checklist that found three defects and one misrouted call
    that no benchmark would have surfaced.

Rough structure

Time Section
5 min The pipeline, the 1h51m baseline, the three things we planned to try
8 min All three results: what the first pass said, and what they really did once measured properly
7 min Why we could not trust any of it: 43% variance and drift
10 min What actually worked: three defects found by reading code (1,402 consecutive HTTP 400s, a 458,752-character runaway output, synchronous prompt reads), and the written-off URI transport that halved runtime once routed through the proxy
5 min Bracket, never sweep: the protocol, the audit checklist, and the before/after

Speaker bio

Shaurya Madukuri. Data Scientist II at Skan.AI, where I lead the team building a
computer-use agent across more than twelve BFSI desktop workflows. B.Tech in
Electrical Engineering from IIT Delhi.

Most of my work is the unglamorous half of shipping models: making them fit,
making them fast, and finding out why they broke. The pipeline in this talk is
one of those. I also maintain forge-kernels (github.com/Shaurya-M002), an
open-source Triton kernel library for memory-efficient LLM training, and
published on glaucoma classification at IEEE ISBI 2024.

Based in Bengaluru.

LinkedIn: https://www.linkedin.com/in/shaurya-madukuri/ · X: https://x.com/shaurya_mk · GitHub: https://github.com/Shaurya-M002

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy