Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Shaurya Madukuri
@shauryamadukuri
Submitted Sep 18, 2026
A 12,858-image vision pipeline took 1h51m, so we tested three optimisations
against it. Our first measurements said none of them worked, because
run-to-run variance was 43%. Once we controlled the noise, two were real
speedups (up to 7x from URI image transport) and one traded latency for cost.
Here’s how we told them apart, and why the rest of the wins came from reading
code rather than benchmarking it.
Anyone putting an LLM pipeline into production hits the same wall: it is too
slow, there are obvious things to try, and none of the usual profiling instincts
work. The provider’s latency moves 2x over minutes. Your effect size is 10%.
Your variance is 43%. Every benchmark you run is a coin flip you mistake for a
result.
We ran the experiment properly and it taught us more about measurement than
about latency:
The real wins came from auditing, not benchmarking:
End to end, the same 2,811 images went from 13.4 minutes to 4.2, with zero
errors. That projects to about 15 minutes for a 10,000-image dataset that took
137 before. Route, model and transport all changed in that number; the 8.7 -> 4.2
step is the transport alone.
Thesis: when run-to-run variance exceeds your effect size, stop benchmarking
and start reading code. Defects are measurable at n=1; optimisations are not. And a null result only
holds for the route you measured it on.
Engineers running inference pipelines in production, and more generally anyone
whose system depends on a service whose latency is noisier than the change they
are trying to make. No ML background needed — the subject is an LLM pipeline,
the content is measurement discipline.
asyncio.gather turns one slow request into a whole-stage stall,| Time | Section |
|---|---|
| 5 min | The pipeline, the 1h51m baseline, the three things we planned to try |
| 8 min | All three results: what the first pass said, and what they really did once measured properly |
| 7 min | Why we could not trust any of it: 43% variance and drift |
| 10 min | What actually worked: three defects found by reading code (1,402 consecutive HTTP 400s, a 458,752-character runaway output, synchronous prompt reads), and the written-off URI transport that halved runtime once routed through the proxy |
| 5 min | Bracket, never sweep: the protocol, the audit checklist, and the before/after |
Shaurya Madukuri. Data Scientist II at Skan.AI, where I lead the team building a
computer-use agent across more than twelve BFSI desktop workflows. B.Tech in
Electrical Engineering from IIT Delhi.
Most of my work is the unglamorous half of shipping models: making them fit,
making them fast, and finding out why they broke. The pipeline in this talk is
one of those. I also maintain forge-kernels (github.com/Shaurya-M002), an
open-source Triton kernel library for memory-efficient LLM training, and
published on glaucoma classification at IEEE ISBI 2024.
Based in Bengaluru.
LinkedIn: https://www.linkedin.com/in/shaurya-madukuri/ · X: https://x.com/shaurya_mk · GitHub: https://github.com/Shaurya-M002
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}