Prudhvi Raj

The GPU Was Waiting on the CPU: Taking the Host Off the LLM Serving Path

Submitted Oct 3, 2026

Session title

The GPU Was Waiting on the CPU: Taking the Host Off the LLM Serving Path

One-line summary

In LLM serving, the CPU launches and synchronises every kernel, and in real-time voice the GPU spends much of each request waiting on it. We compiled models ahead of time into a single GPU launch, ran it in production on 80 on-prem GPUs, and learned where that helps, where it doesn’t, and what it costs to operate.

What problem are you addressing?

Most serving engines are driven from the host: launch a kernel, sync, launch the next. At high batch sizes that overhead hides behind compute. At low concurrency (a voice agent, a single-user coding agent) it shows up as slower first tokens and jitter. That breaks a voice turn that has only a few hundred milliseconds.

  • Why it was hard: removing the host means owning everything it used to do. Scheduling, synchronisation across GPUs and memory placement all move into a compiled plan. We write our own kernels and fall back to cuBLAS only where ours don’t beat it yet, so one compiled plan targets both NVIDIA (CUDA driver API) and AMD (HSA).
  • Constraints: on-prem GPUs the customer already owns, an OpenAI-compatible API, accuracy parity with the existing engine, and several models (speech and LLM) sharing a GPU.
  • Why it matters to platform teams: more teams are moving inference from APIs onto their own or rented GPUs. Where host overhead sits decides how many GPUs you need for a latency target.

Intended audience

Platform engineers, SREs and MLOps engineers running self-hosted inference; engineering leaders deciding between managed APIs and their own GPUs.

Level

Intermediate. No kernel programming needed; familiarity with vLLM or SGLang helps.

Practical takeaways

  1. How to tell whether your serving latency is GPU-bound or host-bound, with the traces to collect and what each pattern looks like.
  2. When ahead-of-time compilation is worth its operational cost, and when a dynamic-batching engine like vLLM is the better choice.

What will you share?

  • Architecture: how a model becomes a precompiled artifact per GPU type, and how counters replace host-side syncs within and across GPUs.
  • Benchmarks, with versions and concurrency stated. On one H100, Gemma-4 26B MoE time to first token falls from 38.3 to 21.4 ms at 1 user and from 71.9 to 36.7 ms at 4 users, against vLLM 0.28. Per-token latency grows 1.4–1.5× from 1k to 128k context on 8× MI350X, vs 1.9–2.9× for the vLLM configurations we tested.
  • Production experience: running a consumer voice assistant on 80 on-prem GPUs.
  • Failure modes and trade-offs (below), plus the open-source runtime and benchmark harness.

Your experience with this problem

Production system; open-source project (github.com/infervisor/plow, Apache-2.0); hard-earned engineering lessons.

What failed, disappointed, or created unexpected problems?

  • High concurrency. Our wins are clearest at 1–4 users. At high batch sizes dynamic-batching engines still do well, and we haven’t closed that gap yet. We’ll show those curves too.
  • Owning the kernels. Writing our own kernels means every new model architecture is a bring-up project. Kimi-K3 on AMD MI355X went from “cannot load” to a correct answer only after new kernels for its hybrid attention.
  • Reproducible but unusable install. A Nix-only build was perfect for our own reproducibility and a blocker for every team evaluating us. [confirm: Docker/pip status]
  • Stale baselines. Comparisons against one engine version went out of date fast; public issues show one engine’s AMD build losing about 38% throughput between two minor releases.

What would you do differently today?

Ship containers before asking anyone to evaluate. Pin and publish baselines with every result. Publish high-concurrency curves from day one rather than after people ask.

Trade-offs considered

  • Static plan vs dynamic scheduling: we optimised for predictable per-request latency and gave up some of the flexibility that helps throughput at high batch sizes.
  • Own kernels vs vendor libraries: our own kernels where they win, keeping one plan across NVIDIA and AMD, and cuBLAS where it is still faster. The cost is more engineering per model.
  • Precompile vs JIT: start-up about 10× faster with no JIT at boot, at the cost of managing one artifact per model and GPU.
  • Replace vs coexist: our control plane runs vLLM and SGLang next to our runtime, so teams route only latency-critical traffic to it.
  • Correctness: Lean 4 proofs gate each compiler stage, and every speed number ships with an accuracy check (GSM8K parity; 0 failed requests out of 3,840 in one H100 run).

How can this help other practitioners?

  • A debugging technique: separating host-bound from GPU-bound latency in your own serving stack.
  • A way to evaluate competing engines: compare curves across concurrency and context length, with pinned versions and an accuracy gate, instead of one headline number.
  • A pattern to adopt: route latency-critical and batch traffic to different engines behind one OpenAI-compatible gateway.

Current state

Production experience (also open source).

Tags

#inference #platformengineering #mlops #scalability #latency #gpu #casestudy #failurestory #opensource

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy