Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Pankaj Merisha
Submitted Oct 5, 2026
Bare-Metal Agents: Building a Low-Latency Inference Stack for In-House Coding LLMs
Serving Qwen3.8-27B in FP8 on a DGX node and wiring it into VS Code coding agents — and the gaps you hit when upstream docs stop at the HTTP API.
What problem are you addressing?
Cloud coding agents bring data sovereignty, cost, and latency problems. Self-hosting an open-weight model sounds easy — until you run it on bare metal and try to make it usable for interactive coding.
Who is the intended audience?
Platform Engineering / SRE / DevOps / Developers
Level:
Advanced
Practical takeaways:
Exact launch configs and runtime flags that make local GPU inference usable for interactive agents.
A working blueprint: DGX node + SGLang + SSH tunnel + VS Code native agents with MCP tools.
What will you share?
Architecture decisions (vLLM → SGLang; terminal chat → Cline → VS Code agents)
Production experience running Qwen3.8-27B FP8 on a DGX node
Trade-offs and failure modes (token latency, KV cache fragmentation)
What is your experience with this problem?
Production system
Internal engineering project
What approaches failed, disappointed, or created unexpected problems?
First iteration: vLLM behind an SSH tunnel. Stable, but high token latency and KV cache fragmentation made multi-turn agent loops too slow for interactive use.
What will you do differently today?
Migrated the engine to SGLang and landed on VS Code’s native agent host as the client — the best of the three we tried. The SSH tunnel stays as transport.
What trade-offs did you consider?
vLLM (mature) vs. SGLang (faster for interactive workloads)
Terminal chat (simple, no tools) vs. Cline (MCP support, bolted-on) vs. VS Code native agents (best integration, newest)
Simplicity vs. bare-metal performance
How can this help other practitioners?
A pattern: invest in the right engine and client on each end
The exact configs that make local inference viable for interactive use
A mistake to avoid: “works at the HTTP API” ≠ “usable as a coding agent”
Current state:
Production experience
Tags:
#inference #agents #mlops #platformengineering #casestudy #demo
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}