Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Alosh Denny
Submitted Sep 29, 2026
One-line summary. Participants quantise a small open weight model to eight bits, watch it break in the specific way that looks like a precision problem but is not, fix it, and then build the evaluation that would have caught the breakage.
Most people meet quantisation as a flag in a serving config. It either works or it does not, and when it does not there is nothing to debug, because the model simply produces worse output. This workshop replaces that with hands on intuition. Participants implement the quantiser themselves in about forty lines, break their model deliberately, and diagnose it from the weights rather than from the output.
Developers, MLOps, platform engineering. Anyone who has enabled a quantisation flag without knowing what it does.
Level: intermediate.
Code and implementation details. A debugging technique. A live build of both the quantiser and the evaluation. Open source tooling that participants leave with.
A laptop with Python 3.10 or newer, PyTorch, NumPy and Matplotlib. No GPU, no cloud account and no dataset download. Everything runs on CPU with a small model. I will provide a setup script and a fallback notebook.
Open source project. Experiment or prototype. Research or investigation.
Every failure in the workshop is one I hit for real: the outlier weights that break a model at any bit width, the initialisation interaction that zeroes an entire layer at low precision, and the evaluation that rewards a model for going silent. The session is built around reproducing each of them deliberately in a few minutes.
Teach range before precision. Everybody’s intuition is that fewer bits means more error, and the failures that actually matter are about range and scale placement instead.
Small model on a laptop against a realistic model on rented GPUs. I chose small and local so that nobody spends the session fighting an environment, at the cost of some realism, which I cover with prepared results from the larger runs.
A design approach for evaluating precision options. A debugging technique for a failure that is otherwise invisible. And a mistake to avoid that costs teams real production quality.
Current state: open source.
Tags: #workshop #quantization #inference #mlops #evaluation #handson
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}