Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Alosh Denny
Submitted Sep 29, 2026
One-line summary. A quantised model scored a perfect 100 on one category of a function calling benchmark while emitting tool calls on 1.2 percent of prompts, and the headline average made it look merely mediocre. This is a talk about validating a model change before it reaches production.
Every platform team that serves models eventually has to approve a change to the model itself: a quantisation, a version bump, a different serving runtime. The change is evaluated with a benchmark, a number comes out, and somebody decides. The problem is that a benchmark number can be produced by a system that has stopped working, and the aggregate can hide it completely.
I hit this while measuring quantised models on a function calling benchmark. One category scores the model on correctly declining to call a function when no function applies. A model that has stopped emitting tool calls at all passes every case in that category for free. Mine scored 100 there, called on 1.2 percent of prompts overall, and would have been reported at about 28.7 percent if I had averaged the categories the way the benchmark invites you to.
That is not a quantisation problem. It is an evaluation design problem, and it generalises to any guardrail metric where the safe behaviour and the broken behaviour look identical from outside.
Platform engineering, SRE, MLOps, developers shipping agents or tool using systems, and engineering leaders who sign off on model changes.
Level: intermediate.
Benchmarks and measurements. A failure mode with the full result table. The reasoning behind the screening strategy. Open source tooling, and a live walkthrough of the same trap in the raw results if the room wants it.
Open source project. Research or investigation. Hard earned engineering lesson.
Reporting a single aggregate number. It was the obvious thing to do and it would have published a broken model as a mediocre one.
Trusting a category whose scoring rewards abstention. I now treat any metric where doing nothing scores well as requiring a companion metric, without exception.
Assuming a result from one model size transfers. At 8B the six bit configuration holds at 88 percent. At 0.6B the same configuration collapses to 27. One number from one model would have produced confident and wrong guidance either way.
Report every category separately with the call rate beside it. Never aggregate across categories with different failure semantics. Screen with a cheap proxy before paying for a benchmark sweep. And run the sweep across sizes, because the interesting behaviour is at the edges.
Cheap proxy against real capability measurement. Loss is fast and correlates well, but it cannot tell you the output is malformed JSON. The answer was to use it as a screen, not as evidence.
Breadth against depth in evaluation. I chose four model sizes crossed with three precisions over one model examined thoroughly. That caught the size dependence and cost me the deeper per category analysis I would otherwise have had.
A specific class of mistake to avoid when approving model changes. A pattern for pairing capability metrics with liveness metrics. And an evaluation budget strategy that uses cheap signals to decide where to spend expensive ones.
Current state: open source.
Tags: #evaluation #observability #mlops #agents #inference #failurestory #modelserving
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}