|
Rootconf 2026 Annual Conference We Burned 48 GPU Hours Finding Out How Few Bits A Model NeedsOne-line summary. I spent a year answering one question, how few bits a language model actually needs, and the infrastructure I answered it on failed in five different ways that had nothing to do with machine learning. This talk is both halves: the research, the day a container spent training on a tenth of its corpus while every dashboard stayed green, and the question underneath both, which is w… more
I am submitting: To speak
I have a submission for: A talk - 30-40 mins
|
|
Rootconf 2026 Annual Conference Your Model Does Not Need Floating Point, But It Does Need One Thing You Probably SkippedOne-line summary. Open weight models quantise to eight bit fixed point after training, with no retraining and no calibration set, and keep their tool calling intact, provided you get one detail right that silently destroys the model if you miss it. more
I am submitting: To speak
I have a submission for: A talk - 30-40 mins
|
|
Rootconf 2026 Annual Conference It Scored 100 Percent Because It Had Stopped AnsweringOne-line summary. A quantised model scored a perfect 100 on one category of a function calling benchmark while emitting tool calls on 1.2 percent of prompts, and the headline average made it look merely mediocre. This is a talk about validating a model change before it reaches production. more
I am submitting: To speak
I have a submission for: A talk - 30-40 mins
|
|
Rootconf 2026 Annual Conference Break A Model, Then Fix It: A Hands-on Session On Quantisation And Proving It WorkedOne-line summary. Participants quantise a small open weight model to eight bits, watch it break in the specific way that looks like a precision problem but is not, fix it, and then build the evaluation that would have caught the breakage. more
I am submitting: To speak
I have a submission for: A talk - 30-40 mins
|
|
Rootconf 2026 Annual Conference Who Owns The Precision Decision, And How Does Your Team Prove It Was Safe?One-line summary. Quantisation changes the model, but the decision usually sits with whoever owns the serving cost. I want to compare notes on who actually makes that call across teams, and what evidence they require before it ships. more
I am submitting: To speak
I have a submission for: A talk - 30-40 mins
|
|
The Fifth Elephant 2026 Winter Edition It Scored 100 Percent Because It Had Stopped AnsweringA model I had compressed scored a perfect 100 on one category of a function calling benchmark. It looked like the best result in the table. It was produced by a model that had stopped emitting function calls almost entirely, on 1.2 percent of prompts, and the category in question scores a model on correctly declining to call a function when none applies. A system that has gone silent passes every… more
I am submitting: To speak
I have a submission for: A talk - 30-40 mins
|
|
The Fifth Elephant 2026 Winter Edition Two Things I Published, And Then KilledI published two empirical claims about training neural networks at low precision. Both were plausible, both matched what I had observed, and both are now wrong in my own repository, with the retraction sitting directly underneath the original text. This session is the story of how each one died, because the interesting part is not the claims themselves but the shape of the mistakes, which I think… more
I am submitting: To speak
I have a submission for: A talk - 30-40 mins
|
|
The Fifth Elephant 2026 Winter Edition Delete The Exponent: What A Model Actually Does With Its BitsA floating point number spends most of its bits buying dynamic range. A trained neural network barely uses that range: weights are bounded and clustered near zero. So I built the format that takes the premise literally. Superfloat keeps the sign bit, spends every remaining bit on the significand, and has no exponent field at all, which makes it plain signed fixed point. Twenty minutes is enough t… more
I am submitting: To speak
I have a submission for: A talk - 30-40 mins
|
|
The Fifth Elephant 2026 Winter Edition How Few Bits Does A Model Need? 878 Training Runs, One Answer, Several SurprisesThis poster presents a year of sweeps asking one question: how few bits does a neural network actually need, and what determines the answer? It covers 878 archived training runs across four scaling tiers and eight follow up experiments, on convolutional networks and transformers, under both training aware and post training compression. more
I am submitting: To speak
I have a submission for: Poster - work I have done that I can share in the alleyways with attendees
|
|
The Fifth Elephant 2026 Winter Edition Why Is The Checkpoint You Were About To Ship The Worst One To Compress?Take one model, compress it after training, and measure the damage. Now do that at seven points along its training run, from 2.1 billion tokens to 300 billion. The damage is not monotone. It falls steeply, bottoms out in the middle of training, and then rises sharply at the end. For a 160 million parameter model at seven bits, the penalty is 0.60 nats early, 0.13 nats at the minimum, and 3.16 nat… more
I am submitting: Have a problem I want to discuss
I have a submission for: Birds of Feather - I have a problem/idea I want like-minded folks to flock around
|
|
The Fifth Elephant 2026 Winter Edition Offering Help On Evaluation Design, Model Compression And Killing Your Own ResultsHow I want to help rewrite vague problem statements more
I am submitting: To help with problem statements
I have a submission for: Volunteering - as facilitator or to help someone with their problem
|
Alosh Denny
@aoxo
I'm an AI researcher that loves making and breaking LLMs
- Joined Sep 2026
- aloshdenny.com