Quantization: Precision vs. Speed
Loading learning experience...
Lecture transcript
Read the narration for Quantization: Precision vs. Speed
From "what is the model thinking?" to "can we run it fast?"
Dr. Lena Hartmann: Every time you deploy a model, you hit a very practical wall: memory bandwidth and latency. Today we are going to discover quantization, the idea of using fewer bits per number to make the same matrix multiplications cheaper. The core idea is surprisingly simple, and we will build it from a tiny numpy experiment. Our goal for now is just the definition and one driving question about what stays almost the same.
Dr. Lena Hartmann: Last lecture, mechanistic interpretability treated activations like vectors: you look at directions, dot products, and how features light up.
Kai: So quantization is like changing the number format, but we still want those feature directions to mean the same thing?
Dr. Lena Hartmann: Exactly. Our guiding question is: when we approximate numbers, which dot products and matrix outputs stay close enough that the model behavior barely changes?
Why quantize? Memory and throughput set the budget
Dr. Lena Hartmann: Let me frame the stakes in one line: big models are often limited not by compute, but by moving weight matrices from memory into the math units.
Dr. Lena Hartmann: This equation says parameter size, meaning just the weights, is roughly the number of parameters times bits per parameter, converted into bytes.
Kai: So cutting from sixteen bits to four bits is like a four times memory win, before we even talk about speed?
Dr. Lena Hartmann: Right. And once memory traffic drops, matrix multiply kernels can run faster too, because the G P U is fed more efficiently.
Kai: Does this also reduce energy use, since you are moving fewer bytes and doing cheaper arithmetic?
Kai: And is it fair to say the real risk is that the model will start behaving differently, even if the speedup looks great on paper?
Dr. Lena Hartmann: Yes. One more practical note: this formula is only for the weights. At inference, total VRAM can be dominated by the K V cache, activations, and other buffers, and we will revisit that on Slide 8. The key tension is that smaller bit width means more approximation, so we need to understand what errors we introduce in linear algebra outputs.
Quantization as "snapping" real numbers to a grid
Dr. Lena Hartmann: Before we talk about bits and hardware, I want a mental picture: quantization takes a real number and snaps it to nearby allowed levels.
Dr. Lena Hartmann: Start with x as any real value. In a neural net, this could be one weight or one activation entry.
Kai: And the quantized hat x is like rounding, but with a limited number of representable levels?
Dr. Lena Hartmann: Exactly. The difference e is the error. In matrix multiplies, lots of small errors can add up, so we will measure how outputs drift.
Let me show you in code: quantize a vector and compare directions
Dr. Lena Hartmann: Before any formulas, let me show you what quantization does to an actual vector, because linear algebra cares about both magnitude and direction.
Dr. Lena Hartmann: We will treat w like a weight vector in a model. Quantization creates a new vector w hat with fewer representable values.
Kai: Why measure cosine here instead of just mean squared error?
Dr. Lena Hartmann: Because many model behaviors depend on dot products, and dot products care a lot about direction. Cosine similarity tells you if the vector points the same way after quantization.
Dr. Lena Hartmann: Before you run it, make a prediction: which will have the higher cosine, eight bit or four bit? And roughly how close do you think they are, like 0.999 versus 0.95?
Dr. Lena Hartmann: Now run it and compare the printed numbers to your guess: eight bit should be closer to 1, and four bit should drop more, matching the idea that int eight is often nearly lossless while int four is a bigger bet that depends on the model and layer.
Dr. Lena Hartmann: For example, if you see something like cos int eight around 0.999 or higher and cos int four around 0.97 to 0.99, that means both keep the direction pretty well, with four bit noticeably noisier. If int four drops closer to 0.9 or lower, that is a warning sign that a few large values are setting the scale and squeezing everything else, so you would consider a different scaling strategy or per channel quantization. And if int four is surprisingly high, almost matching int eight, that usually means the vector has a friendly distribution with no big outliers, so low precision is enough for direction.
What error does rounding introduce, locally?
Dr. Lena Hartmann: Now we build intuition for why the error grows as we use fewer bits: the quantization grid gets coarser.
Dr. Lena Hartmann: This inequality is the key rounding fact: if you round to the nearest grid point with spacing s, you are never more than half a step away.
Kai: So bits basically control how small that step size s can be over the value range we care about?
Dr. Lena Hartmann: Yes. If you want to cover a wider dynamic range with the same number of levels, the steps get bigger and the bound loosens.
Dr. Lena Hartmann: And for linear algebra, remember a dot product is a sum of many multiplications. Those rounding errors can accumulate, so we care about output drift, not just per entry drift.
The standard quantization map used in practice
Dr. Lena Hartmann: Now that you have the snapping picture, here is the exact quantize and dequantize mapping that most deployment toolchains implement.
Dr. Lena Hartmann: The first line says: q equals clip of round of x over s, plus z, clipped between q min and q max. In words, you scale by s, round to an integer, add the zero point z, then clamp to the allowed integer range.
Kai: And the second line is how you get back a float approximation, right? So the model can still do math.
Dr. Lena Hartmann: Exactly: dequantization. You compute x hat as s times the quantity q minus z. The scale sets the step size, and the zero point shifts the grid so that an exact zero is representable even when the integer range is asymmetric.
Dr. Lena Hartmann: We will use this map as the basic definition, and then next we will look at a key practical choice: whether one scale is shared everywhere, or whether you keep multiple scales to better fit different parts of a tensor.
Worked example: quantize $W$ and watch $\mathbf{y} = W\mathbf{x}$ drift
Dr. Lena Hartmann: Quantization matters because models are mostly linear algebra. So the real test is not one number, but a matrix times a vector.
Dr. Lena Hartmann: Here we compute y equals W times x, then repeat with a quantized approximation of W and compare the outputs. Before you look at the numbers, make a prediction: will the relative error for int four be closer to one e minus three, one e minus one, or one?
Kai: Is this basically what happens inside every linear layer in a transformer, just at huge dimensions?
Dr. Lena Hartmann: Yes, same operation, just bigger. Run the code and look at relative error and cosine similarity of the output vector. Int eight usually stays much closer than int four with a single global scale, and the int four relative error is often around one e minus one rather than one e minus three because the noise adds up across the input dimension. If you got something closer to one, that usually means heavy clipping from outliers or a bad scale choice, and the usual fixes are per channel or groupwise scaling, or adding a bit of clipping, which we will connect to in the next slides.
Dr. Lena Hartmann: This bullet is the linear algebra reason: each output entry is a dot product, so it aggregates many small rounding errors across the input dimension.
Dr. Lena Hartmann: And this is the engineering viewpoint: you do not need zero error. You need error small enough that the downstream logits and decisions barely change.
Where quantization helps in transformers
Dr. Lena Hartmann: Now connect the math back to the system: transformers are a pile of matrix multiplies and attention, so quantization targets the largest tensors.
Dr. Lena Hartmann: First, weights. They are fixed at inference time, so you can quantize once and reuse, which is why weight only quantization is so common.
Kai: Activations sound scarier because their values depend on the prompt, so the scale might be wrong?
Dr. Lena Hartmann: Exactly. Activations move around, so you often need calibration or dynamic scaling. Next, the K V cache gets huge for long contexts, so quantizing it can be a major memory win.
Dr. Lena Hartmann: Finally, speed comes from specialized low precision kernels: the same linear algebra, but executed with instructions that move and multiply smaller numbers faster.
Choosing $s$ and $z$: range, clipping, and outliers
Dr. Lena Hartmann: Quantization quality lives or dies on one choice: how you map the floating range into the integer range.
Dr. Lena Hartmann: This is the classic min max rule: pick a scale from the float range, then pick a zero point so zero lands on an integer exactly.
Dr. Lena Hartmann: The first bullet says what it is doing: it tries to represent the full observed range without overflow.
Kai: So one weird outlier value can force the step size to get huge, making everything else low precision?
Dr. Lena Hartmann: Yes, and that is why people talk about outlier channels. Practical toolchains often use per channel scaling or controlled clipping to spend precision where most values actually are.
Going lower: why $INT4$ needs smarter granularity
Dr. Lena Hartmann: Dropping to four bits is where the linear algebra pain becomes obvious, so we usually need more structure than one scale for the whole matrix.
Dr. Lena Hartmann: So far we used symmetric ranges, plus or minus q max. For int four storage, we often use negative eight to seven, which is slightly asymmetric. With int four, your integers only have sixteen possible values, and we set the scale using plus seven, so the negative side gets one extra value.
Kai: Groupwise scales means different parts of W get their own step size, so outliers in one block do not ruin precision everywhere?
Dr. Lena Hartmann: Exactly. This compares global int four versus groupwise int four. You should see the mean squared error drop when each block gets its own scale.
Dr. Lena Hartmann: So the first takeaway is: you can buy back accuracy at the same bit width by increasing the number of scales.
Dr. Lena Hartmann: But nothing is free: you now store extra scale metadata and you need kernels that understand this block structure efficiently.
Quantization: the linear-algebra viewpoint
Dr. Lena Hartmann: Let us compress the whole lecture into a linear algebra story you can reuse in practice.
Dr. Lena Hartmann: First: quantization is a projection. You take a real valued tensor and replace it with the nearest representable value on some grid.
Kai: And the practical question is how that projection changes the outputs of the dot products the model actually uses.
Dr. Lena Hartmann: Exactly. Second: the error is not abstract, it appears in W times x and attention computations. Third: the performance gain comes from moving fewer bytes and using specialized int eight or f p eight kernels.
Dr. Lena Hartmann: Finally: you control the trade off with bit width, the choice of scale and zero point, and granularity choices like per channel and groupwise scaling.
Exit ticket: compute one quantization step
Dr. Lena Hartmann: Time to do one by hand, because once you can compute one quantization, the rest is just repetition across tensors.
Dr. Lena Hartmann: For the practice: compute x over s, so 0.37 over 0.1 is 3.7. Round that to 4, then add the zero point 128 to get q equals 132. That is within 0 to 255, so no clipping.
Dr. Lena Hartmann: Then dequantize: subtract the zero point, 132 minus 128 is 4, and multiply by s. So x hat is 0.4. The error is 0.37 minus 0.4 equals minus 0.03.
Kai: For the reflection: I would avoid int four if that layer has outliers or is really sensitive, like attention projections or small models where a little drift changes outputs a lot.
Dr. Lena Hartmann: Perfect. The real rule is empirical: if output drift flips predictions or hurts generation quality, you either raise precision or use smarter scaling and calibration.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!