Hook

Their other posts in the index, biggest breakout first.
Did you know that there are two ways to shrink a giant AI model and the other just rounds the numbers. Let's start with distillation. It trains a small student model to imitate a big teacher. See, the teacher doesn't just say it's a dog. It says 99% dog, 8% wolf, 2% cat. Not the answer, the whole distribution. Quantization on the other hand is a completely different move. It doesn't make a new model at all. It just takes the one you already have and rounds every weight to fewer bits. 16 bit numbers down to 4 bit for example. The same model, just written more coarsely. About four times smaller in some cases. But here's a clever thing that makes quantization actually work. Only about 1% of all the weights inside of an LLM really matters. So you keep those sharp and then all of the other ones round off. Distillation re-carves, quantization repacks. Together, a giant, in your pocket. Save this. Distillation changes WHO thinks. Quantization changes HOW it's written down. Distill. Quantize. AI shrink - both. Data center - your laptop. Which one? Data + compute, AI can train. Distill. Fit today, re-training. Quantize. AI shrink - both.