What Is Quantization? FP16 vs INT8 vs INT4 and Why LLM Memory Shrinks
Key Takeaway
LLM quantization approximates a model’s weights with fewer bits, reducing both memory usage and the amount of data that needs to be read. Because numeric precision drops at the same time, there is also a chance of accuracy loss, and the extent varies by model and method.
This is part of my AI study notes series as I prepare for the SK hynix AI Hackathon preliminary round (Oct 17–18). Following Day1’s Self-Attention and Day2’s Transformer structure, today’s topic is quantization: reducing the number of bits used to represent numbers. Quantization is closely tied to the Decode process covered in Day3 (its memory-bound characteristic of re-reading weights from memory for every single token).
What is quantization, and what changes when you reduce the number of bits?
One-line answer: It is a technique that approximates weights (and activations, when needed) as numbers with fewer bits for storage and computation.
- Quantization approximates a model’s weights with fewer bits for storage and computation. It works the same way as converting a high-resolution photo into a compressed photo, which reduces the number of color gradations.
- Reducing the number of bits shrinks both memory usage and the amount of data that must be read. However, as numeric precision drops, there is also a chance of accuracy loss, and the extent depends on the model and the quantization method used.
- Related terms include PTQ (Post-Training Quantization), applied after training is complete, and QAT (Quantization-Aware Training), which accounts for quantization during training. Names like GPTQ, AWQ, and GGUF also come up often, but this post only introduces the names without comparing the methods.
How do FP16, INT8, and INT4 differ in memory, speed, and accuracy?
One-line answer: Compared to FP16, INT8 cuts memory roughly in half and INT4 to about a quarter, but the risk of accuracy loss can grow accordingly.
- The bytes needed to store a single parameter are 2 bytes (16 bits) for FP16, 1 byte (8 bits) for INT8, and 0.5 bytes (4 bits) for INT4. Dividing the bit count by 8 gives the byte count.
- As a simple arithmetic example, for a model with 7 billion parameters (7B), the weight memory alone comes out to roughly 14GB for FP16, 7GB for INT8, and 3.5GB for INT4. These figures cover weights only and exclude KV Cache, activations, and runtime overhead. Actual model file sizes vary by storage format and metadata.

- The differences by format can be summarized in the table below.
| Format | Bits | Bytes/Parameter | Memory (vs FP16) | Accuracy Risk (Qualitative) |
|---|---|---|---|---|
| FP16 | 16 | 2 | 1 (baseline) | Low |
| INT8 | 8 | 1 | 1/2 | Medium (can be larger than FP16) |
| INT4 | 4 | 0.5 | 1/4 | Highest |
The accuracy risk in the table above reflects only a qualitative trend; the actual drop depends on the model and method.
- Computation speed does not automatically improve just because the bit count is lower. Whether the hardware supports INT8 and INT4 operations, and which kernels are used, both change how large any speed gain is, so no blanket claim can be made.
How does quantization connect to the memory bandwidth (HBM) problem in Decode?
One-line answer: When weight size shrinks, the amount of data that must be read from HBM per token drops, easing the bandwidth bottleneck.
- As covered in Day3, the Decode stage has a memory-bound characteristic where it re-reads the weights from HBM every time it generates a single token.

- When the weights get smaller, the number of bytes that must be read per token drops, which is more favorable under the same bandwidth. Going from FP16 to INT4 arithmetically reduces the weight bytes read to one quarter.
- From an AI semiconductor perspective, how little you read and how fast you can read it matters just as much as raw compute power. The KV Cache can be a quantization target as well, not only the weights.
[Exam one-liner] Quantization = approximating weights with fewer bits → memory and data-to-read go down, but precision goes down too. Because Decode is memory-bound, this effect is especially large.
Shall we test your knowledge with a quick quiz?
Q1. If you quantize parameters from 16 bits to 4 bits, how does the weight memory change?
A1. It arithmetically shrinks to one quarter of the size (FP16 is 2 bytes, INT4 is 0.5 bytes).
Q2. What is the most important trade-off to watch out for when applying quantization?
A2. The potential loss of model accuracy caused by reduced numeric precision.
Q3. Why is quantization especially effective during the Decode process in LLMs?
A3. Because Decode is memory-bound and must repeatedly read weights from HBM for every token, reducing the amount of data to read significantly eases the bandwidth bottleneck.
Next up: Day4, prompt engineering.