Short definition

Quantization — storing the weights of an AI model with fewer bits per number, so the model needs less memory and can run on smaller hardware, at the cost of a small loss in accuracy.

Also known as: model quantization, LLM quantization, quant.

Quantization makes a large AI model fit on smaller hardware. Instead of storing every weight as a 16-bit number, you store it in fewer bits, often between 2 and 8. The model then needs far less memory, at the cost of a small, measurable loss in accuracy.

How does quantization work?

A language model consists of billions of numbers, its weights. In the original release these are usually 16-bit floating point values. Quantization maps them onto a smaller set of values that can be stored in fewer bits, plus a little extra data such as scales to reconstruct them during inference.

The simplest methods round each weight on its own. Newer methods look at groups of weights at once and use calibration data to decide which details matter most. That is why two quants with the same number of bits per weight can differ in quality: the method matters as much as the bit count.

Common formats

You mostly meet quantization through the file format of a download:

  • GGUF is the format of llama.cpp. It runs on CPUs, Apple Silicon, and most GPUs, which makes it the broadest choice.
  • EXL3 is the format of ExLlamaV3, built for NVIDIA GPUs. It is based on QTIP, a trellis-based method from Cornell RelaxML.
  • GPTQ and AWQ are older GPU methods that are still common in serving frameworks.

In our tutorial on running Qwen3.8-Flash-Next locally, a 3.05 bpw EXL3 quant brings a 125B mixture-of-experts model down to an 85.1 GB download that runs on one consumer GPU plus system RAM.

What should you check when choosing a quant?

  • Memory: the quant has to fit in your VRAM, or in VRAM plus system RAM when part of the model is offloaded to the CPU.
  • Quality: compare published measurements such as KL divergence or perplexity for the same model, not just the bit count.
  • Hardware support: a format is only useful if your inference engine and hardware can run it.