Dictionary entry
sub-1-bit quantization
also sub-one-bit quantization or sub-1-bit LLM compression or ultra-low-bit quantization
Definition of sub-1-bit quantization
-
: a model-compression approach that stores neural-network weights using an average of less than one bit per weight, typically by exploiting shared structure, sparsity, codebooks, or factorized binary representations rather than assigning a separate fractional bit to each weight
- The researchers used sub-1-bit quantization to fit a large language model into far less memory than a conventional 4-bit copy.
- A sub-1-bit result may report bits per weight even though the representation also needs scales, codebooks, or latent factors.
Origin & history
Research on sub-1-bit neural-network compression was established by 2023 in work that pushed average storage below one bit per parameter, including very large mixture-of-experts models. Later methods applied low-rank binary factorization, learned codebooks, and quantization-aware training to dense language models at fractions of a bit per weight.
Test yourself
Which of these is the meaning of sub-1-bit quantization?
Cite this entry
"Sub-1-bit quantization." AI Dictionary, Dadgogo, https://dadgogo.com/dictionary/sub-1-bit-quantization/. Accessed 9 Oct. 2026.