| |
Breaking the 1.58-bit Barrier for Ternary LLMs
Researchers have developed BITCOS, a new compression technique for ternary large language models that breaks the theoretical 1.585-bit barrier by exploiting the actual distribution of weights in these models, which contain up to 51.5% zeros. The method achieves 1.485 bits per weight on sparse models and delivers up to 1.28× faster matrix-vector multiplication and 1.27× faster inference throughput on GPUs compared to standard packing approaches. BITCOS includes optimized unpacking implementations for modern CPUs and GPUs, making it practical for deployment across multiple hardware platforms.
Read Full Article →
← More Tech news