Allowing each quantization group to select among multiple 4-bit grids improves accuracy over single-grid FP4 for both post-training and pre-training of LLMs.
Quantizing for minimum distortion.IRE Transactions on Information Theory, 6(1): 7–12
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 2years
2026 2roles
background 1polarities
background 1representative citing papers
OCTOPUS compresses KV caches via octahedral parametrization of rotated triplets and squared-error-optimized Lloyd-Max quantization, matching or exceeding prior rotation codecs with growing gains at low bit widths.
citing papers explorer
-
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
Allowing each quantization group to select among multiple 4-bit grids improves accuracy over single-grid FP4 for both post-training and pre-training of LLMs.
-
OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization
OCTOPUS compresses KV caches via octahedral parametrization of rotated triplets and squared-error-optimized Lloyd-Max quantization, matching or exceeding prior rotation codecs with growing gains at low bit widths.