XFP introduces quality-targeted adaptive codebook quantization with sparse outlier separation that auto-selects parameters from cosine similarity floors, achieving high throughput and accuracy on Qwen3.5 models at low effective bits without calibration data.
NVFP4 Tensor Core Quantization---Blackwell Architecture Whitepaper
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference
XFP introduces quality-targeted adaptive codebook quantization with sparse outlier separation that auto-selects parameters from cosine similarity floors, achieving high throughput and accuracy on Qwen3.5 models at low effective bits without calibration data.