Opt-GPTQ ports grouped query attention, paging, and ALiBi into vLLM on Hygon DCU chips and measures small throughput gains, but lacks a GQA baseline, error bars, accuracy checks, and code.
Language models are few-shot learners,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
Opt-GPTQ ports grouped query attention, paging, and ALiBi into vLLM on Hygon DCU chips and measures small throughput gains, but lacks a GQA baseline, error bars, accuracy checks, and code.