REVIEW 4 major objections 6 minor 2 cited by
AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read AMXFP4, an asymmetric microscaling 4-bit floating-point format that gives positive and negative elements their own shared scales, enables calibration-free direct-cast 4-bit LLM inference that matches or beats rotation-based INT4 pipelines…
desk verdict Solid format paper with a real insight; abstract overclaims and the emulator/hardware gap needs a check. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the pair of asymmetric shared scales. For each group of 32 elements, AMXFP4 stores a positive shared scale $2^{e_{sp}}\hat{M}_p$ and a negative shared scale $2^{e_{sn}}\hat{M}_n$, both FP8 with 5 exponent bits and 2 mantissa bits, and encodes each element as an FP4 value with sign, 2-bit exponent, and 1-bit mantissa. During multiplication the product of two elements takes one of four scale combinations, $S_{Xp}S_{Wp}$, $S_{Xp}S_{Wn}$, $S_{Xn}S_{Wp}$, or $S_{Xn}S_{Wn}$, selected by a small lookup on the two operand signs; because the scale mantissa is only 2 bits and the same scale serves the whole group, this sign-aware scaling is nearly free in hardware. The design also replaces MX's floor-based power-of-two scale decision with rounding, avoiding the clamping error that otherwise grows as group size shrinks. Together these choices give the format an asymmetric grid of representable values that matches the lopsided group distributions that microscaling itself creates.
What would settle it
Feed the same random and real activation tensors through the software emulation from Appendix B and a detailed simulation of the synthesized AMXFP4 MAC unit, and compare outputs bit by bit; any divergence in scale-pair selection, clamping, or rounding would show that the reported accuracy gains may not hold on the hardware whose cost is claimed.
Extended reading notes
Core claim
The central discovery is a trade-off hidden in microscaling: shrinking the quantization group to 32 elements tames activation outliers (kurtosis falls nearly to zero), but it scatters the group means, meaning each group is more asymmetric the finer the group granularity gets. MXFP4's symmetric representation therefore leaves error on the table, and data rotation, which helps at row-level group sizes, actively hurts when combined with microscaling because it adds still more asymmetry. AMXFP4 counters this with two shared FP8 (E5M2) scales per group, one for positive and one for negative elements, with the scale chosen at multiply time from the signs of the two operands. On Wikitext-2 this lowers perplexity from 6.49 to 6.22 for LLaMA2-7B relative to MXFP4, lifts ChartQA from 46.20 to 49.48, and in Table 10 reaches a WinoGrande accuracy of 67.32 against a best rotation-baseline value of 66.22, without any calibration. The format also narrows the gap to the 16-bit baseline enough that MT-Bench conversational scores recover close to baseline, and a synthesized MAC unit implementing the sign-dependent scale selection is reported at only about 10% area overhead over a compatible MX MAC.
Load-bearing premise
The load-bearing premise is that the software emulator used for all accuracy measurements behaves exactly like the separately synthesized hardware MAC unit, including the sign-dependent choice of shared scale and the treatment of clamping and rounding; the paper reports the two implementations separately and does not show bit-level agreement between them.
Editorial extensions
If this is right
- All attention matrix multiplications, including the softmax-output and query-key products that rotation methods leave in FP16, can be run in 4-bit, which matters most as context length grows because attention FLOPs scale quadratically.
- Models can be deployed at 4-bit by direct casting with no calibration run, removing the multi-hour overhead and the calibration-set overfitting that rotation-based pipelines exhibit.
- A 3-bit variant, AMXFP3, degrades Wikitext-2 perplexity by only about 1.7 on LLaMA2-7B, whereas QuaRot with GPTQ degrades by more than 30, suggesting the approach extends below 4 bits.
- AMXFP4 is compatible with other compression methods: applied to a 20%-pruned LLaMA-7B model it recovers most of the pruning accuracy drop, so its benefits are additive.
- An asymmetric version of the recently deployed commercial MXFP4 variant NVFP4, called ANVFP4, also beats the commercial format, particularly at group size 16, indicating that per-sign scales are a generally useful addition to MX-family formats.
Reading between the lines
- The paper's own group-size sweeps show AMXFP4's advantage over symmetric MXFP4 widening as groups shrink, so a natural extension is to push the asymmetric shared-scale idea to even smaller groups (for example, 16 or 8 elements) or to per-channel granularity.
- Because the mechanism targets distribution shape rather than LLM-specific structure, the same asymmetric shared-scale design could plausibly transfer to other outlier-heavy workloads such as vision transformers, diffusion models, or training-side rescaling, though the paper does not test these.
- The emulator-versus-hardware gap could be closed by a bit-exact co-verification harness; without that artifact, the accuracy story and the 10% overhead story remain two claims about two different implementations.
- A testable prediction of the paper's analysis is that combining row-level rotation with fine-grained asymmetric microscaling would recover the row-level rotation benefit without the destructive interaction seen at group size 32, something the paper's Table 5 suggests but does not directly explore.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes AMXFP4, a 4-bit asymmetric microscaling floating-point format for LLM inference. The format keeps FP4 (E2M1) elements but uses separate FP8 E5M2 shared scales for positive and negative values, with a sign-dependent scale-selection rule during multiplication. The authors argue that microscaling suppresses activation outliers but increases group-wise asymmetry, and that the asymmetric shared scale addresses this without calibration. They evaluate AMXFP4 on Wikitext-2 perplexity, MT-Bench, visual question answering, LongBench-E, MMLU, CSQA, and attention-only settings, reporting consistent improvements over MXFP4 and competitive or better accuracy than rotation-based methods such as QuaRot and SpinQuant. They also implement a custom AMXFP4 MAC unit and claim roughly 10% hardware overhead over an MX-compatible MAC. The paper includes ablations on shared-scale format, group size, rotation interaction, QAT, 3-bit extension, pruning, and a 70B model.
Significance. If the claims hold, AMXFP4 is a practically useful format: it enables calibration-free direct-cast 4-bit inference with fully quantized attention, which is a gap relative to rotation-based methods that require calibration and leave softmax outputs in FP16. The paper is creditable for releasing code, evaluating across a broad set of tasks and model families (including a 70B model and an encoder-decoder model), and for providing extensive ablations, including the interaction of rotation with microscaling and the choice of shared-scale encoding. The central accuracy comparisons (Tables 1, 3, 4, 10, 12) support the qualitative claim that asymmetric shared scales improve over symmetric MXFP4 and often match or exceed rotation-based methods. However, the hardware-cost claim is currently not supported by the reported numbers, and the accuracy results come from a software emulator that is not shown to be bit-exact with the synthesized hardware. These gaps are load-bearing for the paper's central accuracy-plus-cost claim.
major comments (4)
- [§5.5, Table 9] The hardware-cost claim is not supported by the displayed numbers: Table 9 reports AMXFP4 at 8.32x versus MXFP4 at 9.23x on Area-Memory, and AMXFP4 is also lower on Power-Area and Power-Area-Memory. This implies AMXFP4 is roughly 10% cheaper, not 10% more expensive, as the text states in Section 4.3 and Section 5.5 ("adds only 10% overhead"). Please correct the rows, units, or the baseline for the overhead statement, and re-run the synthesis analysis if necessary, because the claimed cost is a central part of the contribution.
- [§5.5 vs. Appendix B.3] The accuracy results are produced by a software emulation path (quantize_mx_op with fp4_e2m1_asym and scale_mode=152 in Appendix B.3), while the hardware cost is measured on a separately synthesized AMXFP4 MAC unit in Section 5.5. The paper provides no bit-exact comparison between these two implementations for the same tensors, including the sign-dependent selection between the positive and negative FP8 E5M2 shared scales, the exponent-rounding rule, and the clamping behavior. Please add such a check, or characterize the numerical divergence if the implementations are not bit-identical, so that the reported accuracy improvements can be attributed to the hardware design whose cost is claimed.
- [Abstract and Tables 3, 4, 10] The abstract's quantitative claims exceed what the tables show. "Outperforms MXFP4 by 3% on VQA" does not match Table 3, where the gains are 1.25 points on VQA-T, 2.72 on DocVQA, 0.50 on OCRBench, and 3.28 on ChartQA; there is no single 3% improvement across the VQA benchmarks. Similarly, "exceeds rotation-based methods by 1.6% on CSQA" is not directly shown: Table 10 reports ARC-Challenge and WinoGrande, and Table 4 compares against NVFP4 rather than rotation-based methods. Please revise the abstract to cite the exact table and metric, or add the missing CSQA comparison.
- [§4.2, Fig. 5, Table 18] The choice of shared-scale format (E5M2 vs. E4M3 vs. PoT with floor or round) is made by minimizing Wikitext-2 perplexity on LLaMA2-7B (Fig. 5, Table 18), and Wikitext-2 perplexity is also a headline evaluation metric throughout the paper (Tables 1, 12, 14). The paper should either report whether the selected configuration remains optimal on a held-out task or model that was not used for selection, or explicitly discuss the potential selection bias. As written, part of the reported advantage is tuned to the evaluation objective, which weakens the claim that E5M2 is intrinsically the best shared-scale choice.
minor comments (6)
- [§4.3] The text refers to "Appendix 5.5" when describing the hardware evaluation; this should be Section 5.5.
- [Appendix B.1, Algorithm 1] Algorithm 1 still describes the original floor-based MX quantization, but Section 4.2 introduces a modified rounding rule for the PoT scale; please present the updated algorithm or state explicitly that the emulator implements the rounding variant.
- [Fig. 1] The caption lists two subfigures labeled "(d)" and the subfigure letters do not match the order in which they are discussed in the text; please renumber the subfigures.
- [Table 9] The column headings "Area-Memory", "Power-Area", and "Power-Area-Memory" need explicit definitions (e.g., whether these are products of normalized ratios) so the reader can interpret the reported multipliers.
- [§5.2] The term "MXFP4-PoT" is used before it is defined in Section 4.2; please define it at first use.
- [§5.1, Table 2] The paper does not report the number of random seeds or trials for the rotation-based comparisons; adding this information would improve the reliability of the overfitting analysis.
Circularity Check
Format choices are selected on the same LLaMA2-7B/Wikitext-2 perplexity that is later reported as evidence, but the central accuracy claim is independently supported on held-out tasks.
-
other
[Sec. 4.1-4.2 (Fig. 5(a), Table 18); reported in Tables 1 and 12]
"To evaluate the benefits of asymmetric formats, we compare the mean-square error (MSE) on activation samples from LLaMA2-7B’s QKV-Proj at layer 5 ... This finding supports the selection of AsymFP4 as the element-wise format, further validated empirically in Table 1. ... Therefore, we select FP8 with a 5-bit exponent (E5M2) as the shared scale, as these scales largely mitigate accuracy degradation caused by the limited resolution and narrower dynamic range (see Table 18 for ablation studies)."
The element format (AsymFP4) and shared scale (E5M2) are chosen by minimizing quantization MSE on LLaMA2-7B activations and by minimizing LLaMA2-7B Wikitext-2 perplexity across candidate formats (Fig. 5(a), Table 18). Tables 1 and 12 then present the lower Wikitext-2 perplexity of AMXFP4 on LLaMA2-7B as empirical validation. On this specific row the comparison is the selection objective: the chosen format's perplexity is lower than the rejected alternatives by construction of the argmin, so it is not an independent confirmation. The circularity is limited because the paper's broader claims are reproduced on held-out models/tasks (VQA, CSQA, MT-Bench, LongBench, LLaMA3-70B) that were not used for format selection.
full rationale
Aside from the in-sample format-selection issue above, the derivation chain is not circular. The proposed AMXFP4 encoding (Eq. 1) is a concrete extension of the cited AsymFP/AFPQ idea to group-wise FP8 shared scales; it is not a renaming of a known result, and it does not import any uniqueness theorem from the authors' prior work. The accuracy comparisons against MXFP4, QuaRot, SpinQuant, and NVFP4 are empirical and mostly on tasks/models outside the selection data. The hardware claim in Sec. 5.5 is a validation gap rather than a circularity: Table 9's numbers appear inconsistent with the stated ~10% overhead, and no bit-exact check links the software emulator (Appendix B.3) to the synthesized MAC, but these are correctness/reporting concerns, not a reduction of the result to its inputs. The limitation paragraph explicitly acknowledges that only MAC-level hardware evaluation was performed.
Assumptions & free parameters
free parameters (3)
- Shared-scale FP8 exponent/mantissa split =
E5M2 (1-5-2)
- Power-of-two shared-scale rounding rule =
Round instead of floor
- Element-wise format =
AsymFP4 (E2M1 with sign-dependent scale)
assumptions (4)
- domain assumption Group-wise kurtosis and mean are sufficient to characterize quantization difficulty under microscaling
- domain assumption Reducing element-wise MSE on sampled activations improves end-task LLM accuracy
- domain assumption The MX emulation library faithfully models the numerical behavior of the proposed AMXFP4 hardware
- domain assumption MAC-level hardware cost is a valid proxy for inference-system-level cost
invented entities (1)
-
AMXFP4 format: asymmetric microscaling FP4 with separate FP8 E5M2 shared scales for positive and negative values
Cite this review
Pith. "Pith review of AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference." pith.science (2026). https://pith.science/paper/VL4S7YSV
@misc{pith2026241109909,
author = {Pith},
title = {Pith review of: AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/VL4S7YSV}},
note = {Machine review of arXiv:2411.09909}
}
read the original abstract
As large language models (LLMs) grow in parameter size and context length, computation precision has been reduced from 16-bit to 4-bit to improve inference efficiency. However, this reduction causes accuracy degradation due to activation outliers. Rotation-based INT4 methods address this via matrix calibration, but they introduce multi-hour overheads and leave key computations in full precision. Microscaling (MX) floating-point (FP) formats offer fine-grained representation with a shared scale, enabling fully quantized matrix multiplications through direct casting without calibration. However, existing research shows unsatisfactory empirical results for MXFP4 inference, and the robustness of MX formats remains largely unexplored. In this work, we uncover the fundamental tradeoffs of the MX format: while it effectively suppresses activation outliers, it does so at the cost of increased group-wise asymmetry. To address this, we propose AMXFP4, a 4-bit asymmetric FP format that handles both issues using asymmetric shared scales, without requiring calibration. Our custom MAC engine adds negligible hardware cost while improving accuracy: AMXFP4 outperforms MXFP4 by 3% on VQA and exceeds rotation-based methods by 1.6% on CSQA. It also surpasses recently deployed commercial MXFP4 variants. Code: https://github.com/aiha-lab/MX-QLLM
Figures
Figures from the paper (9 more)
Forward citations
Cited by 2 Pith papers
-
LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
LightRot uses grouped local rotation plus outlier alignment to make 4-bit LLaMA inference accurate and cheap, claiming 27.4 TOPS/W on a 28nm accelerator.
-
GyRot: Leveraging Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference
Rotation and fine-grained group quantization can work together if rotation spans several quantization groups and outlier channels are permuted onto harmonic Hadamard rows, enabling 4-bit LLM inference with integer-onl...
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
AI@Meta. 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md Llama 3 model card
2024
-
[4]
AMD. 2024. Amd instinct™ mi325x accelerators. https://www.amd.com/content/dam/amd/en/documents/instinct-tech-docs/product-briefs/instinct-mi325x-datasheet.pdf
work page 2024
-
[5]
Michael Andersch, Greg Palmer, Ronny Krashinsky, Nick Stam, Vishal Mehta, Gonzalo Brito, and Sridhar Ramaswamy. 2022. Nvidia hopper architecture in-depth. https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/
work page 2022
-
[6]
Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L Croci, Bo Li, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman. 2024. Quarot: Outlier-free 4-bit inference in rotated llms. arXiv preprint arXiv:2404.00456
arXiv 2024
-
[7]
AzureAI. 2024. Azure maia for the era of ai: From silicon to software to systems. https://azure.microsoft.com/en-us/blog/azure-maia-for-the-era-of-ai-from-silicon-to-software-to-systems/
work page 2024
-
[8]
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609
arXiv 2023
Show all 79 references
-
[9]
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2024. https://doi.org/10.18653/v1/2024.acl-long.172 L ong B ench: A bilingual, multitask benchmark for long context u...
2024 doi
-
[10]
Jihwan Bang, Juntae Lee, Kyuhong Shim, Seunghan Yang, and Simyung Chang. 2024. https://doi.org/10.18653/v1/2024.acl-long.204 Crayon: Customized on-device LLM via instant adapter blending and edge-server hybrid inference . In Proceedings of the 62nd Annual Meeting of the Associ...
2024 doi
-
[11]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2019. https://arxiv.org/abs/1911.11641 Piqa: Reasoning about physical commonsense in natural language . Preprint, arXiv:1911.11641
2019 arXiv
-
[12]
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. Piqa: Reasoning about physical commonsense in natural language. In Thirty-Fourth AAAI Conference on Artificial Intelligence
2020
-
[13]
Gonzalez, Ion Stoica, and Eric P
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. https://lmsys.org/blog/2023-03-30-vicuna/ Vicuna: An open-source chatbot impressing gpt-4 with 90\
2023
-
[14]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311
2022 arXiv
-
[15]
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416
2022 arXiv
-
[16]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1
2018 arXiv
-
[17]
Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e . 2022. Flash A ttention: Fast and memory-efficient exact attention with IO -awareness. In Advances in Neural Information Processing Systems (NeurIPS)
2022
-
[18]
Bita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Ming Liu, Jeremy Fowers, Kalin Ovtcharov, Anna Vinogradsky, Sarah Massengill, Lita Yang, Ray Bittner, et al. 2020. Pushing the limits of narrow precision inferencing at cloud scale with microsoft floating point. Advances in neural...
2020
-
[19]
Bita Darvish Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, et al. 2023. With shared microexponents, a little shifting goes a long way. In Proceedings of the 50th Annua...
2023
-
[20]
Smith, and Matt Gardner
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. 2021. A dataset of information-seeking questions and answers anchored in research papers
2021
-
[21]
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022. Llm.int8(): 8-bit matrix multiplication for transformers at scale. arXiv preprint arXiv:2208.07339
2022 arXiv
-
[22]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. https://openreview.net/forum?id=OUIFPHEgJU QL o RA : Efficient finetuning of quantized LLM s . In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[23]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. https://openreview.net/forum?id=YicbFdNTTy An image is worth 1...
2021
-
[24]
Abdelfattah, and Zhiru Zhang
Jordan Dotzel, Yuzong Chen, Bahaa Kotb, Sushma Prasad, Gang Wu, Sheng Li, Mohamed S. Abdelfattah, and Zhiru Zhang. 2024. Learning from students: Applying t-distributions to explore accurate and efficient formats for llms. International Conference on Machine Learning
2024
-
[25]
Mario Drumond, Tao Lin, Martin Jaggi, and Babak Falsafi. 2018. Training dnns with hybrid block floating point. Advances in Neural Information Processing Systems, 31
2018
-
[26]
Maxim Fishman, Brian Chmiel, Ron Banner, and Daniel Soudry. 2024. https://arxiv.org/abs/2409.12517 Scaling fp8 training to trillion-token llms . Preprint, arXiv:2409.12517
2024 arXiv
-
[27]
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323
2022 arXiv
-
[28]
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027
2020 arXiv
-
[29]
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021. https://doi.org/10.5281/zenodo.53...
2021 doi
-
[30]
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. 2019. https://doi.org/10.18653/v1/D19-5409 SAMS um corpus: A human-annotated dialogue dataset for abstractive summarization . In Proceedings of the 2nd Workshop on New Frontiers in Summarization, pages 70--79, Ho...
2019 doi
-
[31]
Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian McAuley. 2023. https://arxiv.org/abs/2306.14893 Longcoder: A long-range pre-trained language model for code completion . Preprint, arXiv:2306.14893
2023 arXiv
-
[32]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. https://arxiv.org/abs/2009.03300 Measuring massive multitask language understanding . CoRR, abs/2009.03300
2020 arXiv
-
[33]
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. https://doi.org/10.18653/v1/2020.coling-main.580 Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps . In Proceedings of the 28th International Conference on Computational Li...
2020 doi
-
[34]
Mark Horowitz. 2014. Energy table for 45nm process
2014
-
[35]
Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang. 2021. https://doi.org/10.18653/v1/2021.naacl-main.112 Efficient attentions for long document summarization . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computati...
2021 doi
-
[36]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825
2023 arXiv
-
[37]
Mandar Joshi , Eunsol Choi , Daniel Weld , and Luke Zettlemoyer . 2017. https://arxiv.org/abs/1705.03551 triviaqa: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension . arXiv e-prints, arXiv:1705.03551
2017 arXiv
-
[38]
Bryan Klimt and Yiming Yang. 2004. https://api.semanticscholar.org/CorpusID:13451873 The enron corpus: A new dataset for email classi(cid:12)cation research
2004
-
[39]
Janghwan Lee, Minsoo Kim, Seungcheol Baek, Seok Hwang, Wonyong Sung, and Jungwook Choi. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.910 Enhancing computation efficiency in large language models through weight and activation quantization . In Proceedings of the 2023 Confe...
2023 doi
-
[40]
Janghwan Lee, Seongmin Park, Sukjin Hong, Minsoo Kim, Du-Seong Chang, and Jungwook Choi. 2024. https://doi.org/10.18653/v1/2024.acl-long.612 Improving conversational abilities of quantized large language models via direct preference alignment . In Proceedings of the 62nd Annua...
2024 doi
-
[41]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461
2019 arXiv
-
[42]
Xin Li and Dan Roth. 2002. https://aclanthology.org/C02-1150 Learning question classifiers . In COLING 2002: The 19th International Conference on Computational Linguistics
2002
-
[43]
Chin-Yew Lin. 2004. https://aclanthology.org/W04-1013 ROUGE : A package for automatic evaluation of summaries . In Text Summarization Branches Out, pages 74--81, Barcelona, Spain. Association for Computational Linguistics
2004
-
[44]
Haokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui, Yingtao Zhang, Linzhan Mou, Linqi Song, Zhenan Sun, and Ying Wei. 2024. https://arxiv.org/abs/2406.01721 Duquant: Distributing outliers via dual transformation makes stronger quantized llms . Preprint, arXiv:2406.01721
2024 arXiv
-
[45]
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023. Awq: Activation-aware weight quantization for llm compression and acceleration. arXiv
2023
-
[46]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023 a . https://proceedings.neurips.cc/paper_files/paper/2023/file/6dcf277ea32ce3288914faf369fe6de0-Paper-Conference.pdf Visual instruction tuning . In Advances in Neural Information Processing Systems, volume 36, pages...
2023
-
[47]
Tianyang Liu, Canwen Xu, and Julian McAuley. 2023 b . https://arxiv.org/abs/2306.03091 Repobench: Benchmarking repository-level code auto-completion systems . Preprint, arXiv:2306.03091
2023 arXiv
-
[48]
Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xucheng Yin, Cheng lin Liu, Lianwen Jin, and Xiang Bai. 2024 a . https://arxiv.org/abs/2305.07895 Ocrbench: On the hidden mystery of ocr in large multimodal models . Preprint, arXiv:2305.07895
2024 arXiv
-
[49]
Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, and Tijmen Blankevoort. 2024 b . Spinquant--llm quantization with learned rotations. arXiv preprint arXiv:2405.16406
2024 arXiv
-
[50]
S. Lloyd. 1982. https://doi.org/10.1109/TIT.1982.1056489 Least squares quantization in pcm . IEEE Transactions on Information Theory, 28(2):129--137
1982
-
[51]
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023. Llm-pruner: On the structural pruning of large language models. In Advances in Neural Information Processing Systems
2023
-
[52]
Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993. https://aclanthology.org/J93-2004 Building a large annotated corpus of E nglish: The P enn T reebank . Computational Linguistics, 19(2):313--330
1993
-
[53]
Ahmed Masry, Do Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. https://doi.org/10.18653/v1/2022.findings-acl.177 C hart QA : A benchmark for question answering about charts with visual and logical reasoning . In Findings of the Association for Computational Linguisti...
2022 doi
-
[54]
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021. https://arxiv.org/abs/2007.00398 Docvqa: A dataset for vqa on document images . Preprint, arXiv:2007.00398
2021 arXiv
-
[55]
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. https://arxiv.org/abs/1609.07843 Pointer sentinel mixture models . Preprint, arXiv:1609.07843
2016 arXiv
-
[56]
Nvidia. 2017. https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf Nvidia tesla v100 gpu architecture
2017
-
[57]
Nvidia. 2020. Nvidia a100 tensor core gpu architecture. https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/nvidia-ampere-architecture-whitepaper.pdf
2020
-
[58]
Nvidia. 2024. https://resources.nvidia.com/en-us-blackwell-architecture Nvidia blackwell architecture technical brief
2024
-
[59]
NVIDIA. 2024. Tensorrt-llm. https://github.com/NVIDIA/TensorRT-LLM
2024
-
[60]
National Library of Medicine
Courtesy of the U.S. National Library of Medicine. 2023. Pubmed. https://huggingface.co/datasets/ncbi/pubmed
2023
-
[61]
OpenAI. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
2023 arXiv
-
[62]
Bita Darvish Rouhani, Nitin Garegrat, Tom Savell, Ankit More, Kyung-Nam Han, Ritchie Zhao, Mathew Hall, Jasmine Klar, Eric Chung, Yuan Yu, Michael Schulte, Ralph Wittig, Ian Bratt, Nigel Stephens, Jelena Milanovic, John Brothers, Pradeep Dubey, Marius Cornea, Alexander Heineck...
2023
-
[63]
Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, et al. 2023 b . Microscaling data formats for deep learning. arXiv preprint arXiv:2310.10537
2023 arXiv
-
[64]
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. https://arxiv.org/abs/1907.10641 Winogrande: An adversarial winograd schema challenge at scale . Preprint, arXiv:1907.10641
2019 arXiv
-
[65]
Liu, and Christopher D
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. https://doi.org/10.18653/v1/P17-1099 Get to the point: Summarization with pointer-generator networks . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers...
2017 doi
-
[66]
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, and Ping Luo. 2024. https://openreview.net/forum?id=8Wuvhh0LYW Omniquant: Omnidirectionally calibrated quantization for large language models . In The Twelfth Internat...
2024
-
[67]
Amanpreet Singh, Vivek Natarjan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. Towards vqa models that can read. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8317--8326
2019
-
[68]
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://arxiv.org/abs/1811.00937 Commonsenseqa: A question answering challenge targeting commonsense knowledge . Preprint, arXiv:1811.00937
2019 arXiv
-
[69]
Hugo Touvron et al. 2023. https://arxiv.org/abs/2307.09288 Llama 2: Open foundation and fine-tuned chat models . Preprint, arXiv:2307.09288
2023 arXiv
-
[70]
Guangxuan Xiao, Ji Lin, Mickael Seznec, Julien Demouth, and Song Han. 2022. Smoothquant: Accurate and efficient post-training quantization for large language models. arXiv preprint arXiv:2211.10438
2022 arXiv
-
[71]
Jaewoo Yang, Hayun Kim, and Younghoon Kim. 2024. https://arxiv.org/abs/2405.14428 Mitigating quantization errors due to activation spikes in glu-based llms . Preprint, arXiv:2405.14428
2024 arXiv
-
[72]
Cohen, Ruslan Salakhutdinov, and Christopher D
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. https://arxiv.org/abs/1809.09600 Hotpotqa: A dataset for diverse, explainable multi-hop question answering . Preprint, arXiv:1809.09600
2018 arXiv
-
[73]
Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, and Jianfei Chen. 2025 a . Sageattention2: Efficient attention with thorough outlier smoothing and per-thread int4 quantization. In International Conference on Machine Learning (ICML)
2025
-
[74]
Jintao Zhang, Jia Wei, Pengle Zhang, Xiaoming Xu, Haofeng Huang, Haoxu Wang, Kai Jiang, Jun Zhu, and Jianfei Chen. 2025 b . Sageattention3: Microscaling fp4 attention for inference and an exploration of 8-bit training. arXiv preprint arXiv:2505.11594
2025
-
[75]
Jintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu, and Jianfei Chen. 2025 c . Sageattention: Accurate 8-bit attention for plug-and-play inference acceleration. In International Conference on Learning Representations (ICLR)
2025
-
[76]
Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, Jingkang Yang, Chunyuan Li, and Ziwei Liu. 2024 a . https://arxiv.org/abs/2407.12772 Lmms-eval: Reality check on the evaluation of large multimodal models . Preprint, arX...
2024 arXiv
- [77]
-
[78]
Yijia Zhang, Sicheng Zhang, Shijie Cao, DaYou Du, Jianyu Wei, Ting Cao, and Ningyi Xu. 2024 b . https://doi.org/10.18653/v1/2024.findings-acl.3 AFPQ : Asymmetric floating point quantization for LLM s . In Findings of the Association for Computational Linguistics ACL 2024, page...
2024 doi
-
[79]
Gonzalez, and Ion Stoica
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. https://openreview.net/forum?id=uccHPGDlao Judging LLM -as-a-judge with MT -bench and chatbot ...
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.