Pith. sign in

REVIEW 3 cited by

Accuracy is Not All You Need

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.09141 v1 pith:BRVJSD2A submitted 2024-07-12 cs.LG

classification cs.LG
keywords compressedmodelsmodelaccuracybaselinetechniquesmetricswhen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When Large Language Models (LLMs) are compressed using techniques such as quantization, the predominant way to demonstrate the validity of such techniques is by measuring the model's accuracy on various benchmarks.If the accuracies of the baseline model and the compressed model are close, it is assumed that there was negligible degradation in quality.However, even when the accuracy of baseline and compressed model are similar, we observe the phenomenon of flips, wherein answers change from correct to incorrect and vice versa in proportion.We conduct a detailed study of metrics across multiple compression techniques, models and datasets, demonstrating that the behavior of compressed models as visible to end-users is often significantly different from the baseline model, even when accuracy is similar.We further evaluate compressed models qualitatively and quantitatively using MT-Bench and show that compressed models are significantly worse than baseline models in this free-form generative task.Thus, we argue that compression techniques should also be evaluated using distance metrics.We propose two such metrics, KL-Divergence and flips, and show that they are well correlated.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Variable Bit-width Quantization: Learning Per-Group Precision for "Bigger-but-Smaller" Language Models

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Learned per-group bit-widths yield a reusable low-bit recipe that makes language models simultaneously larger in parameters and smaller in storage than FP16 baselines, with growing decode speedups.

  2. HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An integrated black-box audit protocol profiles parametric hindsight in LLMs and shows the date-trigger reflex tracks training generation, not scale, while effective knowledge cutoffs span 22 months.

  3. Reliability Scaling Laws for Quantized Large Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Reliability of quantized LLMs peaks nonlinearly at 4-bit precision under fixed total model bits, while accuracy scales monotonically, and quantization can improve robustness to natural perturbations.

Pith tools