REVIEW 3 cited by
MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As major progress is made in open-ended text generation, measuring how close machine-generated text is to human language remains a critical open problem. We introduce MAUVE, a comparison measure for open-ended text generation, which directly compares the learnt distribution from a text generation model to the distribution of human-written text using divergence frontiers. MAUVE scales up to modern text generation models by computing information divergences in a quantized embedding space. Through an extensive empirical study on three open-ended generation tasks, we find that MAUVE identifies known properties of generated text, scales naturally with model size, and correlates with human judgments, with fewer restrictions than existing distributional evaluation metrics.
Forward citations
Cited by 3 Pith papers
-
Just on Time: Token-Level Early Stopping for Diffusion Language Models
Jot, a token-level early stopping rule using a top-2 confidence ratio and spatial context, speeds up diffusion language model decoding by up to 19.6x with minor quality loss.
-
IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
IDEAlign uses a pick-the-odd-one-out triplet task to measure idea-level similarity between LLMs and expert human annotations, and shows LLM judges using this protocol outperform lexical and vector-based baselines.
-
Multi-Hypothesis Distillation of Multilingual Neural Translation Models for Low-Resource Languages
Generating several candidate translations per source sentence for knowledge distillation yields better small multilingual translators than standard single-hypothesis distillation, especially in low-resource settings.
Discussion (0). Continue with ORCID to comment.