REVIEW 11 cited by
Neural Text Generation with Unlikelihood Training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Neural text generation is a key tool in natural language applications, but it is well known there are major problems at its core. In particular, standard likelihood training and decoding leads to dull and repetitive outputs. While some post-hoc fixes have been proposed, in particular top-$k$ and nucleus sampling, they do not address the fact that the token-level probabilities predicted by the model are poor. In this paper we show that the likelihood objective itself is at fault, resulting in a model that assigns too much probability to sequences containing repeats and frequent words, unlike those from the human training distribution. We propose a new objective, unlikelihood training, which forces unlikely generations to be assigned lower probability by the model. We show that both token and sequence level unlikelihood training give less repetitive, less dull text while maintaining perplexity, giving superior generations using standard greedy or beam search. According to human evaluations, our approach with standard beam search also outperforms the currently popular decoding methods of nucleus sampling or beam blocking, thus providing a strong alternative to existing techniques.
Forward citations
Cited by 11 Pith papers
-
MentalThink: Shaping Thoughts in Mental SVG World
MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.
-
On the Fitness Landscape in the $NK$ Model
For the NK fitness landscape with K/N tending to alpha, exact limits for free energy and maximum fitness are identified, together with the geometry of near-fittest peaks.
-
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
A token-level correctness classifier trained with LoRA and then merged into the model boosts out-of-distribution factuality in summarization and translation.
-
Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models
A closed-form ratio of binomial conditional probabilities, comparing a token's recent count with its corpus rate, defines a self-normalizing logit bias that suppresses repetitive loops in autoregressive language models.
-
LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure
Preserving adaptive non-label local structure from the base model during SFT improves the pass@1 vs pass@k trade-off and reduces catastrophic forgetting versus vanilla cross-entropy and recent SFT variants.
-
Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS
LENS measures narrative unlearning at four resistance levels and shows existing objectives can suppress a target frame at selected middle checkpoints, with transfer to attributed/contrastive prompts and rare residual ...
-
Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding Exploration
Soft Reasoning improves LLM accuracy on reasoning benchmarks by perturbing the embedding of the first generated token and searching the perturbation space with Bayesian optimization guided by the model's own verifier.
-
Avoidance Decoding for Diverse Multi-Branch Story Generation
Avoidance Decoding penalizes token choices that resemble previously generated story branches, using a hybrid concept-level and narrative-level similarity penalty, and reports large diversity gains across several LLMs.
-
Breaking the Likelihood Trap: Consistent Generative Recommendation with Graph-structured Model
CONGRATS uses a DAG-structured positional decoder and evaluator-in-the-loop training to generate more diverse and accurate recommendation lists, showing offline and Kuaishou A/B gains.
-
GOLFer: Smaller LM-Generated Documents Hallucination Filter & Combiner for Query Expansion in Information Retrieval
GOLFer filters hallucinated sentences from small-LM-generated hypothetical documents and reweights the rest into the query, improving retrieval at lower cost than large LLM expansion.
-
Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation
A thesis proposal repurposing two prior papers on LM agents for text games, framed as a path to theory-of-mind AI, with no new theory-of-mind evidence.
Discussion (0). Continue with ORCID to comment.