Pith. sign in

REVIEW 9 cited by

Representation Degeneration Problem in Training Natural Language Generation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.12009 v1 pith:34ZZCEDE submitted 2019-07-28 cs.CL

classification cs.CL
keywords problemlanguagerepresentationtrainingdegenerationgenerationnaturalembeddings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study an interesting problem in training neural network-based models for natural language generation tasks, which we call the \emph{representation degeneration problem}. We observe that when training a model for natural language generation tasks through likelihood maximization with the weight tying trick, especially with big training datasets, most of the learnt word embeddings tend to degenerate and be distributed into a narrow cone, which largely limits the representation power of word embeddings. We analyze the conditions and causes of this problem and propose a novel regularization method to address it. Experiments on language modeling and machine translation show that our method can largely mitigate the representation degeneration problem and achieve better performance than baseline algorithms.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation

    cs.IR 2026-07 conditional novelty 6.0 of 10

    SpecFormer is a spectral-aware Transformer that flattens the singular-value spectrum of embeddings to prevent embedding/attention collapse, outperforming baselines on CTR benchmarks and scaling with layer depth.

  2. Reliability Scaling Laws for Quantized Large Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Reliability of quantized LLMs peaks nonlinearly at 4-bit precision under fixed total model bits, while accuracy scales monotonically, and quantization can improve robustness to natural perturbations.

  3. NITP: Next Implicit Token Prediction for LLM Pre-training

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    NITP augments standard next-token prediction with implicit semantic prediction in representation space using shallow-layer self-supervision, reporting consistent downstream gains on 0.5B-9B models including 5.7% on MM...

  4. Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning

    cs.CL 2026-01 conditional novelty 6.0 of 10

    PALU shows that unlearning only needs local intervention—the first few tokens of the sensitive span and the top-k logits—not full-sequence, full-vocabulary suppression.

  5. Scaling Point-in-Time Language Models

    cs.CL 2026-04 conditional novelty 5.5 of 10

    Scaling point-in-time LLMs to 4B parameters and 1T temporally filtered tokens narrows the gap to unrestricted models to about 8–11 average points and yields positive out-of-sample Sharpe ratios from news embeddings.

  6. ChaosProbe: A Neurochaotic Lens on Frozen Transformer Input-Embedding Spaces

    cs.LG 2026-08 conditional novelty 5.0 of 10

    In a four-model proof of concept, deterministic neurochaotic fingerprints of frozen transformer embedding tables place GPT-2/DistilGPT2 and BERT/RoBERTa as mutual nearest neighbors under Pearson, Spearman, and cosine ...

  7. SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Using multiple LLM-generated summaries of the same abstract as positive pairs trains scientific text embeddings that beat citation-trained baselines on retrieval and clustering, while the new benchmark shares its trai...

  8. Low-Perplexity LLM-Generated Sequences and Where To Find Them

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Only about 40% of low-perplexity 6-token spans generated by Pythia-6.9B can be exactly matched to The Pile, and the authors categorize matched and unmatched spans into four classes.

  9. Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation

    cs.IR 2025-09 conditional novelty 4.0 of 10

    MambaRec improves multimodal recommendation accuracy on Baby, Sports, and Clothing datasets through local dilated-attention alignment and global MMD/contrastive alignment.

Pith tools