REVIEW 3 cited by
Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Topic models extract groups of words from documents, whose interpretation as a topic hopefully allows for a better understanding of the data. However, the resulting word groups are often not coherent, making them harder to interpret. Recently, neural topic models have shown improvements in overall coherence. Concurrently, contextual embeddings have advanced the state of the art of neural models in general. In this paper, we combine contextualized representations with neural topic models. We find that our approach produces more meaningful and coherent topics than traditional bag-of-words topic models and recent neural models. Our results indicate that future improvements in language models will translate into better topic models.
Forward citations
Cited by 3 Pith papers
-
Continual Neural Topic Model
CoNTM is an online neural topic model whose global topics are a running average of time-slice topics; the paper's no-forgetting claim is asserted without a forgetting test.
-
NGTM: Substructure-based Neural Graph Topic Model for Interpretable Graph Generation
NGTM generates graphs by sampling substructures from learned topic-specific distributions and assembling them, achieving competitive quality with interpretable, controllable topics.
-
Understanding Cross-Domain Adaptation in Low-Resource Topic Modeling
DALTA adapts a variational topic model from a high-resource source domain to a low-resource target domain via adversarial latent alignment, separate decoders, and a consistency loss, with a claimed generalization bound.
Discussion (0). Continue with ORCID to comment.