REVIEW 6 cited by
News Category Dataset
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
People rely on news to know what is happening around the world and inform their daily lives. In today's world, when the proliferation of fake news is rampant, having a large-scale and high-quality source of authentic news articles with the published category information is valuable to learning authentic news' Natural Language syntax and semantics. As part of this work, we present a News Category Dataset that contains around 210k news headlines from the year 2012 to 2022 obtained from HuffPost, along with useful metadata to enable various NLP tasks. In this paper, we also produce some novel insights from the dataset and describe various existing and potential applications of our dataset.
Forward citations
Cited by 6 Pith papers
-
Extrapolated Markov Chain Oversampling Method for Imbalanced Text Classification
EMCO oversamples minority text by estimating word-transition probabilities from both minority and majority documents, expanding the synthetic minority vocabulary.
-
TI-StegoAlign: Channel-Guided Post-Training for Generative Text Steganography under Tokenization Inconsistency
A margin-supervision plus receiver-realistic preference post-training method achieves 100% receiver-side bit accuracy and 21.6% lower normalized perplexity deviation than the strongest baseline.
-
MLego: Interactive and Scalable Topic Exploration Through Model Reuse
MLego reuses and merges materialized LDA models to answer ad-hoc topic queries quickly, using hierarchical plan search and batch reordering to keep the cost low.
-
Political Leaning and Politicalness Classification of Texts
The authors compile large multi-dataset benchmarks for political leaning and politicalness classification, show that single-dataset models fail out-of-distribution, and release new models with improved cross-domain F1 scores.
-
Dataset of News Articles with Provenance Metadata for Media Relevance Assessment
A new benchmark dataset and two tasks let researchers test whether AI systems can judge if a news image's recorded location and date match the article, with current chatbots scoring 64-81% on location but 42-58% on date.
-
Investigating Algorithmic Bias in YouTube Shorts
YouTube Shorts recommendations from political seeds drift to entertainment and positive-emotion content within the first few steps, with the drift unchanged by simulated watch-time.
Discussion (0). Continue with ORCID to comment.