REVIEW 6 cited by
TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The experimental landscape in natural language processing for social media is too fragmented. Each year, new shared tasks and datasets are proposed, ranging from classics like sentiment analysis to irony detection or emoji prediction. Therefore, it is unclear what the current state of the art is, as there is no standardized evaluation protocol, neither a strong set of baselines trained on such domain-specific data. In this paper, we propose a new evaluation framework (TweetEval) consisting of seven heterogeneous Twitter-specific classification tasks. We also provide a strong set of baselines as starting point, and compare different language modeling pre-training strategies. Our initial experiments show the effectiveness of starting off with existing pre-trained generic language models, and continue training them on Twitter corpora.
Forward citations
Cited by 6 Pith papers
-
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
An early-exit rule with a zero-shot fallback, calibrated by Learn-then-Test risk control, keeps the average loss from corrupted in-context demonstrations under a preset bound.
-
Brain2Model Transfer: Training sensory and decision models with human neural activity as a teacher
A new transfer-learning framework, Brain2Model, uses human neural recordings to shape the latent representations of artificial networks, improving test accuracy in two tasks.
-
Social Contagion in COVID-19 Discussions within the Belgian Reddit Community: A Statistical and Modeling Study
In r/Belgium, COVID-19 topics were seeded by external events, not by prior posts, but comment sentiment was contagious, and a two-layer bounded confidence model best captured that asymmetry.
-
Confidence Optimization for Probabilistic Encoding
A confidence-aware loss plus a negative L2 variance term gives a small boost to probabilistic encoding classifiers on TweetEval, but gains over the SPC baseline are modest and inconsistent on RoBERTa.
-
From BERT to Qwen: Hate Detection across architectures
Fine-tuned Qwen-0.5B slightly edges out DistilBERT and RoBERTa on a balanced hate speech corpus, and few-shot prompting substantially improves Gemma-3-1B over zero-shot.
-
Involvement drives complexity of language in online debates
Influential Twitter users who are more partisan, negative, or offensive tend to use more lexically complex language, but the causal claim that involvement drives complexity is not supported.
Discussion (0). Continue with ORCID to comment.