REVIEW 6 cited by
Open, Closed, or Small Language Models for Text Classification?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent advancements in large language models have demonstrated remarkable capabilities across various NLP tasks. But many questions remain, including whether open-source models match closed ones, why these models excel or struggle with certain tasks, and what types of practical procedures can improve performance. We address these questions in the context of classification by evaluating three classes of models using eight datasets across three distinct tasks: named entity recognition, political party prediction, and misinformation detection. While larger LLMs often lead to improved performance, open-source models can rival their closed-source counterparts by fine-tuning. Moreover, supervised smaller models, like RoBERTa, can achieve similar or even greater performance in many datasets compared to generative LLMs. On the other hand, closed models maintain an advantage in hard tasks that demand the most generalizability. This study underscores the importance of model selection based on task requirements
Forward citations
Cited by 6 Pith papers
-
Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models
PiFi adds one frozen LLM layer to an SLM and fine-tunes it, reporting consistent but modest gains across NLU and NLG tasks, with larger gains when the LLM matches the target language.
-
IYKYK: Using language models to decode extremist cryptolects
LLMs struggle with extremist in-group jargon, but prompting with example posts and domain-adapting encoders substantially improves detection and decoding.
-
Large Language Models in the Task of Automatic Validation of Text Classifier Predictions
LLM-based annotators using token-probability thresholds, RAG, and reasoning fine-tuning matched or exceeded human annotator quality on a proprietary 250-class intent-validation task.
-
SOI Matters: Analyzing Multi-Setting Training Dynamics in Pretrained Language Models via Subsets of Interest
A new fine-grained taxonomy of example-level learning dynamics (SOI) is applied to multi-task, multi-source, and multi-lingual fine-tuning, showing robust OOD gains for multi-source training and small gains from SOI-g...
-
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
A single BERT-scale model with a game-context token and LLM-assisted label transfer achieves toxicity detection comparable to per-game models while extending to seven languages.
-
Rethinking the Understanding Ability across LLMs through Mutual Information
The paper uses token-level recoverability as a computable lower bound on mutual information to compare LLMs and to fine-tune them, finding encoder-only models preserve information better than decoder-only models.
Discussion (0). Sign in to comment.