REVIEW 14 cited by
Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generative AI offers a simple, prompt-based alternative to fine-tuning smaller BERT-style LLMs for text classification tasks. This promises to eliminate the need for manually labeled training data and task-specific model training. However, it remains an open question whether tools like ChatGPT can deliver on this promise. In this paper, we show that smaller, fine-tuned LLMs (still) consistently and significantly outperform larger, zero-shot prompted models in text classification. We compare three major generative AI models (ChatGPT with GPT-3.5/GPT-4 and Claude Opus) with several fine-tuned LLMs across a diverse set of classification tasks (sentiment, approval/disapproval, emotions, party positions) and text categories (news, tweets, speeches). We find that fine-tuning with application-specific training data achieves superior performance in all cases. To make this approach more accessible to a broader audience, we provide an easy-to-use toolkit alongside this paper. Our toolkit, accompanied by non-technical step-by-step guidance, enables users to select and fine-tune BERT-like LLMs for any classification task with minimal technical and computational effort.
Forward citations
Cited by 14 Pith papers
-
DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain
DLT-Corpus is a 2.98B-token scientific/patent/Twitter corpus for blockchain NLP, plus LedgerBERT (+23% NER vs BERT), a sentiment dataset, and cross-domain innovation-diffusion analyses.
-
PerSoMed: A Large-Scale Balanced Dataset for Persian Social Media Text Classification
PerSoMed, a balanced nine-class dataset of 36,000 Persian social media posts, with benchmark results showing TookaBERT-Large at F1 0.962.
-
Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
A few-shot RLVR method using only rule-based rewards lifts a 2B vision-language model's remote sensing accuracy by double digits, with 128 examples rivaling thousands.
-
Evaluating Apple Intelligence's Writing Tools for Privacy Against Large Language Model-Based Inference Attacks: Insights from Early Datasets
Apple Intelligence's Friendly and Professional text rewrites can substantially reduce LLM-based emotion inference accuracy in small early datasets.
-
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
ReSpace is an autoregressive LLM framework for text-driven 3D indoor scene editing and synthesis, using a structured JSON scene representation and a voxelization-based layout metric.
-
Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification
LLMs show systematic target-dependent sentiment inconsistency that is politically biased: left and center politicians rated more positively, far-right politicians more negatively, with stronger effects in larger model...
-
Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning
Citss combines sentence-level cropping and keyphrase perturbation with contrastive learning to fine-tune both encoder and decoder language models for citation classification.
-
Multi-Lingual Implicit Discourse Relation Recognition with Multi-Label Hierarchical Learning
HArch, a hierarchical multi-task model, is the first to recognize implied discourse relations with multi-label sense distributions in four languages, and it outperforms few-shot LLM prompting.
-
A Multi-Stage Large Language Model Framework for Extracting Suicide-Related Social Determinants of Health
A multi-stage LLM pipeline improves extraction of infrequent suicide-related social determinants from death narratives, but some evaluation results are compromised by using the test set to tune the system.
-
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
For multilingual RAG intent classification, the best translation strategy depends on the model and language; translating instructions into the user's language helps some models, while making the model answer in low-re...
-
Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
On three hate speech datasets, including a new code-mixed IndoHateMix benchmark, fine-tuned LLMs such as LLaMA-3.1 beat multilingual BERT models, with the largest gains on code-mixed Indian text.
-
AIx4Soccer: A Unified Platform Architecture for Football Club Management and Structured Athlete Development
A conceptual multi-tenant football-club SaaS design embeds a PDI development cycle and a 75/25 certified video-analyst marketplace on a proposed event-sourced knowledge-graph substrate.
-
A Scalable and Efficient Signal Integration System for Job Matching
STAR integrates fine-tuned LLM embeddings as node features into a large-scale GNN, improving job matching metrics across three LinkedIn products.
-
DS@GT at CheckThat! 2025: Ensemble Methods for Detection of Scientific Discourse on Social Media
The DS@GT system achieved 0.8611 macro-F1 on the CheckThat! 2025 Task 4a development set by combining a fine-tuned DeBERTa model with GPT-4o few-shot prompting, outperforming the DeBERTaV3 baseline of 0.8375.
Discussion (0). Continue with ORCID to comment.