Pith. sign in

REVIEW 14 cited by

Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08660 v2 pith:6UWFNICG submitted 2024-06-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords classificationllmstextfine-tunedgenerativemodelstrainingchatgpt
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative AI offers a simple, prompt-based alternative to fine-tuning smaller BERT-style LLMs for text classification tasks. This promises to eliminate the need for manually labeled training data and task-specific model training. However, it remains an open question whether tools like ChatGPT can deliver on this promise. In this paper, we show that smaller, fine-tuned LLMs (still) consistently and significantly outperform larger, zero-shot prompted models in text classification. We compare three major generative AI models (ChatGPT with GPT-3.5/GPT-4 and Claude Opus) with several fine-tuned LLMs across a diverse set of classification tasks (sentiment, approval/disapproval, emotions, party positions) and text categories (news, tweets, speeches). We find that fine-tuning with application-specific training data achieves superior performance in all cases. To make this approach more accessible to a broader audience, we provide an easy-to-use toolkit alongside this paper. Our toolkit, accompanied by non-technical step-by-step guidance, enables users to select and fine-tune BERT-like LLMs for any classification task with minimal technical and computational effort.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain

    cs.CL 2026-02 conditional novelty 6.0 of 10

    DLT-Corpus is a 2.98B-token scientific/patent/Twitter corpus for blockchain NLP, plus LedgerBERT (+23% NER vs BERT), a sentiment dataset, and cross-domain innovation-diffusion analyses.

  2. PerSoMed: A Large-Scale Balanced Dataset for Persian Social Media Text Classification

    cs.CL 2026-02 conditional novelty 6.0 of 10

    PerSoMed, a balanced nine-class dataset of 36,000 Persian social media posts, with benchmark results showing TookaBERT-Large at F1 0.962.

  3. Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A few-shot RLVR method using only rule-based rewards lifts a 2B vision-language model's remote sensing accuracy by double digits, with 128 examples rivaling thousands.

  4. Evaluating Apple Intelligence's Writing Tools for Privacy Against Large Language Model-Based Inference Attacks: Insights from Early Datasets

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Apple Intelligence's Friendly and Professional text rewrites can substantially reduce LLM-based emotion inference accuracy in small early datasets.

  5. ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ReSpace is an autoregressive LLM framework for text-driven 3D indoor scene editing and synthesis, using a structured JSON scene representation and a voxelization-based layout metric.

  6. Analyzing Political Bias in LLMs via Target-Oriented Sentiment Classification

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLMs show systematic target-dependent sentiment inconsistency that is politically biased: left and center politicians rated more positively, far-right politicians more negatively, with stronger effects in larger model...

  7. Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Citss combines sentence-level cropping and keyphrase perturbation with contrastive learning to fine-tune both encoder and decoder language models for citation classification.

  8. Multi-Lingual Implicit Discourse Relation Recognition with Multi-Label Hierarchical Learning

    cs.CL 2025-08 conditional novelty 5.0 of 10

    HArch, a hierarchical multi-task model, is the first to recognize implied discourse relations with multi-label sense distributions in four languages, and it outperforms few-shot LLM prompting.

  9. A Multi-Stage Large Language Model Framework for Extracting Suicide-Related Social Determinants of Health

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A multi-stage LLM pipeline improves extraction of infrequent suicide-related social determinants from death narratives, but some evaluation results are compromised by using the test set to tune the system.

  10. How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting

    cs.CL 2025-07 conditional novelty 5.0 of 10

    For multilingual RAG intent classification, the best translation strategy depends on the model and language; translating instructions into the user's language helps some models, while making the model answer in low-re...

  11. Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    On three hate speech datasets, including a new code-mixed IndoHateMix benchmark, fine-tuned LLMs such as LLaMA-3.1 beat multilingual BERT models, with the largest gains on code-mixed Indian text.

  12. AIx4Soccer: A Unified Platform Architecture for Football Club Management and Structured Athlete Development

    cs.CY 2026-07 conditional novelty 4.0 of 10

    A conceptual multi-tenant football-club SaaS design embeds a PDI development cycle and a 75/25 certified video-analyst marketplace on a proposed event-sourced knowledge-graph substrate.

  13. A Scalable and Efficient Signal Integration System for Job Matching

    cs.LG 2025-07 conditional novelty 4.0 of 10

    STAR integrates fine-tuned LLM embeddings as node features into a large-scale GNN, improving job matching metrics across three LinkedIn products.

  14. DS@GT at CheckThat! 2025: Ensemble Methods for Detection of Scientific Discourse on Social Media

    cs.CL 2025-07 conditional novelty 4.0 of 10

    The DS@GT system achieved 0.8611 macro-F1 on the CheckThat! 2025 Task 4a development set by combining a fine-tuned DeBERTa model with GPT-4o few-shot prompting, outperforming the DeBERTaV3 baseline of 0.8375.

Pith tools