Pith. sign in

REVIEW 2 cited by

A thorough benchmark of automatic text classification: From traditional approaches to large language models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.01930 v1 pith:J42P5MKT submitted 2025-04-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmstraditionalapproachesclassificationeffectivenesslargerecentslms
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Automatic text classification (ATC) has experienced remarkable advancements in the past decade, best exemplified by recent small and large language models (SLMs and LLMs), leveraged by Transformer architectures. Despite recent effectiveness improvements, a comprehensive cost-benefit analysis investigating whether the effectiveness gains of these recent approaches compensate their much higher costs when compared to more traditional text classification approaches such as SVMs and Logistic Regression is still missing in the literature. In this context, this work's main contributions are twofold: (i) we provide a scientifically sound comparative analysis of the cost-benefit of twelve traditional and recent ATC solutions including five open LLMs, and (ii) a large benchmark comprising {22 datasets}, including sentiment analysis and topic classification, with their (train-validation-test) partitions based on folded cross-validation procedures, along with documentation, and code. The release of code, data, and documentation enables the community to replicate experiments and advance the field in a more scientifically sound manner. Our comparative experimental results indicate that LLMs outperform traditional approaches (up to 26%-7.1% on average) and SLMs (up to 4.9%-1.9% on average) in terms of effectiveness. However, LLMs incur significantly higher computational costs due to fine-tuning, being, on average 590x and 8.5x slower than traditional methods and SLMs, respectively. Results suggests the following recommendations: (1) LLMs for applications that require the best possible effectiveness and can afford the costs; (2) traditional methods such as Logistic Regression and SVM for resource-limited applications or those that cannot afford the cost of tuning large LLMs; and (3) SLMs like Roberta for near-optimal effectiveness-efficiency trade-off.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Phishing Email Detection Performance of Small Large Language Models

    cs.CL 2025-04 conditional novelty 4.0 of 10

    Explanation-augmented LoRA fine-tuning lets small LLMs detect phishing emails with accuracy and F1 around 0.94 to 0.96 on the SpamAssassin test set, while transferring to unseen datasets.

  2. Ranking-based Fusion Algorithms for Extreme Multi-label Text Classification (XMTC)

    cs.IR 2025-07 reject novelty 3.0 of 10

    Fusing BM25 and BERT rankings with CombMNZ and ZMUV normalization gave the best reported nDCG and precision on four XMTC datasets, though no single-retriever baselines or significance results are reported.

Pith tools