Pith. sign in

REVIEW 6 cited by

How to Fine-Tune BERT for Text Classification?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.05583 v3 pith:VN6PAQ3R submitted 2019-05-14 cs.CL

classification cs.CL
keywords bertlanguageclassificationmodeltextfine-tuningpre-trainingrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language model pre-training has proven to be useful in learning universal language representations. As a state-of-the-art language model pre-training model, BERT (Bidirectional Encoder Representations from Transformers) has achieved amazing results in many language understanding tasks. In this paper, we conduct exhaustive experiments to investigate different fine-tuning methods of BERT on text classification task and provide a general solution for BERT fine-tuning. Finally, the proposed solution obtains new state-of-the-art results on eight widely-studied text classification datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 95 citations worldwide. Full citation record

  1. Political Leaning and Politicalness Classification of Texts

    cs.CL 2025-07 conditional novelty 6.0 of 10

    The authors compile large multi-dataset benchmarks for political leaning and politicalness classification, show that single-dataset models fail out-of-distribution, and release new models with improved cross-domain F1 scores.

  2. Tiny Reward Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    TinyRM shows that 400M-parameter bidirectional masked language models, tuned with FLAN-style prompting, DoRA, and layer freezing, outperform a 70B reward model on RewardBench reasoning and come close on safety.

  3. Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts

    cs.CL 2025-09 conditional novelty 4.0 of 10

    On SemEval 2025 Task 11 short texts, generated data helped some BERT models, continued pretraining was mixed, and classification head changes barely mattered.

  4. Leveraging Machine Learning and Enhanced Parallelism Detection for BPMN Model Generation from Text

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A BPMN extraction pipeline using BERT, RoBERTa, CRF, and CatBoost is evaluated on an augmented PET dataset, with 15 new documents adding 32 AND gateways to improve parallel-structure detection.

  5. Lowering the Barrier of Machine Learning: Achieving Zero Manual Labeling in Review Classification Using LLMs

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Using GPT-4 to generate training labels and a domain-adapted BERT encoder, the pipeline classifies review sentiment with 80.8 to 88.6 percent accuracy and no manual annotation.

  6. Optimising Language Models for Downstream Tasks: A Post-Training Perspective

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A dissertation that repackages the author's previously published papers on continued pre-training, prompt tuning, and instruction modelling into a single narrative.

Pith tools