Pith. sign in

REVIEW 4 cited by

Investigating the Effectiveness of Representations Based on Pretrained Transformer-based Language Models in Active Learning for Labelling Text Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.13138 v1 pith:7ZYQOAFA submitted 2020-04-21 cs.IR cs.CLcs.LG

classification cs.IRcs.CLcs.LG
keywords learningactivemodelsrepresentationseffectivenesslanguagerepresentationtransformer-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Active learning has been shown to be an effective way to alleviate some of the effort required in utilising large collections of unlabelled data for machine learning tasks without needing to fully label them. The representation mechanism used to represent text documents when performing active learning, however, has a significant influence on how effective the process will be. While simple vector representations such as bag-of-words and embedding-based representations based on techniques such as word2vec have been shown to be an effective way to represent documents during active learning, the emergence of representation mechanisms based on the pre-trained transformer-based neural network models popular in natural language processing research (e.g. BERT) offer a promising, and as yet not fully explored, alternative. This paper describes a comprehensive evaluation of the effectiveness of representations based on pre-trained transformer-based language models for active learning. This evaluation shows that transformer-based models, especially BERT-like models, that have not yet been widely used in active learning, achieve a significant improvement over more commonly used vector representations like bag-of-words or other classical word embeddings like word2vec. This paper also investigates the effectiveness of representations based on variants of BERT such as Roberta, Albert as well as comparing the effectiveness of the [CLS] token representation and the aggregated representation that can be generated using BERT-like models. Finally, we propose an approach Adaptive Tuning Active Learning. Our experiments show that the limited label information acquired in active learning can not only be used for training a classifier but can also adaptively improve the embeddings generated by the BERT-like language models as well.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. XOR Games at Full Tilt: The Hardness of Binary Nonlocal Games

    quant-ph 2026-07 accept novelty 7.0 of 10

    Approximating the quantum value of tilted XOR games to constant precision is RE-complete, implying binary nonlocal games are RE-hard to approximate.

  2. Parallel repetition of expanded, and multiplayer, Quantum games: anchoring, optimal values, generalized error bounds, dependency-breaking as symmetry-breaking

    quant-ph 2025-08 reject novelty 4.0 of 10

    States exponential decay bounds for the anchored parallel-repeated value of multiplayer quantum games with N-dependent exponents, but leaves the N-player proof to prior work.

  3. Probability distributions over CSS codes: two-universality, QKD hashing, collision bounds, security

    quant-ph 2025-10 reject novelty 3.0 of 10

    A two-universal QKD hashing protocol is claimed to be 2^{−k/2 + n h_2(r/n) + 35/4 + log_2√C}-secure for an unspecified constant C — a strictly weaker bound than Ostrev's 2^{−k/2 + n h(r/n) + 5/2}, obtained by adapting...

  4. Error correction, authentication, and false acceptance, probabilities for communication over noisy quantum channels: converse upper bounds on the bit transmission rate

    quant-ph 2025-07 reject novelty 2.0 of 10

    The paper asserts a converse upper bound on the bit transmission rate in terms of pruned alphabet sizes, but the proof is a chain of unjustified inequalities and the main theorem reverses the inequality of the result ...

Pith tools