REVIEW 4 cited by
Investigating the Effectiveness of Representations Based on Pretrained Transformer-based Language Models in Active Learning for Labelling Text Datasets
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Active learning has been shown to be an effective way to alleviate some of the effort required in utilising large collections of unlabelled data for machine learning tasks without needing to fully label them. The representation mechanism used to represent text documents when performing active learning, however, has a significant influence on how effective the process will be. While simple vector representations such as bag-of-words and embedding-based representations based on techniques such as word2vec have been shown to be an effective way to represent documents during active learning, the emergence of representation mechanisms based on the pre-trained transformer-based neural network models popular in natural language processing research (e.g. BERT) offer a promising, and as yet not fully explored, alternative. This paper describes a comprehensive evaluation of the effectiveness of representations based on pre-trained transformer-based language models for active learning. This evaluation shows that transformer-based models, especially BERT-like models, that have not yet been widely used in active learning, achieve a significant improvement over more commonly used vector representations like bag-of-words or other classical word embeddings like word2vec. This paper also investigates the effectiveness of representations based on variants of BERT such as Roberta, Albert as well as comparing the effectiveness of the [CLS] token representation and the aggregated representation that can be generated using BERT-like models. Finally, we propose an approach Adaptive Tuning Active Learning. Our experiments show that the limited label information acquired in active learning can not only be used for training a classifier but can also adaptively improve the embeddings generated by the BERT-like language models as well.
Forward citations
Cited by 4 Pith papers
-
XOR Games at Full Tilt: The Hardness of Binary Nonlocal Games
Approximating the quantum value of tilted XOR games to constant precision is RE-complete, implying binary nonlocal games are RE-hard to approximate.
-
Parallel repetition of expanded, and multiplayer, Quantum games: anchoring, optimal values, generalized error bounds, dependency-breaking as symmetry-breaking
States exponential decay bounds for the anchored parallel-repeated value of multiplayer quantum games with N-dependent exponents, but leaves the N-player proof to prior work.
-
Probability distributions over CSS codes: two-universality, QKD hashing, collision bounds, security
A two-universal QKD hashing protocol is claimed to be 2^{−k/2 + n h_2(r/n) + 35/4 + log_2√C}-secure for an unspecified constant C — a strictly weaker bound than Ostrev's 2^{−k/2 + n h(r/n) + 5/2}, obtained by adapting...
-
Error correction, authentication, and false acceptance, probabilities for communication over noisy quantum channels: converse upper bounds on the bit transmission rate
The paper asserts a converse upper bound on the bit transmission rate in terms of pruned alphabet sizes, but the proof is a chain of unjustified inequalities and the main theorem reverses the inequality of the result ...
Discussion (0). Continue with ORCID to comment.