Pith. sign in

REVIEW 2 cited by

LERT: A Linguistically-motivated Pre-trained Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.05344 v1 pith:LPAES6UM submitted 2022-11-10 cs.CL cs.LG

classification cs.CLcs.LG
keywords languagelertmodellinguisticpre-trainedfeaturespre-trainingeffective
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Pre-trained Language Model (PLM) has become a representative foundation model in the natural language processing field. Most PLMs are trained with linguistic-agnostic pre-training tasks on the surface form of the text, such as the masked language model (MLM). To further empower the PLMs with richer linguistic features, in this paper, we aim to propose a simple but effective way to learn linguistic features for pre-trained language models. We propose LERT, a pre-trained language model that is trained on three types of linguistic features along with the original MLM pre-training task, using a linguistically-informed pre-training (LIP) strategy. We carried out extensive experiments on ten Chinese NLU tasks, and the experimental results show that LERT could bring significant improvements over various comparable baselines. Furthermore, we also conduct analytical experiments in various linguistic aspects, and the results prove that the design of LERT is valid and effective. Resources are available at https://github.com/ymcui/LERT

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations

    cs.MM 2025-05 conditional novelty 6.0 of 10

    EmotionTalk provides 19,250 utterances from 744 Chinese dyadic dialogues with emotion, sentiment, and speaking-style caption annotations.

  2. Learning from Impairment: Leveraging Insights from Clinical Linguistics in Language Modelling Research

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Aphasia treatment protocols like CATE offer complexity hierarchies that the paper proposes to reuse for language model evaluation and curriculum learning, without providing empirical evidence.

Pith tools