Pith. sign in

REVIEW 14 cited by

ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.02137 v1 pith:DFGC2T2M submitted 2021-07-05 cs.CL

classification cs.CL
keywords knowledgemodelslanguagelarge-scalemodeltaskstrainedlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. Recent works such as T5 and GPT-3 have shown that scaling up pre-trained language models can improve their generalization abilities. Particularly, the GPT-3 model with 175 billion parameters shows its strong task-agnostic zero-shot/few-shot learning capabilities. Despite their success, these large-scale models are trained on plain texts without introducing knowledge such as linguistic knowledge and world knowledge. In addition, most large-scale models are trained in an auto-regressive way. As a result, this kind of traditional fine-tuning approach demonstrates relatively weak performance when solving downstream language understanding tasks. In order to solve the above problems, we propose a unified framework named ERNIE 3.0 for pre-training large-scale knowledge enhanced models. It fuses auto-regressive network and auto-encoding network, so that the trained model can be easily tailored for both natural language understanding and generation tasks with zero-shot learning, few-shot learning or fine-tuning. We trained the model with 10 billion parameters on a 4TB corpus consisting of plain texts and a large-scale knowledge graph. Empirical results show that the model outperforms the state-of-the-art models on 54 Chinese NLP tasks, and its English version achieves the first place on the SuperGLUE benchmark (July 3, 2021), surpassing the human performance by +0.8% (90.6% vs. 89.8%).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Aligning language model entity embeddings with UMLS knowledge graph subgraph embeddings during pre-training improves biomedical QA and entity linking across PubMedBERT and BioLinkBERT.

  2. MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh

    cs.GR 2025-08 unverdicted novelty 6.0 of 10

    MeshLLM improves LLM-based 3D mesh understanding and generation through primitive decomposition, a 1500k+ sample dataset, and topology-focused training strategies.

  3. Proactive Guidance of Multi-Turn Conversation in Industrial Search

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A framework combining goal-adaptive supervised fine-tuning and click-based reinforcement learning improves proactive guidance quality and speed in an industrial search assistant.

  4. FinRipple: Aligning Large Language Models with Financial Market for Event Ripple Effect Awareness

    cs.SI 2025-05 reject novelty 6.0 of 10

    FinRipple aligns LLMs with financial markets via knowledge-graph adapters and PPO using CAPM residuals as reward, claiming strong ripple-effect prediction, but the evaluation is circular and artifacts are unavailable.

  5. The Graph Language: How Knowledge Graphs Speak to Large Language Models

    cs.AI 2026-08 conditional novelty 5.0 of 10

    GRALAN uses question-focused subgraphs, a graph encoder, and a learned mediator to let a frozen LLM answer KG questions by classifying entities, reporting state-of-the-art or near-SOTA accuracy on several QA benchmarks.

  6. Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    RRP generates semantic and structural reasoning paths, reranks them with a rethinking module, and reports SOTA Hits@1 of 90.0 on WebQSP and 64.5 on CWQ.

  7. Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuned BERT-like models outperform zero-shot and internal-state LLM methods on four of six challenging text classification datasets.

  8. A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization

    cs.CL 2026-07 conditional novelty 4.0 of 10

    An expert-editor stepwise-questioning multi-agent pipeline improves ROUGE/BERTScore/FactCC for long scientific summarization on two datasets relative to direct generation and HERA.

  9. Knowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey

    cs.LG 2026-05 conditional novelty 4.0 of 10

    A two-level taxonomy (KG pipeline stages × GNN architectures) systematically reviews GNN methods for knowledge-graph construction, embedding, reasoning, and applications.

  10. When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A literature review that classifies LLM-for-law research using a dual-lens taxonomy of Toulmin argumentation components and legal practitioner roles.

  11. A Survey on Agent Workflow -- Status and Future

    cs.AI 2025-08 conditional novelty 3.0 of 10

    A review that classifies 24 agent workflow systems along functional and architectural axes and argues for standardization, optimization, and security work.

  12. Data Augmentation for Cognitive Behavioral Therapy: Leveraging ERNIE Language Models using Artificial Intelligence

    cs.AI 2025-06 reject novelty 2.0 of 10

    This paper proposes an AI chatbot for cognitive behavioral therapy, but it provides no implementation or evaluation to support its claims.

  13. Enhancing Large Language Models with Reliable Knowledge Graphs

    cs.CL 2025-06 conditional novelty 2.0 of 10

    A thesis composed of four published papers proposes contrastive KG error detection, attribute-aware error-aware embedding, inductive graph completion, and KG prompting, but adds no new result beyond those papers.

  14. DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models

    cs.AI 2025-07 unverdicted

    A survey of DeepSeek's V3 and R1 models covering MLA, MoE, MTP, GRPO, and training engineering, with no new experimental results.

Pith tools