REVIEW 14 cited by
ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Pre-trained models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks. Recent works such as T5 and GPT-3 have shown that scaling up pre-trained language models can improve their generalization abilities. Particularly, the GPT-3 model with 175 billion parameters shows its strong task-agnostic zero-shot/few-shot learning capabilities. Despite their success, these large-scale models are trained on plain texts without introducing knowledge such as linguistic knowledge and world knowledge. In addition, most large-scale models are trained in an auto-regressive way. As a result, this kind of traditional fine-tuning approach demonstrates relatively weak performance when solving downstream language understanding tasks. In order to solve the above problems, we propose a unified framework named ERNIE 3.0 for pre-training large-scale knowledge enhanced models. It fuses auto-regressive network and auto-encoding network, so that the trained model can be easily tailored for both natural language understanding and generation tasks with zero-shot learning, few-shot learning or fine-tuning. We trained the model with 10 billion parameters on a 4TB corpus consisting of plain texts and a large-scale knowledge graph. Empirical results show that the model outperforms the state-of-the-art models on 54 Chinese NLP tasks, and its English version achieves the first place on the SuperGLUE benchmark (July 3, 2021), surpassing the human performance by +0.8% (90.6% vs. 89.8%).
Forward citations
Cited by 14 Pith papers
-
BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
Aligning language model entity embeddings with UMLS knowledge graph subgraph embeddings during pre-training improves biomedical QA and entity linking across PubMedBERT and BioLinkBERT.
-
MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
MeshLLM improves LLM-based 3D mesh understanding and generation through primitive decomposition, a 1500k+ sample dataset, and topology-focused training strategies.
-
Proactive Guidance of Multi-Turn Conversation in Industrial Search
A framework combining goal-adaptive supervised fine-tuning and click-based reinforcement learning improves proactive guidance quality and speed in an industrial search assistant.
-
FinRipple: Aligning Large Language Models with Financial Market for Event Ripple Effect Awareness
FinRipple aligns LLMs with financial markets via knowledge-graph adapters and PPO using CAPM residuals as reward, claiming strong ripple-effect prediction, but the evaluation is circular and artifacts are unavailable.
-
The Graph Language: How Knowledge Graphs Speak to Large Language Models
GRALAN uses question-focused subgraphs, a graph encoder, and a learned mediator to let a frozen LLM answer KG questions by classifying entities, reporting state-of-the-art or near-SOTA accuracy on several QA benchmarks.
-
Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning with Knowledge Graphs
RRP generates semantic and structural reasoning paths, reranks them with a rethinking module, and reports SOTA Hits@1 of 90.0 on WebQSP and 64.5 on CWQ.
-
Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?
Fine-tuned BERT-like models outperform zero-shot and internal-state LLM methods on four of six challenging text classification datasets.
-
A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization
An expert-editor stepwise-questioning multi-agent pipeline improves ROUGE/BERTScore/FactCC for long scientific summarization on two datasets relative to direct generation and HERA.
-
Knowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey
A two-level taxonomy (KG pipeline stages × GNN architectures) systematically reviews GNN methods for knowledge-graph construction, embedding, reasoning, and applications.
-
When Large Language Models Meet Law: Dual-Lens Taxonomy, Technical Advances, and Ethical Governance
A literature review that classifies LLM-for-law research using a dual-lens taxonomy of Toulmin argumentation components and legal practitioner roles.
-
A Survey on Agent Workflow -- Status and Future
A review that classifies 24 agent workflow systems along functional and architectural axes and argues for standardization, optimization, and security work.
-
Data Augmentation for Cognitive Behavioral Therapy: Leveraging ERNIE Language Models using Artificial Intelligence
This paper proposes an AI chatbot for cognitive behavioral therapy, but it provides no implementation or evaluation to support its claims.
-
Enhancing Large Language Models with Reliable Knowledge Graphs
A thesis composed of four published papers proposes contrastive KG error detection, attribute-aware error-aware embedding, inductive graph completion, and KG prompting, but adds no new result beyond those papers.
-
DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
A survey of DeepSeek's V3 and R1 models covering MLA, MoE, MTP, GRPO, and training engineering, with no new experimental results.
Discussion (0). Continue with ORCID to comment.