Pith. sign in

REVIEW 4 cited by

PanGu-$\alpha$: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.12369 v1 pith:QLK7DOAX submitted 2021-04-26 cs.CL

classification cs.CL
keywords alphapangu-parallelismlanguagemodelchinesefew-shotgeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Large-scale Pretrained Language Models (PLMs) have become the new paradigm for Natural Language Processing (NLP). PLMs with hundreds of billions parameters such as GPT-3 have demonstrated strong performances on natural language understanding and generation with \textit{few-shot in-context} learning. In this work, we present our practice on training large-scale autoregressive language models named PanGu-$\alpha$, with up to 200 billion parameters. PanGu-$\alpha$ is developed under the MindSpore and trained on a cluster of 2048 Ascend 910 AI processors. The training parallelism strategy is implemented based on MindSpore Auto-parallel, which composes five parallelism dimensions to scale the training task to 2048 processors efficiently, including data parallelism, op-level model parallelism, pipeline model parallelism, optimizer model parallelism and rematerialization. To enhance the generalization ability of PanGu-$\alpha$, we collect 1.1TB high-quality Chinese data from a wide range of domains to pretrain the model. We empirically test the generation ability of PanGu-$\alpha$ in various scenarios including text summarization, question answering, dialogue generation, etc. Moreover, we investigate the effect of model scales on the few-shot performances across a broad range of Chinese NLP tasks. The experimental results demonstrate the superior capabilities of PanGu-$\alpha$ in performing various tasks under few-shot or zero-shot settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Chengyu-Bench is a 2,937-example human-verified benchmark showing LLMs are strong at idiom sentiment classification but weak at appropriateness and open cloze generation.

  2. Scalable Complexity Control Facilitates Reasoning Ability of LLMs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Controlling model complexity through smaller initialization rates and stronger weight decay improved LLM benchmark scores and made loss-versus-scale curves descend faster.

  3. EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A retrieval pipeline that augments CLIP queries with LLM-written entity visual descriptions, selected by a retriever-trained rewriter, improves image-text retrieval over CLIP baselines.

  4. Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models

    cs.RO 2025-02 conditional novelty 6.0 of 10

    Occ-LLM tokenizes 4D occupancy with a motion/static separation VAE and uses Llama-2 to forecast occupancy, plan ego motion, and answer scene questions, reporting state-of-the-art results on nuScenes.

Pith tools