Pith. sign in

REVIEW 7 cited by

PanGu-{\Sigma}: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.10845 v1 pith:3GWPVC6O submitted 2023-03-20 cs.CL

classification cs.CL
keywords languagemodelpangu-sigmacomputinggenerationheterogeneousparameter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The scaling of large language models has greatly improved natural language understanding, generation, and reasoning. In this work, we develop a system that trained a trillion-parameter language model on a cluster of Ascend 910 AI processors and MindSpore framework, and present the language model with 1.085T parameters named PanGu-{\Sigma}. With parameter inherent from PanGu-{\alpha}, we extend the dense Transformer model to sparse one with Random Routed Experts (RRE), and efficiently train the model over 329B tokens by using Expert Computation and Storage Separation(ECSS). This resulted in a 6.3x increase in training throughput through heterogeneous computing. Our experimental findings show that PanGu-{\Sigma} provides state-of-the-art performance in zero-shot learning of various Chinese NLP downstream tasks. Moreover, it demonstrates strong abilities when fine-tuned in application data of open-domain dialogue, question answering, machine translation and code generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

    cs.LG 2025-06 reject novelty 7.0 of 10

    Constraining transformer projection weights to a shared low-rank subspace reportedly enables near-lossless compression of pipeline-parallel communication, matching centralized convergence at 80Mbps bandwidth.

  2. Universal Pansharpening Model

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A single pansharpening model works across 4-, 7-, 8-, and 10-band satellite images by projecting arbitrary-band MS data into a fixed latent space and fusing with PAN via a latent diffusion bridge.

  3. DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection

    cs.CL 2025-11 conditional novelty 6.0 of 10

    DEER, a disentangled mixture-of-experts detector with RL-based instance routing, reports F1 gains of about 1.4 in-domain and 5.3 points out-of-domain over prior MGT detectors.

  4. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

    cs.AI 2026-06 conditional novelty 4.0 of 10

    Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.

  5. StackingNet: Collective Inference Across Independent AI Foundation Models

    cs.AI 2026-02 conditional novelty 4.0 of 10

    A lightweight weighted-average 'meta-model' over black-box LLM/VLM outputs improves accuracy, reduces bias, and ranks/prunes unreliable models across regression and classification tasks.

  6. Containerized In-Storage Processing and Computing-Enabled SSD Disaggregation

    cs.AR 2025-06 conditional novelty 4.0 of 10

    The paper presents DockerSSD, a computational SSD that executes Docker containers on firmware via custom NVMe-based Ethernet and virtualized system calls, with simulated and FPGA-prototype evaluations showing speedups...

  7. FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation

    cs.CL 2025-05 reject novelty 3.0 of 10

    FuxiMT combines a frozen BLOOMz model with sparse mixture-of-experts layers, Chinese-first pretraining, and curriculum learning to translate into Chinese from 65 languages, with claimed low-resource gains that the pap...

Pith tools