Pith. sign in

REVIEW 3 cited by

YuLan: An Open-source Large Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19853 v1 pith:2XPTMRNI submitted 2024-06-28 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsyulanlanguagemodelopen-sourcetrainingacrosscapabilities
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Large language models (LLMs) have become the foundation of many applications, leveraging their extensive capabilities in processing and understanding natural language. While many open-source LLMs have been released with technical reports, the lack of training details hinders further research and development. This paper presents the development of YuLan, a series of open-source LLMs with $12$ billion parameters. The base model of YuLan is pre-trained on approximately $1.7$T tokens derived from a diverse corpus, including massive English, Chinese, and multilingual texts. We design a three-stage pre-training method to enhance YuLan's overall capabilities. Subsequent phases of training incorporate instruction-tuning and human alignment, employing a substantial volume of high-quality synthesized data. To facilitate the learning of complex and long-tail knowledge, we devise a curriculum-learning framework throughout across these stages, which helps LLMs learn knowledge in an easy-to-hard manner. YuLan's training is finished on Jan, 2024 and has achieved performance on par with state-of-the-art LLMs across various English and Chinese benchmarks. This paper outlines a comprehensive technical roadmap for developing LLMs from scratch. Our model and codes are available at https://github.com/RUC-GSAI/YuLan-Chat.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FinS-Pilot: A Benchmark for Online Financial RAG System

    cs.CL 2025-05 reject novelty 6.0 of 10

    FinS-Pilot is a small benchmark of real-world financial assistant queries with real-time API data and text corpus, used to compare Chinese LLMs on financial RAG tasks.

  2. GREAT: Guiding Query Generation with a Trie for Recommending Related Search about Video at Kuaishou

    cs.IR 2025-07 conditional novelty 5.0 of 10

    An LLM-based item-to-query recommender with trie-constrained decoding, plus a new dataset, reports modest gains over baselines in Kuaishou's related-search scenario.

  3. YuLan-Mini: An Open Data-efficient Language Model

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A 2.42B-parameter base model trained on 1.08T tokens matches or beats several industry baselines trained on 7T to 18T tokens across math, code, and general benchmarks.

Pith tools