Pith. sign in

REVIEW 8 cited by

The Future of Large Language Model Pre-training is Federated

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.10853 v3 pith:6RSGLDMF submitted 2024-05-17 cs.LG cs.AIcs.DC

classification cs.LGcs.AIcs.DC
keywords llmstrainingdatafederatedpre-trainingphotonresourcescomputational
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative pre-trained large language models (LLMs) have demonstrated impressive performance over a wide range of tasks, thanks to the unprecedented amount of data they have been trained on. As established scaling laws indicate, LLMs' future performance improvement depends on the amount of computing and data sources they can leverage for pre-training. Federated learning (FL) has the potential to unleash the majority of the planet's data and computational resources, which are underutilized by the data-center-focused training methodology of current LLM practice. Our work presents a robust, flexible, reproducible FL approach that enables large-scale collaboration across institutions to train LLMs. We propose a scalable deployment system called Photon to enable the investigation and development of this new training paradigm for LLM pre-training. We show that Photon can be used by organizations interested in collaborating with their private data sources and computational resources for pre-training LLMs with billions of parameters. This paradigm would mobilize more computational and data resources while matching or potentially exceeding centralized performance. We further show the effectiveness of the federated training scales with model size and present our approach for training billion-scale federated LLMs using limited resources. Thus far, we have used Photon to train LLM models to the size of 7B parameters and anticipate larger models being completed in the near future. Finally, we show that LLM training is highly resilient to the classical challenges of federated statistical and hardware heterogeneity. Furthermore, we show that convergence is robust to partial participation, opening the avenue for compute-efficient collaborative training. Photon will help data-rich actors to become the protagonists of LLMs pre-training instead of leaving the stage to compute-rich actors alone.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Collaborative Threshold Watermarking

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A federated-learning watermark that is embedded collectively by all clients and can only be verified by coalitions of at least t clients, demonstrated up to K=128.

  2. DES-LOC: Desynced Low Communication Adaptive Optimizers for Training Foundation Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    DES-LOC synchronizes model parameters and Adam/ADOPT momentum states on separate schedules, matching Local Adam quality with about 2x less communication and 170x less than DDP in tests up to 1.7B parameters.

  3. Incentivizing Permissionless Distributed Learning of LLMs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A deployed incentive mechanism rewarded pseudo-gradient updates with tokens and produced a competitive 1.2B LLM via permissionless distributed training on Bittensor.

  4. A Comprehensive Data-centric Overview of Federated Graph Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A data-centric taxonomy for Federated Graph Learning that classifies 79 studies by data characteristics and data utilization, plus a discussion of integration with pre-trained large models.

  5. Distributed and Decentralised Training: Technical Governance Challenges in a Shifting AI Landscape

    cs.CY 2025-07 conditional novelty 5.0 of 10

    A policy analysis distinguishing distributed and decentralised AI training, arguing decentralised training may erode detectability and shutdownability while compute controls remain relevant.

  6. Federated In-Context Learning: Iterative Refinement for Improved Answer Quality

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Fed-ICL iteratively refines QA answers via federated in-context learning with only label transmission, showing convergence on a linear attention model and gains on MMLU and TruthfulQA.

  7. MuLoCo: Muon is a practical inner optimizer for DiLoCo

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Using Muon instead of AdamW inside DiLoCo improves worker scaling and critical batch size for LLM pre-training across 150M to 15B parameters.

  8. Navigating the Edge-Cloud Continuum: A State-of-Practice Survey

    cs.DC 2025-05 conditional novelty 5.0 of 10

    A state-of-practice survey that maps the edge-cloud continuum through a developer-oriented five-area conceptual framework.

Pith tools