Pith. sign in

REVIEW 1 major objections 1 minor 3 cited by

Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices

T0 review · 1 major / 1 minor · reviewed 2026-05-23 · grok-4.3

Pith's one-line read Massive edge devices can supply the data and compute to keep scaling large language models.

desk verdict This is a position paper that restates the case for federated edge training of LLMs but adds no new evidence, calculations, or technical results. read the letter →

arxiv 2503.08223 v3 submitted 2025-03-11 cs.DC

classification cs.DC
keywords largelanguagemodelsscalinglawsedgecomputingdistributedlearningfederateddatascarcityAIdemocratization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that scaling laws for foundation models are running into two hard limits: exhaustion of high-quality public data and the concentration of required compute power in the hands of a few large organizations. It identifies the vast unused data and processing capacity sitting on billions of edge devices as an alternative resource pool. Recent progress in distributed and federated learning is presented as the practical bridge that turns these scattered devices into a workable training fabric. If the approach holds, ordinary users with small devices could contribute directly to training large models rather than remaining passive consumers.

What carries the argument

Distributed and federated learning applied to the collective data and compute resources of massive edge devices, which together provide both additional training examples and parallel processing capacity without requiring single-site data centers.

What would settle it

A controlled large-scale trial in which models trained via edge collaboration achieve materially lower performance or higher effective cost than centralized training on the same total data and compute volume.

Watch

Extended reading notes

Core claim

By collaborating across massive numbers of edge devices, the two bottlenecks of data scarcity and centralized compute monopolies can be bypassed, enabling continued scaling of large language models through distributed training that lets anyone with a small device participate.

Load-bearing premise

Recent technical advances in distributed and federated learning are now sufficient to make reliable, efficient training across billions of heterogeneous edge devices practical.

Editorial extensions

If this is right

  • High-quality public data no longer sets an absolute ceiling because private data on devices becomes usable.
  • Compute requirements are spread so that participation is no longer restricted to organizations with massive clusters.
  • AI model development can involve a wider community, reducing concentration of control.
  • New coordination mechanisms for data privacy and device incentives become necessary parts of the training pipeline.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Coordination overhead and device heterogeneity could still limit the effective scale even if the basic feasibility claim holds.
  • The same edge resources might also support inference or fine-tuning workloads once the training paradigm is established.
  • Integration with existing cloud infrastructure would likely be required for orchestration rather than replacing it outright.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. This position paper argues that LLM scaling laws face two barriers—depletion of high-quality public data and monopolization of compute by tech giants—and proposes that massive edge devices can overcome them by providing untapped data and compute resources. It reviews recent advances in distributed and federated learning to claim that collaborative training on small edge devices is now viable, enabling broad participation and democratizing AI development.

Significance. If the reviewed literature indeed establishes viability, the position could meaningfully shift AI development toward inclusive, decentralized paradigms by exploiting edge resources. The manuscript contains no new empirical results, derivations, quantitative scaling projections, or falsifiable predictions, so any significance rests entirely on the interpretive synthesis of prior work rather than original technical contributions.

major comments (1)
  1. [Abstract] Abstract: the central claim that 'recent technical advancements in distributed/federated learning ... make this new paradigm viable' is presented without any manuscript-internal quantitative analysis, scaling-law extrapolation, or independent falsifiable prediction; the viability argument therefore reduces to an untested assertion about external literature.
minor comments (1)
  1. The introduction and review sections would benefit from explicit demarcation between synthesized prior results and any original interpretive claims to improve traceability.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the detailed review of our position paper. We appreciate the acknowledgment that the work is a synthesis of prior literature rather than an empirical study. Below we respond directly to the single major comment.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the central claim that 'recent technical advancements in distributed/federated learning ... make this new paradigm viable' is presented without any manuscript-internal quantitative analysis, scaling-law extrapolation, or independent falsifiable prediction; the viability argument therefore reduces to an untested assertion about external literature.

    Authors: We agree that the manuscript performs no new quantitative analysis, scaling-law extrapolation, or falsifiable predictions of its own; this is inherent to its nature as a position paper whose contribution is interpretive synthesis. The viability claim is explicitly grounded in the cited body of recent distributed and federated learning literature that the paper reviews (e.g., advances addressing communication efficiency, heterogeneity, and privacy that were previously limiting factors for edge-scale training). To make this grounding more transparent to readers, we will revise the abstract and expand the relevant sections to include concise, paper-internal summaries of the quantitative results reported in the key referenced works, thereby strengthening the link between external evidence and the position without altering the paper's scope or adding original experiments. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

This is a position paper whose argument consists of an interpretive review of external distributed/federated learning literature. No equations, derivations, fitted parameters, or quantitative predictions are present that could reduce to self-definitions, fitted inputs renamed as predictions, or self-citation chains. The central claim simply asserts viability based on cited external advancements; no load-bearing step matches any enumerated circularity pattern.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on the untested assumption that current federated-learning methods can be extended to the scale, heterogeneity, and incentive structures of consumer edge devices without new fundamental obstacles.

assumptions (2)
  • domain assumption Edge devices collectively possess sufficient high-quality data and idle compute to substitute for centralized resources.
    Invoked in the abstract when stating the vast untapped potential of data and computational resources on massive edge devices.
  • domain assumption Technical advancements in distributed/federated learning are already sufficient to make large-scale edge collaboration practical.
    Stated directly in the abstract as the basis for viability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices." pith.science (2026). https://pith.science/paper/2503.08223

@misc{pith2026250308223,
  author       = {Pith},
  title        = {Pith review of: Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2503.08223}},
  note         = {Machine review of arXiv:2503.08223}
}
read the original abstract

The remarkable success of foundation models has been driven by scaling laws, demonstrating that model performance improves predictably with increased training data and model size. However, this scaling trajectory faces two critical challenges: the depletion of high-quality public data, and the prohibitive computational power required for larger models, which have been monopolized by tech giants. These two bottlenecks pose significant obstacles to the further development of AI. In this position paper, we argue that leveraging massive distributed edge devices can break through these barriers. We reveal the vast untapped potential of data and computational resources on massive edge devices, and review recent technical advancements in distributed/federated learning that make this new paradigm viable. Our analysis suggests that by collaborating on edge devices, everyone can participate in training large language models with small edge devices. This paradigm shift towards distributed training on edge has the potential to democratize AI development and foster a more inclusive AI community.

Figures

Figures reproduced from arXiv: 2503.08223 by the authors.

Figure 1
Figure 1. Trend of Computational Demand for Model Training. (Data source: [38]). Computational demand is growing expo￾nentially. As large-scale AI models like GPT-4 [4], Llama 3 [12], and DeepSeek￾V3 [11] surpass the trillion-parameter scale, the global AI landscape faces severe compu￾tational efficiency challenges. As shown in [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Global data volume from 2014 to 2025 and IoT device data volume in 2015 and 2025. (Data sources: Global data volume from [43]; IoT device data volume from [44].) 5.0 5.6 5.9 6.3 6.6 7.0 7.2 7.4 7.6 7.8 8.0 2.6 3.5 4.7 7.4 11.2 16.5 23.7 33.4 46.7 64.4 87.9 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 4.5 5 5.5 6 6.5 7 7.5 8 8.5 0 20 40 60 80 100 120 Actual Smartphone Data Predicted Smartphone Data Actual E… view at source ↗
Figure 5
Figure 5. Smartphone Market Share and Comput￾ing Power Trends. (Data source: [55]). Edge computing has potential for LLM training. We analyze the performance of smartphone chips, representing typical edge devices, and estimated their overall computing power. To ensure our estimation is as accurate as possible, we based our calculations on the market share data from [55]. We then estimated the total computing power of newly pr… view at source ↗
Figures from the paper (1 more)
Figure 6
Figure 6. Figure 6: Train Large Language Models with Small Edge Devices [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning

    cs.LG 2025-07 unverdicted novelty 7.0 of 10

    A single global merge at the final step of decentralized SGD matches the convergence rate of parallel SGD while improving test accuracy under high data heterogeneity.

  2. Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments

    cs.LG 2025-08 conditional novelty 5.0 of 10

    MetaInf, an XGBoost meta-scheduler with LLM-derived embeddings, selects inference acceleration strategies with reported 89.8% accuracy and 1.55x average acceleration, beating baselines.

  3. LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems

    cs.LG 2026-01 unverdicted novelty 3.0 of 10

    A survey taxonomy of LLMs identifies three scaling crises and six efficiency paradigms while tracing the shift from generation to tool-using agents.

Reference graph

Works this paper leans on

180 extracted references · 180 canonical work pages · cited by 3 Pith papers

  1. [1]

    Scaling Laws for Neural Language Models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020

  2. [2]

    Training Compute-Optimal Large Language Models

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022

  3. [4]

    GPT-4 Technical Report

    OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  4. [5]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhari- wal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877– 1901, 2020

  5. [6]

    PaLM: Scaling Language Modeling with Pathways

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022

  6. [7]

    Carbon Emissions and Large Neural Network Training

    David Patterson, Joseph Gonzalez, Quoc V Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021

  7. [9]

    Imagenet: A large- scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large- scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255. IEEE, 2009

  8. [10]

    LLaMA: Open and Efficient Foundation Language Models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timo- thée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023

Show all 180 references
  1. [11]

    Deepseek-v3 technical report

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024

  2. [12]

    Introducing llama 3.1: Our most capable models to date

    Meta. Introducing llama 3.1: Our most capable models to date. https://ai.meta.com/ blog/meta-llama-3-1/ , 2024. Accessed: 2025-01-22

  3. [13]

    Llama 2: Open foundation and fine-tuned chat models

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023

  4. [14]

    Exploring the limits of transfer learning with a unified text-to-text transformer

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020

  5. [15]

    Introduction to federated learning

    DeepLearning.AI. Introduction to federated learning. https://www.deeplearning.ai/ short-courses/intro-to-federated-learning/ , 2024. Accessed: 2025-02-23

  6. [16]

    Deduplicating training data makes language models better

    Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499, 2021

  7. [17]

    Data governance in the age of large language models

    Stella Biderman, Kieran Schoelkopf, Anthony Weiss, and David Noever. Data governance in the age of large language models. arXiv preprint arXiv:2211.09911, 2022. 10

  8. [18]

    Position: Will we run out of data? limits of llm scaling based on human-generated data

    Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn. Position: Will we run out of data? limits of llm scaling based on human-generated data. In International Conference on Machine Learning, pages 49523–49544. PMLR, 2024

  9. [19]

    On the diversity of synthetic data and its impact on training large language models

    Hao Chen, Abdul Waheed, Xiang Li, Yidong Wang, Jindong Wang, Bhiksha Raj, and Marah I Abdin. On the diversity of synthetic data and its impact on training large language models. arXiv preprint arXiv:2410.15226, 2024

  10. [20]

    Ai produces gibberish when trained on too much ai-generated data, 2024

    Emily Wenger. Ai produces gibberish when trained on too much ai-generated data, 2024

  11. [21]

    Bias of ai-generated content: an examination of news produced by large language models

    Xiao Fang, Shangkun Che, Minjia Mao, Hongzhe Zhang, Ming Zhao, and Xiaohang Zhao. Bias of ai-generated content: an examination of news produced by large language models. Scientific Reports, 14(1):5224, 2024

  12. [22]

    General data protection regulation

    Protection Regulation. General data protection regulation. Intouch, 25:1–5, 2018

  13. [23]

    Are ai scaling laws hitting a wall? https://www.linkedin.com/ pulse/ai-scaling-laws-hitting-wall-dean-hardy-white-xchfe/ , 2024

    Dean Hardy-White. Are ai scaling laws hitting a wall? https://www.linkedin.com/ pulse/ai-scaling-laws-hitting-wall-dean-hardy-white-xchfe/ , 2024. Ac- cessed: 2025-01-22

  14. [24]

    Introducing grok-3

    xAI. Introducing grok-3. https://x.ai/blog/grok-3, 2025. Accessed: 2025-02-23

  15. [25]

    Deep learning’s diminishing returns

    Neil C Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F Manso. Deep learning’s diminishing returns. IEEE Spectrum, 58(10):50–55, 2021

  16. [26]

    The cost of training nlp models: A concise overview

    Or Sharir, Barak Peleg, and Yoav Shoham. The cost of training nlp models: A concise overview. arXiv preprint arXiv:2004.08900, 2020

  17. [27]

    Green ai

    Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. Green ai. Communications of the ACM, 63(12):54–63, 2020

  18. [28]

    Artificial intelligence and competition policy

    Andrei Hagiu and Julian Wright. Artificial intelligence and competition policy. International Journal of Industrial Organization, page 103134, 2025

  19. [29]

    Frontier ai regulation: Managing emerging risks to public safety

    Jack Thompson, Amanda Askell, and Jeffrey Song. Frontier ai regulation: Managing emerging risks to public safety. arXiv preprint arXiv:2207.05257, 2022

  20. [30]

    Trends in training dataset sizes

    Pablo Villalobos and Anson Ho. Trends in training dataset sizes. Epoch AI Blog, 2022

  21. [31]

    Will we run out of data? limits of llm scaling based on human-generated data

    Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn. Will we run out of data? limits of llm scaling based on human-generated data. arXiv preprint arXiv:2211.04325, pages 13–29, 2024

  22. [32]

    Compute trends across three eras of machine learning

    Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos. Compute trends across three eras of machine learning. In 2022 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE, 2022

  23. [33]

    Tinygsm: Achieving 80% on gsm8k with small models

    Bingbin Liu, Sébastien Bubeck, Ronen Eldan, Janardhan Kulkarni, Yuanzhi Li, Anh Nguyen, Rachel Ward, and Yi Zhang. Tinygsm: Achieving 80% on gsm8k with small models. arXiv preprint arXiv:2312.09237, 2023

  24. [34]

    The curse of recursion: Training on generated data makes models forget

    Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. The curse of recursion: Training on generated data makes models forget. arXiv preprint arXiv:2305.17493, 2023

  25. [35]

    Strong model collapse

    Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian, and Julia Kempe. Strong model collapse. arXiv preprint arXiv:2410.04840, 2024

  26. [36]

    Self-consuming generative models go mad

    Sina Alemohammad, Jose Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard G Baraniuk. Self-consuming generative models go mad. arXiv preprint arXiv:2307.01850, 2023

  27. [37]

    Scaling laws of synthetic images for model training

    Li Fan, Kaiming Chen, Dilip Krishnan, Dina Katabi, Phillip Isola, and Yuandong Tian. Scaling laws of synthetic images for model training. arXiv preprint arXiv:2306.09387, 2023. 11

  28. [38]

    Trends in machine learning hardware,

    Marius Hobbhahn, Lennart Heim, and Gökçe Aydos. Trends in machine learning hardware,

  29. [39]

    Accessed: 2025-01-27

  30. [40]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017

  31. [41]

    The end of moore’s law? innovation in computer systems continues

    Henry Kressel. The end of moore’s law? innovation in computer systems continues. Artificial Intelligence in Science: Challenges, Opportunities and the Future of Research, 2023

  32. [42]

    Apple, nvidia secure future with taiwan semi’s advanced chips as ai demand soars

    Benzinga Staff. Apple, nvidia secure future with taiwan semi’s advanced chips as ai demand soars. Benzinga, June 2024

  33. [43]

    Ai’s hardware hunger: The global semiconductor supply chain under pressure

    ScaleFlux Research. Ai’s hardware hunger: The global semiconductor supply chain under pressure. ScaleFlux Insights, 2024. Accessed: 2025-01-27

  34. [44]

    V olume of data/information created, captured, copied, and consumed worldwide from 2010 to 2025, 2023

    Statista global data volume. V olume of data/information created, captured, copied, and consumed worldwide from 2010 to 2025, 2023

  35. [45]

    Internet of things (iot) connected devices data size worldwide from 2019 to 2025, 2023

    Statista IoT device data volume. Internet of things (iot) connected devices data size worldwide from 2019 to 2025, 2023

  36. [46]

    Edge computing market size & share analysis report, 2023-2030, 2023

    Grand View Research. Edge computing market size & share analysis report, 2023-2030, 2023

  37. [47]

    How many smartphones are in the world?, 2023

    BankMyCell. How many smartphones are in the world?, 2023

  38. [48]

    Dataage white paper: The digitization of the world – from edge to core, 2019

    Seagate. Dataage white paper: The digitization of the world – from edge to core, 2019

  39. [49]

    Rethink data report 2020, 2020

    Seagate. Rethink data report 2020, 2020

  40. [50]

    A review on edge analytics: Issues, challenges, opportunities, promises, future directions, and applications

    Sabuzima Nayak, Ripon Patgiri, Lilapati Waikhom, and Arif Ahmed. A review on edge analytics: Issues, challenges, opportunities, promises, future directions, and applications. Digital Communications and Networks, 10(3):783–804, 2024

  41. [51]

    Edge Computing for IoT, Real-Time Data and Low Latency Processing, 2023

    Cavli Wireless. Edge Computing for IoT, Real-Time Data and Low Latency Processing, 2023. Accessed:2025-01-22

  42. [52]

    Small language model as data prospector for large language model

    Shiwen Ni, Haihong Wu, Di Yang, Qiang Qu, Hamid Alinejad-Rokny, and Min Yang. Small language model as data prospector for large language model. arXiv preprint arXiv:2412.09990, 2024

  43. [53]

    iphone 16 pro and 16 pro max - technical specifications, 2024

    Apple Inc. iphone 16 pro and 16 pro max - technical specifications, 2024

  44. [54]

    Nvidia jetson agx orin tflops specifications, 2023

    NVIDIA. Nvidia jetson agx orin tflops specifications, 2023. Forum discussion clarifying sparse vs. dense TFLOPS

  45. [55]

    NanoReview.net - Gadget Specifications and Comparisons

    NanoReview.net. NanoReview.net - Gadget Specifications and Comparisons. https:// nanoreview.net, 2025. Accessed: 2025-02-23

  46. [56]

    Canalys Newsroom - Market Analysis and Research

    Canalys. Canalys Newsroom - Market Analysis and Research. https://canalys.com/ newsroom, 2025. Accessed: 2025-02-23

  47. [57]

    Small language models: Survey, measurements, and insights

    Zhenyan Lu, Xiang Li, Dongqi Cai, Rongjie Yi, Fangming Liu, Xiwen Zhang, Nicholas D Lane, and Mengwei Xu. Small language models: Survey, measurements, and insights. arXiv preprint arXiv:2409.15790, 2024

  48. [58]

    A comprehensive survey of small language mod- els in the era of large language models: Techniques, enhancements, applications, collaboration with llms, and trustworthiness

    Fali Wang, Zhiwei Zhang, Xianren Zhang, Zongyu Wu, Tzuhao Mo, Qiuhao Lu, Wanjing Wang, Rui Li, Junjie Xu, Xianfeng Tang, et al. A comprehensive survey of small language mod- els in the era of large language models: Techniques, enhancements, applications, collaboration with llm...

  49. [59]

    A survey of small language models

    Chien Van Nguyen, Xuan Shen, Ryan Aponte, Yu Xia, Samyadeep Basu, Zhengmian Hu, Jian Chen, Mihir Parmar, Sasidhar Kunapuli, Joe Barrow, et al. A survey of small language models. arXiv preprint arXiv:2410.20011, 2024. 12

  50. [60]

    Tinybert: Distilling bert for natural language understanding

    Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351, 2020

  51. [61]

    Albert: A lite bert for self-supervised learning of language representations

    Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2020

  52. [62]

    Mamba: Linear-time sequence modeling with selective state spaces

    Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023

  53. [63]

    The zamba2 suite: Technical report

    Paolo Glorioso, Quentin Anthony, Yury Tokpanov, Anna Golubeva, Vasudev Shyam, James Whittington, Jonathan Pilault, and Beren Millidge. The zamba2 suite: Technical report. arXiv preprint arXiv:2411.15242, 2024

  54. [64]

    Hymba: A hybrid-head architecture for small language models

    Xin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Sunil Mahabalesh- warkar, Shih-Yang Liu, Matthijs Van Keirsbilck Bilicki, Ziyang Ma, Qingyao Ai, et al. Hymba: A hybrid-head architecture for small language models. arXiv preprint arXiv:2411.13676, 2024

  55. [65]

    xlstm: Extended long short-term memory

    Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prud- nikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xlstm: Extended long short-term memory. arXiv preprint arXiv:2405.04517, 2024

  56. [66]

    SlimPajama: A 627B token cleaned and deduplicated version of RedPajama

    Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, Joel Hestness, and Nolan Dey. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. https://www.cerebras.net/blog/ slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama , 2023

  57. [67]

    Rwkv: Reinventing rnns for the transformer era

    Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Pedro Arcadinho, Eric Cao, Xin Cui, Zihang Dai, Jeff Eissman, Orhan Firat, Sophia Fu, Cong Gao, Yanping Hu, Maarten Hughes, James Kenealy, Maxim Krikun, Sneha Li, Yanping Li, Xiang Liu, Lianmin Luo, David McAllester, Matthe...

  58. [68]

    Paloma: A benchmark for evaluating language model fit

    Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Walsh, Yanai Elazar, Kyle Lo, et al. Paloma: A benchmark for evaluating language model fit. Advances in Neural Information Processing Systems , 37:64338–64376, 2024

  59. [69]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2019

  60. [70]

    Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter

    Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019

  61. [71]

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

    William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022

  62. [72]

    Specializing smaller language models towards multi-step reasoning

    Yao Fu, Hao Peng, Ashish Khotilovich, Liang Chen, and Yan Yang. Specializing smaller language models towards multi-step reasoning. arXiv preprint arXiv:2301.12726, 2023

  63. [73]

    Distilling step-by-step! outperform- ing larger language models with less training data and smaller model sizes

    Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling step-by-step! outperform- ing larger language models with less training data and smaller model sizes. arXiv preprint arXiv...

  64. [74]

    Exo: Run your own ai cluster at home with everyday devices

    Exo Labs. Exo: Run your own ai cluster at home with everyday devices. https://github. com/exo-explore/exo, 2025. Accessed: 2025-01-29. 13

  65. [75]

    Rovinski, Trevor Mudge, Jason Mars, and Lingjia Tang

    Yiping Kang, Johann Hauswald, Cao Gao, Andrew M. Rovinski, Trevor Mudge, Jason Mars, and Lingjia Tang. Neurosurgeon: Collaborative intelligence between the cloud and mobile edge. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programm...

  66. [76]

    Yang, Jian Wu, and Meng Zhang

    Lyudong Jin, Yanning Zhang, Yanhan Li, Shurong Wang, Howard H. Yang, Jian Wu, and Meng Zhang. Moe2: Optimizing collaborative inference for edge large language models.arXiv preprint arXiv:2501.09410, 2025. Submitted to IEEE/ACM Transactions on Networking

  67. [77]

    Edge intelligence: On-demand deep learning model co- inference with device-edge synergy

    En Li, Zhi Zhou, and Xu Chen. Edge intelligence: On-demand deep learning model co- inference with device-edge synergy. In Proceedings of the 2018 ACM/IEEE Symposium on Edge Computing, pages 31–46. IEEE, 2018

  68. [78]

    Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference

    Shengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou, Xiaowen Chu, Yutong Lu, and Xu Chen. Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference. arXiv preprint arXiv:2405.17245, 2024

  69. [79]

    On- device training under 256kb memory

    Ji Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang, Chuang Gan, and Song Han. On- device training under 256kb memory. Advances in Neural Information Processing Systems, 35:22941–22954, 2022

  70. [80]

    Tinytl: Reduce memory, not parameters for efficient on-device learning

    Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han. Tinytl: Reduce memory, not parameters for efficient on-device learning. In Advances in Neural Information Processing Systems, volume 33, pages 11285–11297, 2020

  71. [81]

    Zerofl: Efficient on-device training for federated learning with local sparsity

    Xinchi Qiu, Javier Fernandez-Marques, Pedro PB Gusmao, Yan Gao, Titouan Parcollet, and Nicholas Donald Lane. Zerofl: Efficient on-device training for federated learning with local sparsity. In International Conference on Learning Representations, 2022

  72. [82]

    Elasticzo: A memory-efficient on-device learning with combined zeroth- and first-order optimization

    Keisuke Sugiura and Hiroki Matsutani. Elasticzo: A memory-efficient on-device learning with combined zeroth- and first-order optimization. arXiv preprint arXiv:2501.04287, 2025

  73. [83]

    Communication-efficient learning of deep networks from decentralized data

    Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Ar- cas. Communication-efficient learning of deep networks from decentralized data. Artificial intelligence and statistics, pages 1273–1282, 2017

  74. [84]

    Federated fine-tuning of large language models under heterogeneous language tasks and client resources

    Jiamu Bai, Daoyuan Chen, Bingchen Qian, Liuyi Yao, and Yaliang Li. Federated fine-tuning of large language models under heterogeneous language tasks and client resources. arXiv e-prints, pages arXiv–2402, 2024

  75. [85]

    Federated adapter on foundation models: An out-of-distribution approach

    Yiyuan Yang, Guodong Long, Tianyi Zhou, Qinghua Lu, Shanshan Ye, and Jing Jiang. Federated adapter on foundation models: An out-of-distribution approach. arXiv preprint arXiv:2505.01075, 2025

  76. [86]

    Fedpetuning: When federated learning meets the parameter-efficient tuning methods of pre- trained language models

    Zhuo Zhang, Yuanhang Yang, Yong Dai, Qifan Wang, Yue Yu, Lizhen Qu, and Zenglin Xu. Fedpetuning: When federated learning meets the parameter-efficient tuning methods of pre- trained language models. In Annual Meeting of the Association of Computational Linguistics 2023, pages ...

  77. [87]

    Feddat: an approach for foundation model finetuning in multi-modal heterogeneous federated learning

    Haokun Chen, Yao Zhang, Denis Krompass, Jindong Gu, and V olker Tresp. Feddat: an approach for foundation model finetuning in multi-modal heterogeneous federated learning. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conferenc...

  78. [88]

    Fedmatch: Federated learning over heterogeneous question answering data

    Jiangui Chen, Ruqing Zhang, Jiafeng Guo, Yixing Fan, and Xueqi Cheng. Fedmatch: Federated learning over heterogeneous question answering data. In Proceedings of the 30th ACM international conference on information & knowledge management, pages 181–190, 2021

  79. [89]

    Flower: A friendly federated ai framework, 2025

    Flowerlab. Flower: A friendly federated ai framework, 2025

  80. [90]

    Fate-llm: An industrial grade federated learning framework for large language models

    Tao Fan, Yan Kang, Guoqiang Ma, Weijing Chen, Wenbin Wei, Lixin Fan, and Qiang Yang. Fate-llm: An industrial grade federated learning framework for large language models. arXiv preprint arXiv:2310.10049, 2023. 14

  81. [91]

    Opendiloco: An open-source framework for globally distributed low- communication training, 2025

    PrimeIntellect-ai. Opendiloco: An open-source framework for globally distributed low- communication training, 2025

  82. [92]

    Photon: Federated llm pre-training

    Lorenzo Sani, Alex Iacob, Zeyu Cao, Royson Lee, Bill Marino, Yan Gao, Dongqi Cai, Zexi Li, Wanru Zhao, Xinchi Qiu, et al. Photon: Federated llm pre-training. arXiv preprint arXiv:2411.02908, 2024

  83. [93]

    Biomedlm: A 2.7 b parameter language model trained on biomedical text

    Elliot Bolton, Abhinav Venigalla, Michihiro Yasunaga, David Hall, Betty Xiong, Tony Lee, Roxana Daneshjou, Jonathan Frankle, Percy Liang, Michael Carbin, et al. Biomedlm: A 2.7 b parameter language model trained on biomedical text. arXiv preprint arXiv:2403.18421, 2024

  84. [94]

    Advances and open problems in federated learning.Foundations and Trends® in Machine Learning, 14(1–2):1–210, 2021

    Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Ar- jun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al. Advances and open problems in federated learning.Foundations and Trends® in Machine Learning, 14(1–2)...

  85. [95]

    Fjord: Fair and accurate federated learning under heterogeneous targets with ordered dropout

    Samuel Horvath, Stefanos Laskaridis, Mario Almeida, Ilias Leontiadis, Stylianos Venieris, and Nicholas Lane. Fjord: Fair and accurate federated learning under heterogeneous targets with ordered dropout. Advances in Neural Information Processing Systems, 34:12876–12889, 2021

  86. [96]

    Fedrolex: Model-heterogeneous feder- ated learning with rolling sub-model extraction

    Samiul Alam, Luyang Liu, Ming Yan, and Mi Zhang. Fedrolex: Model-heterogeneous feder- ated learning with rolling sub-model extraction. Advances in neural information processing systems, 35:29677–29690, 2022

  87. [97]

    On the effects of data heterogeneity on the convergence rates of distributed linear system solvers

    Boris Velasevic, Rohit Parasnis, Christopher G Brinton, and Navid Azizan. On the effects of data heterogeneity on the convergence rates of distributed linear system solvers. In 2023 62nd IEEE Conference on Decision and Control (CDC), pages 8394–8399. IEEE, 2023

  88. [98]

    Azizan-Ruhi, F

    N. Azizan-Ruhi, F. Lahouti, A. S. Avestimehr, and B. Hassibi. Distributed solution of large- scale linear systems via accelerated projection-based consensus. IEEE Transactions on Signal Processing, 67(14):3806–3817, July 2019

  89. [99]

    Retrieval-augmented mixture of lora experts for uploadable machine learning

    Ziyu Zhao, Leilei Gan, Guoyin Wang, Yuwei Hu, Tao Shen, Hongxia Yang, Kun Kuang, and Fei Wu. Retrieval-augmented mixture of lora experts for uploadable machine learning. arXiv preprint arXiv:2406.16989, 2024

  90. [100]

    On the opportunities and risks of foundation models

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021

  91. [101]

    Democratising artificial intelligence in healthcare: community-driven approaches for ethical solutions

    Ceilidh Welsh, Susana Román García, Gillian C Barnett, and Raj Jena. Democratising artificial intelligence in healthcare: community-driven approaches for ethical solutions. Future Healthcare Journal, 11(3):100165, 2024

  92. [102]

    Edge-cloud polarization and collaboration: A comprehensive survey for ai

    Jiangchao Yao, Shengyu Zhang, Yang Yao, Feng Wang, Jianxin Ma, Jianwei Zhang, Yunfei Chu, Luo Ji, Kunyang Jia, Tao Shen, et al. Edge-cloud polarization and collaboration: A comprehensive survey for ai. IEEE Transactions on Knowledge and Data Engineering , 35(7):6866–6886, 2022

  93. [103]

    Beyond a single ai cluster: A survey of decentralized llm training

    Haotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo, Jiajun Song, Bowen Li, Ying Shen, and Zhi Wang. Beyond a single ai cluster: A survey of decentralized llm training. arXiv preprint arXiv:2503.11023, 2025

  94. [104]

    Distributed training of large language models

    Fanlong Zeng, Wensheng Gan, Yongheng Wang, and Philip S Yu. Distributed training of large language models. In 2023 IEEE 29th International Conference on Parallel and Distributed Systems (ICPADS), pages 840–847. IEEE, 2023

  95. [105]

    Injecting domain-specific knowledge into large language models: a comprehensive survey

    Zirui Song, Bin Yan, Yuhan Liu, Miao Fang, Mingzhe Li, Rui Yan, and Xiuying Chen. Injecting domain-specific knowledge into large language models: a comprehensive survey. arXiv preprint arXiv:2502.10708, 2025

  96. [106]

    Medicalgpt: Training medical gpt model

    Ming Xu. Medicalgpt: Training medical gpt model. https://github.com/shibing624/ MedicalGPT, 2023. 15

  97. [107]

    Large language models in finance: A survey

    Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen. Large language models in finance: A survey. In Proceedings of the fourth ACM international conference on AI in finance, pages 374–382, 2023

  98. [108]

    Lawyer gpt: A legal large language model with enhanced domain knowledge and reasoning capabilities

    Shunyu Yao, Qingqing Ke, Qiwei Wang, Kangtong Li, and Jie Hu. Lawyer gpt: A legal large language model with enhanced domain knowledge and reasoning capabilities. In Proceedings of the 2024 3rd International Symposium on Robotics, Artificial Intelligence and Information Enginee...

  99. [109]

    Alphaevolve: A gemini-powered coding agent for designing advanced algorithms,

    DeepMind. Alphaevolve: A gemini-powered coding agent for designing advanced algorithms,

  100. [110]

    Accessed: 2025-05-19

  101. [111]

    The pursuit of fairness in artificial intelligence models: A survey, 2024

    Yuxin Yao et al. The pursuit of fairness in artificial intelligence models: A survey, 2024

  102. [112]

    Fair resource allocation in federated learning

    Tian Li, Maziar Sanjabi, Ahmad Beirami, and Virginia Smith. Fair resource allocation in federated learning. In International Conference on Learning Representations, 2020

  103. [113]

    Optimizing federated learning on non-IID data with reinforcement learning

    Hongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris Papailiopoulos, and Yasaman Khaz- aeni. Optimizing federated learning on non-IID data with reinforcement learning. IEEE International Conference on Computer Communications, pages 1698–1707, 2020

  104. [114]

    Agnostic federated learning

    Mehryar Mohri, Gary Sivek, and Ananda Theertha Suresh. Agnostic federated learning. In International Conference on Machine Learning, pages 4615–4625, 2019

  105. [115]

    Incentive design for efficient federated learning in mobile networks: A contract theory approach

    Jiawen Kang, Zehui Xiong, Dusit Niyato, Han Ye, and Dong In Kim. Incentive design for efficient federated learning in mobile networks: A contract theory approach. In IEEE VTS Asia Pacific Wireless Communications Symposium, pages 1–5, 2019

  106. [116]

    Khan, Shashi Raj Pandey, Nguyen H

    Latif U. Khan, Shashi Raj Pandey, Nguyen H. Tran, Walid Saad, Zhu Han, Minh N. H. Nguyen, and Choong Seon Hong. Federated learning for edge networks: Resource optimization and incentive mechanism. IEEE Communications Magazine, 57(10):94–100, 2019

  107. [117]

    Incentive mechanism design for joint resource allocation in blockchain-based federated learning

    Zhilin Wang, Qin Hu, Ruinian Li, Minghui Xu, and Zehui Xiong. Incentive mechanism design for joint resource allocation in blockchain-based federated learning. IEEE Transactions on Parallel and Distributed Systems, 34(5):1536–1547, 2023

  108. [118]

    Federated machine learning: Concept and applications

    Qiang Yang, Yang Liu, Tianjian Chen, and Yongxin Tong. Federated machine learning: Concept and applications. ACM Transactions on Intelligent Systems and Technology (TIST), 10(2):1–19, 2019

  109. [119]

    An edge computing matching framework with guaranteed quality of service

    Nafiseh Sharghivand, Farnaz Derakhshan, Lena Mashayekhy, and Leyli Mohammadkhanli. An edge computing matching framework with guaranteed quality of service. IEEE Transactions on Cloud Computing, 10(3):1557–1570, 2020

  110. [120]

    Environmental burden of united states data centers in the artificial intelli- gence era, 2024

    Yuchen Yang et al. Environmental burden of united states data centers in the artificial intelli- gence era, 2024

  111. [121]

    Carbon footprint reduction for sustainable data centers in real-time, 2024

    Xiaoyu Li et al. Carbon footprint reduction for sustainable data centers in real-time, 2024

  112. [122]

    Cooling systems in data centers: State of art and emerging technologies

    Alfonso Capozzoli and Giulio Primiceri. Cooling systems in data centers: State of art and emerging technologies. Energy Procedia, 83:484–493, 2015

  113. [123]

    Beutel, Taner Topal, Akhil Mathur, and Nicholas D

    Xinchi Qiu, Titouan Parcollet, Daniel J. Beutel, Taner Topal, Akhil Mathur, and Nicholas D. Lane. Can federated learning save the planet? arXiv preprint arXiv:2010.06537, 2021

  114. [124]

    Communication-efficient learning of deep networks from decentralized data

    Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Ar- cas. Communication-efficient learning of deep networks from decentralized data. Artificial Intelligence and Statistics, pages 1273–1282, 2017

  115. [125]

    Nvidia announces jetson tx2: Parker comes to nvidia’s embedded system kit

    Ryan Smith. Nvidia announces jetson tx2: Parker comes to nvidia’s embedded system kit. IEEE Hot Chips, 29, 2017

  116. [126]

    Federated optimization in heterogeneous networks

    Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. arXiv preprint arXiv:1812.06127, 2018. 16

  117. [127]

    Quantifying the carbon emissions of machine learning

    Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700, 2019

  118. [128]

    Imagenet classification with deep convolutional neural networks

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems , volume 25, pages 1097–1105, 2012

  119. [129]

    The computa- tional limits of deep learning

    Neil C Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F Manso. The computa- tional limits of deep learning. arXiv preprint arXiv:2007.05558, 10, 2020

  120. [130]

    The perceptron: a probabilistic model for information storage and organiza- tion in the brain

    Frank Rosenblatt. The perceptron: a probabilistic model for information storage and organiza- tion in the brain. Psychological review, 65(6):386, 1958

  121. [131]

    Gradient-based learning applied to document recognition

    Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998

  122. [132]

    Large-scale deep unsupervised learning using graphics processors

    Rajat Raina, Anand Madhavan, and Andrew Y Ng. Large-scale deep unsupervised learning using graphics processors. In Proceedings of the 26th annual international conference on machine learning, pages 873–880, 2009

  123. [133]

    In-datacenter performance analysis of a tensor processing unit

    Norman P Jouppi, Cliff Young, Nishant Patil, and David Patterson. In-datacenter performance analysis of a tensor processing unit. InProceedings of the 44th annual international symposium on computer architecture, pages 1–12, 2017

  124. [134]

    Benchmarking tpu, gpu, and cpu platforms for deep learning

    Yu Emma Wang, Gu-Yeon Wei, and David Brooks. Benchmarking tpu, gpu, and cpu platforms for deep learning. arXiv preprint arXiv:1907.10701, 2019

  125. [135]

    Ai and compute

    Dario Amodei and Danny Hernandez. Ai and compute. OpenAI Blog, 2, 2018

  126. [136]

    How many smartphones are in the world?, 2021

    CounterPoint. How many smartphones are in the world?, 2021

  127. [137]

    Phonelm: an efficient and capable small language model family through principled pre-training

    Rongjie Yi, Xiang Li, Weikai Xie, Zhenyan Lu, Chenghua Wang, Ao Zhou, Shangguang Wang, Xiwen Zhang, and Mengwei Xu. Phonelm: an efficient and capable small language model family through principled pre-training. arXiv preprint arXiv:2411.05046, 2024

  128. [138]

    Datacomp-lm: In search of the next generation of training sets for language models

    Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Guha, Sedrick Scott Keh, Kushal Arora, et al. Datacomp-lm: In search of the next generation of training sets for language models. Advances in Neural Information Processin...

  129. [139]

    Starcoder: May the source be with you! arXiv preprint arXiv:2305.06161, 2023

    Raymond Li, Daniel Choi, Jordi Chung, et al. Starcoder: May the source be with you! arXiv preprint arXiv:2305.06161, 2023

  130. [140]

    Meta releases llama 3.2

    Meta AI. Meta releases llama 3.2. https://about.fb.com/news/2024/09/ introducing-llama-3-2-1b-3b/ , 2024

  131. [141]

    Qwen2: Technical report

    Jinze Yang, Shuai Wang, Shuohang Ma, Jianbo Zheng, et al. Qwen2: Technical report. arXiv preprint arXiv:2404.05169, 2024

  132. [142]

    Qwen technical report

    Jinze Bai, Shuai Wang, Fei Xiong, Zhenyu Hou, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023

  133. [143]

    Gemma: Open models based on gemini research and technology

    Google Team. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024

  134. [144]

    Smollm2: When smol goes big – data-centric training of a small language model, 2025

    Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíˇcek, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo ...

  135. [145]

    Smollm-corpus, 2024

    Loubna Ben Allal, Anton Lozhkov, Guilherme Penedo, Thomas Wolf, and Leandro von Werra. Smollm-corpus, 2024. 17

  136. [146]

    H2o-danube3 technical report

    Pascal Pfeiffer, Philipp Singer, Yauhen Babakhin, Gabor Fodor, Nischay Dhankhar, and Sri Satish Ambati. H2o-danube3 technical report. arXiv preprint arXiv:2407.09276, 2024

  137. [147]

    Minicpm: Unveiling the potential of small language models

    Edward Hu, Wangchunshu Huang, et al. Minicpm: Unveiling the potential of small language models. arXiv preprint arXiv:2402.03216, 2024

  138. [148]

    Dolma: An open corpus of high-quality english text for language model pre-training

    AI2. Dolma: An open corpus of high-quality english text for language model pre-training. https://huggingface.co/datasets/allenai/dolma, 2023

  139. [149]

    Chinese tiny llm: Pretraining a chinese-centric large language model

    Xinrun Du, Zhouliang Yu, Songyang Gao, Ding Pan, Yuyang Cheng, Ziyang Ma, Ruibin Yuan, Xingwei Qu, Jiaheng Liu, Tianyu Zheng, et al. Chinese tiny llm: Pretraining a chinese-centric large language model. arXiv preprint arXiv:2404.04167, 2024

  140. [150]

    Olmo: Accelerating the science of language models

    Dirk Groeneveld, Iz Beltagy, Pete Davis, et al. Olmo: Accelerating the science of language models. arXiv preprint arXiv:2402.00838, 2024

  141. [151]

    Tinyllama: An open-source small language model

    Peiyuan Zhang, Guangtao Chen, et al. Tinyllama: An open-source small language model. arXiv preprint arXiv:2401.02385, 2024

  142. [152]

    Phi-3 technical report: A highly capable language model locally on your phone

    Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024

  143. [153]

    Phi-2: The surprising power of small language models

    Mojan Javaheripi, Jacob Lobo, et al. Phi-2: The surprising power of small language models. arXiv preprint arXiv:2312.12397, 2023

  144. [154]

    Textbooks are all you need

    Suriya Gunasekar, Yi Zhang, et al. Textbooks are all you need. arXiv preprint arXiv:2306.11644, 2023

  145. [155]

    Openelm: An efficient language model family with open training and inference framework.arXiv preprint arXiv:2404.14619, 2024

    Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, et al. Openelm: An efficient language model family with open training and inference framework.arXiv preprint arXiv:2404...

  146. [156]

    The refinedweb dataset for falcon llm: Outperforming curated corpora with web data

    Guilherme Penedo, Anis Crnisanin, Ethan Shen, et al. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data. arXiv preprint arXiv:2306.01116, 2023

  147. [157]

    The pile: An 800gb dataset of diverse text for language modeling

    Leo Gao, Stella Biderman, Sid Black, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020

  148. [158]

    Mobillama: Towards accurate and lightweight fully transparent gpt

    Omkar Thawakar, Ashmal Vayani, Salman Khan, Hisham Cholakal, Rao M Anwer, Michael Felsberg, Tim Baldwin, Eric P Xing, and Fahad Shahbaz Khan. Mobillama: Towards accurate and lightweight fully transparent gpt. arXiv preprint arXiv:2402.16840, 2024

  149. [159]

    Mobilellm: Optimizing sub-billion parameter language models for on-device use cases

    Zechun Liu, Changsheng Zhao, Forrest Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, et al. Mobilellm: Optimizing sub-billion parameter language models for on-device use cases. In Forty-first International Co...

  150. [160]

    Compact language models via pruning and knowledge distillation

    Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, and Pavlo Molchanov. Compact language models via pruning and knowledge distillation. Advances in Neural Information Processing Sy...

  151. [161]

    Orca 2: Teaching small language models how to reason

    Arindam Mitra, Subhabrata Mukherjee, et al. Orca 2: Teaching small language models how to reason. arXiv preprint arXiv:2311.11045, 2023

  152. [162]

    Orca 2: Teaching small language models how to reason

    Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Vicente Ordonez, and Kai-Wei Chang. Orca 2: Teaching small language models how to reason. arXiv preprint arXiv:2312.02558, 2023

  153. [163]

    The flan collection: Designing data and methods for effective instruction tuning

    Shayne Longpre, Le Hou, Tu Vu, et al. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688, 2023. 18

  154. [164]

    Towards making the most of chatgpt for machine translation

    Kehai Zhang, Zhuocheng Chen, et al. Towards making the most of chatgpt for machine translation. arXiv preprint arXiv:2309.02654, 2023

  155. [165]

    Free dolly: Introducing the world’s first truly open instruction-tuned llm

    Databricks. Free dolly: Introducing the world’s first truly open instruction-tuned llm. https://www.databricks.com/blog/2023/04/12/ dolly-first-open-commercially-viable-instruction-tuned-llm , 2023

  156. [166]

    Lamini-lm: A diverse herd of distilled models from large-scale instructions

    Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, and Alham Fikri Aji. Lamini-lm: A diverse herd of distilled models from large-scale instructions. arXiv preprint arXiv:2304.14402, 2023

  157. [167]

    Sparsegpt: Massive language models can be accurately pruned in one-shot

    Elias Frantar, Saleh Ashkboos, Torsten Hoefler, et al. Sparsegpt: Massive language models can be accurately pruned in one-shot. arXiv preprint arXiv:2301.00774, 2023

  158. [168]

    A simple and effective pruning approach for large language models

    Zongyu Sun, Chen Chen, Zhitao Zhang, et al. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695, 2023

  159. [169]

    Loraprune: Structured pruning meets low-rank parameter-efficient fine-tuning

    Mingyang Zhang, Hao Chen, Chunhua Shen, Zhen Yang, Linlin Ou, Xinyi Yu, and Bohan Zhuang. Loraprune: Structured pruning meets low-rank parameter-efficient fine-tuning. arXiv preprint arXiv:2305.18403, 2023

  160. [170]

    Shortgpt: Layers in large language models are more redundant than you expect

    Yu Men, Xingyu Zhang, Ruiqi Sun, et al. Shortgpt: Layers in large language models are more redundant than you expect. arXiv preprint arXiv:2402.18952, 2024

  161. [171]

    Bitnet: Scaling 1-bit transformers for large language models

    Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Huaijie Wang, Lingxiao Ma, Fan Yang, Ruiping Wang, Yi Wu, and Furu Wei. Bitnet: Scaling 1-bit transformers for large language models. arXiv preprint arXiv:2310.11453, 2023

  162. [172]

    The era of 1-bit llms: All large language models are in 1.58 bits

    Shuming Ma, Hongyu Zhao, Lingxiao Xue, et al. The era of 1-bit llms: All large language models are in 1.58 bits. arXiv preprint arXiv:2402.17764, 2024

  163. [173]

    Qlora: Efficient finetuning of quantized llms

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314, 2023

  164. [174]

    Squeezellm: Dense-and-sparse quantiza- tion

    Sehoon Kim, Coleman Hooper, Amir Gholami, et al. Squeezellm: Dense-and-sparse quantiza- tion. arXiv preprint arXiv:2306.07629, 2023

  165. [175]

    The on-device intelligence update, 2024

    Karan Goel. The on-device intelligence update, 2024

  166. [176]

    Training verifiers to solve math word problems

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021

  167. [177]

    Scaling instruction-finetuned language models

    Hyung Won Chung, Le Hou, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022

  168. [178]

    Together ai: The ai acceleration cloud

    Together AI. Together ai: The ai acceleration cloud. https://www.together.ai/, 2023

  169. [179]

    Flock: Federated machine learning on the blockchain

    FLock. Flock: Federated machine learning on the blockchain. https://www.flock.io/, 2023

  170. [180]

    Federatedscope: An easy-to-use federated learning platform

    alibaba. Federatedscope: An easy-to-use federated learning platform. https://github. com/alibaba/FederatedScope, 2024

  171. [181]

    Fedml: The unified and scalable ml library for large-scale distributed training, model serving, and federated learning

    FedML-AI. Fedml: The unified and scalable ml library for large-scale distributed training, model serving, and federated learning. https://github.com/FedML-AI/FedML, 2024

  172. [182]

    Fedllm-bench: Realistic benchmarks for federated learning of large language models

    Rui Ye, Rui Ge, Xinyu Zhu, Jingyi Chai, Du Yaxin, Yang Liu, Yanfeng Wang, and Siheng Chen. Fedllm-bench: Realistic benchmarks for federated learning of large language models. Advances in Neural Information Processing Systems, 37:111106–111130, 2025. 19 A Impact Statements The ...

Pith tools

Reviewed May 23, 2026 · model on record in the stance chip above.