REVIEW 1 major objections 1 minor 3 cited by
Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices
T0 review · 1 major / 1 minor · reviewed 2026-05-23 · grok-4.3
Pith's one-line read Massive edge devices can supply the data and compute to keep scaling large language models.
desk verdict This is a position paper that restates the case for federated edge training of LLMs but adds no new evidence, calculations, or technical results. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Distributed and federated learning applied to the collective data and compute resources of massive edge devices, which together provide both additional training examples and parallel processing capacity without requiring single-site data centers.
What would settle it
A controlled large-scale trial in which models trained via edge collaboration achieve materially lower performance or higher effective cost than centralized training on the same total data and compute volume.
Extended reading notes
Core claim
By collaborating across massive numbers of edge devices, the two bottlenecks of data scarcity and centralized compute monopolies can be bypassed, enabling continued scaling of large language models through distributed training that lets anyone with a small device participate.
Load-bearing premise
Recent technical advances in distributed and federated learning are now sufficient to make reliable, efficient training across billions of heterogeneous edge devices practical.
Editorial extensions
If this is right
- High-quality public data no longer sets an absolute ceiling because private data on devices becomes usable.
- Compute requirements are spread so that participation is no longer restricted to organizations with massive clusters.
- AI model development can involve a wider community, reducing concentration of control.
- New coordination mechanisms for data privacy and device incentives become necessary parts of the training pipeline.
Reading between the lines
- Coordination overhead and device heterogeneity could still limit the effective scale even if the basic feasibility claim holds.
- The same edge resources might also support inference or fine-tuning workloads once the training paradigm is established.
- Integration with existing cloud infrastructure would likely be required for orchestration rather than replacing it outright.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that LLM scaling laws face two barriers—depletion of high-quality public data and monopolization of compute by tech giants—and proposes that massive edge devices can overcome them by providing untapped data and compute resources. It reviews recent advances in distributed and federated learning to claim that collaborative training on small edge devices is now viable, enabling broad participation and democratizing AI development.
Significance. If the reviewed literature indeed establishes viability, the position could meaningfully shift AI development toward inclusive, decentralized paradigms by exploiting edge resources. The manuscript contains no new empirical results, derivations, quantitative scaling projections, or falsifiable predictions, so any significance rests entirely on the interpretive synthesis of prior work rather than original technical contributions.
major comments (1)
- [Abstract] Abstract: the central claim that 'recent technical advancements in distributed/federated learning ... make this new paradigm viable' is presented without any manuscript-internal quantitative analysis, scaling-law extrapolation, or independent falsifiable prediction; the viability argument therefore reduces to an untested assertion about external literature.
minor comments (1)
- The introduction and review sections would benefit from explicit demarcation between synthesized prior results and any original interpretive claims to improve traceability.
Simulated Author's Rebuttal
We thank the referee for the detailed review of our position paper. We appreciate the acknowledgment that the work is a synthesis of prior literature rather than an empirical study. Below we respond directly to the single major comment.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central claim that 'recent technical advancements in distributed/federated learning ... make this new paradigm viable' is presented without any manuscript-internal quantitative analysis, scaling-law extrapolation, or independent falsifiable prediction; the viability argument therefore reduces to an untested assertion about external literature.
Authors: We agree that the manuscript performs no new quantitative analysis, scaling-law extrapolation, or falsifiable predictions of its own; this is inherent to its nature as a position paper whose contribution is interpretive synthesis. The viability claim is explicitly grounded in the cited body of recent distributed and federated learning literature that the paper reviews (e.g., advances addressing communication efficiency, heterogeneity, and privacy that were previously limiting factors for edge-scale training). To make this grounding more transparent to readers, we will revise the abstract and expand the relevant sections to include concise, paper-internal summaries of the quantitative results reported in the key referenced works, thereby strengthening the link between external evidence and the position without altering the paper's scope or adding original experiments. revision: partial
Circularity Check
No significant circularity
full rationale
This is a position paper whose argument consists of an interpretive review of external distributed/federated learning literature. No equations, derivations, fitted parameters, or quantitative predictions are present that could reduce to self-definitions, fitted inputs renamed as predictions, or self-citation chains. The central claim simply asserts viability based on cited external advancements; no load-bearing step matches any enumerated circularity pattern.
Assumptions & free parameters
assumptions (2)
- domain assumption Edge devices collectively possess sufficient high-quality data and idle compute to substitute for centralized resources.
- domain assumption Technical advancements in distributed/federated learning are already sufficient to make large-scale edge collaboration practical.
Cite this review
Pith. "Pith review of Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices." pith.science (2026). https://pith.science/paper/2503.08223
@misc{pith2026250308223,
author = {Pith},
title = {Pith review of: Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices},
year = {2026},
howpublished = {\url{https://pith.science/paper/2503.08223}},
note = {Machine review of arXiv:2503.08223}
}
read the original abstract
The remarkable success of foundation models has been driven by scaling laws, demonstrating that model performance improves predictably with increased training data and model size. However, this scaling trajectory faces two critical challenges: the depletion of high-quality public data, and the prohibitive computational power required for larger models, which have been monopolized by tech giants. These two bottlenecks pose significant obstacles to the further development of AI. In this position paper, we argue that leveraging massive distributed edge devices can break through these barriers. We reveal the vast untapped potential of data and computational resources on massive edge devices, and review recent technical advancements in distributed/federated learning that make this new paradigm viable. Our analysis suggests that by collaborating on edge devices, everyone can participate in training large language models with small edge devices. This paradigm shift towards distributed training on edge has the potential to democratize AI development and foster a more inclusive AI community.
Figures
Forward citations
Cited by 3 Pith papers
-
On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning
A single global merge at the final step of decentralized SGD matches the convergence rate of parallel SGD while improving test accuracy under high data heterogeneity.
-
Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments
MetaInf, an XGBoost meta-scheduler with LLM-derived embeddings, selects inference acceleration strategies with reported 89.8% accuracy and 1.55x average acceleration, beating baselines.
-
LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems
A survey taxonomy of LLMs identifies three scaling crises and six efficiency paradigms while tracing the shift from generation to tool-using agents.
Reference graph
Works this paper leans on
-
[1]
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020
work page Pith review arXiv 2001
-
[2]
Training Compute-Optimal Large Language Models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022
work page Pith review arXiv 2022
-
[4]
OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023
work page Pith review arXiv 2023
-
[5]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhari- wal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877– 1901, 2020
work page 1901
-
[6]
PaLM: Scaling Language Modeling with Pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022
work page Pith review arXiv 2022
-
[7]
Carbon Emissions and Large Neural Network Training
David Patterson, Joseph Gonzalez, Quoc V Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350, 2021
work page Pith review arXiv 2021
-
[9]
Imagenet: A large- scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large- scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255. IEEE, 2009
work page 2009
-
[10]
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timo- thée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
work page Pith review arXiv 2023
Show all 180 references
-
[11]
Deepseek-v3 technical report
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024
2024 arXiv
-
[12]
Introducing llama 3.1: Our most capable models to date
Meta. Introducing llama 3.1: Our most capable models to date. https://ai.meta.com/ blog/meta-llama-3-1/ , 2024. Accessed: 2025-01-22
2024
-
[13]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[14]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67, 2020
2020
-
[15]
Introduction to federated learning
DeepLearning.AI. Introduction to federated learning. https://www.deeplearning.ai/ short-courses/intro-to-federated-learning/ , 2024. Accessed: 2025-02-23
2024
-
[16]
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499, 2021
2021 arXiv
-
[17]
Data governance in the age of large language models
Stella Biderman, Kieran Schoelkopf, Anthony Weiss, and David Noever. Data governance in the age of large language models. arXiv preprint arXiv:2211.09911, 2022. 10
2022
-
[18]
Position: Will we run out of data? limits of llm scaling based on human-generated data
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn. Position: Will we run out of data? limits of llm scaling based on human-generated data. In International Conference on Machine Learning, pages 49523–49544. PMLR, 2024
2024
-
[19]
On the diversity of synthetic data and its impact on training large language models
Hao Chen, Abdul Waheed, Xiang Li, Yidong Wang, Jindong Wang, Bhiksha Raj, and Marah I Abdin. On the diversity of synthetic data and its impact on training large language models. arXiv preprint arXiv:2410.15226, 2024
2024
-
[20]
Ai produces gibberish when trained on too much ai-generated data, 2024
Emily Wenger. Ai produces gibberish when trained on too much ai-generated data, 2024
2024
-
[21]
Bias of ai-generated content: an examination of news produced by large language models
Xiao Fang, Shangkun Che, Minjia Mao, Hongzhe Zhang, Ming Zhao, and Xiaohang Zhao. Bias of ai-generated content: an examination of news produced by large language models. Scientific Reports, 14(1):5224, 2024
2024
-
[22]
General data protection regulation
Protection Regulation. General data protection regulation. Intouch, 25:1–5, 2018
2018
-
[23]
Are ai scaling laws hitting a wall? https://www.linkedin.com/ pulse/ai-scaling-laws-hitting-wall-dean-hardy-white-xchfe/ , 2024
Dean Hardy-White. Are ai scaling laws hitting a wall? https://www.linkedin.com/ pulse/ai-scaling-laws-hitting-wall-dean-hardy-white-xchfe/ , 2024. Ac- cessed: 2025-01-22
2024
-
[24]
Introducing grok-3
xAI. Introducing grok-3. https://x.ai/blog/grok-3, 2025. Accessed: 2025-02-23
2025
-
[25]
Deep learning’s diminishing returns
Neil C Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F Manso. Deep learning’s diminishing returns. IEEE Spectrum, 58(10):50–55, 2021
2021
-
[26]
The cost of training nlp models: A concise overview
Or Sharir, Barak Peleg, and Yoav Shoham. The cost of training nlp models: A concise overview. arXiv preprint arXiv:2004.08900, 2020
2004
-
[27]
Green ai
Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni. Green ai. Communications of the ACM, 63(12):54–63, 2020
2020
-
[28]
Artificial intelligence and competition policy
Andrei Hagiu and Julian Wright. Artificial intelligence and competition policy. International Journal of Industrial Organization, page 103134, 2025
2025
-
[29]
Frontier ai regulation: Managing emerging risks to public safety
Jack Thompson, Amanda Askell, and Jeffrey Song. Frontier ai regulation: Managing emerging risks to public safety. arXiv preprint arXiv:2207.05257, 2022
2022
-
[30]
Trends in training dataset sizes
Pablo Villalobos and Anson Ho. Trends in training dataset sizes. Epoch AI Blog, 2022
2022
-
[31]
Will we run out of data? limits of llm scaling based on human-generated data
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn. Will we run out of data? limits of llm scaling based on human-generated data. arXiv preprint arXiv:2211.04325, pages 13–29, 2024
2024
-
[32]
Compute trends across three eras of machine learning
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos. Compute trends across three eras of machine learning. In 2022 International Joint Conference on Neural Networks (IJCNN), pages 1–8. IEEE, 2022
2022
-
[33]
Tinygsm: Achieving 80% on gsm8k with small models
Bingbin Liu, Sébastien Bubeck, Ronen Eldan, Janardhan Kulkarni, Yuanzhi Li, Anh Nguyen, Rachel Ward, and Yi Zhang. Tinygsm: Achieving 80% on gsm8k with small models. arXiv preprint arXiv:2312.09237, 2023
2023
-
[34]
The curse of recursion: Training on generated data makes models forget
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson. The curse of recursion: Training on generated data makes models forget. arXiv preprint arXiv:2305.17493, 2023
2023 arXiv
-
[35]
Strong model collapse
Elvis Dohmatob, Yunzhen Feng, Arjun Subramonian, and Julia Kempe. Strong model collapse. arXiv preprint arXiv:2410.04840, 2024
2024
-
[36]
Self-consuming generative models go mad
Sina Alemohammad, Jose Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard G Baraniuk. Self-consuming generative models go mad. arXiv preprint arXiv:2307.01850, 2023
2023
-
[37]
Scaling laws of synthetic images for model training
Li Fan, Kaiming Chen, Dilip Krishnan, Dina Katabi, Phillip Isola, and Yuandong Tian. Scaling laws of synthetic images for model training. arXiv preprint arXiv:2306.09387, 2023. 11
2023
-
[38]
Trends in machine learning hardware,
Marius Hobbhahn, Lennart Heim, and Gökçe Aydos. Trends in machine learning hardware,
-
[39]
Accessed: 2025-01-27
2025
-
[40]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017
2017
-
[41]
The end of moore’s law? innovation in computer systems continues
Henry Kressel. The end of moore’s law? innovation in computer systems continues. Artificial Intelligence in Science: Challenges, Opportunities and the Future of Research, 2023
2023
-
[42]
Apple, nvidia secure future with taiwan semi’s advanced chips as ai demand soars
Benzinga Staff. Apple, nvidia secure future with taiwan semi’s advanced chips as ai demand soars. Benzinga, June 2024
2024
-
[43]
Ai’s hardware hunger: The global semiconductor supply chain under pressure
ScaleFlux Research. Ai’s hardware hunger: The global semiconductor supply chain under pressure. ScaleFlux Insights, 2024. Accessed: 2025-01-27
2024
-
[44]
V olume of data/information created, captured, copied, and consumed worldwide from 2010 to 2025, 2023
Statista global data volume. V olume of data/information created, captured, copied, and consumed worldwide from 2010 to 2025, 2023
2010
-
[45]
Internet of things (iot) connected devices data size worldwide from 2019 to 2025, 2023
Statista IoT device data volume. Internet of things (iot) connected devices data size worldwide from 2019 to 2025, 2023
2019
-
[46]
Edge computing market size & share analysis report, 2023-2030, 2023
Grand View Research. Edge computing market size & share analysis report, 2023-2030, 2023
2023
-
[47]
How many smartphones are in the world?, 2023
BankMyCell. How many smartphones are in the world?, 2023
2023
-
[48]
Dataage white paper: The digitization of the world – from edge to core, 2019
Seagate. Dataage white paper: The digitization of the world – from edge to core, 2019
2019
-
[49]
Rethink data report 2020, 2020
Seagate. Rethink data report 2020, 2020
2020
-
[50]
A review on edge analytics: Issues, challenges, opportunities, promises, future directions, and applications
Sabuzima Nayak, Ripon Patgiri, Lilapati Waikhom, and Arif Ahmed. A review on edge analytics: Issues, challenges, opportunities, promises, future directions, and applications. Digital Communications and Networks, 10(3):783–804, 2024
2024
-
[51]
Edge Computing for IoT, Real-Time Data and Low Latency Processing, 2023
Cavli Wireless. Edge Computing for IoT, Real-Time Data and Low Latency Processing, 2023. Accessed:2025-01-22
2023
-
[52]
Small language model as data prospector for large language model
Shiwen Ni, Haihong Wu, Di Yang, Qiang Qu, Hamid Alinejad-Rokny, and Min Yang. Small language model as data prospector for large language model. arXiv preprint arXiv:2412.09990, 2024
2024
-
[53]
iphone 16 pro and 16 pro max - technical specifications, 2024
Apple Inc. iphone 16 pro and 16 pro max - technical specifications, 2024
2024
-
[54]
Nvidia jetson agx orin tflops specifications, 2023
NVIDIA. Nvidia jetson agx orin tflops specifications, 2023. Forum discussion clarifying sparse vs. dense TFLOPS
2023
-
[55]
NanoReview.net - Gadget Specifications and Comparisons
NanoReview.net. NanoReview.net - Gadget Specifications and Comparisons. https:// nanoreview.net, 2025. Accessed: 2025-02-23
2025
-
[56]
Canalys Newsroom - Market Analysis and Research
Canalys. Canalys Newsroom - Market Analysis and Research. https://canalys.com/ newsroom, 2025. Accessed: 2025-02-23
2025
-
[57]
Small language models: Survey, measurements, and insights
Zhenyan Lu, Xiang Li, Dongqi Cai, Rongjie Yi, Fangming Liu, Xiwen Zhang, Nicholas D Lane, and Mengwei Xu. Small language models: Survey, measurements, and insights. arXiv preprint arXiv:2409.15790, 2024
2024
-
[58]
A comprehensive survey of small language mod- els in the era of large language models: Techniques, enhancements, applications, collaboration with llms, and trustworthiness
Fali Wang, Zhiwei Zhang, Xianren Zhang, Zongyu Wu, Tzuhao Mo, Qiuhao Lu, Wanjing Wang, Rui Li, Junjie Xu, Xianfeng Tang, et al. A comprehensive survey of small language mod- els in the era of large language models: Techniques, enhancements, applications, collaboration with llm...
2024
-
[59]
A survey of small language models
Chien Van Nguyen, Xuan Shen, Ryan Aponte, Yu Xia, Samyadeep Basu, Zhengmian Hu, Jian Chen, Mihir Parmar, Sasidhar Kunapuli, Joe Barrow, et al. A survey of small language models. arXiv preprint arXiv:2410.20011, 2024. 12
2024
-
[60]
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. Tinybert: Distilling bert for natural language understanding. arXiv preprint arXiv:1909.10351, 2020
1909
-
[61]
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:1909.11942, 2020
1909 arXiv
-
[62]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023
2023 arXiv
-
[63]
The zamba2 suite: Technical report
Paolo Glorioso, Quentin Anthony, Yury Tokpanov, Anna Golubeva, Vasudev Shyam, James Whittington, Jonathan Pilault, and Beren Millidge. The zamba2 suite: Technical report. arXiv preprint arXiv:2411.15242, 2024
2024
-
[64]
Hymba: A hybrid-head architecture for small language models
Xin Dong, Yonggan Fu, Shizhe Diao, Wonmin Byeon, Zijia Chen, Ameya Sunil Mahabalesh- warkar, Shih-Yang Liu, Matthijs Van Keirsbilck Bilicki, Ziyang Ma, Qingyao Ai, et al. Hymba: A hybrid-head architecture for small language models. arXiv preprint arXiv:2411.13676, 2024
2024
-
[65]
xlstm: Extended long short-term memory
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prud- nikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xlstm: Extended long short-term memory. arXiv preprint arXiv:2405.04517, 2024
2024
-
[66]
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
Daria Soboleva, Faisal Al-Khateeb, Robert Myers, Jacob R Steeves, Joel Hestness, and Nolan Dey. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. https://www.cerebras.net/blog/ slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama , 2023
2023
-
[67]
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Pedro Arcadinho, Eric Cao, Xin Cui, Zihang Dai, Jeff Eissman, Orhan Firat, Sophia Fu, Cong Gao, Yanping Hu, Maarten Hughes, James Kenealy, Maxim Krikun, Sneha Li, Yanping Li, Xiang Liu, Lianmin Luo, David McAllester, Matthe...
2023 arXiv
-
[68]
Paloma: A benchmark for evaluating language model fit
Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Walsh, Yanai Elazar, Kyle Lo, et al. Paloma: A benchmark for evaluating language model fit. Advances in Neural Information Processing Systems , 37:64338–64376, 2024
2024
-
[69]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2019
2019 arXiv
-
[70]
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019
1910 arXiv
-
[71]
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022
2022
-
[72]
Specializing smaller language models towards multi-step reasoning
Yao Fu, Hao Peng, Ashish Khotilovich, Liang Chen, and Yan Yang. Specializing smaller language models towards multi-step reasoning. arXiv preprint arXiv:2301.12726, 2023
2023
-
[73]
Distilling step-by-step! outperform- ing larger language models with less training data and smaller model sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling step-by-step! outperform- ing larger language models with less training data and smaller model sizes. arXiv preprint arXiv...
2023 arXiv
-
[74]
Exo: Run your own ai cluster at home with everyday devices
Exo Labs. Exo: Run your own ai cluster at home with everyday devices. https://github. com/exo-explore/exo, 2025. Accessed: 2025-01-29. 13
2025
-
[75]
Rovinski, Trevor Mudge, Jason Mars, and Lingjia Tang
Yiping Kang, Johann Hauswald, Cao Gao, Andrew M. Rovinski, Trevor Mudge, Jason Mars, and Lingjia Tang. Neurosurgeon: Collaborative intelligence between the cloud and mobile edge. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programm...
2017
-
[76]
Yang, Jian Wu, and Meng Zhang
Lyudong Jin, Yanning Zhang, Yanhan Li, Shurong Wang, Howard H. Yang, Jian Wu, and Meng Zhang. Moe2: Optimizing collaborative inference for edge large language models.arXiv preprint arXiv:2501.09410, 2025. Submitted to IEEE/ACM Transactions on Networking
2025
-
[77]
Edge intelligence: On-demand deep learning model co- inference with device-edge synergy
En Li, Zhi Zhou, and Xu Chen. Edge intelligence: On-demand deep learning model co- inference with device-edge synergy. In Proceedings of the 2018 ACM/IEEE Symposium on Edge Computing, pages 31–46. IEEE, 2018
2018
-
[78]
Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference
Shengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou, Xiaowen Chu, Yutong Lu, and Xu Chen. Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference. arXiv preprint arXiv:2405.17245, 2024
2024
-
[79]
On- device training under 256kb memory
Ji Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang, Chuang Gan, and Song Han. On- device training under 256kb memory. Advances in Neural Information Processing Systems, 35:22941–22954, 2022
2022
-
[80]
Tinytl: Reduce memory, not parameters for efficient on-device learning
Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han. Tinytl: Reduce memory, not parameters for efficient on-device learning. In Advances in Neural Information Processing Systems, volume 33, pages 11285–11297, 2020
2020
-
[81]
Zerofl: Efficient on-device training for federated learning with local sparsity
Xinchi Qiu, Javier Fernandez-Marques, Pedro PB Gusmao, Yan Gao, Titouan Parcollet, and Nicholas Donald Lane. Zerofl: Efficient on-device training for federated learning with local sparsity. In International Conference on Learning Representations, 2022
2022
-
[82]
Elasticzo: A memory-efficient on-device learning with combined zeroth- and first-order optimization
Keisuke Sugiura and Hiroki Matsutani. Elasticzo: A memory-efficient on-device learning with combined zeroth- and first-order optimization. arXiv preprint arXiv:2501.04287, 2025
2025
-
[83]
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Ar- cas. Communication-efficient learning of deep networks from decentralized data. Artificial intelligence and statistics, pages 1273–1282, 2017
2017
-
[84]
Federated fine-tuning of large language models under heterogeneous language tasks and client resources
Jiamu Bai, Daoyuan Chen, Bingchen Qian, Liuyi Yao, and Yaliang Li. Federated fine-tuning of large language models under heterogeneous language tasks and client resources. arXiv e-prints, pages arXiv–2402, 2024
2024
-
[85]
Federated adapter on foundation models: An out-of-distribution approach
Yiyuan Yang, Guodong Long, Tianyi Zhou, Qinghua Lu, Shanshan Ye, and Jing Jiang. Federated adapter on foundation models: An out-of-distribution approach. arXiv preprint arXiv:2505.01075, 2025
2025
-
[86]
Fedpetuning: When federated learning meets the parameter-efficient tuning methods of pre- trained language models
Zhuo Zhang, Yuanhang Yang, Yong Dai, Qifan Wang, Yue Yu, Lizhen Qu, and Zenglin Xu. Fedpetuning: When federated learning meets the parameter-efficient tuning methods of pre- trained language models. In Annual Meeting of the Association of Computational Linguistics 2023, pages ...
2023
-
[87]
Feddat: an approach for foundation model finetuning in multi-modal heterogeneous federated learning
Haokun Chen, Yao Zhang, Denis Krompass, Jindong Gu, and V olker Tresp. Feddat: an approach for foundation model finetuning in multi-modal heterogeneous federated learning. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conferenc...
2024
-
[88]
Fedmatch: Federated learning over heterogeneous question answering data
Jiangui Chen, Ruqing Zhang, Jiafeng Guo, Yixing Fan, and Xueqi Cheng. Fedmatch: Federated learning over heterogeneous question answering data. In Proceedings of the 30th ACM international conference on information & knowledge management, pages 181–190, 2021
2021
-
[89]
Flower: A friendly federated ai framework, 2025
Flowerlab. Flower: A friendly federated ai framework, 2025
2025
-
[90]
Fate-llm: An industrial grade federated learning framework for large language models
Tao Fan, Yan Kang, Guoqiang Ma, Weijing Chen, Wenbin Wei, Lixin Fan, and Qiang Yang. Fate-llm: An industrial grade federated learning framework for large language models. arXiv preprint arXiv:2310.10049, 2023. 14
2023
-
[91]
Opendiloco: An open-source framework for globally distributed low- communication training, 2025
PrimeIntellect-ai. Opendiloco: An open-source framework for globally distributed low- communication training, 2025
2025
-
[92]
Photon: Federated llm pre-training
Lorenzo Sani, Alex Iacob, Zeyu Cao, Royson Lee, Bill Marino, Yan Gao, Dongqi Cai, Zexi Li, Wanru Zhao, Xinchi Qiu, et al. Photon: Federated llm pre-training. arXiv preprint arXiv:2411.02908, 2024
2024
-
[93]
Biomedlm: A 2.7 b parameter language model trained on biomedical text
Elliot Bolton, Abhinav Venigalla, Michihiro Yasunaga, David Hall, Betty Xiong, Tony Lee, Roxana Daneshjou, Jonathan Frankle, Percy Liang, Michael Carbin, et al. Biomedlm: A 2.7 b parameter language model trained on biomedical text. arXiv preprint arXiv:2403.18421, 2024
2024
-
[94]
Advances and open problems in federated learning.Foundations and Trends® in Machine Learning, 14(1–2):1–210, 2021
Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Ar- jun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al. Advances and open problems in federated learning.Foundations and Trends® in Machine Learning, 14(1–2)...
2021
-
[95]
Fjord: Fair and accurate federated learning under heterogeneous targets with ordered dropout
Samuel Horvath, Stefanos Laskaridis, Mario Almeida, Ilias Leontiadis, Stylianos Venieris, and Nicholas Lane. Fjord: Fair and accurate federated learning under heterogeneous targets with ordered dropout. Advances in Neural Information Processing Systems, 34:12876–12889, 2021
2021
-
[96]
Fedrolex: Model-heterogeneous feder- ated learning with rolling sub-model extraction
Samiul Alam, Luyang Liu, Ming Yan, and Mi Zhang. Fedrolex: Model-heterogeneous feder- ated learning with rolling sub-model extraction. Advances in neural information processing systems, 35:29677–29690, 2022
2022
-
[97]
On the effects of data heterogeneity on the convergence rates of distributed linear system solvers
Boris Velasevic, Rohit Parasnis, Christopher G Brinton, and Navid Azizan. On the effects of data heterogeneity on the convergence rates of distributed linear system solvers. In 2023 62nd IEEE Conference on Decision and Control (CDC), pages 8394–8399. IEEE, 2023
2023
-
[98]
Azizan-Ruhi, F
N. Azizan-Ruhi, F. Lahouti, A. S. Avestimehr, and B. Hassibi. Distributed solution of large- scale linear systems via accelerated projection-based consensus. IEEE Transactions on Signal Processing, 67(14):3806–3817, July 2019
2019
-
[99]
Retrieval-augmented mixture of lora experts for uploadable machine learning
Ziyu Zhao, Leilei Gan, Guoyin Wang, Yuwei Hu, Tao Shen, Hongxia Yang, Kun Kuang, and Fei Wu. Retrieval-augmented mixture of lora experts for uploadable machine learning. arXiv preprint arXiv:2406.16989, 2024
2024
-
[100]
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021
2021 arXiv
-
[101]
Democratising artificial intelligence in healthcare: community-driven approaches for ethical solutions
Ceilidh Welsh, Susana Román García, Gillian C Barnett, and Raj Jena. Democratising artificial intelligence in healthcare: community-driven approaches for ethical solutions. Future Healthcare Journal, 11(3):100165, 2024
2024
-
[102]
Edge-cloud polarization and collaboration: A comprehensive survey for ai
Jiangchao Yao, Shengyu Zhang, Yang Yao, Feng Wang, Jianxin Ma, Jianwei Zhang, Yunfei Chu, Luo Ji, Kunyang Jia, Tao Shen, et al. Edge-cloud polarization and collaboration: A comprehensive survey for ai. IEEE Transactions on Knowledge and Data Engineering , 35(7):6866–6886, 2022
2022
-
[103]
Beyond a single ai cluster: A survey of decentralized llm training
Haotian Dong, Jingyan Jiang, Rongwei Lu, Jiajun Luo, Jiajun Song, Bowen Li, Ying Shen, and Zhi Wang. Beyond a single ai cluster: A survey of decentralized llm training. arXiv preprint arXiv:2503.11023, 2025
2025
-
[104]
Distributed training of large language models
Fanlong Zeng, Wensheng Gan, Yongheng Wang, and Philip S Yu. Distributed training of large language models. In 2023 IEEE 29th International Conference on Parallel and Distributed Systems (ICPADS), pages 840–847. IEEE, 2023
2023
-
[105]
Injecting domain-specific knowledge into large language models: a comprehensive survey
Zirui Song, Bin Yan, Yuhan Liu, Miao Fang, Mingzhe Li, Rui Yan, and Xiuying Chen. Injecting domain-specific knowledge into large language models: a comprehensive survey. arXiv preprint arXiv:2502.10708, 2025
2025
-
[106]
Medicalgpt: Training medical gpt model
Ming Xu. Medicalgpt: Training medical gpt model. https://github.com/shibing624/ MedicalGPT, 2023. 15
2023
-
[107]
Large language models in finance: A survey
Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen. Large language models in finance: A survey. In Proceedings of the fourth ACM international conference on AI in finance, pages 374–382, 2023
2023
-
[108]
Lawyer gpt: A legal large language model with enhanced domain knowledge and reasoning capabilities
Shunyu Yao, Qingqing Ke, Qiwei Wang, Kangtong Li, and Jie Hu. Lawyer gpt: A legal large language model with enhanced domain knowledge and reasoning capabilities. In Proceedings of the 2024 3rd International Symposium on Robotics, Artificial Intelligence and Information Enginee...
2024
-
[109]
Alphaevolve: A gemini-powered coding agent for designing advanced algorithms,
DeepMind. Alphaevolve: A gemini-powered coding agent for designing advanced algorithms,
-
[110]
Accessed: 2025-05-19
2025
-
[111]
The pursuit of fairness in artificial intelligence models: A survey, 2024
Yuxin Yao et al. The pursuit of fairness in artificial intelligence models: A survey, 2024
2024
-
[112]
Fair resource allocation in federated learning
Tian Li, Maziar Sanjabi, Ahmad Beirami, and Virginia Smith. Fair resource allocation in federated learning. In International Conference on Learning Representations, 2020
2020
-
[113]
Optimizing federated learning on non-IID data with reinforcement learning
Hongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris Papailiopoulos, and Yasaman Khaz- aeni. Optimizing federated learning on non-IID data with reinforcement learning. IEEE International Conference on Computer Communications, pages 1698–1707, 2020
2020
-
[114]
Agnostic federated learning
Mehryar Mohri, Gary Sivek, and Ananda Theertha Suresh. Agnostic federated learning. In International Conference on Machine Learning, pages 4615–4625, 2019
2019
-
[115]
Incentive design for efficient federated learning in mobile networks: A contract theory approach
Jiawen Kang, Zehui Xiong, Dusit Niyato, Han Ye, and Dong In Kim. Incentive design for efficient federated learning in mobile networks: A contract theory approach. In IEEE VTS Asia Pacific Wireless Communications Symposium, pages 1–5, 2019
2019
-
[116]
Khan, Shashi Raj Pandey, Nguyen H
Latif U. Khan, Shashi Raj Pandey, Nguyen H. Tran, Walid Saad, Zhu Han, Minh N. H. Nguyen, and Choong Seon Hong. Federated learning for edge networks: Resource optimization and incentive mechanism. IEEE Communications Magazine, 57(10):94–100, 2019
2019
-
[117]
Incentive mechanism design for joint resource allocation in blockchain-based federated learning
Zhilin Wang, Qin Hu, Ruinian Li, Minghui Xu, and Zehui Xiong. Incentive mechanism design for joint resource allocation in blockchain-based federated learning. IEEE Transactions on Parallel and Distributed Systems, 34(5):1536–1547, 2023
2023
-
[118]
Federated machine learning: Concept and applications
Qiang Yang, Yang Liu, Tianjian Chen, and Yongxin Tong. Federated machine learning: Concept and applications. ACM Transactions on Intelligent Systems and Technology (TIST), 10(2):1–19, 2019
2019
-
[119]
An edge computing matching framework with guaranteed quality of service
Nafiseh Sharghivand, Farnaz Derakhshan, Lena Mashayekhy, and Leyli Mohammadkhanli. An edge computing matching framework with guaranteed quality of service. IEEE Transactions on Cloud Computing, 10(3):1557–1570, 2020
2020
-
[120]
Environmental burden of united states data centers in the artificial intelli- gence era, 2024
Yuchen Yang et al. Environmental burden of united states data centers in the artificial intelli- gence era, 2024
2024
-
[121]
Carbon footprint reduction for sustainable data centers in real-time, 2024
Xiaoyu Li et al. Carbon footprint reduction for sustainable data centers in real-time, 2024
2024
-
[122]
Cooling systems in data centers: State of art and emerging technologies
Alfonso Capozzoli and Giulio Primiceri. Cooling systems in data centers: State of art and emerging technologies. Energy Procedia, 83:484–493, 2015
2015
-
[123]
Beutel, Taner Topal, Akhil Mathur, and Nicholas D
Xinchi Qiu, Titouan Parcollet, Daniel J. Beutel, Taner Topal, Akhil Mathur, and Nicholas D. Lane. Can federated learning save the planet? arXiv preprint arXiv:2010.06537, 2021
2010
-
[124]
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Ar- cas. Communication-efficient learning of deep networks from decentralized data. Artificial Intelligence and Statistics, pages 1273–1282, 2017
2017
-
[125]
Nvidia announces jetson tx2: Parker comes to nvidia’s embedded system kit
Ryan Smith. Nvidia announces jetson tx2: Parker comes to nvidia’s embedded system kit. IEEE Hot Chips, 29, 2017
2017
-
[126]
Federated optimization in heterogeneous networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. arXiv preprint arXiv:1812.06127, 2018. 16
2018
-
[127]
Quantifying the carbon emissions of machine learning
Alexandre Lacoste, Alexandra Luccioni, Victor Schmidt, and Thomas Dandres. Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700, 2019
1910 arXiv
-
[128]
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems , volume 25, pages 1097–1105, 2012
2012
-
[129]
The computa- tional limits of deep learning
Neil C Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F Manso. The computa- tional limits of deep learning. arXiv preprint arXiv:2007.05558, 10, 2020
2007
-
[130]
The perceptron: a probabilistic model for information storage and organiza- tion in the brain
Frank Rosenblatt. The perceptron: a probabilistic model for information storage and organiza- tion in the brain. Psychological review, 65(6):386, 1958
1958
-
[131]
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998
1998
-
[132]
Large-scale deep unsupervised learning using graphics processors
Rajat Raina, Anand Madhavan, and Andrew Y Ng. Large-scale deep unsupervised learning using graphics processors. In Proceedings of the 26th annual international conference on machine learning, pages 873–880, 2009
2009
-
[133]
In-datacenter performance analysis of a tensor processing unit
Norman P Jouppi, Cliff Young, Nishant Patil, and David Patterson. In-datacenter performance analysis of a tensor processing unit. InProceedings of the 44th annual international symposium on computer architecture, pages 1–12, 2017
2017
-
[134]
Benchmarking tpu, gpu, and cpu platforms for deep learning
Yu Emma Wang, Gu-Yeon Wei, and David Brooks. Benchmarking tpu, gpu, and cpu platforms for deep learning. arXiv preprint arXiv:1907.10701, 2019
1907
-
[135]
Ai and compute
Dario Amodei and Danny Hernandez. Ai and compute. OpenAI Blog, 2, 2018
2018
-
[136]
How many smartphones are in the world?, 2021
CounterPoint. How many smartphones are in the world?, 2021
2021
-
[137]
Phonelm: an efficient and capable small language model family through principled pre-training
Rongjie Yi, Xiang Li, Weikai Xie, Zhenyan Lu, Chenghua Wang, Ao Zhou, Shangguang Wang, Xiwen Zhang, and Mengwei Xu. Phonelm: an efficient and capable small language model family through principled pre-training. arXiv preprint arXiv:2411.05046, 2024
2024
-
[138]
Datacomp-lm: In search of the next generation of training sets for language models
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Guha, Sedrick Scott Keh, Kushal Arora, et al. Datacomp-lm: In search of the next generation of training sets for language models. Advances in Neural Information Processin...
2024
-
[139]
Starcoder: May the source be with you! arXiv preprint arXiv:2305.06161, 2023
Raymond Li, Daniel Choi, Jordi Chung, et al. Starcoder: May the source be with you! arXiv preprint arXiv:2305.06161, 2023
2023 arXiv
-
[140]
Meta releases llama 3.2
Meta AI. Meta releases llama 3.2. https://about.fb.com/news/2024/09/ introducing-llama-3-2-1b-3b/ , 2024
2024
-
[141]
Qwen2: Technical report
Jinze Yang, Shuai Wang, Shuohang Ma, Jianbo Zheng, et al. Qwen2: Technical report. arXiv preprint arXiv:2404.05169, 2024
2024
-
[142]
Qwen technical report
Jinze Bai, Shuai Wang, Fei Xiong, Zhenyu Hou, et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023
2023 arXiv
-
[143]
Gemma: Open models based on gemini research and technology
Google Team. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295, 2024
2024 arXiv
-
[144]
Smollm2: When smol goes big – data-centric training of a small language model, 2025
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíˇcek, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo ...
2025
-
[145]
Smollm-corpus, 2024
Loubna Ben Allal, Anton Lozhkov, Guilherme Penedo, Thomas Wolf, and Leandro von Werra. Smollm-corpus, 2024. 17
2024
-
[146]
H2o-danube3 technical report
Pascal Pfeiffer, Philipp Singer, Yauhen Babakhin, Gabor Fodor, Nischay Dhankhar, and Sri Satish Ambati. H2o-danube3 technical report. arXiv preprint arXiv:2407.09276, 2024
2024
-
[147]
Minicpm: Unveiling the potential of small language models
Edward Hu, Wangchunshu Huang, et al. Minicpm: Unveiling the potential of small language models. arXiv preprint arXiv:2402.03216, 2024
2024 arXiv
-
[148]
Dolma: An open corpus of high-quality english text for language model pre-training
AI2. Dolma: An open corpus of high-quality english text for language model pre-training. https://huggingface.co/datasets/allenai/dolma, 2023
2023
-
[149]
Chinese tiny llm: Pretraining a chinese-centric large language model
Xinrun Du, Zhouliang Yu, Songyang Gao, Ding Pan, Yuyang Cheng, Ziyang Ma, Ruibin Yuan, Xingwei Qu, Jiaheng Liu, Tianyu Zheng, et al. Chinese tiny llm: Pretraining a chinese-centric large language model. arXiv preprint arXiv:2404.04167, 2024
2024
-
[150]
Olmo: Accelerating the science of language models
Dirk Groeneveld, Iz Beltagy, Pete Davis, et al. Olmo: Accelerating the science of language models. arXiv preprint arXiv:2402.00838, 2024
2024 arXiv
-
[151]
Tinyllama: An open-source small language model
Peiyuan Zhang, Guangtao Chen, et al. Tinyllama: An open-source small language model. arXiv preprint arXiv:2401.02385, 2024
2024 arXiv
-
[152]
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024
2024 arXiv
-
[153]
Phi-2: The surprising power of small language models
Mojan Javaheripi, Jacob Lobo, et al. Phi-2: The surprising power of small language models. arXiv preprint arXiv:2312.12397, 2023
2023
-
[154]
Textbooks are all you need
Suriya Gunasekar, Yi Zhang, et al. Textbooks are all you need. arXiv preprint arXiv:2306.11644, 2023
2023 arXiv
-
[155]
Openelm: An efficient language model family with open training and inference framework.arXiv preprint arXiv:2404.14619, 2024
Sachin Mehta, Mohammad Hossein Sekhavat, Qingqing Cao, Maxwell Horton, Yanzi Jin, Chenfan Sun, Iman Mirzadeh, Mahyar Najibi, Dmitry Belenko, Peter Zatloukal, et al. Openelm: An efficient language model family with open training and inference framework.arXiv preprint arXiv:2404...
2024
-
[156]
The refinedweb dataset for falcon llm: Outperforming curated corpora with web data
Guilherme Penedo, Anis Crnisanin, Ethan Shen, et al. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data. arXiv preprint arXiv:2306.01116, 2023
2023 arXiv
-
[157]
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020
2020 arXiv
-
[158]
Mobillama: Towards accurate and lightweight fully transparent gpt
Omkar Thawakar, Ashmal Vayani, Salman Khan, Hisham Cholakal, Rao M Anwer, Michael Felsberg, Tim Baldwin, Eric P Xing, and Fahad Shahbaz Khan. Mobillama: Towards accurate and lightweight fully transparent gpt. arXiv preprint arXiv:2402.16840, 2024
2024
-
[159]
Mobilellm: Optimizing sub-billion parameter language models for on-device use cases
Zechun Liu, Changsheng Zhao, Forrest Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, et al. Mobilellm: Optimizing sub-billion parameter language models for on-device use cases. In Forty-first International Co...
2024
-
[160]
Compact language models via pruning and knowledge distillation
Saurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Jan Kautz, and Pavlo Molchanov. Compact language models via pruning and knowledge distillation. Advances in Neural Information Processing Sy...
2024
-
[161]
Orca 2: Teaching small language models how to reason
Arindam Mitra, Subhabrata Mukherjee, et al. Orca 2: Teaching small language models how to reason. arXiv preprint arXiv:2311.11045, 2023
2023
-
[162]
Orca 2: Teaching small language models how to reason
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Vicente Ordonez, and Kai-Wei Chang. Orca 2: Teaching small language models how to reason. arXiv preprint arXiv:2312.02558, 2023
2023
-
[163]
The flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, et al. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688, 2023. 18
2023 arXiv
-
[164]
Towards making the most of chatgpt for machine translation
Kehai Zhang, Zhuocheng Chen, et al. Towards making the most of chatgpt for machine translation. arXiv preprint arXiv:2309.02654, 2023
2023
-
[165]
Free dolly: Introducing the world’s first truly open instruction-tuned llm
Databricks. Free dolly: Introducing the world’s first truly open instruction-tuned llm. https://www.databricks.com/blog/2023/04/12/ dolly-first-open-commercially-viable-instruction-tuned-llm , 2023
2023
-
[166]
Lamini-lm: A diverse herd of distilled models from large-scale instructions
Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, and Alham Fikri Aji. Lamini-lm: A diverse herd of distilled models from large-scale instructions. arXiv preprint arXiv:2304.14402, 2023
2023
-
[167]
Sparsegpt: Massive language models can be accurately pruned in one-shot
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, et al. Sparsegpt: Massive language models can be accurately pruned in one-shot. arXiv preprint arXiv:2301.00774, 2023
2023
-
[168]
A simple and effective pruning approach for large language models
Zongyu Sun, Chen Chen, Zhitao Zhang, et al. A simple and effective pruning approach for large language models. arXiv preprint arXiv:2306.11695, 2023
2023 arXiv
-
[169]
Loraprune: Structured pruning meets low-rank parameter-efficient fine-tuning
Mingyang Zhang, Hao Chen, Chunhua Shen, Zhen Yang, Linlin Ou, Xinyi Yu, and Bohan Zhuang. Loraprune: Structured pruning meets low-rank parameter-efficient fine-tuning. arXiv preprint arXiv:2305.18403, 2023
2023
-
[170]
Shortgpt: Layers in large language models are more redundant than you expect
Yu Men, Xingyu Zhang, Ruiqi Sun, et al. Shortgpt: Layers in large language models are more redundant than you expect. arXiv preprint arXiv:2402.18952, 2024
2024
-
[171]
Bitnet: Scaling 1-bit transformers for large language models
Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Huaijie Wang, Lingxiao Ma, Fan Yang, Ruiping Wang, Yi Wu, and Furu Wei. Bitnet: Scaling 1-bit transformers for large language models. arXiv preprint arXiv:2310.11453, 2023
2023 arXiv
-
[172]
The era of 1-bit llms: All large language models are in 1.58 bits
Shuming Ma, Hongyu Zhao, Lingxiao Xue, et al. The era of 1-bit llms: All large language models are in 1.58 bits. arXiv preprint arXiv:2402.17764, 2024
2024 arXiv
-
[173]
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. Qlora: Efficient finetuning of quantized llms. arXiv preprint arXiv:2305.14314, 2023
2023 arXiv
-
[174]
Squeezellm: Dense-and-sparse quantiza- tion
Sehoon Kim, Coleman Hooper, Amir Gholami, et al. Squeezellm: Dense-and-sparse quantiza- tion. arXiv preprint arXiv:2306.07629, 2023
2023
-
[175]
The on-device intelligence update, 2024
Karan Goel. The on-device intelligence update, 2024
2024
-
[176]
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021
2021 arXiv
-
[177]
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022
2022 arXiv
-
[178]
Together ai: The ai acceleration cloud
Together AI. Together ai: The ai acceleration cloud. https://www.together.ai/, 2023
2023
-
[179]
Flock: Federated machine learning on the blockchain
FLock. Flock: Federated machine learning on the blockchain. https://www.flock.io/, 2023
2023
-
[180]
Federatedscope: An easy-to-use federated learning platform
alibaba. Federatedscope: An easy-to-use federated learning platform. https://github. com/alibaba/FederatedScope, 2024
2024
-
[181]
Fedml: The unified and scalable ml library for large-scale distributed training, model serving, and federated learning
FedML-AI. Fedml: The unified and scalable ml library for large-scale distributed training, model serving, and federated learning. https://github.com/FedML-AI/FedML, 2024
2024
-
[182]
Fedllm-bench: Realistic benchmarks for federated learning of large language models
Rui Ye, Rui Ge, Xinyu Zhu, Jingyi Chai, Du Yaxin, Yang Liu, Yanfeng Wang, and Siheng Chen. Fedllm-bench: Realistic benchmarks for federated learning of large language models. Advances in Neural Information Processing Systems, 37:111106–111130, 2025. 19 A Impact Statements The ...
2025
Reviewed May 23, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.